Beatrice Savoldi

dblp:267/2355 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0002-3061-8317ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 11 since 2021
YearPublicationVenuePosition
2026 Does Speech Translation Meet Users' Needs? An English to Portuguese Study Across Demographics
abstract
This paper introduces Ouvia, a research project to assess user-perceived usability and reliability of modern speech translation tools in En\rightarrowPt scenarios. The project centers on a user study in which we simulate real-life daily interactions by recruiting crowdworkers online from different sociodemographic groups. We collect their spoken requests and self-assessments about quality, satisfaction, and reliability. Here, we describe the project’s motivation and objectives, the study design, and the expected outcomes we will provide to speech translation practitioners.
Giuseppe Attanasio, Beatrice Savoldi, Daniel Chechelnitsky, Matteo Negri, Marine Carpuat, André F. T. Martins
EAMT (2)2
2026 SPES: Spectrogram Perturbation for Explainable Speech-to-Text Generation
abstract
Abstract Spurred by the demand for interpretable models, research on explainable AI for language technologies has experienced significant growth, with feature attribution methods emerging as a cornerstone of this progress. While prior work in NLP explored such methods for classification tasks and textual applications, explainability intersecting generation and speech is lagging, with existing techniques failing to account for the autoregressive nature of state-of-the-art models and to provide finegrained, phonetically meaningful explanations. We address this gap by introducing Spectrogram Perturbation for Explainable Speech-to-text Generation (SPES), a feature attribution technique applicable to sequence generation tasks with autoregressive models. SPES provides explanations for each predicted token based on both the input spectrogram and the previously generated tokens. Extensive evaluation on speech recognition and translation demonstrates that SPES generates explanations that are faithful and plausible to humans.
Dennis Fucci, Marco Gaido, Beatrice Savoldi, Matteo Negri, Mauro Cettolo, Luisa Bentivogli
Trans. Assoc. Comput. Linguistics3
2025 Mind the Inclusivity Gap: Multilingual Gender-Neutral Translation Evaluation with mGeNTE
abstract
Beatrice Savoldi, Giuseppe Attanasio, Eleonora Cupin, Eleni Gkovedarou, Janiça Hackenbuchner, Anne Lauscher, Matteo Negri, Andrea Piergentili, Manjinder Thind, Luisa Bentivogli. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Beatrice Savoldi, Giuseppe Attanasio, Eleonora Cupin, Eleni Gkovedarou, Janiça Hackenbuchner, Anne Lauscher, Matteo Negri, Andrea Piergentili, Manjinder Thind, Luisa Bentivogli
EMNLP1
2025 Translation in the Hands of Many: Centering Lay Users in Machine Translation Interactions
abstract
Converging societal and technical factors have transformed language technologies into userfacing applications used by the general public across languages.Machine Translation (MT) has become a global tool, with cross-lingual services now also supported by dialogue systems powered by multilingual Large Language Models (LLMs).Widespread accessibility has extended MT's reach to a vast base of lay users, many with little to no expertise in the languages or the technology itself.And yet, the understanding of MT consumed by such a diverse group of users-their needs, experiences, and interactions with multilingual systemsremains limited.In our position paper, we first trace the evolution of MT user profiles, focusing on non-experts and how their engagement with technology may shift with the rise of LLMs.Building on an interdisciplinary body of work, we identify three factors-usability, trust, and literacy-that are central to shaping user interactions and must be addressed to align MT with user needs.By examining these dimensions, we provide insights to guide the progress of more user-centered MT.
Beatrice Savoldi, Alan Ramponi, Matteo Negri, Luisa Bentivogli
EMNLP1
2025 SHADES: Towards a Multilingual Assessment of Stereotypes in Large Language Models
abstract
Margaret Mitchell, Giuseppe Attanasio, Ioana Baldini, Miruna Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton, Timm Dill, Jad Doughman, Ritam Dutt, Avijit Ghosh, Jessica Zosa Forde, Carolin Holtermann, Lucie-Aimée Kaffee, Tanmay Laud, Anne Lauscher, Roberto L Lopez-Davila, Maraim Masoud, Nikita Nangia, Anaelia Ovalle, Giada Pistilli, Dragomir Radev, Beatrice Savoldi, Vipul Raheja, Jeremy Qin, Esther Ploeger, Arjun Subramonian, Kaustubh Dhole, Kaiser Sun, Amirbek Djanibekov, Jonibek Mansurov, Kayo Yin, Emilio Villa Cueva, Sagnik Mukherjee, Jerry Huang, Xudong Shen, Jay Gala, Hamdan Al-Ali, Tair Djanibekov, Nurdaulet Mukhituly, Shangrui Nie, Shanya Sharma, Karolina Stanczak, Eliza Szczechla, Tiago Timponi Torrent, Deepak Tunuguntla, Marcelo Viridiano, Oskar Van Der Wal, Adina Yakefu, Aurélie Névéol, Mike Zhang, Sydney Zink, Zeerak Talat. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Margaret Mitchell, Giuseppe Attanasio, Ioana Baldini, Miruna-Adriana Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton, Timm Dill, Jad Doughman, Ritam Dutt, Avijit Ghosh, Jessica Zosa Forde, Carolin Holtermann, Lucie-Aimée Kaffee, Tanmay Laud, Anne Lauscher, Roberto L. Lopez-Davila, Maraim Masoud, Nikita Nangia, Anaelia Ovalle, Giada Pistilli, Dragomir R. Radev, Beatrice Savoldi, Vipul Raheja, Jeremy Qin, Esther Ploeger, Arjun Subramonian, Kaustubh D. Dhole, Kaiser Sun, Amirbek Djanibekov, Jonibek Mansurov, Kayo Yin, Emilio Villa Cueva, Sagnik Mukherjee, Jerry Huang, Jay Gala, Hamdan Al-Ali, Tair Djanibekov, Nurdaulet Mukhituly, Shangrui Nie, Shanya Sharma, Karolina Stanczak, Eliza Szczechla, Tiago Timponi Torrent, Deepak Tunuguntla, Marcelo Viridiano, Oskar Van Der Wal, Adina Yakefu, Aurélie Névéol, Mike Zhang, Sydney Zink, Zeerak Talat
NAACL (Long Papers)24
2024 Enhancing Gender-Inclusive Machine Translation with Neomorphemes and Large Language Models
abstract
Machine translation (MT) models are known to suffer from gender bias, especially when translating into languages with extensive gendered morphology. Accordingly, they still fall short in using gender-inclusive language, also representative of non-binary identities. In this paper, we look at gender-inclusive neomorphemes, neologistic elements that avoid binary gender markings as an approach towards fairer MT. In this direction, we explore prompting techniques with large language models (LLMs) to translate from English into Italian using neomorphemes. So far, this area has been under-explored due to its novelty and the lack of publicly available evaluation resources. We fill this gap by releasing NEO-GATE, a resource designed to evaluate gender-inclusive en→it translation with neomorphemes. With NEO-GATE, we assess four LLMs of different families and sizes and different prompt formats, identifying strengths and weaknesses of each on this novel task for MT.
Andrea Piergentili, Beatrice Savoldi, Matteo Negri, Luisa Bentivogli
EAMT (1)2
2024 Twists, Humps, and Pebbles: Multilingual Speech Recognition Models Exhibit Gender Performance Gaps
abstract
Current automatic speech recognition (ASR) models are designed to be used across many languages and tasks without substantial changes.However, this broad language coverage hides performance gaps within languages, for example, across genders.Our study systematically evaluates the performance of two widely used multilingual ASR models on three datasets, encompassing 19 languages from eight language families and two speaking conditions.Our findings reveal clear gender disparities, with the advantaged group varying across languages and models.Surprisingly, those gaps are not explained by acoustic or lexical properties.However, probing internal model states reveals a correlation with gendered performance gap.That is, the easier it is to distinguish speaker gender in a language using probes, the more the gap reduces, favoring female speakers.Our results show that gender disparities persist even in state-of-the-art models.Our findings have implications for the improvement of multilingual ASR systems, underscoring the importance of accessibility to training data and nuanced evaluation to predict and mitigate gender gaps.We release all code and artifacts at https://github.com/g8a9/multilingual -asr-gender-gap.
Giuseppe Attanasio, Beatrice Savoldi, Dennis Fucci, Dirk Hovy
EMNLP2
2024 What the Harm? Quantifying the Tangible Impact of Gender Bias in Machine Translation with a Human-centered Study
abstract
Gender bias in machine translation (MT) is recognized as an issue that can harm people and society.And yet, advancements in the field rarely involve people, the final MT users, or inform how they might be impacted by biased technologies.Current evaluations are often restricted to automatic methods, which offer an opaque estimate of what the downstream impact of gender disparities might be.We conduct an extensive human-centered study to examine if and to what extent bias in MT brings harms with tangible costs, such as quality of service gaps across women and men.To this aim, we collect behavioral data from ∼90 participants, who post-edited MT outputs to ensure correct gender translation.Across multiple datasets, languages, and types of users, our study shows that feminine post-editing demands significantly more technical and temporal effort, also corresponding to higher financial costs.Existing bias measurements, however, fail to reflect the found disparities.Our findings advocate for human-centered approaches that can inform the societal impact of bias.
Beatrice Savoldi, Sara Papi, Matteo Negri, Ana Guerberof Arenas, Luisa Bentivogli
EMNLP1
2023 Hi Guys or Hi Folks? Benchmarking Gender-Neutral Machine Translation with the GeNTE Corpus
abstract
Gender inequality is embedded in our communication practices and perpetuated in translation technologies.This becomes particularly apparent when translating into grammatical gender languages, where machine translation (MT) often defaults to masculine and stereotypical representations by making undue binary gender assumptions.Our work addresses the rising demand for inclusive language by focusing head-on on gender-neutral translation from English to Italian.We start from the essentials: proposing a dedicated benchmark and exploring automated evaluation methods.First, we introduce GeNTE, a natural, bilingual test set for gender-neutral translation, whose creation was informed by a survey on the perception and use of neutral language.Based on GeNTE, we then overview existing reference-based evaluation approaches, highlight their limits, and propose a reference-free method more suitable to assess gender-neutral translation.
Andrea Piergentili, Beatrice Savoldi, Dennis Fucci, Matteo Negri, Luisa Bentivogli
EMNLP2
2022 Under the Morphosyntactic Lens: A Multifaceted Evaluation of Gender Bias in Speech Translation
abstract
Gender bias is largely recognized as a problematic phenomenon affecting language technologies, with recent studies underscoring that it might surface differently across languages.However, most of current evaluation practices adopt a word-level focus on a narrow set of occupational nouns under synthetic conditions.Such protocols overlook key features of grammatical gender languages, which are characterized by morphosyntactic chains of gender agreement, marked on a variety of lexical items and parts-of-speech (POS).To overcome this limitation, we enrich the natural, gender-sensitive MuST-SHE corpus (Bentivogli et al., 2020) with two new linguistic annotation layers (POS and agreement chains), and explore to what extent different lexical categories and agreement phenomena are impacted by gender skews.Focusing on speech translation, we conduct a multifaceted evaluation on three language directions (English-French/Italian/Spanish), with models trained on varying amounts of data and different word segmentation techniques.By shedding light on model behaviours, gender bias, and its detection at several levels of granularity, our findings emphasize the value of dedicated analyses beyond aggregated overall results.
Beatrice Savoldi, Marco Gaido, Luisa Bentivogli, Matteo Negri, Marco Turchi
ACL (1)1
2021 Gender Bias in Machine Translation
abstract
Abstract Machine translation (MT) technology has facilitated our daily tasks by providing accessible shortcuts for gathering, processing, and communicating information. However, it can suffer from biases that harm users and society at large. As a relatively new field of inquiry, studies of gender bias in MT still lack cohesion. This advocates for a unified framework to ease future research. To this end, we: i) critically review current conceptualizations of bias in light of theoretical insights from related disciplines, ii) summarize previous analyses aimed at assessing gender bias in MT, iii) discuss the mitigating strategies proposed so far, and iv) point toward potential directions for future work.
Beatrice Savoldi, Marco Gaido, Luisa Bentivogli, Matteo Negri, Marco Turchi
Trans. Assoc. Comput. Linguistics1
2020 Gender in Danger? Evaluating Speech Translation Technology on the MuST-SHE Corpus
abstract
Translating from languages without productive grammatical gender like English into gender-marked languages is a well-known difficulty for machines.This difficulty is also due to the fact that the training data on which models are built typically reflect the asymmetries of natural languages, gender bias included.Exclusively fed with textual data, machine translation is intrinsically constrained by the fact that the input sentence does not always contain clues about the gender identity of the referred human entities.But what happens with speech translation, where the input is an audio signal?Can audio provide additional information to reduce gender bias?We present the first thorough investigation of gender bias in speech translation, contributing with: i) the release of a benchmark useful for future studies, and ii) the comparison of different technologies (cascade and end-to-end) on two language directions (English-Italian/French).* * These authors contributed equally.The work by Beatrice Savoldi was carried out during an internship at Fondazione Bruno Kessler. 1 We acknowledge that gender is a multifaceted notion, not necessarily constrained within binary assumptions.However, since speech translation is hindered by the scarcity of available data, we rely on the female/male distinction of gender, as it is linguistically reflected in existing natural data.
Luisa Bentivogli, Beatrice Savoldi, Matteo Negri, Mattia Antonino Di Gangi, Roldano Cattoni, Marco Turchi
ACL2
2020 Breeding Gender-aware Direct Speech Translation Systems
abstract
In automatic speech translation (ST), traditional cascade approaches involving separate transcription and translation steps are giving ground to increasingly competitive and more robust direct solutions.In particular, by translating speech audio data without intermediate transcription, direct ST models are able to leverage and preserve essential information present in the input (e.g.speaker's vocal characteristics) that is otherwise lost in the cascade framework.Although such ability proved to be useful for gender translation, direct ST is nonetheless affected by gender bias just like its cascade counterpart, as well as machine translation and numerous other natural language processing applications.Moreover, direct ST systems that exclusively rely on vocal biometric features as a gender cue can be unsuitable and potentially harmful for certain users.Going beyond speech signals, in this paper we compare different approaches to inform direct ST models about the speaker's gender and test their ability to handle gender translation from English into Italian and French.To this aim, we manually annotated large datasets with speakers' gender information and used them for experiments reflecting different possible real-world scenarios.Our results show that gender-aware direct ST solutions can significantly outperform strong -but gender-unaware -direct ST models.In particular, the translation of gender-marked words can increase up to 30 points in accuracy while preserving overall translation quality.
Marco Gaido, Beatrice Savoldi, Luisa Bentivogli, Matteo Negri, Marco Turchi
COLING2