Luisa Bentivogli

dblp:50/1445 · DBLP profile ↗
← Back
56ranked-venue papers
14as first author
28since 2021 · last 2026
0000-0001-7480-2231ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 55 · 14 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021
YearPublicationVenuePosition
2026 Voice, Bias, and Coreference: An Interpretability Study of Gender in Speech Translation
abstract
Unlike text, speech conveys information about the speaker, such as gender, through acoustic cues like pitch. This gives rise to modality-specific bias concerns. For example, in speech translation (ST), when translating from languages with notional gender, such as English, into languages where gender-ambiguous terms referring to the speaker are assigned grammatical gender, the speaker’s vocal characteristics may play a role in gender assignment. This risks misgendering speakers—whether through masculine defaults or vocal-based assumptions—yet how ST models make these decisions remains poorly understood. We investigate the mechanisms ST models use to assign gender to speaker-referring terms across three language pairs (en→es/fr/it). To do so, we examine how training data patterns, internal language model (ILM) biases, and acoustic information interact. We find that models do not simply replicate term-specific gender associations from training data, but learn broader patterns of masculine prevalence. While the ILM exhibits strong masculine bias, models can override these preferences based on acoustic input. Using contrastive feature attribution on spectrograms, we reveal that the model with higher gender accuracy relies on a previously unknown mechanism: using first-person pronouns to link gendered terms back to the speaker, accessing gender information distributed across the frequency spectrum rather than concentrated in pitch.
Lina Conti, Dennis Fucci, Marco Gaido, Matteo Negri, Guillaume Wisniewski, Luisa Bentivogli
LREC6
2026 Phonetic-based Ranking for Improved Pseudo-Labeling in Low-Resource ASR
Marco Matassoni, Roberto Gretter, Falavigna Daniele, Mohamed Nabih Ali, Alessio Brutti, Matteo Negri, Mauro Cettolo, Marco Gaido, Sara Papi, Luisa Bentivogli
LREC10
2026 SPES: Spectrogram Perturbation for Explainable Speech-to-Text Generation
abstract
Abstract Spurred by the demand for interpretable models, research on explainable AI for language technologies has experienced significant growth, with feature attribution methods emerging as a cornerstone of this progress. While prior work in NLP explored such methods for classification tasks and textual applications, explainability intersecting generation and speech is lagging, with existing techniques failing to account for the autoregressive nature of state-of-the-art models and to provide finegrained, phonetically meaningful explanations. We address this gap by introducing Spectrogram Perturbation for Explainable Speech-to-text Generation (SPES), a feature attribution technique applicable to sequence generation tasks with autoregressive models. SPES provides explanations for each predicted token based on both the input spectrogram and the previously generated tokens. Extensive evaluation on speech recognition and translation demonstrates that SPES generates explanations that are faithful and plausible to humans.
Dennis Fucci, Marco Gaido, Beatrice Savoldi, Matteo Negri, Mauro Cettolo, Luisa Bentivogli
Trans. Assoc. Comput. Linguistics6
2025 An Interdisciplinary Approach to Human-Centered Machine Translation
abstract
Marine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli, Fred Blain, Lynne Bowker, Monojit Choudhury, Hal Daumé Iii, Kevin Duh, Ge Gao, Alvin C Grissom II, Marzena Karpinska, Elaine C Khoong, William D. Lewis, Andre Martins, Mary Nurminen, Douglas W. Oard, Maja Popovic, Michel Simard, François Yvon. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Marine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli, Frédéric Blain, Lynne Bowker, Monojit Choudhury, Hal Daumé III, Kevin Duh, Ge Gao 0001, Alvin Grissom II, Marzena Karpinska, Elaine C. Khoong, William D. Lewis, André F. T. Martins, Mary Nurminen, Douglas W. Oard, Maja Popovic, Michel Simard, François Yvon
EMNLP4
2025 Mind the Inclusivity Gap: Multilingual Gender-Neutral Translation Evaluation with mGeNTE
abstract
Beatrice Savoldi, Giuseppe Attanasio, Eleonora Cupin, Eleni Gkovedarou, Janiça Hackenbuchner, Anne Lauscher, Matteo Negri, Andrea Piergentili, Manjinder Thind, Luisa Bentivogli. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Beatrice Savoldi, Giuseppe Attanasio, Eleonora Cupin, Eleni Gkovedarou, Janiça Hackenbuchner, Anne Lauscher, Matteo Negri, Andrea Piergentili, Manjinder Thind, Luisa Bentivogli
EMNLP10
2025 Translation in the Hands of Many: Centering Lay Users in Machine Translation Interactions
abstract
Converging societal and technical factors have transformed language technologies into userfacing applications used by the general public across languages.Machine Translation (MT) has become a global tool, with cross-lingual services now also supported by dialogue systems powered by multilingual Large Language Models (LLMs).Widespread accessibility has extended MT's reach to a vast base of lay users, many with little to no expertise in the languages or the technology itself.And yet, the understanding of MT consumed by such a diverse group of users-their needs, experiences, and interactions with multilingual systemsremains limited.In our position paper, we first trace the evolution of MT user profiles, focusing on non-experts and how their engagement with technology may shift with the rise of LLMs.Building on an interdisciplinary body of work, we identify three factors-usability, trust, and literacy-that are central to shaping user interactions and must be addressed to align MT with user needs.By examining these dimensions, we provide insights to guide the progress of more user-centered MT.
Beatrice Savoldi, Alan Ramponi, Matteo Negri, Luisa Bentivogli
EMNLP4
2025 Echoes of Phonetics: Unveiling Relevant Acoustic Cues for ASR via Feature Attribution
Dennis Fucci, Marco Gaido, Matteo Negri, Mauro Cettolo, Luisa Bentivogli
INTERSPEECH5
2025 How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not
abstract
The remarkable performance achieved by Large Language Models (LLM) has driven research efforts to leverage them for a wide range of tasks and input modalities. In speech-to-text (S2T) tasks, the emerging solution consists of projecting the output of the encoder of a Speech Foundational Model (SFM) into the LLM embedding space through an adapter module. However, no work has yet investigated how much the downstream-task performance depends on each component (SFM, adapter, LLM) nor whether the best design of the adapter depends on the chosen SFM and LLM. To fill this gap, we evaluate the combination of 5 adapter modules, 2 LLMs (Mistral and Llama), and 2 SFMs (Whisper and SeamlessM4T) on two widespread S2T tasks, namely Automatic Speech Recognition and Speech Translation. Our results demonstrate that the SFM plays a pivotal role in downstream performance, while the adapter choice has moderate impact and depends on the SFM and LLM.
Francesco Verdini, Pierfrancesco Melucci, Stefano Perna, Francesco Cariaggi, Marco Gaido, Sara Papi, Szymon Mazurek, Marek Kasztelnik, Luisa Bentivogli, Sébastien Bratières, Paolo Merialdo, Simone Scardapane
INTERSPEECH9
2025 Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison
abstract
Tsz Kin Lam, Marco Gaido, Sara Papi, Luisa Bentivogli, Barry Haddow. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Tsz Kin Lam, Marco Gaido, Sara Papi, Luisa Bentivogli, Barry Haddow
NAACL (Long Papers)4
2024 Speech Translation with Speech Foundation Models and Large Language Models: What is There and What is Missing?
abstract
The field of natural language processing (NLP) has recently witnessed a transformative shift with the emergence of foundation models, particularly Large Language Models (LLMs) that have revolutionized text-based NLP. This paradigm has extended to other modalities, including speech, where researchers are actively exploring the combination of Speech Foundation Models (SFMs) and LLMs into single, unified models capable of addressing multimodal tasks. Among such tasks, this paper focuses on speech-to-text translation (ST). By examining the published papers on the topic, we propose a unified view of the architectural solutions and training strategies presented so far, highlighting similarities and differences among them. Based on this examination, we not only organize the lessons learned but also show how diverse settings and evaluation approaches hinder the identification of the best-performing solution for each architectural building block and training choice. Lastly, we outline recommendations for future works on the topic aimed at better understanding the strengths and weaknesses of the SFM+LLM solutions for ST.
Marco Gaido, Sara Papi, Matteo Negri, Luisa Bentivogli
ACL (1)4
2024 SBAAM! Eliminating Transcript Dependency in Automatic Subtitling
abstract
Subtitling plays a crucial role in enhancing the accessibility of audiovisual content and encompasses three primary subtasks: translating spoken dialogue, segmenting translations into concise textual units, and estimating timestamps that govern their on-screen duration.Past attempts to automate this process rely, to varying degrees, on automatic transcripts, employed diversely for the three subtasks.In response to the acknowledged limitations associated with this reliance on transcripts, recent research has shifted towards transcription-free solutions for translation and segmentation, leaving the direct generation of timestamps as uncharted territory.To fill this gap, we introduce the first direct model capable of producing automatic subtitles, entirely eliminating any dependence on intermediate transcripts also for timestamp prediction.Experimental results, backed by manual evaluation, showcase our solution's new state-of-the-art performance across multiple language pairs and diverse conditions.
Marco Gaido, Sara Papi, Matteo Negri, Mauro Cettolo, Luisa Bentivogli
ACL (1)5
2024 StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History Selection
abstract
Streaming speech-to-text translation (StreamST) is the task of automatically translating speech while incrementally receiving an audio stream.Unlike simultaneous ST (SimulST), which deals with pre-segmented speech, StreamST faces the challenges of handling continuous and unbounded audio streams.This requires additional decisions about what to retain of the previous history, which is impractical to keep entirely due to latency and computational constraints.Despite the real-world demand for real-time ST, research on streaming translation remains limited, with existing works solely focusing on SimulST.To fill this gap, we introduce StreamAtt, the first StreamST policy, and propose StreamLAAL, the first StreamST latency metric designed to be comparable with existing metrics for SimulST.Extensive experiments across all 8 languages of MuST-C v1.0 show the effectiveness of StreamAtt compared to a naive streaming baseline and the related state-of-the-art SimulST policy, providing a first step in StreamST research.
Sara Papi, Marco Gaido, Matteo Negri, Luisa Bentivogli
ACL (1)4
2024 How Do Hyenas Deal with Human Speech? Speech Recognition and Translation with ConfHyena
abstract
The attention mechanism, a cornerstone of state-of-the-art neural models, faces computational hurdles in processing long sequences due to its quadratic complexity. Consequently, research efforts in the last few years focused on finding more efficient alternatives. Among them, Hyena (Poli et al., 2023) stands out for achieving competitive results in both language modeling and image classification, while offering sub-quadratic memory and computational complexity. Building on these promising results, we propose ConfHyena, a Conformer whose encoder self-attentions are replaced with an adaptation of Hyena for speech processing, where the long input sequences cause high computational costs. Through experiments in automatic speech recognition (for English) and translation (from English into 8 target languages), we show that our best ConfHyena model significantly reduces the training time by 27%, at the cost of minimal quality degradation (∼1%), which, in most cases, is not statistically significant.
Marco Gaido, Sara Papi, Matteo Negri, Luisa Bentivogli
LREC/COLING4
2024 Evaluating Automatic Subtitling: Correlating Post-editing Effort and Automatic Metrics
abstract
Systems that automatically generate subtitles from video are gradually entering subtitling workflows, both for supporting subtitlers and for accessibility purposes. Even though robust metrics are essential for evaluating the quality of automatically-generated subtitles and for estimating potential productivity gains, there is limited research on whether existing metrics, some of which directly borrowed from machine translation (MT) evaluation, can fulfil such purposes. This paper investigates how well such MT metrics correlate with measures of post-editing (PE) effort in automatic subtitling. To this aim, we collect and publicly release a new corpus containing product-, process- and participant-based data from post-editing automatic subtitles in two language pairs (en→de,it). We find that different types of metrics correlate with different aspects of PE effort. Specifically, edit distance metrics have high correlation with technical and temporal effort, while neural metrics correlate well with PE speed.
Alina Karakanta, Mauro Cettolo, Matteo Negri, Luisa Bentivogli
LREC/COLING4
2024 Enhancing Gender-Inclusive Machine Translation with Neomorphemes and Large Language Models
abstract
Machine translation (MT) models are known to suffer from gender bias, especially when translating into languages with extensive gendered morphology. Accordingly, they still fall short in using gender-inclusive language, also representative of non-binary identities. In this paper, we look at gender-inclusive neomorphemes, neologistic elements that avoid binary gender markings as an approach towards fairer MT. In this direction, we explore prompting techniques with large language models (LLMs) to translate from English into Italian using neomorphemes. So far, this area has been under-explored due to its novelty and the lack of publicly available evaluation resources. We fill this gap by releasing NEO-GATE, a resource designed to evaluate gender-inclusive en→it translation with neomorphemes. With NEO-GATE, we assess four LLMs of different families and sizes and different prompt formats, identifying strengths and weaknesses of each on this novel task for MT.
Andrea Piergentili, Beatrice Savoldi, Matteo Negri, Luisa Bentivogli
EAMT (1)4
2024 MOSEL: 950, 000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages
abstract
Marco Gaido, Sara Papi, Luisa Bentivogli, Alessio Brutti, Mauro Cettolo, Roberto Gretter, Marco Matassoni, Mohamed Nabih, Matteo Negri. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Marco Gaido, Sara Papi, Luisa Bentivogli, Alessio Brutti, Mauro Cettolo, Roberto Gretter, Marco Matassoni, Mohamed Nabih Ali, Matteo Negri
EMNLP3
2024 What the Harm? Quantifying the Tangible Impact of Gender Bias in Machine Translation with a Human-centered Study
abstract
Gender bias in machine translation (MT) is recognized as an issue that can harm people and society.And yet, advancements in the field rarely involve people, the final MT users, or inform how they might be impacted by biased technologies.Current evaluations are often restricted to automatic methods, which offer an opaque estimate of what the downstream impact of gender disparities might be.We conduct an extensive human-centered study to examine if and to what extent bias in MT brings harms with tangible costs, such as quality of service gaps across women and men.To this aim, we collect behavioral data from ∼90 participants, who post-edited MT outputs to ensure correct gender translation.Across multiple datasets, languages, and types of users, our study shows that feminine post-editing demands significantly more technical and temporal effort, also corresponding to higher financial costs.Existing bias measurements, however, fail to reflect the found disparities.Our findings advocate for human-centered approaches that can inform the societal impact of bias.
Beatrice Savoldi, Sara Papi, Matteo Negri, Ana Guerberof Arenas, Luisa Bentivogli
EMNLP5
2023 No Pitch Left Behind: Addressing Gender Unbalance In Automatic Speech Recognition Through Pitch Manipulation
abstract
Automatic speech recognition (ASR) systems are known to be sensitive to the sociolinguistic variability of speech data, in which gender plays a crucial role. This can result in disparities in recognition accuracy between male and female speakers, primarily due to the under-representation of the latter group in the training data. While in the context of hybrid ASR models several solutions have been proposed, the gender bias issue has not been explicitly addressed in end-to-end neural architectures. To fill this gap, we propose a data augmentation technique that manipulates the fundamental frequency $(f0)$ and formants. This technique reduces the data unbalance among genders by simulating voices of the under-represented female speakers and increases the variability within each gender group. Experiments on spontaneous English speech show that our technique yields a relative WER improvement up to 9.87% for utterances by female speakers, with larger gains for the least-represented $f0$ ranges.
Dennis Fucci, Marco Gaido, Matteo Negri, Mauro Cettolo, Luisa Bentivogli
ASRU5
2023 Integrating Language Models into Direct Speech Translation: An Inference-Time Solution to Control Gender Inflection
abstract
When translating words referring to the speaker, speech translation (ST) systems should not resort to default masculine generics nor rely on potentially misleading vocal traits.Rather, they should assign gender according to the speakers' preference.The existing solutions to do so, though effective, are hardly feasible in practice as they involve dedicated model re-training on gender-labeled ST data.To overcome these limitations, we propose the first inferencetime solution to control speaker-related gender inflections in ST.Our approach partially replaces the (biased) internal language model (LM) implicitly learned by the ST decoder with gender-specific external LMs.Experiments on en→es/fr/it show that our solution outperforms the base models and the best training-time mitigation strategy by up to 31.0 and 1.6 points in gender accuracy, respectively, for feminine forms.The gains are even larger (up to 32.0 and 3.4) in the challenging condition where speakers' vocal traits conflict with their gender.1
Dennis Fucci, Marco Gaido, Sara Papi, Mauro Cettolo, Matteo Negri, Luisa Bentivogli
EMNLP6
2023 Hi Guys or Hi Folks? Benchmarking Gender-Neutral Machine Translation with the GeNTE Corpus
abstract
Gender inequality is embedded in our communication practices and perpetuated in translation technologies.This becomes particularly apparent when translating into grammatical gender languages, where machine translation (MT) often defaults to masculine and stereotypical representations by making undue binary gender assumptions.Our work addresses the rising demand for inclusive language by focusing head-on on gender-neutral translation from English to Italian.We start from the essentials: proposing a dedicated benchmark and exploring automated evaluation methods.First, we introduce GeNTE, a natural, bilingual test set for gender-neutral translation, whose creation was informed by a survey on the perception and use of neutral language.Based on GeNTE, we then overview existing reference-based evaluation approaches, highlight their limits, and propose a reference-free method more suitable to assess gender-neutral translation.
Andrea Piergentili, Beatrice Savoldi, Dennis Fucci, Matteo Negri, Luisa Bentivogli
EMNLP5
2022 Under the Morphosyntactic Lens: A Multifaceted Evaluation of Gender Bias in Speech Translation
abstract
Gender bias is largely recognized as a problematic phenomenon affecting language technologies, with recent studies underscoring that it might surface differently across languages.However, most of current evaluation practices adopt a word-level focus on a narrow set of occupational nouns under synthetic conditions.Such protocols overlook key features of grammatical gender languages, which are characterized by morphosyntactic chains of gender agreement, marked on a variety of lexical items and parts-of-speech (POS).To overcome this limitation, we enrich the natural, gender-sensitive MuST-SHE corpus (Bentivogli et al., 2020) with two new linguistic annotation layers (POS and agreement chains), and explore to what extent different lexical categories and agreement phenomena are impacted by gender skews.Focusing on speech translation, we conduct a multifaceted evaluation on three language directions (English-French/Italian/Spanish), with models trained on varying amounts of data and different word segmentation techniques.By shedding light on model behaviours, gender bias, and its detection at several levels of granularity, our findings emphasize the value of dedicated analyses beyond aggregated overall results.
Beatrice Savoldi, Marco Gaido, Luisa Bentivogli, Matteo Negri, Marco Turchi
ACL (1)3
2022 Extending the MuST-C Corpus for a Comparative Evaluation of Speech Translation Technology
abstract
This project aimed at extending the test sets of the MuST-C speech translation (ST) corpus with new reference translations. The new references were collected from professional post-editors working on the output of different ST systems for three language pairs: English-German/Italian/Spanish. In this paper, we shortly describe how the data were collected and how they are distributed. As an evidence of their usefulness, we also summarise the findings of the first comparative evaluation of cascade and direct ST approaches, which was carried out relying on the collected data. The project was partially funded by the European Association for Machine Translation (EAMT) through its 2020 Sponsorship of Activities programme.
Luisa Bentivogli, Mauro Cettolo, Marco Gaido, Alina Karakanta, Matteo Negri, Marco Turchi
EAMT1
2022 Post-editing in Automatic Subtitling: A Subtitlers' perspective
abstract
Recent developments in machine translation and speech translation are opening up opportunities for computer-assisted translation tools with extended automation functions. Subtitling tools are recently being adapted for post-editing by providing automatically generated subtitles, and featuring not only machine translation, but also automatic segmentation and synchronisation. But what do professional subtitlers think of post-editing automatically generated subtitles? In this work, we conduct a survey to collect subtitlers’ impressions and feedback on the use of automatic subtitling in their workflows. Our findings show that, despite current limitations stemming mainly from speech processing errors, automatic subtitling is seen rather positively and has potential for the future.
Alina Karakanta, Luisa Bentivogli, Mauro Cettolo, Matteo Negri, Marco Turchi
EAMT2
2022 Towards a methodology for evaluating automatic subtitling
abstract
In response to the increasing interest towards automatic subtitling, this EAMT-funded project aimed at collecting subtitle post-editing data in a real use case scenario where professional subtitlers edit automatically generated subtitles. The post-editing setting includes, for the first time, automatic generation of timestamps and segmentation, and focuses on the effect of timing and segmentation edits on the post-editing process. The collected data will serve as the basis for investigating how subtitlers interact with automatic subtitling and for devising evaluation methods geared to the multimodal nature and formal requirements of subtitling.
Alina Karakanta, Luisa Bentivogli, Mauro Cettolo, Matteo Negri, Marco Turchi
EAMT2
2021 Cascade versus Direct Speech Translation: Do the Differences Still Make a Difference?
abstract
Luisa Bentivogli, Mauro Cettolo, Marco Gaido, Alina Karakanta, Alberto Martinelli, Matteo Negri, Marco Turchi. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Luisa Bentivogli, Mauro Cettolo, Marco Gaido, Alina Karakanta, Alberto Martinelli, Matteo Negri, Marco Turchi
ACL/IJCNLP (1)1
2021 Is "moby dick" a Whale or a Bird? Named Entities and Terminology in Speech Translation
abstract
Automatic translation systems are known to struggle with rare words.Among these, named entities (NEs) and domain-specific terms are crucial, since errors in their translation can lead to severe meaning distortions.Despite their importance, previous speech translation (ST) studies have neglected them, also due to the dearth of publicly available resources tailored to their specific evaluation.To fill this gap, we i) present the first systematic analysis of the behavior of state-of-the-art ST systems in translating NEs and terminology, and ii) release NEuRoparl-ST, a novel benchmark built from European Parliament speeches annotated with NEs and terminology.Our experiments on the three language directions covered by our benchmark (en→es/fr/it) show that ST systems correctly translate 75-80% of terms and 65-70% of NEs, with very low performance (37-40%) on person names.
Marco Gaido, Susana Rodríguez, Matteo Negri, Luisa Bentivogli, Marco Turchi
EMNLP (1)4
2021 MuST-C: A multilingual corpus for end-to-end speech translation
Roldano Cattoni, Mattia Antonino Di Gangi, Luisa Bentivogli, Matteo Negri, Marco Turchi
Comput. Speech Lang.3
2021 Gender Bias in Machine Translation
abstract
Abstract Machine translation (MT) technology has facilitated our daily tasks by providing accessible shortcuts for gathering, processing, and communicating information. However, it can suffer from biases that harm users and society at large. As a relatively new field of inquiry, studies of gender bias in MT still lack cohesion. This advocates for a unified framework to ease future research. To this end, we: i) critically review current conceptualizations of bias in light of theoretical insights from related disciplines, ii) summarize previous analyses aimed at assessing gender bias in MT, iii) discuss the mitigating strategies proposed so far, and iv) point toward potential directions for future work.
Beatrice Savoldi, Marco Gaido, Luisa Bentivogli, Matteo Negri, Marco Turchi
Trans. Assoc. Comput. Linguistics3
2020 Gender in Danger? Evaluating Speech Translation Technology on the MuST-SHE Corpus
abstract
Translating from languages without productive grammatical gender like English into gender-marked languages is a well-known difficulty for machines.This difficulty is also due to the fact that the training data on which models are built typically reflect the asymmetries of natural languages, gender bias included.Exclusively fed with textual data, machine translation is intrinsically constrained by the fact that the input sentence does not always contain clues about the gender identity of the referred human entities.But what happens with speech translation, where the input is an audio signal?Can audio provide additional information to reduce gender bias?We present the first thorough investigation of gender bias in speech translation, contributing with: i) the release of a benchmark useful for future studies, and ii) the comparison of different technologies (cascade and end-to-end) on two language directions (English-Italian/French).* * These authors contributed equally.The work by Beatrice Savoldi was carried out during an internship at Fondazione Bruno Kessler. 1 We acknowledge that gender is a multifaceted notion, not necessarily constrained within binary assumptions.However, since speech translation is hindered by the scarcity of available data, we rely on the female/male distinction of gender, as it is linguistically reflected in existing natural data.
Luisa Bentivogli, Beatrice Savoldi, Matteo Negri, Mattia Antonino Di Gangi, Roldano Cattoni, Marco Turchi
ACL1
2020 Breeding Gender-aware Direct Speech Translation Systems
abstract
In automatic speech translation (ST), traditional cascade approaches involving separate transcription and translation steps are giving ground to increasingly competitive and more robust direct solutions.In particular, by translating speech audio data without intermediate transcription, direct ST models are able to leverage and preserve essential information present in the input (e.g.speaker's vocal characteristics) that is otherwise lost in the cascade framework.Although such ability proved to be useful for gender translation, direct ST is nonetheless affected by gender bias just like its cascade counterpart, as well as machine translation and numerous other natural language processing applications.Moreover, direct ST systems that exclusively rely on vocal biometric features as a gender cue can be unsuitable and potentially harmful for certain users.Going beyond speech signals, in this paper we compare different approaches to inform direct ST models about the speaker's gender and test their ability to handle gender translation from English into Italian and French.To this aim, we manually annotated large datasets with speakers' gender information and used them for experiments reflecting different possible real-world scenarios.Our results show that gender-aware direct ST solutions can significantly outperform strong -but gender-unaware -direct ST models.In particular, the translation of gender-marked words can increase up to 30 points in accuracy while preserving overall translation quality.
Marco Gaido, Beatrice Savoldi, Luisa Bentivogli, Matteo Negri, Marco Turchi
COLING3
2020 CEF Data Marketplace: Powering a Long-term Supply of Language Data
abstract
We describe the CEF Data Marketplace project, which focuses on the development of a trading platform of translation data for language professionals: translators, machine translation (MT) developers, language service providers (LSPs), translation buyers and government bodies. The CEF Data Marketplace platform will be designed and built to manage and trade data for all languages and domains. This project will open a continuous and longterm supply of language data for MT and other machine learning applications.
Amir Kamran, Dace Dzeguze, Jaap van der Meer, Milica Panic, Alessandro Cattelan, Daniele Patrioli, Luisa Bentivogli, Marco Turchi
EAMT7
2019 Machine Translation for Machines: the Sentiment Classification Use Case
abstract
Amirhossein Tebbifakhr, Luisa Bentivogli, Matteo Negri, Marco Turchi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Amirhossein Tebbifakhr, Luisa Bentivogli, Matteo Negri, Marco Turchi
EMNLP/IJCNLP (1)2
2019 MAGMATic: A Multi-domain Academic Gold Standard with Manual Annotation of Terminology for Machine Translation Evaluation
Randy Scansani, Luisa Bentivogli, Silvia Bernardini, Adriano Ferraresi
MTSummit (1)2
2019 Do translator trainees trust machine translation? An experiment on post-editing and revision
Randy Scansani, Silvia Bernardini, Adriano Ferraresi, Luisa Bentivogli
MTSummit (2)4
2018 Neural versus phrase-based MT quality: An in-depth analysis on English-German and English-French
Luisa Bentivogli, Arianna Bisazza, Mauro Cettolo, Marcello Federico
Comput. Speech Lang.1
2016 Neural versus Phrase-Based Machine Translation Quality: a Case Study
abstract
Within the field of Statistical Machine Translation (SMT), the neural approach (NMT) has recently emerged as the first technology able to challenge the long-standing dominance of phrase-based approaches (PBMT).In particular, at the IWSLT 2015 evaluation campaign, NMT outperformed well established state-ofthe-art PBMT systems on English-German, a language pair known to be particularly hard because of morphology and syntactic differences.To understand in what respects NMT provides better translation quality than PBMT, we perform a detailed analysis of neural vs. phrase-based SMT outputs, leveraging high quality post-edits performed by professional translators on the IWSLT data.For the first time, our analysis provides useful insights on what linguistic phenomena are best modeled by neural models -such as the reordering of verbs -while pointing out other aspects that remain to be improved.
Luisa Bentivogli, Arianna Bisazza, Mauro Cettolo, Marcello Federico
EMNLP1
2016 WAGS: A Beautiful English-Italian Benchmark Supporting Word Alignment Evaluation on Rare Words
Luisa Bentivogli, Mauro Cettolo, M. Amin Farajian, Marcello Federico
LREC1
2016 The first Automatic Translation Memory Cleaning Shared Task
Eduard Barbu, Carla Parra Escartín, Luisa Bentivogli, Matteo Negri, Marco Turchi, Constantin Orasan, Marcello Federico
Mach. Transl.3
2016 On the Evaluation of Adaptive Machine Translation for Human Post-Editing
abstract
We investigate adaptive machine translation (MT) as a way to reduce human workload and enhance user experience when professional translators operate in real-life conditions. A crucial aspect in our analysis is how to ensure a reliable assessment of MT technologies aimed to support human post-editing. We pay particular attention to two evaluation aspects: i) the design of a sound experimental protocol to reduce the risk of collecting biased measurements, and ii) the use of robust statistical testing methods (linear mixed-effects models) to reduce the risk of under/over-estimating the observed variations. Our adaptive MT technology is integrated in a web-based full-fledged computer-assisted translation (CAT) tool. We report on a post-editing field test that involved 16 professional translators working on two translation directions (English-Italian and English-French), with texts coming from two linguistic domains (legal, information technology). Our contrastive experiments compare user post-editing effort with static vs. adaptive MT in an end-to-end scenario where the system is evaluated as a whole. Our results evidence that adaptive MT leads to an overall reduction in post-editing effort (HTER) up to 10.6% (p <; 0.05). A follow-up manual evaluation of the MT outputs and their corresponding post-edits confirms that the gain in HTER corresponds to higher quality of the adaptive MT system and does not come at the expense of the final human translation quality. Indeed, adaptive MT shows to return better suggestions than static MT (p <; 0.01), and the resulting post-edits do not significantly differ in the two conditions.
Luisa Bentivogli, Nicola Bertoldi, Mauro Cettolo, Marcello Federico, Matteo Negri, Marco Turchi
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 Textual entailment graphs
abstract
Abstract In this work, we present a novel type of graphs for natural language processing (NLP), namely textual entailment graphs (TEGs). We describe the complete methodology we developed for the construction of such graphs and provide some baselines for this task by evaluating relevant state-of-the-art technology. We situate our research in the context of text exploration, since it was motivated by joint work with industrial partners in the text analytics area. Accordingly, we present our motivating scenario and the first gold-standard dataset of TEGs. However, while our own motivation and the dataset focus on the text exploration setting, we suggest that TEGs can have different usages and suggest that automatic creation of such graphs is an interesting task for the community.
Lili Kotlerman, Ido Dagan, Bernardo Magnini, Luisa Bentivogli
Nat. Lang. Eng.4
2014 Assessing the Impact of Translation Errors on Machine Translation Quality with Mixed-effects Models
abstract
Learning from errors is a crucial aspect of improving expertise.Based on this notion, we discuss a robust statistical framework for analysing the impact of different error types on machine translation (MT) output quality.Our approach is based on linear mixed-effects models, which allow the analysis of error-annotated MT output taking into account the variability inherent to the specific experimental setting from which the empirical observations are drawn.Our experiments are carried out on different language pairs involving Chinese, Arabic and Russian as target languages.Interesting findings are reported, concerning the impact of different error types both at the level of human perception of quality and with respect to performance results measured with automatic metrics.
Marcello Federico, Matteo Negri, Luisa Bentivogli, Marco Turchi
EMNLP3
2014 A SICK cure for the evaluation of compositional distributional semantic models
Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, Roberto Zamparelli
LREC4
2013 Comparing two methods for crowdsourcing speech transcription
abstract
This paper presents the results of an experimental study conducted with the aim of comparing two methods for crowdsourcing speech transcription that incorporate two different quality control mechanisms (i.e. explicit versus implicit) and that are based on two different processes (i.e. parallel versus iterative). In the Gold Standard method the same speech segment is transcribed in parallel by multiple contributors whose reliability is checked with respect to some reference transcriptions provided by experts. On the other hand, in the Dual Pathway method two independent groups of contributors work on the same set of transcriptions refining them in an iterative way until they converge, and thus eliminating the need to have reference transcriptions and to check transcription quality in a separate phase. These two methods were tested on about half an hour of broadcast news speech and for two different European languages, namely German and Italian. Both methods obtained good results in terms of Word Error Rate (WER) and compare well with the word disagreement rate of experts on the same data.
Rachele Sprugnoli, Giovanni Moretti, Matteo Fuoli, Diego Giuliani, Luisa Bentivogli, Emanuele Pianta, Roberto Gretter, Fabio Brugnara
ICASSP5
2012 Crowd-based MT Evaluation for non-English Target Languages
Michael Paul, Eiichiro Sumita, Luisa Bentivogli, Marcello Federico
EAMT3
2012 The IWSLT 2011 Evaluation Campaign on Automatic Talk Translation
Marcello Federico, Sebastian Stüker, Luisa Bentivogli, Michael Paul, Mauro Cettolo, Teresa Herrmann, Jan Niehues, Giovanni Moretti
LREC3
2012 Chinese Whispers: Cooperative Paraphrase Acquisition
Matteo Negri, Yashar Mehdad, Alessandro Marchetti, Danilo Giampiccolo, Luisa Bentivogli
LREC5
2011 Divide and Conquer: Crowdsourcing the Creation of Cross-Lingual Textual Entailment Corpora
Matteo Negri, Luisa Bentivogli, Yashar Mehdad, Danilo Giampiccolo, Alessandro Marchetti
EMNLP2
2011 Getting Expert Quality from the Crowd for Machine Translation Evaluation
Luisa Bentivogli, Marcello Federico, Giovanni Moretti, Michael Paul
MTSummit1
2010 A Resource for Investigating the Impact of Anaphora and Coreference on Inference
Azad Abad, Luisa Bentivogli, Ido Dagan, Danilo Giampiccolo, Shachar Mirkin, Emanuele Pianta, Asher Stern 0001
LREC2
2010 Building Textual Entailment Specialized Data Sets: a Methodology for Isolating Linguistic Phenomena Relevant to Inference
Luisa Bentivogli, Elena Cabrio, Ido Dagan, Danilo Giampiccolo, Medea Lo Leggio, Bernardo Magnini
LREC1
2005 Exploiting parallel texts in the creation of multilingual semantically annotated resources: the MultiSemCor Corpus
abstract
In this article we illustrate and evaluate an approach to create high quality linguistically annotated resources based on the exploitation of aligned parallel corpora. This approach is based on the assumption that if a text in one language has been annotated and its translation has not, annotations can be transferred from the source text to the target using word alignment as a bridge. The transfer approach has been tested and extensively applied for the creation of the MultiSemCor corpus, an English/Italian parallel corpus created on the basis of the English SemCor corpus. In MultiSemCor the texts are aligned at the word level and word sense annotated with a shared inventory of senses. A number of experiments have been carried out to evaluate the different steps involved in the methodology and the results suggest that the transfer approach is one promising solution to the resource bottleneck. First, it leads to the creation of a parallel corpus, which represents a crucial resource per se. Second, it allows for the exploitation of existing (mostly English) annotated resources to bootstrap the creation of annotated corpora in new (resource-poor) languages with greatly reduced human effort.
Luisa Bentivogli, Emanuele Pianta
Nat. Lang. Eng.1
2004 Evaluating Cross-Language Annotation Transfer in the MultiSemCor Corpus
Luisa Bentivogli, Pamela Forner, Emanuele Pianta
COLING1
2004 Knowledge Intensive Word Alignment with KNOWA
Emanuele Pianta, Luisa Bentivogli
COLING2
2003 Beyond Lexical Units: Enriching WordNets with Phrasets
Luisa Bentivogli, Emanuele Pianta
EACL1
2002 Opportunistic Semantic Tagging
Luisa Bentivogli, Emanuele Pianta
LREC1
2000 Coping with Lexical Gaps when Building Aligned Multilingual Wordnets
Luisa Bentivogli, Emanuele Pianta, Fabio Pianesi
LREC1