EDBT 2026 Demo / reviewers in the wild / expert
Sara Papi
dblp:277/3949
· DBLP profile ↗
22ranked-venue papers
11as first author
21since 2021 · last 2026
0000-0002-4494-8886ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 10 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Phonetic-based Ranking for Improved Pseudo-Labeling in Low-Resource ASR
Marco Matassoni, Roberto Gretter, Falavigna Daniele, Mohamed Nabih Ali, Alessio Brutti, Matteo Negri, Mauro Cettolo, Marco Gaido, Sara Papi, Luisa Bentivogli |
LREC | 9 |
| 2025 | Granary: Speech Recognition and Translation Dataset in 25 European Languages
Nithin Rao Koluguri, Monica Sekoyan, George Zelenfroynd, Sasha Meister, Shuoyang Ding, Sofia Kostandian, He Huang 0012, Nikolay Karpov, Jagadeesh Balam, Vitaly Lavrukhin, Yifan Peng 0003, Sara Papi, Marco Gaido, Alessio Brutti, Boris Ginsburg |
INTERSPEECH | 12 |
| 2025 | How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does NotabstractThe remarkable performance achieved by Large Language Models (LLM) has driven research efforts to leverage them for a wide range of tasks and input modalities. In speech-to-text (S2T) tasks, the emerging solution consists of projecting the output of the encoder of a Speech Foundational Model (SFM) into the LLM embedding space through an adapter module. However, no work has yet investigated how much the downstream-task performance depends on each component (SFM, adapter, LLM) nor whether the best design of the adapter depends on the chosen SFM and LLM. To fill this gap, we evaluate the combination of 5 adapter modules, 2 LLMs (Mistral and Llama), and 2 SFMs (Whisper and SeamlessM4T) on two widespread S2T tasks, namely Automatic Speech Recognition and Speech Translation. Our results demonstrate that the SFM plays a pivotal role in downstream performance, while the adapter choice has moderate impact and depends on the SFM and LLM. Francesco Verdini, Pierfrancesco Melucci, Stefano Perna, Francesco Cariaggi, Marco Gaido, Sara Papi, Szymon Mazurek, Marek Kasztelnik, Luisa Bentivogli, Sébastien Bratières, Paolo Merialdo, Simone Scardapane |
INTERSPEECH | 6 |
| 2025 | Direct Speech Translation in Constrained Contexts: the Simultaneous and Subtitling Scenarios
Sara Papi |
MTSummit (1) | 1 |
| 2025 | Prepending or Cross-Attention for Speech-to-Text? An Empirical ComparisonabstractTsz Kin Lam, Marco Gaido, Sara Papi, Luisa Bentivogli, Barry Haddow. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Tsz Kin Lam, Marco Gaido, Sara Papi, Luisa Bentivogli, Barry Haddow |
NAACL (Long Papers) | 3 |
| 2025 | How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System?abstractAbstract Simultaneous speech-to-text translation (SimulST) translates source-language speech into target-language text concurrently with the speaker’s speech, ensuring low latency for better user comprehension. Despite its intended application to unbounded speech, most research has focused on human pre-segmented speech, simplifying the task and overlooking significant challenges. This narrow focus, coupled with widespread terminological inconsistencies, is limiting the applicability of research outcomes to real-world applications, ultimately hindering progress in the field. Our extensive literature review of 110 papers not only reveals these critical issues in current research but also serves as the foundation for our key contributions. We: 1) define the steps and core components of a SimulST system, proposing a standardized terminology and taxonomy; 2) conduct a thorough analysis of community trends; and 3) offer concrete recommendations and future directions to bridge the gaps in existing literature, from evaluation frameworks to system architectures, for advancing the field towards more realistic and effective SimulST solutions. Sara Papi, Peter Polak, Dominik Machácek, Ondrej Bojar |
Trans. Assoc. Comput. Linguistics | 1 |
| 2024 | Speech Translation with Speech Foundation Models and Large Language Models: What is There and What is Missing?abstractThe field of natural language processing (NLP) has recently witnessed a transformative shift with the emergence of foundation models, particularly Large Language Models (LLMs) that have revolutionized text-based NLP. This paradigm has extended to other modalities, including speech, where researchers are actively exploring the combination of Speech Foundation Models (SFMs) and LLMs into single, unified models capable of addressing multimodal tasks. Among such tasks, this paper focuses on speech-to-text translation (ST). By examining the published papers on the topic, we propose a unified view of the architectural solutions and training strategies presented so far, highlighting similarities and differences among them. Based on this examination, we not only organize the lessons learned but also show how diverse settings and evaluation approaches hinder the identification of the best-performing solution for each architectural building block and training choice. Lastly, we outline recommendations for future works on the topic aimed at better understanding the strengths and weaknesses of the SFM+LLM solutions for ST. Marco Gaido, Sara Papi, Matteo Negri, Luisa Bentivogli |
ACL (1) | 2 |
| 2024 | SBAAM! Eliminating Transcript Dependency in Automatic SubtitlingabstractSubtitling plays a crucial role in enhancing the accessibility of audiovisual content and encompasses three primary subtasks: translating spoken dialogue, segmenting translations into concise textual units, and estimating timestamps that govern their on-screen duration.Past attempts to automate this process rely, to varying degrees, on automatic transcripts, employed diversely for the three subtasks.In response to the acknowledged limitations associated with this reliance on transcripts, recent research has shifted towards transcription-free solutions for translation and segmentation, leaving the direct generation of timestamps as uncharted territory.To fill this gap, we introduce the first direct model capable of producing automatic subtitles, entirely eliminating any dependence on intermediate transcripts also for timestamp prediction.Experimental results, backed by manual evaluation, showcase our solution's new state-of-the-art performance across multiple language pairs and diverse conditions. Marco Gaido, Sara Papi, Matteo Negri, Mauro Cettolo, Luisa Bentivogli |
ACL (1) | 2 |
| 2024 | StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History SelectionabstractStreaming speech-to-text translation (StreamST) is the task of automatically translating speech while incrementally receiving an audio stream.Unlike simultaneous ST (SimulST), which deals with pre-segmented speech, StreamST faces the challenges of handling continuous and unbounded audio streams.This requires additional decisions about what to retain of the previous history, which is impractical to keep entirely due to latency and computational constraints.Despite the real-world demand for real-time ST, research on streaming translation remains limited, with existing works solely focusing on SimulST.To fill this gap, we introduce StreamAtt, the first StreamST policy, and propose StreamLAAL, the first StreamST latency metric designed to be comparable with existing metrics for SimulST.Extensive experiments across all 8 languages of MuST-C v1.0 show the effectiveness of StreamAtt compared to a naive streaming baseline and the related state-of-the-art SimulST policy, providing a first step in StreamST research. Sara Papi, Marco Gaido, Matteo Negri, Luisa Bentivogli |
ACL (1) | 1 |
| 2024 | When Good and Reproducible Results are a Giant with Feet of Clay: The Importance of Software Quality in NLPabstractDespite its crucial role in research experiments, code correctness is often presumed solely based on the perceived quality of results.This assumption, however, comes with the risk of erroneous outcomes and, in turn, potentially misleading findings.To mitigate this risk, we posit that the current focus on reproducibility should go hand in hand with the emphasis on software quality.We support our arguments with a case study in which we identify and fix three bugs in widely used implementations of the state-ofthe-art Conformer architecture.Through experiments on speech recognition and translation in various languages, we demonstrate that the presence of bugs does not prevent the achievement of good and reproducible results, which however can lead to incorrect conclusions that potentially misguide future research.As countermeasures, we release pangoliNN, a library dedicated to testing neural models, and propose a Code-quality Checklist, with the goal of promoting coding best practices and improving software quality within the NLP community. Sara Papi, Marco Gaido, Andrea Pilzer, Matteo Negri |
ACL (1) | 1 |
| 2024 | How Do Hyenas Deal with Human Speech? Speech Recognition and Translation with ConfHyenaabstractThe attention mechanism, a cornerstone of state-of-the-art neural models, faces computational hurdles in processing long sequences due to its quadratic complexity. Consequently, research efforts in the last few years focused on finding more efficient alternatives. Among them, Hyena (Poli et al., 2023) stands out for achieving competitive results in both language modeling and image classification, while offering sub-quadratic memory and computational complexity. Building on these promising results, we propose ConfHyena, a Conformer whose encoder self-attentions are replaced with an adaptation of Hyena for speech processing, where the long input sequences cause high computational costs. Through experiments in automatic speech recognition (for English) and translation (from English into 8 target languages), we show that our best ConfHyena model significantly reduces the training time by 27%, at the cost of minimal quality degradation (∼1%), which, in most cases, is not statistically significant. Marco Gaido, Sara Papi, Matteo Negri, Luisa Bentivogli |
LREC/COLING | 2 |
| 2024 | MOSEL: 950, 000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU LanguagesabstractMarco Gaido, Sara Papi, Luisa Bentivogli, Alessio Brutti, Mauro Cettolo, Roberto Gretter, Marco Matassoni, Mohamed Nabih, Matteo Negri. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Marco Gaido, Sara Papi, Luisa Bentivogli, Alessio Brutti, Mauro Cettolo, Roberto Gretter, Marco Matassoni, Mohamed Nabih Ali, Matteo Negri |
EMNLP | 2 |
| 2024 | What the Harm? Quantifying the Tangible Impact of Gender Bias in Machine Translation with a Human-centered StudyabstractGender bias in machine translation (MT) is recognized as an issue that can harm people and society.And yet, advancements in the field rarely involve people, the final MT users, or inform how they might be impacted by biased technologies.Current evaluations are often restricted to automatic methods, which offer an opaque estimate of what the downstream impact of gender disparities might be.We conduct an extensive human-centered study to examine if and to what extent bias in MT brings harms with tangible costs, such as quality of service gaps across women and men.To this aim, we collect behavioral data from ∼90 participants, who post-edited MT outputs to ensure correct gender translation.Across multiple datasets, languages, and types of users, our study shows that feminine post-editing demands significantly more technical and temporal effort, also corresponding to higher financial costs.Existing bias measurements, however, fail to reflect the found disparities.Our findings advocate for human-centered approaches that can inform the societal impact of bias. Beatrice Savoldi, Sara Papi, Matteo Negri, Ana Guerberof Arenas, Luisa Bentivogli |
EMNLP | 2 |
| 2024 | Leveraging Timestamp Information for Serialized Joint Streaming Recognition and TranslationabstractThe growing need for instant spoken language transcription and translation is driven by increased global communication and cross-lingual interactions. This has made offering translations in multiple languages essential for user applications. Traditional approaches to automatic speech recognition (ASR) and speech translation (ST) have often relied on separate systems, leading to inefficiencies in computational resources, and increased synchronization complexity in real time. In this paper, we propose a streaming Transformer-Transducer (T-T) model able to jointly produce many-to-one and one-to-many transcription and translation using a single decoder. We introduce a novel method for joint token-level serialized output training based on timestamp information to effectively produce ASR and ST outputs in the streaming setting. Experiments on {it,es,de}↔en prove the effectiveness of our approach, enabling the generation of one-to-many joint outputs with a single decoder for the first time. Sara Papi, Jun-Kun Chen, Naoyuki Kanda, Jinyu Li 0001, Yashesh Gaur |
ICASSP | 1 |
| 2023 | Attention as a Guide for Simultaneous Speech TranslationabstractIn simultaneous speech translation (SimulST), effective policies that determine when to write partial translations are crucial to reach high output quality with low latency.Towards this objective, we propose EDATT (Encoder-Decoder Attention), an adaptive policy that exploits the attention patterns between audio source and target textual translation to guide an offlinetrained ST model during simultaneous inference.EDATT exploits the attention scores modeling the audio-translation relation to decide whether to emit a partial hypothesis or wait for more audio input.This is done under the assumption that, if attention is focused towards the most recently received speech segments, the information they provide can be insufficient to generate the hypothesis (indicating that the system has to wait for additional audio input).Results on en→{de, es} show that EDATT yields better results compared to the SimulST state of the art, with gains respectively up to 7 and 4 BLEU points for the two languages, and with a reduction in computational-aware latency up to 1.4s and 0.7s compared to existing SimulST policies applied to offline-trained models. Sara Papi, Matteo Negri, Marco Turchi |
ACL (1) | 1 |
| 2023 | Token-Level Serialized Output Training for Joint Streaming ASR and ST Leveraging Textual AlignmentsabstractIn real-world applications, users often require both translations and transcriptions of speech to enhance their comprehension, particularly in streaming scenarios where incremental generation is necessary. This paper introduces a streaming Transformer-Transducer that jointly generates automatic speech recognition (ASR) and speech translation (ST) outputs using a single decoder. To produce ASR and ST content effectively with minimal latency, we propose a joint token-level serialized output training method that interleaves source and target words by leveraging an off-the-shelf textual aligner. Experiments in monolingual (it-en) and multilingual ({de,es,it}-en) settings demonstrate that our approach achieves the best quality-latency balance. With an average ASR latency of 1s and ST latency of $1.3 \mathrm{~s}$, our model shows no degradation or even improves output quality compared to separate ASR and ST models, yielding an average improvement of 1.1 WER and 0.4 BLEU in the multilingual case. Sara Papi, Jun-Kun Chen, Jinyu Li 0001, Yashesh Gaur |
ASRU | 1 |
| 2023 | Integrating Language Models into Direct Speech Translation: An Inference-Time Solution to Control Gender InflectionabstractWhen translating words referring to the speaker, speech translation (ST) systems should not resort to default masculine generics nor rely on potentially misleading vocal traits.Rather, they should assign gender according to the speakers' preference.The existing solutions to do so, though effective, are hardly feasible in practice as they involve dedicated model re-training on gender-labeled ST data.To overcome these limitations, we propose the first inferencetime solution to control speaker-related gender inflections in ST.Our approach partially replaces the (biased) internal language model (LM) implicitly learned by the ST decoder with gender-specific external LMs.Experiments on en→es/fr/it show that our solution outperforms the base models and the best training-time mitigation strategy by up to 31.0 and 1.6 points in gender accuracy, respectively, for feminine forms.The gains are even larger (up to 32.0 and 3.4) in the challenging condition where speakers' vocal traits conflict with their gender.1 Dennis Fucci, Marco Gaido, Sara Papi, Mauro Cettolo, Matteo Negri, Luisa Bentivogli |
EMNLP | 3 |
| 2023 | Joint Speech Translation and Named Entity Recognition
Marco Gaido, Sara Papi, Matteo Negri, Marco Turchi |
INTERSPEECH | 2 |
| 2023 | AlignAtt: Using Attention-based Audio-Translation Alignments as a Guide for Simultaneous Speech Translation
Sara Papi, Marco Turchi, Matteo Negri |
INTERSPEECH | 1 |
| 2023 | Direct Speech Translation for Automatic SubtitlingabstractAbstract Automatic subtitling is the task of automatically translating the speech of audiovisual content into short pieces of timed text, i.e., subtitles and their corresponding timestamps. The generated subtitles need to conform to space and time requirements, while being synchronized with the speech and segmented in a way that facilitates comprehension. Given its considerable complexity, the task has so far been addressed through a pipeline of components that separately deal with transcribing, translating, and segmenting text into subtitles, as well as predicting timestamps. In this paper, we propose the first direct speech translation model for automatic subtitling that generates subtitles in the target language along with their timestamps with a single model. Our experiments on 7 language pairs show that our approach outperforms a cascade system in the same data condition, also being competitive with production tools on both in-domain and newly released out-domain benchmarks covering new scenarios. Sara Papi, Marco Gaido, Alina Karakanta, Mauro Cettolo, Matteo Negri, Marco Turchi |
Trans. Assoc. Comput. Linguistics | 1 |
| 2021 | Speechformer: Reducing Information Loss in Direct Speech TranslationabstractTransformer-based models have gained increasing popularity achieving state-of-the-art performance in many research fields including speech translation.However, Transformer's quadratic complexity with respect to the input sequence length prevents its adoption as is with audio signals, which are typically represented by long sequences.Current solutions resort to an initial sub-optimal compression based on a fixed sampling of raw audio features.Therefore, potentially useful linguistic information is not accessible to higher-level layers in the architecture.To solve this issue, we propose Speechformer, an architecture that, thanks to a reduced memory usage in the attention layers, avoids the initial lossy compression and aggregates information only at a higher level according to more informed linguistic criteria.Experiments on three language pairs (en→de/es/nl) show the efficacy of our solution, with gains of up to 0.8 BLEU on the standard MuST-C corpus and of up to 4.0 BLEU in a low resource scenario. Sara Papi, Marco Gaido, Matteo Negri, Marco Turchi |
EMNLP (1) | 1 |
| 2020 | Mixtures of Deep Neural Experts for Automated Speech ScoringabstractThe paper copes with the task of automatic assessment of second language proficiency from the language learners' spoken responses to test prompts. The task has significant relevance to the field of computer assisted language learning. The approach presented in the paper relies on two separate modules: (1) an automatic speech recognition system that yields text transcripts of the spoken interactions involved, and (2) a multiple classifier system based on deep learners that ranks the transcripts into proficiency classes. Different deep neural network architectures (both feed-forward and recurrent) are specialized over diverse representations of the texts in terms of: a reference grammar, the outcome of probabilistic language models, several word embeddings, and two bag-of-word models. Combination of the individual classifiers is realized either via a probabilistic pseudo-joint model, or via a neural mixture of experts. Using the data of the third Spoken CALL Shared Task challenge, the highest values to date were obtained in terms of three popular evaluation metrics. Sara Papi, Edmondo Trentin, Roberto Gretter, Marco Matassoni, Daniele Falavigna |
INTERSPEECH | 1 |