EDBT 2026 Demo / reviewers in the wild / expert
Mauro Cettolo
dblp:67/5075
· DBLP profile ↗
63ranked-venue papers
14as first author
16since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 53 · 11 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 7 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Phonetic-based Ranking for Improved Pseudo-Labeling in Low-Resource ASR
Marco Matassoni, Roberto Gretter, Falavigna Daniele, Mohamed Nabih Ali, Alessio Brutti, Matteo Negri, Mauro Cettolo, Marco Gaido, Sara Papi, Luisa Bentivogli |
LREC | 7 |
| 2026 | SPES: Spectrogram Perturbation for Explainable Speech-to-Text GenerationabstractAbstract Spurred by the demand for interpretable models, research on explainable AI for language technologies has experienced significant growth, with feature attribution methods emerging as a cornerstone of this progress. While prior work in NLP explored such methods for classification tasks and textual applications, explainability intersecting generation and speech is lagging, with existing techniques failing to account for the autoregressive nature of state-of-the-art models and to provide finegrained, phonetically meaningful explanations. We address this gap by introducing Spectrogram Perturbation for Explainable Speech-to-text Generation (SPES), a feature attribution technique applicable to sequence generation tasks with autoregressive models. SPES provides explanations for each predicted token based on both the input spectrogram and the previously generated tokens. Extensive evaluation on speech recognition and translation demonstrates that SPES generates explanations that are faithful and plausible to humans. Dennis Fucci, Marco Gaido, Beatrice Savoldi, Matteo Negri, Mauro Cettolo, Luisa Bentivogli |
Trans. Assoc. Comput. Linguistics | 5 |
| 2025 | Echoes of Phonetics: Unveiling Relevant Acoustic Cues for ASR via Feature Attribution
Dennis Fucci, Marco Gaido, Matteo Negri, Mauro Cettolo, Luisa Bentivogli |
INTERSPEECH | 4 |
| 2024 | SBAAM! Eliminating Transcript Dependency in Automatic SubtitlingabstractSubtitling plays a crucial role in enhancing the accessibility of audiovisual content and encompasses three primary subtasks: translating spoken dialogue, segmenting translations into concise textual units, and estimating timestamps that govern their on-screen duration.Past attempts to automate this process rely, to varying degrees, on automatic transcripts, employed diversely for the three subtasks.In response to the acknowledged limitations associated with this reliance on transcripts, recent research has shifted towards transcription-free solutions for translation and segmentation, leaving the direct generation of timestamps as uncharted territory.To fill this gap, we introduce the first direct model capable of producing automatic subtitles, entirely eliminating any dependence on intermediate transcripts also for timestamp prediction.Experimental results, backed by manual evaluation, showcase our solution's new state-of-the-art performance across multiple language pairs and diverse conditions. Marco Gaido, Sara Papi, Matteo Negri, Mauro Cettolo, Luisa Bentivogli |
ACL (1) | 4 |
| 2024 | Evaluating Automatic Subtitling: Correlating Post-editing Effort and Automatic MetricsabstractSystems that automatically generate subtitles from video are gradually entering subtitling workflows, both for supporting subtitlers and for accessibility purposes. Even though robust metrics are essential for evaluating the quality of automatically-generated subtitles and for estimating potential productivity gains, there is limited research on whether existing metrics, some of which directly borrowed from machine translation (MT) evaluation, can fulfil such purposes. This paper investigates how well such MT metrics correlate with measures of post-editing (PE) effort in automatic subtitling. To this aim, we collect and publicly release a new corpus containing product-, process- and participant-based data from post-editing automatic subtitles in two language pairs (en→de,it). We find that different types of metrics correlate with different aspects of PE effort. Specifically, edit distance metrics have high correlation with technical and temporal effort, while neural metrics correlate well with PE speed. Alina Karakanta, Mauro Cettolo, Matteo Negri, Luisa Bentivogli |
LREC/COLING | 2 |
| 2024 | MOSEL: 950, 000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU LanguagesabstractMarco Gaido, Sara Papi, Luisa Bentivogli, Alessio Brutti, Mauro Cettolo, Roberto Gretter, Marco Matassoni, Mohamed Nabih, Matteo Negri. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Marco Gaido, Sara Papi, Luisa Bentivogli, Alessio Brutti, Mauro Cettolo, Roberto Gretter, Marco Matassoni, Mohamed Nabih Ali, Matteo Negri |
EMNLP | 5 |
| 2023 | No Pitch Left Behind: Addressing Gender Unbalance In Automatic Speech Recognition Through Pitch ManipulationabstractAutomatic speech recognition (ASR) systems are known to be sensitive to the sociolinguistic variability of speech data, in which gender plays a crucial role. This can result in disparities in recognition accuracy between male and female speakers, primarily due to the under-representation of the latter group in the training data. While in the context of hybrid ASR models several solutions have been proposed, the gender bias issue has not been explicitly addressed in end-to-end neural architectures. To fill this gap, we propose a data augmentation technique that manipulates the fundamental frequency $(f0)$ and formants. This technique reduces the data unbalance among genders by simulating voices of the under-represented female speakers and increases the variability within each gender group. Experiments on spontaneous English speech show that our technique yields a relative WER improvement up to 9.87% for utterances by female speakers, with larger gains for the least-represented $f0$ ranges. Dennis Fucci, Marco Gaido, Matteo Negri, Mauro Cettolo, Luisa Bentivogli |
ASRU | 4 |
| 2023 | Integrating Language Models into Direct Speech Translation: An Inference-Time Solution to Control Gender InflectionabstractWhen translating words referring to the speaker, speech translation (ST) systems should not resort to default masculine generics nor rely on potentially misleading vocal traits.Rather, they should assign gender according to the speakers' preference.The existing solutions to do so, though effective, are hardly feasible in practice as they involve dedicated model re-training on gender-labeled ST data.To overcome these limitations, we propose the first inferencetime solution to control speaker-related gender inflections in ST.Our approach partially replaces the (biased) internal language model (LM) implicitly learned by the ST decoder with gender-specific external LMs.Experiments on en→es/fr/it show that our solution outperforms the base models and the best training-time mitigation strategy by up to 31.0 and 1.6 points in gender accuracy, respectively, for feminine forms.The gains are even larger (up to 32.0 and 3.4) in the challenging condition where speakers' vocal traits conflict with their gender.1 Dennis Fucci, Marco Gaido, Sara Papi, Mauro Cettolo, Matteo Negri, Luisa Bentivogli |
EMNLP | 4 |
| 2023 | Direct Speech Translation for Automatic SubtitlingabstractAbstract Automatic subtitling is the task of automatically translating the speech of audiovisual content into short pieces of timed text, i.e., subtitles and their corresponding timestamps. The generated subtitles need to conform to space and time requirements, while being synchronized with the speech and segmented in a way that facilitates comprehension. Given its considerable complexity, the task has so far been addressed through a pipeline of components that separately deal with transcribing, translating, and segmenting text into subtitles, as well as predicting timestamps. In this paper, we propose the first direct speech translation model for automatic subtitling that generates subtitles in the target language along with their timestamps with a single model. Our experiments on 7 language pairs show that our approach outperforms a cascade system in the same data condition, also being competitive with production tools on both in-domain and newly released out-domain benchmarks covering new scenarios. Sara Papi, Marco Gaido, Alina Karakanta, Mauro Cettolo, Matteo Negri, Marco Turchi |
Trans. Assoc. Comput. Linguistics | 4 |
| 2022 | Extending the MuST-C Corpus for a Comparative Evaluation of Speech Translation TechnologyabstractThis project aimed at extending the test sets of the MuST-C speech translation (ST) corpus with new reference translations. The new references were collected from professional post-editors working on the output of different ST systems for three language pairs: English-German/Italian/Spanish. In this paper, we shortly describe how the data were collected and how they are distributed. As an evidence of their usefulness, we also summarise the findings of the first comparative evaluation of cascade and direct ST approaches, which was carried out relying on the collected data. The project was partially funded by the European Association for Machine Translation (EAMT) through its 2020 Sponsorship of Activities programme. Luisa Bentivogli, Mauro Cettolo, Marco Gaido, Alina Karakanta, Matteo Negri, Marco Turchi |
EAMT | 2 |
| 2022 | Post-editing in Automatic Subtitling: A Subtitlers' perspectiveabstractRecent developments in machine translation and speech translation are opening up opportunities for computer-assisted translation tools with extended automation functions. Subtitling tools are recently being adapted for post-editing by providing automatically generated subtitles, and featuring not only machine translation, but also automatic segmentation and synchronisation. But what do professional subtitlers think of post-editing automatically generated subtitles? In this work, we conduct a survey to collect subtitlers’ impressions and feedback on the use of automatic subtitling in their workflows. Our findings show that, despite current limitations stemming mainly from speech processing errors, automatic subtitling is seen rather positively and has potential for the future. Alina Karakanta, Luisa Bentivogli, Mauro Cettolo, Matteo Negri, Marco Turchi |
EAMT | 3 |
| 2022 | Towards a methodology for evaluating automatic subtitlingabstractIn response to the increasing interest towards automatic subtitling, this EAMT-funded project aimed at collecting subtitle post-editing data in a real use case scenario where professional subtitlers edit automatically generated subtitles. The post-editing setting includes, for the first time, automatic generation of timestamps and segmentation, and focuses on the effect of timing and segmentation edits on the post-editing process. The collected data will serve as the basis for investigating how subtitlers interact with automatic subtitling and for devising evaluation methods geared to the multimodal nature and formal requirements of subtitling. Alina Karakanta, Luisa Bentivogli, Mauro Cettolo, Matteo Negri, Marco Turchi |
EAMT | 3 |
| 2022 | Evaluating Subtitle Segmentation for End-to-end Generation SystemsabstractSubtitles appear on screen as short pieces of text, segmented based on formal constraints (length) and syntactic/semantic criteria. Subtitle segmentation can be evaluated with sequence segmentation metrics against a human reference. However, standard segmentation metrics cannot be applied when systems generate outputs different than the reference, e.g. with end-to-end subtitling systems. In this paper, we study ways to conduct reference-based evaluations of segmentation accuracy irrespective of the textual content. We first conduct a systematic analysis of existing metrics for evaluating subtitle segmentation. We then introduce Sigma, a Subtitle Segmentation Score derived from an approximate upper-bound of BLEU on segmentation boundaries, which allows us to disentangle the effect of good segmentation from text quality. To compare Sigma with existing metrics, we further propose a boundary projection method from imperfect hypotheses to the true reference. Results show that all metrics are able to reward high quality output but for similar outputs system ranking depends on each metric’s sensitivity to error type. Our thorough analyses suggest Sigma is a promising segmentation candidate but its reliability over other segmentation metrics remains to be validated through correlations with human judgements. Alina Karakanta, François Buet, Mauro Cettolo, François Yvon |
LREC | 3 |
| 2021 | Cascade versus Direct Speech Translation: Do the Differences Still Make a Difference?abstractLuisa Bentivogli, Mauro Cettolo, Marco Gaido, Alina Karakanta, Alberto Martinelli, Matteo Negri, Marco Turchi. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Luisa Bentivogli, Mauro Cettolo, Marco Gaido, Alina Karakanta, Alberto Martinelli, Matteo Negri, Marco Turchi |
ACL/IJCNLP (1) | 2 |
| 2021 | CTC-based Compression for Direct Speech TranslationabstractPrevious studies demonstrated that a dynamic phone-informed compression of the input audio is beneficial for speech translation (ST).However, they required a dedicated model for phone recognition and did not test this solution for direct ST, in which a single model translates the input audio into the target language without intermediate representations.In this work, we propose the first method able to perform a dynamic compression of the input in direct ST models.In particular, we exploit the Connectionist Temporal Classification (CTC) to compress the input sequence according to its phonetic characteristics.Our experiments demonstrate that our solution brings a 1.3-1.5 BLEU improvement over a strong baseline on two language pairs (English-Italian and English-German), contextually reducing the memory footprint by more than 10%. Marco Gaido, Mauro Cettolo, Matteo Negri, Marco Turchi |
EACL | 2 |
| 2021 | Lexical Modeling of ASR Errors for Robust Speech Translation
Giuseppe Martucci, Mauro Cettolo, Matteo Negri, Marco Turchi |
Interspeech | 2 |
| 2020 | Contextualized Translation of Automatically Segmented SpeechabstractDirect speech-to-text translation (ST) models are usually trained on corpora segmented at sentence level, but at inference time they are commonly fed with audio split by a voice activity detector (VAD). Since VAD segmentation is not syntax-informed, the resulting segments do not necessarily correspond to well-formed sentences uttered by the speaker but, most likely, to fragments of one or more sentences. This segmentation mismatch degrades considerably the quality of ST models' output. So far, researchers have focused on improving audio segmentation towards producing sentence-like splits. In this paper, instead, we address the issue in the model, making it more robust to a different, potentially sub-optimal segmentation. To this aim, we train our models on randomly segmented data and compare two approaches: fine-tuning and adding the previous segment as context. We show that our context-aware solution is more robust to VAD-segmented input, outperforming a strong base model and the fine-tuning on different VAD segmentations of an English-German test set by up to 4.25 BLEU points. Marco Gaido, Mattia Antonino Di Gangi, Matteo Negri, Mauro Cettolo, Marco Turchi |
INTERSPEECH | 4 |
| 2018 | A Comparison of Transformer and Recurrent Neural Networks on Multilingual Neural Machine TranslationabstractRecently, neural machine translation (NMT) has been extended to multilinguality, that is to handle more than one translation direction with a single system. Multilingual NMT showed competitive performance against pure bilingual systems. Notably, in low-resource settings, it proved to work effectively and efficiently, thanks to shared representation space that is forced across languages and induces a sort of transfer-learning. Furthermore, multilingual NMT enables so-called zero-shot inference across language pairs never seen at training time. Despite the increasing interest in this framework, an in-depth analysis of what a multilingual NMT model is capable of and what it is not is still missing. Motivated by this, our work (i) provides a quantitative and comparative analysis of the translations produced by bilingual, multilingual and zero-shot systems; (ii) investigates the translation quality of two of the currently dominant neural architectures in MT, which are the Recurrent and the Transformer ones; and (iii) quantitatively explores how the closeness between languages influences the zero-shot translation. Our analysis leverages multiple professional post-edits of automatic translations by several different systems and focuses both on automatic standard metrics (BLEU and TER) and on widely used error categories, which are lexical, morphology, and word order errors. Surafel Melaku Lakew, Mauro Cettolo, Marcello Federico |
COLING | 2 |
| 2018 | Neural versus phrase-based MT quality: An in-depth analysis on English-German and English-French
Luisa Bentivogli, Arianna Bisazza, Mauro Cettolo, Marcello Federico |
Comput. Speech Lang. | 3 |
| 2016 | Neural versus Phrase-Based Machine Translation Quality: a Case StudyabstractWithin the field of Statistical Machine Translation (SMT), the neural approach (NMT) has recently emerged as the first technology able to challenge the long-standing dominance of phrase-based approaches (PBMT).In particular, at the IWSLT 2015 evaluation campaign, NMT outperformed well established state-ofthe-art PBMT systems on English-German, a language pair known to be particularly hard because of morphology and syntactic differences.To understand in what respects NMT provides better translation quality than PBMT, we perform a detailed analysis of neural vs. phrase-based SMT outputs, leveraging high quality post-edits performed by professional translators on the IWSLT data.For the first time, our analysis provides useful insights on what linguistic phenomena are best modeled by neural models -such as the reordering of verbs -while pointing out other aspects that remain to be improved. Luisa Bentivogli, Arianna Bisazza, Mauro Cettolo, Marcello Federico |
EMNLP | 3 |
| 2016 | WAGS: A Beautiful English-Italian Benchmark Supporting Word Alignment Evaluation on Rare Words
Luisa Bentivogli, Mauro Cettolo, M. Amin Farajian, Marcello Federico |
LREC | 2 |
| 2016 | On the Evaluation of Adaptive Machine Translation for Human Post-EditingabstractWe investigate adaptive machine translation (MT) as a way to reduce human workload and enhance user experience when professional translators operate in real-life conditions. A crucial aspect in our analysis is how to ensure a reliable assessment of MT technologies aimed to support human post-editing. We pay particular attention to two evaluation aspects: i) the design of a sound experimental protocol to reduce the risk of collecting biased measurements, and ii) the use of robust statistical testing methods (linear mixed-effects models) to reduce the risk of under/over-estimating the observed variations. Our adaptive MT technology is integrated in a web-based full-fledged computer-assisted translation (CAT) tool. We report on a post-editing field test that involved 16 professional translators working on two translation directions (English-Italian and English-French), with texts coming from two linguistic domains (legal, information technology). Our contrastive experiments compare user post-editing effort with static vs. adaptive MT in an end-to-end scenario where the system is evaluated as a whole. Our results evidence that adaptive MT leads to an overall reduction in post-editing effort (HTER) up to 10.6% (p <; 0.05). A follow-up manual evaluation of the MT outputs and their corresponding post-edits confirms that the gain in HTER corresponds to higher quality of the adaptive MT system and does not come at the expense of the final human translation quality. Indeed, adaptive MT shows to return better suggestions than static MT (p <; 0.01), and the resulting post-edits do not significantly differ in the two conditions. Luisa Bentivogli, Nicola Bertoldi, Mauro Cettolo, Marcello Federico, Matteo Negri, Marco Turchi |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Online adaptation to post-edits for phrase-based statistical machine translation
Nicola Bertoldi, Patrick Simianer, Mauro Cettolo, Katharina Wäschle, Marcello Federico, Stefan Riezler |
Mach. Transl. | 3 |
| 2014 | Translation project adaptation for MT-enhanced computer assisted translation
Mauro Cettolo, Nicola Bertoldi, Marcello Federico, Holger Schwenk, Loïc Barrault, Christophe Servan |
Mach. Transl. | 1 |
| 2013 | Cache-based Online Adaptation for Machine Translation Enhanced Computer Assisted Translation
Nicola Bertoldi, Mauro Cettolo, Marcello Federico |
MTSummit | 2 |
| 2013 | Project Adaptation for MT-Enhanced Computer Assisted Translation
Mauro Cettolo, Nicola Bertoldi, Marcello Federico |
MTSummit | 1 |
| 2012 | WIT3: Web Inventory of Transcribed and Translated Talks
Mauro Cettolo, Christian Girardi, Marcello Federico |
EAMT | 1 |
| 2012 | The IWSLT 2011 Evaluation Campaign on Automatic Talk Translation
Marcello Federico, Sebastian Stüker, Luisa Bentivogli, Michael Paul, Mauro Cettolo, Teresa Herrmann, Jan Niehues, Giovanni Moretti |
LREC | 5 |
| 2011 | Bootstrapping Arabic-Italian SMT through Comparable Texts and Pivot Translation
Mauro Cettolo, Nicola Bertoldi, Marcello Federico |
EAMT | 1 |
| 2011 | Methods for Smoothing the Optimizer Instability in SMT
Mauro Cettolo, Nicola Bertoldi, Marcello Federico |
MTSummit | 1 |
| 2010 | Online Language Model adaptation via N-gram Mixtures for Statistical Machine Translation
Germán Sanchis-Trilles, Mauro Cettolo |
EAMT | 2 |
| 2010 | Statistical Machine Translation of Texts with Misspelled Words
Nicola Bertoldi, Mauro Cettolo, Marcello Federico |
HLT-NAACL | 2 |
| 2008 | IRSTLM: an open source toolkit for handling large scale language modelsabstractResearch in speech recognition and machine translation is boosting the use of large scale n-gram language models. We present an open source toolkit that permits to efficiently handle language models with billions of n-grams on conventional machines. The IRSTLM toolkit supports distribution of ngram collection and smoothing over a computer cluster, language model compression through probability quantization, lazy-loading of huge language models from disk. IRSTLM has been so far successfully deployed with the Moses toolkit for statistical machine translation and with the FBK-irst speech recognition system. Efficiency of the tool is reported on a speech transcription task of Italian political speeches using a language model of 1.1 billion four-grams. Marcello Federico, Nicola Bertoldi, Mauro Cettolo |
INTERSPEECH | 3 |
| 2007 | The IRST English-Spanish translation system for european parliament speeches
Daniele Falavigna, Nicola Bertoldi, Fabio Brugnara, Roldano Cattoni, Mauro Cettolo, Boxing Chen, Marcello Federico, Diego Giuliani, Roberto Gretter, Dino Seppi |
INTERSPEECH | 5 |
| 2007 | Better n-best translations through generative n-gram language models
Boxing Chen, Marcello Federico, Mauro Cettolo |
MTSummit | 3 |
| 2007 | POS-based reordering models for statistical machine translation
Mauro Cettolo, Marcello Federico |
MTSummit | 2 |
| 2006 | A Web-based Demonstrator of a Multi-lingual Phrase-based Translation System
Roldano Cattoni, Nicola Bertoldi, Mauro Cettolo, Boxing Chen, Marcello Federico |
EACL | 3 |
| 2005 | Integrated n-best re-ranking for spoken language translationabstractThis paper describes the application of N-best lists to a spoken language translation system. Multiple hypotheses are generated both by the speech recognizer and by the statistical machine translator; they are finally re-ranked by optimally weighting recognition and translation scores, estimated in an integrated scheme. We provide experimental results for the Italian-to-English direction on the BTEC corpus, a collection of sentences in the touristic domain developed within the C-STAR project. 1. V. H. Quan, Marcello Federico, Mauro Cettolo |
INTERSPEECH | 3 |
| 2005 | Evaluation of BIC-based algorithms for audio segmentation
Mauro Cettolo, Michele Vescovi, Romeo Rizzi |
Comput. Speech Lang. | 1 |
| 2004 | Advances in the automatic transcription of lecturesabstractTranscribing lectures is a challenging task, both in acoustic and in language modeling. In this work, we present recent results on the automatic transcription of lectures from the Translanguage English Database, which contains the recordings of talks given in English at Eurospeech '93, by mostly non-native speakers. Concerning acoustic modeling, the acoustic model trained for a broadcast news transcription task was adapted on the lectures training data through maximum likelihood linear regression adaptation, including models of spontaneous speech phenomena. Moreover, a normalization procedure was embodied in the training stage, consisting of a cluster-based mean and variance normalization of the static features. Language modeling was based on adaptation of a background language model estimated on broadcast news transcripts, conference proceedings, lecture transcripts, and conversational speech transcripts. Among the examined adaptation techniques, the most effective one was obtained by exploiting the paper presented in each lecture to be processed. The best transcription performance on a 2 hours test set was 32.4% word error rate. Mauro Cettolo, Fabio Brugnara, Marcello Federico |
ICASSP (1) | 1 |
| 2003 | The ITC-irst News on Demand Platform
Nicola Bertoldi, Fabio Brugnara, Mauro Cettolo, Marcello Federico, Diego Giuliani, Erwin Leeuwis, Vanessa Sandrini |
ECIR | 3 |
| 2003 | Efficient audio segmentation algorithms based on the BICabstractA widely adopted algorithm for the audio segmentation is based on the Bayesian information criterion (BIC), applied within a sliding variable-size analysis window. In this work, three different implementations of that algorithm are analyzed in detail: (i) one that keeps updated a pair of sums, that of input vectors and that of square input vectors, in order to save computations in estimating covariance matrixes on partially shared data; (ii) one, recently proposed in the literature, that exploits the encoding of the input signal with cumulative statistics for the efficient estimation of covariance matrixes; and (iii) an original one, that encodes the input stream with the cumulative pair of sums of the first approach. The three approaches have been compared both theoretically and experimentally, and the proposed original approach is shown to be the most efficient. Mauro Cettolo, Michele Vescovi |
ICASSP (6) | 1 |
| 2003 | Language modeling and transcription of the TED corpus lecturesabstractTranscribing lectures is a challenging task, both in acoustic and in language modeling. In this work, we present our first results on the automatic transcription of lectures from the TED corpus, recently released by ELRA and LDC. In particular, we concentrated our effort on language modeling. Baseline acoustic and language models were developed using respectively 8 hours of TED transcripts and various types of texts: conference proceedings, lecture transcripts, and conversational speech transcripts. Then, adaptation of the language model to single speakers was investigated by exploiting different kinds of information: automatic transcripts of the talk, the title of the talk, the abstract and, finally, the paper. In the last case, a 39.2% WER was achieved. Erwin Leeuwis, Marcello Federico, Mauro Cettolo |
ICASSP (1) | 3 |
| 2003 | A DP algorithm for speaker change detection
Michele Vescovi, Mauro Cettolo, Romeo Rizzi |
INTERSPEECH | 2 |
| 2002 | Porting an audio partitioner across domainsabstractPartitioning an audio stream means to segment It In acoustically homogeneous chunks, classify segments into acoustic classes, and cluster speech segments. The process represents the earliest stage of automatic transcription stations, since it allows to filter out portions of the audio not containing speech and to improve recognition accuracy through the use of condition-dependent acoustic models and adaptation techniques. Hence, when transcription systems are applied to new domains, the process of porting involves the partitioner module too. In this work, the porting of the partitioner of the ITC-irst broadcast news transcription system to the domain of historical films is described in detail and experimentally evaluated. Moreover, a new technique that makes the porting easier for the automatic estimation of the working point of the BIC-based segmentation algorithm is introduced. Mauro Cettolo |
ICASSP | 1 |
| 2002 | Issues in automatic transcription of historical audio dataabstractThis work deals with some interesting issues that arose when the ITC-irst broadcast news transcription system was applied to transcribe the audio track of historical documentary films. Due to an evident acoustic and linguistic mismatch between the broadcast news and the new application domain, the initial word error rate was of 46.4%. By exploiting a limited amount of manually annotated training data, adaptation of all components of the transcription system was performed, namely the audio partitioner, the acoustic model, and the language model. This permitted to achieve a word error rate of 30%, which makes automatic transcription of documentary films effective for information retrieval applications. Fabio Brugnara, Mauro Cettolo, Marcello Federico, Diego Giuliani |
INTERSPEECH | 2 |
| 2002 | Cross-task portability of a broadcast news speech recognition system
Nicola Bertoldi, Fabio Brugnara, Mauro Cettolo, Marcello Federico, Diego Giuliani |
Speech Commun. | 3 |
| 2001 | From broadcast news to spontaneous dialogue transcription: portability issuesabstractReports on experiments of porting the ITC-irst Italian broadcast news recognition system to two spontaneous dialogue domains. The trade-off between performance and the required amount of task specific data was investigated. Porting was experimented by applying supervised adaptation methods to acoustic and language models. By using two hours of manually transcribed speech, word error rates of 26.0% and 28.4% were achieved by the adapted systems. Two reference systems, developed on a larger training corpus, achieved word error rates of 22.6% and 21.2%, respectively. Nicola Bertoldi, Fabio Brugnara, Mauro Cettolo, Marcello Federico, Diego Giuliani |
ICASSP | 3 |
| 2000 | A baseline for the transcription of Italian broadcast newsabstractThe paper presents the first achievements in the development of a broadcast news transcription system to be applied for the processing of huge audio archives. In particular, the Italian broadcast news corpus under collection is introduced, and the first implemented baseline system is outlined. The baseline system consists of an audio segmentation module and a speech recognizer featuring a recursive Viterbi beam search, a 64k word lexicon, a tree-based trigram LM representation, and MLLR adaptation. The word error rate of the baseline was 20.9% on planned studio speech and 28.8% on the whole test set. Fabio Brugnara, Mauro Cettolo, Marcello Federico, Diego Giuliani |
ICASSP | 2 |
| 2000 | Advances in automatic transcription of Italian broadcast newsabstractThis paper presents some recent improvements in automatic transcription of Italian broadcast news obtained at ITCirst. A first preliminary activity was carried out in order to develop a suitable speech corpus for the Italian language. The resulting corpus, formed by recordings covering 30 hours of radio news, was exploited for developing a baseline system for transcription of broadcast news. The system performs in different stages: acoustic segmentation and classification, speaker clustering, acoustic model adaptation and speech decoding. Major recent advances allowing performance improvement concern with speech segmentation and clustering, acoustic modeling, acoustic model adaptation and the language model. Fabio Brugnara, Mauro Cettolo, Marcello Federico, Diego Giuliani |
INTERSPEECH | 2 |
| 1999 | Semantic boundaries in multiple languagesabstractThis paper presents the results obtained for the task of detecting Semantic Boundaries (SBs) in spoken language using two different methods on the same data set. Hence we first introduce the two approaches developed by ITC-Irst in Trento (Italy) and the LME of the University Erlangen (Germany) and discuss the individually obtained results. The basis for the decision upon SBs in both cases are textual and prosodic features. The LME has already worked for several years on the computation and application of prosodic features in automatic speech processing within the Verbmobil project. The approaches developed in that project were adapted to work on the data collected at IRST in the Italian language. Finally we compare the results we obtain with the German SB detection against the Italian result with regard to precision and recall. 1. INTRODUCTION For robust spoken language processing it is not always necessary to analyse a user's utterance completely as one coherent segment. Often i... Volker Warnke, Heinrich Niemann, Mauro Cettolo, Anna Corazza, Daniele Falavigna, Gianni Lazzari |
EUROSPEECH | 4 |
| 1998 | Automatic recognition of spontaneous speech dialoguesabstractSome approaches for coping with the problem of recognition of spontaneous speech dialogues are presented. Starting from a HMM-based system, developed for dictation tasks, some modifications are introduced at the acoustic level and in the language model. Acoustic model parameters are modified to account for speaking rate variations, and specific models of extra-linguistic phenomena are defined and added to the language model. Different acoustic models are managed by a single recognizer through their integration into a multi-model search space. Experiments and evaluations were conducted on a spontaneous dialogue corpus collected at our laboratory. 1. INTRODUCTION This paper concerns the assessment of a set of modifications applied to a HMM-based continuous speech recognizer developed at ITC-Irst, for dictation tasks, to improve its performance on a corpus of spontaneously uttered humanhuman dialogues. As the performance of Automatic Speech Recognition (ASR) systems is largely affected ... Mauro Cettolo, Daniele Falavigna |
ICSLP | 1 |
| 1998 | Automatic detection of semantic boundaries based on acoustic and lexical knowledgeabstractIn spoken language systems, the segmentation of utterances into coherent linguistic/semantic units is required when modules following the speech recognizer can only process such units one at a time. In this paper, techniques for semantic boundary prediction, based on both acoustic and lexical knowledge, are presented and tested on a corpus of personto -person dialogues. Best result gives 62.8% recall and 71.8% precision. 1. INTRODUCTION In spoken language systems, the minimal unit of analysis does not necessarily correspond to a full sentence. A possible approach for language processing is that of splitting a given sentence in a sequence of units that can be successively processed by linguistic modules one at a time. The goal of the Semantic Boundary (SB) detector is to locate boundaries inside a sentence in order to obtain such "minimal units". Useful information for SB detection can be extracted either from the waveform of an utterance or from its corresponding word sequence. Some ... Mauro Cettolo, Daniele Falavigna |
ICSLP | 1 |
| 1998 | Language portability of a speech understanding system
Mauro Cettolo, Anna Corazza, Renato De Mori |
Comput. Speech Lang. | 1 |
| 1997 | Multilingual person to person communication at IRSTabstractThis paper refers to a machine-mediated person-to-person multilingual communication system. Stress is put on robustness, that is the ability of the system to preserve communication even in presence of the variability and errors typical of spoken language systems. The statistical approach is adopted not only at the acoustic level, but also for the linguistic processing. Therefore, while an overview of the global architecture is briefly introduced, the focus is put on the acoustic recognizer and the understanding module. Experimental evaluations complete the presentation. Bianca Angelini, Mauro Cettolo, Anna Corazza, Daniele Falavigna, Gianni Lazzari |
ICASSP | 2 |
| 1997 | Automatic detection of semantic boundariesabstractIn spoken language systems, the segmentation of utterances into coherent linguistic/semantic units is very useful, as it makes easier processing after the speech recognition phase. In this paper, a methodology for semantic boundary prediction is presented and tested on a corpus of person-to-person dialogues. The approach is based on binary decision trees and uses text context, including broad classes of silent pauses, filled pauses and human noises. Best results give more than 90% precision, almost 80% recall and about 3% false alarms. 1. INTRODUCTION This work focuses on the automatic segmentation of dialogue turns into homogeneous Semantic Units (SUs) [7]. The approach described below is evaluated in the domain of appointment scheduling, where a system able to deal with this kind of interaction between two persons speaking different languages is being developed [1]. As a working hypothesis, it is assumed that each turn can be represented as a "flat" sequence of concepts, i.e. no nes... Mauro Cettolo, Anna Corazza |
EUROSPEECH | 1 |
| 1996 | A mixed approach to speech understanding
Mauro Cettolo, Anna Corazza, Renato De Mori |
ICSLP | 1 |
| 1995 | Language model representations for beam-search decodingabstractThis paper presents an efficient way of representing a bigram language model for a beam-search based, continuous speech, large vocabulary HMM recognizer. The tree-based topology considered takes advantage of a factorization of the bigram probability derived from the bigram interpolation scheme, and of a tree organization of all the words that can follow a given one. Moreover, an optimization algorithm is used to considerably reduce the space requirements of the language model. Experimental results are provided for two 10,000-word dictation tasks: radiological reporting (perplexity 27) and newspaper dictation (perplexity 120). In the former domain 93% word accuracy is achieved with real-time response and 23 Mb process space. In the newspaper dictation domain, 88.1% word accuracy is achieved with 1.41 real-time response and 38 Mb process space. All recognition tests were performed on an HP-735 workstation. Giuliano Antoniol, Fabio Brugnara, Mauro Cettolo, Marcello Federico |
ICASSP | 3 |
| 1995 | Improvements in tree-based language model representationabstractThis paper describes an efficient way of representing a bigram language model with a finite state network used by a beam-search based and continuous speech HMM recognizer. In a previous paper [1], a compact tree-based organization of the search space was presented, that could be further reduced through an optimization algorithm. There, it was pointed out that for a 10,000-word newspaper dictation task the minimization step could have taken a lot of time and space on a standard workstation. In this paper, a new compilation technique that takes into account the particular tree-based topology is described. Results show that without additional time and space costs, the new technique produces networks equivalent to the tree-based ones but almost as small as the optimized one. 1 INTRODUCTION The most widely used Language Models (LMs) in speech recognition are n-gram models, due to both easy inference from the training corpus and easy integrability with the decoding algorithms commonly used... Fabio Brugnara, Mauro Cettolo |
EUROSPEECH | 2 |
| 1995 | Language modelling for efficient beam-search
Marcello Federico, Mauro Cettolo, Fabio Brugnara, Giuliano Antoniol |
Comput. Speech Lang. | 2 |
| 1994 | Radiological reporting by speech recognition: the a.re.s. systemabstractRadiological reporting has already been identified as a field in which voice technologies can prove to be very useful. Recent progress in automatic speech recognition and in hardware and software technology makes it possible to build large-vocabulary, continuous speech, speaker-independent, real-time systems. In this paper a dictation system for radiology reporting, the A.Re.S. system, is presented. A.Re.S. is a "software only" system which runs in real-time on an HP 715 workstation. It relies on an asynchronous and multi-process architecture in which speech decoding is performed by processes in pipeline. System requirements and architecture will be described, together with the results of a preliminary evaluation based on three months of on-site testing. I. INTRODUCTION Recent progress in Automatic Speech Recognition (ASR) and in hardware and software technology makes it possible to build large-vocabulary, real-time, speaker-independent systems. Medical document generation presents f... Bianca Angelini, Giuliano Antoniol, Fabio Brugnara, Mauro Cettolo, Marcello Federico, Roberto Fiutem, Gianni Lazzari |
ICSLP | 4 |
| 1994 | Language model estimations and representations for real-time continuous speech recognitionabstractThis paper compares different ways of estimating bigram language models and of representing them in a finite state network used by a beam-search based, continuous speech, and speaker independent HMM recognizer. Attention is focused on the n-gram interpolation scheme for which seven models are considered. Among them, the Stacked estimated linear interpolated model favourably compares with the best known ones. Further, two different static representations of the search space are investigated: "linear" and "tree-based". Results show that the latter topology is better suited to the beam-search algorithm. Moreover, this representation can be reduced by a network optimization technique, which allows the dynamic size of the recognition process to be decreased by 60%. Extensive recognition experiments on a 10,000-word dictation task with four speakers are described in which an average word accuracy of 93% is achieved with real-time response. I. INTRODUCTION This paper compares different ways ... Giuliano Antoniol, Fabio Brugnara, Mauro Cettolo, Marcello Federico |
ICSLP | 3 |
| 1993 | Techniques for robust recognition in restricted domainsabstractThis paper describes an Automatic Speech Understanding (ASU) system used in a human-robot interface for the remote control of a mobile robot. The intended application is that of an operator issuing telecontrol commands to one or more robots from a remote workstation. ASU is supposed to be performed with spontaneous continuous speech and quasi real time conditions. Training and testing of the system was based on speech data collected by means of Wizard of Oz simulations. Two kinds of robustness factors are introduced: the first is a recognition error-tolerant approach to semantic interpretation, the second is based on a technique for evaluating the reliability of the ASU system output with respect to the input utterance. Preliminary results are 90.9% of correct semantic interpretations, and 89.1% of correct detection of out-of-domain sentences at the cost of rejecting 16.4% of correct in-domain sentences. 1. INTRODUCTION This paper describes an Automatic Speech Understanding (ASU) sys... Giuliano Antoniol, Mauro Cettolo, Marcello Federico |
EUROSPEECH | 2 |