Marcello Federico

dblp:f/MarcelloFederico · DBLP profile ↗
← Back
105ranked-venue papers
20as first author
12since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 86 · 16 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 45 · 10 first-author · 9 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author
YearPublicationVenuePosition
2025 MEMERAG: A Multilingual End-to-End Meta-Evaluation Benchmark for Retrieval Augmented Generation
abstract
María Andrea Cruz Blandón, Jayasimha Talur, Bruno Charron, Dong Liu, Saab Mansour, Marcello Federico. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
María Andrea Cruz Blandón, Jayasimha Talur, Bruno Charron, Saab Mansour, Marcello Federico
ACL (1)6
2024 Perceptual Evaluation of Audio-Visual Synchrony Grounded in Viewers' Opinion Scores
Lucas Goncalves, Prashant Mathur, Chandrashekhar Lavania, Metehan Cekic, Marcello Federico, Kyu J. Han
ECCV (79)5
2023 End-to-End Single-Channel Speaker-Turn Aware Conversational Speech Translation
abstract
Juan Pablo Zuluaga-Gomez, Zhaocheng Huang, Xing Niu, Rohit Paturi, Sundararajan Srinivasan, Prashant Mathur, Brian Thompson, Marcello Federico. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Juan Zuluaga-Gomez, Zhaocheng Huang, Xing Niu 0001, Rohit Paturi, Sundararajan Srinivasan, Prashant Mathur, Brian Thompson 0001, Marcello Federico
EMNLP8
2023 Improving Isochronous Machine Translation with Target Factors and Auxiliary Counters
Proyag Pal, Brian Thompson 0001, Yogesh Virkar, Prashant Mathur, Alexandra Chronopoulou, Marcello Federico
INTERSPEECH6
2022 Duration Modeling of Neural TTS for Automatic Dubbing
abstract
Automatic dubbing (AD) addresses the problem of translating speech in a video with speech in another language while preserving the viewer experience. A most important requirement of AD is isochrony, i.e. dubbed speech has to closely match the timing of speech and pauses of the original audio. In our automatic dubbing system, isochrony is modeled by controlling the verbosity of machine translation; inserting pauses in the translations, a.k.a. prosodic alignment; and controlling the duration of text-to-speech (TTS) utterances. The latter two steps heavily rely on speech duration information, either to predict or control TTS duration. So far, duration prediction was based on a proxy method while duration control on linear warping of the TTS speech spectrogram. In this study, we propose novel duration models for neural TTS that can be leveraged both to predict and control TTS duration. Experimental results show that compared to previous work, the new models improve or match the performance of prosodic alignment and significantly enhance neural TTS speech quality for both slow and fast speaking rates.
Johanes Effendi, Yogesh Virkar, Roberto Barra-Chicote, Marcello Federico
ICASSP4
2022 ISOMETRIC MT: Neural Machine Translation for Automatic Dubbing
abstract
Automatic dubbing (AD) is among the machine translation (MT) use cases where translations should match a given length to allow for synchronicity between source and target speech. For neural MT, generating translations of length close to the source length (e.g. within ±10% in character count), while preserving quality is a challenging task. Controlling MT output length comes at a cost to translation quality, which is usually mitigated with a two step approach of generating N-best hypotheses and then re-ranking based on length and quality. This work introduces a self-learning approach that allows a transformer model to directly learn to generate outputs that closely match the source length, in short Isometric MT. In particular, our approach does not require to generate multiple hypotheses nor any auxiliary ranking function. We report results on four language pairs (English → French, Italian, German, Spanish) with a publicly available benchmark. Automatic and manual evaluations show that our method for Isometric MT outperforms more complex approaches proposed in the literature.
Surafel Melaku Lakew, Yogesh Virkar, Prashant Mathur, Marcello Federico
ICASSP4
2022 Isochrony-Aware Neural Machine Translation for Automatic Dubbing
abstract
We introduce the task of isochrony-aware machine translation which aims at generating translations suitable for dubbing.Dubbing of a spoken sentence requires transferring the content as well as the speech-pause structure of the source into the target language to achieve audiovisual coherence.Practically, this implies correctly projecting pauses from the source to the target and ensuring that target speech segments have roughly the same duration of the corresponding source speech segments.In this work, we propose implicit and explicit modeling approaches to integrate isochrony information into neural machine translation.Experiments on English-German/French language pairs with automatic metrics show that the simplest of the considered approaches works best.Results are confirmed by human evaluations of translations and dubbed videos.
Derek Tam, Surafel Melaku Lakew, Yogesh Virkar, Prashant Mathur, Marcello Federico
INTERSPEECH5
2022 Prosodic alignment for off-screen automatic dubbing
abstract
The goal of automatic dubbing is to perform speech-to-speech translation while achieving audiovisual coherence.This entails isochrony, i.e., translating the original speech by also matching its prosodic structure into phrases and pauses, especially when the speaker's mouth is visible.In previous work, we introduced a prosodic alignment model to address isochrone or on-screen dubbing.In this work, we extend the prosodic alignment model to also address off-screen dubbing that requires less stringent synchronization constraints.We conduct experiments on four dubbing directions -English to French, Italian, German and Spanish -on a publicly available collection of TED Talks and on publicly available YouTube videos.Empirical results show that compared to our previous work the extended prosodic alignment model provides significantly better subjective viewing experience on videos in which on-screen and off-screen automatic dubbing is applied for sentences with speakers mouth visible and not visible, respectively.
Yogesh Virkar, Marcello Federico, Robert Enyedi, Roberto Barra-Chicote
INTERSPEECH2
2021 Machine Translation Verbosity Control for Automatic Dubbing
abstract
Automatic dubbing aims at seamlessly replacing the speech in a video document with synthetic speech in a different language. The task implies many challenges, one of which is generating translations that not only convey the original content, but also match the duration of the corresponding utterances. In this paper, we focus on the problem of controlling the verbosity of machine translation out-put, so that subsequent steps of our automatic dubbing pipeline can generate dubs of better quality. We propose new methods to control the verbosity of MT output and compare them against the state of the art with both intrinsic and extrinsic evaluations. For our experiments we use a public data set to dub English speeches into French, Italian, German and Spanish. Finally, we report extensive subjective tests that measure the impact of MT verbosity control on the final quality of dubbed video clips.
Surafel Melaku Lakew, Marcello Federico, Yue Wang 0034, Cuong Hoang, Yogesh Virkar, Roberto Barra-Chicote, Robert Enyedi
ICASSP2
2021 Improvements to Prosodic Alignment for Automatic Dubbing
abstract
Automatic dubbing is an extension of speech-to-speech translation such that the resulting target speech is carefully aligned in terms of duration, lip movements, timbre, emotion, prosody, etc. of the speaker in order to achieve audiovisual coherence. Dubbing quality strongly depends on isochrony, i.e., arranging the translation of the original speech to optimally match its sequence of phrases and pauses. To this end, we present improvements to the prosodic alignment component of our recently introduced dubbing architecture. We present empirical results for four dubbing directions – English to French, Italian, German and Spanish – on a publicly available collection of TED Talks. Compared to previous work, our enhanced prosodic alignment model significantly improves prosodic alignment accuracy and provides segmentation perceptibly better or on par with manually annotated reference segmentation.
Yogesh Virkar, Marcello Federico, Robert Enyedi, Roberto Barra-Chicote
ICASSP2
2021 Intra-Sentential Speaking Rate Control in Neural Text-To-Speech for Automatic Dubbing
Yogesh Virkar, Marcello Federico, Roberto Barra-Chicote, Robert Enyedi
Interspeech3
2021 Towards Modeling the Style of Translators in Neural Machine Translation
abstract
One key ingredient of neural machine translation is the use of large datasets from different domains and resources (e.g.Europarl, TED talks).These datasets contain documents translated by professional translators using different but consistent translation styles.Despite that, the model is usually trained in a way that neither explicitly captures the variety of translation styles present in the data nor translates new data in different and controllable styles.In this work, we investigate methods to augment the state-of-the-art Transformer model with translator information that is available in part of the training data.We show that our style-augmented translation models are able to capture the style variations of translators and to generate translations with different styles on new data.Indeed, the generated variations differ significantly, up to +4.5 BLEU score difference.Despite that, human evaluation confirms that the translations are of the same quality.
Yue Wang 0034, Cuong Hoang, Marcello Federico
NAACL-HLT3
2020 Evaluating and Optimizing Prosodic Alignment for Automatic Dubbing
Marcello Federico, Yogesh Virkar, Robert Enyedi, Roberto Barra-Chicote
INTERSPEECH1
2019 Training Neural Machine Translation to Apply Terminology Constraints
abstract
This paper proposes a novel method to inject custom terminology into neural machine translation at run time.Previous works have mainly proposed modifications to the decoding algorithm in order to constrain the output to include run-time-provided target terms.While being effective, these constrained decoding methods add, however, significant computational overhead to the inference step, and, as we show in this paper, can be brittle when tested in realistic conditions.In this paper we approach the problem by training a neural MT system to learn how to use custom terminology when provided with the input.Comparative experiments show that our method is not only more effective than a state-of-the-art implementation of constrained decoding, but is also as fast as constraint-free decoding.
Georgiana Dinu, Prashant Mathur, Marcello Federico, Yaser Al-Onaizan
ACL (1)3
2018 A Comparison of Transformer and Recurrent Neural Networks on Multilingual Neural Machine Translation
abstract
Recently, neural machine translation (NMT) has been extended to multilinguality, that is to handle more than one translation direction with a single system. Multilingual NMT showed competitive performance against pure bilingual systems. Notably, in low-resource settings, it proved to work effectively and efficiently, thanks to shared representation space that is forced across languages and induces a sort of transfer-learning. Furthermore, multilingual NMT enables so-called zero-shot inference across language pairs never seen at training time. Despite the increasing interest in this framework, an in-depth analysis of what a multilingual NMT model is capable of and what it is not is still missing. Motivated by this, our work (i) provides a quantitative and comparative analysis of the translations produced by bilingual, multilingual and zero-shot systems; (ii) investigates the translation quality of two of the currently dominant neural architectures in MT, which are the Recurrent and the Transformer ones; and (iii) quantitatively explores how the closeness between languages influences the zero-shot translation. Our analysis leverages multiple professional post-edits of automatic translations by several different systems and focuses both on automatic standard metrics (BLEU and TER) and on widely used error categories, which are lexical, morphology, and word order errors.
Surafel Melaku Lakew, Mauro Cettolo, Marcello Federico
COLING3
2018 Compositional Source Word Representations for Neural Machine Translation
abstract
The requirement for neural machine translation (NMT) models to use fixed-size input and output vocabularies plays an important role for their accuracy and generalization capability. The conventional approach to cope with this limitation is performing translation based on a vocabulary of sub-word units that are predicted using statistical word segmentation methods. However, these methods have recently shown to be prone to morphological errors, which lead to inaccurate translations. In this paper, we extend the source-language embedding layer of the NMT model with a bi-directional recurrent neural network that generates compositional representations of the source words from embeddings of character n-grams. Our model consistently outperforms conventional NMT with sub-word units on four translation directions with varying degrees of morphological complexity and data sparseness on the source side.
Duygu Ataman, Mattia Antonino Di Gangi, Marcello Federico
EAMT3
2018 The ModernMT Project
abstract
This short presentation introduces ModernMT: an open-source project 1 that integrates real-time adaptive neural machine translation into a single easy-to-use product.
Nicola Bertoldi, Davide Caroselli, Marcello Federico
EAMT3
2018 Evaluation of Terminology Translation in Instance-Based Neural MT Adaptation
abstract
We address the issues arising when a neural machine translation engine trained on generic data receives requests from a new domain that contains many specific technical terms. Given training data of the new domain, we consider two alternative methods to adapt the generic system: corpus-based and instance-based adaptation. While the first approach is computationally more intensive in generating a domain-customized network, the latter operates more efficiently at translation time and can handle on-the-fly adaptation to multiple domains. Besides evaluating the generic and the adapted networks with conventional translation quality metrics, in this paper we focus on their ability to properly handle domain-specific terms. We show that instance-based adaptation, by fine-tuning the model on-the-fly, is capable to significantly boost the accuracy of translated terms, producing translations of quality comparable to the expensive corpusbased method.
M. Amin Farajian, Nicola Bertoldi, Matteo Negri, Marco Turchi, Marcello Federico
EAMT5
2018 Deep Neural Machine Translation with Weakly-Recurrent Units
abstract
Recurrent neural networks (RNNs) have represented for years the state of the art in neural machine translation. Recently, new architectures have been proposed, which can leverage parallel computation on GPUs better than classical RNNs. Faster training and inference combined with different sequence-to-sequence modeling also lead to performance improvements. While the new models completely depart from the original recurrent architecture, we decided to investigate how to make RNNs more efficient. In this work, we propose a new recurrent NMT architecture, called Simple Recurrent NMT, built on a class of fast and weakly-recurrent units that use layer normalization and multiple attentions. Our experiments on the WMT14 English-to-German and WMT16 English-Romanian benchmarks show that our model represents a valid alternative to LSTMs, as it can achieve better results at a significantly lower computational cost.
Mattia Antonino Di Gangi, Marcello Federico
EAMT2
2018 Neural versus phrase-based MT quality: An in-depth analysis on English-German and English-French
Luisa Bentivogli, Arianna Bisazza, Mauro Cettolo, Marcello Federico
Comput. Speech Lang.4
2017 Assessing the Tolerance of Neural Machine Translation Systems Against Speech Recognition Errors
abstract
Machine translation systems are conventionally trained on textual resources that do not model phenomena that occur in spoken language. While the evaluation of neural machine translation systems on textual inputs is actively researched in the literature , little has been discovered about the complexities of translating spoken language data with neural models. We introduce and motivate interesting problems one faces when considering the translation of automatic speech recognition (ASR) outputs on neural machine translation (NMT) systems. We test the robustness of sentence encoding approaches for NMT encoder-decoder modeling, focusing on word-based over byte-pair encoding. We compare the translation of utterances containing ASR errors in state-of-the-art NMT encoder-decoder systems against a strong phrase-based machine translation baseline in order to better understand which phenomena present in ASR outputs are better represented under the NMT framework than approaches that represent translation as a linear model.
Nicholas Ruiz, Mattia Antonino Di Gangi, Nicola Bertoldi, Marcello Federico
INTERSPEECH4
2017 Automatic translation memory cleaning
Matteo Negri, Duygu Ataman, Masoud Jalili Sabet, Marco Turchi, Marcello Federico
Mach. Transl.5
2017 Word from the editors
Constantin Orasan, Marcello Federico
Mach. Transl.2
2016 Neural versus Phrase-Based Machine Translation Quality: a Case Study
abstract
Within the field of Statistical Machine Translation (SMT), the neural approach (NMT) has recently emerged as the first technology able to challenge the long-standing dominance of phrase-based approaches (PBMT).In particular, at the IWSLT 2015 evaluation campaign, NMT outperformed well established state-ofthe-art PBMT systems on English-German, a language pair known to be particularly hard because of morphology and syntactic differences.To understand in what respects NMT provides better translation quality than PBMT, we perform a detailed analysis of neural vs. phrase-based SMT outputs, leveraging high quality post-edits performed by professional translators on the IWSLT data.For the first time, our analysis provides useful insights on what linguistic phenomena are best modeled by neural models -such as the reordering of verbs -while pointing out other aspects that remain to be improved.
Luisa Bentivogli, Arianna Bisazza, Mauro Cettolo, Marcello Federico
EMNLP4
2016 WAGS: A Beautiful English-Italian Benchmark Supporting Word Alignment Evaluation on Rare Words
Luisa Bentivogli, Mauro Cettolo, M. Amin Farajian, Marcello Federico
LREC4
2016 A Survey of Word Reordering in Statistical Machine Translation: Computational Models and Language Phenomena
abstract
Word reordering is one of the most difficult aspects of statistical machine translation (SMT), and an important factor of its quality and efficiency. Despite the vast amount of research published to date, the interest of the community in this problem has not decreased, and no single method appears to be strongly dominant across language pairs. Instead, the choice of the optimal approach for a new translation task still seems to be mostly driven by empirical trials. To orient the reader in this vast and complex research area, we present a comprehensive survey of word reordering viewed as a statistical modeling challenge and as a natural language phenomenon. The survey describes in detail how word reordering is modeled within different string-based and tree-based SMT frameworks and as a stand-alone task, including systematic overviews of the literature in advanced reordering modeling. We then question why some approaches are more successful than others in different language pairs. We argue that besides measuring the amount of reordering, it is important to understand which kinds of reordering occur in a given language pair. To this end, we conduct a qualitative analysis of word reordering phenomena in a diverse sample of language pairs, based on a large collection of linguistic knowledge. Empirical results in the SMT literature are shown to support the hypothesis that a few linguistic facts can be very useful to anticipate the reordering characteristics of a language pair and to select the SMT framework that best suits them.
Arianna Bisazza, Marcello Federico
Comput. Linguistics2
2016 The first Automatic Translation Memory Cleaning Shared Task
Eduard Barbu, Carla Parra Escartín, Luisa Bentivogli, Matteo Negri, Marco Turchi, Constantin Orasan, Marcello Federico
Mach. Transl.7
2016 Word from the editors
Constantin Orasan, Marcello Federico
Mach. Transl.2
2016 On the Evaluation of Adaptive Machine Translation for Human Post-Editing
abstract
We investigate adaptive machine translation (MT) as a way to reduce human workload and enhance user experience when professional translators operate in real-life conditions. A crucial aspect in our analysis is how to ensure a reliable assessment of MT technologies aimed to support human post-editing. We pay particular attention to two evaluation aspects: i) the design of a sound experimental protocol to reduce the risk of collecting biased measurements, and ii) the use of robust statistical testing methods (linear mixed-effects models) to reduce the risk of under/over-estimating the observed variations. Our adaptive MT technology is integrated in a web-based full-fledged computer-assisted translation (CAT) tool. We report on a post-editing field test that involved 16 professional translators working on two translation directions (English-Italian and English-French), with texts coming from two linguistic domains (legal, information technology). Our contrastive experiments compare user post-editing effort with static vs. adaptive MT in an end-to-end scenario where the system is evaluated as a whole. Our results evidence that adaptive MT leads to an overall reduction in post-editing effort (HTER) up to 10.6% (p <; 0.05). A follow-up manual evaluation of the MT outputs and their corresponding post-edits confirms that the gain in HTER corresponds to higher quality of the adaptive MT system and does not come at the expense of the final human translation quality. Indeed, adaptive MT shows to return better suggestions than static MT (p <; 0.01), and the resulting post-edits do not significantly differ in the two conditions.
Luisa Bentivogli, Nicola Bertoldi, Mauro Cettolo, Marcello Federico, Matteo Negri, Marco Turchi
IEEE ACM Trans. Audio Speech Lang. Process.4
2015 Phonetically-oriented word error alignment for speech recognition error analysis in speech translation
abstract
We propose a variation to the commonly used Word Error Rate (WER) metric for speech recognition evaluation which incorporates the alignment of phonemes, in the absence of time boundary information. After computing the Levenshtein alignment on words in the reference and hypothesis transcripts, spans of adjacent errors are converted into phonemes with word and syllable boundaries and a phonetic Levenshtein alignment is performed. The phoneme alignment information is used to correct the word alignment labels in each error region. We demonstrate that our Phonetically-Oriented Word Error Rate (POWER) yields similar scores to WER with the added advantages of better word alignments and the ability to capture one-to-many alignments corresponding to homophonic errors in speech recognition hypotheses. These improved alignments allow us to better trace the impact of Levenshtein error types in speech recognition on downstream tasks such as speech translation.
Nicholas Ruiz, Marcello Federico
ASRU2
2015 Adapting machine translation models toward misrecognized speech with text-to-speech pronunciation rules and acoustic confusability
abstract
In the spoken language translation pipeline, machine translation systems that are trained solely on written bitexts are often unable to recover from speech recognition errors due to the mismatch in training data. We propose a novel technique to simulate the errors generated by an ASR system, using the ASR system’s pronunciation dictionary and language model. Lexical entries in the pronunciation dictionary are converted into phoneme sequences using a text-to-speech (TTS) analyzer and stored in a phoneme-to-word translation model. The translation model and ASR language model are combined into a phonemeto-word MT system that “damages” clean texts to look like ASR outputs based on acoustic confusions. Training texts are TTSconverted and damaged into synthetic ASR data for use as adaptation data for training a speech translation system. Our proposed technique yields consistent improvements in translation quality on English-French lectures.
Nicholas Ruiz, Qin Gao, William Lewis, Marcello Federico
INTERSPEECH4
2015 Topic adaptation for machine translation of e-commerce content
Prashant Mathur, Marcello Federico, Selçuk Köprü, Sharam Khadivi, Hassan Sawaf
MTSummit2
2015 Introduction to the Special Section on Continuous Space and Related Methods in Natural Language Processing
abstract
The articles in this special section discuss some latest findings on research problems related to the application of continuous space and related models in Natural Language Processing (NLP).
Haizhou Li 0001, Marcello Federico, Xiaodong He 0001, Helen M. Meng, Isabel Trancoso
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 MateCat
Marcello Federico
EAMT1
2014 Complexity of spoken versus written language for machine translation
Nicholas Ruiz, Marcello Federico
EAMT2
2014 Assessing the Impact of Translation Errors on Machine Translation Quality with Mixed-effects Models
abstract
Learning from errors is a crucial aspect of improving expertise.Based on this notion, we discuss a robust statistical framework for analysing the impact of different error types on machine translation (MT) output quality.Our approach is based on linear mixed-effects models, which allow the analysis of error-annotated MT output taking into account the variability inherent to the specific experimental setting from which the empirical observations are drawn.Our experiments are carried out on different language pairs involving Chinese, Arabic and Russian as target languages.Interesting findings are reported, concerning the impact of different error types both at the level of human perception of quality and with respect to performance results measured with automatic metrics.
Marcello Federico, Matteo Negri, Luisa Bentivogli, Marco Turchi
EMNLP1
2014 Online adaptation to post-edits for phrase-based statistical machine translation
Nicola Bertoldi, Patrick Simianer, Mauro Cettolo, Katharina Wäschle, Marcello Federico, Stefan Riezler
Mach. Transl.5
2014 Translation project adaptation for MT-enhanced computer assisted translation
Mauro Cettolo, Nicola Bertoldi, Marcello Federico, Holger Schwenk, Loïc Barrault, Christophe Servan
Mach. Transl.3
2014 Data-driven annotation of binary MT quality estimation corpora based on human post-editions
Marco Turchi, Matteo Negri, Marcello Federico
Mach. Transl.3
2014 Editorial: Expanding the Technical Reach of our Transactions
abstract
We thank the strong support of both IEEE Signal Processing Society and the ACM Publication Boards for this successful merger.IEEE's TASLP is a very well-established publication, strong in both quality and quantity.It is closely linked to ICASSP and to a number of workshops, such as ASRU and SLT.The language area is a relatively recent addition to TASLP, incorporated in 2006.ACM's TSLP is a more recent publication.The quality has been very high, but the quantity has only sustained a quarterly publication.There is no ACM Special Interest Group (SIG) or conference connection in this area, so it has been more difficult to maintain a direct link to the research community.One of the main motivations for the merger is that IEEE's TASLP does not yet have a strong profile in the language processing community, and this is reflected in the submissions received and the composition of the Editorial Board.ACM's TSLP has a stronger profile and Editorial Board membership in this area.Thus, it is clear that a joint transactions will be stronger than either publication on its own.For several years, the IEEE Signal Processing Society has recognized the importance of information processing in the work of a wide range of researchers within the Society beyond the traditional scope of signal processing.This led to the technical scope of the IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING being expanded to include Language Processing in 2006.While serving as editor of the IEEE SIGNAL PROCESSING MAGAZINE, one of the authors wrote in 2008 and 2010 two editorials [1], [2] that elaborated on the need for expanding the technical reach of signal processing by including the new "understanding" or "interpretation" component of signals consisting of language/text and bio-sequence data.They are both symbolic in nature, which were outside of the traditional definition of "signal" with numerical values in nature.This editorial on the formation of the joint IEEE/ACM TASLP, which highlights the importance of text or written language processing, is a concrete embodiment of the goal of expanding the technical reach of signal processing.As part of the merger of the two transactions, we have taken the opportunity to restructure the editorial board.In addition to the role of Editor-in-Chief, there will be six Senior Area Editors to advise the Editor-in-Chief across the full range of topics covered by the merged journal.
Li Deng 0001, Steve Renals, Marcello Federico, Mari Ostendorf
IEEE ACM Trans. Audio Speech Lang. Process.3
2013 Cache-based Online Adaptation for Machine Translation Enhanced Computer Assisted Translation
Nicola Bertoldi, Mauro Cettolo, Marcello Federico
MTSummit3
2013 Project Adaptation for MT-Enhanced Computer Assisted Translation
Mauro Cettolo, Nicola Bertoldi, Marcello Federico
MTSummit3
2013 Generative and Discriminative Methods for Online Adaptation in SMT
Katharina Wäschle, Patrick Simianer, Nicola Bertoldi, Stefan Riezler, Marcello Federico
MTSummit5
2013 Dynamically Shaping the Reordering Search Space of Phrase-Based Statistical Machine Translation
abstract
Defining the reordering search space is a crucial issue in phrase-based SMT between distant languages. In fact, the optimal trade-off between accuracy and complexity of decoding is nowadays reached by harshly limiting the input permutation space. We propose a method to dynamically shape such space and, thus, capture long-range word movements without hurting translation quality nor decoding time. The space defined by loose reordering constraints is dynamically pruned through a binary classifier that predicts whether a given input word should be translated right after another. The integration of this model into a phrase-based decoder improves a strong Arabic-English baseline already including state-of-the-art early distortion cost (Moore and Quirk, 2007) and hierarchical phrase orientation models (Galley and Manning, 2008). Significant improvements in the reordering of verbs are achieved by a system that is notably faster than the baseline, while bleu and meteor remain stable, or even increase, at a very high distortion limit.
Arianna Bisazza, Marcello Federico
Trans. Assoc. Comput. Linguistics2
2012 Modified Distortion Matrices for Phrase-Based Statistical Machine Translation
Arianna Bisazza, Marcello Federico
ACL (1)2
2012 Cutting the Long Tail: Hybrid Language Models for Translation Style Adaptation
Arianna Bisazza, Marcello Federico
EACL2
2012 WIT3: Web Inventory of Transcribed and Translated Talks
Mauro Cettolo, Christian Girardi, Marcello Federico
EAMT3
2012 Crowd-based MT Evaluation for non-English Target Languages
Michael Paul, Eiichiro Sumita, Luisa Bentivogli, Marcello Federico
EAMT4
2012 The IWSLT 2011 Evaluation Campaign on Automatic Talk Translation
Marcello Federico, Sebastian Stüker, Luisa Bentivogli, Michael Paul, Mauro Cettolo, Teresa Herrmann, Jan Niehues, Giovanni Moretti
LREC1
2012 Chunk-lattices for verb reordering in Arabic-English statistical machine translation - Special issues on machine translation for Arabic
Arianna Bisazza, Daniele Pighin, Marcello Federico
Mach. Transl.3
2011 Using Bilingual Parallel Corpora for Cross-Lingual Textual Entailment
Yashar Mehdad, Matteo Negri, Marcello Federico
ACL3
2011 Bootstrapping Arabic-Italian SMT through Comparable Texts and Pivot Translation
Mauro Cettolo, Nicola Bertoldi, Marcello Federico
EAMT3
2011 NeMo: A Platform for Multilingual News Monitoring
Christian Girardi, Roberto Gretter, Daniele Falavigna, Fabio Brugnara, Diego Giuliani, Marcello Federico
INTERSPEECH6
2011 Getting Expert Quality from the Crowd for Machine Translation Evaluation
Luisa Bentivogli, Marcello Federico, Giovanni Moretti, Michael Paul
MTSummit2
2011 Methods for Smoothing the Optimizer Instability in SMT
Mauro Cettolo, Nicola Bertoldi, Marcello Federico
MTSummit3
2011 Cross-Language Information Retrieval Jian-Yun Nie (University of Montreal) San Rafael, CA: Morgan & Claypool (Synthesis Lectures on Human Language Technologies, edited by Graeme Hirst, volume 8), 2010, xv+125 pp; paperbound, ISBN 978-1-59829-863-5, $40.00; ebook, ISBN 978-1-59829-864-3, $30.00 or by subscription
abstract
Cross-Language Information Retrieval is a compact book introducing a branch of information retrieval that has gained considerable research interest since the dawn of the World Wide Web in the mid 1990s.Information retrieval is generally concerned with the problem of finding documents within a large collection that are relevant to a given input query.Whereas the original formulation of IR assumes that queries and documents are written in the same language, cross-language IR (CLIR) presumes instead that they are written in two different languages.If the collection contains documents in more languages, then we refer to multi-lingual IR (MLIR), which is typically solved with multiple instances of CLIR.Recently, other variations on the theme have been proposed that address non-textual documents, such as image, music, and speech retrieval.An interesting application of CLIR is the retrieval of images that are provided with textual descriptions in any language.Computational linguistics could be interested in CLIR for several reasons.CLIR is mainly about the optimal integration of machine translation (MT) and IR, and it presents peculiar and difficult translation issues when short queries are involved, which is the most common case.For such problems, interesting approaches have been developed and refined over time, which mainly build on top of core statistical MT techniques (e.g., word alignment models, translation models) and various lexical resources (e.g., WordNet, dictionaries).In recent years, several books on IR have been published (e.g., Grossman and Frieder 2004;Manning, Raghavan, and Sch ütze 2008;B üttcher, Clarke, and Cormack 2010), which devoted at most a section or chapter to CLIR.As specific books on CLIR have been limited so far to edited collections of scientific papers (Grefenstette 1998), it was definitely time for the first monograph on the topic.Jian-Yun Nie's volume is structured as five chapters, which are organized as follows:r Chapter 1, "Introduction," covers IR problems, approaches, and models, language problems in IR with European and East Asian languages, CLIR problems and approaches, needs for CLIR and MLIR, and a brief history of CLIR.r Chapter 2, "Using manually constructed translation systems and resources for CLIR," covers an introduction to MT, basic use of MT in CLIR, and dictionary-based translation for CLIR.
Marcello Federico
Comput. Linguistics1
2010 Statistical Machine Translation of Texts with Misspelled Words
Nicola Bertoldi, Mauro Cettolo, Marcello Federico
HLT-NAACL3
2010 Towards Cross-Lingual Textual Entailment
Yashar Mehdad, Matteo Negri, Marcello Federico
HLT-NAACL3
2009 Coping with out-of-vocabulary words: Open versus huge vocabulary asr
abstract
This paper investigates methods for coping with out-of-vocabulary words in a large vocabulary speech recognition task, namely the automatic transcription of Italian broadcast news. Two alternative ways for augmenting a 64 K(thousand)-word recognition vocabulary and language model are compared: introducing extra words with their phonetic transcription up to 1.2 M (million) words, or extending the language model with so-called graphones, i.e. subword units made of phone-character sequences. Graphones and phonetic transcriptions of words are automatically generated by adapting an off-the-shelf statistical machine translation toolkit. We found that the word-based and graphone based extentions allow both for better recognition performance, with the former performing significantly better than the latter. In addition, the word-based extension approach shows interesting potential even under conditions of little supervision. In fact, by training the grapheme to phoneme translation system with only 2 K manually verified transcriptions, the final word error rate increases by just 3% relative, with respect to starting from a lexicon of 64 K words.
Matteo Gerosa, Marcello Federico
ICASSP2
2008 Fast speech decoding through phone confusion networks
abstract
We present a two stage automatic speech recognition architecture suited for applications, such as spoken document retrieval, where large scale language models can be used and very low out-of-vocabulary rates need to be reached. The proposed system couples a weakly constrained phone-recognizer with a phone-to-word decoder that was originally developed for phrase-based statistical machine translation. The decoder permits to efficiently decode confusion networks in input, and to exploit large scale unpruned language models. Preliminary experiments are reported on the transcription of speeches of the Italian parliament. The use of phone confusion networks as interface between the two decoding steps permits to reduce the WER by 28%, thus making the system perform relatively close to a state-of-the-art baseline using a comparable language model.
Nicola Bertoldi, Marcello Federico, Daniele Falavigna, Matteo Gerosa
INTERSPEECH2
2008 IRSTLM: an open source toolkit for handling large scale language models
abstract
Research in speech recognition and machine translation is boosting the use of large scale n-gram language models. We present an open source toolkit that permits to efficiently handle language models with billions of n-grams on conventional machines. The IRSTLM toolkit supports distribution of ngram collection and smoothing over a computer cluster, language model compression through probability quantization, lazy-loading of huge language models from disk. IRSTLM has been so far successfully deployed with the Moses toolkit for statistical machine translation and with the FBK-irst speech recognition system. Efficiency of the tool is reported on a speech transcription task of Italian political speeches using a language model of 1.1 billion four-grams.
Marcello Federico, Nicola Bertoldi, Mauro Cettolo
INTERSPEECH1
2008 Efficient Speech Translation Through Confusion Network Decoding
abstract
This paper describes advances in the use of confusion networks as interface between automatic speech recognition and machine translation. In particular, it presents a decoding algorithm for confusion networks which results as an extension of a state-of-the-art phrase-based text translation decoder. The confusion network decoder significantly improves both in efficiency and performance over previous work along this direction, and outperforms the background text translation system. Experimental results in terms of translation accuracy and decoding efficiency are reported for the task of translating plenary speeches of the European Parliament from Spanish to english and from english to Spanish.
Nicola Bertoldi, Richard Zens, Marcello Federico, Wade Shen
IEEE Trans. Speech Audio Process.3
2008 System Combination for Machine Translation of Spoken and Written Language
abstract
This paper describes an approach for computing a consensus translation from the outputs of multiple machine translation (MT) systems. The consensus translation is computed by weighted majority voting on a confusion network, similarly to the well-established ROVER approach of Fiscus for combining speech recognition hypotheses. To create the confusion network, pairwise word alignments of the original MT hypotheses are learned using an enhanced statistical alignment algorithm that explicitly models word reordering. The context of a whole corpus of automatic translations rather than a single sentence is taken into account in order to achieve high alignment quality. The confusion network is rescored with a special language model, and the consensus translation is extracted as the best path. The proposed system combination approach was evaluated in the framework of the TC-STAR speech translation project. Up to six state-of-the-art statistical phrase-based translation systems from different project partners were combined in the experiments. Significant improvements in translation quality from Spanish to English and from English to Spanish in comparison with the best of the individual MT systems were achieved under official evaluation conditions.
Evgeny Matusov, Gregor Leusch, Rafael E. Banchs, Nicola Bertoldi, Daniel Déchelotte, Marcello Federico, Muntsin Kolss, Young-Suk Lee 0001, José B. Mariño, Matthias Paulik, Salim Roukos, Holger Schwenk, Hermann Ney
IEEE Trans. Speech Audio Process.6
2007 Moses: Open Source Toolkit for Statistical Machine Translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondrej Bojar, Alexandra Constantin, Evan Herbst
ACL5
2007 Speech Translation by Confusion Network Decoding
abstract
This paper describes advances in the use of confusion networks as interface between automatic speech recognition and machine translation. In particular, it presents an implementation of a confusion network decoder which significantly improves both in efficiency and performance previous work along this direction. The confusion network decoder results as an extension of a state-of-the-art phrase-based text translation system. Experimental results in terms of decoding speed and translation accuracy are reported on a real-data task, namely the translation of plenary speeches at the European Parliament from Spanish to English.
Nicola Bertoldi, Richard Zens, Marcello Federico
ICASSP (4)3
2007 Punctuating confusion networks for speech translation
Roldano Cattoni, Nicola Bertoldi, Marcello Federico
INTERSPEECH3
2007 The IRST English-Spanish translation system for european parliament speeches
Daniele Falavigna, Nicola Bertoldi, Fabio Brugnara, Roldano Cattoni, Mauro Cettolo, Boxing Chen, Marcello Federico, Diego Giuliani, Roberto Gretter, Dino Seppi
INTERSPEECH7
2007 Better n-best translations through generative n-gram language models
Boxing Chen, Marcello Federico, Mauro Cettolo
MTSummit2
2007 POS-based reordering models for statistical machine translation
Mauro Cettolo, Marcello Federico
MTSummit3
2006 A Web-based Demonstrator of a Multi-lingual Phrase-based Translation System
Roldano Cattoni, Nicola Bertoldi, Mauro Cettolo, Boxing Chen, Marcello Federico
EACL5
2006 Exploiting Word Transformation in Statistical Machine Translation from Spanish to English
Marcello Federico
EAMT2
2006 Mark Johnson, Sanjeev P. Khudanpur, Mari Ostendorf and Roni Rosenfeld (eds) Mathematical Foundations of Speech and Language Processing
Marcello Federico
Mach. Transl.1
2005 Integrated n-best re-ranking for spoken language translation
abstract
This paper describes the application of N-best lists to a spoken language translation system. Multiple hypotheses are generated both by the speech recognizer and by the statistical machine translator; they are finally re-ranked by optimally weighting recognition and translation scores, estimated in an integrated scheme. We provide experimental results for the Italian-to-English direction on the BTEC corpus, a collection of sentences in the touristic domain developed within the C-STAR project. 1.
V. H. Quan, Marcello Federico, Mauro Cettolo
INTERSPEECH2
2004 Advances in the automatic transcription of lectures
abstract
Transcribing lectures is a challenging task, both in acoustic and in language modeling. In this work, we present recent results on the automatic transcription of lectures from the Translanguage English Database, which contains the recordings of talks given in English at Eurospeech '93, by mostly non-native speakers. Concerning acoustic modeling, the acoustic model trained for a broadcast news transcription task was adapted on the lectures training data through maximum likelihood linear regression adaptation, including models of spontaneous speech phenomena. Moreover, a normalization procedure was embodied in the training stage, consisting of a cluster-based mean and variance normalization of the static features. Language modeling was based on adaptation of a background language model estimated on broadcast news transcripts, conference proceedings, lecture transcripts, and conversational speech transcripts. Among the examined adaptation techniques, the most effective one was obtained by exploiting the paper presented in each lecture to be processed. The best transcription performance on a 2 hours test set was 32.4% word error rate.
Mauro Cettolo, Fabio Brugnara, Marcello Federico
ICASSP (1)3
2004 Broadcast news LM adaptation over time
Marcello Federico, Nicola Bertoldi
Comput. Speech Lang.1
2004 Statistical Models for Monolingual and Bilingual Information Retrieval
Nicola Bertoldi, Marcello Federico
Inf. Retr.2
2003 The ITC-irst News on Demand Platform
Nicola Bertoldi, Fabio Brugnara, Mauro Cettolo, Marcello Federico, Diego Giuliani, Erwin Leeuwis, Vanessa Sandrini
ECIR4
2003 Spoken Information Extraction from Italian Broadcast News
Vanessa Sandrini, Marcello Federico
ECIR2
2003 Language modeling and transcription of the TED corpus lectures
abstract
Transcribing lectures is a challenging task, both in acoustic and in language modeling. In this work, we present our first results on the automatic transcription of lectures from the TED corpus, recently released by ELRA and LDC. In particular, we concentrated our effort on language modeling. Baseline acoustic and language models were developed using respectively 8 hours of TED transcripts and various types of texts: conference proceedings, lecture transcripts, and conversational speech transcripts. Then, adaptation of the language model to single speakers was investigated by exploiting different kinds of information: automatic transcripts of the talk, the title of the talk, the abstract and, finally, the paper. In the last case, a 39.2% WER was achieved.
Erwin Leeuwis, Marcello Federico, Mauro Cettolo
ICASSP (1)2
2003 Evaluation frameworks for speech translation technologies
abstract
This paper reports on activities carried out under the European project PF-STAR and within the CSTAR consortium, which aim at evaluating speech translation technologies. In PF-STAR, speech translation baselines developed by the partners and off-the-shelf commercial systems will be compared systematically on several language pairs and application scenarios. In CSTAR, evaluation campaigns will be organized, on a regular basis, to compare research baselines developed by the members of the consortium. The first evaluation campaign, which will take place in 2003, will focus on written language translation by exploiting a large phrase-book parallel corpus covering several European and Asiatic languages.
Marcello Federico
INTERSPEECH1
2002 Bootstrapping Named Entity Recognition for Italian Broadcast News
abstract
This paper presents the development of a Named Entity (NE) recognition system for the Italian broadcast news domain. A statistical model is introduced based on a trigram language model defined on words and NE classes. The estimation of the NE model is carried out with a very little list of 2,360 manually tagged NEs and a large untagged newspaper corpus. An iterative training procedure is applied which goes through the estimation of simpler models, whose parameters are used to initialize the complete NE model. In the end, NE recognition experiments are reported, on broadcast news transcripts generated by a speech recognition system.
Marcello Federico, Nicola Bertoldi, Vanessa Sandrini
EMNLP1
2002 Language model adaptation through topic decomposition and MDI estimation
abstract
This work presents a language model adaptation method combining the latent semantic analysis framework with the minimum discrimination information estimation criterion. In particular, an unsupervised topic model decomposition is built which allows to infer topic related word distributions from very short adaptation texts. The resulting word distribution is then used to constraint the estimation of a minimum divergence trigram language. With respect to previous work, implementation details are discussed that make such approach effective for a large scale application. Experimental results are provided for a digital library indexing task, i.e. the speech transcription of five historical documentary films. By adapting a trigram language model from very terse content descriptions, i.e. maximum ten words, available for each film, a word error rate relative reduction of 3.2% was achieved.
Marcello Federico
ICASSP1
2002 Issues in automatic transcription of historical audio data
abstract
This work deals with some interesting issues that arose when the ITC-irst broadcast news transcription system was applied to transcribe the audio track of historical documentary films. Due to an evident acoustic and linguistic mismatch between the broadcast news and the new application domain, the initial word error rate was of 46.4%. By exploiting a limited amount of manually annotated training data, adaptation of all components of the transcription system was performed, namely the audio partitioner, the acoustic model, and the language model. This permitted to achieve a word error rate of 30%, which makes automatic transcription of documentary films effective for information retrieval applications.
Fabio Brugnara, Mauro Cettolo, Marcello Federico, Diego Giuliani
INTERSPEECH3
2002 Statistical cross-language information retrieval using n-best query translations
abstract
This paper presents a novel statistical model for cross-language information retrieval. Given a written query in the source language, documents in the target language are ranked by integrating probabilities computed by two statistical models: a query-translation model, which generates most probable term-by-term translations of the query, and a query-document model, which evaluates the likelihood of each document and translation. Integration of the two scores is performed over the set of N most probable translations of the query. Experimental results with values N=1, 5, 10 are presented on the Italian-English bilingual track data used in the CLEF 2000 and 2001 evaluation campaigns.
Marcello Federico, Nicola Bertoldi
SIGIR1
2002 Cross-task portability of a broadcast news speech recognition system
Nicola Bertoldi, Fabio Brugnara, Mauro Cettolo, Marcello Federico, Diego Giuliani
Speech Commun.4
2001 From broadcast news to spontaneous dialogue transcription: portability issues
abstract
Reports on experiments of porting the ITC-irst Italian broadcast news recognition system to two spontaneous dialogue domains. The trade-off between performance and the required amount of task specific data was investigated. Porting was experimented by applying supervised adaptation methods to acoustic and language models. By using two hours of manually transcribed speech, word error rates of 26.0% and 28.4% were achieved by the adapted systems. Two reference systems, developed on a larger training corpus, achieved word error rates of 22.6% and 21.2%, respectively.
Nicola Bertoldi, Fabio Brugnara, Mauro Cettolo, Marcello Federico, Diego Giuliani
ICASSP4
2001 Broadcast news LM adaptation using contemporary texts
abstract
This paper investigates the problem of dynamically updating the language model (LM) of a broadcast news speech recognition system, in order to cope with language and topic changes, typical of the news domain. Statistical adaptation methods are proposed that exploit written news sources which are daily available on the Internet, i.e. newswires and newspapers. Specifically, LM adaptation is performed by extending the basic lexicon, in order to minimize the out-of-vocabulary (OOV) rate, and by adapting the word probability distribution on the contemporary data. Experiments performed on 19 newscasts showed relative reductions of 58% on the OOV rate, 16% on the perplexity, and 4% on the word error rate.
Marcello Federico, Nicola Bertoldi
INTERSPEECH1
2000 A baseline for the transcription of Italian broadcast news
abstract
The paper presents the first achievements in the development of a broadcast news transcription system to be applied for the processing of huge audio archives. In particular, the Italian broadcast news corpus under collection is introduced, and the first implemented baseline system is outlined. The baseline system consists of an audio segmentation module and a speech recognizer featuring a recursive Viterbi beam search, a 64k word lexicon, a tree-based trigram LM representation, and MLLR adaptation. The word error rate of the baseline was 20.9% on planned studio speech and 28.8% on the whole test set.
Fabio Brugnara, Mauro Cettolo, Marcello Federico, Diego Giuliani
ICASSP3
2000 Advances in automatic transcription of Italian broadcast news
abstract
This paper presents some recent improvements in automatic transcription of Italian broadcast news obtained at ITCirst. A first preliminary activity was carried out in order to develop a suitable speech corpus for the Italian language. The resulting corpus, formed by recordings covering 30 hours of radio news, was exploited for developing a baseline system for transcription of broadcast news. The system performs in different stages: acoustic segmentation and classification, speaker clustering, acoustic model adaptation and speech decoding. Major recent advances allowing performance improvement concern with speech segmentation and clustering, acoustic modeling, acoustic model adaptation and the language model.
Fabio Brugnara, Mauro Cettolo, Marcello Federico, Diego Giuliani
INTERSPEECH3
2000 Development and Evaluation of an Italian Broadcast News Corpus
Marcello Federico, Dimitri Giordani, Paolo Coletti
LREC1
2000 A system for the retrieval of Italian broadcast news
Marcello Federico
Speech Commun.1
1999 Usability field-test of a spoken data-entry system
abstract
This paper reports on the field-test of a speech based data-entry system developed as a follow-up of an EC funded project. The application domain is the data-entry of personnel absence records from a huge historical paper file (about 100,000 records). The application was required by the personnel office of a public administration. The tested system resulted both sufficiently simple to make a detailed analysis feasible, and sufficiently representative of the potentials of spoken data-entry.
Marcello Federico, Fabio Brugnara, Roberto Gretter
ICASSP1
1999 A two-stage speech recognition method for information retrieval applications
abstract
This paper presents a two-stage approach to speech recognition that is suited for information retrieval tasks, e.g. accessing a large telephone directory. The first stage performs a Viterbi beam search to decode the speech input into a sequence of phonemes. The second stage performs a graph search to match the phoneme sequence with a large list of keywords. The key issue is that the first step employs a syllable based language model that does not necessarily depend on the application domain. Experimental results are shown for a telephone directory access task of one million of entries.
Paolo Coletti, Marcello Federico
EUROSPEECH2
1999 Efficient language model adaptation through MDI estimation
abstract
This paper presents a method for n-gram language model adaptation based on the principle of minimum discrimination information. A background language model is adapted to fit constraints on its marginal distributions that are derived from new observed data. This work gives a different derivation of the model by Kneser et al. (1997) and extends its application to interpolated language models. The proposed method has been evaluated on an Italian 60K-word broadcast news task.
Marcello Federico
EUROSPEECH1
1997 Speedata: a prototype for multilingual spoken data-entry
abstract
In this work we describe the development and evaluation of SpeeData, a prototype for multilingual spoken dataentry. The SpeeData project aims at developing a demonstrator that provides a user-friendly interface for spoken data-entry in two languages: Italian and German. A real world application domain is considered, which is the Land Register of an Italian region in which both languages are officially spoken. Original topics of this paper are the interaction modality for spoken data-entry, the evaluation of a data-entry system, bilingual speech recognition, bilingual speaker adaptation. 1. INTRODUCTION Data-entry can be particularly costly when non electronic information - e.g. contained in documents or pictures - has to be interpreted by a domain expert before being stored into the computer - e.g. medical reporting, diagnostics, cataloging, etc. SpeeData [1] aims at exploring how to gain efficiency in this task by employing state-ofthe -art ASR technology. The ideal scenario would b...
Ulla Ackermann, Bianca Angelini, Fabio Brugnara, Marcello Federico, Diego Giuliani, Roberto Gretter, Heinrich Niemann
EUROSPEECH4
1997 Dynamic language models for interactive speech applications
Fabio Brugnara, Marcello Federico
EUROSPEECH2
1996 Speedata: multilingual spoken data entry
abstract
In this paper we present a m ultilingual application for speech t echnology.T h e SpeeData project aims at building a d emonstrator that provides a user-friendly interface for spoken data-entry in two languages: Italian and German.The a p plication domain is the land register of an Italian region in which b o t h languages are ocially spoken.The considered data-entry task is particularly challenging as it considers many dierent t ypes of data -e.g.long t exts, numbers, proper names, tables, etc.-and a v ariety of of pronunciations, since dialects are present a n d users will not always speak in their native language.
Ulla Ackermann, Bianca Angelini, Fabio Brugnara, Marcello Federico, Diego Giuliani, Roberto Gretter, Gianni Lazzari, Heinrich Niemann
ICSLP4
1996 Techniques for approximating a trigram language model
Fabio Brugnara, Marcello Federico
ICSLP2
1996 Bayesian estimation methods for n-gram language model adaptation
Marcello Federico
ICSLP1
1995 Language model representations for beam-search decoding
abstract
This paper presents an efficient way of representing a bigram language model for a beam-search based, continuous speech, large vocabulary HMM recognizer. The tree-based topology considered takes advantage of a factorization of the bigram probability derived from the bigram interpolation scheme, and of a tree organization of all the words that can follow a given one. Moreover, an optimization algorithm is used to considerably reduce the space requirements of the language model. Experimental results are provided for two 10,000-word dictation tasks: radiological reporting (perplexity 27) and newspaper dictation (perplexity 120). In the former domain 93% word accuracy is achieved with real-time response and 23 Mb process space. In the newspaper dictation domain, 88.1% word accuracy is achieved with 1.41 real-time response and 38 Mb process space. All recognition tests were performed on an HP-735 workstation.
Giuliano Antoniol, Fabio Brugnara, Mauro Cettolo, Marcello Federico
ICASSP4
1995 A speech understanding architecture for an information query system
Marcello Federico, Fabrizio Vernesoni
EUROSPEECH1
1995 Language modelling for efficient beam-search
Marcello Federico, Mauro Cettolo, Fabio Brugnara, Giuliano Antoniol
Comput. Speech Lang.1
1994 Radiological reporting by speech recognition: the a.re.s. system
abstract
Radiological reporting has already been identified as a field in which voice technologies can prove to be very useful. Recent progress in automatic speech recognition and in hardware and software technology makes it possible to build large-vocabulary, continuous speech, speaker-independent, real-time systems. In this paper a dictation system for radiology reporting, the A.Re.S. system, is presented. A.Re.S. is a "software only" system which runs in real-time on an HP 715 workstation. It relies on an asynchronous and multi-process architecture in which speech decoding is performed by processes in pipeline. System requirements and architecture will be described, together with the results of a preliminary evaluation based on three months of on-site testing. I. INTRODUCTION Recent progress in Automatic Speech Recognition (ASR) and in hardware and software technology makes it possible to build large-vocabulary, real-time, speaker-independent systems. Medical document generation presents f...
Bianca Angelini, Giuliano Antoniol, Fabio Brugnara, Mauro Cettolo, Marcello Federico, Roberto Fiutem, Gianni Lazzari
ICSLP5
1994 Language model estimations and representations for real-time continuous speech recognition
abstract
This paper compares different ways of estimating bigram language models and of representing them in a finite state network used by a beam-search based, continuous speech, and speaker independent HMM recognizer. Attention is focused on the n-gram interpolation scheme for which seven models are considered. Among them, the Stacked estimated linear interpolated model favourably compares with the best known ones. Further, two different static representations of the search space are investigated: "linear" and "tree-based". Results show that the latter topology is better suited to the beam-search algorithm. Moreover, this representation can be reduced by a network optimization technique, which allows the dynamic size of the recognition process to be decreased by 60%. Extensive recognition experiments on a 10,000-word dictation task with four speakers are described in which an average word accuracy of 93% is achieved with real-time response. I. INTRODUCTION This paper compares different ways ...
Giuliano Antoniol, Fabio Brugnara, Mauro Cettolo, Marcello Federico
ICSLP4
1993 Techniques for robust recognition in restricted domains
abstract
This paper describes an Automatic Speech Understanding (ASU) system used in a human-robot interface for the remote control of a mobile robot. The intended application is that of an operator issuing telecontrol commands to one or more robots from a remote workstation. ASU is supposed to be performed with spontaneous continuous speech and quasi real time conditions. Training and testing of the system was based on speech data collected by means of Wizard of Oz simulations. Two kinds of robustness factors are introduced: the first is a recognition error-tolerant approach to semantic interpretation, the second is based on a technique for evaluating the reliability of the ASU system output with respect to the input utterance. Preliminary results are 90.9% of correct semantic interpretations, and 89.1% of correct detection of out-of-domain sentences at the cost of rejecting 16.4% of correct in-domain sentences. 1. INTRODUCTION This paper describes an Automatic Speech Understanding (ASU) sys...
Giuliano Antoniol, Mauro Cettolo, Marcello Federico
EUROSPEECH3