EDBT 2026 Demo / reviewers in the wild / expert
Daniele Falavigna
dblp:15/5847 · also Daniele Giuseppe Falavigna, Giuseppe Daniele Falavigna
· DBLP profile ↗
63ranked-venue papers
10as first author
10since 2021 · last 2025
0000-0002-4844-5071ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 50 · 9 first-author · 9 since 2021Artificial intelligence and machine learning · 49 · 10 first-author · 7 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EFL-PEFT: A communication Efficient Federated Learning framework using PEFT sparsification for ASRabstractFederated Learning (FL) has garnered substantial interest in training different speech-based tasks (e.g. automatic speech recognition (ASR), and other speech classification tasks): recently, fine-tuning pre-trained self-supervised models for different speech-based tasks has shown promising performance and been successfully applied in FL settings. Nevertheless, fine-tuning these architectures is computationally burdensome and not affordable in several real-time settings. Moreover, the communication costs of transferring all the model parameters for the aggregation stage is critically high. As an alternative approach, parameter-efficient fine-tuning (PEFT) approaches provide promising performance without changing the backbone of the pre-trained model. PEFT has been fruitfully applied, in a variety of flavours, for ASR in central training configurations while only few works investigate its use in FL settings. In this paper, we consolidate the use of PEFT for ASR with pre-trained models, demonstrating that it enables efficient FL reducing the amount of parameters to share with respect to full fine-tuning. We also explore combining PEFT with sparsification methods to further reduce communication cost by transmitting only a fraction of the adapter parameters. Additionally, we show that agglomerating adapters using "FedAvg" is compatible with differential privacy, aligning with trends observed in other domains. Our proposed approach is supported by experimental analysis on ASR using two public datasets, as well as on intent classification tasks. Mohamed Nabih Ali, Daniele Falavigna, Alessio Brutti |
ICASSP | 2 |
| 2025 | Large Language Models are Strong Audio-Visual Speech Recognition LearnersabstractMultimodal large language models (MLLMs) have recently become a focal point of research due to their formidable multimodal understanding capabilities. For example, in the audio and speech domains, an LLM can be equipped with (automatic) speech recognition (ASR) abilities by just concatenating the audio tokens, computed with an audio encoder, and the text tokens to achieve state-of-the-art results. On the contrary, tasks like visual and audio-visual speech recognition (VSR/AVSR), which also exploit noise-invariant lip movement information, have received little or no attention. To bridge this gap, we propose Llama-AVSR, a new MLLM with strong audio-visual speech recognition capabilities. It leverages pre-trained audio and video encoders to produce modality-specific tokens which, together with the text tokens, are processed by a pre-trained LLM (e.g., Llama3.1-8B) to yield the resulting response in an auto-regressive fashion. Llama-AVSR requires a small number of trainable parameters as only modality-specific projectors and LoRA modules are trained whereas the multi-modal encoders and LLM are kept frozen. We evaluate our proposed approach on LRS3, the largest public AVSR benchmark, and we achieve new state-of-the-art results for the tasks of ASR and AVSR with a WER of 0.79% and 0.77%, respectively. To bolster our results, we investigate the key factors that underpin the effectiveness of Llama-AVSR: the choice of the pre-trained encoders and LLM, the efficient integration of LoRA modules, and the optimal performance-efficiency trade-off obtained via modality-aware compression rates. Umberto Cappellazzo, Honglie Chen, Pingchuan Ma 0001, Stavros Petridis, Daniele Falavigna, Alessio Brutti, Maja Pantic |
ICASSP | 6 |
| 2025 | Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach
Umberto Cappellazzo, Stavros Petridis, Daniele Falavigna, Alessio Brutti |
INTERSPEECH | 4 |
| 2024 | Efficient Fine-tuning of Audio Spectrogram Transformers via Soft Mixture of Adapters
Umberto Cappellazzo, Daniele Falavigna, Alessio Brutti |
INTERSPEECH | 2 |
| 2023 | An Investigation of the Combination of Rehearsal and Knowledge Distillation in Continual Learning for Spoken Language Understanding
Umberto Cappellazzo, Daniele Falavigna, Alessio Brutti |
INTERSPEECH | 2 |
| 2023 | Sequence-Level Knowledge Distillation for Class-Incremental End-to-End Spoken Language Understanding
Umberto Cappellazzo, Muqiao Yang, Daniele Falavigna, Alessio Brutti |
INTERSPEECH | 3 |
| 2023 | Direct enhancement of pre-trained speech embeddings for speech processing in noisy conditions
Mohamed Nabih Ali, Alessio Brutti, Daniele Falavigna |
Comput. Speech Lang. | 3 |
| 2022 | End-to-End Low Resource Keyword Spotting Through Character Recognition and Beam-Search Re-ScoringabstractThis paper describes an end-to-end approach to perform keyword spotting with a pre-trained acoustic model that uses recurrent neural networks and connectionist temporal classification loss. Our approach is specifically designed for low-resource keyword spotting tasks where extremely small amounts of in-domain data are available to train the system. The pre-trained model, largely used in ASR tasks, is fine-tuned on in-domain audio recordings. In inference the model output is matched against the set of predefined keywords using a beam-search re-scoring based on the edit distance.We demonstrate that this approach significantly outperforms the best state-of-the art systems on a well known keyword spotting benchmark, namely "google speech commands". Moreover, com-pared against state-of-the-art methods, our proposed approach is extremely robust in case of limited in domain training material. We show that a very small performance reduction is observed when fine tuning with a very small fraction (around 5%) of the training set.We report an extensive set of experiments on two keyword spotting tasks, varying training sizes and correlating keyword classification accuracy with character error rates provided by the system. We also report an ablation study to assess on the contribution of the out-of-domain pre-training and of the beam-search re-scoring. Ephrem Tibebe Mekonnen, Alessio Brutti, Daniele Falavigna |
ICASSP | 3 |
| 2022 | Enhancing Embeddings for Speech Classification in Noisy Conditions
Mohamed Nabih Ali, Alessio Brutti, Daniele Falavigna |
INTERSPEECH | 3 |
| 2021 | ETLT 2021: Shared Task on Automatic Speech Recognition for Non-Native Children's Speech
Roberto Gretter, Marco Matassoni, Daniele Falavigna, A. Misra, Chee Wee Leong, Kate M. Knill |
Interspeech | 3 |
| 2020 | Overview of the Interspeech TLT2020 Shared Task on ASR for Non-Native Children's Speech
Roberto Gretter, Marco Matassoni, Daniele Falavigna, Keelan Evanini, Chee Wee Leong |
INTERSPEECH | 3 |
| 2020 | Mixtures of Deep Neural Experts for Automated Speech ScoringabstractThe paper copes with the task of automatic assessment of second language proficiency from the language learners' spoken responses to test prompts. The task has significant relevance to the field of computer assisted language learning. The approach presented in the paper relies on two separate modules: (1) an automatic speech recognition system that yields text transcripts of the spoken interactions involved, and (2) a multiple classifier system based on deep learners that ranks the transcripts into proficiency classes. Different deep neural network architectures (both feed-forward and recurrent) are specialized over diverse representations of the texts in terms of: a reference grammar, the outcome of probabilistic language models, several word embeddings, and two bag-of-word models. Combination of the individual classifiers is realized either via a probabilistic pseudo-joint model, or via a neural mixture of experts. Using the data of the third Spoken CALL Shared Task challenge, the highest values to date were obtained in terms of three popular evaluation metrics. Sara Papi, Edmondo Trentin, Roberto Gretter, Marco Matassoni, Daniele Falavigna |
INTERSPEECH | 5 |
| 2020 | TLT-school: a Corpus of Non Native Children SpeechabstractThis paper describes “TLT-school” a corpus of speech utterances collected in schools of northern Italy for assessing the performance of students learning both English and German. The corpus was recorded in the years 2017 and 2018 from students aged between nine and sixteen years, attending primary, middle and high school. All utterances have been scored, in terms of some predefined proficiency indicators, by human experts. In addition, most of utterances recorded in 2017 have been manually transcribed carefully. Guidelines and procedures used for manual transcriptions of utterances will be described in detail, as well as results achieved by means of an automatic speech recognition system developed by us. Part of the corpus is going to be freely distributed to scientific community particularly interested both in non-native speech recognition and automatic assessment of second language proficiency. Roberto Gretter, Marco Matassoni, Stefano Bannò, Daniele Falavigna |
LREC | 4 |
| 2019 | Automatic Assessment of Spoken Language Proficiency of Non-native ChildrenabstractThis paper describes technology developed to automatically grade Italian students (ages 9-16) on their English and German spoken language proficiency. The students' spoken answers are first transcribed by an automatic speech recognition (ASR) system and then scored using a feedforward neural network (NN) that processes features extracted from the automatic transcriptions. In-domain acoustic models, employing deep neural networks (DNNs), are derived by adapting the parameters of an original out of domain DNN. Automatic scores are computed for low level proficiency indicators - such as: lexical richness, syntax correctness, quality of pronunciation, discourse fluency, semantic relevance to the prompt, etc - defined by human experts in language proficiency. A set of experiments was carried out on a large set of data collected during proficiency evaluation campaigns involving thousands of students, manually scored by human experts. Obtained results are presented and discussed. Roberto Gretter, Marco Matassoni, Katharina Allgaier, Svetlana Tchistiakova, Daniele Falavigna |
ICASSP | 5 |
| 2018 | Non-Native Children Speech Recognition Through Transfer LearningabstractThis work deals with non-native children's speech and investigates both multi-task and transfer learning approaches to adapt a multi-language Deep Neural Network (DNN) to speakers, specifically children, learning a foreign language. The application scenario is characterized by young students learning English and German and reading sentences in these second-languages, as well as in their mother language. The paper analyzes and discusses techniques for training effective DNN-based acoustic models starting from children's native speech and performing adaptation with limited non-native audio material. A multi -lingual model is adopted as baseline, where a common phonetic lexicon, defined in terms of the units of the International Phonetic Alphabet (IPA), is shared across the three languages at hand (Italian, German and English); DNN adaptation methods based on transfer learning are evaluated on significant non-native evaluation sets. Results show that the resulting non-native models allow a significant improvement with respect to a mono-lingual system adapted to speakers of the target language. Marco Matassoni, Roberto Gretter, Daniele Falavigna, Diego Giuliani |
ICASSP | 3 |
| 2018 | Automatic quality estimation for ASR system combination
Shahab Jalalvand, Matteo Negri, Daniele Falavigna, Marco Matassoni, Marco Turchi |
Comput. Speech Lang. | 3 |
| 2017 | Optimizing DNN Adaptation for Recognition of Enhanced Speech
Marco Matassoni, Alessio Brutti, Daniele Falavigna |
INTERSPEECH | 3 |
| 2017 | DNN adaptation by automatic quality estimation of ASR hypotheses
Daniele Falavigna, Marco Matassoni, Shahab Jalalvand, Matteo Negri, Marco Turchi |
Comput. Speech Lang. | 1 |
| 2016 | DNN adaptation for recognition of children speech through automatic utterance selectionabstractThis paper describes an approach for adapting a DNN trained on adult speech to children voices. The method extends a previous one, based on the Kullback-Leibler divergence between the original (adult) DNN output distribution and the target one, by accounting for the quality of the supervision of the adaptation utterances. In addition, starting from the observation that by gradually removing from the adaptation set the sentences with higher WERs significant performance improvements can be achieved, we also investigate the usage of automatic selection of adaptation utterances. For determining transcription quality we investigate the use of confidence estimates of recognized hypotheses. We present experiments and related results achieved on an Italian data set of children's speech. We show that the proposed DNN adaptation approach allows to significantly reduce the WER on a given test set from 14.2% (corresponding to using the non adapted DNN, trained on adult speech) to 10.6%. It is worth mentioning that the latter result has been achieved without making use of any training data specific of children's speech. Marco Matassoni, Daniele Falavigna, Diego Giuliani |
SLT | 2 |
| 2015 | Driving ROVER with Segment-based ASR Quality EstimationabstractShahab Jalalvand, Matteo Negri, Daniele Falavigna, Marco Turchi. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Shahab Jalalvand, Matteo Negri, Daniele Falavigna, Marco Turchi |
ACL (1) | 3 |
| 2015 | Boosted acoustic model learning and hypotheses rescoring on the CHiME-3 taskabstractSpeech recognition in a realistic noisy environment using multiple microphones is the focal point of the third CHiME challenge. Over the baseline ASR system provided for this challenge, we apply state of the art algorithms for boosting acoustic model learning and hypothesis rescoring to improve the final output. To this aim, we first use the automatic transcription of each channel to re-train the acoustic model for that channel and then we apply linear language model rescoring to find a better solution in the n-best list. LM rescoring is performed using an efficient set of N-gram and Recurrent Neural Network LM (RNNLM) trained on a wisely-selected text set. In the experiments, we show that the proposed approach improves not only the individual channel transcription, but also the enhanced channels produced by MVDR and delay-and-sum beamforming. Shahab Jalalvand, Daniele Falavigna, Marco Matassoni, Piergiorgio Svaizer, Maurizio Omologo |
ASRU | 2 |
| 2015 | Stacked auto-encoder for ASR error detection and word error rate prediction
Shahab Jalalvand, Daniele Falavigna |
INTERSPEECH | 2 |
| 2015 | Multitask Learning for Adaptive Quality Estimation of Automatically Transcribed UtterancesabstractJosé G. C. de Souza, Hamed Zamani, Matteo Negri, Marco Turchi, Daniele Falavigna. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. José Guilherme Camargo de Souza, Hamed Zamani, Matteo Negri, Marco Turchi, Daniele Falavigna |
HLT-NAACL | 5 |
| 2014 | Quality Estimation for Automatic Speech Recognition
Matteo Negri, Marco Turchi, José Guilherme Camargo de Souza, Daniele Falavigna |
COLING | 4 |
| 2014 | Direct word graph rescoring using a* search and RNNLM
Shahab Jalalvand, Daniele Falavigna |
INTERSPEECH | 2 |
| 2013 | Phonetic and anthropometric conditioning of MSA-KST cognitive impairment characterization systemabstractWe explore the impact of speech- and speaker-specific modeling onto the Modulation Spectrum Analysis - Kolmogorov-Smirnov feature Testing (MSA-KST) characterization method in the task of automated prediction of the cognitive impairment diagnosis, namely dysphasia and pervasive development disorder. Phoneme-synchronous capturing of speech dynamics is a reasonable choice for a segmental speech characterization system as it allows comparing speech dynamics in the similar phonetic contexts. Speaker-specific modeling aims at reducing the “within-the-class” variability of the characterized speech or speaker population by removing the effect of speaker properties that should have no relation to the characterization. Specifically the vocal tract length of a speaker has nothing to do with the diagnosis attribution and, thus, the feature set shall be normalized accordingly. The resulting system compares favorably to the baseline system of the Interspeech'2013 Computational Paralinguistics Challenge. Alexei V. Ivanov, Shahab Jalalvand, Roberto Gretter, Daniele Falavigna |
ASRU | 4 |
| 2012 | Analysis of the Characteristics of Talk-show TV Programs
Fabio Brugnara, Daniele Falavigna, Diego Giuliani, Roberto Gretter |
INTERSPEECH | 2 |
| 2011 | Redundancy Reduction in ASR of Spontaneous Speech Through Statistical Machine Translation
Daniele Falavigna |
INTERSPEECH | 1 |
| 2011 | Cheap Bootstrap of Multi-Lingual Hidden Markov Models
Daniele Falavigna, Roberto Gretter |
INTERSPEECH | 1 |
| 2011 | NeMo: A Platform for Multilingual News Monitoring
Christian Girardi, Roberto Gretter, Daniele Falavigna, Fabio Brugnara, Diego Giuliani, Marcello Federico |
INTERSPEECH | 3 |
| 2010 | Evaluation of automatic transcription systems for the judicial domainabstractThis paper describes two different automatic transcription systems developed for judicial application domains for the Polish and Italian languages. The judicial domain requires to cope with several factors which are known to be critical for automatic speech recognition, such as: background noise, reverberation, spontaneous and accented speech, overlapped speech, cross channel effects, etc. The two automatic speech recognition (ASR) systems have been developed independently starting from out-of-domain data and, then, they have been adapted using a certain amount of in-domain audio and text data. The ASR performance have been measured on audio data acquired in the courtrooms of Naples and Wroclaw. The resulting word error rates are around 40%, for Italian, and around between 30% and 50% for Polish. This performance, similar to that reported for other comparable ASR tasks (e.g. meeting transcriptions with distant microphone), suggests that possible applications can address tasks such as indexing and/or information retrieval in multimedia documents recorded during judicial debates. Jonas Lööf, Daniele Falavigna, Ralf Schlüter, Diego Giuliani, Roberto Gretter, Hermann Ney |
SLT | 2 |
| 2009 | Phone-to-word decoding through statistical machine translation and complementary system combinationabstractIn this paper, phone-to-word transduction is first investigated by coupling a speech recognizer, generating for each speech segment a phone sequence or a phone confusion network, with the efficient decoder of confusion networks adopted by MOSES, a popular statistical machine translation toolkit. Then, system combination is investigated by combining the outputs of several conventional ASR systems with the output of a system embedding phone-to-word decoding through statistical machine translation. Experiments are carried out in the context of a large vocabulary speech recognition task consisting of transcription of speeches delivered in English during the European Parliament Plenary Sessions (EPPS). While only a marginal performance improvements is achieved in system combination experiments when the output of the phone-to-word transducer is included in the combination, partial results show a great potential for improvements. Daniele Falavigna, Matteo Gerosa, Roberto Gretter, Diego Giuliani |
ASRU | 1 |
| 2008 | Fast speech decoding through phone confusion networksabstractWe present a two stage automatic speech recognition architecture suited for applications, such as spoken document retrieval, where large scale language models can be used and very low out-of-vocabulary rates need to be reached. The proposed system couples a weakly constrained phone-recognizer with a phone-to-word decoder that was originally developed for phrase-based statistical machine translation. The decoder permits to efficiently decode confusion networks in input, and to exploit large scale unpruned language models. Preliminary experiments are reported on the transcription of speeches of the Italian parliament. The use of phone confusion networks as interface between the two decoding steps permits to reduce the WER by 28%, thus making the system perform relatively close to a state-of-the-art baseline using a comparable language model. Nicola Bertoldi, Marcello Federico, Daniele Falavigna, Matteo Gerosa |
INTERSPEECH | 3 |
| 2007 | The IRST English-Spanish translation system for european parliament speeches
Daniele Falavigna, Nicola Bertoldi, Fabio Brugnara, Roldano Cattoni, Mauro Cettolo, Boxing Chen, Marcello Federico, Diego Giuliani, Roberto Gretter, Dino Seppi |
INTERSPEECH | 1 |
| 2007 | Word duration modeling for word graph rescoring in LVCSR
Dino Seppi, Daniele Falavigna, Georg Stemmer, Roberto Gretter |
INTERSPEECH | 2 |
| 2006 | Design and evaluation of acoustic and language models for large scale telephone services
Andrea Facco, Daniele Falavigna, Roberto Gretter, Marcello Viganò |
Speech Commun. | 2 |
| 2005 | A frame based spoken dialog system for home care
Daniele Falavigna, Toni Giorgino, Roberto Gretter |
INTERSPEECH | 1 |
| 2004 | A multi-modal architecture for cellular phones
Luca Nardelli, Marco Orlandi, Daniele Falavigna |
ICMI | 3 |
| 2004 | On the development of telephone applications: some practical issues and evaluation
Andrea Facco, Daniele Falavigna, Roberto Gretter, Marcello Viganò |
INTERSPEECH | 2 |
| 2003 | On the usage of automatic voice recognition in an interactive Web based medical applicationabstractWe describe the multi-modal browsing system, developed by us, that allows to add automatic speech recognition and text to speech functions to standard Internet browsers. The system is based on the temporal synchronization of HTML and VoiceXML documents. It was developed starting from a real Web application designed for a medical domain (i.e. an electronic patient record adopted in the oncology unit of an Italian hospital). We have introduced the possibility to define the multi-modal interaction by means of a single XML document. System evaluation is being carried out on data collected during the usage of the system in the hospital. Claudio Eccher, Lorenzo Eccher, Daniele Falavigna, Luca Nardelli, Marco Orlandi, Andrea Sboner |
ICASSP (2) | 3 |
| 2003 | Maximum likelihood endpoint detection with time-domain featuresabstractIn this paper we propose an effective, robust and computationally low-cost HMM-based start-endpoint detector for speech recognisers . Our first attempts follow the classical scheme feature extractor-Viterbi classifier (used for voice activity detection) , followed by a post-processing stage, but the ultimate goal we pursue is a pure HMM-based architecture capable of performing the endpointing task. The features used for voice activity detection are energy and zero crossing rate, together with AMDF (Average Magnitude Difference Function), which proves to be a valid alternative to energy; further, we study the impact on performance of grammar structures and training conditions. In the end, we set the basis for the investigation of pure HMM-based architectures. Marco Orlandi, Alfiero Santarelli, Daniele Falavigna |
INTERSPEECH | 3 |
| 2002 | Home Monitoring of Hypertensive Patients through Intelligent Dialog System
Ivano Azzini, Daniele Falavigna, Toni Giorgino, Roberto Gretter, Silvana Quaglini, Carla Rognoni, Mario Stefanelli |
AMIA | 2 |
| 2002 | Acoustic and word lattice based algorithms for confidence scoresabstractWord confidence scores are crucial for unsupervised learning in automatic speech recognition. In the last decade there has been a flourish of work on two fundamentally different approaches to compute confidence scores. The first paradigm is acoustic and the second is based on word lattices. The first approach is dataintensive and it requires to explicitly model the acoustic channel. The second approach is suitable for on-line (unsupervised) learning and requires no training. In this paper we present a comparative analysis of off-the-shelf and new algorithms for computing confidence scores, following the acoustic and lattice-based paradigms. We compare the performance of these algorithms across three tasks for small, medium and large vocabulary speech recognition tasks and for two languages (Italian and English). We show that wordlattice based algorithm provides consistent and effective performance across automatic speech recognition tasks. 1. Daniele Falavigna, Roberto Gretter, Giuseppe Riccardi |
INTERSPEECH | 1 |
| 2002 | The C-ORAL-ROM Project. New methods for spoken language archives in a multilingual romance corpus
Emanuela Cresti, Massimo Moneglia, Fernanda Bacelar do Nascimento, Antonio Moreno-Sandoval, Jean Véronis, Philippe Martin 0003, Khalid Choukri, Valérie Mapelli, Daniele Falavigna, Antonio Cid, Claude Blum |
LREC | 9 |
| 2001 | First steps toward an adaptive spoken dialogue system in medical domainabstractRecently ITC-irst (Interactive Sensory System division) and University of Pavia (Medical Informatics Labs) are working together to realize an intelligent (adaptive) dialogue system with language understanding capabilities. In this framework, some telemedicine services able to handle multimodal interactions (i.e. input/output can be provided by both voice and/or other devices such as: keyboard, mouse, etc ) are going to be investigated and developed. Although the system is placed in medical domains, the basic concepts can be transferred towards other applications. In this papers we define the problem, explain the basic ideas, some theoretical foundation and report the early steps we have done to achieve the goal. Ivano Azzini, Daniele Falavigna, Roberto Gretter, Giordano Lanzola, Marco Orlandi |
INTERSPEECH | 2 |
| 2000 | A mixed language model for a dialogue system over ihe telephone
Daniele Falavigna, Roberto Gretter, Marco Orlandi |
INTERSPEECH | 1 |
| 1999 | Semantic boundaries in multiple languagesabstractThis paper presents the results obtained for the task of detecting Semantic Boundaries (SBs) in spoken language using two different methods on the same data set. Hence we first introduce the two approaches developed by ITC-Irst in Trento (Italy) and the LME of the University Erlangen (Germany) and discuss the individually obtained results. The basis for the decision upon SBs in both cases are textual and prosodic features. The LME has already worked for several years on the computation and application of prosodic features in automatic speech processing within the Verbmobil project. The approaches developed in that project were adapted to work on the data collected at IRST in the Italian language. Finally we compare the results we obtain with the German SB detection against the Italian result with regard to precision and recall. 1. INTRODUCTION For robust spoken language processing it is not always necessary to analyse a user's utterance completely as one coherent segment. Often i... Volker Warnke, Heinrich Niemann, Mauro Cettolo, Anna Corazza, Daniele Falavigna, Gianni Lazzari |
EUROSPEECH | 6 |
| 1999 | Use of simulated data for robust telephone speech recognition
Tarcisio Coianiz, Daniele Falavigna, Roberto Gretter, Marco Orlandi |
EUROSPEECH | 2 |
| 1998 | Automatic recognition of spontaneous speech dialoguesabstractSome approaches for coping with the problem of recognition of spontaneous speech dialogues are presented. Starting from a HMM-based system, developed for dictation tasks, some modifications are introduced at the acoustic level and in the language model. Acoustic model parameters are modified to account for speaking rate variations, and specific models of extra-linguistic phenomena are defined and added to the language model. Different acoustic models are managed by a single recognizer through their integration into a multi-model search space. Experiments and evaluations were conducted on a spontaneous dialogue corpus collected at our laboratory. 1. INTRODUCTION This paper concerns the assessment of a set of modifications applied to a HMM-based continuous speech recognizer developed at ITC-Irst, for dictation tasks, to improve its performance on a corpus of spontaneously uttered humanhuman dialogues. As the performance of Automatic Speech Recognition (ASR) systems is largely affected ... Mauro Cettolo, Daniele Falavigna |
ICSLP | 2 |
| 1998 | Automatic detection of semantic boundaries based on acoustic and lexical knowledgeabstractIn spoken language systems, the segmentation of utterances into coherent linguistic/semantic units is required when modules following the speech recognizer can only process such units one at a time. In this paper, techniques for semantic boundary prediction, based on both acoustic and lexical knowledge, are presented and tested on a corpus of personto -person dialogues. Best result gives 62.8% recall and 71.8% precision. 1. INTRODUCTION In spoken language systems, the minimal unit of analysis does not necessarily correspond to a full sentence. A possible approach for language processing is that of splitting a given sentence in a sequence of units that can be successively processed by linguistic modules one at a time. The goal of the Semantic Boundary (SB) detector is to locate boundaries inside a sentence in order to obtain such "minimal units". Useful information for SB detection can be extracted either from the waveform of an utterance or from its corresponding word sequence. Some ... Mauro Cettolo, Daniele Falavigna |
ICSLP | 2 |
| 1997 | Multilingual person to person communication at IRSTabstractThis paper refers to a machine-mediated person-to-person multilingual communication system. Stress is put on robustness, that is the ability of the system to preserve communication even in presence of the variability and errors typical of spoken language systems. The statistical approach is adopted not only at the acoustic level, but also for the linguistic processing. Therefore, while an overview of the global architecture is briefly introduced, the focus is put on the acoustic recognizer and the understanding module. Experimental evaluations complete the presentation. Bianca Angelini, Mauro Cettolo, Anna Corazza, Daniele Falavigna, Gianni Lazzari |
ICASSP | 4 |
| 1997 | Automatic diphone extraction for an Italian text-to-speech synthesis system
Bianca Angelini, Claudia Barolo, Daniele Falavigna, Maurizio Omologo, Stefano Sandri |
EUROSPEECH | 3 |
| 1997 | On field experiments of continuous digit recognition over the telephone networkabstractIn this paper a continuous digit recognizer over the telephone network in real time will be described. The activity has allowed the realization of a system, installed in some Italian telephone exchanges, for providing semi-automatic collect call services. Data collection has also been performed, and a field database was built. Either a continuous digit recognition task and a confirmation task, requiring rejection, have been defined. Recognition results are presented. INTRODUCTION The activity reported in this paper led to the realization of a system, installed in some Italian telephone exchanges. It provides two semi-automatic collect call services, called "Italy Direct" and "170". These systems require the recognition of digit sequences, as well as of yes/no. In the last case rejection of unforeseen sentences must be used to assure sufficient robustness with respect to user inexperience. To train and test the system some telephone speech databases, later described, have been used. I... Daniele Falavigna, Roberto Gretter |
EUROSPEECH | 1 |
| 1995 | Comparison of different HMM based methods for speaker verificationabstractThree different speaker verification methods are described. All of them are based on Hidden Markov Models (HMM); the first one is of type text independent the other two are of type text prompted. The text independent method makes use of a single state Continuous HMM, to represent each customer in the system, while the text prompted methods require to use speaker dependent phoneme models. To model phonemes both Continuous HMMs and SemiContinuous HMMs are used. Two different normalization methods for the likelihood values provided by the various HMMs are considered: one is based on the posterior probability, the other is based on the application of a mapping function. 1. INTRODUCTION In the paper three different approaches, based on the Hidden Markov Model (HMM) technology, will be described and compared for a speaker verification task. The task consists in accepting or rejecting the identity of a claimed speaker according to the individual information contained in a given input uttera... Daniele Falavigna |
EUROSPEECH | 1 |
| 1995 | Application of Generalized Radial Basis Functions In Speaker Normalization and Identification
Cesare Furlanello, Diego Giuliani, Edmondo Trentin, Daniele Falavigna |
ISCAS | 4 |
| 1995 | Automatic person recognition by acoustic and geometric features
Roberto Brunelli, Daniele Falavigna, Tomaso A. Poggio, Luigi Stringa |
Mach. Vis. Appl. | 2 |
| 1995 | Person Identification Using Multiple CuesabstractThis paper presents a person identification system based on acoustic and visual features. The system is organized as a set of non-homogeneous classifiers whose outputs are integrated after a normalization step. In particular, two classifiers based on acoustic features and three based on visual ones provide data for an integration module whose performance is evaluated. A novel technique for the integration of multiple classifiers at an hybrid rank/measurement level is introduced using HyperBF networks. Two different methods for the rejection of an unknown person are introduced. The performance of the integrated system is shown to be superior to that of the acoustic and visual subsystems. The resulting identification system can be used to log personal access and, with minor modifications, as an identity verification system.> Roberto Brunelli, Daniele Falavigna |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1994 | Speaker independent continuous speech recognition using an acoustic-phonetic Italian corpusabstractThe objective of this paper is to describe the activity that is being carried out at IRST laboratories for the development of an HMM-based speaker independent continuous speech recognition system for the Italian language. The recognition system is trained and tested using the acoustic-phonetic continuous speech portion of the APASCI corpus. Acoustic modeling is based on the use of Continuous Density HMMs with gaussian mixture observation densities. As a baseline, a set of 38 Context Independent Units was evaluated using different numbers of mixture components. Then, two other classes of Context Dependent Unit sets were considered, that provide different performance and system complexity. Performance, expressed in terms of Phone loop recognition accuracy and Word loop recognition accuracy, shows an improvement using both of these classes of unit sets, with respect to the baseline. I. INTRODUCTION A baseline of a speaker independent continuous speech recognition system for the Italian ... Bianca Angelini, Fabio Brugnara, Daniele Falavigna, Diego Giuliani, Roberto Gretter, Maurizio Omologo |
ICSLP | 3 |
| 1993 | Automatic segmentation and labeling of English and Italian speech databases
Bianca Angelini, Fabio Brugnara, Daniele Falavigna, Diego Giuliani, Roberto Gretter, Maurizio Omologo |
EUROSPEECH | 3 |
| 1993 | A baseline of a speaker independent continuous speech recognizer of Italian
Bianca Angelini, Fabio Brugnara, Daniele Falavigna, Diego Giuliani, Roberto Gretter, Maurizio Omologo |
EUROSPEECH | 3 |
| 1993 | Automatic segmentation and labeling of speech based on Hidden Markov Models
Fabio Brugnara, Daniele Falavigna, Maurizio Omologo |
Speech Commun. | 2 |
| 1992 | A HMM-based system for automatic segmentation and labeling of speech
Fabio Brugnara, Daniele Falavigna, Maurizio Omologo |
ICSLP | 2 |
| 1991 | A preliminary statistical evaluation of manual and automatic segmentation discrepancies
Piero Cosi, Daniele Falavigna, Maurizio Omologo |
EUROSPEECH | 2 |