Marco Matassoni

dblp:16/6953 · DBLP profile ↗
← Back
50ranked-venue papers
9as first author
8since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 40 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 30 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Phonetic-based Ranking for Improved Pseudo-Labeling in Low-Resource ASR
Marco Matassoni, Roberto Gretter, Falavigna Daniele, Mohamed Nabih Ali, Alessio Brutti, Matteo Negri, Mauro Cettolo, Marco Gaido, Sara Papi, Luisa Bentivogli
LREC1
2025 Speech LLMs in Low-Resource Scenarios: Data Volume Requirements and the Impact of Pretraining on High-Resource Languages
abstract
Large language models (LLMs) have demonstrated potential in handling spoken inputs for high-resource languages, reaching state-of-the-art performance in various tasks.However, their applicability is still less explored in low-resource settings.This work investigates the use of Speech LLMs for lowresource Automatic Speech Recognition using the SLAM-ASR framework, where a trainable lightweight projector connects a speech encoder and a LLM.Firstly, we assess training data volume requirements to match Whisper-only performance, reemphasizing the challenges of limited data.Secondly, we show that leveraging mono-or multilingual projectors pretrained on high-resource languages reduces the impact of data scarcity, especially with small training sets.Using multilingual LLMs (EuroLLM, Salamandra) with whisper-large-v3-turbo, we evaluate performance on several public benchmarks, providing insights for future research on optimizing Speech LLMs for lowresource languages and multilinguality.
Seraphina Fong, Marco Matassoni, Alessio Brutti
INTERSPEECH2
2025 Automatic detection of speech sound disorders in German-speaking children: augmenting the data with typically developed speech
Darline Monika Marx, Marco Matassoni, Alessio Brutti
INTERSPEECH2
2024 MOSEL: 950, 000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages
abstract
Marco Gaido, Sara Papi, Luisa Bentivogli, Alessio Brutti, Mauro Cettolo, Roberto Gretter, Marco Matassoni, Mohamed Nabih, Matteo Negri. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Marco Gaido, Sara Papi, Luisa Bentivogli, Alessio Brutti, Mauro Cettolo, Roberto Gretter, Marco Matassoni, Mohamed Nabih Ali, Matteo Negri
EMNLP7
2024 Back to grammar: Using grammatical error correction to automatically assess L2 speaking proficiency
Stefano Bannò, Marco Matassoni
Speech Commun.2
2022 Proficiency Assessment of L2 Spoken English Using Wav2Vec 2.0
abstract
The increasing demand for learning English as a second language has led to a growing interest in methods for automatically assessing spoken language proficiency. Most approaches use hand-crafted features, but their efficacy relies on their particular underlying assumptions and they risk discarding potentially salient information about proficiency. Other approaches rely on transcriptions produced by ASR systems which may not provide a faithful rendition of a learner's utterance in specific scenarios (e.g., non-native children's spontaneous speech). Furthermore, transcriptions do not yield any information about relevant aspects such as intonation, rhythm or prosody. In this paper, we investigate the use of wav2vec 2.0 for assessing overall and individual aspects of proficiency on two small datasets, one of which is publicly available. We find that this approach significantly outperforms the BERT-based baseline system trained on ASR and manual transcriptions used for comparison.
Stefano Bannò, Marco Matassoni
SLT2
2021 Learning to Rank Microphones for Distant Speech Recognition
abstract
Fully exploiting ad-hoc microphone networks for distant speech recognition is still an open issue. Empirical evidence shows that being able to select the best microphone leads to significant improvements in recognition without any additional effort on front-end processing. Current channel selection techniques either rely on signal, decoder or posterior-based features. Signal-based features are inexpensive to compute but do not always correlate with recognition performance. Instead decoder and posterior-based features exhibit better correlation but require substantial computational resources. In this work, we tackle the channel selection problem by proposing MicRank, a learning to rank framework where a neural network is trained to rank the available channels using directly the recognition performance on the training set. The proposed approach is agnostic with respect to the array geometry and type of recognition back-end. We investigate different learning to rank strategies using a synthetic dataset developed on purpose and the CHiME-6 data. Results show that the proposed approach is able to considerably improve over previous selection techniques, reaching comparable and in some instances better performance than oracle signal-based measures.
Samuele Cornell, Alessio Brutti, Marco Matassoni, Stefano Squartini
Interspeech3
2021 ETLT 2021: Shared Task on Automatic Speech Recognition for Non-Native Children's Speech
Roberto Gretter, Marco Matassoni, Daniele Falavigna, A. Misra, Chee Wee Leong, Kate M. Knill
Interspeech2
2020 Overview of the Interspeech TLT2020 Shared Task on ASR for Non-Native Children's Speech
Roberto Gretter, Marco Matassoni, Daniele Falavigna, Keelan Evanini, Chee Wee Leong
INTERSPEECH2
2020 Mixtures of Deep Neural Experts for Automated Speech Scoring
abstract
The paper copes with the task of automatic assessment of second language proficiency from the language learners' spoken responses to test prompts. The task has significant relevance to the field of computer assisted language learning. The approach presented in the paper relies on two separate modules: (1) an automatic speech recognition system that yields text transcripts of the spoken interactions involved, and (2) a multiple classifier system based on deep learners that ranks the transcripts into proficiency classes. Different deep neural network architectures (both feed-forward and recurrent) are specialized over diverse representations of the texts in terms of: a reference grammar, the outcome of probabilistic language models, several word embeddings, and two bag-of-word models. Combination of the individual classifiers is realized either via a probabilistic pseudo-joint model, or via a neural mixture of experts. Using the data of the third Spoken CALL Shared Task challenge, the highest values to date were obtained in terms of three popular evaluation metrics.
Sara Papi, Edmondo Trentin, Roberto Gretter, Marco Matassoni, Daniele Falavigna
INTERSPEECH4
2020 TLT-school: a Corpus of Non Native Children Speech
abstract
This paper describes “TLT-school” a corpus of speech utterances collected in schools of northern Italy for assessing the performance of students learning both English and German. The corpus was recorded in the years 2017 and 2018 from students aged between nine and sixteen years, attending primary, middle and high school. All utterances have been scored, in terms of some predefined proficiency indicators, by human experts. In addition, most of utterances recorded in 2017 have been manually transcribed carefully. Guidelines and procedures used for manual transcriptions of utterances will be described in detail, as well as results achieved by means of an automatic speech recognition system developed by us. Part of the corpus is going to be freely distributed to scientific community particularly interested both in non-native speech recognition and automatic assessment of second language proficiency.
Roberto Gretter, Marco Matassoni, Stefano Bannò, Daniele Falavigna
LREC2
2019 Automatic Assessment of Spoken Language Proficiency of Non-native Children
abstract
This paper describes technology developed to automatically grade Italian students (ages 9-16) on their English and German spoken language proficiency. The students' spoken answers are first transcribed by an automatic speech recognition (ASR) system and then scored using a feedforward neural network (NN) that processes features extracted from the automatic transcriptions. In-domain acoustic models, employing deep neural networks (DNNs), are derived by adapting the parameters of an original out of domain DNN. Automatic scores are computed for low level proficiency indicators - such as: lexical richness, syntax correctness, quality of pronunciation, discourse fluency, semantic relevance to the prompt, etc - defined by human experts in language proficiency. A set of experiments was carried out on a large set of data collected during proficiency evaluation campaigns involving thousands of students, manually scored by human experts. Obtained results are presented and discussed.
Roberto Gretter, Marco Matassoni, Katharina Allgaier, Svetlana Tchistiakova, Daniele Falavigna
ICASSP2
2018 Non-Native Children Speech Recognition Through Transfer Learning
abstract
This work deals with non-native children's speech and investigates both multi-task and transfer learning approaches to adapt a multi-language Deep Neural Network (DNN) to speakers, specifically children, learning a foreign language. The application scenario is characterized by young students learning English and German and reading sentences in these second-languages, as well as in their mother language. The paper analyzes and discusses techniques for training effective DNN-based acoustic models starting from children's native speech and performing adaptation with limited non-native audio material. A multi -lingual model is adopted as baseline, where a common phonetic lexicon, defined in terms of the units of the International Phonetic Alphabet (IPA), is shared across the three languages at hand (Italian, German and English); DNN adaptation methods based on transfer learning are evaluated on significant non-native evaluation sets. Results show that the resulting non-native models allow a significant improvement with respect to a mono-lingual system adapted to speakers of the target language.
Marco Matassoni, Roberto Gretter, Daniele Falavigna, Diego Giuliani
ICASSP1
2018 Automatic quality estimation for ASR system combination
Shahab Jalalvand, Matteo Negri, Daniele Falavigna, Marco Matassoni, Marco Turchi
Comput. Speech Lang.4
2017 Optimizing DNN Adaptation for Recognition of Enhanced Speech
Marco Matassoni, Alessio Brutti, Daniele Falavigna
INTERSPEECH1
2017 DNN adaptation by automatic quality estimation of ASR hypotheses
Daniele Falavigna, Marco Matassoni, Shahab Jalalvand, Matteo Negri, Marco Turchi
Comput. Speech Lang.2
2016 DNN adaptation for recognition of children speech through automatic utterance selection
abstract
This paper describes an approach for adapting a DNN trained on adult speech to children voices. The method extends a previous one, based on the Kullback-Leibler divergence between the original (adult) DNN output distribution and the target one, by accounting for the quality of the supervision of the adaptation utterances. In addition, starting from the observation that by gradually removing from the adaptation set the sentences with higher WERs significant performance improvements can be achieved, we also investigate the usage of automatic selection of adaptation utterances. For determining transcription quality we investigate the use of confidence estimates of recognized hypotheses. We present experiments and related results achieved on an Italian data set of children's speech. We show that the proposed DNN adaptation approach allows to significantly reduce the WER on a given test set from 14.2% (corresponding to using the non adapted DNN, trained on adult speech) to 10.6%. It is worth mentioning that the latter result has been achieved without making use of any training data specific of children's speech.
Marco Matassoni, Daniele Falavigna, Diego Giuliani
SLT1
2016 On the relationship between Early-to-Late Ratio of Room Impulse Responses and ASR performance in reverberant environments
Alessio Brutti, Marco Matassoni
Speech Commun.2
2015 Boosted acoustic model learning and hypotheses rescoring on the CHiME-3 task
abstract
Speech recognition in a realistic noisy environment using multiple microphones is the focal point of the third CHiME challenge. Over the baseline ASR system provided for this challenge, we apply state of the art algorithms for boosting acoustic model learning and hypothesis rescoring to improve the final output. To this aim, we first use the automatic transcription of each channel to re-train the acoustic model for that channel and then we apply linear language model rescoring to find a better solution in the n-best list. LM rescoring is performed using an efficient set of N-gram and Recurrent Neural Network LM (RNNLM) trained on a wisely-selected text set. In the experiments, we show that the proposed approach improves not only the individual channel transcription, but also the enhanced channels produced by MVDR and delay-and-sum beamforming.
Shahab Jalalvand, Daniele Falavigna, Marco Matassoni, Piergiorgio Svaizer, Maurizio Omologo
ASRU3
2014 On the use of Early-To-Late Reverberation ratio for ASR in reverberant environments
abstract
This work presents an analysis of distant-talking speech recognition in a variety of reverberant conditions, correlating ASR performance to the acoustic characteristics of a given propagation channel. In particular we show how, for a digit recognition task, the ASR accuracy is directly related to the Early-to-Late Reverberation ratio of the room impulse response, capturing in a single parameter the reverberation properties of a given channel independently of the setup. Consequently, this measure can be successfully considered for acoustic model training either selecting the most suitable model for a given spatial configuration, or defining the subset of RIRs to be used for the creation of multi-condition models. Experimental results on simulated data as well as on data generated with real impulse responses support our claims.
Alessio Brutti, Marco Matassoni
ICASSP2
2014 The DIRHA-GRID corpus: baseline and tools for multi-room distant speech recognition using distributed microphones
abstract
Distant speech recognition in real-world environments is still a challenging problem and a particularly interesting topic is the investigation of multi-channel processing in case of distributed microphones in home environments. This paper presents an initiative oriented to address the challenges of such a scenario; an experimental recognition framework comprising a multi-room, multi-channel corpus and the accompanying evaluation tools is made publicly available. The overall goal is to represent a common platform for comparing state-of-the-art algorithms, share ideas of different research communities and integrate several components in a realistic distant-talking recognition chain, e.g., voice activity detection, speech/feature enhancement, channel selection and fusion, model
Marco Matassoni, Ramón Fernandez Astudillo, Athanasios Katsamanis, Mirco Ravanelli
INTERSPEECH1
2013 The second 'CHiME' speech separation and recognition challenge: An overview of challenge systems and outcomes
abstract
Distant-microphone automatic speech recognition (ASR) remains a challenging goal in everyday environments involving multiple background sources and reverberation. This paper reports on the results of the 2nd ‘CHiME’ Challenge, an initiative designed to analyse and evaluate the performance of ASR systems in a real-world domestic environment. We discuss the rationale for the challenge and provide a summary of the datasets, tasks and baseline systems. The paper overviews the systems that were entered for the two challenge tracks: small-vocabulary with moving talker and medium-vocabulary with stationary talker. We present a summary of the challenge findings including novel results produced by challenge system combination. Possible directions for future challenges are discussed.
Emmanuel Vincent 0001, Jon Barker, Shinji Watanabe 0001, Jonathan Le Roux, Francesco Nesta, Marco Matassoni
ASRU6
2013 The second 'chime' speech separation and recognition challenge: Datasets, tasks and baselines
abstract
Distant-microphone automatic speech recognition (ASR) remains a challenging goal in everyday environments involving multiple background sources and reverberation. This paper is intended to be a reference on the 2nd `CHiME' Challenge, an initiative designed to analyze and evaluate the performance of ASR systems in a real-world domestic environment. Two separate tracks have been proposed: a small-vocabulary task with small speaker movements and a medium-vocabulary task without speaker movements. We discuss the rationale for the challenge and provide a detailed description of the datasets, tasks and baseline performance results for each track.
Emmanuel Vincent 0001, Jon Barker, Shinji Watanabe 0001, Jonathan Le Roux, Francesco Nesta, Marco Matassoni
ICASSP6
2013 Embedding speech recognition to control lights
Alessandro Sosi, Fabio Brugnara, Luca Cristoforetti, Marco Matassoni, Mirco Ravanelli, Maurizio Omologo
INTERSPEECH4
2013 Blind source extraction for robust speech recognition in multisource noisy environments
Francesco Nesta, Marco Matassoni
Comput. Speech Lang.2
2012 Semi-Blind Model Adaptation using Piece-wise Energy Decay Curve for Large Reverberant Environments
abstract
This work presents semi-blind acoustic model adaptation based on a piece-wise energy decay curve. The dual slope representation of the piece-wise curve accurately captures the early and late reflection decay that helps in precisely modeling the smearing effect caused due to reverberation. The slopes are estimated in a semi-blind fashion, late reflection slope is estimated blindly by finding the highest likelihood obtained after matching the test features with Gaussian mixture models trained on reverberant data, while the early reflection slope is empirically computed. Adaptation using piece-wise decay curve leads to robust acoustic models consequently improving the recognition performance. The approach is tested on connected digits recognition task in a lecture room with various large reverberation times. The performance is compared with the exponential decay approach and incremental MLLR, where the proposed technique is found to be robust and consistent across all the cases.
Abdul Waheed Mohammed, Marco Matassoni, Hari Krishna Maganti, Maurizio Omologo
INTERSPEECH2
2011 A Level-Dependent Auditory Filter-Bank for Speech Recognition in Reverberant Environments
Hari Krishna Maganti, Marco Matassoni
INTERSPEECH2
2011 Real-Time Prototype for Integration of Blind Source Extraction and Robust Automatic Speech Recognition
Francesco Nesta, Marco Matassoni, Hari Krishna Maganti
INTERSPEECH2
2010 Experiments on distant-talking speaker verification in TV scenario
abstract
In this work text-independent speaker verification (SV) in a distant-talking noisy scenario is addressed: users can interact with a TV-system able to understand vocal commands and verify simultaneously the identity of the speaker. The main issues with SV under this scenario are related to reverberation, interfering sound sources (TV output) and usually very short utterances; as a consequence, an increasing confusability among speakers models can be observed. To partially cope with this, we propose a system that exploits the processing of signals acquired by a microphone array and a phonetic class segmentation in unsupervised modality. Comparing the proposed system with a GMM-UBM based system we demonstrate the effectiveness of the approach on data acquired with a real prototype.
Christian Zieger, Marco Matassoni, Maurizio Omologo
ICASSP2
2010 An auditory based modulation spectral feature for reverberant speech recognition
Hari Krishna Maganti, Marco Matassoni
INTERSPEECH2
2008 Effective acoustic adaptation for a distant-talking interactive TV system
abstract
In this paper we have studied how to adapt a close-talking baseline acoustic model to a distant-talking application developed in an interactive TV dialogue system: distant-talking interfaces for control of interactive TV (DICIT) project. We have shown that in order to have effective adaptation from the outof-domain data it is better to acquire that data in the same DICIT environment than using contaminated data. By measuring grammar error rate (GER) and action classification error rate (AER) in addition to word error rate (WER), we have shown the best way to adapt the baseline model using available out-of-domain adaptation data (TIMIT) and small amount of in-domain (DICIT) adaptation data. The best approach is to use cascading MAP adaptation. With less than hours of out-ofdomain data and hour of in-domain data, the cascading MAP improves WER/GER/AER by / / relative respectively over the baseline model. The experimental results show that in-domain adaptation data is definitely needed to improve GER and AER. Index Terms: acoustic model adaptation, distant-talking speech recognition, dialog system
Mark Epstein, Marco Matassoni
INTERSPEECH3
2006 Speech Recognition in Reverberant Environments Using Remote Microphones
abstract
This paper addresses distant-talking speech recognition by means of remote sensors in a reverberant room. Recognition performances is investigated for different ways of initializing, steering, and optimizing the related beamformer. Results show how much critical that front-end processing may be in such a challenging setup, according to the different positions and orientations of the speaker
Luca Giulio Brayda, Christian Wellekens, Marco Matassoni, Maurizio Omologo
ISM3
2003 Use of parallel recognizers for robust in-car speech interaction
abstract
This paper refers to an activity under way at the speech recognition technology level for the development of a hands-free dialogue interaction system in the car environment. The use of a set of HMM recognizers, running in parallel, is being investigated in order to ensure low complexity, modularity, fast response, and to allow a real-time reconfiguration of the language models and grammars according to the policy indicated by natural language understanding and dialogue manager modules. A corpus of spontaneous speech interactions was collected using the Wizard-of-Oz method in a real driving situation with a microphone placed far from the driver. The use of parallel recognition units, each specialized on a given geographical domain, was explored using the resulting real corpus. Experiments show the advantage of selecting the recognized sentence according to the maximum likelihood among the active units when compared to the use of a single language model based on a very large vocabulary.
Luca Cristoforetti, Marco Matassoni, Maurizio Omologo, Piergiorgio Svaizer
ICASSP (1)2
2003 Use of a CSP-based voice activity detector for distant-talking ASR
Luca Armani, Marco Matassoni, Maurizio Omologo, Piergiorgio Svaizer
INTERSPEECH2
2003 Evaluation on the Aurora 2 database of acoustic models that are less noise-sensitive
abstract
The Aurora 2 database may be used as a benchmark for evaluation of algorithms under noisy conditions. In particular, the clean training/noisy test mode is aimed at evaluating models that are trained on clean data only without further adjustments on the noisy data, i.e. under severe mismatch between the training and test conditions. While several researchers proposed techniques at the front-end level to improve recognition performance over the reference hideen Markov model (HMM) baseline, investigations at the back-end level are sought. In this respect, the goal is to develop acoustic models that are intrinsically less noise sensitive. This paper presents the word accuracy yielded by a non-parametric HMM with connectionist estimates of the emission probabilities, i.e. a neural network is applied instead of the usual parametric (Gaussian mixture) probability densities. A regularization technique, relying on a maximum-likelihood parameter grouping algorithm, is explicitly introduced to increase the generalization capability of the model and, in turn, its noise-robustness. Results show that a 15,43% relative word error rate reduction w.r.t. the Gaussianmixture HMM is obtained by averaging over the different noises and SNRs of Aurora 2 test set A.
Edmondo Trentin, Marco Matassoni, Marco Gori
INTERSPEECH2
2003 Noise-tolerant speech recognition: the SNN-TA approach
Edmondo Trentin, Marco Matassoni
Inf. Sci.2
2002 On the joint use of noise reduction and MLLR adaptation for in-car hands-free speech recognition
abstract
This paper refers to an activity under way at the speech recognition technology level for the development of a hands-free dialogue interaction system in the car environment. The work here presented concerns the use of two noise reduction techniques, as well as of MLLR adaptation, for recognition error reduction in low and medium complexity tasks, namely connected digits and spelling with or without bigram/trigram statistical constraints. Experiments are based on the use of SpeechDat Car database, a corpus collected under real noisy conditions. Results show the additive improvements in performance, obtained by adopting noise reduction techniques and MLLR adaptation.
Marco Matassoni, Maurizio Omologo, Alfiero Santarelli, Piergiorgio Svaizer
ICASSP1
2002 Hidden Markov model training with contaminated speech material for distant-talking speech recognition
Marco Matassoni, Maurizio Omologo, Diego Giuliani, Piergiorgio Svaizer
Comput. Speech Lang.1
2001 Use of real and contaminated speech for training of a hands-free in-car speech recognizer
abstract
A database of in-car speech for the Italian language was collected under the European projects SpeechDatCar and VODIS II. It consists of 600 sessions recorded under various noise and driving conditions and includes close-talk signals and far microphone signals for hands-free interaction. This paper describes some recognition experiments on two tasks conceived on a portion of this database: connected digit sequences and isolated command words. Recognition rate achieved by means of HMMs trained on real in-car speech is compared with that accomplished by a speech contamination approach, which aims at simulating in-car data starting from a clean speech corpus. Recognition performance is also analyzed as a function of the different noise conditions and of the consequent SNR at the far microphones. Finally, the effect of HMM adaptation is investigated in order to tune the recognizer on the conditions of the various sessions.
Marco Matassoni, Maurizio Omologo, Piergiorgio Svaizer
INTERSPEECH1
2000 Hands-free speech recognition using a filtered clean corpus and incremental HMM adaptation
abstract
A challenging scenario is addressed in which a hands-free speech recognizer operates in a noisy office environment with incremental model adaptation functionalities. The use of a single far microphone as well as that of a microphone array input are investigated. In a previous work it was shown that the acoustic mismatch, remaining after the application of microphone array processing, can be further reduced by conditioning hidden Markov models to operating acoustic conditions. Conditioned HMMs are models trained using the filtered version of a clean corpus, which is speech material better representing noisy real environments. Afterwards, conditioned models are used as initial models for unsupervised incremental adaptation. Experimental results of connected digit recognition show that models trained with filtered clean speech allows to obtain better recognition performance than models trained with clean speech. Furthermore, results show a significant performance increase when incremental adaptation is applied, even after recognition of few utterances.
Marco Matassoni, Maurizio Omologo, Diego Giuliani
ICASSP1
2000 The Regularized SNN-TA Model for Recognition of Noisy Speech
abstract
The segmental neural network (SNN) architecture was introduced at BBN by Zavaliagkos et al. (1994) for rescoring the N-best hypothesis yielded by a standard continuous density hidden Markov model (CDHMM) applied to automatic speech recognition. An enhanced connectionist model, called SNN with trainable amplitude of activation functions (SNN-TA), presented for use instead of the CDHMM to perform the recognition of isolated words. Viterbi-based segmentation is then introduced, relying on the level building algorithm, that can be combined with the SNN-TA to obtain a hybrid framework for continuous speech recognition. The present paradigm is applied to the recognition of isolated digits, collected in a real car environment under several noisy conditions (traffic, speed, road conditions, etc.) using a microphone placed far from the talker. We stress the fact that robustness to noise can be increased by improving the generalization capabilities of the speech recognizer. In this perspective, while CDHMM completely lack of a proper regularization theory, a regularized SNN-TA model is discussed, which yields effective generalization and noise-tolerance, outperforming the CDHMM on the noisy task under consideration.
Edmondo Trentin, Marco Matassoni
IJCNN (5)2
2000 Annotation of a Multichannel Noisy Speech Corpus
Luca Cristoforetti, Marco Matassoni, Maurizio Omologo, Piergiorgio Svaizer, Enrico Zovato
LREC2
1999 Training of HMM with filtered speech material for hands-free recognition
abstract
This paper addresses the problem of hands-free speech recognition in a noisy office environment. An array of six omnidirectional microphones and a corresponding time delay compensation module are used to provide a beamformed signal as input to a HMM-based recognizer. Training of HMMs is performed either using a clean speech database or using a filtered version of the same database. Filtering consists in a convolution with the acoustic impulse response between the speaker and microphone, to reproduce the reverberation effect. Background noise is summed to provide the desired SNR. The paper shows that the new models trained on these data perform better than the baseline ones. Furthermore, the paper investigates on maximum likelihood linear regression (MLLR) adaptation of the new models. It is shown that a further performance improvement is obtained, allowing to reach a 98.7% WRR in a connected digit recognition task, when the talker is at 1.5 m distance from the array.
Diego Giuliani, Marco Matassoni, Maurizio Omologo, Piergiorgio Svaizer
ICASSP2
1998 Experiments of HMM adaptation for hands-free connected digit recognition
abstract
A scenario concerning hands-free connected digit recognition in a noisy office environment is investigated. An array of six omnidirectional microphones and a corresponding time delay compensation module are used to provide a beamformed signal as input to a hidden Markov model (HMM) based recognizer. Two different techniques of phone HMM adaptation have been considered, to reduce the mismatch between training and test conditions. Adaptation material and test material were collected in two different sessions. Results show that a digit accuracy close to 98% can be achieved when the talker is at 1.5 m distance from the array. This result has to be compared with 99.5% accuracy obtained by using a close-talk microphone.
Diego Giuliani, Marco Matassoni, Maurizio Omologo, Piergiorgio Svaizer
ICASSP2
1998 Environmental conditions and acoustic transduction in hands-free speech recognition
Maurizio Omologo, Piergiorgio Svaizer, Marco Matassoni
Speech Commun.3
1997 Microphone array based speech recognition with different talker-array positions
abstract
The use of a microphone array for hands-free continuous speech recognition in noisy and reverberant environment is investigated. An array of eight omnidirectional microphones was placed at different angles and distances from the talker. A time delay compensation module was used to provide a beamformed signal as input to a hidden Markov model (HMM) based recognizer. A phone HMM adaptation, based on a small amount of phonetically rich sentences, further improved the recognition rate obtained by applying only beamforming. These results were confirmed both by experiments conducted in a noisy and reverberant environment and by simulations. In the latter case, different conditions were recreated by using the image method to reproduce synthetic versions of the array microphone signals.
Maurizio Omologo, Marco Matassoni, Piergiorgio Svaizer, Diego Giuliani
ICASSP2
1997 Acoustic source location in a three-dimensional space using crosspower spectrum phase
abstract
A microphone array can be used to locate a dominant acoustic source in a given environment. This capability is successfully employed to locate an active talker in teleconferencing or other multi-speaker applications. In this work the source location is obtained in two steps: (1) a time difference of arrival (TDOA) computation between the signals of the array; (2) an "optimal" source location based on the interchannel delay estimates and on a geometrical description of the sensor arrangement. The crosspower spectrum phase technique was used for TDOA estimation, while a maximum likelihood approach was followed to derive the source coordinates. Source location experiments in a three-dimensional space were performed by means of an array of 8 microphones. For this purpose both a loudspeaker and a real talker were used to collect data in a large noisy and reverberant room.
Piergiorgio Svaizer, Marco Matassoni, Maurizio Omologo
ICASSP2
1997 Use of different microphone array configurations for hands-free speech recognition in noisy and reverberant environment
abstract
In this work hands-free continuous speech recognition based on microphone arrays is investigated. A set of experiments was carried out using arrays having different numbers of omnidirectional microphones as well as different configurations. Both real and simulated array signals, generated by means of the image method, were used. An enhanced input to a recognizer based on Hidden Markov Models was obtained by a time delay compensation module providing a beamformed signal. HMM adaptation was used to realign the recognizer acoustic modeling to the given acoustic condition.
Diego Giuliani, Marco Matassoni, Maurizio Omologo, Piergiorgio Svaizer
EUROSPEECH2
1995 Hands free continuous speech recognition in noisy environment using a four microphone array
abstract
This paper describes advances in the use of HMM based technology for speaker independent continuous speech recognition, in noisy environment, under hands free interaction mode. For this purpose an array of four omnidirectional microphones is employed as the acquisition system. The processing of phase information in the cross-power spectrum provides the capability both of locating the talker position and of reconstructing an enhanced speech spectrum. Two enhancement techniques are described, that provide recognition improvement in the case of clean input speech as well as under different adverse conditions. The results refer to the use of a new multichannel corpus, collected in a real environment by a microphone array as well as a close-talk microphone.
Diego Giuliani, Marco Matassoni, Maurizio Omologo, Piergiorgio Svaizer
ICASSP2
1995 Robust continuous speech recognition using a microphone array
Diego Giuliani, Marco Matassoni, Maurizio Omologo, Piergiorgio Svaizer
EUROSPEECH2