Pietro Laface

dblp:87/2237 · DBLP profile ↗
← Back
91ranked-venue papers
5as first author
0since 2021 · last 2019
0000-0003-2841-7695ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 69 · 5 first-authorArtificial intelligence and machine learning · 51 · 2 first-authorSystems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
15 papers
Speech recognition and synthesis · 91% Probabilistic and Bayesian machine learning · 3% Efficient and distributed learning · 2%
Theoretical computer science
2 papers
Mathematical optimization · 73% Algorithms and data structures · 27%

Topics — the 25 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis
speaker recognition
2.3112018
Scoring Heterogeneous Speaker Vectors Using Nonlinear Transformations and Tied PLDA Models · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Speaker Recognition Using e-Vectors · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Joint Estimation of PLDA and Nonlinear Transformations of Speaker Vectors · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Natural language and speech › Speech recognition and synthesis › speaker recognition
probabilistic linear discriminant analysis
1.142018
Scoring Heterogeneous Speaker Vectors Using Nonlinear Transformations and Tied PLDA Models · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Joint Estimation of PLDA and Nonlinear Transformations of Speaker Vectors · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Nonlinear I-Vector Transformations for PLDA-Based Speaker Recognition · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Natural language and speech › Speech recognition and synthesis › speaker recognition
i-vector
1.052018
Speaker Recognition Using e-Vectors · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Large-Scale Training of Pairwise Support Vector Machines for Speaker Recognition · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Factorized Sub-Space Estimation for Fast and Memory Effective I-vector Extraction · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Natural language and speech › Speech recognition and synthesis › speaker recognition
speaker verification
0.442018
Pairwise Discriminative Speaker Verification in the 𝕀-Vector Space · IEEE Trans. Speech Audio Process. 2013
Scoring Heterogeneous Speaker Vectors Using Nonlinear Transformations and Tied PLDA Models · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Compensation of Nuisance Factors for Speaker and Language Recognition · IEEE Trans. Speech Audio Process. 2007
Natural language and speech › Speech recognition and synthesis › speaker recognition
speaker modeling
0.312018
Speaker Recognition Using e-Vectors · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Natural language and speech › Speech recognition and synthesis › speech analysis
language identification
0.222012
Analysis of Large-Scale SVM Training Algorithms for Language and Speaker Recognition · IEEE Trans. Speech Audio Process. 2012
Compensation of Nuisance Factors for Speaker and Language Recognition · IEEE Trans. Speech Audio Process. 2007
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
maximum likelihood estimation
0.222017
Joint Estimation of PLDA and Nonlinear Transformations of Speaker Vectors · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Nonlinear I-Vector Transformations for PLDA-Based Speaker Recognition · IEEE ACM Trans. Audio Speech Lang. Process. 2017
Machine learning › Optimization for machine learning › regularized risk minimization
support vector machine training
0.112012
Analysis of Large-Scale SVM Training Algorithms for Language and Speaker Recognition · IEEE Trans. Speech Audio Process. 2012
Machine learning › Efficient and distributed learning › distributed training
large-scale training
0.122014
Large-Scale Training of Pairwise Support Vector Machines for Speaker Recognition · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Analysis of Large-Scale SVM Training Algorithms for Language and Speaker Recognition · IEEE Trans. Speech Audio Process. 2012
Natural language and speech › Speech recognition and synthesis › speaker recognition › speaker verification
joint factor analysis
0.112018
Speaker Recognition Using e-Vectors · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Natural language and speech › Speech recognition and synthesis
nonlinear transformation
0.112018
Scoring Heterogeneous Speaker Vectors Using Nonlinear Transformations and Tied PLDA Models · IEEE ACM Trans. Audio Speech Lang. Process. 2018
Natural language and speech › Speech recognition and synthesis
channel compensation
0.112007
Compensation of Nuisance Factors for Speaker and Language Recognition · IEEE Trans. Speech Audio Process. 2007
Machine learning › Efficient and distributed learning
data selection
0.112014
Large-Scale Training of Pairwise Support Vector Machines for Speaker Recognition · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Algorithms and data structures › numerical linear algebra
matrix factorization
0.112014
Factorized Sub-Space Estimation for Fast and Memory Effective I-vector Extraction · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Mathematical optimization › statistical estimation › multivariate estimation
subspace estimation
0.112014
Factorized Sub-Space Estimation for Fast and Memory Effective I-vector Extraction · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Machine learning › Trustworthy machine learning
pairwise classification
0.012013
Pairwise Discriminative Speaker Verification in the 𝕀-Vector Space · IEEE Trans. Speech Audio Process. 2013
Mathematical optimization
iterative methods
0.012013
Memory and Computation Trade-Offs for Efficient I-Vector Extraction · IEEE Trans. Speech Audio Process. 2013
Mathematical optimization
variational inference
0.012013
Memory and Computation Trade-Offs for Efficient I-Vector Extraction · IEEE Trans. Speech Audio Process. 2013
Machine learning › Kernel, tree and ensemble methods › support vector machine
scalable SVM training
0.012012
Analysis of Large-Scale SVM Training Algorithms for Language and Speaker Recognition · IEEE Trans. Speech Audio Process. 2012
Knowledge, reasoning and agents › Knowledge representation and reasoning
expert systems
0.011982
An Expert System for Interpreting Speech Patterns · AAAI 1982
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.011980
Use of Fuzzy Algorithms for Phonetic and Phonemic Labeling of Continuous Speech · IEEE Trans. Pattern Anal. Mach. Intell. 1980
Knowledge, reasoning and agents › Knowledge representation and reasoning › uncertainty reasoning › fuzzy systems › fuzzy logic
fuzzy algorithms
0.011980
Use of Fuzzy Algorithms for Phonetic and Phonemic Labeling of Continuous Speech · IEEE Trans. Pattern Anal. Mach. Intell. 1980
Knowledge, reasoning and agents › Knowledge representation and reasoning › uncertainty reasoning › fuzzy systems
fuzzy logic
0.011980
Use of Fuzzy Algorithms for Phonetic and Phonemic Labeling of Continuous Speech · IEEE Trans. Pattern Anal. Mach. Intell. 1980
Parallel and multicore computing
parallel programming models
0.011985
Parallel Algorithms for Syllable Recognition in Continuous Speech · IEEE Trans. Pattern Anal. Mach. Intell. 1985
Natural language and speech › Speech recognition and synthesis
speech analysis
0.011982
An Expert System for Interpreting Speech Patterns · AAAI 1982

Methods — techniques the papers use, named apart from their topics

maximum likelihood estimation · 0.6total variability subspace · 0.3nonlinear tied-PLDA · 0.3e-vector extraction · 0.3class-dependent nonlinear transformation · 0.3support vector machine · 0.3nonlinear transformation · 0.3joint estimation · 0.3gaussianization · 0.3affine transformation · 0.3factorized sub-space estimation · 0.2eigendecomposition · 0.2variational bayes · 0.2conjugate gradient · 0.2rule-based system · 0.0distributed processing · 0.0
YearPublicationVenuePosition
2019 Tied Normal Variance-Mean Mixtures for Linear Score Calibration
abstract
A speaker verification system decides whether two voice segments belong to the same speaker based on a threshold. An optimal threshold can be set if the recognition scores are well calibrated, i.e., they represent Log-Likelihood Ratios. Logistic Regression (LogReg) is a standard approach for score calibration. While training this discriminative model requires labeled scores, Gaussian and non-Gaussian generative calibration models have been recently proposed. They not only have similar or better performance with respect to LogReg, but also allow for unsupervised or semi-supervised training of the models.The goal of this work is to extend these models. In particular, we show that normal variance-mean mixture distributions are able to model well-calibrated non-Gaussian distributed scores, provided that their parameters for the target and non-target score distributions are properly tied. As for the Gaussian case, a linear calibration model can then be estimated by computing Maximum Likelihood estimates of the distributions parameters and of the score transformation. The quality of all these approaches has been compared on a dataset of segments of variable duration obtained by cutting the NIST 2010 evaluation test data.
Sandro Cumani, Pietro Laface
ICASSP2
2019 Exact memory-constrained UPGMA for large scale speaker clustering
Sandro Cumani, Pietro Laface
Pattern Recognit.2
2018 Speaker Recognition Using e-Vectors
abstract
Systems based on i-vectors represent the current state-of-the-art in text-independent speaker recognition. Unlike joint factor analysis (JFA), which models both speaker and intersession subspaces separately, in the i-vector approach all the important variability is modeled in a single low-dimensional subspace. This paper is based on the observation that JFA estimates a more informative speaker subspace than the “total variability” i- vector subspace, because the latter is obtained by considering each training segment as belonging to a different speaker. We propose a speaker modeling approach that extracts a compact representation of a speech segment, similar to the speaker factors of JFA and to i-vectors, referred to as “e-vector.” Estimating the e-vector subspace follows a procedure similar to i-vector training, but produces a more accurate speaker subspace, as confirmed by the results of a set of tests performed on the NIST 2012 and 2010 Speaker Recognition Evaluations. Simply replacing the i-vectors with e-vectors we get approximately 10% average improvement of the Cprimarycost function, using different systems and classifiers. It is worth noting that these performance gains come without any additional memory or computational costs with respect to the standard i-vector systems.
Sandro Cumani, Pietro Laface
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Scoring Heterogeneous Speaker Vectors Using Nonlinear Transformations and Tied PLDA Models
abstract
Most current state-of-the-art text-independent speaker recognition systems are based on i-vectors, and on probabilistic linear discriminant analysis (PLDA). PLDA assumes that the i-vectors of a trial are homogeneous, i.e., that they have been extracted by the same system. In other words, the enrollment and test i-vectors belong to the same class. However, it is sometimes important to score trials including “heterogeneous” i-vectors, for instance, enrollment i-vectors extracted by an old system, and test i-vectors extracted by a newer, more accurate, system. In this paper, we introduce a PLDA model that is able to score heterogeneous i-vectors independent of their extraction approach, dimensions, and any other characteristics that make a set of i-vectors of the same speaker belong to different classes. The new model, which will be referred to as nonlinear tied-PLDA (NL-Tied-PLDA), is obtained by a generalization of our recently proposed nonlinear PLDA approach, which jointly estimates the PLDA parameters and the parameters of a nonlinear transformation of the i-vectors. The generalization consists of estimating a class-dependent nonlinear transformation of the i-vectors, with the constraint that the transformed i-vectors of the same speaker share the same speaker factor. The resulting model is flexible and accurate, as assessed by the results of a set of experiments performed on the extended core NIST SRE 2012 evaluation. In particular, NL-Tied-PLDA provides better results on heterogeneous trials with respect to the corresponding homogeneous trials scored by the old system, and, in some configurations, it also reaches the accuracy of the new system. Similar results were obtained on the female-extended core NIST SRE 2010 telephone condition.
Sandro Cumani, Pietro Laface
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 e-vectors: JFA and i-vectors revisited
abstract
Systems based on i-vectors represent the current state-of-the-art in text-independent speaker recognition. In this work we introduce a new compact representation of a speech segment, similar to the speaker factors of Joint Factor Analysis (JFA) and to i-vectors, that we call “e-vector”. The e-vectors derive their name from the eigen-voice space of the JFA speaker modeling approach. Our working hypothesis is that JFA estimates a more informative speaker subspace than the “total variability” i-vector subspace, because the latter is obtained by considering each training segment as belonging to a different speaker. We propose, thus, a simple “i-vector style” modeling and training technique that exploits this observation, and estimates a more accurate subspace with respect to the one provided by the classical i-vector approach, as confirmed by the results of a set of tests performed on the extended core NIST 2012 Speaker Recognition Evaluation dataset. Simply replacing the i-vectors with e-vectors we get approximately 10% average improvement of the Cprimarycost function, using different systems and classifiers. These performance gains come without any additional memory or computational costs with respect to the standard i-vector systems.
Sandro Cumani, Pietro Laface
ICASSP2
2017 Nuance - Politecnico di Torino's 2016 NIST Speaker Recognition Evaluation System
abstract
This paper describes the Loquendo – Politecnico di Torino system evaluated on the 2006 NIST speaker recognition evaluation dataset. This system was among the best participants in this evaluation. It combines the results of two independent GMM systems: a Phonetic GMM and a classical GMM. Both systems rely on an intersession variation compensation approach, performed in the feature domain. It allowed a 30% error rate reduction with respect to our 2005 system. The linear combination of the two GMM engines gives a further 10% error rate reduction. We also report the results of a set of post evaluation experiments, related to the training data for the intersession variation evaluation, both for the telephone and microphone datasets. The approach adopted for the two wire tests is also described, showing the effect of the speaker segmentation component of our system. Finally, we describe how we performed the incremental unsupervised adaptation tests.
Daniele Colibro, Claudio Vair, Emanuele Dalmasso, Kevin Farrell, Gennady Karvitsky, Sandro Cumani, Pietro Laface
INTERSPEECH7
2017 Nonlinear I-Vector Transformations for PLDA-Based Speaker Recognition
abstract
This paper proposes to estimate parametric nonlinear transformations of i-vectors for speaker recognition systems based on probabilistic linear discriminant analysis (PLDA) classification. The Gaussian PLDA model assumes that the i-vectors are distributed according to the standard normal distribution. However, it has been shown that the i-vectors are better modeled, for example, by Heavy-Tailed distributions, and that significant improvement of the classification performance can be obtained by whitening and length normalizing the i-vectors. In this paper, we propose to transform the i-vectors so that their distribution becomes more suitable to discriminate speakers using the PLDA model. This is performed by means of a sequence of affine and nonlinear transformations whose parameters are obtained by maximum likelihood estimation on the development set. Another contribution of this paper is the reduction of the mismatch between the development and evaluation i-vector length distributions by means of a scaling factor tuned for the estimated i-vector distribution, rather than by means of a blind length normalization. Relative improvement between 7% and 14% of the detection cost function was obtained with the proposed technique on the NIST SRE-2010 and SRE-2012 evaluation datasets, using both the traditional GMM/UBM and the hybrid DNN/GMM-based systems.
Sandro Cumani, Pietro Laface
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Joint Estimation of PLDA and Nonlinear Transformations of Speaker Vectors
abstract
The Gaussian probabilistic linear discriminant anal-ysis (PLDA) model assumes Gaussian distributed priors for the latent variables that represent the speaker and channel factors. Assuming that each training i-vector belongs to a different speaker, as is usually done in i-vector extraction, i-vectors generated by a PLDA model can be considered independent and identically distributed with Gaussian distribution. Thus, we have recently proposed to transform the development i-vectors so that their distribution becomes more Gaussian-like. This is obtained by means of a sequence of affine and nonlinear transformations whose parameters are trained by maximum likelihood (ML) estimation on the development set. The evaluation i-vectors are then subject to the same transformation. Although the i-vector “gaussianization” has shown to be effective, since the i-vectors extracted from segments of the same speaker are not independent, the original assumption is not satisfactory. In this work, we show that the model can be improved by properly exploiting the information about the speaker labels, which was ignored in the previous model. In particular, a more effective PLDA model can be obtained by jointly estimating the PLDA parameters and the parameters of the nonlinear transformation of the i-vectors. In other words, while the goal of the previous approach was to “gaussianize” the training i-vectors distribution, the objective of this work is to embed the estimation of the nonlinear i-vector transformation in the PLDA model estimation. We will thus refer to this model as the nonlinear PLDA model. We show that this new approach provides significant gain with respect to PLDA, and a small, yet consistent, improvement with respect to our former i-vector “gaussianization” approach, without further additional costs.
Sandro Cumani, Pietro Laface
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 Memory-aware i-vector extraction by means of sub-space factorization
abstract
Most of the state-of-the-art speaker recognition systems use i-vectors, a compact representation of spoken utterances. Since the “standard” i-vector extraction procedure requires large memory structures, we recently presented the Factorized Sub-space Estimation (FSE) approach, an efficient technique that dramatically reduces the memory needs for i-vector extraction, and is also fast and accurate compared to other proposed approaches. FSE is based on the approximation of the matrix T, representing the speaker variability sub-space, by means of the product of appropriately designed matrices. In this work, we introduce and evaluate a further approximation of the matrices that most contribute to the memory costs in the FSE approach, showing that it is possible to obtain comparable system accuracy using less than a half of FSE memory, which corresponds to more than 60 times memory reduction with respect to the standard method of i-vector extraction.
Sandro Cumani, Pietro Laface
ICASSP2
2015 Speaker recognition by means of acoustic and phonetically informed GMMs
abstract
In this work we assess the recently proposed hybrid Deep Neural Network/Gaussian Mixture Model (DNN/GMM) approach for speaker recognition considering the effects of the granularity of the phonetic DNN model, and of the precision of the corresponding GMM models, which will be referred to as the phonetic GMMs. The aim of this work is to better understand the contributions of the phonetic information provided by the DNN model with respect to the accuracy of the acous tic GMMs in fitting the distribution of the features associated to a given context-dependent phone state. The testbed for this work was the text-independent speaker recognition task defined by NIST for the 2012 Speaker Recognition Evaluation. Our experiment confirms that the acoustic and the phonetic GMMs are complementary. Thus, their score combination yields very good results if the DNN is trained on data collected in an environment similar to the one that is used for testing. We show, however, that using a single Gaussian per DNN state is not the best choice: the best single system has been obtained balancing the phonetic and acoustic precision of a DNN/GMM system
Sandro Cumani, Pietro Laface, Farzana Kulsoom
INTERSPEECH2
2014 Training Pairwise Support Vector Machines with large scale datasets
abstract
We recently presented an efficient approach for training a Pairwise Support Vector Machine (PSVM) with a suitable kernel for a quite large speaker recognition task. The PSVM approach, rather than estimating an SVM model per class according to the “one versus all” discriminative paradigm, classifies pairs of examples as belonging or not to the same class. Training a PSVM with large amount of data, however, is a memory and computational expensive task, because the number of training pairs grows quadratically with the number of training patterns. This paper proposes an approach that allows discarding the training pairs that do not essentially contribute to the set of Support Vectors (SVs) of the training set. This selection of training pairs is feasible because we show that the number of SVs does not grow quadratically, with the number of pairs, but only linearly with the number of speakers in the training set. Our approach dramatically reduces the memory and computational complexity of PSVM training, making possible the use of large datasets, including many speakers. It has been assessed on the extended core conditions of the 2012 Speaker Recognition Evaluation. The results show that the accuracy of the trained PSVMs increases with the training set size, and that the Cprimaryof a PSVM trained with a small subset of the i-vectors pairs is 10-30% better than the one obtained by a generative model trained on the complete set of i-vectors.
Sandro Cumani, Pietro Laface
ICASSP2
2014 Factorized Sub-Space Estimation for Fast and Memory Effective I-vector Extraction
abstract
Most of the state–of–the–art speaker recognition systems use a compact representation of spoken utterances referred to as i–vector. Since the “standard” i–vector extraction procedure requires large memory structures and is relatively slow, new approaches have recently been proposed that are able to obtain either accurate solutions at the expense of an increase of the computational load, or fast approximate solutions, which are traded for lower memory costs. We propose a new approach particularly useful for applications that need to minimize their memory requirements. Our solution not only dramatically reduces the memory needs for i–vector extraction, but is also fast and accurate compared to recently proposed approaches. Tested on the female part of the tel-tel extended NIST 2010 evaluation trials, our approach substantially improves the performance with respect to the fastest but inaccurate eigen-decomposition approach, using much less memory than other methods.
Sandro Cumani, Pietro Laface
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Large-Scale Training of Pairwise Support Vector Machines for Speaker Recognition
abstract
State–of–the–art systems for text–independent speaker recognition use as their features a compact representation of a speaker utterance, known as “i–vector.” We recently presented an efficient approach for training a Pairwise Support Vector Machine (PSVM) with a suitable kernel for i–vector pairs for a quite large speaker recognition task. Rather than estimating an SVM model per speaker, according to the “one versus all” discriminative paradigm, the PSVM approach classifies a trial, consisting of a pair of i–vectors, as belonging or not to the same speaker class. Training a PSVM with large amount of data, however, is a memory and computational expensive task, because the number of training pairs grows quadratically with the number of training i–vectors. This paper demonstrates that a very small subset of the training pairs is necessary to train the original PSVM model, and proposes two approaches that allow discarding most of the training pairs that are not essential, without harming the accuracy of the model. This allows dramatically reducing the memory and computational resources needed for training, which becomes feasible with large datasets including many speakers. We have assessed these approaches on the extended core conditions of the NIST 2012 Speaker Recognition Evaluation. Our results show that the accuracy of the PSVM trained with a sufficient number of speakers is 10%-30% better compared to the one obtained by a PLDA model, depending on the testing conditions. Since the PSVM accuracy increases with the training set size, but PSVM training does not scale well for large numbers of speakers, our selection techniques become relevant for training accurate discriminative classifiers.
Sandro Cumani, Pietro Laface
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 On the use of i-vector posterior distributions in Probabilistic Linear Discriminant Analysis
abstract
The i-vector extraction process is affected by several factors such as the noise level, the acoustic content of the observed features, the channel mismatch between the training conditions and the test data, and the duration of the analyzed speech segment. These factors influence both the i–vector estimate and its uncertainty, represented by the i–vector posterior covariance. This paper presents a new PLDA model that, unlike the standard one, exploits the intrinsic i–vector uncertainty. Since the recognition accuracy is known to decrease for short speech segments, and their length is one of the main factors affecting the i–vector covariance, we designed a set of experiments aiming at comparing the standard and the new PLDA models on short speech cuts of variable duration, randomly extracted from the conversations included in the NIST SRE 2010 extended dataset, both from interviews and telephone conversations. Our results on NIST SRE 2010 evaluation data show that in different conditions the new model outperforms the standard PLDA by more than 10% relative when tested on short segments with duration mismatches, and is able to keep the accuracy of the standard model for long enough speaker segments. This technique has also been successfully tested in the NIST SRE 2012 evaluation.
Sandro Cumani, Oldrich Plchot, Pietro Laface
IEEE ACM Trans. Audio Speech Lang. Process.3
2013 Probabilistic linear discriminant analysis of i-vector posterior distributions
abstract
The i-vector extraction process is affected by several factors such as the noise level, the acoustic content of the observed features, and the duration of the analyzed speech segment. These factors influence both the i-vector estimate and its uncertainty, represented by the i-vector posterior covariance. This paper present a new PLDA model that, unlike the standard one, exploits the intrinsic i-vector uncertainty. Since short segments are known to decrease recognition accuracy, and segment duration is the main factor affecting the i-vector covariance, we designed a set of experiments aiming at comparing the standard and the new PLDA models on short speech cuts of variable duration, randomly extracted from the conversations included in the NIST SRE 2010 female telephone extended core condition. Our results show that the new model outperforms the standard PLDA when tested on short segments, and keeps the accuracy of the latter for long enough utterances. In particular, the relative improvement is up to 13% for the EER, 5% for DCF08, and 2.5% for DCF10.
Sandro Cumani, Oldrich Plchot, Pietro Laface
ICASSP3
2013 Nuance - Politecnico di torino's 2012 NIST speaker recognition evaluation system
abstract
This paper describes the Nuance-Politecnico di Torino (NPT) speaker recognition system submitted to the NIST SRE12 evaluation campaign. Included are the results of postevaluation tests, focusing on the analysis of the effects of score normalization and condition-dependent calibration. The submitted system combines the results of five acoustic recognizers all based on Gaussian Mixture Models (GMMs). Each system has its own front end, with features differing by their type and dimension. We illustrate the process of development data selection and configuration of state-of-the-art technology, which contributed to obtaining good performance in all the test conditions proposed in this evaluation.
Daniele Colibro, Claudio Vair, Kevin Farrell, Nir Krause, Gennady Karvitsky, Sandro Cumani, Pietro Laface
INTERSPEECH7
2013 Fast and memory effective i-vector extraction using a factorized sub-space
abstract
Most of the state-of-the-art speaker recognition systems use a compact representation of spoken utterances referred to as i-vectors. Since the "standard" i-vector extraction procedure requires large memory structures and is relatively slow, new approaches have recently been proposed that are able to obtain either accurate solutions at the expense of an increase of the computational load, or fast approximate solutions, which are traded for lower memory costs. We propose a new approach particularly useful for applications that need to minimize their memory requirements. Our solution not only dramatically reduces the storage needs for i-vector extraction, but is also fast. Tested on the female part of the tel-tel extended NIST 2010 evaluation trials, our approach substantially improves the performance with respect to the fastest but inaccurate eigen-decomposition approach, using much less memory than any other known method.
Sandro Cumani, Pietro Laface
INTERSPEECH2
2013 Pairwise Discriminative Speaker Verification in the 𝕀-Vector Space
abstract
This work presents a new and efficient approach to discriminative speaker verification in the${\rm i}$–vector space. We illustrate the development of a linear discriminative classifier that is trained to discriminate between the hypothesis that a pair of feature vectors in a trial belong to the same speaker or to different speakers. This approach is alternative to the usual discriminative setup that discriminates between a speaker and all the other speakers. We use a discriminative classifier based on a Support Vector Machine (SVM) that is trained to estimate the parameters of a symmetric quadratic function approximating a log–likelihood ratio score without explicit modeling of the${\rm i}$–vector distributions as in the generative Probabilistic Linear Discriminant Analysis (PLDA) models. Training these models is feasible because it is not necessary to expand the${\rm i}$–vector pairs, which would be expensive or even impossible even for medium sized training sets. The results of experiments performed on the tel-tel extended core condition of the NIST 2010 Speaker Recognition Evaluation are competitive with the ones obtained by generative models, in terms of normalized Detection Cost Function and Equal Error Rate. Moreover, we show that it is possible to train a gender–independent discriminative model that achieves state–of–the–art accuracy, comparable to the one of a gender–dependent system, saving memory and execution time both in training and in testing.
Sandro Cumani, Niko Brümmer, Lukás Burget, Pietro Laface, Oldrich Plchot, Vasileios Vasilakakis
IEEE Trans. Speech Audio Process.4
2013 Memory and Computation Trade-Offs for Efficient I-Vector Extraction
abstract
This work aims at reducing the memory demand of the data structures that are usually pre-computed and stored for fast computation of the i-vectors, a compact representation of spoken utterances that is used by most state-of-the-art speaker recognition systems. We propose two new approaches allowing accurate i-vector extraction but requiring less memory, showing their relations with the standard computation method introduced for eigenvoices, and with the recently proposed fast eigen-decomposition technique. The first approach computes an i-vector in a Variational Bayes (VB) framework by iterating the estimation of one sub-block of i-vector elements at a time, keeping fixed all the others, and can obtain i-vectors as accurate as the ones obtained by the standard technique but requiring only 25% of its memory. The second technique is based on the Conjugate Gradient solution of a linear system, which is accurate and uses even less memory, but is slower than the VB approach. We analyze and compare the time and memory resources required by all these solutions, which are suited to different applications, and we show that it is possible to get accurate results greatly reducing memory demand compared with the standard solution at almost the same speed.
Sandro Cumani, Pietro Laface
IEEE Trans. Speech Audio Process.2
2012 Gender independent discriminative speaker recognition in i-vector space
abstract
Speaker recognition systems attain their best accuracy when trained with gender dependent features and tested with known gender trials. In real applications, however, gender labels are often not given. In this work we illustrate the design of a system that does not make use of the gender labels both in training and in test, i.e. a completely Gender Independent (GI) system. It relies on discriminative training, where the trials are i-vector pairs, and the discrimination is between the hypothesis that the pair of feature vectors in the trial belong to the same speaker or to different speakers. We demonstrate that this pairwise discriminative training can be interpreted as a procedure that estimates the parameters of the best (second order) approximation of the log-likelihood ratio score function, and that a pairwise SVM can be used for training a gender independent system. Our results show that a pairwise GI SVM, saving memory and execution time, achieves on the last NIST evaluations state-of-the-art performance, comparable to a Gender Dependent(GD) system.
Sandro Cumani, Ondrej Glembek, Niko Brümmer, Edward de Villiers, Pietro Laface
ICASSP5
2012 Analysis of Large-Scale SVM Training Algorithms for Language and Speaker Recognition
abstract
This paper compares a set of large scale support vector machine (SVM) training algorithms for language and speaker recognition tasks. We analyze five approaches for training phonetic and acoustic SVM models for language recognition. We compare the performance of these approaches as a function of the training time required by each of them to reach convergence, and we discuss their scalability towards large corpora. Two of these algorithms can be used in speaker recognition to train a SVM that classifies pairs of utterances as either belonging to the same speaker or to two different speakers. Our results show that the accuracy of these algorithms is asymptotically equivalent, but they have different behavior with respect to the time required to converge. Some of these algorithms not only scale linearly with the training set size, but are also able to give their best results after just a few iterations. State-of-the-art performance has been obtained in the female subset of the NIST 2010 Speaker Recognition Evaluation extended core test using a single SVM system.
Sandro Cumani, Pietro Laface
IEEE Trans. Speech Audio Process.2
2011 Loquendo - Politecnico di Torino's 2010 NIST speaker recognition evaluation system
abstract
This paper describes the improvement introduced in the Loquendo-Politecnico di Torino (LPT) speaker recognition system submitted to the NIST SRE10 evaluation campaign. This system combines the results of eight core acoustic systems all based on Gaussian Mixture Models (GMMs). We illustrate the key factors, in the selection of the development data and in engineering state-of-the art technology, which contributed to the very good performance and calibration of our system in all the test conditions proposed in this evaluation.
Fabio Castaldo, Daniele Colibro, Claudio Vair, Sandro Cumani, Pietro Laface
ICASSP5
2011 Fast discriminative speaker verification in the i-vector space
abstract
This work presents a new approach to discriminative speaker verification. Rather than estimating speaker models, or a model that discriminates between a speaker class and the class of all the other speakers, we directly solve the problem of classifying pairs of utterances as belonging to the same speaker or not.
Sandro Cumani, Niko Brümmer, Lukás Burget, Pietro Laface
ICASSP4
2011 Comparison of Speaker Recognition Approaches for Real Applications
abstract
This paper describes the experimental setup and the results obtained using several state-of-the-art speaker recognition classifiers.The comparison of the different approaches aims at the development of real world applications, taking into account memory and computational constraints, and possible mismatches with respect to the training environment.The NIST SRE 2008 database has been considered our reference dataset, whereas nine commercially available databases of conversational speech in languages different form the ones used for developing the speaker recognition systems have been tested as representative of an application domain.Our results, evaluated on the two domains, show that the classifiers based on i-vectors obtain the best recognition and calibration accuracy.Gaussian PLDA and a recently introduced discriminative SVM together with an adaptive symmetric score normalization achieve the best performance using low memory and processing resources.
Sandro Cumani, Pier Domenico Batzu, Daniele Colibro, Claudio Vair, Pietro Laface, Vasileios Vasilakakis
INTERSPEECH5
2010 Loquendo-Politecnico di Torino system for the 2009 NIST Language Recognition Evaluation
abstract
This paper describes the system submitted by Loquendo and Politecnico di Torino (LPT) for the 2009 NIST Language Recognition Evaluation. The system is a combination of classifiers based on two core acoustic models and on two core phone tokenizers. It exploits several state-of-the-art techniques that have been successfully applied in recent years both in speaker and in language recognition.
Fabio Castaldo, Daniele Colibro, Sandro Cumani, Emanuele Dalmasso, Pietro Laface, Claudio Vair
ICASSP5
2010 Parallel implementation of artificial neural network training
abstract
In this paper we describe the implementation of a complete ANN training procedure for speech recognition using the block mode back-propagation learning algorithm. We exploit the high performance SIMD architecture of GPU using CUDA and its C-like language interface. We also compare the speed-up obtained implementing the training procedure only taking advantage of the multi-thread capabilities of multi-core processors. Our approach has been tested by training acoustic models for large vocabulary speech recognition tasks, showing a 6 times reduction of the time required to train real-world large size networks with respect to an already optimized implementation using the Intel MKL libraries.
Stefano Scanzio, Sandro Cumani, Roberto Gemello, Franco Mana, Pietro Laface
ICASSP5
2010 Parallel implementation of Artificial Neural Network training for speech recognition
Stefano Scanzio, Sandro Cumani, Roberto Gemello, Franco Mana, Pietro Laface
Pattern Recognit. Lett.5
2009 Loquendo - Politecnico di Torino's 2008 NIST speaker recognition evaluation system
abstract
This paper describes the improvements introduced in the Loquendo-Politecnico di Torino (LPT) speaker recognition system submitted to the NIST SRE08 evaluation campaign. This system, which was among the best participants in this evaluation, combines the results of three core acoustic systems, two based on Gaussian mixture models (GMMs), and one on phonetic GMMs. We discuss the results of the experiments performed for the 10sec-10sec condition and for the core condition, including the challenging tasks involving a target speaker and an interviewer. The error rate reduction of our SRE08 system compared to the SRE06 system ranges from 25% of the telephone-interview condition to 57% of the interview-interview condition. On the test with telephone and microphone conversations, the improvements range from 9% to 32%.
Emanuele Dalmasso, Fabio Castaldo, Pietro Laface, Daniele Colibro, Claudio Vair
ICASSP3
2009 Language recognition using language factors
abstract
Language recognition systems based on acoustic models reach state of the art performance using discriminative training techniques.In speaker recognition, eigenvoice modeling of the speaker, and the use of speaker factors as input features to SVMs has recently been demonstrated to give good results compared to the standard GMM-SVM approach, which combines GMMs supervectors and SVMs.In this paper we propose, in analogy to the eigenvoice modeling approach, to estimate an eigen-language space, and to use the language factors as input features to SVM classifiers.Since language factors are low-dimension vectors, training and evaluating SVMs with different kernels and with large training examples becomes an easy task.This approach is demonstrated on the 14 languages of the NIST 2007 language recognition task, and shows performance improvements with respect to the standard GMM-SVM technique.
Fabio Castaldo, Sandro Cumani, Pietro Laface, Daniele Colibro
INTERSPEECH3
2009 Word confidence using duration models
abstract
In this paper, we propose a word confidence measure based on phone durations depending on large contexts.The measure is based on the expected duration of each recognized phone in a word.In the approach here proposed the duration of each phone is in principle context-dependent, and the measure is a function of the distance between the observed and expected phone duration distributions within a word.Our experiments show that, since the "duration confidence" does not make use of any acoustic information, its Equal Error Rate (EER) in terms of False Accept and False Rejection rates is not as good as the one obtained by using the more informed acoustic confidence measure.However, combining the two measures by a simple linear interpolation, the system EER improves by 6% to 10% relative on an isolated word recognition task in several languages.
Stefano Scanzio, Pietro Laface, Daniele Colibro, Roberto Gemello
INTERSPEECH2
2008 Stream-based speaker segmentation using speaker factors and eigenvoices
abstract
This paper presents a stream-based approach for unsupervised multi-speaker conversational speech segmentation. The main idea of this work is to exploit prior knowledge about the speaker space to find a low dimensional vector of speaker factors that summarize the salient speaker characteristics. This new approach produces segmentation error rates that are better than the state of the art ones reported in our previous work on the segmentation task in the NIST 2000 Speaker Recognition Evaluation (SRE). We also show how the performance of a speaker recognition system in the core test of the 2006 NIST SRE is affected, comparing the results obtained using single speaker and automatically segmented test data.
Fabio Castaldo, Daniele Colibro, Emanuele Dalmasso, Pietro Laface, Claudio Vair
ICASSP4
2008 Politecnico di Torino system for the 2007 NIST language recognition evaluation
abstract
This paper describes the system submitted by Politecnico di Torino for the 2007 NIST Language Recognition Evaluation.The system, which was among the best participants in this evaluation, is a combination of classifiers based on three acoustic models and on two sets of Parallel Phone tokenizers.It exploits several state-of-the-art techniques that have been successfully applied in recent years both in speaker and in language recognition.We illustrate the models, the classification techniques and the performance of the system components, and of their combination, in the NIST-07 close-set 30 sec General Language Recognition task.We also highlight the difficulties in setting appropriate decision thresholds whenever the training data of a language are scarce, or the test data are collected through previously unseen channels.
Fabio Castaldo, Emanuele Dalmasso, Pietro Laface, Daniele Colibro, Claudio Vair
INTERSPEECH3
2008 On the use of a multilingual neural network front-end
abstract
This paper presents a front-end consisting of an Artificial Neural Network (ANN) architecture trained with multilingual corpora.The idea is to train an ANN front-end able to integrate the acoustic variations included in databases collected for different languages, through different channels, or even for specific tasks.This ANN front-end produces discriminant features that can be used as observation vectors for language or task dependent recognizers.The approach has been evaluated on three difficult tasks: recognition of non-native speaker sentences, training of a new language with a limited amount of speech data, and training of a model for car environment using a clean microphone corpus of the target language and data collected in car environment in another language.
Stefano Scanzio, Pietro Laface, Luciano Fissore, Roberto Gemello, Franco Mana
INTERSPEECH2
2007 Language Identification using Acoustic Models and Speaker Compensated Cepstral-Time Matrices
abstract
This work presents two contributions to language identification. The first contribution is the definition of a set of properly selected time-frequency features that are a valid alternative to the commonly used shifted delta cepstral features. As a second contribution, we show that significant performance improvement in language recognition can be obtained estimating a subspace that represents the distortions due to inter-speaker variability within the same language, and compensating these distortions in the domain of the features. Experiments on the NIST 1996 and 2003 Language Recognition Evaluation data have been successfully used to validate the effectiveness of the proposed techniques.
Fabio Castaldo, Emanuele Dalmasso, Pietro Laface, Daniele Colibro, Claudio Vair
ICASSP (4)3
2007 Acoustic language identification using fast discriminative training
abstract
Gaussian Mixture Models (GMMs) in combination with Support Vector Machine (SVM) classifiers have been shown to give excellent classification accuracy in speaker recognition.In this work we use this approach for language identification, and we compare its performance with the standard approach based on GMMs.In the GMM-SVM framework, a GMM is trained for each training or test utterance.Since it is difficult to accurately train a model with short utterances, in these conditions the standard GMMs perform better than the GMM-SVM models.To overcome this limitation, we present an extremely fast GMM discriminative training procedure that exploits the information given by the separation hyperplanes estimated by an SVM classifier.We show that our discriminative GMMs provide considerable improvement compared with the standard GMMs and perform better than the GMM-SVM approach for short utterances, achieving state of the art performance for acoustic only systems.
Fabio Castaldo, Daniele Colibro, Emanuele Dalmasso, Pietro Laface, Claudio Vair
INTERSPEECH4
2007 Speeding-up neural network training using sentence and frame selection
abstract
Training Artificial Neural Networks (ANNs) with large amounts of speech data is a time intensive task due to the intrinsically sequential nature of the back-propagation algorithm.This paper presents an approach for training ANNs using sentence and frame selection.The goal is to speed-up the training process, and to balance the phonetic coverage of the selected frames, trying to mitigate the classification problems related to the prior probabilities of the individual phonetic classes.These techniques, together with a three-step training approach and software optimizations, reduced by an order of magnitude the training time of our models.
Stefano Scanzio, Pietro Laface, Roberto Gemello, Franco Mana
INTERSPEECH2
2007 Loquendo - Politecnico di torino's 2006 NIST speaker recognition evaluation system
abstract
This paper describes the Loquendo -Politecnico di Torino system evaluated on the 2006 NIST speaker recognition evaluation dataset.This system was among the best participants in this evaluation.It combines the results of two independent GMM systems: a Phonetic GMM and a classical GMM.Both systems rely on an intersession variation compensation approach, performed in the feature domain.It allowed a 30% error rate reduction with respect to our 2005 system.The linear combination of the two GMM engines gives a further 10% error rate reduction.We also report the results of a set of post evaluation experiments, related to the training data for the intersession variation evaluation, both for the telephone and microphone datasets.The approach adopted for the two wire tests is also described, showing the effect of the speaker segmentation component of our system.Finally, we describe how we performed the incremental unsupervised adaptation tests.
Claudio Vair, Daniele Colibro, Fabio Castaldo, Emanuele Dalmasso, Pietro Laface
INTERSPEECH5
2007 Automatic speech recognition and speech variability: A review
Mohamed Benzeghiba, Renato De Mori, Olivier Deroo, Stéphane Dupont, Teodora Erbes, Denis Jouvet, Luciano Fissore, Pietro Laface, Alfred Mertins, Christophe Ris, Richard Rose, Vivek Tyagi, Christian Wellekens
Speech Commun.8
2007 Linear hidden transformations for adaptation of hybrid ANN/HMM models
Roberto Gemello, Franco Mana, Stefano Scanzio, Pietro Laface, Renato De Mori
Speech Commun.4
2007 Introduction to the Special Issue on Intrinsic Speech Variations
Renato De Mori, Olivier Deroo, Stéphane Dupont, Denis Jouvet, Luciano Fissore, Pietro Laface, Alfred Mertins, Christian Wellekens
Speech Commun.6
2007 Compensation of Nuisance Factors for Speaker and Language Recognition
abstract
The variability of the channel and environment is one of the most important factors affecting the performance of text-independent speaker verification systems. The best techniques for channel compensation are model based. Most of them have been proposed for Gaussian mixture models, while in the feature domain blind channel compensation is usually performed. The aim of this work is to explore techniques that allow more accurate intersession compensation in the feature domain. Compensating the features rather than the models has the advantage that the transformed parameters can be used with models of a different nature and complexity and for different tasks. In this paper, we evaluate the effects of the compensation of the intersession variability obtained by means of the channel factors approach. In particular, we compare channel variability modeling in the usual Gaussian mixture model domain, and our proposed feature domain compensation technique. We show that the two approaches lead to similar results on the NIST 2005 Speaker Recognition Evaluation data with a reduced computation cost. We also report the results of a system, based on the intersession compensation technique in the feature space that was among the best participants in the NIST 2006 Speaker Recognition Evaluation. Moreover, we show how we obtained significant performance improvement in language recognition by estimating and compensating, in the feature domain, the distortions due to interspeaker variability within the same language.
Fabio Castaldo, Daniele Colibro, Emanuele Dalmasso, Pietro Laface, Claudio Vair
IEEE Trans. Speech Audio Process.4
2006 Automatic Speech Recognition and Intrinsic Speech Variation
abstract
This paper briefly reviews state of the art related to the topic of speech variability sources in automatic speech recognition systems. It focuses on some variations within the speech signal that make the ASR task difficult. The variations detailed in the paper are intrinsic to the speech and affect the different levels of the ASR processing chain. For different sources of speech variation, the paper summarizes the current knowledge and highlights specific feature extraction or modeling weaknesses and current trends
Mohamed Benzeghiba, Renato De Mori, Olivier Deroo, Stéphane Dupont, Teodora Erbes, Denis Jouvet, Luciano Fissore, Pietro Laface, Alfred Mertins, Christophe Ris, Richard Rose, Vivek Tyagi, Christian Wellekens
ICASSP (5)8
2006 Adaptation of Hybrid ANN/HMM Models Using Linear Hidden Transformations and Conservative Training
abstract
A technique is proposed for the adaptation of automatic speech recognition systems using hybrid models combining artificial neural networks with hidden Markov models. The application of linear transformations not only to the input features, but also to the outputs of the internal layers is investigated. The motivation is that the outputs of an internal layer represent a projection of the input pattern into a space where it should be easier to learn the classification or transformation expected at the output of the network. A new solution, called conservative training, is proposed that compensates for the lack of adaptation samples in certain classes. Supervised adaptation experiments with different corpora and for different adaptation types are described. The results show that the proposed approach always outperforms the use of transformations in the feature space and yields even better results when combined with linear input transformations
Roberto Gemello, Franco Mana, Stefano Scanzio, Pietro Laface, Renato De Mori
ICASSP (1)4
2006 Adaptation of Hybrid ANN/HMM Using Weights Interpolation
abstract
Many techniques for speaker or channel adaptation have been successfully applied to automatic speech recognition. Most of these techniques have been proposed for the adaptation of hidden Markov models (HMMs). Far less proposals have been made for the adaptation of the artificial neural networks (ANNs) used in the hybrid HMM-ANN approach. This paper presents an adaptation technique for ANNs that, similar to the framework of MAP estimation, tries to exploit in the adaptation process prior information that is particularly useful to deal with the problem of sparse training data. We show that the integration of a priori information can be simply achieved by linear interpolation of the weights of an "a priori" network and of a speaker specific network. Good improvements with respect to the baseline results are reported evaluating this technique on the Wall Street Journal WSJ0 and WSJ1 databases and on TIMIT corpus using different amounts of adaptation data
Stefano Scanzio, Pietro Laface, Roberto Gemello, Franco Mana
ICASSP (5)2
2006 Adaptation of Artificial Neural Networks Avoiding Catastrophic Forgetting
abstract
In connectionist learning, one relevant problem is "catastrophic forgetting" that may occur when a network, trained with a large set of patterns, has to learn new input patterns, or has to be adapted to a different environment. The risk of catastrophic forgetting is particularly high when a network is adapted with new data that do not adequately represent the knowledge included in the original training data. Two original solutions are proposed to reduce the risk that the network focuses on new data only, loosing its generalization capability. The first one, conservative training, is a variant to the target assignment policy, while the second approach, support vector rehearsal, selects from the training set the patterns that lay near the borders of the classes not included in the adaptation set. These patterns are used as sentinels that try to keep unchanged the original boundaries of these classes. Moreover, we investigated the extension of the classical approach consisting in applying linear transformations not only to the input features, but also to the outputs of the internal layers. The motivation is that the outputs of an internal layer represent a projection of the input pattern into a space where it should be easier to learn the classification or transformation expected at the output of the network. We illustrate the problems using an artificial test-bed, and apply our techniques to a set of adaptation tasks in the domain of automatic speech recognition (ASR) based on artificial neural networks. Supervised ASR adaptation experiments with several corpora and for different adaptation types are described. We report on the adaptation potential of different techniques, and on the generalization capability of the adapted networks. The results show that the combination of the proposed approaches mitigates the catastrophic forgetting effects, and always outperforms the use of the classical transformations in the feature space.
Dario Albesano, Roberto Gemello, Pietro Laface, Franco Mana, Stefano Scanzio
IJCNN3
2005 Learning Pronunciation and Formulation Variants in Continuous Speech Applications
abstract
Most voice driven applications are based on recognition grammars. In complex applications it is difficult to exactly predict how the users will formulate their requests even if a careful study of the user's behavior has been performed. Moreover, it is possible that a speaker's word pronunciation does not match the phonetic transcription of the system, mainly in the case of foreign words. Loquendo has developed a tool that collects field data, detects the most significant weaknesses of the application due to pronunciation of formulation mismatches, and filters the collected field corpora. This permits the application designers to perform their analysis only on a reasonable amount of preprocessed and automatically labeled data. This paper presents the approaches that have been devised to detect pronunciation variants of vocabulary words and linguistic formulations not covered by the recognition grammar. Results showing the improvements that have been obtained including automatically detected formulations in three grammars for two languages are also detailed.
Daniele Colibro, Luciano Fissore, Cosmin Popovici, Claudio Vair, Pietro Laface
ICASSP (1)5
2005 A confidence measure invariant to language and grammar
abstract
Confidence measures are necessary in all voiced activated applications to decide whether a recognized word, or a sentence, should be accepted or rejected.A confidence measure should not only be reliable, but possibly application independent, i.e. its dynamic range should be uniform for different languages, grammars, and vocabularies.This is an important practical issue because it allows the application developers to use the same value of the threshold for different applications and to expect comparable rejection rates.This eases their task at least in the first phase of application development.In this paper, we introduce a confidence measure that has these properties.It allows eliminating the cumbersome experimental procedure necessary to tune individually the rejection threshold for every developed recognition object.We present the results of a set of experiments that demonstrate the "normalization" quality of our confidence measure for six different grammars in different languages.
Daniele Colibro, Luciano Fissore, Claudio Vair, Emanuele Dalmasso, Pietro Laface
INTERSPEECH5
2005 Unsupervised segmentation and verification of multi-speaker conversational speech
abstract
This paper presents our approach to unsupervised multispeaker conversational speech segmentation.Speech segmentation is obtained in two steps that employ different techniques.The first step performs a preliminary segmentation of the conversation analyzing fixed length slices, and assumes the presence in every slice of one or two speakers.The second step clusters the segments obtained by the previous step, estimates the number of speaker, and refines the segment boundaries using more accurate models.We evaluated our algorithms on the speaker segmentation tasks proposed by the 2000 NIST speaker recognition evaluation where the proposed approach produces state-of-the art segmentation error rates and on the 2004 NIST multispeaker conversation tests where we compare the verification performance using automatically segmented training data with the one obtained using single speaker data.
Emanuele Dalmasso, Pietro Laface, Daniele Colibro, Claudio Vair
INTERSPEECH2
2003 Incremental learning of new user formulations in automatic directory assistance
Marco Andorno, Luciano Fissore, Pietro Laface, Mario Nigra, Cosmin Popovici, Franco Ravera, Claudio Vair
INTERSPEECH3
2002 Learning new user formulations in automatic Directory Assistance
abstract
Telecom Italia has deployed since the beginning of year 2001 a nationwide automatic Directory Assistance (DA) system that routinely serves customers asking for residential and business listings.
Cosmin Popovici, Marco Andorno, Pietro Laface, Luciano Fissore, Mario Nigra, Claudio Vair
ICASSP3
2002 Experiments in confidence scoring for word and sentence verification
abstract
The successful deployment of a telephone speech application cannot only rely on the accuracy of the recognition results, but also on their reliability.Reliable confidence measures are, thus, necessary in all practical applications to decide whether a recognized wordor sentence -should be accepted or rejected.Since most of the applications are based on continuous speech recognition, controlled by grammars, we present the results of a set of experiments aiming at assessing the quality and the limitations of different confidence measures for six different grammars that can be embedded in several applications.We show that using application independent confidence scoring techniques, good performance are obtained across all six grammars.We introduce also a sentence level confidence measure that allows a significant reduction of the system error rate due to ill-formed sentences.
Marco Andorno, Pietro Laface, Roberto Gemello
INTERSPEECH2
2001 Learning of user formulations for business listings in automatic directory assistance
abstract
Automatic Directory Assistance (DA) for business listings poses many application specific problems.One of the main problem is that customers formulate their requests for the same listing with a great variability.This paper presents the results of a study aiming at the evaluation of an approach towards automatic learning, from field data, of expressions typically used by customers to formulate their requests for the most frequent business listings.We use a clustering procedure that exploits the association of the phonetic string produced by a lexical unconstrained search for a given denomination pronounced by the user and the phone number provided by the system or by the human operator, in case of failure of the automatic DA service.We show that an unsupervised approach allows to detect user formulations that were not foreseen by the designers, and that can be added, as variants, to the denominations already included in the system to reduce its failures.
Cosmin Popovici, Marco Andorno, Pietro Laface, Luciano Fissore, Mario Nigra, Claudio Vair
INTERSPEECH3
2000 Synergy of spectral and perceptual features in multi-source connectionist speech recognition
abstract
The combined use of different set of features extracted from the speech signal with different processing algorithms is a promising approach to improve speech recognition performances.Artificial Neural Networks are well suited to this task since they are able to use directly multiple heterogeneous input features to estimate a near optimal combination of them for classification, without being constrained by a priori assumptions on the stochastic independence of the input sources.This work shows how we have taken advantage of these characteristics of Neural Networks to improve the recognition accuracy of our systems.In particular, three set of input features have been considered as sources in this work: Mel based Cepstral Coefficients derived from the FFT spectrum, RASTA-PLP Cepstral Coefficients, and a set of features that describe the dynamics of the FFT power spectrum along the frequency dimension, instead of the usual time dimension.The experimental results confirm the usefulness of the proposed approach of feature integration that leads to a significant error reduction both on isolated and continuous speech recognition tasks on a large telephone speech test set.
Roberto Gemello, Loreta Moisa, Pietro Laface
INTERSPEECH3
2000 Dynamic adaptation of vocabulary independent HMMs to an application environment
abstract
In this paper, the authors present a software architecture for collecting, selecting, and using speech data and applying the method to a train timetable information system.
Claudio Vair, Luciano Fissore, Pietro Laface
INTERSPEECH3
1999 Connected digit recognition using short and long duration models
abstract
We show that accurate HMMs for connected word recognition can be obtained without context dependent modeling and discriminative training. We train two HMMs for each word that have the same, standard, left to right topology with the possibility of skipping once state, but each model has a different number of states, automatically selected. The two models account for different speaking rates that occur not only in different utterances of the speakers, but also within a connected word utterance of the same speaker. This simple modeling technique has been applied to connected digit recognition using the adult speaker portion of the TI/NIST corpus giving the best results reported so far for this database. It has also been tested on telephone speech using long sequences of Italian digits (credit card numbers), giving better results with respect to classical models with a larger number of densities.
Cristina Chesta, Pietro Laface, Franco Ravera
ICASSP2
1999 Piecewise HMM discriminative training
Cristina Chesta, Pietro Laface, Mario Nigra
EUROSPEECH2
1998 Discriminative training of hidden Markov models using a classification measure criterion
abstract
This paper proposes the optimization of a non-standard objective function in the framework of maximum mutual information estimation (MMIE). In contrast with the classical MMIE estimation, where only misrecognized training utterances contribute to the optimization process, the contributions of near-miss classifications are naturally embedded in the maximization of the proposed function because it takes into account a non-linear combination of the probabilities of the competing models that can be tuned by means of a single parameter. This corrective training procedure has been applied to an isolated word recognition task leading to significant performance improvements with respect to maximum likelihood estimation and MMIE.
Cristina Chesta, Aldo Girardi, Pietro Laface, Mario Nigra
ICASSP3
1998 HMM topology selection for accurate acoustic and duration modeling
abstract
In this paper we show that accurate HMMs for connected word recognition can be obtained without context dependent modeling and discriminative training. To account for di erent speaking rates, we de ne two HMMs for each word that must be trained. The two models have the same, standard, left to right topology with the possibility of skipping one state, but each model has a di erent number of states, automatically selected. Our simple modeling and training technique has been applied to connected digit recognition using the adult speaker portion of the TI/NIST corpus. The obtained results are comparable with the best ones reported in the literature for models with a larger number of densities.
Cristina Chesta, Pietro Laface, Franco Ravera
ICSLP2
1998 Automatic classification of dialogue contexts for dialogue predictions
Cosmin Popovici, Paolo Baggia, Pietro Laface, Loreta Moisa
ICSLP3
1997 Using word temporal structure in HMM speech recognition
abstract
Isolated word speech recognizers with fixed vocabularies are often used to provide vocal services through the telephone line. The paper illustrates a simple postprocessing approach that allows the hypotheses produced by a hidden Markov model recognizer to be rescored taking into account the global temporal structure of the pronounced words. Our approach does not directly rely on state/word duration modeling. It models, instead, the global time variations of the spectral features of each word and their correlation in time: two important perceptual cues that are only partially exploited by standard HMMs. This method has been evaluated using three isolated word speaker independent systems with vocabulary of different size and complexity. We show that, with minimal overhead, the recognition performance improves not only for small vocabulary recognition systems such as the isolated digit one, or for the recognition of 26 Italian spelling names, but also for a system with a 475 city name vocabulary included in a vocal service that provides information about the main railway connections.
Luciano Fissore, Pietro Laface, Franco Ravera
ICASSP2
1997 Bottom-up and top-down state clustering for robust acoustic modeling
Cristina Chesta, Pietro Laface, Franco Ravera
EUROSPEECH2
1996 Segmental search for continuous speech recognition
abstract
The paper illustrates a search strategy for continuous speech recognition based on the recently developed Fast Segmental Viterbi Algorithm (FSVA) [5], a new search strategy particularly eective for very large vocabulary word recognition.The FSVA search has been extended to deal with continuous speech using a network that merges a general lexical tree and a set of bigram subtrees generated on demand during the search.Results are given for a 751-words speaker independent spontaneous speech recognizer of a railway timetable inquiry application, managed by a dialog system.Preliminary tests have been performed on the Wall Street Journal 5K words 1992 evaluation set.
Pietro Laface, Luciano Fissore, A. Maro, Franco Ravera
ICSLP1
1995 A fast segmental Viterbi algorithm for large vocabulary recognition
abstract
The paper presents a fast segmental Viterbi algorithm. A new search strategy particularly effective for very large vocabulary word recognition. It performs a tree based, time synchronous, left-to-right beam search that develops time-dependent acoustic and phonetic hypotheses. At any given time, it makes active a sub-word unit associated to an arc of a lexical tree only if that time is likely to be the boundary between the current and the next unit. This new technique, tested with a vocabulary of 188892 directory entries, achieves the same results obtained with the Viterbi algorithm, with a 35% speedup. Results are also presented for a 718 word, speaker independent continuous speech recognition task.
Pietro Laface, Claudio Vair, Luciano Fissore
ICASSP1
1995 Acoustic-phonetic modeling for flexible vocabulary speech recognition
Luciano Fissore, Franco Ravera, Pietro Laface
EUROSPEECH3
1994 Model topology selection for isolated word recognition
abstract
The paper describes a search procedure that, given a set of alternate models for each word of a small vocabulary isolated words recognizer, selects the set of models that minimizes the expected number of errors. The reported results show that the number of errors that occur on the test set by using the best set of models selected from the training set is less than the one achieved by models with a fixed number of states, or the same number of errors is obtained with less states.>
Pietro Laface, Luciano Fissore
ICASSP (1)1
1994 Automatic generation of words toward flexible vocabulary isolated word recognition
Pietro Laface, Lorenzo Fissore, Franco Ravera
ICSLP1
1994 A Speech Understanding System for Information Retrieval
abstract
This paper describes a Continuous Speech Understanding System that allows information services to be accessed through the telephone line. It accepts queries within a restricted semantic domain, expressed in free but syntactically correct natural language, with a lexicon of the order of 800 words. In the implementation here described, a user can access an electronic mailbox or a train information service through a PABX telephone line. The architecture of the system is based on two main modules that represent and use different knowledge sources. A speaker independent recognition module generates, for each utterance, a lattice of word hypotheses which is the interface to an understanding module that performs the syntactic and semantic analysis. The recognition module is based on Hidden Markov Models of subword units, and performs the acoustic decoding process according to a beam search strategy. The understanding module finds the most likely sequence of words and represents its meaning in a format which facilitates the access to a database. It makes use of a modified caseframe analysis guided by the word hypotheses scores. Experiments were performed with 600 sentences from 10 speakers on the E-Mail application task. Using 15 Gaussian mixtures per state, a word accuracy of 75.7 was obtained with a test vocabulary of 787 words and no linguistic constraints. Linguistic processing of the corresponding lattices achieved a sentence understanding rate of 82%.
Paolo Baggia, Luciano Fissore, Egidio P. Giachin, Giorgio Micca, Claudio Rullent, Pietro Laface
Int. J. Pattern Recognit. Artif. Intell.6
1993 Analysis and improvement of the partial distance search algorithm
Lorenzo Fissore, Pietro Laface, P. Massafra, Franco Ravera
ICASSP (2)2
1993 Using grammars in forward and backward search
Lorenzo Fissore, Egidio P. Giachin, Pietro Laface, P. Massafra
EUROSPEECH3
1992 HMM modeling for speaker independent voice dialing in car environment
abstract
The authors describe the development of a speaker-independent isolated word recognizer for a voice dialing application operating in the car environment. Speaker-dependent and speaker-independent approaches are addressed and compared. Simple continuous hidden Markov models (HMMs) are used for speaker-dependent recognition; multiple codebook discrete and continuous HMMs are trained by speaker-independent reference data derived from a large database of speech collected inside several cars under a wide variety of driving conditions and by a large number of speakers from different Italian regions. By modeling separately two models (one for male and one for female speakers) for each word with 12 state continuous density whole word HMMs with eight diagonal covariance Gaussians per state, and performing a beam search Viterbi decoding a recognition rate of 99% has been obtained (65 errors out of 6423 words).>
Lorenzo Fissore, Pietro Laface, P. Ruscitti
ICASSP2
1992 Channel adaptation for a continuous speech recognizer
Lorenzo Fissore, Pietro Laface, Giorgio Micca, G. Sperto
ICSLP2
1991 Comparison of discrete and continuous HMMs in a CSR task over the telephone
abstract
Attention is given to a comparison of the performance of discrete and continuous density hidden Markov models (DDHMMs and CDHMMs) on a 786-word E-mail inquiry task performed by the speaker-independent word recognition component of a speech understanding system. This comparison between DDHMMs and CDHMMs has also been carried out by training speaker-dependent models. The authors also present the results of a set of experiments carried out with the aim at automatically selecting a suitable set of subword unit models by a clustering procedure. The recognizer gives word accuracy (WA) rates of 67.8% and 75.3% by using DDHMMs and CDHMMs, respectively, without any linguistic constraints. On the same task, WA rates of 87.1% and 85.9% have been obtained in the speaker-dependent mode.>
Lorenzo Fissore, Pietro Laface, Giorgio Micca
ICASSP2
1991 Selection of speech units for a speaker-independent CSR task
Lorenzo Fissore, Egidio P. Giachin, Pietro Laface, Giorgio Micca
EUROSPEECH3
1990 HMM modeling for voice-activated mobile-radio system
Luciano Fissore, Pietro Laface, M. Codogno, Giovanni Venuti
ICSLP2
1989 On the use of neural networks for speaker independent isolated word recognition
abstract
The authors present results obtained by applying the connectionist approach of multilayer perceptrons (MLPs) to three tasks of practical interest: classification of speech in terms of broad phonetic classes, speaker-independent recognition of yes/no answers through the dialed-up telephone line, and speaker-independent recognition of isolated digits through the telephone line. The first task assesses the capability of a simple MLP to generate nonlinear decision surfaces that discriminate among six broad phonetic classes. The MLP performance is actually comparable to that obtained by a hierarchical polynomial classifier. The second task deals with the sequential nature of speech. As short words like SI/NO do not give relevant problems of time alignment, the effects of different parts of the signal are taken into account by means of hidden units. A 98% recognition rate is achieved. For the third task, digital recognition, where the length of the words has a large range variation, a nonlinear time alignment is used that is performed through trace segmentation.>
Piero Demichelis, L. Fissore, Pietro Laface, Giorgio Micca, E. Piccolo
ICASSP3
1989 A word hypothesizer for a large vocabulary continuous speech understanding system
abstract
A continuous-speech recognition and understanding system for a thousand-word vocabulary has been designed and implemented. It is able to answer queries put to a geographical database in natural Italian language. A discussion is presented of the recognition component of the system. It can produce a word lattice that is then processed by a syntactic-semantic component. In addition, a linguistic decoder exploiting word-pair constraints has been investigated. Its results have been compared to those obtained by similar approaches reported in the literature. The system relies on word preselection through lexical access by means of broad phonetic classes and on hidden Markov modeling of subword units. The improvements to the basic approach are presented and system performance is given. Average word accuracy and correct sentence recognition obtained for speaker-dependent tests performed by two speakers pronouncing 214 sentences are 94.5% and 89.3%, respectively. The perplexity of the word-pair language model is 25.>
L. Fissore, Pietro Laface, Giorgio Micca, Roberto Pieraccini
ICASSP2
1988 Experimental results on large-vocabulary continuous speech recognition and understanding
abstract
A continuous speech recognition and understanding system is presented that accepts queries about a restricted geographical domain, expressed in free but syntactically correct natural language, with a lexicon of the order of one thousand words. A lattice of word candidates hypothesized by the speaker dependent recognition level is the interface to an understanding module that performs the syntactic and semantic analysis. The recognition subsystem generates word hypotheses by exploiting hidden Markov models of sub-word units. Bottom-up constraints are also introduced to restrict the set of candidate words. The understanding module determines the most likely sequence of words and represents its meaning in a parse-tree suitable to access a database. It makes use of a modified caseframe analysis driven by the word hypotheses likelihood scores. The results of a set of experiments performed in 150 sentences collected from one speaker are given.>
L. Fissore, Egidio P. Giachin, Pietro Laface, Giorgio Micca, Roberto Pieraccini, Claudio Rullent
ICASSP3
1988 Very large vocabulary isolated utterance recognition: a comparison between one pass and two pass strategies
abstract
A system for recognizing isolated utterances belonging to a very large vocabulary is presented that follows a two-pass strategy. The first step, hypothesization, consists in the selection of a subset of word candidates, starting from the segmentation of speech into six broad phonetic classes. This module is implemented through a dynamic programming algorithm working in a three-dimensional space. The search is performed on a tree representing a coarse description of the lexicon. The second step is the search for the best N candidates according to a maximum-likelihood criterion. Each word candidate is represented by a graph of subword hidden Markov models, and a tree structure of the whole word subset is built on line for an efficient implementation of the Viterbi algorithm. A comparison with a direct approach that does not use the hypothesization module shows that the two-pass approach has the same performance with an 80% reduction in computational complexity.>
L. Fissore, Pietro Laface, Giorgio Micca, Roberto Pieraccini
ICASSP2
1988 Interaction between fast lexical access and word verification in large vocabulary continuous speech recognition
abstract
Recently a two step strategy for large vocabulary isolated word recognition has been successfully experimented. The first step consists in the hypothesization of a reduced set of word candidates on the basis of broad bottom-up features, while the second one is the verification of the hypotheses using more detailed phonetic knowledge. This paper deals with its extension to continuous speech. A tight integration between the two steps rather than a hierarchical approach has been investigated. The hypothesization and the verification modules are implemented as processes running in parallel. Both processes represent lexical knowledge by a tree. Each node of the hypothesization tree is labeled by one of 6 broad phonetic classes. The nodes of the verification tree are, instead, the states of sub-word HMMs. The two processes cooperate to detect word hypotheses along the sentence.>
L. Fissore, Pietro Laface, Giorgio Micca, Roberto Pieraccini
ICASSP2
1987 Experimental results on a large lexicon access task
abstract
In this paper the problem of lexical access to large vocabularies by means of a coarse phonetic description of words is addressed. A generate and test approach is used: first a set of word candidates is extracted from the lexicon by means of a broad phonetic description of the input utterance, then a more detailed stochastic model of each word in this set, based on sub-word phonetic units, is obtained, and the likelihood of the candidate words is estimated using the Viterbi algorithm. Results of the application of the method to a large vocabulary isolated word recognition task are given. The candidate lists produced in the generation phase include the correct word in 98 times out of 100, their average size is of the order of 50 items for a 1011 word lexicon, while they do not exceed 300 units for a 13748 word lexicon.
Pietro Laface, Giorgio Micca, Roberto Pieraccini
ICASSP1
1986 Discrimination of Words in a Large Vocabulary Using Phonetic Descriptions
Attilio Giordana, Lorenza Saitta, Pietro Laface
Int. J. Man Mach. Stud.3
1985 Parallel Algorithms for Syllable Recognition in Continuous Speech
abstract
A distributed rule-based system for automatic speech recognition is described. Acoustic property extraction and feature hypothesization are performed by the application of sequences of operators. These sequences, called plans, are executed by cooperative expert programs. Experimental results on the automatic segmentation and recognition of phrases, made of connected letters and digits, are described and discussed.
Renato De Mori, Pietro Laface, Florence Yu Mong
IEEE Trans. Pattern Anal. Mach. Intell.2
1984 An expert system for mapping acoustic cues into phonetic features
Renato De Mori, Attilio Giordana, Pietro Laface, Lorenza Saitta
Inf. Sci.3
1983 Phonetic feature hypothesization in continuous speech
abstract
An Expert system is introduced for extracting acoustic cues from continuous speech. Part of the knowledge of such a system is a semantic Syntax-Directed Translation algorithm that segments continuous speech into Pseudo-Syllabic segments and generates hypotheses about phonetic features in each segment. Experimental results are provided about the performances of this system.
Renato De Mori, Attilio Giordana, Pietro Laface
ICASSP3
1982 An Expert System for Interpreting Speech Patterns
Renato De Mori, Attilio Giordana, Lorenza Saitta, Pietro Laface
AAAI4
1982 An Expert System for Speech Decoding
Renato De Mori, Attilio Giordana, Pietro Laface, Lorenza Saitta
ECAI3
1982 MODOSK: A modular distributed operating system kernel for real-time process control
Patricia Garetti, Pietro Laface, Silvano Rivoira
Microprocessing and Microprogramming2
1982 Speech segmentation and interpretation using a semantic syntax-directed translation
Renato De Mori, Attilio Giordana, Pietro Laface
Pattern Recognit. Lett.3
1980 Use of Fuzzy Algorithms for Phonetic and Phonemic Labeling of Continuous Speech
abstract
A model for assigning phonetic and phonemic labels to speech segments is presented. The system executes fuzzy algorithms that assign degrees of worthiness to structured interpretations of syllabic segments extracted from the signal of a spoken sentence. The knowledge source is a series of syntactic rules whose syntactic categories are phonetic and phonemic features detected by a precategorical and a categorical classification of speech sounds. Rules inferred from experiments and results for male and female voices are presented.
Renato De Mori, Pietro Laface
IEEE Trans. Pattern Anal. Mach. Intell.2
1979 Computer recognition of stop consonants
abstract
The paper describes a computer program for the automatic recognition of stop consonants in continuous speech. The recognition is performed by a fuzzy algorithm that accounts for the imprecision of the features extracted and of the rules. The rules belong to a fuzzy grammar and account for coarticulation and contextual effects.
Piero Demichelis, Renato De Mori, Pietro Laface, Mary O'Kane
ICASSP3
1977 A syntactic procedure for the recognition of glottal pulses in continuous speech
Renato De Mori, Pietro Laface, V. A. Makhonine, Marco Mezzalama
Pattern Recognit.2