VLDB 2026 Research / reviewers in the wild / expert
Sandro Cumani
dblp:15/8459
· DBLP profile ↗
48ranked-venue papers
31as first author
9since 2021 · last 2025
0000-0001-6036-0065ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 23 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 17 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Analysis of ABC Frontend Audio Systems for the NIST-SRE24abstractSection: Speaker Recognition Sara Barahona, Anna Silnova, Ladislav Mosner, Junyi Peng, Oldrich Plchot, Johan Rohdin, Lin Zhang 0054, Jiangyu Han, Petr Pálka, Federico Landini, Lukás Burget, Themos Stafylakis, Sandro Cumani, Dominik Bobos, Miroslav Hlavácek, Martin Kodovsky, Tomás Pavlícek |
INTERSPEECH | 13 |
| 2025 | A Copula-Based Generative Score-Level Fusion Model for Speaker Verification
Sandro Cumani |
INTERSPEECH | 1 |
| 2025 | Analysis of the ABC Classification Backends for NIST SRE24
Sandro Cumani, Anna Silnova, Sara Barahona, Ladislav Mosner, Oldrich Plchot, Johan Rohdin |
INTERSPEECH | 1 |
| 2024 | Towards Comprehensive Subgroup Performance Analysis in Speech ModelsabstractThe evaluation of spoken language understanding (SLU) systems is often restricted to assessing their global performance or examining predefined subgroups of interest. However, a more detailed analysis at the subgroup level has the potential to uncover valuable insights into how speech system performance differs across various subgroups. In this work, we identify biased data subgroups and describe them at the level of user demographics, recording conditions, and speech targets. We propose a new task-, model- and dataset-agnostic approach to detect significant intra- and cross-model performance gaps. We detect problematic data subgroups in SLU models by leveraging the notion of subgroup divergence. We also compare the outcome of different SLU models on the same dataset and task at the subgroup level. We identify significant gaps in subgroup performance between models different in size, architecture, or pre-training objectives, including multi-lingual and mono-lingual models, yet comparable to each other in overall performance. The results, obtained on two SLU models, four datasets, and three different tasks–intent classification, automatic speech recognition, and emotion recognition–confirm the effectiveness of the proposed approach in providing a nuanced SLU model assessment. Alkis Koudounas, Eliana Pastor, Giuseppe Attanasio, Vittorio Mazzia, Manuel Giollo, Thomas Gueudré, Elisa Reale, Luca Cagliero, Sandro Cumani, Luca de Alfaro, Elena Baralis, Daniele Amberti |
IEEE ACM Trans. Audio Speech Lang. Process. | 9 |
| 2023 | From adaptive score normalization to adaptive data normalization for speaker verification systems
Sandro Cumani, Salvatore Sarni |
INTERSPEECH | 1 |
| 2023 | Description and analysis of the KPT system for NIST Language Recognition Evaluation 2022abstractThis paper presents an analysis of the KPT system for the 2022 NIST Language Recognition Evaluation. The KPT submission focuses on the fixed training condition where only specific speech data can be used to develop all the modules and auxiliary systems used to build the language recognizer. Our solution consists of several sub-systems based on different neural network front-ends and a common back-end for classification and fusion. The goal of each front-end is to extract language-related embeddings. Gaussian linear models are used to classify the embeddings of each front-end, followed by multi-class logistic regression to calibrate and fuse the different sub-systems. Experimental results from the NIST LRE 2022 evaluation task show that our approach achieves competitive performance. Salvatore Sarni, Sandro Cumani, Sabato Marco Siniscalchi, Andrea Bottino |
INTERSPEECH | 2 |
| 2023 | The Distributions of Uncalibrated Speaker Verification Scores: A Generative Model for Domain Mismatch and Trial-Dependent CalibrationabstractSpeaker verification systems that compute log-likelihood ratios (LLR) between the same and different speaker hypotheses allow for cost-effective decisions that depend only on prior information. Domain mismatch, inaccurate model assumptions or the intrinsic nature of non-probabilistic classifiers often result in mis-calibrated scores, and a re-calibration step is required to map the classifier outputs to well-calibrated LLRs. Standard calibration is based on Logistic Regression, often paired with quality measures to provide trial-dependent calibration transformations. More recently, generative methods have been proposed as an alternative to discriminative approaches, which, however, are not yet able to exploit additional side information. In this work we introduce a novel generative approach based on the analysis of the effects of speaker vector distribution mismatch on the distribution of verification scores for PLDA and PLDA-based classifiers. We show that target and non-target scores can be modeled by Variance-Gamma distributions, whose parameters represent effective between and within-class variability. This allows us to introduce utterance-dependent variability models that can incorporate both explicit quality measures, such as the utterance duration, or implicit measures, such as the norm of a speaker embedding. Experimental results on different test sets with different front-ends and classifiers show that the proposed approach improves both calibration and verification accuracy with respect to state-of-the-art calibration models. Sandro Cumani, Salvatore Sarni |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | A Generative Model for Duration-Dependent Score Calibration
Sandro Cumani, Salvatore Sarni |
Interspeech | 1 |
| 2021 | On the Distribution of Speaker Verification Scores: Generative Models for Unsupervised CalibrationabstractSpeaker verification systems whose outputs can be interpreted as log-likelihood ratios (LLR) allow for cost-effective decisions by comparing the system outputs to application-defined thresholds depending only on prior information. Classifiers often produce uncalibrated scores, and require additional processing to produce well-calibrated LLRs. Recently, generative score calibration models have been proposed, which achieve calibration performance close to that of state-of-the-art discriminative techniques for supervised scenarios, while also allowing for unsupervised training. The effectiveness of these methods, however, strongly depends on their capabilities to correctly model the target and non-target score distributions. In this work we propose theoretically grounded and accurate models for characterizing the distribution of scores of speaker verification systems. Our approach is based on tied Generalized Hyperbolic distributions and overcomes many limitations of Gaussian models. Experimental results on different NIST benchmarks, using different utterance representation front-ends and different back-end classifiers, show that our method is effective not only in supervised scenarios, but also in unsupervised tasks characterized by very low proportion of target trials. Sandro Cumani |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Tied Normal Variance-Mean Mixtures for Linear Score CalibrationabstractA speaker verification system decides whether two voice segments belong to the same speaker based on a threshold. An optimal threshold can be set if the recognition scores are well calibrated, i.e., they represent Log-Likelihood Ratios. Logistic Regression (LogReg) is a standard approach for score calibration. While training this discriminative model requires labeled scores, Gaussian and non-Gaussian generative calibration models have been recently proposed. They not only have similar or better performance with respect to LogReg, but also allow for unsupervised or semi-supervised training of the models.The goal of this work is to extend these models. In particular, we show that normal variance-mean mixture distributions are able to model well-calibrated non-Gaussian distributed scores, provided that their parameters for the target and non-target score distributions are properly tied. As for the Gaussian case, a linear calibration model can then be estimated by computing Maximum Likelihood estimates of the distributions parameters and of the score transformation. The quality of all these approaches has been compared on a dataset of segments of variable duration obtained by cutting the NIST 2010 evaluation test data. Sandro Cumani, Pietro Laface |
ICASSP | 1 |
| 2019 | Normal Variance-Mean Mixtures for Unsupervised Score Calibration
Sandro Cumani |
INTERSPEECH | 1 |
| 2019 | Exact memory-constrained UPGMA for large scale speaker clustering
Sandro Cumani, Pietro Laface |
Pattern Recognit. | 1 |
| 2018 | Speaker Recognition Using e-VectorsabstractSystems based on i-vectors represent the current state-of-the-art in text-independent speaker recognition. Unlike joint factor analysis (JFA), which models both speaker and intersession subspaces separately, in the i-vector approach all the important variability is modeled in a single low-dimensional subspace. This paper is based on the observation that JFA estimates a more informative speaker subspace than the “total variability” i- vector subspace, because the latter is obtained by considering each training segment as belonging to a different speaker. We propose a speaker modeling approach that extracts a compact representation of a speech segment, similar to the speaker factors of JFA and to i-vectors, referred to as “e-vector.” Estimating the e-vector subspace follows a procedure similar to i-vector training, but produces a more accurate speaker subspace, as confirmed by the results of a set of tests performed on the NIST 2012 and 2010 Speaker Recognition Evaluations. Simply replacing the i-vectors with e-vectors we get approximately 10% average improvement of the Cprimarycost function, using different systems and classifiers. It is worth noting that these performance gains come without any additional memory or computational costs with respect to the standard i-vector systems. Sandro Cumani, Pietro Laface |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | Scoring Heterogeneous Speaker Vectors Using Nonlinear Transformations and Tied PLDA ModelsabstractMost current state-of-the-art text-independent speaker recognition systems are based on i-vectors, and on probabilistic linear discriminant analysis (PLDA). PLDA assumes that the i-vectors of a trial are homogeneous, i.e., that they have been extracted by the same system. In other words, the enrollment and test i-vectors belong to the same class. However, it is sometimes important to score trials including “heterogeneous” i-vectors, for instance, enrollment i-vectors extracted by an old system, and test i-vectors extracted by a newer, more accurate, system. In this paper, we introduce a PLDA model that is able to score heterogeneous i-vectors independent of their extraction approach, dimensions, and any other characteristics that make a set of i-vectors of the same speaker belong to different classes. The new model, which will be referred to as nonlinear tied-PLDA (NL-Tied-PLDA), is obtained by a generalization of our recently proposed nonlinear PLDA approach, which jointly estimates the PLDA parameters and the parameters of a nonlinear transformation of the i-vectors. The generalization consists of estimating a class-dependent nonlinear transformation of the i-vectors, with the constraint that the transformed i-vectors of the same speaker share the same speaker factor. The resulting model is flexible and accurate, as assessed by the results of a set of experiments performed on the extended core NIST SRE 2012 evaluation. In particular, NL-Tied-PLDA provides better results on heterogeneous trials with respect to the corresponding homogeneous trials scored by the old system, and, in some configurations, it also reaches the accuracy of the new system. Similar results were obtained on the female-extended core NIST SRE 2010 telephone condition. Sandro Cumani, Pietro Laface |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | e-vectors: JFA and i-vectors revisitedabstractSystems based on i-vectors represent the current state-of-the-art in text-independent speaker recognition. In this work we introduce a new compact representation of a speech segment, similar to the speaker factors of Joint Factor Analysis (JFA) and to i-vectors, that we call “e-vector”. The e-vectors derive their name from the eigen-voice space of the JFA speaker modeling approach. Our working hypothesis is that JFA estimates a more informative speaker subspace than the “total variability” i-vector subspace, because the latter is obtained by considering each training segment as belonging to a different speaker. We propose, thus, a simple “i-vector style” modeling and training technique that exploits this observation, and estimates a more accurate subspace with respect to the one provided by the classical i-vector approach, as confirmed by the results of a set of tests performed on the extended core NIST 2012 Speaker Recognition Evaluation dataset. Simply replacing the i-vectors with e-vectors we get approximately 10% average improvement of the Cprimarycost function, using different systems and classifiers. These performance gains come without any additional memory or computational costs with respect to the standard i-vector systems. Sandro Cumani, Pietro Laface |
ICASSP | 1 |
| 2017 | CNN Patch-Based Voting for Fingerprint Liveness DetectionabstractBiometric identification systems based on fingerprints are vulnerable to attacks that use fake replicas of real fingerprints. One possible countermeasure to this issue consists in developing software modules capable of telling the liveness of an input image and, thus, of discarding fakes prior to the recognition step. This paper presents a fingerprint liveness detection method founded on a patch-based voting approach. Fingerprint images are first segmented to discard background information. Then, small-sized foreground patches are extracted and processed by a well-know Convolutional Neural Network model adapted to the problem at hand. Finally, the patch scores are combined to draw the final fingerprint label. Experimental results on well-established benchmarks demonstrate a promising performance of the proposed method compared with several state-of-the-art algorithms. Amirhosein Toosi, Sandro Cumani, Andrea Bottino |
IJCCI | 2 |
| 2017 | Nuance - Politecnico di Torino's 2016 NIST Speaker Recognition Evaluation SystemabstractThis paper describes the Loquendo – Politecnico di Torino system evaluated on the 2006 NIST speaker recognition evaluation dataset. This system was among the best participants in this evaluation. It combines the results of two independent GMM systems: a Phonetic GMM and a classical GMM. Both systems rely on an intersession variation compensation approach, performed in the feature domain. It allowed a 30% error rate reduction with respect to our 2005 system. The linear combination of the two GMM engines gives a further 10% error rate reduction. We also report the results of a set of post evaluation experiments, related to the training data for the intersession variation evaluation, both for the telephone and microphone datasets. The approach adopted for the two wire tests is also described, showing the effect of the speaker segmentation component of our system. Finally, we describe how we performed the incremental unsupervised adaptation tests. Daniele Colibro, Claudio Vair, Emanuele Dalmasso, Kevin Farrell, Gennady Karvitsky, Sandro Cumani, Pietro Laface |
INTERSPEECH | 6 |
| 2017 | Nonlinear I-Vector Transformations for PLDA-Based Speaker RecognitionabstractThis paper proposes to estimate parametric nonlinear transformations of i-vectors for speaker recognition systems based on probabilistic linear discriminant analysis (PLDA) classification. The Gaussian PLDA model assumes that the i-vectors are distributed according to the standard normal distribution. However, it has been shown that the i-vectors are better modeled, for example, by Heavy-Tailed distributions, and that significant improvement of the classification performance can be obtained by whitening and length normalizing the i-vectors. In this paper, we propose to transform the i-vectors so that their distribution becomes more suitable to discriminate speakers using the PLDA model. This is performed by means of a sequence of affine and nonlinear transformations whose parameters are obtained by maximum likelihood estimation on the development set. Another contribution of this paper is the reduction of the mismatch between the development and evaluation i-vector length distributions by means of a scaling factor tuned for the estimated i-vector distribution, rather than by means of a blind length normalization. Relative improvement between 7% and 14% of the detection cost function was obtained with the proposed technique on the NIST SRE-2010 and SRE-2012 evaluation datasets, using both the traditional GMM/UBM and the hybrid DNN/GMM-based systems. Sandro Cumani, Pietro Laface |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | Joint Estimation of PLDA and Nonlinear Transformations of Speaker VectorsabstractThe Gaussian probabilistic linear discriminant anal-ysis (PLDA) model assumes Gaussian distributed priors for the latent variables that represent the speaker and channel factors. Assuming that each training i-vector belongs to a different speaker, as is usually done in i-vector extraction, i-vectors generated by a PLDA model can be considered independent and identically distributed with Gaussian distribution. Thus, we have recently proposed to transform the development i-vectors so that their distribution becomes more Gaussian-like. This is obtained by means of a sequence of affine and nonlinear transformations whose parameters are trained by maximum likelihood (ML) estimation on the development set. The evaluation i-vectors are then subject to the same transformation. Although the i-vector “gaussianization” has shown to be effective, since the i-vectors extracted from segments of the same speaker are not independent, the original assumption is not satisfactory. In this work, we show that the model can be improved by properly exploiting the information about the speaker labels, which was ignored in the previous model. In particular, a more effective PLDA model can be obtained by jointly estimating the PLDA parameters and the parameters of the nonlinear transformation of the i-vectors. In other words, while the goal of the previous approach was to “gaussianize” the training i-vectors distribution, the objective of this work is to embed the estimation of the nonlinear i-vector transformation in the PLDA model estimation. We will thus refer to this model as the nonlinear PLDA model. We show that this new approach provides significant gain with respect to PLDA, and a small, yet consistent, improvement with respect to our former i-vector “gaussianization” approach, without further additional costs. Sandro Cumani, Pietro Laface |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2016 | Discriminative multi-domain PLDA for speaker verificationabstractDomain mismatch occurs when data from application-specific target domain is related to, but cannot be viewed as iid samples from the source domain used for training speaker models. Another problem occurs when several training datasets are available but their domains differ. In this case training on simply merged subsets can lead to suboptimal performance. Existing approaches to cope with these problems employ generative modeling and consist of several separate stages such as training and adaptation. In this work we explore a discriminative approach which naturally incorporates both scenarios in a principled way. To this end, we develop a method that can learn across multiple domains by extending discriminative probabilistic linear discriminant analysis (PLDA) according to multi-task learning paradigm. Our results on the recent JHU Domain Adaptation Challenge (DAC) dataset demonstrate that the proposed multi-task PLDA decreases equal error rate (EER) of the PLDA without domain compensation by more than 35% relative and performs comparable to another competitive domain compensation technique. Alexey Sholokhov, Tomi Kinnunen, Sandro Cumani |
ICASSP | 3 |
| 2015 | On Multiview Analysis for Fingerprint Liveness Detection
Amirhosein Toosi, Sandro Cumani, Andrea Bottino |
CIARP | 2 |
| 2015 | Memory-aware i-vector extraction by means of sub-space factorizationabstractMost of the state-of-the-art speaker recognition systems use i-vectors, a compact representation of spoken utterances. Since the “standard” i-vector extraction procedure requires large memory structures, we recently presented the Factorized Sub-space Estimation (FSE) approach, an efficient technique that dramatically reduces the memory needs for i-vector extraction, and is also fast and accurate compared to other proposed approaches. FSE is based on the approximation of the matrix T, representing the speaker variability sub-space, by means of the product of appropriately designed matrices. In this work, we introduce and evaluate a further approximation of the matrices that most contribute to the memory costs in the FSE approach, showing that it is possible to obtain comparable system accuracy using less than a half of FSE memory, which corresponds to more than 60 times memory reduction with respect to the standard method of i-vector extraction. Sandro Cumani, Pietro Laface |
ICASSP | 1 |
| 2015 | Speaker recognition by means of acoustic and phonetically informed GMMsabstractIn this work we assess the recently proposed hybrid Deep Neural Network/Gaussian Mixture Model (DNN/GMM) approach for speaker recognition considering the effects of the granularity of the phonetic DNN model, and of the precision of the corresponding GMM models, which will be referred to as the phonetic GMMs. The aim of this work is to better understand the contributions of the phonetic information provided by the DNN model with respect to the accuracy of the acous tic GMMs in fitting the distribution of the features associated to a given context-dependent phone state. The testbed for this work was the text-independent speaker recognition task defined by NIST for the 2012 Speaker Recognition Evaluation. Our experiment confirms that the acoustic and the phonetic GMMs are complementary. Thus, their score combination yields very good results if the DNN is trained on data collected in an environment similar to the one that is used for testing. We show, however, that using a single Gaussian per DNN state is not the best choice: the best single system has been obtained balancing the phonetic and acoustic precision of a DNN/GMM system Sandro Cumani, Pietro Laface, Farzana Kulsoom |
INTERSPEECH | 1 |
| 2015 | Exploiting i-vector posterior covariances for short-duration language recognitionabstractLinear models in i-vector space have shown to be an effective solution not only for speaker identification, but also for language recogniton. The i-vector extraction process, however, is affected by several factors, such as noise level, the acoustic content of the utterance and the duration of the spoken segments. These factors influence both the i-vector estimate and its uncertainty, represented by the i-vector posterior covariance matrix. Modeling of i-vector uncertainty with Probabilistic Linear Discriminant Analysis has shown to be effective for short-duration speaker identification. This paper extends the approach to language recognition, analyzing the effects of i-vector covariances on a state-of-the-art Gaussian classifier, and proposes an effective solution for the reduction of the average detection cost (Cavg) for short segments. Sandro Cumani, Oldrich Plchot, Radek Fér |
INTERSPEECH | 1 |
| 2015 | Fast Scoring of Full Posterior PLDA ModelsabstractA low-dimensional representation of a speech segment, the so-called i-vector, in combination with probabilistic linear discriminant analysis (PLDA) models, is the current state-of-the-art in speaker recognition. An i-vector is a compact representation of a Gaussian Mixture Model (GMM) supervector, which captures most of the GMM supervectors variability. It is usually obtained by a MAP estimate of the mean of a posterior distribution. A new PLDA model has been recently presented that, unlike the standard one, exploits the intrinsic i-vector uncertainty. This approach, referred to in this paper as Full Posterior Distribution PLDA (FP-PLDA), is particularly effective for speaker detection of short and variable duration speech segments. It is, however, computationally far more expensive than standard PLDA, making it unattractive for real applications. This paper presents three simplifications of FP-PLDA based on approximate diagonalizations of matrices involved in FP-PLDA scoring. Using in sequence these approximations allows obtaining computational costs comparable to PLDA models, with only a small performance degradation with respect to the more accurate, but less efficient, FP-PLDA models. In particular, up to 10% better performance than PLDA is obtained, with similar computational complexity, on short speech segments of variable duration, randomly extracted from the interviews and telephone conversations included in the NIST SRE 2010 extended dataset. The benefits of the proposed diagonalization approaches have also been confirmed on a short utterance text-independent verification task, where approximately 43% and 34% improvement of the EER and minimum DCF08, respectively, has been obtained with respect to PLDA. Sandro Cumani |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | Training Pairwise Support Vector Machines with large scale datasetsabstractWe recently presented an efficient approach for training a Pairwise Support Vector Machine (PSVM) with a suitable kernel for a quite large speaker recognition task. The PSVM approach, rather than estimating an SVM model per class according to the “one versus all” discriminative paradigm, classifies pairs of examples as belonging or not to the same class. Training a PSVM with large amount of data, however, is a memory and computational expensive task, because the number of training pairs grows quadratically with the number of training patterns. This paper proposes an approach that allows discarding the training pairs that do not essentially contribute to the set of Support Vectors (SVs) of the training set. This selection of training pairs is feasible because we show that the number of SVs does not grow quadratically, with the number of pairs, but only linearly with the number of speakers in the training set. Our approach dramatically reduces the memory and computational complexity of PSVM training, making possible the use of large datasets, including many speakers. It has been assessed on the extended core conditions of the 2012 Speaker Recognition Evaluation. The results show that the accuracy of the trained PSVMs increases with the training set size, and that the Cprimaryof a PSVM trained with a small subset of the i-vectors pairs is 10-30% better than the one obtained by a generative model trained on the complete set of i-vectors. Sandro Cumani, Pietro Laface |
ICASSP | 1 |
| 2014 | Factorized Sub-Space Estimation for Fast and Memory Effective I-vector ExtractionabstractMost of the state–of–the–art speaker recognition systems use a compact representation of spoken utterances referred to as i–vector. Since the “standard” i–vector extraction procedure requires large memory structures and is relatively slow, new approaches have recently been proposed that are able to obtain either accurate solutions at the expense of an increase of the computational load, or fast approximate solutions, which are traded for lower memory costs. We propose a new approach particularly useful for applications that need to minimize their memory requirements. Our solution not only dramatically reduces the memory needs for i–vector extraction, but is also fast and accurate compared to recently proposed approaches. Tested on the female part of the tel-tel extended NIST 2010 evaluation trials, our approach substantially improves the performance with respect to the fastest but inaccurate eigen-decomposition approach, using much less memory than other methods. Sandro Cumani, Pietro Laface |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | Large-Scale Training of Pairwise Support Vector Machines for Speaker RecognitionabstractState–of–the–art systems for text–independent speaker recognition use as their features a compact representation of a speaker utterance, known as “i–vector.” We recently presented an efficient approach for training a Pairwise Support Vector Machine (PSVM) with a suitable kernel for i–vector pairs for a quite large speaker recognition task. Rather than estimating an SVM model per speaker, according to the “one versus all” discriminative paradigm, the PSVM approach classifies a trial, consisting of a pair of i–vectors, as belonging or not to the same speaker class. Training a PSVM with large amount of data, however, is a memory and computational expensive task, because the number of training pairs grows quadratically with the number of training i–vectors. This paper demonstrates that a very small subset of the training pairs is necessary to train the original PSVM model, and proposes two approaches that allow discarding most of the training pairs that are not essential, without harming the accuracy of the model. This allows dramatically reducing the memory and computational resources needed for training, which becomes feasible with large datasets including many speakers. We have assessed these approaches on the extended core conditions of the NIST 2012 Speaker Recognition Evaluation. Our results show that the accuracy of the PSVM trained with a sufficient number of speakers is 10%-30% better compared to the one obtained by a PLDA model, depending on the testing conditions. Since the PSVM accuracy increases with the training set size, but PSVM training does not scale well for large numbers of speakers, our selection techniques become relevant for training accurate discriminative classifiers. Sandro Cumani, Pietro Laface |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | On the use of i-vector posterior distributions in Probabilistic Linear Discriminant AnalysisabstractThe i-vector extraction process is affected by several factors such as the noise level, the acoustic content of the observed features, the channel mismatch between the training conditions and the test data, and the duration of the analyzed speech segment. These factors influence both the i–vector estimate and its uncertainty, represented by the i–vector posterior covariance. This paper presents a new PLDA model that, unlike the standard one, exploits the intrinsic i–vector uncertainty. Since the recognition accuracy is known to decrease for short speech segments, and their length is one of the main factors affecting the i–vector covariance, we designed a set of experiments aiming at comparing the standard and the new PLDA models on short speech cuts of variable duration, randomly extracted from the conversations included in the NIST SRE 2010 extended dataset, both from interviews and telephone conversations. Our results on NIST SRE 2010 evaluation data show that in different conditions the new model outperforms the standard PLDA by more than 10% relative when tested on short segments with duration mismatches, and is able to keep the accuracy of the standard model for long enough speaker segments. This technique has also been successfully tested in the NIST SRE 2012 evaluation. Sandro Cumani, Oldrich Plchot, Pietro Laface |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2013 | Probabilistic linear discriminant analysis of i-vector posterior distributionsabstractThe i-vector extraction process is affected by several factors such as the noise level, the acoustic content of the observed features, and the duration of the analyzed speech segment. These factors influence both the i-vector estimate and its uncertainty, represented by the i-vector posterior covariance. This paper present a new PLDA model that, unlike the standard one, exploits the intrinsic i-vector uncertainty. Since short segments are known to decrease recognition accuracy, and segment duration is the main factor affecting the i-vector covariance, we designed a set of experiments aiming at comparing the standard and the new PLDA models on short speech cuts of variable duration, randomly extracted from the conversations included in the NIST SRE 2010 female telephone extended core condition. Our results show that the new model outperforms the standard PLDA when tested on short segments, and keeps the accuracy of the latter for long enough utterances. In particular, the relative improvement is up to 13% for the EER, 5% for DCF08, and 2.5% for DCF10. Sandro Cumani, Oldrich Plchot, Pietro Laface |
ICASSP | 1 |
| 2013 | Developing a speaker identification system for the DARPA RATS projectabstractThis paper describes the speaker identification (SID) system developed by the Patrol team for the first phase of the DARPA RATS (Robust Automatic Transcription of Speech) program, which seeks to advance state of the art detection capabilities on audio from highly degraded communication channels. We present results using multiple SID systems differing mainly in the algorithm used for voice activity detection (VAD) and feature extraction. We show that (a) unsupervised VAD performs as well supervised methods in terms of downstream SID performance, (b) noise-robust feature extraction methods such as CFCCs out-perform MFCC front-ends on noisy audio, and (c) fusion of multiple systems provides 24% relative improvement in EER compared to the single best system when using a novel SVM-based fusion algorithm that uses side information such as gender, language, and channel id. Oldrich Plchot, Spyridon Matsoukas, Pavel Matejka, Najim Dehak, Jeff Z. Ma, Sandro Cumani, Ondrej Glembek, Hynek Hermansky, Sri Harish Reddy Mallidi, Nima Mesgarani, Richard M. Schwartz, Mehdi Soufifar, Zheng-Hua Tan, Samuel Thomas 0001, Bing Zhang 0004, Xinhui Zhou |
ICASSP | 6 |
| 2013 | Nuance - Politecnico di torino's 2012 NIST speaker recognition evaluation systemabstractThis paper describes the Nuance-Politecnico di Torino (NPT) speaker recognition system submitted to the NIST SRE12 evaluation campaign. Included are the results of postevaluation tests, focusing on the analysis of the effects of score normalization and condition-dependent calibration. The submitted system combines the results of five acoustic recognizers all based on Gaussian Mixture Models (GMMs). Each system has its own front end, with features differing by their type and dimension. We illustrate the process of development data selection and configuration of state-of-the-art technology, which contributed to obtaining good performance in all the test conditions proposed in this evaluation. Daniele Colibro, Claudio Vair, Kevin Farrell, Nir Krause, Gennady Karvitsky, Sandro Cumani, Pietro Laface |
INTERSPEECH | 6 |
| 2013 | Fast and memory effective i-vector extraction using a factorized sub-spaceabstractMost of the state-of-the-art speaker recognition systems use a compact representation of spoken utterances referred to as i-vectors. Since the "standard" i-vector extraction procedure requires large memory structures and is relatively slow, new approaches have recently been proposed that are able to obtain either accurate solutions at the expense of an increase of the computational load, or fast approximate solutions, which are traded for lower memory costs. We propose a new approach particularly useful for applications that need to minimize their memory requirements. Our solution not only dramatically reduces the storage needs for i-vector extraction, but is also fast. Tested on the female part of the tel-tel extended NIST 2010 evaluation trials, our approach substantially improves the performance with respect to the fastest but inaccurate eigen-decomposition approach, using much less memory than any other known method. Sandro Cumani, Pietro Laface |
INTERSPEECH | 1 |
| 2013 | Regularized subspace n-gram model for phonotactic ivector extractionabstractPhonotactic language identification (LID) by means of n-gram statistics and discriminative classifiers is a popular approach for the LID problem. Low-dimensional representation of the n-gram statistics leads to the use of more diverse and efficient machine learning techniques in the LID. Recently, we proposed phototactic iVector as a low-dimensional representation of the n-gram statistics. In this work, an enhanced modeling of the n-gram probabilities along with regularized parameter estimation is proposed. The proposed model consistently improves the LID system performance over all conditions up to 15% relative to the previous state of the art system. The new model also alleviates memory requirement of the iVector extraction and helps to speed up subspace training. Results are presented in terms of Cavg over NIST LRE2009 evaluation set. Mehdi Soufifar, Lukás Burget, Oldrich Plchot, Sandro Cumani, Jan Cernocký |
INTERSPEECH | 4 |
| 2013 | Pairwise Discriminative Speaker Verification in the 𝕀-Vector SpaceabstractThis work presents a new and efficient approach to discriminative speaker verification in the${\rm i}$–vector space. We illustrate the development of a linear discriminative classifier that is trained to discriminate between the hypothesis that a pair of feature vectors in a trial belong to the same speaker or to different speakers. This approach is alternative to the usual discriminative setup that discriminates between a speaker and all the other speakers. We use a discriminative classifier based on a Support Vector Machine (SVM) that is trained to estimate the parameters of a symmetric quadratic function approximating a log–likelihood ratio score without explicit modeling of the${\rm i}$–vector distributions as in the generative Probabilistic Linear Discriminant Analysis (PLDA) models. Training these models is feasible because it is not necessary to expand the${\rm i}$–vector pairs, which would be expensive or even impossible even for medium sized training sets. The results of experiments performed on the tel-tel extended core condition of the NIST 2010 Speaker Recognition Evaluation are competitive with the ones obtained by generative models, in terms of normalized Detection Cost Function and Equal Error Rate. Moreover, we show that it is possible to train a gender–independent discriminative model that achieves state–of–the–art accuracy, comparable to the one of a gender–dependent system, saving memory and execution time both in training and in testing. Sandro Cumani, Niko Brümmer, Lukás Burget, Pietro Laface, Oldrich Plchot, Vasileios Vasilakakis |
IEEE Trans. Speech Audio Process. | 1 |
| 2013 | Memory and Computation Trade-Offs for Efficient I-Vector ExtractionabstractThis work aims at reducing the memory demand of the data structures that are usually pre-computed and stored for fast computation of the i-vectors, a compact representation of spoken utterances that is used by most state-of-the-art speaker recognition systems. We propose two new approaches allowing accurate i-vector extraction but requiring less memory, showing their relations with the standard computation method introduced for eigenvoices, and with the recently proposed fast eigen-decomposition technique. The first approach computes an i-vector in a Variational Bayes (VB) framework by iterating the estimation of one sub-block of i-vector elements at a time, keeping fixed all the others, and can obtain i-vectors as accurate as the ones obtained by the standard technique but requiring only 25% of its memory. The second technique is based on the Conjugate Gradient solution of a linear system, which is accurate and uses even less memory, but is slower than the VB approach. We analyze and compare the time and memory resources required by all these solutions, which are suited to different applications, and we show that it is possible to get accurate results greatly reducing memory demand compared with the standard solution at almost the same speed. Sandro Cumani, Pietro Laface |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | Gender independent discriminative speaker recognition in i-vector spaceabstractSpeaker recognition systems attain their best accuracy when trained with gender dependent features and tested with known gender trials. In real applications, however, gender labels are often not given. In this work we illustrate the design of a system that does not make use of the gender labels both in training and in test, i.e. a completely Gender Independent (GI) system. It relies on discriminative training, where the trials are i-vector pairs, and the discrimination is between the hypothesis that the pair of feature vectors in the trial belong to the same speaker or to different speakers. We demonstrate that this pairwise discriminative training can be interpreted as a procedure that estimates the parameters of the best (second order) approximation of the log-likelihood ratio score function, and that a pairwise SVM can be used for training a gender independent system. Our results show that a pairwise GI SVM, saving memory and execution time, achieves on the last NIST evaluations state-of-the-art performance, comparable to a Gender Dependent(GD) system. Sandro Cumani, Ondrej Glembek, Niko Brümmer, Edward de Villiers, Pietro Laface |
ICASSP | 1 |
| 2012 | Independent component analysis and MLLR transforms for speaker identificationabstractIn this paper, we explore the use of Independent Component Analysis (ICA) and Principal Component Analysis (PCA) techniques to reduce the dimensionality of high-level LVCSR features and at the same time to enable modelling them with state-of-the-art techniques like Probabilistic Linear Discriminant Analysis or Pairwise Support Vector Machines (PSVM). The high-level features are the coefficients from Constrained Maximum-Likelihood Linear Regression (CMLLR) and Maximum-Likelihood Linear Regression (MLLR) transforms estimated in an Automatic Speech Recognition (ASR) system. We also compare a classical approach of modeling every speaker by a single SVM classifier with the recent state-of-the-art modelling techniques in Speaker Identification. We report performance of the systems and score-level combination with a current state-of-the-art acoustic i-vector system on the NIST SRE2010 dataset. Sandro Cumani, Oldrich Plchot, Martin Karafiát |
ICASSP | 1 |
| 2012 | Discriminative classifiers for phonotactic language recognition with iVectorsabstractPhonotactic models based on bags of n-grams representations and discriminative classifiers are a popular approach to the language recognition problem. However, the large size of n-gram count vectors brings about some difficulties in discriminative classifiers. The subspace Multinomial model was recently proposed to effectively represent information contained in the n-grams using low-dimensional iVectors. The availability of a low-dimensional feature vector allows investigating different post-processing techniques and different classifiers to improve recognition performance. In this work, we analyze a set of discriminative classifiers based on Support Vector Machines and Logistic Regression and we propose an iVector post-processing technique which allows to improve recognition performance. The proposed systems are evaluated on the NIST LRE 2009 task. Mehdi Soufifar, Sandro Cumani, Lukás Burget, Jan Cernocký |
ICASSP | 2 |
| 2012 | Analysis of Large-Scale SVM Training Algorithms for Language and Speaker RecognitionabstractThis paper compares a set of large scale support vector machine (SVM) training algorithms for language and speaker recognition tasks. We analyze five approaches for training phonetic and acoustic SVM models for language recognition. We compare the performance of these approaches as a function of the training time required by each of them to reach convergence, and we discuss their scalability towards large corpora. Two of these algorithms can be used in speaker recognition to train a SVM that classifies pairs of utterances as either belonging to the same speaker or to two different speakers. Our results show that the accuracy of these algorithms is asymptotically equivalent, but they have different behavior with respect to the time required to converge. Some of these algorithms not only scale linearly with the training set size, but are also able to give their best results after just a few iterations. State-of-the-art performance has been obtained in the female subset of the NIST 2010 Speaker Recognition Evaluation extended core test using a single SVM system. Sandro Cumani, Pietro Laface |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | Discriminatively trained Probabilistic Linear Discriminant Analysis for speaker verificationabstractRecently, i-vector extraction and Probabilistic Linear Discriminant Analysis (PLDA) have proven to provide state-of-the-art speaker verification performance. In this paper, the speaker verification score for a pair of i-vectors representing a trial is computed with a functional form derived from the successful PLDA generative model. In our case, however, parameters of this function are estimated based on a discriminative training criterion. We propose to use the objective function to directly address the task in speaker verification: discrimination between same-speaker and different-speaker trials. Compared with a baseline which uses a generatively trained PLDA model, discriminative training provides up to 40% relative improvement on the NIST S RE 2010 evaluation task. Lukás Burget, Oldrich Plchot, Sandro Cumani, Ondrej Glembek, Pavel Matejka, Niko Brümmer |
ICASSP | 3 |
| 2011 | Loquendo - Politecnico di Torino's 2010 NIST speaker recognition evaluation systemabstractThis paper describes the improvement introduced in the Loquendo-Politecnico di Torino (LPT) speaker recognition system submitted to the NIST SRE10 evaluation campaign. This system combines the results of eight core acoustic systems all based on Gaussian Mixture Models (GMMs). We illustrate the key factors, in the selection of the development data and in engineering state-of-the art technology, which contributed to the very good performance and calibration of our system in all the test conditions proposed in this evaluation. Fabio Castaldo, Daniele Colibro, Claudio Vair, Sandro Cumani, Pietro Laface |
ICASSP | 4 |
| 2011 | Fast discriminative speaker verification in the i-vector spaceabstractThis work presents a new approach to discriminative speaker verification. Rather than estimating speaker models, or a model that discriminates between a speaker class and the class of all the other speakers, we directly solve the problem of classifying pairs of utterances as belonging to the same speaker or not. Sandro Cumani, Niko Brümmer, Lukás Burget, Pietro Laface |
ICASSP | 1 |
| 2011 | Comparison of Speaker Recognition Approaches for Real ApplicationsabstractThis paper describes the experimental setup and the results obtained using several state-of-the-art speaker recognition classifiers.The comparison of the different approaches aims at the development of real world applications, taking into account memory and computational constraints, and possible mismatches with respect to the training environment.The NIST SRE 2008 database has been considered our reference dataset, whereas nine commercially available databases of conversational speech in languages different form the ones used for developing the speaker recognition systems have been tested as representative of an application domain.Our results, evaluated on the two domains, show that the classifiers based on i-vectors obtain the best recognition and calibration accuracy.Gaussian PLDA and a recently introduced discriminative SVM together with an adaptive symmetric score normalization achieve the best performance using low memory and processing resources. Sandro Cumani, Pier Domenico Batzu, Daniele Colibro, Claudio Vair, Pietro Laface, Vasileios Vasilakakis |
INTERSPEECH | 1 |
| 2010 | Loquendo-Politecnico di Torino system for the 2009 NIST Language Recognition EvaluationabstractThis paper describes the system submitted by Loquendo and Politecnico di Torino (LPT) for the 2009 NIST Language Recognition Evaluation. The system is a combination of classifiers based on two core acoustic models and on two core phone tokenizers. It exploits several state-of-the-art techniques that have been successfully applied in recent years both in speaker and in language recognition. Fabio Castaldo, Daniele Colibro, Sandro Cumani, Emanuele Dalmasso, Pietro Laface, Claudio Vair |
ICASSP | 3 |
| 2010 | Parallel implementation of artificial neural network trainingabstractIn this paper we describe the implementation of a complete ANN training procedure for speech recognition using the block mode back-propagation learning algorithm. We exploit the high performance SIMD architecture of GPU using CUDA and its C-like language interface. We also compare the speed-up obtained implementing the training procedure only taking advantage of the multi-thread capabilities of multi-core processors. Our approach has been tested by training acoustic models for large vocabulary speech recognition tasks, showing a 6 times reduction of the time required to train real-world large size networks with respect to an already optimized implementation using the Intel MKL libraries. Stefano Scanzio, Sandro Cumani, Roberto Gemello, Franco Mana, Pietro Laface |
ICASSP | 2 |
| 2010 | Parallel implementation of Artificial Neural Network training for speech recognition
Stefano Scanzio, Sandro Cumani, Roberto Gemello, Franco Mana, Pietro Laface |
Pattern Recognit. Lett. | 2 |
| 2009 | Language recognition using language factorsabstractLanguage recognition systems based on acoustic models reach state of the art performance using discriminative training techniques.In speaker recognition, eigenvoice modeling of the speaker, and the use of speaker factors as input features to SVMs has recently been demonstrated to give good results compared to the standard GMM-SVM approach, which combines GMMs supervectors and SVMs.In this paper we propose, in analogy to the eigenvoice modeling approach, to estimate an eigen-language space, and to use the language factors as input features to SVM classifiers.Since language factors are low-dimension vectors, training and evaluating SVMs with different kernels and with large training examples becomes an easy task.This approach is demonstrated on the 14 languages of the NIST 2007 language recognition task, and shows performance improvements with respect to the standard GMM-SVM technique. Fabio Castaldo, Sandro Cumani, Pietro Laface, Daniele Colibro |
INTERSPEECH | 2 |