Niko Brümmer

dblp:85/1879 · also Niko Brummer · DBLP profile ↗
← Back
31ranked-venue papers
8as first author
4since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 6 first-author · 3 since 2021Artificial intelligence and machine learning · 24 · 7 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Speech recognition and synthesis · 91% Trustworthy machine learning · 9%
Computer graphics and multimedia
1 paper
Audio and music processing · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis › speaker recognition
i-vector
0.212013
Pairwise Discriminative Speaker Verification in the 𝕀-Vector Space · IEEE Trans. Speech Audio Process. 2013
Natural language and speech › Speech recognition and synthesis
speaker recognition
0.212013
Pairwise Discriminative Speaker Verification in the 𝕀-Vector Space · IEEE Trans. Speech Audio Process. 2013
Natural language and speech › Speech recognition and synthesis › speaker recognition
speaker verification
0.212013
Pairwise Discriminative Speaker Verification in the 𝕀-Vector Space · IEEE Trans. Speech Audio Process. 2013
Audio and music processing › speaker recognition
channel compensation
0.112007
Fusion of Heterogeneous Speaker Recognition Systems in the STBU Submission for the NIST Speaker Recognition Evaluation 2006 · IEEE Trans. Speech Audio Process. 2007
Audio and music processing
speaker recognition
0.112007
Fusion of Heterogeneous Speaker Recognition Systems in the STBU Submission for the NIST Speaker Recognition Evaluation 2006 · IEEE Trans. Speech Audio Process. 2007
Audio and music processing › speaker recognition
speaker verification
0.112007
Fusion of Heterogeneous Speaker Recognition Systems in the STBU Submission for the NIST Speaker Recognition Evaluation 2006 · IEEE Trans. Speech Audio Process. 2007
Machine learning › Trustworthy machine learning
pairwise classification
0.012013
Pairwise Discriminative Speaker Verification in the 𝕀-Vector Space · IEEE Trans. Speech Audio Process. 2013

Methods — techniques the papers use, named apart from their topics

support vector machine · 0.2symmetric quadratic function · 0.2logistic regression calibration · 0.1gaussian mixture model · 0.1eigenchannel adaptation · 0.1MLLR adaptation · 0.1
YearPublicationVenuePosition
2023 Toroidal Probabilistic Spherical Discriminant Analysis
abstract
In speaker recognition, where speech segments are mapped to embeddings on the unit hypersphere, two scoring back-ends are commonly used, namely cosine scoring and PLDA. We have recently proposed PSDA, an analog to PLDA that uses Von Mises-Fisher distributions instead of Gaussians. In this paper, we present toroidal PSDA (T-PSDA). It extends PSDA with the ability to model within and between-speaker variabilities in toroidal submanifolds of the hypersphere. Like PLDA and PSDA, the model allows closed-form scoring and closed-form EM updates for training. On VoxCeleb, we find T-PSDA accu-racy on par with cosine scoring, while PLDA accuracy is inferior. On NIST SRE’21 we find that T-PSDA gives large accuracy gains compared to both cosine scoring and PLDA.1
Anna Silnova, Niko Brümmer, Albert Swart, Lukás Burget
ICASSP2
2022 Probabilistic Spherical Discriminant Analysis: An Alternative to PLDA for length-normalized embeddings
abstract
In speaker recognition, where speech segments are mapped to embeddings on the unit hypersphere, two scoring backends are commonly used, namely cosine scoring or PLDA.Both have advantages and disadvantages, depending on the context.Cosine scoring follows naturally from the spherical geometry, but for PLDA the blessing is mixed-length normalization Gaussianizes the between-speaker distribution, but violates the assumption of a speaker-independent within-speaker distribution.We propose PSDA, an analogue to PLDA that uses Von Mises-Fisher distributions on the hypersphere for both within and between-class distributions.We show how the self-conjugacy of this distribution gives closed-form likelihood-ratio scores, making it a drop-in replacement for PLDA at scoring time.All kinds of trials can be scored, including single-enroll and multienroll verification, as well as more complex likelihood-ratios that could be used in clustering and diarization.Learning is done via an EM-algorithm with closed-form updates.We explain the model and present some first experiments.
Niko Brümmer, Albert Swart, Ladislav Mosner, Anna Silnova, Oldrich Plchot, Themos Stafylakis, Lukás Burget
INTERSPEECH1
2022 A speaker verification backend with robust performance across conditions
Luciana Ferrer, Mitchell McLaren, Niko Brümmer
Comput. Speech Lang.3
2021 Out of a Hundred Trials, How Many Errors Does Your Speaker Verifier Make?
abstract
Out of a hundred trials, how many errors does your speaker verifier make? For the user this is an important, practical question, but researchers and vendors typically sidestep it and supply instead the conditional error-rates that are given by the ROC/DET curve. We posit that the user's question is answered by the Bayes error-rate. We present a tutorial to show how to compute the error-rate that results when making Bayes decisions with calibrated likelihood ratios, supplied by the verifier, and an hypothesis prior, supplied by the user. For perfect calibration, the Bayes error-rate is upper bounded by min(EER,P,1-P), where EER is the equal-error-rate and P, 1-P are the prior probabilities of the competing hypotheses. The EER represents the accuracy of the verifier, while min(P,1-P) represents the hardness of the classification problem. We further show how the Bayes error-rate can be computed also for non-perfect calibration and how to generalize from error-rate to expected cost. We offer some criticism of decisions made by direct score thresholding. Finally, we demonstrate by analyzing error-rates of the recently published DCA-PLDA speaker verifier.
Niko Brümmer, Luciana Ferrer, Albert Swart
Interspeech1
2019 Large-Scale Speaker Diarization of Radio Broadcast Archives
abstract
This paper describes our initial efforts to build a large-scale speaker diarization (SD) and identification system on a recently digitized radio broadcast archive from the Netherlands which has more than 6500 audio tapes with 3000 hours of Frisian-Dutch speech recorded between 1950-2016. The employed large-scale diarization scheme involves two stages: (1) tape-level speaker diarization providing pseudo-speaker identities and (2) speaker linking to relate pseudo-speakers appearing in multiple tapes. Having access to the speaker models of several frequently appearing speakers from the previously collected FAME! speech corpus, we further perform speaker identification by linking these known speakers to the pseudo-speakers identified at the first stage. In this work, we present a recently created longitudinal and multilingual SD corpus designed for large-scale SD research and evaluate the performance of a new speaker linking system using x-vectors with PLDA to quantify cross-tape speaker similarity on this corpus. The performance of this speaker linking system is evaluated on a small subset of the archive which is manually annotated with speaker information. The speaker linking performance reported on this subset (53 hours) and the whole archive (3000 hours) is compared to quantify the impact of scaling up in the amount of speech data.
Emre Yilmaz 0001, Adem Derinel, Kun Zhou 0003, Henk van den Heuvel, Niko Brümmer, Haizhou Li 0001, David A. van Leeuwen
INTERSPEECH5
2018 Fast Variational Bayes for Heavy-tailed PLDA Applied to i-vectors and x-vectors
abstract
The standard state-of-the-art backend for text-independent speaker recognizers that use i-vectors or x-vectors, is Gaussian PLDA (G-PLDA), assisted by a Gaussianization step involving length normalization. G-PLDA can be trained with both generative or discriminative methods. It has long been known that heavy-tailed PLDA (HT-PLDA), applied without length normalization, gives similar accuracy, but at considerable extra computational cost. We have recently introduced a fast scoring algorithm for a discriminatively trained HT-PLDA backend. This paper extends that work by introducing a fast, variational Bayes, generative training algorithm. We compare old and new backends, with and without length-normalization, with i-vectors and x-vectors, on SRE'10, SRE'16 and SITW.
Anna Silnova, Niko Brümmer, Daniel Garcia-Romero, David Snyder, Lukás Burget
INTERSPEECH2
2017 Analysis and Description of ABC Submission to NIST SRE 2016
Oldrich Plchot, Pavel Matejka, Anna Silnova, Ondrej Novotný, Mireia Díez, Johan Rohdin, Ondrej Glembek, Niko Brümmer, Albert Swart, Jesús Jorrín-Prieto, L. Paola García-Perera, Luis Buera, Patrick Kenny, Jahangir Alam 0001, Gautam Bhattacharya
INTERSPEECH8
2017 A Generative Model for Score Normalization in Speaker Recognition
abstract
We propose a theoretical framework for thinking about score normalization, which confirms that normalization is not needed under (admittedly fragile) ideal conditions. If, however, these conditions are not met, e.g. under data-set shift between training and runtime, our theory reveals dependencies between scores that could be exploited by strategies such as score normalization. Indeed, it has been demonstrated over and over experimentally, that various ad-hoc score normalization recipes do work. We present a first attempt at using probability theory to design a generative score-space normalization model which gives similar improvements to ZT-norm on the text-dependent RSR 2015 database.
Albert Swart, Niko Brümmer
INTERSPEECH2
2017 Tied Variational Autoencoder Backends for i-Vector Speaker Recognition
Jesús Villalba 0001, Niko Brümmer, Najim Dehak
INTERSPEECH2
2015 The reddots data collection for speaker recognition
abstract
de niveau recherche, publiés ou non, émanant des établissements d'enseignement et de recherche français ou étrangers, des laboratoires publics ou privés.
Kong-Aik Lee, Anthony Larcher, Guangsen Wang, Patrick Kenny, Niko Brümmer, David A. van Leeuwen, Hagai Aronowitz, Marcel Kockmann, Carlos Vaquero, Bin Ma 0001, Haizhou Li 0001, Themos Stafylakis, Jahangir Alam 0001, Albert Swart, Javier Perez
INTERSPEECH5
2014 Generative modelling for unsupervised score calibration
abstract
Score calibration enables automatic speaker recognizers to make cost-effective accept / reject decisions. Traditional calibration requires supervised data, which is an expensive resource. We propose a 2-component GMM for unsupervised calibration and demonstrate good performance relative to a supervised baseline on NIST SRE'10 and SRE'12. A Bayesian analysis demonstrates that the uncertainty associated with the unsupervised calibration parameter estimates is surprisingly small.
Niko Brümmer, Daniel Garcia-Romero
ICASSP1
2014 Bayesian calibration for forensic evidence reporting
abstract
We introduce a Bayesian solution for the problem in forensic speaker recognition, where there may be very little background material for estimating score calibration parameters. We work within the Bayesian paradigm of evidence reporting and develop a principled probabilistic treatment of the problem, which results in a Bayesian likelihood-ratio as the vehicle for reporting weight of evidence. We show in contrast, that reporting a likelihood-ratio distribution does not solve this problem. Our solution is experimentally exercised on a simulated forensic scenario, using NIST SRE'12 scores, which demonstrates a clear advantage for the proposed method compared to the traditional plugin calibration recipe.
Niko Brümmer, Albert Swart
INTERSPEECH1
2014 Constrained speaker linking
abstract
In this paper we study speaker linking (a.k.a.\ partitioning) given constraints of the distribution of speaker identities over speech recordings. Specifically, we show that the intractable partitioning problem becomes tractable when the constraints pre-partition the data in smaller cliques with non-overlapping speakers. The surprisingly common case where speakers in telephone conversations are known, but the assignment of channels to identities is unspecified, is treated in a Bayesian way. We show that for the Dutch CGN database, where this channel assignment task is at hand, a lightweight speaker recognition system can quite effectively solve the channel assignment problem, with 93% of the cliques solved. We further show that the posterior distribution over channel assignment configurations is well calibrated.
David A. van Leeuwen, Niko Brümmer
INTERSPEECH2
2013 Likelihood-ratio calibration using prior-weighted proper scoring rules
abstract
Prior-weighted logistic regression has become a standard tool for calibration in speaker recognition. Logistic regression is the optimization of the expected value of the logarithmic scoring rule. We generalize this via a parametric family of proper scoring rules. Our theoretical analysis shows how different members of this family induce different relative weightings over a spectrum of applications of which the decision thresholds range from low to high. Special attention is given to the interaction between prior weighting and proper scoring rule parameters. Experiments on NIST SRE'12 suggest that for applications with low false-alarm rate requirements, scoring rules tailored to emphasize higher score thresholds may give better accuracy than logistic regression.
Niko Brümmer, George R. Doddington
INTERSPEECH1
2013 Eigenageing compensation for speaker verification
abstract
Dealing with the effect of vocal ageing on speaker verification is an important challenge. In this paper, a new approach to improving speaker verification performance in the presence of long-term ageing is presented. Analogous to eigenchannel compensation, the proposed eigenageing compensation method operates by adapting a speaker model to a test sample based on a predetermined ageing subspace. An experimental evaluation of the new method, using the Trinity College Dublin Speaker Ageing database, demonstrates it to be very effective at reducing long-term speaker verification error rates, and shows it to compare favourably with our previous stacked classifier technique. Index Terms: speaker verification, ageing, eigenanalysis 1.
Finnian Kelly, Niko Brümmer, Naomi Harte
INTERSPEECH2
2013 The distribution of calibrated likelihood-ratios in speaker recognition
abstract
This paper studies properties of the score distributions of calibrated log-likelihood-ratios that are used in automatic speaker recognition.We derive the essential condition for calibration that the log likelihood ratio of the log-likelihood-ratio is the log-likelihood-ratio.We then investigate what the consequence of this condition is to the probability density functions (PDFs) of the loglikelihood-ratio score.We show that if the PDF of the non-target distribution is Gaussian, then the PDF of the target distribution must be Gaussian as well.The means and variances of these two PDFs are interrelated, and determined completely by the discrimination performance of the recognizer characterized by the equal error rate.These relations allow for a new way of computing the offset and scaling parameters for linear calibration, and we derive closed-form expressions for these and show that for modern i-vector systems with PLDA scoring this leads to good calibration, comparable to traditional logistic regression, over a wide range of system performance.
David A. van Leeuwen, Niko Brümmer
INTERSPEECH2
2013 The albayzin 2012 language recognition evaluation
abstract
The Albayzin 2008 Language Recognition Evaluation was held from May to October 2008, and their results presented and discussed among the participating teams at the 5th Biennial Workshop on Speech Technology [1], or-ganized by the Spanish Network on Speech Technologies [2] in November 2008. In this paper, we present (for the first time) a full description of the Albayzin 2008 LRE and analyze and discuss recognition results. The evalua-tion was designed according to the test procedures, pro-tocols and performance measures used in the NIST 2007 LRE. The KALAKA database [3], consisting of 16 kHz audio signals recorded from TV broadcasts, was created ad-hoc and used for the evaluation. The four official lan-guages spoken in Spain (Basque, Catalan, Galician and
Luis Javier Rodríguez-Fuentes, Niko Brümmer, Mikel Peñagarikano, Amparo Varona, Germán Bordel, Mireia Díez
INTERSPEECH2
2013 Pairwise Discriminative Speaker Verification in the 𝕀-Vector Space
abstract
This work presents a new and efficient approach to discriminative speaker verification in the${\rm i}$–vector space. We illustrate the development of a linear discriminative classifier that is trained to discriminate between the hypothesis that a pair of feature vectors in a trial belong to the same speaker or to different speakers. This approach is alternative to the usual discriminative setup that discriminates between a speaker and all the other speakers. We use a discriminative classifier based on a Support Vector Machine (SVM) that is trained to estimate the parameters of a symmetric quadratic function approximating a log–likelihood ratio score without explicit modeling of the${\rm i}$–vector distributions as in the generative Probabilistic Linear Discriminant Analysis (PLDA) models. Training these models is feasible because it is not necessary to expand the${\rm i}$–vector pairs, which would be expensive or even impossible even for medium sized training sets. The results of experiments performed on the tel-tel extended core condition of the NIST 2010 Speaker Recognition Evaluation are competitive with the ones obtained by generative models, in terms of normalized Detection Cost Function and Equal Error Rate. Moreover, we show that it is possible to train a gender–independent discriminative model that achieves state–of–the–art accuracy, comparable to the one of a gender–dependent system, saving memory and execution time both in training and in testing.
Sandro Cumani, Niko Brümmer, Lukás Burget, Pietro Laface, Oldrich Plchot, Vasileios Vasilakakis
IEEE Trans. Speech Audio Process.2
2012 Gender independent discriminative speaker recognition in i-vector space
abstract
Speaker recognition systems attain their best accuracy when trained with gender dependent features and tested with known gender trials. In real applications, however, gender labels are often not given. In this work we illustrate the design of a system that does not make use of the gender labels both in training and in test, i.e. a completely Gender Independent (GI) system. It relies on discriminative training, where the trials are i-vector pairs, and the discrimination is between the hypothesis that the pair of feature vectors in the trial belong to the same speaker or to different speakers. We demonstrate that this pairwise discriminative training can be interpreted as a procedure that estimates the parameters of the best (second order) approximation of the log-likelihood ratio score function, and that a pairwise SVM can be used for training a gender independent system. Our results show that a pairwise GI SVM, saving memory and execution time, achieves on the last NIST evaluations state-of-the-art performance, comparable to a Gender Dependent(GD) system.
Sandro Cumani, Ondrej Glembek, Niko Brümmer, Edward de Villiers, Pietro Laface
ICASSP3
2011 Discriminatively trained Probabilistic Linear Discriminant Analysis for speaker verification
abstract
Recently, i-vector extraction and Probabilistic Linear Discriminant Analysis (PLDA) have proven to provide state-of-the-art speaker verification performance. In this paper, the speaker verification score for a pair of i-vectors representing a trial is computed with a functional form derived from the successful PLDA generative model. In our case, however, parameters of this function are estimated based on a discriminative training criterion. We propose to use the objective function to directly address the task in speaker verification: discrimination between same-speaker and different-speaker trials. Compared with a baseline which uses a generatively trained PLDA model, discriminative training provides up to 40% relative improvement on the NIST S RE 2010 evaluation task.
Lukás Burget, Oldrich Plchot, Sandro Cumani, Ondrej Glembek, Pavel Matejka, Niko Brümmer
ICASSP6
2011 Fast discriminative speaker verification in the i-vector space
abstract
This work presents a new approach to discriminative speaker verification. Rather than estimating speaker models, or a model that discriminates between a speaker class and the class of all the other speakers, we directly solve the problem of classifying pairs of utterances as belonging to the same speaker or not.
Sandro Cumani, Niko Brümmer, Lukás Burget, Pietro Laface
ICASSP2
2011 Discriminatively Trained i-vector Extractor for Speaker Verification
abstract
We propose a strategy for discriminative training of the ivector extractor in speaker recognition. The original i-vector extractor training was based on the maximum-likelihood generative modeling, where the EM algorithm was used. In our approach, the i-vector extractor parameters are numerically optimized to minimize the discriminative cross-entropy error function. Two versions of the i-vector extraction are studied—the original approach as defined for Joint Factor Analysis, and the simplified version, where orthogonalization of the i-vector extractor matrix is performed. Index Terms: speaker verification, i-vectors, PLDA, discriminative training
Ondrej Glembek, Lukás Burget, Niko Brümmer, Oldrich Plchot, Pavel Matejka
INTERSPEECH3
2011 A Speaker Line-Up for the Likelihood Ratio
abstract
\n Contains fulltext :\n 94223.pdf (author's version ) (Open Access)\n
David A. van Leeuwen, Niko Brümmer
INTERSPEECH2
2011 Mixture of PLDA Models in i-vector Space for Gender-Independent Speaker Recognition
abstract
The Speaker Recognition community that participates in NIST evaluations has concentrated on designing genderand channel-conditioned systems. In the real word, this conditioning is not feasible. Our main purpose in this work is to propose a mixture of Probabilistic Linear Discriminant Analysis models (PLDA) as a solution for making systems independent of speaker gender. In order to show the effectiveness of the mixture model, we first experiment on 2010 NIST telephone speech (det5), where we prove that there is no loss of accuracy compared with a baseline gender-dependent model. We also test with success the mixture model on a more realistic situation where there are cross-gender trials. Furthermore, we report results on microphone speech for the det1, det2, det3 and det4 tasks to confirm the effectiveness of the mixture model.
Mohammed Senoussaoui, Patrick Kenny, Niko Brümmer, Edward de Villiers, Pierre Dumouchel
INTERSPEECH3
2011 Towards Fully Bayesian Speaker Recognition: Integrating Out the Between-Speaker Covariance
abstract
We propose a variational Bayes solution to integrate out the model parameters in a generative i-vector speaker recognizer. The existing state-of-the-art in generative i-vector modelling plugs in fixed maximum-likelihood point-estimates of model parameters. This recipe may suffer from over-fitting of especially the between-speaker covariance. We show how to integrate out the between-speaker covariance and demonstrate dramatic improvements on NIST SRE 2010.
Jesús Villalba 0001, Niko Brümmer
INTERSPEECH2
2009 Comparison of scoring methods used in speaker recognition with Joint Factor Analysis
abstract
The aim of this paper is to compare different log-likelihood scoring methods, that different sites used in the latest state-of-the-art Joint Factor Analysis (JFA) Speaker Recognition systems. The algorithms use various assumptions and have been derived from various approximations of the objective functions of JFA. We compare the techniques in terms of speed and performance. We show, that approximations of the true log-likelihood ratio (LLR) may lead to significant speedup without any loss in performance.
Ondrej Glembek, Lukás Burget, Najim Dehak, Niko Brümmer, Patrick Kenny
ICASSP4
2009 Discriminative acoustic language recognition via channel-compensated GMM statistics
abstract
We propose a novel design for acoustic feature-based automatic spoken language recognizers. Our design is inspired by recent advances in text-independent speaker recognition, where intraclass variability is modeled by factor analysis in Gaussian mixture model (GMM) space. We use approximations to GMMlikelihoods which allow variable-length data sequences to be represented as statistics of fixed size. Our experiments on NIST LRE’07 show that variability-compensation of these statistics can reduce error-rates by a factor of three. Finally, we show that further improvements are possible with discriminative logistic regression training. Index Terms: acoustic language recognition, intersession variability compensation, discriminative training
Niko Brümmer, Albert Strasheim, Valiantsina Hubeika, Pavel Matejka, Lukás Burget, Ondrej Glembek
INTERSPEECH1
2009 Support vector machines versus fast scoring in the low-dimensional total variability space for speaker verification
abstract
This paper presents a new speaker verification system architecture based on Joint Factor Analysis (JFA) as feature extractor. In this modeling, the JFA is used to define a new low-dimensional space named the total variability factor space, instead of both channel and speaker variability spaces for the classical JFA. The main contribution in this approach, is the use of the cosine kernel in the new total factor space to design two different systems: the first system is Support Vector Machines based, and the second one uses directly this kernel as a decision score. This last scoring method makes the process faster and less computation complex compared to others classical methods. We tested several intersession compensation methods in total factors, and we found that the combination of Linear Discriminate Analysis and Within Class Covariance Normalization achieved the best performance. We achieved a remarkable results using fast scoring method based only on cosine kernel especially for male trials, we yield an EER of 1.12% and MinDCF of 0.0094 on the English trials of the NIST 2008 SRE dataset. Index Terms: Total variability space, cosine kernel, fast scoring, support vector machines.
Najim Dehak, Réda Dehak, Patrick Kenny, Niko Brümmer, Pierre Ouellet, Pierre Dumouchel
INTERSPEECH4
2007 STBU System for the NIST 2006 Speaker Recognition Evaluation
abstract
This paper describes STBU 2006 speaker recognition system, which performed well in the NIST 2006 speaker recognition evaluation. STBU is consortium of 4 partners: Spescom DataVoice (South Africa), TNO (Netherlands), BUT (Czech Republic) and University of Stellenbosch (South Africa). The primary system is a combination of three main kinds of systems: (1) GMM, with short-time MFCC or PLP features, (2) GMM-SVM, using GMM mean supervectors as input and (3) MLLR-SVM, using MLLR speaker adaptation coefficients derived from English LVCSR system. In this paper, we describe these sub-systems and present results for each system alone and in combination on the NIST Speaker Recognition Evaluation (SRE) 2006 development and evaluation data sets.
Pavel Matejka, Lukás Burget, Petr Schwarz, Ondrej Glembek, Martin Karafiát, Frantisek Grézl, Jan Cernocký, David A. van Leeuwen, Niko Brümmer, Albert Strasheim
ICASSP (4)9
2007 Fusion of Heterogeneous Speaker Recognition Systems in the STBU Submission for the NIST Speaker Recognition Evaluation 2006
abstract
This paper describes and discusses the "STBU" speaker recognition system, which performed well in the NIST Speaker Recognition Evaluation 2006 (SRE). STBU is a consortium of four partners: Spescom DataVoice (Stellenbosch, South Africa), TNO (Soesterberg, The Netherlands), BUT (Brno, Czech Republic), and the University of Stellenbosch (Stellenbosch, South Africa). The STBU system was a combination of three main kinds of subsystems: 1) GMM, with short-time Mel frequency cepstral coefficient (MFCC) or perceptual linear prediction (PLP) features, 2) Gaussian mixture model-support vector machine (GMM-SVM), using GMM mean supervectors as input to an SVM, and 3) maximum-likelihood linear regression-support vector machine (MLLR-SVM), using MLLR speaker adaptation coefficients derived from an English large vocabulary continuous speech recognition (LVCSR) system. All subsystems made use of supervector subspace channel compensation methods-either eigenchannel adaptation or nuisance attribute projection. We document the design and performance of all subsystems, as well as their fusion and calibration via logistic regression. Finally, we also present a cross-site fusion that was done with several additional systems from other NIST SRE-2006 participants.
Niko Brümmer, Lukás Burget, Jan Cernocký, Ondrej Glembek, Frantisek Grézl, Martin Karafiát, David A. van Leeuwen, Pavel Matejka, Petr Schwarz, Albert Strasheim
IEEE Trans. Speech Audio Process.1
2006 Application-independent evaluation of speaker detection
Niko Brümmer, Johan A. du Preez
Comput. Speech Lang.1