David A. van Leeuwen

dblp:57/3801 · DBLP profile ↗
← Back
71ranked-venue papers
15as first author
5since 2021 · last 2025
0000-0001-9704-6141ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 59 · 14 first-author · 5 since 2021Artificial intelligence and machine learning · 54 · 14 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Learning Strategy with Barlow Twins Objective for Emotion-Robust Speaker Verification System
abstract
A recent trend of the automatic speaker verification (ASV) systems is to use speaker representations extracted from deep learning-based speaker encoders. Those representations are robust against linguistic variation but sensitive to emotional fluctuation that potentially degrades the performance of ASV systems. To tackle this issue, we applied a correlation-based objective function called Barlow Twins objective to speaker representation learning for the ASV task with expressive speech. This helps the learned representations from the same speaker become similar despite their emotional state. Our experiment showed that our representation learning using Barlow Twins objective improves the standard deviation of the EERs between different emotions from 2.43 to 1.52 and the total EER from 10.77% to 6.47% on the CREMA-D dataset. The study also evaluates the system’s generalization performance across out-of-domain datasets, demonstrating improved standard deviation of the EERs and total EER of the baseline system in all datasets while observing the performance drop on synthetic emotional data where emotional mismatch occurs.
Dmitrii V. Mikhailovskii, Aki Kunikoshi, David A. van Leeuwen, Jaebok Kim
ICASSP3
2025 Self-supervised learning of speech representations with Dutch archival data
abstract
Contains fulltext : 326012.pdf (Publisher’s version ) (Open Access)
Nik Vaessen, Roeland Ordelman, David A. van Leeuwen
INTERSPEECH3
2023 Towards Multi-task Learning of Speech and Speaker Recognition
abstract
Contains fulltext : 300714.pdf (Publisher’s version ) (Open Access)
Nik Vaessen, David A. van Leeuwen
INTERSPEECH2
2022 Fine-Tuning Wav2Vec2 for Speaker Recognition
abstract
This paper explores applying the wav2vec2 framework to speaker recognition instead of speech recognition. We study the effectiveness of the pre-trained weights on the speaker recognition task, and how to pool the wav2vec2 output sequence into a fixed-length speaker embedding. To adapt the framework to speaker recognition, we propose a single-utterance classification variant with cross-entropy or additive angular softmax loss, and an utterance-pair classification variant with BCE loss. Our best performing variant achieves a 1.88% EER on the extended voxceleb1 test set compared to 1.69% EER with an ECAPA-TDNN baseline. Code is available at github.com/nikvaessen/w2v2-speaker.
Nik Vaessen, David A. van Leeuwen
ICASSP2
2022 Training speaker recognition systems with limited data
abstract
Contains fulltext : 283068.pdf (Publisher’s version ) (Open Access)
Nik Vaessen, David A. van Leeuwen
INTERSPEECH2
2019 Multi-Graph Decoding for Code-Switching ASR
abstract
\n Contains fulltext :\n 214606.pdf (Publisher’s version ) (Open Access)\n
Emre Yilmaz 0001, Samuel Cohen, Xianghu Yue, David A. van Leeuwen, Haizhou Li 0001
INTERSPEECH4
2019 Large-Scale Speaker Diarization of Radio Broadcast Archives
abstract
This paper describes our initial efforts to build a large-scale speaker diarization (SD) and identification system on a recently digitized radio broadcast archive from the Netherlands which has more than 6500 audio tapes with 3000 hours of Frisian-Dutch speech recorded between 1950-2016. The employed large-scale diarization scheme involves two stages: (1) tape-level speaker diarization providing pseudo-speaker identities and (2) speaker linking to relate pseudo-speakers appearing in multiple tapes. Having access to the speaker models of several frequently appearing speakers from the previously collected FAME! speech corpus, we further perform speaker identification by linking these known speakers to the pseudo-speakers identified at the first stage. In this work, we present a recently created longitudinal and multilingual SD corpus designed for large-scale SD research and evaluate the performance of a new speaker linking system using x-vectors with PLDA to quantify cross-tape speaker similarity on this corpus. The performance of this speaker linking system is evaluated on a small subset of the archive which is manually annotated with speaker information. The speaker linking performance reported on this subset (53 hours) and the whole archive (3000 hours) is compared to quantify the impact of scaling up in the amount of speech data.
Emre Yilmaz 0001, Adem Derinel, Kun Zhou 0003, Henk van den Heuvel, Niko Brümmer, Haizhou Li 0001, David A. van Leeuwen
INTERSPEECH7
2018 Acoustic and Textual Data Augmentation for Improved ASR of Code-Switching Speech
abstract
\n Contains fulltext :\n 196526.pdf (Publisher’s version ) (Open Access)\n
Emre Yilmaz 0001, Henk van den Heuvel, David A. van Leeuwen
INTERSPEECH3
2018 Semi-supervised acoustic model training for speech with code-switching
abstract
In the FAME! project, we aim to develop an automatic speech recognition (ASR) system for Frisian-Dutch code-switching (CS) speech extracted from the archives of a local broadcaster with the ultimate goal of building a spoken document retrieval system. Unlike Dutch, Frisian is a low-resourced language with a very limited amount of manually annotated speech data. In this paper, we describe several automatic annotation approaches to enable using of a large amount of raw bilingual broadcast data for acoustic model training in a semi-supervised setting. Previously, it has been shown that the best-performing ASR system is obtained by two-stage multilingual deep neural network (DNN) training using 11 hours of manually annotated CS speech (reference) data together with speech data from other high-resourced languages. We compare the quality of transcriptions provided by this bilingual ASR system with several other approaches that use a language recognition system for assigning language labels to raw speech segments at the front-end and using monolingual ASR resources for transcription. We further investigate automatic annotation of the speakers appearing in the raw broadcast data by first labeling with (pseudo) speaker tags using a speaker diarization system and then linking to the known speakers appearing in the reference data using a speaker recognition system. These speaker labels are essential for speaker-adaptive training in the proposed setting. We train acoustic models using the manually and automatically annotated data and run recognition experiments on the development and test data of the FAME! speech corpus to quantify the quality of the automatic annotations. The ASR and CS detection results demonstrate the potential of using automatic language and speaker tagging in semi-supervised bilingual acoustic model training.
Emre Yilmaz 0001, Mitchell McLaren, Henk van den Heuvel, David A. van Leeuwen
Speech Commun.4
2017 Language diarization for semi-supervised bilingual acoustic model training
abstract
In this paper, we investigate several automatic transcription schemes for using raw bilingual broadcast news data in semi-supervised bilingual acoustic model training. Specifically, we compare the transcription quality provided by a bilingual ASR system with another system performing language diarization at the front-end followed by two monolingual ASR systems chosen based on the assigned language label. Our research focuses on the Frisian-Dutch code-switching (CS) speech that is extracted from the archives of a local radio broadcaster. Using 11 hours of manually transcribed Frisian speech as a reference, we aim to increase the amount of available training data by using these automatic transcription techniques. By merging the manually and automatically transcribed data, we learn bilingual acoustic models and run ASR experiments on the development and test data of the FAME! speech corpus to quantify the quality of the automatic transcriptions. Using these acoustic models, we present speech recognition and CS detection accuracies. The results demonstrate that applying language diarization to the raw speech data to enable using the monolingual resources improves the automatic transcription quality compared to a baseline system using a bilingual ASR system.
Emre Yilmaz 0001, Mitchell McLaren, Henk van den Heuvel, David A. van Leeuwen
ASRU4
2017 Longitudinal Speaker Clustering and Verification Corpus with Code-Switching Frisian-Dutch Speech
abstract
10.21437/Interspeech.2017-301
Emre Yilmaz 0001, Jelske Dijkstra, Hans Van de Velde, Frederik Kampstra, Jouke Algra, Henk van den Heuvel, David A. van Leeuwen
INTERSPEECH7
2017 Exploiting Untranscribed Broadcast Data for Improved Code-Switching Detection
abstract
We have recently presented an automatic speech recognition (ASR) system operating on Frisian-Dutch code-switched speech.This type of speech requires careful handling of unexpected language switches that may occur in a single utterance.In this paper, we extend this work by using some raw broadcast data to improve multilingually trained deep neural networks (DNN) that have been trained on 11.5 hours of manually annotated bilingual speech.For this purpose, we apply the initial ASR to the untranscribed broadcast data and automatically create transcriptions based on the recognizer output using different language models for rescoring.Then, we train new acoustic models on the combined data, i.e., the manually and automatically transcribed bilingual broadcast data, and investigate the automatic transcription quality based on the recognition accuracies on a separate set of development and test data.Finally, we report code-switching detection performance elaborating on the correlation between the ASR and the code-switching detection performance.
Emre Yilmaz 0001, Henk van den Heuvel, David A. van Leeuwen
INTERSPEECH3
2016 Open Source Speech and Language Resources for Frisian
abstract
10.21437/Interspeech.2016-48
Emre Yilmaz 0001, Henk van den Heuvel, Jelske Dijkstra, Hans Van de Velde, Frederik Kampstra, Jouke Algra, David A. van Leeuwen
INTERSPEECH7
2016 A Longitudinal Bilingual Frisian-Dutch Radio Broadcast Database Designed for Code-Switching Research
Emre Yilmaz 0001, Maaike Andringa, Sigrid Kingma, Jelske Dijkstra, Frits Van der Kuip, Hans Van de Velde, Frederik Kampstra, Jouke Algra, Henk van den Heuvel, David A. van Leeuwen
LREC10
2016 Code-switching detection using multilingual DNNS
abstract
Automatic speech recognition (ASR) of code-switching speech requires careful handling of unexpected language switches that may occur in a single utterance. In this paper, we investigate the feasibility of using multilingually trained deep neural networks (DNN) for the ASR of Frisian speech containing code-switches to Dutch with the aim of building a robust recognizer that can handle this phenomenon. For this purpose, we train several multilingual DNN models on Frisian and two closely related languages, namely English and Dutch, to compare the impact of single-step and two-step multilingual DNN training on the recognition and code-switching detection performance. We apply bilingual DNN retraining on both target languages by varying the amount of training data belonging to the higher-resourced target language (Dutch). The recognition results show that the multilingual DNN training scheme with an initial multilingual training step followed by bilingual retraining provides recognition performance comparable to an oracle baseline recognizer that can employ language-specific acoustic models. We further show that we can detect code-switches at the word level with an equal error rate of around 17% excluding the deletions due to ASR errors.
Emre Yilmaz 0001, Henk van den Heuvel, David A. van Leeuwen
SLT3
2016 A study of speaker clustering for speaker attribution in large telephone conversation datasets
Houman Ghaemmaghami, David Dean, Sridha Sridharan, David A. van Leeuwen
Comput. Speech Lang.4
2015 The reddots data collection for speaker recognition
abstract
de niveau recherche, publiés ou non, émanant des établissements d'enseignement et de recherche français ou étrangers, des laboratoires publics ou privés.
Kong-Aik Lee, Anthony Larcher, Guangsen Wang, Patrick Kenny, Niko Brümmer, David A. van Leeuwen, Hagai Aronowitz, Marcel Kockmann, Carlos Vaquero, Bin Ma 0001, Haizhou Li 0001, Themos Stafylakis, Jahangir Alam 0001, Albert Swart, Javier Perez
INTERSPEECH6
2015 Quality measures based calibration with duration and noise dependency for speaker recognition
Miranti Indar Mandasari, Rahim Saeidi, David A. van Leeuwen
Speech Commun.3
2014 Effect of long-term ageing on i-vector speaker verification
abstract
Assessing the impact of ageing on biometric systems is an important challenge. In this paper, an i-vector speaker verifi-cation framework is used to evaluate the impact of long-term ageing on state-of-the-art speaker verification. Using the Trin-ity College Dublin Speaker Ageing (TCDSA) database, it is ob-served that the performance of the i-vector system, in terms of both discrimination and calibration, degrades progressively as the absolute age difference between training and testing sam-ples increases. In the case of male speakers, the equal error rate (EER) increases from 4.61 % at an ageing difference of 0–1 years to 32.74 % at an age difference of 51–60 years. The performance of a Gaussian Mixture Model- Universal Back-ground Model (GMM-UBM) system is presented for compari-son. It is shown that while the i-vector system outperforms the GMM-UBM system, as absolute age difference increases, the performance of both degrades at a similar rate. It is concluded that long-term ageing variability is distinct from everyday inter-session variability, and therefore must be dealt with via dedi-cated compensation strategies.
Finnian Kelly, Rahim Saeidi, Naomi Harte, David A. van Leeuwen
INTERSPEECH4
2014 Constrained speaker linking
abstract
In this paper we study speaker linking (a.k.a.\ partitioning) given constraints of the distribution of speaker identities over speech recordings. Specifically, we show that the intractable partitioning problem becomes tractable when the constraints pre-partition the data in smaller cliques with non-overlapping speakers. The surprisingly common case where speakers in telephone conversations are known, but the assignment of channels to identities is unspecified, is treated in a Bayesian way. We show that for the Dutch CGN database, where this channel assignment task is at hand, a lightweight speaker recognition system can quite effectively solve the channel assignment problem, with 93% of the cliques solved. We further show that the posterior distribution over channel assignment configurations is well calibrated.
David A. van Leeuwen, Niko Brümmer
INTERSPEECH1
2014 Semi-automatic annotation of the UCU accents speech corpus
Rosemary Orr, Marijn Huijbregts, Roeland van Beek, Lisa Teunissen, Kate Backhouse, David A. van Leeuwen
LREC6
2014 Speaker age estimation using i-vectors
Mohamad Hasan Bahari, Mitchell McLaren, Hugo Van hamme, David A. van Leeuwen
Eng. Appl. Artif. Intell.4
2013 Accent recognition using i-vector, Gaussian Mean Supervector and Gaussian posterior probability supervector for spontaneous telephone speech
abstract
In this paper, three utterance modelling approaches, namely Gaussian Mean Supervector (GMS), i-vector and Gaussian Posterior Probability Supervector (GPPS), are applied to the accent recognition problem. For each utterance modeling method, three different classifiers, namely the Support Vector Machine (SVM), the Naive Bayesian Classifier (NBC) and the Sparse Representation Classifier (SRC), are employed to find out suitable matches between the utterance modelling schemes and the classifiers. The evaluation database is formed by using English utterances of speakers whose native languages are Russian, Hindi, American English, Thai, Vietnamese and Cantonese. These utterances are drawn from the National Institute of Standards and Technology (NIST) 2008 Speaker Recognition Evaluation (SRE) database. The study results show that GPPS and i-vector are more effective than GMS in this accent recognition task. It is also concluded that among the employed classifiers, the best matches for i-vector and GPPS are SVM and SRC, respectively.
Mohamad Hasan Bahari, Rahim Saeidi, Hugo Van hamme, David A. van Leeuwen
ICASSP4
2013 Duration mismatch compensation for i-vector based speaker recognition systems
abstract
Speaker recognition systems trained on long duration utterances are known to perform significantly worse when short test segments are encountered. To address this mismatch, we analyze the effect of duration variability on phoneme distributions of speech utterances and i-vector length. We demonstrate that, as utterance duration is decreased, number of detected unique phonemes and i-vector length approaches zero in a logarithmic and non-linear fashion, respectively. Assuming duration variability as an additive noise in the i-vector space, we propose three different strategies for its compensation: i) multi-duration training in Probabilistic Linear Discriminant Analysis (PLDA) model, ii) score calibration using log duration as a Quality Measure Function (QMF), and iii) multi-duration PLDA training with synthesized short duration i-vectors. Experiments are designed based on the 2012 National Institute of Standards and Technology (NIST) Speaker Recognition Evaluation (SRE) protocol with varying test utterance duration. Experimental results demonstrate the effectiveness of the proposed schemes on short duration test conditions, especially with the QMF calibration approach.
Taufiq Hasan, Rahim Saeidi, John H. L. Hansen, David A. van Leeuwen
ICASSP4
2013 Knowing the non-target speakers: The effect of the i-vector population for PLDA training in speaker recognition
abstract
Inspired by the NIST SRE-2012 evaluation conditions we train the PLDA classifier in an i-vector speaker recognition system with different speaker populations, either including or excluding the target speakers in the evaluation. Including the target speakers in the PLDA training is always beneficial w.r.t. completely excluding them-which is the normal situation in pre-2012 SRE protocols-even in the Pknown= 0 evaluation condition. However, adding other speakers than just the targets speakers can slightly increase performance. We also investigated the effect of adding i-vectors extracted from segments with added noise in the PLDA training. This generally makes the system more robust to noise in the test segments, and doesn't hurt performance in the clean condition. The paper further details the 'simple to compound' log-likelihood-ratio conversion necessary for SRE-2012 style calibration.
David A. van Leeuwen, Rahim Saeidi
ICASSP1
2013 Automatic regularization of cross-entropy cost for speaker recognition fusion
abstract
\n Contains fulltext :\n 116325.pdf (author's version ) (Open Access)\n
Ville Hautamäki, Kong-Aik Lee, David A. van Leeuwen, Rahim Saeidi, Anthony Larcher, Tomi Kinnunen, Taufiq Hasan, Seyed Omid Sadjadi, Gang Liu 0001, Hynek Boril, John H. L. Hansen, Benoit G. B. Fauve
INTERSPEECH3
2013 The distribution of calibrated likelihood-ratios in speaker recognition
abstract
This paper studies properties of the score distributions of calibrated log-likelihood-ratios that are used in automatic speaker recognition.We derive the essential condition for calibration that the log likelihood ratio of the log-likelihood-ratio is the log-likelihood-ratio.We then investigate what the consequence of this condition is to the probability density functions (PDFs) of the loglikelihood-ratio score.We show that if the PDF of the non-target distribution is Gaussian, then the PDF of the target distribution must be Gaussian as well.The means and variances of these two PDFs are interrelated, and determined completely by the discrimination performance of the recognizer characterized by the equal error rate.These relations allow for a new way of computing the offset and scaling parameters for linear calibration, and we derive closed-form expressions for these and show that for modern i-vector systems with PLDA scoring this leads to good calibration, comparable to traditional logistic regression, over a wide range of system performance.
David A. van Leeuwen, Niko Brümmer
INTERSPEECH1
2013 I4u submission to NIST SRE 2012: a large-scale collaborative effort for noise-robust speaker verification
abstract
I4U is a joint entry of nine research Institutes and Universities across 4 continents to NIST SRE 2012. It started with a brief discussion during the Odyssey 2012 workshop in Singapore. An online discussion group was soon set up, providing a discussion platform for different issues surrounding NIST SRE’12. Noisy test segments, uneven multi-session training, variable enrollment duration, and the issue of open-set identification were actively discussed leading to various solutions integrated to the I4U submission. The joint submission and several of its 17 sub-systems were among top-performing systems. We summarize the lessons learnt from this large-scale effort.
Rahim Saeidi, Kong-Aik Lee, Tomi Kinnunen, Tawfik Hasan, Benoit G. B. Fauve, Pierre-Michel Bousquet, Elie Khoury 0001, Pablo Luis Sordo Martinez, Jia Min Karen Kua, Chang Huai You, Hanwu Sun, Anthony Larcher, Padmanabhan Rajan, Ville Hautamäki, Cemal Hanilçi, Billy Braithwaite, Rosa González Hautamäki, Seyed Omid Sadjadi, Gang Liu 0001, Hynek Boril, Navid Shokouhi, Driss Matrouf, Laurent El Shafey, Pejman Mowlaee, Julien Epps, Tharmarajah Thiruvaran, David A. van Leeuwen, Bin Ma 0001, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Sébastien Marcel, John S. D. Mason, Eliathamby Ambikairajah
INTERSPEECH27
2013 Quality Measure Functions for Calibration of Speaker Recognition Systems in Various Duration Conditions
abstract
This paper investigates the effect of utterance duration to the calibration of a modern i-vector speaker recognition system with probabilistic linear discriminant analysis (PLDA) modeling. A calibration approach to deal with these effects using quality measure functions (QMFs) is proposed to include duration in the calibration transformation. Extensive experiments are performed in order to evaluate the robustness of the proposed calibration approach for unseen conditions in the training of calibration parameters. Using the latest NIST corpora for evaluation, results highlight the importance of considering the quality metrics like duration in calibrating the scores for automatic speaker recognition systems.
Miranti Indar Mandasari, Rahim Saeidi, Mitchell McLaren, David A. van Leeuwen
IEEE Trans. Speech Audio Process.4
2012 The effect of noise on modern automatic speaker recognition systems
abstract
Motivated by the application of speaker recognition in forensic area, this paper presents a study on noise robustness of several automatic speaker recognition system approaches, ranging from simple dot-scoring and a standard i-vector system with cosine distance scoring to a state-of-the-art i-vector Probabilistic Linear Discriminant Analysis (PLDA) system. Using the recent NIST 2010 Speaker Recognition Evaluation (SRE) data, the systems are analyzed in added noise conditions with a range of signal to noise ratios. Various experiments were conducted to study the influence of the noise on the speech activity detection and Wiener filtering in the front-end of the system.
Miranti Indar Mandasari, Mitchell McLaren, David A. van Leeuwen
ICASSP3
2012 Gender-independent speaker recognition using source normalisation
abstract
Source-normalisation (SN) was proposed to improve the robustness of i-vector-based speaker recognition for under-resourced and unseen cross-speech-source evaluation conditions. The technique of source-normalisation estimates directions of undesired within-speaker variation more accurately than traditional methods when cross-source variation is not explicitly observed from each speaker in system development data. Incorporated into Within Class Covariance Normalisation (WCCN), source-normalisation provides significant improvements to speaker recognition based on i-vectors. This paper proposes a novel approach to gender-independent Probabilistic LDA (PLDA) through the use of SN-WCCN to normalise for the variation that separates genders as a pre-processing step for i-vector based PLDA classification. Evaluated on the NIST 2010 speaker recognition evaluation (SRE) dataset, the proposed approach demonstrated performance comparable to a typical gender-dependent configuration.
Mitchell McLaren, David A. van Leeuwen
ICASSP2
2012 Age Estimation from Telephone Speech using i-vectors
abstract
Motivated by the success of i-vectors in the field of speaker recognition, this paper proposes a new approach for age estimation from telephone speech patterns based on i-vectors. In this method, each utterance is modeled by its corresponding i-vector. Then, Support Vector Regression (SVR) is applied to estimate the age of speakers. The proposed method is trained and tested on telephone conversations of the National Institute for Standard in Technology (NIST) 2010 and 2008 Speaker Recognition Evaluations databases. Evaluation results show that the proposed method outperforms different conventional methods in speaker age estimation.
Mohamad Hasan Bahari, Mitchell McLaren, Hugo Van hamme, David A. van Leeuwen
INTERSPEECH4
2012 Calibration of probabilistic age recognition
abstract
\n Contains fulltext :\n 102327.pdf (author's version ) (Open Access)\n
David A. van Leeuwen, Mohamad Hasan Bahari
INTERSPEECH1
2012 Speech-based recognition of self-reported and observed emotion in a dimensional space
Khiet P. Truong, David A. van Leeuwen, Franciska de Jong
Speech Commun.2
2012 Large-Scale Speaker Diarization for Long Recordings and Small Collections
abstract
Performing speaker diarization of very long recordings is a problem for most diarization systems that are based on agglomerative clustering with an hidden Markov model (HMM) topology. Performing collection-wide speaker diarization, where each speaker is identified uniquely across the entire collection, is even a more challenging task. In this paper we propose a method with which it is possible to efficiently perform diarization of long recordings. We have also applied this method successfully to a collection of a total duration of approximately 15 hours. The method consists of first segmenting long recordings into smaller chunks on which diarization is performed. Next, a speaker detection system is used to link the speech clusters from each chunk and to assign a unique label to each speaker in the long recording or in the small collection. We show for three different audio collections that it is possible to perform high-quality diarization with this approach. The long meetings from the ICSI corpus are processed 5.5 times faster than the originally needed time and by uniquely labeling each speaker across the entire collection it becomes possible to perform speaker-based information retrieval with high accuracy (mean average precision of 0.57).
Marijn Huijbregts, David A. van Leeuwen
IEEE Trans. Speech Audio Process.2
2012 Speaker Diarization Error Analysis Using Oracle Components
abstract
In this paper, we describe an analysis of our speaker diarization system based on a series of oracle experiments. In this analysis, each system component is substituted by an oracle component that uses the reference transcripts to perform flawlessly. By placing the original components back into the system one at a time, either in a top-down or bottom-up manner, the performance of each individual system component is measured. The analysis approach can be applied to any speaker diarization system that consists of a concatenation of separate components. Our experimental findings are relevant for most RT09s diarization systems that all apply similar techniques. The analysis revealed that three components caused most errors: speech activity detection, the inability to handle overlapping speech, and robustness of the merging component to cluster impurity.
Marijn Huijbregts, David A. van Leeuwen, Chuck Wooters
IEEE Trans. Speech Audio Process.2
2012 Source-Normalized LDA for Robust Speaker Recognition Using i-Vectors From Multiple Speech Sources
Mitchell McLaren, David A. van Leeuwen
IEEE Trans. Speech Audio Process.2
2011 Unsupervised acoustic sub-word unit detection for query-by-example spoken term detection
abstract
In this paper we present a method for automatically generating acoustic sub-word units that can substitute conventional phone models in a query-by-example spoken term detection system. We generate the sub-word units with a modified version of our speaker diarization system. Given a speech recording, the original diarization system generates a set of speaker models in an unsupervised manner without the need for training or development data. Modifying the diarization system to process the speech of a single speaker and decreasing the minimum segment duration constraint allows us to detect speaker-dependent sub-word units. For the task of query-by-example spoken term detection, we show that the pro posed system performs well on both broadcast and non-broadcast recordings, unlike a conventional phone-based system trained solely on broadcast data. A mean average precision of 0.28 and 0.38 was obtained for experiments on broadcast news and on a set of war veteran interviews, respectively.
Marijn Huijbregts, Mitchell McLaren, David A. van Leeuwen
ICASSP3
2011 Source-normalised-and-weighted LDA for robust speaker recognition using i-vectors
abstract
The recently developed i-vector framework for speaker recognition has set a new performance standard in the research field. An i-vector is a compact representation of a speaker utterance extracted from a low-dimensional total variability subspace. Prior to classification using a cosine kernel, i-vectors are projected into an LDA space in order to reduce inter-session variability and enhance speaker discrimination. The accurate estimation of this LDA space from a training dataset is crucial to classification performance. A typical training dataset, however, does not consist of utterances acquired from all sources of interest (ie., telephone, microphone and interview speech sources) for each speaker. This has the effect of introducing source-related variation in the between-speaker covariance matrix and results in an incomplete representation of the within-speaker scatter matrix used for LDA. Proposed is a novel source-normalised-and-weighted LDA algorithm developed to improve the robustness of i-vector-based speaker recognition under both mis-matched evaluation conditions and conditions for which insufficient speech resources are available for adequate system development. Evaluated on the recent NIST 2008 and 2010 Speaker Recognition Evaluations (SRE), the proposed technique demonstrated improvements of up to 31% in minimum DCF and EER under mis-matched and sparsely-resourced conditions.
Mitchell McLaren, David A. van Leeuwen
ICASSP2
2011 Improved speaker recognition when using i-vectors from multiple speech sources
abstract
The concept of speaker recognition using i-vectors was recently introduced offering state-of-the-art performance. An i-vector is a compact representation of a speaker's utterance after projection into a low-dimensional, total variability subspace trained using factor analysis. A secondary process involving linear discriminant analysis (LDA) is then used to improve the discrimination of i-vectors from different speakers. The newness of this technology invokes the question as to the best way to train the total variability subspace and LDA matrix when using speech collected from distinctly different sources. This paper presents a comparative study of a number of subspace training techniques and a novel source-normalised and-weighted LDA algorithm for the purpose of improving i-vector based speaker recognition under mis-matched evaluation conditions. Results from the NIST 2010 speaker recognition evaluation (SRE) suggest that accounting for source conditions in the LDA matrix as opposed to the total variability subspace training regime provides improved robustness to mis-matched evaluation conditions.
Mitchell McLaren, David A. van Leeuwen
ICASSP2
2011 Diarization-Based Speaker Retrieval for Broadcast Television Archives
abstract
\n Contains fulltext :\n 94367.pdf (author's version ) (Open Access)\n
Marijn Huijbregts, David A. van Leeuwen
INTERSPEECH2
2011 A Speaker Line-Up for the Likelihood Ratio
abstract
\n Contains fulltext :\n 94223.pdf (author's version ) (Open Access)\n
David A. van Leeuwen, Niko Brümmer
INTERSPEECH1
2011 Evaluation of i-vector Speaker Recognition Systems for Forensic Application
abstract
This paper contributes a study on i-vector based speaker recognition systems and their application to forensics. The sensitivity of i-vector based speaker recognition is analyzed with respect to the effects of speech duration. This approach is motivated by the potentially limited speech available in a recording for a forensic case. In this context, the classification performance and calibration costs of the i-vector system are analyzed along with the role of normalization in the cosine kernel. Evaluated on the NIST SRE-2010 dataset, results highlight that normalization of the cosine kernel provided improved performance across all speech durations compared to the use of an unnormalized kernel. The normalized kernel was also found to play an important role in reducing miscalibration costs and providing wellcalibrated likelihood ratios with limited speech duration. Index Terms: i-vector, speaker recognition, forensics, calibration, short utterances
Miranti Indar Mandasari, Mitchell McLaren, David A. van Leeuwen
INTERSPEECH3
2011 To Weight or Not to Weight: Source-Normalised LDA for Speaker Recognition Using i-vectors
abstract
Source-normalised Linear Discriminant Analysis (SN-LDA) was recently introduced to improve speaker recognition using i-vectors extracted from multiple speech sources. SN-LDA normalises for the effect of speech source in the calculation of the between-speaker covariance matrix. Sourcenormalised-and-weighted (SNAW) LDA computes a weighted average of source-normalised covariance matrices to better exploit available information. This paper investigates the statistical significance of performance gains offered by SNAW-LDA over SN-LDA. An exhaustive search for optimal scatter weights was conducted to determine the potential benefit of SNAW-LDA. When evaluated on both NIST 2008 and 2010 SRE datasets, scatter-weighting in SNAW-LDA tended to overfit the LDA transform to the evaluation dataset while offering few statistically significant performance improvements over SN-LDA. Index Terms: speaker recognition, linear discriminant analysis, i-vector, source variability
Mitchell McLaren, David A. van Leeuwen
INTERSPEECH2
2011 An International English Speech Corpus for Longitudinal Study of Accent Development
abstract
If English is used intensively as a lingua franca in a multilanguage community, do speakers then converge towards a single common accent? This speech corpus allows for longitudinal study to investigate the question of convergence by means of repeated speech recordings of students at an English-language college over a period of 5 years. We describe the content and collection of the corpus and the type of research that is envisaged, as well as tools used to manage and analyze the recordings, including automatic phone recognition for prosodic analyses; and intelligibility experiments using the SRT method.
Rosemary Orr, Hugo Quené, Roeland van Beek, Thari Diefenbach, David A. van Leeuwen, Marijn Huijbregts
INTERSPEECH5
2009 The majority wins: a method for combining speaker diarization systems
abstract
In this paper we present a method for combining multiple diarization systems into one single system by applying a majority voting scheme. The voting scheme selects the best segmentation purely on basis of the output of each system. On our development set of NIST Rich Transcription evaluation meetings the voting method improves our system on all evaluation conditions. For the single distant microphone condition, DER performance improved by 7:8% (relative) compared to the best input system. For the multiple distant microphone condition the improvement is 3:6%. Index Terms: Speaker diarization
Marijn Huijbregts, David A. van Leeuwen, Franciska de Jong
INTERSPEECH2
2009 Speech overlap detection in a two-pass speaker diarization system
abstract
In this paper we present the two-pass speaker diarization system that we developed for the NIST RT09s evaluation. In the first pass of our system a model for speech overlap detection is generated automatically. This model is used in two ways to reduce the diarization errors due to overlapping speech. First, it is used in a second diarization pass to remove overlapping speech from the data while training the speaker models. Second, it is used to find speech overlap for the final segmentation so that overlapping speech segments can be generated. The experiments show that our overlap detection method improves the performance ofall three of our system configurations. Index Terms: Speaker diarization, speech overlap detection, Benchmark
Marijn Huijbregts, David A. van Leeuwen, Franciska de Jong
INTERSPEECH2
2009 Overall performance metrics for multi-condition speaker recognition evaluations
abstract
In this paper we propose a framework for measuring the overall\nperformance of an automatic speaker recognition system using\na set of trials of a heterogeneous evaluation such as NIST SRE-\n2008, which combines several acoustic conditions in one evalu-\nation. We do this by weighting trials of different conditions ac-\ncording to their relative proportion, and we derive expressions\nfor the basic speaker recognition performance measures Cdet,\nCllr, as well as the DET curve, from which EER and Cmin can det\nbe computed. Examples of pooling of conditions are shown on SRE-2008 data, including speaker sex and microphone type and speaking style.
David A. van Leeuwen
INTERSPEECH1
2009 Results of the n-best 2008 dutch speech recognition evaluation
abstract
In this paper we report the results of a Dutch speech recognition system evaluation held in 2008. The evaluation contained material in two domains: Broadcast News (BN) and Conversational Telephone Speech (CTS) and in two main accent regions (Flemish and Dutch). In total 7 sites submitted recognition results to the evaluation, totalling 58 different submissions in the various conditions. Best performances ranged from 15.9 % word error rate for BN, Flemish to 46.1 % for CTS, Flemish. This evaluation is the first of its kind for the Dutch language.
David A. van Leeuwen, Judith M. Kessens, Eric Sanders, Henk van den Heuvel
INTERSPEECH1
2009 A human benchmark for language recognition
abstract
In this study, we explore a human benchmark in language recognition, for the purpose of comparing human performance to machine performance in the context of the NIST LRE 2007.Humans are categorised in terms of language proficiency, and performance is presented per proficiency. Themain challenge in this work is the design of a test and application of a performance metric which allows a meaningful comparison of humans and machines. The main result of this work is that where subjects have lexical knowledge of a language, even at a low level, they perform as well as the state of the art in language recognition systems in 2007.
Rosemary Orr, David A. van Leeuwen
INTERSPEECH2
2009 Arousal and valence prediction in spontaneous emotional speech: felt versus perceived emotion
abstract
Contains fulltext : 91351.pdf (author's version ) (Open Access)
Khiet P. Truong, David A. van Leeuwen, Mark A. Neerincx, Franciska de Jong
INTERSPEECH2
2008 Assessing agreement of observer- and self-annotations in spontaneous multimodal emotion data
abstract
Contains fulltext : 91352.pdf (Publisher’s version ) (Open Access)
Khiet P. Truong, Mark A. Neerincx, David A. van Leeuwen
INTERSPEECH3
2007 STBU System for the NIST 2006 Speaker Recognition Evaluation
abstract
This paper describes STBU 2006 speaker recognition system, which performed well in the NIST 2006 speaker recognition evaluation. STBU is consortium of 4 partners: Spescom DataVoice (South Africa), TNO (Netherlands), BUT (Czech Republic) and University of Stellenbosch (South Africa). The primary system is a combination of three main kinds of systems: (1) GMM, with short-time MFCC or PLP features, (2) GMM-SVM, using GMM mean supervectors as input and (3) MLLR-SVM, using MLLR speaker adaptation coefficients derived from English LVCSR system. In this paper, we describe these sub-systems and present results for each system alone and in combination on the NIST Speaker Recognition Evaluation (SRE) 2006 development and evaluation data sets.
Pavel Matejka, Lukás Burget, Petr Schwarz, Ondrej Glembek, Martin Karafiát, Frantisek Grézl, Jan Cernocký, David A. van Leeuwen, Niko Brümmer, Albert Strasheim
ICASSP (4)8
2007 N-best: the northern- and southern-dutch benchmark evaluation of speech recognition technology
abstract
In this paper, we describe N-best 2008, the first Large Vocabulary Speech Recognition (LVCSR) benchmark evaluation held for the Dutch language. Both the accent as spoken in the Netherlands (Northern-Dutch) and in Belgium (Southern-Dutch or Flemish), will be evaluated. The evaluation tasks are broadcast news (BN) and conversational telephone speech (CTS). The N-best evaluation will take place in the spring of 2008 and is open to all research institutes and industries on voluntary basis. Thegoals of this first N-best evaluation is to define, set-up and conduct a Dutch LVCSR benchmark evaluation. In this paper, we will describe the state-of-the-art of Dutch LVCSR, recognition problems that are typical for the Dutch language, and the evaluation protocol.
Judith M. Kessens, David A. van Leeuwen
INTERSPEECH2
2007 An open-set detection evaluation methodology applied to language and emotion recognition
abstract
This paper introduces a detection methodology for recognition technologies in speech for which it is dif cult to obtain an abundance of non-target classes. An example is language recognition, where we would like to be able to measure the detection capability of a single target language without confounding with the modeling capability of non-target languages. The evaluation framework is based on a cross validation scheme leaving the non-target class out of the allowed training material for the detector. The framework allows us to use Detection Error Tradeoff curves properly. As another application example we apply the evaluation scheme to emotion recognition in order to obtain single-emotion detection performance assessment. Index Terms: detection methodology, open-set evaluation, language, emotion.
David A. van Leeuwen, Khiet P. Truong
INTERSPEECH1
2007 Design and characterization of the non-native military air traffic communications database (nnMATC)
abstract
This paper describes the speech database that has a central role in the Interspeech 2007 special session "Novel techniques for the NATO non-native Air Traffic Communications database." The rationale for recording and distributing this common research object is given, and details about the acquisition and annotation are given, as well as some statistics. Further, a summary is given of potential uses of the database, in terms of evaluation measures and protocols.
Stéphane Pigeon, Wade Shen, Aaron D. Lawson, David A. van Leeuwen
INTERSPEECH4
2007 Visualizing acoustic similarities between emotions in speech: an acoustic map of emotions
abstract
In this paper, we introduce a visual analysis method to assess the discriminability and confusiability between emotions according to automatic emotion classifiers. The degree of acoustic similarities between emotions can be defined in terms of distances that are based on pair-wise emotion discrimination experiments. By employing Multidimensional Scaling, the discriminability between emotions can then be visualized in a two-dimensional plot that is relatively easy to interpret. This ‘map of emotions’ is compared to the well-known ‘Feeltrace’ two-dimensional mapping of emotions. While there is correlation with the ‘arousal’ dimension of Feeltrace, it appears that the ‘valence’ dimension is difficult to relate to the acoustic map.
Khiet P. Truong, David A. van Leeuwen
INTERSPEECH2
2007 Automatic discrimination between laughter and speech
Khiet P. Truong, David A. van Leeuwen
Speech Commun.2
2007 Fusion of Heterogeneous Speaker Recognition Systems in the STBU Submission for the NIST Speaker Recognition Evaluation 2006
abstract
This paper describes and discusses the "STBU" speaker recognition system, which performed well in the NIST Speaker Recognition Evaluation 2006 (SRE). STBU is a consortium of four partners: Spescom DataVoice (Stellenbosch, South Africa), TNO (Soesterberg, The Netherlands), BUT (Brno, Czech Republic), and the University of Stellenbosch (Stellenbosch, South Africa). The STBU system was a combination of three main kinds of subsystems: 1) GMM, with short-time Mel frequency cepstral coefficient (MFCC) or perceptual linear prediction (PLP) features, 2) Gaussian mixture model-support vector machine (GMM-SVM), using GMM mean supervectors as input to an SVM, and 3) maximum-likelihood linear regression-support vector machine (MLLR-SVM), using MLLR speaker adaptation coefficients derived from an English large vocabulary continuous speech recognition (LVCSR) system. All subsystems made use of supervector subspace channel compensation methods-either eigenchannel adaptation or nuisance attribute projection. We document the design and performance of all subsystems, as well as their fusion and calibration via logistic regression. Finally, we also present a cross-site fusion that was done with several additional systems from other NIST SRE-2006 participants.
Niko Brümmer, Lukás Burget, Jan Cernocký, Ondrej Glembek, Frantisek Grézl, Martin Karafiát, David A. van Leeuwen, Pavel Matejka, Petr Schwarz, Albert Strasheim
IEEE Trans. Speech Audio Process.7
2006 NIST and NFI-TNO evaluations of automatic speaker recognition
David A. van Leeuwen, Alvin F. Martin, Mark A. Przybocki, Jos S. Bouten
Comput. Speech Lang.1
2005 Speaker adaptation in the NIST speaker recognition evaluation 2004
abstract
New in the 2004 edition of the NIST Speaker Recognition Evaluation (SRE) was the condition where unsupervised adaptation of speaker models is allowed. Despite the promising results on development test material, hardly any beneficial results were obtained in the Evaluation itself. An analysis is made why this was the case, and it appears that a mimimum level of performance is essential to obtain results using adaptation that improve on the performance without adaptation. Further, the system should be well calibrated. For the conditions with 8 conversation sides we have been able to find improvement using unsupervised adaptation using the NIST 2004 evaluation, both for an UBM/GMM adaptation methodology, and a novel SVM adaptation methodology. The minimum DCF for a fused system drops from 0.259 for the unadapted condition to 0.231 for the adapted condition. 1.
David A. van Leeuwen
INTERSPEECH1
2005 Automatic detection of laughter
abstract
In the context of detecting ‘paralinguistic events’ with the aim to make classification of the speaker’s emotional state possible, a detector was developed for one of the most obvious ‘paralinguistic events’, namely laughter. Gaussian Mixture Models were trained with Perceptual Linear Prediction features, pitch&energy, pitch&voicing and modulation spectrum features to model laughter and speech. Data from the ICSI Meeting Corpus and the Dutch CGN corpus were used for our classification experiments. The results showed that Gaussian Mixture Models trained with Perceptual Linear Prediction features performed best with Equal Error Rates ranging from 7.1%-20.0%.
Khiet P. Truong, David A. van Leeuwen
INTERSPEECH2
2003 Speaker verification systems and security considerations
abstract
In speaker verification technology, the security considerations are quite dierent from performance measures that are usually studied. The security level of a system is generally expressed in the amount of eort it takes to have a successful break-in attempt. This paper discusses potential weaknesses of speaker verification systems and methods of exploiting these weaknesses, and suggests proper experiments for determining the security level of a speaker verification system.
David A. van Leeuwen
INTERSPEECH1
2002 "Do as I Say! . But Who Says What I Should Say - or Do?" On the Definition of a Standard Spoken Command Vocabulary for ICT Devices and Services
Bruno von Niman, Catriona Chaplin, Jose-Antonio Collado-Vega, Lutz Groh, Scott McGlashan, Wally Mellors, David A. van Leeuwen
Mobile HCI7
2000 Automatic speech recognition of non-native speakers using consonant-vowel-consonant (CVC) words
abstract
In this study we investigate whether non-native speakers using a speech recognition system would benefit from phone models of their own native language. For Dutch as the target recognition language, we found that American speakers do not in general benefit from American English models when speaking Dutch. However, using a CVC test methodology, we can conclude that for a certain level of proficiency, and for certain phones, there is a small beneficial effect of using American acoustic models for American non-native speakers of Dutch.
David A. van Leeuwen, Sander J. van Wijngaarden
INTERSPEECH1
1999 Objective and subjective evaluation of the acoustic models of a continuous speech recognition system
David A. van Leeuwen, Michael de Louwere
EUROSPEECH1
1997 Within-speaker variability of the word error rate for a continuous speech recognition system
abstract
WITHIN-SPEAKERVARIABILITYOFTHEWORDERRORRATEFORACONTINUOUSSPEECHRECOGNITIONSYSTEMDavid A. van Leeuwen and Herman J. M. SteenekenElectronic mail:fvanLeeuwen;[email protected] Human Factors Research Institute.Postbus 23,3769 ZG So esterb erg,The Netherlands.ABSTRACTThevarianceofthep erformanceacontinuoussp eechrecognitionsystemsub jectedtoreplicaut-terancesofthesamesentencesp okenbysp eaker has b een investigated.In an exp eriment withthree di erent sp eech recognition systems in three dif-ferentlanguageswithwodi erengrammarcondi-tions it is shown that the sentence word error rate hasavariance that can b e describ ed in terms of binomialstatistics.The distribution of the measured varianceshows a remarkable corresp ondence to the parameter-free theoretical distribution.It is therefore concludedthatfortheworderrorrateofacontinuoussp eechrecognition system binomial statistics apply.INTRODUCTIONThe word error rate (sometimes expressed in its com-plement, the accuracy) is the most widely used mea-sureofthep erformancesp eechrecognitionsys-tems.Traditionally, for isolated word recognizers thismeasure has b een one which leaves little argument forinterpretation, but for continuous sp eech recognitionsystemsthesituationismorecomplex.Becauseofthe nature of natural sp eech the words are connectedto a long string.This makes it somewhat dicult topinp oint the exact lo cation of an error in case of mis-recognition and consequently makes it hard to countthe numb er of erroneous words.Evaluating the cor-rectness of utterance as a whole, measured in the ut-terance (or sentence) error rate resolves this problem.However, this measure needs much more sp eech mate-rial b efore an accurate gure is found, and researchersoften use the word error rate b ecause it is more sensi-tive to small changes in the p erformance of the sp eechrecognition system.One of the questions wewant to address in thispap er,ishowaccurateameasurementoftheorderror rate is for a continuous sp eech recognition sys-tem.Forarepresentativeevaluation,onegenerallywantstohaveawidecoerageoflanguage,andincaseofasp eakerindep endentsystem,widecov-erageofsp eakers.Becauseb othsetsarevirtuallyin niteinsize,foreachevaluationnewsamplesaredrawn from the sets of language material (sentences)and sp eakers.If there are ways to quantify the accu-racy of a word error rate measure, and ob jectivewaysto calibrate the `diculty' of the test material [1], anewevaluationcansuccessfullybecomparedtoanearlier one.EXPERIMENTAL SETUPIn order to study the inherentvariability of the p er-formance of a continuous sp eech recognizer, we p er-formed a test with no variability in sp eaker and sp o-ken text.This exp erimentwas carried out as an addi-tional test in the pro jectSqale, whichwas a pro jectcompared sp eecrecognition indi erentEuro-p ean languages and for di erent systems [1, 2].Thevariabilityinsp eakerandsp eechcontentwasmadezero byhaving a sp eaker read out the same sentenceseveral times,ofwhichwecalltheindividualutter-ancesreplicasofthesamesentence.(Theserepli-cascaninprinciplebeusedtomeasure thewithin-sp eaker variability.)The replicas were recorded dur-ingarecordingsessionoftheevaluationtestSqale,andwerespreadamongthenormalevalu-ation sentences.The sp eakers were prepared for theo ccurrence of replicas, and were requested to read outa replica as if it was the rst o ccurrence in order tomaketheutterancesasmuchalikp ossible.Wchose for 5 replicas of one sentence for each recordedsp eaker; more replicas mighthave stretched the sub-ject'sacceptancelimitsto ofar,andwedidnotwant thatthereading style of theother (evaluationtest) utterances was inuenced by this test.Table.The numb er of sentences available,for each lan-guage.Each sp eaker, having its own sentence, uttered 5replicas.The number of speech recognition systems avail-able p er language is also indicated, as well as the amountof measurement p oints resulting.LanguageAmericanBritishGermanEnglishsentences3710systems32grammars2data p oints184240Thereplicautteranceswererecordedthreedi erentlanguages, in amounts according to the ta-
David A. van Leeuwen, Herman J. M. Steeneken
EUROSPEECH1
1997 Speaker recognition by humans and machines
Herman J. M. Steeneken, David A. van Leeuwen
EUROSPEECH2
1997 Multilingual large vocabulary speech recognition: the European SQALE project
Steve J. Young, Martine Adda-Decker, Xavier L. Aubert, Christian Dugast, Jean-Luc Gauvain, Dan J. Kershaw, Lori Lamel, David A. van Leeuwen, David Pye, Anthony J. Robinson, Herman J. M. Steeneken, Philip C. Woodland
Comput. Speech Lang.8
1995 Human benchmarks for speaker independent large vocabulary recognition performance
David A. van Leeuwen, Leo-Geert van den Berg, Herman J. M. Steeneken
EUROSPEECH1
1995 Multi-lingual assessment of speaker independent large vocabulary speech-recognition systems: THE SQALE-PROJECT
Herman J. M. Steeneken, David A. van Leeuwen
EUROSPEECH2