EDBT 2026 Demo / reviewers in the wild / expert
Tomi Kinnunen
dblp:94/4754 · also Tomi H. Kinnunen
· DBLP profile ↗
151ranked-venue papers
28as first author
37since 2021 · last 2026
0000-0002-4371-7322ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 119 · 21 first-author · 28 since 2021Artificial intelligence and machine learning · 104 · 18 first-author · 27 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards explainable spoofed speech attribution and detection: A probabilistic approach for characterizing speech synthesizer componentsabstractWe propose an explainable probabilistic framework for characterizing spoofed speech by decomposing it into probabilistic attribute embeddings. Unlike raw high-dimensional countermeasure embeddings, which lack interpretability, the proposed probabilistic attribute embeddings aim to detect specific speech synthesizer components, represented through high-level attributes and their corresponding values. We use these probabilistic embeddings with four classifier back-ends to address two downstream tasks: spoofing detection and spoofing attack attribution. The former is the well-known bonafide-spoof detection task, whereas the latter seeks to identify the source method (generator) of a spoofed utterance. We additionally use Shapley values, a widely used technique in machine learning, to quantify the relative contribution of each attribute value to the decision-making process in each task. Results on the ASVspoof2019 dataset demonstrate the substantial role of waveform generator, conversion model outputs, and inputs in spoofing detection; and inputs, speaker, and duration modeling in spoofing attack attribution. In the detection task, the probabilistic attribute embeddings achieve 99.7% balanced accuracy and 0.22% equal error rate (EER), closely matching the performance of raw embeddings (99.9% balanced accuracy and 0.22% EER). Similarly, in the attribution task, our embeddings achieve 90.23% balanced accuracy and 2.07% EER, compared to 90.16% and 2.11% with raw embeddings. These results demonstrate that the proposed framework is both inherently explainable by design and capable of achieving performance comparable to raw CM embeddings. Jagabandhu Mishra, Manasi Chhibber, Hye-Jin Shim, Tomi Kinnunen |
Comput. Speech Lang. | 4 |
| 2026 | Causal analysis of ASR errors for children: Quantifying the impact of physiological, cognitive, and extrinsic factorsabstractThe increasing use of children’s automatic speech recognition (ASR) systems has spurred research efforts to improve the accuracy of models designed for children’s speech in recent years. The current approach utilizes either open-source speech foundation models (SFMs) directly or fine-tuning them with children’s speech data. These SFMs, whether open-source or fine-tuned for children, often exhibit higher word error rates (WERs) compared to adult speech. However, there is a lack of systemic analysis of the cause of this degraded performance of SFMs. Understanding and addressing the reasons behind this performance disparity is crucial for improving the accuracy of SFMs for children’s speech. Our study addresses this gap by investigating the causes of accuracy degradation and the primary contributors to WER in children’s speech. In the first part of the study, we conduct a comprehensive benchmarking study on two self-supervised SFMs ( Wav2Vec2.0 and Hubert ) and two weakly supervised SFMs ( Whisper and Massively Multilingual Speech (MMS) ) across various age groups on two children speech corpora, establishing the raw data for the causal inference analysis in the second part. In the second part of the study, we analyze the impact of physiological factors (age, gender), cognitive factors (pronunciation ability), and external factors (vocabulary difficulty, background noise, and word count) on SFM accuracy in children’s speech using causal inference. The results indicate that physiology (age) and particular external factor (number of words in audio) have the highest impact on accuracy, followed by background noise and pronunciation ability. Fine-tuning SFMs on children’s speech reduces sensitivity to physiological and cognitive factors, while sensitivity to the number of words in audio persists. Vishwanath Pratap Singh, Md. Sahidullah, Tomi Kinnunen |
Comput. Speech Lang. | 3 |
| 2026 | ASVspoof 5: Design, collection and validation of resources for spoofing, deepfake, and adversarial attack detection using crowdsourced speechabstractASVspoof 5 is the fifth edition in a series of challenges which promote the study of speech spoofing and deepfake attacks as well as the design of detection solutions. We introduce the ASVspoof 5 database which is generated in a crowdsourced fashion from data collected in diverse acoustic conditions (cf. studio-quality data for earlier ASVspoof databases) and from ∼ 2,000 speakers (cf. ∼ 100 earlier). The database contains attacks generated with 32 different algorithms, also crowdsourced, and optimised to varying degrees using new surrogate detection models. Among them are attacks generated with a mix of legacy and contemporary text-to-speech synthesis and voice conversion models, in addition to adversarial attacks which are incorporated for the first time. ASVspoof 5 protocols comprise seven speaker-disjoint partitions. They include two distinct partitions for the training of different sets of attack models, two more for the development and evaluation of surrogate detection models, and then three additional partitions which comprise the ASVspoof 5 training, development and evaluation sets. An auxiliary set of data collected from an additional 30k speakers can also be used to train speaker encoders for the implementation of attack algorithms. Also described herein is an experimental validation of the new ASVspoof 5 database using a set of automatic speaker verification and spoof/deepfake baseline detectors. With the exception of protocols and tools for the generation of spoofed/deepfake speech, the resources described in this paper, already used by participants of the ASVspoof 5 challenge in 2024, are now all freely available to the community. Xin Wang 0037, Héctor Delgado, Hemlata Tak, Jee-Weon Jung, Hye-Jin Shim, Massimiliano Todisco, Ivan Kukanov, Xuechen Liu 0001, Md. Sahidullah, Tomi Kinnunen, Nicholas W. D. Evans, Kong-Aik Lee, Junichi Yamagishi, Myeonghun Jeong, Yongyi Zang, Soumi Maiti, Florian Lux, Nicolas Müller, Wangyou Zhang, Chengzhe Sun 0001, Shuwei Hou, Siwei Lyu, Sébastien Le Maguer, Hanjie Guo, Vishwanath Pratap Singh |
Comput. Speech Lang. | 10 |
| 2025 | Fake-Mamba: Real-Time Speech Deepfake Detection Using Bidirectional Mamba as Self-Attention's AlternativeabstractAdvances in speech synthesis intensify security threats, motivating real-time deepfake detection research. We investigate whether bidirectional Mamba can serve as a competitive alternative to Self-Attention in detecting synthetic speech. Our solution, Fake-Mamba, integrates an XLSR front-end with bidirectional Mamba to capture both local and global artifacts. Our core innovation introduces three efficient encoders: TransBiMamba, ConBiMamba, and PN-BiMamba. Leveraging XLSR’s rich linguistic representations, PN-BiMamba can effectively capture the subtle cues of synthetic speech. Evaluated on ASVspoof 21 LA, 21 DF, and In-The-Wild benchmarks, FakeMamba achieves 0.97 %, 1.74 %, and 5.85% EER, respectively, representing substantial relative gains over SOTA models XLSRConformer and XLSR-Mamba. The framework maintains realtime inference across utterance lengths, demonstrating strong generalization and practical viability. The code is available at https://github.com/xuanxixi/Fake-Mamba. Xi Xuan, Zimo Zhu, Wenxin Zhang 0005, Yi-Cheng Lin, Tomi Kinnunen |
ASRU | 5 |
| 2025 | An Explainable Probabilistic Attribute Embedding Approach for Spoofed Speech CharacterizationabstractWe propose a novel approach for spoofed speech characterization through explainable probabilistic attribute embeddings. In contrast to high-dimensional raw embeddings extracted from a spoofing countermeasure (CM) whose dimensions are not easy to interpret, the probabilistic attributes are designed to gauge the presence or absence of sub-components that make up a specific spoofing attack. These attributes are then applied to two downstream tasks: spoofing detection and attack attribution. To enforce interpretability also to the back-end, we adopt a decision tree classifier. Our experiments on the ASVspoof2019 dataset with spoof CM embeddings extracted from three models (AASIST, Rawboost-AASIST, SSL-AASIST) suggest that the performance of the attribute embeddings are on par with the original raw spoof CM embeddings for both tasks. The best performance achieved with the proposed approach for spoofing detection and attack attribution, in terms of accuracy, is 99.7% and 99.2%, respectively, compared to 99.7% and 94.7% using the raw CM embeddings. To analyze the relative contribution of each attribute, we estimate their Shapley values. Attributes related to acoustic feature prediction, waveform generation (vocoder), and speaker modeling are found important for spoofing detection; while duration modeling, vocoder, and input type play a role in spoofing attack attribution. Manasi Chhibber, Jagabandhu Mishra, Hye-Jin Shim, Tomi Kinnunen |
ICASSP | 4 |
| 2025 | Explaining Speaker and Spoof Embeddings via ProbingabstractThis study investigates the explainability of embedding representations, specifically those used in modern audio spoofing detection systems based on deep neural networks, known as spoof embeddings. Building on established work in speaker embedding explainability, we examine how well these spoof embeddings capture speaker-related information. We train simple neural classifiers using either speaker or spoof embeddings as input, with speaker-related attributes as target labels. These attributes are categorized into two groups: metadata-based traits (e.g., gender, age) and acoustic traits (e.g., fundamental frequency, speaking rate). Our experiments on the ASVspoof 2019 LA evaluation set demonstrate that spoof embeddings preserve several key traits, including gender, speaking rate, F0, and duration. Further analysis of gender and speaking rate indicates that the spoofing detector partially preserves these traits, potentially to ensure the decision process remains robust against them. Xuechen Liu 0001, Junichi Yamagishi, Md. Sahidullah, Tomi Kinnunen |
ICASSP | 4 |
| 2025 | Continuous Learning for Children's ASR: Overcoming Catastrophic Forgetting with Elastic Weight Consolidation and Synaptic Intelligence
Edem Ahadzi, Vishwanath Pratap Singh, Tomi Kinnunen, Ville Hautamäki |
INTERSPEECH | 3 |
| 2025 | STOPA: A Dataset of Systematic VariaTion Of DeePfake Audio for Open-Set Source Tracing and AttributionabstractA key research area in deepfake speech detection is source tracing - determining the origin of synthesised utterances. The approaches may involve identifying the acoustic model (AM), vocoder model (VM), or other generation-specific parameters. However, progress is limited by the lack of a dedicated, systematically curated dataset. To address this, we introduce STOPA, a systematically varied and metadata-rich dataset for deepfake speech source tracing, covering 8 AMs, 6 VMs, and diverse parameter settings across 700k samples from 13 distinct synthesisers. Unlike existing datasets, which often feature limited variation or sparse metadata, STOPA provides a systematically controlled framework covering a broader range of generative factors, such as the choice of the vocoder model, acoustic model, or pretrained weights, ensuring higher attribution reliability. This control improves attribution accuracy, aiding forensic analysis, deepfake detection, and generative model transparency. Anton Firc, Manasi Chhibber, Jagabandhu Mishra, Vishwanath Pratap Singh, Tomi Kinnunen, Kamil Malinka |
INTERSPEECH | 5 |
| 2025 | FROST-EMA: Finnish and Russian Oral Speech Dataset of Electromagnetic Articulography Measurements with L1, L2 and Imitated L2 Accents
Satu Hopponen, Tomi Kinnunen, Alexandre Nikolaev, Rosa González Hautamäki, Lauri Tavi, Einar Meister |
INTERSPEECH | 2 |
| 2025 | Causal Structure Discovery for Error Diagnostics of Children's ASR
Vishwanath Pratap Singh, Md. Sahidullah, Tomi Kinnunen |
INTERSPEECH | 3 |
| 2025 | Optimizing a-DCF for Spoofing-Robust Speaker VerificationabstractAutomatic speaker verification (ASV) systems are vulnerable to spoofing attacks. We propose a spoofing-robust ASV system optimized directly for the recently introduced architecture-agnostic detection cost function (a-DCF), which allows targeting a desired trade-off between the contradicting aims of user convenience and robustness to spoofing. We combine a-DCF and binary cross-entropy (BCE) with a novel straightforward threshold optimization technique. Our results with an embedding fusion system on ASVspoof2019 data demonstrate relative improvement of 13% over a system trained using BCE only (from minimum a-DCF of 0.1445 to 0.1254). Using an alternative non-linear score fusion approach provides relative improvement of 43% (from minimum a-DCF of 0.0508 to 0.0289). Oguzhan Kurnaz, Jagabandhu Mishra, Tomi Kinnunen, Cemal Hanilçi |
IEEE Signal Process. Lett. | 3 |
| 2024 | Revisiting and Improving Scoring Fusion for Spoofing-aware Speaker Verification Using Compositional Data AnalysisabstractInterspeech 2024, 1-5 September 2024, Kos, Greece Xin Wang 0037, Tomi Kinnunen, Kong-Aik Lee, Paul-Gauthier Noé, Junichi Yamagishi |
INTERSPEECH | 2 |
| 2024 | Speaker Detection by the Individual Listener and the Crowd: Parametric Models Applicable to Bonafide and Deepfake Speech
Tomi Kinnunen, Rosa González Hautamäki, Xin Wang 0037, Junichi Yamagishi |
INTERSPEECH | 1 |
| 2024 | ROAR: Reinforcing Original to Augmented Data Ratio Dynamics for Wav2vec2.0 Based ASR
Vishwanath Pratap Singh, Federico Malato, Ville Hautamäki, Md. Sahidullah, Tomi Kinnunen |
INTERSPEECH | 5 |
| 2024 | Meta-Learning Approaches For Improving Detection of Unseen Speech DeepfakesabstractCurrent speech deepfake detection approaches perform satisfactorily against known adversaries; however, generalization to unseen attacks remains an open challenge. The proliferation of speech deepfakes on social media underscores the need for systems that can generalize to unseen attacks not observed during training. We address this problem from the perspective of meta-learning, aiming to learn attack-invariant features to adapt to unseen attacks with very few samples available. This approach is promising since generating of a high-scale training dataset is often expensive or infeasible. Our experiments demonstrated an improvement in the Equal Error Rate (EER) from 21.67% to 10.42% on the InTheWild dataset, using just 96 samples from the unseen dataset. Continuous few-shot adaptation ensures that the system remains up-to-date. Ivan Kukanov, Janne Laakkonen, Tomi Kinnunen, Ville Hautamäki |
SLT | 3 |
| 2024 | t-EER: Parameter-Free Tandem Evaluation of Countermeasures and Biometric ComparatorsabstractPresentation attack (spoofing) detection (PAD) typically operates alongside biometric verification to improve reliablity in the face of spoofing attacks. Even though the two sub-systems operate in tandem to solve the single task of reliable biometric verification, they address different detection tasks and are hence typically evaluated separately. Evidence shows that this approach is suboptimal. We introduce a new metric for the joint evaluation of PAD solutions operating in situ with biometric verification. In contrast to the tandem detection cost function proposed recently, the new tandem equal error rate (t-EER) is parameter free. The combination of two classifiers nonetheless leads to a set of operating points at which false alarm and miss rates are equal and also dependent upon the prevalence of attacks. We therefore introduce the concurrent t-EER, a unique operating point which is invariable to the prevalence of attacks. Using both modality (and even application) agnostic simulated scores, as well as real scores for a voice biometrics application, we demonstrate application of the t-EER to a wide range of biometric system evaluations under attack. The proposed approach is a strong candidate metric for the tandem evaluation of PAD systems and biometric comparators. Tomi Kinnunen, Kong-Aik Lee, Hemlata Tak, Nicholas W. D. Evans, Andreas Nautsch |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Generalizing Speaker Verification for Spoof Awareness in the Embedding SpaceabstractIt is now well-known thatautomatic speaker verification(ASV) systems can be spoofed using various types of adversaries. The usual approach to counteract ASV systems against such attacks is to develop a separate spoofingcountermeasure(CM) module to classify speech input either as a bonafide, or a spoofed utterance. Nevertheless, such a design requires additional computation and utilization efforts at the authentication stage. An alternative strategy involves a single monolithic ASV system designed to handle both zero-effort imposter (non-targets) and spoofing attacks. Suchspoof-awareASV systems have the potential to provide stronger protections and more economic computations. To this end, we propose to generalize the standalone ASV (G-SASV) against spoofing attacks, where we leverage limited training data from CM to enhance a simple backend in the embedding space, without the involvement of a separate CM module during the test (authentication) phase. We propose a novel yet simple backend classifier based on deep neural networks and conduct the study via domain adaptation and multi-task integration of spoof embeddings at the training stage. Experiments are conducted on the ASVspoof 2019 logical access dataset, where we improve the performance of statistical ASV backends on the joint (bonafide and spoofed) and spoofed conditions by a maximum of 36.2% and 49.8% in terms of equal error rates, respectively. Xuechen Liu 0001, Md. Sahidullah, Kong-Aik Lee, Tomi Kinnunen |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Learnable Frontends That Do Not Learn: Quantifying Sensitivity To Filterbank InitialisationabstractWhile much of modern speech and audio processing relies on deep neural networks trained using fixed audio representations, recent studies suggest great potential in acoustic frontends learnt jointly with a backend. In this study, we focus specifically on learnable filterbanks. Prior studies have reported that in frontends using learnable filterbanks initialised to a mel scale, the learned filters do not differ substantially from their initialisation. Using a Gabor-based filterbank, we investigate the sensitivity of a learnable filterbank to its initialisation using several initialisation strategies on two audio tasks: voice activity detection and bird species identification. We use the Jensen-Shannon Distance and analysis of the learned filters before and after training. We show that although performance is overall improved, the filterbanks exhibit strong sensitivity to their initialisation strategy. The limited movement from initialised values suggests that alternate optimisation strategies may allow a learnable frontend to reach better overall performance. Mark Anderson 0006, Tomi Kinnunen, Naomi Harte |
ICASSP | 2 |
| 2023 | Speaker-Aware Anti-spoofing
Xuechen Liu 0001, Md. Sahidullah, Kong-Aik Lee, Tomi Kinnunen |
INTERSPEECH | 4 |
| 2023 | Towards Single Integrated Spoofing-aware Speaker Verification Embeddings
Sung Hwan Mun, Hye-Jin Shim, Hemlata Tak, Xin Wang 0037, Xuechen Liu 0001, Md. Sahidullah, Myeonghun Jeong, Min Hyun Han, Massimiliano Todisco, Kong-Aik Lee, Junichi Yamagishi, Nicholas W. D. Evans, Tomi Kinnunen, Nam Soo Kim, Jee-Weon Jung |
INTERSPEECH | 13 |
| 2023 | How to Construct Perfect and Worse-than-Coin-Flip Spoofing Countermeasures: A Word of Warning on Shortcut Learning
Hye-Jin Shim, Rosa González Hautamäki, Md. Sahidullah, Tomi Kinnunen |
INTERSPEECH | 4 |
| 2023 | Multi-Dataset Co-Training with Sharpness-Aware Optimization for Audio Anti-spoofing
Hye-Jin Shim, Jee-Weon Jung, Tomi Kinnunen |
INTERSPEECH | 3 |
| 2023 | Speaker Verification Across Ages: Investigating Deep Speaker Embedding Sensitivity to Age Mismatch in Enrollment and Test Speech
Vishwanath Pratap Singh, Md. Sahidullah, Tomi Kinnunen |
INTERSPEECH | 3 |
| 2023 | ASVspoof 2021: Towards Spoofed and Deepfake Speech Detection in the WildabstractBenchmarking initiatives support the meaningful comparison of competing solutions to prominent problems in speech and language processing. Successive benchmarking evaluations typically reflect a progressive evolution from ideal lab conditions towards to those encountered in the wild. ASVspoof, the spoofing and deepfake detection initiative and challenge series, has followed the same trend. This article provides a summary of the ASVspoof 2021 challenge and the results of 54 participating teams that submitted to the evaluation phase. For the logical access (LA) task, results indicate that countermeasures are robust to newly introduced encoding and transmission effects. Results for the physical access (PA) task indicate the potential to detect replay attacks in real, as opposed to simulated physical spaces, but a lack of robustness to variations between simulated and real acoustic environments. The Deepfake (DF) task, new to the 2021 edition, targets solutions to the detection of manipulated, compressed speech data posted online. While detection solutions offer some resilience to compression effects, they lack generalization across different source datasets. In addition to a summary of the top-performing systems for each task, new analyses of influential data factors and results for hidden data subsets, the article includes a review of post-challenge results, an outline of the principal challenge limitations and a road-map for the future of ASVspoof. Xuechen Liu 0001, Xin Wang 0037, Md. Sahidullah, Jose Patino 0001, Héctor Delgado, Tomi Kinnunen, Massimiliano Todisco, Junichi Yamagishi, Nicholas W. D. Evans, Andreas Nautsch, Kong-Aik Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2023 | GAN-Aimbots: Using Machine Learning for Cheating in First Person ShootersabstractPlaying games with cheaters is not fun, and in a multi-billion-dollar video game industry with hundreds of millions of players, game developers aim to improve the security and, consequently, the user experience of their games by preventing cheating. Both traditional software-based methods and statistical systems have been successful in protecting against cheating, but recent advances in the automatic generation of content, such as images or speech, threaten the video game industry; they could be used to generate artificial gameplay indistinguishable from that of legitimate human players. To better understand this threat, we begin by reviewing the current state of multiplayer video game cheating, and then proceed to build a proof-of-concept method, GAN-Aimbot. By gathering data from various players in a first-person shooter game we show that the method improves players performance while remaining hidden from automatic and manual protection mechanisms. By sharing this work we hope to raise awareness on this issue and encourage further research into protecting the gaming communities. Anssi Kanervisto, Tomi Kinnunen, Ville Hautamäki |
IEEE Trans. Games | 2 |
| 2022 | Learnable Nonlinear Compression for Robust Speaker VerificationabstractIn this study, we focus on nonlinear compression methods in spectral features for speaker verification based on deep neural network. We consider different kinds of channel-dependent (CD) nonlinear compression methods optimized in a data-driven manner. Our methods are based on power nonlinearities and dynamic range compression (DRC). We also propose multi-regime (MR) design on the nonlinearities, at improving robustness. Results on VoxCeleb1 and VoxMovies data demonstrate improvements brought by proposed compression methods over both the commonly-used logarithm and their static counterparts, especially for ones based on power function. While CD generalization improves performance on VoxCeleb1, MR provides more robustness on VoxMovies, with a maximum relative equal error rate reduction of 21.6%. Xuechen Liu 0001, Md. Sahidullah, Tomi Kinnunen |
ICASSP | 3 |
| 2022 | SASV 2022: The First Spoofing-Aware Speaker Verification ChallengeabstractThe first spoofing-aware speaker verification (SASV) challenge aims to integrate research efforts in speaker verification and anti-spoofing.We extend the speaker verification scenario by introducing spoofed trials to the usual set of target and impostor trials.In contrast to the established ASVspoof challenge where the focus is upon separate, independently optimised spoofing detection and speaker verification sub-systems, SASV targets the development of integrated and jointly optimised solutions.Pre-trained spoofing detection and speaker verification models are provided as open source and are used in two baseline SASV solutions.Both models and baselines are freely available to participants and can be used to develop back-end fusion approaches or end-to-end solutions.Using the provided common evaluation protocol, 23 teams submitted SASV solutions.When assessed with target, bona fide non-target and spoofed non-target trials, the top-performing system reduces the equal error rate of a conventional speaker verification system from 23.83% to 0.13%.SASV challenge results are a testament to the reliability of today's state-of-the-art approaches to spoofing detection and speaker verification. Jee-Weon Jung, Hemlata Tak, Hye-Jin Shim, Hee-Soo Heo, Bong-Jin Lee, Soo-Whan Chung, Ha-Jin Yu, Nicholas W. D. Evans, Tomi Kinnunen |
INTERSPEECH | 9 |
| 2022 | Improving speaker de-identification with functional data analysis of f0 trajectoriesabstractDue to a constantly increasing amount of speech data that is stored in different types of databases, voice privacy has become a major concern. To respond to such concern, speech researchers have developed various methods for speaker de-identification. The state-of-the-art solutions utilize deep learning solutions which can be effective but might be unavailable or impractical to apply for, for example, under-resourced languages. Formant modification is a simpler, yet effective method for speaker de-identification which requires no training data. Still, remaining intonational patterns in formant-anonymized speech may contain speaker-dependent cues. This study introduces a novel speaker de-identification method, which, in addition to simple formant shifts, manipulates f0 trajectories based on functional data analysis. The proposed speaker de-identification method will conceal plausibly identifying pitch characteristics in a phonetically controllable manner and improve formant-based speaker de-identification up to 25%. Lauri Tavi, Tomi Kinnunen, Rosa González Hautamäki |
Speech Commun. | 2 |
| 2022 | Optimizing Tandem Speaker Verification and Anti-Spoofing SystemsabstractAs automatic speaker verification (ASV) systems are vulnerable to spoofing attacks, they are typically used in conjunction with spoofing countermeasure (CM) systems to improve security. For example, the CM can first determine whether the input is human speech, then the ASV can determine whether this speech matches the speakers identity. The performance of such a tandem system can be measured with a tandem detection cost function (t-DCF). However, ASV and CM systems are usually trained separately, using different metrics and data, which does not optimize their combined performance. In this work, we propose to optimize the tandem system directly by creating a differentiable version of t-DCF and employing techniques from reinforcement learning. The results indicate that these approaches offer better outcomes than finetuning, with our method providing a 20\% relative improvement in the t-DCF in the ASVSpoof19 dataset in a constrained setting. Anssi Kanervisto, Ville Hautamäki, Tomi Kinnunen, Junichi Yamagishi |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Voxceleb Enrichment for Age and Gender RecognitionabstractVoxCeleb datasets are widely used in speaker recognition studies. Our work serves two purposes. First, we provide speaker age labels and (an alternative) annotation of speaker gender. Second, we demonstrate the use of this metadata by constructing age and gender recognition models with different features and classifiers. We query different celebrity databases and apply consensus rules to derive age and gender labels. We also compare the original VoxCeleb gender labels with our labels to identify records that might be mislabeled in the original VoxCeleb data. On modeling side, we design a comprehensive study of multiple features and models for recognizing gender and age. Our best system, using i-vector features, achieved an F1-score of 0.9829 for gender recognition task using logistic regression, and the lowest mean absolute error (MAE) in age regression, 9.443 years, is obtained with ridge regression. This indicates challenge in age estimation from in-the-wild style speech data. Khaled Hechmi, Trung Ngo Trong, Ville Hautamäki, Tomi Kinnunen |
ASRU | 4 |
| 2021 | Optimized Power Normalized Cepstral Coefficients Towards Robust Deep Speaker VerificationabstractAfter their introduction to robust speech recognition, power normalized cepstral coefficient (PNCC) features were successfully adopted to other tasks, including speaker verification. However, as a feature extractor with long-term operations on the power spectrogram, its temporal processing and amplitude scaling steps dedicated on environmental compensation may be redundant. Further, they might suppress intrinsic speaker variations that are useful for speaker verification based on deep neural networks (DNN). Therefore, in this study, we revisit and optimize PNCCs by ablating its medium-time processor and by introducing channel energy normalization. Experimental results with a DNN-based speaker verification system indicate substantial improvement over baseline PNCCs on both in-domain and cross-domain scenarios, reflected by relatively 5.8% and 61.2% maximum lower equal error rate on VoxCelebl and VoxMovies, respectively. Xuechen Liu 0001, Md. Sahidullah, Tomi Kinnunen |
ASRU | 3 |
| 2021 | Parameterized Channel Normalization for Far-Field Deep Speaker VerificationabstractWe address far-field speaker verification with deep neural network (DNN) based speaker embedding extractor, where mismatch between enrollment and test data often comes from convolutive effects (e.g. room reverberation) and noise. To mitigate these effects, we focus on two parametric normalization methods: per-channel energy normalization (PCEN) and parameterized cepstral mean normalization (PCMN). Both methods contain differentiable parameters and thus can be conveniently integrated to, and jointly optimized with the DNN using automatic differentiation methods. We consider both fixed and trainable (data-driven) variants of each method. We evaluate the performance on Hi-MIA, a recent large-scale far-field speech corpus, with varied microphone and positional settings. Our methods outperform conventional mel filterbank features, with maximum of 33.5% and 39.5% relative improvement on equal error rate under matched microphone and mismatched microphone conditions, respectively. Xuechen Liu 0001, Md. Sahidullah, Tomi Kinnunen |
ASRU | 3 |
| 2021 | Data Quality as Predictor of Voice Anti-Spoofing GeneralizationabstractVoice anti-spoofing aims at classifying a given utterance either as a bonafide human sample, or a spoofing attack (e.g. synthetic or replayed sample). Many anti-spoofing methods have been proposed but most of them fail to generalize across domains (corpora) -- and we do not know \emph{why}. We outline a novel interpretative framework for gauging the impact of data quality upon anti-spoofing performance. Our within- and between-domain experiments pool data from seven public corpora and three anti-spoofing methods based on Gaussian mixture and convolutive neural network models. We assess the impacts of long-term spectral information, speaker population (through x-vector speaker embeddings), signal-to-noise ratio, and selected voice quality features. Bhusan Chettri, Rosa González Hautamäki, Md. Sahidullah, Tomi Kinnunen |
Interspeech | 4 |
| 2021 | Visualizing Classifier Adjacency Relations: A Case Study in Speaker Verification and Voice Anti-SpoofingabstractWhether it be for results summarization, or the analysis of classifier fusion, some means to compare different classifiers can often provide illuminating insight into their behaviour, (dis)similarity or complementarity. We propose a simple method to derive 2D representation from detection scores produced by an arbitrary set of binary classifiers in response to a common dataset. Based upon rank correlations, our method facilitates a visual comparison of classifiers with arbitrary scores and with close relation to receiver operating characteristic (ROC) and detection error trade-off (DET) analyses. While the approach is fully versatile and can be applied to any detection task, we demonstrate the method using scores produced by automatic speaker verification and voice anti-spoofing systems. The former are produced by a Gaussian mixture model system trained with VoxCeleb data whereas the latter stem from submissions to the ASVspoof 2019 challenge. Tomi Kinnunen, Andreas Nautsch, Md. Sahidullah, Nicholas W. D. Evans, Xin Wang 0037, Massimiliano Todisco, Héctor Delgado, Junichi Yamagishi, Kong-Aik Lee |
Interspeech | 1 |
| 2021 | Learnable MFCCs for Speaker VerificationabstractWe propose a learnable mel-frequency cepstral coefficients (MFCCs) front-end architecture for deep neural network (DNN) based automatic speaker verification. Our architecture retains the simplicity and interpretability of MFCC-based features while allowing the model to be adapted to data flexibly. In practice, we formulate data-driven version of four linear transforms in a standard MFCC extractor - windowing, discrete Fourier transform (DFT), mel filterbank and discrete cosine transform (DCT). Results reported reach up to 6.7% (VoxCeleb1) and 9.7% (SITW) relative improvement in term of equal error rate (EER) from static MFCCs, without additional tuning effort. Xuechen Liu 0001, Md. Sahidullah, Tomi Kinnunen |
ISCAS | 3 |
| 2021 | UIAI System for Short-Duration Speaker Verification Challenge 2020abstractIn this work, we present the system description of the UIAI entry for the short-duration speaker verification (SdSV) challenge 2020. Our focus is on Task 1 dedicated to text-dependent speaker verification. We investigate different feature extraction and modeling approaches for automatic speaker verification (ASV) and utterance verification (UV). We have also studied different fusion strategies for combining UV and ASV modules. Our primary submission to the challenge is the fusion of seven subsystems which yields a normalized minimum detection cost function (minDCF) of 0.072 and an equal error rate (EER) of 2.14% on the evaluation set. The single system consisting of a pass-phrase identification based model with phone-discriminative bottleneck features gives a normalized minDCF of 0.118 and achieves 19% relative improvement over the state-of-the-art challenge baseline. Md. Sahidullah, Achintya Kumar Sarkar, Ville Vestman, Xuechen Liu 0001, Romain Serizel, Tomi Kinnunen, Zheng-Hua Tan, Emmanuel Vincent 0001 |
SLT | 6 |
| 2021 | Optimizing Multi-Taper Features for Deep Speaker VerificationabstractMulti-taper estimators provide low-variance power spectrum estimates that can be used in place of the windowed discrete Fourier transform (DFT) to extract speech features such as mel-frequency cepstral coefficients (MFCCs). Even if past work has reported promising automatic speaker verification (ASV) results with Gaussian mixture model-based classifiers, the performance of multi-taper MFCCs with deep ASV systems remains an open question. Instead of a static-taper design, we propose to optimize the multi-taper estimator jointly with a deep neural network trained for ASV tasks. With a maximum improvement on the SITW corpus of 25.8% in terms of equal error rate over the static-taper, our method helps preserve a balanced level of leakage and variance, providing more robustness. Xuechen Liu 0001, Md. Sahidullah, Tomi Kinnunen |
IEEE Signal Process. Lett. | 3 |
| 2020 | The Attacker's Perspective on Automatic Speaker Verification: An OverviewabstractSecurity of automatic speaker verification (ASV) systems is compromised by various spoofing attacks.While many types of non-proactive attacks (and their defenses) have been studied in the past, attacker's perspective on ASV, represents a far less explored direction.It can potentially help to identify the weakest parts of ASV systems and be used to develop attackeraware systems.We present an overview on this emerging research area by focusing on potential threats of adversarial attacks on ASV, spoofing countermeasures, or both.We conclude the study with discussion on selected attacks and leveraging from such knowledge to improve defense mechanisms against adversarial attacks. Rohan Kumar Das, Xiaohai Tian, Tomi Kinnunen, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2020 | Why Did the x-Vector System Miss a Target Speaker? Impact of Acoustic Mismatch Upon Target Score on VoxCeleb DataabstractModern automatic speaker verification (ASV) relies heavily on machine learning implemented through deep neural networks. It can be difficult to interpret the output of these black boxes. In line with interpretative machine learning, we model the dependency of ASV detection score upon acoustic mismatch of the enrollment and test utterances. We aim to identify mismatch factors that explain target speaker misses (false rejections). We use distance in the first- and second-order statistics of selected acoustic features as the predictors in a linear mixed effects model, while a standard Kaldi x-vector system forms our ASV black-box. Our results on the VoxCeleb data reveal the most prominent mismatch factor to be in F0 mean, followed by mismatches associated with formant frequencies. Our findings indicate that x-vector systems lack robustness to intra-speaker variations. Rosa González Hautamäki, Tomi Kinnunen |
INTERSPEECH | 2 |
| 2020 | A Comparative Re-Assessment of Feature Extractors for Deep Speaker EmbeddingsabstractModern automatic speaker verification relies largely on deep neural networks (DNNs) trained on mel-frequency cepstral coefficient (MFCC) features.While there are alternative feature extraction methods based on phase, prosody and long-term temporal operations, they have not been extensively studied with DNN-based methods.We aim to fill this gap by providing extensive re-assessment of 14 feature extractors on VoxCeleb and SITW datasets.Our findings reveal that features equipped with techniques such as spectral centroids, group delay function, and integrated noise suppression provide promising alternatives to MFCCs for deep speaker embeddings extraction.Experimental results demonstrate up to 16.3% (VoxCeleb) and 25.1% (SITW) relative decrease in equal error rate (EER) to the baseline. Xuechen Liu 0001, Md. Sahidullah, Tomi Kinnunen |
INTERSPEECH | 3 |
| 2020 | Extrapolating False Alarm Rates in Automatic Speaker VerificationabstractAutomatic speaker verification (ASV) vendors and corpus providers would both benefit from tools to reliably extrapolate performance metrics for large speaker populations without collecting new speakers. We address false alarm rate extrapolation under a worst-case model whereby an adversary identifies the closest impostor for a given target speaker from a large population. Our models are generative and allow sampling new speakers. The models are formulated in the ASV detection score space to facilitate analysis of arbitrary ASV systems. Alexey Sholokhov, Tomi Kinnunen, Ville Vestman, Kong-Aik Lee |
INTERSPEECH | 2 |
| 2020 | Introduction to the special issue "Speaker and language characterization and recognition: Voice modeling, conversion, synthesis and ethical aspects"
Jean-François Bonastre, Tomi Kinnunen, Anthony Larcher, Junichi Yamagishi |
Comput. Speech Lang. | 2 |
| 2020 | Deep generative variational autoencoding for replay spoof detection in automatic speaker verification
Bhusan Chettri, Tomi Kinnunen, Emmanouil Benetos |
Comput. Speech Lang. | 2 |
| 2020 | Voice biometrics security: Extrapolating false alarm rate via hierarchical Bayesian modeling of speaker verification scores
Alexey Sholokhov, Tomi Kinnunen, Ville Vestman, Kong-Aik Lee |
Comput. Speech Lang. | 2 |
| 2020 | Voice Mimicry Attacks Assisted by Automatic Speaker Verification
Ville Vestman, Tomi Kinnunen, Rosa González Hautamäki, Md. Sahidullah |
Comput. Speech Lang. | 2 |
| 2020 | ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech
Xin Wang 0037, Junichi Yamagishi, Massimiliano Todisco, Héctor Delgado, Andreas Nautsch, Nicholas W. D. Evans, Md. Sahidullah, Ville Vestman, Tomi Kinnunen, Kong-Aik Lee, Lauri Juvela, Paavo Alku, Yu-Huai Peng, Hsin-Te Hwang, Yu Tsao 0001, Hsin-Min Wang, Sébastien Le Maguer, Zhen-Hua Ling |
Comput. Speech Lang. | 9 |
| 2020 | Tandem Assessment of Spoofing Countermeasures and Automatic Speaker Verification: FundamentalsabstractRecent years have seen growing efforts to develop spoofing countermeasures (CMs) to protect automatic speaker verification (ASV) systems from being deceived by manipulated or artificial inputs. The reliability of spoofing CMs is typically gauged using the equal error rate (EER) metric. The primitive EER fails to reflect application requirements and the impact of spoofing and CMs upon ASV and its use as a primary metric in traditional ASV research has long been abandoned in favour of risk-based approaches to assessment. This paper presents several new extensions to the tandem detection cost function (t-DCF), a recent risk-based approach to assess the reliability of spoofing CMs deployed in tandem with an ASV system. Extensions include a simplified version of the t-DCF with fewer parameters, an analysis of a special case for a fixed ASV system, simulations which give original insights into its interpretation and new analyses using the ASVspoof 2019 database. It is hoped that adoption of the t-DCF for the CM assessment will help to foster closer collaboration between the anti-spoofing and ASV research communities. Tomi Kinnunen, Héctor Delgado, Nicholas W. D. Evans, Kong-Aik Lee, Ville Vestman, Andreas Nautsch, Massimiliano Todisco, Xin Wang 0037, Md. Sahidullah, Junichi Yamagishi, Douglas A. Reynolds |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Towards Controlling False Alarm - Miss Trade-Off in Perceptual Speaker Comparison via Non-Neutral Listening Task FramingabstractSpeaker comparison by listening is a valuable resource, for instance, in human voice discrimination studies, and voice conversion (VC) systems evaluations. Usually, listeners are provided with application-neutral guidelines that encourage retaining overall high speaker discrimination accuracy. Nonetheless, listeners are subject to misses (declaring same-speaker trial as different-speaker) and false alarms (vice versa) with possibly non-symmetric outcomes. In automatic speaker verification (ASV) applications, the consequences of a miss and a false alarm are rarely equal, and decision making policy is adjusted towards a given application with a desired miss/false alarm trade-off. We study whether listener decisions could similarly be controlled to provoke more accept (or reject) decisions, by framing the voice comparison task in different ways. Our neutral, forensic, user-convenient bank and secure bank scenarios are played by disjoint panels (through Amazon's Mechanical Turk), all judging the same speaker trials originated from RedDots and 2018 Voice Conversion Challenge (VCC 2018) data. Our results indicate that listener decisions can be influenced by modifying the task framing. As a subjective task, the challenge is how to drive the panel decisions to the desired direction (to reduce miss or false alarm rate). Our preliminary results suggest potential for novel, application-directed speaker discrimination designs. Rosa González Hautamäki, Tomi Kinnunen |
ASRU | 2 |
| 2019 | Can We Use Speaker Recognition Technology to Attack Itself? Enhancing Mimicry Attacks Using Automatic Target Speaker SelectionabstractWe consider technology-assisted mimicry attacks in the context of automatic speaker verification (ASV). We use ASV itself to select targeted speakers to be attacked by human-based mimicry. We recorded 6 naive mimics for whom we select target celebrities from VoxCeleb1 and VoxCeleb2 corpora (7,365 potential targets) using an i-vector system. The attacker attempts to mimic the selected target, with the utterances subjected to ASV tests using an independently developed x-vector system. Our main finding is negative: even if some of the attacker scores against the target speakers were slightly increased, our mimics did not succeed in spoofing the x-vector system. Interestingly, however, the relative ordering of the selected targets (closest, furthest, median) are consistent between the systems, which suggests some level of transferability between the systems. Tomi Kinnunen, Rosa González Hautamäki, Ville Vestman, Md. Sahidullah |
ICASSP | 1 |
| 2019 | Who Do I Sound like? Showcasing Speaker Recognition Technology by Youtube Voice SearchabstractThe popularization of science can often be disregarded by scientists as it may be challenging to put highly sophisticated research into words that general public can understand. This work aims to help presenting speaker recognition research to public by proposing a publicly appealing concept for showcasing recognition systems. We leverage data from YouTube and use it in a large-scale voice search web application that finds the celebrity voices that best match to the user's voice. The concept was tested in a public event as well as "in the wild" and the received feedback was mostly positive. The i-vector based speaker identification back end was found to be fast (665 ms per request) and had a high identification accuracy (93%) for the YouTube target speakers. To help other researchers to develop the idea further, we share the source codes of the web platform used for the demo at https://github.com/bilalsoomro/speech-demo-platform. Ville Vestman, Bilal Soomro, Anssi Kanervisto, Ville Hautamäki, Tomi Kinnunen |
ICASSP | 5 |
| 2019 | I4U Submission to NIST SRE 2018: Leveraging from a Decade of Shared ExperiencesabstractThe I4U consortium was established to facilitate a joint entry to NIST speaker recognition evaluations (SRE). The latest edition of such joint submission was in SRE 2018, in which the I4U submission was among the best-performing systems. SRE'18 also marks the 10-year anniversary of I4U consortium into NIST SRE series of evaluation. The primary objective of the current paper is to summarize the results and lessons learned based on the twelve sub-systems and their fusion submitted to SRE'18. It is also our intention to present a shared view on the advancements, progresses, and major paradigm shifts that we have witnessed as an SRE participant in the past decade from SRE'08 to SRE'18. In this regard, we have seen, among others, a paradigm shift from supervector representation to deep speaker embedding, and a switch of research challenge from channel compensation to domain adaptation. Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Hitoshi Yamamoto, Koji Okabe, Ville Vestman, Jing Huang 0019, Guo-Hong Ding, Hanwu Sun, Anthony Larcher, Rohan Kumar Das, Haizhou Li 0001, Mickael Rouvier, Pierre-Michel Bousquet, Wei Rao 0002, Qing Wang 0039, Fahimeh Bahmaninezhad, Héctor Delgado, Massimiliano Todisco |
INTERSPEECH | 3 |
| 2019 | ASVspoof 2019: Future Horizons in Spoofed and Fake Audio DetectionabstractASVspoof, now in its third edition, is a series of community-led challenges which promote the development of countermeasures to protect automatic speaker verification (ASV) from the threat of spoofing. Advances in the 2019 edition include: (i) a consideration of both logical access (LA) and physical access (PA) scenarios and the three major forms of spoofing attack, namely synthetic, converted and replayed speech; (ii) spoofing attacks generated with state-of-the-art neural acoustic and waveform models; (iii) an improved, controlled simulation of replay attacks; (iv) use of the tandem detection cost function (t-DCF) that reflects the impact of both spoofing and countermeasures upon ASV reliability. Even if ASV remains the core focus, in retaining the equal error rate (EER) as a secondary metric, ASVspoof also embraces the growing importance of fake audio detection. ASVspoof 2019 attracted the participation of 63 research teams, with more than half of these reporting systems that improve upon the performance of two baseline spoofing countermeasures. This paper describes the 2019 database, protocols and challenge results. It also outlines major findings which demonstrate the real progress made in protecting against the threat of spoofing and fake audio. Massimiliano Todisco, Xin Wang 0037, Ville Vestman, Md. Sahidullah, Héctor Delgado, Andreas Nautsch, Junichi Yamagishi, Nicholas W. D. Evans, Tomi Kinnunen, Kong-Aik Lee |
INTERSPEECH | 9 |
| 2019 | Unleashing the Unused Potential of i-Vectors Enabled by GPU AccelerationabstractSpeaker embeddings are continuous-value vector representations that allow easy comparison between voices of speakers with simple geometric operations. Among others, i-vector and x-vector have emerged as the mainstream methods for speaker embedding. In this paper, we illustrate the use of modern computation platform to harness the benefit of GPU acceleration for i-vector extraction. In particular, we achieve an acceleration of 3000 times in frame posterior computation compared to real time and 25 times in training the i-vector extractor compared to the CPU baseline from Kaldi toolkit. This significant speed-up allows the exploration of ideas that were hitherto impossible. In particular, we show that it is beneficial to update the universal background model (UBM) and re-compute frame alignments while training the i-vector extractor. Additionally, we are able to study different variations of i-vector extractors more rigorously than before. In this process, we reveal some undocumented details of Kaldi's i-vector extractor and show that it outperforms the standard formulation by a margin of 1 to 2% when tested with VoxCeleb speaker verification protocol. All of our findings are asserted by ensemble averaging the results from multiple runs with random start. Ville Vestman, Kong-Aik Lee, Tomi Kinnunen, Takafumi Koshinaka |
INTERSPEECH | 3 |
| 2019 | Vocal effort compensation for MFCC feature extraction in a shouted versus normal speaker recognition task
Emma Jokinen, Rahim Saeidi, Tomi Kinnunen, Paavo Alku |
Comput. Speech Lang. | 3 |
| 2019 | Statistical Regression Models for Noise Robust F0 Estimation Using Recurrent Deep Neural NetworksabstractThe fundamental frequency (F0) in a speech signal, which corresponds to pitch, is one of the key features involved in a variety of speech processing tasks. Therefore, accurate F0 estimation has remained an important problem to be solved over decades. However, this problem is difficult, especially in low signal-to-noise ratio (SNR) conditions with unknown noise. In this work, we propose new approaches to noise-robust F0 estimation using recurrent neural networks (RNNs). Recent F0 estimation studies exploit deep neural networks (DNNs), including RNNs, to classify acoustic features into quantized frequency states. In contrast to these classification approaches, we put forward a regression method for F0 tracking, which is accomplished with RNNs. To this end, we propose two variants. Our first model predicts the (scalar) F0 value directly from a spectrum, while our second model predicts a target sinusoidal waveform (with the desired F0) from the raw speech waveform. Our experiments with the pitch tracking database from Graz University of Technology (PTDB-TUG), contaminated by additive noise (NOISEX-92), demonstrate the improvement of the proposed approaches in terms of the gross pitch error (GPE) and fine pitch error (FPE) rates by more than 35% at SNRs between -10 dB and +10 dB against a well-known, noise-robust F0 tracker, PEFAC. Furthermore, our methods outperform state-of-the-art neural network-based approaches by more than 15% in terms of both the FPE and GPE rates over the abovementioned SNR range. Akihiro Kato, Tomi Kinnunen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Waveform to Single Sinusoid Regression to Estimate the F0 Contour from Noisy Speech Using Recurrent Deep Neural NetworksabstractThe fundamental frequency (F 0) represents pitch in speech that determines prosodic characteristics of speech and is needed in various tasks for speech analysis and synthesis.Despite decades of research on this topic, F 0 estimation at low signal-to-noise ratios (SNRs) in unexpected noise conditions remains difficult.This work proposes a new approach to noise robust F 0 estimation using a recurrent neural network (RNN) trained in a supervised manner.Recent studies employ deep neural networks (DNNs) for F 0 tracking as a frame-by-frame classification task into quantised frequency states but we propose waveform-tosinusoid regression instead to achieve both noise robustness and accurate estimation with increased frequency resolution.Experimental results with PTDB-TUG corpus contaminated by additive noise (NOISEX-92) demonstrate that the proposed method improves gross pitch error (GPE) rate and fine pitch error (FPE) by more than 35 % at SNRs between -10 dB and +10 dB compared with well-known noise robust F 0 tracker, PEFAC.Furthermore, the proposed method also outperforms state-of-the-art DNN-based approaches by more than 15 % in terms of both FPE and GPE rate over the preceding SNR range. Akihiro Kato, Tomi Kinnunen |
INTERSPEECH | 2 |
| 2018 | Integrated Presentation Attack Detection and Automatic Speaker Verification: Common Features and Gaussian Back-end FusionabstractInternational audience Massimiliano Todisco, Héctor Delgado, Kong-Aik Lee, Md. Sahidullah, Nicholas W. D. Evans, Tomi Kinnunen, Junichi Yamagishi |
INTERSPEECH | 6 |
| 2018 | Semi-supervised speech activity detection with an application to automatic speaker verification
Alexey Sholokhov, Md. Sahidullah, Tomi Kinnunen |
Comput. Speech Lang. | 3 |
| 2018 | Speaker recognition from whispered speech: A tutorial survey and an application of time-varying linear prediction
Ville Vestman, Dhananjaya Gowda, Md. Sahidullah, Paavo Alku, Tomi Kinnunen |
Speech Commun. | 5 |
| 2018 | Robust Voice Liveness Detection and Speaker Verification Using Throat MicrophonesabstractWhile having a wide range of applications, automatic speaker verification (ASV) systems are vulnerable to spoofing attacks, in particular, replay attacks that are effective and easy to implement. Most prior work on detecting replay attacks uses audio from a single acoustic microphone only, leading to difficulties in detecting high-end replay attacks close to indistinguishable from live human speech. In this paper, we study the use of a special body-conducted sensor, throat microphone (TM), for combined voice liveness detection (VLD) and ASV in order to improve both robustness and security of ASV against replay attacks. We first investigate the possibility and methods of attacking a TM-based ASV system, followed by a pilot data collection. Second, we study the use of spectral features for VLD using both single-channel and dual-channel ASV systems. We carry out speaker verification experiments using Gaussian mixture model with universal background model (GMM-UBM) and i-vector based systems on a dataset of 38 speakers collected by us. We have achieved considerable improvement in recognition accuracy, with the use of dual-microphone setup. In experiments with noisy test speech, the false acceptance rate (FAR) of the dual-microphone GMM-UBM based system for recorded speech reduces from 69.69% to 18.75%. The FAR of replay condition further drops to 0% when this dual-channel ASV system is integrated with the new dual-channel voice liveness detector. Md. Sahidullah, Dennis Alexander Lehmann Thomsen, Rosa González Hautamäki, Tomi Kinnunen, Zheng-Hua Tan, Robert Parts, Martti Pitkänen |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2017 | Effects of gender information in text-independent and text-dependent speaker verificationabstractIt is well-known that for speaker recognition task, gender-dependent acoustic modeling performs better than gender-independent modeling. The practice is to use the gender ground-truth and to train gender-dependent models. However, such information is not necessarily available, especially if speakers are remotely enrolled. A way to overcome this is to use a gender classification system, which introduces an additional layer of uncertainty. To date, such uncertainty has not been studied. We implement two gender classifier systems and test them with two different corpora and speaker verification systems. We find that estimated gender information can improve speaker verification accuracy over gender-independent methods. Our detailed analysis suggests that gender estimation should have a sufficiently high accuracy to yield improvements in speaker verification performance. Anssi Kanervisto, Ville Vestman, Md. Sahidullah, Ville Hautamäki, Tomi Kinnunen |
ICASSP | 5 |
| 2017 | Non-parallel voice conversion using i-vector PLDA: towards unifying speaker verification and transformationabstractText-independent speaker verification (recognizing speakers regardless of content) and non-parallel voice conversion (transforming voice identities without requiring content-matched training utterances) are related problems. We adopt i-vector method to voice conversion. An i-vector is a fixed-dimensional representation of a speech utterance that enables treating voice conversion in utterance domain, as opposed to frame domain. The high dimensionality (800) and small number of training utterances (24) necessitates using prior information of speakers. We adopt probabilistic linear discriminant analysis (PLDA) for voice conversion. The proposed approach requires neither parallel utterances, transcriptions nor time alignment procedures at any stage. Tomi Kinnunen, Lauri Juvela, Paavo Alku, Junichi Yamagishi |
ICASSP | 1 |
| 2017 | RedDots replayed: A new replay spoofing attack corpus for text-dependent speaker verification researchabstractThis paper describes a new database for the assessment of automatic speaker verification (ASV) vulnerabilities to spoofing attacks. In contrast to other recent data collection efforts, the new database has been designed to support the development of replay spoofing countermeasures tailored towards the protection of text-dependent ASV systems from replay attacks in the face of variable recording and playback conditions. Derived from the re-recording of the original RedDots database, the effort is aligned with that in text-dependent ASV and thus well positioned for future assessments of replay spoofing countermeasures, not just in isolation, but in integration with ASV. The paper describes the database design and re-recording, a protocol and some early spoofing detection results. The new “RedDots Replayed” database is publicly available through a creative commons license. Tomi Kinnunen, Md. Sahidullah, Mauro Falcone, Luca Costantini, Rosa González Hautamäki, Dennis Alexander Lehmann Thomsen, Achintya Kumar Sarkar, Zheng-Hua Tan, Héctor Delgado, Massimiliano Todisco, Nicholas W. D. Evans, Ville Hautamäki, Kong-Aik Lee |
ICASSP | 1 |
| 2017 | The ASVspoof 2017 Challenge: Assessing the Limits of Replay Spoofing Attack DetectionabstractThe ASVspoof initiative was created to promote the development of countermeasures which aim to protect automatic speaker verification (ASV) from spoofing attacks. The first community-led, common evaluation held in 2015 focused on countermeasures for speech synthesis and voice conversion spoofing attacks. Arguably, however, it is replay attacks which pose the greatest threat. Such attacks involve the replay of recordings collected from enrolled speakers in order to provoke false alarms and can be mounted with greater ease using everyday consumer devices. ASVspoof 2017, the second in the series, hence focused on the development of replay attack countermeasures. This paper describes the database, protocols and initial findings. The evaluation entailed highly heterogeneous acoustic recording and replay conditions which increased the equal error rate (EER) of a baseline ASV system from 1.76% to 30.71%. Submissions were received from 49 research teams, 20 of which improved upon a baseline replay spoofing detector EER of 24.65%, in terms of replay/non-replay discrimination. While largely successful, the evaluation indicates that the quest for countermeasures which are resilient in the face of variable replay attacks remains very much alive. Tomi Kinnunen, Md. Sahidullah, Héctor Delgado, Massimiliano Todisco, Nicholas W. D. Evans, Junichi Yamagishi, Kong-Aik Lee |
INTERSPEECH | 1 |
| 2017 | The I4U Mega Fusion and Collaboration for NIST Speaker Recognition Evaluation 2016abstract18th Annual Conference of the International Speech Communication Association, INTERSPEECH 2017, Stockholm, Sweden, 20-24 August 2017 Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Anthony Larcher, Andreas Nautsch, Themos Stafylakis, Gang Liu 0001, Mickael Rouvier, Wei Rao 0002, Federico Alegre, Man-Wai Mak, Achintya Kumar Sarkar, Héctor Delgado, Rahim Saeidi, Hagai Aronowitz, Aleksandr Sizov, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Bin Ma 0001, Ville Vestman, Md. Sahidullah, M. Halonen, Anssi Kanervisto, Gaël Le Lan, Fahimeh Bahmaninezhad, Sergey Isadskiy, Christian Rathgeb, Christoph Busch 0001, Georgios Tzimiropoulos, Q. Qian, Q. Zhao, J. Xue, R. Jin, T. Zhao, Pierre-Michel Bousquet, Moez Ajili, Waad Ben Kheder, Driss Matrouf, Zhi Hao Lim, Chenglin Xu, Haihua Xu 0001, Chng Eng Siong, Benoit G. B. Fauve, Kaavya Sriskandaraja, Vidhyasaharan Sethu, W. W. Lin, Dennis Alexander Lehmann Thomsen, Zheng-Hua Tan, Massimiliano Todisco, Nicholas W. D. Evans, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Eliathamby Ambikairajah |
INTERSPEECH | 3 |
| 2017 | Improving Speaker Verification Performance in Presence of Spoofing Attacks Using Out-of-Domain Spoofed DataabstractAutomatic speaker verification (ASV) systems are vulnerable to spoofing attacks using speech generated by voice conversion and speech synthesis techniques. Commonly, a countermeasure (CM) system is integrated with an ASV system for improved protection against spoofing attacks. But integration of the two systems is challenging and often leads to increased false rejection rates. Furthermore, the performance of CM severely degrades if in-domain development data are unavailable. In this study, therefore, we propose a solution that uses two separate background models – one from human speech and another from spoofed data. During test, the ASV score for an input utterance is computed as the difference of the log-likelihood against the target model and the combination of the log-likelihoods against two background models. Evaluation experiments are conducted using the joint ASV and CM protocol of ASVspoof 2015 corpus consisting of text-independent ASV tasks with short utterances. Our proposed system reduces error rates in the presence of spoofing attacks by using out-of-domain spoofed data for system development, while maintaining the performance for zero-effort imposter attacks compared to the baseline system. Achintya Kumar Sarkar, Md. Sahidullah, Zheng-Hua Tan, Tomi Kinnunen |
INTERSPEECH | 4 |
| 2017 | Time-Varying Autoregressions for Speaker Verification in Reverberant ConditionsabstractAutomatic speaker verification (ASV) systems are vulnerable to spoofing attacks using speech generated by voice conversion and speech synthesis techniques. Commonly, a countermeasure (CM) system is integrated with an ASV system for improved protection against spoofing attacks. But integration of the two systems is challenging and often leads to increased false rejection rates. Furthermore, the performance of CM severely degrades if in-domain development data are unavailable. In this study, therefore, we propose a solution that uses two separate background models — one from human speech and another from spoofed data. During test, the ASV score for an input utterance is computed as the difference of the log-likelihood against the target model and the combination of the log-likelihoods against two background models. Evaluation experiments are conducted using the joint ASV and CM protocol of ASVspoof 2015 corpus consisting of text-independent ASV tasks with short utterances. Our proposed system reduces error rates in the presence of spoofing attacks by using out-of-domain spoofed data for system development, while maintaining the performance for zero-effort imposter attacks compared to the baseline system. Ville Vestman, Dhananjaya Gowda, Md. Sahidullah, Paavo Alku, Tomi Kinnunen |
INTERSPEECH | 5 |
| 2017 | Acoustical and perceptual study of voice disguise by age modification in speaker verification
Rosa González Hautamäki, Md. Sahidullah, Ville Hautamäki, Tomi Kinnunen |
Speech Commun. | 4 |
| 2017 | Direct Optimization of the Detection Cost for I-Vector-Based Spoken Language RecognitionabstractWe explore a method to boost discriminative capabilities of probabilistic linear discriminant analysis (PLDA) model without losing its generative advantages. We show a sequential projection and training steps leading to a classifier that operates in the original i-vector space but is discriminatively trained in a low-dimensional PLDA latent subspace. We use extended Baum-Welch technique to optimize the model with respect to two objective functions for discriminative training. One of them is the well-known maximum mutual information objective, while the other one is a new objective that we propose to approximate the language detection cost. We evaluate the performance on NIST language recognition evaluation (LRE) 2015 and our development dataset comprised of the utterances from previous LREs. We improve the detection cost by 10% and 6% relative compared to our fine-tuned generative and discriminative baselines, and by 10% over the best of our previously reported results. The proposed approximation method of the cost function and PLDA subspace training are applicable for a broad range of tasks. Aleksandr Sizov, Kong-Aik Lee, Tomi Kinnunen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Discriminative multi-domain PLDA for speaker verificationabstractDomain mismatch occurs when data from application-specific target domain is related to, but cannot be viewed as iid samples from the source domain used for training speaker models. Another problem occurs when several training datasets are available but their domains differ. In this case training on simply merged subsets can lead to suboptimal performance. Existing approaches to cope with these problems employ generative modeling and consist of several separate stages such as training and adaptation. In this work we explore a discriminative approach which naturally incorporates both scenarios in a principled way. To this end, we develop a method that can learn across multiple domains by extending discriminative probabilistic linear discriminant analysis (PLDA) according to multi-task learning paradigm. Our results on the recent JHU Domain Adaptation Challenge (DAC) dataset demonstrate that the proposed multi-task PLDA decreases equal error rate (EER) of the PLDA without domain compensation by more than 35% relative and performs comparable to another competitive domain compensation technique. Alexey Sholokhov, Tomi Kinnunen, Sandro Cumani |
ICASSP | 2 |
| 2016 | Utterance Verification for Text-Dependent Speaker Recognition: A Comparative Assessment Using the RedDots CorpusabstractText-dependent automatic speaker verification naturally calls for the simultaneous verification of speaker identity and spoken content. These two tasks can be achieved with automatic speaker verification (ASV) and utterance verification (UV) technologies. While both have been addressed previously in the literature, a treatment of simultaneous speaker and utterance verification with a modern, standard database is so far lacking. This is despite the burgeoning demand for voice biometrics in a plethora of practical security applications. With the goal of improving overall verification performance, this paper reports different strategies for simultaneous ASV and UV in the context of short-duration, text-dependent speaker verification. Experiments performed on the recently released RedDots corpus are reported for three different ASV systems and four different UV systems. Results show that the combination of utterance verification with automatic speaker verification is (almost) universally beneficial with significant performance improvements being observed. Tomi Kinnunen, Md. Sahidullah, Ivan Kukanov, Héctor Delgado, Massimiliano Todisco, Achintya Kumar Sarkar, Nicolai Bæk Thomsen, Ville Hautamäki, Nicholas W. D. Evans, Zheng-Hua Tan |
INTERSPEECH | 1 |
| 2016 | HAPPY Team Entry to NIST OpenSAD Challenge: A Fusion of Short-Term Unsupervised and Segment i-Vector Based Speech Activity DetectorsabstractSpeech activity detection (SAD), the task of locating speech segments from a given recording, remains challenging under acoustically degraded conditions. In 2015, National Institute of Standards and Technology (NIST) coordinated OpenSAD bench-mark. We summarize “HAPPY” team effort to OpenSAD. SADs come in both unsupervised and supervised flavors, the latter requiring a labeled training set. Our solution fuses six base SADs (2 supervised and 4 unsupervised). The individually best SAD, in terms of detection cost function (DCF), is supervised and uses adaptive segmentation with i-vectors to represent the segments. Fusion of the six base SADs yields a relative decrease of 9.3% in DCF over this SAD. Further, relative decrease of 17.4% is obtained by incorporating channel detection side information. Tomi Kinnunen, Alexey Sholokhov, Elie Khoury 0001, Dennis Alexander Lehmann Thomsen, Md. Sahidullah, Zheng-Hua Tan |
INTERSPEECH | 1 |
| 2016 | Integrated Spoofing Countermeasures and Automatic Speaker Verification: An Evaluation on ASVspoof 2015abstractIt is well known that automatic speaker verification (ASV) systems can be vulnerable to spoofing. The community has responded to the threat by developing dedicated countermeasures aimed at detecting spoofing attacks. Progress in this area has accelerated over recent years, partly as a result of the first standard evaluation, ASVspoof 2015, which focused on spoofing detection in isolation from ASV. This paper investigates the integration of state-of-the-art spoofing countermeasures in combination with ASV. Two general strategies to countermeasure integration are reported: cascaded and parallel. The paper reports the first comparative evaluation of each approach performed with the ASVspoof 2015 corpus. Results indicate that, even in the case of varying spoofing attack algorithms, ASV performance remains robust when protected with a diverse set of integrated countermeasures. Md. Sahidullah, Héctor Delgado, Massimiliano Todisco, Hong Yu 0015, Tomi Kinnunen, Nicholas W. D. Evans, Zheng-Hua Tan |
INTERSPEECH | 5 |
| 2016 | Robust Speaker Recognition with Combined Use of Acoustic and Throat Microphone SpeechabstractAccuracy of automatic speaker recognition (ASV) systems degrades severely in the presence of background noise. In this paper, we study the use of additional side information provided by a body-conducted sensor, throat microphone. Throat microphone signal is much less affected by background noise in comparison to acoustic microphone signal. This makes throat microphones potentially useful for feature extraction or speech activity detection. This paper, firstly, proposes a new prototype system for simultaneous data-acquisition of acoustic and throat microphone signals. Secondly, we study the use of this additional information for both speech activity detection, feature extraction and fusion of the acoustic and throat microphone signals. We collect a pilot database consisting of 38 subjects including both clean and noisy sessions. We carry out speaker verification experiments using Gaussian mixture model with universal background model (GMM-UBM) and i-vector based system. We have achieved considerable improvement in recognition accuracy even in highly degraded conditions. Md. Sahidullah, Rosa González Hautamäki, Dennis Alexander Lehmann Thomsen, Tomi Kinnunen, Zheng-Hua Tan, Ville Hautamäki, Robert Parts, Martti Pitkänen |
INTERSPEECH | 4 |
| 2016 | Further optimisations of constant Q cepstral processing for integrated utterance and text-dependent speaker verificationabstractMany authentication applications involving automatic speaker verification (ASV) demand robust performance using short-duration, fixed or prompted text utterances. Text constraints not only reduce the phone-mismatch between enrolment and test utterances, which generally leads to improved performance, but also provide an ancillary level of security. This can take the form of explicit utterance verification (UV). An integrated UV + ASV system should then verify access attempts which contain not just the expected speaker, but also the expected text content. This paper presents such a system and introduces new features which are used for both UV and ASV tasks. Based upon multi-resolution, spectro-temporal analysis and when fused with more traditional parameterisations, the new features not only generally outperform Mel-frequency cepstral coefficients, but also are shown to be complementary when fusing systems at score level. Finally, the joint operation of UV and ASV greatly decreases false acceptances for unmatched text trials. Héctor Delgado, Massimiliano Todisco, Md. Sahidullah, Achintya Kumar Sarkar, Nicholas W. D. Evans, Tomi Kinnunen, Zheng-Hua Tan |
SLT | 6 |
| 2016 | Spoofing detection goes noisy: An analysis of synthetic speech detection in the presence of additive noise
Cemal Hanilçi, Tomi Kinnunen, Md. Sahidullah, Aleksandr Sizov |
Speech Commun. | 2 |
| 2016 | i-Vector Modeling of Speech Attributes for Automatic Foreign Accent RecognitionabstractWe propose a unified approach to automatic foreign accent recognition. It takes advantage of recent technology advances in both linguistics and acoustics based modeling techniques in automatic speech recognition (ASR) while overcoming the issue of a lack of a large set of transcribed data often required in designing state-of-the-art ASR systems. The key idea lies in defining a common set of fundamental units “universally” across all spoken accents such that any given spoken utterance can be transcribed with this set of “accent-universal” units. In this study, we adopt a set of units describing manner and place of articulation as speech attributes. These units exist in most spoken languages and they can be reliably modeled and extracted to represent foreign accent cues. We also propose an i-vector representation strategy to model the feature streams formed by concatenating these units. Testing on both the Finnish national foreign language certificate (FSD) corpus and the English NIST 2008 SRE corpus, the experimental results with the proposed approach demonstrate a significant system performance improvement with p-value over those with the conventional spectrum-based techniques. We observed up to a 15% relative error reduction over the already very strong i-vector accented recognition system when only manner information is used. Additional improvement is obtained by adding place of articulation clues along with context information. Furthermore, diagnostic information provided by the proposed approach can be useful to the designers to further enhance the system performance. Hamid Behravan, Ville Hautamäki, Sabato Marco Siniscalchi, Tomi Kinnunen, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2015 | Exploring ANN back-ends for i-vector based speaker age estimationabstractWe address the problem of speaker age estimation using i-vectors. We first compare different i-vector extraction setups and then focus on (shallow) artificial neural net (ANN) back-ends. We explore ANN architecture, training algorithm and ANN ensembles. The results on NIST 2008 and 2010 SRE data indicate that, after extensive parameter optimization, ANN back-end in combination with i-vectors reaches mean absolute errors (MAEs) of 5.49 (females) and 6.35 (males), which are 4.5 % relative improvement in comparison to our support-vector regression (SVR) baseline. Hence, the choice of back-end did not affect the accuracy much; a suggested future direction is therefore focusing more on front-end processing. Index Terms: age estimation, i-vector, multilayer perceptron 1. Anna Fedorova, Ondrej Glembek, Tomi Kinnunen, Pavel Matejka |
INTERSPEECH | 3 |
| 2015 | Classifiers for synthetic speech detection: a comparisonabstractAutomatic speaker verification (ASV) systems are highly vulnerable against spoofing attacks, also known as imposture. With recent developments in speech synthesis and voice conversion technology, it has become important to detect synthesized or voice-converted speech for the security of ASV systems. In this paper, we compare five different classifiers used in speakerrecognition to detect synthetic speech. Experimental results conducted on the ASVspoof 2015 dataset show that supportvector machines with generalized linear discriminant kernel(GLDS-SVM) yield the best performance on the development set with the EER of 0.12 % whereas Gaussian mixture model (GMM) trained using maximum likelihood (ML) criterion with the EER of 3.01 % is superior for the evaluation set. Cemal Hanilçi, Tomi Kinnunen, Md. Sahidullah, Aleksandr Sizov |
INTERSPEECH | 2 |
| 2015 | Speaker recognition for speech under face coverabstractSpeech under face cover constitute a case that is increasingly met by forensic speech experts. Wearing face cover mostly hap-pens when an individual strives to conceal his or her identity. Based on the material of face cover and the level of contact with speech production organs, speech production becomes affected by face mask and a part of speech energy gets absorbed in the mask. There has been little research on how speech acoustics is affected by different face masks and how face covers might affect performance of automatic speaker recognition systems. In the present paper, we have collected speech under face mask with the aim of studying the effects of wearing different masks on state-of-the-art text-independent automatic speaker recogni-tion system. The preliminary speaker recognition rates along with mask identification experiments are presented in this pa-per. Rahim Saeidi, Tuija Niemi, Hanna Karppelin, Jouni Pohjalainen, Tomi Kinnunen, Paavo Alku |
INTERSPEECH | 5 |
| 2015 | A comparison of features for synthetic speech detectionabstractThe performance of biometric systems based on automatic speaker recognition technology is severely degraded due to spoofing attacks with synthetic speech generated using different voice conversion (VC) and speech synthesis (SS) techniques. Various countermeasures are proposed to detect this type of at-tack, and in this context, choosing an appropriate feature extrac-tion technique for capturing relevant information from speech is an important issue. This paper presents a concise experi-mental review of different features for synthetic speech detec-tion task. A wide variety of features considered in this study include previously investigated features as well as some other potentially useful features for characterizing real and synthetic speech. The experiments are conducted on recently released ASVspoof 2015 corpus containing speech data from a large number of VC and SS technique. Comparative results using two different classifiers indicate that features representing spectral information in high-frequency region, dynamic information of speech, and detailed information related to subband characteris-tics are considerably more useful in detecting synthetic speech. Index Terms: anti-spoofing, ASVspoof 2015, feature extrac-tion, countermeasures Md. Sahidullah, Tomi Kinnunen, Cemal Hanilçi |
INTERSPEECH | 2 |
| 2015 | Automatic speaker verification spoofing and countermeasures (ASVspoof 2015): introductory talk by the organizers
Zhizheng Wu 0001, Tomi Kinnunen |
INTERSPEECH | 2 |
| 2015 | ASVspoof 2015: the first automatic speaker verification spoofing and countermeasures challengeabstractAn increasing number of independent studies have con-firmed the vulnerability of automatic speaker verification (ASV) technology to spoofing. However, in comparison to that involving other biometric modalities, spoofing and countermea-sure research for ASV is still in its infancy. A current barrier to progress is the lack of standards which impedes the comparison of results generated by different researchers. The ASVspoof ini-tiative aims to overcome this bottleneck through the provision of standard corpora, protocols and metrics to support a common evaluation. This paper introduces the first edition, summaries the results and discusses directions for future challenges and re-search. Zhizheng Wu 0001, Tomi Kinnunen, Nicholas W. D. Evans, Junichi Yamagishi, Cemal Hanilçi, Md. Sahidullah, Aleksandr Sizov |
INTERSPEECH | 2 |
| 2015 | Factors affecting i-vector based foreign accent recognition: A case study in spoken Finnish
Hamid Behravan, Ville Hautamäki, Tomi Kinnunen |
Speech Commun. | 3 |
| 2015 | Automatic versus human speaker verification: The case of voice mimicry
Rosa González Hautamäki, Tomi Kinnunen, Ville Hautamäki, Anne-Maria Laukkanen |
Speech Commun. | 2 |
| 2015 | Spoofing and countermeasures for speaker verification: A survey
Zhizheng Wu 0001, Nicholas W. D. Evans, Tomi Kinnunen, Junichi Yamagishi, Federico Alegre, Haizhou Li 0001 |
Speech Commun. | 3 |
| 2015 | Joint Speaker Verification and Antispoofing in the i-Vector SpaceabstractAny biometric recognizer is vulnerable to spoofing attacks and hence voice biometric, also called automatic speaker verification (ASV), is no exception; replay, synthesis, and conversion attacks all provoke false acceptances unless countermeasures are used. We focus on voice conversion (VC) attacks considered as one of the most challenging for modern recognition systems. To detect spoofing, most existing countermeasures assume explicit or implicit knowledge of a particular VC system and focus on designing discriminative features. In this paper, we explore back-end generative models for more generalized countermeasures. In particular, we model synthesis-channel subspace to perform speaker verification and antispoofing jointly in the i-vector space, which is a well-established technique for speaker modeling. It enables us to integrate speaker verification and antispoofing tasks into one system without any fusion techniques. To validate the proposed approach, we study vocoder-matched and vocoder-mismatched ASV and VC spoofing detection on the NIST 2006 speaker recognition evaluation data set. Promising results are obtained for standalone countermeasures as well as their combination with ASV systems using score fusion and joint approach. Aleksandr Sizov, Elie Khoury 0001, Tomi Kinnunen, Zhizheng Wu 0001, Sébastien Marcel |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2014 | Introducing attribute features to foreign accent recognitionabstractWe propose a hybrid approach to foreign accent recognition combining both phonotactic and spectral based systems by treating the problem as a spoken language recognition task. We extract speech attribute features that represent speech and acoustic cues reflecting foreign accents of a speaker to obtain feature streams that are modeled with the i-vector methodology. Testing on the Finnish Language Proficiency exam corpus, we find our proposed technique to achieve a significant performance improvement over the state-of-the-art systems using only spectral based features. Hamid Behravan, Ville Hautamäki, Sabato Marco Siniscalchi, Tomi Kinnunen, Chin-Hui Lee 0001 |
ICASSP | 4 |
| 2014 | Bayesian analysis of similarity matrices for speaker diarizationabstractInspired by recent success of speaker clustering in Total Variability space we propose a new probabilistic model for speaker diarization based on Bayesian modeling of pairwise similarity scores. The recordings are represented by symmetric similarity matrices of likelihood ratio scores from probabilistic linear discriminant analysis (PLDA) trained on short-term i-vectors. We employ Bayesian approach to address the problem of unknown number of speakers in conversation. Diarization error rates on the NIST 2008 SRE telephone data indicate comparable performance with state-of-the-art eigenvoice-based diarization. But unlike the eigenvoice approach, our method finds the number of speakers automatically, making the proposed model more viable for practical applications. Alexey Sholokhov, Timur Pekhovsky, Oleg Kudashev, Andrey Shulipa, Tomi Kinnunen |
ICASSP | 5 |
| 2014 | Dialect levelling in Finnish: a universal speech attribute approachabstractWe adopt automatic language recognition methods to study dialect levelling - a phenomenon that leads to reduced structural differences among dialects in a given spoken language. In terms of dialect characterisation, levelling is a nuisance variable that adversely affects recognition accuracy: The more similar two dialects are, the harder it is to set them apart. We address levelling in Finnish regional dialects using a new SAPU (Satakunta in Speech) corpus containing material from Satakunta (South-Western Finland) between 2007 and 2013. To define a compact and universal set of sound units to characterize dialects, we adopt speech attributes features, namely manner and place of articulation. It will be shown that speech attribute distributions can indeed characterise differences among dialects. Experiments with an i-vector system suggest that (1) the attribute features achieve higher dialect recognition accuracy and (2) they are less sensitive against age-related levelling in comparison to traditional spectral approach. Hamid Behravan, Ville Hautamäki, Sabato Marco Siniscalchi, Elie Khoury 0001, Tommi Kurki, Tomi Kinnunen, Chin-Hui Lee 0001 |
INTERSPEECH | 6 |
| 2014 | Introducing i-vectors for joint anti-spoofing and speaker verificationabstractAny biometric recognizer is vulnerable to direct spoofing attacks and automatic speaker verification (ASV) is no exception; replay, synthesis and conversion attacks all provoke false acceptances unless countermeasures are used.We focus on voice conversion (VC) attacks.Most existing countermeasures use full knowledge of a particular VC system to detect spoofing.We study a potentially more universal approach involving generative modeling perspective.Specifically, we adopt standard ivector representation and probabilistic linear discriminant analysis (PLDA) back-end for joint operation of spoofing attack detector and ASV system.As a proof of concept, we study a vocoder-mismatched ASV and VC attack detection approach on the NIST 2006 speaker recognition evaluation corpus.We report stand-alone accuracy of both the ASV and countermeasure systems as well as their combination using score fusion and joint approach.The method holds promise. Elie Khoury 0001, Tomi Kinnunen, Aleksandr Sizov, Zhizheng Wu 0001, Sébastien Marcel |
INTERSPEECH | 2 |
| 2014 | Mixture Linear Prediction in Speaker Verification Under Vocal Effort MismatchabstractThis paper describes an approach to robust signal analysis using iterative parameter re-estimation of a mixture autoregressive (AR) model. The model's focus can be adjusted by initialization of the target and non-target states. The variant examined in this study uses an i.i.d. mixture AR model and is designed to tackle the spectral biasing effect caused by the voice excitation in speech signals with variable fundamental frequency. In our speaker verification experiments, this method performed competitively against standard spectrum analysis techniques in non-mismatch conditions and showed significant improvements in vocal effort mismatch conditions. Jouni Pohjalainen, Cemal Hanilçi, Tomi Kinnunen, Paavo Alku |
IEEE Signal Process. Lett. | 3 |
| 2013 | Speaker identification from shouted speech: Analysis and compensationabstractText-independent speaker identification is studied using neutral and shouted speech in Finnish to analyze the effect of vocal mode mismatch between training and test utterances. Standard mel-frequency cepstral coefficient (MFCC) features with Gaussian mixture model (GMM) recognizer are used for speaker identification. The results indicate that speaker identification accuracy reduces from perfect (100 %) to 8.71 % under vocal mode mismatch. Because of this dramatic degradation in recognition accuracy, we propose to use a joint density GMM mapping technique for compensating the MFCC features. This mapping is trained on a disjoint emotional speech corpus to create a completely speaker- and speech mode independent emotion-neutralizing mapping. As a result of the compensation, the 8.71 % identification accuracy increases to 32.00 % without degrading the non-mismatched train-test conditions much. Cemal Hanilçi, Tomi Kinnunen, Rahim Saeidi, Jouni Pohjalainen, Paavo Alku, Figen Ertas |
ICASSP | 2 |
| 2013 | A practical, self-adaptive voice activity detector for speaker verification with noisy telephone and microphone dataabstractA voice activity detector (VAD) plays a vital role in robust speaker verification, where energy VAD is most commonly used. Energy VAD works well in noise-free conditions but deteriorates in noisy conditions. One way to tackle this is to introduce speech enhancement preprocessing. We study an alternative, likelihood ratio based VAD that trains speech and nonspeech models on an utterance-by-utterance basis from mel-frequency cepstral coefficients (MFCCs). The training labels are obtained from enhanced energy VAD. As the speech and nonspeech models are re-trained for each utterance, minimum assumptions of the background noise are made. According to both VAD error analysis and speaker verification results utilizing state-of-the-art i-vector system, the proposed method outperforms energy VAD variants by a wide margin. We provide open-source implementation of the method. Tomi Kinnunen, Padmanabhan Rajan |
ICASSP | 1 |
| 2013 | Foreign accent detection from spoken Finnish using i-vectorsabstractI-vector based recognition is a well-established technique in state-of-the-art speaker and language recognition but its use in dialect and accent classification has received less attention. We represent an experimental study of i-vector based dialect classi-fication, with a special focus on foreign accent detection from spoken Finnish. Using the CallFriend corpus, we first study how recognition accuracy is affected by the choices of vari-ous i-vector system parameters, such as the number of Gaus-sians, i-vector dimensionality and reduction method. We then apply the same methods on the Finnish national foreign lan-guage certificate (FSD) corpus and compare the results to tra-ditional Gaussian mixture model- universal background model (GMM-UBM) recognizer. The results, in terms of equal error rate, indicate that i-vectors outperform GMM-UBM as one ex-pects. We also notice that in foreign accent detection, 7 out of 9 accents were more accurately detected by Gaussian scoring than by cosine scoring. Index Terms: Dialect recognition, foreign accent recognition, i-vector, GMM-UBM, Finnish language Hamid Behravan, Ville Hautamäki, Tomi Kinnunen |
INTERSPEECH | 3 |
| 2013 | Spoofing and countermeasures for automatic speaker verificationabstractIt is widely acknowledged that most biometric systems are vulnerable to spoofing, also known as imposture.While vulnerabilities and countermeasures for other biometric modalities have been widely studied, e.g.face verification, speaker verification systems remain vulnerable.This paper describes some specific vulnerabilities studied in the literature and presents a brief survey of recent work to develop spoofing countermeasures.The paper concludes with a discussion on the need for standard datasets, metrics and formal evaluations which are needed to assess vulnerabilities to spoofing in realistic scenarios without prior knowledge. Nicholas W. D. Evans, Tomi Kinnunen, Junichi Yamagishi |
INTERSPEECH | 2 |
| 2013 | Comparison of spectrum estimators in speaker verification: mismatch conditions induced by vocal effortabstractWe study the problem of vocal effort mismatch in speaker verification. Changes in speaker’s vocal effort induce changes in fundamental frequency (F0) and formant structure which introduce unwanted intra-speaker variations to features. We compare seven alternative spectrum estimators in the context of melfrequency cepstral coefficient (MFCC) extraction for speaker verification. The compared variants include traditional FFT spectrum and six parametric all-pole models. Experimental results on the NIST 2010 speaker recognition evaluation (SRE) corpus utilizing both GMM-UBM and more recent GMM supervector classifier indicate that spectrum estimation has a considerable impact on speaker verification accuracy under mismatched vocal effort conditions. The highest recognition accuracy was achieved using a particular variant of temporally weighted all-pole model, stabilized weighted linear prediction (SWLP). Index Terms: speaker recognition, vocal effort mismatch, spectrum estimation Cemal Hanilçi, Tomi Kinnunen, Padmanabhan Rajan, Jouni Pohjalainen, Paavo Alku, Figen Ertas |
INTERSPEECH | 2 |
| 2013 | Merging human and automatic system decisions to improve speaker recognition performanceabstractHuman judgment is the final authority in forensic speaker recognition, but the use of modern speaker verification systems with accurate algorithms to perform the task under various circumstances has a huge potential to help the expert. The ultimate goal is to improve the accuracy of automatic systems when challenging data is provided and find a methodology for human-aided speaker recognition systems. This work presents an evaluation of speaker recognition carried out by human listeners and a gender dependent i-vector recognizer with a strategy for fusion of the decision process. Our experiments with HASR 2010 and HASR 2012 data indicate complementarity in Rosa González Hautamäki, Ville Hautamäki, Padmanabhan Rajan, Tomi Kinnunen |
INTERSPEECH | 4 |
| 2013 | I-vectors meet imitators: on vulnerability of speaker verification systems against voice mimicryabstractVoice imitation is mimicry of another speaker’s voice characteristics and speech behavior. Professional voice mimicry can create entertaining, yet realistic sounding target speaker renditions. As mimicry tends to exaggerate prosodic, idiosyncratic and lexical behavior, it is unclear how modern spectral-feature automatic speaker verification systems respond to mimicry “attacks”. We study the vulnerability of two well-known speaker recognition systems, traditional Gaussian mixture model – universal background model (GMM-UBM) and a state-of-the-art i-vector classifier with cosine scoring. The material consists of one professional Finnish imitator impersonating five wellknown Finnish public figures. In a carefully controlled setting, mimicry attack does slightly increase the false acceptance rate for the i-vector system, but generally this is not alarmingly large in comparison to voice conversion or playback attacks. Index Terms: Voice imitation, speaker recognition, mimicry attack Rosa González Hautamäki, Tomi Kinnunen, Ville Hautamäki, Timo Leino, Anne-Maria Laukkanen |
INTERSPEECH | 2 |
| 2013 | Automatic regularization of cross-entropy cost for speaker recognition fusionabstract\n Contains fulltext :\n 116325.pdf (author's version ) (Open Access)\n Ville Hautamäki, Kong-Aik Lee, David A. van Leeuwen, Rahim Saeidi, Anthony Larcher, Tomi Kinnunen, Taufiq Hasan, Seyed Omid Sadjadi, Gang Liu 0001, Hynek Boril, John H. L. Hansen, Benoit G. B. Fauve |
INTERSPEECH | 6 |
| 2013 | Frequency warping and robust speaker verification: a comparison of alternative mel-scale representationsabstractAccuracy of speaker verification is high under controlled condi-tions but falls off rapidly in the presence of interfering sounds. This is because spectral features, such as Mel-frequency cep-stral coefficients (MFCCs), are sensitive to additive noise. MFCCs are a particular realization of warped-frequency rep-resentation with low-frequency focus. But there are several alternative, potentially more robust, warped-frequency repre-sentations. We provide an experimental comparison of five warped-frequency features. They use exactly the same fre-quency warping function, the same number of coefficients and postprocessing, but differ in their internal computations. The compared variants are (1) conventional MFCCs from discrete Fourier transform (DFT), followed by Mel-scaled filterbank, (2) MFCCs via direct warping of DFT, followed by linear-scale fil-terbank, (3) warped linear prediction features, (4) perceptual minimum variance distortionless features and (5) recently pro-posed sparse Mel-scale histogram features. Experiments car-ried out on a subset of the SRE 10 corpus using a scaled-down i-vector system indicate that direct DFT warping outperforms conventional MFCCs in most of the cases. Index Terms: speaker recognition, noise, frequency warping 1. Tomi Kinnunen, Jahangir Alam 0001, Pavel Matejka, Patrick Kenny, Jan Cernocký, Douglas D. O'Shaughnessy |
INTERSPEECH | 1 |
| 2013 | Effect of multicondition training on i-vector PLDA configurations for speaker recognitionabstractThe i-vector representation and PLDA classifier have shown state-of-the-art performance for speaker recognition systems. The availability of more than one enrollment utterance for a speaker allows a variety of configurations which can be used to enhance robustness to noise. The well-known technique of multicondition training can be utilized at different stages of the system, including enrollment and classifier training. We also study the effect of mismatched training, averaging and length normalization. Our study indicates that multicondition training of the PLDA model, and if possible the enrollment i-vectors are the most important to achieve good performance in noisy evaluation data. Index Terms: Speaker verification, i-vector, PLDA, multicondition training Padmanabhan Rajan, Tomi Kinnunen, Ville Hautamäki |
INTERSPEECH | 2 |
| 2013 | Using group delay functions from all-pole models for speaker recognitionabstractPopular features for speech processing, such as mel-frequency cepstral coefficients (MFCCs), are derived from the short-term magnitude spectrum, whereas the phase spectrum remains un-used. While the common argument to use only the magnitude spectrum is that the human ear is phase-deaf, phase-based fea-tures have remained less explored due to additional signal pro-cessing difficulties they introduce. A useful representation of the phase is the group delay function, but its robust computa-tion remains difficult. This paper advocates the use of group delay functions derived from parametric all-pole models instead of their direct computation from the discrete Fourier transform. Using a subset of the vocal effort data in the NIST 2010 speaker recognition evaluation (SRE) corpus, we show that group delay features derived via parametric all-pole models improve recog-nition accuracy, especially under high vocal effort. Addition-ally, the group delay features provide comparable or improved accuracy over conventional magnitude-based MFCC features. Thus, the use of group delay functions derived from all-pole models provide an effective way to utilize information from the phase spectrum of speech signals. Index Terms: speaker verification, group delay functions, high vocal effort Padmanabhan Rajan, Tomi Kinnunen, Cemal Hanilçi, Jouni Pohjalainen, Paavo Alku |
INTERSPEECH | 2 |
| 2013 | Vulnerability evaluation of speaker verification under voice conversion spoofing: the effect of text constraintsabstractVoice conversion, a technique to change one's voice to sound like that of another, poses a threat to even high performance speaker verification system. Vulnerability of text-independent speaker verification systems under spoofing attack, using statistical voice conversion technique, was evaluated and confirmed in our previous work. In this paper, we further extend the study to text-dependent speaker verification systems. In particular, we compare both joint density Gaussian mixture model (JD-GMM) and unit-selection (US) spoofing methods and, for the first time, the performances of text-independent and text-dependent speaker verification systems in a single study. We conduct the experiments using RSR2015 database which is recorded using multiple mobile devices. The experimental results indicate that text-dependent speaker verification system tolerates spoofing attacks better than the text-independent counterpart. Zhizheng Wu 0001, Anthony Larcher, Kong-Aik Lee, Chng Eng Siong, Tomi Kinnunen, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2013 | Exemplar-based unit selection for voice conversion utilizing temporal informationabstractAlthough temporal information of speech has been shown to play an important role in perception, most of the voice conver-sion approaches assume the speech frames are independent of each other, thereby ignoring the temporal information. In this study, we improve conventional unit selection approach by us-ing exemplars which span multiple frames as base units, and also take temporal information constraint into voice conver-sion by using overlapping frames to generate speech parame-ters. This approach thus provides more stable concatenation cost and avoids discontinuity problem in conventional unit se-lection approach. The proposed method also keeps away from the over-smoothing problem in the mainstream joint density Gaussian mixture model (JD-GMM) based conversion method by directly using target speaker’s training data for synthesizing the converted speech. Both objective and subjective evaluations indicate that our proposed method outperforms JD-GMM and conventional unit selection methods. Index Terms: Voice conversion, unit selection, multi-frame ex-emplar, temporal information Zhizheng Wu 0001, Tuomas Virtanen, Tomi Kinnunen, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2013 | I4u submission to NIST SRE 2012: a large-scale collaborative effort for noise-robust speaker verificationabstractI4U is a joint entry of nine research Institutes and Universities across 4 continents to NIST SRE 2012. It started with a brief discussion during the Odyssey 2012 workshop in Singapore. An online discussion group was soon set up, providing a discussion platform for different issues surrounding NIST SRE’12. Noisy test segments, uneven multi-session training, variable enrollment duration, and the issue of open-set identification were actively discussed leading to various solutions integrated to the I4U submission. The joint submission and several of its 17 sub-systems were among top-performing systems. We summarize the lessons learnt from this large-scale effort. Rahim Saeidi, Kong-Aik Lee, Tomi Kinnunen, Tawfik Hasan, Benoit G. B. Fauve, Pierre-Michel Bousquet, Elie Khoury 0001, Pablo Luis Sordo Martinez, Jia Min Karen Kua, Chang Huai You, Hanwu Sun, Anthony Larcher, Padmanabhan Rajan, Ville Hautamäki, Cemal Hanilçi, Billy Braithwaite, Rosa González Hautamäki, Seyed Omid Sadjadi, Gang Liu 0001, Hynek Boril, Navid Shokouhi, Driss Matrouf, Laurent El Shafey, Pejman Mowlaee, Julien Epps, Tharmarajah Thiruvaran, David A. van Leeuwen, Bin Ma 0001, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Sébastien Marcel, John S. D. Mason, Eliathamby Ambikairajah |
INTERSPEECH | 3 |
| 2013 | Multitaper MFCC and PLP features for speaker verification using i-vectors
Jahangir Alam 0001, Tomi Kinnunen, Patrick Kenny, Pierre Ouellet, Douglas D. O'Shaughnessy |
Speech Commun. | 2 |
| 2013 | Sparse Classifier Fusion for Speaker VerificationabstractState-of-the-art speaker verification systems take advantage of a number of complementary base classifiers by fusing them to arrive at reliable verification decisions. In speaker verification, fusion is typically implemented as a weighted linear combination of the base classifier scores, where the combination weights are estimated using a logistic regression model. An alternative way for fusion is to use classifier ensemble selection, which can be seen as sparse regularization applied to logistic regression. Even though score fusion has been extensively studied in speaker verification, classifier ensemble selection is much less studied. In this study, we extensively study a sparse classifier fusion on a collection of twelve I4U spectral subsystems on the NIST 2008 and 2010 speaker recognition evaluation (SRE) corpora. Ville Hautamäki, Tomi Kinnunen, Filip Sedlak, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2013 | Joint Source-Filter Optimization for Accurate Vocal Tract Estimation Using Differential EvolutionabstractIn this work, we present a joint source-filter optimization approach for separating voiced speech into vocal tract (VT) and voice source components. The presented method is pitch-synchronous and thereby exhibits a high robustness against vocal jitter, shimmer and other glottal variations while covering various voice qualities. The voice source is modeled using the Liljencrants-Fant (LF) model, which is integrated into a time-varying auto-regressive speech production model with exogenous input (ARX). The non-convex optimization problem of finding the optimal model parameters is addressed by a heuristic, evolutionary optimization method called differential evolution. The optimization method is first validated in a series of experiments with synthetic speech. Estimated glottal source and VT parameters are the criteria used for comparison with the iterative adaptive inverse filter (IAIF) method and the linear prediction (LP) method under varying conditions such as jitter, fundamental frequency (f0) as well as environmental and glottal noise. The results show that the proposed method largely reduces the bias and standard deviation of estimated VT coefficients and glottal source parameters. Furthermore, the performance of the source-filter separation is evaluated in experiments using speech generated with a physical model of speech production. The proposed method reliably estimates glottal flow waveforms and lower formant frequencies. Results obtained for higher formant frequencies indicate that research on more accurate voice source models and their interaction with the VT is necessary to improve the source-filter separation. The proposed optimization approach promises to be a useful tool for future research addressing this topic. Olaf Schleusing, Tomi Kinnunen, Brad H. Story, Jean-Marc Vesin |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | SWAN - Scientific Writing AssistaNt. A Tool for Helping Scholars to Write Reader-Friendly Manuscripts
Tomi Kinnunen, Henri Leisma, Monika Machunik, Tuomo Kakkonen, Jean-Luc LeBrun |
EACL | 1 |
| 2012 | Comparing spectrum estimators in speaker verification under additive noise degradationabstractDifferent short-term spectrum estimators for speaker verification under additive noise are considered. Conventionally, mel-frequency cepstral coefficients (MFCCs) are computed from discrete Fourier transform (DFT) spectra of windowed speech frames. Recently, linear prediction (LP) and its temporally weighted variants have been substituted as the spectrum analysis method in speech and speaker recognition. In this paper, 12 different short-term spectrum estimation methods are compared for speaker verification under additive noise contamination. Experimental results conducted on NIST 2002 SRE show that the spectrum estimation method has a large effect on recognition performance and stabilized weighted LP (SWLP) and minimum variance distortionless response (MVDR) methods yield approximately 7 % and 8 % relative improvements over the standard DFT method at −10 dB SNR level of factory and babble noises, respectively in terms of equal error rate (EER). Cemal Hanilçi, Tomi Kinnunen, Rahim Saeidi, Jouni Pohjalainen, Paavo Alku, Figen Ertas, Johan Sandberg, Maria Sandsten |
ICASSP | 2 |
| 2012 | Vulnerability of speaker verification systems against voice conversion spoofing attacks: The case of telephone speechabstractVoice conversion - the methodology of automatically converting one's utterances to sound as if spoken by another speaker - presents a threat for applications relying on speaker verification. We study vulnerability of text-independent speaker verification systems against voice conversion attacks using telephone speech. We implemented a voice conversion systems with two types of features and nonparallel frame alignment methods and five speaker verification systems ranging from simple Gaussian mixture models (GMMs) to state-of-the-art joint factor analysis (JFA) recognizer. Experiments on a subset of NIST 2006 SRE corpus indicate that the JFA method is most resilient against conversion attacks. But even it experiences more than 5-fold increase in the false acceptance rate from 3.24 % to 17.33 %. Tomi Kinnunen, Zhizheng Wu 0001, Kong-Aik Lee, Filip Sedlak, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 1 |
| 2012 | Intonational speaker verification: A study on parameters and performance under noisy conditionsabstractProsody-based speaker verification using fundamental frequency (f0) is considered. Our study consists of two phases. First, we do extensive optimization of parameters to establish a baseline system before dealing with noisy conditions. This includes a study of f0extractor parameters, choice of features (discrete cosine transform, discrete Fourier transform, Legendre polynomials, linear prediction), f0track interpolation (none, linear, Hermite), framing parameters and windowing (none, Hamming), f0representation domain (linear, log), number of transformation coefficients and, finally, use of higher-level delta coefficients. Using the optimized parameters, we then explore the robustness of prosody features under white noise and factory noise degradations. Using a GMM-UBM system on the NIST 2006 SRE corpus, we reach an EER of 28.4 % and 27.6 % for the intonational and MFCC features respectively at -20 dB SNR white noise contamination; fusion of the two yields an EER of 24.38 %. Sadjad Siddiq, Tomi Kinnunen, Martti Vainio, Stefan Werner 0003 |
ICASSP | 2 |
| 2012 | Regularized All-Pole Models for Speaker Verification Under Noisy EnvironmentsabstractRegularization of linear prediction based mel-frequency cepstral coefficient (MFCC) extraction in speaker verification is considered. Commonly, MFCCs are extracted from the discrete Fourier transform (DFT) spectrum of speech frames. In this paper, DFT spectrum estimate is replaced with the recently proposed regularized linear prediction (RLP) method. Regularization of temporally weighted variants, weighted LP (WLP) and stabilized WLP (SWLP) which have earlier shown success in speech and speaker recognition, is also introduced. A novel type of double autocorrelation (DAC) lag windowing is also proposed to enhance robustness. Experiments on the NIST 2002 corpus indicate that regularized all-pole methods (RLP, RWLP and RSWLP) yield large improvement on recognition accuracy under additive factory and babble noise conditions in terms of both equal error rate (EER) and minimum detection cost function (MinDCF). Cemal Hanilçi, Tomi Kinnunen, Figen Ertas, Rahim Saeidi, Jouni Pohjalainen, Paavo Alku |
IEEE Signal Process. Lett. | 2 |
| 2012 | Mixture of Factor Analyzers Using Priors From Non-Parallel Speech for Voice ConversionabstractA robust voice conversion function relies on a large amount of parallel training data, which is difficult to collect in practice. To tackle the sparse parallel training data problem in voice conversion, this paper describes a mixture of factor analyzers method which integrates prior knowledge from non-parallel speech into the training of conversion function. The experiments on CMU ARCTIC corpus show that the proposed method improves the quality and similarity of converted speech. With both objective and subjective evaluations, we show the proposed method outperforms the baseline GMM method. Zhizheng Wu 0001, Tomi Kinnunen, Chng Eng Siong, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 2 |
| 2012 | Low-Variance Multitaper MFCC Features: A Case Study in Robust Speaker VerificationabstractIn speech and audio applications, short-term signal spectrum is often represented using mel-frequency cepstral coefficients (MFCCs) computed from a windowed discrete Fourier transform (DFT). Windowing reduces spectral leakage but variance of the spectrum estimate remains high. An elegant extension to windowed DFT is the so-called multitaper method which uses multiple time-domain windows (tapers) with frequency-domain averaging. Multitapers have received little attention in speech processing even though they produce low-variance features. In this paper, we propose the multitaper method for MFCC extraction with a practical focus. We provide, first, detailed statistical analysis of MFCC bias and variance using autoregressive process simulations on the TIMIT corpus. For speaker verification experiments on the NIST 2002 and 2008 SRE corpora, we consider three Gaussian mixture model based classifiers with universal background model (GMM-UBM), support vector machine (GMM-SVM) and joint factor analysis (GMM-JFA). Multitapers improve MinDCF over the baseline windowed DFT by relative 20.4% (GMM-SVM) and 13.7% (GMM-JFA) on the interview-interview condition in NIST 2008. The GMM-JFA system further reduces MinDCF by 18.7% on the telephone data. With these improvements and generally noncritical parameter selection, multitaper MFCCs are a viable candidate for replacing the conventional MFCCs. Tomi Kinnunen, Rahim Saeidi, Filip Sedlak, Kong-Aik Lee, Johan Sandberg, Maria Sandsten, Haizhou Li 0001 |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | A Joint Approach for Single-Channel Speaker Identification and Speech SeparationabstractIn this paper, we present a novel system for joint speaker identification and speech separation. For speaker identification a single-channel speaker identification algorithm is proposed which provides an estimate of signal-to-signal ratio (SSR) as a by-product. For speech separation, we propose a sinusoidal model-based algorithm. The speech separation algorithm consists of a double-talk/single-talk detector followed by a minimum mean square error estimator of sinusoidal parameters for finding optimal codevectors from pre-trained speaker codebooks. In evaluating the proposed system, we start from a situation where we have prior information of codebook indices, speaker identities and SSR-level, and then, by relaxing these assumptions one by one, we demonstrate the efficiency of the proposed fully blind system. In contrast to previous studies that mostly focus on automatic speech recognition (ASR) accuracy, here, we report the objective and subjective results as well. The results show that the proposed system performs as well as the best of the state-of-the-art in terms of perceived quality while its performance in terms of speaker identification and automatic speech recognition results are generally lower. It outperforms the state-of-the-art in terms of intelligibility showing that the ASR results are not conclusive. The proposed method achieves on average, 52.3% ASR accuracy, 41.2 points in MUSHRA and 85.9% in speech intelligibility. Pejman Mowlaee, Rahim Saeidi, Mads Græsbøll Christensen, Zheng-Hua Tan, Tomi Kinnunen, Pasi Fränti, Søren Holdt Jensen |
IEEE Trans. Speech Audio Process. | 5 |
| 2011 | Multi-taper MFCC features for speaker verification using I-vectorsabstractThis paper studies the low-variance multi-taper mel-frequency cepstral coefficient (MFCC) features in the state-of-the-art speaker verification. The MFCC features are usually computed using a Hamming-windowed DFT spectrum. Windowing reduces the bias of the spectrum but variance remains high. Recently, low-variance multi-taper MFCC features were studied in speaker verification with promising preliminary results on the NIST 2002 SRE data using a simple GMM-UBM recognizer. In this study our goal is to validate those findings using a up-to-date i-vector classifier on the latest NIST 2010 SRE data. Our experiment on the telephone (det5) and microphone speech (det1, det2, det3 and det4) indicate that the multi-taper approaches perform better than the conventional Hamming window technique. Jahangir Alam 0001, Tomi Kinnunen, Patrick Kenny, Pierre Ouellet, Douglas D. O'Shaughnessy |
ASRU | 2 |
| 2011 | Multi-site heterogeneous system fusions for the Albayzin 2010 Language Recognition EvaluationabstractBest language recognition performance is commonly obtained by fusing the scores of several heterogeneous systems. Regardless the fusion approach, it is assumed that different systems may contribute complementary information, either because they are developed on different datasets, or because they use different features or different modeling approaches. Most authors apply fusion as a final resource for improving performance based on an existing set of systems. Though relative performance gains decrease as larger sets of systems are considered, best performance is usually attained by fusing all the available systems, which may lead to high computational costs. In this paper, we aim to discover which technologies combine the best through fusion and to analyse the factors (data, features, modeling methodologies, etc.) that may explain such a good performance. Results are presented and discussed for a number of systems provided by the participating sites and the organizing team of the Albayzin 2010 Language Recognition Evaluation. We hope the conclusions of this work help research groups make better decisions in developing language recognition technology. Luis Javier Rodríguez-Fuentes, Mikel Peñagarikano, Amparo Varona, Mireia Díez, Germán Bordel, David Martínez González, Jesús Villalba 0001, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida, Alberto Abad, Oscar Koller, Isabel Trancoso, Paula Lopez-Otero, Laura Docío Fernández, Carmen García-Mateo, Rahim Saeidi, Mehdi Soufifar, Tomi Kinnunen, Torbjørn Svendsen, Pasi Fränti |
ASRU | 19 |
| 2011 | Shout detection in noiseabstractFor the task of detecting shouted speech in a noisy environment, this paper introduces a system based on mel frequency cepstral coefficient (MFCC) feature extraction, unsupervised frame dropping and Gaussian mixture model (GMM) classification. The evaluation material consists of phonemically identical speech and shouting as well as environmental noise of varying levels. The performance of the shout detection system is analyzed by varying the MFCC feature extraction with respect to 1) the feature vector length and 2) the spectrum estimation method. As for feature vector length, the best performance is obtained using 30 MFCC coefficients, which is more than what is conventionally used. In spectrum estimation, a scheme that combines a linear prediction spectrum envelope with spectral fine structure outperforms the conventional FFT. Jouni Pohjalainen, Paavo Alku, Tomi Kinnunen |
ICASSP | 3 |
| 2011 | Classifier subset selection and fusion for speaker verificationabstractState-of-the-art speaker verification systems consists of a number of complementary subsystems whose outputs are fused, to arrive at more accurate and reliable verification decision. In speaker verification, fusion is typically implemented as a linear combination of the subsystem scores. Parameters of the linear model are commonly estimated using the logistic regression method, as implemented in the popular FoCal toolkit. In this paper, we study simultaneous use of classifier selection and fusion. We study four alternative fusion strategies, three score warping techniques, and provide interesting experimental bounds on optimal classifier subset selection. Detailed experiments are carried out on the NIST 2008 and 2010 SRE corpora. Filip Sedlak, Tomi Kinnunen, Ville Hautamäki, Kong-Aik Lee, Haizhou Li 0001 |
ICASSP | 2 |
| 2011 | Regularized Logistic Regression Fusion for Speaker VerificationabstractFusion of the base classifiers is seen as the way to achieve stateof-the art performance in the speaker verfication systems. Standard approach is to pose the fusion problem as the linear binary classification task. Most successful loss function in speaker verification fusion has been the weighted logistic regression popularized by the FoCal toolkit. However, it is known that optimizing logistic regression can overfit severely without appropriate regularization. In addition, subset classifier selection can be achieved by using an external 0/1 loss function on the best subset. In this work, we propose to use LASSO based regularization on the FoCal cost function to achive improved performance and classifier subset selection method integrated into one optimization task. Proposed method is able to achieve 51 % relative improvement in Actual DCF over the FoCal baseline. Index Terms: logistic regression, regularization, compressed sensing, linear fusion, speaker verification Ville Hautamäki, Kong-Aik Lee, Tomi Kinnunen, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2011 | Sinusoidal Approach for the Single-Channel Speech Separation and Recognition ChallengeabstractMost of the single-channel speech separation (SCSS) systems use the short-time Fourier transform as their parametric features. Recent studies have shown that employing sinusoidal features for the SCSS application results in a high perceived speech quality. In this paper, we make a systematic study on automatic speech recognition results for a SCSS system that uses sinusoidal features composed of amplitude and frequency. We compare the speech recognition results with those already reported by other participants in the single-channel speech separation and recognition challenge. Our results show that a newly proposed system achieves an overall recognition accuracy of 52.3%, ranges at the median over all other participants in the challenge. Index Terms: sinusoidal modeling, single-channel speech separation and recognition challenge. Pejman Mowlaee, Rahim Saeidi, Zheng-Hua Tan, Mads Græsbøll Christensen, Tomi Kinnunen, Pasi Fränti, Søren Holdt Jensen |
INTERSPEECH | 5 |
| 2011 | Comparison of clustering methods: A case study of text-independent speaker modeling
Tomi Kinnunen, Ilja Sidoroff, Marko Tuononen, Pasi Fränti |
Pattern Recognit. Lett. | 1 |
| 2011 | Using Discrete Probabilities With Bhattacharyya Measure for SVM-Based Speaker VerificationabstractSupport vector machines (SVMs), and kernel classifiers in general, rely on the kernel functions to measure the pairwise similarity between inputs. This paper advocates the use of discrete representation of speech signals in terms of the probabilities of discrete events as feature for speaker verification and proposes the use of Bhattacharyya coefficient as the similarity measure for this type of inputs to SVM. We analyze the effectiveness of the Bhattacharyya measure from the perspective of feature normalization and distribution warping in the SVM feature space. Experiments conducted on the NIST 2006 speaker verification task indicate that the Bhattacharyya measure outperforms the Fisher kernel, term frequency log-likelihood ratio (TFLLR) scaling, and rank normalization reported earlier in literature. Moreover, the Bhattacharyya measure is computed using a data-independent square-root operation instead of data-driven normalization, which simplifies the implementation. The effectiveness of the Bhattacharyya measure becomes more apparent when channel compensation is applied at the model and score levels. The performance of the proposed method is close to that of the popular GMM supervector with a small margin. Kong-Aik Lee, Chang Huai You, Haizhou Li 0001, Tomi Kinnunen, Khe Chai Sim |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2010 | Towards task-independent person authentication using eye movement signalsabstractWe propose a person authentication system using eye movement signals. In security scenarios, eye-tracking has earlier been used for gaze-based password entry. A few authors have also used physical features of eye movement signals for authentication in a task-dependent scenario with matched training and test samples. We propose and implement a task-independent scenario whereby the training and test samples can be arbitrary. We use short-term eye gaze direction to construct feature vectors which are modeled using Gaussian mixtures. The results suggest that there are personspecific features in the eye movements that can be modeled in a task-independent manner. The range of possible applications extends beyond the security-type of authentication to proactive and user-convenience systems. Tomi Kinnunen, Filip Sedlak, Roman Bednarik |
ETRA | 1 |
| 2010 | Joint frame and Gaussian selection for text independent speaker verificationabstractGaussian selection is a technique applied in the GMM-UBM framework to accelerate score calculation. We have recently introduced a novel Gaussian selection method known as sorted GMM (SGMM). SGMM uses scalar-indexing of the universal background model mean vectors to achieve fast search of the top-scoring Gaussians. In the present work we extend this method by using 2-dimensional indexing, which leads to simultaneous frame and Gaussian selection. Our results on the NIST 2002 speaker recognition evaluation corpus indicate that both the 1- and 2- dimensional SGMMs outperform frame decimation and temporal tracking of top-scoring Gaussians by a wide margin (in terms of Gaussian computations relative to GMM-UBM as baseline). Rahim Saeidi, Tomi Kinnunen, Hamid Reza Sadegh Mohammadi, Robert D. Rodman, Pasi Fränti |
ICASSP | 2 |
| 2010 | Signal-to-Signal Ratio Independent Speaker Identification for Co-channel Speech SignalsabstractIn this paper, we consider speaker identification for the co-channel scenario in which speech mixture from speakers is recorded by one microphone only. The goal is to identify both of the speakers from their mixed signal. High recognition accuracies have already been reported when an accurately estimated signal-to-signal ratio (SSR) is available. In this paper, we approach the problem without estimating SSR. We show that a simple method based on fusion of adapted Gaussian mixture models and Kullback-Leibler divergence calculated between models, achieves an accuracy of 97% and 93% when the two target speakers enlisted as three and two most probable speakers, respectively. Rahim Saeidi, Pejman Mowlaee, Tomi Kinnunen, Zheng-Hua Tan, Mads Græsbøll Christensen, Søren Holdt Jensen, Pasi Fränti |
ICPR | 3 |
| 2010 | Approaching human listener accuracy with modern speaker verificationabstractBeing able to recognize people from their voice is a natural ability that we take for granted. Recent advances have shown significant improvement in automatic speaker recognition performance. Besides being able to process large amount of data in a fraction of time required by human, automatic systems are now able to deal with diverse channel effects. The goal of this paper is to examine how state-of-the-art automatic system performs in comparison with human listeners, and to investigate the strategy for human-assisted form of automatic speaker recognition, which is useful in forensic investigation. We set up an experimental protocol using data from the NIST SRE 2008 core set. A total of 36 listeners have participated in the listening experiments from three sites, namely Australia, Finland and Singapore. State-of-the-art automatic system achieved 20 % error rate, whereas fusion of human listeners achieved 22%. 1. Ville Hautamäki, Tomi Kinnunen, Mohaddeseh Nosratighods, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2010 | What else is new than the hamming window? robust MFCCs for speaker recognition via multitaperingabstractUsually the mel-frequency cepstral coefficients (MFCCs) are derived via Hamming windowed DFT spectrum. In this paper, we advocate to use a so-called multitaper method instead. Mul-titaper methods form a spectrum estimate using multiple win-dow functions and frequency-domain averaging. Multitapers provide a robust spectrum estimate but have not received much attention in speech processing. Our speaker recognition exper-iment on NIST 2002 yields equal error rates (EERs) of 9.66 % (clean data) and 16.41 % (-10 dB SNR) for the conventional Hamming method and 8.13 % (clean data) and 14.63 % (-10 dB SNR) using multitapers. Multitapering is a simple and robust alternative to the Hamming window method. Index Terms: speaker verification, multiple window method 1. Tomi Kinnunen, Rahim Saeidi, Johan Sandberg, Maria Sandsten |
INTERSPEECH | 1 |
| 2010 | Extended weighted linear prediction (XLP) analysis of speech and its application to speaker verification in adverse conditionsabstractThis paper introduces a generalized formulation of linear prediction (LP), including both conventional and temporally weighted LP analysis methods as special cases. The temporally weighted methods have recently been successfully applied to noise robust spectrum analysis in speech and speaker recognition applications. In comparison to those earlier methods, the new generalized approach allows more versatility in weighting different parts of the data in the LP analysis. Two such weighted methods are evaluated and compared to the conventional spectrum modeling methods FFT and LP, as well as the temporally weighted methods WLP and SWLP, by substituting each of them in turn as the spectrum estimation method of the MFCC feature extraction stage of a GMM-UBM based speaker verification system. The new methods are shown to lead to performance improvement in several cases involving channel distortion and additive noise mismatch between the training and recognition conditions. Index Terms: linear prediction, speaker verification, mel frequency cepstral coefficients Jouni Pohjalainen, Rahim Saeidi, Tomi Kinnunen, Paavo Alku |
INTERSPEECH | 3 |
| 2010 | Improving monaural speaker identification by double-talk detectionabstractThis paper describes a novel approach to improve monoaural speaker identification where two speakers are present in a single-microphone recording. The goal is to identify both of the underlying speakers in the given mixture. The proposed approach is composed of a double-talk detector (DTD) as a preprocessor and speaker identification back-end. We demonstrate that including the double-talk detector improves the speaker identification accuracy. Experiments on GRID corpus show that including the DTD improves average recognition accuracy from 96.53% to 97.43%. Rahim Saeidi, Pejman Mowlaee, Tomi Kinnunen, Zheng-Hua Tan, Mads Græsbøll Christensen, Søren Holdt Jensen, Pasi Fränti |
INTERSPEECH | 3 |
| 2010 | Text-independent F0 transformation with non-parallel data for voice conversionabstractIn voice conversion, frame-level mean and variance normal- ization is typically used for fundamental frequency (F0) trans- formation, which is text-independent and requires no parallel training data. Some advanced methods transform pitch con- tours instead, but require either parallel training data or syllabic annotations. We propose a method which retains the simplic- ity and text-independence of the frame-level conversion while yielding high-quality conversion. We achieve these goals by (1) introducing a text-independent tri-frame alignment method, (2) including delta features of F0 into Gaussian mixture model (GMM) conversion and (3) reducing the well-known GMM oversmoothing effect by F0 histogram equalization. Our ob- jective and subjective experiments on the CMU Arctic corpus indicate improvements over both the mean/variance normaliza- tion and the baseline GMM conversion. Index Terms: Voice conversion, F0 transformation, GMM, his- togram equalization, text-independence Zhizheng Wu 0001, Tomi Kinnunen, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2010 | An overview of text-independent speaker recognition: From features to supervectors
Tomi Kinnunen, Haizhou Li 0001 |
Speech Commun. | 1 |
| 2010 | Temporally Weighted Linear Prediction Features for Tackling Additive Noise in Speaker VerificationabstractText-independent speaker verification under additive noise corruption is considered. In the popular mel-frequency cepstral coefficient (MFCC) front-end, the conventional Fourier-based spectrum estimation is substituted with weighted linear predictive methods, which have earlier shown success in noise-robust speech recognition. Two temporally weighted variants of linear predictive modeling are introduced to speaker verification and they are compared to FFT, which is normally used in computing MFCCs, and to conventional linear prediction. The effect of speech enhancement (spectral subtraction) on the system performance with each of the four feature representations is also investigated. Experiments by the authors on the NIST 2002 SRE corpus indicate that the accuracy of the conventional and proposed features are close to each other on clean data. For factory noise at 0 dB SNR level, baseline FFT and the better of the proposed features give EERs of 17.4% and 15.6%, respectively. These accuracies improve to 11.6% and 11.2%, respectively, when spectral subtraction is included as a preprocessing method. The new features hold a promise for noise-robust speaker verification. Rahim Saeidi, Jouni Pohjalainen, Tomi Kinnunen, Paavo Alku |
IEEE Signal Process. Lett. | 3 |
| 2010 | Multitaper Estimation of Frequency-Warped Cepstra With Application to Speaker VerificationabstractUsually the mel-frequency cepstral coefficients are estimated either from a periodogram or from a windowed periodogram. We state a general estimator which also includes multitaper estimators. We propose approximations of the variance and bias of the estimate of each coefficient. By using Monte Carlo computations, we demonstrate that the approximations are accurate. Using the proposed formulas, the peak matched multitaper estimator is shown to have low mean square error (squared bias + variance) on speech-like processes. It is also shown to perform slightly better in the NIST 2006 speaker verification task as compared to the Hamming window conventionally used in this context. Johan Sandberg, Maria Sandsten, Tomi Kinnunen, Rahim Saeidi, Patrick Flandrin, Pierre Borgnat |
IEEE Signal Process. Lett. | 3 |
| 2009 | On separating glottal source and vocal tract information in telephony speaker verificationabstractThe popular mel-frequency cepstral coefficients (MFCCs) capture a mixture of speaker-related, phonemic and channel information. Speaker-related information could be further broken down according to articulatory criteria. How these underlying components are exactly mixed in the features is not well understood. To this end, in this paper we aim at separating the spectra of glottal source and vocal tract using glottal inverse filtering, with an application to speaker recognition over telephone lines. Our experiments on the 10 sec-10 sec condition of the NIST 2006 SRE corpus suggest that the mel-frequency cepstrum of the voice source is not too useful for recognizing speakers. On the contrary, fusing the vocal tract spectrum with conventional MFCCs improves accuracy, suggesting that vocal tract information should be enhanced. Tomi Kinnunen, Paavo Alku |
ICASSP | 1 |
| 2009 | Comparing maximum a posteriori vector quantization and Gaussian mixture models in speaker verificationabstractGaussian mixture model - universal background model (GMM-UBM) is a standard reference classifier in speaker verification. We have proposed a simplified model using vector quantization (VQ-UBM). In this study, we extensively compare these two classifiers on NIST 2005, 2006 and 2008 SRE corpora, while having a standard discriminative classifier (GLDS-SVM) as a reference point. We focus on parameter setting for N-top scoring, model order, and performance for different amounts of training data. The most interesting result, against a general belief, is that GMM-UBM yields better results for short segments whereas VQ-UBM is good for long utterances. The results also suggest that maximum likelihood training of the UBM is sub-optimal, and hence, alternative ways to train the UBM should be considered. Tomi Kinnunen, Juhani Saastamoinen, Ville Hautamäki, Mikko Vinni, Pasi Fränti |
ICASSP | 1 |
| 2009 | Comparative evaluation of maximum a Posteriori vector quantization and gaussian mixture models in speaker verification
Tomi Kinnunen, Juhani Saastamoinen, Ville Hautamäki, Mikko Vinni, Pasi Fränti |
Pattern Recognit. Lett. | 1 |
| 2008 | Characterizing speech utterances for speaker verification with sequence kernel SVMabstractSupport vector machine (SVM) equipped with sequence kernel has been proven to be a powerful technique for speaker verification. A number of sequence kernels have been recently proposed, each being motivated from different perspectives with diverse mathematical derivations. Analytical comparison of kernels becomes difficult. To facilitate such comparisons, we propose a generic structure showing how different levels of cues conveyed by speech utterances, ranging from low-level acoustic features to highlevel speaker cues, are being characterized within a sequence kernel. We then identify the similarities and differences between the popular generalized linear discriminant sequence (GLDS) and GMM supervector kernels, as well as our own probabilistic sequence kernel (PSK). Furthermore, we enhance the PSK in terms of accuracy and computational complexity. The enhanced PSK gives competitive accuracy with the other two kernels. Fusing all the three kernels yields an EER of 4.83 % on the 2006 NIST SRE core test. Index Terms: speaker verification, characteristic vector, support vector machine, sequence kernel Kong-Aik Lee, Chang Huai You, Haizhou Li 0001, Tomi Kinnunen, Donglai Zhu |
INTERSPEECH | 4 |
| 2008 | Text-independent speaker recognition using graph matching
Ville Hautamäki, Tomi Kinnunen, Pasi Fränti |
Pattern Recognit. Lett. | 2 |
| 2008 | Maximum a Posteriori Adaptation of the Centroid Model for Speaker VerificationabstractMaximum a posteriori adapted Gaussian mixture model (GMM-MAP) is widely used in speaker verification. GMMs have three sets of parameters to be adapted: means, covariances, and weights. However, practice has shown that it is sufficient to adapt the means only. Motivated by this, we formulate maximum a posteriori vector quantization (VQ-MAP) procedure which stores and adapts the mean vectors (centroids) only. Experiments on the NIST 2001 and NIST 2006 corpora indicate that VQ-MAP gives comparable accuracy with GMM-MAP with simpler implementation and faster adaptation. Ville Hautamäki, Tomi Kinnunen, Ismo Kärkkäinen, Juhani Saastamoinen, Marko Tuononen, Pasi Fränti |
IEEE Signal Process. Lett. | 2 |
| 2007 | A New Segmentation Algorithm Combined with Transient Frames Power for Text Independent Speaker VerificationabstractIn this paper we propose a new segmentation algorithm called delta MFCC based speech segmentation (DMFCC-SS), with application to speaker recognition systems. We show that DMFCC-SS can separate the regions of speech that result from similar likelihood scores using models such as a Gaussian mixture model (GMM), and can therefore be used to identify the regions of speech between two transitional states in a speech signal. By combining this segmentation algorithm with the discriminative power of transient frames in speaker recognition, we can investigate the tradeoff in speed-up rates that result from DMFCC-SS, with speaker verification equal error rates that result from representatives of each segment. We use a universal background model Gaussian mixture model (UBM-GMM) as a baseline system. The proposed speed-up algorithm, working in the pre-processing stage, performs well while having no computational load compared to the main GMM system. Experimental results show the superior performance of this pre-processing method in comparison with other algorithms working in the pre-processing stage of a UBM-GMM system. Rahim Saeidi, Hamid Reza Sadegh Mohammadi, Robert D. Rodman, Tomi Kinnunen |
ICASSP (4) | 4 |
| 2007 | A GMM-based probabilistic sequence kernel for speaker verificationabstractThis paper describes the derivation of a sequence kernel that transforms speech utterances into probabilistic vectors for classification in an expanded feature space. The sequence kernel is built upon a set of Gaussian basis functions, where half of the basis functions contain speaker specific information while the other half implicates the common characteristics of the competing background speakers. The idea is similar to that in the Gaussian mixture model – universal background model (GMM-UBM) system, except that the Gaussian densities are treated individually in our proposed sequence kernel, as opposed to two mixtures of Gaussian densities in the GMM-UBM system. The motivation is to exploit the individual Gaussian components for better speaker discrimination. Experiments on NIST 2001 SRE corpus show convincing results for the probabilistic sequence kernel approach. Kong-Aik Lee, Chang Huai You, Haizhou Li 0001, Tomi Kinnunen |
INTERSPEECH | 4 |
| 2006 | Joint Acoustic-Modulation Frequency for Speaker RecognitionabstractWe propose a method for computing joint acoustic-modulation frequency feature for speaker recognition. This feature describes the amplitude modulation spectrum of each subband, and results in a single feature vector per utterance. This vector is directly used as the speaker's modulation frequency template, excluding the need for a separate training phase. The effects of analysis parameters and pattern matching are studied using the NIST 2001 corpus. When fusing the proposed feature with the baseline MFCC/GMM system, EER is reduced from 18.2% to 16.7% Tomi Kinnunen |
ICASSP (1) | 1 |
| 2006 | Real-time speaker identification and verificationabstractIn speaker identification, most of the computation originates from the distance or likelihood computations between the feature vectors of the unknown speaker and the models in the database. The identification time depends on the number of feature vectors, their dimensionality, the complexity of the speaker models and the number of speakers. In this paper, we concentrate on optimizing vector quantization (VQ) based speaker identification. We reduce the number of test vectors by pre-quantizing the test sequence prior to matching, and the number of speakers by pruning out unlikely speakers during the identification process. The best variants are then generalized to Gaussian mixture model (GMM) based modeling. We apply the algorithms also to efficient cohort set search for score normalization in speaker verification. We obtain a speed-up factor of 16:1 in the case of VQ-based modeling with minor degradation in the identification accuracy, and 34:1 in the case of GMM-based modeling. An equal error rate of 7% can be reached in 0.84 s on average when the length of test utterance is 30.4 s. Tomi Kinnunen, Evgeny Karpov, Pasi Fränti |
IEEE Trans. Speech Audio Process. | 1 |
| 2004 | Real-time speaker identificationabstractIn speaker identification, most of the computation originates from distance or likelihood computations between the feature vectors of the unknown speaker and the models in the database. The identification time depends on the number of feature vectors, their dimensionality, the complexity of the speaker models and the number of speakers. In this paper, we focus on optimizing vector quantization (VQ) based speaker identification. We reduce the number of test vectors by pre-quantizing the test sequence prior to matching, and the number of speakers by pruning out unlikely speakers during the identification process. The best variants are then generalized to Gaussian mixture model (GMM) based modeling also. We obtain a speed-up factor of 16:1 with VQ-based system, and 34:1 with GMM-based system with a minor degradation in the identification error rate. Pasi Fränti, Evgeny Karpov, Tomi Kinnunen |
INTERSPEECH | 3 |
| 2004 | Efficient online cohort selection method for speaker verificationabstractCohort normalization is a method for normalizing the scores in speaker verification in order to reduce undesirable variation arising from acoustically mismatched conditions. A particular form of cohort normalization, unconstrained cohort normalization (UCN) is addressed in this study. The UCN method has been shown to give excellent results but its major drawback is the huge computational load arising from the search of the cohort speakers. In this paper, we propose a fast cohort search algorithm, that quantizes the test vector sequence and uses the quantized data for both impostor and claimant scoring. Results on the NIST-1999 corpus show a speed-up factor of 23:1 compared to full search. Furthermore, the equal error rates are decreased from those of the full search. Tomi Kinnunen, Evgeny Karpov, Pasi Fränti |
INTERSPEECH | 1 |
| 2003 | On the fusion of dissimilarity-based classifiers for speaker identificationabstractIn this work, we describe a speaker identification system that uses multiple supplementary information sources for computing a combined match score for the unknown speaker. Each speaker profile in the database consists of multiple feature vector sets that can vary in their scale, dimensionality, and the number of vectors. The evidence from a given feature set is weighted by its reliability that is set in a priori fashion. The confidence of the identification result is also estimated. The system is evaluated with a corpus of 110 Finnish speakers. The evaluated feature sets include mel-cepstrum, LPC-cepstrum, dynamic cepstrum, long-term averaged spectrum of /A/ vowel, and F0. Tomi Kinnunen, Ville Hautamäki, Pasi Fränti |
INTERSPEECH | 1 |
| 2002 | Designing a speaker-discriminative adaptive filter bank for speaker recognitionabstractA new filter bank approach for speaker recognition front-end is proposed. The conventional mel-scaled filter bank is replaced with a speaker-discriminative filter bank. Filter bank is selected from a library in adaptive basis, based on the broad phoneme class of the input frame. Each phoneme class is associated with its own filter bank. Each filter bank is designed in a way that emphasizes discriminative subbands that are characteristic for that phoneme. Experiments on TIMIT corpus show that the proposed method outperforms traditional MFCC features. Tomi Kinnunen |
INTERSPEECH | 1 |
| 2001 | Is speech data clustered? - statistical analysis of cepstral featuresabstractSpeech analysis applications are typically based on short-term spectral analysis of the speech signal. Feature extraction process outputs one feature vector per frame. The features are further processed by application-dependent techniques, such as hidden Markov models or vector quantization. Independent from the application, it is often desirable that the feature vectors form separable clusters in the feature space. In this work, we study whether data is really clustered in the feature space and, if so, what is the number of the clusters in typical speech data. We consider different forms of the widely used cepstral features. Tomi Kinnunen, Ismo Kärkkäinen, Pasi Fränti |
INTERSPEECH | 1 |