EDBT 2026 Demo / reviewers in the wild / expert
Kong-Aik Lee
dblp:35/4621 · also Kong Aik Lee
· DBLP profile ↗
170ranked-venue papers
22as first author
73since 2021 · last 2026
0000-0001-9133-3000ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 132 · 19 first-author · 52 since 2021Artificial intelligence and machine learning · 96 · 16 first-author · 35 since 2021Security and privacy · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ASVspoof 5: Design, collection and validation of resources for spoofing, deepfake, and adversarial attack detection using crowdsourced speechabstractASVspoof 5 is the fifth edition in a series of challenges which promote the study of speech spoofing and deepfake attacks as well as the design of detection solutions. We introduce the ASVspoof 5 database which is generated in a crowdsourced fashion from data collected in diverse acoustic conditions (cf. studio-quality data for earlier ASVspoof databases) and from ∼ 2,000 speakers (cf. ∼ 100 earlier). The database contains attacks generated with 32 different algorithms, also crowdsourced, and optimised to varying degrees using new surrogate detection models. Among them are attacks generated with a mix of legacy and contemporary text-to-speech synthesis and voice conversion models, in addition to adversarial attacks which are incorporated for the first time. ASVspoof 5 protocols comprise seven speaker-disjoint partitions. They include two distinct partitions for the training of different sets of attack models, two more for the development and evaluation of surrogate detection models, and then three additional partitions which comprise the ASVspoof 5 training, development and evaluation sets. An auxiliary set of data collected from an additional 30k speakers can also be used to train speaker encoders for the implementation of attack algorithms. Also described herein is an experimental validation of the new ASVspoof 5 database using a set of automatic speaker verification and spoof/deepfake baseline detectors. With the exception of protocols and tools for the generation of spoofed/deepfake speech, the resources described in this paper, already used by participants of the ASVspoof 5 challenge in 2024, are now all freely available to the community. Xin Wang 0037, Héctor Delgado, Hemlata Tak, Jee-Weon Jung, Hye-Jin Shim, Massimiliano Todisco, Ivan Kukanov, Xuechen Liu 0001, Md. Sahidullah, Tomi Kinnunen, Nicholas W. D. Evans, Kong-Aik Lee, Junichi Yamagishi, Myeonghun Jeong, Yongyi Zang, Soumi Maiti, Florian Lux, Nicolas Müller, Wangyou Zhang, Chengzhe Sun 0001, Shuwei Hou, Siwei Lyu, Sébastien Le Maguer, Hanjie Guo, Vishwanath Pratap Singh |
Comput. Speech Lang. | 12 |
| 2026 | A Study of the Removability of Speaker-Adversarial PerturbationsabstractRecent advancements in adversarial attacks have demonstrated their effectiveness in misleading speaker recognition models, making wrong predictions about speaker identities. On the other hand, defense techniques against speaker-adversarial attacks focus on reducing the effects of speaker-adversarial perturbations on speaker attribute extraction. These techniques do not seek to fully remove the perturbations and restore the original speech. To this end, this paper studies the removability of speaker-adversarial perturbations. Specifically, the investigation is conducted assuming various degrees of awareness of the perturbation generator across three scenarios: ignorant, semi-informed, and well-informed. Besides, we consider both the optimization-based and feedforward perturbation generation methods. Experiments conducted on the LibriSpeech dataset demonstrated that: 1) in the ignorant scenario, speaker-adversarial perturbations cannot be eliminated, although their impact on speaker attribute extraction is reduced, 2) in the semi-informed scenario, the speaker-adversarial perturbations cannot be fully removed, while those generated by the feedforward model can be considerably reduced, and 3) in the well-informed scenario, speaker-adversarial perturbations are nearly eliminated, allowing for the restoration of the original speech. Chenyang Guo, Kong-Aik Lee, Zhen-Hua Ling, Wu Guo |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | Grouped Knowledge Distillation with Adaptive Logit Softening for Speaker RecognitionabstractRecent works suggest that decoupling the information of non-target speakers from that of the target speaker in knowledge distillation (KD) and subsequently emphasizing the former can lead to significant performance improvement. However, a well-trained teacher model typically produces almost zero non-target speaker posteriors with limited contribution to knowledge transfer, resulting in a less effective KD. To address this problem, we advocate a dual-group knowledge distillation framework, wherein the primary group with top-k speaker posteriors captures most of the speaker discrimination knowledge in an utterance. The non-primary group contributes to the KD through a binary classification (distillation) between the primary and non-primary groups. In addition, adaptive logit softening is proposed to adjust the teacher’s and student’s logits in the binary distillation, further facilitating effective knowledge transfer. The proposed method trained with a simple x-vector pipeline obtains an impressive equal error rate of 1.46%, 1.47%, and 2.70% on three VoxCeleb1 test sets, outperforming the state-of-the-art methods with a noticeable margin. Chong-Xin Gan, Youzhi Tu, Zezhong Jin, Man-Wai Mak, Kong-Aik Lee |
ICASSP | 5 |
| 2025 | LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation GenerationabstractPrevious fake speech datasets were constructed from a defender’s perspective to develop countermeasure (CM) systems without considering diverse motivations of attackers. To better align with real-life scenarios, we created LlamaPartialSpoof, a 130-hour dataset that contains both fully and partially fake speech, using a large language model (LLM) and voice cloning technologies to evaluate the robustness of CMs. By examining valuable information for both attackers and defenders, we identify several key vulnerabilities in current CM systems, which can be exploited to enhance attack success rates, including biases toward certain text-to-speech models or concatenation methods. Our experimental results indicate that the current fake speech detection system struggle to generalize to unseen scenarios, achieving a best performance of 24.49% equal error rate. Hieu-Thi Luong, Haoyang Li 0018, Lin Zhang 0054, Kong-Aik Lee, Chng Eng Siong |
ICASSP | 4 |
| 2025 | Text-dependent Speaker Verification Challenge 2024: Exploring Shared and User-defined PassphrasesabstractIn contrast to text-independent speaker verification, which has received significant attention from researchers and has many competitions dedicated to it, text-dependent speaker verification (TdSV) has been less explored recently. The TdSV Challenge 2024 was organized to analyze and explore novel methods for this type of speaker verification and aims to motivate participants to develop new approaches to TdSV, conduct comprehensive analyses, and investigate advanced techniques such as self-supervised learning. This challenge builds on the achievements of the short-duration speaker verification (SdSV) Challenges held in 2020 and 2021 and focuses specifically on TdSV in two distinct scenarios. The first scenario involves conventional TdSV, while the second focuses on speaker enrollment using user-defined passphrases. This paper provides a detailed description of both tasks, introduces the evaluation rules, and presents a comprehensive analysis of the results obtained from this challenge. Hossein Zeinali, Kong-Aik Lee, Jahangir Alam 0001, Lukás Burget |
ICASSP | 2 |
| 2025 | MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual CuesabstractAudio-visual Target Speaker Extraction (AV-TSE) aims to isolate the speech of a specific target speaker from an audio mixture using time-synchronized visual cues. In real-world scenarios, visual cues are not always available due to various impairments, which undermines the stability of AV-TSE. Despite this challenge, humans can maintain attentional momentum over time, even when the target speaker is not visible. In this paper, we introduce the Momentum Multi-modal target Speaker Extraction (MoMuSE), which retains a speaker identity momentum in memory, enabling the model to continuously track the target speaker. Designed for real-time inference, MoMuSE extracts the current speech window with guidance from both visual cues and dynamically updated speaker momentum. Experimental results demonstrate that MoMuSE exhibits significant improvement, particularly in scenarios with severe impairment of visual cues. Shuai Wang 0016, Kong-Aik Lee, Man-Wai Mak, Haizhou Li 0001 |
ICME | 4 |
| 2025 | IDIR: Identifying and Distilling Informative Relations for Speaker Verification
Chong-Xin Gan, Zhe Li 0030, Zezhong Jin, Man-Wai Mak, Kong-Aik Lee |
INTERSPEECH | 6 |
| 2025 | Bayesian Learning for Domain-Invariant Speaker Verification and Anti-Spoofing
Man-Wai Mak, Johan Rohdin, Kong-Aik Lee, Hynek Hermansky |
INTERSPEECH | 4 |
| 2025 | The Sub-3Sec Problem: From Text-Independent to Text-Dependent Corpus
Ruichen Zuo, Kong-Aik Lee, Man-Wai Mak |
INTERSPEECH | 2 |
| 2025 | Quantifying prediction uncertainties in automatic speaker verification systemsabstractFor modern automatic speaker verification (ASV) systems, explicitly quantifying the confidence for each prediction strengthens the system’s reliability by indicating in which case the system is with trust. However, current paradigms do not take this into consideration. We thus propose to express confidence in the prediction by quantifying the uncertainty in ASV predictions. This is achieved by developing a novel Bayesian framework to obtain a score distribution for each input. The mean of the distribution is used to derive the decision while the spread of the distribution represents the uncertainty arising from the plausible choices of the model parameters. To capture the plausible choices, we sample the probabilistic linear discriminant analysis (PLDA) back-end model posterior through Hamiltonian Monte-Carlo (HMC) and approximate the embedding model posterior through stochastic Langevin dynamics (SGLD) and Bayes-by-backprop. Given the resulting score distribution, a further quantification and decomposition of the prediction uncertainty are achieved by calculating the score variance, entropy, and mutual information. The quantified uncertainties include the aleatoric uncertainty and epistemic uncertainty (model uncertainty). We evaluate them by observing how they change while varying the amount of training speech, the duration, and the noise level of testing speech. The experiments indicate that the behaviour of those quantified uncertainties reflects the changes we made to the training and testing data, demonstrating the validity of the proposed method as a measure of uncertainty. • The paper emphasises the need for quantifying and separating uncertainties in ASV. • The paper proposes a novel framework incorporating various Bayesian learning methods. • The major cause of epistemic uncertainty is training data size and test data length. • The noise level in the test utterance increases the aleatoric uncertainty in ASV. Miao Jing, Vidhyasaharan Sethu, Beena Ahmed, Kong-Aik Lee |
Comput. Speech Lang. | 4 |
| 2025 | Adversarially adaptive temperatures for decoupled knowledge distillation with applications to speaker verificationabstract202502 bcch Zezhong Jin, Youzhi Tu, Chong-Xin Gan, Man-Wai Mak, Kong-Aik Lee |
Neurocomputing | 5 |
| 2025 | ConFusionformer: Locality-enhanced Conformer through multi-resolution attention fusion for speaker verification
Youzhi Tu, Man-Wai Mak, Kong-Aik Lee, Weiwei Lin 0002 |
Neurocomputing | 3 |
| 2025 | Make full use of your data: On copy-based augmentation in speech anti-spoofing
Linjuan Zhang, Baoning Niu, Kong-Aik Lee, Longbiao Wang |
Neurocomputing | 3 |
| 2025 | Pinhole Effect on Linkability and Dispersion in Speaker AnonymizationabstractSpeaker anonymization aims to conceal speaker-specific attributes in speech signals, making the anonymized speech unlinkable to the original speaker identity. Recent approaches achieve this by disentangling speech into content and speaker components, replacing the latter with pseudo- speakers. The anonymized speech can be mapped either to a common pseudo-speaker shared across instances or to distinct pseudo-speakers unique to each instance. This paper investigates the impact of these mapping strategies on three key dimensions: speaker linkability, dispersion in the anonymized speaker space, and de-identification from the original identity. Our findings show that using distinct pseudo-speakers increases speaker dispersion and reduces linkability compared to common pseudo-speaker mapping, while maintaining de-identification, thereby enhancing overall privacy preservation. These observations are interpreted through the proposedpinhole effect, a conceptual framework introduced to explain the relationship between mapping strategies and anonymization performance. The hypothesis is validated through empirical evaluation. Kong-Aik Lee, Zeyan Liu, Zhen-Hua Ling |
IEEE Signal Process. Lett. | 1 |
| 2025 | Asynchronous Voice Anonymization by Learning From Speaker-Adversarial SpeechabstractThis paper focuses on asynchronous voice anonymization, wherein machine-discernible speaker attributes in a speech utterance are obscured while human perception is preserved. We propose to transfer the voice-protection capability of speaker-adversarial speech to speaker embedding, thereby facilitating the modification of speaker embedding extracted from original speech to generate anonymized speech. Experiments conducted on the LibriSpeech dataset demonstrated that compared to the speaker-adversarial utterances, the generated anonymized speech demonstrates improved transferability and voice-protection capability. Furthermore, the proposed method enhances the human perception preservation capability of anonymized speech within the generative asynchronous voice anonymization framework. Kong-Aik Lee, Zhen-Hua Ling |
IEEE Signal Process. Lett. | 3 |
| 2025 | Any-to-Any Speaker Attribute Perturbation for Asynchronous Voice AnonymizationabstractSpeaker attribute perturbation offers a feasible approach to asynchronous voice anonymization by employing adversarially perturbed speech as anonymized output. In order to enhance the identity unlinkability among anonymized utterances from the same original speaker, the targeted attack training strategy is usually applied to anonymize the utterances to a common designated speaker. However, this strategy may violate the privacy of the designated speaker who is an actual speaker. To mitigate this risk, this paper proposes an any-to-any training strategy. It is accomplished by defining a batch mean loss to anonymize the utterances from various speakers within a training mini-batch to a common pseudo-speaker, which is approximated as the average speaker in the mini-batch. Based on this, a speaker-adversarial speech generation model is proposed, incorporating the supervision from both the untargeted attack and the any-to-any strategies. The speaker attribute perturbations are generated and incorporated into the original speech to produce its anonymized version. The effectiveness of the proposed model was justified in asynchronous voice anonymization through experiments conducted on the LibriSpeech datasets. Additional experiments were carried out to explore the potential limitations of speaker-adversarial speech in voice privacy protection. With them, we aim to provide insights for future research on its protective efficacy against black-box speaker extractors and adaptive attacks, as well as generalization to out-of-domain datasets and stability. Audio samples and open-source code are published in https://github.com/VoicePrivacy/any-to-any-speaker-attribute-perturbation. Chenyang Guo, Kong-Aik Lee, Zhen-Hua Ling |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Nes2Net: A Lightweight Nested Architecture for Foundation Model Driven Speech Anti-SpoofingabstractSpeech foundation models have significantly advanced various speech-related tasks by providing exceptional representation capabilities. However, their high-dimensional output features often create a mismatch with downstream task models, which typically require lower-dimensional inputs. A common solution is to apply a dimensionality reduction (DR) layer, but this approach increases parameter overhead, computational costs, and risks losing valuable information. To address these issues, we propose Nested Res2Net (Nes2Net), a lightweight back-end architecture designed to directly process high-dimensional features without DR layers. The nested structure enhances multi-scale feature extraction, improves feature interaction, and preserves high-dimensional information. We first validate Nes2Net on CtrSVDD, a singing voice deepfake detection dataset, and report a 22% performance improvement and an 87% back-end computational cost reduction over the state-of-the-art baseline. Additionally, extensive testing across four diverse datasets: ASVspoof 2021, ASVspoof 5, PartialSpoof, and In-the-Wild, covering fully spoofed speech, adversarial attacks, partial spoofing, and real-world scenarios, consistently highlights Nes2Net’s superior robustness and generalization capabilities. The code package and pre-trained models are available at https://github.com/Liu-Tianchi/Nes2Net. Tianchi Liu 0004, Duc-Tuan Truong, Rohan Kumar Das, Kong-Aik Lee, Haizhou Li 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Adversarial Speech for Voice Privacy Protection from Personalized Speech GenerationabstractThe rapid progress in personalized speech generation technology, including personalized text-to-speech (TTS) and voice conversion (VC), poses a challenge in distinguishing between generated and real speech for human listeners, resulting in an urgent demand in protecting speakers' voices from malicious misuse. In this regard, we propose a speaker protection method based on adversarial attacks. The proposed method perturbs speech signals by minimally altering the original speech while rendering downstream speech generation models unable to accurately generate the voice of the target speaker. For validation, we employ the open-source pre-trained YourTTS model for speech generation and protect the target speaker's speech in the white-box scenario. Automatic speaker verification (ASV) evaluations were carried out on the generated speech as the assessment of the voice protection capability. Our experimental results show that we successfully perturbed the speaker encoder of the YourTTS model using the gradient-based I-FGSM adversarial perturbation method. Furthermore, the adversarial perturbation is effective in preventing the YourTTS model from generating the speech of the target speaker. Audio samples can be found in https://voiceprivacy.github.io/Adeversarial-Speech-with-YourTTS. Shihao Chen, Jie Zhang 0042, Kong-Aik Lee, Zhen-Hua Ling, Li-Rong Dai 0001 |
ICASSP | 4 |
| 2024 | Modeling Pseudo-Speaker Uncertainty in Voice AnonymizationabstractVoice anonymization refers to the goal of suppressing personally identifiable voice attributes in speech. State-of-the-art models based on the voice conversion framework accomplish this goal by replacing the voice attributes of the speaker with those of a pseudo-speaker. This paper proposes to exploit the uncertainty estimate of pseudo-speaker in voice anonymization. For each target speaker, a pseudo-speaker distribution, characterized by a point estimate and its uncertainty, is estimated from a selected set of cohort speakers. Based on this distribution, a pseudo-speaker vector is sampled and used to replace the voice attributes in an anonymized speech. The efficacy of the proposed method was validated in the framework as provided by VoicePrivacy Challenge 2022. Audio samples can be found in https://voiceprivacy.github.io/pseudo-speaker-vector/. Kong-Aik Lee, Wu Guo, Zhen-Hua Ling |
ICASSP | 2 |
| 2024 | Gradient Weighting for Speaker Verification in Extremely Low Signal-to-Noise RatioabstractSpeaker verification is hampered by background noise, particularly at extremely low Signal-to-Noise Ratio (SNR) under 0 dB. It is difficult to suppress noise without introducing unwanted artifacts, which adversely affects speaker verification. We proposed the mechanism called Gradient Weighting (Grad-W), which dynamically identifies and reduces artifact noise during prediction. The mechanism is based on the property that the gradient indicates which parts of the input the model is paying attention to. Specifically, when the speaker network focuses on a region in the denoised utterance but not on the clean counterpart, we consider it artifact noise and assign higher weights for this region during optimization of enhancement. We validate it by training an enhancement model and testing the enhanced utterance on speaker verification. The experimental results show that our approach effectively reduces artifact noise, improving speaker verification across various SNR levels. Kong-Aik Lee, Ville Hautamäki, Meng Ge, Haizhou Li 0001 |
ICASSP | 2 |
| 2024 | Emphasized Non-Target Speaker Knowledge in Knowledge Distillation for Automatic Speaker VerificationabstractKnowledge distillation (KD) is used to enhance automatic speaker verification performance by ensuring consistency between large teacher networks and lightweight student networks at the embedding level or label level. However, the conventional label-level KD overlooks the significant knowledge from non-target speakers, particularly their classification probabilities, which can be crucial for automatic speaker verification. In this paper, we first demonstrate that leveraging a larger number of training non-target speakers improves the performance of automatic speaker verification models. Inspired by this finding about the importance of non-target speakers’ knowledge, we modified the conventional label-level KD by disentangling and emphasizing the classification probabilities of non-target speakers during knowledge distillation. The proposed method is applied to three different student model architectures and achieves an average of 13.67% improvement in EER on the VoxCeleb dataset compared to embedding-level and conventional label-level KD methods.1 Duc-Tuan Truong, Ruijie Tao, Jia Qi Yip, Kong-Aik Lee, Chng Eng Siong |
ICASSP | 4 |
| 2024 | CPAUG: Refining Copy-Paste Augmentation for Speech Anti-SpoofingabstractConventional copy-paste augmentations generate new training instances by concatenating existing utterances to increase the amount of data for neural network training. However, the direct application of copy-paste augmentation for anti-spoofing is problematic. This paper refines the copy-paste augmentation for speech anti-spoofing, dubbed CpAug, to generate more training data with rich intra-class diversity. The CpAug employs two policies: concatenation to merge utterances with identical labels, and substitution to replace segments in an anchor utterance. Besides, considering the impacts of speakers and spoofing attack types, we craft four blending strategies for the CpAug. Furthermore, we explore how CpAug complements the Rawboost augmentation method. Experimental results reveal that the proposed CpAug significantly improves the performance of speech anti-spoofing. Particularly, CpAug with substitution policy leads to relative improvements of 43% and 38% on the ASVspoof’ 19LA and 21LA, respectively. Notably, the CpAug and Rawboost synergize effectively, achieving an EER of 2.91% on ASVspoof’ 21LA. Linjuan Zhang, Kong-Aik Lee, Lin Zhang 0054, Longbiao Wang, Baoning Niu |
ICASSP | 2 |
| 2024 | Two-stage Semi-supervised Speaker Recognition with Gated Label Learning
Xingmei Wang 0002, Jiaxiang Meng, Kong-Aik Lee, Boquan Li 0002, Jinghan Liu |
IJCAI | 3 |
| 2024 | Revisiting and Improving Scoring Fusion for Spoofing-aware Speaker Verification Using Compositional Data AnalysisabstractInterspeech 2024, 1-5 September 2024, Kos, Greece Xin Wang 0037, Tomi Kinnunen, Kong-Aik Lee, Paul-Gauthier Noé, Junichi Yamagishi |
INTERSPEECH | 3 |
| 2024 | MM-NodeFormer: Node Transformer Multimodal Fusion for Emotion Recognition in ConversationabstractInterspeech 2024, 1-5 September 2024, Kos, Greece Man-Wai Mak, Kong-Aik Lee |
INTERSPEECH | 3 |
| 2024 | Temporal-Channel Modeling in Multi-head Self-Attention for Synthetic Speech DetectionabstractRecent synthetic speech detectors leveraging the Transformer model have superior performance compared to the convolutional neural network counterparts.This improvement could be due to the powerful modeling ability of the multi-head selfattention (MHSA) in the Transformer model, which learns the temporal relationship of each input token.However, artifacts of synthetic speech can be located in specific regions of both frequency channels and temporal segments, while MHSA neglects this temporal-channel dependency of the input sequence.In this work, we proposed a Temporal-Channel Modeling (TCM) module to enhance MHSA's capability for capturing temporalchannel dependencies.Experimental results on the ASVspoof 2021 show that with only 0.03M additional parameters, the TCM module can outperform the state-of-the-art system by 9.25% in EER.Further ablation study reveals that utilizing both temporal and channel information yields the most improvement for detecting synthetic speech 1 . Duc-Tuan Truong, Ruijie Tao, Hieu-Thi Luong, Kong-Aik Lee, Chng Eng Siong |
INTERSPEECH | 5 |
| 2024 | Asynchronous Voice Anonymization Using Adversarial Perturbation On Speaker EmbeddingabstractInterspeech 2024, 1-5 September 2024, Kos, Greece Kong-Aik Lee, Zhen-Hua Ling |
INTERSPEECH | 3 |
| 2024 | On The Generation and Removal of Speaker Adversarial Perturbation For Voice-Privacy ProtectionabstractNeural networks are commonly known to be vulnerable to adversarial attacks mounted through subtle perturbation on the input data. Recent development in voice-privacy protection has shown the positive use cases of the same technique to conceal speaker’s voice attribute with additive perturbation signal generated by an adversarial network. This paper examines the reversibility property where an entity generating the adversarial perturbations is authorized to remove them and restore original speech (e.g., the speaker him/herself). A similar technique could also be used by an investigator to deanonymize a voice-protected speech to restore criminals’ identities in security and forensic analysis. In this setting, the perturbation generative module is assumed to be known in the removal process. To this end, a joint training of perturbation generation and removal modules is proposed. Experimental results on the LibriSpeech dataset demonstrated that the subtle perturbations added to the original speech can be predicted from the anonymized speech while achieving the goal of privacy protection. By removing these perturbations from the anonymized sample, the original speech can be restored. Audio samples can be found in https://voiceprivacy.github.io/Perturbation-Generation-Removal/. Chenyang Guo, Zhuhai Li, Kong-Aik Lee, Zhen-Hua Ling, Wu Guo |
SLT | 4 |
| 2024 | On the Effectiveness of Enrollment Speech Augmentation For Target Speaker ExtractionabstractDeep learning technologies have significantly advanced the performance of target speaker extraction (TSE) tasks. To enhance the generalization and robustness of these algorithms when training data is insufficient, data augmentation is a commonly adopted technique. Unlike typical data augmentation applied to speech mixtures, this work thoroughly investigates the effectiveness of augmenting the enrollment speech space. We found that for both pretrained and jointly optimized speaker encoders, directly augmenting the enrollment speech leads to consistent performance improvement. In addition to conventional methods such as noise and reverberation addition, we propose a novel augmentation method called self-estimated speech augmentation (SSA). Experimental results on the Libri2Mix test set show that our proposed method can achieve an improvement of up to 2.5 dB. Shuai Wang 0016, Haizhou Li 0001, Man-Wai Mak, Kong-Aik Lee |
SLT | 6 |
| 2024 | Towards Quantifying and Reducing Language Mismatch Effects in Cross-Lingual Speech Anti-SpoofingabstractThe effects of language mismatch impact speech anti-spoofing systems, while investigations and quantification of these effects remain limited. Existing anti-spoofing datasets are mainly in English, and the high cost of acquiring multilingual datasets hinders training language-independent models. We initiate this work by evaluating top-performing speech anti-spoofing systems that are trained on English data but tested on other languages, observing notable performance declines. We propose an innovative approach - Accent-based data expansion via TTS (ACCENT), which introduces diverse linguistic knowledge to monolingual-trained models, improving their cross-lingual capabilities. We conduct experiments on a large-scale dataset consisting of over 3 million samples, including 1.8 million training samples and nearly 1.2 million testing samples across 12 languages. The language mismatch effects are preliminarily quantified and remarkably reduced over 15% by applying the proposed ACCENT. This easily implementable method shows promise for multilingual and low-resource language scenarios. Tianchi Liu 0004, Ivan Kukanov, Zihan Pan, Qiongqiong Wang, Hardik B. Sailor, Kong-Aik Lee |
SLT | 6 |
| 2024 | Room Impulse Responses Help Attackers to Evade Deep Fake DetectionabstractThe ASVspoof 2021 benchmark, a widely-used evaluation framework for anti-spoofing, consists of two subsets: Logical Access (LA) and Deepfake (DF), featuring samples with varied coding characteristics and compression artifacts. Notably, the current state-of-the-art (SOTA) system boasts impressive performance, achieving an Equal Error Rate (EER) of 0.87% on the LA subset and 2.58% on the DF. However, benchmark accuracy is no guarantee of robustness in real-world scenarios. This paper investigates the effectiveness of utilizing room impulse responses (RIRs) to enhance fake speech and increase their likelihood of evading fake speech detection systems. Our findings reveal that this simple approach significantly improves the evasion rate, doubling the SOTA system’s EER. To counter this type of attack, We augmented training data with a large-scale synthetic/simulated RIR dataset. The results demonstrate significant improvement on both reverberated fake speech and original samples, reducing DF task EER to 2.13%. Hieu-Thi Luong, Duc-Tuan Truong, Kong-Aik Lee, Chng Eng Siong |
SLT | 3 |
| 2024 | t-EER: Parameter-Free Tandem Evaluation of Countermeasures and Biometric ComparatorsabstractPresentation attack (spoofing) detection (PAD) typically operates alongside biometric verification to improve reliablity in the face of spoofing attacks. Even though the two sub-systems operate in tandem to solve the single task of reliable biometric verification, they address different detection tasks and are hence typically evaluated separately. Evidence shows that this approach is suboptimal. We introduce a new metric for the joint evaluation of PAD solutions operating in situ with biometric verification. In contrast to the tandem detection cost function proposed recently, the new tandem equal error rate (t-EER) is parameter free. The combination of two classifiers nonetheless leads to a set of operating points at which false alarm and miss rates are equal and also dependent upon the prevalence of attacks. We therefore introduce the concurrent t-EER, a unique operating point which is invariable to the prevalence of attacks. Using both modality (and even application) agnostic simulated scores, as well as real scores for a voice biometrics application, we demonstrate application of the t-EER to a wide range of biometric system evaluations under attack. The proposed approach is a strong candidate metric for the tandem evaluation of PAD systems and biometric comparators. Tomi Kinnunen, Kong-Aik Lee, Hemlata Tak, Nicholas W. D. Evans, Andreas Nautsch |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Cosine Scoring With Uncertainty for Neural Speaker EmbeddingabstractUncertainty modeling in speaker representation aims to learn the variability present in speech utterances. While the conventional cosine-scoring is computationally efficient and prevalent in speaker recognition, it lacks the capability to handle uncertainty. To address this challenge, this paper proposes an approach for estimating uncertainty at the speaker embedding front-end and propagating it to the cosine scoring back-end. Experiments conducted on the VoxCeleb and SITW datasets confirmed the efficacy of the proposed method in handling uncertainty arising from embedding estimation. It achieved improvement with 8.5% and 9.8% average reductions in EER and minDCF compared to the conventional cosine similarity. It is also computationally efficient in practice. Qiongqiong Wang, Kong-Aik Lee |
IEEE Signal Process. Lett. | 2 |
| 2024 | Golden Gemini is All You Need: Finding the Sweet Spots for Speaker VerificationabstractThe residual neural networks (ResNet) demonstrate the impressive performance in automatic speaker verification (ASV). They treat the time and frequency dimensions equally, following the default stride configuration designed for image recognition, where the horizontal and vertical axes exhibit similarities. This approach ignores the fact that time and frequency are asymmetric in speech representation. We address this issue and postulateGolden-Gemini Hypothesis,which posits the prioritization of temporal resolution over frequency resolution for ASV. The hypothesis is verified by conducting a systematic study on the impact of temporal and frequency resolutions on the performance, using a trellis diagram to represent the stride space. We further identify two optimal points, namelyGolden Gemini, which serves as a guiding principle for designing 2D ResNet-based ASV models. By following the principle, a state-of-the-art ResNet baseline model gains a significant performance improvement on VoxCeleb, SITW, and CNCeleb datasets with 7.70%/11.76% average EER/minDCF reductions, respectively, across different network depths (ResNet18, 34, 50, and 101), while reducing the number of parameters by 16.5% and FLOPs by 4.1%. We refer to it asGeminiResNet. Further investigation reveals the efficacy of the proposedGolden Geminioperating points across various training conditions and architectures. Furthermore, we present a new benchmark, namely theGeminiDF-ResNet, using a cutting-edge model.Codes and pre-trained models are available athttps://github.com/Tianchi-Liu9/Golden-Gemini-for-Speaker-Verification. Tianchi Liu 0004, Kong-Aik Lee, Qiongqiong Wang, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | Generalizing Speaker Verification for Spoof Awareness in the Embedding SpaceabstractIt is now well-known thatautomatic speaker verification(ASV) systems can be spoofed using various types of adversaries. The usual approach to counteract ASV systems against such attacks is to develop a separate spoofingcountermeasure(CM) module to classify speech input either as a bonafide, or a spoofed utterance. Nevertheless, such a design requires additional computation and utilization efforts at the authentication stage. An alternative strategy involves a single monolithic ASV system designed to handle both zero-effort imposter (non-targets) and spoofing attacks. Suchspoof-awareASV systems have the potential to provide stronger protections and more economic computations. To this end, we propose to generalize the standalone ASV (G-SASV) against spoofing attacks, where we leverage limited training data from CM to enhance a simple backend in the embedding space, without the involvement of a separate CM module during the test (authentication) phase. We propose a novel yet simple backend classifier based on deep neural networks and conduct the study via domain adaptation and multi-task integration of spoof embeddings at the training stage. Experiments are conducted on the ASVspoof 2019 logical access dataset, where we improve the performance of statistical ASV backends on the joint (bonafide and spoofed) and spoofed conditions by a maximum of 36.2% and 49.8% in terms of equal error rates, respectively. Xuechen Liu 0001, Md. Sahidullah, Kong-Aik Lee, Tomi Kinnunen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | Overview of Speaker Modeling and Its Applications: From the Lens of Deep Speaker Representation LearningabstractSpeaker individuality information is among the most critical elements within speech signals. By thoroughly and accurately modeling this information, it can be utilized in various intelligent speech applications, such as speaker recognition, speaker diarization, speech synthesis, and target speaker extraction. In this overview, we present a comprehensive review of neural approaches to speaker representation learning from both theoretical and practical perspectives. Theoretically, we discuss speaker encoders ranging from supervised to self-supervised learning algorithms, standalone models to large pretrained models, pure speaker embedding learning to joint optimization with downstream tasks, and efforts toward interpretability. Practically, we systematically examine approaches for robustness and effectiveness, introduce and compare various open-source toolkits in the field. Through the systematic and comprehensive review of the relevant literature, research activities, and resources, we provide a clear reference for researchers in the speaker characterization and modeling field, as well as for those who wish to apply speaker modeling techniques to specific downstream tasks. Shuai Wang 0016, Zhengyang Chen, Kong-Aik Lee, Yanmin Qian, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | The Second Multi-Channel Multi-Party Meeting Transcription Challenge (M2MeT 2.0): A Benchmark for Speaker-Attributed ASRabstractWith the success of the first Multi-channel Multi-party Meeting Transcription challenge (M2MeT), the second M2MeT challenge (M2MeT 2.0) held in ASRU2023 particularly aims to tackle the complex task of speaker-attributed ASR (SAASR), which directly addresses the practical and challenging problem of “who spoke what at when” at typical meeting scenario. We particularly established two sub-tracks. The fixed training condition sub-track, where the training data is constrained to predetermined datasets, but participants can use any open-source pre-trained model. The open training condition sub-track, which allows for the use of all available data and models without limitation. In addition, we release a new 10-hour test set for challenge ranking. This paper provides an overview of the dataset, track settings, results, and analysis of submitted systems, as a benchmark to show the current state of speaker-attributed ASR. Yuhao Liang, Mohan Shi, Fan Yu 0002, Yangze Li, Shiliang Zhang, Zhihao Du, Qian Chen 0003, Lei Xie 0001, Yanmin Qian, Jian Wu 0027, Zhuo Chen 0006, Kong-Aik Lee, Zhijie Yan, Hui Bu |
ASRU | 12 |
| 2023 | Self-Supervised Audio-Visual Speaker Representation with Co-Meta LearningabstractIn self-supervised speaker verification, the quality of pseudo labels determines the upper bound of its performance and it is not uncommon to end up with massive amount of unreliable pseudo labels. We observe that the complementary information in different modalities ensures a robust supervisory signal for audio and visual representation learning. This motivates us to propose an audio-visual self-supervised learning framework named Co-Meta Learning. Inspired by the Coteaching+, we design a strategy that allows the information of two modalities to be coordinated through the Update by Disagreement. Moreover, we use the idea of modelagnostic meta learning (MAML) to update the network parameters, which makes the hard samples of two modalities to be better resolved by the other modality through gradient regularization. Compared to the baseline, our proposed method achieves a 29.8%, 11.7% and 12.9% relative improvement on Vox-O, Vox-E and Vox-H trials of Voxceleb1 evaluation dataset respectively. Hanyi Zhang, Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001 |
ICASSP | 4 |
| 2023 | Leveraging Positional-Related Local-Global Dependency for Synthetic Speech DetectionabstractAutomatic speaker verification (ASV) systems are vulnerable to spoofing attacks. As synthetic speech exhibits local and global artifacts compared to natural speech, incorporating local-global dependency would lead to better anti-spoofing performance. To this end, we propose the Rawformer that leverages positional-related local-global dependency for synthetic speech detection. The two-dimensional convolution and Transformer are used in our method to capture local and global dependency, respectively. Specifically, we design a novel positional aggregator that integrates local-global dependency by adding positional information and flattening strategy with less information loss. Furthermore, we propose the squeeze-and-excitation Rawformer (SE-Rawformer), which introduces squeeze-and-excitation operation to acquire local dependency better. The results demonstrate that our proposed SE-Rawformer leads to 37% relative improvement compared to the single state-of-the-art system on ASVspoof 2019 LA and generalizes well on ASVspoof 2021 LA. Especially, using the positional aggregator in the SE-Rawformer brings a 43% improvement on average. Meng Liu 0017, Longbiao Wang, Kong-Aik Lee, Hanyi Zhang, Jianwu Dang 0001 |
ICASSP | 4 |
| 2023 | Cross-Modal Audio-Visual Co-Learning for Text-Independent Speaker VerificationabstractVisual speech (i.e., lip motion) is highly related to auditory speech due to the co-occurrence and synchronization in speech production. This paper investigates this correlation and proposes a cross-modal speech co-learning paradigm. The primary motivation of our cross-modal co-learning method is modeling one modality aided by exploiting knowledge from another modality. Specifically, two cross-modal boosters are introduced based on an audio-visual pseudo-siamese structure to learn the modality-transformed correlation. Inside each booster, a max-feature-map embedded Transformer variant is proposed for modality alignment and enhanced feature generation. The network is co-learned both from scratch and with pretrained models. Experimental results on the test scenarios demonstrate that our proposed method achieves around 60% and 20% average relative performance improvement over baseline unimodal and fusion systems, respectively. Meng Liu 0017, Kong-Aik Lee, Longbiao Wang, Hanyi Zhang, Chang Zeng, Jianwu Dang 0001 |
ICASSP | 2 |
| 2023 | Probabilistic Back-ends for Online Speaker Recognition and ClusteringabstractThis paper focuses on multi-enrollment speaker recognition which naturally occurs in the task of online speaker clustering, and studies the properties of different scoring back-ends in this scenario. First, we show that popular cosine scoring suffers from poor score calibration with a varying number of enrollment utterances. Second, we propose a simple replacement for cosine scoring based on an extremely constrained version of probabilistic linear discriminant analysis (PLDA). The proposed model improves over the cosine scoring for multi-enrollment recognition while keeping the same performance in the case of one-to-one comparisons. Finally, we consider an online speaker clustering task where each step naturally involves multi-enrollment recognition. We propose an online clustering algorithm allowing us to take benefits from the PLDA model such as the ability to handle uncertainty and better score calibration. Our experiments demonstrate the effectiveness of the proposed algorithm. Alexey Sholokhov, Nikita Kuzmin, Kong-Aik Lee, Chng Eng Siong |
ICASSP | 3 |
| 2023 | Noise-Disentanglement Metric Learning for Robust Speaker VerificationabstractAutomatic speaker verification (ASV) suffers from performance degradation in noisy environments. To solve this problem, we propose the noise-disentanglement metric learning to reduce the speaker-irrelevant noisy components and build a noise-invariant embedding space. Specifically, the disentanglement module, including the speaker encoder and re-construction module, is dedicated to decoupling speech signals. The speaker encoder is used to disentangle speaker-related components, and the reconstruction module increases the model’s ability to constrain the noise information by re-constructing the signal. In addition, distribution optimization is introduced to supervise the spatial structure of speaker embeddings under noisy environments. Experiments on Vox-Celeb1 indicate that the proposed method improves the performance of the speaker verification system in both clean and noisy conditions. Hanyi Zhang, Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001 |
ICASSP | 4 |
| 2023 | Speaker Recognition with Two-Step Multi-Modal Deep CleansingabstractNeural network-based speaker recognition has achieved significant improvement in recent years. A robust speaker representation learns meaningful knowledge from both hard and easy samples in the training set to achieve good performance. However, noisy samples (i.e., with wrong labels) in the training set induce confusion and cause the network to learn the incorrect representation. In this paper, we propose a two-step audio-visual deep cleansing framework to eliminate the effect of noisy labels in speaker representation learning. This framework contains a coarse-grained cleansing step to search for the complex samples, followed by a fine-grained cleansing step to filter out the noisy labels. Our study starts from an efficient audio-visual speaker recognition system, which achieves a close to perfect equal-error-rate (EER) of 0.01%, 0.07% and 0.13% on the Vox-O, E and H test sets. With the proposed multi-modal cleansing mechanism, four different speaker recognition networks achieve an average improvement of 5.9%. Code has been made available at: https://github.com/TaoRuijie/AVCleanse. Ruijie Tao, Kong-Aik Lee, Haizhou Li 0001 |
ICASSP | 2 |
| 2023 | Incorporating Uncertainty from Speaker Embedding Estimation to Speaker VerificationabstractSpeech utterances recorded under differing conditions exhibit varying degrees of confidence in their embedding estimates, i.e., uncertainty, even if they are extracted using the same neural network. This paper aims to incorporate the uncertainty estimate produced in the xi-vector network front-end with a probabilistic linear discriminant analysis (PLDA) back-end scoring for speaker verification. To achieve this we derive a posterior covariance matrix, which measures the uncertainty, from the frame-wise precisions to the embedding space. We propose a log-likelihood ratio function for the PLDA scoring with the uncertainty propagation. We also propose to replace the length normalization pre-processing technique with a length scaling technique for the application of uncertainty propagation in the back-end. Experimental results on the VoxCeleb-1, SITW test sets as well as a domain-mismatched CNCeleb1-E set show the effectiveness of the proposed techniques with 14.5%–41.3% EER reductions and 4.6%–25.3% minDCF reductions. Qiongqiong Wang, Kong-Aik Lee, Tianchi Liu 0004 |
ICASSP | 2 |
| 2023 | Speaker-Aware Anti-spoofing
Xuechen Liu 0001, Md. Sahidullah, Kong-Aik Lee, Tomi Kinnunen |
INTERSPEECH | 3 |
| 2023 | Towards Single Integrated Spoofing-aware Speaker Verification Embeddings
Sung Hwan Mun, Hye-Jin Shim, Hemlata Tak, Xin Wang 0037, Xuechen Liu 0001, Md. Sahidullah, Myeonghun Jeong, Min Hyun Han, Massimiliano Todisco, Kong-Aik Lee, Junichi Yamagishi, Nicholas W. D. Evans, Tomi Kinnunen, Nam Soo Kim, Jee-Weon Jung |
INTERSPEECH | 10 |
| 2023 | Disentangling Voice and Content with Self-Supervision for Speaker RecognitionabstractFor speaker recognition, it is difficult to extract an accurate speaker representation from speech because of its mixture of speaker traits and content. This paper proposes a disentanglement framework that simultaneously models speaker traits and content variability in speech. It is realized with the use of three Gaussian inference layers, each consisting of a learnable transition model that extracts distinct speech components. Notably, a strengthened transition model is specifically designed to model complex speech dynamics. We also propose a self-supervision method to dynamically disentangle content without the use of labels other than speaker identities. The efficacy of the proposed framework is validated via experiments conducted on the VoxCeleb and SITW datasets with 9.56\% and 8.24\% average reductions in EER and minDCF, respectively. Since neither additional model training nor data is specifically needed, it is easily applicable in practical use. Tianchi Liu 0004, Kong-Aik Lee, Qiongqiong Wang, Haizhou Li 0001 |
NeurIPS | 2 |
| 2023 | Partially Randomizing Transformer Weights for Dialogue Response Diversity
Jing Yang Lee, Kong-Aik Lee, Woon-Seng Gan |
PACLIC | 2 |
| 2023 | ASVspoof 2021: Towards Spoofed and Deepfake Speech Detection in the WildabstractBenchmarking initiatives support the meaningful comparison of competing solutions to prominent problems in speech and language processing. Successive benchmarking evaluations typically reflect a progressive evolution from ideal lab conditions towards to those encountered in the wild. ASVspoof, the spoofing and deepfake detection initiative and challenge series, has followed the same trend. This article provides a summary of the ASVspoof 2021 challenge and the results of 54 participating teams that submitted to the evaluation phase. For the logical access (LA) task, results indicate that countermeasures are robust to newly introduced encoding and transmission effects. Results for the physical access (PA) task indicate the potential to detect replay attacks in real, as opposed to simulated physical spaces, but a lack of robustness to variations between simulated and real acoustic environments. The Deepfake (DF) task, new to the 2021 edition, targets solutions to the detection of manipulated, compressed speech data posted online. While detection solutions offer some resilience to compression effects, they lack generalization across different source datasets. In addition to a summary of the top-performing systems for each task, new analyses of influential data factors and results for hidden data subsets, the article includes a review of post-challenge results, an outline of the principal challenge limitations and a road-map for the future of ASVspoof. Xuechen Liu 0001, Xin Wang 0037, Md. Sahidullah, Jose Patino 0001, Héctor Delgado, Tomi Kinnunen, Massimiliano Todisco, Junichi Yamagishi, Nicholas W. D. Evans, Andreas Nautsch, Kong-Aik Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 11 |
| 2023 | Self-Supervised Training of Speaker Encoder With Multi-Modal Diverse Positive PairsabstractWe study a novel neural speaker encoder and its training strategies for speaker recognition without using any identity labels. The speaker encoder is trained to extract a fixed dimensional speaker embedding from a spoken utterance of variable length. Contrastive learning is a typical self-supervised learning technique. However, the contrastive learning of the speaker encoder depends very much on the sampling strategy of positive and negative pairs. It is common that we sample a positive pair of segments from the same utterance. Unfortunately, such a strategy, denoted as poor-man's positive pairs (PPP), lacks the necessary diversity. In this work, we propose a multi-modal contrastive learning technique with novel sampling strategies. By cross-referencing between speech and face data, we find diverse positive pairs (DPP) for contrastive learning, thus improving the robustness of speaker encoder. We train the speaker encoder on the VoxCeleb2 dataset without any speaker labels, and achieve an equal error rate (EER) of 2.89%, 3.17% and 6.27% under the proposed progressive clustering strategy, and an EER of 1.44%, 1.77% and 3.27% under the two-stage learning strategy with pseudo labels, on the three test sets of VoxCeleb1. This novel solution outperforms the state-of-the-art self-supervised learning methods by a large margin, at the same time, achieves comparable results with the supervised learning counterpart. We also evaluate our self-supervised learning technique on the LRS2 and LRW datasets, where speaker information is unavailable. All experiments suggest that the proposed neural architecture and sampling strategies are robust across datasets. Ruijie Tao, Kong-Aik Lee, Rohan Kumar Das, Ville Hautamäki, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Meta-Generalization for Domain-Invariant Speaker VerificationabstractAutomatic speaker verification (ASV) exhibits unsatisfactory performance under domain mismatch conditions owing to intrinsic and extrinsic factors, such as variations in speaking styles and recording devices encountered in real-world applications. To ensure robust performance under unseen conditions, domain generalization has been explored. However, an inherent contradiction exists between model discrimination and domain generalization, in which the discrimination ability may be reduced while learning to generalize. In this paper, to extract discriminative yet domain-invariant representations, we propose the meta-generalized speaker verification (MGSV) via meta-learning. Specifically, we propose a metric-based distribution optimization and a gradient-based meta-optimization to simultaneously supervise the spatial relationship between embeddings and improve the generalization ability of the model on unseen domains. In addition, we design multiple-single (MS) and simulated speaker verification (SSV) sampling strategies based on single-domain (SD) and single-single (SS) strategies to simulate the train/test domain mismatch more relevantly, thereby mining transferable speaker-related knowledge. SSV is chosen as the most effective method, as it substantially improves the domain generalization by ensuring that the model has learned to discriminate efficiently. Additionally, to intuitively reflect the model performance on the unseen domains, the proposed method is validated on cross-genre, cross-device, and cross-dataset tasks. The experimental results demonstrate that our proposed method achieves remarkable performance in handling domain mismatch issues in speaker verification. Hanyi Zhang, Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001, Helen M. Meng |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Generalized Domain Adaptation Framework for Parametric Back-End in Speaker RecognitionabstractState-of-the-art speaker recognition systems comprise a speaker embedding front-end followed by a probabilistic linear discriminant analysis (PLDA) back-end. The effectiveness of these components relies on the availability of a large amount of labeled training data. In practice, it is common for domains (e.g., language, channel, demographic) in which a system is deployed to differ from that in which a system has been trained. To close the resulting gap, domain adaptation is often essential for PLDA models. Among two of its variants are Heavy-tailed PLDA (HT-PLDA) and Gaussian PLDA (G-PLDA). Though the former better fits real feature spaces than does the latter, its popularity has been severely limited by its computational complexity and, especially, by the difficulty, it presents in domain adaptation, which results from its non-Gaussian property. Various domain adaptation methods have been proposed for G-PLDA. This paper proposes a generalized framework for domain adaptation that can be applied to both of the above variants of PLDA for speaker recognition. It not only includes several existing supervised and unsupervised domain adaptation methods but also makes possible more flexible usage of available data in different domains. In particular, we introduce here two new techniques: (1) correlation-alignment in the model level, and (2) covariance regularization. To the best of our knowledge, this is the first proposed application of such techniques for domain adaptation w.r.t. HT-PLDA. The efficacy of the proposed techniques has been experimentally validated on NIST 2016, 2018, and 2019 Speaker Recognition Evaluation (SRE’16, SRE’18 and SRE’19) datasets. Qiongqiong Wang, Koji Okabe, Kong-Aik Lee, Takafumi Koshinaka |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | DLVGen: A Dual Latent Variable Approach to Personalized Dialogue GenerationabstractThe generation of personalized dialogue is vital to natural and human-like conversation. Typically, personalized dialogue generation models involve conditioning the generated response on the dialogue history and a representation of the persona/personality of the interlocutor. As it is impractical to obtain the persona/personality representations for every interlocutor, recent works have explored the possibility of generating personalized dialogue by finetuning the model with dialogue examples corresponding to a given persona instead. However, in real-world implementations, a sufficient number of corresponding dialogue examples are also rarely available. Hence, in this paper, we propose a Dual Latent Variable Generator (DLVGen) capable of generating personalized dialogue in the absence of any persona/personality information or any corresponding dialogue examples. Unlike prior work, DLVGen models the latent distribution over potential responses as well as the latent distribution over the agent's potential persona. During inference, latent variables are sampled from both distributions and fed into the decoder. Empirical results show that DLVGen is capable of generating diverse responses which accurately incorporate the agent's persona. Jing Yang Lee, Kong-Aik Lee, Woon-Seng Gan |
ICAART (2) | 2 |
| 2022 | Improving Contextual Coherence in Variational Personalized and Empathetic Dialogue AgentsabstractIn recent years, latent variable models, such as the Conditional Variational Auto Encoder (CVAE), have been applied to both personalized and empathetic dialogue generation. Prior work have largely focused on generating diverse dialogue responses that exhibit persona consistency and empathy. However, when it comes to the contextual coherence of the generated responses, there is still room for improvement. Hence, to improve the contextual coherence, we propose a novel Uncertainty Aware CVAE (UA-CVAE) framework. The UA-CVAE framework involves approximating and incorporating the aleatoric uncertainty during response generation. We apply our framework to both personalized and empathetic dialogue generation. Empirical results show that our framework significantly improves the contextual coherence of the generated response. Additionally, we introduce a novel automatic metric for measuring contextual coherence, which was found to correlate positively with human judgement. Jing Yang Lee, Kong-Aik Lee, Woon-Seng Gan |
ICASSP | 2 |
| 2022 | MFA: TDNN with Multi-Scale Frequency-Channel Attention for Text-Independent Speaker Verification with Short UtterancesabstractThe time delay neural network (TDNN) represents one of the state-of-the-art of neural solutions to text-independent speaker verification. However, they require a large number of filters to capture the speaker characteristics at any local frequency region. In addition, the performance of such systems may degrade under short utterance scenarios. To address these issues, we propose a multi-scale frequency-channel attention (MFA), where we characterize speakers at different scales through a novel dual-path design which consists of a convolutional neural network and TDNN. We evaluate the proposed MFA on the VoxCeleb database and observe that the proposed framework with MFA can achieve state-of-the-art performance while reducing parameters and computation complexity. Further, the MFA mechanism is found to be effective for speaker verification with short test utterances. Tianchi Liu 0004, Rohan Kumar Das, Kong-Aik Lee, Haizhou Li 0001 |
ICASSP | 3 |
| 2022 | Self-Supervised Speaker Recognition with Loss-Gated LearningabstractIn self-supervised learning for speaker recognition, pseudo labels are useful as the supervision signals. It is a known fact that a speaker recognition model doesn’t always benefit from pseudo labels due to their unreliability. In this work, we observe that a speaker recognition network tends to model the data with reliable labels faster than those with unreliable labels. This motivates us to study a loss-gated learning (LGL) strategy, which extracts the reliable labels through the fitting ability of the neural network during training. With the proposed LGL, our speaker recognition model obtains a 46.3% performance gain over the system without it. Further, the proposed self-supervised speaker recognition with LGL trained on the VoxCeleb2 dataset without any labels achieves an equal error rate of 1.66% on the VoxCeleb1 original test set. Ruijie Tao, Kong-Aik Lee, Rohan Kumar Das, Ville Hautamäki, Haizhou Li 0001 |
ICASSP | 2 |
| 2022 | Summary on the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand ChallengeabstractThe ICASSP 2022 Multi-channel Multi-party Meeting Transcription Grand Challenge (M2MeT) focuses on one of the most valuable and the most challenging scenarios of speech technologies. The M2MeT challenge has particularly set up two tracks, speaker diarization (track 1) and multi-speaker automatic speech recognition (ASR) (track 2). Along with the challenge, we released 120 hours of real-recorded Mandarin meeting speech data with manual annotation, including far-field data collected by 8-channel micro-phone array as well as near-field data collected by each participants’ headset microphone. We briefly describe the released dataset, track setups, baselines and summarize the challenge results and major techniques used in the submissions. Fan Yu 0002, Shiliang Zhang, Yihui Fu, Zhihao Du, Weilong Huang, Lei Xie 0001, Zheng-Hua Tan, DeLiang Wang, Yanmin Qian, Kong-Aik Lee, Zhijie Yan, Bin Ma 0001, Hui Bu |
ICASSP | 12 |
| 2022 | Learning Domain-Invariant Transformation for Speaker VerificationabstractAutomatic speaker verification (ASV) faces domain shift caused by the mismatch of intrinsic and extrinsic factors such as recording device and speaking style in real-world applications, which leads to unsatisfactory performance. To this end, we propose the meta generalized transformation via meta-learning to build a domain-invariant embedding space. Specifically, the transformation module is motivated to learn the domain generalization knowledge by executing meta-optimization on the meta-train and meta-test sets which are designed to simulate domain shift. Furthermore, distribution optimization is incorporated to supervise the metric structure of embeddings. In terms of the transformation module, we investigate various instantiations and observe the multilayer perceptron with gating (gMLP) is the most effective given its extrapolation capability. The experimental results on cross-genre and cross-dataset settings demonstrate that the meta generalized transformation dramatically improves the robustness of ASV systems to domain shift, while outperforms the state-of-the-art methods. Hanyi Zhang, Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001 |
ICASSP | 3 |
| 2022 | Scoring of Large-Margin Embeddings for Speaker Verification: Cosine or PLDA?abstractThe emergence of large-margin softmax cross-entropy losses in training deep speaker embedding neural networks has triggered a gradual shift from parametric back-ends to a simpler cosine similarity measure for speaker verification. Popular parametric back-ends include the probabilistic linear discriminant analysis (PLDA) and its variants. This paper investigates the properties of margin-based cross-entropy losses leading to such a shift and aims to find scoring back-ends best suited for speaker verification. In addition, we revisit the pre-processing techniques which have been widely used in the past and assess their effectiveness on large-margin embeddings. Experiments on the state-of-the-art ECAPA-TDNN networks trained with various large-margin softmax cross-entropy losses show a substantial increment in intra-speaker compactness making the conventional PLDA superfluous. In this regard, we found that constraining the within-speaker covariance matrix could improve the performance of the PLDA. It is demonstrated through a series of experiments on the VoxCeleb-1 and SITW core-core test sets with 40.8% equal error rate (EER) reduction and 35.1% minimum detection cost (minDCF) reduction. It also outperforms cosine scoring consistently with reductions in EER and minDCF by 10.9% and 4.9%, respectively. Qiongqiong Wang, Kong-Aik Lee, Tianchi Liu 0004 |
INTERSPEECH | 2 |
| 2022 | Noise-Robust Semi-supervised Multi-modal Machine Translation
Lin Li 0001, Kaixi Hu, Turghun Tayir, Jianquan Liu, Kong-Aik Lee |
PRICAI (2) | 5 |
| 2022 | Discriminative speaker embedding with serialized multi-layer multi-head attention
Hongning Zhu, Kong-Aik Lee, Haizhou Li 0001 |
Speech Commun. | 2 |
| 2022 | Neural Acoustic-Phonetic Approach for Speaker Verification With Phonetic Attention MaskabstractTraditional acoustic-phonetic approach makes use of both spectral and phonetic information when comparing the voice of speakers. While phonetic units are not equally informative, the phonetic context of speech plays an important role in speaker verification (SV). In this paper, we propose a neural acoustic-phonetic approach that learns to dynamically assign differentiated weights to spectral features for SV. Such differentiated weights form a phonetic attention mask (PAM). The neural acoustic-phonetic framework consists of two training pipelines, one for SV and another for speech recognition. Through the PAM, we leverage the phonetic information for SV. We evaluate the proposed neural acoustic-phonetic framework on the RSR2015 database Part III corpus, that consists of random digit strings. We show that the proposed framework with PAM consistently outperforms baseline with an equal error rate reduction of 13.45% and 10.20% for female and male data, respectively. Tianchi Liu 0004, Rohan Kumar Das, Kong-Aik Lee, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 3 |
| 2021 | DeepLip: A Benchmark for Deep Learning-Based Audio-Visual Lip BiometricsabstractAudio-visual lip biometrics (AV-LB) has been an emerging biometrics technology that straddles auditory and visual speech processing. Previous works mainly focused on the front-end lip-based feature engineering combined with a shallow statistical back-end model. Over the past decade, convolutional neural network (CNN, or ConvNet) has been widely used and achieved good performance in computer vision and speech processing tasks. However, the lack of a sizeable public AV-LB database led to a stagnation in deep-learning exploration on AV-LB tasks. In addition to the dual audio-visual streams, one essential requirement on the video stream is the region of interest (ROI) around the lips has to be of sufficient resolution. To this end, we compile a moderate-size database using existing public databases. Using this database, we present a deep learning-based AV-LB benchmark, dubbed DeepLip11https://github.com/DanielMengLiu/DeepLip, realized with convolutional video and audio unimodal modules, and a multimodal fusion module. Our experiments show that DeepLip outperforms the traditional lip-biometrics system in context modeling and achieves over 50% relative improvements compared with its unimodal system, with an equal error rate of 0.75% and 1.11% on the test datasets, respectively. Meng Liu 0017, Longbiao Wang, Kong-Aik Lee, Hanyi Zhang, Chang Zeng, Jianwu Dang 0001 |
ASRU | 3 |
| 2021 | PL-EESR: Perceptual Loss Based End-to-End Robust Speaker Representation ExtractionabstractSpeech enhancement aims to improve the perceptual quality of the speech signal by suppression of the background noise. However, excessive suppression may lead to speech distortion and speaker information loss, which degrades the performance of speaker embedding extraction. To alleviate this problem, we propose an end-to-end deep learning framework, dubbed PL-EESR, for robust speaker representation extraction. This framework is optimized based on the feedback of the speaker identification task and the high-level perceptual deviation between the raw speech signal and its noisy version. We conducted speaker verification tasks in both noisy and clean environment respectively to evaluate our system. Compared to the baseline, our method shows better performance in both clean and noisy environments, which means our method can not only enhance the speaker relative information but also avoid adding distortions. Kong-Aik Lee, Ville Hautamäki, Haizhou Li 0001 |
ASRU | 2 |
| 2021 | COOPNet: Multi-Modal Cooperative Gender Prediction in Social Media User ProfilingabstractThe principal way of performing user profiling is to investigate accumulated social media data. However, the problem of information asymmetry generally exists in user generated contents since users post multi-modal contents in social media freely. In this paper, we propose a novel text-image cooperation framework (COOPNet), a bridge connection network architecture that exchanges information between texts and images. First, we map the representations of both visual and sentiment enriched textual modalities into a cooperative semantic space to derive a cooperative representation. Next, the representations of texts and images are combined with their cooperative representation to exchange knowledge in the learning process. Finally, a multi-modal regression is leveraged to make cooperative decisions. Extensive experiments on the public PAN-2018 dataset demonstrate the efficacy of our framework over the state-of-the-art methods on the premise of automatic feature learning. Lin Li 0001, Kaixi Hu, Yunpei Zheng, Jianquan Liu, Kong-Aik Lee |
ICASSP | 5 |
| 2021 | Replay-Attack Detection Using Features With Adaptive Spectro-Temporal ResolutionabstractVariable-resolution processing aims to improve the feature representation ability by enlarging the local discriminative details. In previous anti-spoofing studies, different phones and frequency regions were both proven to have various levels of sensitivity to replay distortion. In this paper, an adaptive spectro-temporal resolution is proposed to obtain the optimal scale in the feature space: the frequency resolution is adaptive to frequency discrimination, while the temporal resolution is adaptive to continuous phones. In the process, phone-frequency F-ratio analysis is applied to investigate the sensitivity divergences to replay distortion among phones and frequencies. Then, attentive filters are designed to automatically adapt to the phone-frequency discrimination. Validation experiments for the proposed method are conducted on two well-acknowledged magnitude and phase features. A comparative analysis on the ASVspoof 2017 V2.0 database demonstrates that our proposed adaptive spectro-temporal resolution method attains considerably higher error reduction rates than the approaches involving the corresponding original resolution features. Meng Liu 0017, Longbiao Wang, Kong-Aik Lee, Xuanda Chen, Jianwu Dang 0001 |
ICASSP | 3 |
| 2021 | Meta-Learning for Cross-Channel Speaker VerificationabstractAutomatic speaker verification (ASV) has been successfully deployed for identity recognition. With increasing use of ASV technology in real-world applications, channel mismatch caused by the recording devices and environments severely degrade its performance, especially in the case of unseen channels. To this end, we propose a meta speaker embedding network (MSEN) via meta-learning to generate channel-invariant utterance embeddings. Specifically, we optimize the differences between the embeddings of a support set and a query set in order to learn a channel-invariant embedding space for utterances. Furthermore, we incorporate distribution optimization (DO) to stabilize the performance of MSEN. To quantitatively measure the effect of MSEN on unseen channels, we specially design the generalized cross-channel (GCC) evaluation. The experimental results on the HI-MIA corpus demonstrate that the proposed MSEN reduce considerably the impact of channel mismatch, while significantly outperforms other state-of-the-art methods. Hanyi Zhang, Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001 |
ICASSP | 3 |
| 2021 | Visualizing Classifier Adjacency Relations: A Case Study in Speaker Verification and Voice Anti-SpoofingabstractWhether it be for results summarization, or the analysis of classifier fusion, some means to compare different classifiers can often provide illuminating insight into their behaviour, (dis)similarity or complementarity. We propose a simple method to derive 2D representation from detection scores produced by an arbitrary set of binary classifiers in response to a common dataset. Based upon rank correlations, our method facilitates a visual comparison of classifiers with arbitrary scores and with close relation to receiver operating characteristic (ROC) and detection error trade-off (DET) analyses. While the approach is fully versatile and can be applied to any detection task, we demonstrate the method using scores produced by automatic speaker verification and voice anti-spoofing systems. The former are produced by a Gaussian mixture model system trained with VoxCeleb data whereas the latter stem from submissions to the ASVspoof 2019 challenge. Tomi Kinnunen, Andreas Nautsch, Md. Sahidullah, Nicholas W. D. Evans, Xin Wang 0037, Massimiliano Todisco, Héctor Delgado, Junichi Yamagishi, Kong-Aik Lee |
Interspeech | 9 |
| 2021 | Joint Feature Enhancement and Speaker Recognition with Multi-Objective Task-Oriented Network
Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001 |
Interspeech | 3 |
| 2021 | Multi-Level Transfer Learning from Near-Field to Far-Field Speaker VerificationabstractIn far-field speaker verification, the performance of speaker embeddings is susceptible to degradation when there is a mismatch between the conditions of enrollment and test speech.To solve this problem, we propose the feature-level and instancelevel transfer learning in the teacher-student framework to learn a domain-invariant embedding space.For the feature-level knowledge transfer, we develop the contrastive loss to transfer knowledge from teacher model to student model, which can not only decrease the intra-class distance, but also enlarge the inter-class distance.Moreover, we propose the instance-level pairwise distance transfer method to force the student model to preserve pairwise instances distance from the well optimized embedding space of the teacher model.On FFSVC 2020 evaluation set, our EER on Full-eval trials is relatively reduced by 13.9% compared with the fusion system result on Partialeval trials of Task2.On Task1, compared with the winner's DenseNet result on Partial-eval trials, our minDCF on Full-eval trials is relatively reduced by 6.3%.On Task3, the EER and minDCF of our proposed method on Full-eval trials are very close to the result of the fusion system on Partial-eval trials.Our results also outperform other competitive domain adaptation methods. Li Zhang 0084, Qing Wang 0039, Kong-Aik Lee, Lei Xie 0001, Haizhou Li 0001 |
Interspeech | 3 |
| 2021 | Serialized Multi-Layer Multi-Head Attention for Neural Speaker EmbeddingabstractThis paper proposes a serialized multi-layer multi-head attention for neural speaker embedding in text-independent speaker verification.In prior works, frame-level features from one layer are aggregated to form an utterance-level representation.Inspired by the Transformer network, our proposed method utilizes the hierarchical architecture of stacked self-attention mechanisms to derive refined features that are more correlated with speakers.Serialized attention mechanism contains a stack of self-attention modules to create fixed-dimensional representations of speakers.Instead of utilizing multi-head attention in parallel, the proposed serialized multi-layer multi-head attention is designed to aggregate and propagate attentive statistics from one layer to the next in a serialized manner.In addition, we employ an input-aware query for each utterance with the statistics pooling.With more layers stacked, the neural network can learn more discriminative speaker embeddings.Experiment results on VoxCeleb1 dataset and SITW dataset show that our proposed method outperforms other baseline methods, including x-vectors and other x-vectors + conventional attentive pooling approaches by 9.7% in EER and 8.1% in DCF10 -2 . Hongning Zhu, Kong-Aik Lee, Haizhou Li 0001 |
Interspeech | 2 |
| 2021 | Replay attack detection using variable-frequency resolution phase and magnitude features
Meng Liu 0017, Longbiao Wang, Jianwu Dang 0001, Kong-Aik Lee, Seiichi Nakagawa |
Comput. Speech Lang. | 4 |
| 2021 | Xi-Vector Embedding for Speaker RecognitionabstractWe present a Bayesian formulation for deep speaker embedding, wherein the xi-vector is the Bayesian counterpart of the x-vector, taking into account the uncertainty estimate. On the technology front, we offer a simple and straightforward extension to the now widely used x-vector. It consists of an auxiliary neural net predicting the frame-wise uncertainty of the input sequence. We show that the proposed extension leads to substantial improvement across all operating points, with a significant reduction in error rates and detection cost. On the theoretical front, our proposal integrates the Bayesian formulation of linear Gaussian model to speaker-embedding neural networks via the pooling layer. In one sense, our proposal integrates the Bayesian formulation of the i-vector to that of the x-vector. Hence, we refer to the embedding as the xi-vector, which is pronounced as /zai/ vector. Experimental results on the SITW evaluation set show a consistent improvement of over 17.5% in equal-error-rate and 10.9% in minimum detection cost. Kong-Aik Lee, Qiongqiong Wang, Takafumi Koshinaka |
IEEE Signal Process. Lett. | 1 |
| 2020 | A Generalized Framework for Domain Adaptation of PLDA in Speaker RecognitionabstractThis paper proposes a generalized framework for domain adaptation of Probabilistic Linear Discriminant Analysis (PLDA) in speaker recognition. It not only includes several existing supervised and unsupervised domain adaptation methods but also makes possible more flexible usage of available data in different domains. In particular, we introduce here the two new techniques described below. (1) Correlation-alignment-based interpolation and (2) covariance regularization. The proposed correlation-alignment-based-interpolation method decreases minCprimaryup to 30.5% as compared with that from an out-of-domain PLDA model before adaptation, and minCprimaryis also 5.5% lower than with a conventional linear interpolation method with optimal interpolation weights. Further, the proposed regularization technique ensures robustness in interpolations w.r.t. varying interpolation weights, which in practice is essential. Qiongqiong Wang, Koji Okabe, Kong-Aik Lee, Takafumi Koshinaka |
ICASSP | 3 |
| 2020 | Deep Discriminative Embedding with Ranked Weight for Speaker Verification
Dao Zhou, Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001 |
ICONIP (5) | 3 |
| 2020 | POCO: A Voice Spoofing and Liveness Detection Corpus Based on Pop Noise
Kosuke Akimoto, Seng Pei Liew, Sakiko Mishima, Ryo Mizushima, Kong-Aik Lee |
INTERSPEECH | 5 |
| 2020 | NEC-TT Speaker Verification System for SRE'19 CTS Challenge
Kong-Aik Lee, Koji Okabe, Hitoshi Yamamoto, Qiongqiong Wang, Takafumi Koshinaka, Jiacen Zhang, Keisuke Ishikawa, Koichi Shinoda |
INTERSPEECH | 1 |
| 2020 | Extrapolating False Alarm Rates in Automatic Speaker VerificationabstractAutomatic speaker verification (ASV) vendors and corpus providers would both benefit from tools to reliably extrapolate performance metrics for large speaker populations without collecting new speakers. We address false alarm rate extrapolation under a worst-case model whereby an adversary identifies the closest impostor for a given target speaker from a large population. Our models are generative and allow sampling new speakers. The models are formulated in the ASV detection score space to facilitate analysis of arbitrary ASV systems. Alexey Sholokhov, Tomi Kinnunen, Ville Vestman, Kong-Aik Lee |
INTERSPEECH | 4 |
| 2020 | SdSV Challenge 2020: Large-Scale Evaluation of Short-Duration Speaker Verification
Hossein Zeinali, Kong-Aik Lee, Jahangir Alam 0001, Lukás Burget |
INTERSPEECH | 2 |
| 2020 | Adversarial Separation Network for Speaker Recognition
Hanyi Zhang, Longbiao Wang, Yunchun Zhang, Meng Liu 0017, Kong-Aik Lee, Jianguo Wei |
INTERSPEECH | 5 |
| 2020 | Dynamic Margin Softmax Loss for Speaker Verification
Dao Zhou, Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001, Jianguo Wei |
INTERSPEECH | 3 |
| 2020 | Two decades into Speaker Recognition Evaluation - are we there yet?
Kong-Aik Lee, Seyed Omid Sadjadi, Haizhou Li 0001, Douglas A. Reynolds |
Comput. Speech Lang. | 1 |
| 2020 | NEC-TT System for Mixed-Bandwidth and Multi-Domain Speaker Recognition
Kong-Aik Lee, Hitoshi Yamamoto, Koji Okabe, Qiongqiong Wang, Takafumi Koshinaka, Jiacen Zhang, Koichi Shinoda |
Comput. Speech Lang. | 1 |
| 2020 | Voice biometrics security: Extrapolating false alarm rate via hierarchical Bayesian modeling of speaker verification scores
Alexey Sholokhov, Tomi Kinnunen, Ville Vestman, Kong-Aik Lee |
Comput. Speech Lang. | 4 |
| 2020 | ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech
Xin Wang 0037, Junichi Yamagishi, Massimiliano Todisco, Héctor Delgado, Andreas Nautsch, Nicholas W. D. Evans, Md. Sahidullah, Ville Vestman, Tomi Kinnunen, Kong-Aik Lee, Lauri Juvela, Paavo Alku, Yu-Huai Peng, Hsin-Te Hwang, Yu Tsao 0001, Hsin-Min Wang, Sébastien Le Maguer, Zhen-Hua Ling |
Comput. Speech Lang. | 10 |
| 2020 | Tandem Assessment of Spoofing Countermeasures and Automatic Speaker Verification: FundamentalsabstractRecent years have seen growing efforts to develop spoofing countermeasures (CMs) to protect automatic speaker verification (ASV) systems from being deceived by manipulated or artificial inputs. The reliability of spoofing CMs is typically gauged using the equal error rate (EER) metric. The primitive EER fails to reflect application requirements and the impact of spoofing and CMs upon ASV and its use as a primary metric in traditional ASV research has long been abandoned in favour of risk-based approaches to assessment. This paper presents several new extensions to the tandem detection cost function (t-DCF), a recent risk-based approach to assess the reliability of spoofing CMs deployed in tandem with an ASV system. Extensions include a simplified version of the t-DCF with fewer parameters, an analysis of a special case for a fixed ASV system, simulations which give original insights into its interpretation and new analyses using the ASVspoof 2019 database. It is hoped that adoption of the t-DCF for the CM assessment will help to foster closer collaboration between the anti-spoofing and ASV research communities. Tomi Kinnunen, Héctor Delgado, Nicholas W. D. Evans, Kong-Aik Lee, Ville Vestman, Andreas Nautsch, Massimiliano Todisco, Xin Wang 0037, Md. Sahidullah, Junichi Yamagishi, Douglas A. Reynolds |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Maximal Figure-of-Merit Framework to Detect Multi-Label Phonetic Features for Spoken Language RecognitionabstractBottleneck features (BNFs) generated with a deep neural network (DNN) have proven to boost spoken language recognition accuracy over basic spectral features significantly. However, BNFs are commonly extracted using language-dependent tied-context phone states as learning targets. Moreover, BNFs are less phonetically expressive than the output layer in a DNN, which is usually not used as a speech feature because of its very high dimensionality hindering further post-processing. In this article, we put forth a novel deep learning framework to overcome all of the above issues and evaluate it on the 2017 NIST Language Recognition Evaluation (LRE) challenge. We use manner and place of articulation as speech attributes, which lead to low-dimensional “universal” phonetic features that can be defined across all spoken languages. To model the asynchronous nature of the speech attributes while capturing their intrinsic relationships in a given speech segment, we introduce a new training scheme for deep architectures based on a Maximal Figure of Merit (MFoM) objective. MFoM introduces non-differentiable metrics into the backpropagation-based approach, which is elegantly solved in the proposed framework. The experimental evidence collected on the recent NIST LRE 2017 challenge demonstrates the effectiveness of our solution. In fact, the performance of speech language recognition (SLR) systems based on spectral features is improved for more than 5% absolute Cavg. Finally, the F1 metric can be brought from 77.6% up to 78.1% by combining the conventional baseline phonetic BNFs with the proposed articulatory attribute features. Ivan Kukanov, Trung Ngo Trong, Ville Hautamäki, Sabato Marco Siniscalchi, Valerio Mario Salerno, Kong-Aik Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2019 | The CORAL+ Algorithm for Unsupervised Domain Adaptation of PLDAabstractState-of-the-art speaker recognition systems comprise an x-vector (or i-vector) speaker embedding front-end followed by a probabilistic linear discriminant analysis (PLDA) backend. The effectiveness of these components relies on the availability of a large collection of labeled training data. In practice, it is common that the domains (e.g., language, demographic) in which the system is deployed differ from that we trained the system. To close the gap due to the domain mismatch, we propose an unsupervised PLDA adaptation algorithm to learn from a small amount of unlabeled in-domain data. The proposed method was inspired by a prior work on feature-based domain adaptation technique known as the correlation alignment (CORAL). We refer to the model-based adaptation technique proposed in this paper as CORAL+. The efficacy of the proposed technique is experimentally validated on the recent NIST 2016 and 2018 Speaker Recognition Evaluation (SRE'16, SRE'18) datasets. Kong-Aik Lee, Qiongqiong Wang, Takafumi Koshinaka |
ICASSP | 1 |
| 2019 | I4U Submission to NIST SRE 2018: Leveraging from a Decade of Shared ExperiencesabstractThe I4U consortium was established to facilitate a joint entry to NIST speaker recognition evaluations (SRE). The latest edition of such joint submission was in SRE 2018, in which the I4U submission was among the best-performing systems. SRE'18 also marks the 10-year anniversary of I4U consortium into NIST SRE series of evaluation. The primary objective of the current paper is to summarize the results and lessons learned based on the twelve sub-systems and their fusion submitted to SRE'18. It is also our intention to present a shared view on the advancements, progresses, and major paradigm shifts that we have witnessed as an SRE participant in the past decade from SRE'08 to SRE'18. In this regard, we have seen, among others, a paradigm shift from supervector representation to deep speaker embedding, and a switch of research challenge from channel compensation to domain adaptation. Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Hitoshi Yamamoto, Koji Okabe, Ville Vestman, Jing Huang 0019, Guo-Hong Ding, Hanwu Sun, Anthony Larcher, Rohan Kumar Das, Haizhou Li 0001, Mickael Rouvier, Pierre-Michel Bousquet, Wei Rao 0002, Qing Wang 0039, Fahimeh Bahmaninezhad, Héctor Delgado, Massimiliano Todisco |
INTERSPEECH | 1 |
| 2019 | The NEC-TT 2018 Speaker Verification System
Kong-Aik Lee, Hitoshi Yamamoto, Koji Okabe, Qiongqiong Wang, Takafumi Koshinaka, Jiacen Zhang, Koichi Shinoda |
INTERSPEECH | 1 |
| 2019 | ASVspoof 2019: Future Horizons in Spoofed and Fake Audio DetectionabstractASVspoof, now in its third edition, is a series of community-led challenges which promote the development of countermeasures to protect automatic speaker verification (ASV) from the threat of spoofing. Advances in the 2019 edition include: (i) a consideration of both logical access (LA) and physical access (PA) scenarios and the three major forms of spoofing attack, namely synthetic, converted and replayed speech; (ii) spoofing attacks generated with state-of-the-art neural acoustic and waveform models; (iii) an improved, controlled simulation of replay attacks; (iv) use of the tandem detection cost function (t-DCF) that reflects the impact of both spoofing and countermeasures upon ASV reliability. Even if ASV remains the core focus, in retaining the equal error rate (EER) as a secondary metric, ASVspoof also embraces the growing importance of fake audio detection. ASVspoof 2019 attracted the participation of 63 research teams, with more than half of these reporting systems that improve upon the performance of two baseline spoofing countermeasures. This paper describes the 2019 database, protocols and challenge results. It also outlines major findings which demonstrate the real progress made in protecting against the threat of spoofing and fake audio. Massimiliano Todisco, Xin Wang 0037, Ville Vestman, Md. Sahidullah, Héctor Delgado, Andreas Nautsch, Junichi Yamagishi, Nicholas W. D. Evans, Tomi Kinnunen, Kong-Aik Lee |
INTERSPEECH | 10 |
| 2019 | Unleashing the Unused Potential of i-Vectors Enabled by GPU AccelerationabstractSpeaker embeddings are continuous-value vector representations that allow easy comparison between voices of speakers with simple geometric operations. Among others, i-vector and x-vector have emerged as the mainstream methods for speaker embedding. In this paper, we illustrate the use of modern computation platform to harness the benefit of GPU acceleration for i-vector extraction. In particular, we achieve an acceleration of 3000 times in frame posterior computation compared to real time and 25 times in training the i-vector extractor compared to the CPU baseline from Kaldi toolkit. This significant speed-up allows the exploration of ideas that were hitherto impossible. In particular, we show that it is beneficial to update the universal background model (UBM) and re-compute frame alignments while training the i-vector extractor. Additionally, we are able to study different variations of i-vector extractors more rigorously than before. In this process, we reveal some undocumented details of Kaldi's i-vector extractor and show that it outperforms the standard formulation by a margin of 1 to 2% when tested with VoxCeleb speaker verification protocol. All of our findings are asserted by ensemble averaging the results from multiple runs with random start. Ville Vestman, Kong-Aik Lee, Tomi Kinnunen, Takafumi Koshinaka |
INTERSPEECH | 2 |
| 2019 | Speaker Augmentation and Bandwidth Extension for Deep Speaker Embedding
Hitoshi Yamamoto, Kong-Aik Lee, Koji Okabe, Takafumi Koshinaka |
INTERSPEECH | 2 |
| 2018 | Maximal Figure-of-Merit Embedding for Multi-Label Audio ClassificationabstractThis work tackles the problem of the domestic audio tagging or environmental sound classification, where one audio recording can contain one or more acoustic events and a recognizer should output all of those tags. A baseline model for this task is a convolutional recurrent neural network (CRNN) with sigmoid output nodes optimized using the binary cross-entropy objective. Traditional error metrics, such as classification error, are not suitable for this type of task. In this work, we show that the maximal figure-of-merit (MFoM) framework helps to separate the multi-label classes in terms of equal error rate (EER). We embed MFoM into the deep learning objective function and gain more than 9% relative improvement, compared to the baseline model with binary cross-entropy. Ivan Kukanov, Ville Hautamäki, Kong-Aik Lee |
ICASSP | 3 |
| 2018 | Speaker-Phonetic Vector Estimation for Short Duration Speaker VerificationabstractPhonetic variability is one of the primary challenges in short duration speaker verification. This paper proposes a novel method that modifies the standard normal distribution prior in the total variability model to use a mixture of Gaussians as the prior distribution. The proposed speaker-phonetic vectors are then estimated from the posterior probability of latent variables, and each vector has a phonetic meaning. Unlike the standard total variability model, the proposed method can incorporate a phoneme classifier to perform soft content matching, which has the potential to solve the phonetic variability problem. Parameter estimation and scoring formulae for speaker-phonetic vectors method are presented. Experimental results obtained using NIST 2010 data show that the proposed technique leads to relative improvements of more than 30% when fused with total variability model and tested on 3 second duration test files. Vidhyasaharan Sethu, Eliathamby Ambikairajah, Kong-Aik Lee |
ICASSP | 4 |
| 2018 | On the Importance of Analytic Phase of Speech Signals in Spoken Language RecognitionabstractIn this paper, we study the role of long-time analytic phase of speech signals in spoken language recognition (SLR) and employ a set of features termed as instantaneous frequency cepstral coefficients (IFCC). We extract IFCC from long-time analytic phase, in an effort to capture long range acoustic features from speech signals. These features are used in combination with the traditional shifted delta cepstral coefficients (SDCC) for SLR. As the SDCC are extracted from spectral magnitude and IFCC are from analytic phase, they characterize long-time information of speech in different ways. The experiments conducted with NIST LRE 2017 task reveals the complementary effects of IFCC features to SDCC and deep bottleneck (DBN) features. The fusion of IFCC with SDCC/DBN features delivered relative improvements of 23.23% and 16.78% in average equal error rate over the SDCC and DBN features, respectively, indicating the benefits of information from analytic phase in SLR. Karthika Vijayan, Haizhou Li 0001, Hanwu Sun, Kong-Aik Lee |
ICASSP | 4 |
| 2018 | Integrated Presentation Attack Detection and Automatic Speaker Verification: Common Features and Gaussian Back-end FusionabstractInternational audience Massimiliano Todisco, Héctor Delgado, Kong-Aik Lee, Md. Sahidullah, Nicholas W. D. Evans, Tomi Kinnunen, Junichi Yamagishi |
INTERSPEECH | 3 |
| 2018 | Co-whitening of I-vectors for Short and Long Duration Speaker Verification
Longting Xu, Kong-Aik Lee, Haizhou Li 0001, Zhen Yang 0001 |
INTERSPEECH | 2 |
| 2018 | Attention Mechanism in Speaker Recognition: What Does it Learn in Deep Speaker Embedding?abstractThis paper presents an experimental study on deep speaker embedding with an attention mechanism that has been found to be a powerful representation learning technique in speaker recognition. In this framework, an attention model works as a frame selector that computes an attention weight for each frame-level feature vector, in accord with which an utterance-level representation is produced at the pooling layer in a speaker embedding network. In general, an attention model is trained together with the speaker embedding network on a single objective function, and thus those two components are tightly bound to one another. In this paper, we consider the possibility that the attention model might be decoupled from its parent network and assist other speaker embedding networks and even conventional i-vector extractors. This possibility is demonstrated through a series of experiments on a NIST Speaker Recognition Evaluation (SRE) task, with 9.0% EER reduction and 3.8% minCprimaryreduction when the attention weights are applied to i-vector extraction. Another experiment shows that DNN-based soft voice activity detection (VAD) can be effectively combined with the attention mechanism to yield further reduction of minCprimaryby 6.6% and 1.6% in deep speaker embedding and i-vector systems, respectively. Qiongqiong Wang, Koji Okabe, Kong-Aik Lee, Hitoshi Yamamoto, Takafumi Koshinaka |
SLT | 3 |
| 2018 | Generalized Variability Model for Speaker VerificationabstractIn this letter, we propose a generalized variability model as an extension to the total variability model. While the total variability model employs a standard normal prior distribution in its typical setup, the proposed generalized variability model relaxes this assumption and allows the latent variable distribution to be a mixture of Gaussians. The conventional total variability model can then be viewed as a special case of this generalized version where the number of mixture components is constrained to one. This proposed model is validated in the context of speaker verification tasks on both the standard and extended NIST SRE 2010 datasets. Experimental results show that modeling the distribution of the latent variables as a mixture of Gaussians leads to a better performance under all conditions and a greater gain can be expected for speaker verification using short utterances. Vidhyasaharan Sethu, Eliathamby Ambikairajah, Kong-Aik Lee |
IEEE Signal Process. Lett. | 4 |
| 2018 | Generalizing I-Vector Estimation for Rapid Speaker RecognitionabstractAn i-vector is a compact representation that captures both the speaker and session variabilities rendered in a spoken utterance. Over the past years, it has prevailed over other techniques and is now the de facto representation for text-independent speaker recognition. Standard i-vector extraction requires intense computation at run-time. Reducing the computation will allow effective use of i-vector in more applications. Such intense computation arises from the posterior covariance matrix, when estimating the i-vector. There have been studies on how to simplify the computation of posterior covariance matrix with modest success. In this paper, we propose a novel approach to i-vector extraction without the need to evaluate the full posterior covariance thereby speeding up the run-time extraction process. This is achieved by generalizing the i-vector estimation in two ways. First, we introduce the use of occupancy reweighting in conjunction with whitening over the Baum-Welch statistics as part of the preprocessing step. Second, we introduce the so-called subspace-orthogonalizing prior (SOP) to replace the standard Gaussian prior in i-vector formulation. Experiments conducted on the extended-core task of NIST SRE'10 show that the proposed rapid SOP approach achieves considerable speed-up over the standard i-vector with comparable equal error rates. Longting Xu, Kong-Aik Lee, Haizhou Li 0001, Zhen Yang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Adaptation of PLDA for multi-source text-independent speaker verificationabstractProbabilistic linear discriminant analysis (PLDA) is widely described as an effective model for text-independent speaker verification in the i-vector space. The PLDA scoring function is typically formulated as the likelihood ratio between the speaker-adapted and the universal PLDAs. In this case, the adaptation of PLDA was performed through the speaker factors. In this paper, we show that the channel factors of the PLDA could be equivalently exploited to deal with the multi-source conditions. In speaker verification, with the proposed method, a PLDAmodel trained on conversational telephone speech could be adequately adapted for interview-style microphone recordings. Experimental results on NIST SRE'08 and SRE'10 datasets confirm that the proposed method is effective, especially for the case whereby enrollment and test utterances were captured from different sources. Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001, Li-Rong Dai 0001 |
ICASSP | 2 |
| 2017 | RedDots replayed: A new replay spoofing attack corpus for text-dependent speaker verification researchabstractThis paper describes a new database for the assessment of automatic speaker verification (ASV) vulnerabilities to spoofing attacks. In contrast to other recent data collection efforts, the new database has been designed to support the development of replay spoofing countermeasures tailored towards the protection of text-dependent ASV systems from replay attacks in the face of variable recording and playback conditions. Derived from the re-recording of the original RedDots database, the effort is aligned with that in text-dependent ASV and thus well positioned for future assessments of replay spoofing countermeasures, not just in isolation, but in integration with ASV. The paper describes the database design and re-recording, a protocol and some early spoofing detection results. The new “RedDots Replayed” database is publicly available through a creative commons license. Tomi Kinnunen, Md. Sahidullah, Mauro Falcone, Luca Costantini, Rosa González Hautamäki, Dennis Alexander Lehmann Thomsen, Achintya Kumar Sarkar, Zheng-Hua Tan, Héctor Delgado, Massimiliano Todisco, Nicholas W. D. Evans, Ville Hautamäki, Kong-Aik Lee |
ICASSP | 13 |
| 2017 | The ASVspoof 2017 Challenge: Assessing the Limits of Replay Spoofing Attack DetectionabstractThe ASVspoof initiative was created to promote the development of countermeasures which aim to protect automatic speaker verification (ASV) from spoofing attacks. The first community-led, common evaluation held in 2015 focused on countermeasures for speech synthesis and voice conversion spoofing attacks. Arguably, however, it is replay attacks which pose the greatest threat. Such attacks involve the replay of recordings collected from enrolled speakers in order to provoke false alarms and can be mounted with greater ease using everyday consumer devices. ASVspoof 2017, the second in the series, hence focused on the development of replay attack countermeasures. This paper describes the database, protocols and initial findings. The evaluation entailed highly heterogeneous acoustic recording and replay conditions which increased the equal error rate (EER) of a baseline ASV system from 1.76% to 30.71%. Submissions were received from 49 research teams, 20 of which improved upon a baseline replay spoofing detector EER of 24.65%, in terms of replay/non-replay discrimination. While largely successful, the evaluation indicates that the quest for countermeasures which are resilient in the face of variable replay attacks remains very much alive. Tomi Kinnunen, Md. Sahidullah, Héctor Delgado, Massimiliano Todisco, Nicholas W. D. Evans, Junichi Yamagishi, Kong-Aik Lee |
INTERSPEECH | 7 |
| 2017 | The I4U Mega Fusion and Collaboration for NIST Speaker Recognition Evaluation 2016abstract18th Annual Conference of the International Speech Communication Association, INTERSPEECH 2017, Stockholm, Sweden, 20-24 August 2017 Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Anthony Larcher, Andreas Nautsch, Themos Stafylakis, Gang Liu 0001, Mickael Rouvier, Wei Rao 0002, Federico Alegre, Man-Wai Mak, Achintya Kumar Sarkar, Héctor Delgado, Rahim Saeidi, Hagai Aronowitz, Aleksandr Sizov, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Bin Ma 0001, Ville Vestman, Md. Sahidullah, M. Halonen, Anssi Kanervisto, Gaël Le Lan, Fahimeh Bahmaninezhad, Sergey Isadskiy, Christian Rathgeb, Christoph Busch 0001, Georgios Tzimiropoulos, Q. Qian, Q. Zhao, J. Xue, R. Jin, T. Zhao, Pierre-Michel Bousquet, Moez Ajili, Waad Ben Kheder, Driss Matrouf, Zhi Hao Lim, Chenglin Xu, Haihua Xu 0001, Chng Eng Siong, Benoit G. B. Fauve, Kaavya Sriskandaraja, Vidhyasaharan Sethu, W. W. Lin, Dennis Alexander Lehmann Thomsen, Zheng-Hua Tan, Massimiliano Todisco, Nicholas W. D. Evans, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Eliathamby Ambikairajah |
INTERSPEECH | 1 |
| 2017 | Gain Compensation for Fast i-Vector Extraction Over Short Duration
Kong-Aik Lee, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2017 | Incorporating Local Acoustic Variability Information into Short Duration Speaker Verification
Vidhyasaharan Sethu, Eliathamby Ambikairajah, Kong-Aik Lee |
INTERSPEECH | 4 |
| 2017 | Direct Optimization of the Detection Cost for I-Vector-Based Spoken Language RecognitionabstractWe explore a method to boost discriminative capabilities of probabilistic linear discriminant analysis (PLDA) model without losing its generative advantages. We show a sequential projection and training steps leading to a classifier that operates in the original i-vector space but is discriminatively trained in a low-dimensional PLDA latent subspace. We use extended Baum-Welch technique to optimize the model with respect to two objective functions for discriminative training. One of them is the well-known maximum mutual information objective, while the other one is a new objective that we propose to approximate the language detection cost. We evaluate the performance on NIST language recognition evaluation (LRE) 2015 and our development dataset comprised of the utterances from previous LREs. We improve the detection cost by 10% and 6% relative compared to our fine-tuned generative and discriminative baselines, and by 10% over the best of our previously reported results. The proposed approximation method of the cost function and PLDA subspace training are applicable for a broad range of tasks. Aleksandr Sizov, Kong-Aik Lee, Tomi Kinnunen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Content-aware local variability vector for speaker verification with short utteranceabstractI-vector has shown to be very effective in speaker verification with long-duration speech utterances. But when test utterances are of short duration, content mismatch between the enrollment and test utterances limit the performance of i-vector system. This paper proposes to extract local session variability vectors on different phonetic classes from the utterances instead of estimating the session variability across the whole utterance as i-vector does. Using the posteriors given by a deep neural network (DNN) trained for phone state classification, the local vectors represent the session variability contained in specific phonetic content. Our experiments show that the content-aware local vectors are better at coping with the content mismatch between training and test utterances of short durations for text-independent, text-constrained and text-dependent tasks. Kong-Aik Lee, Chng Eng Siong, Bin Ma 0001, Haizhou Li 0001, Li-Rong Dai 0001 |
ICASSP | 2 |
| 2016 | An extensible speaker identification sidekit in PythonabstractSIDEKIT is a new open-source Python toolkit that includes a large panel of state-of-the-art components and allow a rapid prototyping of an end-to-end speaker recognition system. For each step from front-end feature extraction, normalization, speech activity detection, modelling, scoring and visualization, SIDEKIT offers a wide range of standard algorithms and flexible interfaces. The use of a single efficient programming and scripting language (Python in this case), and the limited dependencies, facilitate the deployment for industrial applications and extension to include new algorithms as part of the whole tool-chain provided by SIDEKIT. Performance of SIDEKIT is demonstrated on two standard evaluation tasks, namely the RSR2015 and NIST-SRE 2010. Anthony Larcher, Kong-Aik Lee, Sylvain Meignier |
ICASSP | 2 |
| 2016 | The 2015 NIST Language Recognition Evaluation: The Shared View of I2R, Fantastic4 and SingaMSabstractTechnical report for NIST LRE 2015 Workshop Kong-Aik Lee, Haizhou Li 0001, Li Deng 0001, Ville Hautamäki, Wei Rao 0002, Anthony Larcher, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Aleksandr Sizov, Jianshu Chen, Ivan Kukanov, Amir Hossein Poorjam, Trung Ngo Trong, Chenglin Xu, Haihua Xu 0001, Bin Ma 0001, Chng Eng Siong, Sylvain Meignier |
INTERSPEECH | 1 |
| 2016 | Twin Model G-PLDA for Duration Mismatch Compensation in Text-Independent Speaker Verification
Vidhyasaharan Sethu, Eliathamby Ambikairajah, Kong-Aik Lee |
INTERSPEECH | 4 |
| 2016 | Joint Speaker and Lexical Modeling for Short-Term Characterization of Speaker
Guangsen Wang, Kong-Aik Lee, Trung Hieu Nguyen 0001, Hanwu Sun, Bin Ma 0001 |
INTERSPEECH | 2 |
| 2016 | Total Variability Modeling Using Source-Specific PriorsabstractIn total variability modeling, variable length speech utterances are mapped to fixed low-dimensional i-vectors. Central to computing the total variability matrix and i-vector extraction, is the computation of the posterior distribution for a latent variable conditioned on an observed feature sequence of an utterance. In both cases the prior for the latent variable is assumed to be non-informative, since for homogeneous datasets there is no gain in generality in using an informative prior. This work shows in the heterogeneous case, that using informative priors for computing the posterior, can lead to favorable results. We focus on modeling the priors using minimum divergence criterion or factor analysis techniques. Tests on the NIST 2008 and 2010 Speaker Recognition Evaluation (SRE) dataset show that our proposed method beats four baselines: For i-vector extraction using an already trained matrix, for the short2-short3 task in SRE'08, five out of eight female and four out of eight male common conditions, were improved. For the core-extended task in SRE'10, four out of nine female and six out of nine male common conditions were improved. When incorporating prior information into the training of the T matrix itself, the proposed method beats the baselines for six out of eight female and five out of eight male common conditions, for SRE'08, and five and six out of nine conditions, for the male and female case, respectively, for SRE'10. Tests using factor analysis for estimating priors show that two priors do not offer much improvement, but in the case of three separate priors (sparse data), considerable improvements were gained. Sven Ewan Shepstone, Kong-Aik Lee, Haizhou Li 0001, Zheng-Hua Tan, Søren Holdt Jensen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Channel adaptation of plda for text-independent speaker verificationabstractProbabilistic linear discriminant analysis (PLDA) has shown to be effective for modeling channel variability in the i-vector space for text-independent speaker verification. Speaker verification is a binary hypothesis testing. Given a test segment, the verification score could be computed as the log-likelihood ratio between a speaker-adapted PLDA and the universal PLDA model. This work proposes to infer the channel factor specific to each test segment and to include the channel estimate in the PLDA models, which essentially shifts the scoring function to better match that of the test channel. We also explore the influence of covariance adaptation in both speaker and channel adaptations. Experimental results on NIST SRE'08 and SRE'10 dataset confirm that the proposed channel adaptation can be effective when the covariance is kept un-adapted, while the covariance adaptation is necessary in the speaker adaptation. Kong-Aik Lee, Bin Ma 0001, Wu Guo, Haizhou Li 0001, Li-Rong Dai 0001 |
ICASSP | 2 |
| 2015 | Normalization of total variability matrix for i-vector/PLDA speaker verificationabstractGaussian PLDA with uncertainty propagation is effective for i-vector based speaker verification. The idea is to propagate the uncertainty of i-vectors caused by the duration variability of utterances to the PLDA model. However, a limitation of the method is the difficulty of performing length normalization on the posterior covariance matrix of an i-vector. This paper proposes a method to avoid performing length normalization on i-vectors in Gaussian PLDA modeling so that uncertainty propagation can be directly applied without transforming the posterior covariance matrices of i-vectors. Instead of performing length normalization on i-vectors independently, the proposed method normalizes the column vectors of the total variability matrix. Because the i-vectors of all utterances are derived from the same normalized total variability matrix, they will be subject to the same degree of normalization, thereby avoiding the undesirable distortion introduced by the utterance-dependent length-normalization process. Experimental results on both NIST 2010 and 2012 SREs demonstrate that the proposed method achieves a performance similar to (and in some situations better than) that of Gaussian PLDA with length normalization. The method has the potential of improving the performance of uncertainty propagation for i-vector/PLDA speaker verification. Wei Rao 0002, Man-Wai Mak, Kong-Aik Lee |
ICASSP | 3 |
| 2015 | Source-specific informative prior for i-vector extractionabstractAn i-vector is a low-dimensional fixed-length representation of a variable-length speech utterance, and is defined as the posterior mean of a latent variable conditioned on the observed feature sequence of an utterance. The assumption is that the prior for the latent variable is non-informative, since for homogeneous datasets there is no gain in generality in using an informative prior. This work shows that extracting i-vectors for a heterogeneous dataset, containing speech samples recorded from multiple sources, using informative priors instead is applicable, and leads to favorable results. Tests carried out on the NIST 2008 and 2010 Speaker Recognition Evaluation (SRE) dataset show that our proposed method beats three baselines: For the short2-short3 core-task in SRE'08, for the female and male cases, five and six respectively, out of eight common conditions were beaten, and for the core-core task in SRE'10, for both genders, five out of nine common conditions were beaten. Sven Ewan Shepstone, Kong-Aik Lee, Haizhou Li 0001, Zheng-Hua Tan, Søren Holdt Jensen |
ICASSP | 2 |
| 2015 | A new study of GMM-SVM system for text-dependent speaker recognitionabstractThis paper presents a new approach and the study of GMM-SVM system for text-dependent speaker recognition on scenario of the fixed pass-phrases. The uniform-split content-based GMM-SVM system is proposed and applied to text-dependent speaker evaluation. We conducted detailed study of the proposed method compared to the baseline GMM-SVM system on the RSR2015 database, which has been designed and collected for the evaluation of text-dependent speaker verification system. The experiment results show that the new approach can significantly reduce the detection error of the target-wrong error type (i.e., target speaker with wrong pass-phrase) while maintaining a low detection error for both imposter-correct and imposter-wrong error types (i.e., imposter with correct pass-phrase and imposter with wrong pass-phrase). We also show that score normalization could be applied with respect to the imposter-wrong distribution as opposed to the imposter-correct distribution. Hanwu Sun, Kong-Aik Lee, Bin Ma 0001 |
ICASSP | 2 |
| 2015 | Phone-centric local variability vector for text-constrained speaker verification
Kong-Aik Lee, Bin Ma 0001, Wu Guo, Haizhou Li 0001, Li-Rong Dai 0001 |
INTERSPEECH | 2 |
| 2015 | The reddots data collection for speaker recognitionabstractde niveau recherche, publiés ou non, émanant des établissements d'enseignement et de recherche français ou étrangers, des laboratoires publics ou privés. Kong-Aik Lee, Anthony Larcher, Guangsen Wang, Patrick Kenny, Niko Brümmer, David A. van Leeuwen, Hagai Aronowitz, Marcel Kockmann, Carlos Vaquero, Bin Ma 0001, Haizhou Li 0001, Themos Stafylakis, Jahangir Alam 0001, Albert Swart, Javier Perez |
INTERSPEECH | 1 |
| 2015 | The reddots platform for mobile crowd-sourcing of speech data
Kong-Aik Lee, Guangsen Wang, Kam Pheng Ng, Hanwu Sun, Trung Hieu Nguyen 0001, Ngoc Thuy Huong Thai, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2015 | Sparse coding of total variability matrix
Longting Xu, Kong-Aik Lee, Haizhou Li 0001, Zhen Yang 0001 |
INTERSPEECH | 2 |
| 2015 | Relevance factor of maximum a posteriori adaptation for GMM-NAP-SVM in speaker and language recognition
Chang Huai You, Haizhou Li 0001, Kong-Aik Lee |
Comput. Speech Lang. | 3 |
| 2015 | Quasi-Factorial Prior for i-vector ExtractionabstractWe analyze the i-vector extraction from the perspective of the prior distribution exerted on the mean supervector of Gaussian mixture model (GMM). To this end, we start off with the analysis of the subspace prior which leads to the compressed representation in the standard i-vector extraction. We then propose the use of quasi-factorial prior and show how it impacts the total variability space and its application for i-vector extraction. The quasi-factorial prior could be used in a standalone manner, or in combination with a subspace prior. In the latter context, we found that the performance of the standard i-vector can be greatly improved with the use of quasi-factorial prior followed by a subspace prior. This assertion is confirmed through experiments conducted on the NIST 2010 Speaker Recognition Evaluation (SRE10) dataset. Kong-Aik Lee, Li-Rong Dai 0001, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 2 |
| 2014 | Minimum divergence estimation of speaker prior in multi-session PLDA scoringabstractProbabilistic linear discriminant analysis (PLDA) has shown to be effective for modeling speaker and channel variability in the i-vector space for text-independent speaker verification. This paper shows that the PLDA scoring function could be formulated as model comparison between an adapted PLDA model and the universal PLDA. Based on this formulation, we show that a more robust adaptation could be attained by adapting the PLDA model through the use of minimum divergence estimate of speaker prior in the latent subspace. Experimental results on NIST SRE'10 and SRE'12 dataset confirm that the proposed method is effective in handling multi-session task. Notably, it is free from the covariance shrinkage problem typically found in the standard multi-session PLDA scoring. Kong-Aik Lee, Bin Ma 0001, Wu Guo, Haizhou Li 0001, Li-Rong Dai 0001 |
ICASSP | 2 |
| 2014 | Modelling the alternative hypothesis for text-dependent speaker verificationabstractThis paper describes text-dependent speaker verification as a task involving four classes of trials depending on whether the target speaker or an impostor pronounces the expected pass-phrase or not. These four classes are used to reformulate the log-likelihood ratio traditionally used in text-independent speaker verification. Three formulations of the alternative hypothesis are considered, leading to three new expressions of the verification score. Experiments performed on the publicly available RSR2015 database show a significant improvement compared to existing baseline scores. A relative gain up to 61% in term of minimum cost is achieved when considering that the alternative hypothesis is the union of three sub-hypotheses corresponding to the three existing classes of impostures. Anthony Larcher, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 2 |
| 2014 | Imposture classification for text-dependent speaker verificationabstractThis work focuses on text-dependent speaker verification, where a user is required to chose and pronounce a customized pass-phrase to get authenticated. In this context, there are three types of impostures: an impostor pronouncing the correct pass-phrase, an impostor pronouncing a wrong pass-phrase and the most difficult one: an impostor playing back a recording of the target speaker pronouncing a wrong pass-phrase. Detecting and classifying different types of impostures can help to prevent future impostures of the same type. In this work, we first propose a new verification score to reject Playback impostures. This score allows a relative reduction of 90% of the equal error rate against Playback impostures while offering performance similar to the baseline text-dependent score against other types of impostures. As a second contribution, we show that the new score can be combined with an existing text-dependent verification score to improve the classification of the different types of impostures. The performance of the speaker verification engine for imposture classification is significantly improved with the Cllrdecreasing by at least 29% compared to the original system. Anthony Larcher, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 2 |
| 2014 | Extended RSR2015 for text-dependent speaker verification over VHF channelabstractInternational audience Anthony Larcher, Kong-Aik Lee, Pablo Luis Sordo Martinez, Trung Hieu Nguyen 0001, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2014 | Text-dependent speaker verification: Classifiers, databases and RSR2015abstractThe RSR2015 database, designed to evaluate text-dependent speaker verification systems under different durations and lexical constraints has been collected and released by the Human Language Technology (HLT) department at Institute for Infocomm Research (I2R) in Singapore. English speakers were recorded with a balanced diversity of accents commonly found in Singapore. More than 151 h of speech data were recorded using mobile devices. The pool of speakers consists of 300 participants (143 female and 157 male speakers) between 17 and 42 years old making the RSR2015 database one of the largest publicly available database targeted for text-dependent speaker verification. We provide evaluation protocol for each of the three parts of the database, together with the results of two speaker verification system: the HiLAM system, based on a three layer acoustic architecture, and an i-vector/PLDA system. We thus provide a reference evaluation scheme and a reference performance on RSR2015 database to the research community. The HiLAM outperforms the state-of-the-art i-vector system in most of the scenarios. Anthony Larcher, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001 |
Speech Commun. | 2 |
| 2013 | Phonetically-constrained PLDA modeling for text-dependent speaker verification with multiple short utterancesabstractThe importance of phonetic variability for short duration speaker verification is widely acknowledged. This paper assesses the performance of Probabilistic Linear Discriminant Analysis (PLDA) and i-vector normalization for a text-dependent verification task. We show that using a class definition based on both speaker and phonetic content significantly improves the performance of a state-of-the-art system. We also compare four models for computing the verification scores using multiple enrollment utterances and show that using PLDA intrinsic scoring obtains the best performance in this context. This study suggests that such scoring regime remains to be optimized. Anthony Larcher, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 2 |
| 2013 | Anti-model KL-SVM-NAP system for NIST SRE 2012 evaluationabstractThis paper presents an anti-model based speaker recognition system for NIST SRE 2012 evaluation, which is one of subsystems in IIR SRE12 submission. We apply the anti-model approach for the SRE12 evaluation. The KL-SVM-NAP based speaker recognition system is adopted to evaluate the performance. We present detailed comparison study of the classical KL-SVM-NAP based speaker recognition system and anti-model based KL-SVM-NAP system for NIST 2012 speaker recognition evaluation. The results are reported on in-house pre-SRE12 development set and NIST SRE12 core task. The clear advantages of the anti-model approach over that the traditional KL-SVM-NAP approach are presented and discussed. Hanwu Sun, Kong-Aik Lee, Bin Ma 0001 |
ICASSP | 2 |
| 2013 | A study on GMM-SVM with adaptive relevance factor and its comparison with i-vector and JFA for speaker recognitionabstractRecently, joint factor analysis (JFA) and identity-vector (i-vector) represent the dominant techniques used for speaker recognition due to their superior performance. Developed relatively earlier, the Gaussian mixture model - support vector machine (GMM-SVM) with nuisance attribute projection (NAP) has gradually become less popular. However, when developing the relevance factor in maximum a posteriori (MAP) estimation of GMM to be adapted by application data in place of the conventional fixed value, it is noted that GMM-SVM demonstrates some advantages. In this paper, we conduct a comparative study between GMM-SVM with adaptive relevance factor and JFA/i-vector under the framework of Speaker Recognition Evaluation (SRE) formulated by the National Institute of Standards and Technology (NIST). Chang Huai You, Haizhou Li 0001, Bin Ma 0001, Kong-Aik Lee |
ICASSP | 4 |
| 2013 | Automatic regularization of cross-entropy cost for speaker recognition fusionabstract\n Contains fulltext :\n 116325.pdf (author's version ) (Open Access)\n Ville Hautamäki, Kong-Aik Lee, David A. van Leeuwen, Rahim Saeidi, Anthony Larcher, Tomi Kinnunen, Taufiq Hasan, Seyed Omid Sadjadi, Gang Liu 0001, Hynek Boril, John H. L. Hansen, Benoit G. B. Fauve |
INTERSPEECH | 2 |
| 2013 | ALIZE 3.0 - open source toolkit for state-of-the-art speaker recognitionabstractInternational audience Anthony Larcher, Jean-François Bonastre, Benoit G. B. Fauve, Kong-Aik Lee, Christophe Lévy, Haizhou Li 0001, John S. D. Mason, Jean-Yves Parfait |
INTERSPEECH | 4 |
| 2013 | Multi-session PLDA scoring of i-vector for partially open-set speaker detectionabstractInternational audience Kong-Aik Lee, Anthony Larcher, Chang Huai You, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2013 | Vulnerability evaluation of speaker verification under voice conversion spoofing: the effect of text constraintsabstractVoice conversion, a technique to change one's voice to sound like that of another, poses a threat to even high performance speaker verification system. Vulnerability of text-independent speaker verification systems under spoofing attack, using statistical voice conversion technique, was evaluated and confirmed in our previous work. In this paper, we further extend the study to text-dependent speaker verification systems. In particular, we compare both joint density Gaussian mixture model (JD-GMM) and unit-selection (US) spoofing methods and, for the first time, the performances of text-independent and text-dependent speaker verification systems in a single study. We conduct the experiments using RSR2015 database which is recorded using multiple mobile devices. The experimental results indicate that text-dependent speaker verification system tolerates spoofing attacks better than the text-independent counterpart. Zhizheng Wu 0001, Anthony Larcher, Kong-Aik Lee, Chng Eng Siong, Tomi Kinnunen, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2013 | I4u submission to NIST SRE 2012: a large-scale collaborative effort for noise-robust speaker verificationabstractI4U is a joint entry of nine research Institutes and Universities across 4 continents to NIST SRE 2012. It started with a brief discussion during the Odyssey 2012 workshop in Singapore. An online discussion group was soon set up, providing a discussion platform for different issues surrounding NIST SRE’12. Noisy test segments, uneven multi-session training, variable enrollment duration, and the issue of open-set identification were actively discussed leading to various solutions integrated to the I4U submission. The joint submission and several of its 17 sub-systems were among top-performing systems. We summarize the lessons learnt from this large-scale effort. Rahim Saeidi, Kong-Aik Lee, Tomi Kinnunen, Tawfik Hasan, Benoit G. B. Fauve, Pierre-Michel Bousquet, Elie Khoury 0001, Pablo Luis Sordo Martinez, Jia Min Karen Kua, Chang Huai You, Hanwu Sun, Anthony Larcher, Padmanabhan Rajan, Ville Hautamäki, Cemal Hanilçi, Billy Braithwaite, Rosa González Hautamäki, Seyed Omid Sadjadi, Gang Liu 0001, Hynek Boril, Navid Shokouhi, Driss Matrouf, Laurent El Shafey, Pejman Mowlaee, Julien Epps, Tharmarajah Thiruvaran, David A. van Leeuwen, Bin Ma 0001, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Sébastien Marcel, John S. D. Mason, Eliathamby Ambikairajah |
INTERSPEECH | 2 |
| 2013 | Spoken Language Recognition: From Fundamentals to PracticeabstractSpoken language recognition refers to the automatic process through which we determine or verify the identity of the language spoken in a speech sample. We study a computational framework that allows such a decision to be made in a quantitative manner. In recent decades, we have made tremendous progress in spoken language recognition, which benefited from technological breakthroughs in related areas, such as signal processing, pattern recognition, cognitive science, and machine learning. In this paper, we attempt to provide an introductory tutorial on the fundamentals of the theory and the state-of-the-art solutions, from both phonological and computational aspects. We also give a comprehensive review of current trends and future research directions using the language recognition evaluation (LRE) formulated by the National Institute of Standards and Technology (NIST) as the case studies. Haizhou Li 0001, Bin Ma 0001, Kong-Aik Lee |
Proc. IEEE | 3 |
| 2013 | Sparse Classifier Fusion for Speaker VerificationabstractState-of-the-art speaker verification systems take advantage of a number of complementary base classifiers by fusing them to arrive at reliable verification decisions. In speaker verification, fusion is typically implemented as a weighted linear combination of the base classifier scores, where the combination weights are estimated using a logistic regression model. An alternative way for fusion is to use classifier ensemble selection, which can be seen as sparse regularization applied to logistic regression. Even though score fusion has been extensively studied in speaker verification, classifier ensemble selection is much less studied. In this study, we extensively study a sparse classifier fusion on a collection of twelve I4U spectral subsystems on the NIST 2008 and 2010 speaker recognition evaluation (SRE) corpora. Ville Hautamäki, Tomi Kinnunen, Filip Sedlak, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001 |
IEEE Trans. Speech Audio Process. | 4 |
| 2012 | Vulnerability of speaker verification systems against voice conversion spoofing attacks: The case of telephone speechabstractVoice conversion - the methodology of automatically converting one's utterances to sound as if spoken by another speaker - presents a threat for applications relying on speaker verification. We study vulnerability of text-independent speaker verification systems against voice conversion attacks using telephone speech. We implemented a voice conversion systems with two types of features and nonparallel frame alignment methods and five speaker verification systems ranging from simple Gaussian mixture models (GMMs) to state-of-the-art joint factor analysis (JFA) recognizer. Experiments on a subset of NIST 2006 SRE corpus indicate that the JFA method is most resilient against conversion attacks. But even it experiences more than 5-fold increase in the false acceptance rate from 3.24 % to 17.33 %. Tomi Kinnunen, Zhizheng Wu 0001, Kong-Aik Lee, Filip Sedlak, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 3 |
| 2012 | I-vectors in the context of phonetically-constrained short utterances for speaker verificationabstractShort speech duration remains a critical factor of performance degradation when deploying a speaker verification system. To overcome this difficulty, a large number of commercial applications impose the use of fixed pass-phrases. In this context, we show that the performance of the popular i-vector approach can be greatly improved by taking advantage of the phonetic information that they convey. Moreover, as i-vectors require a conditioning process to reach high accuracy, we show that further improvements are possible by taking advantage of this phonetic information within the normalisation process. We compare two methods, Within Class Covariance Normalization (WCCN) and Eigen Factor Radial (EFR), both relying on parameters estimated on the same development data. Our study suggests that WCCN is more robust to data mismatch but less efficient than EFR when the development data has a better match with the test data. Anthony Larcher, Pierre-Michel Bousquet, Kong-Aik Lee, Driss Matrouf, Haizhou Li 0001, Jean-François Bonastre |
ICASSP | 3 |
| 2012 | PLDA Modeling in I-Vector and Supervector Space for Speaker VerificationabstractIn this paper, we advocate the use of uncompressed form of i-vector. We employ the probabilistic linear discriminant analysis (PLDA) to handle speaker and session variability for speaker verification task. An i-vector is a low-dimensional vector containing both speaker and channel information acquired from a speech segment. When PLDA is used on i-vector, dimension reduction is performed twice – first in the i-vector extraction process and second in the PLDA model. Keeping the full dimensionality of i-vector in the supervector space for PLDA modeling and scoring would avoid unnecessary loss of information. The drawback of using PLDA on uncompressed i-vector is the inversion of large matrices, which we show can be solved rather efficiently by portioning large matrix into smaller blocks. We also introduce the Gaussianized rank-norm, as an alternative to whitening, for feature normalization prior to PLDA modeling. Index Terms: speaker verification, i-vector, probabilistic LDA 1. Kong-Aik Lee, Zhenmin Tang, Bin Ma 0001, Anthony Larcher, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2012 | RSR2015: Database for Text-Dependent Speaker Verification using Multiple Pass-PhrasesabstractInternational audience Anthony Larcher, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2012 | Effect of Relevance Factor of Maximum a posteriori Adaptation for GMM-SVM in Speaker and Language Recognition
Chang Huai You, Haizhou Li 0001, Bin Ma 0001, Kong-Aik Lee |
INTERSPEECH | 4 |
| 2012 | Low-Variance Multitaper MFCC Features: A Case Study in Robust Speaker VerificationabstractIn speech and audio applications, short-term signal spectrum is often represented using mel-frequency cepstral coefficients (MFCCs) computed from a windowed discrete Fourier transform (DFT). Windowing reduces spectral leakage but variance of the spectrum estimate remains high. An elegant extension to windowed DFT is the so-called multitaper method which uses multiple time-domain windows (tapers) with frequency-domain averaging. Multitapers have received little attention in speech processing even though they produce low-variance features. In this paper, we propose the multitaper method for MFCC extraction with a practical focus. We provide, first, detailed statistical analysis of MFCC bias and variance using autoregressive process simulations on the TIMIT corpus. For speaker verification experiments on the NIST 2002 and 2008 SRE corpora, we consider three Gaussian mixture model based classifiers with universal background model (GMM-UBM), support vector machine (GMM-SVM) and joint factor analysis (GMM-JFA). Multitapers improve MinDCF over the baseline windowed DFT by relative 20.4% (GMM-SVM) and 13.7% (GMM-JFA) on the interview-interview condition in NIST 2008. The GMM-JFA system further reduces MinDCF by 18.7% on the telephone data. With these improvements and generally noncritical parameter selection, multitaper MFCCs are a viable candidate for replacing the conventional MFCCs. Tomi Kinnunen, Rahim Saeidi, Filip Sedlak, Kong-Aik Lee, Johan Sandberg, Maria Sandsten, Haizhou Li 0001 |
IEEE Trans. Speech Audio Process. | 4 |
| 2011 | Classifier subset selection and fusion for speaker verificationabstractState-of-the-art speaker verification systems consists of a number of complementary subsystems whose outputs are fused, to arrive at more accurate and reliable verification decision. In speaker verification, fusion is typically implemented as a linear combination of the subsystem scores. Parameters of the linear model are commonly estimated using the logistic regression method, as implemented in the popular FoCal toolkit. In this paper, we study simultaneous use of classifier selection and fusion. We study four alternative fusion strategies, three score warping techniques, and provide interesting experimental bounds on optimal classifier subset selection. Detailed experiments are carried out on the NIST 2008 and 2010 SRE corpora. Filip Sedlak, Tomi Kinnunen, Ville Hautamäki, Kong-Aik Lee, Haizhou Li 0001 |
ICASSP | 4 |
| 2011 | Factored covariance modeling for text-independent speaker verificationabstractGaussian mixture models (GMMs) are commonly used to model the spectral distribution of speech signals for text-independent speaker verification. Mean vectors of the GMM, used in conjunction with support vector machine (SVM), have shown to be effective in characterizing speaker information. In addition to the mean vectors, covariance matrices capture the correlation between spectral features, which also represent some salient information about speaker identity. This paper investigates the use of local correlation between different dimensions of acoustic vector by using factor analysis and linear Gaussian model. Log-Euclidean inner product kernel is used to measure the similarity between two speech utterances in the form of covariance matrices. Experiments carried on NIST 2006 speaker verification tasks shows promising results. Eryu Wang, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001, Wu Guo, Li-Rong Dai 0001 |
ICASSP | 2 |
| 2011 | Speech enhancement with masking properties in eigen-domain for colored noiseabstractIn this paper, we study speech enhancement in eigen-domain. In our previous work on audible noise reduction, we use masking properties of the human auditory system to define the audible noise quantity in the eigen-domain. We then derived the speech enhancement algorithm using white noise model. Without loss of generality, in this paper, we introduce the noise reduction method for colored noise. Through many simulations, we show that colored noise modeling is superior over other existing eigen-decomposition methods in terms of objective and subjective evaluations. Chang Huai You, Kong-Aik Lee, Cheung-Chi Leung |
ICASSP | 2 |
| 2011 | Regularized Logistic Regression Fusion for Speaker VerificationabstractFusion of the base classifiers is seen as the way to achieve stateof-the art performance in the speaker verfication systems. Standard approach is to pose the fusion problem as the linear binary classification task. Most successful loss function in speaker verification fusion has been the weighted logistic regression popularized by the FoCal toolkit. However, it is known that optimizing logistic regression can overfit severely without appropriate regularization. In addition, subset classifier selection can be achieved by using an external 0/1 loss function on the best subset. In this work, we propose to use LASSO based regularization on the FoCal cost function to achive improved performance and classifier subset selection method integrated into one optimization task. Proposed method is able to achieve 51 % relative improvement in Actual DCF over the FoCal baseline. Index Terms: logistic regression, regularization, compressed sensing, linear fusion, speaker verification Ville Hautamäki, Kong-Aik Lee, Tomi Kinnunen, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2011 | Joint Application of Speech and Speaker Recognition for Automation and Security in Smart Home
Kong-Aik Lee, Anthony Larcher, Helen Thai, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2011 | Spoken Language Recognition in the Latent Topic SimplexabstractInternational audience Kong-Aik Lee, Chang Huai You, Ville Hautamäki, Anthony Larcher, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2011 | Study on the Relevance Factor of Maximum a Posteriori with GMM for Language RecognitionabstractIn this paper, the relevance factor in maximum a posteriori (MAP) adaptation of Gaussian mixture model (GMM) from universal background model (UBM) is studied for language recognition. In conventional MAP, relevance factor is typically set as a constant empirically. Knowing that relevance factor determines how much the observed training data influence the model adaptation, thus the resulting GMM models, we believe that the relevance factor should be dependent to the data for more effective modeling. We formulate the estimation of relevance factor in a systematic manner and study its role in characterizing spoken languages with supervectors. We use a Bhattacharyya-based language recognition system on National Institute of Standards and Technology (NIST) language recognition evaluation (LRE) 2009 task to investigate the validate of the data-dependent relevance factor. Experimental results show that we achieve improved performance by using the proposed relevance factor. Index Terms: maximum a posteriori, supervector, Gaussian mixture model, support vector machine Chang Huai You, Haizhou Li 0001, Kong-Aik Lee |
INTERSPEECH | 3 |
| 2011 | Using Discrete Probabilities With Bhattacharyya Measure for SVM-Based Speaker VerificationabstractSupport vector machines (SVMs), and kernel classifiers in general, rely on the kernel functions to measure the pairwise similarity between inputs. This paper advocates the use of discrete representation of speech signals in terms of the probabilities of discrete events as feature for speaker verification and proposes the use of Bhattacharyya coefficient as the similarity measure for this type of inputs to SVM. We analyze the effectiveness of the Bhattacharyya measure from the perspective of feature normalization and distribution warping in the SVM feature space. Experiments conducted on the NIST 2006 speaker verification task indicate that the Bhattacharyya measure outperforms the Fisher kernel, term frequency log-likelihood ratio (TFLLR) scaling, and rank normalization reported earlier in literature. Moreover, the Bhattacharyya measure is computed using a data-independent square-root operation instead of data-driven normalization, which simplifies the implementation. The effectiveness of the Bhattacharyya measure becomes more apparent when channel compensation is applied at the model and score levels. The performance of the proposed method is close to that of the popular GMM supervector with a small margin. Kong-Aik Lee, Chang Huai You, Haizhou Li 0001, Tomi Kinnunen, Khe Chai Sim |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2010 | Adaptive score fusion using Weighted Logistic Linear Regression for spoken language recognitionabstractState-of-the-art spoken language recognition systems typically consist of a combination of sub-systems. These sub-systems generate language detection scores for each speech segment, which will be fused (combined) to yield the overall detection scores. Typically, score fusion is achieved using a linear model and Logistic Linear Regression (LLR) is commonly used to estimate the model parameters. This paper proposes an extension to the LLR model, known as the Weighted LLR (WLLR). WLLR is obtained using a weighted combination of multiple LLRs where the weights are obtained as a nonlinear function of the speech segments. Although the resultant score is still linear with respect to the scores of the individual sub-systems, the linear function depends on the speech segment. Hence, the overall score fusion model can be regarded as an adaptive model. Experimental results shows that WLLR outperforms LLR by approximately 10% relative for PPRLM system fusion on the NIST 2003 and 2005 language recognition evaluation sets. Khe Chai Sim, Kong-Aik Lee |
ICASSP | 2 |
| 2010 | Approaching human listener accuracy with modern speaker verificationabstractBeing able to recognize people from their voice is a natural ability that we take for granted. Recent advances have shown significant improvement in automatic speaker recognition performance. Besides being able to process large amount of data in a fraction of time required by human, automatic systems are now able to deal with diverse channel effects. The goal of this paper is to examine how state-of-the-art automatic system performs in comparison with human listeners, and to investigate the strategy for human-assisted form of automatic speaker recognition, which is useful in forensic investigation. We set up an experimental protocol using data from the NIST SRE 2008 core set. A total of 36 listeners have participated in the listening experiments from three sites, namely Australia, Finland and Singapore. State-of-the-art automatic system achieved 20 % error rate, whereas fusion of human listeners achieved 22%. 1. Ville Hautamäki, Tomi Kinnunen, Mohaddeseh Nosratighods, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2010 | Incorporating MAP estimation and covariance transform for SVM based speaker recognition
Cheung-Chi Leung, Donglai Zhu, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2010 | The estimation and kernel metric of spectral correlation for text-independent speaker verification
Eryu Wang, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001, Wu Guo, Li-Rong Dai 0001 |
INTERSPEECH | 2 |
| 2010 | A hybrid modeling strategy for GMM-SVM speaker recognition with adaptive relevance factor
Chang Huai You, Haizhou Li 0001, Kong-Aik Lee |
INTERSPEECH | 3 |
| 2010 | MAP estimation of subspace transform for speaker recognition
Donglai Zhu, Bin Ma 0001, Kong-Aik Lee, Cheung-Chi Leung, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2010 | GMM-SVM Kernel With a Bhattacharyya-Based Distance for Speaker RecognitionabstractAmong conventional methods for text-independent speaker recognition, Gaussian mixture model (GMM) is known for its effectiveness and scalability in modeling the spectral distribution of speech. A GMM-supervector characterizes a speaker's voice by the GMM parameters such as the mean vectors, covariance matrices and mixture weights. Besides the first-order statistics, it is generally believed that speaker's cues are partly conveyed by the second-order statistics. In this paper, we introduce a Bhattacharyya-based GMM-distance to measure the distance between two GMM distributions. Subsequently, the GMM-UBM mean interval (GUMI) concept is introduced to derive a GUMI kernel which can be used in conjunction with support vector machine (SVM) for speaker recognition. The GUMI kernel allows us to exploit the speaker's information not only from the mean vectors of GMM but also from the covariance matrices. Moreover, by analyzing the Bhattacharyya-based GMM-distance measure, we extend the Bhattacharyya-based kernel by involving both the mean and covariance statistical dissimilarities. We demonstrate the effectiveness of the new kernel on the National Institute of Standards and Technology (NIST) speaker recognition evaluation (SRE) 2006 dataset. Chang Huai You, Kong-Aik Lee, Haizhou Li 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | The I4U system in NIST 2008 speaker recognition evaluationabstractThis paper describes the performance of the I4U speaker recognition system in the NIST 2008 Speaker Recognition Evaluation. The system consists of seven subsystems, each with different cepstral features and classifiers. We describe the I4U Primary system and report on its core test results as they were submitted, which were among the best-performing submissions. The I4U effort was led by the Institute for Infocomm Research, Singapore (IIR), with contributions from the University of Science and Technology of China (USTC), the University of New South Wales, Australia (UNSW), Nanyang Technological University, Singapore (NTU) and Carnegie Mellon University, USA (CMU). Haizhou Li 0001, Bin Ma 0001, Kong-Aik Lee, Hanwu Sun, Donglai Zhu, Khe Chai Sim, Chang Huai You, Rong Tong, Ismo Kärkkäinen, Chien-Lin Huang, Vladimir Pervouchine, Wu Guo, Yijie Li 0001, Li-Rong Dai 0001, Mohaddeseh Nosratighods, Tharmarajah Thiruvaran, Julien Epps, Eliathamby Ambikairajah, Chng Eng Siong, Tanja Schultz, Qin Jin |
ICASSP | 3 |
| 2009 | A GMM supervector Kernel with the Bhattacharyya distance for SVM based speaker recognitionabstractGaussian mixture model (GMM) supervector is one of the effective techniques in text independent speaker recognition. In our previous work, we introduce the GMM-UBM mean interval (GUMI) concept based on the Bhattacharyya distance. Subsequently GUMI kernel was successfully used in conjunction with support vector machine (SVM) for speaker recognition. Besides the first order statistics, it is generally believed that speaker cues are also partly conveyed by second order statistics. In this paper, we extend the Bhattacharyya-based SVM kernel by constructing the supervector with the mean statistical vector and the covariance statistical vector. Comparing with the Kullback-Leibler (KL) kernel, we demonstrate the effectiveness of the new kernel on the 2006 National Institute of Standards and Technology (NIST) speaker recognition evaluation (SRE) dataset. Chang Huai You, Kong-Aik Lee, Haizhou Li 0001 |
ICASSP | 2 |
| 2009 | Target-aware language models for spoken language recognition
Rong Tong, Bin Ma 0001, Haizhou Li 0001, Chng Eng Siong, Kong-Aik Lee |
INTERSPEECH | 5 |
| 2009 | An SVM Kernel With GMM-Supervector Based on the Bhattacharyya Distance for Speaker RecognitionabstractGaussian mixture model (GMM) and support vector machine (SVM) have become popular classifiers in text-independent speaker recognition. A GMM-supervector characterizes a speaker's voice with the parameters of GMM, which include mean vectors, covariance matrices, and mixture weights. GMM-supervector SVM benefits from both GMM and SVM frameworks to achieve the state-of-the-art performance. Conventional Kullback-Leibler (KL) kernel in GMM-supervector SVM classifier limits the adaptation of GMM to mean value and leaves covariance unchanged. In this letter, we introduce the GMM-UBM mean interval (GUMI) concept based on the Bhattacharyya distance. This leads to a new kernel for SVM classifier. Comparing with the KL kernel, the new kernel allows us to exploit the information not only from the mean but also from the covariance. We demonstrate the effectiveness of the new kernel on the 2006 National Institute of Standards and Technology (NIST) speaker recognition evaluation (SRE) dataset. Chang Huai You, Kong-Aik Lee, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 2 |
| 2008 | Spoken Language recognition using support vector machines with generative front-endabstractThis paper introduces a spoken language recognition system with a generative front-end and a discriminative backend. The generative front-end is built upon an ensemble of Gaussian densities. These Gaussian densities are trained to represent elementary speech sound units characterizing a wide variety of languages. We formulate the generative front-end in a form of sequence kernel. This sequence kernel transforms a spoken utterance into a feature vector with its attributes representing the occurrence statistics of the speech sound units. A discriminative support vector machine (SVM) then operates on the feature vectors to make classification decision. The proposed language recognition system demonstrates competitive performance on NIST 1996, 2003 and 2005 LRE corpora. Kong-Aik Lee, Chang Huai You, Haizhou Li 0001 |
ICASSP | 1 |
| 2008 | Characterizing speech utterances for speaker verification with sequence kernel SVMabstractSupport vector machine (SVM) equipped with sequence kernel has been proven to be a powerful technique for speaker verification. A number of sequence kernels have been recently proposed, each being motivated from different perspectives with diverse mathematical derivations. Analytical comparison of kernels becomes difficult. To facilitate such comparisons, we propose a generic structure showing how different levels of cues conveyed by speech utterances, ranging from low-level acoustic features to highlevel speaker cues, are being characterized within a sequence kernel. We then identify the similarities and differences between the popular generalized linear discriminant sequence (GLDS) and GMM supervector kernels, as well as our own probabilistic sequence kernel (PSK). Furthermore, we enhance the PSK in terms of accuracy and computational complexity. The enhanced PSK gives competitive accuracy with the other two kernels. Fusing all the three kernels yields an EER of 4.83 % on the 2006 NIST SRE core test. Index Terms: speaker verification, characteristic vector, support vector machine, sequence kernel Kong-Aik Lee, Chang Huai You, Haizhou Li 0001, Tomi Kinnunen, Donglai Zhu |
INTERSPEECH | 1 |
| 2008 | NIST 2007 Language Recognition Evaluation: From the Perspective of IIR
Haizhou Li 0001, Bin Ma 0001, Kong-Aik Lee, Khe Chai Sim, Hanwu Sun, Rong Tong, Donglai Zhu, Chang Huai You |
PACLIC | 3 |
| 2007 | On Delayless Architecture for the Normalized Subband Adaptive FilterabstractDelayless architecture for the recently proposed normalized subband adaptive filter (NSAF) is described and analyzed in this paper. The NSAF has a unique weight-control mechanism whereby error signals estimated in subbands are used to adapt a fullband filter. In the delayless architecture, we implement the subband weight adaptation in an auxiliary loop, and place only the fullband filter along the input signal path. By so doing, delay due to the filter banks is moved to the auxiliary loop out of the signal path, thereby making the algorithm attractive for applications where excessive signal path delay is intolerable. Simulation results demonstrate that the proposed delayless NSAF outperforms other delayless approaches in terms of convergence rate. Kong-Aik Lee, Woon-Seng Gan |
ICME | 1 |
| 2007 | A GMM-based probabilistic sequence kernel for speaker verificationabstractThis paper describes the derivation of a sequence kernel that transforms speech utterances into probabilistic vectors for classification in an expanded feature space. The sequence kernel is built upon a set of Gaussian basis functions, where half of the basis functions contain speaker specific information while the other half implicates the common characteristics of the competing background speakers. The idea is similar to that in the Gaussian mixture model – universal background model (GMM-UBM) system, except that the Gaussian densities are treated individually in our proposed sequence kernel, as opposed to two mixtures of Gaussian densities in the GMM-UBM system. The motivation is to exploit the individual Gaussian components for better speaker discrimination. Experiments on NIST 2001 SRE corpus show convincing results for the probabilistic sequence kernel approach. Kong-Aik Lee, Chang Huai You, Haizhou Li 0001, Tomi Kinnunen |
INTERSPEECH | 1 |
| 2004 | Improving convergence of the NLMS algorithm using constrained subband updatesabstractWe propose a new design criterion for subband adaptive filters (SAFs). The proposed multiple-constraint optimization criterion is based on the principle of minimal disturbance, where the multiple constraints are imposed on the updated subband filter outputs. Compared to the classical fullband least-mean-square (LMS) algorithm, the subband adaptive filtering algorithm derived from the proposed criterion exhibits faster convergence under colored excitation. Furthermore, the recursive tap-weight adaptation can be expressed in a simple form comparable to that of the normalized LMS (NLMS) algorithm. We also show that the proposed multiple-constraint optimization criterion is related to another known weighted criterion. The efficacy of the proposed criterion and algorithm are examined and validated via mathematical analysis and simulation. Kong-Aik Lee, Woon-Seng Gan |
IEEE Signal Process. Lett. | 1 |