Man-Wai Mak

dblp:16/5121 · DBLP profile ↗
← Back
154ranked-venue papers
25as first author
49since 2021 · last 2026
0000-0001-8854-3760ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 93 · 13 first-author · 31 since 2021Artificial intelligence and machine learning · 88 · 13 first-author · 32 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1
YearPublicationVenuePosition
2026 Class unbiasing for generalization in medical diagnosis
Lishi Zuo, Lu Yi 0001, Youzhi Tu, Man-Wai Mak
Pattern Recognit.4
2025 TrInk: Ink Generation with Transformer Network
abstract
Zezhong Jin, Shubhang Desai, Xu Chen, Biyi Fang, Zhuoyi Huang, Zhe Li, Chong-Xin Gan, Xiao Tu, Man-Wai Mak, Yan Lu, Shujie Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zezhong Jin, Shubhang Desai, Biyi Fang, Zhuoyi Huang, Zhe Li 0030, Chong-Xin Gan, Xiao Tu, Man-Wai Mak, Yan Lu 0001, Shujie Liu 0001
EMNLP9
2025 Grouped Knowledge Distillation with Adaptive Logit Softening for Speaker Recognition
abstract
Recent works suggest that decoupling the information of non-target speakers from that of the target speaker in knowledge distillation (KD) and subsequently emphasizing the former can lead to significant performance improvement. However, a well-trained teacher model typically produces almost zero non-target speaker posteriors with limited contribution to knowledge transfer, resulting in a less effective KD. To address this problem, we advocate a dual-group knowledge distillation framework, wherein the primary group with top-k speaker posteriors captures most of the speaker discrimination knowledge in an utterance. The non-primary group contributes to the KD through a binary classification (distillation) between the primary and non-primary groups. In addition, adaptive logit softening is proposed to adjust the teacher’s and student’s logits in the binary distillation, further facilitating effective knowledge transfer. The proposed method trained with a simple x-vector pipeline obtains an impressive equal error rate of 1.46%, 1.47%, and 2.70% on three VoxCeleb1 test sets, outperforming the state-of-the-art methods with a noticeable margin.
Chong-Xin Gan, Youzhi Tu, Zezhong Jin, Man-Wai Mak, Kong-Aik Lee
ICASSP4
2025 Denoising Student Features with Diffusion Models for Knowledge Distillation in Speaker Verification
abstract
In recent years, there has been a surge in the use of a pre-trained speech model as a feature extractor for speaker verification (SV). To reduce model complexity, researchers transfer knowledge from a pre-trained model to a lightweight student model, enabling the latter to reach a performance level not attainable by conventional methods. However, due to the differences in model capacity, the student features contain more noise. This results in discrepancies between the teacher and student features at the intermediate layers, negatively impacting feature-level knowledge distillation (KD). To address this issue, we employ a diffusion model to denoise the student features for KD (DenoKD). This approach enables more effective feature-level distillation. Our method, trained with a small ECAPA-TDNN, achieved a 13% improvement over the baseline on the VoxCeleb1-O test set. Further more, the DenoKD mechanism is found to be effective for SV on short test utterances.
Zezhong Jin, Youzhi Tu, Zhe Li 0030, Chong-Xin Gan, Man-Wai Mak
ICASSP6
2025 Spectral-Aware Low-Rank Adaptation for Speaker Verification
abstract
Previous research has shown that the principal singular vectors of a pre-trained model’s weight matrices capture critical knowledge. In contrast, those associated with small singular values may contain noise or less reliable information. As a result, the LoRA-based parameter-efficient fine-tuning (PEFT) approach, which does not constrain the use of the spectral space, may not be effective for tasks that demand high representation capacity. In this study, we enhance existing PEFT techniques by incorporating the spectral information of pre-trained weight matrices into the fine-tuning process. We investigate spectral adaptation strategies with a particular focus on the additive adjustment of top singular vectors. This is accomplished by applying singular value decomposition (SVD) to the pre-trained weight matrices and restricting the fine-tuning within the top spectral space. Extensive speaker verification experiments on VoxCeleb1 and CN-Celeb1 demonstrate enhanced tuning performance with the proposed approach. Code is released at https://github.com/lizhepolyu/SpectralFT.
Zhe Li 0030, Man-Wai Mak, Mert Pilanci, Hung-yi Lee, Helen M. Meng
ICASSP2
2025 MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues
abstract
Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate the speech of a specific target speaker from an audio mixture using time-synchronized visual cues. In real-world scenarios, visual cues are not always available due to various impairments, which undermines the stability of AV-TSE. Despite this challenge, humans can maintain attentional momentum over time, even when the target speaker is not visible. In this paper, we introduce the Momentum Multi-modal target Speaker Extraction (MoMuSE), which retains a speaker identity momentum in memory, enabling the model to continuously track the target speaker. Designed for real-time inference, MoMuSE extracts the current speech window with guidance from both visual cues and dynamically updated speaker momentum. Experimental results demonstrate that MoMuSE exhibits significant improvement, particularly in scenarios with severe impairment of visual cues.
Shuai Wang 0016, Kong-Aik Lee, Man-Wai Mak, Haizhou Li 0001
ICME5
2025 Disentangling Speaker and Content in Pre-trained Speech Models with Latent Diffusion for Robust Speaker Verification
Zhe Li 0030, Man-Wai Mak, Jen-Tzung Chien, Mert Pilanci, Zezhong Jin, Helen M. Meng
INTERSPEECH2
2025 IDIR: Identifying and Distilling Informative Relations for Speaker Verification
Chong-Xin Gan, Zhe Li 0030, Zezhong Jin, Man-Wai Mak, Kong-Aik Lee
INTERSPEECH5
2025 Optimizing Pause Context in Fine-Tuning Pre-trained Large Language Models for Dementia Detection
Xiaoquan Ke, Man-Wai Mak, Helen M. Meng
INTERSPEECH2
2025 Bayesian Learning for Domain-Invariant Speaker Verification and Anti-Spoofing
Man-Wai Mak, Johan Rohdin, Kong-Aik Lee, Hynek Hermansky
INTERSPEECH2
2025 The Sub-3Sec Problem: From Text-Independent to Text-Dependent Corpus
Ruichen Zuo, Kong-Aik Lee, Man-Wai Mak
INTERSPEECH4
2025 Leveraging Ordinal Information for Speech-based Depression Classification
Lishi Zuo, Man-Wai Mak
INTERSPEECH2
2025 Wi-Fi CSI fingerprinting-based indoor positioning using deep learning and vector embedding for temporal stability
Josyl Mariela Rocamora Reyes, Ivan Wang-Hei Ho, Man-Wai Mak
Expert Syst. Appl.3
2025 Adversarially adaptive temperatures for decoupled knowledge distillation with applications to speaker verification
abstract
202502 bcch
Zezhong Jin, Youzhi Tu, Chong-Xin Gan, Man-Wai Mak, Kong-Aik Lee
Neurocomputing4
2025 ConFusionformer: Locality-enhanced Conformer through multi-resolution attention fusion for speaker verification
Youzhi Tu, Man-Wai Mak, Kong-Aik Lee, Weiwei Lin 0002
Neurocomputing2
2025 Vector Quantization-Based Counterfactual Augmentation for Speech-Based Depression Detection Under Data Scarcity
abstract
Data scarcity is a common and serious problem in depression detection, often leading to overfitting and bias that degrade the performance of depression detectors. We propose a counterfactual augmentation (CF-aug) framework that generates latent features for speech-based depression detection under data-scarce conditions. The generation method is based on exploring how feature changes affect the outcomes. To this end, we introduce a counterfactual layer to a deep network to transform the representation of the original data to its opposite class, while a group-wise vector quantization module helps the model explore how the changes in vectors (or entries) sampled from codebooks affect the outcome. Experimental results demonstrate that CF-aug can alleviate the overfitting and bias problems caused by data scarcity. Our CF-aug framework achieves competitive performance compared to state-of-the-art methods on two depression datasets. We also demonstrate the potential of CF-aug in other domains and modalities for medical diagnosis under data-scarce settings.
Lishi Zuo, Man-Wai Mak
IEEE J. Biomed. Health Informatics2
2024 Asymmetric Clean Segments-Guided Self-Supervised Learning for Robust Speaker Verification
abstract
Contrastive self-supervised learning (CSL) for speaker verification (SV) has drawn increasing interest recently due to its ability to exploit unlabeled data. Performing data augmentation on raw waveforms, such as adding noise or reverberation, plays a pivotal role in achieving promising results in SV. Data augmentation, however, demands meticulous calibration to ensure intact speaker-specific information, which is difficult to achieve without speaker labels. To address this issue, we introduce a novel framework by incorporating clean and augmented segments into the contrastive training pipeline. The clean segments are repurposed to pair with noisy segments to form additional positive and negative pairs. Moreover, the contrastive loss is weighted to increase the difference between the clean and augmented embeddings of different speakers. Experimental results on Voxceleb1 suggest that the proposed framework can achieve a remarkable 19% improvement over the conventional methods, and it surpasses many existing state-of-the-art techniques.
Chong-Xin Gan, Man-Wai Mak, Weiwei Lin 0002, Jen-Tzung Chien
ICASSP2
2024 Dual Parameter-Efficient Fine-Tuning for Speaker Representation Via Speaker Prompt Tuning and Adapters
abstract
Fine-tuning a pre-trained Transformer model (PTM) for speech applications in a parameter-efficient manner offers the dual benefits of reducing memory and leveraging the rich feature representations in massive unlabeled datasets. However, existing parameter-efficient fine-tuning approaches either adapt the classification head or the whole PTM. The former is unsuitable when the PTM is used as a feature extractor, and the latter does not leverage the different degrees of feature abstraction at different Transformer layers. We propose two solutions to address these limitations. First, we apply speaker prompt tuning to update the task-specific embeddings of a PTM. The tuning enhances speaker feature relevance in the speaker embeddings through the cross-attention between prompt and speaker features. Second, we insert adapter blocks into the Transformer encoders and their outputs. This novel arrangement enables the fine-tuned PTM to determine the most suitable layers to extract relevant information for the downstream task. Extensive speaker verification experiments on Voxceleb and CU-MARVEL demonstrate higher parameter efficiency and better model adaptability of the proposed methods than the existing ones.
Zhe Li 0030, Man-Wai Mak, Helen M. Meng
ICASSP2
2024 Contrastive Speaker Embedding With Sequential Disentanglement
abstract
Contrastive speaker embedding assumes that the contrast between the positive and negative pairs of speech segments is attributed to speaker identity only. However, this assumption is incorrect because speech signals contain not only speaker identity but also linguistic content. In this paper, we propose a contrastive learning framework with sequential disentanglement to remove linguistic content by incorporating a disentangled sequential variational autoencoder (DSVAE) into the conventional SimCLR framework. The DSVAE aims to disentangle speaker factors from content factors in an embedding space so that only the speaker factors are used for constructing a contrastive loss objective. Because content factors have been removed from the contrastive learning, the resulting speaker embeddings will be content-invariant. Experimental results on VoxCeleb1-test show that the proposed method consistently outperforms SimCLR. This suggests that applying sequential disentanglement is beneficial to learning speaker-discriminative embeddings.
Youzhi Tu, Man-Wai Mak, Jen-Tzung Chien
ICASSP2
2024 Promoting Independence of Depression and Speaker Features for Speaker Disentanglement in Speech-Based Depression Detection
abstract
Recent studies have demonstrated the effectiveness of speaker disentanglement in mitigating the interference caused by speaker features in speech-based depression detection. However, the inherent entanglement between depression features and speaker features poses challenges to depression detection. In this study, we propose a mutual information-based speaker-invariant depression detector (MI-SIDD) that aims to promote independence between depression and speaker features to facilitate speaker disentanglement. Specifically, we disentangle the speaker features using a vanilla autoencoder with a well-tuned bottleneck layer and minimize the mutual information between depression and speaker features using a conditional mutual information constraint. Experimental results demonstrate the effectiveness of speaker disentanglement and the promotion of independence between depression and speaker features. Our MI-SIDD model achieves competitive performance compared to state-of-the-art methods on the DAIC-WOZ dataset.
Lishi Zuo, Man-Wai Mak, Youzhi Tu
ICASSP2
2024 Collaborative Contrastive Learning for Hypothesis Domain Adaptation
abstract
Interspeech 2024, 1-5 September 2024, Kos, Greece
Jen-Tzung Chien, I-Ping Yeh, Man-Wai Mak
INTERSPEECH3
2024 MM-NodeFormer: Node Transformer Multimodal Fusion for Emotion Recognition in Conversation
abstract
Interspeech 2024, 1-5 September 2024, Kos, Greece
Man-Wai Mak, Kong-Aik Lee
INTERSPEECH2
2024 W-GVKT: Within-Global-View Knowledge Transfer for Speaker Verification
abstract
Interspeech 2024, 1-5 September 2024, Kos, Greece
Zezhong Jin, Youzhi Tu, Man-Wai Mak
INTERSPEECH3
2024 Self-Supervised Learning with Multi-Head Multi-Mode Knowledge Distillation for Speaker Verification
abstract
Interspeech 2024, 1-5 September 2024, Kos, Greece
Zezhong Jin, Youzhi Tu, Man-Wai Mak
INTERSPEECH3
2024 Parameter-efficient Fine-tuning of Speaker-Aware Dynamic Prompts for Speaker Verification
abstract
Interspeech 2024, 1-5 September 2024, Kos, Greece
Zhe Li 0030, Man-Wai Mak, Hung-yi Lee, Helen M. Meng
INTERSPEECH2
2024 On the Effectiveness of Enrollment Speech Augmentation For Target Speaker Extraction
abstract
Deep learning technologies have significantly advanced the performance of target speaker extraction (TSE) tasks. To enhance the generalization and robustness of these algorithms when training data is insufficient, data augmentation is a commonly adopted technique. Unlike typical data augmentation applied to speech mixtures, this work thoroughly investigates the effectiveness of augmenting the enrollment speech space. We found that for both pretrained and jointly optimized speaker encoders, directly augmenting the enrollment speech leads to consistent performance improvement. In addition to conventional methods such as noise and reverberation addition, we propose a novel augmentation method called self-estimated speech augmentation (SSA). Experimental results on the Libri2Mix test set show that our proposed method can achieve an improvement of up to 2.5 dB.
Shuai Wang 0016, Haizhou Li 0001, Man-Wai Mak, Kong-Aik Lee
SLT5
2024 DITA: DETR with improved queries for end-to-end temporal action detection
Chongkai Lu, Man-Wai Mak
Neurocomputing2
2024 Automatic selection of spoken language biomarkers for dementia detection
Xiaoquan Ke, Man-Wai Mak, Helen M. Meng
Neural Networks2
2024 Contrastive Self-Supervised Speaker Embedding With Sequential Disentanglement
abstract
Contrastive self-supervised learning has been widely used in speaker embedding to address the labeling challenge. Contrastive speaker embedding assumes that the contrast between the positive and negative pairs of speech segments is attributed to speaker identity only. However, this assumption is incorrect because speech signals contain not only speaker identity but also linguistic content. In this paper, we propose a contrastive learning framework with sequential disentanglement to remove linguistic content by incorporating a disentangled sequential variational autoencoder (DSVAE) into the conventional contrastive learning framework. The DSVAE aims to disentangle speaker factors from content factors in an embedding space so that the speaker factors become the main contributor to the contrastive loss. Because content factors have been removed from contrastive learning, the resulting speaker embeddings will be content-invariant. The learned embeddings are also robust to language mismatch. It is shown that the proposed method consistently outperforms the conventional contrastive speaker embedding on the VoxCeleb1 and CN-Celeb datasets. This finding suggests that applying sequential disentanglement is beneficial to learning speaker-discriminative embeddings.
Youzhi Tu, Man-Wai Mak, Jen-Tzung Chien
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Feature Selection and Text Embedding for Detecting Dementia from Spontaneous Cantonese
abstract
Dementia is a severe cognitive impairment that affects the health of older adults and creates a burden on their families and caretakers. This paper analyzes diverse hand-crafted features extracted from spoken languages and selects the most discriminative ones for dementia detection. Recently, the performance of dementia detection has been significantly improved by utilizing Transformer-based models that automatically capture the structural and linguistic properties of spoken languages. We investigate Transformer-based features and propose an end-to-end system for dementia detection. We also explore recent ASR and representation learning frameworks, such as Wav2vec 2.0 and Hubert, for transcribing a Cantonese corpus that contains recordings of older adults describing the rabbit story. We investigate using disfluency patterns (DP) in spontaneous speech to enhance the recognized word sequences for the Transformer-based feature extractor. Results show that fine-tuning the feature extractor using the enhanced word sequences can improve dementia detection performance.
Xiaoquan Ke, Man-Wai Mak, Helen M. Meng
ICASSP2
2023 Discriminative Speaker Representation Via Contrastive Learning with Class-Aware Attention in Angular Space
abstract
The challenges in applying contrastive learning to speaker verification (SV) are that the softmax-based contrastive loss lacks discriminative power and that the hard negative pairs can easily influence learning. To overcome the first challenge, we propose a contrastive learning SV framework incorporating an additive angular margin into the supervised contrastive loss in which the margin improves the speaker representation’s discrimination ability. For the second challenge, we introduce a class-aware attention mechanism through which hard negative samples contribute less significantly to the supervised contrastive loss. We also employed gradient-based multi-objective optimization to balance the classification and contrastive loss. Experimental results on CN-Celeb and Voxceleb1 show that this new learning objective can cause the encoder to find an embedding space that exhibits great speaker discrimination across languages.
Zhe Li 0030, Man-Wai Mak, Helen M. Meng
ICASSP2
2023 Self-supervised Neural Factor Analysis for Disentangling Utterance-level Speech Representations
abstract
Self-supervised learning (SSL) speech models such as wav2vec and HuBERT have demonstrated state-of-the-art performance on automatic speech recognition (ASR) and proved to be extremely useful in low label-resource settings. However, the success of SSL models has yet to transfer to utterance-level tasks such as speaker, emotion, and language recognition, which still require supervised fine-tuning of the SSL models to obtain good performance. We argue that the problem is caused by the lack of disentangled representations and an utterance-level learning objective for these tasks. Inspired by how HuBERT uses clustering to discover hidden acoustic units, we formulate a factor analysis (FA) model that uses the discovered hidden acoustic units to align the SSL features. The underlying utterance-level representations are disentangled using probabilistic inference on the aligned features. Furthermore, the variational lower bound derived from the FA model provides an utterance-level objective, allowing error gradients to be backpropagated to the Transformer layers to learn highly discriminative acoustic units. When used in conjunction with HuBERT's masked prediction training, our models outperform the current best model, WavLM, on all utterance-level non-semantic tasks on the SUPERB benchmark with only 20% of labeled data.
Weiwei Lin 0002, Chenhang He, Man-Wai Mak, Youzhi Tu
ICML3
2023 Integrated and Enhanced Pipeline System to Support Spoken Language Analytics for Screening Neurocognitive Disorders
abstract
24th Annual Conference of the International Speech Communication Association, INTERSPEECH 2023, Dublin, Ireland, August 20-24, 2023
Helen M. Meng, Brian Kan-Wing Mak, Man-Wai Mak, Helene H. Fung, Xianmin Gong, Timothy C. Y. Kwok, Xunying Liu, Vincent C. T. Mok, Patrick C. M. Wong, Jean Woo, Xixin Wu, Ka-Ho Wong, Sean Shensheng Xu, Naijun Zheng, Ranzo Huang, Jiawen Kang 0002, Xiaoquan Ke, Junan Li, Jinchao Li
INTERSPEECH3
2023 Avoiding dominance of speaker features in speech-based depression detection
Lishi Zuo, Man-Wai Mak
Pattern Recognit. Lett.2
2023 Cluster-Guided Unsupervised Domain Adaptation for Deep Speaker Embedding
abstract
Recent studies have shown that pseudo labels can contribute to unsupervised domain adaptation (UDA) for speaker verification. Inspired by the self-training strategies that use an existing classifier to label the unlabeled data for retraining, we propose a cluster-guided UDA framework that labels the target domain data by clustering and combines the labeled source domain data and pseudo-labeled target domain data to train a speaker embedding network. To improve the cluster quality, we train a speaker embedding network dedicated for clustering by minimizing the contrastive center loss. The goal is to reduce the distance between an embedding and its assigned cluster center while enlarging the distance between the embedding and the other cluster centers. Using VoxCeleb2 as the source domain and CN-Celeb1 as the target domain, we demonstrate that the proposed method can achieve an equal error rate (EER) of 8.10% on the CN-Celeb1 evaluation set without using any labels from the target domain. This result outperforms the supervised baseline by 39.6% and is the state-of-the-art UDA performance on this corpus.
Haiquan Mao, Feng Hong 0002, Man-Wai Mak
IEEE Signal Process. Lett.3
2023 Robust Speaker Verification Using Deep Weight Space Ensemble
abstract
Domain shift is one of the most challenging problems in speaker verification. Although numerous methods have been proposed to address domain shift, most approaches optimize the performance of one domain at the sacrifice of the other. As a result, to obtain the best performance, each domain requires a dedicated model. However, deploying multiple models is resource-demanding and impractical, particularly when the deployment domains are not known in advance. Recent studies in deep neural networks (DNNs) suggest that near the low error surface of the DNN's weight space, there exists a linear path connecting a base model and a fine-tuned model. This finding inspires us to combine the strength of the fine-tuned models and the base models to solve challenging SV problems. Specifically, we aim to develop models that can handle 1) mixed text-dependent (TD) and text-independent (TI) speaker verification where the speech content can be either unconstrained or constrained, 2) cross-channel speaker verification where the recording can be 16 kHz high-fidelity microphone speech or 8 kHz telephone speech, and 3) bi-lingual speaker verification where the enrollment and test speech can be one of the two languages. With weight space ensemble, we show that we can substantially improve the tasks mentioned above, with a 39.6% improvement in mixing TD and TI SV, a 17.4% improvement in bi-lingual SV, and an 18.4% improvement in cross-channel SV. Moreover, we show that the weight space ensemble can also enhance the performance in the target domain, thanks to the regularization effect of the interpolation.
Weiwei Lin 0002, Man-Wai Mak
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Model-Agnostic Meta-Learning for Fast Text-Dependent Speaker Embedding Adaptation
abstract
By constraining the lexical content of input speech, text-dependent speaker verification (TD-SV) offers more reliable performance than text-independent speaker verification (TI-SV) when dealing with short utterances. Because speech with constrained lexical content is harder to collect, often TD models are fine-tuned from a TI model using a small target phrase dataset. However, sometimes the target phrase dataset is too tiny for fine-tuning, which is the main obstacle for deploying TD-SV. One solution is to fine-tune the model using medium-size multi-phrase TD data and then deploy the model on the target phrase. Although this strategy does help in some cases, the performance is still sub-optimal because the model is not optimized for the target phrase. Inspired by the recent progress in meta-learning, we propose a three-stage pipeline for adapting a TI model to a TD model for the target phrase. Firstly, a TI model is trained using a large amount of speech data. Then, we use a multi-phrase TD dataset to tune the TI model via model-agnostic meta-learning. Finally, we perform fast adaptation using a small target phrase dataset. Results show that the three-stage pipeline consistently outperforms multi-phrase and target phrase fine-tuning.
Weiwei Lin 0002, Man-Wai Mak
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Robust Speaker Verification Using Population-Based Data Augmentation
abstract
Speaker recognition under environments with a low signal-to-noise ratio (SNR) and high reverberation level has always been challenging. Data augmentation can be applied to simulate the adverse environments that a speaker recognition system may encounter. Typically, the augmentation parameters are manually set. Recently, automatic hyper-parameter optimization using population-based learning has shown promising results. This paper proposes a population-based searching strategy for optimizing the augmentation parameters. We refer to the resulting augmentation as population-based augmentation (PBA). Instead of finding a fixed set of hyper-parameters, PBA learns a scheduler for setting the hyper-parameters. This strategy offers a considerable computation advantage over the grid search. We obtained high-performance augmentation policies using a population of six networks only. With PBA, we achieved an EER of 3.98% on the VOiCES19 evaluation set.
Weiwei Lin 0002, Man-Wai Mak
ICASSP2
2022 Disentangled Speaker Embedding for Robust Speaker Verification
abstract
Entanglement of speaker features and redundant features may lead to poor performance when evaluating speaker verification systems on an unseen domain. To address this issue, we propose an InfoMax domain separation and adaptation network (InfoMax–DSAN) to disentangle the domain-specific features and domain-invariant speaker features based on domain adaptation techniques. A frame-based mutual information neural estimator is proposed to maximize the mutual information between frame-level features and input acoustic features, which can help retain more useful information. Furthermore, we propose adopting triplet loss based on the idea of self-supervised learning to overcome the label mismatch problem. Experimental results on VOiCES Challenge 2019 demonstrate that our proposed method can help learn more discriminative and robust speaker embeddings.
Lu Yi 0001, Man-Wai Mak
ICASSP2
2022 UNet-DenseNet for Robust Far-Field Speaker Verification
abstract
23rd Annual Conference of the International Speech Communication Association, INTERSPEECH 2022, Incheon, Korea, September 18-22, 2022
Zhenke Gao, Man-Wai Mak, Weiwei Lin 0002
INTERSPEECH2
2022 Automatic Selection of Discriminative Features for Dementia Detection in Cantonese-Speaking People
abstract
Interspeech 2022, Incheon, Korea, 18-22 September 2022
Xiaoquan Ke, Man-Wai Mak, Helen M. Meng
INTERSPEECH2
2022 Inter-patient ECG classification with i-vector based unsupervised patient adaptation
Sean Shensheng Xu, Man-Wai Mak, Chunqi Chang
Expert Syst. Appl.2
2022 Mixture Representation Learning for Deep Speaker Embedding
abstract
How to effectively convert a sequence of variable-length acoustic features to a fixed-dimension representation has always been a research focus in speaker recognition. In state-of-the-art speaker recognition systems, the conversion is implemented by concatenating the mean and the standard deviation of a sequence of frame-level features. However, a single mean and a single standard deviation are limited descriptive statistics for an acoustic sequence even with powerful feature extractors such as convolutional neural networks. In this paper, we propose a novel statistics pooling method that can produce more descriptive statistics through a mixture representation. Our approach is inspired by the expectation–maximization (EM) algorithm in Gaussian mixture models (GMMs). Instead of using traditional GMM style alignment, we novelly leverage modern deep learning tools to produce a more powerful mixture representation. The novelty includes: (1) unlike GMMs, the mixture assignments are determined by an attention network instead of the Euclidean distances between the frame-level features and explicit centers; (2) instead of using a single frame as input to the attention network, contextual frames are included to smooth out attention transition; and (3) soft-attention assignments are replaced by hard-attention assignments via the Gumbel-Softmax with straight-through estimators. With the proposed attention mechanism, we obtained a 13.7% relative improvement over vanilla mean and standard deviation pooling in the VOiCES19-eval set.
Weiwei Lin 0002, Man-Wai Mak
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Aggregating Frame-Level Information in the Spectral Domain With Self-Attention for Speaker Embedding
abstract
Most pooling methods in state-of-the-art speaker embedding networks are implemented in the temporal domain. However, due to the high non-stationarity in the feature maps produced from the last frame-level layer, it is not advantageous to use the global statistics (e.g., means and standard deviations) of the temporal feature maps as aggregated embeddings. This motivates us to explore stationary spectral representations and perform aggregation in the spectral domain. In this paper, we propose attentive short-time spectral pooling (attentive STSP) from a Fourier perspective to exploit the local stationarity of the feature maps. In attentive STSP, for each utterance, we compute the spectral representations through a weighted average of the windowed segments within each spectrogram by attention weights and aggregate their lowest spectral components to form the speaker embedding. Because most of the feature map energy is concentrated in the low-frequency region of the spectral domain, attentive STSP facilitates the information aggregation by retaining the low spectral components only. Attentive STSP is shown to consistently outperform attentive pooling on VoxCeleb1, VOiCES19-eval, SRE16-eval, and SRE18-CMN2-eval. This observation suggests that applying segment-level attention and leveraging low spectral components can produce discriminative speaker embeddings.
Youzhi Tu, Man-Wai Mak
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Contrastive Adversarial Domain Adaptation Networks for Speaker Recognition
abstract
Domain adaptation aims to reduce the mismatch between the source and target domains. A domain adversarial network (DAN) has been recently proposed to incorporate adversarial learning into deep neural networks to create a domain-invariant space. However, DAN's major drawback is that it is difficult to find the domain-invariant space by using a single feature extractor. In this article, we propose to split the feature extractor into two contrastive branches, with one branch delegating for the class-dependence in the latent space and another branch focusing on domain-invariance. The feature extractor achieves these contrastive goals by sharing the first and last hidden layers but possessing decoupled branches in the middle hidden layers. For encouraging the feature extractor to produce class-discriminative embedded features, the label predictor is adversarially trained to produce equal posterior probabilities across all of the outputs instead of producing one-hot outputs. We refer to the resulting domain adaptation network as "contrastive adversarial domain adaptation network (CADAN)." We evaluated the embedded features' domain-invariance via a series of speaker identification experiments under both clean and noisy conditions. Results demonstrate that the embedded features produced by CADAN lead to a 33% improvement in speaker identification accuracy compared with the conventional DAN.
Longxin Li, Man-Wai Mak, Jen-Tzung Chien
IEEE Trans. Neural Networks Learn. Syst.2
2022 Improving Speech Emotion Recognition With Adversarial Data Augmentation Network
abstract
When training data are scarce, it is challenging to train a deep neural network without causing the overfitting problem. For overcoming this challenge, this article proposes a new data augmentation network-namely adversarial data augmentation network (ADAN)- based on generative adversarial networks (GANs). The ADAN consists of a GAN, an autoencoder, and an auxiliary classifier. These networks are trained adversarially to synthesize class-dependent feature vectors in both the latent space and the original feature space, which can be augmented to the real training data for training classifiers. Instead of using the conventional cross-entropy loss for adversarial training, the Wasserstein divergence is used in an attempt to produce high-quality synthetic samples. The proposed networks were applied to speech emotion recognition using EmoDB and IEMOCAP as the evaluation data sets. It was found that by forcing the synthetic latent vectors and the real latent vectors to share a common representation, the gradient vanishing problem can be largely alleviated. Also, results show that the augmented data generated by the proposed networks are rich in emotion information. Thus, the resulting emotion classifiers are competitive with state-of-the-art speech emotion recognition systems.
Lu Yi 0001, Man-Wai Mak
IEEE Trans. Neural Networks Learn. Syst.2
2021 A Comparative Study of Acoustic and Linguistic Features Classification for Alzheimer's Disease Detection
abstract
With the global population ageing rapidly, Alzheimer's disease (AD) is particularly prominent in older adults, which has an insidious onset followed by gradual, irreversible deterioration in cognitive domains (memory, communication, etc). Thus the detection of Alzheimer's disease is crucial for timely intervention to slow down disease progression. This paper presents a comparative study of different acoustic and linguistic features for the AD detection using various classifiers. Experimental results on ADReSS dataset reflect that the proposed models using ComParE, X-vector, Linguistics, TFIDF and BERT features are able to detect AD with high accuracy and sensitivity, and are comparable with the state-of-the-art results reported. While most previous work used manual transcripts, our results also indicate that similar or even better performance could be obtained using automatically recognized transcripts over manually collected ones. This work achieves accuracy scores at 0.67 for acoustic features and 0.88 for linguistic features on either manual or ASR transcripts on the ADReSS Challenge1test set.
Jinchao Li, Jianwei Yu 0001, Zi Ye 0001, Simon Wong, Man-Wai Mak, Brian Kan-Wing Mak, Xunying Liu, Helen M. Meng
ICASSP5
2021 Short-Time Spectral Aggregation for Speaker Embedding
abstract
State-of-the-art speaker verification systems take frame-level acoustics features as input and produce fixed-dimensional embeddings as utterance-level representations. Thus, how to aggregate information from frame-level features is vital for achieving high performance. This paper introduces short-time spectral pooling (STSP) for better aggregation of frame-level information. STSP transforms the temporal feature maps of a speaker embedding network into the spectral domain and extracts the lowest spectral components of the averaged spectrograms for aggregation. Benefiting from the low-pass characteristic of the averaged spectrograms, STSP is able to preserve most of the speaker information in the feature maps using a few spectral components only. We show that statistics pooling is a special case of STSP where only the DC spectral components are used. Experiments on VoxCeleb1 and VOiCES 2019 show that STSP outperforms statistics pooling and multi-head attentive pooling, which suggests that leveraging more spectral information in the CNN feature maps can produce highly discriminative speaker embeddings.
Youzhi Tu, Man-Wai Mak
ICASSP2
2021 Mutual Information Enhanced Training for Speaker Embedding
abstract
22nd Annual Conference of the International Speech Communication Association, INTERSPEECH 2021, Brno, Czechia, August 30 - September 3, 2021
Youzhi Tu, Man-Wai Mak
Interspeech2
2020 Gaussian Models for CSI Fingerprinting in Practical Indoor Environment Identification
abstract
It is not uncommon to experience highly dynamic channels in indoor environments due to time-varying signals as well as moving reflectors and scatterers. This greatly affects the performance of wireless sensing systems that use received signal strength indicator (RSSI) and channel state information (CSI) fingerprints for indoor positioning and event detection. Solutions to this dynamic channel problem often involve laborintensive database maintenance and customized hardware. With this, we present Gaussian models that can withstand temporal and environmental dynamics in practical indoor environments using off-the-shelf devices in this paper. Although systems employing Gaussian models have been previously proposed in the literature, most systems use RSSI instead of CSI to represent the wireless channel. By using a Gaussian distribution to model CSI fingerprints, which offer more abundant information regarding the channel dynamics than RSSI, we can exploit the variance inherent in the wireless channels. Our experiments demonstrate that the Gaussian classifier incurs minimal delay of less than 4 seconds and achieves high classification accuracy compared to other techniques. In particular, it achieves up to 50% and 150% performance improvement over the time-reversal resonating strength (TRRS) and the support vector machines (SVM) methods, respectively.
Josyl Mariela B. Rocamora, Ivan Wang-Hei Ho, Man-Wai Mak
GLOBECOM3
2020 Multi-Level Deep Neural Network Adaptation for Speaker Verification Using MMD and Consistency Regularization
abstract
Adapting speaker verification (SV) systems to a new environment is a very challenging task. Current adaptation methods in SV mainly focus on the backend, i.e, adaptation is carried out after the speaker embeddings have been created. In this paper, we present a DNN-based adaptation method using maximum mean discrepancy (MMD). Our method exploits two important aspects neglected by previous research. First, instead of minimizing domain discrepancy at utterance-level alone, our method minimizes domain discrepancy at both frame-level and utterance-level, which we believe will make the adaptation more robust to the duration discrepancy between training data and test data. Second, we introduce a consistency regularization for unlabelled target-domain data. The consistency regularization encourages the target speaker embeddings robust to adverse perturbations. Experiments on NIST SRE 2016 and 2018 show that our DNN adaptation works significantly better than the previously proposed DNN adaptation methods. What's more, our method works well with backend adaptation. By combining the proposed method with backend adaptation, we achieve a 9% improvement over backend adaptation in SRE18.
Weiwei Lin 0002, Man-Wai Mak, Na Li 0012, Dan Su 0002, Dong Yu 0001
ICASSP2
2020 Information Maximized Variational Domain Adversarial Learning for Speaker Verification
abstract
Domain mismatch is a common problem in speaker verification. This paper proposes an information-maximized variational domain adversarial neural network (InfoVDANN) to reduce domain mismatch by incorporating an InfoVAE into domain adversarial training (DAT). DAT aims to produce speaker discriminative and domain-invariant features. The InfoVAE has two roles. First, it performs variational regularization on the learned features so that they follow a Gaussian distribution, which is essential for the standard PLDA backend. Second, it preserves mutual information between the features and the training set to extract extra speaker discriminative information. Experiments on both SRE16 and SRE18-CMN2 show that the InfoVDANN outperforms the recent VDANN, which suggests that increasing the mutual information between the latent features and input features enables the InfoVDANN to extract extra speaker information that is otherwise not possible.
Youzhi Tu, Man-Wai Mak, Jen-Tzung Chien
ICASSP2
2020 Wav2Spk: A Simple DNN Architecture for Learning Speaker Embeddings from Waveforms
abstract
21st Annual Conference of the International Speech Communication Association, INTERSPEECH 2020, 25-29 October 2020, Shanghai, China
Weiwei Lin 0002, Man-Wai Mak
INTERSPEECH2
2020 Strategies for End-to-End Text-Independent Speaker Verification
abstract
21st Annual Conference of the International Speech Communication Association, INTERSPEECH 2020, 25-29 October 2020, Shanghai, China
Weiwei Lin 0002, Man-Wai Mak, Jen-Tzung Chien
INTERSPEECH2
2020 Adversarial Separation and Adaptation Network for Far-Field Speaker Verification
abstract
21st Annual Conference of the International Speech Communication Association, INTERSPEECH 2020, 25-29 October 2020, Shanghai, China
Lu Yi 0001, Man-Wai Mak
INTERSPEECH2
2020 A Framework for Adapting DNN Speaker Embedding Across Languages
abstract
Language mismatch remains a major hindrance to the extensive deployment of speaker verification (SV) systems. Current language adaptation methods in SV mainly rely on linear projection in embedding space; i.e., adaptation is carried out after the speaker embeddings have been created, which underutilizes the powerful representation of deep neural networks. This article proposes a maximum mean discrepancy (MMD) based framework for adapting deep neural network (DNN) speaker embedding across languages, featuring multi-level domain loss, separate batch normalization, and consistency regularization. We refer to the framework as MSC. We show that (1) minimizing domain discrepancy at both frame- and utterance-levels performs significantly better than at utterance-level alone; (2) separating the source-domain data from the target-domain in batch normalization improves adaptation performance; and (3) data augmentation can be utilized in the unlabelled target-domain through consistency regularization. By combining these findings, we achieve an EER of 8.69% and 7.95% in NIST SRE 2016 and 2018, respectively, which are significantly better than the previously proposed DNN adaptation methods. Our framework also works well with backend adaptation. By combining the proposed framework with backend adaptation, we achieve an 11.8% improvement over the backend adaptation in SRE18. When applying our framework to a 121-layer Densenet, we achieved an EER of 7.81% and 7.02% in NIST SRE 2016 and 2018, respectively.
Weiwei Lin 0002, Man-Wai Mak, Na Li 0012, Dan Su 0002, Dong Yu 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Variational Domain Adversarial Learning With Mutual Information Maximization for Speaker Verification
abstract
Domain mismatch is a common problem in speaker verification (SV) and often causes performance degradation. For the system relying on the Gaussian PLDA backend to suppress the channel variability, the performance would be further limited if there is no Gaussianity constraint on the learned embeddings. This paper proposes an information-maximized variational domain adversarial neural network (InfoVDANN) that incorporates an InfoVAE into domain adversarial training (DAT) to reduce domain mismatch and simultaneously meet the Gaussianity requirement of the PLDA backend. Specifically, DAT is applied to produce speaker discriminative and domain-invariant features, while the InfoVAE performs variational regularization on the embedded features so that they follow a Gaussian distribution. Another benefit of the InfoVAE is that it avoids posterior collapse in VAEs by preserving the mutual information between the embedded features and the training set so that extra speaker information can be retained in the features. Experiments on both SRE16 and SRE18-CMN2 show that the InfoVDANN outperforms the recent VDANN, which suggests that increasing the mutual information between the embedded features and input features enables the InfoVDANN to extract extra speaker information that is otherwise not possible.
Youzhi Tu, Man-Wai Mak, Jen-Tzung Chien
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 I-Vector-Based Patient Adaptation of Deep Neural Networks for Automatic Heartbeat Classification
abstract
Automatic classification of electrocardiogram (ECG) signals is important for diagnosing heart arrhythmias. A big challenge in automatic ECG classification is the variation in the waveforms and characteristics of ECG signals among different patients. To address this issue, this paper proposes adapting a patient-independent deep neural network (DNN) using the information in the patient-dependent identity vectors (i-vectors). The adapted networks, namely i-vector adapted patient-specific DNNs (iAP-DNNs), are tuned toward the ECG characteristics of individual patients. For each patient, his/her ECG waveforms are compressed into an i-vector using a factor analysis model. Then, this i-vector is injected into the middle hidden layer of the patient-independent DNN. Stochastic gradient descent is then applied to fine-tune the whole network to form a patient-specific classifier. As a result, the adaptation makes use of not only the raw ECG waveforms from the specific patient but also the compact representation of his/her ECG characteristics through the i-vector. Analysis on the hidden-layer activations shows that by leveraging the information in the i-vectors, the iAP-DNNs are more capable of discriminating normal heartbeats against arrhythmic heartbeats than the networks that use the patient-specific ECG only for the adaptation. Experimental results based on the MIT-BIH database suggest that the iAP-DNNs perform better than existing patient-specific classifiers in terms of various performance measures. In particular, the sensitivity and specificity of the existing methods are all under the receiver operating characteristic curves of the iAP-DNNs.
Sean Shensheng Xu, Man-Wai Mak, Chi-Chung Cheung
IEEE J. Biomed. Health Informatics2
2019 Semi-supervised Nuisance-attribute Networks for Domain Adaptation
abstract
How to overcome the training and test data mismatch in speaker verification systems has been a focus of research recently. In this paper, we propose a semi-supervised nuisance attribute network (SNAN) to reduce the domain mismatch in i-vectors and x-vectors. SNANs are based on the idea of nuisance attribute removal in inter-dataset variability compensation (IDVC). But instead of measuring the domain variability through the dataset means, SNANs use the maximum mean discrepancy (MMD) as part of their loss function, which enables the network to find nuisance directions in which domain variability is measured up to infinite moment. The architecture of SNANs also allows us to incorporate the out-of-domain speaker labels into the semi-supervised training process through the center loss and triplet loss. Using SNANs as a preprocessing step for PLDA training, we achieve a relative improvement of 11.8% in EER on NIST 2016 SRE compared to PLDA without adaptation. We also found that the semi-supervised approach can further improve SNANs' performance.
Weiwei Lin 0002, Man-Wai Mak, Youzhi Tu, Jen-Tzung Chien
ICASSP2
2019 Variational Domain Adversarial Learning for Speaker Verification
abstract
20th Annual Conference of the International Speech Communication Association: Crossroads of Speech and Language, INTERSPEECH 2019, Graz, Austria, 15-19 September 2019
Youzhi Tu, Man-Wai Mak, Jen-Tzung Chien
INTERSPEECH2
2019 Towards End-to-End ECG Classification With Raw Signal Extraction and Deep Neural Networks
abstract
This paper proposes deep learning methods with signal alignment that facilitate the end-to-end classification of raw electrocardiogram (ECG) signals into heartbeat types, i.e., normal beat or different types of arrhythmias. Time-domain sample points are extracted from raw ECG signals, and consecutive vectors are extracted from a sliding time-window covering these sample points. Each of these vectors comprises the consecutive sample points of a complete heartbeat cycle, which includes not only the QRS complex but also the P and T waves. Unlike existing heartbeat classification methods in which medical doctors extract handcrafted features from raw ECG signals, the proposed end-to-end method leverages a deep neural network for both feature extraction and classification based on aligned heartbeats. This strategy not only obviates the need to handcraft the features but also produces optimized ECG representation for heartbeat classification. Evaluations on the MIT-BIH arrhythmia database show that at the same specificity, the proposed patient-independent classifier can detect supraventricular- and ventricular-ectopic beats at a sensitivity that is at least 10% higher than current state-of-the-art methods. More importantly, there is a wide range of operating points in which both the sensitivity and specificity of the proposed classifier are higher than those achieved by state-of-the-art classifiers. The proposed classifier can also perform comparable to patient-specific classifiers, but at the same time enjoys the advantage of patient independence.
Sean Shensheng Xu, Man-Wai Mak, Chi-Chung Cheung
IEEE J. Biomed. Health Informatics2
2018 Patient-Specific Heartbeat Classification Based on I-Vector Adapted Deep Neural Networks
Sean Shensheng Xu, Man-Wai Mak, Chi-Chung Cheung
BIBM2
2018 Unsupervised Domain Adaptation for Gender-Aware PLDA Mixture Models
abstract
Probabilistic linear discriminant analysis (PLDA) is a state-of-art back-end for i-vector based speaker verification. However, this backend is still problematic when (1) the model is deployed to new environment (in-domain) that is very different from the training one (out-of-domain) and (2) there are insufficient labeled data from the new environment. To address these problems, this paper proposes using out-of-domain training data to pre-train a PLDA mixture model and applying the mixture model on the in-domain training data to compute a pairwise score matrix for spectral clustering. The hypothesized speaker labels produced by spectral clustering are then used for re-training the mixture model to fit the new environment. To refine the mixture model, the spectral clustering and re-training processes are repeated a number of times. To make the mixture model amenable to both genders, a deep neural network (DNN) is trained to produce gender posteriors given an i-vector. The gender posteriors then replace the posterior probabilities of the indicator variables in the PLDA mixture model. Evaluations based on NIST 2016 SRE suggest that at the end of the iterative re-training, the PLDA mixture model becomes fully adapted to the new domain. Results also show that the PLDA scores can be readily incorporated into spectral clustering, resulting in high quality speaker clusters that could not be possibly achieved by agglomerative hierarchical clustering.
Longxin Li, Man-Wai Mak
ICASSP2
2018 SNR-Invariant Multitask Deep Neural Networks for Robust Speaker Verification
abstract
A major challenge in speaker verification is to achieve low error rates under noisy environments. We observed that background noise in utterances will not only enlarge the speaker-dependenti-vector clusters but also shift the clusters, with the amount of shift depending on the signal-to-noise ratio (SNR) of the utterances. To overcome this SNR-dependent clustering phenomenon, we propose two deep neural network (DNN) architectures: hierarchical regression DNN (H-RDNN) and multitask DNN (MT-DNN). The H-RDNN is formed by stacking two regression DNNs in which the lower DNN is trained to map noisyi-vectors to their respective speaker-dependent cluster means of cleani-vectors and the upper DNN aims to regularize the outliers that cannot be denoised properly by the lower DNN. The MT-DNN is trained to denoisei-vectors (main task) and classify speakers (auxiliary task). The network leverages the auxiliary task to retain speaker information in the denoisedi-vectors. Experimental results suggest that these two DNN architectures together with the PLDA backend significantly outperform the multicondition PLDA model and mixtures of PLDA, and that multitask learning helps to boost verification performance.
Man-Wai Mak
IEEE Signal Process. Lett.2
2018 Multisource I-Vectors Domain Adaptation Using Maximum Mean Discrepancy Based Autoencoders
abstract
Like many machine learning tasks, the performance of speaker verification (SV) systems degrades when training and test data come from very different distributions. What's more, both training and test data themselves could be composed of heterogeneous subsets. These multisource mismatches are detrimental to SV performance. This paper proposes incorporating maximum mean discrepancy (MMD) into the loss function of autoencoders to reduce these mismatches. MMD is a nonparametric method for measuring the distance between two probability distributions. With a properly chosen kernel, MMD can match up to infinite moments of data distributions. We generalize MMD to measure the discrepancies of multiple distributions. We call the generalized MMD domainwise MMD. Using domainwise MMD as an objective function, we propose two autoencoders, namely nuisance-attribute autoencoder (NAE) and domain-invariant autoencoder (DAE), for multisource i-vector adaptation. NAE encodes the features that cause most of the multisource mismatch measured by domainwise MMD. DAE directly encodes the features that minimize the multisource mismatch. Using these MMD-based autoencoders as a preprocessing step for PLDA training, we achieve a relative improvement of 19.2% EER on the NIST 2016 SRE compared to PLDA without adaptation. We also found that MMD-based autoencoders are more robust to unseen domains. In the domain robustness experiments, MMD-based autoencoders show 6.8% and 5.2% improvements over IDVC on female and male Cantonese speakers, respectively.
Weiwei Lin 0002, Man-Wai Mak, Jen-Tzung Chien
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 DNN-Based Score Calibration With Multitask Learning for Noise Robust Speaker Verification
abstract
This paper proposes and investigates several deep neural network (DNN) based score compensation, transformation, and calibration algorithms for enhancing the noise robustness of i-vector speaker verification systems. Unlike conventional calibration methods where the required score shift is a linear function of SNR or log-duration, the DNN approach learns the complex relationship between the score shifts and the combination of i-vector pairs and uncalibrated scores. Furthermore, with the flexibility of DNNs, it is possible to explicitly train a DNN to recover the clean scores without having to estimate the score shifts. To alleviate the overfitting problem, multitask learning is applied to incorporate auxiliary information such as SNRs and speaker ID of training utterances into the DNN. Experiments on NIST 2012 SRE show that score calibration derived from multitask DNNs can improve the performance of the conventional score-shift approch significantly, especially under noisy conditions.
Zhili Tan, Man-Wai Mak, Brian Kan-Wing Mak
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Denoised Senone I-Vectors for Robust Speaker Verification
abstract
Recently, it has been shown that senone i-vectors, whose posteriors are produced by senone deep neural networks (DNNs), outperform the conventional Gaussian mixture model (GMM) i-vectors in both speaker and language recognition tasks. The success of senone i-vectors relies on the capability of the DNN to incorporate phonetic information into the i-vector extraction process. In this paper, we argue that to apply senone i-vectors in noisy environments, it is important to robustify the phonetically discriminative acoustic features and senone posteriors estimated by the DNN. To this end, we propose a deep architecture formed by stacking a deep belief network on top of a denoising autoencoder (DAE). After backpropagation fine-tuning, the network, referred to as denoising autoencoder-deep neural network (DAE-DNN), facilitates the extraction of robust phonetically-discriminitive bottleneck (BN) features and senone posteriors for i-vector extraction. We refer to the resulting i-vectors as denoised BN-based senone i-vectors. Results on NIST 2012 SRE show that senone i-vectors outperform the conventional GMM i-vectors. More interestingly, the BN features are not only phonetically discriminative, results suggest that they also contain sufficient speaker information to produce BN-based senone i-vectors that outperform the conventional senone i-vectors. This work also shows that DAE training is more beneficial to BN feature extraction than senone posterior estimation.
Zhili Tan, Man-Wai Mak, Brian Kan-Wing Mak, Yingke Zhu
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 The I4U Mega Fusion and Collaboration for NIST Speaker Recognition Evaluation 2016
abstract
18th Annual Conference of the International Speech Communication Association, INTERSPEECH 2017, Stockholm, Sweden, 20-24 August 2017
Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Anthony Larcher, Andreas Nautsch, Themos Stafylakis, Gang Liu 0001, Mickael Rouvier, Wei Rao 0002, Federico Alegre, Man-Wai Mak, Achintya Kumar Sarkar, Héctor Delgado, Rahim Saeidi, Hagai Aronowitz, Aleksandr Sizov, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Bin Ma 0001, Ville Vestman, Md. Sahidullah, M. Halonen, Anssi Kanervisto, Gaël Le Lan, Fahimeh Bahmaninezhad, Sergey Isadskiy, Christian Rathgeb, Christoph Busch 0001, Georgios Tzimiropoulos, Q. Qian, Q. Zhao, J. Xue, R. Jin, T. Zhao, Pierre-Michel Bousquet, Moez Ajili, Waad Ben Kheder, Driss Matrouf, Zhi Hao Lim, Chenglin Xu, Haihua Xu 0001, Chng Eng Siong, Benoit G. B. Fauve, Kaavya Sriskandaraja, Vidhyasaharan Sethu, W. W. Lin, Dennis Alexander Lehmann Thomsen, Zheng-Hua Tan, Massimiliano Todisco, Nicholas W. D. Evans, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Eliathamby Ambikairajah
INTERSPEECH13
2017 i-Vector DNN Scoring and Calibration for Noise Robust Speaker Verification
abstract
18th Annual Conference of the International Speech Communication Association, INTERSPEECH 2017, Stockholm, Sweden, 20-24 August 2017
Zhili Tan, Man-Wai Mak
INTERSPEECH2
2017 FUEL-mLoc: feature-unified prediction and explanation of multi-localization of cellular proteins in multiple organisms
abstract
Although many web-servers for predicting protein subcellular localization have been developed, they often have the following drawbacks: (i) lack of interpretability or interpreting results with heterogenous information which may confuse users; (ii) ignoring multi-location proteins and (iii) only focusing on specific organism. To tackle these problems, we present an interpretable and efficient web-server, namely FUEL-mLoc, using eature- nified prediction and xplanation of m ulti- oc alization of cellular proteins in multiple organisms. Compared to conventional localization predictors, FUEL-mLoc has the following advantages: (i) using unified features (i.e. essential GO terms) to interpret why a prediction is made; (ii) being capable of predicting both single- and multi-location proteins and (iii) being able to handle proteins of multiple organisms, including Eukaryota, Homo sapiens, Viridiplantae, Gram-positive Bacteria, Gram-negative Bacteria and Virus . Experimental results demonstrate that FUEL-mLoc outperforms state-of-the-art subcellular-localization predictors. Availability and Implementation: http://bioinfo.eie.polyu.edu.hk/FUEL-mLoc/. Contacts: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Shibiao Wan, Man-Wai Mak, Sun-Yuan Kung
Bioinform.2
2017 Discriminative subspace modeling of SNR and duration variabilities for robust speaker verification
Na Li 0012, Man-Wai Mak, Weiwei Lin 0002, Jen-Tzung Chien
Comput. Speech Lang.2
2017 Fast scoring for PLDA with uncertainty propagation via i-vector grouping
Weiwei Lin 0002, Man-Wai Mak, Jen-Tzung Chien
Comput. Speech Lang.2
2017 DNN-Driven Mixture of PLDA for Robust Speaker Verification
abstract
The mismatch between enrollment and test utterances due to different types of variabilities is a great challenge in speaker verification. Based on the observation that the SNR-level variability or channel-type variability causes heterogeneous clusters in i-vector space, this paper proposes to apply supervised learning to drive or guide the learning of probabilistic linear discriminant analysis (PLDA) mixture models. Specifically, a deep neural network (DNN) is trained to produce the posterior probabilities of different SNR levels or channel types given i-vectors as input. These posteriors then replace the posterior probabilities of indicator variables in the mixture of PLDA. The discriminative training causes the mixture model to perform more reasonable soft divisions of the i-vector space as compared to the conventional mixture of PLDA. During verification, given a test i-vector and a target-speaker's i-vector, the marginal likelihood for the same-speaker hypothesis is obtained by summing the component likelihoods weighted by the component posteriors produced by the DNN, and likewise for the different-speaker hypothesis. Results based on NIST 2012 SRE demonstrate that the proposed scheme leads to better performance under more realistic situations where both training and test utterances cover a wide range of SNRs and different channel types. Unlike the previous SNR-dependent mixture of PLDA which only focuses on SNR mismatch, the proposed model is more general and is potentially applicable to addressing different types of variability in speech.
Na Li 0012, Man-Wai Mak, Jen-Tzung Chien
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Transductive Learning for Multi-Label Protein Subchloroplast Localization Prediction
abstract
Predicting the localization of chloroplast proteins at the sub-subcellular level is an essential yet challenging step to elucidate their functions. Most of the existing subchloroplast localization predictors are limited to predicting single-location proteins and ignore the multi-location chloroplast proteins. While recent studies have led to some multi-location chloroplast predictors, they usually perform poorly. This paper proposes an ensemble transductive learning method to tackle this multi-label classification problem. Specifically, given a protein in a dataset, its composition-based sequence information and profile-based evolutionary information are respectively extracted. These two kinds of features are respectively compared with those of other proteins in the dataset. The comparisons lead to two similarity vectors which are weighted-combined to constitute an ensemble feature vector. A transductive learning model based on the least squares and nearest neighbor algorithms is proposed to process the ensemble features. We refer to the resulting predictor to as EnTrans-Chlo. Experimental results on a stringent benchmark dataset and a novel dataset demonstrate that EnTrans-Chlo significantly outperforms state-of-the-art predictors and particularly gains more than 4% (absolute) improvement on the overall actual accuracy. For readers' convenience, EnTrans-Chlo is freely available online at http://bioinfo.eie.polyu.edu.hk/EnTransChloServer/.
Shibiao Wan, Man-Wai Mak, Sun-Yuan Kung
IEEE ACM Trans. Comput. Biol. Bioinform.2
2016 SNR-invariant PLDA with multiple speaker subspaces
abstract
To deal with the mismatch between the enrollment and test utterances caused by noise with different signal-to-noise ratios (SNR), we have recently proposed an SNR-invariant PLDA model for robust speaker verification. In the model, SNR-specific information were separated from speaker-specific information through marginalizing out the SNR factors during the scoring process. However, this modeling approach assumes that speaker variabilities can be captured by a single speaker subspace regardless of the noise level of the utterances. We will show in this paper that i-vectors extracted from utterances with different noise levels will shift to different regions of the i-vector space and that i-vectors extracted from utterances having similar SNR tend to cluster together. In view of this observation, we propose introducing multiple speaker subspaces to the SNR-invariance PLDA model and use multiple covariance matrices to represent SNR-dependent channel variability. Through NIST 2012 SRE, this paper demonstrates that this finer and more precise modeling of speaker and SNR variabilities leads to better performance when compared with the conventional PLDA and SNR-invariant PLDA.
Na Li 0012, Man-Wai Mak
ICASSP2
2016 Deep neural network driven mixture of PLDA for robust i-vector speaker verification
abstract
In speaker recognition, the mismatch between the enrollment and test utterances due to noise with different signal-to-noise ratios (SNRs) is a great challenge. Based on the observation that noise-level variability causes the i-vectors to form heterogeneous clusters, this paper proposes using an SNR-aware deep neural network (DNN) to guide the training of PLDA mixture models. Specifically, given an i-vector, the SNR posterior probabilities produced by the DNN are used as the posteriors of indicator variables of the mixture model. As a result, the proposed model provides a more reasonable soft division of the i-vector space compared to the conventional mixture of PLDA. During verification, given a test trial, the marginal likelihoods from individual PLDA models are linearly combined by the posterior probabilities of SNR levels computed by the DNN. Experimental results for SNR mismatch tasks based on NIST 2012 SRE suggest that the proposed model is more effective than PLDA and conventional mixture of PLDA for handling heterogeneous corpora.
Na Li 0012, Man-Wai Mak, Jen-Tzung Chien
SLT2
2016 Sparse regressions for predicting and interpreting subcellular localization of multi-label proteins
abstract
BACKGROUND: Predicting protein subcellular localization is indispensable for inferring protein functions. Recent studies have been focusing on predicting not only single-location proteins, but also multi-location proteins. Almost all of the high performing predictors proposed recently use gene ontology (GO) terms to construct feature vectors for classification. Despite their high performance, their prediction decisions are difficult to interpret because of the large number of GO terms involved. RESULTS: This paper proposes using sparse regressions to exploit GO information for both predicting and interpreting subcellular localization of single- and multi-location proteins. Specifically, we compared two multi-label sparse regression algorithms, namely multi-label LASSO (mLASSO) and multi-label elastic net (mEN), for large-scale predictions of protein subcellular localization. Both algorithms can yield sparse and interpretable solutions. By using the one-vs-rest strategy, mLASSO and mEN identified 87 and 429 out of more than 8,000 GO terms, respectively, which play essential roles in determining subcellular localization. More interestingly, many of the GO terms selected by mEN are from the biological process and molecular function categories, suggesting that the GO terms of these categories also play vital roles in the prediction. With these essential GO terms, not only where a protein locates can be decided, but also why it resides there can be revealed. CONCLUSIONS: Experimental results show that the output of both mEN and mLASSO are interpretable and they perform significantly better than existing state-of-the-art predictors. Moreover, mEN selects more features and performs better than mLASSO on a stringent human benchmark dataset. For readers' convenience, an online server called SpaPredictor for both mLASSO and mEN is available at http://bioinfo.eie.polyu.edu.hk/SpaPredictorServer/.
Shibiao Wan, Man-Wai Mak, Sun-Yuan Kung
BMC Bioinform.2
2016 Sparse kernel machines with empirical kernel maps for PLDA speaker verification
Wei Rao 0002, Man-Wai Mak
Comput. Speech Lang.2
2016 Robust scream sound detection via sound event partitioning
Bai Ying Lei, Man-Wai Mak
Multim. Tools Appl.2
2016 Mixture of PLDA for Noise Robust I-Vector Speaker Verification
abstract
In real-world environments, noisy utterances with variable noise levels are recorded and then converted to i-vectors for cosine distance or PLDA scoring. This paper investigates the effect of noise-level variability on i-vectors. It demonstrates that noise-level variability causes the i-vectors to shift, causing the noise contaminated i-vectors to form clusters in the i-vector space. It also demonstrates that optimal subspaces for discriminating speakers are noise-level dependent. Based on these observations, this paper proposes using signal-to-noise ratio (SNR) of utterances as guidance for training mixture of PLDA models. To maximize the coordination among the PLDA models, mixtures of PLDA models are trained simultaneously via an EM algorithm using the utterances contaminated with noise at various levels. For scoring, given a test i-vector, the marginal likelihoods from individual PLDA models are linearly combined by the posterior probabilities of the test utterance's SNR. Verification scores are the ratio of the marginal likelihoods. Results based on NIST 2012 SRE suggest that the SNR-dependent mixture of PLDA is not only suitable for the situations where the test utterances exhibit a wide range of SNR, but also beneficial for the test utterances with unknown SNR distribution. Supplementary materials containing full derivations of the EM algorithms and scoring functions can be found in http://bioinfo.eie.polyu.edu.hk/mPLDA/SuppMaterials.pdf.
Man-Wai Mak, Xiaomin Pang, Jen-Tzung Chien
IEEE ACM Trans. Audio Speech Lang. Process.1
2016 Mem-mEN: Predicting Multi-Functional Types of Membrane Proteins by Interpretable Elastic Nets
abstract
Membrane proteins play important roles in various biological processes within organisms. Predicting the functional types of membrane proteins is indispensable to the characterization of membrane proteins. Recent studies have extended to predicting single- and multi-type membrane proteins. However, existing predictors perform poorly and more importantly, they are often lack of interpretability. To address these problems, this paper proposes an efficient predictor, namely Mem-mEN, which can produce sparse and interpretable solutions for predicting membrane proteins with single- and multi-label functional types. Given a query membrane protein, its associated gene ontology (GO) information is retrieved by searching a compact GO-term database with its homologous accession number, which is subsequently classified by a multi-label elastic net (EN) classifier. Experimental results show that Mem-mEN significantly outperforms existing state-of-the-art membrane-protein predictors. Moreover, by using Mem-mEN, 338 out of more than 7,900 GO terms are found to play more essential roles in determining the functional types. Based on these 338 essential GO terms, Mem-mEN can not only predict the functional type of a membrane protein, but also explain why it belongs to that type. For the reader's convenience, the Mem-mEN server is available online at http://bioinfo.eie.polyu.edu.hk/MemmENServer/.
Shibiao Wan, Man-Wai Mak, Sun-Yuan Kung
IEEE ACM Trans. Comput. Biol. Bioinform.2
2015 Normalization of total variability matrix for i-vector/PLDA speaker verification
abstract
Gaussian PLDA with uncertainty propagation is effective for i-vector based speaker verification. The idea is to propagate the uncertainty of i-vectors caused by the duration variability of utterances to the PLDA model. However, a limitation of the method is the difficulty of performing length normalization on the posterior covariance matrix of an i-vector. This paper proposes a method to avoid performing length normalization on i-vectors in Gaussian PLDA modeling so that uncertainty propagation can be directly applied without transforming the posterior covariance matrices of i-vectors. Instead of performing length normalization on i-vectors independently, the proposed method normalizes the column vectors of the total variability matrix. Because the i-vectors of all utterances are derived from the same normalized total variability matrix, they will be subject to the same degree of normalization, thereby avoiding the undesirable distortion introduced by the utterance-dependent length-normalization process. Experimental results on both NIST 2010 and 2012 SREs demonstrate that the proposed method achieves a performance similar to (and in some situations better than) that of Gaussian PLDA with length normalization. The method has the potential of improving the performance of uncertainty propagation for i-vector/PLDA speaker verification.
Wei Rao 0002, Man-Wai Mak, Kong-Aik Lee
ICASSP2
2015 SNR-invariant PLDA modeling for robust speaker verification
abstract
16th Annual Conference of the International Speech Communication Association, INTERSPEECH 2015, Dresden, Germany, September 6-10, 2015
Na Li 0012, Man-Wai Mak
INTERSPEECH2
2015 SNR-Invariant PLDA Modeling in Nonparametric Subspace for Robust Speaker Verification
abstract
While i-vector/PLDA framework has achieved great success, its performance still degrades dramatically under noisy conditions. To compensate for the variability of i-vectors caused by different levels of background noise, this paper proposes an SNR-invariant PLDA framework for robust speaker verification. First, nonparametric feature analysis (NFA) is employed to suppress intra-speaker variation and emphasize the discriminative information inherited in the boundaries between speakers in the i-vector space. Then, in the NFA-projected subspace, SNR-invariant PLDA is applied to separate the SNR-specific information from speaker-specific information using an identity factor and an SNR factor. Accordingly, a projected i-vector in the NFA subspace can be represented as a linear combination of three components: speaker, SNR, and channel. During verification, the variability due to SNR and channels are integrated out when computing the marginal likelihood ratio. Experiments based on NIST 2012 SRE show that the proposed framework achieves superior performance when compared with the conventional PLDA and SNR-dependent mixture of PLDA.
Na Li 0012, Man-Wai Mak
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Construction of discriminative Kernels from known and unknown non-targets for PLDA-SVM scoring
abstract
Conventional PLDA scoring in i-vector speaker verification involves the i-vectors of target speakers and claimants only. We have previously demonstrated that better performance can be achieved by incorporating the information of background speakers in the scoring process via speaker-dependent SVMs. This is achieved by defining a PLDA score space with dimension equal to the number of training i-vectors for each target speaker. The new protocol in NIST 2012 SRE permits systems to use the information of other target-speakers (called known non-targets) in each verification trial. In this paper, we exploit this new protocol to enhance the performance of PLDA-SVM scoring by using the score vectors of both known and unknown non-targets as the impostor class data to train the speaker-dependent SVMs. Because some target speakers have one enrollment utterance only, which results in severe imbalance in the speaker- and impostor-class data for SVM training. This paper shows that if the enrollment utterance is sufficiently long, a number of target-speaker i-vectors can be generated by an utterance partitioning and resampling technique, resulting in much better scoring SVMs. Results on NIST 2012 SRE demonstrate the advantages of pooling the known and unknown non-targets for training the SVMs and that the resampling techniques can help the SVM training algorithm to find better decision boundaries for those speakers with only a small number of enrollment utterances.
Wei Rao 0002, Man-Wai Mak
ICASSP2
2014 Ensemble random projection for multi-label classification with application to protein subcellular localization
abstract
The curse of dimensionality severely restricts the predictive power of multi-label classification systems. High-dimensional feature vectors may contain redundant or irrelevant information, causing the classification systems suffer from overfitting. To address this problem, this paper proposes a dimensionality-reduction method that applies random projection (RP) to construct an ensemble of multilabel classifiers. The merits of the proposed method are demonstrated through a multi-label protein classification task. Specifically, high-dimensional feature vectors are extracted from protein sequences using the gene ontology (GO) and Swiss-Prot databases. The feature vectors are then projected onto lower-dimensional spaces by random projection matrices whose elements conform to a distribution with zero mean and unit variance. The transformed low-dimensional vectors are classified by an ensemble of one-vs-rest multi-label support vector machine (SVM) classifiers, each corresponding to one of the RP matrices. The scores obtained from the ensemble are then fused for predicting the subcellular localization of proteins. Experimental results suggest that the proposed method can reduce the dimensions by seven folds and impressively improve the classification performance.
Shibiao Wan, Man-Wai Mak, Bai Zhang, Yue Joseph Wang, Sun-Yuan Kung
ICASSP2
2014 SNR-dependent mixture of PLDA for noise robust speaker verification
abstract
15th Annual Conference of the International Speech Communication Association: Celebrating the Diversity of Spoken Languages, INTERSPEECH 2014, 14-18 September 2014
Man-Wai Mak
INTERSPEECH1
2014 PLDA modeling in the fishervoice subspace for speaker verification
abstract
15th Annual Conference of the International Speech Communication Association: Celebrating the Diversity of Spoken Languages, INTERSPEECH 2014, 14-18 September 2014
Jinghua Zhong, Weiwu Jiang, Wei Rao 0002, Man-Wai Mak, Helen M. Meng
INTERSPEECH4
2014 A study of voice activity detection techniques for NIST speaker recognition evaluations
Man-Wai Mak, Hon-Bill Yu
Comput. Speech Lang.1
2013 An ensemble classifier with random projection for predicting multi-label protein subcellular localization
abstract
In protein subcellular localization prediction, a predominant scenario is that the number of available features is much larger than the number of data samples. Among the large number of features, many of them may contain redundant or irrelevant information, causing the prediction systems suffer from overfitting. To address this problem, this paper proposes a dimensionality-reduction method that applies random projection (RP) to construct an ensemble multi-label classifier for predicting protein subcellular localization. Specifically, the frequencies of occurrences of gene-ontology terms are used as feature vectors, which are projected onto lower-dimensional spaces by random projection matrices whose elements conform to a distribution with zero mean and unit variance. The transformed low-dimensional vectors are classified by an ensemble of one-vs-rest multi-label support vector machine (SVM) classifiers, each corresponding to one of the RP matrices. The scores obtained from the ensemble are then fused for making the final decision. Experimental results on two recent datasets suggest that the proposed method can reduce the dimensions by six folds and remarkably improve the classification performance.
Shibiao Wan, Man-Wai Mak, Bai Zhang, Yue Joseph Wang, Sun-Yuan Kung
BIBM2
2013 Likelihood-ratio empirical kernels for i-vector based PLDA-SVM scoring
abstract
Likelihood ratio (LR) scoring in PLDA speaker verification systems only uses the information of background speakers implicitly. This paper exploits the notion of empirical kernel maps to incorporate background speaker information into the scoring process explicitly. This is achieved by training a scoring SVM for each target speaker based on a kernel in the empirical feature space. More specially, given a test i-vector and the identity of the target under test, a score vector is constructed by computing the LR scores of the test i-vector with respect to the target-speaker's i-vectors and a set of background-speakers' i-vectors. While in most situations, only one target-speaker i-vector is available for training the SVM, this paper demonstrates that if the enrollment utterance is sufficiently long, a number of target-speaker i-vectors can be generated by an utterance partitioning and resampling technique, resulting in much better scoring SVMs. Results on NIST 2010 SRE suggests that the idea of incorporating background speaker information into PLDA scoring through training speaker-dependent SVMs together with the utterance partitioning techniques can boost the performance of i-vector based PLDA systems significantly.
Man-Wai Mak, Wei Rao 0002
ICASSP1
2013 Adaptive thresholding for multi-label SVM classification with application to protein subcellular localization prediction
abstract
Multi-label classification has received increasing attention in computational proteomics, especially in protein subcellular localization. Many existing multi-label protein predictors suffer from over-prediction because they use a fixed decision threshold to determine the number of labels to which a query protein should be assigned. To address this problem, this paper proposes an adaptive thresholding scheme for multi-label support vector machine (SVM) classifiers. Specifically, each one-vs-rest SVM has an adaptive threshold that is a fraction of the maximum score of the one-vs-rest SVMs in the classifier. Therefore, the number of class labels of the query protein depends on the confidence of the SVMs in the classification. This scheme is integrated into our recently proposed subcellular localization predictor that uses the frequency of occurrences of gene-ontology terms as feature vectors and one-vs-rest SVMs as classifiers. Experimental results on two recent datasets suggest that the scheme can effectively avoid both over-prediction and under-prediction, resulting in performance significantly better than other gene-ontology based subcellular localization predictors.
Shibiao Wan, Man-Wai Mak, Sun-Yuan Kung
ICASSP2
2013 Boosting the Performance of I-Vector Based Speaker Verification via Utterance Partitioning
abstract
The success of the recent i-vector approach to speaker verification relies on the capability of i-vectors to capture speaker characteristics and the subsequent channel compensation methods to suppress channel variability. Typically, given an utterance, an i-vector is determined from the utterance regardless of its length. This paper investigates how the utterance length affects the discriminative power of i-vectors and demonstrates that the discriminative power of i-vectors reaches a plateau quickly when the utterance length increases. This observation suggests that it is possible to make the best use of a long conversation by partitioning it into a number of sub-utterances so that more i-vectors can be produced for each conversation. To increase the number of sub-utterances without scarifying the representation power of the corresponding i-vectors, repeated applications of frame-index randomization and utterance partitioning are performed. Results on NIST 2010 speaker recognition evaluation (SRE) suggest that (1) using more i-vectors per conversation can help to find more robust linear discriminant analysis (LDA) and within-class covariance normalization (WCCN) transformation matrices, especially when the number of conversations per training speaker is limited; and (2) increasing the number of i-vectors per target speaker helps the i-vector based support vector machines (SVM) to find better decision boundaries, thus making SVM scoring outperforms cosine distance scoring by 19% and 9% in terms of minimum normalized DCF and EER.
Wei Rao 0002, Man-Wai Mak
IEEE Trans. Speech Audio Process.2
2012 Low-power SVM classifiers for sound event classification on mobile devices
abstract
With the high processing power of today's smartphones, it becomes possible to turn a smartphone into a personal audio surveillance and monitoring system. Ideally, such a system should be able to detect and classify a variety of sound events 24 hours a day and trigger an emergence phone call or message once a specified sound event (e.g., screaming) occurs. To prolong battery life, it is important to trade off the detection accuracy against power consumption. This paper investigates the power consumption of different stages of a sound-event classification system, including segmentation, feature extraction, and SVM scoring. The performance and power consumption of various acoustic features and SVM kernels are compared. This paper advocates the notion of intrinsic complexity through which the scoring function of polynomial SVMs can be written in a matrix-vector-multiplication form so that the resulting complexity becomes independent of the number of support vectors. Results show that this intrinsic complexity can reduce the CPU utilization of polynomial SVMs by 28 times without reducing classification accuracy.
Man-Wai Mak, Sun-Yuan Kung
ICASSP1
2012 GOASVM: Protein subcellular localization prediction based on Gene ontology annotation and SVM
abstract
Protein subcellular localization is an essential step to annotate proteins and to design drugs. This paper proposes a functional-domain based method-GOASVM-by making full use of Gene Ontology Annotation (GOA) database to predict the subcellular locations of proteins. GOASVM uses the accession number (AC) of a query protein and the accession numbers (ACs) of homologous proteins returned from PSI-BLAST as the query strings to search against the GOA database. The occurrences of a set of predefined GO terms are used to construct the GO vectors for classification by support vector machines (SVMs). The paper investigated two different approaches to constructing the GO vectors. Experimental results suggest that using the ACs of homologous proteins as the query strings can achieve an accuracy of 94.68%, which is significantly higher than all published results based on the same dataset. As a user-friendly web-server, GOASVM is freely accessible to the public at http://bioinfo.eie.polyu.edu.hk/mGoaSvmServer/GOASVM.html.
Shibiao Wan, Man-Wai Mak, Sun-Yuan Kung
ICASSP2
2012 mGOASVM: Multi-label protein subcellular localization based on gene ontology and support vector machines
abstract
BACKGROUND: Although many computational methods have been developed to predict protein subcellular localization, most of the methods are limited to the prediction of single-location proteins. Multi-location proteins are either not considered or assumed not existing. However, proteins with multiple locations are particularly interesting because they may have special biological functions, which are essential to both basic research and drug discovery. RESULTS: This paper proposes an efficient multi-label predictor, namely mGOASVM, for predicting the subcellular localization of multi-location proteins. Given a protein, the accession numbers of its homologs are obtained via BLAST search. Then, the original accession number and the homologous accession numbers of the protein are used as keys to search against the Gene Ontology (GO) annotation database to obtain a set of GO terms. Given a set of training proteins, a set of T relevant GO terms is obtained by finding all of the GO terms in the GO annotation database that are relevant to the training proteins. These relevant GO terms then form the basis of a T-dimensional Euclidean space on which the GO vectors lie. A support vector machine (SVM) classifier with a new decision scheme is proposed to classify the multi-label GO vectors. The mGOASVM predictor has the following advantages: (1) it uses the frequency of occurrences of GO terms for feature representation; (2) it selects the relevant GO subspace which can substantially speed up the prediction without compromising performance; and (3) it adopts an efficient multi-label SVM classifier which significantly outperforms other predictors. Briefly, on two recently published virus and plant datasets, mGOASVM achieves an actual accuracy of 88.9% and 87.4%, respectively, which are significantly higher than those achieved by the state-of-the-art predictors such as iLoc-Virus (74.8%) and iLoc-Plant (68.1%). CONCLUSIONS: mGOASVM can efficiently predict the subcellular locations of multi-label proteins. The mGOASVM predictor is available online at http://bioinfo.eie.polyu.edu.hk/mGoaSvmServer/mGOASVM.html.
Shibiao Wan, Man-Wai Mak, Sun-Yuan Kung
BMC Bioinform.2
2011 The HKCUPU system for the NIST 2010 speaker recognition evaluation
abstract
This paper presents the HKCUPU speaker recognition system submitted to NIST 2010 speaker recognition evaluation (SRE). The system comprises five subsystems, each with different acoustic features, session-variability reduction methods, speaker modeling and scoring methods and classifiers. This paper reports the results of individual and fusion systems for the core test and highlights the improvements made by our newly proposed JFA-Fishervoice (FSH) subsystem. Results show that FSH outperforms JFA when its projection matrix is channel-dependent (telephone or microphone) and that FSH is complementary to other state-of-the-art techniques. It was also found that VAD is an important pre-processing step for interview speech.
Weiwu Jiang, Man-Wai Mak, Wei Rao 0002, Helen M. Meng
ICASSP2
2011 Addressing the Data-Imbalance Problem in Kernel-Based Speaker Verification via Utterance Partitioning and Speaker Comparison
abstract
GMM-SVM has become a promising approach to textindependent speaker verification. However, a problematic issue of this approach is the extremely serious imbalance between the numbers of speaker-class and impostor-class utterances available for training the speaker-dependent SVMs. This data-imbalance problem can be addressed by (1) creating more speaker-class supervectors for SVM training through utterance partitioning with acoustic vector resampling (UP-AVR) and (2) avoiding the SVM training so that speaker scores are formulated as an inner product discriminant function (IPDF) between the target-speaker’s supervector and test supervector. This paper highlights the differences between these two approaches and compares the effect of using different kernels – including the KL divergence kernel, GMM-UBM mean interval (GUMI) kernel and geometric-mean-comparison kernel – on their performance. Experiments on the NIST 2010 Speaker Recognition Evaluation suggest that GMM-SVM with UP-AVR is superior to speaker comparison and that the GUMI kernel is slightly better than the KL kernel in speaker comparison. Index Terms: speaker verification, GMM-SVM, speaker comparison, NIST SRE, utterance partitioning, data imbalance.
Wei Rao 0002, Man-Wai Mak
INTERSPEECH2
2011 Comparison of Voice Activity Detectors for Interview Speech in NIST Speaker Recognition Evaluation
abstract
12th Annual Conference of the International Speech Communication Association, INTERSPEECH 2011, Florence, Italy, August 27-31, 2011
Hon-Bill Yu, Man-Wai Mak
INTERSPEECH2
2011 Utterance partitioning with acoustic vector resampling for GMM-SVM speaker verification
Man-Wai Mak, Wei Rao 0002
Speech Commun.1
2011 Optimized Discriminative Kernel for SVM Scoring and Its Application to Speaker Verification
abstract
The decision-making process of many binary classification systems is based on the likelihood ratio (LR) scores of test patterns. This paper shows that LR scores can be expressed in terms of the similarity between the supervectors (SVs) formed by stacking the mean vectors of Gaussian mixture models corresponding to the test patterns, the target model, and the background model. By interpreting the support vector machine (SVM) kernels as a specific similarity (or discriminant) function between SVs, this paper shows that LR scoring is a special case of SVM scoring and that most sequence kernels can be obtained by assuming a specific form for the similarity function of SVs. This paper further shows that this assumption can be relaxed to derive a new general kernel. The kernel function is general in that it is a linear combination of any kernels belonging to the reproducing kernel Hilbert space. The combination weights are obtained by optimizing the ability of a discriminant function to separate the positive and negative classes using either regression analysis or SVM training. The idea was applied to both high-and low-level speaker verification. In both cases, results show that the proposed kernels achieve better performance than several state-of-the-art sequence kernels. Further performance enhancement was also observed when the high-level scores were combined with acoustic scores.
Shixiong Zhang 0001, Man-Wai Mak
IEEE Trans. Neural Networks2
2010 Truncation of protein sequences for fast profile alignment with application to subcellular localization
abstract
We have recently found that the computation time of homology-based subcellular localization can be substantially reduced by aligning profiles up to the cleavage site positions of signal peptides, mitochondrial targeting peptides, and chloro-plast transit peptides [1]. While the method can reduce the profile alignment time by as much as 20 folds, it cannot reduce the computation time spent on creating the profiles. In this paper, we propose a new approach that can reduce both the profile creation time and profile alignment time. In the new approach, instead of cutting the profiles, we shorten the sequences by cutting them at the cleavage site locations. The shortened sequences are then presented to PSI-BLAST to compute the profiles. Experimental results and analysis of profile-alignment score matrices suggest that both profile creation time and profile alignment time can be reduced without sacrificing subcellular localization accuracy. Once a pairwise profile-alignment score matrix has been obtained, a one-vs-rest SVM classifier can be trained. To further reduce the training and recognition time of the classifier, we propose a perturbation discriminant analysis (PDA) technique. It was found that PDA enjoys a short training time as compared to the conventional SVM.
Man-Wai Mak, Sun-Yuan Kung
BIBM1
2010 Speeding up subcellular localization by extracting informative regions of protein sequences for profile alignment
abstract
The functions of proteins are closely related to their subcellular locations. In the post-proteomics era, the amount of gene and protein data grows exponentially, which necessitates the prediction of subcellular localization by computational means. This paper proposes mitigating the computation burden of alignment-based approaches to subcellular localization prediction by using the information provided by the N-terminal sorting signals. To this end, a cascaded fusion of cleavage site prediction and profile alignment is proposed. Specifically, the informative segments of protein sequences are identified by a cleavage site predictor. Then, only the informative segments are applied to a homology-based classifier for predicting the subcellular locations. Experimental results on a newly constructed dataset show that the method can make use of the best property of both approaches and can attain an accuracy higher than using the full-length sequences. Moreover, the method can reduce the computation time by 20 folds. We advocate that the method will be important for biologists to conduct large-scale protein annotation or for bioinformaticians to perform preliminary investigations on new algorithms that involve pairwise alignments.
Man-Wai Mak, Sun-Yuan Kung
CIBCB2
2010 Acoustic vector resampling for GMMSVM-based speaker verification
abstract
Using GMM-supervectors as the input to SVM classifiers (namely, GMM-SVM) is one of the promising approaches to text-independent speaker verification. However, one unaddressed issue of this approach is the severe imbalance between the numbers of speaker-class utterances and impostor-class utterances available for training a speaker-dependent SVM. This paper proposes a resampling technique – namely utterance partitioning with acoustic vector resampling (UP-AVR) – to mitigate the data imbalance problem. Specifically, the sequence order of acoustic vectors in an enrollment utterance is first randomized; then the randomized sequence is partitioned into a number of segments. Each of these segments is then used to produce a GMM-supervector via MAP adaptation and mean vector concatenation. A desirable number of speaker-class supervectors can be produced by repeating this randomization and partitioning process a number of times. Experimental evaluations suggest that UP-AVR can reduce the EER of GMM-SVM systems by about 10%. 1.
Man-Wai Mak, Wei Rao 0002
INTERSPEECH1
2009 Conditional random fields for the prediction of signal peptide cleavage sites
abstract
Correct prediction of signal peptide cleavage sites has a significant impact on drug design. State-of-the-art approaches to cleavage site prediction typically use generative models (such as HMMs) to represent the statistics of amino acid sequences or use neural networks to detect the changes in short amino-acid segments along a query sequence. By formulating cleavage site prediction as a sequence labeling problem, this paper demonstrates how conditional random fields (CRFs) can be applied to cleavage site prediction. The paper also demonstrates how amino acid properties can be exploited and incorporated into the CRFs to boost prediction performance. Results show that the performance of CRFs is comparable to that of a state-of-the-art predictor (SignalP V3.0). Further performance improvement was observed when the decisions of SignalP and the CRF-based predictor are fused.
Man-Wai Mak, Sun-Yuan Kung
ICASSP1
2009 Fast GMM computation for speaker verification using scalar quantization and discrete densities
abstract
Most of current state-of-the-art speaker verification (SV) sys-tems use Gaussian mixture model (GMM) to represent the uni-versal background model (UBM) and the speaker models (SM). For an SV system that employs log-likelihood ratio between SM and UBM to make the decision, its computational effi-ciency is largely determined by the GMM computation. This paper attempts to speedup GMM computation by converting a continuous-density GMM to a single or a mixture of discrete densities using scalar quantization. We investigated a spectrum of such discrete models: from high-density discrete models to discrete mixture models, and their combination called high-density discrete-mixture models. For the NIST 2002 SV task, we obtained an overall speedup by a factor of 2–100 with little loss in EER performance. Index Terms: speaker verification, scalar quantization, high density discrete HMM, discrete mixture HMM
Guoli Ye, Brian Kan-Wing Mak, Man-Wai Mak
INTERSPEECH3
2009 Optimization of discriminative kernels in SVM speaker verification
abstract
An important aspect of SVM-based speaker verification systems is the design of sequence kernels. These kernels should be able to map variable-length observation sequences to fixed-size supervectors that capture the dynamic characteristics of speech utterances and allow speakers to be easily distinguished. Most existing kernels in SVM speaker verification are obtained by assuming a specific form for the similarity function of supervectors. This paper relaxes this assumption to derive a new general kernel. The kernel function is general in that it is a linear combination of any kernels belonging to the reproducing kernel Hilbert space. The combination weights are obtained by optimizing the ability of a discriminant function to separate a target speaker from impostors using either regression analysis or SVM training. The idea was applied to both low- and high-level speaker verification. In both cases, results show that the proposed kernels outperform the state-of-the-art sequence kernels. Further performance enhancement was also observed when the high-level scores were combined with acoustic scores. Index Terms — speaker verification; optimal kernels; sequence kernels; SVM; high-level features. 1.
Shixiong Zhang 0001, Man-Wai Mak
INTERSPEECH2
2009 A new adaptation approach to high-level speaker-model creation in speaker verification
Shixiong Zhang 0001, Man-Wai Mak
Speech Commun.2
2008 Fusion of cleavage site detection and pairwise alignment for fast subcellular localization
abstract
In recent years, homology-based and signal-based methods have been proposed for predicting the subcellular localization of proteins. While it has been known that homology-based methods can detect more subcellular locations than signal-based methods, the former generally requires a lot more computational resources during both training and prediction. The problem will become intractable for annotating large databases. One possible solution is to reduce the sequence length. This paper proposes to use the cleavage sites detected by signal-based methods (e.g., TargetP) to extract the sequence or profile segments that contain the most localization information for alignment. It was found that the method can reduce computation time of full-length alignment by 27-fold at a cost of only 8% reduction in prediction accuracy. Moreover, the method can increase the accuracy by 0.8% and at the same time reduce the computation time by 41%. Results also show that cutting the sequences at the cleavage sites detected by TargetP is better than cutting them at a fixed position.
Man-Wai Mak, Sun-Yuan Kung
ICASSP1
2008 High-level speaker verification via articulatory-feature based sequence kernels and SVM
abstract
Interspeech 2008, Brisbane, Australia, 22-26 September 2008
Shixiong Zhang 0001, Man-Wai Mak
INTERSPEECH2
2008 Fusion of feature selection methods for pairwise scoring SVM
Man-Wai Mak, Sun-Yuan Kung
Neurocomputing1
2008 PairProSVM: Protein Subcellular Localization Based on Local Pairwise Profile Alignment and SVM
abstract
The subcellular locations of proteins are important functional annotations. An effective and reliable subcellular localization method is necessary for proteomics research. This paper introduces a new method---PairProSVM---to automatically predict the subcellular locations of proteins. The profiles of all protein sequences in the training set are constructed by PSI-BLAST and the pairwise profile-alignment scores are used to form feature vectors for training a support vector machine (SVM) classifier. It was found that PairProSVM outperforms the methods that are based on sequence alignment and amino-acid compositions even if most of the homologous sequences have been removed. This paper also demonstrates that the performance of PairProSVM is sensitive (and somewhat proportional) to the degree of its kernel matrix meeting the Mercer's condition. PairProSVM was evaluated on Reinhardt and Hubbard's, Huang and Li's, and Gardy et al.'s protein datasets. The overall accuracies on these three datasets reach 99.3\\%, 76.5\\%, and 91.9\\%, respectively, which are higher than or comparable to those obtained by sequence alignment and by the methods compared in this paper.
Man-Wai Mak, Jian Guo 0002, Sun-Yuan Kung
IEEE ACM Trans. Comput. Biol. Bioinform.1
2007 Adaptive Weight Estimation in Multi-Biometric Verification using Fuzzy Logic Decision Fusion
abstract
This paper describes a multi-biometric verification system that is fully adaptive to variability in data acquisition using fuzzy logic decision fusion. The system uses fuzzy logic to dynamically alter the weight of three biometrics (face, fingerprint and speech), taking into account the variations during data acquisition (e.g. lighting, noise and user-device interactions). A specific decision boundary can be determined by this dynamic weight assignment to make the authentication decisions. An overall EER improvement of 42.1% relative to weighted average fusion has been achieved.
Henry Pak-Sum Hui, Helen M. Meng, Man-Wai Mak
ICASSP (1)3
2007 Feature Selection for Pairwise Scoring Kernels with Applications to Protein Subcellular Localization
abstract
In biological sequence classification, it is common to convert variable-length sequences into fixed-length vectors via pairwise sequence comparison. This pairwise approach, however, can lead to feature vectors with dimension equal to the training set size, causing the curse of dimensionality. This calls for feature selection methods that can weed out irrelevant features to reduce training and recognition time. In this paper, we propose to train an SVM using the full-feature column vectors of a pairwise scoring matrix and select the relevant features based on the support vectors of the SVM. The idea stems from the fact that pairwise scoring matrices are symmetric and support vectors are important for classification. We refer to this approach as vector-index-adaptive SVM (VIA-SVM). We compare VIA-SVM with other feature selection schemes-including SVM-RFE, R-SVM, and a filter method based on symmetric divergence (SD)-in protein subcellular localization. Results show that VIA-SVM is able to automatically bound the number of selected features within a small range. We also found that fusion of VIA-SVM and SD can produce more compact feature subsets without decreasing prediction accuracy, and that while VIA-SVM is superior for large feature-set size, the combination of SD and VIA-SVM performs better at small feature-set size.
Sun-Yuan Kung, Man-Wai Mak
ICASSP (2)2
2007 Effects of Device Mismatch, Language Mismatch and Environmental Mismatch on Speaker Verification
abstract
Device, language and environmental mismatch adversely affect speaker verification (SV) performance. We investigate such effects empirically based on the M3 (multibiometric, multilingual and multi-device) corpus (H. Meng et al., 2006). Device mismatch (among 3G phone, PocketPC and a desktop PC plug-in microphone) brings relative performance degradation of 523%; language mismatch (between English and Cantonese) brings 284% and environmental mismatch (between office environment and recording studio) brings 109%. In particular, verification with wide-band models on narrow-band test data outperforms narrow-band models on wide-band test data. The 3G phone's SV performance is generally low, but remains stable across environments. Additionally, durational variations within two-second utterances may cause a relative change of 633% in SV performance.
Bin Ma 0001, Helen M. Meng, Man-Wai Mak
ICASSP (4)3
2007 High-level feature-based speaker verification via articulatory phonetic-class pronunciation modeling
abstract
Although articulatory feature-based conditional pronunciation models (AFCPMs) can capture the pronunciation characteristics of speakers, they requires one discrete density function for each phoneme, which may lead to inaccurate models when the amount of training data is limited. This paper proposes a phonetic-class based AFCPM in which the density functions in speaker models are conditioned on phonetic classes instead of phonemes. Phonemes are mapped to phonetic classes by (1) vector quantizing the phoneme-dependent universal background models, (2) grouping phonemes according to the classical phoneme tree, and (3) combination of (1) and (2). A new scoring method that uses an SVM to combine the scores of phonetic-class models is also proposed. Evaluations based on 2000 NIST SRE show that the proposed approach can effectively solve the data sparseness problem encountered in conventional AFCPM. 1.
Shixiong Zhang 0001, Man-Wai Mak, Helen M. Meng
INTERSPEECH2
2007 Environment adaptation for robust speaker verification by cascading maximum likelihood linear regression and reinforced learning
Kwok-Kwong Yiu, Man-Wai Mak, Sun-Yuan Kung
Comput. Speech Lang.2
2007 Probabilistic feature-based transformation for speaker verification over telephone networks
Man-Wai Mak, Kwok-Kwong Yiu, Sun-Yuan Kung
Neurocomputing1
2007 Speaker Verification via High-Level Feature Based Phonetic-Class Pronunciation Modeling
abstract
It has been shown recently that the pronunciation characteristics of speakers can be represented by articulatory feature-based conditional pronunciation models (AFCPMs). However, the pronunciation models are phoneme-dependent, which may lead to speaker models with low discriminative power when the amount of enrollment data is limited. This paper proposes to mitigate this problem by grouping similar phonemes into phonetic classes and representing background and speaker models as phonetic-class dependent density functions. Phonemes are grouped by (1) vector quantizing the discrete densities in the phoneme-dependent universal background models, (2) using the phone properties specified in the classical phoneme tree, or (3) combining vector quantization and phone properties. Evaluations based on 2000 NIST SRE show that this phonetic-class approach effectively alleviates the data spareness problem encountered in conventional AFCPM, which results in better performance when fused with acoustic features.
Shixiong Zhang 0001, Man-Wai Mak, Helen M. Meng
IEEE Trans. Computers2
2006 On Consistent Fusion of Multimodal Biometrics
abstract
Audio-visual (AV) biometrics offer complementary information sources, and the use of both voice and facial images for biometric authentication has recently become economically feasible. Therefore, multi-modality adaptive fusion, combining audio and visual information, offers an efficient tool for substantially improving the classification performance. In terms of implementation, we propose to integrate an audio classifier (based on Gaussian mixture models) and a visual classifier (based on FaceIT, a commercially available software) into a well-established mixture-of-expert fusion architecture. In addition, a consistent fusion strategy is introduced as a baseline fusion scheme, which establishes the lower bound of the "consistent region” in the FAR-FRR ROC. Our simulation results indicate that the prediction performance of the proposed adaptive fusion schemes fall in the consistent region. More importantly, the notion of consistent fusion can also facilitate the selection of the best modalities to fuse.
Sun-Yuan Kung, Man-Wai Mak
ICASSP (5)2
2006 A Comparison of Various Adaptation Methods for Speaker Verification With Limited Enrollment Data
abstract
One key factor that hinders the widespread deployment of speaker verification technologies is the requirement of long enrollment utterances to guarantee low error rate during verification. To gain user acceptance of speaker verification technologies, adaptation algorithms that can enroll speakers with short utterances are highly essential. To this end, this paper applies kernel eigenspace-based MLLR (KEMLLR) for speaker enrollment and compares its performance against three state-of-the-art model adaptation techniques: maximum a posteriori (MAP), maximum-likelihood linear regression (MLLR), and reference speaker weighting (RSW). The techniques were compared under the NIST2001 SRE framework, with enrollment data vary from 2 to 32 seconds. Experimental results show that KEMLLR is most effective for short enrollment utterances (between 2 to 4 seconds) and that MAP performs better when long utterances (32 seconds) are available.*This work was supported by the Research Grant Council of the Hong Kong SAR (Project Nos. CUHK 1/02C and PolyU 5214/04E).†Roger completed this work while he was with the Hong Kong University of Science and Technology before he left for CMU.‡This research is partially supported by the Research Grants Council of the Hong Kong SAR under the grant number CA02/03.EG04.
Man-Wai Mak, Roger Hsiao, Brian Kan-Wing Mak
ICASSP (1)1
2006 A Solution to the Curse of Dimensionality Problem in Pairwise Scoring Techniques
Man-Wai Mak, Sun-Yuan Kung
ICONIP (1)1
2006 Adaptive articulatory feature-based conditional pronunciation modeling for speaker verification
Ka-Yee Leung, Man-Wai Mak, Man-Hung Siu, Sun-Yuan Kung
Speech Commun.2
2005 A two-level fusion approach to multimodal biometric verification
abstract
This paper proposes a two-level fusion strategy for audio-visual biometric authentication. Specifically, fusion is performed at two levels: intramodal and intermodal. In intramodal fusion, the scores of multiple samples (e.g. utterances or video shots) obtained from the same modality are linearly combined, where the combination weights depend on the difference between the score values and a client-dependent reference score obtained during enrollment. This is followed by intermodal fusion in which the means of intramodal fused scores obtained from different modalities are either linearly combined or fused by a support vector machine (SVM). Experimental results based on the XM2VTSDB corpus show that intramodal and intermodal fusion are complementary to each other and that SVM-based intermodal fusion is superior to linear combination. 1.
Ming-Cheung Cheung, Man-Wai Mak, Sun-Yuan Kung
ICASSP (5)2
2005 Speaker Verification Using Adapted Articulatory Feature-based Conditional Pronunciation Modeling
abstract
The paper proposes an articulatory feature-based conditional pronunciation modeling (AFCPM) technique for speaker verification. The technique captures the pronunciation characteristics of speakers by modeling the linkage between the actual phones produced by the speakers and the state of articulations during speech production. The speaker models, which consist of conditional probabilities of two articulatory classes, are adapted from a set of universal background models (UBMs) via MAP adaptation. This creates a direct coupling between the speaker and background models, which prevents over-fitting the speaker models when the amount of speaker data is limited. Experimental results demonstrate that MAP adaptation not only enhances the discriminative power of the speaker models but also improves their robustness against handset mismatches. Results also show that fusing the scores derived from an AFCPM-based system and a conventional spectral-based system achieves an error rate that is significantly lower than that which can be achieved by the individual systems. This suggests that AFCPM and spectral features are complementary to each other.
Ka-Yee Leung, Man-Wai Mak, Man-Hung Siu, Sun-Yuan Kung
ICASSP (1)2
2005 Speaker verification via articulatory feature-based conditional pronunciation modeling with vowel and consonant mixture models
abstract
Articulatory feature-based conditional pronunciation modeling (AFCPM) aims to capture the pronunciation characteristics of speakers by modeling the linkage between the states of articulation during speech production and the actual phones produced by a speaker. Previous AFCPM systems use one discrete density function for each phoneme to model the pronunciation characteristics of speakers. This paper proposes using a mixture of discrete density functions for AFCPM. In particular, the pronunciation characteristics of each phoneme is modeled by two density functions: one responsible for describing the articulatory features that are more relevant to vowels and the other for consonants. Verification scores are the weighted sum of the outputs of the two models. To enhance the resolution of the pronunciation models, four articulatory properties (front-back, liprounding, place of articulation, and manner of articulation) are used for pronunciation modeling. The proposed AFCPM is applied to a speaker verification task. Results show that using four articulatory features achieves a lower error rate as compared to using two features (manner and place of articulation) only. It was also found that dividing the articulatory properties into two groups is an effective means of solving the data-sparseness problem encountered in the training phase of AFCPM systems. 1.
Ka-Yee Leung, Man-Wai Mak, Man-Hung Siu, Sun-Yuan Kung
INTERSPEECH2
2005 Channel robust speaker verification via Bayesian blind stochastic feature transformation
abstract
In telephone-based speaker verification, the channel conditions can be varied significantly from sessions to sessions. Therefore, it is desirable to estimate the channel conditions online and compensate the acoustic distortion without prior knowledge of the channel characteristics. Because no a priori knowledge is used, the estimation accuracy depends greatly on the length of the verification utterances. This paper extends the Blind Stochastic Feature Transformation (BSFT) algorithm that we recently proposed to handle the short-utterance scenario. The idea is to estimate a set of prior transformation parameters from a development set in which a wide variety of channel conditions exists in the verification utterances. The prior transformations are then incorporated into the online estimation of the BSFT parameters in a Bayesian (maximum a posteriori) fashion. The resulting transformation parameters are therefore dependent on both the prior transformations and the verification utterances. For short (long) utterances, the prior transformations play a more (less) important role. We referred the extended algorithm to as Bayesian BSFT (BBSFT) and applied it to the 2001 NIST SRE task. Results show that Bayesian BSFT outperforms BSFT for utterances shorter than or equal to 4 seconds. 1.
Kwok-Kwong Yiu, Man-Wai Mak, Sun-Yuan Kung
INTERSPEECH2
2004 Multi-sample data-dependent fusion of sorted score sequences for biometric verification
abstract
In many biometric systems, the scores of multiple samples (e.g. utterances) are averaged and the average score is compared against a decision threshold for decision making. The average score, however, may not be optimal because the distribution of the scores is ignored. To address this limitation, we have recently proposed a fusion model that incorporates the score distribution by making the fusion weights dependent on the dispersion between the frame-based scores and the prior score statistics obtained from training data. As the fusion weights are data-dependent, the positions of scores in the score sequences become detrimental to the final fused scores. We propose to enhance the fusion model by sorting the score sequences before fusion takes place. The fusion model was evaluated on a speaker verification task where each claimant utters two utterances in a verification session. Results demonstrate that fusion of sorted scores has the effect of maximizing the dispersion between the client scores and the impostor scores, making the verification process more reliable. Compared with our previous work, where no sorting was applied, the new approach reduces the equal error rate by 11 %.
Ming-Cheung Cheung, Man-Wai Mak, Sun-Yuan Kung
ICASSP (5)2
2004 Applying articulatory features to telephone-based speaker verification
abstract
This paper presents an approach that uses articulatory features (AF) derived from spectral features for telephone-based speaker verification. To minimize the acoustic mismatch caused by different handsets, handset-specific normalization is applied to the spectral features before the AF are extracted. Experimental results based on 150 speakers using 10 different handsets show that AF contain useful speaker-specific information for speaker verification and the use of handset-specific normalization significantly lowers the error rates under the handset mismatched conditions. Results also demonstrate that fusing the scores obtained from an AF-based system with those obtained from a spectral feature-based (MFCC) system helps lower the error rates of the individual systems.
Ka-Yee Leung, Man-Wai Mak, Sun-Yuan Kung
ICASSP (1)2
2004 Multi-sample fusion with constrained feature transformation for robust speaker verification
abstract
This paper proposes a single-source multi-sample fusion approach to text-independent speaker verification. In conventional speaker verification systems, the scores obtained from claimant's utterances are averaged and the resulting mean score is used for decision making. Instead of using an equal weight for all scores, this paper proposes assigning a different weight to each score, where the weights are made dependent on the difference between the score values and a speaker-dependent reference score obtained during enrollment. Because the fusion weights depend on the verification scores, a technique called constrained stochastic feature transformation is applied to minimize the mismatch between enrollment and verification data in order to enhance the scores' reliability. Experimental results based on the 2001 NIST evaluation set show that the proposed fusion approach outperforms the equal-weight approach by 22% in terms of equal error rate and 16% in terms of minimum detection cost.
Ming-Cheung Cheung, Kwok-Kwong Yiu, Man-Wai Mak, Sun-Yuan Kung
INTERSPEECH3
2004 Articulatory feature-based conditional pronunciation modeling for speaker verification
Ka-Yee Leung, Man-Wai Mak, Sun-Yuan Kung
INTERSPEECH2
2004 A new approach to channel robust speaker verification via constrained stochastic feature transformation
abstract
This paper proposes a constrained stochastic feature transformation algorithm for robust speaker verification. The algorithm computes the feature transformation parameters based on the statistical difference between a test utterance and a composite GMM formed by combining the speaker and background models. The transformation is then used to transform the test utterance to fit the clean speaker model and background model before verification. By implicitly constraining the transformation, the transformed features can fit both models simultaneously. Experimental results based on the 2001 NIST evaluation set show that the proposed algorithms achieves significant improvement in both equal error rate and minimum detection cost when compared to cepstral mean subtraction and Z-norm. The performance of the proposed transformation approach is also slightly better than the short-time Gaussianization method proposed in [1].
Man-Wai Mak, Kwok-Kwong Yiu, Ming-Cheung Cheung, Sun-Yuan Kung
INTERSPEECH1
2003 Robust speaker verification from GSM-transcoded speech based on decision fusion and feature transformation
abstract
In speaker verification, a claimant may produce two or more utterances. Typically, the scores of the speech patterns extracted from these utterances are averaged and the resulting mean score is compared with a decision threshold. Rather than simply computing the mean score, we propose to compute the optimal weights for fusing the scores based on the score distribution of the independent utterances and our prior knowledge about the score statistics. More specifically, we use enrollment data to compute the mean scores of client speakers and impostors and consider them to be the prior scores. During verification, we set the fusion weights for individual speech patterns to be a function of the dispersion between the scores of these speech patterns and the prior scores. Experimental results based on the GSM-transcoded speech of 150 speakers from the HTIMIT corpus demonstrate that the proposed fusion algorithm can increase the dispersion between the mean speaker scores and the mean impostor scores. Compared with a baseline approach where equal weights are assigned to all scores, the proposed approach provides a relative error reduction of 19%.
Man-Wai Mak, Ming-Cheung Cheung, Sun-Yuan Kung
ICASSP (2)1
2003 Adaptive decision fusion for multi-sample speaker verification over GSM networks
abstract
In speaker verification, a claimant may produce two or more utterances. In our previous study [1], we proposed to compute the optimal weights for fusing the scores of these utterances based on their score distribution and our prior knowledge about the score statistics estimated from the mean scores of the corresponding client speaker and some pseudo-impostors during enrollment. As the fusion weights depend on the prior scores, in this paper, we propose to adapt the prior scores during verification based on the likelihood of the claimant being an impostor. To this end, a pseudo-imposter GMM score model is created for each speaker. During verification, the claimant's scores are fed to the score model to obtain a likelihood for adapting the prior score. Experimental results based on the GSM-transcoded speech of 150 speakers from the HTIMIT corpus demonstrate that the proposed prior score adaptation approach provides a relative error reduction of 15% when compared with our previous approach where the prior scores are non-adaptive.
Ming-Cheung Cheung, Man-Wai Mak, Sun-Yuan Kung
INTERSPEECH2
2003 Environment adaptation for robust speaker verification
abstract
In speaker verification over public telephone networks, utterances can be obtained from different types of handsets. Different handsets may introduce different degrees of distortion to the speech signals. This paper attempts to combine a handset selector with (1) handset-specific transformations and (2) handset-dependent speaker models to reduce the effect caused by the acoustic distortion. Specifically, a number of Gaussian mixture models are independently trained to identify the most likely handset given a test utterance; then during recognition, the speaker model and background model are either transformed by MLLR-based handset-specific transformation or respectively replaced by a handset-dependent speaker model and a handset-dependent background model whose parameters were adapted by reinforced learning to fit the new environment. Experimental results based on 150 speakers of the HTIMIT corpus show that environment adaptation based on both MLLR and reinforced learning outperforms the classical CMS, Hnorm and Tnorm approaches, with MLLR adaptation achieves the best performance.
Kwok-Kwong Yiu, Man-Wai Mak, Sun-Yuan Kung
INTERSPEECH2
2003 Speaker verification based on g.729 and g.723.1 coder parameters and handset mismatch compensation
abstract
A novel technique for speaker verification over a communication network is proposed. The technique employs cepstral coefficients (LPCCs) derived from G.729 and G.723.1 coder parameters as feature vectors. Based on the LP coefficients derived from the coder parameters, LP residuals are reconstructed, and the verification performance is improved by taking account of the additional speaker-dependent information contained in the reconstructed residuals. This is achieved by adding the LPCCs of the LP residuals to the LPCCs derived from the coder parameters. To reduce the acoustic mismatch between different handsets, a technique combining a handset selector with stochastic feature transformation is employed. Experimental results based on 150 speakers show that the proposed technique outperforms the approaches that only utilize the coder-derived LPCCs.
Eric W. M. Yu, Man-Wai Mak, Chin-Hung Sit, Sun-Yuan Kung
INTERSPEECH2
2002 Combining stochastic feature transformation and handset identification for telephone-based speaker verification
abstract
The performance of telephone-based speaker verification systems can be severely degraded by the acoustic mismatch caused by telephone handsets. This paper proposes to combine a handset selector with stochastic feature transformation to reduce the mismatch. Specifically, a GMM-based handset selector is trained to identify the most likely handset used by the claimants, and then handset-specific stochastic feature transformations are applied to the distorted feature vectors. To overcome the non-linear distortion introduced by telephone handsets, a 2nd-order stochastic feature transformation is proposed. Estimation algorithms based on the stochastic matching technique and the EM algorithm are derived. Experimental results based on 150 speakers of the HTIMIT corpus show that the handset selector is able to identify the handsets accurately (98.3%), and that both linear and non-linear transformation reduce the error rate significantly (from 12.37% to 5.49%).
Man-Wai Mak, Sun-Yuan Kung
ICASSP1
2002 Divergence-based out-of-class rejection for telephone handset identification
abstract
Research has shown that handset selectors can be used to assist telephone-based speech/speaker recognition. Most handset selectors, however, simply select the most likely handset from a set of known handsets even for speech coming from an ‘unseen’ handset. This paper proposes a divergence-based handset selector with out-of-handset (OOH) rejection capability to identify the ‘unseen’ handsets. This is achieved by measuring the Jensen difference between the selector’s output and a constant vector with identical elements. The resulting handset selector is combined with a feature-based channel compensation algorithm for telephonebased speaker verification. Utterances whose handsets were identified as ‘unseen’ are either transformed by a global bias vector or normalized by cepstral mean subtraction (CMS). On the other hand, if the handset can be identified (considered as ‘seen’), its corresponding transformation parameters will be used to transform the utterances. Experiments based on ten handsets of the HTIMIT corpus show that using the transformation parameters of the ‘seen’ handsets to transform the utterances with correctly identified handsets and processing those utterances with ‘unseen’ handsets by CMS achieve the best result.
Chi-Leung Tsang, Man-Wai Mak, Sun-Yuan Kung
INTERSPEECH2
2002 A Comparative Study on Kernel-Based Probabilistic Neural Networks for Speaker Verification
abstract
This paper compares kernel-based probabilistic neural networks for speaker verification based on 138 speakers of the YOHO corpus. Experimental evaluations using probabilistic decision-based neural networks (PDBNNs), Gaussian mixture models (GMMs) and elliptical basis function networks (EBFNs) as speaker models were conducted. The original training algorithm of PDBNNs was also modified to make PDBNNs appropriate for speaker verification. Results show that the equal error rate obtained by PDBNNs and GMMs is less than that of EBFNs (0.33% vs. 0.48%), suggesting that GMM- and PDBNN-based speaker models outperform the EBFN ones. This work also finds that the globally supervised learning of PDBNNs is able to find decision thresholds that not only maintain the false acceptance rates to a low level but also reduce their variation, whereas the ad-hoc threshold-determination approach used by the EBFNs and GMMs causes a large variation in the error rates. This property makes the performance of PDBNN-based systems more predictable.
Kwok-Kwong Yiu, Man-Wai Mak, Sun-Yuan Kung
Int. J. Neural Syst.2
2000 A two-stage scoring method combining world and cohort models for speaker verification
abstract
The cohort and world models are commonly used for scoring normalization in speaker verification. As these models represent different regions of the feature space, a better solution could be obtained by integrating them into a single framework. In this paper, we embed the two models in elliptical basis function networks and propose a two-stage decision procedure for improving verification performance. In the first stage, the score of an unknown utterance is normalized by a world model. If the difference between the resulting normalized score and a world threshold is sufficiently large, the claimant is accepted or rejected immediately. Otherwise, the score will be normalized by a cohort model and compared with a cohort threshold to make a final accept/reject decision. Experimental evaluations based on the YOHO corpus suggest that the two-stage method achieves a lower error rate as compared to the case where only one background model is used.
W. D. Zhang, Man-Wai Mak, M. X. He
ICASSP2
2000 A study of the Lamarckian evolution of recurrent neural networks
abstract
Training neural networks by evolutionary search can require a long computation time. In certain situations, using Lamarckian evolution, local search and evolutionary search can complement each other to yield a better training algorithm. This paper demonstrates the potential of this evolutionary-learning synergy by applying it to train recurrent neural networks in an attempt to resolve a long-term dependency problem and the inverted pendulum problem. This work also aims at investigating the interaction between local search and evolutionary search when they are combined; it is found that the combinations are particularly efficient when the local search is simple. In the case where no teacher signal is available for the local search to learn the desired task directly, the paper proposes a related local task for the local search to learn, and finds that this approach is able to reduce the training time considerably.
Kim W. C. Ku, Man-Wai Mak, Wan-Chi Siu
IEEE Trans. Evol. Comput.2
2000 Estimation of elliptical basis function parameters by the EM algorithm with application to speaker verification
abstract
This paper proposes to incorporate full covariance matrices into the radial basis function (RBF) networks and to use the expectation-maximization (EM) algorithm to estimate the basis function parameters. The resulting networks, referred to as elliptical basis function (EBF) networks, are evaluated through a series of text-independent speaker verification experiments involving 258 speakers from a phonetically balanced, continuous speech corpus (TIMIT).We propose a verification procedure using RBF and EBF networks as speaker models and show that the networks are readily applicable to verifying speakers using LP-derived cepstral coefficients as features. Experimental results show that small EBF networks with basis function parameters estimated by the EM algorithm outperform the large RBF networks trained in the conventional approach. The results also show that the equal error rate achieved by the EBF networks is about two-third of that achieved by the vetor quantization (VQ)-based speaker models.
Man-Wai Mak, Sun-Yuan Kung
IEEE Trans. Neural Networks Learn. Syst.1
1999 Elliptical basis function networks and radial basis function networks for speaker verification: a comparative study
abstract
It is well known that radial basis function (RBF) networks require a large number of function centers if the data to be modeled contain clusters with complicated shape. This paper proposes to overcome this problem by incorporating full covariance matrices into the RBF structure and to use the expectation-maximization (EM) algorithm to estimate the network parameters. The resulting networks, referred to as the elliptical basis function (EBF) networks, are applied to text-independent speaker verification. Experimental evaluations based on 258 speakers of the TIMIT corpus show that smaller size EBF networks with basis function parameters determined by the EM algorithm outperform the large RBF networks trained by the conventional approach.
Man-Wai Mak
IJCNN1
1999 A new cepstrum-based channel compensation method for speaker verification
T. F. Lo, Man-Wai Mak, Kwok-Kwong Yiu
EUROSPEECH2
1999 A priori threshold determination for phrase-prompted speaker verification
abstract
1 Now work at Intel China Research Center, Beijing. Email: [email protected] ABSTRACT DDBHMM solved the defects of traditional HMM. Based on DDBHMM, the problem of how to effectively utilize the duration information is studied in detail. The approach on estimating the duration distribution is introduced firstly, then the data file is classified according to the speak rate. The recognition experiment shows that, the duration information behaves best on the data of low speak rate, behaves normal on the data of medium speak rate and has little effect on the data of fast speak rate. Therefore, the most importance of duration is that by it the more accurate state segmentation point could be obtained and then the recognition rate can be improved. At the same time, the robustness of the system to speaking rate is improved with the employment of the duration information. Furthermore, the method of classified duration and normalized duration is also put forward and studied in detail, it shows that both of the two method can improve the effect. In order to study the dependency between the duration, the method of using the Bigram of the duration is proposed and analyzed. At last, the approach of post processing duration is studied, it shows that only based on DDBHMM, and utilizing the duration information synchronously in the recognition process, then the performance can be improved greatly.
W. D. Zhang, Kwok-Kwong Yiu, Man-Wai Mak, M. X. He
EUROSPEECH3
1999 A conjugate gradient learning algorithm for recurrent neural networks
Wing-Fai Chang, Man-Wai Mak
Neurocomputing2
1999 On the improvement of the real time recurrent learning algorithm for recurrent neural networks
Man-Wai Mak, Kim-Wing Ku, Yee-Ling Lu
Neurocomputing1
1999 Gaussian Mixture Models and Probabilistic Decision-Based Neural Networks for Pattern Classification: A Comparative Study
Kwok-Kwong Yiu, Man-Wai Mak, Chi-Kwong Li
Neural Comput. Appl.2
1999 Adding learning to cellular genetic algorithms for training recurrent neural networks
abstract
This paper proposes a hybrid optimization algorithm which combines the efforts of local search (individual learning) and cellular genetic algorithms (GA's) for training recurrent neural networks (RNN's). Each weight of an RNN is encoded as a floating point number, and a concatenation of the numbers forms a chromosome. Reproduction takes place locally in a square grid with each grid point representing a chromosome. Two approaches, Lamarckian and Baldwinian mechanisms, for combining cellular GA's and learning have been compared. Different hill-climbing algorithms are incorporated into the cellular GA's as learning methods. These include the real-time recurrent learning (RTRL) and its simplified versions, and the delta rule. The RTRL algorithm has been successively simplified by freezing some of the weights to form simplified versions. The delta rule, which is the simplest form of learning, has been implemented by considering the RNN's as feedforward networks during learning. The hybrid algorithms are used to train the RNN's to solve a long-term dependency problem. The results show that Baldwinian learning is inefficient in assisting the cellular GA. It is conjectured that the more difficult it is for genetic operations to produce the genotypic changes that match the phenotypic changes due to learning, the poorer is the convergence of Baldwinian learning. Most of the combinations using the Lamarckian mechanism show an improvement in reducing the number of generations required for an optimum network; however, only a few can reduce the actual time taken. Embedding the delta rule in the cellular GA's has been found to be the fastest method. It is also concluded that learning should not be too extensive if the hybrid algorithm is to be benefit from learning.
Kim W. C. Ku, Man-Wai Mak, Wan-Chi Siu
IEEE Trans. Neural Networks2
1998 Empirical Analysis of the Factors that Affect the Baldwin Effect
Kim W. C. Ku, Man-Wai Mak
PPSN2
1995 A learning algorithm for Recurrent Radial Basis Function Networks
Man-Wai Mak
Neural Process. Lett.1
1994 Speaker identification using multilayer perceptrons and radial basis function networks
Man-Wai Mak, William G. Allen, G. G. Sexton
Neurocomputing1
1994 Lip-motion analysis for speech segmentation in noise
Man-Wai Mak, William G. Allen
Speech Commun.1
1994 A lip-tracking system based on morphological processing and block matching techniques
Man-Wai Mak, William G. Allen
Signal Process. Image Commun.1