Guoli Ye

dblp:53/9230 · DBLP profile ↗
← Back
24ranked-venue papers
5as first author
10since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 13 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Customizing Speech Recognition Model with Large Language Model Feedback
abstract
Automatic speech recognition (ASR) systems have achieved strong performance on general transcription tasks. However, they continue to struggle with recognizing rare named entities and adapting to domain mismatches. In contrast, large language models (LLMs), trained on massive internet-scale datasets, are often more effective across a wide range of domains. In this work, we propose a reinforcement learning based approach for unsupervised domain adaptation, leveraging unlabeled data to enhance transcription quality—particularly the named entities affected by domain mismatch—through feedback from a LLM. Given contextual information, our framework employs a LLM as the reward model to score the hypotheses from the ASR model. These scores serve as reward signals to fine-tune the ASR model via reinforcement learning. Our method achieves a 21% improvement on entity word error rate over conventional self-training methods.
Shaoshi Ling, Guoli Ye
ASRU2
2025 Efficient Long-Form Speech Recognition for General Speech In-Context Learning
abstract
We propose a novel approach to end-to-end automatic speech recognition (ASR) to achieve efficient speech in-context learning (SICL) for (i) long-form speech decoding, (ii) test-time speaker adaptation, and (iii) test-time contextual biasing. Specifically, we introduce an attention-based encoder-decoder (AED) model with SICL capability (referred to as SICL-AED), where the decoder utilizes an utterance-level cross-attention to integrate information from the encoder’s output efficiently, and a document-level self-attention to learn contextual information. Evaluated on the benchmark TEDLIUM3 dataset, SICL-AED achieves an 8.64% relative word error rate (WER) reduction compared to a baseline utterance-level AED model by leveraging previously decoded outputs as in-context examples. It also demonstrates comparable performance to conventional long-form AED systems with significantly reduced runtime and memory complexity. Additionally, we introduce an in-context fine-tuning (ICFT) technique that further enhances SICL effectiveness during inference. Experiments on speaker adaptation and contextual biasing highlight the general speech in-context learning capabilities of our system, achieving effective results with provided contexts. Without specific fine-tuning, SICL-AED matches the performance of supervised AED baselines for speaker adaptation and improves entity recall by 64% for contextual biasing task.
Hao Yen, Shaoshi Ling, Guoli Ye
ICASSP3
2024 Adapting Large Language Model with Speech for Fully Formatted End-to-End Speech Recognition
abstract
Most end-to-end (E2E) speech recognition models are composed of encoder and decoder blocks that perform acoustic and language modeling functions. Pretrained large language models (LLMs) have the potential to improve the performance of E2E ASR. However, integrating a pretrained language model into an E2E speech recognition model has shown limited benefits due to the mismatches between text-based LLMs and those used in E2E ASR. In this paper, we explore an alternative approach by adapting a pretrained LLMs to speech. Our experiments on fully-formatted E2E ASR transcription tasks across various domains demonstrate that our approach can effectively leverage the strengths of pretrained LLMs to produce more readable ASR transcriptions. Our model, which is based on the pretrained large language models with either an encoder-decoder or decoder-only structure, surpasses strong ASR models such as Whisper1, in terms of recognition error rate, considering formats like punctuation and capitalization as well.
Shaoshi Ling, Yuxuan Hu 0003, Shuangbei Qian, Guoli Ye, Yao Qian, Yifan Gong 0001, Ed Lin, Michael Zeng 0001
ICASSP4
2024 Hybrid Attention-Based Encoder-Decoder Model for Efficient Language Model Adaptation
abstract
The attention-based encoder-decoder (AED) speech recognition model has been widely successful in recent years. However, the joint optimization of acoustic model and language model in end-to-end manner has created challenges for text adaptation. In particular, effective, quick and inexpensive adaptation with text input has become a primary concern for deploying AED systems in the industry. To address this issue, we propose a novel model, the hybrid attention-based encoder-decoder (HAED) speech recognition model that preserves the modularity of conventional hybrid automatic speech recognition systems. Our HAED model separates the acoustic and language models, allowing for the use of conventional text-based language model adaptation techniques. We demonstrate that the proposed HAED model yields 23% relative Word Error Rate (WER) improvements when out-of-domain text data is used for language model adaptation, with only a minor degradation in WER on a general test set compared with the conventional AED model.
Shaoshi Ling, Guoli Ye, Rui Zhao 0017, Yifan Gong 0001
SLT2
2023 Multi Transcription-Style Speech Transcription Using Attention-Based Encoder-Decoder Model
abstract
Human professional transcription services provide a variety of transcription styles to customize different needs. To accommodate different users and facilitate seamless integration with downstream applications, we propose a framework to generate multi-style transcription in an attention-based encoder-decoder model (AED) using three different architectures: (A) style-dependent layers; (B) mixed-style output; (C) style-dependent prompt. In this framework, both the verbatim lexical transcription and the readable transcription of various styles can be generated simultaneously or separately, through a single decoding pass or multiple decoding passes on-demand. We conduct experiments in a large-scale AED-based speech transcription system trained with 50k hours speech. The proposed framework can achieve nearly on-par performance compared to the single-style AED with significant savings in model footprint and decoding cost. Moreover, it provides an efficient data sharing mechanism across different styles through knowledge transfer.
Yan Huang 0028, Piyush Behre, Guoli Ye, Shawn Chang, Yifan Gong 0001
ASRU3
2022 Have Best of Both Worlds: Two-Pass Hybrid and E2E Cascading Framework for Speech Recognition
abstract
Hybrid and end-to-end (E2E) systems have their individual advantages, with different error patterns in the speech recognition results. By jointly modeling audio and text, the E2E model performs better in matched scenarios and scales well with a large amount of paired audio-text training data. The modularized hybrid model is easier for customization, and better to make use of a massive amount of unpaired text data. This paper proposes a two-pass hybrid and E2E cascading (HEC) framework to combine the hybrid and E2E model in order to take advantage of both sides, with hybrid in the first pass and E2E in the second pass. We show that the proposed system achieves 8-10% relative word error rate reduction with respect to each individual system. More importantly, compared with the pure E2E system, we show the proposed system has the potential to keep the advantages of hybrid system, e.g., customization and segmentation capabilities. We also show the second pass E2E model in HEC is robust with respect to the change in the first pass hybrid model.
Guoli Ye, Vadim Mazalov, Jinyu Li 0001, Yifan Gong 0001
ICASSP1
2021 Rapid Speaker Adaptation for Conformer Transducer: Attention and Bias Are All You Need
Yan Huang 0028, Guoli Ye, Jinyu Li 0001, Yifan Gong 0001
Interspeech2
2021 End-to-End Speaker-Attributed ASR with Transformer
abstract
This paper presents our recent effort on end-to-end speakerattributed automatic speech recognition, which jointly performs speaker counting, speech recognition and speaker identification for monaural multi-talker audio.Firstly, we thoroughly update the model architecture that was previously designed based on a long short-term memory (LSTM)-based attention encoder decoder by applying transformer architectures.Secondly, we propose a speaker deduplication mechanism to reduce speaker identification errors in highly overlapped regions.Experimental results on the LibriSpeechMix dataset shows that the transformer-based architecture is especially good at counting the speakers and that the proposed model reduces the speakerattributed word error rate by 47% over the LSTM-based baseline.Furthermore, for the LibriCSS dataset, which consists of real recordings of overlapped speech, the proposed model achieves concatenated minimum-permutation word error rates of 11.9% and 16.3% with and without target speaker profiles, respectively, both of which are the state-of-the-art results for LibriCSS with the monaural setting.
Naoyuki Kanda, Guoli Ye, Yashesh Gaur, Xiaofei Wang 0009, Zhong Meng, Zhuo Chen 0006, Takuya Yoshioka
Interspeech2
2021 Large-Scale Pre-Training of End-to-End Multi-Talker ASR for Meeting Transcription with Single Distant Microphone
abstract
Transcribing meetings containing overlapped speech with only a single distant microphone (SDM) has been one of the most challenging problems for automatic speech recognition (ASR).While various approaches have been proposed, all previous studies on the monaural overlapped speech recognition problem were based on either simulation data or small-scale real data.In this paper, we extensively investigate a two-step approach where we first pre-train a serialized output training (SOT)-based multi-talker ASR by using large-scale simulation data and then fine-tune the model with a small amount of real meeting data.Experiments are conducted by utilizing 75 thousand (K) hours of our internal single-talker recording to simulate a total of 900K hours of multi-talker audio segments for supervised pretraining.With fine-tuning on the 70 hours of the AMI-SDM training data, our SOT ASR model achieves a word error rate (WER) of 21.2% for the AMI-SDM evaluation set while automatically counting speakers in each test segment.This result is not only significantly better than the previous state-of-the-art WER of 36.4% with oracle utterance boundary information but also better than a result by a similarly fine-tuned single-talker ASR model applied to beamformed audio.
Naoyuki Kanda, Guoli Ye, Yu Wu 0012, Yashesh Gaur, Xiaofei Wang 0009, Zhong Meng, Zhuo Chen 0006, Takuya Yoshioka
Interspeech2
2021 Minimum Word Error Rate Training with Language Model Fusion for End-to-End Speech Recognition
abstract
Integrating external language models (LMs) into end-to-end (E2E) models remains a challenging task for domain-adaptive speech recognition.Recently, internal language model estimation (ILME)-based LM fusion has shown significant word error rate (WER) reduction from Shallow Fusion by subtracting a weighted internal LM score from an interpolation of E2E model and external LM scores during beam search.However, on different test sets, the optimal LM interpolation weights vary over a wide range and have to be tuned extensively on well-matched validation sets.In this work, we perform LM fusion in the minimum WER (MWER) training of an E2E model to obviate the need for LM weights tuning during inference.Besides MWER training with Shallow Fusion (MWER-SF), we propose a novel MWER training with ILME (MWER-ILME) where the ILME-based fusion is conducted to generate N-best hypotheses and their posteriors.Additional gradient is induced when internal LM is engaged in MWER-ILME loss computation.During inference, LM weights pre-determined in MWER training enable robust LM integrations on test sets from different domains.Experimented with 30K-hour trained transformer transducers, MWER-ILME achieves on average 8.8% and 5.8% relative WER reductions from MWER and MWER-SF training, respectively, on 6 different test sets.
Zhong Meng, Yu Wu 0012, Naoyuki Kanda, Liang Lu 0001, Xie Chen 0001, Guoli Ye, Eric Sun, Jinyu Li 0001, Yifan Gong 0001
Interspeech6
2020 Adaptation of RNN Transducer with Text-To-Speech Technology for Keyword Spotting
abstract
With the advent of recurrent neural network transducer (RNN-T) model, the performance of keyword spotting (KWS) systems has greatly improved. However, the KWS systems, employed for wake-word detection, still rely on the availability of keyword specific training data for achieving reasonable performance on each keyword. With a goal to improve the KWS performance for these keywords without having to collect additional natural speech data, we explore Text-To-Speech (TTS) technology to synthetically generate training data for such keywords. Employing an RNN-T based KWS model, already well trained on large keyword-independent natural speech dataset, as a seed model, we run adaptation experiments using the generated keyword-specific TTS data. Besides observing a considerable improvement in the overall performance for the low-resource keywords, we find that the performance improvement with TTS-generated training data, similar to natural speech data, depends on speaker diversity, amount of data per speaker and data simulation. We get additional improvement in performance by selectively adapting specific parts of the RNN-T model and gain key insights into different architectural constructs of RNN-T model.
Eva Sharma, Guoli Ye, Wenning Wei, Rui Zhao 0017, Jian Wu 0027, Lei He 0005, Ed Lin, Yifan Gong 0001
ICASSP2
2020 Semantic Mask for Transformer Based End-to-End Speech Recognition
abstract
Attention-based encoder-decoder model has achieved impressive results for both automatic speech recognition (ASR) and text-to-speech (TTS) tasks.This approach takes advantage of the memorization capacity of neural networks to learn the mapping from the input sequence to the output sequence from scratch, without the assumption of prior knowledge such as the alignments.However, this model is prone to overfitting, especially when the amount of training data is limited.Inspired by SpecAugment and BERT, in this paper, we propose a semantic mask based regularization for training such kind of end-toend (E2E) model.The idea is to mask the input features corresponding to a particular output token, e.g., a word or a wordpiece, in order to encourage the model to fill the token based on the contextual information.While this approach is applicable to the encoder-decoder framework with any type of neural network architecture, we study the transformer-based model for ASR in this work.We perform experiments on Librispeech 960h and TedLium2 data sets, and achieve the state-of-the-art performance on the test set in the scope of E2E models.
Chengyi Wang 0002, Yu Wu 0012, Yujiao Du, Jinyu Li 0001, Shujie Liu 0001, Liang Lu 0001, Shuo Ren 0002, Guoli Ye, Sheng Zhao 0002, Ming Zhou 0001
INTERSPEECH8
2020 Low Latency End-to-End Streaming Speech Recognition with a Scout Network
abstract
The attention-based Transformer model has achieved promising results for speech recognition (SR) in the offline mode.However, in the streaming mode, the Transformer model usually incurs significant latency to maintain its recognition accuracy when applying a fixed-length look-ahead window in each encoder layer.In this paper, we propose a novel low-latency streaming approach for Transformer models, which consists of a scout network and a recognition network.The scout network detects the whole word boundary without seeing any future frames, while the recognition network predicts the next subword by utilizing the information from all the frames before the predicted boundary.Our model achieves the best performance (2.7/6.4WER) with only 639 ms latency on the test-clean and test-other data sets of Librispeech.
Chengyi Wang 0002, Yu Wu 0012, Liang Lu 0001, Shujie Liu 0001, Jinyu Li 0001, Guoli Ye, Ming Zhou 0001
INTERSPEECH6
2019 Towards Code-switching ASR for End-to-end CTC Models
abstract
Although great progress has been made on end-to-end (E2E) models for monolingual and multilingual automatic speech recognition (ASR), there is no successful study for E2E models on the challenging intra-sentential code-switching (CS) ASR task to our best knowledge. In this paper, we propose an approach for CS ASR using E2E connectionist temporal classification (CTC) models. We use a frame-level language identification model to linearly adjust the posteriors of an E2E CTC model. We evaluate the proposed method on Microsoft live Chinese Cortana data with 7000 hours Chinese and English monolingual data and 300 hours CS data as the training data. Trained with only monolingual data without observing any CS data, the proposed method can obtain up to 6.3% relative word error rate (WER) reduction. In the scenario of training with both monolingual and CS data, the proposed method can get up to 4.2% relative WER improvement. This approach can also maintain comparable performance on a Chinese test set compared with baseline models.
Jinyu Li 0001, Guoli Ye, Rui Zhao 0017, Yifan Gong 0001
ICASSP3
2019 Advancing Acoustic-to-Word CTC Model With Attention and Mixed-Units
abstract
The acoustic-to-word model based on the Connectionist Temporal Classification (CTC) criterion is a natural end-to-end (E2E) system directly targeting word as output unit. Two issues exist in the system: first, the current output of the CTC model relies on the current input and does not account for context weighted inputs. This is the hard alignment issue. Second, the word-based CTC model suffers from the out-of-vocabulary (OOV) issue. This means it can model only frequently occurring words while tagging the remaining words as OOV. Hence, such a model is limited in its capacity in recognizing only a fixed set of frequent words. In this study, we propose addressing these problems using a combination of attention mechanism and mixed-units. In particular, we introduce Attention CTC, Self-Attention CTC, Hybrid CTC, and Mixed-unit CTC. First, we blend attention modeling capabilities directly into the CTC network using Attention CTC and Self-Attention CTC. Second, to alleviate the OOV issue, we present Hybrid CTC which uses a word and letter CTC with shared hidden layers. The Hybrid CTC consults the letter CTC when the word CTC emits an OOV. Then, we propose a much better solution by training a Mixed-unit CTC which decomposes all the OOV words into sequences of frequent words and multi-letter units. Evaluated on a 3400 hours Microsoft Cortana voice assistant task, our final acoustic-to-word solution using attention and mixed-units achieves a relative reduction in word error rate (WER) over the vanilla word CTC by 12.09%. Such an E2E model without using any language model (LM) or complex decoder also outperforms a traditional context-dependent (CD) phoneme CTC with strong LM and decoder by 6.79% relative.
Amit Das 0007, Jinyu Li 0001, Guoli Ye, Rui Zhao 0017, Yifan Gong 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2018 Exploring Sequential Characteristics in Speaker Bottleneck Feature for Text-Dependent Speaker Verification
abstract
In this paper, given the speaker bottleneck feature vectors extracted with speaker discriminant neural networks, we focus on using the sequential speaker characteristics for text-dependent speaker verification. In each evaluation trial, speaker supervectors are used as the representations of the sequential speaker characteristics rendered in the compared speech utterances. To this end, dynamic time warping is used to warp the variable-length speaker feature vector sequences of the utterances to the same length. Thereafter for every utterance, a speaker supervector can be obtained as the concatenation of its speaker feature vectors. We use Euclidean distance and support vector machine (SVM) to compute the decision score on the speaker supervectors. Our experiments on a Microsoft internal keyword-spotting database showed the effectiveness of the proposed speaker supervector for text-dependent speaker verification. Moreover, when SVM backend was used in scoring, the speaker supervector achieved the best EER performance 1.627%, better than the combination of i-vector and probabilistic linear discriminant analysis.
Yong Zhao 0008, Shixiong Zhang 0001, Guoli Ye, Frank K. Soong
ICASSP5
2018 Advancing Acoustic-to-Word CTC Model
abstract
The acoustic-to-word model based on the connectionist temporal classification (CTC) criterion was shown as a natural end-to-end (E2E) model directly targeting words as output units. However, the word-based CTC model suffers from the out-of-vocabulary (OOV) issue as it can only model limited number of words in the output layer and maps all the remaining words into an OOV output node. Hence, such a word-based CTC model can only recognize the frequent words modeled by the network output nodes. Our first attempt to improve the acoustic-to-word model is a hybrid CTC model which consults a letter-based CTC when the word-based CTC model emits OOV tokens during testing time. Then, we propose a much better solution by training a mixed-unit CTC model which decomposes all the OOV words into sequences of frequent words and multi-letter units. Evaluated on a 3400 hours Microsoft Cortana voice assistant task, the final acoustic-to-word solution improves the baseline word-based CTC by relative 12.09% word error rate (WER) reduction when combined with our proposed attention CTC. Such an E2E model without using any language model (LM) or complex decoder outperforms the traditional context-dependent phoneme CTC which has strong LM and decoder by relative 6.79%.
Jinyu Li 0001, Guoli Ye, Amit Das 0007, Rui Zhao 0017, Yifan Gong 0001
ICASSP2
2018 Developing Far-Field Speaker System Via Teacher-Student Learning
abstract
In this study, we develop the keyword spotting (KWS) and acoustic model (AM) components in a far-field speaker system. Specifically, we use teacher-student (T/S) learning to adapt a close-talk well-trained production AM to far-field by using parallel close-talk and simulated far-field data. We also use T/S learning to compress a large-size KWS model into a small-size one to fit the device computational cost. Without the need of transcription, T/S learning well utilizes untranscribed data to boost the model performance in both the AM adaptation and KWS model compression. We further optimize the models with sequence discriminative training and live data to reach the best performance of systems. The adapted AM improved from the baseline by 72.60% and 57.16% relative word error rate reduction on play-back and live test data, respectively. The final KWS model size was reduced by 27 times from a large-size KWS model without losing accuracy.
Jinyu Li 0001, Rui Zhao 0017, Zhuo Chen 0006, Changliang Liu, Guoli Ye, Yifan Gong 0001
ICASSP6
2017 Acoustic-to-word model without OOV
abstract
Recently, the acoustic-to-word model based on the Connectionist Temporal Classification (CTC) criterion was shown as a natural end-to-end model directly targeting words as output units. However, this type of word-based CTC model suffers from the out-of-vocabulary (OOV) issue as it can only model limited number of words in the output layer and maps all the remaining words into an OOV output node. Therefore, such word-based CTC model can only recognize the frequent words modeled by the network output nodes. It also cannot easily handle the hot-words which emerge after the model is trained. In this study, we improve the acoustic-to-word model with a hybrid CTC model which can predict both words and characters at the same time. With a shared-hidden-layer structure and modular design, the alignments of words generated from the word-based CTC and the character-based CTC are synchronized. Whenever the acoustic-to-word model emits an OOV token, we back off that OOV segment to the word output generated from the character-based CTC, hence solving the OOV or hot-words issue. Evaluated on a Microsoft Cortana voice assistant task, the proposed model can reduce the errors introduced by the OOV output token in the acoustic-to-word model by 30%.
Jinyu Li 0001, Guoli Ye, Rui Zhao 0017, Jasha Droppo, Yifan Gong 0001
ASRU2
2016 Geo-location dependent deep neural network acoustic model for speech recognition
abstract
Users from the same geo-location region exhibit similar acoustic characteristics, e.g., they have similar accent; even more, they may have similar preference to device. In this paper, we propose to build geo-location dependent deep neural network for speech recognition, where the geo-location signal is inferred from users' GPS. During runtime, the server will base on a user's geo-location to select the right model to recognize his voice. We tackle three major issues associated with this model: high train/deployment cost, large model size, and train data sparsity. Our solution is featured by its low cost, thus practical for production modeling. We also discuss the reliability of GPS signal in practical use. The proposed model is evaluated on Microsoft Chinese voice search and Cortana live test set. Among 12 provinces, it shows an overall 4.8% relative character error rate reduction, over a strong baseline production-level model, with only 50% model size increase. The gain is larger for the low-resource provinces, with relative error rate reduction up to 9%.
Guoli Ye, Chaojun Liu, Yifan Gong 0001
ICASSP1
2016 Deep Convolutional Neural Networks with Layer-Wise Context Expansion and Attention
abstract
In this paper, we propose a deep convolutional neural network (CNN) with layer-wise context expansion and location-based attention, for large vocabulary speech recognition. In our model each higher layer uses information from broader contexts, along both the time and frequency dimensions, than its immediate lower layer. We show that both the layer-wise context expansion and the location-based attention can be implemented using the element-wise matrix product and the convolution operation. For this reason, contrary to other CNNs, no pooling operation is used in our model. Experiments on the 309hr Switchboard task and the 375hr short message dictation task indicates that our model outperforms both the DNN and LSTM significantly.
Dong Yu 0001, Wayne Xiong, Jasha Droppo, Andreas Stolcke, Guoli Ye, Jinyu Li 0001, Geoffrey Zweig
INTERSPEECH5
2012 Transition probabilities are more important than we once thought
abstract
It is generally believed that the transition probabilities in a hidden Markov model (HMM) have a limited role in the speech decoding process. In this paper, through a series of recognition experiments on Wall Street Journal (WSJ) read speech and SVitchboard (SVB) conversational telephone speech, we find that the HMM transition probabilities may be more important than we once thought. The experiments include: (1) setting or not setting all outgoing transition probabilities equal; (2) the introduction of word-final triphones and the re-estimation of their transition probabilities; (3) besides grammar factor and insertion penalty, the addition of a third decoding parameter called transition factor to scale the transition probability score during decoding. The results of the above three experiments enable us to improve the the word accuracy of the WSJ and SVB speech recognition task by 0.7% and 5.3% absolute respectively when compared to their baseline model in which all transition probabilities are simply set to 0.5.
Guoli Ye, Dongpeng Chen, Brian Kan-Wing Mak
ICASSP1
2010 The use of subvector quantization and discrete densities for fast GMM computation for speaker verification
abstract
Last year, we showed that the computation of a GMM-UBM-based speaker verification (SV) system may be sped up by 30 times by using a high-density discrete model (HDDM) on the NIST 2002 evaluation task. The speedup was obtained using a special case of the product-code vector quantization in which each dimension is scalar-quantized in the construction of the discrete model. However, the speedup resulted in a drop of an absolute 1.5% in equal-error rate (EER). In this paper, our previous work is generalized to the use of subvector quantization (SVQ) in the construction of HDDM. For the same NIST 2002 SV task, the use of SVQ leads to an overall speedup by a factor of 8-25 with no significant loss in EER performance.
Guoli Ye, Brian Kan-Wing Mak
INTERSPEECH1
2009 Fast GMM computation for speaker verification using scalar quantization and discrete densities
abstract
Most of current state-of-the-art speaker verification (SV) sys-tems use Gaussian mixture model (GMM) to represent the uni-versal background model (UBM) and the speaker models (SM). For an SV system that employs log-likelihood ratio between SM and UBM to make the decision, its computational effi-ciency is largely determined by the GMM computation. This paper attempts to speedup GMM computation by converting a continuous-density GMM to a single or a mixture of discrete densities using scalar quantization. We investigated a spectrum of such discrete models: from high-density discrete models to discrete mixture models, and their combination called high-density discrete-mixture models. For the NIST 2002 SV task, we obtained an overall speedup by a factor of 2–100 with little loss in EER performance. Index Terms: speaker verification, scalar quantization, high density discrete HMM, discrete mixture HMM
Guoli Ye, Brian Kan-Wing Mak, Man-Wai Mak
INTERSPEECH1