VLDB 2026 Research / reviewers in the wild / expert
Khe Chai Sim
dblp:78/6873
· DBLP profile ↗
124ranked-venue papers
31as first author
24since 2021 · last 2024
0000-0002-0866-2223ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 102 · 25 first-author · 22 since 2021Artificial intelligence and machine learning · 66 · 19 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 2Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Improving Speech Recognition for African American English with Audio ClassificationabstractAutomatic speech recognition (ASR) systems have been shown to have large quality disparities between the language varieties they are intended or expected to recognize. One way to mitigate this is to train or fine-tune models with more representative datasets. But this approach can be hindered by limited in-domain data for training and evaluation. We propose a new way to improve the robustness of a US English short-form speech recognizer using a small amount of out-of-domain (long-form) African American English (AAE) data. We use CORAAL, YouTube and Mozilla Common Voice to train an audio classifier to approximately output whether an utterance is AAE or some other variety including Mainstream American English (MAE). By combining the classifier output with coarse geographic information, we can select a subset of utterances from a large corpus of untranscribed short-form queries for semi-supervised learning at scale. Fine-tuning on this data results in a 38.5% relative word error rate disparity reduction between AAE and MAE without reducing MAE quality. Shefali Garg, Zhouyuan Huo, Khe Chai Sim, Suzan Schwartz, Mason Chua, Alëna Aksënova, Tsendsuren Munkhdalai, Levi King, Darryl Wright, Zion Mengesha, Dongseong Hwang, Tara N. Sainath, Françoise Beaufays, Pedro J. Moreno 0001 |
ICASSP | 3 |
| 2024 | A Comparison of Parameter-Efficient ASR Domain Adaptation Methods for Universal Speech and Language ModelsabstractA recent paradigm shift in artificial intelligence has seen the rise of foundation models, such as the large language models and the universal speech models. With billions of model parameters and trained with a wide range of data, these foundation models are expected to have a better generalization to different downstream tasks. Efficient adaptation is the key to leveraging these foundation models in a new task or domain. In this paper, we compare several popular parameter-efficient tuning methods, such as vector adaptation, residual adapters, low-rank adapter (LoRA) and prompt-tuning, for automatic speech recognition (ASR) domain adaptation. We use the connectionist temporal classification (CTC) model with Conformer encoder and fused it with a universal language model. We study the effect of adapting either or both of the Conformer encoder and the universal language model. We carry out extensive experiments to study these methods under different hyper-parameter settings and the effect of combining some of these methods. We find that combining vector adaptation and residual adapters with increasing bottleneck dimension achieved the best performance. Khe Chai Sim, Zhouyuan Huo, Tsendsuren Munkhdalai, Nikhil Siddhartha, Adam Stooke, Zhong Meng, Bo Li 0028, Tara N. Sainath |
ICASSP | 1 |
| 2024 | AdaRA: Adaptive Rank Allocation of Residual Adapters for Speech Foundation Model
Zhouyuan Huo, Dongseong Hwang, Gan Song, Khe Chai Sim |
INTERSPEECH | 4 |
| 2024 | Contextual Biasing with the Knuth-Morris-Pratt Matching Algorithm
Zelin Wu, Diamantino Caseiro, Tsendsuren Munkhdalai, Khe Chai Sim, Pat Rondon, Golan Pundak, Gan Song, Rohit Prabhavalkar, Zhong Meng, Ding Zhao, Tara Sainath, Yanzhang He, Pedro J. Moreno 0001 |
INTERSPEECH | 5 |
| 2024 | Massive End-to-end Speech Recognition Models with Time ReductionabstractWeiran Wang, Rohit Prabhavalkar, Haozhe Shan, Zhong Meng, Dongseong Hwang, Qiujia Li, Khe Chai Sim, Bo Li, James Qin, Xingyu Cai, Adam Stooke, Chengjian Zheng, Yanzhang He, Tara Sainath, Pedro Moreno Mengibar. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Rohit Prabhavalkar, Haozhe Shan, Zhong Meng, Dongseong Hwang, Qiujia Li, Khe Chai Sim, Bo Li 0028, James Qin, Xingyu Cai, Adam Stooke, Chengjian Zheng, Yanzhang He, Tara N. Sainath, Pedro J. Moreno 0001 |
NAACL-HLT | 7 |
| 2024 | Aligner-Encoders: Self-Attention Transformers Can Be Self-TransducersabstractModern systems for automatic speech recognition, including the RNN-Transducer and Attention-based Encoder-Decoder (AED), are designed so that the encoder is not required to alter the time-position of information from the audio sequence into the embedding; alignment to the final text output is processed during decoding. We discover that the transformer-based encoder adopted in recent years is actually capable of performing the alignment internally during the forward pass, prior to decoding. This new phenomenon enables a simpler and more efficient model, the ''Aligner-Encoder''. To train it, we discard the dynamic programming of RNN-T in favor of the frame-wise cross-entropy loss of AED, while the decoder employs the lighter text-only recurrence of RNN-T without learned cross-attention---it simply scans embedding frames in order from the beginning, producing one token each until predicting the end-of-message. We conduct experiments demonstrating performance remarkably close to the state of the art, including a special inference configuration enabling long-form recognition. In a representative comparison, we measure the total inference time for our model to be 2x faster than RNN-T and 16x faster than AED. Lastly, we find that the audio-text alignment is clearly visible in the self-attention weights of a certain layer, which could be said to perform ''self-transduction''. Adam Stooke, Rohit Prabhavalkar, Khe Chai Sim, Pedro J. Moreno 0001 |
NeurIPS | 3 |
| 2023 | Contextual Spelling Correction with Large Language ModelsabstractContextual Spelling Correction (CSC) models are used to improve automatic speech recognition (ASR) quality given userspecific context. Typically, context is modeled as a large set of text spans to compare against a given ASR hypothesis using some distance measure (text, phonetic, or neural embedding). In this work we propose a CSC system based on a single Large Language Model (LLM) adapted with prompt tuning. Our approach is shown to be data efficient, and does not require dedicated serving. Our system exhibits advanced contextualization capabilities, such as support for phonetic spellings, cross-lingual scripts, and context specified as topics, with little to no data engineering. On voice assistant datasets, our system achieves $7.8 \%$ absolute word error rate reduction from a reference ASR system with relevant context and improving upon other contextualization solutions. Finally, we test our system in a prompt-injection attack scenario and report vulnerabilities and mitigations. Gan Song, Zelin Wu, Golan Pundak, Angad Chandorkar, Kandarp Joshi, Xavier Velez, Diamantino Caseiro, Ben Haynor, Nikhil Siddhartha, Pat Rondon, Khe Chai Sim |
ASRU | 12 |
| 2023 | Resource-Efficient Transfer Learning from Speech Foundation Model Using Hierarchical Feature FusionabstractSelf-supervised pre-training of a speech foundation model, followed by supervised fine-tuning, has shown impressive quality improvements on automatic speech recognition (ASR) tasks. Fine-tuning separate foundation models for many downstream tasks are expensive since the foundation model is usually very big. Parameter-efficient fine-tuning methods (e.g. adapter, sparse update methods) offer an alternative paradigm where a small set of parameters are updated to adapt the foundation model to new tasks. However, these methods still suffer from a high computational memory cost and slow training speed because they require backpropagation through the entire neural network at each step. In the paper, we analyze the performance of features at different layers of a foundation model on the speech recognition task and propose a novel hierarchical feature fusion method for resource-efficient transfer learning from speech foundation models. Experimental results show that the proposed method can achieve better performance on speech recognition task than existing algorithms with fewer number of trainable parameters, less computational memory cost and faster training speed. After combining with Adapters at all layers, the proposed method can achieve the same performance as fine-tuning the whole model with 97% fewer trainable encoder parameters and 53% faster training speed. Zhouyuan Huo, Khe Chai Sim, Bo Li 0028, Dongseong Hwang, Tara N. Sainath, Trevor Strohman |
ICASSP | 2 |
| 2023 | Comparison of Soft and Hard Target RNN-T Distillation for Large-Scale ASRabstractKnowledge distillation is an effective machine learning technique to transfer knowledge from a teacher model to a smaller student model, especially with unlabeled data. In this paper, we focus on knowledge distillation for the RNN-T model, which is widely used in state-of-the-art (SoTA) automatic speech recognition (ASR). Specifically, we compared using soft and hard target distillation to train large-scale RNN-T models on the LibriSpeech/LibriLight public dataset (60k hours) and our in-house data (600k hours). We found that hard targets are more effective when the teacher and student have different architecture, such as large teacher and small streaming student. On the other hand, soft target distillation works better in self-training scenario like iterative large teacher training. For a large model with 0.6B weights, we achieve a new SoTA word error rate (WER) on LibriSpeech (8% relative improvement on dev-other) using Noisy Student Training with soft target distillation. It also allows our production teacher to adapt new data domain continuously. Dongseong Hwang, Khe Chai Sim, Yu Zhang 0033, Trevor Strohman |
ICASSP | 2 |
| 2023 | Efficient Domain Adaptation for Speech Foundation ModelsabstractFoundation models (FMs), that are trained on broad data at scale and are adaptable to a wide range of downstream tasks, have brought large interest in the research community. Benefiting from the diverse data sources such as different modalities, languages and application domains, foundation models have demonstrated strong generalization and knowledge transfer capabilities. In this paper, we present a pioneering study towards building an efficient solution for FM-based speech recognition systems. We adopt the recently developed self-supervised BEST-RQ for pretraining, and extend the joint training strategy JUST Hydra for finetuning using both source and unsuper-vised target domain data. The FM encoder adapter and decoder are then finetuned to the target domain with a small amount of super-vised in-domain data. On a large-scale YouTube and Voice Search task, our method is shown to be both data and model parameter efficient. It achieves the same quality with only 21.6M supervised in-domain data and 130.8M finetuned parameters, compared to the 731.1M model trained from scratch on additional 300M supervised in-domain data. Bo Li 0028, Dongseong Hwang, Zhouyuan Huo, Junwen Bai, Guru Prakash Arumugam, Tara N. Sainath, Khe Chai Sim, Yu Zhang 0033, Wei Han 0002, Trevor Strohman, Françoise Beaufays |
ICASSP | 7 |
| 2023 | Re-investigating the Efficient Transfer Learning of Speech Foundation Model using Feature Fusion Methods
Zhouyuan Huo, Khe Chai Sim, Dongseong Hwang, Tsendsuren Munkhdalai, Tara N. Sainath, Pedro J. Moreno 0001 |
INTERSPEECH | 2 |
| 2023 | Dual-Mode NAM: Effective Top-K Context Injection for End-to-End ASR
Zelin Wu, Tsendsuren Munkhdalai, Pat Rondon, Golan Pundak, Khe Chai Sim, Christopher Li |
INTERSPEECH | 5 |
| 2022 | Joint Unsupervised and Supervised Training for Multilingual ASRabstractSelf-supervised training has shown promising gains in pretraining models and facilitating the downstream finetuning for speech recognition, like multilingual ASR. Most existing methods adopt a 2-stage scheme where the self-supervised loss is optimized in the first pretraining stage, and the standard supervised finetuning resumes in the second stage. In this paper, we propose an end-to-end (E2E) Joint Unsupervised and Supervised Training (JUST) method to combine the supervised RNN-T loss and the self-supervised contrastive and masked language modeling (MLM) losses. We validate its performance on the public dataset Multilingual LibriSpeech (MLS), which includes 8 languages and is extremely imbalanced. On MLS, we explore (1) JUST trained from scratch, and (2) JUST finetuned from a pretrained checkpoint. Experiments show that JUST can consistently outperform other existing state-of-the-art methods, and beat the monolingual baseline by a significant margin, demonstrating JUST’s capability of handling low-resource languages in multilingual ASR. Our average WER of all languages outperforms average monolingual baseline by 33.3%, and the state-of-the-art 2-stage XLSR by 32%. On low-resource languages like Polish, our WER is less than half of the monolingual baseline and even beats the supervised transfer learning method which uses external supervision. Junwen Bai, Bo Li 0028, Yu Zhang 0033, Ankur Bapna, Nikhil Siddhartha, Khe Chai Sim, Tara N. Sainath |
ICASSP | 6 |
| 2022 | Large-Scale ASR Domain Adaptation Using Self- and Semi-Supervised LearningabstractSelf- and semi-supervised learning methods have been actively investigated to reduce labeled training data or enhance model performance. However, these approaches mostly focus on in-domain performance for public datasets. In this study, we utilize the combination of self- and semi-supervised learning methods to solve unseen domain adaptation problems in a large-scale production setting for online ASR model. This approach demonstrates that using the source domain data with a small fraction of the target domain data (3%) can recover the performance gap compared to a full data baseline: 13.5% relative WER improvement for target domain data. Dongseong Hwang, Ananya Misra, Zhouyuan Huo, Nikhil Siddhartha, Shefali Garg, David Qiu, Khe Chai Sim, Trevor Strohman, Françoise Beaufays, Yanzhang He |
ICASSP | 7 |
| 2022 | Fast Contextual Adaptation with Neural Associative Memory for On-Device Personalized Speech RecognitionabstractFast contextual adaptation has shown to be effective in improving Automatic Speech Recognition (ASR) of rare words and when combined with an on-device personalized training, it can yield an even better recognition result. However, the traditional re-scoring approaches based on an external language model is prone to diverge during the personalized training. In this work, we introduce a model-based end-to-end contextual adaptation approach that is decoder-agnostic and amenable to on-device personalization. Our on-device simulation experiments demonstrate that the proposed approach outperforms the traditional re-scoring technique by 12% relative WER and 15.7% entity mention specific F1-score in a continuous personalization scenario. Tsendsuren Munkhdalai, Khe Chai Sim, Angad Chandorkar, Mason Chua, Trevor Strohman, Françoise Beaufays |
ICASSP | 2 |
| 2022 | UserLibri: A Dataset for ASR Personalization Using Only Text
Theresa Breiner, Swaroop Ramaswamy, Ehsan Variani, Shefali Garg, Rajiv Mathews, Khe Chai Sim, Kilol Gupta, Mingqing Chen, Lara McConnaughey |
INTERSPEECH | 6 |
| 2022 | Incremental Layer-Wise Self-Supervised Learning for Efficient Unsupervised Speech Domain Adaptation On Device
Zhouyuan Huo, Dongseong Hwang, Khe Chai Sim, Shefali Garg, Ananya Misra, Nikhil Siddhartha, Trevor Strohman, Françoise Beaufays |
INTERSPEECH | 3 |
| 2022 | Pseudo Label Is Better Than Human LabelabstractState-of-the-art automatic speech recognition (ASR) systems are trained with tens of thousands of hours of labeled speech data.Human transcription is expensive and time consuming.Factors such as the quality and consistency of the transcription can greatly affect the performance of the ASR models trained with these data.In this paper, we show that we can train a strong teacher model to produce high quality pseudo labels by utilizing recent self-supervised and semi-supervised learning techniques.Specifically, we use JUST (Joint Unsupervised/Supervised Training) and iterative noisy student teacher training to train a 600 million parameter bi-directional teacher model.This model achieved 4.0% word error rate (WER) on a voice search task, 11.1% relatively better than a baseline.We further show that by using this strong teacher model to generate high-quality pseudo labels for training, we can achieve 13.6% relative WER reduction (5.9% to 5.1%) for a streaming model compared to using human labels. Dongseong Hwang, Khe Chai Sim, Zhouyuan Huo, Trevor Strohman |
INTERSPEECH | 2 |
| 2022 | On-the-fly ASR Corrections with Audio Exemplars
Golan Pundak, Tsendsuren Munkhdalai, Khe Chai Sim |
INTERSPEECH | 3 |
| 2022 | NAM+: Towards Scalable End-to-End Contextual Biasing for Adaptive ASRabstractAttention-based biasing techniques for end-to-end ASR systems are able to achieve large accuracy gains without requiring the inference algorithm adjustments and parameter tuning common to fusion approaches. However, it is challenging to simultaneously scale up attention-based biasing to realistic numbers of biased phrases; maintain in-domain WER gains, while minimizing out-of-domain losses; and run in real time. We present NAM+, an attention-based biasing approach which achieves a 16X inference speedup per acoustic frame over prior work when run with 3,000 biasing entities, as measured on a typical mobile CPU. NAM+ achieves these run-time gains through a combination of Two-Pass Hierarchical Attention and Dilated Context Update. Compared to the adapted baseline, NAM+ further decreases the in-domain WER by up to 12.6% relative, while incurring an out-of-domain WER regression of 20% relative. Compared to the non-adapted baseline, the out-of-domain WER regression is 7.1 % relative. Tsendsuren Munkhdalai, Zelin Wu, Golan Pundak, Khe Chai Sim, Pat Rondon, Tara N. Sainath |
SLT | 4 |
| 2022 | Context-Aware Neural Confidence Estimation for Rare Word Speech RecognitionabstractConfidence estimation for automatic speech recognition (ASR) is important for many downstream tasks. Recently, neural confidence estimation models (CEMs) have been shown to produce accurate confidence scores for predicting word-level errors. These models are built on top of an end-to-end (E2E) ASR and the acoustic embeddings are part of the input features. However, practical E2E ASR systems often incorporate contextual information in the decoder to improve rare word recognition. The CEM is not aware of this and underestimates the confidence of the rare words that have been corrected by the context. In this paper, we propose a context-aware CEM by incorporating context into the encoder using a neural associative memory (NAM) model. It uses attention to detect for presence of the biasing phrases and modify the encoder features. Experiments show that the proposed context-aware CEM using NAM augmented training can improve the AUC-ROC for word error prediction from 0.837 to 0.892. David Qiu, Tsendsuren Munkhdalai, Yanzhang He, Khe Chai Sim |
SLT | 4 |
| 2022 | Internal Language Model Personalization of E2E Automatic Speech Recognition Using Random Encoder FeaturesabstractEnd-to-end (E2E) speech-to-text models generally require transcribed audio for training and personalization. We introduce the use of random audio encoder features, rather than speech, to fine-tune the final model layers and acquire new vocabulary from text-only data. This technique can be used for on-device personalization before the user has provided any speech data. We show improvements in the recall of new vocabulary and word error rate (WER) on held-out test sets using simulated user experiments on hybrid autoregressive transducer (HAT) models using conformer-based encoders and simple text embeddings for label processing. We compare this approach to the use of synthetic audio, finding random encoder features to be more beneficial with lower computational cost. Experiments show that the maximum benefit is gained by updating specific network components comprising a subset of those expressing the internal language model. Adam Stooke, Khe Chai Sim, Mason Chua, Tsendsuren Munkhdalai, Trevor Strohman |
SLT | 2 |
| 2021 | A Comparison of Supervised and Unsupervised Pre-Training of End-to-End Models
Ananya Misra, Dongseong Hwang, Zhouyuan Huo, Shefali Garg, Nikhil Siddhartha, Arun Narayanan, Khe Chai Sim |
Interspeech | 7 |
| 2021 | Robust Continuous On-Device Personalization for Automatic Speech Recognition
Khe Chai Sim, Angad Chandorkar, Mason Chua, Tsendsuren Munkhdalai, Françoise Beaufays |
Interspeech | 1 |
| 2020 | Low-Rank Gradient Approximation for Memory-Efficient on-Device Training of Deep Neural NetworkabstractTraining machine learning models on mobile devices has the potential of improving both privacy and accuracy of the models. However, one of the major obstacles to achieving this goal is the memory limitation of mobile devices. Reducing training memory enables models with high-dimensional weight matrices, like automatic speech recognition (ASR) models, to be trained on-device. In this paper, we propose approximating the gradient matrices of deep neural networks using a low-rank parameterization as an avenue to save training memory. The low-rank gradient approximation enables more advanced, memory-intensive optimization techniques to be run on device. Our experimental results show that we can reduce the training memory by about 33.0% for Adam optimization. It uses comparable memory to momentum optimization and achieves a 4.5% relative lower word error rate on an ASR personalization task. Mary Gooneratne, Khe Chai Sim, Petr Zadrazil, Andreas Kabel, Françoise Beaufays, Giovanni Motta |
ICASSP | 2 |
| 2019 | Personalization of End-to-End Speech Recognition on Mobile Devices for Named EntitiesabstractWe study the effectiveness of several techniques to personalize end-to-end speech models and improve the recognition of proper names relevant to the user. These techniques differ in the amounts of user effort required to provide supervision, and are evaluated on how they impact speech recognition performance. We propose using keyword-dependent precision and recall metrics to measure vocabulary acquisition performance. We evaluate the algorithms on a dataset that we designed to contain names of persons that are difficult to recognize. Therefore, the baseline recall rate for proper names in this dataset is very low: 2.4%. A data synthesis approach we developed brings it to 48.6%, with no need for speech input from the user. With speech input, if the user corrects only the names, the name recall rate improves to 64.4%. If the user corrects all the recognition errors, we achieve the best recall of 73.5%. To eliminate the need to upload user data and store personalized models on a server, we focus on performing the entire personalization workflow on a mobile device. Khe Chai Sim, Leif Johnson, Giovanni Motta, Lillian Zhou, Françoise Beaufays, Arnaud Benard, Dhruv Guliani, Andreas Kabel, Nikhil Khare, Tamar Lucassen, Petr Zadrazil, Harry Zhang |
ASRU | 1 |
| 2019 | Streaming End-to-end Speech Recognition for Mobile DevicesabstractEnd-to-end (E2E) models, which directly predict output character sequences given input speech, are good candidates for on-device speech recognition. E2E models, however, present numerous challenges: In order to be truly useful, such models must decode speech utterances in a streaming fashion, in real time; they must be robust to the long tail of use cases; they must be able to leverage user-specific context (e.g., contact lists); and above all, they must be extremely accurate. In this work, we describe our efforts at building an E2E speech recog-nizer using a recurrent neural network transducer. In experimental evaluations, we find that the proposed approach can outperform a conventional CTC-based model in terms of both latency and accuracy in a number of evaluation categories. Yanzhang He, Tara N. Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Ruoming Pang, Qiao Liang 0001, Deepti Bhatia, Yuan Shangguan, Bo Li 0028, Golan Pundak, Khe Chai Sim, Tom Bagby, Shuo-Yiin Chang, Kanishka Rao, Alexander Gruenstein |
ICASSP | 16 |
| 2019 | Improving CTC Using Stimulated Learning for Sequence ModelingabstractConnectionist temporal classification (CTC) is a sequence-level loss that has been successfully applied to train recurrent neural network (RNN) models for automatic speech recognition. However, one major weakness of CTC is the conditional independence assumption that makes it difficult for the model to learn label dependencies. In this paper, we propose stimulated CTC, which uses stimulated learning to help CTC models learn label dependencies implicitly by using an auxiliary RNN to generate the appropriate stimuli. This stimuli comes in the form of an additional stimulation loss term which encourages the model to learn said label dependencies. The auxiliary network is only used during training and the inference model has the same structure as a standard CTC model. The proposed stimulated CTC model achieves about 35 % relative character error rate improvements on a synthetic gesture keyboard recognition task and over 30 % relative word error rate improvements on the Librispeech automatic speech recognition tasks over a baseline model trained with CTC only. Jahn Heymann, Khe Chai Sim, Bo Li 0028 |
ICASSP | 2 |
| 2019 | An Investigation into On-Device Personalization of End-to-End Automatic Speech Recognition ModelsabstractSpeaker-independent speech recognition systems trained with data from many users are generally robust against speaker variability and work well for a large population of speakers. However, these systems do not always generalize well for users with very different speech characteristics. This issue can be addressed by building personalized systems that are designed to work well for each specific user. In this paper, we investigate the idea of securely training personalized end-to-end speech recognition models on mobile devices so that user data and models never leave the device and are never stored on a server. We study how the mobile training environment impacts performance by simulating on-device data consumption. We conduct experiments using data collected from speech impaired users for personalization. Our results show that personalization achieved 63.7\% relative word error rate reduction when trained in a server environment and 58.1% in a mobile environment. Moving to on-device personalization resulted in 18.7% performance degradation, in exchange for improved scalability and data privacy. To train the model on device, we split the gradient computation into two and achieved 45% memory reduction at the expense of 42% increase in training time. Khe Chai Sim, Petr Zadrazil, Françoise Beaufays |
INTERSPEECH | 1 |
| 2018 | Understanding Recurrent Neural State Using Memory SignaturesabstractWe demonstrate a network visualization technique to analyze the recurrent state inside the LSTMs/GRUs used commonly in language and acoustic models. Interpreting intermediate state and network activations inside end-to-end models remains an open challenge. Our method allows users to understand exactly how much and what history is encoded inside recurrent state in grapheme sequence models. Our procedure trains multiple decoders that predict prior input history. Compiling results from these decoders, a user can obtain a signature of the recurrent kernel that characterizes its memory behavior. We demonstrate this method's usefulness in revealing information divergence in the bases of recurrent factorized kernels, visualizing the character-level differences between the memory of n-gram and recurrent language models, and extracting knowledge of history encoded in the layers of grapheme-based end-to-end ASR networks. Skanda Koppula, Khe Chai Sim, Kean K. Chin |
ICASSP | 2 |
| 2018 | Multi-Dialect Speech Recognition with a Single Sequence-to-Sequence ModelabstractSequence-to-sequence models provide a simple and elegant solution for building speech recognition systems by folding separate components of a typical system, namely acoustic (AM), pronunciation (PM) and language (LM) models into a single neural network. In this work, we look at one such sequence-to-sequence model, namely listen, attend and spell (LAS) [1], and explore the possibility of training a single model to serve different English dialects, which simplifies the process of training multi-dialect systems without the need for separate AM, PM and LMs for each dialect. We show that simply pooling the data from all dialects into one LAS model falls behind the performance of a model fine-tuned on each dialect. We then look at incorporating dialect-specific information into the model, both by modifying the training targets by inserting the dialect symbol at the end of the original grapheme sequence and also feeding a 1-hot representation of the dialect information into all layers of the model. Experimental results on seven English dialects show that our proposed system is effective in modeling dialect variations within a single LAS model, outperforming a LAS model trained individually on each of the seven dialects by 3.1~16.5% relative. Bo Li 0028, Tara N. Sainath, Khe Chai Sim, Michiel Bacchiani, Eugene Weinstein, Patrick Nguyen, Yanghui Wu, Kanishka Rao |
ICASSP | 3 |
| 2018 | learning Effective Factorized Hidden Layer Bases Using Student-Teacher Training for LSTM Acoustic Model AdaptationabstractFactorized Hidden Layer (FHL) has been proposed for the adaptation of deep neural network (DNN) and Long Short-Term Memory (LSTM) based acoustic models (AMs). In FHL, a speaker-dependent (SD) transformation matrix and an SD bias are included in addition to the standard affine transformation. The SD transformation is a linear combination of rank- l matrices whereas the SD bias is a linear combination of vectors. However, the adaptation of LSTMs is challenging and often reports modest gains. In this paper, we propose to use student-teacher training to estimate more efficient FHL bases for LSTM AMs using an FHL adapted DNN as the teacher model. For both AMI IHM and AMI SDM tasks, FHL achieves 3.2% absolute improvement over the frame-level cross entropy trained LSTM baselines. Moreover, FHL results 3.0% and 3.8% absolute improvements over sequentially trained LSTM baselines for the AMI IHM and AMI SDM tasks respectively. Lahiru Samarakoon, Brian Kan-Wing Mak, Khe Chai Sim |
ICASSP | 3 |
| 2018 | Domain Adaptation Using Factorized Hidden Layer for Robust Automatic Speech Recognition
Khe Chai Sim, Arun Narayanan, Ananya Misra, Anshuman Tripathi, Golan Pundak, Tara N. Sainath, Parisa Haghani, Bo Li 0028, Michiel Bacchiani |
INTERSPEECH | 1 |
| 2018 | Efficient Implementation of Recurrent Neural Network Transducer in TensorflowabstractRecurrent neural network transducer (RNN-T) has been successfully applied to automatic speech recognition to jointly learn the acoustic and language model components. The RNN-T loss and its gradient with respect to the softmax outputs can be computed efficiently using a forward-backward algorithm. In this paper, we present an efficient implementation of the RNN-T forward-backward and Viterbi algorithms using standard matrix operations. This allows us to easily implement the algorithm in TensorFlow by making use of the existing hardware-accelerated implementations of these operations. This work is based on a similar technique used in our previous work for computing the connectionist temporal classification and lattice-free maximum mutual information losses, where the forward and backward recursions are viewed as a bi-directional RNN whose states represent the forward and backward probabilities. Our benchmark results on graphic processing unit (GPU) and tensor processing unit (TPU) show that our implementation can achieve better throughput performance by increasing the batch size to maximize parallel computation. Furthermore, our implementation is about twice as fast on TPU compared to GPU for batch. Tom Bagby, Kanishka Rao, Khe Chai Sim |
SLT | 3 |
| 2018 | Toward Domain-Invariant Speech Recognition via Large Scale TrainingabstractCurrent state-of-the-art automatic speech recognition systems are trained to work in specific `domains', defined based on factors like application, sampling rate and codec. When such recognizers are used in conditions that do not match the training domain, performance significantly drops. This work explores the idea of building a single domain-invariant model for varied use-cases by combining large scale training data from multiple application domains. Our final system is trained using 162,000 hours of speech. Additionally, each utterance is artificially distorted during training to simulate effects like background noise, codec distortion, and sampling rates. Our results show that, even at such a scale, a model thus trained works almost as well as those fine-tuned to specific subsets: A single model can be robust to multiple application domains, and variations like codecs and noise. More importantly, such models generalize better to unseen conditions and allow for rapid adaptation - we show that by using as little as 10 hours of data from a new domain, an adapted domain-invariant model can match performance of a domain-specific model trained from scratch using 70 times as much data. We also highlight some of the limitations of such models and areas that need addressing in future work. Arun Narayanan, Ananya Misra, Khe Chai Sim, Golan Pundak, Anshuman Tripathi, Mohamed G. Elfeky, Parisa Haghani, Trevor Strohman, Michiel Bacchiani |
SLT | 3 |
| 2018 | Improving Interpretability and Regularization in Deep LearningabstractDeep learning approaches yield state-of-the-art performance in a range of tasks, including automatic speech recognition. However, the highly distributed representation in a deep neural network (DNN) or other network variations is difficult to analyze, making further parameter interpretation and regularization challenging. This paper presents a regularization scheme acting on the activation function output to improve the network interpretability and regularization. The proposed approach, referred to as activation regularization, encourages activation function outputs to satisfy a target pattern. By defining appropriate target patterns, different learning concepts can be imposed on the network. This method can aid network interpretability and also has the potential to reduce overfitting. The scheme is evaluated on several continuous speech recognition tasks: the Wall Street Journal continuous speech recognition task, eight conversational telephone speech tasks from the IARPA Babel program and a U.S. English broadcast news task. On all the tasks, the activation regularization achieved consistent performance gains over the standard DNN baselines. Chunyang Wu, Mark J. F. Gales, Anton Ragni, Panagiota Karanasou, Khe Chai Sim |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2017 | Improving the efficiency of forward-backward algorithm using batched computation in TensorFlowabstractSequence-level losses are commonly used to train deep neural network acoustic models for automatic speech recognition. The forward-backward algorithm is used to efficiently compute the gradients of the sequence loss with respect to the model parameters. Gradient-based optimization is used to minimize these losses. Recent work has shown that the forward-backward algorithm can be efficiently implemented as a series of matrix operations. This paper further improves the forward-backward algorithm via batched computation, a technique commonly used to improve training speed by exploiting the parallel computation of matrix multiplication. Specifically, we show how batched computation of the forward-backward algorithm can be efficiently implemented using TensorFlow to handle variable-length sequences within a mini batch. Furthermore, we also show how the batched forward-backward computation can be used to compute the gradients of the connectionist temporal classification (CTC) and maximum mutual information (MMI) losses with respect to the logits. We show, via empirical benchmarks, that the batched forward-backward computation can speed up the CTC loss and gradient computation by about 183 times when run on GPU with a batch size of 256 compared to using a batch size of 1; and by about 22 times for lattice-free MMI using a trigram phone language model for the denominator. Khe Chai Sim, Arun Narayanan, Tom Bagby, Tara N. Sainath, Michiel Bacchiani |
ASRU | 1 |
| 2017 | An investigation into learning effective speaker subspaces for robust unsupervised DNN adaptationabstractSubspace methods are used for deep neural network (DNN)-based acoustic model adaptation. These methods first construct a subspace and then perform the speaker adaptation as a point in the subspace. This paper aims to investigate the effectiveness of subspace methods for robust unsupervised adaptation. For the analysis, we compare two state-of-the-art subspace methods, namely, the singular value decomposition (SVD)-based bottleneck adaptation and the factorized hidden layer (FHL) adaptation. Both of these methods perform speaker adaptation as a linear combination of rank-1 bases. The main difference between the subspace construction is that FHL adaptation constructs a speaker subspace separate from the phoneme classification space while SVD-based bottleneck adaptation shares the same subspace for both the phoneme classification and the speaker adaptation. So far, no direct comparisons between these two methods are reported. In this work, we compare these two methods for their robustness to unsupervised adaptation on Aurora 4, AMI IHM and AMI SDM tasks. Our findings show that the FHL adaptation outperforms the SVD-based bottleneck adaptation especially in challenging conditions where the adaptation data is limited, or the quality of the adaptation alignments are low. Lahiru Samarakoon, Khe Chai Sim, Brian Kan-Wing Mak |
ICASSP | 2 |
| 2017 | Acoustic Modeling for Google Home
Bo Li 0028, Tara N. Sainath, Arun Narayanan, Joe Caroselli, Michiel Bacchiani, Ananya Misra, Izhak Shafran, Hasim Sak, Golan Pundak, Kean K. Chin, Khe Chai Sim, Ron J. Weiss, Kevin W. Wilson, Ehsan Variani, Chanwoo Kim 0001, Olivier Siohan, Mitch Weintraub, Erik McDermott, Richard Rose, Matt Shannon |
INTERSPEECH | 11 |
| 2017 | Learning Factorized Transforms for Unsupervised Adaptation of LSTM-RNN Acoustic ModelsabstractFactorized Hidden Layer (FHL) adaptation has been proposed for speaker adaptation of deep neural network (DNN) based acoustic models. In FHL adaptation, a speaker-dependent (SD) transformation matrix and an SD bias are included in addition to the standard affine transformation. The SD transformation is a linear combination of rank-1 matrices whereas the SD bias is a linear combination of vectors. Recently, the Long Short- Term Memory (LSTM) Recurrent Neural Networks (RNNs) have shown to outperform DNN acoustic models in many Automatic Speech Recognition (ASR) tasks. In this work, we investigate the effectiveness of SD transformations for LSTM-RNN acoustic models. Experimental results show that when combined with scaling of LSTM cell states' outputs, SD transformations achieve 2.3% and 2.1% absolute improvements over the baseline LSTM systems for the AMI IHM and AMI SDM tasks respectively. Lahiru Samarakoon, Brian Kan-Wing Mak, Khe Chai Sim |
INTERSPEECH | 3 |
| 2017 | An Efficient Phone N-Gram Forward-Backward Computation Using Dense Matrix Multiplication
Khe Chai Sim, Arun Narayanan |
INTERSPEECH | 1 |
| 2016 | Joint acoustic factor learning for robust deep neural network based automatic speech recognitionabstractDeep neural networks (DNNs) for acoustic modeling have been shown to provide impressive results on many state-of-the-art automatic speech recognition (ASR) applications. However, DNN performance degrades due to mismatches in training and testing conditions and thus adaptation is necessary. In this paper, we explore the use of discriminative auxiliary input features obtained using joint acoustic factor learning for DNN adaptation. These features are derived from a bottleneck (BN) layer of a DNN and are referred to as BN vectors. To derive these BN vectors, we explore the use of two types of joint acoustic factor learning which capture speaker and auxiliary information such as noise, phone and articulatory information of speech. In this paper, we show that these BN vectors can be used for adaptation and thereby improve the performance of an ASR system. We also show that the performance can be further improved on augmenting these BN vectors to conventional i-vectors. In this paper, experiments are performed on Aurora-4, REVERB challenge and AMI databases. Souvik Kundu 0003, Gautam Mantena, Yanmin Qian, Tian Tan 0002, Marc Delcroix, Khe Chai Sim |
ICASSP | 6 |
| 2016 | On combining i-vectors and discriminative adaptation methods for unsupervised speaker normalization in DNN acoustic modelsabstractIn automatic speech recognition (ASR), adaptation and adaptive training techniques are used to perform speaker normalization. Previous methods mainly focus on using these techniques in isolation. In contrast, this paper investigates two approaches to improve the ASR performance by combining i-vector based speaker adaptive training in deep neural network (DNN) acoustic models with discriminative adaptation techniques. First, we combine these techniques by interpolating the decoding lattices of i-vector based systems with the decoding lattices of a discriminatively adapted model. Then, we combine these methods by discriminatively adapting the i-vector based system in unsupervised fashion. Our experiments on TED-LIUM dataset show that compared with a strong speaker independent baseline, lattice interpolation and adaptation of the i-vector systems achieve 12.0% and 15.6% relative improvements, respectively. Moreover, in comparison to the i-vector based systems, lattice interpolation reported a 4.5% relative improvement while discriminatively adapting the i-vector system reported a 8.3% relative improvement. Lahiru Samarakoon, Khe Chai Sim |
ICASSP | 2 |
| 2016 | Speaker-aware training of LSTM-RNNS for acoustic modellingabstractLong Short-Term Memory (LSTM) is a particular type of recurrent neural network (RNN) that can model long term temporal dynamics. Recently it has been shown that LSTM-RNNs can achieve higher recognition accuracy than deep feed-forword neural networks (DNNs) in acoustic modelling. However, speaker adaption for LSTM-RNN based acoustic models has not been well investigated. In this paper, we study the LSTM-RNN speaker-aware training that incorporates the speaker information during model training to normalise the speaker variability. We first present several speaker-aware training architectures, and then empirically evaluate three types of speaker representation: I-vectors, bottleneck speaker vectors and speaking rate. Furthermore, to factorize the variability in the acoustic signals caused by speakers and phonemes respectively, we investigate the speaker-aware and phone-aware joint training under the framework of multi-task learning. In AMI meeting speech transcription task, speaker-aware training of LSTM-RNNs reduces word error rates by 6.5% relative to a very strong LSTM-RNN baseline, which uses FMLLR features. Tian Tan 0002, Yanmin Qian, Dong Yu 0001, Souvik Kundu 0003, Liang Lu 0001, Khe Chai Sim, Yu Zhang 0033 |
ICASSP | 6 |
| 2016 | Towards implicit complexity control using variable-depth deep neural networks for automatic speech recognitionabstractIn speech recognition, a trade-off can be made between transcription accuracy and computation time. In this paper, we empirically measure the performance of using the softmax outputs connected to different hidden layers of an already fine-tuned deep neural network (DNN) and explore decoding strategies that do not require computing all the hidden layers of the DNN. We find that selecting the specific outputs from a variable-depth DNN achieves better Phoneme Error Rates (PER) on the TIMIT task than directly training a fixed-depth DNN with the same number of layers. We experimented with different ways of stopping the forward-propagation early, first by using a threshold on the entropy of the respective outputs, and formulate a `gating' system on the hidden layers to predict when to stop the forward propagation. Shawn Tan, Khe Chai Sim |
ICASSP | 2 |
| 2016 | Incorporating a Generative Front-End Layer to Deep Neural Network for Noise Robust Automatic Speech Recognition
Souvik Kundu 0003, Khe Chai Sim, Mark J. F. Gales |
INTERSPEECH | 2 |
| 2016 | Microphone Distance Adaptation Using Cluster Adaptive Training for Robust Far Field Speech Recognition
Animesh Prasad, Khe Chai Sim |
INTERSPEECH | 2 |
| 2016 | Subspace LHUC for Fast Adaptation of Deep Neural Network Acoustic Models
Lahiru Samarakoon, Khe Chai Sim |
INTERSPEECH | 2 |
| 2016 | Multi-Attribute Factorized Hidden Layer Adaptation for DNN Acoustic Models
Lahiru Samarakoon, Khe Chai Sim |
INTERSPEECH | 2 |
| 2016 | Stimulated Deep Neural Network for Speech RecognitionabstractDeep neural networks (DNNs) and deep learning approaches yield state-of-the-art performance in a range of tasks, including speech recognition.However, the parameters of the network are hard to analyze, making network regularization and robust adaptation challenging.Stimulated training has recently been proposed to address this problem by encouraging the node activation outputs in regions of the network to be related.This kind of information aids visualization of the network, but also has the potential to improve regularization and adaptation.This paper investigates stimulated training of DNNs for both of these options.These schemes take advantage of the smoothness constraints that stimulated training offers.The approaches are evaluated on two large vocabulary speech recognition tasks: a U.S. English broadcast news (BN) task and a Javanese conversational telephone speech task from the IARPA Babel program.Stimulated DNN training acquires consistent performance gains on both tasks over unstimulated baselines.On the BN task, the proposed smoothing approach is also applied to rapid adaptation, again outperforming the standard adaptation scheme. Chunyang Wu, Panagiota Karanasou, Mark J. F. Gales, Khe Chai Sim |
INTERSPEECH | 4 |
| 2016 | Entropy-based pruning of hidden units to reduce DNN parametersabstractFor acoustic modeling, the use of DNN has become popular due to its superior performance improvements observed in many automatic speech recognition (ASR) tasks. Typically, DNNs with deep (many layers) and wide (many hidden units per layer) architectures are chosen in order to achieve good gains. An issue with such approaches is that there is an explosion in the number of learnable parameters. Thus, it is often difficult to build models in cases where there is no sufficient amount of training data (or data for adaptation), and also limits the usage of ASR systems on hand-held devices such as mobile phones. A method to overcome this issue is to reduce the number of parameters. In this work, we provide a framework to effectively reduce the number of parameters by removing the hidden units. Each hidden unit is represented by an activity vector associated with speech attributes such as phones. A normalized entropy-based measure is computed from these activity vectors which reflects the significance of these units in the DNN model. For comparison we also use low-rank matrix factorization to reduce the number of parameters. We show that low-rank matrix factorization can reduce the number of parameters only to a certain extent. Thus, we extend the pruning technique in combination with low-rank matrix factorization to further reduce the model. In this work, we provide detailed experimental results on the Aurora-4 and TEDLIUM databases and show that the models can be reduced to approximately 20 - 30% of its initial size without much loss in the ASR performance. Gautam Mantena, Khe Chai Sim |
SLT | 2 |
| 2016 | Low-rank bases for factorized hidden layer adaptation of DNN acoustic modelsabstractRecently, the factorized hidden layer (FHL) adaptation method is proposed for speaker adaptation of deep neural network (DNN) acoustic models. An FHL contains a speaker-dependent (SD) transformation matrix using a linear combination of rank-1 matrices and an SD bias using a linear combination of vectors, in addition to the standard affine transformation. On the other hand, full-rank bases are used with a similar DNN adaptation method which is based on cluster adaptive training (CAT). Therefore, it is interesting to investigate the effect of the rank of the bases used for adaptation. The increase of the rank of the bases improves the speaker subspace representation, without increasing the number of learnable speaker parameters. In this work, we investigate the effect of using various ranks for the bases of the SD transformation of FHLs on Aurora 4, AMI IHM and AMI SDM tasks. Experimental results have shown that when one FHL layer is used, it is optimal to use low-ranked bases of rank-50, instead of full-rank bases. Furthermore, when multiple FHLs are used, rank-1 bases are sufficient. Lahiru Samarakoon, Khe Chai Sim |
SLT | 2 |
| 2016 | Learning utterance-level normalisation using Variational Autoencoders for robust automatic speech recognitionabstractThis paper presents a Variational Autoencoder (VAE) based framework for modelling utterances. In this model, a mapping from an utterance to a distribution over the latent space, the VAE-utterance feature, is defined. This is in addition to a frame-level mapping, the VAE-frame feature. Using the Aurora-4 dataset, we train and perform some analysis on these models based on their detection of speaker and utterance variability, and also use combinations of LDA, i-vector, and VAE-frame and utterance features for speech recognition training. We find that it works equally well using VAE-frame + VAE-utterance features alone, and by using an LDA + VAE-frame +VAE-utterance feature combination, we obtain a word-errorrate (WER) of 9.59%, a gain over the 9.72% baseline which uses an LDA + i-vector combination. Shawn Tan, Khe Chai Sim |
SLT | 2 |
| 2016 | Factorized Hidden Layer Adaptation for Deep Neural Network Based Acoustic ModelingabstractIn this paper, we propose the factorized hidden layer (FHL) approach to adapt the deep neural network (DNN) acoustic models for automatic speech recognition (ASR). FHL aims at modeling speaker dependent (SD) hidden layers by representing an SD affine transformation as a linear combination of bases. The combination weights are low-dimensional speaker parameters that can be initialized using speaker representations like i-vectors and then reliably refined in an unsupervised adaptation fashion. Therefore, our method provides an efficient way to perform both adaptive training and (test-time) adaptation. Experimental results have shown that the FHL adaptation improves the ASR performance significantly, compared to the standard DNN models, as well as other state-of-the-art DNN adaptation approaches, such as training with the speaker-normalized CMLLR features, speaker-aware training using i-vector and learning hidden unit contributions (LHUC). For Aurora 4, FHL achieves 3.8% and 2.3% absolute improvements over the standard DNNs trained on the LDA + STC and CMLLR features, respectively. It also achieves 1.7% absolute performance improvement over a system that combines the i-vector adaptive training with LHUC adaptation. For the AMI dataset, FHL achieved 1.4% and 1.9% absolute improvements over the sequence-trained CMLLR baseline systems, for the IHM and SDM tasks, respectively. Lahiru Samarakoon, Khe Chai Sim |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Learning factorized feature transforms for speaker normalizationabstractThis paper proposes an approach to improve automatic speech recognition (ASR) by normalizing the speaker variability of a well trained Deep Neural Network (DNN) acoustic model using i-vectors. Our approach learns a speaker dependent transformation of the acoustic features combined with the standard speaker dependent bias, to minimize the mismatch due to the inter-speaker variability. Speaker normalization experiments on the Aurora 4 task show 10.9% relative improvement over the baseline. Moreover, the proposed approach reported 4.5% relative improvement over the standard i-vector based method where only a speaker dependent bias is used. Furthermore, we report an analysis to compare our approach with the Constrained Maximum Likelihood Linear Regression (CMLLR) method. Lahiru Samarakoon, Khe Chai Sim |
ASRU | 2 |
| 2015 | On constructing and analysing an interpretable brain model for the DNN based on hidden activity patternsabstractDeep Neural Network (DNN) has been well received as a powerful machine learning model in a wide range of pattern classification tasks. Despite its superior performance in handling complex real-world problems, DNNs have been used pretty much as a black box, without offering much insights in terms of how and why high quality classification performance has been achieved. To address this problem, this paper studies the DNN hidden unit activities and presents a novel interpretable DNN visualisation technique that projects the hidden units of the DNN onto a meaningful 2-dimensional subspace. The projected points are displayed with colours to reflect the activation values for the purpose of visualisation. In this paper, the proposed technique is used to visualise two DNN acoustic models trained on the multi-condition data from the Aurora 4 corpus. The technique is able to produce a two dimensional representation of the DNN "brain" with interpretable regions. It also accentuates the effect of how the behaviour of the hidden units changes across different layers. Khe Chai Sim |
ASRU | 1 |
| 2015 | Improving the interpretability of deep neural networks with stimulated learningabstractDeep Neural Networks (DNNs) have demonstrated improvements in acoustic modelling for automatic speech recognition. However, they are often used as a black box, and not much is understood about what each of the hidden layers does. We seek to understand how the activations in the hidden layers change with different input, and how we can leverage such knowledge to modify the behaviour of the model. To this end, we propose stimulated deep learning where stimuli are introduced during the DNN training process to influence the behaviour of the hidden units. Specifically, constraints are applied so that the hidden units of each layer will exhibit phone-dependent regional activities when arranged in a 2-dimensional grid. We demonstrate that such constraints are able to yield visible activation regions without compromising the classification of the network and suppressing the activations for a region affects the classification accuracy of the corresponding phone more than the others. Shawn Tan, Khe Chai Sim, Mark J. F. Gales |
ASRU | 2 |
| 2015 | An investigation of augmenting speaker representations to improve speaker normalisation for DNN-based speech recognitionabstractThe conventional short-term interval features used by the Deep Neural Networks (DNNs) lack the ability to learn longer term information. This poses a challenge for training a speaker-independent (SI) DNN since the short-term features do not provide sufficient information for the DNN to estimate the real robust factors of speaker-level variations. The key to this problem is to obtain a sufficiently robust and informative speaker representation. This paper compares several speaker representations. Firstly, a DNN speaker classifier is used to extract the bottleneck features as the speaker representation, called the Bottleneck Speaker Vector (BSV). To further improve the robustness of this representation, a first-order Bottleneck Speaker Super Vector (BSSV) is also proposed, where the BSV is expanded into a super vector space by incorporating the phoneme posterior probabilities. Finally, a more fine-grain speaker representation based on the FMLLR-shifted features is examined. The experimental results on the WSJ0 and WSJ1 datasets show that the proposed speaker representations are useful in normalising the speaker effects for robust DNN-based automatic speech recognition. The best performance is achieved by augmenting both the BSSV and the FMLLR-shifted representations, yielding 10.0% - 15.3% relatively performance gains over the SI DNN baseline. Hengguan Huang, Khe Chai Sim |
ICASSP | 2 |
| 2014 | A Beam-Search Decoder for Disfluency Detection
Xuancong Wang, Hwee Tou Ng, Khe Chai Sim |
COLING | 3 |
| 2014 | Combining Punctuation and Disfluency Prediction: An Empirical StudyabstractPunctuation prediction and disfluency prediction can improve downstream natural language processing tasks such as machine translation and information extraction.Combining the two tasks can potentially improve the efficiency of the overall pipeline system and reduce error propagation.In this work 1 , we compare various methods for combining punctuation prediction (PU) and disfluency prediction (DF) on the Switchboard corpus.We compare an isolated prediction approach with a cascade approach, a rescoring approach, and three joint model approaches.For the cascade approach, we show that the soft cascade method is better than the hard cascade method.We also use the cascade models to generate an n-best list, use the bi-directional cascade models to perform rescoring, and compare that with the results of the cascade models.For the joint model approach, we compare mixedlabel Linear-chain Conditional Random Field (LCRF), cross-product LCRF and 2layer Factorial Conditional Random Field (FCRF) with soft-cascade LCRF.Our results show that the various methods linking the two tasks are not significantly different from one another, although they perform better than the isolated prediction method by 0.5-1.5% in the F1 score.Moreover, the clique order of features also shows a marked difference. Xuancong Wang, Khe Chai Sim, Hwee Tou Ng |
EMNLP | 2 |
| 2014 | Second order vector taylor series based robust speech recognitionabstractVector Taylor Series (VTS) model based compensation approach has been successfully applied to various robust speech recognition tasks. In this paper, a novel method to derive the formula to calculate the static and dynamic statistics based on second-order VTS (sVTS) is presented, which provides a new insight on the VTS approximation. Lengthy derivation could therefore be avoided when high order VTS is used and the proposed approach is more compact and easier to implement compared to previous high order VTS approaches. Experiments on Aurora 4 showed that the proposed sVTS based model compensation approach obtained 16.7% relative WER reduction over traditional first-order VTS (fVTS) approach. Suliang Bu, Yanmin Qian, Khe Chai Sim, Yongbin You, Kai Yu 0004 |
ICASSP | 3 |
| 2014 | An ideal hidden-activation mask for deep neural networks based noise-robust speech recognitionabstractDeep neural networks (DNNs) are capable of modeling large acoustic variations. However, the performance on noisy data is still below humans' expectations. In this work, we present an ideal hidden-activation masking (IHM) approach to improve their noise robustness. This IHM is inspired by the existing spectral masking techniques. Instead of masking away the noise-dominant components in the spectral domain, we propose to discard DNNs' inconsistent hidden activations. The IHM is computed from the parallel data to identify hidden units that are immune to environment noise. DNNs then utilize it to improve their prediction robustness with the noise-invariant activations. Experimental results on the Aurora4 task have shown that the proposed IHM is both effective in reducing noise variations and robust to mask estimation errors. Bo Li 0028, Khe Chai Sim |
ICASSP | 2 |
| 2014 | On combining DNN and GMM with unsupervised speaker adaptation for robust automatic speech recognitionabstractRecently, context-dependent Deep Neural Network (CD-DNN) has been found to significantly outperform Gaussian Mixture Model (GMM) for various large vocabulary continuous speech recognition tasks. Unlike the GMM approach, there is no meaningful interpretation of the DNN parameters, which makes it difficult to devise effective adaptation methods for DNNs. Furthermore, DNN parameter estimation is based on discriminative criteria, which is more sensitive to label errors and therefore less reliable for unsupervised adaptation. Many effective adaptation techniques that have been developed and proven to work well for GMM/HMM systems cannot be easily applied to DNNs. Therefore, this paper proposes a novel method of combining DNN and GMM using the Temporally Varying Weight Regression framework to take advantage of the superior performance of the DNNs and the robust adaptability of the GMMs. This paper addresses the issue of incorporating the high-dimensional CD-DNN posteriors into this framework without dramatically increasing the system complexity. Experimental results on a broadcast news large vocabulary transcription task show that the proposed GMM+DNN/HMM system achieved significant performance gain over the baseline DNN/HMM system. With additional unsupervised speaker adaptation, the best GMM+DNN/HMM system obtained about 20% relative improvements over the DNN/HMM baseline. Khe Chai Sim |
ICASSP | 2 |
| 2014 | Refinements of regression-based context-dependent modelling of deep neural networks for automatic speech recognitionabstractThe data sparsity problem of context-dependent (CD) acoustic modelling of deep neural networks (DNNs) in speech recognition is addressed by using the decision tree state clusters as the training targets. The CD states within a cluster cannot be distinguished during decoding. This problem, referred to as the clustering problem, is not explicitly addressed in the current literature. In our previous work, a regression-based CD-DNN framework was proposed to address both the data sparsity and the clustering problems. This paper investigates several refinements for the regression-based CD-DNN including two more representative state approximation schemes and the incorporation of sequential learning. The two approximations are obtained based on the statistics learned from the training data. Sequential learning is applied to both broad phone DNN detectors and the regression NN. The proposed refinements are evaluated on a broadcast news transcription task. For the cross-entropy systems, the two approximations perform consistently better than our previous work. Consistent performance gain over the corresponding cross-entropy trained systems is also observed for both the baseline CD-DNN and the regression model with sequential learning. Guangsen Wang, Khe Chai Sim |
ICASSP | 2 |
| 2014 | Modeling long temporal contexts for robust DNN-based speech recognition
Bo Li 0028, Khe Chai Sim |
INTERSPEECH | 2 |
| 2014 | Joint adaptation and adaptive training of TVWR for robust automatic speech recognition
Khe Chai Sim |
INTERSPEECH | 2 |
| 2014 | A multimodal stroke-based predictive input for efficient Chinese text entry on mobile devicesabstractHandwriting input method is particularly useful for languages with a logographic writing system. This paper introduces a multimodal stroked-based predictive input for the Chinese language. The proposed method requires users to write only the first few strokes of each character and the system will intelligently infer the intended characters by making use of contextual information. Specifically, a statistical n-gram language model is used. Motivated by the work on Haptic Voice Recognition, this paper also incorporates voice input as an additional modality to further enhance the prediction accuracy. Empirical simulation results show that the predictive handwriting input method with 3 initial strokes outperforms the predictive Pinyin input method. Further improvements can be obtained by considering both the initial stroke order and the corresponding stroke layout. Finally, with voice overlay, the proposed multimodal stroke-based predictive input method achieved more than 85% and 95% 1-best prediction accuracies with 2 and 3 initial strokes respectively. Khe Chai Sim |
SLT | 1 |
| 2014 | A Spectral Masking Approach to Noise-Robust Speech Recognition Using Deep Neural NetworksabstractImproving the noise robustness of automatic speech recognition systems has been a challenging task for many years. Recently, it was found that Deep Neural Networks (DNNs) yield large performance gains over conventional GMM-HMM systems, when used in both hybrid and tandem systems. However, they are still far from the level of human expectations especially under adverse environments. Motivated by the separation-prior-to-recognition process of the human auditory system, we propose a robust spectral masking system where power spectral domain masks are predicted using a DNN trained on the same filter-bank features used for acoustic modeling. To further improve performance, Linear Input Network (LIN) adaptation is applied to both the mask estimator and the acoustic model DNNs. Since the estimation of LINs for the mask estimator requires stereo data, which is not available during testing, we proposed using the LINs estimated for the acoustic model DNNs to adapt the mask estimators. Furthermore, we used the same set of weights obtained from pre-training for the input layers of both the mask estimator and the acoustic model DNNs to ensure a better consistency for sharing LINs. Experimental results on benchmark Aurora2 and Aurora4 tasks demonstrated the effectiveness of our system, which yielded Word Error Rates (WERs) of 4.6% and 11.8% respectively. Furthermore, the simple averaging of posteriors from systems with and without spectral masking can further reduce the WERs to 4.3% on Aurora2 and 11.4% on Aurora4. Bo Li 0028, Khe Chai Sim |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Temporally Varying Weight Regression: A Semi-Parametric Trajectory Model for Automatic Speech RecognitionabstractStandard Hidden Markov Model (HMM) assumes that successive observations are independent to one another given the state sequence. This leads to a poor trajectory model for speech. Many explicit trajectory modeling techniques have been studied in the past to improve trajectory modeling for HMM. However, these techniques do not yield promising improvements over conventional HMM systems where differential parameters and Gaussian Mixture Model have been used implicitly to circumvent the poor trajectory modeling issue of HMM. Recently, semi-parametric trajectory modeling techniques based on temporally varying model parameters such as fMPE and pMPE have been shown to yield promising improvements over state-of-the-art systems on large vocabulary continuous speech recognition tasks. These techniques use high dimensional posterior features derived from a long span of acoustic features to model temporally varying attributes of the speech signal. Bases corresponding to these posterior features are then discriminatively estimated to yield temporally varying mean (fMPE) and precision matrix (pMPE) parameters. Motivated by the success of fMPE and pMPE, Temporally Varying Weight Regression (TVWR) was recently proposed to model HMM trajectory implicitly using time-varying Gaussian weights. In this paper, a complete formulation of TVWR is given based on a probabilistic modeling framework. Parameter estimation formulae using both maximum likelihood (ML) and minimum phone error (MPE) criteria are derived. Experimental results based on the Wall Street Journal ( CSR-WSJ0 + WSJ1) and Aurora 4 corpora show that consistent promising improvements over the standard HMM systems can be obtained in both the 20 k open vocabulary recognition task (NIST Nov'92 WSJ0) and 5 k closed vocabulary noisy speech recognition for both ML and MPE criteria. Khe Chai Sim |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Regression-Based Context-Dependent Modeling of Deep Neural Networks for Speech RecognitionabstractThe data sparsity problem is addressed by using the decision tree state clusters as the training targets for the state-of-the- art context-dependent (CD) deep neural network (DNN) systems. The CD states within a cluster cannot be distinguished at the frame level. We surmise that the state clustering may cause an issue for the standard CD-DNNs, which has so far not been addressed in the literature. In this paper, a logistic regression framework is proposed for the CD-DNNs based on a set of broad phone classes to address both the data sparsity and the clustering problems. To address the data sparsity issue, the triphones are clustered into shorter biphones with broad phone contexts under multiple articulatory categories. A DNN is trained to discriminate the disjoint biphone clusters within each articulatory category. The regression bases are formed by the concatenated log posterior probabilities of all the broad phone DNNs. Logistic regression is used to transform the regression bases into the triphone state posteriors. Clustering of the regression parameters is used to reduce the regression model complexity while still achieving unique acoustic scores for all possible triphones. Based on some approximations, the regression model can be trained as a sparse softmax layer and its parameters can be learned by optimizing the cross-entropy criterion. The experimental results on a broadcast news transcription task reveal that the proposed regression-based CD-DNN significantly outperforms the standard CD-DNN. The best system provides a 1.3% absolute word error rate reduction compared to the best standard CD-DNN system. Guangsen Wang, Khe Chai Sim |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | Improving robustness of deep neural networks via spectral masking for automatic speech recognitionabstractThe performance of human listeners degrades rather slowly compared to machines in noisy environments. This has been attributed to the ability of performing auditory scene analysis which separates the speech prior to recognition. In this work, we investigate two mask estimation approaches, namely the state dependent and the deep neural network (DNN) based estimations, to separate speech from noises for improving DNN acoustic models' noise robustness. The second approach has been experimentally shown to outperform the first one. Due to the stereo data based training and ill-defined masks for speech with channel distortions, both methods do not generalize well to unseen conditions and fail to beat the performance of the multi-style trained baseline system. However, the model trained on masked features demonstrates strong complementariness to the baseline model. The simple average of the two system's posteriors yields word error rates of 4.4% on Aurora2 and 12.3% on Aurora4. Bo Li 0028, Khe Chai Sim |
ASRU | 2 |
| 2013 | Multi-stream temporally varying weight regression for cross-lingual speech recognitionabstractBuilding a good Automatic Speech Recognition (ASR) system with limited resources is a very challenging task due to the existing many speech variations. Multilingual and cross-lingual speech recognition techniques are commonly used for this task. This paper investigates the recently proposed Temporally Varying Weight Regression (TVWR) method for cross-lingual speech recognition. TVWR uses posterior features to implicitly model the long-term temporal structures in acoustic patterns. By leveraging on the well-trained foreign recognizers, high quality monophone/state posteriors can be easily incorporated into TVWR to boost the ASR performance on low-resource languages. Furthermore, multi-stream TVWR is proposed, where multiple sets of posterior features are used to incorporate richer (temporal and spatial) context information. Finally, a separate state-tying for the TVWR regression parameters is used to better utilize the more reliable posterior features. Experimental results are evaluated for English and Malay speech recognition with limited resources. By using the Czech, Hungarian and Russian posterior features, TVWR was found to consistently outperform the tandem systems trained on the same features. Khe Chai Sim |
ASRU | 2 |
| 2013 | Context-dependent modelling of deep neural network using logistic regressionabstractThe data sparsity problem of context-dependent acoustic modelling in automatic speech recognition is addressed by using the decision tree state clusters as the training targets in the standard context-dependent (CD) deep neural network (DNN) systems. As a result, the CD states within a cluster cannot be distinguished during decoding. This problem, referred to as the clustering problem, is not explicitly addressed in the current literature. In this paper, we formulate the CD DNN as an instance of the canonical state modelling technique based on a set of broad phone classes to address both the data sparsity and the clustering problems. The triphone is clustered into multiple sets of shorter biphones using broad phone contexts to address the data sparsity issue. A DNN is trained to discriminate the biphones within each set. The canonical states are represented by the concatenated log posteriors of all the broad phone DNNs. Logistic regression is used to transform the canonical states into the triphone state output probability. Clustering of the regression parameters is used to reduce model complexity while still achieving unique acoustic scores for all possible triphones. The experimental results on a broadcast news transcription task reveal that the proposed regression-based CD DNN significantly outperforms the standard CD DNN. The best system provides a 2.7% absolute WER reduction compared to the best standard CD DNN system. Guangsen Wang, Khe Chai Sim |
ASRU | 2 |
| 2013 | Noise adaptive front-end normalization based on Vector Taylor Series for Deep Neural Networks in robust speech recognitionabstractDeep Neural Networks (DNNs) have been successfully applied to various speech tasks during recent years. In this paper, we investigate the use of DNNs for noise-robust speech recognition and demonstrate their superior capabilities of modeling acoustic variations over the conventional Gaussian Mixture Models (GMMs). We then propose to compensate the normalization front-end of the DNNs using the GMM-based Vector Taylor Series (VTS) model compensation technique, which has been successfully applied in the GMM-based ASR systems to handle noisy speech. To fully benefit from both the powerful modeling capability of the DNN and the effective noise compensation of the VTS, an adaptive training algorithm is further developed. The preliminary experimental results on the AURORA 2 task have demonstrated the effectiveness of our approach. The adaptively trained system has been shown to outperform the GMM-based VTS adaptive training by relatively 18.8% using the MFCC features and 21.9% using the FBank features. Bo Li 0028, Khe Chai Sim |
ICASSP | 2 |
| 2013 | Approximated Parallel Model Combination for efficient noise-robust speech recognitionabstractParallel Model Combination (PMC) and Vector Taylor Series (VTS) are two model-based approaches for noise-robust speech recognition. The latter is more popular because of its simple compensation formulae for both the static and dynamic parameters. Furthermore, this VTS compensation formulation can be easily extended to noise adaptive training where the parameters of the underlying pseudo-clean speech and distortion models can be optimized. PMC lacks the above benefits because of its nonlinear variance compensation formula. In this paper, the Approximated PMC (APMC) method is proposed where linearized PMC variance compensation is used. The same approximation has also been applied to Trajectory-based APMC (TAPMC) to achieve a four-time computational saving over the Trajectory-based PMC (TPMC). The dynamic parameter compensation and noise re-estimation formulae for APMC are also derived. Experimental results on AURORA 4 show that APMC and TAPMC consistently outperformed the standard VTS and Trajectory-based VTS (TVTS) by 6.3% and 5.3% relative respectively. Khe Chai Sim |
ICASSP | 1 |
| 2013 | An investigation of spectral restoration algorithms for deep neural networks based noise robust speech recognitionabstractDeep Neural Networks (DNNs) are becoming widely accepted in automatic speech recognition (ASR) systems. The deep structured nonlinear processing greatly improves the model’s generalization capability, but the performance under adverse environments is still unsatisfactory. In the literature, there have been many techniques successfully developed to improve Gaussian mixture models’ robustness. Investigating the effectiveness of these techniques for the DNN is an important step to thoroughly understand its superiority, pinpoint its limitations and most importantly to further improve it towards the ultimate human-level robustness. In this paper, we investigate the effectiveness of speech enhancement using spectral restoration algorithms for DNNs. Four approaches are evaluated, namely minimum mean-square error spectral estimator (MMSE), maximum likelihood spectral amplitude estimator (MLSA), maximum a posteriori spectral amplitude estimator (MAPA), and generalized maximum a posteriori spectral amplitude algorithm (GMAPA). The preliminary experimental results on the Aurora 2 speech database show that with multi-condition training data the DNN itself is capable of learning robust representations. However, if only clean data is available, the MLSA algorithm is the best spectral restoration training method for DNNs. Bo Li 0028, Yu Tsao 0001, Khe Chai Sim |
INTERSPEECH | 3 |
| 2013 | Parameter clustering for temporally varying weight regression for automatic speech recognition
Khe Chai Sim |
INTERSPEECH | 2 |
| 2013 | An investigation of temporally varying weight regression for noise robust speech recognition
Khe Chai Sim |
INTERSPEECH | 2 |
| 2013 | Integrating conditional random fields and joint multi-gram model with syllabic features for grapheme-to-phone conversion
Khe Chai Sim |
INTERSPEECH | 2 |
| 2012 | Probabilistic Integration of Partial Lexical Information for Noise Robust Haptic Voice Recognition
Khe Chai Sim |
ACL (1) | 1 |
| 2012 | Implicit trajectory modelling using temporally varying weight regression for automatic speech recognitionabstractRecently, implicit trajectory modelling using temporally varying model parameters has achieved promising gains over the discriminatively trained standard HMM system. However, these works only focus on the temporally varying means or precisions explicitly. It is interesting to explore the capability of temporally varying weights, since the effect of time varying Gaussian parameters can be achieved by adjusting the weights of Gaussian Mixture Models (GMM) for different observation. This paper proposes a Temporally Varying Weight Regression (TVWR) model to learn the importance of different Gaussian components under different temporal contexts. Technically, TVWR factorizes the HMM state likelihood such that the contextual information can be modelled using time varying weights. Additionally, approximate constraints are derived to ensure a valid probabilistic model for TVWR. Experimental results for continuous speech recognition on Wall Street Journal show consistent improvements with varying system complexity and about 12% relative significant improvements in the best case. Khe Chai Sim |
ICASSP | 2 |
| 2012 | An investigation of tied-mixture GMM based triphone state clusteringabstractParameter tying is a crucial scheme for robust context dependent acoustic modeling since it takes a major role in balancing the desired model complexity and the amount of data available. In this paper, a modified decision tree state clustering scheme based on tied-mixture Gaussian Mixture Model (GMM) is proposed. Instead of using a single Gaussian untied triphone system, a tied-mixture GMM triphone system is adopted as a better acoustic model for state clustering. Meanwhile, the proposed scheme allows easy incorporation of discriminative training during clustering. Experimental results show that for a varying number of state clusters, the proposed approach consistently outperforms the standard single Gaussian based state tying. The best WER performance has a 10.5% relative improvement over the conventional decision tree clustering and the proposed scheme achieves its best performance using a much smaller number of state clusters. Moreover, detailed analyses reveal that the proposed GMM clustering has a better state distribution which leads to 1) better frame-state alignments 2) better phonetic question selections. These two factors may make the proposed approach superior for clustering. Guangsen Wang, Khe Chai Sim |
ICASSP | 2 |
| 2012 | Design and implementation of the note-taking style haptic voice recognition for mobile devicesabstractThis research proposes the "note-taking style" Haptic Voice Recognition (HVR) technology which incorporates speech and touch sensory inputs in a note-like form to enhance the performance of speech recognition. A note is taken from a user via two different haptic input methods - handwriting and a keyboard. A note consists of some of the keywords in the given utterance, either partially spelled or fully spelled. In order to facilitate fast input, the interface allows a shorthand writing system such as Gregg Shorthand. Using this haptic note sequence as an additional knowledge source, the algorithm re-ranks the n-best list generated by a speech engine. The simulation and experimental results show that the proposed HVR method improves the Word Error Rate (WER) and Keyword Error Rate (KER) performance in comparison to an Automatic Speech Recognition (ASR) system. Although it generates an inevitable increase in speech duration due to disfluency and occasional mistakes in haptic input, the compensation is shown to be less than conventional HVR methods. As such, this new note-taking style HVR interaction has the potential to be both natural and effective in increasing the recognition performance by choosing the most likely utterance among multiple hypotheses. This paper discusses the algorithm for the proposed system, the results from the simulation and the experiments, and the possible applications of this new technology such as aiding spoken document retrieval with haptic notes. Seungwhan Moon, Khe Chai Sim |
ICMI | 2 |
| 2012 | Speak-as-you-swipe (SAYS): a multimodal interface combining speech and gesture keyboard synchronously for continuous mobile text entryabstractModern mobile devices, such as the smartphones and tablets, are becoming increasingly popular amongst users of all ages. Text entry is one of the most important modes of interaction between human and their mobile devices. Although typing on a touchscreen display using a soft keyboard remains the most common text input method for many users, the process can be frustratingly slow, especially on smartphones with a much smaller screen. Voice input offers an attractive alternative that completely eliminates the need for typing. However, voice input relies on automatic speech recognition technology whose performance degrades significantly in noisy environment or for non-native users. This paper presents Speak-As-You-Swipe (SAYS), a novel multimodal interface that enables efficient continuous text entry on mobile devices. SAYS integrates a gesture keyboard with speech recognition to improve the efficiency and accuracy of text entry. The swipe gesture and voice inputs provide complementary information that can be very effective in disambiguating confusions in word predictions. The word prediction hypotheses from a gesture keyboard are directly incorporated into the speech recognition process so that the SAYS interface can handle continuous input. Experimental results show that for a 20k vocabulary, the proposed SAYS interface can achieve prediction accuracy of 96.4% in clean condition and about 94.0% in noisy environment, compared to 92.2% using a gesture keyboard alone. Khe Chai Sim |
ICMI | 1 |
| 2012 | ICMI'12 grand challenge: haptic voice recognitionabstractThis paper describes the Haptic Voice Recognition (HVR) Grand Challenge 2012 and its datasets. The HVR Grand Challenge 2012 is a research oriented competition designed to bring together researchers across multiple disciplines to work on novel multimodal text entry methods involving speech and touch inputs. Annotated datasets were collected and released for this grand challenge as well as future research purposes. A simple recipe for building an HVR system using the Hidden Markov Model Toolkit (HTK) was also provided. In this paper, detailed analyses of the datasets will be given. Experimental results obtained using these data will also be presented. Khe Chai Sim, Shengdong Zhao 0001, Kai Yu 0004, Hank Liao |
ICMI | 1 |
| 2012 | Improving mandarin predictive text input by augmenting pinyin initials with speech and tonal informationabstractRecently, a new technology called Haptic Voice Recognition (HVR) was proposed to enhance the speech recognition efficiency and accuracy for modern mobile devices, which has been successfully applied for robust English voice recognition. As both Pinyin and handwriting input methods work quite slow in mobile devices because of typing errors and ambiguity, it is interesting to apply this technology to assist Mandarin predictive text input. However, it is not straightforward because the characteristics of Mandarin significantly differ from alphabetic western languages. In this paper, we investigated what possible haptic inputs are important and how these information can be incorporated to improve Mandarin text input. Various experiments were conducted and results have shown that with the help of acoustic and tonal information, the ambiguity of Pinyin Initial based Mandarin predictive text input is largely reduced and an oracle character error rate of 3.8% for the top 4 candidates could be achieved, which is usually the number of word candidates displayed on mobile devices. Our Mandarin HVR system has also shown its robustness in noisy environments. Guangsen Wang, Bo Li 0028, Xuancong Wang, Khe Chai Sim |
ICMI | 6 |
| 2012 | A Weighted Combination of Speech with Text-based Models for Arabic Diacritization
Aisha S. Azim, Khe Chai Sim |
INTERSPEECH | 3 |
| 2012 | A Two-stage Speaker Adaptation Approach for Subspace Gaussian Mixture Model based Nonnative Speech Recognitionabstract13th Annual Conference of the International Speech Communication Association 2012, INTERSPEECH 2012 Bo Li 0028, Khe Chai Sim |
INTERSPEECH | 2 |
| 2012 | Dynamic Conditional Random Fields for Joint Sentence Boundary and Punctuation Predictionabstract13th Annual Conference of the International Speech Communication Association 2012, INTERSPEECH 2012 Xuancong Wang, Hwee Tou Ng, Khe Chai Sim |
INTERSPEECH | 3 |
| 2012 | MOGAT: mobile games with auditory training for children with cochlear implantsabstractCochlear implants have improved the lives of tens of thousands of the hearing impaired by providing sufficient auditory perception for speech, but these devices are far from satisfactory for music perception. Many cochlear implant recipients, especially pre-lingually deafened children, have difficulty recognizing and producing specific pitches. To improve musical auditory habilitation for children post cochlear implantation, we developed MOGAT: MObile Games with Auditory Training. The system includes three musical games built with off-the-shelf mobile devices to train their pitch perception and intonation skills respectively, and a cloud-based web service which allows music therapists to monitor and design individual training for children. The design of the games and web service was informed by a pilot survey (N=60 children). To ensure widespread use with low-cost mobile devices, we minimized the computation load while retaining highly accurate audio analysis. A 6-week user study (N=15 children) showed that the music habilitation with MOGAT was intuitive, enjoyable and motivating. It has improved most children's pitch discrimination and production, and several children's improvement was statistically significant (p<0.05). Yinsheng Zhou, Khe Chai Sim, Patsy Tan, Ye Wang 0007 |
ACM Multimedia | 2 |
| 2011 | A Trajectory-based Parallel Model Combination with a unified static and dynamic parameter compensation for noisy speech recognitionabstractParallel Model Combination (PMC) is widely used as a technique to compensate Gaussian parameters of a clean speech model for noisy speech recognition. The basic principle of PMC uses a log normal approximation to transform statistics of the data distribution between the cepstral domain and the linear spectral domain. Typically, further approximations are needed to compensate the dynamic parameters separately. In this paper, Trajectory PMC (TPMC) is proposed to compensate both the static and dynamic parameters. TPMC uses the explicit relationships between the static and dynamic features to transform the static and dynamic parameters into a sequence (trajectory) of static parameters, so that the log normal approximation can be applied. Experimental results on WSJCAM0 database corrupted with additive babble noise reveals that the proposed TPMC method gives promising improvements over PMC and VTS. Khe Chai Sim, Minh-Thang Luong |
ASRU | 1 |
| 2011 | Sequential Classification Criteria for NNs in Automatic Speech RecognitionabstractProceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH Guangsen Wang, Khe Chai Sim |
INTERSPEECH | 2 |
| 2011 | Comparison of Smoothing Techniques for Robust Context Dependent Acoustic Modelling in Hybrid NN/HMM SystemsabstractProceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH Guangsen Wang, Khe Chai Sim |
INTERSPEECH | 2 |
| 2011 | Using Discrete Probabilities With Bhattacharyya Measure for SVM-Based Speaker VerificationabstractSupport vector machines (SVMs), and kernel classifiers in general, rely on the kernel functions to measure the pairwise similarity between inputs. This paper advocates the use of discrete representation of speech signals in terms of the probabilities of discrete events as feature for speaker verification and proposes the use of Bhattacharyya coefficient as the similarity measure for this type of inputs to SVM. We analyze the effectiveness of the Bhattacharyya measure from the perspective of feature normalization and distribution warping in the SVM feature space. Experiments conducted on the NIST 2006 speaker verification task indicate that the Bhattacharyya measure outperforms the Fisher kernel, term frequency log-likelihood ratio (TFLLR) scaling, and rank normalization reported earlier in literature. Moreover, the Bhattacharyya measure is computed using a data-independent square-root operation instead of data-driven normalization, which simplifies the implementation. The effectiveness of the Bhattacharyya measure becomes more apparent when channel compensation is applied at the model and score levels. The performance of the proposed method is close to that of the popular GMM supervector with a small margin. Kong-Aik Lee, Chang Huai You, Haizhou Li 0001, Tomi Kinnunen, Khe Chai Sim |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2010 | A minimum variance asynchronous Detection Error Trade-off performance analysis for multi-class detection problemsabstractIn a detection problem, the trade-off between the miss and false alarm probabilities are often shown as a Detection Error Trade-off (DET) curve. The DET curve is obtained by adjusting a decision threshold to vary the compromise between these probabilities. For a multi-class detection problem, each class has its own decision threshold, which leads to a multivariate detection error trade-off. In order to plot a DET curve, the decision thresholds are constrained to a single degree of freedom. Typically, they are synchronously constrained to use the same values. In this paper, a minimum variance DET analysis is proposed where the decision thresholds are asynchronously constrained such that the variances of the miss and false alarm probabilities are minimised. If the scores are normally distributed, the decision thresholds can be approximated as a linear univariate function and the resulting DET curve is also a straight line. In more general cases where the scores are not normally distributed, piecewise linear functions can be estimated iteratively, instead. The minimum variance asynchronous DET analysis is applied to a phone verification task on TIMIT database. Khe Chai Sim |
ICASSP | 1 |
| 2010 | Adaptive score fusion using Weighted Logistic Linear Regression for spoken language recognitionabstractState-of-the-art spoken language recognition systems typically consist of a combination of sub-systems. These sub-systems generate language detection scores for each speech segment, which will be fused (combined) to yield the overall detection scores. Typically, score fusion is achieved using a linear model and Logistic Linear Regression (LLR) is commonly used to estimate the model parameters. This paper proposes an extension to the LLR model, known as the Weighted LLR (WLLR). WLLR is obtained using a weighted combination of multiple LLRs where the weights are obtained as a nonlinear function of the speech segments. Although the resultant score is still linear with respect to the scores of the individual sub-systems, the linear function depends on the speech segment. Hence, the overall score fusion model can be regarded as an adaptive model. Experimental results shows that WLLR outperforms LLR by approximately 10% relative for PPRLM system fusion on the NIST 2003 and 2005 language recognition evaluation sets. Khe Chai Sim, Kong-Aik Lee |
ICASSP | 1 |
| 2010 | Comparison of discriminative input and output transformations for speaker adaptation in the hybrid NN/HMM systemsabstractProceedings of the 11th Annual Conference of the International Speech Communication Association, INTERSPEECH 2010 Bo Li 0028, Khe Chai Sim |
INTERSPEECH | 2 |
| 2010 | Hidden logistic linear regression for support vector machine based phone verificationabstractProceedings of the 11th Annual Conference of the International Speech Communication Association, INTERSPEECH 2010 Bo Li 0028, Khe Chai Sim |
INTERSPEECH | 2 |
| 2010 | Probabilistic state clustering using conditional random field for context-dependent acoustic modellingabstractProceedings of the 11th Annual Conference of the International Speech Communication Association, INTERSPEECH 2010 Khe Chai Sim |
INTERSPEECH | 1 |
| 2010 | Semi-parametric trajectory modelling using temporally varying feature mapping for speech recognitionabstractProceedings of the 11th Annual Conference of the International Speech Communication Association, INTERSPEECH 2010 Khe Chai Sim |
INTERSPEECH | 1 |
| 2010 | Haptic Voice Recognition: Augmenting speech modality with touch events for efficient speech recognitionabstractThis paper proposes the Haptic Voice Recognition (HVR), a multi-modal interface that combines speech and touch sensory inputs to perform voice recognition. These touch inputs form a series of haptic events that provide cues or `landmarks' for word boundaries. These word boundary cues greatly reduce the search space for speech recognition, thereby making the decoding process more efficient and suitable for portable devices with limited compute and memory resources. Furthermore, having the knowledge of word boundaries also suppresses insertion and deletion errors. This is particularly helpful when recognition is performed in noisy environment. In this paper, a series of experiments were conducted to study the feasibility of augmenting touch events to automatic speech recognition and explore its potential benefits. Experiments were conducted with syntactically simulated haptic events on the Wall Street Journal database as well as realistic haptic events acquired using a prototype HVR interface implemented on a touchscreen device. Khe Chai Sim |
SLT | 1 |
| 2010 | Statistical lattice-based spoken document retrievalabstractRecent research efforts on spoken document retrieval have tried to overcome the low quality of 1-best automatic speech recognition transcripts, especially in the case of conversational speech, by using statistics derived from speech lattices containing multiple transcription hypotheses as output by a speech recognizer. We present a method for lattice-based spoken document retrieval based on a statistical n -gram modeling approach to information retrieval. In this statistical lattice-based retrieval (SLBR) method, a smoothed statistical model is estimated for each document from the expected counts of words given the information in a lattice, and the relevance of each document to a query is measured as a probability under such a model. We investigate the efficacy of our method under various parameter settings of the speech recognition and lattice processing engines, using the Fisher English Corpus of conversational telephone speech. Experimental results show that our method consistently achieves better retrieval performance than using only the 1-best transcripts in statistical retrieval, outperforms a recently proposed lattice-based vector space retrieval method, and also compares favorably with a lattice-based retrieval method based on the Okapi BM25 model. Tee Kiah Chia, Khe Chai Sim, Haizhou Li 0001, Hwee Tou Ng |
ACM Trans. Inf. Syst. | 2 |
| 2010 | Word level automatic alignment of music and lyrics using vocal synthesisabstractWe propose a signal-based approach instead of the commonly used model-based approach, to automatically align vocal music with text lyrics at the word level. In this approach, we use a text-to-speech system to synthesize the singing voice according to the lyrics. In this way, aligning the music signal with the corresponding text lyrics becomes the alignment of two audio signals. This study uses the results of music information modeling and singing voice synthesis. In music information modeling, we study different music representation strategies for music segmentation, music region indexing and region content descriptions; in singing voice synthesis, we generate singing voice by making use of music knowledge to approximate the target vocal line in terms of tempo. The experimental results on a 20-song database show 26.3% and 36.1% word level alignment error rates at eighth note and sixteenth note alignment tolerances respectively. The proposed approach presents an alternative and effective solution to music-lyrics alignment which may require less training dataset. Namunu Chinthaka Maddage, Khe Chai Sim, Haizhou Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2009 | Discriminative Product-of-Expert acoustic mapping for cross-lingual phone recognitionabstractThis paper presents a product-of-expert framework to perform probabilistic acoustic mapping for cross-lingual phone recognition. Under this framework, the posterior probabilities of the target HMM states are modelled as the weighted product of experts, where the experts or their weights are modelled as functions of the posterior probabilities of the source HMM states generated by a foreign phone recogniser. Careful choice of these functions leads to the product-of-posterior and posterior weighted product-of-expert models, which can be conveniently represented as 2-layer and 3-layer feed-forward neural networks respectively. Therefore, the commonly used error back-propagation method can be used to discriminatively train the model parameters. Experimental results are presented on the NTIMIT database using the Czech, Hungarian and Russian hybrid NN/HMM recognisers as the foreign phone recognisers to recognise English phones. With only about 15.6 minutes of training data, the best acoustic mapping model achieved 46.00% phone error rate, which is not far behind the 43.55% performance of the NN/HMM system trained directly on the full 3.31 hours of data. Khe Chai Sim |
ASRU | 1 |
| 2009 | The I4U system in NIST 2008 speaker recognition evaluationabstractThis paper describes the performance of the I4U speaker recognition system in the NIST 2008 Speaker Recognition Evaluation. The system consists of seven subsystems, each with different cepstral features and classifiers. We describe the I4U Primary system and report on its core test results as they were submitted, which were among the best-performing submissions. The I4U effort was led by the Institute for Infocomm Research, Singapore (IIR), with contributions from the University of Science and Technology of China (USTC), the University of New South Wales, Australia (UNSW), Nanyang Technological University, Singapore (NTU) and Carnegie Mellon University, USA (CMU). Haizhou Li 0001, Bin Ma 0001, Kong-Aik Lee, Hanwu Sun, Donglai Zhu, Khe Chai Sim, Chang Huai You, Rong Tong, Ismo Kärkkäinen, Chien-Lin Huang, Vladimir Pervouchine, Wu Guo, Yijie Li 0001, Li-Rong Dai 0001, Mohaddeseh Nosratighods, Tharmarajah Thiruvaran, Julien Epps, Eliathamby Ambikairajah, Chng Eng Siong, Tanja Schultz, Qin Jin |
ICASSP | 6 |
| 2009 | Stream-based context-sensitive phone mapping for cross-lingual speech recognition
Khe Chai Sim, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2008 | Robust phone set mapping using decision tree clustering for cross-lingual phone recognitionabstractRecently, research related to multi-lingual and cross-lingual speech has gained increasing popularity. One of the major problems when dealing with multi-lingual speech data is the mapping of the phone sets between different languages. Phone mapping is useful for cross-lingual speech recognition, cross-lingual pronunciation modelling and mixed language speech synthesis, to name a few. In this paper, an automatic context sensitive phone set mapping method is presented to improve the mapping accuracy. A training methodology that allows the mapping to be learned automatically from parallel time-aligned phone transcriptions is also described. In particular, a decision tree clustering technique is used to tie unseen contexts for robustness. The quality of the proposed mapping method is evaluated on a cross-lingual phone recognition task where the Hungarian and Russian phone recognisers are used to recognise Czech speech and produce Czech phone sequences through phone set mapping. The mapping was trained on only a small amount of data. A consistent relative improvement of 5 – 7% is reported when contextual information is added to phone set mapping. Khe Chai Sim, Haizhou Li 0001 |
ICASSP | 1 |
| 2008 | Context-sensitive probabilistic phone mapping model for cross-lingual speech recognition
Khe Chai Sim, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2008 | NIST 2007 Language Recognition Evaluation: From the Perspective of IIR
Haizhou Li 0001, Bin Ma 0001, Kong-Aik Lee, Khe Chai Sim, Hanwu Sun, Rong Tong, Donglai Zhu, Chang Huai You |
PACLIC | 4 |
| 2008 | A lattice-based approach to query-by-example spoken document retrievalabstractRecent efforts on the task of spoken document retrieval (SDR) have made use of speech lattices: speech lattices contain information about alternative speech transcription hypotheses other than the 1-best transcripts, and this information can improve retrieval accuracy by overcoming recognition errors present in the 1-best transcription. In this paper, we look at using lattices for the query-by-example spoken document retrieval task - retrieving documents from a speech corpus, where the queries are themselves in the form of complete spoken documents (query exemplars). We extend a previously proposed method for SDR with short queries to the query-by-example task. Specifically, we use a retrieval method based on statistical modeling: we compute expected word counts from document and query lattices, estimate statistical models from these counts, and compute relevance scores as divergences between these models. Experimental results on a speech corpus of conversational English show that the use of statistics from lattices for both documents and query exemplars results in better retrieval accuracy than using only 1-best transcripts for either documents, or queries, or both. In addition, we investigate the effect of stop word removal which further improves retrieval accuracy. To our knowledge, our work is the first to have used a lattice-based approach to query-by-example spoken document retrieval. Tee Kiah Chia, Khe Chai Sim, Haizhou Li 0001, Hwee Tou Ng |
SIGIR | 2 |
| 2008 | On Acoustic Diversification Front-End for Spoken Language IdentificationabstractThe parallel phone recognition followed by language model (PPRLM) architecture represents one of the state-of-the-art spoken language identification systems. A PPRLM system comprises multiple parallel subsystems, where each subsystem employs a phone recognizer with a different phone set for a particular language. The phone recognizer extracts phonotactic attributes from the speech input to characterize a language. The multiple parallel subsystems are devised to capture the phonetic diversification available in the speech input. Alternatively, this paper investigates a new approach for building a PPRLM system that aims at improving the acoustic diversification among its parallel subsystems by using multiple acoustic models. These acoustic models are trained on the same speech data with the same phone set but using different model structures and training paradigms. We examine the use of various structured precision (inverse covariance) matrix modeling techniques as well as the maximum likelihood and maximum mutual information training paradigms to produce complementary acoustic models. The results show that acoustic diversification, which requires only one set of phonetically transcribed speech data, yields similar performance improvements compared to phonetic diversification. In addition, further improvements were obtained by combining both diversification factors. The best performing system reported in this paper combined phonetic and acoustic diversifications to achieve EERs of 4.71% and 8.61% on the 2003 and 2005 NIST LRE sets, respectively, compared to 5.77% and 9.94% using phonetic diversification alone. Khe Chai Sim, Haizhou Li 0001 |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | Semantic Transliteration of Personal Names
Haizhou Li 0001, Khe Chai Sim, Jin-Shea Kuo, Minghui Dong |
ACL | 2 |
| 2007 | Consensus Network Decoding for Statistical Machine Translation System CombinationabstractThis paper presents a simple and robust consensus decoding approach for combining multiple machine translation (MT) system outputs. A consensus network is constructed from an N-best list by aligning the hypotheses against an alignment reference, where the alignment is based on minimising the translation edit rate (TER). The minimum Bayes risk (MBR) decoding technique is investigated for the selection of an appropriate alignment reference. Several alternative decoding strategies proposed to retain coherent phrases in the original translations. Experimental results are presented primarily based on three-way combination of Chinese-English translation outputs, and also presents results for six-way system combination. It is shown that worthwhile improvements in translation performance can be obtained using the methods discussed. Khe Chai Sim, William J. Byrne, Mark J. F. Gales, Hichem Sahbi, Philip C. Woodland |
ICASSP (4) | 1 |
| 2007 | Improving Speech Transcription for Mandarin-English TranslationabstractThis paper describes the development of the CU-HTK Mandarin speech-to-text (STT) system and assesses its performance as part of a transcription-translation pipeline which converts broadcast Mandarin audio into English text. Recent improvements to the STT system are described and these give character error rate (CER) gains of 14.3% absolute for a broadcast conversation (BC) task and 5.1% absolute for a broadcast news (BN) task. The output of these STT systems is then post-processed, so that it consists of sentence-like segments, and translated into English text using a statistical machine translation (SMT) system. The performance of the transcription-translation pipeline is evaluated using the translation edit rate (TER) and BLEU metrics. It is shown that improving both the STT system and the post-STT segmentations can lower the TER scores by up to 5.3% absolute and increase the BLEU scores by up to 2.7% absolute. Marcus Tomalin, Mark J. F. Gales, Xunying Liu, Khe Chai Sim, Rohit Sinha 0003, Philip C. Woodland, Kai Yu 0004 |
ICASSP (4) | 4 |
| 2007 | Fusion of contrastive acoustic models for parallel phonotactic spoken language identification
Khe Chai Sim, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2007 | Discriminative semi-parametric trajectory model for speech recognition
Khe Chai Sim, Mark J. F. Gales |
Comput. Speech Lang. | 1 |
| 2006 | The Cu-Htk Mandarin Broadcast News Transcription SystemabstractThis paper discusses the development of the CU-HTK Mandarin broadcast news (BN) transcription system. The Mandarin BN task includes a significant amount of English data. Hence techniques have been investigated to allow the same system to handle both Mandarin and English by augmenting the Mandarin training sets with English acoustic and language model training data. A range of acoustic models were built including models based on Gaussianised features, speaker adaptive training and feature-space MPE. A multi-branch system architecture is described in which multiple acoustic model types, alternate phone sets and segmentations can be used in a system combination framework to generate the final output. The final system shows state-of-the-art performance over a range of test sets Rohit Sinha 0003, Mark J. F. Gales, Do Yeong Kim, Xunying Liu, Khe Chai Sim, Philip C. Woodland |
ICASSP (1) | 5 |
| 2006 | Minimum phone error training of precision matrix modelsabstractGaussian mixture models (GMMs) are commonly used as the output density function for large-vocabulary continuous speech recognition (LVCSR) systems. A standard problem when using multivariate GMMs to classify data is how to accurately represent the correlations in the feature vector. Full covariance matrices yield a good model, but dramatically increase the number of model parameters. Hence, diagonal covariance matrices are commonly used. Structured precision matrix approximations provide an alternative, flexible, and compact representation. Schemes in this category include the extended maximum likelihood linear transform and subspace for precision and mean models. This paper examines how these precision matrix models can be discriminatively trained and used on state-of-the-art speech recognition tasks. In particular, the use of the minimum phone error criterion is investigated. Implementation issues associated with building LVCSR systems are also addressed. These models are evaluated and compared using large vocabulary continuous telephone speech and broadcast news English tasks. Khe Chai Sim, Mark J. F. Gales |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Development of the CUHTK 2004 Mandarin Conversational Telephone Speech Transcription SystemabstractThe paper details all aspects of the CUHTK 2004 Mandarin conversational telephone speech transcription system, but concentrates on the development of the acoustic models. As there are significant differences between the available training corpora, both in terms of topics of conversation and accents, forms of data normalisation and adaptive training techniques are investigated. The baseline discriminatively trained acoustic models are compared to a system built with a Gaussianisation front-end, a speaker adaptively trained system and an adaptively trained structured precision matrix system. The models are finally evaluated within a multi-pass, multi-branch, system combination framework. Mark J. F. Gales, Xunying Liu, Khe Chai Sim, Philip C. Woodland, Kai Yu 0004 |
ICASSP (1) | 4 |
| 2005 | Development of the CU-HTK 2004 Broadcast News Transcription SystemsabstractThe paper describes our recent work on improving broadcast news transcription and presents details of the CU-HTK broadcast news English (BN-E) transcription system for the DARPA/NIST rich transcription 2004 speech-to-text (RT04) evaluation. A key focus has been building a system using an order of magnitude more acoustic training data than we have previously attempted. We have also investigated a range of techniques to improve both minimum phone error (MPE) training and the efficient creation of MPE-based narrow-band models. The paper describes two alternative system structures that run in under 10/spl times/RT and a further system that runs in less than 1/spl times/RT. This final system gives lower word error rates than our 2003 system that ran in 10/spl times/RT. Do Yeong Kim, Ricky Ho Yin Chan, Gunnar Evermann, Mark J. F. Gales, David Mrva, Khe Chai Sim, Philip C. Woodland |
ICASSP (1) | 6 |
| 2005 | Investigation of Acoustic Modeling Techniques for LVCSR SystemsabstractThe paper describes the use of several advanced acoustic modeling techniques for the 2004 CU-HTK large vocabulary speech recognition systems. These techniques include Gaussianization for speaker normalization, discriminative cluster adaptive training (CAT), subspace for precision and mean (SPAM) modeling of inverse covariances, and discriminative complexity control. Acoustic models featuring these techniques were integrated into a state-of-the-art 10 real-time multi-pass system with sophisticated adaptation for performance evaluation. Experimental results are presented on both broadcast news (BN) and conversational telephone speech (CTS) transcription tasks. Xunying Liu, Mark J. F. Gales, Khe Chai Sim, Kai Yu 0004 |
ICASSP (1) | 3 |
| 2005 | Adaptation of Precision Matrix Models on Large Vocabulary Continuous Speech RecognitionabstractRecently, structured precision matrix models were found to outperform the conventional diagonal covariance matrix models. Minimum phone error discriminative training of these models gave very good unadapted performance on large vocabulary continuous speech recognition systems. To obtain state-of-the-art performance, it is important to apply adaptation techniques efficiently to these models. In this paper, simple row-by-row iterative formulae are described for both MLLR mean and constrained MLLR transform estimations of these models. These update formulae are derived within the standard expectation maximisation framework and are guaranteed to increase the likelihood of the adaptation data. Efficient approximate schemes for these adaptation methods are also investigated to further reduce the computation. Experimental results are presented based on the MPE trained subspace for precision and mean models, evaluated on both broadcast news and conversational telephone speech English tasks. Khe Chai Sim, Mark J. F. Gales |
ICASSP (1) | 1 |
| 2005 | Temporally varying model parameters for large vocabulary continuous speech recognitionabstractMany forms of time varying acoustic models have been applied to the area of speech recognition. However, there has been little success in applying these models to Large Vocabulary Continuous Speech Recognition (LVCSR). Recently, fMPE was introduced as a discriminative feature space estimation scheme for the HMM-based LVCSR. This method estimates a projection matrix from a high dimensional space ( ∼ 100,000) down to a standard feature space (typically 39). This projection is then added on to the original feature vector (e.g. MFCC or PLP) to yield a feature vector to train the final model. This paper considers fMPE as a time varying model for the mean vectors by applying the time varying feature offset to the Gaussian mean vectors. This approach naturally yields the update formulae for fMPE and motivates an alternative style of training systems. This concept is then extended to the temporal precision matrix modelling (pMPE). In pMPE, a temporally varying positive scale is applied to each element of the diagonal precision matrices. Experimental results are presented on a conversational telephone speech English task. 1. Khe Chai Sim, Mark J. F. Gales |
INTERSPEECH | 1 |
| 2004 | Basis superposition precision matrix modelling for large vocabulary continuous speech recognitionabstractAn important aspect of using Gaussian mixture models in a HMM-based speech recognition systems is the form of the covariance matrix. One successful approach has been to model the inverse covariance, precision, matrix by superimposing multiple bases. This paper presents a general framework of basis superposition. Models are described in terms of parameter tying of the basis coefficients and restrictions in the number of basis. Two forms of parameter tying are described which provide a compact model structure. The first constrains the basis coefficients over multiple basis vectors (or matrices). This is related to the Subspace for Precision and Mean (SPAM) model. The second constrains the basis coefficients over multiple components, yielding as one example heteroscedastic LDA (HLDA). Both maximum likelihood and minimum phone error training of these models are discussed. The performance of various configurations is examined on a conversational telephone speech task, SwitchBoard. Khe Chai Sim, Mark J. F. Gales |
ICASSP (1) | 1 |