EDBT 2026 Demo / reviewers in the wild / expert
Hainan Xu
dblp:120/4014
· DBLP profile ↗
36ranked-venue papers
13as first author
17since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 9 first-author · 15 since 2021Artificial intelligence and machine learning · 24 · 6 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HAINAN: Fast and Accurate Transducer for Hybrid-Autoregressive ASRabstractWe present Hybrid-Autoregressive INference TrANsducers (HAINAN), a novel architecture for speech recognition that extends the Token-and-Duration Transducer (TDT) model. Trained with randomly masked predictor network outputs, HAINAN supports both autoregressive inference with all network components and non-autoregressive inference without the predictor. Additionally, we propose a novel semi-autoregressive inference method that first generates an initial hypothesis using non-autoregressive inference, followed by refinement steps where each token prediction is regenerated using parallelized autoregression on the initial hypothesis. Experiments on multiple datasets across different languages demonstrate that HAINAN achieves efficiency parity with CTC in non-autoregressive mode and with TDT in autoregressive mode. In terms of accuracy, autoregressive HAINAN achieves parity with TDT and RNN-T, while non-autoregressive HAINAN significantly outperforms CTC. Semi-autoregressive inference further enhances the model's accuracy with minimal computational overhead, and even outperforms TDT results in some cases. These results highlight HAINAN's flexibility in balancing accuracy and speed, positioning it as a strong candidate for real-world speech recognition applications. Hainan Xu, Travis M. Bartley, Vladimir Bataev, Boris Ginsburg |
ICLR | 1 |
| 2025 | Pushing the Limits of Beam Search Decoding for Transducer-based ASR models
Lilit Grigoryan, Vladimir Bataev, Andrei Andrusenko, Hainan Xu, Vitaly Lavrukhin, Boris Ginsburg |
INTERSPEECH | 4 |
| 2025 | Word Level Timestamp Generation for Automatic Speech Recognition and Translation
Krishna C. Puvvada, Elena Rastorgueva, Zhehuai Chen, He Huang 0012, Shuoyang Ding, Kunal Dhawan, Hainan Xu, Jagadeesh Balam, Boris Ginsburg |
INTERSPEECH | 8 |
| 2025 | WIND: Accelerated RNN-T Decoding with Windowed Inference for Non-blank Detection
Hainan Xu, Vladimir Bataev, Lilit Grigoryan, Boris Ginsburg |
INTERSPEECH | 1 |
| 2024 | TDT-KWS: Fast and Accurate Keyword Spotting Using Token-and-Duration TransducerabstractDesigning an efficient keyword spotting (KWS) system that delivers exceptional performance on resource-constrained edge devices has long been a subject of significant attention. Existing KWS search algorithms typically follow a frame-synchronous approach, where search decisions are made repeatedly at each frame despite the fact that most frames are keyword-irrelevant. In this paper, we propose TDT-KWS, which leverages token-and-duration Transducers (TDT) for KWS tasks. We also propose a novel KWS task-specific decoding algorithm for Transducer-based models, which supports highly effective frame-asynchronous keyword search in streaming speech scenarios. With evaluations conducted on both the public Hey Snips and self-constructed LibriKWS-20 datasets, our proposed KWS-decoding algorithm produces more accurate results than conventional ASR decoding algorithms. Additionally, TDTKWS achieves on-par or better wake word detection performance than both RNN-T and traditional TDT-ASR systems while achieving significant inference speed-up. Furthermore, experiments show that TDT-KWS is more robust to noisy environments compared to RNN-T KWS. Yu Xi, Baochen Yang, Hainan Xu |
ICASSP | 5 |
| 2024 | Transducers with Pronunciation-Aware Embeddings for Automatic Speech RecognitionabstractThis paper proposes Transducers with Pronunciation-aware Embeddings (PET). Unlike conventional Transducers where the decoder embeddings for different tokens are trained independently, the PET model’s decoder embedding incorporates shared components for text tokens with the same or similar pronunciations. With experiments conducted in multiple datasets in Mandarin Chinese and Korean, we show that PET models consistently improve speech recognition accuracy compared to conventional Transducers. Our investigation also uncovers a phenomenon that we call error chain reactions. Instead of recognition errors being evenly spread throughout an utterance, they tend to group together, with subsequent errors often following earlier ones. Our analysis shows that PET models effectively mitigate this issue by substantially reducing the likelihood of the model generating additional errors following a prior one. Our implementation will be open-sourced with the NeMo toolkit. Hainan Xu, Zhehuai Chen, Fei Jia, Boris Ginsburg |
ICASSP | 1 |
| 2024 | Speed of Light Exact Greedy Decoding for RNN-T Speech Recognition Models on GPU
Daniel Galvez, Vladimir Bataev, Hainan Xu, Tim Kaldewey |
INTERSPEECH | 3 |
| 2024 | Label-Looping: Highly Efficient Decoding For TransducersabstractThis paper introduces a highly efficient greedy decoding algorithm for Transducer-based speech recognition models. We redesign the standard nested-loop design for RNN-T decoding, swapping loops over frames and labels: the outer loop iterates over labels, while the inner loop iterates over frames searching for the next non-blank symbol. Additionally, we represent partial hypotheses in a special structure using CUDA tensors, supporting parallelized hypotheses manipulations. Experiments show that the label-looping algorithm is up to 2.0X faster than conventional batched decoding when using batch size 32. It can be further combined with other compiler or GPU call-related techniques to achieve even more speedup. Our algorithm is general-purpose and can work with both conventional Transducers and Token-and-Duration Transducers. We open-source our implementation to benefit the research community. Vladimir Bataev, Hainan Xu, Daniel Galvez, Vitaly Lavrukhin, Boris Ginsburg |
SLT | 2 |
| 2024 | Romanization Encoding For Multilingual ASRabstractWe introduce romanization encoding for script-heavy languages to optimize multilingual and code-switching Automatic Speech Recognition (ASR) systems. By adopting romanization encoding alongside a balanced concatenated tokenizer within a FastConformer-RNNT framework equipped with a Roman2Char module, we significantly reduce vocabulary and output dimensions, enabling larger training batches and reduced memory consumption. Our method decouples acoustic modeling and language modeling, enhancing the flexibility and adaptability of the system. In our study, applying this method to Mandarin-English ASR resulted in a remarkable 63.51% vocabulary reduction and notable performance gains of 13.72% and 15.03% on SEAME code-switching benchmarks. Ablation studies on MandarinKorean and Mandarin-Japanese highlight our method’s strong capability to address the complexities of other script-heavy languages, paving the way for more versatile and effective multilingual ASR systems. Wen Ding 0005, Fei Jia, Hainan Xu, Yu Xi, Junjie Lai, Boris Ginsburg |
SLT | 3 |
| 2024 | Longer is (Not Necessarily) Stronger: Punctuated Long-Sequence Training for Enhanced Speech Recognition and TranslationabstractThis paper presents a new method for training sequence-to-sequence models for speech recognition and translation tasks. Instead of the traditional approach of training models on short segments containing only lowercase or partial punctuation and capitalization (PnC) sentences, we propose training on longer utterances that include complete sentences with proper punctuation and capitalization. We achieve this by using the FastConformer architecture which allows training 1 Billion parameter models with sequences up to 60 seconds long with full attention. However, while training with PnC enhances the overall performance, we observed that accuracy plateaus when training on sequences longer than 40 seconds across various evaluation settings. Our proposed method significantly improves punctuation and capitalization accuracy, showing a 25% relative word error rate (WER) improvement on the Earnings-21 and Earnings-22 benchmarks. Additionally, training on longer audio segments increases the overall model accuracy across speech recognition and translation benchmarks. The model weights and training code are open-sourced though NVIDIA NeMo.121https://github.com/NVIDIA/NeMo2https://huggingface.co/nvidia/parakeet-tdt_ctc-1.1b Nithin Rao Koluguri, Travis M. Bartley, Hainan Xu, Oleksii Hrinchuk, Jagadeesh Balam, Boris Ginsburg, Georg Kucsko |
SLT | 3 |
| 2023 | Learning From Flawed Data: Weakly Supervised Automatic Speech RecognitionabstractTraining automatic speech recognition (ASR) systems requires large amounts of well-curated paired data. However, human annotators usually perform “non-verbatim” transcription, which can result in poorly trained models. In this paper, we propose Omni-temporal Classification (OTC), a novel training criterion that explicitly incorporates label uncertainties originating from such weak supervision. This allows the model to effectively learn speech-text alignments while accommodating errors present in the training transcripts. OTC extends the conventional CTC objective for imperfect transcripts by leveraging weighted finite state transducers. Through experiments conducted on the LibriSpeech and LibriVox datasets, we demonstrate that training ASR models with OTC avoids performance degradation even with transcripts containing up to 70% errors, a scenario where CTC models fail completely. Our implementation is available at https://github.com/k2-fsa/icefall. Dongji Gao, Hainan Xu, Desh Raj, L. Paola García-Perera, Daniel Povey, Sanjeev Khudanpur |
ASRU | 2 |
| 2023 | Multi-Blank Transducers for Speech RecognitionabstractThis paper proposes a modification to RNN-Transducer (RNN-T) models for automatic speech recognition (ASR). In standard RNN-T, the emission of a blank symbol consumes exactly one input frame; in our proposed method, we introduce additional blank symbols, which consume two or more input frames when emitted. We refer to the added symbols as big blanks, and the method multi-blank RNN-T. For training multi-blank RNN-Ts, we propose a novel logit under-normalization method in order to prioritize emissions of big blanks. With experiments on multiple languages and datasets, we show that multi-blank RNN-T methods could bring relative speedups of over +90%/+139% to model inference for English Librispeech and German Multilingual Librispeech datasets, respectively. The multi-blank RNN-T method also improves ASR accuracy consistently. We will release our implementation of the method in the NeMo (https://github.com/NVIDIA/NeMo) toolkit. Hainan Xu, Fei Jia, Somshubra Majumdar, Shinji Watanabe 0001, Boris Ginsburg |
ICASSP | 1 |
| 2023 | Efficient Sequence Transduction by Jointly Predicting Tokens and DurationsabstractThis paper introduces a novel Token-and-Duration Transducer (TDT) architecture for sequence-to-sequence tasks. TDT extends conventional RNN-Transducer architectures by jointly predicting both a token and its duration, i.e. the number of input frames covered by the emitted token. This is achieved by using a joint network with two outputs which are independently normalized to generate distributions over tokens and durations. During inference, TDT models can skip input frames guided by the predicted duration output, which makes them significantly faster than conventional Transducers which process the encoder output frame by frame. TDT models achieve both better accuracy and significantly faster inference than conventional Transducers on different sequence transduction tasks. TDT models for Speech Recognition achieve better accuracy and up to 2.82X faster inference than conventional Transducers. TDT models for Speech Translation achieve an absolute gain of over 1 BLEU on the MUST-C test compared with conventional Transducers, and its inference is 2.27X faster. In Speech Intent Classification and Slot Filling tasks, TDT models improve the intent accuracy by up to over 1% (absolute) over conventional Transducers, while running up to 1.28X faster. Our implementation of the TDT model will be open-sourced with the NeMo (https://github.com/NVIDIA/NeMo) toolkit. Hainan Xu, Fei Jia, Somshubra Majumdar, He Huang 0012, Shinji Watanabe 0001, Boris Ginsburg |
ICML | 1 |
| 2023 | Bypass Temporal Classification: Weakly Supervised Automatic Speech Recognition with Imperfect Transcripts
Dongji Gao, Matthew Wiesner, Hainan Xu, L. Paola García-Perera, Daniel Povey, Sanjeev Khudanpur |
INTERSPEECH | 3 |
| 2021 | An Asynchronous WFST-Based Decoder for Automatic Speech RecognitionabstractWe introduce asynchronous dynamic decoder, which adopts an efficient A* algorithm to incorporate big language models in the one-pass decoding for large vocabulary continuous speech recognition. Unlike standard one-pass decoding with on-the-fly composition decoder which might induce a significant computation overhead, the asynchronous dynamic decoder has a novel design where it has two fronts, with one performing "exploration" and the other "backfill". The computation of the two fronts alternates in the decoding process, resulting in more effective pruning than the standard one-pass decoding with an on-the-fly composition decoder. Experiments show that the proposed decoder works notably faster than the standard one-pass decoding with on-the-fly composition decoder, while the acceleration will be more obvious with the increment of data complexity. Hang Lv 0001, Zhehuai Chen, Hainan Xu, Daniel Povey, Lei Xie 0001, Sanjeev Khudanpur |
ICASSP | 3 |
| 2021 | Convolutional Dropout and Wordpiece Augmentation for End-to-End Speech RecognitionabstractRegularization and data augmentation are crucial to training end-to-end automatic speech recognition systems. Dropout is a popular regularization technique, which operates on each neuron independently by multiplying it with a Bernoulli random variable. We propose a generalization of dropout, called "convolutional dropout", where each neuron’s activation is replaced with a randomly-weighted linear combination of neuron values in its neighborhood. We believe that this formulation combines the regularizing effect of dropout with the smoothing effects of the convolution operation. In addition to convolutional dropout, this paper also proposes using random word-piece segmentations as a data augmentation scheme during training, inspired by results in neural machine translation. We adopt both these methods during the training of transformer-transducer speech recognition models, and show consistent WER improvements on Librispeech as well as across different languages. Hainan Xu, Kartik Audhkhasi, Bhuvana Ramabhadran |
ICASSP | 1 |
| 2021 | Regularizing Word Segmentation by Creating Misspellings
Hainan Xu, Kartik Audhkhasi, Jesse Emond, Bhuvana Ramabhadran |
Interspeech | 1 |
| 2020 | Multilingual Speech Recognition with Self-Attention Structured Parameterization
Parisa Haghani, Anshuman Tripathi, Bhuvana Ramabhadran, Brian Farris, Hainan Xu, Han Lu 0003, Hasim Sak, Isabel Leal, Neeraj Gaur, Pedro J. Moreno 0001 |
INTERSPEECH | 6 |
| 2019 | Incremental Lattice Determinization for WFST DecodersabstractWe introduce a lattice determinization algorithm that can operate incrementally. That is, a word-level lattice can be generated for a partial utterance and then, once we have processed more audio, we can obtain a word-level lattice for the extended utterance without redoing all the work of lattice determinization. This is relevant for ASR decoders such as those used in Kaldi, which first generate a state-level lattice and then convert it to a word-level lattice using a determinization algorithm in a special semiring. Our incremental determinization algorithm is useful when word-level lattices are needed prior to the end of the utterance, and also reduces the latency due to determinization at the end of the utterance. Zhehuai Chen, Mahsa Yarmohammadi, Hainan Xu, Hang Lv 0001, Lei Xie 0001, Daniel Povey, Sanjeev Khudanpur |
ASRU | 3 |
| 2019 | Espresso: A Fast End-to-End Neural Speech Recognition ToolkitabstractWe present Espresso, an open-source, modular, extensible end-to-end neural automatic speech recognition (ASR) toolkit based on the deep learning library PyTorch and the popular neural machine translation toolkit FAIRSEQ. ESRESSO supports distributed training across GPUs and computing nodes, and features various decoding approaches commonly employed in ASR, including look-ahead word-based language model fusion, for which a fast, parallelized decoder is implemented. Espresso achieves state-of-the-art ASR performance on the WSJ, LibriSpeech, and Switchboard data sets among other end-to-end systems without data augmentation, and is 4-11x faster for decoding than similar systems (e.g. ESPNET). Yiming Wang 0006, Sanjeev Khudanpur, Tongfei Chen, Hainan Xu, Shuoyang Ding, Hang Lv 0001, Yiwen Shao, Nanyun Peng 0001, Lei Xie 0001, Shinji Watanabe 0001 |
ASRU | 4 |
| 2019 | Improving End-to-end Speech Recognition with Pronunciation-assisted Sub-word ModelingabstractMost end-to-end speech recognition systems model text directly as a sequence of characters or sub-words. Current approaches to sub-word extraction only consider character sequence frequencies, which at times produce inferior sub-word segmentation that might lead to erroneous speech recognition output. We propose pronunciation-assisted sub-word modeling (PASM), a sub-word extraction method that leverages the pronunciation information of a word. Experiments show that the proposed method can greatly improve upon the character-based baseline, and also outperform commonly used byte-pair encoding methods. Hainan Xu, Shuoyang Ding, Shinji Watanabe 0001 |
ICASSP | 1 |
| 2019 | The JHU ASR System for VOiCES from a Distance Challenge 2019
Yiming Wang 0006, David Snyder, Hainan Xu, Vimal Manohar, Phani S. Nidadavolu, Daniel Povey, Sanjeev Khudanpur |
INTERSPEECH | 3 |
| 2019 | Robust Document Representations for Cross-Lingual Information Retrieval in Low-Resource Settings
Mahsa Yarmohammadi, Xutai Ma, Sorami Hisamoto, Muhammad Mahbubur Rahman 0001, Yiming Wang 0006, Hainan Xu, Daniel Povey, Philipp Koehn, Kevin Duh |
MTSummit (1) | 6 |
| 2018 | A Pruned Rnnlm Lattice-Rescoring Algorithm for Automatic Speech RecognitionabstractLattice-rescoring is a common approach to take advantage of recurrent neural language models in ASR, where a word-lattice is generated from 1st-pass decoding and the lattice is then rescored with a neural model, and ann-gram approximation method is usually adopted to limit the search space. In this work, we describe a pruned lattice-rescoring algorithm for ASR, improving the n-gram approximation method. The pruned algorithm further limits the search space and uses heuristic search to pick better histories when expanding the lattice. Experiments show that the proposed algorithm achieves better ASR accuracies while running much faster than the standard algorithm. In particular, it brings a 4x speedup for lattice-rescoring with 4-gram approximation while giving better recognition accuracies than the standard algorithm. Hainan Xu, Tongfei Chen, Dongji Gao, Yiming Wang 0006, Ke Li 0018, Nagendra K. Goel, Yishay Carmiel, Daniel Povey, Sanjeev Khudanpur |
ICASSP | 1 |
| 2018 | Neural Network Language Modeling with Letter-Based Features and Importance SamplingabstractIn this paper we describe an extension of the Kaldi software toolkit to support neural-based language modeling, intended for use in automatic speech recognition (ASR) and related tasks. We combine the use of subword features (letter n-grams) and one-hot encoding of frequent words so that the models can handle large vocabularies containing infrequent words. We propose a new objective function that allows for training of unnormalized probabilities. An importance sampling based method is supported to speed up training when the vocabulary is large. Experimental results on five corpora show that Kaldi-RNNLM rivals other recurrent neural network language model toolkits both on performance and training speed. Hainan Xu, Ke Li 0018, Yiming Wang 0006, Shiyin Kang, Xie Chen 0001, Daniel Povey, Sanjeev Khudanpur |
ICASSP | 1 |
| 2018 | A GPU-based WFST Decoder with Exact Lattice GenerationabstractWe describe initial work on an extension of the Kaldi toolkit that supports weighted finite-state transducer (WFST) decoding on Graphics Processing Units (GPUs). We implement token recombination as an atomic GPU operation in order to fully parallelize the Viterbi beam search, and propose a dynamic load balancing strategy for more efficient token passing scheduling among GPU threads. We also redesign the exact lattice generation and lattice pruning algorithms for better utilization of the GPUs. Experiments on the Switchboard corpus show that the proposed method achieves identical 1-best results and lattice quality in recognition and confidence measure tasks, while running 3 to 15 times faster than the single process Kaldi decoder. The above results are reported on different GPU architectures. Additionally we obtain a 46-fold speedup with sequence parallelism and multi-process service (MPS) in GPU. Zhehuai Chen, Justin Luitjens, Hainan Xu, Yiming Wang 0006, Daniel Povey, Sanjeev Khudanpur |
INTERSPEECH | 3 |
| 2018 | Building State-of-the-art Distant Speech Recognition Using the CHiME-4 Challenge with a Setup of Speech Enhancement BaselineabstractThis paper describes a new baseline system for automatic speech recognition (ASR) in the CHiME-4 challenge to promote the development of noisy ASR in speech processing communities by providing 1) state-of-the-art system with a simplified single system comparable to the complicated top systems in the challenge, 2) publicly available and reproducible recipe through the main repository in the Kaldi speech recognition toolkit.The proposed system adopts generalized eigenvalue beamforming with bidirectional long short-term memory (LSTM) mask estimation.We also propose to use a time delay neural network (TDNN) based on the lattice-free version of the maximum mutual information (LF-MMI) trained with augmented all six microphones plus the enhanced data after beamforming.Finally, we use a LSTM language model for lattice and n-best re-scoring.The final system achieved 2.74% WER for the real test set in the 6-channel track, which corresponds to the 2nd place in the challenge.In addition, the proposed baseline recipe includes four different speech enhancement measures, short-time objective intelligibility measure (STOI), extended STOI (eSTOI), perceptual evaluation of speech quality (PESQ) and speech distortion ratio (SDR) for the simulation test set.Thus, the recipe also provides an experimental platform for speech enhancement studies with these performance measures. Szu-Jui Chen, Aswin Shanmugam Subramanian, Hainan Xu, Shinji Watanabe 0001 |
INTERSPEECH | 3 |
| 2018 | Recurrent Neural Network Language Model Adaptation for Conversational Speech Recognition
Ke Li 0018, Hainan Xu, Yiming Wang 0006, Daniel Povey, Sanjeev Khudanpur |
INTERSPEECH | 2 |
| 2018 | Semi-Orthogonal Low-Rank Matrix Factorization for Deep Neural Networks
Daniel Povey, Gaofeng Cheng, Yiming Wang 0006, Ke Li 0018, Hainan Xu, Mahsa Yarmohammadi, Sanjeev Khudanpur |
INTERSPEECH | 5 |
| 2017 | Zipporah: a Fast and Scalable Data Cleaning System for Noisy Web-Crawled Parallel CorporaabstractWe introduce Zipporah, a fast and scalable data cleaning system.We propose a novel type of bag-of-words translation feature, and train logistic regression models to classify good data and synthetic noisy data in the proposed feature space.The trained model is used to score parallel sentences in the data pool for selection.As shown in experiments, Zipporah selects a high-quality parallel corpus from a large, mixed quality data pool.In particular, for one noisy dataset, Zipporah achieves a 2.1 BLEU score improvement with using 1/5 of the data over using the entire corpus. Hainan Xu, Philipp Koehn |
EMNLP | 1 |
| 2017 | The Kaldi OpenKWS System: Improving Low Resource Keyword Search
Jan Trmal, Matthew Wiesner, Vijayaditya Peddinti, Xiaohui Zhang 0007, Pegah Ghahremani, Yiming Wang 0006, Vimal Manohar, Hainan Xu, Daniel Povey, Sanjeev Khudanpur |
INTERSPEECH | 8 |
| 2017 | Backstitch: Counteracting Finite-Sample Bias via Negative Steps
Yiming Wang 0006, Vijayaditya Peddinti, Hainan Xu, Xiaohui Zhang 0007, Daniel Povey, Sanjeev Khudanpur |
INTERSPEECH | 3 |
| 2015 | Pronunciation and silence probability modeling for ASRabstractIn this paper we evaluate the WER improvement from modeling pronunciation probabilities and word-specific silence probabilities in speech recognition. We do this in the context of Finite State Transducer (FST)-based decoding, where pronunciation and silence probabilities are encoded in the lexicon (L) transducer. We describe a novel way to model word-dependent silence probabilities, where in addition to modeling the probability of silence following each individual word, we also model the probability of each word appearing after silence. All of these probabilities are estimated from aligned training data, with suitable smoothing. We conduct our experiments on four commonly used automatic speech recognition datasets, namely Wall Street Journal, Switchboard, TED-LIUM, and Librispeech. The improvement from modeling pronunciation and silence probabilities is small but fairly consistent across datasets. Guoguo Chen, Hainan Xu, Minhua Wu, Daniel Povey, Sanjeev Khudanpur |
INTERSPEECH | 2 |
| 2015 | Modeling phonetic context with non-random forests for speech recognitionabstractModern speech recognition systems typically cluster triphone phonetic contexts using decision trees. In this paper we describe a way to build multiple complementary decision trees from the same data, for the purpose of system combination. We do this by jointly building the decision trees using an objective function that has an added entropy term to encourage diversity among the decision trees. After the trees are built, the systems are built in the standard way and the emission probabilities are combined during decoding. Experiments on multiple datasets show gains from the use of multiple trees, at the expense of evaluating multiple models in test time. Hainan Xu, Guoguo Chen, Daniel Povey, Sanjeev Khudanpur |
INTERSPEECH | 1 |
| 2013 | Cluster adaptive training with factorized decision trees for speech recognitionabstractCluster adaptive training (CAT) is a popular approach to train multiple-cluster HMMs for fast speaker adaptation in speech recognition. Traditionally, a cluster-independent decision tree is shared among all clusters, which could limit the modelling power of multiple-cluster HMMs. In this paper, each cluster is allowed to have its own decision tree. The intersections between the triphones subsets, corresponding to the leaf nodes of each cluster-dependent trees, are used to define a finer state sharing structure. The parameters of these intersections are constructed from the parameters of the leaf nodes of each individual decision tree. This is referred to as CAT with factorized decision trees (FD-CAT). FD-CAT significantly increases the modelling power without introducing additional free parameters. A novel iterative mean cluster update approach and a robust covariance matrix update method with united statistics are proposed to efficiently train FD-CAT. Experiments showed that using multiple decision trees can yield better performance than single decision tree. Furthermore, FD-CAT significantly outperformed traditional CAT system. Index Terms Adaptation, factorized decision trees, cluster adaptive training, eigenvoices Hainan Xu |
INTERSPEECH | 2 |
| 2012 | Development of the 2012 SJTU HVR systemabstractHaptic voice recognition (HVR) is a multi-modal text entry method for smart mobile devices. It employs haptic events generated by speakers during speaking to achieve better efficiency and robustness for automatic speech recognition. This paper describes the detailed design of the 2012 SJTU submission for the HVR Grand Challenge. During the design, a new perplexity metric using conditional entropy is proposed to evaluate the potential search space reduction of a haptic event without speech input. A number of new haptic events are evaluated both theoretically and experimentally in detail. The final submission system uses the haptic event of initial letter plus final letter and reduces word error rate by 76% compared to the baseline initial letter event. Hainan Xu, Yuchen Fan 0001, Kai Yu 0004 |
ICMI | 1 |