VLDB 2026 Research / reviewers in the wild / expert
Yifan Gong 0001
dblp:49/3073-1
· DBLP profile ↗
206ranked-venue papers
34as first author
30since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 189 · 29 first-author · 27 since 2021Artificial intelligence and machine learning · 106 · 21 first-author · 16 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Influence-based Online Experience Selection for Effective RLHFabstractReinforcement Learning from Human Feedback (RLHF) has emerged as a crucial technique for aligning large language models (LLMs) with human preferences.However, existing RLHF methods face key challenges, including poor sample efficiency, high computational overhead, and slow convergence.Recent studies highlight the importance of data selection in RL, but how to effectively select the most beneficial experiences for RL training remains an open problem.Existing data selection methods for RL rely on heuristic metrics, failing to establish an interpretable connection between data and optimization objectives.To address this problem, we propose InfOES (Influence-based Online Experience Selection), a novel data selection method for RLHF that dynamically estimates the influence of individual training samples on policy optimization.By incorporating data attribution into the policy gradient, InfOES can identify and filter out detrimental samples on the fly, ensuring effective convergence toward alignment objectives.Our approach is compatible with various RL algorithms (e.g., PPO, GRPO, RE-INFORCE++).Extensive experiments demonstrate that InfOES significantly enhances training effectiveness, achieving superior alignment performance with fewer optimization steps. Yifan Gong 0001, Jing Yao 0003, Xiting Wang, Xunlong Wang, Xiaoyuan Yi, Xing Xie 0001 |
ACL (1) | 1 |
| 2025 | Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios
Aswin Shanmugam Subramanian, Amit Das 0007, Naoyuki Kanda, Jinyu Li 0001, Xiaofei Wang 0007, Yifan Gong 0001 |
INTERSPEECH | 6 |
| 2024 | Adapting Large Language Model with Speech for Fully Formatted End-to-End Speech RecognitionabstractMost end-to-end (E2E) speech recognition models are composed of encoder and decoder blocks that perform acoustic and language modeling functions. Pretrained large language models (LLMs) have the potential to improve the performance of E2E ASR. However, integrating a pretrained language model into an E2E speech recognition model has shown limited benefits due to the mismatches between text-based LLMs and those used in E2E ASR. In this paper, we explore an alternative approach by adapting a pretrained LLMs to speech. Our experiments on fully-formatted E2E ASR transcription tasks across various domains demonstrate that our approach can effectively leverage the strengths of pretrained LLMs to produce more readable ASR transcriptions. Our model, which is based on the pretrained large language models with either an encoder-decoder or decoder-only structure, surpasses strong ASR models such as Whisper1, in terms of recognition error rate, considering formats like punctuation and capitalization as well. Shaoshi Ling, Yuxuan Hu 0003, Shuangbei Qian, Guoli Ye, Yao Qian, Yifan Gong 0001, Ed Lin, Michael Zeng 0001 |
ICASSP | 6 |
| 2024 | NOTSOFAR-1 Challenge: New Datasets, Baseline, and Tasks for Distant Meeting Transcription
Alon Vinnikov, Amir Ivry, Aviv Hurvitz, Igor Abramovski, Sharon Koubi, Ilya Gurvich, Shai Pe'er, Benjamin Elizalde, Naoyuki Kanda, Xiaofei Wang 0009, Shalev Shaer, Stav Yagev, Yossi Asher, Sunit Sivasankaran, Yifan Gong 0001, Huaming Wang, Eyal Krupka |
INTERSPEECH | 16 |
| 2024 | Value FULCRA: Mapping Large Language Models to the Multidimensional Spectrum of Basic Human ValueabstractJing Yao, Xiaoyuan Yi, Yifan Gong, Xiting Wang, Xing Xie. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jing Yao 0003, Xiaoyuan Yi, Yifan Gong 0001, Xiting Wang, Xing Xie 0001 |
NAACL-HLT | 3 |
| 2024 | Hybrid Attention-Based Encoder-Decoder Model for Efficient Language Model AdaptationabstractThe attention-based encoder-decoder (AED) speech recognition model has been widely successful in recent years. However, the joint optimization of acoustic model and language model in end-to-end manner has created challenges for text adaptation. In particular, effective, quick and inexpensive adaptation with text input has become a primary concern for deploying AED systems in the industry. To address this issue, we propose a novel model, the hybrid attention-based encoder-decoder (HAED) speech recognition model that preserves the modularity of conventional hybrid automatic speech recognition systems. Our HAED model separates the acoustic and language models, allowing for the use of conventional text-based language model adaptation techniques. We demonstrate that the proposed HAED model yields 23% relative Word Error Rate (WER) improvements when out-of-domain text data is used for language model adaptation, with only a minor degradation in WER on a general test set compared with the conventional AED model. Shaoshi Ling, Guoli Ye, Rui Zhao 0017, Yifan Gong 0001 |
SLT | 4 |
| 2023 | Multi Transcription-Style Speech Transcription Using Attention-Based Encoder-Decoder ModelabstractHuman professional transcription services provide a variety of transcription styles to customize different needs. To accommodate different users and facilitate seamless integration with downstream applications, we propose a framework to generate multi-style transcription in an attention-based encoder-decoder model (AED) using three different architectures: (A) style-dependent layers; (B) mixed-style output; (C) style-dependent prompt. In this framework, both the verbatim lexical transcription and the readable transcription of various styles can be generated simultaneously or separately, through a single decoding pass or multiple decoding passes on-demand. We conduct experiments in a large-scale AED-based speech transcription system trained with 50k hours speech. The proposed framework can achieve nearly on-par performance compared to the single-style AED with significant savings in model footprint and decoding cost. Moreover, it provides an efficient data sharing mechanism across different styles through knowledge transfer. Yan Huang 0028, Piyush Behre, Guoli Ye, Shawn Chang, Yifan Gong 0001 |
ASRU | 5 |
| 2023 | Building High-Accuracy Multilingual ASR With Gated Language Experts and Curriculum TrainingabstractWe propose gated language experts and curriculum training to enhance multilingual transformer transducer models without requiring user input for language identification (LID) during inference. Our method incorporates a gating mechanism and LID loss to enable transformer experts to learn language-specific information. Linear experts are applied on joint network to stabilize training. The curriculum training scheme leverages LID to guide gated experts in improving their respective language-specific performance. Experimental results on an English and Spanish bilingual task show significant average relative word error reductions of 12.5 % and 7.3 % compared to the baseline bilingual and monolingual models, respectively. Our models even perform similarly to upper-bound models with oracle LID. Extending our approach to trilingual, quadrilingual, and pentalingual models reveals similar advantages to those seen in the bilingual models, highlighting its ease of extension to multiple languages. Eric Sun, Jinyu Li 0001, Yuxuan Hu 0003, Yimeng Zhu, Linquan Liu, Shujie Liu 0001, Edward Lin, Yifan Gong 0001 |
ASRU | 11 |
| 2022 | Endpoint Detection for Streaming End-to-End Multi-Talker ASRabstractStreaming end-to-end multi-talker speech recognition aims at transcribing the overlapped speech from conversations or meetings with an all-neural model in a streaming fashion, which is fundamentally different from a modular-based approach that usually cascades the speech separation and the speech recognition models trained independently. Previously, we proposed the Streaming Unmixing and Recognition Transducer (SURT) model based on recurrent neural network transducer (RNN-T) for this problem and presented promising results. However, for real applications, the speech recognition system is also required to determine the times-tamp when a speaker finishes speaking for prompt system response. This problem, known as endpoint (EP) detection, has not been studied previously for multi-talker end-to-end models. In this work, we address the EP detection problem in the SURT framework by introducing an end-of-sentence token as an output unit, following the practice of single-talker end-to-end models. Furthermore, we also present a latency penalty approach that can significantly cut down the EP detection latency. Our experimental results based on the 2-speaker LibrispeechMix dataset show that the SURT model can achieve promising EP detection without significantly degradation of the recognition accuracy. Liang Lu 0001, Jinyu Li 0001, Yifan Gong 0001 |
ICASSP | 3 |
| 2022 | Have Best of Both Worlds: Two-Pass Hybrid and E2E Cascading Framework for Speech RecognitionabstractHybrid and end-to-end (E2E) systems have their individual advantages, with different error patterns in the speech recognition results. By jointly modeling audio and text, the E2E model performs better in matched scenarios and scales well with a large amount of paired audio-text training data. The modularized hybrid model is easier for customization, and better to make use of a massive amount of unpaired text data. This paper proposes a two-pass hybrid and E2E cascading (HEC) framework to combine the hybrid and E2E model in order to take advantage of both sides, with hybrid in the first pass and E2E in the second pass. We show that the proposed system achieves 8-10% relative word error rate reduction with respect to each individual system. More importantly, compared with the pure E2E system, we show the proposed system has the potential to keep the advantages of hybrid system, e.g., customization and segmentation capabilities. We also show the second pass E2E model in HEC is robust with respect to the change in the first pass hybrid model. Guoli Ye, Vadim Mazalov, Jinyu Li 0001, Yifan Gong 0001 |
ICASSP | 4 |
| 2022 | Internal Language Model Adaptation with Text-Only Data for End-to-End Speech RecognitionabstractText-only adaptation of an end-to-end (E2E) model remains a challenging task for automatic speech recognition (ASR).Language model (LM) fusion-based approaches require an additional external LM during inference, significantly increasing the computation cost.To overcome this, we propose an internal LM adaptation (ILMA) of the E2E model using text-only data.Trained with audio-transcript pairs, an E2E model implicitly learns an internal LM that characterizes the token sequence probability which is approximated by the E2E model output after zeroing out the encoder contribution.During ILMA, we fine-tune the internal LM, i.e., the E2E components excluding the encoder, to minimize a cross-entropy loss.To make ILMA effective, it is essential to train the E2E model with an internal LM loss besides the standard E2E loss.Furthermore, we propose to regularize ILMA by minimizing the Kullback-Leibler divergence between the output distributions of the adapted and unadapted internal LMs.ILMA is the most effective when we update only the last linear layer of the joint network.ILMA enables a fast text-only adaptation of the E2E model without increasing the run-time computational cost.Experimented with 30K-hour trained transformer transducer models, ILMA achieves up to 34.9% relative word error rate reduction from the unadapted baseline. Zhong Meng, Yashesh Gaur, Naoyuki Kanda, Jinyu Li 0001, Xie Chen 0001, Yu Wu 0012, Yifan Gong 0001 |
INTERSPEECH | 7 |
| 2022 | Streaming, Fast and Accurate on-Device Inverse Text Normalization for Automatic Speech RecognitionabstractAutomatic Speech Recognition (ASR) systems typically yield output in lexical form. However, humans prefer a written form output. To bridge this gap, ASR systems usually employ Inverse Text Normalization (ITN). In previous works, Weighted Finite State Transducers (WFST) have been employed to do ITN. WFSTs are nicely suited to this task but their size and run-time costs can make deployment on embedded applications challenging. In this paper, we describe the development of an on-device ITN system that is streaming, lightweight & accurate. At the core of our system is a streaming transformer tagger, that tags lexical tokens from ASR. The tag informs which ITN category might be applied, if at all. Following that, we apply an ITN-category-specific WFST, only on the tagged text, to reliably perform the ITN conversion. We show that the proposed ITN solution performs equivalent to strong base-lines, while being significantly smaller in size and retaining customization capabilities. Yashesh Gaur, Nick Kibre, Kangyuan Shu, Issac Alphanso, Jinyu Li 0001, Yifan Gong 0001 |
SLT | 8 |
| 2022 | Diarisation Using Location Tracking with Agglomerative ClusteringabstractPrevious works have shown that spatial location information can be complementary to speaker embeddings for a speaker diarisation task. However, the models used often assume that speakers are fairly stationary throughout a meeting. This paper proposes to relax this assumption, by explicitly modelling the movements of speakers within an Agglomerative Hierarchical Clustering (AHC) diarisation framework. Kalman filters, which track the locations of speakers, are used to compute log-likelihood ratios that contribute to the cluster affinity computations for the AHC merging and stopping decisions. Experiments show that the proposed approach is able to yield improvements on a Microsoft rich meeting transcription task, compared to methods that do not use location information or that make stationarity assumptions. Jeremy H. M. Wong, Igor Abramovski, Yifan Gong 0001 |
SLT | 4 |
| 2022 | Joint Speaker Diarisation and Tracking in Switching State-Space ModelabstractSpeakers may move around while diarisation is being performed. When a microphone array is used, the instantaneous locations of where the sounds originated from can be estimated, and previous investigations have shown that such information can be complementary to speaker embeddings in the diarisation task. However, these approaches often assume that speakers are fairly stationary throughout a meeting. This paper relaxes this assumption, by proposing to explicitly track the movements of speakers while jointly performing diarisation within a unified model. A state-space model is proposed, where the hidden state expresses the identity of the current active speaker and the predicted locations of all speakers. The model is implemented as a particle filter. Experiments on a Microsoft rich meeting transcription task show that the proposed joint location tracking and diarisation approach is able to perform comparably with other methods that use location information. Jeremy H. M. Wong, Yifan Gong 0001 |
SLT | 2 |
| 2021 | On Addressing Practical Challenges for RNN-TransducerabstractIn this paper, several works are proposed to address practi-cal challenges for deploying RNN Transducer (RNN-T) based speech recognition systems. These challenges are adapting a well-trained RNN-T model to a new domain without col-lecting the audio data, obtaining time stamps and confidence scores at word level. We solve the first challenge with a splicing data method which concatenates the speech segments ex-tracted from the source domain data. To get time stamps, a phone prediction branch is added to the RNN-T model by sharing the encoder for the purpose of forced alignment. Fi-nally, we obtain word level confidence scores by utilizing sev-eral types of features calculated during decoding and from a confusion network. Evaluated with Microsoft production data, the splicing data adaptation method improves the base-line and adaptation with the text to speech method by 58.03% and 15.25% relative word error rate reduction, respectively. The proposed time stamping method can get less than 50 mil-lisecond word timing difference from the ground truth align-ment on average while maintaining the recognition accuracy. We also obtain high confidence annotation performance with limited computation cost. Rui Zhao 0017, Jinyu Li 0001, Wenning Wei, Lei He 0005, Yifan Gong 0001 |
ASRU | 6 |
| 2021 | Internal Language Model Training for Domain-Adaptive End-To-End Speech RecognitionabstractThe efficacy of external language model (LM) integration with existing end-to-end (E2E) automatic speech recognition (ASR) systems can be improved significantly using the internal language model estimation (ILME) method [1]. In this method, the internal LM score is subtracted from the score obtained by interpolating the E2E score with the external LM score, during inference. To improve the ILME-based inference, we propose an internal LM training (ILMT) method to minimize an additional internal LM loss by updating only the E2E model components that affect the internal LM estimation. ILMT encourages the E2E model to form a standalone LM inside its existing components, without sacrificing ASR accuracy. After ILMT, the more modular E2E model with matched training and inference criteria enables a more thorough elimination of the source-domain internal LM, and therefore leads to a more effective integration of the target-domain external LM. Experimented with 30K-hour trained recurrent neural network transducer and attention-based encoder- decoder models, ILMT with ILME-based inference achieves up to 31.5% and 11.4% relative word error rate reductions from standard E2E training with Shallow Fusion on out-of-domain LibriSpeech and in-domain Microsoft production test sets, respectively. Zhong Meng, Naoyuki Kanda, Yashesh Gaur, Sarangarajan Parthasarathy, Eric Sun, Liang Lu 0001, Xie Chen 0001, Jinyu Li 0001, Yifan Gong 0001 |
ICASSP | 9 |
| 2021 | Sequence-Level Self-Teaching RegularizationabstractIn our previous research, we proposed a frame-level self-teaching network to regularize the deep neural network during training. In this paper, we extend the previous approach and propose a sequence self-teaching network to regularize the sequence-level information in speech recognition. The idea is to generate the sequence-level soft supervision labels from the top layer of the network to supervise the training of lower layer parameters. The network is trained with an auxiliary criterion in order to reduce the sequence-level Kullback-Leibler (KL) divergence between the top layer and lower layers, where the posterior probabilities in the KL-divergence term is computed from a lattice at the sequence-level. We evaluated the sequence-level self-teaching regularization approach with bidirectional long short-term memory models on LibriSpeech task, and show consistent improvements over the discriminative sequence maximum mutual information trained baseline. Eric Sun, Liang Lu 0001, Zhong Meng, Yifan Gong 0001 |
ICASSP | 4 |
| 2021 | Ensemble Combination between Different Time SegmentationsabstractHypothesis-level combination between multiple models can often yield gains in speech recognition. However, all models in the ensemble are usually restricted to use the same audio segmentation times. This paper proposes to generalise hypothesis-level combination, allowing the use of different audio segmentation times between the models, by splitting and re-joining the hypothesised N-best lists in time. A hypothesis tree method is also proposed to distribute hypothesis posteriors among the constituent words, to facilitate such splitting when per-word scores are not available. The approach is assessed on a Microsoft meeting transcription task, by performing combination between a streaming first-pass recognition and an offline second-pass recognition. The experimental results show that the proposed approach can yield gains when combining over different segmentation times. Furthermore, the results also show that a combination between a hybrid model and an end-to-end neural network model yields a greater improvement than a combination between two hybrid models. Jeremy H. M. Wong, Dimitrios Dimitriadis, Ken'ichi Kumatani, Yashesh Gaur, George Polovets, Partha Parthasarathy, Eric Sun, Jinyu Li 0001, Yifan Gong 0001 |
ICASSP | 9 |
| 2021 | Hidden Markov Model Diarisation with Speaker Location InformationabstractSpeaker diarisation methods often rely on speaker embeddings to cluster together the segments of audio that are uttered by the same speaker. When the audio is captured using a microphone array, it is possible to estimate the locations of where the sounds originate from. This location information may be complementary to the speaker embeddings in the diarisation processes. This report proposes to extend the Hidden Markov Model (HMM) clustering method, to enable the use of speaker location information. The HMM observation log-likelihood for the speaker location can take the form of a KL-divergence, when the speaker location is represented as a discrete posterior distribution of the probabilities that the sound originated from each possible location. Experimental results on a Microsoft rich meeting transcription task show that using speaker location information with the proposed HMM modification can yield performance improvements over using speaker embeddings alone. Jeremy H. M. Wong, Yifan Gong 0001 |
ICASSP | 3 |
| 2021 | Microsoft Speaker Diarization System for the Voxceleb Speaker Recognition Challenge 2020abstractThis paper describes the Microsoft speaker diarization system for monaural multi-talker recordings in the wild, evaluated at the diarization track of the VoxCeleb Speaker Recognition Challenge (VoxSRC) 2020. We will first explain our system design to address issues in handling real multi-talker recordings. We then present the details of the components, which include Res2Net-based speaker embedding extractor, conformer-based continuous speech separation with leakage filtering, and a modified DOVER (short for Diarization Output Voting Error Reduction) method for system fusion. We evaluate the systems with the data set provided by VoxSRC challenge 2020, which contains real-life multi-talker audio collected from YouTube. Our best system achieves 3.71% and 6.23% of the diarization error rate (DER) on development set and evaluation set, respectively, being ranked the 1st at the diarization track of the challenge. Naoyuki Kanda, Zhuo Chen 0006, Tianyan Zhou, Takuya Yoshioka, Sanyuan Chen, Yong Zhao 0008, Gang Liu 0001, Yu Wu 0012, Jian Wu 0027, Shujie Liu 0001, Jinyu Li 0001, Yifan Gong 0001 |
ICASSP | 13 |
| 2021 | Streaming Multi-Talker Speech Recognition with Joint Speaker IdentificationabstractIn multi-talker scenarios such as meetings and conversations, speech processing systems are usually required to transcribe the audio as well as identify the speakers for downstream applications. Since overlapped speech is common in this case, conventional approaches usually address this problem in a cascaded fashion that involves speech separation, speech recognition and speaker identification that are trained independently. In this paper, we propose Streaming Unmixing, Recognition and Identification Transducer (SURIT) -- a new framework that deals with this problem in an end-to-end streaming fashion. SURIT employs the recurrent neural network transducer (RNN-T) as the backbone for both speech recognition and speaker identification. We validate our idea on the LibrispeechMix dataset -- a multi-talker dataset derived from Librispeech, and present encouraging results. Liang Lu 0001, Naoyuki Kanda, Jinyu Li 0001, Yifan Gong 0001 |
Interspeech | 4 |
| 2021 | On Minimum Word Error Rate Training of the Hybrid Autoregressive TransducerabstractHybrid Autoregressive Transducer (HAT) is a recently proposed end-to-end acoustic model that extends the standard Recurrent Neural Network Transducer (RNN-T) for the purpose of the external language model (LM) fusion.In HAT, the blank probability and the label probability are estimated using two separate probability distributions, which provides a more accurate solution for internal LM score estimation, and thus works better when combining with an external LM.Previous work mainly focuses on HAT model training with the negative log-likelihood loss, while in this paper, we study the minimum word error rate (MWER) training of HAT -a criterion that is closer to the evaluation metric for speech recognition, and has been successfully applied to other types of end-to-end models such as sequenceto-sequence (S2S) and RNN-T models.From experiments with around 30,000 hours of training data, we show that MWER training can improve the accuracy of HAT models, while at the same time, improving the robustness of the model against the decoding hyper-parameters such as length normalization and decoding beam during inference. Liang Lu 0001, Zhong Meng, Naoyuki Kanda, Jinyu Li 0001, Yifan Gong 0001 |
Interspeech | 5 |
| 2021 | Rapid Speaker Adaptation for Conformer Transducer: Attention and Bias Are All You Need
Yan Huang 0028, Guoli Ye, Jinyu Li 0001, Yifan Gong 0001 |
Interspeech | 4 |
| 2021 | Improving RNN-T for Domain Scaling Using Semi-Supervised Training with Neural TTS
Rui Zhao 0017, Zhong Meng, Xie Chen 0001, Jinyu Li 0001, Yifan Gong 0001, Lei He 0005 |
Interspeech | 7 |
| 2021 | Multiple Softmax Architecture for Streaming Multilingual End-to-End ASR Systems
Vikas Joshi, Amit Das 0007, Eric Sun, Rupesh R. Mehta, Jinyu Li 0001, Yifan Gong 0001 |
Interspeech | 6 |
| 2021 | Minimum Word Error Rate Training with Language Model Fusion for End-to-End Speech RecognitionabstractIntegrating external language models (LMs) into end-to-end (E2E) models remains a challenging task for domain-adaptive speech recognition.Recently, internal language model estimation (ILME)-based LM fusion has shown significant word error rate (WER) reduction from Shallow Fusion by subtracting a weighted internal LM score from an interpolation of E2E model and external LM scores during beam search.However, on different test sets, the optimal LM interpolation weights vary over a wide range and have to be tuned extensively on well-matched validation sets.In this work, we perform LM fusion in the minimum WER (MWER) training of an E2E model to obviate the need for LM weights tuning during inference.Besides MWER training with Shallow Fusion (MWER-SF), we propose a novel MWER training with ILME (MWER-ILME) where the ILME-based fusion is conducted to generate N-best hypotheses and their posteriors.Additional gradient is induced when internal LM is engaged in MWER-ILME loss computation.During inference, LM weights pre-determined in MWER training enable robust LM integrations on test sets from different domains.Experimented with 30K-hour trained transformer transducers, MWER-ILME achieves on average 8.8% and 5.8% relative WER reductions from MWER and MWER-SF training, respectively, on 6 different test sets. Zhong Meng, Yu Wu 0012, Naoyuki Kanda, Liang Lu 0001, Xie Chen 0001, Guoli Ye, Eric Sun, Jinyu Li 0001, Yifan Gong 0001 |
Interspeech | 9 |
| 2021 | Improving Multilingual Transformer Transducer Models by Reducing Language Confusions
Eric Sun, Jinyu Li 0001, Zhong Meng, Yu Wu 0012, Shujie Liu 0001, Yifan Gong 0001 |
Interspeech | 7 |
| 2021 | Internal Language Model Estimation for Domain-Adaptive End-to-End Speech RecognitionabstractThe external language models (LM) integration remains a challenging task for end-to-end (E2E) automatic speech recognition (ASR) which has no clear division between acoustic and language models. In this work, we propose an internal LM estimation (ILME) method to facilitate a more effective integration of the external LM with all pre-existing E2E models with no additional model training, including the most popular recurrent neural network transducer (RNN-T) and attention-based encoder-decoder (AED) models. Trained with audio-transcript pairs, an E2E model implicitly learns an internal LM that characterizes the training data in the source domain. With ILME, the internal LM scores of an E2E model are estimated and subtracted from the log-linear interpolation between the scores of the E2E model and the external LM. The internal LM scores are approximated as the output of an E2E model when eliminating its acoustic components. ILME can alleviate the domain mismatch between training and testing, or improve the multi-domain E2E ASR. Experimented with 30K-hour trained RNN-T and AED models, ILME achieves up to 15.5% and 6.8% relative word error rate reductions from Shallow Fusion on out-of-domain LibriSpeech and in-domain Microsoft production test sets, respectively. Zhong Meng, Sarangarajan Parthasarathy, Eric Sun, Yashesh Gaur, Naoyuki Kanda, Liang Lu 0001, Xie Chen 0001, Rui Zhao 0017, Jinyu Li 0001, Yifan Gong 0001 |
SLT | 10 |
| 2021 | Streaming End-to-End Multi-Talker Speech RecognitionabstractEnd-to-end multi-talker speech recognition is an emerging research trend in the speech community due to its vast potential in applications such as conversation and meeting transcriptions. To the best of our knowledge, all existing research works are constrained in the offline scenario. In this work, we propose the Streaming Unmixing and Recognition Transducer (SURT) for end-to-end multi-talker speech recognition. Our model employs the Recurrent Neural Network Transducer (RNN-T) as the backbone that can meet various latency constraints. We study two different model architectures that are based on a speaker-differentiator encoder and a mask encoder respectively. To train this model, we investigate the widely used Permutation Invariant Training (PIT) approach and the Heuristic Error Assignment Training (HEAT) approach. Based on experiments on the publicly available LibriSpeechMix dataset, we show that HEAT can achieve better accuracy compared with PIT, and the SURT model with 150 milliseconds algorithmic latency constraint compares favorably with the offline sequence-to-sequence based baseline model in terms of accuracy. Liang Lu 0001, Naoyuki Kanda, Jinyu Li 0001, Yifan Gong 0001 |
IEEE Signal Process. Lett. | 4 |
| 2021 | Speaker Separation Using Speaker Inventories and Estimated SpeechabstractWe propose speaker separation using speaker inventories and estimated speech (SSUSIES), a framework leveraging speaker profiles and estimated speech for speaker separation. SSUSIES contains two methods, speaker separation using speaker inventories (SSUSI) and speaker separation using estimated speech (SSUES). SSUSI performs speaker separation with the help of speaker inventory. By combining the advantages of permutation invariant training (PIT) and speech extraction, SSUSI significantly outperforms conventional approaches. SSUES is a widely applicable technique that can substantially improve speaker separation performance using the output of first-pass separation. We evaluate the models on both speaker separation and speech recognition metrics. Zhuo Chen 0006, DeLiang Wang, Jinyu Li 0001, Yifan Gong 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2020 | Acoustic Model Adaptation for Presentation Transcription and Intelligent Meeting Assistant SystemsabstractWe present our solution for unsupervised rapid speaker adaptation in a state-of-art presentation and intelligent meeting transcription system. We adopt the Kullback-Leibler (KL) divergence regularized model adaptation paradigm. For the adaptation architecture, we found that the linear projection layer adaptation yields competitive performance with the additional benefit in its simplicity and robustness to small amount of adaptation data. To address the imperfect supervision, we use a supervision committee formed by multiple systems or single-system n-best to mask possibly mislabeled frames. To relieve the data sparsity issue, we apply noise and speaking rate perturbation data augmentation techniques to create a richer adaptation data set. In summary, the proposed solution consists of the KL-divergence regularized linear projection layer adaptation with frame masking and data augmentation. On a presentation transcription and a meeting transcription task, our proposed methodology yields 7.3 % and 7.9 % relative word error rate (WER) reduction against a strong baseline model trained from tens of thousand hour speech. To the best of our knowledge, this is a first reported work on rapid speaker adaptation on a state-of-art production system. Yan Huang 0028, Yifan Gong 0001 |
ICASSP | 2 |
| 2020 | Using Personalized Speech Synthesis and Neural Language Generator for Rapid Speaker AdaptationabstractWe propose to use the personalized speech synthesis and the neural language generator to synthesize content relevant personalized speech for rapid speaker adaptation. It has two distinct aspects: First, it relieves the general data sparsity issue in rapid adaptation via making use of additional synthesized personalized speech; Second, it circumvents the obstacle of the explicit labeling error in unsupervised adaptation by converting it to pseudo-supervised adaptation. In this setup, the labeling error is implicitly rendered as less damaging speech distortion in the personalized synthesized speech. This results in significant performance breakthrough in the rapid unsupervised speaker adaptation. We apply the proposed methodology to a speaker adaptation task in a state-of-art speech transcription system. With 1 minute (min) adaptation data, our proposed approach yields 9.19 % or 5.98 % relative word error rate (WER) reduction for the supervised and the unsupervised adaptation, comparing to the negligible gain when adapting only with 1 min original speech. With 10 min adaptation data, it yields 12.53 % or 7.89 % relative WER reduction, doubling the gain of the baseline adaptation. The proposed approach is particularly suitable for unsupervised adaptation. Yan Huang 0028, Lei He 0005, Wenning Wei, William Gale, Jinyu Li 0001, Yifan Gong 0001 |
ICASSP | 6 |
| 2020 | Exploring Pre-Training with Alignments for RNN Transducer Based End-to-End Speech RecognitionabstractRecently, the recurrent neural network transducer (RNN-T) architecture has become an emerging trend in end-to-end automatic speech recognition research due to its advantages of being capable for online streaming speech recognition. However, RNN-T training is made difficult by the huge memory requirements, and complicated neural structure. A common solution to ease the RNN-T training is to employ connectionist temporal classification (CTC) model along with RNN language model (RNNLM) to initialize the RNN-T parameters. In this work, we conversely leverage external alignments to seed the RNN-T model. Two different pre-training solutions are explored, referred to as encoder pre-training, and whole-network pre-training respectively. Evaluated on Microsoft 65,000 hours anonymized production data with personally identifiable information removed, our proposed methods can obtain significant improvement. In particular, the encoder pre-training solution achieved a 10% and a 8% relative word error rate reduction when compared with random initialization and the widely used CTC+RNNLM initialization strategy, respectively. Our solutions also significantly reduce the RNN-T model latency from the baseline. Hu Hu, Rui Zhao 0017, Jinyu Li 0001, Liang Lu 0001, Yifan Gong 0001 |
ICASSP | 5 |
| 2020 | Minimum Latency Training Strategies for Streaming Sequence-to-Sequence ASRabstractRecently, a few novel streaming attention-based sequence-to-sequence (S2S) models have been proposed to perform online speech recognition with linear-time decoding complexity. However, in these models, the decisions to generate tokens are delayed compared to the actual acoustic boundaries since their unidirectional encoders lack future information. This leads to an inevitable latency during inference. To alleviate this issue and reduce latency, we propose several strategies during training by leveraging external hard alignments extracted from the hybrid model. We investigate to utilize the alignments in both the encoder and the decoder. On the encoder side, (1) multi-task learning and (2) pre-training with the framewise classification task are studied. On the decoder side, we (3) remove inappropriate alignment paths beyond an acceptable latency during the alignment marginalization, and (4) directly min-imize the differentiable expected latency loss. Experiments on the Cortana voice search task demonstrate that our proposed methods can significantly reduce the latency, and even improve the recognition accuracy in certain cases on the decoder side. We also present some analysis to understand the behaviors of streaming S2S models. Hirofumi Inaguma, Yashesh Gaur, Liang Lu 0001, Jinyu Li 0001, Yifan Gong 0001 |
ICASSP | 5 |
| 2020 | High-Accuracy and Low-Latency Speech Recognition with Two-Head Contextual Layer Trajectory LSTM ModelabstractWhile the community keeps promoting end-to-end models over conventional hybrid models, which usually are long short-term memory (LSTM) models trained with a cross entropy criterion followed by a sequence discriminative training criterion, we argue that such conventional hybrid models can still be significantly improved. In this paper, we detail our recent efforts to improve conventional hybrid LSTM acoustic models for high-accuracy and low-latency automatic speech recognition. To achieve high accuracy, we use a contextual layer trajectory LSTM (cltLSTM), which decouples the temporal modeling and target classification tasks, and incorporates future context frames to get more information for accurate acoustic modeling. We further improve the training strategy with sequence-level teacher-student learning. To obtain low latency, we design a two-head cltLSTM, in which one head has zero latency and the other head has a small latency, compared to an LSTM. When trained with Microsoft's 65 thousand hours of anonymized training data and evaluated with test sets with 1.8 million words, the proposed two-head cltLSTM model with the proposed training strategy yields a 28.2% relative WER reduction over the conventional LSTM acoustic model, with a similar perceived latency. Jinyu Li 0001, Rui Zhao 0017, Eric Sun, Jeremy H. M. Wong, Amit Das 0007, Zhong Meng, Yifan Gong 0001 |
ICASSP | 7 |
| 2020 | L-Vector: Neural Label Embedding for Domain AdaptationabstractWe propose a novel neural label embedding (NLE) scheme for the domain adaptation of a deep neural network (DNN) acoustic model with unpaired data samples from source and target domains. With NLE method, we distill the knowledge from a powerful source-domain DNN into a dictionary of label embeddings, or l-vectors, one for each senone class. Each l-vector is a representation of the senone-specific output distributions of the source-domain DNN and is learned to minimize the average L2, Kullback-Leibler (KL) or symmetric KL distance to the output vectors with the same label through simple averaging or standard back-propagation. During adaptation, the l-vectors serve as the soft targets to train the target-domain model with cross-entropy loss. Without parallel data constraint as in the teacher-student learning, NLE is specially suited for the situation where the paired target-domain data cannot be simulated from the source-domain data. We adapt a 6400 hours multi-conditional US English acoustic model to each of the 9 accented English (80 to 830 hours) and kids' speech (80 hours). NLE achieves up to 14.1% relative word error rate reduction over direct re-training with one-hot labels. Zhong Meng, Hu Hu, Jinyu Li 0001, Changliang Liu, Yan Huang 0028, Yifan Gong 0001, Chin-Hui Lee 0001 |
ICASSP | 6 |
| 2020 | Adaptation of RNN Transducer with Text-To-Speech Technology for Keyword SpottingabstractWith the advent of recurrent neural network transducer (RNN-T) model, the performance of keyword spotting (KWS) systems has greatly improved. However, the KWS systems, employed for wake-word detection, still rely on the availability of keyword specific training data for achieving reasonable performance on each keyword. With a goal to improve the KWS performance for these keywords without having to collect additional natural speech data, we explore Text-To-Speech (TTS) technology to synthetically generate training data for such keywords. Employing an RNN-T based KWS model, already well trained on large keyword-independent natural speech dataset, as a seed model, we run adaptation experiments using the generated keyword-specific TTS data. Besides observing a considerable improvement in the overall performance for the low-resource keywords, we find that the performance improvement with TTS-generated training data, similar to natural speech data, depends on speaker diversity, amount of data per speaker and data simulation. We get additional improvement in performance by selectively adapting specific parts of the RNN-T model and gain key insights into different architectural constructs of RNN-T model. Eva Sharma, Guoli Ye, Wenning Wei, Rui Zhao 0017, Jian Wu 0027, Lei He 0005, Ed Lin, Yifan Gong 0001 |
ICASSP | 9 |
| 2020 | Rapid RNN-T Adaptation Using Personalized Speech Synthesis and Neural Language GeneratorabstractRapid unsupervised speaker adaptation in an E2E system posits us new challenges due to its end-to-end unified structure in addition to its intrinsic difficulty of data sparsity and imperfect label [1]. Previously we proposed utilizing the content relevant personalized speech synthesis for rapid speaker adaptation and achieved significant performance breakthrough in a hybrid system [2]. In this paper, we answer the following two questions: First, how to effectively perform rapid speaker adaptation in an RNN-T. Second, whether our previously proposed approach is still beneficial for the RNN-T and what are the modification and distinct observations. We apply the proposed methodology to a speaker adaptation task in a state-of-art presentation transcription RNN-T system. In the 1 min setup, it yields 11.58 % or 7.95 % relative word error rate (WER) reduction for the sup/unsup adaptation, comparing to the negligible gain when adapting with 1 min source speech. In the 10 min setup, it yields 15.71 % or 8.00 % relative WER reduction, doubling the gain of the source speech adaptation. We further apply various data filtering techniques and significantly bridge the gap between sup/unsup adaptation. Yan Huang 0028, Jinyu Li 0001, Lei He 0005, Wenning Wei, William Gale, Yifan Gong 0001 |
INTERSPEECH | 6 |
| 2020 | 1-D Row-Convolution LSTM: Fast Streaming ASR at Accuracy Parity with LC-BLSTM
Kshitiz Kumar, Chaojun Liu, Yifan Gong 0001, Jian Wu 0027 |
INTERSPEECH | 3 |
| 2020 | Bandpass Noise Generation and Augmentation for Unified ASR
Kshitiz Kumar, Yifan Gong 0001, Jian Wu 0027 |
INTERSPEECH | 3 |
| 2020 | Developing RNN-T Models Surpassing High-Performance Hybrid Models with Customization CapabilityabstractBecause of its streaming nature, recurrent neural network transducer (RNN-T) is a very promising end-to-end (E2E) model that may replace the popular hybrid model for automatic speech recognition.In this paper, we describe our recent development of RNN-T models with reduced GPU memory consumption during training, better initialization strategy, and advanced encoder modeling with future lookahead.When trained with Microsoft's 65 thousand hours of anonymized training data, the developed RNN-T model surpasses a very well trained hybrid model with both better recognition accuracy and lower latency.We further study how to customize RNN-T models to a new domain, which is important for deploying E2E models to practical scenarios.By comparing several methods leveraging text-only data in the new domain, we found that updating RNN-T's prediction and joint networks using text-to-speech generated from domain-specific text is the most effective. Jinyu Li 0001, Rui Zhao 0017, Zhong Meng, Wenning Wei, Sarangarajan Parthasarathy, Vadim Mazalov, Lei He 0005, Sheng Zhao 0002, Yifan Gong 0001 |
INTERSPEECH | 11 |
| 2020 | Exploring Transformers for Large-Scale Speech RecognitionabstractWhile recurrent neural networks still largely define state-of-theart speech recognition systems, the Transformer network has been proven to be a competitive alternative, especially in the offline condition.Most studies with Transformers have been constrained in a relatively small scale setting, and some forms of data argumentation approaches are usually applied to combat the data sparsity issue.In this paper, we aim at understanding the behaviors of Transformers in the large-scale speech recognition setting, where we have used around 65,000 hours of training data.We investigated various aspects on scaling up Transformers, including model initialization, warmup training as well as different Layer Normalization strategies.In the streaming condition, we compared the widely used attention mask based future context lookahead approach to the Transformer-XL network.From our experiments, we show that Transformers can achieve around 6% relative word error rate (WER) reduction compared to the BLSTM baseline in the offline fashion, while in the streaming fashion, Transformer-XL is comparable to LC-BLSTM with 800 millisecond latency constraint. Liang Lu 0001, Changliang Liu, Jinyu Li 0001, Yifan Gong 0001 |
INTERSPEECH | 4 |
| 2020 | Combination of End-to-End and Hybrid Models for Speech RecognitionabstractRecent studies suggest that it may now be possible to construct end-to-end Neural Network (NN) models that perform on-par with, or even outperform, hybrid models in speech recognition. These models differ in their designs, and as such, may exhibit diverse and complementary error patterns. A combination between the predictions of these models may therefore yield significant gains. This paper studies the feasibility of performing hypothesis-level combination between hybrid and end-to-end NN models. The end-to-end NN models often exhibit a bias in their posteriors toward short hypotheses, and this may adversely affect Minimum Bayes’ Risk (MBR) combination methods. MBR training and length normalisation can be used to reduce this bias. Models are trained on Microsoft’s 75 thousand hours of anonymised data and evaluated on test sets with 1.8 million words. The results show that significant gains can be obtained by combining the hypotheses of hybrid and end-to-end NN models together. Jeremy H. M. Wong, Yashesh Gaur, Rui Zhao 0017, Liang Lu 0001, Eric Sun, Jinyu Li 0001, Yifan Gong 0001 |
INTERSPEECH | 7 |
| 2019 | Improving RNN Transducer Modeling for End-to-End Speech RecognitionabstractIn the last few years, an emerging trend in automatic speech recognition research is the study of end-to-end (E2E) systems. Connectionist Temporal Classification (CTC), Attention Encoder-Decoder (AED), and RNN Transducer (RNN-T) are the most popular three methods. Among these three methods, RNN-T has the advantages to do online streaming which is challenging to AED and it doesn't have CTC's frame-independence assumption. In this paper, we improve the RNN-T training in two aspects. First, we optimize the training algorithm of RNN-T to reduce the memory consumption so that we can have larger training minibatch for faster training speed. Second, we propose better model structures so that we obtain RNN-T models with the very good accuracy but small footprint. Trained with 30 thousand hours anonymized and transcribed Microsoft production data, the best RNN-T model with even smaller model size (216 Megabytes) achieves up-to 11.8% relative word error rate (WER) reduction from the baseline RNN-T model. This best RNN-T model is significantly better than the device hybrid model with similar size by achieving up-to 15.0% relative WER reduction, and obtains similar WERs as the server hybrid model of 5120 Megabytes in size. Jinyu Li 0001, Rui Zhao 0017, Hu Hu, Yifan Gong 0001 |
ASRU | 4 |
| 2019 | Character-Aware Attention-Based End-to-End Speech RecognitionabstractPredicting words and subword units (WSUs) as the output has shown to be effective for the attention-based encoder-decoder (AED) model in end-to-end speech recognition. However, as one input to the decoder recurrent neural network (RNN), each WSU embedding is learned independently through context and acoustic information in a purely data-driven fashion. Little effort has been made to explicitly model the morphological relationships among WSUs. In this work, we propose a novel character-aware (CA) AED model in which each WSU embedding is computed by summarizing the embeddings of its constituent characters using a CA-RNN. This WSU-independent CA-RNN is jointly trained with the encoder, the decoder and the attention network of a conventional AED to predict WSUs. With CA-AED, the embeddings of morphologically similar WSUs are naturally and directly correlated through the CA-RNN in addition to the semantic and acoustic relations modeled by a traditional AED. Moreover, CA-AED significantly reduces the model parameters in a traditional AED by replacing the large pool of WSU embeddings with a much smaller set of character embeddings. On a 3400 hours Microsoft Cortana dataset, CA-AED achieves up to 11.9% relative WER improvement over a strong AED baseline with 27.1% fewer model parameters. Zhong Meng, Yashesh Gaur, Jinyu Li 0001, Yifan Gong 0001 |
ASRU | 4 |
| 2019 | Domain Adaptation via Teacher-Student Learning for End-to-End Speech RecognitionabstractTeacher-student (T/S) has shown to be effective for domain adaptation of deep neural network acoustic models in hybrid speech recognition systems. In this work, we extend the T/S learning to large-scale unsupervised domain adaptation of an attention-based end-to-end (E2E) model through two levels of knowledge transfer: teacher's token posteriors as soft labels and one-best predictions as decoder guidance. To further improve T/S learning with the help of ground-truth labels, we propose adaptive T/S (AT/S) learning. Instead of conditionally choosing from either the teacher's soft token posteriors or the one-hot ground-truth label, in AT/S, the student always learns from both the teacher and the ground truth with a pair of adaptive weights assigned to the soft and one-hot labels quantifying the confidence on each of the knowledge sources. The confidence scores are dynamically estimated at each decoder step as a function of the soft and one-hot labels. With 3400 hours parallel close-talk and far-field Microsoft Cortana data for domain adaptation, T/S and AT/S achieves 6.3% and 10.3% relative word error rate improvement over a strong E2E model trained with the same amount of far-field data. Zhong Meng, Jinyu Li 0001, Yashesh Gaur, Yifan Gong 0001 |
ASRU | 4 |
| 2019 | Advances in Online Audio-Visual Meeting TranscriptionabstractThis paper describes a system that generates speaker-annotated transcripts of meetings by using a microphone array and a 360-degree camera. The hallmark of the system is its ability to handle overlapped speech, which has been an unsolved problem in realistic settings for over a decade. We show that this problem can be addressed by using a continuous speech separation approach. In addition, we describe an online audio-visual speaker diarization method that leverages face tracking and identification, sound source localization, speaker identification, and, if available, prior speaker information for robustness to various real world challenges. All components are integrated in a meeting transcription framework called SRD, which stands for “separate, recognize, and diarize”. Experimental results using recordings of natural meetings involving up to 11 attendees are reported. The continuous speech separation improves a word error rate (WER) by 16.1% compared with a highly tuned beamformer. When a complete list of meeting attendees is available, the discrepancy between WER and speaker-attributed WER is only 1.0%, indicating accurate word-to-speaker association. This increases marginally to 1.6% when 50% of the attendees are unknown to the system. Takuya Yoshioka, Yan Huang 0028, Aviv Hurvitz, Sharon Koubi, Eyal Krupka, Ido Leichter, Changliang Liu, Partha Parthasarathy, Alon Vinnikov, Lingfeng Wu, Igor Abramovski, Wayne Xiong, Huaming Wang, Jun Zhang 0066, Yong Zhao 0008, Tianyan Zhou, Cem Aksoylar, Zhuo Chen 0006, Moshe David, Dimitrios Dimitriadis, Yifan Gong 0001, Ilya Gurvich, Xuedong Huang 0001 |
ASRU | 24 |
| 2019 | CNN with Phonetic Attention for Text-Independent Speaker VerificationabstractText-independent speaker verification imposes no constraints on the spoken content and usually needs long observations to make reliable prediction. In this paper, we propose two speaker embedding approaches by integrating the phonetic information into the attention-based residual convolutional neural network (CNN). Phonetic features are extracted from the bottleneck layer of a pretrained acoustic model. In implicit phonetic attention (IPA), the phonetic features are projected by a transformation network into multi-channel feature maps, and then combined with the raw acoustic features as the input of the CNN network. In explicit phonetic attention (EPA), the phonetic features are directly connected to the attentive pooling layer through a separate 1-dim CNN to generate the attention weights. With the incorporation of spoken content and attention mechanism, the system can not only distill the speaker-discriminant frames but also actively normalize the phonetic variations. Multi-head attention and discriminative objectives are further studied to improve the system. Experiments on the VoxCeleb corpus show our proposed system could outperform the state-of-the-art by around 43% relative. Tianyan Zhou, Yong Zhao 0008, Jinyu Li 0001, Yifan Gong 0001, Jian Wu 0027 |
ASRU | 4 |
| 2019 | Universal Acoustic Modeling Using Neural Mixture ModelsabstractAcoustic models are domain dependent and do not perform well if there is a mismatch between training and test conditions. As an alternative, the Mixture of Experts (MoE) model was introduced for multi-domain modeling. It combines the outputs of several domain specific models (or experts) using a gating network. However, one drawback is that the gating network directly uses raw features and is unaware of the state of the experts. In this work, we propose several alternatives to improve the MoE model. First, to make our MoE model state-aware, we use outputs of experts as inputs to the gating network. Then we show that vector based interpolation of the mixture weights is more effective than scalar interpolation. Second, we show that directly learning the mixture weights without using any complex gating is still effective. Finally, we introduce a hybrid attention model that uses the logits and mixture weights from the previous time step to generate the mixture weights at the current time. Our best proposed model outperforms a baseline model using LSTM based gating achieving about 20.48% relative reduction in word error rate (WER). Moreover, it beats an oracle model which picks the best expert for a given test condition. Amit Das 0007, Jinyu Li 0001, Changliang Liu, Yifan Gong 0001 |
ICASSP | 4 |
| 2019 | Word Characters and Phone Pronunciation Embedding for ASR Confidence ClassifierabstractConfidence classifier is an integral component of an automatic speech recognition (ASR) system. These classifiers predict the accuracy of an ASR hypothesis by associating a confidence score in [0,1] range, where larger score implies higher probability of the hypothesis being correct. Confidence scores have significant applications in ASR system design, training data selection, model adaptation, and other ASR applications. In this work we focus on word embedding features to improve confidence classifier, and introduce character and phone embeddings as confidence features. We motivate these features in the context of representing and factorizing acoustic scores along the proposed features. We evaluate our work on large scale ASR tasks, and demonstrate significant improvement in the confidence performance with the proposed features. At our typical operating point, we report 8% relative reduction in false alarm (FA) for limited vocabulary enUS Xbox task, and 9.9% relative reduction in FA for large vocabulary enUS server task. We also conducted server experiments for our proposed features in combination with natural language Glove embeddings, and improved the overall relative reduction in FA to 16%. Kshitiz Kumar, Tasos Anastasakos, Yifan Gong 0001 |
ICASSP | 3 |
| 2019 | Static and Dynamic State Predictions for Acoustic Model CombinationabstractAcoustic model combination (AMOC) is an active research area. Model combination techniques are critical for many automatic speech recognition (ASR) scenarios, and provide frameworks to combine diverse acoustic models to boost ASR performance. We scope this work in the broad framework of AMOC, and present static and dynamic state combinations of acoustic models. We motivate and rationalize the benefits from our combination techniques, and present many applications and extensions. We apply our work in the context of combining a generic and a scenario-specific (dedicated) acoustic model; we train the proposed model with an ASR objective to best align with ASR performance. We conduct our experiments on large-vocabulary ASR task with over 30k hours of training data. Compared to generic model, we demonstrate a strong 6% word error relative reduction (WERR) in average across a variety of tasks, and specifically 25% and 8% WERR for far-field speaker and an emerging car scenario. Kshitiz Kumar, Yifan Gong 0001 |
ICASSP | 2 |
| 2019 | Improving Layer Trajectory LSTM with Future Context FramesabstractIn our recent work, we proposed a layer trajectory long short-term memory (ltLSTM) model which decouples the tasks of temporal modeling and senone classification with time-LSTMs and depth-LSTMs. The ltLSTM model achieved significant accuracy improvement over the traditional multi-layer LSTM models from our previous study. Considering the future context frames carrying valuable information for predicting the target label evidenced by the success of bi-directional LSTMs, in this work we investigate how to incorporate this kind of information with hidden vectors from either time-LSTM or depth-LSTM. Trained with 30 thousand hours of EN-US Microsoft internal data, the best ltLSTM model with future context frames can improve the baseline ltLSTM with up to 11.5% relative word error rate (WER) reduction and improve the baseline LSTM with up to 24.6% relative WER reduction across different tasks. Jinyu Li 0001, Liang Lu 0001, Changliang Liu, Yifan Gong 0001 |
ICASSP | 4 |
| 2019 | Towards Code-switching ASR for End-to-end CTC ModelsabstractAlthough great progress has been made on end-to-end (E2E) models for monolingual and multilingual automatic speech recognition (ASR), there is no successful study for E2E models on the challenging intra-sentential code-switching (CS) ASR task to our best knowledge. In this paper, we propose an approach for CS ASR using E2E connectionist temporal classification (CTC) models. We use a frame-level language identification model to linearly adjust the posteriors of an E2E CTC model. We evaluate the proposed method on Microsoft live Chinese Cortana data with 7000 hours Chinese and English monolingual data and 300 hours CS data as the training data. Trained with only monolingual data without observing any CS data, the proposed method can obtain up to 6.3% relative word error rate (WER) reduction. In the scenario of training with both monolingual and CS data, the proposed method can get up to 4.2% relative WER improvement. This approach can also maintain comparable performance on a Chinese test set compared with baseline models. Jinyu Li 0001, Guoli Ye, Rui Zhao 0017, Yifan Gong 0001 |
ICASSP | 5 |
| 2019 | Adversarial Speaker AdaptationabstractWe propose a novel adversarial speaker adaptation (ASA) scheme, in which adversarial learning is applied to regularize the distribution of deep hidden features in a speaker-dependent (SD) deep neural network (DNN) acoustic model to be close to that of a fixed speaker-independent (SI) DNN acoustic model during adaptation. An additional discriminator network is introduced to distinguish the deep features generated by the SD model from those produced by the SI model. In ASA, with a fixed SI model as the reference, an SD model is jointly optimized with the discriminator network to minimize the senone classification loss, and simultaneously to mini-maximize the SI/SD discrimination loss on the adaptation data. With ASA, a senone-discriminative deep feature is learned in the SD model with a similar distribution to that of the SI model. With such a regularized and adapted deep feature, the SD model can perform improved automatic speech recognition on the target speaker's speech. Evaluated on the Microsoft short message dictation dataset, ASA achieves 14.4% and 7.9% relative word error rate improvements for supervised and unsupervised adaptation, respectively, over an SI model trained from 2600 hours data, with 200 adaptation utterances per speaker. Zhong Meng, Jinyu Li 0001, Yifan Gong 0001 |
ICASSP | 3 |
| 2019 | Attentive Adversarial Learning for Domain-invariant TrainingabstractAdversarial domain-invariant training (ADIT) proves to be effective in suppressing the effects of domain variability in acoustic modeling and has led to improved performance in automatic speech recognition (ASR). In ADIT, an auxiliary domain classifier takes in equally-weighted deep features from a deep neural network (DNN) acoustic model and is trained to improve their domain-invariance by optimizing an adversarial loss function. In this work, we propose an attentive ADIT (AADIT) in which we advance the domain classifier with an attention mechanism to automatically weight the input deep features according to their importance in domain classification. With this attentive re-weighting, ADDIT can focus on the domain normalization of phonetic components that are more susceptible to domain variability and generates deep features with improved domain-invariance and senone-discriminativity over ADIT. Most importantly, the attention block serves only as an external component to the DNN acoustic model and is not involved in ASR, so AADIT can be used to improve the acoustic modeling with any DNN architectures. More generally, the same methodology can improve any adversarial learning system with an auxiliary discriminator. Evaluated on CHiME-3 dataset, the AADIT achieves 13.6% and 9.3% relative WER improvements, respectively, over a multi-conditional model and a strong ADIT baseline. Zhong Meng, Jinyu Li 0001, Yifan Gong 0001 |
ICASSP | 3 |
| 2019 | Conditional Teacher-student LearningabstractThe teacher-student (T/S) learning has been shown to be effective for a variety of problems such as domain adaptation and model compression. One shortcoming of the T/S learning is that a teacher model, not always perfect, sporadically produces wrong guidance in form of posterior probabilities that misleads the student model towards a suboptimal performance. To overcome this problem, we propose a conditional T/S learning scheme, in which a "smart" student model selectively chooses to learn from either the teacher model or the ground truth labels conditioned on whether the teacher can correctly predict the ground truth. Unlike a naive linear combination of the two knowledge sources, the conditional learning is exclusively engaged with the teacher model when the teacher model's prediction is correct, and otherwise backs off to the ground truth. Thus, the student model is able to learn effectively from the teacher and even potentially surpass the teacher. We examine the proposed learning scheme on two tasks: domain adaptation on CHiME-3 dataset and speaker adaptation on Microsoft short message dictation dataset. The proposed method achieves 9.8% and 12.8% relative word error rate reductions, respectively, over T/S learning for environment adaptation and speaker-independent model for speaker adaptation. Zhong Meng, Jinyu Li 0001, Yong Zhao 0008, Yifan Gong 0001 |
ICASSP | 4 |
| 2019 | Adversarial Speaker VerificationabstractThe use of deep networks to extract embeddings for speaker recognition has proven successfully. However, such embeddings are susceptible to performance degradation due to the mismatches among the training, enrollment, and test conditions. In this work, we propose an adversarial speaker verification (ASV) scheme to learn the condition-invariant deep embedding via adversarial multi-task training. In ASV, a speaker classification network and a condition identification network are jointly optimized to minimize the speaker classification loss and simultaneously mini-maximize the condition loss. The target labels of the condition network can be categorical (environment types) and continuous (SNR values). We further propose multi-factorial ASV to simultaneously suppress multiple factors that constitute the condition variability. Evaluated on a Microsoft Cortana text-dependent speaker verification task, the ASV achieves 8.8% and 14.5% relative improvements in equal error rates (EER) for known and unknown conditions, respectively. Zhong Meng, Yong Zhao 0008, Jinyu Li 0001, Yifan Gong 0001 |
ICASSP | 4 |
| 2019 | Single-channel Speech Extraction Using Speaker Inventory and Attention NetworkabstractNeural network-based speech separation has received a surge of interest in recent years. Previously proposed methods either are speaker independent or extract a target speaker's voice by using his or her voice snippet. In applications such as home devices or office meeting transcriptions, a possible speaker list is available, which can be leveraged for speech separation. This paper proposes a novel speech extraction method that utilizes an inventory of voice snippets of possible interfering speakers, or speaker enrollment data, in addition to that of the target speaker. Furthermore, an attention-based network architecture is proposed to form time-varying masks for both the target and other speakers during the separation process. This architecture does not reduce the enrollment audio of each speaker into a single vector, thereby allowing each short time frame of the input mixture signal to be aligned and accurately compared with the enrollment signals. We evaluate the proposed system on a speaker extraction task derived from the Libri corpus and show the effectiveness of the method. Zhuo Chen 0006, Takuya Yoshioka, Hakan Erdogan, Changliang Liu, Dimitrios Dimitriadis, Jasha Droppo, Yifan Gong 0001 |
ICASSP | 8 |
| 2019 | Encrypted Speech Recognition Using Deep Polynomial NetworksabstractThe cloud-based speech recognition/API provides developers or enterprises an easy way to create speech-enabled features in their applications. However, sending audios about personal or company internal information to the cloud, raises concerns about the privacy and security issues. The recognition results generated in cloud may also reveal some sensitive information. This paper proposes a deep polynomial network (DPN) that can be applied to the encrypted speech as an acoustic model. It allows clients to send their data in an encrypted form to the cloud to ensure that their data remains confidential, at mean while the DPN can still make frame-level predictions over the encrypted speech and return them in encrypted form. One good property of the DPN is that it can be trained on unencrypted speech features in the traditional way. To keep the cloud away from the raw audio and recognition results, a cloud-local joint decoding framework is also proposed. We demonstrate the effectiveness of model and framework on the Switchboard and Cortana voice assistant tasks with small performance degradation and latency increased comparing with the traditional cloud-based DNNs. Shixiong Zhang 0001, Yifan Gong 0001, Dong Yu 0001 |
ICASSP | 2 |
| 2019 | Acoustic-to-Phrase Models for Speech RecognitionabstractDirectly emitting words and sub-words from speech spectrogram has been shown to produce good results using end-to-end (E2E) trained models. Connectionist Temporal Classification (CTC) and Sequence-to-Sequence attention (Seq2Seq) models have both shown better success when directly targeting words or sub-words. In this work, we ask the question: Can an E2E model go beyond words and transcribe directly to phrases (i.e., a group of words)? Directly modeling frequent phrases might be better than modeling its constituent words. Also, emitting multiple words together might speed up inference in models like Seq2Seq where decoding is inherently sequential. To answer this, we undertake a study on a 3400-hour Microsoft Cortana voice assistant task. We present a side-by-side comparison for CTC and Seq2Seq models that have been trained to target a variety of tokens including letters, sub-words, words and phrases. We show that an E2E model can indeed transcribe directly to phrases. We see that while CTC has difficulty in accurately modeling phrases, a more powerful model like Seq2Seq can effortlessly target phrases that are up to 4 words long, with only a reasonable degradation in the final word error rate. Yashesh Gaur, Jinyu Li 0001, Zhong Meng, Yifan Gong 0001 |
INTERSPEECH | 4 |
| 2019 | Self-Teaching NetworksabstractWe propose self-teaching networks to improve the generalization capacity of deep neural networks. The idea is to generate soft supervision labels using the output layer for training the lower layers of the network. During the network training, we seek an auxiliary loss that drives the lower layer to mimic the behavior of the output layer. The connection between the two network layers through the auxiliary loss can help the gradient flow, which works similar to the residual networks. Furthermore, the auxiliary loss also works as a regularizer, which improves the generalization capacity of the network. We evaluated the self-teaching network with deep recurrent neural networks on speech recognition tasks, where we trained the acoustic model using 30 thousand hours of data. We tested the acoustic model using data collected from 4 scenarios. We show that the self-teaching network can achieve consistent improvements and outperform existing methods such as label smoothing and confidence penalization. Liang Lu 0001, Eric Sun, Yifan Gong 0001 |
INTERSPEECH | 3 |
| 2019 | Speaker Adaptation for Attention-Based End-to-End Speech RecognitionabstractWe propose three regularization-based speaker adaptation approaches to adapt the attention-based encoder-decoder (AED) model with very limited adaptation data from target speakers for end-to-end automatic speech recognition. The first method is Kullback-Leibler divergence (KLD) regularization, in which the output distribution of a speaker-dependent (SD) AED is forced to be close to that of the speaker-independent (SI) model by adding a KLD regularization to the adaptation criterion. To compensate for the asymmetric deficiency in KLD regularization, an adversarial speaker adaptation (ASA) method is proposed to regularize the deep-feature distribution of the SD AED through the adversarial learning of an auxiliary discriminator and the SD AED. The third approach is the multi-task learning, in which an SD AED is trained to jointly perform the primary task of predicting a large number of output units and an auxiliary task of predicting a small number of output units to alleviate the target sparsity issue. Evaluated on a Microsoft short message dictation task, all three methods are highly effective in adapting the AED model, achieving up to 12.2% and 3.0% word error rate improvement over an SI AED trained from 3400 hours data for supervised and unsupervised adaptation, respectively. Zhong Meng, Yashesh Gaur, Jinyu Li 0001, Yifan Gong 0001 |
INTERSPEECH | 4 |
| 2019 | Layer Trajectory BLSTM
Eric Sun, Jinyu Li 0001, Yifan Gong 0001 |
INTERSPEECH | 3 |
| 2019 | Advancing Acoustic-to-Word CTC Model With Attention and Mixed-UnitsabstractThe acoustic-to-word model based on the Connectionist Temporal Classification (CTC) criterion is a natural end-to-end (E2E) system directly targeting word as output unit. Two issues exist in the system: first, the current output of the CTC model relies on the current input and does not account for context weighted inputs. This is the hard alignment issue. Second, the word-based CTC model suffers from the out-of-vocabulary (OOV) issue. This means it can model only frequently occurring words while tagging the remaining words as OOV. Hence, such a model is limited in its capacity in recognizing only a fixed set of frequent words. In this study, we propose addressing these problems using a combination of attention mechanism and mixed-units. In particular, we introduce Attention CTC, Self-Attention CTC, Hybrid CTC, and Mixed-unit CTC. First, we blend attention modeling capabilities directly into the CTC network using Attention CTC and Self-Attention CTC. Second, to alleviate the OOV issue, we present Hybrid CTC which uses a word and letter CTC with shared hidden layers. The Hybrid CTC consults the letter CTC when the word CTC emits an OOV. Then, we propose a much better solution by training a Mixed-unit CTC which decomposes all the OOV words into sequences of frequent words and multi-letter units. Evaluated on a 3400 hours Microsoft Cortana voice assistant task, our final acoustic-to-word solution using attention and mixed-units achieves a relative reduction in word error rate (WER) over the vanilla word CTC by 12.09%. Such an E2E model without using any language model (LM) or complex decoder also outperforms a traditional context-dependent (CD) phoneme CTC with strong LM and decoder by 6.79% relative. Amit Das 0007, Jinyu Li 0001, Guoli Ye, Rui Zhao 0017, Yifan Gong 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2018 | Efficient Integration of Fixed Beamformers and Speech Separation Networks for Multi-Channel Far-Field Speech SeparationabstractSpeech separation research has significantly progressed in recent years thanks to the rapid advances in deep learning technology. However the performance of recently proposed single-channel neural network-based speech separation methods is still limited especially in reverberant environments. To push the performance limit, we recently developed a method of integrating beamforming and single-channel speech separation approaches. This paper proposes a novel architecture that integrates multi -channel beamforming and speech separation in a much more efficient way than our previous method. The proposed architecture comprises a set of fixed beamformers, a beam prediction network, and a speech separation network based on permutation invariant training (PIT). The beam prediction network takes in the beamformed audio signals and estimates the best beam for each speaker constituting the input mixture. Two variants of PIT-based speech separation networks are proposed. Our approach is evaluated on reverberant speech mixtures under three different mixing conditions, covering cases where speakers partially overlap or one speaker's utterance is very short. The experimental results show that the proposed system significantly outperforms the conventional single-channel PIT system, producing the same performance as a single-channel system using oracle masks. Zhuo Chen 0006, Takuya Yoshioka, Linyu Li 0007, Michael L. Seltzer, Yifan Gong 0001 |
ICASSP | 6 |
| 2018 | Advancing Connectionist Temporal Classification with Attention ModelingabstractIn this study, we propose advancing all-neural speech recognition by directly incorporating attention modeling within the Connectionist Temporal Classification (CTC) framework. In particular, we derive new context vectors using time convolution features to model attention as part of the CTC network. To further improve attention modeling, we utilize content information extracted from a network representing an implicit language model. Finally, we introduce vector based attention weights that are applied on context vectors across both time and their individual components. We evaluate our system on a 3400 hours Microsoft Cortana voice assistant task and demonstrate that our proposed model consistently outperforms the baseline model achieving about 20% relative reduction in word error rates. Amit Das 0007, Jinyu Li 0001, Rui Zhao 0017, Yifan Gong 0001 |
ICASSP | 4 |
| 2018 | Advancing Acoustic-to-Word CTC ModelabstractThe acoustic-to-word model based on the connectionist temporal classification (CTC) criterion was shown as a natural end-to-end (E2E) model directly targeting words as output units. However, the word-based CTC model suffers from the out-of-vocabulary (OOV) issue as it can only model limited number of words in the output layer and maps all the remaining words into an OOV output node. Hence, such a word-based CTC model can only recognize the frequent words modeled by the network output nodes. Our first attempt to improve the acoustic-to-word model is a hybrid CTC model which consults a letter-based CTC when the word-based CTC model emits OOV tokens during testing time. Then, we propose a much better solution by training a mixed-unit CTC model which decomposes all the OOV words into sequences of frequent words and multi-letter units. Evaluated on a 3400 hours Microsoft Cortana voice assistant task, the final acoustic-to-word solution improves the baseline word-based CTC by relative 12.09% word error rate (WER) reduction when combined with our proposed attention CTC. Such an E2E model without using any language model (LM) or complex decoder outperforms the traditional context-dependent phoneme CTC which has strong LM and decoder by relative 6.79%. Jinyu Li 0001, Guoli Ye, Amit Das 0007, Rui Zhao 0017, Yifan Gong 0001 |
ICASSP | 5 |
| 2018 | Developing Far-Field Speaker System Via Teacher-Student LearningabstractIn this study, we develop the keyword spotting (KWS) and acoustic model (AM) components in a far-field speaker system. Specifically, we use teacher-student (T/S) learning to adapt a close-talk well-trained production AM to far-field by using parallel close-talk and simulated far-field data. We also use T/S learning to compress a large-size KWS model into a small-size one to fit the device computational cost. Without the need of transcription, T/S learning well utilizes untranscribed data to boost the model performance in both the AM adaptation and KWS model compression. We further optimize the models with sequence discriminative training and live data to reach the best performance of systems. The adapted AM improved from the baseline by 72.60% and 57.16% relative word error rate reduction on play-back and live test data, respectively. The final KWS model size was reduced by 27 times from a large-size KWS model without losing accuracy. Jinyu Li 0001, Rui Zhao 0017, Zhuo Chen 0006, Changliang Liu, Guoli Ye, Yifan Gong 0001 |
ICASSP | 7 |
| 2018 | Speaker-Invariant Training Via Adversarial LearningabstractWe propose a novel adversarial multi-task learning scheme, aiming at actively curtailing the inter-talker feature variability while maximizing its senone discriminability so as to enhance the performance of a deep neural network (DNN) based ASR system. We call the scheme speaker-invariant training (SIT). In SIT, a DNN acoustic model and a speaker classifier network are jointly optimized to minimize the senone (tied triphone state) classification loss, and simultaneously mini-maximize the speaker classification loss. A speaker-invariant and senone-discriminative deep feature is learned through this adversarial multi-task learning. With SIT, a canonical DNN acoustic model with significantly reduced variance in its output probabilities is learned with no explicit speaker-independent (SI) transformations or speaker-specific representations used in training or testing. Evaluated on the CHiME-3 dataset, the SIT achieves 4.99% relative word error rate (WER) improvement over the conventional SI acoustic model. With additional unsupervised speaker adaptation, the speaker-adapted (SA) SIT model achieves 4.86% relative WER gain over the SA SI acoustic model. Zhong Meng, Jinyu Li 0001, Zhuo Chen 0006, Yang Zhao 0002, Vadim Mazalov, Yifan Gong 0001, Biing-Hwang Juang |
ICASSP | 6 |
| 2018 | Adversarial Teacher-Student Learning for Unsupervised Domain AdaptationabstractThe teacher-student (T/S) learning has been shown effective in unsupervised domain adaptation [1]. It is a form of transfer learning, not in terms of the transfer of recognition decisions, but the knowledge of posteriori probabilities in the source domain as evaluated by the teacher model. It learns to handle the speaker and environment variability inherent in and restricted to the speech signal in the target domain without proactively addressing the robustness to other likely conditions. Performance degradation may thus ensue. In this work, we advance T/S learning by proposing adversarial T/S learning to explicitly achieve condition-robust unsupervised domain adaptation. In this method, a student acoustic model and a condition classifier are jointly optimized to minimize the Kullback-Leibler divergence between the output distributions of the teacher and student models, and simultaneously, to min-maximize the condition classification loss. A condition-invariant deep feature is learned in the adapted student model through this procedure. We further propose multi-factorial adversarial T/S learning which suppresses condition variabilities caused by multiple factors simultaneously. Evaluated with the noisy CHiME-3 test set, the proposed methods achieve relative word error rate improvements of 44.60% and 5.38%, respectively, over a clean source model and a strong T/S learning baseline model. Zhong Meng, Jinyu Li 0001, Yifan Gong 0001, Biing-Hwang Juang |
ICASSP | 3 |
| 2018 | Domain and Speaker Adaptation for Cortana Speech RecognitionabstractVoice assistant represents one of the most popular and important scenarios for speech recognition. In this paper, we propose two adaptation approaches to customize a multi-style well-trained acoustic model towards its subsidiary domain of Cortana assistant. First, we present anchor-based speaker adaptation by extracting the speaker information, i-vector or d-vector embeddings, from the anchor segments of ‘Hey Cortana’. The anchor embeddings are mapped to layer-wise parameters to control the transformations of both weight matrices and biases of multiple layers. Second, we directly update the existing model parameters for domain adaptation. We demonstrate that prior distribution should be updated along with the network adaptation to compensate the label bias from the development data. Updating the priors may have a significant impact when the target domain features high occurrence of anchor words. Experiments on Hey Cortana desktop test set show that both approaches improve the recognition accuracy significantly. The anchor-based adaptation using the anchor d-vector and the prior interpolation achieves 32% relative reduction in WER over the generic model. Yong Zhao 0008, Jinyu Li 0001, Shixiong Zhang 0001, Yifan Gong 0001 |
ICASSP | 5 |
| 2018 | Layer Trajectory LSTMabstractIt is popular to stack LSTM layers to get better modeling power, especially when large amount of training data is available. However, an LSTM-RNN with too many vanilla LSTM layers is very hard to train and there still exists the gradient vanishing issue if the network goes too deep. This issue can be partially solved by adding skip connections between layers, such as residual LSTM. In this paper, we propose a layer trajectory LSTM (ltLSTM) which builds a layer-LSTM using all the layer outputs from a standard multi-layer time-LSTM. This layer-LSTM scans the outputs from time-LSTMs, and uses the summarized layer trajectory information for final senone classification. The forward-propagation of time-LSTM and layer-LSTM can be handled in two separate threads in parallel so that the network computation time is the same as the standard time-LSTM. With a layer-LSTM running through layers, a gated path is provided from the output layer to the bottom layer, alleviating the gradient vanishing issue. Trained with 30 thousand hours of EN-US Microsoft internal data, the proposed ltLSTM performed significantly better than the standard multi-layer LSTM and residual LSTM, with up to 9.0% relative word error rate reduction across different tasks. Jinyu Li 0001, Changliang Liu, Yifan Gong 0001 |
INTERSPEECH | 3 |
| 2018 | Cycle-Consistent Speech EnhancementabstractFeature mapping using deep neural networks is an effective approach for single-channel speech enhancement. Noisy features are transformed to the enhanced ones through a mapping network and the mean square errors between the enhanced and clean features are minimized. In this paper, we propose a cycle-consistent speech enhancement (CSE) in which an additional inverse mapping network is introduced to reconstruct the noisy features from the enhanced ones. A cycle-consistent constraint is enforced to minimize the reconstruction loss. Similarly, a backward cycle of mappings is performed in the opposite direction with the same networks and losses. With cycle-consistency, the speech structure is well preserved in the enhanced features while noise is effectively reduced such that the feature-mapping network generalizes better to unseen data. In cases where only unparalleled noisy and clean data is available for training, two discriminator networks are used to distinguish the enhanced and noised features from the clean and noisy ones. The discrimination losses are jointly optimized with reconstruction losses through adversarial multi-task learning. Evaluated on the CHiME-3 dataset, the proposed CSE achieves 19.60% and 6.69% relative word error rate improvements respectively when using or without using parallel clean and noisy speech data. Zhong Meng, Jinyu Li 0001, Yifan Gong 0001, Biing-Hwang Juang |
INTERSPEECH | 3 |
| 2018 | Adversarial Feature-Mapping for Speech EnhancementabstractFeature-mapping with deep neural networks is commonly used for single-channel speech enhancement, in which a feature-mapping network directly transforms the noisy features to the corresponding enhanced ones and is trained to minimize the mean square errors between the enhanced and clean features. In this paper, we propose an adversarial feature-mapping (AFM) method for speech enhancement which advances the feature-mapping approach with adversarial learning. An additional discriminator network is introduced to distinguish the enhanced features from the real clean ones. The two networks are jointly optimized to minimize the feature-mapping loss and simultaneously mini-maximize the discrimination loss. The distribution of the enhanced features is further pushed towards that of the clean features through this adversarial multi-task training. To achieve better performance on ASR task, senone-aware (SA) AFM is further proposed in which an acoustic model network is jointly trained with the feature-mapping and discriminator networks to optimize the senone classification loss in addition to the AFM losses. Evaluated on the CHiME-3 dataset, the proposed AFM achieves 16.95% and 5.27% relative word error rate (WER) improvements over the real noisy data and the feature-mapping baseline respectively and the SA-AFM achieves 9.85% relative WER improvement over the multi-conditional acoustic model. Zhong Meng, Jinyu Li 0001, Yifan Gong 0001, Biing-Hwang Juang |
INTERSPEECH | 3 |
| 2018 | Multi-Channel Overlapped Speech Recognition with Location Guided Speech Extraction NetworkabstractAlthough advances in close-talk speech recognition have resulted in relatively low error rates, the recognition performance in far-field environments is still limited due to low signal-to-noise ratio, reverberation, and overlapped speech from simultaneous speakers which is especially more difficult. To solve these problems, beamforming and speech separation networks were previously proposed. However, they tend to suffer from leakage of interfering speech or limited generalizability. In this work, we propose a simple yet effective method for multi-channel far-field overlapped speech recognition. In the proposed system, three different features are formed for each target speaker, namely, spectral, spatial, and angle features. Then a neural network is trained using all features with a target of the clean speech of the required speaker. An iterative update procedure is proposed in which the mask-based beamforming and mask estimation are performed alternatively. The proposed system were evaluated with real recorded meetings with different levels of overlapping ratios. The results show that the proposed system achieves more than 24% relative word error rate (WER) reduction than fixed beamforming with oracle selection. Moreover, as overlap ratio rises from 20% to 70+%, only 3.8% WER increase is observed for the proposed system. Zhuo Chen 0006, Takuya Yoshioka, Hakan Erdogan, Jinyu Li 0001, Yifan Gong 0001 |
SLT | 6 |
| 2018 | Exploring Layer Trajectory LSTM with Depth Processing Units and AttentionabstractTraditional LSTM model and its variants normally work in a frame-by-frame and layer-by-layer fashion, which deals with the temporal modeling and target classification problems at the same time. In this paper, we extend our recently proposed layer trajectory LSTM (ltLSTM) and present a generalized framework, which is equipped with a depth processing block that scans the hidden states of each time-LSTM layer, and uses the summarized layer trajectory information for final senone classification. We explore different modeling units used in the depth processing block to have a good tradeoff between accuracy and runtime cost. Furthermore, we integrate an attention module into this framework to explore wide context information, which is especially beneficial for uni-directional LSTMs. Trained with 30 thousand hours of EN-US Microsoft internal data and cross entropy criterion, the proposed generalized ltLSTM performed significantly better than the standard multi-layer time-LSTM, with up to 12.8% relative word error rate (WER) reduction across different tasks. With attention modeling, the relative WER reduction can be up to 17.9%. We observed similar gain when the models were trained with sequence discriminative training criterion. Jinyu Li 0001, Liang Lu 0001, Changliang Liu, Yifan Gong 0001 |
SLT | 4 |
| 2018 | Speaker Adaptation for End-to-End CTC ModelsabstractWe propose two approaches for speaker adaptation in end-to-end (E2E) automatic speech recognition systems. One is Kullback-Leibler divergence (KLD) regularization and the other is multi-task learning (MTL). Both approaches aim to address the data sparsity especially output target sparsity issue of speaker adaptation in E2E systems. The KLD regularization adapts a model by forcing the output distribution from the adapted model to be close to the unadapted one. The MTL utilizes a jointly trained auxiliary task to improve the performance of the main task. We investigated our approaches on E2E connectionist temporal classification (CTC) models with three different types of output units. Experiments on the Microsoft short message dictation task demonstrated that MTL outperforms KLD regularization. In particular, the MTL adaptation obtained 8.8% and 4.0% relative word error rate reductions (WERRs) for supervised and unsupervised adaptations for the word CTC model, and 9.6% and 3.8% relative WERRs for the mix-unit CTC model, respectively. Jinyu Li 0001, Yong Zhao 0008, Kshitiz Kumar, Yifan Gong 0001 |
SLT | 5 |
| 2017 | Cracking the cocktail party problem by multi-beam deep attractor networkabstractWhile recent progresses in neural network approaches to singlechannel speech separation, or more generally the cocktail party problem, achieved significant improvement, their performance for complex mixtures is still not satisfactory. In this work, we propose a novel multi-channel framework for multi-talker separation. In the proposed model, an input multi-channel mixture signal is firstly converted to a set of beamformed signals using fixed beam patterns. For this beamforming, we propose to use differential beamformers as they are more suitable for speech separation. Then each beamformed signal is fed into a single-channel anchored deep attractor network to generate separated signals. And the final separation is acquired by post selecting the separating output for each beams. To evaluate the proposed system, we create a challenging dataset comprising mixtures of 2, 3 or 4 speakers. Our results show that the proposed system largely improves the state of the art in speech separation, achieving 11.5 dB, 11.76 dB and 11.02 dB average signal-to-distortion ratio improvement for 4, 3 and 2 overlapped speaker mixtures, which is comparable to the performance of a minimum variance distortionless response beamformer that uses oracle location, source, and noise information. We also run speech recognition with a clean trained acoustic model on the separated speech, achieving relative word error rate (WER) reduction of 45.76%, 59.40% and 62.80% on fully overlapped speech of 4, 3 and 2 speakers, respectively. With a far talk acoustic model, the WER is further reduced. Zhuo Chen 0006, Jinyu Li 0001, Takuya Yoshioka, Huaming Wang, Yifan Gong 0001 |
ASRU | 7 |
| 2017 | Acoustic-to-word model without OOVabstractRecently, the acoustic-to-word model based on the Connectionist Temporal Classification (CTC) criterion was shown as a natural end-to-end model directly targeting words as output units. However, this type of word-based CTC model suffers from the out-of-vocabulary (OOV) issue as it can only model limited number of words in the output layer and maps all the remaining words into an OOV output node. Therefore, such word-based CTC model can only recognize the frequent words modeled by the network output nodes. It also cannot easily handle the hot-words which emerge after the model is trained. In this study, we improve the acoustic-to-word model with a hybrid CTC model which can predict both words and characters at the same time. With a shared-hidden-layer structure and modular design, the alignments of words generated from the word-based CTC and the character-based CTC are synchronized. Whenever the acoustic-to-word model emits an OOV token, we back off that OOV segment to the word output generated from the character-based CTC, hence solving the OOV or hot-words issue. Evaluated on a Microsoft Cortana voice assistant task, the proposed model can reduce the errors introduced by the OOV output token in the acoustic-to-word model by 30%. Jinyu Li 0001, Guoli Ye, Rui Zhao 0017, Jasha Droppo, Yifan Gong 0001 |
ASRU | 5 |
| 2017 | Unsupervised adaptation with domain separation networks for robust speech recognitionabstractUnsupervised domain adaptation of speech signal aims at adapting a well-trained source-domain acoustic model to the unlabeled data from target domain. This can be achieved by adversarial training of deep neural network (DNN) acoustic models to learn an intermediate deep representation that is both senone-discriminative and domain-invariant. Specifically, the DNN is trained to jointly optimize the primary task of senone classification and the secondary task of domain classification with adversarial objective functions. In this work, instead of only focusing on learning a domain-invariant feature (i.e. the shared component between domains), we also characterize the difference between the source and target domain distributions by explicitly modeling the private component of each domain through a private component extractor DNN. The private component is trained to be orthogonal with the shared component and thus implicitly increases the degree of domain-invariance of the shared component. A reconstructor DNN is used to reconstruct the original speech feature from the private and shared components as a regularization. This domain separation framework is applied to the unsupervised environment adaptation task and achieved 11.08% relative WER reduction from the gradient reversal layer training, a representative adversarial training method, for automatic speech recognition on CHiME-3 dataset. Zhong Meng, Zhuo Chen 0006, Vadim Mazalov, Jinyu Li 0001, Yifan Gong 0001 |
ASRU | 5 |
| 2017 | Improved cepstra minimum-mean-square-error noise reduction algorithm for robust speech recognitionabstractIn the era of deep learning, although beam-forming multi-channel signal processing is still very helpful, it was reported that single-channel robust front-ends usually cannot benefit deep learning models because the layer-by-layer structure of deep learning models provides a feature extraction strategy that automatically derives powerful noise-resistant features from primitive raw data for senone classification. In this study, we show that the single-channel robust front-end is still very beneficial to deep learning modelling as long as it is well designed. We improve a robust front-end, cepstra minimum mean square error (CMMSE), by using more reliable voice activity detector, refined prior SNR estimation, better gain smoothing and two-stage processing. This new front-end, improved CMMSE (ICMMSE), is evaluated on the standard Aurora 2 and Chime 3 tasks, and a 3400 hour Microsoft Cortana digital assistant task using Gaussian mixture models, feed-forward deep neural networks, and long short-term memory recurrent neural networks, respectively. It is shown that ICMMSE is superior regardless of the underlying acoustic models and the scale of evaluation tasks, with 25.46% relative WER reduction on Aurora 2, up to 11.98% relative WER reduction on Chime 3, and up to 11.01% relative WER reduction on Cortana digital assistant task, respectively. Jinyu Li 0001, Yan Huang 0028, Yifan Gong 0001 |
ICASSP | 3 |
| 2017 | Extended low-rank plus diagonal adaptation for deep and recurrent neural networksabstractRecently, the low-rank plus diagonal (LRPD) adaptation was proposed for speaker adaptation of deep neural network (DNN) models. The LRPD restructures the adaptation matrix as a superposition of a diagonal matrix and a product of two low-rank matrices. In this paper, we extend the LRPD adaptation into the subspace-based approach to further reduce the speaker-dependent (SD) footprint. We apply the extended LRPD (eLRPD) adaptation for the DNN and LSTM models with emphasis placed on the applicability of the adaptation to large-scale speech recognition systems. To speed up the adaptation in test time, we propose the bottleneck (BN) caching approach to eliminate the redundant computations during multiple sweeps of development data. Experimental results on the short message dictation (SMD) task show that the eLRPD adaptation can reduce the SD footprints by 82% for the SVD DNN and 96% for the LSTM-RNN over the linear adaptation, while maintaining the comparable accuracy. The BN caching achieves up to 3.5 times speedup in adaptation at no loss of recognition accuracy. Yong Zhao 0008, Jinyu Li 0001, Kshitiz Kumar, Yifan Gong 0001 |
ICASSP | 4 |
| 2017 | Improving Mask Learning Based Speech Enhancement System with Restoration Layers and Residual ConnectionabstractFor single-channel speech enhancement, mask learning based approach through neural network has been shown to outperform the feature mapping approach, and to be effective as a pre-processor for automatic speech recognition. However, its assumption that the mixture and clean reference must have the correspondent scale doesn’t hold in data collected from real world, and thus leads to significant performance degradation on parallel recorded data. In this paper, we first extend the mask learning based speech enhancement by integrating two types of restoration layer to address the scale mismatch problem. We further propose a novel residual learning based speech enhancement model via adding different shortcut connections to a feature mapping network. We show such a structure can benefit from both the mask learning and the feature mapping. We evaluate the proposed speech enhancement models on CHiME 3 data. Without retraining the acoustic model, the best bidirection LSTM with residue connections yields 24.90% relative WER reduction on real data and 34.57% WER on simulated data. Zhuo Chen 0006, Yan Huang 0028, Jinyu Li 0001, Yifan Gong 0001 |
INTERSPEECH | 4 |
| 2017 | Don't Count on ASR to Transcribe for You: Breaking Bias with Two CrowdsabstractA crowdsourcing approach for collecting high-quality speech transcriptions is presented. The approach addresses typical weakness of traditional semi-supervised transcription strategies that show ASR hypotheses to transcribers to help them cope with unclear or ambiguous audio and speed up transcriptions. We explain how the traditional methods introduce bias into transcriptions that make it difficult to objectively measure system improvements against existing baselines, and suggest a two-stage crowdsourcing alternative that, first, iteratively collects transcription hypotheses and, then, asks a different crowd to pick the best of them. We show that this alternative not only outperforms the traditional method in a side-by-side comparison, but it also leads to ASR improvements due to superior quality of acoustic and language models trained on the transcribed data. Michael Levit, Yan Huang 0028, Shuangyu Chang, Yifan Gong 0001 |
INTERSPEECH | 4 |
| 2017 | Large-Scale Domain Adaptation via Teacher-Student LearningabstractHigh accuracy speech recognition requires a large amount of transcribed data for supervised training.In the absence of such data, domain adaptation of a well-trained acoustic model can be performed, but even here, high accuracy usually requires significant labeled data from the target domain.In this work, we propose an approach to domain adaptation that does not require transcriptions but instead uses a corpus of unlabeled parallel data, consisting of pairs of samples from the source domain of the well-trained model and the desired target domain.To perform adaptation, we employ teacher/student (T/S) learning, in which the posterior probabilities generated by the source-domain model can be used in lieu of labels to train the target-domain model.We evaluate the proposed approach in two scenarios, adapting a clean acoustic model to noisy speech and adapting an adults' speech acoustic model to children's speech.Significant improvements in accuracy are obtained, with reductions in word error rate of up to 44% over the original source model without the need for transcribed data in the target domain.Moreover, we show that increasing the amount of unlabeled data results in additional model robustness, which is particularly beneficial when using simulated training data in the target-domain. Jinyu Li 0001, Michael L. Seltzer, Rui Zhao 0017, Yifan Gong 0001 |
INTERSPEECH | 5 |
| 2016 | Non-negative intermediate-layer DNN adaptation for a 10-KB speaker adaptation profileabstractPreviously we demonstrated that speaker adaptation of acoustic models (AM) can provide significant improvement in the accuracy of large-scale speech recognition systems. In this work we discuss numerous challenges in scaling speaker adaptation to millions of speakers, where the size of speaker-dependent (SD) parameters is a critical challenge. Subsequently, we formulate an intermediate-layer adaptation framework for adaptation, upon which we build a non-negative adaptation for a very sparse set of non-negative SD parameters. We further improve this work with, (a) non-negative adaptation with a small-positive threshold, (b) setting small-positive weights in an already trained non-negative model to zero. We also discuss effective methods to store the non-negative SD parameters. We show that our methods reduce the SD parameters from 86KB for our previous best adaptation approach to 8.8KB, thus about 90% relative reduction in the size of SD parameters, and still retain 10+% word-error-rate-relative (WERR) gain over the baseline speaker-independent (SI) model. Kshitiz Kumar, Chaojun Liu, Yifan Gong 0001 |
ICASSP | 3 |
| 2016 | Exploring multidimensional lstms for large vocabulary ASRabstractLong short-term memory (LSTM) recurrent neural networks (RNNs) have recently shown significant performance improvements over deep feed-forward neural networks. A key aspect of these models is the use of time recurrence, combined with a gating architecture that allows them to track the long-term dynamics of speech. Inspired by human spectrogram reading, we recently proposed the frequency LSTM (F-LSTM) that performs 1-D recurrence over the frequency axis and then performs 1-D recurrence over the time axis. In this study, we further improve the acoustic model by proposing a 2-D, time-frequency (TF) LSTM. The TF-LSTM jointly scans the input over the time and frequency axes to model spectro-temporal warping, and then uses the output activations as the input to a time LSTM (T-LSTM). The joint time-frequency modeling better normalizes the features for the upper layer T-LSTMs. Evaluated on a 375-hour short message dictation task, the proposed TF-LSTM obtained a 3.4% relative WER reduction over the best T-LSTM. The invariance property achieved by joint time-frequency analysis is demonstrated on a mismatched test set, where the TF-LSTM achieves a 14.2% relative WER reduction over the best T-LSTM. Jinyu Li 0001, Abdel-rahman Mohamed, Geoffrey Zweig, Yifan Gong 0001 |
ICASSP | 4 |
| 2016 | Investigations on speaker adaptation of LSTM RNN models for speech recognitionabstractRecently Long Short-Term Memory (LSTM) Recurrent Neural Networks (RNN) acoustic models have demonstrated superior performance over deep neural networks (DNN) models in speech recognition and many other tasks. Although a lot of work have been reported on DNN model adaptation, very little has been done on LSTM model adaptation. In this paper we present our extensive studies of speaker adaptation of LSTM-RNN models for speech recognition. We investigated different adaptation methods combined with KL-divergence based regularization, where and which network component to adapt, supervised versus unsupervised adaptation and asymptotic analysis. We made a few distinct and important observations. In a large vocabulary speech recognition task, by adapting only 2.5% of the LSTM model parameters using 50 utterances per speaker, we obtained 12.6% WERR on the dev set and 9.1% WERR on the evaluation set over a strong LSTM baseline model. Chaojun Liu, Kshitiz Kumar, Yifan Gong 0001 |
ICASSP | 4 |
| 2016 | Simplifying long short-term memory acoustic models for fast training and decodingabstractOn acoustic modeling, recurrent neural networks (RNNs) using Long Short-Term Memory (LSTM) units have recently been shown to outperform deep neural networks (DNNs) models. This paper focuses on resolving two challenges faced by LSTM models: high model complexity and poor decoding efficiency. Motivated by our analysis of the gates activation and function, we present two LSTM simplifications: deriving input gates from forget gates, and removing recurrent inputs from output gates. To accelerate decoding of LSTMs, we propose to apply frame skipping during training, and frame skipping and posterior copying (FSPC) during decoding. In the experiments, model simplifications reduce the size of LSTM models by 26%, resulting in a simpler model structure. Meanwhile, the application of FSPC speeds up model computation by 2 times during LSTM decoding. All these improvements are achieved at the cost of 1% WER degradation. Yajie Miao, Jinyu Li 0001, Shixiong Zhang 0001, Yifan Gong 0001 |
ICASSP | 5 |
| 2016 | Geo-location dependent deep neural network acoustic model for speech recognitionabstractUsers from the same geo-location region exhibit similar acoustic characteristics, e.g., they have similar accent; even more, they may have similar preference to device. In this paper, we propose to build geo-location dependent deep neural network for speech recognition, where the geo-location signal is inferred from users' GPS. During runtime, the server will base on a user's geo-location to select the right model to recognize his voice. We tackle three major issues associated with this model: high train/deployment cost, large model size, and train data sparsity. Our solution is featured by its low cost, thus practical for production modeling. We also discuss the reliability of GPS signal in practical use. The proposed model is evaluated on Microsoft Chinese voice search and Cortana live test set. Among 12 provinces, it shows an overall 4.8% relative character error rate reduction, over a strong baseline production-level model, with only 50% model size increase. The gain is larger for the low-resource provinces, with relative error rate reduction up to 9%. Guoli Ye, Chaojun Liu, Yifan Gong 0001 |
ICASSP | 3 |
| 2016 | Recurrent support vector machines for speech recognitionabstractRecurrent Neural Networks (RNNs) using Long-Short Term Memory (LSTM) architecture have demonstrated the state-of-the-art performances on speech recognition. Most of deep RNNs use the softmax activation function in the last layer for classification. This paper illustrates small but consistent advantages of replacing the softmax layer in RNN with Support Vector Machines (SVMs). The parameters of RNNs and SVMs are jointly learned using a sequence-level max-margin criteria, instead of cross-entropy. The resulting model is termed Recurrent SVM. The conventional SVMs need to predefine a feature space and do not have internal states to deal with arbitrary long-term dependencies in sequences. The proposed recurrent SVM uses LSTMs to learn the feature space and to capture temporal dependencies, while using the SVM (in the last layer) for sequence classification. The model is evaluated on the Windows phone task for large vocabulary continuous speech recognition. Shixiong Zhang 0001, Rui Zhao 0017, Chaojun Liu, Jinyu Li 0001, Yifan Gong 0001 |
ICASSP | 5 |
| 2016 | Low-rank plus diagonal adaptation for deep neural networksabstractIn this paper, we propose a scalable adaptation technique that adapts the deep neural network (DNN) model through the low-rank plus diagonal (LRPD) decomposition. It is desired that an adaptation method can properly accommodate the available development data with a variable amount of adaptation parameters. Thus, the resulting models neither over-fit nor under-fit as the development data vary in size for different speakers. The technique developed in this paper is inspired by observing that adaptation matrices are very close to an identity matrix or diagonally dominant. The LRPD restructures the adaptation matrix as a superposition of a diagonal matrix and a low-rank matrix. By varying the low-rank values, the LRPD contains the full and the diagonal adaptation matrix as its special cases. Experimental results demonstrated that the LRPD adaptation of the full-size DNN obtains improved accuracy over the standard linear adaptation. The LRPD bottleneck adaptation can reduce the speaker-specific footprint by 82% over an already very compact SVD bottleneck adaptation, at an expense of 1% relative WER increase. Yong Zhao 0008, Jinyu Li 0001, Yifan Gong 0001 |
ICASSP | 3 |
| 2016 | Semi-Supervised Training in Deep Learning Acoustic ModelabstractWe studied the semi-supervised training in a fully connected deep neural network (DNN), unfolded recurrent neural network (RNN), and long short-term memory recurrent neural network (LSTM-RNN) with respect to the transcription quality, the importance data sampling, and the training data amount. We found that DNN, unfolded RNN, and LSTM-RNN are increasingly more sensitive to labeling errors. For example, with the simulated erroneous training transcription at 5%, 10%, or 15% word error rate (WER) level, the semi-supervised DNN yields 2.37%, 4.84%, or 7.46% relative WER increase against the baseline model trained with the perfect transcription; in comparison, the corresponding WER increase is 2.53%, 4.89%, or 8.85% in an unfolded RNN and 4.47%, 9.38%, or 14.01% in an LSTM-RNN. We further found that the importance sampling has similar impact on all three models with 2~3% relative WER reduction comparing to the random sampling. Lastly, we compared the modeling capability with increased training data. Experimental results suggested that LSTM-RNN can benefit more from enlarged training data comparing to unfolded RNN and DNN. We trained a semi-supervised LSTM-RNN using 2600 hr transcribed and 10000 hr untranscribed data on a mobile speech task. The semi-supervised LSTM-RNN yields 7.9\% relative WER reduction against the supervised baseline. Yan Huang 0028, Yifan Gong 0001 |
INTERSPEECH | 3 |
| 2016 | End-to-End attention based text-dependent speaker verificationabstractA new type of End-to-End system for text-dependent speaker verification is presented in this paper. Previously, using the phonetic discriminate/speaker discriminate DNN as a feature extractor for speaker verification has shown promising results. The extracted frame-level (bottleneck, posterior or d-vector) features are equally weighted and aggregated to compute an utterance-level speaker representation (d-vector or i-vector). In this work we use a speaker discriminate CNN to extract the noise-robust frame-level features. These features are smartly combined to form an utterance-level speaker vector through an attention mechanism. The proposed attention model takes the speaker discriminate information and the phonetic information to learn the weights. The whole system, including the CNN and attention model, is joint optimized using an end-to-end criterion. The training algorithm imitates exactly the evaluation process — directly mapping a test utterance and a few target speaker utterances into a single verification score. The algorithm can smartly select the most similar impostor for each target speaker to train the network. We demonstrated the effectiveness of the proposed end-to-end system on Windows 10 “Hey Cortana” speaker verification task. Shixiong Zhang 0001, Zhuo Chen 0006, Yong Zhao 0008, Jinyu Li 0001, Yifan Gong 0001 |
SLT | 5 |
| 2015 | LSTM time and frequency recurrence for automatic speech recognitionabstractLong short-term memory (LSTM) recurrent neural networks (RNNs) have recently shown significant performance improvements over deep feed-forward neural networks (DNNs). A key aspect of these models is the use of time recurrence, combined with a gating architecture that ameliorates the vanishing gradient problem. Inspired by human spectrogram reading, in this paper we propose an extension to LSTMs that performs the recurrence in frequency as well as in time. This model first scans the frequency bands to generate a summary of the spectral information, and then uses the output layer activations as the input to a traditional time LSTM (T-LSTM). Evaluated on a Microsoft short message dictation task, the proposed model obtained a 3.6% relative word error rate reduction over the T-LSTM. Jinyu Li 0001, Abdel-rahman Mohamed, Geoffrey Zweig, Yifan Gong 0001 |
ASRU | 4 |
| 2015 | An analysis of convolutional neural networks for speech recognitionabstractDespite the fact that several sites have reported the effectiveness of convolutional neural networks (CNNs) on some tasks, there is no deep analysis regarding why CNNs perform well and in which case we should see CNNs' advantage. In the light of this, this paper aims to provide some detailed analysis of CNNs. By visualizing the localized filters learned in the convolutional layer, we show that edge detectors in varying directions can be automatically learned. We then identify four domains we think CNNs can consistently provide advantages over fully-connected deep neural networks (DNNs): channel-mismatched training-test conditions, noise robustness, distant speech recognition, and low-footprint models. For distant speech recognition, a CNN trained on 1000 hours of Kinect distant speech data obtains relative 4% word error rate reduction (WERR) over a DNN of a similar size. To our knowledge, this is the largest corpus so far reported in the literature for CNNs to show its effectiveness. Lastly, we establish that the CNN structure combined with maxout units is the most effective model under small-sizing constraints for the purpose of deploying small-footprint models to devices. This setup gives relative 9.3% WERR from DNNs with sigmoid units. Jui-Ting Huang, Jinyu Li 0001, Yifan Gong 0001 |
ICASSP | 3 |
| 2015 | Estimating confidence scores on ASR results using recurrent neural networksabstractIn this paper we present a confidence estimation system using recurrent neural networks (RNN) and compare it to a traditional multilayered perception (MLP) based system. The ability of RNN to capture sequence information and improve decisions using processed history was main motivation to explore RNN's for confidence estimation. In this paper we also explore two subtle variations of confidence estimator: one that uses objective extracted over the entire sequence for training, and other that uses dynamic programming to decode and estimate confidence on all the words of the sequence jointly. In our experiments, we observed that for a constant false positive (FP) rate of 3% we can secure a relative reduction of 10% in false negative (FN) rate when we replaced a MLP in confidence estimator with a RNN.We also observed that relative gains achieved by a RNN based confidence estimator are directly proportional to the number of word in the utterances. Kaustubh Kalgaonkar, Chaojun Liu, Yifan Gong 0001, Kaisheng Yao |
ICASSP | 3 |
| 2015 | Small-footprint high-performance deep neural network-based speech recognition using split-VQabstractDue to a large number of parameters in deep neural networks (DNNs), it is challenging to design a small-footprint DNN-based speech recognition system while maintaining a high recognition performance. Even with a singular value matrix decomposition (SVD) method and scalar quantization, the DNN model is still too large to be deployed on many mobile devices. Common practices like reducing the number of hidden nodes often result in significant accuracy loss. In this work, we propose to split each row vector of weight matrices into sub-vectors, and quantize them into a set of codewords using a split vector quantization (split-VQ) algorithm. The codebook can be fine-tuned using back-propagation when an aggressive quantization is performed. Experimental results demonstrate that the proposed method can further reduce the model size by 75% to 80% and save 10% to 50% computation on top of an already very compact SVD-DNN without a noticeable performance degradation. This results in a 3.2 MB-footprint DNN giving similar recognition performance as what a 59.1 MB standard DNN can achieve. Jinyu Li 0001, Yifan Gong 0001 |
ICASSP | 3 |
| 2015 | Deep neural support vector machines for speech recognitionabstractA new type of deep neural networks (DNNs) is presented in this paper. Traditional DNNs use the multinomial logistic regression (softmax activation) at the top layer for classification. The new DNN instead uses a support vector machine (SVM) at the top layer. Two training algorithms are proposed at the frame and sequence-level to learn parameters of SVM and DNN in the maximum-margin criteria. In the frame-level training, the new model is shown to be related to the multiclass SVM with DNN features; In the sequence-level training, it is related to the structured SVM with DNN features and HMM state transition features. Its decoding process is similar to the DNN-HMM hybrid system but with frame-level posterior probabilities replaced by scores from the SVM. We term the new model deep neural support vector machine (DNSVM). We have verified its effectiveness on the TIMIT task for continuous speech recognition. Shixiong Zhang 0001, Chaojun Liu, Kaisheng Yao, Yifan Gong 0001 |
ICASSP | 4 |
| 2015 | Investigating online low-footprint speaker adaptation using generalized linear regression and click-through dataabstractTo develop speaker adaptation algorithms for deep neural network (DNN) that are suitable for large-scale online deployment, it is desirable that the adaptation model be represented in a compact form and learned in an unsupervised fashion. In this paper, we propose a novel low-footprint adaptation technique for DNN that adapts the DNN model through node activation functions. The approach introduces slope and bias parameters in the sigmoid activation functions for each speaker, allowing the adaptation model to be stored in a small-sized storage space. We show that this adaptation technique can be formulated in a linear regression fashion, analogous to other speak adaptation algorithms that apply additional linear transformations to the DNN layers. We further investigate semi-supervised online adaptation by making use of the user click-through data as a supervision signal. The proposed method is evaluated on short message dictation and voice search tasks in supervised, unsupervised, and semi-supervised setups. Compared with the singular value decomposition (SVD) bottleneck adaptation, the proposed adaptation method achieves comparable accuracy improvements with much smaller footprint. Yong Zhao 0008, Jinyu Li 0001, Yifan Gong 0001 |
ICASSP | 4 |
| 2015 | Regularized sequence-level deep neural network model adaptationabstractWe propose a regularized sequence-level (SEQ) deep neural network (DNN) model adaptation methodology as an extension of the previous KL-divergence regularized cross-entropy (CE) adaptation [1]. In this approach, the negative KL-divergence between the baseline and the adapted model is added to the maximum mutual information (MMI) as regularization in the sequence-level adaptation. We compared eight different adaptation setups specified by the baseline training criterion, the adaptation criterion, and the regularization methodology. We found that the proposed sequence-level adaptation consistently outperforms the crossentropy adaptation. For both of them, regularization is critical. We further introduced a unified formulation in which the regularized CE and SEQ adaptation are the special cases. We applied the proposed approach to speaker adaptation and accent adaptation in a mobile short message dictation task. For the speaker adaptation, with 25 or 100 utterances, the proposed approach yields 13.72% or 23.18% WER reduction when adapting from the CE baseline, comparing to 11.87% or 20.18% for the CE adaptation. For the accent adaptation, with 1K utterances, the proposed approach yields 18.74% or 19.50% WER reduction when adapting from the CE-DNN or the SEQ-DNN. The WER reduction using the regularized CE adaptation is 15.98% and 15.69%, respectively. Yan Huang 0028, Yifan Gong 0001 |
INTERSPEECH | 2 |
| 2015 | Confidence-features and confidence-scores for ASR applications in arbitration and DNN speaker adaptation
Kshitiz Kumar, Ziad Al Bawab, Yong Zhao 0008, Chaojun Liu, Benoît Dumoulin, Yifan Gong 0001 |
INTERSPEECH | 6 |
| 2015 | Delta-melspectra features for noise robustness to DNN-based ASR systems
Kshitiz Kumar, Chaojun Liu, Yifan Gong 0001 |
INTERSPEECH | 3 |
| 2015 | Intermediate-layer DNN adaptation for offline and session-based iterative speaker adaptation
Kshitiz Kumar, Chaojun Liu, Kaisheng Yao, Yifan Gong 0001 |
INTERSPEECH | 4 |
| 2015 | SVD-based universal DNN modeling for multiple scenariosabstractSpeech recognition scenarios (aka tasks) differ from each other in acoustic transducers, acoustic environments, and speaking style etc. Building one acoustic model per task is one common practice in industry. However, this limits training data sharing across scenarios thus may not give highest possible accuracy. Based on the deep neural network (DNN) technique, we propose to build a universal acoustic model for all scenarios by utilizing all the data together. Two advantages are obtained: 1) leveraging more data sources to improve the recognition accuracy, 2) reducing substantially service deployment and maintenance costs. We achieve this by extending the singular value decomposition (SVD) structure of DNNs. The data from all scenarios are used to first train a single SVD-DNN model. Then a series of scenario-dependent linear square matrices are added on top of each SVD layer and updated with only scenario-related data. At the recognition time, a flag indicates different scenarios and guides the recognizer to use the scenario-dependent matrices together with the scenario-independent matrices in the universal DNN for acoustic score evaluation. In our experiments on Microsoft Winphone/Skype/Xbox data sets, the universal DNN model is better than traditional trained isolated models, with up to 15.5% relative word error rate reduction. Changliang Liu, Jinyu Li 0001, Yifan Gong 0001 |
INTERSPEECH | 3 |
| 2014 | Factorized adaptation for deep neural networkabstractIn this paper, we propose a novel method to adapt context-dependent deep neural network hidden Markov model (CD-DNN-HMM) with only limited number of parameters by taking into account the underlying factors that contribute to the distorted speech signal. We derive this factorized adaptation method from the perspectives of joint factor analysis and vector Taylor series expansion, respectively. Evaluated on Aurora 4, the proposed method can get 19.0% and 10.6% relative word error rate reduction on test set B and D with only 20 adaptation utterances, and can have decent improvement with as few as two adaptation utterances. We also show that the proposed method is better than feature discriminative linear regression (fDLR), an existing DNN adaptation method. Its small number of parameters and short training time offer an attractive solution to low-footprint speech applications. Jinyu Li 0001, Jui-Ting Huang, Yifan Gong 0001 |
ICASSP | 3 |
| 2014 | Singular value decomposition based low-footprint speaker adaptation and personalization for deep neural networkabstractThe large number of parameters in deep neural networks (DNN) for automatic speech recognition (ASR) makes speaker adaptation very challenging. It also limits the use of speaker personalization due to the huge storage cost in large-scale deployments. In this paper we address DNN adaptation and personalization issues by presenting two methods based on the singular value decomposition (SVD). The first method uses an SVD to replace the weight matrix of a speaker independent DNN by the product of two low rank matrices. Adaptation is then performed by updating a square matrix inserted between the two low-rank matrices. In the second method, we adapt the full weight matrix but only store the delta matrix - the difference between the original and adapted weight matrices. We decrease the footprint of the adapted model by storing a reduced rank version of the delta matrix via an SVD. The proposed methods were evaluated on short message dictation task. Experimental results show that we can obtain similar accuracy improvements as the previously proposed Kullback-Leibler divergence (KLD) regularized method with far fewer parameters, which only requires 0.89% of the original model storage. Jinyu Li 0001, Dong Yu 0001, Mike Seltzer, Yifan Gong 0001 |
ICASSP | 5 |
| 2014 | Towards better performance with heterogeneous training data in acoustic modeling using deep neural networksabstractModeling heterogeneous data sources remains a fundamental chal-lenge of acoustic modeling in speech recognition. We call this the multi-condition problem because the speech data come from many different conditions. In this paper, we introduce the fundamen-tal confusability problem in multi-condition learning, then discuss the problem formalization, the taxonomy, and the architectures for multi-condition learning. While the ideas presented are applicable to all classifiers, we focus our attention in this work on acoustic models based on deep neural networks (DNN). We propose four different strategies for multi-condition learning of a DNN that we refer to as a mixed-condition model, a condition-dependent model, a condition-normalizing model, and a condition-aware model. Based on the experimental results on the voice search and short message dictation task and the Aurora 4 task, we show that the confusabil-ity introduced when modeling heterogeneous data depends on the source of acoustic distortion itself, the front-end feature extractor, and the classifier. We also demonstrate the best approach for dealing with heterogeneous data may not be to let the model sort it out blindly, even with a classifier as sophisticated as a DNN. Index Terms — Multi-task learning, deep learning, CD-DNN-HMM, noise robustness, channel compensation Yan Huang 0028, Malcolm Slaney, Michael L. Seltzer, Yifan Gong 0001 |
INTERSPEECH | 4 |
| 2014 | A comparative analytic study on the Gaussian mixture and context dependent deep neural network hidden Markov modelsabstractWe conducted a comparative analytic study on the contextdependent Gaussian mixture hiddenMarkov model (CD-GMMHMM) and deep neural network hidden Markov model (CDDNN-HMM) with respect to the phone discrimination and the robustness performance. We found that the DNN can significantly improve the phone recognition performance for every phoneme with 15.6% to 39.8% relative phone error rate reduction (PERR). It is particularly good at discriminating certain consonants, which are found to be “hard” in the GMM. On the robustness side, the DNN outperforms the GMM at all SNR levels, across different devices, and under all speaking rate with nearly uniform improvement. The performance gap with respect to different SNR levels, distinct channels, and varied speaking rate remains large. For example, in CD-DNNHMM, we observed 1∼2% performance degradation per 1dB SNR drop; 20∼25% performance gap between the best and least well performed devices; 15∼30% relative word error rate increase when the speaking rate speeds up or slows down by 30% from the “sweet” spot. Therefore, we conclude the robustness remains to be a major challenge in the deep learning acoustic model. Speech enhancement, channel normalization, and speaking rate compensation are important research areas in order to further improve the DNN model accuracy. Yan Huang 0028, Dong Yu 0001, Chaojun Liu, Yifan Gong 0001 |
INTERSPEECH | 4 |
| 2014 | Multi-accent deep neural network acoustic model with accent-specific top layer using the KLD-regularized model adaptationabstractWe propose a multi-accent deep neural network acoustic model with an accent-specific top layer and shared bottom hidden layers. The accent-specific top layer is used to model the distinct accent specific patterns. The shared bottom hidden layers allow maximum knowledge sharing between the native and the accent models. This design is particularly attractive when considering deploying such a system to a live speech service due to its computational efficiency. We applied the KL-divergence (KLD) regularized model adaptation to train the accent-specific top layer. On the mobile short message dictation task (SMD), with 1K, 10K, and 100K British or Indian accent adaptation utterances, the proposed approach achieves 18.1%, 26.0%, and 28.5% or 16.1%, 25.4%, and 30.6% word error rate reduction (WERR) for the British and the Indian accent respectively against a baseline cross entropy (CE) model trained from 400 hour data. On the 100K utterance accent adaptation setup, comparable performance gain can be obtained against a baseline CE model trained with 2000 hour data. We observe smaller yet significant WER reduction on a baseline model trained using the MMI sequence-level criterion. Yan Huang 0028, Dong Yu 0001, Chaojun Liu, Yifan Gong 0001 |
INTERSPEECH | 4 |
| 2014 | Normalization of ASR confidence classifier scores via confidence mapping
Kshitiz Kumar, Chaojun Liu, Yifan Gong 0001 |
INTERSPEECH | 3 |
| 2014 | Learning small-size DNN with output-distribution-based criteriaabstractDeep neural network (DNN) obtains significant accuracy improvements on many speech recognition tasks and its power comes from the deep and wide network structure with a very large number of parameters. It becomes challenging when we deploy DNN on devices which have limited computational and storage resources. The common practice is to train a DNN with a small number of hidden nodes and a small senone set using the standard training process, leading to significant accuracy loss. In this study, we propose to better address these issues by utilizing the DNN output distribution. To learn a DNN with small number of hidden nodes, we minimize the Kullback–Leibler divergence between the output distributions of the small-size DNN and a standard large-size DNN by utilizing a large number of un-transcribed data. For better senone set generation, we cluster the senones in the large set into a small one by directly relating the clustering process to DNN parameters, as opposed to decoupling the senone generation and DNN training process in the standard training. Evaluated on a short message dictation task, the proposed two methods get 5.08% and 1.33% relative word error rate reduction from the standard training method, respectively. Jinyu Li 0001, Rui Zhao 0017, Jui-Ting Huang, Yifan Gong 0001 |
INTERSPEECH | 4 |
| 2014 | Variable-component deep neural network for robust speech recognition
Rui Zhao 0017, Jinyu Li 0001, Yifan Gong 0001 |
INTERSPEECH | 3 |
| 2014 | Variable-activation and variable-input deep neural network for robust speech recognitionabstractIn a previous study, we proposed variable-component deep neural network (VCDNN) to improve the robustness of context-dependent deep neural network hidden Markov model (CD-DNN-HMM). We model the components of DNN a set of polynomial functions of environmental variables, more specifically signal-to-noise ratio (SNR). We refined VCDNN on two types of DNN components: (1) weighting matrix and bias (2) the output of each layer. These two methods are called variable-parameter DNN (VPDNN) and variable-output DNN (VODNN). Although both methods got good gain over the standard DNN, they doubled the number of parameters even with only the first-order environment variable. In this study, we propose two new types of VCDNN, namely variable activation DNN (VADNN) and variable input DNN (VIDNN). The environment variable is applied to the hidden layer activation function in VADNN, and is applied directly to the input in VIDNN. Both DNNs only increase a negligible number of parameters compared to the standard DNN. Experimental results on Aurora4 task show that both methods are effective, and VIDNN can beat all other variations of VCDNN with relative 7.69% word error reduction from the standard DNN with the least increase in number of parameters. Rui Zhao 0017, Jinyu Li 0001, Yifan Gong 0001 |
SLT | 3 |
| 2014 | A fast maximum likelihood nonlinear feature transformation method for GMM-HMM speaker adaptation
Kaisheng Yao, Dong Yu 0001, Li Deng 0001, Yifan Gong 0001 |
Neurocomputing | 4 |
| 2014 | An Overview of Noise-Robust Automatic Speech RecognitionabstractNew waves of consumer-centric applications, such as voice search and voice interaction with mobile devices and home entertainment systems, increasingly require automatic speech recognition (ASR) to be robust to the full range of real-world noise and other acoustic distorting conditions. Despite its practical importance, however, the inherent links between and distinctions among the myriad of methods for noise-robust ASR have yet to be carefully studied in order to advance the field further. To this end, it is critical to establish a solid, consistent, and common mathematical foundation for noise-robust ASR, which is lacking at present. This article is intended to fill this gap and to provide a thorough overview of modern noise-robust techniques for ASR developed over the past 30 years. We emphasize methods that are proven to be successful and that are likely to sustain or expand their future applicability. We distill key insights from our comprehensive overview in this field and take a fresh look at a few old problems, which nevertheless are still highly relevant today. Specifically, we have analyzed and categorized a wide range of noise-robust techniques using five different criteria: 1) feature-domain vs. model-domain processing, 2) the use of prior knowledge about the acoustic environment distortion, 3) the use of explicit environment-distortion models, 4) deterministic vs. uncertainty processing, and 5) the use of acoustic models trained jointly with the same feature enhancement or model adaptation process used in the testing stage. With this taxonomy-oriented review, we equip the reader with the insight to choose among techniques and with the awareness of the performance-complexity tradeoffs. The pros and cons of using different noise-robust ASR techniques in practical application scenarios are provided as a guide to interested practitioners. The current challenges and future research directions in this field is also carefully analyzed. Jinyu Li 0001, Li Deng 0001, Yifan Gong 0001, Reinhold Häb-Umbach |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2013 | Recent advances in deep learning for speech research at MicrosoftabstractDeep learning is becoming a mainstream technology for speech recognition at industrial scale. In this paper, we provide an overview of the work by Microsoft speech researchers since 2009 in this area, focusing on more recent advances which shed light to the basic capabilities and limitations of the current deep learning technology. We organize this overview along the feature-domain and model-domain dimensions according to the conventional approach to analyzing speech systems. Selected experimental results, including speech recognition and related applications such as spoken dialogue and language modeling, are presented to demonstrate and analyze the strengths and weaknesses of the techniques described in the paper. Potential improvement of these techniques and future research directions are discussed. Li Deng 0001, Jinyu Li 0001, Jui-Ting Huang, Kaisheng Yao, Dong Yu 0001, Frank Seide, Michael L. Seltzer, Geoffrey Zweig, Xiaodong He 0001, Jason D. Williams, Yifan Gong 0001, Alex Acero |
ICASSP | 11 |
| 2013 | Predicting speech recognition confidence using deep learning with word identity and score featuresabstractConfidence classifiers for automatic speech recognition (ASR) provide a quantitative representation for the reliability of ASR decoding. In this paper, we improve the ASR confidence measure performance for an utterance using two distinct approaches: (1) to define and incorporate additional predictors in the confidence classifier including those based on the word identity and on the aggregated words, and (2) to train the confidence classifier built on deep learning architectures including the deep neural network (DNN) and the kernel deep convex network (K-DCN). Our experiments show that adding the new predictors to our multi-layer perceptron (MLP)-based baseline classifier provides 38.6% relative reduction in the correct-reject rate as our measure of the classifier performance. Further, replacing the MLP with the DNN and K-DCN provides an additional 14.5% and 47.5% in the relative performance gain, respectively. Po-Sen Huang, Kshitiz Kumar, Chaojun Liu, Yifan Gong 0001, Li Deng 0001 |
ICASSP | 4 |
| 2013 | Cross-language knowledge transfer using multilingual deep neural network with shared hidden layersabstractIn the deep neural network (DNN), the hidden layers can be considered as increasingly complex feature transformations and the final softmax layer as a log-linear classifier making use of the most abstract features computed in the hidden layers. While the loglinear classifier should be different for different languages, the feature transformations can be shared across languages. In this paper we propose a shared-hidden-layer multilingual DNN (SHL-MDNN), in which the hidden layers are made common across many languages while the softmax layers are made language dependent. We demonstrate that the SHL-MDNN can reduce errors by 3-5%, relatively, for all the languages decodable with the SHL-MDNN, over the monolingual DNNs trained using only the language specific data. Further, we show that the learned hidden layers sharing across languages can be transferred to improve recognition accuracy of new languages, with relative error reductions ranging from 6% to 28% against DNNs trained without exploiting the transferred hidden layers. It is particularly interesting that the error reduction can be achieved for the target language that is in different families of the languages used to learn the hidden layers. Jui-Ting Huang, Jinyu Li 0001, Dong Yu 0001, Li Deng 0001, Yifan Gong 0001 |
ICASSP | 5 |
| 2013 | Semi-supervised GMM and DNN acoustic model training with multi-system combination and confidence re-calibrationabstractWe present our study on semi-supervised Gaussian mixture model (GMM) hidden Markov model (HMM) and deep neural network (DNN) HMM acoustic model training. We analyze the impact of transcription quality and data sampling approach on the performance of the resulting model, and propose a multisystem combination and confidence re-calibration approach to improve the transcription inference and data selection. Compared to using a single system recognition result and confidence score, our proposed approach reduces the phone error rate of the inferred transcription by 23.8% relatively when top 60% of data are selected. Experiments were conducted on the mobile short message dictation (SMD) task. For the GMM-HMM model, we achieved 7.2% relative word error rate reduction (WERR) against a well-trained narrow-band fMPE+bMMI system by adding 2100 hours of untranscribed data, and 28.2% relative WERR over a wide-band MLE model trained from transcribed out-of-domain voice search data after adding 10K hours of untranscribed SMD data. For the CD-DNN-HMM model, 11.7% and 15.0% relative WERRs are achieved after adding 1K hours of untranscribed data using random and importance sampling, respectively. We also found using large amount of untranscribed data for pretraining does not help. Index Terms: semi-supervised acoustic model training, system combination, confidence re-calibration, importance sampling Yan Huang 0028, Dong Yu 0001, Yifan Gong 0001, Chaojun Liu |
INTERSPEECH | 3 |
| 2013 | Restructuring of deep neural network acoustic models with singular value decompositionabstractRecently proposed deep neural network (DNN) obtains significant accuracy improvements in many large vocabulary continuous speech recognition (LVCSR) tasks. However, DNN requires much more parameters than traditional systems, which brings huge cost during online evaluation, and also limits the application of DNN in a lot of scenarios. In this paper we present our new effort on DNN aiming at reducing the model size while keeping the accuracy improvements. We apply singular value decomposition (SVD) on the weight matrices in DNN, and then restructure the model based on the inherent sparseness of the original matrices. After restructuring we can reduce the DNN model size significantly with negligible accuracy loss. We also fine-tune the restructured model using the regular back-propagation method to get the accuracy back when reducing the DNN model size heavily. The proposed method has been evaluated on two LVCSR tasks, with context-dependent DNN hidden Markov model (CD-DNN-HMM). Experimental results show that the proposed approach dramatically reduces the DNN model size by more than 80% without losing any accuracy. Index Terms: deep neural network, singular value decomposition, model restructuring Jinyu Li 0001, Yifan Gong 0001 |
INTERSPEECH | 3 |
| 2012 | Improvements to VTS feature enhancementabstractBy explicitly modelling the distortion of speech signals, model adaptation based on vector Taylor series (VTS) approaches have been shown to significantly improve the robustness of speech recognizers to environmental noise. However, the computational cost of VTS model adaptation (MVTS) methods hinders them from being widely used because they need to adapt all the HMM parameters for every utterance at runtime. In contrast, VTS feature enhancement (FVTS) methods have more computation advantages because they do not need multiple decoding passes and do not adapt all the HMM model parameters. In this paper, we propose two improvements to VTS feature enhancement: updating all of the environment distortion parameters and noise adaptive training of the front-end GMM. In addition, we investigate some other performance-related issues such as the selection of FVTS algorithms and the spectrum domain that MFCC is extracted from. As an important result of our investigation, we established the FVTS method can achieve comparable accuracy as the MVTS method with a smaller runtime cost. This makes FVTS method an ideal candidate for real world tasks. Jinyu Li 0001, Michael L. Seltzer, Yifan Gong 0001 |
ICASSP | 3 |
| 2012 | Efficient VTS Adaptation Using Jacobian ApproximationabstractBy exploiting a model of environmental distortion, model adaptation based on vector Taylor series (VTS) approaches have been shown to significantly improve the robustness of speech recognizers to environmental noise. However, the computational cost of VTS model adaptation (MVTS) methods hinders them from being more widely used. In this paper, we propose to reduce the computational cost of MVTS by replacing the Jacobian matrix used in the vector Taylor series approximation with a diagonal Jacobian matrix (DJVTS). We verify this approximation by showing that the Jacobian matrices are dominated by their diagonal elements and therefore the model distortion introduced by this approximation is very small. DJVTS gives similar accuracy as the standard MVTS method with significant reduction in computational cost. The proposed method also achieves higher accuracy than VTS-based feature enhancement. Jinyu Li 0001, Michael L. Seltzer, Yifan Gong 0001 |
INTERSPEECH | 3 |
| 2012 | A Feature Space Transformation Method for Personalization using Generalized I-Vector Clustering
Kaisheng Yao, Yifan Gong 0001, Chaojun Liu |
INTERSPEECH | 2 |
| 2012 | Improving wideband speech recognition using mixed-bandwidth training data in CD-DNN-HMMabstractContext-dependent deep neural network hidden Markov model (CD-DNN-HMM) is a recently proposed acoustic model that significantly outperformed Gaussian mixture model (GMM)-HMM systems in many large vocabulary speech recognition (LVSR) tasks. In this paper we present our strategy of using mixed-bandwidth training data to improve wideband speech recognition accuracy in the CD-DNN-HMM framework. We show that DNNs provide the flexibility of using arbitrary features. By using the Mel-scale log-filter bank features we not only achieve higher recognition accuracy than using MFCCs, but also can formulate the mixed-bandwidth training problem as a missing feature problem, in which several feature dimensions have no value when narrowband speech is presented. This treatment makes training CD-DNN-HMMs with mixed-bandwidth data an easy task since no bandwidth extension is needed. Our experiments on voice search data indicate that the proposed solution not only provides higher recognition accuracy for the wideband speech but also allows the same CD-DNN-HMM to recognize mixed-bandwidth speech. By exploiting mixed-bandwidth training data CD-DNN-HMM outperforms fMPE+BMMI trained GMM-HMM, which cannot benefit from using narrowband data, by 18.4%. Jinyu Li 0001, Dong Yu 0001, Jui-Ting Huang, Yifan Gong 0001 |
SLT | 4 |
| 2012 | Adaptation of context-dependent deep neural networks for automatic speech recognitionabstractIn this paper, we evaluate the effectiveness of adaptation methods for context-dependent deep-neural-network hidden Markov models (CD-DNN-HMMs) for automatic speech recognition. We investigate the affine transformation and several of its variants for adapting the top hidden layer. We compare the affine transformations against direct adaptation of the softmax layer weights. The feature-space discriminative linear regression (fDLR) method with the affine transformations on the input layer is also evaluated. On a large vocabulary speech recognition task, a stochastic gradient ascent implementation of the fDLR and the top hidden layer adaptation is shown to reduce word error rates (WERs) by 17% and 14%, respectively, compared to the baseline DNN performances. With a batch update implementation, the softmax layer adaptation technique reduces WERs by 10%. We observe that using bias shift performs as well as doing scaling plus bias shift. Kaisheng Yao, Dong Yu 0001, Frank Seide, Li Deng 0001, Yifan Gong 0001 |
SLT | 6 |
| 2010 | Unscented transform with online distortion estimation for HMM adaptationabstractIn this paper, we propose to improve our previously developed method for joint compensation of additive and convolutive distortions (JAC) applied to model adaptation. The improvement entails replacing the vector Taylor series (VTS) approximation with unscented transform (UT) in formulating both the static and dynamic model parameter adaptation. Our new JAC-UT method differentiates itself from other UT-based approaches in that it combines the online noise and channel distortion estimation and model parameter adaptation in a unified UT framework. Experimental results on the standard Aurora 2 task show that the new algorithm enjoys 20.0 % and 16.9 % relative word error rate reductions over the previous JAC-VTS algorithm when using the simple and complex backend models, respectively. Index Terms: unscented transform, vector Taylor series, additive and convolutive distortions, robust ASR, adaptation Jinyu Li 0001, Dong Yu 0001, Yifan Gong 0001, Li Deng 0001 |
INTERSPEECH | 3 |
| 2009 | A study on multilingual acoustic modeling for large vocabulary ASRabstractWe study key issues related to multilingual acoustic modeling for automatic speech recognition (ASR) through a series of large-scale ASR experiments. Our study explores shared structures embedded in a large collection of speech data spanning over a number of spoken languages in order to establish a common set of universal phone models that can be used for large vocabulary ASR of all the languages seen or unseen during training. Language-universal and language-adaptive models are compared with language-specific models, and the comparison results show that in many cases it is possible to build general-purpose language-universal and language-adaptive acoustic models that outperform language-specific ones if the set of shared units, the structure of shared states, and the shared acoustic-phonetic properties among different languages can be properly utilized. Specifically, our results demonstrate that when the context coverage is poor in language-specific training, we can use one tenth of the adaptation data to achieve equivalent performance in cross-lingual speech recognition. Li Deng 0001, Dong Yu 0001, Yifan Gong 0001, Alex Acero, Chin-Hui Lee 0001 |
ICASSP | 4 |
| 2009 | Cross-lingual speech recognition under runtime resource constraintsabstractThis paper proposes and compares four cross-lingual and bilingual automatic speech recognition techniques under the constraint that only the acoustic model (AM) of the native language is used at runtime. The first three techniques fall into the category of lexicon conversion where each phoneme sequence (PHS) in the foreign language (FL) lexicon is mapped into the native language (NL) phoneme sequence. The first technique determines the PHS mapping through the international phonetic alphabet (IPA) features; The second and third techniques are data-driven. They determine the mapping by converting the PHS into corresponding context-independent and context-dependent hidden Markov models (HMMs) respectively and searching for the NL PHS with the least Kullback-Leibler divergence (KLD) between the HMMs. The fourth technique falls into the category of AM merging where the FL's AM is merged into the NL's AM by mapping each senone in the FL's AM to the senone in the NL's AM with the minimum KLD. We discuss the strengths and limitations of each technique developed, report empirical evaluation results on recognizing English utterances with a Korean recognizer, and demonstrate the high correlation between the average KLD and the word error rate (WER). The results show that the AM merging technique performs the best, achieving 60% relative WER reduction over the IPA-based technique. Dong Yu 0001, Li Deng 0001, Jian Wu 0027, Yifan Gong 0001, Alex Acero |
ICASSP | 5 |
| 2009 | A unified framework of HMM adaptation with joint compensation of additive and convolutive distortions
Jinyu Li 0001, Li Deng 0001, Dong Yu 0001, Yifan Gong 0001, Alex Acero |
Comput. Speech Lang. | 4 |
| 2009 | A Novel Framework and Training Algorithm for Variable-Parameter Hidden Markov ModelsabstractWe propose a new framework and the associated maximum-likelihood and discriminative training algorithms for the variable-parameter hidden Markov model (VPHMM) whose mean and variance parameters vary as functions of additional environment-dependent conditioning parameters. Our framework differs from the VPHMM proposed by Cui and Gong (2007) in that piecewise spline interpolation instead of global polynomial regression is used to represent the dependency of the HMM parameters on the conditioning parameters, and a more effective functional form is used to model the variances. Our framework unifies and extends the conventional discrete VPHMM. It no longer requires quantization in estimating the model parameters and can support both parameter sharing and instantaneous conditioning parameters naturally. We investigate the strengths and weaknesses of the model on the Aurora-3 corpus. We show that under the well-matched condition the proposed discriminatively trained VPHMM outperforms the conventional HMM trained in the same way with relative word error rate (WER) reduction of 19% and 15%, respectively, when only mean is updated and when both mean and variances are updated. Dong Yu 0001, Li Deng 0001, Yifan Gong 0001, Alex Acero |
IEEE Trans. Speech Audio Process. | 3 |
| 2008 | HMM adaptation using a phase-sensitive acoustic distortion model for environment-robust speech recognitionabstractIn this paper, we present a new approach to HMM adaptation that jointly compensates for additive and convolutive acoustic distortion in environment-robust speech recognition. The hallmark of our new approach is the use of a nonlinear, phase-sensitive model of acoustic distortion that captures phase asynchrony between clean speech and the mixing noise. In the first step of the developed algorithm, both the static and dynamic portions of the noise and channel parameters are estimated in the cepstral domain, using the speech recognizer’s “feedback” information and the vector-Taylor-series linearization technique on the nonlinear phase-sensitive model. In the second step, the estimated noise and channel parameters are used to effectively adapt the static and dynamic portions of the HMM means and variances also using the linearized phase-sensitive acoustic distortion model. In the experimental evaluation using the standard Aurora 2 task, the proposed new algorithm achieves 93.3% accuracy using the clean-trained complex HMM backend as the baseline system for unsupervised HMM adaptation. This reaches the highest performance number in the literature on this task with clean-trained HMM model. The experimental results show that the phase term, which was missing in all previous HMM-adaptation work, contributes significantly to the achieved high recognition accuracy. Jinyu Li 0001, Li Deng 0001, Dong Yu 0001, Yifan Gong 0001, Alex Acero |
ICASSP | 4 |
| 2008 | Adaptation of compressed HMM parameters for resource-constrained speech recognitionabstractRecently, we successfully developed and reported a new unsupervised online adaptation technique, which jointly compensates for additive and convolutive distortions with vector Taylor series (JAC/VTS), to adjust (uncompressed) HMMs under acoustically distorted environments [1]. In this paper, we extend that technique to adapt compressed HMMs using JAC/VTS where limited computation and/or memory resources are available for speech recognition (e.g., on mobile devices). Subspace coding (SSC) is developed and used to quantize each dimension of the multivariate Gaussians in the compressed HMMs. Three algorithmic design options are proposed and evaluated that combine SSC with JAC/VTS, where three different types of tradeoffs are made between recognition accuracy and the required computation/memory/storage resources. The strengths and weaknesses of these three options are discussed and shown on the Aurora2 task of noise-robust speech recognition. The first option greatly reduces the storage space and gives 93.2% accuracy, which is the same as the baseline accuracy but with little reduction in the run-time computation/memory cost. The second option reduces about 79.9% of the computation cost and about 33.5% of the memory requirement at a very small price of 0.5% decrease of accuracy (to 92.7%). The third option cuts about 89.2% of the computation cost and about 65.5% of the memory requirement while reducing recognition accuracy by 2.7% (to 90.5%). Jinyu Li 0001, Li Deng 0001, Dong Yu 0001, Jian Wu 0027, Yifan Gong 0001, Alex Acero |
ICASSP | 5 |
| 2008 | A minimum-mean-square-error noise reduction algorithm on Mel-frequency cepstra for robust speech recognitionabstractWe present a non-linear feature-domain noise reduction algorithm based on the minimum mean square error (MMSE) criterion on Mel-frequency cepstra (MFCC) for environment-robust speech recognition. Distinguishing from the MMSE enhancement in log spectral amplitude proposed by Ephraim and Malah (E&M) [7], the new algorithm presented in this paper develops the suppression rule that applies to power spectral magnitude of the filter-banks’ outputs and to MFCC directly, making it demonstrably more effective in noise-robust speech recognition. The noise variance in the new algorithm contains a significant term resulting from instantaneous phase asynchrony between clean speech and mixing noise, missing in the E&M algorithm. Speech recognition experiments on the standard Aurora-3 task demonstrate a reduction of word error rate by 48% against the ICSLP02 baseline, by 26% against the cepstral mean normalization baseline, and by 13% against the conventional E&M log-MMSE noise suppressor. The new algorithm is also much more efficient than E&M noise suppressor since the number of the channels in the Mel-frequency filter bank is much smaller (23 in our case) than the number of bins in the FFT domain (256). The results also show that our algorithm performs slightly better than the ETSI AFE on the well-matched and mid-mismatched settings. Dong Yu 0001, Li Deng 0001, Jasha Droppo, Jian Wu 0027, Yifan Gong 0001, Alex Acero |
ICASSP | 5 |
| 2008 | Discriminative training of variable-parameter HMMs for noise robust speech recognitionabstractWe propose a new type of variable-parameter hidden Markov model (VPHMM) whose mean and variance parameters vary each as a continuous function of additional environmentdependent parameters. Different from the polynomialfunction-based VPHMM proposed by Cui and Gong (2007), the new VPHMM uses cubic splines to represent the dependency of the means and variances of Gaussian mixtures on the environment parameters. Importantly, the new model no longer requires quantization in estimating the model parameters and it supports parameter sharing and instantaneous conditioning parameters directly. We develop and describe a growth-transformation algorithm that discriminatively learns the parameters in our cubic-splinebased VPHMM (CS-VPHMM), and evaluate the model on the Aurora-3 corpus with our recently developed MFCC-MMSE noise suppressor applied. Our experiments show that the proposed CS-VPHMM outperforms the discriminatively trained and maximum-likelihood trained conventional HMMs with relative word error rate (WER) reduction of 14 % and 20 % respectively under the well-matched conditions when both mean and variances are updated. Index Terms: speech recognition, variable-parameter hidden Markov model, discriminative training, cubic spline, growth transformation 1. Dong Yu 0001, Li Deng 0001, Yifan Gong 0001, Alex Acero |
INTERSPEECH | 3 |
| 2008 | Parameter clustering and sharing in variable-parameter HMMs for noise robust speech recognitionabstractRecently we proposed a cubic-spline-based variableparameter hidden Markov model (CS-VPHMM) whose mean and variance parameters vary according to some cubic spline functions of additional environment-dependent parameters. We have shown good properties of the CS-VPHMM and demonstrated on the Aurora-3 corpus that MCE-trained CSVPHMM greatly outperforms the MCE-trained conventional HMM at the cost of increased total number of model parameters. In this paper, we propose to share spline functions across different Gaussian mixture components to reduce the total number of model parameters and develop a clustering algorithm to do so. We demonstrate the effectiveness of our parameter clustering and sharing algorithm for the CSVPHMM on Aurora-3 corpus and show that proper parameter sharing can reduce the number of parameters from 4 times of that used in the conventional HMM to 1.13 times and still get 18% relative WER reduction over the MCE trained conventional HMM under the well-matched condition. Effective parameter sharing makes the CS-VPHMM an attractive model for noise robustness. Dong Yu 0001, Li Deng 0001, Yifan Gong 0001, Alex Acero |
INTERSPEECH | 3 |
| 2008 | Robust Speech Recognition Using a Cepstral Minimum-Mean-Square-Error-Motivated Noise SuppressorabstractWe present an efficient and effective nonlinear feature-domain noise suppression algorithm, motivated by the minimum-mean-square-error (MMSE) optimization criterion, for noise-robust speech recognition. Distinguishing from the log-MMSE spectral amplitude noise suppressor proposed by Ephraim and Malah (E&M), our new algorithm is aimed to minimize the error expressed explicitly for the Mel-frequency cepstra instead of discrete Fourier transform (DFT) spectra, and it operates on the Mel-frequency filter bank's output. As a consequence, the statistics used to estimate the suppression factor become vastly different from those used in the E&M log-MMSE suppressor. Our algorithm is significantly more efficient than the E&M's log-MMSE suppressor since the number of the channels in the Mel-frequency filter bank is much smaller (23 in our case) than the number of bins (256) in DFT. We have conducted extensive speech recognition experiments on the standard Aurora-3 task. The experimental results demonstrate a reduction of the recognition word error rate by 48% over the standard ICSLP02 baseline, 26% over the cepstral mean normalization baseline, and 13% over the popular E&M's log-MMSE noise suppressor. The experiments also show that our new algorithm performs slightly better than the ETSI advanced front end (AFE) on the well-matched and mid-mismatched settings, and has 8% and 10% fewer errors than our earlier SPLICE (stereo-based piecewise linear compensation for environments) system on these settings, respectively. Dong Yu 0001, Li Deng 0001, Jasha Droppo, Jian Wu 0027, Yifan Gong 0001, Alex Acero |
IEEE Trans. Speech Audio Process. | 5 |
| 2007 | High-performance hmm adaptation with joint compensation of additive and convolutive distortions via Vector Taylor SeriesabstractIn this paper, we present our recent development of a model-domain environment-robust adaptation algorithm, which demonstrates high performance in the standard Aurora 2 speech recognition task. The algorithm consists of two main steps. First, the noise and channel parameters are estimated using a nonlinear environment distortion model in the cepstral domain, the speech recognizer’s “feedback” information, and the Vector-Taylor-Series (VTS) linearization technique collectively. Second, the estimated noise and channel parameters are used to adapt the static and dynamic portions of the HMM means and variances. This two-step algorithm enables Joint compensation of both Additive and Convolutive distortions (JAC). In the experimental evaluation using the standard Aurora 2 task, the proposed JAC/VTS algorithm achieves 91.11% accuracy using the clean-trained simple HMM backend as the baseline system for the model adaptation. This represents high recognition performance on this task without discriminative training of the HMM system. Detailed analysis on the experimental results shows that adaptation of the dynamic portion of the HMM mean and variance parameters is critical to the success of our algorithm. Jinyu Li 0001, Li Deng 0001, Dong Yu 0001, Yifan Gong 0001, Alex Acero |
ASRU | 4 |
| 2007 | A Study of Variable-Parameter Gaussian Mixture Hidden Markov Modeling for Noisy Speech RecognitionabstractTo improve recognition performance in noisy environments, multicondition training is usually applied in which speech signals corrupted by a variety of noise are used in acoustic model training. Published hidden Markov modeling of speech uses multiple Gaussian distributions to cover the spread of the speech distribution caused by noise, which distracts the modeling of speech event itself and possibly sacrifices the performance on clean speech. In this paper, we propose a novel approach which extends the conventional Gaussian mixture hidden Markov model (GMHMM) by modeling state emission parameters (mean and variance) as a polynomial function of a continuous environment-dependent variable. At the recognition time, a set of HMMs specific to the given value of the environment variable is instantiated and used for recognition. The maximum-likelihood (ML) estimation of the polynomial functions of the proposed variable-parameter GMHMM is given within the expectation-maximization (EM) framework. Experiments on the Aurora 2 database show significant improvements of the variable-parameter Gaussian mixture HMMs compared to the conventional GMHMMs Yifan Gong 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Modeling Variance Variation in a Variable Parameter HMM Framework for Noise Robust Speech RecognitionabstractVariance variation with respect to a continuous environment-department variable is investigated in this paper in a variable parameter Gaussian mixture HMM (VP-GMHMM) for noisy speech recognition. The variation is modeled by a scaling polynomial applied to the variances in the conventional hidden Markov acoustic models. The maximum likelihood estimation of the scaling polynomial is performed under an SNR quantization approximation. Experiments on the Aurora 2 database show significant improvements by incorporating the variance scaling scheme into the previous VP-GMHMM where only mean variation is considered. Yifan Gong 0001 |
ICASSP (1) | 2 |
| 2005 | A Method of Joint Compensation of Additive and Convolutive Distortions for Speaker-Independent Speech RecognitionabstractA speech recognizer operating in a mobile environment has to be robust to two distortion sources: ambient noise (additive distortion) and microphone changes (convolutive distortion). Explicitly and simultaneously modeling the two distortion sources has been a great challenge for speech recognition in adverse environments. In this paper, two log-spectral domain components are introduced in speech acoustic models to represent additive and convolutive distortions. A method, called JAC, jointly compensates both additive and convolutive distortions. For each utterance to be recognized, it adapts HMM mean vectors with a noise estimate and a channel estimate. The noise estimate is calculated from the pre-utterance pause and the channel estimate is calculated using an EM algorithm from speech utterances produced in the distortion environment. The algorithm is evaluated on a noisy speech database recorded in-vehicle with a hands-free distant microphone in several sessions, including parked, stop-and-go, and highway driving conditions. Experiments show that the method typically reduces recognition word error rate by an order of magnitude. The method makes it possible to obtain high performance for speaker-independent recognition in changing noisy environments without collecting any noisy speech for training. Yifan Gong 0001 |
IEEE Trans. Speech Audio Process. | 1 |
| 2004 | Can back-ends be more robust than front-ends? Investigation over the Aurora-2 databaseabstractWe present a back-end solution developed at Texas Instruments for noise robust speech recognition. The solution consists of three techniques: 1) a joint additive and convolutive noise compensation (JAC) which adapts speech acoustic models; 2) an enhanced channel estimation procedure which extends JAC performance towards lower SNR ranges; 3) an N-pass decoding algorithm. The performance of the proposed back-end is evaluated on the Aurora-2 database. With 20% fewer model parameters and without the need for the second order derivative of the recognition features, the performance of the proposed solution is 91.86%, which outperforms that of the ETSI advanced front-end standard (88.19%) by more than 30% relative word error rate reduction. Alexis Bernard, Yifan Gong 0001 |
ICASSP (1) | 2 |
| 2003 | Variable parameter Gaussian mixture hidden Markov modeling for speech recognitionabstractTo improve recognition, speech signals corrupted by a variety of noises can be used in speech model training. Published hidden Markov modeling of speech uses multiple Gaussian distributions to cover the spread of the speech distribution caused by the noises, which distracts the modeling of speech event itself and and possibly sacrifices the performance on clean speech. We extend GMHMM by allowing state emission parameters to change as function of an environment-dependent continuous variable. At the recognition time, a set of HMMs specific to the given the environment is instantiated and used for recognition. Variable parameter (VP) HMM with parameters modeled as a polynomial function of the environment variable is developed. Parameter estimation based on EM-algorithm is given. With the same number of mixtures, VPHMM reduces WER by 40% compared to conventional multi-condition training. Yifan Gong 0001 |
ICASSP (1) | 2 |
| 2003 | Model-space compensation of microphone and noise for speaker-independent speech recognitionabstractAmbient noise (additive distortion) and microphone changes (convolutive distortion) are two sources of distortion that may severely degrade speech recognition performance in real operating environments. Simultaneously modeling the two distortion sources has been a great challenge for robust speech recognition. A method, called JAC (Joint compensation of Additive and Convolutive distortions), is presented. It uses two log-spectral domain components in speech acoustic models to represent additive and convolutive distortions. The method adapts HMM mean vectors with a noise estimate and a channel estimate. The noise estimate is calculated from the pre-utterance pause and the channel estimate is calculated using an EM algorithm from speech utterances produced in the distortion environment. Evaluated on a noisy speech database recorded in-vehicle with a hands-free distant microphone in several driving conditions, the algorithm reduces recognition word error rate in typical operating conditions by an order of magnitude. Yifan Gong 0001 |
ICASSP (1) | 1 |
| 2002 | Noise-robust open-set speaker recognition using noise-dependent Gaussian mixture classifierabstractSpeaker recognition makes a decision to either accept or reject a recognized speaker candidate, based on some score (e.g. likelihood) associated to the item. Model-based classification can be used to make the decision. In mobile device applications, the background noise level may affect the score distributions and cause a decision failure. We describe a new decision procedure, which treats the scores as the outcome of Gaussian mixture distributions, where mean and covariance parameters are modeled as polynomial functions of noise level. We evaluate the procedure on a speaker recognition task in a mobile and noisy environment, using a hands-free microphone. Experiments show that the system delivers an equal error rate of 0.30%, 0.80% and 3.53% for parked, stop-and-go and highway driving conditions. The method maintains a balance between false acceptance and false rejection under all driving conditions, making any empirical threshold adjustment unnecessary. Yifan Gong 0001 |
ICASSP | 1 |
| 2002 | A comparative study of approximations for parallel model combination of static and dynamic parameters
Yifan Gong 0001 |
INTERSPEECH | 1 |
| 2002 | Experiments on speaker-independent voice command recognition using in-vehicle hands free speech
Yifan Gong 0001, Lorin Netsch |
INTERSPEECH | 1 |
| 2002 | The effects of speech compression on speech recognition and text-to-speech synthesis
Yeshwant K. Muthusamy, Yifan Gong 0001, Roshan Gupta |
INTERSPEECH | 2 |
| 2002 | Noise-dependent Gaussian mixture classifiers for robust rejection decisionabstractSpeech or speaker recognizers need to make a decision on either accepting or rejecting a recognized item, based on some measurement (e.g., likelihood) associated to the item. Distribution-based classification can be used to make the decision. In practical applications, the background noise level may adversely affect the distributions of the likelihoods and cause classification failure. A new decision mechanism is described, which treats the likelihoods as outcome of multidimensional Gaussian distributions with noise-dependent mean and covariance. The dependence on the noise is explicitly modeled as a polynomial function of noise level. The steps of estimating the decision parameters using the EM algorithm are given. Experimental results on in-car speech data show that the procedure, for noise ranging from a parked car (/spl sim/30 dB SNR) to highway (/spl sim/0 dB SNR) driving conditions, maintains a well-balanced decision performance between false rejection and false acceptance. Yifan Gong 0001 |
IEEE Trans. Speech Audio Process. | 1 |
| 2000 | Implementing a high accuracy speaker-independent continuous speech recognizer on a fixed-point DSPabstractContinuous speech recognition is a resource-intensive algorithm. Commercial dictation software requires more than 10 Mbytes to install on the disk and 32 Mbytes RAM to run the application. A typical embedded system can not afford this much RAM because of its high cost and power consumption; it also lacks disk to store the large amount of static data (e.g. acoustic models). We have been working on optimization of a small vocabulary speech recognizer suitable for implementation on a 16-bit fixed-point DSP. This recognizer supports sophisticated continuous density, tied-mixtures Gaussians, parallel model combination, and a noise-robust utterance detection algorithm. The fixed-point version achieves the same performance as the floating-point version. The algorithm runs real-time on a 100 MHz, 16-bit, fixed-point Texas Instruments TMS320C5410 even for the most challenging continuous digit dialing with hands-free microphone in driving conditions. Yifan Gong 0001, Yu-Hung Kao |
ICASSP | 1 |
| 2000 | HMM adaptation and microphone array processing for distant speech recognitionabstractConnected strings of seven digits from the TIDIGITS database were recorded in a reverberant office room for evaluation using microphone array processing and HMM, hidden Markov model, adaptation. A sixteen-channel linear microphone array records a distance speech database useful for further experimentation. The adaptation techniques of parallel model combination (PMC) and maximum likelihood linear regression (MLLR) are evaluated and compared. The effect of the number of adaptation utterances and number of vectors per class for the regression tree in order to optimize MLLR results are studied. Results show, compared to no adaptation, 40% word error reduction (improvement to 4.2%) for PMC and 60% word error reduction (improvement to 3.0%) for MLLR. Jim Kleban, Yifan Gong 0001 |
ICASSP | 2 |
| 1999 | Transforming HMMs for speaker-independent hands-free speech recognition in the carabstractIn the absence of HMMs trained with speech collected in the target environment, one may use HMMs trained with a large amount of speech collected in another recording condition (e.g., quiet office, with high quality microphone). However, this may result in poor performance because of the mismatch between the two acoustic conditions. We propose a linear regression-based model adaptation procedure to reduce such a mismatch. With some adaptation utterances collected for the target environment, the procedure transforms the HMMs trained in a quiet condition to maximize the likelihood of observing the adaptation utterances. The transformation must be designed to maintain speaker-independence of the HMM. Our speaker-independent test results show that with this procedure about 1% digit error rate can be achieved for hands-free recognition, using target environment speech from only 20 speakers. Yifan Gong 0001, John J. Godfrey |
ICASSP | 1 |
| 1999 | Speech-enabled information retrieval in the automobile environmentabstractWith the advances in speech recognition and wireless communications, the possibilities for information access in the automobile have expanded significantly. We describe four system prototypes for (i) voice-dialing, (ii) Internet information retrieval-called InfoPhone, (iii) voice e-mail, and (iv) car navigation. These systems are designed primarily for hands-busy, eyes-busy conditions, use speaker-independent speech recognizers, and can be used with a restricted display or no display at all. The voice-dialing prototype incorporates our hands-free speech recognition engine that is very robust in noisy car environments (1% WER and 3% string error rate on the continuous digit recognition task at 0 db SNR). The InfoPhone, voice e-mail, and car navigation prototypes use a client-server architecture with the client designed to be resident on a phone or other hand-held device. Yeshwant K. Muthusamy, Rajeev Agarwal, Yifan Gong 0001, Vishu Viswanathan |
ICASSP | 3 |
| 1999 | Speaker-dependent name dialing in a car environment with out-of-vocabulary rejectionabstractWe describe a system for name dialing in the car and present results under three driving conditions using real-life data. The names are enrolled in the parked car condition (engine off) and we describe two approaches for endpointing them-energy-based and recognition-based schemes-which result in word-based and phone-based models, respectively. We outline a simple algorithm to reject out-of-vocabulary names. PMC is used for noise compensation. When tested on an internally collected twenty-speaker database, for a list size of 50 and a hand-held microphone, the performance averaged over all driving conditions and speakers was 98%/92% (IV accuracy/OOV rejection); for the hands-free data, it was 98%180%. Coimbatore S. Ramalingam, Yifan Gong 0001, Lorin Netsch, Wallace W. Anderson, John J. Godfrey, Yu-Hung Kao |
ICASSP | 2 |
| 1999 | A minimum cross-entropy approach to hidden Markov model adaptationabstractAn adaptation algorithm using the theoretically optimal maximum a posteriori (MAP) formulation, and at the same time accounting for parameter correlation between different classes is desirable, especially when using sparse adaptation data. However, a direct implementation of such an approach may be prohibitive in many practical situations. We present an algorithm that approximates the above mentioned correlated MAP algorithm by iteratively maximizing the set of posterior marginals. With some simplifying assumptions, expressions for these marginals are then derived, using the principle of minimum cross-entropy. The resulting algorithm is simple, and includes conventional MAP estimation as a special case. The utility of the proposed method is tested in adaptation experiments for an alphabet recognition task. Mohamed Afify, Yifan Gong 0001, Jean Paul Haton |
IEEE Signal Process. Lett. | 2 |
| 1998 | Environment normalization training and environment adaptation using mixture stochastic trajectory model
Irina Illina, Mohamed Afify, Yifan Gong 0001 |
Speech Commun. | 3 |
| 1998 | Assessing the importance of the segmentation probability in segment-based speech recognition
Jan P. Verhasselt, Irina Illina, Jean-Pierre Martens, Yifan Gong 0001, Jean Paul Haton |
Speech Commun. | 4 |
| 1998 | A general joint additive and convolutive bias compensation approach applied to noisy Lombard speech recognitionabstractA unified approach to the acoustic mismatch problem is proposed. A maximum likelihood state-based additive bias compensation algorithm is developed for the continuous density hidden Markov model (CDHMM). Based on this technique, specific bias models in the mel cepstral and the linear spectral domains are presented. Among these models, a new polynomial trend bias model in the mel cepstral domain is derived, which proved effective for Lombard speech compensation. In addition, a joint estimation algorithm for additive and convolutive bias compensation is proposed. This algorithm is based on applying the expectation maximization (EM) technique in both above-mentioned domains, in conjunction with a parallel model combination (PMC) based transformation. The compensation of the dynamic (difference) coefficients in the proposed framework is also studied. The evaluation data base consists of a 21 confusable word vocabulary uttered by 24 speakers. Three mismatched versions of the data base are considered, i.e., Lombard speech, 15 dB noisy Lombard speech, and 5 dB noisy Lombard speech. The proposed techniques result in 50.9%, 74.6%, and 67.3% reduction in the performance difference between matched and uncompensated word error rates for the three mismatch conditions, respectively. When dynamic coefficients are considered the corresponding reductions are 46.8%, 72.4%, and 70.9%. Mohamed Afify, Yifan Gong 0001, Jean Paul Haton |
IEEE Trans. Speech Audio Process. | 2 |
| 1997 | A unified maximum likelihood approach to acoustic mismatch compensation: application to noisy Lombard speech recognitionabstractIn the context of continuous density hidden Markov model (CDHMM) we present a unified maximum likelihood (ML) approach to acoustic mismatch compensation. This is achieved by introducing additive Gaussian biases at the state level in both the mel cepstral and linear spectral domains. Flexible modelling of different mismatch effects can be obtained through appropriate bias tying. A maximum likelihood approach for joint estimation of both mel cepstral and linear spectral biases from the observed mismatched speech given only one set of clean speech models is presented, where the obtained bias estimates are used for the compensation of clean speech models during decoding. The proposed approach is applied to the recognition of noisy Lombard speech, and significant improvement in the word recognition rate is achieved. Mohamed Afify, Yifan Gong 0001, Jean Paul Haton |
ICASSP | 2 |
| 1997 | Elimination of trajectory folding phenomenon: HMM, trajectory mixture HMM and mixture stochastic trajectory modelabstractIn this paper, a study of topology of hidden Markov model (HMM) used in speech recognition is addressed. Our main contribution is the introduction of the notion of trajectory folding phenomenon of HMM. In complex phonetic contexts and in speaker-variability, this phenomenon degrades the discriminability of HMM. The goal of this paper is to give some explanation and experimental evidence suggesting the existence of this phenomenon. The systems eliminating (partially or entirely) the trajectory folding are HMM with a special topology, called trajectory mixture HMM (TMHMM), and a mixture stochastic trajectory model (MSTM), proposed recently. HMM, TMHMM and MSTM have been tested on a 1011 words vocabulary, speaker dependent and multi-speaker continuous French speech recognition task. With similar number of model parameters, TMHMM and MSTM cuts down the error rate produced by the HMM, which confirms our hypothesis. Irina Illina, Yifan Gong 0001 |
ICASSP | 2 |
| 1997 | The importance of segmentation probability in segment based speech recognizersabstractIn segment based recognizers, variable length speech segments are mapped to the basic speech units (phones, diphones, ...). We address the acoustical modeling of these basic units in the framework of segmental posterior distribution models (SPDM). The joint posterior probability of a unit sequence u_ and a segmentation s_, Pr(u_,s_|X_) can be written as the product of the segmentation probability Pr(s_|X_) and the unit classification probability Pr(u_|s_,X_), where X_ is the sequence of acoustic observation parameter vectors. In particular, we point out the role of the segmentation probability and demonstrate that it does improve the recognition accuracy. We present evidence for this in two different tasks (speaker dependent continuous word recognition in French and speaker independent phone recognition in American English) in combination with two different unit classification models. Jan P. Verhasselt, Irina Illina, Jean-Pierre Martens, Yifan Gong 0001, Jean Paul Haton |
ICASSP | 4 |
| 1997 | Correlation based predictive adaptation of hidden Markov models
Mohamed Afify, Yifan Gong 0001, Jean Paul Haton |
EUROSPEECH | 2 |
| 1997 | An acoustic subword unit approach to non-linguistic speech feature identification
Mohamed Afify, Yifan Gong 0001, Jean Paul Haton |
EUROSPEECH | 2 |
| 1997 | Source normalization training for HMM applied to noisy telephone speech recognition
Yifan Gong 0001 |
EUROSPEECH | 1 |
| 1997 | Speaker normalization training for mixture stochastic trajectory modelabstractIn this paper we are interested in speaker and environment adaptation techniques for speaker independent (SI) continuous speech recognition. These techniques are used to reduce mismatch between training and the testing conditions, using a small amount of adaptation data. In addition to reducing this mismatch during the adaptation, we propose to reduce the variation due to speakers or environments during the training itself in the context of Speaker Normalisation (SN) approach, using MLLR transformation. SN also includes a combination of the context-dependent, phone dependent and broad phonetic class dependent information. The use of linear regression to model broad phonetic class dependent information assures our model to be used in the case that the adaptation data or training data is not given for some phonetic symbols. SN is developed for Mixture Stochastic Trajectory Model, a segment based model. The approach can be used for speaker, gender or environment normalization. We show the performance of SN compared to SI recognition and to MLLR speaker adaptation, through experiments on continuous speech recognition. Irina Illina, Yifan Gong 0001 |
EUROSPEECH | 2 |
| 1997 | Stochastic trajectory modeling and sentence searching for continuous speech recognitionabstractThe paper first points out a defect in hidden Markov modeling (HMM) of continuous speech, referred as trajectory folding phenomenon. A new approach to modeling phoneme-based speech units is then proposed, which represents the acoustic observations of a phoneme as clusters of trajectories in a parameter space. The trajectories are modeled by a mixture of probability density functions of a random sequence of states. Each state is associated with a multivariate Gaussian density function, optimized at the state sequence level. Conditional trajectory duration probability is integrated in the modeling. An efficient sentence search procedure based on trajectory modeling is also formulated. Experiments with a speaker-dependent, 2010-word continuous speech recognition application with a word-pair perplexity of 50, using vocabulary-independent acoustic training, monophone models trained with 80 sentences per speaker, reported about a 1% word error rate. The new models were experimentally compared to continuous density mixture HMM (CDHMM) on the same recognition task, and gave significantly smaller word error rates. These results suggest that the stochastic trajectory model provides a more in-depth modeling of continuous speech signals. Yifan Gong 0001 |
IEEE Trans. Speech Audio Process. | 1 |
| 1996 | Probabilistic mapping networks for speaker recognitionabstractThe expectation-maximization (EM) algorithm is a general technique for maximum likelihood estimation (MLE). In this paper, we present two important theoretical issues concerning Gaussian mixture modeling (GMM) within the EM framework. First, we propose an EM algorithm for estimating the parameters of a GMM structure dedicated to speaker recognition, the probabilistic mapping network (PMN), where the Gaussian probability density function is realized as an internal node. Hence, the EM algorithm is extended to deal with the supervised learning of a multicategory classification problem and serves as a parameter estimator of the neural network classifier. Then, a generalized EM (GEM) algorithm is developed as an alternative to the MLE problem of PMN. The effectiveness of the proposed PMN architecture and developed EM algorithms are assessed by conducting a set of speaker recognition experiments. It is shown that GEM converges faster than EM to the same solution space. Haizhou Li 0001, Yifan Gong 0001, Jean Paul Haton |
ICASSP | 2 |
| 1996 | A semi-continuous stochastic trajectory model for phoneme-based continuous speech recognitionabstractWe propose a model of phoneme-based speech unit, called semi-continuous stochastic trajectory model (SC-STM), which generalizes our stochastic trajectory models (STM). As STMs, the SC-STMs focus on the modeling of speech segments (called trajectories) in their parameter space, and can therefore handle segmental information, which is critical for large vocabulary continuous speech recognition. Compared to the STMs, the SC-STMs improve the resolution of the trajectory modeling, while keeping a moderate number of free parameters by sharing state probability density functions. The SC-STM can therefore maintain a good trade-off between detailed acoustic modeling and limited training data. We tested the idea on a 2010 words, speaker-dependent, continuous speech database. Preliminary results show that SC-STM gives a word accuracy close to that of STM, without using heuristic techniques that enhanced STM. Olivier Siohan, Yifan Gong 0001 |
ICASSP | 2 |
| 1996 | Modelling long term variability information in mixture stochastic trajectory framework
Yifan Gong 0001, Irina Illina, Jean Paul Haton |
ICSLP | 1 |
| 1996 | Stochastic trajectory model with state-mixture for continuous speech recognition
Irina Illina, Yifan Gong 0001 |
ICSLP | 2 |
| 1996 | Improvement in n-best search for continuous speech recognition
Irina Illina, Yifan Gong 0001 |
ICSLP | 2 |
| 1996 | A study on continuous Chinese speech recognition based on stochastic trajectory models
Yifan Gong 0001, Yuqing Fu, Jiren Lu, Jean Paul Haton |
ICSLP | 2 |
| 1996 | Estimation of mixtures of stochastic dynamic trajectories: application to continuous speech recognition
Mohamed Afify, Yifan Gong 0001, Jean Paul Haton |
Comput. Speech Lang. | 2 |
| 1996 | Comparative experiments of several adaptation approaches to noisy speech recognition using stochastic trajectory models
Olivier Siohan, Yifan Gong 0001, Jean Paul Haton |
Speech Commun. | 2 |
| 1995 | Stochastic trajectory modeling for recognition of unconstrained handwritten wordsabstractIn this paper we describe an off-line handwritten word recognition system applied to the identification of literal french check amounts. It consists of three successive levels denoted as character, word and phrase level, each of them being related to the previous ones via conditional probability distributions. Training is done on character samples extracted from amount images which are modeled as trajectories in some feature space. At word level, guided by a dictionary, an internal character segmentation algorithm is used in order to maximize a global word probability measure. A stochastic grammar for a priori grammar generation probability of a phrase is proposed at the last level. Results obtained on a 1779 amounts data base provided by the SRTP are encouraging, showing our system open to further improvements. George Saon, Abdel Belaïd, Yifan Gong 0001 |
ICDAR | 3 |
| 1995 | Stochastic trajectory models for speech recognition: an extension to modelling time correlation
Mohamed Afify, Yifan Gong 0001, Jean Paul Haton |
EUROSPEECH | 2 |
| 1995 | Evaluation of Bayes decision approach to automatic determination of thresholds for speaker verification
Yifan Gong 0001 |
EUROSPEECH | 1 |
| 1995 | On MMI learning of Gaussian mixture for speaker models
Haizhou Li 0001, Jean Paul Haton, Yifan Gong 0001 |
EUROSPEECH | 3 |
| 1995 | Speaker recognition with temporal transition models
Haizhou Li 0001, Jean Paul Haton, Jian Su 0002, Yifan Gong 0001 |
EUROSPEECH | 4 |
| 1995 | Noise adaptation using linear regression for continuous noisy speech recognition
Olivier Siohan, Yifan Gong 0001, Jean Paul Haton |
EUROSPEECH | 2 |
| 1995 | Speech recognition in noisy environments: A survey
Yifan Gong 0001 |
Speech Commun. | 1 |
| 1994 | Stochastic trajectory modeling for speech recognitionabstractModels observations of phoneme-based speech units as clusters of trajectories in their parameter space. The trajectories are modeled by a mixture of state sequences of multi-variate Gaussian density functions, optimized at the state sequence level. The duration of trajectories are integrated in the modeling. The authors also provide an algorithm for sentence recognition based on the modeling. In an alphabet recognition task the resulting system trained in context-independent mode demonstrated substantially better recognition accuracy, compared to a conventional context-dependent, whole word HMM.> Yifan Gong 0001, Jean Paul Haton |
ICASSP (1) | 1 |
| 1994 | Noise independent speech recognition for a variety of noise typesabstractBy a base transformation technique, we previously reported a recognizer which gives a noise-adapted recognition rate of 90% under 10 dB SNR on a vocabulary of 206 words. This rate is 97% of the recognition rate for clean speech. The technique is extended here so that the input noise is first recognized as one of a set reference noises, and the noise reference is used for the base transformation of the noisy utterance. Using 32 reference noise classes, for speech signals corrupted by noises of unknown natures (either Gaussian, bus or aircrafts with SNR randomly from 10 to 40 dB), we obtained a noise-independent recognition rate of about 95.5% of the recognition rate for clean speech.> William C. Treurniet, Yifan Gong 0001 |
ICASSP (1) | 2 |
| 1994 | Nonlinear time alignment in stochastic trajectory models for speech recognition
Mohamed Afify, Yifan Gong 0001, Jean Paul Haton |
ICSLP | 2 |
| 1994 | A comparison of three noisy speech recognition approaches
Olivier Siohan, Yifan Gong 0001, Jean Paul Haton |
ICSLP | 2 |
| 1993 | Base transformation for environment adaptation in continuous speech recognition
Yifan Gong 0001 |
EUROSPEECH | 1 |
| 1993 | Iterative transformation and alignment for speech labeling
Yifan Gong 0001, Jean Paul Haton |
EUROSPEECH | 1 |
| 1993 | Duration of phones as function of utterance length and its use in automatic speech recognitionabstract3rd european conference on speech communication and technology, Berlin, Germany, 21-23 September 1993 - 1 vol Yifan Gong 0001, William C. Treurniet |
EUROSPEECH | 1 |
| 1993 | Use of explicit context-dependent phonemic model in continuous speech recognition
Feriel Mouria, Yifan Gong 0001, Jean Paul Haton |
EUROSPEECH | 2 |
| 1993 | A Bayesian approach to phone duration adaptation for lombard speech recognition
Olivier Siohan, Yifan Gong 0001, Jean Paul Haton |
EUROSPEECH | 2 |
| 1993 | Plausibility functions in continuous speech recognition: The VINICS system
Yifan Gong 0001, Jean Paul Haton |
Speech Commun. | 1 |
| 1992 | Nonlinear vectorial interpolation for speaker recognitionabstractThe authors address the problem of speaker recognition using very short utterances, both for training and for recognition. The authors propose to exploit speaker-specific correlations between two suitably defined parameter vector sequences. A nonlinear vectorial interpolation technique is used to capture speaker-specific information, through least-square-error minimization. The experiments show the feasibility of recognizing a speaker among a population of about 100 persons using only an utterance of one word both for training and for recognition.> Yifan Gong 0001, Jean Paul Haton |
ICASSP | 1 |
| 1992 | Hand-written text recognition based on a new formulationabstractThe hand-written word recognition problem is formulated in two steps. In the first step, the plausibility of observing each character is computed as a function of sample index of a line. In the second step, word recognition is achieved by finding the word (sequence of characters) which maximizes the sum of plausibilities of individual characters which make up the word. The authors propose an efficient algorithm for the second step which makes use of peaks of plausibility functions and solves the maximization process by two embedded search processes: finding the best path connecting peaks of the plausibility functions of two successive characters, and finding the best transition sample index for two given peaks. In a preliminary experimentation, using template matching for character plausibility estimation, a 96% recognition rate was obtained for a cursive-writing-like font.> Yifan Gong 0001, Anne Boyer |
ICPR (2) | 1 |
| 1992 | DTW-based phonetic labeling using explicit phoneme duration constraintsabstractcommunication to : Proc. Intern. Conf. on Spoken Language Processing, Banff (Alberta, Canada), October 1992 Yifan Gong 0001, Jean Paul Haton |
ICSLP | 1 |
| 1992 | Minimization of speech alignment error by iterative transformation for speaker adaptationabstractExtrait de : Proc. International Conference on Spoken Language Processing, Banff (Alberta, Canada), October 1992 Yifan Gong 0001, Olivier Siohan, Jean Paul Haton |
ICSLP | 1 |
| 1991 | Neural network coupled with IIR sequential adapter for phoneme recognition in continuous speechabstractThe authors present an NN-IIR (neural network/infinite impulse response filter) system for phoneme recognition in continuous speech based on the idea of modeling the recognition process by state evolution and interpretation equations. This work gives a solution to temporal information representation in phoneme recognition using neural networks and recursive filters, yielding better recognition results for continuous speech. This recognition system has two promising properties, i.e., capabilities for dealing with sequential properties and for interpreting speech signals by means of a training process. It was shown experimentally that the NN-IIR network obtained good performance for continuous speech recognition. Preliminary experiments with limited training data indicate that the NN-IIR provides good discrimination power for plosives, which are highly context-dependent.> Yifan Gong 0001, Jean Paul Haton |
ICASSP | 1 |
| 1991 | Non-linear vector interpolation by neural network for phoneme identification in continuous speechabstractThe correlations between vectors in a sequence of analysis frames are supposed to be specific to phonetic units in acoustic-phonetic decoding of speech. The authors propose nonlinear vector interpolation techniques to represent this correlation and to recognize phonemes. The interpolation is based on the decomposition of a frame sequence into two parts and on the construction of a function that interpolates one part using information from the second part. According to quantities to be interpolated, three families of interpolator models are developed. In a recognition system, each phonemic symbol is associated with a nonlinear vector interpolator which is trained to give minimum interpolation error for that specific phoneme. Multilayer feedforward neural networks are used to implement the nonlinear vector interpolators. For continuous speech under the phoneme spotting test using 16 PLCC-derived cepstrum coefficients as parametric vectors, the three categories of models gave compatible results.> Yifan Gong 0001, Jean Paul Haton |
ICASSP | 1 |
| 1991 | Continuous speech recognition based on high plausibility regionsabstractThe authors propose an approach to phoneme-based continuous speech recognition when a time function of the plausibility of observing each phoneme (spotting result) is given. They introduce a criterion for the best sentence, based on the sum of plausibilities of individual symbols composing the sentence. Based on the idea of making use of high plausibility regions to reduce the computational load while maintaining optimality, the method finds the most plausible sentences relating to the input speech. Two optimization procedures are defined to deal with the following embedded search processes: (1) finding the best path connecting peaks of the plausibility functions of two successive symbols, and (2) finding the best time transition slot index for two given peaks. Experimental results show that the method gives better recognition precision while requiring about 1/20 of the computing time of the traditional DP-based methods. The experimental system obtained a 95% sentence recognition rate on a multispeaker test.> Yifan Gong 0001, Jean Paul Haton, Feriel Mouria |
ICASSP | 1 |
| 1991 | Comparing two phoneme identification methods using a continuous speech recognizer
Yifan Gong 0001, Jean Paul Haton |
EUROSPEECH | 1 |
| 1991 | VINICS: a continuous speech recognizer based on a new robust formulation
Yifan Gong 0001, Jean Paul Haton |
EUROSPEECH | 1 |
| 1991 | Signal-to-String Conversion Based on High Likelihood Regions Using Embedded Dynamic ProgrammingabstractA method of signal-to-string conversion based on embedded dynamic programming (DP) which can adapt its search to the variation of the input signal is proposed. The optimizing process is guided by high-valued portions of the likelihood function of symbols composing the string and is solved by two embedded dynamic programming processes. Algorithms in a Pascal-like language relating to the solution are given. When applied to continuous speech recognition on a 100-word vocabulary using the phoneme as the basic recognition unit, the method is shown to achieve a 4% improvement in the recognition rate compared to a classical DP-based method.> Yifan Gong 0001, Jean Paul Haton |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1990 | Text-independent speaker recognition by trajectory space comparisonabstractThe principle of trajectory space comparison for text-independent speaker recognition and some solutions to the space comparison problem based on vector quantization are presented. The comparison of the recognition rates of different solutions is reported. The experimental system achieves a 99.5% text-independent speaker recognition rate for 23 speakers, using five phrases for training and five for test. A speaker-independent continuous speech recognition system is built in which this principle is used for speaker adaptation.> Yifan Gong 0001, Jean Paul Haton |
ICASSP | 1 |
| 1990 | Towards a general signal interpretation system-signal-to-symbol conversion levelabstractThe signal-to-symbol conversion level of a generic signal interpretation system is presented. Using an application-independent structure, the system can compile user-supplied application-specific task descriptions to generate programs capable of performing statistic and structural pattern recognition, and reasoning in a given application domain. The computation of three types of symbols, quantitative symbols, qualitative symbols, and compound symbols, is outlined. The notion of context of symbols and its use in the system are discussed. In this system, symbolic processing, signal processing routine, and similarity measure can be specified in terms of situations. An evaluation of the system performance in an application of nondestructive testing of a steam generator using eddy current signals is given.> Yifan Gong 0001, Jean Paul Haton |
ICPR (2) | 1 |
| 1990 | A multiknowledge base system for continuous speech understandingabstractA spoken Chinese understanding system which accepts continuous speech and produces Lisp-executable functions is presented. A multiple-knowledge-source organization model, syntactic structure construction by forward-backward inference, a phoneme model of speech profiles, and knowledge-based tone classification are proposed and implemented in order to remove the uncertainty of speech signals. Experimental results are given and discussed. In the single-speaker mode, 90% of the test sentences were recognized as first proposition by the system. 3% were rejected, and the rest were given within the five propositions of best quality.> Yifan Gong 0001, Jean Paul Haton |
ICPR (2) | 1 |
| 1989 | Parallel construction of syntactic structure for continuous speech recognitionabstractPublie dans : Proceedings EUROSPEECH 89 (European conference on speech communication and technology), Paris, September 1989 Yifan Gong 0001, Anne Boyer, Jean Paul Haton |
EUROSPEECH | 1 |
| 1988 | A specialist society for continuous speech understandingabstractThe authors propose a homogeneous architecture for organizing and controlling multiple knowledge sources in signal interpretation systems based on the decomposition of the interpretation problem into subproblems at successive conceptual levels. Partial interpretations and knowledge about the problem are partitioned into associations each of which consisting of several specialists. Information exchange is assured by a common data structure within knowledge sources at each given level and by a message passing mechanism between two different levels. Different strategies adapted to the nature of the problem may be implemented at each level. The interpretation process consists in executing the groups of knowledge sources in multiple phases controlled by data and models in each concept domain and allowing incremental construction of solutions. The authors illustrate the architecture by an application to continuous spoken Chinese understanding.> Yifan Gong 0001, Jean Paul Haton |
ICASSP | 1 |