Frank Seide

dblp:75/4172 · also Frank Torsten Bernd Seide · DBLP profile ↗
← Back
85ranked-venue papers
18as first author
13since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 75 · 16 first-author · 13 since 2021Artificial intelligence and machine learning · 48 · 13 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Directional Source Separation for Robust Speech Recognition on Smart Glasses
abstract
Modern smart glasses leverage machine learning to offer real-time transcriptions, considerably enriching human communication experiences. However, such systems frequently encounter challenges related to environmental noises, leading to decreased speech recognition. To improve voice quality, this work investigates directional source separation using the multi-microphone array. We explore multiple beamformers to assist source separation by strengthening the directional properties of speech signals. In addition to relying on predetermined beamformers, we investigate neural beamforming in multi-channel source separation, demonstrating that automatic learning directional characteristics effectively improves separation quality. Furthermore, we investigate the training strategies for ASR when utilizing separated outputs. Our results suggest that jointly training a directional speech separation and ASR model achieves the best overall performance while balancing the wearer and conversation partner’s performance.
Tiantian Feng, Ju Lin, Yiteng Huang, Weipeng He, Kaustubh Kalgaonkar, Niko Moritz, Ming Sun 0013, Frank Seide
ICASSP10
2025 Efficient Streaming LLM for Speech Recognition
abstract
Recent works have shown that prompting large language models with audio encodings can unlock speech recognition capabilities. However, existing techniques do not scale efficiently, especially while handling long form streaming audio inputs — not only do they extrapolate poorly beyond the audio length seen during training, but they are also computationally inefficient due to the quadratic cost of attention.In this work, we introduce SpeechLLM-XL, a linear scaling decoder-only model for streaming speech recognition. We process audios in configurable chunks using limited attention window for reduced computation, and the text tokens for each audio chunk are generated auto-regressively until an EOS is predicted. During training, the transcript is segmented into chunks, using a CTC forced alignment estimated from encoder output. SpeechLLM-XL with 1.28 seconds chunk size achieves 2.7%/6.7% WER on LibriSpeech test clean/other, and it shows no quality degradation on long form utterances 10x longer than the training utterances.
Junteng Jia, Gil Keren, Egor Lakomkin, Xiaohui Zhang 0007, Chunyang Wu, Frank Seide, Jay Mahadeokar, Ozlem Kalinli
ICASSP7
2025 Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition
abstract
We propose the joint speech translation and recognition (JSTAR) model that leverages the fast-slow cascaded encoder architecture for simultaneous end-to-end automatic speech recognition (ASR) and speech translation (ST). The model is transducer-based and uses a multi-objective training strategy that optimizes both ASR and ST objectives simultaneously. This allows JSTAR to produce high-quality streaming ASR and ST results. We apply JSTAR in a bilingual conversational speech setting with smart-glasses, where the model is also trained to distinguish speech from different directions corresponding to the wearer and a conversational partner. Different model pre-training strategies are studied to further improve results, including training of a transducer-based streaming machine translation (MT) model for the first time and applying it for parameter initialization of JSTAR. We demonstrate superior performances of JSTAR compared to a strong cascaded ST model in both BLEU scores and latency.
Niko Moritz, Ruiming Xie, Yashesh Gaur, Ke Li 0023, Simone Merello, Frank Seide, Christian Fügen
ICASSP7
2025 Directional Speech Recognition with Full-Duplex Capability
Ju Lin, Yiteng Huang, Ming Sun 0013, Frank Seide, Florian Metze
INTERSPEECH4
2024 Effective Internal Language Model Training and Fusion for Factorized Transducer Model
abstract
The internal language model (ILM) of the neural transducer has been widely studied. In most prior work, it is mainly used for estimating the ILM score and is subsequently subtracted during inference to facilitate improved integration with external language models. Recently, various of factorized transducer models have been proposed, which explicitly embrace a standalone internal language model for non-blank token prediction. However, even with the adoption of factorized transducer models, limited improvement has been observed compared to shallow fusion. In this paper, we propose a novel ILM training and decoding strategy for factorized transducer models, which effectively combines the blank, acoustic and ILM scores. Our experiments show a 17% relative improvement over the standard decoding method when utilizing a well-trained ILM and the proposed decoding strategy on LibriSpeech datasets. Furthermore, when compared to a strong RNN-T baseline enhanced with external LM fusion, the proposed model yields a 5.5% relative improvement on general-sets and an 8.9% WER reduction for rare words. The proposed model can achieve superior performance without relying on external language models, rendering it highly efficient for production use-cases. To further improve the performance, we propose a novel and memory-efficient ILM-fusion-aware minimum word error rate (MWER) training method which improves ILM integration significantly.
Jinxi Guo, Niko Moritz, Yingyi Ma, Frank Seide, Chunyang Wu, Jay Mahadeokar, Ozlem Kalinli, Christian Fügen, Mike Seltzer
ICASSP4
2024 AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition
abstract
Wearable devices like smart glasses are approaching the compute capability to seamlessly generate real-time closed captions for live conversations. We build on our recently introduced directional Automatic Speech Recognition (ASR) for smart glasses that have microphone arrays, which fuses multi-channel ASR with serialized output training, for wearer/conversation-partner disambiguation as well as suppression of cross-talk speech from non-target directions and noise.When ASR work is part of a broader system-development process, one may be faced with changes to microphone geometries as system development progresses.This paper aims to make multi-channel ASR insensitive to limited variations of microphone-array geometry. We show that a model trained on multiple similar geometries is largely agnostic and generalizes well to new geometries, as long as they are not too different. Furthermore, training the model this way improves accuracy for seen geometries by 15 to 28% relative. Lastly, we refine the beamforming by a novel Non-Linearly Constrained Minimum Variance criterion.
Ju Lin, Niko Moritz, Yiteng Huang, Ruiming Xie, Ming Sun 0013, Christian Fügen, Frank Seide
ICASSP7
2024 Navigating the Minefield of MT Beam Search in Cascaded Streaming Speech Translation
Rastislav Rabatin, Frank Seide, Ernie Chang
INTERSPEECH2
2024 Speech ReaLLM - Real-time Speech Recognition with Multimodal Language Models by Teaching the Flow of Time
Frank Seide, Yangyang Shi, Morrie Doulaty, Yashesh Gaur, Junteng Jia, Chunyang Wu
INTERSPEECH1
2023 Joint Federated Learning and Personalization for on-Device ASR
abstract
In this paper, we propose a joint federated learning (FL) and personalization method for on-device ASR adaptation. Starting with a Conformer-based RNN-T as the ASR model backbone that is pretrained on public data and shared across devices, we propose to adapt to user data on-device by ① collectively finetuning the backbone on all user data by FL, and ② for each user, augmenting the backbone with a personalized adapter that is trained on on-device data and stored locally on their devices. As ground-truth transcriptions are not available, we use pseudo-label training, which can be completely performed on-device. Our joint method combines the best of both FL and personalization and achieves maximum effect for both heavy and light users: For users with 50+ adaptation utterances, our recipe achieves $27.4 \%$ relative WER reduction, largely due to personalization; for users with no adaptation utterance, our recipe achieves $8.9 \%$ relative WER reduction purely due to FL.
Junteng Jia, Ke Li 0023, Mani Malek 0001, Kshitiz Malik, Jay Mahadeokar, Ozlem Kalinli, Frank Seide
ASRU7
2023 Factorized Blank Thresholding for Improved Runtime Efficiency of Neural Transducers
abstract
We show how factoring the RNN-T’s output distribution can significantly reduce the computation cost and power consumption for on-device ASR inference with no loss in accuracy. With the rise in popularity of neural-transducer type models like the RNN-T for on-device ASR, optimizing RNN-T’s runtime efficiency is of great interest. While previous work has primarily focused on the optimization of RNN-T’s acoustic encoder and predictor, this paper focuses the attention on the joiner. We show that despite being only a small part of RNN-T, the joiner has a large impact on the overall model’s runtime efficiency. We propose to utilize HAT-style joiner factorization for the purpose of skipping the more expensive non-blank computation when the blank probability exceeds a certain threshold. Since the blank probability can be computed very efficiently and the RNN-T output is dominated by blanks, our proposed method leads to a 26-30% decoding speed-up and 43-53% reduction in on-device power consumption, all the while incurring no accuracy degradation and being relatively simple to implement.
Frank Seide, Yang Li 0183, Kjell Schubert, Ozlem Kalinli, Michael L. Seltzer
ICASSP2
2023 Directional Speech Recognition for Speaker Disambiguation and Cross-talk Suppression
Ju Lin, Niko Moritz, Ruiming Xie, Kaustubh Kalgaonkar, Christian Fügen, Frank Seide
INTERSPEECH6
2022 Federated Domain Adaptation for ASR with Full Self-Supervision
Junteng Jia, Jay Mahadeokar, Weiyi Zheng, Yuan Shangguan, Ozlem Kalinli, Frank Seide
INTERSPEECH6
2022 An Investigation of Monotonic Transducers for Large-Scale Automatic Speech Recognition
abstract
The two most popular loss functions for streaming end-to-end automatic speech recognition (ASR) are RNN-Transducer (RNN-T) and connectionist temporal classification (CTC). Between these two loss types we can classify the monotonic RNN-T (MonoRNN-T) and the recently proposed CTC-like Transducer (CTC-T). Monotonic transducers have a few advantages. First, RNN-T can suffer from runaway hallucination, where a model keeps emitting non-blank symbols without advancing in time. Secondly, monotonic transducers consume exactly one model score per time step and are therefore more compatible with traditional FST-based ASR decoders. However, the MonoRNN-T so far has been found to have worse accuracy than RNN-T. It does not have to be that way: By regularizing the training via joint LAS training or parameter initialization from RNN-T, both MonoRNN-T and CTC-T perform as well or better than RNN-T. This is demonstrated for LibriSpeech and for a large-scale in-house data set.
Niko Moritz, Frank Seide, Jay Mahadeokar, Christian Fügen
SLT2
2017 The microsoft 2016 conversational speech recognition system
abstract
We describe Microsoft's conversational speech recognition system, in which we combine recent developments in neural-network-based acoustic and language modeling to advance the state of the art on the Switchboard recognition task. Inspired by machine learning ensemble techniques, the system uses a range of convolutional and recurrent neural networks. I-vector modeling and lattice-free MMI training provide significant gains for all acoustic model architectures. Language model rescoring with multiple forward and backward running RNNLMs, and word posterior-based system combination provide a 20% boost. The best single system uses a ResNet architecture acoustic model with RNNLM rescoring, and achieves a word error rate of 6.9% on the NIST 2000 Switchboard task. The combined system has an error rate of 6.2%, representing an improvement over previously reported results on this benchmark task.
Wayne Xiong, Jasha Droppo, Xuedong Huang 0001, Frank Seide, Mike Seltzer, Andreas Stolcke, Dong Yu 0001, Geoffrey Zweig
ICASSP4
2017 Toward Human Parity in Conversational Speech Recognition
abstract
Conversational speech recognition has served as a flagship speech recognition task since the release of the Switchboard corpus in the 1990s. In this paper, we measure a human error rate on the widely used NIST 2000 test set for commercial bulk transcription. The error rate of professional transcribers is 5.9% for the Switchboard portion of the data, in which newly acquainted pairs of people discuss an assigned topic, and 11.3% for the CallHome portion, where friends and family members have open-ended conversations. In both cases, our automated system edges past the human benchmark, achieving error rates of 5.8% and 11.0%, respectively. The key to our system's performance is the use of various convolutional and long-short-term memory acoustic model architectures, combined with a novel spatial smoothing method and lattice-free discriminative acoustic training, multiple recurrent neural network language modeling approaches, and a systematic use of system combination. Comparing frequent errors in our human and machine transcripts, we find them to be remarkably similar, and highly correlated as a function of the speaker. Human subjects find it very difficult to tell which errorful transcriptions come from humans. Overall, this suggests that, given sufficient matched training data, conversational speech transcription engines are approximating human parity in both quantitative and qualitative terms.
Wayne Xiong, Jasha Droppo, Xuedong Huang 0001, Frank Seide, Michael L. Seltzer, Andreas Stolcke, Dong Yu 0001, Geoffrey Zweig
IEEE ACM Trans. Audio Speech Lang. Process.4
2016 CNTK: Microsoft's Open-Source Deep-Learning Toolkit
abstract
This tutorial will introduce the Computational Network Toolkit, or CNTK, Microsoft's cutting-edge open-source deep-learning toolkit for Windows and Linux. CNTK is a powerful computation-graph based deep-learning toolkit for training and evaluating deep neural networks. Microsoft product groups use CNTK, for example to create the Cortana speech models and web ranking. CNTK supports feed-forward, convolutional, and recurrent networks for speech, image, and text workloads, also in combination. Popular network types are supported either natively (convolution) or can be described as a CNTK configuration (LSTM, sequence-to-sequence). CNTK scales to multiple GPU servers and is designed around efficiency. The tutorial will give an overview of CNTK's general architecture and describe the specific methods and algorithms used for automatic differentiation, recurrent-loop inference and execution, memory sharing, on-the-fly randomization of large corpora, and multi-server parallelization. We will then show how typical uses looks like for relevant tasks like image recognition, sequence-to-sequence modeling, and speech recognition.
Frank Seide
KDD1
2015 Deep bi-directional recurrent networks over spectral windows
abstract
Long short-term memory (LSTM) acoustic models have recently achieved state-of-the-art results on speech recognition tasks. As a type of recurrent neural network, LSTMs potentially have the ability to model long-span phenomena relating the spectral input to linguistic units. However, it has not been clear whether their observed performance is actually due to this capability, or instead if it is due to a better modeling of short term dynamics through the recurrence. In this paper. we answer this question by applying a windowed (truncated) LSTM to conversational speech transcription, and find that a limited context is adequate, and that it is not necessaary to scan the entire utterance. The sliding window approach allows not only incremental (online) recognition with a bidirectional model, but also frame-wise randomization (as opposed to utterance randomization), which results in faster convergence. On the SWBD/Fisher corpus, applying bidirectional LSTM RNNs to spectral windows of about 0.5s improves WER on the Hub5'00 benchmark set by 16% relative compared to our best sequence-trained DNN. On an extended 3850h training set that that also includes lectures, the relative gain becomes 28% (Hub5'00 WER 9.2%). In-house conversational data improves by 12 to 17% relative.
Abdel-rahman Mohamed, Frank Seide, Dong Yu 0001, Jasha Droppo, Andreas Stolcke, Geoffrey Zweig, Gerald Penn
ASRU2
2014 On parallelizability of stochastic gradient descent for speech DNNS
abstract
This paper compares the theoretical efficiency of model-parallel and data-parallel distributed stochastic gradient descent training of DNNs. For a typical Switchboard DNN with 46M parameters, the results are not pretty: With modern GPUs and interconnects, model parallelism is optimal with only 3 GPUs in a single server, while data parallelism with a minibatch size of 1024 does not even scale to 2 GPUs. We further show that data-parallel training efficiency can be improved by increasing the minibatch size (through a combination of AdaGrad and automatic adjustments of learning rate and minibatch size) and data compression. We arrive at an estimated possible end-to-end speed-up of 5 times or more. We do not address issues of robustness to process failure or other issues that might occur during training, nor of speed of convergence differences between ASGD and SGD parameter update patterns.
Frank Seide, Jasha Droppo, Gang Li 0012, Dong Yu 0001
ICASSP1
2014 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs
abstract
We show empirically that in SGD training of deep neural networks, one can, at no or nearly no loss of accuracy, quantize the gradients aggressively—to but one bit per value—if the quantization error is carried forward across minibatches (error feedback). This size reduction makes it feasible to parallelize SGD through data-parallelism with fast processors like recent GPUs. We implement data-parallel deterministically distributed SGD by combining this finding with AdaGrad, automatic minibatch-size selection, double buffering, and model parallelism. Unexpectedly, quantization benefits AdaGrad, giving a small accuracy gain. For a typical Switchboard DNN with 46M parameters, we reach computation speeds of 27k frames per second (kfps) when using 2880 samples per minibatch, and 51kfps with 16k, on a server with 8 K20X GPUs. This corresponds to speed-ups over a single GPU of 3.6 and 6.3, respectively. 7 training passes over 309h of data complete in under 7h. A 160M-parameter model training processes 3300h of data in under 16h on 20 dual-GPU servers—a 10 times speed-up—albeit at a small accuracy loss.
Frank Seide, Jasha Droppo, Gang Li 0012, Dong Yu 0001
INTERSPEECH1
2014 An introduction to computational networks and the computational network toolkit (invited talk)
Dong Yu 0001, Adam Eversole, Michael L. Seltzer, Kaisheng Yao, Brian Guenter, Oleksii Kuchaiev, Frank Seide, Huaming Wang, Jasha Droppo, Zhiheng Huang, Geoffrey Zweig, Christopher J. Rossbach, Jon Currey
INTERSPEECH7
2013 Recent advances in deep learning for speech research at Microsoft
abstract
Deep learning is becoming a mainstream technology for speech recognition at industrial scale. In this paper, we provide an overview of the work by Microsoft speech researchers since 2009 in this area, focusing on more recent advances which shed light to the basic capabilities and limitations of the current deep learning technology. We organize this overview along the feature-domain and model-domain dimensions according to the conventional approach to analyzing speech systems. Selected experimental results, including speech recognition and related applications such as spoken dialogue and language modeling, are presented to demonstrate and analyze the strengths and weaknesses of the techniques described in the paper. Potential improvement of these techniques and future research directions are discussed.
Li Deng 0001, Jinyu Li 0001, Jui-Ting Huang, Kaisheng Yao, Dong Yu 0001, Frank Seide, Michael L. Seltzer, Geoffrey Zweig, Xiaodong He 0001, Jason D. Williams, Yifan Gong 0001, Alex Acero
ICASSP6
2013 Error back propagation for sequence training of Context-Dependent Deep NetworkS for conversational speech transcription
abstract
We investigate back-propagation based sequence training of Context-Dependent Deep-Neural-Network HMMs, or CD-DNN-HMMs, for conversational speech transcription. Theoretically, sequence training integrates with backpropagation in a straight-forward manner. However, we find that to get reasonable results, heuristics are needed that point to a problem with lattice sparseness: The model must be adjusted to the updated numerator lattices by additional iterations of frame-based cross-entropy (CE) training; and to avoid distortions from “runaway” models, we can either add artificial silence arcs to the denominator lattices, or smooth the sequence objective with the frame-based one (F-smoothing). With the 309h Switchboard training set, the MMI objective achieves a relative word-error rate reduction of 11-15% over CE for matched test sets, and 10-17% for mismatched ones. This includes gains of 4-7% from realigned CE iterations. The BMMI and sMBR objectives gain less. With 2000h of data, gains are 2-9% after realigned CE iterations. Using GPGPUs, MMI is about 70% slower than CE training.
Gang Li 0012, Dong Yu 0001, Frank Seide
ICASSP4
2013 KL-divergence regularized deep neural network adaptation for improved large vocabulary speech recognition
abstract
We propose a novel regularized adaptation technique for context dependent deep neural network hidden Markov models (CD-DNN-HMMs). The CD-DNN-HMM has a large output layer and many large hidden layers, each with thousands of neurons. The huge number of parameters in the CD-DNN-HMM makes adaptation a challenging task, esp. when the adaptation set is small. The technique developed in this paper adapts the model conservatively by forcing the senone distribution estimated from the adapted model to be close to that from the unadapted model. This constraint is realized by adding Kullback-Leibler divergence (KLD) regularization to the adaptation criterion. We show that applying this regularization is equivalent to changing the target distribution in the conventional backpropagation algorithm. Experiments on Xbox voice search, short message dictation, and Switchboard and lecture speech transcription tasks demonstrate that the proposed adaptation technique can provide 2%-30% relative error reduction against the already very strong speaker independent CD-DNN-HMM systems using different adaptation sets under both supervised and unsupervised adaptation setups.
Dong Yu 0001, Kaisheng Yao, Gang Li 0012, Frank Seide
ICASSP5
2013 A new language independent, photo-realistic talking head driven by voice only
abstract
We propose a new photo-realistic, voice driven only (i.e. no linguistic info of the voice input is needed) talking head. The core of the new talking head is a context-dependent, multilayer, Deep Neural Network (DNN), which is discriminatively trained over hundreds of hours, speaker independent speech data. The trained DNN is then used to map acoustic speech input to 9,000 tied “senone” states probabilistically. For each photo-realistic talking head, an HMM-based lips motion synthesizer is trained over the speaker’s audio/visual training data where states are statistically mapped to the corresponding lips images. In test, for given speech input, DNN predicts the likely states in their posterior probabilities and photo-realistic lips animation is then rendered through the DNN predicted state lattice. The DNN trained on English, speaker independent data has also been tested with other language input, e.g. Mandarin, Spanish, etc. to mimic the lips movements cross-lingually. Subjective experiments show that lip motions thus rendered for 15 non-English languages are highly synchronized with the audio input and photo-realistic to human eyes perceptually.
Xinjian Zhang, Gang Li 0012, Frank Seide, Frank K. Soong
INTERSPEECH4
2013 The Deep Tensor Neural Network With Applications to Large Vocabulary Speech Recognition
abstract
The recently proposed context-dependent deep neural network hidden Markov models (CD-DNN-HMMs) have been proved highly promising for large vocabulary speech recognition. In this paper, we develop a more advanced type of DNN, which we call the deep tensor neural network (DTNN). The DTNN extends the conventional DNN by replacing one or more of its layers with a double-projection (DP) layer, in which each input vector is projected into two nonlinear subspaces, and a tensor layer, in which two subspace projections interact with each other and jointly predict the next layer in the deep architecture. In addition, we describe an approach to map the tensor layers to the conventional sigmoid layers so that the former can be treated and trained in a similar way to the latter. With this mapping we can consider a DTNN as the DNN augmented with DP layers so that not only the BP learning algorithm of DTNNs can be cleanly derived but also new types of DTNNs can be more easily developed. Evaluation on Switchboard tasks indicates that DTNNs can outperform the already high-performing DNNs with 4-5% and 3% relative word error reduction, respectively, using 30-hr and 309-hr training sets.
Dong Yu 0001, Li Deng 0001, Frank Seide
IEEE Trans. Speech Audio Process.3
2012 Exploiting sparseness in deep neural networks for large vocabulary speech recognition
abstract
Recently, we developed context-dependent deep neural network (DNN) hidden Markov models for large vocabulary speech recognition. While reducing errors by 33% compared to its discriminatively trained Gaussian-mixture counterpart on the switchboard benchmark task, DNN requires much more parameters. In this paper, we report our recent work on DNN for improved generalization, model size, and computation speed by exploiting parameter sparseness. We formulate the goal of enforcing sparseness as soft regularization and convex constraint optimization problems, and propose solutions under the stochastic gradient ascent setting. We also propose novel data structures to exploit the random sparseness patterns to reduce model size and computation time. The proposed solutions have been evaluated on the voice-search and switchboard datasets. They have decreased the number of nonzero connections to one third while reducing the error rate by 0.2-0.3% over the fully connected model on both datasets. The nonzero connections have been further reduced to only 12% and 19% on the two respective datasets without sacrificing speech recognition performance. Under these conditions we can reduce the model size to 18% and 29%, and computation to 14% and 23%, respectively, on these two datasets.
Dong Yu 0001, Frank Seide, Gang Li 0012, Li Deng 0001
ICASSP2
2012 Conversational Speech Transcription Using Context-Dependent Deep Neural Networks
Dong Yu 0001, Frank Seide, Gang Li 0012
ICML2
2012 Pipelined Back-Propagation for Context-Dependent Deep Neural Networks
abstract
The Context-Dependent Deep-Neural-Network HMM, or CD-DNN-HMM, is a recently proposed acoustic-modeling tech-nique for HMM-based speech recognition that can greatly out-perform conventional Gaussian-mixture based HMMs. For ex-ample, a CD-DNN-HMM trained on the 2000h Fisher corpus achieves 14.4 % word error rate on the Hub5’00-FSH speaker-independent phone-call transcription task, compared to 19.6% obtained by a state-of-the-art, conventional discriminatively trained GMM-based HMM. That CD-DNN-HMM, however, took 59 days to train on a modern GPGPU—the immense computational cost of the mini-batch based back-propagation (BP) training is a major road-block. Unlike the familiar Baum-Welch training for conven-tional HMMs, BP cannot be efficiently parallelized across data. In this paper we show that the pipelined approximation to BP, which parallelizes computation with respect to layers, is an efficient way of utilizing multiple GPGPU cards in a single server. Using 2 and 4 GPGPUs, we achieve a 1.9 and 3.3 times end-to-end speed-up, at parallelization efficiency of 0.95 and 0.82, respectively, at no loss of recognition accuracy. Index Terms: speech recognition, deep neural networks, paral-lelization, GPGPU
Xie Chen 0001, Adam Eversole, Gang Li 0012, Dong Yu 0001, Frank Seide
INTERSPEECH5
2012 ClippyScript: A Programming Language for Multi-Domain Dialogue Systems
Frank Seide, Sean McDirmid
INTERSPEECH1
2012 Voice Activity Detection Using Speech Recognizer Feedback
Kit Thambiratnam, Weiwu Zhu, Frank Seide
INTERSPEECH3
2012 Large Vocabulary Speech Recognition Using Deep Tensor Neural Networks
abstract
Recently, we proposed and developed the context-dependent deep neural network hidden Markov models (CD-DNN-HMMs) for large vocabulary speech recognition and achieved highly promising recognition results including over one third fewer word errors than the discriminatively trained, conventional HMM-based systems on the 300hr Switchboard benchmark task. In this paper, we extend DNNs to deep tensor neural networks (DTNNs) in which one or more layers are double-projection and tensor layers. The basic idea of the DTNN comes from our realization that many factors interact with each other to predict the output. To represent these interactions, we project the input to two nonlinear subspaces through the double-projection layer and model the interactions between these two subspaces and the output neurons through a tensor with three-way connections. Evaluation on 30hr Switchboard task indicates that DTNNs can outperform DNNs with similar number of parameters with 5% relative word error reduction. Index Terms: automatic speech recognition, tensor deep neural networks, CD-DNN-HMM, large vocabulary 1.
Dong Yu 0001, Li Deng 0001, Frank Seide
INTERSPEECH3
2012 Context-dependent Deep Neural Networks for audio indexing of real-life data
abstract
We apply Context-Dependent Deep-Neural-Network HMMs, or CD-DNN-HMMs, to the real-life problem of audio indexing of data across various sources. Recently, we had shown that on the Switchboard benchmark on speaker-independent transcription of phone calls, CD-DNN-HMMs with 7 hidden layers reduce the word error rate by as much as one-third, compared to discriminatively trained Gaussian-mixture HMMs, and by one-fourth if the GMM-HMM also uses fMPE features. This paper takes CD-DNN-HMM based recognition into a real-life deployment for audio indexing. We find that for our best speaker-independent CD-DNN-HMM, with 32k senones trained on 2000h of data, the one-fourth reduction does carry over to inhomogeneous field data (video podcasts and talks). Compared to a speaker-adaptive GMM system, the relative improvement is 18%, at very similar end-to-end runtime. In system building, we find that DNNs can benefit from a larger number of senones than the GMM-HMM; and that DNN likelihood evaluation is a sizeable runtime factor even in our wide-beam context of generating rich lattices: Cutting the model size by 60% reduces runtime by one-third at a 5% relative WER loss.
Gang Li 0012, Huifeng Zhu, Gong Cheng 0005, Kit Thambiratnam, Behrooz Chitsaz, Dong Yu 0001, Frank Seide
SLT7
2012 Adaptation of context-dependent deep neural networks for automatic speech recognition
abstract
In this paper, we evaluate the effectiveness of adaptation methods for context-dependent deep-neural-network hidden Markov models (CD-DNN-HMMs) for automatic speech recognition. We investigate the affine transformation and several of its variants for adapting the top hidden layer. We compare the affine transformations against direct adaptation of the softmax layer weights. The feature-space discriminative linear regression (fDLR) method with the affine transformations on the input layer is also evaluated. On a large vocabulary speech recognition task, a stochastic gradient ascent implementation of the fDLR and the top hidden layer adaptation is shown to reduce word error rates (WERs) by 17% and 14%, respectively, compared to the baseline DNN performances. With a batch update implementation, the softmax layer adaptation technique reduces WERs by 10%. We observe that using bias shift performs as well as doing scaling plus bias shift.
Kaisheng Yao, Dong Yu 0001, Frank Seide, Li Deng 0001, Yifan Gong 0001
SLT3
2011 Subword-based multi-span pronunciation adaptation for recognizing accented speech
abstract
We investigate automatic pronunciation adaptation for non-native accented speech by using statistical models trained on multi-span lingustic parse tables to generate candidate mispronunciations for a target language. Compared to traditional phone re-writing rules, parse table modeling captures more context in the form of phone-clusters or syllables, and encodes abstract features such as word-internal position or syllable structure. The proposed approach is attractive because it gives a unified method for combining multiple levels of linguistic information. The reported experiments demonstrate word error rate reductions of up to 7.9% and 3.3% absolute on Italian and German accented English using lexicon adaptation alone, and 12.4% and 11.3% absolute when combined with acoustic adaptation.
Timo Mertens, Kit Thambiratnam, Frank Seide
ASRU3
2011 Feature engineering in Context-Dependent Deep Neural Networks for conversational speech transcription
abstract
We investigate the potential of Context-Dependent Deep-Neural-Network HMMs, or CD-DNN-HMMs, from a feature-engineering perspective. Recently, we had shown that for speaker-independent transcription of phone calls (NIST RT03S Fisher data), CD-DNN-HMMs reduced the word error rate by as much as one third-from 27.4%, obtained by discriminatively trained Gaussian-mixture HMMs with HLDA features, to 18.5%-using 300+ hours of training data (Switchboard), 9000+ tied triphone states, and up to 9 hidden network layers.
Frank Seide, Gang Li 0012, Xie Chen 0001, Dong Yu 0001
ASRU1
2011 Leveraging the Web for automatically generating indexable and browsable keywords for speech files
abstract
This paper presents a method for generating indexable and browsable keyword metadata from ASR transcripts by leveraging the Web. Search engine queries are built from an ASR transcript and used to retrieve similar text from the Web. The keyword meta information embedded in those pages for search engines is then ranked using a mutual information criteria to derive a keyword set. The proposed method is training-free, allows phrase keyword generation, and can generate words that were not spoken in the ASR transcript, alleviating the impact of ASR out-of-vocabulary. Subjective evaluations on technical presentations demonstrate a clear preference for this approach. Additionally an objective measure of keyword generation performance is proposed and shown to be a useful guide for tuning compared to more onerous subjective evaluations.
Kishan Thambiratnam, Gang Li 0012, Sha Meng, Frank Seide
ICASSP4
2011 Conversational Speech Transcription Using Context-Dependent Deep Neural Networks
abstract
We apply the recently proposed Context-Dependent Deep-Neural-Network HMMs, or CD-DNN-HMMs, to speech-to-text transcription. For single-pass speaker-independent recognition on the RT03S Fisher portion of phone-call transcription benchmark (Switchboard), the word-error rate is reduced from 27.4%, obtained by discriminatively trained Gaussian-mixture HMMs, to 18.5%—a 33 % relative improvement. CD-DNN-HMMs combine classic artificial-neural-network HMMs with traditional tied-state triphones and deep-beliefnetwork pre-training. They had previously been shown to reduce errors by 16 % relatively when trained on tens of hours of data using hundreds of tied states. This paper takes CD-DNN-HMMs further and applies them to transcription using over 300 hours of training data, over 9000 tied states, and up to 9 hidden layers, and demonstrates how sparseness can be exploited. On four less well-matched transcription tasks, we observe relative error reductions of 22–28%. Index Terms: speech recognition, deep belief networks, deep neural networks
Frank Seide, Gang Li 0012, Dong Yu 0001
INTERSPEECH1
2010 Music rhythm characterization with application to workout-mix generation
abstract
In this paper, we present approaches to musical rhythm pattern extraction, rhythm-based music retrieval, and rhythm-synchronized music mixing. A probabilistic model is used to jointly estimate tempo and time signature as a basis for beat tracking and measure detection. A representative rhythm pattern is then extracted through clustering to characterize the rhythm of a song. Based on this, a probabilistic approach is used for retrieving songs with similar rhythmic patterns. These are then mixed rhythm-synchronously with transitions maintaining continuity and regularity of beats. We apply the presented methods into workout-mix generation, which aims at automatically selecting rhythmically similar music given a seed song and a user-defined tempo profile. Our probabilistic approaches achieve accuracies similar to best published results, but avoid manually tuned parameters and “fudge factors”.
Lie Lu, Christopher Weare, Frank Seide
ICASSP4
2010 Vocabulary and language model adaptation using just one speech file
abstract
This paper investigates unsupervised vocabulary and language model self-adaptation (VLA) from just one speech file using the web as a knowledge source and without prior knowledge of topic or domain beyond optional file metadata. Single-file self adaptation is regularly used for acoustic adaptation, but to date, is rarely used for VLA. The method investigated here uses a first-pass transcript or file metadata to generate web search queries for retrieving texts for adaptation. Various strategies for building queries, retrieving web texts and maximizing out-of-vocabulary (OOV) recovery while constraining vocabulary growth are examined. Significant improvements are demonstrated for transcribing and searching recorded lectures and telephone calls. The proposed method is orthogonal with acoustic adaptation and system combination and integrates well in multi-pass recognition architectures.
Sha Meng, Kishan Thambiratnam, Yimeng Lin, Gang Li 0012, Frank Seide
ICASSP6
2010 On using missing-feature theory with cepstral features - approximations to the multivariate integral
Frank Seide
INTERSPEECH1
2009 Automatic punctuation generation for speech
abstract
Automatic generation of punctuation is an essential feature for many speech-to-text transcription tasks. This paper describes a maximum a-posteriori (MAP) approach for inserting punctuation marks into raw word sequences obtained from automatic speech recognition (ASR). The system consists of an ¿acoustic model¿ (AM) for prosodic features (actually pause duration) and a ¿language model¿ (LM) for text-only features. The LM combines three components: an MLP-based trigger-word model and a forward and a backward trigram punctuation predictor. The separation into acoustic and language model allows to learn these models on different corpora, especially allowing the LM to be trained on large amounts of data (text) for which no acoustic information is available. We find that the trigger-word LM is very useful, and further improvement can be achieved when combining both prosodic and lexical information. We achieve an F-measure of 81.0% and 56.5% for voicemails and podcasts, respectively, on reference transcripts, and 69.6% for voicemails on ASR transcripts.
Wenzhu Shen, Roger Peng Yu, Frank Seide, Ji Wu 0002
ASRU3
2009 Unsupervised speaker adaptation for telephone call transcription
abstract
The use of the PC and Internet for placing telephone calls will present new opportunities to capture vast amounts of un-transcribed speech for a particular speaker. This paper investigates how to best exploit this data for speaker-dependent speech recognition. Supervised and unsupervised experiments in acoustic model and language model adaptation are presented. Using one hour of automatically transcribed speech per speaker with a word error rate of 36.0%, unsupervised adaptation resulted in an absolute gain of 6.3%, equivalent to 70% of the gain from the supervised case, with additional adaptation data likely to yield further improvements. LM adaptation experiments suggested that although there seems to be a small degree of speaker idiolect, adaptation to the speaker alone, without considering the topic of the conversation, is in itself unlikely to improve transcription accuracy.
R. Wallace, Kishan Thambiratnam, Frank Seide
ICASSP3
2009 Learning a music similarity measure on automatic annotations with application to playlist generation
abstract
This paper presents an approach to learn a better music similarity measure and presents an application to music playlist generation. Different from previous work, in our approach, automatically detected music attributes are used to represent each song. A set of kernels is employed in similarity measure, with each kernel measuring on a subset of music attributes and having a different importance weight. In automatic music playlist generation, a ranking method is presented, which considers multiple seed songs and possible outlier seed. Experiments show the effectiveness of the proposed approach, and the quality of the playlist generated based on automatic annotations is comparable to that based on manual annotations.
Linxing Xiao, Lie Lu, Frank Seide, Jie Zhou 0001
ICASSP3
2009 Unsupervised lattice-based acoustic model adaptation for speaker-dependent conversational telephone speech transcription
abstract
This paper examines the application of lattice adaptation tech-niques to speaker-dependent models for the purpose of conver-sational telephone speech transcription. Given sufficient train-ing data per speaker, it is feasible to build adapted speaker-dependent models using lattice MLLR and lattice MAP. Experi-ments on iterative and cascaded adaptation are presented. Addi-tionally various strategies for thresholding frame posteriors are investigated, and it is shown that accumulating statistics from the local best-confidence path is sufficient to achieve optimal adaptation. Overall, an iterative cascaded lattice system was able to reduce WER by 7.0 % abs., which was a 0.8 % abs. gain over transcript-based adaptation. Lattice adaptation reduced the unsupervised/supervised adaptation gap from 2.5 % to 1.7%.
Kishan Thambiratnam, Frank Seide
INTERSPEECH2
2008 Mobile ringtone search through query by humming
abstract
In the context of voice-based mobile search, this paper presents a new approach to mobile ringtone search through query by humming: A user can call a service, hum a part of melody through the mobile phone, and obtain the ringtones or songs he or she is looking for. Correspondingly, we propose a method of query by humming tailored to this scenario. A robust front-end processing is first presented to deal with the mobile phone recording, which is distorted due to GSM codec, environment and wireless transmission. Then, a systematic probabilistic model and matching procedure inspired by Hidden Markov Model (HMM) is presented, by considering the alignment and error tolerance in the matching between query and songs. A rescoring heuristic is finally employed to further improve matching accuracy. Moreover, our system is evaluated on realistic mobile recordings from the field. Experiments show our approach can achieve 83% accuracy on a database with 3000 songs in this realistic scenario.
Lie Lu, Frank Seide
ICASSP2
2008 Fusing multiple systems into a compact lattice index for chinese spoken term detection
abstract
We examine the task of spoken term detection in Chinese spontaneous speech with a lattice-based approach. We first compare lattices generated with different units: word, character, tonal and toneless syllables, and also lattices converted from one unit to another unit. Then we combine lattices from multiple systems into a single lattice. By fully exploiting the redundant information in the combined lattice with a time-based node/arc merging, we achieve the result of a compact lattice index with the accuracy improved to 79.2% from 73.9% using the best subsystem.
Sha Meng, Jia Liu 0001, Frank Seide
ICASSP4
2008 Approximateword-lattice indexing with text indexers: Time-Anchored Lattice Expansion
abstract
We address the problem of how to represent or approximate speech lattices to be indexed with existing text indexers. We present a method named Time-Anchored Lattice Expansion (TALE), which can be implemented by a Standard Text Indexer (STI). On a 170-hour lecture set, we compare TALE with other lattice indexing methods: confusion networks, Position-Specific Posterior Lattices (PSPL), and Time-based Merging for Index (TMI). All methods achieve accuracies comparable to searching raw lattices when the corresponding index structures and phrase matching algorithms are used. However, when implemented with an STI, TALE significantly outperforms all other methods. Compared to indexing linear text, TALE improves accuracy by 30–60% for multi-word phrase searches and by 130% for two-term AND queries.
Yu Shi 0001, Frank Seide
ICASSP3
2008 Addressing the out-of-vocabulary problem for large-scale Chinese spoken term detection
Sha Meng, Jian Shao 0001, Roger Peng Yu, Jia Liu 0001, Frank Seide
INTERSPEECH5
2008 Towards vocabulary-independent speech indexing for large-scale repositories
Jian Shao 0001, Roger Peng Yu, Qingwei Zhao, Yonghong Yan 0002, Frank Seide
INTERSPEECH5
2008 GPU-accelerated Gaussian clustering for fMPE discriminative training
abstract
The Graphics Processing Unit (GPU) has extended its applications from its original graphic rendering to more general scientific computation. Through massive parallelization, state-ofthe-art GPUs can deliver 200 billion floating-point operations per second (0.2 TFLOPS) on a single consumer-priced graphics card. This paper describes our attempt in leveraging GPUs for efficient HMM model training. We show that using GPUs for a specific example of Gaussian clustering, as required in fMPE, or feature-domain Minimum Phone Error discriminative training, can be highly desirable. The clustering of huge number of Gaussians is very time consuming due to the enormous model size in current LVCSR systems. Comparing an NVidia Geforce 8800 Ultra GPU against an Intel Pentium 4 implementation, we find that our brute-force GPU implementation is 14 times faster overall than a CPU implementation that uses approximate speed-up heuristics. GPU accelerated fMPE reduces the WER 6% relatively, compared to the maximumlikelihood trained baseline on two conversational-speech recognition tasks.
Yu Shi 0001, Frank Seide, Frank K. Soong
INTERSPEECH2
2008 Fragmented context-dependent syllable acoustic models
abstract
Though touted as an excellent candidate, past work has yet to demonstrate the value of the syllable for acoustic modeling. One reason is that critical factors such as context-dependency and model clustering are typically neglected in syllable works. This paper presents fragmented syllable models, a means to realize context-dependency for the syllable while constraining the implied explosion in training data requirements. Fragmented syllables only expose their head/tail phones as context, and thus limit the context space for triphone expansion. Furthermore, decision-tree clustering can be used to share data between parts, or fragments, of syllables, to better exploit training data for data-sparse syllables. The best resulting system achieves a 1.8% absolute (5.4% relative) reduction in WER over a baseline triphone acoustic model on a Switchboard-1 conversational telephone speech task.
Kishan Thambiratnam, Frank Seide
INTERSPEECH2
2008 Word-lattice based spoken-document indexing with standard text indexers
abstract
Indexing the spoken content of audio recordings requires automatic speech recognition, which is as of today not reliable. Unlike indexing text, we cannot reliably know from a speech recognizer whether a word is present at a given point in the audio; we can only obtain a probability for it. Correct use of these probabilities significantly improves spoken-document search accuracy.
Frank Seide, Kishan Thambiratnam, Roger Peng Yu
SLT1
2008 Mobile Search With Multimodal Queries
abstract
The popularity of mobile devices, such as PDAs and SmartPhones, has grown rapidly over the last couple of years. Though most users still perform searches using desktop computers, it is expected that more and more people will also search the Web while they are on the move. In addition to text-based keyword queries, mobile devices can support richer and hybrid queries such as images, audio, video, and their combinations. In this paper, we will discuss mobile search systems that support image queries and audio queries, covering typical designs for mobile visual and audio search, as well as the opportunities and challenges. Specifically, we will present an in-depth study of two real systems we have developed: product image categorization and mobile ringtone search, which use image queries and audio queries, respectively. Experimental results on real-life data demonstrate their effectiveness and efficiency.
Xing Xie 0001, Lie Lu, Menglei Jia, Hua Li 0001, Frank Seide, Wei-Ying Ma
Proc. IEEE5
2007 A study of lattice-based spoken term detection for Chinese spontaneous speech
abstract
We examine the task of spoken term detection in Chinese spontaneous speech with a lattice-based approach. We compare lattices generated with different units: word, character, tonal syllable and toneless syllable, and also look into methods of converting lattices from one unit to another one. We find the best system is with toneless-syllable lattices converted from word lattices. Further improvement is achieved by lattice post-processing and system combination. Our best system has an accuracy of 80.2% on a keyword spotting task.
Sha Meng, Frank Seide, Jia Liu 0001
ASRU3
2007 Towards spoken-document retrieval for the enterprise: Approximate word-lattice indexing with text indexers
abstract
Enterprise-scale search engines are generally designed for linear text. Linear text is suboptimal for audio search, where accuracy can be significantly improved if the search includes alternate recognition candidates, commonly represented as word lattices. We propose two methods to enable text indexers to approximately index lattices with little or no code change: “TMI” (Time-based Merging for Indexing) aims at lattice-index size reduction, and the “sausage”-like “TALE” (Time-Anchored Lattice Expansion) approximation requires no indexer-code or data-format changes at all. On four enterprise-type data sets (meetings, phone calls, lectures, and voicemail), TMI and TALE improve accuracy by 30–60% for multi-word phrase searches and by 130% for two-term AND queries, compared to indexing linear text.
Frank Seide, Yu Shi 0001
ASRU1
2007 A Hidden-State Maximum Entropy Model Forword Confidence Estimation
abstract
We propose a probabilistic model for estimating word confidence by fusing predictor features. Starting from the maximum entropy (ME) method, we first prove that ME model is equivalent to the best model with certain form to the minimum expected cross entropy (MECE) criterion. Under the MECE criterion, We extend the form of ME model by introducing a hidden state. We call the new model hidden-state maximum entropy (HSME) model. In a keyword-spotting task, we combine predictor features from both phonetic and word-level systems. Compared to lattice posterior alone, recall at 80% precision is improved from 38.1% to 49.5% on voicemail and from 37.1% to 51.9% on Switchboard. Compared with other fusion methods, HSME consistently outperforms decision tree, and most cases SVM.
Yuchou Chang, Frank Seide
ICASSP (4)5
2007 Online vocabulary adaptation using limited adaptation data
abstract
This paper presents a study of low-latency domain-independent online vocabulary adaptation using limited amounts of support-ing text data. The target applications include blind indexing of Internet content, indexing of new content with low latency, and domains where Out-Of-Vocabulary (OOV) words are prob-lematic. A number of methods to perform document-specific adaptation using a small amount of support metadata and the Internet are examined. It is shown that a combination of word feature fusion and cross-file statistics pooling provides robust adaptation. The best evaluated method achieved an absolute re-duction of 27.6 % in OOV detection false alarm rate over the baseline word feature thresholding methods.
C. E. Liu, Kishan Thambiratnam, Frank Seide
INTERSPEECH3
2007 Learning spoken document similarity and recommendation using supervised probabilistic latent semantic analysis
abstract
This paper presents a model-based approach to spoken docu-ment similarity called Supervised Probabilistic Latent Seman-tic Analysis (PLSA). The method differs from traditional spo-ken document similarity techniques in that it allows similarity to be learned rather than approximated. The ability to learn similarity is desirable in applications such as Internet video recommendation, in which complex relationships like user-preference or speaking style need to be predicted. The pro-posed method exploits prior knowledge of document relation-ships to learn similarity. Experiments on broadcast news and Internet video corpora yielded 16.2 % and 9.7 % absolute mAP gains over traditional PLSA. Additionally, a cascaded Super-vised+Discriminative PLSA system achieved a 3.0 % absolute mAP gain over a Discriminative PLSA system, demonstrat-ing the complementary nature of Supervised and Discriminative
Kishan Thambiratnam, Frank Seide
INTERSPEECH2
2006 Maximum Entropy Based Normalization Of Word Posteriors For Phonetic And Lvcsr Lattice Search
abstract
In many KEWORD-spotting systems, the word posterior probability is an elementary quantity. In theory, the posterior of a KEWORD match denotes the probability of the match being correct. However, posteriors estimated on lattices, in particular phoneme lattices, are often off by orders of magnitude. This paper investigates the problem of providing "correct" posteriors, in the context of lattice-based word spotting. Unlike other work on word posteriors that focusses on relative ranking of posteriors, we emphasize relevance of the absolute value of the posterior in our user scenario. We stipulate that the posteriors should approach empirical precisions in a limit sense. Using this as a constraint, we estimate a mapping function based on Maximum Entropy. We find that for posteriors generated from phonetic lattices, mapped posteriors are satisfyingly consistent with empirical precision. In a joint search task, where different words are ranked together by posterior, FOM (Figure Of Merit) improved from 11.2% to 57.8%, which demonstrated the effectiveness of the method. Applied to searching LVCSR-based word lattices, the improvement is neglectable, but it is still effective when combining phonetic and word-lattice search in a hybrid mode, yielding an improvement from 46.7% to 65.8%.
Frank Seide
ICASSP (1)3
2006 Towards Spoken-Document Retrieval for the Internet: Lattice Indexing For Large-Scale Web-Search Architectures
Zheng-Yu Zhou, Ciprian Chelba, Frank Seide
HLT-NAACL4
2006 Discriminatively Trained spoken Document Similarity Models and their Application to Probabilistic Latent Semantic Analysis
abstract
This paper presents a novel framework for discriminatively training spoken document similarity models. Traditional similarity methods such as Vector Space Modeling and Probabilistic Latent Semantic Analysis suffer from a mismatch in modeling and evaluation objective functions. This work proposes reconciling this mismatch by using a discriminative training process in conjunction with prior knowledge of known document relationships to train an ensemble of spoken document similarity models. The reported experiments demonstrate dramatic improvements in mAP performance for the tasks of related document search and query-by-document retrieval, and highlight the ability of the resulting models to better generalize to unseen topics and unseen documents.
Kishan Thambiratnam, Frank Seide
SLT2
2005 Fast Two-Stage Vocabulary-Independent Search In Spontaneous Speech
abstract
For efficient organization of speech recordings - meetings, interviews, voice mails, lectures - the ability to search for spoken keywords is an essential capability. In Seide et al. (2004) and Yu et al. (2004), we presented our work on vocabulary-independent search in spontaneous speech. That method involved linear scanning of phonetic lattices, and thus did not scale up to large collections. In this paper, we present a two-stage approach to fast search: first we retrieve segments from an index-like structure that are promising to contain the keyword, then we locate individual keyword occurrences by a detailed linear lattice scan. However, designing an efficient vocabulary-independent indexing structure is non-trivial. We use a "soft" index, similar to Allauzen et al., that provides expected term frequencies (ETF) of query terms. We propose to approximate ETF by M-gram phoneme language models estimated on the lattices (one per segment). Our index stores these language models in an inverted structure. Word spotting experiments on voicemails show that with this two-stage method, we lose under 4% FOM (figure of merit) relative at a 25-times speed-up compared with a full linear search.
Frank Seide
ICASSP (1)2
2005 The use of virtual hypothesis copies in decoding of large-vocabulary continuous speech
abstract
High computational effort hinders wide-spread deployment of large-vocabulary continuous-speech recognition (LVCSR), for example in home or mobile devices. To this end, we developed a novel approach to LVCSR Viterbi decoding with significantly reduced effort. By a novel search-space organization called virtual hypothesis copies, we eliminate search-space copies that are approximately redundant: 1) Word-lattice generation and (M+1)-gram lattice rescoring are integrated into a single-pass time-synchronous beam search. Hypothesis copying becomes independent from the language-model order. 2) The word-pair approximation is replaced by the novel phone-history approximation (PHA). Tree copies are shared among multiple linguistic histories that end in the same phone(s). 3) Copies of individual tree arcs are shared by recombining within-word hypotheses at phone boundaries according to the PHA. At no loss of accuracy, we achieve a search-space reduction of 60-80% for Mandarin LVCSR, and of 40-50% for English (NAB 64 K). The method is exact under certain model assumptions. A formal specification is derived. In addition, we propose an extremely effective syllable lookahead for Mandarin. Together with the methods above, search space was reduced 12-15 times and state likelihood evaluations 4-9 times without significant error increase.
Frank Seide
IEEE Trans. Speech Audio Process.1
2005 Vocabulary-Independent Indexing of Spontaneous Speech
abstract
We present a system for vocabulary-independent indexing of spontaneous speech, i.e., neither do we know the vocabulary of a speech recording nor can we predict which query terms for which a user is going to search. The technique can be applied to information retrieval, information extraction, and data mining. Our specific target is search in recorded conversations in the office/information-worker scenario-teleconferences, meetings, presentations, and voice mails. The focus of this paper is on how to index phonetic lattices. We will show that an index should provide expected term frequencies (ETFs) of query terms. Since, at indexing time, it is unknown which phoneme sequences constitute valid query terms, we will introduce an approximation of ETFs of a query's phoneme sequence by M-gram phoneme language models, which are estimated on lattices and organized in an inverted index-like structure for fast access. We will discuss ranking, estimation, and integration of phoneme/word hybrid approaches. Compared with an unindexed baseline without approximation, our approximation leads only to a 3.4% relative loss of search accuracy on the Linguistic Data Consortium (LDC) voicemail task. We also propose a two-stage method for locating individual keyword occurrences using the above method as a fast match. A 20-times speedup is achieved over unindexed search at under a 2-point accuracy loss. Last, we will briefly introduce a prototype applet based on the above techniques.
Kaijiang Chen, Chengyuan Ma, Frank Seide
IEEE Trans. Speech Audio Process.4
2004 Vocabulary-independent search in spontaneous speech
abstract
For efficient organization of speech recordings - meetings, interviews, voice mails, lectures - the ability to search for spoken keywords is an essential capability. Today, most spoken-document retrieval systems use large-vocabulary recognition. For the above scenarios, such systems suffer from both the unpredictable vocabulary/domain and generally high word-error rates (WER). We present a vocabulary-independent system to index and to search rapidly spontaneous speech. A speech recognizer generates lattices of phonetic word fragments, against which keywords are matched phonetically. We first show the need to use recognition alternatives (lattices) in a high-WER context, on a word-based baseline. Then we introduce our new method of phonetic word-fragment lattice generation, which uses longer-span language knowledge than a phoneme recognizer. Last we introduce heuristics to compact the lattices to feasible sizes that can be searched efficiently. On the LDC voice mail corpus, we show that vocabulary/domain-independent phonetic search is as accurate as a vocabulary/domain-dependent word-lattice based baseline system for in-vocabulary keywords (FOMs of 74-75%), but nearly maintains this accuracy also for out-of-vocabulary keywords.
Frank Seide, Chengyuan Ma, Eric Chang
ICASSP (1)1
2004 A hybrid word / phoneme-based approach for improved vocabulary-independent search in spontaneous speech
abstract
For efficient organization of speech recordings – meetings, interviews, voice mails, and lectures – being able to search for spoken keywords is essential. Today, most spoken document retrieval systems use large-vocabulary recognition. For the above scenarios, such systems suffer from the unpredictable domain, out-ofvocabulary queries, and generally high word-error rate (WER). In [1], we presented a system for phonetic indexing and searching of spontaneous speech. It is vocabulary-independent and based on phoneme lattices. In the present paper, we propose to combine it with word-based search into a hybrid approach. We explore two methods of combination: posterior combination (merging search results of a word-based and a phoneme-based system) and prior combination (combining word and phoneme language models and vocabularies to form a hybrid recognizer). The search accuracy of our best purely phonetic baseline is 64% (Figure of Merit), and our purely word-based baselines are below 50%. The new hybrid approach achieves 73%, if the recognizer uses a language model that matches the test-set domain. With a mismatched language model, 71 % is achieved. Our results show that the proposed hybrid model benefits from the best of two worlds: Word-level language context and robustness of phonetic search to unknown words and domain mismatch. 1.
Frank Seide
INTERSPEECH2
2003 Coarticulation modeling by embedding a target-directed hidden trajectory model into HMM - MAP decoding and evaluation
abstract
The hidden dynamic model (HDM) has been an attractive acoustic modeling approach because it provides a computational model for coarticulation and the dynamics of human speech. However, the lack of a direct decoding algorithm has been a barrier to research progress on HDM. We have developed a new HDM-based acoustic model, the hidden-trajectory HMM (HTHMM), which combines the state/mixture topology of a traditional monophone HMM with a target-directed hidden-trajectory model (a special form of HDM) for coarticulation modeling. Because the classical Viterbi algorithm is not admissible, we have developed a novel MAP decoding algorithm for HTHMM that correctly takes the hidden continuous trajectory into account. This paper introduces our new HTHMM decoder that allows us for the first time to evaluate an HDM-type model by direct decoding instead of N-best rescoring. Using direct decoding, we demonstrate that the coarticulatory mechanism of our HTHMM matches traditional context-dependent modeling (enumeration of model parameters): The context-independent HTHMM has slightly better accuracy than a crossword-triphone HMM on the Aurora2 task. The decoder also enables us to include state-boundary optimization into the HDM/HTHMM training procedure. This paper presents the detailed decoding algorithm and evaluation results, while in Zhou et al. (2003) we present the HTHMM model itself and parameter training.
Frank Seide, Jian-Lai Zhou, Li Deng 0001
ICASSP (1)1
2003 Coarticulation modeling by embedding a target-directed hidden trajectory model into HMM - model and training
abstract
We propose and evaluate a new acoustic model that combines HMM and a special type of the hidden dynamic model (HDM) a target-directed hidden trajectory model - into a single integrated model named HTHMM. The new model provides a computational model of coarticulation by representing the internal dynamics of human speech based on the hidden trajectory of the vocal-tract resonances. This paper focuses on the general structure of the new model and the EM training procedure. The corresponding MAP decoding algorithm and more detailed evaluation are given in Seide et al. (2003). Speech recognition experimental results on the Aurora2 task demonstrated that the new model, although using only context-independent phoneme units (no context-dependent parameters), is still slightly superior in word error rate to the corresponding crossword triphone HMM. This provides the evidence that the coarticulatory mechanism represented by the HTHMM via the model structure matches the traditional context-dependent modeling approach based on enumeration of model parameters.
Jian-Lai Zhou, Frank Seide, Li Deng 0001
ICASSP (1)2
2003 An improved model-based speaker segmentation system
Frank Seide, Chengyuan Ma, Eric Chang
INTERSPEECH2
2002 A system for spoken query information retrieval on mobile devices
abstract
With the proliferation of handheld devices, information access on mobile devices is a topic of growing relevance. This paper presents a system that allows the user to search for information on mobile devices using spoken natural-language queries. We explore several issues related to the creation of this system, which combines state-of-the-art speech-recognition and information-retrieval technologies. This is the first work that we are aware of which evaluates spoken query based information retrieval on a commonly available and well researched text database, the Chinese news corpus used in the National Institute of Standards and Technology (NIST)s TREC-5 and TREC-6 benchmarks. To compare spoken-query retrieval performance for different relevant scenarios and recognition accuracies, the benchmark queries-read verbatim by 20 speakers-were recorded simultaneously through three channels: headset microphone, PDA microphone, and cellular phone. Our results show that for mobile devices with high-quality microphones, spoken-query retrieval based on existing technologies yields retrieval precisions that come close to that for perfect text input (mean average precision 0.459 and 0.489, respectively, on TREC-6).
Eric Chang, Frank Seide, Helen M. Meng, Zhuoran Chen, Yu Shi 0001, Yuk-Chi Li
IEEE Trans. Speech Audio Process.2
2001 Rapid speaker adaptation using a priori knowledge by eigenspace analysis of MLLR parameters
abstract
This paper considers the problem of rapid speaker adaptation in speech recognition. In particular, we exploit an approach based on combination of transformations, which utilizes the concepts of both maximum likelihood linear regression (MLLR) and eigenvoice adaptation. We analyze three different possible methods to realize the concept, and formulate a fast algorithm of maximum likelihood coefficient estimation for test speakers. It is found that the best approach can properly utilize the a priori knowledge of speaker-independent models in constructing the eigenspace for speaker characteristics, while using MLLR matrices in representing the specific speakers so as to reduce the on-line memory and computation requirement of the adaptation phase. This best approach leads to identical models relative to eigenvoice adaptation that is based on MLLR-adapted speaker models. The experimental results and discussions also provide a good analysis towards integration of the MLLR and eigenvoice approaches.
Nick Jui-Chang Wang, Sammy S.-M. Lee, Frank Seide, Lin-Shan Lee
ICASSP3
2000 Pitch tracking and tone features for Mandarin speech recognition
abstract
Tone modeling is a critical component for Mandarin large-vocabulary continuous-speech recognition systems. This paper presents an efficient real-time pitch tracker and a set of tone features that achieve a vast 30% reduction of the character error rate (CER), compared to the non-tonal baseline. To our knowledge, this is the highest improvement from tones ever reported for Mandarin. The paper first discusses adapting a known pitch-tracking algorithm for real-time operation. Second, we study the derivation of tone features for Mandarin LVCSR. Compared to the baseline vector (F/sub 0/, /spl Delta/F/sub 0/), our best tone features lead to a 28% reduction of tone errors. Results are shown for three LVCSR databases, including the Chinese 1998 National Performance Assessment (Project 863) and the Taiwan telephony database "MAT." Performance of Western-language systems is reached, and for the "863 System Performance Test," our system achieves 1.5% CER.
Hank Chang-Han Huang, Frank Seide
ICASSP2
2000 Improvements of the Philips 2000 Taiwan Mandarin benchmark system
abstract
In this paper, we present the Philips large vocabulary continuous Mandarin speech recognition system developed for the 2000 Taiwan Speech Input Technology Assessment. We systematically integrated key Mandarin components with up-todate Western-language techniques to build up a state-of-the-art Mandarin speech recognition system. These technologies include robust pitch extraction/tone modeling, context-dependent preme/core-final units, Chinese phrase/syllable trigram language model, linear discriminant analysis ( LDA), cross-syllable modeling/decoding, speaker clustering and maximum likelihood linear regression (MLLR) adaptation. Among them, the major breakthroughs were our robust pitch extraction/tone modeling technology and the treatment of coarticulation across syllable boundaries. For the development set, we dramatically reduced last year’s best error rates by relative 44.8%~67.8% on all three categories we participated. Moreover, for the evaluation set, we achieved the lowest unit error rates on all three categories.
Yuan-Fu Liao, Nick Jui-Chang Wang, Max Huang, Hank Huang, Frank Seide
INTERSPEECH5
2000 Two-stream modeling of Mandarin tones
Frank Seide, Nick Jui-Chang Wang
INTERSPEECH1
2000 MAT-2000 - design, collection, and validation of a Mandarin 2000-speaker telephone speech database
Hsiao-Chuan Wang, Frank Seide, Chiu-yu Tseng, Lin-Shan Lee
INTERSPEECH2
2000 The thoughtful elephant: strategies for spoken dialog systems
abstract
We present technology used in spoken dialog systems for applications of a wide range. They include tasks from the travel domain and automatic switchboards as well as large scale directory assistance. The overall goal in developing spoken dialog systems is to allow for a natural and flexible dialog flow similar to human-human interaction. This imposes the challenging task to recognize and interpret user input, where he/she is allowed to choose from an unrestricted vocabulary and an infinite set of possible formulations. We therefore put emphasis on strategies that make the system more robust while still maintaining a high level of naturalness and flexibility. In view of this paradigm, we found that two fundamental principles characterize many of the proposed methods: to consider available sources of information as early as possible; and to keep alternative hypotheses and delay the decision for a single option as long as possible. We describe how our system architecture caters to incorporating application specific knowledge, including, for example, database constraints, in the determination of the best sentence hypothesis for a user turn. On the next higher level, we use the dialog history to assess the plausibility of a sentence hypothesis by applying consistency checks with information items from previous user turns. In particular, we demonstrate how combination decisions over several turns can be exploited to boost the recognition performance of the system.
Bernd Souvignier, Andreas Kellner, Bernhard Rüber, Hauke Schramm, Frank Seide
IEEE Trans. Speech Audio Process.5
1999 Development of the philips 1999 taiwan Mandarin benchmark system
abstract
An automatic system for detection of pronunciation errors by adult learners of English is embedded in a language–learning package. Four main features are: (1) a recognizer robust to non–native speech; (2) localization of phone– and word–level errors; (3) diagnosis of what sorts of phone–level errors took place; and (4) a lexical–stress detector. These tools together allow robust, consistent, and specific feedback on pronunciation errors, unlike many previous systems that provide feedbaconly at a more general level. The diagnosis technique searches for errors expected based on the student’s mother tongue and uses a separate bias for each error in order to maintain a particular desired global false alarm rate. Results are presented here for non–native recognition on tasks of differing complexity and for diagnosis, based on a data set of artificial errors, showing that this method can detect many contrasts with a high hit rate and a low false alarm rate.
Chiwei Che, Nick Jui-Chang Wang, Max Huang, Hank Huang, Frank Seide
EUROSPEECH5
1997 Towards an automated directory information system
abstract
This paper describes a design and feasibility study for a large-scale automatic directory information system with a scalable architecture. The current demonstrator, called PADIS-XL 1, operates in realtime and handles a database of a medium-size German city with 130,000 listings. The system uses a new technique of taking a combined decision on the joint probability over multiple dialogue turns, and a dialogue strategy that strives to restrict the search space more and more with every dialogue turn. During the course of the dialogue, the last name of the desired subscriber must be spelled out. The spelling recognizer permits continuous spelling and uses a context-free grammar to parse common spelling expressions. This paper describes the system architecture, our maximum a-posteriori (MAP) decision rule, the spelling grammar, and the dialogue strategy. We give results on the SPEECHDAT and SIETILL databases on recognition of first names by spelling and on jointly deciding on the spelled and the spoken name. In a 35,000-names setup, the joint decision reduced name-recognition errors by 31%. 1.
Frank Seide, Andreas Kellner
EUROSPEECH1
1997 PADIS - An automatic telephone switchboard and directory information system
Andreas Kellner, Bernhard Rüber, Frank Seide, Bach-Hiep Tran
Speech Commun.3
1996 A comparison of time conditioned and word conditioned search techniques for large vocabulary speech recognition
abstract
In this paper, we compare the search effort of the word conditioned and the time conditioned tree search methods.Both methods are based on a time-synchronous, left-to-right beam search using a treeorganized lexicon.Whereas the word conditioned method is well known and widely used, the time conditioned method is novel in the context of 20 000-word vocabulary recognition.We extend both methods to handle trigram language models in a one-pass strategy.Both methods were tested on a train schedule inquiry task (1 850 words, telephone speech) and on the North American Business (Nov.'94)development corpus (20 000 words).
Stefan Ortmanns, Hermann Ney, Frank Seide, Ingo Lindam
ICSLP3
1996 Improving speech understanding by incorporating database constraints and dialogue history
abstract
In the course of a (man-machine) dialogue, the system's belief concerning the user's intention is continuously being built up.Moreover, restricting the discourse to a narrow application domain further constrains the variety of possible user reactions.In this paper, we will show h o w these knowledge sources may be utilized in a stochastic framework to improve speech understanding.On eld-test data collected with our automatic exchange board prototype PADIS 1 , a relative reduction of attribute errors by 27% has been obtained.
Frank Seide, Bernhard Rüber, Andreas Kellner
ICSLP1
1996 A word graph based n-best search in continuous speech recognition
abstract
In this paper, we i n troduce an ecient algorithm for the exhaustive search of N best sentence hypotheses in a word graph.The search procedure is based on a two-pass algorithm.In the rst pass, a word graph is constructed with standard time-synchronous beam search.The actual extraction of N best word sequences from the word graph takes place during the second pass.With our implementation of a tree-organized N-Best list, the search is performed directly on the resulting word graph.Therefore, the parallel bookkeeping of N hypotheses at each processing step during the search is not necessary.It is important to point out that the proposed N-Best search algorithm produces an exact N-Best list as dened by the word graph structure.Possible errors can only result from pruning during the construction of the word graph.In a postprocessing step, the N candidates can be rescored with a more complex language model with highly reduced computational cost.This algorithm is also applied in speech understanding to select the most likely sentence hypothesis that satises some additional constraints.
Bach-Hiep Tran, Frank Seide, Volker Steinbiss
ICSLP2
1995 Fast likelihood computation for continuous-mixture densities using a tree-based nearest neighbor search
Frank Seide
EUROSPEECH1
1995 The Philips automatic train timetable information system
Harald Aust, Martin Oerder, Frank Seide, Volker Steinbiss
Speech Commun.3
1994 Non-linear regression based feature extraction for connected-word recognition in noise
abstract
This paper shows the application of non-linear regression to robust feature extraction for noisy speech recognition. In this approach, a non-linear estimator is used to compute noise invariant features from non-linear combinations of noise contaminated observations. The observations may be short-term subband-energies obtained from a filter bank analysis, cepstral coefficients of linear prediction coefficients. Instead of training the hidden Markov models (HMMs) under various noise conditions, they can be trained with clean data. The results show that this method leads to error rates comparable to those achieved by training in the presence of noise.>
Frank Seide, Alfred Mertins
ICASSP (2)1