VLDB 2026 Research / reviewers in the wild / expert
George Saon
dblp:52/6787
· DBLP profile ↗
126ranked-venue papers
41as first author
30since 2021 · last 2025
0009-0004-6837-5009ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 112 · 35 first-author · 28 since 2021Artificial intelligence and machine learning · 73 · 24 first-author · 17 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Granite-speech: open-source speech-aware LLMs with strong English ASR capabilitiesabstractGranite-speech LLMs are compact and efficient speech language models specifically designed for English ASR1and automatic speech translation (AST). The models were trained by modality aligning granite-3.3-instruct to speech on publicly available open-source corpora. Comprehensive benchmarking on English ASR shows that they outperform several competitors’ models that were trained on orders of magnitude more proprietary data, and they keep pace on English-to-X AST for major European languages, Japanese, and Mandarin. The speech-specific components are: a conformer acoustic encoder using block attention and self-conditioning trained with connectionist temporal classification, a windowed query-transformer speech modality adapter used to do temporal downsampling of the acoustic embeddings and map them to the LLM text embedding space, and LoRA adapters to further fine-tune the text LLM. The models are freely available on HuggingFace2under a permissive Apache 2.0 license.1The latest models (revision 3.3.2) support multilingual ASR in English, French, German, Spanish and Portuguese and bidirectional speech translation to and from English. This paper covers the initial English-only release.2https://huggingface.co/ibm-granite/granite-speech-3.3-2b (and…-8b). George Saon, Avihu Dekel, Alexi Brooks, Tohru Nagano, Abraham Daniels, Aharon Satt, Ashish R. Mittal, Brian Kingsbury, David Haws, Edmilson da Silva Morais, Gakuto Kurata, Hagai Aronowitz, Ibrahim Ibrahim, Hong-Kwang Jeff Kuo, Kate Soule, Luis A. Lastras, Masayuki Suzuki, Ron Hoory, Samuel Thomas 0001, Sashi Novitasari, Takashi Fukuda, Vishal Sunder, Zvi Kons |
ASRU | 1 |
| 2025 | Knowledge Distillation Based Training of Unified Conformer CTC Models for Multi-form ASRabstractThere is an on-going body of research on training separate dedicated models for either short-form or long-form utterances. Multi-form acoustic models that are simply trained on combined data from long-form and short-form utterances often suffer from various negative impacts due to the diversity of a speaking style, an accent, and a recording condition. In addition, a linguistic mismatch that comes from an utterance length is also another factor of the degradation. In this paper we investigate novel techniques for training unified Conformer-based models on multi-form speech data obtained from diverse domains and sources to serve multiple downstream applications with a single model. Our approach incorporates chunk-wise short-term discriminative knowledge distillation with an encoder embedding masking and mitigates the aforementioned problems that appear for single unified models. We show the benefit of our proposed technique on long and short-form ASR test sets by comparing our models against several variants trained by mixing utterances with various audio lengths. The proposed technique provides a significant improvement of up to 8.5% relative WER reduction over baseline systems that operate at a similar decoding cost. Takashi Fukuda, Gakuto Kurata, George Saon |
ICASSP | 3 |
| 2025 | LLM based Text Generation for Improved Low-resource Speech Recognition ModelsabstractLimited transcribed spoken style data is a critical bottleneck in building automatic speech recognition (ASR) systems for low-resource languages. Prompting a large language model (LLM) to paraphrase input text can generate novel text data that is constrained to be semantically similar to the source data. We leverage this capability of LLMs to improve the performance of low-resource ASR systems by increasing the limited text training data while keeping the same spoken style. Since word sequences in the training data are now more diverse and the vocabulary of the ASR model is also expanded, this approach allows for building general purpose ASR without prior knowledge of various domains in the low-resource language. In our experiments with Brazilian Portuguese as a low-resource language, paraphrased data enhanced the n-gram language model (LM) used to build the weighted finite state transducer (WFST) for decoding with a Conformer-CTC speech recognition model, resulting in improvement of word error rate (WER) by 15.6% over the baseline model. Synthesizing the paraphrased text into speech and using it to fine-tune the acoustic model (AM) component helped to further improve the WER by 2.9%, achieving a combined improvement of 18.5%. We also demonstrate the usefulness of our proposed approach for high-resource languages like English. Tohru Nagano, Gakuto Kurata, Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Daniel Bolaños, Hyun Jung, George Saon |
ICASSP | 7 |
| 2025 | A Non-autoregressive Model for Joint STT and TTSabstractIn this paper, we take a step towards jointly modeling automatic speech recognition (STT) and speech synthesis (TTS) in a fully non-autoregressive way. We develop a novel multimodal framework capable of handling the speech and text modalities as input either individually or together. The proposed model can also be trained with unpaired speech or text data owing to its multimodal nature. We further propose an iterative refinement strategy to improve the STT and TTS performance of our model such that the partial hypothesis at the output can be fed back to the input of our model, thus iteratively improving both STT and TTS predictions. We show that our joint model can effectively perform both STT and TTS tasks, outperforming the STT-specific baseline in all tasks and performing competitively with the TTS-specific baseline across a wide range of evaluation metrics. Vishal Sunder, Brian Kingsbury, George Saon, Samuel Thomas 0001, Slava Shechtman, Hagai Aronowitz, Eric Fosler-Lussier, Luis A. Lastras |
ICASSP | 3 |
| 2025 | Exploring the Limits of Conformer CTC-Encoder for Speech Emotion Recognition using Large Language Models
Edmilson da Silva Morais, Hagai Aronowitz, Aharon Satt, Ron Hoory, Avihu Dekel, Brian Kingsbury, George Saon |
INTERSPEECH | 7 |
| 2024 | Semi-Autoregressive Streaming ASR with Label ContextabstractNon-autoregressive (NAR) modeling has gained significant interest in speech processing since these models achieve dramatically lower inference time than autoregressive (AR) models while also achieving good transcription accuracy. Since NAR automatic speech recognition (ASR) models must wait for the completion of the entire utterance before processing, some works explore streaming NAR models based on blockwise attention for low-latency applications. However, streaming NAR models significantly lag in accuracy compared to streaming AR and non-streaming NAR models. To address this, we propose a streaming "semi-autoregressive" ASR model that incorporates the labels emitted in previous blocks as additional context using a Language Model (LM) subnetwork. We also introduce a novel greedy decoding algorithm that addresses insertion and deletion errors near block boundaries while not significantly increasing the inference time. Experiments show that our method outperforms the existing streaming NAR model by 19% relative on Tedlium2, 16%/8% on Librispeech-100 clean/other test sets, and 19%/8% on the Switchboard(SWB)/Callhome(CH) test sets. It also reduced the accuracy gap with streaming AR and non-streaming NAR models while achieving 2.5x lower latency. We also demonstrate that our approach can effectively utilize external text data to pre-train the LM subnetwork to further improve streaming ASR accuracy. Siddhant Arora, George Saon, Shinji Watanabe 0001, Brian Kingsbury |
ICASSP | 2 |
| 2024 | Multiple Representation Transfer from Large Language Models to End-to-End ASR SystemsabstractTransferring the knowledge of large language models (LLMs) is a promising technique to incorporate linguistic knowledge into end-to-end automatic speech recognition (ASR) systems. However, existing works only transfer a single representation of LLM (e.g. the last layer of pretrained BERT), while the representation of a text is inherently non-unique and can be obtained variously from different layers, contexts and models. In this work, we explore a wide range of techniques to obtain and transfer multiple representations of LLMs into a transducer-based ASR system. While being conceptually simple, we show that transferring multiple representations of LLMs can be an effective alternative to transferring only a single LLM representation. Takuma Udagawa, Masayuki Suzuki, Gakuto Kurata, Masayasu Muraoka, George Saon |
ICASSP | 5 |
| 2024 | Exploring the limits of decoder-only models trained on public speech recognition corpora
Ankit Gupta 0001, George Saon, Brian Kingsbury |
INTERSPEECH | 2 |
| 2023 | Speech-enriched Memory for Inference-time Adaptation of ASR Models to Word DictionariesabstractDespite the impressive performance of ASR models on mainstream benchmarks, their performance on rare words is unsatisfactory.In enterprise settings, often a focused list of entities (such as locations, names, etc) are available which can be used to adapt the model to the terminology of specific domains.In this paper, we present a novel inference algorithm that improves the prediction of state-of-the-art ASR models using nearest-neighbor-based matching on an inference-time word list.We consider both the Transducer architecture that is useful in the streaming setting, and state-of-the-art encoder-decoder models such as Whisper.In our approach, a list of rare entities is indexed in a memory by synthesizing speech for each entry, and then storing the internal acoustic and language model states obtained from the best possible alignment on the ASR model.The memory is organized as a trie which we harness to perform a stateful lookup during inference.A key property of our extension is that we prevent spurious matches by restricting to only word-level matches.In our experiments on publicly available datasets and private benchmarks, we show that our method is effective in significantly improving rare word recognition. Ashish R. Mittal, Sunita Sarawagi, Preethi Jyothi, George Saon, Gakuto Kurata |
EMNLP | 4 |
| 2023 | Diagonal State Space Augmented Transformers for Speech RecognitionabstractWe improve on the popular conformer architecture by replacing the depthwise temporal convolutions with diagonal state space (DSS) models. DSS is a recently introduced variant of linear RNNs obtained by discretizing a linear dynamical system with a diagonal state transition matrix. DSS layers project the input sequence onto a space of orthogonal polynomials where the choice of basis functions, metric and support is controlled by the eigenvalues of the transition matrix. We compare neural transducers with either conformer or our proposed DSS-augmented transformer (DSSformer) encoders on three public corpora: Switchboard English conversational telephone speech 300 hours, Switchboard+Fisher 2000 hours, and a spoken archive of holocaust survivor testimonials called MALACH 176 hours. On Switchboard 300/2000 hours, we reach a single model performance of 8.9%/6.7% WER on the combined test set of the Hub5 2000 evaluation, respectively, and on MALACH we improve the WER by 7% relative over the previous best published result. In addition, we present empirical evidence suggesting that DSS layers learn damped Fourier basis functions where the attenuation coefficients are layer specific whereas the frequency coefficients converge to almost identical linearly-spaced values across all layers. George Saon, Ankit Gupta 0001 |
ICASSP | 1 |
| 2023 | Multi-Speaker Data Augmentation for Improved end-to-end Automatic Speech RecognitionabstractPublicly available datasets traditionally used to train E2E ASR models for conversational telephone speech recognition are based on clean, short duration, single speaker utterances collected on separate channels. While E2E ASR models achieve state-of-the-art performance on recognition tasks that match well with such training data, they are observed to fail on test recordings that contain multiple speakers, significant channel or background noise or span longer durations than training data utterances. To mitigate these issues, we propose an on-the-fly data augmentation strategy that transforms single speaker training data into multiple speaker data by appending together multiple single speaker utterances. The proposed technique encourages the E2E model to become robust to speaker changes and also process longer utterances effectively. During training, the model is also guided by a teacher model trained on single speaker utterances to map its multi-speaker encoder embeddings to better performing single speaker representations. With the proposed technique we obtain 7-14% relative improvement on various single speaker and multiple speaker test sets. We also show how this technique is able to improve recognition performance by up to 14% by capturing useful information from preceding spoken utterances used as dialog history. Samuel Thomas 0001, Hong-Kwang Jeff Kuo, George Saon, Brian Kingsbury |
ICASSP | 3 |
| 2023 | Improving RNN Transducer Acoustic Models for English Conversational Speech Recognition
George Saon, Brian Kingsbury |
INTERSPEECH | 2 |
| 2022 | Speech Recognition Using Biologically-Inspired Neural NetworksabstractAutomatic speech recognition systems (ASR), such as the recurrent neural network transducer (RNN-T), have reached close to human-like performance and are deployed in commercial applications. However, their core operations depart from the powerful biological counterpart, the human brain. On the other hand, the current developments in biologically-inspired ASR models lag behind in terms of accuracy and focus primarily on small-scale applications. In this work, we revisit the incorporation of biologically-plausible models into deep learning and enhance their capabilities, by taking inspiration from the brain’s diverse neural and synaptic dynamics. In particular, we propose novel deep learning units by introducing neural connectivity concepts emulating the axo-somatic and the axo-axonic synapses and integrate them into the RNN-T architecture. We demonstrate for the first time that such a model can yield performance levels competitive to the state-of-the-art. Moreover, our implementation has a significantly reduced computational cost and a lower latency. Thomas Bohnstingl, Ayush Garg 0006, Stanislaw Wozniak, George Saon, Evangelos Eleftheriou, Angeliki Pantazi |
ICASSP | 4 |
| 2022 | Improving End-to-end Models for Set Prediction in Spoken Language UnderstandingabstractThe goal of spoken language understanding (SLU) systems is to determine the meaning of the input speech signal, unlike speech recognition which aims to produce verbatim transcripts. Advances in end-to-end (E2E) speech modeling have made it possible to train solely on semantic entities, which are far cheaper to collect than verbatim transcripts. We focus on this set prediction problem, where entity order is unspecified. Using two classes of E2E models, RNN transducers and attention based encoder-decoders, we show that these models work best when the training entity sequence is arranged in spoken order. To improve E2E SLU models when entity spoken order is unknown, we propose a novel data augmentation technique along with an implicit attention based alignment method to infer the spoken order. F1 scores significantly increased by more than 11% for RNN-T and about 2% for attention based encoder-decoder SLU models, outperforming previously reported results. Hong-Kwang Jeff Kuo, Zoltán Tüske, Samuel Thomas 0001, Brian Kingsbury, George Saon |
ICASSP | 5 |
| 2022 | Towards Reducing the Need for Speech Training Data to Build Spoken Language Understanding SystemsabstractThe lack of speech data annotated with labels required for spoken language understanding (SLU) is often a major hurdle in building end-to-end (E2E) systems that can directly process speech inputs. In contrast, large amounts of text data with suitable labels are usually available. In this paper, we propose a novel text representation and training methodology that allows E2E SLU systems to be effectively constructed using these text resources. With very limited amounts of additional speech, we show that these models can be further improved to perform at levels close to similar systems built on the full speech datasets. The efficacy of our proposed approach is demonstrated on both intent and entity tasks using three different SLU datasets. With text-only training, the proposed system achieves up to 90% of the performance possible with full speech training. With just an additional 10% of speech data, these models significantly improve further to 97% of full performance. Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Brian Kingsbury, George Saon |
ICASSP | 4 |
| 2022 | Integrating Text Inputs for Training and Adapting RNN Transducer ASR ModelsabstractCompared to hybrid automatic speech recognition (ASR) systems that use a modular architecture in which each component can be in-dependently adapted to a new domain, recent end-to-end (E2E) ASR system are harder to customize due to their all-neural monolithic construction. In this paper, we propose a novel text representation and training framework for E2E ASR models. With this approach, we show that a trained RNN Transducer (RNN-T) model’s internal LM component can be effectively adapted with text-only data. An RNN-T model trained using both speech and text inputs improves over a baseline model trained on just speech with close to 13% word error rate (WER) reduction on the Switchboard and CallHome test sets of the NIST Hub5 2000 evaluation. The usefulness of the proposed approach is further demonstrated by customizing this general purpose RNN-T model to three separate datasets. We observe 20-45% relative word error rate (WER) reduction in these settings with this novel LM style customization technique using only unpaired text data from the new domains. Samuel Thomas 0001, Brian Kingsbury, George Saon, Hong-Kwang Jeff Kuo |
ICASSP | 3 |
| 2022 | Improving Generalization of Deep Neural Network Acoustic Models with Length Perturbation and N-best Based Label SmoothingabstractWe introduce two techniques, length perturbation and n-best based label smoothing, to improve generalization of deep neural network (DNN) acoustic models for automatic speech recognition (ASR).Length perturbation is a data augmentation algorithm that randomly drops and inserts frames of an utterance to alter the length of the speech feature sequence.N-best based label smoothing randomly injects noise to ground truth labels during training in order to avoid overfitting, where the noisy labels are generated from n-best hypotheses.We evaluate these two techniques extensively on the 300-hour Switchboard (SWB300) dataset and an in-house 500-hour Japanese (JPN500) dataset using recurrent neural network transducer (RNNT) acoustic models for ASR.We show that both techniques improve the generalization of RNNT models individually and they can also be complementary.In particular, they yield good improvements over a strong SWB300 baseline and give state-of-art performance on SWB300 using RNNT models. George Saon, Tohru Nagano, Masayuki Suzuki, Takashi Fukuda, Brian Kingsbury, Gakuto Kurata |
INTERSPEECH | 2 |
| 2022 | Accelerating Inference and Language Model Fusion of Recurrent Neural Network Transducers via End-to-End 4-bit QuantizationabstractWe report on aggressive quantization strategies that greatly accelerate inference of Recurrent Neural Network Transducers (RNN-T).We use a 4 bit integer representation for both weights and activations and apply Quantization Aware Training (QAT) to retrain the full model (acoustic encoder and language model) and achieve near-iso-accuracy.We show that customized quantization schemes that are tailored to the local properties of the network are essential to achieve good performance while limiting the computational overhead of QAT.Density ratio Language Model fusion has shown remarkable accuracy gains on RNN-T workloads but it severely increases the computational cost of inference.We show that our quantization strategies enable using large beam widths for hypothesis search while achieving streaming-compatible runtimes and a full model compression ratio of 7.6× compared to the full precision model.Via hardware simulations, we estimate a 3.4× acceleration from FP16 to INT4 for the end-to-end quantized RNN-T inclusive of LM fusion, resulting in a Real Time Factor (RTF) of 0.06.On the NIST Hub5 2000, Hub5 2001, and RT-03 test sets, we retain most of the gains associated with LM fusion, improving the average WER by >1.5%. Andrea Fasoli, Chia-Yu Chen, Mauricio J. Serrano, Swagath Venkataramani, George Saon, Brian Kingsbury, Kailash Gopalakrishnan |
INTERSPEECH | 5 |
| 2022 | Global RNN Transducer Models For Multi-dialect Speech Recognition
Takashi Fukuda, Samuel Thomas 0001, Masayuki Suzuki, Gakuto Kurata, George Saon, Brian Kingsbury |
INTERSPEECH | 5 |
| 2022 | Extending RNN-T-based speech recognition systems with emotion and language classificationabstractSpeech transcription, emotion recognition, and language identification are usually considered to be three different tasks.Each one requires a different model with a different architecture and training process.We propose using a recurrent neural network transducer (RNN-T)-based speech-to-text (STT) system as a common component that can be used for emotion recognition and language identification as well as for speech recognition.Our work extends the STT system for emotion classification through minimal changes, and shows successful results on the IEMOCAP and MELD datasets.In addition, we demonstrate that by adding a lightweight component to the RNN-T module, it can also be used for language identification.In our evaluations, this new classifier demonstrates state-of-the-art accuracy for the NIST-LRE-07 dataset. Zvi Kons, Hagai Aronowitz, Edmilson da Silva Morais, Matheus Damasceno, Hong-Kwang Jeff Kuo, Samuel Thomas 0001, George Saon |
INTERSPEECH | 7 |
| 2022 | VQ-T: RNN Transducers using Vector-Quantized Prediction Network StatesabstractBeam search, which is the dominant ASR decoding algorithm for end-to-end models, generates tree-structured hypotheses.However, recent studies have shown that decoding with hypothesis merging can achieve a more efficient search with comparable or better performance.But, the full context in recurrent networks is not compatible with hypothesis merging.We propose to use vector-quantized long short-term memory units (VQ-LSTM) in the prediction network of RNN transducers.By training the discrete representation jointly with the ASR network, hypotheses can be actively merged for lattice generation.Our experiments on the Switchboard corpus show that the proposed VQ RNN transducers improve ASR performance over transducers with regular prediction networks while also producing denser lattices with a very low oracle word error rate (WER) for the same beam size.Additional language model rescoring experiments also demonstrate the effectiveness of the proposed lattice generation scheme. Jiatong Shi, George Saon, David Haws, Shinji Watanabe 0001, Brian Kingsbury |
INTERSPEECH | 2 |
| 2022 | Effect and Analysis of Large-scale Language Model Rescoring on Competitive ASR Systems
Takuma Udagawa, Masayuki Suzuki, Gakuto Kurata, Nobuyasu Itoh, George Saon |
INTERSPEECH | 5 |
| 2021 | RNN Transducer Models for Spoken Language UnderstandingabstractWe present a comprehensive study on building and adapting RNN transducer (RNN-T) models for spoken language understanding (SLU). These end-to-end (E2E) models are constructed in three practical settings: a case where verbatim transcripts are available, a constrained case where the only available annotations are SLU labels and their values, and a more restrictive case where transcripts are available but not corresponding audio. We show how RNN-T SLU models can be developed starting from pre-trained automatic speech recognition (ASR) systems, followed by an SLU adaptation step. In settings where real audio data is not available, artificially synthesized speech is used to successfully adapt various SLU models. When evaluated on two SLU data sets, the ATIS corpus and a customer call center data set, the proposed models closely track the performance of other E2E models and achieve state-of-the-art results. Samuel Thomas 0001, Hong-Kwang Jeff Kuo, George Saon, Zoltán Tüske, Brian Kingsbury, Gakuto Kurata, Zvi Kons, Ron Hoory |
ICASSP | 3 |
| 2021 | Advancing RNN Transducer Technology for Speech RecognitionabstractWe investigate a set of techniques for RNN Transducers (RNN-Ts) that were instrumental in lowering the word error rate on three different tasks (Switchboard 300 hours, conversational Spanish 780 hours and conversational Italian 900 hours). The techniques pertain to architectural changes, speaker adaptation, language model fusion, model combination and general training recipe. First, we introduce a novel multiplicative integration of the encoder and prediction network vectors in the joint network (as opposed to additive). Second, we discuss the applicability of i-vector speaker adaptation to RNN-Ts in conjunction with data perturbation. Third, we explore the effectiveness of the recently proposed density ratio language model fusion for these tasks. Last but not least, we describe the other components of our training recipe and their effect on recognition performance. We report a 5.9% and 12.5% word error rate on the Switchboard and CallHome test sets of the NIST Hub5 2000 evaluation and a 12.7% WER on the Mozilla CommonVoice Italian test set. George Saon, Zoltán Tüske, Daniel Bolaños, Brian Kingsbury |
ICASSP | 1 |
| 2021 | Reducing Exposure Bias in Training Recurrent Neural Network TransducersabstractWhen recurrent neural network transducers (RNNTs) are trained using the typical maximum likelihood criterion, the prediction network is trained only on ground truth label sequences.This leads to a mismatch during inference, known as exposure bias, when the model must deal with label sequences containing errors.In this paper we investigate approaches to reducing exposure bias in training to improve the generalization of RNNT models for automatic speech recognition (ASR).A label-preserving input perturbation to the prediction network is introduced.The input token sequences are perturbed using SwitchOut and scheduled sampling based on an additional token language model.Experiments conducted on the 300-hour Switchboard dataset demonstrate their effectiveness.By reducing the exposure bias, we show that we can further improve the accuracy of a high-performance RNNT ASR model and obtain state-of-the-art results on the 300-hour Switchboard dataset. Brian Kingsbury, George Saon, David Haws, Zoltán Tüske |
Interspeech | 3 |
| 2021 | 4-Bit Quantization of LSTM-Based Speech Recognition ModelsabstractWe investigate the impact of aggressive low-precision representations of weights and activations in two families of large LSTM-based architectures for Automatic Speech Recognition (ASR): hybrid Deep Bidirectional LSTM -Hidden Markov Models (DBLSTM-HMMs) and Recurrent Neural Network -Transducers (RNN-Ts).Using a 4-bit integer representation, a naïve quantization approach applied to the LSTM portion of these models results in significant Word Error Rate (WER) degradation.On the other hand, we show that minimal accuracy loss is achievable with an appropriate choice of quantizers and initializations.In particular, we customize quantization schemes depending on the local properties of the network, improving recognition performance while limiting computational time.We demonstrate our solution on the Switchboard (SWB) and CallHome (CH) test sets of the NIST Hub5-2000 evaluation.DBLSTM-HMMs trained with 300 or 2000 hours of SWB data achieves <0.5% and <1% average WER degradation, respectively.On the more challenging RNN-T models, our quantization strategy limits degradation in 4-bit inference to 1.3%. Andrea Fasoli, Chia-Yu Chen, Mauricio J. Serrano, Xiao Sun 0013, Naigang Wang, Swagath Venkataramani, George Saon, Brian Kingsbury, Wei Zhang 0022, Zoltán Tüske, Kailash Gopalakrishnan |
Interspeech | 7 |
| 2021 | Integrating Dialog History into End-to-End Spoken Language Understanding SystemsabstractEnd-to-end spoken language understanding (SLU) systems that process human-human or human-computer interactions are often context independent and process each turn of a conversation independently. Spoken conversations on the other hand, are very much context dependent, and dialog history contains useful information that can improve the processing of each conversational turn. In this paper, we investigate the importance of dialog history and how it can be effectively integrated into end-to-end SLU systems. While processing a spoken utterance, our proposed RNN transducer (RNN-T) based SLU model has access to its dialog history in the form of decoded transcripts and SLU labels of previous turns. We encode the dialog history as BERT embeddings, and use them as an additional input to the SLU model along with the speech features for the current utterance. We evaluate our approach on a recently released spoken dialog data set, the HarperValleyBank corpus. We observe significant improvements: 8% for dialog action and 30% for caller intent recognition tasks, in comparison to a competitive context independent end-to-end baseline system. Jatin Ganhotra, Samuel Thomas 0001, Hong-Kwang Jeff Kuo, Sachindra Joshi, George Saon, Zoltán Tüske, Brian Kingsbury |
Interspeech | 5 |
| 2021 | Improving Customization of Neural Transducers by Mitigating Acoustic Mismatch of Synthesized Audio
Gakuto Kurata, George Saon, Brian Kingsbury, David Haws, Zoltán Tüske |
Interspeech | 2 |
| 2021 | On the Limit of English Conversational Speech RecognitionabstractIn our previous work we demonstrated that a single headed attention encoder-decoder model is able to reach state-of-the-art results in conversational speech recognition. In this paper, we further improve the results for both Switchboard 300 and 2000. Through use of an improved optimizer, speaker vector embeddings, and alternative speech representations we reduce the recognition errors of our LSTM system on Switchboard-300 by 4% relative. Compensation of the decoder model with the probability ratio approach allows more efficient integration of an external language model, and we report 5.9% and 11.5% WER on the SWB and CHM parts of Hub5'00 with very simple LSTM models. Our study also considers the recently proposed conformer, and more advanced self-attention based language models. Overall, the conformer shows similar performance to the LSTM; nevertheless, their combination and decoding with an improved LM reaches a new record on Switchboard-300, 5.0% and 10.0% WER on SWB and CHM. Our findings are also confirmed on Switchboard-2000, and a new state of the art is reported, practically reaching the limit of the benchmark. Zoltán Tüske, George Saon, Brian Kingsbury |
Interspeech | 2 |
| 2021 | Asynchronous Decentralized Distributed Training of Acoustic ModelsabstractLarge-scale distributed training of deep acoustic models plays an important role in today's high-performance automatic speech recognition (ASR). In this paper we investigate a variety of asynchronous decentralized distributed training strategies based on data parallel stochastic gradient descent (SGD) to show their superior performance over the commonly-used synchronous distributed training via allreduce, especially when dealing with large batch sizes. Specifically, we study three variants of asynchronous decentralized parallel SGD (ADPSGD), namely, fixed and randomized communication patterns on a ring as well as a delay-by-one scheme. We introduce a mathematical model of ADPSGD, give its theoretical convergence rate, and compare the empirical convergence behavior and straggler resilience properties of the three variants. Experiments are carried out on an IBM supercomputer for training deep long short-term memory (LSTM) acoustic models on the 2000-hour Switchboard dataset. Recognition and speedup performance of the proposed strategies are evaluated under various training configurations. We show that ADPSGD with fixed and randomized communication patterns cope well with slow learners. When learners are equally fast, ADPSGD with the delay-by-one strategy has the fastest convergence with large batches. In particular, using the delay-by-one strategy, we can train the acoustic model in less than 2 hours using 128 V100 GPUs with competitive word error rates. Wei Zhang 0022, Abdullah Kayi, Ulrich Finkler, Brian Kingsbury, George Saon, David S. Kung 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2020 | Alignment-Length Synchronous Decoding for RNN TransducerabstractWe present a beam decoding strategy for recurrent neural network transducers which has the characteristic that all competing hypotheses within the beam have the same alignment length (number of output symbols plus BLANK symbols). We contrast the proposed technique with time-synchronous decoding where the competing hypotheses within the beam correspond to the same input frames (but can have different length output sequences). Experiments on the Switchboard 2000 hours corpus show that alignment-length synchronous decoding (ALSD) is 25% faster than time-synchronous decoding (TSD) for the same accuracy because ALSD performs 42% fewer joint network evaluations and hypothesis expansions during the search. Additionally, we discuss the benefit of caching and batching the prediction and joint network evaluations, of using prefix trees instead of full output vocabulary expansions, and of performing hypothesis recombination after pruning. With open beam decoding, we reach a 6.2% / 10.9% word error rate on the Switchboard and CallHome Hub5 2000 evaluation testsets which compares favorably to other published single-model results on this corpus. George Saon, Zoltán Tüske, Kartik Audhkhasi |
ICASSP | 1 |
| 2020 | Improving Efficiency in Large-Scale Decentralized Distributed TrainingabstractDecentralized Parallel SGD (D-PSGD) and its asynchronous variant Asynchronous Parallel SGD (AD-PSGD) is a family of distributed learning algorithms that have been demonstrated to perform well for large-scale deep learning tasks. One drawback of (A)D-PSGD is that the spectral gap of the mixing matrix decreases when the number of learners in the system increases, which hampers convergence. In this paper, we investigate techniques to accelerate (A)D-PSGD based training by improving the spectral gap while minimizing the communication cost. We demonstrate the effectiveness of our proposed techniques by running experiments on the 2000-hour Switchboard speech recognition task and the ImageNet computer vision task. On an IBM P9 supercomputer, our system is able to train an LSTM acoustic model in 2.28 hours with 7.5% WER on the Hub5-2000 Switchboard (SWB) test set and 13.3% WER on the CallHome (CH) test set using 64 V100 GPUs and in 1.98 hours with 7.7% WER on SWB and 13.3% WER on CH using 128 V100 GPUs, the fastest training time reported to date. Wei Zhang 0022, Abdullah Kayi, Ulrich Finkler, Brian Kingsbury, George Saon, Youssef Mroueh, Alper Buyuktosunoglu, David S. Kung 0001, Michael Picheny |
ICASSP | 7 |
| 2020 | Knowledge Distillation from Offline to Streaming RNN Transducer for End-to-End Speech Recognition
Gakuto Kurata, George Saon |
INTERSPEECH | 2 |
| 2020 | Single Headed Attention Based Sequence-to-Sequence Model for State-of-the-Art Results on SwitchboardabstractIt is generally believed that direct sequence-to-sequence (seq2seq) speech recognition models are competitive with hybrid models only when a large amount of data, at least a thousand hours, is available for training. In this paper, we show that state-of-the-art recognition performance can be achieved on the Switchboard-300 database using a single headed attention, LSTM based model. Using a cross-utterance language model, our single-pass speaker independent system reaches 6.4% and 12.5% word error rate (WER) on the Switchboard and CallHome subsets of Hub5'00, without a pronunciation lexicon. While careful regularization and data augmentation are crucial in achieving this level of performance, experiments on Switchboard-2000 show that nothing is more useful than more data. Overall, the combination of various regularizations and a simple but fairly large model results in a new state of the art, 4.7% and 7.8% WER on the Switchboard and CallHome sets, using SWB-2000 without any external data resources. Zoltán Tüske, George Saon, Kartik Audhkhasi, Brian Kingsbury |
INTERSPEECH | 2 |
| 2019 | Simplified LSTMS for Speech RecognitionabstractIn this paper we explore new variants of Long Short-Term Memory (LSTM) networks for sequential modeling of acoustic features. In particular, we show that: (i) removing the output gate, (ii) replacing the hyperbolic tangent nonlinearity at the cell output with hard tanh, and (iii) collapsing the cell and hidden state vectors leads to a model that is conceptually simpler than and comparable in effectiveness to a regular LSTM for speech recognition. The proposed model has 25% fewer parameters than an LSTM with the same number of cells, trains faster because it has larger gradients leading to larger steps in weight space, and reaches a better optimum because there are fewer nonlinearities to traverse across layers. We report experimental results for both hybrid and CTC acoustic models on three publicly available English datasets: Switchboard 300 hours telephone conversations, 400 hours broadcast news transcription, and the MALACH 176 hours corpus of Holocaust survivor testimonies. In all cases the proposed models achieve similar or better accuracy than regular LSTMs while being conceptually simpler. George Saon, Zoltán Tüske, Kartik Audhkhasi, Brian Kingsbury, Michael Picheny, Samuel Thomas 0001 |
ASRU | 1 |
| 2019 | Sequence Noise Injected Training for End-to-end Speech RecognitionabstractWe present a simple noise injection algorithm for training end-to-end ASR models which consists in adding to the spectra of training utterances the scaled spectra of random utterances of comparable length. We conjecture that the sequence information of the "noise" utterances is important and verify this via a contrast experiment where the frames of the utterances to be added are randomly shuffled. Experiments for both CTC and attention-based models show that the pro-posed scheme results in up to 9% relative word error rate improvements (depending on the model and test set) on the Switchboard 300 hours English conversational telephony database. Additionally, we set a new benchmark for attention-based encoder-decoder models on this corpus. George Saon, Zoltán Tüske, Kartik Audhkhasi, Brian Kingsbury |
ICASSP | 1 |
| 2019 | English Broadcast News Speech Recognition by Humans and MachinesabstractWith recent advances in deep learning, considerable attention has been given to achieving automatic speech recognition performance close to human performance on tasks like conversational telephone speech (CTS) recognition. In this paper we evaluate the usefulness of these proposed techniques on broadcast news (BN), a similar challenging task. We also perform a set of recognition measurements to understand how close the achieved automatic speech recognition results are to human performance on this task. On two publicly available BN test sets, DEV04F and RT04, our speech recognition system using LSTM and residual network based acoustic models with a combination of n-gram and neural network language models performs at 6.5% and 5.9% word error rate. By achieving new performance milestones on these test sets, our experiments show that techniques developed on other related tasks, like CTS, can be transferred to achieve similar performance. In contrast, the best measured human recognition performance on these test sets is much lower, at 3.6% and 2.8% respectively, indicating that there is still room for new techniques and improvements in this space, to reach human performance levels. Samuel Thomas 0001, Masayuki Suzuki, Gakuto Kurata, Zoltán Tüske, George Saon, Brian Kingsbury, Michael Picheny, Tom Dibert, Alice Kaiser-Schatzlein, Bern Samko |
ICASSP | 6 |
| 2019 | Distributed Deep Learning Strategies for Automatic Speech RecognitionabstractIn this paper, we propose and investigate a variety of distributed deep learning strategies for automatic speech recognition (ASR) and evaluate them with a state-of-the-art Long short-term memory (LSTM) acoustic model on the 2000-hour Switchboard (SWB2000), which is one of the most widely used datasets for ASR performance benchmark. We first investigate what are the proper hyper-parameters (e.g., learning rate) to enable the training with sufficiently large batch size without impairing the model accuracy. We then implement various distributed strategies, including Synchronous (SYNC) , Asynchronous Decentralized Parallel SGD (ADPSGD) and the hybrid of the two HYBRID, to study their runtime/accuracy trade-off. We show that we can train the LSTM model using ADPSGD in 14 hours with 16 NVIDIA P100 GPUs to reach a 7.6% WER on the Hub5-2000 Switchboard (SWB) test set and a 13.1% WER on the Call-Home (CH) test set. Furthermore, we can train the model using HYBRID in 11.5 hours with 32 NVIDIA V100 GPUs without loss in accuracy. Wei Zhang 0022, Ulrich Finkler, Brian Kingsbury, George Saon, David S. Kung 0001, Michael Picheny |
ICASSP | 5 |
| 2019 | Forget a Bit to Learn Better: Soft Forgetting for CTC-Based Automatic Speech Recognition
Kartik Audhkhasi, George Saon, Zoltán Tüske, Brian Kingsbury, Michael Picheny |
INTERSPEECH | 2 |
| 2019 | Challenging the Boundaries of Speech Recognition: The MALACH CorpusabstractThere has been huge progress in speech recognition over the last several years. Tasks once thought extremely difficult, such as SWITCHBOARD, now approach levels of human performance. The MALACH corpus (LDC catalog LDC2012S05), a 375-Hour subset of a large archive of Holocaust testimonies collected by the Survivors of the Shoah Visual History Foundation, presents significant challenges to the speech community. The collection consists of unconstrained, natural speech filled with disfluencies, heavy accents, age-related coarticulations, un-cued speaker and language switching, and emotional speech - all still open problems for speech recognition systems. Transcription is challenging even for skilled human annotators. This paper proposes that the community place focus on the MALACH corpus to develop speech recognition systems that are more robust with respect to accents, disfluencies and emotional speech. To reduce the barrier for entry, a lexicon and training and testing setups have been created and baseline results using current deep learning technologies are presented. The metadata has just been released by LDC (LDC2019S11). It is hoped that this resource will enable the community to build on top of these baselines so that the extremely important information in these and related oral histories becomes accessible to a wider audience. Michael Picheny, Zoltán Tüske, Brian Kingsbury, Kartik Audhkhasi, George Saon |
INTERSPEECH | 6 |
| 2019 | Advancing Sequence-to-Sequence Based Speech Recognition
Zoltán Tüske, Kartik Audhkhasi, George Saon |
INTERSPEECH | 3 |
| 2019 | A Highly Efficient Distributed Deep Learning System for Automatic Speech RecognitionabstractModern Automatic Speech Recognition (ASR) systems rely on distributed deep learning to for quick training completion.To enable efficient distributed training, it is imperative that the training algorithms can converge with a large mini-batch size.In this work, we discovered that Asynchronous Decentralized Parallel Stochastic Gradient Descent (ADPSGD) can work with much larger batch size than commonly used Synchronous SGD (SSGD) algorithm.On commonly used public SWB-300 and SWB-2000 ASR datasets, ADPSGD can converge with a batch size 3X as large as the one used in SSGD, thus enable training at a much larger scale.Further, we proposed a Hierarchical-ADPSGD (H-ADPSGD) system in which learners on the same computing node construct a super learner via a fast allreduce implementation, and super learners deploy ADPSGD algorithm among themselves.On a 64 Nvidia V100 GPU cluster connected via a 100Gb/s Ethernet network, our system is able to train SWB-2000 to reach a 7.6% WER on the Hub5-2000 Switchboard (SWB) test-set and a 13.2% WER on the Callhome (CH) test-set in 5.2 hours.To the best of our knowledge, this is the fastest ASR training system that attains this level of model accuracy for SWB-2000 task to be ever reported in the literature. Wei Zhang 0022, Ulrich Finkler, George Saon, Abdullah Kayi, Alper Buyuktosunoglu, Brian Kingsbury, David S. Kung 0001, Michael Picheny |
INTERSPEECH | 4 |
| 2018 | Building Competitive Direct Acoustics-to-Word Models for English Conversational Speech RecognitionabstractDirect acoustics-to-word (A2W) models in the end-to-end paradigm have received increasing attention compared to conventional subword based automatic speech recognition models using phones, characters, or context-dependent hidden Markov model states. This is because A2W models recognize words from speech without any decoder, pronunciation lexicon, or externally-trained language model, making training and decoding with such models simple. Prior work has shown that A2W models require orders of magnitude more training data in order to perform comparably to conventional models. Our work also showed this accuracy gap when using the English Switchboard-Fisher data set. This paper describes a recipe to train an A2W model that closes this gap and is at-par with state-of-the-art sub-word based models. We achieve a word error rate of 8.8.8%/13.9% on the Hub5-2000 Switchboard/CallHome test sets without any decoder or language model. We find that model initialization, training data order, and regularization have the most impact on the A2W model performance. Next, we present a joint word-character A2W model that learns to first spell the word and then recognize it. This model provides a rich output to the user instead of simple word hypotheses, making it especially useful in the case of words unseen or rarely-seen during training. Kartik Audhkhasi, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, Michael Picheny |
ICASSP | 4 |
| 2017 | Language modeling with highway LSTMabstractLanguage models (LMs) based on Long Short Term Memory (LSTM) have shown good gains in many automatic speech recognition tasks. In this paper, we extend an LSTM by adding highway networks inside an LSTM and use the resulting Highway LSTM (HW-LSTM) model for language modeling. The added highway networks increase the depth in the time dimension. Since a typical LSTM has two internal states, a memory cell and a hidden state, we compare various types of HW-LSTM by adding highway networks onto the memory cell and/or the hidden state. Experimental results on English broadcast news and conversational telephone speech recognition show that the proposed HW-LSTM LM improves speech recognition accuracy on top of a strong LSTM LM baseline. We report 5.1% and 9.9% on the Switchboard and CallHome subsets of the Hub5 2000 evaluation, which reaches the best performance numbers reported on these tasks to date. Gakuto Kurata, Bhuvana Ramabhadran, George Saon, Abhinav Sethy |
ASRU | 3 |
| 2017 | Knowledge distillation across ensembles of multilingual models for low-resource languagesabstractThis paper investigates the effectiveness of knowledge distillation in the context of multilingual models. We show that with knowledge distillation, Long Short-Term Memory(LSTM) models can be used to train standard feed-forward Deep Neural Network (DNN) models for a variety of low-resource languages. We then examine how the agreement between the teacher's best labels and the original labels affects the student model's performance. Next, we show that knowledge distillation can be easily applied to semi-supervised learning to improve model performance. We also propose a promising data selection method to filter un-transcribed data. Then we focus on knowledge transfer among DNN models with multilingual features derived from CNN+DNN, LSTM, VGG, CTC and attention models. We show that a student model equipped with better input features not only learns better from the teacher's labels, but also outperforms the teacher. Further experiments suggest that by learning from each other, the original ensemble of various models is able to evolve into a new ensemble with even better combined performance. Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, Tom Sercu, Kartik Audhkhasi, Abhinav Sethy, Markus Nußbaum-Thom, Andrew Rosenberg |
ICASSP | 4 |
| 2017 | Network architectures for multilingual speech representation learningabstractMultilingual (ML) representations play a key role in building speech recognition systems for low resource languages. The IARPA sponsored BABEL program focuses on building speech recognition (ASR) and keyword search (KWS) systems in over 24 languages with limited training data. The most common mechanism to derive ML representations in the BABEL program has been with the use of a two-stage network, the first stage being a convolutional network (CNN) from where multilingual features are extracted, expanded contextually and used as input to the second stage which can be a feed-forward DNN or a CNN. The final multilingual representations are derived from the second network. This paper presents two novel methods for deriving ML representations. The first is based on Long-Short Term Memory (LSTM) networks and the second is based on a very deep CNN (VGG-net). We demonstrate that ML features extracted from both models show significant improvement over the baseline CNN-DNN based ML representations, in terms of both speech recognition and keyword search performance and draw the comparison between the LSTM model itself and the ML representations derived from it on Georgian, the surprise language for the OpenKWS evaluation. Tom Sercu, George Saon, Jia Cui, Bhuvana Ramabhadran, Brian Kingsbury, Abhinav Sethy |
ICASSP | 2 |
| 2017 | Direct Acoustics-to-Word Models for English Conversational Speech RecognitionabstractRecent work on end-to-end automatic speech recognition (ASR) has shown that the connectionist temporal classification (CTC) loss can be used to convert acoustics to phone or character sequences.Such systems are used with a dictionary and separately-trained Language Model (LM) to produce word sequences.However, they are not truly end-to-end in the sense of mapping acoustics directly to words without an intermediate phone representation.In this paper, we present the first results employing direct acoustics-to-word CTC models on two well-known public benchmark tasks: Switchboard and Call-Home.These models do not require an LM or even a decoder at run-time and hence recognize speech with minimal complexity.However, due to the large number of word output units, CTC word models require orders of magnitude more data to train reliably compared to traditional systems.We present some techniques to mitigate this issue.Our CTC word model achieves a word error rate of 13.0%/18.8%on the Hub5-2000 Switchboard/CallHome test sets without any LM or decoder compared with 9.6%/16.0%for phone-based CTC with a 4-gram LM.We also present rescoring results on CTC word model lattices to quantify the performance benefits of a LM, and contrast the performance of word and phone CTC models. Kartik Audhkhasi, Bhuvana Ramabhadran, George Saon, Michael Picheny, David Nahamoo |
INTERSPEECH | 3 |
| 2017 | Embedding-Based Speaker Adaptive Training of Deep Neural NetworksabstractAn embedding-based speaker adaptive training (SAT) approach is proposed and investigated in this paper for deep neural network acoustic modeling.In this approach, speaker embedding vectors, which are a constant given a particular speaker, are mapped through a control network to layer-dependent elementwise affine transformations to canonicalize the internal feature representations at the output of hidden layers of a main network.The control network for generating the speaker-dependent mappings is jointly estimated with the main network for the overall speaker adaptive acoustic modeling.Experiments on large vocabulary continuous speech recognition (LVCSR) tasks show that the proposed SAT scheme can yield superior performance over the widely-used speaker-aware training using i-vectors with speaker-adapted input features. Vaibhava Goel, George Saon |
INTERSPEECH | 3 |
| 2017 | Empirical Exploration of Novel Architectures and Objectives for Language Models
Gakuto Kurata, Abhinav Sethy, Bhuvana Ramabhadran, George Saon |
INTERSPEECH | 4 |
| 2017 | English Conversational Telephone Speech Recognition by Humans and MachinesabstractOne of the most difficult speech recognition tasks is accurate recognition of human to human communication. Advances in deep learning over the last few years have produced major speech recognition improvements on the representative Switchboard conversational corpus. Word error rates that just a few years ago were 14% have dropped to 8.0%, then 6.6% and most recently 5.8%, and are now believed to be within striking range of human performance. This then raises two issues - what IS human performance, and how far down can we still drive speech recognition error rates? A recent paper by Microsoft suggests that we have already achieved human performance. In trying to verify this statement, we performed an independent set of human performance measurements on two conversational tasks and found that human performance may be considerably better than what was earlier reported, giving the community a significantly harder goal to achieve. We also report on our own efforts in this area, presenting a set of acoustic and language modeling techniques that lowered the word error rate of our own English conversational telephone LVCSR system to the level of 5.5%/10.3% on the Switchboard/CallHome subsets of the Hub5 2000 evaluation, which - at least at the writing of this paper - is a new performance milestone (albeit not at what we measure to be human performance!). On the acoustic side, we use a score fusion of three models: one LSTM with multiple feature inputs, a second LSTM trained with speaker-adversarial multi-task learning and a third residual net (ResNet) with 25 convolutional layers and time-dilated convolutions. On the language modeling side, we use word and character LSTMs and convolutional WaveNet-style language models. George Saon, Gakuto Kurata, Tom Sercu, Kartik Audhkhasi, Samuel Thomas 0001, Dimitrios Dimitriadis, Bhuvana Ramabhadran, Michael Picheny, Lynn-Li Lim, Bergul Roomi, Phil Hall |
INTERSPEECH | 1 |
| 2016 | On the importance of event detection for ASRabstractThe performance of modern large vocabulary continuous speech recognition (LVCSR) systems is heavily affected by segment boundaries, proper speaker identification of the segments, as well as removal of spurious data. We propose to use Long Short Term Memory (LSTM) recurrent neural networks to partition audio into speech segments as well as track speaker turns. Additionally, we train an LSTM to also identify music segments. We show that the accurate detection of events, along with removal of silence and music, using our LSTM yields a 9-10% relative improvement in ASR performance. Secondary processing by speaker clustering provides an additional boost in accuracy. Event detection accuracy of the LSTM approach is also described. David Haws, Dimitrios Dimitriadis, George Saon, Samuel Thomas 0001, Michael Picheny |
ICASSP | 3 |
| 2016 | The IBM 2016 English Conversational Telephone Speech Recognition SystemabstractWe describe the latest improvements to the IBM English conversational telephone speech recognition system.Some of the techniques that were found beneficial are: maxout networks with annealed dropout rates; networks with a very large number of outputs trained on 2000 hours of data; joint modeling of partially unfolded recurrent neural networks and convolutional nets by combining the bottleneck and output layers and retraining the resulting model; and lastly, sophisticated language model rescoring with exponential and neural network LMs.These techniques result in an 8.0% word error rate on the Switchboard part of the Hub5-2000 evaluation test set which is 23% relative better than our previous best published result. George Saon, Tom Sercu, Steven J. Rennie, Hong-Kwang Jeff Kuo |
INTERSPEECH | 1 |
| 2016 | Domain Adaptation of CNN Based Acoustic Models Under Limited Resource Settings
Masayuki Suzuki, Ryuki Tachibana, Samuel Thomas 0001, Bhuvana Ramabhadran, George Saon |
INTERSPEECH | 5 |
| 2015 | A nonmonotone learning rate strategy for SGD training of deep neural networksabstractThe algorithm of choice for cross-entropy training of deep neural network (DNN) acoustic models is mini-batch stochastic gradient descent (SGD). One of the important decisions for this algorithm is the learning rate strategy (also called stepsize selection). We investigate several existing schemes and propose a new learning rate strategy which is inspired by nonmonotone linesearch techniques in nonlinear optimization and the NewBob algorithm. This strategy was found to be relatively insensitive to poorly tuned parameters and resulted in lower word error rates compared to Newbob on two different LVCSR tasks (English broadcast news transcription 50 hours and Switchboard telephone conversations 300 hours). Further, we discuss some justifications for the method by briefly linking it to results in optimization theory. Nitish Shirish Keskar, George Saon |
ICASSP | 2 |
| 2015 | Order-free spoken term detectionabstractIn this paper, we propose Time-Marked Word (TMW) lists as a replacement for the lattices and Confusion Networks (CNs) widely used as indexing vehicles for Spoken Term Detection (STD). In a TMW list, candidates are simply tagged with posterior probabilities and time information and stored as a large list of words: the additional ordering present in a lattice or CN is discarded. TMW lists compactly summarize a large ASR search space. Representing a large search space is critical for STD metrics such as ATWV that heavily penalize misses of rare keywords. Comparisons on the OpenKWS 2014 Tamil limited language pack task [1] show that the new TMW-based indexing results in better performance while being faster and having a smaller footprint. Lidia Mangu, George Saon, Michael Picheny, Brian Kingsbury |
ICASSP | 2 |
| 2015 | Improvements to the IBM speech activity detection system for the DARPA RATS programabstractIn this paper we describe improvements to the IBM speech activity detection (SAD) system for the third phase of the DARPA RATS program. The progress during this final phase comes from jointly training convolutional and regular deep neural networks with rich time-frequency representations of speech. With these additions, the phase 3 system reduces the equal error rate (EER) significantly on both of the program's development sets (relative improvements of 20% on dev1 and 7% on dev2) compared to an earlier phase 2 system. For the final program evaluation, the newly developed system also performs well past the program target of 3% Pmissat 1% Pfawith a performance of 1.2% Pmissat 1% Pfaand 0.3% Pfaat 3% Pmiss. Samuel Thomas 0001, George Saon, Maarten Van Segbroeck, Shri Narayanan |
ICASSP | 2 |
| 2015 | A multi-region deep neural network model in speech recognition
Jia Cui, George Saon, Bhuvana Ramabhadran, Brian Kingsbury |
INTERSPEECH | 2 |
| 2015 | The IBM 2015 English conversational telephone speech recognition systemabstractWe describe the latest improvements to the IBM English conversational telephone speech recognition system. Some of the techniques that were found beneficial are: maxout networks with annealed dropout rates; networks with a very large number of outputs trained on 2000 hours of data; joint modeling of partially unfolded recurrent neural networks and convolutional nets by combining the bottleneck and output layers and retraining the resulting model; and lastly, sophisticated language model rescoring with exponential and neural network LMs. These techniques result in an 8.0% word error rate on the Switchboard part of the Hub5-2000 evaluation test set which is 23% relative better than our previous best published result. George Saon, Hong-Kwang Jeff Kuo, Steven J. Rennie, Michael Picheny |
INTERSPEECH | 1 |
| 2015 | The IBM BOLT speech transcription systemabstractWe describe the IBM automatic speech recognition (ASR) sys-tem for the DARPA Broad Operational Language Translation (BOLT) program. The system is used to transcribe conversa-tional telephone speech (CTS) prior to machine translation for Phase 3 of the program’s Activity A. The ASR system is a com-bination of novel sequence trained ensemble deep neural net-work acoustic models on speaker adapted features and convolu-tional neural network models on two kinds of spectro-temporal representations of speech, in conjunction with a variety of class, neural network and n-gram based language models. Acoustic and language models for the recognition system are built on transcribed audio released under the program and further opti-mized for the final machine translation task as well. The evalua-tion system has a word error rate of 32.7 % on a 2 hour Egyptian Arabic development set for this task. Index Terms: Automatic speech recognition, conversational telephone speech, deep neural networks, machine translation Samuel Thomas 0001, George Saon, Hong-Kwang Jeff Kuo, Lidia Mangu |
INTERSPEECH | 2 |
| 2015 | Deep Convolutional Neural Networks for Large-scale Speech Tasks
Tara N. Sainath, Brian Kingsbury, George Saon, Hagen Soltau, Abdel-rahman Mohamed, George E. Dahl, Bhuvana Ramabhadran |
Neural Networks | 3 |
| 2014 | Improvements to filterbank and delta learning within a deep neural network frameworkabstractMany features used in speech recognition tasks are hand-crafted and are not always related to the objective at hand, that is minimizing word error rate. Recently, we showed that replacing a perceptually motivated mel-filter bank with a filter bank layer that is learned jointly with the rest of a deep neural network was promising. In this paper, we extend filter learning to a speaker-adapted, state-of-the-art system. First, we incorporate delta learning into the filter learning framework. Second, we incorporate various speaker adaptation techniques, including VTLN warping and speaker identity features. On a 50-hour English Broadcast News task, we show that we can achieve a 5% relative improvement in word error rate (WER) using the filter and delta learning, compared to having a fixed set of filters and deltas. Furthermore, after speaker adaptation, we find that filter and delta learning allows for a 3% relative improvement in WER compared to a state-of-the-art CNN. Tara N. Sainath, Brian Kingsbury, Abdel-rahman Mohamed, George Saon, Bhuvana Ramabhadran |
ICASSP | 4 |
| 2014 | A comparison of two optimization techniques for sequence discriminative training of deep neural networksabstractWe compare two optimization methods for lattice-based sequence discriminative training of neural network acoustic models: distributed Hessian-free (DHF) and stochastic gradient descent (SGD). Our findings on two different LVCSR tasks suggest that SGD running on a single GPU machine achieves the best accuracy 2.5 times faster than DHF running on multiple non-GPU machines; however, DHF training achieves a higher accuracy at the end of the optimization. In addition, we present an improved modified forward-backward algorithm for computing lattice-based expected loss functions and gradients that results in a 34% speedup for SGD. George Saon, Hagen Soltau |
ICASSP | 1 |
| 2014 | Joint training of convolutional and non-convolutional neural networksabstractWe describe a simple modification of neural networks which consists in extending the commonly used linear layer structure to an arbitrary graph structure. This allows us to combine the benefits of convolutional neural networks with the benefits of regular networks. The joint model has only a small increase in parameter size and training and decoding time are virtually unaffected. We report significant improvements over very strong baselines on two LVCSR tasks and one speech activity detection task. Hagen Soltau, George Saon, Tara N. Sainath |
ICASSP | 2 |
| 2014 | Analyzing convolutional neural networks for speech activity detection in mismatched acoustic conditionsabstractConvolutional neural networks (CNN) are extensions to deep neural networks (DNN) which are used as alternate acoustic models with state-of-the-art performances for speech recognition. In this paper, CNNs are used as acoustic models for speech activity detection (SAD) on data collected over noisy radio communication channels. When these SAD models are tested on audio recorded from radio channels not seen during training, there is severe performance degradation. We attribute this degradation to mismatches between the two dimensional filters learnt in the initial CNN layers and the novel channel data. Using a small amount of supervised data from the novel channels, the filters can be adapted to provide significant improvements in SAD performance. In mismatched acoustic conditions, the adapted models provide significant improvements (about 10-25%) relative to conventional DNN-based SAD systems. These results illustrate that CNNs have a considerable advantage in fast adaptation for acoustic modeling in these settings. Samuel Thomas 0001, Sriram Ganapathy, George Saon, Hagen Soltau |
ICASSP | 3 |
| 2014 | Parallel deep neural network training for LVCSR tasks using blue gene/QabstractWhile Deep Neural Networks (DNNs) have achieved tremendous success for LVCSR tasks, training these networks is slow. To date, the most common approach to train DNNs is via stochastic gradient descent (SGD), serially on a single GPU machine. Serial training, coupled with the large number of training parameters and speech data set sizes, makes DNN training very slow for LVCSR tasks. While 2nd order, data-parallel methods have also been explored, these methods are not always faster on CPU clusters due to the large communication cost between processors. In this work, we explore using a specialized hardware/software approach, utilizing a Blue Gene/Q (BG/Q) system, which has thousands of processors and excellent interprocessor communication. We explore using the 2nd order Hessian-free (HF) algorithm for DNN training with BG/Q, for both cross-entropy and sequence training of DNNs. Results on three LVCSR tasks indicate that using HF with BG/Q offers up to an 11x speedup, as well as an improved word error rate (WER), compared to SGD on a GPU. Tara N. Sainath, I-Hsin Chung, Bhuvana Ramabhadran, Michael Picheny, John A. Gunnels, Brian Kingsbury, George Saon, Vernon Austel, Upendra V. Chaudhari |
INTERSPEECH | 7 |
| 2014 | Unfolded recurrent neural networks for speech recognitionabstractWe introduce recurrent neural networks (RNNs) for acoustic modeling which are unfolded in time for a fixed number of time steps. The proposed models are feedforward networks with the property that the unfolded layers which correspond to the recurrent layer have time-shifted inputs and tied weight matrices. Besides the temporal depth due to unfolding, hierarchical processing depth is added by means of several non-recurrent hidden layers inserted between the unfolded layers and the output layer. The training of these models: (a) has a complexity that is comparable to deep neural networks (DNNs) with the same number of layers; (b) can be done on frame-randomized minibatches; (c) can be implemented efficiently through matrix-matrix operations on GPU architectures which makes it scalable for large tasks. Experimental results on the Switchboard 300 hours English conversational telephony task show a 5% relative improvement in word error rate over state-of-the-art DNNs trained on FMLLR features with i-vector speaker adaptation and hessianfree sequence discriminative training. Index Terms: recurrent neural networks, speech recognition George Saon, Hagen Soltau, Ahmad Emami, Michael Picheny |
INTERSPEECH | 1 |
| 2014 | A distributed architecture for fast SGD sequence discriminative training of DNN acoustic modelsabstractWe describe a hybrid GPU/CPU architecture for stochastic gradient descent training of neural network acoustic models under a lattice-based minimum Bayes risk (MBR) criterion. The crux of the method is to run SGD on a GPU card which consumes frame-randomized mini-batches produced by multiple workers running on a cluster of multi-core CPU nodes which compute HMM state MBR occupancies. To minimize communication cost, a separate thread running on the GPU host receives minibatches from and sends updated models to the workers, and communicates with the SGD thread via a producer-consumer queue of minibatches. Using this architecture, it is possible to match the speed of GPU-based SGD cross-entropy (CE) training (1 hour of processing per 100 hours of audio on Switchboard). Additionally, we compare different ways of doing frame randomization and discuss experimental results on three LVCSR tasks (Switchboard 300 hours, English broadcast news 50 hours, and noisy Levantine telephone conversations 300 hours). George Saon |
SLT | 1 |
| 2013 | The IBM keyword search system for the DARPA RATS programabstractThe paper describes a state-of-the-art keyword search (KWS) system in which significant improvements are obtained by using Convolutional Neural Network acoustic models, a two-step speech segmentation approach and a simplified ASR architecture optimized for KWS. The system described in this paper had the best performance in the 2013 DARPA RATS evaluation for both Levantine and Farsi. Lidia Mangu, Hagen Soltau, Hong-Kwang Jeff Kuo, George Saon |
ASRU | 4 |
| 2013 | Improvements to Deep Convolutional Neural Networks for LVCSRabstractDeep Convolutional Neural Networks (CNNs) are more powerful than Deep Neural Networks (DNN), as they are able to better reduce spectral variation in the input signal. This has also been confirmed experimentally, with CNNs showing improvements in word error rate (WER) between 4-12% relative compared to DNNs across a variety of LVCSR tasks. In this paper, we describe different methods to further improve CNN performance. First, we conduct a deep analysis comparing limited weight sharing and full weight sharing with state-of-the-art features. Second, we apply various pooling strategies that have shown improvements in computer vision to an LVCSR speech task. Third, we introduce a method to effectively incorporate speaker adaptation, namely fMLLR, into log-mel features. Fourth, we introduce an effective strategy to use dropout during Hessian-free sequence training. We find that with these improvements, particularly with fMLLR and dropout, we are able to achieve an additional 2-3% relative improvement in WER on a 50-hour Broadcast News task over our previous best CNN baseline. On a larger 400-hour BN task, we find an additional 4-5% relative improvement over our previous best CNN baseline. Tara N. Sainath, Brian Kingsbury, Abdel-rahman Mohamed, George E. Dahl, George Saon, Hagen Soltau, Tomás Beran, Aleksandr Y. Aravkin, Bhuvana Ramabhadran |
ASRU | 5 |
| 2013 | Speaker adaptation of neural network acoustic models using i-vectorsabstractWe propose to adapt deep neural network (DNN) acoustic models to a target speaker by supplying speaker identity vectors (i-vectors) as input features to the network in parallel with the regular acoustic features for ASR. For both training and test, the i-vector for a given speaker is concatenated to every frame belonging to that speaker and changes across different speakers. Experimental results on a Switchboard 300 hours corpus show that DNNs trained on speaker independent features and i-vectors achieve a 10% relative improvement in word error rate (WER) over networks trained on speaker independent features only. These networks are comparable in performance to DNNs trained on speaker-adapted features (with VTLN and FMLLR) with the advantage that only one decoding pass is needed. Furthermore, networks trained on speaker-adapted features and i-vectors achieve a 5-6% relative improvement in WER after hessian-free sequence training over networks trained on speaker-adapted features only. George Saon, Hagen Soltau, David Nahamoo, Michael Picheny |
ASRU | 1 |
| 2013 | Exploiting diversity for spoken term detectionabstractThe paper describes a state-of-the-art spoken term detection system in which significant improvements are obtained by diversifying the ASR engines used for indexing and combining the search results. First, we describe the design factors that, when varied, produce complementary STD systems and show that the performance of the combined system is 3 times better than the best individual component. Next, we describe different strategies for system combination and show that significant improvements can be achieved by normalizing the combined scores. We propose a classifier-based system combination strategy which outperforms a highly optimized baseline. The system described in this paper had the highest accuracy in the 2012 DARPA RATS evaluation. Lidia Mangu, Hagen Soltau, Hong-Kwang Jeff Kuo, Brian Kingsbury, George Saon |
ICASSP | 5 |
| 2013 | The IBM speech activity detection system for the DARPA RATS program
George Saon, Samuel Thomas 0001, Hagen Soltau, Sriram Ganapathy, Brian Kingsbury |
INTERSPEECH | 1 |
| 2013 | Neural network acoustic models for the DARPA RATS program
Hagen Soltau, Hong-Kwang Jeff Kuo, Lidia Mangu, George Saon, Tomás Beran |
INTERSPEECH | 4 |
| 2012 | Sparse Bayesian Factor Analysis for Stereo-based Stochastic Mapping
Mohamed Afify, George Saon, Vaibhava Goel |
INTERSPEECH | 3 |
| 2012 | Discriminative feature-space transforms using deep neural networksabstractWe present a deep neural network (DNN) architecture which learns time-dependent offsets to acoustic feature vectors according to a discriminative objective function such as maximum mutual information (MMI) between the reference words and the transformed acoustic observation sequence. A key ingredient in this technique is a greedy layer-wise pretraining of the network based on minimum squared error between the DNN outputs and the offsets provided by a linear feature-space MMI (FMMI) transform. Next, the weights of the pretrained network are updated with stochastic gradient ascent by backpropagating the MMI gradient through the DNN layers. Experiments on a 50 hour English broadcast news transcription task show a 4% relative improvement using a 6-layer DNN transform over a state-of-the-art speaker-adapted system with FMMI and modelspace discriminative training. George Saon, Brian Kingsbury |
INTERSPEECH | 1 |
| 2012 | Boosting systems for large vocabulary continuous speech recognition
George Saon, Hagen Soltau |
Speech Commun. | 1 |
| 2012 | Bayesian Sensing Hidden Markov ModelsabstractIn this paper, we introduce Bayesian sensing hidden Markov models (BS-HMMs) to represent sequential data based on a set of state-dependent basis vectors. The goal of this work is to perform Bayesian sensing and model regularization for heterogeneous training data. By incorporating a prior density on sensing weights, the relevance of different bases to a feature vector is determined by the corresponding precision parameters. The BS-HMM parameters, consisting of the basis vectors, the precision matrices of sensing weights and the precision matrices of reconstruction errors, are jointly estimated by maximizing the likelihood function, which is marginalized over the weight priors. We derive recursive solutions for the three parameters, which are expressed via maximum a posteriori estimates of the sensing weights. We specifically optimize BS-HMMs for large-vocabulary continuous speech recognition (LVCSR) by introducing a mixture model of BS-HMMs and by adapting the basis vectors to different speakers. Discriminative training of BS-HMMs in the model domain and the feature domain is also proposed. Experimental results on an LVCSR task show consistent improvements due to the three sets of BS-HMM parameters and demonstrate how the extensions of mixture models, speaker adaptation, and discriminative training achieve better recognition results compared to those of conventional HMMs based on Gaussian mixture models. George Saon, Jen-Tzung Chien |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | Minimum Bayes risk discriminative language models for Arabic speech recognitionabstractIn this paper we explore discriminative language modeling (DLM) on highly optimized state-of-the-art large vocabulary Arabic broadcast speech recognition systems used for the Phase 5 DARPA GALE Evaluation. In particular, we study in detail a minimum Bayes risk (MBR) criterion for DLM. MBR training outperforms perceptron training. Interestingly, we found that our DLMs generalized to mismatched conditions, such as using a different acoustic model during testing. We also examine the interesting problem of unsupervised DLM training using a Bayes risk metric as a surrogate for word error rate (WER). In some experiments, we were able to obtain about half of the gain of the supervised DLM. Hong-Kwang Jeff Kuo, Ebru Arisoy, Lidia Mangu, George Saon |
ASRU | 4 |
| 2011 | The IBM 2011 GALE Arabic speech transcription systemabstractWe describe the Arabic broadcast transcription system fielded by IBM in the GALE Phase 5 machine translation evaluation. Key advances over our Phase 4 system include a new Bayesian Sensing HMM acoustic model; multistream neural network features; a MADA vowelized acoustic model; and the use of a variety of language model techniques with significant additive gains. These advances were instrumental in achieving a word error rate of 7.4% on the Phase 5 evaluation set, and an absolute improvement of 0.9% word error rate over our 2009 system on the unsequestered Phase 4 evaluation data. Lidia Mangu, Hong-Kwang Jeff Kuo, Stephen M. Chu, Brian Kingsbury, George Saon, Hagen Soltau, Fadi Biadsy |
ASRU | 5 |
| 2011 | Some properties of Bayesian sensing hidden Markov modelsabstractIn Bayesian sensing hidden Markov models (BSHMMs) the acoustic feature vectors are represented by a set of state-dependent basis vectors and by time-dependent sensing weights. The Bayesian formulation comes from assuming state-dependent zero mean Gaussian priors for the weights and from using marginal likelihood functions obtained by integrating out the weights. Here, we discuss two properties of BSHMMs. The first property is that the marginal likelihood is Gaussian with a factor analyzed covariance matrix with the basis providing a low-rank correction to the diagonal covariance of the reconstruction errors. The second property, termed automatic relevance determination, provides a method for discarding basis vectors that are not relevant for encoding feature vectors. This allows model complexity control where one can initially train a large model and then prune it to a smaller size by removing the basis vectors which correspond to the largest precision values of the sensing weights. The last property turned out to be useful in successfully deploying models trained on 1800 hours of data during the 2011 DARPA GALE Arabic broadcast news transcription evaluation. George Saon, Jen-Tzung Chien |
ASRU | 1 |
| 2011 | The IBM 2009 GALE Arabic speech transcription systemabstractWe describe the Arabic broadcast transcription system fielded by IBM in the GALE Phase 4 machine translation evaluation. Key advances over our Phase 3.5 system include improvements to context-dependent modeling in vowelized Arabic acoustic models; the use of neural-network features provided by the International Computer Science Institute; Model M language models; a neural network language model that uses syntactic and morphological features; and improvements to our system combination strategy. These advances were instrumental in achieving a word error rate of 8.9% on the Phase 4 evaluation set, and an absolute improvement of 1.6% word error rate over our 2008 system on the unsequestered Phase 3.5 evaluation data. Brian Kingsbury, Hagen Soltau, George Saon, Stephen M. Chu, Hong-Kwang Jeff Kuo, Lidia Mangu, Suman V. Ravuri, Nelson Morgan, Adam Janin |
ICASSP | 3 |
| 2011 | Bayesian sensing hidden Markov models for speech recognitionabstractWe introduce Bayesian sensing hidden Markov models (BS-HMMs) to represent speech data based on a set of state-dependent basis vectors. By incorporating the prior density of sensing weights, the relevance of a feature vector to different bases is determined by the corresponding precision parameters. The BS-HMM parameters, consisting of the basis vectors, the precision matrices of sensing weights and the precision matrices of reconstruction errors, are jointly estimated by maximizing the likelihood function, which is marginalized over the weight priors. We derive recursive solutions for the three parameters, which are expressed via maximum a posteriori estimates of the sensing weights. Experimental results on an LVCSR task show consistent gains over conventional HMMs with Gaussian mixture models for both ML and discriminative training scenarios. George Saon, Jen-Tzung Chien |
ICASSP | 1 |
| 2011 | Discriminative training for Bayesian sensing hidden Markov modelsabstractWe describe feature space and model space discriminative training for a new class of acoustic models called Bayesian sensing hidden Markov models (BS-HMMs). In BS-HMMs, speech data is represented by a set of state-dependent basis vectors. The relevance of a feature vector to different bases is determined by the precision matrices of the sensing weights. The basis vectors and the precision matrices of the reconstruction errors are jointly estimated by optimizing a maximum mutual information (MMI) criterion. Additionally, we discuss the training of an fMPE-style discriminative feature transformation under the same criterion given these models. Experimental results on an LVCSR task show that the proposed models outperform discriminatively trained conventional HMMs with Gaussian mixture models (GMMs). Cross-adapting the baseline GMM-HMMs to the BS-HMM output yields a 6% relative gain which indicates that the two systems make different errors. George Saon, Jen-Tzung Chien |
ICASSP | 1 |
| 2010 | The IBM 2008 GALE Arabic speech transcription systemabstractThis paper describes the Arabic broadcast transcription system fielded by IBM in the GALE Phase 3.5 machine translation evaluation. Key advances compared to our Phase 2.5 system include improved discriminative training, the use of Subspace Gaussian Mixture Models (SGMM), neural network acoustic features, variable frame rate decoding, training data partitioning experiments, unpruned n-gram language models and neural network language models. These advances were instrumental in achieving a word error rate of 8.9% on the evaluation test set. George Saon, Hagen Soltau, Upendra V. Chaudhari, Stephen M. Chu, Brian Kingsbury, Hong-Kwang Jeff Kuo, Lidia Mangu, Daniel Povey |
ICASSP | 1 |
| 2010 | Boosting systems for LVCSR
George Saon, Hagen Soltau |
INTERSPEECH | 1 |
| 2010 | The IBM Attila speech recognition toolkitabstractWe describe the design of IBM's Attila speech recognition toolkit. We show how the combination of a highly modular and efficient library of low-level C++ classes with simple interfaces, an interconnection layer implemented in a modern scripting language (Python), and a standardized collection of scripts for system-building produce a flexible and scalable toolkit that is useful both for basic research and for construction of large transcription systems for competitive evaluations. Hagen Soltau, George Saon, Brian Kingsbury |
SLT | 2 |
| 2009 | Dynamic network decoding revisitedabstractWe present a dynamic network decoder capable of using large cross-word context models and large n-gram histories. Our method for constructing the search network is designed to process large cross-word context models very efficiently and we address the optimization of the search network to minimize any overhead during run-time for the dynamic network decoder. The search procedure uses the full LM history for lookahead, and path recombination is done as early as possible. In our systematic comparison to a static FSM based decoder, we find the dynamic decoder can run at comparable speed as the static decoder when large language models are used, while the static decoder performs best for small language models. We discuss the use of very large vocabularies of up to 2.5 million words for both decoding approaches and analyze the effect of weak acoustic models for pruning. Hagen Soltau, George Saon |
ASRU | 2 |
| 2009 | Large margin semi-tied covariance transforms for discriminative trainingabstractWe discuss the applicability of large margin techniques to the problem of estimating linear transforms for discriminative training of a semi-tied covariance (STC) model. Since STC models are good proxies for full-covariance (FC) Gaussian models, the idea is to combine the benefit of the latest discriminative training techniques and the modeling advantage of FC Gaussians at a much lower computational cost. We study the interaction of these transforms with feature-space and model-space discriminative training on state-of-the-art speaker adapted systems built for a large-scale Arabic broadcast news transcription task. George Saon, Daniel Povey, Hagen Soltau |
ICASSP | 1 |
| 2009 | Advances in Arabic Speech Transcription at IBM Under the DARPA GALE ProgramabstractThis paper describes the Arabic broadcast transcription system fielded by IBM in the GALE Phase 2.5 machine translation evaluation. Key advances include the use of additional training data from the Linguistic Data Consortium (LDC), use of a very large vocabulary comprising 737 K words and 2.5 M pronunciation variants, automatic vowelization using flat-start training, cross-adaptation between unvowelized and vowelized acoustic models, and rescoring with a neural-network language model. The resulting system achieves word error rates below 10% on Arabic broadcasts. Very large scale experiments with unsupervised training demonstrate that the utility of unsupervised data depends on the amount of supervised data available. While unsupervised training improves system performance when a limited amount (135 h) of supervised data is available, these gains disappear when a greater amount (848 h) of supervised data is used, even with a very large (7069 h) corpus of unsupervised data. We also describe a method for modeling Arabic dialects that avoids the problem of data sparseness entailed by dialect-specific acoustic models via the use of non-phonetic, dialect questions in the decision trees. We show how this method can be used with a statically compiled decoding graph by partitioning the decision trees into a static component and a dynamic component, with the dynamic component being replaced by a mapping that is evaluated at run-time. Hagen Soltau, George Saon, Brian Kingsbury, Hong-Kwang Jeff Kuo, Lidia Mangu, Daniel Povey, Ahmad Emami |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | Boosted MMI for model and feature-space discriminative trainingabstractWe present a modified form of the maximum mutual information (MMI) objective function which gives improved results for discriminative training. The modification consists of boosting the likelihoods of paths in the denominator lattice that have a higher phone error relative to the correct transcript, by using the same phone accuracy function that is used in Minimum Phone Error (MPE) training. We combine this with another improvement to our implementation of the Extended Baum-Welch update equations for MMI, namely the canceling of any shared part of the numerator and denominator statistics on each frame (a procedure that is already done in MPE). This change affects the Gaussian-specific learning rate. We also investigate another modification whereby we replace I-smoothing to the ML estimate with I-smoothing to the previous iteration's value. Boosted MMI gives better results than MPE in both model and feature-space discriminative training, although not consistently. Daniel Povey, Dimitri Kanevsky, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, Karthik Visweswariah |
ICASSP | 5 |
| 2008 | Penalty function maximization for large margin HMM trainingabstractWe perform large margin training of HMM acoustic parameters by maximizing a penalty function which combines two terms. The first term is a scale which gets multiplied with the Hamming distance between HMM state sequences to form a multi-label (or sequence) margin. The second term arises from constraints on the training data that the joint log-likelihoods of acoustic and correct word sequences exceed the joint log-likelihoods of acoustic and incorrect word sequences by at least the multi-label margin between the corresponding Viterbi state sequences. Using the softmax trick, we collapse these constraints into a boosted MMI-like term. The resulting objective function can be efficiently maximized using extended Baum-Welch updates. Experimental results on multiple LVCSR tasks show a good correlation between the objective function and the word error rate. George Saon, Daniel Povey |
INTERSPEECH | 1 |
| 2007 | Lattice-based Viterbi decoding techniques for speech translationabstractWe describe a cardinal-synchronous Viterbi decoder for statistical phrase-based machine translation which can operate on general ASR lattices (as opposed to confusion networks). The decoder implements constrained source reordering on the input lattice and makes use of an outbound distortion model to score the possible reorderings. The phrase table, representing the decoding search space, is encoded as a weighted finite state acceptor which is determined and minimized. At a high level, the search proceeds by performing simultaneous transitions in two pairs of automata: (input lattice, phrase table FSM) and (phrase table FSM, target language model). An alternative decoding strategy that we explore is to break the search into two independent subproblems: first, we perform monotone lattice decoding and find the best foreign path through the ASR lattice and then, we decode this path with reordering using standard sentence-based SMT. We report experimental results on several testsets of a large scale Arabic-to-English speech translation task in the context of the global autonomous language exploitation (or GALE) DARPA project. The results indicate that, for monotone search, lattice-based decoding outperforms 1-best decoding whereas for search with reordering, only the second decoding strategy was found to be superior to 1-best decoding. In both cases, the improvements hold only for shallow lattices. George Saon, Michael Picheny |
ASRU | 1 |
| 2007 | The IBM 2006 Gale Arabic ASR SystemabstractThis paper describes the advances made in IBM's Arabic broadcast news transcription system which was fielded in the 2006 GALE ASR and machine translation evaluation. These advances were instrumental in lowering the word error rate by 42% relative over the course of one year and include: training on additional LDC data, large-scale discriminative training on 1800 hours of unsupervised data, automatic vowelization using a flat-start approach, use of a large vocabulary with 617K words and 2 million pronunciations and lastly, a system architecture based on cross-adaptation between unvowelized and vowelized acoustic models. Hagen Soltau, George Saon, Brian Kingsbury, Hong-Kwang Jeff Kuo, Lidia Mangu, Daniel Povey, Geoffrey Zweig |
ICASSP (4) | 2 |
| 2006 | A Non-Linear Speaker Adaptation Technique using Kernel Ridge RegressionabstractWe propose a non-linear model space transformation for speaker or environment adaptation based on weighted kernel ridge regression (KRR). The transformation is given by a generalized least squares linear regression in a kernel-induced feature space operating on Gaussian mixture model means and having as targets the adaptation frames. Using the "kernel trick", the solution to the optimization problem is obtained by solving a system of linear equations involving the Gram matrix of the input variables. We show that MLLR is a special case of KRR when a linear kernel is employed. Furthermore, we study an efficient low-rank approximation to the kernel matrix termed "rectangle method", where the regressors are chosen to be a small set of clustered adaptation frames. Experiments conducted on the EARS database (English conversational telephone speech) indicate that KRR with a Gaussian RBF kernel outperforms standard regression class-based MLLR George Saon |
ICASSP (1) | 1 |
| 2006 | Automated Quality Monitoring in the Call Center with ASR and Maximum EntropyabstractThis paper describes an automated system for assigning quality scores to recorded call center conversations. The system combines speech recognition, pattern matching, and maximum entropy classification to rank calls according to their measured quality. Calls at both end of the spectrum are flagged as "interesting" and made available for further human monitoring. In this process, pattern matching on the ASR transcript is used to answer a set of standard quality control questions such as "did the agent use courteous words and phrases," and to generate a question-based score. This is interpolated with the probability of a call being "bad," as determined by maximum entropy operating on a set of ASR-derived features such as "maximum silence length" and the occurrence of selected n-gram word sequences. The system is trained on a set of calls with associated manual evaluation forms. We present precision and recall results from IBM's North American Help Desk indicating that for a given amount of listening effort, this system triples the number of bad calls that are identified, over the current policy of randomly sampling calls Geoffrey Zweig, Olivier Siohan, George Saon, Bhuvana Ramabhadran, Daniel Povey, Lidia Mangu, Brian Kingsbury |
ICASSP (1) | 3 |
| 2006 | Feature and model space speaker adaptation with full covariance GaussiansabstractFull covariance models can give better results for speech recognition than diagonal models, yet they introduce complications for standard speaker adaptation techniques such as MLLR and fMLLR. Here we introduce efficient update methods to train adaptation matrices for the full covariance case. We also experiment with a simplified technique in which we pretend that the full covariance Gaussians are diagonal and obtain adaptation matrices under that assumption. We show that this approximate method works almost as well as the exact method. Daniel Povey, George Saon |
INTERSPEECH | 2 |
| 2006 | Automated Quality Monitoring for Call Centers using Speech and NLP Technologies
Geoffrey Zweig, Olivier Siohan, George Saon, Bhuvana Ramabhadran, Daniel Povey, Lidia Mangu, Brian Kingsbury |
HLT-NAACL | 3 |
| 2006 | On the Effect Ofword Error Rate on Automated Quality MonitoringabstractThis paper studies the effect of word-error-rate (WER) on an automated quality monitoring application for call centers. The system consists of a speech recognition module and a call ranking module. The call ranking module combines direct question answering with a maximum-entropy classifier to automatically monitor the calls that enter a call center, and label them as "good" or "bad". We find that, in the monitoring regime where only a small fraction of the calls are monitored, we achieve 80% precision and 50% recall in classifying whether a call belongs to the bottom 20%. Additionally, the correlation between human and computer-generated scores turns out to be highly sensitive to word error rate. George Saon, Bhuvana Ramabhadran, Geoffrey Zweig |
SLT | 1 |
| 2006 | Advances in speech transcription at IBM under the DARPA EARS programabstractThis paper describes the technical and system building advances made in IBM's speech recognition technology over the course of the Defense Advanced Research Projects Agency (DARPA) Effective Affordable Reusable Speech-to-Text (EARS) program. At a technical level, these advances include the development of a new form of feature-based minimum phone error training (fMPE), the use of large-scale discriminatively trained full-covariance Gaussian models, the use of septaphone acoustic context in static decoding graphs, and improvements in basic decoding algorithms. At a system building level, the advances include a system architecture based on cross-adaptation and the incorporation of 2100 h of training data in every system component. We present results on English conversational telephony test data from the 2003 and 2004 NIST evaluations. The combination of technical advances and an order of magnitude more training data in 2004 reduced the error rate on the 2003 test set by approximately 21% relative-from 20.4% to 16.1%-over the most accurate system in the 2003 evaluation and produced the most accurate results on the 2004 test sets in every speed category. Stanley F. Chen, Brian Kingsbury, Lidia Mangu, Daniel Povey, George Saon, Hagen Soltau, Geoffrey Zweig |
IEEE Trans. Speech Audio Process. | 5 |
| 2005 | fMPE: Discriminatively Trained Features for Speech RecognitionabstractMPE (minimum phone error) is a previously introduced technique for discriminative training of HMM parameters. fMPE applies the same objective function to the features, transforming the data with a kernel-like method and training millions of parameters, comparable to the size of the acoustic model. Despite the large number of parameters, fMPE is robust to over-training. The method is to train a matrix projecting from posteriors of Gaussians to a normal size feature space, and then to add the projected features to normal features such as PLP. The matrix is trained from a zero start using a linear method. Sparsity of posteriors ensures speed in both training and test time. The technique gives similar improvements to MPE (around 10% relative). MPE on top of fMPE results in error rates up to 6.5% relative better than MPE alone, or more if multiple layers of transform are trained. Daniel Povey, Brian Kingsbury, Lidia Mangu, George Saon, Hagen Soltau, Geoffrey Zweig |
ICASSP (1) | 4 |
| 2005 | The IBM 2004 Conversational Telephony System for Rich TranscriptionabstractThis paper describes the technical advances in IBM's conversational telephony submission to the DARPA-sponsored 2004 rich transcription evaluation (RT-04). These advances include a system architecture based on cross-adaptation; a new form of feature-based MPE training; the use of a full-scale discriminatively trained full covariance Gaussian system; the use of septaphone cross-word acoustic context in static decoding graphs; and the incorporation of 2100 hours of training data in every system component. These advances reduced the error rate by approximately 21% relative, on the 2003 test set, over the best-performing system in last year's evaluation, and produced the best results on the RT-04 current and progress CTS data. Hagen Soltau, Brian Kingsbury, Lidia Mangu, Daniel Povey, George Saon, Geoffrey Zweig |
ICASSP (1) | 5 |
| 2005 | Anatomy of an extremely fast LVCSR decoderabstractWe report in detail the decoding strategy that we used for the past two Darpa Rich Transcription evaluations (RT’03 and RT’04) which is based on finite state automata (FSA). We discuss the format of the static decoding graphs, the particulars of our Viterbi implementation, the lattice generation and the likelihood evaluation. This paper is intended to familiarize the reader with some of the design issues encountered when building an FSA decoder. Experimental results are given on the EARS database (English conversational telephone speech) with emphasis on our faster than real-time system. 1. George Saon, Daniel Povey, Geoffrey Zweig |
INTERSPEECH | 1 |
| 2004 | Feature space GaussianizationabstractWe propose a non-linear feature space transformation for speaker/environment adaptation which forces the individual dimensions of the acoustic data for every speaker to be Gaussian distributed. The transformation is given by the preimage under the Gaussian cumulative distribution function (CDF) of the empirical CDF on a per dimension basis. We show that, for a given dimension, this transformation achieves minimum divergence between the density function of the transformed adaptation data and the normal density with zero mean and unit variance. Experimental results on both small and large vocabulary tasks show consistent improvements over the application of linear adaptation transforms only. George Saon, Satya Dharanipragada, Daniel Povey |
ICASSP (1) | 1 |
| 2004 | Fractional Fourier transform features for speech recognitionabstractIn this paper a novel speech signal representation method is presented. The proposed method is based on the fractional Fourier transform (FrFT), which is a generalization of the classical Fourier transform (FT). Even though we use FrFT in feature extraction for speech recognition, it can very well be used in other areas such as enhancement, verification, and synthesis, where parametric representation of speech is needed. Experimental results conducted on the Aurora 2 database show significant improvements over MFCC at high SNR conditions. Ruhi Sarikaya, George Saon |
ICASSP (1) | 3 |
| 2004 | Arc minimization in finite-state decoding graphs with cross-word acoustic context
François Yvon, Geoffrey Zweig, George Saon |
Comput. Speech Lang. | 3 |
| 2003 | Toward domain-independent conversational speech recognitionabstractWe describe a multi-domain, conversational test set developed for IBM’s Superhuman speech recognition project and our 2002 benchmark system for this task. Through the use of multipass decoding, unsupervised adaptation and combination of hypotheses from systems using diverse feature sets and acoustic models, we achieve a word error rate of 32.0 % on data drawn from voicemail messages, two-person conversations and multiple-person meetings. 1. Brian Kingsbury, Lidia Mangu, George Saon, Geoffrey Zweig, Scott Axelrod, Vaibhava Goel, Karthik Visweswariah, Michael Picheny |
INTERSPEECH | 3 |
| 2003 | An architecture for rapid decoding of large vocabulary conversational speechabstractThis paper addresses the question of how to design a large vocabulary recognition system so that it can simultaneously handle a sophisticated language model, perform state-ofthe-art speaker adaptation, and run in one times real time 1 (1 RT). The architecture we propose is based on classical HMM Viterbi decoding, but uses an extremely fast initial speaker-independent decoding to estimate VTL warp factors, feature-space and model-space MLLR transformations that are used in a final speaker-adapted decoding. We present results on past Switchboard evaluation data that indicate that this strategy compares favorably to published unlimited-time systems (running in several hundred times real-time). Coincidentally, this is the system that IBM fielded in the 2003 EARS Rich Transcription evaluation. 1. George Saon, Geoffrey Zweig, Brian Kingsbury, Lidia Mangu, Upendra V. Chaudhari |
INTERSPEECH | 1 |
| 2002 | Digit recognition in noisy environments via a sequential GMM/SVM systemabstractThis paper exploits the fact that when GMM and SVM classifiers with roughly the same level of performance exhibit uncorrelated errors they can be combined to produce a better classifier. The gain accrues from combining the descriptive strength of GMM models with the discriminative power of SVM classifiers. This idea, first exploited in the context of speaker recognition [1, 2], is applied to speech recognition - specifically to a digit recognition task in a noisy environment - with significant gains in performance. Shai Fine, George Saon, Ramesh A. Gopinath |
ICASSP | 2 |
| 2002 | Robust speech recognition in Noisy Environments: The 2001 IBM spine evaluation systemabstractWe report on the system IBM fielded in the second SPeech In Noisy Environments (SPINE-2) evaluation, conducted by the Naval Research Laboratory in October 2001. The key components of the system include an HMM-based automatic segmentation module using a novel set of LDA-transformed voicing and energy features, a multiple-pass decoding strategy that uses several speaker-and environment-normalization operations to deal with the highly variable acoustics of the evaluation, the combination of hypotheses from decoders operating on three distinct acoustic feature sets, and a class-based language model that uses both the SPINE-1 and SPINE-2 training data to estimate reliable probabilities for the new SPINE-2 vocabulary. Brian Kingsbury, George Saon, Lidia Mangu, Mukund Padmanabhan, Ruhi Sarikaya |
ICASSP | 2 |
| 2002 | Improvements to the IBM Aurora 2 multi-condition system
George Saon, Juan M. Huerta |
INTERSPEECH | 1 |
| 2002 | Arc minimization in finite state decoding graphs with cross-word acoustic contextabstractRecent approaches to large vocabulary decoding with finite state graphs have focused on the use of state minimization algorithms to produce relatively compact graphs. This paper extends the finite state approach by developing complementary arc-minimization techniques. The use of these techniques in concert with state minimization allows us to statically compile decoding graphs in which the acoustic models utilize a full word of cross-word context. This is in significant contrast to typical systems which use only a single phone. We show that the particular arc-minimization problem that arises is in fact an NP-complete combinatorial optimization problem, and describe the reduction from 3-SAT. We present experimental results that illustrate the moderate sizes and runtimes of graphs for the Switchboard task. 1. Geoffrey Zweig, George Saon, François Yvon |
INTERSPEECH | 2 |
| 2002 | Automatic speech recognition performance on a voicemail transcription taskabstractWe report on the performance of automatic speech recognition (ASR) systems on voicemail transcription. Voicemail is spontaneous telephone speech recorded over a variety of channels; consequently, it is representative of many challenging problems in speech recognition. In the course of working on this task, several algorithms were developed that focus on different components of an ASR system, including lexicon design, feature extraction, hypothesis search, and adaptation. We report the improvements provided by these techniques, as well as other standard techniques, on a voicemail test set. Although the techniques are benchmarked on voicemail test data, their scope is not restricted to this domain as they address fundamental aspects of the speech recognition process. Mukund Padmanabhan, George Saon, Jing Huang 0019, Brian Kingsbury, Lidia Mangu |
IEEE Trans. Speech Audio Process. | 2 |
| 2001 | Speech recognition for DARPA CommunicatorabstractWe report the results of investigations in acoustic modeling, language modeling and decoding techniques, for the DARPA Communicator, a speaker-independent, telephone-based dialog system. By a combination of methods, including enlarging the acoustic model, augmenting the recognizer vocabulary, conditioning the language model upon the dialog state, and applying a post-processing decoding method, we lowered the overall word error rate from 21.9% to 15.0%, a gain of 6.9% absolute and 31.5% relative. Andrew Aaron, Scott Saobing Chen, Paul S. Cohen, Satya Dharanipragada, Ellen Eide, Martin Franz, Jean-Michel LeRoux, X. Luo, Benoît Maison, Lidia Mangu, T. Mathes, Miroslav Novak, Peder A. Olsen, Michael Picheny, Harry Printz, Bhuvana Ramabhadran, Andrej Sakrajda, George Saon, Borivoj Tydlitát, Karthik Visweswariah, D. Yuk |
ICASSP | 18 |
| 2001 | Linear feature space projections for speaker adaptationabstractWe extend the well-known technique of constrained maximum likelihood linear regression (MLLR) to compute a projection (instead of a full rank transformation) on the feature vectors of the adaptation data. We model the projected features with phone-dependent Gaussian distributions and also model the complement of the projected space with a single class-independent, speaker-specific Gaussian distribution. Subsequently, we compute the projection and its complement using maximum likelihood techniques. The resulting ML transformation is shown to be equivalent to performing a speaker-dependent heteroscedastic discriminant (or HDA) projection. Our method is in contrast to traditional approaches which use a single speaker-independent projection, and execute speaker adaptation in the resulting subspace. Experimental results on Switchboard show a 3% relative improvement in the word error rate over constrained MLLR in the projected subspace only. George Saon, Geoffrey Zweig, Mukund Padmanabhan |
ICASSP | 1 |
| 2001 | Robust digit recognition in noisy environments: the IBM Aurora 2 systemabstractABSTRACTIn this paper we describe some experiments on the Aurora 2 noisydigits database. The algorithms that we used can be broadly clas-sified into noise robustness techniques based on a linear-channelmodel of the acoustic environment such as CDCN [1] and its novelvariant termed Alignment-based CDCN( ACDCN , proposed here),and techniques which do not assume any particular knowledgeabout thestructure of the environment or noise conditions affectingthe speech signal such as discriminant feature space transforma-tions and speaker/channel adaptation. We present recognition ex-periments for both the clean training data and the multi-conditiontraining data scenarios.1. INTRODUCTIONIn this paper we describe the system and techniques for the Aurora2 noisy digits database and the results obtained. We developed twosets of acoustic models: the first set of models was trained on cleandata only and the second set was trained on the multi-conditiontraining data. For the system trained on clean data only, we appliedCDCN [1] and a novel variant of this technique called Alignment-based CDCN ( George Saon, Juan M. Huerta, Ea-Ee Jan |
INTERSPEECH | 1 |
| 2001 | Data-driven approach to designing compound words for continuous speech recognitionabstractWe present a new approach to deriving compound words from a training corpus. The motivation for making compound words is because under some assumptions, speech recognition errors occur less frequently in longer words. Furthermore, they also enable more accurate modeling of pronunciation variability at the boundary between adjacent words in a continuously spoken utterance. We introduce a measure based on the product between the direct and the reverse bigram probability of a pair of words for finding candidate pairs in order to create compound words. Our experimental results show that by augmenting both the acoustic vocabulary and the language model with these new tokens, the word recognition accuracy can be improved by absolute 2.8% (7% relative) on a voice mail continuous speech recognition task. We also compare the proposed measure for selecting compound words with other measures that have been described in the literature. George Saon, Mukund Padmanabhan |
IEEE Trans. Speech Audio Process. | 1 |
| 2000 | Maximum likelihood discriminant feature spacesabstractLinear discriminant analysis (LDA) is known to be inappropriate for the case of classes with unequal sample covariances. There has been an interest in generalizing LDA to heteroscedastic discriminant analysis (HDA) by removing the equal within-class covariance constraint. This paper presents a new approach to HDA by defining an objective function which maximizes the class discrimination in the projected subspace while ignoring the rejected dimensions. Moreover, we investigate the link between discrimination and the likelihood of the projected samples and show that HDA can be viewed as a constrained ML projection for a full covariance Gaussian model, the constraint being given by the maximization of the projected between-class scatter volume. It is shown that, under diagonal covariance Gaussian modeling constraints, applying a diagonalizing linear transformation (MLLT) to the HDA space results in increased classification accuracy even though HDA alone actually degrades the recognition performance. Experiments performed on the Switchboard and Voicemail databases show a 10%-13% relative improvement in the word error rate over standard cepstral processing. George Saon, Mukund Padmanabhan, Ramesh A. Gopinath, Scott Saobing Chen |
ICASSP | 1 |
| 2000 | Recent improvements in speech recognition performance on large vocabulary conversational speech (voicemail and switchboard)abstractIn this paper we report recent improvements in word error performance on a voicemail transcription task. Last year, the speaker independent word error rate (WER) on the dev test set of the Voicemail Transcription task was reported at 35.45% [1]. This year, we report a relative 20% gain over this number. The improvements were obtained using several new algorithms and an increased amount of training data. In addition to benchmarking the performance of these algorithms on the Voicemail task, we have also evaluated them on the Switchboard task, and we report these results here as well. Finally, we also present the result of crossdomain experiments to evaluate the domain-independence of the constructed systems. 1. INTRODUCTION In this paper we report recent improvements in transcribing conversational telephone speech, as typified by the Voicemail and Switchboard transcription tasks. These improvements are a result of some new algorithms and, in the case of Voicemail, also due to an increa... Jing Huang 0019, Brian Kingsbury, Lidia Mangu, Mukund Padmanabhan, George Saon, Geoffrey Zweig |
INTERSPEECH | 5 |
| 2000 | Real-time multilingual HMM training robust to channel variations
Ea-Ee Jan, Jaime Botella Ordinas, George Saon, Salim Roukos |
INTERSPEECH | 3 |
| 2000 | Minimum Bayes error feature selectionabstractWe consider the problem of designing a linear transformation 2 IR pn , of rank p n, which projects the features of a classier x 2 IR n onto y = x 2 IR p such as to achieve minimum Bayes error (or probability of misclassication). Two avenues will be explored: the rst is to maximize the -average divergence between the class densities and the second is to minimize the union Bhattacharyya bound in the range of . While both approaches yield similar performance in practice, they outperform standard LDA features and show a 10% relative improvement in the word error rate over state-of-the-art cepstral features on a large vocabulary telephony speech recognition task. 1 Introduction Modern speech recognition systems use cepstral features characterizing the short-term spectrum of the speech signal for classifying frames into phonetic classes. These features are augmented with dynamic information from the adjacent frames to capture transient spectral events in the signal. What ... George Saon, Mukund Padmanabhan |
INTERSPEECH | 1 |
| 2000 | Minimum Bayes Error Feature Selection for Continuous Speech RecognitionabstractWe consider the problem of designing a linear transformation () E lRPx n, of rank p ~ n, which projects the features of a classifier x E lRn onto y = ()x E lRP such as to achieve minimum Bayes error (or probabil(cid:173) ity of misclassification). Two avenues will be explored: the first is to maximize the ()-average divergence between the class densities and the second is to minimize the union Bhattacharyya bound in the range of (). While both approaches yield similar performance in practice, they out(cid:173) perform standard LDA features and show a 10% relative improvement in the word error rate over state-of-the-art cepstral features on a large vocabulary telephony speech recognition task. George Saon, Mukund Padmanabhan |
NIPS | 1 |
| 1999 | Recent improvements in voicemail transcriptionabstractIn this paper we report recent improvements in voicemail transcription. Last year, the speaker independent and speaker adapted word error rates (WER) on the Voicemail Transcription task were reported at 41.94% and 38.18% respectively. This year, we report a relative improvement of 18% in the speaker independent performance and 11% in the speaker adapted performance over last year. This improvement is a result of some new algorithms and an increase in the amount of training data. In the following sections, we describe the contribution of several components to improving the word error rate. 1. INTRODUCTION In this paper we report recent improvements in voicemail transcription. The voicemail transcription task was introduced last year [1] as representing a style of conversational telephone speech that is somewhat different from the Switchboard and CallHome databases. Last year, the speaker independent and speaker adapted word error rates (WER) on this task were reported at 41.94% and 38... Mukund Padmanabhan, George Saon, Sankar Basu, Jing Huang 0019, Geoffrey Zweig |
EUROSPEECH | 2 |
| 1999 | Cursive word recognition using a random field based hidden Markov model
George Saon |
Int. J. Document Anal. Recognit. | 1 |
| 1997 | Binary pattern recognition using Markov random fields and HMMsabstractWe present a stochastic framework for the recognition of binary random patterns which advantageously combine HMMs and Markov random fields (MRFs). The HMM component of the model analyzes the image along one direction, in a specific state observation probability given by the product of causal MRF-like pixel conditional probabilities. Aspects concerning definition, training and recognition via this type of model are developed throughout the paper. Experiments were performed on handwritten digits and words in a small lexicon. For the latter, we report a 89.68% average word recognition rate on the SRTP French postal cheque database (7057 words, 1779 scriptors). George Saon, Abdel Belaïd |
ICASSP | 1 |
| 1997 | High Performance Unconstrained Word Recognition System Combining HMMs and Markov Random FieldsabstractIn this paper we present a system for the recognition of handwritten words on literal check amounts which advantageously combine HMMs and Markov random fields (MRFs). It operates at pixel level, in a holistic manner, on height normalized word images which are viewed as random field realizations. The HMM analyzes the image along the horizontal writing direction, in a specific state observation probability given by the column product of causal MRF-like pixel conditional probabilities. Aspects concerning definition, training and recognition via this type of model are developed throughout the paper. We report a 90.08% average word recognition rate on 2378 words and a 79.52% amount rate on 579 amounts of the SRTP* French postal check database (7031 words, 1779 amounts, different scriptors). George Saon, Abdel Belaïd |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1995 | Stochastic trajectory modeling for recognition of unconstrained handwritten wordsabstractIn this paper we describe an off-line handwritten word recognition system applied to the identification of literal french check amounts. It consists of three successive levels denoted as character, word and phrase level, each of them being related to the previous ones via conditional probability distributions. Training is done on character samples extracted from amount images which are modeled as trajectories in some feature space. At word level, guided by a dictionary, an internal character segmentation algorithm is used in order to maximize a global word probability measure. A stochastic grammar for a priori grammar generation probability of a phrase is proposed at the last level. Results obtained on a 1779 amounts data base provided by the SRTP are encouraging, showing our system open to further improvements. George Saon, Abdel Belaïd, Yifan Gong 0001 |
ICDAR | 1 |