VLDB 2026 Research / reviewers in the wild / expert
Michael Picheny
dblp:73/4588 · also Michael A. Picheny
· DBLP profile ↗
130ranked-venue papers
3as first author
10since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 111 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 66 · 3 first-author · 7 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A Comparison of Semi-Supervised Learning Techniques for Streaming ASR at ScaleabstractUnpaired text and audio injection have emerged as dominant methods for improving ASR performance in the absence of a large labeled corpus. However, little guidance exists on deploying these methods to improve production ASR systems that are trained on very large supervised corpora and with realistic requirements like a constrained model size and CPU budget, streaming capability, and a rich lattice for rescoring and for downstream NLU tasks. In this work, we compare three state-of-the-art semi-supervised methods encompassing both unpaired text and audio as well as several of their combinations in a controlled setting using joint training. We find that in our setting these methods offer many improvements beyond raw WER, including substantial gains in tail-word WER, decoder computation during inference, and lattice density. Cal Peyser, Michael Picheny, Kyunghyun Cho, Rohit Prabhavalkar, W. Ronny Huang, Tara N. Sainath |
ICASSP | 2 |
| 2023 | Improving Joint Speech-Text Representations Without Alignment
Cal Peyser, Zhong Meng, Rohit Prabhavalkar, Andrew Rosenberg, Tara N. Sainath, Michael Picheny, Kyunghyun Cho |
INTERSPEECH | 6 |
| 2023 | The MALACH Corpus: Results with End-to-End Architectures and Pretraining
Michael Picheny, Daiheng Zhang, Lining Zhang |
INTERSPEECH | 1 |
| 2022 | Towards Measuring Fairness in Speech Recognition: Casual Conversations Dataset TranscriptionsabstractThe problem of machine learning systems demonstrating bias towards specific groups of individuals has been studied extensively, particularly in the Facial Recognition area, but much less so in Automatic Speech Recognition (ASR). This paper presents initial Speech Recognition results on “Casual Conversations” – a publicly released 846 hour corpus designed to help researchers evaluate their computer vision and audio models for accuracy across a diverse set of metadata, including age, gender, and skin tone. The entire corpus has been manually transcribed, allowing for detailed ASR evaluations across these metadata. Multiple ASR models are evaluated, including models trained on LibriSpeech, 14,000 hour transcribed, and over 2 million hour untranscribed social media videos. Significant differences in word error rate across gender and skin tone are observed at times for all models. We are releasing human transcripts from the Casual Conversations dataset to encourage the community to develop a variety of techniques to reduce these statistical biases. Chunxi Liu, Michael Picheny, Leda Sari, Pooja Chitkara, Alex Xiao, Xiaohui Zhang 0007, Mark Chou, Andres Alvarado, Caner Hazirbas, Yatharth Saraf |
ICASSP | 2 |
| 2022 | Towards Disentangled Speech RepresentationsabstractThe careful construction of audio representations has become a dominant feature in the design of approaches to many speech tasks.Increasingly, such approaches have emphasized "disentanglement", where a representation contains only parts of the speech signal relevant to transcription while discarding irrelevant information.In this paper, we construct a representation learning task based on joint modeling of ASR and TTS, and seek to learn a representation of audio that disentangles that part of the speech signal that is relevant to transcription from that part which is not.We present empirical evidence that successfully finding such a representation is tied to the randomness inherent in training.We then make the observation that these desired, disentangled solutions to the optimization problem possess unique statistical properties.Finally, we show that enforcing these properties during training improves WER by 24.5% relative on average for our joint modeling task.These observations motivate a novel approach to learning effective audio representations. Cal Peyser, W. Ronny Huang, Andrew Rosenberg, Tara N. Sainath, Michael Picheny, Kyunghyun Cho |
INTERSPEECH | 5 |
| 2022 | Dual Learning for Large Vocabulary On-Device ASRabstractDual learning is a paradigm for semi-supervised machine learning that seeks to leverage unsupervised data by solving two opposite tasks at once. In this scheme, each model is used to generate pseudo-labels for unlabeled examples that are used to train the other model. Dual learning has seen some use in speech processing by pairing ASR and TTS as dual tasks. However, these results mostly address only the case of using unpaired examples to compensate for very small supervised datasets, and mostly on large, non-streaming models. Dual learning has not yet been proven effective for using unsupervised data to improve realistic on-device streaming models that are already trained on large supervised corpora. We provide this missing piece though an analysis of an on-device-sized streaming conformer trained on the entirety of Librispeech, showing relative WER improvements of 10.7%/5.2% without an LM and 11.7%/16.4% with an LM. Cal Peyser, W. Ronny Huang, Tara N. Sainath, Rohit Prabhavalkar, Michael Picheny, Kyunghyun Cho |
SLT | 5 |
| 2021 | Multimodal Clustering Networks for Self-supervised Learning from Unlabeled VideosabstractMultimodal self-supervised learning is getting more and more attention as it allows not only to train large networks without human supervision but also to search and retrieve data across various modalities. In this context, this paper proposes a framework that, starting from a pre-trained backbone, learns a common multimodal embedding space that, in addition to sharing representations across different modalities, enforces a grouping of semantically similar instances. To this end, we extend the concept of instance-level contrastive learning with a multimodal clustering step in the training pipeline to capture semantic similarities across modalities. The resulting embedding space enables retrieval of samples across all modalities, even from unseen datasets and different domains. To evaluate our approach, we train our model on the HowTo100M dataset and evaluate its zero-shot retrieval capabilities in two challenging domains, namely text-to-video retrieval, and temporal action localization, showing state-of-the-art results on four different datasets. Brian Chen 0001, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne, Samuel Thomas 0001, Angie W. Boggust, Rameswar Panda, Brian Kingsbury, Rogério Feris, David F. Harwath, James R. Glass, Michael Picheny, Shih-Fu Chang |
ICCV | 12 |
| 2021 | Speak or Chat with Me: End-to-End Spoken Language Understanding System with Flexible InputsabstractA major focus of recent research in spoken language understanding (SLU) has been on the end-to-end approach where a single model can predict intents directly from speech inputs without intermediate transcripts.However, this approach presents some challenges.First, since speech can be considered as personally identifiable information, in some cases only automatic speech recognition (ASR) transcripts are accessible.Second, intent-labeled speech data is scarce.To address the first challenge, we propose a novel system that can predict intents from flexible types of inputs: speech, ASR transcripts, or both.We demonstrate strong performance for either modality separately, and when both speech and ASR transcripts are available, through system combination, we achieve better results than using a single input modality.To address the second challenge, we leverage a semantically robust pre-trained BERT model and adopt a cross-modal system that co-trains text embeddings and acoustic embeddings in a shared latent space.We further enhance this system by utilizing an acoustic module pre-trained on LibriSpeech and domain-adapting the text module on our target datasets.Our experiments show significant advantages for these pre-training and fine-tuning strategies, resulting in a system that achieves competitive intent-classification performance on Snips SLU and Fluent Speech Commands datasets. Sujeong Cha, Wangrui Hou, Hyun Jung, My Phung, Michael Picheny, Hong-Kwang Jeff Kuo, Samuel Thomas 0001, Edmilson da Silva Morais |
Interspeech | 5 |
| 2021 | Cascaded Multilingual Audio-Visual Learning from VideosabstractIn this paper, we explore self-supervised audio-visual models that learn from instructional videos.Prior work has shown that these models can relate spoken words and sounds to visual content after training on a large-scale dataset of videos, but they were only trained and evaluated on videos in English.To learn multilingual audio-visual representations, we propose a cascaded approach that leverages a model trained on English videos and applies it to audio-visual data in other languages, such as Japanese videos.With our cascaded approach, we show an improvement in retrieval performance of nearly 10x compared to training on the Japanese videos solely.We also apply the model trained on English videos to Japanese and Hindi spoken captions of images, achieving state-of-the-art performance. Andrew Rouditchenko, Angie W. Boggust, David F. Harwath, Samuel Thomas 0001, Hilde Kuehne, Brian Chen 0001, Rameswar Panda, Rogério Feris, Brian Kingsbury, Michael Picheny, James R. Glass |
Interspeech | 10 |
| 2021 | AVLnet: Learning Audio-Visual Language Representations from Instructional VideosabstractCurrent methods for learning visually grounded language from videos often rely on text annotation, such as human generated captions or machine generated automatic speech recognition (ASR) transcripts. In this work, we introduce the Audio-Video Language Network (AVLnet), a self-supervised network that learns a shared audio-visual embedding space directly from raw video inputs. To circumvent the need for text annotation, we learn audio-visual representations from randomly segmented video clips and their raw audio waveforms. We train AVLnet on HowTo100M, a large corpus of publicly available instructional videos, and evaluate on image retrieval and video retrieval tasks, achieving state-of-the-art performance. We perform analysis of AVLnet's learned representations, showing our model utilizes speech and natural sounds to learn audio-visual concepts. Further, we propose a tri-modal model that jointly processes raw audio, video, and text captions from videos to learn a multi-modal semantic embedding space useful for text-video retrieval. Our code, data, and trained models will be released at avlnet.csail.mit.edu Andrew Rouditchenko, Angie W. Boggust, David F. Harwath, Brian Chen 0001, Dhiraj Joshi, Samuel Thomas 0001, Kartik Audhkhasi, Hilde Kuehne, Rameswar Panda, Rogério Feris, Brian Kingsbury, Michael Picheny, Antonio Torralba 0001, James R. Glass |
Interspeech | 12 |
| 2020 | Leveraging Unpaired Text Data for Training End-To-End Speech-to-Intent SystemsabstractTraining an end-to-end (E2E) neural network speech-to-intent (S2I) system that directly extracts intents from speech requires large amounts of intent-labeled speech data, which is time consuming and expensive to collect. Initializing the S2I model with an ASR model trained on copious speech data can alleviate data sparsity. In this paper, we attempt to leverage NLU text resources. We implemented a CTC-based S2I system that matches the performance of a state-of-the-art, traditional cascaded SLU system. We performed controlled experiments with varying amounts of speech and text training data. When only a tenth of the original data is available, intent classification accuracy degrades by 7.6% absolute. Assuming we have additional text-to-intent data (without speech) available, we investigated two techniques to improve the S2I system: (1) transfer learning, in which acoustic embeddings for intent classification are tied to fine-tuned BERT text embeddings; and (2) data augmentation, in which the text-to-intent data is converted into speech-to-intent data using a multi-speaker text-to-speech system. The proposed approaches recover 80% of performance lost due to using limited intent-labeled speech. Hong-Kwang Jeff Kuo, Samuel Thomas 0001, Zvi Kons, Kartik Audhkhasi, Brian Kingsbury, Ron Hoory, Michael Picheny |
ICASSP | 8 |
| 2020 | Improving Efficiency in Large-Scale Decentralized Distributed TrainingabstractDecentralized Parallel SGD (D-PSGD) and its asynchronous variant Asynchronous Parallel SGD (AD-PSGD) is a family of distributed learning algorithms that have been demonstrated to perform well for large-scale deep learning tasks. One drawback of (A)D-PSGD is that the spectral gap of the mixing matrix decreases when the number of learners in the system increases, which hampers convergence. In this paper, we investigate techniques to accelerate (A)D-PSGD based training by improving the spectral gap while minimizing the communication cost. We demonstrate the effectiveness of our proposed techniques by running experiments on the 2000-hour Switchboard speech recognition task and the ImageNet computer vision task. On an IBM P9 supercomputer, our system is able to train an LSTM acoustic model in 2.28 hours with 7.5% WER on the Hub5-2000 Switchboard (SWB) test set and 13.3% WER on the CallHome (CH) test set using 64 V100 GPUs and in 1.98 hours with 7.7% WER on SWB and 13.3% WER on CH using 128 V100 GPUs, the fastest training time reported to date. Wei Zhang 0022, Abdullah Kayi, Ulrich Finkler, Brian Kingsbury, George Saon, Youssef Mroueh, Alper Buyuktosunoglu, David S. Kung 0001, Michael Picheny |
ICASSP | 12 |
| 2019 | Semi-Supervised Training and Data Augmentation for Adaptation of Automatic Broadcast News Captioning SystemsabstractIn this paper we present a comprehensive study on building and adapting deep neural network based speech recognition systems for automatic closed captioning. We develop the proposed systems by first building base automatic speech recognition (ASR) systems that are not specific to any particular show or station. These models are trained on nearly 6000 hours of broadcast news data using conventional hybrid and more recent attention based end-to-end acoustic models. We then employ various adaptation and data augmentation strategies to further improve the trained base models. We use 535 hours of data from two independent BN sources to study how the base models can be customized. We observe up to 32% relative improvement using the proposed techniques on test sets related to, but independent of the adaptation data. At these low word error rates (WERs), we believe the customized BN ASR systems can be used effectively for automatic closed captioning. Samuel Thomas 0001, Masayuki Suzuki, Zoltán Tüske, Larry Sansone, Michael Picheny |
ASRU | 6 |
| 2019 | Simplified LSTMS for Speech RecognitionabstractIn this paper we explore new variants of Long Short-Term Memory (LSTM) networks for sequential modeling of acoustic features. In particular, we show that: (i) removing the output gate, (ii) replacing the hyperbolic tangent nonlinearity at the cell output with hard tanh, and (iii) collapsing the cell and hidden state vectors leads to a model that is conceptually simpler than and comparable in effectiveness to a regular LSTM for speech recognition. The proposed model has 25% fewer parameters than an LSTM with the same number of cells, trains faster because it has larger gradients leading to larger steps in weight space, and reaches a better optimum because there are fewer nonlinearities to traverse across layers. We report experimental results for both hybrid and CTC acoustic models on three publicly available English datasets: Switchboard 300 hours telephone conversations, 400 hours broadcast news transcription, and the MALACH 176 hours corpus of Holocaust survivor testimonies. In all cases the proposed models achieve similar or better accuracy than regular LSTMs while being conceptually simpler. George Saon, Zoltán Tüske, Kartik Audhkhasi, Brian Kingsbury, Michael Picheny, Samuel Thomas 0001 |
ASRU | 5 |
| 2019 | Pre-training of Speaker Embeddings for Low-latency Speaker Change Detection in Broadcast NewsabstractIn this work, we investigate pre-training of neural network based speaker embeddings for low-latency speaker change detection. Our proposed system takes two speech segments, generates embeddings using shared Siamese layers and then classifies the concatenated embeddings depending on whether they are spoken by the same speaker. We investigate gender classification, contrastive loss and triplet loss based pre-training of the embedding layers and also joint training of the embedding layers along with a same/different classifier. Training is performed on 2-second single speaker segments based on ground truth speaker segmentation of broadcast news data. However, during test, we use the detection system in a practical low-latency setting for annotating automatic closed captions. In contrast to training, test pairs are now created around automatic speech recognition (ASR) based segmentation boundaries. The ASR segments are often shorter than 2 seconds causing duration mismatch during testing. In our experiments, although the baseline i-vector based classifier performs well, the proposed triplet loss based pre-training followed by joint training provides 7-50% relative F-measure improvement in matched and mismatched conditions. In addition, the degradation in performance is less severe for network based embeddings as compared to using i-vectors in the variable duration test conditions. Leda Sari, Samuel Thomas 0001, Mark Hasegawa-Johnson, Michael Picheny |
ICASSP | 4 |
| 2019 | Acoustically Grounded Word Embeddings for Improved Acoustics-to-word Speech RecognitionabstractDirect acoustics-to-word (A2W) systems for end-to-end automatic speech recognition are simpler to train, and more efficient to decode with, than sub-word systems. However, A2W systems can have difficulties at training time when data is limited, and at decoding time when recognizing words outside the training vocabulary. To address these shortcomings, we investigate the use of recently proposed acoustic and acoustically grounded word embedding techniques in A2W systems. The idea is based on treating the final pre-softmax weight matrix of an AWE recognizer as a matrix of word embedding vectors, and using an externally trained set of word embeddings to improve the quality of this matrix. In particular we introduce two ideas: (1) Enforcing similarity at training time between the external embeddings and the recognizer weights, and (2) using the word embeddings at test time for predicting out-of-vocabulary words. Our word embedding model is acoustically grounded, that is it is learned jointly with acoustic embeddings so as to encode the words' acoustic-phonetic content; and it is parametric, so that it can embed any arbitrary (potentially out-of-vocabulary) sequence of characters. We find that both techniques improve the performance of an A2W recognizer on conversational telephone speech. Shane Settle, Kartik Audhkhasi, Karen Livescu, Michael Picheny |
ICASSP | 4 |
| 2019 | English Broadcast News Speech Recognition by Humans and MachinesabstractWith recent advances in deep learning, considerable attention has been given to achieving automatic speech recognition performance close to human performance on tasks like conversational telephone speech (CTS) recognition. In this paper we evaluate the usefulness of these proposed techniques on broadcast news (BN), a similar challenging task. We also perform a set of recognition measurements to understand how close the achieved automatic speech recognition results are to human performance on this task. On two publicly available BN test sets, DEV04F and RT04, our speech recognition system using LSTM and residual network based acoustic models with a combination of n-gram and neural network language models performs at 6.5% and 5.9% word error rate. By achieving new performance milestones on these test sets, our experiments show that techniques developed on other related tasks, like CTS, can be transferred to achieve similar performance. In contrast, the best measured human recognition performance on these test sets is much lower, at 3.6% and 2.8% respectively, indicating that there is still room for new techniques and improvements in this space, to reach human performance levels. Samuel Thomas 0001, Masayuki Suzuki, Gakuto Kurata, Zoltán Tüske, George Saon, Brian Kingsbury, Michael Picheny, Tom Dibert, Alice Kaiser-Schatzlein, Bern Samko |
ICASSP | 8 |
| 2019 | Distributed Deep Learning Strategies for Automatic Speech RecognitionabstractIn this paper, we propose and investigate a variety of distributed deep learning strategies for automatic speech recognition (ASR) and evaluate them with a state-of-the-art Long short-term memory (LSTM) acoustic model on the 2000-hour Switchboard (SWB2000), which is one of the most widely used datasets for ASR performance benchmark. We first investigate what are the proper hyper-parameters (e.g., learning rate) to enable the training with sufficiently large batch size without impairing the model accuracy. We then implement various distributed strategies, including Synchronous (SYNC) , Asynchronous Decentralized Parallel SGD (ADPSGD) and the hybrid of the two HYBRID, to study their runtime/accuracy trade-off. We show that we can train the LSTM model using ADPSGD in 14 hours with 16 NVIDIA P100 GPUs to reach a 7.6% WER on the Hub5-2000 Switchboard (SWB) test set and a 13.1% WER on the Call-Home (CH) test set. Furthermore, we can train the model using HYBRID in 11.5 hours with 32 NVIDIA V100 GPUs without loss in accuracy. Wei Zhang 0022, Ulrich Finkler, Brian Kingsbury, George Saon, David S. Kung 0001, Michael Picheny |
ICASSP | 7 |
| 2019 | Identifying Mood Episodes Using Dialogue Features from Clinical InterviewsabstractBipolar disorder, a severe chronic mental illness characterized by pathological mood swings from depression to mania, requires ongoing symptom severity tracking to both guide and measure treatments that are critical for maintaining long-term health. Mental health professionals assess symptom severity through semi-structured clinical interviews. During these interviews, they observe their patients' spoken behaviors, including both what the patients say and how they say it. In this work, we move beyond acoustic and lexical information, investigating how higher-level interactive patterns also change during mood episodes. We then perform a secondary analysis, asking if these interactive patterns, measured through dialogue features, can be used in conjunction with acoustic features to automatically recognize mood episodes. Our results show that it is beneficial to consider dialogue features when analyzing and building automated systems for predicting and monitoring mood. Zakaria Aldeneh, Mimansa Jaiswal, Michael Picheny, Melvin G. McInnis, Emily Mower Provost |
INTERSPEECH | 3 |
| 2019 | Forget a Bit to Learn Better: Soft Forgetting for CTC-Based Automatic Speech Recognition
Kartik Audhkhasi, George Saon, Zoltán Tüske, Brian Kingsbury, Michael Picheny |
INTERSPEECH | 5 |
| 2019 | Acoustic Model Optimization Based on Evolutionary Stochastic Gradient Descent with Anchors for Automatic Speech RecognitionabstractEvolutionary stochastic gradient descent (ESGD) was proposed as a population-based approach that combines the merits of gradient-aware and gradient-free optimization algorithms for superior overall optimization performance.In this paper we investigate a variant of ESGD for optimization of acoustic models for automatic speech recognition (ASR).In this variant, we assume the existence of a well-trained acoustic model and use it as an anchor in the parent population whose good "gene" will propagate in the evolution to the offsprings.We propose an ESGD algorithm leveraging the anchor models such that it guarantees the best fitness of the population will never degrade from the anchor model.Experiments on 50-hour Broadcast News (BN50) and 300-hour Switchboard (SWB300) show that the ESGD with anchors can further improve the loss and ASR performance over the existing well-trained acoustic models. Michael Picheny |
INTERSPEECH | 2 |
| 2019 | Large-Scale Mixed-Bandwidth Deep Neural Network Acoustic Modeling for Automatic Speech RecognitionabstractIn automatic speech recognition (ASR), wideband (WB) and narrowband (NB) speech signals with different sampling rates typically use separate acoustic models. Therefore mixed-bandwidth (MB) acoustic modeling has important practical values for ASR system deployment. In this paper, we extensively investigate large-scale MB deep neural network acoustic modeling for ASR using 1,150 hours of WB data and 2,300 hours of NB data. We study various MB strategies including downsampling, upsampling and bandwidth extension for MB acoustic modeling and evaluate their performance on 8 diverse WB and NB test sets from various application domains. To deal with the large amounts of training data, distributed training is carried out on multiple GPUs using synchronous data parallelism. Khoi-Nguyen C. Mac, Wei Zhang 0022, Michael Picheny |
INTERSPEECH | 4 |
| 2019 | Challenging the Boundaries of Speech Recognition: The MALACH CorpusabstractThere has been huge progress in speech recognition over the last several years. Tasks once thought extremely difficult, such as SWITCHBOARD, now approach levels of human performance. The MALACH corpus (LDC catalog LDC2012S05), a 375-Hour subset of a large archive of Holocaust testimonies collected by the Survivors of the Shoah Visual History Foundation, presents significant challenges to the speech community. The collection consists of unconstrained, natural speech filled with disfluencies, heavy accents, age-related coarticulations, un-cued speaker and language switching, and emotional speech - all still open problems for speech recognition systems. Transcription is challenging even for skilled human annotators. This paper proposes that the community place focus on the MALACH corpus to develop speech recognition systems that are more robust with respect to accents, disfluencies and emotional speech. To reduce the barrier for entry, a lexicon and training and testing setups have been created and baseline results using current deep learning technologies are presented. The metadata has just been released by LDC (LDC2019S11). It is hoped that this resource will enable the community to build on top of these baselines so that the extremely important information in these and related oral histories becomes accessible to a wider audience. Michael Picheny, Zoltán Tüske, Brian Kingsbury, Kartik Audhkhasi, George Saon |
INTERSPEECH | 1 |
| 2019 | Detection and Recovery of OOVs for Improved English Broadcast News Captioning
Samuel Thomas 0001, Kartik Audhkhasi, Zoltán Tüske, Michael Picheny |
INTERSPEECH | 5 |
| 2019 | A Highly Efficient Distributed Deep Learning System for Automatic Speech RecognitionabstractModern Automatic Speech Recognition (ASR) systems rely on distributed deep learning to for quick training completion.To enable efficient distributed training, it is imperative that the training algorithms can converge with a large mini-batch size.In this work, we discovered that Asynchronous Decentralized Parallel Stochastic Gradient Descent (ADPSGD) can work with much larger batch size than commonly used Synchronous SGD (SSGD) algorithm.On commonly used public SWB-300 and SWB-2000 ASR datasets, ADPSGD can converge with a batch size 3X as large as the one used in SSGD, thus enable training at a much larger scale.Further, we proposed a Hierarchical-ADPSGD (H-ADPSGD) system in which learners on the same computing node construct a super learner via a fast allreduce implementation, and super learners deploy ADPSGD algorithm among themselves.On a 64 Nvidia V100 GPU cluster connected via a 100Gb/s Ethernet network, our system is able to train SWB-2000 to reach a 7.6% WER on the Hub5-2000 Switchboard (SWB) test-set and a 13.2% WER on the Callhome (CH) test-set in 5.2 hours.To the best of our knowledge, this is the fastest ASR training system that attains this level of model accuracy for SWB-2000 task to be ever reported in the literature. Wei Zhang 0022, Ulrich Finkler, George Saon, Abdullah Kayi, Alper Buyuktosunoglu, Brian Kingsbury, David S. Kung 0001, Michael Picheny |
INTERSPEECH | 9 |
| 2019 | Kernel Approximation Methods for Speech RecognitionabstractWe study the performance of kernel methods on the acoustic modeling task for automatic speech recognition, and compare their performance to deep neural networks (DNNs). To scale the kernel methods to large data sets, we use the random Fourier feature method of Rahimi and Recht (2007). We propose two novel techniques for improving the performance of kernel acoustic models. First, we propose a simple but effective feature selection method which reduces the number of random features required to attain a fixed level of performance. Second, we present a number of metrics which correlate strongly with speech recognition performance when computed on the heldout set; we attain improved performance by using these metrics to decide when to stop training. Additionally, we show that the linear bottleneck method of Sainath et al. (2013a) improves the performance of our kernel models significantly, in addition to speeding up training and making the models more compact. Leveraging these three methods, the kernel methods attain token error rates between $0.5\%$ better and $0.1\%$ worse than fully-connected DNNs across four speech recognition data sets, including the TIMIT and Broadcast News benchmark tasks. Avner May, Alireza Bagheri Garakani, Zhiyun Lu, Aurélien Bellet, Linxi Fan, Michael Collins 0001, Daniel Hsu 0001, Brian Kingsbury, Michael Picheny, Fei Sha |
J. Mach. Learn. Res. | 11 |
| 2018 | Building Competitive Direct Acoustics-to-Word Models for English Conversational Speech RecognitionabstractDirect acoustics-to-word (A2W) models in the end-to-end paradigm have received increasing attention compared to conventional subword based automatic speech recognition models using phones, characters, or context-dependent hidden Markov model states. This is because A2W models recognize words from speech without any decoder, pronunciation lexicon, or externally-trained language model, making training and decoding with such models simple. Prior work has shown that A2W models require orders of magnitude more training data in order to perform comparably to conventional models. Our work also showed this accuracy gap when using the English Switchboard-Fisher data set. This paper describes a recipe to train an A2W model that closes this gap and is at-par with state-of-the-art sub-word based models. We achieve a word error rate of 8.8.8%/13.9% on the Hub5-2000 Switchboard/CallHome test sets without any decoder or language model. We find that model initialization, training data order, and regularization have the most impact on the A2W model performance. Next, we present a joint word-character A2W model that learns to first spell the word and then recognize it. This model provides a rich output to the user instead of simple word hypotheses, making it especially useful in the case of words unseen or rarely-seen during training. Kartik Audhkhasi, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, Michael Picheny |
ICASSP | 5 |
| 2018 | Evolutionary Stochastic Gradient Descent for Optimization of Deep Neural NetworksabstractWe propose a population-based Evolutionary Stochastic Gradient Descent (ESGD) framework for optimizing deep neural networks. ESGD combines SGD and gradient-free evolutionary algorithms as complementary algorithms in one framework in which the optimization alternates between the SGD step and evolution step to improve the average fitness of the population. With a back-off strategy in the SGD step and an elitist strategy in the evolution step, it guarantees that the best fitness in the population will never degrade. In addition, individuals in the population optimized with various SGD-based optimizers using distinct hyper-parameters in the SGD step are considered as competing species in a coevolution setting such that the complementarity of the optimizers is also taken into account. The effectiveness of ESGD is demonstrated across multiple applications including speech recognition, image recognition and language modeling, using networks with a variety of deep architectures. Wei Zhang 0022, Zoltán Tüske, Michael Picheny |
NeurIPS | 4 |
| 2017 | Training variance and performance evaluation of neural networks in speechabstractIn this work we study variance in the results of neural network training on a wide variety of configurations in automatic speech recognition. Although this variance itself is well known, this is, to the best of our knowledge, the first paper that performs an extensive empirical study on its effects in speech recognition. We view training as sampling from a distribution and show that these distributions can have a substantial variance. These results show the urgent need to rethink the way in which results in the literature are reported and interpreted. Ewout van den Berg, Bhuvana Ramabhadran, Michael Picheny |
ICASSP | 3 |
| 2017 | End-to-end speech recognition and keyword search on low-resource languagesabstractIn recent years, so-called, “end-to-end” speech recognition systems have emerged as viable alternatives to traditional ASR frameworks. Keyword search, localizing an orthographic query in a speech corpus, is typically performed by using automatic speech recognition (ASR) to generate an index. Previous work has evaluated the use of end-to-end systems for ASR on well known corpora (WSJ, Switchboard, TIMIT, etc.) in high-resource languages like English and Mandarin. In this work, we investigate the use of Connectionist Temporal Classification (CTC) networks, recurrent encoder-decoders with attention, two end-to-end ASR systems for keyword search and speech recognition on low resource languages. We find end-to-end systems can generate high quality 1-best transcripts on low-resource languages, but, because they generate very sharp posteriors, their utility is limited for KWS. We explore a number of ways to address this limitation with modest success. Experimental results reported are based on the IARPA BABEL OP3 languages and evaluation framework. This paper represents the first results using “end-to-end” techniques for speech recognition and keyword search on low-resource languages. Andrew Rosenberg, Kartik Audhkhasi, Abhinav Sethy, Bhuvana Ramabhadran, Michael Picheny |
ICASSP | 5 |
| 2017 | Direct Acoustics-to-Word Models for English Conversational Speech RecognitionabstractRecent work on end-to-end automatic speech recognition (ASR) has shown that the connectionist temporal classification (CTC) loss can be used to convert acoustics to phone or character sequences.Such systems are used with a dictionary and separately-trained Language Model (LM) to produce word sequences.However, they are not truly end-to-end in the sense of mapping acoustics directly to words without an intermediate phone representation.In this paper, we present the first results employing direct acoustics-to-word CTC models on two well-known public benchmark tasks: Switchboard and Call-Home.These models do not require an LM or even a decoder at run-time and hence recognize speech with minimal complexity.However, due to the large number of word output units, CTC word models require orders of magnitude more data to train reliably compared to traditional systems.We present some techniques to mitigate this issue.Our CTC word model achieves a word error rate of 13.0%/18.8%on the Hub5-2000 Switchboard/CallHome test sets without any LM or decoder compared with 9.6%/16.0%for phone-based CTC with a 4-gram LM.We also present rescoring results on CTC word model lattices to quantify the performance benefits of a LM, and contrast the performance of word and phone CTC models. Kartik Audhkhasi, Bhuvana Ramabhadran, George Saon, Michael Picheny, David Nahamoo |
INTERSPEECH | 4 |
| 2017 | English Conversational Telephone Speech Recognition by Humans and MachinesabstractOne of the most difficult speech recognition tasks is accurate recognition of human to human communication. Advances in deep learning over the last few years have produced major speech recognition improvements on the representative Switchboard conversational corpus. Word error rates that just a few years ago were 14% have dropped to 8.0%, then 6.6% and most recently 5.8%, and are now believed to be within striking range of human performance. This then raises two issues - what IS human performance, and how far down can we still drive speech recognition error rates? A recent paper by Microsoft suggests that we have already achieved human performance. In trying to verify this statement, we performed an independent set of human performance measurements on two conversational tasks and found that human performance may be considerably better than what was earlier reported, giving the community a significantly harder goal to achieve. We also report on our own efforts in this area, presenting a set of acoustic and language modeling techniques that lowered the word error rate of our own English conversational telephone LVCSR system to the level of 5.5%/10.3% on the Switchboard/CallHome subsets of the Hub5 2000 evaluation, which - at least at the writing of this paper - is a new performance milestone (albeit not at what we measure to be human performance!). On the acoustic side, we use a score fusion of three models: one LSTM with multiple feature inputs, a second LSTM trained with speaker-adversarial multi-task learning and a third residual net (ResNet) with 25 convolutional layers and time-dilated convolutions. On the language modeling side, we use word and character LSTMs and convolutional WaveNet-style language models. George Saon, Gakuto Kurata, Tom Sercu, Kartik Audhkhasi, Samuel Thomas 0001, Dimitrios Dimitriadis, Bhuvana Ramabhadran, Michael Picheny, Lynn-Li Lim, Bergul Roomi, Phil Hall |
INTERSPEECH | 9 |
| 2017 | Parallel Deep Neural Network Training for Big Data on Blue Gene/QabstractDeep Neural Networks (DNNs) have recently been shown to significantly outperform existing machine learning techniques in several pattern recognition tasks. DNNs are the state-of-the-art models used in image recognition, object detection, classification and tracking, and speech and language processing applications. The biggest drawback to DNNs has been the enormous cost in computation and time taken to train the parameters of the networks-often a tenfold increase relative to conventional technologies. Such training time costs can be mitigated by the application of parallel computing algorithms and architectures. However, these algorithms often run into difficulties because of the cost of inter-processor communication bottlenecks. In this paper, we describe how to enable Parallel Deep Neural Network Training on the IBM Blue Gene/Q (BG/Q) computer system. Specifically, we explore DNN training using the data-parallel Hessian-free 2nd order optimization algorithm. Such an algorithm is particularly well-suited to parallelization across a large set of loosely coupled processors. BG/Q, with its excellent inter-processor communication characteristics, is an ideal match for this type of algorithm. The paper discusses how issues regarding programming model and data-dependent imbalances are addressed. Results on large-scale speech tasks show that the performance on BG/Q scales linearly up to 4,096 processes with no loss in accuracy. This allows us to train neural networks using billions of training examples in a few hours. I-Hsin Chung, Tara N. Sainath, Bhuvana Ramabhadran, Michael Picheny, John A. Gunnels, Vernon Austel, Upendra V. Chaudhari, Brian Kingsbury |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | On the importance of event detection for ASRabstractThe performance of modern large vocabulary continuous speech recognition (LVCSR) systems is heavily affected by segment boundaries, proper speaker identification of the segments, as well as removal of spurious data. We propose to use Long Short Term Memory (LSTM) recurrent neural networks to partition audio into speech segments as well as track speaker turns. Additionally, we train an LSTM to also identify music segments. We show that the accurate detection of events, along with removal of silence and music, using our LSTM yields a 9-10% relative improvement in ASR performance. Secondary processing by speaker clustering provides an additional boost in accuracy. Event detection accuracy of the LSTM approach is also described. David Haws, Dimitrios Dimitriadis, George Saon, Samuel Thomas 0001, Michael Picheny |
ICASSP | 5 |
| 2016 | A comparison between deep neural nets and kernel acoustic models for speech recognitionabstractWe study large-scale kernel methods for acoustic modeling and compare to DNNs on performance metrics related to both acoustic modeling and recognition. Measuring perplexity and frame-level classification accuracy, kernel-based acoustic models are as effective as their DNN counterparts. However, on token-error-rates DNN models can be significantly better. We have discovered that this might be attributed to DNN's unique strength in reducing both the perplexity and the entropy of the predicted posterior probabilities. Motivated by our findings, we propose a new technique, entropy regularized perplexity, for model selection. This technique can noticeably improve the recognition performance of both types of models, and reduces the gap between them. While effective on Broadcast News, this technique could be also applicable to other tasks. Zhiyun Lu, Alireza Bagheri Garakani, Avner May, Aurélien Bellet, Linxi Fan, Michael Collins 0001, Brian Kingsbury, Michael Picheny, Fei Sha |
ICASSP | 10 |
| 2015 | Multilingual representations for low resource speech recognition and keyword searchabstractThis paper examines the impact of multilingual (ML) acoustic representations on Automatic Speech Recognition (ASR) and keyword search (KWS) for low resource languages in the context of the OpenKWS15 evaluation of the IARPA Babel program. The task is to develop Swahili ASR and KWS systems within two weeks using as little as 3 hours of transcribed data. Multilingual acoustic representations proved to be crucial for building these systems under strict time constraints. The paper discusses several key insights on how these representations are derived and used. First, we present a data sampling strategy that can speed up the training of multilingual representations without appreciable loss in ASR performance. Second, we show that fusion of diverse multilingual representations developed at different LORELEI sites yields substantial ASR and KWS gains. Speaker adaptation and data augmentation of these representations improves both ASR and KWS performance (up to 8.7% relative). Third, incorporating un-transcribed data through semi-supervised learning, improves WER and KWS performance. Finally, we show that these multilingual representations significantly improve ASR and KWS performance (relative 9% for WER and 5% for MTWV) even when forty hours of transcribed audio in the target language is available. Multilingual representations significantly contributed to the LORELEI KWS systems winning the OpenKWS15 evaluation. Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran, Abhinav Sethy, Kartik Audhkhasi, Ellen Eide, Lidia Mangu, Markus Nußbaum-Thom, Michael Picheny, Zoltán Tüske, Pavel Golik, Ralf Schlüter, Hermann Ney, Mark J. F. Gales, Kate M. Knill, Anton Ragni, Philip C. Woodland |
ASRU | 10 |
| 2015 | Order-free spoken term detectionabstractIn this paper, we propose Time-Marked Word (TMW) lists as a replacement for the lattices and Confusion Networks (CNs) widely used as indexing vehicles for Spoken Term Detection (STD). In a TMW list, candidates are simply tagged with posterior probabilities and time information and stored as a large list of words: the additional ordering present in a lattice or CN is discarded. TMW lists compactly summarize a large ASR search space. Representing a large search space is critical for STD metrics such as ATWV that heavily penalize misses of rare keywords. Comparisons on the OpenKWS 2014 Tamil limited language pack task [1] show that the new TMW-based indexing results in better performance while being faster and having a smaller footprint. Lidia Mangu, George Saon, Michael Picheny, Brian Kingsbury |
ICASSP | 3 |
| 2015 | The IBM 2015 English conversational telephone speech recognition systemabstractWe describe the latest improvements to the IBM English conversational telephone speech recognition system. Some of the techniques that were found beneficial are: maxout networks with annealed dropout rates; networks with a very large number of outputs trained on 2000 hours of data; joint modeling of partially unfolded recurrent neural networks and convolutional nets by combining the bottleneck and output layers and retraining the resulting model; and lastly, sophisticated language model rescoring with exponential and neural network LMs. These techniques result in an 8.0% word error rate on the Switchboard part of the Hub5-2000 evaluation test set which is 23% relative better than our previous best published result. George Saon, Hong-Kwang Jeff Kuo, Steven J. Rennie, Michael Picheny |
INTERSPEECH | 4 |
| 2014 | Efficient spoken term detection using confusion networksabstractIn this paper, we present a fast, vocabulary independent algorithm for spoken term detection (STD) that demonstrates a word-based index is sufficient to achieve good performance for both in-vocabulary (IV) and out-of-vocabulary (OOV) terms. Previous approaches have required that a separate index be built at the sub-word level and then expanded to allow for matching OOV terms. Such a process, while accurate, is expensive in both time and memory. In the proposed architecture, a word-level confusion network (CN) based index is used for both IV and OOV search. This is implemented using a flexible WFST framework. Comparisons on 3 Babel languages (Tagalog, Pashto and Turkish) show that CN-based indexing results in better performance compared with the lattice approach while being orders of magnitude faster and having a much smaller footprint. Lidia Mangu, Brian Kingsbury, Hagen Soltau, Hong-Kwang Jeff Kuo, Michael Picheny |
ICASSP | 5 |
| 2014 | Parallel deep neural network training for LVCSR tasks using blue gene/QabstractWhile Deep Neural Networks (DNNs) have achieved tremendous success for LVCSR tasks, training these networks is slow. To date, the most common approach to train DNNs is via stochastic gradient descent (SGD), serially on a single GPU machine. Serial training, coupled with the large number of training parameters and speech data set sizes, makes DNN training very slow for LVCSR tasks. While 2nd order, data-parallel methods have also been explored, these methods are not always faster on CPU clusters due to the large communication cost between processors. In this work, we explore using a specialized hardware/software approach, utilizing a Blue Gene/Q (BG/Q) system, which has thousands of processors and excellent interprocessor communication. We explore using the 2nd order Hessian-free (HF) algorithm for DNN training with BG/Q, for both cross-entropy and sequence training of DNNs. Results on three LVCSR tasks indicate that using HF with BG/Q offers up to an 11x speedup, as well as an improved word error rate (WER), compared to SGD on a GPU. Tara N. Sainath, I-Hsin Chung, Bhuvana Ramabhadran, Michael Picheny, John A. Gunnels, Brian Kingsbury, George Saon, Vernon Austel, Upendra V. Chaudhari |
INTERSPEECH | 4 |
| 2014 | Unfolded recurrent neural networks for speech recognitionabstractWe introduce recurrent neural networks (RNNs) for acoustic modeling which are unfolded in time for a fixed number of time steps. The proposed models are feedforward networks with the property that the unfolded layers which correspond to the recurrent layer have time-shifted inputs and tied weight matrices. Besides the temporal depth due to unfolding, hierarchical processing depth is added by means of several non-recurrent hidden layers inserted between the unfolded layers and the output layer. The training of these models: (a) has a complexity that is comparable to deep neural networks (DNNs) with the same number of layers; (b) can be done on frame-randomized minibatches; (c) can be implemented efficiently through matrix-matrix operations on GPU architectures which makes it scalable for large tasks. Experimental results on the Switchboard 300 hours English conversational telephony task show a 5% relative improvement in word error rate over state-of-the-art DNNs trained on FMLLR features with i-vector speaker adaptation and hessianfree sequence discriminative training. Index Terms: recurrent neural networks, speech recognition George Saon, Hagen Soltau, Ahmad Emami, Michael Picheny |
INTERSPEECH | 4 |
| 2014 | Parallel Deep Neural Network Training for Big Data on Blue Gene/QabstractDeep Neural Networks (DNNs) have recently been shown to significantly outperform existing machine learning techniques in several pattern recognition tasks. DNNs are the state-of-the-art models used in image recognition, object detection, classification and tracking, and speech and language processing applications. The biggest drawback to DNNs has been the enormous cost in computation and time taken to train the parameters of the networks - often a tenfold increase relative to conventional technologies. Such training time costs can be mitigated by the application of parallel computing algorithms and architectures. However, these algorithms often run into difficulties because of the cost of inter-processor communication bottlenecks. In this paper, we describe how to enable Parallel Deep Neural Network Training on the IBM Blue Gene/Q (BG/Q) computer system. Specifically, we explore DNN training using the data parallel Hessian-free 2nd order optimization algorithm. Such an algorithm is particularly well-suited to parallelization across a large set of loosely coupled processors. BG/Q, with its excellent inter-processor communication characteristics, is an ideal match for this type of algorithm. The paper discusses how issues regarding programming model and data-dependent imbalances are addressed. Results on large-scale speech tasks show that the performance on BG/Q scales linearly up to 4096 processes with no loss in accuracy. This allows us to train neural networks using billions of training examples in a few hours. I-Hsin Chung, Tara N. Sainath, Bhuvana Ramabhadran, Michael Picheny, John A. Gunnels, Vernon Austel, Upendra V. Chaudhari, Brian Kingsbury |
SC | 4 |
| 2013 | Speaker adaptation of neural network acoustic models using i-vectorsabstractWe propose to adapt deep neural network (DNN) acoustic models to a target speaker by supplying speaker identity vectors (i-vectors) as input features to the network in parallel with the regular acoustic features for ASR. For both training and test, the i-vector for a given speaker is concatenated to every frame belonging to that speaker and changes across different speakers. Experimental results on a Switchboard 300 hours corpus show that DNNs trained on speaker independent features and i-vectors achieve a 10% relative improvement in word error rate (WER) over networks trained on speaker independent features only. These networks are comparable in performance to DNNs trained on speaker-adapted features (with VTLN and FMLLR) with the advantage that only one decoding pass is needed. Furthermore, networks trained on speaker-adapted features and i-vectors achieve a 5-6% relative improvement in WER after hessian-free sequence training over networks trained on speaker-adapted features only. George Saon, Hagen Soltau, David Nahamoo, Michael Picheny |
ASRU | 4 |
| 2013 | Developing speech recognition systems for corpus indexing under the IARPA Babel programabstractAutomatic speech recognition is a core component of many applications, including keyword search. In this paper we describe experiments on acoustic modeling, language modeling, and decoding for keyword search on a Cantonese conversational telephony corpus collected as part of the IARPA Babel program. We show that acoustic modeling techniques such as the bootstrapped-and-restructured model and deep neural network acoustic model significantly outperform a state-of-the-art baseline GMM/HMM model, in terms of both recognition performance and keyword search performance, with improvements of up to 11% relative character error rate reduction and 31% relative maximum term weighted value improvement. We show that while an interpolated Model M and neural network LM improve recognition performance, they do not improve keyword search results; however, the advanced LM does reduce the size of the keyword search index. Finally, we show that a simple form of automatically adapted keyword search performs 16% better than a preindexed search system, indicating that out-of-vocabulary search is still a challenge. Jia Cui, Bhuvana Ramabhadran, Janice Kim, Brian Kingsbury, Jonathan Mamou, Lidia Mangu, Michael Picheny, Tara N. Sainath, Abhinav Sethy |
ICASSP | 8 |
| 2013 | A high-performance Cantonese keyword search systemabstractWe present a system for keyword search on Cantonese conversational telephony audio, collected for the IARPA Babel program, that achieves good performance by combining postings lists produced by diverse speech recognition systems from three different research groups. We describe the keyword search task, the data on which the work was done, four different speech recognition systems, and our approach to system combination for keyword search. We show that the combination of four systems outperforms the best single system by 7%, achieving an actual term-weighted value of 0.517. Brian Kingsbury, Jia Cui, Mark J. F. Gales, Kate M. Knill, Jonathan Mamou, Lidia Mangu, David Nolden, Michael Picheny, Bhuvana Ramabhadran, Ralf Schlüter, Abhinav Sethy, Philip C. Woodland |
ICASSP | 9 |
| 2013 | System combination and score normalization for spoken term detectionabstractSpoken content in languages of emerging importance needs to be searchable to provide access to the underlying information. In this paper, we investigate the problem of extending data fusion methodologies from Information Retrieval for Spoken Term Detection on low-resource languages in the framework of the IARPA Babel program. We describe a number of alternative methods improving keyword search performance. We apply these methods to Cantonese, a language that presents some new issues in terms of reduced resources and shorter query lengths. First, we show score normalization methodology that improves in average by 20% keyword search performance. Second, we show that properly combining the outputs of diverse ASR systems performs 14% better than the best normalized ASR system. Jonathan Mamou, Jia Cui, Mark J. F. Gales, Brian Kingsbury, Kate M. Knill, Lidia Mangu, David Nolden, Michael Picheny, Bhuvana Ramabhadran, Ralf Schlüter, Abhinav Sethy, Philip C. Woodland |
ICASSP | 9 |
| 2012 | Matching Criteria for Vocabulary-Independent SearchabstractThis paper investigates a variety of progressively more complex similarity measures for vocabulary independent search in phone based audio transcripts. English audio data is segmented and decoded to produce a sequence of phones that represent the data. These sequences are then parsed intoN-grams which are used to index the data. The audio segments define the documents to be retrieved and are thus localized in time. Search is performed by expanding text based queries into phone sequences andN-grams, followed by matching these against the index. The baseline similarity measure combines elements found in the literature and uses edit distance with a phonetic confusion matrix to determine the similarity of query and indexN-grams. Comparable performance to other approaches in the literature is achieved. Extensions to the baseline are developed using a constrained form of the similarity measure together with the ability to account for higher order confusions, namely of phone bi-grams and tri-grams. Results show improved performance across a variety of system configurations. We then generalize further and use the framework of conditional random fields (CRFs) to model confusions. Whereas others in the literature have used CRFs to model parameters of an edit distance that incorporates deletions, substitutions, and insertions, our approach focuses on using CRFs to model context dependent phone level confusions directly. The CRF is trained on parallel phonetic transcripts, which provides a general framework for modeling the errors that a recognition system may make, taking contextual effects into consideration. Results obtained on both in and out of vocabulary (OOV) search tasks improve most notably for OOV, showing 5%-6% relative improvement. Finally, we investigate the degree to which the information captured in the three approaches is complementary and show that system combination can further improve performance. Upendra V. Chaudhari, Michael Picheny |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Deep Belief Networks using discriminative features for phone recognitionabstractDeep Belief Networks (DBNs) are multi-layer generative models. They can be trained to model windows of coefficients extracted from speech and they discover multiple layers of features that capture the higher-order statistical structure of the data. These features can be used to initialize the hidden units of a feed-forward neural network that is then trained to predict the HMM state for the central frame of the window. Initializing with features that are good at generating speech makes the neural network perform much better than initializing with random weights. DBNs have already been used successfully for phone recognition with input coefficients that are MFCCs or filterbank outputs. In this paper, we demonstrate that they work even better when their inputs are speaker adaptive, discriminative features. On the standard TIMIT corpus, they give phone error rates of 19.6% using monophone HMMs and a bigram language model and 19.4% using monophone HMMs and a trigram language model. Abdel-rahman Mohamed, Tara N. Sainath, George E. Dahl, Bhuvana Ramabhadran, Geoffrey E. Hinton, Michael Picheny |
ICASSP | 6 |
| 2011 | Exemplar-Based Sparse Representation Features: From TIMIT to LVCSRabstractThe use of exemplar-based methods, such as support vector machines (SVMs), k-nearest neighbors (kNNs) and sparse representations (SRs), in speech recognition has thus far been limited. Exemplar-based techniques utilize information about individual training examples and are computationally expensive, making it particularly difficult to investigate these methods on large-vocabulary continuous speech recognition (LVCSR) tasks. While research in LVCSR provides a good testbed to tackle real-world speech recognition problems, research in this area suffers from two main drawbacks. First, the overall complexity of an LVCSR system makes error analysis quite difficult. Second, exploring new research ideas on LVCSR tasks involves training and testing state-of-the-art LVCSR systems, which can render a large turnaround time. This makes a small vocabulary task such as TIMIT more appealing. TIMIT provides a phonetically rich and hand-labeled corpus that allows easy insight into new algorithms. However, research ideas explored for small vocabulary tasks do not always provide gains on LVCSR systems. In this paper, we combine the advantages of using both small and large vocabulary tasks by taking well-established techniques used in LVCSR systems and applying them on TIMIT to establish a new baseline. We then utilize these existing LVCSR techniques in creating a novel set of exemplar-based sparse representation (SR) features. Using these existing LVCSR techniques, we achieve a phonetic error rate (PER) of 19.4% on the TIMIT task. The additional use of SR features reduce the PER to 18.6%. We then explore applying the SR features to a large vocabulary Broadcast News task, where we achieve a 0.3% absolute reduction in word error rate (WER). Tara N. Sainath, Bhuvana Ramabhadran, Michael Picheny, David Nahamoo, Dimitri Kanevsky |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2010 | Effects of automated transcription quality on non-native speakers' comprehension in real-time computer-mediated communicationabstractReal-time transcription has been shown to be valuable in facilitating non-native speakers' comprehension in real-time communication. Automated speech recognition (ASR) technology is a critical ingredient for its practical deployment. This paper presents a series of studies investigating how the quality of transcripts generated by an ASR system impacts user comprehension and subjective evaluation. Experiments are first presented comparing performance across three different transcription conditions: no transcript, a perfect transcript, and a transcript with Word Error Rate (WER) =20%. We found 20% WER was the most likely critical point for transcripts to be just acceptable and useful. Then we further examined a lower WER of 10% (a lower bound for today's state-of-the-art systems) employing the same experimental design. The results indicated that at 10% WER comprehension performance was significantly improved compared to the no-transcript condition. Finally, implications for further system development and design are discussed. Yingxin Pan, Danning Jiang, Michael Picheny |
CHI | 4 |
| 2009 | Articulatory feature detection with Support Vector Machines for integration into ASR and phone recognitionabstractWe study the use of support vector machines (SVM) for detecting the occurrence of articulatory features in speech audio data and using the information contained in the detector outputs to improve phone and speech recognition. Our expectation is that an SVM should be able to appropriately model the separation of the classes which may have complex distributions in feature space. We show that performance improves markedly when using discriminatively trained speaker dependent parameters for the SVM inputs, and compares quite well to results in the literature using other classifiers, namely artificial neural networks (ANN). Further, we show that the resulting detector outputs can be successfully integrated into a state of the art speech recognition system, with consequent performance gains. Notably, we test our system on English broadcast news data from dev04f. Upendra V. Chaudhari, Michael Picheny |
ASRU | 2 |
| 2009 | Improved vocabulary independent search with approximate match based on Conditional Random FieldsabstractWe investigate the use of Conditional Random Fields (CRF) to model confusions and account for errors in the phonetic decoding derived from Automatic Speech Recognition output. The goal is to improve the accuracy of approximate phonetic match, given query terms and an indexed database of documents, in a vocabulary independent audio search system. Audio data is ingested, segmented, decoded to produce a sequence of phones, and subsequently indexed using phone N-grams. Search is performed by expanding queries into phone sequences and matching against the index. The approximate match score is derived from a CRF, trained on parallel transcripts, which provides a general framework for modeling the errors that a recognition system may make taking contextual effects into consideration. Our approach differs from other work in the field in that we focus on using CRFs to model context dependent phone level confusions, rather than on explicitly modeling parameters of an edit distance. While, the results we obtain on both in and out of vocabulary (OOV) search tasks improve on previous work which incorporated high order phone confusions, the gains for OOV are more impressive. Upendra V. Chaudhari, Michael Picheny |
ASRU | 2 |
| 2009 | An exploration of large vocabulary tools for small vocabulary phonetic recognitionabstractWhile research in large vocabulary continuous speech recognition (LVCSR) has sparked the development of many state of the art research ideas, research in this domain suffers from two main drawbacks. First, because of the large number of parameters and poorly labeled transcriptions, gaining insight into further improvements based on error analysis is very difficult. Second, LVCSR systems often take a significantly longer time to train and test new research ideas compared to small vocabulary tasks. A small vocabulary task like TIMIT provides a phonetically rich and hand-labeled corpus and offers a good test bed to study algorithmic improvements. However, oftentimes research ideas explored for small vocabulary tasks do not always provide gains on LVCSR systems. In this paper, we address these issues by taking the standard "recipe" used in typical LVCSR systems and applying it to the TIMIT phonetic recognition corpus, which provides a standard benchmark to compare methods. We find that at the speaker-independent (SI) level, our results offer comparable performance to other SI HMM systems. By taking advantage of speaker adaptation and discriminative training techniques commonly used in LVCSR systems, we achieve an error rate of 20%, the best results reported on the TIMIT task to date, moving us closer to the human reported phonetic recognition error rate of 15%. We propose the use of this system as the baseline for future research and believe that it will serve as a good framework to explore ideas that will carry over to LVCSR systems. Tara N. Sainath, Bhuvana Ramabhadran, Michael Picheny |
ASRU | 3 |
| 2009 | Effects of real-time transcription on non-native speaker's comprehension in computer-mediated communicationsabstractWe performed an empirical study to understand the relative contributions of real-time transcription to a non-native speaker's comprehension in audio/video meetings. 48 participants were assigned to 2 presentation modes (audio, audio+video) and 3 transcription modes (no transcript, real-time transcripts in the streaming mode, transcripts with all past records) in a 3x2 factorial experimental design. The results suggest that comprehension can be significantly improved for both audio and audio+video conditions when real-time transcription is provided. Also, the participants reported positive subjective responses to the presence of real-time transcription in terms of usefulness, preference, and willingness to use such a feature if provided. No cognitive load issues were reported by the participants in the ability to synthesize across modalities. Implications for system development and design, as well as future work utilizing automation speech recognition to provide the transcripts are discussed. Yingxin Pan, Danning Jiang, Michael Picheny |
CHI | 3 |
| 2007 | Improvements in phone based audio search via constrained match with high order confusion estimatesabstractThis paper investigates an approximate similarity measure for searching in phone based audio transcripts. The baseline method combines elements found in the literature to form an approach based on a phonetic confusion matrix that is used to determine the similarity of an audio document and a query, both of which are parsed into phoneN-grams. Experimental results show comparable performance to other approaches in the literature. Extensions of the approach are developed based on a constrained form of the similarity measure that can take into consideration the system dependent errors that can occur. This is done by accounting for higher order confusions, namely of phone bi-grams and tri-grams. Results show improved performance across a variety of system configurations. Upendra V. Chaudhari, Michael Picheny |
ASRU | 2 |
| 2007 | Lattice-based Viterbi decoding techniques for speech translationabstractWe describe a cardinal-synchronous Viterbi decoder for statistical phrase-based machine translation which can operate on general ASR lattices (as opposed to confusion networks). The decoder implements constrained source reordering on the input lattice and makes use of an outbound distortion model to score the possible reorderings. The phrase table, representing the decoding search space, is encoded as a weighted finite state acceptor which is determined and minimized. At a high level, the search proceeds by performing simultaneous transitions in two pairs of automata: (input lattice, phrase table FSM) and (phrase table FSM, target language model). An alternative decoding strategy that we explore is to break the search into two independent subproblems: first, we perform monotone lattice decoding and find the best foreign path through the ASR lattice and then, we decode this path with reordering using standard sentence-based SMT. We report experimental results on several testsets of a large scale Arabic-to-English speech translation task in the context of the global autonomous language exploitation (or GALE) DARPA project. The results indicate that, for monotone search, lattice-based decoding outperforms 1-best decoding whereas for search with reordering, only the second decoding strategy was found to be superior to 1-best decoding. In both cases, the improvements hold only for shallow lattices. George Saon, Michael Picheny |
ASRU | 2 |
| 2007 | Voice-Melody Transcription Under a Speech Recognition FrameworkabstractThis paper presents a robust voice-melody transcription system using a speech recognition framework. While many previous voice-melody transcription systems have utilized non-statistical approaches, statistical recognition technology can potentially achieve more robust results. A cepstrum-based acoustic model is employed to avoid the hard-decisions that have to be made when using explicit voiced-unvoiced segmentation and pitch extraction, and a key-independent 4-gram language model is employed to capture prior probabilities of different melodic sequences. Evaluations are done from the perspective of both note recognition error rate and query-by-humming end-to-end performance. The results are compared with three other voice-melody transcription systems. Experiments have shown that our system is state-of-the-art: it is much more robust than other systems on data containing noise, and close to the best of all the systems on the clean data set. Danning Jiang, Michael Picheny |
ICASSP (4) | 2 |
| 2006 | Towards Pooled-Speaker Concatenative Text-to-SpeechabstractIn this paper we explore the merging of data from various speakers in building a concatenative text-to-speech system. First, we investigate the pooling of data from multiple speakers for building statistical models to predict pitch and duration, and present listening test results which show that the expressiveness of our TTS system is improved using these techniques. Additionally, we describe an experiment in which we merged databases from several speakers to form an enlarged database from which our concatenative text-to-speech system draws segments. We present listening test results which show that pooling data from several speakers yields higher quality synthetic speech in general domains than restricting ourselves to the data from just one speaker in our repertoire Ellen Eide, Michael Picheny |
ICASSP (1) | 2 |
| 2006 | Concept-based speech-to-speech translation using maximum entropy models for statistical natural concept generationabstractThe IBM Multilingual Automatic Speech-To-Speech TranslatOR (MASTOR) system is a research prototype developed for the Defense Advanced Research Projects Agency (DARPA) Babylon/CAST speech-to-speech machine translation program. The system consists of cascaded components of large-vocabulary conversational spontaneous speech recognition, statistical machine translation, and concatenative text-to-speech synthesis. To achieve highly accurate and robust conversational spoken language translation, a unique concept-based speech-to-speech translation approach is proposed that performs the translation by first understanding the meaning of the automatically recognized text. A decision-tree based statistical natural language understanding algorithm extracts the semantic information from the input sentences, while a natural language generation (NLG) algorithm predicts the translated text via maximum-entropy-based statistical models. One critical component in our statistical NLG approach is natural concept generation (NCG). The goal of NCG is not only to generate the correct set of concepts in the target language, but also to produce them in an appropriate order. To improve maximum-entropy-based concept generation, a set of new approaches is proposed. One approach improves concept sequence generation in the target language via forward–backward modeling, which selects the hypothesis with the highest combined conditional probability based on both the forward and backward generation models. This paradigm allows the exploration of both the left and right context information in the source and target languages during concept generation. Another approach selects bilingual features that enable maximum-entropy-based model training on the preannotated parallel corpora. This feature is augmented with word-level information in order to achieve higher NCG accuracy while minimizing the total number of distinct concepts and, hence, greatly reducing the concept annotation and natural language understanding effort. These features are further expanded to multiple sets to enhance model robustness. Finally, a confidence threshold is introduced to alleviate data sparseness problems in our training corpora. Experiments show a dramatic concept generation error rate reduction of more than 40% in our speech translation corpus within limited domains. Significant improvements of both word error rate and BiLingual Evaluation Understudy (BLEU) score are also achieved in our experiments on speech-to-speech translation. Liang Gu, Fu-Hua Liu, Michael Picheny |
IEEE Trans. Speech Audio Process. | 4 |
| 2006 | The IBM expressive text-to-speech synthesis system for American EnglishabstractExpressive text-to-speech (TTS) synthesis should contribute to the pleasantness, intelligibility, and speed of speech-based human-machine interactions which use TTS. We describe a TTS engine which can be directed, via text markup, to use a variety of expressive styles, here, questioning, contrastive emphasis, and conveying good and bad news. Differences in these styles lead us to investigate two approaches for expressive TTS, a "corpus-driven" and a "prosodic-phonology" approach. Each speaker records 11 h (excluding silences) of "neutral" sentences. In the corpus-driven approach, the speaker also records 1-h corpora in each expressive style; these segments are tagged by style for use during search, and decision trees for determining f0contours and timing are trained separately for each of the neutral and expressive corpora. In the prosodic-phonology approach, rules translating certain expressive markup elements to tones and break indices (ToBI) are manually determined, and the ToBI elements are used in single f0and duration trees for all expressions. Tests show that listeners identify synthesis in particular styles ranging from 70% correctly for "conveying bad news" to 85% for "yes-no questions". Further improvements are demonstrated through the use of speaker-pooled f0and duration models John F. Pitrelli, Raimo Bakis, Ellen Eide, Raul Fernandez, Wael Hamza, Michael Picheny |
IEEE Trans. Speech Audio Process. | 6 |
| 2005 | Toward multiple-language TTS: experiments in English and MandarinabstractText-to-speech systems have dramatically improved in recent years through the use of corpus-based concatenative approaches, and we are beginning to see an interest in endowing them with the ability to handle more than the native language for which they have been developed. In this paper we present ongoing work at IBM in text-to-speech systems that can produce high-quality synthesis in more than one language. We illustrate the discussion with a case study in which two systems, originally developed to support English and Mandarin respectively, have been extended to support each other’s languages. We describe the challenges faced when adapting one system to a different target language, propose adaptation solutions, and present the results of perceptual tests carried out to evaluate how the approaches compare with the performance of the native systems. Raul Fernandez, Wei Zhang 0022, Ellen Eide, Raimo Bakis, Wael Hamza, Michael Picheny, John F. Pitrelli, Yong Qing, Zhiwei Shuang, Li Qin Shen |
INTERSPEECH | 7 |
| 2005 | Using semantic analysis to improve speech recognition performance
Hakan Erdogan, Ruhi Sarikaya, Stanley F. Chen, Michael Picheny |
Comput. Speech Lang. | 5 |
| 2005 | Semantic confidence measurement for spoken dialog systemsabstractThis paper proposes two methods to incorporate semantic information into word and concept level confidence measurement. The first method uses tag and extension probabilities obtained from a statistical classer and parser. The second method uses a maximum entropy based semantic structured language model to assign probabilities to each word. Incorporation of semantic features into a lattice posterior probability based confidence measure provides significant improvements compared to posterior probability when used together in an air travel reservation task. At 5% False Alarm (FA) rate relative improvements of 28% and 61% in Correct Acceptance (CA) rate are achieved for word level and concept level confidence measurements, respectively. Ruhi Sarikaya, Michael Picheny, Hakan Erdogan |
IEEE Trans. Speech Audio Process. | 3 |
| 2004 | The IBM expressive speech synthesis systemabstractThis paper introduces the IBM Expressive Speech Synthesis system. We describe recent work in improving the quality of our baseline text-to-speech system as well as extending our capabilities to generate expressive synthetic speech. We present results showing improved base quality, especially for sentences drawn from a limited domain. We also demonstrate our ability to convey good news and bad news, produce contrastive emphasis, and ask a question appropriately. In order to facilitate access to the expressive capabilities, we use some of our proposed extensions to the Speech Synthesis Markup Language (SSML). 1. Wael Hamza, Ellen Eide, Raimo Bakis, Michael Picheny, John F. Pitrelli |
INTERSPEECH | 4 |
| 2004 | Automatic recognition of spontaneous speech for access to multilingual oral history archivesabstractMuch is known about the design of automated systems to search broadcast news, but it has only recently become possible to apply similar techniques to large collections of spontaneous speech. This paper presents initial results from experiments with speech recognition, topic segmentation, topic categorization, and named entity detection using a large collection of recorded oral histories. The work leverages a massive manual annotation effort on 10 000 h of spontaneous speech to evaluate the degree to which automatic speech recognition (ASR)-based segmentation and categorization techniques can be adapted to approximate decisions made by human annotators. ASR word error rates near 40% were achieved for both English and Czech for heavily accented, emotional and elderly spontaneous speech based on 65-84 h of transcribed speech. Topical segmentation based on shifts in the recognized English vocabulary resulted in 80% agreement with manually annotated boundary positions at a 0.35 false alarm rate. Categorization was considerably more challenging, with a nearest-neighbor technique yielding F=0.3. This is less than half the value obtained by the same technique on a standard newswire categorization benchmark, but replication on human-transcribed interviews showed that ASR errors explain little of that difference. The paper concludes with a description of how these capabilities could be used together to search large collections of recorded oral histories. William J. Byrne, David S. Doermann, Martin Franz, Samuel Gustman, Jan Hajic 0001, Douglas W. Oard, Michael Picheny, Josef Psutka, Bhuvana Ramabhadran, Dagobert Soergel, Todd Ward, Wei-Jing Zhu |
IEEE Trans. Speech Audio Process. | 7 |
| 2003 | Recent improvements to the IBM trainable speech synthesis systemabstractIn this paper we describe the current status of the trainable text-to-speech system at IBM. Recent algorithmic and database changes to the system have led to significant gains in the output quality. On the algorithms side, we have introduced statistical models for predicting pitch and duration targets which replace the rule-based target generation previously employed. Additionally, we have changed the cost function and the search strategy, introduced a post-search pitch smoothing algorithm, and improved our method of preselection. Through the combined data and algorithmic contributions, we have been able to significantly improve (p < 0.0001) the mean opinion score (MOS) of our female voice, from 3.68 to 4.85 when heard over loudspeakers and to 5.42 when heard over the telephone (seven point scale). Ellen Eide, Andrew Aaron, Raimo Bakis, Paul S. Cohen, Robert E. Donovan, Wael Hamza, T. Mathes, Michael Picheny, M. Polkosky, Mahesh Viswanathan 0002 |
ICASSP (1) | 8 |
| 2003 | Use of statistical N-gram models in natural language generation for machine translationabstractVarious language modeling issues in a speech-to-speech translation system are described in this paper. First, the language models for the speech recognizer need to be adapted to the specific domain to improve the recognition performance for in-domain utterances, while keeping the domain coverage as broad as possible. Second, when a maximum entropy based statistical natural language generation model is used to generate target language sentence as the translation output, serious inflection and synonym issues arise, because the compromised solution is used in semantic representation to avoid the data sparseness problem. We use N-gram models as a postprocessing step to enhance the generation performance. When an interpolated language model is applied to a Chinese-to-English translation task, the translation performance, measured by an objective metric of BLEU, improves substantially to 0.514 from 0.318 when we use the correct transcription as input. Similarly, the BLEU score is improved to 0.300 from 0.194 for the same task when the input is speech data. Fu-Hua Liu, Liang Gu, Michael Picheny |
ICASSP (1) | 4 |
| 2003 | Towards automatic transcription of large spoken archives - English ASR for the MALACH projectabstractDigital archives have emerged as the pre-eminent method for capturing the human experience. Before such archives can be used efficiently, their contents must be described. The NSF-funded MALACH project aims to provide improved access to large spoken archives by advancing the state-of-the-art in automated speech recognition (ASR), Information Retrieval (IR) and related technologies [1,2] for multiple languages. This paper describes the ASR research for the English speech in the MALACH corpus. The MALACH corpus consists of unconstrained, natural speech filled with disfluencies, heavy accents, age-related coarticulation, uncued speaker and language switching, and emotional speech collected in the form of interviews from over 52000 speakers in 32 languages. In this paper, we describe this new testbed for developing speech recognition algorithms and report on the performance of well-known techniques for building better acoustic models for the speaking styles seen in this corpus. The best English ASR system to date has a word error rate of 43.8% on this corpus. Bhuvana Ramabhadran, Jing Huang 0019, Michael Picheny |
ICASSP (1) | 3 |
| 2003 | Word level confidence measurement using semantic featuresabstractThis paper proposes two principled methods to incorporate semantic information into word level confidence measurement. The first technique uses tag and arc probabilities obtained from a statistical classer and parser tree. The second technique uses a maximum entropy based semantic structured language model to use semantic structure of a sentence to assign semantic probabilities to each word. Semantic features provide significant improvements over a posterior probability based confidence measure when used together in an air travel reservation task. Ruhi Sarikaya, Michael Picheny |
ICASSP (1) | 3 |
| 2003 | Automated transcription and topic segmentation of large spoken archivesabstractDigital archives have emerged as the pre-eminent method for capturing the human experience. Before such archives can be used efficiently, their contents must be described. The scale of such archives along with the associated content mark up cost make it impractical to provide access via purely manual means, but automatic technologies for search in spoken materials still have relatively limited capabilities. The NSF-funded MALACH project will use the world’s largest digital archive of video oral histories, collected by the Survivors of the Shoah Visual History Foundation (VHF) to make a quantum leap in the ability to access such archives by advancing the state-of-the-art in Automated Speech Recognition (ASR), Natural Language Processing (NLP) and related technologies [1, 2]. This corpus consists of over 115,000 hours of unconstrained, natural speech from 52,000 speakers in 32 different languages, filled with disfluencies, heavy accents, age-related coarticulations, and un-cued speaker and language switching. This paper discusses some of the ASR and NLP tools and technologies that we have been building for the English speech in the MALACH corpus. We also discuss this new test bed while emphasizing the unique characteristics of this corpus. 1. Martin Franz, Bhuvana Ramabhadran, Todd Ward, Michael Picheny |
INTERSPEECH | 4 |
| 2003 | Improving statistical natural concept generation in interlingua-based speech-to-speech translation
Liang Gu, Michael Picheny |
INTERSPEECH | 3 |
| 2003 | Toward domain-independent conversational speech recognitionabstractWe describe a multi-domain, conversational test set developed for IBM’s Superhuman speech recognition project and our 2002 benchmark system for this task. Through the use of multipass decoding, unsupervised adaptation and combination of hypotheses from systems using diverse feature sets and acoustic models, we achieve a word error rate of 32.0 % on data drawn from voicemail messages, two-person conversations and multiple-person meetings. 1. Brian Kingsbury, Lidia Mangu, George Saon, Geoffrey Zweig, Scott Axelrod, Vaibhava Goel, Karthik Visweswariah, Michael Picheny |
INTERSPEECH | 8 |
| 2003 | Noise robustness in speech to speech translationabstractThis paper describes various noise robustness issues in a speech-to-speech translation system. We present quantitative measures for noise robustness in the context of speech recognition accuracy and speech-to-speech translation performance. To enhance noise immunity, we explore two approaches to improve the overall speech-to-speech translation performance. First, a multi-style training technique is used to tackle the issue of environmental degradation at the acoustic model level. Second, a pre-processing technique, CDCN, is exploited to compensate for the acoustic distortion at the signal level. Further improvement can be obtained by combining both schemes. In addition to recognition accuracy for speech recognition, this paper studies and examines how closely speech recognition accuracy is related the overall speech-to-speech recognition. When we apply the proposed schemes to an English-to-Chinese translation task, the word error rate for our speech recognition subsystem is substantially reduced by 28% relative, to 13.2% from 18.9% for test data of 15dB SNR. The corresponding BLEU score improves to 0.478 from 0.43 for the overall speech-to-speech translation. Similar improvements are also observed for a lower SNR condition. Full Paper Fu-Hua Liu, Liang Gu, Michael Picheny |
INTERSPEECH | 4 |
| 2002 | Turn-Based Language Modeling for spoken dialog systemsabstractIn this paper I we propose a turn-based language modeling (TurnLM) technique for spoken dialog systems. This technique utilizes the time dependent nature of a dialog aimed at accomplishing a task. As opposed to the dialog state based language modeling techniques which depend on the information in the system prompt, TurnLM does not require any information from the dialog manager. As such, TurnLM can be used not only for human-machine dialogs but also human-human dialogs. We report performance improvement compared to the baseline system on the IBM DARPA Communicator spoken dialog system. Experimental results also suggest that TurnLM is a viable alternative to dialog state based language modeling technique. Ruhi Sarikaya, Hakan Erdogan, Michael Picheny |
ICASSP | 4 |
| 2002 | Semantic structured language modelsabstractIn this study, we propose two novel semantic language modeling techniques for spoken dialog systems. These methods are called semantic concept based language modeling and semantic structured language modeling. In the concept based language modeling, we propose to use long span semantic units to model meaning sequences in spoken utterances. In the latter technique, we use statistical semantic parsers to extract information from a sentence. This information is then utilized in a maximum entropy based language model. The language models are trained and evaluated in the air travel reservation domain. We obtain improvement over a sophisticated class based N-gram language model both in terms of recognition accuracy and perplexity. Interpolation of the proposed techniques with the class-based N-gram LM provides additional improvement. 1. Hakan Erdogan, Ruhi Sarikaya, Michael Picheny |
INTERSPEECH | 4 |
| 2002 | Statistical natural language generation for speech-to-speech machine translation systems
Bowen Zhou 0006, Jeffrey S. Sorensen, Zijian Diao, Michael Picheny |
INTERSPEECH | 5 |
| 2002 | MARS: A Statistical Semantic Parsing and Generation-Based Multilingual Automatic tRanslation System
Bowen Zhou 0006, Zijian Diao, Jeffrey S. Sorensen, Michael Picheny |
Mach. Transl. | 5 |
| 2001 | Speech recognition for DARPA CommunicatorabstractWe report the results of investigations in acoustic modeling, language modeling and decoding techniques, for the DARPA Communicator, a speaker-independent, telephone-based dialog system. By a combination of methods, including enlarging the acoustic model, augmenting the recognizer vocabulary, conditioning the language model upon the dialog state, and applying a post-processing decoding method, we lowered the overall word error rate from 21.9% to 15.0%, a gain of 6.9% absolute and 31.5% relative. Andrew Aaron, Scott Saobing Chen, Paul S. Cohen, Satya Dharanipragada, Ellen Eide, Martin Franz, Jean-Michel LeRoux, X. Luo, Benoît Maison, Lidia Mangu, T. Mathes, Miroslav Novak, Peder A. Olsen, Michael Picheny, Harry Printz, Bhuvana Ramabhadran, Andrej Sakrajda, George Saon, Borivoj Tydlitát, Karthik Visweswariah, D. Yuk |
ICASSP | 14 |
| 2001 | Rapid adaptation using penalized-likelihood methodsabstractWe introduce rapid adaptation techniques that extend and improve two successful methods previously introduced, cluster weighting (CW) and MAPLR. First, we introduce an adaptation scheme called CWB which extends the cluster weighting adaptation method by including a bias term and a reference speaker model. CWB is shown to improve the adaptation performance as compared to CW. Second, we introduce an extension of cluster weighting that uses penalized-likelihood objective functions to stabilize the estimation and provide soft constraints. Third, we propose a variant of MAPLR adaptation that uses prior speaker information. Previously, prior distributions of transforms in MAPLR were obtained using the same adaptation data, speaker independent HMM means or by some heuristics. We propose to use the prior information of speaker variability to obtain the priors, by using CW or CWB weights. Penalized-likelihood or Bayesian theory serves as a tool to combine transformation based and prior speaker information based adaptation methods resulting in effective rapid adaptation techniques. The techniques are shown to outperform full, block diagonal and diagonal MLLR as well as some other recently proposed methods for rapid adaptation. Hakan Erdogan, Michael Picheny |
ICASSP | 3 |
| 2001 | Innovative approaches for large vocabulary name recognitionabstractAutomatic name dialing is a practical and interesting application of speech recognition on telephony systems. The IBM name recognition system is a large vocabulary, speaker independent system currently in use for reaching IBM employees in the United States. We present some innovative algorithms that improve name recognition accuracy. Unlike transcription tasks, such as the Switchboard task, recognition of names poses a variety of different problems. Several of these problems arise from the fact that foreign names are hard to pronounce for speakers who are not familiar with the names and that there are no standardized methods for pronouncing proper names. Noise robustness is another very important factor as these calls are typically made in noisy environments, such as from a car, cafeteria, airport, etc. and over different kinds of cellular and land-line telephone channels. We have performed a systematic analysis of the speech recognition errors and tackled the issues separately with techniques ranging from weighted speaker clustering, massive adaptation, rapid and unsupervised adaptation methods to pronunciation modeling methods. We find that the decoding accuracy can be improved significantly (28% relative) in this manner. Bhuvana Ramabhadran, C. Julian Chen, Hakan Erdogan, Michael Picheny |
ICASSP | 5 |
| 2001 | Recent advances in speech recognition system for IBM DARPA communicatorabstractIn this paper, we present methods to improve speech recognition performance of the IBM DARPA Communicator system. Our efforts for acoustic modeling include training a domain specific yet broad acoustic model, speaker clustering and speaker adaptation using feature space transforms. For language modeling, we achieved improvements by using compound words, carefully designed LM classes and adjusting the within class probabilities, using NLU state information to enhance the language model and building a language model with embedded grammar objects. Our efforts produced a relative error rate reduction of 34.6 % on the test set that consists of 1173 utterances that IBM received during the NIST evaluation of the DARPA Communicator systems in June 2000. We also tested our decoding on the data from some other sites to further demonstrate the robustness of the system improvements. 1. Hakan Erdogan, Vaibhava Goel, Michael Picheny |
INTERSPEECH | 5 |
| 2000 | Rapid likelihood calculation of subspace clustered Gaussian componentsabstractIn speech recognition systems, computing the likelihoods of the acoustic models is an intensive task. One approach to reduce this cost is to use subspace distributed clustering HMM. Here individual Gaussian components are stored as indices to, and their likelihoods computed from, a set of subspace Gaussian components. This paper examines a scheme for reducing the computational cost of the likelihood calculation when such an HMM system is used. The proposed method identifies and stores frequently occurring partial sums called meta-atom elements and thus avoids computing them repeatedly. The resultant savings in the number of additions is 50% when all Gaussian components are computed or 20% when a Gaussian selection scheme is used. Anuradha Aiyer, Mark J. F. Gales, Michael Picheny |
ICASSP | 3 |
| 2000 | Maximal rank likelihood as an optimization function for speech recognition
Michael Picheny |
INTERSPEECH | 3 |
| 2000 | Speed improvement of the tree-based time asynchronous search
Miroslav Novak, Michael Picheny |
INTERSPEECH | 2 |
| 2000 | Heredity and environment in speech recognition: the role of a priori information vs. data
Michael Picheny |
INTERSPEECH | 1 |
| 2000 | Dynamic selection of feature spaces for robust speech recognition
Bhuvana Ramabhadran, Michael Picheny |
INTERSPEECH | 3 |
| 2000 | Impact of bucketing on performance of linearly interpolated language models
Karthik Visweswariah, Harry Printz, Michael Picheny |
INTERSPEECH | 3 |
| 1999 | HMM training based on quality measurementabstractTwo discriminant measures for HMM states to improve the effectiveness on HMM training are presented. In HMM based speech recognition, the context-dependent states are usually modeled by Gaussian mixture distributions. In general, the number of Gaussian mixtures for each state is fixed or proportional to the amount of training data. From our study, some of the states are "non-aggressive" compared to others, and a higher acoustic resolution is required for them. Two methods are presented in this paper to determine those non-aggressive states. The first approach uses the recognition accuracy of the states and the second method is based on a rank distribution of states. Baseline systems, trained by a fixed number of Gaussian mixtures for each state, having 33 K and 120 K Gaussians, yield 14.57% and 13.04% word error rates, respectively. Using our approach, a 38 K Gaussian system was constructed that reduces the error rate to 13.95%. The average ranks of non-aggressive states in rank lists of testing data were also seen to dramatic improve compared to the baseline systems. Ea-Ee Jan, Mukund Padmanabhan, Michael Picheny |
ICASSP | 4 |
| 1999 | Speed improvement of the time-asynchronous acoustic fast matchabstractThis paper describes an algorithm for improvement of the speed of a time-asynchronous fast match, which is a part of a stack-search based recognition system. This fast match uses a phonetic tree to represent the entire vocabulary of the recognizer. Evaluation of the tree (in a depthrst manner), can be done much more e ciently using the fact that under certain conditions, the results of branch evaluations can be used to approximate the scores of other branches of the tree. Miroslav Novak, Michael Picheny |
EUROSPEECH | 2 |
| 1999 | Enhanced likelihood computation using regressionabstractIn this paper, a new auditory spectrum based speech feature is proposed using sinusoidal representation and auditory model. The feature is optimized using the properties of auditory perception and masking. After quantizing and encoding the optimized feature parameters, a new speech-coding algorithm with average bit-rate of 3.25kbps is developed. The experimental results show that the synthetic speech retains most of the intelligibility and clearness of articulation of the original speech. Compared with the conventional algorithms, no voiced/unvoiced decision and pitch estimation are needed, complexity of the algorithm is much reduced, robustness and adaptation are both raised. The algorithm makes it possible to be realized with single DSP chip. Peter V. de Souza, Bhuvana Ramabhadran, Michael Picheny |
EUROSPEECH | 4 |
| 1998 | Improvements in children's speech recognition performanceabstractThere are several reasons why conventional speech recognition systems modeled on adult data fail to perform satisfactorily on children's speech input. For instance, children's vocal characteristics differ significantly from those of adults. In addition, their choices of vocabulary and sentence construction modalities usually do not conform to adult patterns. We describe comparative studies demonstrating the performance gain realized by adopting to children's acoustic and language model data to construct a children's speech recognition system. Subrata K. Das, Don Nix, Michael Picheny |
ICASSP | 3 |
| 1998 | Telephone band LVCSR for hearing-impaired usersabstractLarge vocabulary automatic speech recognition might assist hearing impaired telephone users by displaying a transcription of the incoming side of the conversation, but the system would have to achieve su cient accuracy on conversationalstyle, telephone-bandwidth speech. We describe our development work toward such a system. This work comprised three phases: Experiments with clean data ltered to 200-3500Hz, experiments with real telephone data, and language model development. In the rst phase, the speaker independent error rate was reduced from 25% to 12% by using MLLT, increasing the number of cepstral components from 9 to 13, and increasing the number of Gaussians from 30,000 to 120,000. The resulting system, however, performed less well on actual telephony, producing an error rate of 28.4%. By additional adaptation and the use of an LDA and CDCN combination, the error rate was reduced to 19.1%. Speaker adaptation reduces the error rate to 10.96%. These results were obtained with read speech. To explore the language-model requirements in a more realistic situation, we collected some conversational speech with an arrangement in which one participant could not hear the conversation but only saw recognizer output on a screen. We found that a mixture of language models, one derived from the Switchboard corpus and the other from prepared texts, resulted in approximately 10% fewer errors than either model alone. Ea-Ee Jan, Raimo Bakis, Fu-Hua Liu, Michael Picheny |
ICSLP | 4 |
| 1998 | A new confidence measure based on rank-ordering subphone scoresabstractThis paper describes a new con dence measure based on rank-ordering subphone likelihood scores. The approach consists of three major steps. The rst step is standard decoding which hypothesizes a word string given an utterance. Then forced Viterbi alignment is made at the second step. (This step is of course not needed for a Viterbi decoder.) From the aligned sentence, the third step computes likelihood scores of the hypothesized subphone and all other competing subphones to generate a list of subphones in the descending order of the likelihood scores. A rank is assigned to the hypothesized subphone according to its positioning in the list. The rank value is then merged to obtain the corresponding rank at the phone level. After the merge, selective weighting is applied such that contribution of phones having large acoustic variations is de-emphasized. Additional upper-bound limiting is also made to guarantee a rank computation not to be contaminated by a very bad segment. The new con dence measure has been favorably evaluated on word conrmation/rejection experiments with a small vocabulary of many confusable and short words. More speci cally, experimental results show that the new approach outperforms other measures such as whole-word scores by reducing the equal error rate from 32% to 20%. Qiguang Lin, Subrata K. Das, David M. Lubensky, Michael Picheny |
ICSLP | 4 |
| 1998 | On variable sampling frequencies in speech recognition
Fu-Hua Liu, Michael Picheny |
ICSLP | 2 |
| 1998 | Speaker clustering and transformation for speaker adaptation in speech recognition systemsabstractA speaker adaptation strategy is described that is based on finding a subset of speakers, from the training set, who are acoustically close to the test speaker, and using only the data from these speakers (rather than the complete training corpus) to reestimate the system parameters. Further, a linear transformation is computed for every one of the selected training speakers to better map the training speaker's data to the test speaker's acoustic space. Finally, the system parameters (Gaussian means) are reestimated specifically for the test speaker using the transformed data from the selected training speakers. Experiments showed that this scheme is capable of providing an 18% relative improvement in the error rate on a large-vocabulary task with the use of as little as three sentences of adaptation data. Mukund Padmanabhan, Lalit R. Bahl, David Nahamoo, Michael Picheny |
IEEE Trans. Speech Audio Process. | 4 |
| 1997 | New methods in continuous Mandarin speech recognition
C. Julian Chen, Ramesh A. Gopinath, Michael D. Monkowski, Michael Picheny, Katherine Shen |
EUROSPEECH | 4 |
| 1997 | Speaker adaptation based on pre-clustering training speakers
Mukund Padmanabhan, Michael Picheny |
EUROSPEECH | 3 |
| 1997 | Key-phrase spotting using an integrated language model of n-grams and finite-state grammar
Qiguang Lin, David M. Lubensky, Michael Picheny, P. Srinivasa Rao |
EUROSPEECH | 3 |
| 1996 | Speech recognition on Mandarin Call Home: a large-vocabulary, conversational, and telephone speech corpusabstractWe describe IBM's most recent efforts for speech recognition on a conversational-speech database, the Mandarin Call Home corpus. While it is similar to the well-known Switchboard corpus, the Call Home task addresses several major challenges in the domain of spoken language systems, including spontaneous dialogue with no pre-specified topics, limited-bandwidth telephone signal, and recognition of other languages than English. We particularly describe the methodology used in Mandarin Call Home corpus to address language-specific issues. We also examine and compare our results with those of the English Switchboard corpus. Preliminary experiments show that a 58.7% character error rate can be achieved in the context of April 95 Mandarin Call Home data set. The experimental results are comparable to those of the state-of-the-art IBM Switchboard system with similar amount of training data. Fu-Hua Liu, Michael Picheny, Patibandla Srinivasa, Michael D. Monkowski, C. Julian Chen |
ICASSP | 2 |
| 1996 | Speaker clustering and transformation for speaker adaptation in large-vocabulary speech recognition systemsabstractA speaker adaptation strategy is described that is based on finding a subset of speakers, from the training set, who are acoustically close to the test speaker, and using only the data from these speakers (rather than the complete training corpus) to re-estimate the system parameters. Further, a linear transformation is computed for every one of the selected training speakers to better map the training speaker's data to the test speaker's acoustic space. Finally, the system parameters (Gaussian means) are re-estimated specifically for the test speaker using the transformed data from the selected training speakers. Experiments showed that this scheme is capable of reducing the error rate by 10-15% with the use of as little as 3 sentences of adaptation data. Mukund Padmanabhan, Lalit R. Bahl, David Nahamoo, Michael Picheny |
ICASSP | 4 |
| 1995 | Performance of the IBM large vocabulary continuous speech recognition system on the ARPA Wall Street Journal taskabstractIn this paper we discuss various experimental results using our continuous speech recognition system on the Wall Street Journal task. Experiments with different feature extraction methods, varying amounts and type of training data, and different vocabulary sizes are reported. Lalit R. Bahl, S. Balakrishnan-Aiyer, Jerome R. Bellegarda, Martin Franz, Ponani S. Gopalakrishnan, David Nahamoo, Miroslav Novak, Mukund Padmanabhan, Michael Picheny, Salim Roukos |
ICASSP | 9 |
| 1995 | Experiments using data augmentation for speaker adaptationabstractSpeaker adaptation typically involves customizing some existing (reference) models in order to account for the characteristics of a new speaker. This work considers the slightly different paradigm of customizing some reference data for the purpose of populating the new speaker's space, and then using the resulting (augmented) data to derive the customized models. The data augmentation technique is based on the metamorphic algorithm first proposed in Bellegarda et al. [1992], assuming that a relatively modest amount of data (100 sentences) is available from each new speaker. This contraint requires that reference speakers be selected with some care. The performance of this method is illustrated on a portion of the Wall Street Journal task. Jerome R. Bellegarda, Peter V. de Souza, David Nahamoo, Mukund Padmanabhan, Michael Picheny, Lalit R. Bahl |
ICASSP | 5 |
| 1995 | Context dependent phonetic duration models for decoding conversational speechabstractConversational speech provides a particularly difficult task for speech recognition. It provides much more variability than either dictation, read speech, or isolated commands. Phonetic context was used to predict the durations of phones using a decision tree. These predictions were used to calculate context dependent HMM transition probabilities for these phone models, which were used to decode telephone conversations from the SwitchBoard corpus. We observed that the duration models do not appreciably improve the word error rate; that more can be gained by modeling phone durations within words than by adjusting for local average speaking rates; and conclude that local or global variations in speaking rate are not major contributors to the observed high error rates for SwitchBoard. Michael D. Monkowski, Michael Picheny, P. Srinivasa Rao |
ICASSP | 2 |
| 1994 | Robust methods for using context-dependent features and models in a continuous speech recognizerabstractIn this paper we describe the method we use to derive acoustic features that reflect some of the dynamics of frame-based parameter vectors. Models for such observations must be context dependent. Such models were outlined in an earlier paper. Here we describe a method for using these models in a recognition system. The method is more robust than using continuous parameter models in recognition. At the same time it does not suffer from the possible information loss in vector quantization based systems.> Lalit R. Bahl, Peter V. de Souza, Ponani S. Gopalakrishnan, David Nahamoo, Michael Picheny |
ICASSP (1) | 5 |
| 1994 | Adaptation techniques for ambience and microphone compensation in the IBM Tangora speech recognition systemabstractConventional speech recognition systems such as the IBM Tangora tend to be adversely influenced by external factors such as background noise and microphone characteristics. This paper discusses some adaptation techniques to counteract such influences. We also consider ways to reduce computation so that the methods may be implemented for real time Tangora operation. The adaptation strategy utilized two sets of VQ codebooks derived from the training data, one representing the ambience and the other characterizing the speech domain. A speech versus ambience decision was made by examining several factors, such as a comparison of the overall energy in a frame with a percentile point of a running energy histogram. We include results to demonstrate the effectiveness of our approach.> Subrata K. Das, Arthur Nádas, David Nahamoo, Michael Picheny |
ICASSP (1) | 4 |
| 1994 | A channel-bank-based phone detection strategyabstractThis paper presents a channel-bank based phone detection algorithm, that can be used in greatly cut down the search space in the process of mapping a set of acoustic features to a phone sequence. The algorithm involves a low cost preprocessing, and relies on using local information (in time) about the acoustic features, to predict, whether or not a phone can exist over a certain time interval; hence, it may be thought of as producing a short-list of possible phones at all times, and it is only necessary to search in this short list, rather than in the entire phone alphabet, for the correct phone. The technique was applied in the acoustic fast match of the IBM speech recognition system, and resulted in a significant reduction in the execution time of the fast match.> Ponani S. Gopalakrishnan, David Nahamoo, Mukund Padmanabhan, Michael Picheny |
ICASSP (2) | 4 |
| 1994 | The metamorphic algorithm: a speaker mapping approach to data augmentationabstractLarge vocabulary speaker-dependent speech recognition systems adjust to the acoustic peculiarities of each new speaker based on some enrolment data provided by this speaker. As the amount of data required increases with the sophistication of the underlying acoustic models, the enrolment may get lengthy. To streamline it, it is therefore desirable to make use of previously acquired speech data. The authors describe a data augmentation strategy based on a piecewise linear mapping between the feature space of a new speaker and that of a reference speaker. This speaker-normalizing mapping is used to transform the previously acquired data of the reference speaker onto the space of the new speaker. The performance of the resulting procedure, dubbed the metamorphic algorithm, is illustrated on an isolated utterance speech recognition task with a vocabulary of 20000 words. Results show that the metamorphic algorithm can substantially reduce the word error rate when only a limited amount of enrolment data is available. Alternatively, it leads to a level of performance comparable to that obtained when a much greater amount of enrolment data is required from the new speaker. In addition, it can also be used for tracking spectral evolution over time, thus providing a possible means for robust speaker self-adaptation.> Jerome R. Bellegarda, Peter V. de Souza, Arthur Nádas, David Nahamoo, Michael Picheny, Lalit R. Bahl |
IEEE Trans. Speech Audio Process. | 5 |
| 1994 | Speaker adaptation via VQ prototype modificationabstractA statistical technique for vector quantizer (VQ) prototype adaptation, based on tied-mixture continuous-parameter HMM's, is derived and evaluated on the basis of experimental evidence. Performance on difficult adaptation tasks indicates that VQ-prototype adaptation via tied-mixture HMM's constitutes a useful mechanism for speaker adaptation, particularly when there are substantial channel differences or when there is a large mismatch between reference and target speaker characteristics. Dimitry Rtischev, David Nahamoo, Michael Picheny |
IEEE Trans. Speech Audio Process. | 3 |
| 1993 | Context dependent vector quantization for continuous speech recognition
Lalit R. Bahl, Peter V. de Souza, Ponani S. Gopalakrishnan, Michael Picheny |
ICASSP (2) | 4 |
| 1993 | A supervised approach to the construction of context-sensitive acoustic prototypes
Jerome R. Bellegarda, Peter V. de Souza, David Nahamoo, Michael Picheny, Lalit R. Bahl |
ICASSP (2) | 4 |
| 1993 | Influence of background noise and microphone on the performance of the IBM Tangora speech recognition system
Subrata K. Das, Raimo Bakis, Arthur Nádas, David Nahamoo, Michael Picheny |
ICASSP (2) | 5 |
| 1993 | Word lookahead scheme for cross-word right context models in a stack decoder
Lalit R. Bahl, Peter V. de Souza, Ponani S. Gopalakrishnan, David Nahamoo, Michael Picheny |
EUROSPEECH | 5 |
| 1993 | Multonic Markov word models for large vocabulary continuous speech recognitionabstractA new class of hidden Markov models is proposed for the acoustic representation of words in an automatic speech recognition system. The models, built from combinations of acoustically based sub-word units called fenones, are derived automatically from one or more sample utterances of a word. Because they are more flexible than previously reported fenone-based word models, they lead to an improved capability of modeling variations in pronunciation. They are therefore particularly useful in the recognition of continuous speech. In addition, their construction is relatively simple, because it can be done using the well-known forward-backward algorithm for parameter estimation of hidden Markov models. Appropriate reestimation formulas are derived for this purpose. Experimental results obtained on a 5000-word vocabulary natural language continuous speech recognition task are presented to illustrate the enhanced power of discrimination of the new models.> Lalit R. Bahl, Jerome R. Bellegarda, Peter V. de Souza, Ponani S. Gopalakrishnan, David Nahamoo, Michael Picheny |
IEEE Trans. Speech Audio Process. | 6 |
| 1993 | A method for the construction of acoustic Markov models for wordsabstractA technique for constructing Markov models for the acoustic representation of words is described. Word models are constructed from models of subword units called fenones. Fenones represent very short speech events and are obtained automatically through the use of a vector quantizer. The fenonic baseform for a word-i.e., the sequence of fenones used to represent the word-is derived automatically from one or more utterances of that word. Since the word models are all composed from a small inventory of subword models, training for large-vocabulary speech recognition systems can be accomplished with a small training script. A method for combining phonetic and fenonic models is presented. Results of experiments with speaker-dependent and speaker-independent models on several isolated-word recognition tasks are reported. The results are compared with those for phonetics-based Markov models and template-based dynamic programming (DP) matching.> Lalit R. Bahl, Peter F. Brown, Peter V. de Souza, Robert L. Mercer, Michael Picheny |
IEEE Trans. Speech Audio Process. | 5 |
| 1992 | A fast match for continuous speech recognition using allophonic modelsabstractIn a large vocabulary real-time speech recognition system, there is a need for a fast method for selecting a list of candidate words from the vocabulary that match well with a given acoustic input. The authors describe a highly accurate fast acoustic match for continuous speech recognition. The algorithm uses allophonic models and efficient search techniques to select a set of candidate words. The allophonic models are derived by constructing decision trees that query the context in which each phone occurs to arrive at an allophone in a given context. The models for all the words in the vocabulary are arranged in a tree structure and efficient tree search algorithms are used to select a list of candidate words using these models. Using this method, the authors are able to obtain over 99% accuracy in the fast match for a continuous speech recognition task which has a vocabulary of 5000 words.> Lalit R. Bahl, Peter V. de Souza, Ponani S. Gopalakrishnan, David Nahamoo, Michael Picheny |
ICASSP | 5 |
| 1992 | Adaptation of large vocabulary recognition system parametersabstractThe authors report on a series of experiments in which the hidden Markov model baseforms and the language model probabilities were updated from spontaneously dictated speech captured during recognition sessions with the IBM Tangora system. The basic technique for baseform modification consisted of constructing new fenonic baseforms for all recognized words. To modify the language model probabilities, a simplified version of a cache language model was implemented. The word error rate across six talkers was 3.7%. Baseform adaptation reduced the average error rate to 3.5%, and using the cache language model reduced the error rate to 3.2%. Combining both techniques further reduced the error rate to 3.1%-a respectable improvement over the original error rate, especially given that the system was speaker-trained prior to adaptation.> Lalit R. Bahl, Peter V. de Souza, David Nahamoo, Michael Picheny, Salim Roukos |
ICASSP | 4 |
| 1992 | Robust speaker adaptation using a piecewise linear acoustic mappingabstractIn a large vocabulary speech recognition system, it is desirable to make use of previously acquired speech data when encountering new speakers. The authors describe an adaptation strategy based on a piecewise linear mapping between the feature space of a new speaker and that of a reference speaker. This speaker-normalizing mapping is used to transform the previously acquired parameters of the reference speaker onto the space of the new speaker. This results in a robust speaker adaptation procedure which allows for a drastic reduction in the amount of training data required from the new speaker. The performance of this method is illustrated on an isolated utterance speech recognition task with a vocabulary of 20000 words.> Jerome R. Bellegarda, Peter V. de Souza, Arthur Nádas, David Nahamoo, Michael Picheny, Lalit R. Bahl |
ICASSP | 5 |
| 1991 | A new class of fenonic Markov word models for large vocabulary continuous speech recognitionabstractA technique for constructing hidden Markov models for the acoustic representation of words is described. The models, built from combinations of acoustically based subword units called fenones, are derived automatically from one or more sample utterances of words. They are more flexible than previously reported fenone-based word models and lead to an improved capability of modeling variations in pronunciation. In addition, their construction is simplified, because it can be done using the well-known forward-backward algorithm for the parameter estimation of hidden Markov models. Experimental results obtained on a 5000-word vocabulary continuous speech recognition task are presented to illustrate some of the benefits associated with the new models. Multonic baseforms resulted in a reduction of 16% in the average error rate obtained for ten speakers.> Lalit R. Bahl, Jerome R. Bellegarda, Peter V. de Souza, Ponani S. Gopalakrishnan, David Nahamoo, Michael Picheny |
ICASSP | 6 |
| 1991 | Automatic phonetic baseform determinationabstractThe authors describe a series of experiments in which the phonetic baseform is deduced automatically for new words by utilizing actual utterances of the new word in conjunction with a set of automatically derived spelling-to-sound rules. Recognition performance was evaluated on new words spoken by two different speakers when the phonetic baseforms were extracted via the above approach. The error rates on these new words were found to be comparable to or better than when the phonetic baseforms were derived by hand, thus validating the basic approach.> Lalit R. Bahl, Subrata K. Das, Peter DeSouza, M. Epstein, Robert L. Mercer, Bernard Mérialdo, David Nahamoo, Michael Picheny, J. Powell |
ICASSP | 8 |
| 1991 | Decision trees for phonological rules in continuous speechabstractThe authors present an automatic method for modeling phonological variation using decision trees. For each phone they construct a decision tree that specifies the acoustic realization of the phone as a function of the context in which it appears. Several-thousand sentences from a natural language corpus spoken by several speakers are used to construct these decision trees. Experimental results on a 5000-word vocabulary natural language speech recognition task are presented.> Lalit R. Bahl, Peter V. de Souza, Ponani S. Gopalakrishnan, David Nahamoo, Michael Picheny |
ICASSP | 5 |
| 1991 | An iterative 'flip-flop' approximation of the most informative split in the construction of decision treesabstractThe authors seek a fast algorithm for finding the best question to ask (i.e., best split of predictor values) about a predictor variable when predicting membership in more than two categories. They give a fast iterative algorithm for finding a suboptimal question in the N category problem by exploiting a fast algorithm for finding the optimal question in the two-category problem. The algorithm has been used in a number of speech recognition applications.> Arthur Nádas, David Nahamoo, Michael Picheny, J. Powell |
ICASSP | 3 |
| 1989 | Large vocabulary natural language continuous speech recognitionabstractA description is presented of the authors' current research on automatic speech recognition of continuously read sentences from a naturally-occurring corpus: office correspondence. The recognition system combines features from their current isolated-word recognition system and from their previously developed continuous-speech recognition system. It consists of an acoustic processor, an acoustic channel model, a language model, and a linguistic decoder. Some new features in the recognizer relative to the isolated-word speech recognition system include the use of a fast match to prune rapidly to a manageable number the candidates considered by the detailed match, multiple pronunciations of all function words, and modeling of interphone coarticulatory behavior. The authors recorded training and test data from a set of ten male talkers. The perplexity of the test sentences was found to be 93; none of sentences was part of the data used to generate the language model. Preliminary (speaker-dependent) recognition results on these talkers yielded an average word error rate of 11.0%.> Lalit R. Bahl, Raimo Bakis, Jerome R. Bellegarda, Peter F. Brown, David Burshtein, Subrata K. Das, Peter V. de Souza, Ponani S. Gopalakrishnan, Frederick Jelinek, Dimitri Kanevsky, Robert L. Mercer, Arthur Nádas, David Nahamoo, Michael Picheny |
ICASSP | 14 |
| 1988 | Acoustic Markov models used in the Tangora speech recognition systemabstractThe Speech Recognition Group at IBM Research has developed a real-time, isolated-word speech recognizer called Tangora, which accepts natural English sentences drawn from a vocabulary of 20000 words. Despite its large vocabulary, the Tangora recognizer requires only about 20 minutes of speech from each new user for training purposes. The accuracy of the system and its ease of training are largely attributable to the use of hidden Markov models in its acoustic match component. An automatic technique for constructing Markov word models is described and results are included of experiments with speaker-dependent and speaker-independent models on several isolated-word recognition tasks.> Lalit R. Bahl, Peter F. Brown, Peter V. de Souza, Robert L. Mercer, Michael Picheny |
ICASSP | 5 |
| 1988 | Decoder selection based on cross-entropiesabstractThe authors generalize the maximum likelihood and related optimization criteria for training and decoding with a speech recognizer. The generalizations are constructed by considering weighted linear combinations of the logarithms of the likelihoods of words, of acoustics, and of (word, acoustic) pairs. The utility of various patterns of weights are examined.> Ponani S. Gopalakrishnan, Dimitri Kanevsky, Arthur Nádas, David Nahamoo, Michael Picheny |
ICASSP | 5 |
| 1988 | Speech recognition using noise-adaptive prototypesabstractA probabilistic mixture model is described for a frame (the short-term spectrum) of each to be used in speech recognition. Each component of the mixture is regarded as a prototype for the labeling phase of a hidden Markov model based speech recognition system. Since the ambient noise during recognition can differ from the ambient noise present in the training data, the model is designed for convenient updating in changing noise. Based on the observation that the energy in a frequency band is at any fixed time dominated either by signal energy or by noise energy, the authors model the energy as the larger of the separate energies of signal and noise in the band. Statistical algorithms are given for training this as a hidden variables model. The hidden variables are the prototype identities and the separate signal and noise components. A series of speech recognition experiments that successfully utilize this model is also discussed.> Arthur Nádas, David Nahamoo, Michael Picheny |
ICASSP | 3 |
| 1988 | Adaptive labeling: normalization of speech by adaptive transformations based on vector quantizationabstractA general technique termed adaptive labeling is presented for the normalization of the speech signal. In principle, adaptive labeling is applicable to any sequence of feature vectors of a given dimension. It combines the familiar labeling process executed by a vector quantizer with an adaptive renormalization transformation of the feature vectors proposed here. Adaptive labeling is applied to speech recognition, where the particular interest lies in diminishing the degradation of performance that occurs as a result of changes in the signal characteristics following changes in ambient noise and other recording environment conditions or in response to a change in the characteristics of the talker. Results are presented for a series of experiments using soft and loud noises as well as environments in which microphone-to-speaker distances were allowed to vary. A 5000-word vocabulary with isolated word input was used.> Arthur Nádas, David Nahamoo, Michael Picheny |
ICASSP | 3 |
| 1987 | Experiments with the Tangora 20, 000 word speech recognizerabstractThe Speech Recognition Group at IBM Research in Yorktown Heights has developed a real-time, isolated-utterance speech recognizer for natural language based on the IBM Personal Computer AT and IBM Signal Processors. The system has recently been enhanced by expanding the vocabulary from 5,000 words to 20,000 words and by the addition of a speech workstation to support usability studies on document creation by voice. The system supports spelling and interactive personalization to augment the vocabularies. This paper describes the implementation, user interface, and comparative performance of the recognizer. Amir Averbuch, Lalit R. Bahl, Raimo Bakis, Peter F. Brown, Gregg Daggett, Ken Davies, Steven V. De Gennaro, Peter V. de Souza, E. Epstein, D. Fraleigh, Frederick Jelinek, B. Lewis, Robert L. Mercer, J. Moorhead, Arthur Nádas, David Nahamoo, Michael Picheny, G. Shichman, P. Spinelli, Dirk Van Compernolle, H. Wilkens |
ICASSP | 18 |
| 1986 | An IBM PC based large-vocabulary isolated-utterance speech recognizerabstractThe Speech Recognition Group at IBM Research in Yorktown Heights has designed a real-time, isolated-utterance speech recognizer for natural language with a 5,000-word vocabulary based on the IBM Personal Computer (PC) AT model and two IBM Signal Processors realized in VLSI technology. The enrollment period for a new user is approximately 20 minutes. The basic vocabulary is chosen from the most common words in several collections of documents such as office memoranda and business letters. The system supports spelling and interactive personalization to augment this vocabulary. Signal processing, vector quantization, and acoustic matching algorithms are programmed on the IBM Signal Processors which fit into the PC AT chassis. The PC AT controls the Processors and implements the decoder stack search and the language model, as well as the application-specific interface. The modular architecture of the design is expandable to a 20,000-word vocabulary system by the addition of two more IBM Signal Processors housed in a PC Expansion Unit. Amir Averbuch, Lalit R. Bahl, Raimo Bakis, Peter F. Brown, A. G. Cole, Gregg Daggett, Subrata K. Das, Ken Davies, S. DeGennaro, Peter V. de Souza, E. Epstein, D. Fraleigh, Frederick Jelinek, Slava M. Katz, B. Lewis, Robert L. Mercer, Arthur Nádas, David Nahamoo, Michael Picheny, G. Shichman, P. Spinelli |
ICASSP | 19 |
| 1984 | Some experiments with large-vocabulary isolated-word sentence recognitionabstractThis paper deals with two experiments with a large vocabulary isolated word recognizer. The first compares word error rates for 1) meaningful sentences belonging to actual documents and 2) random word lists from the same vocabulary. The error rate is considerably lower for random word lists. The second experiment investigates the performance of the recognition system on sentences containing words outside the vocabulary of the recognizer. Sentences from a 5000 word vocabulary task are recognized with a recognizer limited to a 2000 word subvocabulary. The error rate is only slightly higher than it would be if recognition of the full 5000 word vocabulary was allowed. Lalit R. Bahl, Subrata K. Das, Peter V. de Souza, Frederick Jelinek, Slava M. Katz, Robert L. Mercer, Michael Picheny |
ICASSP | 7 |
| 1983 | Recognition of isolated-word sentences from a 5000-word vocabulary office correspondence taskabstractRecognition results on sentences from a 5000-word vocabulary drawn from office correspondence are presented. The sentences were read with pauses between the words. The vocabulary comprises the 5000 most frequently occurring words in a data-base of 14,000 office memoranda and letters, and has a perplexity of 90, measured from a trigram language model. Experiments were carried out with 6 speakers (4 male, 2 female) in an office environment using a close-talking microphone. The recognition system was automatically trained to each speaker by having the speaker read 100 typical sentences from the office correspondence data-base. Recognition was carried out for each speaker on 20 test sentences, consisting of 299 words. The recognition rate (% words correct) averaged across the 6 speakers was 94.5%. Lalit R. Bahl, A. G. Cole, Frederick Jelinek, Robert L. Mercer, Arthur Nádas, David Nahamoo, Michael Picheny |
ICASSP | 7 |