VLDB 2026 Research / reviewers in the wild / expert
Bhuvana Ramabhadran
dblp:39/1849
· DBLP profile ↗
218ranked-venue papers
12as first author
41since 2021 · last 2025
0000-0002-8049-2345ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 203 · 12 first-author · 41 since 2021Artificial intelligence and machine learning · 115 · 10 first-author · 21 since 2021Databases, data management, data science and information retrieval · 4Systems, architecture and hardware · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MADASR 2.0: Multi-Lingual Multi-Dialect ASR Challenge in 8 Indian LanguagesabstractWe present MADASR 2.0, a challenge at ASRU 2025 aimed at advancing multilingual and multidialectal automatic speech recognition (ASR) in low-resource Indian languages. Building on the 2023 edition, it introduces a subset of the RESPIN corpus, over 1200 hours of read speech across 8 languages and 33 dialects, with test sets including both read and spontaneous speech. The challenge comprises four tracks varying by training data size and external resource usage, and supports auxiliary tasks like language and dialect identification. We detail the dataset, tasks, baselines, and submissions and analyse trends across tracks and speech styles. Results highlight the continued difficulty of spontaneous ASR, the benefits of multitask and transfer learning, and effective strategies for building dialect-aware ASR systems. MADASR 2.0 offers a standardised benchmark to support future research on inclusive and scalable ASR for linguistically diverse populations. Sumit Sharma 0016, Deekshitha G, Abhayjeet Singh, Amartyaveer, Sathvik Udupa, Sandhya Badiger, Sanjeev Khudanpur, Sunayana Sitaram, Srinivasan Umesh, Bhuvana Ramabhadran, Brian Kingsbury, Hema A. Murthy, Srikanth S. Narayanan, Howard Lakougna, Prasanta Kumar Ghosh |
ASRU | 11 |
| 2025 | Audio Diffusion with Large Language ModelsabstractIn this paper, we explore an alternate approach to the popular method of using large language models (LLMs) as a second decoder for Automated Speech Recognition (ASR) and speech understanding tasks. We propose to employ diffusion networks to generate a correction signal that can be applied on the original input audio features to improve performance. Specifically, the diffusion network is trained to predict the gradient of any ASR objective with respect to the input audio features conditioned on LLM embeddings. Our experiments are conducted on public corpora, namely, Librispeech and Common Voice. We show that the diffusion model is able to improve ASR performance on noisy and accented speech, with the addition of knowledge from the LLM, and also helps improve generalization to out-of-domain test sets. Kyle Kastner, Kartik Audhkhasi, Bhuvana Ramabhadran, Andrew Rosenberg |
ICASSP | 4 |
| 2025 | Speech Re-Painting for Robust ASRabstractSynthetic speech is a useful source for augmentation of automatic speech recognition (ASR) systems, but there is a "sim-to-real" gap between synthetic and real speech that can limit generalization. The natural variability of real speech is essential to the training of robust ASR systems. While synthetic data augmentation can be used to approximate the variability of natural speech, however, not all aspects of variation are equally relevant for augmentation. In this work, we introduce speech re-painting, a method for in-context augmented synthesis, using target training datasets to generate new utterances guided by speech and text on the fly in a zero-shot manner. We evaluate this technique using downstream ASR word error rate (WER) using the VCTK and LibriSpeech datasets. These represent unique speaker and lexical challenges that are addressed by re-painting, realizing a reduction of WER more than 50% in particular settings. Kyle Kastner, Gary Wang, Isaac Elias, Takaaki Saeki, Pedro J. Moreno 0001, Françoise Beaufays, Andrew Rosenberg, Bhuvana Ramabhadran |
ICASSP | 8 |
| 2025 | Weak-to-Strong Generalization in Speech RecognitionabstractTo surpass human-level accuracy, speech recognition models must go beyond relying solely on human labels. To this end, we must build stronger models from weaker supervisors and this is the main goal in weak-to-strong generalization (WSG). WSG methods normally incorporate additional information into weak teacher models to improve their performance for example reliability of teacher-generated labels. In this research, we investigate two sources of additional information to implement WSG for speech recognition: unsupervised data from the target language and supervised data from other similar languages. We study scenarios where unsupervised data boosts performance, and propose a new clustering method leveraging supervised data from similar languages for further gains. Our clustering method yields an average 8% reduction in word error rate compared to a universal speech model trained on 182 languages. By incorporating both supervised data from similar languages and unsupervised data from the target language, we further enhance the USM model by 10%. This improvement reaches 15% for the top-performing languages with a WER below 50%. Soheil Khorram, Rohit Prabhavalkar, Kartik Audhkhasi, Bhuvana Ramabhadran |
ICASSP | 5 |
| 2025 | Identifying and Mitigating Mismatched Language Code in Multilingual ASRabstractMultilingual speech recognition systems often use an input language code in order to prompt the transcription in the target language. However, the spoken language in the input audio may not always match the language code, as often prevalent in multilingual societies. This language mismatch can significantly reduce ASR quality. We present a technique to identify and mitigate this issue. We combine off-the-shelf language-ID and language verification models to determine the language code input to the ASR model. The language verification model acts as a gate that decides when to trust the provided language code or use the output of the language-ID model. We compare these approaches with baselines that include vanilla language-ID based and language-independent ASR models. Our experiments on YouTube, SPRING-INX and FLEURS datasets shows the efficacy of the proposed model especially in the mismatched language code setting. Sepand Mavandadi, Kartik Audhkhasi, Shikhar Bharadwaj, Brian Farris, Tongzhou Chen, Bhuvana Ramabhadran, Sriram Ganapathy |
ICASSP | 7 |
| 2024 | Task Vector Algebra for ASR ModelsabstractVector representations of text and speech signals such as word2vec and wav2vec are used commonly in automatic speech recognition (ASR) and spoken language understanding systems. Recent results in natural language processing have proposed a task vector, defined as the difference vector between a model’s converged parameters and its initial parameters. Task vector algebra provides a simple and computationally-efficient way to solve several modeling problems, including model editing to reduce undesirable behavior, multi-tasking, and improving domain generalization. We apply task vectors to ASR models for the first time. Our experiments with Conformer-RNNT models trained on the SpeechStew corpora show that task vectors retain their scaling and multi-tasking applications. We propose two novel applications of task vectors to ASR. First, we show that task vectors can perform zero-shot adaptation of ASR models to unseen domains without using any supervised data. Second, we present a novel "task analogy" formulation that enables us to use models trained on high-resource tasks to improve performance on low-resource tasks. We also explore a technique to improve the performance of task vector arithmetic for ASR models. Gowtham Ramesh, Kartik Audhkhasi, Bhuvana Ramabhadran |
ICASSP | 3 |
| 2024 | Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed DataabstractCollecting high-quality studio recordings of audio is challenging, which limits the language coverage of text-to-speech (TTS) systems. This paper proposes a framework for scaling a multilingual TTS model to 100+ languages using found data without supervision. The proposed framework combines speech-text encoder pretraining with unsupervised training using untranscribed speech and unspoken text data sources, thereby leveraging massively multilingual joint speech and text representation learning. Without any transcribed speech in a new language, this TTS model can generate intelligible speech in >30 unseen languages (CER difference of <10% to ground truth). With just 15 minutes of transcribed, found data, we can reduce the intelligibility difference to 1% or less from the ground-truth, and achieve naturalness scores that match the ground-truth in several languages. Takaaki Saeki, Gary Wang, Nobuyuki Morioka, Isaac Elias, Kyle Kastner, Andrew Rosenberg, Bhuvana Ramabhadran, Heiga Zen, Françoise Beaufays, Hadar Shemtov |
ICASSP | 7 |
| 2024 | Speech Prefix-Tuning with RNNT Loss for Improving LLM Predictions
Murali Karthick Baskar, Andrew Rosenberg, Bhuvana Ramabhadran, Neeraj Gaur, Zhong Meng |
INTERSPEECH | 3 |
| 2024 | ASTRA: Aligning Speech and Text Representations for Asr without Sampling
Neeraj Gaur, Rohan Agrawal, Gary Wang, Parisa Haghani, Andrew Rosenberg, Bhuvana Ramabhadran |
INTERSPEECH | 6 |
| 2024 | Text Injection for Neural Contextual Biasing
Zhong Meng, Zelin Wu, Rohit Prabhavalkar, Cal Peyser, Nanxin Chen, Tara N. Sainath, Bhuvana Ramabhadran |
INTERSPEECH | 8 |
| 2024 | Contemplative Mechanism for Speech Recognition: Speech Encoders can Think
Tien-Ju Yang, Andrew Rosenberg, Bhuvana Ramabhadran |
INTERSPEECH | 3 |
| 2024 | Speculative Speech Recognition by Audio-Prefixed Low-Rank Adaptation of Language Models
Bolaji Yusuf, Murali Karthick Baskar, Andrew Rosenberg, Bhuvana Ramabhadran |
INTERSPEECH | 4 |
| 2023 | Mask-Conformer: Augmenting Conformer with Mask-Predict DecoderabstractMuch of the recent progress in automatic speech recognition (ASR) lies in developing an acoustic encoder, such as enlarging its capacity and designing a refined architecture for speech processing. With these highly optimized encoders, the decoder has become less influential in its role as a language model (LM). In this work, we explore an effective approach for employing the LM structure in an ASR model. The proposed Mask-Conformer augments a Conformer-based model with a mask-predict decoder, which learns output context via the masked LM objective. The mask-predict decoder is applied to stacks of encoder layers, where the decoder output explicitly conditions the subsequent layers using cross-attention. We also propose a fill-mask decoding algorithm that refines a sequence using the decoder’s linguistic information. Experimental results show that Mask-Conformer outperforms strong baselines on some tasks. In addition, our analyses validate the effectiveness of the proposed model design. Yosuke Higuchi, Andrew Rosenberg, Murali Karthick Baskar, Bhuvana Ramabhadran |
ASRU | 5 |
| 2023 | Modular Conformer Training for Flexible End-to-End ASRabstractThe state-of-the-art conformer used in automatic speech recognition combines feed-forward, convolution and multi-headed self-attention layers in a single model that is trained end-to-end with a decoder network. While this end-to-end training is simple and beneficial for word error rate, it restricts the ability to perform inference with the model at different operating points of word error rate and latency. Existing approaches to overcome this limitation include cascaded encoders and variable attention context models. We propose an alternative approach, called Modular Conformer training, which splits the Conformer model into a backbone convolutional model and attention submodels, which are added at each layer. We conduct experiments with a few training techniques on the Librispeech and Librilight corpus. We show that dropping-out the attention layers during the training of the backbone model allows for the largest WER improvements upon adding fine-tuned attention submodels, without impacting the WER of the backbone model itself. Kartik Audhkhasi, Brian Farris, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
ICASSP | 3 |
| 2023 | Large-Scale Language Model Rescoring on Long-Form DataabstractIn this work, we study the impact of Large-scale Language Models (LLM) on Automated Speech Recognition (ASR) of YouTube videos, which we use as a source for long-form ASR. We demonstrate up to 8% relative reduction in Word Error Eate (WER) on US English (en-us) and code-switched Indian English (en-in) long-form ASR test sets and a reduction of up to 30% relative on Salient Term Error Rate (STER) over a strong first-pass baseline that uses a maximum-entropy based language model. Improved lattice processing that results in a lattice with a proper (non-tree) digraph topology and carrying context from the 1-best hypothesis of the previous segment(s) results in significant wins in rescoring with LLMs. We also find that the gains in performance from the combination of LLMs trained on vast quantities of available data (such as C4 [1]) and conventional neural LMs is additive and significantly outperforms a strong first-pass baseline with a maximum entropy LM. Tongzhou Chen, Cyril Allauzen, Daniel S. Park, David Rybach, W. Ronny Huang, Rodrigo Cabrera, Kartik Audhkhasi, Bhuvana Ramabhadran, Pedro J. Moreno 0001, Michael Riley 0001 |
ICASSP | 9 |
| 2023 | JEIT: Joint End-to-End Model and Internal Language Model Training for Speech RecognitionabstractWe propose JEIT, a joint end-to-end (E2E) model and internal language model (ILM) training method to inject large-scale unpaired text into ILM during E2E training which improves rare-word speech recognition. With JEIT, the E2E model computes an E2E loss on audio-transcript pairs while its ILM estimates a cross-entropy loss on unpaired text. The E2E model is trained to minimize a weighted sum of E2E and ILM losses. During JEIT, ILM absorbs knowledge from unpaired text while the E2E training serves as regularization. Unlike ILM adaptation methods, JEIT does not require a separate adaptation step and avoids the need for Kullback-Leibler divergence regularization of ILM. We also show that modular hybrid autoregressive transducer (MHAT) performs better than HAT in the JEIT framework, and is much more robust than HAT during ILM adaptation. To push the limit of unpaired text injection, we further propose a combined JEIT and JOIST training (CJJT) that benefits from modality matching, encoder text injection and ILM training. Both JEIT and CJJT can foster a more effective LM fusion. With 100B unpaired sentences, JEIT/CJJT improves rare-word recognition accuracy by up to 16.4% over a model trained without unpaired text. Zhong Meng, Rohit Prabhavalkar, Tara N. Sainath, Tongzhou Chen, Ehsan Variani, Yu Zhang 0033, Bo Li 0028, Andrew Rosenberg, Bhuvana Ramabhadran |
ICASSP | 10 |
| 2023 | Virtuoso: Massive Multilingual Speech-Text Joint Semi-Supervised Learning for Text-to-SpeechabstractThis paper proposes Virtuoso, a massively multilingual speech–text joint semi-supervised learning framework for text-to-speech synthesis (TTS) models. Existing multilingual TTS typically supports tens of languages, which are a small fraction of the thousands of languages in the world. One difficulty to scale multilingual TTS to hundreds of languages is collecting high-quality speech–text paired data in low-resource languages. This study extends Maestro, a speech–text joint pretraining framework for automatic speech recognition (ASR), to speech generation tasks. To train a TTS model from various types of speech and text data, different training schemes are designed to handle supervised (paired TTS and ASR data) and unsupervised (untranscribed speech and unspoken text) datasets. Experimental evaluation shows that 1) multilingual TTS models trained on Virtuoso can achieve significantly better naturalness and intelligibility than baseline ones in seen languages, and 2) they can synthesize reasonably intelligible and naturally sounding speech for unseen languages where no high-quality paired TTS data is available. Takaaki Saeki, Heiga Zen, Zhehuai Chen, Nobuyuki Morioka, Gary Wang, Yu Zhang 0033, Ankur Bapna, Andrew Rosenberg, Bhuvana Ramabhadran |
ICASSP | 9 |
| 2023 | Understanding Shared Speech-Text RepresentationsabstractRecently, a number of approaches to train speech models by incorporating text into end-to-end models have been developed, with Maestro advancing state-of-the-art automatic speech recognition (ASR) and Speech Translation (ST) performance. In this paper, we expand our understanding of the resulting shared speech-text representations with two types of analyses. First we examine the limits of speech-free domain adaptation, finding that a corpus-specific duration model for speech-text alignment is the most important component for learning a shared speech-text representation. Second, we inspect the similarities between activations of unimodal (speech or text) encoders as compared to the activations of a shared encoder. We find that the shared encoder learns a more compact and overlapping speech-text representation than the uni-modal encoders. We hypothesize that this partially explains the effectiveness of the Maestro shared speech-text representations. Gary Wang, Kyle Kastner, Ankur Bapna, Zhehuai Chen, Andrew Rosenberg, Bhuvana Ramabhadran, Yu Zhang 0033 |
ICASSP | 6 |
| 2023 | Robust Knowledge Distillation from RNN-T Models with Noisy Training Labels Using Full-Sum LossabstractThis work studies knowledge distillation (KD) and addresses its constraints for recurrent neural network transducer (RNN-T) models. In hard distillation, a teacher model transcribes large amounts of unlabelled speech to train a student model. Soft distillation is another popular KD method that distills the output logits of the teacher model. Due to the nature of RNN-T alignments, applying soft distillation between RNNT architectures having different posterior distributions is challenging. In addition, bad teachers having high word-error-rate (WER) reduce the efficacy of KD. We investigate how to effectively distill knowledge from variable quality ASR teachers, which has not been studied before to the best of our knowledge. We show that a sequence-level KD, full-sum distillation, outperforms other distillation methods for RNN-T models, especially for bad teachers. We also propose a variant of full-sum distillation that distills the sequence discriminative knowledge of the teacher leading to further improvement in WER. We conduct experiments on public datasets namely SpeechStew and LibriSpeech, and on in-house production data. Mohammad Zeineldeen, Kartik Audhkhasi, Murali Karthick Baskar, Bhuvana Ramabhadran |
ICASSP | 4 |
| 2023 | O-1: Self-training with Oracle and 1-best Hypothesis
Murali Karthick Baskar, Andrew Rosenberg, Bhuvana Ramabhadran, Kartik Audhkhasi |
INTERSPEECH | 3 |
| 2023 | Using Text Injection to Improve Recognition of Personal Identifiers in Speech
Yochai Blau, Rohan Agrawal, Lior Madmony, Gary Wang, Andrew Rosenberg, Zhehuai Chen, Zorik Gekhman, Genady Beryozkin, Parisa Haghani, Bhuvana Ramabhadran |
INTERSPEECH | 10 |
| 2022 | Tts4pretrain 2.0: Advancing the use of Text and Speech in ASR Pretraining with Consistency and Contrastive LossesabstractAn effective way to learn representations from untranscribed speech and unspoken text with linguistic/lexical representations derived from synthesized speech was introduced in tts4pretrain [1]. However, the representations learned from synthesized and real speech are likely to be different, potentially limiting the improvements from incorporating unspoken text. In this paper, we introduce learning from supervised speech earlier on in the training process with consistency-based regularization between real and synthesized speech. This allows for better learning of shared speech and text representations. Thus, we introduce a new objective, with encoder and decoder consistency and contrastive regularization between real and synthesized speech derived from the labeled corpora during the pretraining stage. We show that the new objective leads to more similar representations derived from speech and text that help downstream ASR. The proposed pretraining method yields Word Error Rate (WER) reductions of 7-21% relative on six public corpora, Librispeech, AMI, TEDLIUM, Common Voice, Switchboard, CHiME-6, over a state-of-the-art baseline pretrained with wav2vec2.0 and 2-17% over the previously proposed tts4pretrain. The proposed method outperforms the supervised SpeechStew by up to 17%. Moreover, we show that the proposed method also yields WER reductions on larger data sets by evaluating on a large resource, in-house Voice Search task and streaming ASR. Zhehuai Chen, Yu Zhang 0033, Andrew Rosenberg, Bhuvana Ramabhadran, Pedro J. Moreno 0001, Gary Wang |
ICASSP | 4 |
| 2022 | Multilingual Second-Pass Rescoring for Automatic Speech Recognition SystemsabstractSecond-pass rescoring is a well known technique to improve the performance of Automatic Speech Recognition (ASR) systems. Neural Oracle Search (NOS), which selects the most likely hypothesis from an N-best hypothesis list by integrating information from multiple sources, such as the input acoustic representations, N-best hypotheses, additional first-pass statistics, and unpaired textual information through an external language model, has shown success in rescoring for RNN-T first-pass models. Multilingual first-pass speech recognition models often outperform their monolingual counterparts when trained on related or low-resource languages. In this paper, we investigate the use of the NOS rescoring model on a first-pass multilingual model and show that similar to the first-pass model, the rescoring model can be made multilingual. Our first-pass multilingual model does not require a language-id and we make a realistic assumption that an estimate of the language-id would be available for second-pass rescoring. We conduct comprehensive experiments on two sets of languages, one consisting of related low-resource languages, and the other with a high-resource language added to the first set to analyze the performance of the multilingual NOS rescorer under different settings. Our experimental results show that, multilingual NOS can improve the first-pass multilingual model resulting in average word error rate reduction of 9.4% in the first case, and 8.4% in the second, and out-performing the monolingual counterparts in both cases. Neeraj Gaur, Tongzhou Chen, Ehsan Variani, Parisa Haghani, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
ICASSP | 5 |
| 2022 | Analysis of Self-Attention Head Diversity for Conformer-based Automatic Speech RecognitionabstractAttention layers are an integral part of modern end-to-end automatic speech recognition systems, for instance as part of the Transformer or Conformer architecture.Attention is typically multi-headed, where each head has an independent set of learned parameters and operates on the same input feature sequence.The output of multi-headed attention is a fusion of the outputs from the individual heads.We empirically analyze the diversity between representations produced by the different attention heads and demonstrate that the heads become highly correlated during the course of training.We investigate a few approaches to increasing attention head diversity, including using different attention mechanisms for each head and auxiliary training loss functions to promote head diversity.We show that introducing diversity-promoting auxiliary loss functions during training is a more effective approach, and obtain WER improvements of up to 6% relative on the Librispeech corpus.Finally, we draw a connection between the diversity of attention heads and the similarity of the gradients of head parameters. Kartik Audhkhasi, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
INTERSPEECH | 3 |
| 2022 | Reducing Domain mismatch in Self-supervised speech pre-training
Murali Karthick Baskar, Andrew Rosenberg, Bhuvana Ramabhadran, Yu Zhang 0033 |
INTERSPEECH | 3 |
| 2022 | MAESTRO: Matched Speech Text Representations through Modality MatchingabstractWe present Maestro, a self-supervised training method to unify representations learnt from speech and text modalities.Self-supervised learning from speech signals aims to learn the latent structure inherent in the signal, while self-supervised learning from text attempts to capture lexical information.Learning aligned representations from unpaired speech and text sequences is a challenging task.Previous work either implicitly enforced the representations learnt from these two modalities to be aligned in the latent space through multitasking and parameter sharing or explicitly through conversion of modalities via speech synthesis.While the former suffers from interference between the two modalities, the latter introduces additional complexity.In this paper, we propose Maestro, a novel algorithm to learn unified representations from both these modalities simultaneously that can transfer to diverse downstream tasks such as Automated Speech Recognition (ASR) and Speech Translation (ST).Maestro learns unified representations through sequence alignment, duration prediction and matching embeddings in the learned space through an aligned masked-language model loss.We establish a new state-of-the-art (SOTA) on VoxPopuli multilingual ASR with a 8% relative reduction in Word Error Rate (WER), multidomain SpeechStew ASR (3.7% relative) and 21 languages to English multilingual ST on CoVoST 2 with an improvement of 2.8 BLEU averaged over 21 languages. Zhehuai Chen, Yu Zhang 0033, Andrew Rosenberg, Bhuvana Ramabhadran, Pedro J. Moreno 0001, Ankur Bapna, Heiga Zen |
INTERSPEECH | 4 |
| 2022 | On Adaptive Weight Interpolation of the Hybrid Autoregressive Transducer
Ehsan Variani, Michael Riley 0001, David Rybach, Cyril Allauzen, Tongzhou Chen, Bhuvana Ramabhadran |
INTERSPEECH | 6 |
| 2022 | Improving Rare Word Recognition with LM-aware MWER TrainingabstractLanguage models (LMs) significantly improve the recognition accuracy of end-to-end (E2E) models on words rarely seen during training, when used in either the shallow fusion or the rescoring setups. In this work, we introduce LMs in the learning of hybrid autoregressive transducer (HAT) models in the discriminative training framework, to mitigate the training versus inference gap regarding the use of LMs. For the shallow fusion setup, we use LMs during both hypotheses generation and loss computation, and the LM-aware MWER-trained model achieves 10\% relative improvement over the model trained with standard MWER on voice search test sets containing rare words. For the rescoring setup, we learn a small neural module to generate per-token fusion weights in a data-dependent manner. This model achieves the same rescoring WER as regular MWER-trained model, but without the need for sweeping fusion weights. Tongzhou Chen, Tara N. Sainath, Ehsan Variani, Rohit Prabhavalkar, W. Ronny Huang, Bhuvana Ramabhadran, Neeraj Gaur, Sepand Mavandadi, Cal Peyser, Trevor Strohman, Yanzhang He, David Rybach |
INTERSPEECH | 7 |
| 2022 | Non-Parallel Voice Conversion for ASR AugmentationabstractAutomatic speech recognition (ASR) needs to be robust to speaker differences.Voice Conversion (VC) modifies speaker characteristics of input speech.This is an attractive feature for ASR data augmentation.In this paper, we demonstrate that voice conversion can be used as a data augmentation technique to improve ASR performance, even on LibriSpeech, which contains 2,456 speakers.For ASR augmentation, it is necessary that the VC model be robust to a wide range of input speech.This motivates the use of a non-autoregressive, non-parallel VC model, and the use of a pretrained ASR encoder within the VC model.This work suggests that despite including many speakers, speaker diversity may remain a limitation to ASR quality.Finally, interrogation of our VC performance has provided useful metrics for objective evaluation of VC quality. Gary Wang, Andrew Rosenberg, Bhuvana Ramabhadran, Fadi Biadsy, Jesse Emond, Pedro J. Moreno 0001 |
INTERSPEECH | 3 |
| 2022 | Maestro-U: Leveraging Joint Speech-Text Representation Learning for Zero Supervised Speech ASRabstractTraining state-of-the-art Automated Speech Recognition (ASR) models typically requires a substantial amount of transcribed speech. In this work, we demonstrate that a modality-matched joint speech and text model introduced in [1] can be leveraged to train a massively multilingual ASR model without any supervised (manually transcribed) speech for some languages. This paper explores the use of jointly learnt speech and text representations in a massively multilingual, zero supervised speech, real-world setting to expand the set of languages covered by ASR with only unlabeled speech and text in the target languages. Using the FLEURS dataset, we define the task to cover 102 languages, where transcribed speech is available in 52 of these languages and can be used to improve end-to-end ASR quality on the remaining 50. First, we show that by combining speech representations with byte-level text representations and use of language embeddings, we can dramatically reduce the Character Error Rate (CER) on languages with no supervised speech from 64.8% to 30.8%, a relative reduction of 53%. Second, using a subset of South Asian languages we show that Maestro-U can promote knowledge transfer from languages with supervised speech even when there is limited to no graphemic overlap. Overall, Maestro-U closes the gap to oracle performance by 68.5% relative and reduces the CER of 19 languages below 15%. Zhehuai Chen, Ankur Bapna, Andrew Rosenberg, Yu Zhang 0033, Bhuvana Ramabhadran, Pedro J. Moreno 0001, Nanxin Chen |
SLT | 5 |
| 2022 | Modular Hybrid Autoregressive TransducerabstractText-only adaptation of a transducer model remains challenging for end-to-end speech recognition since the transducer has no clearly separated acoustic model (AM), language model (LM) or blank model. In this work, we propose a modular hybrid autoregressive transducer (MHAT) that has structurally separated label and blank decoders to predict label and blank distributions, respectively, along with a shared acoustic encoder. The encoder and label decoder outputs are directly projected to AM and internal LM scores and then added to compute label posteriors. We train MHAT with an internal LM loss and a HAT loss to ensure that its internal LM becomes a standalone neural LM that can be effectively adapted to text. Moreover, text adaptation of MHAT fosters a much better LM fusion than internal LM subtraction-based methods. On Google's large-scale production data, a multi-domain MHAT adapted with 100B sentences achieves relative WER reductions of up to 12.4% without LM fusion and 21.5% with LM fusion from 400K-hour trained HAT. Zhong Meng, Tongzhou Chen, Rohit Prabhavalkar, Yu Zhang 0033, Gary Wang, Kartik Audhkhasi, Jesse Emond, Trevor Strohman, Bhuvana Ramabhadran, W. Ronny Huang, Ehsan Variani, Pedro J. Moreno 0001 |
SLT | 9 |
| 2022 | G-Augment: Searching for the Meta-Structure of Data Augmentation Policies for ASRabstractData augmentation is a ubiquitous technique used to provide robustness to automatic speech recognition (ASR) training. However, even as so much of the ASR training process has become automated and more “end-to-end,” the data augmentation policy (what augmentation functions to use, and how to apply them) remains hand-crafted. We present G(raph)-Augment, a technique to define the augmentation space as directed acyclic graphs (DAGs) and search over this space to optimize the augmentation policy itself. We show that given the same computational budget, policies produced by G-Augment are able to perform better than SpecAugment policies obtained by random search on fine-tuning tasks on CHiME-6 and AMI. G-Augment is also able to establish a new state-of-the-art ASR performance on the CHiME-6 evaluation set (30.7% WER). We further demonstrate that G- Augment policies show better transfer properties across warm-start to cold-start training and model size compared to random-searched SpecAugment policies. Gary Wang, Ekin Dogus Cubuk, Andrew Rosenberg, Shuyang Cheng, Ron J. Weiss, Bhuvana Ramabhadran, Pedro J. Moreno 0001, Quoc V. Le, Daniel S. Park |
SLT | 6 |
| 2021 | Injecting Text in Self-Supervised Speech PretrainingabstractSelf-supervised pretraining for Automated Speech Recognition (ASR) has shown varied degrees of success. In this paper, we propose to jointly learn representations during pretraining from two different modalities: speech and text. The proposed method, tts4pretrain complements the power of contrastive learning in self-supervision with linguistic/lexical representations derived from synthesized speech, effectively learning from untranscribed speech and unspoken text. Lexical learning in the speech encoder is enforced through an additional sequence loss term that is coupled with contrastive loss during pretraining. We demonstrate that this novel pretraining method yields Word Error Rate (WER) reductions of 10% relative on the well-benchmarked, Librispeech task over a state-of-the-art baseline pretrained with wav2vec2.0 only. The proposed method also serves as an effective strategy to compensate for the lack of transcribed speech, effectively matching the performance of 5000 hours of transcribed speech with just 100 hours of transcribed speech on the AMI meeting transcription task. Finally, we demonstrate WER reductions of up to 15% on an inhouse Voice Search task over traditional pretraining. Incorporating text into encoder pretraining is complimentary to rescoring with a larger or in-domain language model, resulting in additional 6% relative reduction in WER. Zhehuai Chen, Yu Zhang 0033, Andrew Rosenberg, Bhuvana Ramabhadran, Gary Wang, Pedro J. Moreno 0001 |
ASRU | 4 |
| 2021 | Extending Parrotron: An End-to-End, Speech Conversion and Speech Recognition Model for Atypical SpeechabstractWe present an extended Parrotron model: a single, end-to-end network that enables voice conversion and recognition simultaneously. Input spectrograms are transformed to output spectrograms in the voice of a predetermined target speaker while also generating hypotheses in a target vocabulary. We study the performance of this novel architecture, which jointly predicts speech and text, on atypical (e.g. dysarthric) speech. We show that with as little as an hour of atypical speech, speaker adaptation can yield a 77% relative reduction in Word Error Rate (WER), measured by ASR performance on the converted speech. We also show that data augmentation using a customized synthesizer built on atypical speech can provide an additional 10% relative improvement over the best speaker-adapted model. Finally, we show how these methods generalize across 8 types of atypical speech for a range of speech impairment severities. Rohan Doshi, Youzheng Chen, Liyang Jiang, Fadi Biadsy, Bhuvana Ramabhadran, Fang Chu, Andrew Rosenberg, Pedro J. Moreno 0001 |
ICASSP | 6 |
| 2021 | Mixture of Informed Experts for Multilingual Speech RecognitionabstractWhen trained on related or low-resource languages, multilingual speech recognition models often outperform their monolingual counterparts. However, these models can suffer from loss in performance for high resource or unrelated languages. We investigate the use of a mixture-of-experts approach to assign per-language parameters in the model to increase network capacity in a structured fashion. We introduce a novel variant of this approach, ‘informed experts’, which attempts to tackle inter-task conflicts by eliminating gradients from other tasks in these task-specific parameters. We conduct experiments on a real-world task with English, French and four dialects of Arabic to show the effectiveness of our approach. Our model matches or outperforms the monolingual models for almost all languages, with gains of as much as 31% relative. Our model also outperforms the baseline multilingual model for all languages by up to 9% relative. Neeraj Gaur, Brian Farris, Parisa Haghani, Isabel Leal, Pedro J. Moreno 0001, Manasa Prasad, Bhuvana Ramabhadran |
ICASSP | 7 |
| 2021 | Convolutional Dropout and Wordpiece Augmentation for End-to-End Speech RecognitionabstractRegularization and data augmentation are crucial to training end-to-end automatic speech recognition systems. Dropout is a popular regularization technique, which operates on each neuron independently by multiplying it with a Bernoulli random variable. We propose a generalization of dropout, called "convolutional dropout", where each neuron’s activation is replaced with a randomly-weighted linear combination of neuron values in its neighborhood. We believe that this formulation combines the regularizing effect of dropout with the smoothing effects of the convolution operation. In addition to convolutional dropout, this paper also proposes using random word-piece segmentations as a data augmentation scheme during training, inspired by results in neural machine translation. We adopt both these methods during the training of transformer-transducer speech recognition models, and show consistent WER improvements on Librispeech as well as across different languages. Hainan Xu, Kartik Audhkhasi, Bhuvana Ramabhadran |
ICASSP | 5 |
| 2021 | Mixture Model Attention: Flexible Streaming and Non-Streaming Automatic Speech Recognition
Kartik Audhkhasi, Tongzhou Chen, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
Interspeech | 3 |
| 2021 | Conformer Parrotron: A Faster and Stronger End-to-End Speech Conversion and Recognition Model for Atypical Speech
Zhehuai Chen, Bhuvana Ramabhadran, Fadi Biadsy, Youzheng Chen, Liyang Jiang, Fang Chu, Rohan Doshi, Pedro J. Moreno 0001 |
Interspeech | 2 |
| 2021 | Semi-Supervision in ASR: Sequential MixMatch and Factorized TTS-Based Augmentation
Zhehuai Chen, Andrew Rosenberg, Yu Zhang 0033, Heiga Zen, Mohammadreza Ghodsi, Jesse Emond, Gary Wang, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
Interspeech | 9 |
| 2021 | Self-Adaptive Distillation for Multilingual Speech Recognition: Leveraging Student Independence
Isabel Leal, Neeraj Gaur, Parisa Haghani, Brian Farris, Pedro J. Moreno 0001, Manasa Prasad, Bhuvana Ramabhadran |
Interspeech | 7 |
| 2021 | Regularizing Word Segmentation by Creating Misspellings
Hainan Xu, Kartik Audhkhasi, Jesse Emond, Bhuvana Ramabhadran |
Interspeech | 5 |
| 2020 | Language-Agnostic Multilingual ModelingabstractMultilingual Automated Speech Recognition (ASR) systems allow for the joint training of data-rich and data-scarce languages in a single model. This enables data and parameter sharing across languages, which is especially beneficial for the data-scarce languages. However, most state-of-the-art multilingual models require the encoding of language information and therefore are not as flexible or scalable when expanding to newer languages. Language-independent multilingual models help to address this issue, and are also better suited for multicultural societies where several languages are frequently used together (but often rendered with different writing systems). In this paper, we propose a new approach to building a language-agnostic multilingual ASR system which transforms all languages to one writing system through a many-to-one transliteration transducer. Thus, similar sounding acoustics are mapped to a single, canonical target sequence of graphemes, effectively separating the modeling and rendering problems. We show with four Indic languages, namely, Hindi, Bengali, Tamil and Kannada, that the language-agnostic multilingual model achieves up to 10% relative reduction in Word Error Rate (WER) over a language-dependent multilingual model. Arindrima Datta, Bhuvana Ramabhadran, Jesse Emond, Anjuli Kannan, Brian Roark |
ICASSP | 2 |
| 2020 | Generating Diverse and Natural Text-to-Speech Samples Using a Quantized Fine-Grained VAE and Autoregressive Prosody PriorabstractRecent neural text-to-speech (TTS) models with fine-grained latent features enable precise control of the prosody of synthesized speech. Such models typically incorporate a fine-grained variational autoencoder (VAE) structure, extracting latent features at each input token (e.g., phonemes). However, generating samples with the standard VAE prior often results in unnatural and discontinuous speech, with dramatic prosodic variation between tokens. This paper proposes a sequential prior in a discrete latent space which can generate more naturally sounding samples. This is accomplished by discretizing the latent features using vector quantization (VQ), and separately training an autoregressive (AR) prior model over the result. We evaluate the approach using listening tests, objective metrics of automatic speech recognition (ASR) performance, and measurements of prosody attributes. Experimental results show that the proposed model significantly improves the naturalness in random sample generation. Furthermore, initial experiments demonstrate that randomly sampling from the proposed model can be used as data augmentation to improve the ASR performance. Guangzhi Sun, Yu Zhang 0033, Ron J. Weiss, Yuan Cao 0007, Heiga Zen, Andrew Rosenberg, Bhuvana Ramabhadran |
ICASSP | 7 |
| 2020 | Neural Oracle Search on N-BEST HypothesesabstractIn this paper, we propose a neural search algorithm to select the most likely hypothesis using a sequence of acoustic representations and multiple hypotheses as input. The algorithm provides a sequence level score for each audio-hypothesis pair that is obtained by integrating information from multiple sources, such as the input acoustic representations, N-best hypotheses, additional 1st-pass statistics, and unpaired textual information through an external language model. These scores are then used to map the search problem of identifying the most likely hypothesis to a sequence classification problem. The definition of the proposed algorithm is broad enough to allow its use as an alternative to beam search in the 1st-pass or as a 2nd-pass, rescoring step. This algorithm achieves up to 12% relative reductions in Word Error Rate (WER) across several languages over state-of-the-art baselines with relatively few additional parameters. We also propose the use of a binary classifier gating function that can learn to trigger the 2nd-pass neural search model when the 1-best hypothesis is not the oracle hypothesis, thereby avoiding extra computation. Ehsan Variani, Tongzhou Chen, James Apfel, Bhuvana Ramabhadran, Seungji Lee, Pedro J. Moreno 0001 |
ICASSP | 4 |
| 2020 | Improving Speech Recognition Using Consistent Predictions on Synthesized SpeechabstractSpeech synthesis has advanced to the point of being close to indistinguishable from human speech. However, efforts to train speech recognition systems on synthesized utterances have not been able to show that synthesized data can be effectively used to augment or replace human speech. In this work, we demonstrate that promoting consistent predictions in response to real and synthesized speech enables significantly improved speech recognition performance. We also find that training on 460 hours of LibriSpeech augmented with 500 hours of transcripts (without audio) performance is within 0.2% WER of a system trained on 960 hours of transcribed audio. This suggests that with this approach, when there is sufficient text available, reliance on transcribed audio can be cut nearly in half. Gary Wang, Andrew Rosenberg, Zhehuai Chen, Yu Zhang 0033, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
ICASSP | 5 |
| 2020 | Improving Speech Recognition Using GAN-Based Speech Synthesis and Contrastive Unspoken Text Selection
Zhehuai Chen, Andrew Rosenberg, Yu Zhang 0033, Gary Wang, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
INTERSPEECH | 5 |
| 2020 | SCADA: Stochastic, Consistent and Adversarial Data Augmentation to Improve ASR
Gary Wang, Andrew Rosenberg, Zhehuai Chen, Yu Zhang 0033, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
INTERSPEECH | 5 |
| 2020 | Multilingual Speech Recognition with Self-Attention Structured Parameterization
Parisa Haghani, Anshuman Tripathi, Bhuvana Ramabhadran, Brian Farris, Hainan Xu, Han Lu 0003, Hasim Sak, Isabel Leal, Neeraj Gaur, Pedro J. Moreno 0001 |
INTERSPEECH | 4 |
| 2019 | Speech Recognition with Augmented Synthesized SpeechabstractRecent success of the Tacotron speech synthesis architecture and its variants in producing natural sounding multi-speaker synthesized speech has raised the exciting possibility of replacing expensive, manually transcribed, domain-specific, human speech that is used to train speech recognizers. The multi-speaker speech synthesis architecture can learn latent embedding spaces of prosody, speaker and style variations derived from input acoustic representations thereby allowing for manipulation of the synthesized speech. In this paper, we evaluate the feasibility of enhancing speech recognition performance using speech synthesis using two corpora from different domains. We explore algorithms to provide the necessary acoustic and lexical diversity needed for robust speech recognition. Finally, we demonstrate the feasibility of this approach as a data augmentation strategy for domain-transfer. We find that improvements to speech recognition performance is achievable by augmenting training data with synthesized material. However, there remains a substantial gap in performance between recognizers trained on human speech those trained on synthesized speech. Andrew Rosenberg, Yu Zhang 0033, Bhuvana Ramabhadran, Ye Jia, Pedro J. Moreno 0001, Zelin Wu |
ASRU | 3 |
| 2019 | Comparison of Data Augmentation and Adaptation Strategies for Code-switched Automatic Speech RecognitionabstractCode-switching occurs when the speaker alternates between two or more languages or dialects. It is a pervasive phenomenon in most Indic spoken languages. Code-switching poses a challenge in language modeling as it complicates the orthographic realization of text, and generally, there is a shortage of code-switched data. In this paper, we investigate data augmentation and adaptation strategies for language modeling. Using Bengali and English as an example, we study augmenting the code-switched transcripts with separate transliterated Bengali and English corpora. We present results on two speech recognition tasks, namely, voice search and dictation. We show improvements on both tasks with Maximum Entropy (MaxEnt) and Long Short-Term Memory (LSTM) language models (LMs). We also explore different adaptation strategies for MaxEnt LM and LSTM LM, demonstrating that the transliteration-based data-augmented LSTM LM matches the adapted MaxEnt LM which is trained on more Bengali-English data. Bhuvana Ramabhadran, Jesse Emond, Andrew Rosenberg, Fadi Biadsy |
ICASSP | 2 |
| 2019 | Large-Scale Multilingual Speech Recognition with a Streaming End-to-End ModelabstractMultilingual end-to-end (E2E) models have shown great promise in expansion of automatic speech recognition (ASR) coverage of the world's languages.They have shown improvement over monolingual systems, and have simplified training and serving by eliminating language-specific acoustic, pronunciation, and language models.This work presents an E2E multilingual system which is equipped to operate in low-latency interactive applications, as well as handle a key challenge of real world data: the imbalance in training data across languages.Using nine Indic languages, we compare a variety of techniques, and find that a combination of conditioning on a language vector and training language-specific adapter layers produces the best model.The resulting E2E multilingual model achieves a lower word error rate (WER) than both monolingual E2E models (eight of nine languages) and monolingual conventional systems (all nine languages). Anjuli Kannan, Arindrima Datta, Tara N. Sainath, Eugene Weinstein, Bhuvana Ramabhadran, Ankur Bapna, Seungji Lee |
INTERSPEECH | 5 |
| 2019 | Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice CloningabstractWe present a multispeaker, multilingual text-to-speech (TTS) synthesis model based on Tacotron that is able to produce high quality speech in multiple languages.Moreover, the model is able to transfer voices across languages, e.g.synthesize fluent Spanish speech using an English speaker's voice, without training on any bilingual or parallel examples.Such transfer works across distantly related languages, e.g.English and Mandarin.Critical to achieving this result are: 1. using a phonemic input representation to encourage sharing of model capacity across languages, and 2. incorporating an adversarial loss term to encourage the model to disentangle its representation of speaker identity (which is perfectly correlated with language in the training data) from the speech content.Further scaling up the model by training on multiple speakers of each language, and incorporating an autoencoding input to help stabilize attention during training, results in a model which can be used to consistently synthesize intelligible speech for training speakers in all languages seen during training, and in native or foreign accents. Yu Zhang 0033, Ron J. Weiss, Heiga Zen, R. J. Skerry-Ryan, Ye Jia, Andrew Rosenberg, Bhuvana Ramabhadran |
INTERSPEECH | 9 |
| 2018 | Building Competitive Direct Acoustics-to-Word Models for English Conversational Speech RecognitionabstractDirect acoustics-to-word (A2W) models in the end-to-end paradigm have received increasing attention compared to conventional subword based automatic speech recognition models using phones, characters, or context-dependent hidden Markov model states. This is because A2W models recognize words from speech without any decoder, pronunciation lexicon, or externally-trained language model, making training and decoding with such models simple. Prior work has shown that A2W models require orders of magnitude more training data in order to perform comparably to conventional models. Our work also showed this accuracy gap when using the English Switchboard-Fisher data set. This paper describes a recipe to train an A2W model that closes this gap and is at-par with state-of-the-art sub-word based models. We achieve a word error rate of 8.8.8%/13.9% on the Hub5-2000 Switchboard/CallHome test sets without any decoder or language model. We find that model initialization, training data order, and regularization have the most impact on the A2W model performance. Next, we present a joint word-character A2W model that learns to first spell the word and then recognize it. This model provides a rich output to the user instead of simple word hypotheses, making it especially useful in the case of words unseen or rarely-seen during training. Kartik Audhkhasi, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, Michael Picheny |
ICASSP | 3 |
| 2018 | Whole Sentence Neural Language ModelsabstractRecurrent neural networks have become increasingly popular for the task of language modeling achieving impressive gains in state-of-the-art speech recognition and natural language processing (NLP) tasks. Recurrent models exploit word dependencies over a much longer context window (as retained by the history states) than what is feasible with n-gram language models. However the training criterion of choice for recurrent language models continues to be the local conditional likelihood of generating the current word given the (pos-sibly long) word context, thus making local decisions at each word. This locally-conditional design fundamentally limits the ability of the model in exploiting whole sentence structures. In this paper, we present our initial results at whole sentence neural language models which assign a probability to the entire word sequence. We extend the previous work on whole sentence maximum entropy models to recurrent language models while using Noise Contrastive Estimation (NCE) for training, as these sentence models are fundamentally un-normalizable. We present results on a range of tasks: from sequence identification tasks such as, palindrome detection to large vocabulary automatic speech recognition (LVCSR) and demonstrate the modeling power of this approach. Abhinav Sethy, Kartik Audhkhasi, Bhuvana Ramabhadran |
ICASSP | 4 |
| 2018 | Measuring the Effect of Linguistic Resources on Prosody Modeling for Speech SynthesisabstractThe generation of natural and expressive prosodic contours is an important component of a text-to-speech (TTS) system which, in most classical architectures, relies on the existence of a text-analysis processor that can extract prosody-predictive features and pass them to a statistical learning model. These features can range from basic properties of the input string to rich high-level features which may not be always available when developing a TTS system in a new language with sparse computational resources. In this work we investigate how the prosody model of a speech-synthesis system performs as a function of different predictive feature sets that assume access to a certain amount of rich resources. We investigate, using objective metrics, the effect of relaxing the assumptions on input representations for prosody prediction for 5 languages, and evaluate the perceptual implications for US English. Andrew Rosenberg, Raul Fernandez, Bhuvana Ramabhadran |
ICASSP | 3 |
| 2018 | Joint Modeling of Accents and Acoustics for Multi-Accent Speech RecognitionabstractThe performance of automatic speech recognition systems degrades with increasing mismatch between the training and testing scenarios. Differences in speaker accents are a significant source of such mismatch. The traditional approach to deal with multiple accents involves pooling data from several accents during training and building a single model in multi-task fashion, where tasks correspond to individual accents. In this paper, we explore an alternate model where we jointly learn an accent classifier and a multi-task acoustic model. Experiments on the American English Wall Street Journal and British English Cambridge corpora demonstrate that our joint model outperforms the strong multi-task acoustic model baseline. We obtain a 5.94% relative improvement in word error rate on British English, and 9.47% relative improvement on American English. This illustrates that jointly modeling with accent information improves acoustic model performance. Xuesong Yang, Kartik Audhkhasi, Andrew Rosenberg, Samuel Thomas 0001, Bhuvana Ramabhadran, Mark Hasegawa-Johnson |
ICASSP | 5 |
| 2018 | Data Augmentation Improves Recognition of Foreign Accented Speech
Takashi Fukuda, Raul Fernandez, Andrew Rosenberg, Samuel Thomas 0001, Bhuvana Ramabhadran, Alexander Sorin, Gakuto Kurata |
INTERSPEECH | 5 |
| 2018 | Open Problems in Speech Recognition
Bhuvana Ramabhadran |
INTERSPEECH | 1 |
| 2018 | Transliteration Based Approaches to Improve Code-Switched Speech Recognition PerformanceabstractCode-switching is a commonly occurring phenomenon in many multilingual communities, wherein a speaker switches between languages within a single utterance. Conventional Word Error Rate (WER) is not sufficient for measuring the performance of code-mixed languages due to ambiguities in transcription, misspellings and borrowing of words from two different writing systems. These rendering errors artificially inflate the WER of an Automated Speech Recognition (ASR) system and complicate its evaluation. Furthermore, these errors make it harder to accurately evaluate modeling errors originating from code-switched language and acoustic models. In this work, we propose the use of a new metric, transliteration-optimized Word Error Rate (toWER) that smoothes out many of these irregularities by mapping all text to one writing system and demonstrate a correlation with the amount of code-switching present in a language. We also present a novel approach to acoustic and language modeling for bilingual code-switched Indic languages using the same transliteration approach to normalize the data for three types of language models, namely, a conventional n-gram language model, a maximum entropy based language model and a Long Short Term Memory (LSTM) language model, and a state-of-the-art Connectionist Temporal Classification (CTC) acoustic model. We demonstrate the robustness of the proposed approach on several Indic languages from Google Voice Search traffic with significant gains in ASR performance up to 10% relative over the state-of-the-art baseline. Jesse Emond, Bhuvana Ramabhadran, Brian Roark, Pedro J. Moreno 0001 |
SLT | 2 |
| 2017 | Language modeling with highway LSTMabstractLanguage models (LMs) based on Long Short Term Memory (LSTM) have shown good gains in many automatic speech recognition tasks. In this paper, we extend an LSTM by adding highway networks inside an LSTM and use the resulting Highway LSTM (HW-LSTM) model for language modeling. The added highway networks increase the depth in the time dimension. Since a typical LSTM has two internal states, a memory cell and a hidden state, we compare various types of HW-LSTM by adding highway networks onto the memory cell and/or the hidden state. Experimental results on English broadcast news and conversational telephone speech recognition show that the proposed HW-LSTM LM improves speech recognition accuracy on top of a strong LSTM LM baseline. We report 5.1% and 9.9% on the Switchboard and CallHome subsets of the Hub5 2000 evaluation, which reaches the best performance numbers reported on these tasks to date. Gakuto Kurata, Bhuvana Ramabhadran, George Saon, Abhinav Sethy |
ASRU | 2 |
| 2017 | End-to-end ASR-free keyword search from speechabstractEnd-to-end (E2E) systems have achieved competitive results compared to conventional hybrid hidden Markov model (HMM)-deep neural network based automatic speech recognition (ASR) systems. Such E2E systems are attractive due to the lack of dependence on alignments between input acoustic and output grapheme or HMM state sequence during training. This paper explores the design of an ASR-free end-to-end system for text query-based keyword search (KWS) from speech trained with minimal supervision. Our E2E KWS system consists of three sub-systems. The first sub-system is a recurrent neural network (RNN)-based acoustic auto-encoder trained to reconstruct the audio through a finite-dimensional representation. The second sub-system is a character-level RNN language model using embeddings learned from a convolutional neural network. Since the acoustic and text query embeddings occupy different representation spaces, they are input to a third feed-forward neural network that predicts whether the query occurs in the acoustic utterance or not. This E2E ASR-free KWS system performs respectably despite lacking a conventional ASR system and trains much faster. Kartik Audhkhasi, Andrew Rosenberg, Abhinav Sethy, Bhuvana Ramabhadran, Brian Kingsbury |
ICASSP | 4 |
| 2017 | Training variance and performance evaluation of neural networks in speechabstractIn this work we study variance in the results of neural network training on a wide variety of configurations in automatic speech recognition. Although this variance itself is well known, this is, to the best of our knowledge, the first paper that performs an extensive empirical study on its effects in speech recognition. We view training as sampling from a distribution and show that these distributions can have a substantial variance. These results show the urgent need to rethink the way in which results in the literature are reported and interpreted. Ewout van den Berg, Bhuvana Ramabhadran, Michael Picheny |
ICASSP | 2 |
| 2017 | Knowledge distillation across ensembles of multilingual models for low-resource languagesabstractThis paper investigates the effectiveness of knowledge distillation in the context of multilingual models. We show that with knowledge distillation, Long Short-Term Memory(LSTM) models can be used to train standard feed-forward Deep Neural Network (DNN) models for a variety of low-resource languages. We then examine how the agreement between the teacher's best labels and the original labels affects the student model's performance. Next, we show that knowledge distillation can be easily applied to semi-supervised learning to improve model performance. We also propose a promising data selection method to filter un-transcribed data. Then we focus on knowledge transfer among DNN models with multilingual features derived from CNN+DNN, LSTM, VGG, CTC and attention models. We show that a student model equipped with better input features not only learns better from the teacher's labels, but also outperforms the teacher. Further experiments suggest that by learning from each other, the original ensemble of various models is able to evolve into a new ensemble with even better combined performance. Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, Tom Sercu, Kartik Audhkhasi, Abhinav Sethy, Markus Nußbaum-Thom, Andrew Rosenberg |
ICASSP | 3 |
| 2017 | Voice-transformation-based data augmentation for prosodic classificationabstractIn this work we explore data-augmentation techniques for the task of improving the performance of a supervised recurrent-neural-network classifier tasked with predicting prosodic-boundary and pitch-accent labels. The technique is based on applying voice transformations to the training data that modify the pitch baseline and range, as well as the vocal-tract and vocal-source characteristics of the speakers to generate further training examples. We demonstrate the validity of the approach by improving performance when the amount of base labeled examples is small (showing reductions in the range of 7%–12% for reduced-data conditions) as well as in terms of its generalization to speakers unseen in the training set (showing a relative reduction in the error rate of 8.74% and 4.75%, on the average, for boundaries and accent tasks respectively, in leave-one-speaker-out validation). Raul Fernandez, Andrew Rosenberg, Alexander Sorin, Bhuvana Ramabhadran, Ron Hoory |
ICASSP | 4 |
| 2017 | Effective joint training of denoising feature space transforms and Neural Network based acoustic modelsabstractNeural Network (NN) based acoustic frontends, such as denoising autoencoders, are actively being investigated to improve the robustness of NN based acoustic models to various noise conditions. In recent work the joint training of such frontends with backend NNs has been shown to significantly improve speech recognition performance. In this paper, we propose an effective algorithm to jointly train such a denoising feature space transform and a NN based acoustic model with various kinds of data. Our proposed method first pretrains a Convolutional Neural Network (CNN) based denoising frontend and then jointly trains this frontend with a NN backend acoustic model. In the unsupervised pretraining stage, the frontend is designed to estimate clean log Mel-filterbank features from noisy log-power spectral input features. A subsequent multi-stage training of the proposed frontend, with the dropout technique applied only at the joint layer between the frontend and backend NNs, leads to significant improvements in the overall performance. On the Aurora-4 task, our proposed system achieves an average WER of 9.98%. This is a 9.0% relative improvement over one of the best reported speaker independent baseline system's performance. A final semi-supervised adaptation of the frontend NN, similar to feature space adaptation, reduces the average WER to 7.39%, a further relative WER improvement of 25%. Takashi Fukuda, Osamu Ichikawa, Gakuto Kurata, Ryuki Tachibana, Samuel Thomas 0001, Bhuvana Ramabhadran |
ICASSP | 6 |
| 2017 | Harmonic feature fusion for robust neural network-based acoustic modelingabstractAcoustic modeling with deep learning has drastically improved the performance of automatic speech recognition (ASR) where the main stream of the acoustic feature is still log-Mel filtered one. While the log-Mel filtered features lose harmonic-structure information, they still include useful information for ASR. Several attempts have been made to integrate higher-resolution information into the network. In order to improve the ASR accuracy in noisy conditions, we propose new features integrated into acoustic modeling to represent which parts in the time-frequency domain have a distinct harmonic structure, since it is partially observed in noisy environments. The new features are combined with the standard acoustic features, and the network is trained with them using various noisy data. Through these operations, it learns the acoustic features with a kind of quality tag describing which parts are clean or degraded. Our model reduced the word error rate in an Aurora-4 task by 10.3% in DNN compared with the strong baseline while retaining the high accuracy in clean test cases. Osamu Ichikawa, Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata, Bhuvana Ramabhadran |
ICASSP | 5 |
| 2017 | End-to-end speech recognition and keyword search on low-resource languagesabstractIn recent years, so-called, “end-to-end” speech recognition systems have emerged as viable alternatives to traditional ASR frameworks. Keyword search, localizing an orthographic query in a speech corpus, is typically performed by using automatic speech recognition (ASR) to generate an index. Previous work has evaluated the use of end-to-end systems for ASR on well known corpora (WSJ, Switchboard, TIMIT, etc.) in high-resource languages like English and Mandarin. In this work, we investigate the use of Connectionist Temporal Classification (CTC) networks, recurrent encoder-decoders with attention, two end-to-end ASR systems for keyword search and speech recognition on low resource languages. We find end-to-end systems can generate high quality 1-best transcripts on low-resource languages, but, because they generate very sharp posteriors, their utility is limited for KWS. We explore a number of ways to address this limitation with modest success. Experimental results reported are based on the IARPA BABEL OP3 languages and evaluation framework. This paper represents the first results using “end-to-end” techniques for speech recognition and keyword search on low-resource languages. Andrew Rosenberg, Kartik Audhkhasi, Abhinav Sethy, Bhuvana Ramabhadran, Michael Picheny |
ICASSP | 4 |
| 2017 | Network architectures for multilingual speech representation learningabstractMultilingual (ML) representations play a key role in building speech recognition systems for low resource languages. The IARPA sponsored BABEL program focuses on building speech recognition (ASR) and keyword search (KWS) systems in over 24 languages with limited training data. The most common mechanism to derive ML representations in the BABEL program has been with the use of a two-stage network, the first stage being a convolutional network (CNN) from where multilingual features are extracted, expanded contextually and used as input to the second stage which can be a feed-forward DNN or a CNN. The final multilingual representations are derived from the second network. This paper presents two novel methods for deriving ML representations. The first is based on Long-Short Term Memory (LSTM) networks and the second is based on a very deep CNN (VGG-net). We demonstrate that ML features extracted from both models show significant improvement over the baseline CNN-DNN based ML representations, in terms of both speech recognition and keyword search performance and draw the comparison between the LSTM model itself and the ML representations derived from it on Georgian, the surprise language for the OpenKWS evaluation. Tom Sercu, George Saon, Jia Cui, Bhuvana Ramabhadran, Brian Kingsbury, Abhinav Sethy |
ICASSP | 5 |
| 2017 | Direct Acoustics-to-Word Models for English Conversational Speech RecognitionabstractRecent work on end-to-end automatic speech recognition (ASR) has shown that the connectionist temporal classification (CTC) loss can be used to convert acoustics to phone or character sequences.Such systems are used with a dictionary and separately-trained Language Model (LM) to produce word sequences.However, they are not truly end-to-end in the sense of mapping acoustics directly to words without an intermediate phone representation.In this paper, we present the first results employing direct acoustics-to-word CTC models on two well-known public benchmark tasks: Switchboard and Call-Home.These models do not require an LM or even a decoder at run-time and hence recognize speech with minimal complexity.However, due to the large number of word output units, CTC word models require orders of magnitude more data to train reliably compared to traditional systems.We present some techniques to mitigate this issue.Our CTC word model achieves a word error rate of 13.0%/18.8%on the Hub5-2000 Switchboard/CallHome test sets without any LM or decoder compared with 9.6%/16.0%for phone-based CTC with a 4-gram LM.We also present rescoring results on CTC word model lattices to quantify the performance benefits of a LM, and contrast the performance of word and phone CTC models. Kartik Audhkhasi, Bhuvana Ramabhadran, George Saon, Michael Picheny, David Nahamoo |
INTERSPEECH | 2 |
| 2017 | Efficient Knowledge Distillation from an Ensemble of Teachers
Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata, Samuel Thomas 0001, Jia Cui, Bhuvana Ramabhadran |
INTERSPEECH | 6 |
| 2017 | Fast Neural Network Language Model Lookups at N-Gram Speeds
Abhinav Sethy, Bhuvana Ramabhadran |
INTERSPEECH | 3 |
| 2017 | Empirical Exploration of Novel Architectures and Objectives for Language Models
Gakuto Kurata, Abhinav Sethy, Bhuvana Ramabhadran, George Saon |
INTERSPEECH | 3 |
| 2017 | Weakly-Supervised Phrase Assignment from Text in a Speech-Synthesis System Using Noisy Labels
Asaf Rendel, Raul Fernandez, Zvi Kons, Andrew Rosenberg, Ron Hoory, Bhuvana Ramabhadran |
INTERSPEECH | 6 |
| 2017 | Bias and Statistical Significance in Evaluating Speech Synthesis with Mean Opinion Scores
Andrew Rosenberg, Bhuvana Ramabhadran |
INTERSPEECH | 2 |
| 2017 | English Conversational Telephone Speech Recognition by Humans and MachinesabstractOne of the most difficult speech recognition tasks is accurate recognition of human to human communication. Advances in deep learning over the last few years have produced major speech recognition improvements on the representative Switchboard conversational corpus. Word error rates that just a few years ago were 14% have dropped to 8.0%, then 6.6% and most recently 5.8%, and are now believed to be within striking range of human performance. This then raises two issues - what IS human performance, and how far down can we still drive speech recognition error rates? A recent paper by Microsoft suggests that we have already achieved human performance. In trying to verify this statement, we performed an independent set of human performance measurements on two conversational tasks and found that human performance may be considerably better than what was earlier reported, giving the community a significantly harder goal to achieve. We also report on our own efforts in this area, presenting a set of acoustic and language modeling techniques that lowered the word error rate of our own English conversational telephone LVCSR system to the level of 5.5%/10.3% on the Switchboard/CallHome subsets of the Hub5 2000 evaluation, which - at least at the writing of this paper - is a new performance milestone (albeit not at what we measure to be human performance!). On the acoustic side, we use a score fusion of three models: one LSTM with multiple feature inputs, a second LSTM trained with speaker-adversarial multi-task learning and a third residual net (ResNet) with 25 convolutional layers and time-dilated convolutions. On the language modeling side, we use word and character LSTMs and convolutional WaveNet-style language models. George Saon, Gakuto Kurata, Tom Sercu, Kartik Audhkhasi, Samuel Thomas 0001, Dimitrios Dimitriadis, Bhuvana Ramabhadran, Michael Picheny, Lynn-Li Lim, Bergul Roomi, Phil Hall |
INTERSPEECH | 8 |
| 2017 | Symbol Sequence Search from Telephone Conversation
Masayuki Suzuki, Gakuto Kurata, Abhinav Sethy, Bhuvana Ramabhadran, Kenneth Church 0001, Mark Drake |
INTERSPEECH | 4 |
| 2017 | Parallel Deep Neural Network Training for Big Data on Blue Gene/QabstractDeep Neural Networks (DNNs) have recently been shown to significantly outperform existing machine learning techniques in several pattern recognition tasks. DNNs are the state-of-the-art models used in image recognition, object detection, classification and tracking, and speech and language processing applications. The biggest drawback to DNNs has been the enormous cost in computation and time taken to train the parameters of the networks-often a tenfold increase relative to conventional technologies. Such training time costs can be mitigated by the application of parallel computing algorithms and architectures. However, these algorithms often run into difficulties because of the cost of inter-processor communication bottlenecks. In this paper, we describe how to enable Parallel Deep Neural Network Training on the IBM Blue Gene/Q (BG/Q) computer system. Specifically, we explore DNN training using the data-parallel Hessian-free 2nd order optimization algorithm. Such an algorithm is particularly well-suited to parallelization across a large set of loosely coupled processors. BG/Q, with its excellent inter-processor communication characteristics, is an ideal match for this type of algorithm. The paper discusses how issues regarding programming model and data-dependent imbalances are addressed. Results on large-scale speech tasks show that the performance on BG/Q scales linearly up to 4,096 processes with no loss in accuracy. This allows us to train neural networks using billions of training examples in a few hours. I-Hsin Chung, Tara N. Sainath, Bhuvana Ramabhadran, Michael Picheny, John A. Gunnels, Vernon Austel, Upendra V. Chaudhari, Brian Kingsbury |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2016 | Semantic word embedding neural network language models for automatic speech recognitionabstractSemantic word embeddings have become increasingly important in natural language processing tasks over the last few years. This popularity is due to their ability to easily capture rich semantic information through a distributed representation and the availability of fast and scalable algorithms for learning them from large text corpora. State-of-the-art neural network language models (NNLMs) used in automatic speech recognition (ASR) and natural language processing also learn word embeddings optimized to model local N-gram dependencies given training text but are not optimized to capture semantic information. We hypothesize that semantic word embeddings provide diverse information compared to the word embeddings learned by NNLMs. We propose novel feedforward NNLM architectures that incorporate semantic word embeddings. We apply the resulting NNLMs to ASR on broadcast news and show improvements in both perplexity and word error rate. Kartik Audhkhasi, Abhinav Sethy, Bhuvana Ramabhadran |
ICASSP | 3 |
| 2016 | Efficient one-vs-one kernel ridge regression for speech recognitionabstractRecent evidences suggest that the performance of kernel methods may match that of deep neural networks (DNNs), which have been the state-of-the-art approach for speech recognition. In this work, we present an improvement of the kernel ridge regression studied in Huang et al., ICASSP 2014, and show that our proposal is computationally advantageous. Our approach performs classifications by using the one-vs-one scheme, which, under certain assumptions, reduces the costs of the one-vs-rest scheme by asymptotically a factor of c2 in training time and c in memory consumption. Here, c is the number of classes and it is typically on the order of hundreds and thousands for speech recognition. We demonstrate empirical results on the benchmark corpus TIMIT. In particular, the classification accuracy is one to two percentages higher (in the absolute term) than the best of the kernel methods and of the DNNs reported by Huang et al, and the speech recognition accuracy is highly comparable. Jie Chen 0007, Lingfei Wu 0001, Kartik Audhkhasi, Brian Kingsbury, Bhuvana Ramabhadran |
ICASSP | 5 |
| 2016 | Using continuous lexical embeddings to improve symbolic-prosody prediction in a text-to-speech front-endabstractThe prediction of symbolic prosodic categories from text is an important, but challenging, natural-language processing task given the various ways in which an input can be realized, and the fact that knowledge about what features determine this realization is incomplete or inaccessible to the model. In this work, we look at augmenting baseline features with lexical representations that are derived from text, providing continuous embeddings of the lexicon in a lower-dimensional space. Although learned in an unsupervised fashion, such features capture semantic and syntactic properties that make them amenable for prosody prediction. We deploy various embedding models on prominence- and phrase-break prediction tasks, showing substantial gains, particularly for prominence prediction. Asaf Rendel, Raul Fernandez, Ron Hoory, Bhuvana Ramabhadran |
ICASSP | 4 |
| 2016 | Acoustic Modeling Using Bidirectional Gated Recurrent Convolutional Units
Markus Nußbaum-Thom, Jia Cui, Bhuvana Ramabhadran, Vaibhava Goel |
INTERSPEECH | 3 |
| 2016 | Domain Adaptation of CNN Based Acoustic Models Under Limited Resource Settings
Masayuki Suzuki, Ryuki Tachibana, Samuel Thomas 0001, Bhuvana Ramabhadran, George Saon |
INTERSPEECH | 4 |
| 2016 | Multilingual Data Selection for Low Resource Speech RecognitionabstractAbstract : Feature representations extracted from deep neural network-based multilingual frontends provide significant improvements to speech recognition systems in low resource settings. To effectively train these frontends, we introduce a data selection technique that discovers language groups from an available set of training languages. This data selection method reduces the required amount of training data and training time by approximately 40 , with minimal performance degradation. We present speech recognition results on 7 very limited language pack (VLLP) languages from the second option period of the IARPA Babel program using multilingual features trained on up to 10 languages. The proposed multilingual features provide up to 15 relative improvement over baseline acoustic features on the VLLP languages. Samuel Thomas 0001, Kartik Audhkhasi, Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran |
INTERSPEECH | 5 |
| 2015 | Multilingual representations for low resource speech recognition and keyword searchabstractThis paper examines the impact of multilingual (ML) acoustic representations on Automatic Speech Recognition (ASR) and keyword search (KWS) for low resource languages in the context of the OpenKWS15 evaluation of the IARPA Babel program. The task is to develop Swahili ASR and KWS systems within two weeks using as little as 3 hours of transcribed data. Multilingual acoustic representations proved to be crucial for building these systems under strict time constraints. The paper discusses several key insights on how these representations are derived and used. First, we present a data sampling strategy that can speed up the training of multilingual representations without appreciable loss in ASR performance. Second, we show that fusion of diverse multilingual representations developed at different LORELEI sites yields substantial ASR and KWS gains. Speaker adaptation and data augmentation of these representations improves both ASR and KWS performance (up to 8.7% relative). Third, incorporating un-transcribed data through semi-supervised learning, improves WER and KWS performance. Finally, we show that these multilingual representations significantly improve ASR and KWS performance (relative 9% for WER and 5% for MTWV) even when forty hours of transcribed audio in the target language is available. Multilingual representations significantly contributed to the LORELEI KWS systems winning the OpenKWS15 evaluation. Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran, Abhinav Sethy, Kartik Audhkhasi, Ellen Eide, Lidia Mangu, Markus Nußbaum-Thom, Michael Picheny, Zoltán Tüske, Pavel Golik, Ralf Schlüter, Hermann Ney, Mark J. F. Gales, Kate M. Knill, Anton Ragni, Philip C. Woodland |
ASRU | 3 |
| 2015 | Bidirectional recurrent neural network language models for automatic speech recognitionabstractRecurrent neural network language models have enjoyed great success in speech recognition, partially due to their ability to model longer-distance context than word n-gram models. In recurrent neural networks (RNNs), contextual information from past inputs is modeled with the help of recurrent connections at the hidden layer, while Long Short-Term Memory (LSTM) neural networks are RNNs that contain units that can store values for arbitrary amounts of time. While conventional unidirectional networks predict outputs from only past inputs, one can build bidirectional networks that also condition on future inputs. In this paper, we propose applying bidirectional RNNs and LSTM neural networks to language modeling for speech recognition. We discuss issues that arise when utilizing bidirectional models for speech, and compare unidirectional and bidirectional models on an English Broadcast News transcription task. We find that bidirectional RNNs significantly outperform unidirectional RNNs, but bidirectional LSTMs do not provide any further gain over their unidirectional counterparts. Ebru Arisoy, Abhinav Sethy, Bhuvana Ramabhadran, Stanley F. Chen |
ICASSP | 3 |
| 2015 | Unnormalized exponential and neural network language modelsabstractModel M, an exponential class-based language model, and neural network language models (NNLM's) have outperformed word n-gram language models over a wide range of tasks. However, these gains come at the cost of vastly increased computation when calculating word probabilities. For both models, the bulk of this computation involves evaluating the softmax function over a large word or class vocabulary to ensure that probabilities sum to 1. In this paper, we study unnormalized variants of Model M and NNLM's, whereby the softmax function is simply omitted. Accordingly, model training must be modified to encourage scores to sum close to 1. In this paper, we demonstrate up to a factor of 35 faster n-gram lookups with unnormalized models over their normalized counterparts, while still yielding state-of-the-art performance in WER (10.2 on the English broadcast news rt04 set). Abhinav Sethy, Stanley F. Chen, Ebru Arisoy, Bhuvana Ramabhadran |
ICASSP | 4 |
| 2015 | Efficient GPU implementation of convolutional neural networks for speech recognition
Ewout van den Berg, Daniel Brand, Rajesh Bordawekar, Leonid Rachevsky, Bhuvana Ramabhadran |
INTERSPEECH | 5 |
| 2015 | A multi-region deep neural network model in speech recognition
Jia Cui, George Saon, Bhuvana Ramabhadran, Brian Kingsbury |
INTERSPEECH | 3 |
| 2015 | Using deep bidirectional recurrent neural networks for prosodic-target prediction in a unit-selection text-to-speech system
Raul Fernandez, Asaf Rendel, Bhuvana Ramabhadran, Ron Hoory |
INTERSPEECH | 3 |
| 2015 | Modeling phrasing and prominence using deep recurrent learning
Andrew Rosenberg, Raul Fernandez, Bhuvana Ramabhadran |
INTERSPEECH | 3 |
| 2015 | Deep Convolutional Neural Networks for Large-scale Speech Tasks
Tara N. Sainath, Brian Kingsbury, George Saon, Hagen Soltau, Abdel-rahman Mohamed, George E. Dahl, Bhuvana Ramabhadran |
Neural Networks | 7 |
| 2014 | Semi-supervised term-weighted value rescoring for keyword searchabstractWe present a semi-supervised algorithm for rescoring the output of a speech keyword search (KWS) system. Conventional loss functions such as squared-error and logistic loss are not suitable for optimizing the commonly-used KWS term-weighted value (TWV) performance metric. We derive a novel concave modified logistic log-likelihood function which lower-bounds TWV. We then use a manifold-regularized kernel classifier that maximizes this lower-bound. A manifold regularization term in our objective function uses available unlabeled speech data and makes our approach semi-supervised. This term is particularly useful for KWS in low-resource languages and ensures that the predicted keyword confidence scores are smooth on a low-dimensional manifold in the feature space. We conduct KWS experiments on the IARPA Babel Vietnamese task and show performance improvements in terms of the maximum TWV (MTWV). Our estimated confidence score is complementary with respect to the ASR posterior score and gives MTWV improvement upon interpolation with it. Kartik Audhkhasi, Abhinav Sethy, Bhuvana Ramabhadran, Shri Narayanan |
ICASSP | 3 |
| 2014 | Automatic keyword selection for keyword search development and tuningabstractIn this paper, we investigate the problem of automatically selecting textual keywords for keyword search development and tuning on audio data for any language. Briefly, the method samples candidate keywords in the training data while trying to match a set of target marginal distributions for keyword features such as keyword frequency in the training or development audio, keyword length, frequency of out-of-vocabulary words, and TF-IDF scores. The method is evaluated on four IARPA Babel program base period languages. We show the use of the automatically selected keywords for the keyword search system development and tuning. We show also that search performance is improved by tuning the decision threshold on the automatically selected keywords. Jia Cui, Jonathan Mamou, Brian Kingsbury, Bhuvana Ramabhadran |
ICASSP | 4 |
| 2014 | Kernel methods match Deep Neural Networks on TIMITabstractDespite their theoretical appeal and grounding in tractable convex optimization techniques, kernel methods are often not the first choice for large-scale speech applications due to their significant memory requirements and computational expense. In recent years, randomized approximate feature maps have emerged as an elegant mechanism to scale-up kernel methods. Still, in practice, a large number of random features is required to obtain acceptable accuracy in predictive tasks. In this paper, we develop two algorithmic schemes to address this computational bottleneck in the context of kernel ridge regression. The first scheme is a specialized distributed block coordinate descent procedure that avoids the explicit materialization of the feature space data matrix, while the second scheme gains efficiency by combining multiple weak random feature models in an ensemble learning framework. We demonstrate that these schemes enable kernel methods to match the performance of state of the art Deep Neural Networks on TIMIT for speech recognition and classification tasks. In particular, we obtain the best classification error rates reported on TIMIT using kernel methods. Po-Sen Huang, Haim Avron, Tara N. Sainath, Vikas Sindhwani, Bhuvana Ramabhadran |
ICASSP | 5 |
| 2014 | Deep Scattering Spectrum with deep neural networksabstractState-of-the-art convolutional neural networks (CNNs) typically use a log-mel spectral representation of the speech signal. However, this representation is limited by the spectro-temporal resolution afforded by log-mel filter-banks. A novel technique known as Deep Scattering Spectrum (DSS) addresses this limitation and preserves higher resolution information, while ensuring time warp stability, through the cascaded application of the wavelet-modulus operator. The first order scatter is equivalent to log-mel features and standard CNN modeling techniques can directly be used with these features. However the higher order scatter, which preserves the higher resolution information, presents new challenges in modelling. This paper explores how to effectively use DSS features with CNN acoustic models. Specifically, we identify the effective normalization, neural network topology and regularization techniques to effectively model higher order scatter. The use of these higher order scatter features, in conjunction with CNNs, results in relative improvement of 7% compared to log-mel features on TIMIT, providing a phonetic error rate (PER) of 17.4%, one of the lowest reported PERs to date on this task. Vijayaditya Peddinti, Tara N. Sainath, Shay Maymon, Bhuvana Ramabhadran, David Nahamoo, Vaibhava Goel |
ICASSP | 4 |
| 2014 | Improvements to filterbank and delta learning within a deep neural network frameworkabstractMany features used in speech recognition tasks are hand-crafted and are not always related to the objective at hand, that is minimizing word error rate. Recently, we showed that replacing a perceptually motivated mel-filter bank with a filter bank layer that is learned jointly with the rest of a deep neural network was promising. In this paper, we extend filter learning to a speaker-adapted, state-of-the-art system. First, we incorporate delta learning into the filter learning framework. Second, we incorporate various speaker adaptation techniques, including VTLN warping and speaker identity features. On a 50-hour English Broadcast News task, we show that we can achieve a 5% relative improvement in word error rate (WER) using the filter and delta learning, compared to having a fixed set of filters and deltas. Furthermore, after speaker adaptation, we find that filter and delta learning allows for a 3% relative improvement in WER compared to a state-of-the-art CNN. Tara N. Sainath, Brian Kingsbury, Abdel-rahman Mohamed, George Saon, Bhuvana Ramabhadran |
ICASSP | 5 |
| 2014 | Static interpolation of exponential n-gram models using features of featuresabstractThe best language model performance for a task is often achieved by interpolating language models built separately on corpora from multiple sources. While common practice is to use a single set of fixed interpolation weights to combine models, past work has found that gains can be had by allowing weights to vary by n-gram, when linearly interpolating word n-gram models. In this work, we investigate whether similar ideas can be used to improve log-linear interpolation for Model M, an exponential class-based n-gram model with state-of-the-art performance. We focus on log-linear interpolation as Model M's combined via (regular) linear interpolation cannot be statically compiled into a single model, as is required for many applications due to resource constraints. We present a general parameter interpolation framework in which a weight prediction model is used to compute the interpolation weights for each n-gram. The weight prediction model takes a rich representation of n-gram features as input, and is trained to optimize the perplexity of a held-out set. In experiments on Broadcast News, we show that a mixture of experts weight prediction model yields significant perplexity and word-error rate improvements as compared to static linear interpolation. Abhinav Sethy, Stanley F. Chen, Bhuvana Ramabhadran, Paul Vozila |
ICASSP | 3 |
| 2014 | Dictionary-based pitch tracking with dynamic programming
Ewout van den Berg, Bhuvana Ramabhadran |
INTERSPEECH | 2 |
| 2014 | Improving deep neural network acoustic modeling for audio corpus indexing under the IARPA babel program
Brian Kingsbury, Jia Cui, Bhuvana Ramabhadran, Andrew Rosenberg, Mohammad Sadegh Rasooli, Owen Rambow, Nizar Habash, Vaibhava Goel |
INTERSPEECH | 4 |
| 2014 | Recent improvements in neural network acoustic modeling for LVCSR in low resource languages
Jia Cui, Bhuvana Ramabhadran, Andrew Rosenberg, Brian Kingsbury, Abhinav Sethy |
INTERSPEECH | 2 |
| 2014 | Exploiting vocal-source features to improve ASR accuracy for low-resource languages
Raul Fernandez, Jia Cui, Andrew Rosenberg, Bhuvana Ramabhadran |
INTERSPEECH | 4 |
| 2014 | Prosody contour prediction with long short-term memory, bi-directional, deep recurrent neural networksabstractDeep Neural Networks (DNNs) have been shown to provide state-of-the-art performance over other baseline models in the task of predicting prosodic targets from text in a speechsynthesis system. However, prosody prediction can be affected by an interaction of short- and long-term contextual factors that a static model that depends on a fixed-size context window can fail to properly capture. In this work, we look at a recurrent formulation of neural networks (RNNs) that are deep in time and can store state information from an arbitrarily large input history when making a prediction. We show that RNNs provide improved performance over DNNs of comparable size in terms of various objective metrics for a variety of prosodic streams (notably, a relative reduction of about 6% in F0 mean-square error accompanied by a relative increase of about 14% in F0 variance), as well as in terms of perceptual quality assessed through mean-opinion-score listening tests. Index Terms: speech synthesis,text-to-speech, prosody prediction, recurrent neural networks, deep learning Raul Fernandez, Asaf Rendel, Bhuvana Ramabhadran, Ron Hoory |
INTERSPEECH | 3 |
| 2014 | Parallel deep neural network training for LVCSR tasks using blue gene/QabstractWhile Deep Neural Networks (DNNs) have achieved tremendous success for LVCSR tasks, training these networks is slow. To date, the most common approach to train DNNs is via stochastic gradient descent (SGD), serially on a single GPU machine. Serial training, coupled with the large number of training parameters and speech data set sizes, makes DNN training very slow for LVCSR tasks. While 2nd order, data-parallel methods have also been explored, these methods are not always faster on CPU clusters due to the large communication cost between processors. In this work, we explore using a specialized hardware/software approach, utilizing a Blue Gene/Q (BG/Q) system, which has thousands of processors and excellent interprocessor communication. We explore using the 2nd order Hessian-free (HF) algorithm for DNN training with BG/Q, for both cross-entropy and sequence training of DNNs. Results on three LVCSR tasks indicate that using HF with BG/Q offers up to an 11x speedup, as well as an improved word error rate (WER), compared to SGD on a GPU. Tara N. Sainath, I-Hsin Chung, Bhuvana Ramabhadran, Michael Picheny, John A. Gunnels, Brian Kingsbury, George Saon, Vernon Austel, Upendra V. Chaudhari |
INTERSPEECH | 3 |
| 2014 | Deep scattering spectra with deep neural networks for LVCSR tasksabstractLog-mel filterbank features, which are commonly used features for CNNs, can remove higher-resolution information from the speech signal. A novel technique, known as Deep Scattering Spectrum (DSS), addresses this issue and looks to preserve this information. DSS features have shown promise on TIMIT, both for classification and recognition. In this paper, we extend the use of DSS features for LVCSR tasks. First, we explore the optimal multi-resolution time and frequency scattering operations for LVCSR tasks. Next, we explore techniques to reduce the dimension of the DSS features. We also incorporate speaker adaptation techniques into the DSS features. Results on a 50 and 430 hour English Broadcast News task show that the DSS features provide between a 4-7% relative improvement in WER over log-mel features, within a state-of-the-art CNN framework which incorporates speaker-adaptation and sequence training. Finally, we show that DSS features are similar to multi-resolution log-mel + MFCCs, and similar improvements can be obtained with this representation. Tara N. Sainath, Vijayaditya Peddinti, Brian Kingsbury, Petr Fousek, Bhuvana Ramabhadran, David Nahamoo |
INTERSPEECH | 5 |
| 2014 | Parallel Deep Neural Network Training for Big Data on Blue Gene/QabstractDeep Neural Networks (DNNs) have recently been shown to significantly outperform existing machine learning techniques in several pattern recognition tasks. DNNs are the state-of-the-art models used in image recognition, object detection, classification and tracking, and speech and language processing applications. The biggest drawback to DNNs has been the enormous cost in computation and time taken to train the parameters of the networks - often a tenfold increase relative to conventional technologies. Such training time costs can be mitigated by the application of parallel computing algorithms and architectures. However, these algorithms often run into difficulties because of the cost of inter-processor communication bottlenecks. In this paper, we describe how to enable Parallel Deep Neural Network Training on the IBM Blue Gene/Q (BG/Q) computer system. Specifically, we explore DNN training using the data parallel Hessian-free 2nd order optimization algorithm. Such an algorithm is particularly well-suited to parallelization across a large set of loosely coupled processors. BG/Q, with its excellent inter-processor communication characteristics, is an ideal match for this type of algorithm. The paper discusses how issues regarding programming model and data-dependent imbalances are addressed. Results on large-scale speech tasks show that the performance on BG/Q scales linearly up to 4096 processes with no loss in accuracy. This allows us to train neural networks using billions of training examples in a few hours. I-Hsin Chung, Tara N. Sainath, Bhuvana Ramabhadran, Michael Picheny, John A. Gunnels, Vernon Austel, Upendra V. Chaudhari, Brian Kingsbury |
SC | 3 |
| 2014 | Editorial for the special issue on spoken content retrieval
Murat Saraclar, Ciprian Chelba, Bhuvana Ramabhadran |
Comput. Speech Lang. | 3 |
| 2014 | Converting Neural Network Language Models into Back-off Language Models for Efficient Decoding in Automatic Speech RecognitionabstractNeural network language models (NNLMs) have achieved very good performance in large-vocabulary continuous speech recognition (LVCSR) systems. Because decoding with NNLMs is computationally expensive, there is interest in developing methods to approximate NNLMs with simpler language models that are suitable for fast decoding. In this work, we propose an approximate method for converting a feedforward NNLM into a back-off n-gram language model that can be used directly in existing LVCSR decoders. We convert NNLMs of increasing order to pruned back-off language models, using lower-order models to constrain the n-grams allowed in higher-order models. In experiments on Broadcast News data, we find that the resulting back-off models retain the bulk of the gain achieved by NNLMs over conventional n-gram language models, and give accuracy improvements as compared to existing methods for converting NNLMs to back-off models. In addition, the proposed approach can be applied to any type of non-back-off language model to enable efficient decoding. Ebru Arisoy, Stanley F. Chen, Bhuvana Ramabhadran, Abhinav Sethy |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2013 | Accelerating Hessian-free optimization for Deep Neural Networks by implicit preconditioning and samplingabstractHessian-free training has become a popular parallel second order optimization technique for Deep Neural Network training. This study aims at speeding up Hessian-free training, both by means of decreasing the amount of data used for training, as well as through reduction of the number of Krylov subspace solver iterations used for implicit estimation of the Hessian. In this paper, we develop an L-BFGS based preconditioning scheme that avoids the need to access the Hessian explicitly. Since L-BFGS cannot be regarded as a fixed-point iteration, we further propose the employment of flexible Krylov subspace solvers that retain the desired theoretical convergence guarantees of their conventional counterparts. Second, we propose a new sampling algorithm, which geometrically increases the amount of data utilized for gradient and Krylov subspace iteration calculations. On a 50-hr English Broadcast News task, we find that these methodologies provide roughly a 1.5× speed-up, whereas, on a 300-hr Switchboard task, these techniques provide over a 2.3× speedup, with no loss in WER. These results suggest that even further speed-up is expected, as problems scale and complexity grows. Tara N. Sainath, Lior Horesh, Brian Kingsbury, Aleksandr Y. Aravkin, Bhuvana Ramabhadran |
ASRU | 5 |
| 2013 | Improvements to Deep Convolutional Neural Networks for LVCSRabstractDeep Convolutional Neural Networks (CNNs) are more powerful than Deep Neural Networks (DNN), as they are able to better reduce spectral variation in the input signal. This has also been confirmed experimentally, with CNNs showing improvements in word error rate (WER) between 4-12% relative compared to DNNs across a variety of LVCSR tasks. In this paper, we describe different methods to further improve CNN performance. First, we conduct a deep analysis comparing limited weight sharing and full weight sharing with state-of-the-art features. Second, we apply various pooling strategies that have shown improvements in computer vision to an LVCSR speech task. Third, we introduce a method to effectively incorporate speaker adaptation, namely fMLLR, into log-mel features. Fourth, we introduce an effective strategy to use dropout during Hessian-free sequence training. We find that with these improvements, particularly with fMLLR and dropout, we are able to achieve an additional 2-3% relative improvement in WER on a 50-hour Broadcast News task over our previous best CNN baseline. On a larger 400-hour BN task, we find an additional 4-5% relative improvement over our previous best CNN baseline. Tara N. Sainath, Brian Kingsbury, Abdel-rahman Mohamed, George E. Dahl, George Saon, Hagen Soltau, Tomás Beran, Aleksandr Y. Aravkin, Bhuvana Ramabhadran |
ASRU | 9 |
| 2013 | Learning filter banks within a deep neural network frameworkabstractMel-filter banks are commonly used in speech recognition, as they are motivated from theory related to speech production and perception. While features derived from mel-filter banks are quite popular, we argue that this filter bank is not really an appropriate choice as it is not learned for the objective at hand, i.e. speech recognition. In this paper, we explore replacing the filter bank with a filter bank layer that is learned jointly with the rest of a deep neural network. Thus, the filter bank is learned to minimize cross-entropy, which is more closely tied to the speech recognition objective. On a 50-hour English Broadcast News task, we show that we can achieve a 5% relative improvement in word error rate (WER) using the filter bank learning approach, compared to having a fixed set of filters. Tara N. Sainath, Brian Kingsbury, Abdel-rahman Mohamed, Bhuvana Ramabhadran |
ASRU | 4 |
| 2013 | An empirical study of confusion modeling in keyword search for low resource languagesabstractKeyword search, in the context of low resource languages, has emerged as a key area of research. The dominant approach in keyword search is to use Automatic Speech Recognition (ASR) as a front end to produce a representation of audio that can be indexed. The biggest drawback of this approach lies in its the inability to deal with out-of-vocabulary words and query terms that are not in the ASR system output. In this paper we present an empirical study evaluating various approaches based on using confusion models as query expansion techniques to address this problem. We present results across four languages using a range of confusion models which lead to significant improvements in keyword search performance as measured by the Maximum Term Weighted Value (MTWV) metric. Murat Saraclar, Abhinav Sethy, Bhuvana Ramabhadran, Lidia Mangu, Jia Cui, Brian Kingsbury, Jonathan Mamou |
ASRU | 3 |
| 2013 | Joint training of interpolated exponential n-gram modelsabstractFor many speech recognition tasks, the best language model performance is achieved by collecting text from multiple sources or domains, and interpolating language models built separately on each individual corpus. When multiple corpora are available, it has also been shown that when using a domain adaptation technique such as feature augmentation [1], the performance on each individual domain can be improved by training a joint model across all of the corpora. In this paper, we explore whether improving each domain model via joint training also improves performance when interpolating the models together. We show that the diversity of the individual models is an important consideration, and propose a method for adjusting diversity to optimize overall performance. We present results using word n-gram models and Model M, a class-based n-gram model, and demonstrate improvements in both perplexity and word-error rate relative to state-of-the-art results on a Broadcast News transcription task. Abhinav Sethy, Stanley F. Chen, Ebru Arisoy, Bhuvana Ramabhadran, Kartik Audhkhasi, Shri Narayanan, Paul Vozila |
ASRU | 4 |
| 2013 | Converting Neural Network Language Models into back-off language models for efficient decoding in automatic speech recognitionabstractNeural Network Language Models (NNLMs) have achieved very good performance in large-vocabulary continuous speech recognition (LVCSR) systems. Because decoding with NNLMs is very computationally expensive, there is interest in developing methods to approximate NNLMs with simpler language models that are suitable for fast decoding. In this work, we propose an approximate method for converting a feedforward NNLM into a back-off n-gram language model that can be used directly in existing LVCSR decoders. We convert NNLMs of increasing order to pruned back-off language models, using lower-order models to constrain the n-grams allowed in higher-order models. In experiments on Broadcast News data, we find that the resulting back-off models retain the bulk of the gain achieved by NNLMs over conventional n-gram language models, and give significant accuracy improvements as compared to existing methods for converting NNLMs to back-off models. In addition, the proposed approach can be applied to any type of non-back-off language model to enable efficient decoding. Ebru Arisoy, Stanley F. Chen, Bhuvana Ramabhadran, Abhinav Sethy |
ICASSP | 3 |
| 2013 | Developing speech recognition systems for corpus indexing under the IARPA Babel programabstractAutomatic speech recognition is a core component of many applications, including keyword search. In this paper we describe experiments on acoustic modeling, language modeling, and decoding for keyword search on a Cantonese conversational telephony corpus collected as part of the IARPA Babel program. We show that acoustic modeling techniques such as the bootstrapped-and-restructured model and deep neural network acoustic model significantly outperform a state-of-the-art baseline GMM/HMM model, in terms of both recognition performance and keyword search performance, with improvements of up to 11% relative character error rate reduction and 31% relative maximum term weighted value improvement. We show that while an interpolated Model M and neural network LM improve recognition performance, they do not improve keyword search results; however, the advanced LM does reduce the size of the keyword search index. Finally, we show that a simple form of automatically adapted keyword search performs 16% better than a preindexed search system, indicating that out-of-vocabulary search is still a challenge. Jia Cui, Bhuvana Ramabhadran, Janice Kim, Brian Kingsbury, Jonathan Mamou, Lidia Mangu, Michael Picheny, Tara N. Sainath, Abhinav Sethy |
ICASSP | 3 |
| 2013 | F0 contour prediction with a deep belief network-Gaussian process hybrid modelabstractIn this work we look at using non-parametric, exemplar-based regression for the prediction of prosodic contour targets from textual features in a speech synthesis system. We investigate the performance of Gaussian Process regression on this task when the covariance kernel operates on a variety of input feature spaces. In particular, we consider non-linear features extracted via Deep Belief Networks. We motivate the use of this hybrid model by considering the initial deep-layer model as a feature extractor that can summarize high-level structure from the raw inputs to improve the regression of an exemplar-based model in the second part of the approach. By looking at both objective metrics and perceptual listening tests, we evaluate these proposals against each other, and against the standard clustering-tree techniques implemented in parametric synthesis for the prediction of prosodic targets. Raul Fernandez, Asaf Rendel, Bhuvana Ramabhadran, Ron Hoory |
ICASSP | 3 |
| 2013 | A high-performance Cantonese keyword search systemabstractWe present a system for keyword search on Cantonese conversational telephony audio, collected for the IARPA Babel program, that achieves good performance by combining postings lists produced by diverse speech recognition systems from three different research groups. We describe the keyword search task, the data on which the work was done, four different speech recognition systems, and our approach to system combination for keyword search. We show that the combination of four systems outperforms the best single system by 7%, achieving an actual term-weighted value of 0.517. Brian Kingsbury, Jia Cui, Mark J. F. Gales, Kate M. Knill, Jonathan Mamou, Lidia Mangu, David Nolden, Michael Picheny, Bhuvana Ramabhadran, Ralf Schlüter, Abhinav Sethy, Philip C. Woodland |
ICASSP | 10 |
| 2013 | System combination and score normalization for spoken term detectionabstractSpoken content in languages of emerging importance needs to be searchable to provide access to the underlying information. In this paper, we investigate the problem of extending data fusion methodologies from Information Retrieval for Spoken Term Detection on low-resource languages in the framework of the IARPA Babel program. We describe a number of alternative methods improving keyword search performance. We apply these methods to Cantonese, a language that presents some new issues in terms of reduced resources and shorter query lengths. First, we show score normalization methodology that improves in average by 20% keyword search performance. Second, we show that properly combining the outputs of diverse ASR systems performs 14% better than the best normalized ASR system. Jonathan Mamou, Jia Cui, Mark J. F. Gales, Brian Kingsbury, Kate M. Knill, Lidia Mangu, David Nolden, Michael Picheny, Bhuvana Ramabhadran, Ralf Schlüter, Abhinav Sethy, Philip C. Woodland |
ICASSP | 10 |
| 2013 | An evaluation of posterior modeling techniques for phonetic recognitionabstractSeveral methods have been proposed recently for modeling posterior representations derived from local classifiers [1, 2]. In recent work, Sainath et al. have proposed the use of a tied-mixture-based posterior modeling approach [3] to enhance exemplar-based posterior representations for phone recognition tasks. In this work, we conduct a detailed evaluation to determine the effectiveness of this technique on three representative posterior systems. In addition, we propose and evaluate an alternative discriminative formulation of the posterior modeling objective function that seeks to minimize framelevel errors. In experimental evaluations on the TIMIT corpus, we find that posterior modeling results in relative phone error rate (PER) reductions of between 1.1-5.5% across the systems tested. In fact, using Spif-NN [4, 3] posteriors, we are able to achieve a PER of 18.5; to the best of our knowledge, this is the best result reported in the literature to date. minimize framelevel errors. Rohit Prabhavalkar, Tara N. Sainath, David Nahamoo, Bhuvana Ramabhadran, Dimitri Kanevsky |
ICASSP | 4 |
| 2013 | Low-rank matrix factorization for Deep Neural Network training with high-dimensional output targetsabstractWhile Deep Neural Networks (DNNs) have achieved tremendous success for large vocabulary continuous speech recognition (LVCSR) tasks, training of these networks is slow. One reason is that DNNs are trained with a large number of training parameters (i.e., 10–50 million). Because networks are trained with a large number of output targets to achieve good performance, the majority of these parameters are in the final weight layer. In this paper, we propose a low-rank matrix factorization of the final weight layer. We apply this low-rank technique to DNNs for both acoustic modeling and language modeling. We show on three different LVCSR tasks ranging between 50–400 hrs, that a low-rank factorization reduces the number of parameters of the network by 30–50%. This results in roughly an equivalent reduction in training time, without a significant loss in final recognition accuracy, compared to a full-rank representation. Tara N. Sainath, Brian Kingsbury, Vikas Sindhwani, Ebru Arisoy, Bhuvana Ramabhadran |
ICASSP | 5 |
| 2013 | Deep convolutional neural networks for LVCSRabstractConvolutional Neural Networks (CNNs) are an alternative type of neural network that can be used to reduce spectral variations and model spectral correlations which exist in signals. Since speech signals exhibit both of these properties, CNNs are a more effective model for speech compared to Deep Neural Networks (DNNs). In this paper, we explore applying CNNs to large vocabulary speech tasks. First, we determine the appropriate architecture to make CNNs effective compared to DNNs for LVCSR tasks. Specifically, we focus on how many convolutional layers are needed, what is the optimal number of hidden units, what is the best pooling strategy, and the best input feature type for CNNs. We then explore the behavior of neural network features extracted from CNNs on a variety of LVCSR tasks, comparing CNNs to DNNs and GMMs. We find that CNNs offer between a 13-30% relative improvement over GMMs, and a 4-12% relative improvement over DNNs, on a 400-hr Broadcast News and 300-hr Switchboard task. Tara N. Sainath, Abdel-rahman Mohamed, Brian Kingsbury, Bhuvana Ramabhadran |
ICASSP | 4 |
| 2013 | Optimization Techniques to Improve Training Speed of Deep Neural Networks for Large Speech TasksabstractWhile Deep Neural Networks (DNNs) have achieved tremendous success for large vocabulary continuous speech recognition (LVCSR) tasks, training these networks is slow. Even to date, the most common approach to train DNNs is via stochastic gradient descent, serially on one machine. Serial training, coupled with the large number of training parameters (i.e., 10-50 million) and speech data set sizes (i.e., 20-100 million training points) makes DNN training very slow for LVCSR tasks. In this work, we explore a variety of different optimization techniques to improve DNN training speed. This includes parallelization of the gradient computation during cross-entropy and sequence training, as well as reducing the number of parameters in the network using a low-rank matrix factorization. Applying the proposed optimization techniques, we show that DNN training can be sped up by a factor of 3 on a 50-hour English Broadcast News (BN) task with no loss in accuracy. Furthermore, using the proposed techniques, we are able to train DNNs on a 300-hr Switchboard (SWB) task and a 400-hr English BN task, showing improvements between 9-30% relative over a state-of-the art GMM/HMM system while the number of parameters of the DNN is smaller than the GMM/HMM system. Tara N. Sainath, Brian Kingsbury, Hagen Soltau, Bhuvana Ramabhadran |
IEEE Trans. Speech Audio Process. | 4 |
| 2012 | Creating ensemble of diverse maximum entropy modelsabstractDiversity of a classifier ensemble has been shown to benefit overall classification performance. But most conventional methods of training ensembles offer no control on the extent of diversity and are meta-learners. We present a method for creating an ensemble of diverse maximum entropy (∂MaxEnt) models, which are popular in speech and language processing. We modify the objective function for conventional training of a MaxEnt model such that its output posterior distribution is diverse with respect to a reference model. Two diversity scores are explored - KL divergence and posterior cross-correlation. Experiments on the CoNLL-2003 Named Entity Recognition task and the IEMOCAP emotion recognition database show the benefits of a ∂MaxEnt ensemble. Kartik Audhkhasi, Abhinav Sethy, Bhuvana Ramabhadran, Shri Narayanan |
ICASSP | 3 |
| 2012 | Prediction of F0 contours from symbolic and numerical variables using continuous conditional random fieldsabstractRegression of continuous-valued variables as a function of both categorical and continuous predictors arises in some areas of speech processing, such as when predicting prosodic targets in a text-to-speech system. In this work we investigate the use of Continuous Conditional Random Fields (CCRF) to conditionally predict F0 targets from a series of s symbolic and numerical predictive features derived from text. We derive the training equations for the model using a Least-Squares-Error criterion within a supervised framework, and evaluate the proposed system using this objective criterion against other baseline models that can handle mixed inputs, such as regression trees and ensemble of regression trees. Raul Fernandez, Steve Minnis, Bhuvana Ramabhadran |
ICASSP | 3 |
| 2012 | Constructing ensembles of dissimilar acoustic models using hidden attributes of training dataabstractOne of the objectives in acoustic modeling is to realize robust statistical models against the wide variety of acoustic conditions that are present in real world environments. As large amounts of training data become available, modeling subsets of the data with similar acoustic qualities can be done accurately and multiple acoustic models are jointly used as a form of system combination or model selection. In this paper, we propose a method to partition the training data for constructing ensembles of acoustic models using metadata attributes such as SNR, speaking rate, and duration via a binary tree. The metadata attribute used at each binary split in the decision tree is obtained using a metric proposed in this paper that is cosine-similarity based. The resulting multiple models are combined using voting techniques such as n-best ROVER. The proposed method improved the recognition accuracy by up to 4% relative over the state-of-the-art system on a large vocabulary continuous speech recognition voice search task. Takashi Fukuda, Ryuki Tachibana, Upendra V. Chaudhari, Bhuvana Ramabhadran, Puming Zhan |
ICASSP | 4 |
| 2012 | N-best entropy based data selection for acoustic modelingabstractThis paper presents a strategy for efficiently selecting informative data from large corpora of untranscribed speech. Confidence-based selection methods (i.e., selecting utterances we are least confident about) have been a popular approach, though they only look at the top hypothesis when selecting utterances and tend to select outliers, therefore, not always improving overall recognition accuracy. Alternatively, we propose a method for selecting data looking at competing hypothesis by computing entropy of N-best hypothesis decoded by the baseline acoustic model. In addition we address the issue of outliers by calculating how representative a specific utterance is to all other unselected utterances via a tf-idf score. Experiments show that N-best entropy based selection (%relative 5.8 in 400-hour corpus) outperformed other conventional selection strategies; confidence based and lattice entropy based, and that tf-idf based representativeness improved the model further (%relative 6.2). A comparison with random selection is also presented. Finally model size impact is discussed. Nobuyasu Itoh, Tara N. Sainath, Dan-Ning Jiang, Bhuvana Ramabhadran |
ICASSP | 5 |
| 2012 | Improved pre-training of Deep Belief Networks using Sparse Encoding Symmetric MachinesabstractRestricted Boltzmann Machines (RBM) continue to be a popular methodology to pre-train weights of Deep Belief Networks (DBNs). However, the RBM objective function cannot be maximized directly. Therefore, it is not clear what function to monitor when deciding to stop the training, leading to a challenge in managing the computational costs. The Sparse Encoding Symmetric Machine (SESM) has been suggested as an alternative method for pre-training. By placing a sparseness term on the NN output codebook, SESM allows the objective function to be optimized directly and reliably be monitored as an indicator to stop the training. In this paper, we explore SESM to pre-train DBNs and apply this the first time to speech recognition. First, we provide a detailed analysis comparing the behavior of SESM and RBM. Second, we compare the performance of SESM pre-trained and RBM pre-trained DBNs on TIMIT and a 50 hour English Broadcast News task. Results indicate that pre-trained DBNs using SESM and RBMs achieve comparable performance and outperform randomly initialized DBNs with SESM providing a much easier stopping criterion relative to RBM. Christian Plahl, Tara N. Sainath, Bhuvana Ramabhadran, David Nahamoo |
ICASSP | 3 |
| 2012 | Auto-encoder bottleneck features using deep belief networksabstractNeural network (NN) bottleneck (BN) features are typically created by training a NN with a middle bottleneck layer. Recently, an alternative structure was proposed which trains a NN with a constant number of hidden units to predict output targets, and then reduces the dimensionality of these output probabilities through an auto-encoder, to create auto-encoder bottleneck (AE-BN) features. The benefit of placing the BN after the posterior estimation network is that it avoids the loss in frame classification accuracy incurred by networks that place the BN before the softmax. In this work, we investigate the use of pre-training when creating AE-BN features. Our experiments indicate that with the AE-BN architecture, pre-trained and deeper NNs produce better AE-BN features. On a 50-hour English Broadcast News task, the AE-BN features provide over a 1% absolute improvement compared to a state-of-the-art GMM/HMM with a WER of 18.8% and pre-trained NN hybrid system with a WER of 18.4%. In addition, on a larger 430-hour Broadcast News task, AE-BN features provide a 0.5% absolute improvement over a strong GMM/HMM baseline with a WER of 16.0%. Finally, system combination with the GMM/HMM baseline and AE-BN systems provides an additional 0.5% absolute on 430 hours over the AE-BN system alone, yielding a final WER of 15.0%. Tara N. Sainath, Brian Kingsbury, Bhuvana Ramabhadran |
ICASSP | 3 |
| 2012 | Phrase Boundary Assignment from Text in Multiple DomainsabstractDetecting and modeling proper phrasing from an input text string is an important aspect when producing synthesis that sounds intelligible and natural. Knowledge of proper phrase structure influences, e.g., the placement and length of pauses, and the realization of phrase-final boundary contours, both of which can have an effect in a listener’s percepts ranging from naturalness to semantic interpretation. In this work, we look at modeling the occurrence, and types, of phrase breaks from purely textual features, paying close attention to how the performance of the systems generalizes inand out-of-domain for corpora of various types (such as broadcast news, spontaneous speech, and synthesis databases), and as a function of various subsets of syntactical and lexical features investigated. Andrew Rosenberg, Raul Fernandez, Bhuvana Ramabhadran |
INTERSPEECH | 3 |
| 2012 | Enhancing Exemplar-Based Posteriors for Speech Recognition Tasks
Tara N. Sainath, David Nahamoo, Dimitri Kanevsky, Bhuvana Ramabhadran |
INTERSPEECH | 4 |
| 2012 | Leveraging word confusion networks for named entity modeling and detection from conversational telephone speech
Gakuto Kurata, Nobuyasu Itoh, Masafumi Nishimura, Abhinav Sethy, Bhuvana Ramabhadran |
Speech Commun. | 5 |
| 2012 | Acoustically discriminative language model training with pseudo-hypothesis
Gakuto Kurata, Abhinav Sethy, Bhuvana Ramabhadran, Ariya Rastrow, Nobuyasu Itoh, Masafumi Nishimura |
Speech Commun. | 3 |
| 2011 | Pruning exponential language modelsabstractLanguage model pruning is an essential technology for speech applications running on resource-constrained devices, and many pruning algorithms have been developed for conventional word n-gram models. However, while exponential language models can give superior performance, there has been little work on the pruning of these models. In this paper, we propose several pruning algorithms for general exponential language models. We show that our best algorithm applied to an exponential n-gram model outperforms existing n-gram model pruning algorithms by up to 0.4% absolute in speech recognition word-error rate on Wall Street Journal and Broadcast News data sets. In addition, we show that Model M, an exponential class-based language model, retains its performance improvement over conventional word n-gram models when pruned to equal size, with gains of up to 2.5% absolute in word-error rate. Stanley F. Chen, Abhinav Sethy, Bhuvana Ramabhadran |
ASRU | 3 |
| 2011 | Making Deep Belief Networks effective for large vocabulary continuous speech recognitionabstractTo date, there has been limited work in applying Deep Belief Networks (DBNs) for acoustic modeling in LVCSR tasks, with past work using standard speech features. However, a typical LVCSR system makes use of both feature and model-space speaker adaptation and discriminative training. This paper explores the performance of DBNs in a state-of-the-art LVCSR system, showing improvements over Multi-Layer Perceptrons (MLPs) and GMM/HMMs across a variety of features on an English Broadcast News task. In addition, we provide a recipe for data parallelization of DBN training, showing that data parallelization can provide linear speed-up in the number of machines, without impacting WER. Tara N. Sainath, Brian Kingsbury, Bhuvana Ramabhadran, Petr Fousek, Abdel-rahman Mohamed |
ASRU | 3 |
| 2011 | A convex hull approach to sparse representations for exemplar-based speech recognitionabstractIn this paper, we propose a novel exemplar based technique for classification problems where for every new test sample the classification model is re-estimated from a subset of relevant samples of the training data.We formulate the exemplar-based classification paradigm as a sparse representation (SR) problem, and explore the use of convex hull constraints to enforce both regularization and sparsity. Finally, we utilize the Extended Baum-Welch (EBW) optimization technique to solve the SR problem. We explore our proposed methodology on the TIMIT phonetic classification task, showing that our proposed method offers statistically significant improvements over common classification methods, and provides an accuracy of 82.9%, the best single-classifier number reported to date. Tara N. Sainath, David Nahamoo, Dimitri Kanevsky, Bhuvana Ramabhadran, Parikshit M. Shah |
ASRU | 4 |
| 2011 | Frame-level AnyBoost for LVCSR with the MMI CriterionabstractThis paper propose a variant of AnyBoost for a large vocabulary continuous speech recognition (LVCSR) task. AnyBoost is an efficient algorithm to train an ensemble of weak learners by gradient descent for an objective function.We present a novel training procedure that trains acoustic models via the MMI criterion using data that is weighted proportional to the summation of the posterior functions of previous round of weak learners. Optimized for system combination by n-best ROVER at runtime, data weights for a new weak learner are computed as a weighted summation of posteriors of previous weak learners. We compare a frame-based version and a sentence-based version of our proposed algorithm with a frame-based AdaBoost algorithm. We will present results on a voice search task trained with different amounts of data with gains of 5.1% to 7.5% relative in WER can be obtained by three rounds of boosting. Ryuki Tachibana, Takashi Fukuda, Upendra V. Chaudhari, Bhuvana Ramabhadran, Puming Zhan |
ASRU | 4 |
| 2011 | Exploiting active-learning strategies for annotating prosodic events with limited labeled dataabstractMany applications of spoken-language systems can benefit from having access to annotations of prosodic events. Unfortunately, obtaining human annotations of these events, even sensible amounts to train a supervised system, can become a laborious and costly effort. Given these constraints, this task serves as a good case study for approaches that judiciously guide the selection of data in order to maximize the gain from the human-labeling process or which minimize the size of the training set. To address this, we explore active learning techniques with the objective of reducing the amount of human-annotated data needed to attain a given level of performance. We review strategies that can be used to guide the selection of sequences by combining the output of a classifier and information about the structure of the data into a criterion that can be used during the learning process to query the label of data points that are both informative and representative of the task, and show that for most of the cases considered, active selection strategies when labeling pitch accents and prosodic boundaries are as good as or exceed the performance of random data selection. Raul Fernandez, Bhuvana Ramabhadran |
ICASSP | 2 |
| 2011 | A-Functions: A generalization of Extended Baum-Welch transformations to convex optimizationabstractWe introduce the Line Search A-Function (LSAF) technique that generalizes the Extended-Baum Welch technique in order to provide an effective optimization technique for a broader set of functions. We show how LSAF can be applied to functions of various probability density and distribution functions by demonstrating that these probability functions have an A-function. We also show that sparse representation problems (SR) that use 11 or combination of 11/12 regularization norms can also be efficiently optimized through an A-function derived for their objective functions. We will demonstrate the efficiency of LSAF for SR problems through simulations by comparing it with Approximate Bayesian Compressive Sensing method that we recently applied to speech recognition. Dimitri Kanevsky, David Nahamoo, Tara N. Sainath, Bhuvana Ramabhadran, Peder A. Olsen |
ICASSP | 4 |
| 2011 | Named entity recognition from Conversational Telephone Speech leveraging Word Confusion Networks for training and recognitionabstractNamed Entity (NE) recognition from the results of Automatic Speech Recognition (ASR) is challenging because of ASR errors. To detect NEs, one of the options is to use a statistical NE model that is usually trained with ASR one-best results. In order to make NE recognition more robust to ASR errors, we propose using Word Confusion Networks (WCNs), sequences of bundled words, for both NE modeling and recognition by regarding the word bundles as units instead of the independent words. This is done by clustering similar word bundles that may originate from the same word. We trained the NE models with the maximum entropy principle and evaluated the performance using real-life call-center data. The results showed that by using the WCNs, the error of NE recognition was relatively reduced by up to 33.0%. Gakuto Kurata, Nobuyasu Itoh, Masafumi Nishimura, Abhinav Sethy, Bhuvana Ramabhadran |
ICASSP | 5 |
| 2011 | Deep Belief Networks using discriminative features for phone recognitionabstractDeep Belief Networks (DBNs) are multi-layer generative models. They can be trained to model windows of coefficients extracted from speech and they discover multiple layers of features that capture the higher-order statistical structure of the data. These features can be used to initialize the hidden units of a feed-forward neural network that is then trained to predict the HMM state for the central frame of the window. Initializing with features that are good at generating speech makes the neural network perform much better than initializing with random weights. DBNs have already been used successfully for phone recognition with input coefficients that are MFCCs or filterbank outputs. In this paper, we demonstrate that they work even better when their inputs are speaker adaptive, discriminative features. On the standard TIMIT corpus, they give phone error rates of 19.6% using monophone HMMs and a bigram language model and 19.4% using monophone HMMs and a trigram language model. Abdel-rahman Mohamed, Tara N. Sainath, George E. Dahl, Bhuvana Ramabhadran, Geoffrey E. Hinton, Michael Picheny |
ICASSP | 4 |
| 2011 | Hill climbing on speech lattices: A new rescoring frameworkabstractWe describe a new approach for rescoring speech lattices - with long-span language models or wide-context acoustic models - that does not entail computationally intensive lattice expansion or limited rescoring of only an N-best list. We view the set of word-sequences in a lattice as a discrete space equipped with the edit-distance metric, and develop a hill climbing technique to start with, say, the 1-best hypothesis under the lattice-generating model(s) and iteratively search a local neighborhood for the highest-scoring hypothesis under the rescoring model(s); such neighborhoods are efficiently constructed via finite state techniques. We demonstrate empirically that to achieve the same reduction in error rate using a better estimated, higher order language model, our technique evaluates fewer utterance-length hypotheses than conventional N-best rescoring by two orders of magnitude. For the same number of hypotheses evaluated, our technique results in a significantly lower error rate. Ariya Rastrow, Markus Dreyer, Abhinav Sethy, Sanjeev Khudanpur, Bhuvana Ramabhadran, Mark Dredze |
ICASSP | 5 |
| 2011 | Exemplar-based Sparse Representation phone identification featuresabstractExemplar-based techniques, such as k-nearest neighbors (kNNs) and Sparse Representations (SRs), can be used to model a test sample from a few training points in a dictionary set. In past work, we have shown that using a SR approach for phonetic classification allows for a higher accuracy than other classification techniques. These phones are the basic units of speech to be recognized. Motivated by this result, we create a new dictionary which is a function of the phonetic labels of the original dictionary. The SR method now selects relevant samples from this new dictionary to create a new feature representation of the test sample, where the new feature is better linked to the actual units to be recognized. We will refer to these new features as Spif. We present results using these new Spiffeatures in a Hidden Markov Model (HMM) framework for speech recognition. We find that the Spiffeatures allow for a 2.9% relative reduction in Phonetic Error Rate (PER) on the TIMIT phonetic recognition task. Furthermore, we find that the Spiffeatures allow for a 4.8% relative improvement in Word Error Rate (WER) on a large vocabulary 50 hour Broadcast News task. Tara N. Sainath, David Nahamoo, Bhuvana Ramabhadran, Dimitri Kanevsky, Vaibhava Goel, Parikshit M. Shah |
ICASSP | 3 |
| 2011 | Deep belief nets for natural language call-routingabstractThis paper considers application of Deep Belief Nets (DBNs) to natural language call routing. DBNs have been successfully applied to a number of tasks, including image, audio and speech classification, thanks to the recent discovery of an efficient learning technique. DBNs learn a multi-layer generative model from unlabeled data and the features discovered by this model are then used to initialize a feed-forward neural network which is fine-tuned with backpropagation. We compare a DBN-initialized neural network to three widely used text classification algorithms; Support Vector machines (SVM), Boosting and Maximum Entropy (MaxEnt). The DBN-based model gives a call-routing classification accuracy that is equal to the best of the other models even though it currently uses an impoverished representation of the input. Ruhi Sarikaya, Geoffrey E. Hinton, Bhuvana Ramabhadran |
ICASSP | 3 |
| 2011 | Distributed training of large scale exponential language modelsabstractShrinkage-based exponential language models, such as the recently introduced Model M, have provided significant gains over a range of tasks [1]. Training such models requires a large amount of computational resources in terms of both time and memory. In this paper, we present a distributed training algorithm for such models based on the idea of cluster expansion [2]. Cluster expansion allows us to efficiently calculate the normalization and expectations terms required for Model M training by minimizing the computation needed between consecutive n-grams. We also show how the algorithm can be implemented in a distributed environment, greatly reducing the memory required per process and training time. Abhinav Sethy, Stanley F. Chen, Bhuvana Ramabhadran |
ICASSP | 3 |
| 2011 | Speech processing and retrieval in a personal memory aid system for the elderlyabstractThe paper presents a new application of automatic speech processing in the Ambient Assisted Living area, developed in the course of a three year research project. Recording and automatic processing of spoken conversations plays a major role in this solution enabling effective search in a personal audio archive and fast browsing of conversations. Processing of elderly conversational speech recorded by a distant PDA microphone poses a great challenge. The speech processing flow includes transcription, speaker tracking and combined indexing and search of spoken terms and participating speakers identity extracted from the audio. We present the entire application and individual speech processing components as well as evaluation results of the individual components and of the end-to-end spoken information retrieval solution. Alexander Sorin, Hagai Aronowitz, Jonathan Mamou, Orith Toledo-Ronen, Ron Hoory, Michael Kuritzky, Yael Erez, Bhuvana Ramabhadran, Abhinav Sethy |
ICASSP | 8 |
| 2011 | Application specific loss minimization using gradient boostingabstractGradient boosting is a flexible machine learning technique that produces accurate predictions by combining many weak learners. In this work, we investigate its use in two applications, where we show the advantage of loss functions that are designed specifically for optimizing application objectives. We also extend the original gradient boosting algorithm with Newton-Raphson method to speed up learning. In the experiments, we demonstrate that the use of gradient boosting and application specific loss functions results in a relative improvement of 0.8% over an 82.6% baseline on the CoNLL 2003 named entity recognition task. We also show that this novel framework is useful in identifying regions of high word error rate (WER) and can provide up to 20% relative improvement depending on the chosen operating point. Bin Zhang 0009, Abhinav Sethy, Tara N. Sainath, Bhuvana Ramabhadran |
ICASSP | 4 |
| 2011 | Feature Combination Approaches for Discriminative Language Models
Ebru Arisoy, Bhuvana Ramabhadran, Hong-Kwang Jeff Kuo |
INTERSPEECH | 2 |
| 2011 | Convergence of Line Search A-Function Methods
Dimitri Kanevsky, David Nahamoo, Tara N. Sainath, Bhuvana Ramabhadran |
INTERSPEECH | 4 |
| 2011 | Improved Spoken Query Transcription Using Co-Occurrence InformationabstractSpoken queries are a natural medium for searching the Mobile Web. Language modeling for voice search recognition offers different challenges compared to more conventional speech applications. The challenges arise from the fact that spoken queries are usually a set of keywords and do not have a syntactic and grammatical structure. This paper describes a cooccurrence based approach to improve the accuracy of voice queries automatic transcription. With the right choice of scoring function and co-occurrence level, we show that co-occurrence information gives a 2% relative accuracy improvement over a state of the art system. Jonathan Mamou, Abhinav Sethy, Bhuvana Ramabhadran, Ron Hoory, Paul Vozila |
INTERSPEECH | 3 |
| 2011 | Clustering with Modified Cosine Distance Learned from Constraints
Leonid Rachevsky, Dimitri Kanevsky, Ruhi Sarikaya, Bhuvana Ramabhadran |
INTERSPEECH | 4 |
| 2011 | "What is... Dengue Fever?" - Modeling and Predicting Pronunciation Errors in a Text-to-Speech SystemabstractWe propose a system to predict baseform-generation errors in a text-to-speech (TTS) front-end, and aid in the process of cus-tomizing the synthesis engine to a novel application with a large, open-ended vocabulary. We motivate the use of the sys-tem by using data collected during the deployment of the IBM TTS engine in the Watson Deep Question-Answering system customized to play a game of Jeopardy!. We propose a set of features derived from a lexeme’s orthography and candidate baseform, and use a variety of learning schemes and data sam-pling algorithms to address the issue of skewed class priors in the training data. We show that 1) these different approaches provide complementary information that can then be exploited by fusion schemes to improve on the baseline performances, and 2) it is possible to use these techniques to retrieve a list of likely incorrect lexemes so as to reduce the number of tokens that must be vetted before finding and fixing an error. Index Terms: front-end error modeling, speech synthesis 1. Andrew Rosenberg, Raul Fernandez, Bhuvana Ramabhadran |
INTERSPEECH | 3 |
| 2011 | Reducing Computational Complexities of Exemplar-Based Sparse Representations with Applications to Large Vocabulary Speech Recognition
Tara N. Sainath, Bhuvana Ramabhadran, David Nahamoo, Dimitri Kanevsky |
INTERSPEECH | 2 |
| 2011 | Shrinkage-Based Features for Natural Language Call Routing
Ruhi Sarikaya, Stanley F. Chen, Bhuvana Ramabhadran |
INTERSPEECH | 3 |
| 2011 | Exemplar-Based Sparse Representation Features: From TIMIT to LVCSRabstractThe use of exemplar-based methods, such as support vector machines (SVMs), k-nearest neighbors (kNNs) and sparse representations (SRs), in speech recognition has thus far been limited. Exemplar-based techniques utilize information about individual training examples and are computationally expensive, making it particularly difficult to investigate these methods on large-vocabulary continuous speech recognition (LVCSR) tasks. While research in LVCSR provides a good testbed to tackle real-world speech recognition problems, research in this area suffers from two main drawbacks. First, the overall complexity of an LVCSR system makes error analysis quite difficult. Second, exploring new research ideas on LVCSR tasks involves training and testing state-of-the-art LVCSR systems, which can render a large turnaround time. This makes a small vocabulary task such as TIMIT more appealing. TIMIT provides a phonetically rich and hand-labeled corpus that allows easy insight into new algorithms. However, research ideas explored for small vocabulary tasks do not always provide gains on LVCSR systems. In this paper, we combine the advantages of using both small and large vocabulary tasks by taking well-established techniques used in LVCSR systems and applying them on TIMIT to establish a new baseline. We then utilize these existing LVCSR techniques in creating a novel set of exemplar-based sparse representation (SR) features. Using these existing LVCSR techniques, we achieve a phonetic error rate (PER) of 19.4% on the TIMIT task. The additional use of SR features reduce the PER to 18.6%. We then explore applying the SR features to a large vocabulary Broadcast News task, where we achieve a 0.3% absolute reduction in word error rate (WER). Tara N. Sainath, Bhuvana Ramabhadran, Michael Picheny, David Nahamoo, Dimitri Kanevsky |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2010 | Kalman filtering for compressed sensing
Dimitri Kanevsky, Avishy Carmi, Lior Horesh, Pini Gurfil, Bhuvana Ramabhadran, Tara N. Sainath |
FUSION | 5 |
| 2010 | The Use of isometric transformations and bayesian estimation in compressive sensing for fMRI classificationabstractCompressive sensing (CS) is a popular technique used to reconstruct a signal from few training examples, a problem which arises in many machine learning applications. In this paper, we introduce a technique to guarantee that our data obeys certain isometric properties. In addition, we introduce a bayesian approach to compressive sensing, which we call ABCS, allowing us to obtain complete statistics for estimated parameters. We apply these ideas to fMRI classification and find that by isometrically transforming our data, significant improvements in classification accuracy can be achieved using the LASSO and Dantzig selector methods, two standard techniques used in CS. In addition, applying the ABCS method offers improvements in classification accuracy over both LASSO and Dantzig. Finally, we find that applying both the ABCS method together with isometric transformations, we are able to achieve an error rate of 0.0%. Avishy Carmi, Tara N. Sainath, Pini Gurfil, Dimitri Kanevsky, David Nahamoo, Bhuvana Ramabhadran |
ICASSP | 6 |
| 2010 | Improved language modeling for conversational applications using sentence qualityabstractIn this paper, we propose a new approach to build language models for conversationals system using a a corpus of text as a opposed to a live or a Wizard-of-Oz collection. Each sentence in the corpus is assigned a “quality” that reflects the developer's intuition for how likely that sentence is to be spoken by a real user to the live system. Language Models (LM) are built for each sentence quality and these are subsequently interpolated to produce the final model. We also have built a classifier that assigns sentence qualities to the data, and whose subsequent language models achive similar improvements in word and turn error rate. Mark E. Epstein, Bhuvana Ramabhadran, Rajesh Balchandran |
ICASSP | 2 |
| 2010 | Balancing false alarms and hits in Spoken Term DetectionabstractThis paper presents methods to improve retrieval of Out-Of-Vocabulary (OOV) terms in a Spoken Term Detection (STD) system. We demonstrate that automated tagging of OOV regions helps to reduce false alarms while incorporating phonetic confusability increases the hits. Additional features that boost the probability of a hit in accordance with the number of neighboring hits for the same query and query-length normalization also improve the overall performance of the spoken-term detection system. We show that these methods can be combined effectively to provide a relative improvement of 21% in Average Term Weighted Value (ATWV) on a 100-hour corpus with 1290 OOV-only queries and 2% relative on the NIST 2006 STD task, where only 16 of the 1107 queries were OOV terms. Lastly, we present results to show that the proposed methods are general enough to work well in query-by-example based spoken-term detection, and in mismatched situations when the representation of the index being searched through and the queries are not generated by the same system. Carolina Parada, Abhinav Sethy, Bhuvana Ramabhadran |
ICASSP | 3 |
| 2010 | Bayesian compressive sensing for phonetic classificationabstractIn this paper, we introduce a novel Bayesian compressive sensing (CS) technique for phonetic classification. CS is often used to characterize a signal from a few support training examples, similar to k-nearest neighbor (kNN) and Support Vector Machines (SVMs). However, unlike SVMs and kNNs, CS allows the number of supports to be adapted to the specific signal being characterized. On the TIMIT phonetic classification task, we find that our CS method outperforms the SVM, kNN and Gaussian Mixture Model (GMM) methods. Our CS method achieves an accuracy of 80.01%, one of the best reported result in the literature to date. Tara N. Sainath, Avishy Carmi, Dimitri Kanevsky, Bhuvana Ramabhadran |
ICASSP | 4 |
| 2010 | Continuous space language modeling techniquesabstractThis paper focuses on comparison of two continuous space language modeling techniques, namely Tied-Mixture Language modeling (TMLM) and Neural Network Based Language Modeling (NNLM). Additionally, we report on using alternative feature representations for words and histories used in TMLM. Besides bigram co-occurrence based features we consider using NNLM based input features for training TMLMs. We also describe how we improve certain steps in building TMLMs. We demonstrate that TMLMs provide significant improvements of over 16% relative and 10% relative in Character Error Rate (CER) for Mandarin speech recognition, over the trigram and NNLM models, respectively in a speech to speech translation task. Ruhi Sarikaya, Ahmad Emami, Mohamed Afify, Bhuvana Ramabhadran |
ICASSP | 4 |
| 2010 | An autoencoder neural-network based low-dimensionality approach to excitation modeling for HMM-based text-to-speechabstractHMM-TTS synthesis is a popular approach toward flexible, low-footprint, data driven systems that produce highly intelligible speech. In spite of these strengths, speech generated by these systems exhibit some degradation in quality, attributable to an inadequacy in modeling the excitation signal that drives the parametric models of the vocal tract. This paper proposes a novel method for modeling the excitation as a low-dimensional set of coefficients, based on a non-linear map learned through an autoencoder. Through analysis-and-resynthesis experiments, and a formal listening test, we show that this model produces speech of higher perceptual quality compared to conventional pulse-excited speech signals at the p <; 0.01 significance level. Srikanth Vishnubhotla, Raul Fernandez, Bhuvana Ramabhadran |
ICASSP | 3 |
| 2010 | Techniques for topic detection based processing in spoken dialog systems
Rajesh Balchandran, Leonid Rachevsky, Bhuvana Ramabhadran, Miroslav Novak |
INTERSPEECH | 3 |
| 2010 | Discriminative training and unsupervised adaptation for labeling prosodic events with limited training data
Raul Fernandez, Bhuvana Ramabhadran |
INTERSPEECH | 2 |
| 2010 | Incorporating sparse representation phone identification features in automatic speech recognition using exponential familiesabstractSparse representation phone identification features (SPIF) is a recently developed technique to obtain an estimate of phone posterior probabilities conditioned on an acoustic feature vector. In this paper, we explore incorporating SPIF phone posterior probability estimates in large vocabulary continuous speech recognition (LVCSR) task by including them as additional features of exponential densities that model the HMM state emission likelihoods. We compare our proposed approach to a number of other well known methods of combining feature streams or multiple LVCSR systems. Our experiments show that using exponential models to combine features results in a word error rate reduction of 0.5% absolute (18.7% down to 18.2%); this is comparable to best error rate reduction obtained from system combination methods, but without having to build multiple systems or tune the system combination weights. Vaibhava Goel, Tara N. Sainath, Bhuvana Ramabhadran, Peder A. Olsen, David Nahamoo, Dimitri Kanevsky |
INTERSPEECH | 3 |
| 2010 | An analysis of sparseness and regularization in exemplar-based methods for speech classificationabstractThe use of exemplar-based techniques for both speech classification and recognition tasks has become increasingly popular in recent years. However, the notion of why sparseness is important for exemplar-based speech processing has been relatively unexplored. In addition, little analysis has been done in speech processing on the appropriateness of different types of sparsity regularization constraints. The goal of this paper is to answer the above two questions, both through mathematically analyzing different sparseness methods and also comparing these approaches for phonetic classification in TIMIT. Dimitri Kanevsky, Tara N. Sainath, Bhuvana Ramabhadran, David Nahamoo |
INTERSPEECH | 3 |
| 2010 | Sparse representations for text categorizationabstractSparse representations (SRs) are often used to characterize a test signal using few support training examples, and allow the number of supports to be adapted to the specific signal being categorized. Given the good performance of SRs compared to other classifiers for both image classification and phonetic clas-sification, in this paper, we extended the use of SRs for text classification, a method which has thus far not been explored for this domain. Specifically, we demonstrate how sparse repre-sentations can be used for text classification and how their per-formance varies with the vocabulary size of the documents. In addition, we also show that this method offers promising results over the Naive Bayes (NB) classifier, a standard baseline classi-fier used for text categorization, thus introducing an alternative class of methods for text categorization. 1. Tara N. Sainath, Sameer Maskey, Dimitri Kanevsky, Bhuvana Ramabhadran, David Nahamoo, Julia Hirschberg |
INTERSPEECH | 4 |
| 2010 | Sparse representation features for speech recognitionabstractIn this paper, we explore the use of exemplar-based sparse representations (SRs) to map test features into the linear span of training examples. We show that the frame classification ac-curacy with these new features is 1.3 % higher than a Gaussian Mixture Model (GMM), showing that not only do SRs move test features closer to training, but also move the features closer to the correct class. Given these new SR features, we train up a Hidden Markov Model (HMM) on these features and per-form recognition. On the TIMIT corpus, we show that apply-ing the SR features on top of our best discriminatively trained system allows for a 0.7 % absolute reduction in phonetic er-ror rate (PER), from 19.9 % to 19.2%. In fact, after applying model adaptation we reduce the PER to 19.0%, the best re-sults on TIMIT to date. Furthermore, on a large vocabulary 50 hour broadcast news task, we achieve a reduction in word error rate (WER) of 0.3 % absolute, demonstrating the benefit of this method for large vocabulary speech recognition. 1. Tara N. Sainath, Bhuvana Ramabhadran, David Nahamoo, Dimitri Kanevsky, Abhinav Sethy |
INTERSPEECH | 2 |
| 2010 | Impact of word classing on shrinkage-based language models
Ruhi Sarikaya, Stanley F. Chen, Abhinav Sethy, Bhuvana Ramabhadran |
INTERSPEECH | 4 |
| 2010 | Data selection for language modeling using sparse representations
Abhinav Sethy, Tara N. Sainath, Bhuvana Ramabhadran, Dimitri Kanevsky |
INTERSPEECH | 3 |
| 2010 | Unsupervised Model Adaptation using Information-Theoretic Criterion
Ariya Rastrow, Frederick Jelinek, Abhinav Sethy, Bhuvana Ramabhadran |
HLT-NAACL | 4 |
| 2009 | Scaling shrinkage-based language modelsabstractIn we show that a novel class-based language model, Model M, and the method of regularized minimum discrimination information (rMDI) models outperform comparable methods on moderate amounts of Wall Street Journal data. Both of these methods are motivated by the observation that shrinking the sum of parameter magnitudes in an exponential language model tends to improve performance. In this paper, we investigate whether these shrinkage-based techniques also perform well on larger training sets and on other domains. First, we explain why good performance on large data sets is uncertain, by showing that gains relative to a baseline n-gram model tend to decrease as training set size increases. Next, we evaluate several methods for data/model combination with Model M and rMDI models on limited-scale domains, to uncover which techniques should work best on large domains. Finally, we apply these methods on a variety of medium-to-large-scale domains covering several languages, and show that Model M consistently provides significant gains over existing language models for state-of-the-art systems in both speech recognition and machine translation. Stanley F. Chen, Lidia Mangu, Bhuvana Ramabhadran, Ruhi Sarikaya, Abhinav Sethy |
ASRU | 3 |
| 2009 | Query-by-example Spoken Term Detection For OOV termsabstractThe goal of spoken term detection (STD) technology is to allow open vocabulary search over large collections of speech content. In this paper, we address cases where search term(s) of interest (queries) are acoustic examples. This is provided either by identifying a region of interest in a speech stream or by speaking the query term. Queries often relate to named-entities and foreign words, which typically have poor coverage in the vocabulary of large vocabulary continuous speech recognition (LVCSR) systems. Throughout this paper, we focus on query-by-example search for such out-of-vocabulary (OOV) query terms. We build upon a finite state transducer (FST) based search and indexing system to address the query by example search for OOV terms by representing both the query and the index as phonetic lattices from the output of an LVCSR system. We provide results comparing different representations and generation mechanisms for both queries and indexes built with word and combined word and subword units. We also present a two-pass method which uses query-by-example search using the best hit identified in an initial pass to augment the STD search results. The results demonstrate that query-by-example search can yield a significantly better performance, measured using actual term-weighted value (ATWV), of 0.479 when compared to a baseline ATWV of 0.325 that uses reference pronunciations for OOVs. Further improvements can be obtained with the proposed two pass approach and filtering using the expected unigram counts from the LVCSR system's lexicon. Carolina Parada, Abhinav Sethy, Bhuvana Ramabhadran |
ASRU | 3 |
| 2009 | Constrained discriminative training of N-gram language modelsabstractIn this paper, we present a novel version of discriminative training for N-gram language models. Language models impose language specific constraints on the acoustic hypothesis and are crucial in discriminating between competing acoustic hypotheses. As reported in the literature, discriminative training of acoustic models has yielded significant improvements in the performance of a speech recognition system, however, discriminative training for N-gram language models (LMs) has not yielded the same impact. In this paper, we present three techniques to improve the discriminative training of LMs, namely updating the back-off probability of unseen events, normalization of the N-gram updates to ensure a probability distribution and a relative-entropy based global constraint on the N-gram probability updates. We also present a framework for discriminative adaptation of LMs to a new domain and compare it to existing linear interpolation methods. Results are reported on the Broadcast News and the MIT lecture corpora. A modest improvement of 0.2% absolute (on Broadcast News) and 0.3% absolute (on MIT lectures) was observed with discriminatively trained LMs over state-of-the-art systems. Ariya Rastrow, Abhinav Sethy, Bhuvana Ramabhadran |
ASRU | 3 |
| 2009 | An exploration of large vocabulary tools for small vocabulary phonetic recognitionabstractWhile research in large vocabulary continuous speech recognition (LVCSR) has sparked the development of many state of the art research ideas, research in this domain suffers from two main drawbacks. First, because of the large number of parameters and poorly labeled transcriptions, gaining insight into further improvements based on error analysis is very difficult. Second, LVCSR systems often take a significantly longer time to train and test new research ideas compared to small vocabulary tasks. A small vocabulary task like TIMIT provides a phonetically rich and hand-labeled corpus and offers a good test bed to study algorithmic improvements. However, oftentimes research ideas explored for small vocabulary tasks do not always provide gains on LVCSR systems. In this paper, we address these issues by taking the standard "recipe" used in typical LVCSR systems and applying it to the TIMIT phonetic recognition corpus, which provides a standard benchmark to compare methods. We find that at the speaker-independent (SI) level, our results offer comparable performance to other SI HMM systems. By taking advantage of speaker adaptation and discriminative training techniques commonly used in LVCSR systems, we achieve an error rate of 20%, the best results reported on the TIMIT task to date, moving us closer to the human reported phonetic recognition error rate of 15%. We propose the use of this system as the baseline for future research and believe that it will serve as a good framework to explore ideas that will carry over to LVCSR systems. Tara N. Sainath, Bhuvana Ramabhadran, Michael Picheny |
ASRU | 2 |
| 2009 | Map approach to learning sparse Gaussian Markov networksabstractRecently proposed l1-regularized maximum-likelihood optimization methods for learning sparse Markov networks result into convex problems that can be solved optimally and efficiently. However, the accuracy of such methods can be very sensitive to the choice of regularization parameter, and optimal selection of this parameter remains an open problem. Herein, we propose a maximum a posteriori probability (MAP) approach that investigates different priors on the regularization parameter and yields promising empirical results on both synthetic data and real-life application such as brain imaging data (fMRI). Narges Bani Asadi, Irina Rish, Katya Scheinberg, Dimitri Kanevsky, Bhuvana Ramabhadran |
ICASSP | 5 |
| 2009 | Effect of pronounciations on OOV queries in spoken term detectionabstractThe spoken term detection (STD) task aims to return relevant segments from a spoken archive that contain the query terms whether or not they are in the system vocabulary. This paper focuses on pronunciation modeling for out-of-vocabulary (OOV) terms which frequently occur in STD queries. The STD system described in this paper indexes word-level and sub-word level lattices or confusion networks produced by an LVCSR system using weighted finite state transducers (WFST).We investigate the inclusion of n-best pronunciation variants for OOV terms (obtained from letter-to-sound rules) into the search and present the results obtained by indexing confusion networks as well as lattices. The following observations are worth mentioning: phone indexes generated from sub-words represent OOVs well and too many variants for the OOV terms degrade performance if pronunciations are not weighted. Dogan Can, Erica Cooper, Abhinav Sethy, Christopher M. White, Bhuvana Ramabhadran, Murat Saraclar |
ICASSP | 5 |
| 2009 | A generalized family of parameter estimation techniquesabstractThe extended Baum-Welch (EBW) transformations is one of a variety of techniques to estimate parameters of Gaussian mixture models. In this paper, we provide a theoretical framework for general parameter estimation and show the relationship between these different techniques. We introduce a general family of model parameter updates that generalizes a Baum-Welch (BW) recursive process to an arbitrary objective function of Gaussian Mixture Models, and show how other common parameter estimation techniques belong to this family of model update rules. Furthermore, we formulate the construction of an even more general family of update rules that has any specified value as a gradient steepness which belongs to the family of EBW gradient steepness, measuring how much an initial model is moved to an estimated updated model. Dimitri Kanevsky, Tara N. Sainath, Bhuvana Ramabhadran |
ICASSP | 3 |
| 2009 | A new method for OOV detection using hybrid word/fragment systemabstractIn this paper, we propose a new method for detecting regions with out-of-vocabulary (OOV) words in the output of a large vocabulary continuous speech recognition (LVCSR) system. The proposed method uses a hybrid system combining words and data-driven variable length sub word units. With the use of a single feature, the posterior probability of sub word units, this method outperforms existing methods published in the literature. We also presents a recipe to discriminatively train a hybrid language model to improve OOV detection rate. Results are presented on the RT04 broadcast news task. Ariya Rastrow, Abhinav Sethy, Bhuvana Ramabhadran |
ICASSP | 3 |
| 2009 | Unsupervised pronunciation validationabstractThis paper addresses selecting between candidate pronunciations for out-of-vocabulary words in speech processing tasks. We introduce a simple, unsupervised method that outperforms the conventional supervised method of forced alignment with a reference. The success of this method is independently demonstrated using three metrics from large-scale speech tasks: word error rates for large vocabulary continuous speech recognition, decision error tradeoff curves for spoken term detection, and phone error rates compared to a handcrafted pronunciation lexicon. The experiments were conducted using state-of-the-art recognition, indexing, and retrieval systems. The results were compared across many terms, hundreds of hours of speech, and well known data sets. Christopher M. White, Abhinav Sethy, Bhuvana Ramabhadran, Patrick J. Wolfe, Erica Cooper, Murat Saraclar, James K. Baker |
ICASSP | 3 |
| 2009 | Towards using hybrid word and fragment units for vocabulary independent LVCSR systemsabstractThis paper presents the advantages of augmenting a word-based system with sub-word units as a step towards building open vocabulary speech recognition systems. We show that a hybrid system which combines words and data-driven, variable length sub word units has a better phone accuracy than word only systems. In addition the hybrid system is better in detecting Out-Of-Vocabulary (OOV) terms and representing them phonetically. Results are presented on the RT-04 broadcast news and MIT Lecture data sets. An FSM-based approach to recover OOV words from the hybrid lattices is also presented. At an OOV rate of 2.5 % on RT-04 we observed a 8 % relative improvement in phone error rate (PER), 7.3 % relative improvement in oracle PER and 7 % relative improvement in WER after recovering the OOV terms. A significant reduction of 33 % relative in PER is seen in the OOV regions. Index Terms: hybrid LVCSR system, out-of-vocabulary, subword unit selection, phone recognition, OOV detection Ariya Rastrow, Abhinav Sethy, Bhuvana Ramabhadran, Frederick Jelinek |
INTERSPEECH | 3 |
| 2009 | Iterative sentence-pair extraction from quasi-parallel corpora for machine translationabstractThis paper addresses parallel data extraction from the quasi–parallel corpora generated in a crowd-sourcing project where ordinary people watch tv shows and movies and transcribe/translate what they hear, creating document pools in different languages. Since they do not have guidelines for naming and performing translations, it is often not clear which documents are the translations of the same show/movie and which sentences are the translations of the each other in a given document pair. We introduce a method for automatically pairing documents in two languages and extracting parallel sentences from the paired documents. The method consists of three steps: i) document pairing, ii) sentence pair alignment of the paired documents, and iii) context extrapolation to boost the sentence pair coverage. Human evaluation of the extracted data shows that 95 % of the extracted sentences carry useful information for translation. Experimental results also show that using the extracted data provides significant gains over the baseline statistical machine translation system built with manually annotated data. Index Terms: data extraction, comparable data, machine translation Ruhi Sarikaya, Sameer Maskey, Ea-Ee Jan, Bhuvana Ramabhadran, Salim Roukos |
INTERSPEECH | 6 |
| 2009 | Web derived pronunciations for spoken term detectionabstractIndexing and retrieval of speech content in various forms such as broadcast news, customer care data and on-line media has gained a lot of interest for a wide range of applications, from customer analytics to on-line media search. For most retrieval applications, the speech content is typically first converted to a lexical or phonetic representation using automatic speech recognition (ASR). The first step in searching through indexes built on these representations is the generation of pronunciations for named entities and foreign language query terms. This paper summarizes the results of the work conducted during the 2008 JHU Summer Workshop by the Multilingual Spoken Term Detection team, on mining the web for pronunciations and analyzing their impact on spoken term detection. We will first present methods to use the vast amount of pronunciation information available on the Web, in the form of IPA and ad-hoc transcriptions. We describe techniques for extracting candidate pronunciations from Web pages and associating them with orthographic words, filtering out poorly extracted pronunciations, normalizing IPA pronunciations to better conform to a common transcription standard, and generating phonemic representations from ad-hoc transcriptions. We then present an analysis of the effectiveness of using these pronunciations to represent Out-Of-Vocabulary (OOV) query terms on the performance of a spoken term detection (STD) system. We will provide comparisons of Web pronunciations against automated techniques for pronunciation generation as well as pronunciations generated by human experts. Our results cover a range of speech indexes based on lattices, confusion networks and one-best transcriptions at both word and word fragments levels. Dogan Can, Erica Cooper, Arnab Ghoshal, Martin Jansche, Sanjeev Khudanpur, Bhuvana Ramabhadran, Michael Riley 0001, Murat Saraclar, Abhinav Sethy, Morgan Ulinski, Christopher M. White |
SIGIR | 6 |
| 2009 | An Iterative Relative Entropy Minimization-Based Data Selection Approach for n-Gram Model AdaptationabstractPerformance of statistical n-gram language models depends heavily on the amount of training text material and the degree to which the training text matches the domain of interest. The language modeling community is showing a growing interest in using large collections of text (obtainable, for example, from a diverse set of resources on the Internet) to supplement sparse in-domain resources. However, in most cases the style and content of the text harvested from the web differs significantly from the specific nature of these domains. In this paper, we present a relative entropy based method to select subsets of sentences whose n-gram distribution matches the domain of interest. We present results on language model adaptation using two speech recognition tasks: a medium vocabulary medical domain doctor-patient dialog system and a large vocabulary transcription system for European parliamentary plenary speeches (EPPS). We show that the proposed subset selection scheme leads to performance improvements over state of the art speech recognition systems in terms of both speech recognition word error rate (WER) and language model perplexity (PPL). Abhinav Sethy, Panayiotis G. Georgiou, Bhuvana Ramabhadran, Shri Narayanan |
IEEE Trans. Speech Audio Process. | 3 |
| 2008 | Boosted MMI for model and feature-space discriminative trainingabstractWe present a modified form of the maximum mutual information (MMI) objective function which gives improved results for discriminative training. The modification consists of boosting the likelihoods of paths in the denominator lattice that have a higher phone error relative to the correct transcript, by using the same phone accuracy function that is used in Minimum Phone Error (MPE) training. We combine this with another improvement to our implementation of the Extended Baum-Welch update equations for MMI, namely the canceling of any shared part of the numerator and denominator statistics on each frame (a procedure that is already done in MPE). This change affects the Gaussian-specific learning rate. We also investigate another modification whereby we replace I-smoothing to the ML estimate with I-smoothing to the previous iteration's value. Boosted MMI gives better results than MPE in both model and feature-space discriminative training, although not consistently. Daniel Povey, Dimitri Kanevsky, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, Karthik Visweswariah |
ICASSP | 4 |
| 2008 | Gradient steepness metrics using extended Baum-Welch transformations for universal pattern recognition tasksabstractIn many pattern recognition tasks, given some input data and a family of models, the "best" model is defined as the one which maximizes the likelihood of the data given the model. Extended Baum- Welch (EBW) transformations are most commonly used as a discriminative technique for estimating parameters of Gaussian mixtures. In this paper, we use the EBW transformations to derive a novel gradient steepness measurement to find which model best explains the data. We use this gradient measurement to derive a variety of EBW metrics to explain model fit to the data. We apply these EBW metrics to audio segmentation via Hidden Markov Models (HMMs) and show that our gradient steepness measurement is robust across different EBW metrics and model complexities. Tara N. Sainath, Dimitri Kanevsky, Bhuvana Ramabhadran |
ICASSP | 3 |
| 2008 | A study of unsupervised clustering techniques for language modeling
Sangyun Hahn, Abhinav Sethy, Hong-Kwang Jeff Kuo, Bhuvana Ramabhadran |
INTERSPEECH | 4 |
| 2008 | Generalization of extended baum-welch parameter estimation for discriminative training and decodingabstractWe demonstrate the generalizability of the Extended Baum-Welch (EBW) algorithm not only for HMM parameter estimation but for decoding as well. We show that there can exist a general function associated with the objective function under EBW that reduces to the well-known auxiliary function used in the Baum-Welch algorithm for maximum likelihood estimates. We generalize representation for the updates of model parameters by making use of a differentiable function (such as arithmetic or geometric mean) on the updated and current model parameters and describe their effect on the learning rate during HMM parameter estimation. Improvements on speech recognition tasks are also presented here. 1. Dimitri Kanevsky, Tara N. Sainath, Bhuvana Ramabhadran, David Nahamoo |
INTERSPEECH | 3 |
| 2008 | Phonetic query expansion for spoken document retrievalabstractWe are interested in retrieving information from speech data using phonetic search. We show improvement by expanding the query phonetically using a joint maximum entropy N-gram model. The value of this approach is demonstrated on Broadcast News data from NIST 2006 Spoken Term Detection evaluation. Jonathan Mamou, Bhuvana Ramabhadran |
INTERSPEECH | 2 |
| 2008 | Bag-of-word normalized n-gram models
Abhinav Sethy, Bhuvana Ramabhadran |
INTERSPEECH | 2 |
| 2007 | Fast audio search using vector space modellingabstractMany techniques for retrieving arbitrary content from audio have been developed to leverage the important challenge of providing fast access to very large volumes of multimedia data. We present a two-stage method for fast audio search, where a vector-space modelling approach is first used to retrieve a short list of candidate audio segments for a query. The list of candidate segments is then searched using a word-based index for known words and a phone-based index for out-of-vocabulary words. We explore various system configurations and examine trade-offs between speed and accuracy. We evaluate our audio search system according to the NIST 2006 Spoken Term Detection evaluation initiative. We find that we can obtain a 30-times speedup for the search phase of our system with a 10% relative loss in accuracy. Brett Matthews, Upendra V. Chaudhari, Bhuvana Ramabhadran |
ASRU | 3 |
| 2007 | The IBM 2007 speech transcription system for European parliamentary speechesabstractTC-STAR is an European Union funded speech to speech translation project to transcribe, translate and synthesize European Parliamentary Plenary Speeches (EPPS). This paper describes IBM's English speech recognition system submitted to the TC-STAR 2007 Evaluation. Language model adaptation based on clustering and data selection using relative entropy minimization provided significant gains in the 2007 evaluation. The additional advances over the 2006 system that we present in this paper include unsupervised training of acoustic and language models; a system architecture that is based on cross-adaptation across complementary systems and system combination through generation of an ensemble of systems using randomized decision tree state-tying. These advances reduced the error rate by 30% relative over the best-performing system in the TC-STAR 2006 evaluation on the 2006 English development and evaluation test sets, and produced one of the best performing systems on the 2007 evaluation in English with a word error rate of 7.1%. Bhuvana Ramabhadran, Olivier Siohan, Abhinav Sethy |
ASRU | 1 |
| 2007 | Broad phonetic class recognition in a Hidden Markov model framework using extended Baum-Welch transformationsabstractIn many pattern recognition tasks, given some input data and a model, a probabilistic likelihood score is often computed to measure how well the model describes the data. Extended Baum-Welch (EBW) transformations are most commonly used as a discriminative technique for estimating parameters of Gaussian mixtures, though recently they have been used to derive a gradient steepness measurement to evaluate the quality of the model to match the distribution of the data. In this paper, we explore applying the EBW gradient steepness metric in the context of Hidden Markov Models (HMMs) for recognition of broad phonetic classes and present a detailed analysis and results on the use of this gradient metric on the TIMIT corpus. We find that our gradient metric is able to outperform the baseline likelihood method, and offers improvements in noisy conditions. Tara N. Sainath, Dimitri Kanevsky, Bhuvana Ramabhadran |
ASRU | 3 |
| 2007 | Data Driven Approach for Language Model Adaptation using Stepwise Relative Entropy MinimizationabstractThe ability to build domain and task specific language models from large generic text corpora is of considerable interest to the language modeling community. One of the key challenges is to identify the relevant text material in the collection. The text selection problem can be cast in a semi-supervised learning framework. Motivated by recent advancements in semi-supervised learning which emphasize the need of balanced label assignments, we present a stepwise relative entropy minimization scheme which focuses on selection of a set of sentences instead of selecting sentences solely on their individual merit. Our results on the IBM European Parliament Plenary Speech (EPPS) transcription system, show significant performance improvement (0.5% on an 8.9% baseline), with just a seventh of the out-of-domain data. The IBM EPPS LVCSR system which has a 60K vocabulary is a particularly hard baseline for out-of-domain adaptation because of low WER with in-domain training data. Abhinav Sethy, Shri Narayanan, Bhuvana Ramabhadran |
ICASSP (4) | 3 |
| 2007 | Vocabulary independent spoken term detectionabstractWe are interested in retrieving information from speech data like broadcast news, telephone conversations and roundtable meetings. Today, most systems use large vocabulary continuous speech recognition tools to produce word transcripts; the transcripts are indexed and query terms are retrieved from the index. However, query terms that are not part of the recognizer's vocabulary cannot be retrieved, and the recall of the search is affected. In addition to the output word transcript, advanced systems provide also phonetic transcripts, against which query terms can be matched phonetically. Such phonetic transcripts suffer from lower accuracy and cannot be an alternative to word transcripts.We present a vocabulary independent system that can handle arbitrary queries, exploiting the information provided by having both word transcripts and phonetic transcripts. A speech recognizer generates word confusion networks and phonetic lattices. The transcripts are indexed for query processing and ranking purpose.The value of the proposed method is demonstrated by the relative high performance ofour system, which received the highest overall ranking for US English speech data in the recent NIST Spoken Term Detection evaluation. Jonathan Mamou, Bhuvana Ramabhadran, Olivier Siohan |
SIGIR | 2 |
| 2006 | Automated Quality Monitoring in the Call Center with ASR and Maximum EntropyabstractThis paper describes an automated system for assigning quality scores to recorded call center conversations. The system combines speech recognition, pattern matching, and maximum entropy classification to rank calls according to their measured quality. Calls at both end of the spectrum are flagged as "interesting" and made available for further human monitoring. In this process, pattern matching on the ASR transcript is used to answer a set of standard quality control questions such as "did the agent use courteous words and phrases," and to generate a question-based score. This is interpolated with the probability of a call being "bad," as determined by maximum entropy operating on a set of ASR-derived features such as "maximum silence length" and the occurrence of selected n-gram word sequences. The system is trained on a set of calls with associated manual evaluation forms. We present precision and recall results from IBM's North American Help Desk indicating that for a given amount of listening effort, this system triples the number of bad calls that are identified, over the current policy of randomly sampling calls Geoffrey Zweig, Olivier Siohan, George Saon, Bhuvana Ramabhadran, Daniel Povey, Lidia Mangu, Brian Kingsbury |
ICASSP (1) | 4 |
| 2006 | The IBM 2006 speech transcription system for european parliamentary speechesabstractTC-STAR is an European Union funded speech to speech translation project to transcribe, translate and synthesize European Parliamentary Plenary Speeches (EPPS). This paper describes IBM’s English and Spanish speech recognition systems submitted to the TC-STAR 2006 Evaluation. The technical advances in this submission include two different algorithms for automatic segmentation and speaker clustering of the input audio; a system architecture that is based on cross-adaptation across these two segmentation schemes and system combination through generation of an ensemble of systems using randomized decision tree state-tying; automatic punctuation of the speech recognition output; and the incorporation of an additional 35 hours of in-domain EPPS acoustic training data. These advances reduced the error rate by 30% relative over the best-performing system in the TC-STAR 2005 Evaluation on the 2006 English development test set, and produced one of the best performing systems on the 2006 evaluation in English with a word error rate of 8.3%. Index Terms: speech recognition, automatic segmentation, crossadaptation, randomized decision trees, TC-STAR. Bhuvana Ramabhadran, Olivier Siohan, Lidia Mangu, Geoffrey Zweig, Martin Westphal, Henrik Schulz 0002, Alvaro Soneiro |
INTERSPEECH | 1 |
| 2006 | Automated Quality Monitoring for Call Centers using Speech and NLP Technologies
Geoffrey Zweig, Olivier Siohan, George Saon, Bhuvana Ramabhadran, Daniel Povey, Lidia Mangu, Brian Kingsbury |
HLT-NAACL | 4 |
| 2006 | On the Effect Ofword Error Rate on Automated Quality MonitoringabstractThis paper studies the effect of word-error-rate (WER) on an automated quality monitoring application for call centers. The system consists of a speech recognition module and a call ranking module. The call ranking module combines direct question answering with a maximum-entropy classifier to automatically monitor the calls that enter a call center, and label them as "good" or "bad". We find that, in the monitoring regime where only a small fraction of the calls are monitored, we achieve 80% precision and 50% recall in classifying whether a call belongs to the bottom 20%. Additionally, the correlation between human and computer-generated scores turns out to be highly sensitive to word error rate. George Saon, Bhuvana Ramabhadran, Geoffrey Zweig |
SLT | 2 |
| 2005 | Contructing Ensembles of ASR Systems Using Randomized Decision TreesabstractBuilding multiple automatic speech recognition (ASR) systems and combining their outputs using voting techniques such as ROVER is an effective technique for lowering the overall word error rate. A successful system combination approach requires the construction of multiple systems with complementary errors, or the combination will not outperform any of the individual systems. In general, this is achieved empirically, for example by building systems on different input features. In this paper, we present a systematic approach for building multiple ASR systems in which the decision tree state-tying procedure that is used to specify context-dependent acoustic models is randomized. Experiments carried out on two large vocabulary recognition tasks, MALACH and DARPA EARS, illustrate the effectiveness of the approach. Olivier Siohan, Bhuvana Ramabhadran, Brian Kingsbury |
ICASSP (1) | 2 |
| 2005 | Exploiting large quantities of spontaneous speech for unsupervised training of acoustic models
Bhuvana Ramabhadran |
INTERSPEECH | 1 |
| 2004 | Use of metadata to improve recognition of spontaneous speech and named entitiesabstractWith improved recognition accuracies for LVCSR tasks, it has become possible to search large collections of spontaneous speech for a variety of information. The MALACH corpus of Holocaust testimonials is one such collection, in which we are interested in automatically transcribing and retrieving portions that are relevant to named entities such as people, places, and organizations. Since the testimonials were gathered from thousands of people in countries throughout Europe, an extremely large number of potential named entities are possible, and this causes a well-known dilemma: increasing the size of the vocabulary allows for more of these words to be recognized, but also increases confusability, and can harm recognition performance. However, the MALACH corpus, like many other collections, includes side information or metadata that can be exploited to provide prior information on exactly which named entities are likely to appear. This paper proposes a method that capitalizes on this prior information to reduce named-entity recognition errors by over 50 % relative, and simultaneously decrease the overall word error rate by 7 % relative. The metadata we use derives from a pre-interview questionaire that includes the names of friends, relatives, places visited, membership of organizations, synonyms of place names, and similar information. By augmenting the lexicon and language model with this information on a speaker-by-speaker basis, we are able to exploit the textual information that is already available in the corpus to facilitate much improved speech recognition. 1. Bhuvana Ramabhadran, Olivier Siohan, Geoffrey Zweig |
INTERSPEECH | 1 |
| 2004 | Measuring convergence in language model estimation using relative entropyabstractLanguage models are generally estimated using smoothed counting techniques. These counting schemes can be viewed as non linear functions operating on a Bernoulli process which converge asymptotically to the true density. The rate at which these counting schemes converge to the true density is constrained by the training data set available and the nature of the language model (LM) being estimated. In this paper we look at language model estimates as random variables and present an efficient relative entropy (R.E) based approach to study their convergence with increasing training data size. We present experimental results for language modeling in a generic LVCSR system and a medical domain dialogue task. We also present an efficient recursive R.E computation method which can be used as a LM distance measure for a number of tasks including LM clustering. Abhinav Sethy, Shri Narayanan, Bhuvana Ramabhadran |
INTERSPEECH | 3 |
| 2004 | Speech recognition error analysis on the English MALACH corpusabstractThis paper presents an analysis of the word recognition error rate on an English subset of the MALACH corpus. The MALACH project is an NSF-funded research program related to the development of multilingual access to large audio archives. The archive of interest is a large collection of testimonies from 52,000 survivors, liberators, rescuers and witnesses of the Nazi Holocaust, assembled by the Shoah Visual History Foundation. This data has some unique characteristics that make it quite unusual in the speech recognition community such as elderly speech, noisy conditions, heavily accented speech. Hence, it is a challenging task for automatic speech recognition (ASR). This paper attempts to identify the factors affecting the ASR performance on that task. It was found that the signal-to-noise ratio and syllable Olivier Siohan, Bhuvana Ramabhadran, Geoffrey Zweig |
INTERSPEECH | 2 |
| 2004 | Building an information retrieval test collection for spontaneous conversational speechabstractTest collections model use cases in ways that facilitate evaluation of information retrieval systems. This paper describes the use of search-guided relevance assessment to create a test collection for retrieval of spontaneous conversational speech. Approximately 10,000 thematically coherent segments were manually identified in 625 hours of oral history interviews with 246 individuals. Automatic speech recognition results, manually prepared summaries, controlled vocabulary indexing, and name authority control are available for every segment. Those features were leveraged by a team of four relevance assessors to identify topically relevant segments for 28 topics developed from actual user requests. Search-guided assessment yielded sufficient inter-annotator agreement to support formative evaluation during system development. Baseline results for ranked retrieval are presented to illustrate use of the collection. Douglas W. Oard, Dagobert Soergel, David S. Doermann, G. Craig Murray, Jianqiang Wang 0002, Bhuvana Ramabhadran, Martin Franz, Samuel Gustman, James Mayfield, Liliya Kharevych, Stephanie M. Strassel |
SIGIR | 7 |
| 2004 | Automatic recognition of spontaneous speech for access to multilingual oral history archivesabstractMuch is known about the design of automated systems to search broadcast news, but it has only recently become possible to apply similar techniques to large collections of spontaneous speech. This paper presents initial results from experiments with speech recognition, topic segmentation, topic categorization, and named entity detection using a large collection of recorded oral histories. The work leverages a massive manual annotation effort on 10 000 h of spontaneous speech to evaluate the degree to which automatic speech recognition (ASR)-based segmentation and categorization techniques can be adapted to approximate decisions made by human annotators. ASR word error rates near 40% were achieved for both English and Czech for heavily accented, emotional and elderly spontaneous speech based on 65-84 h of transcribed speech. Topical segmentation based on shifts in the recognized English vocabulary resulted in 80% agreement with manually annotated boundary positions at a 0.35 false alarm rate. Categorization was considerably more challenging, with a nearest-neighbor technique yielding F=0.3. This is less than half the value obtained by the same technique on a standard newswire categorization benchmark, but replication on human-transcribed interviews showed that ASR errors explain little of that difference. The paper concludes with a description of how these capabilities could be used together to search large collections of recorded oral histories. William J. Byrne, David S. Doermann, Martin Franz, Samuel Gustman, Jan Hajic 0001, Douglas W. Oard, Michael Picheny, Josef Psutka, Bhuvana Ramabhadran, Dagobert Soergel, Todd Ward, Wei-Jing Zhu |
IEEE Trans. Speech Audio Process. | 9 |
| 2003 | Towards automatic transcription of large spoken archives - English ASR for the MALACH projectabstractDigital archives have emerged as the pre-eminent method for capturing the human experience. Before such archives can be used efficiently, their contents must be described. The NSF-funded MALACH project aims to provide improved access to large spoken archives by advancing the state-of-the-art in automated speech recognition (ASR), Information Retrieval (IR) and related technologies [1,2] for multiple languages. This paper describes the ASR research for the English speech in the MALACH corpus. The MALACH corpus consists of unconstrained, natural speech filled with disfluencies, heavy accents, age-related coarticulation, uncued speaker and language switching, and emotional speech collected in the form of interviews from over 52000 speakers in 32 languages. In this paper, we describe this new testbed for developing speech recognition algorithms and report on the performance of well-known techniques for building better acoustic models for the speaking styles seen in this corpus. The best English ASR system to date has a word error rate of 43.8% on this corpus. Bhuvana Ramabhadran, Jing Huang 0019, Michael Picheny |
ICASSP (1) | 1 |
| 2003 | Automated transcription and topic segmentation of large spoken archivesabstractDigital archives have emerged as the pre-eminent method for capturing the human experience. Before such archives can be used efficiently, their contents must be described. The scale of such archives along with the associated content mark up cost make it impractical to provide access via purely manual means, but automatic technologies for search in spoken materials still have relatively limited capabilities. The NSF-funded MALACH project will use the world’s largest digital archive of video oral histories, collected by the Survivors of the Shoah Visual History Foundation (VHF) to make a quantum leap in the ability to access such archives by advancing the state-of-the-art in Automated Speech Recognition (ASR), Natural Language Processing (NLP) and related technologies [1, 2]. This corpus consists of over 115,000 hours of unconstrained, natural speech from 52,000 speakers in 32 different languages, filled with disfluencies, heavy accents, age-related coarticulations, and un-cued speaker and language switching. This paper discusses some of the ASR and NLP tools and technologies that we have been building for the English speech in the MALACH corpus. We also discuss this new test bed while emphasizing the unique characteristics of this corpus. 1. Martin Franz, Bhuvana Ramabhadran, Todd Ward, Michael Picheny |
INTERSPEECH | 2 |
| 2003 | Impact of audio segmentation and segment clustering on automated transcription accuracy of large spoken archivesabstractThis paper addresses the influence of audio segmentation and segment clustering on automatic transcription accuracy for large spoken archives. The work forms part of the ongoing MALACH project, which is developing advanced techniques for supporting access to the world’s largest digital archive of video oral histories collected in many languages from over 52000 survivors and witnesses of the Holocaust. We present several audio-only and audio-visual segmentation schemes, including two novel schemes: the first is iterative and audio-only, the second uses audio-visual synchrony. Unlike most previous work, we evaluate these schemes in terms of their impact upon recognition accuracy. Results on English interviews show the automatic segmentation schemes give performance comparable to (exhorbitantly expensive and impractically lengthy) manual segmentation when using a single pass decoding strategy based on speaker-independent models. However, when using a multiple pass decoding strategy with adaptation, results are sensitive to both initial audio segmentation and the scheme for clustering segments prior to adaptation: the combination of our best automatic segmentation and clustering scheme has an error rate 8% worse (relative) to manual audio segmentation and clustering due to the occurrence of “speaker-impure ” segments. 1. Bhuvana Ramabhadran, Jing Huang 0019, Upendra V. Chaudhari, Giridharan Iyengar, Harriet J. Nock |
INTERSPEECH | 1 |
| 2001 | Speech recognition for DARPA CommunicatorabstractWe report the results of investigations in acoustic modeling, language modeling and decoding techniques, for the DARPA Communicator, a speaker-independent, telephone-based dialog system. By a combination of methods, including enlarging the acoustic model, augmenting the recognizer vocabulary, conditioning the language model upon the dialog state, and applying a post-processing decoding method, we lowered the overall word error rate from 21.9% to 15.0%, a gain of 6.9% absolute and 31.5% relative. Andrew Aaron, Scott Saobing Chen, Paul S. Cohen, Satya Dharanipragada, Ellen Eide, Martin Franz, Jean-Michel LeRoux, X. Luo, Benoît Maison, Lidia Mangu, T. Mathes, Miroslav Novak, Peder A. Olsen, Michael Picheny, Harry Printz, Bhuvana Ramabhadran, Andrej Sakrajda, George Saon, Borivoj Tydlitát, Karthik Visweswariah, D. Yuk |
ICASSP | 16 |
| 2001 | Innovative approaches for large vocabulary name recognitionabstractAutomatic name dialing is a practical and interesting application of speech recognition on telephony systems. The IBM name recognition system is a large vocabulary, speaker independent system currently in use for reaching IBM employees in the United States. We present some innovative algorithms that improve name recognition accuracy. Unlike transcription tasks, such as the Switchboard task, recognition of names poses a variety of different problems. Several of these problems arise from the fact that foreign names are hard to pronounce for speakers who are not familiar with the names and that there are no standardized methods for pronouncing proper names. Noise robustness is another very important factor as these calls are typically made in noisy environments, such as from a car, cafeteria, airport, etc. and over different kinds of cellular and land-line telephone channels. We have performed a systematic analysis of the speech recognition errors and tackled the issues separately with techniques ranging from weighted speaker clustering, massive adaptation, rapid and unsupervised adaptation methods to pronunciation modeling methods. We find that the decoding accuracy can be improved significantly (28% relative) in this manner. Bhuvana Ramabhadran, C. Julian Chen, Hakan Erdogan, Michael Picheny |
ICASSP | 2 |
| 2000 | Decision tree based rate of speech modeling for speech recognition
Bhuvana Ramabhadran |
INTERSPEECH | 1 |
| 2000 | Dynamic selection of feature spaces for robust speech recognition
Bhuvana Ramabhadran, Michael Picheny |
INTERSPEECH | 1 |
| 1999 | Acoustics-based baseform generation with pronunciation and/or phonotactic models
Bhuvana Ramabhadran, Sabine Deligne, Abraham Ittycheriah |
EUROSPEECH | 1 |
| 1999 | Enhanced likelihood computation using regressionabstractIn this paper, a new auditory spectrum based speech feature is proposed using sinusoidal representation and auditory model. The feature is optimized using the properties of auditory perception and masking. After quantizing and encoding the optimized feature parameters, a new speech-coding algorithm with average bit-rate of 3.25kbps is developed. The experimental results show that the synthetic speech retains most of the intelligibility and clearness of articulation of the original speech. Compared with the conventional algorithms, no voiced/unvoiced decision and pitch estimation are needed, complexity of the algorithm is much reduced, robustness and adaptation are both raised. The algorithm makes it possible to be realized with single DSP chip. Peter V. de Souza, Bhuvana Ramabhadran, Michael Picheny |
EUROSPEECH | 2 |
| 1998 | Speech recognition performance on a voicemail transcription taskabstractWe describe a new testbed for developing speech recognition algorithms-the ARRPA-sponsored voicemail transcription task, analogous to other tasks such as the Switchboard, CallHome and the Hub 4 tasks. The task involves the transcription of voicemail conversations. Voicemail represents a very large volume of real-world speech data, which is however not particularly well represented in existing databases. For instance, the Switchboard and CallHome databases contain telephone conversations between two humans, representing telephone-bandwidth spontaneous speech; the Hub 4 database contains radio broadcasts which represents different kinds of speech data such as spontaneous speech from a well-trained speaker, conversations between two humans possibly over the telephone, etc. The voicemail database on the other hand also represents telephone bandwidth spontaneous speech, however the difference with respect to the Switchboard and CallHome tasks is that the interaction is not between two humans, but rather between a human and a machine-consequently, the speech is expected to be a little more formal in its nature, without the problems of crosstalk, barge-in etc. This eliminates some of the variables and provides more controlled conditions enabling one to concentrate on the aspects of spontaneous speech and effects of the telephone channel. We describe the modality of collection of the speech data, and some algorithmic techniques that were devised based on this data. We also describe the initial results of the transcription performance on this task. Mukund Padmanabhan, Ellen Eide, Bhuvana Ramabhadran, Ganesh N. Ramaswamy, Lalit R. Bahl |
ICASSP | 3 |
| 1998 | Acoustics-only based automatic phonetic baseform generationabstractPhonetic baseforms are the basic recognition units in most speech recognition systems. These baseforms are usually determined by linguists once a vocabulary is chosen and not modified thereafter. However, several applications, such as name dialing, require the user be able to add new words to the vocabulary. These new words are often names, or task-specific jargon, that have user-specific pronunciations. This paper describes a novel method for generating phonetic transcriptions (baseforms) of words based on acoustic evidence alone. It does not require either the spelling or any prior acoustic representation of the new word, is vocabulary independent, and does not have any linguistic constraints (pronunciation rules). Our experiments demonstrate the high decoding accuracies obtained when baseforms deduced using this approach are incorporated into our speech recognizer. Also, the error rates on the added words were found to be comparable to or better than when the baseforms were derived by hand. Bhuvana Ramabhadran, Lalit R. Bahl, Peter DeSouza, Mukund Padmanabhan |
ICASSP | 1 |
| 1998 | Factor analysis invariant to linear transformations of dataabstractModeling data with Gaussian distributions is an important statistical problem. To obtain robust models one imposes constraints the means and covariances of these distributions [6, 4, 10, 8]. Constrained ML modeling implies the existence of optimal feature spaces where the constraints are more valid [2, 3]. This paper introduces one such constrained ML modeling technique called factor analysis invariant to linear transformations (FACILT) which is essentially factor analysis in optimal feature spaces. FACILT is a generalization of several existing methods for modeling covariances. This paper presents an EM algorithm for FACILT modeling. 1. Ramesh A. Gopinath, Bhuvana Ramabhadran, Satya Dharanipragada |
ICSLP | 2 |
| 1998 | Speech recognition performance on a new voicemail transcription task
Mukund Padmanabhan, Bhuvana Ramabhadran, Sankar Basu |
ICSLP | 2 |
| 1998 | Phonological rules for enhancing acoustic enrollment of unknown words
Bhuvana Ramabhadran, Abraham Ittycheriah |
ICSLP | 1 |