EDBT 2026 Demo / reviewers in the wild / expert
Pedro J. Moreno 0001
dblp:34/3391 · also Pedro Moreno Mengibar
· DBLP profile ↗
98ranked-venue papers
12as first author
28since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 87 · 11 first-author · 25 since 2021Artificial intelligence and machine learning · 51 · 6 first-author · 16 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Theory of computation · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Doc2Doc: Structure-Aware Generative Rendering for Bi-directional Document Translation
Fahad Al-Otaibi, Daulet Toibazar, Renad A. Alnuaim, Ranya A. Alkahtani, Haneen A. Alhomoud, Asma A. Ibrahim, Yazeed Alharbi, Murtadha Al-Jubran, Pedro J. Moreno 0001 |
ICDAR (2) | 9 |
| 2025 | Speech Re-Painting for Robust ASRabstractSynthetic speech is a useful source for augmentation of automatic speech recognition (ASR) systems, but there is a "sim-to-real" gap between synthetic and real speech that can limit generalization. The natural variability of real speech is essential to the training of robust ASR systems. While synthetic data augmentation can be used to approximate the variability of natural speech, however, not all aspects of variation are equally relevant for augmentation. In this work, we introduce speech re-painting, a method for in-context augmented synthesis, using target training datasets to generate new utterances guided by speech and text on the fly in a zero-shot manner. We evaluate this technique using downstream ASR word error rate (WER) using the VCTK and LibriSpeech datasets. These represent unique speaker and lexical challenges that are addressed by re-painting, realizing a reduction of WER more than 50% in particular settings. Kyle Kastner, Gary Wang, Isaac Elias, Takaaki Saeki, Pedro J. Moreno 0001, Françoise Beaufays, Andrew Rosenberg, Bhuvana Ramabhadran |
ICASSP | 5 |
| 2024 | Improving Speech Recognition for African American English with Audio ClassificationabstractAutomatic speech recognition (ASR) systems have been shown to have large quality disparities between the language varieties they are intended or expected to recognize. One way to mitigate this is to train or fine-tune models with more representative datasets. But this approach can be hindered by limited in-domain data for training and evaluation. We propose a new way to improve the robustness of a US English short-form speech recognizer using a small amount of out-of-domain (long-form) African American English (AAE) data. We use CORAAL, YouTube and Mozilla Common Voice to train an audio classifier to approximately output whether an utterance is AAE or some other variety including Mainstream American English (MAE). By combining the classifier output with coarse geographic information, we can select a subset of utterances from a large corpus of untranscribed short-form queries for semi-supervised learning at scale. Fine-tuning on this data results in a 38.5% relative word error rate disparity reduction between AAE and MAE without reducing MAE quality. Shefali Garg, Zhouyuan Huo, Khe Chai Sim, Suzan Schwartz, Mason Chua, Alëna Aksënova, Tsendsuren Munkhdalai, Levi King, Darryl Wright, Zion Mengesha, Dongseong Hwang, Tara N. Sainath, Françoise Beaufays, Pedro J. Moreno 0001 |
ICASSP | 14 |
| 2024 | Extreme Encoder Output Frame Rate Reduction: Improving Computational Latencies of Large End-to-End ModelsabstractThe accuracy of end-to-end (E2E) automatic speech recognition (ASR) models continues to improve as they are scaled to larger sizes, with some now reaching billions of parameters. Widespread deployment and adoption of these models, however, requires computationally efficient strategies for decoding. In the present work, we study one such strategy: applying multiple frame reduction layers in the encoder to compress encoder outputs into a small number of output frames. While similar techniques have been investigated in previous work, we achieve dramatically more reduction than has previously been demonstrated through the use of multiple funnel reduction layers. Through ablations, we study the impact of various architectural choices in the encoder to identify the most effective strategies. We demonstrate that we can generate one encoder output frame for every 2.56 sec of input speech, without significantly affecting word error rate on a large-scale voice search task, while improving encoder and decoder latencies by 48% and 92% respectively, relative to a strong but computationally expensive baseline. Rohit Prabhavalkar, Zhong Meng, Adam Stooke, Xingyu Cai, Yanzhang He, Arun Narayanan, Dongseong Hwang, Tara N. Sainath, Pedro J. Moreno 0001 |
ICASSP | 10 |
| 2024 | Contextual Biasing with the Knuth-Morris-Pratt Matching Algorithm
Zelin Wu, Diamantino Caseiro, Tsendsuren Munkhdalai, Khe Chai Sim, Pat Rondon, Golan Pundak, Gan Song, Rohit Prabhavalkar, Zhong Meng, Ding Zhao, Tara Sainath, Yanzhang He, Pedro J. Moreno 0001 |
INTERSPEECH | 14 |
| 2024 | Massive End-to-end Speech Recognition Models with Time ReductionabstractWeiran Wang, Rohit Prabhavalkar, Haozhe Shan, Zhong Meng, Dongseong Hwang, Qiujia Li, Khe Chai Sim, Bo Li, James Qin, Xingyu Cai, Adam Stooke, Chengjian Zheng, Yanzhang He, Tara Sainath, Pedro Moreno Mengibar. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Rohit Prabhavalkar, Haozhe Shan, Zhong Meng, Dongseong Hwang, Qiujia Li, Khe Chai Sim, Bo Li 0028, James Qin, Xingyu Cai, Adam Stooke, Chengjian Zheng, Yanzhang He, Tara N. Sainath, Pedro J. Moreno 0001 |
NAACL-HLT | 15 |
| 2024 | Aligner-Encoders: Self-Attention Transformers Can Be Self-TransducersabstractModern systems for automatic speech recognition, including the RNN-Transducer and Attention-based Encoder-Decoder (AED), are designed so that the encoder is not required to alter the time-position of information from the audio sequence into the embedding; alignment to the final text output is processed during decoding. We discover that the transformer-based encoder adopted in recent years is actually capable of performing the alignment internally during the forward pass, prior to decoding. This new phenomenon enables a simpler and more efficient model, the ''Aligner-Encoder''. To train it, we discard the dynamic programming of RNN-T in favor of the frame-wise cross-entropy loss of AED, while the decoder employs the lighter text-only recurrence of RNN-T without learned cross-attention---it simply scans embedding frames in order from the beginning, producing one token each until predicting the end-of-message. We conduct experiments demonstrating performance remarkably close to the state of the art, including a special inference configuration enabling long-form recognition. In a representative comparison, we measure the total inference time for our model to be 2x faster than RNN-T and 16x faster than AED. Lastly, we find that the audio-text alignment is clearly visible in the self-attention weights of a certain layer, which could be said to perform ''self-transduction''. Adam Stooke, Rohit Prabhavalkar, Khe Chai Sim, Pedro J. Moreno 0001 |
NeurIPS | 4 |
| 2023 | Audio-Adapterfusion: A Task-Id-Free Approach for Efficient and Non-Destructive Multi-Task Speech RecognitionabstractAdapters are an efficient, composable alternative to full fine-tuning of pre-trained models and help scale the deployment of large ASR models to many tasks. In practice, a task ID is commonly prepended to the input during inference to route to single-task adapters for the specified task. However, one major limitation of this approach is that the task ID may not be known during inference, rendering it unsuitable for most multi-task settings. To address this, we propose three novel task-ID-free methods to combine single-task adapters in multi-task ASR and investigate two learning algorithms for training. We evaluate our methods on 10 test sets from 4 diverse ASR tasks and show that our methods are non-destructive and parameter-efficient. While only updating 17 % of the model parameters, our methods can achieve an 8 % mean WER improvement relative to full fine-tuning and are on-par with task-ID adapter routing. Hillary Ngai, Rohan Agrawal, Neeraj Gaur, W. Ronny Huang, Parisa Haghani, Pedro J. Moreno 0001 |
ASRU | 6 |
| 2023 | Modular Conformer Training for Flexible End-to-End ASRabstractThe state-of-the-art conformer used in automatic speech recognition combines feed-forward, convolution and multi-headed self-attention layers in a single model that is trained end-to-end with a decoder network. While this end-to-end training is simple and beneficial for word error rate, it restricts the ability to perform inference with the model at different operating points of word error rate and latency. Existing approaches to overcome this limitation include cascaded encoders and variable attention context models. We propose an alternative approach, called Modular Conformer training, which splits the Conformer model into a backbone convolutional model and attention submodels, which are added at each layer. We conduct experiments with a few training techniques on the Librispeech and Librilight corpus. We show that dropping-out the attention layers during the training of the backbone model allows for the largest WER improvements upon adding fine-tuned attention submodels, without impacting the WER of the backbone model itself. Kartik Audhkhasi, Brian Farris, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
ICASSP | 4 |
| 2023 | Large-Scale Language Model Rescoring on Long-Form DataabstractIn this work, we study the impact of Large-scale Language Models (LLM) on Automated Speech Recognition (ASR) of YouTube videos, which we use as a source for long-form ASR. We demonstrate up to 8% relative reduction in Word Error Eate (WER) on US English (en-us) and code-switched Indian English (en-in) long-form ASR test sets and a reduction of up to 30% relative on Salient Term Error Rate (STER) over a strong first-pass baseline that uses a maximum-entropy based language model. Improved lattice processing that results in a lattice with a proper (non-tree) digraph topology and carrying context from the 1-best hypothesis of the previous segment(s) results in significant wins in rescoring with LLMs. We also find that the gains in performance from the combination of LLMs trained on vast quantities of available data (such as C4 [1]) and conventional neural LMs is additive and significantly outperforms a strong first-pass baseline with a maximum entropy LM. Tongzhou Chen, Cyril Allauzen, Daniel S. Park, David Rybach, W. Ronny Huang, Rodrigo Cabrera, Kartik Audhkhasi, Bhuvana Ramabhadran, Pedro J. Moreno 0001, Michael Riley 0001 |
ICASSP | 10 |
| 2023 | Re-investigating the Efficient Transfer Learning of Speech Foundation Model using Feature Fusion Methods
Zhouyuan Huo, Khe Chai Sim, Dongseong Hwang, Tsendsuren Munkhdalai, Tara N. Sainath, Pedro J. Moreno 0001 |
INTERSPEECH | 6 |
| 2023 | Modular Domain Adaptation for Conformer-Based Streaming ASR
Qiujia Li, Bo Li 0028, Dongseong Hwang, Tara N. Sainath, Pedro J. Moreno 0001 |
INTERSPEECH | 5 |
| 2022 | Tts4pretrain 2.0: Advancing the use of Text and Speech in ASR Pretraining with Consistency and Contrastive LossesabstractAn effective way to learn representations from untranscribed speech and unspoken text with linguistic/lexical representations derived from synthesized speech was introduced in tts4pretrain [1]. However, the representations learned from synthesized and real speech are likely to be different, potentially limiting the improvements from incorporating unspoken text. In this paper, we introduce learning from supervised speech earlier on in the training process with consistency-based regularization between real and synthesized speech. This allows for better learning of shared speech and text representations. Thus, we introduce a new objective, with encoder and decoder consistency and contrastive regularization between real and synthesized speech derived from the labeled corpora during the pretraining stage. We show that the new objective leads to more similar representations derived from speech and text that help downstream ASR. The proposed pretraining method yields Word Error Rate (WER) reductions of 7-21% relative on six public corpora, Librispeech, AMI, TEDLIUM, Common Voice, Switchboard, CHiME-6, over a state-of-the-art baseline pretrained with wav2vec2.0 and 2-17% over the previously proposed tts4pretrain. The proposed method outperforms the supervised SpeechStew by up to 17%. Moreover, we show that the proposed method also yields WER reductions on larger data sets by evaluating on a large resource, in-house Voice Search task and streaming ASR. Zhehuai Chen, Yu Zhang 0033, Andrew Rosenberg, Bhuvana Ramabhadran, Pedro J. Moreno 0001, Gary Wang |
ICASSP | 5 |
| 2022 | Multilingual Second-Pass Rescoring for Automatic Speech Recognition SystemsabstractSecond-pass rescoring is a well known technique to improve the performance of Automatic Speech Recognition (ASR) systems. Neural Oracle Search (NOS), which selects the most likely hypothesis from an N-best hypothesis list by integrating information from multiple sources, such as the input acoustic representations, N-best hypotheses, additional first-pass statistics, and unpaired textual information through an external language model, has shown success in rescoring for RNN-T first-pass models. Multilingual first-pass speech recognition models often outperform their monolingual counterparts when trained on related or low-resource languages. In this paper, we investigate the use of the NOS rescoring model on a first-pass multilingual model and show that similar to the first-pass model, the rescoring model can be made multilingual. Our first-pass multilingual model does not require a language-id and we make a realistic assumption that an estimate of the language-id would be available for second-pass rescoring. We conduct comprehensive experiments on two sets of languages, one consisting of related low-resource languages, and the other with a high-resource language added to the first set to analyze the performance of the multilingual NOS rescorer under different settings. Our experimental results show that, multilingual NOS can improve the first-pass multilingual model resulting in average word error rate reduction of 9.4% in the first case, and 8.4% in the second, and out-performing the monolingual counterparts in both cases. Neeraj Gaur, Tongzhou Chen, Ehsan Variani, Parisa Haghani, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
ICASSP | 6 |
| 2022 | Analysis of Self-Attention Head Diversity for Conformer-based Automatic Speech RecognitionabstractAttention layers are an integral part of modern end-to-end automatic speech recognition systems, for instance as part of the Transformer or Conformer architecture.Attention is typically multi-headed, where each head has an independent set of learned parameters and operates on the same input feature sequence.The output of multi-headed attention is a fusion of the outputs from the individual heads.We empirically analyze the diversity between representations produced by the different attention heads and demonstrate that the heads become highly correlated during the course of training.We investigate a few approaches to increasing attention head diversity, including using different attention mechanisms for each head and auxiliary training loss functions to promote head diversity.We show that introducing diversity-promoting auxiliary loss functions during training is a more effective approach, and obtain WER improvements of up to 6% relative on the Librispeech corpus.Finally, we draw a connection between the diversity of attention heads and the similarity of the gradients of head parameters. Kartik Audhkhasi, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
INTERSPEECH | 4 |
| 2022 | A Scalable Model Specialization Framework for Training and Inference using Submodels and its Application to Speech Model PersonalizationabstractModel fine-tuning and adaptation have become a common approach for model specialization for downstream tasks or domains. Fine-tuning the entire model or a subset of the parameters using light-weight adaptation has shown considerable success across different specialization tasks. Fine-tuning a model for a large number of domains typically requires starting a new training job for every domain posing scaling limitations. Once these models are trained, deploying them also poses significant scalability challenges for inference for real-time applications. In this paper, building upon prior light-weight adaptation techniques, we propose a modular framework that enables us to substantially improve scalability for model training and inference. We introduce Submodels that can be quickly and dynamically loaded for on-the-fly inference. We also propose multiple approaches for training those Submodels in parallel using an embedding space in the same training job. We test our framework on an extreme use-case which is speech model personalization for atypical speech, requiring a Submodel for each user. We obtain 128x Submodel throughput with a fixed computation budget without a loss of accuracy. We also show that learning a speaker-embedding space can scale further and reduce the amount of personalization training data required per speaker. Fadi Biadsy, Youzheng Chen, Oleg Rybakov, Andrew Rosenberg, Pedro J. Moreno 0001 |
INTERSPEECH | 6 |
| 2022 | MAESTRO: Matched Speech Text Representations through Modality MatchingabstractWe present Maestro, a self-supervised training method to unify representations learnt from speech and text modalities.Self-supervised learning from speech signals aims to learn the latent structure inherent in the signal, while self-supervised learning from text attempts to capture lexical information.Learning aligned representations from unpaired speech and text sequences is a challenging task.Previous work either implicitly enforced the representations learnt from these two modalities to be aligned in the latent space through multitasking and parameter sharing or explicitly through conversion of modalities via speech synthesis.While the former suffers from interference between the two modalities, the latter introduces additional complexity.In this paper, we propose Maestro, a novel algorithm to learn unified representations from both these modalities simultaneously that can transfer to diverse downstream tasks such as Automated Speech Recognition (ASR) and Speech Translation (ST).Maestro learns unified representations through sequence alignment, duration prediction and matching embeddings in the learned space through an aligned masked-language model loss.We establish a new state-of-the-art (SOTA) on VoxPopuli multilingual ASR with a 8% relative reduction in Word Error Rate (WER), multidomain SpeechStew ASR (3.7% relative) and 21 languages to English multilingual ST on CoVoST 2 with an improvement of 2.8 BLEU averaged over 21 languages. Zhehuai Chen, Yu Zhang 0033, Andrew Rosenberg, Bhuvana Ramabhadran, Pedro J. Moreno 0001, Ankur Bapna, Heiga Zen |
INTERSPEECH | 5 |
| 2022 | Non-Parallel Voice Conversion for ASR AugmentationabstractAutomatic speech recognition (ASR) needs to be robust to speaker differences.Voice Conversion (VC) modifies speaker characteristics of input speech.This is an attractive feature for ASR data augmentation.In this paper, we demonstrate that voice conversion can be used as a data augmentation technique to improve ASR performance, even on LibriSpeech, which contains 2,456 speakers.For ASR augmentation, it is necessary that the VC model be robust to a wide range of input speech.This motivates the use of a non-autoregressive, non-parallel VC model, and the use of a pretrained ASR encoder within the VC model.This work suggests that despite including many speakers, speaker diversity may remain a limitation to ASR quality.Finally, interrogation of our VC performance has provided useful metrics for objective evaluation of VC quality. Gary Wang, Andrew Rosenberg, Bhuvana Ramabhadran, Fadi Biadsy, Jesse Emond, Pedro J. Moreno 0001 |
INTERSPEECH | 7 |
| 2022 | Maestro-U: Leveraging Joint Speech-Text Representation Learning for Zero Supervised Speech ASRabstractTraining state-of-the-art Automated Speech Recognition (ASR) models typically requires a substantial amount of transcribed speech. In this work, we demonstrate that a modality-matched joint speech and text model introduced in [1] can be leveraged to train a massively multilingual ASR model without any supervised (manually transcribed) speech for some languages. This paper explores the use of jointly learnt speech and text representations in a massively multilingual, zero supervised speech, real-world setting to expand the set of languages covered by ASR with only unlabeled speech and text in the target languages. Using the FLEURS dataset, we define the task to cover 102 languages, where transcribed speech is available in 52 of these languages and can be used to improve end-to-end ASR quality on the remaining 50. First, we show that by combining speech representations with byte-level text representations and use of language embeddings, we can dramatically reduce the Character Error Rate (CER) on languages with no supervised speech from 64.8% to 30.8%, a relative reduction of 53%. Second, using a subset of South Asian languages we show that Maestro-U can promote knowledge transfer from languages with supervised speech even when there is limited to no graphemic overlap. Overall, Maestro-U closes the gap to oracle performance by 68.5% relative and reduces the CER of 19 languages below 15%. Zhehuai Chen, Ankur Bapna, Andrew Rosenberg, Yu Zhang 0033, Bhuvana Ramabhadran, Pedro J. Moreno 0001, Nanxin Chen |
SLT | 6 |
| 2022 | Modular Hybrid Autoregressive TransducerabstractText-only adaptation of a transducer model remains challenging for end-to-end speech recognition since the transducer has no clearly separated acoustic model (AM), language model (LM) or blank model. In this work, we propose a modular hybrid autoregressive transducer (MHAT) that has structurally separated label and blank decoders to predict label and blank distributions, respectively, along with a shared acoustic encoder. The encoder and label decoder outputs are directly projected to AM and internal LM scores and then added to compute label posteriors. We train MHAT with an internal LM loss and a HAT loss to ensure that its internal LM becomes a standalone neural LM that can be effectively adapted to text. Moreover, text adaptation of MHAT fosters a much better LM fusion than internal LM subtraction-based methods. On Google's large-scale production data, a multi-domain MHAT adapted with 100B sentences achieves relative WER reductions of up to 12.4% without LM fusion and 21.5% with LM fusion from 400K-hour trained HAT. Zhong Meng, Tongzhou Chen, Rohit Prabhavalkar, Yu Zhang 0033, Gary Wang, Kartik Audhkhasi, Jesse Emond, Trevor Strohman, Bhuvana Ramabhadran, W. Ronny Huang, Ehsan Variani, Pedro J. Moreno 0001 |
SLT | 13 |
| 2022 | G-Augment: Searching for the Meta-Structure of Data Augmentation Policies for ASRabstractData augmentation is a ubiquitous technique used to provide robustness to automatic speech recognition (ASR) training. However, even as so much of the ASR training process has become automated and more “end-to-end,” the data augmentation policy (what augmentation functions to use, and how to apply them) remains hand-crafted. We present G(raph)-Augment, a technique to define the augmentation space as directed acyclic graphs (DAGs) and search over this space to optimize the augmentation policy itself. We show that given the same computational budget, policies produced by G-Augment are able to perform better than SpecAugment policies obtained by random search on fine-tuning tasks on CHiME-6 and AMI. G-Augment is also able to establish a new state-of-the-art ASR performance on the CHiME-6 evaluation set (30.7% WER). We further demonstrate that G- Augment policies show better transfer properties across warm-start to cold-start training and model size compared to random-searched SpecAugment policies. Gary Wang, Ekin Dogus Cubuk, Andrew Rosenberg, Shuyang Cheng, Ron J. Weiss, Bhuvana Ramabhadran, Pedro J. Moreno 0001, Quoc V. Le, Daniel S. Park |
SLT | 7 |
| 2021 | Injecting Text in Self-Supervised Speech PretrainingabstractSelf-supervised pretraining for Automated Speech Recognition (ASR) has shown varied degrees of success. In this paper, we propose to jointly learn representations during pretraining from two different modalities: speech and text. The proposed method, tts4pretrain complements the power of contrastive learning in self-supervision with linguistic/lexical representations derived from synthesized speech, effectively learning from untranscribed speech and unspoken text. Lexical learning in the speech encoder is enforced through an additional sequence loss term that is coupled with contrastive loss during pretraining. We demonstrate that this novel pretraining method yields Word Error Rate (WER) reductions of 10% relative on the well-benchmarked, Librispeech task over a state-of-the-art baseline pretrained with wav2vec2.0 only. The proposed method also serves as an effective strategy to compensate for the lack of transcribed speech, effectively matching the performance of 5000 hours of transcribed speech with just 100 hours of transcribed speech on the AMI meeting transcription task. Finally, we demonstrate WER reductions of up to 15% on an inhouse Voice Search task over traditional pretraining. Incorporating text into encoder pretraining is complimentary to rescoring with a larger or in-domain language model, resulting in additional 6% relative reduction in WER. Zhehuai Chen, Yu Zhang 0033, Andrew Rosenberg, Bhuvana Ramabhadran, Gary Wang, Pedro J. Moreno 0001 |
ASRU | 6 |
| 2021 | Extending Parrotron: An End-to-End, Speech Conversion and Speech Recognition Model for Atypical SpeechabstractWe present an extended Parrotron model: a single, end-to-end network that enables voice conversion and recognition simultaneously. Input spectrograms are transformed to output spectrograms in the voice of a predetermined target speaker while also generating hypotheses in a target vocabulary. We study the performance of this novel architecture, which jointly predicts speech and text, on atypical (e.g. dysarthric) speech. We show that with as little as an hour of atypical speech, speaker adaptation can yield a 77% relative reduction in Word Error Rate (WER), measured by ASR performance on the converted speech. We also show that data augmentation using a customized synthesizer built on atypical speech can provide an additional 10% relative improvement over the best speaker-adapted model. Finally, we show how these methods generalize across 8 types of atypical speech for a range of speech impairment severities. Rohan Doshi, Youzheng Chen, Liyang Jiang, Fadi Biadsy, Bhuvana Ramabhadran, Fang Chu, Andrew Rosenberg, Pedro J. Moreno 0001 |
ICASSP | 9 |
| 2021 | Mixture of Informed Experts for Multilingual Speech RecognitionabstractWhen trained on related or low-resource languages, multilingual speech recognition models often outperform their monolingual counterparts. However, these models can suffer from loss in performance for high resource or unrelated languages. We investigate the use of a mixture-of-experts approach to assign per-language parameters in the model to increase network capacity in a structured fashion. We introduce a novel variant of this approach, ‘informed experts’, which attempts to tackle inter-task conflicts by eliminating gradients from other tasks in these task-specific parameters. We conduct experiments on a real-world task with English, French and four dialects of Arabic to show the effectiveness of our approach. Our model matches or outperforms the monolingual models for almost all languages, with gains of as much as 31% relative. Our model also outperforms the baseline multilingual model for all languages by up to 9% relative. Neeraj Gaur, Brian Farris, Parisa Haghani, Isabel Leal, Pedro J. Moreno 0001, Manasa Prasad, Bhuvana Ramabhadran |
ICASSP | 5 |
| 2021 | Mixture Model Attention: Flexible Streaming and Non-Streaming Automatic Speech Recognition
Kartik Audhkhasi, Tongzhou Chen, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
Interspeech | 4 |
| 2021 | Conformer Parrotron: A Faster and Stronger End-to-End Speech Conversion and Recognition Model for Atypical Speech
Zhehuai Chen, Bhuvana Ramabhadran, Fadi Biadsy, Youzheng Chen, Liyang Jiang, Fang Chu, Rohan Doshi, Pedro J. Moreno 0001 |
Interspeech | 9 |
| 2021 | Semi-Supervision in ASR: Sequential MixMatch and Factorized TTS-Based Augmentation
Zhehuai Chen, Andrew Rosenberg, Yu Zhang 0033, Heiga Zen, Mohammadreza Ghodsi, Jesse Emond, Gary Wang, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
Interspeech | 10 |
| 2021 | Self-Adaptive Distillation for Multilingual Speech Recognition: Leveraging Student Independence
Isabel Leal, Neeraj Gaur, Parisa Haghani, Brian Farris, Pedro J. Moreno 0001, Manasa Prasad, Bhuvana Ramabhadran |
Interspeech | 5 |
| 2020 | Neural Oracle Search on N-BEST HypothesesabstractIn this paper, we propose a neural search algorithm to select the most likely hypothesis using a sequence of acoustic representations and multiple hypotheses as input. The algorithm provides a sequence level score for each audio-hypothesis pair that is obtained by integrating information from multiple sources, such as the input acoustic representations, N-best hypotheses, additional 1st-pass statistics, and unpaired textual information through an external language model. These scores are then used to map the search problem of identifying the most likely hypothesis to a sequence classification problem. The definition of the proposed algorithm is broad enough to allow its use as an alternative to beam search in the 1st-pass or as a 2nd-pass, rescoring step. This algorithm achieves up to 12% relative reductions in Word Error Rate (WER) across several languages over state-of-the-art baselines with relatively few additional parameters. We also propose the use of a binary classifier gating function that can learn to trigger the 2nd-pass neural search model when the 1-best hypothesis is not the oracle hypothesis, thereby avoiding extra computation. Ehsan Variani, Tongzhou Chen, James Apfel, Bhuvana Ramabhadran, Seungji Lee, Pedro J. Moreno 0001 |
ICASSP | 6 |
| 2020 | Improving Speech Recognition Using Consistent Predictions on Synthesized SpeechabstractSpeech synthesis has advanced to the point of being close to indistinguishable from human speech. However, efforts to train speech recognition systems on synthesized utterances have not been able to show that synthesized data can be effectively used to augment or replace human speech. In this work, we demonstrate that promoting consistent predictions in response to real and synthesized speech enables significantly improved speech recognition performance. We also find that training on 460 hours of LibriSpeech augmented with 500 hours of transcripts (without audio) performance is within 0.2% WER of a system trained on 960 hours of transcribed audio. This suggests that with this approach, when there is sufficient text available, reliance on transcribed audio can be cut nearly in half. Gary Wang, Andrew Rosenberg, Zhehuai Chen, Yu Zhang 0033, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
ICASSP | 7 |
| 2020 | Improving Speech Recognition Using GAN-Based Speech Synthesis and Contrastive Unspoken Text Selection
Zhehuai Chen, Andrew Rosenberg, Yu Zhang 0033, Gary Wang, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
INTERSPEECH | 6 |
| 2020 | SCADA: Stochastic, Consistent and Adversarial Data Augmentation to Improve ASR
Gary Wang, Andrew Rosenberg, Zhehuai Chen, Yu Zhang 0033, Bhuvana Ramabhadran, Pedro J. Moreno 0001 |
INTERSPEECH | 6 |
| 2020 | Multilingual Speech Recognition with Self-Attention Structured Parameterization
Parisa Haghani, Anshuman Tripathi, Bhuvana Ramabhadran, Brian Farris, Hainan Xu, Han Lu 0003, Hasim Sak, Isabel Leal, Neeraj Gaur, Pedro J. Moreno 0001 |
INTERSPEECH | 11 |
| 2019 | Speech Recognition with Augmented Synthesized SpeechabstractRecent success of the Tacotron speech synthesis architecture and its variants in producing natural sounding multi-speaker synthesized speech has raised the exciting possibility of replacing expensive, manually transcribed, domain-specific, human speech that is used to train speech recognizers. The multi-speaker speech synthesis architecture can learn latent embedding spaces of prosody, speaker and style variations derived from input acoustic representations thereby allowing for manipulation of the synthesized speech. In this paper, we evaluate the feasibility of enhancing speech recognition performance using speech synthesis using two corpora from different domains. We explore algorithms to provide the necessary acoustic and lexical diversity needed for robust speech recognition. Finally, we demonstrate the feasibility of this approach as a data augmentation strategy for domain-transfer. We find that improvements to speech recognition performance is achievable by augmenting training data with synthesized material. However, there remains a substantial gap in performance between recognizers trained on human speech those trained on synthesized speech. Andrew Rosenberg, Yu Zhang 0033, Bhuvana Ramabhadran, Ye Jia, Pedro J. Moreno 0001, Zelin Wu |
ASRU | 5 |
| 2019 | Leveraging Language ID in Multilingual End-to-End Speech RecognitionabstractRecent advances in end-to-end speech recognition have made it possible to build multilingual models, capable of recognizing speech in multiple languages. Multilingual models can outperform their monolingual counterparts, depending on the amount of training data and the relatedness of languages. However, in some cases, these models rely on having perfect knowledge of the language being spoken; that is, they expect to be provided with an external language ID that augments the input features or modulates internal layers of the network. In this paper, we introduce a novel technique for inferring the language ID in a streaming fashion using RNN-T, and a novel loss function that pressures the model to identify the language after as few frames as possible. The output of this streaming language-ID model is used in training and inference of a multilingual recognition model. We show the effectiveness of our approach through experiments on two sets of languages, one consisting of different dialects of Arabic, and the other consisting of Nordic languages, Finnish and Dutch. Austin Waters, Neeraj Gaur, Parisa Haghani, Pedro J. Moreno 0001, Zhongdi Qu |
ASRU | 4 |
| 2019 | Parrotron: An End-to-End Speech-to-Speech Conversion Model and its Applications to Hearing-Impaired Speech and Speech SeparationabstractWe describe Parrotron, an end-to-end-trained speech-to-speech conversion model that maps an input spectrogram directly to another spectrogram, without utilizing any intermediate discrete representation.The network is composed of an encoder, spectrogram and phoneme decoders, followed by a vocoder to synthesize a time-domain waveform.We demonstrate that this model can be trained to normalize speech from any speaker regardless of accent, prosody, and background noise, into the voice of a single canonical target speaker with a fixed accent and consistent articulation and prosody.We further show that this normalization model can be adapted to normalize highly atypical speech from a deaf speaker, resulting in significant improvements in intelligibility and naturalness, measured via a speech recognizer and listening tests.Finally, demonstrating the utility of this model on other speech tasks, we show that the same model architecture can be trained to perform a speech separation task. Fadi Biadsy, Ron J. Weiss, Pedro J. Moreno 0001, Dimitri Kanvesky, Ye Jia |
INTERSPEECH | 3 |
| 2018 | Modeling Non-Linguistic Contextual Signals in LSTM Language Models Via Domain AdaptationabstractLanguage Models (LMs) for Automatic Speech Recognition (ASR) can benefit from utilizing non-linguistic contextual signals in modeling. Examples of these signals include the geographical location of the user speaking to the system and/or the identity of the application (app) being spoken to. In practice, the vast majority of input speech queries typically lack annotations of such signals, which poses a challenge to directly train domain-specific LMs. To obtain robust domain LMs, generally an LM which has been pre-trained on general data will be adapted to specific domains. We propose four domain adaptation schemes to improve the domain performance of Long Short-Term Memory (LSTM) LMs, by incorporating app based contextual signals of voice search queries. We show that most of our adaptation strategies are effective, reducing word perplexity up to 21 % relative to a fine-tuned baseline on a held-out domain-specific development set. Initial experiments using a state-of-the-art Italian ASR system show a 3 % relative reduction in WER on top of an unadapted 5-gram LM. In addition, human evaluations show significant improvements on sub-domains from using app signals. Shankar Kumar, Fadi Biadsy, Michael Nirschl, Tomas Vykruta, Pedro J. Moreno 0001 |
ICASSP | 6 |
| 2018 | Hybrid Lstm-Fsmn Networks for Acoustic ModelingabstractThis paper describes a series of experiments with neural networks containing long short-term memory (LSTM) [1] and feedforward sequential memory network (FSMN) [2]-[4] layers trained with the connectionist temporal classification (CTC) [5] criteria for acoustic modeling. We propose using a hybrid LSTM/FSMN (FLMN) architecture as an enhancement to conventional LSTM-only acoustic models. The addition of FSMN layers allows the network to model a fixed size representation of future context suitable for online speech recognition. Our experiments show that FLMN acoustic models significantly outperform conventional LSTM. We also compare the FLMN architecture with other methods of modeling future context. Finally, we present a modification of the FSMN architecture that improves performance by reducing the width of the FSMN output. Asa Oines, Eugene Weinstein, Pedro J. Moreno 0001 |
ICASSP | 3 |
| 2018 | Multilingual Speech Recognition with a Single End-to-End ModelabstractTraining a conventional automatic speech recognition (ASR) system to support multiple languages is challenging because the sub-word unit, lexicon and word inventories are typically language specific. In contrast, sequence-to-sequence models are well suited for multilingual ASR because they encapsulate an acoustic, pronunciation and language model jointly in a single network. In this work we present a single sequence-to-sequence ASR model trained on 9 different Indian languages, which have very little overlap in their scripts. Specifically, we take a union of language-specific grapheme sets and train a grapheme-based sequence-to-sequence model jointly on data from all languages. We find that this model, which is not explicitly given any information about language identity, improves recognition performance by 21% relative compared to analogous sequence-to-sequence models trained on each language individually. By modifying the model to accept a language identifier as an additional input feature, we further improve performance by an additional 7% relative and eliminate confusion between different languages. Shubham Toshniwal, Tara N. Sainath, Ron J. Weiss, Bo Li 0028, Pedro J. Moreno 0001, Eugene Weinstein, Kanishka Rao |
ICASSP | 5 |
| 2018 | Semantic Lattice Processing in Contextual Automatic Speech Recognition for Google Assistant
Leonid Velikovich, Justin Scheiner, Petar S. Aleksic, Pedro J. Moreno 0001, Michael Riley 0001 |
INTERSPEECH | 5 |
| 2018 | Transliteration Based Approaches to Improve Code-Switched Speech Recognition PerformanceabstractCode-switching is a commonly occurring phenomenon in many multilingual communities, wherein a speaker switches between languages within a single utterance. Conventional Word Error Rate (WER) is not sufficient for measuring the performance of code-mixed languages due to ambiguities in transcription, misspellings and borrowing of words from two different writing systems. These rendering errors artificially inflate the WER of an Automated Speech Recognition (ASR) system and complicate its evaluation. Furthermore, these errors make it harder to accurately evaluate modeling errors originating from code-switched language and acoustic models. In this work, we propose the use of a new metric, transliteration-optimized Word Error Rate (toWER) that smoothes out many of these irregularities by mapping all text to one writing system and demonstrate a correlation with the amount of code-switching present in a language. We also present a novel approach to acoustic and language modeling for bilingual code-switched Indic languages using the same transliteration approach to normalize the data for three types of language models, namely, a conventional n-gram language model, a maximum entropy based language model and a Long Short Term Memory (LSTM) language model, and a state-of-the-art Connectionist Temporal Classification (CTC) acoustic model. We demonstrate the robustness of the proposed approach on several Indic languages from Google Voice Search traffic with significant gains in ASR performance up to 10% relative over the state-of-the-art baseline. Jesse Emond, Bhuvana Ramabhadran, Brian Roark, Pedro J. Moreno 0001 |
SLT | 4 |
| 2018 | From Audio to Semantics: Approaches to End-to-End Spoken Language UnderstandingabstractConventional spoken language understanding systems consist of two main components: an automatic speech recognition module that converts audio to a transcript, and a natural language understanding module that transforms the resulting text (or top N hypotheses) into a set of domains, intents, and arguments. These modules are typically optimized independently. In this paper, we formulate audio to semantic understanding as a sequence-to-sequence problem [1]. We propose and compare various encoder-decoder based approaches that optimize both modules jointly, in an end-to-end manner. Evaluations on a real-world task show that 1) having an intermediate text representation is crucial for the quality of the predicted semantics, especially the intent arguments and 2) jointly optimizing the full system improves overall accuracy of prediction. Compared to independently trained models, our best jointly trained model achieves similar domain and intent prediction F1 scores, but improves argument word error rate by 18% relative. Parisa Haghani, Arun Narayanan, Michiel Bacchiani, Galen Chuang, Neeraj Gaur, Pedro J. Moreno 0001, Rohit Prabhavalkar, Zhongdi Qu, Austin Waters |
SLT | 6 |
| 2017 | Syllable-based acoustic modeling with CTC-SMBR-LSTMabstractWe explore the feasibility of training long short-term memory (LSTM) recurrent neural networks (RNNs) with syllables, rather than phonemes, as outputs. Syllables are a natural choice of linguistic unit for modeling the acoustics of languages such as Mandarin Chinese, due to the inherent nature of the syllable as an elemental pronunciation construct and the limited size of the syllable set for such languages (around 1400 syllables for Mandarin). Our models are trained with Connectionist Temporal Classification (CTC) and state-level minimum Bayes risk (sMBR) loss using asynchronous stochastic gradient descent (ASGD) utilizing a parallel computation infrastructure for large-scale training. Our acoustic models operate on feature frames computed every 30ms, which makes them well suited for modeling syllables rather than phonemes, which can have a shorter duration. Additionally, when compared to wordlevel modeling, syllables have the advantage of avoiding out-of-vocabulary (OOV) model outputs. Our experiments on a Mandarin voice search task show that syllable-output models can perform better than context-independent (CI) phone-output models, and can give similar performance as our state-of-the-art context-dependent (CD) models. Additionally, decoding with syllable-output models is substantially faster than with CI models or with CD models. We demonstrate that these improvements are maintained when the model is trained to recognize both Mandarin syllables and English phonemes. Zhongdi Qu, Parisa Haghani, Eugene Weinstein, Pedro J. Moreno 0001 |
ASRU | 4 |
| 2016 | Selection and combination of hypotheses for dialectal speech recognitionabstractWhile research has often shown that building dialect-specific Automatic Speech Recognizers is the optimal approach to dealing with dialectal variations of the same language, we have observed that dialect-specific recognizers do not always output the best recognitions. Often enough, another dialectal recognizer outputs a better recognition than the dialect-specific one. In this paper, we present two methods to select and combine the best decoded hypothesis from a pool of dialectal recognizers. We follow a Machine Learning approach and extract features from the Speech Recognition output along with Word Embeddings and use Shallow Neural Networks for classification. Our experiments using Dictation and Voice Search data from the main four Arabic dialects show good WER improvements for the hypothesis selection scheme, reducing the WER by 2.1 to 12.1% depending on the test set, and promising results for the hypotheses combination scheme. Victor Soto, Olivier Siohan, Mohamed G. Elfeky, Pedro J. Moreno 0001 |
ICASSP | 4 |
| 2016 | Towards acoustic model unification across dialectsabstractAcoustic model performance typically decreases when evaluated on a dialectal variation of the same language that was not used during training. Similarly, models simultaneously trained on a group of dialects tend to underperform dialect-specific models. In this paper, we report on our efforts towards building a unified acoustic model that can serve a multi-dialectal language. Two techniques are presented: Distillation and MultiTask Learning (MTL). In Distillation, we use an ensemble of dialect-specific acoustic models and distill its knowledge in a single model. In MTL, we utilize multitask learning to train a unified acoustic model that learns to distinguish dialects as a side task. We show that both techniques are superior to the jointly-trained model that is trained on all dialectal data, reducing word error rates by 4:2% and 0:6%, respectively. While achieving this improvement, neither technique degrades the performance of the dialect-specific models by more than 3:4%. Mohamed G. Elfeky, Meysam Bastani, Xavier Velez, Pedro J. Moreno 0001, Austin Waters |
SLT | 4 |
| 2016 | High quality agreement-based semi-supervised training data for acoustic modelingabstractThis paper describes a new technique to automatically obtain large high-quality training speech corpora for acoustic modeling. Traditional approaches select utterances based on confidence thresholds and other heuristics. We propose instead to use an ensemble approach: we transcribe each utterance using several recognizers, and only keep those on which they agree. The recognizers we use are trained on data from different dialects of the same language, and this diversity leads them to make different mistakes in transcribing speech utterances. In this work we show, however, that when they agree, this is an extremely strong signal that the transcript is correct. This allows us to produce automatically transcribed speech corpora that are superior in transcript correctness even to those manually transcribed by humans. Furthermore, we show that using the produced semi-supervised data sets, we can train new acoustic models which outperform those trained solely on previously available data sets. Félix de Chaumont Quitry, Asa Oines, Pedro J. Moreno 0001, Eugene Weinstein |
SLT | 3 |
| 2016 | On the use of deep feedforward neural networks for automatic language identificationabstractIn this work, we present a comprehensive study on the use of deep neural networks (DNNs) for automatic language identification (LID). Motivated by the recent success of using DNNs in acoustic modeling for speech recognition, we adapt DNNs to the problem of identifying the language in a given utterance from its short-term acoustic features. We propose two different DNN-based approaches. In the first one, the DNN acts as an end-to-end LID classifier, receiving as input the speech features and providing as output the estimated probabilities of the target languages. In the second approach, the DNN is used to extract bottleneck features that are then used as inputs for a state-of-the-art i-vector system. Experiments are conducted in two different scenarios: the complete NIST Language Recognition Evaluation dataset 2009 (LRE'09) and a subset of the Voice of America (VOA) data from LRE'09, in which all languages have the same amount of training data. Results for both datasets demonstrate that the DNN-based systems significantly outperform a state-of-art i-vector system when dealing with short-duration utterances. Furthermore, the combination of the DNN-based and the classical i-vector system leads to additional performance improvements (up to 45% of relative improvement in both EER and Cavg on 3s and 10s conditions, respectively). Ignacio López-Moreno, Javier Gonzalez-Dominguez, David Martinez, Oldrich Plchot, Joaquín González-Rodríguez, Pedro J. Moreno 0001 |
Comput. Speech Lang. | 6 |
| 2015 | Improved recognition of contact names in voice commandsabstractThe recognition of contact names in mobile-device voice commands is a challenging problem. Some of the difficulties include potentially infinite vocabularies, low probability of contact tokens in the language model (LM), increased false triggering of contact voice commands when none are spoken, and very large and noisy contact name lists. In this paper we suggest solutions for each of these difficulties. We address low prior probability and out-of-vocabulary contact name problems by using class-based language models, and creating on-the-fly user dependent small language models containing only relevant names. These models are compiled dynamically based on analysis of the mobile device state. Since these solutions can increase biasing towards contact names during recognition, it is crucial to monitor false triggering. To properly balance this bias we introduce the concept of a contacts insertion reward. This reward is tuned using both positive and negative test sets. We show significant recognition performance improvements on data sets in three languages, without negatively impacting the overall system performance. The improvements are obtained in both offline evaluations as well as on live traffic experiments. Petar S. Aleksic, Cyril Allauzen, David Elson, Aleksandar Kracun, Diego Melendo Casado, Pedro J. Moreno 0001 |
ICASSP | 6 |
| 2015 | Bringing contextual information to google speech recognitionabstractIn automatic speech recognition on mobile devices, very often what a user says strongly depends on the particular context he or she is in. The n-grams relevant to the context are often not known in advance. The context can depend on, for example, particular dialog state, options presented to the user, conversation topic, location, etc. Speech recognition of sentences that include these n-grams can be challenging, as they are often not well represented in a language model (LM) or even include out-of-vocabulary (OOV) words. In this paper, we propose a solution for using contextual information to improve speech recognition accuracy. We utilize an on-the-fly rescoring mechanism to adjust the LM weights of a small set of n-grams relevant to the particular context during speech decoding. Our solution handles out of vocabulary words. It also addresses efficient combination of multiple sources of context and it even allows biasing class based language models. We show significant speech recognition accuracy improvements on several datasets, using various types of contexts, without negatively impacting the overall system. The improvements are obtained in both offline and live experiments. Petar S. Aleksic, Mohammadreza Ghodsi, Assaf Hurwitz Michaely, Cyril Allauzen, Keith B. Hall, Brian Roark, David Rybach, Pedro J. Moreno 0001 |
INTERSPEECH | 8 |
| 2015 | Frame-by-frame language identification in short utterances using deep neural networks
Javier Gonzalez-Dominguez, Ignacio López-Moreno, Pedro J. Moreno 0001, Joaquín González-Rodríguez |
Neural Networks | 3 |
| 2014 | Automatic language identification using deep neural networksabstractThis work studies the use of deep neural networks (DNNs) to address automatic language identification (LID). Motivated by their recent success in acoustic modelling, we adapt DNNs to the problem of identifying the language of a given spoken utterance from short-term acoustic features. The proposed approach is compared to state-of-the-art i-vector based acoustic systems on two different datasets: Google 5M LID corpus and NIST LRE 2009. Results show how LID can largely benefit from using DNNs, especially when a large amount of training data is available. We found relative improvements up to 70%, in Cavg, over the baseline system. Ignacio López-Moreno, Javier Gonzalez-Dominguez, Oldrich Plchot, David Martinez, Joaquín González-Rodríguez, Pedro J. Moreno 0001 |
ICASSP | 6 |
| 2014 | Backoff inspired features for maximum entropy language modelsabstractMaximum Entropy (MaxEnt) language models [1, 2] are linear models that are typically regularized via well-known L1 or L2 terms in the likelihood objective, hence avoiding the need for the kinds of backoff or mixture weights used in smoothed n-gram language models using Katz backoff [3] and similar tech-niques. Even though backoff cost is not required to regularize the model, we investigate the use of backoff features in Max-Ent models, as well as some backoff-inspired variants. These features are shown to improve model quality substantially, as shown in perplexity and word-error rate reductions, even in very large scale training scenarios of tens or hundreds of billions of words and hundreds of millions of features. Index Terms: maximum entropy modeling, language model-ing, n-gram models, linear models Fadi Biadsy, Keith B. Hall, Pedro J. Moreno 0001, Brian Roark |
INTERSPEECH | 3 |
| 2014 | Automatic language identification using long short-term memory recurrent neural networksabstractThis work explores the use of Long Short-Term Memory (LSTM) recurrent neural networks (RNNs) for automatic lan-guage identification (LID). The use of RNNs is motivated by their better ability in modeling sequences with respect to feed forward networks used in previous works. We show that LSTM RNNs can effectively exploit temporal dependencies in acoustic data, learning relevant features for language discrimination pur-poses. The proposed approach is compared to baseline i-vector and feed forward Deep Neural Network (DNN) systems in the NIST Language Recognition Evaluation 2009 dataset. We show LSTM RNNs achieve better performance than our best DNN system with an order of magnitude fewer parameters. Further, the combination of the different systems leads to significant per-formance improvements (up to 28%). 1. Javier Gonzalez-Dominguez, Ignacio López-Moreno, Hasim Sak, Joaquín González-Rodríguez, Pedro J. Moreno 0001 |
INTERSPEECH | 5 |
| 2014 | A big data approach to acoustic model training corpus selectionabstractDeep neural networks (DNNs) have recently become the state of the art technology in speech recognition systems. In this pa-per we propose a new approach to constructing large high qual-ity unsupervised sets to train DNN models for large vocabulary speech recognition. The core of our technique consists of two steps. We first redecode speech logged by our production rec-ognizer with a very accurate (and hence too slow for real-time usage) set of speech models to improve the quality of ground truth transcripts used for training alignments. Using confidence scores, transcript length and transcript flattening heuristics de-signed to cull salient utterances from three decades of speech per language, we then carefully select training data sets consist-ing of up to 15K hours of speech to be used to train acoustic models without any reliance on manual transcription. We show that this approach yields models with approximately 18K con-text dependent states that achieve 10 % relative improvement in large vocabulary dictation and voice-search systems for Brazil-ian Portuguese, French, Italian and Russian languages. Index Terms: large unsupervised training sets, data selection, Olga Kapralova, John Alex, Eugene Weinstein, Pedro J. Moreno 0001, Olivier Siohan |
INTERSPEECH | 4 |
| 2014 | Asynchronous stochastic optimization for sequence training of deep neural networks: towards big dataabstractPrevious work presented a proof of concept for sequence training of deep neural networks (DNNs) using asynchronous stochastic optimization, mainly focusing on a small-scale task. The approach offers the potential to leverage both the efficiency of stochastic gradient descent and the scalability of parallel computation. This study presents results for four different voice search tasks to confirm the effectiveness and efficiency of the proposed framework across different conditions: amount of data (from 60 hours to 20,000 hours), type of speech (read speech vs. spontaneous speech), quality of data (supervised vs. unsupervised data), and language. Significant gains over baselines (DNNs trained at the frame level) are found to hold across these conditions. The experimental results are analyzed, and additional practical details for the approach are provided. Furthermore, different sequence training criteria are compared. Erik McDermott, Georg Heigold, Pedro J. Moreno 0001, Andrew W. Senior, Michiel Bacchiani |
INTERSPEECH | 3 |
| 2012 | Google's cross-dialect Arabic voice searchabstractWe present a large scale effort to build a commercial Automatic Speech Recognition (ASR) product for Arabic. Our goal is to support voice search, dictation, and voice control for the general Arabic-speaking public, including support for multiple Arabic dialects. We describe our ASR system design and compare recognizers for five Arabic dialects, with the potential to reach more than 125 million people in Egypt, Jordan, Lebanon, Saudi Arabia, and the United Arab Emirates (UAE). We compare systems built on diacritized vs. non-diacritized text. We also conduct cross-dialect experiments, where we train on one dialect and test on the others. Our average word error rate (WER) is 24.8% for voice search. Fadi Biadsy, Pedro J. Moreno 0001, Martin Jansche |
ICASSP | 2 |
| 2011 | Deploying Google Search by Voice in CantoneseabstractWe describe our efforts in deploying Google search by voice for Cantonese, a southern Chinese dialect widely spoken in and around Hong Kong and Guangzhou. We collected audio data from local Cantonese speakers in Hong Kong and Guangzhou by using our DataHound smartphone application. This data was used to create appropriate acoustic models. Language models were trained on anonymized query logs from Google Web Search for Hong Kong. Because users in Hong Kong frequently mix English and Cantonese in their queries, we designed our system from the ground up to handle both languages. We report on experiments with different techniques for mapping the phoneme inventories for both languages into a common space. Based on extensive experiments we report word error rates and web scores for both Hong Kong and Guangzhou data. Cantonese Google search by voice was launched in December 2010. Index Terms: voice search, Cantonese speech recognition, multilingual speech recognition Yun-Hsuan Sung, Martin Jansche, Pedro J. Moreno 0001 |
INTERSPEECH | 3 |
| 2010 | Voice search for developmentabstractIn light of the serious problems with both illiteracy and information access in the developing world, there is a widespread belief that speech technology can play a significant role in improving the quality of life of developing-world citizens. We review the main reasons why this impact has not occurred to date, and propose that voice-search systems may be a useful tool in delivering on the original promise. The challenges that must be addressed to realize this vision are analyzed, and initial experimental results in developing voice search for two languages of South Africa (Zulu and Afrikaans) are summarized. Index Terms: voice search, zulu, afrikaans Etienne Barnard, Johan Schalkwyk, Charl Johannes van Heerden, Pedro J. Moreno 0001 |
INTERSPEECH | 4 |
| 2010 | Building transcribed speech corpora quickly and cheaply for many languagesabstractWe present a system for quickly and cheaply building transcribed speech corpora containing utterances from many speakers in a variety of acoustic conditions. The system consists of a client application running on an Android mobile device with an intermittent Internet connection to a server. The client application collects demographic information about the speaker, fetches textual prompts from the server for the speaker to read, records the speaker’s voice, and uploads the audio and associated metadata to the server. The system has so far been used to collect over 3000 hours of transcribed audio in 17 languages around the world. Index Terms: speech corpora, speech recognition, internationalization Thad Hughes, Kaisuke Nakajima, Linne Ha, Atul Vasu, Pedro J. Moreno 0001, Mike LeBeau |
INTERSPEECH | 5 |
| 2010 | Search by voice in Mandarin ChineseabstractIn this paper we describe our efforts to build a Mandarin Chinese voice search system. We describe our strategies for data collection, language, lexicon and acoustic modeling, as well as issues related to text normalization that are an integral part of building voice search systems. We show excellent performance on typical spoken search queries under a variety of accents and acoustic conditions. The system has been in operation since October 2009 and has received very positive user reviews. Jiulong Shan, Genqing Wu, Zhihong Hu, Xiliu Tang, Martin Jansche, Pedro J. Moreno 0001 |
INTERSPEECH | 6 |
| 2010 | Efficient and Robust Music Identification With Weighted Finite-State TransducersabstractWe present an approach to music identification based on weighted finite-state transducers and Gaussian mixture models, inspired by techniques used in large-vocabulary speech recognition. Our modeling approach is based on learning a set of elementary music sounds in a fully unsupervised manner. While the space of possible music sound sequences is very large, our method enables the construction of a compact and efficient representation for the song collection using finite-state transducers. This paper gives a novel and substantially faster algorithm for the construction offactor transducers, the key representation of song snippets supporting our music identification technique. The complexity of our algorithm is linear with respect to the size of the suffix automaton constructed. Our experiments further show that it helps speed up the construction of the weighted suffix automaton in our task by a factor of 17 with respect to our previous method using the intermediate steps of determinization and minimization. We show that, using these techniques, a large-scale music identification system can be constructed for a database of over 15 000 songs while achieving an identification accuracy of 99.4% on undistorted test data, and performing robustly in the presence of noise and distortions. Mehryar Mohri, Pedro J. Moreno 0001, Eugene Weinstein |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | An audio indexing system for election video materialabstractIn the 2008 presidential election race in the United States, the prospective candidates made extensive use of YouTube to post video material. We developed a scalable system that transcribes this material and makes the content searchable (by indexing the meta-data and transcripts of the videos) and allows the user to navigate through the video material based on content. The system is available as an iGoogle gadget1as well as a Labs product (labs.google.com/gaudi). Given the large exposure, special emphasis was put on the scalability and reliability of the system. This paper describes the design and implementation of this system. Christopher Alberti, Michiel Bacchiani, Ari Bezman, Ciprian Chelba, Anastassia Drofa, Hank Liao, Pedro J. Moreno 0001, Ted Power, Arnaud Sahuguet, Maria Shugrina, Olivier Siohan |
ICASSP | 7 |
| 2009 | A factor automaton approach for the forced alignment of long speech recordingsabstractThis paper addresses the problem of aligning long speech recordings to their transcripts. Previous work has focused on using highly tuned language models trained on the transcripts to reduce the search space. In this paper we propose the use of a factor automaton, a well known method to represent all substrings from a string. This automaton encodes a highly constrained language model trained on the transcripts. We show competitive results with n-gram models in several testing scenarios. Preliminary experiments show perfect alignments at a reduced computational load and with a smaller memory footprint when compared to n-gram models. Pedro J. Moreno 0001, Christopher Alberti |
ICASSP | 1 |
| 2009 | Audiovisual celebrity recognition in unconstrained web videosabstractThe number of video clips available online is growing at a tremendous pace. Conventionally, user-supplied metadata text, such as the title of the video and a set of keywords, has been the only source of indexing information for user-uploaded videos. Automated extraction of video content for unconstrained and large scale video databases is a challenging and yet unsolved problem. In this paper, we present an audiovisual celebrity recognition system towards automatic tagging of unconstrained Web videos. Prior work on audiovisual person recognition relied on the fact that the person in the video is speaking and the features extracted from audio and visual domain are associated with each other throughout the video. However, this assumption is not valid on unconstrained Web videos. Proposed method finds the audiovisual mapping and hence improve upon the association assumption. Considering the scale of the application, all pieces of the system are trained automatically without any human supervision. We present the results on 26,000 videos and show the effectiveness of the method per-celebrity basis. Mehmet Emre Sargin, Hrishikesh B. Aradhye, Pedro J. Moreno 0001 |
ICASSP | 3 |
| 2009 | A new quality measure for topic segmentation of text and speechabstractThe recent proliferation of large multimedia collections has gathered immense attention from the speech research community, because speech recognition enables the transcription and indexing of such collections. Topicality information can be used to improve transcription quality and enable content navigation. In this paper, we give a novel quality measure for topic segmentation algorithms that improves over previously used measures. Our measure takes into account not only the presence or absence of topic boundaries but also the content of the text or speech segments labeled as topic-coherent. Additionally, we demonstrate that topic segmentation quality of spoken language can be improved using speech recognition lattices. Using lattices, improvements over the baseline one-best topic model are observed when measured with the previously existing topic segmentation quality measure, as well as the new measure proposed in this paper (9.4 % and 7.0 % relative error reduction, respectively). Index Terms: Topic segmentation, speech recognition lattices, text similarity, speech processing. Mehryar Mohri, Pedro J. Moreno 0001, Eugene Weinstein |
INTERSPEECH | 2 |
| 2009 | General suffix automaton construction algorithm and space bounds
Mehryar Mohri, Pedro J. Moreno 0001, Eugene Weinstein |
Theor. Comput. Sci. | 2 |
| 2007 | Music Identification with Weighted Finite-State TransducersabstractMusic identification is the process of matching an audio stream to a particular song. Previous work has relied on hashing, where an exact or almost-exact match between local features of the test and reference recordings is required. In this work we present a new approach to music identification based on finite-state transducers and Gaussian mixture models. We apply an unsupervised training process to learn an inventory of music phone units similar to phonemes in speech. We also learn a unique sequence of music units characterizing each song. We further propose a novel application of transducers for recognition of music phone sequences. Preliminary experiments demonstrate an identification accuracy of 99.5% on a database of over 15,000 songs running faster than real time. Eugene Weinstein, Pedro J. Moreno 0001 |
ICASSP (2) | 2 |
| 2007 | Factor Automata of Automata and Applications
Mehryar Mohri, Pedro J. Moreno 0001, Eugene Weinstein |
CIAA | 2 |
| 2007 | Supervised Learning of Semantic Classes for Image Annotation and RetrievalabstractA probabilistic formulation for semantic image annotation and retrieval is proposed. Annotation and retrieval are posed as classification problems where each class is defined as the group of database images labeled with a common semantic label. It is shown that, by establishing this one-to-one correspondence between semantic labels and semantic classes, a minimum probability of error annotation and retrieval are feasible with algorithms that are 1) conceptually simple, 2) computationally efficient, and 3) do not require prior semantic segmentation of training images. In particular, images are represented as bags of localized feature vectors, a mixture density estimated for each image, and the mixtures associated with all images annotated with a common semantic label pooled into a density estimate for the corresponding semantic class. This pooling is justified by a multiple instance learning argument and performed efficiently with a hierarchical extension of expectation-maximization. The benefits of the supervised formulation over the more complex, and currently popular, joint modeling of semantic label and visual feature distributions are illustrated through theoretical arguments and extensive experiments. The supervised formulation is shown to achieve higher accuracy than various previously published methods at a fraction of their computational cost. Finally, the proposed method is shown to be fairly robust to parameter tuning. Gustavo Carneiro 0001, Antoni B. Chan, Pedro J. Moreno 0001, Nuno Vasconcelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2007 | Bridging the Gap: Query by Semantic ExampleabstractA combination of query-by-visual-example (QBVE) and semantic retrieval (SR), denoted as query-by-semantic-example (QBSE), is proposed. Images are labeled with respect to a vocabulary of visual concepts, as is usual in SR. Each image is then represented by a vector, referred to as a semantic multinomial, of posterior concept probabilities. Retrieval is based on the query-by-example paradigm: the user provides a query image, for which 1) a semantic multinomial is computed and 2) matched to those in the database. QBSE is shown to have two main properties of interest, one mostly practical and the other philosophical. From a practical standpoint, because it inherits the generalization ability of SR inside the space of known visual concepts (referred to as the semantic space) but performs much better outside of it, QBSE produces retrieval systems that are more accurate than what was previously possible. Philosophically, because it allows a direct comparison of visual and semantic representations under a common query paradigm, QBSE enables the design of experiments that explicitly test the value of semantic representations for image retrieval. An implementation of QBSE under the minimum probability of error (MPE) retrieval framework, previously applied with success to both QBVE and SR, is proposed, and used to demonstrate the two properties. In particular, an extensive objective comparison of QBSE with QBVE is presented, showing that the former significantly outperforms the latter both inside and outside the semantic space. By carefully controlling the structure of the semantic space, it is also shown that this improvement can only be attributed to the semantic nature of the representation on which QBSE is based. Nikhil Rasiwasia, Pedro J. Moreno 0001, Nuno Vasconcelos |
IEEE Trans. Multim. | 2 |
| 2005 | Approaches to reduce the effects of OOV queries on indexed spoken audioabstractWe present several novel approaches to the Out of Vocabulary (OOV) query problem for spoken audio: indexing based on syllable-like units called particles and query expansion according to acoustic confusability for a word index. We also examine linear and OOV-based combination of indexing schemes. We experiment on 75 h of broadcast news, comparing our techniques to a word index, a phoneme index and a phoneme index queried with phoneme sequences. Our results show that our approaches are superior to both a word index and a phoneme index for OOV words, and have comparable performance to the sequence of phonemes scheme. The particle system has worse performance than the acoustic query expansion scheme. The best system uses word queries for in-vocabulary words and a linear combination of the phoneme sequence scheme and acoustic query expansion for OOV words. Using the best possible weights for linear combination, this system improves the average precision from 0.35 for a word index to 0.40, a result only obtainable if the weights could be learnt on a development query set. The next best system used a word index for in-vocabulary words and the phoneme sequence system otherwise and had average precision of 0.39. Beth Logan, Jean-Manuel Van Thong, Pedro J. Moreno 0001 |
IEEE Trans. Multim. | 3 |
| 2004 | The Kullback-Leibler Kernel as a Framework for Discriminant and Localized Representations for Visual Recognition
Nuno Vasconcelos, Purdy Ho, Pedro J. Moreno 0001 |
ECCV (3) | 3 |
| 2004 | Semantic analysis of song lyricsabstractWe explore the use of song lyrics for automatic indexing of music. Using lyrics mined from the Web, we apply a standard text processing technique to characterize their semantic content. We then determine artist similarity in this space. We found lyrics can be used to discover natural genre clusters. Experiments on a publicly available set of 399 artists showed that determining artist similarity using lyrics is better than random, but inferior to a state-of-the-art acoustic similarity technique. However the approaches made different errors, suggesting they could be profitably combined. Beth Logan, A. Kositsky, Pedro J. Moreno 0001 |
ICME | 3 |
| 2004 | News Tuner: a simple interface for searching and browsing radio archivesabstractWe present A new Web-based application, called the News Tuner, for searching and browsing large radio archives. While popular search engines provide means for finding text and images, our approach combines semantic and acoustic search for efficient retrieval of audio documents. Semantic search allows the user to retrieve stories for a given concept, while acoustic search allows random access within stored audio files. Our experiments on over 1700 programs show that our method is effective at quickly retrieving stories that would be difficult to find otherwise. The News Tuner paradigm is intended primarily for news and talk radio programs, however it may be applied to browsing and searching any spoken word audio content. J. Marston, G. MacCarthy, Beth Logan, Pedro J. Moreno 0001, Jean-Manuel Van Thong |
ICME | 4 |
| 2004 | SVM kernel adaptation in speaker classification and verification
Purdy Ho, Pedro J. Moreno 0001 |
INTERSPEECH | 2 |
| 2003 | A new SVM approach to speaker identification and verification using probabilistic distance kernelsabstractsupport vector machine, SVM, speaker identification, speaker verification, KL divergence, Kullback-Leibler divergence, probabilistic distance kernels, multimedia One major SVM weakness has been the use of generic kernel functions to compute distances among data points. Polynomial, linear, and Gaussian are typical examples. They do not take full advantage of the inherent probability distributions of the data. Focusing on audio speaker identification and verification, we propose to explore the use of novel kernel functions that take full advantage of good probabilistic and descriptive models of audio data. We explore the use of generative speaker identification models such as Gaussian Mixture Models and derive a kernel distance based on the Kullback-Leibler (KL) divergence between generative models. In effect our approach combines the best of both generative and discriminative methods. Our results show that these new kernels perform as well as baseline GMM classifiers and outperform generic kernel based SVM’s in both speaker identification and verification on two different audio databases. Pedro J. Moreno 0001, Purdy Ho |
INTERSPEECH | 1 |
| 2003 | A Kullback-Leibler Divergence Based Kernel for SVM Classification in Multimedia ApplicationsabstractOver the last years significant efforts have been made to develop kernels that can be applied to sequence data such as DNA, text, speech, video and images. The Fisher Kernel and similar variants have been suggested as good ways to combine an underlying generative model in the feature space and discriminant classifiers such as SVM’s. In this paper we sug- gest an alternative procedure to the Fisher kernel for systematically find- ing kernel functions that naturally handle variable length sequence data in multimedia domains. In particular for domains such as speech and images we explore the use of kernel functions that take full advantage of well known probabilistic models such as Gaussian Mixtures and sin- gle full covariance Gaussian models. We derive a kernel distance based on the Kullback-Leibler (KL) divergence between generative models. In effect our approach combines the best of both generative and discrim- inative methods and replaces the standard SVM kernels. We perform experiments on speaker identification/verification and image classifica- tion tasks and show that these new kernels have the best performance in speaker verification and mostly outperform the Fisher kernel based SVM’s and the generative classifiers in speaker identification and image classification. Pedro J. Moreno 0001, Purdy Ho, Nuno Vasconcelos |
NIPS | 1 |
| 2002 | Speechbot: an experimental speech-based search engine for multimedia content on the webabstractAs the Web transforms from a text-only medium into a more multimedia-rich medium, the need arises to perform searches based on the multimedia content. In this paper, we present an audio and video search engine to tackle this problem. The engine uses speech recognition technology to index spoken audio and video files from the World Wide Web (WWW) when no transcriptions are available. If transcriptions (even imperfect ones) are available, we can also take advantage of them to improve the indexing process. Our engine indexes several thousand talk and news radio shows covering a wide range of topics and speaking styles from a selection of public Web sites with multimedia archives. Our Web site is similar in spirit to normal Web search sites; it contains an index, not the actual multimedia content. The audio from these shows suffers in acoustic quality due to bandwidth limitations, coding, compression, and poor acoustic conditions. Our word error rate (WER) results using appropriately trained acoustic models show remarkable resilience to the high compression, although many factors combine to increase the average WERs over standard broadcast news benchmarks. We show that, even if the transcription is inaccurate, we can still achieve good retrieval performance for typical user queries (77.5%). Jean-Manuel Van Thong, Pedro J. Moreno 0001, Beth Logan, Blair Fidler, K. Maffey, M. Moores |
IEEE Trans. Multim. | 2 |
| 2001 | A boosting approach for confidence scoringabstractIn this paper we present the application of a boosting classification algorithm to confidence scoring. We derive feature vectors from speech recognition lattices and feed them into a boosting classifier. This classifier combines hundreds of very simple `weak learners' and derives classification rules that can reduce the confidence error rate by up to 34%. We compare our results to those obtained using two other standard classification techniques, Support Vector Machines (SVMs) and Classification and Regression Trees (CART), and show significant improvements. Furthermore, the nature of the boosting algorithm allows us to combine the best single classifier and improve its performance. We present experimental results on real world corpora derived from our SpeechBot Web index http://www.speechbot.com and from the HUB4 DARPA evaluation sets. We believe these results have wide applicability to audio indexing and to acoustic and language modeling adaptation where word confidence scores can be used in iterative adaptation schemes. 1. Pedro J. Moreno 0001, Beth Logan, Bhiksha Raj |
INTERSPEECH | 1 |
| 2001 | Topic Segmentation with an Aspect Hidden Markov ModelabstractWe present a novel probabilistic method for topic segmentation on unstructured text. One previous approach to this problem utilizes the hidden Markov model (HMM) method for probabilistically modeling sequence data [7]. The HMM treats a document as mutually independent sets of words generated by a latent topic variable in a time series. We extend this idea by embedding Hofmann's aspect model for text [5] into the segmenting HMM to form an aspect HMM (AHMM). In doing so, we provide an intuitive topical dependency between words and a cohesive segmentation model. We apply this method to segment unbroken streams of New York Times articles as well as noisy transcripts of radio programs on SpeechBot, an online audio archive indexed by an automatic speech recognition engine. We provide experimental comparisons which show that the AHMM outperforms the HMM for this task. David M. Blei, Pedro J. Moreno 0001 |
SIGIR | 2 |
| 2000 | Using the Fisher kernel method for Web audio classificationabstractAs the multimedia content of the Web increases techniques to automatically classify this content become more important. We present a system to classify audio files collected from the Web. The system classifies any audio file as belonging to one of three categories: speech, music and other. To classify the audio files, we use the technique of Fisher kernels. The technique as proposed by Jaakkola (1998) assumes a probabilistic generative model for the data, in our case a Gaussian mixture model. Then a discriminative classifier uses the GMM as an intermediate step to produce appropriate feature vectors. Support vector machines are our choice of discriminative classifier. We present classification results on a collection of more than 173 hours of Web audio randomly collected. We believe our results represent one of the first realistic studies of audio classification performance on found data. Our final system yielded a classification rate of 81.8%. Pedro J. Moreno 0001, Ryan Rifkin |
ICASSP | 1 |
| 2000 | An experimental study of an audio indexing system for the webabstractWe have developed a speech recognition based audio search engine for indexing spoken documents found on the World Wide Web. Our site (http://www.compaq.com/speechbot) indexes around 20 news and talk radio shows covering a wide range of topics, speaking styles and acoustic conditions from a selection of public Web sites with multimedia archives. In this paper, we describe our system and its performance, focusing on the speech recognition and retrieval aspects. We describe our training procedure in some detail and report our historical error rate since the site launch. We also investigate the impact of Out Of Vocabulary (OOV) words. Finally we report the results of retrieval experiments which demonstrate that our system can index effectively. Beth Logan, Pedro J. Moreno 0001, Jean-Manuel Van Thong, Edward W. D. Whittaker |
INTERSPEECH | 2 |
| 1999 | On the use of support vector machines for phonetic classificationabstractSupport vector machines (SVMs) represent a new approach to pattern classification which has attracted a great deal of interest in the machine learning community. Their appeal lies in their strong connection to the underlying statistical learning theory, in particular the theory of structural risk minimization. SVMs have been shown to be particularly successful in fields such as image identification and face recognition; in many problems SVM classifiers have been shown to perform much better than other nonlinear classifiers such as artificial neural networks and k-nearest neighbors. This paper explores the issues involved in applying SVMs to phonetic classification as a first step to speech recognition. We present results on several standard vowel and phonetic classification tasks and show better performance than Gaussian mixture classifiers. We also present an analysis of the difficulties we foresee in applying SVMs to continuous speech recognition problems. Philip Clarkson, Pedro J. Moreno 0001 |
ICASSP | 2 |
| 1998 | Factorial HMMs for acoustic modelingabstractIn the machine learning research field several extensions of hidden Markov models (HMMs) have been proposed. In this paper we study their possibilities and potential benefits for the field of acoustic modeling. We describe preliminary experiments using an alternative modeling approach known as factorial hidden Markov models (FHMMs). We present these models as extensions of HMMs and detail a modification to the original formulation which seems to allow a more natural fit to speech. We present experimental results on the phonetically balanced TIMIT database comparing the performance of FHMMs with HMMs. We also study alternative feature representations that might be more suited to FHMMs. Beth Logan, Pedro J. Moreno 0001 |
ICASSP | 2 |
| 1998 | A recursive algorithm for the forced alignment of very long audio segmentsabstractIn this paper we address the problem of aligning very long (of-ten more than one hour) audio files to their corresponding textual transcripts in an effective manner. We present an efficient recur-sive technique to solve this problem that works well even on noisy speech signals. The key idea of this algorithm is to turn the forced alignment problem into a recursive speech recognition problem with a gradually restricting dictionary and language model. The algorithm is tolerant to acoustic noise and errors or gaps in the text transcript or audio tracks. We report experimental results on a 3 hour audio file containing TV and radio broadcasts. We will show accurate alignments on speech under a variety of real acoustic conditions such as speech over music and speech over telephone lines. We also report re-sults when the same audio stream has been corrupted with white additive noise or compressed using a popular web encoding for-mat such as RealAudio. This algorithm has been used in our internal multimedia indexing project. It has processed more than 200 hours of audio from var-ied sources, such as WGBH NOVA documentaries and NPR web audio files. The system aligns speech media content in about one to five times realtime, depending on the acoustic conditions of the audio signal. 1. Pedro J. Moreno 0001, Christopher F. Joerg, Jean-Manuel Van Thong, Oren Glickman |
ICSLP | 1 |
| 1998 | Data-driven environmental compensation for speech recognition: A unified approach
Pedro J. Moreno 0001, Bhiksha Raj, Richard M. Stern |
Speech Commun. | 1 |
| 1997 | Delta vector taylor series environment compensation for speaker recognition
Brian S. Eberman, Pedro J. Moreno 0001 |
EUROSPEECH | 2 |
| 1997 | A new algorithm for robust speech recognition: the delta vector taylor series approachabstractA New Algorithm for Robust Sp eech Recognition: The Delta VectorTaylor Series ApproachPedroJ.Moreno and Brian Ebermanemail: [email protected], [email protected] Equipment Corp orationCambridge Research Lab oratoryABSTRACTIn this pap er we present a new mo del-based comp ensationtechnique called Delta Vector Taylor Series (DVTS). Thisnew technique is an extension and improvementoer theVector Taylor Series (VTS) approach [7] that addressesseveral of its limitations .In particular, we presentanew statistical representation for the distribution of cleansp eech feature vectors based on a weighted vector co de-b o ok. This change to the underlying probabili ty densityfunction (PDF) allows us to pro duce more accurate andstable solutions for our algorithm. The algorithm is alsopresented in a EM-MAP framework where some the en-vironmental parameters are treated as random variableswith known PDF's. Finally,we explore a new comp ensa-tion approach based on the use of convex hulls.Weevaluate our algorithm in a phonetic classi cati on taskon the TIMIT [5] database and also in a small vo cabu-lary size sp eech recognition database. In b oth databasesarti cial and natural noise is injected at several signal tonoise ratios (SNR). The algorithm achieves matched p er-formance at all SNR's ab ove 10 dB.1.Intro ductionOver the last years several techniques have b een prop osedto deal with the problem of sp eech recognition in noisy en-vironments. Some of them such as PMC [3], or MLLR [6]have used the recognition engine and its rich statisticalrepresentation (more than 90,000 Gaussians in systemslike SPHINX-3 and HTK [9]) to mo del and comp ensatefor the e ects of the environment on sp eech recognitionsystems. Other techniques like CDCN [1] and POF [8]among others have used a reduced set of Gaussian mix-tures (typically 256 or less) to mo del the sp eechfeature vectors and prepro cess the noisy sp eech featuresvectors to e ectively clean the features b efore b eing pro-cessed by the recognition engine.The use of a rich statistical representation improves p er-formance, but has the drawback of using the whole sp eechrecognition engine with its asso ciated complexity.Anideal robust recognition technique should have the advan-tages of a rich statistical representation and at the sametime b eing simple and fast in its op eration.The Delta Vector Taylor Series (DVTS) approachis anattempt in this direction.It tries to gain the b ene tsof a rich statistical representation and a low complexitytechnique for robust sp eech recognition. It tries to achievethese goals by using a di erent statistical representationfor the sp eech feature vectors.The outline of the pap er is as follows. In section 2 wedescrib e the DVTS algorithm.In section 3 we brieydescrib e the necessary mo di cations to the algorithm tomakeit work as a lter.In section 4 we describ e ourexp erimental results and nally in section 5 we presentour conclusions.2.New Algorithm: Delta-VTSDVTS mo dels the sp eech feature vectors as a weightedsum of multidimensi onal Dirac deltasp(x)=M1Xk=0P[k])(1)where eachvector function(xk) is mo deled as(xk)=D1Yi=0ii;k)(2)P[k]is ana prioriprobability of observing a particulardelta. The sum of these probabili ties must add up to one.This novel representation of the PDF ofxhas several ad-vantages. First of all it greatly simpli es the mathemat-ical assumptions of the VTS [7] algorithm. It pro ducesa simple, fast, robust and direct formulation of the EMsolutions already presented in [7].In this pap er we assume a mo del of the environmentinwhich sp eech is corrupted by unknown additive stationarynoise and unknown linear lteringZ(!)=X)jH2+N(3)whereZ(!) represents the p ower sp ectrum of the de-graded sp eech,X(!) is the p ower sp ectrum of the cleansp eech,jH(!)2is the transfer function of the linear lter,andN(!) is the p ower sp ectrum of the additive noise.In the log-mel-sp ectral domain this can b e expressed asz=x+ log (exp (q) + exp (n))(4)or in more general termsz=x+f(;nq)(5) Pedro J. Moreno 0001, Brian S. Eberman |
EUROSPEECH | 1 |
| 1996 | A vector Taylor series approach for environment-independent speech recognitionabstractIn this paper we introduce a new analytical approach to environment compensation for speech recognition. Previous attempts at solving analytically the problem of noisy speech recognition have either used an overly-simplified mathematical description of the effects of noise on the statistics of speech or they have relied on the availability of large environment-specific adaptation sets. Some of the previous methods required the use of adaptation data that consists of simultaneously-recorded or "stereo" recordings of clean and degraded speech. In this work we introduce the use of a vector Taylor series (VTS) expansion to characterize efficiently and accurately the effects on speech statistics of unknown additive noise and unknown linear filtering in a transmission channel. The VTS approach is computationally efficient. It can be applied either to the incoming speech feature vectors, or to the statistics representing these vectors. In the first case the speech is compensated and then recognized; in the second case HMM statistics are modified using the VTS formulation. Both approaches use only the actual speech segment being recognized to compute the parameters required for environmental compensation. We evaluate the performance of two implementations of VTS algorithms using the CMU SPHINX-II system on the 100-word alphanumeric CENSUS database and on the 1993 5000-word ARPA Wall Street Journal database. Artificial white Gaussian noise is added to both databases. The VTS approaches provide significant improvements in recognition accuracy compared to previous algorithms. Pedro J. Moreno 0001, Bhiksha Raj, Richard M. Stern |
ICASSP | 1 |
| 1996 | Cepstral compensation by polynomial approximation for environment-independent speech recognition
Bhiksha Raj, Evandro B. Gouvêa, Pedro J. Moreno 0001, Richard M. Stern |
ICSLP | 3 |
| 1995 | Multivariate-Gaussian-based cepstral normalization for robust speech recognitionabstractWe introduce a new family of environmental compensation algorithms called multivariate gaussian based cepstral normalization (RATZ). RATZ assumes that the effects of unknown noise and filtering on speech features can be compensated by corrections to the mean and variance of components of Gaussian mixtures, and an efficient procedure for estimating the correction factors is provided. The RATZ algorithm can be implemented to work with or without the use of "stereo" development data that had been simultaneously recorded in the training and testing environments. "Blind" RATZ partially overcomes the loss of information that would have been provided by stereo training through the use of a more accurate description of how noisy environments affect clean speech. We evaluate the performance of the two RATZ algorithms using the CMU SPHINX-II system on the alphanumeric census database and compare their performance with that of previous environmental-robustness developed at CMU. Pedro J. Moreno 0001, Bhiksha Raj, Evandro B. Gouvêa, Richard M. Stern |
ICASSP | 1 |
| 1995 | A unified approach for robust speech recognitionabstractThere are two major structural approaches to robust speech recognition.In the first approach to the problem, compensation is performed by modifying the incoming cepstral stream using ML or MMSE methods to estimate parameters characterizing environmental degradation, from direct frame-by-frame comparisons between speech recorded in high-quality and degraded acoustical environments, or by signal processing techniques such as spectral subtraction.The second approach tackles the problem by modifying the statistics of the internal representation of speech cepstra in the classifier to make them more closely resemble the statistics of degraded speech.This paper attempts to unify these approaches to robust speech recognition by presenting three techniques that share the same basic assumptions and internal structure but differ in whether they modify the incoming speech cepstra or whether they modify the classifier statistics.We present SNR-dependent multi-vaRiate gAussian-based cepsTral normaliZation (SNR-RATZ) and SNR-based Blind RATZ (SNR-BRATZ), which modify incoming cepstra, along with STAR (STAtistical Re-estimation), which modifies the internal statistics of the classifier.The algorithms were tested using the SPHINX-II speech recognition system on the CENSUS database, a database of strings of letters and numbers to which unknown added and unknown linear filtering was introduced artificially.While all the algorithms showed good performance, STAR was observed to provide lower error rates as SNR decreases than any of the algorithms that modify incoming cepstra. Pedro J. Moreno 0001, Bhiksha Raj, Richard M. Stern |
EUROSPEECH | 1 |
| 1994 | Environment normalization for robust speech recognition using direct cepstral comparisonabstractIn this paper we describe and evaluate a series of new algorithms that compensate for the effects of unknown acoustical environments or changes in environment. The algorithms use compensation vectors that are added to the cepstral representations of speech that is input to a speech recognition system. While these vectors are computed from direct frame-by-frame comparisons of cepstra of speech simultaneously recorded in the training environment and various prototype testing environments, the compensation algorithms do not assume that the acoustical characteristics of the actual testing environment are known. The specific compensation vector applied in a given frame depends on either physical attributes such as SNR or presumed phonetic identity. The compensation algorithms are evaluated using the 1992 ARPA 5000 word WSJ/CSR corpus. The best system combines phoneme-based and SNR-based cepstral compensation with cepstral mean normalization, and provides a 66.8% reduction in error rate over baseline processing when tested using a standard suite of unknown microphones.> Fu-Hua Liu, Richard M. Stern, Alex Acero, Pedro J. Moreno 0001 |
ICASSP (2) | 4 |
| 1994 | Sources of degradation of speech recognition in the telephone networkabstractWe compare speech recognition accuracy for high-quality speech recorded under controlled conditions with speech as it appears over long-distance telephone lines. In addition to comparing recognition accuracy we use telephone-channel simulation to identify the sources of degradation of speech over telephone lines that have the greatest impact on speech recognition accuracy. We first compare the performance of the CMU SPHINX-I system on the TIMIT and NTIMIT databases. We found that other factors beyond a mere decrease in bandwidth cause the observed degradation in recognition accuracy, and that the environmental compensation algorithms RASTA and CDCN fail to compensate completely for degradations introduced by the telephone network. We identify the most problematic telephone-channel impairments using a commercial telephone channel simulator and the SPHINX-II system. Of the various effects considered, additive noise and linear filtering appear to have the greatest impact on recognition accuracy. Finally, we examined the performance of three cepstral compensation algorithms in the presence of the most damaging conditions. We found the compensation algorithms to be effective except for the worst 1% of the telephone channels.> Pedro J. Moreno 0001, Richard M. Stern |
ICASSP (1) | 1 |
| 1994 | Signal processing for robust speech recognitionabstractThis paper describes several new cepstral-based compensation procedures that render the SPHINX-II system more robust with respect to acoustical environment.The first algorithm, phonedependent cepstral compensation, is similar in concept to the previously-described MFCDCN method, except that cepstral compensation vectors are selected according to the current phonetic hypothesis, rather than on the basis of SNR or VQ codeword identity.We also describe two procedures to accomplish adaptation of the VQ codebook for new environments.Use of the various compensation algorithms in consort produces a reduction of error rates for SPHINX-II by as much as 40 percent relative to the rate achieved with cepstral mean normalization alone. Richard M. Stern, Fu-Hua Liu, Pedro J. Moreno 0001, Alex Acero |
ICSLP | 3 |
| 1992 | Efficient grammar processing for a spoken language translation systemabstractA problem with many speech understanding systems is that grammars that are more suitable for representing the relation between sentences and their meanings, such as context free grammars (CFGs) and augmented phrase structure grammars (APSGs), are computationally very demanding. On the other hand, finite state grammars are efficient, but cannot represent directly the sentence-meaning relation. The authors describe how speech recognition and language analysis can be tightly coupled by developing an APSG for the analysis component and deriving automatically from it a finite-state approximation that is used as the recognition language model. Using this technique, the authors have built an efficient translation system that is fast compared to others with comparably sized language models.> David B. Roe, Fernando Pereira 0003, Richard Sproat, Michael Riley 0001, Pedro J. Moreno 0001, Alejandro Macarrón Larumbe |
ICASSP | 5 |
| 1992 | A spoken language translator for restricted-domain context-free languages
David B. Roe, Pedro J. Moreno 0001, Richard Sproat, Fernando Pereira 0003, Michael Riley 0001, Alejandro Macarrón Larumbe |
Speech Commun. | 2 |
| 1991 | Toward a spoken language translator for restricted-domain context-free languages
David B. Roe, Fernando Pereira 0003, Richard Sproat, Michael Riley 0001, Pedro J. Moreno 0001, Alejandro Macarrón Larumbe |
EUROSPEECH | 5 |