EDBT 2026 Demo / reviewers in the wild / expert
Françoise Beaufays
dblp:18/686
· DBLP profile ↗
62ranked-venue papers
10as first author
20since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 55 · 7 first-author · 19 since 2021Artificial intelligence and machine learning · 32 · 6 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Speech Re-Painting for Robust ASRabstractSynthetic speech is a useful source for augmentation of automatic speech recognition (ASR) systems, but there is a "sim-to-real" gap between synthetic and real speech that can limit generalization. The natural variability of real speech is essential to the training of robust ASR systems. While synthetic data augmentation can be used to approximate the variability of natural speech, however, not all aspects of variation are equally relevant for augmentation. In this work, we introduce speech re-painting, a method for in-context augmented synthesis, using target training datasets to generate new utterances guided by speech and text on the fly in a zero-shot manner. We evaluate this technique using downstream ASR word error rate (WER) using the VCTK and LibriSpeech datasets. These represent unique speaker and lexical challenges that are addressed by re-painting, realizing a reduction of WER more than 50% in particular settings. Kyle Kastner, Gary Wang, Isaac Elias, Takaaki Saeki, Pedro J. Moreno 0001, Françoise Beaufays, Andrew Rosenberg, Bhuvana Ramabhadran |
ICASSP | 6 |
| 2024 | Improving Speech Recognition for African American English with Audio ClassificationabstractAutomatic speech recognition (ASR) systems have been shown to have large quality disparities between the language varieties they are intended or expected to recognize. One way to mitigate this is to train or fine-tune models with more representative datasets. But this approach can be hindered by limited in-domain data for training and evaluation. We propose a new way to improve the robustness of a US English short-form speech recognizer using a small amount of out-of-domain (long-form) African American English (AAE) data. We use CORAAL, YouTube and Mozilla Common Voice to train an audio classifier to approximately output whether an utterance is AAE or some other variety including Mainstream American English (MAE). By combining the classifier output with coarse geographic information, we can select a subset of utterances from a large corpus of untranscribed short-form queries for semi-supervised learning at scale. Fine-tuning on this data results in a 38.5% relative word error rate disparity reduction between AAE and MAE without reducing MAE quality. Shefali Garg, Zhouyuan Huo, Khe Chai Sim, Suzan Schwartz, Mason Chua, Alëna Aksënova, Tsendsuren Munkhdalai, Levi King, Darryl Wright, Zion Mengesha, Dongseong Hwang, Tara N. Sainath, Françoise Beaufays, Pedro J. Moreno 0001 |
ICASSP | 13 |
| 2024 | Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed DataabstractCollecting high-quality studio recordings of audio is challenging, which limits the language coverage of text-to-speech (TTS) systems. This paper proposes a framework for scaling a multilingual TTS model to 100+ languages using found data without supervision. The proposed framework combines speech-text encoder pretraining with unsupervised training using untranscribed speech and unspoken text data sources, thereby leveraging massively multilingual joint speech and text representation learning. Without any transcribed speech in a new language, this TTS model can generate intelligible speech in >30 unseen languages (CER difference of <10% to ground truth). With just 15 minutes of transcribed, found data, we can reduce the intelligibility difference to 1% or less from the ground-truth, and achieve naturalness scores that match the ground-truth in several languages. Takaaki Saeki, Gary Wang, Nobuyuki Morioka, Isaac Elias, Kyle Kastner, Andrew Rosenberg, Bhuvana Ramabhadran, Heiga Zen, Françoise Beaufays, Hadar Shemtov |
ICASSP | 9 |
| 2023 | Lego-Features: Exporting Modular Encoder Features for Streaming and Deliberation ASRabstractIn end-to-end (E2E) speech recognition models, a representational tight-coupling inevitably emerges between the encoder and the decoder. We build upon recent work that has begun to explore building encoders with modular encoded representations, such that encoders and decoders from different models can be stitched together in a zero-shot manner without further fine-tuning. While previous research only addresses full-context speech models, we explore the problem in a streaming setting as well. Our framework builds on top of existing encoded representations, converting them to modular features, dubbed as Lego-Features, without modifying the pre-trained model. The features remain interchangeable when the model is retrained with distinct initializations. Though sparse, we show that the Lego-Features are powerful when tested with RNN-T or LAS decoders, maintaining high-quality downstream performance. They are also rich enough to represent the first-pass prediction during two-pass deliberation. In this scenario, they outperform the N-best hypotheses, since they do not need to be supplemented with acoustic features to deliver the best results. Moreover, generating the Lego-Features does not require beam search or auto-regressive computation. Overall, they present a modular, powerful and cheap alternative to the standard encoder output, as well as the N-best hypotheses. Rami Botros, Rohit Prabhavalkar, Johan Schalkwyk, Ciprian Chelba, Tara N. Sainath, Françoise Beaufays |
ICASSP | 6 |
| 2023 | Efficient Domain Adaptation for Speech Foundation ModelsabstractFoundation models (FMs), that are trained on broad data at scale and are adaptable to a wide range of downstream tasks, have brought large interest in the research community. Benefiting from the diverse data sources such as different modalities, languages and application domains, foundation models have demonstrated strong generalization and knowledge transfer capabilities. In this paper, we present a pioneering study towards building an efficient solution for FM-based speech recognition systems. We adopt the recently developed self-supervised BEST-RQ for pretraining, and extend the joint training strategy JUST Hydra for finetuning using both source and unsuper-vised target domain data. The FM encoder adapter and decoder are then finetuned to the target domain with a small amount of super-vised in-domain data. On a large-scale YouTube and Voice Search task, our method is shown to be both data and model parameter efficient. It achieves the same quality with only 21.6M supervised in-domain data and 130.8M finetuned parameters, compared to the 731.1M model trained from scratch on additional 300M supervised in-domain data. Bo Li 0028, Dongseong Hwang, Zhouyuan Huo, Junwen Bai, Guru Prakash Arumugam, Tara N. Sainath, Khe Chai Sim, Yu Zhang 0033, Wei Han 0002, Trevor Strohman, Françoise Beaufays |
ICASSP | 11 |
| 2023 | Online Model Compression for Federated Learning with Large ModelsabstractThis paper addresses the challenges of training large neural networks under federated learning settings: high on-device memory usage and communication cost. The proposed Online Model Compression (OMC) provides a framework that stores model parameters in a compressed format and decompresses them only when needed. We use quantization as the compression method in this paper and propose three methods, (1) per-variable transformation, (2) weight-matrix-only quantization, and (3) partial variable quantization, to minimize its impact on model accuracy. Our experiments on two recent neural networks for speech recognition and two different datasets show that OMC can reduce memory usage and communication cost of model parameters by up to 59% while attaining comparable accuracy and training speed when compared with full-precision federated learning. Tien-Ju Yang, Yonghui Xiao, Giovanni Motta, Françoise Beaufays, Rajiv Mathews, Mingqing Chen |
ICASSP | 4 |
| 2023 | Mixture-of-Expert Conformer for Streaming Multilingual ASR
Bo Li 0028, Tara N. Sainath, Yu Zhang 0033, Françoise Beaufays |
INTERSPEECH | 5 |
| 2022 | A Method to Reveal Speaker Identity in Distributed ASR Training, and How to Counter ITabstractEnd-to-end Automatic Speech Recognition (ASR) models are commonly trained over spoken utterances using optimization methods like Stochastic Gradient Descent (SGD). In distributed settings like Federated Learning, model training requires transmission of gradients over a network. In this work, we design the first method for revealing the identity of the speaker of a training utterance with access only to a gradient. We propose Hessian-Free Gradients Matching, an input reconstruction technique that operates without second derivatives of the loss function (required in prior works), which can be expensive to compute. We show the effectiveness of our method using the DeepSpeech model architecture, demonstrating that it is possible to reveal the speaker’s identity with 34% top-1 accuracy (51% top-5 accuracy) on the LibriSpeech dataset. Further, we study the effect of Dropout on the success of our method. We show that a dropout rate of 0.2 can reduce the speaker identity accuracy to 0% top-1 (0.5% top-5). Trung Dang 0002, Om Thakkar 0001, Swaroop Ramaswamy, Rajiv Mathews, Sang (Peter) Chin, Françoise Beaufays |
ICASSP | 6 |
| 2022 | Enabling On-Device Training of Speech Recognition Models With Federated DropoutabstractFederated learning can be used to train machine learning models on the edge on local data that never leave devices, providing privacy by default. This presents a challenge pertaining to the communication and computation costs associated with clients’ devices. These costs are strongly correlated with the size of the model being trained, and are significant for state-of-the-art automatic speech recognition models.We propose using federated dropout to reduce the size of client models while training a full-size model server-side. We provide empirical evidence of the effectiveness of federated dropout, and propose a novel approach to vary the dropout rate applied at each layer. Furthermore, we find that federated dropout enables a set of smaller sub-models within the larger model to independently have low word error rates, making it easier to dynamically adjust the size of the model deployed for inference. Dhruv Guliani, Lillian Zhou, Changwan Ryu, Tien-Ju Yang, Harry Zhang, Yonghui Xiao, Françoise Beaufays, Giovanni Motta |
ICASSP | 7 |
| 2022 | Large-Scale ASR Domain Adaptation Using Self- and Semi-Supervised LearningabstractSelf- and semi-supervised learning methods have been actively investigated to reduce labeled training data or enhance model performance. However, these approaches mostly focus on in-domain performance for public datasets. In this study, we utilize the combination of self- and semi-supervised learning methods to solve unseen domain adaptation problems in a large-scale production setting for online ASR model. This approach demonstrates that using the source domain data with a small fraction of the target domain data (3%) can recover the performance gap compared to a full data baseline: 13.5% relative WER improvement for target domain data. Dongseong Hwang, Ananya Misra, Zhouyuan Huo, Nikhil Siddhartha, Shefali Garg, David Qiu, Khe Chai Sim, Trevor Strohman, Françoise Beaufays, Yanzhang He |
ICASSP | 9 |
| 2022 | Fast Contextual Adaptation with Neural Associative Memory for On-Device Personalized Speech RecognitionabstractFast contextual adaptation has shown to be effective in improving Automatic Speech Recognition (ASR) of rare words and when combined with an on-device personalized training, it can yield an even better recognition result. However, the traditional re-scoring approaches based on an external language model is prone to diverge during the personalized training. In this work, we introduce a model-based end-to-end contextual adaptation approach that is decoder-agnostic and amenable to on-device personalization. Our on-device simulation experiments demonstrate that the proposed approach outperforms the traditional re-scoring technique by 12% relative WER and 15.7% entity mention specific F1-score in a continuous personalization scenario. Tsendsuren Munkhdalai, Khe Chai Sim, Angad Chandorkar, Mason Chua, Trevor Strohman, Françoise Beaufays |
ICASSP | 7 |
| 2022 | Partial Variable Training for Efficient on-Device Federated LearningabstractThis paper aims to address the major challenges of Federated Learning (FL) on edge devices: limited memory and expensive communication. We propose a novel method, called Partial Variable Training (PVT), that only trains a small subset of variables on edge devices to reduce memory usage and communication cost. With PVT, we show that network accuracy can be maintained by utilizing more local training steps and devices, which is favorable for FL involving a large population of devices. According to our experiments on two state-of-the-art neural networks for speech recognition and two different datasets, PVT can reduce memory usage by up to 1.9× and communication cost by up to 593× while attaining comparable accuracy when compared with full network training. Tien-Ju Yang, Dhruv Guliani, Françoise Beaufays, Giovanni Motta |
ICASSP | 3 |
| 2022 | Exploring Heterogeneous Characteristics of Layers in ASR Models for More Efficient TrainingabstractTransformer-based architectures have been the subject of research aimed at understanding their overparameterization and the non-uniform importance of their layers. Applying these approaches to Automatic Speech Recognition, we demonstrate that the state-of-the-art Conformer models generally have multiple ambient layers. We study the stability of these layers across runs and model sizes, propose that group normalization may be used without disrupting their formation, and examine their correlation with model weight updates in each layer. Finally, we apply these findings to Federated Learning in order to improve the training procedure, by targeting Federated Dropout to layers by importance. This allows us to reduce the model size optimized by clients without quality degradation, and shows potential for future exploration. Lillian Zhou, Dhruv Guliani, Andreas Kabel, Giovanni Motta, Françoise Beaufays |
ICASSP | 5 |
| 2022 | Extracting Targeted Training Data from ASR Models, and How to Mitigate ItabstractRecent work has designed methods to demonstrate that model updates in ASR training can leak potentially sensitive attributes of the utterances used in computing the updates.In this work, we design the first method to demonstrate information leakage about training data from trained ASR models.We design Noise Masking, a fill-in-the-blank style method for extracting targeted parts of training data from trained ASR models.We demonstrate the success of Noise Masking by using it in four settings for extracting names from the LibriSpeech dataset used for training a state-of-the-art Conformer model.In particular, we show that we are able to extract the correct names from masked training utterances with 11.8% accuracy, while the model outputs some name from the train set 55.2% of the time.Further, we show that even in a setting that uses synthetic audio and partial transcripts from the test set, our method achieves 2.5% correct name accuracy (47.7% any name success rate).Lastly, we design Word Dropout, a data augmentation method that we show when used in training along with Multistyle TRaining (MTR), provides comparable utility as the baseline, along with significantly mitigating extraction via Noise Masking across the four evaluated settings. Ehsan Amid, Om Thakkar 0001, Arun Narayanan, Rajiv Mathews, Françoise Beaufays |
INTERSPEECH | 5 |
| 2022 | Production federated keyword spotting via distillation, filtering, and joint federated-centralized training
Andrew Hard, Kurt Partridge, Neng Chen, Sean Augenstein, Aishanee Shah, Hyun Jin Park, Alex Park 0001, Sara Ng, Jessica Nguyen, Ignacio López-Moreno, Rajiv Mathews, Françoise Beaufays |
INTERSPEECH | 12 |
| 2022 | Incremental Layer-Wise Self-Supervised Learning for Efficient Unsupervised Speech Domain Adaptation On Device
Zhouyuan Huo, Dongseong Hwang, Khe Chai Sim, Shefali Garg, Ananya Misra, Nikhil Siddhartha, Trevor Strohman, Françoise Beaufays |
INTERSPEECH | 8 |
| 2022 | Federated Pruning: Improving Neural Network Efficiency with Federated LearningabstractAutomatic Speech Recognition models require large amount of speech data for training, and the collection of such data often leads to privacy concerns.Federated learning has been widely used and is considered to be an effective decentralized technique by collaboratively learning a shared prediction model while keeping the data local on different clients devices.However, the limited computation and communication resources on clients devices present practical difficulties for large models.To overcome such challenges, we propose Federated Pruning to train a reduced model under the federated setting, while maintaining similar performance compared to the full model.Moreover, the vast amount of clients data can also be leveraged to improve the pruning results compared to centralized training.We explore different pruning schemes and provide empirical evidence of the effectiveness of our methods. Rongmei Lin, Yonghui Xiao, Tien-Ju Yang, Ding Zhao, Li Xiong 0001, Giovanni Motta, Françoise Beaufays |
INTERSPEECH | 7 |
| 2021 | Training Speech Recognition Models with Federated Learning: A Quality/Cost FrameworkabstractWe propose using federated learning, a decentralized on-device learning paradigm, to train speech recognition models. By performing epochs of training on a per-user basis, federated learning must incur the cost of dealing with non-IID data distributions, which are expected to negatively affect the quality of the trained model. We propose a framework by which the degree of non-IID-ness can be varied, consequently illustrating a trade-off between model quality and the computational cost of federated training, which we capture through a novel metric. Finally, we demonstrate that hyper-parameter optimization and appropriate use of variational noise are sufficient to compensate for the quality impact of non-IID distributions, while decreasing the cost. Dhruv Guliani, Françoise Beaufays, Giovanni Motta |
ICASSP | 2 |
| 2021 | Robust Continuous On-Device Personalization for Automatic Speech Recognition
Khe Chai Sim, Angad Chandorkar, Mason Chua, Tsendsuren Munkhdalai, Françoise Beaufays |
Interspeech | 6 |
| 2021 | Revealing and Protecting Labels in Distributed TrainingabstractDistributed learning paradigms such as federated learning often involve transmission of model updates, or gradients, over a network, thereby avoiding transmission of private data. However, it is possible for sensitive information about the training data to be revealed from such gradients. Prior works have demonstrated that labels can be revealed analytically from the last layer of certain models (e.g., ResNet), or they can be reconstructed jointly with model inputs by using Gradients Matching [Zhu et al.] with additional knowledge about the current state of the model. In this work, we propose a method to discover the set of labels of training samples from only the gradient of the last layer and the id to label mapping. Our method is applicable to a wide variety of model architectures across multiple domains. We demonstrate the effectiveness of our method for model training in two domains - image classification, and automatic speech recognition. Furthermore, we show that existing reconstruction techniques improve their efficacy when used in conjunction with our method. Conversely, we demonstrate that gradient quantization and sparsification can significantly reduce the success of the attack. Trung Dang 0002, Om Thakkar 0001, Swaroop Ramaswamy, Rajiv Mathews, Sang (Peter) Chin, Françoise Beaufays |
NeurIPS | 6 |
| 2020 | Low-Rank Gradient Approximation for Memory-Efficient on-Device Training of Deep Neural NetworkabstractTraining machine learning models on mobile devices has the potential of improving both privacy and accuracy of the models. However, one of the major obstacles to achieving this goal is the memory limitation of mobile devices. Reducing training memory enables models with high-dimensional weight matrices, like automatic speech recognition (ASR) models, to be trained on-device. In this paper, we propose approximating the gradient matrices of deep neural networks using a low-rank parameterization as an avenue to save training memory. The low-rank gradient approximation enables more advanced, memory-intensive optimization techniques to be run on device. Our experimental results show that we can reduce the training memory by about 33.0% for Adam optimization. It uses comparable memory to momentum optimization and achieves a 4.5% relative lower word error rate on an ASR personalization task. Mary Gooneratne, Khe Chai Sim, Petr Zadrazil, Andreas Kabel, Françoise Beaufays, Giovanni Motta |
ICASSP | 5 |
| 2020 | Analyzing the Quality and Stability of a Streaming End-to-End On-Device Speech RecognizerabstractThe demand for fast and accurate incremental speech recognition increases as the applications of automatic speech recognition (ASR) proliferate.Incremental speech recognizers output chunks of partially recognized words while the user is still talking.Partial results can be revised before the ASR finalizes its hypothesis, causing instability issues.We analyze the quality and stability of on-device streaming end-to-end (E2E) ASR models.We first introduce a novel set of metrics that quantify the instability at word and segment levels.We study the impact of several model training techniques that improve E2E model qualities but degrade model stability.We categorize the causes of instability and explore various solutions to mitigate them in a streaming E2E ASR system. Yuan Shangguan, Kate Knister, Yanzhang He, Ian McGraw, Françoise Beaufays |
INTERSPEECH | 5 |
| 2019 | Personalization of End-to-End Speech Recognition on Mobile Devices for Named EntitiesabstractWe study the effectiveness of several techniques to personalize end-to-end speech models and improve the recognition of proper names relevant to the user. These techniques differ in the amounts of user effort required to provide supervision, and are evaluated on how they impact speech recognition performance. We propose using keyword-dependent precision and recall metrics to measure vocabulary acquisition performance. We evaluate the algorithms on a dataset that we designed to contain names of persons that are difficult to recognize. Therefore, the baseline recall rate for proper names in this dataset is very low: 2.4%. A data synthesis approach we developed brings it to 48.6%, with no need for speech input from the user. With speech input, if the user corrects only the names, the name recall rate improves to 64.4%. If the user corrects all the recognition errors, we achieve the best recall of 73.5%. To eliminate the need to upload user data and store personalized models on a server, we focus on performing the entire personalization workflow on a mobile device. Khe Chai Sim, Leif Johnson, Giovanni Motta, Lillian Zhou, Françoise Beaufays, Arnaud Benard, Dhruv Guliani, Andreas Kabel, Nikhil Khare, Tamar Lucassen, Petr Zadrazil, Harry Zhang |
ASRU | 5 |
| 2019 | Federated Learning of N-Gram Language ModelsabstractMingqing Chen, Ananda Theertha Suresh, Rajiv Mathews, Adeline Wong, Cyril Allauzen, Françoise Beaufays, Michael Riley. Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL). 2019. Mingqing Chen, Ananda Theertha Suresh, Rajiv Mathews, Adeline Wong, Cyril Allauzen, Françoise Beaufays, Michael Riley 0001 |
CoNLL | 6 |
| 2019 | An Investigation into On-Device Personalization of End-to-End Automatic Speech Recognition ModelsabstractSpeaker-independent speech recognition systems trained with data from many users are generally robust against speaker variability and work well for a large population of speakers. However, these systems do not always generalize well for users with very different speech characteristics. This issue can be addressed by building personalized systems that are designed to work well for each specific user. In this paper, we investigate the idea of securely training personalized end-to-end speech recognition models on mobile devices so that user data and models never leave the device and are never stored on a server. We study how the mobile training environment impacts performance by simulating on-device data consumption. We conduct experiments using data collected from speech impaired users for personalization. Our results show that personalization achieved 63.7\% relative word error rate reduction when trained in a server environment and 58.1% in a mobile environment. Moving to on-device personalization resulted in 18.7% performance degradation, in exchange for improved scalability and data privacy. To train the model on device, we split the gradient computation into two and achieved 45% memory reduction at the expense of 42% increase in training time. Khe Chai Sim, Petr Zadrazil, Françoise Beaufays |
INTERSPEECH | 3 |
| 2017 | Pronunciation Learning with RNN-Transducers
Antoine Bruguier, Danushen Gnanapragasam, Leif Johnson, Kanishka Rao, Françoise Beaufays |
INTERSPEECH | 5 |
| 2017 | Recurrent Neural Aligner: An Encoder-Decoder Neural Network Model for Sequence to Sequence Mapping
Hasim Sak, Matt Shannon, Kanishka Rao, Françoise Beaufays |
INTERSPEECH | 4 |
| 2016 | Personalized speech recognition on mobile devicesabstractWe describe a large vocabulary speech recognition system that is accurate, has low latency, and yet has a small enough memory and computational footprint to run faster than real-time on a Nexus 5 Android smartphone. We employ a quantized Long Short-Term Memory (LSTM) acoustic model trained with connectionist temporal classification (CTC) to directly predict phoneme targets, and further reduce its memory footprint using an SVD-based compression scheme. Additionally, we minimize our memory footprint by using a single language model for both dictation and voice command domains, constructed using Bayesian interpolation. Finally, in order to properly handle device-specific information, such as proper names and other context-dependent information, we inject vocabulary items into the decoder graph and bias the language model on-the-fly. Our system achieves 13.5% word error rate on an open-ended dictation task, running with a median speed that is seven times faster than real-time. Ian McGraw, Rohit Prabhavalkar, Raziel Alvarez, Montse Gonzalez Arenas, Kanishka Rao, David Rybach, Ouais Alsharif, Hasim Sak, Alexander Gruenstein, Françoise Beaufays, Carolina Parada |
ICASSP | 10 |
| 2016 | Learning Personalized Pronunciations for Contact Name Recognition
Antoine Bruguier, Fuchun Peng, Françoise Beaufays |
INTERSPEECH | 3 |
| 2015 | Long short term memory neural network for keyboard gesture decodingabstractGesture typing is an efficient input method for phones and tablets using continuous traces created by a pointed object (e.g., finger or stylus). Translating such continuous gestures into textual input is a challenging task as gesture inputs exhibit many features found in speech and handwriting such as high variability, co-articulation and elision. In this work, we address these challenges with a hybrid approach, combining a variant of recurrent networks, namely Long Short Term Memories [1] with conventional Finite State Transducer decoding [2]. Results using our approach show considerable improvement relative to a baseline shape-matching-based system, amounting to 4% and 22% absolute improvement respectively for small and large lexicon decoding on real datasets and 2% on a synthetic large scale dataset. Ouais Alsharif, Tom Ouyang, Françoise Beaufays, Shumin Zhai, Thomas M. Breuel, Johan Schalkwyk |
ICASSP | 3 |
| 2015 | Fix it where it fails: Pronunciation learning by mining error corrections from speech logsabstractThe pronunciation dictionary, or lexicon, is an essential component in an automatic speech recognition (ASR) system in that incorrect pronunciations cause systematic misrecognitions. It typically consists of a list of word-pronunciation pairs written by linguists, and a grapheme-to-phoneme (G2P) engine to generate pronunciations for words not in the list. The hand-generated list can never keep pace with the growing vocabulary of a live speech recognition system, and the G2P is usually of limited accuracy. This is especially true for proper names whose pronunciations may be influenced by various historical or foreign-origin factors. In this paper, we propose a language-independent approach to detect misrecognitions and their corrections from voice search logs. We learn previously unknown pronunciations from this data, and demonstrate that they significantly improve the quality of a production-quality speech recognition system. Zhenzhen Kou, Daisy Stanton, Fuchun Peng, Françoise Beaufays, Trevor Strohman |
ICASSP | 4 |
| 2015 | Automatic pronunciation verification for speech recognitionabstractPronunciations for words are a critical component in an automated speech recognition system (ASR) as mis-recognitions may be caused by missing or inaccurate pronunciations. The need for high quality pronunciations has recently motivated data-driven techniques to generate them [1]. We propose a data-driven and language-independent framework for verification of such pronunciations to further improve the lexicon quality in ASR. New candidate pronunciations are verified by re-recognizing historical audio logs and examining the associated recognition costs. We build an additional pronunciation quality feature from word and pronunciation frequencies in logs. A machine learned classifier trained on these features achieves nearly 90% accuracy in labeling good vs bad pronunciations across all languages we tested. New pronunciations verified as good may be added to a dictionary, while bad pronunciations may be discarded or sent to experts for further evaluation. We simultaneously verify 5,000 to 30,000 new pronunciations within a few hours and show improvements in the ASR performance as a result of including pronunciations verified by this system. Kanishka Rao, Fuchun Peng, Françoise Beaufays |
ICASSP | 3 |
| 2015 | Grapheme-to-phoneme conversion using Long Short-Term Memory recurrent neural networksabstractGrapheme-to-phoneme (G2P) models are key components in speech recognition and text-to-speech systems as they describe how words are pronounced. We propose a G2P model based on a Long Short-Term Memory (LSTM) recurrent neural network (RNN). In contrast to traditional joint-sequence based G2P approaches, LSTMs have the flexibility of taking into consideration the full context of graphemes and transform the problem from a series of grapheme-to-phoneme conversions to a word-to-pronunciation conversion. Training joint-sequence based G2P require explicit grapheme-to-phoneme alignments which are not straightforward since graphemes and phonemes don't correspond one-to-one. The LSTM based approach forgoes the need for such explicit alignments. We experiment with unidirectional LSTM (ULSTM) with different kinds of output delays and deep bidirectional LSTM (DBLSTM) with a connectionist temporal classification (CTC) layer. The DBLSTM-CTC model achieves a word error rate (WER) of 25.8% on the public CMU dataset for US English. Combining the DBLSTM-CTC model with a joint n-gram model results in a WER of 21.3%, which is a 9% relative improvement compared to the previous best WER of 23.4% from a hybrid system. Kanishka Rao, Fuchun Peng, Hasim Sak, Françoise Beaufays |
ICASSP | 4 |
| 2015 | Learning acoustic frame labeling for speech recognition with recurrent neural networksabstractWe explore alternative acoustic modeling techniques for large vocabulary speech recognition using Long Short-Term Memory recurrent neural networks. For an acoustic frame labeling task, we compare the conventional approach of cross-entropy (CE) training using fixed forced-alignments of frames and labels, with the Connectionist Temporal Classification (CTC) method proposed for labeling unsegmented sequence data. We demonstrate that the latter can be implemented with finite state transducers. We experiment with phones and context dependent HMM states as acoustic modeling units. We also investigate the effect of context in acoustic input by training unidirectional and bidirectional LSTM RNN models. We show that a bidirectional LSTM RNN CTC model using phone units can perform as well as an LSTM RNN model trained with CE using HMM state alignments. Finally, we also show the effect of sequence discriminative training on these models and show the first results for sMBR training of CTC models. Hasim Sak, Andrew W. Senior, Kanishka Rao, Ozan Irsoy, Alex Graves, Françoise Beaufays, Johan Schalkwyk |
ICASSP | 6 |
| 2015 | Garbage modeling for on-device speech recognitionabstractUser interactions with mobile devices increasingly depend on voice as a primary input modality. Due to the disadvantages of sending audio across potentially spotty network connections for speech recognition, in recent years there has been growing attention to performing recognition on-device. The limited computational resources, however, typically require additional model constraints. In this work, we explore the task of on-device utterance verification, wherein the recognizer must transcribe an utterance if it is in a target set or reject it as being out of domain. We present a data-driven methodology for mining tens of thousands of target phrases from an existing corpus. We then compare two common garbage-modeling approaches to utterance verification: a sub-word rejection model and a white-listed n-gram model. We examine a deficiency of the sub-word modeling approach and introduce a novel modification that makes use of common prefixes between targeted phrases and non-targeted phrases. We show good performance in the trade-off between recall and word error rate using both the prefix and white-listed n-gram approaches. Finally, we evaluate the prefix-based approach in a hybrid setting where rejected instances are sent to a server-side recognizer. Christophe Van Gysel, Leonid Velikovich, Ian McGraw, Françoise Beaufays |
INTERSPEECH | 4 |
| 2015 | Composition-based on-the-fly rescoring for salient n-gram biasingabstractWe introduce a technique for dynamically applying contextually-derived language models to a state-of-the-art speech recognition system. These generally small-footprint models can be seen as a generalization of cache-based models [1], whereby contextually salient n-grams are derived from relevant sources (not just user generated language) to produce a model intended for combination with the baseline language model. The derived models are applied during first-pass decoding as a form of on-the-fly composition between the decoder search graph and the set of weighted contextual n-grams. We present a construction algorithm which takes a trie representing the contextual n-grams and produces a weighted finite state automaton which is more compact than a standard n-gram machine. Finally, we present a set of empirical results on the recognition of spoken search queries where a contextual model encoding recent trending queries is applied using the proposed technique. Keith B. Hall, Eunjoon Cho, Cyril Allauzen, Françoise Beaufays, Noah Coccaro, Kaisuke Nakajima, Michael Riley 0001, Brian Roark, David Rybach, Linda Zhang 0004 |
INTERSPEECH | 4 |
| 2015 | Large vocabulary automatic speech recognition for childrenabstractRecently, Google launched YouTube Kids, a mobile application for children, that uses a speech recognizer built specifically for recognizing children’s speech. In this paper we present techniques we explored to build such a system. We describe the use of a neural network classifier to identify matched acoustic training data, filtering data for language modeling to reduce the chance of producing offensive results. We also compare long short-term memory (LSTM) recurrent networks to convolutional, LSTM, deep neural networks (CLDNN). We found that a CLDNN acoustic model outperforms an LSTM across a variety of different conditions, but does not specifically model child speech relatively better than adult. Overall, these findings allow us to build a successful, state-of-the-art large vocabulary speech recognizer for both children and adults. Hank Liao, Golan Pundak, Olivier Siohan, Melissa K. Carroll, Noah Coccaro, Qi-Ming Jiang, Tara N. Sainath, Andrew W. Senior, Françoise Beaufays, Michiel Bacchiani |
INTERSPEECH | 9 |
| 2015 | Fast and accurate recurrent neural network acoustic models for speech recognitionabstractWe have recently shown that deep Long Short-Term Memory (LSTM) recurrent neural networks (RNNs) outperform feed forward deep neural networks (DNNs) as acoustic models for speech recognition. More recently, we have shown that the performance of sequence trained context dependent (CD) hidden Markov model (HMM) acoustic models using such LSTM RNNs can be equaled by sequence trained phone models initialized with connectionist temporal classification (CTC). In this paper, we present techniques that further improve performance of LSTM RNN acoustic models for large vocabulary speech recognition. We show that frame stacking and reduced frame rate lead to more accurate models and faster decoding. CD phone modeling leads to further improvements. We also present initial results for LSTM RNN models outputting words directly. Hasim Sak, Andrew W. Senior, Kanishka Rao, Françoise Beaufays |
INTERSPEECH | 4 |
| 2014 | Pronunciation learning for named-entities through crowd-sourcingabstractObtaining good pronunciations for named-entities poses a challenge for automated speech recognition because named-entities are diverse in nature and origin, and new entities come up every day. In this paper, we investigate the feasibility of learning named-entity pronunciations using crowd-sourcing. By collecting audio samples from non-linguistic-expert speak-ers with Mechanical Turk and learning from them, we can quickly derive pronunciations that are more accurate in speech recognition tests than manual pronunciations generated by lin-guistic experts. Compared to traditional approaches of generat-ing pronunciations, this new approach proves to be cheap, fast, and quite accurate. 1. Attapol Rutherford, Fuchun Peng, Françoise Beaufays |
INTERSPEECH | 3 |
| 2014 | Long short-term memory recurrent neural network architectures for large scale acoustic modelingabstractLong Short-Term Memory (LSTM) is a specific recurrent neural network (RNN) architecture that was designed to model temporal sequences and their long-range dependencies more accurately than conventional RNNs. In this paper, we explore LSTM RNN architectures for large scale acoustic modeling in speech recognition. We recently showed that LSTM RNNs are more effective than DNNs and conventional RNNs for acoustic modeling, considering moderately-sized models trained on a single machine. Here, we introduce the first distributed training of LSTM RNNs using asynchronous stochastic gradient descent optimization on a large cluster of machines. We show that a two-layer deep LSTM RNN where each LSTM layer has a linear recurrent projection layer can exceed state-of-the-art speech recognition performance. This architecture makes more effective use of model parameters than the others considered, converges quickly, and outperforms a deep feed forward neural network having an order of magnitude more parameters. Index Terms: Long Short-Term Memory, LSTM, recurrent neural network, RNN, speech recognition, acoustic modeling. Hasim Sak, Andrew W. Senior, Françoise Beaufays |
INTERSPEECH | 3 |
| 2013 | Search results based N-best hypothesis rescoring with maximum entropy classificationabstractWe propose a simple yet effective method for improving speech recognition by reranking the N-best speech recognition hypotheses using search results. We model N-best reranking as a binary classification problem and select the hypothesis with the highest classification confidence. We use query-specific features extracted from the search results to encode domain knowledge and use it with a maximum entropy classifier to rescore the N-best list. We show that rescoring even only the top 2 hypotheses, we can obtain a significant 3% absolute sentence accuracy (SACC) improvement over a strong baseline on production traffic from an entertainment domain. Fuchun Peng, Scott Roy, Ben Shahshahani, Françoise Beaufays |
ASRU | 4 |
| 2013 | Mixture of mixture n-gram language modelsabstractThis paper presents a language model adaptation technique to build a single static language model from a set of language models each trained on a separate text corpus while aiming to maximize the likelihood of an adaptation data set given as a development set of sentences. The proposed model can be considered as a mixture of mixture language models. The mixture model at the top level is a sentence-level mixture model where each sentence is assumed to be drawn from one of a discrete set of topic or task clusters. After selecting a cluster, each n-gram is assumed to be drawn from one of the given n-gram language models. We estimate cluster mixture weights and n-gram language model mixture weights for each cluster using the expectation-maximization (EM) algorithm to seek the parameter estimates maximizing the likelihood of the development sentences. This mixture of mixture models can be represented efficiently as a static n-gram language model using the previously proposed Bayesian language model interpolation technique. We show a significant improvement with this technique (both perplexity and WER) compared to the standard one level interpolation scheme. Hasim Sak, Cyril Allauzen, Kaisuke Nakajima, Françoise Beaufays |
ASRU | 4 |
| 2013 | Language model capitalizationabstractIn many speech recognition systems, capitalization is not an inherent component of the language model: training corpora are down cased, and counts are accumulated for sequences of lower-cased words. This level of modeling is sufficient for automating voice commands or otherwise enabling users to communicate with a machine, but when the recognized speech is intended to be read by a person, such as in email dictation or even some web search applications, the lack of capitalization of the user's input can add an extra cognitive load on the reader. For these cases, speech recognition systems often post-process the recognized text to restore capitalization. We propose folding capitalization directly in the recognition language model. Instead of post-processing, we take the approach that language should be represented in all its richness, with capitalization, diacritics, and other special symbols. With that perspective, we describe a strategy to handle poorly capitalized or uncapitalized training corpora for language modeling. The resulting recognition system retains the accuracy/latency/memory tradeoff of our uncapitalized production recognizer, while providing properly cased outputs. Françoise Beaufays, Brian Strope |
ICASSP | 1 |
| 2013 | Language model verbalization for automatic speech recognitionabstractTranscribing speech in properly formatted written language presents some challenges for automatic speech recognition systems. The difficulty arises from the conversion ambiguity between verbal and written language in both directions. Non-lexical vocabulary items such as numeric entities, dates, times, abbreviations and acronyms are particularly ambiguous. This paper describes a finite-state transducer based approach that improves proper transcription of these entities. The approach involves training a language model in the written language domain, and integrating verbal expansions of vocabulary items as a finite-state model into the decoding graph construction. We build an inverted finite-state transducer to map written vocabulary items to alternate verbal expansions using rewrite rules. Then, this verbalizer transducer is composed with the n-gram language model to obtain a verbalized language model, whose input labels are in the verbal language domain while output labels are in the written language domain. We show that the proposed approach is very effective in improving the recognition accuracy of numeric entities. Hasim Sak, Françoise Beaufays, Kaisuke Nakajima, Cyril Allauzen |
ICASSP | 2 |
| 2013 | Written-domain language modeling for automatic speech recognitionabstractLanguage modeling for automatic speech recognition (ASR) systems has been traditionally in the verbal domain. In this paper, we present finite-state modeling techniques that we developed for language modeling in the written domain. The first technique we describe is for the verbalization of written-domain vocabulary items, which include lexical and non-lexical entities. The second technique is the decomposition–recomposition approach to address the out-of-vocabulary (OOV) and the data sparsity problems with non-lexical entities such as URLs, email addresses, phone numbers, and dollar amounts. We evaluate the proposed written-domain language modeling approaches on a very large vocabulary speech recognition system for English. We show that the written-domain language modeling improves the speech recognition and the ASR transcript rendering accuracy in the written domain over a baseline system using a verbal-domain language model. In addition, the writtendomain system is much simpler since it does not require complex and error-prone text normalization and denormalization rules, which are generally required for verbal-domain language modeling. Hasim Sak, Yun-Hsuan Sung, Françoise Beaufays, Cyril Allauzen |
INTERSPEECH | 3 |
| 2012 | Recognition of multilingual speech in mobile applicationsabstractWe evaluate different architectures to recognize multilingual speech for real-time mobile applications. In particular, we show that combining the results of several recognizers greatly outperforms other solutions such as training a single large multilingual system or using an explicit language identification system to select the appropriate recognizer. Experiments are conducted on a trilingual English-French-Mandarin mobile speech task. The data set includes Google searches, Maps queries, as well as more general inputs such as email and short message dictation. Without pre-specifying the input language, the combined system achieves comparable accuracy to that of the monolingual systems when the input language is known. The combined system is also roughly 5% absolute better than an explicit language identification approach, and 10% better than a single large multilingual system. Jui-Ting Huang, Françoise Beaufays, Brian Strope, Yun-Hsuan Sung |
ICASSP | 3 |
| 2011 | Recognizing English queries in Mandarin Voice SearchabstractRecent improvements in speech recognition technology, along with increased computing power and bigger datasets, have considerably improved the state of the art in the field, making it possible for commercial apps such as Google Voice Search to serve users in their everyday mobile search needs. Deploying such systems in various countries has shown us the extent to which multilingualism is present in some cultures, and the need for better solutions to handle it in our speech recognition systems. In this paper, we describe a few early data sharing and model combination experiments we did to improve the recognition of English queries made to Mandarin Voice Search, in Taiwan. We obtained a 12% relative sentence accuracy improvement over a baseline system already including some support for English queries. Hung-An Chang, Yun-Hsuan Sung, Brian Strope, Françoise Beaufays |
ICASSP | 4 |
| 2010 | Unsupervised discovery and training of maximally dissimilar cluster modelsabstractOne of the difficult problems of acoustic modeling for Automatic Speech Recognition (ASR) is how to adequately model the wide variety of acoustic conditions which may be present in the data. The problem is especially acute for tasks such as Google Search by Voice, where the amount of speech available per transaction is small, and adaptation techniques start showing their limitations. As training data from a very large user population is available however, it is possible to identify and jointly model subsets of the data with similar acoustic qualities. We describe a technique which allows us to perform this modeling at scale on large amounts of data by learning a treestructured partition of the acoustic space, and we demonstrate that we can significantly improve recognition accuracy in various conditions through unsupervised Maximum Mutual Information (MMI) training. Being fully unsupervised, this technique scales easily to increasing numbers of conditions. Index Terms: acoustic modeling, unsupervised learning, clustering, too much data. Françoise Beaufays, Vincent Vanhoucke, Brian Strope |
INTERSPEECH | 1 |
| 2009 | Revisiting graphemes with increasing amounts of dataabstractLetter units, or graphemes, have been reported in the literature as a surprisingly effective substitute to the more traditional phoneme units, at least in languages that enjoy a strong correspondence between pronunciation and orthography. For English however, where letter symbols have less acoustic consistency, previously reported results fell short of systems using highly-tuned pronunciation lexicons. Grapheme units simplify system design, but since graphemes map to a wider set of acoustic realizations than phonemes, we should expect grapheme-based acoustic models to require more training data to capture these variations. In this paper, we compare the rate of improvement of grapheme and phoneme systems trained with datasets ranging from 450 to 1200 hours of speech. We consider various grapheme unit configurations, including using letter-specific, onset, and coda units. We show that the grapheme systems improve faster and, depending on the lexicon, reach or surpass the phoneme baselines with the largest training set. Yun-Hsuan Sung, Thad Hughes, Françoise Beaufays, Brian Strope |
ICASSP | 3 |
| 2008 | Deploying GOOG-411: Early lessons in data, measurement, and testingabstractWe describe our early experience building and optimizing GOOG-411, a fully automated, voice-enabled, business finder. We show how taking an iterative approach to system development allows us to optimize the various components of the system, thereby progressively improving user-facing metrics. We show the contributions of different data sources to recognition accuracy. For business listing language models, we see a nearly linear performance increase with the logarithm of the amount of training data. To date, we have improved our correct accept rate by 25% absolute, and increased our transfer rate by 35% absolute. Michiel Bacchiani, Françoise Beaufays, Johan Schalkwyk, Mike Schuster, Brian Strope |
ICASSP | 2 |
| 2003 | Using speech/non-speech detection to bias recognition search on noisy dataabstractThis paper focuses on the recognition of noisy speech. We show that the decoding of a noisy speech waveform can be facilitated if the recognizer has explicit knowledge of where it should hypothesize speech phones, and where it should map the acoustics to non-speech phones. We build a speech/non-speech detector and use its output as an additional front-end feature. We show that by appropriately weighting the contribution of this feature in the decoder and by modifying the acoustic models accordingly, we can penalize speech/non-speech confusions and consequently reduce the recognition error rate. This approach gives a 12% overall error rate reduction on a wide variety of recognition tasks and noise characteristics without degrading performance on clean test data. A simple extension of the approach boosts recognition improvements on noisy test sets to 14% overall. Françoise Beaufays, Daniel Boies, Mitch Weintraub, Qifeng Zhu 0001 |
ICASSP (1) | 1 |
| 2003 | Learning Name Pronunciations in Automatic Speech Recognition SystemsabstractMany speech recognition systems that provide over-the phone services, e.g. name dialers, stock quote providers, location finders, rely on the accurate recognition of proper names. For this to happen, the systems need to know how their users will pronounce these words. However, predicting the pronunciation of a proper name is a notoriously difficult problem as it depends on the origin of the name, the linguistic background of the speaker, and other cultural and sociological factors, in addition of course to the word spelling. In this paper, we describe a data-driven method that learns proper name pronunciations from audio samples of these words. The algorithm relies on the machinery of a general purpose speech recognizer to find the phone sequence that best matches the sample speech waveforms. In addition, it incorporates linguistic knowledge automatically acquired from a pronunciation dictionary to ensure that the learned pronunciations are "reasonable" from a linguistic viewpoint. We show on a corporate name dialing database that the proposed algorithm reduces the call routing error rate by 40% compared to a reference letter-to-phone pronunciation engine. Françoise Beaufays, Ananth Sankar, Shaun Williams, Mitch Weintraub |
ICTAI | 1 |
| 2003 | Learning linguistically valid pronunciations from acoustic dataabstractWe describe an algorithm to learn word pronunciations from acoustic data. The algorithm jointly optimizes the pronunciation of a word using (a) the acoustic match of this pronunciation to the observed data, and (b) how “linguistically reasonable ” the pronunciation is. Variations of word pronunciations in the recognition dictionary (which was created by linguists), are used to train a model of whether new hypothesized pronunciations are reasonable or not. The algorithm is well-suited for proper name pronunciation learning. Experiments on a corporate name dialing database show 40 % error rate reduction with respect to a letter-to-phone pronunciation engine. 1. Françoise Beaufays, Ananth Sankar, Shaun Williams, Mitch Weintraub |
INTERSPEECH | 1 |
| 2002 | Porting channel robustness across languages
Françoise Beaufays, Daniel Boies, Mitch Weintraub |
INTERSPEECH | 1 |
| 1999 | Discriminative mixture weight estimation for large Gaussian mixture modelsabstractThis paper describes a new approach to acoustic modeling for large vocabulary continuous speech recognition (LVCSR) systems. Each phone is modeled with a large Gaussian mixture model (GMM) whose context-dependent mixture weights are estimated with a sentence-level discriminative training criterion. The estimation problem is cast in a neural network framework, which enables the incorporation of the appropriate constraints on the mixture weight vectors, and allows a straight-forward training procedure, based on steepest descent. Experiments conducted on the Callhome-English and Switchboard databases show a significant improvement of the acoustic model performance, and a somewhat lesser improvement with the combined acoustic and language models. Françoise Beaufays, Mitch Weintraub, Yochai Konig |
ICASSP | 1 |
| 1999 | Robust text-independent speaker identification over telephone channelsabstractThis paper addresses the issue of closed-set text-independent speaker identification from samples of speech recorded over the telephone. It focuses on the effects of acoustic mismatches between training and testing data, and concentrates on two approaches: (1) extracting features that are robust against channel variations and (2) transforming the speaker models to compensate for channel effects. First, an experimental study shows that optimizing the front end processing of the speech signal can significantly improve speaker recognition performance. A new filterbank design is introduced to improve the robustness of the speech spectrum computation in the front-end unit. Next, a new feature based on spectral slopes is described. Its ability to discriminate between speakers is shown to be superior to that of the traditional cepstrum. This feature can be used alone or combined with the cepstrum. The second part of the paper presents two model transformation methods that further reduce channel effects. These methods make use of a locally collected stereo database to estimate a speaker-independent variance transformation for each speech feature used by the classifier. The transformations constructed on this stereo database can then be applied to speaker models derived from other databases. Combined, the methods developed in this paper resulted in a 38% relative improvement on the closed-set 30-s training 5-s testing condition of the NIST'95 Evaluation task, after cepstral mean removal. Hema A. Murthy, Françoise Beaufays, Larry Heck, Mitch Weintraub |
IEEE Trans. Speech Audio Process. | 2 |
| 1997 | Model transformation for robust speaker recognition from telephone dataabstractIn the context of automatic speaker recognition, we propose a model transformation technique that renders speaker models more robust to acoustic mismatches and to data scarcity by appropriately increasing their variances. We use a stereo database containing speech recorded simultaneously under different acoustic conditions to derive a synthetic variance distribution. This distribution is then used to modify the variances of other speaker models from other telephone databases. The technique is illustrated with experiments conducted on a locally collected database and on the NIST'95 and '96 subsets of the Switchboard corpus. Françoise Beaufays, Mitch Weintraub |
ICASSP | 1 |
| 1997 | Neural-network based measures of confidence for word recognitionabstractThis paper proposes a probabilistic framework to define and evaluate confidence measures for word recognition. We describe a novel method to combine different knowledge sources and estimate the confidence in a word hypothesis, via a neural network. We also propose a measure of the joint performance of the recognition and confidence systems. The definitions and algorithms are illustrated with results on the Switchboard Corpus. Mitch Weintraub, Françoise Beaufays, Zeév Rivlin, Yochai Konig, Andreas Stolcke |
ICASSP | 2 |
| 1996 | Diagrammatic Derivation of Gradient Algorithms for Neural NetworksabstractDeriving gradient algorithms for time-dependent neural network structures typically requires numerous chain rule expansions, diligent bookkeeping, and careful manipulation of terms. In this paper, we show how to derive such algorithms via a set of simple block diagram manipulation rules. The approach provides a common framework to derive popular algorithms including backpropagation and backpropagation-through-time without a single chain rule expansion. Additional examples are provided for a variety of complicated architectures to illustrate both the generality and the simplicity of the approach. Eric A. Wan, Françoise Beaufays |
Neural Comput. | 2 |
| 1995 | Training data clustering for improved speech recognitionabstractWe present an approach to cluster the training data for automatic speech recognition (ASR). A relativeentropy based distance metric between training data clusters is defined. This metric is used to hierarchically cluster the training data. The metric can also be used to select the closest training data clusters given a small amount of data from the test speaker. The selected clusters are then used to estimate a set of hidden Markov models (HMMs) for recognizing the speech from the test speaker. We present preliminary experimental results of the clustering algorithm and its application to ASR. 1 Introduction While progress in ASR has been encouraging, it has become increasingly clear that ASR systems must perform well in the presence of mismatches between the training and testing environments. ASR systems trained in one environment often perform poorly in a new environment due to mismatches between the training and testing conditions. Common sources of mismatches include different tran... Ananth Sankar, Françoise Beaufays, Vassilios Digalakis |
EUROSPEECH | 2 |
| 1994 | Relating Real-Time Backpropagation and Backpropagation-Through-Time: An Application of Flow Graph InterreciprocityabstractWe show that signal flow graph theory provides a simple way to relate two popular algorithms used for adapting dynamic neural networks, real-time backpropagation and backpropagation-through-time. Starting with the flow graph for real-time backpropagation, we use a simple transposition to produce a second graph. The new graph is shown to be interreciprocal with the original and to correspond to the backpropagation-through-time algorithm. Interreciprocity provides a theoretical argument to verify that both flow graphs implement the same overall weight update. Françoise Beaufays, Eric A. Wan |
Neural Comput. | 1 |
| 1994 | Application of neural networks to load-frequency control in power systemsabstractThis paper describes an application of layered neural networks to nonlinear power systems control. A single generator unit feeds a power line to various users whose power demand can vary over time. As a consequence of load variations, the frequency of the generator changes over time. A feedforward neural network is trained to control the steam admission valve of the turbine that drives the generator, thereby restoring the frequency to its nominal value. Frequency transients are minimized and zero steady-state error is obtained. The same technique is then applied to control a system composed of two single units tied together through a power line. Electric load variations can happen independently in both units. Both neural controllers are trained with the back propagation-through-time algorithm. Use of a neural network to model the dynamic system is avoided by introducing the Jacobian matrices of the system in the back propagation chain used in controller training. Françoise Beaufays, Youssef Lotfy Abdel-Magid, Bernard Widrow |
Neural Networks | 1 |