EDBT 2026 Demo / reviewers in the wild / expert
Michiel Bacchiani
dblp:39/2964 · also Michiel Adriaan Unico Bacchiani
· DBLP profile ↗
61ranked-venue papers
14as first author
7since 2021 · last 2024
0000-0003-4527-0197ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 57 · 13 first-author · 7 since 2021Artificial intelligence and machine learning · 33 · 5 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | FLEURS-R: A Restored Multilingual Speech Corpus for Generation Tasks
Yuma Koizumi, Shigeki Karita, Heiga Zen, Jason Riesa, Haruko Ishikawa, Michiel Bacchiani |
INTERSPEECH | 7 |
| 2023 | LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus
Yuma Koizumi, Heiga Zen, Shigeki Karita, Yifan Ding 0004, Kohei Yatabe, Nobuyuki Morioka, Michiel Bacchiani, Yu Zhang 0033, Wei Han 0002, Ankur Bapna |
INTERSPEECH | 7 |
| 2022 | Knowledge Transfer from Large-Scale Pretrained Language Models to End-To-End Speech RecognizersabstractEnd-to-end speech recognition is a promising technology for enabling compact automatic speech recognition (ASR) systems since it can unify the acoustic and language model into a single neural network. However, as a drawback, training of end-to-end speech recognizers always requires transcribed utterances. Since end-to-end models are also known to be severely data hungry, this constraint is crucial especially because obtaining transcribed utterances is costly and can possibly be impractical or impossible. This paper proposes a method for alleviating this issue by transferring knowledge from a language model neural network that can be pretrained with text-only data. Specifically, this paper attempts to transfer semantic knowledge acquired in embedding vectors of large-scale language models. Since embedding vectors can be assumed as implicit representations of linguistic information such as part-of-speech, intent, and so on, those are also expected to be useful modeling cues for ASR decoders. This paper extends two types of ASR decoders, attention-based decoders and neural transducers, by modifying training loss functions to include embedding prediction terms. The proposed systems were shown to be effective for error rate reduction without incurring extra computational costs in the decoding phase. Yotaro Kubo, Shigeki Karita, Michiel Bacchiani |
ICASSP | 3 |
| 2022 | SNRi Target Training for Joint Speech Enhancement and Recognition
Yuma Koizumi, Shigeki Karita, Arun Narayanan, Sankaran Panchapagesan, Michiel Bacchiani |
INTERSPEECH | 5 |
| 2022 | SpecGrad: Diffusion Probabilistic Model based Neural Vocoder with Adaptive Noise Spectral ShapingabstractNeural vocoder using denoising diffusion probabilistic model (DDPM) has been improved by adaptation of the diffusion noise distribution to given acoustic features.In this study, we propose SpecGrad that adapts the diffusion noise so that its timevarying spectral envelope becomes close to the conditioning log-mel spectrogram.This adaptation by time-varying filtering improves the sound quality especially in the high-frequency bands.It is processed in the time-frequency domain to keep the computational cost almost the same as the conventional DDPMbased neural vocoders.Experimental results showed that Spec-Grad generates higher-fidelity speech waveform than conventional DDPM-based neural vocoders in both analysis-synthesis and speech enhancement scenarios.Audio demos are available at wavegrad.github.io/specgrad/. Yuma Koizumi, Heiga Zen, Kohei Yatabe, Nanxin Chen, Michiel Bacchiani |
INTERSPEECH | 5 |
| 2022 | Wavefit: an Iterative and Non-Autoregressive Neural Vocoder Based on Fixed-Point IterationabstractDenoising diffusion probabilistic models (DDPMs) and generative adversarial networks (GANs) are popular generative models for neural vocoders. The DDPMs and GANs can be characterized by the iterative denoising framework and adversarial training, respectively. This study proposes a fast and high-quality neural vocoder called WaveFit, which integrates the essence of GANs into a DDPM-like iterative framework based on fixed-point iteration. WaveFit iteratively denoises an input signal, and trains a deep neural network (DNN) for minimizing an adversarial loss calculated from intermediate outputs at all iterations. Subjective (side-by-side) listening tests showed no statistically significant differences in naturalness between human natural speech and those synthesized by WaveFit with five iterations. Furthermore, the inference speed of WaveFit was more than 240 times faster than WaveRNN. Audio demos are available at google.github.io/df-conformer/wavefit/. Yuma Koizumi, Kohei Yatabe, Heiga Zen, Michiel Bacchiani |
SLT | 4 |
| 2021 | A Comparative Study on Neural Architectures and Training Methods for Japanese Speech RecognitionabstractEnd-to-end (E2E) modeling is advantageous for automatic speech recognition (ASR) especially for Japanese since word-based tokenization of Japanese is not trivial, and E2E modeling is able to model character sequences directly. This paper focuses on the latest E2E modeling techniques, and investigates their performances on character-based Japanese ASR by conducting comparative experiments. The results are analyzed and discussed in order to understand the relative advantages of long short-term memory (LSTM), and Conformer models in combination with connectionist temporal classification, transducer, and attention-based loss functions. Furthermore, the paper investigates on effectivity of the recent training techniques such as data augmentation (SpecAugment), variational noise injection, and exponential moving average. The best configuration found in the paper achieved the state-of-the-art character error rates of 4.1%, 3.2%, and 3.5% for Corpus of Spontaneous Japanese (CSJ) eval1, eval2, and eval3 tasks, respectively. The system is also shown to be computationally efficient thanks to the efficiency of Conformer transducers. Shigeki Karita, Yotaro Kubo, Michiel Bacchiani, Llion Jones |
Interspeech | 3 |
| 2020 | Joint Phoneme-Grapheme Model for End-To-End Speech RecognitionabstractThis paper proposes methods to improve a commonly used end-to-end speech recognition model, Listen-Attend-Spell (LAS). The methods we propose use multi-task learning to improve generalization of the model by leveraging information from multiple labels. The focus in this paper is on multi-task models for simultaneous signal-to-grapheme and signal-to-phoneme conversions while sharing the encoder parameters. Since phonemes are designed to be a precise description of the linguistic aspects of the speech signal, using phoneme recognition as an auxiliary task can help guiding the early stages of training to be more stable. In addition to conventional multi-task learning, we obtain further improvements by introducing a method that can exploit dependencies between labels in different tasks. Specifically, the dependencies between phonemes and grapheme sequences are considered. In conventional multi-task learning these sequences are assumed to be independent. Instead, in this paper, a joint model is proposed based on "iterative refinement" where dependency modeling is achieved by a multi-pass strategy. The proposed method is evaluated on a 28000h corpus of Japanese speech data. Performance of a conventional multi-task approach is contrasted with that of the joint model with iterative refinement. Yotaro Kubo, Michiel Bacchiani |
ICASSP | 2 |
| 2018 | State-of-the-Art Speech Recognition with Sequence-to-Sequence ModelsabstractAttention-based encoder-decoder architectures such as Listen, Attend, and Spell (LAS), subsume the acoustic, pronunciation and language model components of a traditional automatic speech recognition (ASR) system into a single neural network. In previous work, we have shown that such architectures are comparable to state-of-the-art ASR systems on dictation tasks, but it was not clear if such architectures would be practical for more challenging tasks such as voice search. In this work, we explore a variety of structural and optimization improvements to our LAS model which significantly improve performance. On the structural side, we show that word piece models can be used instead of graphemes. We also introduce a multi-head attention architecture, which offers improvements over the commonly-used single-head attention. On the optimization side, we explore synchronous training, scheduled sampling, label smoothing, and minimum word error rate optimization, which are all shown to improve accuracy. We present results with a unidirectional LSTM encoder for streaming recognition. On a 12, 500 hour voice search task, we find that the proposed changes improve the WER from 9.2% to 5.6%, while the best conventional system achieves 6.7%; on a dictation task our model achieves a WER of 4.1% compared to 5% for the conventional system. Chung-Cheng Chiu, Tara N. Sainath, Rohit Prabhavalkar, Patrick Nguyen, Anjuli Kannan, Ron J. Weiss, Kanishka Rao, Ekaterina Gonina, Navdeep Jaitly, Bo Li 0028, Jan Chorowski, Michiel Bacchiani |
ICASSP | 14 |
| 2018 | Performance of Mask Based Statistical Beamforming in a Smart Home ScenarioabstractMask based statistical beamforming, where signal statistics for the target and the interference gained from masking are used for beamforming, has shown great effectiveness in the two recent CHiME challenges. This idea has sparked interest in the research community and resulted in numerous proposed approaches based on the idea. At the same time, the advent of voice controlled smart home devices, such as Google Home and Amazon Alexa, has strengthened the need for robust far-field automatic speech recognition. In this paper, we evaluate if mask based beamforming can live up to the expectations created by the CHiME challenges and provide similar gains in a smart home scenario. To this extend, we pinpoint the main differences between the scenarios, review the recent developments and conduct extensive experiments on large scale data. These experiments show that, while a 10 % relative reduction of the word error rate can be achieved, the gains are not as high as those seen in the CHiME challenge. We also show that approaches where the frontend and back-end is trained jointly do not reach the performance level of their independently trained counterparts. On the plus side, we see a 20 % relative improvement for an evaluation set with crosstalk. Jahn Heymann, Michiel Bacchiani, Tara N. Sainath |
ICASSP | 2 |
| 2018 | Sound Source Separation Using Phase Difference and Reliable Mask Selection SelectionabstractIn this paper, we present an algorithm called Reliable Mask Selection-Phase Difference Channel Weighting (RMS-PDCW) which selects the target source masked by a noise source using the Angle of Arrival (AoA) information calculated using the phase difference information. The RMS-PDCW algorithm selects masks to apply using the information about the localized sound source and the onset detection of speech. We demonstrate that this algorithm shows relatively 5.3 percent improvement over the baseline acoustic model, which was multistyle-trained using 22 million utterances on the simulated test set consisting of real-world and interfering-speaker noise with reverberation time distribution between 0 ms and 900 ms and SNR distribution between 0 dB up to clean. Chanwoo Kim 0001, Anjali Menon, Michiel Bacchiani, Richard M. Stern |
ICASSP | 3 |
| 2018 | Spectral Distortion Model for Training Phase-Sensitive Deep-Neural Networks for Far-Field Speech RecognitionabstractIn this paper, we present an algorithm which introduces phase-perturbation to the training database when training phase-sensitive deep neural-network models. Traditional features such as log-mel or cepstral features do not have have any phase-relevant information. However features such as raw-waveform or complex spectra features contain phase-relevant information. Phase-sensitive features have the advantage of being able to detect differences in time of arrival across different microphone channels or frequency bands. However, compared to magnitude-based features, phase information is more sensitive to various kinds of distortions such as variations in microphone characteristics, reverberation, and so on. For traditional magnitude-based features, it is widely known that adding noise or reverberation, often called Multistyle-TRaining (MTR), improves robustness. In a similar spirit, we propose an algorithm which introduces spectral distortion to make the deep-learning models more robust to phase-distortion. We call this approach Spectral-Distortion TRaining (SDTR). In our experiments using a training set consisting of 22-million utterances with and without MTR, this approach reduces Word Error Rates (WERs) relatively by 3.2 % and 8.48 % respectively on test sets recorded on Google Home. Chanwoo Kim 0001, Tara N. Sainath, Arun Narayanan, Ananya Misra, Rajeev C. Nongpiur, Michiel Bacchiani |
ICASSP | 6 |
| 2018 | Multi-Dialect Speech Recognition with a Single Sequence-to-Sequence ModelabstractSequence-to-sequence models provide a simple and elegant solution for building speech recognition systems by folding separate components of a typical system, namely acoustic (AM), pronunciation (PM) and language (LM) models into a single neural network. In this work, we look at one such sequence-to-sequence model, namely listen, attend and spell (LAS) [1], and explore the possibility of training a single model to serve different English dialects, which simplifies the process of training multi-dialect systems without the need for separate AM, PM and LMs for each dialect. We show that simply pooling the data from all dialects into one LAS model falls behind the performance of a model fine-tuned on each dialect. We then look at incorporating dialect-specific information into the model, both by modifying the training targets by inserting the dialect symbol at the end of the original grapheme sequence and also feeding a 1-hot representation of the dialect information into all layers of the model. Experimental results on seven English dialects show that our proposed system is effective in modeling dialect variations within a single LAS model, outperforming a LAS model trained individually on each of the seven dialects by 3.1~16.5% relative. Bo Li 0028, Tara N. Sainath, Khe Chai Sim, Michiel Bacchiani, Eugene Weinstein, Patrick Nguyen, Yanghui Wu, Kanishka Rao |
ICASSP | 4 |
| 2018 | Sampled Connectionist Temporal ClassificationabstractThis article introduces and evaluates Sampled Connectionist Temporal Classification (CTC) which connects the CTC criterion to the Cross Entropy (CE) objective through sampling. Instead of computing the logarithm of the sum of the alignment path likelihoods, at each training step the sampled CTC only computes the CE loss between the sampled alignment path and model posteriors. It is shown that the sampled CTC objective is an unbiased estimator of an upper bound for the CTC loss, thus minimization of the sampled CTC is equivalent to the minimization of the upper bound of the CTC objective. The definition of the sampled CTC objective has the advantage that it is scalable computationally to the massive datasets using accelerated computation machines. The sampled CTC is compared with CTC in two large-scale speech recognition tasks and it is shown that sampled CTC can achieve similar WER performance of the best CTC baseline in about one fourth of the training time of the CTC baseline. Ehsan Variani, Tom Bagby, Kamel Lahouel, Erik McDermott, Michiel Bacchiani |
ICASSP | 5 |
| 2018 | Efficient Implementation of the Room Simulator for Training Deep Neural Network Acoustic ModelsabstractIn this paper, we describe how to efficiently implement an acoustic room simulator to generate large-scale simulated data for training deep neural networks.Even though Google Room Simulator in [1] was shown to be quite effective in reducing the Word Error Rates (WERs) for far-field applications by generating simulated far-field training sets, it requires a very large number of FFTs.Room Simulator used approximately 80 % of CPU usage in our CPU/GPU training architecture [2].In this work, we implement an efficient OverLap Addition (OLA) based filtering using the open-source FFTW3 library.Further, we investigate the effects of the Room Impulse Response (RIR) lengths.Experimentally, we conclude that we can cut the tail portions of RIRs whose power is less than 20 dB below the maximum power without sacrificing the speech recognition accuracy.However, we observe that cutting RIR tail more than this threshold harms the speech recognition accuracy for rerecorded test sets.Using these approaches, we were able to reduce CPU usage for the room simulator portion down to 9.69 % in CPU/GPU training architecture.Profiling result shows that we obtain 22.4 times speed-up on a single machine and 37.3 times speed up on Google's distributed training infrastructure. Chanwoo Kim 0001, Ehsan Variani, Arun Narayanan, Michiel Bacchiani |
INTERSPEECH | 4 |
| 2018 | Domain Adaptation Using Factorized Hidden Layer for Robust Automatic Speech Recognition
Khe Chai Sim, Arun Narayanan, Ananya Misra, Anshuman Tripathi, Golan Pundak, Tara N. Sainath, Parisa Haghani, Bo Li 0028, Michiel Bacchiani |
INTERSPEECH | 9 |
| 2018 | From Audio to Semantics: Approaches to End-to-End Spoken Language UnderstandingabstractConventional spoken language understanding systems consist of two main components: an automatic speech recognition module that converts audio to a transcript, and a natural language understanding module that transforms the resulting text (or top N hypotheses) into a set of domains, intents, and arguments. These modules are typically optimized independently. In this paper, we formulate audio to semantic understanding as a sequence-to-sequence problem [1]. We propose and compare various encoder-decoder based approaches that optimize both modules jointly, in an end-to-end manner. Evaluations on a real-world task show that 1) having an intermediate text representation is crucial for the quality of the predicted semantics, especially the intent arguments and 2) jointly optimizing the full system improves overall accuracy of prediction. Compared to independently trained models, our best jointly trained model achieves similar domain and intent prediction F1 scores, but improves argument word error rate by 18% relative. Parisa Haghani, Arun Narayanan, Michiel Bacchiani, Galen Chuang, Neeraj Gaur, Pedro J. Moreno 0001, Rohit Prabhavalkar, Zhongdi Qu, Austin Waters |
SLT | 3 |
| 2018 | Toward Domain-Invariant Speech Recognition via Large Scale TrainingabstractCurrent state-of-the-art automatic speech recognition systems are trained to work in specific `domains', defined based on factors like application, sampling rate and codec. When such recognizers are used in conditions that do not match the training domain, performance significantly drops. This work explores the idea of building a single domain-invariant model for varied use-cases by combining large scale training data from multiple application domains. Our final system is trained using 162,000 hours of speech. Additionally, each utterance is artificially distorted during training to simulate effects like background noise, codec distortion, and sampling rates. Our results show that, even at such a scale, a model thus trained works almost as well as those fine-tuned to specific subsets: A single model can be robust to multiple application domains, and variations like codecs and noise. More importantly, such models generalize better to unseen conditions and allow for rapid adaptation - we show that by using as little as 10 hours of data from a new domain, an adapted domain-invariant model can match performance of a domain-specific model trained from scratch using 70 times as much data. We also highlight some of the limitations of such models and areas that need addressing in future work. Arun Narayanan, Ananya Misra, Khe Chai Sim, Golan Pundak, Anshuman Tripathi, Mohamed G. Elfeky, Parisa Haghani, Trevor Strohman, Michiel Bacchiani |
SLT | 9 |
| 2017 | Improving the efficiency of forward-backward algorithm using batched computation in TensorFlowabstractSequence-level losses are commonly used to train deep neural network acoustic models for automatic speech recognition. The forward-backward algorithm is used to efficiently compute the gradients of the sequence loss with respect to the model parameters. Gradient-based optimization is used to minimize these losses. Recent work has shown that the forward-backward algorithm can be efficiently implemented as a series of matrix operations. This paper further improves the forward-backward algorithm via batched computation, a technique commonly used to improve training speed by exploiting the parallel computation of matrix multiplication. Specifically, we show how batched computation of the forward-backward algorithm can be efficiently implemented using TensorFlow to handle variable-length sequences within a mini batch. Furthermore, we also show how the batched forward-backward computation can be used to compute the gradients of the connectionist temporal classification (CTC) and maximum mutual information (MMI) losses with respect to the logits. We show, via empirical benchmarks, that the batched forward-backward computation can speed up the CTC loss and gradient computation by about 183 times when run on GPU with a batch size of 256 compared to using a batch size of 1; and by about 22 times for lattice-free MMI using a trigram phone language model for the denominator. Khe Chai Sim, Arun Narayanan, Tom Bagby, Tara N. Sainath, Michiel Bacchiani |
ASRU | 5 |
| 2017 | Generation of Large-Scale Simulated Utterances in Virtual Rooms to Train Deep-Neural Networks for Far-Field Speech Recognition in Google Home
Chanwoo Kim 0001, Ananya Misra, Kean K. Chin, Thad Hughes, Arun Narayanan, Tara N. Sainath, Michiel Bacchiani |
INTERSPEECH | 7 |
| 2017 | Acoustic Modeling for Google Home
Bo Li 0028, Tara N. Sainath, Arun Narayanan, Joe Caroselli, Michiel Bacchiani, Ananya Misra, Izhak Shafran, Hasim Sak, Golan Pundak, Kean K. Chin, Khe Chai Sim, Ron J. Weiss, Kevin W. Wilson, Ehsan Variani, Chanwoo Kim 0001, Olivier Siohan, Mitch Weintraub, Erik McDermott, Richard Rose, Matt Shannon |
INTERSPEECH | 5 |
| 2017 | End-to-End Training of Acoustic Models for Large Vocabulary Continuous Speech Recognition with TensorFlow
Ehsan Variani, Tom Bagby, Erik McDermott, Michiel Bacchiani |
INTERSPEECH | 4 |
| 2017 | Multichannel Signal Processing With Deep Neural Networks for Automatic Speech RecognitionabstractMultichannel automatic speech recognition (ASR) systems commonly separate speech enhancement, including localization, beamforming, and postfiltering, from acoustic modeling. In this paper, we perform multichannel enhancement jointly with acoustic modeling in a deep neural network framework. Inspired by beamforming, which leverages differences in the fine time structure of the signal at different microphones to filter energy arriving from different directions, we explore modeling the raw time-domain waveform directly. We introduce a neural network architecture, which performs multichannel filtering in the first layer of the network, and show that this network learns to be robust to varying target speaker direction of arrival, performing as well as a model that is given oracle knowledge of the true target speaker direction. Next, we show how performance can be improved by factoring the first layer to separate the multichannel spatial filtering operation from a single channel filterbank which computes a frequency decomposition. We also introduce an adaptive variant, which updates the spatial filter coefficients at each time frame based on the previous inputs. Finally, we demonstrate that these approaches can be implemented more efficiently in the frequency domain. Overall, we find that such multichannel neural networks give a relative word error rate improvement of more than 5% compared to a traditional beamforming-based multichannel ASR system and more than 10% compared to a single channel waveform model. Tara N. Sainath, Ron J. Weiss, Kevin W. Wilson, Bo Li 0028, Arun Narayanan, Ehsan Variani, Michiel Bacchiani, Izhak Shafran, Andrew W. Senior, Kean K. Chin, Ananya Misra, Chanwoo Kim 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2016 | Factored spatial and spectral multichannel raw waveform CLDNNsabstractMultichannel ASR systems commonly separate speech enhancement, including localization, beamforming and postfiltering, from acoustic modeling. Recently, we explored doing multichannel enhancement jointly with acoustic modeling, where beamforming and frequency decomposition was folded into one layer of the neural network [1, 2]. In this paper, we explore factoring these operations into separate layers in the network. Furthermore, we explore using multi-task learning (MTL) as a proxy for postfiltering, where we train the network to predict "clean" features as well as context-dependent states. We find that with the factored architecture, we can achieve a 10% relative improvement in WER over a single channel and a 5% relative improvement over the unfactored model from [1] on a 2,000-hour Voice Search task. In addition, by incorporating MTL, we can achieve 11% and 7% relative improvements over single channel and unfactored multichannel models, respectively. Tara N. Sainath, Ron J. Weiss, Kevin W. Wilson, Arun Narayanan, Michiel Bacchiani |
ICASSP | 5 |
| 2016 | Neural Network Adaptive Beamforming for Robust Multichannel Speech Recognition
Bo Li 0028, Tara N. Sainath, Ron J. Weiss, Kevin W. Wilson, Michiel Bacchiani |
INTERSPEECH | 5 |
| 2016 | Reducing the Computational Complexity of Multimicrophone Acoustic Models with Integrated Feature Extraction
Tara N. Sainath, Arun Narayanan, Ron J. Weiss, Ehsan Variani, Kevin W. Wilson, Michiel Bacchiani, Izhak Shafran |
INTERSPEECH | 6 |
| 2016 | Complex Linear Projection (CLP): A Discriminative Approach to Joint Feature Extraction and Acoustic Modeling
Ehsan Variani, Tara N. Sainath, Izhak Shafran, Michiel Bacchiani |
INTERSPEECH | 4 |
| 2015 | Speaker location and microphone spacing invariant acoustic modeling from raw multichannel waveformsabstractMultichannel ASR systems commonly use separate modules to perform speech enhancement and acoustic modeling. In this paper, we present an algorithm to do multichannel enhancement jointly with the acoustic model, using a raw waveform convolutional LSTM deep neural network (CLDNN). We will show that our proposed method offers ~5% relative improvement in WER over a log-mel CLDNN trained on multiple channels. Analysis shows that the proposed network learns to be robust to varying angles of arrival for the target speaker, and performs as well as a model that is given oracle knowledge of the true location. Finally, we show that training such a network on inputs captured using multiple (linear) array configurations results in a model that is robust to a range of microphone spacings. Tara N. Sainath, Ron J. Weiss, Kevin W. Wilson, Arun Narayanan, Michiel Bacchiani, Andrew W. Senior |
ASRU | 5 |
| 2015 | Large vocabulary automatic speech recognition for childrenabstractRecently, Google launched YouTube Kids, a mobile application for children, that uses a speech recognizer built specifically for recognizing children’s speech. In this paper we present techniques we explored to build such a system. We describe the use of a neural network classifier to identify matched acoustic training data, filtering data for language modeling to reduce the chance of producing offensive results. We also compare long short-term memory (LSTM) recurrent networks to convolutional, LSTM, deep neural networks (CLDNN). We found that a CLDNN acoustic model outperforms an LSTM across a variety of different conditions, but does not specifically model child speech relatively better than adult. Overall, these findings allow us to build a successful, state-of-the-art large vocabulary speech recognizer for both children and adults. Hank Liao, Golan Pundak, Olivier Siohan, Melissa K. Carroll, Noah Coccaro, Qi-Ming Jiang, Tara N. Sainath, Andrew W. Senior, Françoise Beaufays, Michiel Bacchiani |
INTERSPEECH | 10 |
| 2014 | Context dependent state tying for speech recognition using deep neural network acoustic modelsabstractThis paper proposes an algorithm to design a tied-state inventory for a context dependent, neural network-based acoustic model for speech recognition. Rather than relying on a GMM/HMM system that operates on a different feature space and is of a different model family, the proposed algorithm optimizes state tying on the activation vectors of the neural network directly. Experiments show the viability of the proposed algorithm reducing the WER from 36.3% for a context independent system to 16.0% for a 15000 tied-state system. Michiel Bacchiani, David Rybach |
ICASSP | 1 |
| 2014 | Asynchronous stochastic optimization for sequence training of deep neural networksabstractThis paper explores asynchronous stochastic optimization for sequence training of deep neural networks. Sequence training requires more computation than frame-level training using pre-computed frame data. This leads to several complications for stochastic optimization, arising from significant asynchrony in model updates under massive parallelization, and limited data shuffling due to utterance-chunked processing. We analyze the impact of these two issues on the efficiency and performance of sequence training. In particular, we suggest a framework to formalize the reasoning about the asynchrony and present experimental results on both small and large scale Voice Search tasks to validate the effectiveness and efficiency of asynchronous stochastic optimization. Georg Heigold, Erik McDermott, Vincent Vanhoucke, Andrew W. Senior, Michiel Bacchiani |
ICASSP | 5 |
| 2014 | GMM-free DNN acoustic model trainingabstractWhile deep neural networks (DNNs) have become the dominant acoustic model (AM) for speech recognition systems, they are still dependent on Gaussian mixture models (GMMs) for alignments both for supervised training and for context dependent (CD) tree building. Here we explore bootstrapping DNN AM training without GMM AMs and show that CD trees can be built with DNN alignments which are better matched to the DNN model and its features. We show that these trees and alignments result in better models than from the GMM alignments and trees. By removing the GMM acoustic model altogether we simplify the system required to train a DNN from scratch. Andrew W. Senior, Georg Heigold, Michiel Bacchiani, Hank Liao |
ICASSP | 3 |
| 2014 | Asynchronous, online, GMM-free training of a context dependent acoustic model for speech recognitionabstractWe propose an algorithm that allows online training of a con-text dependent DNN model. It designs a state inventory based on DNN features and jointly optimizes the DNN parameters and alignment of the training data. The process allows flat starting a model from scratch and avoids any dependency on a GMM/HMM model to bootstrap the training process. A 15k state model trained with the proposed algorithm reduced the er-ror rate on a mobile speech task with 24 % compared to a system bootstrapped from a CI HMM/GMM and with 16 % compared to a system bootstrapped from a CD HMM/GMM system. Index Terms: Deep Neural Networks, online training 1. Michiel Bacchiani, Andrew W. Senior, Georg Heigold |
INTERSPEECH | 1 |
| 2014 | Robust speech recognition using temporal masking and thresholding algorithmabstractIn this paper, we present a new dereverberation algorithm called Temporal Masking and Thresholding (TMT) to enhance the temporal spectra of spectral features for robust speech recognition in reverberant environments. This algorithm is motivated by the precedence effect and temporal masking of human auditory perception. This work is an improvement of our previous dereverberation work called Suppression of Slowlyvarying components and the falling edge of the power envelope (SSF). The TMT algorithm uses a different mathematical model to characterize temporal masking and thresholding compared to the model that had been used to characterize the SSF algorithm. Specifically, the nonlinear highpass filtering used in the SSF algorithm has been replaced by a masking mechanism based on a combination of peak detection and dynamic thresholding. Speech recognition results show that the TMT algorithm provides superior recognition accuracy compared to other algorithms such as LTLSS, VTS, or SSF in reverberant environments. Chanwoo Kim 0001, Kean K. Chin, Michiel Bacchiani, Richard M. Stern |
INTERSPEECH | 3 |
| 2014 | Asynchronous stochastic optimization for sequence training of deep neural networks: towards big dataabstractPrevious work presented a proof of concept for sequence training of deep neural networks (DNNs) using asynchronous stochastic optimization, mainly focusing on a small-scale task. The approach offers the potential to leverage both the efficiency of stochastic gradient descent and the scalability of parallel computation. This study presents results for four different voice search tasks to confirm the effectiveness and efficiency of the proposed framework across different conditions: amount of data (from 60 hours to 20,000 hours), type of speech (read speech vs. spontaneous speech), quality of data (supervised vs. unsupervised data), and language. Significant gains over baselines (DNNs trained at the frame level) are found to hold across these conditions. The experimental results are analyzed, and additional practical details for the approach are provided. Furthermore, different sequence training criteria are compared. Erik McDermott, Georg Heigold, Pedro J. Moreno 0001, Andrew W. Senior, Michiel Bacchiani |
INTERSPEECH | 5 |
| 2013 | Rapid adaptation for mobile speech applicationsabstractWe investigate the use of iVector-based rapid adaptation for recognition in mobile speech applications. We show that on this task, the proposed approach has two merits over a linear-transform based approach. First it provides larger error reductions (11% vs. 6%) as it is better suited for the short utterances and varied recording conditions. Second it omits the need for speaker data pooling and/or clustering and the very large infrastructure complexity that accompanies that. Empirical results show that although the proposed utterance-based training algorithm leads to large data fragmentation, the resulting model re-estimation performs well. Our implementation within the MapReduce framework allows processing of the large statistics that this approach gives rise to when applied on a database of thousands of hours. Michiel Bacchiani |
ICASSP | 1 |
| 2013 | ivector-based acoustic data selectionabstractThis paper presents a data selection approach where spoken ut-terances are selected in a sequential fashion from a large out-of-domain data set to match the utterance distribution of an in-domain data set. We propose to represent each utterance by its iVector [1], a low dimensional vector indicating the coordi-nate of that utterance in a subspace acoustic model. We show that the distribution of iVectors can characterize a data set and enables distinguishing subsets of utterances from different do-mains. Last, we present experimental speech recognition results based on a system trained on a data set constructed by the pro-posed algorithm and a comparison with random data selection. Index Terms: speech recognition, data selection, acoustic mod-eling Olivier Siohan, Michiel Bacchiani |
INTERSPEECH | 2 |
| 2011 | Discriminative Features for Language IdentificationabstractIn this paper we investigate the use of discriminatively trained feature transforms to improve the accuracy of a MAP-SVM language recognition system. We train the feature transforms by alternatively solving an SVM optimization on MAP super-vectors estimated from transformed features, and performing a small step on the transforms in the direction of the antigradi-ent of the SVM objective function. We applied this method on the LRE2003 dataset, and obtained an 5:9 % relative reduction of pooled equal error rate. Index Terms — Language recognition, support vector machines, discriminative feature transforms. Christopher Alberti, Michiel Bacchiani |
INTERSPEECH | 2 |
| 2010 | Decision tree state clustering with word and syllable featuresabstractIn large vocabulary continuous speech recognition, decision trees are widely used to cluster triphone states. In addition to commonly used phonetically based questions, others have proposed additional questions such as phone position within word or syllable. This paper examines using the word or syllable context itself as a feature in the decision tree, providing an elegant way of introducing word- or syllable-specific models into the system. Positive results are reported on two state-of-the-art systems: voicemail transcription and a search by voice tasks across av ariety of acoustic model and training set sizes. Index Terms :d ecision tree state clustering, large vocabulary continuous speech recognition, tagged clustering. Hank Liao, Christopher Alberti, Michiel Bacchiani, Olivier Siohan |
INTERSPEECH | 3 |
| 2009 | An audio indexing system for election video materialabstractIn the 2008 presidential election race in the United States, the prospective candidates made extensive use of YouTube to post video material. We developed a scalable system that transcribes this material and makes the content searchable (by indexing the meta-data and transcripts of the videos) and allows the user to navigate through the video material based on content. The system is available as an iGoogle gadget1as well as a Labs product (labs.google.com/gaudi). Given the large exposure, special emphasis was put on the scalability and reliability of the system. This paper describes the design and implementation of this system. Christopher Alberti, Michiel Bacchiani, Ari Bezman, Ciprian Chelba, Anastassia Drofa, Hank Liao, Pedro J. Moreno 0001, Ted Power, Arnaud Sahuguet, Maria Shugrina, Olivier Siohan |
ICASSP | 2 |
| 2009 | Restoring punctuation and capitalization in transcribed speechabstractAdding punctuation and capitalization greatly improves the readability of automatic speech transcripts. We discuss an approach for performing both tasks in a single pass using a purely text-based n-gram language model. We study the effect on performance of varying the n-gram order (from n = 3 to n = 6) and the amount of training data (from 58 million to 55 billion tokens). Our results show that using larger training data sets consistently improves performance, while increasing the n-gram order does not help nearly as much. Agustín Gravano, Martin Jansche, Michiel Bacchiani |
ICASSP | 3 |
| 2008 | Deploying GOOG-411: Early lessons in data, measurement, and testingabstractWe describe our early experience building and optimizing GOOG-411, a fully automated, voice-enabled, business finder. We show how taking an iterative approach to system development allows us to optimize the various components of the system, thereby progressively improving user-facing metrics. We show the contributions of different data sources to recognition accuracy. For business listing language models, we see a nearly linear performance increase with the logarithm of the amount of training data. To date, we have improved our correct accept rate by 25% absolute, and increased our transfer rate by 35% absolute. Michiel Bacchiani, Françoise Beaufays, Johan Schalkwyk, Mike Schuster, Brian Strope |
ICASSP | 1 |
| 2008 | Confidence scores for acoustic model adaptationabstractThis paper focuses on confidence scores for use in acoustic model adaptation. Frame-based confidence estimates are used in linear transform (CMLLR and MLLR) and MAP adaptation. We show that adaptation approaches with a limited number of free parameters such as linear transform-based approaches are robust in the face of frame labeling errors whereas adaptation approaches with a large number of free parameters such as MAP are sensitive to the quality of the supervision and hence benefit most from use of confidences. Different approaches for using confidence information in adaptation are investigated. This analysis shows that a thresholding approach is effective in that it improves the frame labeling accuracy with little detrimental effect on frame recall. Experimental results show an absolute WER reduction of 2.1% over a CMLLR adapted system on a video transcription task. Christian Gollan, Michiel Bacchiani |
ICASSP | 2 |
| 2006 | MAP adaptation of stochastic grammars
Michiel Bacchiani, Michael Riley 0001, Brian Roark, Richard Sproat |
Comput. Speech Lang. | 1 |
| 2005 | Fast vocabulary-independent audio search using path-based graph indexingabstractClassical audio retrieval techniques consist in transcribing audio documents using a large vocabulary speech recognition system and indexing the resulting transcripts. However, queries that are not part of the recognizer’s vocabulary or have a large probability of getting misrecognized can significantly impair the performance of the retrieval system. Instead, we propose a fast vocabulary independent audio search approach that operates on phonetic lattices and is suitable for any query. However, indexing phonetic lattices so that any arbitrary phone sequence query can be processed efficiently is a challenge, as the choice of the indexing unit is unclear. We propose an inverted index structure on lattices that uses paths as indexing features. The approach is inspired by a general graph indexing method that defines an automatic procedure to select a small number of paths as indexing features, keeping the index size small while allowing fast retrieval of the lattices matching a given query. The effectiveness of the proposed approach is illustrated on broadcast news and Switchboard databases. Olivier Siohan, Michiel Bacchiani |
INTERSPEECH | 2 |
| 2004 | Meta-data conditional language modelingabstractAutomatic speech recognition (ASR) often occurs in circumstances in which knowledge external to the speech signal, or meta-data, is given. For example, a company receiving a call from a customer might have access to a database record of that customer. Conditioning the ASR models directly on this information to improve the transcription accuracy is hampered because, generally, the meta-data takes on many values and a training corpus has little data for each meta-data condition. The paper presents an algorithm to construct language models conditioned on such metadata. It uses tree-based clustering of the the training data to derive automatically meta-data projections, useful as language model conditioning contexts. The algorithm was tested on a multiple domain voice mail transcription task. We compare the performance of an adapted system aware of the domain shift to a system that only has meta-data to infer that fact. The meta-data used were the caller ID strings associated with the voice mail messages. The meta-data adapted system matched the performance of the system adapted using the domain knowledge explicitly. Michiel Bacchiani, Brian Roark |
ICASSP (1) | 1 |
| 2004 | Improved name recognition with meta-data dependent name networksabstractA transcription system that requires accurate general name transcription is faced with the problem of covering the large number of names it may encounter, Without any prior knowledge, this requires a large increase in the size and complexity of the system due to the expansion of the lexicon. Furthermore, this increase will adversely affect the system performance due to the increased confusability. Here we propose a method that uses meta-data, available at runtime to ensure better name coverage without significantly increasing the system complexity. We tested this approach on a voicemail transcription task and assumed meta-data to be available in the form of a caller ID string (as it would show up on a caller ID enabled telephone) and the name of the mailbox owner. Networks representing possible spoken realization of those names are generated at runtime and included in the network of the decoder. The decoder network is built at training time using a class-dependent language model, with caller and mailbox name instances modeled as class tokens. The class tokens are replaced at test time with the name networks built from the meta-data. The proposed algorithm showed a reduction in the error rate of name tokens of 22.1%. Sameer Maskey, Michiel Bacchiani, Brian Roark, Richard Sproat |
ICASSP (1) | 2 |
| 2003 | Unsupervised language model adaptationabstractThis paper investigates unsupervised language model adaptation, from ASR transcripts. N-gram counts from these transcripts can be used either to adapt an existing n-gram model or to build an n-gram model from scratch. Various experimental results are reported on a particular domain adaptation task, namely building a customer care application starting from a general voicemail transcription system. The experiments investigate the effectiveness of various adaptation strategies, including iterative adaptation and self-adaptation on the test data. They show an error rate reduction of 3.9% over the unadapted baseline performance, from 28% to 24.1%, using 17 hours of unsupervised adaptation material. This is 51% of the 7.7% adaptation gain obtained by supervised adaptation. Self-adaptation on the test data resulted in a 1.3% improvement over the baseline. Michiel Bacchiani, Brian Roark |
ICASSP (1) | 1 |
| 2003 | Supervised and unsupervised PCFG adaptation to novel domains
Brian Roark, Michiel Bacchiani |
HLT-NAACL | 2 |
| 2002 | SCANMail: a voicemail interface that makes speech browsable, readable and searchableabstractIncreasing amounts of public, corporate, and private speech data are now available on-line. These are limited in their usefulness, however, by the lack of tools to permit their browsing and search. The goal of our research is to provide tools to overcome the inherent difficulties of speech access, by supporting visual scanning, search, and information extraction. We describe a novel principle for the design of UIs to speech data: What You See Is Almost What You Hear (WYSIAWYH). In WYSIAWYH, automatic speech recognition (ASR) generates a transcript of the speech data. The transcript is then used as a visual analogue to that underlying data. A graphical user interface allows users to visually scan, read, annotate and search these transcripts. Users can also use the transcript to access and play specific regions of the underlying message. We first summarize previous studies of voicemail usage that motivated the WYSIAWYH principle, and describe a voicemail UI, SCANMail, that embodies WYSIAWYH. We report on a laboratory experiment and a two-month field trial evaluation. SCANMail outperformed a state of the art voicemail system on core voicemail tasks. This was attributable to SCANMail's support for visual scanning, search and information extraction. While the ASR transcripts contain errors, they nevertheless improve the efficiency of voicemail processing. Transcripts either provide enough information for users to extract key points or to navigate to important regions of the underlying speech, which they can then play directly Steve Whittaker 0001, Julia Hirschberg, Brian Amento, Litza A. Stark, Michiel Bacchiani, Philip L. Isenhour, Larry Stead, Gary Zamchick, Aaron E. Rosenberg |
CHI | 5 |
| 2002 | Combining maximum likelihood and maximum a posteriori estimation for detailed acoustic modeling of context dependency
Michiel Bacchiani |
INTERSPEECH | 1 |
| 2001 | Automatic transcription of voicemail at AT&TabstractReports on the automatic transcription accuracy of voicemail messages. It shows that vocal tract length normalization and adaptation using linear transformations, proven to improve accuracy on the Switchboard task, provide similar accuracy improvements on this task. Direct application of the normalization techniques is complicated by the fragmentation of the data. However, unsupervised clustering was found to be effective in ensuring robust estimation of normalization parameters. Variance adaptation resulted in larger accuracy improvements than adaptation of only mean parameters, probably due to a large variability in channel conditions. The use of semi-tied covariances provides additional gains over using speaker and channel normalization. The combined gain of using various compensation techniques improves the system word error rate from 34.9% for the baseline system to 28.7%. Michiel Bacchiani |
ICASSP | 1 |
| 2001 | SCANMail: browsing and searching speech data by contentabstractIncreasing amounts of public, corporate, and private audio data are available for use, but limited in usefulness by the lack of tools to permit their browsing and search. In this paper, we describe SCANMail, a system that employs automatic speech recognition, information retrieval, information extraction, and human computer interaction technology to permit users to browse and search their voicemail messages by content through a graphical user interface interface. The SCANMail client also provides note-taking capabilities as well as browsing and querying features. A CallerId server also proposes caller names from existing caller acoustic models and is trained from user feedback. An Email server sends the original message plus its transcription to a mailing address specified in the user's profile. 1. Julia Hirschberg, Michiel Bacchiani, Donald Hindle, Philip L. Isenhour, Aaron E. Rosenberg, Litza A. Stark, Larry Stead, Steve Whittaker 0001, Gary Zamchick |
INTERSPEECH | 2 |
| 2001 | Caller identification for the SCANMail voicemail browserabstractSCANMail is a prototype system developed at AT&T Labs for the purpose of providing useful tools for managing and searching through voicemail messages. Content is extracted from voicemail messages using various speech and text processing tools. One such content category is the identity of the message caller. This paper describes CallerID, the server tool attached to SCANMail for the purpose of providing caller labels for voicemail messages. CallerID make use of text independent speaker recognition techniques. Two kinds of requests are handled by the CallerID server. A request triggered by the arrival of a new voicemail message results in the processing of the message to score it against the models of callers assigned to the user (recipient) in order to propose the identity of the caller. A second request is initiated by a user who provides a caller label for a message he/she has reviewed. CallerID processes the message and uses it to train or adapt a speaker model for the caller whose label is provided. The paper describes in detail the CallerID functions and provides some results of performance evaluations of the caller identification capability. Aaron E. Rosenberg, Julia Hirschberg, Michiel Bacchiani, Sarangarajan Parthasarathy, Philip L. Isenhour, Larry Stead |
INTERSPEECH | 3 |
| 2000 | Using maximum likelihood linear regression for segment clustering and speaker identificationabstractMany adaptation scenarios rely on clustering of either the test or training data. Although consistency between the clustering and adaptation objective functions is desired, most previous approaches have not implemented such consistency. This paper shows that the statistics used in Maximum Likelihood Linear Regression (MLLR) adaptation are sufficient to cluster data with a consistent Maximum Likelihood (ML) criterion. In addition, as the algorithm uses the same statistics for both adaptation and clustering, it is computationally efficient. Clustering experiments contrasting the performance of this algorithm with the widely used text independent Gaussian mixture model approach show increased adaptation likelihoods and consistency of within-cluster speaker identity. In a speaker identification experiment the adaptation-based scoring showed improved classification performance compared to the mixture model-based scoring. Michiel Bacchiani |
INTERSPEECH | 1 |
| 1999 | Joint lexicon, acoustic unit inventory and model designabstractAlthough most parameters in a speech recognition system are estimated from data by the use of an objective function, the unit inventory and lexicon are generally hand crafted and therefore unlikely to be optimal. This paper proposes a joint solution to the related problems of learning a unit inventory and corresponding lexicon from data. On a speaker-independent read speech task with a 1k vocabulary, the proposed algorithm outperforms phone-based systems at both high and low complexities. Obwohl die meisten Parameter eines Spracherkennungssystems aus Daten geschätzt werden, ist die Wahl der akustischen Grundeinheiten und des Lexikons normalerweise nicht automatisch und deshalb wahrscheinlich nicht optimal. Dieser Artikel stellt einen kombinierten Ansatz für die Lösung dieser verwandten Probleme dar – das Lernen von akustischen Grundeinheiten und des zugehörigen Lexikons aus Daten. Experimente mit sprecher-unabhängigen gelesenen Sprachdaten mit einem Vokabular von 1000 Wörtern zeigen, daß der vorgestellte Ansatz besser ist als ein System niedriger oder höherer Komplexität, das auf Phonemen basiert ist. Bien que la plupart des paramètres dans un système de reconnaissance de la parole soient estimés à partie des données en utilisant une fonction objective, l'inventaire des unités acoustiques et le lexique sont généralement créés à la main, et donc susceptibles de ne pas être optimeux. Cette étude propose une solution conjointe aux problèmes interdépendants que sont l'apprentissage à partir des données d'un inventaire des unités acoustiques et du lexique correspondant. Nous avons testé l'algorithme proposé sur des échantillons lus, en reconnaissance indépendantes du locuteur avec un vocabulaire de 1k: il surpassé les systèmes phonétiques en faible ou forte complexité. Michiel Bacchiani, Mari Ostendorf |
Speech Commun. | 1 |
| 1998 | Using automatically-derived acoustic sub-word units in large vocabulary speech recognition
Michiel Bacchiani, Mari Ostendorf |
ICSLP | 1 |
| 1996 | Design of a speech recognition system based on acoustically derived segmental unitsabstractThe design of a speech recognition system based on acoustically-derived, segmental units can be divided in three steps: unit design, lexicon building and pronunciation modeling. We formulate an iterative unit design procedure which consistently uses a maximum likelihood (ML) objective in successive application of resegmentation and model re-estimation. The lexicon building allows multi-word entries in the lexicon but restricts the number of these entries in order to avoid a too costly search. Selected multi-word lexical entries are those with high frequency (such as function words) and those which consistently exhibit cross-word phone assimilation. The stochastic pronunciation model represents the likelihood of a particular acoustic segment sequence given the phonetic baseform of a lexical item, where the sequence of baseform phones are treated as a Markov state sequence and each state can emit multiple segments. Michiel Bacchiani, Mari Ostendorf, Yoshinori Sagisaka, Kuldip K. Paliwal |
ICASSP | 1 |
| 1996 | Speech recognition based on acoustically derived segment units
Toshiaki Fukada, Michiel Bacchiani, Kuldip K. Paliwal, Yoshinori Sagisaka |
ICSLP | 2 |
| 1995 | Minimum classification error training algorithm for feature extractor and pattern classifier in speech recognition
Kuldip K. Paliwal, Michiel Bacchiani, Yoshinori Sagisaka |
EUROSPEECH | 2 |
| 1994 | Optimization of time-frequency masking filters using the minimum classification error criterionabstractThe dynamic cepstrum parameter representing a masked spectrum performed extremely well in continuous speech recognition. This paper proposes a new algorithm for optimizing the dynamic cepstrum lifter array. The masking filter is represented by a set of Gaussian-shaped lifters. The standard deviation and the gain of the Gaussians are trained in order to improve the performance of the time-frequency filter. Parameterizing the lifter shape provides robustness against unknown speech samples. Because of the parameterized lifter's small degree of freedom, it can avoid over-learning. The gradient descent optimizing algorithm is formulated for both a neural network classifier and an HMM classifier. The optimized dynamic cepstrum successfully improved the speech recognition performance for the speech spoken even in a different speaking style.> Michiel Bacchiani, Kiyoaki Aikawa |
ICASSP (2) | 1 |