VLDB 2026 Research / reviewers in the wild / expert
Michael L. Seltzer
dblp:69/172
· DBLP profile ↗
91ranked-venue papers
23as first author
21since 2021 · last 2024
0000-0003-3474-2451ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 84 · 20 first-author · 21 since 2021Artificial intelligence and machine learning · 45 · 12 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | End-to-End Speech Recognition Contextualization with Large Language ModelsabstractIn recent years, Large Language Models (LLMs) have garnered significant attention from the research community due to their exceptional performance and generalization capabilities. In this paper, we introduce a novel method for contextualizing speech recognition models incorporating LLMs. Our approach casts speech recognition as a mixed-modal language modeling task based on a pretrained LLM. We use audio features, along with optional text tokens for context, to train the system to complete transcriptions in a decoder-only fashion. As a result, the system implicitly learns how to leverage unstructured contextual information during training. Our empirical results demonstrate a significant improvement in performance, with a 6% WER reduction when additional textual context is provided. Moreover, we find that our method performs competitively, improving by 7.5% WER overall and 17% WER on rare words, compared to a baseline contextualized RNN-T system that has been trained on a speech dataset more than twenty-five times larger. Overall, we demonstrate that by adding only a handful of trainable parameters via adapters, we can unlock the contextualized speech recognition capability of the pretrained LLM while maintaining the same text-only input functionality. Egor Lakomkin, Chunyang Wu, Yassir Fathullah, Ozlem Kalinli, Michael L. Seltzer, Christian Fügen |
ICASSP | 5 |
| 2024 | Towards measuring fairness in speech recognition: Fair-Speech datasetabstractThe current public datasets for speech recognition (ASR) tend not to focus specifically on the fairness aspect, such as performance across different demographic groups.This paper introduces a novel dataset, Fair-Speech, a publicly released corpus to help researchers evaluate their ASR models for accuracy across a diverse set of self-reported demographic information, such as age, gender, ethnicity, geographic variation and whether the participants consider themselves native English speakers.Our dataset includes approximately 26.5K utterances in recorded speech by 593 people in the United States, who were paid to record and submit audios of themselves saying voice commands.We also provide ASR baselines, including on models trained on transcribed and untranscribed social media videos and open source models. Irina-Elena Veliche, Zhuangqun Huang, Vineeth Ayyat Kochaniyan, Fuchun Peng, Ozlem Kalinli, Michael L. Seltzer |
INTERSPEECH | 6 |
| 2023 | Factorized Blank Thresholding for Improved Runtime Efficiency of Neural TransducersabstractWe show how factoring the RNN-T’s output distribution can significantly reduce the computation cost and power consumption for on-device ASR inference with no loss in accuracy. With the rise in popularity of neural-transducer type models like the RNN-T for on-device ASR, optimizing RNN-T’s runtime efficiency is of great interest. While previous work has primarily focused on the optimization of RNN-T’s acoustic encoder and predictor, this paper focuses the attention on the joiner. We show that despite being only a small part of RNN-T, the joiner has a large impact on the overall model’s runtime efficiency. We propose to utilize HAT-style joiner factorization for the purpose of skipping the more expensive non-blank computation when the blank probability exceeds a certain threshold. Since the blank probability can be computed very efficiently and the RNN-T output is dominated by blanks, our proposed method leads to a 26-30% decoding speed-up and 43-53% reduction in on-device power consumption, all the while incurring no accuracy degradation and being relatively simple to implement. Frank Seide, Yang Li 0183, Kjell Schubert, Ozlem Kalinli, Michael L. Seltzer |
ICASSP | 7 |
| 2023 | Improving fast-slow Encoder based Transducer with Streaming DeliberationabstractThis paper introduces a fast-slow encoder based transducer with streaming deliberation for end-to-end automatic speech recognition. We aim to improve the recognition accuracy of the fast-slow encoder based transducer while keeping its latency low by integrating a streaming deliberation model. Specifically, the deliberation model leverages partial hypotheses from the streaming fast encoder and implicitly learns to correct recognition errors. We modify the parallel beam search algorithm for fast-slow encoder based transducer to be efficient and compatible with the deliberation model. In addition, the deliberation model is designed to process streaming data. To further improve the deliberation performance, a simple text augmentation approach is explored. We also compare LSTM and Conformer models for encoding partial hypotheses. Experiments on Librispeech and in-house data show relative WER reductions (WERRs) from 3% to 5% with a slight increase in model size and negligible extra token emission latency compared with fast-slow encoder based transducer. The deliberation model also reduces rare WERs by 3-7% on large-scale in-house data. Compared with vanilla neural transducers, the proposed deliberation model together with fast-slow encoder based transducer obtains relative 10-11% WERRs on Librispeech and around relative 6% WERR on in-house data with smaller emission delays. Ke Li 0023, Jay Mahadeokar, Jinxi Guo, Yangyang Shi, Gil Keren, Ozlem Kalinli, Michael L. Seltzer |
ICASSP | 7 |
| 2023 | Massively Multilingual ASR on 70 Languages: Tokenization, Architecture, and Generalization CapabilitiesabstractEnd-to-end multilingual ASR has become more appealing because of several reasons such as simplifying the training and deployment process and positive performance transfer from high-resource to low-resource languages. However, scaling up the number of languages, total hours, and number of unique tokens is not a trivial task. This paper explores large-scale multilingual ASR models on 70 languages. We inspect two architectures: (1) Shared embedding and output and (2) Multiple embedding and output model. In the shared model experiments, we show the importance of tokenization strategy across different languages. Later, we use our optimal tokenization strategy to train multiple embedding and output model to further improve our result. Our multilingual ASR achieves 13.9%-15.6% average WER relative improvement compared to monolingual models. We show that our multilingual ASR generalizes well on an unseen dataset and domain, achieving 9.5% and 7.5% WER on Multilingual Librispeech (MLS) with zero-shot and finetuning, respectively. Andros Tjandra, Nayan Singhal, Ozlem Kalinli, Abdel-rahman Mohamed, Michael L. Seltzer |
ICASSP | 7 |
| 2023 | Modality Confidence Aware Training for Robust End-to-End Spoken Language Understanding
Suyoun Kim, Akshat Shrivastava, Ju Lin, Ozlem Kalinli, Michael L. Seltzer |
INTERSPEECH | 6 |
| 2022 | Neural-FST Class Language Model for End-to-End Speech RecognitionabstractWe propose Neural-FST Class Language Model (NFCLM) for end-to-end speech recognition, a novel method that combines neural network language models (NNLMs) and finite state transducers (FSTs) in a mathematically consistent framework. Our method utilizes a background NNLM which models generic background text together with a collection of domain-specific entities modeled as individual FSTs. Each output token is generated by a mixture of these components; the mixture weights are estimated with a separately trained neural decider. We show that NFCLM significantly outperforms NNLM by 15.8% relative in terms of Word Error Rate. NFCLM achieves similar performance as traditional NNLM and FST shallow fusion while being less prone to overbiasing and 12 times more compact, making it more suitable for on-device usage. Antoine Bruguier, Rohit Prabhavalkar, Dangna Li, Zhe Liu 0011, Eun Chang, Fuchun Peng, Ozlem Kalinli, Michael L. Seltzer |
ICASSP | 10 |
| 2022 | Evaluating User Perception of Speech Recognition System Quality with Semantic Distance MetricabstractMeasuring automatic speech recognition (ASR) system quality is critical for creating user-satisfying voice-driven applications. Word Error Rate (WER) has been traditionally used to evaluate ASR system quality; however, it sometimes correlates poorly with user perception/judgement of transcription quality. This is because WER weighs every word equally and does not consider semantic correctness which has a higher impact on user perception. In this work, we propose evaluating ASR output hypotheses quality with SemDist that can measure semantic correctness by using the distance between the semantic vectors of the reference and hypothesis extracted from a pre-trained language model. Our experimental results of 71K and 36K user annotated ASR output quality show that SemDist achieves higher correlation with user perception than WER. We also show that SemDist has higher correlation with downstream Natural Language Understanding (NLU) tasks than WER. Suyoun Kim, Weiyi Zheng, Tarun Singh, Abhinav Arora, Xiaoyu Zhai, Christian Fügen, Ozlem Kalinli, Michael L. Seltzer |
INTERSPEECH | 9 |
| 2022 | Deliberation Model for On-Device Spoken Language UnderstandingabstractWe propose a novel deliberation-based approach to end-to-end (E2E) spoken language understanding (SLU), where a streaming automatic speech recognition (ASR) model produces the first-pass hypothesis and a second-pass natural language understanding (NLU) component generates the semantic parse by conditioning on both ASR's text and audio embeddings.By formulating E2E SLU as a generalized decoder, our system is able to support complex compositional semantic structures.Furthermore, the sharing of parameters between ASR and NLU makes the system especially suitable for resource-constrained (on-device) environments; our proposed approach consistently outperforms strong pipeline NLU baselines by 0.60% to 0.65% on the spoken version of the TOPv2 dataset (STOP).We demonstrate that the fusion of text and audio features, coupled with the system's ability to rewrite the first-pass hypothesis, makes our approach more robust to ASR errors.Finally, we show that our approach can significantly reduce the degradation when moving from natural speech to synthetic speech training, but more work is required to make text-to-speech (TTS) a viable solution for scaling up E2E SLU. Akshat Shrivastava, Paden Tomasello, Suyoun Kim, Aleksandr Livshits, Ozlem Kalinli, Michael L. Seltzer |
INTERSPEECH | 7 |
| 2022 | Streaming parallel transducer beam search with fast slow cascaded encoders
Jay Mahadeokar, Yangyang Shi, Ke Li 0023, Jiedan Zhu, Vikas Chandra, Ozlem Kalinli, Michael L. Seltzer |
INTERSPEECH | 8 |
| 2021 | Improved Neural Language Model Fusion for Streaming Recurrent Neural Network TransducerabstractRecurrent Neural Network Transducer (RNN-T), like most end-to-end speech recognition model architectures, has an implicit neural network language model (NNLM) and cannot easily leverage unpaired text data during training. Previous work has proposed various fusion methods to incorporate external NNLMs into end-to-end ASR to address this weakness. In this paper, we propose extensions to these techniques that allow RNN-T to exploit external NNLMs during both training and inference time, resulting in 13-18% relative Word Error Rate improvement on Librispeech compared to strong baselines. Furthermore, our methods do not incur extra algorithmic latency and allow for flexible plug- and-play of different NNLMs without re-training. We also share in-depth analysis to better understand the benefits of the different NNLM fusion methods. Our work provides a reliable technique for leveraging unpaired text data to significantly improve RNN-T while keeping the system streamable, flexible, and lightweight. Suyoun Kim, Yuan Shangguan, Jay Mahadeokar, Antoine Bruguier, Christian Fügen, Michael L. Seltzer |
ICASSP | 6 |
| 2021 | Memory-Efficient Speech Recognition on Smart DevicesabstractRecurrent transducer models have emerged as a promising solution for speech recognition on the current and next generation smart devices. The transducer models provide competitive accuracy within a reasonable memory footprint alleviating the memory capacity constraints in these devices. However, these models access parameters from off-chip memory for every input time step which adversely effects device battery life and limits their usability on low-power devices.We address transducer model’s memory access concerns by optimizing their model architecture and designing novel recurrent cell designs. We demonstrate that i) model’s energy cost is dominated by accessing model weights from off-chip memory, ii) transducer model architecture is pivotal in determining the number of accesses to off-chip memory and just model size is not a good proxy, iii) our transducer model optimizations and novel recurrent cell reduces off-chip memory accesses by 4.5× and model size by 2× with minimal accuracy impact. Ganesh Venkatesh, Alagappan Valliappan, Jay Mahadeokar, Yuan Shangguan, Christian Fügen, Michael L. Seltzer, Vikas Chandra |
ICASSP | 6 |
| 2021 | Semantic Distance: A New Metric for ASR Performance Analysis Towards Spoken Language UnderstandingabstractWord Error Rate (WER) has been the predominant metric used to evaluate the performance of automatic speech recognition (ASR) systems. However, WER is sometimes not a good indicator for downstream Natural Language Understanding (NLU) tasks, such as intent recognition, slot filling, and semantic parsing in task-oriented dialog systems. This is because WER takes into consideration only literal correctness instead of semantic correctness, the latter of which is typically more important for these downstream tasks. In this study, we propose a novel Semantic Distance (SemDist) measure as an alternative evaluation metric for ASR systems to address this issue. We define SemDist as the distance between a reference and hypothesis pair in a sentence-level embedding space. To represent the reference and hypothesis as a sentence embedding, we exploit RoBERTa, a state-of-the-art pre-trained deep contextualized language model based on the transformer architecture. We demonstrate the effectiveness of our proposed metric on various downstream tasks, including intent recognition, semantic parsing, and named entity recognition. Suyoun Kim, Abhinav Arora, Ching-Feng Yeh, Christian Fügen, Ozlem Kalinli, Michael L. Seltzer |
Interspeech | 7 |
| 2021 | Contextualized Streaming End-to-End Speech Recognition with Trie-Based Deep Biasing and Shallow FusionabstractHow to leverage dynamic contextual information in end-toend speech recognition has remained an active research area.Previous solutions to this problem were either designed for specialized use cases that did not generalize well to open-domain scenarios, did not scale to large biasing lists, or underperformed on rare long-tail words.We address these limitations by proposing a novel solution that combines shallow fusion, trie-based deep biasing, and neural network language model contextualization.These techniques result in significant 19.5% relative Word Error Rate improvement over existing contextual biasing approaches and 5.4%-9.3%improvement compared to a strong hybrid baseline on both open-domain and constrained contextualization tasks, where the targets consist of mostly rare long-tail words.Our final system remains lightweight and modular, allowing for quick modification without model re-training. Mahaveer Jain, Gil Keren, Suyoun Kim, Yangyang Shi, Jay Mahadeokar, Julian Chan, Yuan Shangguan, Christian Fügen, Ozlem Kalinli, Yatharth Saraf, Michael L. Seltzer |
Interspeech | 12 |
| 2021 | Flexi-Transducer: Optimizing Latency, Accuracy and Compute for Multi-Domain On-Device Scenarios
Jay Mahadeokar, Yangyang Shi, Yuan Shangguan, Chunyang Wu, Alex Xiao, Ozlem Kalinli, Christian Fügen, Michael L. Seltzer |
Interspeech | 10 |
| 2021 | Collaborative Training of Acoustic Encoders for Speech RecognitionabstractOn-device speech recognition requires training models of different sizes for deploying on devices with various computational budgets. When building such different models, we can benefit from training them jointly to take advantage of the knowledge shared between them. Joint training is also efficient since it reduces the redundancy in the training procedure's data handling operations. We propose a method for collaboratively training acoustic encoders of different sizes for speech recognition. We use a sequence transducer setup where different acoustic encoders share a common predictor and joiner modules. The acoustic encoders are also trained using co-distillation through an auxiliary task for frame level chenone prediction, along with the transducer loss. We perform experiments using the LibriSpeech corpus and demonstrate that the collaboratively trained acoustic encoders can provide up to a 11% relative improvement in the word error rate on both the test partitions. Varun Nagaraja, Yangyang Shi, Ganesh Venkatesh, Ozlem Kalinli, Michael L. Seltzer, Vikas Chandra |
Interspeech | 5 |
| 2021 | Dissecting User-Perceived Latency of On-Device E2E Speech RecognitionabstractAs speech-enabled devices such as smartphones and smart speakers become increasingly ubiquitous, there is growing interest in building automatic speech recognition (ASR) systems that can run directly on-device; end-to-end (E2E) speech recognition models such as recurrent neural network transducers and their variants have recently emerged as prime candidates for this task.Apart from being accurate and compact, such systems need to decode speech with low user-perceived latency (UPL), producing words as soon as they are spoken.This work examines the impact of various techniques -model architectures, training criteria, decoding hyperparameters, and endpointer parameters -on UPL.Our analyses suggest that measures of model size (parameters, input chunk sizes), or measures of computation (e.g., FLOPS, RTF) that reflect the model's ability to process input frames are not always strongly correlated with observed UPL.Thus, conventional algorithmic latency measurements might be inadequate in accurately capturing latency observed when models are deployed on embedded devices.Instead, we find that factors affecting token emission latency, and endpointing behavior have a larger impact on UPL.We achieve the best trade-off between latency and word error rate when performing ASR jointly with endpointing, while utilizing the recently proposed alignment regularization mechanism. Yuan Shangguan, Rohit Prabhavalkar, Jay Mahadeokar, Yangyang Shi, Jiatong Zhou, Chunyang Wu, Ozlem Kalinli, Christian Fügen, Michael L. Seltzer |
Interspeech | 11 |
| 2021 | Dynamic Encoder Transducer: A Flexible Solution for Trading Off Accuracy for LatencyabstractWe propose a dynamic encoder transducer (DET) for on-device speech recognition. One DET model scales to multiple devices with different computation capacities without retraining or finetuning. To trading off accuracy and latency, DET assigns different encoders to decode different parts of an utterance. We apply and compare the layer dropout and the collaborative learning for DET training. The layer dropout method that randomly drops out encoder layers in the training phase, can do on-demand layer dropout in decoding. Collaborative learning jointly trains multiple encoders with different depths in one single model. Experiment results on Librispeech and in-house data show that DET provides a flexible accuracy and latency trade-off. Results on Librispeech show that the full-size encoder in DET relatively reduces the word error rate of the same size baseline by over 8%. The lightweight encoder in DET trained with collaborative learning reduces the model size by 25% but still gets similar WER as the full-size baseline. DET gets similar accuracy as a baseline model with better latency on a large in-house data set by assigning a lightweight encoder for the beginning part of one utterance and a full-size encoder for the rest. Yangyang Shi, Varun Nagaraja, Chunyang Wu, Jay Mahadeokar, Rohit Prabhavalkar, Alex Xiao, Ching-Feng Yeh, Julian Chan, Christian Fügen, Ozlem Kalinli, Michael L. Seltzer |
Interspeech | 12 |
| 2021 | Deep Shallow Fusion for RNN-T PersonalizationabstractEnd-to-end models in general, and Recurrent Neural Network Transducer (RNN-T) in particular, have gained significant traction in the automatic speech recognition community in the last few years due to their simplicity, compactness, and excellent performance on generic transcription tasks. However, these models are more challenging to personalize compared to traditional hybrid systems due to the lack of external language models and difficulties in recognizing rare long-tail words, specifically entity names. In this work, we present novel techniques to improve RNN-T's ability to model rare WordPieces, infuse extra information into the encoder, enable the use of alternative graphemic pronunciations, and perform deep fusion with personalized language models for more robust biasing. We show that these combined techniques result in 15.4%-34.5% relative Word Error Rate improvement compared to a strong RNN-T baseline which uses shallow fusion and text-to-speech augmentation. Our work helps push the boundary of RNN-T personalization and close the gap with hybrid systems on use cases where biasing and entity recognition are crucial. Gil Keren, Julian Chan, Jay Mahadeokar, Christian Fügen, Michael L. Seltzer |
SLT | 6 |
| 2021 | Alignment Restricted Streaming Recurrent Neural Network TransducerabstractThere is a growing interest in the speech community in developing Recurrent Neural Network Transducer (RNN-T) models for automatic speech recognition (ASR) applications. RNN-T is trained with a loss function that does not enforce temporal alignment of the training transcripts and audio. As a result, RNN-T models built with uni-directional long short term memory (LSTM) encoders tend to wait for longer spans of input audio, before streaming already decoded ASR tokens. In this work, we propose a modification to the RNN-T loss function and develop Alignment Restricted RNN-T (Ar-RNN-T) models, which utilize audio-text alignment in-formation to guide the loss computation. We compare the proposed method with existing works, such as monotonic RNN-T, on LibriSpeech and in-house datasets. We show that the Ar-RNN-T loss provides a refined control to navigate the trade-offs between the token emission delays and the Word Error Rate (WER). The Ar-RNN-T models also improve downstream applications such as the ASR End-pointing by guaranteeing token emissions within any given range of latency. Moreover, the Ar-RNN-T loss allows for bigger batch sizes and 4 times higher throughput for our LSTM model architecture, enabling faster training and convergence on GPUs. Jay Mahadeokar, Yuan Shangguan, Gil Keren, Thong Le, Ching-Feng Yeh, Christian Fügen, Michael L. Seltzer |
SLT | 9 |
| 2021 | Streaming Attention-Based Models with Augmented Memory for End-To-End Speech RecognitionabstractAttention-based models have been gaining popularity recently for their strong performance demonstrated in fields such as machine translation [1] and automatic speech recognition [2]. One major challenge of attention-based models is the need of access to the full sequence and the quadratically growing computational cost concerning the sequence length. These characteristics pose challenges, especially for low-latency scenarios, where the system is often required to be streaming. In this paper, we build a compact and streaming speech recognition system on top of the end-to-end neural transducer architecture [3] with attention-based modules augmented with convolution [2]. The proposed system equips the end-to-end models with the streaming capability and reduces the large footprint from the streaming attention-based model using augmented memory [4], [5]. On the LibriSpeech [6] dataset, our proposed system achieves word error rates 2.7% on test-clean and 5.8% on test-other, to our best knowledge the lowest among streaming approaches reported so far. Ching-Feng Yeh, Yongqiang Wang 0005, Yangyang Shi, Chunyang Wu, Frank Zhang 0001, Julian Chan, Michael L. Seltzer |
SLT | 7 |
| 2020 | Aipnet: Generative Adversarial Pre-Training of Accent-Invariant Networks for End-To-End Speech RecognitionabstractAs one of the major sources in speech variability, accents have posed a grand challenge to the robustness of speech recognition systems. In this paper, our goal is to build a unified end-to-end speech recognition system that generalizes well across accents. For this purpose, we propose a novel pre-training framework AIPNetbased on generative adversarial nets (GAN) for accent-invariant representation learning: Accent Invariant Pre-training Networks. We pre-train AIPNetto disentangle accent-invariant and accent-specific characteristics from acoustic features through adversarial training on accented data for which transcriptions are not necessarily available. We further fine-tune AIPNetby connecting the accent-invariant module with an attention-based encoder-decoder model for multi-accent speech recognition. In the experiments, our approach is compared against four baselines including both accent-dependent and accent-independent models. Experimental results on 9 English accents show that the proposed approach outperforms all the baselines by 2.3 ~ 4.5% relative reduction on average WER when transcriptions are available in all accents and by 1.6 ~ 6.1% relative reduction when transcriptions are only available in US accent. Ching-Feng Yeh, Mahaveer Jain, Michael L. Seltzer |
ICASSP | 5 |
| 2020 | G2G: TTS-Driven Pronunciation Learning for Graphemic Hybrid ASRabstractGrapheme-based acoustic modeling has recently been shown to outperform phoneme-based approaches in both hybrid and end-to-end automatic speech recognition (ASR), even on non-phonemic languages like English. However, graphemic ASR still has problems with low-frequency words that do not follow the standard spelling conventions seen in training, such as entity names. In this work, we present a novel method to train a statistical grapheme-to-grapheme (G2G) model on text-to-speech data that can rewrite an arbitrary character sequence into more phonetically consistent forms. We show that using G2G to provide alternative pronunciations during decoding reduces Word Error Rate by 3% to 11% relative over a strong graphemic baseline and bridges the gap on rare name recognition with an equivalent phonetic setup. Unlike many previously proposed methods, our method does not require any change to the acoustic model training procedure. This work reaffirms the efficacy of grapheme-based modeling and shows that specialized linguistic knowledge, when available, can be leveraged to improve graphemic ASR. Thilo Köhler, Christian Fügen, Michael L. Seltzer |
ICASSP | 4 |
| 2020 | Transformer-Based Acoustic Modeling for Hybrid Speech RecognitionabstractWe propose and evaluate transformer-based acoustic models (AMs) for hybrid speech recognition. Several modeling choices are discussed in this work, including various positional embedding methods and an iterated loss to enable training deep transformers. We also present a preliminary study of using limited right context in transformer models, which makes it possible for streaming applications. We demonstrate that on the widely used Librispeech benchmark, our transformer-based AM outperforms the best published hybrid result by 19% to 26% relative when the standard n-gram language model (LM) is used. Combined with neural network LM for rescoring, our proposed approach achieves state-of-the-art results on Librispeech. Our findings are also confirmed on a much larger internal dataset. Yongqiang Wang 0005, Abdel-rahman Mohamed, Chunxi Liu, Alex Xiao, Jay Mahadeokar, Hongzhao Huang, Andros Tjandra, Xiaohui Zhang 0007, Frank Zhang 0001, Christian Fügen, Geoffrey Zweig, Michael L. Seltzer |
ICASSP | 13 |
| 2020 | Weak-Attention Suppression for Transformer Based Speech RecognitionabstractTransformers, originally proposed for natural language processing (NLP) tasks, have recently achieved great success in automatic speech recognition (ASR). However, adjacent acoustic units (i.e., frames) are highly correlated, and long-distance dependencies between them are weak, unlike text units. It suggests that ASR will likely benefit from sparse and localized attention. In this paper, we propose Weak-Attention Suppression (WAS), a method that dynamically induces sparsity in attention probabilities. We demonstrate that WAS leads to consistent Word Error Rate (WER) improvement over strong transformer baselines. On the widely used LibriSpeech benchmark, our proposed method reduced WER by 10%$ on test-clean and 5% on test-other for streamable transformers, resulting in a new state-of-the-art among streaming models. Further analysis shows that WAS learns to suppress attention of non-critical and redundant continuous acoustic frames, and is more likely to suppress past frames rather than future ones. It indicates the importance of lookahead in attention-based ASR models. Yangyang Shi, Yongqiang Wang 0005, Chunyang Wu, Christian Fügen, Frank Zhang 0001, Ching-Feng Yeh, Michael L. Seltzer |
INTERSPEECH | 8 |
| 2019 | From Senones to Chenones: Tied Context-Dependent Graphemes for Hybrid Speech RecognitionabstractThere is an implicit assumption that traditional hybrid approaches for automatic speech recognition (ASR) cannot directly model graphemes and need to rely on phonetic lexicons to get competitive performance, especially on English which has poor grapheme-phoneme correspondence. In this work, we show for the first time that, on English, hybrid ASR systems can in fact model graphemes effectively by leveraging tied context-dependent graphemes, i.e., chenones. Our chenone-based systems significantly outperform equivalent senone baselines by 4.5% to 11.1% relative on three different English datasets. Our results on Librispeech are state-of-the-art compared to other hybrid approaches and competitive with previously published end-to-end numbers. Further analysis shows that chenones can better utilize powerful acoustic models and large training data, and require context- and position-dependent modeling to work well. Chenone-based systems also outperform senone baselines on proper noun and rare word recognition, an area where the latter is traditionally thought to have an advantage. Our work provides an alternative for end-to-end ASR and establishes that hybrid systems can be improved by dropping the reliance on phonetic knowledge. Xiaohui Zhang 0007, Weiyi Zheng, Christian Fügen, Geoffrey Zweig, Michael L. Seltzer |
ASRU | 6 |
| 2019 | End-to-end Contextual Speech Recognition Using Class Language Models and a Token Passing DecoderabstractEnd-to-end modeling (E2E) of automatic speech recognition (ASR) blends all the components of a traditional speech recognition system into a single, unified model. Although it simplifies the ASR systems, the unified model is hard to adapt when training and testing data mismatches. In this work, we focus on contextual speech recognition, which is particularly challenging for E2E models because contextual information is only available in inference time. To improve the performance in the presence of contextual information during training, we propose to use class-based language models (CLM) that can populate context-dependent information during inference. To enable this approach to scale to a large number of class members and minimize search errors, we propose a token passing algorithm with an efficient token recombination for E2E systems. We evaluate the proposed system on general and contextual ASR tasks, and achieve relative 62% Word Error Rate (WER) reduction for the contextual ASR task without hurting recognition performance for the general ASR task. We also show that the proposed method performs well without modification of the decoding hyper-parameters across tasks, making it a desirable solution for E2E ASR. Zhehuai Chen, Mahaveer Jain, Yongqiang Wang 0005, Michael L. Seltzer, Christian Fügen |
ICASSP | 4 |
| 2019 | Joint Grapheme and Phoneme Embeddings for Contextual End-to-End ASR
Zhehuai Chen, Mahaveer Jain, Yongqiang Wang 0005, Michael L. Seltzer, Christian Fügen |
INTERSPEECH | 4 |
| 2018 | Efficient Integration of Fixed Beamformers and Speech Separation Networks for Multi-Channel Far-Field Speech SeparationabstractSpeech separation research has significantly progressed in recent years thanks to the rapid advances in deep learning technology. However the performance of recently proposed single-channel neural network-based speech separation methods is still limited especially in reverberant environments. To push the performance limit, we recently developed a method of integrating beamforming and single-channel speech separation approaches. This paper proposes a novel architecture that integrates multi -channel beamforming and speech separation in a much more efficient way than our previous method. The proposed architecture comprises a set of fixed beamformers, a beam prediction network, and a speech separation network based on permutation invariant training (PIT). The beam prediction network takes in the beamformed audio signals and estimates the best beam for each speaker constituting the input mixture. Two variants of PIT-based speech separation networks are proposed. Our approach is evaluated on reverberant speech mixtures under three different mixing conditions, covering cases where speakers partially overlap or one speaker's utterance is very short. The experimental results show that the proposed system significantly outperforms the conventional single-channel PIT system, producing the same performance as a single-channel system using oracle masks. Zhuo Chen 0006, Takuya Yoshioka, Linyu Li 0007, Michael L. Seltzer, Yifan Gong 0001 |
ICASSP | 5 |
| 2018 | Towards Language-Universal End-to-End Speech RecognitionabstractBuilding speech recognizers in multiple languages typically involves replicating a monolingual training recipe for each language, or utilizing a multi-task learning approach where models for different languages have separate output labels but share some internal parameters. In this work, we exploit recent progress in end-to-end speech recognition to create a single multilingual speech recognition system capable of recognizing any of the languages seen in training. To do so, we propose the use of a universal character set that is shared among all languages. We also create a language-specific gating mechanism within the network that can modulate the network's internal representations in a language-specific way. We evaluate our proposed approach on the Microsoft Cortana task across three languages and show that our system outperforms both the individual monolingual systems and systems built with a multi-task learning approach. We also show that this model can be used to initialize a monolingual speech recognizer, and can be used to create a bilingual model for use in code-switching scenarios. Suyoun Kim, Michael L. Seltzer |
ICASSP | 2 |
| 2018 | Improved Training for Online End-to-end Speech Recognition SystemsabstractAchieving high accuracy with end-to-end speech recognizers requires careful parameter initialization prior to training.Otherwise, the networks may fail to find a good local optimum.This is particularly true for online networks, such as unidirectional LSTMs.Currently, the best strategy to train such systems is to bootstrap the training from a tied-triphone system.However, this is time consuming, and more importantly, is impossible for languages without a high-quality pronunciation lexicon.In this work, we propose an initialization strategy that uses teacher-student learning to transfer knowledge from a large, well-trained, offline end-to-end speech recognition model to an online end-to-end model, eliminating the need for a lexicon or any other linguistic resources.We also explore curriculum learning and label smoothing and show how they can be combined with the proposed teacher-student learning for further improvements.We evaluate our methods on a Microsoft Cortana personal assistant task and show that the proposed method results in a 19% relative improvement in word error rate compared to a randomly-initialized baseline system. Suyoun Kim, Michael L. Seltzer, Jinyu Li 0001, Rui Zhao 0017 |
INTERSPEECH | 2 |
| 2017 | May I take your order? A Neural Model for Extracting Structured Information from ConversationsabstractBaolin Peng, Michael Seltzer, Y.C. Ju, Geoffrey Zweig, Kam-Fai Wong. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Baolin Peng, Michael L. Seltzer, Y. C. Ju, Geoffrey Zweig, Kam-Fai Wong |
EACL (1) | 2 |
| 2017 | A study on data augmentation of reverberant speech for robust speech recognitionabstractThe environmental robustness of DNN-based acoustic models can be significantly improved by using multi-condition training data. However, as data collection is a costly proposition, simulation of the desired conditions is a frequently adopted strategy. In this paper we detail a data augmentation approach for far-field ASR. We examine the impact of using simulated room impulse responses (RIRs), as real RIRs can be difficult to acquire, and also the effect of adding point-source noises. We find that the performance gap between using simulated and real RIRs can be eliminated when point-source noises are added. Further we show that the trained acoustic models not only perform well in the distant-talking scenario but also provide better results in the close-talking scenario. We evaluate our approach on several LVCSR tasks which can adequately represent both scenarios. Tom Ko, Vijayaditya Peddinti, Daniel Povey, Michael L. Seltzer, Sanjeev Khudanpur |
ICASSP | 4 |
| 2017 | Large-Scale Domain Adaptation via Teacher-Student LearningabstractHigh accuracy speech recognition requires a large amount of transcribed data for supervised training.In the absence of such data, domain adaptation of a well-trained acoustic model can be performed, but even here, high accuracy usually requires significant labeled data from the target domain.In this work, we propose an approach to domain adaptation that does not require transcriptions but instead uses a corpus of unlabeled parallel data, consisting of pairs of samples from the source domain of the well-trained model and the desired target domain.To perform adaptation, we employ teacher/student (T/S) learning, in which the posterior probabilities generated by the source-domain model can be used in lieu of labels to train the target-domain model.We evaluate the proposed approach in two scenarios, adapting a clean acoustic model to noisy speech and adapting an adults' speech acoustic model to children's speech.Significant improvements in accuracy are obtained, with reductions in word error rate of up to 44% over the original source model without the need for transcribed data in the target domain.Moreover, we show that increasing the amount of unlabeled data results in additional model robustness, which is particularly beneficial when using simulated training data in the target-domain. Jinyu Li 0001, Michael L. Seltzer, Rui Zhao 0017, Yifan Gong 0001 |
INTERSPEECH | 2 |
| 2017 | Toward Human Parity in Conversational Speech RecognitionabstractConversational speech recognition has served as a flagship speech recognition task since the release of the Switchboard corpus in the 1990s. In this paper, we measure a human error rate on the widely used NIST 2000 test set for commercial bulk transcription. The error rate of professional transcribers is 5.9% for the Switchboard portion of the data, in which newly acquainted pairs of people discuss an assigned topic, and 11.3% for the CallHome portion, where friends and family members have open-ended conversations. In both cases, our automated system edges past the human benchmark, achieving error rates of 5.8% and 11.0%, respectively. The key to our system's performance is the use of various convolutional and long-short-term memory acoustic model architectures, combined with a novel spatial smoothing method and lattice-free discriminative acoustic training, multiple recurrent neural network language modeling approaches, and a systematic use of system combination. Comparing frequent errors in our human and machine transcripts, we find them to be remarkably similar, and highly correlated as a function of the speaker. Human subjects find it very difficult to tell which errorful transcriptions come from humans. Overall, this suggests that, given sufficient matched training data, conversational speech transcription engines are approximating human parity in both quantitative and qualitative terms. Wayne Xiong, Jasha Droppo, Xuedong Huang 0001, Frank Seide, Michael L. Seltzer, Andreas Stolcke, Dong Yu 0001, Geoffrey Zweig |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2016 | Linearly augmented deep neural networkabstractDeep neural networks (DNN) are a powerful tool for many large vocabulary continuous speech recognition (LVCSR) tasks. Training a very deep network is a challenging problem and pre-training techniques are needed in order to achieve the best results. In this paper, we propose a new type of network architecture, Linear Augmented Deep Neural Network (LA-DNN). This type of network augments each non-linear layer with a linear connection from layer input to layer output. The resulting LA-DNN model eliminates the need for pre-training, addresses the gradient vanishing problem for deep networks, has higher capacity in modeling linear transformations, trains significantly faster than normal DNN, and produces better acoustic models. The proposed model has been evaluated on TIMIT phoneme recognition and AMI speech recognition tasks. Experimental results show that the LA-DNN models can have 70% fewer parameters than a DNN, while still improving accuracy. On the TIMIT phoneme recognition task, the smaller LA-DNN model improves TIMIT phone accuracy by 2% absolute, and AMI word accuracy by 1.7% absolute. Pegah Ghahremani, Jasha Droppo, Michael L. Seltzer |
ICASSP | 3 |
| 2016 | Deep beamforming networks for multi-channel speech recognitionabstractDespite the significant progress in speech recognition enabled by deep neural networks, poor performance persists in some scenarios. In this work, we focus on far-field speech recognition which remains challenging due to high levels of noise and reverberation in the captured speech signals. We propose to represent the stages of acoustic processing including beamforming, feature extraction, and acoustic modeling, as three components of a single unified computational network. The parameters of a frequency-domain beam-former are first estimated by a network based on features derived from the microphone channels. These filter coefficients are then applied to the array signals to form an enhanced signal. Conventional features are then extracted from this signal and passed to a second network that performs acoustic modeling for classification. The parameters of both the beamforming and acoustic modeling networks are trained jointly using back-propagation with a common cross-entropy objective function. In experiments on the AMI meeting corpus, we observed improvements by pre-training each sub-network with a network-specific objective function before joint training of both networks. The proposed method obtained a 3.2% absolute word error rate reduction compared to a conventional pipeline of independent processing stages. Shinji Watanabe 0001, Hakan Erdogan, Liang Lu 0001, John R. Hershey, Michael L. Seltzer, Guoguo Chen, Yu Zhang 0033, Michael I. Mandel, Dong Yu 0001 |
ICASSP | 6 |
| 2016 | On the Role of Nonlinear Transformations in Deep Neural Network Acoustic Models
Tasha Nagamine, Michael L. Seltzer, Nima Mesgarani |
INTERSPEECH | 2 |
| 2015 | Improving speech recognition in reverberation using a room-aware deep neural network and multi-task learningabstractIn this paper, we propose two approaches to improve deep neural network (DNN) acoustic models for speech recognition in reverberant environments. Both methods utilize auxiliary information in training the DNN but differ in the type of information and the manner in which it is used. The first method uses parallel training data for multi-task learning, in which the network is trained to perform both a primary senone classification task and a secondary feature enhancement task using a shared representation. The second method uses a parameterization of the reverberant environment extracted from the observed signal to train a room-aware DNN. Experiments were performed on the single microphone task of the REVERB Challenge corpus. The proposed approach obtained a word error rate of 7.8% on the SimData test set, which is lower than all reported systems using the same training data and evaluation conditions, and 27.5% on the mismatched RealData test set, which is lower than all but two systems. Ritwik Giri, Michael L. Seltzer, Jasha Droppo, Dong Yu 0001 |
ICASSP | 2 |
| 2015 | Speech recognition with prediction-adaptation-correction recurrent neural networksabstractWe propose the prediction-adaptation-correction RNN (PAC-RNN), in which a correction DNN estimates the state posterior probability based on both the current frame and the prediction made on the past frames by a prediction DNN. The result from the main DNN is fed back to the prediction DNN to make better predictions for the future frames. In the PAC-RNN, we can consider that, given the new, current frame information, the main DNN makes a correction on the prediction made by the prediction DNN. Alternatively, it can be viewed as adapting the main DNN's behavior based on the prediction DNN's prediction. Experiments on the TIMIT phone recognition task indicate that the PAC-RNN outperforms DNN, RNN, and LSTM with 2.4%, 2.1%, and 1.9% absolute phone accuracy improvement, respectively. We found that incorporating the prediction objective and including the recurrent loop are both important to boost the performance of the PAC-RNN. Yu Zhang 0033, Dong Yu 0001, Michael L. Seltzer, Jasha Droppo |
ICASSP | 3 |
| 2015 | Exploring how deep neural networks form phonemic categoriesabstractDeep neural networks (DNNs) have become the dominant technique for acoustic-phonetic modeling due to their markedly improved performance over other models. Despite this, little is understood about the computation they implement in creating phonemic categories from highly variable acoustic signals. In this paper, we analyzed a DNN trained for phoneme recognition and characterized its representational properties, both at the single node and population level in each layer. At the single node level, we found strong selectivity to distinct phonetic features in all layers. Node selectivity to specific manners and places of articulation appeared from the first hidden layer and became more explicit in deeper layers. Furthermore, we found that nodes with similar phonetic feature selectivity were differentially activated to different exemplars of these features. Thus, each node becomes tuned to a particular acoustic manifestation of the same feature, providing an effective representational basis for the formation of invariant phonemic categories. This study reveals that phonetic features organize the activations in different layers of a DNN, a result that mirrors the recent findings of feature encoding in the human auditory system. These insights may provide better understanding of the limitations of current models, leading to new strategies to improve their performance. Index Terms: Deep neural networks, deep learning, automatic speech recognition. Tasha Nagamine, Michael L. Seltzer, Nima Mesgarani |
INTERSPEECH | 2 |
| 2015 | Deep Neural Networks for Single-Channel Multi-Talker Speech RecognitionabstractWe investigate techniques based on deep neural networks (DNNs) for attacking the single-channel multi-talker speech recognition problem. Our proposed approach contains five key ingredients: a multi-style training strategy on artificially mixed speech data, a separate DNN to estimate senone posterior probabilities of the louder and softer speakers at each frame, a weighted finite-state transducer (WFST)-based two-talker decoder to jointly estimate and correlate the speaker and speech, a speaker switching penalty estimated from the energy pattern change in the mixed-speech, and a confidence based system combination strategy. Experiments on the 2006 speech separation and recognition challenge task demonstrate that our proposed DNN-based system has remarkable noise robustness to the interference of a competing speaker. The best setup of our proposed systems achieves an average word error rate (WER) of 18.8% across different SNRs and outperforms the state-of-the-art IBM superhuman system by 2.8% absolute with fewer assumptions. Chao Weng, Dong Yu 0001, Michael L. Seltzer, Jasha Droppo |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Factored adaptation of speaker and environment using orthogonal subspace transformsabstractThis paper presents a subspace-based acoustic factorization framework to transform-based adaptation in speech recognition. In the proposed method, adaptation transforms are projected onto factor-dependent low-rank subspaces in a way that decouples the combined extrinsic factors affecting the speech signals. Usually, mismatch between the observed speech and the acoustic models is caused by multiple acoustic factors simultaneously, such as the speaker and environment. Data-driven adaptation methods, such as constrained MLLR, compensate for all sources of mismatch jointly. In many scenarios, however, it is highly desirable to separate the sources of mismatch in order to adapt to speaker and environment variability independently. This adds flexibility to the model adaptation framework. For example, a speaker transform obtained in one environment can be reused when the same speaker is in different environments. Or, an environment transform obtained during training, independently of speaker identities, can be applied to a speaker in deployment. One way to achieve this factorization is to construct each set of transforms such that they are orthogonal to each other, so that any change in one acoustic factor keeps other factors intact. The proposed subspace approach provides a straightforward factor analysis framework while allows us to explicitly formulate the independence among the estimated factor transforms. A series of experiments performed on the Aurora 4 corpus validates our approach. Hyunson Seo, Hong-Goo Kang, Michael L. Seltzer |
ICASSP | 3 |
| 2014 | Single-channel mixed speech recognition using deep neural networksabstractIn this work, we study the problem of single-channel mixed speech recognition using deep neural networks (DNNs). Using a multi-style training strategy on artificially mixed speech data, we investigate several different training setups that enable the DNN to generalize to corresponding similar patterns in the test data. We also introduce a WFST-based two-talker decoder to work with the trained DNNs. Experiments on the 2006 speech separation and recognition challenge task demonstrate that the proposed DNN-based system has remarkable noise robustness to the interference of a competing speaker. The best setup of our proposed systems achieves an overall WER of 19.7% which improves upon the results obtained by the state-of-the-art IBM superhuman system by 1.9% absolute, with fewer assumptions and lower computational complexity. Chao Weng, Dong Yu 0001, Michael L. Seltzer, Jasha Droppo |
ICASSP | 3 |
| 2014 | Towards better performance with heterogeneous training data in acoustic modeling using deep neural networksabstractModeling heterogeneous data sources remains a fundamental chal-lenge of acoustic modeling in speech recognition. We call this the multi-condition problem because the speech data come from many different conditions. In this paper, we introduce the fundamen-tal confusability problem in multi-condition learning, then discuss the problem formalization, the taxonomy, and the architectures for multi-condition learning. While the ideas presented are applicable to all classifiers, we focus our attention in this work on acoustic models based on deep neural networks (DNN). We propose four different strategies for multi-condition learning of a DNN that we refer to as a mixed-condition model, a condition-dependent model, a condition-normalizing model, and a condition-aware model. Based on the experimental results on the voice search and short message dictation task and the Aurora 4 task, we show that the confusabil-ity introduced when modeling heterogeneous data depends on the source of acoustic distortion itself, the front-end feature extractor, and the classifier. We also demonstrate the best approach for dealing with heterogeneous data may not be to let the model sort it out blindly, even with a classifier as sophisticated as a DNN. Index Terms — Multi-task learning, deep learning, CD-DNN-HMM, noise robustness, channel compensation Yan Huang 0028, Malcolm Slaney, Michael L. Seltzer, Yifan Gong 0001 |
INTERSPEECH | 3 |
| 2014 | The influence of pitch and noise on the discriminability of filterbank featuresabstractMost features used for speech recognition are derived from the out-put of a filterbank inspired by the auditory system. The two most commonly used filter shapes are the triangular filters used in MFCC (mel-frequency cepstral coefficients) and the gammatone filters that model psychoacoustic critical bands. However, for both of these fil-terbanks there are free parameters that must be chosen by the system designer. In this paper, we explore the effect that different parameter settings have on the discriminability of speech sound classes. Specif-ically, we focus our attention on two primary parameters: the filter shape (triangular or gammatone) and the filter bandwidth. We use variations in the noise level and the pitch to explore the behavior of different filterbanks. We use the Fisher linear discriminant to give us insight about why some filterbanks perform better than others. We observe three things: 1) there are significant differences even among different implementations of the same filterbank, 2) wider filters help remove the non-informative pitch information, and 3) the Fisher cri-teria helps us understand why. We validate the Fisher measure with speech recognition experiments on the Aurora-4 speech corpus. Malcolm Slaney, Michael L. Seltzer |
INTERSPEECH | 2 |
| 2014 | An introduction to computational networks and the computational network toolkit (invited talk)
Dong Yu 0001, Adam Eversole, Michael L. Seltzer, Kaisheng Yao, Brian Guenter, Oleksii Kuchaiev, Frank Seide, Huaming Wang, Jasha Droppo, Zhiheng Huang, Geoffrey Zweig, Christopher J. Rossbach, Jon Currey |
INTERSPEECH | 3 |
| 2013 | Recent advances in deep learning for speech research at MicrosoftabstractDeep learning is becoming a mainstream technology for speech recognition at industrial scale. In this paper, we provide an overview of the work by Microsoft speech researchers since 2009 in this area, focusing on more recent advances which shed light to the basic capabilities and limitations of the current deep learning technology. We organize this overview along the feature-domain and model-domain dimensions according to the conventional approach to analyzing speech systems. Selected experimental results, including speech recognition and related applications such as spoken dialogue and language modeling, are presented to demonstrate and analyze the strengths and weaknesses of the techniques described in the paper. Potential improvement of these techniques and future research directions are discussed. Li Deng 0001, Jinyu Li 0001, Jui-Ting Huang, Kaisheng Yao, Dong Yu 0001, Frank Seide, Michael L. Seltzer, Geoffrey Zweig, Xiaodong He 0001, Jason D. Williams, Yifan Gong 0001, Alex Acero |
ICASSP | 7 |
| 2013 | Multi-task learning in deep neural networks for improved phoneme recognitionabstractIn this paper we demonstrate how to improve the performance of deep neural network (DNN) acoustic models using multi-task learning. In multi-task learning, the network is trained to perform both the primary classification task and one or more secondary tasks using a shared representation. The additional model parameters associated with the secondary tasks represent a very small increase in the number of trained parameters, and can be discarded at runtime. In this paper, we explore three natural choices for the secondary task: the phone label, the phone context, and the state context. We demonstrate that, even on a strong baseline, multi-task learning can provide a significant decrease in error rate. Using phone context, the phonetic error rate (PER) on TIMIT is reduced from 21.63% to 20.25% on the core test set, and surpassing the best performance in the literature for a DNN that uses a standard feed-forward network architecture. Michael L. Seltzer, Jasha Droppo |
ICASSP | 1 |
| 2013 | An investigation of deep neural networks for noise robust speech recognitionabstractRecently, a new acoustic model based on deep neural networks (DNN) has been introduced. While the DNN has generated significant improvements over GMM-based systems on several tasks, there has been no evaluation of the robustness of such systems to environmental distortion. In this paper, we investigate the noise robustness of DNN-based acoustic models and find that they can match state-of-the-art performance on the Aurora 4 task without any explicit noise compensation. This performance can be further improved by incorporating information about the environment into DNN training using a new method called noise-aware training. When combined with the recently proposed dropout training technique, a 7.5% relative improvement over the previously best published result on this task is achieved using only a single decoding pass and no additional decoding complexity compared to a standard DNN. Michael L. Seltzer, Dong Yu 0001, Yongqiang Wang 0005 |
ICASSP | 1 |
| 2013 | Deep neural network features and semi-supervised training for low resource speech recognitionabstractWe propose a new technique for training deep neural networks (DNNs) as data-driven feature front-ends for large vocabulary continuous speech recognition (LVCSR) in low resource settings. To circumvent the lack of sufficient training data for acoustic modeling in these scenarios, we use transcribed multilingual data and semi-supervised training to build the proposed feature front-ends. In our experiments, the proposed features provide an absolute improvement of 16% in a low-resource LVCSR setting with only one hour of in-domain training data. While close to three-fourths of these gains come from DNN-based features, the remaining are from semi-supervised training. Samuel Thomas 0001, Michael L. Seltzer, Kenneth Church 0001, Hynek Hermansky |
ICASSP | 2 |
| 2012 | Improvements to VTS feature enhancementabstractBy explicitly modelling the distortion of speech signals, model adaptation based on vector Taylor series (VTS) approaches have been shown to significantly improve the robustness of speech recognizers to environmental noise. However, the computational cost of VTS model adaptation (MVTS) methods hinders them from being widely used because they need to adapt all the HMM parameters for every utterance at runtime. In contrast, VTS feature enhancement (FVTS) methods have more computation advantages because they do not need multiple decoding passes and do not adapt all the HMM model parameters. In this paper, we propose two improvements to VTS feature enhancement: updating all of the environment distortion parameters and noise adaptive training of the front-end GMM. In addition, we investigate some other performance-related issues such as the selection of FVTS algorithms and the spectrum domain that MFCC is extracted from. As an important result of our investigation, we established the FVTS method can achieve comparable accuracy as the MVTS method with a smaller runtime cost. This makes FVTS method an ideal candidate for real world tasks. Jinyu Li 0001, Michael L. Seltzer, Yifan Gong 0001 |
ICASSP | 2 |
| 2012 | Efficient VTS Adaptation Using Jacobian ApproximationabstractBy exploiting a model of environmental distortion, model adaptation based on vector Taylor series (VTS) approaches have been shown to significantly improve the robustness of speech recognizers to environmental noise. However, the computational cost of VTS model adaptation (MVTS) methods hinders them from being more widely used. In this paper, we propose to reduce the computational cost of MVTS by replacing the Jacobian matrix used in the vector Taylor series approximation with a diagonal Jacobian matrix (DJVTS). We verify this approximation by showing that the Jacobian matrices are dominated by their diagonal elements and therefore the model distortion introduced by this approximation is very small. DJVTS gives similar accuracy as the standard MVTS method with significant reduction in computational cost. The proposed method also achieves higher accuracy than VTS-based feature enhancement. Jinyu Li 0001, Michael L. Seltzer, Yifan Gong 0001 |
INTERSPEECH | 2 |
| 2012 | Factored adaptation using a combination of feature-space and model-space transforms
Michael L. Seltzer, Alex Acero |
INTERSPEECH | 1 |
| 2011 | Factored adaptation for separable compensation of speaker and environmental variabilityabstractWhile many algorithms for speaker or environment adaptation have been proposed, far less attention has been paid to approaches which address both factors. We recently proposed a method called factored adaptation that can jointly compensate for speaker and environmental mismatch using a cascade of CMLLR transforms that separately compensate for the environment and speaker variability. Performing adaptation in this manner enables a speaker transform estimated in one environment to be be applied when the same user is in different environments. While this algorithm performed well, it relied on knowledge of the operating environment in both training and test. In this paper, we show how unsupervised environment clustering can be used to eliminate this requirement. The improved factored adaptation algorithm achieves relative improvements of 10-18% over conventional CMLLR when applying speaker transforms across environments without needing any additional a priori knowledge. Michael L. Seltzer, Alex Acero |
ASRU | 1 |
| 2011 | Joint encoding of the waveform and speech recognition features using a transform codecabstractWe propose a new transform speech codec that jointly encodes a wideband waveform and its corresponding wideband and narrowband speech recognition features. For distributed speech recognition, wideband features are compressed and transmitted as side information. The waveform is then encoded in a manner that exploits the information already captured by the speech features. Narrowband speech acoustic features can be synthesized at the server by applying a transformation to the decoded wideband features. An evaluation conducted on an in-car speech recognition task show that at 16 kbps our new system typically shows essentially no impact in word error rate compared to uncompressed audio, whereas the standard transform codec produces up to a 20% increase in word error rate. In addition, good quality speech is obtained for playback and transcription, with PESQ scores ranging from 3.2 to 3.4. Michael L. Seltzer, Jasha Droppo, Henrique S. Malvar, Alex Acero |
ICASSP | 2 |
| 2011 | CROWDMOS: An approach for crowdsourcing mean opinion score studiesabstractMOS (mean opinion score) subjective quality studies are used to evaluate many signal processing methods. Since laboratory quality studies are time consuming and expensive, researchers often run small studies with less statistical significance or use objective measures which only approximate human perception. We propose a cost-effective and convenient measure called crowdMOS, obtained by having internet users participate in a MOS-like listening study. Workers listen and rate sentences at their leisure, using their own hardware, in an environment of their choice. Since these individuals cannot be supervised, we propose methods for detecting and discarding inaccurate scores. To automate crowdMOS testing, we offer a set of freely distributable, open-source tools for Amazon Mechanical Turk, a platform designed to facilitate crowdsourcing. These tools implement the MOS testing methodology described in this pa per, providing researchers with a user-friendly means of performing subjective quality evaluations without the overhead associated with laboratory studies. Finally, we demonstrate the use of crowdMOS using data from the Blizzard text-to-speech competition, showing that it delivers accurate and repeatable results. Flavio P. Ribeiro, Dinei A. F. Florêncio, Cha Zhang, Michael L. Seltzer |
ICASSP | 4 |
| 2011 | Separating Speaker and Environmental Variability Using Factored TransformsabstractTwo primary sources of variability that degrade accuracy in speech recognition systems are the speaker and the environment. While many algorithms for speaker or environment adaptation have been proposed to improve performance, far less attention has been paid to approaches which address for both factors. In this paper, we present a method for compensating for speaker and environmental mismatch using a cascade of CMLLR transforms. The proposed approach enables speaker transforms estimated in one environment to be effectively applied to speech from the same user in a different environment. This approach can be further improved using a new training method called speaker and environment adaptive training method. When applying speaker transforms to new environments, the proposed approach results in a 13% relative improvement over conventional CMLLR. Michael L. Seltzer, Alex Acero |
INTERSPEECH | 1 |
| 2011 | Improved Bottleneck Features Using Pretrained Deep Neural NetworksabstractBottleneck features have been shown to be effective in improving the accuracy of automatic speech recognition (ASR) systems. Conventionally, bottleneck features are extracted from a multi-layer perceptron (MLP) trained to predict context-independent monophone states. The MLP typically has three hidden layers and is trained using the backpropagation algorithm. In this paper, we propose two improvements to the training of bottleneck features motivated by recent advances in the use of deep neural networks (DNNs) for speech recognition. First, we show how the use of unsupervised pretraining of a DNN enhances the network’s discriminative power and improves the bottleneck features it generates. Second, we show that a neural network trained to predict context-dependent senone targets produces better bottleneck features than one trained to predict monophone states. Bottleneck features trained using the proposed methods produced a 16% relative reduction in sentence error rate over conventional bottleneck features on a large vocabulary business search task. Dong Yu 0001, Michael L. Seltzer |
INTERSPEECH | 2 |
| 2010 | Acoustic model adaptation via Linear Spline Interpolation for robust speech recognitionabstractWe recently proposed a new algorithm to perform acoustic model adaptation to noisy environments called Linear Spline Interpolation (LSI). In this method, the nonlinear relationship between clean and noisy speech features is modeled using linear spline regression. Linear spline parameters that minimize the error the between the predicted noisy features and the actual noisy features are learned from training data. A variance associated with each spline segment captures the uncertainty in the assumed model. In this work, we extend the LSI algorithm in two ways. First, the adaptation scheme is extended to compensate for the presence of linear channel distortion. Second, we show how the noise and channel parameters can be updated during decoding in an unsupervised manner within the LSI framework. Using LSI, we obtain an average relative improvement in word error rate of 10.8% over VTS adaptation on the Aurora 2 task with improvements of 15-18% at SNRs between 10 and 15 dB. Michael L. Seltzer, Alex Acero, Kaustubh Kalgaonkar |
ICASSP | 1 |
| 2010 | Binary coding of speech spectrograms using a deep auto-encoderabstractThis paper reports our recent exploration of the layer-by-layer learning strategy for training a multi-layer generative model of patches of speech spectrograms. The top layer of the generative model learns binary codes that can be used for efficient compression of speech and could also be used for scalable speech recognition or rapid speech content retrieval. Each layer of the generative model is fully connected to the layer below and the weights on these connections are pretrained efficiently by using the contrastive divergence approximation to the log likelihood gradient. After layer-bylayer pre-training we “unroll” the generative model to form a deep auto-encoder, whose parameters are then fine-tuned using back-propagation. To reconstruct the full-length speech spectrogram, individual spectrogram segments predicted by their respective binary codes are combined using an overlapand-add method. Experimental results on speech spectrogram coding demonstrate that the binary codes produce a logspectral distortion that is approximately 2 dB lower than a subband vector quantization technique over the entire frequency range of wide-band speech. Index Terms: deep learning, speech feature extraction, neural networks, auto-encoder, binary codes, Boltzmann machine Li Deng 0001, Michael L. Seltzer, Dong Yu 0001, Alex Acero, Abdel-rahman Mohamed, Geoffrey E. Hinton |
INTERSPEECH | 2 |
| 2010 | HMM adaptation using linear spline interpolation with integrated spline parameter training for robust speech recognitionabstractWe recently proposed a method for HMM adaptation to noisy environments called Linear Spline Interpolation (LSI). LSI uses linear spline regression to model the relationship between clean and noisy speech features. In the original algorithm, stereo training data was used to learn the spline parameters that minimize the error between the predicted and actual noisy speech features. The estimated splines are then used at runtime to adapt the clean HMMs to the current environment. While good results can be obtained with this approach, the performance is limited by the fact that the splines are trained independently from the speech recognizer and as such, they may actually be suboptimal for adaptation. In this work, we introduce a new Generalized EM algorithm for estimating the spline parameters using the speech recognizer itself. Experiments on the Aurora 2 task show that using LSI adaptation with splines trained in this manner results in a 20% improvement over the original LSI algorithm that used splines estimated from stereo data and a 28% improvement over VTS adaptation. Michael L. Seltzer, Alex Acero |
INTERSPEECH | 1 |
| 2010 | Noise Adaptive Training for Robust Automatic Speech RecognitionabstractIn traditional methods for noise robust automatic speech recognition, the acoustic models are typically trained using clean speech or using multi-condition data that is processed by the same feature enhancement algorithm expected to be used in decoding. In this paper, we propose a noise adaptive training (NAT) algorithm that can be applied to all training data that normalizes the environmental distortion as part of the model training. In contrast to feature enhancement methods, NAT estimates the underlying “pseudo-clean” model parameters directly without relying on point estimates of the clean speech features as an intermediate step. The pseudo-clean model parameters learned with NAT are later used with vector Taylor series (VTS) model adaptation for decoding noisy utterances at test time. Experiments performed on the Aurora 2 and Aurora 3 tasks demonstrate that the proposed NAT method obtain relative improvements of 18.83% and 32.02%, respectively, over VTS model adaptation. Ozlem Kalinli, Michael L. Seltzer, Jasha Droppo, Alex Acero |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Noise robust model adaptation using linear spline interpolationabstractThis paper presents a novel data-driven technique for performing acoustic model adaptation to noisy environments. In the presence of additive noise, the relationship between log mel spectra of speech, noise and noisy speech is nonlinear. Traditional methods linearize this relationship using the mode of the nonlinearity or use some other approximation. The approach presented in this paper models this nonlinear relationship using linear spline regression. In this method, the set of spline parameters that minimizes the error between the predicted and actual noisy speech features is learned from training data, and used at runtime to adapt clean acoustic model parameters to the current noise conditions. Experiments were performed to evaluate the performance of the system on the Aurora 2 task. Results show that the proposed adaptation algorithm (word accuracy 89.22%) outperforms VTS model adaptation (word accuracy 88.38%). Kaustubh Kalgaonkar, Michael L. Seltzer, Alex Acero |
ASRU | 2 |
| 2009 | Noise adaptive training using a vector taylor series approach for noise robust automatic speech recognitionabstractIn traditional methods for noise robust automatic speech recognition, the acoustic models are typically trained using clean speech or using multi-condition data that is processed by the same feature enhancement algorithm expected to be used in decoding. In this paper, we propose a noise adaptive training (NAT) algorithm that can be applied to all training data that normalizes the environmental distortion as part of the model training. In contrast to the feature enhancement methods, NAT estimates the underlying “pseudo-clean” model parameters directly without relying on point estimates of the clean speech features as an intermediate step. The pseudo-clean model parameters learned with NAT are later used with vector Taylor series (VTS) model adaptation for decoding noisy utterances at test time. Experiments performed on the Aurora 2 and Aurora 3 tasks, demonstrate that the proposed NAT method obtain relative improvements of 18.83% and 32.02%, respectively, over VTS model adaptation. Ozlem Kalinli, Michael L. Seltzer, Alex Acero |
ICASSP | 2 |
| 2009 | The data deluge: Challenges and opportunities of unlimited data in statistical signal processingabstractRecently, there has been a dramatic increase of the amount of audio, video, and images created and shared on the internet by users around the world. Much of this content is publicly available and free of cost. When viewed through the lens of pattern classification, this content can be seen as a virtually unlimited supply of training data for various statistical modeling and labeling tasks such as speech recognition and computer vision. In order to effectively exploit this data resource, significant research challenges must be addressed. In this paper, we present three significant challenges that must be solved to harness the potential of this “data deluge”. We then describe recent work in spoken language processing and image processing that has begun to address these challenges in order to tackle large-scale classification tasks. By bringing together the work of these two communities, we hope to stimulate the cross-pollination of ideas and methods among different signal processing communities. Michael L. Seltzer, Lei Zhang 0001 |
ICASSP | 1 |
| 2009 | Voice search of structured media dataabstractThis paper addresses the problem of using unstructured queries to search a structured database in voice search applications. By incorporating structural information in music metadata, the end-to-end search error has been reduced by 15% on text queries and up to 11% on spoken queries. Based on that, an HMM sequential rescoring model has reduced the error rate by 28% on text queries and up to 23% on spoken queries compared to the baseline system. Furthermore, a phonetic similarity model has been introduced to compensate speech recognition errors, which has improved the end-to-end search accuracy consistently across different levels of speech recognition accuracy. Young-In Song, Ye-Yi Wang, Yun-Cheng Ju, Michael L. Seltzer, Ivan Tashev, Alex Acero |
ICASSP | 4 |
| 2009 | Improving perceived accuracy for in-car media searchabstractSpeech recognition technology is prone to mistakes, but this is not the only source of errors that cause speech recognition systems to fail; sometimes the user simply does not utter the command correctly. Usually, user mistakes are not considered when a system is designed and evaluated. This creates a gap between the claimed accuracy of the system and the actual accuracy perceived by the users. We address this issue quantitatively in our in-car infotainment media search task and propose expanding the capability of voice command to accommodate user mistakes while retaining a high percentage of the performance for queries with correct syntax. As a result, failures caused by user mistakes were reduced by an absolute 70 % at the cost of a drop in accuracy of only 0.28%. Yun-Cheng Ju, Michael L. Seltzer, Ivan Tashev |
INTERSPEECH | 2 |
| 2008 | Robust design of wideband loudspeaker arraysabstractLoudspeaker arrays usually are used in professional sound reinforcement systems to provide uniform sound coverage of the listening area. They can also be used for focusing the sound to the user for semi-private communication and reducing the overall noise pollution. In this paper we describe a procedure for designing broadband beamformers for loudspeaker arrays that is robust to the manufacturing tolerances of the loudspeakers - the limiting factor for achieving high directivity. The designed beamformer is evaluated using simulations and measurements of actual loudspeaker array. Ivan Tashev, Jasha Droppo, Michael L. Seltzer, Alex Acero |
ICASSP | 3 |
| 2008 | Maximum a posteriori ICA: Applying prior knowledge to the separation of acoustic sourcesabstractIndependent component analysis (ICA) for convolutive mixtures is often applied in the frequency domain due to the desirable decoupling into independent instantaneous mixtures per frequency bin. This approach suffers from a well-known scaling and permutation ambiguity. Existing methods perform a computation-heavy and sometimes unreliable phase of post-processing which typically makes use of knowledge regarding the geometry of the sensors post-ICA. In this paper, we propose a natural way to incorporate a priori knowledge of the unmixing matrix in the form of a prior distribution. This softly constrains ICA in a manner that avoids the permutation problem, and also allows us to integrate information about the environment, such as likely user configurations, into ICA using a unified statistical framework. Maximum a priori ICA easily follows from the maximum likelihood derivation of ICA. Its effectiveness is demonstrated through a series of experiments on convolutive mixtures of speech signals. Graham W. Taylor, Michael L. Seltzer, Alex Acero |
ICASSP | 2 |
| 2008 | Towards a non-parametric acoustic model: an acoustic decision tree for observation probability calculationabstractModern automatic speech recognition systems use Gaussian mixture models (GMM) on acoustic observations to model the probability of producing a given observation under any one of many hidden discrete phonetic states. This paper investigates the feasibility of using an acoustic decision tree to directly model these probabilities. Unlike the more common phonetic decision tree, which asks questions about phonetic context, an acoustic decision tree asks questions about the vectorvalued observations. Three different types of acoustic questions are proposed and evaluated, including LDA, PCA, and MMI questions. Frame classification experiments are run on a subset of the Switchboard corpus. On these experiments, the acoustic decision tree produces slightly better results than maximum likelihood trained GMMs, with significantly less computation. Some theoretical advantages of the acoustic decision tree are discussed, including more economical use of the training data and reduced mismatch between the acoustic model and the true probability distribution of the phonetic labels. Jasha Droppo, Michael L. Seltzer, Alex Acero, Yu-Hsiang Bosco Chiu |
INTERSPEECH | 2 |
| 2007 | Microphone Array Post-Filter using Incremental Bayes Learning to Track the Spatial Distributions of Speech and NoiseabstractWhile current post-filtering algorithms for microphone array applications can enhance beamformer output signals, they assume that the noise is either incoherent or diffuse, and make no allowances for point noise sources which may be strongly correlated across the microphones. In this paper, we present a novel post-filtering algorithm that alleviates this assumption by tracking the spatial as well as spectral distribution of the speech and noise sources present. A generative statistical model is employed to model the speech and noise sources at distinct regions in the soundfield, and incremental Bayesian learning is used to track the model parameters over time. This approach allows a post-filter derived from these parameters to effectively suppress both diffuse ambient noise and interfering point sources. The performance of the proposed approach is evaluated on multiple recordings made in a realistic office environment. Michael L. Seltzer, Ivan Tashev, Alex Acero |
ICASSP (1) | 1 |
| 2007 | Robust location understanding in spoken dialog systems using intersectionsabstractThe availability of digital maps and mapping software has led to significant growth in location-based software and services. To safely use these applications in mobile and automotive scenarios, users must be able to input precise locations using speech. In this paper, we propose a novel method for location understanding based on spoken intersections. The proposed approach utilizes a rich, automatically-generated grammar for street names that maps all street name variations into a single canonical semantic representation. This representation is then transformed to a sequence of position-dependent subword units. This sequence is used by a classifier based on the vector space model to reliably recognize spoken intersections in the presence of recognition errors and incomplete street names. The efficacy of the proposed approach is demonstrated using data collected from users of a deployed spoken dialog system. Michael L. Seltzer, Yun-Cheng Ju, Ivan Tashev, Alex Acero |
INTERSPEECH | 1 |
| 2007 | Automatic Removal of Typed Keystrokes From Speech SignalsabstractComputers are increasingly being used to capture audio in various applications such as video conferencing and meeting recording. In many of these applications, user may be simultaneously typing on the keyboard, e.g., to take notes or search for information. As a result, the captured speech signals are significantly corrupted by sounds generated by the user's keystrokes. In this paper we propose an algorithm to automatically detect and remove keystrokes from speech signals. The proposed method does not require any user training or enrollment and is computationally efficient. The keystroke removal algorithm generates significantly enhanced speech as measured by both user listening tests and speech recognition experiments Amarnag Subramanya, Michael L. Seltzer, Alex Acero |
IEEE Signal Process. Lett. | 2 |
| 2007 | Training Wideband Acoustic Models Using Mixed-Bandwidth Training Data for Speech RecognitionabstractOne serious difficulty in the deployment of wideband speech recognition systems for new tasks is the expense in both time and cost of obtaining sufficient training data. A more economical approach is to collect telephone speech and then restrict the application to operate at the telephone bandwidth. However, this generally results in suboptimal performance compared to a wideband recognition system. In this paper, we propose a novel expectation-maximization (EM) algorithm in which wideband acoustic models are trained using a small amount of wideband speech and a larger amount of narrowband speech. We show how this algorithm can be incorporated into the existing training schemes of hidden Markov model (HMM) speech recognizers. Experiments performed using wideband speech and telephone speech demonstrate that the proposed mixed-bandwidth training algorithm results in significant improvements in recognition accuracy over conventional training strategies when the amount of wideband data is limited Michael L. Seltzer, Alex Acero |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Automatic removal of typed keystrokes from speech signalsabstractComputers are increasingly being used to capture audio in various applications such as video conferencing and meeting recording. In many of these applications, user may be simultaneously typing on the keyboard, e.g., to take notes or search for information. As a result, the captured speech signals are significantly corrupted by sounds generated by the user's keystrokes. In this paper we propose an algorithm to automatically detect and remove keystrokes from speech signals. The proposed method does not require any user training or enrollment and is computationally efficient. The keystroke removal algorithm generates significantly enhanced speech as measured by both user listening tests and speech recognition experiments Amarnag Subramanya, Michael L. Seltzer, Alex Acero |
INTERSPEECH | 2 |
| 2006 | Subband Likelihood-Maximizing Beamforming for Speech Recognition in Reverberant EnvironmentsabstractIn this paper, we introduce subband likelihood-maximizing beamforming (S-LIMABEAM), a new microphone-array processing algorithm specifically designed for speech recognition applications. The proposed algorithm is an extension of the previously developed LIMABEAM array processing algorithm. Unlike most array processing algorithms which operate according to some waveform-level objective function, the goal of LIMABEAM is to find the set of array parameters that maximizes the likelihood of the correct recognition hypothesis. Optimizing the array parameters in this manner results in significant improvements in recognition accuracy over conventional array processing methods when speech is corrupted by additive noise and moderate levels of reverberation. Despite the success of the LIMABEAM algorithm in such environments, little improvement was achieved in highly reverberant environments. In such situations where the noise is highly correlated to the speech signal and the number of filter parameters to estimate is large, subband processing has been used to improve the performance of LMS-type adaptive filtering algorithms. We use subband processing principles to design a novel array processing architecture in which select groups of subbands are processed jointly to maximize the likelihood of the resulting speech recognition features, as measured by the recognizer itself. By creating a subband filtering architecture that explicitly accounts for the manner in which recognition features are computed, we can effectively apply the LIMABEAM framework to highly reverberant environments. By doing so, we are able to achieve improvements in word error rate of over 20% compared to conventional methods in highly reverberant environments Michael L. Seltzer, Richard M. Stern |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Training Wideband Acoustic Models using Mixed-Bandwidth Training Data via Feature Bandwidth ExtensionabstractOne serious difficulty in the deployment of wideband speech recognition systems for new tasks is the expense in both time and cost of obtaining sufficient training data. A more economical approach is to collect telephone speech and then restrict the application to operate at the telephone bandwidth. However, this generally results in sub-optimal performance. We propose a new algorithm for training wideband acoustic models that requires only a small amount of wideband speech augmented by a larger amount of narrowband speech. The algorithm operates by first converting the narrowband features to wideband features through a process called feature bandwidth extension. The bandwidth-extended features are then combined with available wideband data to train the acoustic models using a modified version of the conventional forward-backward algorithm. Experiments performed using wideband speech and telephone speech demonstrate that the proposed mixed-bandwidth training algorithm results in significant improvements in recognition accuracy over conventional training strategies when the amount of wideband data is limited. Michael L. Seltzer, Alex Acero |
ICASSP (1) | 1 |
| 2005 | Robust bandwidth extension of noise-corrupted narrowband speechabstractWe present a new bandwidth extension algorithm for converting narrowband telephone speech into wideband speech using a transformation in the mel cepstral domain. Unlike previous approaches, the proposed method is designed specifically for bandwidth extension of narrowband speech that has been corrupted by environmental noise. We show that by exploiting previous research in mel cepstrum feature enhancement, we can create a unified probabilistic framework under which the feature denoising and bandwidth extension processes are tightly integrated using a single shared statistical model. By doing so, we are able to both denoise the observed narrowband speech and robustly extend its bandwidth in a jointly optimal manner. A series of experiments on clean and noise-corrupted narrowband speech is performed to validate our approach. Michael L. Seltzer, Alex Acero, Jasha Droppo |
INTERSPEECH | 1 |
| 2004 | Parameter sharing in subband likelihood-maximizing beamforming for speech recognition using microphone arraysabstractIn this paper, we present methods to improve the computational efficiency of our previously proposed algorithm for microphone array processing for speech recognition, called subband likelihood-maximizing beamforming (S-LIMABEAM). In S-LIMABEAM, the parameters of a subband filter-and-sum beamformer are optimized to maximize the likelihood of the correct transcription of the utterance, as measured by the speech recognizer itself. This approach has been shown to produce significant improvements in recognition accuracy over conventional array processing methods in a variety of noisy and reverberant environments. However, because of the manner in which recognition features are computed, the number of subband parameters that have to be jointly optimized may be large, which slows the convergence of the algorithm. To address this problem, we present two methods of sharing parameters among multiple subband filters in order to significantly reduce the number of parameters to be optimized. Both of these methods exploit the spectral smoothing that occurs in the feature extraction process, but do so in different ways. By sharing parameters in the proposed manner, we are able to obtain a significant reduction in the time to convergence of S-LIMABEAM with a minimal degradation in speech recognition accuracy. Michael L. Seltzer, Richard M. Stern |
ICASSP (1) | 1 |
| 2004 | Reconstruction of missing features for robust speech recognition
Bhiksha Raj, Michael L. Seltzer, Richard M. Stern |
Speech Commun. | 2 |
| 2004 | A Bayesian classifier for spectrographic mask estimation for missing feature speech recognition
Michael L. Seltzer, Bhiksha Raj, Richard M. Stern |
Speech Commun. | 1 |
| 2004 | Likelihood-maximizing beamforming for robust hands-free speech recognitionabstractSpeech recognition performance degrades significantly in distant-talking environments, where the speech signals can be severely distorted by additive noise and reverberation. In such environments, the use of microphone arrays has been proposed as a means of improving the quality of captured speech signals. Currently, microphone-array-based speech recognition is performed in two independent stages: array processing and then recognition. Array processing algorithms, designed for signal enhancement, are applied in order to reduce the distortion in the speech waveform prior to feature extraction and recognition. This approach assumes that improving the quality of the speech waveform will necessarily result in improved recognition performance and ignores the manner in which speech recognition systems operate. In this paper a new approach to microphone-array processing is proposed in which the goal of the array processing is not to generate an enhanced output waveform but rather to generate a sequence of features which maximizes the likelihood of generating the correct hypothesis. In this approach, called likelihood-maximizing beamforming, information from the speech recognition system itself is used to optimize a filter-and-sum beamformer. Speech recognition experiments performed in a real distant-talking environment confirm the efficacy of the proposed approach. Michael L. Seltzer, Bhiksha Raj, Richard M. Stern |
IEEE Trans. Speech Audio Process. | 1 |
| 2003 | Subband parameter optimization of microphone arrays for speech recognition in reverberant environmentsabstractWe present a new subband microphone array processing algorithm specifically designed for speech recognition applications. We previously proposed a speech recognizer-based array processing algorithm which resulted in significant improvements in recognition accuracy when the speech was corrupted by additive noise and moderate levels of reverberation. However, little improvement was achieved over conventional beamforming methods in highly reverberant environments. Subband processing has been used to improve the poor performance of LMS-type algorithms when the number of filter parameters to estimate is large and the noise is highly correlated to the speech signal, e.g. in highly reverberant environments. We apply a subband approach to a new array processing architecture in which select groups of subbands are processed jointly to maximize the likelihood of the resulting speech recognition features, as measured by the recognition system itself. By incorporating the recognizer into the filter optimization scheme we ensure that signal components important for recognition are emphasized without undue emphasis on less critical components. By utilizing a subband approach, we can effectively apply this framework to highly reverberant environments. In doing so, we are able to achieve improvements in word error rate of over 20% compared to conventional methods in highly reverberant environments. Michael L. Seltzer, Richard M. Stern |
ICASSP (1) | 1 |
| 2003 | A harmonic-model-based front end for robust speech recognitionabstractSpeech recognition accuracy degrades significantly when the speech has been corrupted by noise, especially when the system has been trained on clean speech. Many compensation algorithms have been developed which require reliable online noise estimates or a priori knowledge of the noise. In situations where such estimates or knowledge is difficult to obtain, these methods fail. We present a new robustness algorithm which avoids these problems by making no assumptions about the corrupting noise. Instead, we exploit properties inherent to the speech signal itself to denoise the recognition features. In this method, speech is decomposed into harmonic and noise-like components, which are then processed independently and recombined. By processing noise-corrupted speech in this manner we achieve significant improvements in recognition accuracy on the Aurora 2 task. 1. Michael L. Seltzer, Jasha Droppo, Alex Acero |
INTERSPEECH | 1 |
| 2003 | Speech-recognizer-based filter optimization for microphone array processingabstractConventional microphone array processing schemes used for speech recognition enhance the output waveform using optimization criteria that are independent of the recognition system. We present a new filter-and-sum array processing algorithm in which the filter parameters are calibrated to maximize recognizer likelihoods. The proposed method provides significant improvement in recognition accuracy over conventional methods. Michael L. Seltzer, Bhiksha Raj |
IEEE Signal Process. Lett. | 1 |
| 2002 | Speech recognizer-based microphone array processing for robust hands-free speech recognitionabstractWe present a new array processing algorithm for microphone array speech recognition. Conventionally, the goal of array processing is to take distorted signals captured by the array and generate a cleaner output waveform. However, speech recognition systems operate on a set of features derived from the waveform, rather than the waveform itself. The goal of an array processor used in conjunction with a recognition system is to generate a waveform which produces a set of recognition features which maximize die likelihood for the words that are spoken, rather than to minimize the waveform distortion. We propose a new array processing algorithm which maximizes the likelihood of the recognition features. This is accomplished through the use of a new objective function which utilizes information from the recognition system itself, obtained in an unsupervised manner, to optimize the parameters of a filter-and-sum array processor. Using the proposed method, improvements in word error rate of up to 36% over conventional methods are achieved on real microphone array tasks in a wide range of environments. Michael L. Seltzer, Bhiksha Raj, Richard M. Stern |
ICASSP | 1 |
| 2001 | Speech in Noisy Environments: robust automatic segmentation, feature extraction, and hypothesis combinationabstractThe first evaluation for Speech in Noisy Environments (SPINE1) was conducted by the Naval Research Labs (NRL) in August, 2000. The purpose of the evaluation was to test existing core speech recognition technologies for speech in the presence of varying types and levels of noise. In this case the noises were taken from military settings. Among the strategies used by Carnegie Mellon University's successful systems designed for this task were session-adaptive segmentation, robust mel-scale filtering for the computation of cepstra, the use of parallel front-end features and noise-compensation algorithms, and parallel hypotheses combination through word-graphs. This paper describes the motivations behind the design decisions taken for these components, supported by observations and experiments. Rita Singh, Michael L. Seltzer, Bhiksha Raj, Richard M. Stern |
ICASSP | 2 |
| 2001 | Calibration of microphone arrays for improved speech recognitionabstractWe present a new microphone array calibration algorithm specifically designed for speech recognition. Currently, microphone-array-based speech recognition is performed in two independent stages: array processing, and then recognition. Array processing algorithms designed for speech enhancement are used to process the waveforms before recognition. These systems make the assumption that the best array processing methods will result in the best recognition performance. However, recognition systems interpret a set of features extracted from the speech waveform, not the waveform itself. In our calibration method, the filter parameters of a filter-and-sum array processing scheme are optimized to maximize the likelihood of the recognition features extracted from the resulting output signal. By incorporating the speech recognition system into the design of the array processing algorithm, we are able to achieve improvements in word error rate of up to 37 % over conventional array processing methods on both simulated and actual microphone array data. 1. Michael L. Seltzer, Bhiksha Raj |
INTERSPEECH | 1 |
| 2000 | Reconstruction of damaged spectrographic features for robust speech recognitionabstractWe present two missing-feature based algorithms that recover noise-corrupted regions of spectrographic representations of speech for noise-robust speech recognition. These algorithms modify the incoming feature vector without any changes to the speech recognition system, in contrast to previously-described approaches. The first approach clusters the feature vectors representing clean speech. Missing data are recovered by estimating the spectral cluster in each analysis frame based on the uncorrupted feature values. The second approach uses MAP procedures to estimate the values of missing data elements based on their correlations with the features that are present. Both methods take into account bounds on the clean spectrogram implied by the noisy spectrogram. Large improvements in recognition accuracy are observed when these methods are used on speech corrupted by non-stationary noise when the locations of the corrupt regions of the spectrogram are known. We also present a new method of estimating the locations of corrupt regions in spectrograms that treats the problem of identifying these regions as one of Bayesian classification. This method, when used along with the best method to reconstruct them, results in recognition accuracies comparable with the best previous data compensation algorithm on speech corrupted by white noise. It also provides significant improvement on speech corrupted by music when the global SNR of the corrupted signal is known a priori. 1. Bhiksha Raj, Michael L. Seltzer, Richard M. Stern |
INTERSPEECH | 2 |
| 2000 | Classifier-based mask estimation for missing feature methods of robust speech recognitionabstractMissing feature methods of noise compensation for speech recognition operate by removing components of a spectrographic representation of speech that are considered to be corrupt, as indicated by a low signal-to-noise ratio. Recognition is either performed directly on the incomplete spectrograms or the missing components are reconstructed prior to recognition. These methods require a spectrographic mask which accurately labels the reliable and corrupt regions of the spectrogram. Current methods of mask estimation rely on assumptions about the corrupting noise such as stationarity. This is a significant drawback since the missing feature methods themselves have no such restrictions. We present a new mask estimation technique that uses a Bayesian classifier to determine the reliability of spectrographic elements. Features were designed that make no assumptions about the corrupting noise signal, but rather exploit characteristics of the speech signal itself. Missing feature compensation experiments were performed on speech corrupted by a variety of noises. In all cases, classifier-based mask estimation resulted in significantly better recognition accuracy than conventional mask estimation methods. 1. Michael L. Seltzer, Bhiksha Raj, Richard M. Stern |
INTERSPEECH | 1 |