VLDB 2026 Research / reviewers in the wild / expert
Arun Narayanan
dblp:89/9874
· DBLP profile ↗
67ranked-venue papers
16as first author
34since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 56 · 12 first-author · 30 since 2021Artificial intelligence and machine learning · 36 · 7 first-author · 15 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic deployment of UAVs for temporary networks using multi-criteria decision-makingabstractUnmanned Aerial Vehicle Base Stations (UAV-BSs) are effective to support mobile wireless systems in situations where an unusually high density of users require enhanced coverage and capacity such as in large sport events or music festivals. However, the joint deployment problem of UAV-BSs is NP-hard, and the optimization methods used to solve such a class of problems are often too slow for (quasi-)real-time applications when the density of users and the number of UAV-BS are high. This inefficiency creates a need for more adaptable methods to solve the UAV-BS positioning problem. This paper proposes a solution by transforming the UAV-BSs’ placement problem into a decision-making process using the Analytic Hierarchy Process (AHP) method. Our solution considers that all UAV-BSs move along scanning points with a predetermined path, each with a minimal distance from one another. The proposed algorithm, named UAV-AHP, efficiently determines the UAV-BSs’ positions, thereby improving the network performance based on (quasi-) real-time acquisition of user signals at scanning points. For the high-density scenarios under investigation, our numerical results demonstrate that UAV-AHP outperforms commonly used heuristics that sub-optimally solve NP-hard problems, namely, cuckoo search (CS), particle swarm optimization (PSO), and a genetic algorithm (NSGA-II). The proposed method for UAV-BS deployment requires considerably lower running times to find satisfactory solutions than CS, PSO, and NSGA-II. Flávio Henry Ferreira, Fabrício J. B. Barros, Miércio Cardoso De Alcântara Neto, Arun Narayanan, Pedro Henrique Juliano Nardelli, Jasmine Priscyla Leite De Araújo |
Ad Hoc Networks | 4 |
| 2025 | Improving Streaming ASR via Differentially Private Fusion of Data from Multiple SourcesabstractWith increasing regulatory constraints, combining potentially sensitive data from diverse distributions (or domains) to train ML models is not always possible. In such settings, domain adaptation (DA) methods are very popular. DA pretrains a base model on certain domain data and adapts it using data from each of target domains to obtain per-domain models. Unfortunately, traditional DA does not use, and hence, cannot benefit from the diverse multi-domain (MD) data and produces sub-par models. To leverage MD data, when training on combined MD data is prohibited, we propose a novel differential privacy (DP) based solution. Our DP MD training improves the quality of the base model, thereby improving the quality of the downstream DA models. With the same number of perdomain parameters, our DP-based DA produces significantly better performing per-domain models for short, medium and long speech query domains, while ensuring reasonable privacy with respect to data of each domain. Virat Shejwalkar, Om Thakkar 0001, Steve Chien, Nicole Rafidi, Arun Narayanan |
ASRU | 5 |
| 2025 | Bone Conducted Signal Guided Speech Enhancement For Voice Assistant on EarbudsabstractIn this work we present a multi-modal, streaming enhancement network to improve speech recognition for voice assistants on earbuds. The proposed model is guided by a bone conducted signal (BCS) to separate the interfering sources from the target speaker signal. We train the model on a simulated speech enhancement training set with a simulated BCS and finetune it on a small earbuds specific training set, consisting of about 6 hours of speech. To account for distorted BCS the enhancement module is complemented by a voice activity-based decision to discard the enhanced output for BCS without speech information. A possibility to preprocess the BCS to account for the low-pass characteristic of the bone conduction is evaluated to lower the required transmission bandwidth from the earbuds to the recognition device. The results show that the BCS bandwidth can be reduced to 500 Hz with only small losses in word error rate. In comparison with a larger state-of-the-art multi-channel enhancement method, the systems, with and without bandwidth reduction, demonstrate superior performance on most of the considered realistic test sets. Jens Heitkaemper, Joe Caroselli, Max McKinnon, Arun Narayanan, Nathan Howard |
ICASSP | 4 |
| 2024 | Improving Acoustic Echo Cancellation for Voice Assistants Using Neural Echo Suppression and Multi-Microphone Noise ReductionabstractKeyword spotting (KS) and automatic speech recognition (ASR) on smart speakers in a home environment with interfering signals from loudspeakers are challenging tasks to this day, despite improvements in acoustic echo cancellation (AEC) systems. In this work we propose to combine a single microphone AEC system, consisting of an adaptive linear filter (linear AEC) and a neural echo suppressor (NES), with an adaptive filter developed for multi-microphone noise reduction, called Cleaner. This additional enhancement step allows the AEC system to profit from spatial information to remove residual echo. The single microphone NES model improves upon the waveform domain counterpart proposed in [1] using a frequency domain representation that helps with generalization. Furthermore, we show that using multiple linear AEC configurations during model training provides large gains over a fixed configuration. On the hardest considered test condition, the proposed system outperforms the baseline model [1] for single microphone input by 66 % (relative) in KS false reject rate (FRR) and 52 % (relative) in ASR word error rate (WER). Using the multi-microphone setting, the FRR is reduced by an additional 52 % and the WER by an additional 32 %. Jens Heitkaemper, Arun Narayanan, Turaj Zakizadeh Shabestary, Sankaran Panchapagesan, James Walker, Bhalchandra Gajare, Shlomi Regev, Ajay Dudani, Alexander Gruenstein |
ICASSP | 2 |
| 2024 | Extreme Encoder Output Frame Rate Reduction: Improving Computational Latencies of Large End-to-End ModelsabstractThe accuracy of end-to-end (E2E) automatic speech recognition (ASR) models continues to improve as they are scaled to larger sizes, with some now reaching billions of parameters. Widespread deployment and adoption of these models, however, requires computationally efficient strategies for decoding. In the present work, we study one such strategy: applying multiple frame reduction layers in the encoder to compress encoder outputs into a small number of output frames. While similar techniques have been investigated in previous work, we achieve dramatically more reduction than has previously been demonstrated through the use of multiple funnel reduction layers. Through ablations, we study the impact of various architectural choices in the encoder to identify the most effective strategies. We demonstrate that we can generate one encoder output frame for every 2.56 sec of input speech, without significantly affecting word error rate on a large-scale voice search task, while improving encoder and decoder latencies by 48% and 92% respectively, relative to a strong but computationally expensive baseline. Rohit Prabhavalkar, Zhong Meng, Adam Stooke, Xingyu Cai, Yanzhang He, Arun Narayanan, Dongseong Hwang, Tara N. Sainath, Pedro J. Moreno 0001 |
ICASSP | 7 |
| 2024 | Efficiently Train ASR Models that Memorize Less and Perform Better with Per-core Clipping
Lun Wang 0001, Om Thakkar 0001, Zhong Meng, Nicole Rafidi, Rohit Prabhavalkar, Arun Narayanan |
INTERSPEECH | 6 |
| 2024 | TfCleanformer: A streaming, array-agnostic, full- and sub-band modeling front-end for robust ASR
Jens Heitkaemper, Joe Caroselli, Arun Narayanan, Nathan Howard |
INTERSPEECH | 3 |
| 2024 | Quantifying Unintended Memorization in BEST-RQ ASR Encoders
Virat Shejwalkar, Om Thakkar 0001, Arun Narayanan |
INTERSPEECH | 3 |
| 2024 | Training Large ASR Encoders With Differential PrivacyabstractSelf-supervised learning (SSL) methods for large speech models have proven to be highly effective at ASR. With the interest in public deployment of large pre-trained models, there is a rising concern for unintended memorization and leakage of sensitive data points from the training data. In this paper, we apply differentially private (DP) pre-training to a SOTA Conformer-based encoder, and study its performance on a downstream ASR task assuming the fine-tuning data is public. This paper is the first to apply DP to SSL for ASR, investigating the DP noise tolerance of the BEST-RQ pre-training method. Notably, we introduce a novel variant of model pruning called gradient-based layer freezing that provides strong improvements in privacy-utility-compute trade-offs. Our approach yields a LibriSpeech test-clean/other WER (%) of 3.78/ 8.41 with ($10,1 \mathrm{e}-9$)-DP for extrapolation towards low dataset scales, and 2.81/5.89 with ($10,7.9 \mathrm{e}-11$)DP for extrapolation towards high scales. Geeticka Chauhan, Steve Chien, Om Thakkar 0001, Abhradeep Thakurta, Arun Narayanan |
SLT | 5 |
| 2024 | Image-based intrusion detection system for GPS spoofing cyberattacks in unmanned aerial vehiclesabstractThe operations of unmanned aerial vehicles (UAVs) are susceptible to cybersecurity risks, mainly because of their firm reliance on the Global Positioning System (GPS) and radio frequency (RF) sensors. GPS and RF sensors are vulnerable to potential threats, such as spoofing attacks that can cause the UAVs to behave erratically. Since these threats are widespread and potent, it is imperative to develop effective intrusion detection systems. In this paper, we propose an image-based intrusion detection system for detecting GPS spoofing cyberattacks based on a deep learning methodology. We combine convolutional neural networks with Principal Component Analysis (PCA) to reduce the dimensionality of the dataset features, data augmentation to increase the size and diversity of the training dataset, and transfer learning to improve the proposed model’s performance with limited data to design a fast, accurate, and general method. Extensive numerical experiments demonstrate the effectiveness of the proposed solution carried out using benchmark datasets. We achieved an accuracy of 100% within a running time of 120.64 s at 0.3529 ms latency and a detection time of 2.035 s in the case of the training dataset. Further, using this trained model, we achieved an accuracy of 99.25% within a detection time of 2.721 s on an unseen dataset that was unrelated to the one used for training the model. In contrast, other models, such as Inception-v3, showed lower accuracy on unseen datasets. However, Inception-v3 performance improved significantly after Bayesian optimization, with the Tree-structured Parzen Estimator reaching 99.06% accuracy. Our results demonstrate that the proposed image-based intrusion detection method outperforms the existing solutions while providing a general model for detecting cyberattacks included in unseen datasets. Mohamed Selim Korium, Ahmed Mahmoud Ahmed, Arun Narayanan, Pedro Henrique Juliano Nardelli |
Ad Hoc Networks | 4 |
| 2024 | Intrusion detection system for cyberattacks in the Internet of Vehicles environmentabstractThis paper presents a novel framework for intrusion detection specially designed for cyberattacks, such as Denial-of-Service, Distributed Denial-of-Service, Distributed Reflection Denial-of-Service, Brute Force, Botnets, and Sniffing, on vehicles that are situated in the Internet of Vehicles environment. We propose an intrusion detection system based on machine learning that is capable of detecting abnormal behavior by examining network traffic to find unusual data flows. In this paper, we have presented a strategy for intrusion detection through a careful evaluation and selection of the most effective techniques for the following steps of the machine learning process: (i) data preprocessing by using Z-score normalization that preserves the data distribution for the proposed method and handles outliers; (ii) feature selection by using a regression model that simplifies the model complexity and reduces the execution time; and (iii) model selection and training – Random Forest, Extreme Gradient Boosting, Categorical Boosting, Light Gradient Boosting Machine – with hyperparameter optimization to control the behavior in the training phase and to prevent overfitting. The effectiveness of the proposed solution is demonstrated by extensive numerical experiments carried out using the well-known standard datasets CIC-IDS-2017, CSE-CIC-IDS-2018, and CIC-DDoS-2019, both separately and merged. We achieved a high accuracy above 99.8% within a running time of 46.9 s and 0.24 s detection time for the three combined intrusion detection system datasets, thereby showing that the proposed intrusion detection system outperforms the previous methods introduced in the literature. Mohamed Selim Korium, Alexander Beattie, Arun Narayanan, Subham Sahoo, Pedro Henrique Juliano Nardelli |
Ad Hoc Networks | 4 |
| 2023 | Cleanformer: A Multichannel Array Configuration-Invariant Neural Enhancement Frontend for ASR in Smart SpeakersabstractThis work introduces Cleanformer —a streaming multichannel neural enhancement frontend for automatic speech recognition (ASR). This model has a Conformer-based architecture which takes as inputs a single channel each of raw and enhanced signals, and uses self-attention to derive a time-frequency mask. The enhanced input is generated by a multichannel adaptive noise cancellation algorithm known as Speech Cleaner. The time-frequency mask is applied to the noisy input to produce enhanced features for ASR. Detailed evaluations are presented with speech- and non-speech-based noise that show significant reduction in word error rate (WER) – about 80% for -6 dB SNR – over a state-of-the-art ASR model alone. It also significantly outperforms enhancement using a beamformer with ideal steering. The enhancement model can be used with different microphone arrays without the need for retraining. Joseph Caroselli, Arun Narayanan, Nathan Howard, Tom O'Malley |
ICASSP | 2 |
| 2023 | Conditional Conformer: Improving Speaker Modulation For Single And Multi-User Speech EnhancementabstractRecently, Feature-wise Linear Modulation (FiLM) has been shown to outperform other approaches to incorporate speaker embedding into speech separation and VoiceFilter models. We propose an improved method of incorporating such embeddings into a Voice- Filter frontend for automatic speech recognition (ASR) and text- independent speaker verification (TI-SV). We extend the widely- used Conformer architecture to construct a FiLM Block with additional feature processing before and after the FiLM layers. Apart from its application to single-user VoiceFilter, we show that our system can be easily extended to multi-user VoiceFilter models via element-wise max pooling of the speaker embeddings in a projected space. The final architecture, which we call Conditional Conformer, tightly integrates the speaker embeddings into a Conformer backbone. We improve TI-SV equal error rates by as much as 56% over prior multi-user VoiceFilter models, and our element-wise max pooling reduces relative WER compared to an attention mechanism by as much as 10%. Tom O'Malley, Shaojin Ding, Arun Narayanan, Rajeev Rikhye, Qiao Liang 0001, Yanzhang He, Ian McGraw |
ICASSP | 3 |
| 2023 | On Training a Neural Residual Acoustic Echo Suppressor for Improved ASR
Sankaran Panchapagesan, Turaj Zakizadeh Shabestary, Arun Narayanan |
INTERSPEECH | 3 |
| 2022 | Transducer-Based Streaming Deliberation for Cascaded EncodersabstractPrevious research on applying deliberation networks to automatic speech recognition has achieved excellent results. The attention decoder based deliberation model often works as a rescorer to improve first-pass recognition results, and requires the full first-pass hypothesis for second-pass deliberation. In this work, we propose a transducer-based streaming deliberation model. The joint network of a transducer decoder often receives inputs from the encoder and the prediction network. We propose to use attention to the first-pass text hypothesis as the third input to the joint network. The proposed transducer based deliberation model naturally streams, making it more desirable for on-device applications. We also show that the model improves rare word recognition compared to cascaded encoders, with relative WER reductions ranging from 3.6% to 10.4% for a variety of test sets. Our model does not use any additional text data for training. Tara N. Sainath, Arun Narayanan, Ruoming Pang, Trevor Strohman |
ICASSP | 3 |
| 2022 | Improving The Latency And Quality Of Cascaded EncodersabstractIn this paper, we explore reducing computational latency of the 2-pass cascaded encoder model [1]. Specifically, we experiment with reducing the size of the causal 1st-pass and adding capacity to the non-causal 2nd-pass, such that the overall latency can be reduced without loss of quality. In addition, we explore using a confidence model for deciding to stop 2nd-pass recognition if we are confident in the 1st-pass hypothesis. Overall, we are able to reduce latency by a factor of 1.7X, compared to the baseline cascaded encoder from [1]. Secondly, with the added capacity in the non-causal 2nd-pass, we find that we can improve WER by up to 7% relative using wav2vec and minimum word-error-rate (MWER) training. Tara N. Sainath, Yanzhang He, Arun Narayanan, Rami Botros, David Qiu, Chung-Cheng Chiu, Rohit Prabhavalkar, Alexander Gruenstein, Anmol Gulati, Bo Li 0028, David Rybach, Emmanuel Guzman, Ian McGraw, James Qin, Krzysztof Choromanski, Qiao Liang 0001, Robert David 0002, Ruoming Pang, Shuo-Yiin Chang, Trevor Strohman, W. Ronny Huang, Wei Han 0002, Yu Zhang 0033 |
ICASSP | 3 |
| 2022 | Extracting Targeted Training Data from ASR Models, and How to Mitigate ItabstractRecent work has designed methods to demonstrate that model updates in ASR training can leak potentially sensitive attributes of the utterances used in computing the updates.In this work, we design the first method to demonstrate information leakage about training data from trained ASR models.We design Noise Masking, a fill-in-the-blank style method for extracting targeted parts of training data from trained ASR models.We demonstrate the success of Noise Masking by using it in four settings for extracting names from the LibriSpeech dataset used for training a state-of-the-art Conformer model.In particular, we show that we are able to extract the correct names from masked training utterances with 11.8% accuracy, while the model outputs some name from the train set 55.2% of the time.Further, we show that even in a setting that uses synthetic audio and partial transcripts from the test set, our method achieves 2.5% correct name accuracy (47.7% any name success rate).Lastly, we design Word Dropout, a data augmentation method that we show when used in training along with Multistyle TRaining (MTR), provides comparable utility as the baseline, along with significantly mitigating extraction via Noise Masking across the four evaluated settings. Ehsan Amid, Om Thakkar 0001, Arun Narayanan, Rajiv Mathews, Françoise Beaufays |
INTERSPEECH | 3 |
| 2022 | Personal VAD 2.0: Optimizing Personal Voice Activity Detection for On-Device Speech RecognitionabstractPersonalization of on-device speech recognition (ASR) has seen explosive growth in recent years, largely due to the increasing popularity of personal assistant features on mobile devices and smart home speakers. In this work, we present Personal VAD 2.0, a personalized voice activity detector that detects the voice activity of a target speaker, as part of a streaming on-device ASR system. Although previous proof-of-concept studies have validated the effectiveness of Personal VAD, there are still several critical challenges to address before this model can be used in production: first, the quality must be satisfactory in both enrollment and enrollment-less scenarios; second, it should operate in a streaming fashion; and finally, the model size should be small enough to fit a limited latency and CPU/Memory budget. To meet the multi-faceted requirements, we propose a series of novel designs: 1) advanced speaker embedding modulation methods; 2) a new training paradigm to generalize to enrollment-less conditions; 3) architecture and runtime optimizations for latency and resource restrictions. Extensive experiments on a realistic speech recognition system demonstrated the state-of-the-art performance of our proposed method. Shaojin Ding, Rajeev Rikhye, Qiao Liang 0001, Yanzhang He, Arun Narayanan, Tom O'Malley, Ian McGraw |
INTERSPEECH | 6 |
| 2022 | SNRi Target Training for Joint Speech Enhancement and Recognition
Yuma Koizumi, Shigeki Karita, Arun Narayanan, Sankaran Panchapagesan, Michiel Bacchiani |
INTERSPEECH | 3 |
| 2022 | A universally-deployable ASR frontend for joint acoustic echo cancellation, speech enhancement, and voice separationabstractRecent work has shown that it is possible to train a single model to perform joint acoustic echo cancellation (AEC), speech enhancement, and voice separation, thereby serving as a unified frontend for robust automatic speech recognition (ASR).The joint model uses contextual information, such as a reference of the playback audio, noise context, and speaker embedding.In this work, we propose a number of novel improvements to such a model.First, we improve the architecture of the Cross-Attention Conformer that is used to ingest noise context into the model.Second, we generalize the model to be able to handle varying lengths of noise context.Third, we propose Signal Dropout, a novel strategy that models missing contextual information.In the absence of one or more signals, the proposed model performs nearly as well as task-specific models trained without these signals; and when such signals are present, our system compares well against systems that require all context signals.Over the baseline, the final model retains a relative word error rate reduction of 25.0% on background speech when speaker embedding is absent, and 61.2% on AEC when device playback is absent. Thomas R. O'Malley, Arun Narayanan |
INTERSPEECH | 2 |
| 2022 | A Conformer-based Waveform-domain Neural Acoustic Echo Canceller Optimized for ASR AccuracyabstractAcoustic Echo Cancellation (AEC) is essential for accurate recognition of queries spoken to a smart speaker that is playing out audio.Previous work has shown that a neural AEC model operating on log-mel spectral features (denoted "logmel" hereafter) can greatly improve Automatic Speech Recognition (ASR) accuracy when optimized with an auxiliary loss utilizing a pre-trained ASR model encoder.In this paper, we develop a conformer-based waveform-domain neural AEC model inspired by the "TasNet" architecture.The model is trained by jointly optimizing Negative Scale-Invariant SNR (SISNR) and ASR losses on a large speech dataset.On a realistic rerecorded test set, we find that cascading a linear adaptive AEC and a waveform-domain neural AEC is very effective, giving 56-59% word error rate (WER) reduction over the linear AEC alone.On this test set, the 1.6M parameter waveform-domain neural AEC also improves over a larger 6.5M parameter logmeldomain neural AEC model by 20-29% in easy to moderate conditions.By operating on smaller frames, the waveform neural model is able to perform better at smaller sizes and is better suited for applications where memory is limited. Sankaran Panchapagesan, Arun Narayanan, Turaj Zakizadeh Shabestary, Nathan Howard, Alex Park 0001, James Walker, Alexander Gruenstein |
INTERSPEECH | 2 |
| 2022 | Learning Mask Scalars for Improved Robust Automatic Speech RecognitionabstractImproving robustness of streaming automatic speech recognition (ASR) systems using neural network based acoustic frontends is challenging because of causality constraints and the speech-distortions introduced by the frontend. Time-frequency masking based approaches are commonly used, but they need additional hyperparameters – mask scalars – to limit distortion. Mask scalars are typically hand-tuned and chosen conservatively. In this work, we present a technique to predict mask scalars using ASR loss in an end-to-end fashion, with minimal increase in model size and complexity. We evaluate the approach on two robust ASR tasks: multichannel enhancement in the presence of speech and non-speech noise, and acoustic echo cancellation (AEC). Results show that the presented algorithm consistently improves word error rate (WER) over strong baselines that use hand-tuned hyperparameters: up to 16% in noisy conditions, and up to 7% for AEC. Arun Narayanan, James Walker, Sankaran Panchapagesan, Nathan Howard, Yuma Koizumi |
SLT | 1 |
| 2022 | Performance evaluation of machine learning for fault selection in power transmission linesabstractAbstract Learning methods have been increasingly used in power engineering to perform various tasks. In this paper, a fault selection procedure in double-circuit transmission lines employing different learning methods is accordingly proposed. In the proposed procedure, the discrete Fourier transform (DFT) is used to pre-process raw data from the transmission line before it is fed into the learning algorithm, which will detect and classify any fault based on a training period. The performance of different machine learning algorithms is then numerically compared through simulations. The comparison indicates that an artificial neural network (ANN) achieves remarkable accuracy of 98.47%. As a drawback, the ANN method cannot provide explainable results and is also not robust against noisy measurements. Subsequently, it is demonstrated that explainable results can be obtained with high accuracy by using rule-based learners such as the recently developed quantitative association rule mining algorithm (QARMA). The QARMA algorithm outperforms other explainable schemes, while attaining an accuracy of 98%. Besides, it was shown that QARMA leads to a very high accuracy of 97% for highly noisy data. The proposed method was also validated using data from an actual transmission line fault. In summary, the proposed two-step procedure using the DFT combined with either deep learning or rule-based algorithms can accurately and successfully perform fault selection tasks but indicating remarkable advantages of the QARMA due to its explainability and robustness against noise. Those aspects are extremely important if machine learning and other data-driven methods are to be employed in critical engineering applications. Daniel Gutierrez-Rojas, Ioannis T. Christou, Daniel Dantas, Arun Narayanan, Pedro Henrique Juliano Nardelli, Yongheng Yang |
Knowl. Inf. Syst. | 4 |
| 2021 | Cross-Attention Conformer for Context Modeling in Speech Enhancement for ASRabstractThis work introduces cross-attention conformer, an attention-based architecture for context modeling in speech enhancement. Given that the context information can often be sequential, and of different length as the audio that is to be enhanced, we make use of cross-attention to summarize and merge contextual information with input features. Building upon the recently proposed conformer model that uses self attention layers as building blocks, the proposed cross-attention conformer can be used to build deep contextual models. As a concrete example, we show how noise context, i.e., short noise-only audio segment preceding an utterance, can be used to build a speech enhancement feature frontend using cross-attention conformer layers for improving noise robustness of automatic speech recognition. Arun Narayanan, Chung-Cheng Chiu, Tom O'Malley, Yanzhang He |
ASRU | 1 |
| 2021 | A Conformer-Based ASR Frontend for Joint Acoustic Echo Cancellation, Speech Enhancement and Speech SeparationabstractWe present a frontend for improving robustness of automatic speech recognition (ASR), that jointly implements three modules within a single model: acoustic echo cancellation, speech enhancement, and speech separation. This is achieved by using a contextual enhancement neural network that can optionally make use of different types of side inputs: (1) a reference signal of the playback audio, which is necessary for echo cancellation; (2) a noise context, which is useful for speech enhancement; and (3) an embedding vector representing the voice characteristic of the target speaker of interest, which is not only critical in speech separation, but also helpful for echo cancellation and speech enhancement. We present detailed evaluations to show that the joint model performs almost as well as the task-specific models, and significantly reduces word error rate in noisy conditions even when using a large-scale state-of-the-art ASR model. Compared to the noisy baseline, the joint model reduces the word error rate in low signal-to-noise ratio conditions by at least 71% on our echo cancellation dataset, 10% on our noisy dataset, and 26% on our multi-speaker dataset. Compared to task-specific models, the joint model performs within 10% on our echo cancellation dataset, 2% on the noisy dataset, and 3% on the multi-speaker dataset. Tom O'Malley, Arun Narayanan, Alex Park 0001, James Walker, Nathan Howard |
ASRU | 2 |
| 2021 | Improving Streaming Automatic Speech Recognition with Non-Streaming Model Distillation on Unsupervised DataabstractStreaming end-to-end automatic speech recognition (ASR) models are widely used on smart speakers and on-device applications. Since these models are expected to transcribe speech with minimal latency, they are constrained to be causal with no future context, compared to their non-streaming counterparts. Consequently, streaming models usually perform worse than non-streaming models. We propose a novel and effective learning method by leveraging a non-streaming ASR model as a teacher to generate transcripts on an arbitrarily large data set, which is then used to distill knowledge into streaming ASR models. This way, we scale the training of streaming models to up to 3 million hours of YouTube audio. Experiments show that our approach can significantly reduce the word error rate (WER) of RNN-T models not only on LibriSpeech but also on YouTube data in four languages. For example, in French, we are able to reduce the WER by 16.4% relatively to a baseline streaming model by leveraging a non-streaming teacher model trained on the same amount of labeled data as the baseline. Thibault Doutre, Wei Han 0002, Zhiyun Lu, Chung-Cheng Chiu, Ruoming Pang, Arun Narayanan, Ananya Misra, Yu Zhang 0033, Liangliang Cao |
ICASSP | 7 |
| 2021 | A Better and Faster end-to-end Model for Streaming ASRabstractEnd-to-end (E2E) models have shown to outperform state-of-the-art conventional models for streaming speech recognition [1] across many dimensions, including quality (as measured by word error rate (WER)) and endpointer latency [2]. However, the model still tends to delay the predictions towards the end and thus has much higher partial latency compared to a conventional ASR model. To address this issue, we look at encouraging the E2E model to emit words early, through an algorithm called FastEmit [3]. Naturally, improving on latency results in a quality degradation. To address this, we explore replacing the LSTM layers in the encoder of our E2E model with Conformer layers [4], which has shown good improvements for ASR. Secondly, we also explore running a 2nd-pass beam search to improve quality. In order to ensure the 2nd-pass completes quickly, we explore non-causal Conformer layers that feed into the same 1st-pass RNN-T decoder, an algorithm called Cascaded Encoders [5]. Overall, the Conformer RNN-T with Cascaded Encoders offers a better quality and latency tradeoff for streaming ASR. Bo Li 0028, Anmol Gulati, Tara N. Sainath, Chung-Cheng Chiu, Arun Narayanan, Shuo-Yiin Chang, Ruoming Pang, Yanzhang He, James Qin, Wei Han 0002, Qiao Liang 0001, Yu Zhang 0033, Trevor Strohman |
ICASSP | 6 |
| 2021 | Cascaded Encoders for Unifying Streaming and Non-Streaming ASRabstractEnd-to-end (E2E) automatic speech recognition (ASR) models, by now, have shown competitive performance on several benchmarks. These models are structured to either operate in streaming or non-streaming mode. This work presents cascaded encoders for building a single E2E ASR model that can operate in both these modes simultaneously. The proposed model consists of streaming and non-streaming encoders. Input features are first processed by the streaming encoder; the non-streaming encoder operates exclusively on the output of the streaming encoder. A single decoder then learns to decode either using the output of the streaming or the non-streaming encoder. Results show that this model achieves similar word error rates (WER) as a standalone streaming model when operating in streaming mode, and obtains 10% – 27% relative improvement when operating in non-streaming mode. Our results also show that the proposed approach outperforms existing E2E two-pass models, especially on long-form speech. Arun Narayanan, Tara N. Sainath, Ruoming Pang, Chung-Cheng Chiu, Rohit Prabhavalkar, Ehsan Variani, Trevor Strohman |
ICASSP | 1 |
| 2021 | Less is More: Improved RNN-T Decoding Using Limited Label Context and Path MergingabstractEnd-to-end models that condition the output sequence on all previously predicted labels have emerged as popular alternatives to conventional systems for automatic speech recognition (ASR). Since distinct label histories correspond to distinct models states, such models are decoded using an approximate beam-search which produces a tree of hypotheses.In this work, we study the influence of the amount of label context on the model’s accuracy, and its impact on the efficiency of the decoding process. We find that we can limit the context of the recurrent neural network transducer (RNN-T) during training to just four previous word-piece labels, without degrading word error rate (WER) relative to the full-context baseline. Limiting context also provides opportunities to improve decoding efficiency by removing redundant paths from the active beam, and instead retaining them in the final lattice. This path-merging scheme can also be applied when decoding the baseline full-context model through an approximation. Overall, we find that the proposed path-merging scheme is extremely effective, allowing us to improve oracle WERs by up to 36% over the baseline, while simultaneously reducing the number of model evaluations by up to 5.3% without any degradation in WER, or up to 15.7% when lattice rescoring is applied. Rohit Prabhavalkar, Yanzhang He, David Rybach, Sean Campbell, Arun Narayanan, Trevor Strohman, Tara N. Sainath |
ICASSP | 5 |
| 2021 | FastEmit: Low-Latency Streaming ASR with Sequence-Level Emission RegularizationabstractStreaming automatic speech recognition (ASR) aims to emit each hypothesized word as quickly and accurately as possible. However, emitting fast without degrading quality, as measured by word error rate (WER), is highly challenging. Existing approaches including Early and Late Penalties [1] and Constrained Alignments [2], [3] penalize emission delay by manipulating per-token or per-frame probability prediction in sequence transducer models [4]. While being successful in reducing delay, these approaches suffer from significant accuracy regression and also require additional word alignment information from an existing model. In this work, we propose a sequence-level emission regularization method, named FastEmit, that applies latency regularization directly on per-sequence probability in training transducer models, and does not require any alignment. We demonstrate that FastEmit is more suitable to the sequence-level optimization of transducer models [4] for streaming ASR by applying it on various end-to-end streaming ASR networks including RNN-Transducer [5], Transformer-Transducer [6], [7], ConvNet-Transducer [8] and Conformer-Transducer [9]. We achieve 150 ~ 300ms latency reduction with significantly better accuracy over previous techniques on a Voice Search test set. FastEmit also improves streaming ASR accuracy from 4.4%/8.9% to 3.1%/7.5% WER, meanwhile reduces 90th percentile latency from 210ms to only 30ms on LibriSpeech. Chung-Cheng Chiu, Bo Li 0028, Shuo-Yiin Chang, Tara N. Sainath, Yanzhang He, Arun Narayanan, Wei Han 0002, Anmol Gulati, Ruoming Pang |
ICASSP | 7 |
| 2021 | A Comparison of Supervised and Unsupervised Pre-Training of End-to-End Models
Ananya Misra, Dongseong Hwang, Zhouyuan Huo, Shefali Garg, Nikhil Siddhartha, Arun Narayanan, Khe Chai Sim |
Interspeech | 6 |
| 2021 | Personalized Keyphrase Detection Using Speaker and Environment InformationabstractIn this paper, we introduce a streaming keyphrase detection system that can be easily customized to accurately detect any phrase composed of words from a large vocabulary. The system is implemented with an end-to-end trained automatic speech recognition (ASR) model and a text-independent speaker verification model. To address the challenge of detecting these keyphrases under various noisy conditions, a speaker separation model is added to the feature frontend of the speaker verification model, and an adaptive noise cancellation (ANC) algorithm is included to exploit cross-microphone noise coherence. Our experiments show that the text-independent speaker verification model largely reduces the false triggering rate of the keyphrase detection, while the speaker separation model and adaptive noise cancellation largely reduce false rejections. Rajeev Rikhye, Qiao Liang 0001, Yanzhang He, Ding Zhao, Yiteng Huang, Arun Narayanan, Ian McGraw |
Interspeech | 7 |
| 2021 | An Efficient Streaming Non-Recurrent On-Device End-to-End Model with Improvements to Rare-Word Modeling
Tara N. Sainath, Yanzhang He, Arun Narayanan, Rami Botros, Ruoming Pang, David Rybach, Cyril Allauzen, Ehsan Variani, James Qin, Quoc-Nam Le-The, Shuo-Yiin Chang, Bo Li 0028, Anmol Gulati, Chung-Cheng Chiu, Diamantino Caseiro, Wei Li 0133, Qiao Liang 0001, Pat Rondon |
Interspeech | 3 |
| 2021 | RNN-T Models Fail to Generalize to Out-of-Domain Audio: Causes and SolutionsabstractIn recent years, all-neural end-to-end approaches have obtained state-of-the-art results on several challenging automatic speech recognition (ASR) tasks. However, most existing works focus on building ASR models where train and test data are drawn from the same domain. This results in poor generalization characteristics on mismatched-domains: e.g., end-to-end models trained on short segments perform poorly when evaluated on longer utterances. In this work, we analyze the generalization properties of streaming and non-streaming recurrent neural network transducer (RNN-T) based end-to-end models in order to identify model components that negatively affect generalization performance. We propose two solutions: combining multiple regularization techniques during training, and using dynamic overlapping inference. On a long-form YouTube test set, when the non-streaming RNN-T model is trained with shorter segments of data, the proposed combination improves word error rate (WER) from 22.3% to 14.8%; when the streaming RNN-T model trained on short Search queries, the proposed techniques improve WER on the YouTube set from 67.0% to 25.3%. Finally, when trained on Librispeech, we find that dynamic overlapping inference improves WER on YouTube from 99.8% to 33.0%. Chung-Cheng Chiu, Arun Narayanan, Wei Han 0002, Rohit Prabhavalkar, Yu Zhang 0033, Navdeep Jaitly, Ruoming Pang, Tara N. Sainath, Patrick Nguyen, Liangliang Cao |
SLT | 2 |
| 2020 | A Streaming On-Device End-To-End Model Surpassing Server-Side Conventional Model Quality and LatencyabstractThus far, end-to-end (E2E) models have not been shown to outperform state-of-the-art conventional models with respect to both quality, i.e., word error rate (WER), and latency, i.e., the time the hypothesis is finalized after the user stops speaking. In this paper, we develop a first-pass Recurrent Neural Network Transducer (RNN-T) model and a second-pass Listen, Attend, Spell (LAS) rescorer that surpasses a conventional model in both quality and latency. On the quality side, we incorporate a large number of utterances across varied domains [1] to increase acoustic diversity and the vocabulary seen by the model. We also train with accented English speech to make the model more robust to different pronunciations. In addition, given the increased amount of training data, we explore a varied learning rate schedule. On the latency front, we explore using the end-of-sentence decision emitted by the RNN-T model to close the microphone, and also introduce various optimizations to improve the speed of LAS rescoring. Overall, we find that RNN-T+LAS offers a better WER and latency tradeoff compared to a conventional model. For example, for the same latency, RNN-T+LAS obtains a 8% relative improvement in WER, while being more than 400-times smaller in model size. Tara N. Sainath, Yanzhang He, Bo Li 0028, Arun Narayanan, Ruoming Pang, Antoine Bruguier, Shuo-Yiin Chang, Wei Li 0133, Raziel Alvarez, Chung-Cheng Chiu, Alexander Gruenstein, Anjuli Kannan, Qiao Liang 0001, Ian McGraw, Cal Peyser, Rohit Prabhavalkar, Golan Pundak, David Rybach, Yuan Shangguan, Yash Sheth, Trevor Strohman, Mirkó Visontai, Yu Zhang 0033, Ding Zhao |
ICASSP | 4 |
| 2020 | Anti-Aliasing Regularization in Stacking Layers
Antoine Bruguier, Ananya Misra, Arun Narayanan, Rohit Prabhavalkar |
INTERSPEECH | 3 |
| 2020 | Profit Allocation in Renewables Based Community Microgrids with Aggregation and Self-SufficiencyabstractPeer-to-peer (p2p) electricity exchange can reduce energy wastage and improve capacity utilization in microgrids with renewable energy based electricity production. In such microgrids, called community microgrids, an important problem is to enable customers to share and transact their energy in a rational and acceptable manner. In this paper, electricity exchanges are modeled using co-operative game-theoretic concepts. We assume that all the participants of a community microgrid, which is connected to an external grid, co-operate to form a single (grand) coalition. Using the concept of marginal contributions (MC), we had previously proposed a fair methodology to distribute the total profits of the coalition among its participants. Here, we extend the methodology by considering two scenarios. First, we analyze the case where the community microgrid sometimes collectively produces more electricity than its total load and sells this excess production to an aggregator. Secondly, we consider self-sufficiency as an important objective of the community microgrid. We then use MC to derive a new, simple, and scalable allocation formula to distribute the profits among the participants fairly. To demonstrate our methodology, we applied it to a microgrid in Austin, Texas, USA. The newly proposed methodology saved ≈4.55% as compared to the old originally proposed methodology. Arun Narayanan, Pedro Henrique Juliano Nardelli |
PIMRC | 1 |
| 2019 | A Comparison of End-to-End Models for Long-Form Speech RecognitionabstractEnd-to-end automatic speech recognition (ASR) models, including both attention-based models and the recurrent neural network transducer (RNN-T), have shown superior performance compared to conventional systems [1], [2]. However, previous studies have focused primarily on short utterances that typically last for just a few seconds or, at most, a few tens of seconds. Whether such architectures are practical on long utterances that last from minutes to hours remains an open question. In this paper, we both investigate and improve the performance of end-to-end models on long-form transcription. We first present an empirical comparison of different end-to-end models on a real world long-form task and demonstrate that the RNN-T model is much more robust than attention-based systems in this regime. We next explore two improvements to attention-based systems that significantly improve its performance: restricting the attention to be monotonic, and applying a novel decoding algorithm that breaks long utterances into shorter overlapping segments. Combining these two improvements, we show that attention-based end-to-end models can be very competitive to RNN-T on long-form speech recognition. Chung-Cheng Chiu, Anjuli Kannan, Rohit Prabhavalkar, Tara N. Sainath, Wei Han 0002, Yu Zhang 0033, Ruoming Pang, Sergey Kishchenko, Patrick Nguyen, Arun Narayanan, Hank Liao, Shuyuan Zhang 0002 |
ASRU | 12 |
| 2019 | Recognizing Long-Form Speech Using Streaming End-to-End ModelsabstractAll-neural end-to-end (E2E) automatic speech recognition (ASR) systems that use a single neural network to transduce audio to word sequences have been shown to achieve state-of-the-art results on several tasks. In this work, we examine the ability of E2E models to generalize to unseen domains, where we find that models trained on short utterances fail to generalize to long-form speech. We propose two complementary solutions to address this: training on diverse acoustic data, and LSTM state manipulation to simulate long-form audio when training using short utterances. On a synthesized long-form test set, adding data diversity improves word error rate (WER) by 90% relative, while simulating long-form training improves it by 67% relative, though the combination doesn't improve over data diversity alone. On a real long-form call-center test set, adding data diversity improves WER by 40% relative. Simulating long-form training on top of data diversity improves performance by an additional 27% relative. Arun Narayanan, Rohit Prabhavalkar, Chung-Cheng Chiu, David Rybach, Tara N. Sainath, Trevor Strohman |
ASRU | 1 |
| 2018 | Spectral Distortion Model for Training Phase-Sensitive Deep-Neural Networks for Far-Field Speech RecognitionabstractIn this paper, we present an algorithm which introduces phase-perturbation to the training database when training phase-sensitive deep neural-network models. Traditional features such as log-mel or cepstral features do not have have any phase-relevant information. However features such as raw-waveform or complex spectra features contain phase-relevant information. Phase-sensitive features have the advantage of being able to detect differences in time of arrival across different microphone channels or frequency bands. However, compared to magnitude-based features, phase information is more sensitive to various kinds of distortions such as variations in microphone characteristics, reverberation, and so on. For traditional magnitude-based features, it is widely known that adding noise or reverberation, often called Multistyle-TRaining (MTR), improves robustness. In a similar spirit, we propose an algorithm which introduces spectral distortion to make the deep-learning models more robust to phase-distortion. We call this approach Spectral-Distortion TRaining (SDTR). In our experiments using a training set consisting of 22-million utterances with and without MTR, this approach reduces Word Error Rates (WERs) relatively by 3.2 % and 8.48 % respectively on test sets recorded on Google Home. Chanwoo Kim 0001, Tara N. Sainath, Arun Narayanan, Ananya Misra, Rajeev C. Nongpiur, Michiel Bacchiani |
ICASSP | 3 |
| 2018 | Efficient Implementation of the Room Simulator for Training Deep Neural Network Acoustic ModelsabstractIn this paper, we describe how to efficiently implement an acoustic room simulator to generate large-scale simulated data for training deep neural networks.Even though Google Room Simulator in [1] was shown to be quite effective in reducing the Word Error Rates (WERs) for far-field applications by generating simulated far-field training sets, it requires a very large number of FFTs.Room Simulator used approximately 80 % of CPU usage in our CPU/GPU training architecture [2].In this work, we implement an efficient OverLap Addition (OLA) based filtering using the open-source FFTW3 library.Further, we investigate the effects of the Room Impulse Response (RIR) lengths.Experimentally, we conclude that we can cut the tail portions of RIRs whose power is less than 20 dB below the maximum power without sacrificing the speech recognition accuracy.However, we observe that cutting RIR tail more than this threshold harms the speech recognition accuracy for rerecorded test sets.Using these approaches, we were able to reduce CPU usage for the room simulator portion down to 9.69 % in CPU/GPU training architecture.Profiling result shows that we obtain 22.4 times speed-up on a single machine and 37.3 times speed up on Google's distributed training infrastructure. Chanwoo Kim 0001, Ehsan Variani, Arun Narayanan, Michiel Bacchiani |
INTERSPEECH | 3 |
| 2018 | Domain Adaptation Using Factorized Hidden Layer for Robust Automatic Speech Recognition
Khe Chai Sim, Arun Narayanan, Ananya Misra, Anshuman Tripathi, Golan Pundak, Tara N. Sainath, Parisa Haghani, Bo Li 0028, Michiel Bacchiani |
INTERSPEECH | 2 |
| 2018 | From Audio to Semantics: Approaches to End-to-End Spoken Language UnderstandingabstractConventional spoken language understanding systems consist of two main components: an automatic speech recognition module that converts audio to a transcript, and a natural language understanding module that transforms the resulting text (or top N hypotheses) into a set of domains, intents, and arguments. These modules are typically optimized independently. In this paper, we formulate audio to semantic understanding as a sequence-to-sequence problem [1]. We propose and compare various encoder-decoder based approaches that optimize both modules jointly, in an end-to-end manner. Evaluations on a real-world task show that 1) having an intermediate text representation is crucial for the quality of the predicted semantics, especially the intent arguments and 2) jointly optimizing the full system improves overall accuracy of prediction. Compared to independently trained models, our best jointly trained model achieves similar domain and intent prediction F1 scores, but improves argument word error rate by 18% relative. Parisa Haghani, Arun Narayanan, Michiel Bacchiani, Galen Chuang, Neeraj Gaur, Pedro J. Moreno 0001, Rohit Prabhavalkar, Zhongdi Qu, Austin Waters |
SLT | 2 |
| 2018 | Toward Domain-Invariant Speech Recognition via Large Scale TrainingabstractCurrent state-of-the-art automatic speech recognition systems are trained to work in specific `domains', defined based on factors like application, sampling rate and codec. When such recognizers are used in conditions that do not match the training domain, performance significantly drops. This work explores the idea of building a single domain-invariant model for varied use-cases by combining large scale training data from multiple application domains. Our final system is trained using 162,000 hours of speech. Additionally, each utterance is artificially distorted during training to simulate effects like background noise, codec distortion, and sampling rates. Our results show that, even at such a scale, a model thus trained works almost as well as those fine-tuned to specific subsets: A single model can be robust to multiple application domains, and variations like codecs and noise. More importantly, such models generalize better to unseen conditions and allow for rapid adaptation - we show that by using as little as 10 hours of data from a new domain, an adapted domain-invariant model can match performance of a domain-specific model trained from scratch using 70 times as much data. We also highlight some of the limitations of such models and areas that need addressing in future work. Arun Narayanan, Ananya Misra, Khe Chai Sim, Golan Pundak, Anshuman Tripathi, Mohamed G. Elfeky, Parisa Haghani, Trevor Strohman, Michiel Bacchiani |
SLT | 1 |
| 2017 | Improving the efficiency of forward-backward algorithm using batched computation in TensorFlowabstractSequence-level losses are commonly used to train deep neural network acoustic models for automatic speech recognition. The forward-backward algorithm is used to efficiently compute the gradients of the sequence loss with respect to the model parameters. Gradient-based optimization is used to minimize these losses. Recent work has shown that the forward-backward algorithm can be efficiently implemented as a series of matrix operations. This paper further improves the forward-backward algorithm via batched computation, a technique commonly used to improve training speed by exploiting the parallel computation of matrix multiplication. Specifically, we show how batched computation of the forward-backward algorithm can be efficiently implemented using TensorFlow to handle variable-length sequences within a mini batch. Furthermore, we also show how the batched forward-backward computation can be used to compute the gradients of the connectionist temporal classification (CTC) and maximum mutual information (MMI) losses with respect to the logits. We show, via empirical benchmarks, that the batched forward-backward computation can speed up the CTC loss and gradient computation by about 183 times when run on GPU with a batch size of 256 compared to using a batch size of 1; and by about 22 times for lattice-free MMI using a trigram phone language model for the denominator. Khe Chai Sim, Arun Narayanan, Tom Bagby, Tara N. Sainath, Michiel Bacchiani |
ASRU | 2 |
| 2017 | Adaptive Multichannel Dereverberation for Automatic Speech RecognitionabstractUtilizing an adaptive multichannel technique to mitigate reverberation present in received audio signals, prior to providing corresponding audio data to one or more additional component(s), such as automatic speech recognition (ASR) components. Implementations disclosed herein are “adaptive”, in that they utilize a filter, in the reverberation mitigation, that is online, causal and varies depending on characteristics of the input. Implementations disclosed herein are “multichannel”, in that a corresponding audio signal is received from each of multiple audio transducers (also referred to herein as “microphones”) of a client device, and the multiple audio signals (e.g., frequency domain representations thereof) are utilized in updating of the filter—and dereverberation occurs for audio data corresponding to each of the audio signals (e.g., frequency domain representations thereof) prior to the audio data being provided to ASR component(s) and/or other component(s). Joe Caroselli, Izhak Shafran, Arun Narayanan, Richard Rose |
INTERSPEECH | 3 |
| 2017 | Generation of Large-Scale Simulated Utterances in Virtual Rooms to Train Deep-Neural Networks for Far-Field Speech Recognition in Google Home
Chanwoo Kim 0001, Ananya Misra, Kean K. Chin, Thad Hughes, Arun Narayanan, Tara N. Sainath, Michiel Bacchiani |
INTERSPEECH | 5 |
| 2017 | Acoustic Modeling for Google Home
Bo Li 0028, Tara N. Sainath, Arun Narayanan, Joe Caroselli, Michiel Bacchiani, Ananya Misra, Izhak Shafran, Hasim Sak, Golan Pundak, Kean K. Chin, Khe Chai Sim, Ron J. Weiss, Kevin W. Wilson, Ehsan Variani, Chanwoo Kim 0001, Olivier Siohan, Mitch Weintraub, Erik McDermott, Richard Rose, Matt Shannon |
INTERSPEECH | 3 |
| 2017 | Annealed f-Smoothing as a Mechanism to Speed up Neural Network Training
Tara N. Sainath, Vijayaditya Peddinti, Olivier Siohan, Arun Narayanan |
INTERSPEECH | 4 |
| 2017 | An Efficient Phone N-Gram Forward-Backward Computation Using Dense Matrix Multiplication
Khe Chai Sim, Arun Narayanan |
INTERSPEECH | 2 |
| 2017 | Multichannel Signal Processing With Deep Neural Networks for Automatic Speech RecognitionabstractMultichannel automatic speech recognition (ASR) systems commonly separate speech enhancement, including localization, beamforming, and postfiltering, from acoustic modeling. In this paper, we perform multichannel enhancement jointly with acoustic modeling in a deep neural network framework. Inspired by beamforming, which leverages differences in the fine time structure of the signal at different microphones to filter energy arriving from different directions, we explore modeling the raw time-domain waveform directly. We introduce a neural network architecture, which performs multichannel filtering in the first layer of the network, and show that this network learns to be robust to varying target speaker direction of arrival, performing as well as a model that is given oracle knowledge of the true target speaker direction. Next, we show how performance can be improved by factoring the first layer to separate the multichannel spatial filtering operation from a single channel filterbank which computes a frequency decomposition. We also introduce an adaptive variant, which updates the spatial filter coefficients at each time frame based on the previous inputs. Finally, we demonstrate that these approaches can be implemented more efficiently in the frequency domain. Overall, we find that such multichannel neural networks give a relative word error rate improvement of more than 5% compared to a traditional beamforming-based multichannel ASR system and more than 10% compared to a single channel waveform model. Tara N. Sainath, Ron J. Weiss, Kevin W. Wilson, Bo Li 0028, Arun Narayanan, Ehsan Variani, Michiel Bacchiani, Izhak Shafran, Andrew W. Senior, Kean K. Chin, Ananya Misra, Chanwoo Kim 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2016 | Factored spatial and spectral multichannel raw waveform CLDNNsabstractMultichannel ASR systems commonly separate speech enhancement, including localization, beamforming and postfiltering, from acoustic modeling. Recently, we explored doing multichannel enhancement jointly with acoustic modeling, where beamforming and frequency decomposition was folded into one layer of the neural network [1, 2]. In this paper, we explore factoring these operations into separate layers in the network. Furthermore, we explore using multi-task learning (MTL) as a proxy for postfiltering, where we train the network to predict "clean" features as well as context-dependent states. We find that with the factored architecture, we can achieve a 10% relative improvement in WER over a single channel and a 5% relative improvement over the unfactored model from [1] on a 2,000-hour Voice Search task. In addition, by incorporating MTL, we can achieve 11% and 7% relative improvements over single channel and unfactored multichannel models, respectively. Tara N. Sainath, Ron J. Weiss, Kevin W. Wilson, Arun Narayanan, Michiel Bacchiani |
ICASSP | 4 |
| 2016 | Reducing the Computational Complexity of Multimicrophone Acoustic Models with Integrated Feature Extraction
Tara N. Sainath, Arun Narayanan, Ron J. Weiss, Ehsan Variani, Kevin W. Wilson, Michiel Bacchiani, Izhak Shafran |
INTERSPEECH | 2 |
| 2015 | Speaker location and microphone spacing invariant acoustic modeling from raw multichannel waveformsabstractMultichannel ASR systems commonly use separate modules to perform speech enhancement and acoustic modeling. In this paper, we present an algorithm to do multichannel enhancement jointly with the acoustic model, using a raw waveform convolutional LSTM deep neural network (CLDNN). We will show that our proposed method offers ~5% relative improvement in WER over a log-mel CLDNN trained on multiple channels. Analysis shows that the proposed network learns to be robust to varying angles of arrival for the target speaker, and performs as well as a model that is given oracle knowledge of the true location. Finally, we show that training such a network on inputs captured using multiple (linear) array configurations results in a model that is robust to a range of microphone spacings. Tara N. Sainath, Ron J. Weiss, Kevin W. Wilson, Arun Narayanan, Michiel Bacchiani, Andrew W. Senior |
ASRU | 4 |
| 2015 | Large-scale, sequence-discriminative, joint adaptive training for masking-based robust ASRabstractRecently, it was shown that the performance of supervised timefrequency masking based robust automatic speech recognition techniques can be improved by training them jointly with the acoustic model [1]. The system in [1], termed deep neural network based joint adaptive training, used fully-connected feedforward deep neural networks for estimating time-frequency masks and for acoustic modeling; stacked log mel spectra was used as features and training minimized cross entropy loss. In this work, we extend such jointly trained systems in several ways. First, we use recurrent neural networks based on long short-term memory (LSTM) units – this allows the use of unstacked features, simplifying joint optimization. Next, we use a sequence discriminative training criterion for optimizing parameters. Finally, we conduct experiments on large scale data and show that joint adaptive training can provide gains over a strong baseline. Systematic evaluations on noisy voice-search data show relative improvements ranging from 2% at 15 dB to 5.4% at -5 dB over a sequence discriminative, multi-condition trained LSTM acoustic model. Arun Narayanan, Ananya Misra, Kean K. Chin |
INTERSPEECH | 1 |
| 2015 | Improving Robustness of Deep Neural Network Acoustic Models via Speech Separation and Joint Adaptive TrainingabstractAlthough deep neural network (DNN) acoustic models are known to be inherently noise robust, especially with matched training and testing data, the use of speech separation as a frontend and for deriving alternative feature representations has been shown to improve performance in challenging environments. We first present a supervised speech separation system that significantly improves automatic speech recognition (ASR) performance in realistic noise conditions. The system performs separation via ratio time-frequency masking; the ideal ratio mask (IRM) is estimated using DNNs. We then propose a framework that unifies separation and acoustic modeling via joint adaptive training. Since the modules for acoustic modeling and speech separation are implemented using DNNs, unification is done by introducing additional hidden layers with fixed weights and appropriate network architecture. On the CHiME-2 medium-large vocabulary ASR task, and with log mel spectral features as input to the acoustic model, an independently trained ratio masking frontend improves word error rates by 10.9% (relative) compared to the noisy baseline. In comparison, the jointly trained system improves performance by 14.4%. We also experiment with alternative feature representations to augment the standard log mel features, like the noise and speech estimates obtained from the separation module, and the standard feature set used for IRM estimation. Our best system obtains a word error rate of 15.4% (absolute), an improvement of 4.6 percentage points over the next best result on this corpus. Arun Narayanan, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | Analysis-by-synthesis feature estimation for robust automatic speech recognition using spectral masksabstractSpectral masking is a promising method for noise suppression in which regions of the spectrogram that are dominated by noise are attenuated while regions dominated by speech are preserved. It is not clear, however, how best to combine spectral masking with the non-linear processing necessary to compute automatic speech recognition features. We propose an analysis-by-synthesis approach to automatic speech recognition, which, given a spectral mask, poses the estimation of mel frequency cepstral coefficients (MFCCs) of the clean speech as an optimization problem. MFCCs are found that minimize a combination of the distance from the resynthesized clean power spectrum to the regions of the noisy spectrum selected by the mask and the negative log likelihood under an unmodified large vocabulary continuous speech recognizer. In evaluations on the Aurora4 noisy speech recognition task with both ideal and estimated masks, analysis-by-synthesis decreases both word error rates and distances to clean speech as compared to traditional approaches. Michael I. Mandel, Arun Narayanan |
ICASSP | 2 |
| 2014 | Joint noise adaptive training for robust automatic speech recognitionabstractWe explore time-frequency masking to improve noise robust automatic speech recognition. Apart from its use as a frontend, we use it for providing smooth estimates of speech and noise which are then passed as additional features to a deep neural network (DNN) based acoustic model. Such a system improves performance on the Aurora-4 dataset by 10.5% (relative) compared to the previous best published results. By formulating separation as a supervised mask estimation problem, we develop a unified DNN framework that jointly improves separation and acoustic modeling. Our final system outperforms the previous best system on CHiME-2 corpus by 22.1% (relative). Arun Narayanan, DeLiang Wang |
ICASSP | 1 |
| 2014 | Investigation of Speech Separation as a Front-End for Noise Robust Speech RecognitionabstractRecently, supervised classification has been shown to work well for the task of speech separation. We perform an in-depth evaluation of such techniques as a front-end for noise-robust automatic speech recognition (ASR). The proposed separation front-end consists of two stages. The first stage removes additive noise via time-frequency masking. The second stage addresses channel mismatch and the distortions introduced by the first stage; a non-linear function is learned that maps the masked spectral features to their clean counterpart. Results show that the proposed front-end substantially improves ASR performance when the acoustic models are trained in clean conditions. We also propose a diagonal feature discriminant linear regression (dFDLR) adaptation that can be performed on a per-utterance basis for ASR systems employing deep neural networks and HMM. Results show that dFDLR consistently improves performance in all test conditions. Surprisingly, the best average results are obtained when dFDLR is applied to models trained using noisy log-Mel spectral features from the multi-condition training set. With no channel mismatch, the best results are obtained when the proposed speech separation front-end is used along with multi-condition training using log-Mel features followed by dFDLR adaptation. Both these results are among the best on the Aurora-4 dataset. Arun Narayanan, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | On training targets for supervised speech separationabstractFormulation of speech separation as a supervised learning problem has shown considerable promise. In its simplest form, a supervised learning algorithm, typically a deep neural network, is trained to learn a mapping from noisy features to a time-frequency representation of the target of interest. Traditionally, the ideal binary mask (IBM) is used as the target because of its simplicity and large speech intelligibility gains. The supervised learning framework, however, is not restricted to the use of binary targets. In this study, we evaluate and compare separation results by using different training targets, including the IBM, the target binary mask, the ideal ratio mask (IRM), the short-time Fourier transform spectral magnitude and its corresponding mask (FFT-MASK), and the Gammatone frequency power spectrum. Our results in various test conditions reveal that the two ratio mask targets, the IRM and the FFT-MASK, outperform the other targets in terms of objective intelligibility and quality metrics. In addition, we find that masking based targets, in general, are significantly better than spectral envelope based targets. We also present comparisons with recent methods in non-negative matrix factorization and speech enhancement, which show clear performance advantages of supervised speech separation. Yuxuan Wang 0002, Arun Narayanan, DeLiang Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | Coupling binary masking and robust ASRabstractWe present a novel framework for performing speech separation and robust automatic speech recognition (ASR) in a unified fashion. Separation is performed by estimating the ideal binary mask (IBM), which identifies speech dominant and noise dominant units in a time-frequency (T-F) representation of the noisy signal. ASR is performed on extracted cepstral features after binary masking. Previous systems perform these steps in a sequential fashion - separation followed by recognition. The proposed framework, which we call bidirectional speech decoding (BSD), unifies these two stages. It does this by using multiple IBM estimators each of which is designed specifically for a back-end acoustic phonetic unit (BPU) of the recognizer. The standard ASR decoder is modified to use these IBM estimators to obtain BPU-specific cepstra during likelihood calculation. On the Aurora-4 robust ASR task, the proposed framework obtains a relative improvement of 17% in word error rate over the noisy baseline. It also obtains significant improvements in the quality of the estimated IBM. Arun Narayanan, DeLiang Wang |
ICASSP | 1 |
| 2013 | Ideal ratio mask estimation using deep neural networks for robust speech recognitionabstractWe propose a feature enhancement algorithm to improve robust automatic speech recognition (ASR). The algorithm estimates a smoothed ideal ratio mask (IRM) in the Mel frequency domain using deep neural networks and a set of time-frequency unit level features that has previously been used to estimate the ideal binary mask. The estimated IRM is used to filter out noise from a noisy Mel spectrogram before performing cepstral feature extraction for ASR. On the noisy subset of the Aurora-4 robust ASR corpus, the proposed enhancement obtains a relative improvement of over 38% in terms of word error rates using ASR models trained in clean conditions, and an improvement of over 14% when the models are trained using the multi-condition training data. In terms of instantaneous SNR estimation performance, the proposed system obtains a mean absolute error of less than 4 dB in most frequency channels. Arun Narayanan, DeLiang Wang |
ICASSP | 1 |
| 2013 | A Direct Masking Approach to Robust ASRabstractRecently, much work has been devoted to the computation of binary masks for speech segregation. Conventional wisdom in the field of ASR holds that these binary masks cannot be used directly; the missing energy significantly affects the calculation of the cepstral features commonly used in ASR. We show that this commonly held belief may be a misconception; we demonstrate the effectiveness of directly using the masked data on both a small and large vocabulary dataset. In fact, this approach, which we term the direct masking approach, performs comparably to two previously proposed missing feature techniques. We also investigate the reasons why other researchers may have not come to this conclusion; variance normalization of the features is a significant factor in performance. This work suggests a much better baseline than unenhanced speech for future work in missing feature ASR. William Hartmann, Arun Narayanan, Eric Fosler-Lussier, DeLiang Wang |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | On the Role of Binary Mask Pattern in Automatic Speech RecognitionabstractProcessing noisy signals using the ideal binary mask has been shown to improve automatic speech recognition (ASR) performance. In this paper, we present the first study that investigates the role of mask patterns in ASR under varying signalto-noise ratios (SNR), noise conditions and mask definitions. Binary masks are typically computed either by comparing the local SNR within a time-frequency unit of a mixture signal with a threshold termed the local criterion (LC), or by comparing the local target energy with the long-term average energy of speech. Results show that: (i) Akin to human speech recognition, binary masking can significantly improve ASR even when the mixture SNR is as low as -60 dB. (ii) The difference between the LC and the mixture SNR is more correlated to the recognition accuracy than LC. (iii) The performance profiles in ASR are qualitatively similar to those obtained for human speech recognition. (iv) The LC at which the peak performance is obtained is lower than 0 dB, which is the optimal threshold as far as the SNR gain of processed signals is concerned. This indicates that maximizing SNR gain may not be the optimal criterion to improve either human or machine recognition of noisy speech. Arun Narayanan, DeLiang Wang |
INTERSPEECH | 1 |
| 2012 | A CASA-Based System for Long-Term SNR EstimationabstractWe present a system for robust signal-to-noise ratio (SNR) estimation based on computational auditory scene analysis (CASA). The proposed algorithm uses an estimate of the ideal binary mask to segregate a time-frequency representation of the noisy signal into speech dominated and noise dominated regions. Energy within each of these regions is summated to derive the filtered global SNR. An SNR transform is introduced to convert the estimated filtered SNR to the true broadband SNR of the noisy signal. The algorithm is further extended to estimate subband SNRs. Evaluations are done using the TIMIT speech corpus and the NOISEX92 noise database. Results indicate that both global and subband SNR estimates are superior to those of existing methods, especially at low SNR conditions. Arun Narayanan, DeLiang Wang |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | On the use of ideal binary masks for improving phonetic classificationabstractIdeal binary masks are binary patterns that encode the masking characteristics of speech in noise. Recent evidence in speech perception suggests that such binary patterns provide sufficient information for human speech recognition. Motivated by these findings, we propose to use ideal binary masks to improve phonetic modeling. We show that by combining the outputs of classifiers trained on the traditional MFCC features and this novel speech pattern, statistically significant improvements over the baseline MFCC based classifier can be achieved for the task of phonetic classification. Using the combined classifiers, we achieve an error rate of 19.5% on the TIMIT phonetic classification task using multilayer perceptrons as the underlying classifier. Arun Narayanan, DeLiang Wang |
ICASSP | 1 |
| 2011 | Robust speech recognition using multiple prior models for speech reconstructionabstractPrior models of speech have been used in robust automatic speech recognition to enhance noisy speech. Typically, a single prior model is trained by pooling the entire training data. In this paper we propose to train multiple prior models of speech instead of a single prior model. The prior models can be trained based on distinct characteristics of speech. In this study, they are trained based on voicing characteristics. The trained prior models are then used to reconstruct noisy speech. Significant improvements are obtained on the Aurora-4 robust speech recognition task when multiple priors are used; in conjunction with an uncertainty transform technique, multiple priors yield a 13.7% absolute improvement in the average word error rate over directly recognizing noisy speech. Arun Narayanan, Xiaojia Zhao, DeLiang Wang, Eric Fosler-Lussier |
ICASSP | 1 |