VLDB 2026 Research / reviewers in the wild / expert
Nathan Howard
dblp:26/4494
· DBLP profile ↗
8ranked-venue papers
1as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bone Conducted Signal Guided Speech Enhancement For Voice Assistant on EarbudsabstractIn this work we present a multi-modal, streaming enhancement network to improve speech recognition for voice assistants on earbuds. The proposed model is guided by a bone conducted signal (BCS) to separate the interfering sources from the target speaker signal. We train the model on a simulated speech enhancement training set with a simulated BCS and finetune it on a small earbuds specific training set, consisting of about 6 hours of speech. To account for distorted BCS the enhancement module is complemented by a voice activity-based decision to discard the enhanced output for BCS without speech information. A possibility to preprocess the BCS to account for the low-pass characteristic of the bone conduction is evaluated to lower the required transmission bandwidth from the earbuds to the recognition device. The results show that the BCS bandwidth can be reduced to 500 Hz with only small losses in word error rate. In comparison with a larger state-of-the-art multi-channel enhancement method, the systems, with and without bandwidth reduction, demonstrate superior performance on most of the considered realistic test sets. Jens Heitkaemper, Joe Caroselli, Max McKinnon, Arun Narayanan, Nathan Howard |
ICASSP | 5 |
| 2024 | TfCleanformer: A streaming, array-agnostic, full- and sub-band modeling front-end for robust ASR
Jens Heitkaemper, Joe Caroselli, Arun Narayanan, Nathan Howard |
INTERSPEECH | 4 |
| 2023 | Cleanformer: A Multichannel Array Configuration-Invariant Neural Enhancement Frontend for ASR in Smart SpeakersabstractThis work introduces Cleanformer —a streaming multichannel neural enhancement frontend for automatic speech recognition (ASR). This model has a Conformer-based architecture which takes as inputs a single channel each of raw and enhanced signals, and uses self-attention to derive a time-frequency mask. The enhanced input is generated by a multichannel adaptive noise cancellation algorithm known as Speech Cleaner. The time-frequency mask is applied to the noisy input to produce enhanced features for ASR. Detailed evaluations are presented with speech- and non-speech-based noise that show significant reduction in word error rate (WER) – about 80% for -6 dB SNR – over a state-of-the-art ASR model alone. It also significantly outperforms enhancement using a beamformer with ideal steering. The enhancement model can be used with different microphone arrays without the need for retraining. Joseph Caroselli, Arun Narayanan, Nathan Howard, Tom O'Malley |
ICASSP | 3 |
| 2022 | A Conformer-based Waveform-domain Neural Acoustic Echo Canceller Optimized for ASR AccuracyabstractAcoustic Echo Cancellation (AEC) is essential for accurate recognition of queries spoken to a smart speaker that is playing out audio.Previous work has shown that a neural AEC model operating on log-mel spectral features (denoted "logmel" hereafter) can greatly improve Automatic Speech Recognition (ASR) accuracy when optimized with an auxiliary loss utilizing a pre-trained ASR model encoder.In this paper, we develop a conformer-based waveform-domain neural AEC model inspired by the "TasNet" architecture.The model is trained by jointly optimizing Negative Scale-Invariant SNR (SISNR) and ASR losses on a large speech dataset.On a realistic rerecorded test set, we find that cascading a linear adaptive AEC and a waveform-domain neural AEC is very effective, giving 56-59% word error rate (WER) reduction over the linear AEC alone.On this test set, the 1.6M parameter waveform-domain neural AEC also improves over a larger 6.5M parameter logmeldomain neural AEC model by 20-29% in easy to moderate conditions.By operating on smaller frames, the waveform neural model is able to perform better at smaller sizes and is better suited for applications where memory is limited. Sankaran Panchapagesan, Arun Narayanan, Turaj Zakizadeh Shabestary, Nathan Howard, Alex Park 0001, James Walker, Alexander Gruenstein |
INTERSPEECH | 5 |
| 2022 | Learning Mask Scalars for Improved Robust Automatic Speech RecognitionabstractImproving robustness of streaming automatic speech recognition (ASR) systems using neural network based acoustic frontends is challenging because of causality constraints and the speech-distortions introduced by the frontend. Time-frequency masking based approaches are commonly used, but they need additional hyperparameters – mask scalars – to limit distortion. Mask scalars are typically hand-tuned and chosen conservatively. In this work, we present a technique to predict mask scalars using ASR loss in an end-to-end fashion, with minimal increase in model size and complexity. We evaluate the approach on two robust ASR tasks: multichannel enhancement in the presence of speech and non-speech noise, and acoustic echo cancellation (AEC). Results show that the presented algorithm consistently improves word error rate (WER) over strong baselines that use hand-tuned hyperparameters: up to 16% in noisy conditions, and up to 7% for AEC. Arun Narayanan, James Walker, Sankaran Panchapagesan, Nathan Howard, Yuma Koizumi |
SLT | 4 |
| 2021 | A Conformer-Based ASR Frontend for Joint Acoustic Echo Cancellation, Speech Enhancement and Speech SeparationabstractWe present a frontend for improving robustness of automatic speech recognition (ASR), that jointly implements three modules within a single model: acoustic echo cancellation, speech enhancement, and speech separation. This is achieved by using a contextual enhancement neural network that can optionally make use of different types of side inputs: (1) a reference signal of the playback audio, which is necessary for echo cancellation; (2) a noise context, which is useful for speech enhancement; and (3) an embedding vector representing the voice characteristic of the target speaker of interest, which is not only critical in speech separation, but also helpful for echo cancellation and speech enhancement. We present detailed evaluations to show that the joint model performs almost as well as the task-specific models, and significantly reduces word error rate in noisy conditions even when using a large-scale state-of-the-art ASR model. Compared to the noisy baseline, the joint model reduces the word error rate in low signal-to-noise ratio conditions by at least 71% on our echo cancellation dataset, 10% on our noisy dataset, and 26% on our multi-speaker dataset. Compared to task-specific models, the joint model performs within 10% on our echo cancellation dataset, 2% on the noisy dataset, and 3% on the multi-speaker dataset. Tom O'Malley, Arun Narayanan, Alex Park 0001, James Walker, Nathan Howard |
ASRU | 6 |
| 2021 | A Neural Acoustic Echo Canceller Optimized Using An Automatic Speech Recognizer and Large Scale Synthetic DataabstractWe consider the problem of recognizing speech utterances spoken to a device which is generating a known sound waveform; for example, recognizing queries issued to a digital assistant which is generating responses to previous user inputs. Previous work has proposed building acoustic echo cancellation (AEC) models for this task that optimize speech enhancement metrics using both neural network as well as signal processing approaches.Since our goal is to recognize the input speech, we consider enhancements which improve word error rates (WERs) when the predicted speech signal is passed to an automatic speech recognition (ASR) model. First, we augment the loss function with a term that produces outputs useful to a pre-trained ASR model and show that this augmented loss function improves WER metrics. Second, we demonstrate that augmenting our training dataset of real world examples with a large synthetic dataset improves performance. Crucially, applying SpecAugment style masks to the reference channel during training aids the model in adapting from synthetic to real domains. In experimental evaluations, we find the proposed approaches improve performance, on average, by 57% over a signal processing baseline and 45% over the neural AEC model without the proposed changes. Nathan Howard, Alex Park 0001, Turaj Zakizadeh Shabestary, Alexander Gruenstein, Rohit Prabhavalkar |
ICASSP | 1 |
| 2005 | A space-based end-to-end prototype geographic information network for lunar and planetary exploration and emergency response (2002 and 2003 field experiments)
Richard A. Beck, Robert K. Vincent, Doyle W. Watts, Marc A. Seibert, David P. Pleva, Michael A. Cauley, Calvin T. Ramos, Theresa M. Scott, Dean W. Harter, Mary Vickerman, David Irmies, Al Tucholski, Brian Frantz, Glenn Lindamood, Isaac Lopez, Gregory J. Follen, Thaddeus J. Kollar, Jay Horowitz, Robert Griffin, Raymond Gilstrap, Marjory J. Johnson, Kenneth Freeman, Celeste Banaag, Joseph Kosmo, Amy Ross, Kevin Groneman, Jeffrey Graham, Kim Shillcutt, Robert Hirsh, Nathan Howard, Dean B. Eppler |
Comput. Networks | 30 |