VLDB 2026 Research / reviewers in the wild / expert
Sankaran Panchapagesan
dblp:74/7160
· DBLP profile ↗
16ranked-venue papers
8as first author
6since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 10 · 6 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Improving Acoustic Echo Cancellation for Voice Assistants Using Neural Echo Suppression and Multi-Microphone Noise ReductionabstractKeyword spotting (KS) and automatic speech recognition (ASR) on smart speakers in a home environment with interfering signals from loudspeakers are challenging tasks to this day, despite improvements in acoustic echo cancellation (AEC) systems. In this work we propose to combine a single microphone AEC system, consisting of an adaptive linear filter (linear AEC) and a neural echo suppressor (NES), with an adaptive filter developed for multi-microphone noise reduction, called Cleaner. This additional enhancement step allows the AEC system to profit from spatial information to remove residual echo. The single microphone NES model improves upon the waveform domain counterpart proposed in [1] using a frequency domain representation that helps with generalization. Furthermore, we show that using multiple linear AEC configurations during model training provides large gains over a fixed configuration. On the hardest considered test condition, the proposed system outperforms the baseline model [1] for single microphone input by 66 % (relative) in KS false reject rate (FRR) and 52 % (relative) in ASR word error rate (WER). Using the multi-microphone setting, the FRR is reduced by an additional 52 % and the WER by an additional 32 %. Jens Heitkaemper, Arun Narayanan, Turaj Zakizadeh Shabestary, Sankaran Panchapagesan, James Walker, Bhalchandra Gajare, Shlomi Regev, Ajay Dudani, Alexander Gruenstein |
ICASSP | 4 |
| 2023 | On Training a Neural Residual Acoustic Echo Suppressor for Improved ASR
Sankaran Panchapagesan, Turaj Zakizadeh Shabestary, Arun Narayanan |
INTERSPEECH | 1 |
| 2022 | SNRi Target Training for Joint Speech Enhancement and Recognition
Yuma Koizumi, Shigeki Karita, Arun Narayanan, Sankaran Panchapagesan, Michiel Bacchiani |
INTERSPEECH | 4 |
| 2022 | A Conformer-based Waveform-domain Neural Acoustic Echo Canceller Optimized for ASR AccuracyabstractAcoustic Echo Cancellation (AEC) is essential for accurate recognition of queries spoken to a smart speaker that is playing out audio.Previous work has shown that a neural AEC model operating on log-mel spectral features (denoted "logmel" hereafter) can greatly improve Automatic Speech Recognition (ASR) accuracy when optimized with an auxiliary loss utilizing a pre-trained ASR model encoder.In this paper, we develop a conformer-based waveform-domain neural AEC model inspired by the "TasNet" architecture.The model is trained by jointly optimizing Negative Scale-Invariant SNR (SISNR) and ASR losses on a large speech dataset.On a realistic rerecorded test set, we find that cascading a linear adaptive AEC and a waveform-domain neural AEC is very effective, giving 56-59% word error rate (WER) reduction over the linear AEC alone.On this test set, the 1.6M parameter waveform-domain neural AEC also improves over a larger 6.5M parameter logmeldomain neural AEC model by 20-29% in easy to moderate conditions.By operating on smaller frames, the waveform neural model is able to perform better at smaller sizes and is better suited for applications where memory is limited. Sankaran Panchapagesan, Arun Narayanan, Turaj Zakizadeh Shabestary, Nathan Howard, Alex Park 0001, James Walker, Alexander Gruenstein |
INTERSPEECH | 1 |
| 2022 | Learning Mask Scalars for Improved Robust Automatic Speech RecognitionabstractImproving robustness of streaming automatic speech recognition (ASR) systems using neural network based acoustic frontends is challenging because of causality constraints and the speech-distortions introduced by the frontend. Time-frequency masking based approaches are commonly used, but they need additional hyperparameters – mask scalars – to limit distortion. Mask scalars are typically hand-tuned and chosen conservatively. In this work, we present a technique to predict mask scalars using ASR loss in an end-to-end fashion, with minimal increase in model size and complexity. We evaluate the approach on two robust ASR tasks: multichannel enhancement in the presence of speech and non-speech noise, and acoustic echo cancellation (AEC). Results show that the presented algorithm consistently improves word error rate (WER) over strong baselines that use hand-tuned hyperparameters: up to 16% in noisy conditions, and up to 7% for AEC. Arun Narayanan, James Walker, Sankaran Panchapagesan, Nathan Howard, Yuma Koizumi |
SLT | 3 |
| 2021 | Efficient Knowledge Distillation for RNN-Transducer ModelsabstractKnowledge Distillation is an effective method of transferring knowledge from a large model to a smaller model. Distillation can be viewed as a type of model compression, and has played an important role for on-device ASR applications. In this paper, we develop a distillation method for RNN-Transducer (RNN-T) models, a popular end-to-end neural network architecture for streaming speech recognition. Our proposed distillation loss is simple and efficient, and uses only the "y" and "blank" posterior probabilities from the RNN-T output probability lattice. We study the effectiveness of the proposed approach in improving the accuracy of sparse RNN-T models obtained by gradually pruning a larger uncompressed model, which also serves as the teacher during distillation. With distillation of 60% and 90% sparse multi-domain RNN-T models, we obtain WER reductions of 4.3% and 12.1% respectively, on a noisy FarField eval set. We also present results of experiments on LibriSpeech, where the introduction of the distillation loss yields a 4.8% relative WER reduction on the test-other dataset for a small Conformer model. Sankaran Panchapagesan, Daniel S. Park, Chung-Cheng Chiu, Yuan Shangguan, Qiao Liang 0001, Alexander Gruenstein |
ICASSP | 1 |
| 2018 | Monophone-Based Background Modeling for Two-Stage On-Device Wake Word DetectionabstractAccurate on-device wake word detection is crucial to products with far-field voice control such as the Amazon Echo. It is quite challenging to build a wake word system with both low False Reject Rate (FRR) and low False Alarm Rate (FAR) in real scenarios where there are various types of background speech, music or noise, especially when computational resources on the device is limited. In this paper, we introduce a two-stage wake word system based on Deep Neural Network (DNN) acoustic modeling, propose a new way to model the non-keyword background events using monophone-based units and present how richer information can be extracted from those monophone units for final wake word detection. Under the new system, we could get around 16% relative reduction in FRR when fixing the false alarm level, and about 37% relative reduction in FAR on the other hand if we maintain the miss rate. For the 2nd stage classifier itself, it is able to reduce the false alarm rate relatively by about 67% on top of 1st stage hypothesis with very few computational resources. Minhua Wu, Sankaran Panchapagesan, Ming Sun 0007, Jiacheng Gu, Ryan Thomas, Shiv Vitaladevuni, Björn Hoffmeister, Arindam Mandal |
ICASSP | 2 |
| 2017 | Direct modeling of raw audio with DNNS for wake word detectionabstractIn this work, we develop a technique for training features directly from the single-channel speech waveform in order to improve wake word (WW) detection performance. Conventional speech recognition systems typically extract a compact feature representation based on prior knowledge such as log-mel filter bank energy (LFBE). Such a feature is then used for training a deep neural network (DNN) acoustic model (AM). In contrast, we directly train the WW DNN AM from the single-channel audio data in a stage-wise manner. We first build a feature extraction DNN with a small hidden bottleneck layer, and train this bottleneck feature representation using the same multi-task cross-entropy objective function as we use to train our WW DNNs. Then, the WW classification DNN is trained with input bottleneck features, keeping the feature extraction layers fixed. Finally, the feature extraction and classification DNNs are combined and then jointly optimized. We show the effectiveness of this stage-wise training technique through a set of experiments on real beam-formed far-field data. The experiment results show that the audioinput DNN provides significantly lower miss rates for a range of false alarm rates over the LFBE when a sufficient amount of training data is available, yielding approximately 12 % relative improvement in the area under the curve (AUC). Ken'ichi Kumatani, Sankaran Panchapagesan, Minhua Wu, Nikko Strom, Gautam Tiwari, Arindam Mandal |
ASRU | 2 |
| 2017 | Compressed Time Delay Neural Network for Small-Footprint Keyword Spotting
Ming Sun 0007, David Snyder, Varun K. Nagaraja, Mike Rodehorst, Sankaran Panchapagesan, Nikko Strom, Spyridon Matsoukas, Shiv Vitaladevuni |
INTERSPEECH | 6 |
| 2016 | Multi-Task Learning and Weighted Cross-Entropy for DNN-Based Keyword Spotting
Sankaran Panchapagesan, Ming Sun 0007, Aparna Khare, Spyridon Matsoukas, Arindam Mandal, Björn Hoffmeister, Shiv Vitaladevuni |
INTERSPEECH | 1 |
| 2016 | Model Compression Applied to Small-Footprint Keyword Spotting
George Tucker, Minhua Wu, Ming Sun 0007, Sankaran Panchapagesan, Gengshen Fu, Shiv Vitaladevuni |
INTERSPEECH | 4 |
| 2016 | Max-pooling loss training of long short-term memory networks for small-footprint keyword spottingabstractWe propose a max-pooling based loss function for training Long Short-Term Memory (LSTM) networks for small-footprint keyword spotting (KWS), with low CPU, memory, and latency requirements. The max-pooling loss training can be further guided by initializing with a cross-entropy loss trained network. A posterior smoothing based evaluation approach is employed to measure keyword spotting performance. Our experimental results show that LSTM models trained using cross-entropy loss or max-pooling loss outperform a cross-entropy loss trained baseline feed-forward Deep Neural Network (DNN). In addition, max-pooling loss trained LSTM with randomly initialized network performs better compared to cross-entropy loss trained LSTM. Finally, the max-pooling loss trained LSTM initialized with a cross-entropy pre-trained network shows the best performance, which yields 67:6% relative reduction compared to baseline feed-forward DNN in Area Under the Curve (AUC) measure. Ming Sun 0007, Anirudh Raju, George Tucker, Sankaran Panchapagesan, Gengshen Fu, Arindam Mandal, Spyridon Matsoukas, Nikko Strom, Shiv Vitaladevuni |
SLT | 4 |
| 2009 | Frequency warping for VTLN and speaker adaptation by linear transformation of standard MFCC
Sankaran Panchapagesan, Abeer Alwan |
Comput. Speech Lang. | 1 |
| 2008 | Vocal tract inversion by cepstral analysis-by-synthesis using chain matricesabstractAcoustic-to-articulatory inversion for vowels is performed by cepstral analysis-by-synthesis, using chain-matrix calculation of vocal tract (VT) acoustics and the Maeda articulatory model. The derivative of the VT chain matrix with respect to the area function was calculated in a novel efficient manner, and used in the BFGS quasi-Newton method for optimizing a distance measure between input and synthesized cepstral features over the entire articulatory trajectory. The optimization is initialized by a fast search of an articulatory codebook with a bin structure in formant space and the cost function also includes regularization and continuity terms to obtain realistic inverted VT shapes and smooth articulatory trajectories. Inversion is evaluated on the three diphthongs /ai/, /oi/ and /au/ of two speakers, one male and one female, from the University of Wisconsin X-ray microbeam (XRMB) database, and good agreement was achieved between inverted midsagittal vocal tract outlines and measured XRMB tongue and lip pellet positions, with an average relative error of less than 3% in the first three formants. Sankaran Panchapagesan, Abeer Alwan |
INTERSPEECH | 1 |
| 2006 | Multi-Parameter Frequency Warping for Vtln by Gradient SearchabstractThe current method for estimating frequency warping (FW) functions for vocal tract length normalization (VTLN) is by maximizing the ASR likelihood score by an exhaustive search over a grid of FW parameters. Exhaustive search is inefficient when estimating multi-parameter FWs, which have been shown to give improvements in recognition accuracy over single parameter FWs (J.W. McDonough, 2000). Here we develop a gradient search algorithm to obtain the optimal FW parameters for MFCC features, since previous work focussed on PLP cepstral features (J.W. McDonough, 2000). The novel calculation involved was that of the gradient of the Mel filterbank with respect to the FW parameters. Even for a single parameter, the gradient search method was more efficient than grid search by a factor of around 1.6 on the average for male children speakers tested on models trained from adult males. When used to estimate multi-parameter sine-log allpass transform (SLAPT, (J.W. McDonough, 2000)) FWs for VTLN, more than 50% reduction in word error rate was obtained with five parameter SLAPT compared to single-parameter piecewise linear FW Sankaran Panchapagesan, Abeer Alwan |
ICASSP (1) | 1 |
| 2006 | Frequency warping by linear transformation of standard MFCCabstractpanchap @ icsl.ucla.edu A novel linear transform (LT) is proposed for frequency warp-ing (FW) with standard filterbank based MFCC features. Here, we use the idea of spectral interpolation of [9] to perform a continuous warping in the log filterbank output domain, and incorporate both interpolation and warping into a single warped IDCT matrix. The new transformation matrix is thus mathematically simpler than in [9], and no modification of standard MFCC feature extraction is required like the previous approach. In VTLN experiments with maximum likelihood score (MLS) estimation of the FW parameter, the new LT outperformed regular VTLN implemented by warping the Mel filterbank. In speaker adaptation experiments using the new LT to transform HMM means, the results were significantly better than MLLR for limited adaptation data and comparable to those in [8], while using the computationally simpler MLS FW estimation. Index Terms: speech recognition, speaker normalization, fre-quency warping, linear transformation, speaker adaptation Sankaran Panchapagesan |
INTERSPEECH | 1 |