EDBT 2026 Demo / reviewers in the wild / expert
Kaustubh Kalgaonkar
dblp:79/2771
· DBLP profile ↗
27ranked-venue papers
14as first author
10since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 14 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 5 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Directional Source Separation for Robust Speech Recognition on Smart GlassesabstractModern smart glasses leverage machine learning to offer real-time transcriptions, considerably enriching human communication experiences. However, such systems frequently encounter challenges related to environmental noises, leading to decreased speech recognition. To improve voice quality, this work investigates directional source separation using the multi-microphone array. We explore multiple beamformers to assist source separation by strengthening the directional properties of speech signals. In addition to relying on predetermined beamformers, we investigate neural beamforming in multi-channel source separation, demonstrating that automatic learning directional characteristics effectively improves separation quality. Furthermore, we investigate the training strategies for ASR when utilizing separated outputs. Our results suggest that jointly training a directional speech separation and ASR model achieves the best overall performance while balancing the wearer and conversation partner’s performance. Tiantian Feng, Ju Lin, Yiteng Huang, Weipeng He, Kaustubh Kalgaonkar, Niko Moritz, Ming Sun 0013, Frank Seide |
ICASSP | 5 |
| 2025 | SMARTMOS: Modeling Subjective Audio Quality Evaluation for Real-Time Applications
Jose Antonio Jimenez Amador, Kaustubh Kalgaonkar, King-Wei Hor, Sriram Srinivasan 0003 |
INTERSPEECH | 3 |
| 2024 | Interference Aware Training Target for DNN based joint Acoustic Echo Cancellation and Noise Suppression
Vahid Khanagha, Dimitris Koutsaidis, Kaustubh Kalgaonkar, Sriram Srinivasan 0003 |
INTERSPEECH | 3 |
| 2023 | SCA: Streaming Cross-Attention Alignment For Echo CancellationabstractEnd-to-End deep learning has shown promising results for speech enhancement tasks, such as noise suppression, dereverberation, and speech separation. However, most state-of-the-art methods for echo cancellation are either classical DSP-based or hybrid DSP-ML algorithms. Components such as the delay estimator and adaptive linear filter are based on traditional signal processing concepts, and deep learning algorithms typically only serve to replace the non-linear residual echo suppressor. This paper introduces an end-to-end echo cancellation network with a streaming cross-attention alignment (SCA). Our proposed method can handle unaligned inputs without requiring external alignment and generate high-quality speech without echoes. At the same time, the end-to-end algorithm simplifies the current echo cancellation pipeline for time-variant echo path cases. We test our proposed method on the ICASSP2022 and Inter-speech2021 Microsoft deep echo cancellation challenge evaluation dataset, where our method outperforms some of the other hybrid and end-to-end methods. Yang Liu 0175, Yangyang Shi, Kaustubh Kalgaonkar, Sriram Srinivasan 0003 |
ICASSP | 4 |
| 2023 | Egocentric Audio-Visual Noise SuppressionabstractThis paper studies audio-visual noise suppression for egocentric videos -where the speaker is not captured in the video. Instead, potential noise sources are visible on screen with the camera emulating the off-screen speaker’s view of the outside world. This setting is different from prior work in audio-visual speech enhancement that relies on lip and facial visuals. In this paper, we first demonstrate that egocentric visual information is helpful for noise suppression. We compare object recognition and action classification-based visual feature extractors and investigate methods to align audio and visual representations. Then, we examine different fusion strategies for the aligned features, and locations within the noise suppression model to incorporate visual information. Experiments demonstrate that visual features are most helpful when used to generate additive correction masks. Finally, in order to ensure that the visual features are discriminative with respect to different noise types, we introduce a multi-task learning framework that jointly optimizes audio-visual noise suppression and video-based acoustic event detection. This proposed multi-task framework outperforms the audio-only baseline on all metrics, including a 0.16 PESQ improvement. Extensive ablations reveal the improved performance of the proposed model with multiple active distractors, overall noise types, and across different SNRs. Weipeng He, Ju Lin, Egor Lakomkin, Kaustubh Kalgaonkar |
ICASSP | 6 |
| 2023 | Directional Speech Recognition for Speaker Disambiguation and Cross-talk Suppression
Ju Lin, Niko Moritz, Ruiming Xie, Kaustubh Kalgaonkar, Christian Fügen, Frank Seide |
INTERSPEECH | 4 |
| 2022 | Architecture for Variable Bitrate Neural Speech Codec with Configurable Computation ComplexityabstractLow bitrate speech codecs have become an area of intense research. Traditional speech codecs, which use signal processing methods to encode and decode speech, often suffer from quality issues at low bitrates. A neural speech codec, which uses a deep neural network in the compression pipeline, can help alleviate this issue. In this paper we present a new neural speech codec that: 1) supports variable bitrates 2) supports packet losses of up to 120 ms and 3) can operate at low-compute and high-compute modes. Our codec uses a hierarchical VQ-VAE (HVQVAE) for encoding and decoding spectral features at different bitrates. The decoded features are fed to a vocoder for speech synthesis. Depending upon the end user’s computing resources, the decoder either uses a powerful WaveRNN or a parametric vocoder for speech synthesis. Our experiments demonstrate that our HVQVAE + WaveRNN setup achieves high audio quality. Tejas Jayashankar, Thilo Köhler, Kaustubh Kalgaonkar, Zhiping Xiu, Jilong Wu, Ju Lin, Prabhav Agrawal |
ICASSP | 3 |
| 2022 | Speech Enhancement for Low Bit Rate Speech CodecabstractSpeech codec compresses the input signal into compact bit stream, which is then decoded at the receiver to generate the best possible perceptual quality. This compression makes storing and transmitting speech efficient. In this work, we propose a neural extension to low bit rate speech codec (e.g., Codec2) that aims to improve the perceptual quality of synthesized speech. Our proposed framework combines decoded audio with neural embeddings without breaking the existing speech coders. In addition to embeddings, we also use the least-square generative adversarial network (LSGAN) to reduce artifacts and prevent over-smoothing in the reconstructed audio. The Mean Opinion Scores (MOS) from the listening tests show that our framework can boost the audio quality of speech encoded at 3.6kbps to outperform that of speech encoded at 6kbps using Opus. Ju Lin, Kaustubh Kalgaonkar |
ICASSP | 2 |
| 2021 | A Time-Domain Convolutional Recurrent Network for Packet Loss ConcealmentabstractPacket loss may affect a wide range of applications that use voice over IP (VoIP), e.g. video conferencing. In this paper, we investigate a time-domain convolutional recurrent network (CRN) for online packet loss concealment. The CRN comprises a convolutional encoder-decoder structure and long short-term memory (LSTM) layers, which have been shown to be suitable for real-time speech enhancement applications. Moreover, we propose lookahead and masked training to further improve the performance of the CRN framework. Experimental results show that the proposed system outperforms a baseline system using only LSTM layers in terms of two objective metrics – perceptual evaluation of speech quality (PESQ) and short-term objective intelligibility (STOI); it also reduces the word error rate (WER) more than the baseline when used as a frontend for speech recognition. The advantage of the proposed system is also verified in a subjective evaluation by the mean opinion score (MOS). Ju Lin, Kaustubh Kalgaonkar, Gil Keren, Didi Zhang, Christian Fügen |
ICASSP | 3 |
| 2021 | A Two-Stage Approach to Speech Bandwidth Extension
Ju Lin, Kaustubh Kalgaonkar, Gil Keren, Didi Zhang, Christian Fügen |
Interspeech | 3 |
| 2020 | Spatial Attention for Far-Field Speech Recognition with Deep Beamforming Neural NetworksabstractIn this paper, we introduce spatial attention for refining the information in multi-direction neural beamformer for far-field automatic speech recognition. Previous approaches of neural beamformers with multiple look directions, such as the factored complex linear projection, have shown promising results. However, the features extracted by such methods contain redundant information, as only the direction of the target speech is relevant. We propose using a spatial attention subnet to weigh the features from different directions, so that the subsequent acoustic model could focus on the most relevant features for the speech recognition. Our experimental results show that spatial attention achieves up to 9% relative word error rate improvement over methods without the attention. Weipeng He, Biqiao Zhang, Jay Mahadeokar, Kaustubh Kalgaonkar, Christian Fügen |
ICASSP | 5 |
| 2015 | Estimating confidence scores on ASR results using recurrent neural networksabstractIn this paper we present a confidence estimation system using recurrent neural networks (RNN) and compare it to a traditional multilayered perception (MLP) based system. The ability of RNN to capture sequence information and improve decisions using processed history was main motivation to explore RNN's for confidence estimation. In this paper we also explore two subtle variations of confidence estimator: one that uses objective extracted over the entire sequence for training, and other that uses dynamic programming to decode and estimate confidence on all the words of the sequence jointly. In our experiments, we observed that for a constant false positive (FP) rate of 3% we can secure a relative reduction of 10% in false negative (FN) rate when we replaced a MLP in confidence estimator with a RNN.We also observed that relative gains achieved by a RNN based confidence estimator are directly proportional to the number of word in the utterances. Kaustubh Kalgaonkar, Chaojun Liu, Yifan Gong 0001, Kaisheng Yao |
ICASSP | 1 |
| 2010 | HMM adaptation using sparse Probabilistic Space Mapping for noisy speechabstractThis paper presents an extension of Probabilistic Space Maps (PS-MAPS) to adapt clean acoustic models to a noisy environment. In the presence of noise, the relationship between noisy and clean speech features (MFCC's, HDLA, etc.) is either nonlinear or unknown. Given the relationship between features, traditional methods try to linearize it using approximations. These methods cannot be used for systems where the mapping model for clean and noisy features is missing. Given sufficient training data, PS-MAPS provides an excellent framework for extracting and modeling this relationship. The PS-MAP based approach to model adaptation is completely data driven. Experiments were performed on Aurora 2 dataset to evaluate the effectiveness of the algorithm. Kaustubh Kalgaonkar, Mark A. Clements |
ICASSP | 1 |
| 2010 | Acoustic model adaptation via Linear Spline Interpolation for robust speech recognitionabstractWe recently proposed a new algorithm to perform acoustic model adaptation to noisy environments called Linear Spline Interpolation (LSI). In this method, the nonlinear relationship between clean and noisy speech features is modeled using linear spline regression. Linear spline parameters that minimize the error the between the predicted noisy features and the actual noisy features are learned from training data. A variance associated with each spline segment captures the uncertainty in the assumed model. In this work, we extend the LSI algorithm in two ways. First, the adaptation scheme is extended to compensate for the presence of linear channel distortion. Second, we show how the noise and channel parameters can be updated during decoding in an unsupervised manner within the LSI framework. Using LSI, we obtain an average relative improvement in word error rate of 10.8% over VTS adaptation on the Aurora 2 task with improvements of 15-18% at SNRs between 10 and 15 dB. Michael L. Seltzer, Alex Acero, Kaustubh Kalgaonkar |
ICASSP | 3 |
| 2010 | Synthesizing speech from Doppler signalsabstractIt has long been considered a desirable goal to be able to construct an intelligible speech signal merely by observing the talker in the act of speaking. Past methods at performing this have been based on camera-based observations of the talker's face, combined with statistical methods that infer the speech signal from the facial motion captured by the camera. Other methods have included synthesis of speech from measurements taken by electro-myelo graphs and other devices that are tethered to the talker - an undesirable setup. In this paper we present a new device for synthesizing speech from characterizations of facial motion associated with speech - a Doppler sonar. Facial movement is characterized through Doppler frequency shifts in a tone that is incident on the talker's face. These frequency shifts are used to infer the underlying speech signal. The setup is farfield and untethered, with the sonar acting from the distance of a regular desktop microphone. Preliminary experimental evaluations show that the mechanism is very promising - we are able to synthesize reasonable speech signals, comparable to those obtained from tethered devices such as EMGs. Arthur R. Toth, Kaustubh Kalgaonkar, Bhiksha Raj, Tony Ezzat |
ICASSP | 2 |
| 2009 | Noise robust model adaptation using linear spline interpolationabstractThis paper presents a novel data-driven technique for performing acoustic model adaptation to noisy environments. In the presence of additive noise, the relationship between log mel spectra of speech, noise and noisy speech is nonlinear. Traditional methods linearize this relationship using the mode of the nonlinearity or use some other approximation. The approach presented in this paper models this nonlinear relationship using linear spline regression. In this method, the set of spline parameters that minimizes the error between the predicted and actual noisy speech features is learned from training data, and used at runtime to adapt clean acoustic model parameters to the current noise conditions. Experiments were performed to evaluate the performance of the system on the Aurora 2 task. Results show that the proposed adaptation algorithm (word accuracy 89.22%) outperforms VTS model adaptation (word accuracy 88.38%). Kaustubh Kalgaonkar, Michael L. Seltzer, Alex Acero |
ASRU | 1 |
| 2009 | Sparse probabilistic state mapping and its application to speech bandwidth expansionabstractIn this paper we present a probabilistic algorithm that extracts a mapping between two subspaces by representing each subspace as a collection of states. An arbitrary increase in number of states results in over-fitting the training data without exploring the underlying structure of the map. This paper suggests a method to impose sparsity constraints on the state map by using entropic priors. This probabilistic model is applied to the problem of artificial bandwidth expansion that involves estimating the missing frequency components (3.7 - 8 kHz and 0 - 0.3 kHz) of speech given the narrowband speech signal (0.3 - 3.7 kHz). Kaustubh Kalgaonkar, Mark A. Clements |
ICASSP | 1 |
| 2009 | One-handed gesture recognition using ultrasonic Doppler sonarabstractThis paper presents a new device based on ultrasonic sensors to recognize one-handed gestures. The device uses three ultrasonic receivers and a single transmitter. Gestures are characterized through the Doppler frequency shifts they generate in reflections of an ultrasonic tone emitted by the transmitter. We show that this setup can be used to classify simple one-handed gestures with high accuracy. The ultrasonic doppler based device is very inexpensive - $20 USD for the whole setup including the acquisition system, and computationally efficient as compared to most traditional devices (e.g. video). These gestures, could potentially be used to control and drive a device. Kaustubh Kalgaonkar, Bhiksha Raj |
ICASSP | 1 |
| 2009 | Constrained probabilistic subspace maps applied to speech enhancement
Kaustubh Kalgaonkar, Mark A. Clements |
INTERSPEECH | 1 |
| 2008 | Recognizing talking faces from acoustic Doppler reflectionsabstractFace recognition algorithms typically deal with the classification of static images of faces that are obtained using a camera. In this paper we propose a new sensing mechanism based on the Doppler effect to capture the patterns of motion of talking faces. We incident an ultrasonic tone on subjects' faces and capture the reflected signal. When the subject talks, different parts of their face move with different velocities in a characteristic manner. Each of these velocities imparts a different Doppler shift to the reflected ultrasonic signal. Thus, the set of frequencies in the reflected ultrasonic signal is characteristic of the subject. We show that even using a simple feature computation scheme to characterize the spectrum of the reflected signal, and a simple GMM based Bayesian classifier, we are able to recognize talkers with an accuracy of over 90%. Interestingly, we are also able to identify the gender of the talker with an accuracy of over 90%. Kaustubh Kalgaonkar, Bhiksha Raj |
FG | 1 |
| 2008 | Vocal tract area based formant tracking using particle filterabstractThis paper presents a novel method for estimating formant frequencies and bandwidths based on an underlying vocal tract model. A novel statistical model for vocal tract cross-sectional areas is developed which allows computation of full likelihood functions. Modifications to the basic particle filter algorithm have also been developed to help combat both diversity depletion and convergence problems. The performance of the method is evaluated against hand labeled formant database (L. Deng et al., 2003). Kaustubh Kalgaonkar, Mark A. Clements |
ICASSP | 1 |
| 2008 | Ultrasonic Doppler sensor for speaker recognitionabstractIn this paper we present a novel use of an acoustic Doppler sonar for multi-modal speaker identification. An ultrasonic emitter directs a 40 kHz tone toward the speaker. Reflections from the speaker's face are recorded as the speaker talks. The frequency of the tone is modified by the velocity of the facial structures it is reflected by. The received ultrasonic signal thus contains an entire spectrum of frequencies representing the set of all velocities of facial components. The pattern of frequencies in the reflected signal is observed to be typical of the speaker. The captured ultrasonic signal is synchronously analyzed with the corresponding voice signal to extract specific characteristics that can be used to identify the speaker. Experiments show that the information this can result in significant improvements in speaker identification accuracy both under clean conditions and in noise. Kaustubh Kalgaonkar, Bhiksha Raj |
ICASSP | 1 |
| 2007 | Acoustic Doppler sonar for gait recoginationabstractA person's gait is a characteristic that might be employed to identify him/her automatically. Conventionally, automatic for gait-based identification of subjects employ video and image processing to characterize gait. In this paper we present an Acoustic Doppler Sensor(ADS) based technique for the characterization of gait. The ADS is very inexpensive sensor that can be built using off-the-shelf components, for under $20 USD at today's prices. We show that remarkably good gait recognition is possible with the ADS sensor. Kaustubh Kalgaonkar, Bhiksha Raj |
AVSS | 1 |
| 2007 | Sensor and Data Systems, Audio-Assisted Cameras and Acoustic Doppler SensorsabstractIn this chapter we present two technologies for sensing and surveillance -audio-assisted cameras and acoustic Doppler sensors for gait recognition. Kaustubh Kalgaonkar, Paris Smaragdis, Bhiksha Raj |
CVPR | 1 |
| 2007 | Vocal tract and area function estimation with both lip and glottal losses
Kaustubh Kalgaonkar, Mark A. Clements |
INTERSPEECH | 1 |
| 2007 | Ultrasonic Doppler Sensor for Voice Activity DetectionabstractThis letter describes a robust voice activity detector using an ultrasonic Doppler sonar device. An ultrasonic beam is incident on the talker's face. Facial movements result in Doppler frequency shifts in the reflected signal that are sensed by an ultrasonic sensor. Speech-related facial movements result in identifiable patterns in the spectrum of the received signal that can be used to identify speech activity. These sensors are not affected by even high levels of ambient audio noise. Unlike most other non-acoustic sensors, the device need not be taped to a talker. A simple yet robust method of extracting the voice activity information from the ultrasonic Doppler signal is developed and presented in this letter. The algorithm is seen to be very effective and robust to noise, and it can be implemented in real time. Kaustubh Kalgaonkar, Rongquiang Hu, Bhiksha Raj |
IEEE Signal Process. Lett. | 1 |
| 2006 | An acoustic Doppler-Based Front End for Hands Free spoken User InterfacesabstractTwo major problems facing hands-free spoken user interfaces are first, to be able to identify accurately when it is addressed and second to effectively determine the start and end points of utterances so that proper sections of speech can be used for further processing and recognition tasks. Accurate determination of end points not only aids accurate noise estimation but also improves speech recognition accuracy. This paper presents an acoustic Doppler sensor based front end for a hands-free spoken user device which provides an efficient and effective solution to both the aforementioned problems. Kaustubh Kalgaonkar, Bhiksha Raj |
SLT | 1 |