VLDB 2026 Research / reviewers in the wild / expert
Srikanth Vishnubhotla
dblp:99/8055
· DBLP profile ↗
16ranked-venue papers
5as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech TranslationabstractAudio-Visual Speech-to-Speech Translation (AVS2S) typically prioritizes improving translation quality and naturalness. However, an equally critical aspect in audio-visual content is lip-synchrony—ensuring that the movements of the lips match the spoken content—essential for maintaining realism in dubbed videos. Despite its importance, the inclusion of lip-synchrony constraints in AVS2S models has been largely overlooked. This study addresses this gap by integrating a lip-synchrony loss into the training process of AVS2S models. Our proposed method significantly enhances lip-synchrony in direct audio-visual speechto-speech translation, achieving an average LSE-D score of 10.67, representing a 9.2% reduction in LSE-D over a strong baseline across four language pairs. Additionally, it maintains the naturalness and high quality of the translated speech when overlaid onto the original video, without any degradation in translation quality. Lucas Goncalves, Prashant Mathur, Xing Niu 0001, Chandrashekhar Lavania, Brady Houston, Srikanth Vishnubhotla, Lijia Sun, Anthony Ferritto |
ICASSP | 6 |
| 2025 | Knowledge Distillation From Ensemble for Spoken Language IdentificationabstractSpoken language identification (LID) has seen substantial performance gains with the rise of large-scale models. However, these models are often computationally expensive and impractical for many real-world applications. In this work, we propose a novel knowledge distillation from ensemble framework to address this challenge. By distilling an ensemble of large LID models into a single, more efficient student, we achieve comparable or even superior performance while reducing computational cost by 67%. Our approach yields a student model with less than 10% the size of a 200M+ parameter teacher ensemble, yet outperforming a 140M parameter teacher by 13% relative. Additionally, combining our distillation technique with decoupled knowledge distillation leads to substantial gains (50% relative), especially for confusable and low-resource languages in the FLEURS dataset. Raghuveer Peri, Seyed Omid Sadjadi, Daniel Garcia-Romero, Srikanth Vishnubhotla, Kyu J. Han |
ICASSP | 4 |
| 2025 | Defending Speech-enabled LLMs Against Adversarial Jailbreak Threats
Antonios Alexos, Raghuveer Peri, Sai Muralidhar Jayanthi, Metehan Cekic, Srikanth Vishnubhotla, Kyu J. Han, Srikanth Ronanki |
INTERSPEECH | 5 |
| 2024 | Improving Multilingual ASR Robustness to Errors in Language Input
Brady Houston, Omid Sadjadi, Zejiang Hou, Srikanth Vishnubhotla, Kyu J. Han |
INTERSPEECH | 4 |
| 2024 | SWAN: SubWord Alignment Network for HMM-free word timing estimation in end-to-end automatic speech recognition
Woo Hyun Kang, Srikanth Vishnubhotla, Rudolf Braun, Yogesh Virkar, Raghuveer Peri, Kyu J. Han |
INTERSPEECH | 2 |
| 2021 | Knowledge Transfer for Efficient on-Device False Trigger MitigationabstractIn this paper, we address the task of determining whether a given utterance is directed towards a voice-enabled smart-assistant device or not. An undirected utterance is termed as a "false trigger" and false trigger mitigation (FTM) is essential for designing a privacy-centric non-intrusive smart assistant. The directedness of an utterance can be identified by running automatic speech recognition (ASR) and determining the user intent by analyzing the ASR transcript. Yet, in case of a false trigger, transcribing the audio using ASR itself is strongly undesirable. To alleviate this issue, we propose an LSTM-based FTM architecture which determines the user intent from acoustic features directly without explicitly generating ASR transcripts from the audio. The proposed models are small-footprint and can be run on-device with limited computational resources. During training, the model parameters are optimized using a knowledge transfer approach where a more accurate self-attention graph neural network model [1] serves as the teacher. Given the whole audio snippets, our approach mitigates 87% of false triggers at 99% true positive rate (TPR), and in a streaming audio scenario, the system listens to only 1.69s of the false trigger audio before rejecting it while achieving the same TPR. Pranay Dighe, Erik Marchi, Srikanth Vishnubhotla, Sachin Kajarekar, Devang Naik |
ICASSP | 3 |
| 2020 | Lattice-Based Improvements for Voice Triggering Using Graph Neural NetworksabstractVoice-triggered smart assistants often rely on detection of a trigger-phrase before they start listening for the user request. Mitigation of false triggers is an important aspect of building a privacy-centric non-intrusive smart assistant. In this paper, we address the task of false trigger mitigation (FTM) using a novel approach based on analyzing automatic speech recognition (ASR) lattices using graph neural networks (GNN). The proposed approach uses the fact that decoding lattice of a falsely triggered audio exhibits uncertainties in terms of many alternative paths and unexpected words on the lattice arcs as compared to the lattice of a correctly triggered audio. A pure trigger-phrase detector model doesn't fully utilize the intent of the user speech whereas by using the complete decoding lattice of user audio, we can effectively mitigate speech not intended for the smart assistant. We deploy two variants of GNNs in this paper based on 1) graph convolution layers and 2) self-attention mechanism respectively. Our experiments demonstrate that GNNs are highly accurate in FTM task by mitigating ~87% of false triggers at 99% true positive rate (TPR). Furthermore, the proposed models are fast to train and efficient in parameter requirements. Pranay Dighe, Saurabh Adya, Nuoyu Li, Srikanth Vishnubhotla, Devang Naik, Adithya Sagar, Stephen G. Pulman, Jason D. Williams |
ICASSP | 4 |
| 2020 | Complementary Language Model and Parallel Bi-LRNN for False Trigger MitigationabstractFalse triggers in voice assistants are unintended invocations of the assistant, which not only degrade the user experience but may also compromise privacy. False trigger mitigation (FTM) is a process to detect the false trigger events and respond appropriately to the user. In this paper, we propose a novel solution to the FTM problem by introducing a parallel ASR decoding process with a special language model trained from "out-of-domain" data sources. Such language model is complementary to the existing language model optimized for the assistant task. A bidirectional lattice RNN (Bi-LRNN) classifier trained from the lattices generated by the complementary language model shows a $38.34\%$ relative reduction of the false trigger (FT) rate at the fixed rate of $0.4\%$ false suppression (FS) of correct invocations, compared to the current Bi-LRNN model. In addition, we propose to train a parallel Bi-LRNN model based on the decoding lattices from both language models, and examine various ways of implementation. The resulting model leads to further reduction in the false trigger rate by $10.8\%$. Rishika Agarwal, Xiaochuan Niu, Pranay Dighe, Srikanth Vishnubhotla, Sameer Badaskar, Devang Naik |
INTERSPEECH | 4 |
| 2013 | A novel single channel speech enhancement approach by combining Wiener filter and dictionary learningabstractIn this paper, a novel algorithm named Sparsity-based Wiener plus Dictionary Learning (SWDL) is proposed for single channel speech enhancement. SWDL combines both Wiener filter and dictionary learning technique. The Wiener filter is used to ensure the enhanced speech is statistically optimal, while the dictionary learning technique is used to improve the enhanced speech quality and intelligibility by utilizing speech-specific information. Such information is incorporated in the pre-trained speech dictionary that can sparsely represent the clean speech spectra. When applied to the TIM-IT database, SWDL outperforms the Log Mean Square-Error Short-Time Spectra Amplitude estimator (LSTSA) according to four different objective metrics measuring speech quality and intelligibility. Subjective tests also show that SWDL produces better speech quality and intelligibility than LSTSA. Hung-Wei Tseng 0004, Srikanth Vishnubhotla, Mingyi Hong 0001, Jinjun Xiao, Zhi-Quan Luo, Tao Zhang 0024 |
ICASSP | 2 |
| 2013 | A single channel speech enhancement approach by combining statistical criterion and multi-frame sparse dictionary learningabstractIn this paper, we consider the single-channel speech enhancement problem, in which a clean speech signal needs to be estimated from a noisy observation. To capture the characteristics of both the noise and speech signals, we combine the well-known Short-Time-Spectrum-Amplitude (STSA) estimator with a machine learning based technique called Multi-frame Sparse Dictionary Learning (MSDL). The former utilizes statistical information for denoising, while the latter helps better preserve speech, especially its temporal structure. The proposed algorithm, named STSA-MSDL, outperforms standard statistical algorithms such as the Wiener filter, STSA estimator, as well as dictionary based algorithms when applied to the TIMIT database, using four different objective metrics that measure speech intelligibility, speech distortion, background noise reduction, and the overall quality. Hung-Wei Tseng 0004, Srikanth Vishnubhotla, Mingyi Hong 0001, Xiangfeng Wang 0001, Jinjun Xiao, Zhi-Quan Luo, Tao Zhang 0024 |
INTERSPEECH | 2 |
| 2012 | Annoyance perception and modeling for hearing-impaired listenersabstractPerceptual annoyance of environmental sounds is measured for normal-hearing and hearing-impaired listeners under iso-level and iso-loudness conditions. Data from the hearing-impaired listeners shows similar trends to that from normal-hearing listeners, but with greater variability across individuals. A regression model based on the statistics of specific loudness and other perceptual features is fit to the data from the normal-hearing listeners, and is used to predict annoyance for the hearing-impaired listeners. Differences across the subject populations are discussed. Srikanth Vishnubhotla, Jinjun Xiao, Buye Xu, Martin F. McKinney, Tao Zhang 0024 |
ICASSP | 1 |
| 2010 | An autoencoder neural-network based low-dimensionality approach to excitation modeling for HMM-based text-to-speechabstractHMM-TTS synthesis is a popular approach toward flexible, low-footprint, data driven systems that produce highly intelligible speech. In spite of these strengths, speech generated by these systems exhibit some degradation in quality, attributable to an inadequacy in modeling the excitation signal that drives the parametric models of the vocal tract. This paper proposes a novel method for modeling the excitation as a low-dimensional set of coefficients, based on a non-linear map learned through an autoencoder. Through analysis-and-resynthesis experiments, and a formal listening test, we show that this model produces speech of higher perceptual quality compared to conventional pulse-excited speech signals at the p <; 0.01 significance level. Srikanth Vishnubhotla, Raul Fernandez, Bhuvana Ramabhadran |
ICASSP | 1 |
| 2009 | An algorithm for speech segregation of co-channel speechabstractThis paper introduces an algorithm to separate speech streams from a single-channel speech mixture. Most current speech segregation algorithms allocate speech regions to participating speakers depending on which speaker dominates in which spectro-temporal region. The proposed method is a different approach to speech segregation, in that it separates the participating speaker streams rather than decide in the favor of the dominating speaker. The algorithm depends on a lease-squares fitting approach to model the speech mixture as a sum of complex exponentials. The algorithm gives results that are better than an existent algorithm when tested on the same task. The performance on a different database yielded good segregation results, even for Target-to-Masker ratios as low as -15 dB. The algorithm has immense promise for improvement and practical implementation. Srikanth Vishnubhotla, Carol Y. Espy-Wilson |
ICASSP | 1 |
| 2008 | An algorithm for multi-pitch tracking in co-channel speechabstractMost multi-pitch algorithms are tested for performance only in voiced regions of speech, and are prone to yield pitch estimates even when the participating speakers are unvoiced. This paper presents a multi-pitch algorithm that detects the voiced and unvoiced regions in a mixture of two speakers, identifies the number of speakers in voiced regions, and yields the pitch estimates of each speaker in those regions. The algorithm relies on the 2-Dimensional AMDF for estimating the periodicity of the signal, and uses the temporal evolution of the 2-D AMDF to estimate the number of speakers present in periodic regions. Evaluation of this algorithm on a frame-wise basis demonstrates accurate voiced / unvoiced decisions and also gives pitch estimation results comparable to the state of the art. The pitch estimation errors are quantitatively analyzed and shown to be resulting partly from speaker domination & pitch matching between speakers. Srikanth Vishnubhotla, Carol Y. Espy-Wilson |
INTERSPEECH | 1 |
| 2006 | A new set of features for text-independent speaker identificationabstractThe success of a speaker identification system depends largely on the set of features used to characterize speaker-specific information. In this paper, we discuss a small set of low-level acoustic parameters that capture information about the speaker’s source, vocal tract size and vocal tract shape. We demonstrate that the set of eight acoustic parameters has comparable performance to the standard sets of 26 or 39 MFCCs for the speaker identification task. Gaussian Mixture Models were used for constructing speaker models. Index Terms: speaker identification, acoustic parameters, Carol Y. Espy-Wilson, Sandeep Manocha, Srikanth Vishnubhotla |
INTERSPEECH | 3 |
| 2006 | Automatic detection of irregular phonation in continuous speechabstractVoice quality is one of the most important source characteristics of a speaker’s speech production process, and creakiness is one of the variations of voice quality. This paper describes the development of an algorithm to automatically detect irregular phonation, including creakiness and other variations, in continuous running speech. The algorithm is an extension of the Aperiodicity, Periodicity and Pitch (APP) Detector. The algorithm has been run on 485 files of the TIMIT database, which contained 677 instances of irregular phonation. The test set comprised of 97 speakers, of which 57 were male and 40 were female. The algorithm has been found to give an accuracy of 86.7 % on average, with performance being almost the same for both male and female speakers. Automatic detection of irregular phonation should help characterize speakers for speaker identification applications. Srikanth Vishnubhotla, Carol Y. Espy-Wilson |
INTERSPEECH | 1 |