Srikanth Raj Chetupalli

dblp:15/10648 · also Ch. Srikanth Raj, Srikanth Raj Chetupally · DBLP profile ↗
← Back
24ranked-venue papers
9as first author
17since 2021 · last 2026
0000-0002-2186-5420ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 7 first-author · 14 since 2021Artificial intelligence and machine learning · 17 · 7 first-author · 12 since 2021
YearPublicationVenuePosition
2026 Speaker Anonymization for Children's Oral Reading Assessment
abstract
Speaker anonymization aims to modify the speech signal in order to protect the identity of a speaker while preserving the linguistic content. Despite the increasing use of children's voices in educational applications, such as oral reading fluency (ORF) assessment, there is little work on the anonymization aspects. In this work, we investigate the effectiveness of available speaker anonymization methods drawing from traditional speech-production based approaches and a neural codec based method. We investigate the trade-off between privacy protection, measured as the degree of anonymity, and utility preservation, which in the current context of ORF assessment, includes the segmental and suprasegmental features of children’s read speech utterances. We report objective and subjective evaluations using two child-speaker datasets: MPS and SpeechOcean. Our objective evaluation results indicate that the speech-production based method of vocal tract length normalization coupled with pitch-transposition achieves the best balance between privacy and utility. Subjective listening results indicate that naturalness is achievable across methods while the neural method fails to preserve age characteristics, which are more easily controlled by the speech-production driven methods.
Sandipan Dhar, Srikanth Raj Chetupalli, Preeti Rao
AAAI2
2025 AdaBit-TasNet: Speech Separation with Inference Adaptable Precision
abstract
Deploying advanced neural network-based speech separation (SS) models on resource-constrained devices is challenging due to their high computational and memory demands. Conventional network compression techniques, such as pruning and quantization, can alleviate these demands without significantly compromising performance. However, they lack the flexibility to select the compression factor at run-time to suit varying operating conditions, such as changing computational and energy budgets in battery-powered devices. In this paper, we introduce AdaBit-TasNet, an adaptable-precision network (APN) for SS that enables flexible bit-width selection during inference. Experimental evaluation on the Libri2Mix dataset demonstrates that AdaBit-TasNet achieves comparable performance to that of individually trained fixed-precision networks at several bitwidths.
Mohamed Elminshawi, Srikanth Raj Chetupalli, Emanuël A. P. Habets
ASRU2
2025 Low-Complexity Neural Speech Dereverberation With Adaptive Target Control
abstract
Existing neural network-based speech dereverberation approaches use a fixed-length early reflection part of the reverberant signal as the target for estimation, irrespective of the severity of reverberation. Such an approach often leads to distortions in the enhanced signals in highly reverberant scenarios. In practice, while some listeners prefer minimal speech distortions, others have a higher tolerance for distortions and prefer a clean signal. To address these points, we propose a novel target definition and a low-complexity neural network for user-controlled single-channel dereverberation. Our target definition is parameterized by the relative amount of reduction in the late reverberation energy. Further, the same parameter is passed as a control input to the dereverberation network for adaptability during inference. Objective and subjective evaluation shows the feasibility of the proposed dereverberation approach.
Nagashree K. S. Rao, Srikanth Raj Chetupalli, Shrishti Saha Shetu, Emanuël A. P. Habets, Oliver Thiergart
ICASSP2
2024 Dynamic Slimmable Network for Speech Separation
abstract
Neural networks for speech separation generally exhibit high computational costs and large memory footprints. Moreover, typical separation networks have a fixed computational graph that processes all input frames at a uniform computational cost, even though intensive processing may not be necessary for frames containing silence or a single active speaker. Addressing this computational inefficiency is especially crucial when these networks are deployed on resource-constrained devices. In this letter, we propose a dynamic slimmable network for speech separation that mitigates the computational inefficiency of existing networks. We introduce slimmable layers with a gating mechanism that can adapt their computational complexity based on the input characteristics. As an example, we propose to use the slimmable layers in the intra-chunk blocks of a dual-path structure-based network to facilitate adaptation based on the local characteristics of the input signal. Experimental evaluation on simulated two-speaker mixtures from the WSJ0-2mix dataset demonstrates that the proposed method substantially reduces the computational cost while maintaining comparable performance to fully utilized static networks.
Mohamed Elminshawi, Srikanth Raj Chetupalli, Emanuël A. P. Habets
IEEE Signal Process. Lett.2
2023 Beamformer-Guided Target Speaker Extraction
abstract
We propose a Beamformer-guided Target Speaker Extraction (BG-TSE) method to extract a target speaker’s voice from a multi-channel recording informed by the direction of arrival of the target. The proposed method employs a front-end beamformer steered towards the target speaker to pro-vide an auxiliary signal to a single-channel TSE system. By allowing for time-varying embeddings in the single-channel TSE block, the proposed method fully exploits the correspondence between the front-end beamformer output and the tar-get speech in the microphone signal. Experimental evaluation on simulated multi-channel 2-speaker mixtures, in both anechoic and reverberant conditions, demonstrates the advantage of the proposed method compared to recent single-channel and multi-channel baselines.
Mohamed Elminshawi, Srikanth Raj Chetupalli, Emanuël A. P. Habets
ICASSP2
2023 Multi-Microphone Speaker Separation by Spatial Regions
abstract
We consider the task of region-based source separation of reverberant multi-microphone recordings. We assume pre-defined spatial regions with a single active source per region. The objective is to estimate the signals from the individual spatial regions as captured by a reference microphone while retaining a correspondence between signals and spatial regions. We propose a data-driven approach using a modified version of a state-of-the-art network, where different layers model spatial and spectro-temporal information. The network is trained to enforce a fixed mapping of regions to network outputs. Using speech from LibriMix, we construct a data set specifically designed to contain the region information. Additionally, we train the network with permutation invariant training. We show that both training methods result in a fixed mapping of regions to network outputs, achieve comparable performance, and that the networks exploit spatial information. The proposed network outperforms a baseline network by 1.5 dB in scale-invariant signal-to-distortion ratio.
Julian Wechsler, Srikanth Raj Chetupalli, Wolfgang Mack, Emanuël A. P. Habets
ICASSP2
2023 Speaker Counting and Separation From Single-Channel Noisy Mixtures
abstract
We address the problem of speaker counting and separation from a noisy, single-channel, multi-source, recording. Most of the works in the literature assume mixtures containing two to five speakers. In this work, we consider noisy speech mixtures with one to five speakers and noise-only recordings. We propose a deep neural network (DNN) architecture, that predicts a speaker count of zero for noise-only recordings and predicts the individual clean speaker signals and speaker count for mixtures of one to five speakers. The DNN is composed of transformer layers and processes the recordings using the long-time and short-time sequence modeling approach to masking in a learned time-feature domain. The network uses an encoder-decoder attractor module with long-short term memory units to generate a variable number of outputs. The network is trained with simulated noisy speech mixtures composed of the speech recordings from WSJ0 corpus, and noise recordings from the WHAM! corpus. We show that the network achieves 99% speaker counting accuracy and more than 19 dB improvement in the scale-invariant signal-to-noise ratio for mixtures of up to three speakers.
Srikanth Raj Chetupalli, Emanuël A. P. Habets
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 AmbiSep: Joint Ambisonic-to-Ambisonic Speech Separation and Noise Reduction
abstract
Blind separation of the sounds in an Ambisonic sound scene is a challenging problem, especially when the spatial impression of these sounds needs to be preserved. In this work, we consider Ambisonic-to-Ambisonic separation of reverberant speech mixtures, optionally containing noise. A supervised learning approach is adopted utilizing a transformer-based deep neural network, denoted by AmbiSep. AmbiSep takes mutichannel Ambisonic signals as input and estimates separate multichannel Ambisonic signals for each speaker while preserving their spatial images including reverberation. The GPU memory requirement of AmbiSep during training increases with the number of Ambisonic channels. To overcome this issue, we propose different aggregation methods. The model is trained and evaluated for first-order and second-order Ambisonics using simulated speech mixtures. Experimental results show that the model performs well on clean and noisy reverberant speech mixtures, and also generalizes to mixtures generated with measured Ambisonic impulse responses.
Adrian Herzog, Srikanth Raj Chetupalli, Emanuël A. P. Habets
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 The Second Dicova Challenge: Dataset and Performance Analysis for Diagnosis of Covid-19 Using Acoustics
abstract
The Second Diagnosis of COVID-19 using Acoustics (DiCOVA) Challenge aimed at accelerating the research in acoustics based detection of COVID-19, a topic at the intersection of acoustics, signal processing, machine learning, and healthcare. This paper presents the details of the challenge, which was an open call for researchers to analyze a dataset of audio recordings consisting of breathing, cough and speech signals. This data was collected from individuals with and without COVID-19 infection, and the task in the challenge was a two-class classification. The development set audio recordings were collected from 965 (172 COVID-19 positive) individuals, while the evaluation set contained data from 471 individuals (71 COVID-19 positive). The challenge featured four tracks, one associated with each sound category of cough, speech and breathing, and a fourth fusion track. A baseline system was also released to benchmark the participants. In this paper, we present an overview of the challenge, the rationale for the data collection and the baseline system. Further, a performance analysis for the systems submitted by the 21 participating teams in the leaderboard is also presented.
Neeraj Kumar Sharma 0001, Srikanth Raj Chetupalli, Debarpan Bhattacharya, Debottam Dutta, Pravin Mote, Sriram Ganapathy
ICASSP2
2022 Coswara: A website application enabling COVID-19 screening by analysing respiratory sound samples and health symptoms
Debarpan Bhattacharya, Debottam Dutta, Neeraj Kumar Sharma 0001, Srikanth Raj Chetupalli, Pravin Mote, Sriram Ganapathy, Chandrakiran C, Sahiti Nori, Suhail K. K, Sadhana Gonuguntla, Murali Alagesan
INTERSPEECH4
2022 Analyzing the impact of SARS-CoV-2 variants on respiratory sound signals
abstract
The COVID-19 outbreak resulted in multiple waves of infections that have been associated with different SARS-CoV-2 variants.Studies have reported differential impact of the variants on respiratory health of patients.We explore whether acoustic signals, collected from COVID-19 subjects, show computationally distinguishable acoustic patterns suggesting a possibility to predict the underlying virus variant.We analyze the Coswara dataset which is collected from three subject pools, namely, i) healthy, ii) COVID-19 subjects recorded during the delta variant dominant period, and iii) data from COVID-19 subjects recorded during the omicron surge.Our findings suggest that multiple sound categories, such as cough, breathing, and speech, indicate significant acoustic feature differences when comparing COVID-19 subjects with omicron and delta variants.The classification areas-under-the-curve are significantly above chance for differentiating subjects infected by omicron from those infected by delta.Using a score fusion from multiple sound categories, we obtained an area-under-the-curve of 89% and 52.4% sensitivity at 95% specificity.Additionally, a hierarchical three class approach was used to classify the acoustic data into healthy and COVID-19 positive, and further COVID-19 subjects into delta and omicron variants providing high level of 3-class classification accuracy.These results suggest new ways for designing sound based COVID-19 diagnosis approaches.
Debarpan Bhattacharya, Debottam Dutta, Neeraj Kumar Sharma 0001, Srikanth Raj Chetupalli, Pravin Mote, Sriram Ganapathy, Chandrakiran C, Sahiti Nori, Suhail K. K, Sadhana Gonuguntla, Murali Alagesan
INTERSPEECH4
2022 Speaker conditioned acoustic modeling for multi-speaker conversational ASR
abstract
In this paper, we propose a novel approach for the transcription of speech conversations with natural speaker overlap, from single channel speech recordings.The proposed model is a combination of a speaker diarization system and a hybrid automatic speech recognition (ASR) system.The speaker conditioned acoustic model (SCAM) in the ASR system consists of a series of embedding layers which use the speaker activity inputs from the diarization system to derive speaker specific embeddings.The output of the SCAM are speaker specific senones that are used for decoding the transcripts for each speaker in the conversation.In this work, we experiment with the automatic speaker activity decisions generated using an endto-end speaker diarization system.A joint learning approach is also proposed where the diarization model and the ASR acoustic model are jointly optimized.The experiments are performed on the mixed-channel two speaker recordings from the Switchboard corpus of telephone conversations.In these experiments, we show that the proposed acoustic model, incorporating speaker activity decisions and joint optimization, improves significantly over the ASR system with explicit source filtering (relative improvements of 12% in word error rate (WER) over the baseline system).
Srikanth Raj Chetupalli, Sriram Ganapathy
INTERSPEECH1
2022 Speech Separation for an Unknown Number of Speakers Using Transformers With Encoder-Decoder Attractors
abstract
5393
Srikanth Raj Chetupalli, Emanuël A. P. Habets
INTERSPEECH1
2022 Towards sound based testing of COVID-19 - Summary of the first Diagnostics of COVID-19 using Acoustics (DiCOVA) Challenge
Neeraj Kumar Sharma 0001, Ananya Muguli, Prashant Krishnan V, Srikanth Raj Chetupalli, Sriram Ganapathy
Comput. Speech Lang.5
2021 Investigating Feature Selection and Explainability for COVID-19 Diagnostics from Cough Sounds
abstract
In this paper, we propose an approach to automatically classify COVID-19 and non-COVID-19 cough samples based on the combination of both feature engineering and deep learning models. In the feature engineering approach, we develop a support vector machine classifier over high dimensional (6373D) space of acoustic features. In the deep learning-based approach, on the other hand, we apply a convolutional neural network trained on the log-mel spectrograms. These two methodologically diverse models are then combined by fusing the probability scores of the models. The proposed system, which ranked 9th on the 2021 Diagnosing COVID-19 using Acoustics (Di- COVA) challenge leaderboard, obtained an area under the receiver operating characteristic curve (AUC) of 0:81 on the blind test data set, which is a 10:9 absolute improvement compared to the baseline. Moreover, we analyze the explainability of the deep learning-based model when detecting COVID-19 from cough signals. Copyright © 2021 ISCA.
Flávio Ávila, Amir Hossein Poorjam, Deepak Mittal, Charles Dognin, Ananya Muguli, Srikanth Raj Chetupalli, Sriram Ganapathy, Maneesh Kumar Singh 0001
Interspeech7
2021 DiCOVA Challenge: Dataset, Task, and Baseline System for COVID-19 Diagnosis Using Acoustics
abstract
The DiCOVA challenge aims at accelerating research in diagnosing COVID-19 using acoustics (DiCOVA), a topic at the intersection of speech and audio processing, respiratory health diagnosis, and machine learning.This challenge is an open call for researchers to analyze a dataset of sound recordings, collected from COVID-19 infected and non-COVID-19 individuals, for a two-class classification.These recordings were collected via crowdsourcing from multiple countries, through a website application.The challenge features two tracks, one focusing on cough sounds, and the other on using a collection of breath, sustained vowel phonation, and number counting speech recordings.In this paper, we introduce the challenge and provide a detailed description of the task, and present a baseline system for the task.
Ananya Muguli, Lancelot Pinto, Nirmala R., Neeraj Kumar Sharma 0001, Prashant Krishnan V, Prasanta Kumar Ghosh, Shrirama Bhat, Srikanth Raj Chetupalli, Sriram Ganapathy, Shreyas Ramoji, Viral Nanda
Interspeech9
2021 LEAP Submission for the Third DIHARD Diarization Challenge
abstract
The LEAP submission for DIHARD-III challenge is described in this paper. The proposed system is composed of a speech bandwidth classifier, and diarization systems fine-tuned for narrowband and wideband speech separately. We use an end-to-end speaker diarization system for the narrowband conversational telephone speech recordings. For the wideband multi-speaker recordings, we use a neural embedding based clustering approach, similar to the baseline system. The embeddings are extracted from a time-delay neural network (called x-vectors) followed by the graph based path integral clustering (PIC) approach. The LEAP system showed 24% and 18% relative improvements for Track-1 and Track-2 respectively over the baseline system provided by the organizers. This paper describes the challenge submission, the post-evaluation analysis and improvements observed on the DIHARD-III dataset.
Prachi Singh, Rajat Varma, Venkat Krishnamohan, Srikanth Raj Chetupalli, Sriram Ganapathy
Interspeech4
2020 Context Dependent RNNLM for Automatic Transcription of Conversations
abstract
Conversational speech, while being unstructured at an utterance level, typically has a macro topic which provides larger context spanning multiple utterances. The current language models in speech recognition systems using recurrent neural networks (RNNLM) rely mainly on the local context and exclude the larger context. In order to model the long term dependencies of words across multiple sentences, we propose a novel architecture where the words from prior utterances are converted to an embedding. The relevance of these embeddings for the prediction of next word in the current sentence is found using a gating network. The relevance weighted context embedding vector is combined in the language model to improve the next word prediction, and the entire model including the context embedding and the relevance weighting layers is jointly learned for a conversational language modeling task. Experiments are performed on two conversational datasets - AMI corpus and the Switchboard corpus. In these tasks, we illustrate that the proposed approach yields significant improvements in language model perplexity over the RNNLM baseline. In addition, the use of proposed conversational LM for ASR rescoring results in absolute WER reduction of $1.2$\% on Switchboard dataset and $1.0$\% on AMI dataset over the RNNLM based ASR baseline.
Srikanth Raj Chetupalli, Sriram Ganapathy
INTERSPEECH1
2020 Coswara - A Database of Breathing, Cough, and Voice Sounds for COVID-19 Diagnosis
abstract
The COVID-19 pandemic presents global challenges transcending boundaries of country, race, religion, and economy.The current gold standard method for COVID-19 detection is the reverse transcription polymerase chain reaction (RT-PCR) testing.However, this method is expensive, time-consuming, and violates social distancing.Also, as the pandemic is expected to stay for a while, there is a need for an alternate diagnosis tool which overcomes these limitations, and is deployable at a large scale.The prominent symptoms of COVID-19 include cough and breathing difficulties.We foresee that respiratory sounds, when analyzed using machine learning techniques, can provide useful insights, enabling the design of a diagnostic tool.Towards this, the paper presents an early effort in creating (and analyzing) a database, called Coswara, of respiratory sounds, namely, cough, breath, and voice.The sound samples are collected via worldwide crowdsourcing using a website application.The curated dataset is released as open access.As the pandemic is evolving, the data collection and analysis is a work in progress.We believe that insights from analysis of Coswara can be effective in enabling sound based technology solutions for point-of-care diagnosis of respiratory infection, and in the near future this can help to diagnose COVID-19.
Neeraj Kumar Sharma 0001, Prashant Krishnan V, Shreyas Ramoji, Srikanth Raj Chetupalli, Nirmala R., Prasanta Kumar Ghosh, Sriram Ganapathy
INTERSPEECH5
2019 Late Reverberation Cancellation Using Bayesian Estimation of Multi-Channel Linear Predictors and Student's t-Source Prior
abstract
Multi-channel linear prediction (MCLP) can model the late reverberation in the short-time Fourier transform domain using a delayed linear predictor and the prediction residual is taken as the desired early reflection component. Traditionally, a Gaussian source model with time-dependent precision (inverse of variance) is considered for the desired signal. In this paper, we propose a Student's t-distribution model for the desired signal, which is realized as a Gaussian source with a Gamma distributed precision. Further, since the choice of a proper MCLP order is critical, we also incorporate a Gaussian distribution prior for the prediction coefficients and a higher order. We consider a batch estimation scenario and develop variational Bayes expectation maximization (VBEM) algorithm for joint posterior inference and hyper-parameter estimation. This has lead to more accurate and robust estimation of the late reverb component and hence its cancellation, benefitting the desired residual signal estimation. Along with these stochastic models, we formulate single channel output (MISO) and multi channel output (MIMO) schemes using shared priors for the desired signal precision and the estimated MCLP coefficients at each microphone. Experiments using real room impulse responses show improved late reverberation suppression with the proposed VBEM approach over the traditional methods, for different room conditions. Additionally, we achieve a sparse coefficient vector for the MCLP avoiding the criticality of manually choosing the model order. The MIMO formulation is easily extended to include spatial filtering of the enhanced signals, which further improves the estimation of the desired signal.
Srikanth Raj Chetupalli, Thippur V. Sreenivas
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 Joint Bayesian Estimation of Time-Varying LP Parameters and Excitation for Speech
abstract
We consider the joint estimation of time-varying linear prediction (TVLP) filter coefficients and the excitation signal parameters for the analysis of long-term speech segments. Traditional approaches to TVLP estimation assume linear expansion of the coefficients in a set of known basis functions only. But, excitation signal is also time-varying, which affects the estimation of TVLP filter parameters. In this letter, we propose a Bayesian approach, to incorporate the nature of excitation signal and also adapt regularization of the filter parameters. Since the order of the system is not known a priori, we formulate a Gaussian prior for the filter parameters, and the excitation signal is modeled as Gaussian with time-varying Gamma distributed precision. We develop an iterative algorithm for the maximum-likelihood estimation of the posterior distribution of filter parameters and the time-varying precision of the excitation signal, along with the parameters of the prior distribution. We show that the proposed method adapts to different types of excitation signals in speech, and also the time-varying system with unknown model order. The spectral modeling performance for synthetic speech-like signals, quantified using the absolute spectral difference shows that the proposed method estimates the system function more accurately compared to several of the traditional methods.
Srikanth Raj Chetupalli, Thippur V. Sreenivas
IEEE Signal Process. Lett.1
2014 Time varying linear prediction using sparsity constraints
abstract
Time-varying linear prediction has been studied in the context of speech signals, in which the auto-regressive (AR) coefficients of the system function are modeled as a linear combination of a set of known bases. Traditionally, least squares minimization is used for the estimation of model parameters of the system. Motivated by the sparse nature of the excitation signal for voiced sounds, we explore the time-varying linear prediction modeling of speech signals using sparsity constraints. Parameter estimation is posed as a 0-norm minimization problem. The re-weighted 1-norm minimization technique is used to estimate the model parameters. We show that for sparsely excited time-varying systems, the formulation models the underlying system function better than the least squares error minimization approach. Evaluation with synthetic and real speech examples show that the estimated model parameters track the formant trajectories closer than the least squares approach.
Srikanth Raj Chetupalli, Thippur V. Sreenivas
ICASSP1
2012 Joint Pitch-Analysis Formant-Synthesis framework for CS recovery of speech
abstract
A joint analysis-synthesis framework is developed for the compressive sensing (CS) recovery of speech signals. The signal is assumed to be sparse in the residual domain with the linear prediction filter used as the sparse transformation. Importantly this transform is not known apriori, since estimating the predictor filter requires the knowledge of the signal. Two prediction filters, one comb filter for pitch and another all pole formant filter are needed to induce maximum sparsity. An iterative method is proposed for the estimation of both the prediction filters and the signal itself. Formant prediction filter is used as the synthesis transform, while the pitch filter is used to model the periodicity in the residual excitation signal, in the analysis mode. Significant improvement in the LLR measure is seen over the previously reported formant filter estimation.
Srikanth Raj Chetupalli, Thippur V. Sreenivas
INTERSPEECH1
2011 Time-Varying Signal Adaptive Transform and IHT Recovery of Compressive Sensed Speech
Srikanth Raj Chetupalli, Thippur V. Sreenivas
INTERSPEECH1