VLDB 2026 Research / reviewers in the wild / expert
Sunil Kumar Kopparapu
dblp:49/6239
· DBLP profile ↗
40ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-0502-527XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 27 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HamaraAwaz: Advancing Low-Latency Streaming TTS for Multilingual Speech in Indian LanguagesabstractWe present a multilingual, multi-speaker, low-latency speech synthesis system developed by the HamaraAwaz team for Track 1 of the LIMMITS’25 challenge. To improve speaker similarity and naturalness in Indic languages, we build on ParrotTTS. We utilize disentangled self-supervised speech representations and incorporate enhancements such as Byte-Pair Encoding for text representation to reduce latency and relative positional representations to enhance speech quality. The proposed model achieved a naturalness Mean Opinion Score (MOS) of 3.51 and a speaker similarity score of 3.53 in the LIMMITS’25 grand challenge held as part of ICASSP-25. Speech samples are available at https://parrot-tts.github.io/HamaraAwaz/ Neil Kumar Shah, Parth Khadse, Shirish S. Karande, Sunil Kumar Kopparapu |
ICASSP | 4 |
| 2025 | EmbedAug: An Augmentation Scheme for End-to-End Automatic Speech Recognition
Ashish Panda, Sunil Kumar Kopparapu |
INTERSPEECH | 2 |
| 2024 | A Cost Minimization Approach to Fix the Vocabulary Size in a Tokenizer for an End-to-End ASR System
Sunil Kumar Kopparapu, Ashish Panda |
ICPR (31) | 1 |
| 2024 | Reinforcement Learning based Data Augmentation for Noise Robust Speech Emotion Recognition
Sumit Ranjan, Rupayan Chakraborty, Sunil Kumar Kopparapu |
INTERSPEECH | 3 |
| 2023 | A Novel Metric For Evaluating Audio Caption SimilarityabstractAutomatic Audio Captioning (AAC) refers to the task of describing an audio sample in a natural language (NL) text. Unlike NL text generation tasks, which rely on lexical semantic metrics like BLEU for evaluation, the AAC evaluation metric requires acoustic semantics to map NL text corresponding to similar sounds in addition to lexical semantics. In this paper, we propose a novel metric based on Text-to-Audio Grounding (TAG), to incorporate acoustic semantics. Experiments demonstrate our evaluation metric to perform better compared to existing metrics used in NL text and image captioning literature for AAC. Swapnil Bhosale, Rupayan Chakraborty, Sunil Kumar Kopparapu |
ICASSP | 3 |
| 2021 | Deep Lung Auscultation Using Acoustic Biomarkers for Abnormal Respiratory Sound Event DetectionabstractLung Auscultation is a non-invasive process of distinguishing normal respiratory sounds from abnormal ones by analyzing the airflow along the respiratory tract. With developments in the Deep Learning (DL) techniques and wider access to anonymized medical data, automatic detection of specific sounds such as crackles and wheezes have been gaining popularity. In this paper, we propose to use two sets of diversified acoustic biomarkers extracted using Discrete Wavelet Transform (DWT) and deep encoded features from the intermediate layer of a pre-trained Audio Event Detection (AED) model trained using sounds from daily activities. First set of biomarkers highlight the time frequency localization characteristics obtained from DWT coefficients. However, the second set of deep encoded biomarkers captures a generalized reliable representation, and thus indemnifies the scarcity of training samples and the class imbalance in dataset. The model trained using these features achieves a 15.05% increase in terms of the specificity over the baseline model that uses spectrogram features. Moreover, ensemble of DWT features and deep encoded feature based models show absolute improvements of 8.32%, 6.66% and 7.40% in terms of sensitivity, specificity and ICBHI-score, respectively, and clearly outperforms the state-of-the-art with a significant margin. Upasana Tiwari, Swapnil Bhosale, Rupayan Chakraborty, Sunil Kumar Kopparapu |
ICASSP | 4 |
| 2021 | Contrastive Learning of Cough Descriptors for Automatic COVID-19 Preliminary DiagnosisabstractCough sounds as a descriptor have been used for detecting various respiratory ailments based on its intensity, duration of intermediate phase between two cough sounds, repetitions, dryness etc.However, COVID-19 diagnosis using only cough sounds is challenging because of cough being a common symptom among many non COVID-19 health diseases and inherent data imbalance within the available datasets.As one of the approach in this direction, we explore the robustness of multi-domain representation by performing the early fusion over a wide set of temporal, spectral and tempo-spectral handcrafted features, followed by training a Support Vector Machine (SVM) classifier.In our second approach, using a contrastive loss function we learn a latent space from Mel Filter Cepstral Coefficients (MFCCs) where representations belonging to samples having similar cough characteristics are closer.This helps learn representations for the highly varied COVIDnegative class (healthy and symptomatic COVID-negative), by learning multiple smaller clusters.Using only the DiCOVA data, multi-domain features yields an absolute improvement of 0.74% and 1.07%, whereas our second approach shows an improvement of 2.09% and 3.98%, over the blind test and validation set, respectively, when compared with challenge baseline. Swapnil Bhosale, Upasana Tiwari, Rupayan Chakraborty, Sunil Kumar Kopparapu |
Interspeech | 4 |
| 2021 | Automatic speaker independent dysarthric speech intelligibility assessment system
Ayush Tripathi, Swapnil Bhosale, Sunil Kumar Kopparapu |
Comput. Speech Lang. | 3 |
| 2020 | Deep Encoded Linguistic and Acoustic Cues for Attention Based End to End Speech Emotion RecognitionabstractAn End-to-End model with convolutional layers and multi-head self attention mechanism is proposed for Speech Emotion Recognition (SER) task. As inputs, we propose to use both the deep encoded linguistic features that carry the language related context of emotion and the audio spectrogram that are representatives of acoustic cues. To facilitate the deep linguistic feature representation, we use outputs from the intermediate layers of a pre-trained Automatic Speech Recognition (ASR) model, where the layer is selected empirically. The influence of both acoustic and linguistic features, both separately and in combination, for emotion recognition in different scenarios (scripted and spontaneous recording of emotional speech samples) have been studied. Extensive experiments on the standard IEMOCAP database are conducted to investigate the efficacy of our proposed approach. To address the class imbalance, we carried out down sampling and ensembling, which further improved the SER accuracy. Overall, we observe that the acoustic features perform best for improvised recordings which is due to the spontaneity in speech with less linguistic correlation. But the linguistic features are found to be effective for the scripted as well as for the combined (scripted and improvised recordings together) scenario that reflects more linguistic information in spoken utterances. Swapnil Bhosale, Rupayan Chakraborty, Sunil Kumar Kopparapu |
ICASSP | 3 |
| 2020 | Multi-Conditioning and Data Augmentation Using Generative Noise Model for Speech Emotion Recognition in Noisy ConditionsabstractDegradation due to additive noise is a significant road block in the real-life deployment of Speech Emotion Recognition (SER) systems. Most of the previous work in this field dealt with the noise degradation either at the signal or at the feature level. In this paper, to address the robustness aspect of the SER in additive noise scenarios, we propose multi-conditioning and data augmentation using an utterance level parametric Generative noise model. The Generative noise model is designed to generate noise types which can span the entire noise space in the mel-filterbank energy domain. This characteristic of the model renders the system robust against unseen noise conditions. The generated noise types can be used to create multiconditioned data for training the SER systems. Multi-conditioning approach can also be used to increase the training data by many folds where such data is limited. We report the performance of the proposed method on two datasets, namely EmoDB and IEMOCAP. We also explore multi-conditioning and data augmentation using noise samples from NOISEX-92 database. Upasana Tiwari, Meet H. Soni, Rupayan Chakraborty, Ashish Panda, Sunil Kumar Kopparapu |
ICASSP | 5 |
| 2020 | Improved Speaker Independent Dysarthria Intelligibility Classification Using Deepspeech PosteriorsabstractIndividuals with dysarthria are unable to control rapid movement of the velum leading to reduction in intelligibility, audibility, naturalness and efficiency of vocal communication. Automatic intelligibility assessment of dysarthric patients allows clinicians diagnose the impact of therapy and medication and also to plan future course of action. Earlier works have concentrated on building speaker dependent machine learning systems for intelligibility assessment, due to limited availability of data. However, a speaker independent assessment system is of greater use by clinicians. Motivated by this observation, we propose a speaker independent intelligibility assessment system which relies on a novel set of features obtained by processing the output of DeepSpeech, an end to end Speech-to-Text engine. All experiments have been performed on the Universal Access Speech database. An accuracy of 53.9% was obtained using Support Vector Machine based four-class classification system for the speaker independent scenario while the accuracy obtained for the speaker dependent scenario is 97.4%. Ayush Tripathi, Swapnil Bhosale, Sunil Kumar Kopparapu |
ICASSP | 3 |
| 2020 | A Novel Approach for Intelligibility Assessment in Dysarthric SubjectsabstractDysarthria is a motor speech impairment caused by muscle weakness. Individuals, with this condition, are unable to control rapid movement of the velum leading to reduction in intelligibility, audibility, naturalness and efficiency of vocal communication. Systems that can assess intelligibility of dysarthric speech can help clinicians diagnose the impact of therapy and medication. In the paper, we propose a usable novel method to assess intelligibility of dysarthric speakers. The approach is based on the observation that the performance of a speech recognition engine deteriorates with increase in severity of the disorder. The mismatch between the original word and the recognized string is exploited to compute the dysarthria intelligibility score. Experiments on UA speech corpus show that the computed intelligibility score exhibits a significant correlation with perceptually assessed intelligibility scores. We further show that a small set of words spoken by the dysarthric subject is sufficient to assess the speech intelligibility reliably. Ayush Tripathi, Swapnil Bhosale, Sunil Kumar Kopparapu |
ICASSP | 3 |
| 2020 | A Novel Adaptive Minority Oversampling Technique for Improved Classification in Data Imbalanced ScenariosabstractImbalance in the proportion of training samples belonging to different classes often poses performance degradation of conventional classifiers. This is primarily due to the tendency of the classifier to be biased towards the majority classes in the imbalanced dataset. In this paper, we propose a novel three step technique to address imbalanced data. As a first step we significantly oversample the minority class distribution by employing the traditional Synthetic Minority Oversampling Technique (SMOTE) algorithm using the neighborhood of the minority class samples and in the next step we partition the generated samples using a Gaussian-Mixture Model based clustering algorithm. In the final step synthetic data samples are chosen based on the weight associated with the cluster, the weight itself being determined by the distribution of the majority class samples. Extensive experiments on several standard datasets from diverse domains show the usefulness of the proposed technique in comparison with the original SMOTE and its state-of-the-art variants algorithms. Ayush Tripathi, Rupayan Chakraborty, Sunil Kumar Kopparapu |
ICPR | 3 |
| 2020 | Effect of Microphone Position Measurement Error on RIR and its Impact on Speech Intelligibility and Quality
Aditya Raikar, Karan Nathwani, Ashish Panda, Sunil Kumar Kopparapu |
INTERSPEECH | 4 |
| 2019 | Improving ASR Robustness to Perturbed Speech Using Cycle-consistent Generative Adversarial NetworksabstractNaturally introduced perturbations in audio signal, caused by emotional and physical states of the speaker, can significantly degrade the performance of Automatic Speech Recognition (ASR) systems. In this paper, we propose a front-end based on Cycle-Consistent Generative Adversarial Network (CycleGAN) which transforms naturally perturbed speech into normal speech, and hence improves the robustness of an ASR system. The CycleGAN model is trained on non-parallel examples of perturbed and normal speech. Experiments on spontaneous laughter-speech and creaky voice datasets show that the performance of four different ASR systems improve by using speech obtained from CycleGAN based front-end, as compared to directly using the original perturbed speech. Visualization of the features of the laughter perturbed speech and those generated by the proposed front-end further demonstrates the effectiveness of our approach. Sri Harsha Dumpala, Imran A. Sheikh, Rupayan Chakraborty, Sunil Kumar Kopparapu |
ICASSP | 4 |
| 2019 | End-to-End Spoken Language Understanding: Bootstrapping in Low Resource Scenarios
Swapnil Bhosale, Imran A. Sheikh, Sri Harsha Dumpala, Sunil Kumar Kopparapu |
INTERSPEECH | 4 |
| 2019 | Front-End Feature Compensation and Denoising for Noise Robust Speech Emotion Recognition
Rupayan Chakraborty, Ashish Panda, Meghna Pandharipande, Sonal Joshi, Sunil Kumar Kopparapu |
INTERSPEECH | 5 |
| 2018 | FEMH Voice Data Challenge: Voice disorder Detection and Classification using Acoustic DescriptorsabstractThis paper describes the participation of TCS Research and Innovation, Mumbai in the FEMH voice data challenge. The goal of the FEMH voice data challenge is detection of pathological voice and classification into three different categories using voice samples. In this work, we use a mix of speech processing and machine learning techniques to not only automatically detect pathological speech but also classify into one of the three categories namely, Neoplasm, Phonotrauma and Vocal Palsy. Chitralekha Bhat, Sunil Kumar Kopparapu |
IEEE BigData | 2 |
| 2018 | A Novel Data Representation for Effective Learning in Class Imbalanced ScenariosabstractClass imbalance refers to the scenario where certain classes are highly under-represented compared to other classes in terms of the availability of training data. This situation hinders the applicability of conventional machine learning algorithms to most of the classification problems where class imbalance is prominent. Most existing methods addressing class imbalance either rely on sampling techniques or cost-sensitive learning methods; thus inheriting their shortcomings. In this paper, we introduce a novel approach that is different from sampling or cost-sensitive learning based techniques, to address the class imbalance problem, where two samples are simultaneously considered to train the classifier. Further, we propose a mechanism to use a single base classifier, instead of an ensemble of classifiers, to obtain the output label of the test sample using majority voting method. Experimental results on several benchmark datasets clearly indicate the usefulness of the proposed approach over the existing state-of-the-art techniques. Sri Harsha Dumpala, Rupayan Chakraborty, Sunil Kumar Kopparapu |
IJCAI | 3 |
| 2018 | Dysarthric Speech Recognition Using Time-delay Neural Network Based Denoising Autoencoder
Chitralekha Bhat, Biswajit Das, Bhavik Vachhani, Sunil Kumar Kopparapu |
INTERSPEECH | 4 |
| 2018 | Analysis of the Effect of Speech-Laugh on Speaker Recognition System
Sri Harsha Dumpala, Ashish Panda, Sunil Kumar Kopparapu |
INTERSPEECH | 3 |
| 2018 | Data Augmentation Using Healthy Speech for Dysarthric Speech Recognition
Bhavik Vachhani, Chitralekha Bhat, Sunil Kumar Kopparapu |
INTERSPEECH | 3 |
| 2018 | Sentiment Classification on Erroneous ASR Transcripts: A Multi View Learning ApproachabstractSentiment classification on spoken language transcriptions has received less attention. A practical system employing the spoken language modality will have to use a language transcription from an Automatic Speech Recognition (ASR) engine which is inherently prone to errors. The main interest of this paper lies in improvement of sentiment classification on erroneous ASR transcriptions. Our aim is to improve the representation of the ASR transcripts using the manual transcripts and other modalities, like audio and visual, that are available during training but not necessarily during test conditions. We adopt an approach based on Deep Canonical Correlation Analysis (DCCA) and propose two new extensions of DCCA to enhance the ASR view using multiple modalities. We present a detailed evaluation of the performance of our approach on datasets of opinion videos (CMU-MOSI and CMU-MOSEI) collected from Youtube. Sri Harsha Dumpala, Imran A. Sheikh, Rupayan Chakraborty, Sunil Kumar Kopparapu |
SLT | 4 |
| 2017 | Automatic assessment of dysarthria severity level using audio descriptorsabstractDysarthria is a motor speech impairment, often characterized by speech that is generally indiscernible by human listeners. Assessment of the severity level of dysarthria provides an understanding of the patient's progression in the underlying cause and is essential for planning therapy, as well as improving automatic dysarthric speech recognition. In this paper, we propose a non-linguistic manner of automatic assessment of severity levels using audio descriptors or a set of features traditionally used to define timbre of musical instruments and have been modified to suit this purpose. Multitapered spectral estimation based features were computed and used for classification, in addition to the audio descriptors for timbre. An Artificial Neural Network (ANN) was trained to classify speech into various severity levels within Universal Access dysarthric speech corpus and the TORGO database. An average classification accuracy of 96.44% and 98.7% was obtained for UA speech corpus and TORGO database respectively. Chitralekha Bhat, Bhavik Vachhani, Sunil Kumar Kopparapu |
ICASSP | 3 |
| 2017 | Improved speaker recognition system for stressed speech using deep neural networksabstractGood speaker recognition systems should identify the speaker irrespective of what is spoken, including non-speech sounds that are often produced during natural conversations. In this work, the inclusion of breath sounds in the training phase of the speaker recognition is analyzed using the popular Gaussian mixture model-universal background model (GMM-UBM) and deep neural network (DNN) based systems. It is shown that the DNN-based systems have a better learning capability to perform well even on unseen data compared to GMM-UBM-based systems. Specifically, enhancement in speaker recognition performance is obtained on unseen stressed speech data by training systems with both breath sounds and modal speech. Experimental results show that inclusion of breath sounds in training data reduces the equal error rate (EER) of the speaker recognition system on stressed speech by 40% to 50% in absolute terms. It is also shown that increasing the number of hidden layers help DNNs to improve the performance even on unseen data. Sri Harsha Dumpala, Sunil Kumar Kopparapu |
IJCNN | 2 |
| 2017 | MoPAReST - Mobile Phone Assisted Remote Speech Therapy Platform
Chitralekha Bhat, Anjali Kant, Bhavik Vachhani, Sarita Rautara, Ashok Kumar Sinha, Sunil Kumar Kopparapu |
INTERSPEECH | 6 |
| 2017 | Deep Autoencoder Based Speech Features for Improved Dysarthric Speech Recognition
Bhavik Vachhani, Chitralekha Bhat, Biswajit Das, Sunil Kumar Kopparapu |
INTERSPEECH | 4 |
| 2017 | A spoof resistant multibiometric system based on the physiological and behavioral characteristics of fingerprint
Ishan Bhardwaj, Narendra D. Londhe, Sunil Kumar Kopparapu |
Pattern Recognit. | 3 |
| 2016 | Spontaneous speech emotion recognition using prior knowledgeabstractAutomatic and spontaneous speech emotion recognition is an important part of a human-computer interactive system. However, emotion identification in spontaneous speech is difficult because most often the emotion expressed by the speaker are not necessarily as prominent as in acted speech. In this paper, we propose a spontaneous speech emotion recognition framework that makes use of the associated knowledge. The framework is motivated by the observation that there is significant disagreement amongst human annotators when they annotate spontaneous speech; the disagreement largely reduces when they are provided with additional knowledge related to the conversation. The proposed framework makes use of the contexts (derived from linguistic contents) and the knowledge regarding the time lapse of the spoken utterances in the context of an audio call to reliably recognize the current emotion of the speaker in spontaneous audio conversations. Our experimental results demonstrate that there is a significant improvement in the performance of spontaneous speech emotion recognition using the proposed framework. Rupayan Chakraborty, Meghna Pandharipande, Sunil Kumar Kopparapu |
ICPR | 3 |
| 2016 | Repairing General-Purpose ASR Output to Improve Accuracy of Spoken Sentences in Specific Domains Using Artificial Development Approach
C. Anantaram, Sunil Kumar Kopparapu, Chiragkumar Patel, Aditya Mittal |
IJCAI | 2 |
| 2016 | Recognition of Dysarthric Speech Using Voice Parameters for Speaker Adaptation and Multi-Taper Spectral Estimation
Chitralekha Bhat, Bhavik Vachhani, Sunil Kumar Kopparapu |
INTERSPEECH | 3 |
| 2016 | Knowledge-based Framework for Intelligent Emotion Recognition in Spontaneous SpeechabstractAutomatic speech emotion recognition plays an important role in intelligent human computer interaction. Identifying emotion in natural, day to day, spontaneous conversational speech is difficult because most often the emotion expressed by the speaker are not necessarily as prominent as in acted speech. In this paper, we propose a novel spontaneous speech emotion recognition framework that makes use of the available knowledge. The framework is motivated by the observation that there is significant disagreement amongst human annotators when they annotate spontaneous speech; the disagreement largely reduces when they are provided with additional knowledge related to the conversation. The proposed framework makes use of the contexts (derived from linguistic contents) and the knowledge regarding the time lapse of the spoken utterances in the context of an audio call to reliably recognize the current emotion of the speaker in spontaneous audio conversations. Our experimental results demonstrate that there is a significant improvement in the performance of spontaneous speech emotion recognition using the proposed framework. Rupayan Chakraborty, Meghna Pandharipande, Sunil Kumar Kopparapu |
KES | 3 |
| 2016 | Mining Call Center Conversations exhibiting Similar Affective States
Rupayan Chakraborty, Meghna Pandharipande, Sunil Kumar Kopparapu |
PACLIC | 3 |
| 2016 | Validating "Is ECC-ANN combination equivalent to DNN?" for speech emotion recognitionabstractUse of the error correcting codes (ECC) in a multiclass audio emotion recognition problem is proposed to improve the emotion recognition accuracy. We visualize the emotion recognition system as a noisy communication channel, thus motivating the use of ECC. We assume the emotion recognition process consists of an audio feature extractor followed by an artificial neural network (ANN) for emotion classification. In our formulation, the noise in the communication channel is a result of insufficiently learnt ANN classifier which results in an erroneous emotion classification. We first show that the ECC-ANN combination performs better than the ANN classifier, justifying the use of ECC-ANN combination. We further make the conjecture that ECC in ECC-ANN combination can be visualized as a part of Deep Neural Network (DNN) where the intelligence is under control. We show through rigorous experimentation, on Emo-DB database, that the use of ECC-ANN combination is equivalent to the DNN; in terms of the improved recognition accuracies over an ANN. Our experimental results show that both ECC-ANN and DNN give a minimum absolute improvement of around 13.75%. Rupayan Chakraborty, Sunil Kumar Kopparapu |
SMC | 2 |
| 2015 | Viseme comparison based on phonetic cues for varying speech accents
Chitralekha Bhat, Sunil Kumar Kopparapu |
INTERSPEECH | 2 |
| 2013 | Technique for automatic sentence level alignment of long speech and transcripts
Imran Ahmed 0007, Sunil Kumar Kopparapu |
INTERSPEECH | 2 |
| 2006 | Lighting design for machine vision application
Sunil Kumar Kopparapu |
Image Vis. Comput. | 1 |
| 2001 | The Effect of Noise on Camera Calibration Parameters
Sunil Kumar Kopparapu, Peter I. Corke |
Graph. Model. | 1 |
| 2000 | Behaviour of image degradation model in multiresolution
Sunil Kumar Kopparapu, Uday B. Desai, Peter I. Corke |
Signal Process. | 1 |
| 1999 | The Effect of Measurement Noise on Intrinsic Camera Calibration ParametersabstractThe camera calibration matrix captures the transformation of a 3D point onto a 2D image plane as occurs in the camera being used for imaging. Given the world coordinates of a number of precisely placed points in a 3D space, camera calibration requires the measurement of the 2D projection of those scene points on the image plane. While the coordinates of the points in 3D space can be measured precisely, the image coordinates that are determined from the digital image are not precise mainly because of the measurement noise. In this paper, we derive analytical relationships between the errors in the intrinsic camera parameters (ICPs) due to error in measurements. We show that the errors in the ICPs are Gaussian distributed when the measurement noise is Gaussian. We verify these results experimentally. Sunil Kumar Kopparapu, Peter I. Corke |
ICRA | 1 |