VLDB 2026 Research / reviewers in the wild / expert
Rupayan Chakraborty
dblp:09/8618
· DBLP profile ↗
22ranked-venue papers
10as first author
6since 2021 · last 2024
0000-0002-3566-0784ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 11 · 7 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Identifying Indian Cattle Behaviour Using Acoustic Biomarkers
Ruturaj Patil, Hemavathy B, Sanat Sarangi, Dinesh Kumar Singh, Rupayan Chakraborty, Sanket Junagade, Srinivasu Pappula |
ICPRAM | 5 |
| 2024 | Interpreting Cattle Behaviour and Specific Behaviour Intensities with Acoustic Biomarkers
Ruturaj Patil, Hemavathy B, Sanat Sarangi, Dinesh Kumar Singh, Rupayan Chakraborty, Sanket Junagade, Srinivasu Pappula |
ICPRAM | 5 |
| 2024 | Reinforcement Learning based Data Augmentation for Noise Robust Speech Emotion Recognition
Sumit Ranjan, Rupayan Chakraborty, Sunil Kumar Kopparapu |
INTERSPEECH | 2 |
| 2023 | A Novel Metric For Evaluating Audio Caption SimilarityabstractAutomatic Audio Captioning (AAC) refers to the task of describing an audio sample in a natural language (NL) text. Unlike NL text generation tasks, which rely on lexical semantic metrics like BLEU for evaluation, the AAC evaluation metric requires acoustic semantics to map NL text corresponding to similar sounds in addition to lexical semantics. In this paper, we propose a novel metric based on Text-to-Audio Grounding (TAG), to incorporate acoustic semantics. Experiments demonstrate our evaluation metric to perform better compared to existing metrics used in NL text and image captioning literature for AAC. Swapnil Bhosale, Rupayan Chakraborty, Sunil Kumar Kopparapu |
ICASSP | 2 |
| 2021 | Deep Lung Auscultation Using Acoustic Biomarkers for Abnormal Respiratory Sound Event DetectionabstractLung Auscultation is a non-invasive process of distinguishing normal respiratory sounds from abnormal ones by analyzing the airflow along the respiratory tract. With developments in the Deep Learning (DL) techniques and wider access to anonymized medical data, automatic detection of specific sounds such as crackles and wheezes have been gaining popularity. In this paper, we propose to use two sets of diversified acoustic biomarkers extracted using Discrete Wavelet Transform (DWT) and deep encoded features from the intermediate layer of a pre-trained Audio Event Detection (AED) model trained using sounds from daily activities. First set of biomarkers highlight the time frequency localization characteristics obtained from DWT coefficients. However, the second set of deep encoded biomarkers captures a generalized reliable representation, and thus indemnifies the scarcity of training samples and the class imbalance in dataset. The model trained using these features achieves a 15.05% increase in terms of the specificity over the baseline model that uses spectrogram features. Moreover, ensemble of DWT features and deep encoded feature based models show absolute improvements of 8.32%, 6.66% and 7.40% in terms of sensitivity, specificity and ICBHI-score, respectively, and clearly outperforms the state-of-the-art with a significant margin. Upasana Tiwari, Swapnil Bhosale, Rupayan Chakraborty, Sunil Kumar Kopparapu |
ICASSP | 3 |
| 2021 | Contrastive Learning of Cough Descriptors for Automatic COVID-19 Preliminary DiagnosisabstractCough sounds as a descriptor have been used for detecting various respiratory ailments based on its intensity, duration of intermediate phase between two cough sounds, repetitions, dryness etc.However, COVID-19 diagnosis using only cough sounds is challenging because of cough being a common symptom among many non COVID-19 health diseases and inherent data imbalance within the available datasets.As one of the approach in this direction, we explore the robustness of multi-domain representation by performing the early fusion over a wide set of temporal, spectral and tempo-spectral handcrafted features, followed by training a Support Vector Machine (SVM) classifier.In our second approach, using a contrastive loss function we learn a latent space from Mel Filter Cepstral Coefficients (MFCCs) where representations belonging to samples having similar cough characteristics are closer.This helps learn representations for the highly varied COVIDnegative class (healthy and symptomatic COVID-negative), by learning multiple smaller clusters.Using only the DiCOVA data, multi-domain features yields an absolute improvement of 0.74% and 1.07%, whereas our second approach shows an improvement of 2.09% and 3.98%, over the blind test and validation set, respectively, when compared with challenge baseline. Swapnil Bhosale, Upasana Tiwari, Rupayan Chakraborty, Sunil Kumar Kopparapu |
Interspeech | 3 |
| 2020 | Deep Encoded Linguistic and Acoustic Cues for Attention Based End to End Speech Emotion RecognitionabstractAn End-to-End model with convolutional layers and multi-head self attention mechanism is proposed for Speech Emotion Recognition (SER) task. As inputs, we propose to use both the deep encoded linguistic features that carry the language related context of emotion and the audio spectrogram that are representatives of acoustic cues. To facilitate the deep linguistic feature representation, we use outputs from the intermediate layers of a pre-trained Automatic Speech Recognition (ASR) model, where the layer is selected empirically. The influence of both acoustic and linguistic features, both separately and in combination, for emotion recognition in different scenarios (scripted and spontaneous recording of emotional speech samples) have been studied. Extensive experiments on the standard IEMOCAP database are conducted to investigate the efficacy of our proposed approach. To address the class imbalance, we carried out down sampling and ensembling, which further improved the SER accuracy. Overall, we observe that the acoustic features perform best for improvised recordings which is due to the spontaneity in speech with less linguistic correlation. But the linguistic features are found to be effective for the scripted as well as for the combined (scripted and improvised recordings together) scenario that reflects more linguistic information in spoken utterances. Swapnil Bhosale, Rupayan Chakraborty, Sunil Kumar Kopparapu |
ICASSP | 2 |
| 2020 | Multi-Conditioning and Data Augmentation Using Generative Noise Model for Speech Emotion Recognition in Noisy ConditionsabstractDegradation due to additive noise is a significant road block in the real-life deployment of Speech Emotion Recognition (SER) systems. Most of the previous work in this field dealt with the noise degradation either at the signal or at the feature level. In this paper, to address the robustness aspect of the SER in additive noise scenarios, we propose multi-conditioning and data augmentation using an utterance level parametric Generative noise model. The Generative noise model is designed to generate noise types which can span the entire noise space in the mel-filterbank energy domain. This characteristic of the model renders the system robust against unseen noise conditions. The generated noise types can be used to create multiconditioned data for training the SER systems. Multi-conditioning approach can also be used to increase the training data by many folds where such data is limited. We report the performance of the proposed method on two datasets, namely EmoDB and IEMOCAP. We also explore multi-conditioning and data augmentation using noise samples from NOISEX-92 database. Upasana Tiwari, Meet H. Soni, Rupayan Chakraborty, Ashish Panda, Sunil Kumar Kopparapu |
ICASSP | 3 |
| 2020 | A Novel Adaptive Minority Oversampling Technique for Improved Classification in Data Imbalanced ScenariosabstractImbalance in the proportion of training samples belonging to different classes often poses performance degradation of conventional classifiers. This is primarily due to the tendency of the classifier to be biased towards the majority classes in the imbalanced dataset. In this paper, we propose a novel three step technique to address imbalanced data. As a first step we significantly oversample the minority class distribution by employing the traditional Synthetic Minority Oversampling Technique (SMOTE) algorithm using the neighborhood of the minority class samples and in the next step we partition the generated samples using a Gaussian-Mixture Model based clustering algorithm. In the final step synthetic data samples are chosen based on the weight associated with the cluster, the weight itself being determined by the distribution of the majority class samples. Extensive experiments on several standard datasets from diverse domains show the usefulness of the proposed technique in comparison with the original SMOTE and its state-of-the-art variants algorithms. Ayush Tripathi, Rupayan Chakraborty, Sunil Kumar Kopparapu |
ICPR | 2 |
| 2019 | Improving ASR Robustness to Perturbed Speech Using Cycle-consistent Generative Adversarial NetworksabstractNaturally introduced perturbations in audio signal, caused by emotional and physical states of the speaker, can significantly degrade the performance of Automatic Speech Recognition (ASR) systems. In this paper, we propose a front-end based on Cycle-Consistent Generative Adversarial Network (CycleGAN) which transforms naturally perturbed speech into normal speech, and hence improves the robustness of an ASR system. The CycleGAN model is trained on non-parallel examples of perturbed and normal speech. Experiments on spontaneous laughter-speech and creaky voice datasets show that the performance of four different ASR systems improve by using speech obtained from CycleGAN based front-end, as compared to directly using the original perturbed speech. Visualization of the features of the laughter perturbed speech and those generated by the proposed front-end further demonstrates the effectiveness of our approach. Sri Harsha Dumpala, Imran A. Sheikh, Rupayan Chakraborty, Sunil Kumar Kopparapu |
ICASSP | 3 |
| 2019 | Front-End Feature Compensation and Denoising for Noise Robust Speech Emotion Recognition
Rupayan Chakraborty, Ashish Panda, Meghna Pandharipande, Sonal Joshi, Sunil Kumar Kopparapu |
INTERSPEECH | 1 |
| 2018 | A Novel Data Representation for Effective Learning in Class Imbalanced ScenariosabstractClass imbalance refers to the scenario where certain classes are highly under-represented compared to other classes in terms of the availability of training data. This situation hinders the applicability of conventional machine learning algorithms to most of the classification problems where class imbalance is prominent. Most existing methods addressing class imbalance either rely on sampling techniques or cost-sensitive learning methods; thus inheriting their shortcomings. In this paper, we introduce a novel approach that is different from sampling or cost-sensitive learning based techniques, to address the class imbalance problem, where two samples are simultaneously considered to train the classifier. Further, we propose a mechanism to use a single base classifier, instead of an ensemble of classifiers, to obtain the output label of the test sample using majority voting method. Experimental results on several benchmark datasets clearly indicate the usefulness of the proposed approach over the existing state-of-the-art techniques. Sri Harsha Dumpala, Rupayan Chakraborty, Sunil Kumar Kopparapu |
IJCAI | 2 |
| 2018 | Sentiment Classification on Erroneous ASR Transcripts: A Multi View Learning ApproachabstractSentiment classification on spoken language transcriptions has received less attention. A practical system employing the spoken language modality will have to use a language transcription from an Automatic Speech Recognition (ASR) engine which is inherently prone to errors. The main interest of this paper lies in improvement of sentiment classification on erroneous ASR transcriptions. Our aim is to improve the representation of the ASR transcripts using the manual transcripts and other modalities, like audio and visual, that are available during training but not necessarily during test conditions. We adopt an approach based on Deep Canonical Correlation Analysis (DCCA) and propose two new extensions of DCCA to enhance the ASR view using multiple modalities. We present a detailed evaluation of the performance of our approach on datasets of opinion videos (CMU-MOSI and CMU-MOSEI) collected from Youtube. Sri Harsha Dumpala, Imran A. Sheikh, Rupayan Chakraborty, Sunil Kumar Kopparapu |
SLT | 3 |
| 2016 | Spontaneous speech emotion recognition using prior knowledgeabstractAutomatic and spontaneous speech emotion recognition is an important part of a human-computer interactive system. However, emotion identification in spontaneous speech is difficult because most often the emotion expressed by the speaker are not necessarily as prominent as in acted speech. In this paper, we propose a spontaneous speech emotion recognition framework that makes use of the associated knowledge. The framework is motivated by the observation that there is significant disagreement amongst human annotators when they annotate spontaneous speech; the disagreement largely reduces when they are provided with additional knowledge related to the conversation. The proposed framework makes use of the contexts (derived from linguistic contents) and the knowledge regarding the time lapse of the spoken utterances in the context of an audio call to reliably recognize the current emotion of the speaker in spontaneous audio conversations. Our experimental results demonstrate that there is a significant improvement in the performance of spontaneous speech emotion recognition using the proposed framework. Rupayan Chakraborty, Meghna Pandharipande, Sunil Kumar Kopparapu |
ICPR | 1 |
| 2016 | Knowledge-based Framework for Intelligent Emotion Recognition in Spontaneous SpeechabstractAutomatic speech emotion recognition plays an important role in intelligent human computer interaction. Identifying emotion in natural, day to day, spontaneous conversational speech is difficult because most often the emotion expressed by the speaker are not necessarily as prominent as in acted speech. In this paper, we propose a novel spontaneous speech emotion recognition framework that makes use of the available knowledge. The framework is motivated by the observation that there is significant disagreement amongst human annotators when they annotate spontaneous speech; the disagreement largely reduces when they are provided with additional knowledge related to the conversation. The proposed framework makes use of the contexts (derived from linguistic contents) and the knowledge regarding the time lapse of the spoken utterances in the context of an audio call to reliably recognize the current emotion of the speaker in spontaneous audio conversations. Our experimental results demonstrate that there is a significant improvement in the performance of spontaneous speech emotion recognition using the proposed framework. Rupayan Chakraborty, Meghna Pandharipande, Sunil Kumar Kopparapu |
KES | 1 |
| 2016 | Mining Call Center Conversations exhibiting Similar Affective States
Rupayan Chakraborty, Meghna Pandharipande, Sunil Kumar Kopparapu |
PACLIC | 1 |
| 2016 | Validating "Is ECC-ANN combination equivalent to DNN?" for speech emotion recognitionabstractUse of the error correcting codes (ECC) in a multiclass audio emotion recognition problem is proposed to improve the emotion recognition accuracy. We visualize the emotion recognition system as a noisy communication channel, thus motivating the use of ECC. We assume the emotion recognition process consists of an audio feature extractor followed by an artificial neural network (ANN) for emotion classification. In our formulation, the noise in the communication channel is a result of insufficiently learnt ANN classifier which results in an erroneous emotion classification. We first show that the ECC-ANN combination performs better than the ANN classifier, justifying the use of ECC-ANN combination. We further make the conjecture that ECC in ECC-ANN combination can be visualized as a part of Deep Neural Network (DNN) where the intelligence is under control. We show through rigorous experimentation, on Emo-DB database, that the use of ECC-ANN combination is equivalent to the DNN; in terms of the improved recognition accuracies over an ANN. Our experimental results show that both ECC-ANN and DNN give a minimum absolute improvement of around 13.75%. Rupayan Chakraborty, Sunil Kumar Kopparapu |
SMC | 1 |
| 2014 | Sound-model-based acoustic source localization using distributed microphone arraysabstractAcoustic source localization and sound recognition are common acoustic scene analysis tasks that are usually considered separately. In this paper, a new source localization technique is proposed that works jointly with an acoustic event detection system. Given the identities and the end-points of simultaneous sounds, the proposed technique uses the statistical models of those sounds to compute a likelihood score for each model and for each signal at the output of a set of null-steering beamformers per microphone array. Those scores are subsequently combined to find the MAP-optimal event source positions in the room. Experimental work is reported for a scenario consisting of meeting-room acoustic events, either isolated or overlapped with speech. From the localization results, which are compared with those from the SRP-PHAT technique, it seems that the proposed model-based approach can be an alternative to current techniques for event-based localization. Rupayan Chakraborty, Climent Nadeu |
ICASSP | 1 |
| 2013 | Real-time multi-microphone recognition of simultaneous sounds in a room environmentabstractTime overlapping of acoustic signals, which so often occurs in real life, is a challenge for current state-of-the-art sound recognition systems. In this work, we propose an approach for detecting, identifying and positioning a set of simultaneous acoustic events in a room environment, using multiple arbitrarily-located microphone arrays, and working in real time. Assuming a set of estimated acoustic source positions, the use of a frequency invariant null-steering beamformer for each position and each array yields a set of signals which show different balances among the various acoustic sources. For each signal, a model-based likelihood computation is carried out to obtain a matrix of likelihood scores. Then a MAP criterion is used to jointly detect the event classes and assign each of them to a given source position. Experimental results with two sources, one of which is speech, and two three-microphone linear arrays are reported, and a comparison with alternatives approaches is carried out. Rupayan Chakraborty, Climent Nadeu |
ICASSP | 1 |
| 2013 | Joint recognition and direction-of-arrival estimation of simultaneous meeting-room acoustic eventsabstractAcoustic scene analysis usually requires several sub-systems working in parallel for carrying out the various required functionalities. Focusing to a more integrated approach, in this paper we present an attempt to jointly recognize and localize several simultaneous acoustic events that take place in a meeting room environment, by developing a computationally efficient technique that employs multiple arbitrarily-located small microphone arrays. Assuming a set of simultaneous sounds, for each array a matrix is computed whose elements are likelihoods along the set of classes and a set of discretized directions of arrival. MAP estimation is used to decide about both the recognized events and the estimated directions. Experimental results with two sources, one of which is speech, and two three-microphone linear arrays are reported. The recognition results compare favorably with the ones obtained by assuming that the positions are known. Rupayan Chakraborty, Climent Nadeu |
INTERSPEECH | 1 |
| 2012 | Detection and Positioning of Overlapped Sounds in a Room Environment
Rupayan Chakraborty, Climent Nadeu, Taras Butko |
INTERSPEECH | 1 |
| 2010 | Role of Synthetically Generated Samples on Speech Recognition in a Resource-Scarce LanguageabstractSpeech recognition systems that make use of statistical classifiers require a large number of training samples. However, collection of real samples has always been a difficult problem due to the involvement of substantial amount of human intervention and cost. Considering this problem, this paper presents a novel method for generating synthetic samples from a handful of real samples and investigates the role of these samples in designing a speech recognition system. Speaker dependent limited vocabulary isolated word recognition in an Indian language (i.e. Bengali) has been taken a reference to demonstrate the potential of the proposed framework. The role of synthetic samples is demonstrated by showing a significant improvement in recognition accuracy. A maximum improvement of 10% is achieved using the proposed approach. Rupayan Chakraborty, Utpal Garain |
ICPR | 1 |