VLDB 2026 Research / reviewers in the wild / expert
Sri Harsha Dumpala
dblp:148/9851
· DBLP profile ↗
18ranked-venue papers
15as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 13 first-author · 5 since 2021Artificial intelligence and machine learning · 14 · 11 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Test-Time Training for Speech-based Depression Detection
Sri Harsha Dumpala, Chandramouli Shama Sastry, Rudolf Uher, Sageev Oore |
INTERSPEECH | 1 |
| 2024 | Self-Supervised Embeddings for Detecting Individual Symptoms of Depression
Sri Harsha Dumpala, Katerina Dikaios, Abraham Nunes, Frank Rudzicz, Rudolf Uher, Sageev Oore |
INTERSPEECH | 1 |
| 2024 | XANE: eXplainable Acoustic Neural Embeddings
Sri Harsha Dumpala, Dushyant Sharma, Chandramouli Shama Sastry, Stanislav Yu. Kruchinin, James Fosburgh, Patrick A. Naylor |
INTERSPEECH | 1 |
| 2024 | SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical AlterationsabstractDespite their remarkable successes, state-of-the-art large language models (LLMs), including vision-and-language models (VLMs) and unimodal language models (ULMs), fail to understand precise semantics. For example, semantically equivalent sentences expressed using different lexical compositions elicit diverging representations. The degree of this divergence and its impact on encoded semantics is not very well understood. In this paper, we introduce the SUGARCREPE++ dataset to analyze the sensitivity of VLMs and ULMs to lexical and semantic alterations. Each sample in SUGARCREPE++ dataset consists of an image and a corresponding triplet of captions: a pair of semantically equivalent but lexically different positive captions and one hard negative caption. This poses a 3-way semantic (in)equivalence problem to the language models. We comprehensively evaluate VLMs and ULMs that differ in architecture, pre-training objectives and datasets to benchmark the performance of SUGARCREPE++ dataset. Experimental results highlight the difficulties of VLMs in distinguishing between lexical and semantic variations, particularly to object attributes and spatial relations. Although VLMs with larger pre-training datasets, model sizes, and multiple pre-training objectives achieve better performance on SUGARCREPE++, there is a significant opportunity for improvement. We demonstrate that models excelling on compositionality datasets may not perform equally well on SUGARCREPE++. This indicates that compositionality alone might not be sufficient to fully understand semantic and lexical alterations. Given the importance of the property that the SUGARCREPE++ dataset targets, it serves as a new challenge to the vision-and-language community. Data and code is available at https://github.com/Sri-Harsha/scpp. Sri Harsha Dumpala, Aman Jaiswal, Chandramouli Shama Sastry, Evangelos E. Milios, Sageev Oore, Hassan Sajjad 0001 |
NeurIPS | 1 |
| 2024 | DiffAug: A Diffuse-and-Denoise Augmentation for Training Robust ClassifiersabstractWe introduce DiffAug, a simple and efficient diffusion-based augmentation technique to train image classifiers for the crucial yet challenging goal of improved classifier robustness. Applying DiffAug to a given example consists of one forward-diffusion step followed by one reverse-diffusion step. Using both ResNet-50 and Vision Transformer architectures, we comprehensively evaluate classifiers trained with DiffAug and demonstrate the surprising effectiveness of single-step reverse diffusion in improving robustness to covariate shifts, certified adversarial accuracy and out of distribution detection. When we combine DiffAug with other augmentations such as AugMix and DeepAugment we demonstrate further improved robustness. Finally, building on this approach, we also improve classifier-guided diffusion wherein we observe improvements in: (i) classifier-generalization, (ii) gradient quality (i.e., improved perceptual alignment) and (iii) image generation performance. We thus introduce a computationally efficient technique for training with improved robustness that does not require any additional data, and effectively complements existing augmentation approaches. Chandramouli Shama Sastry, Sri Harsha Dumpala, Sageev Oore |
NeurIPS | 2 |
| 2022 | On Combining Global and Localized Self-Supervised Models of Speech
Sri Harsha Dumpala, Chandramouli Shama Sastry, Rudolf Uher, Sageev Oore |
INTERSPEECH | 1 |
| 2021 | Estimating Severity of Depression From Acoustic Features and Embeddings of Natural SpeechabstractMajor depressive disorder, referred to as depression, is a leading cause of disability, absence from work, and premature death. Automatic assessment of depression from speech is a critical step towards improving diagnosis and treatment of depression. Previous works on depression assessment from speech considered various acoustic features extracted from speech to estimate depression severity. But performance of these approaches is not at clinical standards, and thus requires further improvement. In this work, we examine two novel approaches for improving depression severity estimation from short audio recordings of speech. Specifically, in audio recordings of a narrative by individuals diagnosed with major depressive disorder, we analyze spectral-based and excitation source-based features extracted from speech, and significance of sentiment and emotion classification in estimation of depression severity. Initial results indicate synchrony between depression scores and the sentiment and emotion labels. We propose the use of sentiment and emotion based embeddings obtained using machine learning techniques in estimation of depression severity. We also propose use of multi-task training to better estimate depression severity. We show that the proposed approaches provide additive improvements in the estimation of depression severity. Sri Harsha Dumpala, Sheri Rempel, Katerina Dikaios, Mehri Sajjadian, Rudolf Uher, Sageev Oore |
ICASSP | 1 |
| 2019 | Improving ASR Robustness to Perturbed Speech Using Cycle-consistent Generative Adversarial NetworksabstractNaturally introduced perturbations in audio signal, caused by emotional and physical states of the speaker, can significantly degrade the performance of Automatic Speech Recognition (ASR) systems. In this paper, we propose a front-end based on Cycle-Consistent Generative Adversarial Network (CycleGAN) which transforms naturally perturbed speech into normal speech, and hence improves the robustness of an ASR system. The CycleGAN model is trained on non-parallel examples of perturbed and normal speech. Experiments on spontaneous laughter-speech and creaky voice datasets show that the performance of four different ASR systems improve by using speech obtained from CycleGAN based front-end, as compared to directly using the original perturbed speech. Visualization of the features of the laughter perturbed speech and those generated by the proposed front-end further demonstrates the effectiveness of our approach. Sri Harsha Dumpala, Imran A. Sheikh, Rupayan Chakraborty, Sunil Kumar Kopparapu |
ICASSP | 1 |
| 2019 | End-to-End Spoken Language Understanding: Bootstrapping in Low Resource Scenarios
Swapnil Bhosale, Imran A. Sheikh, Sri Harsha Dumpala, Sunil Kumar Kopparapu |
INTERSPEECH | 3 |
| 2019 | Excitation Source and Vocal Tract System Based Acoustic Features for Detection of Nasals in Continuous Speech
Bhanu Teja Nellore, Sri Harsha Dumpala, Karan Nathwani, Suryakanth V. Gangashetty |
INTERSPEECH | 2 |
| 2018 | A Novel Data Representation for Effective Learning in Class Imbalanced ScenariosabstractClass imbalance refers to the scenario where certain classes are highly under-represented compared to other classes in terms of the availability of training data. This situation hinders the applicability of conventional machine learning algorithms to most of the classification problems where class imbalance is prominent. Most existing methods addressing class imbalance either rely on sampling techniques or cost-sensitive learning methods; thus inheriting their shortcomings. In this paper, we introduce a novel approach that is different from sampling or cost-sensitive learning based techniques, to address the class imbalance problem, where two samples are simultaneously considered to train the classifier. Further, we propose a mechanism to use a single base classifier, instead of an ensemble of classifiers, to obtain the output label of the test sample using majority voting method. Experimental results on several benchmark datasets clearly indicate the usefulness of the proposed approach over the existing state-of-the-art techniques. Sri Harsha Dumpala, Rupayan Chakraborty, Sunil Kumar Kopparapu |
IJCAI | 1 |
| 2018 | Analysis of the Effect of Speech-Laugh on Speaker Recognition System
Sri Harsha Dumpala, Ashish Panda, Sunil Kumar Kopparapu |
INTERSPEECH | 1 |
| 2018 | Sentiment Classification on Erroneous ASR Transcripts: A Multi View Learning ApproachabstractSentiment classification on spoken language transcriptions has received less attention. A practical system employing the spoken language modality will have to use a language transcription from an Automatic Speech Recognition (ASR) engine which is inherently prone to errors. The main interest of this paper lies in improvement of sentiment classification on erroneous ASR transcriptions. Our aim is to improve the representation of the ASR transcripts using the manual transcripts and other modalities, like audio and visual, that are available during training but not necessarily during test conditions. We adopt an approach based on Deep Canonical Correlation Analysis (DCCA) and propose two new extensions of DCCA to enhance the ASR view using multiple modalities. We present a detailed evaluation of the performance of our approach on datasets of opinion videos (CMU-MOSI and CMU-MOSEI) collected from Youtube. Sri Harsha Dumpala, Imran A. Sheikh, Rupayan Chakraborty, Sunil Kumar Kopparapu |
SLT | 1 |
| 2017 | Improved speaker recognition system for stressed speech using deep neural networksabstractGood speaker recognition systems should identify the speaker irrespective of what is spoken, including non-speech sounds that are often produced during natural conversations. In this work, the inclusion of breath sounds in the training phase of the speaker recognition is analyzed using the popular Gaussian mixture model-universal background model (GMM-UBM) and deep neural network (DNN) based systems. It is shown that the DNN-based systems have a better learning capability to perform well even on unseen data compared to GMM-UBM-based systems. Specifically, enhancement in speaker recognition performance is obtained on unseen stressed speech data by training systems with both breath sounds and modal speech. Experimental results show that inclusion of breath sounds in training data reduces the equal error rate (EER) of the speaker recognition system on stressed speech by 40% to 50% in absolute terms. It is also shown that increasing the number of hidden layers help DNNs to improve the performance even on unseen data. Sri Harsha Dumpala, Sunil Kumar Kopparapu |
IJCNN | 1 |
| 2016 | Use of Vowels in Discriminating Speech-Laugh from Laughter and Neutral Speech
Sri Harsha Dumpala, P. Gangamohan, Suryakanth V. Gangashetty, Bayya Yegnanarayana |
INTERSPEECH | 1 |
| 2016 | Robust Vowel Landmark Detection Using Epoch-Based Features
Sri Harsha Dumpala, Bhanu Teja Nellore, Raghu Ram Nevali, Suryakanth V. Gangashetty, Bayya Yegnanarayana |
INTERSPEECH | 1 |
| 2015 | Robust features for sonorant segmentation in continuous speech
Sri Harsha Dumpala, Bhanu Teja Nellore, Raghu Ram Nevali, Suryakanth V. Gangashetty, Bayya Yegnanarayana |
INTERSPEECH | 1 |
| 2014 | Analysis of laughter and speech-laugh signals using excitation source informationabstractSpeech-laugh is a speech-synchronous form of laughter that often occurs in natural conversation. However, there are deviations in features of speech-laugh when compared with laughter and neutral speech individually. The objective of this study is to analyse the excitation source features to capture the deviations between laughter and speech-laughs in voiced regions. The features used in this analysis are based on instantaneous fundamental frequency and strength of excitation (β) at epochs. Modified zero frequency filtering (ZFF) method is used to extract the features. Kullback-Leibler (KL) distances obtained show that there are deviations in excitation source features which can be exploited to develop a method to discriminate speech-laughs from laughter. Experimental results show that features used are robust and speaker independent in discriminating speech-laughs from laughter. Results showing deviations of laughter and speech-laughs from neutral speech were also presented. Sri Harsha Dumpala, Karthik Venkat Sridaran, Suryakanth V. Gangashetty, Bayya Yegnanarayana |
ICASSP | 1 |