VLDB 2026 Research / reviewers in the wild / expert
Hira Dhamyal
dblp:251/9095
· DBLP profile ↗
10ranked-venue papers
6as first author
8since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?abstractReference summaries for abstractive speech summarization require human annotation, which can be performed by listening to an audio recording or by reading textual transcripts of the recording.In this paper, we examine whether summaries based on annotators listening to the recordings differ from those based on annotators reading transcripts.Using existing intrinsic evaluation based on human evaluation, automatic metrics, LLM-based evaluation, and a retrieval-based reference-free method.We find that summaries are indeed different based on the source modality, and that speechbased summaries are more factually consistent and information-selective than transcript-based summaries.Meanwhile, transcript-based summaries are impacted by recognition errors in the source, and expert-written summaries are more informative and reliable.We make all the collected data and analysis code public 1 to facilitate the reproduction of our work and advance research in this area. Roshan S. Sharma, Suwon Shon, Mark Lindsey, Hira Dhamyal, Bhiksha Raj |
ACL (1) | 4 |
| 2024 | Prompting Audios Using Acoustic Properties for Emotion RepresentationabstractEmotions lie on a continuum, but current models treat emotions as a finite valued discrete variable. This representation does not capture the diversity in the expression of emotion. To better represent emotions we propose the use of natural language descriptions (or prompts). In this work, we address the challenge of automatically generating these prompts and training a model to better learn emotion representations from audio and prompt pairs. We use acoustic properties that are correlated to emotion like pitch, intensity, speech rate, and articulation rate to automatically generate prompts i.e. ‘acoustic prompts’. We use a contrastive learning objective to map speech to their respective acoustic prompts. We evaluate our model on Emotion Audio Retrieval and Speech Emotion Recognition. Our results show that the acoustic prompts significantly improve the model’s performance in EAR, in various Precision@K metrics. In SER, we observe a 3.8% relative accuracy improvement on the Ravdess dataset. Hira Dhamyal, Benjamin Elizalde, Soham Deshmukh, Huaming Wang, Bhiksha Raj, Rita Singh |
ICASSP | 1 |
| 2024 | SELM: Enhancing Speech Emotion Recognition for Out-of-Domain Scenarios
Hazim T. Bukhari, Soham Deshmukh, Hira Dhamyal, Bhiksha Raj, Rita Singh |
INTERSPEECH | 3 |
| 2024 | PDAF: A Phonetic Debiasing Attention Framework For Speaker VerificationabstractSpeaker verification systems are crucial for authenticating identity through voice. Traditionally, these systems focus on comparing feature vectors, overlooking the speech’s content. However, this paper challenges this by highlighting the importance of phonetic dominance, a measure of the frequency or duration of phonemes, as a crucial cue in speaker verification. A novel Phoneme-Debiasing Attention Framework (PDAF) is introduced, integrating with existing attention frameworks to mitigate biases caused by phonetic dominance. PDAF adjusts the weighting for each phoneme and influences feature extraction, allowing for a more nuanced analysis of speech. This approach paves the way for more accurate and reliable identity authentication through voice. Furthermore, by employing various weighting strategies, we evaluate the influence of phonetic features on the efficacy of the speaker verification system. Massa Baali, Abdulhamid Aldoobi, Hira Dhamyal, Rita Singh, Bhiksha Raj |
SLT | 3 |
| 2022 | Positional Encoding for Capturing Modality Specific Cadence for Emotion Detection
Hira Dhamyal, Bhiksha Raj, Rita Singh |
INTERSPEECH | 1 |
| 2021 | Using Self Attention DNNs to Discover Phonemic Features for Audio Deep Fake DetectionabstractWith the advancement in natural-sounding speech production models, it is becoming important to develop models that can detect spoofed audios. Synthesized speech models do not explicitly account for all factors affecting speech production, such as the shape, size and structure of a speaker's vocal tract. In this paper, we hypothesize that due to practical limitations of audio corpora (including size, distribution, and balance of variables like gender, age, and accents), there exist certain phonemes that synthesized models are not able to replicate as well as the human articulation system and such phonemes differ in their spectral characteristics from bonafide speech. To discover such phonemes and quantify their effectiveness in distinguishing between spoofed and bonafide speech, we use a deep learning model with self-attention, and analyze the attention weights of the trained model. We use the ASVSpoof2019 dataset for our analysis and find that the attention mechanism picks most on fricatives: /S/,/SH/, nasals: /M/,/N/, vowels: /Y/, and stops: /D/. Furthermore, we obtain 7.54% EER on train and 11.98% on dev data when using only the top-16 most attended phonemes from input audio, better than when any other phoneme classes are used. Hira Dhamyal, Ayesha Ali, Ihsan Ayyub Qazi, Agha Ali Raza |
ASRU | 1 |
| 2021 | Fake Audio Detection in Resource-Constrained Settings Using Microfeatures
Hira Dhamyal, Ayesha Ali, Ihsan Ayyub Qazi, Agha Ali Raza |
Interspeech | 1 |
| 2021 | Masked Proxy Loss for Text-Independent Speaker VerificationabstractOpen-set speaker recognition can be regarded as a metric learning problem, which is to maximize inter-class variance and minimize intra-class variance. Supervised metric learning can be categorized into entity-based learning and proxy-based learning. Most of the existing metric learning objectives like Contrastive, Triplet, Prototypical, GE2E, etc all belong to the former division, the performance of which is either highly dependent on sample mining strategy or restricted by insufficient label information in the mini-batch. Proxy-based losses mitigate both shortcomings, however, fine-grained connections among entities are either not or indirectly leveraged. This paper proposes a Masked Proxy (MP) loss which directly incorporates both proxy-based relationships and pair-based relationships. We further propose Multinomial Masked Proxy (MMP) loss to leverage the hardness of speaker pairs. These methods have been applied to evaluate on VoxCeleb test set and reach state-of-the-art Equal Error Rate(EER). Jiachen Lian, Aiswarya Vinod Kumar, Hira Dhamyal, Bhiksha Raj, Rita Singh |
Interspeech | 3 |
| 2020 | The Phonetic Bases of Vocal Expressed Emotion: Natural versus ActedabstractCan vocal emotions be emulated? This question has been a recurrent concern of the speech community, and has also been vigorously investigated. It has been fueled further by its link to the issue of validity of acted emotion databases. Much of the speech and vocal emotion research has relied on acted emotion databases as valid proxies for studying natural emotions. To create models that generalize to natural settings, it is crucial to work with valid prototypes -- ones that can be assumed to reliably represent natural emotions. More concretely, it is important to study emulated emotions against natural emotions in terms of their physiological, and psychological concomitants. In this paper, we present an on-scale systematic study of the differences between natural and acted vocal emotions. We use a self-attention based emotion classification model to understand the phonetic bases of emotions by discovering the most 'attended' phonemes for each class of emotions. We then compare these attended-phonemes in their importance and distribution across acted and natural classes. Our tests show significant differences in the manner and choice of phonemes in acted and natural speech, concluding moderate to low validity and value in using acted speech databases for emotion classification tasks. Hira Dhamyal, Shahan Ali Memon, Bhiksha Raj, Rita Singh |
INTERSPEECH | 1 |
| 2019 | Optimizing Neural Network Embeddings Using a Pair-Wise Loss for Text-Independent Speaker VerificationabstractThis paper proposes a new loss function called the “quartet” loss for the better optimization of the neural networks for matching tasks. For such tasks, where neural network embeddings are the key component, the optimization of the network for better embeddings is critical. The embeddings are required to be class discriminative, resulting in minimal inter-class variation and maximal intra-class variation even for unseen classes for better generalization of the network. The quartet loss explicitly computes the distance metric between pairs of inputs and increases the gap between the similarity score distributions between the same class pairs and the different class pairs. We evaluate on the speaker verification task and demonstrate the performance of the loss on our proposed neural network. Hira Dhamyal, Tianyan Zhou, Bhiksha Raj, Rita Singh |
ASRU | 1 |