VLDB 2026 Research / reviewers in the wild / expert
Shekhar Nayak
dblp:158/4180
· DBLP profile ↗
13ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0002-4277-4851ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 10 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AMuSeD: An Attentive Deep Neural Network for Multimodal Sarcasm Detection Incorporating Bimodal Data AugmentationabstractDetecting sarcasm effectively requires a nuanced understanding of context, including vocal tones and facial expressions. The progression towards multimodal computational methods in sarcasm detection, however, faces challenges due to the scarcity of data. To address this, we present AMuSeD (Attentive deep neural network for MUltimodal Sarcasm dEtection incorporating bi-modal Data augmentation). This approach utilizes the Multimodal Sarcasm Detection Dataset (MUStARD) and introduces a two-phase bimodal data augmentation strategy. The first phase involves generating varied text samples through Back-Translation from several secondary languages. The second phase involves the refinement of a FastSpeech2-based speech synthesis system, tailored specifically for sarcasm to retain sarcastic intonations. Alongside a cloud-based Text-to-Speech (TTS) service, this Fine-tuned FastSpeech2 system produces corresponding audio for the text augmentations. We also evaluate various attention mechanisms for selectively enhancing sarcasm-relevant features, finding self-attention to be the most efficient. Our experiments reveal that the proposed approach achieves a significant F1-score of 81.0% in text-audio modalities, surpassing even models that use three modalities from the MUStARD dataset. Xiyuan Gao, Shubhi Bansal, Kushaan Gowda, Shekhar Nayak, Nagendra Kumar 0001, Matt Coler |
IEEE Trans. Affect. Comput. | 5 |
| 2025 | Intra-modal Relation and Emotional Incongruity Learning using Graph Attention Networks for Multimodal Sarcasm DetectionabstractSarcasm detection poses unique challenges due to the complex nature of sarcastic expressions often embedded across multiple modalities. Current methods frequently fall short in capturing the incongruent emotional cues that are essential for identifying sarcasm in multimodal contexts. In this paper, we present a novel method to capture the pair-wise emotional incongruities between modalities through a cross-modal Contrastive Attention Mechanism (CAM), leveraging advanced data augmentation techniques to enhance data diversity and Supervised Contrastive Learning (SCL) to obtain discriminative embeddings. Additionally, we employ Graph Attention Networks (GATs) to construct modality-specific graphs, capturing intra-modal dependencies. Experiments conducted on the MUStARD++ dataset demonstrate the efficacy of our approach, achieving a macro F1 score of 74.96%, which outperforms state-of-the-art methods. Devraj Raghuvanshi, Xiyuan Gao, Shubhi Bansal, Matt Coler, Nagendra Kumar 0001, Shekhar Nayak |
ICASSP | 7 |
| 2025 | A Multimodal Chinese Dataset for Cross-lingual Sarcasm DetectionabstractSarcasm is expressed through subtle cues like pitch, speech rate, and facial expressions, with patterns varying across languages, e.g., English speakers lower the pitch while Cantonese speakers raise it. While humans readily interpret these signals, computational models struggle, creating challenges for Human-Machine Interaction. Most multimodal sarcasm recognition research focuses on English and the lack of high-quality datasets for other languages hinders cross-lingual and cross-cultural studies. We introduce the Multimodal Chinese Sarcasm Dataset (MCSD), containing 10.57 hours of video. We propose a standardized annotation framework that captures annotator certainty to reflect the subjectivity of sarcasm, achieving a Fleiss'kappa of 0.74 (unweighted) and 0.79 (certainty-weighted). Validation of our dataset using SVM achieves a 76.64% F1-score in sarcasm detection. MCSD lays the foundation for robust cross-lingual sarcasm detection, contributing to advanced, human-centric systems. Xiyuan Gao, Bruce Xiao Wang, Shuming Huang, Shekhar Nayak, Matt Coler |
INTERSPEECH | 6 |
| 2025 | Leveraging Large Language Models for Sarcastic Speech Annotation in Sarcasm DetectionabstractSarcasm fundamentally alters meaning through tone and context, yet detecting it in speech remains a challenge due to data scarcity. In addition, existing detection systems often rely on multimodal data, limiting their applicability in contexts where only speech is available. To address this, we propose an annotation pipeline that leverages large language models (LLMs) to generate a sarcasm dataset. Using a publicly available sarcasm-focused podcast, we employ GPT-4o and LLaMA 3 for initial sarcasm annotations, followed by human verification to resolve disagreements. We validate this approach by comparing annotation quality and detection performance on a publicly available sarcasm dataset using a collaborative gating architecture. Finally, we introduce PodSarc, a large-scale sarcastic speech dataset created through this pipeline. The detection model achieves a 73.63% F1 score, demonstrating the dataset's potential as a benchmark for sarcasm detection research. Yuqing Zhang 0003, Xiyuan Gao, Shekhar Nayak, Matt Coler |
INTERSPEECH | 4 |
| 2025 | Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition - Multimodal Fusion, Challenges, and Future ProspectsabstractSarcasm, a common feature of human communication, poses challenges in interpersonal interactions and human-machine interactions. Linguistic research has highlighted the importance of prosodic cues, such as variations in pitch, speaking rate, and intonation, in conveying sarcastic intent. Although previous work has focused on text-based sarcasm detection, the role of speech data in recognizing sarcasm has been underexplored. Recent advancements in speech technology emphasize the growing importance of leveraging speech data for automatic sarcasm recognition, which can enhance social interactions for individuals with neurodegenerative conditions and improve machine understanding of complex human language use, leading to more nuanced interactions. This systematic review is the first to focus on speech-based sarcasm recognition, charting the evolution from unimodal to multimodal approaches. It covers datasets, feature extraction, and classification methods, and aims to bridge gaps across diverse research domains. The findings include limitations in datasets for sarcasm recognition in speech, the evolution of feature extraction techniques from traditional acoustic features to deep learning-based representations, and the progression of classification methods from unimodal approaches to multimodal fusion techniques. In so doing, we identify the need for greater emphasis on cross-cultural and multilingual sarcasm recognition, as well as the importance of addressing sarcasm as a multimodal phenomenon, rather than a text-based challenge. Xiyuan Gao, Shekhar Nayak, Matt Coler |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | Fine-Tuning Strategies for Dutch Dysarthric Speech Recognition: Evaluating the Impact of Healthy, Disease-Specific, and Speaker-Specific DataabstractDespite significant advancements in automatic speech recognition technology (ASR) the performance of such systems on dysarthric speech is still inadequate for widespread use. One key reason is the lack of sufficiently rich and diverse dysarthric speech datasets to train machine learning models that could handle all types and varieties of such speech. Motivated by the data scarcity problem, as well as by successful applications of self-supervised learning (SSL) in ASR for low-resource languages, this paper investigates and evaluates the effectiveness of three different data-centric SSL training strategies in improving Dutch dysarthric speech recognition. The first strategy involves fine-tuning with both dysarthric and healthy speech data, the second with disease-specific data and the third with speaker-specific data. The first and third strategies are proven effective, while the second one, though ineffective, provides valuable insights for further research. Spyretta Leivaditi, Tatsunari Matsushima, Matt Coler, Shekhar Nayak, Vass Verkhodanova |
INTERSPEECH | 4 |
| 2024 | A Functional Trade-off between Prosodic and Semantic Cues in Conveying SarcasmabstractThis study investigates the acoustic features of sarcasm and disentangles the interplay between the propensity of an utterance being used sarcastically and the presence of prosodic cues signaling sarcasm. Using a dataset of sarcastic utterances compiled from television shows, we analyze the prosodic features within utterances and key phrases belonging to three distinct sarcasm categories (embedded, propositional, and illocutionary), which vary in the degree of semantic cues present, and compare them to neutral expressions. Results show that in phrases where the sarcastic meaning is salient from the semantics, the prosodic cues are less relevant than when the sarcastic meaning is not evident from the semantics, suggesting a trade-off between prosodic and semantic cues of sarcasm at the phrase level. These findings highlight a lessened reliance on prosodic modulation in semantically dense sarcastic expressions and a nuanced interaction that shapes the communication of sarcastic intent. Xiyuan Gao, Yuqing Zhang 0003, Shekhar Nayak, Matt Coler |
INTERSPEECH | 4 |
| 2022 | Deep CNN-based Inductive Transfer Learning for Sarcasm Detection in SpeechabstractSarcasm is a frequently used linguistic device which is expressed in a multitude of ways, both with acoustic cues (including pitch, intonation, intensity, etc.) and visual cues (including facial expression, eye gaze, etc.). While cues used in the expression of sarcasm are well-described in the literature, there is a striking paucity of attempts to perform automatic sarcasm detection in speech. To explore this gap, we elaborate a methodology of implementing Inductive Transfer Learning (ITL) based on pre-trained Deep Convolutional Neural Networks (DCNNs) to detect sarcasm in speech. To those ends, the multimodal dataset MUStARD is used as a target dataset in this study. The two selected pre-trained DCNN models used are Xception and VGGish, which we trained on visual and audio datasets. Results show that VGGish, which is applied as a feature extractor in the experiment, performs better than Xception, which has its convolutional layers and pooling layers retrained. Both models achieve a higher F-score compared to the baseline Support Vector Machines (SVM) model by 7% and 5% in unimodal sarcasm detection in speech. Xiyuan Gao, Shekhar Nayak, Matt Coler |
INTERSPEECH | 2 |
| 2022 | Improving Luxembourgish Speech Recognition with Cross-Lingual Speech RepresentationsabstractLuxembourgish is a West Germanic language spoken by roughly 390,000 people, mainly in Luxembourg. It is one of Europe's under-described and under-resourced languages, not extensively investigated in the context of speech recognition. We explore the self-supervised multilingual learning of Luxembourgish speech representations for the speech recognition downstream task. We show that learning cross-lingual representations is essential for low-resourced languages such as Luxembourgish. Learning cross-lingual representations and rescoring the output transcriptions with language modelling while using only 4 hours of labelled speech achieves a word error rate of 15.1% and improves our Transfer Learning baseline model relatively by 33.1% and absolutely by 7.5%. Increasing the amount of labelled speech to 14 hours yields a significant performance gain resulting in a 9.3% word error rate.11Models and datasets are available at https://hugging£ace.co/lemswasabi Le Minh Nguyen 0002, Shekhar Nayak, Matt Coler |
SLT | 2 |
| 2019 | Zero Resource Speaking Rate Estimation from Change Point Detection of Syllable-like UnitsabstractSpeaking rate is an important attribute of the speech signal which plays a crucial role in the performance of automatic speech processing systems. In this paper, we propose to estimate the speaking rate by segmenting the speech into syllable-like units using end point detection algorithms which do not require any training and fine-tuning. Also, there are no predefined constraints on the expected number of syllabic segments. The acoustic subword units are obtained only from speech signal to estimate the speaking rate without any requirement of transcriptions or phonetic knowledge of the speech data. A recent theta-rate oscillator based syllabification algorithm is also employed for speaking rate estimation. The performance is evaluated on TIMIT corpus and spontaneous speech from Switchboard corpus. The correlation results are comparable to recent algorithms which are trained with specific training set and/or make use of the available transcriptions. Shekhar Nayak, Saurabhchand Bhati, K. Sri Rama Murty |
ICASSP | 1 |
| 2019 | Unsupervised Acoustic Segmentation and Clustering Using Siamese Network EmbeddingsabstractUnsupervised discovery of acoustic units from the raw speech signal forms the core objective of zero-resource speech processing. It involves identifying the acoustic segment boundaries and consistently assigning unique labels to acoustically similar segments. In this work, the possible candidates for segment boundaries are identified in an unsupervised manner from the kernel Gram matrix computed from the Mel-frequency cepstral coefficients (MFCC). These segment boundary candidates are used to train a siamese network, that is intended to learn embeddings that minimize intrasegment distances and maximize the intersegment distances. The siamese embeddings capture phonetic information from longer contexts of the speech signal and enhance the intersegment discriminability. These properties make the siamese embeddings better suited for acoustic segmentation and clustering than the raw MFCC features. The Gram matrix computed from the siamese embeddings provides unambiguous evidence for boundary locations. The initial candidate boundaries are refined using this evidence, and siamese embeddings are extracted for the new acoustic segments. A graph growing approach is used to cluster the siamese embeddings, and a unique label is assigned to acoustically similar segments. The performance of the proposed method for acoustic segmentation and clustering is evaluated on Zero Resource 2017 database. Saurabhchand Bhati, Shekhar Nayak, K. Sri Rama Murty, Najim Dehak |
INTERSPEECH | 2 |
| 2017 | Unsupervised Speech Signal to Symbol Transformation for Zero Resource Speech ApplicationsabstractZero resource speech processing refers to a scenario where no or minimal transcribed data is available. In this paper, we propose a three-step unsupervised approach to zero resource speech processing, which does not require any other information/dataset. In the first step, we segment the speech signal into phonemelike units, resulting in a large number of varying length segments. The second step involves clustering the varying-length segments into a finite number of clusters so that each segment can be labeled with a cluster index. The unsupervised transcriptions, thus obtained, can be thought of as a sequence of virtual phone labels. In the third step, a deep neural network classifier is trained to map the feature vectors extracted from the signal to its corresponding virtual phone label. The virtual phone posteriors extracted from the DNN are used as features in the zero resource speech processing. The effectiveness of the proposed approach is evaluated on both ABX and spoken term discovery tasks (STD) using spontaneous American English and Tsonga language datasets, provided as part of zero resource 2015 challenge. It is observed that the proposed system outperforms baselines, supplied along the datasets, in both the tasks without any task specific modifications Saurabhchand Bhati, Shekhar Nayak, K. Sri Rama Murty |
INTERSPEECH | 2 |
| 2014 | Unsupervised spoken word retrieval using Gaussian-bernoulli restricted boltzmann machinesabstractThe objective of this work is to explore a novel unsupervised framework, using Restricted Boltzmann machines, for Spoken Word Retrieval (SWR). In the absence of labelled speech data, SWR is typically performed by matching sequence of feature vectors of query and test utterances using dynamic time warping (DTW). In such a scenario, performance of SWR system critically depends on representation of the speech signal. Typical features, like mel-frequency cepstral coefficients (MFCC), carry significant speaker-specific information, and hence may not be used directly in SWR system. To overcome this issue, we propose to capture the joint density of the acoustic space spanned by MFCCs using Gaussian-Bernoulli restricted Boltzmann machine (GBRBM). In this work, we have used hidden activations of the GBRBM as features for SWR system. Since the GBRBM is trained with speech data collected from large number of speakers, the hidden activations are more robust to the speaker-variability compared to MFCCs. The performance of the proposed features is evaluated on Telugu broadcast news data, and an absolute improvement of 12% was observed compared to MFCCs. Raghavendra Reddy Pappagari, Shekhar Nayak, K. Sri Rama Murty |
INTERSPEECH | 2 |