Raghavendra Pappagari

dblp:201/7113 · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
6since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 8 first-author · 5 since 2021Artificial intelligence and machine learning · 9 · 5 first-author · 5 since 2021
YearPublicationVenuePosition
2022 Non-contrastive self-supervised learning of utterance-level speech representations
Raghavendra Pappagari, Piotr Zelasko, Laureano Moro-Velázquez, Jesús Villalba 0001, Najim Dehak
INTERSPEECH2
2021 Joint Prediction of Truecasing and Punctuation for Conversational Speech in Low-Resource Scenarios
abstract
Capitalization and punctuation are important cues for comprehending written texts and conversational transcripts. Yet, many ASR systems do not produce punctuated and case-formatted speech transcripts. We propose to use a multi-task system that can exploit the relations between casing and punctuation to improve their prediction performance. Whereas text data for predicting punctuation and truecasing is seemingly abundant, we argue that written text resources are inadequate as training data for conversational models. We quantify the mismatch between written and conversational text domains by comparing the joint distributions of punctuation and word cases, and by testing our model cross-domain. Further, we show that by training the model in the written text domain and then transfer learning to conversations, we can achieve reasonable performance with less data.
Raghavendra Pappagari, Piotr Zelasko, Agnieszka Mikolajczyk, Piotr Pezik, Najim Dehak
ASRU1
2021 Beyond Isolated Utterances: Conversational Emotion Recognition
abstract
Speech emotion recognition is the task of recognizing the speaker's emotional state given a recording of their utterance. While most of the current approaches focus on inferring emotion from isolated utterances, we argue that this is not sufficient to achieve conversational emotion recognition (CER) which deals with recognizing emotions in conversations. In this work, we propose several approaches for CER by treating it as a sequence labeling task. We investigated transformer architecture for CER and, compared it with ResNet-34 and BiLSTM architectures in both contextual and contextless scenarios using IEMOCAP corpus. Based on the inner workings of the self-attention mechanism, we proposed DiverseCatAugment (DCA), an augmentation scheme, which improved the transformer model performance by an absolute 3.3% micro-f1 on conversations and 3.6% on isolated utterances. We further enhanced the performance by introducing an interlocutor-aware transformer model where we learn a dictionary of interlocutor index embeddings to exploit diarized conversations.
Raghavendra Pappagari, Piotr Zelasko, Jesús Villalba 0001, Laureano Moro-Velázquez, Najim Dehak
ASRU1
2021 CopyPaste: An Augmentation Method for Speech Emotion Recognition
abstract
Data augmentation is a widely used strategy for training robust machine learning models. It partially alleviates the problem of limited data for tasks like speech emotion recognition (SER), where collecting data is expensive and challenging. This study proposes CopyPaste, a perceptually motivated novel augmentation procedure for SER. Assuming that the presence of emotions other than neutral dictates a speaker’s overall perceived emotion in a recording, concatenation of an emotional (emotion E) and a neutral utterance can still be labeled with emotion E. We hypothesize that SER performance can be improved using these concatenated utterances in model training. To verify this, three CopyPaste schemes are tested on two deep learning models: one trained independently and another using transfer learning from an x-vector model, a speaker recognition model. We observed that all three CopyPaste schemes improve SER performance on all the three datasets considered: MSP-Podcast, Crema-D, and IEMOCAP. Additionally, CopyPaste performs better than noise augmentation and, using them together improves the SER performance further. Our experiments on noisy test sets suggested that CopyPaste is effective even in noisy test conditions.
Raghavendra Pappagari, Jesús Villalba 0001, Piotr Zelasko, Laureano Moro-Velázquez, Najim Dehak
ICASSP1
2021 Automatic Detection and Assessment of Alzheimer Disease Using Speech and Language Technologies in Low-Resource Scenarios
Raghavendra Pappagari, Sonal Joshi, Laureano Moro-Velázquez, Piotr Zelasko, Jesús Villalba 0001, Najim Dehak
Interspeech1
2021 What Helps Transformers Recognize Conversational Structure? Importance of Context, Punctuation, and Labels in Dialog Act Recognition
abstract
Abstract Dialog acts can be interpreted as the atomic units of a conversation, more fine-grained than utterances, characterized by a specific communicative function. The ability to structure a conversational transcript as a sequence of dialog acts—dialog act recognition, including the segmentation—is critical for understanding dialog. We apply two pre-trained transformer models, XLNet and Longformer, to this task in English and achieve strong results on Switchboard Dialog Act and Meeting Recorder Dialog Act corpora with dialog act segmentation error rates (DSER) of 8.4% and 14.2%. To understand the key factors affecting dialog act recognition, we perform a comparative analysis of models trained under different conditions. We find that the inclusion of a broader conversational context helps disambiguate many dialog act classes, especially those infrequent in the training data. The presence of punctuation in the transcripts has a massive effect on the models’ performance, and a detailed analysis reveals specific segmentation patterns observed in its absence. Finally, we find that the label set specificity does not affect dialog act segmentation performance. These findings have significant practical implications for spoken language understanding applications that depend heavily on a good-quality segmentation being available.
Piotr Zelasko, Raghavendra Pappagari, Najim Dehak
Trans. Assoc. Comput. Linguistics2
2020 X-Vectors Meet Emotions: A Study On Dependencies Between Emotion and Speaker Recognition
abstract
In this work, we explore the dependencies between speaker recognition and emotion recognition. We first show that knowledge learned for speaker recognition can be reused for emotion recognition through transfer learning. Then, we show the effect of emotion on speaker recognition. For emotion recognition, we show that using a simple linear model is enough to obtain good performance on the features extracted from pre-trained models such as the x-vector model. Then, we improve emotion recognition performance by finetuning for emotion classification. We evaluated our experiments on three different types of datasets: IEMOCAP, MSP-Podcast, and Crema-D. By fine-tuning, we obtained 30.40%, 7.99%, and 8.61% absolute improvement on IEMOCAP, MSP-Podcast, and Crema-D respectively over baseline model with no pre-training. Finally, we present results on the effect of emotion on speaker verification. We observed that speaker verification performance is prone to changes in test speaker emotions. We found that trials with angry utterances performed worst in all three datasets. We hope our analysis will initiate a new line of research in the speaker recognition community.
Raghavendra Pappagari, Tianzi Wang, Jesús Villalba 0001, Nanxin Chen, Najim Dehak
ICASSP1
2020 Using State of the Art Speaker Recognition and Natural Language Processing Technologies to Detect Alzheimer's Disease and Assess its Severity
Raghavendra Pappagari, Laureano Moro-Velázquez, Najim Dehak
INTERSPEECH1
2019 Can You Tell Me How to Get Past Sesame Street? Sentence-Level Pretraining Beyond Language Modeling
abstract
Alex Wang, Jan Hula, Patrick Xia, Raghavendra Pappagari, R. Thomas McCoy, Roma Patel, Najoung Kim, Ian Tenney, Yinghui Huang, Katherin Yu, Shuning Jin, Berlin Chen, Benjamin Van Durme, Edouard Grave, Ellie Pavlick, Samuel R. Bowman. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Jan Hula, Patrick Xia 0002, Raghavendra Pappagari, Tom McCoy 0001, Roma Patel, Najoung Kim, Ian Tenney, Katherin Yu, Shuning Jin, Berlin Chen, Benjamin Van Durme, Edouard Grave, Ellie Pavlick, Samuel R. Bowman
ACL (1)4
2019 Hierarchical Transformers for Long Document Classification
abstract
BERT, which stands for Bidirectional Encoder Representations from Transformers, is a recently introduced language representation model based upon the transfer learning paradigm. We extend its fine-tuning procedure to address one of its major limitations - applicability to inputs longer than a few hundred words, such as transcripts of human call conversations. Our method is conceptually simple. We segment the input into smaller chunks and feed each of them into the base model. Then, we propagate each output through a single recurrent layer, or another transformer, followed by a softmax activation. We obtain the final classification decision after the last segment has been consumed. We show that both BERT extensions are quick to fine-tune and converge after as little as 1 epoch of training on a small, domain-specific data set. We successfully apply them in three different tasks involving customer call satisfaction prediction and topic classification, and obtain a significant improvement over the baseline models in two of them.
Raghavendra Pappagari, Piotr Zelasko, Jesús Villalba 0001, Yishay Carmiel, Najim Dehak
ASRU1
2018 Joint Verification-Identification in end-to-end Multi-Scale CNN Framework for Topic Identification
abstract
We present an end-to-end multi-scale Convolutional Neural Network (CNN) framework for topic identification (topic ID). In this work, we examined multi -scale CNN for classification using raw text input. Topical word embeddings are learnt at multiple scales using parallel convolutional layers. A technique to integrate verification and identification objectives is examined to improve topic ID performance. With this approach, we achieved significant improvement in identification task. We evaluated our framework on two contrasting datasets: 20 newsgroups and Fisher. We obtained 92.93% accuracy on Fisher and 86.12% on 20 newsgroups, which to our know ledge are the best published results on these datasets at the moment.
Raghavendra Pappagari, Jesús Villalba 0001, Najim Dehak
ICASSP1
2018 Deep Neural Networks for Emotion Recognition Combining Audio and Transcripts
abstract
In this paper, we propose to improve emotion recognition by combining acoustic information and conversation transcripts.On the one hand, a LSTM network was used to detect emotion from acoustic features like f0, shimmer, jitter, MFCC, etc.On the other hand, a multi-resolution CNN was used to detect emotion from word sequences.This CNN consists of several parallel convolutions with different kernel sizes to exploit contextual information at different levels.A temporal pooling layer aggregates the hidden representations of different words into a unique sequence level embedding, from which we computed the emotion posteriors.We optimized a weighted sum of classification and verification losses.The verification loss tries to bring embeddings from same emotions closer while separating embeddings from different emotions.We also compared our CNN with state-of-the-art text-based hand-crafted features (e-vector).We evaluated our approach on the USC-IEMOCAP dataset as well as the dataset consisting of US English telephone speech.In the former, we used human-annotated transcripts while in the latter, we used ASR transcripts.The results showed fusing audio and transcript information improved unweighted accuracy by relative 24% for IEMOCAP and relative 3.4% for the telephone data compared to a single acoustic system.
Raghavendra Pappagari, Purva Kulkarni, Jesús Villalba 0001, Yishay Carmiel, Najim Dehak
INTERSPEECH2
2017 Topic identification of spoken documents using unsupervised acoustic unit discovery
abstract
This paper investigates the application of unsupervised acoustic unit discovery for topic identification (topic ID) of spoken audio documents. The acoustic unit discovery method is based on a non-parametric Bayesian phone-loop model that segments a speech utterance into phone-like categories. The discovered phone-like (acoustic) units are further fed into the conventional topic ID framework. Using multilingual bottleneck features for the acoustic unit discovery, we show that the proposed method outperforms other systems that are based on cross-lingual phoneme recognizer.
Santosh Kesiraju, Raghavendra Pappagari, Lucas Ondel Yang, Lukás Burget, Najim Dehak, Sanjeev Khudanpur, Jan Cernocký, Suryakanth V. Gangashetty
ICASSP2