VLDB 2026 Research / reviewers in the wild / expert
Emily Mower Provost
dblp:58/4610 · also Emily Mower
· DBLP profile ↗
92ranked-venue papers
16as first author
23since 2021 · last 2025
0000-0003-1870-6063ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 63 · 12 first-author · 12 since 2021Artificial intelligence and machine learning · 50 · 7 first-author · 18 since 2021Human-computer interaction and ubiquitous computing · 12 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | More Similar than Dissimilar: Modeling Annotators for Cross-Corpus Speech Emotion RecognitionabstractSpeech emotion recognition systems often predict a consensus value generated from the ratings of multiple annotators. However, these models have limited ability to predict the annotation of any one person. Alternatively, models can learn to predict the annotations of all annotators. Adapting such models to new annotators is difficult as new annotators must individually provide sufficient labeled training data. We propose to leverage inter-annotator similarity by using a model pre-trained on a large annotator population to identify a similar, previously seen annotator. Given a new, previously unseen, annotator and limited enrollment data, we can make predictions for a similar annotator, enabling off-the-shelf annotation of unseen data in target datasets, providing a mechanism for extremely low-cost personalization. We demonstrate our approach significantly outperforms other off-the-shelf approaches, paving the way for lightweight emotion adaptation, practical for real-world deployment. James Tavernor, Emily Mower Provost |
ASRU | 2 |
| 2025 | Efficient Finetuning for Dimensional Speech Emotion Recognition in the Age of TransformersabstractAccurate speech emotion recognition is essential for developing human-facing systems. Recent advancements have included finetuning large, pretrained transformer models like Wav2Vec 2.0. However, the finetuning process requires substantial computational resources, including high-memory GPUs and significant processing time. As the demand for accurate emotion recognition continues to grow, efficient finetuning approaches are needed to reduce the computational burden. Our study focuses on dimensional emotion recognition, predicting attributes such as activation (calm to excited) and valence (negative to positive). We present various finetuning techniques, including full finetuning, partial finetuning of transformer layers, finetuning with mixed precision, partial finetuning with caching, and low-rank adaptation (LoRA) on the Wav2Vec 2.0 base model. We find that partial finetuning with mixed precision achieves performance comparable to full finetuning while increasing training speed by 67%. Caching intermediate representations further boosts efficiency, yielding an 88% speedup and a 71% reduction in learnable parameters. We recommend finetuning the final three transformer layers in mixed precision to balance performance and training efficiency, and adding intermediate representation caching for optimal speed with minimal performance trade-offs. These findings lower the barriers to finetuning speech emotion recognition systems, making accurate emotion recognition more accessible to a broader range of researchers and practitioners. Aneesha Sampath, James Tavernor, Emily Mower Provost |
ICASSP | 3 |
| 2025 | Large language models accurately identify immunosuppression in intensive care unit patientsabstractOBJECTIVE: Rule-based structured data algorithms and natural language processing (NLP) approaches applied to unstructured clinical notes have limited accuracy and poor generalizability for identifying immunosuppression. Large language models (LLMs) may effectively identify patients with heterogenous types of immunosuppression from unstructured clinical notes. We compared the performance of LLMs applied to unstructured notes for identifying patients with immunosuppressive conditions or immunosuppressive medication use against 2 baselines: (1) structured data algorithms using diagnosis codes and medication orders and (2) NLP approaches applied to unstructured notes. MATERIALS AND METHODS: We used hospital admission notes from a primary cohort of 827 intensive care unit (ICU) patients at Northwestern Memorial Hospital and a validation cohort of 200 ICU patients at Beth Israel Deaconess Medical Center, along with diagnosis codes and medication orders from the primary cohort. We evaluated the performance of structured data algorithms, NLP approaches, and LLMs in identifying 7 immunosuppressive conditions and 6 immunosuppressive medications. RESULTS: In the primary cohort, structured data algorithms achieved peak F1 scores ranging from 0.30 to 0.97 for identifying immunosuppressive conditions and medications. NLP approaches achieved peak F1 scores ranging from 0 to 1. GPT-4o outperformed or matched structured data algorithms and NLP approaches across all conditions and medications, with F1 scores ranging from 0.51 to 1. GPT-4o also performed impressively in our validation cohort (F1 = 1 for 8/13 variables). DISCUSSION: LLMs, particularly GPT-4o, outperformed structured data algorithms and NLP approaches in identifying immunosuppressive conditions and medications with robust external validation. CONCLUSION: LLMs can be applied for improved cohort identification for research purposes. Vijeeth Guggilla, Mengjia Kang, Melissa J. Bak, Steven D. Tran, Anna Pawlowski, Prasanth Nannapaneni, Luke V. Rasmussen, Helen K. Donnelly, Ankit Agrawal 0001, David M. Liebovitz, Alexander V. Misharin, G. R. Scott Budinger, Richard G. Wunderink, Theresa Walunas, Catherine A. Gao, Alan R. Hauser, Alec Peltekian, Alexis Rose Wolfe, Alison L. Szabo, Alok N. Choudhary, Amy Ludwig, Anahid Amani Moghadam, Anjana V. Yeldandi, Ankit Bharat, Anna E. Pawlowski, Anthony M. Joudi, Arjun Prakash Tambe, Ashley J. Smith-Nunez, Benjamin D. Singer, Benjamin J. Ulrich, Betty Tran, Cara J. Gottardi, Chiagozie O. Pickens, Clara J. Schroedl, Daniel Meza, Dulce Sarai Garcia, Egon A. Ozer, Elen Gusman, Elisheva D. Shanes, Emily Mower Provost, Emily M. Olson, Erica Marie Hartmann, Erin A. Korth, Estefani Diaz, Estefany R. Guzman, Francisco J. Martinez, Gabrielle Matias, Hiam Abdala-Valencia, Jack T. Sumner, Jacob I Sznajder, Jacqueline M. Kruser, Jakub Glowala, James M. Walter, Jamie H. Rowell, Jason M. Arnold, John Coleman, Jon W. Lomasney, Joseph Isaac Bailey, Judd F. Hultquist, Justin A. Fiala, Justin Starren, Karen M. Ridge, Karolina Senkow, Kathryn A. Helmin, Khalilah L. Gates, Lacy Simmons, Lesley Pinzon, Lindsey D. Gradone, Lisa F. Wolfe, Lucy Luo, Luisa Morales-Nebreda, Manu Jain, Marc Sala, Maxwell Schleck, Melissa H. Ross, Melissa Querrey, Michael J. Cuttica, Michelle Hinsch Prickett, Nandita R. Nadig, Nathaniel Rhodes, Navdeep S. Chandel, Nikolay S. Markov, Peter H. S. Sporn, Qianli Liu, Rachel B. Kadar, Rachel L. Medernach, Ramon Lorenzo-Redondo, Ravi Kalhan, Rebecca K. Clepp, Richard I. Morimoto, Rogan A. Grant, Ruben J. Mylvaganam, Samuel Fenske, Scott A. Laurenzo, Seung Hye Han, Sophia Nozick, Srinivas Panchamukhi, Stephanie C. Eisenbarth, Suchitra Swaminathan, Susan R. Russell, Taylor A. Poor, Thaddeus Cybulski, Theresa A. Lombardo, Thomas Bolig, Thomas Stoeger, Tien Doan, Timothy Rowe, Wan-Ting Liao, Yuan Luo 0001, Yuliana Sokolenko, Ziyan Lu |
J. Am. Medical Informatics Assoc. | 42 |
| 2025 | Rethinking Emotion Annotations in the Era of Large Language ModelsabstractModern affective computing systems rely heavily on datasets with human-annotated emotion labels for both training and evaluation. However, human annotations are expensive to obtain, sensitive to study design, and difficult to quality control, because of the subjective nature of emotions. Meanwhile, Large Language Models (LLMs) have shown remarkable performance on many Natural Language Understanding tasks, emerging as a promising tool for text annotation. In this work, we analyze the complexities of emotion annotation in the context of LLMs, focusing on GPT-4 as a leading model. In our experiments, GPT-4 achieves high ratings in a human evaluation study, painting a more positive picture than previous work, in which human labels served as the only ground truth. On the other hand, we observe differences between human and GPT-4 emotion perception, underscoring the importance of human input in annotation studies. To harness GPT-4's strength while preserving human perspective, we explore two ways of integrating GPT-4 into emotion annotation pipelines, showing its potential to flag low-quality labels, reduce the workload of human annotators, and improve downstream model learning performance and efficiency. Together, our findings highlight opportunities for new emotion labeling practices and suggest the use of LLMs as a promising tool to aid human annotation. Minxue Niu, Yara El-Tawil, Amrit Romana, Emily Mower Provost |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | Emotion Recognition in the Real World: Passively Collecting and Estimating Emotions From Natural Speech Data of Individuals With Bipolar DisorderabstractEmotions provide critical information regarding a person's health and well-being. Therefore, the ability to track emotion and patterns in emotion over time could provide new opportunities in measuring health longitudinally. This is of particular importance for individuals with bipolar disorder (BD), where emotion dysregulation is a hallmark symptom of increasing mood severity. However, measuring emotions typically requires self-assessment, a willful action outside of one's daily routine. In this paper, we describe a novel approach for collecting real-world natural speech data from daily life and measuring emotions from these data. The approach combines a novel data collection pipeline and validated robust emotion recognition models. We describe a deployment of this pipeline that included parallel clinical and self-report measures of mood and self-reported measures of emotion. Finally, we present approaches to estimate clinical and self-reported mood measures using a combination of passive and self-reported emotion measures. The results demonstrate that both passive and self-reported measures of emotion contribute to our ability to accurately estimate mood symptom severity for individuals with BD. Emily Mower Provost, Sarah H. Sperry, James Tavernor, Steve Anderau, Anastasia Yocum, Melvin G. McInnis |
IEEE Trans. Affect. Comput. | 1 |
| 2024 | From Text to Emotion: Unveiling the Emotion Annotation Capabilities of LLMs
Minxue Niu, Mimansa Jaiswal, Emily Mower Provost |
INTERSPEECH | 3 |
| 2024 | Beyond Binary: Multiclass Paraphasia Detection with Generative Pretrained Transformers and End-to-End Models
Matthew Perez, Aneesha Sampath, Minxue Niu, Emily Mower Provost |
INTERSPEECH | 4 |
| 2024 | The Whole Is Bigger Than the Sum of Its Parts: Modeling Individual Annotators to Capture Emotional VariabilityabstractEmotion expression and perception are nuanced, complex, and highly subjective processes. When multiple annotators label emotional data, the resulting labels contain high variability. Most speech emotion recognition tasks address this by averaging annotator labels as ground truth. However, this process omits the nuance of emotion and inter-annotator variability, which are important signals to capture. Previous work has attempted to learn distributions to capture emotion variability, but these methods also lose information about the individual annotators. We address these limitations by learning to predict individual annotators and by introducing a novel method to create distributions from continuous model outputs that permit the learning of emotion distributions during model training. We show that this combined approach can result in emotion distributions that are more accurate than those seen in prior work, in both within- and cross-corpus settings. James Tavernor, Yara El-Tawil, Emily Mower Provost |
INTERSPEECH | 3 |
| 2024 | Guest Editorial Best of ACII 2021abstractThe 9TH AAAC Conference on Affective Computing and Intelligent Interaction 2021 was held in a virtual format in the fall of 2021. It was technically co-sponsored by the IEEE Computer Society and featured the recent work on Affective Computing. The six best papers from this conference were selected by the technical program chairs. They were invited to submit their extended version to be considered for this special section at the IEEE Transactions on Affective Computing. Each submission was reviewed by at least three expert reviewers and was evaluated in terms of overall contribution and the adequacy of the additional content to warrant a new article. This special section features five accepted submissions whose major contributions are summarized below. Mohammad Soleymani 0001, Shiro Kumano, Emily Mower Provost, Nadia Bianchi-Berthouze, Akane Sano, Kenji Suzuki 0002 |
IEEE Trans. Affect. Comput. | 3 |
| 2024 | Automatic Disfluency Detection From Untranscribed SpeechabstractSpeech disfluencies, such as filled pauses or repetitions, are disruptions in the typical flow of speech. All speakers experience disfluencies at times, and the rate at which we produce disfluencies may be increased by certain speaker or environmental characteristics. Modeling disfluencies has been shown to be useful for a range of downstream tasks, and as a result, disfluency detection has many potential applications. In this work, we investigate language, acoustic, and multimodal methods for frame-level automatic disfluency detection and categorization. Each of these methods relies on audio as an input. First, we evaluate several automatic speech recognition (ASR) systems in terms of their ability to transcribe disfluencies, measured using disfluency error rates. We then use these ASR transcripts as input to a language-based disfluency detection model. We find that disfluency detection performance is largely limited by the quality of transcripts and alignments. We find that an acoustic-based approach that does not require transcription as an intermediate step outperforms the ASR language approach. Finally, we present multimodal architectures which we find improve disfluency detection performance over the unimodal approaches. Ultimately, this work introduces novel approaches for automatic frame-level disfluency and categorization. In the long term, this will help researchers incorporate automatic disfluency detection into a range of applications. Amrit Romana, Kazuhito Koishida, Emily Mower Provost |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Capturing Mismatch between Textual and Acoustic Emotion Expressions for Mood Identification in Bipolar Disorder
Minxue Niu, Amrit Romana, Mimansa Jaiswal, Melvin G. McInnis, Emily Mower Provost |
INTERSPEECH | 5 |
| 2023 | Episodic Memory For Domain-Adaptable, Robust Speech Emotion Recognition
James Tavernor, Matthew Perez, Emily Mower Provost |
INTERSPEECH | 3 |
| 2023 | An Engineering View on Emotions and Speech: From Analysis and Predictive Models to Responsible Human-Centered ApplicationsabstractThe substantial growth of Internet-of-Things technology and the ubiquity of smartphone devices has increased the public and industry focus on speech emotion recognition (SER) technologies. Yet, conceptual, technical, and societal challenges restrict the wide adoption of these technologies in various domains, including, healthcare, and education. These challenges are amplified when automated emotion recognition systems are called to function “in-the-wild” due to the inherent complexity and subjectivity of human emotion, the difficulty of obtaining reliable labels at high temporal resolution, and the diverse contextual and environmental factors that confound the expression of emotion in real life. In addition, societal and ethical challenges hamper the wide acceptance and adoption of these technologies, with the public raising questions about user privacy, fairness, and explainability. This article briefly reviews the history of affective speech processing, provides an overview of current state-of-the-art approaches to SER, and discusses algorithmic approaches to render these technologies accessible to all, maximizing their benefits and leading to responsible human-centered computing applications. Chi-Chun Lee, Theodora Chaspari, Emily Mower Provost, Shri Narayanan |
Proc. IEEE | 3 |
| 2023 | You're Not You When You're Angry: Robust Emotion Features Emerge by Recognizing SpeakersabstractThe robustness of an acoustic emotion recognition system hinges on first having access to features that represent an acoustic input signal. These representations should abstract extraneous low-level variations present in acoustic signals and only capture speaker characteristics relevant for emotion recognition. Previous research has demonstrated that, in other classification tasks, when large labeled datasets are available, neural networks trained on these data learn to extract robust features from the input signal. However, the datasets used for developing emotion recognition systems remain significantly smaller than those used for developing other speech systems. Thus, acoustic emotion recognition systems remain in need of robust feature representations. In this article, we study the utility of speaker embeddings, representations extracted from a trained speaker recognition network, as robust features for detecting emotions. We first study the relationship between emotions and speaker embeddings and demonstrate how speaker embeddings highlight the differences that exist between neutral speech and emotionally expressive speech. We quantify the modulations that variations in emotional expression incur on speaker embeddings and show how these modulations are greater than those incurred from lexical variations in an utterance. Finally, we demonstrate how speaker embeddings can be used as a replacement for traditional low-level acoustic features for emotion recognition. Zakaria Aldeneh, Emily Mower Provost |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Mind the gap: On the value of silence representations to lexical-based speech emotion recognition
Matthew Perez, Mimansa Jaiswal, Minxue Niu, Cristina Gorrostieta, Matthew Roddy, Kye Taylor, Reza Lotfian, Emily Mower Provost |
INTERSPEECH | 9 |
| 2022 | Enabling Off-the-Shelf Disfluency Detection and Categorization for Pathological SpeechabstractA speech disfluency, such as a filled pause, repetition, or revision, disrupts the typical flow of speech. Disfluency modeling has grown as a research area, as recent work has shown that these disfluencies may help in assessing health conditions. For example, for individuals with cognitive impairment, changes in disfluencies may indicate worsening symptoms. However, work on disfluency modeling has focused heavily on detection and less on categorization. Work that has focused on categorization has suffered with two specific classes: repetitions and revisions. In this paper, we evaluate how BERT (Bidirectional Encoder Representations from Transformers) compares to other models on disfluency detection and categorization. We also propose adding a second fine-tuning task where BERT learns to distance repetitions and revisions from their repairs with triplet loss. We find that BERT and BERT with triplet loss outperform previous work on disfluency detection and categorization, particularly for repetitions and revisions. In this paper we present the first analysis of how these models can be fine-tuned on widely available disfluency data, and then used in an off-the-shelf manner on small corpora of pathological speech. Amrit Romana, Minxue Niu, Matthew Perez, Angela Roberts 0001, Emily Mower Provost |
INTERSPEECH | 5 |
| 2021 | Towards Noise Robust Speech Emotion Recognition Using Dynamic Layer CustomizationabstractRobustness to environmental noise is important to creating automatic speech emotion recognition systems that are deployable in the real world. In this work, we experiment with two paradigms, one where we can anticipate noise sources that will be seen at test time and one where we cannot. In our first experiment, we assume that we have advance knowledge of the noise conditions that will be seen at test time. We show that we can use this knowledge to create "expert" feature encoders for each noise condition. If the noise condition is unchanging, data can be routed to a single encoder to improve robustness. However, if the noise source is variant, this paradigm is too restrictive. In-stead, we introduce a new approach, dynamic layer customization (DLC), that allows the data to be dynamically routed to noise-matched encoders and then recombined. Critically, this process maintains temporal order, enabling extensions for multimodal models that generally benefit from long-term context. In our second experiment, we investigate whether partial knowledge of noise seen at test time can still be used to train systems that generalize well to unseen noise conditions using state-of-the-art domain adaptation algorithms. We find that DLC enables performance increases in both cases, highlighting the utility of mixture-of-expert approaches, domain adaptation methods and DLC to noise robust automatic speech emotion recognition. Alex Wilf, Emily Mower Provost |
ACII | 2 |
| 2021 | Articulatory Coordination for Speech Motor Tracking in Huntington DiseaseabstractHuntington Disease (HD) is a progressive disorder which often manifests in motor impairment. Motor severity (captured via motor score) is a key component in assessing overall HD severity. However, motor score evaluation involves in-clinic visits with a trained medical professional, which are expensive and not always accessible. Speech analysis provides an attractive avenue for tracking HD severity because speech is easy to collect remotely and provides insight into motor changes. HD speech is typically characterized as having irregular articulation. With this in mind, acoustic features that can capture vocal tract movement and articulatory coordination are particularly promising for characterizing motor symptom progression in HD. In this paper, we present an experiment that uses Vocal Tract Coordination (VTC) features extracted from read speech to estimate a motor score. When using an elastic-net regression model, we find that VTC features significantly outperform other acoustic features across varied-length audio segments, which highlights the effectiveness of these features for both short- and long-form reading tasks. Lastly, we analyze the F-value scores of VTC features to visualize which channels are most related to motor score. This work enables future research efforts to consider VTC features for acoustic analyses which target HD motor symptomatology tracking. Matthew Perez, Amrit Romana, Angela Roberts 0001, Noelle Carlozzi, Jennifer Ann Miner, Praveen Dayalu, Emily Mower Provost |
Interspeech | 7 |
| 2021 | Automatically Detecting Errors and Disfluencies in Read Speech to Predict Cognitive Impairment in People with Parkinson's DiseaseabstractParkinson's disease (PD) is a central nervous system disorder that causes motor impairment. Recent studies have found that people with PD also often suffer from cognitive impairment (CI). While a large body of work has shown that speech can be used to predict motor symptom severity in people with PD, much less has focused on cognitive symptom severity. Existing work has investigated if acoustic features, derived from speech, can be used to detect CI in people with PD. However, these acoustic features are general and are not targeted toward capturing CI. Speech errors and disfluencies provide additional insight into CI. In this study, we focus on read speech, which offers a controlled template from which we can detect errors and disfluencies, and we analyze how errors and disfluencies vary with CI. The novelty of this work is an automated pipeline, including transcription and error and disfluency detection, capable of predicting CI in people with PD. This will enable efficient analyses of how cognition modulates speech for people with PD, leading to scalable speech assessments of CI. Amrit Romana, John Bandon, Matthew Perez, Stephanie Gutierrez, Richard Richter, Angela Roberts 0001, Emily Mower Provost |
Interspeech | 7 |
| 2021 | Learning Paralinguistic Features from Audiobooks through Style Voice ConversionabstractZakaria Aldeneh, Matthew Perez, Emily Mower Provost. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Zakaria Aldeneh, Matthew Perez, Emily Mower Provost |
NAACL-HLT | 3 |
| 2021 | Read speech voice quality and disfluency in individuals with recent suicidal ideation or suicide attempt
Brian Stasak, Julien Epps, Heather T. Schatten, Ivan W. Miller, Emily Mower Provost, Michael F. Armey |
Speech Commun. | 5 |
| 2021 | Improving Cross-Corpus Speech Emotion Recognition with Adversarial Discriminative Domain Generalization (ADDoG)abstractAutomatic speech emotion recognition provides computers with critical context to enable user understanding. While methods trained and tested within the same dataset have been shown successful, they often fail when applied to unseen datasets. To address this, recent work has focused on adversarial methods to find more generalized representations of emotional speech. However, many of these methods have issues converging, and only involve datasets collected in laboratory conditions. In this paper, we introduce Adversarial Discriminative Domain Generalization (ADDoG), which follows an easier to train "meet in the middle" approach. The model iteratively moves representations learned for each dataset closer to one another, improving cross-dataset generalization. We also introduce Multiclass ADDoG, or MADDoG, which is able to extend the proposed method to more than two datasets, simultaneously. Our results show consistent convergence for the introduced methods, with significantly improved results when not using labels from the target dataset. We also show how, in most cases, ADDoG and MADDoG can be used to improve upon baseline state-of-the-art methods when target dataset labels are added and in-the-wild data are considered. Even though our experiments focus on cross-corpus speech emotion, these methods could be used to remove unwanted factors of variation in other settings. John Gideon, Melvin G. McInnis, Emily Mower Provost |
IEEE Trans. Affect. Comput. | 3 |
| 2021 | Jointly Aligning and Predicting Continuous Emotion AnnotationsabstractTime-continuous dimensional descriptions of emotions (e.g., arousal, valence) allow researchers to characterize short-time changes and to capture long-term trends in emotion expression. However, continuous emotion labels are generally not synchronized with the input speech signal due to delays caused by reaction-time, which is inherent in human evaluations. To deal with this challenge, we introduce a new convolutional neural network (multi-delay sinc network) that is able to simultaneously align and predict labels in an end-to-end manner. The proposed network is a stack of convolutional layers followed by an aligner network that aligns the speech signal and emotion labels. This network is implemented using a new convolutional layer that we introduce, thedelayed sinc layer. It is a time-shifted low-pass (sinc) filter that uses a gradient-based algorithm to learn a single delay. Multiple delayed sinc layers can be used to compensate for a non-stationary delay that is a function of the acoustic space. We test the efficacy of this system on two common emotion datasets, RECOLA and SEWA, and show that this approach obtains state-of-the-art speech-only results by learning time-varying delays while predicting dimensional descriptors of emotions. Soheil Khorram, Melvin G. McInnis, Emily Mower Provost |
IEEE Trans. Affect. Comput. | 3 |
| 2020 | Privacy Enhanced Multimodal Neural Representations for Emotion RecognitionabstractMany mobile applications and virtual conversational agents now aim to recognize and adapt to emotions. To enable this, data are transmitted from users' devices and stored on central servers. Yet, these data contain sensitive information that could be used by mobile applications without user's consent or, maliciously, by an eavesdropping adversary. In this work, we show how multimodal representations trained for a primary task, here emotion recognition, can unintentionally leak demographic information, which could override a selected opt-out option by the user. We analyze how this leakage differs in representations obtained from textual, acoustic, and multimodal data. We use an adversarial learning paradigm to unlearn the private information present in a representation and investigate the effect of varying the strength of the adversarial component on the primary task and on the privacy metric, defined here as the inability of an attacker to predict specific demographic information. We evaluate this paradigm on multiple datasets and show that we can improve the privacy metric while not significantly impacting the performance on the primary task. To the best of our knowledge, this is the first work to analyze how the privacy metric differs across modalities and how multiple privacy concerns can be tackled while still maintaining performance on emotion recognition. Mimansa Jaiswal, Emily Mower Provost |
AAAI | 2 |
| 2020 | Aphasic Speech Recognition Using a Mixture of Speech Intelligibility ExpertsabstractRobust speech recognition is a key prerequisite for semantic feature extraction in automatic aphasic speech analysis. However, standard one-size-fits-all automatic speech recognition models perform poorly when applied to aphasic speech. One reason for this is the wide range of speech intelligibility due to different levels of severity (i.e., higher severity lends itself to less intelligible speech). To address this, we propose a novel acoustic model based on a mixture of experts (MoE), which handles the varying intelligibility stages present in aphasic speech by explicitly defining severity-based experts. At test time, the contribution of each expert is decided by estimating speech intelligibility with a speech intelligibility detector (SID). We show that our proposed approach significantly reduces phone error rates across all severity stages in aphasic speech compared to a baseline approach that does not incorporate severity information into the modeling process. Matthew Perez, Zakaria Aldeneh, Emily Mower Provost |
INTERSPEECH | 3 |
| 2020 | Classification of Manifest Huntington Disease Using Vowel Distortion MeasuresabstractHuntington disease (HD) is a fatal autosomal dominant neurocognitive disorder that causes cognitive disturbances, neuropsychiatric symptoms, and impaired motor abilities (e.g., gait, speech, voice). Due to its progressive nature, HD treatment requires ongoing clinical monitoring of symptoms. Individuals with the Huntingtin gene mutation, which causes HD, may exhibit a range of speech symptoms as they progress from premanifest to manifest HD. Speech-based passive monitoring has the potential to augment clinical information by more continuously tracking manifestation symptoms. Differentiating between premanifest and manifest HD is an important yet under-studied problem, as this distinction marks the need for increased treatment. In this work we present the first demonstration of how changes in speech can be measured to differentiate between premanifest and manifest HD. To do so, we focus on one speech symptom of HD: distorted vowels. We introduce a set of Filtered Vowel Distortion Measures (FVDM) which we extract from read speech. We show that FVDM, coupled with features from existing literature, can differentiate between premanifest and manifest HD with 80% accuracy. Amrit Romana, John Bandon, Noelle Carlozzi, Angela Roberts 0001, Emily Mower Provost |
INTERSPEECH | 5 |
| 2020 | MuSE: a Multimodal Dataset of Stressed EmotionabstractEndowing automated agents with the ability to provide support, entertainment and interaction with human beings requires sensing of the users’ affective state. These affective states are impacted by a combination of emotion inducers, current psychological state, and various conversational factors. Although emotion classification in both singular and dyadic settings is an established area, the effects of these additional factors on the production and perception of emotion is understudied. This paper presents a new dataset, Multimodal Stressed Emotion (MuSE), to study the multimodal interplay between the presence of stress and expressions of affect. We describe the data collection protocol, the possible areas of use, and the annotations for the emotional content of the recordings. The paper also presents several baselines to measure the performance of multimodal features for emotion and stress classification. Mimansa Jaiswal, Cristian-Paul Bara, Yuanhang Luo, Mihai Burzo, Rada Mihalcea, Emily Mower Provost |
LREC | 6 |
| 2019 | f-Similarity Preservation Loss for Soft Labels: A Demonstration on Cross-Corpus Speech Emotion RecognitionabstractIn this paper, we propose a Deep Metric Learning (DML) approach that supports soft labels. DML seeks to learn representations that encode the similarity between examples through deep neural networks. DML generally presupposes that data can be divided into discrete classes using hard labels. However, some tasks, such as our exemplary domain of speech emotion recognition (SER), work with inherently subjective data, data for which it may not be possible to identify a single hard label. We propose a family of loss functions, fSimilarity Preservation Loss (f-SPL), based on the dual form of f-divergence for DML with soft labels. We show that the minimizer of f-SPL preserves the pairwise label similarities in the learned feature embeddings. We demonstrate the efficacy of the proposed loss function on the task of cross-corpus SER with soft labels. Our approach, which combines f-SPL and classification loss, significantly outperforms a baseline SER system with the same structure but trained with only classification loss in most experiments. We show that the presented techniques are more robust to over-training and can learn an embedding space in which the similarity between examples is meaningful. Biqiao Zhang, Yuqing Kong, Georg Essl, Emily Mower Provost |
AAAI | 4 |
| 2019 | Muse-ing on the Impact of Utterance Ordering on Crowdsourced Emotion AnnotationsabstractEmotion recognition algorithms rely on data annotated with high quality labels. However, emotion expression and perception are inherently subjective. There is generally not a single annotation that can be unambiguously declared "correct." As a result, annotations are colored by the manner in which they were collected. In this paper, we conduct crowdsourcing experiments to investigate this impact on both the annotations themselves and on the performance of these algorithms. We focus on one critical question: the effect of context. We present a new emotion dataset, Multimodal Stressed Emotion (MuSE), and annotate the dataset using two conditions: randomized, in which annotators are presented with clips in random order, and contextualized, in which annotators are presented with clips in order. We find that contextual labeling schemes result in annotations that are more similar to a speaker's own self-reported labels and that labels generated from randomized schemes are most easily predictable by automated systems. Mimansa Jaiswal, Zakaria Aldeneh, Cristian-Paul Bara, Yuanhang Luo, Mihai Burzo, Rada Mihalcea, Emily Mower Provost |
ICASSP | 7 |
| 2019 | Trainable Time Warping: Aligning Time-series in the Continuous-time DomainabstractDTW calculates the similarity or alignment between two signals, subject to temporal warping. However, its computational complexity grows exponentially with the number of time-series. Although there have been algorithms developed that are linear in the number of time-series, they are generally quadratic in time-series length. The exception is generalized time warping (GTW), which has linear computational cost. Yet, it can only identify simple time warping functions. There is a need for a new fast, high-quality multisequence alignment algorithm. We introduce trainable time warping (TTW), whose complexity is linear in both the number and the length of time-series. TTW performs alignment in the continuoustime domain using a sinc convolutional kernel and a gradient-based optimization technique. We compare TTW and GTW on S5 UCR datasets in time-series averaging and classification. TTW outperforms GTW on 67.1% of the datasets for the averaging tasks, and 61.2% of the datasets for the classification tasks. Soheil Khorram, Melvin G. McInnis, Emily Mower Provost |
ICASSP | 3 |
| 2019 | Exploiting Acoustic and Lexical Properties of Phonemes to Recognize Valence from SpeechabstractEmotions modulate speech acoustics as well as language. The latter influences the sequences of phonemes that are produced, which in turn further modulate the acoustics. Therefore, phonemes impact emotion recognition in two ways: (1) they introduce an additional source of variability in speech signals and (2) they provide information about the emotion expressed in speech content. Previous work in speech emotion recognition has considered (1) or (2), individually. In this paper, we investigate how we can jointly consider both factors to improve the prediction of emotional valence (positive vs. negative), and the relationship between improved prediction and the emotion elicitation process (e.g., fixed script, improvisation, natural interaction). We present a network that exploits both the acoustic and the lexical properties of phonetic information using multi-stage fusion. Our results on the IEMOCAP and MSP-Improv datasets show that our approach outperforms systems that either do not consider the influence of phonetic information or that only consider a single aspect of this influence. Biqiao Zhang, Soheil Khorram, Emily Mower Provost |
ICASSP | 3 |
| 2019 | Controlling for Confounders in Multimodal Emotion Classification via Adversarial LearningabstractVarious psychological factors affect how individuals express emotions. Yet, when we collect data intended for use in building emotion recognition systems, we often try to do so by creating paradigms that are designed just with a focus on eliciting emotional behavior. Algorithms trained with these types of data are unlikely to function outside of controlled environments because our emotions naturally change as a function of these other factors. In this work, we study how the multimodal expressions of emotion change when an individual is under varying levels of stress. We hypothesize that stress produces modulations that can hide the true underlying emotions of individuals and that we can make emotion recognition algorithms more generalizable by controlling for variations in stress. To this end, we use adversarial networks to decorrelate stress modulations from emotion representations. We study how stress alters acoustic and lexical emotional predictions, paying special attention to how modulations due to stress affect the transferability of learned emotion recognition models across domains. Our results show that stress is indeed encoded in trained emotion classifiers and that this encoding varies across levels of emotions and across the lexical and acoustic modalities. Our results also show that emotion recognition models that control for stress during training have better generalizability when applied to new domains, compared to models that do not control for stress during training. We conclude that is is necessary to consider the effect of extraneous psychological factors when building and testing emotion recognition models. Mimansa Jaiswal, Zakaria Aldeneh, Emily Mower Provost |
ICMI | 3 |
| 2019 | Identifying Mood Episodes Using Dialogue Features from Clinical InterviewsabstractBipolar disorder, a severe chronic mental illness characterized by pathological mood swings from depression to mania, requires ongoing symptom severity tracking to both guide and measure treatments that are critical for maintaining long-term health. Mental health professionals assess symptom severity through semi-structured clinical interviews. During these interviews, they observe their patients' spoken behaviors, including both what the patients say and how they say it. In this work, we move beyond acoustic and lexical information, investigating how higher-level interactive patterns also change during mood episodes. We then perform a secondary analysis, asking if these interactive patterns, measured through dialogue features, can be used in conjunction with acoustic features to automatically recognize mood episodes. Our results show that it is beneficial to consider dialogue features when analyzing and building automated systems for predicting and monitoring mood. Zakaria Aldeneh, Mimansa Jaiswal, Michael Picheny, Melvin G. McInnis, Emily Mower Provost |
INTERSPEECH | 5 |
| 2019 | Emotion Recognition from Natural Phone Conversations in Individuals with and without Recent Suicidal Ideation
John Gideon, Heather T. Schatten, Melvin G. McInnis, Emily Mower Provost |
INTERSPEECH | 4 |
| 2019 | Into the Wild: Transitioning from Recognizing Mood in Clinical Interactions to Personal Conversations for Individuals with Bipolar Disorder
Katie Matton, Melvin G. McInnis, Emily Mower Provost |
INTERSPEECH | 3 |
| 2019 | ISLA: Temporal Segmentation and Labeling for Audio-Visual Emotion RecognitionabstractEmotion is an essential part of human interaction. Automatic emotion recognition can greatly benefit human-centered interactive technology, since extracted emotion can be used to understand and respond to user needs. However, real-world emotion recognition faces a central challenge when a user is speaking: facial movements due to speech are often confused with facial movements related to emotion. Recent studies have found that the use of phonetic information can reduce speech-related variability in the lower face region. However, methods to differentiate upper face movements due to emotion and due to speech have been underexplored. This gap leads us to the proposal of the Informed Segmentation and Labeling Approach (ISLA). ISLA uses speech signals that alter the dynamics of the lower and upper face regions. We demonstrate how pitch can be used to improve estimates of emotion from the upper face, and how this estimate can be combined with emotion estimates from the lower face and speech in a multimodal classification system. Our emotion classification results on the IEMOCAP and SAVEE datasets show that ISLA improves overall classification performance. We also demonstrate how emotion estimates from different modalities correlate with each other, providing insights into the differences between posed and spontaneous expressions. Yelin Kim, Emily Mower Provost |
IEEE Trans. Affect. Comput. | 2 |
| 2019 | Cross-Corpus Acoustic Emotion Recognition with Multi-Task Learning: Seeking Common Ground While Preserving DifferencesabstractThere is growing interest in emotion recognition due to its potential in many applications. However, a pervasive challenge is the presence of data variability caused by factors such as differences across corpora, speaker's gender, and the “domain” of expression (e.g., whether the expression is spoken or sung). Prior work has addressed this challenge by combining data across corpora and/or genders, or by explicitly controlling for these factors. In this work, we investigate the influence of corpus, domain, and gender on the cross-corpus generalizability of emotion recognition systems. We use a multi-task learning approach, where we define the tasks according to these factors. We find that incorporating variability caused by corpus, domain, and gender through multi-task learning outperforms approaches that treat the tasks as either identical or independent. Domain is a larger differentiating factor than gender for multi-domain data. When considering only the speech domain, gender and corpus are similarly influential. Defining tasks by gender is more beneficial than by either corpus or corpus and gender for valence, while the opposite holds for activation. On average, cross-corpus performance increases with the number of training corpora. The results demonstrate that effective cross-corpus modeling requires that we understand how emotion expression patterns change as a function of non-emotional factors. Biqiao Zhang, Emily Mower Provost, Georg Essl |
IEEE Trans. Affect. Comput. | 2 |
| 2018 | Improving End-of-Turn Detection in Spoken Dialogues by Detecting Speaker Intentions as a Secondary TaskabstractThis work focuses on the use of acoustic cues for modeling turn-taking in dyadic spoken dialogues. Previous work has shown that speaker intentions (e.g., asking a question, uttering a backchannel, etc.) can influence turn-taking behavior and are good predictors of turn-transitions in spoken dialogues. However, speaker intentions are not readily available for use by automated systems at run-time; making it difficult to use this information to anticipate a turn-transition. To this end, we propose a multi-task neural approach for predicting turn-transitions and speaker intentions simultaneously. Our results show that adding the auxiliary task of speaker intention prediction improves the performance of turn-transition prediction in spoken dialogues, without relying on additional input features during run-time. Zakaria Aldeneh, Dimitrios Dimitriadis, Emily Mower Provost |
ICASSP | 3 |
| 2018 | The PRIORI Emotion Dataset: Linking Mood to Emotion Detected In-the-WildabstractBipolar Disorder is a chronic psychiatric illness characterized by pathological mood swings associated with severe disruptions in emotion regulation. Clinical monitoring of mood is key to the care of these dynamic and incapacitating mood states. Frequent and detailed monitoring improves clinical sensitivity to detect mood state changes, but typically requires costly and limited resources. Speech characteristics change during both depressed and manic states, suggesting automatic methods applied to the speech signal can be effectively used to monitor mood state changes. However, speech is modulated by many factors, which renders mood state prediction challenging. We hypothesize that emotion can be used as an intermediary step to improve mood state prediction. This paper presents critical steps in developing this pipeline, including (1) a new in the wild emotion dataset, the PRIORI Emotion Dataset, collected from everyday smartphone conversational speech recordings, (2) activation/valence emotion recognition baselines on this dataset (PCC of 0.71 and 0.41, respectively), and (3) significant correlation between predicted emotion and mood state for individuals with bipolar disorder. This provides evidence and a working baseline for the use of emotion as a meta-feature for mood state monitoring. Soheil Khorram, Mimansa Jaiswal, John Gideon, Melvin G. McInnis, Emily Mower Provost |
INTERSPEECH | 5 |
| 2018 | Classification of Huntington Disease Using Acoustic and Lexical FeaturesabstractSpeech is a critical biomarker for Huntington Disease (HD), with changes in speech increasing in severity as the disease progresses. Speech analyses are currently conducted using either transcriptions created manually by trained professionals or using global rating scales. Manual transcription is both expensive and time-consuming and global rating scales may lack sufficient sensitivity and fidelity [1]. Ultimately, what is needed is an unobtrusive measure that can cheaply and continuously track disease progression. We present first steps towards the development of such a system, demonstrating the ability to automatically differentiate between healthy controls and individuals with HD using speech cues. The results provide evidence that objective analyses can be used to support clinical diagnoses, moving towards the tracking of symptomatology outside of laboratory and clinical environments. Matthew Perez, Wenyu Jin 0001, Noelle Carlozzi, Praveen Dayalu, Angela Roberts 0001, Emily Mower Provost |
INTERSPEECH | 7 |
| 2018 | Automatic quantitative analysis of spontaneous aphasic speech
Keli Licata, Emily Mower Provost |
Speech Commun. | 3 |
| 2017 | Using regional saliency for speech emotion recognitionabstractIn this paper, we show that convolutional neural networks can be directly applied to temporal low-level acoustic features to identify emotionally salient regions without the need for defining or applying utterance-level statistics. We show how a convolutional neural network can be applied to minimally hand-engineered features to obtain competitive results on the IEMOCAP and MSP-IMPROV datasets. In addition, we demonstrate that, despite their common use across most categories of acoustic features, utterance-level statistics may obfuscate emotional information. Our results suggest that convolutional neural networks with Mel Filterbanks (MFBs) can be used as a replacement for classifiers that rely on features obtained from applying utterance-level statistics. Zakaria Aldeneh, Emily Mower Provost |
ICASSP | 2 |
| 2017 | Pooling acoustic and lexical features for the prediction of valenceabstractIn this paper, we present an analysis of different multimodal fusion approaches in the context of deep learning, focusing on pooling intermediate representations learned for the acoustic and lexical modalities. Traditional approaches to multimodal feature pooling include: concatenation, element-wise addition, and element-wise multiplication. We compare these traditional methods to outer-product and compact bilinear pooling approaches, which consider more comprehensive interactions between features from the two modalities. We also study the influence of each modality on the overall performance of a multimodal system. Our experiments on the IEMOCAP dataset suggest that: (1) multimodal methods that combine acoustic and lexical features outperform their unimodal counterparts; (2) the lexical modality is better for predicting valence than the acoustic modality; (3) outer-product-based pooling strategies outperform other pooling strategies. Zakaria Aldeneh, Soheil Khorram, Dimitrios Dimitriadis, Emily Mower Provost |
ICMI | 4 |
| 2017 | Predicting the distribution of emotion perception: capturing inter-rater variabilityabstractEmotion perception is person-dependent and variable. Dimensional characterizations of emotion can capture this variability by describing emotion in terms of its properties (e.g., valence, positive vs. negative, and activation, calm vs. excited). However, in many emotion recognition systems, this variability is often considered "noise" and is attenuated by averaging across raters. Yet, inter-rater variability provides information about the subtlety or clarity of an emotional expression and can be used to describe complex emotions. In this paper, we investigate methods that can effectively capture the variability across evaluators by predicting emotion perception as a discrete probability distribution in the valence-activation space. We propose: (1) a label processing method that can generate two-dimensional discrete probability distributions of emotion from a limited number of ordinal labels; (2) a new approach that predicts the generated probabilistic distributions using dynamic audio-visual features and Convolutional Neural Networks (CNNs). Our experimental results on the MSP-IMPROV corpus suggest that the proposed approach is more effective than the conventional Support Vector Regressions (SVRs) approach with utterance-level statistical features, and that feature-level fusion of the audio and video modalities outperforms decision-level fusion. The proposed CNN model predominantly improves the prediction accuracy for the valence dimension and brings a consistent performance improvement over data recorded from natural interactions. The results demonstrate the effectiveness of generating emotion distributions from limited number of labels and predicting the distribution using dynamic features and neural networks. Biqiao Zhang, Georg Essl, Emily Mower Provost |
ICMI | 3 |
| 2017 | Progressive Neural Networks for Transfer Learning in Emotion RecognitionabstractMany paralinguistic tasks are closely related and thus representations learned in one domain can be leveraged for another. In this paper, we investigate how knowledge can be transferred between three paralinguistic tasks: speaker, emotion, and gender recognition. Further, we extend this problem to cross-dataset tasks, asking how knowledge captured in one emotion dataset can be transferred to another. We focus on progressive neural networks and compare these networks to the conventional deep learning method of pre-training and fine-tuning. Progressive neural networks provide a way to transfer knowledge and avoid the forgetting effect present when pre-training neural networks on different tasks. Our experiments demonstrate that: (1) emotion recognition can benefit from using representations originally learned for different paralinguistic tasks and (2) transfer learning can effectively leverage additional datasets to improve the performance of emotion recognition systems. John Gideon, Soheil Khorram, Zakaria Aldeneh, Dimitrios Dimitriadis, Emily Mower Provost |
INTERSPEECH | 5 |
| 2017 | Capturing Long-Term Temporal Dependencies with Convolutional Networks for Continuous Emotion RecognitionabstractThe goal of continuous emotion recognition is to assign an emotion value to every frame in a sequence of acoustic features. We show that incorporating long-term temporal dependencies is critical for continuous emotion recognition tasks. To this end, we first investigate architectures that use dilated convolutions. We show that even though such architectures outperform previously reported systems, the output signals produced from such architectures undergo erratic changes between consecutive time steps. This is inconsistent with the slow moving ground-truth emotion labels that are obtained from human annotators. To deal with this problem, we model a downsampled version of the input signal and then generate the output signal through upsampling. Not only does the resulting downsampling/upsampling network achieve good performance, it also generates smooth output trajectories. Our method yields the best known audio-only performance on the RECOLA dataset. Soheil Khorram, Zakaria Aldeneh, Dimitrios Dimitriadis, Melvin G. McInnis, Emily Mower Provost |
INTERSPEECH | 5 |
| 2017 | Discretized Continuous Speech Emotion Recognition with Multi-Task Deep Recurrent Neural Network
Zakaria Aldeneh, Emily Mower Provost |
INTERSPEECH | 3 |
| 2017 | Automatic Paraphasia Detection from Aphasic Speech: A Preliminary Study
Keli Licata, Emily Mower Provost |
INTERSPEECH | 3 |
| 2017 | MSP-IMPROV: An Acted Corpus of Dyadic Interactions to Study Emotion PerceptionabstractWe present the MSP-IMPROV corpus, a multimodal emotional database, where the goal is to have control over lexical content and emotion while also promoting naturalness in the recordings. Studies on emotion perception often require stimuli with fixed lexical content, but that convey different emotions. These stimuli can also serve as an instrument to understand how emotion modulates speech at the phoneme level, in a manner that controls for coarticulation. Such audiovisual data are not easily available from natural recordings. A common solution is to record actors reading sentences that portray different emotions, which may not produce natural behaviors. We propose an alternative approach in which we define hypothetical scenarios for each sentence that are carefully designed to elicit a particular emotion. Two actors improvise these emotion-specific situations, leading them to utter contextualized, non-read renditions of sentences that have fixed lexical content and convey different emotions. We describe the context in which this corpus was recorded, the key features of the corpus, the areas in which this corpus can be useful, and the emotional content of the recordings. The paper also provides the performance for speech and facial emotion classifiers. The analysis brings novel classification evaluations where we study the performance in terms of inter-evaluator agreement and naturalness perception, leveraging the large size of the audiovisual database. Carlos Busso, Srinivas Parthasarathy, Alec Burmania, Mohammed Abdel-Wahab 0001, Najmeh Sadoughi, Emily Mower Provost |
IEEE Trans. Affect. Comput. | 6 |
| 2016 | Mood state prediction from speech of varying acoustic quality for individuals with bipolar disorderabstractSpeech contains patterns that can be altered by the mood of an individual. There is an increasing focus on automated and distributed methods to collect and monitor speech from large groups of patients suffering from mental health disorders. However, as the scope of these collections increases, the variability in the data also increases. This variability is due in part to the range in the quality of the devices, which in turn affects the quality of the recorded data, negatively impacting the accuracy of automatic assessment. It is necessary to mitigate variability effects in order to expand the impact of these technologies. This paper explores speech collected from phone recordings for analysis of mood in individuals with bipolar disorder. Two different phones with varying amounts of clipping, loudness, and noise are employed. We describe methodologies for use during preprocessing, feature extraction, and data modeling to correct these differences and make the devices more comparable. The results demonstrate that these pipeline modifications result in statistically significantly higher performance, which highlights the potential of distributed mental health systems. John Gideon, Emily Mower Provost, Melvin G. McInnis |
ICASSP | 2 |
| 2016 | Cross-corpus acoustic emotion recognition from singing and speaking: A multi-task learning approachabstractEmotion is expressed over both speech and song. Previous works have found that although spoken and sung emotion recognition are different tasks, they are related. Classifiers that explicitly utilize this relatedness can achieve better performance than classifiers that do not. Further, research in speech emotion recognition has demonstrated that emotion is more accurately modeled when gender is taken into account. However, it is not yet clear how domain (speech or song) and gender can be jointly leveraged in emotion recognition systems nor how systems leveraging this information can perform in cross-corpus settings. In this paper, we explore a multi-task emotion recognition framework and compare the performance across different classification models and output selection/fusion methods using cross-corpus evaluation. Our results show the classification accuracy is the highest when information is shared only between closely related tasks and when the output of disparate models are fused. Biqiao Zhang, Emily Mower Provost, Georg Essl |
ICASSP | 2 |
| 2016 | Wild wild emotion: a multimodal ensemble approachabstractAutomatic emotion recognition from audio-visual data is a topic that has been broadly explored using data captured in the laboratory. However, these data are not necessarily representative of how emotion is manifested in the real-world. In this paper, we describe our system for the 2016 Emotion Recognition in the Wild challenge. We use the Acted Facial Expressions in the Wild database 6.0 (AFEW 6.0), which contains short clips of popular TV shows and movies and has more variability in the data compared to laboratory recordings. We explore a set of features that incorporate information from facial expressions and speech, in addition to cues from the background music and overall scene. In particular, we propose the use of a feature set composed of dimensional emotion estimates trained from outside acoustic corpora. We design sets of multiclass and pairwise (one-versus-one) classifiers and fuse the resulting systems. Our fusion increases the performance from a baseline of 38.81% to 43.86% and from 40.47% to 46.88%, for validation and test sets, respectively. While the video features perform better than audio features alone, a combination of the two modalities achieves the greatest performance, with gains of 4.4% and 1.4%, with and without information gain, respectively. Because of the flexible design of the fusion, it is easily adaptable to other multimodal learning problems. John Gideon, Biqiao Zhang, Zakaria Aldeneh, Yelin Kim, Soheil Khorram, Emily Mower Provost |
ICMI | 7 |
| 2016 | Emotion spotting: discovering regions of evidence in audio-visual emotion expressionsabstractResearch has demonstrated that humans require different amounts of information, over time, to accurately perceive emotion expressions. This varies as a function of emotion classes. For example, recognition of happiness requires a longer stimulus than recognition of anger. However, previous automatic emotion recognition systems have often overlooked these differences. In this work, we propose a data-driven framework to explore patterns (timings and durations) of emotion evidence, specific to individual emotion classes. Further, we demonstrate that these patterns vary as a function of which modality (lower face, upper face, or speech) is examined, and consistent patterns emerge across different folds of experiments. We also show similar patterns across emotional corpora (IEMOCAP and MSP-IMPROV). In addition, we show that our proposed method, which uses only a portion of the data (59% for the IEMOCAP), achieves comparable accuracy to a system that uses all of the data within each utterance. Our method has a higher accuracy when compared to a baseline method that randomly chooses a portion of the data. We show that the performance gain of the method is mostly from prototypical emotion expressions (defined as expressions with rater consensus). The innovation in this study comes from its understanding of how multimodal cues reveal emotion over time. Yelin Kim, Emily Mower Provost |
ICMI | 2 |
| 2016 | Automatic recognition of self-reported and perceived emotion: does joint modeling help?abstractEmotion labeling is a central component of automatic emotion recognition. Evaluators are asked to estimate the emotion label given a set of cues, produced either by themselves (self-report label) or others (perceived label). This process is complicated by the mismatch between the intentions of the producer and the interpretation of the perceiver. Traditionally, emotion recognition systems use only one of these types of labels when estimating the emotion content of data. In this paper, we explore the impact of jointly modeling both an individual's self-report and the perceived label of others. We use deep belief networks (DBN) to learn a representative feature space, and model the potentially complementary relationship between intention and perception using multi-task learning. We hypothesize that the use of DBN feature-learning and multi-task learning of self-report and perceived emotion labels will improve the performance of emotion recognition systems. We test this hypothesis on the IEMOCAP dataset, an audio-visual and motion-capture emotion corpus. We show that both DBN feature learning and multi-task learning offer complementary gains. The results demonstrate that the perceived emotion tasks see greatest performance gain for emotionally subtle utterances, while the self-report emotion tasks see greatest performance gain for emotionally clear utterances. Our results suggest that the combination of knowledge from the self-report and perceived emotion labels lead to more effective emotion recognition systems. Biqiao Zhang, Georg Essl, Emily Mower Provost |
ICMI | 3 |
| 2016 | Experiences with Shared Resources for Research and Education in Speech and Language Processingabstract\n Contains fulltext :\n 161889.pdf (Publisher’s version ) (Open Access)\n Rebecca Bates 0001, Eric Fosler-Lussier, Florian Metze, Martha A. Larson, Gina-Anne Levow, Emily Mower Provost |
INTERSPEECH | 6 |
| 2016 | Recognition of Depression in Bipolar Disorder: Leveraging Cohort and Person-Specific Knowledge
Soheil Khorram, John Gideon, Melvin G. McInnis, Emily Mower Provost |
INTERSPEECH | 4 |
| 2016 | Improving Automatic Recognition of Aphasic Speech with AphasiaBank
Emily Mower Provost |
INTERSPEECH | 2 |
| 2016 | Automatic Assessment of Speech Intelligibility for Individuals With AphasiaabstractTraditional in-person therapy may be difficult to access for individuals with aphasia due to the shortage of speech-language pathologists and high treatment cost. Computerized exercises offer a promising low-cost and constantly accessible supplement to in-person therapy. Unfortunately, the lack of feedback for verbal expression in existing programs hinders the applicability and effectiveness of this form of treatment. A prerequisite for producing meaningful feedback is speech intelligibility assessment. In this work, we investigate the feasibility of an automated system to assess three aspects of aphasic speech intelligibility: clarity, fluidity, and prosody. We introduce our aphasic speech corpus, which contains speech-based interaction between individuals with aphasia and a tablet-based application designed for therapeutic purposes. We present our method for eliciting reliable ground-truth labels for speech intelligibility based on the perceptual judgment of nonexpert human evaluators. We describe and analyze our feature set engineered for capturing pronunciation, rhythm, and intonation. We investigate the classification performance of our system under two conditions, one using human-labeled transcripts to drive feature extraction, and another using transcripts generated automatically. We show that some aspects of aphasic speech intelligibility can be estimated at human-level performance. Our results demonstrate the potential for the computerized treatment of aphasia and lay the groundwork for bridging the gap between human and automatic intelligibility assessment. Keli Licata, Carol Persad, Emily Mower Provost |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2015 | Predicting Emotion Perception Across Domains: A Study of Singing and SpeakingabstractEmotion affects our understanding of the opinions and sentiments of others. Research has demonstrated that humans are able to recognize emotions in various domains, including speech and music, and that there are potential shared features that shape the emotion in both domains. In this paper, we investigate acoustic and visual features that are relevant to emotion perception in the domains of singing and speaking. We train regression models using two paradigms: (1) within-domain, in which models are trained and tested on the same domain and (2) cross-domain, in which models are trained on one domain and tested on the other domain. This strategy allows us to analyze the similarities and differences underlying the relationship between audio-visual feature expression and emotion perception and how this relationship is affected by domain of expression. We use kernel density estimation to model emotion as a probability distribution over the perception associated with multiple evaluators on the valence-activation space. This allows us to model the variation inherent in the reported perception. Results suggest that activation can be modeled more accurately across domains, compared to valence. Furthermore, visual features capture cross-domain emotion more accurately than acoustic features. The results provide additional evidence for a shared mechanism underlying spoken and sung emotion perception. Biqiao Zhang, Emily Mower Provost, Robert Swedberg, Georg Essl |
AAAI | 2 |
| 2015 | Leveraging inter-rater agreement for audio-visual emotion recognitionabstractHuman expressions are often ambiguous and unclear, resulting in disagreement or confusion among different human evaluators. In this paper, we investigate how audiovisual emotion recognition systems can leverage prototypicality, the level of agreement or confusion among human evaluators. We propose the use of a weighted Support Vector Machine to explicitly model the relationship between the prototypicality of training instances and evaluated emotion from the IEMOCAP corpus. We choose weights of prototypical and non-prototypical instances based on the maximal accuracy of each speaker. We then provide per-speaker analysis to understand specific speech characteristics associated with the information gain of emotion given prototypicality information. Our experimental results show that neutrality, one of the most challenging emotion to recognize, has the highest performance gain from prototypicality information, compared to other emotion classes: Angry, Happy, and Sad. We also show that the proposed method improves the overall multi-class classification accuracy significantly over traditional methods that do not leverage prototypicality. Yelin Kim, Emily Mower Provost |
ACII | 2 |
| 2015 | Data selection for acoustic emotion recognition: Analyzing and comparing utterance and sub-utterance selection strategiesabstractData selection is an important component of cross-corpus training and semi-supervised/active learning. However, its effect on acoustic emotion recognition is still not well understood. In this work, we perform an in-depth exploration of various data selection strategies for emotion classification from speech using classifier agreement as the selection metric. Our methods span both the traditional utterance as well as the less explored sub-utterance level. A median unweighted average recall of 70.68%, comparable to the winner of the 2009 INTERSPEECH Emotion Challenge, was achieved on the FAU Aibo 2-class problem using less than 50% of the training data. Our results indicate that sub-utterance selection leads to slightly faster convergence and significantly more stable learning. In addition, diversifying instances in terms of classifier agreement produces a faster learning rate, whereas selecting those near the median results in higher stability. We show that the selected data instances can be explained intuitively based on their acoustic properties and position within an utterance. Our work helps provide a deeper understanding of the strengths, weaknesses, and trade-offs of different data selection strategies for speech emotion recognition. Emily Mower Provost |
ACII | 2 |
| 2015 | EmoShapelets: Capturing local dynamics of audio-visual affective speechabstractAutomatic recognition of emotion in speech is an active area of research. One of the important open challenges relates to how the emotional characteristics of speech change in time. Past research has demonstrated the importance of capturing global dynamics (across an entire utterance) and local dynamics (within segments of an utterance). In this paper, we propose a novel concept, EmoShapelets, to capture the local dynamics in speech. EmoShapelets capture changes in emotion that occur within utterances. We propose a framework to generate, update, and select EmoShapelets. We also demonstrate the discriminative power of EmoShapelets by using them with various classifiers to achieve comparable results with the state-of-the-art systems on the IEMOCAP dataset. EmoShapelets can serve as basic units of emotion expression and provide additional evidence supporting the existence of local patterns of emotion underlying human communication. Yuan Shangguan, Emily Mower Provost |
ACII | 2 |
| 2015 | Recognizing emotion from singing and speaking using shared modelsabstractSpeech and song are two types of vocal communications that are closely related to each other. While significant progress has been made in both speech and music emotion recognition, few works have concentrated on building a shared emotion recognition model for both speech and song. In this paper, we propose three shared emotion recognition models for speech and song: a simple model, a single-task hierarchical model, and a multi-task hierarchical model. We study the commonalities and differences present in emotion expression across these two communication domains. We compare the performance across different settings, investigate the relationship between evaluator agreement rate and classification accuracy, and analyze the classification performance of individual feature groups. Our results show that the multi-task model classifies emotion more accurately compared to single-task models when the same set of features is used. This suggests that although spoken and sung emotion recognition tasks are different, they are related, and can be considered together. The results demonstrate that utterances with lower agreement rate and emotions with low activation benefit the most from multi-task learning. Visual features appear to be more similar across spoken and sung emotion expression, compared to acoustic features. Biqiao Zhang, Georg Essl, Emily Mower Provost |
ACII | 3 |
| 2015 | UMEME: University of Michigan Emotional McGurk Effect Data SetabstractEmotion is central to communication; it colors our interpretation of events and social interactions. Emotion expression is generally multimodal, modulating our facial movement, vocal behavior, and body gestures. The method through which this multimodal information is integrated and perceived is not well understood. This knowledge has implications for the design of multimodal classification algorithms, affective interfaces, and even mental health assessment. We present a novel data set designed to support research into the emotion perception process, the University of Michigan Emotional McGurk Effect Data set (UMEME). UMEME has a critical feature that differentiates it from currently existing data sets; it contains not only emotionally congruent stimuli (emotionally matched faces and voices), but also emotionally incongruent stimuli (emotionally mismatched faces and voices). The inclusion of emotionally complex and dynamic stimuli provides an opportunity to study how individuals make assessments of emotion content in the presence of emotional incongruence, or emotional noise. We describe the collection, annotation, and statistical properties of the data and present evidence illustrating how audio and video interact to result in specific types of emotion perception. The results demonstrate that there exist consistent patterns underlying emotion evaluation, even given incongruence, positioning UMEME as an important new tool for understanding emotion perception. Emily Mower Provost, Yuan Shangguan, Carlos Busso |
IEEE Trans. Affect. Comput. | 1 |
| 2015 | Emotion Recognition During Speech Using Dynamics of Multiple Regions of the FaceabstractThe need for human-centered, affective multimedia interfaces has motivated research in automatic emotion recognition. In this article, we focus on facial emotion recognition. Specifically, we target a domain in which speakers produce emotional facial expressions while speaking. The main challenge of this domain is the presence of modulations due to both emotion and speech. For example, an individual's mouth movement may be similar when he smiles and when he pronounces the phoneme /IY/, as in “cheese”. The result of this confusion is a decrease in performance of facial emotion recognition systems. In our previous work, we investigated the joint effects of emotion and speech on facial movement. We found that it is critical to employ proper temporal segmentation and to leverage knowledge of spoken content to improve classification performance. In the current work, we investigate the temporal characteristics of specific regions of the face, such as the forehead, eyebrow, cheek, and mouth. We present methodology that uses the temporal patterns of specific regions of the face in the context of a facial emotion recognition system. We test our proposed approaches on two emotion datasets, the IEMOCAP and SAVEE datasets. Our results demonstrate that the combination of emotion recognition systems based on different facial regions improves overall accuracy compared to systems that do not leverage different characteristics of individual regions. Yelin Kim, Emily Mower Provost |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2014 | Ecologically valid long-term mood monitoring of individuals with bipolar disorder using speechabstractSpeech patterns are modulated by the emotional and neurophysiological state of the speaker. There exists a growing body of work that computationally examines this modulation in patients suffering from depression, autism, and post-traumatic stress disorder. However, the majority of the work in this area focuses on the analysis of structured speech collected in controlled environments. Here we expand on the existing literature by examining bipolar disorder (BP). BP is characterized by mood transitions, varying from a healthy euthymic state to states characterized by mania or depression. The speech patterns associated with these mood states provide a unique opportunity to study the modulations characteristic of mood variation. We describe methodology to collect unstructured speech continuously and unobtrusively via the recording of day-to-day cellular phone conversations. Our pilot investigation suggests that manic and depressive mood states can be recognized from this speech data, providing new insight into the feasibility of unobtrusive, unstructured, and continuous speech-based wellness monitoring for individuals with BP. Zahi N. Karam, Emily Mower Provost, Satinder Singh 0001, Jennifer Montgomery, Christopher Archer, Gloria Harrington, Melvin G. McInnis |
ICASSP | 2 |
| 2014 | Automatic analysis of speech quality for aphasia treatmentabstractAphasia is a common language disorder which can severely affect an individual's ability to communicate with others. Aphasia rehabilitation requires intensive practice accompanied by appropriate feedback, the latter of which is difficult to satisfy outside of therapy. In this paper we take a first step towards developing an intelligent system capable of providing feedback to patients with aphasia through the automation of two typical therapeutic exercises, sentence building and picture description. We describe the natural speech corpus collected from our interaction with clients in the University of Michigan Aphasia Program (UMAP). We develop classifiers to automatically estimate speech quality based on human perceptual judgment. Our automatic prediction yields accuracies comparable to the average human evaluator. Our feature selection process gives insights into the factors that influence human evaluation. The results presented in this work provide support for the feasibility of this type of system. Keli Licata, Elizabeth Mercado, Carol Persad, Emily Mower Provost |
ICASSP | 5 |
| 2014 | Modeling pronunciation, rhythm, and intonation for automatic assessment of speech quality in aphasia rehabilitationabstractPatients with aphasia often have impaired speech-language pro-duction skills, resulting in tremendous difficulties in tasks that require verbal communication. To facilitate rehabilitation out-side of therapy, we are collaborating with the University of Michigan Aphasia Program (UMAP) to develop an automated system capable of providing feedback regarding the patient’s verbal output. In this paper we introduce a robust method for extracting rhythm and intonation features from aphasic speech based on template matching. These features, combined with Goodness of Pronunciation (GOP) scores and our previous fea-ture set, help our system achieve human-level performance in classifying the quality of speech produced by patients attend-ing UMAP. The results presented in this work demonstrate the efficacy of our technique and the potential of this system for handling natural speech data recorded in non-ideal conditions as well as the unpredictability in aphasic speech patterns. Index Terms: aphasia, speech-language disorder, clinical ap-plication, speech quality analysis, machine learning Emily Mower Provost |
INTERSPEECH | 2 |
| 2014 | Say Cheese vs. Smile: Reducing Speech-Related Variability for Facial Emotion RecognitionabstractFacial movement is modulated both by emotion and speech articulation. Facial emotion recognition systems aim to discriminate between emotions, while reducing the speech-related variability in facial cues. This aim is often achieved using two key features: (1) phoneme segmentation: facial cues are temporally divided into units with a single phoneme and (2) phoneme-specific classification: systems learn patterns associated with groups of visually similar phonemes (visemes), e.g. P, B, and M. In this work, we empirically compare the effects of different temporal segmentation and classification schemes for facial emotion recognition. We propose an unsupervised segmentation method that does not necessitate costly phonetic transcripts. We show that the proposed method bridges the accuracy gap between a traditional sliding window method and phoneme segmentation, achieving a statistically significant performance gain. We also demonstrate that the segments derived from the proposed unsupervised and phoneme segmentation strategies are similar to each other. This paper provides new insight into unsupervised facial motion segmentation and the impact of speech variability on emotion classification. Yelin Kim, Emily Mower Provost |
ACM Multimedia | 2 |
| 2013 | Emotion recognition from spontaneous speech using Hidden Markov models with deep belief networksabstractResearch in emotion recognition seeks to develop insights into the temporal properties of emotion. However, automatic emotion recognition from spontaneous speech is challenging due to non-ideal recording conditions and highly ambiguous ground truth labels. Further, emotion recognition systems typically work with noisy high-dimensional data, rendering it difficult to find representative features and train an effective classifier. We tackle this problem by using Deep Belief Networks, which can model complex and non-linear high-level relationships between low-level features. We propose and evaluate a suite of hybrid classifiers based on Hidden Markov Models and Deep Belief Networks. We achieve state-of-the-art results on FAU Aibo, a benchmark dataset in emotion recognition [1]. Our work provides insights into important similarities and differences between speech and emotion. Emily Mower Provost |
ASRU | 2 |
| 2013 | Deep learning for robust feature generation in audiovisual emotion recognitionabstractAutomatic emotion recognition systems predict high-level affective content from low-level human-centered signal cues. These systems have seen great improvements in classification accuracy, due in part to advances in feature selection methods. However, many of these feature selection methods capture only linear relationships between features or alternatively require the use of labeled data. In this paper we focus on deep learning techniques, which can overcome these limitations by explicitly capturing complex non-linear feature interactions in multimodal data. We propose and evaluate a suite of Deep Belief Network models, and demonstrate that these models show improvement in emotion classification performance over baselines that do not employ deep learning. This suggests that the learned high-order non-linear relationships are effective for emotion recognition. Yelin Kim, Honglak Lee, Emily Mower Provost |
ICASSP | 3 |
| 2013 | Emotion classification via utterance-level dynamics: A pattern-based approach to characterizing affective expressionsabstractHuman emotion changes continuously and sequentially. This results in dynamics intrinsic to affective communication. One of the goals of automatic emotion recognition research is to computationally represent and analyze these dynamic patterns. In this work, we focus on the global utterance-level dynamics. We are motivated by the hypothesis that global dynamics have emotion-specific variations that can be used to differentiate between emotion classes. Consequently, classification systems that focus on these patterns will be able to make accurate emotional assessments. We quantitatively represent emotion flow within an utterance by estimating short-time affective characteristics. We compare time-series estimates of these characteristics using Dynamic Time Warping, a time-series similarity measure. We demonstrate that this similarity can effectively recognize the affective label of the utterance. The similarity-based pattern modeling outperforms both a feature-based baseline and static modeling. It also provides insight into typical high-level patterns of emotion. We visualize these dynamic patterns and the similarities between the patterns to gain insight into the nature of emotion expression. Yelin Kim, Emily Mower Provost |
ICASSP | 2 |
| 2013 | Identifying salient sub-utterance emotion dynamics using flexible units and estimates of affective flowabstractEmotion recognition is the process of identifying the affective characteristics of an utterance given either static or dynamic descriptions of its signal content. This requires the use of units, windows over which the emotion variation is quantified. However, the appropriate time scale for these units is still an open question. Traditionally, emotion recognition systems have relied upon units of fixed length, whose variation is then modeled over time. This paper takes the view that emotion is expressed over units of variable length. In this paper, variable-length units are introduced and used to capture the local dynamics of emotion at the sub-utterance scale. The results demonstrate that subsets of these local dynamics are salient with respect to emotion class. These salient units provide insight into the natural variation in emotional speech and can be used in a classification framework to achieve performance comparable to the state-of-the-art. This hints at the existence of building blocks that may underlie natural human emotional communication. Emily Mower Provost |
ICASSP | 1 |
| 2013 | Using emotional noise to uncloud audio-visual emotion perceptual evaluationabstractEmotion perception underlies communication and social interaction, shaping how we interpret our world. However, there are many aspects of this process that we still do not fully understand. Notably, we have not yet identified how audio and video information are integrated during the perception of emotion. In this work we present an approach to enhance our understanding of this process using the McGurk effect paradigm, a framework in which stimuli composed of mismatched audio and video cues are presented to human evaluators. Our stimuli set contain sentence-level emotional stimuli with either the same emotion on each channel (“matched”) or different emotions on each channel (“mismatched”, for example, an angry face with a happy voice). We obtain dimensional evaluations (valence and activation) of these emotionally consistent and noisy stimuli using crowd sourcing via Amazon Mechanical Turk. We use these data to investigate the audio-visual feature bias that underlies the evaluation process. We demonstrate that both audio and video information individually contribute to the perception of these dimensional properties. We further demonstrate that the change in perception from the emotionally matched to emotionally mismatched stimuli can be modeled using only unimodal feature variation. These results provide insight into the nature of audio-visual feature integration in emotion perception. Emily Mower Provost, Irene Zhu, Shri Narayanan |
ICME | 1 |
| 2013 | Analyzing the structure of parent-moderated narratives from children with ASD using an entity-based approachabstractStorytelling is a commonly used technique for rating linguistic and communicative abilities of children with Autism Spectrum Disorders (ASD). It highlights their language use beyond sentence-level production, and their ability to cohesively link events into a plot, including incorporating social context. A key scenario of interest we consider is spoken narrative creation in interactive settings, where confederates such as parents can offer scaffolding to their children’s narratives by eliciting answers with appropriate questions, shaping the structure of the resulting narrative. We analyze the structure of children’s stories narrated with the help of their parents using entity-based feature-level patterns in order to see how there are influenced by the parents’ narrative elicitation techniques. The frequency distribution and evolution of entities -meaning the co-referent people, objects and ideas- can capture the main axis of the story plot. Our results indicate that the type of questions the parents ask can be reflected in the entity-based features of a narrative, affecting its underlying structure and coherence. Index Terms: Narrative Structure, Coherence, Text Entities, Autism Spectrum Disorders Theodora Chaspari, Emily Mower Provost, Shri Narayanan |
INTERSPEECH | 2 |
| 2012 | An acoustic analysis of shared enjoyment in ECA interactions of children with autismabstractThe quality of shared enjoyment in interactions is a key aspect related to Autism Spectrum Disorders (ASD). This paper discusses two types of enjoyment: the first refers to humorous events and is associated with one's positive affective state and the second is used to facilitate social interactions between people. These types of shared enjoyment are objectively specified by their proximity to a voiced and unvoiced laughter instance, respectively. The goal of this work is to study the acoustic differences of areas surrounding the two kinds of shared enjoyment instances, called “social zones”, using data collected from children with autism, and their parents, interacting with an Embodied Conversational Agent (ECA). A classification task was performed to predict whether a “social zone” surrounds a voiced or an unvoiced laughter instance. Our results indicate that humorous events are more easily recognized than events acting as social facilitators and that related speech patterns vary more across children compared to other interlocutors. Theodora Chaspari, Emily Mower Provost, Athanasios Katsamanis, Shri Narayanan |
ICASSP | 2 |
| 2011 | A hierarchical static-dynamic framework for emotion classificationabstractThe goal of emotion classification is to estimate an emotion label, given representative data and discriminative features. Humans are very good at deriving high-level representations of emotion state and integrating this information over time to arrive at a final judgment. However, currently, most emotion classification algorithms do not use this technique. This paper presents a hierarchical static dynamic emotion classification framework that estimates high-level emotional judgments and locally integrates this information over time to arrive at a final estimate of the affective label. The results suggest that this framework for emotion classification leads to more accurate results than either purely static or purely dynamic strategies. Emily Mower Provost, Shri Narayanan |
ICASSP | 1 |
| 2011 | Rachel: Design of an emotionally targeted interactive agent for children with autismabstractIncreasingly, multimodal human-computer interactive tools are leveraged in both autism research and therapies. Embodied conversational agents (ECAs) are employed to facilitate the collection of socio-emotional interactive data from children with autism. In this paper we present an overview of the Rachel system developed at the University of Southern California. The Rachel ECA is designed to elicit and analyze complex, structured, and naturalistic interactions and to encourage affective and social behavior. The pilot studies suggest that this tool can be used to effectively elicit social conversational behavior. This paper presents a description of the multimodal human-computer interaction system and an overview of the collected data. Future work includes utilizing signal processing techniques to provide a quantitative description of the interaction patterns. Emily Mower Provost, Matthew Black, Elisa Flores, Marian E. Williams, Shri Narayanan |
ICME | 1 |
| 2011 | Analyzing the Nature of ECA Interactions in Children with AutismabstractEmbodied conversational agents (ECA) offer platforms for the collection of structured interaction and communication data. This paper discusses the data collected from the Rachel system, an ECA developed at the University of Southern California, for interactions with children with autism. Two dyads each com-posed of a child with autism and his parent participated in an experiment with two modes: interactions with and without the ECA present. The goal of this work is to assess the naturalness of the data recorded in the ECA interaction. This analysis was carried out using a classification framework with a prediction variable of the presence or absence of the ECA in the inter-action. The results demonstrate that it is possible to estimate whether or not a parent is interacting with the ECA using their speech data. However, it is not generally possible to do so for the child suggesting that the Rachel system is eliciting commu-nication data that is similar to that elicited through interactions between the child and his parent. Index Terms: Embodied conversational agent, multimodal in-terface, audio-video recording, autism, children’s speech Emily Mower Provost, Chi-Chun Lee, James Gibson, Theodora Chaspari, Marian E. Williams, Shri Narayanan |
INTERSPEECH | 1 |
| 2011 | Emotion recognition using a hierarchical binary decision tree approach
Chi-Chun Lee, Emily Mower Provost, Carlos Busso, Sungbok Lee, Shri Narayanan |
Speech Commun. | 2 |
| 2011 | A Framework for Automatic Human Emotion Classification Using Emotion ProfilesabstractAutomatic recognition of emotion is becoming an increasingly important component in the design process for affect-sensitive human-machine interaction (HMI) systems. Well-designed emotion recognition systems have the potential to augment HMI systems by providing additional user state details and by informing the design of emotionally relevant and emotionally targeted synthetic behavior. This paper describes an emotion classification paradigm, based on emotion profiles (EPs). This paradigm is an approach to interpret the emotional content of naturalistic human expression by providing multiple probabilistic class labels, rather than a single hard label. EPs provide an assessment of the emotion content of an utterance in terms of a set of simple categorical emotions: anger; happiness; neutrality; and sadness. This method can accurately capture the general emotional label (attaining an accuracy of 68.2% in our experiment on the IEMOCAP data) in addition to identifying underlying emotional properties of highly emotionally ambiguous utterances. This capability is beneficial when dealing with naturalistic human emotional expressions, which are often not well described by a single semantic label. Emily Mower Provost, Maja J. Mataric, Shri Narayanan |
IEEE Trans. Speech Audio Process. | 1 |
| 2010 | Speech emotion estimation in 3D spaceabstractSpeech processing is an important element of affective computing. Most research in this direction has focused on classifying emotions into a small number of categories. However, numerical representations of emotions in a multi-dimensional space can be more appropriate to reflect the gradient nature of emotion expressions, and can be more convenient in the sense of dealing with a small set of emotion primitives. This paper presents three approaches (robust regression, support vector regression, and locally linear reconstruction) for emotion primitives estimation in 3D space (valence/activation/dominance), and two approaches (average fusion and locally weighted fusion) to fuse the three elementary estimators for better overall recognition accuracy. The three elementary estimators are diverse and complementary because they cover both linear and nonlinear models, and both global and local models. These five approaches are compared with the state-of-the-art estimator on the same spontaneously elicited emotion dataset. Our results show that all of our three elementary estimators are suitable for speech emotion estimation. Moreover, it is possible to boost the estimation performance by fusing them properly since they appear to leverage complementary speech features. Dongrui Wu, Thomas D. Parsons, Emily Mower Provost, Shri Narayanan |
ICME | 3 |
| 2010 | A cluster-profile representation of emotion using agglomerative hierarchical clusteringabstractThe proper representation of emotion is critical to automatic classification systems. In previous research, we demonstrated that emotion profile (EP) based representations are effective for this task. In EP-based representations, emotions are expressed in terms of underlying affective components from the subset of anger, happiness, neutrality, and sadness. The current study explores cluster profiles (CP), an alternate profile representa-tion in which the components are no longer semantic labels, but clusters inherent in the feature space. This unsupervised clus-tering of the feature space permits the application of a system-level semi-supervised learning paradigm. The results demon-strate that CPs are similarly discriminative to EPs (EP classifica-tion accuracy: 68.37 % vs. 69.25 % for the CP-based classifica-tion). This suggests that exhaustive labeling of a representative training corpus may not be necessary for emotion classification tasks. Emily Mower Provost, Kyu Jeong Han, Sungbok Lee, Shri Narayanan |
INTERSPEECH | 1 |
| 2010 | Robust representations for out-of-domain emotions using Emotion ProfilesabstractThe proper representation of emotion is of vital importance for human-machine interaction. A correct understanding of emotion would allow interactive technology to appropriately respond and adapt to users. In human-machine interaction scenarios it is likely that over the course of an interaction, the human interaction partner will express an emotion not seen during the training of the machine's emotion models. It is therefore crucial to prepare for such eventualities by developing robust representations of emotion that can distinctly represent emotions regardless of whether the data were seen during training of the representation. This novel work demonstrates that an Emotion Profile (EP) representation introduced in [1], a representation composed of the confidences of four binary emotion-specific classifiers, can distinctly represent emotions unseen during training. The classification accuracy increases by only 0.35% over the full dataset when the data excluded from the EP training is included. The results demonstrate that EPs are a robust method for emotion representation. Emily Mower Provost, Maja J. Mataric, Shri Narayanan |
SLT | 1 |
| 2009 | Emotion recognition using a hierarchical binary decision tree approachabstractAutomated emotion state tracking is a crucial element in the computational study of human communication behaviors. It is important to design robust and reliable emotion recognition systems that are suitable for real-world applications both to enhance analytical abilities to support human decision making and to design human-machine interfaces that facilitate efficient communication. We introduce a hierarchical computational structure to recognize emotions. The proposed structure maps an input speech utterance into one of the multiple emotion classes through subsequent layers of binary classifications. The key idea is that the levels in the tree are designed to solve the easiest classification tasks first, allowing us to mitigate error propagation. We evaluated the classification framework on two different emotional databases using acoustic features, the AIBO database and the USC IEMOCAP database. In the case of the AIBO database, we obtain a balanced recall on each of the individual emotion classes using this hierarchical structure. The performance measure of the average unweighted recall on the evaluation data set improves by 3.37% absolute (8.82% relative) over a Support Vector Machine baseline model. In the USC IEMOCAP database, we obtain an absolute improvement of 7.44% (14.58%) over a baseline Support Vector Machine modeling. The results demonstrate that the presented hierarchical approach is effective for classifying emotional utterances in multiple database contexts. Chi-Chun Lee, Emily Mower Provost, Carlos Busso, Sungbok Lee, Shri Narayanan |
INTERSPEECH | 2 |
| 2009 | Evaluating evaluators: a case study in understanding the benefits and pitfalls of multi-evaluator modelingabstractEmotion perception is a complex process, often measured using stimuli presentation experiments that query evaluators for their perceptual ratings of emotional cues. These evaluations contain large amounts of variability both related and unrelated to the evaluated utterances. One approach to handling this variability is to model emotion perception at the individual level. However, the perceptions of specific users may not adequately capture the emotional acoustic properties of an utterance. This problem can be mitigated by the common technique of averaging evalu-ations from multiple users. We demonstrate that this averaging procedure improves classification performance when compared to classification results from models created using individual-specific evaluations. We also demonstrate that the performance increases are related to the consistency with which evaluators label data. These results suggest that the acoustic properties of emotional speech are better captured using models formed from averaged evaluations rather than from individual-specific eval-uations. Emily Mower Provost, Maja J. Mataric, Shri Narayanan |
INTERSPEECH | 1 |
| 2009 | Human Perception of Audio-Visual Synthetic Character Emotion Expression in the Presence of Ambiguous and Conflicting InformationabstractComputer simulated avatars and humanoid robots have an increasingly prominent place in today's world. Acceptance of these synthetic characters depends on their ability to properly and recognizably convey basic emotion states to a user population. This study presents an analysis of the interaction between emotional audio (human voice) and video (simple animation) cues. The emotional relevance of the channels is analyzed with respect to their effect on human perception and through the study of the extracted audio-visual features that contribute most prominently to human perception. As a result of the unequal level of expressivity across the two channels, the audio was shown to bias the perception of the evaluators. However, even in the presence of a strong audio bias, the video data were shown to affect human perception. The feature sets extracted from emotionally matched audio-visual displays contained both audio and video features while feature sets resulting from emotionally mismatched audio-visual displays contained only audio information. This result indicates that observers integrate natural audio cues and synthetic video cues only when the information expressed is in congruence. It is therefore important to properly design the presentation of audio-visual cues as incorrect design may cause observers to ignore the information conveyed in one of the channels. Emily Mower Provost, Maja J. Mataric, Shri Narayanan |
IEEE Trans. Multim. | 1 |
| 2008 | Human perception of synthetic character emotions in the presence of conflicting and congruent vocal and facial expressionsabstractAudio-visual emotion expression by synthetic agents is widely employed in research, industrial, and commercial applications. However, the mechanism through which people judge the multimodal emotional display of these agents is not yet well understood. This study is an attempt to provide a better understanding of the interaction between video and audio channels through the use of a continuous dimensional evaluation framework of valence, activation, and dominance. The results indicate that the congruent audio-visual presentation contains information allowing users to differentiate between happy and angry emotional expressions to a greater degree than either of the two channels individually. Interestingly, however, sad and neutral emotions which exhibit a lesser degree of activation show more confusion when presented using both channels. Furthermore, when faced with a conflicting emotional presentation, users predominantly attended to the vocal channel. It is speculated that this is most likely due to the limited level of facial emotion expression inherent in the current animated face. The results also indicate that there is no clear integration of audio and visual channels in emotion perception as in speech perception indicated by the McGurk effect. The final judgments were biased toward the modality with stronger expression power. Emily Mower Provost, Sungbok Lee, Maja J. Mataric, Shri Narayanan |
ICASSP | 1 |
| 2008 | Joint-processing of audio-visual signals in human perception of conflicting synthetic character emotionsabstractExpressive audio-visual synthetic characters are increasingly employed in research and commercial applications. However, the mechanism that people employ to interpret conflicting or uncertain multimodal emotional displays of these agents is not yet well understood. This study is an attempt to provide a better understanding of the interpretation of conflicting expressive displays in video and audio channels through the use of a continuous dimensional evaluation framework of emotional valence, activation, and dominance. The results indicate that when two conflicting emotions are presented to subjects using audio and video channels, the means of the dimensional evaluations of the resulting emotional judgments by the subjects is located in between the audio-only and video-only emotion perceptual centers. Furthermore, the deviation from the audio-only center is proportional to the distance between the audio and video centers. This indicates that the perceptual judgment of conflicting emotions involves the joint processing of both the audio and the video information irrespective of the perceptual bias toward the audio channel. In general the amount of interaction between audio and video channel seems proportional to the emotional disparity of the two channels in the continuous emotional space considered in this study. Emily Mower Provost, Sungbok Lee, Maja J. Mataric, Shri Narayanan |
ICME | 1 |
| 2008 | Selection of Emotionally Salient Audio-Visual Features for Modeling Human Evaluations of Synthetic Character Emotion DisplaysabstractComputer simulated avatars and humanoid robots have an increasingly prominent place in today's world. Acceptance of these synthetic characters depends on their ability to properly and recognizably convey basic emotion states to a user population. This study presents an analysis of audio-visual features that can be used to predict user evaluations of synthetic character emotion displays. These features include prosodic, spectral, and semantic properties of audio signals in addition to FACS-inspired video features. The goal of this paper is to identify the audio-visual features that explain the variance in the emotional evaluations of naive listeners through the utilization of information gain feature selection in conjunction with support vector machines. These results suggest that there exists an emotionally salient subset of the audio-visual feature space. The features that contribute most to the explanation of evaluator variance are the prior knowledge audio statistics (e.g., average valence rating), the high energy band spectral components, and the quartile pitch range. This feature subset should be correctly modeled and implemented in the design of synthetic expressive displays to convey the desired emotions. Emily Mower Provost, Maja J. Mataric, Shri Narayanan |
ISM | 1 |
| 2007 | Investigating Implicit Cues for User State Estimation in Human-Robot Interaction Using Physiological MeasurementsabstractAchieving and maintaining user engagement is a key goal of human-robot interaction. This paper presents a method for determining user engagement state from physiological data (including galvanic skin response and skin temperature). In the reported study, physiological data were measured while participants played a wire puzzle game moderated by either a simulated or embodied robot, both with varying personalities. The resulting physiological data were segmented and classified based on position within trial using the K-Nearest Neighbors algorithm. We found it was possible to estimate the user's engagement state for trials of variable length with an accuracy of 84.73%. In future experiments, this ability would allow assistive robot moderators to estimate the user's likelihood of ending an interaction at any given point during the interaction. This knowledge could then be used to adapt the behavior of the robot in an attempt to re-engage the user. Emily Mower Provost, David Feil-Seifer, Maja J. Mataric, Shri Narayanan |
RO-MAN | 1 |
| 2007 | Primitives-based evaluation and estimation of emotions in speech
Michael Grimm, Kristian Kroschel, Emily Mower Provost, Shri Narayanan |
Speech Commun. | 3 |