Carlos Busso

dblp:87/5752 · DBLP profile ↗
← Back
196ranked-venue papers
15as first author
84since 2021 · last 2026
0000-0002-4075-4072ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 119 · 9 first-author · 47 since 2021Artificial intelligence and machine learning · 110 · 9 first-author · 52 since 2021Human-computer interaction and ubiquitous computing · 22 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021
YearPublicationVenuePosition
2026 RankList - a Listwise Preference Learning Framework for Predicting Subjective Preferences
abstract
Preference learning has gained significant attention in tasks involving subjective human judgments, such as speech emotion recognition (SER) and image aesthetic assessment. While pairwise frameworks such as RankNet offer robust modeling of relative preferences, they are inherently limited to local comparisons and struggle to capture global ranking consistency. To address these limitations, we propose RankList, a novel listwise preference learning framework that generalizes RankNet to structured list-level supervision. Our formulation explicitly models local and non-local ranking constraints within a probabilistic framework. The paper introduces a log-sum-exp approximation to improve training efficiency. We further extend RankList with skip-wise comparisons, enabling progressive exposure to complex list structures and enhancing global ranking fidelity. Extensive experiments demonstrate the superiority of our method across diverse modalities. On benchmark SER datasets (MSP-Podcast, IEMOCAP, BIIC Podcast), RankList achieves consistent improvements in Kendall's Tau and ranking accuracy compared to standard listwise baselines. We also validate our approach on aesthetic image ranking using the Artistic Image Aesthetics dataset, highlighting its broad applicability. Through ablation and cross-domain studies, we show that RankList not only improves in-domain ranking but also generalizes better across datasets. Our framework offers a unified, extensible approach for modeling ordered preferences in subjective learning scenarios.
Abinay Reddy Naini, Carlos Busso
AAAI3
2026 FedMLAC: Mutual learning driven heterogeneous federated audio classification
abstract
Federated Learning (FL) offers a privacy-preserving framework for training audio classification (AC) models across decentralized clients without sharing raw data. However, Federated Audio Classification faces three major challenges: data heterogeneity , model heterogeneity , and data corruption , which degrade performance in real-world settings. While existing methods often address these issues separately, a unified solution remains underexplored. We propose FedMLAC, a mutual learning-based FL framework that tackles all three challenges simultaneously. Each client maintains a personalized local AC model and a lightweight, globally shared Plug-in model. These models interact via bidirectional knowledge distillation, enabling global knowledge sharing while adapting to local data distributions, thus addressing both data and model heterogeneity. To counter data corruption, we introduce a Layer-wise Pruning Aggregation (LPA) strategy that filters anomalous Plug-in updates based on parameter deviations during aggregation. Extensive experiments on four diverse AC benchmarks, including both speech and non-speech tasks, show that FedMLAC consistently outperforms state-of-the-art baselines in classification accuracy and robustness to noisy data.
Rajib Rana, Di Wu 0050, Youyang Qu, Xiaohui Tao 0001, Ji Zhang 0001, Carlos Busso, Palaiahnakote Shivakumara
Pattern Recognit.7
2026 Harnessing Multimodal Unlabeled Data for Enhanced Speech Emotion Recognition
abstract
Speech emotion recognition(SER) often faces challenges due to the lack of large, annotated datasets. The presence of abundance unlabeled data offers a chance to explore methods that could significantly improve SER systems. This study explores the feasibility of enhancing general speech models by incorporating unimodal and multimodal training objectives derived from unlabeled data, specifically tailored to extract emotional content. These multimodal objectives aim to refineself-supervised learning(SSL)-based representations that, while effective in SER, were not originally created to extract emotional cues from speech. Our methodology introduces a set of multimodal objectives focused on capturing information from three primary sources: acoustic signals, through a representation objective based on the extendedGeneva Minimalistic Acoustic Parameter Set(eGEMAPS); facial expressions, via visual representations obtained from a pre-trained facial expression recognition system; and textual content, through pseudo-labels generated by a pre-trained emotion sentiment model. These objectives are automatically generated from 70.7 hours of unlabeled emotional content captured in naturalistic settings. We apply our strategy to four state-of-the-art SSL-based speech models, aiming to enhance their capabilities in SER tasks with multimodal signals while still keeping inference strictly audio-only. Our experimental evaluations across the CREMA-D, MSP-IMPROV, and MSP-Podcast datasets demonstrate that our approach significantly improves SER performance, especially in settings with limited labeled data.
Lucas Goncalves, Carlos Busso
IEEE Trans. Affect. Comput.2
2026 Describe Where You Are: Improving Noise-Robustness for Speech Emotion Recognition With Text Description of the Environment
abstract
Speech emotion recognition(SER) systems often struggle in real-world environments, where ambient noise severely degrades their performance. This paper explores a novel approach that exploits prior knowledge of testing environments to maximize SER performance under noisy conditions. To address this task, we propose a text-guided, environment-aware training where an SER model is trained with contaminated speech samples and their paired noise description. We use a pre-trained text encoder to extract the text-based environment embedding and then fuse it to a transformer-based SER model during training and inference. We demonstrate the effectiveness of our approach through our experiment with the MSP-Podcast corpus and real-world additive noise samples collected from the Freesound and DEMAND repositories. Our experiment indicates that the text-based environment descriptions processed by alarge language model(LLM) produce representations that improve the noise-robustness of the SER system. With acontrastive learning(CL)-based representation, our proposed method can be improved by jointly fine-tuning the text encoder with the emotion recognition model. Under the -5dBsignal-to-noise ratio(SNR) level, fine-tuning the text encoder improves our CL-based representation method by 76.4% (arousal), 100.0% (dominance), and 27.7% (valence).
Seong-Gyun Leem, Daniel Fulford, Jukka-Pekka Onnela, David Gard, Carlos Busso
IEEE Trans. Affect. Comput.5
2025 Speaker Style-Aware Phoneme Anchoring For Improved Cross-Lingual Speech Emotion Recognition
abstract
Cross-lingual speech emotion recognition (SER) remains a challenging task due to differences in phonetic variability and speaker-specific expressive styles across languages. Effectively capturing emotion under such diverse conditions requires a framework that can align the externalization of emotions across different speakers and languages. To address this problem, we propose a speaker-style aware phoneme anchoring framework that aligns emotional expression at the phonetic and speaker levels. Our method builds emotionspecific speaker communities via graph-based clustering to capture shared speaker traits. Using these groups, we apply dual-space anchoring in speaker and phonetic spaces to enable better emotion transfer across languages. Evaluations on the MSP-Podcast (English) and BIIC-Podcast (Taiwanese Mandarin) corpora demonstrate improved generalization over competitive baselines and provide valuable insights into the commonalities in cross-lingual emotion representation.
Shreya G. Upadhyay, Carlos Busso, Chi-Chun Lee
ASRU2
2025 Efficient Fusion of Computationally Diverse Modalities Using Chunking and Cross-Attention
abstract
Emotion recognition is inherently a multimodal problem. Humans use both audible and visual cues to determine a person’s emotions. There has been extensive improvement in the methods we use to fuse audio and visual representations between two unimodal deep-learning models. However, there is a lack of accommodation for modalities that have a disparity in the amount of computational resources needed to provide the same amount of temporal information. As the sequence length increases, current methods often make simplifications such as discarding frames or cropping the sequence. This paper introduces a chunking methodology designed for cross-attention-based multimodal transformer architectures. The approach involves segmenting the visual input—the more computationally demanding modality—into chunks. Cross-attention is then performed between the encoded audio and visual features instead of the original sequence lengths of the unimodal backbones. Our method achieves significant improvements over conventional cross-attention techniques in the audio-visual domain for a six-class emotional recognition problem, demonstrating better F1 score, precision, and recall on the CREMA-D database while reducing computational overhead.
Christian Flores, Lucas Goncalves, Carlos Busso
ICASSP3
2025 Domain-Specific Adaptation in Speech Emotion Recognition Using Emotional Distribution Alignment
abstract
This work addresses the challenge of building speech emotion recognition models that generalize effectively across different domains, particularly when only limited target domain data is available with or without emotional label information. Traditional models often struggle with cross-domain performance due to the variability in emotional expressions and the lack of alignment between the training and target domains. We propose a novel approach that prioritizes aligning the emotional label distribution of the training data with that of the target domain by undersampling the source domain. Even though we intentionally reduce the size of the training set from the source domain, the emotional content alignment leads to clear performance improvements, outperforming models trained with the complete training set. This strategy highlights the importance of aligning emotional attributes during training, helping to create robust emotion recognition models across diverse applications. Our findings also reveal that performance significantly improves when even a small amount of labeled target domain data is available, allowing for a more accurate assessment of the emotional distribution in the target domain.
Abinay Reddy Naini, Donita Robinson, Elizabeth Richerson, Carlos Busso
ICASSP4
2025 Noise-Robust Speech Emotion Recognition Using Shared Self-Supervised Representations with Integrated Speech Enhancement
abstract
Recent studies have demonstrated the effectiveness of fine-tuning self-supervised speech representation models for speech emotion recognition (SER). However, applying SER in real-world environments remains challenging due to pervasive noise. Relying on low-accuracy predictions due to noisy speech can undermine the user’s trust. This paper proposes a unified self-supervised speech representation framework for enhanced speech emotion recognition designed to increase noise robustness in SER while generating enhanced speech. Our framework integrates speech enhancement (SE) and SER tasks, leveraging shared self-supervised learning (SSL)-derived features to improve emotion classification performance in noisy environments. This strategy encourages the SE module to enhance discriminative information for SER tasks. Additionally, we introduce a cascade unfrozen training strategy, where the SSL model is gradually unfrozen and fine-tuned alongside the SE and SER heads, ensuring training stability and preserving the generalizability of SSL representations. This approach demonstrates improvements in SER performance under unseen noisy conditions without compromising SE quality. When tested at a 0 dB signal-to-noise ratio (SNR) level, our proposed method outperforms the original baseline by 3.7% in F1-Macro and 2.7% in F1-Micro scores, where the differences are statistically significant.
Jing-Tong Tzeng, Seong-Gyun Leem, Ali N. Salman, Chi-Chun Lee, Carlos Busso
ICASSP5
2025 Mouth Articulation-Based Anchoring for Improved Cross-Corpus Speech Emotion Recognition
abstract
Cross-corpus speech emotion recognition (SER) plays a vital role in numerous practical applications. Traditional approaches to cross-corpus emotion transfer often concentrate on adapting acoustic features to align with different corpora, domains, or labels. However, acoustic features are inherently variable and error-prone due to factors like speaker differences, domain shifts, and recording conditions. To address these challenges, this study adopts a novel contrastive approach by focusing on emotion-specific articulatory gestures as the core elements for analysis. By shifting the emphasis on the more stable and consistent articulatory gestures, we aim to enhance emotion transfer learning in SER tasks. Our research leverages the CREMA-D and MSP-IMPROV corpora as benchmarks and it reveals valuable insights into the commonality and reliability of these articulatory gestures. The findings highlight mouth articulatory gesture potential as a better constraint for improving emotion recognition across different settings or domains.
Shreya G. Upadhyay, Ali N. Salman, Carlos Busso, Chi-Chun Lee
ICASSP3
2025 DifussionCleft: Facial Anomaly Synthesis Guided by Text
Karen Rosero, Lucas M. Harrison, Alex A. Kane, Rami R. Hallac, Carlos Busso
ICMI5
2025 Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset
Rui Liu 0008, Pu Gao, Jiatian Xi, Berrak Sisman, Carlos Busso, Haizhou Li 0001
INTERSPEECH5
2025 EmotionRankCLAP: Bridging Natural Language Speaking Styles and Ordinal Speech Emotion via Rank-N-Contrast
Shreeram Suresh Chandra, Lucas Goncalves, Junchen Lu, Carlos Busso, Berrak Sisman
INTERSPEECH4
2025 Can Emotion Fool Anti-spoofing?
Aurosweta Mahapatra, Ismail Rasim Ülgen, Abinay Reddy Naini, Carlos Busso, Berrak Sisman
INTERSPEECH4
2025 Analysis of Phonetic Level Similarities Across Languages in Emotional Speech
Pravin Mote, Abinay Reddy Naini, Donita Robinson, Elizabeth Richerson, Carlos Busso
INTERSPEECH5
2025 Vector Quantized Cross-lingual Unsupervised Domain Adaptation for Speech Emotion Recognition
Pravin Mote, Donita Robinson, Elizabeth Richerson, Carlos Busso
INTERSPEECH4
2025 The Interspeech 2025 Challenge on Speech Emotion Recognition in Naturalistic Conditions
Abinay Reddy Naini, Lucas Goncalves, Ali N. Salman, Pravin Mote, Ismail Rasim Ülgen, Thomas Thebaud, Laureano Moro-Velázquez, L. Paola García-Perera, Najim Dehak, Berrak Sisman, Carlos Busso
INTERSPEECH11
2025 Advancing Pediatric ASR: The Role of Voice Generation in Disordered Speech
Karen Rosero, Ali N. Salman, Shreeram Suresh Chandra, Berrak Sisman, Cortney Van't Slot, Alex A. Kane, Rami R. Hallac, Carlos Busso
INTERSPEECH8
2025 Speech emotion recognition in real static and dynamic human-robot interaction scenarios
abstract
The use of speech-based solutions is an appealing alternative to communicate in human-robot interaction (HRI). An important challenge in this area is processing distant speech which is often noisy, and affected by reverberation and time-varying acoustic channels. It is important to investigate effective speech solutions, especially in dynamic environments where the robots and the users move, changing the distance and orientation between a speaker and the microphone. This paper addresses this problem in the context of speech emotion recognition (SER), which is an important task to understand the intention of the message and the underlying mental state of the user. We propose a novel setup with a PR2 robot that moves as target speech and ambient noise are simultaneously recorded. Our study not only analyzes the detrimental effect of distance speech in this dynamic robot-user setting for speech emotion recognition but also provides solutions to attenuate its effect. We evaluate the use of two beamforming schemes to spatially filter the speech signal using either delay-and-sum (D&S) or minimum variance distortionless response (MVDR). We consider the original training speech recorded in controlled situations, and simulated conditions where the training utterances are processed to simulate the target acoustic environment. We consider the case where the robot is moving (dynamic case) and not moving (static case). For speech emotion recognition, we explore two state-of-the-art classifiers using hand-crafted features implemented with the ladder network strategy and learned features implemented with the wav2vec 2.0 feature representation. MVDR led to a signal-to-noise ratio higher than the basic D&S method. However, both approaches provided very similar average concordance correlation coefficient (CCC) improvements equal to 116% with the HRI subsets using the ladder network trained with the original MSP-Podcast training utterances. For the wav2vec 2.0-based model, only D&S led to improvements. Surprisingly, the static and dynamic HRI testing subsets resulted in a similar average concordance correlation coefficient. Finally, simulating the acoustic environment in the training dataset provided the highest average concordance correlation coefficient scores with the HRI subsets that are just 29% and 22% lower than those obtained with the original training/testing utterances, with ladder network and wav2vec 2.0, respectively.
Nicolás Grágeda, Carlos Busso, Eduardo Alvarado, Ricardo García, Rodrigo Mahú, Fernando Huenupán, Néstor Becerra Yoma
Comput. Speech Lang.2
2025 Multitask Transformer for Cross-Corpus Speech Emotion Recognition
abstract
Deep learning has significantly advanced the field of Speech Emotion Recognition (SER), yet its efficacy in cross-corpus scenarios remains a challenge. To overcome this limitation, recent studies demonstrate the success of multitask learning, which uses auxiliary tasks to reduce difference between source and target dataset (or transfer knowledge from source to target datasets). Despite the efforts, the overall accuracy for cross-corpus SER is still relatively low and needs attention. To improve performance, we propose a multitask framework with SER as the primary task and contrastive learning and information maximization as auxiliary tasks. We design the auxiliary tasks innovatively to use the target data without emotional labels to develop a better understanding of the target data. The core of our multitask framework is a pre-trained transformer. While transformers have gained attention in SER, their application to cross-corpus scenarios is still limited. Multimodal approaches for cross-corpus scenario is substantially limited as well. We use text as the second modality, developing separate multitask transformers for audio and text and conduct decision-level fusion during inference. We use publicly available and widely used speech corpora, including the IEMOCAP, MSP-IMPROV and EMO-DB databases. The results demonstrate the benefits of the proposed approach, achieving improved performance on the benchmark databases in cross-corpus settings.
Chung Soo Ahn, Rajib Rana, Carlos Busso, Jagath C. Rajapakse
IEEE Trans. Affect. Comput.3
2025 Differential Impacts of Monologue and Conversation on Speech Emotion Recognition
abstract
The advancement ofSpeech Emotion Recognition(SER) is significantly dependent on the quality of emotional speech corpora used for model training. Researchers in the field of SER have developed various corpora by adjusting design parameters to enhance the reliability of the training source. For this study, we focus on exploring communication modes of collection, specifically analyzing spontaneous emotional speech patterns gathered during conversation or monologue. While conversations are acknowledged as effective for eliciting authentic emotional expressions, systematic analyses are necessary to confirm their reliability as a better source of emotional speech data. We investigate this research question from perceptual differences and acoustic variability present in both emotional speeches. Our analyses on multi-lingual corpora show that, first, raters exhibit higher consistency for conversation recordings when evaluating categorical emotions, and second, perceptions and acoustic patterns observed in conversational samples align more closely with expected trends discussed in relevant emotion literature. We further examine the impact of these differences on SER modeling, which shows that we can train a more robust and stable SER model by using conversation data. This work provides comprehensive evidence suggesting that conversation may offer a better source compared to monologue for developing an SER model.
Woan-Shiuan Chien, Shreya G. Upadhyay, Carlos Busso, Chi-Chun Lee
IEEE Trans. Affect. Comput.4
2025 Minority Views Matter: Evaluating Speech Emotion Classifiers With Human Subjective Annotations by an All-Inclusive Aggregation Rule
abstract
When selecting test data for subjective tasks, most studies define ground truth labels using aggregation methods such as the majority or plurality rules. These methods discard data points without consensus, making the test set easier than practical tasks where a prediction is needed for each sample. However, the discarded data points often express ambiguous cues that elicit coexisting traits perceived by annotators. This paper addresses the importance of considering all the annotations and samples in the data, highlighting that only showing the model's performance on an incomplete test set selected by using the majority or plurality rules can lead to bias in the models’ performances. We focus onspeech-emotion recognition(SER) tasks. We observe that traditional aggregation rules have a data loss ratio ranging from 5.63% to 89.17%. From this observation, we propose a flexible method named the all-inclusive aggregation rule to evaluate SER systems on the complete test data. We contrast traditional single-label formulations with a multi-label formulation to consider the coexistence of emotions. We show that training an SER model with the data selected by the all-inclusive aggregation rule shows consistently higher macro-F1 scores when tested in the entire test set, including ambiguous samples without agreement.
Huang-Cheng Chou, Lucas Goncalves, Seong-Gyun Leem, Ali N. Salman, Chi-Chun Lee, Carlos Busso
IEEE Trans. Affect. Comput.6
2025 Versatile Audio-Visual Learning for Emotion Recognition
abstract
Most current audio-visual emotion recognition models lack the flexibility needed for deployment in practical applications. We envision a multimodal system that works even when only one modality is available and can be implemented interchangeably for either predicting emotional attributes or recognizing categorical emotions. Achieving such flexibility in a multimodal emotion recognition system is difficult due to the inherent challenges in accurately interpreting and integrating varied data sources. It is also a challenge to robustly handle missing or partial information while allowing direct switch between regression or classification tasks. This study proposes a versatile audio-visual learning (VAVL) framework for handling unimodal and multimodal systems for emotion regression or emotion classification tasks. We implement an audio-visual framework that can be trained even when audio and visual paired data is not available for part of the training set (i.e., audio only or only video is present). We achieve this effective representation learning with audio-visual shared layers, residual connections over shared layers, and a unimodal reconstruction task. Our experimental results reveal that our architecture significantly outperforms strong baselines on the CREMA-D, MSP-IMPROV, and CMU-MOSEI corpora. Notably, VAVL attains a new state-of-the-art performance in the emotional attribute prediction task on the MSP-IMPROV corpus.
Lucas Goncalves, Seong-Gyun Leem, Berrak Sisman, Carlos Busso
IEEE Trans. Affect. Comput.5
2025 Phonetically-Anchored Domain Adaptation for Cross-Lingual Speech Emotion Recognition
abstract
The prevalence of cross-lingualspeech emotion recognition(SER) modeling has significantly increased due to its wide range of applications. Previous studies have primarily focused on technical strategies to adapt features, domains, and labels across languages, often overlooking the underlying commonalities between the languages. In this study, we address the language adaptation challenge in cross-lingual scenarios by incorporating vowel-phonetic constraints. Our approach is structured in two main parts. First, we investigate the vowel-phonetic commonalities associated with specific emotions across languages, particularly focusing on common vowels that prove to be valuable for SER modeling. Second, we utilize these identified common vowels as anchors to facilitate cross-lingual SER. To demonstrate the effectiveness of our approach, we conduct case studies usingAmerican EnglishandTaiwanese Mandarinwith two naturalistic emotional speech corpora: the MSP-Podcast and BIIC-Podcast corpora. The approach leverages evidence that certain vowels, including monophthongs and diphthongs, exhibit emotion-specific commonality across languages, serving as phonetic anchors to enhance unsupervised cross-lingual SER learning. The proposed model surpasses baseline performance, highlighting the importance of phonetic similarities for effective language adaptation in cross-lingual SER scenarios.
Shreya G. Upadhyay, Luz Martinez-Lucas, William F. Katz, Carlos Busso, Chi-Chun Lee
IEEE Trans. Affect. Comput.4
2024 Enhanced Facial Landmarks Detection for Patients with Repaired Cleft Lip and Palate
abstract
Cleft lip and palate (CLP) is a congenital condition causing deformities in the oral and labial tissues. Post-surgery, patients often experience residual issues like facial asymmetry, and speech disorders. Tracking points in the orofacial area using a facial landmark detector (FLD) contributes to the assessment of speech development and movement impairments. However, off-the-shelf FLDs fail at delineating the lips of patients with repaired CLP. To address this need, our study introduces the CLP-Trans strategy, a domain transfer solution that employs tailor-made affine transformations to modify facial images sourced from publicly available datasets, which constitute our source domain, whereas images of patients with repaired CLP form our target domain. We aim to reduce distribution disparities between the source and target domains for FLD by simulating common outcomes of CLP repair surgery. The system utilizes a deep convolutional neural network (CNN) to learn from transformed images, therefore, preserving the privacy and facilitating the reproducibility of the findings. The strategy achieves statistically significant improvements in the normalized mean square error (NMSE), reducing it from 2.417 to 2.086 (i.e., 13.7% error reduction) by using the proposed strategy when evaluating images of patients with CLP.
Karen Rosero, Ali N. Salman, Berrak Sisman, Rami R. Hallac, Carlos Busso
FG5
2024 Dynamic Speech Emotion Recognition Using A Conditional Neural Process
abstract
The problem of predicting emotional attributes from speech has often focused on predicting a single value from a sentence or short speaking turn. These methods often ignore that natural emotions are both dynamic and dependent on context. To model the dynamic nature of emotions, we can treat the prediction of emotion from speech as a time-series problem. We refer to the problem of predicting these emotional traces as dynamic speech emotion recognition. Previous studies in this area have used models that treat all emotional traces as coming from the same underlying distribution. Since emotions are dependent on contextual information, these methods might obscure the context of an emotional interaction. This paper uses a neural process model with a segment-level speech emotion recognition (SER) model for this problem. This type of model leverages information from the time-series and predictions from the SER model to learn a prior that defines a distribution over emotional traces. Our proposed model performs 21% better than a bidirectional long short-term memory (BiLSTM) baseline when predicting emotional traces for valence.
Luz Martinez-Lucas, Carlos Busso
ICASSP2
2024 Generalization of Self-Supervised Learning-Based Representations for Cross-Domain Speech Emotion Recognition
abstract
Self-supervised learning (SSL) from unlabelled speech data has revolutionized speech representation learning. Among them, wavLM, wav2vec2, HuBERT, and Data2vec have produced benchmark performances on automatic speech recognition. However, few studies have explored the generalization of SSL-based representations to different tasks based on paralinguistic information in speech such as emotion recognition. This paper explores the generalization of all four popular SSL models for speech emotion recognition (SER) when trained and tested in different domains. We aim to understand how adaptable these SSL representations are when using simple domain adaptation techniques. The evaluation considers emotional speech databases that deviate in language, recording conditions, and emotional distribution, providing very different target domains. The results reveal the necessity to fine-tune the representations for the SER downstream. As the differences between the source and target domain increase, we observe that the unsupervised domain adaptation techniques are more effective. The analysis in this study provides useful insights to understand the advantages of different representations for domain adaptation in SER.
Abinay Reddy Naini, Mary A. Kohler, Elizabeth Richerson, Donita Robinson, Carlos Busso
ICASSP5
2024 Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition
abstract
Speaker embeddings carry valuable emotion-related information, which makes them a promising resource for enhancing speech emotion recognition (SER), especially with limited labeled data. Traditionally, it has been assumed that emotion information is indirectly embedded within speaker embeddings, leading to their under-utilization. Our study reveals a direct and useful link between emotion and state-of-the-art speaker embeddings in the form of intra-speaker clusters. By conducting a thorough clustering analysis, we demonstrate that emotion information can be readily extracted from speaker embeddings. In order to leverage this information, we introduce a novel contrastive pretraining approach applied to emotion-unlabeled data for speech emotion recognition. The proposed approach involves the sampling of positive and the negative examples based on the intra-speaker clusters of speaker embeddings. The proposed strategy, which leverages extensive emotion-unlabeled data, leads to a significant improvement in SER performance, whether employed as a standalone pretraining task or integrated into a multi-task pretraining setting.
Ismail Rasim Ülgen, Zongyang Du, Carlos Busso, Berrak Sisman
ICASSP3
2024 Lip Abnormality Detection for Patients with Repaired Cleft Lip and Palate: A Lip Normalization Approach
abstract
The cleft lip condition arises from the incomplete fusion of oral and labial structures during fetal development, impacting vital functions. After surgical closure, patients commonly present with abnormal lip shape, which may require secondary revision surgery for both aesthetic and functional improvement. However, a lack of standardized evaluation methods complicates decision-making for secondary surgery. To address this limitation, we propose a transformer-based lip normalization approach that filters out abnormalities and achieves a standardized appearance while preserving individual anatomy. An innovation of our approach is a lip transformation method using available face datasets to mimic repaired cleft lip shapes, enabling the training of deep learning models without using patients’ data. We employ a Siamese convolutional neural network that processes pre- and post-normalization images to detect lip abnormalities with an accuracy of 89%. We compare our approach with a single-branch model without lip normalization, which reached an accuracy of 60%. Our approach has the potential to provide an impartial view to determine the need for revision surgery while also assisting in the selection of healthcare tools specialized for patients with repaired cleft lip. The code for this work is available in our anonymous repository 1.
Karen Rosero, Ali N. Salman, Rami R. Hallac, Carlos Busso
ICMI4
2024 MSP-GEO Corpus: A Multimodal Database for Understanding Video-Learning Experience
abstract
Video-based learning has become a popular, scalable, and effective approach for students to learn new skills. Many of the challenges for video-based learning can be addressed with machine learning models. However, the available datasets often lack the rich source of data that is needed to accurately predict students’ learning experiences and outcomes. To address this limitation, we introduce the MSP-GEO corpus, a new multimodal database that contains detailed demographic and educational data, recordings of the students and their screens, and meta-data about the lecture during the learning experience. The MSP-GEO corpus was collected using a quasi-experimental pre-test/post-test design. It consists of more than 39,600 seconds (11 hours) of continuous facial footage from 76 participants watching one of three experimental videos on the topic of fossil formation, resulting in over one million facial images. The data collected includes 21 gaze synchronization points, webcam and monitor recordings, and metadata for pauses, plays, and timeline navigation. Additionally, we annotated the recordings for engagement, boredom, and confusion using human evaluators. The MSP-GEO corpus has the potential to improve the accuracy of video-based learning outcomes and experience predictions, facilitate research on the psychological processes of video-based learning, inform the design of instructional videos, and advance the development of learning analytics methods.
Ali N. Salman, Ning Wang 0104, Luz Martinez-Lucas, Andrea Vidal, Carlos Busso
ICMI5
2024 Speech emotion recognition with deep learning beamforming on a distant human-robot interaction scenario
Ricardo García, Rodrigo Mahú, Nicolás Grágeda, Alejandro Luzanto, Nicolas Bohmer, Carlos Busso, Néstor Becerra Yoma
INTERSPEECH6
2024 Bridging Emotions Across Languages: Low Rank Adaptation for Multilingual Speech Emotion Recognition
Lucas Goncalves, Donita Robinson, Elizabeth Richerson, Carlos Busso
INTERSPEECH4
2024 Keep, Delete, or Substitute: Frame Selection Strategy for Noise-Robust Speech Emotion Recognition
Seong-Gyun Leem, Daniel Fulford, Jukka-Pekka Onnela, David Gard, Carlos Busso
INTERSPEECH5
2024 Unsupervised Domain Adaptation for Speech Emotion Recognition using K-Nearest Neighbors Voice Conversion
Pravin Mote, Berrak Sisman, Carlos Busso
INTERSPEECH3
2024 WHiSER: White House Tapes Speech Emotion Recognition Corpus
Abinay Reddy Naini, Lucas Goncalves, Mary A. Kohler, Donita Robinson, Elizabeth Richerson, Carlos Busso
INTERSPEECH6
2024 Towards Naturalistic Voice Conversion: NaturalVoices Dataset with an Automatic Processing Pipeline
Ali N. Salman, Zongyang Du, Shreeram Suresh Chandra, Ismail Rasim Ülgen, Carlos Busso, Berrak Sisman
INTERSPEECH5
2024 A Layer-Anchoring Strategy for Enhancing Cross-Lingual Speech Emotion Recognition
Shreya G. Upadhyay, Carlos Busso, Chi-Chun Lee
INTERSPEECH2
2024 Driver Head Pose Estimation with Multimodal Temporal Fusion of Color and Depth Modeling Networks
abstract
For in-vehicle systems, head pose estimation (HPE) is a primitive task for many safety indicators, including driver attention modeling, visual awareness estimation, behavior detection, and gaze detection. The driver’s head pose information is also used to augment human-vehicle interfaces for infotainment and navigation. HPE is challenging, especially in the context of driving, due to the sudden variations in illumination, extreme poses, and occlusions. Due to these challenges, driver HPE based only on 2D color data is unreliable. These challenges can be addressed by 3D-depth data to an extent. We observe that features from 2D and 3D data complement each other. The 2D data provides detailed localized features, but is sensitive to illumination variations, whereas 3D data provides topological geometrical features and is robust to lighting conditions. Motivated by these observations, we propose a robust HPE model which fuses data obtained from color and depth cameras (i.e., 2D and 3D). The depth feature representation is obtained with a model based on PointNet++. The color images are processed with the ResNet-50 model. In addition, we add temporal modeling to our framework to exploit the time-continuous nature of head pose trajectories. We implement our proposed model using the multimodal driving monitoring (MDM) corpus, which is a naturalistic driving database. We present our model results with a detailed ablation study with unimodal and multimodal implementations, showing improvement in head pose estimation. We compare our results with baseline HPE models using regular cameras, including OpenFace 2.0 and HopeNet. Our fusion model achieves the best performance, obtaining an average root mean square error (RMSE) equal to 4.38 degrees.
Susmitha Gogineni, Carlos Busso
IV2
2024 Embracing Ambiguity And Subjectivity Using The All-Inclusive Aggregation Rule For Evaluating Multi-Label Speech Emotion Recognition Systems
abstract
Speech Emotion Recognition (SER) faces a distinct challenge compared to other speech-related tasks because the annotations will show the subjective emotional perceptions of different annotators. Previous SER studies often view the subjectivity of emotion perception as noise by using the majority rule or plurality rule to obtain the consensus labels. However, these standard approaches overlook the valuable information of labels that do not agree with the consensus and make it easier for the test set. Emotion perception can have co-occurring emotions in realistic conditions, and it is unnecessary to regard the disagreement between raters as noise. To bridge the SER into a multi-label task, we introduced an “all-inclusive rule,” which considers all available data, ratings, and distributional labels as multi-label targets and a complete test set. We demonstrated that models trained with multi-label targets generated by the proposed AR outperform conventional single-label methods across incomplete and complete test sets.
Huang-Cheng Chou, Lucas Goncalves, Seong-Gyun Leem, Ali N. Salman, Carlos Busso, Hung-yi Lee, Chi-Chun Lee
SLT6
2024 Deep temporal clustering features for speech emotion recognition
abstract
Deep clustering is a popular unsupervised technique for feature representation learning. We recently proposed the chunk-based DeepEmoCluster framework for speech emotion recognition (SER) to adopt the concept of deep clustering as a novel semi-supervised learning (SSL) framework, which achieved improved recognition performances over conventional reconstruction-based approaches. However, the vanilla DeepEmoCluster lacks critical sentence-level temporal information that is useful for SER tasks. This study builds upon the DeepEmoCluster framework, creating a powerful SSL approach that leverages temporal information within a sentence. We propose two sentence-level temporal modeling alternatives using either the temporal-net or the triplet loss function, resulting in a novel temporal-enhanced DeepEmoCluster framework to capture essential temporal information. The key contribution to achieving this goal is the proposed sentence-level uniform sampling strategy, which preserves the original temporal order of the data for the clustering process. An extra network module (e.g., gated recurrent unit) is utilized for the temporal-net option to encode temporal information across the data chunks. Alternatively, we can impose additional temporal constraints by using the triplet loss function while training the DeepEmoCluster framework, which does not increase model complexity. Our experimental results based on the MSP-Podcast corpus demonstrate that the proposed temporal-enhanced framework significantly outperforms the vanilla DeepEmoCluster framework and other existing SSL approaches in regression tasks for the emotional attributes arousal, dominance, and valence. The improvements are observed in fully-supervised learning or SSL implementations. Further analyses validate the effectiveness of the proposed temporal modeling, showing (1) high temporal consistency in the cluster assignment, and (2) well-separated emotional patterns in the generated clusters.
Carlos Busso
Speech Commun.2
2024 Analyzing Continuous-Time and Sentence-Level Annotations for Speech Emotion Recognition
abstract
The emotional content of several databases are annotated withcontinuous-time(CT) annotations, providing traces with frame-by-frame scores describing the instantaneous value of an emotional attribute. However, having a single score describing the global emotion of a short segment is more convenient for several emotion recognition formulations. A common approach is to derivesentence-level(SL) labels from CT annotations by aggregating the values of the emotional traces across time and annotators. How similar are these aggregated SL labels from labels originally collected at the sentence level? The release of the MSP-Podcast (SL annotations) and MSP-Conversation (CT annotations) corpora provides the resources to explore the validity of aggregating SL labels from CT annotations. There are 2,884 speech segments that belong to both corpora. Using this set, this study (1) compares both types of annotations using statistical metrics, (2) evaluates their inter-evaluator agreements, and (3) explores the effect of these SL labels onspeech emotion recognition(SER) tasks. The analysis reveals benefits of using SL labels derived from CT annotations in the estimation of valence. This analysis also provides insights on how the two types of labels differ and how that could affect a model.
Luz Martinez-Lucas, Carlos Busso
IEEE Trans. Affect. Comput.3
2024 Selective Acoustic Feature Enhancement for Speech Emotion Recognition With Noisy Speech
abstract
Aspeech emotion recognition(SER) system deployed on a real-world application is highly likely to encounter speech contaminated with unconstrained background noise. To deal with this issue, aspeech enhancement(SE) module can be attached to the SER system to compensate for the environmental difference of an input. Although the SE module can improve the quality and intelligibility of a given speech, there is a risk of affecting discriminative acoustic features for SER that are resilient to environmental differences. Exploring this idea, we propose to enhance only weak features that degrade the emotion recognition performance, while keeping strong features that are resilient to environmental differences. Our model first identifies weak feature sets by using multiple models trained with one acoustic feature at a time using clean speech. After training the single-feature models, we rank each speech feature by measuring three criteria: performance, robustness, and a joint rank ranking that combines performance and robustness. We group the weak features by cumulatively incrementing the features from the bottom to the top of each rank. Once the weak feature set is defined, we only enhance those weak features, keeping the resilient features unchanged. We implement these ideas with thelow-level descriptors(LLDs). We show that extracting LLDs from an enhanced speech signal does not improve the performance of weak features. Instead, directly enhancing the LLDs lead to better performance. Our experiment with clean and noisy versions of the MSP-Podcast corpus shows that the selective feature enhancement approach proposed in this study yields a 17.7% (arousal), 21.2% (dominance), and 3.3% (valence) performance gains over a system that enhances all the LLDs for the 10dBsignal-to-noise ratio(SNR) condition.
Seong-Gyun Leem, Daniel Fulford, Jukka-Pekka Onnela, David Gard, Carlos Busso
IEEE ACM Trans. Audio Speech Lang. Process.5
2024 An Interpretable Deep Mutual Information Curriculum Metric for a Robust and Generalized Speech Emotion Recognition System
abstract
It is difficult to achieve robust and well-generalized models for tasks involving subjective concepts such as emotion. It is inevitable to deal with noisy labels, given the ambiguous nature of human perception. Methodologies relying onsemi-supervised learning(SSL) and curriculum learning have been proposed to enhance the generalization of the models. This study proposes a noveldeep mutual information(DeepMI) metric, built with the SSL pre-trained DeepEmoCluster framework to establish the difficulty of samples. The DeepMI metric quantifies the relationship between the acoustic patterns and emotional attributes (e.g., arousal, valence, and dominance). The DeepMI metric provides a better curriculum, achieving state-of-the-art performance that is higher than results obtained with existing curriculum metrics forspeech emotion recognition(SER). We evaluate the proposed method with three emotional datasets in matched and mismatched testing conditions. The experimental evaluations systematically show that a model trained with the DeepMI metric not only obtains competitive generalization performances, but also maintains convergence stability. Furthermore, the extracted DeepMI values are highly interpretable, reflecting information ranks of the training samples.
Kusha Sridhar, Carlos Busso
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Analyzing the Effect of Affective Priming on Emotional Annotations
abstract
In the field of affective computing, emotional annotations are highly important for both the recognition and synthesis of human emotions. Researchers must ensure that these emotional labels are adequate for modeling general human perception. An unavoidable part of obtaining such labels is that human annotators are exposed to known and unknown stimuli before and during the annotation process that can affect their perception. Emotional stimuli cause an affective priming effect, which is a pre-conscious phenomenon in which previous emotional stimuli affect the emotional perception of a current target stimulus. In this paper, we use sequences of emotional annotations during a perceptual evaluation to study the effect of affective priming on emotional ratings of speech. We observe that previous emotional sentences with extreme emotional content push annotations of current samples to the same extreme. We create a sentence-level bias metric to study the effect of affective priming on speech emotion recognition(SER) modeling. The metric is used to identify subsets in the database with more affective priming bias intentionally creating biased datasets. We train and test SER models using the full and biased datasets. Our results show that although the biased datasets have low inter-evaluator agreements, SER models for arousal and dominance trained with those datasets perform the best. For valence, the models trained with the less-biased datasets perform the best.
Luz Martinez-Lucas, Ali N. Salman, Seong-Gyun Leem, Shreya G. Upadhyay, Chi-Chun Lee, Carlos Busso
ACII6
2023 An Intelligent Infrastructure Toward Large Scale Naturalistic Affective Speech Corpora Collection
abstract
The field of speech emotion recognition (SER) aims to create scientifically rigorous systems that can reliably characterize emotional behaviors expressed in speech. A key aspect for building SER systems is to obtain emotional data that is both reliable and reproducible for practitioners. However, academic researchers encounter difficulties in accessing or collecting naturalistic large-scale, reliable emotional recordings. Also, the best practices for data collection are not necessarily described or shared when presenting emotional corpora. To address this issue, the paper proposes the creation of an affective naturalistic database consortium (AndC) that can encourage multidisciplinary cooperation among researchers and practitioners in the field of affective computing. This paper’s contribution is twofold. First, it proposes the design of the AndC with a customizable-standard framework for intelligently-controlled emotional data collection. The focus is on leveraging naturalistic spontaneous recordings available on audio-sharing websites. Second, it presents as a case study the development of a naturalistic large-scale Taiwanese Mandarin podcast corpus using the customizable-standard intelligently-controlled framework. The AndC will enable research groups to effectively collect data using the provided pipeline and to contribute with alternative algorithms or data collection protocols.
Shreya G. Upadhyay, Woan-Shiuan Chien, Bo-Hao Su, Lucas Goncalves, Ya-Tse Wu, Ali N. Salman, Carlos Busso, Chi-Chun Lee
ACII7
2023 Combining Relative and Absolute Learning Formulations to Predict Emotional Attributes From Speech
abstract
Predicting absolute scores is the most common speech-emotion recognition (SER) task when predicting emotional attributes (i.e., valence, arousal, and dominance). However, studies have shown that emotion has an ordinal nature where it is more reliable to establish a preference between speech samples (e.g., one sample is more positive than the other). This paper pursues a novel direction to combine absolute and relative learning formulations for SER. The proposed multitask formulation can simultaneously estimate preference between speech samples and predict their absolute score, providing a flexible tool to analyze emotional content in speech. Both tasks mutually complement each other, allowing the model to outperform SER systems that are exclusively trained to either predict absolute scores or estimate preferences. The multitask weights can be set according to the intended applications, prioritizing one task while slightly compromising the performance of the other task.
Abinay Reddy Naini, Shruthi Subramanium, Seong-Gyun Leem, Carlos Busso
ASRU4
2023 Learning Cross-Modal Audiovisual Representations with Ladder Networks for Emotion Recognition
abstract
Representation learning is a challenging, but essential task in audiovisual learning. A key challenge is to generate strong cross-modal representations while still capturing discriminative information contained in unimodal features. Properly capturing this information is important to increase accuracy and robustness in audiovisual tasks. Focusing on emotion recognition, this study proposes novel cross-modal ladder networks to capture modality-specific information while building strong cross-modal representations. Our method utilizes representations from a backbone network to implement unsupervised auxiliary tasks to reconstruct intermediate layer representations across the acoustic and visual networks. The skip connections between the cross-modal encoder and decoder provide powerful modality-specific and multimodal representations for emotion recognition. Our model on the CREMA-D corpus achieves high performance with precision, recall, and F1 scores over 80% on a six-class problem.
Lucas Goncalves, Carlos Busso
ICASSP2
2023 Adapting a Self-Supervised Speech Representation for Noisy Speech Emotion Recognition by Using Contrastive Teacher-Student Learning
abstract
Studies have shown high performance in the speech emotion recognition (SER) task by fine-tuning a self-supervised speech representation model. Although this model can provide emotionally discriminative embedding in clean conditions, adapting it to a noisy target environment is still required when deployed on real-world applications. For adaptation, it is essential to balance between acquiring new knowledge from noisy speech and keeping the previous knowledge acquired during the pre-training and fine-tuning of the model. Therefore, we propose a contrastive teacher-student learning framework to retrain a self-supervised speech representation model for noisy SER. To keep the knowledge of the original model, we minimize the root mean square error between the clean embeddings from the original SER model and the noisy embeddings from the retrained model. To acquire the discriminative knowledge in the target noisy condition, we also minimize the InfoNCE loss by selecting the corresponding clean embedding as a positive sample and other noisy embeddings with different emotional labels as negative samples. Our experiment with the clean and noisy version of the MSP-Podcast corpus demonstrates that the contrastive teacher-student learning framework can significantly improve the performance of the model only trained with the clean speech in the target noisy condition for all the emotional attributes.
Seong-Gyun Leem, Daniel Fulford, Jukka-Pekka Onnela, David Gard, Carlos Busso
ICASSP5
2023 Role of Lexical Boundary Information in Chunk-Level Segmentation for Speech Emotion Recognition
abstract
Chunk-level speech emotion recognition (SER) is a common modeling scheme to obtain better recognition performance than sentence-level formulations. A key open question is the role of lexical boundary information in the process of splitting a sentence into small chunks. Is there any benefit in providing precise lexical boundary information to segment the speech into chunks (e.g., word-level alignments)? This study analyzes the role of lexical boundary information by exploring alternative segmentation strategies for chunk-level SER. We compare six chunk-level segmentation strategies that either consider word-level alignments or traditional time-based segmentation methods by varying the number of chunks and the duration of the chunks. We conduct extensive experiments to evaluate these chunk-level segmentation approaches using multiples corpora, and multiple acoustic feature sets. The results show a minor contribution of the word-level timing boundaries, where centering the chunks around words does not lead to significant performance gains. Instead, the critical factor to effectively segment a sentence into data chunks is to define the number of chunks according to the number of spoken words in the sentence.
Carlos Busso
ICASSP2
2023 Unsupervised Domain Adaptation for Preference Learning Based Speech Emotion Recognition
abstract
Retrieving speech samples that have specific expressive content has many applications. It is desirable to build a preference learning framework that ranks speech samples according to emotional attribute values that generalize well to new domains. A popular architecture for preference learning is the RankNet framework, which uses a function to obtain the preference between pairs of speech sentences. This study explores implementing this function with alternative feature representations that are explicitly selected to reduce the mismatch between source and target domains. In particular, we implement our preference-learning based speech emotion recognition (SER) system using ladder networks and adversarial domain adaptation. The study also proposes a novel combination of these two unsupervised domain adaptation strategies. The experimental results in cross-corpus evaluations using the MSP-Podcast and MSP-IMPROV datasets reveal that the proposed adversarial domain adaptation on a ladder network-based feature representation performs the best across different conditions. The results also show that preference learning leads to better precision for retrieval tasks than comparable SER systems built to directly predict absolute emotional attribute scores.
Abinay Reddy Naini, Mary A. Kohler, Carlos Busso
ICASSP3
2023 Phonetic Anchor-Based Transfer Learning to Facilitate Unsupervised Cross-Lingual Speech Emotion Recognition
abstract
Modeling cross-lingual speech emotion recognition (SER) has become more prevalent because of its diverse applications. Existing studies have mostly focused on technical approaches that adapt the feature, domain, or label across languages, without considering in detail the similarities between the languages. This study focuses on domain adaptation in cross-lingual scenarios using phonetic constraints. This work is framed in a twofold manner. First, we analyze emotion-specific phonetic commonality across languages by identifying common vowels that are useful for SER modeling. Second, we leverage these common vowels as an anchoring mechanism to facilitate cross-lingual SER. We consider American English and Taiwanese Mandarin as a case study to demonstrate the potential of our approach. This work uses two in-the-wild natural emotional speech corpora: MSP-Podcast (American English), and BIIC-Podcast (Taiwanese Mandarin). The proposed unsupervised cross-lingual SER model using these phonetical anchors outperforms the baselines with a 58.64% of unweighted average recall (UAR).
Shreya G. Upadhyay, Luz Martinez-Lucas, Bo-Hao Su, Woan-Shiuan Chien, Ya-Tse Wu, William F. Katz, Carlos Busso, Chi-Chun Lee
ICASSP8
2023 The Importance of Calibration: Rethinking Confidence and Performance of Speech Multi-label Emotion Classifiers
Huang-Cheng Chou, Lucas Goncalves, Seong-Gyun Leem, Chi-Chun Lee, Carlos Busso
INTERSPEECH5
2023 Distant Speech Emotion Recognition in an Indoor Human-robot Interaction Scenario
Nicolás Grágeda, Eduardo Alvarado, Rodrigo Mahú, Carlos Busso, Néstor Becerra Yoma
INTERSPEECH4
2023 Computation and Memory Efficient Noise Adaptation of Wav2Vec2.0 for Noisy Speech Emotion Recognition with Skip Connection Adapters
Seong-Gyun Leem, Daniel Fulford, Jukka-Pekka Onnela, David Gard, Carlos Busso
INTERSPEECH5
2023 Preference Learning Labels by Anchoring on Consecutive Annotations
Abinay Reddy Naini, Ali N. Salman, Carlos Busso
INTERSPEECH3
2023 Seatbelt Segmentation Using Synthetic Images
abstract
Recent advancement in deep learning has led to an increased interest in image processing and computer vision applications for driver monitoring systems. One of the applications where these techniques can be useful is in segmenting and tracking seatbelts. A seatbelt is an important safety feature in the vehicle that if properly used can save lives. Efficient segmentation of the seatbelts in an image provides important information about the correct use of seatbelts. The challenge in developing deep learning algorithms for seatbelt detection and segmentation is the manual annotations required for this task, which is cumbersome. This paper explores a novel formulation to efficiently train a seatbelt model with minimal supervision. We exploit the textureless and shape characteristics of the seatbelts to programmatically synthesize images. Our proposed method synthetically creates images that resemble seatbelt patterns. After training a model exclusively with synthetic images, we iteratively fine-tune it using naturalistic images extracted from online video-sharing websites. The labels for these images are pseudo-labels assigned by the model to confident predictions. Fine-tuning helps adapt the model to better work on real naturalistic images, improving the performance of the system. We obtain an F1-score of 0.55 in segmenting the seatbelt with this approach. We also experiment with fine-tuning the model with a small number of naturalistic images with annotated labels. After pretraining on synthetic samples and pseudo-labeled naturalistic images, we achieve an F1-score of 0.67 using only 200 annotated images.
Isaac Brooks, Soumitry J. Ray, Rajesh Narasimha, Naofal Al-Dhahir, Carlos Busso
IV6
2023 Example-Based Query To Identify Causes of Driving Anomaly with Few Labeled Samples
abstract
Driving anomaly detection is important for advanced driver assistance systems (ADAS) to increase driving safety and avoid traffic accidents. However, driving anomaly detection faces many challenges such as numerous and uncertain abnormal patterns observed on the road, sparsity of real anomaly cases documented with accurate labels, and rigid existing systems that rely on manually set thresholds and rules. Previous studies have proposed unsupervised methods for driving anomaly detection in the driver’s behaviors or the road condition by identifying deviations from normal driving conditions. A challenge with unsupervised models is the lack of interpretability, where the cause of the anomaly is not always clear. We address this problem with an example-based query method that combines unsupervised anomaly detection methods with the multi-label k-nearest neighbors (ML-KNN) algorithm to interpret the detected driving anomalies by identifying their possible causes (e.g., surrounding objects or driver’s errors). Our approach relies on a few manually labeled driving segments that are efficiently used as anchors to retrieve the causes of driving anomalies in a given driving segment. These anchors are projected into the embedding created by unsupervised driving anomaly detection systems. The experimental results show that this method can effectively identify the causes of driving anomalies, even for abnormal driving segments triggered by multiple causes. The evaluation shows the flexibility of our proposed solution, where we successfully implement the ML-KNN approach with three alternative feature representations.
Yuning Qiu, Teruhisa Misu, Carlos Busso
IV3
2023 Multimodal attention for lip synthesis using conditional generative adversarial networks
Andrea Vidal, Carlos Busso
Speech Commun.2
2023 Quantifying Emotional Similarity in Speech
abstract
This study proposes the novel formulation of measuring emotional similarity between speech recordings. This formulation explores the ordinal nature of emotions by comparing emotional similarities instead of predicting an emotional attribute, or recognizing an emotional category. The proposed task determines which of two alternative samples has the most similar emotional content to the emotion of a given anchor. This task raises some interesting questions. Which is the emotional descriptor that provide the most suitable space to assess emotional similarities Can deep neural networks (DNNs learn representations to robustly quantify emotional similarities We address these questions by exploring alternative emotional spaces created with attribute-based descriptors and categorical emotions. We create the representation using a DNN trained with the triplet loss function, which relies on triplets formed with an anchor, a positive example, and a negative example. We select a positive sample that has similar emotion content to the anchor, and a negative sample that has dissimilar emotion to the anchor. The task of our DNN is to identify the positive sample. The experimental evaluations demonstrate that we can learn a meaningful embedding to assess emotional similarities, achieving higher performance than human evaluators asked to complete the same task.
John B. Harvill, Seong-Gyun Leem, Mohammed Abdel-Wahab 0001, Reza Lotfian, Carlos Busso
IEEE Trans. Affect. Comput.5
2023 Chunk-Level Speech Emotion Recognition: A General Framework of Sequence-to-One Dynamic Temporal Modeling
abstract
A critical issue of current speech-based sequence-to-one learning tasks, such as speech emotion recognition(SER), is the dynamic temporal modeling for speech sentences with different durations. The goal is to extract an informative representation vector of the sentence from acoustic feature sequences with varied length. Traditional methods rely on static descriptions such as statistical functions or a universal background model (UBM), which are not capable of characterizing dynamic temporal changes. Recent advances in deep learning architectures provide promising results, directly extracting sentence-level representations from frame-level features. However, conventional cropping and padding techniques that deal with varied length sequences are not optimal, since they truncate or artificially add sentence-level information. Therefore, we propose a novel dynamic chunking approach, which maps the original sequences of different lengths into a fixed number of chunks that have the same duration by adjusting their overlap. This simple chunking procedure creates a flexible framework that can incorporate different feature extractions and sentence-level temporal aggregation approaches to cope, in a principled way, with different sequence-to-one tasks. Our experimental results based on three databases demonstrate that the proposed framework provides: 1) improvement in recognition accuracy, 2) robustness toward different temporal length predictions, and 3) high model computational efficiency advantages.
Carlos Busso
IEEE Trans. Affect. Comput.2
2023 Sequential Modeling by Leveraging Non-Uniform Distribution of Speech Emotion
abstract
The expression and perception of human emotions are not uniformly distributed over time. Therefore, tracking local changes of emotion within a segment can lead to better models forspeech emotion recognition(SER), even when the task is to provide a sentence-level prediction of the emotional content. A challenge to exploring local emotional changes within a sentence is that most existing emotional corpora only provide sentence-level annotations (i.e., one label per sentence). This labeling approach is not appropriate for leveraging the dynamic emotional trends within a sentence. We propose a framework that splits a sentence into a fixed number of chunks, generating chunk-level emotional patterns. The approach relies on emotion rankers to unveil the emotional pattern within a sentence, creating continuous emotional curves. Our approach trains the sentence-level SER model with asequence-to-sequenceformulation by leveraging the retrieved emotional curves. The proposed method achieves the bestconcordance correlation coefficient(CCC) prediction performance for arousal (0.7120), valence (0.3125), and dominance (0.6324) on the MSP-Podcast corpus. In addition, we validate the approach with experiments on the IEMOCAP and MSP-IMPROV databases. We further compare the retrieved curves with time-continuous emotional traces. The evaluation demonstrates that these retrieved chunk-label curves can effectively capture emotional trends within a sentence, displaying a time-consistency property that is similar to time-continuous traces annotated by human listeners. The proposed SER model learns meaningful, complementary, local information that contributes to the improvement of sentence-level predictions of emotional attributes.
Carlos Busso
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Aligning Small Datasets Using Domain Adversarial Learning: Applications in Automated in Vivo Oral Cancer Diagnosis
abstract
Deep learning approaches for medical image analysis are limited by small data set size due to factors such as patient privacy and difficulties in obtaining expert labelling for each image. In medical imaging system development pipelines, phases for system development and classification algorithms often overlap with data collection, creating small disjoint data sets collected at numerous locations with differing protocols. In this setting, merging data from different data collection centers increases the amount of training data. However, a direct combination of datasets will likely fail due to domain shifts between imaging centers. In contrast to previous approaches that focus on a single data set, we add a domain adaptation module to a neural network and train using multiple data sets. Our approach encourages domain invariance between two multispectral autofluorescence imaging (maFLIM) data sets of in vivo oral lesions collected with an imaging system currently in development. The two data sets have differences in the sub-populations imaged and in the calibration procedures used during data collection. We mitigate these differences using a gradient reversal layer and domain classifier. Our final model trained with two data sets substantially increases performance, including a significant increase in specificity. We also achieve a significant increase in average performance over the best baseline model train with two domains (p = 0.0341). Our approach lays the foundation for faster development of computer-aided diagnostic systems and presents a feasible approach for creating a robust classifier that aligns images from multiple data centers in the presence of domain shifts.
Kayla Caughlin, Elvis Duran-Sierra, Shuna Cheng, Rodrigo Cuenca-Martinez, Beena Ahmed, Jim Xiuquan Ji, Mathias Martinez, Moustafa Al-Khalil, Hussain Al-Enazi, Yi-Shing Lisa Cheng, Javier A. Jo, Carlos Busso
IEEE J. Biomed. Health Informatics13
2022 Monologue versus Conversation: Differences in Emotion Perception and Acoustic Expressivity
abstract
Advancing speech emotion recognition (SER) depends highly on the source used to train the model, i.e., the emotional speech corpora. By permuting different design parameters, researchers have released versions of corpora that attempt to provide a better-quality source for training SER. In this work, we focus on studying communication modes of collection. In particular, we analyze the patterns of emotional speech collected during interpersonal conversations or monologues. While it is well known that conversation provides a better protocol for eliciting authentic emotion expressions, there is a lack of systematic analyses to determine whether conversational speech provide a “better-quality” source. Specifically, we examine this research question from three perspectives: perceptual differences, acoustic variability and SER model learning. Our analyses on the MSP-Podcast corpus show that: 1) rater's consistency for conversation recordings is higher when evaluating categorical emotions, 2) the perceptions and acoustic patterns observed on conversations have properties that are better aligned with expected trends discussed in emotion literature, and 3) a more robust SER model can be trained from conversational data. This work brings initial evidences stating that samples of conversations may provide a better-quality source than samples from monologues for building a SER model.
Woan-Shiuan Chien, Shreya G. Upadhyay, Ya-Tse Wu, Bo-Hao Su, Carlos Busso, Chi-Chun Lee
ACII6
2022 Exploiting Annotators' Typed Description of Emotion Perception to Maximize Utilization of Ratings for Speech Emotion Recognition
abstract
The decision of ground truth for speech emotion recognition (SER) is still a critical issue in affective computing tasks. Previous studies on emotion recognition often rely on consensus labels after aggregating the classes selected by multiple annotators. It is common for a perceptual evaluation conducted to annotate emotional corpora to include the class “other,” allowing the annotators the opportunity to describe the emotion with their own words. This practice provides valuable emotional information, which, however, is ignored in most emotion recognition studies. This paper utilizes easy-accessed natural language processing toolkits to mine the sentiment of these typed descriptions, enriching and maximizing the information obtained from the annotators. The polarity information is combined with primary and secondary annotations provided by individual evaluators under a label distribution framework, creating a complete representation of the emotional content of the spoken sentences. Finally, we train multitask learning SER models with existing learning methods (soft-label, multi-label, and distribution-label) to show the performance of the novel ground truth in the MSP-Podcast corpus.
Huang-Cheng Chou, Chi-Chun Lee, Carlos Busso
ICASSP4
2022 AuxFormer: Robust Approach to Audiovisual Emotion Recognition
abstract
A challenging task in audiovisual emotion recognition is to implement neural network architectures that can leverage and fuse multimodal information while temporally aligning modalities, handling missing modalities, and capturing information from all modalities without losing information during training. These requirements are important to achieve model robustness and to increase accuracy on the emotion recognition task. A recent approach to perform multimodal fusion is to use the transformer architecture to properly fuse and align the modalities. This study proposes the AuxFormer framework, which addresses in a principled way the aforementioned challenges. AuxFormer combines the transformer framework with auxiliary networks. It uses shared losses to infuse information from single-modality networks that are separately embedded. The extra layer of audiovisual information added to our main network retains information that would otherwise be lost during training. The results show that the AuxFormer architecture achieves macro and micro F1Scores of 71.3% and 71.7%, respectively, on the CREMA-D corpus. For the MSP-IMPROV corpus, AuxFormer achieves a macro and micro F1-Scores of 70.4% and 76.5%, respectively. The results for both corpora are significantly better than strong baselines, indicating that our framework benefits from auxiliary networks. We also show that under non-ideal conditions (e.g., missing modalities) our architecture is able to sustain strong performance under audio-only and video-only scenarios, benefiting from a optimized training strategy.
Lucas Goncalves, Carlos Busso
ICASSP2
2022 Not All Features are Equal: Selection of Robust Features for Speech Emotion Recognition in Noisy Environments
abstract
Speech emotion recognition (SER) system deployed in real-world applications often encounters noisy speech. While most noise compensation techniques consider all acoustic features to have equal impact on the SER model, some acoustic features may be more sensitive to noisy conditions. This paper investigates the noise robustness of each feature in the acoustic feature set. We focus on low-level descriptors (LLDs) commonly used in SER systems. We firstly train SER models with clean speech by only using a single LLD. Then, we rank each LLD with respect to the absolute performance on a development set contaminated with noise, and the relative performance decrease from the results from the models trained with the clean set. Our experiment shows that using all the LLDs leads to worse performance than training the system with a single robust LLD. We propose to select a group of robust features according to their performance and robustness in noisy condition. Without using any compensation method, our feature selection methods improve the performance by 24.4% (arousal), 23.9% (dominance), and 43.2% (valence) in the 10dB noisy condition. Moreover, even though the selection is conducted with the 10dB condition, our selection methods also yield performance improvements in unseen noisy recording conditions.
Seong-Gyun Leem, Daniel Fulford, Jukka-Pekka Onnela, David Gard, Carlos Busso
ICASSP5
2022 Incorporating Gaze Behavior Using Joint Embedding With Scene Context for Driver Takeover Detection
abstract
Despite the recent advancement in driver assistance systems, most existing solutions and partial automation systems such as SAE Level 2 driving automation systems assume that the driver is in the loop; the human driver must continuously monitor the driving environment. Frequent transition of maneuver control is expected between the driver and the car while using such automation in difficult traffic conditions. In this work, we aim to predict driver takeover timing in order for the system to prepare transition from automation to driver control. While previous studies indicated that eye gaze is an important cue to predict driver takeover, we hypothesize that traffic condition as well as the reliability of the driving automation also have a strong impact. Therefore, we propose an algorithm that jointly consider the driver’s gaze information and contextual driving environment, which is complemented with the vehicle operational and driver physiological signals. Specifically, we consider joint embedding of traffic scene information and gaze behavior using 3DConvolutional Neural Network (3D-CNN). We demonstrate that our algorithm is successfully able to predict driver takeover intent, using user study data from 28 participants collected in simulated driving environments.
Yuning Qiu, Carlos Busso, Teruhisa Misu, Kumar Akash
ICASSP2
2022 Privacy Preserving Personalization for Video Facial Expression Recognition Using Federated Learning
abstract
The increased ubiquitousness of small smart devices, such as cellphones, tablets, smart watches and laptops, has led to unique user data, which can be locally processed. The sensors (e.g., microphones and webcam) and improved hardware of the new devices have allowed running deep learning models that 20 years ago would have been exclusive to high-end expensive machines. In spite of this progress, state-of-the-art algorithms for facial expression recognition (FER) rely on architectures that cannot be implemented on these devices due to computational and memory constraints. Alternatives involving cloud-based solutions impose privacy barriers that prevent their adoption or user acceptance in wide range of applications. This paper proposes a lightweight model that can run in real-time for image facial expression recognition (IFER) and video facial expression recognition (VFER). The approach relies on a personalization mechanism locally implemented for each subject by fine-tuning a central VFER model with unlabeled videos from a target subject. We train the IFER model to generate pseudo labels and we select the videos with the highest confident predictions to be used for adaptation. The adaptation is performed by implementing a federated learning strategy where the weights of the local model are averaged and used by the central VFER model. We demonstrate that this approach can improve not only the performance on the edge device providing personalized models to the users, but also the central VFER model. We implement a federated learning strategy where the weights of the local models are averaged and used by the central VFER. Within corpus and cross-corpus evaluations on two emotional databases demonstrate that edge models adapted with our personalization strategy achieve up to 13.1% gains in F1-scores. Furthermore, the federated learning implementation improves the mean micro F1-score of the central VFER model by up to 3.4%. The proposed lightweight solution is ideal for interactive user interfaces that preserve the data of the users.
Ali N. Salman, Carlos Busso
ICMI2
2022 Exploiting Co-occurrence Frequency of Emotions in Perceptual Evaluations To Train A Speech Emotion Classifier
Huang-Cheng Chou, Chi-Chun Lee, Carlos Busso
INTERSPEECH3
2022 Improving Speech Emotion Recognition Using Self-Supervised Learning with Domain-Specific Audiovisual Tasks
Lucas Goncalves, Carlos Busso
INTERSPEECH2
2022 Driving Anomaly Detection Using Contrastive Multiview Coding to Interpret Cause of Anomaly
abstract
Modern advanced driver assistant systems (ADAS) rely on various types of sensors to monitor the vehicle status, driver's behaviors and road condition. The multimodal systems in the vehicle include sensors, such as accelerometers, pressure sensors, cameras, lidar and radars. When looking at a given scene with multiple modalities, there should be congruent in-formation among different modalities. Exploring the congruent information across modalities can lead to appealing solutions to create robust multimodal representations. This work proposes an unsupervised approach based on contrastive multiview coding (CMC) to capture the correlations in representations extracted from different modalities, learning a more discriminative rep-resentation space for unsupervised anomaly driving detection. We use CMC to train our model to extract view-invariant factors by maximizing the mutual information between mul-tiple representations from a given view, and increasing the distance of views from unrelated segments. We consider the vehicle driving data, driver's physiological data, and external environment data consisting of distances to nearby pedestrians, bicycles, and vehicles. The experimental results on the driving anomaly dataset (DAD) indicate that the CMC representation is effective for driving anomaly detection. The approach is efficient, scalable and interpretable, where the distances in the contrastive embedding for each view can be used to understand potential causes of the detected anomalies.
Yuning Qiu, Teruhisa Misu, Carlos Busso
IROS3
2022 Robust Audiovisual Emotion Recognition: Aligning Modalities, Capturing Temporal Information, and Handling Missing Features
abstract
Emotion recognition using audiovisual features is a challenging task for human-machine interaction systems. Under ideal conditions (perfect illumination, clean speech signals, and non-occluded visual data) many systems are able to achieve reliable results. However, few studies have considered developing multimodal systems and training strategies to build systems that can perform well under non ideal conditions. Audiovisual models still face challenging problems such as misalignment of modalities, lack of temporal modeling, and missing features due to noise or occlusions. In this article, we implement a model that combines auxiliary networks, a transformer architecture, and an optimized training mechanism to achieve a robust system for audiovisual emotion recognition that addresses, in a principled way, these challenges. Our evaluation analyzes how well this model performs in ideal conditions and when modalities are missing. We contrast this method with other multimodal fusion methods for emotion recognition. Our experimental results based on two audiovisual databases demonstrate that the proposed framework achieves: 1) improvements in emotion recognition accuracy, 2) better alignment and fusion of audiovisual features at the model level, 3) awareness of temporal information, and 4) robustness to non-ideal scenarios.
Lucas Goncalves, Carlos Busso
IEEE Trans. Affect. Comput.2
2022 Unsupervised Personalization of an Emotion Recognition System: The Unique Properties of the Externalization of Valence in Speech
abstract
The prediction of valence from speech is an important, but challenging problem. The expression of valence in speech has speaker-dependent cues, which contribute to performances that are often significantly lower than the prediction of other emotional attributes such as arousal and dominance. A practical approach to improve valence prediction from speech is to adapt the models to the target speakers in the test set. Adapting aspeech emotion recognition(SER) system to a particular speaker is a hard problem, especially withdeep neural networks(DNNs), since it requires optimizing millions of parameters. This study proposes an unsupervised approach to address this problem by searching for speakers in the train set with similar acoustic patterns as the speaker in the test set. Speech samples from the selected speakers are used to create the adaptation set. This approach leverages transfer learning using pre-trained models, which are adapted with these speech samples. We propose three alternative adaptation strategies: unique speaker, oversampling and weighting approaches. These methods differ on the use of the adaptation set in the personalization of the valence models. The results demonstrate that a valence prediction model can be efficiently personalized with these unsupervised approaches, leading to relative improvements as high as 13.52%.
Kusha Sridhar, Carlos Busso
IEEE Trans. Affect. Comput.2
2022 Temporal Head Pose Estimation From Point Cloud in Naturalistic Driving Conditions
abstract
Head pose estimation is an important problem as it facilitates tasks such as gaze estimation and attention modeling. In the automotive context, head pose provides crucial information about the driver’s mental state, including drowsiness, distraction and attention. It can also be used for interaction with in-vehicle infotainment systems. While computer vision algorithms using RGB cameras are reliable in controlled environments, head pose estimation is a challenging problem in the car due to sudden illumination changes, occlusions and large head rotations that are common in a vehicle. These issues can be partially alleviated by using depth cameras. Head rotation trajectories are continuous with important temporal dependencies. Our study leverages this observation, proposing a novel temporal deep learning model for head pose estimation from point cloud. The approach extracts discriminative feature representation directly from point cloud data, leveraging the 3D spatial structure of the face. The frame-based representations are then combined withbidirectional long short term memory(BLSTM) layers. We train this model on the newly collectedmultimodal driver monitoring(MDM) dataset, achieving better results compared to non-temporal algorithms using point cloud data, and state-of-the-art models using RGB images. We further show quantitatively and qualitatively that incorporating temporal information provides large improvements not only in accuracy, but also in the smoothness of the predictions.
Tiancheng Hu, Carlos Busso
IEEE Trans. Intell. Transp. Syst.3
2022 The Multimodal Driver Monitoring Database: A Naturalistic Corpus to Study Driver Attention
abstract
A smart vehicle should be able to monitor the actions and behaviors of the human driver to provide critical warnings or intervene when necessary. Recent advancements in deep learning and computer vision have shown great promise in monitoring human behavior and activities. While these algorithms work well in a controlled environment, naturalistic driving conditions add new challenges such as illumination variations, occlusions, and extreme head poses. A vast amount of in-domain data is required to train models that provide high performance in predicting driving related tasks to effectively monitor driver actions and behaviors. Toward building the required infrastructure, this paper presents themultimodal driver monitoring(MDM) dataset, which was collected with 59 subjects that were recorded performing various tasks. We use the Fi-Cap device that continuously tracks the head movement of the driver using fiducial markers, providing frame-based annotations to train head pose algorithms in naturalistic driving conditions. We ask the driver to look at predetermined gaze locations to obtain accurate correlation between the driver’s facial image and visual attention. We also collect data when the driver performs common secondary activities such as navigation using a smart phone and operating the in-car infotainment system. All of the driver’s activities are recorded with high definition RGB cameras and a time-of-flight depth camera. We also record thecontroller area network-bus(CAN-Bus), extracting important information. These high quality recordings serve as the ideal resource to train various efficient algorithms for monitoring the driver, providing further advancements in the field of in-vehicle safety systems.
Mohamed F. Marzban, Tiancheng Hu, Mohamed Hany Mahmoud, Naofal Al-Dhahir, Carlos Busso
IEEE Trans. Intell. Transp. Syst.6
2021 Generative Approach Using Soft-Labels to Learn Uncertainty in Predicting Emotional Attributes *
abstract
This paper presents a novel speech emotion recognition (SER) method to capture the uncertainty in predicting emotional attributes using the true distribution of scores provided by annotators as ground truth (i.e., soft-labels). Reliable, generalizable, and scalable SER systems are important in areas such as healthcare, customer service, security, and defense. A barrier to build these systems is the lack of quality labels due to the expensive annotation process, leading to poor generalization. To address this limitation, this study proposes a semi-supervised generative modeling approach using a variational autoencoder (VAE) with an emotional regressor at the bottleneck trained with soft-labels of emotional attributes. We demonstrate that estimating uncertainties in predicting emotional attribute scores is possible with soft-labels. We analyze the benefits of uncertainty estimation with a reject option formulation, where the model can abstain from predicting emotion when it is less confident. At 60% test coverage, we achieve relative improvements in concordance correlation coefficient (CCC) up to 16.85% for valence, 7.12% for arousal, and 8.01% for dominance. Furthermore, we propose an uncertainty transfer learning strategy where uncertainties learned from one attribute are used as a sample re-ordering criterion for another attribute, achieving additional improvements in prediction performance for valence. We also demonstrate the generalization power of our method in comparison to other uncertainty estimating methods using cross-corpus evaluations. Finally, we demonstrate that our method has lower computational complexity than alternative approaches.
Kusha Sridhar, Carlos Busso
ACII3
2021 Deepemocluster: a Semi-Supervised Framework for Latent Cluster Representation of Speech Emotions
abstract
Semi-supervised learning (SSL) is an appealing approach to resolve generalization problem for speech emotion recognition (SER) systems. By utilizing large amounts of unlabeled data, SSL is able to gain extra information about the prior distribution of the data. Typically, it can lead to better and robust recognition performance. Existing SSL approaches for SER include variations of encoder-decoder model structures such as autoencoder (AE) and variational autoecoders (VAEs), where it is difficult to interpret the learning mechanism behind the latent space. In this study, we introduce a new SSL framework, which we refer to as the DeepEmoCluster framework, for attribute-based SER tasks. The DeepEmoCluster framework is an end-to-end model with mel-spectrogram inputs, which combines a self-supervised pseudo labeling classification network with a supervised emotional attribute regressor. The approach encourages the model to learn latent representations by maximizing the emotional separation of K-means clusters. Our experimental results based on the MSP-Podcast corpus indicate that the DeepEmoCluster framework achieves competitive prediction performances in fully supervised scheme, outperforming baseline methods in most of the conditions. The approach can be further improved by incorporating extra unlabeled set. Moreover, our experimental results explicitly show that the latent clusters have emotional dependencies, enriching the geometric interpretation of the clusters.
Kusha Sridhar, Carlos Busso
ICASSP3
2021 Separation of Emotional and Reconstruction Embeddings on Ladder Network to Improve Speech Emotion Recognition Robustness in Noisy Conditions
Seong-Gyun Leem, Daniel Fulford, Jukka-Pekka Onnela, David Gard, Carlos Busso
Interspeech5
2021 Voice Activity Detection with Teacher-Student Domain Emulation
Jarrod Luckenbaugh, Samuel Abplanalp, Rachel Gonzalez, Daniel Fulford, David Gard, Carlos Busso
Interspeech6
2021 Over-Sampling Emotional Speech Data Based on Subjective Evaluations Provided by Multiple Individuals
abstract
A common step in the area of speech emotion recognition is to obtain ground-truth labels describing the emotional content of a sentence. The underlying emotion of a given recording is usually unknown, so perceptual evaluations are conducted to annotate its perceived emotion. Each sentence is often annotated by multiple raters, which are aggregated with methods such as majority vote rules. This paper argues that several labels provided by different individuals convey more information than the consensus labels. We demonstrate that leveraging the information provided by separate evaluations collected by multiple raters can help in building more robust classifiers which maximize the utilization of labeled data. Motivated by thesynthetic minority over-sampling technique(SMOTE), we present a novel over-sampling approach during training, where the samples with categorical emotion labels are over-sampled according to the labels assigned by multiple individuals. This approach (1) increases the number of sentences from classes with underrepresented consensus labels, and (2) utilizes sentences with ambiguous emotional content even if they do not reach consensus agreement. The experimental evaluation shows the benefits of the approach over a baseline classifier trained with consensus labels, which increases the F1-score by 5.2 percent (absolute) for the USC-IEMOCAP corpus, and 5.4 percent (absolute) for the MSP-IMPROV corpus.
Reza Lotfian, Carlos Busso
IEEE Trans. Affect. Comput.2
2021 Predicting Emotionally Salient Regions Using Qualitative Agreement of Deep Neural Network Regressors
abstract
Automatic emotion recognition plays a crucial role in various fields such as healthcare, human-computer interaction (HCI) and security and defense. While most of previous studies have focused on the recognition of emotion in isolated utterances, a more natural approach is to continuously track emotions during human interaction, identifying regions that are highly emotional. This study proposes a framework to define emotionally salient regions (hotspots), which we then attempt to dynamically detect. Our proposed approach defines hotspots relying on the qualitative agreement (QA) method, which searches for trends across continuous-time evaluations provided by different raters for arousal and valence. We illustrate the benefits of the QA method over averaging absolute values of the traces without considering trends across evaluators. After defining hotspot regions, we propose a deep learning framework to automatically detect these emotional hotspots. The proposed method relies on an ensemble of bidirectional long short term memory (BLSTM) regressors, trained on individual emotional traces provided by the evaluators, which are combined to automatically detect emotional hotspots. An appealing fusion approach to combine these regressors is to rely again on the QA method, which detects emotional salient regions with F1-scores as high as 60.9 percent for arousal and 50.4 percent for valence on the RECOLA dataset.
Srinivas Parthasarathy, Carlos Busso
IEEE Trans. Affect. Comput.2
2021 Speech-Driven Expressive Talking Lips with Conditional Sequential Generative Adversarial Networks
abstract
Articulation, emotion, and personality play strong roles in the orofacial movements. To improve the naturalness and expressiveness ofvirtual agents(VAs), it is important that we carefully model the complex interplay between these factors. This paper proposes a conditional generative adversarial network, calledconditional sequential GAN(CSG), which learns the relationship between emotion, lexical content and lip movements in a principled manner. This model uses a set of spectral and emotional speech features directly extracted from the speech signal as conditioning inputs, generating realistic movements. A key feature of the approach is that it is a speech-driven framework that does not require transcripts. Our experiments show the superiority of this model over three state-of-the-art baselines in terms of objective and subjective evaluations. When the target emotion is known, we propose to create emotionally dependent models by either adapting the base model with the target emotional data (CSG-Emo-Adapted), or adding emotional conditions as the input of the model (CSG-Emo-Aware). Objective evaluations of these models show improvements for the CSG-Emo-Adapted compared with the CSG model, as the trajectory sequences are closer to the original sequences. Subjective evaluations show significantly better results for this model compared with the CSG model when the target emotion is happiness.
Najmeh Sadoughi, Carlos Busso
IEEE Trans. Affect. Comput.2
2021 The Ordinal Nature of Emotions: An Emerging Approach
abstract
Computational representation of everyday emotional states is a challenging task and, arguably, one of the most fundamental for affective computing. Standard practice in emotion annotation is to ask people to assign a value of intensity or a class value to each emotional behavior they observe. Psychological theories and evidence from multiple disciplines including neuroscience, economics and artificial intelligence, however, suggest that the task of assigning reference-based values to subjective notions is better aligned with the underlying representations. This paper draws together the theoretical reasons to favor ordinal labels for representing and annotating emotion, reviewing the literature across several disciplines. We go on to discuss good and bad practices of treating ordinal and other forms of annotation data and make the case for preference learning methods as the appropriate approach for treating ordinal labels. We finally discuss the advantages of ordinal annotation with respect to both reliability and validity through a number of case studies in affective computing, and address common objections to the use of ordinal data. More broadly, the thesis that emotions are by nature ordinal is supported by both theoretical arguments and evidence, and opens new horizons for the way emotions are viewed, represented and analyzed computationally.
Georgios N. Yannakakis, Roddy Cowie, Carlos Busso
IEEE Trans. Affect. Comput.3
2021 Guided Generative Adversarial Neural Network for Representation Learning and Audio Generation Using Fewer Labelled Audio Data
abstract
The Generation power of Generative Adversarial Neural Networks (GANs) has shown great promise to learn representations from unlabelled data while guided by a small amount of labelled data. We aim to utilise the generation power of GANs to learn Audio Representations. Most existing studies are, however, focused on images. Some studies use GANs for speech generation, but they are conditioned on text or acoustic features, limiting their use for other audio, such as instruments, and even for speech where transcripts are limited. This paper proposes a novel GAN-based model that we named Guided Generative Adversarial Neural Network (GGAN), which can learn powerful representations and generate good-quality samples using a small amount of labelled data as guidance. Experimental results based on a speech [Speech Command Dataset (S09)] and a non-speech [Musical Instrument Sound dataset (Nsyth)] dataset demonstrate that using only 5% of labelled data as guidance, GGAN learns significantly better representations than the state-of-the-art models.
Kazi Nazmul Haque, Rajib Rana, Jiajun Liu 0013, John H. L. Hansen, Nicholas Cummins, Carlos Busso, Björn W. Schuller
IEEE ACM Trans. Audio Speech Lang. Process.6
2021 End-to-End Audiovisual Speech Recognition System With Multitask Learning
abstract
An automatic speech recognition (ASR) system is a key component in current speech-based systems. However, the surrounding acoustic noise can severely degrade the performance of an ASR system. An appealing solution to address this problem is to augment conventional audio-based ASR systems with visual features describing lip activity. This paper proposes a novel end-to-end, multitask learning (MTL), audiovisual ASR (AV-ASR) system. A key novelty of the approach is the use of MTL, where the primary task is AV-ASR, and the secondary task is audiovisual voice activity detection (AV-VAD). We obtain a robust and accurate audiovisual system that generalizes across conditions. By detecting segments with speech activity, the AV-ASR performance improves as its connectionist temporal classification (CTC) loss function can leverage from the AV-VAD alignment information. Furthermore, the end-to-end system learns from the raw audiovisual inputs a discriminative high-level representation for both speech tasks, providing the flexibility to mine information directly from the data. The proposed architecture considers the temporal dynamics within and across modalities, providing an appealing and practical fusion scheme. We evaluate the proposed approach on a large audiovisual corpus (over 60 hours), which contains different channel and environmental conditions, comparing the results with competitive single task learning (STL) and MTL baselines. Although our main goal is to improve the performance of our ASR task, the experimental results show that the proposed approach can achieve the best performance across all conditions for both speech tasks. In addition to state-of-the-art performance in AV-ASR, the proposed solution can also provide valuable information about speech activity, solving two of the most important tasks in speech-based applications.
Fei Tao 0003, Carlos Busso
IEEE Trans. Multim.2
2020 Dynamic versus Static Facial Expressions in the Presence of Speech
abstract
Face analysis is an important area in affective computing. While studies have reported important progress in detecting emotions from still images, an open challenge is to determine emotions from videos, leveraging the dynamic nature in the externalization of emotions. A common approach in earlier studies is to individually process each frame of a video, aggregating the results obtained across frames. This study questions this approach, especially when the subjects are speaking. Speech articulation affects the face appearance, which may lead to misleading emotional perceptions when the isolated frames are taken out-of-context. The analysis in this study explores the similarities and differences in emotion perceptions between (1) videos of speaking segments (without audio), and (2) isolated frames from the same videos evaluated out-of-context. We consider the emotions happiness, sadness, anger and neutral state, and emotional attributes valence, arousal, and dominance using the MSP-IMPROV corpus. The results consistently reveal that the emotional perception of static representations of emotion in isolated frames is significantly different from the overall emotional perception of dynamic representation in videos in the presence of speech. The results reveal the intrinsic limitations of the common frame-by-frame analysis of videos, highlighting the importance of explicitly modeling temporal and lexical information in face emotion recognition from videos.
Ali N. Salman, Carlos Busso
FG2
2020 Modeling Uncertainty in Predicting Emotional Attributes from Spontaneous Speech
abstract
A challenging task in affective computing is to build reliable speech emotion recognition (SER) systems that can accurately predict emotional attributes from spontaneous speech. To increase the trust in these SER systems, it is important to predict not only their accuracy, but also their confidence. An intriguing approach to predict uncertainty is Monte Carlo (MC) dropout, which obtains predictions from multiple feed-forward passes through a deep neural network (DNN) by using dropout regularization in both training and inference. This study evaluates this approach with regression models to predict emotional attribute scores for valence, arousal and dominance. The analysis illustrates that predicting uncertainty in this problem is possible, where the performance is higher for samples in the test set with lower uncertainty. The study evaluates uncertainty estimation as a function of the emotional attributes, showing that samples with extreme values have lower uncertainty. Finally, we demonstrate the benefits of uncertainty estimation with reject option, where a classifier can decline to give a prediction when its confidence is low. By rejecting only 25% of the test set with the highest uncertainty, we achieve relative performance gains of 7.34% for arousal, 13.73% for valence and 8.79% for dominance.
Kusha Sridhar, Carlos Busso
ICASSP2
2020 Style Extractor For Facial Expression Recognition in the Presence of Speech
abstract
The performance of facial expression recognition (FER) systems has improved with recent advances in machine learning. While studies have reported impressive accuracies in detecting emotion from posed expressions in static images, there are still important challenges in developing FER systems for videos, especially in the presence of speech. Speech articulation modulates the orofacial area, changing the facial appearance. These facial movements induced by speech introduce noise, reducing the performance of an FER system. Solving this problem is important if we aim to study more naturalistic environment or applications in the wild. We propose a novel approach to compensate for lexical information that does not require phonetic information during inference. The approach relies on a style extractor model, which creates emotional-to-neutral transformations. The transformed facial representations are spatially contrasted with the original faces, highlighting the emotional information conveyed in the video. The results demonstrate that adding the proposed style extractor model to a dynamic FER system improves the performance by 7% (absolute) compared to a similar model with no style extractor. This novel feature representation also improves the generalization of the model.
Ali N. Salman, Carlos Busso
ICIP2
2020 MSP-Face Corpus: A Natural Audiovisual Emotional Database
abstract
Expressive behaviors conveyed during daily interactions are difficult to determine, because they often consist of a blend of different emotions. The complexity in expressive human communication is an important challenge to build and evaluate automatic systems that can reliably predict emotions. Emotion recognition systems are often trained with limited databases, where the emotions are either elicited or recorded by actors. These approaches do not necessarily reflect real emotions, creating a mismatch when the same emotion recognition systems are applied to practical applications. Developing rich emotional databases that reflect the complexity in the externalization of emotion is an important step to build better models to recognize emotions. This study presents the MSP-Face database, a natural audiovisual database obtained from video-sharing websites, where multiple individuals discuss various topics expressing their opinions and experiences. The natural recordings convey a broad range of emotions that are difficult to obtain with other alternative data collection protocols. A feature of the corpus is the addition of two sets. The first set includes videos that have been annotated with emotional labels using a crowd-sourcing protocol (9,370 recordings -- 24 hrs, 41 m). The second set includes similar videos without emotional labels (17,955 recordings -- 45 hrs, 57 m), offering the perfect infrastructure to explore semi-supervised and unsupervised machine-learning algorithms on natural emotional videos. This study describes the process of collecting and annotating the corpus. It also provides baselines over this new database using unimodal (audio, video) and multimodal emotional recognition systems.
Andrea Vidal, Ali N. Salman, Carlos Busso
ICMI4
2020 An Efficient Temporal Modeling Approach for Speech Emotion Recognition by Mapping Varied Duration Sentences into Fixed Number of Chunks
Carlos Busso
INTERSPEECH2
2020 The MSP-Conversation Corpus
Luz Martinez-Lucas, Mohammed Abdel-Wahab 0001, Carlos Busso
INTERSPEECH3
2020 Ensemble of Students Taught by Probabilistic Teachers to Improve Speech Emotion Recognition
Kusha Sridhar, Carlos Busso
INTERSPEECH2
2020 Robust Driver Head Pose Estimation in Naturalistic Conditions from Point-Cloud Data
abstract
Head pose estimation has been a key task in computer vision since a broad range of applications often requires accurate information about the orientation of the head. Achieving this goal with regular RGB cameras faces challenges in automotive applications due to occlusions, extreme head poses and sudden changes in illumination. Most of these challenges can be attenuated with algorithms relying on depth cameras. This paper proposes a novel point-cloud based deep learning approach to estimate the driver's head pose from depth camera data, addressing these challenges. The proposed algorithm is inspired by the PointNet++ framework, where points are sampled and grouped before extracting discriminative features. We demonstrate the effectiveness of our algorithm by evaluating our approach on a naturalistic driving database consisting of 22 drivers, where the benchmark for the orientation of the driver's head is obtained with the Fi-Cap device. The experimental evaluation demonstrates that our proposed approach relying on point-cloud data achieves predictions that are almost always more reliable than state-of-the-art head pose estimation methods based on regular cameras. Furthermore, our approach provides predictions even for extreme rotations, which is not the case for the baseline methods. To the best of our knowledge, this is the first study to propose head pose estimation using deep learning on point-cloud data.
Tiancheng Hu, Carlos Busso
IV3
2020 Semi-Supervised Speech Emotion Recognition With Ladder Networks
abstract
Speech emotion recognition (SER) systems find applications in various fields such as healthcare, education, and security and defense. A major drawback of these systems is their lack of generalization across different conditions. For example, systems that show superior performance on certain databases show poor performance when tested on other corpora. This problem can be solved by training models on large amounts of labeled data from the target domain, which is expensive and time-consuming. Another approach is to increase the generalization of the models. An effective way to achieve this goal is by regularizing the models through multitask learning (MTL), where auxiliary tasks are learned along with the primary task. These methods often require the use of labeled data which is computationally expensive to collect for emotion recognition (gender, speaker identity, age or other emotional descriptors). This study proposes the use of ladder networks for emotion recognition, which utilizes an unsupervised auxiliary task. The primary task is a regression problem to predict emotional attributes. The auxiliary task is the reconstruction of intermediate feature representations using a denoising autoencoder. This auxiliary task does not require labels so it is possible to train the framework in a semi-supervised fashion with abundant unlabeled data from the target domain. This study shows that the proposed approach creates a powerful framework for SER, achieving superior performance than fully supervised single-task learning (STL) and MTL baselines. We implement the approach with sentence-level or frame-level features, demonstrating the flexibility of our approach. Additionally, the generalization of the ladder networks is evaluated in cross-corpus settings using sentence-level features, obtaining important improvements. Compared to the STL baselines, the proposed approach achieves relative gains in concordance correlation coefficient (CCC) between 3.0% and 3.5% for within corpus evaluations, and between 16.1% and 74.1% for cross corpus evaluations, highlighting the power of the architecture.
Srinivas Parthasarathy, Carlos Busso
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 Active Learning for Speech Emotion Recognition Using Deep Neural Network
abstract
Deep neural networks (DNNs) have consistently pushed the state-of-the-art performance in many fields, including speech emotion recognition. However, DNN-based solutions require vast amounts of labeled data for training. In speech emotion recognition, the cost and time needed to annotate data with emotional labels can be prohibitive. The available corpora normally have a few thousand recordings collected by a limited number of speakers. As a result, models trained on such corpora fail to generalize to samples from new domains. This study explores practical solutions to train DNNs for speech emotion recognition with limited resources by using active learning (AL). We assume that data without emotional labels from a new domain are available and we have resources to select a limited number of recordings to be annotated with emotional labels. We actively select samples using greedy sampling (GS) and uncertainty-based methods, evaluating the performance on regression problems where the goal is to predict scores for arousal and valence. We show that the use of active learning leads to competitive performance with limited training data.
Mohammed Abdel-Wahab 0001, Carlos Busso
ACII2
2019 Retrieving Speech Samples with Similar Emotional Content Using a Triplet Loss Function
abstract
The ability to identify speech with similar emotional content is valuable to many applications, including speech retrieval, surveillance, and emotional speech synthesis. While current formulations in speech emotion recognition based on classification or regression are not appropriate for this task, solutions based on preference learning offer appealing approaches for this task. This paper aims to find speech samples that are emotionally similar to an anchor speech sample provided as a query. This novel formulation opens interesting research questions. How well can a machine complete this task? How does the accuracy of automatic algorithms compare to the performance of a human performing this task? This study addresses these questions by training a deep learning model using a triplet loss function, mapping the acoustic features into an embedding that is discriminative for this task. The network receives an anchor speech sample and two competing speech samples, and the task is to determine which of the candidate speech sample conveys the closest emotional content to the emotion conveyed by the anchor. By comparing the results from our model with human perceptual evaluations, this study demonstrates that the proposed approach has performance very close to human performance in retrieving samples with similar emotional content.
John B. Harvill, Mohammed Abdel-Wahab 0001, Reza Lotfian, Carlos Busso
ICASSP4
2019 Estimation of Gaze Region Using Two Dimensional Probabilistic Maps Constructed Using Convolutional Neural Networks
abstract
Predicting the gaze of a user can have important applications in human computer interactions (HCI). They find applications in areas such as social interaction, driver distraction, human robot interaction and education. Appearance based models for gaze estimation have significantly improved due to recent advances in convolutional neural network (CNN). This paper proposes a method to predict the gaze of a user with deep models purely based on CNNs. A key novelty of the proposed model is that it produces a probabilistic map describing the gaze distribution (as opposed to predicting a single gaze direction). This approach is achieved by converting the regression problem into a classification problem, predicting the probability at the output instead of a single direction. The framework relies in a sequence of downsampling followed by upsampling to obtain the probabilistic gaze map. We observe that our proposed approach works better than a regression model in terms of prediction accuracy. The average mean squared error between the predicted gaze and the true gaze is observed to be 6.89° in a model trained and tested on the MSP-Gaze database, without any calibration or adaptation to the target user.
Carlos Busso
ICASSP2
2019 Driving Anomaly Detection with Conditional Generative Adversarial Network using Physiological and CAN-Bus Data
abstract
New developments in advanced driver assistance systems (ADAS) can help drivers deal with risky driving maneuvers, preventing potential hazard scenarios. A key challenge in these systems is to determine when to intervene. While there are situations where the needs for intervention or feedback is clear (e.g., lane departure), it is often difficult to determine scenarios that deviate from normal driving conditions. These scenarios can appear due to errors by the drivers, presence of pedestrian or bicycles, or maneuvers from other vehicles. We formulate this problem as a driving anomaly detection, where the goal is to automatically identify cases that require intervention. Towards addressing this challenging but important goal, we propose a multimodal system that considers (1) physiological signals from the driver, and (2) vehicle information obtained from the controller area network (CAN) bus sensor. The system relies on conditional generative adversarial networks (GAN) where the models are constrained by the signals previously observed. The difference of the scores in the discriminator between the predicted and actual signals is used as a metric for detecting driving anomalies. We collected and annotated a novel dataset for driving anomaly detection tasks, which is used to validate our proposed models. We present the analysis of the results, and perceptual evaluations which demonstrate the discriminative power of this unsupervised approach for detecting driving anomalies.
Yuning Qiu, Teruhisa Misu, Carlos Busso
ICMI3
2019 Speech Emotion Recognition with a Reject Option
Kusha Sridhar, Carlos Busso
INTERSPEECH2
2019 Speech-driven animation with meaningful behaviors
Najmeh Sadoughi, Carlos Busso
Speech Commun.2
2019 End-to-end audiovisual speech activity detection with bimodal recurrent neural models
Fei Tao 0003, Carlos Busso
Speech Commun.2
2019 Building Naturalistic Emotionally Balanced Speech Corpus by Retrieving Emotional Speech from Existing Podcast Recordings
abstract
The lack of a large, natural emotional database is one of the key barriers to translate results on speech emotion recognition in controlled conditions into real-life applications. Collecting emotional databases is expensive and time demanding, which limits the size of existing corpora. Current approaches used to collect spontaneous databases tend to provide unbalanced emotional content, which is dictated by the given recording protocol (e.g., positive for colloquial conversations, negative for discussion or debates). The size and speaker diversity are also limited. This paper proposes a novel approach to effectively build a large, naturalistic emotional database with balanced emotional content, reduced cost and reduced manual labor. It relies on existing spontaneous recordings obtained from audio-sharing websites. The proposed approach combines machine learning algorithms to retrieve recordings conveying balanced emotional content with a cost effective annotation process using crowdsourcing, which make it possible to build a large scale speech emotional database. This approach provides natural emotional renditions from multiple speakers, with different channel conditions and conveying balanced emotional content that are difficult to obtain with alternative data collection protocols.
Reza Lotfian, Carlos Busso
IEEE Trans. Affect. Comput.2
2019 Curriculum Learning for Speech Emotion Recognition From Crowdsourced Labels
abstract
This study introduces a method to design a curriculum for machine-learning to maximize the efficiency during the training process of deep neural networks (DNNs) for speech emotion recognition. Previous studies in other machine-learning problems have shown the benefits of training a classifier following a curriculum where samples are gradually presented in increasing level of difficulty. For speech emotion recognition, the challenge is to establish a natural order of difficulty in the training set to create the curriculum. We address this problem by assuming that, ambiguous samples for humans are also ambiguous for computers. Speech samples are often annotated by multiple evaluators to account for differences in emotion perception across individuals. While some sentences with clear emotional content are consistently annotated, sentences with more ambiguous emotional content present important disagreement between individual evaluations. We propose to use the disagreement between evaluators as a measure of difficulty for the classification task. We propose metrics that quantify the inter-evaluation agreement to define the curriculum for regression problems and binary and multi-class classification problems. The experimental results consistently show that relying on a curriculum based on agreement between human judgments leads to statistically significant improvements over baselines trained without a curriculum.
Reza Lotfian, Carlos Busso
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Expressive Speech-Driven Lip Movements with Multitask Learning
abstract
The orofacial area conveys a range of information, including speech articulation and emotions. These two factors add constraints to the facial movements, creating non-trivial integrations and interplays. To generate more expressive and naturalistic movements for conversational agents (CAs) the relationship between these factors should be carefully modeled. Data-driven models are more appropriate for this task than rule-based systems. This paper provides two deep learning speech-driven structures to integrate speech articulation and emotional cues. The proposed approaches rely on multitask learning (MTL) strategies, where related secondary tasks are jointly solved when synthesizing orofacial movements. In particular, we evaluate emotion recognition and viseme recognition as secondary tasks. The approach creates shared representations that generate behaviors that not only are closer to the original orofacial movements, but also are perceived more natural than the results from single task learning.
Najmeh Sadoughi, Carlos Busso
FG2
2018 Study of Dense Network Approaches for Speech Emotion Recognition
abstract
Deep neural networks have been proven to be very effective in various classification problems and show great promise for emotion recognition from speech. Studies have proposed various architectures that further improve the performance of emotion recognition systems. However, there are still various open questions regarding the best approach to building a speech emotion recognition system. Would the system's performance improve if we have more labeled data? How much do we benefit from data augmentation? What activation and regularization schemes are more beneficial? How does the depth of the network affect the performance? We are collecting the MSP-Podcast corpus, a large dataset with over 30 hours of data, which provides an ideal resource to address these questions. This study explores various dense architectures to predict arousal, valence and dominance scores. We investigate varying the training set size, width, and depth of the network, as well as the activation functions used during training. We also study the effect of data augmentation on the network's performance. We find that bigger training set improves the performance. Batch normalization is crucial to achieving a good performance for deeper networks. We do not observe significant differences in the performance in residual networks compared to dense networks.
Mohammed Abdel-Wahab 0001, Carlos Busso
ICASSP2
2018 Novel Realizations of Speech-Driven Head Movements with Generative Adversarial Networks
abstract
Head movement is an integral part of face-to-face communications. It is important to investigate methodologies to generate naturalistic movements for conversational agents (CAs). The predominant method for head movement generation is using rules based on the meaning of the message. However, the variations of head movements by these methods are bounded by the predefined dictionary of gestures. Speech-driven methods offer an alternative approach, learning the relationship between speech and head movements from real recordings. However, previous studies do not generate novel realizations for a repeated speech signal. Conditional generative adversarial network (GAN) provides a framework to generate multiple realizations of head movements for each speech segment by sampling from a conditioned distribution. We build a conditional GAN with bidirectional long-short term memory (BLSTM), which is suitable for capturing the long-short term dependencies of time-continuous signals. This model learns the distribution of head movements conditioned on speech prosodic features. We compare this model with a dynamic Bayesian network (DBN) and BLSTM models optimized to reduce mean squared error (MSE) or to increase concordance correlation. The objective evaluations and subjective evaluations of the results showed better performance for the conditional GAN model compared with these baseline systems.
Najmeh Sadoughi, Carlos Busso
ICASSP2
2018 FI-CAP: Robust Framework to Benchmark Head Pose Estimation in Challenging Environments
abstract
Head pose estimation is challenging in a naturalistic environment. To effectively train machine-learning algorithms, we need datasets with reliable ground truth labels from diverse environments. We present Fi-Cap, a helmet with fiducial markers designed for head pose estimation. The relative position and orientation of the tags from a reference camera can be automatically obtained from a subset of the tags. Placed at the back of the head, it provides a reference system without interfering with sensors that record frontal face. We quantify the performance of the Fi-Cap by (1) rendering the 3D model of the design, evaluating its accuracy under various rotation, image resolution and illumination conditions, and (2) comparing the predicted head pose with the location of the projected beam of a laser mounted on glasses worn by the subjects in controlled experiments conducted in our laboratory. Fi-Cap provides ideal benchmark information to evaluate automatic algorithms and alternative sensors for head pose estimation in a variety of challenging environments, including our target application for advanced driver assistance systems (ADAS).
Carlos Busso
ICME2
2018 Aligning Audiovisual Features for Audiovisual Speech Recognition
abstract
Visual information can improve the performance of automatic speech recognition (ASR), especially in the presence of background noise or different speech modes. A key problem is how to fuse the acoustic and visual features leveraging their complementary information and overcoming the alignment differences between modalities. Current audiovisual ASR (AV-ASR) systems rely on linear interpolation or extrapolation as a pre-processing technique to align audio and visual features, assuming that the feature sequences are aligned frame-by-frame. These pre-processing methods oversimplify the phase difference between lip motion and speech, lacking flexibility and impairing the performance of the system. This paper addresses the fusion of audiovisual features with an alignment neural network (AliNN), relying on recurrent neural network (RNN) with attention model. The proposed front-end model can automatically learn the alignment from the data. The resulting aligned features are concatenated and fed to conventional back-end ASR systems. The proposed front-end system is evaluated with matched and mismatch channel conditions, under clean and noisy recordings. The results show that our proposed approach can relatively outperform the baseline by 24.9% with Gaussian mixture model with hidden Markov model (GMM-HMM) back-end and 2.4% with deep neural network with hidden Markov model (DNN-HMM) back-end.
Fei Tao 0003, Carlos Busso
ICME2
2018 Predicting Categorical Emotions by Jointly Learning Primary and Secondary Emotions through Multitask Learning
Reza Lotfian, Carlos Busso
INTERSPEECH2
2018 Preference-Learning with Qualitative Agreement for Sentence Level Emotional Annotations
Srinivas Parthasarathy, Carlos Busso
INTERSPEECH2
2018 Ladder Networks for Emotion Recognition: Using Unsupervised Auxiliary Tasks to Improve Predictions of Emotional Attributes
abstract
Recognizing emotions using few attribute dimensions such as arousal, valence and dominance provides the flexibility to effectively represent complex range of emotional behaviors. Conventional methods to learn these emotional descriptors primarily focus on separate models to recognize each of these attributes. Recent work has shown that learning these attributes together regularizes the models, leading to better feature representations. This study explores new forms of regularization by adding unsupervised auxiliary tasks to reconstruct hidden layer representations. This auxiliary task requires the denoising of hidden representations at every layer of an auto-encoder. The framework relies on ladder networks that utilize skip connections between encoder and decoder layers to learn powerful representations of emotional dimensions. The results show that ladder networks improve the performance of the system compared to baselines that individually learn each attribute, and conventional denoising autoencoders. Furthermore, the unsupervised auxiliary tasks have promising potential to be used in a semi-supervised setting, where few labeled sentences are available.
Srinivas Parthasarathy, Carlos Busso
INTERSPEECH2
2018 Role of Regularization in the Prediction of Valence from Speech
abstract
Regularization plays a key role in improving the prediction of emotions using attributes such as arousal, valence and dominance. Regularization is particularly important with deep neural networks (DNNs), which have millions of parameters. While previous studies have reported competitive performance for arousal and dominance, the prediction results for valence using acoustic features are significantly lower. We hypothesize that higher regularization can lead to better results for valence. This study focuses on exploring the role of dropout as a form of regularization for valence, suggesting the need for higher regularization. We analyze the performance of regression models for valence, arousal and dominance as a function of the dropout probability. We observe that the optimum dropout rates are consistent for arousal and dominance. However, the optimum dropout rate for valence is higher. To understand the need for higher regularization for valence, we perform an empirical analysis to explore the nature of emotional cues conveyed in speech. We compare regression models with speakerdependent and speaker-independent partitions for training and testing. The experimental evaluation suggests stronger speaker dependent traits for valence. We conclude that higher regularization is needed for valence to force the network to learn global patterns that generalize across speakers.
Kusha Sridhar, Srinivas Parthasarathy, Carlos Busso
INTERSPEECH3
2018 Audiovisual Speech Activity Detection with Advanced Long Short-Term Memory
Fei Tao 0003, Carlos Busso
INTERSPEECH2
2018 Calibration free, user-independent gaze estimation with tensor analysis
Nanxiang Li, Carlos Busso
Image Vis. Comput.2
2018 Domain Adversarial for Acoustic Emotion Recognition
abstract
The performance of speech emotion recognition is affected by the differences in data distributions between train (source domain) and test (target domain) sets used to build and evaluate the models. This is a common problem, as multiple studies have shown that the performance of emotional classifiers drops when they are exposed to data that do not match the distribution used to build the emotion classifiers. The difference in data distributions becomes very clear when the training and testing data come from different domains, causing a large performance gap between development and testing performance. Due to the high cost of annotating new data and the abundance of unlabeled data, it is crucial to extract as much useful information as possible from the available unlabeled data. This study looks into the use of adversarial multitask training to extract a common representation between train and test domains. The primary task is to predict emotional-attribute-based descriptors for arousal, valence, or dominance. The secondary task is to learn a common representation, where the train and test domains cannot be distinguished. By using a gradient reversal layer, the gradients coming from the domain classifier are used to bring the source and target domain representations closer. We show that exploiting unlabeled data consistently leads to better emotion recognition performance across all emotional dimensions. We visualize the effect of adversarial training on the feature representation across the proposed deep learning architecture. The analysis shows that the data representations for the train and test domains converge as the data are passed to deeper layers of the network. We also evaluate the difference in performance when we use a shallow neural network versus a deep neural network and the effect of the number of shared layers used by the task and domain classifiers.
Mohammed Abdel-Wahab 0001, Carlos Busso
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Gating Neural Network for Large Vocabulary Audiovisual Speech Recognition
abstract
Audio-based automatic speech recognition (A-ASR) systems are affected by noisy conditions in real-world applications. Adding visual cues to the ASR system is an appealing alternative to improve the robustness of the system, replicating the audiovisual perception process used during human interactions. A common problem observed when using audiovisual automatic speech recognition (AV-ASR) is the drop in performance when speech is clean. In this case, visual features may not provide complementary information, introducing variability that negatively affects the performance of the system. The experimental evaluation in this study clearly demonstrates this problem when we train an audiovisual state-of-the-art hybrid system with a deep neural network (DNN) and hidden Markov models (HMMs). This study proposes a framework that addresses this problem, improving, or at least, maintaining the performance when visual features are used. The proposed approach is a deep learning solution with a gating layer that diminishes the effect of noisy or uninformative visual features, keeping only useful information. The framework is implemented with a subset of the audiovisual CRSS-4ENGLISH-14 corpus which consists of 61 h of speech from 105 subjects simultaneously collected with multiple cameras and microphones. The proposed framework is compared with conventional HMMs with observation models implemented with either a Gaussian mixture model or DNNs. We also compare the system with a multi-stream HMM system. The experimental evaluation indicates that the proposed framework outperforms alternative methods under all configurations, showing the robustness of the gating-based framework for AV-ASR.
Fei Tao 0003, Carlos Busso
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Formulating emotion perception as a probabilistic model with application to categorical emotion classification
abstract
Automatic recognition of emotions is an important part of affect-sensitive human-computer interaction (HCI). Expressive behaviors tend to be ambiguous with blended emotions during natural spontaneous conversations. Therefore, evaluators disagree on the perceived emotion, assigning multiple emotional classes to the same stimuli (e.g., sadness, anger, surprise). These observations have clear implications on emotion classification, where assigning a single descriptor per stimuli oversimplifies the intrinsic subjectivity in emotion perception. This study proposes a new formulation, where the emotional perception of a stimuli is a multidimensional Gaussian random variable with an unobserved distribution. Each dimension corresponds to an emotion characterized by a numerical scale. The covariance matrix of this distribution captures the intrinsic dependencies between different emotional categories. The process where an evaluator judges the stimuli is equivalent to sampling a point from this distribution, reporting the class with the highest value. The proposed approach recursively estimates this multimodal distribution using numerical methods. The mean of the Gaussian distribution is used as a soft label to train a deep neural network (DNN). Our experimental results show that the proposed training method leads to improvements in F-score over training with (1) hard-labels based on majority vote, and (2) soft-label framework proposed by other studies.
Reza Lotfian, Carlos Busso
ACII2
2017 Predicting speaker recognition reliability by considering emotional content
abstract
Studies have shown that emotional variability in speech degrades the performance of speaker recognition tasks. Of particular interest is the error produced due to mismatch between training speaker recognition models with neutral speech and testing them with expressive speech. While previous studies have considered categorical emotions, expressive speech during human interaction conveys subtle behaviors that are better characterized with continuous descriptors (e.g., attributes such as arousal, valence, dominance). As the emotion becomes more intense, we expect the performance of speaker recognition tasks to drop. Can we define emotional regions for which the speaker recognition performance is expected to be reliable? This study focuses on automatically predicting reliable regions for speaker recognition by analyzing and predicting the emotional content. We collected a unique emotional database from 80 speakers. We estimate speaker recognition performance as a function of arousal and valence, creating regions in this space where we can reliably recognize the identity of a speaker. Then, we train speech emotion recognizers designed to predict whether the emotional content in a sentence is within the reliable region. The experimental evaluation demonstrates that sentences that are classified as reliable for speaker recognition tasks have lower equal error rate (EER) than sentences that are considered unreliable.
Srinivas Parthasarathy, Carlos Busso
ACII2
2017 The ordinal nature of emotions
abstract
Representing computationally everyday emotional states is a challenging task and, arguably, one of the most fundamental for affective computing. Standard practice in emotion annotation is to ask humans to assign an absolute value of intensity to each emotional behavior they observe. Psychological theories and evidence from multiple disciplines including neuroscience, economics and artificial intelligence, however, suggest that the task of assigning reference-based (relative) values to subjective notions is better aligned with the underlying representations than assigning absolute values. Evidence also shows that we use reference points, or else anchors, against which we evaluate values such as the emotional state of a stimulus; suggesting again that ordinal labels are a more suitable way to represent emotions. This paper draws together the theoretical reasons to favor relative over absolute labels for representing and annotating emotion, reviewing the literature across several disciplines. We go on to discuss good and bad practices of treating ordinal and other forms of annotation data, and make the case for preference learning methods as the appropriate approach for treating ordinal labels. We finally discuss the advantages of relative annotation with respect to both reliability and validity through a number of case studies in affective computing, and address common objections to the use of ordinal data. Overall, the thesis that emotions are by nature relative is supported by both theoretical arguments and evidence, and opens new horizons for the way emotions are viewed, represented and analyzed computationally.
Georgios N. Yannakakis, Roddy Cowie, Carlos Busso
ACII3
2017 Ensemble feature selection for domain adaptation in speech emotion recognition
abstract
When emotion recognition systems are used in new domains, the classification performance usually drops due to mismatches between training and testing conditions. Annotations of new data in the new domain is expensive and time demanding. Therefore, it is important to design strategies that efficiently use limited amount of new data to improve the robustness of the classification system. The use of ensembles is an attractive solution, since they can be built to perform well across different mismatches. The key challenge is to create ensembles that are diverse. This paper proposes the use of active learning along with feature selection to build a diverse ensemble that performs well in the new domain. The diversity and accuracy of the ensemble are achieved by (1) training emotional classifiers with bias toward specific emotions, (2) eliminating overlap in the feature sets of the ensemble, and (3) conducting feature selection by maximizing the performance over the new labeled data. We study various data selection criteria, and different sample sizes to determine the best approach toward building a stable diverse ensemble that generalize well on new domains.
Mohammed Abdel-Wahab 0001, Carlos Busso
ICASSP2
2017 Incremental adaptation using active learning for acoustic emotion recognition
abstract
The performance of speech emotion classifiers greatly degrade when the training conditions do not match the testing conditions. This problem is observed in cross-corpora evaluations, even when the corpora are similar. The lack of generalization is particularly problematic when the emotion classifiers are used in real applications. This study addresses this problem by combining active learning (AL) and supervised domain adaptation (DA) using an elegant approach for support vector machine (SVM). Active learning selects samples in the new domain that are used to adapt the speech classification models using domain adaptation. This paper demonstrates that we can increase the performance of the speech recognition system by incrementally adapting the models using carefully selected samples available after active learning. We propose a novel iterative fast converging incremental adaptation algorithm that only uses correctly classified samples at each iteration. This conservative framework creates sequences of smooth changes in the decision hyperplane, resulting in statistically significant improvements over conventional schemes that adapt the models at once using all the available data.
Mohammed Abdel-Wahab 0001, Carlos Busso
ICASSP2
2017 Ranking emotional attributes with deep neural networks
abstract
Studies have shown that ranking emotional attributes through preference learning methods has significant advantages over conventional emotional classification/regression frameworks. Preference learning is particularly appealing for retrieval tasks, where the goal is to identify speech conveying target emotional behaviors (e.g., positive samples with low arousal). With recent advances in deep neural networks (DNNs), this study explores whether a preference learning framework relying on deep learning can outperform conventional ranking algorithms. We use a deep learning ranker implemented with the RankNet algorithm to evaluate preference between emotional sentences in terms of dimensional attributes (arousal, valence and dominance). The results show improved performance over ranking algorithms trained with support vector machine (SVM) (i.e., RankSVM). The results are significantly better than performance reported in previous work, demonstrating the potential of RankNet to retrieve speech with target emotional behaviors.
Srinivas Parthasarathy, Reza Lotfian, Carlos Busso
ICASSP3
2017 A study of speaker verification performance with expressive speech
abstract
Expressive speech introduces variations in the acoustic features affecting the performance of speech technology such as speaker verification systems. It is important to identify the range of emotions for which we can reliably estimate speaker verification tasks. This paper studies the performance of a speaker verification system as a function of emotions. Instead of categorical classes such as happiness or anger, which have important intra-class variability, we use the continuous attributes arousal, valence, and dominance which facilitate the analysis. We evaluate an speaker verification system trained with the i-vector framework with a probabilistic linear discriminant analysis (PLDA) back-end. The study relies on a subset of the MSP-PODCAST corpus, which has naturalistic recordings from 40 speakers. We train the system with neutral speech, creating mismatches on the testing set. The results show that speaker verification errors increase when the values of the emotional attributes increase. For neutral/moderate values of arousal, valence and dominance, the speaker verification performance are reliable. These results are also observed when we artificially force the sentences to have the same duration.
Srinivas Parthasarathy, John H. L. Hansen, Carlos Busso
ICASSP4
2017 A Stepwise Analysis of Aggregated Crowdsourced Labels Describing Multimodal Emotional Behaviors
Alec Burmania, Carlos Busso
INTERSPEECH2
2017 Jointly Predicting Arousal, Valence and Dominance with Multi-Task Learning
Srinivas Parthasarathy, Carlos Busso
INTERSPEECH2
2017 Bimodal Recurrent Neural Network for Audiovisual Voice Activity Detection
Fei Tao 0003, Carlos Busso
INTERSPEECH2
2017 Joint Learning of Speech-Driven Facial Motion with Bidirectional Long-Short Term Memory
Najmeh Sadoughi, Carlos Busso
IVA2
2017 Assessment and classification of singing quality based on audio-visual features
abstract
The process of speech production changes between speaking and singing due to excitation, vocal tract articulatory positioning, and cognitive motor planning while singing. Singing does not only deviate from typical spoken speech, but it varies across various styles of singing. This is due to alternative genres of music, singing quality of an individual, as well as different languages and cultures. Because of this variation, it is important to establish a baseline system for differentiating between certain aspects of singing. In this study, we establish a classification system that automatically estimates singing quality of candidates from an American TV singing show based on their singing speech acoustics, lip and eye movements. We employ three classifiers that include: Logistic Regression, Naive Bayes and K-nearest neighbor (k-NN) and compare performance of each using unimodal and multimodal features. We also compare performance based on different modalities (speech, lip, eye structure). The results show that audio content performs the best, with modest gains when lip and eye content are fused. An interesting outcome is that lip and eye content achieve an 82% quality assessment while audio achieves 95%. The ability to assess singing quality from lip and eye content at this level is remarkable.
Mangona Bokshi, Fei Tao 0003, Carlos Busso, John H. L. Hansen
VCIP3
2017 Meaningful head movements driven by emotional synthetic speech
Najmeh Sadoughi, Yang Liu 0004, Carlos Busso
Speech Commun.3
2017 MSP-IMPROV: An Acted Corpus of Dyadic Interactions to Study Emotion Perception
abstract
We present the MSP-IMPROV corpus, a multimodal emotional database, where the goal is to have control over lexical content and emotion while also promoting naturalness in the recordings. Studies on emotion perception often require stimuli with fixed lexical content, but that convey different emotions. These stimuli can also serve as an instrument to understand how emotion modulates speech at the phoneme level, in a manner that controls for coarticulation. Such audiovisual data are not easily available from natural recordings. A common solution is to record actors reading sentences that portray different emotions, which may not produce natural behaviors. We propose an alternative approach in which we define hypothetical scenarios for each sentence that are carefully designed to elicit a particular emotion. Two actors improvise these emotion-specific situations, leading them to utter contextualized, non-read renditions of sentences that have fixed lexical content and convey different emotions. We describe the context in which this corpus was recorded, the key features of the corpus, the areas in which this corpus can be useful, and the emotional content of the recordings. The paper also provides the performance for speech and facial emotion classifiers. The analysis brings novel classification evaluations where we study the performance in terms of inter-evaluator agreement and naturalness perception, leveraging the large size of the audiovisual database.
Carlos Busso, Srinivas Parthasarathy, Alec Burmania, Mohammed Abdel-Wahab 0001, Najmeh Sadoughi, Emily Mower Provost
IEEE Trans. Affect. Comput.1
2017 The Cost of Dichotomizing Continuous Labels for Binary Classification Problems: Deriving a Bayesian-Optimal Classifier
abstract
Many pattern recognition problems involve characterizing samples with continuous labels instead of discrete categories. While regression models are suitable for these learning tasks, these labels are often discretized into binary classes to formulate the problem as a conventional classification task (e.g., classes with low versus high values). This methodology brings intrinsic limitations on the classification performance. The continuous labels are typically normally-distributed, with many samples close to the boundary threshold, resulting in poor classification rates. Previous studies only use the discretized labels to train binary classifiers, neglecting the original, continuous labels. This study demonstrates that, even in binary classification problems, exploiting the original labels before splitting the classes can lead to better classification performance. This work proposes an optimal classifier based on the Bayesian maximum a posterior (MAP) criterion for these problems, which effectively utilizes the real-valued labels. We derive the theoretical average performance of this classifier, which can be considered as the expected upper bound performance for the task. Experimental evaluations on synthetic and real data sets show the improvement achieved by the proposed classifier, in contrast to conventional classifiers trained with binary labels. These evaluations clearly demonstrate the optimality of the proposed classifier, and the precision of the expected upper bound obtained by our derivation.
Soroosh Mariooryad, Carlos Busso
IEEE Trans. Affect. Comput.2
2016 Tradeoff between quality and quantity of emotional annotations to characterize expressive behaviors
abstract
Emotional descriptors collected from perceptual evaluations are important in the study of emotions. Many studies on emotion recognition depend on these labels to train classifiers. The reliability of the emotion descriptors vary with the number and quality of the raters. Conducting perceptual evaluations used to be an expensive and time demanding task, resulting in emotional databases with poor labels annotated by few raters. Nowadays, crowdsourcing services have simplified the process, reducing the cost, facilitating more evaluations per stimuli. The key challenge in using crowdsourcing for perceptual evaluation is the quality which significantly varies across workers. Is it better to have multiple annotations with lower inter-evaluator agreement or to have few annotations with higher inter-evaluator agreement? This study explores this tradeoff between quality and quantity in emotional annotations to characterize expressive behaviors. The analysis relies on emotional labels from the MSP-IMPROV database, where each video was evaluated by over 20 workers. We discuss the theoretical concept of effective reliability to address this problem. We demonstrate that a reduced set of labels with higher inter-evaluator agreement can provide similar classification performance than unfiltered set of labels from multiple workers. We discuss best practices to collecting annotations for emotion recognition tasks using crowdsourcing.
Alec Burmania, Mohammed Abdel-Wahab 0001, Carlos Busso
ICASSP3
2016 Automatic composition of broadcast news summaries using rank classifiers trained with acoustic and lexical features
abstract
Research on automatic speech summarization typically focuses on optimizing objective evaluation criteria, such as the ROUGE metric, which depend on word and phrase overlaps between automatic and manually generated summary documents. However, the actual quality of the speech summarizer largely depends on how the end-users perceive the audio output. This work focuses on the task of composing summarized audio streams with the aim of improving the quality and interest perceived by the end-user. First, using crowd-sourced summary annotations on a broadcast news corpus, we train a rank-SVM classifier to learn the relative importance of each sentence in a news story. Acoustic, lexical and structural features are used for training. In addition, we investigate the perceived emotion level in each sentence to aid the summarizer in selecting interesting sentences, yielding an emotion-aware summarizer. Next, we propose several methods to combine these sentences to generate a compressed audio stream. Subjective evaluations are performed to evaluate the quality of the generated summaries on the following criterion: interest, abruptness, informativeness, attractiveness, and overall quality. The results indicate that users are most sensitive to the linguistic coherence and continuity of the audio stream.
Taufiq Hasan, Mohammed Abdel-Wahab 0001, Srinivas Parthasarathy, Carlos Busso, Yang Liu 0004
ICASSP4
2016 A multimodal analysis of synchrony during dyadic interaction using a metric based on sequential pattern mining
abstract
In human-human interaction, people tend to adapt to each other as the conversation progresses, mirroring their intonation, speech rate, fundamental frequency, word selection, hand gestures, and head movements. This phenomenon is known as synchrony, convergence, entrainment, and adaptation. Recent studies have investigated this phenomenon at different dimensions and levels for single modalities. However, the interplay between modalities at a local level to study synchrony between conversational partners is an open question. This paper studies synchrony using a multimodal approach based on sequential pattern mining in dyadic conversations. This analysis deals with both acoustic and text-based features at a local level. The proposed data-driven framework identifies frequent sequences containing events from multiple modalities that can quantify the synchrony between conversational partners (e.g., a speaker reduces speech rate when the other utters disfluencies). The evaluation relies on 90 sessions from the Fishers corpus, which comprises telephone conversations between two people. We develop a multimodal metric to quantify synchrony between conversational partners using this framework. We report initial results on this metric by comparing actual dyadic conversations with sessions artificially created by randomly pairing the speakers.
Anil Jakkam, Carlos Busso
ICASSP2
2016 Practical considerations on the use of preference learning for ranking emotional speech
abstract
A speech emotion retrieval system aims to detect a subset of data with specific expressive content. Preference learning represents an appealing framework to rank speech samples in terms of continuous attributes such as arousal and valence. The training of ranking classifiers usually requires pairwise samples where one is preferred over the other according to a specific criterion. For emotional databases, these relative labels are not available and are very difficult to collect. As an alternative, they can be derived from existing absolute emotional labels. For continuous attributes, we can create relative rankings by forming pairs with high and low values of a specific attribute which are separated by a predefined margin. This approach raises questions about efficient approaches for building such a training set, which is important to improve the performance of the emotional retrieval system. This paper analyzes practical considerations in training ranking classifiers including optimum number of pairs used during training, and the margin used to define the relative labels. We compare the preference learning approach to binary classifier and regression models. The experimental results on a spontaneous emotional database indicate that a rank-based classifier with fine-tuned parameters outperforms the other two approaches in both arousal and valence dimensions.
Reza Lotfian, Carlos Busso
ICASSP2
2016 Retrieving Categorical Emotions Using a Probabilistic Framework to Define Preference Learning Samples
Reza Lotfian, Carlos Busso
INTERSPEECH2
2016 Defining Emotionally Salient Regions Using Qualitative Agreement Method
Srinivas Parthasarathy, Carlos Busso
INTERSPEECH2
2016 Head Motion Generation with Synthetic Speech: A Data Driven Approach
Najmeh Sadoughi, Carlos Busso
INTERSPEECH2
2016 A Portable Automatic PA-TA-KA Syllable Detection System to Derive Biomarkers for Neurological Disorders
Fei Tao 0003, Louis Daudet, Christian Poellabauer, Sandra L. Schneider, Carlos Busso
INTERSPEECH5
2016 Improving Boundary Estimation in Audiovisual Speech Activity Detection Using Bayesian Information Criterion
Fei Tao 0003, John H. L. Hansen, Carlos Busso
INTERSPEECH3
2016 Increasing the Reliability of Crowdsourcing Evaluations Using Online Quality Assessment
abstract
Manual annotations and transcriptions have an ever-increasing importance in areas such as behavioral signal processing, image processing, computer vision, and speech signal processing. Conventionally, this metadata has been collected through manual annotations by experts. With the advent of crowdsourcing services, the scientific community has begun to crowdsource many tasks that researchers deem tedious, but can be easily completed by many human annotators. While crowdsourcing is a cheaper and more efficient approach, the quality of the annotations becomes a limitation in many cases. This paper investigates the use of reference sets with predetermined ground-truth to monitor annotators' accuracy and fatigue, all in real-time. The reference set includes evaluations that are identical in form to the relevant questions that are collected, so annotators are blind to whether or not they are being graded on performance on a specific question. We explore these ideas on the emotional annotation of the MSP-IMPROV database. We present promising results which suggest that our system is suitable for collecting accurate annotations.
Alec Burmania, Srinivas Parthasarathy, Carlos Busso
IEEE Trans. Affect. Comput.3
2016 The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for Voice Research and Affective Computing
abstract
Work on voice sciences over recent decades has led to a proliferation of acoustic parameters that are used quite selectively and are not always extracted in a similar fashion. With many independent teams working in different research areas, shared standards become an essential safeguard to ensure compliance with state-of-the-art methods allowing appropriate comparison of results across studies and potential integration and combination of extraction and recognition systems. In this paper we propose a basic standard acoustic parameter set for various areas of automatic voice analysis, such as paralinguistic or clinical speech analysis. In contrast to a large brute-force parameter set, we present a minimalistic set of voice parameters here. These were selected based on a) their potential to index affective physiological changes in voice production, b) their proven value in former studies as well as their automatic extractability, and c) their theoretical significance. The set is intended to provide a common baseline for evaluation of future research and eliminate differences caused by varying parameter sets or even different implementations of the same parameters. Our implementation is publicly available with the openSMILE toolkit. Comparative evaluations of the proposed feature set and large baseline feature sets of INTERSPEECH challenges show a high performance of the proposed set in relation to its size.
Florian Eyben, Klaus R. Scherer, Björn W. Schuller, Johan Sundberg, Elisabeth André, Carlos Busso, Laurence Devillers, Julien Epps, Petri Laukka, Shri Narayanan, Khiet P. Truong
IEEE Trans. Affect. Comput.6
2016 Facial Expression Recognition in the Presence of Speech Using Blind Lexical Compensation
abstract
During spontaneous conversations the articulation process as well as the internal emotional states influence the facial configurations. Inferring the conveyed emotions from the information presented in facial expressions requires decoupling the linguistic and affective messages in the face. Normalizing and compensating for the underlying lexical content have shown improvement in recognizing facial expressions. However, this requires the transcription and phoneme alignment information, which is not available in broad range of applications. This study uses the asymmetric bilinear factorization model to perform the decoupling of linguistic and affective information when they are not given. The emotion recognition evaluations on the IEMOCAP database show the capability of the proposed approach in separating these factors in facial expressions, yielding statistically significant performance improvements. The achieved improvement is similar to the case when the ground truth phonetic transcription is known. Similarly, experiments on the SEMAINE database using image-based features demonstrate the effectiveness of the proposed technique in practical scenarios.
Soroosh Mariooryad, Carlos Busso
IEEE Trans. Affect. Comput.2
2016 Using Agreement on Direction of Change to Build Rank-Based Emotion Classifiers
abstract
Automatic emotion recognition in realistic domains is a challenging task given the subtle expressive behaviors that occur during human interactions. The challenges start with noisy emotional descriptors provided by multiple evaluators, which are characterized by low interevaluator agreement. Studies have suggested that evaluators are more consistent in detecting qualitative relations between episodes (i.e., emotional contrasts), rather than absolute scores (i.e., the actual emotion). Based on these observations, this study explores the use of relative labels to train machine learning algorithms that can rank expressive behaviors. Instead of deriving relative labels from expensive and time-consuming subjective evaluations, the labels are extracted from existing time-continuous evaluations over expressive attributes annotated with FEELTRACE. We rely on the qualitative agreement (QA) analysis to estimate relative labels which are used to train rank-based classifiers (rankers). The experimental evaluation on the SEMAINE database demonstrates the benefits of the proposed approach. The ranking performance using the QA-based labels compare favorably against preference learning rankers trained with relative labels obtained by simply aggregating the absolute values of the emotional traces across evaluators, which is the common approach used by other studies.
Srinivas Parthasarathy, Roddy Cowie, Carlos Busso
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 Detecting Drivers' Mirror-Checking Actions and Its Application to Maneuver and Secondary Task Recognition
abstract
This study explores the feasibility of detecting drivers' mirror-checking actions using noninvasive sensors. Checking the mirrors is an important primary driving action that allows drivers to maintain their situational awareness, particularly when they are planning to turn or change lanes. Recognizing when drivers are checking the mirrors can facilitate the detection of hazard scenarios by considering contextual information (e.g., turning without checking mirrors, lack of mirror-checking actions signaling cognitive distractions, or distinction between gazes due to primary or secondary tasks). This study analyzes drivers' mirror-checking actions under various real driving conditions. We analyze the drivers' mirror-checking actions under normal conditions, as well as when the drivers are engaged in secondary tasks such as tuning the radio or operating a cell phone. We also compare mirror-checking behaviors observed during different maneuver actions: driving straight, turning, and switching lanes. This study reveals statistically significant differences in mirror-checking actions among most of the comparisons. The results suggest that mirror-checking actions can be useful indicators in recognizing drivers engaged in secondary tasks, as well as in detecting driving maneuvers. We propose to detect mirror-checking actions using features extracted from multiple noninvasive sensors (CAN-Bus and cameras facing the driver and the road). We consider three machine learning algorithms for unbalanced data sets, achieving an F-score of 91%. The recognized mirror-checking actions are used as additional features to improve the performance of secondary task detection and maneuver recognition. These promising results suggest that it is possible to detect mirror-checking actions, providing contextual information to improve new driver monitoring systems.
Nanxiang Li, Carlos Busso
IEEE Trans. Intell. Transp. Syst.2
2015 Supervised domain adaptation for emotion recognition from speech
abstract
One of the main barriers in the deployment of speech emotion recognition systems in real applications is the lack of generalization of the emotion classifiers. The recognition performance achieved in controlled recordings drops when the models are tested with different speakers, channels, environments and domain conditions. This paper explores supervised model adaptation, which can improve the performance of systems evaluated with mismatched training and testing conditions. We address the following key questions in the context of supervised adaptation for speech emotion recognition: (a) how much labeled data is needed for adaptation to achieve good performance? (b) how important is speaker diversity in the labeled set? (c) can spontaneous acted data provide similar performance than naturalistic non-acted recordings? and (d) what is the best approach to adapt the models (domain adaptation versus incremental/online training)? We address these problems by using a multi-corpus framework where the models are trained and tested with different databases. The results indicate that even small portion of data used for adaptation can significantly improve the performance. Increasing the speaker diversity in the labeled data used for adaptation does not provide significant gain in performance. Also, we observe similar performance when the classifiers are trained with naturalistic non-acted data and spontaneous acted data.
Mohammed Abdel-Wahab 0001, Carlos Busso
ICASSP2
2015 Emotion recognition using synthetic speech as neutral reference
abstract
A common approach to recognize emotion from speech is to estimate multiple acoustic features at sentence or turn level. These features are derived independent of the underlying lexical content. Studies have demonstrated that lexical dependent models improve emotion recognition accuracy. However, current practical approaches can only model small lexical units like phonemes, syllables or few key words, which limits these systems. We believe that building longer lexical models (i.e., sentence level model) is feasible by leveraging the advances in speech synthesis. Assuming that the transcript of the target speech is available, we synthesize speech conveying the same lexical information. The synthetic speech is used as a neutral reference model to contrast different acoustic features, unveiling local emotional changes. This paper introduces this novel framework and provides insights on how to compare the target and synthetic speech signals. Our evaluations demonstrate the benefits of synthetic speech as neutral reference to incorporate lexical dependencies in emotion recognition. The experimental results show that adding features derived from contrasting expressive speech with the proposed synthetic speech reference increases the accuracy in 2.1% and 2.8% (absolute) in classifying low versus high levels of arousal and valence, respectively.
Reza Lotfian, Carlos Busso
ICASSP2
2015 Adjacent Vehicle Collision Warning System using Image Sensor and Inertial Measurement Unit
abstract
Advanced driver assistance systems are the newest addition to vehicular technology. Such systems use a wide array of sensors to provide a superior driving experience. Vehicle safety and driver alert are important parts of these system. This paper proposes a driver alert system to prevent and mitigate adjacent vehicle collisions by proving warning information of on-road vehicles and possible collisions. A dynamic Bayesian network (DBN) is utilized to fuse multiple sensors to provide driver awareness. It detects oncoming adjacent vehicles and gathers ego vehicle motion characteristics using an on-board camera and inertial measurement unit (IMU). A histogram of oriented gradient feature based classifier is used to detect any adjacent vehicles. Vehicles front-rear end and side faces were considered in training the classifier. Ego vehicles heading, speed and acceleration are captured from the IMU and feed into the DBN. The network parameters were learned from data via expectation maximization(EM) algorithm. The DBN is designed to provide two type of warning to the driver, a cautionary warning and a brake alert for possible collision with other vehicles. Experiments were completed on multiple public databases, demonstrating successful warnings and brake alerts in most situations.
Asif Iqbal 0009, Carlos Busso, Nicholas R. Gans
ICMI2
2015 Retrieving Target Gestures Toward Speech Driven Animation with Meaningful Behaviors
abstract
Creating believable behaviors for conversational agents (CAs) is a challenging task, given the complex relationship between speech and various nonverbal behaviors. The two main approaches are rule-based systems, which tend to produce behaviors with limited variations compared to natural interactions, and data-driven systems, which tend to ignore the underlying semantic meaning of the message (e.g., gestures without meaning). We envision a hybrid system, acting as the behavior realization layer in rule-based systems, while exploiting the rich variation in natural interactions. Constrained on a given target gesture (e.g., head nod) and speech signal, the system will generate novel realizations learned from the data, capturing the timely relationship between speech and gestures. An important task in this research is identifying multiple examples of the target gestures in the corpus. This paper proposes a data mining framework for detecting gestures of interest in a motion capture database. First, we train One-class support vector machines (SVMs) to detect candidate segments conveying the target gesture. Second, we use dynamic time alignment kernel (DTAK) to compare the similarity between the examples (i.e., target gesture) and the given segments. We evaluate the approach for five prototypical hand and head gestures showing reasonable performance. These retrieved gestures are then used to train a speech-driven framework based on dynamic Bayesian networks (DBNs) to synthesize these target behaviors.
Najmeh Sadoughi, Carlos Busso
ICMI2
2015 An unsupervised visual-only voice activity detection approach using temporal orofacial features
abstract
Detecting the presence or absence of speech is an important step toward building robust speech-based interfaces. While previous studies have made progress on voice activity detection (VAD), the performance of these systems significantly degrades when subjects employ challenging speech modes that deviate from normal acoustic patterns (e.g., whisper speech), or in noisy/adverse conditions. An appealing approach under these conditions is visual voice activity detection (VVAD), which detects speech using features characterizing the orofacial activity. This study proposes an unsupervised approach that relies only on visual features, and, therefore, is insensitive to vocal style or time-varying background noise. This study proposes an unsupervised approach that relies on visual features. We estimate optical flow variance and geometrical features around lips, extracting the short-time zero crossing rates, short-time variances, and delta features over a small temporal window. These variables are fused using principal component analysis (PCA) to obtain a “combo” feature, which displays a bimodal distributions (speech versus silence). A threshold is automatically determine using the expectation-maximization (EM) algorithm. The approach can be easily transformed into a supervised VVAD, if needed. We evaluate the system in neutral and whisper speech. While speech based VADs generally fail to detect speech activity in whisper speech, given its important acoustic differences, the proposed VVAD achieves near 80% accuracy in both neutral and whisper speech, highlighting the benefits of the system. Index Terms: Visual voice activity detection, whisper speech
Fei Tao 0003, John H. L. Hansen, Carlos Busso
INTERSPEECH3
2015 Correcting Time-Continuous Emotional Labels by Modeling the Reaction Lag of Evaluators
abstract
An appealing scheme to characterize expressive behaviors is the use of emotional dimensions such as activation (calm versus active) and valence (negative versus positive). These descriptors offer many advantages to describe the wide spectrum of emotions. Due to the continuous nature of fast-changing expressive vocal and gestural behaviors, it is desirable to continuously track these emotional traces, capturing subtle and localized events (e.g., with FEELTRACE). However, time-continuous annotations introduce challenges that affect the reliability of the labels. In particular, an important issue is the evaluators' reaction lag caused by observing, appraising, and responding to the expressive behaviors. An empirical analysis demonstrates that this delay varies from 1 to 6 seconds, depending on the annotator, expressive dimension, and actual behaviors. Our experiments show accuracy improvements even with fixed delays (1-3 seconds). This paper proposes to compensate for this reaction lag by finding the time-shift that maximizes the mutual information between the expressive behaviors and the time-continuous annotations. The approach is implemented by making different assumptions about the evaluators' reaction lag. The benefits of compensating for the delay is demonstrated with emotion classification experiments. On average, the classifiers trained with facial and speech features show more than 7 percent relative improvements over baseline classifiers trained and tested without shifting the time-continuous annotations.
Soroosh Mariooryad, Carlos Busso
IEEE Trans. Affect. Comput.2
2015 UMEME: University of Michigan Emotional McGurk Effect Data Set
abstract
Emotion is central to communication; it colors our interpretation of events and social interactions. Emotion expression is generally multimodal, modulating our facial movement, vocal behavior, and body gestures. The method through which this multimodal information is integrated and perceived is not well understood. This knowledge has implications for the design of multimodal classification algorithms, affective interfaces, and even mental health assessment. We present a novel data set designed to support research into the emotion perception process, the University of Michigan Emotional McGurk Effect Data set (UMEME). UMEME has a critical feature that differentiates it from currently existing data sets; it contains not only emotionally congruent stimuli (emotionally matched faces and voices), but also emotionally incongruent stimuli (emotionally mismatched faces and voices). The inclusion of emotionally complex and dynamic stimuli provides an opportunity to study how individuals make assessments of emotion content in the presence of emotional incongruence, or emotional noise. We describe the collection, annotation, and statistical properties of the data and present evidence illustrating how audio and video interact to result in specific types of emotion perception. The results demonstrate that there exist consistent patterns underlying emotion evaluation, even given incongruence, positioning UMEME as an important new tool for understanding emotion perception.
Emily Mower Provost, Yuan Shangguan, Carlos Busso
IEEE Trans. Affect. Comput.3
2015 Predicting Perceived Visual and Cognitive Distractions of Drivers With Multimodal Features
abstract
A driver's behaviors can be affected by visual, cognitive, auditory, and manual distractions. While it is important to identify the patterns associated with particular secondary tasks, it is more general and useful to define distraction modes that capture the general behaviors induced by various sources of distractions. By explicitly modeling the distinction between types of distractions, we can assess the detrimental effects induced by new in-vehicle technology. This study investigates drivers' behaviors associated with visual and cognitive distractions, both separately and jointly. External observers assessed the perceived cognitive and visual distractions from real-world driving recordings, showing high interevaluator agreement in both dimensions. The scores from the perceptual evaluation are used to define regression models with elastic net regularization and binary classifiers to separately estimate the cognitive and visual distraction levels. The analysis reveals multimodal features that are discriminative of cognitive and visual distractions. Furthermore, the study proposes a novel joint visual-cognitive distraction space to characterize driver behaviors. A data-driven clustering approach identifies four distraction modes that provide insights to better understand the deviation in driving behaviors induced by secondary tasks. Binary and multiclass recognition problems demonstrate the effectiveness of the proposed multimodal features to infer these distraction modes defined in the visual-cognitive space.
Nanxiang Li, Carlos Busso
IEEE Trans. Intell. Transp. Syst.2
2014 User Independent Gaze Estimation by Exploiting Similarity Measures in the Eye Pair Appearance Eigenspace
abstract
The design of gaze-based computer interfaces has been an active research area for over 40 years. One challenge of using gaze detectors is the repetitive calibration process required to adjust the parameters of the systems, and the constrained conditions imposed on the user for robust gaze estimation. We envision user-independent gaze detectors that do not require calibration, or any cooperation from the user. Toward this goal, we investigate an appearance-based approach, where we estimate the eigenspace for the gaze using principal component analysis (PCA). The projections are used as features of regression models that estimate the screen's coordinates. As expected, the performance of the approach decreases when the models are trained without data from the target user (i.e., user-independent condition). This study proposes an appealing training approach to bridge the gap in performance between user-dependent and user-independent conditions. Using the projections onto the eigenspace, the scheme identifies samples in training set that are similar to the testing images. We build the sample covariance matrix and the regression models only with these samples. We consider either similar frames or data from subjects with similar eye appearance. The promising results suggest that the proposed training approach is a feasible and convenient scheme for gaze-based multimodal interfaces.
Nanxiang Li, Carlos Busso
ICMI2
2014 Speech-Driven Animation Constrained by Appropriate Discourse Functions
abstract
Conversational agents provide powerful opportunities to interact and engage with the users. The challenge is how to create naturalistic behaviors that replicate the complex gestures observed during human interactions. Previous studies have used rule-based frameworks or data-driven models to generate appropriate gestures, which are properly synchronized with the underlying discourse functions. Among these methods, speech-driven approaches are especially appealing given the rich information conveyed on speech. It captures emotional cues and prosodic patterns that are important to synthesize behaviors (i.e., modeling the variability and complexity of the timings of the behaviors). The main limitation of these models is that they fail to capture the underlying semantic and discourse functions of the message (e.g., nodding). This study proposes a speech-driven framework that explicitly model discourse functions, bridging the gap between speech-driven and rule-based models. The approach is based on dynamic Bayesian Network (DBN), where an additional node is introduced to constrain the models by specific discourse functions. We implement the approach by synthesizing head and eyebrow motion. We conduct perceptual evaluations to compare the animations generated using the constrained and unconstrained models.
Najmeh Sadoughi, Yang Liu 0004, Carlos Busso
ICMI3
2014 Building a naturalistic emotional speech corpus by retrieving expressive behaviors from existing speech corpora
abstract
A key element in affective computing is to have large corpora of genuine emotional samples collected during natural conversations. Recording natural interactions through telephone is an appealing approach to build emotional databases. However, collecting real conversational data with expressive reactions is a challenging task, especially if the recordings are to be shared with the community (e.g., privacy concerns). This study explores a novel approach consisting in retrieving emotional reactions from existing spontaneous speech databases collected for general speech processing problems. Although most of the recordings in these databases are expected to have non-emotional expressions, given the naturalness of the interactions, the flow of the conversation can lead to emotional responses from conversation partners which we aim to retrieve. We use the IEMOCAP and SEMAINE databases to build emotion detector systems. We use these classifiers to identify emotional behaviors from the FISHER database, which is a large conversational speech corpus recorded over the phone. Subjective evaluations over the retrieved samples demonstrate the potential of the proposed scheme to build naturalistic emotional speech database. Index Terms: emotion recognition, expressive speech, information retrieval, emotional databases
Soroosh Mariooryad, Reza Lotfian, Carlos Busso
INTERSPEECH3
2014 Lipreading approach for isolated digits recognition under whisper and neutral speech
abstract
Whisper is a speech production mode normally used to protect confidential information. Given the differences in the acoustic domain, the performance of automatic speech recognition (ASR) systems decreases with whisper speech. An appealing approach to improve the performance is the use of lipreading. This study explores the use of visual features characterizing the lips’ geometry and appearance to recognize digits under normal and whisper speech conditions using hidden Markov models (HMMs). We evaluate the proposed features on the digit part of the audiovisual whisper (AVW) corpus. While the proposed system achieves high accuracy in speaker dependent conditions (80.8%), the performance decreases when we evaluate speaker independent models (52.9%). We propose supervised adaptation schemes to reduce the mismatch between speakers. Across all conditions, the performance of the classifiers remain competitive even in the presence of whisper speech, highlighting the benefits of using visual features. Index Terms: Lipreading, whisper speech, multimodal corpus
Fei Tao 0003, Carlos Busso
INTERSPEECH2
2014 Evaluation of syllable rate estimation in expressive speech and its contribution to emotion recognition
abstract
It is commonly accepted that speaking rate is an important aspect characterizing expressive speech. The speaking rate increases for emotions such as happiness and anger, and decreases for emotions such as sadness. In spite of these observations, most of the current speech emotion classifiers do not explicitly use speaking rate features. This study explores two interrelated questions to evaluate the role of speaking rate in emotion recognition: Can we reliably estimate syllable rate from emotional speech? Does syllable rate provide complementary emotional information over other acoustic features? We consider two syllable rate estimation algorithms, as well as reference values derived from forced alignment. We evaluate the performance of these syllable rate estimation methods in expressive speech (SEMAINE database). The analysis reveals a drop in performance as the intensity of the emotion increases. Next, we conduct emotion recognition experiments to evaluate the contribution of syllable rate in recognizing emotions. The emotion classification experiments demonstrate that features conveying accurate syllable rate estimations complement features that are commonly used in current emotion recognition system.
Mohammed Abdel-Wahab 0001, Carlos Busso
SLT2
2014 Shape-based modeling of the fundamental frequency contour for emotion detection in speech
Juan Pablo Arias, Carlos Busso, Néstor Becerra Yoma
Comput. Speech Lang.2
2014 Compensating for speaker or lexical variabilities in speech for emotion recognition
Soroosh Mariooryad, Carlos Busso
Speech Commun.2
2013 Analysis and Compensation of the Reaction Lag of Evaluators in Continuous Emotional Annotations
abstract
Defining useful emotional descriptors to characterize expressive behaviors is an important research area in affective computing. Recent studies have shown the benefits of using continuous emotional evaluations to annotate spontaneous corpora. Instead of assigning global labels per segments, this approach captures the temporal dynamic evolution of the emotions. A challenge of continuous assessments is the inherent reaction lag of the evaluators. During the annotation process, an observer needs to sense the stimulus, perceive the emotional message, and define his/her judgment, all this in real time. As a result, we expect a reaction lag between the annotation and the underlying emotional content. This paper uses mutual information to quantify and compensate for this reaction lag. Classification experiments on the SEMAINE database demonstrate that the performance of emotion recognition systems improve when the evaluator reaction lag is considered. We explore annotator-dependent and annotator-independent compensation schemes.
Soroosh Mariooryad, Carlos Busso
ACII2
2013 Audiovisual corpus to analyze whisper speech
abstract
Current automatic speech recognition (ASR) systems cannot recognize whisper speech with high accuracy. ASR systems are trained with neutral speech, which have significant acoustic differences with whisper speech (i.e., energy, duration, harmonics structure, and spectral slope). Given the limitations of speech-based systems to process whisper speech, we propose to explore the benefits of visual features describing the orofacial area. We hypothesize that the lips' articulation between whisper and neutral speech is similar, providing a valuable whisper-invariant modality. This paper introduces the first audiovisual corpus of whisper speech. While we are targeting over 40 speakers, the current corpus has recordings from eleven subjects who were asked to read TIMIT sentences, and isolated digits alternating between neutral and whisper speech. The corpus also includes spontaneous recordings, in which the subject answered a series of general questions. The paper also analyzes an exhaustive set of audiovisual features, including action units (AUs), lip spreading, fundamental frequency, intensity, MFCCs, and formants. We study the differences in the features' distributions between whisper and neutral speech using Kullback-Leibler divergence (KLD). Then, we conducted statistical test to determine whether the differences in the features are statistically significant. The results support our hypothesis that visual features are less affected by whisper speech.
Tam Tran, Soroosh Mariooryad, Carlos Busso
ICASSP3
2013 Analysis of facial features of drivers under cognitive and visual distractions
abstract
Drivers are exposed to a growing risk of being distracted with the recent development of in-vehicle systems for navigation, communication and infotainment. As a result, there is a need for tracking systems that can monitor the drivers' attention. This study investigates driver distractions using a multimodal corpus collected from real world driving scenarios. The paper focuses on facial cues automatically extracted from a frontal camera facing the driver. We conducted subjective evaluations by external observers to assess the perceived visual and cognitive distraction of drivers performing secondary tasks. The data is divided into two classes - distracted and normal. This partition is separately created for visual and cognitive scores. Binary classifiers are built with features describing action units (AU) and gaze (e.g., head poses). The classifiers achieve 80.8% F-score for visual distractions, and 73.8% F-score for cognitive distractions. The study identifies features that are relevant for detecting both types of distractions. Furthermore, the paper presents a logistic regression analysis to identify facial features that are useful for detecting samples in which cognitive distraction scores are not related to visual distraction scores. The analysis reveals the benefits of using AU in cognitive related distraction detection.
Nanxiang Li, Carlos Busso
ICME2
2013 Evaluating the robustness of an appearance-based gaze estimation method for multimodal interfaces
abstract
Given the crucial role of eye movements on visual attention, tracking gaze behaviors is an important research problem in various applications including biometric identification, attention modeling and human-computer interaction. Most of the existing gaze tracking methods require a repetitive system calibration process and are sensitive to the user's head movements. Therefore, they cannot be easily implemented in current multimodal interfaces. This paper investigates an appearance-based approach for gaze estimation that requires minimum calibration and is robust against head motion. The approach consists in building an orthonormal basis, or eigenspace, of the eye appearance with principal component analysis (PCA). Unlike previous studies, we build the eigenspace using image patches displaying both eyes. The projections into the basis are used to train regression models which predict the gaze location. The approach is trained and tested with a new multimodal corpus introduced in this paper. We consider several variables such as the distance between user and the computer monitor, and head movement. The evaluation includes the performance of the proposed gaze estimation system with and without head movement. It also evaluates the results in subject-dependent versus subject-independent conditions under different distances. We report promising results which suggest that the proposed gaze estimation approach is a feasible and flexible scheme to facilitate gaze-based multimodal interfaces.
Nanxiang Li, Carlos Busso
ICMI2
2013 Energy and F0 contour modeling with functional data analysis for emotional speech detection
abstract
This paper proposes the use of reference models to detect emotional prominence in the energy and F0 contours. The proposed framework aims to model the intrinsic variability of these prosodic features. We present a novel approach based on Functional Data Analysis (FDA) to build reference models using a family of energy and F0 contours, which are implemented with lexicon-independent models. The neutral models are represented by bases of functions and the testing energy and F0 contours are characterized by their projections onto the corresponding bases. The proposed system can lead to accuracies as high as 80.4% in binary emotion classification in the EMODB corpus, which is 17.6% higher than the one achieved by a benchmark classifier trained with sentence level prosodic features. The approach is also evaluated with the SEMAINE corpus, showing that it can be effectively used in real applications. Index Terms: Emotion detection, prosody modeling, emotional speech analysis, expressive speech, functional data analysis.
Juan Pablo Arias, Carlos Busso, Néstor Becerra Yoma
INTERSPEECH2
2013 Iterative Feature Normalization Scheme for Automatic Emotion Detection from Speech
abstract
The externalization of emotion is intrinsically speaker-dependent. A robust emotion recognition system should be able to compensate for these differences across speakers. A natural approach is to normalize the features before training the classifiers. However, the normalization scheme should not affect the acoustic differences between emotional classes. This study presents the iterative feature normalization (IFN) framework, which is an unsupervised front-end, especially designed for emotion detection. The IFN approach aims to reduce the acoustic differences, between the neutral speech across speakers, while preserving the inter-emotional variability in expressive speech. This goal is achieved by iteratively detecting neutral speech for each speaker, and using this subset to estimate the feature normalization parameters. Then, an affine transformation is applied to both neutral and emotional speech. This process is repeated till the results from the emotion detection system are consistent between consecutive iterations. The IFN approach is exhaustively evaluated using the IEMOCAP database and a data set obtained under free uncontrolled recording conditions with different evaluation configurations. The results show that the systems trained with the IFN approach achieve better performance than systems trained either without normalization or with global normalization.
Carlos Busso, Soroosh Mariooryad, Angeliki Metallinou, Shri Narayanan
IEEE Trans. Affect. Comput.1
2013 Exploring Cross-Modality Affective Reactions for Audiovisual Emotion Recognition
abstract
Psycholinguistic studies on human communication have shown that during human interaction individuals tend to adapt their behaviors mimicking the spoken style, gestures, and expressions of their conversational partners. This synchronization pattern is referred to as entrainment. This study investigates the presence of entrainment at the emotion level in cross-modality settings and its implications on multimodal emotion recognition systems. The analysis explores the relationship between acoustic features of the speaker and facial expressions of the interlocutor during dyadic interactions. The analysis shows that 72 percent of the time the speakers displayed similar emotions, indicating strong mutual influence in their expressive behaviors. We also investigate the cross-modality, cross-speaker dependence, using mutual information framework. The study reveals a strong relation between facial and acoustic features of one subject with the emotional state of the other subject. It also shows strong dependence between heterogeneous modalities across conversational partners. These findings suggest that the expressive behaviors from one dialog partner provide complementary information to recognize the emotional state of the other dialog partner. The analysis motivates classification experiments exploiting cross-modality, cross-speaker information. The study presents emotion recognition experiments using the IEMOCAP and SEMAINE databases. The results demonstrate the benefit of exploiting this emotional entrainment effect, showing statistically significant improvements.
Soroosh Mariooryad, Carlos Busso
IEEE Trans. Affect. Comput.2
2013 Modeling of Driver Behavior in Real World Scenarios Using Multiple Noninvasive Sensors
abstract
With the development of new in-vehicle technology, drivers are exposed to more sources of distraction, which can lead to an unintentional accident. Monitoring the driver attention level has become a relevant research problem. This is the precise aim of this study. A database containing 20 drivers was collected in real-driving scenarios. The drivers were asked to perform common secondary tasks such as operating the radio, phone and a navigation system. The collected database comprises of various noninvasive sensors including the controller area network-bus (CAN-Bus), video cameras and microphone arrays. The study analyzes the effects in driver behaviors induced by secondary tasks. The corpus is analyzed to identify multimodal features that can be used to discriminate between normal and task driving conditions. Separate binary classifiers are trained to distinguish between normal and each of the secondary tasks, achieving an average accuracy of 77.2%. When a joint, multi-class classifier is trained, the system achieved accuracies of 40.8%, which is significantly higher than chances (12.5%). We observed that the classifiers' accuracy varies across secondary tasks, suggesting that certain tasks are more distracting than others. Motivated by these results, the study builds statistical models in the form of Gaussian Mixture Models (GMMs) to quantify the actual deviations in driver behaviors from the expected normal driving patterns. The study includes task independent and task dependent models. Building upon these results, a regression model is proposed to obtain a metric that characterizes the attention level of the driver. This metric can be used to signal alarms, preventing collision and improving the overall driving experience.
Nanxiang Li, Jinesh J. Jain, Carlos Busso
IEEE Trans. Multim.3
2012 A personalized emotion recognition system using an unsupervised feature adaptation scheme
abstract
A personalized emotion recognition system aims to tune the model to recognize the expressive behaviors of a targeted person. Such a system can play an important role in various domains including call center and health care applications. Adapting any general emotion recognition system for a particular individual requires speech samples and prior knowledge about their emotional content. These assumptions constrain the use of these techniques in many real scenarios in which no annotated data is available to train or adapt the models. To address this problem, this paper introduces an unsupervised feature adaptation scheme that aims to reduce the mismatch between the acoustic features used to train the system and the acoustic features extracted from the unknown targeted speaker. The adaptation scheme uses our recently proposed iterative feature normalization (IFN) framework. An emotion detection system is trained with the IEMOCAP database. For testing, a database was created by downloading videos from a video-sharing website, containing various interviews from a targeted subject (1.5 hours). The detection system is used to identify emotional speech with and without the proposed feature adaptation scheme. The experimental results indicate that the proposed approach improves the unweighted accuracy from 50.8% to 70.0%.
Tauhidur Rahman, Carlos Busso
ICASSP2
2012 Factorizing speaker, lexical and emotional variabilities observed in facial expressions
abstract
An effective human computer interaction system should be equipped with mechanisms to recognize and respond to the affective state of the user. However, spoken message conveys different communicative aspects such as the verbal content, emotional state and idiosyncrasy of the speaker. Each of these aspects introduces variability that will affect the performance of an emotion recognition system. If the models used to capture the expressive behaviors are constrained by the lexical content and speaker identity, it is expected that the observed uncertainty in the channel will decrease, improving the accuracy of the system. Motivated by these observations, this study aims to quantify and localize the speaker, lexical and emotional variabilities observed in the face during human interaction. A metric inspired in mutual information theory is proposed to quantify the dependency of facial features on these factors. This metric uses the trace of the covariance matrix of facial motion trajectories to measure the uncertainty. The experimental results confirm the strong influence of the lexical information in the lower part of the face. For this facial region, the results demonstrate the benefit of constraining the emotional model on the lexical content. The ultimate goal of this research is to utilize this information to constrain the emotional models on the underlying lexical units to improve the accuracy of emotion recognition systems.
Soroosh Mariooryad, Carlos Busso
ICIP2
2012 Indoor robotic terrain classification via angular velocity based hierarchical classifier selection
abstract
This paper proposes a novel approach to terrain classification by wheeled mobile robots, which utilizes vibration data. In our proposed approach, a mobile robot has the ability to categorize terrain types simply by driving over them. Classification of terrain is based on measurements obtained from an inertial measurement unit strapped directly to the robot's chassis. In contrast to the previous approaches, we use acceleration and angular velocity measurements in all cardinal directions to extract over 800 features. Sequential Forward Floating Feature Selection is used to narrow down this large group of features to a set of 15 to 20 that are the most useful. The reduced set of features is used by a Linear Bayes Normal Classifier to classify terrain. Furthermore, different feature sets are generated for different velocity conditions, and the classifier switches based on the current robot velocity. Experimental results are presented that show the strong performance of the proposed system, including 90% accuracy over 20 continuous minutes of driving across different terrains.
David Tick, Tauhidur Rahman, Carlos Busso, Nicholas R. Gans
ICRA3
2012 Unveiling the Acoustic Properties that Describe the Valence Dimension
Carlos Busso, Tauhidur Rahman
INTERSPEECH1
2012 Generating Human-Like Behaviors Using Joint, Speech-Driven Models for Conversational Agents
abstract
During human communication, every spoken message is intrinsically modulated within different verbal and nonverbal cues that are externalized through various aspects of speech and facial gestures. These communication channels are strongly interrelated, which suggests that generating human-like behavior requires a careful study of their relationship. Neglecting the mutual influence of different communicative channels in the modeling of natural behavior for a conversational agent may result in unrealistic behaviors that can affect the intended visual perception of the animation. This relationship exists both between audiovisual information and within different visual aspects. This paper explores the idea of using joint models to preserve the coupling not only between speech and facial expression, but also within facial gestures. As a case study, the paper focuses on building a speech-driven facial animation framework to generate natural head and eyebrow motions. We propose three dynamic Bayesian networks (DBNs), which make different assumptions about the coupling between speech, eyebrow and head motion. Synthesized animations are produced based on the MPEG-4 facial animation standard, using the audiovisual IEMOCAP database. The experimental results based on perceptual evaluations reveal that the proposed joint models (speech/eyebrow/head) outperform audiovisual models that are separately trained (speech/head and speech/eyebrow).
Soroosh Mariooryad, Carlos Busso
IEEE Trans. Speech Audio Process.2
2011 Iterative feature normalization for emotional speech detection
abstract
Contending with signal variability due to source and channel effects is a critical problem in automatic emotion recognition. Any approach in mitigating these effects however has to be done so as to not compromise emotion-relevant information in the signal. A promising approach to this problem has been through feature normalization using features drawn from non-emotional ("neutral") speech samples. This paper considers a scheme for minimizing the inter-speaker differences while still preserving the emotional discrimination of the acoustic features. This can be achieved by estimating the normalization parameters using only neutral speech, and then applying the coefficients to the entire corpus (including emotional set). Specifically, this paper introduces a feature normalization scheme that implements these ideas by iteratively detecting neutral speech and normalizing the features. As the approximation error of the normalization parameters is reduced, the accuracy of the emotion detection system increases. The accuracy of the proposed iterative approach, evaluated across three databases, is only 2.5% lower than the one trained with optimal normalization parameters, and 9.7% higher than the one trained without any normalization scheme.
Carlos Busso, Angeliki Metallinou, Shri Narayanan
ICASSP1
2011 Analysis of driver behaviors during common tasks using frontal video camera and CAN-Bus information
abstract
Even a small distraction in drivers can lead to life-threatening accidents that affect the life of many. Monitoring distraction is a key aspect of any feedback system intended to keep the driver attention. Toward this goal, this paper studies the behaviors observed when the driver is performing in-vehicle common tasks such as operating a cellphone, radio or navigation system. The study employs the UTDrive platform -a car equipped with multiple sensors, including cameras, microphones, and Controller Area Network-Bus (CAN-Bus) information. The purpose of the analysis is to identify relevant features extracted from a frontal video camera and the car CAN-Bus data that can be used to distinguish between normal and task driving conditions. Statistical hypothesis tests are used to assess whether the differences observed in the selected features are significant. Then, these features are used in binary classification tasks (normal versus task). For most of the considered tasks, features extracted from the frontal video camera are found to be the most prominent indicators to distinguish between normal and task driving conditions (e.g., head pitch and yaw). The features from the car CAN-Bus data slightly improve the classification accuracy, from 76.7% (using features only from the frontal video) to 78.9% (using all features).
Jinesh J. Jain, Carlos Busso
ICME2
2011 Detecting Sleepiness by Fusing Classifiers Trained with Novel Acoustic Features
abstract
Automatic sleepiness detection is a challenging task that can lead to advances in various domains including traffic safety, medicine and human-machine interaction. This paper analyzes the discriminative power of different acoustic features to detect sleepiness. The study uses the sleepy language corpus (SLC). Along with standard acoustic features, novel features are proposed including functionals across voiced segment statistics in the F0 contour, likelihoods of reference models used to contrast non-neutral speech, and a set of robust to noise spectral features. These feature sets, which have performed well in other paralinguistic tasks such as emotion recognition, are used to train classifiers that are combined at the feature and decision levels. The best unweighted accuracy (UA) is obtained by combining the classifiers at the decision level under a maximum likelihood framework (UA = 70.97%). This performance is higher than the best results reported in the corpus. Index Terms: Speaker State Recognition, Paralinguistics, Affective Computing, Sleepiness
Tauhidur Rahman, Soroosh Mariooryad, Shalini Keshavamurthy, Gang Liu 0001, John H. L. Hansen, Carlos Busso
INTERSPEECH6
2011 Emotion recognition using a hierarchical binary decision tree approach
Chi-Chun Lee, Emily Mower Provost, Carlos Busso, Sungbok Lee, Shri Narayanan
Speech Commun.3
2010 Visual emotion recognition using compact facial representations and viseme information
abstract
Emotion expression is an essential part of human interaction. Rich emotional information is conveyed through the human face. In this study, we analyze detailed motion-captured facial information of ten speakers of both genders during emotional speech. We derive compact facial representations using methods motivated by Principal Component Analysis and speaker face normalization. Moreover, we model emotional facial movements by conditioning on knowledge of speech-related movements (articulation). We achieve average classification accuracies on the order of 75% for happiness, 50-60% for anger and sadness and 35% for neutrality in speaker independent experiments. We also find that dynamic modeling and the use of viseme information improves recognition accuracy for anger, happiness and sadness, as well as for the overall unweighted performance.
Angeliki Metallinou, Carlos Busso, Sungbok Lee, Shri Narayanan
ICASSP2
2009 Modeling mutual influence of interlocutor emotion states in dyadic spoken interactions
abstract
In dyadic human interactions, mutual influence- a person’s in-fluence on the interacting partner’s behaviors- is shown to be important and could be incorporated into the modeling frame-work in characterizing, and automatically recognizing the par-ticipants ’ states. We propose a Dynamic Bayesian Network (DBN) to explicitly model the conditional dependency between two interacting partners ’ emotion states in a dialog using data from the IEMOCAP corpus of expressive dyadic spoken in-teractions. Also, we focus on automatically computing the Valence-Activation emotion attributes to obtain a continuous characterization of the participants ’ emotion flow. Our pro-posed DBNmodels the temporal dynamics of the emotion states as well as the mutual influence between speakers in a dialog. With speech based features, the proposed network improves classification accuracy by 3.67 % absolute and 7.12 % relative over the Gaussian Mixture Model (GMM) baseline on isolated turn-by-turn emotion classification. Index Terms: emotion recognition, mutual influence, Dynamic Bayesian Network, dyadic interaction
Chi-Chun Lee, Carlos Busso, Sungbok Lee, Shri Narayanan
INTERSPEECH2
2009 Emotion recognition using a hierarchical binary decision tree approach
abstract
Automated emotion state tracking is a crucial element in the computational study of human communication behaviors. It is important to design robust and reliable emotion recognition systems that are suitable for real-world applications both to enhance analytical abilities to support human decision making and to design human-machine interfaces that facilitate efficient communication. We introduce a hierarchical computational structure to recognize emotions. The proposed structure maps an input speech utterance into one of the multiple emotion classes through subsequent layers of binary classifications. The key idea is that the levels in the tree are designed to solve the easiest classification tasks first, allowing us to mitigate error propagation. We evaluated the classification framework on two different emotional databases using acoustic features, the AIBO database and the USC IEMOCAP database. In the case of the AIBO database, we obtain a balanced recall on each of the individual emotion classes using this hierarchical structure. The performance measure of the average unweighted recall on the evaluation data set improves by 3.37% absolute (8.82% relative) over a Support Vector Machine baseline model. In the USC IEMOCAP database, we obtain an absolute improvement of 7.44% (14.58%) over a baseline Support Vector Machine modeling. The results demonstrate that the presented hierarchical approach is effective for classifying emotional utterances in multiple database contexts.
Chi-Chun Lee, Emily Mower Provost, Carlos Busso, Sungbok Lee, Shri Narayanan
INTERSPEECH3
2009 Analysis of Emotionally Salient Aspects of Fundamental Frequency for Emotion Detection
abstract
During expressive speech, the voice is enriched to convey not only the intended semantic message but also the emotional state of the speaker. The pitch contour is one of the important properties of speech that is affected by this emotional modulation. Although pitch features have been commonly used to recognize emotions, it is not clear what aspects of the pitch contour are the most emotionally salient. This paper presents an analysis of the statistics derived from the pitch contour. First, pitch features derived from emotional speech samples are compared with the ones derived from neutral speech, by using symmetric Kullback-Leibler distance. Then, the emotionally discriminative power of the pitch features is quantified by comparing nested logistic regression models. The results indicate that gross pitch contour statistics such as mean, maximum, minimum, and range are more emotionally prominent than features describing the pitch shape. Also, analyzing the pitch statistics at the utterance level is found to be more accurate and robust than analyzing the pitch statistics for shorter speech regions (e.g., voiced segments). Finally, the best features are selected to build a binary emotion detection system for distinguishing between emotional versus neutral speech. A new two-step approach is proposed. In the first step, reference models for the pitch features are trained with neutral speech, and the input features are contrasted with the neutral model. In the second step, a fitness measure is used to assess whether the input speech is similar to, in the case of neutral speech, or different from, in the case of emotional speech, the reference models. The proposed approach is tested with four acted emotional databases spanning different emotional categories, recording settings, speakers and languages. The results show that the recognition accuracy of the system is over 77% just with the pitch features (baseline 50%). When compared to conventional classification schemes, the proposed approach performs better in terms of both accuracy and robustness.
Carlos Busso, Sungbok Lee, Shri Narayanan
IEEE Trans. Speech Audio Process.1
2008 The expression and perception of emotions: comparing assessments of self versus others
abstract
In the study of expressive speech communication, it is commonly accepted that the emotion perceived by the listener is a good approximation of the intended emotion conveyed by the speaker. This paper analyzes the validity of this assumption by comparing the mismatches between the assessments made by naive listeners and by the speakers that generated the data. The analysis is based on the hypothesis that people are better decoders of their own emotions. Therefore, self-assessments will be closer to the intended emotions. Using the IEMOCAP database, discrete (categorical) and continuous (attribute) emotional assessments evaluated by the actors and naive listeners are compared. The results indicate that there is a mismatch between the expression and perception of emotion. The speakers in the database assigned their own emotions to more specific emotional categories, which led to more extreme values in the activation-valence space.
Carlos Busso, Shri Narayanan
INTERSPEECH1
2008 Scripted dialogs versus improvisation: lessons learned about emotional elicitation techniques from the IEMOCAP database
Carlos Busso, Shri Narayanan
INTERSPEECH1
2007 Real-Time Monitoring of Participants' Interaction in a Meeting using Audio-Visual Sensors
abstract
Intelligent environments equipped with audio-visual sensors provide suitable means for automatically monitoring and tracking the behavior, strategies and engagement of the participants in multiperson meetings. In this paper, high-level features are calculated from active speaker segmentations, automatically annotated by our smart room system, to infer the interaction dynamics between the participants. These features include the number and the average duration of each turn, statistics of turn-taking such as time as active speaker, and turn-taking transition patterns between participants. The results show that it is possible to accurately estimate in real-time not only the flow of the interaction, but also how dominant and engaged each participant was during the discussion. These high-level features, which cannot be inferred from any of the individual modalities by themselves, can be useful for summarization, classification, retrieval and (after action) analysis of meetings.
Carlos Busso, Panayiotis G. Georgiou, Shri Narayanan
ICASSP (2)1
2007 Using neutral speech models for emotional speech analysis
abstract
Abstract Since emotional speech can be regarded as a variation onneutral (non-emotional) speech, it is expected that a robust neu-tral speech model can be useful in contrasting different emo-tions expressed in speech. This study explores this idea by cre-ating acoustic models trained with spectral features, using theemotionally-neutral TIMIT corpus. The performance is testedwith two emotional speech databases: one recorded with a mi-crophone (acted), and another recorded from a telephone ap-plication (spontaneous). It is found that accuracy up to 78%and 65% can be achieved in the binary and category emotiondiscriminations, respectively. Raw Mel Filter Bank (MFB) out-put was found to perform better than conventional MFCC, withboth broad-band and telephone-band speech. These results sug-gest that well-trained neutral acoustic models can be effectivelyused as a front-end for emotion recognition, and once trainedwith MFB, it may reasonably work well regardless of the chan-nel characteristics.Index Terms: Emotion recognition, Neutral speech, HMMs,Mel filter bank (MFB), TIMIT
Carlos Busso, Sungbok Lee, Shri Narayanan
INTERSPEECH1
2007 Joint Analysis of the Emotional Fingerprint in the Face and Speech: A single subject study
abstract
In daily human interaction, speech and gestures are used to express an intended message, enriched with verbal and non-verbal information. Although many communicative goals are simultaneously encoded using the same modalities such as the face or the voice, listeners are generally good at decoding each aspect of the message. This encoding process includes an underlying interplay between communicative goals and channels, which is yet not well understood. In this direction, this paper explores the interplay between linguistic and affective goals in speech and facial expression. We hypothesize that when one modality is constrained by the articulatory speech process, other channels with more degrees of freedom are used to convey the emotions. The results presented here support this hypothesis, since it is observed that facial expression and prosodic speech tend to have a stronger emotional modulation when the vocal tract is physically constrained by the articulation to convey other linguistic communicative goals.
Carlos Busso, Shri Narayanan
MMSP1
2007 Multimodal Meeting Monitoring: Improvements on Speaker Tracking and Segmentation through a Modified Mixture Particle Filter
abstract
In this paper we address improvements to our multimodal system for tracking of meeting participants and speaker segmentation with a focus on the microphone array modality. We propose an algorithm that uses Directions-of-Arrival estimated for each microphone pair as observations and performs tracking of an unknown number of acoustically-active meeting participants and subsequent speaker segmentation. We propose modified mixture particle filter (mMPF) for tracking of acoustic sources in the track-before-detection (TbD) framework. Trajectories of sound sources are reconstructed by the optimal assignment of posterior mixture components produced by mMPF in consecutive frames. Further, we propose a sequential optimal change-point detection algorithm which discovers speech segments in the reconstructed trajectories i.e., performs speaker segmentation. The algorithm is tested on a multi-participant meeting dataset both separately and as a part of the multimodal system. On the task of speaker detection in the multimodal setup we report significant improvement over our previous state of the art implementation.
Viktor Rozgic, Carlos Busso, Panayiotis G. Georgiou, Shri Narayanan
MMSP2
2007 Rigid Head Motion in Expressive Speech Animation: Analysis and Synthesis
abstract
Rigid head motion is a gesture that conveys important nonverbal information in human communication, and hence it needs to be appropriately modeled and included in realistic facial animations to effectively mimic human behaviors. In this paper, head motion sequences in expressive facial animations are analyzed in terms of their naturalness and emotional salience in perception. Statistical measures are derived from an audiovisual database, comprising synchronized facial gestures and speech, which revealed characteristic patterns in emotional head motion sequences. Head motion patterns with neutral speech significantly differ from head motion patterns with emotional speech in motion activation, range, and velocity. The results show that head motion provides discriminating information about emotional categories. An approach to synthesize emotional head motion sequences driven by prosodic features is presented, expanding upon our previous framework on head motion synthesis. This method naturally models the specific temporal dynamics of emotional head motion sequences by building hidden Markov models for each emotional category (sadness, happiness, anger, and neutral state). Human raters were asked to assess the naturalness and the emotional content of the facial animations. On average, the synthesized head motion sequences were perceived even more natural than the original head motion sequences. The results also show that head motion modifies the emotional perception of the facial animation especially in the valence and activation domain. These results suggest that appropriate head motion not only significantly improves the naturalness of the animation but can also be used to enhance the emotional content of the animation to effectively engage the users
Carlos Busso, Zhigang Deng 0001, Michael Grimm, Ulrich Neumann, Shri Narayanan
IEEE Trans. Speech Audio Process.1
2007 Interrelation Between Speech and Facial Gestures in Emotional Utterances: A Single Subject Study
abstract
The verbal and nonverbal channels of human communication are internally and intricately connected. As a result, gestures and speech present high levels of correlation and coordination. This relationship is greatly affected by the linguistic and emotional content of the message. The present paper investigates the influence of articulation and emotions on the interrelation between facial gestures and speech. The analyses are based on an audio-visual database recorded from an actress with markers attached to her face, who was asked to read semantically neutral sentences, expressing four emotion states (neutral, sadness, happiness, and anger). A multilinear regression framework is used to estimate facial features from acoustic speech parameters. The levels of coupling between the communication channels are quantified by using Pearson's correlation between the recorded and estimated facial features. The results show that facial and acoustic features are strongly interrelated, showing levels of correlation higher than r = 0.8 when the mapping is computed at sentence-level using spectral envelope speech features. The results reveal that the lower face region provides the highest activeness and correlation levels. Furthermore, the correlation levels present significant interemo- tional differences, which suggest that emotional content affect the relationship between facial gestures and speech. Principal component analysis (PCA) shows that the audiovisual mapping parameters are grouped in a smaller subspace, which suggests that there is an emotion-dependent structure that is preserved from across sentences. The results suggest that this internal structure seems to be easy to model when prosodic-features are used to estimate the audiovisual mapping. The results also reveal that the correlation levels within a sentence vary according to broad phonetic properties presented in the sentence. Consonants, especially unvoiced and fricative sounds, present the lowest correlation levels. Likewise, the results show that facial gestures are linked at different resolutions. While the orofacial area is locally connected with the speech, other facial gestures such as eyebrow motion are linked only at the sentence-level. The results presented here have important implications for applications such as facial animation and multimodal emotion recognition.
Carlos Busso, Shri Narayanan
IEEE Trans. Speech Audio Process.1
2006 Modeling, estimating, and compensating low-bit rate coding distortion in speech recognition
abstract
A solution to the problem of speech recognition with signals distorted by low-bit rate coders is presented in this paper. A model for the coding-decoding distortion, a HMM compensation method to include this model, and an EM-based adaptation algorithm to estimate this distortion are proposed here. Medium vocabulary continuous-speech speaker-independent recognition experiments with 8 kbps G.729(CS-CELP), 13 kbps RPE-LTP (GSM), 5.3 kbps G723.1, 4.8 kbps FS-1016 and 32 kbps G.726(ADPCM) coders show that the approach described in this paper is able to dramatically reduce the effect of the coding distortion and, in some cases, gives a word accuracy higher than the baseline system with uncoded speech. Finally, the EM estimation algorithm requires only one adapting utterance and the approach described is certainly suitable for dialogue systems where just a few adapting utterances are available.
Néstor Becerra Yoma, Jorge F. Silva, Carlos Busso
IEEE Trans. Speech Audio Process.4
2005 Smart room: participant and speaker localization and identification
abstract
Our long-term objective is to create smart room technologies that are aware of the users presence and their behavior and can become an active, but not an intrusive, part of the interaction. In this work, we present a multimodal approach for estimating and tracking the location and identity of the participants including the active speaker. Our smart room design contains three user-monitoring systems: four CCD cameras, an omnidirectional camera and a 16 channel microphone array. The various sensory modalities are processed both individually and jointly and it is shown that the multimodal approach results in significantly improved performance in spatial localization, identification and speech activity detection of the participants.
Carlos Busso, Sergi Hernanz, Chi-Wei Chu, Soonil Kwon, Sung Lee, Panayiotis G. Georgiou, Isaac Cohen, Shri Narayanan
ICASSP (2)1
2005 Investigating the role of phoneme-level modifications in emotional speech resynthesis
abstract
Recent studies in our lab show that emotions in speech are manifested as, besides supra-segmental trends, distinct variations in phoneme-level prosodic and spectral parameters. In this paper, we further investigate the significance of this finding in the context of emotional speech synthesis. Specifically, we study phoneme-level signal property manipulation in transforming the emotional information conveyed in a speech utterance. We analyze the effect of individual and combined modifications of F0, duration, energy and spectrum using data recorded by a professional actress with happy, angry, sad and neutral expressiveness. We use content matched source-target pairs and apply TDPSOLA for prosody and LPC for spectrum modifications by directly extracting the required parameters from the target speech. Listening tests conducted with 10 naive raters show that modification of prosody and spectral envelope parameters by themselves is not sufficient. However, when applied together, modifying spectrum and prosody at the phone level gives successful results for most emotion pairs, except conversion to happy targets. We also observe that at the phoneme level, spectral envelope modifications are more effective than local prosodic modifications; and that, duration modifications are more effective than pitch modifications. The results confirm our hypothesis that phoneme level modifications can be used to fine tune the ensuing suprasegmental-parameter-based modifications to improve the overall quality of synthesized emotions.
Murtaza Bulut, Carlos Busso, Serdar Yildirim, Abe Kazemzadeh, Chul Min Lee, Sungbok Lee, Shri Narayanan
INTERSPEECH2
2005 Natural head motion synthesis driven by acoustic prosodic features
abstract
Abstract Natural head motion is important to realistic facial animation and engaging human–computer interactions. In this paper, we present a novel data‐driven approach to synthesize appropriate head motion by sampling from trained hidden markov models (HMMs). First, while an actress recited a corpus specifically designed to elicit various emotions, her 3D head motion was captured and further processed to construct a head motion database that included synchronized speech information. Then, an HMM for each discrete head motion representation (derived directly from data using vector quantization) was created by using acoustic prosodic features derived from speech. Finally, first‐order Markov models and interpolation techniques were used to smooth the synthesized sequence. Our comparison experiments and novel synthesis results show that synthesized head motions follow the temporal dynamic behavior of real human subjects. Copyright © 2005 John Wiley & Sons, Ltd.
Carlos Busso, Zhigang Deng 0001, Ulrich Neumann, Shri Narayanan
Comput. Animat. Virtual Worlds1
2004 Analysis of emotion recognition using facial expressions, speech and multimodal information
abstract
The interaction between human beings and computers will be more natural if computers are able to perceive and respond to human non-verbal communication such as emotions. Although several approaches have been proposed to recognize human emotions based on facial expressions or speech, relatively limited work has been done to fuse these two, and other, modalities to improve the accuracy and robustness of the emotion recognition system. This paper analyzes the strengths and the limitations of systems based only on facial expressions or acoustic information. It also discusses two approaches used to fuse these two modalities: decision level and feature level integration. Using a database recorded from an actress, four emotions were classified: sadness, anger, happiness, and neutral state. By the use of markers on her face, detailed facial motions were captured with motion capture, in conjunction with simultaneous speech recordings. The results reveal that the system based on facial expression gave better performance than the system based on just acoustic information for the emotions considered. Results also show the complementarily of the two modalities and that when these two modalities are fused, the performance and the robustness of the emotion recognition system improve measurably.
Carlos Busso, Zhigang Deng 0001, Serdar Yildirim, Murtaza Bulut, Chul Min Lee, Abe Kazemzadeh, Sungbok Lee, Ulrich Neumann, Shri Narayanan
ICMI1
2004 Emotion recognition based on phoneme classes
abstract
Recognizing human emotions/attitudes from speech cues has gained increased attention recently. Most previous work has focused primarily on suprasegmental prosodic features calcu-lated at the utterance level for modeling against details at the segmental phoneme level. Based on the hypothesis that dif-ferent emotions have varying effects on the properties of the different speech sounds, this paper investigates the usefulness of phoneme-level modeling for the classification of emotional states from speech. Hidden Markov models (HMM) based on short-term spectral features are used for this purpose using data obtained from a recording of an actress ’ expressing 4 different emotional states- anger, happiness, neutral, and sadness. We designed and compared two sets of HMM classifiers: a generic set of “emotional speech ” HMMs (one for each emotion) and a set of broad phonetic-class based HMMs for each emotion type considered. Five broad phonetic classes were used to explore the effect of emotional coloring on different phoneme classes, and it was found that spectral properties of vowel sounds were the best indicator of emotions in terms of the classification per-formance. The experiments also showed that the better per-formance can be obtained by using phoneme-class classifiers than generic “emotional ” HMM classifier and classifiers based on global prosodic features. To see the complementary effect of the prosodic and spectral features, the two classifiers were combined at the decision level. The improvement was 0.55% in absolute (0.7 % relatively) compared with the result from phoneme-class based HMM classifier. 1.
Chul Min Lee, Serdar Yildirim, Murtaza Bulut, Abe Kazemzadeh, Carlos Busso, Zhigang Deng 0001, Sungbok Lee, Shri Narayanan
INTERSPEECH5
2004 An acoustic study of emotions expressed in speech
abstract
In this study, we investigate acoustic properties of speech associ-ated with four different emotions (sadness, anger, happiness, and neutral) intentionally expressed in speech by an actress. The aim is to obtain detailed acoustic knowledge on how speech is modulated when speaker’s emotion changes from neutral to a certain emotional state. It is based on measurements of acoustic parameters related to speech prosody, vowel articulation and spectral energy distribution. Acoustic similarities and differences among the emotions are then explored with mutual information computation, multidimensional scaling, and comparison of acoustic likelihoods relative to the neu-tral emotion. In addition, acoustic separability of the emotions is tested using the discriminant analysis at the utterance level and the result is compared with human evaluation. Results show that hap-piness/anger and neutral/sadness share similar acoustic properties in this speaker. Speech associated with anger and happiness are characterized by longer utterance duration, shorter inter-word si-lence, higher pitch and energy values with wider ranges, showing the characteristics of exaggerated or hyperarticulated speech. The discriminant analysis indicates that within-group acoustic separa-bility is relatively poor, suggesting that conventional acoustic pa-rameters examined in this study are not effective in describing the emotions along the valence (or pleasure) dimension. It is noted that RMS energy, inter-word silence and speaking rate are useful in dis-tinguishing sadness from others. Interestingly, the between-group difference in formant patterns seems better reflected in back vowels such as /a / (/father/) than in the front vowels. Larger lip opening and/or more tongue constriction at the mid or rear part of the vocal tract could be underlying reasons. 1.
Serdar Yildirim, Murtaza Bulut, Chul Min Lee, Abe Kazemzadeh, Zhigang Deng 0001, Sungbok Lee, Shri Narayanan, Carlos Busso
INTERSPEECH8
2004 A real-time protocol for the Internet based on the least mean square algorithm
abstract
Generally, real-time applications based on the User Datagram Protocol (UDP) generate large volumes of data and are not sensitive to network congestion. In contrast, Transmission Control Protocol (TCP) traffic is considered "well-behaved" because it prevents the network becoming congested by means of closed-loop control of packet-loss and round-trip-time. The integration of both sorts of traffic is a complex problem, and depends on solutions such as admission control that have not yet been deployed on the Internet. Moreover, the problem of quality-of-service (QoS) and resource allocation is extremely relevant from the point of view of convergence of streaming media and data transmission on the Internet. In this paper an adaptive real-time protocol based on the least mean square (LMS) algorithm is proposed to estimate the application UDP bandwidth in order to reduce the quadratic error between the packet loss and a target. Moreover, the LMS algorithm is also applied to make sure that the reduction in the average bandwidth allocated to each TCP process will not be higher than a given percentage of the average bandwidth allocated before the beginning of the UDP application.
Néstor Becerra Yoma, Juan Hood, Carlos Busso
IEEE Trans. Multim.3