Chi-Chun Lee

dblp:99/8062 · also Chi-Chun Jeremy Lee · DBLP profile ↗
← Back
151ranked-venue papers
12as first author
70since 2021 · last 2026
0000-0003-0186-4321ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 116 · 9 first-author · 49 since 2021Artificial intelligence and machine learning · 81 · 8 first-author · 36 since 2021Human-computer interaction and ubiquitous computing · 13 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Only Subsets Matters: The Effect of Dual Fairness Constraints in Speech Emotion Recognition
abstract
Speech Emotion Recognition (SER) systems are increasingly deployed in voice-centric applications, yet often suffer from fairness concerns due to speaker-induced variability. In particular, speaker-gender bias can cause systematic disparities in performance across demographic groups (group fairness), while emotionally expressive or acoustically unique speakers may be treated inconsistently despite similar input (individual fairness). Although prior work suggests that group and individual fairness objectives may be inherently incompatible, our proposed two-stage debiasing framework aims to address both: an in-processing approach first mitigates speaker-gender bias (group fairness), followed by a post-processing calibration step that improves consistency across similar instances (individual fairness). While most samples benefit from this dual intervention, we identify a smallsubsetof speech samples that remain difficult to classify fairly. This work focuses on systematically analyzing these fairness-ambiguous samples to understand what makes them challenging. We examine this question from two perspectives: emotion perception and acoustic expressivity. Our analyses on thesesubsetsindicate that: (1) exhibit extreme or atypical emotional ratings, (2) show high acoustic variability, and (3) tend to come disproportionately from specific individuals. These findings suggest that some speakers inherently present greater challenges to fairness optimization, due to the uniqueness of their emotional or acoustic expression. By characterizing thesesubsets, our work contributes to a deeper understanding of fairness conflicts in SER and offers new directions for developing more robust and inclusive emotion recognition systems.
Woan-Shiuan Chien, Chi-Chun Lee
IEEE Trans. Affect. Comput.2
2025 RE-LLM: Refining Empathetic Speech-LLM Responses by Integrating Emotion Nuance
abstract
With generative AI advancing, empathy in human-AI interaction is essential. While prior work focuses on emotional reflection, emotional exploration—key to deeper engagement—remains overlooked. Existing LLMs rely on text which captures limited emotion nuances. To address this, we propose RE-LLM, a speech-LLM integrating dimensional emotion embeddings and auxiliary learning. Experiments show statistically significant gains in empathy metrics almost across three datasets. RE-LLM relatively improves the Emotional Reaction score by 14.79% and $\mathbf{6. 7 6} \boldsymbol{\%}$ compared to text-only and speech-LLM baselines on ESD. Notably, it raises the Exploration score by 35.42% and 3.91% on IEMOCAP, 139.28% and 9.83% on ESD and 60.95% and 22.64% on MSP-PODCAST relatively. It also boosts unweighted accuracy by 5.4% on IEMOCAP, 2.3% on ESD and $\mathbf{6. 9 \%}$ on MSP-PODCAST in speech emotion recognition. These results highlight the enriched emotional understanding and improved empathetic response generation of RE-LLM.
Jing-Han Chen, Bo-Hao Su, Ya-Tse Wu, Chi-Chun Lee
ASRU4
2025 Speaker Style-Aware Phoneme Anchoring For Improved Cross-Lingual Speech Emotion Recognition
abstract
Cross-lingual speech emotion recognition (SER) remains a challenging task due to differences in phonetic variability and speaker-specific expressive styles across languages. Effectively capturing emotion under such diverse conditions requires a framework that can align the externalization of emotions across different speakers and languages. To address this problem, we propose a speaker-style aware phoneme anchoring framework that aligns emotional expression at the phonetic and speaker levels. Our method builds emotionspecific speaker communities via graph-based clustering to capture shared speaker traits. Using these groups, we apply dual-space anchoring in speaker and phonetic spaces to enable better emotion transfer across languages. Evaluations on the MSP-Podcast (English) and BIIC-Podcast (Taiwanese Mandarin) corpora demonstrate improved generalization over competitive baselines and provide valuable insights into the commonalities in cross-lingual emotion representation.
Shreya G. Upadhyay, Carlos Busso, Chi-Chun Lee
ASRU3
2025 ASR for Affective Speech: Investigating Impact of Emotion and Speech Generative Strategy
abstract
This work investigates how emotional speech and generative strategies affect ASR performance. We analyze speech synthesized from three emotional TTS models and find that substitution errors dominate, with emotional expressiveness varying across models. Based on these insights, we introduce two generative strategies: one using transcription correctness and another using emotional salience, to construct fine-tuning subsets. Results show consistent WER improvements on real emotional datasets without noticeable degradation on clean LibriSpeech utterances. The combined strategy achieves the strongest gains, particularly for expressive speech. These findings highlight the importance of targeted augmentation for building emotion-aware ASR systems.
Ya-Tse Wu, Chi-Chun Lee
ASRU2
2025 Personalized Federated Learning with Fuzzy Clustering for Dysarthric Speech Recognition
abstract
Pathological speech recognition is challenging because clinical datasets are scarce, variable, and subject to strict privacy constraints preventing cross-institutional data sharing. These regulations necessitate federated learning (FL) for collaborative training without sharing raw data. However, FL degrades under non-IID data. Hard-clustering FL addresses this by partitioning clients into groups but imposes rigid boundaries, discards boundary samples, and suffers performance drops as cluster numbers increase. We propose Fuzzy Cluster-Based Personalized Federated Learning (FCPFL), using fuzzy C-means to softly group clients and pseudo-label-guided feature selection to identify discriminative features. FCPFL weights client updates by membership degree, allowing boundary samples to participate in multiple clusters and increasing training data by $25 \%$. Experiments show FCPFL reduces word error rate (WER) by $\mathbf{4 . 8 2 \%}$ and $\mathbf{1 . 7 4 \%}$ on ADReSS and TORGO, compared to hardclustered FL baselines.
Jie-Shiang Yang, Jing-Tong Tzeng, Chi-Chun Lee
ASRU3
2025 Valve Token Masked Autoencoder for Missing Recordings on Cardiac Abnormality Classification
abstract
Automated auscultation and cardiovascular screening systems for cardiac abnormalities have received growing interest in clinical applications. Still, they face challenges due to missing or invalid recordings caused by technical issues. To address this, we introduce a novel framework leveraging the masked autoencoder strategy, uniquely treating each heart valve recording as a distinct token. Our approach reconstructs missing valve data using existing representation and learnable mask tokens, achieving inter-valve integration through positional embeddings and TCNs. We demonstrate state-of-the-art (SOTA) performance on the CirCor DigiScope dataset, outperforming top participants and SOTA imputation methods in terms of mean cost of patient outcome, accuracy, F1-measure, and macro F1 score. Furthermore, our analysis highlights improved predictive accuracy on limited input data, while generative results indicate our capability to provide comprehensive reconstruction of auscultation recordings for further clinical evaluations.
An-Yan Chang, Jing-Tong Tzeng, Huan-Yu Chen, Chun-Hsiang Huang, Edward Pei-Chuan Huang, Chi-Chun Lee
ICASSP6
2025 SocialRecNet: A Multimodal LLM-Based Framework for Assessing Social Reciprocity in Autism Spectrum Disorder
abstract
Accurate assessment of social reciprocity is crucial for early diagnosis and intervention in Autism Spectrum Disorder (ASD). Traditional methods, often relying on unimodal data or lacking in cross-modal alignment, do not fully capture the complexity of social reciprocity. To address these limitations, we developed SocialRecNet, a novel Multimodal Large Language Model (MLLM) utilizing the Autism Diagnostic Observation Schedule (ADOS) dataset. SocialRecNet integrates conversational speech and text with the textual reasoning capabilities of LLMs to analyze social reciprocity across multiple dimensions. By effectively aligning speech and text, enhanced by properly designed prompts, SocialRecNet achieves an average Pearson correlation of 0.711 in predicting ADOS scores, marking a significant improvement of approximately 26.24% over the best-performing baseline method. This state of the art framework not only improves the prediction of social reciprocity scores but also provides deeper insights into ASD diagnosis and intervention strategies.
Chin-Po Chen, Bo-Hao Su, Susan Shur-Fen Gau, Chi-Chun Lee
ICASSP6
2025 Disentangle Heart Rate Signals for Improved Stress Detection
abstract
Accurate stress detection from physiological signals is often complicated by individual identity traits, which must first be identified before they can be effectively removed to improve model performance. To address this, we propose a method that combines Detrended Fluctuation Analysis (DFA) and Augmented Dickey-Fuller (ADF) Analysis to extract stable identity-related features from heart rate signals without relying on explicit labels. By masking these identity features and raw content features, we can effectively eliminate their impact on task-relevant stress signals using guided contrastive learning. Validated on the TILES-2018 and Firefighters datasets, our approach significantly improves stress detection accuracy, achieving F1 score gains of 5.3% and 8.1%, respectively, compared to baseline models. These results highlight the model’s enhanced ability to generalize to diverse populations while minimizing identity bias, ultimately improving the robustness and precision of stress detection systems.
Pin-Jhao Chen, Woan-Shiuan Chien, Chi-Chun Lee
ICASSP3
2025 Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
abstract
Speech Emotion Recognition (SER) systems rely on speech input and emotional labels annotated by humans. However, various emotion databases collect perceptional evaluations in different ways. For instance, the IEMOCAP dataset uses video clips with sounds for annotators to provide their emotional perceptions. However, the most significant English emotion dataset, the MSP-PODCAST, only provides speech for raters to choose the emotional ratings. Nevertheless, using speech as input is the standard approach to training SER systems. Therefore, the open question is the emotional labels elicited by which scenarios are the most effective for training SER systems. We comprehensively compare the effectiveness of SER systems trained with labels elicited by different modality stimuli and evaluate the SER systems on various testing conditions. Also, we introduce an all-inclusive label that combines all labels elicited by various modalities. We show that using labels elicited by voice-only stimuli for training yields better performance on the test set, whereas labels elicited by voice-only stimuli.
Huang-Cheng Chou, Chi-Chun Lee
ICASSP3
2025 Toward Zero-Shot Speech Emotion Recognition Using LLMs in the Absence of Target Data
abstract
In generalized Speech Emotion Recognition (SER), traditional generalization techniques like transfer learning and domain adaptation rely on access to some amount of unlabeled target domain data. However, with increasing privacy concerns, building SER systems under zero-shot scenarios, where no target domain data is available, poses a significant challenge. In such cases, conventional methods become impractical without access to target samples or features. To leverage any available target information to bridge this gap, this work explores the potential of Large Language Models (LLMs), with their powerful generative capabilities, to generate target corpora based on documented scenario settings and published research, enabling SER under zero-shot conditions. We assess the effectiveness of LLMs in SER tasks across both text and speech modalities under challenging zero-shot conditions, using IEMOCAP and MSP-PODCAST as unseen target corpora. To ensure a fair comparison, we validate the performance of the synthetic data against real source data from MELD and MSP-IMPROV. Our experimental results reveal that, on average, the synthetic data not only matches but often surpasses the performance of real data in both text and speech modalities.
Bo-Hao Su, Shreya G. Upadhyay, Chi-Chun Lee
ICASSP3
2025 Noise-Robust Speech Emotion Recognition Using Shared Self-Supervised Representations with Integrated Speech Enhancement
abstract
Recent studies have demonstrated the effectiveness of fine-tuning self-supervised speech representation models for speech emotion recognition (SER). However, applying SER in real-world environments remains challenging due to pervasive noise. Relying on low-accuracy predictions due to noisy speech can undermine the user’s trust. This paper proposes a unified self-supervised speech representation framework for enhanced speech emotion recognition designed to increase noise robustness in SER while generating enhanced speech. Our framework integrates speech enhancement (SE) and SER tasks, leveraging shared self-supervised learning (SSL)-derived features to improve emotion classification performance in noisy environments. This strategy encourages the SE module to enhance discriminative information for SER tasks. Additionally, we introduce a cascade unfrozen training strategy, where the SSL model is gradually unfrozen and fine-tuned alongside the SE and SER heads, ensuring training stability and preserving the generalizability of SSL representations. This approach demonstrates improvements in SER performance under unseen noisy conditions without compromising SE quality. When tested at a 0 dB signal-to-noise ratio (SNR) level, our proposed method outperforms the original baseline by 3.7% in F1-Macro and 2.7% in F1-Micro scores, where the differences are statistically significant.
Jing-Tong Tzeng, Seong-Gyun Leem, Ali N. Salman, Chi-Chun Lee, Carlos Busso
ICASSP4
2025 Is It Still Fair? Investigating Gender Fairness in Cross-Corpus Speech Emotion Recognition
abstract
Speech emotion recognition (SER) is a vital component in various everyday applications. Cross-corpus SER models are increasingly recognized for their ability to generalize performance. However, concerns arise regarding fairness across demographics in diverse corpora. Existing fairness research often focuses solely on corpus-specific fairness, neglecting its generalizability in cross-corpus scenarios. Our study focuses on this underexplored area, examining the gender fairness generalizability in cross-corpus SER scenarios. We emphasize that the performance of cross-corpus SER models and their fairness are two distinct considerations. Moreover, we propose the approach of a combined fairness adaptation mechanism to enhance gender fairness in the SER transfer learning tasks by addressing both source and target genders. Our findings bring one of the first insights into the generalizability of gender fairness in cross-corpus SER systems.
Shreya G. Upadhyay, Woan-Shiuan Chien, Chi-Chun Lee
ICASSP3
2025 Mouth Articulation-Based Anchoring for Improved Cross-Corpus Speech Emotion Recognition
abstract
Cross-corpus speech emotion recognition (SER) plays a vital role in numerous practical applications. Traditional approaches to cross-corpus emotion transfer often concentrate on adapting acoustic features to align with different corpora, domains, or labels. However, acoustic features are inherently variable and error-prone due to factors like speaker differences, domain shifts, and recording conditions. To address these challenges, this study adopts a novel contrastive approach by focusing on emotion-specific articulatory gestures as the core elements for analysis. By shifting the emphasis on the more stable and consistent articulatory gestures, we aim to enhance emotion transfer learning in SER tasks. Our research leverages the CREMA-D and MSP-IMPROV corpora as benchmarks and it reveals valuable insights into the commonality and reliability of these articulatory gestures. The findings highlight mouth articulatory gesture potential as a better constraint for improving emotion recognition across different settings or domains.
Shreya G. Upadhyay, Ali N. Salman, Carlos Busso, Chi-Chun Lee
ICASSP4
2025 A Dynamic Edge-Selection Mechanism in HRV Hypergraph Learning for Improved Stress Detection
abstract
Studies show that individual attributes such as age and gender significantly influence physiological responses and their correlation with stress, often forming complex and overlapping relationships. These attributes are essential for enhancing physiological signal-based stress detection. Our work leverages hypergraph in multi-attribute representation and tackles the challenge of redundant attributes that misguide embeddings. Our dynamic edge-selection mechanism for hypergraph-based metric learning (DESHM) enables the hypergraph to focus on selected stress-related attribute connections within groups. This batch-wise selection adapts to varying connections between batches and maximizes the effectiveness of metric learning. Evaluation of the TILES-2018 and Firefighter datasets shows promising results, further improving 3.99% in F1, 3.52% in BACC, and 25.27% in MCC compared to the best pairwise result on the TILES-2018 dataset. Our analysis shows that our mechanism prioritizes attributes that effectively represent stress levels, guiding hyper-graph clustering to achieve improved discriminability.
Jing-Chun Wang, Woan-Shiuan Chien, Chi-Chun Lee
ICASSP3
2025 Mask Augmentation For Tumor Classification In Medical Images
abstract
Tumor detection and classification in medical images are critical for guiding patient management and treatment decisions. However, accurate segmentation and classification of tumors remain challenging due to their small size relative to the overall image. Existing approaches often face difficulties with limited data and potential segmentation errors, resulting in suboptimal performance in real-world applications. To address these challenges, we propose a novel approach leveraging the Segmentation Mask Augmentation (SMA) framework. Our framework enhances the robustness of tumor classification models by generating diverse and imprecise segmentation masks during training, thereby simulating real-world scenarios. Experimental results across four distinct datasets demonstrate the effectiveness of our approach. Our framework presents a promising solution for robust tumor classification, with potential implications for improving clinical diagnosis and patient management.
Chun-Chieh Weng, Huan-Yu Chen, Jing-Tong Tzeng, Ching-Heng Lin, Po-Chih Kuo, Chi-Chun Lee
ICASSP6
2025 ZSDEVC: Zero-Shot Diffusion-based Emotional Voice Conversion with Disentangled Mechanism
Hsing-Hang Chou, Yun-Shao Lin, Ching-Chin Sung, Yu Tsao 0001, Chi-Chun Lee
INTERSPEECH5
2025 Defend for Self-Vocoding: A Novel Enhanced Decoder Network for Watermark Recovery
Ching-Yu Yang, Hsing-Hang Chou, Ya-Tse Wu, Bo-Hao Su, Chi-Chun Lee
INTERSPEECH6
2025 Lessons Learnt: Revisit Key Training Strategies for Effective Speech Emotion Recognition in the Wild
abstract
In this study, we revisit key training strategies in machine learning often overlooked in favor of deeper architectures. Specifically, we explore balancing strategies, activation functions, and fine-tuning techniques to enhance speech emotion recognition (SER) in naturalistic conditions. Our findings show that simple modifications improve generalization with minimal architectural changes. Our multi-modal fusion model, integrating these optimizations, achieves a valence CCC of 0.6953, the best valence score in Task 2: Emotional Attribute Regression. Notably, fine-tuning RoBERTa and WavLM separately in a single-modality setting, followed by feature fusion without training the backbone extractor, yields the highest valence performance. Additionally, focal loss and activation functions significantly enhance performance without increasing complexity. These results suggest that refining core components, rather than deepening models, leads to more robust SER in-the-wild.
Jing-Tong Tzeng, Bo-Hao Su, Ya-Tse Wu, Hsing-Hang Chou, Chi-Chun Lee
INTERSPEECH5
2025 Differential Impacts of Monologue and Conversation on Speech Emotion Recognition
abstract
The advancement ofSpeech Emotion Recognition(SER) is significantly dependent on the quality of emotional speech corpora used for model training. Researchers in the field of SER have developed various corpora by adjusting design parameters to enhance the reliability of the training source. For this study, we focus on exploring communication modes of collection, specifically analyzing spontaneous emotional speech patterns gathered during conversation or monologue. While conversations are acknowledged as effective for eliciting authentic emotional expressions, systematic analyses are necessary to confirm their reliability as a better source of emotional speech data. We investigate this research question from perceptual differences and acoustic variability present in both emotional speeches. Our analyses on multi-lingual corpora show that, first, raters exhibit higher consistency for conversation recordings when evaluating categorical emotions, and second, perceptions and acoustic patterns observed in conversational samples align more closely with expected trends discussed in relevant emotion literature. We further examine the impact of these differences on SER modeling, which shows that we can train a more robust and stable SER model by using conversation data. This work provides comprehensive evidence suggesting that conversation may offer a better source compared to monologue for developing an SER model.
Woan-Shiuan Chien, Shreya G. Upadhyay, Carlos Busso, Chi-Chun Lee
IEEE Trans. Affect. Comput.5
2025 Minority Views Matter: Evaluating Speech Emotion Classifiers With Human Subjective Annotations by an All-Inclusive Aggregation Rule
abstract
When selecting test data for subjective tasks, most studies define ground truth labels using aggregation methods such as the majority or plurality rules. These methods discard data points without consensus, making the test set easier than practical tasks where a prediction is needed for each sample. However, the discarded data points often express ambiguous cues that elicit coexisting traits perceived by annotators. This paper addresses the importance of considering all the annotations and samples in the data, highlighting that only showing the model's performance on an incomplete test set selected by using the majority or plurality rules can lead to bias in the models’ performances. We focus onspeech-emotion recognition(SER) tasks. We observe that traditional aggregation rules have a data loss ratio ranging from 5.63% to 89.17%. From this observation, we propose a flexible method named the all-inclusive aggregation rule to evaluate SER systems on the complete test data. We contrast traditional single-label formulations with a multi-label formulation to consider the coexistence of emotions. We show that training an SER model with the data selected by the all-inclusive aggregation rule shows consistently higher macro-F1 scores when tested in the entire test set, including ambiguous samples without agreement.
Huang-Cheng Chou, Lucas Goncalves, Seong-Gyun Leem, Ali N. Salman, Chi-Chun Lee, Carlos Busso
IEEE Trans. Affect. Comput.5
2025 Phonetically-Anchored Domain Adaptation for Cross-Lingual Speech Emotion Recognition
abstract
The prevalence of cross-lingualspeech emotion recognition(SER) modeling has significantly increased due to its wide range of applications. Previous studies have primarily focused on technical strategies to adapt features, domains, and labels across languages, often overlooking the underlying commonalities between the languages. In this study, we address the language adaptation challenge in cross-lingual scenarios by incorporating vowel-phonetic constraints. Our approach is structured in two main parts. First, we investigate the vowel-phonetic commonalities associated with specific emotions across languages, particularly focusing on common vowels that prove to be valuable for SER modeling. Second, we utilize these identified common vowels as anchors to facilitate cross-lingual SER. To demonstrate the effectiveness of our approach, we conduct case studies usingAmerican EnglishandTaiwanese Mandarinwith two naturalistic emotional speech corpora: the MSP-Podcast and BIIC-Podcast corpora. The approach leverages evidence that certain vowels, including monophthongs and diphthongs, exhibit emotion-specific commonality across languages, serving as phonetic anchors to enhance unsupervised cross-lingual SER learning. The proposed model surpasses baseline performance, highlighting the importance of phonetic similarities for effective language adaptation in cross-lingual SER scenarios.
Shreya G. Upadhyay, Luz Martinez-Lucas, William F. Katz, Carlos Busso, Chi-Chun Lee
IEEE Trans. Affect. Comput.5
2024 Stress Detection Using HRV Features Augmentation Based on Heart Rate Signal Transformation
abstract
Heart Rate Variability (HRV) features are recognized as powerful indicators of various diseases, including heart failure, diabetes, and mental health disorders. Besides, HRV features are robust against noise, making them ideal for wearable devices. Despite their potential, HRV feature sets are limited by sample quantity. Direct augmentation often distorts underlying signal properties and interpretability. This study addresses this limitation by introducing spatial-domain, temporal-domain, and serial signal transformations for HRV feature-based augmentations. We further analyze the transformed signal using the Poincare plots to understand the effect of the transformation effect from the clinical perspective. Lastly, we design a stress detection deep learning model using the TILES-2018 database to verify the effectiveness of the augmented HRV features, which shows performance enhancements ranging from 1.05 % to 4.52%. Among the three domains of transformations, serial transformations yield the best results.
Pin-Jhao Chen, Woan-Shiuan Chien, Huan-Yu Chen, Chi-Chun Lee
BSN4
2024 Understanding Missing Data Bias in Longitudinal Mental Stress Detection
abstract
Mental stress has become a growing concern in contemporary society; fortunately, recent developments in wear-able technology now offer a promising solution. However, a common issue in longitudinal tracking with wearable sensors is missing data, which can introduce biases during model training, affecting predictions and leading to unfair outcomes for users. In this work, we explore the impact of missing data on stress detection performance across two longitudinal datasets. Our analysis reveals that biases stemming from missing data can result in the unfair treatment of individuals with higher levels of missing data, detrimentally affecting model performance. Additionally, we assess various imputation methods to mitigate these issues. Our findings indicate that while imputation generally improves model overall performances, performance decreases significantly when missing data exceeds half of the total data. This research provides initial insights into the challenges of missing data in longitudinal studies.
Woan-Shiuan Chien, Chi-Chun Lee
BSN2
2024 In-The-Wild HRV-Based Stress Detection Using Individual-Aware Metric Learning
abstract
Advanced wearable tracking shows potential for identifying psychological and emotional stress relevant to the mental health of high-intensity emergency responders. Heart rate variability (HRV) captured by wearable devices can indi-cate the correlation between intra-subject daily variations and stress. HRV also varies due to various demographic attributes, representing inter-subject relationships. This work introduces an individual-aware metric learning approach that leverages HRV features to train intra-subject representations, considering inter-subject effects based on attribute similarity through stress label clustering. We use the multi-similarity loss within the metric learning framework to consider various personal attributes, thereby improving discriminability. Evaluation of the TILES-2018 and Firefighter database shows promising results in binary stress classification: F1 score of 68.15 % with BACC of 59.13 % and MCC of 0.186, and F1 score of 73.07% with BACC of 56.52% and MCC of 0.136. resnectively.
Jing-Chun Wang, Woan-Shiuan Chien, Huan-Yu Chen, Chi-Chun Lee
BSN4
2024 GaP-Aug: Gamma Patch-Wise Correction Augmentation Method for Respiratory Sound Classification
abstract
Automated auscultation analysis using electronic stethoscope has received growing interest in clinical applications. Recently, researchers showed successes by using deep learning methods to distinguish between pathological respiratory sound classes. Nevertheless, the challenge persists due to the scarcity of abnormal samples, and the distinct characteristics between low-pitched and discontinuous crackles and high-pitched and continuous wheezes. In this study, we proposed a novel augmentation method, namely gamma patch-wise correction augmentation, which directly operates on spectrograms to handle with these two challenges. We achieved state-of-the-art performances on both 60-40 official split and 80-20 cross-validation of the public ICBHI dataset, outperforming previous top-performing studies by 11.82% in sensitivity and 5.27% in ICBHI score. Furthermore, Grad-CAM analysis shows that our approach better preserves the distinctive characteristics of crackles and wheezes than SpecAug.
An-Yan Chang, Jing-Tong Tzeng, Huan-Yu Chen, Chih-Wei Sung, Chun-Hsiang Huang, Edward Pei-Chuan Huang, Chi-Chun Lee
ICASSP7
2024 Balancing Speaker-Rater Fairness for Gender-Neutral Speech Emotion Recognition
abstract
Speech emotion recognition (SER) adds to the humane aspects of voice technologies to enhance user experiences. The ground truth emotion annotations provided by human raters and attributes related to the speakers themselves arise a compounded fairness issue in SER. While there exist works in fair SER, our work presents one of the first studies in addressing the unique joint speaker-rater (two-sided) bias, focusing on the issue of gender fairness. Our cross-reference evaluation demonstrates that the SER fair model, which merely mitigates one-sided bias introduces biases when examining from another viewpoint. Furthermore, in order to handle model stability when optimizing for these compounded speaker-rater constraints, we introduce a flexible controlled mechanism that dynamically balances the contribution of each viewpoint. Our analyses show the efficacy of our approach in achieving a fair SER that meets the dual speaker-rater gender neutrality criterion.
Woan-Shiuan Chien, Shreya G. Upadhyay, Chi-Chun Lee
ICASSP3
2024 Concealing Medical Condition by Node Toggling in ASR for Dementia Patients
abstract
It is important to make automatic speech recognition (ASR) be inclusive to all users, including those with disorders. Besides model performances, privacy concerns, such as leakage of medical condition, are severe and harmful for this already vulnerable population. Hence, developing privacy-preserving machine learning (PPML) algorithms is important. Recent node cancellation strategies, while repeatedly showing their privacy protection efficacy, involve complex multi-branched structures with manually-tuned thresholds. In this work, we focus on learning ASR for dementia patients without revealing their medical condition. Specifically, we present a dementia attribute cancellation strategy (DACS) that trains a single toggling network in an end-to-end manner to toggle off particular node dimensions at ASR decoding, concealing a subject’s dementia status. We show that using DACS can achieve 33% dementia protection efficacy (DPE), and further configuring for higher protection efficacy achieves 44% DPE, with only a slight decrease of 0.1% WER in ASR performance.
Wei-Tung Hsu, Chin-Po Chen, Chi-Chun Lee
ICASSP3
2024 In-The-Wild Physiological-Based Stress Detection Using Federated Strategy
abstract
Continuously identifying day-to-day mental stress can be realized by accessing wearable devices to measure physiological indicators. However, the nature of bodily signals raises issues of privacy and data heterogeneity. Recent federated learning scheme provides a promising direction to alleviate the privacy concern, but the large inter-client differences can lead to a sub-optimal model performance. In this work, we propose a client-aware aggregation strategy to customize the global model forked by each client to conduct mutual learning in federated setting. Our proposed mixture Federated Mutual Learning (mixFML) weighs the distances of local models to generate a unique mixture of global model per client. We evaluated our method on the public TILES-2018 and an in-house Firefighters dataset for stress detection using HRV. Our proposed mixFML achieved 8.0% and 1.8% MCC improvement on two datasets compared to federated mutual learning.
Po-Chen Lin, Jeng-Lin Li, Woan-Shiuan Chien, Chi-Chun Lee
ICASSP4
2024 An Investigation of Group versus Individual Fairness in Perceptually Fair Speech Emotion Recognition
Woan-Shiuan Chien, Chi-Chun Lee
INTERSPEECH2
2024 An Inter-Speaker Fairness-Aware Speech Emotion Regression Framework
Hsing-Hang Chou, Woan-Shiuan Chien, Ya-Tse Wu, Chi-Chun Lee
INTERSPEECH4
2024 A Cluster-based Personalized Federated Learning Strategy for End-to-End ASR of Dementia Patients
Wei-Tung Hsu, Chin-Po Chen, Yun-Shao Lin, Chi-Chun Lee
INTERSPEECH4
2024 SWiBE: A Parameterized Stochastic Diffusion Process for Noise-Robust Bandwidth Expansion
Yin-Tse Lin, Shreya G. Upadhyay, Bo-Hao Su, Chi-Chun Lee
INTERSPEECH4
2024 Emo-bias: A Large Scale Evaluation of Social Bias on Speech Emotion Recognition
Yi-Cheng Lin, Huang-Cheng Chou, Chi-Chun Lee, Hung-yi Lee
INTERSPEECH4
2024 A Layer-Anchoring Strategy for Enhancing Cross-Lingual Speech Emotion Recognition
Shreya G. Upadhyay, Carlos Busso, Chi-Chun Lee
INTERSPEECH3
2024 Can Modelling Inter-Rater Ambiguity Lead To Noise-Robust Continuous Emotion Predictions?
abstract
There has been increasing attention drawn to modelling interrater ambiguity in Continuous Emotion Recognition (CER) systems using probability distributions for arousal and valence.However, the relationship between modelling label ambiguity and robustness to noise, and more broadly, the impact of realworld noise on CER systems remains insufficiently explored.In this study, we argue that incorporating inter-rater ambiguity during training can regularize the noise response, leading to noise robustness.To this end, we propose a novel loss function that incorporates inter-rater ambiguity into model training.Experiments conducted on the RECOLA dataset demonstrate that our proposed method achieves a maximum Concordance Correlation Coefficient (CCC) improvement of 0.117 and 0.077 for mean and standard deviation predictions, respectively, across all noise conditions.We further integrate traditional noisy augmentation strategies with our proposed method and observe promising results.
Ya-Tse Wu, Jingyao Wu 0002, Vidhyasaharan Sethu, Chi-Chun Lee
INTERSPEECH4
2024 RW-VoiceShield: Raw Waveform-based Adversarial Attack on One-shot Voice Conversion
Ching-Yu Yang, Shreya G. Upadhyay, Ya-Tse Wu, Bo-Hao Su, Chi-Chun Lee
INTERSPEECH5
2024 Embracing Ambiguity And Subjectivity Using The All-Inclusive Aggregation Rule For Evaluating Multi-Label Speech Emotion Recognition Systems
abstract
Speech Emotion Recognition (SER) faces a distinct challenge compared to other speech-related tasks because the annotations will show the subjective emotional perceptions of different annotators. Previous SER studies often view the subjectivity of emotion perception as noise by using the majority rule or plurality rule to obtain the consensus labels. However, these standard approaches overlook the valuable information of labels that do not agree with the consensus and make it easier for the test set. Emotion perception can have co-occurring emotions in realistic conditions, and it is unnecessary to regard the disagreement between raters as noise. To bridge the SER into a multi-label task, we introduced an “all-inclusive rule,” which considers all available data, ratings, and distributional labels as multi-label targets and a complete test set. We demonstrated that models trained with multi-label targets generated by the proposed AR outperform conventional single-label methods across incomplete and complete test sets.
Huang-Cheng Chou, Lucas Goncalves, Seong-Gyun Leem, Ali N. Salman, Carlos Busso, Hung-yi Lee, Chi-Chun Lee
SLT8
2024 Open-Emotion: A Reproducible EMO-Superb For Speech Emotion Recognition Systems
abstract
Speech emotion recognition (SER) is an essential technology for human-computer interaction systems. However, the previous study reveals that 80.77% of SER papers yield results that cannot be reproduced on the well-known IEMOCAP dataset. The main reason for reproducibility challenges is that the database did not provide standard data splits (e.g., train, development, and test sets). Prior papers could define its partition, but they did not provide details of the partition or source code for processing the partition. Therefore, this work aims to make SER open and reproducible to everyone. We develop the EMO-SUPERB, shorted for EMOtion Speech Universal PERformance Benchmark, including a user-friendly codebase to leverage 16 state-of-the-art (SOTA) speech self-supervised learning models for exhaustive evaluation plus one SOTA SER model across 6 open-source SER datasets in English and Chinese. We make all resources open-source to facilitate future developments in SER. Researchers can easily upload their systems or datasets to EMO-SUPERB, and we name the project “Open-Emotion”.
Huang-Cheng Chou, Kai-Wei Chang 0001, Lucas Goncalves, Jiawei Du 0003, Jyh-Shing Roger Jang, Chi-Chun Lee, Hung-yi Lee
SLT7
2024 Learning With Rater-Expanded Label Space to Improve Speech Emotion Recognition
abstract
Automatic sensing of emotional information in speech is important for numerous everyday applications. Conventional Speech Emotion Recognition (SER) models rely on averaging or consensus of human annotations for training, but emotions and raters' interpretations are subjective in nature, leading to diverse variations in perceptions. To address this, our proposed approach integrates the rater's subjectivity by forming the Perception-Coherent Clusters (PCC) of raters to be used to derive expanded label space for learning to improve SER. We evaluate our method on the IEMOCAP and the MSP-Podcast corpora, considering scenarios of fixed and variable raters, respectively. The proposed architecture, Rater Perception Coherency (RPC)-based SER surpasses single-task models with consensus labels by achieving UAR improvements of 3.39% for the IEMOCAP and 2.03% for the MSP-Podcast. Further analysis provides comprehensive insights into the contributions of these perception consistency clusters in SER learning.
Shreya G. Upadhyay, Woan-Shiuan Chien, Bo-Hao Su, Chi-Chun Lee
IEEE Trans. Affect. Comput.4
2024 Using Measures of Vowel Space for Autistic Traits Characterization
abstract
Autism Spectrum Disorder (ASD) is a neurodevelopmental disorder that is prevalent and heterogeneous. Autistic traits describe a wide heterogeneity of behavior symptoms of ASD, and these traits are reflections of core neurodevelopment function deficits. Researchers have predominantly taken a clinical angle to understand autistic traits. They have been developing various clinical-grade instruments with behavioral codes to quantify autistic traits for diagnostic and research purposes. However, the need for highly trained professionals and the inevitable subjectivity limit their usage. Hence, researchers have been developing computational methods to address these issues. Among many efforts, methods based on computing speech have emerged rapidly due to their ability to characterize communicative behaviors and social interactions. Our work addresses one particular under-studied speech aspect: articulation-related acoustics, one of the broad autism spectrum symptoms. In this paper, we examine the articulatory information in a natural spoken interaction through measures of vowel space characteristics (VSCs) to understand autistic traits. Specifically, we approach by modeling statistical relationships of the corner vowel distributions and the interpersonal correlation of these relationships in conversation. Our method is evaluated by deriving VSC features and using them in ASD classification and regression tasks. We found these features predict autism-related communication assessment and add additional information to classification tasks. Furthermore, our analyses show a relationship between VSCs and autism-related communication deficit and also imply differences in VSCs between typical developing people and each ASD subgroup.
Chin-Po Chen, Ho-hsien Pan, Susan Shur-Fen Gau, Chi-Chun Lee
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Analyzing the Effect of Affective Priming on Emotional Annotations
abstract
In the field of affective computing, emotional annotations are highly important for both the recognition and synthesis of human emotions. Researchers must ensure that these emotional labels are adequate for modeling general human perception. An unavoidable part of obtaining such labels is that human annotators are exposed to known and unknown stimuli before and during the annotation process that can affect their perception. Emotional stimuli cause an affective priming effect, which is a pre-conscious phenomenon in which previous emotional stimuli affect the emotional perception of a current target stimulus. In this paper, we use sequences of emotional annotations during a perceptual evaluation to study the effect of affective priming on emotional ratings of speech. We observe that previous emotional sentences with extreme emotional content push annotations of current samples to the same extreme. We create a sentence-level bias metric to study the effect of affective priming on speech emotion recognition(SER) modeling. The metric is used to identify subsets in the database with more affective priming bias intentionally creating biased datasets. We train and test SER models using the full and biased datasets. Our results show that although the biased datasets have low inter-evaluator agreements, SER models for arousal and dominance trained with those datasets perform the best. For valence, the models trained with the less-biased datasets perform the best.
Luz Martinez-Lucas, Ali N. Salman, Seong-Gyun Leem, Shreya G. Upadhyay, Chi-Chun Lee, Carlos Busso
ACII5
2023 An Intelligent Infrastructure Toward Large Scale Naturalistic Affective Speech Corpora Collection
abstract
The field of speech emotion recognition (SER) aims to create scientifically rigorous systems that can reliably characterize emotional behaviors expressed in speech. A key aspect for building SER systems is to obtain emotional data that is both reliable and reproducible for practitioners. However, academic researchers encounter difficulties in accessing or collecting naturalistic large-scale, reliable emotional recordings. Also, the best practices for data collection are not necessarily described or shared when presenting emotional corpora. To address this issue, the paper proposes the creation of an affective naturalistic database consortium (AndC) that can encourage multidisciplinary cooperation among researchers and practitioners in the field of affective computing. This paper’s contribution is twofold. First, it proposes the design of the AndC with a customizable-standard framework for intelligently-controlled emotional data collection. The focus is on leveraging naturalistic spontaneous recordings available on audio-sharing websites. Second, it presents as a case study the development of a naturalistic large-scale Taiwanese Mandarin podcast corpus using the customizable-standard intelligently-controlled framework. The AndC will enable research groups to effectively collect data using the provided pipeline and to contribute with alternative algorithms or data collection protocols.
Shreya G. Upadhyay, Woan-Shiuan Chien, Bo-Hao Su, Lucas Goncalves, Ya-Tse Wu, Ali N. Salman, Carlos Busso, Chi-Chun Lee
ACII8
2023 Achieving Fair Speech Emotion Recognition via Perceptual Fairness
abstract
Speech emotion recognition (SER) is a key technological module to be integrated into many voice-based solutions. One of the unique fairness issues in SER is caused by the inherently biased emotion perception given by the raters as ground truth labels. Mitigating rater biases are at core for SER to move toward optimizing both recognition and fairness performance. In this work, we proposed a two-stage framework, which produces debiased representations by using a fairness constraint adversarial framework in the first stage. Then, users are endued with the right to toggle between specified gender-wise perceptions on-demand after the gender-wise perceptual learning in the second stage. We further evaluate our results on two important fairness metrics to show that the distributions and predictions across different gender are fair.
Woan-Shiuan Chien, Chi-Chun Lee
ICASSP2
2023 Phonetic Anchor-Based Transfer Learning to Facilitate Unsupervised Cross-Lingual Speech Emotion Recognition
abstract
Modeling cross-lingual speech emotion recognition (SER) has become more prevalent because of its diverse applications. Existing studies have mostly focused on technical approaches that adapt the feature, domain, or label across languages, without considering in detail the similarities between the languages. This study focuses on domain adaptation in cross-lingual scenarios using phonetic constraints. This work is framed in a twofold manner. First, we analyze emotion-specific phonetic commonality across languages by identifying common vowels that are useful for SER modeling. Second, we leverage these common vowels as an anchoring mechanism to facilitate cross-lingual SER. We consider American English and Taiwanese Mandarin as a case study to demonstrate the potential of our approach. This work uses two in-the-wild natural emotional speech corpora: MSP-Podcast (American English), and BIIC-Podcast (Taiwanese Mandarin). The proposed unsupervised cross-lingual SER model using these phonetical anchors outperforms the baselines with a 58.64% of unweighted average recall (UAR).
Shreya G. Upadhyay, Luz Martinez-Lucas, Bo-Hao Su, Woan-Shiuan Chien, Ya-Tse Wu, William F. Katz, Carlos Busso, Chi-Chun Lee
ICASSP9
2023 The Importance of Calibration: Rethinking Confidence and Performance of Speech Multi-label Emotion Classifiers
Huang-Cheng Chou, Lucas Goncalves, Seong-Gyun Leem, Chi-Chun Lee, Carlos Busso
INTERSPEECH4
2023 Noise-Robust Bandwidth Expansion for 8K Speech Recordings
Yin-Tse Lin, Bo-Hao Su, Chi-Han Lin, Shih-Chan Kuo, Jyh-Shing Roger Jang, Chi-Chun Lee
INTERSPEECH6
2023 Speaking State Decoder with Transition Detection for Next Speaker Prediction
Shao-Hao Lu, Yun-Shao Lin, Chi-Chun Lee
INTERSPEECH3
2023 A Context-Constrained Sentence Modeling for Deception Detection in Real Interrogation
Ya-Tse Wu, Yuan-Ting Chang, Shao-Hao Lu, Jing-Yi Chuang, Chi-Chun Lee
INTERSPEECH5
2023 MetricAug: A Distortion Metric-Lead Augmentation Strategy for Training Noise-Robust Speech Emotion Recognizer
Ya-Tse Wu, Chi-Chun Lee
INTERSPEECH2
2023 An Engineering View on Emotions and Speech: From Analysis and Predictive Models to Responsible Human-Centered Applications
abstract
The substantial growth of Internet-of-Things technology and the ubiquity of smartphone devices has increased the public and industry focus on speech emotion recognition (SER) technologies. Yet, conceptual, technical, and societal challenges restrict the wide adoption of these technologies in various domains, including, healthcare, and education. These challenges are amplified when automated emotion recognition systems are called to function “in-the-wild” due to the inherent complexity and subjectivity of human emotion, the difficulty of obtaining reliable labels at high temporal resolution, and the diverse contextual and environmental factors that confound the expression of emotion in real life. In addition, societal and ethical challenges hamper the wide acceptance and adoption of these technologies, with the public raising questions about user privacy, fairness, and explainability. This article briefly reviews the history of affective speech processing, provides an overview of current state-of-the-art approaches to SER, and discusses algorithmic approaches to render these technologies accessible to all, maximizing their benefits and leading to responsible human-centered computing applications.
Chi-Chun Lee, Theodora Chaspari, Emily Mower Provost, Shri Narayanan
Proc. IEEE1
2023 Enforcing Semantic Consistency for Cross Corpus Emotion Prediction Using Adversarial Discrepancy Learning in Emotion
abstract
Mismatch between databases entails a challenge in performing emotion recognition on a practical-condition unlabeled database with labeled source data. The alignment between the source and target is crucial for conventional neural network; therefore, many studies have mapped two domains in a common feature space. However, the effect of distortion in emotion semantics across different conditions has been neglected in such work, and a sample from the target may be considered a high emotional annotation in the target but as low in the source. In this article, we propose the maximum regression discrepancy (MRD) network, which enforces semantic consistency in a source and target by adjusting the acoustic feature encoder to minimize discrepancy in maximally distorted samples through adversarial training. We show our framework in several experiments using three databases (the USC IEMOCAP, MSP-Improv, and MSP-Podcast) for cross corpus emotion prediction. Compared to the Source-only neural network and DANN, MRD network demonstrates a significant improvement between 5% and 10% in the concordance correlation coefficient (CCC) in cross-corpus prediction and between 3% and 10% for evaluation on MSP-PODCAST. We also visualize the effect of MRD on feature representation to shows the efficacy of the MRD structure we designed.
Chun-Min Chang, Gao-Yi Chao, Chi-Chun Lee
IEEE Trans. Affect. Comput.3
2023 Learning Enhanced Acoustic Latent Representation for Small Scale Affective Corpus with Adversarial Cross Corpora Integration
abstract
Achieving robust cross contexts speech emotion recognition (SER) has become a critical next direction of research for wide adoption of SER technology. The core challenge is in the large variability of affective speech that is highly contextualized. Prior works have worked on this as a transfer learning problem that mostly focuses on developing domain adaptation strategy. However, many of the existing speech emotion corpora, even those considered as large scale, are still limited in size resulting in an unsatisfactory transfer result. On the other hand, directly collecting context-specific corpus often results in an even smaller data size leading to an inevitably non-robust accuracy. In order to mitigate this issue, we propose the concept of enhancing the affect-related variability when learning thein-contextacoustic latent representation by integratingout-of-contextemotion data. Specifically, we utilize adversarial autoencoder network as our backbone with multipleout-of-contextemotion labels derived for eachin-contextsamples that serve as an auxiliary constraint in learning the latent representation. We extensively evaluate our framework using threein-contextdatabases with threeout-of-contextdatabases. In this work, we demonstrate not only an improved recognition accuracy but also a comprehensive analysis on the effectiveness of this representation learning strategy.
Chun-Min Chang, Chi-Chun Lee
IEEE Trans. Affect. Comput.2
2023 An Enroll-to-Verify Approach for Cross-Task Unseen Emotion Class Recognition
abstract
Most speech emotion recognition studies often focus on recognizing pre-set emotion classes. However, the task definition may change due to a shift in focus to a previously unseen class in real-world applications. This cross-task modeling has not been addressed previously. Lengthy data re-collection, model retraining, and the traditional adaptation and transfer learning approaches are not applicable to this cross-task setting. This study proposes an enroll-to-verify framework to avoid model retraining and rapidly perform a new task prediction using only a handful of enrolled samples. Specifically, we use negative angular margin prototypical loss in a pretrained multiclass network as an emotion encoder. Then, we enroll a few samples corresponding to emotion classes in the new task definition and simply compare the encoded embedding distance to perform recognition. In the experiments on the IEMOCAP dataset, given a four-class pretrained emotion encoder, we achieved a 71.9% unweighted average recall in the frustration (unseen) recognition task. The MELD dataset was used where the unseen class was surprise, fear, or disgust. The results revealed that enrolling only 20 samples without retraining was comparable to supervised training using the complete dataset. Further analyses were conducted to demonstrate the working mechanism of our proposed enroll-to-verify approach.
Jeng-Lin Li, Chi-Chun Lee
IEEE Trans. Affect. Comput.2
2023 Unsupervised Cross-Corpus Speech Emotion Recognition Using a Multi-Source Cycle-GAN
abstract
Speech emotion recognition (SER) plays a crucial role in understanding user feelings when developing artificial intelligence services. However, the data mismatch and label distortion between the training (source) set and the testing (target) set significantly degrade the performances when developing the SER systems. Additionally, most emotion-related speech datasets are highly contextualized and limited in size. The manual annotation cost is often too high leading to an active investigation of unsupervised cross-corpus SER techniques. In this paper, we propose a framework in unsupervised cross-corpus emotion recognition using multi-source corpus in a data augmentation manner. We introduced Corpus-Aware Emotional CycleGAN (CAEmoCyGAN) including a corpus-aware attention mechanism to aggregate each source datasets to generate the synthetic target sample. We choose the widely used speech emotion corpora the IEMOCAP and the VAM as sources and the MSP-Podcast as the target. By generating synthetic target-aware samples to augment source datasets and by directly training on this augmented dataset, our proposed multi-source target-aware augmentation method outperforms other baseline models in activation and valence classification.
Bo-Hao Su, Chi-Chun Lee
IEEE Trans. Affect. Comput.2
2023 A Media-Guided Attentive Graphical Network for Personality Recognition Using Physiology
abstract
Physiological automatic personality recognition has been largely developed to model an individual’s personality trait from a variety of signals. However, few studies have tackled the problems of integration methodology from multiple observations into a single personality prediction. In this study, we focus on finding a novel learning architecture to model the personality trait under aMany-to-Onescenario. We propose to integrate not only the information on the user but also consider the effect of the affective multimedia stimulus. Specifically, we present a novel Acoustic-Visual Guided Attentive Graph Convolutional Network for enhanced personality recognition. The emotional multimedia content guides the formation of the physiological responses into a graph-like structure to integrate latent inter-correlation among all responses toward affective multimedia. Then these graphs would be further processed by the Graph Convolutional Network (GCN) to jointly model instances and inter-correlation levels of the subject’s responses. We show that our model outperforms the current state of the art on two large public corpora for personality recognition. Further analysis reveals that there indeed exists a multimedia preference for inferring personality from physiology, and several frequency-domain descriptors in ECG and the tonic component in EDA are shown to be robust for automatic personality recognition.
Hao-Chun Yang, Chi-Chun Lee
IEEE Trans. Affect. Comput.2
2023 Lung Cancer Prediction Using Electronic Claims Records: A Transformer-Based Approach
abstract
Electronic claims records (ECRs) are large scale and longitudinal collections of individual's medical service seeking actions. Compared to in-hospital medical records (EMRs), ECRs are more standardized and cross-sites. Recently, there has been studies showing promising results on modeling claims data for a wide range of medical applications. However, few of them address the exclusion criteria on cohort selection to extract new incidence without prior signs and also often lack of emphasis on predicting cancer in early stages. In this work, we aim to design a lung cancer prediction framework using ECRs with rigorous exclusion design using state-of-the-art sequence-based transformer. Furthermore, this work presents one of the first results by applying disease prediction model to the entire population in Taiwan. The result shows over 2.1 predictive power, 5 average positive predictive value (PPV), and 0.668 area under curve (AUC) in all-stage lung cancer and around 2.0 predictive power, 1 average PPV and 0.645 AUC in early-stage in our dataset. Sub-cohort analysis could funnel high precision selective group into prioritized clinical examination. Onset analysis validates the effect of our exclusion criteria. This work presents comprehensive analyses on lung cancer prediction, and the proposed approach can serve as a state-of-the-art disease risk prediction framework on claims data.
Huan-Yu Chen, Hui-Min Wang, Ching-Heng Lin, Rob Yang, Chi-Chun Lee
IEEE J. Biomed. Health Informatics5
2023 An Interaction-process-guided Framework for Small-group Performance Prediction
abstract
A small group is a fundamental interaction unit for achieving a shared goal. Group performance can be automatically predicted using computational methods to analyze members’ verbal behavior in task-oriented interactions, as has been proven in several recent works. Most of the prior works focus on lower-level verbal behaviors, such as acoustics and turn-taking patterns, using either hand-crafted features or even advanced end-to-end methods. However, higher-level group-based communicative functions used between group members during conversations have not yet been considered. In this work, we propose a two-stage training framework that effectively integrates the communication function, as defined using Bales’s interaction process analysis (IPA) coding system, with the embedding learned from the low-level features in order to improve the group performance prediction. Our result shows a significant improvement compared to the state-of-the-art methods (4.241 MSE and 0.341 Pearson’s correlation on NTUBA-task1 and 3.794 MSE and 0.291 Pearson’s correlation on NTUBA-task2) on the National Taiwan University Business Administration (NTUBA) small-group interaction database. Furthermore, based on the design of IPA, our computational framework can provide a time-grained analysis of the group communication process and interpret the beneficial communicative behaviors for achieving better group performance.
Yun-Shao Lin, Yi-Ching Liu, Chi-Chun Lee
ACM Trans. Multim. Comput. Commun. Appl.3
2022 Monologue versus Conversation: Differences in Emotion Perception and Acoustic Expressivity
abstract
Advancing speech emotion recognition (SER) depends highly on the source used to train the model, i.e., the emotional speech corpora. By permuting different design parameters, researchers have released versions of corpora that attempt to provide a better-quality source for training SER. In this work, we focus on studying communication modes of collection. In particular, we analyze the patterns of emotional speech collected during interpersonal conversations or monologues. While it is well known that conversation provides a better protocol for eliciting authentic emotion expressions, there is a lack of systematic analyses to determine whether conversational speech provide a “better-quality” source. Specifically, we examine this research question from three perspectives: perceptual differences, acoustic variability and SER model learning. Our analyses on the MSP-Podcast corpus show that: 1) rater's consistency for conversation recordings is higher when evaluating categorical emotions, 2) the perceptions and acoustic patterns observed on conversations have properties that are better aligned with expected trends discussed in emotion literature, and 3) a more robust SER model can be trained from conversational data. This work brings initial evidences stating that samples of conversations may provide a better-quality source than samples from monologues for building a SER model.
Woan-Shiuan Chien, Shreya G. Upadhyay, Ya-Tse Wu, Bo-Hao Su, Carlos Busso, Chi-Chun Lee
ACII7
2022 Romantic and Family Movie Database: Towards Understanding Human Emotion and Relationship via Genre-Dependent Movies
abstract
Movie gains popularity by successfully immersing viewers into affective contents under an emotion perception and elicitation process. Film genre sets the tone of a movie to shape key emotional context during the storytelling procedure. In romantic and family movies, character relationship is a key contextual attribute guiding the whole storyline, which is expected to evoke a sense of being moved in the audience. In this work, we propose Romantic and Family Movie Database, which consists of 1029 movie clips from 10 romantic and family movies. We provide annotations including character relationship type, character relationship status, perceived emotions, induced emotions, and degree of being moved. Our analysis elaborates the inconsistency between perceived and induced valence with the support of annotations including character relationship status and degree of being moved. We also find that being moved and induced arousal are positively correlated in romantic and family movies. We present comprehensive baseline results in recognizing relationship status and emotions using different models and modalities. The database is publicly available at https://rfmd.ee.nthu.edu.tw/.
Po-Chien Hsu, Jeng-Lin Li, Chi-Chun Lee
ACII3
2022 Exploiting Annotators' Typed Description of Emotion Perception to Maximize Utilization of Ratings for Speech Emotion Recognition
abstract
The decision of ground truth for speech emotion recognition (SER) is still a critical issue in affective computing tasks. Previous studies on emotion recognition often rely on consensus labels after aggregating the classes selected by multiple annotators. It is common for a perceptual evaluation conducted to annotate emotional corpora to include the class “other,” allowing the annotators the opportunity to describe the emotion with their own words. This practice provides valuable emotional information, which, however, is ignored in most emotion recognition studies. This paper utilizes easy-accessed natural language processing toolkits to mine the sentiment of these typed descriptions, enriching and maximizing the information obtained from the annotators. The polarity information is combined with primary and secondary annotations provided by individual evaluators under a label distribution framework, creating a complete representation of the emotional content of the spoken sentences. Finally, we train multitask learning SER models with existing learning methods (soft-label, multi-label, and distribution-label) to show the performance of the novel ground truth in the MSP-Podcast corpus.
Huang-Cheng Chou, Chi-Chun Lee, Carlos Busso
ICASSP3
2022 An Audio-Saliency Masking Transformer for Audio Emotion Classification in Movies
abstract
The process of perception to affective response of humans is gated by a bottom-up saliency mechanism at the sensory level. In specifics, auditory saliency emphasizes audio segments that need to be attended to cognitively appraise and experience emotion. In this work, inspired by this mechanism, we propose an end-to-end feature masking network for audio emotion recognition in movies. Our proposed Audio-Saliency Masking Transformer (ASTM) adjusts feature embedding using two learnable masks; one of them cross-refers to an auditory saliency map, and the other one is through self-reference. By joint training for front-end mask gating and the transformer as the back-end emotion classifier, we achieve three-class UARs improvement of 1.74%, 1.27%, 0.95%, 0.82% when comparing to the best of the other models on experienced arousal, experienced valence, intended arousal, and intended valence, respectively. We further analyze which acoustic feature categories that our saliency mask attends to the most.
Ya-Tse Wu, Jeng-Lin Li, Chi-Chun Lee
ICASSP3
2022 Emotion-Shift Aware CRF for Decoding Emotion Sequence in Conversation
Yun-Shao Lin, Chi-Chun Lee
INTERSPEECH3
2022 Exploiting Co-occurrence Frequency of Emotions in Perceptual Evaluations To Train A Speech Emotion Classifier
Huang-Cheng Chou, Chi-Chun Lee, Carlos Busso
INTERSPEECH2
2022 An Attention-Based Method for Guiding Attribute-Aligned Speech Representation Learning
Yu-Lin Huang, Bo-Hao Su, Yao-Win Peter Hong, Chi-Chun Lee
INTERSPEECH4
2022 Vaccinating SER to Neutralize Adversarial Attacks with Self-Supervised Augmentation Strategy
Bo-Hao Su, Chi-Chun Lee
INTERSPEECH2
2022 A Chunking-for-Pooling Strategy for Cytometric Representation Learning for Automatic Hematologic Malignancy Classification
abstract
Differentiating types of hematologic malignancies is vital to determine therapeutic strategies for the newly diagnosed patients. Flow cytometry (FC) can be used as diagnostic indicator by measuring the multi-parameter fluorescent markers on thousands of antibody-bound cells, but the manual interpretation of large scale flow cytometry data has long been a time-consuming and complicated task for hematologists and laboratory professionals. Past studies have led to the development of representation learning algorithms to perform sample-level automatic classification. In this work, we propose a chunking-for-pooling strategy to include large-scale FC data into a supervised deep representation learning procedure for automatic hematologic malignancy classification. The use of discriminatively-trained representation learning strategy and the fixed-size chunking and pooling design are key components of this framework. It improves the discriminative power of the FC sample-level embedding and simultaneously addresses the robustness issue due to an inevitable use of down-sampling in conventional distribution based approaches for deriving FC representation. We evaluated our framework on two datasets. Our framework outperformed other baseline methods and achieved 92.3% unweighted average recall (UAR) for four-class recognition on the UPMC dataset and 85.0% UAR for five-class recognition on the hema.to dataset. We further compared the robustness of our proposed framework with that of the traditional downsampling approach. Analysis of the effects of the chunk size and the error cases revealed further insights about different hematologic malignancy characteristics in the FC data.
Jeng-Lin Li, Yun-Chun Lin, Yu-Fen Wang, Sara A. Monaghan, Bor-Sheng Ko, Chi-Chun Lee
IEEE J. Biomed. Health Informatics6
2022 A Social Condition-Enhanced Network for Recognizing Power Distance Using Expressive Prosody and Intrinsic Brain Connectivity
abstract
Culture is the social norm that often dictates a person's thoughts, decision-making, and social behaviors during interaction at an individual level. In this study, we present a computational framework that automatically assesses an individual culture attribute of power distance (PDI), i.e., the measure to describe one's acceptance of social status, power and authority in organizations through multimodal modeling of a participant's expressive prosodic structures and brain connectivity using a social condition-enhanced network. In specific, we propose a joint learning approach of center-loss embedding network architecture that learns to "centerize" the embedding space given a particular social interaction condition to enhance the PDI discriminability of the representation. Our proposed method achieves 88.5% and 73.1% in binary classification task of recognizing low versus high power distance on prosodic and fMRI modality separately. After performing multimodal fusion, it improves to 96.2% of 2-class recognition rate (7.7% relative improvement). Further analyses reveal that average and standard deviation of speech energy are significantly correlated with power distance index; the right middle cingulate cortex (MCC) of brain region achieves the best recognition accuracy demonstrating its role in processing a person's belief about power distance.
Fu-Sheng Tsai, Wei-Wen Chang, Chi-Chun Lee
IEEE Trans. Multim.3
2021 An Attribute-Aligned Strategy for Learning Speech Representation
abstract
Advancement in speech technology has brought convenience to our life. However, the concern is on the rise as speech signal contains multiple personal attributes, which would lead to either sensitive information leakage or bias toward decision. In this work, we propose an attribute-aligned learning strategy to derive speech representation that can flexibly address these issues by attribute-selection mechanism. Specifically, we propose a layered-representation variational autoencoder (LR-VAE), which factorizes speech representation into attribute-sensitive nodes, to derive an identity-free representation for speech emotion recognition (SER), and an emotionless representation for speaker verification (SV). Our proposed method achieves competitive performances on identity-free SER and a better performance on emotionless SV, comparing to the current state-of-the-art method of using adversarial learning applied on a large emotion corpora, the MSP-Podcast. Also, our proposed learning strategy reduces the model and training process needed to achieve multiple privacy-preserving tasks.
Yu-Lin Huang, Bo-Hao Su, Yao-Win Peter Hong, Chi-Chun Lee
Interspeech4
2021 Through the Words of Viewers: Using Comment-Content Entangled Network for Humor Impression Recognition
abstract
Research into understanding humor has been investigated over centuries. It has recently attracted various technical effort in computing humor automatically from data, especially for humor in speech. Comprehension on the same speech and the ability to realize a humor event vary depending on each individual audience's background and experience. Most previous works on automatic humor detection or impression recognition mainly model the produced textual content only without considering audience responses. We collect a corpus of TED Talks including audience comments for each of the presented TED speech. We propose a novel network architecture that considers the natural entanglement between speech transcripts and user's online feedbacks as an integrative graph structure, where the content speech and online feedbacks are nodes where the edges are connected though their common words. Our model achieves 61.2% of accuracy in a three-class classification on humor impression recognition on TED talks; our experiments further demonstrate viewers comments are essential in improving the recognition tasks, and a joint content-comment modeling achieves the best recognition.
Huan-Yu Chen, Yun-Shao Lin, Chi-Chun Lee
SLT3
2021 A Conditional Cycle Emotion Gan for Cross Corpus Speech Emotion Recognition
abstract
Speech emotion recognition (SER) is important in enabling personalized services and multimedia applications in our life. It also becomes a prevalent topic of research with its potential in creating a better user experience across many modern technologies. However, the highly contextualized scenario and expensive emotion labeling required cause a severe mismatch between already limited-in-scale speech emotional corpora; this hinders the wide adoption of SER. In this work, instead of conventionally learning a common feature space between corpora, we take a novel approach in enhancing the variability of the source (labeled) corpus that is target (unlabeled) data-aware by generating synthetic source domain data using a conditional cycle emotion generative adversarial network (CCEmoGAN). Note that no target samples with label are used during whole training process. We evaluate our framework in cross corpus emotion recognition tasks and obtain a three classes valence recognition accuracy of 47.56%, 50.11% and activation accuracy of 51.13%, 65.7% when transferring from the IEMOCAP to the CIT dataset, and the IEMOCAP to the MSP-IMPROV dataset respectively. The benefit of increasing target domain-aware variability in the source domain to improve emotion discriminability in cross corpus emotion recognition is further visualized in our augmented data space.
Bo-Hao Su, Chi-Chun Lee
SLT2
2020 Conditional Domain Adversarial Transfer for Robust Cross-Site ADHD Classification Using Functional MRI
abstract
There is a growing number of large scale cross-site database collection of resting-state functional magnetic resonance imaging (rs-fMRI) for studying neurobehavioral diseases, such as ADHD. Although a large amount of data benefits machine learning-based classification methods, the idiosyncratic variability of each site can deteriorate cross-site generalization ability. This challenge creates a bottleneck in requiring a large number of labeled samples of each site. Hence in this research, we utilize an approach of conditional adversarial domain adaptation network (CDAN) to learn a discriminative fMRI representation that is site-invariant for unsupervised transfer of ADHD classification. We evaluate our framework on a multi-site ADHD dataset and achieve improvement in transferring between sites. Further visualization reveals that there indeed exists a substantial site discrepancy and statistically analysis indicates that male’s rs-fMRI could be more vulnerable toward site-specific effects.
Ya-Lin Huang, Wan-Ting Hsieh, Hao-Chun Yang, Chi-Chun Lee
ICASSP4
2020 Predicting Performance Outcome with a Conversational Graph Convolutional Network for Small Group Interactions
abstract
Studying behaviors of members during small group interaction provides objective insights in improving the efficiency of the decision making process in our daily working life. By introducing the use of the graph structure in modeling the natural inter-member conversational ties during such an interaction, we aim to advance the state-of-art computational approach in predicting group performance scores. Specifically, we proposed a Conversational Graph Convolutional Network (CGCN) that utilizes conversation dynamic as the graph to aggregate group member's speech and lexical behaviors in predicting the group performance. Our result shows that Speech CGCN achieves the state-of-the-art performance at MSE 3.896 (0.323 Pearson correlation) outperform the current best method in ELEA dataset. Our model additionally reveals that an imbalance conversational graph structure is positively correlated to group performances.
Yun-Shao Lin, Chi-Chun Lee
ICASSP2
2020 A Siamese Content-Attentive Graph Convolutional Network for Personality Recognition Using Physiology
abstract
Affective multimedia content has long been used as stimulation to study an individual's personality using physiology. In this work, we propose a novel Siamese Content-Attentive Graph Convolutional Network (SCA-GCN) to learn a discriminative physiology representation jointly guided by the actual video content of the emotional stimuli. The visual content of the stimuli is integrated into learning to weight the importance of physiology in the task of personality recognition. We evaluate our framework on a large public corpus of physiological data. Our method achieves the state of the art unweighted accuracy of 72.1%,69.5%, and 68.2% in a binary classification for dimensions of Openness, Emotion Stability, and Extraversion, which improves over the baseline DNN by 20.4%, 9%, and 13.9%. Further analysis reveals that there indeed exists a substantial effect from the media content in affecting the subject's internal physiological responses that result in an improved personality recognition performances.
Hao-Chun Yang, Chi-Chun Lee
ICASSP2
2020 A Dialogical Emotion Decoder for Speech Motion Recognition in Spoken Dialog
abstract
Developing a robust emotion speech recognition (SER) system for human dialog is important in advancing conversational agent design. In this paper, we proposed a novel inference algorithm, a dialogical emotion decoding (DED) algorithm, that treats a dialog as a sequence and consecutively decode the emotion states of each utterance over time with a given recognition engine. This decoder is trained by incorporating intra- and inter-speakers emotion influences within a conversation. Our approach achieves a 70.1% in four class emotion on the IEMOCAP database, which is 3% over the state-of-art model. The evaluation is further conducted on a multi-party interaction database, the MELD, which shows a similar effect. Our proposed DED is in essence a conversational emotion rescoring decoder that can also be flexibly combined with different SER engines.
Sung-Lin Yeh, Yun-Shao Lin, Chi-Chun Lee
ICASSP3
2020 Learning Converse-Level Multimodal Embedding to Assess Social Deficit Severity for Autism Spectrum Disorder
abstract
Developing algorithms to automatically assess mental constructs using media data of human behaviors is becoming important, especially relevant for mental health applications. In this work, we focus on a critically prevalent neurodevelopment disorder, i.e., the autism spectrum disorder (ASD). While researchers have worked on automatic differentiation of ASD from healthy control using a variety of behavior modalities, few works have modeled the severity of ASD behavior symptoms in the existing clinical practice. Thus, we propose to learn a converse-level multimodal (speech and text) embedding derived during a severity assessment interview, i.e., the Autism Diagnosis Observation Schedule (ADOS), that considers the intricate interaction behaviors between the investigator and the participant. Further by fusing two attentional GRUs with this multimodal embedding, our approach achieves an averaged regression score of 0.567 on four items of socio-communicative constructs in the ADOS. Our analysis results suggest that the number of words uttered by both the investigator and the participant is a major predictor.
Chin-Po Chen, Susan Shur-Fen Gau, Chi-Chun Lee
ICME3
2020 Learning to Recognize Per-Rater's Emotion Perception Using Co-Rater Training Strategy with Soft and Hard Labels
Huang-Cheng Chou, Chi-Chun Lee
INTERSPEECH2
2020 Using Speaker-Aligned Graph Memory Block in Multimodally Attentive Emotion Recognition Network
Jeng-Lin Li, Chi-Chun Lee
INTERSPEECH2
2020 Improving Speech Emotion Recognition Using Graph Attentive Bi-Directional Gated Recurrent Unit Network
Bo-Hao Su, Chun-Min Chang, Yun-Shao Lin, Chi-Chun Lee
INTERSPEECH4
2020 Attentive Convolutional Recurrent Neural Network Using Phoneme-Level Acoustic Representation for Rare Sound Event Detection
Shreya G. Upadhyay, Bo-Hao Su, Chi-Chun Lee
INTERSPEECH3
2020 Speech Representation Learning for Emotion Recognition Using End-to-End ASR with Factorized Adaptation
Sung-Lin Yeh, Yun-Shao Lin, Chi-Chun Lee
INTERSPEECH3
2020 Predicting Collaborative Task Performance Using Graph Interlocutor Acoustic Network in Small Group Interaction
Shun-Chang Zhong, Bo-Hao Su, Yi-Ching Liu, Chi-Chun Lee
INTERSPEECH5
2020 Cross Corpus Physiological-based Emotion Recognition Using a Learnable Visual Semantic Graph Convolutional Network
abstract
Affective media videos have been used as stimulus to investigate an individual's affective-physio responses. In this study, we aim to develop a network learning strategy for robust cross-corpus emotion recognition using physiological features jointly with affective video content. Specifically, we present a novel framework of Visual Semantic Graph Learning Convolutional Network (VGLCN) for individual emotional state recognition using physiology on transfer learning tasks. The stimulus of videos content is integrated into learnable graph structure to weight the importance of physiology on the two emotion dimensions, valence and arousal. Furthermore, we evaluate our proposed framework on two public emotion databases with a rigorous cross validation method, and our model achieves the best unweighted average recall (UAR), which is 67.9%, 56.9% for arousal and 79.8%, 70.4% for valence on the cross datasets recognition experiments respectively. Further analyses reveal that 1) VGLCN is especially effective on transfer valence binary-task, 2) the physiological features (ECG, EDA) are very informative features for emotion recognition and 3) the affective media videos are important constraint to be included in the framework to stabilize the performance power.
Woan-Shiuan Chien, Hao-Chun Yang, Chi-Chun Lee
ACM Multimedia3
2020 Computational Analyses of Thin-Sliced Behavior Segments in Session-Level Affect Perception
abstract
The ability to accurately judge another person's emotional states with a short duration of observations is a unique perceptual mechanism of humans, termed as the thin-sliced judgment. In this work, we propose a computational framework based on mutual information to identify the thin-sliced emotion-rich behavior segments within each session and further use these segments to train the session-level affect regressors. Our proposed thin-sliced framework obtains regression accuracies measured in Spearman correlations of 0.605, 0.633, and 0.672 on session-level attributes of activation, dominance, and valence, respectively. It outperforms framework using data of the entire session as baseline. The significant improvement in the regression correlations reinforces the thin-sliced nature of human emotion perception. By properly extracting these emotion-rich behavior segments, we obtain not only an improved overall accuracy but also bring additional insights. Specifically, our detailed analyses indicate that this thin-sliced nature in emotion perception is more evident for attributes of activation and valence, and the within-session time distribution of emotion-salient behavior is located more toward the ending portion. Lastly, we observe that there indeed exists a certain set of behavior types that carry high emotion-related content, and this is especially apparent in the extreme emotion levels.
Chi-Chun Lee
IEEE Trans. Affect. Comput.2
2019 A Dual-Complementary Acoustic Embedding Network Learned from Raw Waveform for Speech Emotion Recognition
abstract
Speech emotion recognition (SER) technology has recently become a trend in a broader field and has achieved remarkable recognition performances using deep learning technique. However, the recognition performances obtained using end-to-end learning directly from raw audio waveform still hardly exceed those based on hand-crafted acoustic descriptors. Instead of solely rely on raw waveform or acoustic descriptors for SER, we propose an acoustic space augmentation network, termed as Dual-Complementary Acoustic Embedding Network (DCaEN), that combines knowledge-based features with raw waveform embedding learned with a novel complementary constraint. DCaEN includes representations from eGeMAPS acoustic feature and raw waveform by specifying a negative cosine distance loss to explicitly constrain the raw waveform embedding to be different from eGeMAPS. Our experimental results demonstrate an improved emotion discriminative power on the IEMOCAP database, which achieves 59.31% in a four class emotion recognition. Our analysis also demonstrates that the learned raw waveform embedding of DCaEN converges close to reverse mirroring of the original eGeMAPS space.
Tzu-Yun Huang, Jeng-Lin Li, Chun-Min Chang, Chi-Chun Lee
ACII4
2019 Attention Learning with Retrievable Acoustic Embedding of Personality for Emotion Recognition
abstract
Modeling multimodal behavior streams to automatically identify emotion states of an individual has progressed extensively especially with the advancement of deep learning algorithms. Emotion, being an abstract internal state, creates substantial differences in an individual's behavior expressivity, the development of personalized recognition framework is a critical next step to improve algorithm's modeling capacity. In this work, we propose to integrate the target speaker's personality embedding into the learning of multimodal (speech and language) attention based network architecture to improve recognition performances. Specifically, we propose a Personal Attribute-Aware Attention Network (PAaAN) that learns its multimodal attention weights jointly with the target speaker's retrievable acoustic embedding of personality. Our acoustic domain adapted personality retrieval strategy mitigates the common issue on the lack of personality scores in the current available emotion databases, and our proposed PAaAN then learns its attention weight by jointly considering an individual target speaker's personality profile with his or her multimodal acoustic and lexical modalties. In this work, we achieve a 70% unweighted accuracy in the IEMOCAP 4-class multimodal emotion recognition task. Further analysis shows the effect of integrating personality on the variation of our attention weights of each acoustic and lexical behavior modality for each speaker in the IEMOCAP database.
Jeng-Lin Li, Chi-Chun Lee
ACII2
2019 Annotation Matters: A Comprehensive Study on Recognizing Intended, Self-reported, and Observed Emotion Labels using Physiology
abstract
Studies have shown that the formation of emotion as self-awareness and cognitive appraisal process is complicated and can lead to idiosyncratic differences. Subject's self emotion evaluation process could be biased due to factors of environment, personal experience, and one's own cognitive ability, and the true affective state may be neglected (un-noticeable) due to an unconscious mental process. In this work, we present a comprehensive study to investigate the emotion recognition accuracy obtained using physiology with respect to different annotation schemes, i.e., intended, self-reported, and observed emotion labels. We found that when performing recognition across these three different labeling schemes using the same physiological parameters, the accuracy of the self-reported emotion labels results in about 10.3% and 3.1% drop when compared to two other annotation schemes. It indicates that self-assessed emotion labels may be noisier and induces a larger mismatch with respect to the affect-stimulated physiological responses. Further analysis shows that the electrodermal activity signal has the highest recognition rate with respect to the intended emotion of the stimuli. Finally, our error analysis reveals that there may exist a bias in the self-annotated label that is conditioned on the intended stimuli's valence polarity.
Hao-Chun Yang, Chi-Chun Lee
ACII2
2019 Adversarially-enriched Acoustic Code Vector Learned from Out-of-context Affective Corpus for Robust Emotion Recognition
abstract
Advancement in speech emotion recognition technology has brought tremendous potential in designing human-centered applications across a wide range of scenarios. However, due to the difficulty in obtaining large-scale labeled emotion corpus for every application domains, most of the existing databases are collected within disparate and limited contexts. This contextualization often undermines the variability in the emotional acoustic manifestation due to the limitation in the amount of labeled data that can be collected for each particular context. This, hence, creates a robustness issue across emotional scenarios. In this work, we propose to learn an enhanced acoustic code vector for in-context emotion database through adversarially learning from large out-of-context emotion corpus to obtain robust emotion recognition. We demonstrate that our framework can obtain improved recognition accuracy using low dimensional representations on two different databases, and it maintains its modeling power even when given very limited in-context training samples.
Chun-Min Chang, Chi-Chun Lee
ICASSP2
2019 Learning Semantic-preserving Space Using User Profile and Multimodal Media Content from Political Social Network
abstract
The use of social media in politics has dramatically changed the way campaigns are run and how elected officials interact with their constituents. An advanced algorithm is required to analyze and understand this large amount of heterogeneous social media data to investigate several key issues, such as stance and strategy, in political science. Most of previous works concentrate their studies using text-as-data approach, where the rich yet heterogeneous information in the user profile, social relationship, and multimodal media content is largely ignored. In this work, we propose a two-branch network that jointly maps the post contents and politician profile into the same latent space, which is trained using a large-margin objective that combines a cross-instance distance constraint with a within-instance semantic-preserving constraint. Our proposed political embedding space can be utilized not only in reliably identifying political spectrum and message type but also in providing a political representation space for interpretable ease-of-visualization.
Wei-Hao Chang, Jeng-Lin Li, Chi-Chun Lee
ICASSP3
2019 Every Rating Matters: Joint Learning of Subjective Labels and Individual Annotators for Speech Emotion Classification
abstract
Emotion perception is subjective and vary with respect to each individual due to the natural bias of human, such as gender, culture, and age. Conventionally, emotion recognition relies on the consensus, e.g., majority of annotations (hard label) or the distribution of annotations (soft label), and do not include rater-specific model. In this paper, we propose a joint learning methodology that simultaneously considers the label uncertainty and annotator idiosyncrasy using hard and soft emotion label annotation accompanying with individual and crowd annotator modeling. Our proposed model achieves unweighted average recall (UAR) 61.48% on the benchmark emotion corpus. Further analyses reveal that emotion perception is indeed rater-dependent, using the hard label and soft emotion distribution provides complementary affect modeling information, and finally joint learning of subjective emotion perception and individual rater model provides the best discriminative power.
Huang-Cheng Chou, Chi-Chun Lee
ICASSP2
2019 An Event-contrastive Connectome Network for Automatic Assessment of Individual Face Processing and Memory Ability
abstract
Human adapt their behaviors by continuously monitoring one another to function socially in our society. The ability to process face identity from memory is a crucial basic capability. In this work, we propose an event-contrastive connectome network (E-cCN) in representing brain's functional connectivity with contrastive loss to handle layers of fMRI data variabilities exists under different controlled stimuli events to achieve improved automatic assessing of an individual's face processing and memory ability. Our proposed connectome network achieves an overall recognition accuracy of 80.20% and 82.05% in binary classification of separating high versus low scoring subjects on tasks of Taiwanese Face Memory Test (TFMT) and component inverse efficiency score (cIE) respectively. Further, our network embedding representation demonstrate distinct connectivity patterns in key face processing brain regions (ROIs) when comparing between high and low face processing and memory ability.
Wan-Ting Hsieh, Hao-Chun Yang, Fu-Sheng Tsai, Chon-Wen Shyi, Chi-Chun Lee
ICASSP5
2019 An Attribute-invariant Variational Learning for Emotion Recognition Using Physiology
abstract
Studies have shown that people with different personalities would result in a different physiological reaction when encountering emotional stimulus. In this work, we propose an attribute-invariance loss embedded variational autoencoder (AI-VAE) to learn the personality-invariant physiological signal representation. The AI-VAE includes an additional loss aiming to perturb features from different personality polarity to obtain emotion discriminative representation. We evaluate our framework on a large emotion corpus of physiological data. Our method achieves a state of the art unweighted accuracy of 68.8% and 67.0% in a binary classification of arousal and valence, which improves over the baseline vanilla VAE by 5.5% and 6.5%. Further analysis reveals that several EEG features are statistically relevant between different personalities types across emotional states, and ECG features are also specifically correlated to personality dimension of "Creativeness", underscoring the importance of personality in modulating psychophysiological processes.
Hao-Chun Yang, Chi-Chun Lee
ICASSP2
2019 An Interaction-aware Attention Network for Speech Emotion Recognition in Spoken Dialogs
abstract
Obtaining robust speech emotion recognition (SER) in scenarios of spoken interactions is critical to the developments of next generation human-machine interface. Previous research has largely focused on performing SER by modeling each utterance of the dialog in isolation without considering the transactional and dependent nature of the human-human conversation. In this work, we propose an interaction-aware attention network (IAAN) that incorporate contextual information in the learned vocal representation through a novel attention mechanism. Our proposed method achieves 66.3% accuracy (7.9% over baseline methods) in four class emotion recognition and is also the current state-of-art recognition rates obtained on the benchmark database.
Sung-Lin Yeh, Yun-Shao Lin, Chi-Chun Lee
ICASSP3
2019 Learning Minimal Intra-Genre Multimodal Embedding from Trailer Content and Reactor Expressions for Box Office Prediction
abstract
Movie watching is one of the most popular leisure activities in our daily life. The box office revenue, especially in the first week, is critical for financial planning in the movie industry. Most existing movie box office prediction relies on meta data, viewer's comments, and trailer content. However, when viewers are immersed in a movie experience, they would naturally manifest expressions invoked by the media content. In this work, we propose a novel movie box office prediction framework by joint modeling meta attributes, trailer content, and viewer's natural expressions gathered from YouTube reactor videos. The proposed network learns a discriminabilityenhanced content and expression embeddings using a minimal intra-genre distance loss function. The proposed architecture achieves 79.07%, 73.79% and 76.82% for low/high movie box office tier classification (top 30%, top 10% and top 5%) on a large scale trailer-reactor database. Furthermore, we provide an analysis on the effectiveness of viewer's reaction and our intra-genre projection. Most existing movie box office prediction relies on meta data, viewer's comments, and trailer content. However, when viewers are immersed in a movie experience, they would naturally manifest expressions invoked by the media content. In this work, we propose a novel movie box office prediction framework by joint modeling meta attributes, trailer content, and viewer's natural expressions gathered from YouTube reactor videos. The proposed network learns a discriminability-enhanced content and expression embeddings using a minimal intra-genre distance loss function. The proposed architecture achieves 79.07%, 73.79% and 76.82% for low/high movie box office tier classification (top 30%, top 10% and top 5%) on a large scale trailer-reactor database. Furthermore, we provide an analysis on the effectiveness of viewer's reaction and our intra-genre projection.
Ming-Ya Ko, Jeng-Lin Li, Chi-Chun Lee
ICME3
2019 Enforcing Semantic Consistency for Cross Corpus Valence Regression from Speech Using Adversarial Discrepancy Learning
Gao-Yi Chao, Yun-Shao Lin, Chun-Min Chang, Chi-Chun Lee
INTERSPEECH4
2019 Investigating the Variability of Voice Quality and Pain Levels as a Function of Multiple Clinical Parameters
Hui-Ting Hong, Jeng-Lin Li, Yi-Ming Weng, Chip-Jin Ng, Chi-Chun Lee
INTERSPEECH5
2019 Acoustic Indicators of Deception in Mandarin Daily Conversations Recorded from an Interactive Game
Chih-Hsiang Huang, Huang-Cheng Chou, Yi-Tong Wu, Chi-Chun Lee, Yi-Wen Liu
INTERSPEECH4
2019 Attentive to Individual: A Multimodal Emotion Recognition Network with Personalized Attention Profile
Jeng-Lin Li, Chi-Chun Lee
INTERSPEECH2
2019 Using Attention Networks and Adversarial Augmentation for Styrian Dialect Continuous Sleepiness and Baby Sound Recognition
Sung-Lin Yeh, Gao-Yi Chao, Bo-Hao Su, Yu-Lin Huang, Meng-Han Lin, Yin-Chun Tsai, Yu-Wen Tai, Zheng-Chi Lu, Chieh-Yu Chen, Tsung-Ming Tai, Chiu-Wang Tseng, Cheng-Kuang Lee, Chi-Chun Lee
INTERSPEECH13
2019 Predicting Group Performances Using a Personality Composite-Network Architecture During Collaborative Task
Shun-Chang Zhong, Yun-Shao Lin, Chun-Min Chang, Yi-Ching Liu, Chi-Chun Lee
INTERSPEECH5
2019 Toward differential diagnosis of autism spectrum disorder using multimodal behavior descriptors and executive functions
Chin-Po Chen, Susan Shur-Fen Gau, Chi-Chun Lee
Comput. Speech Lang.3
2019 Toward Automating Oral Presentation Scoring During Principal Certification Program Using Audio-Video Low-Level Behavior Profiles
abstract
Effective leadership bears strong relationship to attributes of emotion contagion, positive mood, and social intelligence. In fact, leadership quality has been shown to be manifested in the exhibited communicative behaviors, especially in settings of public speaking. While studies on the theories of leadership has received much attention, little has progressed in terms of the computational development in its measurements. In this work, we present a behavioral signal processing (BSP) research to assess the qualities of oral presentations in the domain of education, in specific, we propose a multimodal framework toward automating the scoring process of pre-service school principals' oral presentations given at the yearly certification program. We utilize a dense unit-level audio-video feature extraction approach with session-level behavior profile representation techniques based on bag-of-word and Fisher-vector encoding. Furthermore, we design a scoring framework, inspired by the psychological evidences of human's decision-making mechanism, to use confidence measures outputted from support vector machine classifier trained on the distinctive set of data samples as the regressed scores. Our proposed approach achieves an absolute improvement of 0.049 (9.8 percent relative) on average over support vector regression. We further demonstrate that the framework is reliable and consistent compared to human experts.
Shan-Wen Hsiao, Hung-Ching Sun, Ming-Chuan Hsieh, Ming-Hsueh Tsai, Yu Tsao 0001, Chi-Chun Lee
IEEE Trans. Affect. Comput.6
2018 Integrating Perceivers Neural-Perceptual Responses Using a Deep Voting Fusion Network for Automatic Vocal Emotion Decoding
abstract
Understanding neuro-perceptual mechanism of vocal emotion perception continues to be an important research direction not only in advancing scientific knowledge but also in inspiring more robust affective computing technologies. The large variabilities in the manifested fMRI signals among subjects has been shown to be due to the effect of individual difference, i.e., inter-subject variability. However, relatively few works have developed modeling techniques in task of automatic neuro-perceptual decoding to handle such idiosyncrasies. In our work, we propose a novel computation method of deep voting fusion neural network architecture by learning an adjusted weight matrix applied at the fusion layer. The framework achieves an unweighted average recall of 53.10% in a four-class vocal emotion states decoding task, i.e., a relative improvement of 8.9% over a two-stage SVM decision-level fusion. Our framework demonstrates its effectiveness in handling individual differences. Further analysis is conducted to study the properties of the learned adjusted weight matrix as a function of emotion classification accuracy.
Wan-Ting Hsieh, Hao-Chun Yang, Ya-Tse Wu, Fu-Sheng Tsai, Li-Wei Kuo, Chi-Chun Lee
ICASSP6
2018 Learning Lexical Coherence Representation Using LSTM Forget Gate for Children with Autism Spectrum Disorder During Story-Telling
abstract
Inability to carry out cohesive narratives has been identified in children with autism spectrum disorder (ASD). However, deriving cohesion measures is often done using manual labeling or relying on expert-crafted features. In this work, we develop a novel LSTM framework to learn the embedded narrative cohesion representation from data directly. Our lexical coherence representation achieves a promising recognition accuracy of 92% in classifying between typically-developing (TD) and ASD children, as compared to 73% by using conventional coherence measures computed from syntactic, word usage, and latent semantic analysis. We perform additional validity analyses on our proposed representation. By experimentally introducing incoherence in the TD's story-telling narratives through word and sentence-level shuffling, the derived lexical coherence representation from these incoherent TD data samples result in a representation closer to those of ASD data samples.
Yu-Shuo Liu, Chin-Po Chen, Susan Shur-Fen Gau, Chi-Chun Lee
ICASSP4
2018 A Triplet-Loss Embedded Deep Regressor Network for Estimating Blood Pressure Changes Using Prosodic Features
abstract
Studies have shown that measures of personal physiology, e.g., blood pressure (BP) variation and heart rate variability (HRV), is closely related to a subject's psychological states and are being used regularly to track patients' health conditions in medical settings. The conventional method of monitoring physiology requires wearing specialized sensors or utilizing medical instruments, which hinders the ability of scalable and just-in-time monitoring of patients. In this study, we propose a triplet-loss embedded deep regressor network to predict changes of BP using expressive prosodic features for on-boarding emergency room patients between pre- and post-triage sessions. The framework achieves correlations of 0.419 and 0.386 in predicting changes in SBP (systolic blood pressure) and DBP (diastolic blood pressure) respectively, which is 26.1% and 17.3% relative improvement compared to DNN-regressors without triplet-loss embedding. Further correlation analyses on the relationship between prosodic features and BP changes are presented.
Hao-Chun Yang, Fu-Sheng Tsai, Yi-Ming Weng, Chip-Jin Ng, Chi-Chun Lee
ICASSP5
2018 A Genre-Affect Relationship Network with Task-Specific Uncertainty Weighting foR Recognizing Induced Emotion in Music
abstract
Emotion is a core fundamental attribute of humans. Using music to induce emotional responses from subjects to better facilitate human behavior shaping have been effective across domains of health, education, and retail. Computationally model the musically-induced emotion provides necessary content-based analytics for large-scale and wide-applicability of such human-centered applications. In this work, we propose a relationship neural network architecture to learn to regress the induced emotion attributes with an auxiliary task of genre classification. Our proposed Genre-Affect Relationship Network with homoscedastic uncertainty weighting embeds the relationship between affect and genre as tensor normal prior within task-specific layers; the architecture is optimized further by incorporating task-specific uncertainty. The proposed architecture achieves a state-of-art 0.564 average Pearson correlation computed over nine induced emotion ratings in the Emotify database. Furthermore, we provide an analysis to understand the relationship between the induced emotions of these musical pieces and their associated genres.
Wei-Hao Chang, Jeng-Lin Li, Yun-Shao Lin, Chi-Chun Lee
ICME4
2018 Generating fMRI-Enriched Acoustic Vectors using a Cross-Modality Adversarial Network for Emotion Recognition
abstract
Automatic emotion recognition has long been developed by concentrating on modeling human expressive behavior. At the same time, neuro-scientific evidences have shown that the varied neuro-responses (i.e., blood oxygen level-dependent (BOLD) signals measured from the functional magnetic resonance imaging (fMRI)) is also a function on the types of emotion perceived. While past research has indicated that fusing acoustic features and fMRI improves the overall speech emotion recognition performance, obtaining fMRI data is not feasible in real world applications. In this work, we propose a cross modality adversarial network that jointly models the bi-directional generative relationship between acoustic features of speech samples and fMRI signals of human percetual responses by leveraging a parallel dataset. We encode the acoustic descriptors of a speech sample using the learned cross modality adversarial network to generate the fMRI-enriched acoustic vectors to be used in the emotion classifier. The generated fMRI-enriched acoustic vector is evaluated not only in the parallel dataset but also in an additional dataset without fMRI scanning. Our proposed framework significantly outperform using acoustic features only in a four-class emotion recognition task for both datasets, and the use of cyclic loss in learning the bi-directional mapping is also demonstrated to be crucial in achieving improved recognition rates.
Gao-Yi Chao, Chun-Min Chang, Jeng-Lin Li, Ya-Tse Wu, Chi-Chun Lee
ICMI5
2018 Using Interlocutor-Modulated Attention BLSTM to Predict Personality Traits in Small Group Interaction
abstract
Small group interaction occurs often in workplace and education settings. Its dynamic progression is an essential factor in dictating the final group performance outcomes. The personality of each individual within the group is reflected in his/her interpersonal behaviors with other members of the group as they engage in these task-oriented interactions. In this work, we propose an interlocutor-modulated attention BSLTM (IM-aBLSTM) architecture that models an individual's vocal behaviors during small group interactions in order to automatically infer his/her personality traits. The interlocutor-modulated attention mechanism jointly optimize the relevant interpersonal vocal behaviors of other members of group during interactions. In specifics, we evaluate our proposed IM-aBLSTM in one of the largest small group interaction database, the ELEA corpus. Our framework achieves a promising unweighted recall accuracy of 87.9% in ten different binary personality trait prediction tasks, which outperforms the best results previously reported on the same database by 10.4% absolute. Finally, by analyzing the interpersonal vocal behaviors in the region of high attention weights, we observe several distinct intra- and inter-personal vocal behavior patterns that vary as a function of personality traits.
Yun-Shao Lin, Chi-Chun Lee
ICMI2
2018 Encoding Individual Acoustic Features Using Dyad-Augmented Deep Variational Representations for Dialog-level Emotion Recognition
Jeng-Lin Li, Chi-Chun Lee
INTERSPEECH2
2018 Learning Conditional Acoustic Latent Representation with Gender and Age Attributes for Automatic Pain Level Recognition
Jeng-Lin Li, Yi-Ming Weng, Chip-Jin Ng, Chi-Chun Lee
INTERSPEECH4
2018 An Interlocutor-Modulated Attentional LSTM for Differentiating between Subgroups of Autism Spectrum Disorder
Yun-Shao Lin, Susan Shur-Fen Gau, Chi-Chun Lee
INTERSPEECH3
2018 Self-Assessed Affect Recognition Using Fusion of Attentional BLSTM and Static Acoustic Features
Bo-Hao Su, Sung-Lin Yeh, Ming-Ya Ko, Huan-Yu Chen, Shun-Chang Zhong, Jeng-Lin Li, Chi-Chun Lee
INTERSPEECH7
2018 Automatic Assessment of Individual Culture Attribute of Power Distance Using a Social Context-Enhanced Prosodic Network Representation
Fu-Sheng Tsai, Hao-Chun Yang, Wei-Wen Chang, Chi-Chun Lee
INTERSPEECH4
2017 A bootstrapped multi-view weighted Kernel fusion framework for cross-corpus integration of multimodal emotion recognition
abstract
Recently the development of robust emotion recognition has been increasingly emphasized in order to handle situations of different cultures and languages. This has become critical due to the potential applicability of emotion recognizers across a wide range of application scenarios. Instead of conventional approach in deriving a single universal emotion recognition module across all languages, we have previously demonstrated a method based on integrating other database's useful information to improve the emotion recognition of the current data with fusion of multiple emotion perspectives. In this paper, we present an improved framework, i.e., a bootstrapped multi-view weighted kernel fusion, to further advance the recognition accuracies. We have also extended the modeling of speech-only modality to include video information. In specifics, we utilize two emotional corpora of different languages. Our proposed framework obtains improved recognition in regressing activation and valence attributes using audio and video modalities across both of the databases. We not only demonstrate that the weighted kernel fusion can provide additional modeling power but also present analyses on the complementary emotionally-relevant acoustic and visual behaviors computed from the multiple emotion perspectives.
Chun-Min Chang, Bo-Hao Su, Shih-Chen Lin, Jeng-Lin Li, Chi-Chun Lee
ACII5
2017 NNIME: The NTHU-NTUA Chinese interactive multimodal emotion corpus
abstract
The increasing availability of large-scale emotion corpus along with the advancement in emotion recognition algorithms have enabled the emergence of next-generation human-machine interfaces. This paper describes a newly-collected multimodal corpus, i.e., the NTHU-NTUA Chinese Interactive Emotion Corpus (NNIME). The database is a result of the collaborative work between engineers and drama experts. This database includes recordings of 44 subjects engaged in spontaneous dyadic spoken interactions. The multimodal data includes approximately 11-hour worth of audio, video, and electrocardiogram data recorded continuously and synchronously. The database is also completed with a rich set of emotion annotations on discrete and continuous-in-time annotation by a total of 49 annotators. Thees emotion annotations include a diverse perspectives: peer-report, director-report, self-report, and observer-report. This carefully-engineered data collection and annotation processes provide an additional valuable resource to quantify and investigate various aspects of affective phenomenon and human communication. To our best knowledge, the NNIME is one of the few large-scale Chinese affective dyadic interaction database that have been systematically collected, organized, and to be publicly-released to the research community.
Huang-Cheng Chou, Lien-Chiang Chang, Chyi-Chang Li, Hsi-Pin Ma, Chi-Chun Lee
ACII6
2017 Embedding stacked bottleneck vocal features in a LSTM architecture for automatic pain level classification during emergency triage
abstract
In order to effectively allocate healthcare resource, a proper triage classification system plays an important role in assessing the severity of on-boarding patients at the emergency department. One of the major items in the current triage system is to assess the level of pain intensity, which relies solely on patients self-report numerical-rating scale (NRS) at the moment. The nature of self-report on pain level poses a challenge in maintaining the validity and consistency of the triage classification outcome. While there has been algorithms developed to automatically detect pain from expressive behaviors, most of them concentrate only on facial or body gestural expressions within the context of physical exercises. In this work, we propose to utilize stacked bottleneck acoustic representations in a long-short term memory neural networks (LSTMs) architecture as features for pain severity classification in a database consists of patients during real triage sessions. Our proposed framework achieves accuracy of 72.3% and 54.2% in binary and three-class pain intensity classification tasks. Our results further demonstrate that the severity of pain can largely be captured in the patients prosodic characteristics.
Fu-Sheng Tsai, Yi-Ming Weng, Chip-Jin Ng, Chi-Chun Lee
ACII4
2017 Fusion of multiple emotion perspectives: Improving affect recognition through integrating cross-lingual emotion information
abstract
Developing cross-corpus, cross-domain, and cross-language emotion recognition algorithm has becoming more prevalent recently to ensure the wide applicability of robust emotion recognizer. In this work, we propose a computational framework on fusing multiple emotion perspectives by integrating cross-lingual emotion information. By assuming that each data is `perceived' not only by a main perspective but additional derived perspectives (from a corpus of a different language), we can then combine each of the perspective-dependent features via kernel fusion technique. In specifics, we utilize two emotional corpora of different languages (Chinese and English). Our experiments demonstrate that our proposed framework achieves significant improvement over single perspective baseline across both databases.
Chun-Min Chang, Chi-Chun Lee
ICASSP2
2017 Computing Multimodal Dyadic Behaviors During Spontaneous Diagnosis Interviews Toward Automatic Categorization of Autism Spectrum Disorder
Chin-Po Chen, Xian-Hong Tseng, Susan Shur-Fen Gau, Chi-Chun Lee
INTERSPEECH4
2017 Deriving Dyad-Level Interaction Representation Using Interlocutors Structural and Expressive Multimodal Behavior Features
Yun-Shao Lin, Chi-Chun Lee
INTERSPEECH2
2017 Modeling Perceivers Neural-Responses Using Lobe-Dependent Convolutional Neural Network to Improve Speech Emotion Recognition
Ya-Tse Wu, Yu-Hsien Liao, Li-Wei Kuo, Chi-Chun Lee
INTERSPEECH5
2016 A Gaussian mixture regression approach toward modeling the affective dynamics between acoustically-derived vocal arousal score (VC-AS) and internal brain fMRI bold signal response
abstract
Understanding the underlying neuro-perceptual mechanism of humans' ability to decode emotional content in vocal signal is an important research direction. In this paper, we describe our initial research effort into quantitatively modeling the joint dynamics between measures of vocal arousal and blood oxygen level-dependent (BOLD) signals. We utilize Gaussian mixture regression approach to predict the invoked BOLD signal response as the subject is exposed to various levels of continuous vocal arousal stimuli. The proposed framework is built upon measures of vocal arousal from acoustically-derived features, and we obtain a reasonable predictive correlation to the true BOLD signal for the seven emotionally-related brain regions. Further experiment also demonstrates that there exists a more explanatory power of using signal-derived arousal measure to the internal BOLD signal responses compared to using human annotated arousal in the construction of Gaussian mixture regression modeling.
Yu-Hsien Liao, Heng-Tai Jan, Li-Wei Kuo, Chi-Chun Lee
ICASSP5
2016 A thin-slice perception of emotion? An information theoretic-based framework to identify locally emotion-rich behavior segments for global affect recognition
abstract
Human's judgment has been shown to be thin-sliced in nature, i.e., accurate perception can often be achieved for a short duration of exposure to expressive behaviors. In this work, we develop a mutual information-based framework to select the most emotion-rich 20% of local multimodal behavior segments within a 3-minute long affective dyadic interaction in the USC CreativeIT database. We obtain a prediction accuracy of 0.597, 0.728, and 0.772 (measured by Spearman correlation) for an actor's global (session-level) emotion attributes (activation, dominance, and valence) using Fisher-vector encoding and support vector regression built on these 20% of multimodal emotion-rich behavior segments. Our framework achieves a better accuracy over using the interaction in its entirety and a variety of other data selection baseline methods by a significant margin. Furthermore, our analysis indicates that the highest prediction accuracy can be obtained using only 20%-30% of data within each session, i.e., additional evidences for the thin-slice nature of affect perception.
Chi-Chun Lee
ICASSP2
2016 Enhancement of Automatic Oral Presentation Assessment System Using Latent N-Grams Word Representation and Part-of-Speech Information
Wen-Yu Huang, Shan-Wen Hsiao, Hung-Ching Sun, Ming-Chuan Hsieh, Ming-Hsueh Tsai, Chi-Chun Lee
INTERSPEECH6
2016 Minimization of Regression and Ranking Losses with Shallow Neural Networks on Automatic Sincerity Evaluation
Hung-Shin Lee, Yu Tsao 0001, Chi-Chun Lee, Hsin-Min Wang, Wei-Chen Chen, Shan-Wen Hsiao, Shyh-Kang Jeng
INTERSPEECH3
2016 Toward Development and Evaluation of Pain Level-Rating Scale for Emergency Triage based on Vocal Characteristics and Facial Expressions
Fu-Sheng Tsai, Ya-Ling Hsu, Wei-Chen Chen, Yi-Ming Weng, Chip-Jin Ng, Chi-Chun Lee
INTERSPEECH6
2015 Multimodal arousal rating using unsupervised fusion technique
abstract
Arousal is essential in understanding human behavior and decision-making. In this work, we present a multimodal arousal rating framework that incorporates minimal set of vocal and non-verbal behavior descriptors. The rating framework and fusion techniques are unsupervised in nature to ensure that it can be readily-applicable and interpretable. Our proposed multimodal framework improves correlation to human judgment from 0.66 (vocal-only) to 0.68 (multimodal); analysis shows that the supervised fusion framework does not improve correlation. Lastly, an interesting empirical evidence demonstrates that the signal-based quantification of arousal achieves a higher agreement with each individual rater than the agreement among raters themselves. This further strengthens that machine-based rating is a viable way of measuring subjective humans' internal states through observing behavior features objectively.
Wei-Chen Chen, Po-Tsun Lai, Yu Tsao 0001, Chi-Chun Lee
ICASSP4
2015 A multimodal approach for automatic assessment of school principals' oral presentation during pre-service training program
Shan-Wen Hsiao, Hung-Ching Sun, Ming-Chuan Hsieh, Ming-Hsueh Tsai, Hsin-Chih Lin, Chi-Chun Lee
INTERSPEECH6
2015 An analysis of the relationship between signal-derived vocal arousal score and human emotion production and perception
abstract
Bone et al. recently proposed an unsupervised signal-derived vocal arousal score (VC-AS) based on fusion of three intuitive acoustic features, i.e., pitch, intensity, and HF500, and have shown the effectiveness of quantifying humans perceptual rat-ings of arousal robustly across multiple corpora. Due to the readily-applicable nature of the system, this objective quantifi-cation scheme could foresee-ably be used in multiple fields of behavioral science as an objective measure of affect. In this work, we investigate in detail the relationship of this signal-derived measure to both intended arousal expression (i.e., pro-duction aspect) and perceived arousal rating (i.e., perception as-pect). On the perception side, our results in three databases (EMA, VAM, and IEMOCAP) indicate that VC-AS agrees with mean perception at least as well as an average individ-ual rater does. Regarding production, we demonstrate that in-tended arousal correlates more with VC-AS than mean percep-tion (EMA and IEMOCAP). We also show that VC-AS corre-lates more with intended arousal than perceived arousal (EMA); this finding is quite surprising given that the framework is sup-ported by extensive affective perception studies, although there is physiological motivation as well. Implication for utilizing VC-AS for novel scientific study in engineering (e.g., to miti-gate subjectivity) is further discussed. Index Terms: vocal arousal rating, affective perception, affec-tive production
Chi-Chun Lee, Daniel Bone, Shri Narayanan
INTERSPEECH1
2014 An investigation of vocal arousal dynamics in child-psychologist interactions using synchrony measures and a conversation-based model
abstract
Researchers from various disciplines are concerned with the study of affective phenomena, especially arousal. Expressed affective modulations, which reflect both an individual’s in-ternal state and external factors, are central to the commu-nicative process. Bone et al. developed a robust, unsuper-vised (rule-based) method which provides a scale-continuous, bounded arousal rating from the vocal signal. In this study, we investigate the joint-dynamics of child and psychologist vocal arousal in autism spectrum disorder (ASD) diagnostic interac-tions. Arousal synchrony is assessed with multiple methods. Results indicate that children with higher ASD severity tend to lead the arousal dynamics more, seemingly because the children aren’t as responsive to the psychologist’s affective modulations. A vocal arousal model is also proposed which incorporates so-cial and conversational constructs. The model captures conver-sational signal relations, and is able to distinguish between high and low ASD severity at accuracies well-above chance.
Daniel Bone, Chi-Chun Lee, Alexandros Potamianos, Shri Narayanan
INTERSPEECH2
2014 Ensemble of machine learning algorithms for cognitive and physical speaker load detection
abstract
We present our methods and results on participating in the Interspeech 2014 Computational Paralinguistics ChallengE (ComParE) of which the goal is to detect certain type of load of a speaker using acoustic features. There are in total seven classification models contributing to our final prediction, namely, neural network with rectified linear unit and dropout (ReLUNet), conditional restricted Boltzmann machine (CRBM), logistic regression (LR), support vector machine (SVM), Gaussian discriminant analysis (GDA), k-nearest neighbors (KNN), and random forest (RF). When linearly blending the predictions of these models, we are able to get significant improvements over the challenge baseline. Index Terms: Physical Load Detection, Cognitive Load Detection, Neural Network, Classification Models
How Jing, Ting-Yao Hu, Hung-Shin Lee, Wei-Chen Chen, Chi-Chun Lee, Yu Tsao 0001, Hsin-Min Wang
INTERSPEECH5
2014 Computing vocal entrainment: A signal-derived PCA-based quantification scheme with application to affect analysis in married couple interactions
Chi-Chun Lee, Athanasios Katsamanis, Matthew Black, Brian R. Baucom, Andrew Christensen, Panayiotis G. Georgiou, Shri Narayanan
Comput. Speech Lang.1
2014 Robust Unsupervised Arousal Rating: A Rule-Based Framework withKnowledge-Inspired Vocal Features
abstract
Studies in classifying affect from vocal cues have produced exceptional within-corpus results, especially for arousal (activation or stress); yet cross-corpora affect recognition has only recently garnered attention. An essential requirement of many behavioral studies is affect scoring that generalizes across different social contexts and data conditions. We present a robust, unsupervised (rule-based) method for providing a scale-continuous, bounded arousal rating operating on the vocal signal. The method incorporates just three knowledge-inspired features chosen based on empirical and theoretical evidence. It constructs a speaker's baseline model for each feature separately, and then computes single-feature arousal scores. Lastly, it advantageously fuses the single-feature arousal scores into a final rating without knowledge of the true affect. The baseline data is preferably labeled as neutral, but some initial evidence is provided to suggest that no labeled data is required in certain cases. The proposed method is compared to a state-of-the-art supervised technique which employs a high-dimensional feature set. The proposed framework achieves highly-competitive performance with additional benefits. The measure is interpretable, scale-continuous as opposed to discrete, and can operate without any affective labeling. An accompanying Matlab tool is made available with the paper.
Daniel Bone, Chi-Chun Lee, Shri Narayanan
IEEE Trans. Affect. Comput.2
2013 Using physiology and language cues for modeling verbal response latencies of children with ASD
abstract
Signal-derived measures can provide effective ways towards quantifying human behavior. Verbal Response Latencies (VRLs) of children with Autism Spectrum Disorders (ASD) during conversational interactions are able to convey valuable information about their cognitive and social skills. Motivated by the inherent gap between the external behavior and inner affective state of children with ASD, we study their VRLs in relation to their explicit but also implicit behavioral cues. Explicit cues include the children's language use, while implicit cues are based on physiological signals. Using these cues, we perform classification and regression tasks to predict the duration type (short/long) and value of VRLs of children with ASD while they interacted with an Embodied Conversational Agent (ECA) and their parents. Since parents are active participants in these triadic interactions, we also take into account their linguistic and physiological behaviors. Our results suggest an association between VRLs and these externalized and internalized signal information streams, providing complementary views of the same problem.
Theodora Chaspari, Daniel Bone, James Gibson, Chi-Chun Lee, Shri Narayanan
ICASSP4
2013 Head motion synchrony and its correlation to affectivity in dyadic interactions
abstract
Behavioral synchrony, or entrainment, is a phenomenon of great interest to psychologists and a challenging construct to quantify. In this work we study the synchrony behavior of head motion in human dyadic interactions. We model head motion using Gaussian Mixture Model (GMM) of line spectral frequencies extracted from the motion vectors of the head. We quantify interlocutor head motion similarity through the Kullback-Leibler divergence of the GMM posteriors of their respective motion sequences. We use an audiovisual database of distressed couple interactions, extensively annotated by psychologists, to test two hypotheses using the derived similarity measure. We validate the first hypothesis — that people are more likely to increase their degree of synchrony as the interaction progresses — by comparing the first and second halves of the interaction. The second hypothesis tests if the relative change of the similarity measure from these two halves is significantly correlated with the behavioral annotation by the domain experts. This work underscores the importance of head motion as an interaction cue, and the feasibility of using it in a computational model for synchrony behavior.
Bo Xiao 0003, Panayiotis G. Georgiou, Chi-Chun Lee, Brian R. Baucom, Shri Narayanan
ICME3
2013 Acoustic-prosodic, turn-taking, and language cues in child-psychologist interactions for varying social demand
abstract
Impaired social communication and social reciprocity are the primary phenotypic distinctions between autism spectrum dis-orders (ASD) and other developmental disorders. We investi-gate quantitative conversational cues in child-psychologist in-teractions using acoustic-prosodic, turn-taking, and language features. Results indicate the conversational quality degraded for children with higher ASD severity, as the child exhibited difficulties conversing and the psychologist varied her speech and language strategies to engage the child. When interacting with children with increasing ASD severity, the psychologist exhibited higher prosodic variability, increased pausing, more speech, atypical voice quality, and less use of conventional con-versational cue such as assents and non-fluencies. Children with increasing ASD severity spoke less, spoke slower, responded later, had more variable prosody, and used personal pronouns, affect language, and fillers less often. We also investigated the predictive power of features from interaction subtasks with varying social demands placed on the child. We found that acoustic prosodic and turn-taking features were more predictive during higher social demand tasks, and that the most predictive features vary with context of interaction. We also observed that psychologist language features may be robust to the amount of speech in a subtask, showing significance even when the child is participating in minimal-speech, low social-demand tasks. Index Terms: autism spectrum disorders, atypical prosody, so-cial reciprocity, turn-taking, language cues
Daniel Bone, Chi-Chun Lee, Theodora Chaspari, Matthew Black, Marian E. Williams, Sungbok Lee, Pat Levitt, Shri Narayanan
INTERSPEECH2
2013 Analyzing eye-voice coordination in rapid automatized naming
abstract
Rapid Automatized Naming (RAN) is a powerful tool for pre-dicting future reading skill. A person’s ability to quickly name symbols as they scan a table is related to higher-level reading proficiency in adults and is predictive of future literacy gains in children. However, noticeable differences are present in the strategies or patterns within groups having similar task comple-tion times. Thus, a further stratification of RAN dynamics may lead to better characterization and later intervention to support reading skill acquisition. In this work, we analyze the dynamics of the eyes, voice, and the coordination between the two during performance. It is shown that fast performers are more similar to each other than to slow performers in their patterns, but not vice versa. Further insights are provided about the patterns of more proficient subjects. For instance, fast performers tended to exhibit smoother behavior contours, suggesting a more sta-ble perception-production process.
Daniel Bone, Chi-Chun Lee, Vikram Ramanarayanan, Shri Narayanan, Renske S. Hoedemaker, Peter C. Gordon
INTERSPEECH2
2013 Toward automating a human behavioral coding system for married couples' interactions using speech acoustic features
Matthew Black, Athanasios Katsamanis, Brian R. Baucom, Chi-Chun Lee, Adam C. Lammert, Andrew Christensen, Panayiotis G. Georgiou, Shri Narayanan
Speech Commun.4
2012 Classification of emotional content of sighs in dyadic human interactions
abstract
Emotions are an important part of human communication and are expressed both verbally and non-verbally. Common nonverbal vocalizations such as laughter, cries and sighs carry important emotional content in conversations. Sighs often are associated with negative emotion. In this work, we show that emotional sighs exist along both ends of the valence axis (positive-emotion vs. negative-emotion sighs) in spontaneous affective dialogs and that they have certain distinct multimodal characteristics. Classification results show that it is possible to differentiate between the two types of emotionally valenced sighs, using a combination of acoustic and gestural features with an overall unweighted accuracy of 58.26%.
Rahul Gupta 0001, Chi-Chun Lee, Shri Narayanan
ICASSP2
2012 Spontaneous-Speech Acoustic-Prosodic Features of Children with Autism and the Interacting Psychologist
abstract
Atypical prosody, often reported in children with Autism Spectrum Disorders, is described by a range of qualitative terms that reflect the eccentricities and variability among persons in the spectrum. We investigate various word-and phonetic-level features from spontaneous speech that may quantify the cues reflecting prosody. Furthermore, we introduce the importance of jointly modeling the psy-chologist’s vocal behavior in this dyadic interaction. We demonstrate that acoustic-prosodic features of both par-ticipants correlate with the children’s rated autism sever-ity. For increasing perceived atypicality, we find chil-dren’s prosodic features that suggest ‘monotonic ’ speech, variable volume, atypical voice quality, and slower rate of speech. Additionally, we find the psychologist’s features inform their perception of a child’s atypical behavior– e.g., the psychologist’s pitch slope and jitter are increas-ingly variable and their speech rate generally decreases. Index Terms: atypical prosody, autism spectrum disor-der, intonation, psychologist, voice quality, ADOS
Daniel Bone, Matthew Black, Chi-Chun Lee, Marian E. Williams, Pat Levitt, Sungbok Lee, Shri Narayanan
INTERSPEECH3
2012 A Robust Unsupervised Arousal Rating Framework using Prosody with Cross-Corpora Evaluation
abstract
This paper presents an unsupervised method for produc-ing a bounded rating of affective arousal from speech. One of the major challenges in such behavioral signal classification is the design of methods that generalize well across domains and datasets. We propose a frame-work that provides robustness across databases by: se-lecting coherent features based on empirical and theoret-ical evidence, fusing activation confidences from mul-tiple features, and effectively weighting the soft-labels without knowing the true labels. Spearman’s rank-correlation (and binary classification accuracy) on four
Daniel Bone, Chi-Chun Lee, Shri Narayanan
INTERSPEECH2
2012 Interplay between verbal response latency and physiology of children with autism during ECA interactions
Theodora Chaspari, Chi-Chun Lee, Shri Narayanan
INTERSPEECH2
2012 Based on Isolated Saliency or Causal Integration? Toward a Better Understanding of Human Annotation Process using Multiple Instance Learning and Sequential Probability Ratio Test
abstract
Human perception is capable of integrating local events to generate an overall impression at the global level; this is evident in daily life and is utilized repeatedly in behavioral science studies to bring objective measures into studies of human behavior. In this work, we explore two hypotheses considering whether it is the isolated-saliency or the causal-integration of information that can trigger the global perceptual behavioral ratings as trained annotators engage in tasks of observational coding. We carry out analyses using Multiple Instance Learning and Sequential Probability Ratio Test in a corpus of real and spontaneous distressed couples’ interaction with global sessionlevel abstract behavioral coding done by trained human annotators. We present various analyses based on different behavioral detection schemes demonstrating the potential of utilizing these algorithms in bringing insights into the human annotation process. We further show that while annotating behaviors with more positive impression, annotators gather information throughout the session compared to behaviors with more negative impression, where a single salient instance is enough to trigger the final global decision.
Chi-Chun Lee, Athanasios Katsamanis, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH1
2011 Affective State Recognition in Married Couples' Interactions Using PCA-Based Vocal Entrainment Measures with Multiple Instance Learning
Chi-Chun Lee, Athanasios Katsamanis, Matthew Black, Brian R. Baucom, Panayiotis G. Georgiou, Shri Narayanan
ACII (2)1
2011 An Analysis of PCA-Based Vocal Entrainment Measures in Married Couples' Affective Spoken Interactions
abstract
Entrainment has played a crucial role in analyzing marital couples interactions. In this work, we introduce a novel technique for quantifying vocal entrainment based on Principal Component Analysis (PCA). The entrainment measure, as we define in this work, is the amount of preserved variability of one interlocutor’s speaking characteristic when projected onto representing space of the other’s speaking characteristics. Our analysis on real couples interactions shows that when a spouse is rated as having positive emotion, he/she has a higher value of vocal entrainment compared when rated as having negative emotion. We further performed various statistical analyses on the strength and the directionality of vocal entrainment under different affective interaction conditions to bring quantitative insights into the entrainment phenomenon. These analyses along with a baseline prediction model demonstrate the validity and utility of the proposed PCA-based vocal entrainment measure. Index Terms: vocal entrainment, couples therapy, behavioral signal processing, principal component analysis
Chi-Chun Lee, Athanasios Katsamanis, Matthew Black, Brian R. Baucom, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH1
2011 Analyzing the Nature of ECA Interactions in Children with Autism
abstract
Embodied conversational agents (ECA) offer platforms for the collection of structured interaction and communication data. This paper discusses the data collected from the Rachel system, an ECA developed at the University of Southern California, for interactions with children with autism. Two dyads each com-posed of a child with autism and his parent participated in an experiment with two modes: interactions with and without the ECA present. The goal of this work is to assess the naturalness of the data recorded in the ECA interaction. This analysis was carried out using a classification framework with a prediction variable of the presence or absence of the ECA in the inter-action. The results demonstrate that it is possible to estimate whether or not a parent is interacting with the ECA using their speech data. However, it is not generally possible to do so for the child suggesting that the Rachel system is eliciting commu-nication data that is similar to that elicited through interactions between the child and his parent. Index Terms: Embodied conversational agent, multimodal in-terface, audio-video recording, autism, children’s speech
Emily Mower Provost, Chi-Chun Lee, James Gibson, Theodora Chaspari, Marian E. Williams, Shri Narayanan
INTERSPEECH2
2011 Emotion recognition using a hierarchical binary decision tree approach
Chi-Chun Lee, Emily Mower Provost, Carlos Busso, Sungbok Lee, Shri Narayanan
Speech Commun.1
2010 Predicting interruptions in dyadic spoken interactions
abstract
Interruptions occur frequently in spontaneous conversations, and they are often associated with changes in the flow of conversation. Predicting interruption is essential in the design of natural human-machine spoken dialog interface. The modeling can bring insights into the dynamics of human-human conversation. This work utilizes Hidden Condition Random Field (HCRF) to predict occurrences of interruption in dyadic spoken interactions by modeling both speakers' behaviors before a turn change takes place. Our prediction model, using both the foreground speaker's acoustic cues and the listener's gestural cues, achieves an F-measure of 0.54, accuracy of 70.68%, and unweighted accuracy of 66.05% on a multimodal database of dyadic interactions. The experimental results also show that listener's behaviors provides an indication of his/her intention of interruption.
Chi-Chun Lee, Shri Narayanan
ICASSP1
2010 Automatic classification of married couples' behavior using audio features
abstract
In this work, we analyzed a 96-hour corpus of married couples spontaneously interacting about a problem in their relationship. Each spouse was manually coded with relevant session-level perceptual observations (e.g., level of blame toward other spouse, global positive affect), and our goal was to classify the spouses’ behavior using features derived from the audio signal. Based on automatic segmentation, we extracted prosodic/spectral features to capture global acoustic properties for each spouse. We then trained gender-specific classifiers to predict the behavior of each spouse for six codes. We compare performance for the various factors (across codes, gender, classifier type, and feature type) and discuss future work for this novel and challenging corpus. Index Terms: behavioral signal processing, human behavior analysis, couples therapy, prosody, emotion recognition
Matthew Black, Athanasios Katsamanis, Chi-Chun Lee, Adam C. Lammert, Brian R. Baucom, Andrew Christensen, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH3
2010 Quantification of prosodic entrainment in affective spontaneous spoken interactions of married couples
abstract
Interaction synchrony among interlocutors happens naturally as people adapt their speaking style gradually to promote efficient communication. In this work, we quantify one aspect of interaction synchrony prosodic entrainment, specifically pitch and energy, in married couples’ problem-solving interactions using speech signal-derived measures. Statistical testings demonstrate that some of these measures capture useful information; they show higher values in interactions with couples having high positive attitude compared to high negative attitude. Further, by using quantized entrainment measures employed with statistical symbol sequence matching in a maximum likelihood framework, we obtained 76% accuracy in predicting positive affect vs. negative affect.
Chi-Chun Lee, Matthew Black, Athanasios Katsamanis, Adam C. Lammert, Brian R. Baucom, Andrew Christensen, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH1
2009 Modeling mutual influence of interlocutor emotion states in dyadic spoken interactions
abstract
In dyadic human interactions, mutual influence- a person’s in-fluence on the interacting partner’s behaviors- is shown to be important and could be incorporated into the modeling frame-work in characterizing, and automatically recognizing the par-ticipants ’ states. We propose a Dynamic Bayesian Network (DBN) to explicitly model the conditional dependency between two interacting partners ’ emotion states in a dialog using data from the IEMOCAP corpus of expressive dyadic spoken in-teractions. Also, we focus on automatically computing the Valence-Activation emotion attributes to obtain a continuous characterization of the participants ’ emotion flow. Our pro-posed DBNmodels the temporal dynamics of the emotion states as well as the mutual influence between speakers in a dialog. With speech based features, the proposed network improves classification accuracy by 3.67 % absolute and 7.12 % relative over the Gaussian Mixture Model (GMM) baseline on isolated turn-by-turn emotion classification. Index Terms: emotion recognition, mutual influence, Dynamic Bayesian Network, dyadic interaction
Chi-Chun Lee, Carlos Busso, Sungbok Lee, Shri Narayanan
INTERSPEECH1
2009 Emotion recognition using a hierarchical binary decision tree approach
abstract
Automated emotion state tracking is a crucial element in the computational study of human communication behaviors. It is important to design robust and reliable emotion recognition systems that are suitable for real-world applications both to enhance analytical abilities to support human decision making and to design human-machine interfaces that facilitate efficient communication. We introduce a hierarchical computational structure to recognize emotions. The proposed structure maps an input speech utterance into one of the multiple emotion classes through subsequent layers of binary classifications. The key idea is that the levels in the tree are designed to solve the easiest classification tasks first, allowing us to mitigate error propagation. We evaluated the classification framework on two different emotional databases using acoustic features, the AIBO database and the USC IEMOCAP database. In the case of the AIBO database, we obtain a balanced recall on each of the individual emotion classes using this hierarchical structure. The performance measure of the average unweighted recall on the evaluation data set improves by 3.37% absolute (8.82% relative) over a Support Vector Machine baseline model. In the USC IEMOCAP database, we obtain an absolute improvement of 7.44% (14.58%) over a baseline Support Vector Machine modeling. The results demonstrate that the presented hierarchical approach is effective for classifying emotional utterances in multiple database contexts.
Chi-Chun Lee, Emily Mower Provost, Carlos Busso, Sungbok Lee, Shri Narayanan
INTERSPEECH1
2008 An analysis of multimodal cues of interruption in dyadic spoken interactions
abstract
Interruptions are integral elements of natural spontaneous human interaction. Both competitive and cooperative interruption serve a distinct role in the flow of conversation. This paper analyzes their differences with features, change and activeness, employing audio, visual, and disfluency data. These features are able to capture differences between the two types of interruptions better than average feature values of any single modality. Also, discriminant analysis shows that the use of multimodal cues provides a 21% improvement in classification accuracy between the two types of interruptions relative to the baseline while any individual single modality cue does not provide significant improvement.
Chi-Chun Lee, Sungbok Lee, Shri Narayanan
INTERSPEECH1