Woan-Shiuan Chien

dblp:276/3164 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0003-2235-4080ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Only Subsets Matters: The Effect of Dual Fairness Constraints in Speech Emotion Recognition
abstract
Speech Emotion Recognition (SER) systems are increasingly deployed in voice-centric applications, yet often suffer from fairness concerns due to speaker-induced variability. In particular, speaker-gender bias can cause systematic disparities in performance across demographic groups (group fairness), while emotionally expressive or acoustically unique speakers may be treated inconsistently despite similar input (individual fairness). Although prior work suggests that group and individual fairness objectives may be inherently incompatible, our proposed two-stage debiasing framework aims to address both: an in-processing approach first mitigates speaker-gender bias (group fairness), followed by a post-processing calibration step that improves consistency across similar instances (individual fairness). While most samples benefit from this dual intervention, we identify a smallsubsetof speech samples that remain difficult to classify fairly. This work focuses on systematically analyzing these fairness-ambiguous samples to understand what makes them challenging. We examine this question from two perspectives: emotion perception and acoustic expressivity. Our analyses on thesesubsetsindicate that: (1) exhibit extreme or atypical emotional ratings, (2) show high acoustic variability, and (3) tend to come disproportionately from specific individuals. These findings suggest that some speakers inherently present greater challenges to fairness optimization, due to the uniqueness of their emotional or acoustic expression. By characterizing thesesubsets, our work contributes to a deeper understanding of fairness conflicts in SER and offers new directions for developing more robust and inclusive emotion recognition systems.
Woan-Shiuan Chien, Chi-Chun Lee
IEEE Trans. Affect. Comput.1
2025 Disentangle Heart Rate Signals for Improved Stress Detection
abstract
Accurate stress detection from physiological signals is often complicated by individual identity traits, which must first be identified before they can be effectively removed to improve model performance. To address this, we propose a method that combines Detrended Fluctuation Analysis (DFA) and Augmented Dickey-Fuller (ADF) Analysis to extract stable identity-related features from heart rate signals without relying on explicit labels. By masking these identity features and raw content features, we can effectively eliminate their impact on task-relevant stress signals using guided contrastive learning. Validated on the TILES-2018 and Firefighters datasets, our approach significantly improves stress detection accuracy, achieving F1 score gains of 5.3% and 8.1%, respectively, compared to baseline models. These results highlight the model’s enhanced ability to generalize to diverse populations while minimizing identity bias, ultimately improving the robustness and precision of stress detection systems.
Pin-Jhao Chen, Woan-Shiuan Chien, Chi-Chun Lee
ICASSP2
2025 Is It Still Fair? Investigating Gender Fairness in Cross-Corpus Speech Emotion Recognition
abstract
Speech emotion recognition (SER) is a vital component in various everyday applications. Cross-corpus SER models are increasingly recognized for their ability to generalize performance. However, concerns arise regarding fairness across demographics in diverse corpora. Existing fairness research often focuses solely on corpus-specific fairness, neglecting its generalizability in cross-corpus scenarios. Our study focuses on this underexplored area, examining the gender fairness generalizability in cross-corpus SER scenarios. We emphasize that the performance of cross-corpus SER models and their fairness are two distinct considerations. Moreover, we propose the approach of a combined fairness adaptation mechanism to enhance gender fairness in the SER transfer learning tasks by addressing both source and target genders. Our findings bring one of the first insights into the generalizability of gender fairness in cross-corpus SER systems.
Shreya G. Upadhyay, Woan-Shiuan Chien, Chi-Chun Lee
ICASSP2
2025 A Dynamic Edge-Selection Mechanism in HRV Hypergraph Learning for Improved Stress Detection
abstract
Studies show that individual attributes such as age and gender significantly influence physiological responses and their correlation with stress, often forming complex and overlapping relationships. These attributes are essential for enhancing physiological signal-based stress detection. Our work leverages hypergraph in multi-attribute representation and tackles the challenge of redundant attributes that misguide embeddings. Our dynamic edge-selection mechanism for hypergraph-based metric learning (DESHM) enables the hypergraph to focus on selected stress-related attribute connections within groups. This batch-wise selection adapts to varying connections between batches and maximizes the effectiveness of metric learning. Evaluation of the TILES-2018 and Firefighter datasets shows promising results, further improving 3.99% in F1, 3.52% in BACC, and 25.27% in MCC compared to the best pairwise result on the TILES-2018 dataset. Our analysis shows that our mechanism prioritizes attributes that effectively represent stress levels, guiding hyper-graph clustering to achieve improved discriminability.
Jing-Chun Wang, Woan-Shiuan Chien, Chi-Chun Lee
ICASSP2
2025 Differential Impacts of Monologue and Conversation on Speech Emotion Recognition
abstract
The advancement ofSpeech Emotion Recognition(SER) is significantly dependent on the quality of emotional speech corpora used for model training. Researchers in the field of SER have developed various corpora by adjusting design parameters to enhance the reliability of the training source. For this study, we focus on exploring communication modes of collection, specifically analyzing spontaneous emotional speech patterns gathered during conversation or monologue. While conversations are acknowledged as effective for eliciting authentic emotional expressions, systematic analyses are necessary to confirm their reliability as a better source of emotional speech data. We investigate this research question from perceptual differences and acoustic variability present in both emotional speeches. Our analyses on multi-lingual corpora show that, first, raters exhibit higher consistency for conversation recordings when evaluating categorical emotions, and second, perceptions and acoustic patterns observed in conversational samples align more closely with expected trends discussed in relevant emotion literature. We further examine the impact of these differences on SER modeling, which shows that we can train a more robust and stable SER model by using conversation data. This work provides comprehensive evidence suggesting that conversation may offer a better source compared to monologue for developing an SER model.
Woan-Shiuan Chien, Shreya G. Upadhyay, Carlos Busso, Chi-Chun Lee
IEEE Trans. Affect. Comput.1
2024 Stress Detection Using HRV Features Augmentation Based on Heart Rate Signal Transformation
abstract
Heart Rate Variability (HRV) features are recognized as powerful indicators of various diseases, including heart failure, diabetes, and mental health disorders. Besides, HRV features are robust against noise, making them ideal for wearable devices. Despite their potential, HRV feature sets are limited by sample quantity. Direct augmentation often distorts underlying signal properties and interpretability. This study addresses this limitation by introducing spatial-domain, temporal-domain, and serial signal transformations for HRV feature-based augmentations. We further analyze the transformed signal using the Poincare plots to understand the effect of the transformation effect from the clinical perspective. Lastly, we design a stress detection deep learning model using the TILES-2018 database to verify the effectiveness of the augmented HRV features, which shows performance enhancements ranging from 1.05 % to 4.52%. Among the three domains of transformations, serial transformations yield the best results.
Pin-Jhao Chen, Woan-Shiuan Chien, Huan-Yu Chen, Chi-Chun Lee
BSN2
2024 Understanding Missing Data Bias in Longitudinal Mental Stress Detection
abstract
Mental stress has become a growing concern in contemporary society; fortunately, recent developments in wear-able technology now offer a promising solution. However, a common issue in longitudinal tracking with wearable sensors is missing data, which can introduce biases during model training, affecting predictions and leading to unfair outcomes for users. In this work, we explore the impact of missing data on stress detection performance across two longitudinal datasets. Our analysis reveals that biases stemming from missing data can result in the unfair treatment of individuals with higher levels of missing data, detrimentally affecting model performance. Additionally, we assess various imputation methods to mitigate these issues. Our findings indicate that while imputation generally improves model overall performances, performance decreases significantly when missing data exceeds half of the total data. This research provides initial insights into the challenges of missing data in longitudinal studies.
Woan-Shiuan Chien, Chi-Chun Lee
BSN1
2024 In-The-Wild HRV-Based Stress Detection Using Individual-Aware Metric Learning
abstract
Advanced wearable tracking shows potential for identifying psychological and emotional stress relevant to the mental health of high-intensity emergency responders. Heart rate variability (HRV) captured by wearable devices can indi-cate the correlation between intra-subject daily variations and stress. HRV also varies due to various demographic attributes, representing inter-subject relationships. This work introduces an individual-aware metric learning approach that leverages HRV features to train intra-subject representations, considering inter-subject effects based on attribute similarity through stress label clustering. We use the multi-similarity loss within the metric learning framework to consider various personal attributes, thereby improving discriminability. Evaluation of the TILES-2018 and Firefighter database shows promising results in binary stress classification: F1 score of 68.15 % with BACC of 59.13 % and MCC of 0.186, and F1 score of 73.07% with BACC of 56.52% and MCC of 0.136. resnectively.
Jing-Chun Wang, Woan-Shiuan Chien, Huan-Yu Chen, Chi-Chun Lee
BSN2
2024 Balancing Speaker-Rater Fairness for Gender-Neutral Speech Emotion Recognition
abstract
Speech emotion recognition (SER) adds to the humane aspects of voice technologies to enhance user experiences. The ground truth emotion annotations provided by human raters and attributes related to the speakers themselves arise a compounded fairness issue in SER. While there exist works in fair SER, our work presents one of the first studies in addressing the unique joint speaker-rater (two-sided) bias, focusing on the issue of gender fairness. Our cross-reference evaluation demonstrates that the SER fair model, which merely mitigates one-sided bias introduces biases when examining from another viewpoint. Furthermore, in order to handle model stability when optimizing for these compounded speaker-rater constraints, we introduce a flexible controlled mechanism that dynamically balances the contribution of each viewpoint. Our analyses show the efficacy of our approach in achieving a fair SER that meets the dual speaker-rater gender neutrality criterion.
Woan-Shiuan Chien, Shreya G. Upadhyay, Chi-Chun Lee
ICASSP1
2024 In-The-Wild Physiological-Based Stress Detection Using Federated Strategy
abstract
Continuously identifying day-to-day mental stress can be realized by accessing wearable devices to measure physiological indicators. However, the nature of bodily signals raises issues of privacy and data heterogeneity. Recent federated learning scheme provides a promising direction to alleviate the privacy concern, but the large inter-client differences can lead to a sub-optimal model performance. In this work, we propose a client-aware aggregation strategy to customize the global model forked by each client to conduct mutual learning in federated setting. Our proposed mixture Federated Mutual Learning (mixFML) weighs the distances of local models to generate a unique mixture of global model per client. We evaluated our method on the public TILES-2018 and an in-house Firefighters dataset for stress detection using HRV. Our proposed mixFML achieved 8.0% and 1.8% MCC improvement on two datasets compared to federated mutual learning.
Po-Chen Lin, Jeng-Lin Li, Woan-Shiuan Chien, Chi-Chun Lee
ICASSP3
2024 An Investigation of Group versus Individual Fairness in Perceptually Fair Speech Emotion Recognition
Woan-Shiuan Chien, Chi-Chun Lee
INTERSPEECH1
2024 An Inter-Speaker Fairness-Aware Speech Emotion Regression Framework
Hsing-Hang Chou, Woan-Shiuan Chien, Ya-Tse Wu, Chi-Chun Lee
INTERSPEECH2
2024 Learning With Rater-Expanded Label Space to Improve Speech Emotion Recognition
abstract
Automatic sensing of emotional information in speech is important for numerous everyday applications. Conventional Speech Emotion Recognition (SER) models rely on averaging or consensus of human annotations for training, but emotions and raters' interpretations are subjective in nature, leading to diverse variations in perceptions. To address this, our proposed approach integrates the rater's subjectivity by forming the Perception-Coherent Clusters (PCC) of raters to be used to derive expanded label space for learning to improve SER. We evaluate our method on the IEMOCAP and the MSP-Podcast corpora, considering scenarios of fixed and variable raters, respectively. The proposed architecture, Rater Perception Coherency (RPC)-based SER surpasses single-task models with consensus labels by achieving UAR improvements of 3.39% for the IEMOCAP and 2.03% for the MSP-Podcast. Further analysis provides comprehensive insights into the contributions of these perception consistency clusters in SER learning.
Shreya G. Upadhyay, Woan-Shiuan Chien, Bo-Hao Su, Chi-Chun Lee
IEEE Trans. Affect. Comput.2
2023 An Intelligent Infrastructure Toward Large Scale Naturalistic Affective Speech Corpora Collection
abstract
The field of speech emotion recognition (SER) aims to create scientifically rigorous systems that can reliably characterize emotional behaviors expressed in speech. A key aspect for building SER systems is to obtain emotional data that is both reliable and reproducible for practitioners. However, academic researchers encounter difficulties in accessing or collecting naturalistic large-scale, reliable emotional recordings. Also, the best practices for data collection are not necessarily described or shared when presenting emotional corpora. To address this issue, the paper proposes the creation of an affective naturalistic database consortium (AndC) that can encourage multidisciplinary cooperation among researchers and practitioners in the field of affective computing. This paper’s contribution is twofold. First, it proposes the design of the AndC with a customizable-standard framework for intelligently-controlled emotional data collection. The focus is on leveraging naturalistic spontaneous recordings available on audio-sharing websites. Second, it presents as a case study the development of a naturalistic large-scale Taiwanese Mandarin podcast corpus using the customizable-standard intelligently-controlled framework. The AndC will enable research groups to effectively collect data using the provided pipeline and to contribute with alternative algorithms or data collection protocols.
Shreya G. Upadhyay, Woan-Shiuan Chien, Bo-Hao Su, Lucas Goncalves, Ya-Tse Wu, Ali N. Salman, Carlos Busso, Chi-Chun Lee
ACII2
2023 Achieving Fair Speech Emotion Recognition via Perceptual Fairness
abstract
Speech emotion recognition (SER) is a key technological module to be integrated into many voice-based solutions. One of the unique fairness issues in SER is caused by the inherently biased emotion perception given by the raters as ground truth labels. Mitigating rater biases are at core for SER to move toward optimizing both recognition and fairness performance. In this work, we proposed a two-stage framework, which produces debiased representations by using a fairness constraint adversarial framework in the first stage. Then, users are endued with the right to toggle between specified gender-wise perceptions on-demand after the gender-wise perceptual learning in the second stage. We further evaluate our results on two important fairness metrics to show that the distributions and predictions across different gender are fair.
Woan-Shiuan Chien, Chi-Chun Lee
ICASSP1
2023 Phonetic Anchor-Based Transfer Learning to Facilitate Unsupervised Cross-Lingual Speech Emotion Recognition
abstract
Modeling cross-lingual speech emotion recognition (SER) has become more prevalent because of its diverse applications. Existing studies have mostly focused on technical approaches that adapt the feature, domain, or label across languages, without considering in detail the similarities between the languages. This study focuses on domain adaptation in cross-lingual scenarios using phonetic constraints. This work is framed in a twofold manner. First, we analyze emotion-specific phonetic commonality across languages by identifying common vowels that are useful for SER modeling. Second, we leverage these common vowels as an anchoring mechanism to facilitate cross-lingual SER. We consider American English and Taiwanese Mandarin as a case study to demonstrate the potential of our approach. This work uses two in-the-wild natural emotional speech corpora: MSP-Podcast (American English), and BIIC-Podcast (Taiwanese Mandarin). The proposed unsupervised cross-lingual SER model using these phonetical anchors outperforms the baselines with a 58.64% of unweighted average recall (UAR).
Shreya G. Upadhyay, Luz Martinez-Lucas, Bo-Hao Su, Woan-Shiuan Chien, Ya-Tse Wu, William F. Katz, Carlos Busso, Chi-Chun Lee
ICASSP5
2022 Monologue versus Conversation: Differences in Emotion Perception and Acoustic Expressivity
abstract
Advancing speech emotion recognition (SER) depends highly on the source used to train the model, i.e., the emotional speech corpora. By permuting different design parameters, researchers have released versions of corpora that attempt to provide a better-quality source for training SER. In this work, we focus on studying communication modes of collection. In particular, we analyze the patterns of emotional speech collected during interpersonal conversations or monologues. While it is well known that conversation provides a better protocol for eliciting authentic emotion expressions, there is a lack of systematic analyses to determine whether conversational speech provide a “better-quality” source. Specifically, we examine this research question from three perspectives: perceptual differences, acoustic variability and SER model learning. Our analyses on the MSP-Podcast corpus show that: 1) rater's consistency for conversation recordings is higher when evaluating categorical emotions, 2) the perceptions and acoustic patterns observed on conversations have properties that are better aligned with expected trends discussed in emotion literature, and 3) a more robust SER model can be trained from conversational data. This work brings initial evidences stating that samples of conversations may provide a better-quality source than samples from monologues for building a SER model.
Woan-Shiuan Chien, Shreya G. Upadhyay, Ya-Tse Wu, Bo-Hao Su, Carlos Busso, Chi-Chun Lee
ACII1
2020 Cross Corpus Physiological-based Emotion Recognition Using a Learnable Visual Semantic Graph Convolutional Network
abstract
Affective media videos have been used as stimulus to investigate an individual's affective-physio responses. In this study, we aim to develop a network learning strategy for robust cross-corpus emotion recognition using physiological features jointly with affective video content. Specifically, we present a novel framework of Visual Semantic Graph Learning Convolutional Network (VGLCN) for individual emotional state recognition using physiology on transfer learning tasks. The stimulus of videos content is integrated into learnable graph structure to weight the importance of physiology on the two emotion dimensions, valence and arousal. Furthermore, we evaluate our proposed framework on two public emotion databases with a rigorous cross validation method, and our model achieves the best unweighted average recall (UAR), which is 67.9%, 56.9% for arousal and 79.8%, 70.4% for valence on the cross datasets recognition experiments respectively. Further analyses reveal that 1) VGLCN is especially effective on transfer valence binary-task, 2) the physiological features (ECG, EDA) are very informative features for emotion recognition and 3) the affective media videos are important constraint to be included in the framework to stabilize the performance power.
Woan-Shiuan Chien, Hao-Chun Yang, Chi-Chun Lee
ACM Multimedia1