EDBT 2026 Demo / reviewers in the wild / expert
Ya-Tse Wu
dblp:212/6444
· DBLP profile ↗
16ranked-venue papers
6as first author
13since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RE-LLM: Refining Empathetic Speech-LLM Responses by Integrating Emotion NuanceabstractWith generative AI advancing, empathy in human-AI interaction is essential. While prior work focuses on emotional reflection, emotional exploration—key to deeper engagement—remains overlooked. Existing LLMs rely on text which captures limited emotion nuances. To address this, we propose RE-LLM, a speech-LLM integrating dimensional emotion embeddings and auxiliary learning. Experiments show statistically significant gains in empathy metrics almost across three datasets. RE-LLM relatively improves the Emotional Reaction score by 14.79% and $\mathbf{6. 7 6} \boldsymbol{\%}$ compared to text-only and speech-LLM baselines on ESD. Notably, it raises the Exploration score by 35.42% and 3.91% on IEMOCAP, 139.28% and 9.83% on ESD and 60.95% and 22.64% on MSP-PODCAST relatively. It also boosts unweighted accuracy by 5.4% on IEMOCAP, 2.3% on ESD and $\mathbf{6. 9 \%}$ on MSP-PODCAST in speech emotion recognition. These results highlight the enriched emotional understanding and improved empathetic response generation of RE-LLM. Jing-Han Chen, Bo-Hao Su, Ya-Tse Wu, Chi-Chun Lee |
ASRU | 3 |
| 2025 | ASR for Affective Speech: Investigating Impact of Emotion and Speech Generative StrategyabstractThis work investigates how emotional speech and generative strategies affect ASR performance. We analyze speech synthesized from three emotional TTS models and find that substitution errors dominate, with emotional expressiveness varying across models. Based on these insights, we introduce two generative strategies: one using transcription correctness and another using emotional salience, to construct fine-tuning subsets. Results show consistent WER improvements on real emotional datasets without noticeable degradation on clean LibriSpeech utterances. The combined strategy achieves the strongest gains, particularly for expressive speech. These findings highlight the importance of targeted augmentation for building emotion-aware ASR systems. Ya-Tse Wu, Chi-Chun Lee |
ASRU | 1 |
| 2025 | Defend for Self-Vocoding: A Novel Enhanced Decoder Network for Watermark Recovery
Ching-Yu Yang, Hsing-Hang Chou, Ya-Tse Wu, Bo-Hao Su, Chi-Chun Lee |
INTERSPEECH | 4 |
| 2025 | Lessons Learnt: Revisit Key Training Strategies for Effective Speech Emotion Recognition in the WildabstractIn this study, we revisit key training strategies in machine learning often overlooked in favor of deeper architectures. Specifically, we explore balancing strategies, activation functions, and fine-tuning techniques to enhance speech emotion recognition (SER) in naturalistic conditions. Our findings show that simple modifications improve generalization with minimal architectural changes. Our multi-modal fusion model, integrating these optimizations, achieves a valence CCC of 0.6953, the best valence score in Task 2: Emotional Attribute Regression. Notably, fine-tuning RoBERTa and WavLM separately in a single-modality setting, followed by feature fusion without training the backbone extractor, yields the highest valence performance. Additionally, focal loss and activation functions significantly enhance performance without increasing complexity. These results suggest that refining core components, rather than deepening models, leads to more robust SER in-the-wild. Jing-Tong Tzeng, Bo-Hao Su, Ya-Tse Wu, Hsing-Hang Chou, Chi-Chun Lee |
INTERSPEECH | 3 |
| 2024 | An Inter-Speaker Fairness-Aware Speech Emotion Regression Framework
Hsing-Hang Chou, Woan-Shiuan Chien, Ya-Tse Wu, Chi-Chun Lee |
INTERSPEECH | 3 |
| 2024 | Can Modelling Inter-Rater Ambiguity Lead To Noise-Robust Continuous Emotion Predictions?abstractThere has been increasing attention drawn to modelling interrater ambiguity in Continuous Emotion Recognition (CER) systems using probability distributions for arousal and valence.However, the relationship between modelling label ambiguity and robustness to noise, and more broadly, the impact of realworld noise on CER systems remains insufficiently explored.In this study, we argue that incorporating inter-rater ambiguity during training can regularize the noise response, leading to noise robustness.To this end, we propose a novel loss function that incorporates inter-rater ambiguity into model training.Experiments conducted on the RECOLA dataset demonstrate that our proposed method achieves a maximum Concordance Correlation Coefficient (CCC) improvement of 0.117 and 0.077 for mean and standard deviation predictions, respectively, across all noise conditions.We further integrate traditional noisy augmentation strategies with our proposed method and observe promising results. Ya-Tse Wu, Jingyao Wu 0002, Vidhyasaharan Sethu, Chi-Chun Lee |
INTERSPEECH | 1 |
| 2024 | RW-VoiceShield: Raw Waveform-based Adversarial Attack on One-shot Voice Conversion
Ching-Yu Yang, Shreya G. Upadhyay, Ya-Tse Wu, Bo-Hao Su, Chi-Chun Lee |
INTERSPEECH | 3 |
| 2023 | An Intelligent Infrastructure Toward Large Scale Naturalistic Affective Speech Corpora CollectionabstractThe field of speech emotion recognition (SER) aims to create scientifically rigorous systems that can reliably characterize emotional behaviors expressed in speech. A key aspect for building SER systems is to obtain emotional data that is both reliable and reproducible for practitioners. However, academic researchers encounter difficulties in accessing or collecting naturalistic large-scale, reliable emotional recordings. Also, the best practices for data collection are not necessarily described or shared when presenting emotional corpora. To address this issue, the paper proposes the creation of an affective naturalistic database consortium (AndC) that can encourage multidisciplinary cooperation among researchers and practitioners in the field of affective computing. This paper’s contribution is twofold. First, it proposes the design of the AndC with a customizable-standard framework for intelligently-controlled emotional data collection. The focus is on leveraging naturalistic spontaneous recordings available on audio-sharing websites. Second, it presents as a case study the development of a naturalistic large-scale Taiwanese Mandarin podcast corpus using the customizable-standard intelligently-controlled framework. The AndC will enable research groups to effectively collect data using the provided pipeline and to contribute with alternative algorithms or data collection protocols. Shreya G. Upadhyay, Woan-Shiuan Chien, Bo-Hao Su, Lucas Goncalves, Ya-Tse Wu, Ali N. Salman, Carlos Busso, Chi-Chun Lee |
ACII | 5 |
| 2023 | Phonetic Anchor-Based Transfer Learning to Facilitate Unsupervised Cross-Lingual Speech Emotion RecognitionabstractModeling cross-lingual speech emotion recognition (SER) has become more prevalent because of its diverse applications. Existing studies have mostly focused on technical approaches that adapt the feature, domain, or label across languages, without considering in detail the similarities between the languages. This study focuses on domain adaptation in cross-lingual scenarios using phonetic constraints. This work is framed in a twofold manner. First, we analyze emotion-specific phonetic commonality across languages by identifying common vowels that are useful for SER modeling. Second, we leverage these common vowels as an anchoring mechanism to facilitate cross-lingual SER. We consider American English and Taiwanese Mandarin as a case study to demonstrate the potential of our approach. This work uses two in-the-wild natural emotional speech corpora: MSP-Podcast (American English), and BIIC-Podcast (Taiwanese Mandarin). The proposed unsupervised cross-lingual SER model using these phonetical anchors outperforms the baselines with a 58.64% of unweighted average recall (UAR). Shreya G. Upadhyay, Luz Martinez-Lucas, Bo-Hao Su, Woan-Shiuan Chien, Ya-Tse Wu, William F. Katz, Carlos Busso, Chi-Chun Lee |
ICASSP | 6 |
| 2023 | A Context-Constrained Sentence Modeling for Deception Detection in Real Interrogation
Ya-Tse Wu, Yuan-Ting Chang, Shao-Hao Lu, Jing-Yi Chuang, Chi-Chun Lee |
INTERSPEECH | 1 |
| 2023 | MetricAug: A Distortion Metric-Lead Augmentation Strategy for Training Noise-Robust Speech Emotion Recognizer
Ya-Tse Wu, Chi-Chun Lee |
INTERSPEECH | 1 |
| 2022 | Monologue versus Conversation: Differences in Emotion Perception and Acoustic ExpressivityabstractAdvancing speech emotion recognition (SER) depends highly on the source used to train the model, i.e., the emotional speech corpora. By permuting different design parameters, researchers have released versions of corpora that attempt to provide a better-quality source for training SER. In this work, we focus on studying communication modes of collection. In particular, we analyze the patterns of emotional speech collected during interpersonal conversations or monologues. While it is well known that conversation provides a better protocol for eliciting authentic emotion expressions, there is a lack of systematic analyses to determine whether conversational speech provide a “better-quality” source. Specifically, we examine this research question from three perspectives: perceptual differences, acoustic variability and SER model learning. Our analyses on the MSP-Podcast corpus show that: 1) rater's consistency for conversation recordings is higher when evaluating categorical emotions, 2) the perceptions and acoustic patterns observed on conversations have properties that are better aligned with expected trends discussed in emotion literature, and 3) a more robust SER model can be trained from conversational data. This work brings initial evidences stating that samples of conversations may provide a better-quality source than samples from monologues for building a SER model. Woan-Shiuan Chien, Shreya G. Upadhyay, Ya-Tse Wu, Bo-Hao Su, Carlos Busso, Chi-Chun Lee |
ACII | 4 |
| 2022 | An Audio-Saliency Masking Transformer for Audio Emotion Classification in MoviesabstractThe process of perception to affective response of humans is gated by a bottom-up saliency mechanism at the sensory level. In specifics, auditory saliency emphasizes audio segments that need to be attended to cognitively appraise and experience emotion. In this work, inspired by this mechanism, we propose an end-to-end feature masking network for audio emotion recognition in movies. Our proposed Audio-Saliency Masking Transformer (ASTM) adjusts feature embedding using two learnable masks; one of them cross-refers to an auditory saliency map, and the other one is through self-reference. By joint training for front-end mask gating and the transformer as the back-end emotion classifier, we achieve three-class UARs improvement of 1.74%, 1.27%, 0.95%, 0.82% when comparing to the best of the other models on experienced arousal, experienced valence, intended arousal, and intended valence, respectively. We further analyze which acoustic feature categories that our saliency mask attends to the most. Ya-Tse Wu, Jeng-Lin Li, Chi-Chun Lee |
ICASSP | 1 |
| 2018 | Integrating Perceivers Neural-Perceptual Responses Using a Deep Voting Fusion Network for Automatic Vocal Emotion DecodingabstractUnderstanding neuro-perceptual mechanism of vocal emotion perception continues to be an important research direction not only in advancing scientific knowledge but also in inspiring more robust affective computing technologies. The large variabilities in the manifested fMRI signals among subjects has been shown to be due to the effect of individual difference, i.e., inter-subject variability. However, relatively few works have developed modeling techniques in task of automatic neuro-perceptual decoding to handle such idiosyncrasies. In our work, we propose a novel computation method of deep voting fusion neural network architecture by learning an adjusted weight matrix applied at the fusion layer. The framework achieves an unweighted average recall of 53.10% in a four-class vocal emotion states decoding task, i.e., a relative improvement of 8.9% over a two-stage SVM decision-level fusion. Our framework demonstrates its effectiveness in handling individual differences. Further analysis is conducted to study the properties of the learned adjusted weight matrix as a function of emotion classification accuracy. Wan-Ting Hsieh, Hao-Chun Yang, Ya-Tse Wu, Fu-Sheng Tsai, Li-Wei Kuo, Chi-Chun Lee |
ICASSP | 3 |
| 2018 | Generating fMRI-Enriched Acoustic Vectors using a Cross-Modality Adversarial Network for Emotion RecognitionabstractAutomatic emotion recognition has long been developed by concentrating on modeling human expressive behavior. At the same time, neuro-scientific evidences have shown that the varied neuro-responses (i.e., blood oxygen level-dependent (BOLD) signals measured from the functional magnetic resonance imaging (fMRI)) is also a function on the types of emotion perceived. While past research has indicated that fusing acoustic features and fMRI improves the overall speech emotion recognition performance, obtaining fMRI data is not feasible in real world applications. In this work, we propose a cross modality adversarial network that jointly models the bi-directional generative relationship between acoustic features of speech samples and fMRI signals of human percetual responses by leveraging a parallel dataset. We encode the acoustic descriptors of a speech sample using the learned cross modality adversarial network to generate the fMRI-enriched acoustic vectors to be used in the emotion classifier. The generated fMRI-enriched acoustic vector is evaluated not only in the parallel dataset but also in an additional dataset without fMRI scanning. Our proposed framework significantly outperform using acoustic features only in a four-class emotion recognition task for both datasets, and the use of cyclic loss in learning the bi-directional mapping is also demonstrated to be crucial in achieving improved recognition rates. Gao-Yi Chao, Chun-Min Chang, Jeng-Lin Li, Ya-Tse Wu, Chi-Chun Lee |
ICMI | 4 |
| 2017 | Modeling Perceivers Neural-Responses Using Lobe-Dependent Convolutional Neural Network to Improve Speech Emotion Recognition
Ya-Tse Wu, Yu-Hsien Liao, Li-Wei Kuo, Chi-Chun Lee |
INTERSPEECH | 1 |