Yuan Gao 0040

dblp:76/2452-40 · DBLP profile ↗
← Back
11ranked-venue papers
8as first author
10since 2021 · last 2026
0000-0002-2147-1835ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 5 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Bridging Speech Emotion Recognition and Personality: Dataset and Temporal Interaction Condition Network
abstract
This study investigates the interaction between personality traits and emotion expression, exploring how personality information can improve speech emotion recognition (SER). We collect the personality annotation for the IEMOCAP dataset, making it the first speech dataset that contains both emotion and personality annotations (PA-IEMOCAP), and enabling direct integration of personality traits into SER. Statistical analysis on this dataset identified significant correlations between per sonality traits and emotional expressions. To extract finegrained personality features, we propose a temporal interaction condition network (TICN), in which personality features are integrated with HuBERT-based acoustic features for SER. Experiments show that incorporating ground-truth personality traits significantly enhances valence recognition, improving the concordance correlation coefficient (CCC) from 0.698 to 0.785 compared to the baseline without personality information. For practical applications in dialogue systems where personality information about the user is unavailable, we develop a front-end module of automatic personality recognition. Using these automatically predicted traits as inputs to our proposed TICN model, we achieve a CCC of 0.776 for valence recognition, representing an 11.17% relative improvement over the baseline. These findings confirm the effectiveness of personality-aware SER and provide a solid foundation for further exploration in personality-aware speech processing applications.
Yuan Gao 0040, Yahui Fu 0001, Chenhui Chu, Tatsuya Kawahara
IEEE Trans. Affect. Comput.1
2024 Enhancing Two-Stage Finetuning for Speech Emotion Recognition Using Adapters
abstract
This study investigates the effective finetuning of a pretrained model using adapters for speech emotion recognition (SER). Since emotion is related with linguistic and prosodic information and also other attributes such as gender and speaking style, a framework of multi-task learning (MTL) has been shown to be effective for SER. However, the learning targets of automatic speech recognition (ASR) and other attribute recognition are apparently in conflict. Therefore, we propose to employ different adaptation methods for different tasks in multiple finetuning stages. Since ASR is the most challenging and also influential for SER, in the first stage, we finetune all parameters of the pretrained model for ASR and SER. In the second stage, we incorporate adapters to finetune the model for gender and style recognition in addition to SER by freezing the parameters of the main Transformer model tuned for ASR. Experimental evaluations which extensively compare different adaptation methods using the IEMOCAP dataset demonstrate that the proposed approach achieves a significant improvement from the simple MTL.
Yuan Gao 0040, Chenhui Chu, Tatsuya Kawahara
ICASSP1
2024 Speech Emotion Recognition with Multi-level Acoustic and Semantic Information Extraction and Interaction
Yuan Gao 0040, Chenhui Chu, Tatsuya Kawahara
INTERSPEECH1
2024 Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
abstract
Serialized output training (SOT) attracts increasing attention due to its convenience and flexibility for multi-speaker automatic speech recognition (ASR). However, it is not easy to train with attention loss only. In this paper, we propose the overlapped encoding separation (EncSep) to fully utilize the benefits of the connectionist temporal classification (CTC) and attention (CTC-Attention) hybrid loss. This additional separator is inserted after the encoder to extract the multi-speaker information with CTC losses. Furthermore, we propose the serialized speech information guidance SOT (GEncSep) to further utilize the separated encodings. The separated streams are concatenated to provide single-speaker information to guide attention during decoding. The experimental results on Libri2Mix and Libri3Mix show that the single-speaker encoding can be separated from the overlapped encoding. The CTC loss helps to improve the encoder representation under complex scenarios (three-speaker and noisy conditions), which makes the EncSep have a relative improvement of more than 8% and 6% on the noisy Libri2Mix and Libri3Mix evaluation sets, respectively. GEncSep further improved performance, which was more than 12% and 9% relative improvement for the noisy Libri2Mix and Libri3Mix evaluation sets.
Yuan Gao 0040, Zhaoheng Ni, Tatsuya Kawahara
SLT2
2024 Adversarial Domain Generalized Transformer for Cross-Corpus Speech Emotion Recognition
abstract
Speech emotion recognition (SER) promotes the development of intelligent devices, which enable natural and friendly human-computer interactions. However, the recognition performance of existing approaches is significantly reduced on unseen datasets, and the lack of sufficient training data limits the generalizability of deep learning models. In this work, we analyze the impact of the domain generalization method on cross-corpus SER and propose an adversarial domain generalized transformer (ADoGT), which is aimed at learning a shared feature distribution for the source and target domains. Specifically, we investigate the effect of domain adversarial learning by eliminating nonaffective information. We also combine the center loss with the softmax function as joint supervision to learn discriminative features. Moreover, we introduce unsupervised transfer learning to extract additional features, and incorporate a gated fusion model to learn the complementary information of the features learned by the supervised feature extractor and pretrained model. The proposed transformer based domain generalization method is evaluated using four emotional datasets. We also provide an ablation study of different domain adversarial model structures and feature fusion models. The results of comparative experiments demonstrate the effectiveness of the proposed ADoGT.
Yuan Gao 0040, Longbiao Wang, Jiaxing Liu 0001, Jianwu Dang 0001, Shogo Okada
IEEE Trans. Affect. Comput.1
2023 FedCPC: An Effective Federated Contrastive Learning Method for Privacy Preserving Early-Stage Alzheimers Speech Detection
abstract
The early-stage Alzheimer’s disease (AD) detection has been considered an important field of medical studies. Like traditional machine learning methods, speech-based automatic detection also suffers from data privacy risks because the data of specific patients are exclusive to each medical institution. A common practice is to use federated learning to protect the patients’ data privacy. However, its distributed learning process also causes performance reduction. To alleviate this problem while protecting user privacy, we propose a federated contrastive pre-training (FedCPC) performed before federated training for AD speech detection, which can learn a better representation from raw data and enables different clients to share data in the pre-training and training stages. Experimental results demonstrate that the proposed methods can achieve satisfactory performance while preserving data privacy.
Wenqing Wei, Zhengdong Yang, Yuan Gao 0040, Jiyi Li, Chenhui Chu, Shogo Okada, Sheng Li 0010
ASRU3
2023 Two-stage Finetuning of Wav2vec 2.0 for Speech Emotion Recognition with ASR and Gender Pretraining
Yuan Gao 0040, Chenhui Chu, Tatsuya Kawahara
INTERSPEECH1
2022 Domain-Invariant Feature Learning for Cross Corpus Speech Emotion Recognition
abstract
To deal with speech emotion recognition (SER) in real-life applications, researchers have to focus on cross corpus SER, where the feature distribution of source and target datasets are different. In this paper, we propose an efficient domain adversarial training method to cope with the non-affective information during feature extraction. Through the proposed domain-adversarial learning, we can reduce the domain divergence between train and test data. Furthermore, we incorporate center loss with the emotion classifier to reduce the intra-class variation of features learned from the same emotion. We conduct experiments on four emotional benchmark datasets to verify the performance of the proposed method. The experimental results demonstrate that our proposed model outperform the baseline system in both cross-corpus and multi-corpus evaluation.
Yuan Gao 0040, Shogo Okada, Longbiao Wang, Jiaxing Liu 0001, Jianwu Dang 0001
ICASSP1
2021 Domain-Adversarial Autoencoder with Attention Based Feature Level Fusion for Speech Emotion Recognition
abstract
Over the past two decades, although speech emotion recognition (SER) has garnered considerable attention, the problem of insufficient training data has been unresolved. A potential solution for this problem is to pre-train a model and transfer knowledge from large amounts of audio data. However, the data used for pre-training and testing originate from different domains, resulting in the latent representations to contain non-affective information. In this paper, we propose a domain-adversarial autoencoder to extract discriminative representations for SER. Through domain-adversarial learning, we can reduce the mismatch between domains while retaining discriminative information for emotion recognition. We also introduce multi-head attention to capture emotion information from different subspaces of input utterances. Experiments on IEMOCAP show that the proposed model outperforms the state-of-the-art systems by improving the unweighted accuracy by 4.15%, thereby demonstrating the effectiveness of the proposed model.
Yuan Gao 0040, Jiaxing Liu 0001, Longbiao Wang, Jianwu Dang 0001
ICASSP1
2021 Metric Learning Based Feature Representation with Gated Fusion Model for Speech Emotion Recognition
Yuan Gao 0040, Jiaxing Liu 0001, Longbiao Wang, Jianwu Dang 0001
Interspeech1
2020 Temporal Attention Convolutional Network for Speech Emotion Recognition with Latent Representation
Jiaxing Liu 0001, Zhilei Liu, Longbiao Wang, Yuan Gao 0040, Lili Guo 0001, Jianwu Dang 0001
INTERSPEECH4