Changzeng Fu

dblp:260/0086 · DBLP profile ↗
← Back
27ranked-venue papers
12as first author
27since 2021 · last 2026
0000-0003-1083-9486ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 8 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Personality-guided Public-Private Domain Disentangled Hypergraph-Former Network for Multimodal Depression Detection
abstract
Depression represents a global mental health challenge requiring efficient and reliable automated detection methods. Current Transformer- or Graph Neural Networks (GNNs)-based multimodal depression detection methods face significant challenges in modeling individual differences and cross-modal temporal dependencies across diverse behavioral contexts. Therefore, we propose P³HF (Personality-guided Public-Private Domain Disentangled Hypergraph-Former Network) with three key innovations: (1) personality-guided representation learning using LLMs to transform discrete individual features into contextual descriptions for personalized encoding; (2) Hypergraph-Former architecture modeling high-order cross-modal temporal relationships; (3) event-level domain disentanglement with contrastive learning for improved generalization across behavioral contexts. Experiments on MPDD-Young dataset show P³HF achieves around 10% improvement on accuracy and weighted F1 for binary and ternary depression classification task over existing methods. Extensive ablation studies validate the independent contribution of each architectural component, confirming that personality-guided representation learning and high-order hypergraph reasoning are both essential for generating robust, individual-aware depression-related representations.
Changzeng Fu, Shiwen Zhao, Yunze Zhang, Zhongquan Jian, Shiqi Zhao 0001
AAAI1
2026 Explainable Depression Assessment from Face Videos by Weakly Supervised Learning
abstract
Existing video-based automatic depression assessment (ADA) approaches frequently achieve video-level depression assessment by aggregating features or predictions of individual frames or equal-length segments within the given video. While their performances have been largely enhanced by recent advanced deep learning models, they typically fail to explicitly consider the varied importance of depression-related behavioural cues across different video segments, i.e., segments within one video may contain behaviours reflecting varying levels of depression. Underestimating segment-level variations can obscure the detection of facial behaviour cues associated with depression, thereby undermining the accuracy and interpretability of video-based depression detection systems. In this paper, we propose a novel video-based ADA approach that specifically identifies and differentiates video segments that exhibit depression-related facial behaviours across varying temporal durations, providing clear insights into how each segment contributes to the video-level depression prediction. To achieve this, a novel weakly supervised strategy is proposed to compare segment-level behaviours with video-level depression label, enabling the model to assign depression-relevant scores to multiple temporal scale video segments and attend selectively to those most indicative of depressive states. Extensive experiments on the AVEC 2013 and AVEC 2014 face video depression datasets demonstrate the effectiveness of our approach.
Rongfan Liao, Xiangyu Kong 0001, Shiqing Tang, Changzeng Fu, Weicheng Xie 0001, Lu Liu 0001, Siyang Song
AAAI5
2026 A plug-and-play hybrid pruning framework for spike-driven transformers via adaptive spiking statistical importance scoring
Hanfei Liu, Shiqi Zhao 0001, Changzeng Fu, Jie Yang 0033, Mohamad Sawan
Neurocomputing3
2026 HAQ-ViT: A hardware-aware post-training quantization for efficient vision transformer inference
Shiqi Zhao 0001, Haozu Sun, Changzeng Fu, Jie Yang 0033, Mohamad Sawan
Knowl. Based Syst.6
2025 DepMGNN: Matrixial Graph Neural Network for Video-based Automatic Depression Assessment
abstract
Depression can be reflected by long-term human spatio-temporal facial behaviours. While human face videos recorded in real-world usually have long and variable lengths, existing video-based depression assessment approaches frequently re-sample/down-sample such videos to short and equal-length videos, or split each video into several equal-length segments, where segment-level spatio-temporal facial behaviours are suppressed as a vector-style representations for RNN-based long-term (video-level) modelling. Both strategies lead to crucial information loss and distortion. In this paper, we propose a novel graph-style data structure called Matrixial Graph and an effective Matrixial Graph Neural Network (MGNN) for face video-based depression assessment, which can directly and end-to-end model long-term depression-specific spatio-temporal facial cues from variable-length videos without resampling/splitting videos or suppressing video segments to vectors. Importantly, the nodes in our matrixial graph are capable of including matrices of different shapes, and thus nodes of a matrix graph can directly represent all frame-level 2D facial feature maps (or images themselves) of an entire video regardless of its length. Then, our MGNN is the first GNN that can jointly process matrixial graphs containing varying numbers of nodes, which further learns matrix-style edge features, thereby facilitating to explicit model video-level multi-scale spatio-temporal facial behaviours among matrixial graph nodes for depression assessment. Experiments show that the explicit spatio-temporal modeling on 2D facial feature maps, facilitated by our matrixial graph/MGNN, provided significant benefits, leading our approach to achieve new state-of-the-art performances on AVEC2013 and AVEC2014 datasets with large advantages.
Leijing Zhou, Shuanglin Li, Changzeng Fu, Jun Lu 0006, Jing Han 0009, Yi Zhang 0036, Siyang Song
AAAI4
2025 In-Context Multitask Learning for Few-shot Fine-tuning of Large Language Models in Traditional Chinese Medicine Tongue Diagnosis
abstract
Tongue diagnosis is integral to Traditional Chinese Medicine (TCM) for evaluating a patient’s body constitution. Yet, this field faces challenges such as indirect constitution diagnosis, a dearth of labeled datasets, and the complexities of few-shot learning. Existing studies focus mainly on analyzing tongue color and coating, rather than directly analyzing the body constitution from the patient’s tongue. Moreover, the lack of publicly available datasets with constitution labels impedes model development for constitution diagnosis. The resource-intensive process of creating large-scale labeled datasets calls for efficient training methods on small datasets. Addressing these issues, this paper presents an In-Context Multitask Learning approach to improve few-shot fine-tuning accuracy for Large Language Models (LLMs) in constitution diagnosis. We created a dataset for color-coating analysis and a few-shot dataset with constitution labels and a structured prompt-label framework, allowing LLMs to learn from varied datasets. Our method shows improved performance over traditional methods in constitution diagnosis and significantly boosts LLMs’ generalization and robustness in TCM’s complex diagnostic tasks, providing a viable path for automating tongue diagnosis.
Changzeng Fu, Zelin Fu, Shaojun Yan, Xiaoyong Lyu, Yuliang Zhao
ICASSP1
2025 M3ADD: A Novel Benchmark for Physiology Signal-based Automatic Depression Detection with Multimodal Multitask Multievent Framework
abstract
The prevalence of depression is escalating, especially among youth, which has become a critical mental health concern. Current assessment methods, relying heavily on questionnaires, clinical observations, and AI-driven analyses, are limited by their focus on single-event data, failing to encapsulate the nuanced expressions of depressive symptoms. Moreover, a significant oversight in existing research is the underutilization of electromyogram (EMG) alongside electroencephalogram (EEG) data, which could provide a more holistic view of unconscious body behaviors. Additionally, given the high variability of depression among individuals, traditional analysis models are in urgent need of refinement to accommodate the personality of different individuals. To address these limitations, we propose M3ADD, a novel benchmark for Automatic Depression Detection that employs a Multimodal, Multitask, and Multievent framework. We collected EEG and EMG data from 97 participants across varied events (interview, reading tasks, walking), coupled with standardized questionnaires assessing depression, wellbeing, and personality, enriching our multitask learning approach. Our benchmark recognition algorithm leverages multitask learning, channel and interactive attention mechanisms to synthesize event-specific and modal-specific features, enhancing adaptability to individual differences and improving data utilization efficiency. M3ADD surpasses existing models by achieving 87% accuracy in detecting depression and 95% accuracy in assessing wellbeing, providing a promising avenue for early identification.
Changzeng Fu, Kaifeng Su, Yikai Su, Fengkui Qian, Siyang Song, Le Yang 0004, Xiaoyong Lv, Yuliang Zhao
ICASSP1
2025 Hierarchical Similarity Loss Enhanced Depth and Structural Fidelity in Monocular RGB-to-Depth Mapping with Adversarial Training
abstract
The conversion of monocular RGB images to depth maps is crucial in robotic applications. Current supervised learning approaches, dependent on high-quality RGB-Depth pairs, struggle with indoor environments characterized by multiple objects and fluctuating lighting, leading to inaccurate and unstable depth estimations. Additionally, commercial depth cameras on robots produce incomplete and light-sensitive depth maps, further degrading estimation quality. To surmount these issues, We created a comprehensive dataset of 9,600 RGB-Depth image pairs, capturing a range of indoor scenes under various lighting (dim, normal, and strong lighting), and interference conditions (local strong light interference, specular reflection interference, background similarity interference, and combinations of these factors). This dataset serves as a foundation for our proposed monocular RGB-to-depth mapping framework, which employs a hierarchical similarity loss to enhance the model’s structural feature learning from reference depth maps, improving the fidelity of estimated depth maps. We also integrated adversarial training and attention mechanisms to refine the depth consistency between the estimated and original depth maps. Experiments show our method surpasses current benchmarks in metrics like Threshold Accuracy, SILog, and MSE, demonstrating its robustness and potential for real-world applications, even with poor-quality inputs.
Changzeng Fu, Yikai Su, Kaifeng Su, Le Yang 0004, Xiaoyong Lv, Yuliang Zhao
ICASSP1
2025 The First MPDD Challenge: Multimodal Personality-aware Depression Detection
abstract
Depression is a widespread mental health issue affecting diverse age groups, with notable prevalence among college students and the elderly. However, existing datasets and detection methods primarily focus on young adults, neglecting the broader age spectrum and individual differences that influence depression manifestation. Current approaches often establish a direct mapping between multimodal data and depression indicators, failing to capture the complexity and diversity of depression across individuals. This challenge includes two tracks based on age-specific subsets: Track 1 uses the MPDD-Elderly dataset for detecting depression in older adults, and Track 2 uses the MPDD-Young dataset for detecting depression in younger participants. The Multimodal Personality-aware Depression Detection (MPDD) Challenge aims to address this gap by incorporating multimodal data alongside individual difference factors. We provide a baseline model that fuses audio and video modalities with individual difference information to detect depression manifestations in diverse populations. This challenge aims to promote the development of more personalized and accurate de pression detection methods, advancing mental health research and fostering inclusive detection systems. More details are available on the official challenge website: https://hacilab.github.io/MPDDChallenge.github.io.
Changzeng Fu, Zelin Fu, Qi Zhang 0124, Xinhe Kuang, Jiacheng Dong, Kaifeng Su, Yikai Su, Junfeng Yao, Yuliang Zhao, Shiqi Zhao 0001, Siyang Song, Yuichiro Yoshikawa, Björn W. Schuller, Hiroshi Ishiguro
ACM Multimedia1
2025 HAM-GNN: A hierarchical attention-based multi-dimensional edge graph neural network for dialogue act classification
Changzeng Fu, Yikai Su, Kaifeng Su, Yinghao Liu, Bowen Wu 0002, Carlos Toshinori Ishi, Hiroshi Ishiguro
Expert Syst. Appl.1
2025 Facial action units guided graph representation learning for multimodal depression detection
Changzeng Fu, Fengkui Qian, Yikai Su, Kaifeng Su, Siyang Song, Mingyue Niu, Zhigang Liu 0014, Carlos Toshinori Ishi, Hiroshi Ishiguro
Neurocomputing1
2025 Disease and personality information enhanced depression detection based on the TransGCL framework
Yuliang Zhao, Jian Li 0063, Chao Lian, Kaixuan Tian, Changzeng Fu
Neurocomputing9
2025 Incorporating image representation and texture feature for sensor-based gymnastics activity recognition
Chao Lian, Yuliang Zhao, Tianang Sun, Jin-Liang Shao, Yinghao Liu, Changzeng Fu, Xiaoyong Lyu, Zhikun Zhan
Knowl. Based Syst.6
2025 HiMul-LGG: A hierarchical decision fusion-based local-global graph neural network for multimodal emotion recognition in conversation
Changzeng Fu, Fengkui Qian, Kaifeng Su, Yikai Su, Zhigang Liu 0014, Carlos Toshinori Ishi
Neural Networks1
2025 Image Encoding and Fusion of Multi-Modal Data Enhance Depression Diagnosis in Parkinson's Disease Patients
abstract
The diagnosis of depression in individuals with Parkinson's Disease (PD) through the utilization of multimodal fusion techniques represents a significant domain. The primary challenge involves the creation of a robust fusion framework to address the heterogeneity among different modalities effectively. However, previous studies primarily focused on interactions between heterogeneous data, neglecting the structural similarities among isomorphic data, resulting in a substantial loss of feature information when merging heterogeneous data. In this study, we introduced a multi-modal data image encoding and fusion approach for diagnosing depression in PD patients. Additionally, we proposed a multi-modal dataset encompassing motion, facial expression, and audio data. First, we designed an RGB and sparse coding method to encode the multi-modal data, achieving the isomorphic transformation of multi-modal information and extracting feature information from lower-dimensional spaces. Furthermore, we introduced a Spatial-Temporal Network (STN) to fuse the three types of encoded images. We incorporated the Relation Global Attention (RGA) to enhance feature extraction and leverage all encoded image location feature nodes for balanced decision attention. Finally, recognizing the limitations of traditional machine learning algorithms in handling multi-tasks in medical diagnosis, we established a multi-task weighted loss function to achieve depression identification and severity prediction through Multi-Task learning (MTL).
Jian Li 0063, Yuliang Zhao, Wayne Jason Li, Changzeng Fu, Chao Lian
IEEE Trans. Affect. Comput.5
2025 Multimodal Depression Assessment Framework Integrating Personality and Gait for Older Adults With Medical Conditions
abstract
Elderly individuals often suffer from underlying medical conditions, resulting in a significant decline in quality of life and a heightened susceptibility to depression. Presently, AI screening tools based on behavioral indicators offer an objective and effective approach to diagnosing depression. However, current AI depression screening tools are primarily tailored to adolescents and adults, exhibiting shortcomings in their applicability and accuracy for elderly individuals with underlying medical conditions. To address the above issues, first, this paper constructs a depression dataset for elderly people with underlying diseases by using semi-structured interviews. Second, based on cognitive science insights, it is recognized that personality factors significantly influence behavioral expressions and also determine the attitudes of elderly individuals toward current life circumstances/health issues. Therefore, besides annotating depression severity, the Big Five-10 personality scale was utilized to annotate participant personalities. Finally, a late fusion-based multi-task learning framework was proposed, and the effects of introducing gait information and personality annotation on the performance of depression assessment were investigated. The experimental findings affirm the importance of integrating gait information and personality assessment in improving depression detection effectiveness. This study provides valuable foundational resources, as well as beneficial references and insights, for the research on depression in the elderly.
Yuliang Zhao, Jian Li 0063, Siyang Song, Chao Lian, Yinghao Liu, Changzeng Fu
IEEE Trans. Affect. Comput.8
2025 BoostViT: Booth-Serial Skipping and Tunable Scaling for Vision Transformers
abstract
Vision Transformers (ViTs) have emerged as a dominant architecture in computer vision (CV), surpassing conventional neural network counterparts across diverse visual tasks. Despite their exceptional performance, ViTs incur substantial computational overhead characterized by high memory footprint, long inference latency, and elevated energy consumption. Current acceleration strategies for ViTs primarily focus on pruning operations or leveraging the inherent sparsity, requiring complex address control logic or position encoding. Alternatively, some software-based approaches attempt to pre-compute and separate dense and sparse matrix position encoding, while hardware solutions typically spend additional time and resources to obtain position encoding, decomposing matrix multiplications into structured forms. Through an analysis of ViTs’ parameters, we found that approximately 91.03% of the most significant bits (MSBs) are either 111s or 000s, and nearly 45% of 3 adjacent bits are identical. To leverage this characteristic of ViTs, we propose the Booth-Serial Skipping algorithm, which transforms the computation of consecutive 111 or 000 sequences into skip steps that require no additional computation time. Furthermore, the 4th to 6th bits of ViT weights can undergo aggressive scaling, enhancing the likelihood of Booth-skip operations with minimal impact on accuracy. The key innovation of this paper lies in exploiting the high proportion of naturally consecutive 0s or 1s in 8-bit weights during ViT inference and further expanding the skippable range through the Tunable Scaling strategy. At the hardware level, we develop a specialized accelerator to coordinate the proposed acceleration strategies. The processing element array in the accelerator is optimized for general matrix multiplication, it not only significantly improves the computation of multi-head self-attention but also enables resource reuse for linear transformations, ultimately optimizing end-to-end inference. Our design achieves$50.3\times $,$21.9\times $,$17.37\times $,$7.47\times $, and$1.49\times $an average end-to-end speedup on DeiT over CPU (Intel Xeon Gold 6152), EdgeGPU (NVIDIA Jetson Xavier NX), GPU (TITAN Xp), ViTCoD, and ViT-slice, respectively.
Shiqi Zhao 0001, Chaoming Fang, Fengshi Tian, Jinbo Chen 0002, Changzeng Fu, Jie Yang 0033, Mohamad Sawan
IEEE Trans. Circuits Syst. I Regul. Pap.7
2025 Predictive Body Awareness in Soft Robots: A Bayesian Variational Autoencoder Fusing Multimodal Sensory Data
abstract
Predicting the causal flow by fusing multimodal perception is fundamental for constructing the bodily awareness of soft robots. However, forming such a predictive model while fusing the multimodal sensory data of soft robots remains challenging and less explored. In this study, we leverage the free energy principle within a Bayesian probabilistic deep learning framework to merge visual, pressure, and flex sensing signals. Our proposed multimodal association mechanism enhances the fusion process, establishing a robust computational methodology. We train the model using a newly collected dataset that captures the grasping dynamics of a soft gripper equipped with multimodal perception capabilities. By incorporating the current state and image differences, the forward model can predict the soft gripper's physical interaction and movement in the image flow, which amounts to imagining future motion events. Moreover, we showcase effective predictions across modalities as well as for grasping outcomes. Notably, our enhanced variational autoencoder approach can pave the way for unprecedented possibilities of bodily awareness in soft robotics.
Dongling Liu, Changzeng Fu, Xiaoming Yuan 0002, Victor C. M. Leung
IEEE Trans. Robotics3
2024 Global joint information extraction convolution neural network for Parkinson's disease diagnosis
Yuliang Zhao, Yinghao Liu, Jian Li 0063, Xiaoai Wang, Ruige Yang, Chao Lian, Zhikun Zhan, Changzeng Fu
Expert Syst. Appl.10
2024 Loss Relaxation Strategy for Noisy Facial Video-based Automatic Depression Recognition
abstract
Automatic depression analysis has been widely investigated on face videos that have been carefully collected and annotated in lab conditions. However, videos collected under real-world conditions may suffer from various types of noise due to challenging data acquisition conditions and lack of annotators. Although deep learning (DL) models frequently show excellent depression analysis performances on datasets collected in controlled lab conditions, such noise may degrade their generalization abilities for real-world depression analysis tasks. In this article, we uncovered that noisy facial data and annotations consistently change the distribution of training losses for facial depression DL models; i.e., noisy data–label pairs cause larger loss values compared to clean data–label pairs. Since different loss functions could be applied depending on the employed model and task, we propose a generic loss function relaxation strategy that can jointly reduce the negative impact of various noisy data and annotation problems occurring in both classification and regression loss functions for face video-based depression analysis, where the parameters of the proposed strategy can be automatically adapted during depression model training. The experimental results on 25 different artificially created noisy depression conditions (i.e., five noise types with five different noise levels) show that our loss relaxation strategy can clearly enhance both classification and regression loss functions, enabling the generation of superior face video-based depression analysis models under almost all noisy conditions. Our approach is robust to its main variable settings and can adaptively and automatically obtain its parameters during training.
Siyang Song, Tugba Tümer, Changzeng Fu, Michel F. Valstar, Hatice Gunes
ACM Trans. Comput. Heal.4
2024 PointTransform Networks for automatic depression level prediction via facial keypoints
Mingyue Niu, Ming Li 0065, Changzeng Fu
Knowl. Based Syst.3
2023 HAG: Hierarchical Attention with Graph Network for Dialogue Act Classification in Conversation
abstract
The prediction of dialogue acts (DA) labels on utterance-level in conversations can be treated as a sequence labeling problem, which requires context- and speaker-aware semantic comprehension, especially for Japanese. In this study, we pro-posed a hierarchical attention with the graph neural network (HAG) to consider the contextual interconnections as well as the semantics carried by the sentence itself. Concretely, the model use long-short term memory networks (LSTMs) to perform a context-aware encoding within a dialogue window. Then, we construct the context graph by aggregating the neighboring utterances. Subsequently, a speaker feature transformation is executed with a graph attention network (GAT) to calculate the interconnections, while a context-level feature selection is performed with a gated graph convolutional network (GatedGCN) to select the salient utterances that contribute to the DA classification. Finally, we merge the representations of different levels and conduct a classification with two dense layers. We evaluate the proposed model on Japanese dialogue act dataset (JPS-DA). The experimental results show that our method outperforms the baselines.
Changzeng Fu, Zhenghan Chen, Bowen Wu 0002, Carlos Toshinori Ishi, Hiroshi Ishiguro
ICASSP1
2023 LGFat-RGCN: Faster Attention with Heterogeneous RGCN for Medical ICD Coding Generation
abstract
With the increasing volume of healthcare data, automated International Classification of Diseases (ICD) has become increasingly relevant and is frequently regarded as a medical multi-label prediction problem. Current methods struggle to accurately classify medical diagnosis texts that represent deep and sparse categories. Unlike these works that model the label with code hierarchy or description for label prediction, we argue that the label generation with structural information can provide more comprehensive knowledge based on the observation that label synonyms and parent-child relationships in vary from their context in clinical contexts. In this study, we introduce \tool, a heterogeneous graph model with improved attention for automated ICD coding. Notably, our approach represents the model to consider this task as a labelled graph generation problem. Our enhanced attention mechanism boosts the model's capacity to learn from multi-relational heterogeneous graph representations. Additionally, we propose a discriminator for labelled graphs (LG) that computes the reward for each ICD code in the labelled graph generator. Our experimental findings demonstrate that our proposed model significantly outperforms all existing strong baseline methods and attains the best performance on three benchmark datasets.
Zhenghan Chen, Changzeng Fu, Ruoxue Wu, Ye Wang 0023, Xunzhu Tang, Xiaoxuan Liang 0002
ACM Multimedia2
2023 An Adversarial Training Based Speech Emotion Classifier With Isolated Gaussian Regularization
abstract
Speaker individual bias may cause emotion-related features to form clusters with irregular borders (non-Gaussian distributions), making the model sensitive to local irregularities of pattern distributions, resulting in the model over-fit of the in-domain dataset. This problem may cause a decrease in the validation scores in cross-domain (i.e., speaker-independent, channel-variant) implementation. To mitigate this problem, in this paper, we propose an adversarial training-based classifier to regularize the distribution of latent representations to further smooth the boundaries among different categories. In the regularization phase, the representations are mapped into Gaussian distributions in an unsupervised manner to improve the discriminative ability of the latent representations. A single Gaussian distribution is used for mapping the latent representations in our previous study. In this presented work, we adopt a mixture of isolated Gaussian distributions. Moreover, multi-instance learning was adopted by dividing speech into a bag of segments to capture the most salient part of presenting an emotion. The model was evaluated on the IEMOCAP and MELD datasets with in-corpus speaker-independent sittings. In addition, we investigated the accuracy of cross-corpus sittings in simulating speaker-independent and channel-variants. In the experiment, the proposed model was compared not only with baseline models but also with different configurations of our model. The results show that the proposed model is competitive with respect to the baseline, as demonstrated both by in-corpus and cross-corpus validation.
Changzeng Fu, Carlos Toshinori Ishi, Hiroshi Ishiguro
IEEE Trans. Affect. Comput.1
2022 First Attempt of Gender-free Speech Style Transfer for Genderless Robot
abstract
Some robots for human-robot interaction are designed with female or male physical appearance. Other robots are endowed with no gender characteristics, namely genderless robots, such as Pepper and NAO robot. A robot with male or female physical appearance should possess the mapped speech gender style during a natural human-robot interaction, which can be learned from humans' male or female speech. In this paper, we make a new trial to synthesis gender-free speeches for physically genderless robots, which is promising in order to improve a more natural human-robot interaction with genderless robots. Our gender style-controlled speech synthesizer takes the speech text and gender style embedding as inputs to generate speech audio. A speech gender encoder network is used to extract the embedding of the speech gender style with female and male speeches as input. Based on the distribution of the female and male gender style embedding, we explore the gender-free speech style embedding space where we sample some gender-free embedding vectors to generate genderless speech audio. This is a preliminary work where we show how the genderless speech audio wave will be synthesized from text.
Chuang Yu 0001, Changzeng Fu, Adriana Tapus
HRI2
2022 An improved CycleGAN-based emotional voice conversion model by augmenting temporal dependency with a transformer
Changzeng Fu, Carlos Toshinori Ishi, Hiroshi Ishiguro
Speech Commun.1
2021 MAEC: Multi-Instance Learning with an Adversarial Auto-Encoder-Based Classifier for Speech Emotion Recognition
abstract
In this paper, we propose an adversarial auto-encoder-based classifier, which can regularize the distribution of latent representation to smooth the boundaries among categories. Moreover, we adopt multi-instance learning by dividing speech into a bag of segments to capture the most salient moments for presenting an emotion. The proposed model was trained on the IEMOCAP dataset and evaluated on the in-corpus validation set (IEMOCAP) and the cross-corpus validation set (MELD). The experiment results show that our model outperforms the baseline on in-corpus validation and increases the scores on cross-corpus validation with regularization.
Changzeng Fu, Carlos Toshinori Ishi, Hiroshi Ishiguro
ICASSP1