EDBT 2026 Demo / reviewers in the wild / expert
Jiaxing Liu 0001
dblp:220/1068-1 · also Jia-Xing Liu 0001
· DBLP profile ↗
14ranked-venue papers
6as first author
10since 2021 · last 2024
0000-0001-9691-8470ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Adversarial Domain Generalized Transformer for Cross-Corpus Speech Emotion RecognitionabstractSpeech emotion recognition (SER) promotes the development of intelligent devices, which enable natural and friendly human-computer interactions. However, the recognition performance of existing approaches is significantly reduced on unseen datasets, and the lack of sufficient training data limits the generalizability of deep learning models. In this work, we analyze the impact of the domain generalization method on cross-corpus SER and propose an adversarial domain generalized transformer (ADoGT), which is aimed at learning a shared feature distribution for the source and target domains. Specifically, we investigate the effect of domain adversarial learning by eliminating nonaffective information. We also combine the center loss with the softmax function as joint supervision to learn discriminative features. Moreover, we introduce unsupervised transfer learning to extract additional features, and incorporate a gated fusion model to learn the complementary information of the features learned by the supervised feature extractor and pretrained model. The proposed transformer based domain generalization method is evaluated using four emotional datasets. We also provide an ablation study of different domain adversarial model structures and feature fusion models. The results of comparative experiments demonstrate the effectiveness of the proposed ADoGT. Yuan Gao 0040, Longbiao Wang, Jiaxing Liu 0001, Jianwu Dang 0001, Shogo Okada |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | Domain-Invariant Feature Learning for Cross Corpus Speech Emotion RecognitionabstractTo deal with speech emotion recognition (SER) in real-life applications, researchers have to focus on cross corpus SER, where the feature distribution of source and target datasets are different. In this paper, we propose an efficient domain adversarial training method to cope with the non-affective information during feature extraction. Through the proposed domain-adversarial learning, we can reduce the domain divergence between train and test data. Furthermore, we incorporate center loss with the emotion classifier to reduce the intra-class variation of features learned from the same emotion. We conduct experiments on four emotional benchmark datasets to verify the performance of the proposed method. The experimental results demonstrate that our proposed model outperform the baseline system in both cross-corpus and multi-corpus evaluation. Yuan Gao 0040, Shogo Okada, Longbiao Wang, Jiaxing Liu 0001, Jianwu Dang 0001 |
ICASSP | 4 |
| 2022 | Multi-Stage Graph Representation Learning for Dialogue-Level Speech Emotion RecognitionabstractWith the development of speech emotion recognition (SER), most of current research is utterance-level and cannot fit the need of actual scenarios. In this paper, we propose a novel strategy that focuses on capturing dialogue-level contextual information. On the basis of utterance-level representation learned by convolutional neural network (CNN) which is followed by the bidirectional long short-term memory network (BLSTM), the proposed dialogue-level method consists of two modules. The first module is Dialogue Multi-stage Graph Representation Learning Algorithm (DialogMSG). The multi-stage graph that modeling from different dialogue scope is introduced to capture more effective information. The other one is a double-constrained module. This module includes not only an utterance-level classifier but also a dialogue-level graph classifier which is named as Atmosphere. The results of extensive experiments show that the proposed method outperforms the current state of the art on the IEMOCAP benchmark dataset. Yaodong Song, Jiaxing Liu 0001, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 2 |
| 2021 | Talking Head Generation with Audio and Speech Related Facial Action Units
Zhilei Liu, Jiaxing Liu 0001, Zhengxiang Yan, Longbiao Wang |
BMVC | 3 |
| 2021 | Domain-Adversarial Autoencoder with Attention Based Feature Level Fusion for Speech Emotion RecognitionabstractOver the past two decades, although speech emotion recognition (SER) has garnered considerable attention, the problem of insufficient training data has been unresolved. A potential solution for this problem is to pre-train a model and transfer knowledge from large amounts of audio data. However, the data used for pre-training and testing originate from different domains, resulting in the latent representations to contain non-affective information. In this paper, we propose a domain-adversarial autoencoder to extract discriminative representations for SER. Through domain-adversarial learning, we can reduce the mismatch between domains while retaining discriminative information for emotion recognition. We also introduce multi-head attention to capture emotion information from different subspaces of input utterances. Experiments on IEMOCAP show that the proposed model outperforms the state-of-the-art systems by improving the unweighted accuracy by 4.15%, thereby demonstrating the effectiveness of the proposed model. Yuan Gao 0040, Jiaxing Liu 0001, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 2 |
| 2021 | Multimodal Emotion Recognition with Capsule Graph Convolutional Based Representation FusionabstractDue to the more robust characteristics compared to unimodal, audio-video multimodal emotion recognition (MER) has attracted a lot of attention. The efficiency of representation fusion algorithm often determines the performance of MER. Although there are many fusion algorithms, information redundancy and information complementarity are usually ignored. In this paper, we propose a novel representation fusion method, Capsule Graph Convolutional Network (CapsGCN). Firstly, after unimodal representation learning, the extracted audio and video representations are distilled by capsule network and encapsulated into multimodal capsules respectively. Multimodal capsules can effectively reduce data redundancy by the dynamic routing algorithm. Secondly, the multimodal capsules with their inter-relations and intra-relations are treated as a graph structure. The graph structure is learned by Graph Convolutional Network (GCN) to get hidden representation which is a good supplement for information complementarity. Finally, the multimodal capsules and hidden relational representation learned by CapsGCN are fed to multihead self-attention to balance the contributions of source representation and relational representation. To verify the performance, visualization of representation, the results of commonly used fusion methods, and ablation studies of the proposed CapsGCN are provided. Our proposed fusion method achieves 80.83% accuracy and 80.23% F1 score on eNTERFACE05’. Jiaxing Liu 0001, Longbiao Wang, Zhilei Liu, Yahui Fu 0001, Lili Guo 0001, Jianwu Dang 0001 |
ICASSP | 1 |
| 2021 | CONSK-GCN: Conversational Semantic- and Knowledge-Oriented Graph Convolutional Network for Multimodal Emotion RecognitionabstractEmotion recognition in conversations (ERC) has received significant attention in recent years due to its widespread applications in diverse areas, such as social media, health care, and artificial intelligence interactions. However, different from nonconversational text, it is particularly challenging to model the effective context-aware dependence for the task of ERC. To address this problem, we propose a new Conversational Semantic- and Knowledge-oriented Graph Convolutional Network (ConSK-GCN) approach that leverages both semantic dependence and commonsense knowledge. First, we construct the contextual inter-interaction and intradependence of the interlocutors via a conversational graph-based convolutional network based on multimodal representations. Second, we incorporate commonsense knowledge to guide ConSK-GCN to model the semantic-sensitive and knowledge-sensitive contextual dependence. The results of extensive experiments show that the proposed method outperforms the current state of the art on the IEMOCAP dataset. Yahui Fu 0001, Shogo Okada, Longbiao Wang, Lili Guo 0001, Yaodong Song, Jiaxing Liu 0001, Jianwu Dang 0001 |
ICME | 6 |
| 2021 | Metric Learning Based Feature Representation with Gated Fusion Model for Speech Emotion Recognition
Yuan Gao 0040, Jiaxing Liu 0001, Longbiao Wang, Jianwu Dang 0001 |
Interspeech | 2 |
| 2021 | Time-Frequency Representation Learning with Graph Convolutional Network for Dialogue-Level Speech Emotion Recognition
Jiaxing Liu 0001, Yaodong Song, Longbiao Wang, Jianwu Dang 0001 |
Interspeech | 1 |
| 2021 | A Sentiment Similarity-Oriented Attention Model with Multi-task Learning for Text-Based Emotion Recognition
Yahui Fu 0001, Lili Guo 0001, Longbiao Wang, Zhilei Liu, Jiaxing Liu 0001, Jianwu Dang 0001 |
MMM (1) | 5 |
| 2020 | Speech Emotion Recognition with Local-Global Aware Deep Representation LearningabstractConvolutional neural network (CNN) based deep representation learning methods for speech emotion recognition (SER) have demonstrated great success. The basic design of CNN restricts the ability to model only local information well. Capsule network (CapsNet) can overcome the shortages of CNNs to capture the shallow global features from the spectrogram, although CapsNet cannot learn the local and deep global information. In this paper, we propose a local-global aware deep representation learning system that mainly includes two modules. One module contains a multi-scale CNN, time- frequency CNN (TFCNN) to learn the local representation. In the other module, we introduce a structure with dense connections of multiple blocks to learn shallow and deep global information. Every block in this structure is a complete CapsNet improved by a new routing algorithm. The local and global representations are fed to the classifier and achieve an absolute increase of at least 4.25% than benchmarks on IEMOCAP. Jiaxing Liu 0001, Zhilei Liu, Longbiao Wang, Lili Guo 0001, Jianwu Dang 0001 |
ICASSP | 1 |
| 2020 | Temporal Attention Convolutional Network for Speech Emotion Recognition with Latent Representation
Jiaxing Liu 0001, Zhilei Liu, Longbiao Wang, Yuan Gao 0040, Lili Guo 0001, Jianwu Dang 0001 |
INTERSPEECH | 1 |
| 2019 | Time-Frequency Deep Representation Learning for Speech Emotion Recognition Integrating Self-attention
Jiaxing Liu 0001, Zhilei Liu, Longbiao Wang, Lili Guo 0001, Jianwu Dang 0001 |
ICONIP (4) | 1 |
| 2017 | A robust approach of watermarking in contourlet domain based on probabilistic neural network
Jiaxing Liu 0001, Xianbin Wen, Haixia Xu 0003 |
Multim. Tools Appl. | 1 |