Shaokai Li

dblp:301/5544 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-1684-043XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Enhancing Emotion Recognition in Incomplete Data: A Novel Cross-Modal Alignment, Reconstruction, and Refinement Framework
abstract
Multimodal emotion recognition systems rely heavily on the full availability of modalities, suffering significant performance declines when modal data is incomplete. To tackle this issue, we present the Cross-Modal Alignment, Reconstruction, and Refinement (CM-ARR) framework, an innovative approach that sequentially engages in cross-modal alignment, reconstruction, and refinement phases to handle missing modalities and enhance emotion recognition. This framework utilizes unsupervised distribution-based contrastive learning to align heterogeneous modal distributions, reducing discrepancies and modeling semantic uncertainty effectively. The reconstruction phase applies normalizing flow models to transform these aligned distributions and recover missing modalities. The refinement phase employs supervised point-based contrastive learning to disrupt semantic correlations and accentuate emotional traits, thereby enriching the affective content of the reconstructed representations. Extensive experiments confirm the superior performance of CM-ARR. Notably, averaged across six scenarios of missing modalities, CM-ARR achieves absolute improvements of 2.11%/2.12% (WAR/UAR), and 1.71%/1.96% (WAR/UAR), respectively, on IEMOCAP and MSP-IMPROV datasets.
Haoqin Sun, Shiwan Zhao, Shaokai Li, Xiangyu Kong 0001, Xuechen Wang, Jiaming Zhou 0001, Aobo Kong, Wenjia Zeng
ICASSP3
2024 Multi-Source Unsupervised Transfer Components Learning for Cross-Domain Speech Emotion Recognition
abstract
As an important research direction in the field of speech signal processing, cross-domain speech emotion recognition (SER) has attracted extensive attention. In practice, it is challenging to collect enough labeled samples from single source domain to train robust classifiers. To this end, this paper presents a novel method named multi-source unsupervised transfer components learning (MUTCL) for cross-domain SER. In MUTCL, we first adopt a PCA-like strategy and apply it to multi-source domains, aiming to preserve both intra-domain individuality and inter-domain commonality principal components within each domain. Simultaneously, a simple alignment strategy is developed to guide cross-domain samples to have similar structures, thus preserving more transfer components. Moreover, an adaptive weight strategy is utilized to determine the contribution of each source domain. We conduct experiments on five benchmark datasets, and the results show that MUTCL achieves excellent performance compared with some state-of-the-art methods.
Shenjie Jiang, Peng Song 0002, Shaokai Li, Wenming Zheng
ICASSP3
2024 Common Latent Embedding Space for Cross-Domain Facial Expression Recognition
abstract
In practical facial expression recognition (FER), the training data and test data are often obtained from different domains. It is obvious that the domain disparity could significantly degrade the recognition performance. To tackle this challenging cross-domain FER problem, we put forward a novel method termed common latent embedding space (CLES). To be specific, first, we obtain a common embedding space for cross-domain samples by matrix factorization (MF). Then, the dual-graph Laplacian is applied to this common embedding space to narrow the gap across distinct domains and, meanwhile, explores the inherent geometric information. Furthermore, to characterize the global relationship of the cross-domain samples, the self-representation strategy is used to guide the learning of the common embedding space. Finally, comprehensive experiments on four benchmark databases indicate that the proposed method can achieve better performance in comparison with the state-of-the-art methods on cross-domain FER tasks.
Peng Song 0002, Shaokai Li, Wenming Zheng
IEEE Trans. Comput. Soc. Syst.3
2023 A Generalized Subspace Distribution Adaptation Framework for Cross-Corpus Speech Emotion Recognition
abstract
In this paper, we propose a novel transfer learning framework, named generalized subspace distribution adaptation (GSDA), to tackle the challenging cross-corpus speech emotion recognition problem. First, we learn a common low-dimensional feature subspace by utilizing a generalized subspace learning method. Second, we develop a novel distance metric to reduce the divergence between the source and target corpora, which can efficiently explore the similarity and dissimilarity information in the process of knowledge transfer. Third, to demonstrate the effectiveness of our framework, we apply GSDA to the traditional subspace learning algorithms. Finally, we conduct extensive experiments by using the low-level features and deep features on three popular emotional databases, i.e., Berlin, IEMOCAP, and CVE. The results demonstrate that the proposed framework can achieve better performance than several state-of-the-art transfer learning approaches.
Shaokai Li, Peng Song 0002, Yun Jin, Wenming Zheng
ICASSP1
2023 Unsupervised Transfer Components Learning for Cross-Domain Speech Emotion Recognition
Shenjie Jiang, Peng Song 0002, Shaokai Li, Keke Zhao, Wenming Zheng
INTERSPEECH3
2023 Joint Instance Reconstruction and Feature Subspace Alignment for Cross-Domain Speech Emotion Recognition
Keke Zhao, Peng Song 0002, Shaokai Li, Wenming Zheng
INTERSPEECH3
2023 Dual-graph regularized concept factorization for multi-view clustering
Jinshuai Mu, Peng Song 0002, Shaokai Li
Expert Syst. Appl.4
2023 Multi-Source Discriminant Subspace Alignment for Cross-Domain Speech Emotion Recognition
abstract
Cross-domain speech emotion recognition (SER) is an effective strategy to improve the generalization ability of emotion classification models, which is an important research direction in speech signal processing. However, since the speech signals are non-stationary, it is difficult to train a robust classifier from single-source emotional corpus. To solve this shortcoming, we propose a novel method named multi-source discriminant subspace alignment (MDSA) for cross-domain SER. In MDSA, we first conduct linear discriminant analysis (LDA) in the multi-source domain. Then, the instances in the multi-source discriminant subspace are used to linearly reconstruct the instances in the target subspace. At the same time, the reconstruction contribution of each source discriminant subspace is determined by adaptive weights. Furthermore, the multi-source discriminant subspace is aligned by reducing the loss between projections, which can make our model more robust. In this way, MDSA considers both the alignment of cross-domain data distribution and the structural information of cross-domain instances. Finally, extensive experiments are conducted on five standard emotional corpora, i.e., Berlin, IEMOCAP, CVE, EMOVO, and TESS, and the results demonstrate the proposed MDSA is superior to several state-of-the-art transfer learning algorithms in terms of performance. The codes are available athttps://github.com/shaokai1209/MDSA.
Shaokai Li, Peng Song 0002, Wenming Zheng
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 Coupled Discriminant Subspace Alignment for Cross-database Speech Emotion Recognition
Shaokai Li, Peng Song 0002, Keke Zhao, Wenjing Zhang 0003, Wenming Zheng
INTERSPEECH1