VLDB 2026 Research / reviewers in the wild / expert
Yan Zhao 0037
dblp:88/5320-37
· DBLP profile ↗
22ranked-venue papers
5as first author
22since 2021 · last 2026
0000-0003-4577-7078ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Two-stream attentive spatial-temporal graph convolutional network for P300 detection in brain-computer interface
Jincen Wang, Yan Zhao 0037, Cunhang Fan, Yong Li 0032, Fan Liu 0003, Hailun Lian, Cheng Lu 0005 |
Expert Syst. Appl. | 2 |
| 2026 | Weighted Macro-Expression Guided Transitional Cross-Domain Learning for Micro-Expression RecognitionabstractDespite progress in micro-expression recognition (MER), the low intensity of micro-expressions remains a key challenge, hindering effective learning of discriminative features. Given that macro-expressions share emotional components with micro-expressions and offer stronger discriminative cues, we propose a method that leverages macro-expression knowledge to enhance MER. A generative model is designed to synthesize transitional-expression features based on weighted representations, enabling one-to-many guidance from macro-expressions to micro-expressions. To align these transitional features with micro-expression semantics, we apply consistency regularization using micro-expression labels. Finally, feature alignment is conducted to increase inter-class separability and reduce intra-class variation. Experiments demonstrate that our method significantly improves MER performance. Yan Zhao 0037, Dong Wei 0007 |
IEEE Signal Process. Lett. | 2 |
| 2026 | An Interpretable Collaborative Acoustic Parameter Modeling Network for Speech Emotion RecognitionabstractAcoustic parameters can collaboratively represent speech emotions. However, existing speech emotion recognition (SER) methods neglect explicit collaborative modeling among acoustic parameters during representation learning. This not only limits their ability to extract the information contained in acoustic parameters that collaboratively represent emotions, but also makes it difficult to trace back which acoustic parameters are closely associated with emotional states (i.e., lack of interpretability). To address these challenges, we propose the interpretable collaborative acoustic parameter modeling network (ICAPM-Net), which innovatively introduces Graph Convolutional Networks (GCN) to model the relationships among acoustic parameters. Specifically, ICAPM-Net represents different acoustic parameters as nodes, with edges characterizing their collaborative relationships. The GCN utilizes edges to explicitly model the collaborative dependencies between acoustic parameters. Importantly, visualizing the edges in the graph (i.e., adjacency matrix) can trace which acoustic parameters (or their combinations) play dominant roles in emotional expression. The overall architecture of ICAPM-Net consists of three modules: 1) an encoding module, which is used to unify feature dimensions of acoustic parameters; 2) a graph-based collaborative modeling module (GBCM), which is responsible for modeling the collaborative relationships among acoustic parameters; and 3) an emotion adjacency matrix alignment loss module (EAMAL), which is designed to further enhance collaborative modeling in GBCM. Experiments on IEMOCAP, ABC, and EMO-DB show that ICAPM-Net outperforms state-of-the-art methods. Additionally, visualization results reveal that certain acoustic parameters, e.g., MFCC and RASTA, have a strong correlation with emotions and exhibit significant collaborative patterns in different emotional contexts. Hailun Lian, Cheng Lu 0005, Hao Yang 0028, Yan Zhao 0037, Sunan Li, Yuan Zong |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2025 | Low-rank joint distribution adaptation for cross-corpus speech emotion recognition
Sunan Li, Cheng Lu 0005, Yan Zhao 0037, Hailun Lian, Tianhua Qi, Yuan Zong |
Knowl. Based Syst. | 3 |
| 2025 | AMGCN: An adaptive multi-graph convolutional network for speech emotion recognition
Hailun Lian, Cheng Lu 0005, Hongli Chang, Yan Zhao 0037, Sunan Li, Yang Li 0019, Yuan Zong |
Speech Commun. | 4 |
| 2025 | Towards Domain-Specific Cross-Corpus Speech Emotion Recognition ApproachabstractCross-corpus speech emotion recognition (SER) poses a challenge due to feature distribution mismatch between the training and testing speech samples, potentially degrading the performance of established SER methods. In this article, we tackle this challenge by proposing a novel transfer subspace learning method called acoustic knowledge-guided transfer linear regression (AKTLR). Unlike existing approaches, which often overlook domain-specific knowledge related to SER and simply treat cross-corpus SER as a generic transfer learning task, our AKTLR method is built upon a well-designed acoustic knowledge-guided dual sparsity constraint mechanism. This mechanism emphasizes the potential of minimalistic acoustic parameter feature sets to alleviate classifier over-adaptation, which is empirically validated acoustic knowledge in SER, enabling superior generalization in cross-corpus SER tasks compared to using large feature sets. Through this mechanism, we extend a simple transfer linear regression model to AKTLR. This extension harnesses its full capability to seek emotion-discriminative and corpus-invariant features from established acoustic parameter feature sets used for describing speech signals across two scales: contributive acoustic parameter groups and constituent elements within each contributive group. We evaluate our method through extensive cross-corpus SER experiments on three widely used speech emotion corpora: EmoDB, eNTERFACE, and CASIA. The proposed AKTLR achieves an average UAR of 42.12% across six tasks using the eGeMAPS feature set, outperforming many recent state-of-the-art transfer subspace learning and deep transfer learning methods. This demonstrates the effectiveness and superior performance of our approach. Furthermore, our work provides experimental evidence supporting the feasibility and superiority of incorporating domain-specific knowledge into the transfer learning model to address cross-corpus SER tasks. Yan Zhao 0037, Yuan Zong, Hailun Lian, Cheng Lu 0005, Jingang Shi, Wenming Zheng |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2024 | Improving Speaker-Independent Speech Emotion Recognition using Dynamic Joint Distribution AdaptationabstractIn speaker-independent speech emotion recognition, the training and testing samples are collected from diverse speakers, leading to a multi-domain shift challenge across the feature distributions of data from different speakers. Consequently, when the trained model is confronted with data from new speakers, its performance tends to degrade. To address the issue, we propose a Dynamic Joint Distribution Adaptation (DJDA) method under the framework of multi-source domain adaptation. DJDA firstly utilizes joint distribution adaptation (JDA), involving marginal distribution adaptation (MDA) and conditional distribution adaptation (CDA), to more precisely measure the multi-domain distribution shifts caused by different speakers. This helps eliminate speaker bias in emotion features, allowing for learning discriminative and speaker-invariant speech emotion features from coarse-level to fine-level. Furthermore, we quantify the adaptation contributions of MDA and CDA within JDA by using a dynamic balance factor based on $\mathcal{A}$-Distance, promoting to effectively handle the unknown distributions encountered in data from new speakers. Experimental results demonstrate the superior performance of our DJDA as compared to other state-of-the-art (SOTA) methods. Cheng Lu 0005, Yuan Zong, Hailun Lian, Yan Zhao 0037, Björn W. Schuller, Wenming Zheng |
ICASSP | 4 |
| 2024 | Speech Swin-Transformer: Exploring a Hierarchical Transformer with Shifted Windows for Speech Emotion RecognitionabstractSwin-Transformer has demonstrated remarkable success in computer vision by leveraging its hierarchical feature representation based on Transformer. In speech signals, emotional information is distributed across different scales of speech features, e. g., word, phrase, and utterance. Drawing above inspiration, this paper presents a hierarchical speech Transformer with shifted windows to aggregate multi-scale emotion features for speech emotion recognition (SER), called Speech Swin-Transformer. Specifically, we first divide the speech spectrogram into segment-level patches in the time domain, composed of multiple frame patches. These segment-level patches are then encoded using a stack of Swin blocks, in which a local window Transformer is utilized to explore local inter-frame emotional information across frame patches of each segment patch. After that, we also design a shifted window Transformer to compensate for patch correlations near the boundaries of segment patches. Finally, we employ a patch merging operation to aggregate segment-level emotional features for hierarchical speech representation by expanding the receptive field of Transformer from frame-level to segment-level. Experimental results demonstrate that our proposed Speech Swin-Transformer outperforms the state-of-the-art methods. Yong Wang 0073, Cheng Lu 0005, Hailun Lian, Yan Zhao 0037, Björn W. Schuller, Yuan Zong, Wenming Zheng |
ICASSP | 4 |
| 2024 | Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion RecognitionabstractCross-corpus speech emotion recognition (SER) aims to transfer emotional knowledge from a labeled source corpus to an unlabeled corpus. However, prior methods require access to source data during adaptation, which is unattainable in real-life scenarios due to data privacy protection concerns. This paper tackles a more practical task, namely source-free cross-corpus SER, where a pre-trained source model is adapted to the target domain without access to source data. To address the problem, we propose a novel method called emotion-aware contrastive adaptation network (ECAN). The core idea is to capture local neighborhood information between samples while considering the global class-level adaptation. Specifically, we propose a nearest neighbor contrastive learning to promote local emotion consistency among features of highly similar samples. Furthermore, relying solely on nearest neighborhoods may lead to ambiguous boundaries between clusters. Thus, we incorporate supervised contrastive learning to encourage greater separation between clusters representing different emotions, thereby facilitating improved class-level adaptation. Extensive experiments indicate that our proposed ECAN significantly outperforms state-of-the-art methods under the source-free cross-corpus SER setting on several speech emotion corpora. Yan Zhao 0037, Jincen Wang, Cheng Lu 0005, Sunan Li, Björn W. Schuller, Yuan Zong, Wenming Zheng |
ICASSP | 1 |
| 2024 | Hierarchical Distribution Adaptation for Unsupervised Cross-corpus Speech Emotion Recognition
Cheng Lu 0005, Yuan Zong, Yan Zhao 0037, Hailun Lian, Tianhua Qi, Björn W. Schuller, Wenming Zheng |
INTERSPEECH | 3 |
| 2024 | Towards Realistic Emotional Voice Conversion using Controllable Emotional Intensity
Tianhua Qi, Shiyan Wang, Cheng Lu 0005, Yan Zhao 0037, Yuan Zong, Wenming Zheng |
INTERSPEECH | 4 |
| 2024 | Boosting Cross-Corpus Speech Emotion Recognition using CycleGAN with Contrastive Learning
Jincen Wang, Yan Zhao 0037, Cheng Lu 0005, Chuangao Tang, Sunan Li, Yuan Zong, Wenming Zheng |
INTERSPEECH | 2 |
| 2024 | Confidence-aware Hypothesis Transfer Networks for Source-Free Cross-Corpus Speech Emotion Recognition
Jincen Wang, Yan Zhao 0037, Cheng Lu 0005, Hailun Lian, Hongli Chang, Yuan Zong, Wenming Zheng |
INTERSPEECH | 2 |
| 2024 | Exploring corpus-invariant emotional acoustic feature for cross-corpus speech emotion recognition
Hailun Lian, Cheng Lu 0005, Yan Zhao 0037, Sunan Li, Tianhua Qi, Yuan Zong |
Expert Syst. Appl. | 3 |
| 2024 | CFEW: A Large-Scale Database for Understanding Child Facial Expression in Real WorldabstractCurrently, much progress has been achieved on adult facial expressions recognition. Few attentions have been paid to child facial expression analysis. A lack of publicly available large-scale child facial expression databases hinders the development of automatic coding for child facial expression behaviors. In this work, we constructed a new face database for understandingChildFacialExpression in realWorld (CFEW). The database contains three novelties: (1) the largest publicly available child facial expression database (11,000+ images); (2) covering full developmental range of 0–18-year-old child subjects; (3) rich annotations for facial expression labels, including discrete expression categories aka happy, neutral, disgust, angry, sad, cry, fear, surprise, sleepy and others, intensity of arousal and valence, and several types of facial action units (AUs). In addition, the images in this database cover several challenging conditions in real world, including frontal and non-frontal head poses, facial occlusions, various illuminations and low image resolution. Three dominant deep convolutional neural networks (i.e., VGG11bn, ResNet18 and DenseNet121) were used to conduct extensive baseline experiments for discrete facial expression classification, arousal and valence estimation and facial action units detection within database, and cross-database seven facial expressions recognition. Chuangao Tang, Sunan Li, Wenming Zheng, Yuan Zong, Su Zhang 0004, Cheng Lu 0005, Yan Zhao 0037 |
IEEE Trans. Affect. Comput. | 7 |
| 2024 | Layer-Adapted Implicit Distribution Alignment Networks for Cross-Corpus Speech Emotion RecognitionabstractIn this article, we propose a new unsupervised domain adaptation (DA) method called layer-adapted implicit distribution alignment networks (LIDANs) to address the challenge of cross-corpus speech emotion recognition (SER). LIDAN extends our previous ICASSP work, deep implicit distribution alignment networks (DIDANs), whose key contribution lies in the introduction of a novel regularization term called implicit distribution alignment (IDA). This term allows DIDAN trained on source (training) speech samples to remain applicable to predicting emotion labels for target (testing) speech samples, regardless of corpus variance in cross-corpus SER. To further enhance this method, we extend IDA to layer-adapted IDA (LIDA), resulting in LIDAN. This layer-adapted extension consists of three modified IDA terms that consider emotion labels at different levels of granularity. These terms are strategically arranged within different fully connected layers in LIDAN, aligning with the increasing emotion-discriminative abilities with respect to the layer depth. This arrangement enables LIDAN to more effectively learn emotion-discriminative and corpus-invariant features for SER across various corpora compared to DIDAN. It is also worthy to mention that unlike most existing methods that rely on estimating statistical moments to describe preassumed explicit distributions, both IDA and LIDA take a different approach. They utilize an idea of target sample reconstruction to directly bridge the feature distribution gap without making assumptions about their distribution type. As a result, DIDAN and LIDAN can be viewed as implicit cross-corpus SER methods. To evaluate LIDAN, we conducted extensive cross-corpus SER experiments on EmoDB, eNTERFACE, and CASIA corpora. The experimental results demonstrate that LIDAN surpasses recent state-of-theart explicit unsupervised DA methods in tackling cross-corpus SER tasks. Yan Zhao 0037, Yuan Zong, Jincen Wang, Hailun Lian, Cheng Lu 0005, Li Zhao 0003, Wenming Zheng |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2023 | Deep Implicit Distribution Alignment Networks for cross-Corpus Speech Emotion RecognitionabstractIn this paper, we propose a novel deep transfer learning method called deep implicit distribution alignment networks (DIDAN) to deal with cross-corpus speech emotion recognition (SER) problem, in which the labeled training (source) and unlabeled testing (target) speech signals come from different corpora. Specifically, DIDAN first adopts a simple deep regression network consisting of a set of convolutional and fully connected layers to directly regress the source speech spectrums into the emotional labels such that the proposed DIDAN can own the emotion discriminative ability. Then, such ability is transferred to be also applicable to the target speech samples regardless of corpus variance by resorting to a well-designed regularization term called implicit distribution alignment (IDA). Unlike widely-used maximum mean discrepancy (MMD) and its variants, the proposed IDA absorbs the idea of sample reconstruction to implicitly align the distribution gap, which enables DIDAN to learn both emotion discriminative and corpus invariant features from speech spectrums. To evaluate the proposed DIDAN, extensive cross-corpus SER experiments on widely-used speech emotion corpora are carried out. Experimental results show that the proposed DIDAN can outperform lots of recent state-of-the-art methods in coping with the cross-corpus SER tasks. Yan Zhao 0037, Jincen Wang, Yuan Zong, Wenming Zheng, Hailun Lian, Li Zhao 0003 |
ICASSP | 1 |
| 2023 | Time-Frequency Transformer: A Novel Time Frequency Joint Learning Method for Speech Emotion Recognition
Yong Wang 0073, Cheng Lu 0005, Yuan Zong, Hailun Lian, Yan Zhao 0037, Sunan Li |
ICONIP (9) | 5 |
| 2023 | Learning Local to Global Feature Aggregation for Speech Emotion Recognition
Cheng Lu 0005, Hailun Lian, Wenming Zheng, Yuan Zong, Yan Zhao 0037, Sunan Li |
INTERSPEECH | 5 |
| 2023 | Multimodal Emotion Recognition in Noisy Environment Based on Progressive Label RevisionabstractThe multimodal emotion recognition has attracted more attention in recent decades. Though remarkable progress has been achieved with the rapid development of deep learning, existing methods are still hard to tackle noise problems that occurred commonly in emotion recognition's practical application. To improve the robustness of the multimodal emotion recognition algorithm, we propose an MLP-based label revision algorithm. The framework consists of three complementary feature extraction networks that were verified in MER2023. After that, an MLP-based attention network with specially designed loss functions was used to fuse features from different modalities. Finally, the scheme that used the output probability of each emotion to revise the sample's output category was employed to revise the test set's label obtained by classifier. The samples that are most likely to be affected by noise and misclassified have a chance to get correct classification. The best experimental result shows that the F1-score of our algorithm on the test dataset of the MER 2023 Noise subchallenge is 86.35 and combined metric is 0.6694, which ranks 2nd at the MER 2023 NOISE subchallenge. Sunan Li, Hailun Lian, Cheng Lu 0005, Yan Zhao 0037, Chuangao Tang, Yuan Zong, Wenming Zheng |
ACM Multimedia | 4 |
| 2023 | Speech Emotion Recognition via an Attentive Time-Frequency Neural NetworkabstractSpectrogram is commonly used as the input feature of deep neural networks to learn the high(er)-level time–frequency pattern of speech signal for speech emotion recognition (SER). Generally, different emotions correspond to specific energy activations both within frequency bands and time frames on spectrogram, which indicates the frequency and time domains are both essential to represent the emotion for SER. However, recent spectrogram-based works mainly focus on modeling the long-term dependency in time domain, which makes these methods suffer from the following issues: 1) neglecting to model the emotion-related correlations within frequency domain during the time–frequency joint learning and 2) ignoring to capture the specific frequency bands associated with emotions. To cope with the issues, we propose an attentive time–frequency neural network (ATFNN) for SER, including a time–frequency neural network (TFNN) and time–frequency attention. Specifically, aiming at the first issue, we design a TFNN with a frequency-domain encoder (F-Encoder) based on the Transformer encoder and a time-domain encoder (T-Encoder) based on the bidirectional long short-term memory (Bi-LSTM). The F-Encoder and T-Encoder model the correlations within frequency bands and time frames, respectively, and they are embedded into a time–frequency joint learning strategy to obtain the time–frequency patterns of speech emotions. Moreover, to handle the second issue, we adopt the time–frequency attention with a frequency-attention network (F-Attention) and a time-attention network (T-Attention) to focus on the emotion-related long-range dependencies between frequency bands and across time frames, which can enhance the emotional discrimination of speech features. Extensive experimental results on three public emotional databases, i.e., IEMOCAP, ABC, and CASIA, show that our proposed ATFNN outperforms the state-of-the-art methods. Cheng Lu 0005, Wenming Zheng, Hailun Lian, Yuan Zong, Chuangao Tang, Sunan Li, Yan Zhao 0037 |
IEEE Trans. Comput. Soc. Syst. | 7 |
| 2022 | Deep Transductive Transfer Regression Network for Cross-Corpus Speech Emotion Recognition
Yan Zhao 0037, Jincen Wang, Ru Ye, Yuan Zong, Wenming Zheng, Li Zhao 0003 |
INTERSPEECH | 1 |