Hailun Lian

dblp:235/0669 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Two-stream attentive spatial-temporal graph convolutional network for P300 detection in brain-computer interface
Jincen Wang, Yan Zhao 0037, Cunhang Fan, Yong Li 0032, Fan Liu 0003, Hailun Lian, Cheng Lu 0005
Expert Syst. Appl.6
2026 Speaker-independent speech emotion recognition using group sparse-based adversarial local fisher discriminant analysis
Cheng Lu 0005, Kaifei Zhang, Hailun Lian, Sunan Li, Tianhua Qi, Yuan Zong, Wenming Zheng
Pattern Recognit.3
2026 An Interpretable Collaborative Acoustic Parameter Modeling Network for Speech Emotion Recognition
abstract
Acoustic parameters can collaboratively represent speech emotions. However, existing speech emotion recognition (SER) methods neglect explicit collaborative modeling among acoustic parameters during representation learning. This not only limits their ability to extract the information contained in acoustic parameters that collaboratively represent emotions, but also makes it difficult to trace back which acoustic parameters are closely associated with emotional states (i.e., lack of interpretability). To address these challenges, we propose the interpretable collaborative acoustic parameter modeling network (ICAPM-Net), which innovatively introduces Graph Convolutional Networks (GCN) to model the relationships among acoustic parameters. Specifically, ICAPM-Net represents different acoustic parameters as nodes, with edges characterizing their collaborative relationships. The GCN utilizes edges to explicitly model the collaborative dependencies between acoustic parameters. Importantly, visualizing the edges in the graph (i.e., adjacency matrix) can trace which acoustic parameters (or their combinations) play dominant roles in emotional expression. The overall architecture of ICAPM-Net consists of three modules: 1) an encoding module, which is used to unify feature dimensions of acoustic parameters; 2) a graph-based collaborative modeling module (GBCM), which is responsible for modeling the collaborative relationships among acoustic parameters; and 3) an emotion adjacency matrix alignment loss module (EAMAL), which is designed to further enhance collaborative modeling in GBCM. Experiments on IEMOCAP, ABC, and EMO-DB show that ICAPM-Net outperforms state-of-the-art methods. Additionally, visualization results reveal that certain acoustic parameters, e.g., MFCC and RASTA, have a strong correlation with emotions and exhibit significant collaborative patterns in different emotional contexts.
Hailun Lian, Cheng Lu 0005, Hao Yang 0028, Yan Zhao 0037, Sunan Li, Yuan Zong
IEEE Trans. Comput. Soc. Syst.1
2026 Feature Evaluation and Joint Interaction for Audio-Visual Emotion Recognition
abstract
Automatic emotion recognition has attracted significant attention due to its potential applications in various real-world scenarios. Methods that integrate visual and audio modalities have become increasingly prominent because of their superior information-carrying capacity and complementarity. Despite advancements in feature fusion between video and audio, existing modality fusion-based methods struggle to effectively address the dynamic changes in feature quality caused by interference, which is common in emotion recognition in the wild tasks. To overcome this limitation, we propose a Parameter-Free Feature Evaluation and Interaction (PFFEI) model based on information quality assessment. The model leverages the scaling factor γ of the normalization layer to evaluate information quality and dynamically adjusts the degree of interaction between modalities, suppressing the impact of low-quality features affected by interference. Additionally, the norm constraint integrated into the model ensures that the γ value consistently measures feature quality across different modalities. This approach effectively mitigates the effects of modality imbalance and significantly enhances the model’s accuracy. The effectiveness of our method is demonstrated through experiments on three challenging real-world emotion datasets: DFEW, AFEW, and Ekman6. The results show that the PFFEI model outperforms state-of-the-art methods, achieving significant improvements of 8.71% (UAR) and 8.61% (WAR) on the AFEW database.
Sunan Li, Cheng Lu 0005, Yuan Zong, Hailun Lian, Wenming Zheng
IEEE Trans. Circuits Syst. Video Technol.4
2025 Low-rank joint distribution adaptation for cross-corpus speech emotion recognition
Sunan Li, Cheng Lu 0005, Yan Zhao 0037, Hailun Lian, Tianhua Qi, Yuan Zong
Knowl. Based Syst.4
2025 AMGCN: An adaptive multi-graph convolutional network for speech emotion recognition
Hailun Lian, Cheng Lu 0005, Hongli Chang, Yan Zhao 0037, Sunan Li, Yang Li 0019, Yuan Zong
Speech Commun.1
2025 Towards Domain-Specific Cross-Corpus Speech Emotion Recognition Approach
abstract
Cross-corpus speech emotion recognition (SER) poses a challenge due to feature distribution mismatch between the training and testing speech samples, potentially degrading the performance of established SER methods. In this article, we tackle this challenge by proposing a novel transfer subspace learning method called acoustic knowledge-guided transfer linear regression (AKTLR). Unlike existing approaches, which often overlook domain-specific knowledge related to SER and simply treat cross-corpus SER as a generic transfer learning task, our AKTLR method is built upon a well-designed acoustic knowledge-guided dual sparsity constraint mechanism. This mechanism emphasizes the potential of minimalistic acoustic parameter feature sets to alleviate classifier over-adaptation, which is empirically validated acoustic knowledge in SER, enabling superior generalization in cross-corpus SER tasks compared to using large feature sets. Through this mechanism, we extend a simple transfer linear regression model to AKTLR. This extension harnesses its full capability to seek emotion-discriminative and corpus-invariant features from established acoustic parameter feature sets used for describing speech signals across two scales: contributive acoustic parameter groups and constituent elements within each contributive group. We evaluate our method through extensive cross-corpus SER experiments on three widely used speech emotion corpora: EmoDB, eNTERFACE, and CASIA. The proposed AKTLR achieves an average UAR of 42.12% across six tasks using the eGeMAPS feature set, outperforming many recent state-of-the-art transfer subspace learning and deep transfer learning methods. This demonstrates the effectiveness and superior performance of our approach. Furthermore, our work provides experimental evidence supporting the feasibility and superiority of incorporating domain-specific knowledge into the transfer learning model to address cross-corpus SER tasks.
Yan Zhao 0037, Yuan Zong, Hailun Lian, Cheng Lu 0005, Jingang Shi, Wenming Zheng
IEEE Trans. Comput. Soc. Syst.3
2024 Improving Speaker-Independent Speech Emotion Recognition using Dynamic Joint Distribution Adaptation
abstract
In speaker-independent speech emotion recognition, the training and testing samples are collected from diverse speakers, leading to a multi-domain shift challenge across the feature distributions of data from different speakers. Consequently, when the trained model is confronted with data from new speakers, its performance tends to degrade. To address the issue, we propose a Dynamic Joint Distribution Adaptation (DJDA) method under the framework of multi-source domain adaptation. DJDA firstly utilizes joint distribution adaptation (JDA), involving marginal distribution adaptation (MDA) and conditional distribution adaptation (CDA), to more precisely measure the multi-domain distribution shifts caused by different speakers. This helps eliminate speaker bias in emotion features, allowing for learning discriminative and speaker-invariant speech emotion features from coarse-level to fine-level. Furthermore, we quantify the adaptation contributions of MDA and CDA within JDA by using a dynamic balance factor based on $\mathcal{A}$-Distance, promoting to effectively handle the unknown distributions encountered in data from new speakers. Experimental results demonstrate the superior performance of our DJDA as compared to other state-of-the-art (SOTA) methods.
Cheng Lu 0005, Yuan Zong, Hailun Lian, Yan Zhao 0037, Björn W. Schuller, Wenming Zheng
ICASSP3
2024 PAVITS: Exploring Prosody-Aware VITS for End-to-End Emotional Voice Conversion
abstract
In this paper, we propose Prosody-aware VITS (PAVITS) for emotional voice conversion (EVC), aiming to achieve two major objectives of EVC: high content naturalness and high emotional naturalness, which are crucial for meeting the demands of human perception. To improve the content naturalness of converted audio, we have developed an end-to-end EVC architecture inspired by the high audio quality of VITS. By seamlessly integrating an acoustic converter and vocoder, we effectively address the common issue of mismatch between emotional prosody training and run-time conversion that is prevalent in existing EVC models. To further enhance the emotional naturalness, we introduce an emotion descriptor to model the subtle prosody variations of different speech emotions. Additionally, we propose a prosody predictor, which predicts prosody features from text based on the provided emotion label. Notably, we introduce a prosody alignment loss to establish a connection between latent prosody features from two distinct modalities, ensuring effective training. Experimental results show that the performance of PAVITS is superior to the state-of-the-art EVC methods. Speech Samples are available at https://jeremychee4.github.io/pavits4EVC/.
Tianhua Qi, Wenming Zheng, Cheng Lu 0005, Yuan Zong, Hailun Lian
ICASSP5
2024 Speech Swin-Transformer: Exploring a Hierarchical Transformer with Shifted Windows for Speech Emotion Recognition
abstract
Swin-Transformer has demonstrated remarkable success in computer vision by leveraging its hierarchical feature representation based on Transformer. In speech signals, emotional information is distributed across different scales of speech features, e. g., word, phrase, and utterance. Drawing above inspiration, this paper presents a hierarchical speech Transformer with shifted windows to aggregate multi-scale emotion features for speech emotion recognition (SER), called Speech Swin-Transformer. Specifically, we first divide the speech spectrogram into segment-level patches in the time domain, composed of multiple frame patches. These segment-level patches are then encoded using a stack of Swin blocks, in which a local window Transformer is utilized to explore local inter-frame emotional information across frame patches of each segment patch. After that, we also design a shifted window Transformer to compensate for patch correlations near the boundaries of segment patches. Finally, we employ a patch merging operation to aggregate segment-level emotional features for hierarchical speech representation by expanding the receptive field of Transformer from frame-level to segment-level. Experimental results demonstrate that our proposed Speech Swin-Transformer outperforms the state-of-the-art methods.
Yong Wang 0073, Cheng Lu 0005, Hailun Lian, Yan Zhao 0037, Björn W. Schuller, Yuan Zong, Wenming Zheng
ICASSP3
2024 Hierarchical Distribution Adaptation for Unsupervised Cross-corpus Speech Emotion Recognition
Cheng Lu 0005, Yuan Zong, Yan Zhao 0037, Hailun Lian, Tianhua Qi, Björn W. Schuller, Wenming Zheng
INTERSPEECH4
2024 Confidence-aware Hypothesis Transfer Networks for Source-Free Cross-Corpus Speech Emotion Recognition
Jincen Wang, Yan Zhao 0037, Cheng Lu 0005, Hailun Lian, Hongli Chang, Yuan Zong, Wenming Zheng
INTERSPEECH4
2024 Exploring corpus-invariant emotional acoustic feature for cross-corpus speech emotion recognition
Hailun Lian, Cheng Lu 0005, Yan Zhao 0037, Sunan Li, Tianhua Qi, Yuan Zong
Expert Syst. Appl.1
2024 Layer-Adapted Implicit Distribution Alignment Networks for Cross-Corpus Speech Emotion Recognition
abstract
In this article, we propose a new unsupervised domain adaptation (DA) method called layer-adapted implicit distribution alignment networks (LIDANs) to address the challenge of cross-corpus speech emotion recognition (SER). LIDAN extends our previous ICASSP work, deep implicit distribution alignment networks (DIDANs), whose key contribution lies in the introduction of a novel regularization term called implicit distribution alignment (IDA). This term allows DIDAN trained on source (training) speech samples to remain applicable to predicting emotion labels for target (testing) speech samples, regardless of corpus variance in cross-corpus SER. To further enhance this method, we extend IDA to layer-adapted IDA (LIDA), resulting in LIDAN. This layer-adapted extension consists of three modified IDA terms that consider emotion labels at different levels of granularity. These terms are strategically arranged within different fully connected layers in LIDAN, aligning with the increasing emotion-discriminative abilities with respect to the layer depth. This arrangement enables LIDAN to more effectively learn emotion-discriminative and corpus-invariant features for SER across various corpora compared to DIDAN. It is also worthy to mention that unlike most existing methods that rely on estimating statistical moments to describe preassumed explicit distributions, both IDA and LIDA take a different approach. They utilize an idea of target sample reconstruction to directly bridge the feature distribution gap without making assumptions about their distribution type. As a result, DIDAN and LIDAN can be viewed as implicit cross-corpus SER methods. To evaluate LIDAN, we conducted extensive cross-corpus SER experiments on EmoDB, eNTERFACE, and CASIA corpora. The experimental results demonstrate that LIDAN surpasses recent state-of-theart explicit unsupervised DA methods in tackling cross-corpus SER tasks.
Yan Zhao 0037, Yuan Zong, Jincen Wang, Hailun Lian, Cheng Lu 0005, Li Zhao 0003, Wenming Zheng
IEEE Trans. Comput. Soc. Syst.4
2023 Deep Implicit Distribution Alignment Networks for cross-Corpus Speech Emotion Recognition
abstract
In this paper, we propose a novel deep transfer learning method called deep implicit distribution alignment networks (DIDAN) to deal with cross-corpus speech emotion recognition (SER) problem, in which the labeled training (source) and unlabeled testing (target) speech signals come from different corpora. Specifically, DIDAN first adopts a simple deep regression network consisting of a set of convolutional and fully connected layers to directly regress the source speech spectrums into the emotional labels such that the proposed DIDAN can own the emotion discriminative ability. Then, such ability is transferred to be also applicable to the target speech samples regardless of corpus variance by resorting to a well-designed regularization term called implicit distribution alignment (IDA). Unlike widely-used maximum mean discrepancy (MMD) and its variants, the proposed IDA absorbs the idea of sample reconstruction to implicitly align the distribution gap, which enables DIDAN to learn both emotion discriminative and corpus invariant features from speech spectrums. To evaluate the proposed DIDAN, extensive cross-corpus SER experiments on widely-used speech emotion corpora are carried out. Experimental results show that the proposed DIDAN can outperform lots of recent state-of-the-art methods in coping with the cross-corpus SER tasks.
Yan Zhao 0037, Jincen Wang, Yuan Zong, Wenming Zheng, Hailun Lian, Li Zhao 0003
ICASSP5
2023 Time-Frequency Transformer: A Novel Time Frequency Joint Learning Method for Speech Emotion Recognition
Yong Wang 0073, Cheng Lu 0005, Yuan Zong, Hailun Lian, Yan Zhao 0037, Sunan Li
ICONIP (9)4
2023 Learning Local to Global Feature Aggregation for Speech Emotion Recognition
Cheng Lu 0005, Hailun Lian, Wenming Zheng, Yuan Zong, Yan Zhao 0037, Sunan Li
INTERSPEECH2
2023 Multimodal Emotion Recognition in Noisy Environment Based on Progressive Label Revision
abstract
The multimodal emotion recognition has attracted more attention in recent decades. Though remarkable progress has been achieved with the rapid development of deep learning, existing methods are still hard to tackle noise problems that occurred commonly in emotion recognition's practical application. To improve the robustness of the multimodal emotion recognition algorithm, we propose an MLP-based label revision algorithm. The framework consists of three complementary feature extraction networks that were verified in MER2023. After that, an MLP-based attention network with specially designed loss functions was used to fuse features from different modalities. Finally, the scheme that used the output probability of each emotion to revise the sample's output category was employed to revise the test set's label obtained by classifier. The samples that are most likely to be affected by noise and misclassified have a chance to get correct classification. The best experimental result shows that the F1-score of our algorithm on the test dataset of the MER 2023 Noise subchallenge is 86.35 and combined metric is 0.6694, which ranks 2nd at the MER 2023 NOISE subchallenge.
Sunan Li, Hailun Lian, Cheng Lu 0005, Yan Zhao 0037, Chuangao Tang, Yuan Zong, Wenming Zheng
ACM Multimedia2
2023 Speech Emotion Recognition via an Attentive Time-Frequency Neural Network
abstract
Spectrogram is commonly used as the input feature of deep neural networks to learn the high(er)-level time–frequency pattern of speech signal for speech emotion recognition (SER). Generally, different emotions correspond to specific energy activations both within frequency bands and time frames on spectrogram, which indicates the frequency and time domains are both essential to represent the emotion for SER. However, recent spectrogram-based works mainly focus on modeling the long-term dependency in time domain, which makes these methods suffer from the following issues: 1) neglecting to model the emotion-related correlations within frequency domain during the time–frequency joint learning and 2) ignoring to capture the specific frequency bands associated with emotions. To cope with the issues, we propose an attentive time–frequency neural network (ATFNN) for SER, including a time–frequency neural network (TFNN) and time–frequency attention. Specifically, aiming at the first issue, we design a TFNN with a frequency-domain encoder (F-Encoder) based on the Transformer encoder and a time-domain encoder (T-Encoder) based on the bidirectional long short-term memory (Bi-LSTM). The F-Encoder and T-Encoder model the correlations within frequency bands and time frames, respectively, and they are embedded into a time–frequency joint learning strategy to obtain the time–frequency patterns of speech emotions. Moreover, to handle the second issue, we adopt the time–frequency attention with a frequency-attention network (F-Attention) and a time-attention network (T-Attention) to focus on the emotion-related long-range dependencies between frequency bands and across time frames, which can enhance the emotional discrimination of speech features. Extensive experimental results on three public emotional databases, i.e., IEMOCAP, ABC, and CASIA, show that our proposed ATFNN outperforms the state-of-the-art methods.
Cheng Lu 0005, Wenming Zheng, Hailun Lian, Yuan Zong, Chuangao Tang, Sunan Li, Yan Zhao 0037
IEEE Trans. Comput. Soc. Syst.3