Sunan Li

dblp:210/7709 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0003-1494-4873ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Speaker-independent speech emotion recognition using group sparse-based adversarial local fisher discriminant analysis
Cheng Lu 0005, Kaifei Zhang, Hailun Lian, Sunan Li, Tianhua Qi, Yuan Zong, Wenming Zheng
Pattern Recognit.4
2026 An Interpretable Collaborative Acoustic Parameter Modeling Network for Speech Emotion Recognition
abstract
Acoustic parameters can collaboratively represent speech emotions. However, existing speech emotion recognition (SER) methods neglect explicit collaborative modeling among acoustic parameters during representation learning. This not only limits their ability to extract the information contained in acoustic parameters that collaboratively represent emotions, but also makes it difficult to trace back which acoustic parameters are closely associated with emotional states (i.e., lack of interpretability). To address these challenges, we propose the interpretable collaborative acoustic parameter modeling network (ICAPM-Net), which innovatively introduces Graph Convolutional Networks (GCN) to model the relationships among acoustic parameters. Specifically, ICAPM-Net represents different acoustic parameters as nodes, with edges characterizing their collaborative relationships. The GCN utilizes edges to explicitly model the collaborative dependencies between acoustic parameters. Importantly, visualizing the edges in the graph (i.e., adjacency matrix) can trace which acoustic parameters (or their combinations) play dominant roles in emotional expression. The overall architecture of ICAPM-Net consists of three modules: 1) an encoding module, which is used to unify feature dimensions of acoustic parameters; 2) a graph-based collaborative modeling module (GBCM), which is responsible for modeling the collaborative relationships among acoustic parameters; and 3) an emotion adjacency matrix alignment loss module (EAMAL), which is designed to further enhance collaborative modeling in GBCM. Experiments on IEMOCAP, ABC, and EMO-DB show that ICAPM-Net outperforms state-of-the-art methods. Additionally, visualization results reveal that certain acoustic parameters, e.g., MFCC and RASTA, have a strong correlation with emotions and exhibit significant collaborative patterns in different emotional contexts.
Hailun Lian, Cheng Lu 0005, Hao Yang 0028, Yan Zhao 0037, Sunan Li, Yuan Zong
IEEE Trans. Comput. Soc. Syst.5
2026 Feature Evaluation and Joint Interaction for Audio-Visual Emotion Recognition
abstract
Automatic emotion recognition has attracted significant attention due to its potential applications in various real-world scenarios. Methods that integrate visual and audio modalities have become increasingly prominent because of their superior information-carrying capacity and complementarity. Despite advancements in feature fusion between video and audio, existing modality fusion-based methods struggle to effectively address the dynamic changes in feature quality caused by interference, which is common in emotion recognition in the wild tasks. To overcome this limitation, we propose a Parameter-Free Feature Evaluation and Interaction (PFFEI) model based on information quality assessment. The model leverages the scaling factor γ of the normalization layer to evaluate information quality and dynamically adjusts the degree of interaction between modalities, suppressing the impact of low-quality features affected by interference. Additionally, the norm constraint integrated into the model ensures that the γ value consistently measures feature quality across different modalities. This approach effectively mitigates the effects of modality imbalance and significantly enhances the model’s accuracy. The effectiveness of our method is demonstrated through experiments on three challenging real-world emotion datasets: DFEW, AFEW, and Ekman6. The results show that the PFFEI model outperforms state-of-the-art methods, achieving significant improvements of 8.71% (UAR) and 8.61% (WAR) on the AFEW database.
Sunan Li, Cheng Lu 0005, Yuan Zong, Hailun Lian, Wenming Zheng
IEEE Trans. Circuits Syst. Video Technol.1
2025 Low-rank joint distribution adaptation for cross-corpus speech emotion recognition
Sunan Li, Cheng Lu 0005, Yan Zhao 0037, Hailun Lian, Tianhua Qi, Yuan Zong
Knowl. Based Syst.1
2025 AMGCN: An adaptive multi-graph convolutional network for speech emotion recognition
Hailun Lian, Cheng Lu 0005, Hongli Chang, Yan Zhao 0037, Sunan Li, Yang Li 0019, Yuan Zong
Speech Commun.5
2024 Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition
abstract
Cross-corpus speech emotion recognition (SER) aims to transfer emotional knowledge from a labeled source corpus to an unlabeled corpus. However, prior methods require access to source data during adaptation, which is unattainable in real-life scenarios due to data privacy protection concerns. This paper tackles a more practical task, namely source-free cross-corpus SER, where a pre-trained source model is adapted to the target domain without access to source data. To address the problem, we propose a novel method called emotion-aware contrastive adaptation network (ECAN). The core idea is to capture local neighborhood information between samples while considering the global class-level adaptation. Specifically, we propose a nearest neighbor contrastive learning to promote local emotion consistency among features of highly similar samples. Furthermore, relying solely on nearest neighborhoods may lead to ambiguous boundaries between clusters. Thus, we incorporate supervised contrastive learning to encourage greater separation between clusters representing different emotions, thereby facilitating improved class-level adaptation. Extensive experiments indicate that our proposed ECAN significantly outperforms state-of-the-art methods under the source-free cross-corpus SER setting on several speech emotion corpora.
Yan Zhao 0037, Jincen Wang, Cheng Lu 0005, Sunan Li, Björn W. Schuller, Yuan Zong, Wenming Zheng
ICASSP4
2024 Boosting Cross-Corpus Speech Emotion Recognition using CycleGAN with Contrastive Learning
Jincen Wang, Yan Zhao 0037, Cheng Lu 0005, Chuangao Tang, Sunan Li, Yuan Zong, Wenming Zheng
INTERSPEECH5
2024 Exploring corpus-invariant emotional acoustic feature for cross-corpus speech emotion recognition
Hailun Lian, Cheng Lu 0005, Yan Zhao 0037, Sunan Li, Tianhua Qi, Yuan Zong
Expert Syst. Appl.4
2024 CFEW: A Large-Scale Database for Understanding Child Facial Expression in Real World
abstract
Currently, much progress has been achieved on adult facial expressions recognition. Few attentions have been paid to child facial expression analysis. A lack of publicly available large-scale child facial expression databases hinders the development of automatic coding for child facial expression behaviors. In this work, we constructed a new face database for understandingChildFacialExpression in realWorld (CFEW). The database contains three novelties: (1) the largest publicly available child facial expression database (11,000+ images); (2) covering full developmental range of 0–18-year-old child subjects; (3) rich annotations for facial expression labels, including discrete expression categories aka happy, neutral, disgust, angry, sad, cry, fear, surprise, sleepy and others, intensity of arousal and valence, and several types of facial action units (AUs). In addition, the images in this database cover several challenging conditions in real world, including frontal and non-frontal head poses, facial occlusions, various illuminations and low image resolution. Three dominant deep convolutional neural networks (i.e., VGG11bn, ResNet18 and DenseNet121) were used to conduct extensive baseline experiments for discrete facial expression classification, arousal and valence estimation and facial action units detection within database, and cross-database seven facial expressions recognition.
Chuangao Tang, Sunan Li, Wenming Zheng, Yuan Zong, Su Zhang 0004, Cheng Lu 0005, Yan Zhao 0037
IEEE Trans. Affect. Comput.2
2023 Time-Frequency Transformer: A Novel Time Frequency Joint Learning Method for Speech Emotion Recognition
Yong Wang 0073, Cheng Lu 0005, Yuan Zong, Hailun Lian, Yan Zhao 0037, Sunan Li
ICONIP (9)6
2023 Learning Local to Global Feature Aggregation for Speech Emotion Recognition
Cheng Lu 0005, Hailun Lian, Wenming Zheng, Yuan Zong, Yan Zhao 0037, Sunan Li
INTERSPEECH6
2023 Multimodal Emotion Recognition in Noisy Environment Based on Progressive Label Revision
abstract
The multimodal emotion recognition has attracted more attention in recent decades. Though remarkable progress has been achieved with the rapid development of deep learning, existing methods are still hard to tackle noise problems that occurred commonly in emotion recognition's practical application. To improve the robustness of the multimodal emotion recognition algorithm, we propose an MLP-based label revision algorithm. The framework consists of three complementary feature extraction networks that were verified in MER2023. After that, an MLP-based attention network with specially designed loss functions was used to fuse features from different modalities. Finally, the scheme that used the output probability of each emotion to revise the sample's output category was employed to revise the test set's label obtained by classifier. The samples that are most likely to be affected by noise and misclassified have a chance to get correct classification. The best experimental result shows that the F1-score of our algorithm on the test dataset of the MER 2023 Noise subchallenge is 86.35 and combined metric is 0.6694, which ranks 2nd at the MER 2023 NOISE subchallenge.
Sunan Li, Hailun Lian, Cheng Lu 0005, Yan Zhao 0037, Chuangao Tang, Yuan Zong, Wenming Zheng
ACM Multimedia1
2023 Speech Emotion Recognition via an Attentive Time-Frequency Neural Network
abstract
Spectrogram is commonly used as the input feature of deep neural networks to learn the high(er)-level time–frequency pattern of speech signal for speech emotion recognition (SER). Generally, different emotions correspond to specific energy activations both within frequency bands and time frames on spectrogram, which indicates the frequency and time domains are both essential to represent the emotion for SER. However, recent spectrogram-based works mainly focus on modeling the long-term dependency in time domain, which makes these methods suffer from the following issues: 1) neglecting to model the emotion-related correlations within frequency domain during the time–frequency joint learning and 2) ignoring to capture the specific frequency bands associated with emotions. To cope with the issues, we propose an attentive time–frequency neural network (ATFNN) for SER, including a time–frequency neural network (TFNN) and time–frequency attention. Specifically, aiming at the first issue, we design a TFNN with a frequency-domain encoder (F-Encoder) based on the Transformer encoder and a time-domain encoder (T-Encoder) based on the bidirectional long short-term memory (Bi-LSTM). The F-Encoder and T-Encoder model the correlations within frequency bands and time frames, respectively, and they are embedded into a time–frequency joint learning strategy to obtain the time–frequency patterns of speech emotions. Moreover, to handle the second issue, we adopt the time–frequency attention with a frequency-attention network (F-Attention) and a time-attention network (T-Attention) to focus on the emotion-related long-range dependencies between frequency bands and across time frames, which can enhance the emotional discrimination of speech features. Extensive experimental results on three public emotional databases, i.e., IEMOCAP, ABC, and CASIA, show that our proposed ATFNN outperforms the state-of-the-art methods.
Cheng Lu 0005, Wenming Zheng, Hailun Lian, Yuan Zong, Chuangao Tang, Sunan Li, Yan Zhao 0037
IEEE Trans. Comput. Soc. Syst.6
2019 Human Movements Classification Using Multi-channel Surface EMG Signals and Deep Learning Technique
abstract
Electromyography (EMG) signals can be used for human movements classification. Nonetheless, due to their nonlinear and time-varying properties, it is difficult to classify the EMG signals and it is critical to use appropriate algorithms for EMG feature extraction and pattern classification. In literature various machine learning (ML) methods have been applied to the EMG signal classification problem in question. In this paper, we extracted four time-domain features of the EMG signals and use a generative graphical model, Deep Belief Network (DBN), to classify the EMG signals. A DBN is a fast, greedy deep learning algorithm that can rapidly find a set of optimal weights of a deep network with many hidden layers. To evaluate the DBN model, we acquired EMG signals, extracted their time-domain features, and then utilized the DBN model to classify human movements. The real data analysis results are presented to show the effectiveness of the proposed deep learning technique for both binary and 4-class recognition of human movements using the measured 8-channel EMG signals. The proposed DBN model may find applications in design of EMG-based user interfaces.
Sunan Li
CW3
2019 Bi-modality Fusion for Emotion Recognition in the Wild
abstract
The emotion recognition in the wild has been a hot research topic in the field of affective computing. Though some progresses have been achieved, the emotion recognition in the wild is still an unsolved problem due to the challenge of head movement, face deformation, illumination variation etc. To deal with these unconstrained challenges, we propose a bi-modality fusion method for video based emotion recognition in the wild. The proposed framework takes advantages of the visual information from facial expression sequences and the speech information from audio. The state-of-the-art CNN based object recognition models are employed to facilitate the facial expression recognition performance. A bi-direction long short term Memory (Bi-LSTM) is employed to capture dynamic information of the learned features. Additionally, to take full advantages of the facial expression information, the VGG16 network is trained on AffectNet dataset to learn a specialized facial expression recognition model. On the other hand, the audio based features, like low level descriptor (LLD) and deep features obtained by spectrogram image, are also developed to improve the emotion recognition performance. The best experimental result shows that the overall accuracy of our algorithm on the Test dataset of the EmotiW challenge is 62.78, which outperforms the best result of EmotiW2018 and ranks 2nd at the EmotiW2019 challenge.
Sunan Li, Wenming Zheng, Yuan Zong, Cheng Lu 0005, Chuangao Tang, Xingxun Jiang, Jiateng Liu, Wanchuang Xia
ICMI1