EDBT 2026 Demo / reviewers in the wild / expert
Cheng Lu 0005
dblp:91/1482-5
· DBLP profile ↗
44ranked-venue papers
7as first author
40since 2021 · last 2026
0000-0002-1477-1020ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 3 first-author · 23 since 2021Artificial intelligence and machine learning · 22 · 4 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Two-stream attentive spatial-temporal graph convolutional network for P300 detection in brain-computer interface
Jincen Wang, Yan Zhao 0037, Cunhang Fan, Yong Li 0032, Fan Liu 0003, Hailun Lian, Cheng Lu 0005 |
Expert Syst. Appl. | 7 |
| 2026 | Speaker-independent speech emotion recognition using group sparse-based adversarial local fisher discriminant analysis
Cheng Lu 0005, Kaifei Zhang, Hailun Lian, Sunan Li, Tianhua Qi, Yuan Zong, Wenming Zheng |
Pattern Recognit. | 1 |
| 2026 | FAME: Frequency and motion extrapolation for robust multimodal facial action unit detection
Mengxin Shi, Cheng Lu 0005, Zhangfeng Hu, Hongli Chang, Yuan Zong |
Pattern Recognit. | 2 |
| 2026 | An Interpretable Collaborative Acoustic Parameter Modeling Network for Speech Emotion RecognitionabstractAcoustic parameters can collaboratively represent speech emotions. However, existing speech emotion recognition (SER) methods neglect explicit collaborative modeling among acoustic parameters during representation learning. This not only limits their ability to extract the information contained in acoustic parameters that collaboratively represent emotions, but also makes it difficult to trace back which acoustic parameters are closely associated with emotional states (i.e., lack of interpretability). To address these challenges, we propose the interpretable collaborative acoustic parameter modeling network (ICAPM-Net), which innovatively introduces Graph Convolutional Networks (GCN) to model the relationships among acoustic parameters. Specifically, ICAPM-Net represents different acoustic parameters as nodes, with edges characterizing their collaborative relationships. The GCN utilizes edges to explicitly model the collaborative dependencies between acoustic parameters. Importantly, visualizing the edges in the graph (i.e., adjacency matrix) can trace which acoustic parameters (or their combinations) play dominant roles in emotional expression. The overall architecture of ICAPM-Net consists of three modules: 1) an encoding module, which is used to unify feature dimensions of acoustic parameters; 2) a graph-based collaborative modeling module (GBCM), which is responsible for modeling the collaborative relationships among acoustic parameters; and 3) an emotion adjacency matrix alignment loss module (EAMAL), which is designed to further enhance collaborative modeling in GBCM. Experiments on IEMOCAP, ABC, and EMO-DB show that ICAPM-Net outperforms state-of-the-art methods. Additionally, visualization results reveal that certain acoustic parameters, e.g., MFCC and RASTA, have a strong correlation with emotions and exhibit significant collaborative patterns in different emotional contexts. Hailun Lian, Cheng Lu 0005, Hao Yang 0028, Yan Zhao 0037, Sunan Li, Yuan Zong |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2026 | Feature Evaluation and Joint Interaction for Audio-Visual Emotion RecognitionabstractAutomatic emotion recognition has attracted significant attention due to its potential applications in various real-world scenarios. Methods that integrate visual and audio modalities have become increasingly prominent because of their superior information-carrying capacity and complementarity. Despite advancements in feature fusion between video and audio, existing modality fusion-based methods struggle to effectively address the dynamic changes in feature quality caused by interference, which is common in emotion recognition in the wild tasks. To overcome this limitation, we propose a Parameter-Free Feature Evaluation and Interaction (PFFEI) model based on information quality assessment. The model leverages the scaling factor γ of the normalization layer to evaluate information quality and dynamically adjusts the degree of interaction between modalities, suppressing the impact of low-quality features affected by interference. Additionally, the norm constraint integrated into the model ensures that the γ value consistently measures feature quality across different modalities. This approach effectively mitigates the effects of modality imbalance and significantly enhances the model’s accuracy. The effectiveness of our method is demonstrated through experiments on three challenging real-world emotion datasets: DFEW, AFEW, and Ekman6. The results show that the PFFEI model outperforms state-of-the-art methods, achieving significant improvements of 8.71% (UAR) and 8.61% (WAR) on the AFEW database. Sunan Li, Cheng Lu 0005, Yuan Zong, Hailun Lian, Wenming Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Enhancing Task-Specific Feature Learning with LLMs for Multimodal Emotion and Intent Joint UnderstandingabstractThis paper introduces our solution, the Task-Specific Feature Learning (TSFL) method, designed to address the second track of the MEIJU Challenge at ICASSP 2025, namely, Imbalanced Emotion and Intent Recognition (English). The TSFL method incorporates three core components: the use of LLM features to represent multimodal signals, coarse-grained task-specific feature decomposition, and fine-grained task-specific feature learning. These components enable the effective joint learning of emotion-discriminative and intent-discriminative features. As a result, our method achieved a JRBM score of 0.6230, significantly outperforming the official baseline result and surpassing all other competing teams to win the championship. Cheng Lu 0005, Kaifei Zhang, Yujia Gu, Banghua Li, Yuan Zong, Wenming Zheng |
ICASSP | 2 |
| 2025 | Enhancing Zero-Shot Emotional Voice Conversion via Speaker Adaptation and Duration PredictionabstractZero-shot Emotional Voice Conversion (EVC) aims to transform a speaker’s emotional state to match a target emotion, even for speakers and emotion categories that were not encountered during training, thereby enhancing the generalization ability of traditional EVC systems. Despite advancements in the field, existing methods often face challenges in preserving speaker identity and ensuring the naturalness of emotional expression, particularly in the context of rhythm modeling. To this end, we propose the Zero-Shot Emotion Voice Conversion (ZSEVC) model, which leverages self-supervised learning for speaker adaptation and duration prediction. To adjust speech rhythm in alignment with the target emotional state, we introduce a rhythm-aware content encoder that captures and refines discrete speech units at a finer granularity. Additionally, a hierarchical emotion fusion scheme is employed to integrate emotional features with content features, enhancing both pronunciation accuracy and emotional expressiveness. Moreover, a residual speaker-emotion fusion module is incorporated to better adapt speaker characteristics to emotional prosodic variation. Experimental results show ZSEVC’s superior performance in terms of naturalness and speaker similarity in zero-shot scenario, successfully generating emotional speeches for unseen emotions and speakers. Speech samples are available at https://wosyoo.github.io/ZSEVC. Shiyan Wang, Tianhua Qi, Cheng Lu 0005, Zhaojie Luo, Wenming Zheng |
ICASSP | 3 |
| 2025 | Reliable Learning From LLM Features for Multimodal Emotion and Intent Joint UnderstandingabstractThis paper describes a Reliable Learning Framework (RLF) for the 1st Multimodal Emotion and Intent Joint Understanding (MEIJU) Challenge at ICASSP 2025. Our proposed RLF includes a Hierarchical Interaction Network and a Reliable Fusion Strategy. The former can excavate emotion and intent cues from the high-level semantic features of multimodal data (video, audio, and text) generated by pretrained Large Language Models (LMMs), to enhance their representations, and the latter reliably integrates multiple predictions to further improve the robustness of emotion and intent understanding. Our RLF method achieved first place on Track 2 (Mandarin) of MEIJU, with performance scores for emotion, intent, and joint recognition reaching 0.7285, 0.7456, and 0.7370. Cheng Lu 0005, Yuyun Liu, Yinghao Ma, Jiahao Luo, Yuan Zong, Wenming Zheng |
ICASSP | 2 |
| 2025 | DisenEmo: Learning disentangled emotional representation from facial motion for 3D talking head generationabstractEmotional 3D talking head generation synthesizes vivid facial expressions with precise lip synchronization for immersive interactions. We introduce DisenEMO, a novel framework designed to disentangle emotion and content from facial motions, thereby facilitating the synthesis of personalized and expressive audio-driven facial animations. To achieve precise emotional disentanglement, we incorporate an intensity perception constraint which improves the accurate perception of categorized emotion and its intensities, leading to the generation of subtle emotional expressions. To ensure the temporal consistency of facial expressions, we introduce facial dynamic modeling, which refines motion trajectories to better capture emotional nuances. Finally, a motion decoder integrates emotional features with audio features extracted from driving speech, producing 3D talking heads with enhanced emotional expressiveness and realism. Experimental results demonstrate that our method outperforms state-of-the-art approaches. Synthesis samples are available at https://c295bw.github.io/DisenEMO-icip25.github.io/. Tianhua Qi, Cheng Lu 0005, Wenming Zheng |
ICIP | 3 |
| 2025 | Interactive Fusion of Multi-View Speech Embeddings via Pretrained Large-Scale Speech Models for Speech Emotional Attribute Prediction in Naturalistic Conditions
Yuyun Liu, Yujia Gu, Jiahao Luo, Wenming Zheng, Cheng Lu 0005, Yuan Zong |
INTERSPEECH | 5 |
| 2025 | PromptEVC: Controllable Emotional Voice Conversion with Natural Language Prompts
Tianhua Qi, Shiyan Wang, Cheng Lu 0005, Tengfei Song, Hao Yang 0006, Zhanglin Wu, Wenming Zheng |
INTERSPEECH | 3 |
| 2025 | Multi-Level Segment Fusion Based on Adaptive Time-Window Selection for Multimodal Personality-Aware Elderly Depression DetectionabstractMajor Depressive Disorder (MDD) is a prevalent and severe psychiatric disorder, and its detection remains challenging due to the complexity and variability of its symptoms. Traditional single-modality methods often fail to capture the full spectrum of depressive cues, which has led to the rise of multimodal methods. The ACM Multimedia 2025 ''Multimodal Personality-Aware Depression Detection Challenge'' (MPDD 2025) aims to advance the development of more accurate depression detection models by incorporating multimodal data. In this paper, we proposed a Multi-Level Segment Fusion Based on Adaptive Time-Window Selection (MSF-ATS) method for the MPDD-Elderly Track. To address the challenge of sparse and transient depressive symptoms, we fuse segment-level classifications to obtain subject-level classifications. An adaptive time-window selection based on mean class variance is employed to choose the window with the smallest variance for more stable detection results. Our method achieved an average score of 0.8576 on the MPDD 2025 official test set, significantly outperforming the baseline score of 0.6675. Yuyun Liu, Kaifei Zhang, Yinghao Ma, Tianhua Qi, Wenming Zheng, Cheng Lu 0005, Yuan Zong |
ACM Multimedia | 7 |
| 2025 | Low-rank joint distribution adaptation for cross-corpus speech emotion recognition
Sunan Li, Cheng Lu 0005, Yan Zhao 0037, Hailun Lian, Tianhua Qi, Yuan Zong |
Knowl. Based Syst. | 2 |
| 2025 | AMGCN: An adaptive multi-graph convolutional network for speech emotion recognition
Hailun Lian, Cheng Lu 0005, Hongli Chang, Yan Zhao 0037, Sunan Li, Yang Li 0019, Yuan Zong |
Speech Commun. | 2 |
| 2025 | Learning to Rank Onset-Occurring-Offset Representations for Micro-Expression RecognitionabstractThis paper focuses on the research of micro-expression recognition (MER) and proposes a flexible and reliable deep learning method called learning to rank onset-occurring-offset representations (LTR3O). The LTR3O method introduces a dynamic and reduced-size sequence structure known as 3O, which consists of onset, occurring, and offset frames, for representing micro-expressions (MEs). This structure facilitates the subsequent learning of ME-discriminative features. A noteworthy advantage of the 3O structure is its flexibility, as the occurring frame is randomly extracted from the original ME sequence without the need for accurate frame spotting methods. Based on the 3O structures, LTR3O generates multiple 3O representation candidates for each ME sample and incorporates well-designed modules based on learning to rank (LTR) to measure and calibrate their emotional expressiveness. This calibration process implicitly enhances the visibility of MEs by amplifying the originally narrow emotional expressiveness gap among ME frames caused by their low-intensity characteristics, thereby facilitating the reliable learning of more discriminative features for MER. Extensive experiments were conducted to evaluate the performance of LTR3O using four widely-used ME databases: CASME II, SMIC, SAMM, and MEVIEW. The experimental results demonstrate the effectiveness and superior performance of LTR3O, particularly in terms of its flexibility and reliability, when compared to recent state-of-the-art MER methods. Yuan Zong, Jingang Shi, Cheng Lu 0005, Hongli Chang, Wenming Zheng |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | Towards Domain-Specific Cross-Corpus Speech Emotion Recognition ApproachabstractCross-corpus speech emotion recognition (SER) poses a challenge due to feature distribution mismatch between the training and testing speech samples, potentially degrading the performance of established SER methods. In this article, we tackle this challenge by proposing a novel transfer subspace learning method called acoustic knowledge-guided transfer linear regression (AKTLR). Unlike existing approaches, which often overlook domain-specific knowledge related to SER and simply treat cross-corpus SER as a generic transfer learning task, our AKTLR method is built upon a well-designed acoustic knowledge-guided dual sparsity constraint mechanism. This mechanism emphasizes the potential of minimalistic acoustic parameter feature sets to alleviate classifier over-adaptation, which is empirically validated acoustic knowledge in SER, enabling superior generalization in cross-corpus SER tasks compared to using large feature sets. Through this mechanism, we extend a simple transfer linear regression model to AKTLR. This extension harnesses its full capability to seek emotion-discriminative and corpus-invariant features from established acoustic parameter feature sets used for describing speech signals across two scales: contributive acoustic parameter groups and constituent elements within each contributive group. We evaluate our method through extensive cross-corpus SER experiments on three widely used speech emotion corpora: EmoDB, eNTERFACE, and CASIA. The proposed AKTLR achieves an average UAR of 42.12% across six tasks using the eGeMAPS feature set, outperforming many recent state-of-the-art transfer subspace learning and deep transfer learning methods. This demonstrates the effectiveness and superior performance of our approach. Furthermore, our work provides experimental evidence supporting the feasibility and superiority of incorporating domain-specific knowledge into the transfer learning model to address cross-corpus SER tasks. Yan Zhao 0037, Yuan Zong, Hailun Lian, Cheng Lu 0005, Jingang Shi, Wenming Zheng |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | Improving Speaker-Independent Speech Emotion Recognition using Dynamic Joint Distribution AdaptationabstractIn speaker-independent speech emotion recognition, the training and testing samples are collected from diverse speakers, leading to a multi-domain shift challenge across the feature distributions of data from different speakers. Consequently, when the trained model is confronted with data from new speakers, its performance tends to degrade. To address the issue, we propose a Dynamic Joint Distribution Adaptation (DJDA) method under the framework of multi-source domain adaptation. DJDA firstly utilizes joint distribution adaptation (JDA), involving marginal distribution adaptation (MDA) and conditional distribution adaptation (CDA), to more precisely measure the multi-domain distribution shifts caused by different speakers. This helps eliminate speaker bias in emotion features, allowing for learning discriminative and speaker-invariant speech emotion features from coarse-level to fine-level. Furthermore, we quantify the adaptation contributions of MDA and CDA within JDA by using a dynamic balance factor based on $\mathcal{A}$-Distance, promoting to effectively handle the unknown distributions encountered in data from new speakers. Experimental results demonstrate the superior performance of our DJDA as compared to other state-of-the-art (SOTA) methods. Cheng Lu 0005, Yuan Zong, Hailun Lian, Yan Zhao 0037, Björn W. Schuller, Wenming Zheng |
ICASSP | 1 |
| 2024 | PAVITS: Exploring Prosody-Aware VITS for End-to-End Emotional Voice ConversionabstractIn this paper, we propose Prosody-aware VITS (PAVITS) for emotional voice conversion (EVC), aiming to achieve two major objectives of EVC: high content naturalness and high emotional naturalness, which are crucial for meeting the demands of human perception. To improve the content naturalness of converted audio, we have developed an end-to-end EVC architecture inspired by the high audio quality of VITS. By seamlessly integrating an acoustic converter and vocoder, we effectively address the common issue of mismatch between emotional prosody training and run-time conversion that is prevalent in existing EVC models. To further enhance the emotional naturalness, we introduce an emotion descriptor to model the subtle prosody variations of different speech emotions. Additionally, we propose a prosody predictor, which predicts prosody features from text based on the provided emotion label. Notably, we introduce a prosody alignment loss to establish a connection between latent prosody features from two distinct modalities, ensuring effective training. Experimental results show that the performance of PAVITS is superior to the state-of-the-art EVC methods. Speech Samples are available at https://jeremychee4.github.io/pavits4EVC/. Tianhua Qi, Wenming Zheng, Cheng Lu 0005, Yuan Zong, Hailun Lian |
ICASSP | 3 |
| 2024 | Speech Swin-Transformer: Exploring a Hierarchical Transformer with Shifted Windows for Speech Emotion RecognitionabstractSwin-Transformer has demonstrated remarkable success in computer vision by leveraging its hierarchical feature representation based on Transformer. In speech signals, emotional information is distributed across different scales of speech features, e. g., word, phrase, and utterance. Drawing above inspiration, this paper presents a hierarchical speech Transformer with shifted windows to aggregate multi-scale emotion features for speech emotion recognition (SER), called Speech Swin-Transformer. Specifically, we first divide the speech spectrogram into segment-level patches in the time domain, composed of multiple frame patches. These segment-level patches are then encoded using a stack of Swin blocks, in which a local window Transformer is utilized to explore local inter-frame emotional information across frame patches of each segment patch. After that, we also design a shifted window Transformer to compensate for patch correlations near the boundaries of segment patches. Finally, we employ a patch merging operation to aggregate segment-level emotional features for hierarchical speech representation by expanding the receptive field of Transformer from frame-level to segment-level. Experimental results demonstrate that our proposed Speech Swin-Transformer outperforms the state-of-the-art methods. Yong Wang 0073, Cheng Lu 0005, Hailun Lian, Yan Zhao 0037, Björn W. Schuller, Yuan Zong, Wenming Zheng |
ICASSP | 2 |
| 2024 | Progressively Learning from Macro-Expressions for Micro-Expression RecognitionabstractMicro-expression (ME) recognition is challenging due to the low-intensity facial motions. An idea to overcome this is learning assisted by macro-expressions (MaEs). However, the intensity gap between MaE and ME is so huge that related works fail to effectively leverage MaE’s assistance in overcoming low-intensity interference, which attempt to directly force ME knowledge to mimic MaE knowledge. In this paper, we propose that the knowledge transfer from MaE to ME can be converted into a progressive process for better implementation. Thus, we construct a progressive multi-step learning framework, which accomplishes two tasks: first, we dissect the huge intensity gap into multiple segments that are easier to bridge by constructing multiple learning steps, each corresponding to various intensity levels of expression recognition tasks. Second, through a designed self-knowledge distillation (self-KD) model, each dissected gap can be bridged, enabling the MaE knowledge to progressively transfer to guide the ME learning. Experiments carried out on three widely used databases demonstrated that the proposed PLMaM achieves state-of-the-art results. Yuan Zong, Mengting Wei, Cheng Lu 0005, Wenming Zheng |
ICASSP | 6 |
| 2024 | Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion RecognitionabstractCross-corpus speech emotion recognition (SER) aims to transfer emotional knowledge from a labeled source corpus to an unlabeled corpus. However, prior methods require access to source data during adaptation, which is unattainable in real-life scenarios due to data privacy protection concerns. This paper tackles a more practical task, namely source-free cross-corpus SER, where a pre-trained source model is adapted to the target domain without access to source data. To address the problem, we propose a novel method called emotion-aware contrastive adaptation network (ECAN). The core idea is to capture local neighborhood information between samples while considering the global class-level adaptation. Specifically, we propose a nearest neighbor contrastive learning to promote local emotion consistency among features of highly similar samples. Furthermore, relying solely on nearest neighborhoods may lead to ambiguous boundaries between clusters. Thus, we incorporate supervised contrastive learning to encourage greater separation between clusters representing different emotions, thereby facilitating improved class-level adaptation. Extensive experiments indicate that our proposed ECAN significantly outperforms state-of-the-art methods under the source-free cross-corpus SER setting on several speech emotion corpora. Yan Zhao 0037, Jincen Wang, Cheng Lu 0005, Sunan Li, Björn W. Schuller, Yuan Zong, Wenming Zheng |
ICASSP | 3 |
| 2024 | Hierarchical Distribution Adaptation for Unsupervised Cross-corpus Speech Emotion Recognition
Cheng Lu 0005, Yuan Zong, Yan Zhao 0037, Hailun Lian, Tianhua Qi, Björn W. Schuller, Wenming Zheng |
INTERSPEECH | 1 |
| 2024 | Towards Realistic Emotional Voice Conversion using Controllable Emotional Intensity
Tianhua Qi, Shiyan Wang, Cheng Lu 0005, Yan Zhao 0037, Yuan Zong, Wenming Zheng |
INTERSPEECH | 3 |
| 2024 | Boosting Cross-Corpus Speech Emotion Recognition using CycleGAN with Contrastive Learning
Jincen Wang, Yan Zhao 0037, Cheng Lu 0005, Chuangao Tang, Sunan Li, Yuan Zong, Wenming Zheng |
INTERSPEECH | 3 |
| 2024 | Confidence-aware Hypothesis Transfer Networks for Source-Free Cross-Corpus Speech Emotion Recognition
Jincen Wang, Yan Zhao 0037, Cheng Lu 0005, Hailun Lian, Hongli Chang, Yuan Zong, Wenming Zheng |
INTERSPEECH | 3 |
| 2024 | Exploring corpus-invariant emotional acoustic feature for cross-corpus speech emotion recognition
Hailun Lian, Cheng Lu 0005, Yan Zhao 0037, Sunan Li, Tianhua Qi, Yuan Zong |
Expert Syst. Appl. | 2 |
| 2024 | Exploring holistic discriminative representation for micro-expression recognition via contrastive learning
Wanyuan He, Hongli Chang, Cheng Lu 0005, Yuan Zong |
Image Vis. Comput. | 5 |
| 2024 | CFEW: A Large-Scale Database for Understanding Child Facial Expression in Real WorldabstractCurrently, much progress has been achieved on adult facial expressions recognition. Few attentions have been paid to child facial expression analysis. A lack of publicly available large-scale child facial expression databases hinders the development of automatic coding for child facial expression behaviors. In this work, we constructed a new face database for understandingChildFacialExpression in realWorld (CFEW). The database contains three novelties: (1) the largest publicly available child facial expression database (11,000+ images); (2) covering full developmental range of 0–18-year-old child subjects; (3) rich annotations for facial expression labels, including discrete expression categories aka happy, neutral, disgust, angry, sad, cry, fear, surprise, sleepy and others, intensity of arousal and valence, and several types of facial action units (AUs). In addition, the images in this database cover several challenging conditions in real world, including frontal and non-frontal head poses, facial occlusions, various illuminations and low image resolution. Three dominant deep convolutional neural networks (i.e., VGG11bn, ResNet18 and DenseNet121) were used to conduct extensive baseline experiments for discrete facial expression classification, arousal and valence estimation and facial action units detection within database, and cross-database seven facial expressions recognition. Chuangao Tang, Sunan Li, Wenming Zheng, Yuan Zong, Su Zhang 0004, Cheng Lu 0005, Yan Zhao 0037 |
IEEE Trans. Affect. Comput. | 6 |
| 2024 | Layer-Adapted Implicit Distribution Alignment Networks for Cross-Corpus Speech Emotion RecognitionabstractIn this article, we propose a new unsupervised domain adaptation (DA) method called layer-adapted implicit distribution alignment networks (LIDANs) to address the challenge of cross-corpus speech emotion recognition (SER). LIDAN extends our previous ICASSP work, deep implicit distribution alignment networks (DIDANs), whose key contribution lies in the introduction of a novel regularization term called implicit distribution alignment (IDA). This term allows DIDAN trained on source (training) speech samples to remain applicable to predicting emotion labels for target (testing) speech samples, regardless of corpus variance in cross-corpus SER. To further enhance this method, we extend IDA to layer-adapted IDA (LIDA), resulting in LIDAN. This layer-adapted extension consists of three modified IDA terms that consider emotion labels at different levels of granularity. These terms are strategically arranged within different fully connected layers in LIDAN, aligning with the increasing emotion-discriminative abilities with respect to the layer depth. This arrangement enables LIDAN to more effectively learn emotion-discriminative and corpus-invariant features for SER across various corpora compared to DIDAN. It is also worthy to mention that unlike most existing methods that rely on estimating statistical moments to describe preassumed explicit distributions, both IDA and LIDA take a different approach. They utilize an idea of target sample reconstruction to directly bridge the feature distribution gap without making assumptions about their distribution type. As a result, DIDAN and LIDAN can be viewed as implicit cross-corpus SER methods. To evaluate LIDAN, we conducted extensive cross-corpus SER experiments on EmoDB, eNTERFACE, and CASIA corpora. The experimental results demonstrate that LIDAN surpasses recent state-of-theart explicit unsupervised DA methods in tackling cross-corpus SER tasks. Yan Zhao 0037, Yuan Zong, Jincen Wang, Hailun Lian, Cheng Lu 0005, Li Zhao 0003, Wenming Zheng |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2023 | CMNet: Contrastive Magnification Network for Micro-Expression RecognitionabstractMicro-Expression Recognition (MER) is challenging because the Micro-Expressions' (ME) motion is too weak to distinguish. This hurdle can be tackled by enhancing intensity for a more accurate acquisition of movements. However, existing magnification strategies tend to use the features of facial images that include not only intensity clues as intensity features, leading to the intensity representation deficient of credibility. In addition, the intensity variation over time, which is crucial for encoding movements, is also neglected. To this end, we provide a reliable scheme to extract intensity clues while considering their variation on the time scale. First, we devise an Intensity Distillation (ID) loss to acquire the intensity clues by contrasting the difference between frames, given that the difference in the same video lies only in the intensity. Then, the intensity clues are calibrated to follow the trend of the original video. Specifically, due to the lack of truth intensity annotation of the original video, we build the intensity tendency by setting each intensity vacancy an uncertain value, which guides the extracted intensity clues to converge towards this trend rather some fixed values. A Wilcoxon rank sum test (Wrst) method is enforced to implement the calibration. Experimental results on three public ME databases i.e. CASME II, SAMM, and SMIC-HS validate the superiority against state-of-the-art methods. Mengting Wei, Xingxun Jiang, Wenming Zheng, Yuan Zong, Cheng Lu 0005, Jiateng Liu |
AAAI | 5 |
| 2023 | Time-Frequency Transformer: A Novel Time Frequency Joint Learning Method for Speech Emotion Recognition
Yong Wang 0073, Cheng Lu 0005, Yuan Zong, Hailun Lian, Yan Zhao 0037, Sunan Li |
ICONIP (9) | 2 |
| 2023 | Learning Local to Global Feature Aggregation for Speech Emotion Recognition
Cheng Lu 0005, Hailun Lian, Wenming Zheng, Yuan Zong, Yan Zhao 0037, Sunan Li |
INTERSPEECH | 1 |
| 2023 | Multimodal Emotion Recognition in Noisy Environment Based on Progressive Label RevisionabstractThe multimodal emotion recognition has attracted more attention in recent decades. Though remarkable progress has been achieved with the rapid development of deep learning, existing methods are still hard to tackle noise problems that occurred commonly in emotion recognition's practical application. To improve the robustness of the multimodal emotion recognition algorithm, we propose an MLP-based label revision algorithm. The framework consists of three complementary feature extraction networks that were verified in MER2023. After that, an MLP-based attention network with specially designed loss functions was used to fuse features from different modalities. Finally, the scheme that used the output probability of each emotion to revise the sample's output category was employed to revise the test set's label obtained by classifier. The samples that are most likely to be affected by noise and misclassified have a chance to get correct classification. The best experimental result shows that the F1-score of our algorithm on the test dataset of the MER 2023 Noise subchallenge is 86.35 and combined metric is 0.6694, which ranks 2nd at the MER 2023 NOISE subchallenge. Sunan Li, Hailun Lian, Cheng Lu 0005, Yan Zhao 0037, Chuangao Tang, Yuan Zong, Wenming Zheng |
ACM Multimedia | 3 |
| 2023 | Speech Emotion Recognition via an Attentive Time-Frequency Neural NetworkabstractSpectrogram is commonly used as the input feature of deep neural networks to learn the high(er)-level time–frequency pattern of speech signal for speech emotion recognition (SER). Generally, different emotions correspond to specific energy activations both within frequency bands and time frames on spectrogram, which indicates the frequency and time domains are both essential to represent the emotion for SER. However, recent spectrogram-based works mainly focus on modeling the long-term dependency in time domain, which makes these methods suffer from the following issues: 1) neglecting to model the emotion-related correlations within frequency domain during the time–frequency joint learning and 2) ignoring to capture the specific frequency bands associated with emotions. To cope with the issues, we propose an attentive time–frequency neural network (ATFNN) for SER, including a time–frequency neural network (TFNN) and time–frequency attention. Specifically, aiming at the first issue, we design a TFNN with a frequency-domain encoder (F-Encoder) based on the Transformer encoder and a time-domain encoder (T-Encoder) based on the bidirectional long short-term memory (Bi-LSTM). The F-Encoder and T-Encoder model the correlations within frequency bands and time frames, respectively, and they are embedded into a time–frequency joint learning strategy to obtain the time–frequency patterns of speech emotions. Moreover, to handle the second issue, we adopt the time–frequency attention with a frequency-attention network (F-Attention) and a time-attention network (T-Attention) to focus on the emotion-related long-range dependencies between frequency bands and across time frames, which can enhance the emotional discrimination of speech features. Extensive experimental results on three public emotional databases, i.e., IEMOCAP, ABC, and CASIA, show that our proposed ATFNN outperforms the state-of-the-art methods. Cheng Lu 0005, Wenming Zheng, Hailun Lian, Yuan Zong, Chuangao Tang, Sunan Li, Yan Zhao 0037 |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2023 | EEG-Based Parkinson's Disease Recognition via Attention-Based Sparse Graph Convolutional Neural NetworkabstractParkinson's disease (PD) is a complicated neurological ailment that affects both the physical and mental wellness of elderly individuals which makes it problematic to diagnose in its initial stages. Electroencephalogram (EEG) promises to be an efficient and cost-effective method for promptly detecting cognitive impairment in PD. Nevertheless, prevailing diagnostic practices utilizing EEG features have failed to examine the functional connectivity among EEG channels and the response of associated brain areas causing an unsatisfactory level of precision. Here, we construct an attention-based sparse graph convolutional neural network (ASGCNN) for diagnosing PD. Our ASGCNN model uses a graph structure to represent channel relationships, the attention mechanism for selecting channels, and the L1 norm to capture channel sparsity. We conduct extensive experiments on the publicly available PD auditory oddball dataset, which consists of 24 PD patients (under ON/OFF drug status) and 24 matched controls, to validate the effectiveness of our method. Our results show that the proposed method provides better results compared to the publicly available baselines. The achieved scores for Recall, Precision, F1-score, Accuracy and Kappa measures are 90.36%, 88.43%, 88.41%, 87.67%, and 75.24%, respectively. Our study reveals that the frontal and temporal lobes show significant differences between PD patients and healthy individuals. In addition, EEG features extracted by ASGCNN demonstrate significant asymmetry in the frontal lobe among PD patients. These findings can offer a basis for the establishment of a clinical system for intelligent diagnosis of PD by using auditory cognitive impairment features. Hongli Chang, Yuan Zong, Cheng Lu 0005, Xuenan Wang |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | A Novel Micro-Expression Recognition Approach Using Attention-Based Magnification-Adaptive NetworksabstractMicro-Expression recognition (MER) is a challenging task due to the short duration and low intensity of Micro-Expressions. A popular method to tackle this is magnifying MEs so as to enlarge the expression intensity to make recognition easier. However, the single fixed magnification strategy, widely used in existing works of MER, is not appropriate for different subjects, because each subject has specific expression intensity corresponding to different MEs. To cope with this issue, we propose a novel Attention-based Magnification-Adaptive Network (AMAN) to learn adaptive magnification levels for the ME representation. The network consists of two modules: magnification attention (MA module) to adaptively focus on appropriate magnification levels of different MEs, and frame attention (FA module) to focus on discriminative aggregated frames in a ME video. Extensive experiments on three widely used databases manifest that our method yields state-of-art results compared with other methods. Mengting Wei, Wenming Zheng, Yuan Zong, Xingxun Jiang, Cheng Lu 0005, Jiateng Liu |
ICASSP | 5 |
| 2022 | A Novel Magnification-Robust Network with Sparse Self-Attention for Micro-expression RecognitionabstractExisting works for spontaneous Micro-Expression Recognition (MER) tend to encode Micro-Expression (ME) movements to get more discriminative features. However, MEs’ low intensity makes the capture for motion extremely difficult, and the widely adopted unified-magnification strategy is prone to noise and lacks flexibility. To this end, this paper provides a new insight to encode ME motion and tackle magnification noise. Specifically, we reconstruct a new sequence via magnification techniques to make subtle ME movements more distinguishable. Afterward, Sparse Self-Attention (SSA) rectifies self-attention with Locality Sensitive Hashing (LSH), cutting the space into several hush buckets of related features. Only keys in the same bucket are operated in the attention term for every query feature. The resulting sparsity in the attention matrix prevents the network from attending features stemming from less-informative magnification degrees which could be regarded as noise, while retains the sequence modelling capability of standard self-attention. Extensive experiments on three public MER databases demonstrate our superiority against the state-of-the-art methods. Mengting Wei, Wenming Zheng, Xingxun Jiang, Yuan Zong, Cheng Lu 0005, Jiateng Liu |
ICPR | 5 |
| 2022 | Sample Self-Revised Network for Cross-Dataset Facial Expression RecognitionabstractFacial images with low quality, subjective annotation, severe occlusion, and rare subject identity can lead to the existence of outlier samples in facial expression datasets. These outlier samples are usually far from the center of the dataset in the feature space, resulting in huge differences in feature distribution, which severely restricts the performance of cross-dataset facial expression recognition (FER). To eliminate the influence of outlier samples on cross-dataset FER, we propose an unsupervised domain adaptation (UDA) method called Sample Self-Revised Network (SSRN), which 1) dynamically detects the outlier level of each sample in the source domain to reduce the disturbance of outlier samples to the model training, as well as 2) adaptively revises outlier samples in the source domain to improve transferability of the learned features. Experimental results show that our SSRN outperforms both classic deep UDA methods and state-of-the-art cross-dataset FER results. Wenming Zheng, Yuan Zong, Cheng Lu 0005, Xingxun Jiang |
IJCNN | 4 |
| 2022 | Domain Invariant Feature Learning for Speaker-Independent Speech Emotion RecognitionabstractIn this paper, we propose a novel domain invariant feature learning (DIFL) method to deal with speaker-independent speech emotion recognition (SER). The basic idea of DIFL is to learn the speaker-invariant emotion feature by eliminating domain shifts between the training and testing data caused by different speakers from the perspective of multi-source unsupervised domain adaptation (UDA). Specifically, we embed a hierarchical alignment layer with the strong-weak distribution alignment strategy into the feature extraction block to firstly reduce the discrepancy in feature distributions of speech samples across different speakers as much as possible. Furthermore, multiple discriminators in the discriminator block are utilized to confuse the speaker information of emotion features both inside the training data and between the training and testing data. Through them, a multi-domain invariant representation of emotional speech can be gradually and adaptively achieved by updating network parameters. We conduct extensive experiments on three public datasets, i. e., Emo-DB, eNTERFACE, and CASIA, to evaluate the SER performance of the proposed method, respectively. The experimental results show that the proposed method is superior to the state-of-the-art methods. Cheng Lu 0005, Yuan Zong, Wenming Zheng, Yang Li 0019, Chuangao Tang, Björn W. Schuller |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | Attention-based Spatio-Temporal Graphic LSTM for EEG Emotion RecognitionabstractAutomatic emotion recognition based on electroencephalogram (EEG) is a challenging task in Brain Machine Interfaces (BMI). Since it is still not very clear about the intrinsic connection relationship among the various EEG channels, it is still a challenging task of how to better represent the topology of EEG channels for emotion recognition. On the other hand, the intensity of the emotion may vary along the different time instants, which would affect the recognition accuracy of emotion. To tackle the above issues, in this paper, we propose a novel multichannel EEG emotion recognition method called attention-based spatiotemporal graphic long short-term memory (ASTG-LSTM), in which a Dynamic Structured Learning (DSL) branch that focuses on the most emotion-relevant connectivity of the brain is incorporated to represent inter-channel connections of the EEG signals. In addition, a specific spatio-temporal attention is embedded into the DSL branch to improve the invariance ability against the emotional intensity fluctuation. Extensive experiments on: DEAP and DREAMER are conducted and the experimental results indicate that the proposed ASTG-LSTM model improves the EEG emotion recognition performance compared with many state-of-the-art approaches. Wenming Zheng, Yuan Zong, Hongli Chang, Cheng Lu 0005 |
IJCNN | 5 |
| 2020 | DFEW: A Large-Scale Database for Recognizing Dynamic Facial Expressions in the WildabstractRecently, facial expression recognition (FER) in the wild has gained a lot of researchers' attention because it is a valuable topic to enable the FER techniques to move from the laboratory to the real applications. In this paper, we focus on this challenging but interesting topic and make contributions from three aspects. First, we present a new large-scale 'in-the-wild' dynamic facial expression database, DFEW (Dynamic Facial Expression in the Wild), consisting of over 16,000 video clips from thousands of movies. These video clips contain various challenging interferences in practical scenarios such as extreme illumination, occlusions, and capricious pose changes. Second, we propose a novel method called Expression-Clustered Spatiotemporal Feature Learning (EC-STFL) framework to deal with dynamic FER in the wild. Third, we conduct extensive benchmark experiments on DFEW using a lot of spatiotemporal deep feature learning methods as well as our proposed EC-STFL. Experimental results show that DFEW is a well-designed and challenging database, and the proposed EC-STFL can promisingly improve the performance of existing spatiotemporal deep neural networks in coping with the problem of dynamic FER in the wild. Our DFEW database is publicly available and can be freely downloaded from https://dfew-dataset.github.io/. Xingxun Jiang, Yuan Zong, Wenming Zheng, Chuangao Tang, Wanchuang Xia, Cheng Lu 0005, Jiateng Liu |
ACM Multimedia | 6 |
| 2019 | Bi-modality Fusion for Emotion Recognition in the WildabstractThe emotion recognition in the wild has been a hot research topic in the field of affective computing. Though some progresses have been achieved, the emotion recognition in the wild is still an unsolved problem due to the challenge of head movement, face deformation, illumination variation etc. To deal with these unconstrained challenges, we propose a bi-modality fusion method for video based emotion recognition in the wild. The proposed framework takes advantages of the visual information from facial expression sequences and the speech information from audio. The state-of-the-art CNN based object recognition models are employed to facilitate the facial expression recognition performance. A bi-direction long short term Memory (Bi-LSTM) is employed to capture dynamic information of the learned features. Additionally, to take full advantages of the facial expression information, the VGG16 network is trained on AffectNet dataset to learn a specialized facial expression recognition model. On the other hand, the audio based features, like low level descriptor (LLD) and deep features obtained by spectrogram image, are also developed to improve the emotion recognition performance. The best experimental result shows that the overall accuracy of our algorithm on the Test dataset of the EmotiW challenge is 62.78, which outperforms the best result of EmotiW2018 and ranks 2nd at the EmotiW2019 challenge. Sunan Li, Wenming Zheng, Yuan Zong, Cheng Lu 0005, Chuangao Tang, Xingxun Jiang, Jiateng Liu, Wanchuang Xia |
ICMI | 4 |
| 2019 | ℓ1-Norm Heteroscedastic Discriminant Analysis Under Mixture of Gaussian DistributionsabstractFisher’s criterion is one of the most popular discriminant criteria for feature extraction. It is defined as the generalized Rayleigh quotient of the between-class scatter distance to the within-class scatter distance. Consequently, Fisher’s criterion does not take advantage of the discriminant information in the class covariance differences, and hence, its discriminant ability largely depends on the class mean differences. If the class mean distances are relatively large compared with the within-class scatter distance, Fisher’s criterion-based discriminant analysis methods may achieve a good discriminant performance. Otherwise, it may not deliver good results. Moreover, we observe that the between-class distance of Fisher’s criterion is based on the$\ell _{2}$-norm, which would be disadvantageous to separate the classes with smaller class mean distances. To overcome the drawback of Fisher’s criterion, in this paper, we first derive a new discriminant criterion, expressed as amixture of absolute generalized Rayleigh quotients, based on a Bayes error upper bound estimation, where mixture of Gaussians is adopted to approximate the real distribution of data samples. Then, the criterion is further modified by replacing$\ell _{2}$-norm with$\ell _{1}$one to better describe the between-class scatter distance, such that it would be more effective to separate the different classes. Moreover, we propose a novel$\ell _{1}$-norm heteroscedastic discriminant analysis method based on the new discriminant analysis (L1-HDA/GM) for heteroscedastic feature extraction, in which the optimization problem of L1-HDA/GM can be efficiently solved by using the eigenvalue decomposition approach. Finally, we conduct extensive experiments on four real data sets and demonstrate that the proposed method achieves much competitive results compared with the state-of-the-art methods. Wenming Zheng, Cheng Lu 0005, Zhouchen Lin, Tong Zhang 0021, Zhen Cui 0001, Wankou Yang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Multiple Spatio-temporal Feature Learning for Video-based Emotion Recognition in the WildabstractThe difficulty of emotion recognition in the wild (EmotiW) is how to train a robust model to deal with diverse scenarios and anomalies. The Audio-video Sub-challenge in EmotiW contains audio-video short clips with several emotional labels and the task is to distinguish which label the video belongs to. For the better emotion recognition in videos, we propose a multiple spatio-temporal feature fusion (MSFF) framework, which can more accurately depict emotional information in spatial and temporal dimensions by two mutually complementary sources, including the facial image and audio. The framework is consisted of two parts: the facial image model and the audio model. With respect to the facial image model, three different architectures of spatial-temporal neural networks are employed to extract discriminative features about different emotions in facial expression images. Firstly, the high-level spatial features are obtained by the pre-trained convolutional neural networks (CNN), including VGG-Face and ResNet-50 which are all fed with the images generated by each video. Then, the features of all frames are sequentially input to the Bi-directional Long Short-Term Memory (BLSTM) so as to capture dynamic variations of facial appearance textures in a video. In addition to the structure of CNN-RNN, another spatio-temporal network, namely deep 3-Dimensional Convolutional Neural Networks (3D CNN) by extending the 2D convolution kernel to 3D, is also applied to attain evolving emotional information encoded in multiple adjacent frames. For the audio model, the spectrogram images of speech generated by preprocessing audio, are also modeled in a VGG-BLSTM framework to characterize the affective fluctuation more efficiently. Finally, a fusion strategy with the score matrices of different spatio-temporal networks gained from the above framework is proposed to boost the performance of emotion recognition complementally. Extensive experiments show that the overall accuracy of our proposed MSFF is 60.64%, which achieves a large improvement compared with the baseline and outperform the result of champion team in 2017. Cheng Lu 0005, Wenming Zheng, Chaolong Li, Chuangao Tang, Suyuan Liu, Simeng Yan, Yuan Zong |
ICMI | 1 |