EDBT 2026 Demo / reviewers in the wild / expert
Yuan Zong
dblp:79/6395
· DBLP profile ↗
95ranked-venue papers
9as first author
68since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 52 · 5 first-author · 39 since 2021Artificial intelligence and machine learning · 41 · 1 first-author · 31 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Speaker-independent speech emotion recognition using group sparse-based adversarial local fisher discriminant analysis
Cheng Lu 0005, Kaifei Zhang, Hailun Lian, Sunan Li, Tianhua Qi, Yuan Zong, Wenming Zheng |
Pattern Recognit. | 6 |
| 2026 | FAME: Frequency and motion extrapolation for robust multimodal facial action unit detection
Mengxin Shi, Cheng Lu 0005, Zhangfeng Hu, Hongli Chang, Yuan Zong |
Pattern Recognit. | 6 |
| 2026 | An Interpretable Collaborative Acoustic Parameter Modeling Network for Speech Emotion RecognitionabstractAcoustic parameters can collaboratively represent speech emotions. However, existing speech emotion recognition (SER) methods neglect explicit collaborative modeling among acoustic parameters during representation learning. This not only limits their ability to extract the information contained in acoustic parameters that collaboratively represent emotions, but also makes it difficult to trace back which acoustic parameters are closely associated with emotional states (i.e., lack of interpretability). To address these challenges, we propose the interpretable collaborative acoustic parameter modeling network (ICAPM-Net), which innovatively introduces Graph Convolutional Networks (GCN) to model the relationships among acoustic parameters. Specifically, ICAPM-Net represents different acoustic parameters as nodes, with edges characterizing their collaborative relationships. The GCN utilizes edges to explicitly model the collaborative dependencies between acoustic parameters. Importantly, visualizing the edges in the graph (i.e., adjacency matrix) can trace which acoustic parameters (or their combinations) play dominant roles in emotional expression. The overall architecture of ICAPM-Net consists of three modules: 1) an encoding module, which is used to unify feature dimensions of acoustic parameters; 2) a graph-based collaborative modeling module (GBCM), which is responsible for modeling the collaborative relationships among acoustic parameters; and 3) an emotion adjacency matrix alignment loss module (EAMAL), which is designed to further enhance collaborative modeling in GBCM. Experiments on IEMOCAP, ABC, and EMO-DB show that ICAPM-Net outperforms state-of-the-art methods. Additionally, visualization results reveal that certain acoustic parameters, e.g., MFCC and RASTA, have a strong correlation with emotions and exhibit significant collaborative patterns in different emotional contexts. Hailun Lian, Cheng Lu 0005, Hao Yang 0028, Yan Zhao 0037, Sunan Li, Yuan Zong |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2026 | Trend-Aware Multiscale Spatial-Temporal Graph Convolution Network for P300 DetectionabstractP300-based brain–computer interfaces (BCIs) enable direct communication between the brain and external devices by decoding P300 potentials, playing a vital role in rehabilitation and cognitive neuroscience research. Accurate detection of P300 potentials is essential for the successful implementation of P300-based BCIs. Current mainstream P300 detection algorithms are based on multichannel electroencephalogram (EEG) signals, and while promising detection results have been achieved, they still suffer from the following issues: 1) insufficient exploitation of the distinct and characteristic overall temporal trend of P300 potentials; and 2) oversimplified aggregation of multichannel EEG information without effectively utilizing the non-Euclidean topological relationships between EEG channels. To address the above issues, we propose a trend-aware multiscale spatial-temporal graph convolutional neural network (TMSGCN) for P300 detection. Specifically, to effectively capture the long-term temporal trend of P300 potentials to improve detection robustness, we explicitly extract the trend component from raw EEG signals along time dimension and utilize it to assist the identification of P300 potentials. Subsequently, a multiscale temporal convolution module (MTCM) is applied to extract multitime scale amplitude features from the processed EEG signals, serving as input for an adaptive graph convolution module (AGCM). The AGCM effectively models the intricate inter-channel relationships as a graph by complementarily considering the structural and functional connectivity of human brain, thereby capturing more discriminative and informative spatial features related to P300 potentials. Extensive experiments on three datasets demonstrate the superiority of TMSGCN over other state-of-the-art P300 detection methods. Furthermore, the results of ablation study and visualization experiments indicate the effectiveness of each component in TMSGCN. Jincen Wang, Wenming Zheng, Yuan Zong |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2026 | Feature Evaluation and Joint Interaction for Audio-Visual Emotion RecognitionabstractAutomatic emotion recognition has attracted significant attention due to its potential applications in various real-world scenarios. Methods that integrate visual and audio modalities have become increasingly prominent because of their superior information-carrying capacity and complementarity. Despite advancements in feature fusion between video and audio, existing modality fusion-based methods struggle to effectively address the dynamic changes in feature quality caused by interference, which is common in emotion recognition in the wild tasks. To overcome this limitation, we propose a Parameter-Free Feature Evaluation and Interaction (PFFEI) model based on information quality assessment. The model leverages the scaling factor γ of the normalization layer to evaluate information quality and dynamically adjusts the degree of interaction between modalities, suppressing the impact of low-quality features affected by interference. Additionally, the norm constraint integrated into the model ensures that the γ value consistently measures feature quality across different modalities. This approach effectively mitigates the effects of modality imbalance and significantly enhances the model’s accuracy. The effectiveness of our method is demonstrated through experiments on three challenging real-world emotion datasets: DFEW, AFEW, and Ekman6. The results show that the PFFEI model outperforms state-of-the-art methods, achieving significant improvements of 8.71% (UAR) and 8.61% (WAR) on the AFEW database. Sunan Li, Cheng Lu 0005, Yuan Zong, Hailun Lian, Wenming Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Fine-Grained Recognition of Arteriovenous Fistula Stenosis Using Blood Flow Sounds: An Animal Model-Based Dataset and a Frequency-Aware Decoupling Network
Shanlin Xiao, Yangyi Zhou, Jincen Wang, Yuan Zong |
ICANN (4) | 6 |
| 2025 | Enhancing Task-Specific Feature Learning with LLMs for Multimodal Emotion and Intent Joint UnderstandingabstractThis paper introduces our solution, the Task-Specific Feature Learning (TSFL) method, designed to address the second track of the MEIJU Challenge at ICASSP 2025, namely, Imbalanced Emotion and Intent Recognition (English). The TSFL method incorporates three core components: the use of LLM features to represent multimodal signals, coarse-grained task-specific feature decomposition, and fine-grained task-specific feature learning. These components enable the effective joint learning of emotion-discriminative and intent-discriminative features. As a result, our method achieved a JRBM score of 0.6230, significantly outperforming the official baseline result and surpassing all other competing teams to win the championship. Cheng Lu 0005, Kaifei Zhang, Yujia Gu, Banghua Li, Yuan Zong, Wenming Zheng |
ICASSP | 7 |
| 2025 | Reliable Learning From LLM Features for Multimodal Emotion and Intent Joint UnderstandingabstractThis paper describes a Reliable Learning Framework (RLF) for the 1st Multimodal Emotion and Intent Joint Understanding (MEIJU) Challenge at ICASSP 2025. Our proposed RLF includes a Hierarchical Interaction Network and a Reliable Fusion Strategy. The former can excavate emotion and intent cues from the high-level semantic features of multimodal data (video, audio, and text) generated by pretrained Large Language Models (LMMs), to enhance their representations, and the latter reliably integrates multiple predictions to further improve the robustness of emotion and intent understanding. Our RLF method achieved first place on Track 2 (Mandarin) of MEIJU, with performance scores for emotion, intent, and joint recognition reaching 0.7285, 0.7456, and 0.7370. Cheng Lu 0005, Yuyun Liu, Yinghao Ma, Jiahao Luo, Yuan Zong, Wenming Zheng |
ICASSP | 7 |
| 2025 | Interactive Fusion of Multi-View Speech Embeddings via Pretrained Large-Scale Speech Models for Speech Emotional Attribute Prediction in Naturalistic Conditions
Yuyun Liu, Yujia Gu, Jiahao Luo, Wenming Zheng, Cheng Lu 0005, Yuan Zong |
INTERSPEECH | 6 |
| 2025 | FEALLM: Advancing Facial Emotion Analysis in Multimodal Large Language Models with Emotional Synergy and Reasoning
Zhuozhao Hu, Kaishen Yuan, Xin Liu 0012, Zitong Yu, Yuan Zong, Jingang Shi, Huanjing Yue, Jing-Yu Yang 0002 |
ACM Multimedia | 5 |
| 2025 | NaME: A Natural Micro-expression Dataset for Micro-expression Recognition in the WildabstractMicro-expressions (MEs) are involuntary facial expressions that reveal genuine emotions and have significant applications in fields such as psychology, security, and human-computer interaction. However, previous ME datasets are mainly collected in controlled laboratory environments, such as fixed views, single illumination and head movements, limited subjects and the lack of background. There are significant gaps between them and the real world. To handle this issue, we introduce a novel Natural Micro-Expression (NaME) dataset, a natural dataset collected under unconstrained real-world conditions. It encompasses (1) diverse subjects, multiple views and varying head movements ; (2) rich background information, providing a more realistic benchmark for the micro-expression recognition (MER) research. Furthermore, we propose a MER benchmark for natural environments, named MixFormer. MixFormer includes an efficient sparse attention mechanism to capture subtle facial motions from various factors, and a face-background mix of attention module to model the environment context to help MER. Extensive experiments are conducted to analyze our NaME dataset and benchmark. We believe that our dataset and benchmark will pave the way for future research in MER beyond controlled settings, facilitating the deployment of MER in practical applications. NaME is available at github.com/real-ljt/NAMEdataset. Jiateng Liu, Hengcan Shi, Haiwen Liang, Yuan Zong, Yaonan Wang 0001, Wenming Zheng |
ACM Multimedia | 5 |
| 2025 | Multi-Level Segment Fusion Based on Adaptive Time-Window Selection for Multimodal Personality-Aware Elderly Depression DetectionabstractMajor Depressive Disorder (MDD) is a prevalent and severe psychiatric disorder, and its detection remains challenging due to the complexity and variability of its symptoms. Traditional single-modality methods often fail to capture the full spectrum of depressive cues, which has led to the rise of multimodal methods. The ACM Multimedia 2025 ''Multimodal Personality-Aware Depression Detection Challenge'' (MPDD 2025) aims to advance the development of more accurate depression detection models by incorporating multimodal data. In this paper, we proposed a Multi-Level Segment Fusion Based on Adaptive Time-Window Selection (MSF-ATS) method for the MPDD-Elderly Track. To address the challenge of sparse and transient depressive symptoms, we fuse segment-level classifications to obtain subject-level classifications. An adaptive time-window selection based on mean class variance is employed to choose the window with the smallest variance for more stable detection results. Our method achieved an average score of 0.8576 on the MPDD 2025 official test set, significantly outperforming the baseline score of 0.6675. Yuyun Liu, Kaifei Zhang, Yinghao Ma, Tianhua Qi, Wenming Zheng, Cheng Lu 0005, Yuan Zong |
ACM Multimedia | 8 |
| 2025 | Assessing Personality Traits and Interview Performance from Asynchronous Video InterviewsabstractAsynchronous Video Interviews (AVIs) allow candidates to record responses to predefined questions using digital devices, offering both flexibility and remote accessibility. Assessing personality traits and interview performance via AVIs provides organizations with valuable insights into candidate profiles and facilitates the prediction of future job performance. However, prior benchmark challenges, whose datasets were predominantly sourced from social media, suffer from suboptimal construct and methodological validity, limiting their utility for model development and real-world applications. To address these limitations, we introduce the AVI Grand Challenge at ACM Multimedia 2025, featuring a novel dataset of mock AVIs comprising 3,876 videos from 646 participants in a simulated job application procedure. Interview questions were carefully designed to reflect real-world selection contexts and elicit personality expressions grounded in Trait Activation Theory. Personality traits and job competencies were annotated by trained evaluators and professional recruiters, ensuring both methodological rigor and ecological validity. The solutions and algorithms developed in this challenge are analyzed and summarized in this paper to foster the development of fair, reliable, and AI-driven hiring assessments. Tianyi Zhang 0013, Tianhua Qi, Antonis Koutsoumpis, Yuan Zong, Wenming Zheng, Janneke K. Oostrom, Djurre Holtrop, Zhaojie Luo, Reinout E. de Vries |
ACM Multimedia | 4 |
| 2025 | Low-rank joint distribution adaptation for cross-corpus speech emotion recognition
Sunan Li, Cheng Lu 0005, Yan Zhao 0037, Hailun Lian, Tianhua Qi, Yuan Zong |
Knowl. Based Syst. | 6 |
| 2025 | AMGCN: An adaptive multi-graph convolutional network for speech emotion recognition
Hailun Lian, Cheng Lu 0005, Hongli Chang, Yan Zhao 0037, Sunan Li, Yang Li 0019, Yuan Zong |
Speech Commun. | 7 |
| 2025 | RF-REN: RGB-Frequency Relation Exploration Network for Micro-Expression RecognitionabstractMicro-expression recognition (MER) has drawn increasing attention in recent years due to its ability to reveal the true feelings people want to hide. The key challenge in MER is subtle motions, which are hard to capture but crucial for MER. Existing methods usually solve this problem by magnifying all motions in the whole face and temporal sequence. However, micro-expressions (MEs) only involve a few facial areas and several temporal snippets. The all-motion magnification in previous methods cannot precisely capture these local ME motion patterns, and can easily cause spatial as well as temporal distortions, which significantly decrease the MER accuracy. In this paper, we propose an RGB-Frequency Relation Exploration Network (RF-REN), which enhances the subtle motions in refined local ME cues by exploring spatial and temporal relations in both RGB and frequency domains. Specifically, we first decompose the ME video into RGB as well as frequency domains, and conduct temporal division according to different motion stages to cover various ME local patterns. Secondly, we construct an adaptive local-global relation exploration (LGRE) module to explore the local relation cues in the spatial appearance and temporal dynamics in both domains. Finally, we propose an RGB-Frequency routing strategy to fuse the RGB and frequency cues, aiming to aggregate spatial-temporal local-global information and enhance subtle motions for MER. Extensive experiments on three databases (CASME II, SAMM and SMIC) show that the proposed model outperforms other state-of-the-art methods. Jiateng Liu, Hengcan Shi, Yaonan Wang 0001, Yuan Zong |
IEEE Signal Process. Lett. | 4 |
| 2025 | Learning to Rank Onset-Occurring-Offset Representations for Micro-Expression RecognitionabstractThis paper focuses on the research of micro-expression recognition (MER) and proposes a flexible and reliable deep learning method called learning to rank onset-occurring-offset representations (LTR3O). The LTR3O method introduces a dynamic and reduced-size sequence structure known as 3O, which consists of onset, occurring, and offset frames, for representing micro-expressions (MEs). This structure facilitates the subsequent learning of ME-discriminative features. A noteworthy advantage of the 3O structure is its flexibility, as the occurring frame is randomly extracted from the original ME sequence without the need for accurate frame spotting methods. Based on the 3O structures, LTR3O generates multiple 3O representation candidates for each ME sample and incorporates well-designed modules based on learning to rank (LTR) to measure and calibrate their emotional expressiveness. This calibration process implicitly enhances the visibility of MEs by amplifying the originally narrow emotional expressiveness gap among ME frames caused by their low-intensity characteristics, thereby facilitating the reliable learning of more discriminative features for MER. Extensive experiments were conducted to evaluate the performance of LTR3O using four widely-used ME databases: CASME II, SMIC, SAMM, and MEVIEW. The experimental results demonstrate the effectiveness and superior performance of LTR3O, particularly in terms of its flexibility and reliability, when compared to recent state-of-the-art MER methods. Yuan Zong, Jingang Shi, Cheng Lu 0005, Hongli Chang, Wenming Zheng |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | Towards Domain-Specific Cross-Corpus Speech Emotion Recognition ApproachabstractCross-corpus speech emotion recognition (SER) poses a challenge due to feature distribution mismatch between the training and testing speech samples, potentially degrading the performance of established SER methods. In this article, we tackle this challenge by proposing a novel transfer subspace learning method called acoustic knowledge-guided transfer linear regression (AKTLR). Unlike existing approaches, which often overlook domain-specific knowledge related to SER and simply treat cross-corpus SER as a generic transfer learning task, our AKTLR method is built upon a well-designed acoustic knowledge-guided dual sparsity constraint mechanism. This mechanism emphasizes the potential of minimalistic acoustic parameter feature sets to alleviate classifier over-adaptation, which is empirically validated acoustic knowledge in SER, enabling superior generalization in cross-corpus SER tasks compared to using large feature sets. Through this mechanism, we extend a simple transfer linear regression model to AKTLR. This extension harnesses its full capability to seek emotion-discriminative and corpus-invariant features from established acoustic parameter feature sets used for describing speech signals across two scales: contributive acoustic parameter groups and constituent elements within each contributive group. We evaluate our method through extensive cross-corpus SER experiments on three widely used speech emotion corpora: EmoDB, eNTERFACE, and CASIA. The proposed AKTLR achieves an average UAR of 42.12% across six tasks using the eGeMAPS feature set, outperforming many recent state-of-the-art transfer subspace learning and deep transfer learning methods. This demonstrates the effectiveness and superior performance of our approach. Furthermore, our work provides experimental evidence supporting the feasibility and superiority of incorporating domain-specific knowledge into the transfer learning model to address cross-corpus SER tasks. Yan Zhao 0037, Yuan Zong, Hailun Lian, Cheng Lu 0005, Jingang Shi, Wenming Zheng |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2025 | Decoupled Doubly Contrastive Learning for Cross-Domain Facial Action Unit DetectionabstractDespite the impressive performance of current vision-based facial action unit (AU) detection approaches, they are heavily susceptible to the variations across different domains and the cross-domain AU detection methods are under-explored. In response to this challenge, we propose a decoupled doubly contrastive adaptation (D2CA) approach to learn a purified AU representation that is semantically aligned for the source and target domains. Specifically, we decompose latent representations into AU-relevant and AU-irrelevant components, with the objective of exclusively facilitating adaptation within the AU-relevant subspace. To achieve the feature decoupling, D2CA is trained to disentangle AU and domain factors by assessing the quality of synthesized faces in cross-domain scenarios when either AU or domain attributes are modified. To further strengthen feature decoupling, particularly in scenarios with limited AU data diversity, D2CA employs a doubly contrastive learning mechanism comprising image and feature-level contrastive learning to ensure the quality of synthesized faces and mitigate feature ambiguities. This new framework leads to an automatically learned, dedicated separation of AU-relevant and domain-relevant factors, and it enables intuitive, scale-specific control of the cross-domain facial image synthesis. Extensive experiments demonstrate the efficacy of D2CA in successfully decoupling AU and domain factors, yielding visually pleasing cross-domain synthesized facial images. Meanwhile, D2CA consistently outperforms state-of-the-art cross-domain AU detection approaches, achieving an average F1 score improvement of 6%-14% across various cross-domain scenarios. Yong Li 0032, Menglin Liu, Zhen Cui 0001, Yi Ding 0012, Yuan Zong, Wenming Zheng, Shiguang Shan, Cuntai Guan |
IEEE Trans. Image Process. | 5 |
| 2024 | Improving Speaker-Independent Speech Emotion Recognition using Dynamic Joint Distribution AdaptationabstractIn speaker-independent speech emotion recognition, the training and testing samples are collected from diverse speakers, leading to a multi-domain shift challenge across the feature distributions of data from different speakers. Consequently, when the trained model is confronted with data from new speakers, its performance tends to degrade. To address the issue, we propose a Dynamic Joint Distribution Adaptation (DJDA) method under the framework of multi-source domain adaptation. DJDA firstly utilizes joint distribution adaptation (JDA), involving marginal distribution adaptation (MDA) and conditional distribution adaptation (CDA), to more precisely measure the multi-domain distribution shifts caused by different speakers. This helps eliminate speaker bias in emotion features, allowing for learning discriminative and speaker-invariant speech emotion features from coarse-level to fine-level. Furthermore, we quantify the adaptation contributions of MDA and CDA within JDA by using a dynamic balance factor based on $\mathcal{A}$-Distance, promoting to effectively handle the unknown distributions encountered in data from new speakers. Experimental results demonstrate the superior performance of our DJDA as compared to other state-of-the-art (SOTA) methods. Cheng Lu 0005, Yuan Zong, Hailun Lian, Yan Zhao 0037, Björn W. Schuller, Wenming Zheng |
ICASSP | 2 |
| 2024 | PAVITS: Exploring Prosody-Aware VITS for End-to-End Emotional Voice ConversionabstractIn this paper, we propose Prosody-aware VITS (PAVITS) for emotional voice conversion (EVC), aiming to achieve two major objectives of EVC: high content naturalness and high emotional naturalness, which are crucial for meeting the demands of human perception. To improve the content naturalness of converted audio, we have developed an end-to-end EVC architecture inspired by the high audio quality of VITS. By seamlessly integrating an acoustic converter and vocoder, we effectively address the common issue of mismatch between emotional prosody training and run-time conversion that is prevalent in existing EVC models. To further enhance the emotional naturalness, we introduce an emotion descriptor to model the subtle prosody variations of different speech emotions. Additionally, we propose a prosody predictor, which predicts prosody features from text based on the provided emotion label. Notably, we introduce a prosody alignment loss to establish a connection between latent prosody features from two distinct modalities, ensuring effective training. Experimental results show that the performance of PAVITS is superior to the state-of-the-art EVC methods. Speech Samples are available at https://jeremychee4.github.io/pavits4EVC/. Tianhua Qi, Wenming Zheng, Cheng Lu 0005, Yuan Zong, Hailun Lian |
ICASSP | 4 |
| 2024 | Speech Swin-Transformer: Exploring a Hierarchical Transformer with Shifted Windows for Speech Emotion RecognitionabstractSwin-Transformer has demonstrated remarkable success in computer vision by leveraging its hierarchical feature representation based on Transformer. In speech signals, emotional information is distributed across different scales of speech features, e. g., word, phrase, and utterance. Drawing above inspiration, this paper presents a hierarchical speech Transformer with shifted windows to aggregate multi-scale emotion features for speech emotion recognition (SER), called Speech Swin-Transformer. Specifically, we first divide the speech spectrogram into segment-level patches in the time domain, composed of multiple frame patches. These segment-level patches are then encoded using a stack of Swin blocks, in which a local window Transformer is utilized to explore local inter-frame emotional information across frame patches of each segment patch. After that, we also design a shifted window Transformer to compensate for patch correlations near the boundaries of segment patches. Finally, we employ a patch merging operation to aggregate segment-level emotional features for hierarchical speech representation by expanding the receptive field of Transformer from frame-level to segment-level. Experimental results demonstrate that our proposed Speech Swin-Transformer outperforms the state-of-the-art methods. Yong Wang 0073, Cheng Lu 0005, Hailun Lian, Yan Zhao 0037, Björn W. Schuller, Yuan Zong, Wenming Zheng |
ICASSP | 6 |
| 2024 | Progressively Learning from Macro-Expressions for Micro-Expression RecognitionabstractMicro-expression (ME) recognition is challenging due to the low-intensity facial motions. An idea to overcome this is learning assisted by macro-expressions (MaEs). However, the intensity gap between MaE and ME is so huge that related works fail to effectively leverage MaE’s assistance in overcoming low-intensity interference, which attempt to directly force ME knowledge to mimic MaE knowledge. In this paper, we propose that the knowledge transfer from MaE to ME can be converted into a progressive process for better implementation. Thus, we construct a progressive multi-step learning framework, which accomplishes two tasks: first, we dissect the huge intensity gap into multiple segments that are easier to bridge by constructing multiple learning steps, each corresponding to various intensity levels of expression recognition tasks. Second, through a designed self-knowledge distillation (self-KD) model, each dissected gap can be bridged, enabling the MaE knowledge to progressively transfer to guide the ME learning. Experiments carried out on three widely used databases demonstrated that the proposed PLMaM achieves state-of-the-art results. Yuan Zong, Mengting Wei, Cheng Lu 0005, Wenming Zheng |
ICASSP | 2 |
| 2024 | Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion RecognitionabstractCross-corpus speech emotion recognition (SER) aims to transfer emotional knowledge from a labeled source corpus to an unlabeled corpus. However, prior methods require access to source data during adaptation, which is unattainable in real-life scenarios due to data privacy protection concerns. This paper tackles a more practical task, namely source-free cross-corpus SER, where a pre-trained source model is adapted to the target domain without access to source data. To address the problem, we propose a novel method called emotion-aware contrastive adaptation network (ECAN). The core idea is to capture local neighborhood information between samples while considering the global class-level adaptation. Specifically, we propose a nearest neighbor contrastive learning to promote local emotion consistency among features of highly similar samples. Furthermore, relying solely on nearest neighborhoods may lead to ambiguous boundaries between clusters. Thus, we incorporate supervised contrastive learning to encourage greater separation between clusters representing different emotions, thereby facilitating improved class-level adaptation. Extensive experiments indicate that our proposed ECAN significantly outperforms state-of-the-art methods under the source-free cross-corpus SER setting on several speech emotion corpora. Yan Zhao 0037, Jincen Wang, Cheng Lu 0005, Sunan Li, Björn W. Schuller, Yuan Zong, Wenming Zheng |
ICASSP | 6 |
| 2024 | A Novel Decoupled Prototype Completion Network for Incomplete Multimodal Emotion RecognitionabstractReconstructing missing modality based on available modalities is widely used to address inevitable modality-missing for Multimodal Emotion Recognition (MER). However, due to explicit distribution gap across heterogeneous modalities, they fail to guarantee the consistency between the reconstructed data and the ground truth. To mitigate this problem, we propose a novel method to restore the missing modality using its weighted prototypes rather than other modalities. Specifically, prototypes of different classes of missing modality are used to encapsulate its representative knowledge. Then sample-to-prototype affinity measuring class similarity is used as weights to combine these prototypes for reconstruction, thereby effectively restoring the distribution-consistent modality. Furthermore, to improve the efficacy of prototype-based completion under seriously missing, we devise an adaptive knowledge distillation from the strong modality to the weaker ones. This reinforces the representation ability of weak modality features. Extensive experiments on CMU-MOSI and IEMOCAP datasets demonstrate the superiority of our method. Zhangfeng Hu, Wenming Zheng, Yuan Zong, Mengting Wei, Xingxun Jiang, Mengxin Shi |
ICME | 3 |
| 2024 | Missing Customized Distillation Network for Incomplete Multimodal Sentiment Analysis
Zhangfeng Hu, Wenming Zheng, Mengting Wei, Mengxin Shi, Yuan Zong |
ICPR (8) | 5 |
| 2024 | Hierarchical Distribution Adaptation for Unsupervised Cross-corpus Speech Emotion Recognition
Cheng Lu 0005, Yuan Zong, Yan Zhao 0037, Hailun Lian, Tianhua Qi, Björn W. Schuller, Wenming Zheng |
INTERSPEECH | 2 |
| 2024 | Towards Realistic Emotional Voice Conversion using Controllable Emotional Intensity
Tianhua Qi, Shiyan Wang, Cheng Lu 0005, Yan Zhao 0037, Yuan Zong, Wenming Zheng |
INTERSPEECH | 5 |
| 2024 | Boosting Cross-Corpus Speech Emotion Recognition using CycleGAN with Contrastive Learning
Jincen Wang, Yan Zhao 0037, Cheng Lu 0005, Chuangao Tang, Sunan Li, Yuan Zong, Wenming Zheng |
INTERSPEECH | 6 |
| 2024 | Confidence-aware Hypothesis Transfer Networks for Source-Free Cross-Corpus Speech Emotion Recognition
Jincen Wang, Yan Zhao 0037, Cheng Lu 0005, Hailun Lian, Hongli Chang, Yuan Zong, Wenming Zheng |
INTERSPEECH | 6 |
| 2024 | A Novel Non-Stationary Channel Emulator for 6G MIMO Wireless ChannelsabstractThe performance evaluation of sixth generation (6G) communication systems is anticipated to be a controlled and repeatable process in the lab, which brings up the demand for wireless channel emulators. However, channel emulation for 6G space-time-frequency (STF) non-stationary channels is missing currently. In this paper, a non-stationary multiple-input multiple-output (MIMO) geometry-based stochastic model (GBSM) that accurately characterizes the channel STF properties is introduced firstly. Then, a subspace-based method is proposed for reconstructing the channel fading obtained from the GBSM and a channel emulator architecture with frequency domain processing is presented for 6G MIMO systems. Moreover, the spatial time-varying channel transfer functions (CTFs) of the channel simulation and the channel emulation are compared and analyzed. The Doppler power spectral density (PSD) and delay PSD are further derived and compared between the channel model simulation and subspace-based emulation. The results demonstrate that the proposed channel emulator is capable of reproducing the non-stationary channel characteristics. Yuan Zong, Lijian Xin, Jie Huang 0004, Cheng-Xiang Wang 0001 |
WCNC | 1 |
| 2024 | Exploring corpus-invariant emotional acoustic feature for cross-corpus speech emotion recognition
Hailun Lian, Cheng Lu 0005, Yan Zhao 0037, Sunan Li, Tianhua Qi, Yuan Zong |
Expert Syst. Appl. | 6 |
| 2024 | Exploring holistic discriminative representation for micro-expression recognition via contrastive learning
Wanyuan He, Hongli Chang, Cheng Lu 0005, Yuan Zong |
Image Vis. Comput. | 6 |
| 2024 | CFEW: A Large-Scale Database for Understanding Child Facial Expression in Real WorldabstractCurrently, much progress has been achieved on adult facial expressions recognition. Few attentions have been paid to child facial expression analysis. A lack of publicly available large-scale child facial expression databases hinders the development of automatic coding for child facial expression behaviors. In this work, we constructed a new face database for understandingChildFacialExpression in realWorld (CFEW). The database contains three novelties: (1) the largest publicly available child facial expression database (11,000+ images); (2) covering full developmental range of 0–18-year-old child subjects; (3) rich annotations for facial expression labels, including discrete expression categories aka happy, neutral, disgust, angry, sad, cry, fear, surprise, sleepy and others, intensity of arousal and valence, and several types of facial action units (AUs). In addition, the images in this database cover several challenging conditions in real world, including frontal and non-frontal head poses, facial occlusions, various illuminations and low image resolution. Three dominant deep convolutional neural networks (i.e., VGG11bn, ResNet18 and DenseNet121) were used to conduct extensive baseline experiments for discrete facial expression classification, arousal and valence estimation and facial action units detection within database, and cross-database seven facial expressions recognition. Chuangao Tang, Sunan Li, Wenming Zheng, Yuan Zong, Su Zhang 0004, Cheng Lu 0005, Yan Zhao 0037 |
IEEE Trans. Affect. Comput. | 4 |
| 2024 | Layer-Adapted Implicit Distribution Alignment Networks for Cross-Corpus Speech Emotion RecognitionabstractIn this article, we propose a new unsupervised domain adaptation (DA) method called layer-adapted implicit distribution alignment networks (LIDANs) to address the challenge of cross-corpus speech emotion recognition (SER). LIDAN extends our previous ICASSP work, deep implicit distribution alignment networks (DIDANs), whose key contribution lies in the introduction of a novel regularization term called implicit distribution alignment (IDA). This term allows DIDAN trained on source (training) speech samples to remain applicable to predicting emotion labels for target (testing) speech samples, regardless of corpus variance in cross-corpus SER. To further enhance this method, we extend IDA to layer-adapted IDA (LIDA), resulting in LIDAN. This layer-adapted extension consists of three modified IDA terms that consider emotion labels at different levels of granularity. These terms are strategically arranged within different fully connected layers in LIDAN, aligning with the increasing emotion-discriminative abilities with respect to the layer depth. This arrangement enables LIDAN to more effectively learn emotion-discriminative and corpus-invariant features for SER across various corpora compared to DIDAN. It is also worthy to mention that unlike most existing methods that rely on estimating statistical moments to describe preassumed explicit distributions, both IDA and LIDA take a different approach. They utilize an idea of target sample reconstruction to directly bridge the feature distribution gap without making assumptions about their distribution type. As a result, DIDAN and LIDAN can be viewed as implicit cross-corpus SER methods. To evaluate LIDAN, we conducted extensive cross-corpus SER experiments on EmoDB, eNTERFACE, and CASIA corpora. The experimental results demonstrate that LIDAN surpasses recent state-of-theart explicit unsupervised DA methods in tackling cross-corpus SER tasks. Yan Zhao 0037, Yuan Zong, Jincen Wang, Hailun Lian, Cheng Lu 0005, Li Zhao 0003, Wenming Zheng |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2024 | Adaptive Multi-Scale Iterative Optimized Video Object Segmentation Based on Correlation EnhancementabstractSemi-supervised video object segmentation (VOS) is a highly challenging task, which relies on the initial frame’s mask as a segmentation reference in a video sequence to classify each pixel in subsequent frames. However, the guidance provided by the first frame is limited due to the diverse types of segmentation targets and uncertain appearance changes. Consequently, it is crucial to retain useful information during the segmentation process and employ this information for model iteration optimization, enabling the model to better adapt to rapidly changing segmentation objectives. In this work, we propose a multi-scale adaptive model optimization strategy, which incorporates a contextual relevance enhancement module to enforce object correlation by emphasizing feature similarity across adjacent frames. Additionally, we introduce a keyframe discrimination module to deal with the segmentation challenges in scenarios involving significant target changes. Moreover, we also introduce a multi-scale memory screening module to automatically screen and select global-local optimization features for ensuring the model’s generalization performance. Extensive experiments show that the proposed method achieves state-of-the-art performance on DAVIS and large-scale Youtube-VOS 2018/2019 datasets without relying on synthetic training data or first-frame fine-tuning. Yuan Zong, Wenming Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | MaskFusionNet: A Dual-Stream Fusion Model With Masked Pre-Training Mechanism for rPPG MeasurementabstractRemote photoplethysmography (rPPG) has considerable significance in areas such as disease diagnosis and emotion analysis. Recent rPPG models have demonstrated excellent performance due to their powerful heart rate information extraction capabilities. However, these models often focus on limited regions of interest (ROI) on facial image, which makes them sensitive to interference. If the ROI is affected by muscle movement, lighting variation and noise, the model’s performance would degrade significantly. To address this limitation, we propose a two-stage model called MaskFusionNet. The model includes two stages: 1) During the pre-training stage, the mask-reconstruction mechanism drives MaskFusionNet to learn rPPG information from various facial regions by applying a tube masking strategy. This enhances the model’s ability to resist interference. Based on the periodicity and continuity of the heart rate signal, we also design a novel spatio-temporal reconstruction loss function that focuses on the data’s spatial features and temporal continuity. 2) In the fine-tuning stage, we propose the Multi-Scale Fusion Block (MFB) to combine multi-scale features from the dual-stream network. It allows the model to detect subtle heart rate variations in adjacent frames while minimizing the impact of interference by extracting features within longer segments. The transformer-based MaskFusionNet can extract multi-scale fused heart rate features from a wide range of skin regions while preserving the modeling capability of long-range sequence information. To validate its effectiveness, we extensively evaluate our model on three benchmark datasets (VIPL-HR, COHFACE, and PURE), demonstrating its superior performance in both intra-dataset and cross-dataset testing scenarios. Yizhu Zhang, Jingang Shi, Yuan Zong, Wenming Zheng, Guoying Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Convolutional Transformer-Based Cross Subject Model for SSVEP-Based BCI ClassificationabstractSteady-state visual evoked potential (SSVEP) is a commonly used brain-computer interface (BCI) paradigm. The performance of cross-subject SSVEP classification has a strong impact on SSVEP-BCI. This study designed a cross subject generalization SSVEP classification model based on an improved transformer structure that uses domain generalization (DG). The global receptive field of multi-head self-attention is used to learn the global generalized SSVEP temporal information across subjects. This is combined with a parallel local convolution module, designed to avoid oversmoothing the oscillation characteristics of temporal SSVEP data and better fit the feature. Moreover, to improve the cross-subject calibration-free SSVEP classification performance, an DG method named StableNet is combined with the proposed convolutional transformer structure to form the DG-Conformer method, which can eliminate spurious correlations between SSVEP discriminative information and background noise to improve cross-subject generalization. Experiments on two public datasets, Benchmark and BETA, demonstrated the outstanding performance of the proposed DG-Conformer compared with other calibration-free methods, FBCCA, tt-CCA, Compact-CNN, FB-tCNN, and SSVEPNet. Additionally, DG-Conformer outperforms the classic calibration-required algorithms eCCA, eTRCA and eSSCOR when calibration is used. An incomplete partial stimulus calibration scheme was also explored on the Benchmark dataset, and it was demonstrated to be a potential solution for further high-performance personalized SSVEP-BCI with quick calibration. Yuankui Yang, Yuan Zong, Yue Leng, Wenming Zheng, Sheng Ge |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | Novel Sinusoidal Signal Assisted Multivariate Variational Mode Decomposition Combined With Task-Related Component Analysis for Enhancing SSVEP-Based BCI PerformanceabstractBrain-computer interfaces (BCIs) based on steady-state visually evoked potential (SSVEP) have a broad application prospect owing to their multiple command output and high performance. Each harmonic component of SSVEP individually contains unique features, which can be utilized to enhance the recognition performance of SSVEP-based BCIs. However, the existing subband analysis methods for SSVEP, including those based on filter banks and existing mode decomposition methods, have limitations in extracting and utilizing independent harmonic components. This study proposes a sinusoidal signal assisted multivariate variational mode decomposition (SA-MVMD) algorithm that allows the constraint of the center frequencies and narrowband filtering structures of the intrinsic mode functions (IMFs) based on the prior frequency knowledge of the signal. It preserves the target information of the signal during decomposition while avoiding mode mixing and incorrect decomposition, thereby enabling the effective extraction of each independent harmonic component of SSVEP. Building on this, a SA-MVMD based task-related component analysis (SA-MVMD-TRCA) method is further proposed to fully utilize the features within the overall SSVEP as well as its independent harmonics, thereby enhancing the recognition performance. Testing on the public SSVEP Benchmark dataset demonstrates that the proposed method significantly outperforms the filter bank-based control methods. This study confirms the effectiveness of SA-MVMD and the potential of this approach, which analyzes and utilizes each independent harmonic of SSVEP, providing new strategies and perspectives for performance enhancement in SSVEP-based BCIs. Jinpeng Lyu, Yuankui Yang, Yuan Zong, Yue Leng, Wenming Zheng, Sheng Ge |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | CMNet: Contrastive Magnification Network for Micro-Expression RecognitionabstractMicro-Expression Recognition (MER) is challenging because the Micro-Expressions' (ME) motion is too weak to distinguish. This hurdle can be tackled by enhancing intensity for a more accurate acquisition of movements. However, existing magnification strategies tend to use the features of facial images that include not only intensity clues as intensity features, leading to the intensity representation deficient of credibility. In addition, the intensity variation over time, which is crucial for encoding movements, is also neglected. To this end, we provide a reliable scheme to extract intensity clues while considering their variation on the time scale. First, we devise an Intensity Distillation (ID) loss to acquire the intensity clues by contrasting the difference between frames, given that the difference in the same video lies only in the intensity. Then, the intensity clues are calibrated to follow the trend of the original video. Specifically, due to the lack of truth intensity annotation of the original video, we build the intensity tendency by setting each intensity vacancy an uncertain value, which guides the extracted intensity clues to converge towards this trend rather some fixed values. A Wilcoxon rank sum test (Wrst) method is enforced to implement the calibration. Experimental results on three public ME databases i.e. CASME II, SAMM, and SMIC-HS validate the superiority against state-of-the-art methods. Mengting Wei, Xingxun Jiang, Wenming Zheng, Yuan Zong, Cheng Lu 0005, Jiateng Liu |
AAAI | 4 |
| 2023 | Deep Implicit Distribution Alignment Networks for cross-Corpus Speech Emotion RecognitionabstractIn this paper, we propose a novel deep transfer learning method called deep implicit distribution alignment networks (DIDAN) to deal with cross-corpus speech emotion recognition (SER) problem, in which the labeled training (source) and unlabeled testing (target) speech signals come from different corpora. Specifically, DIDAN first adopts a simple deep regression network consisting of a set of convolutional and fully connected layers to directly regress the source speech spectrums into the emotional labels such that the proposed DIDAN can own the emotion discriminative ability. Then, such ability is transferred to be also applicable to the target speech samples regardless of corpus variance by resorting to a well-designed regularization term called implicit distribution alignment (IDA). Unlike widely-used maximum mean discrepancy (MMD) and its variants, the proposed IDA absorbs the idea of sample reconstruction to implicitly align the distribution gap, which enables DIDAN to learn both emotion discriminative and corpus invariant features from speech spectrums. To evaluate the proposed DIDAN, extensive cross-corpus SER experiments on widely-used speech emotion corpora are carried out. Experimental results show that the proposed DIDAN can outperform lots of recent state-of-the-art methods in coping with the cross-corpus SER tasks. Yan Zhao 0037, Jincen Wang, Yuan Zong, Wenming Zheng, Hailun Lian, Li Zhao 0003 |
ICASSP | 3 |
| 2023 | Geometric Magnification-based Attention Graph Convolutional Network for Skeleton-based Micro-Gesture RecognitionabstractMicro-Gesture (MG) recognition is an emerging and challenging task due to the short duration and small amplitude of joints. MGs indicate subtle movements of the body in response to stress, which are more difficult to recognize than regular gestures. To solve the above problems, for the modeling of micro-gesture skeleton data, we propose a Geometric Magnification-Based Attention Graph Convolutional Network (MA-GCN) to magnify and select features. The network mainly consists of two modules: the geometric magnification module (GM module) controls the magnification of different joints, and the spatial temporal attention graph convolution module (STA module) selects valid information by weighting different joints and frames to focus on subtle movements. Extensive experiments on two MG datasets prove that our method achieves remarkable performance. Haolin Jiang, Wenming Zheng, Yuan Zong, Xingxun Jiang, Yunlong Xue |
ICIP | 3 |
| 2023 | Time-Frequency Transformer: A Novel Time Frequency Joint Learning Method for Speech Emotion Recognition
Yong Wang 0073, Cheng Lu 0005, Yuan Zong, Hailun Lian, Yan Zhao 0037, Sunan Li |
ICONIP (9) | 3 |
| 2023 | Learning Attention from Attention: Efficient Self-Refinement Transformer for Face Super-ResolutionabstractRecently, Transformer-based architecture has been introduced into face super-resolution task due to its advantage in capturing long-range dependencies. However, these approaches tend to integrate global information in a large searching region, which neglect to focus on the most relevant information and induce blurry effect by the irrelevant textures. Some improved methods simply constrain self-attention in a local window to suppress the useless information. But it also limits the capability of recovering high-frequency details when flat areas dominate the local searching window. To improve the above issues, we propose a novel self-refinement mechanism which could adaptively achieve texture-aware reconstruction in a coarse-to-fine procedure. Generally, the primary self-attention is first conducted to reconstruct the coarse-grained textures and detect the fine-grained regions required further compensation. Then, region selection attention is performed to refine the textures on these key regions. Since self-attention considers the channel information on tokens equally, we employ a dual-branch feature integration module to privilege the important channels in feature extraction. Furthermore, we design the wavelet fusion module which integrate shallow-layer structure and deep-layer detailed feature to recover realistic face images in frequency domain. Extensive experiments demonstrate the effectiveness on a variety of datasets. Guanxin Li, Jingang Shi, Yuan Zong, Fei Wang 0037, Tian Wang 0002, Yihong Gong |
IJCAI | 3 |
| 2023 | Learning Local to Global Feature Aggregation for Speech Emotion Recognition
Cheng Lu 0005, Hailun Lian, Wenming Zheng, Yuan Zong, Yan Zhao 0037, Sunan Li |
INTERSPEECH | 4 |
| 2023 | Multimodal Emotion Recognition in Noisy Environment Based on Progressive Label RevisionabstractThe multimodal emotion recognition has attracted more attention in recent decades. Though remarkable progress has been achieved with the rapid development of deep learning, existing methods are still hard to tackle noise problems that occurred commonly in emotion recognition's practical application. To improve the robustness of the multimodal emotion recognition algorithm, we propose an MLP-based label revision algorithm. The framework consists of three complementary feature extraction networks that were verified in MER2023. After that, an MLP-based attention network with specially designed loss functions was used to fuse features from different modalities. Finally, the scheme that used the output probability of each emotion to revise the sample's output category was employed to revise the test set's label obtained by classifier. The samples that are most likely to be affected by noise and misclassified have a chance to get correct classification. The best experimental result shows that the F1-score of our algorithm on the test dataset of the MER 2023 Noise subchallenge is 86.35 and combined metric is 0.6694, which ranks 2nd at the MER 2023 NOISE subchallenge. Sunan Li, Hailun Lian, Cheng Lu 0005, Yan Zhao 0037, Chuangao Tang, Yuan Zong, Wenming Zheng |
ACM Multimedia | 6 |
| 2023 | An Efficient and Consistent Solution to the PnP Problem
Zhengfeng Xie, Qida Yu, Yuan Zong |
PRCV (2) | 4 |
| 2023 | Variational Instance-Adaptive Graph for EEG Emotion RecognitionabstractThe individual differences and the dynamic uncertain relationships among different electroencephalogram (EEG) regions are essential factors that limit EEG emotion recognition. To address these issues, in this article, we propose a variational instance-adaptive graph method (V-IAG) that simultaneously captures the individual dependencies among different EEG electrodes and estimates the underlying uncertain information. Specifically, we employ two branches, i.e., instance-adaptive branch and variational branch, to construct the graph. Inspired by the attention mechanism, the instance-adaptive branch generates the graph based on the input so as to characterize the individual dependencies among EEG channels. The variational branch generates the probabilistic graph, which quantifies the uncertainties. We combine these two types of graphs to extract more discriminative features. To present more precise graph representation, we propose a new operation named the multi-level and multi-graph convolution operation, which aggregates the features of EEG channels from different frequencies with different graphs. Furthermore, we design the graph coarsening and employ the sparse constraint to obtain more robust features. We conduct extensive experiments on three widely-used EEG emotion recognition databases, i.e., SJTU emotion EEG dataset (SEED), multi-modal physiological emotion recognition dataset (MPED) and DREAMER. The results demonstrate that the proposed model achieves the-state-of-the-art performance. Tengfei Song, Suyuan Liu, Wenming Zheng, Yuan Zong, Zhen Cui 0001, Yang Li 0019 |
IEEE Trans. Affect. Comput. | 4 |
| 2023 | FENP: A Database of Neonatal Facial Expression for Pain AnalysisabstractIn this article, we introduce a new neonatal facial expression database for pain analysis. This database, called facial expression of neonatal pain (FENP), contains 11,000 neonatal facial expression images associated with 106 Chinese neonates from two children's hospitals, i.e., the Children's Hospital Affiliated to Nanjing Medical University and Second Affiliated Hospital Affiliated to Nanjing Medical University in China. The facial expression images cover four categories of facial expressions, i.e., severe pain expression, mild pain expression, crying expression and calmness expression, where each category contains 2750 neonatal facial expression images. Based on this database, we also investigate the pain facial expression recognition problem using several state-of-the-art facial expression features and expression recognition methods, such as Gabor+SVM, LBP+SVM, HOG+SVM, LBP+HOG+SVM, and several Convolutional Neural Network (CNN) methods (including AlexNet, VGGNet, GoogLeNet, ResNet and DenseNet). The experimental results indicate that the proposed neonatal pain facial expression database is very suitable for the study of both neonatal pain and facial expression recognition. Moreover, the FENP database is publicly available after signing a license agreement (the users can contact Jingjie Yan ([email protected]), Guanming Lu ([email protected])) or Xiaonan Li ([email protected]). Jingjie Yan, Guanming Lu, Wenming Zheng, Chengwei Huang, Zhen Cui 0001, Yuan Zong, Mengying Chen, Jindu Zhu, Haibo Li 0001 |
IEEE Trans. Affect. Comput. | 7 |
| 2023 | Speech Emotion Recognition via an Attentive Time-Frequency Neural NetworkabstractSpectrogram is commonly used as the input feature of deep neural networks to learn the high(er)-level time–frequency pattern of speech signal for speech emotion recognition (SER). Generally, different emotions correspond to specific energy activations both within frequency bands and time frames on spectrogram, which indicates the frequency and time domains are both essential to represent the emotion for SER. However, recent spectrogram-based works mainly focus on modeling the long-term dependency in time domain, which makes these methods suffer from the following issues: 1) neglecting to model the emotion-related correlations within frequency domain during the time–frequency joint learning and 2) ignoring to capture the specific frequency bands associated with emotions. To cope with the issues, we propose an attentive time–frequency neural network (ATFNN) for SER, including a time–frequency neural network (TFNN) and time–frequency attention. Specifically, aiming at the first issue, we design a TFNN with a frequency-domain encoder (F-Encoder) based on the Transformer encoder and a time-domain encoder (T-Encoder) based on the bidirectional long short-term memory (Bi-LSTM). The F-Encoder and T-Encoder model the correlations within frequency bands and time frames, respectively, and they are embedded into a time–frequency joint learning strategy to obtain the time–frequency patterns of speech emotions. Moreover, to handle the second issue, we adopt the time–frequency attention with a frequency-attention network (F-Attention) and a time-attention network (T-Attention) to focus on the emotion-related long-range dependencies between frequency bands and across time frames, which can enhance the emotional discrimination of speech features. Extensive experimental results on three public emotional databases, i.e., IEMOCAP, ABC, and CASIA, show that our proposed ATFNN outperforms the state-of-the-art methods. Cheng Lu 0005, Wenming Zheng, Hailun Lian, Yuan Zong, Chuangao Tang, Sunan Li, Yan Zhao 0037 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2023 | EEG-Based Parkinson's Disease Recognition via Attention-Based Sparse Graph Convolutional Neural NetworkabstractParkinson's disease (PD) is a complicated neurological ailment that affects both the physical and mental wellness of elderly individuals which makes it problematic to diagnose in its initial stages. Electroencephalogram (EEG) promises to be an efficient and cost-effective method for promptly detecting cognitive impairment in PD. Nevertheless, prevailing diagnostic practices utilizing EEG features have failed to examine the functional connectivity among EEG channels and the response of associated brain areas causing an unsatisfactory level of precision. Here, we construct an attention-based sparse graph convolutional neural network (ASGCNN) for diagnosing PD. Our ASGCNN model uses a graph structure to represent channel relationships, the attention mechanism for selecting channels, and the L1 norm to capture channel sparsity. We conduct extensive experiments on the publicly available PD auditory oddball dataset, which consists of 24 PD patients (under ON/OFF drug status) and 24 matched controls, to validate the effectiveness of our method. Our results show that the proposed method provides better results compared to the publicly available baselines. The achieved scores for Recall, Precision, F1-score, Accuracy and Kappa measures are 90.36%, 88.43%, 88.41%, 87.67%, and 75.24%, respectively. Our study reveals that the frontal and temporal lobes show significant differences between PD patients and healthy individuals. In addition, EEG features extracted by ASGCNN demonstrate significant asymmetry in the frontal lobe among PD patients. These findings can offer a basis for the establishment of a clinical system for intelligent diagnosis of PD by using auditory cognitive impairment features. Hongli Chang, Yuan Zong, Cheng Lu 0005, Xuenan Wang |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | A Novel Micro-Expression Recognition Approach Using Attention-Based Magnification-Adaptive NetworksabstractMicro-Expression recognition (MER) is a challenging task due to the short duration and low intensity of Micro-Expressions. A popular method to tackle this is magnifying MEs so as to enlarge the expression intensity to make recognition easier. However, the single fixed magnification strategy, widely used in existing works of MER, is not appropriate for different subjects, because each subject has specific expression intensity corresponding to different MEs. To cope with this issue, we propose a novel Attention-based Magnification-Adaptive Network (AMAN) to learn adaptive magnification levels for the ME representation. The network consists of two modules: magnification attention (MA module) to adaptively focus on appropriate magnification levels of different MEs, and frame attention (FA module) to focus on discriminative aggregated frames in a ME video. Extensive experiments on three widely used databases manifest that our method yields state-of-art results compared with other methods. Mengting Wei, Wenming Zheng, Yuan Zong, Xingxun Jiang, Cheng Lu 0005, Jiateng Liu |
ICASSP | 3 |
| 2022 | Seeking Salient Facial Regions for Cross-Database Micro-Expression RecognitionabstractCross-Database Micro-Expression Recognition (CD-MER) aims to develop the Micro-Expression Recognition (MER) methods with strong domain adaptability, i.e., the ability to recognize the Micro-Expressions (MEs) of different subjects captured by different imaging devices in different scenes. The development of CDMER is faced with two key problems: 1) the severe feature distribution gap between the source and target databases; 2) the feature representation bottleneck of ME such local and subtle facial expressions. To solve these problems, this paper proposes a novel Transfer Group Sparse Regression method, namely TGSR, which aims to 1) optimize the measurement and better alleviate the difference between the source and target databases, and 2) highlight the valid facial regions to enhance extracted features, by the operation of selecting the group features from the raw face feature, where each region is associated with a group of raw face feature, i.e., the salient facial region selection. Compared with previous transfer group sparse methods, our proposed TGSR has the ability to select the salient facial regions, which is effective in alleviating aforementioned problems for better performance and reducing the computational cost at the same time. We use two public ME databases, i.e., CASME II and SMIC, to evaluate our proposed TGSR method. Experimental results show that our proposed TGSR learns the discriminative and explicable regions, and outperforms most state-of-the-art subspace-learning-based domain-adaptive methods for CDMER. Xingxun Jiang, Yuan Zong, Wenming Zheng, Jiateng Liu, Mengting Wei |
ICPR | 2 |
| 2022 | A Novel Magnification-Robust Network with Sparse Self-Attention for Micro-expression RecognitionabstractExisting works for spontaneous Micro-Expression Recognition (MER) tend to encode Micro-Expression (ME) movements to get more discriminative features. However, MEs’ low intensity makes the capture for motion extremely difficult, and the widely adopted unified-magnification strategy is prone to noise and lacks flexibility. To this end, this paper provides a new insight to encode ME motion and tackle magnification noise. Specifically, we reconstruct a new sequence via magnification techniques to make subtle ME movements more distinguishable. Afterward, Sparse Self-Attention (SSA) rectifies self-attention with Locality Sensitive Hashing (LSH), cutting the space into several hush buckets of related features. Only keys in the same bucket are operated in the attention term for every query feature. The resulting sparsity in the attention matrix prevents the network from attending features stemming from less-informative magnification degrees which could be regarded as noise, while retains the sequence modelling capability of standard self-attention. Extensive experiments on three public MER databases demonstrate our superiority against the state-of-the-art methods. Mengting Wei, Wenming Zheng, Xingxun Jiang, Yuan Zong, Cheng Lu 0005, Jiateng Liu |
ICPR | 4 |
| 2022 | Sample Self-Revised Network for Cross-Dataset Facial Expression RecognitionabstractFacial images with low quality, subjective annotation, severe occlusion, and rare subject identity can lead to the existence of outlier samples in facial expression datasets. These outlier samples are usually far from the center of the dataset in the feature space, resulting in huge differences in feature distribution, which severely restricts the performance of cross-dataset facial expression recognition (FER). To eliminate the influence of outlier samples on cross-dataset FER, we propose an unsupervised domain adaptation (UDA) method called Sample Self-Revised Network (SSRN), which 1) dynamically detects the outlier level of each sample in the source domain to reduce the disturbance of outlier samples to the model training, as well as 2) adaptively revises outlier samples in the source domain to improve transferability of the learned features. Experimental results show that our SSRN outperforms both classic deep UDA methods and state-of-the-art cross-dataset FER results. Wenming Zheng, Yuan Zong, Cheng Lu 0005, Xingxun Jiang |
IJCNN | 3 |
| 2022 | Adaptive Hierarchical Graph Convolutional Network for EEG Emotion RecognitionabstractHuman emotion is closely related to multiple distributed brain regions, and functional connections exist between the regions. However, how to abstract the region-level information to improve electroencephalograph (EEG) emotion recognition performance has not been well considered. To address this problem, we proposed a novel Adaptive Hierarchical Graph Convolutional Network (AHGCN), which includes the basic channel-level graph of EEG channels and the region-level graph of brain regions. Different from previous methods, we propose an adaptive pooling operation to automatically partition brain regions rather than manually define them. To capture the intrinsic functional connections between the brain regions or EEG channels, we design a gated adaptive graph convolution operation. Besides, we develop a graph unpooling operation to integrate the region-level graph and channel-level graph to extract more discrimination features for classification. Experiments on two widely-used datasets show that our proposed method is superior to many state-of-the-art methods on EEG emotion recognition and could find some interesting combinations of EEG channels. Yunlong Xue, Wenming Zheng, Yuan Zong, Hongli Chang, Xingxun Jiang |
IJCNN | 3 |
| 2022 | Deep Transductive Transfer Regression Network for Cross-Corpus Speech Emotion Recognition
Yan Zhao 0037, Jincen Wang, Ru Ye, Yuan Zong, Wenming Zheng, Li Zhao 0003 |
INTERSPEECH | 4 |
| 2022 | Motion cues guided feature aggregation and enhancement for video object segmentation
Wenming Zheng, Yuan Zong |
Neurocomputing | 3 |
| 2022 | Learning two groups of discriminative features for micro-expression recognition
Jinsheng Wei, Guanming Lu, Jingjie Yan, Yuan Zong |
Neurocomputing | 4 |
| 2022 | Cross-database micro-expression recognition based on transfer double sparse learning
Jiateng Liu, Yuan Zong, Wenming Zheng |
Multim. Tools Appl. | 2 |
| 2022 | A Sparse-Based Transformer Network With Associated Spatiotemporal Feature for Micro-Expression RecognitionabstractDespite a lot of work in excavating the emotion descriptor from the hidden information, learning an effective spatiotemporal feature is a challenging issue for micro-expression recognition due to the fact that the micro-expression has a small difference in dynamic change and occurs in localized facial regions. Therefore, these properties of micro-expression suggest that the representation is sparse in the spatiotemporal domain. In this letter, a high-performance spatiotemporal feature learning based on sparse transformer is presented to solve the above issue. We extract the strong associated spatiotemporal feature by distinguishing the spatial attention map and attentively fusing the temporal feature. Thus, the feature map extracted from the critical relation will be fully utilized, while the superfluous relation will be masked. Our proposed method achieves remarkable results compared to state-of-the-art methods, proving that the sparse representation can be successfully integrated into the self-attention mechanism for micro-expression recognition. Yuan Zong, Hongli Chang, Yushun Xiao, Li Zhao 0003 |
IEEE Signal Process. Lett. | 2 |
| 2022 | From Regional to Global Brain: A Novel Hierarchical Spatial-Temporal Neural Network Model for EEG Emotion RecognitionabstractIn this paper, we propose a novel Electroencephalograph (EEG) emotion recognition method inspired by neuroscience with respect to the brain response to different emotions. The proposed method, denoted by R2G-STNN, consists of spatial and temporal neural network models with regional to global hierarchical feature learning process to learn discriminative spatial-temporal EEG features. To learn the spatial features, a bidirectional long short term memory (BiLSTM) network is adopted to capture the intrinsic spatial relationships of EEG electrodes within brain region and between brain regions, respectively. Considering that different brain regions play different roles in the EEG emotion recognition, a region-attention layer into the R2G-STNN model is also introduced to learn a set of weights to strengthen or weaken the contributions of brain regions. Based on the spatial feature sequences, BiLSTM is adopted to learn both regional and global spatial-temporal features and the features are fitted into a classifier layer for learning emotion-discriminative features, in which a domain discriminator working corporately with the classifier is used to decrease the domain shift between training and testing data. Finally, to evaluate the proposed method, we conduct both subject-dependent and subject-independent EEG emotion recognition experiments on SEED database, and the experimental results show that the proposed method achieves state-of-the-art performance. Yang Li 0019, Wenming Zheng, Lei Wang 0001, Yuan Zong, Zhen Cui 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2022 | Domain Invariant Feature Learning for Speaker-Independent Speech Emotion RecognitionabstractIn this paper, we propose a novel domain invariant feature learning (DIFL) method to deal with speaker-independent speech emotion recognition (SER). The basic idea of DIFL is to learn the speaker-invariant emotion feature by eliminating domain shifts between the training and testing data caused by different speakers from the perspective of multi-source unsupervised domain adaptation (UDA). Specifically, we embed a hierarchical alignment layer with the strong-weak distribution alignment strategy into the feature extraction block to firstly reduce the discrepancy in feature distributions of speech samples across different speakers as much as possible. Furthermore, multiple discriminators in the discriminator block are utilized to confuse the speaker information of emotion features both inside the training data and between the training and testing data. Through them, a multi-domain invariant representation of emotional speech can be gradually and adaptively achieved by updating network parameters. We conduct extensive experiments on three public datasets, i. e., Emo-DB, eNTERFACE, and CASIA, to evaluate the SER performance of the proposed method, respectively. The experimental results show that the proposed method is superior to the state-of-the-art methods. Cheng Lu 0005, Yuan Zong, Wenming Zheng, Yang Li 0019, Chuangao Tang, Björn W. Schuller |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Cross-Database Micro-Expression Recognition: A BenchmarkabstractCross-database micro-expression recognition (CDMER) is one of recently emerging and interesting problem in micro-expression analysis. CDMER is more challenging than the conventional micro-expression recognition (MER), because the training and testing samples in CDMER come from different micro-expression databases, resulting in inconsistency of the feature distributions between the training and testing sets. In this paper, we contribute to this topic from three aspects. First, we establish a CDMER experimental evaluation protocol aiming to allow the researchers to conveniently work on this topic and evaluate their proposed methods under the same standard. Second, we conduct benchmark experiments by using NINE state-of-the-art domain adaptation (DA) methods and SIX popular spatiotemporal descriptors for investigating CDMER problem from two different perspectives. Third, we propose a novel DA method called region selective transfer regression (RSTR) to deal with the CDMER task. The overall superior performance of RSTR over the state-of-the-art DA methods demonstrates that taking into consideration the facial local region information used in RSTR contributes to developing effective DA methods for dealing with CDMER problem. Tong Zhang 0015, Yuan Zong, Wenming Zheng, C. L. Philip Chen, Xiaopeng Hong, Chuangao Tang, Zhen Cui 0001, Guoying Zhao 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | Cross-Corpus Speech Emotion Recognition Using Joint Distribution Adaptive RegressionabstractIn this paper, we focus on the research of cross-corpus speech emotion recognition (SER), in which the training and testing speech signals in cross-corpus SER belong to dierent speech corpus. Due to this fact, mismatched feature distributions may exist between the training and testing speech feature sets degrading the performance of most originally well-performing SER methods. To deal with cross-corpus SER, we propose a novel domain adaptation (DA) method called joint distribution adaptive regression (JDAR). The basic idea of JDAR is to learn a regression matrix by jointly considering the marginal and conditional probability distribution between the training and testing speech signals and hence their feature distribution dierence can be alleviated in the subspace spanned by the learned regression matrix. To evaluate the proposed JDAR, we conduct extensive cross-corpus SER experiments on EmoDB, eNTERFACE, and CASIA speech databases. Experimental results show that the proposed JDAR achieves satisfactory performance and outperforms most of state-of-the-art subspace learning based DA methods. Lin Jiang 0007, Yuan Zong, Wenming Zheng, Li Zhao 0003 |
ICASSP | 3 |
| 2021 | Attention-based Spatio-Temporal Graphic LSTM for EEG Emotion RecognitionabstractAutomatic emotion recognition based on electroencephalogram (EEG) is a challenging task in Brain Machine Interfaces (BMI). Since it is still not very clear about the intrinsic connection relationship among the various EEG channels, it is still a challenging task of how to better represent the topology of EEG channels for emotion recognition. On the other hand, the intensity of the emotion may vary along the different time instants, which would affect the recognition accuracy of emotion. To tackle the above issues, in this paper, we propose a novel multichannel EEG emotion recognition method called attention-based spatiotemporal graphic long short-term memory (ASTG-LSTM), in which a Dynamic Structured Learning (DSL) branch that focuses on the most emotion-relevant connectivity of the brain is incorporated to represent inter-channel connections of the EEG signals. In addition, a specific spatio-temporal attention is embedded into the DSL branch to improve the invariance ability against the emotional intensity fluctuation. Extensive experiments on: DEAP and DREAMER are conducted and the experimental results indicate that the proposed ASTG-LSTM model improves the EEG emotion recognition performance compared with many state-of-the-art approaches. Wenming Zheng, Yuan Zong, Hongli Chang, Cheng Lu 0005 |
IJCNN | 3 |
| 2021 | A Bi-Hemisphere Domain Adversarial Neural Network Model for EEG Emotion RecognitionabstractIn this paper, we propose a novel neural network model, called bi-hemisphere domain adversarial neural network (BiDANN) model, for electroencephalograph (EEG) emotion recognition. The BiDANN model is inspired by the neuroscience findings that the left and right hemispheres of human's brain are asymmetric to the emotional response. It contains a global and two local domain discriminators that work adversarially with a classifier to learn discriminative emotional features for each hemisphere. At the same time, it tries to reduce the possible domain differences in each hemisphere between the source and target domains so as to improve the generality of the recognition model. In addition, we also propose an improved version of BiDANN, denoted by BiDANN-S, for subject-independent EEG emotion recognition problem by lowering the influences of the personal information of subjects to the EEG emotion recognition. Extensive experiments on the SEED database are conducted to evaluate the performance of both BiDANN and BiDANN-S. The experimental results have shown that the proposed BiDANN and BiDANN models achieve state-of-the-art performance in the EEG emotion recognition. Yang Li 0019, Wenming Zheng, Yuan Zong, Zhen Cui 0001, Tong Zhang 0021 |
IEEE Trans. Affect. Comput. | 3 |
| 2021 | Multi-scale discrepancy adversarial network for crosscorpus speech emotion recognitionabstractOne of the most critical issues in human-computer interaction applications is recognizing human emotions based on speech. In recent years, the challenging problem of cross-corpus speech emotion recognition (SER) has generated extensive research. Nevertheless, the domain discrepancy between training data and testing data remains a major challenge to achieving improved system performance. This paper introduces a novel multi-scale discrepancy adversarial (MSDA) network for conducting multiple timescales domain adaptation for cross-corpus SER, i.e.,integrating domain discriminators of hierarchical levels into the emotion recognition framework to mitigate the gap between the source and target domains. Specifically, we extract two kinds of speech features, i.e., handcraft features and deep features, from three timescales of global, local, and hybrid levels. In each timescale, the domain discriminator and the emotion classifier compete against each other to learn features that minimize the discrepancy between the two domains by fooling the discriminator. Extensive experiments on cross-corpus and cross-language SER were conducted on a combination dataset that combines one Chinese dataset and two English datasets commonly used in SER. The MSDA is affected by the strong discriminate power provided by the adversarial process, where three discriminators are working in tandem with an emotion classifier. Accordingly, the MSDA achieves the best performance over all other baseline methods. The proposed architecture was tested on a combination of one Chinese and two English datasets. The experimental results demonstrate the superiority of our powerful discriminative model for solving cross-corpus SER. Wanlu Zheng, Wenming Zheng, Yuan Zong |
Virtual Real. Intell. Hardw. | 3 |
| 2020 | Instance-Adaptive Graph for EEG Emotion RecognitionabstractTo tackle the individual differences and characterize the dynamic relationships among different EEG regions for EEG emotion recognition, in this paper, we propose a novel instance-adaptive graph method (IAG), which employs a more flexible way to construct graphic connections so as to present different graphic representations determined by different input instances. To fit the different EEG pattern, we employ an additional branch to characterize the intrinsic dynamic relationships between different EEG channels. To give a more precise graphic representation, we design the multi-level and multi-graph convolutional operation and the graph coarsening. Furthermore, we present a type of sparse graphic representation to extract more discriminative features. Experiments on two widely-used EEG emotion recognition datasets are conducted to evaluate the proposed model and the experimental results show that our method achieves the state-of-the-art performance. Tengfei Song, Suyuan Liu, Wenming Zheng, Yuan Zong, Zhen Cui 0001 |
AAAI | 4 |
| 2020 | DFEW: A Large-Scale Database for Recognizing Dynamic Facial Expressions in the WildabstractRecently, facial expression recognition (FER) in the wild has gained a lot of researchers' attention because it is a valuable topic to enable the FER techniques to move from the laboratory to the real applications. In this paper, we focus on this challenging but interesting topic and make contributions from three aspects. First, we present a new large-scale 'in-the-wild' dynamic facial expression database, DFEW (Dynamic Facial Expression in the Wild), consisting of over 16,000 video clips from thousands of movies. These video clips contain various challenging interferences in practical scenarios such as extreme illumination, occlusions, and capricious pose changes. Second, we propose a novel method called Expression-Clustered Spatiotemporal Feature Learning (EC-STFL) framework to deal with dynamic FER in the wild. Third, we conduct extensive benchmark experiments on DFEW using a lot of spatiotemporal deep feature learning methods as well as our proposed EC-STFL. Experimental results show that DFEW is a well-designed and challenging database, and the proposed EC-STFL can promisingly improve the performance of existing spatiotemporal deep neural networks in coping with the problem of dynamic FER in the wild. Our DFEW database is publicly available and can be freely downloaded from https://dfew-dataset.github.io/. Xingxun Jiang, Yuan Zong, Wenming Zheng, Chuangao Tang, Wanchuang Xia, Cheng Lu 0005, Jiateng Liu |
ACM Multimedia | 2 |
| 2020 | Toward Bridging Microexpressions From Different DomainsabstractRecently, microexpression recognition has attracted a lot of researchers' attention due to its challenges and valuable applications. However, it is noticed that currently most of the existing proposed methods are often evaluated and tested on the single database and, hence, this brings us a question whether these methods are still effective if the training and testing samples belong to different domains, for example, different microexpression databases. In this case, a large feature distribution difference may exist between training (source) and testing (target) samples and, hence, microexpression recognition tasks would become more difficult. To solve this challenging problem, that is, cross-domain microexpression recognition, in this paper, we propose an effective method consisting of an auxiliary set selection model (ASSM) and a transductive transfer regression model (TTRM). In our method, an ASSM is designed to automatically select an optimal set of samples from the target domain to serve as the auxiliary set, which is used for subsequent TTRM training. As for TTRM, it aims at bridging the feature distribution gap between the source and target domains by learning a joint regression model with the source domain samples and the auxiliary set selected from the target domain. We evaluate the proposed TTRM plus ASSM by extensive cross-domain microexpression recognition experiments on SMIC and CASME II databases. Compared with the recent state-of-the-art domain adaptation methods, our proposed method has a more satisfactory performance in dealing with the cross-domain microexpression recognition tasks. Yuan Zong, Wenming Zheng, Zhen Cui 0001, Guoying Zhao 0001, Bin Hu 0001 |
IEEE Trans. Cybern. | 1 |
| 2020 | Deep Manifold-to-Manifold Transforming Network for Skeleton-Based Action RecognitionabstractIn this paper, we will investigate skeleton-based action recognition by employing high-order statistics feature and first-order statistics feature, where the high-order statistics feature is characterized by symmetric positive definite (SPD) matrices. Noting that SPD matrices are theoretically embedded on Riemannian manifolds, we propose an end-to-end deep manifold-to-manifold transforming network (DMT-Net), which can make SPD matrices flow from one Riemannian manifold to another one for facilitating the action recognition task. To learn discriminative SPD features from both spatial and temporal dependencies, we propose a neural network model with three novel layers on manifolds: i.e., (1) the local SPD convolutional layer, (2) the non-linear SPD activation layer, and (3) the Riemannian-preserved recursive layer. The SPD property is preserved through all layers without the singular value decomposition (SVD) operation, which has to be conducted in the existing methods with expensive computation cost. Furthermore, a diagonalizing SPD layer is designed to efficiently calculate the final metric for the classification task. Finally, DMT-Net is further fused with a first order layer to capture temporal evolution information. To evaluate our proposed method, we conduct extensive experiments on the task of action recognition, where the input signals are represented as SPD matrices. The experimental results demonstrate that the proposed method is competitive over state-of-the-art methods. Tong Zhang 0021, Wenming Zheng, Zhen Cui 0001, Yuan Zong, Chaolong Li, Jian Yang 0003 |
IEEE Trans. Multim. | 4 |
| 2019 | Bi-modality Fusion for Emotion Recognition in the WildabstractThe emotion recognition in the wild has been a hot research topic in the field of affective computing. Though some progresses have been achieved, the emotion recognition in the wild is still an unsolved problem due to the challenge of head movement, face deformation, illumination variation etc. To deal with these unconstrained challenges, we propose a bi-modality fusion method for video based emotion recognition in the wild. The proposed framework takes advantages of the visual information from facial expression sequences and the speech information from audio. The state-of-the-art CNN based object recognition models are employed to facilitate the facial expression recognition performance. A bi-direction long short term Memory (Bi-LSTM) is employed to capture dynamic information of the learned features. Additionally, to take full advantages of the facial expression information, the VGG16 network is trained on AffectNet dataset to learn a specialized facial expression recognition model. On the other hand, the audio based features, like low level descriptor (LLD) and deep features obtained by spectrogram image, are also developed to improve the emotion recognition performance. The best experimental result shows that the overall accuracy of our algorithm on the Test dataset of the EmotiW challenge is 62.78, which outperforms the best result of EmotiW2018 and ranks 2nd at the EmotiW2019 challenge. Sunan Li, Wenming Zheng, Yuan Zong, Cheng Lu 0005, Chuangao Tang, Xingxun Jiang, Jiateng Liu, Wanchuang Xia |
ICMI | 3 |
| 2019 | Sparse Graphic Attention LSTM for EEG Emotion Recognition
Suyuan Liu, Wenming Zheng, Tengfei Song, Yuan Zong |
ICONIP (4) | 4 |
| 2019 | Cross-Database Micro-Expression Recognition: A BenchmarkabstractCross-database micro-expression recognition (CDMER) is one of recently emerging and interesting problems in micro-expression analysis. CDMER is more challenging than the conventional micro-expression recognition (MER), because the training and testing samples in CDMER come from different micro-expression databases, resulting in inconsistency of the feature distributions between the training and testing sets. In this paper, we contribute to this topic from two aspects. First, we establish a CDMER experimental evaluation protocol and provide a standard platform for evaluating their proposed methods. Second, we conduct extensive benchmark experiments by using NINE state-of-the-art domain adaptation (DA) methods and SIX popular spatiotemporal descriptors for investigating the CDMER problem from two different perspectives and deeply analyze and discuss the experimental results. In addition, all the data and codes involving CDMER in this paper are released on our project website: http://aip.seu.edu.cn/cdmer. Yuan Zong, Wenming Zheng, Xiaopeng Hong, Chuangao Tang, Zhen Cui 0001, Guoying Zhao 0001 |
ICMR | 1 |
| 2019 | EEG Emotion Recognition Based on Graph Regularized Sparse Linear Regression
Yang Li 0019, Wenming Zheng, Zhen Cui 0001, Yuan Zong, Sheng Ge |
Neural Process. Lett. | 4 |
| 2019 | Spatial-Temporal Recurrent Neural Network for Emotion RecognitionabstractIn this paper, we propose a novel deep learning framework, called spatial-temporal recurrent neural network (STRNN), to integrate the feature learning from both spatial and temporal information of signal sources into a unified spatial-temporal dependency model. In STRNN, to capture those spatially co-occurrent variations of human emotions, a multidirectional recurrent neural network (RNN) layer is employed to capture long-range contextual cues by traversing the spatial regions of each temporal slice along different directions. Then a bi-directional temporal RNN layer is further used to learn the discriminative features characterizing the temporal dependencies of the sequences, where sequences are produced from the spatial RNN layer. To further select those salient regions with more discriminative ability for emotion recognition, we impose sparse projection onto those hidden states of spatial and temporal domains to improve the model discriminant ability. Consequently, the proposed two-layer RNN model provides an effective way to make use of both spatial and temporal dependencies of the input signals for emotion recognition. Experimental results on the public emotion datasets of electroencephalogram and facial expression demonstrate the proposed STRNN method is more competitive over those state-of-the-art methods. Tong Zhang 0021, Wenming Zheng, Zhen Cui 0001, Yuan Zong, Yang Li 0019 |
IEEE Trans. Cybern. | 4 |
| 2018 | Super Wide Regression Network for Unsupervised Cross-Database Facial Expression RecognitionabstractUnsupervised cross-database facial expression recognition (FER) is a challenging problem, in which the training and testing samples belong to different facial expression databases. For this reason, the training (source) and testing (target) facial expression samples would have different feature distributions and hence the performance of lots of existing FER methods may decrease. To solve this problem, in this paper we propose a novel super wide regression network (SWiRN) model, which serves as the regression parameter to bridge the original feature space and the label space and herein in each layer the maximum mean discrepancy (MMD) criterion is used to enforce the source and target facial expression samples to share the same or similar feature distributions. Consequently, the learned SWiRN is able to predict the expression categories of the target samples although we have no access to any label information of target samples. We conduct extensive cross-database FER experiments on CK+, eNTERFACE, and Oulu-CASIA VIS facial expression databases to evaluate the proposed SWiRN. Experimental results show that our SWiRN model achieves more promising performance than recent proposed cross-database emotion recognition methods. Baofeng Zhang, Yuan Zong, Li Liu 0002, Jie Chen 0001, Guoying Zhao 0001, Junchao Zhu |
ICASSP | 3 |
| 2018 | Unsupervised Cross-Corpus Speech Emotion Recognition Using Domain-Adaptive Subspace LearningabstractIn this paper, we investigate an interesting problem, i.e., unsupervised cross-corpus speech emotion recognition (SER), in which the training and testing speech signals come from two different speech emotion corpora. Meanwhile, the training speech signals are labeled, while the label information of the testing speech signals is entirely unknown. Due to this setting, the training (source) and testing (target) speech signals may have different feature distributions and therefore lots of existing SER methods would not work. To deal with this problem, we propose a domain-adaptive subspace learning (DoSL) method for learning a projection matrix with which we can transform the source and target speech signals from the original feature space to the label space. The transformed source and target speech signals in the label space would have similar feature distributions. Consequently, the classifier learned on the labeled source speech signals can effectively predict the emotional states of the unlabeled target speech signals. To evaluate the performance of the proposed DoSL method, we carry out extensive cross-corpus SER experiments on three speech emotion corpora including EmoDB, eNTERFACE, and AFEW 4.0. Compared with recent state-of-the-art cross-corpus SER methods, the proposed DoSL can achieve more satisfactory overall results. Yuan Zong, Baofeng Zhang, Li Liu 0002, Jie Chen 0001, Guoying Zhao 0001, Junchao Zhu |
ICASSP | 2 |
| 2018 | Multiple Spatio-temporal Feature Learning for Video-based Emotion Recognition in the WildabstractThe difficulty of emotion recognition in the wild (EmotiW) is how to train a robust model to deal with diverse scenarios and anomalies. The Audio-video Sub-challenge in EmotiW contains audio-video short clips with several emotional labels and the task is to distinguish which label the video belongs to. For the better emotion recognition in videos, we propose a multiple spatio-temporal feature fusion (MSFF) framework, which can more accurately depict emotional information in spatial and temporal dimensions by two mutually complementary sources, including the facial image and audio. The framework is consisted of two parts: the facial image model and the audio model. With respect to the facial image model, three different architectures of spatial-temporal neural networks are employed to extract discriminative features about different emotions in facial expression images. Firstly, the high-level spatial features are obtained by the pre-trained convolutional neural networks (CNN), including VGG-Face and ResNet-50 which are all fed with the images generated by each video. Then, the features of all frames are sequentially input to the Bi-directional Long Short-Term Memory (BLSTM) so as to capture dynamic variations of facial appearance textures in a video. In addition to the structure of CNN-RNN, another spatio-temporal network, namely deep 3-Dimensional Convolutional Neural Networks (3D CNN) by extending the 2D convolution kernel to 3D, is also applied to attain evolving emotional information encoded in multiple adjacent frames. For the audio model, the spectrogram images of speech generated by preprocessing audio, are also modeled in a VGG-BLSTM framework to characterize the affective fluctuation more efficiently. Finally, a fusion strategy with the score matrices of different spatio-temporal networks gained from the above framework is proposed to boost the performance of emotion recognition complementally. Extensive experiments show that the overall accuracy of our proposed MSFF is 60.64%, which achieves a large improvement compared with the baseline and outperform the result of champion team in 2017. Cheng Lu 0005, Wenming Zheng, Chaolong Li, Chuangao Tang, Suyuan Liu, Simeng Yan, Yuan Zong |
ICMI | 7 |
| 2018 | A Novel Neural Network Model based on Cerebral Hemispheric Asymmetry for EEG Emotion RecognitionabstractIn this paper, we propose a novel neural network model, called bi-hemispheres domain adversarial neural network (BiDANN), for EEG emotion recognition. BiDANN is motivated by the neuroscience findings, i.e., the emotional brain's asymmetries between left and right hemispheres. The basic idea of BiDANN is to map the EEG feature data of both left and right hemispheres into discriminative feature spaces separately, in which the data representations can be classified easily. For further precisely predicting the class labels of testing data, we narrow the distribution shift between training and testing data by using a global and two local domain discriminators, which work adversarially to the classifier to encourage domain-invariant data representations to emerge. After that, the learned classifier from labeled training data can be applied to unlabeled testing data naturally. We conduct two experiments to verify the performance of our BiDANN model on SEED database. The experimental results show that the proposed model achieves the state-of-the-art performance. Yang Li 0019, Wenming Zheng, Zhen Cui 0001, Tong Zhang 0021, Yuan Zong |
IJCAI | 5 |
| 2018 | Multi-cue fusion for emotion recognition in the wild
Jingwei Yan, Wenming Zheng, Zhen Cui 0001, Chuangao Tang, Tong Zhang 0021, Yuan Zong |
Neurocomputing | 6 |
| 2018 | Unsupervised facial expression recognition using domain adaptation based dictionary learning approach
Wenming Zheng, Zhen Cui 0001, Yuan Zong, Tong Zhang 0021, Chuangao Tang |
Neurocomputing | 4 |
| 2018 | Cross-Domain Color Facial Expression Recognition Using Transductive Transfer Subspace LearningabstractFacial expression recognition across domains, e.g., training and testing facial images come from different facial poses, is very challenging due to the different marginal distributions between training and testing facial feature vectors. To deal with such challenging cross-domain facial expression recognition problem, a novel transductive transfer subspace learning method is proposed in this paper. In this method, a labelled facial image set from source domain is combined with an unlabelled auxiliary facial image set from target domain to jointly learn a discriminative subspace and make the class labels prediction of the unlabelled facial images, where a transductive transfer regularized least-squares regression (TTRLSR) model is proposed to this end. Then, based on the auxiliary facial image set, we train a SVM classifier for classifying the expressions of other facial images in the target domain. Moreover, we also investigate the use of color facial features to evaluate the recognition performance of the proposed facial expression recognition method, where color scale invariant feature transform (CSIFT) features associated with 49 landmark facial points are extracted to describe each color facial image. Finally, extensive experiments on BU-3DFE and Multi-PIE multiview color facial expression databases are conducted to evaluate the cross-database & cross-view facial expression recognition performance of the proposed method. Comparisons with state-of-the-art domain adaption methods are also included in the experiments. The experimental results demonstrate that the proposed method achieves much better recognition performance compared with the state-of-the-art methods. Wenming Zheng, Yuan Zong, Minghai Xin |
IEEE Trans. Affect. Comput. | 2 |
| 2018 | Hallucinating Face Image by Regularization Models in High-Resolution Feature SpaceabstractIn this paper, we propose two novel regularization models in patch-wise and pixel-wise respectively, which are efficient to reconstruct high-resolution (HR) face image from low-resolution (LR) input. Unlike the conventional patch-based models which depend on the assumption of local geometry consistency in LR and HR spaces, the proposed method directly regularizes the relationship between the target patch and corresponding training set in the HR space. It avoids to deal with the tough problem of preserving local geometry in various resolutions. Taking advantage of kernel function in efficiently describing intrinsic features, we further conduct the patch-based reconstruction model in the high-dimensional kernel space for capturing nonlinear characteristics. Meanwhile, a pixel-based model is proposed to regularize the relationship of pixels in the local neighborhood, which can be employed to enhance the fuzzy details in the target HR face image. It privileges the reconstruction of pixels along the dominant orientation of structure, which is useful for preserving high-frequency information on complex edges. Finally, we combine the two reconstruction models into a unified framework. The output HR face image can be finally optimized by performing an iterative procedure. Experimental results demonstrate that the proposed face hallucination method produces superior performance than the state-of-the-art methods. Jingang Shi, Xin Liu 0012, Yuan Zong, Chun Qi, Guoying Zhao 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | Domain Regeneration for Cross-Database Micro-Expression RecognitionabstractRecently, micro-expression recognition has attracted lots of researchers' attention due to its potential value in many practical applications, e.g., lie detection. In this paper, we investigate an interesting and challenging problem in micro-expression recognition, i.e., cross-database micro-expression recognition, in which the training and testing samples come from different micro-expression databases. Under this problem setting, the consistent feature distribution between the training and testing samples originally existing in conventional micro-expression recognition would be seriously broken and hence the performance of most current well-performing micro-expression recognition methods may sharply drop. In order to overcome it, we propose a simple yet effective framework called Domain Regeneration (DR) in this paper. DR framework aims at learning a domain regenerator to regenerate the micro-expression samples from source and target databases respectively such that they can abide by the same or similar feature distributions. Thus, we are able to use the classifier learned based on the labeled source micro-expression samples to predict the label information of the unlabeled target micro-expression samples. To evaluate the proposed DR framework, we conduct extensive cross-database micro-expression recognition experiments designed based on SMIC and CASME II databases. Experimental results show that compared with recent state-of-the-art cross-database emotion recognition methods, the proposed DR framework has more promising performance. Yuan Zong, Wenming Zheng, Xiaohua Huang 0003, Jingang Shi, Zhen Cui 0001, Guoying Zhao 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Learning From Hierarchical Spatiotemporal Descriptors for Micro-Expression RecognitionabstractMicro-expression recognition aims to infer genuine emotions that people try to conceal from facial video clips. It is a very challenging task because micro-expressions have a very low intensity and short duration, which makes micro-expressions difficult to observe. Recently, researchers have designed various spatiotemporal descriptors to describe micro-expressions. It is notable that for better capturing the low-intensity facial muscle movement, a fixed spatial division grid, 8× 8 for example, is commonly used to partition the facial images into a few facial blocks before extracting descriptors. However, it is hard to choose an ideal division grid for different micro-expression samples because the division grids affect the discriminative ability of spatiotemporal descriptors to distinguish micro-expressions. To address this problem, in this paper, we design a hierarchical spatial division scheme for spatiotemporal descriptor extraction. By using the proposed scheme, it would not be a problem to determine which division grid is most suitable regarding different micro-expression samples. Furthermore, we propose a kernelized group sparse learning (KGSL) model to process hierarchical scheme based spatiotemporal descriptors such that they are more effective for micro-expression recognition tasks. To evaluate the performance of the proposed micro-expression recognition method consisting of the hierarchical scheme based spatiotemporal descriptors and KGSL, extensive experiments are conducted on two public micro-expression databases: CASME II and SMIC. Compared with many recent state-of-the-art approaches, our method achieves more promising recognition results. Yuan Zong, Xiaohua Huang 0003, Wenming Zheng, Zhen Cui 0001, Guoying Zhao 0001 |
IEEE Trans. Multim. | 1 |
| 2017 | Learning a Target Sample Re-Generator for Cross-Database Micro-Expression RecognitionabstractIn this paper, we investigate the cross-database micro-expression recognition problem, where the training and testing samples are from two different micro-expression databases. Under this setting, the training and testing samples would have different feature distributions and hence the performance of most existing micro-expression recognition methods may decrease greatly. To solve this problem, we propose a simple yet effective method called Target Sample Re-Generator (TSRG) in this paper. By using TSRG, we are able to re-generate the samples from target micro-expression database and the re-generated target samples would share same or similar feature distributions with the original source samples. For this reason, we can then use the classifier learned based on the labeled source samples to accurately predict the micro-expression categories of the unlabeled target samples. To evaluate the performance of the proposed TSRG method, extensive cross-database micro-expression recognition experiments designed based on SMIC and CASME II databases are conducted. Compared with recent state-of-the-art cross-database emotion recognition methods, the proposed TSRG achieves more promising results. Yuan Zong, Xiaohua Huang 0003, Wenming Zheng, Zhen Cui 0001, Guoying Zhao 0001 |
ACM Multimedia | 1 |
| 2016 | Multi-clue fusion for emotion recognition in the wildabstractIn the past three years, Emotion Recognition in the Wild (EmotiW) Grand Challenge has drawn more and more attention due to its huge potential applications. In the fourth challenge, aimed at the task of video based emotion recognition, we propose a multi-clue emotion fusion (MCEF) framework by modeling human emotion from three mutually complementary sources, facial appearance texture, facial action, and audio. To extract high-level emotion features from sequential face images, we employ a CNN-RNN architecture, where face image from each frame is first fed into the fine-tuned VGG-Face network to extract face feature, and then the features of all frames are sequentially traversed in a bidirectional RNN so as to capture dynamic changes of facial textures. To attain more accurate facial actions, a facial landmark trajectory model is proposed to explicitly learn emotion variations of facial components. Further, audio signals are also modeled in a CNN framework by extracting low-level energy features from segmented audio clips and then stacking them as an image-like map. Finally, we fuse the results generated from three clues to boost the performance of emotion recognition. Our proposed MCEF achieves an overall accuracy of 56.66% with a large improvement of 16.19% with respect to the baseline. Jingwei Yan, Wenming Zheng, Zhen Cui 0001, Chuangao Tang, Tong Zhang 0021, Yuan Zong, Ning Sun 0001 |
ICMI | 6 |
| 2016 | Cross-Database Facial Expression Recognition via Unsupervised Domain Adaptive Dictionary Learning
Wenming Zheng, Zhen Cui 0001, Yuan Zong |
ICONIP (2) | 4 |
| 2016 | WFID: Passive Device-free Human Identification Using WiFi SignalabstractWe present WFID, a passive device-free indoor human identification system with one pair of WiFi signal transmitter and receiver. WFID design is motivated by the observation that PHY layer Channel State Information (CSI) is capable of capturing the frequency diversity of wideband channel, such that the human body curve may be uniquely identified by learning the feature pattern of CSI. Different from many CSI-based techniques focusing on phase shift, we propose a novel feature of subcarrier-amplitude frequency (SAF). Based on this feature, WFID realizes human identification through a linear-kernel SVM. We have implemented a prototype of WFID with a commercial AP and a computer equipped with one Intel 5300 NIC. WFID is evaluated in two typical indoor scenarios. The results confirm that WFID achieves high classification accuracy which is permanent over several days under two typical indoor scenarios, with low computation cost. This reveals the potential for WFID to realize real-time indoor human identification. Feng Hong 0001, Yuan Zong, Zhongwen Guo |
MobiQuitous | 4 |
| 2016 | Cross-Corpus Speech Emotion Recognition Based on Domain-Adaptive Least-Squares RegressionabstractIn this letter, a novel cross-corpus speech emotion recognition (SER) method using domain-adaptive least-squares regression (DaLSR) model is proposed. In this method, an additional unlabeled data set from target speech corpus is used to serve as an auxiliary data set and combined with the labeled training data set from source speech corpus for jointly training the DaLSR model. In contrast to the traditional least-squares regression (LSR) method, the major novelty of DaLSR is that it is able to handle the mismatch problem between source and target speech corpora. Hence, the proposed DaLSR method is very suitable for coping with cross-corpus SER problem. For evaluating the performance of the proposed method in dealing with the cross-corpus SER problem, we conduct extensive experiments on three emotional speech corpora and compare the results with several state-of-the-art transfer learning methods that are widely used for cross-corpus SER problem. The experimental results show that the proposed method achieves better recognition accuracies than the state-of-the-art methods. Yuan Zong, Wenming Zheng, Tong Zhang 0021, Xiaohua Huang 0003 |
IEEE Signal Process. Lett. | 1 |
| 2016 | A Deep Neural Network-Driven Feature Learning Method for Multi-view Facial Expression RecognitionabstractIn this paper, a novel deep neural network (DNN)-driven feature learning method is proposed and applied to multi-view facial expression recognition (FER). In this method, scale invariant feature transform (SIFT) features corresponding to a set of landmark points are first extracted from each facial image. Then, a feature matrix consisting of the extracted SIFT feature vectors is used as input data and sent to a well-designed DNN model for learning optimal discriminative features for expression classification. The proposed DNN model employs several layers to characterize the corresponding relationship between the SIFT feature vectors and their corresponding high-level semantic information. By training the DNN model, we are able to learn a set of optimal features that are well suitable for classifying the facial expressions across different facial views. To evaluate the effectiveness of the proposed method, two nonfrontal facial expression databases, namely BU-3DFE and Multi-PIE, are respectively used to testify our method and the experimental results show that our algorithm outperforms the state-of-the-art methods. Tong Zhang 0021, Wenming Zheng, Zhen Cui 0001, Yuan Zong, Jingwei Yan |
IEEE Trans. Multim. | 4 |
| 2015 | Transductive Transfer LDA with Riesz-based Volume LBP for Emotion Recognition in The WildabstractIn this paper, we propose the method using Transductive Transfer Linear Discriminant Analysis (TTLDA) and Riesz-based Volume Local Binary Patterns (RVLBP) for image based static facial expression recognition challenge of the Emotion Recognition in the Wild Challenge (EmotiW 2015). The task of this challenge is to assign facial expression labels to frames of some movies containing a face under the real word environment. In our method, we firstly employ a multi-scale image partition scheme to divide each face image into some image blocks and use RVLBP features extracted from each block to describe each facial image. Then, we adopt the TTLDA approach based on RVLBP to cope with the expression recognition task. The experiments on the testing data of SFEW 2.0 database, which is used for image based static facial expression challenge, demonstrate that our method achieves the accuracy of 50%. This result has a 10.87% improvement over the baseline provided by this challenge organizer. Yuan Zong, Wenming Zheng, Xiaohua Huang 0003, Jingwei Yan, Tong Zhang 0021 |
ICMI | 1 |
| 2000 | Emotion Expression Functions in Multimodal Presentation
Yuan Zong, Hiroshi Dohi, Helmut Prendinger, Mitsuru Ishizuka |
ICMI | 1 |