VLDB 2026 Research / reviewers in the wild / expert
Wenming Zheng
dblp:64/2253
· DBLP profile ↗
203ranked-venue papers
28as first author
112since 2021 · last 2026
0000-0002-7764-5179ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 116 · 21 first-author · 62 since 2021Graphics, computer vision, multimedia, augmented reality and games · 86 · 8 first-author · 47 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 15 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-authorDatabases, data management, data science and information retrieval · 4 · 2 since 2021Systems, architecture and hardware · 3 · 2 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SF-STACK: Streamlining RDMA for Heterogeneous Telecom Storage
Wenming Zheng, Xiaoping Fan, Fangfang Yan, Luren Liu, Xingling Han, Anran Xu 0003 |
INFOCOM | 2 |
| 2026 | Budget-Constrained Federated Bandits for Mobile Applications
Anran Xu 0003, Zhenzhe Zheng 0001, Wenming Zheng, Fan Wu 0006 |
INFOCOM | 3 |
| 2026 | FPS: Frequency prompt synchronization for micro-expression recognition
Jiateng Liu, Hengcan Shi, Yaonan Wang 0001, Wenming Zheng |
Pattern Recognit. | 4 |
| 2026 | Speaker-independent speech emotion recognition using group sparse-based adversarial local fisher discriminant analysis
Cheng Lu 0005, Kaifei Zhang, Hailun Lian, Sunan Li, Tianhua Qi, Yuan Zong, Wenming Zheng |
Pattern Recognit. | 7 |
| 2026 | Hypergraph regularization-based anchor learning for multi-view clustering
Yunpeng Zeng, Peng Song 0002, Beihua Yang, Changjia Wang, Guanghao Du, Yanwei Yu, Wenming Zheng |
Pattern Recognit. | 7 |
| 2026 | Towards Identity-Independent Facial Action Unit Detection: Integrating Decoupled 3D Geometry With Textural FeaturesabstractOne challenge in facial action unit detection lies in the variability of facial attributes across individuals. This variability leads to diminished performance of existing AU detection methods on unseen facial identities. The primary reason for this is that the AU representations, as acquired by these methods, are closely entangled with facial identities in the training data. Consequently, the detection model tends to learn identity-specific features, which are not exclusively related to AUs. To alleviate the entanglement problem, we propose a facial geometric decoupling method inspired by 3D Morphable Model. Our approach consists of a Texture Learning (TL) module for extracting AU-related textural features, a Geometry Learning (GL) module for identity-independent geometric features via decoupling losses, and a Feature Fusion (FF) module. FF module is used to capture the interrelationship between 2D textural and 3D geometric features, then it further learns the relationship between AUs through a graph neural network. We evaluated our model on the DISFA and BP4D datasets, achieving state-of-the-art F1 scores. In addition, compared to existing methods that do not explicitly learn identity-independent features, our method maintains effectiveness on DISFA dataset with fewer facial identities, showing that our method is capable of learning more generalized AU detection features. Mengxin Shi, Wenming Zheng |
IEEE Trans. Affect. Comput. | 2 |
| 2026 | Robust Multimodal Sentiment Analysis Based on Adaptive Information Distillation and Adversarial Learning
Ning Sun 0005, Weiliang Zhang, Wenming Zheng, Jixin Liu 0001, Lei Chai, Cong Wu 0007 |
IEEE Trans. Affect. Comput. | 3 |
| 2026 | Mixture-of-Expert Large Language Models for Text-Based Personality Assessment From Asynchronous Video InterviewsabstractIn selection and assessment, Large Language Models (LLMs) are deemed suitable for personality assessment of Asynchronous Video Interviews (AVI) due to their advanced linguistic understanding and semantic interpretation capacities. However, most of the previous works have focused on personality traitclassificationrather thanregression. Since current LLMs are trained on massive corpora, they are more attuned to text-based structures than to numeric data. The classification-centered approach fails to take into account the fact that personality traits are continuous, instead of categorical, variables. In addition, LLMs suffer from low rating validity due to their tendency to be over-lenient, assigning higher than average personality scores to a large number of individuals (i.e., issues ofPositivity Biases). To address these challenges, we designed a text-based, two-stage, Mixture-of-Experts based personality assessment framework (MoE-Personality) to provide fine-grained personality ratings and regulate the positivity biases of LLMs. The designed model first rates the coarse-grained personality category of the individual (classification stage). After that, the model rates fine-grained personality scores by merging the obtained personality category and empirical score ranges of different personality categories. Inspired by the the way human annotators rate personality traits, each stage comprises multiple annotator LLM experts and one aggregator LLM expert to promote validity. Our experiments in two datasets show that the designed framework outperformed open-sourced, medium-sized LLMs (e.g., Llama 3.1-8B, qwen2-7B) and achieved comparable results with close-sourced, large-sized LLMs (e.g., GPT-3.5 and GPT-4). Tianyi Zhang 0013, Shan Liang 0003, Wenming Zheng, Antonis Koutsoumpis, Janneke K. Oostrom, Reinout E. de Vries |
IEEE Trans. Affect. Comput. | 3 |
| 2026 | Adaptive Key Role Guided Hierarchical Relation Inference for Enhanced Group-Level Emotion RecognitionabstractIn this paper, we propose a novel hierarchical relational network, termed Key Role Guided Hierarchical Relation Inference (KR-HRI), for enhanced group-level emotion recognition (GER). Unlike existing methods that adopt a coarse-grained approach to model interactions among all individuals, our approach adaptively identifies and emphasizes key individuals who play a crucial role in conveying group-level emotions. By integrating coarse-grained relationship modeling with fine-grained key individual enhancement and leveraging global scene information, our method effectively refines discriminative feature generation while minimizing irrelevant interference. We introduce a Multi-branch Interaction Module (MIM) to dynamically fuse features from both the global scene and local individual branches using a localized mask integration strategy. This comprehensive approach enhances the interaction between global and local features, resulting in robust group-level emotion representations. Extensive experiments on three widely adopted GER datasets demonstrate that our framework consistently outperforms state-of-the-art methods, validating the effectiveness and robustness of our proposed approach. Qing Zhu 0002, Qirong Mao, Wenlong Dong, Xiuyan Shao, Xiaohua Huang 0003, Wenming Zheng |
IEEE Trans. Affect. Comput. | 6 |
| 2026 | Dual-Hypergraph Based Symmetrical Self-Representation Learning for Cross-Domain Facial Expression Recognition
Yuhan Cheng, Peng Song 0002, Siqi Fu, Xingxin Wan, Changjia Wang, Wenming Zheng |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2026 | Coupled Sparse Subspace Alignment-Based Domain Adaptation for Speech Emotion RecognitionabstractSpeech emotion recognition (SER) is crucial for human–computer interaction (HCI), yet remains a challenge in cross-domain scenarios. Emotional expressions vary significantly across speakers, languages, and recording conditions, leading to serious domain shifts. Coupled subspace learning has recently attracted considerable attention in domain adaptation (DA) as an effective approach to mitigating domain discrepancies by capturing common and domain-specific information. However, existing algorithms suffer from two limitations: 1) most methods directly adopt classifiers [e.g., support vector machine (SVM)] to incorporate discriminative information of the target domain, but such strategies lack flexibility and adaptability; and 2) the features learned from coupled projection matrices are redundant and poorly discriminative. To address these issues, we propose a novel DA approach named coupled sparse subspace alignment (CSSA) for cross-domain SER. Specifically, CSSA first performs latent representation learning on the unlabeled target domain data, in which the latent representation matrix is then optimized into a pseudolabel matrix to provide emotional guidance. Meanwhile, it models the source and target domains separately by sparse regression, thereby learning both discriminative and domain-specific information. Subsequently, CSSA performs coupled subspace alignment to reduce the domain discrepancy, where the dual projection matrices are progressively aligned to enhance their similarity. Additionally, a graph Laplacian regularization is applied to the cross-domain data to capture the local geometric structure. Extensive experiments on five public SER datasets demonstrate the superiority of CSSA over state-of-the-art DA methods. Siqi Fu, Peng Song 0002, Wenming Zheng |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2026 | A Survey on Deep Learning for Group-Level Emotion RecognitionabstractWith the rapid advancement of artificial intelligence, group-level emotion recognition (GER) has emerged as an important domain in human behavior analysis. Early GER methods primarily relied on handcrafted features. However, the recent success of deep learning has shifted the focus toward neural network-based solution, enabling more effective exploitation of the rich visual and contextual cues in group images and videos. Unlike individual-level emotion recognition, GER must account for the diversity and dynamics of multiple individuals within varied social contexts. Over the past decade, numerous deep learning-based methods have been proposed, achieving substantial performance gains. This survey provides a comprehensive review of deep learning-centric review of GER, introducing a new taxonomy that spans representation learning, graph-based modeling, attention and transformer architectures, and multimodal fusion strategies. We summarize benchmark datasets, outline prevailing GER pipelines, and consolidate performance trends from recent state-of-the-art approaches. In addition, we discuss the integration of foundation models and large language model-guided multimodal reasoning into GER. Key challenges are identified, and potential research directions are proposed to support the development of robust, real-world GER systems. This work aims to serve as a pivotal reference for future research in this evolving field. Xiaohua Huang 0003, Xiaopeng Hong, Qirong Mao, Wenming Zheng, Abhinav Dhall |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2026 | Dynamic Graph Consistent Weighted Subspace Learning for Cross-Domain Speech Emotion RecognitionabstractIn recent years, cross-domain speech emotion recognition (SER) has attracted considerable interest. Most transfer subspace learning based SER methods lack unified adaptive constraints, making it difficult to balance discriminative capability and domain alignment, which limits their cross-domain generalization. To address these problems, we propose a novel domain adaptation (DA) approach called dynamic graph consistent weighted subspace learning (DGCWSL). Specifically, DGCWSL first projects samples from the source and target domains into a shared low-dimensional discriminative subspace, then performs cross-domain instance reconstruction, representing each target as a weighted combination of source instances. In parallel, a dynamic graph is constructed to capture local structural information between domains while preserving the data manifold. Subsequently, label supervision and discriminative learning between domains are achieved through linear regression. Furthermore, we introduce an adaptive weighted matrix that enforces consistent feature contributions across the distance metric, instance alignment, and discriminative regression, thereby mitigating both overfitting and underfitting. Finally, extensive experiments are conducted on four public datasets. The results confirm the superiority of DGCWSL over several state-of-the-art DA methods. Peng Song 0002, Siqi Fu, Zhaowei Liu 0001, Changjia Wang, Wenming Zheng |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2026 | Trend-Aware Multiscale Spatial-Temporal Graph Convolution Network for P300 DetectionabstractP300-based brain–computer interfaces (BCIs) enable direct communication between the brain and external devices by decoding P300 potentials, playing a vital role in rehabilitation and cognitive neuroscience research. Accurate detection of P300 potentials is essential for the successful implementation of P300-based BCIs. Current mainstream P300 detection algorithms are based on multichannel electroencephalogram (EEG) signals, and while promising detection results have been achieved, they still suffer from the following issues: 1) insufficient exploitation of the distinct and characteristic overall temporal trend of P300 potentials; and 2) oversimplified aggregation of multichannel EEG information without effectively utilizing the non-Euclidean topological relationships between EEG channels. To address the above issues, we propose a trend-aware multiscale spatial-temporal graph convolutional neural network (TMSGCN) for P300 detection. Specifically, to effectively capture the long-term temporal trend of P300 potentials to improve detection robustness, we explicitly extract the trend component from raw EEG signals along time dimension and utilize it to assist the identification of P300 potentials. Subsequently, a multiscale temporal convolution module (MTCM) is applied to extract multitime scale amplitude features from the processed EEG signals, serving as input for an adaptive graph convolution module (AGCM). The AGCM effectively models the intricate inter-channel relationships as a graph by complementarily considering the structural and functional connectivity of human brain, thereby capturing more discriminative and informative spatial features related to P300 potentials. Extensive experiments on three datasets demonstrate the superiority of TMSGCN over other state-of-the-art P300 detection methods. Furthermore, the results of ablation study and visualization experiments indicate the effectiveness of each component in TMSGCN. Jincen Wang, Wenming Zheng, Yuan Zong |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2026 | Feature Evaluation and Joint Interaction for Audio-Visual Emotion RecognitionabstractAutomatic emotion recognition has attracted significant attention due to its potential applications in various real-world scenarios. Methods that integrate visual and audio modalities have become increasingly prominent because of their superior information-carrying capacity and complementarity. Despite advancements in feature fusion between video and audio, existing modality fusion-based methods struggle to effectively address the dynamic changes in feature quality caused by interference, which is common in emotion recognition in the wild tasks. To overcome this limitation, we propose a Parameter-Free Feature Evaluation and Interaction (PFFEI) model based on information quality assessment. The model leverages the scaling factor γ of the normalization layer to evaluate information quality and dynamically adjusts the degree of interaction between modalities, suppressing the impact of low-quality features affected by interference. Additionally, the norm constraint integrated into the model ensures that the γ value consistently measures feature quality across different modalities. This approach effectively mitigates the effects of modality imbalance and significantly enhances the model’s accuracy. The effectiveness of our method is demonstrated through experiments on three challenging real-world emotion datasets: DFEW, AFEW, and Ekman6. The results show that the PFFEI model outperforms state-of-the-art methods, achieving significant improvements of 8.71% (UAR) and 8.61% (WAR) on the AFEW database. Sunan Li, Cheng Lu 0005, Yuan Zong, Hailun Lian, Wenming Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | SFCC: A Scalable and Flexible RDMA Congestion Control AlgorithmabstractThe rapid evolution of cloud computing, big data, and Artificial Intelligence (AI) technologies has created an urgent demand for ultra-high bandwidth and ultra-low latency in modern data centers. Although Remote Direct Memory Access (RDMA) technology has been widely adopted, traditional congestion control algorithms such as Priority-based Flow Control (PFC) and Data Center Quantized Congestion Notification (DCQCN) face significant challenges in scalability, deployment flexibility, and tail latency management in complex network environments. To address these limitations, we propose SFCC, a sender-driven rate control solution based on Round-Trip Time (RTT) that requires no switch configuration. SFCC incorporates an adaptive Additive Increase Multiplicative Decrease (AIMD) control loop for enhanced scalability and latency reduction, an RTT re-estimation mechanism to correct deviations, dynamic RTT calculation for optimized queue management, and Negative Acknowledgment (NAK) signal utilization for fast convergence during packet loss. Implemented on commercial RDMA Network Interface Cards (NICs) using a Programmable Congestion Control (PCC) platform, SFCC has been evaluated through small-scale testbeds, storage network testbeds, and NS-3 simulations of ultra-large-scale incast scenarios. Experimental results demonstrate that SFCC significantly improves throughput while reducing switch queue lengths, decreases the 99th percentile tail latency by 55.2 %, maintains high reliability under packet loss conditions, and achieves up to$\mathbf{7 5. 4 \%}$reduction in flow completion time for small flows compared to DCQCN, along with faster convergence speed. Anran Xu 0003, Wenming Zheng, Biyao Che, Yonghang Zhang, Xiaoping Fan, Luren Liu |
HiPC | 3 |
| 2025 | APSCC: Adaptive Congestion Control for Packet-Sprayed RDMA Networks in AI ClustersabstractLarge Language Model (LLM) training increasingly relies on Remote Direct Memory Access (RDMA) to enable ultra-efficient networking. However, the unique traffic characteristics—sparse yet bandwidth-intensive—often lead to severe load imbalance under Equal-Cost Multi-Path (ECMP) routing. Packet Spraying (PS) offers a promising solution by distributing traffic across multiple paths, but its impact on congestion dynamics remains insufficiently studied. This paper presents a comprehensive study of PS in Artificial Intelligence (AI) clusters using NS-3 simulations, analyzing its effects on congestion distribution, packet reordering, and flow completion time. Our findings show that congestion patterns vary significantly with workload intensity and oversubscription ratios, and existing congestion control schemes are inadequate for general PS networks, where both routing paths and congestion hotspots frequently change. To address this gap, we propose APSCC, a congestion control algorithm that infers congestion locations from out-of-order packets and aggregates Explicit Congestion Notification (ECN) signals across paths for precise rate adaptation. Compared to state-of-the-art mechanisms, APSCC reduces Job Completion Time (JCT) by up to 30 %. The implementation is publicly available at https://github.com/tangjianback/APSCC. Wenming Zheng, Fangfang Yan, Xiaoping Fan, Luren Liu, Anran Xu 0003 |
HPCC | 2 |
| 2025 | Enhancing Task-Specific Feature Learning with LLMs for Multimodal Emotion and Intent Joint UnderstandingabstractThis paper introduces our solution, the Task-Specific Feature Learning (TSFL) method, designed to address the second track of the MEIJU Challenge at ICASSP 2025, namely, Imbalanced Emotion and Intent Recognition (English). The TSFL method incorporates three core components: the use of LLM features to represent multimodal signals, coarse-grained task-specific feature decomposition, and fine-grained task-specific feature learning. These components enable the effective joint learning of emotion-discriminative and intent-discriminative features. As a result, our method achieved a JRBM score of 0.6230, significantly outperforming the official baseline result and surpassing all other competing teams to win the championship. Cheng Lu 0005, Kaifei Zhang, Yujia Gu, Banghua Li, Yuan Zong, Wenming Zheng |
ICASSP | 8 |
| 2025 | Unsupervised Motion-Robust Self-Distillation Framework for Remote Physiological MeasurementabstractRemote photoplethysmography (rPPG) holds great potential in medical surveillance. However, head movements commonly encountered in real-world scenarios often degrade physiological estimation performance, particularly for unsupervised learning methods based on physiological frequency band priors, which usually struggle to detect dynamic signals occurring within this band, thereby limiting their performance ceiling. In this paper, we propose a novel strategy to endow unsupervised learning methods with motion robustness. Specifically, we introduce a simple motion simulation technique, Sliding Crop, to incorporate dynamic signals. Based on this, we develop an unsupervised motion-robust self-distillation framework (UMoRo) with the existing unsupervised learning method, where the model leverages its own high-quality mappings of simple samples as pseudo-labels to guide the learning process of suppressing simulated motion artifacts in challenging samples, thus enhancing motion robustness in real-world movements. Experimental results on three public datasets show that our method achieves superior or competitive performance compared to state-of-the-art supervised methods, demonstrating outstanding motion robustness. Anbang Liu, Shanlin Xiao, Wenming Zheng |
ICASSP | 3 |
| 2025 | Enhancing Zero-Shot Emotional Voice Conversion via Speaker Adaptation and Duration PredictionabstractZero-shot Emotional Voice Conversion (EVC) aims to transform a speaker’s emotional state to match a target emotion, even for speakers and emotion categories that were not encountered during training, thereby enhancing the generalization ability of traditional EVC systems. Despite advancements in the field, existing methods often face challenges in preserving speaker identity and ensuring the naturalness of emotional expression, particularly in the context of rhythm modeling. To this end, we propose the Zero-Shot Emotion Voice Conversion (ZSEVC) model, which leverages self-supervised learning for speaker adaptation and duration prediction. To adjust speech rhythm in alignment with the target emotional state, we introduce a rhythm-aware content encoder that captures and refines discrete speech units at a finer granularity. Additionally, a hierarchical emotion fusion scheme is employed to integrate emotional features with content features, enhancing both pronunciation accuracy and emotional expressiveness. Moreover, a residual speaker-emotion fusion module is incorporated to better adapt speaker characteristics to emotional prosodic variation. Experimental results show ZSEVC’s superior performance in terms of naturalness and speaker similarity in zero-shot scenario, successfully generating emotional speeches for unseen emotions and speakers. Speech samples are available at https://wosyoo.github.io/ZSEVC. Shiyan Wang, Tianhua Qi, Cheng Lu 0005, Zhaojie Luo, Wenming Zheng |
ICASSP | 5 |
| 2025 | Reliable Learning From LLM Features for Multimodal Emotion and Intent Joint UnderstandingabstractThis paper describes a Reliable Learning Framework (RLF) for the 1st Multimodal Emotion and Intent Joint Understanding (MEIJU) Challenge at ICASSP 2025. Our proposed RLF includes a Hierarchical Interaction Network and a Reliable Fusion Strategy. The former can excavate emotion and intent cues from the high-level semantic features of multimodal data (video, audio, and text) generated by pretrained Large Language Models (LMMs), to enhance their representations, and the latter reliably integrates multiple predictions to further improve the robustness of emotion and intent understanding. Our RLF method achieved first place on Track 2 (Mandarin) of MEIJU, with performance scores for emotion, intent, and joint recognition reaching 0.7285, 0.7456, and 0.7370. Cheng Lu 0005, Yuyun Liu, Yinghao Ma, Jiahao Luo, Yuan Zong, Wenming Zheng |
ICASSP | 8 |
| 2025 | Heterogeneous Graph Convolutional Neural Networks for EEG-fNIRS Bimodal Emotion RecognitionabstractLeveraging multimodal brain signals, such as electroencephalogram (EEG) and functional near-infrared spectroscopy (fNIRS), for the objective detection of brain activity is regarded as a promising approach for affective brain-computer interface. Existing EEG-fNIRS bimodal methods primarily focus on data alignment and basic feature fusion, neglecting the intrinsic connections between different signals and brain regions. To this end, we propose a novel graph-based method, the Heterogeneous Graph Convolutional Neural Network (HGCN), which integrates multimodal complementary information and constructs a heterogeneous EEG-fNIRS interaction graph within a coordinated hyperspace to model brain networks. Specifically, various directed edges types and node feature aggregation strategies are employed to dynamically update and enhance the representation of brain signals and emotional states, thereby improving the coherence and consistency of cross-modal signals. Furthermore, this approach provides a robust framework for exploring cross-modal spatial connectivity. In this paper, we also developed a novel EEG-fNIRS emotion database collected from 17 subjects by video stimuli. Extensive experimental results demonstrate the superiority of our method and the effectiveness of the introduced heterogeneous EEG-fNIRS interaction graph. Yunlong Xue, Wenming Zheng |
ICASSP | 4 |
| 2025 | DisenEmo: Learning disentangled emotional representation from facial motion for 3D talking head generationabstractEmotional 3D talking head generation synthesizes vivid facial expressions with precise lip synchronization for immersive interactions. We introduce DisenEMO, a novel framework designed to disentangle emotion and content from facial motions, thereby facilitating the synthesis of personalized and expressive audio-driven facial animations. To achieve precise emotional disentanglement, we incorporate an intensity perception constraint which improves the accurate perception of categorized emotion and its intensities, leading to the generation of subtle emotional expressions. To ensure the temporal consistency of facial expressions, we introduce facial dynamic modeling, which refines motion trajectories to better capture emotional nuances. Finally, a motion decoder integrates emotional features with audio features extracted from driving speech, producing 3D talking heads with enhanced emotional expressiveness and realism. Experimental results demonstrate that our method outperforms state-of-the-art approaches. Synthesis samples are available at https://c295bw.github.io/DisenEMO-icip25.github.io/. Tianhua Qi, Cheng Lu 0005, Wenming Zheng |
ICIP | 4 |
| 2025 | A Hybrid Graph Neural Network for Enhanced EEG-Based Depression DetectionabstractGraph neural networks (GNNs) are gaining increasing popularity for EEG-based depression detection. However, previous GNN-based methods inadequately consider the characteristics of depression, which limit their performance. First, neuroscience studies indicate that patients with depression exhibit both common and individualized brain abnormalities. Previous GNN-based approaches typically focus either on common graph connections to capture common brain abnormalities or on individualized connections to capture individualized patterns, which is insufficient for depression detection. Second, brain network exhibits a hierarchical structure, ranging from channel-level graphs to region-level graphs. This hierarchical structure varies across individuals and contains significant information relevant to detecting depression. However, previous GNN-based methods overlook this individualized hierarchical information. To address these issues, we propose a Hybrid GNN (HybGNN) that combines a Common Graph Neural Network (CGNN) branch using common connections and an Individualized Graph Neural Network (IGNN) branch employing individualized connections. The two branches capture common and individualized depression patterns, respectively, complementing each other. Furthermore, we enhance the HybGNN with a Cross-Branch Hierarchical Information Extractor (CB-HIE) to extract more task-relevant individualized hierarchical information. Extensive experiments on the MODMA and HUSM datasets demonstrate that the proposed HybGNN achieves state-of-the-art performance. Yiye Wang, Wenming Zheng, Yang Li 0019, Hao Yang 0006 |
IJCNN | 2 |
| 2025 | Interactive Fusion of Multi-View Speech Embeddings via Pretrained Large-Scale Speech Models for Speech Emotional Attribute Prediction in Naturalistic Conditions
Yuyun Liu, Yujia Gu, Jiahao Luo, Wenming Zheng, Cheng Lu 0005, Yuan Zong |
INTERSPEECH | 4 |
| 2025 | PromptEVC: Controllable Emotional Voice Conversion with Natural Language Prompts
Tianhua Qi, Shiyan Wang, Cheng Lu 0005, Tengfei Song, Hao Yang 0006, Zhanglin Wu, Wenming Zheng |
INTERSPEECH | 7 |
| 2025 | NaME: A Natural Micro-expression Dataset for Micro-expression Recognition in the WildabstractMicro-expressions (MEs) are involuntary facial expressions that reveal genuine emotions and have significant applications in fields such as psychology, security, and human-computer interaction. However, previous ME datasets are mainly collected in controlled laboratory environments, such as fixed views, single illumination and head movements, limited subjects and the lack of background. There are significant gaps between them and the real world. To handle this issue, we introduce a novel Natural Micro-Expression (NaME) dataset, a natural dataset collected under unconstrained real-world conditions. It encompasses (1) diverse subjects, multiple views and varying head movements ; (2) rich background information, providing a more realistic benchmark for the micro-expression recognition (MER) research. Furthermore, we propose a MER benchmark for natural environments, named MixFormer. MixFormer includes an efficient sparse attention mechanism to capture subtle facial motions from various factors, and a face-background mix of attention module to model the environment context to help MER. Extensive experiments are conducted to analyze our NaME dataset and benchmark. We believe that our dataset and benchmark will pave the way for future research in MER beyond controlled settings, facilitating the deployment of MER in practical applications. NaME is available at github.com/real-ljt/NAMEdataset. Jiateng Liu, Hengcan Shi, Haiwen Liang, Yuan Zong, Yaonan Wang 0001, Wenming Zheng |
ACM Multimedia | 7 |
| 2025 | Multi-Level Segment Fusion Based on Adaptive Time-Window Selection for Multimodal Personality-Aware Elderly Depression DetectionabstractMajor Depressive Disorder (MDD) is a prevalent and severe psychiatric disorder, and its detection remains challenging due to the complexity and variability of its symptoms. Traditional single-modality methods often fail to capture the full spectrum of depressive cues, which has led to the rise of multimodal methods. The ACM Multimedia 2025 ''Multimodal Personality-Aware Depression Detection Challenge'' (MPDD 2025) aims to advance the development of more accurate depression detection models by incorporating multimodal data. In this paper, we proposed a Multi-Level Segment Fusion Based on Adaptive Time-Window Selection (MSF-ATS) method for the MPDD-Elderly Track. To address the challenge of sparse and transient depressive symptoms, we fuse segment-level classifications to obtain subject-level classifications. An adaptive time-window selection based on mean class variance is employed to choose the window with the smallest variance for more stable detection results. Our method achieved an average score of 0.8576 on the MPDD 2025 official test set, significantly outperforming the baseline score of 0.6675. Yuyun Liu, Kaifei Zhang, Yinghao Ma, Tianhua Qi, Wenming Zheng, Cheng Lu 0005, Yuan Zong |
ACM Multimedia | 6 |
| 2025 | Assessing Personality Traits and Interview Performance from Asynchronous Video InterviewsabstractAsynchronous Video Interviews (AVIs) allow candidates to record responses to predefined questions using digital devices, offering both flexibility and remote accessibility. Assessing personality traits and interview performance via AVIs provides organizations with valuable insights into candidate profiles and facilitates the prediction of future job performance. However, prior benchmark challenges, whose datasets were predominantly sourced from social media, suffer from suboptimal construct and methodological validity, limiting their utility for model development and real-world applications. To address these limitations, we introduce the AVI Grand Challenge at ACM Multimedia 2025, featuring a novel dataset of mock AVIs comprising 3,876 videos from 646 participants in a simulated job application procedure. Interview questions were carefully designed to reflect real-world selection contexts and elicit personality expressions grounded in Trait Activation Theory. Personality traits and job competencies were annotated by trained evaluators and professional recruiters, ensuring both methodological rigor and ecological validity. The solutions and algorithms developed in this challenge are analyzed and summarized in this paper to foster the development of fair, reliable, and AI-driven hiring assessments. Tianyi Zhang 0013, Tianhua Qi, Antonis Koutsoumpis, Yuan Zong, Wenming Zheng, Janneke K. Oostrom, Djurre Holtrop, Zhaojie Luo, Reinout E. de Vries |
ACM Multimedia | 5 |
| 2025 | AEP: An adaptive ensemble P300-BCI classifier based on user-feedback and knowledge-transfer
Qingzhi Chen, Xuewei Chen, Wenming Zheng, Zhixiong Lin |
Appl. Intell. | 4 |
| 2025 | High-order correlation preserved multi-view unsupervised feature selection
Meng Duan, Peng Song 0002, Shixuan Zhou, Yuanbo Cheng, Jinshuai Mu, Wenming Zheng |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | Essential anchor graph learning for incomplete multi-view clustering
Peng Song 0002, Jinshuai Mu, Yuanbo Cheng, Zhaohu Liu, Wenming Zheng |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Weighted tensor-based consistent anchor graph learning for multi-view clustering
Guanghao Du, Peng Song 0002, Yuanbo Cheng, Zhaowei Liu 0001, Yanwei Yu, Wenming Zheng |
Neurocomputing | 6 |
| 2025 | Enhancing bone-conducted speech with spectrum similarity metric in adversarial learning
Jian Zhou 0006, Wenming Zheng, Hon Keung Kwan |
Speech Commun. | 4 |
| 2025 | Multimodal Sentimental Privileged Information Embedding for Improving Facial Expression RecognitionabstractFacial expression recognition (FER) has always been one of the key task in affective computing. Over the years, researchers have worked to improve the performance of FER by designing models with more powerful feature extraction, embedding attention mechanism, and reconstructing missing information, etc. Different from the paradigms above, we attempt to improve FER performance by using multimodal sentiment data, such as audio and text, as privileged information (PI) for facial images. To this end, a multimodal privileged information embedded facial expression recognition network (MPI-FER) is proposed in this paper. During the training phase, this model achieves the PI embedding of multimodal data for FER by developing cross-modality translation between multimodal sentiment data. During the test phase, input images alone are sufficient for the model inference to accomplish the FER task input. The MPI-FER is a large-scale, heterogeneous deep neural network. To achieve effective training of this model with limited training samples, we design a multi-stage training strategy of module-wise pre-training followed by end-to-end fine-tuning. In addition, a strategy of filling the multimodal sentiment quaternion is proposed for implementing our method on a facial expression database consisting only of face images. We conducted extensive experiments to evaluate the proposed method on two databases of multimodal sentiment analysis (CH-SIMS and CMU-MOSI) and two databases of FER in the wild (RAF-DB and AffectNet). The results show that embedding multimodal sentiment data as privileged information into the FER task based on face images can significantly improve the accuracy of FER. Furthermore, by only using image in the test phase, the proposed method can achieve better results of multimodal sentiment analysis than those methods achieved by using multimodal sentimental data fusion. Ning Sun 0005, Changwei You, Wenming Zheng, Jixin Liu 0001, Lei Chai, Haian Sun |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | Dynamical Causal Graph Neural Network for EEG Emotion RecognitionabstractRecently, topological graphs based on structural or functional connectivity of brain network have been utilized to construct graph neural networks (GNN) for Electroencephalogram (EEG) emotion recognition. In this paper, we propose a novel dynamical causal graph neural network (DCGNN) based on the effective causal connectivity of brain function network for EEG emotion recognition, in which Greedy Equivalence Search (GES) is used to find the optimal causal graph topology associated with the adjacent matrix of DCGNN. To this end, we firstly construct a skeleton graph using canonical correlation analysis (CCA) and then use GES to optimize the directional graph topology of DCGNN. Then, learnable weight parameters associated with the adjacent matrix are learnt during the model training of DCGNN. Additionally, a sparse graphic constraint is employed to enhance the efficacy of emotion recognition, while a Conditional Domain Adversarial Network (CDAN) is used to integrate features with emotion labels for improving subjectindependent validation. Extensive experiments and ablation studies are conducted on five public datasets, i.e., SEED, SEED-IV, SEED-V, MPED, and FACED, demonstrating that the proposed model surpasses recent causal based (Granger causality) and domain adaptation based GNN models across all experimental settings. Yushun Xiao, Wenming Zheng, Guoying Zhao 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | Decoupled Multi-Perspective Fusion for Speech Depression DetectionabstractSpeechDepressionDetection (SDD) has garnered attention from researchers due to its low cost and convenience. However, current algorithms lack methods for extracting interpretable acoustic features based on clinical manifestations. In addition, effectively fusing these features to overcome individual heterogeneity remains a challenge. This study proposes a decoupled multi-perspective fusion (DMPF) model. The model extracts five key features of voiceprint, emotion, pause, energy, and tremor based on the multi-perspective clinical manifestations. These features are then decoupled into common and private features, which fused through graph attention network to obtain the comprehensive depression representation. Notably, this study has collected a depression speech dataset, which includes standardized and comprehensive tasks along with diagnostic labels provided by psychologists. Extensive subject-independent experiments were conducted on the DAIC-WOZ, MODMA and MPSC datasets. The voiceprint features can automatically cluster the depressed and non-depressed populations. Furthermore, DMPF can effectively fuse common and private features from different perspectives, achieving AUC of 84.20%, 85.34%, 86.13% on three datasets. The results illustrate the interpretability of multi-perspective features and demonstrate that the combination of speech manifestations can enhance the detection ability, which can provide a multi-perspective observational tool for physicians and clinical practice. Hongxiang Gao, Fei Wang 0064, Wenming Zheng, Jianqing Li 0002, Chengyu Liu 0001 |
IEEE Trans. Affect. Comput. | 6 |
| 2025 | Controllable Multi-Speaker Emotional Speech Synthesis With an Emotion Representation of High Generalization CapabilityabstractThe aim of multi-speaker emotional speech synthesis is to generate speech for a designated speaker in a desired emotional state. The task is challenging due to the presence of speech variations, such as noise, content, and timbre, which can obstruct emotion extraction and transfer. This paper proposes a new approach to performing multi-speaker emotional speech synthesis. The proposed method, which is based on a seq2seq synthesizer, integrates emotion embedding as a conditioned variable to convey exact emotional information from reference audio to the synthesized speech. To boost emotion representation capability, we utilize a three-dimensional acoustic feature as input. And an emotion generalization module with adaptive instance normalization (AdaIN) is proposed to obtain emotion embedding with high generalization ability, which also results in improved controllability. The derived emotion embedding from the generalization module can be readily conditioned by affine parameters, allowing for control both the emotion category and the emotion intensity of synthesized speech. Various emotional speech synthesis experimental results of the propposed method demonstrate its state-of-the-art performance in multi-speaker emotional speech synthesis, coupled with its advantage of high emotion controllability. Jian Zhou 0006, Wenming Zheng, Hon Keung Kwan |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | An Empirical Study of Super-Resolution on Low-Resolution Micro-Expression Recognition
Ling Zhou 0005, Mingpei Wang, Xiaohua Huang 0003, Wenming Zheng, Qirong Mao, Guoying Zhao 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | Towards a Robust Group-Level Emotion Recognition via Uncertainty-Aware LearningabstractGroup-level emotion recognition (GER) is an inseparable part of human behavior analysis, aiming to recognize an overall emotion in a multi-person scene. However, the existing methods are devoted to combing diverse emotion cues while ignoring the inherent uncertainties under unconstrained environments, such as congestion and occlusion occurring within a group. Additionally, since only group-level labels are available, inconsistent emotion predictions among individuals in one group can confuse the network. In this paper, we propose an uncertainty-aware learning (UAL) method to extract more robust representations for GER. By explicitly modeling the uncertainty, we adopt stochastic embedding sourced from a Gaussian distribution instead of deterministic point embedding. It helps capture the probabilities of emotions and facilitates diverse inferences. Additionally, we adaptively assign uncertainty-sensitive scores as the fusion weights for individuals’ faces within a group. Moreover, we developed an image enhancement module to evaluate and filter samples, strengthening the model’s data-level robustness against uncertainties. The overall three-branch model, encompassing face, object, and scene components, is guided by a proportional-weighted fusion strategy and integrates the proposed uncertainty-aware method to produce the final group-level output. Experimental results demonstrate the effectiveness and generalization ability of our method across three widely used databases. Qing Zhu 0002, Qirong Mao, Xiaohua Huang 0003, Wenming Zheng |
IEEE Trans. Affect. Comput. | 5 |
| 2025 | Learning to Rank Onset-Occurring-Offset Representations for Micro-Expression RecognitionabstractThis paper focuses on the research of micro-expression recognition (MER) and proposes a flexible and reliable deep learning method called learning to rank onset-occurring-offset representations (LTR3O). The LTR3O method introduces a dynamic and reduced-size sequence structure known as 3O, which consists of onset, occurring, and offset frames, for representing micro-expressions (MEs). This structure facilitates the subsequent learning of ME-discriminative features. A noteworthy advantage of the 3O structure is its flexibility, as the occurring frame is randomly extracted from the original ME sequence without the need for accurate frame spotting methods. Based on the 3O structures, LTR3O generates multiple 3O representation candidates for each ME sample and incorporates well-designed modules based on learning to rank (LTR) to measure and calibrate their emotional expressiveness. This calibration process implicitly enhances the visibility of MEs by amplifying the originally narrow emotional expressiveness gap among ME frames caused by their low-intensity characteristics, thereby facilitating the reliable learning of more discriminative features for MER. Extensive experiments were conducted to evaluate the performance of LTR3O using four widely-used ME databases: CASME II, SMIC, SAMM, and MEVIEW. The experimental results demonstrate the effectiveness and superior performance of LTR3O, particularly in terms of its flexibility and reliability, when compared to recent state-of-the-art MER methods. Yuan Zong, Jingang Shi, Cheng Lu 0005, Hongli Chang, Wenming Zheng |
IEEE Trans. Affect. Comput. | 6 |
| 2025 | Common Discriminative Latent Space Learning for Cross-Domain Speech Emotion RecognitionabstractCross-domain speech emotion recognition (SER) has received increasing attention in recent years. Existing transfer subspace learning and regression-based SER methods have the following drawbacks. The features in the subspace are still insufficiently representative and discriminative, and direct regression would lead to information loss. To address these problems, we present a novel common discriminative latent space learning (CDLSL) method for cross-domain SER. To be specific, we first obtain a common latent space by imposing a projection matrix on the cross-domain data. Meanwhile, we impose an uncorrelated constraint on the projection matrix to ensure that the features are representative and discriminative after dimension reduction. Then, we implement a graph regularization term on the latent representations of the samples to capture the local similarity information. Furthermore, to obtain a more discriminative common latent space, we introduce the label information by aligning the latent space with the relaxed label space, while mitigating the information loss for regression. Extensive experimental results validate the superiority of the proposed method over the state-of-the-art competitors. Siqi Fu, Peng Song 0002, Hao Wang 0269, Zhaowei Liu 0001, Wenming Zheng |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2025 | Towards Domain-Specific Cross-Corpus Speech Emotion Recognition ApproachabstractCross-corpus speech emotion recognition (SER) poses a challenge due to feature distribution mismatch between the training and testing speech samples, potentially degrading the performance of established SER methods. In this article, we tackle this challenge by proposing a novel transfer subspace learning method called acoustic knowledge-guided transfer linear regression (AKTLR). Unlike existing approaches, which often overlook domain-specific knowledge related to SER and simply treat cross-corpus SER as a generic transfer learning task, our AKTLR method is built upon a well-designed acoustic knowledge-guided dual sparsity constraint mechanism. This mechanism emphasizes the potential of minimalistic acoustic parameter feature sets to alleviate classifier over-adaptation, which is empirically validated acoustic knowledge in SER, enabling superior generalization in cross-corpus SER tasks compared to using large feature sets. Through this mechanism, we extend a simple transfer linear regression model to AKTLR. This extension harnesses its full capability to seek emotion-discriminative and corpus-invariant features from established acoustic parameter feature sets used for describing speech signals across two scales: contributive acoustic parameter groups and constituent elements within each contributive group. We evaluate our method through extensive cross-corpus SER experiments on three widely used speech emotion corpora: EmoDB, eNTERFACE, and CASIA. The proposed AKTLR achieves an average UAR of 42.12% across six tasks using the eGeMAPS feature set, outperforming many recent state-of-the-art transfer subspace learning and deep transfer learning methods. This demonstrates the effectiveness and superior performance of our approach. Furthermore, our work provides experimental evidence supporting the feasibility and superiority of incorporating domain-specific knowledge into the transfer learning model to address cross-corpus SER tasks. Yan Zhao 0037, Yuan Zong, Hailun Lian, Cheng Lu 0005, Jingang Shi, Wenming Zheng |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2025 | Decoupled Doubly Contrastive Learning for Cross-Domain Facial Action Unit DetectionabstractDespite the impressive performance of current vision-based facial action unit (AU) detection approaches, they are heavily susceptible to the variations across different domains and the cross-domain AU detection methods are under-explored. In response to this challenge, we propose a decoupled doubly contrastive adaptation (D2CA) approach to learn a purified AU representation that is semantically aligned for the source and target domains. Specifically, we decompose latent representations into AU-relevant and AU-irrelevant components, with the objective of exclusively facilitating adaptation within the AU-relevant subspace. To achieve the feature decoupling, D2CA is trained to disentangle AU and domain factors by assessing the quality of synthesized faces in cross-domain scenarios when either AU or domain attributes are modified. To further strengthen feature decoupling, particularly in scenarios with limited AU data diversity, D2CA employs a doubly contrastive learning mechanism comprising image and feature-level contrastive learning to ensure the quality of synthesized faces and mitigate feature ambiguities. This new framework leads to an automatically learned, dedicated separation of AU-relevant and domain-relevant factors, and it enables intuitive, scale-specific control of the cross-domain facial image synthesis. Extensive experiments demonstrate the efficacy of D2CA in successfully decoupling AU and domain factors, yielding visually pleasing cross-domain synthesized facial images. Meanwhile, D2CA consistently outperforms state-of-the-art cross-domain AU detection approaches, achieving an average F1 score improvement of 6%-14% across various cross-domain scenarios. Yong Li 0032, Menglin Liu, Zhen Cui 0001, Yi Ding 0012, Yuan Zong, Wenming Zheng, Shiguang Shan, Cuntai Guan |
IEEE Trans. Image Process. | 6 |
| 2025 | Decoding SSVEP Via Calibration-Free TFA-Net: A Novel Network Using Time-Frequency FeaturesabstractBrain-computer interfaces (BCIs) based on steady-state visual evoked potential (SSVEP) signals offer high information transfer rates and non-invasive brain-to-device connectivity, making them highly practical. In recent years, deep learning techniques, particularly convolutional neural network (CNN) architectures, have gained prominence in EEG (e.g., SSVEP) decoding because of their nonlinear modeling capabilities and autonomy from manual feature extraction. However, most studies using CNNs employ temporal signals as the input and cannot directly mine the implicit frequency information, which may cause crucial frequency details to be lost and challenges in decoding. By contrast, the prevailing supervised recognition algorithms rely on a lengthy calibration phase to enhance algorithm performance, which could impede the popularization of SSVEP based BCIs. To address these problems, this study proposes the Time-Frequency Attention Network (TFA-Net), a novel CNN model tailored for SSVEP signal decoding without the calibration phase. Additionally, we introduce the Frequency Attention and Channel Recombination modules to enhance ability of TFA-Net to infer finer frequency-wise attention and extract features efficiently from SSVEP in the time-frequency domain. Classification results on a public dataset demonstrated that the proposed TFA-Net outperforms all the compared models, achieving an accuracy of 79.00% $\pm$ 0.27% and information transfer rate of 138.82 $\pm$ 0.78 bits/min with a 1-s data length. TFA-Net represents a novel approach to SSVEP identification as well as time-frequency signal analysis, offering a calibration-free solution that enhances the generalizability and practicality of SSVEP based BCIs. Pan Lin, Yuankui Yang, Yue Leng, Wenming Zheng, Sheng Ge |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | Multi-Source Unsupervised Transfer Components Learning for Cross-Domain Speech Emotion RecognitionabstractAs an important research direction in the field of speech signal processing, cross-domain speech emotion recognition (SER) has attracted extensive attention. In practice, it is challenging to collect enough labeled samples from single source domain to train robust classifiers. To this end, this paper presents a novel method named multi-source unsupervised transfer components learning (MUTCL) for cross-domain SER. In MUTCL, we first adopt a PCA-like strategy and apply it to multi-source domains, aiming to preserve both intra-domain individuality and inter-domain commonality principal components within each domain. Simultaneously, a simple alignment strategy is developed to guide cross-domain samples to have similar structures, thus preserving more transfer components. Moreover, an adaptive weight strategy is utilized to determine the contribution of each source domain. We conduct experiments on five benchmark datasets, and the results show that MUTCL achieves excellent performance compared with some state-of-the-art methods. Shenjie Jiang, Peng Song 0002, Shaokai Li, Wenming Zheng |
ICASSP | 5 |
| 2024 | Improving Speaker-Independent Speech Emotion Recognition using Dynamic Joint Distribution AdaptationabstractIn speaker-independent speech emotion recognition, the training and testing samples are collected from diverse speakers, leading to a multi-domain shift challenge across the feature distributions of data from different speakers. Consequently, when the trained model is confronted with data from new speakers, its performance tends to degrade. To address the issue, we propose a Dynamic Joint Distribution Adaptation (DJDA) method under the framework of multi-source domain adaptation. DJDA firstly utilizes joint distribution adaptation (JDA), involving marginal distribution adaptation (MDA) and conditional distribution adaptation (CDA), to more precisely measure the multi-domain distribution shifts caused by different speakers. This helps eliminate speaker bias in emotion features, allowing for learning discriminative and speaker-invariant speech emotion features from coarse-level to fine-level. Furthermore, we quantify the adaptation contributions of MDA and CDA within JDA by using a dynamic balance factor based on $\mathcal{A}$-Distance, promoting to effectively handle the unknown distributions encountered in data from new speakers. Experimental results demonstrate the superior performance of our DJDA as compared to other state-of-the-art (SOTA) methods. Cheng Lu 0005, Yuan Zong, Hailun Lian, Yan Zhao 0037, Björn W. Schuller, Wenming Zheng |
ICASSP | 6 |
| 2024 | PAVITS: Exploring Prosody-Aware VITS for End-to-End Emotional Voice ConversionabstractIn this paper, we propose Prosody-aware VITS (PAVITS) for emotional voice conversion (EVC), aiming to achieve two major objectives of EVC: high content naturalness and high emotional naturalness, which are crucial for meeting the demands of human perception. To improve the content naturalness of converted audio, we have developed an end-to-end EVC architecture inspired by the high audio quality of VITS. By seamlessly integrating an acoustic converter and vocoder, we effectively address the common issue of mismatch between emotional prosody training and run-time conversion that is prevalent in existing EVC models. To further enhance the emotional naturalness, we introduce an emotion descriptor to model the subtle prosody variations of different speech emotions. Additionally, we propose a prosody predictor, which predicts prosody features from text based on the provided emotion label. Notably, we introduce a prosody alignment loss to establish a connection between latent prosody features from two distinct modalities, ensuring effective training. Experimental results show that the performance of PAVITS is superior to the state-of-the-art EVC methods. Speech Samples are available at https://jeremychee4.github.io/pavits4EVC/. Tianhua Qi, Wenming Zheng, Cheng Lu 0005, Yuan Zong, Hailun Lian |
ICASSP | 2 |
| 2024 | Speech Swin-Transformer: Exploring a Hierarchical Transformer with Shifted Windows for Speech Emotion RecognitionabstractSwin-Transformer has demonstrated remarkable success in computer vision by leveraging its hierarchical feature representation based on Transformer. In speech signals, emotional information is distributed across different scales of speech features, e. g., word, phrase, and utterance. Drawing above inspiration, this paper presents a hierarchical speech Transformer with shifted windows to aggregate multi-scale emotion features for speech emotion recognition (SER), called Speech Swin-Transformer. Specifically, we first divide the speech spectrogram into segment-level patches in the time domain, composed of multiple frame patches. These segment-level patches are then encoded using a stack of Swin blocks, in which a local window Transformer is utilized to explore local inter-frame emotional information across frame patches of each segment patch. After that, we also design a shifted window Transformer to compensate for patch correlations near the boundaries of segment patches. Finally, we employ a patch merging operation to aggregate segment-level emotional features for hierarchical speech representation by expanding the receptive field of Transformer from frame-level to segment-level. Experimental results demonstrate that our proposed Speech Swin-Transformer outperforms the state-of-the-art methods. Yong Wang 0073, Cheng Lu 0005, Hailun Lian, Yan Zhao 0037, Björn W. Schuller, Yuan Zong, Wenming Zheng |
ICASSP | 7 |
| 2024 | Progressively Learning from Macro-Expressions for Micro-Expression RecognitionabstractMicro-expression (ME) recognition is challenging due to the low-intensity facial motions. An idea to overcome this is learning assisted by macro-expressions (MaEs). However, the intensity gap between MaE and ME is so huge that related works fail to effectively leverage MaE’s assistance in overcoming low-intensity interference, which attempt to directly force ME knowledge to mimic MaE knowledge. In this paper, we propose that the knowledge transfer from MaE to ME can be converted into a progressive process for better implementation. Thus, we construct a progressive multi-step learning framework, which accomplishes two tasks: first, we dissect the huge intensity gap into multiple segments that are easier to bridge by constructing multiple learning steps, each corresponding to various intensity levels of expression recognition tasks. Second, through a designed self-knowledge distillation (self-KD) model, each dissected gap can be bridged, enabling the MaE knowledge to progressively transfer to guide the ME learning. Experiments carried out on three widely used databases demonstrated that the proposed PLMaM achieves state-of-the-art results. Yuan Zong, Mengting Wei, Cheng Lu 0005, Wenming Zheng |
ICASSP | 7 |
| 2024 | Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion RecognitionabstractCross-corpus speech emotion recognition (SER) aims to transfer emotional knowledge from a labeled source corpus to an unlabeled corpus. However, prior methods require access to source data during adaptation, which is unattainable in real-life scenarios due to data privacy protection concerns. This paper tackles a more practical task, namely source-free cross-corpus SER, where a pre-trained source model is adapted to the target domain without access to source data. To address the problem, we propose a novel method called emotion-aware contrastive adaptation network (ECAN). The core idea is to capture local neighborhood information between samples while considering the global class-level adaptation. Specifically, we propose a nearest neighbor contrastive learning to promote local emotion consistency among features of highly similar samples. Furthermore, relying solely on nearest neighborhoods may lead to ambiguous boundaries between clusters. Thus, we incorporate supervised contrastive learning to encourage greater separation between clusters representing different emotions, thereby facilitating improved class-level adaptation. Extensive experiments indicate that our proposed ECAN significantly outperforms state-of-the-art methods under the source-free cross-corpus SER setting on several speech emotion corpora. Yan Zhao 0037, Jincen Wang, Cheng Lu 0005, Sunan Li, Björn W. Schuller, Yuan Zong, Wenming Zheng |
ICASSP | 7 |
| 2024 | A Novel Decoupled Prototype Completion Network for Incomplete Multimodal Emotion RecognitionabstractReconstructing missing modality based on available modalities is widely used to address inevitable modality-missing for Multimodal Emotion Recognition (MER). However, due to explicit distribution gap across heterogeneous modalities, they fail to guarantee the consistency between the reconstructed data and the ground truth. To mitigate this problem, we propose a novel method to restore the missing modality using its weighted prototypes rather than other modalities. Specifically, prototypes of different classes of missing modality are used to encapsulate its representative knowledge. Then sample-to-prototype affinity measuring class similarity is used as weights to combine these prototypes for reconstruction, thereby effectively restoring the distribution-consistent modality. Furthermore, to improve the efficacy of prototype-based completion under seriously missing, we devise an adaptive knowledge distillation from the strong modality to the weaker ones. This reinforces the representation ability of weak modality features. Extensive experiments on CMU-MOSI and IEMOCAP datasets demonstrate the superiority of our method. Zhangfeng Hu, Wenming Zheng, Yuan Zong, Mengting Wei, Xingxun Jiang, Mengxin Shi |
ICME | 2 |
| 2024 | Missing Customized Distillation Network for Incomplete Multimodal Sentiment Analysis
Zhangfeng Hu, Wenming Zheng, Mengting Wei, Mengxin Shi, Yuan Zong |
ICPR (8) | 2 |
| 2024 | Hierarchical Distribution Adaptation for Unsupervised Cross-corpus Speech Emotion Recognition
Cheng Lu 0005, Yuan Zong, Yan Zhao 0037, Hailun Lian, Tianhua Qi, Björn W. Schuller, Wenming Zheng |
INTERSPEECH | 7 |
| 2024 | Towards Realistic Emotional Voice Conversion using Controllable Emotional Intensity
Tianhua Qi, Shiyan Wang, Cheng Lu 0005, Yan Zhao 0037, Yuan Zong, Wenming Zheng |
INTERSPEECH | 6 |
| 2024 | Boosting Cross-Corpus Speech Emotion Recognition using CycleGAN with Contrastive Learning
Jincen Wang, Yan Zhao 0037, Cheng Lu 0005, Chuangao Tang, Sunan Li, Yuan Zong, Wenming Zheng |
INTERSPEECH | 7 |
| 2024 | Confidence-aware Hypothesis Transfer Networks for Source-Free Cross-Corpus Speech Emotion Recognition
Jincen Wang, Yan Zhao 0037, Cheng Lu 0005, Hailun Lian, Hongli Chang, Yuan Zong, Wenming Zheng |
INTERSPEECH | 7 |
| 2024 | Comprehensive multi-view self-representations for clusteringabstractSubspace learning-based methods have shown excellent performance for multi-view clustering, yet have the following problems: (1) most existing methods obtain the subspace representation from the original space, which might contain noises and cannot guarantee a clean enough subspace representation; (2) existing methods mainly focus on the consistency of the subspace representation, while the unique information of each view is not sufficiently exploited. To solve these two problems, we propose a novel multi-view subspace clustering method called comprehensive multi-view self-representations (CMSR). Specifically, we learn the original coefficient matrix of each view through the self-representation, which can reduce the noise of the original space to some extent. Then, we learn the subspace representation of the original coefficient matrix and decompose it into a consistent coefficient matrix and multiple diverse coefficient matrices, which can exploit the consistent and complementary information of multi-view data. Further, we impose the Schatten p -norm constraint on the consistent coefficient matrix to capture robust consistent information. Finally, the comprehensive results on eight real datasets demonstrate the versatility and effectiveness of the proposed method. Yuanbo Cheng, Peng Song 0002, Jinshuai Mu, Yanwei Yu, Wenming Zheng |
Expert Syst. Appl. | 5 |
| 2024 | Clean affinity matrix induced hyper-Laplacian regularization for unsupervised multi-view feature selection
Peng Song 0002, Shixuan Zhou, Jinshuai Mu, Meng Duan, Yanwei Yu, Wenming Zheng |
Inf. Sci. | 6 |
| 2024 | Wasserstein Discriminant Dictionary Learning for Graph RepresentationabstractMining discriminative graph topological information plays an important role in promoting graph representation ability. However, it suffers from two main issues: (1) the difficulty/complexity of computing global inter-class/intra-class scatters, commonly related to mean and covariance of graph samples, for discriminant learning; (2) the huge complexity and variety of graph topological structure that is rather challenging to robustly characterize. In this paper, we propose the Wasserstein Discriminant Dictionary Learning (WDDL) framework to achieve discriminant learning on graphs with robust graph topology modeling, and hence facilitate graph-based pattern analysis tasks. Considering the difficulty of calculating global inter-class/intra-class scatters, a reference set of graphs (aka graph dictionary) is first constructed by generating representative graph samples (aka graph keys) with expressive topological structure. Then, a Wasserstein Graph Representation (WGR) process is proposed to project input graphs into a succinct dictionary space through the graph dictionary lookup. To further achieve discriminant graph learning, a Wasserstein discriminant loss (WD-loss) is defined on the graph dictionary, in which the graph keys are optimizable, to make the intra-class keys more compact and inter-class keys more dispersed. Hence, the calculation of global Wasserstein metric (W-metric) centers can be bypassed. For sophisticated topology mining in the WGR process, a joint-Wasserstein graph embedding module is constructed to model both between-node and between-edge relationships across inputs and graph keys by encapsulating both the Wasserstein metric (between cross-graph nodes) and proposed novel Kron-Gromov-Wasserstein (KGW) metric (between cross-graph adjacencies). Specifically, the KGW-metric comprehensively characterizes the cross-graph connection patterns with the Kronecker operation, then adaptively captures those salient patterns through connection pooling. To evaluate the proposed framework, we study two graph-based pattern analysis problems, i.e. graph classification and cross-modal retrieval, with the graph dictionary flexibly adjusted to cater to these two tasks. Extensive experiments are conducted to comprehensively compare with existing advanced methods, as well as dissect the critical component of our proposed architecture. The experimental results validate the effectiveness of the WDDL framework. Tong Zhang 0021, Guangbu Liu, Zhen Cui 0001, Wei Liu 0005, Wenming Zheng, Jian Yang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | CFEW: A Large-Scale Database for Understanding Child Facial Expression in Real WorldabstractCurrently, much progress has been achieved on adult facial expressions recognition. Few attentions have been paid to child facial expression analysis. A lack of publicly available large-scale child facial expression databases hinders the development of automatic coding for child facial expression behaviors. In this work, we constructed a new face database for understandingChildFacialExpression in realWorld (CFEW). The database contains three novelties: (1) the largest publicly available child facial expression database (11,000+ images); (2) covering full developmental range of 0–18-year-old child subjects; (3) rich annotations for facial expression labels, including discrete expression categories aka happy, neutral, disgust, angry, sad, cry, fear, surprise, sleepy and others, intensity of arousal and valence, and several types of facial action units (AUs). In addition, the images in this database cover several challenging conditions in real world, including frontal and non-frontal head poses, facial occlusions, various illuminations and low image resolution. Three dominant deep convolutional neural networks (i.e., VGG11bn, ResNet18 and DenseNet121) were used to conduct extensive baseline experiments for discrete facial expression classification, arousal and valence estimation and facial action units detection within database, and cross-database seven facial expressions recognition. Chuangao Tang, Sunan Li, Wenming Zheng, Yuan Zong, Su Zhang 0004, Cheng Lu 0005, Yan Zhao 0037 |
IEEE Trans. Affect. Comput. | 3 |
| 2024 | Common Latent Embedding Space for Cross-Domain Facial Expression RecognitionabstractIn practical facial expression recognition (FER), the training data and test data are often obtained from different domains. It is obvious that the domain disparity could significantly degrade the recognition performance. To tackle this challenging cross-domain FER problem, we put forward a novel method termed common latent embedding space (CLES). To be specific, first, we obtain a common embedding space for cross-domain samples by matrix factorization (MF). Then, the dual-graph Laplacian is applied to this common embedding space to narrow the gap across distinct domains and, meanwhile, explores the inherent geometric information. Furthermore, to characterize the global relationship of the cross-domain samples, the self-representation strategy is used to guide the learning of the common embedding space. Finally, comprehensive experiments on four benchmark databases indicate that the proposed method can achieve better performance in comparison with the state-of-the-art methods on cross-domain FER tasks. Peng Song 0002, Shaokai Li, Wenming Zheng |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2024 | Graph-Diffusion-Based Domain-Invariant Representation Learning for Cross-Domain Facial Expression RecognitionabstractThe precondition that most of the existing facial expression recognition (FER) algorithms have succeeded lies in that the training (source) and test (target) samples are independent of each other and identically distributed. However, it is too strict to satisfy this precondition in the real-world. To this end, we propose a novel graph-diffusion-based domain-invariant representation learning (GDRL) model for the cross-domain FER scenario where there exist distribution shifts between various domains. Specifically, a low-dimensional space mapping strategy is first adopted to diminish the domain mismatch. Then, by skillfully combining the local graph embedding and affinity graph diffusion, the local geometric structures can be effectively modeled and the deeper higher-order relationships of samples from various domains can be captured. In addition, in order to better guide the transfer process and learn a more discriminative and invariant representation, we take into account the label consistency. Experimental results on four laboratory-controlled databases and two in-the-wild databases demonstrate that our proposed model can yield better recognition performance compared with state-of-the-art domain adaptation methods. Peng Song 0002, Wenming Zheng |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | Adaptive Dual-Space Network With Multigraph Fusion for EEG-Based Emotion RecognitionabstractMost of the work on electroencephalogram (EEG)-based emotion recognition aims to extract the distinguishing features from high-dimensional EEG signals, ignoring the complementarity of information between EEG latent space and graph space. Furthermore, the influence of brain connectivity on emotions encompasses both physical structure and functional connectivity, which may have varying degrees of importance for different individuals. To address these issues, this article introduces an adaptive dual-space network (ADS-Net) with multigraph fusion aimed at capturing more comprehensive information by integrating dual-space representations. Specifically, ADS-Net models the spatial correlation of EEG channels in graph topological space, while exploring long-range dependencies and frequency relationships from EEG data in latent space. Subsequently, these representations are adaptively combined through an innovative gated fusion approach to extract complementary corepresentations. Moreover, drawing on the principles of brain connectivity theory, the proposed method constructs a multigraph to indicate the associativity of EEG channels. To further capture individual differences, an adaptive multigraph fusion mechanism is developed for the dynamic integration of physical and functional connectivity graphs. When compared to state-of-the-art methods, the superior experimental results underscore the effectiveness and broad applicability of the proposed method. Mengqing Ye, C. L. Philip Chen, Wenming Zheng, Tong Zhang 0015 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | Layer-Adapted Implicit Distribution Alignment Networks for Cross-Corpus Speech Emotion RecognitionabstractIn this article, we propose a new unsupervised domain adaptation (DA) method called layer-adapted implicit distribution alignment networks (LIDANs) to address the challenge of cross-corpus speech emotion recognition (SER). LIDAN extends our previous ICASSP work, deep implicit distribution alignment networks (DIDANs), whose key contribution lies in the introduction of a novel regularization term called implicit distribution alignment (IDA). This term allows DIDAN trained on source (training) speech samples to remain applicable to predicting emotion labels for target (testing) speech samples, regardless of corpus variance in cross-corpus SER. To further enhance this method, we extend IDA to layer-adapted IDA (LIDA), resulting in LIDAN. This layer-adapted extension consists of three modified IDA terms that consider emotion labels at different levels of granularity. These terms are strategically arranged within different fully connected layers in LIDAN, aligning with the increasing emotion-discriminative abilities with respect to the layer depth. This arrangement enables LIDAN to more effectively learn emotion-discriminative and corpus-invariant features for SER across various corpora compared to DIDAN. It is also worthy to mention that unlike most existing methods that rely on estimating statistical moments to describe preassumed explicit distributions, both IDA and LIDA take a different approach. They utilize an idea of target sample reconstruction to directly bridge the feature distribution gap without making assumptions about their distribution type. As a result, DIDAN and LIDAN can be viewed as implicit cross-corpus SER methods. To evaluate LIDAN, we conducted extensive cross-corpus SER experiments on EmoDB, eNTERFACE, and CASIA corpora. The experimental results demonstrate that LIDAN surpasses recent state-of-theart explicit unsupervised DA methods in tackling cross-corpus SER tasks. Yan Zhao 0037, Yuan Zong, Jincen Wang, Hailun Lian, Cheng Lu 0005, Li Zhao 0003, Wenming Zheng |
IEEE Trans. Comput. Soc. Syst. | 7 |
| 2024 | Adaptive Multi-Scale Iterative Optimized Video Object Segmentation Based on Correlation EnhancementabstractSemi-supervised video object segmentation (VOS) is a highly challenging task, which relies on the initial frame’s mask as a segmentation reference in a video sequence to classify each pixel in subsequent frames. However, the guidance provided by the first frame is limited due to the diverse types of segmentation targets and uncertain appearance changes. Consequently, it is crucial to retain useful information during the segmentation process and employ this information for model iteration optimization, enabling the model to better adapt to rapidly changing segmentation objectives. In this work, we propose a multi-scale adaptive model optimization strategy, which incorporates a contextual relevance enhancement module to enforce object correlation by emphasizing feature similarity across adjacent frames. Additionally, we introduce a keyframe discrimination module to deal with the segmentation challenges in scenarios involving significant target changes. Moreover, we also introduce a multi-scale memory screening module to automatically screen and select global-local optimization features for ensuring the model’s generalization performance. Extensive experiments show that the proposed method achieves state-of-the-art performance on DAVIS and large-scale Youtube-VOS 2018/2019 datasets without relying on synthetic training data or first-frame fine-tuning. Yuan Zong, Wenming Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | MaskFusionNet: A Dual-Stream Fusion Model With Masked Pre-Training Mechanism for rPPG MeasurementabstractRemote photoplethysmography (rPPG) has considerable significance in areas such as disease diagnosis and emotion analysis. Recent rPPG models have demonstrated excellent performance due to their powerful heart rate information extraction capabilities. However, these models often focus on limited regions of interest (ROI) on facial image, which makes them sensitive to interference. If the ROI is affected by muscle movement, lighting variation and noise, the model’s performance would degrade significantly. To address this limitation, we propose a two-stage model called MaskFusionNet. The model includes two stages: 1) During the pre-training stage, the mask-reconstruction mechanism drives MaskFusionNet to learn rPPG information from various facial regions by applying a tube masking strategy. This enhances the model’s ability to resist interference. Based on the periodicity and continuity of the heart rate signal, we also design a novel spatio-temporal reconstruction loss function that focuses on the data’s spatial features and temporal continuity. 2) In the fine-tuning stage, we propose the Multi-Scale Fusion Block (MFB) to combine multi-scale features from the dual-stream network. It allows the model to detect subtle heart rate variations in adjacent frames while minimizing the impact of interference by extracting features within longer segments. The transformer-based MaskFusionNet can extract multi-scale fused heart rate features from a wide range of skin regions while preserving the modeling capability of long-range sequence information. To validate its effectiveness, we extensively evaluate our model on three benchmark datasets (VIPL-HR, COHFACE, and PURE), demonstrating its superior performance in both intra-dataset and cross-dataset testing scenarios. Yizhu Zhang, Jingang Shi, Yuan Zong, Wenming Zheng, Guoying Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Convolutional Transformer-Based Cross Subject Model for SSVEP-Based BCI ClassificationabstractSteady-state visual evoked potential (SSVEP) is a commonly used brain-computer interface (BCI) paradigm. The performance of cross-subject SSVEP classification has a strong impact on SSVEP-BCI. This study designed a cross subject generalization SSVEP classification model based on an improved transformer structure that uses domain generalization (DG). The global receptive field of multi-head self-attention is used to learn the global generalized SSVEP temporal information across subjects. This is combined with a parallel local convolution module, designed to avoid oversmoothing the oscillation characteristics of temporal SSVEP data and better fit the feature. Moreover, to improve the cross-subject calibration-free SSVEP classification performance, an DG method named StableNet is combined with the proposed convolutional transformer structure to form the DG-Conformer method, which can eliminate spurious correlations between SSVEP discriminative information and background noise to improve cross-subject generalization. Experiments on two public datasets, Benchmark and BETA, demonstrated the outstanding performance of the proposed DG-Conformer compared with other calibration-free methods, FBCCA, tt-CCA, Compact-CNN, FB-tCNN, and SSVEPNet. Additionally, DG-Conformer outperforms the classic calibration-required algorithms eCCA, eTRCA and eSSCOR when calibration is used. An incomplete partial stimulus calibration scheme was also explored on the Benchmark dataset, and it was demonstrated to be a potential solution for further high-performance personalized SSVEP-BCI with quick calibration. Yuankui Yang, Yuan Zong, Yue Leng, Wenming Zheng, Sheng Ge |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | Novel Sinusoidal Signal Assisted Multivariate Variational Mode Decomposition Combined With Task-Related Component Analysis for Enhancing SSVEP-Based BCI PerformanceabstractBrain-computer interfaces (BCIs) based on steady-state visually evoked potential (SSVEP) have a broad application prospect owing to their multiple command output and high performance. Each harmonic component of SSVEP individually contains unique features, which can be utilized to enhance the recognition performance of SSVEP-based BCIs. However, the existing subband analysis methods for SSVEP, including those based on filter banks and existing mode decomposition methods, have limitations in extracting and utilizing independent harmonic components. This study proposes a sinusoidal signal assisted multivariate variational mode decomposition (SA-MVMD) algorithm that allows the constraint of the center frequencies and narrowband filtering structures of the intrinsic mode functions (IMFs) based on the prior frequency knowledge of the signal. It preserves the target information of the signal during decomposition while avoiding mode mixing and incorrect decomposition, thereby enabling the effective extraction of each independent harmonic component of SSVEP. Building on this, a SA-MVMD based task-related component analysis (SA-MVMD-TRCA) method is further proposed to fully utilize the features within the overall SSVEP as well as its independent harmonics, thereby enhancing the recognition performance. Testing on the public SSVEP Benchmark dataset demonstrates that the proposed method significantly outperforms the filter bank-based control methods. This study confirms the effectiveness of SA-MVMD and the potential of this approach, which analyzes and utilizes each independent harmonic of SSVEP, providing new strategies and perspectives for performance enhancement in SSVEP-based BCIs. Jinpeng Lyu, Yuankui Yang, Yuan Zong, Yue Leng, Wenming Zheng, Sheng Ge |
IEEE J. Biomed. Health Informatics | 5 |
| 2023 | CMNet: Contrastive Magnification Network for Micro-Expression RecognitionabstractMicro-Expression Recognition (MER) is challenging because the Micro-Expressions' (ME) motion is too weak to distinguish. This hurdle can be tackled by enhancing intensity for a more accurate acquisition of movements. However, existing magnification strategies tend to use the features of facial images that include not only intensity clues as intensity features, leading to the intensity representation deficient of credibility. In addition, the intensity variation over time, which is crucial for encoding movements, is also neglected. To this end, we provide a reliable scheme to extract intensity clues while considering their variation on the time scale. First, we devise an Intensity Distillation (ID) loss to acquire the intensity clues by contrasting the difference between frames, given that the difference in the same video lies only in the intensity. Then, the intensity clues are calibrated to follow the trend of the original video. Specifically, due to the lack of truth intensity annotation of the original video, we build the intensity tendency by setting each intensity vacancy an uncertain value, which guides the extracted intensity clues to converge towards this trend rather some fixed values. A Wilcoxon rank sum test (Wrst) method is enforced to implement the calibration. Experimental results on three public ME databases i.e. CASME II, SAMM, and SMIC-HS validate the superiority against state-of-the-art methods. Mengting Wei, Xingxun Jiang, Wenming Zheng, Yuan Zong, Cheng Lu 0005, Jiateng Liu |
AAAI | 3 |
| 2023 | A Generalized Subspace Distribution Adaptation Framework for Cross-Corpus Speech Emotion RecognitionabstractIn this paper, we propose a novel transfer learning framework, named generalized subspace distribution adaptation (GSDA), to tackle the challenging cross-corpus speech emotion recognition problem. First, we learn a common low-dimensional feature subspace by utilizing a generalized subspace learning method. Second, we develop a novel distance metric to reduce the divergence between the source and target corpora, which can efficiently explore the similarity and dissimilarity information in the process of knowledge transfer. Third, to demonstrate the effectiveness of our framework, we apply GSDA to the traditional subspace learning algorithms. Finally, we conduct extensive experiments by using the low-level features and deep features on three popular emotional databases, i.e., Berlin, IEMOCAP, and CVE. The results demonstrate that the proposed framework can achieve better performance than several state-of-the-art transfer learning approaches. Shaokai Li, Peng Song 0002, Yun Jin, Wenming Zheng |
ICASSP | 5 |
| 2023 | Deep Implicit Distribution Alignment Networks for cross-Corpus Speech Emotion RecognitionabstractIn this paper, we propose a novel deep transfer learning method called deep implicit distribution alignment networks (DIDAN) to deal with cross-corpus speech emotion recognition (SER) problem, in which the labeled training (source) and unlabeled testing (target) speech signals come from different corpora. Specifically, DIDAN first adopts a simple deep regression network consisting of a set of convolutional and fully connected layers to directly regress the source speech spectrums into the emotional labels such that the proposed DIDAN can own the emotion discriminative ability. Then, such ability is transferred to be also applicable to the target speech samples regardless of corpus variance by resorting to a well-designed regularization term called implicit distribution alignment (IDA). Unlike widely-used maximum mean discrepancy (MMD) and its variants, the proposed IDA absorbs the idea of sample reconstruction to implicitly align the distribution gap, which enables DIDAN to learn both emotion discriminative and corpus invariant features from speech spectrums. To evaluate the proposed DIDAN, extensive cross-corpus SER experiments on widely-used speech emotion corpora are carried out. Experimental results show that the proposed DIDAN can outperform lots of recent state-of-the-art methods in coping with the cross-corpus SER tasks. Yan Zhao 0037, Jincen Wang, Yuan Zong, Wenming Zheng, Hailun Lian, Li Zhao 0003 |
ICASSP | 4 |
| 2023 | Geometric Magnification-based Attention Graph Convolutional Network for Skeleton-based Micro-Gesture RecognitionabstractMicro-Gesture (MG) recognition is an emerging and challenging task due to the short duration and small amplitude of joints. MGs indicate subtle movements of the body in response to stress, which are more difficult to recognize than regular gestures. To solve the above problems, for the modeling of micro-gesture skeleton data, we propose a Geometric Magnification-Based Attention Graph Convolutional Network (MA-GCN) to magnify and select features. The network mainly consists of two modules: the geometric magnification module (GM module) controls the magnification of different joints, and the spatial temporal attention graph convolution module (STA module) selects valid information by weighting different joints and frames to focus on subtle movements. Extensive experiments on two MG datasets prove that our method achieves remarkable performance. Haolin Jiang, Wenming Zheng, Yuan Zong, Xingxun Jiang, Yunlong Xue |
ICIP | 2 |
| 2023 | Learning Local to Global Feature Aggregation for Speech Emotion Recognition
Cheng Lu 0005, Hailun Lian, Wenming Zheng, Yuan Zong, Yan Zhao 0037, Sunan Li |
INTERSPEECH | 3 |
| 2023 | Unsupervised Transfer Components Learning for Cross-Domain Speech Emotion Recognition
Shenjie Jiang, Peng Song 0002, Shaokai Li, Keke Zhao, Wenming Zheng |
INTERSPEECH | 5 |
| 2023 | Joint Instance Reconstruction and Feature Subspace Alignment for Cross-Domain Speech Emotion Recognition
Keke Zhao, Peng Song 0002, Shaokai Li, Wenming Zheng |
INTERSPEECH | 4 |
| 2023 | Multimodal Emotion Recognition in Noisy Environment Based on Progressive Label RevisionabstractThe multimodal emotion recognition has attracted more attention in recent decades. Though remarkable progress has been achieved with the rapid development of deep learning, existing methods are still hard to tackle noise problems that occurred commonly in emotion recognition's practical application. To improve the robustness of the multimodal emotion recognition algorithm, we propose an MLP-based label revision algorithm. The framework consists of three complementary feature extraction networks that were verified in MER2023. After that, an MLP-based attention network with specially designed loss functions was used to fuse features from different modalities. Finally, the scheme that used the output probability of each emotion to revise the sample's output category was employed to revise the test set's label obtained by classifier. The samples that are most likely to be affected by noise and misclassified have a chance to get correct classification. The best experimental result shows that the F1-score of our algorithm on the test dataset of the MER 2023 Noise subchallenge is 86.35 and combined metric is 0.6694, which ranks 2nd at the MER 2023 NOISE subchallenge. Sunan Li, Hailun Lian, Cheng Lu 0005, Yan Zhao 0037, Chuangao Tang, Yuan Zong, Wenming Zheng |
ACM Multimedia | 7 |
| 2023 | Tensor-based consensus learning for incomplete multi-view clustering
Jinshuai Mu, Peng Song 0002, Yanwei Yu, Wenming Zheng |
Expert Syst. Appl. | 4 |
| 2023 | Progressive graph convolution network for EEG emotion recognition
Yijin Zhou, Fu Li 0002, Yang Li 0019, Youshuo Ji, Guangming Shi, Wenming Zheng, Lijian Zhang, Yuanfang Chen, Rui Cheng 0010 |
Neurocomputing | 6 |
| 2023 | Structural regularization based discriminative multi-view unsupervised feature selection
Shixuan Zhou, Peng Song 0002, Yanwei Yu, Wenming Zheng |
Knowl. Based Syst. | 4 |
| 2023 | Window-Adjusted Common Spatial Pattern for Detecting Error-Related Potentials in P300 BCIabstractAbstract Under certain task conditions, error-related potential (ErrP) will be elicited, meaning that the subject is perceiving an error, responding to an external error, or engaging in a cognitive process of reinforcement learning. The detection of ErrP on a single trial basis has been studied and applied to improve all kinds of brain–computer interfaces (BCIs). However, the performance of this kind of detection is not currently good enough. In the paper, we proposed a novel method, called window-adjusted common spatial pattern (WACSP), for detecting ErrP in P300 BCI. In this method, the coefficient of determination was introduced to measure the difference of Electroencephalogram (EEG) signals on a channel at a moment and to guide the search of time windows in which EEG differences are significant, and common spatial pattern (CSP) was further used to capture the stable spatial patterns of EEG differences between correct and incorrect responses in each time window. WACSP and the commonly used methods were tested on the data sets that were built using the EEG signals acquired during the P300 BCI experiments with different feedback. The comparisons of accuracy, area under receiver operating characteristics curve (AUC) and F-measure show that WACSP significantly outperforms the commonly used methods. The proposed method can improve ErrP detection based on a single trial. Minghong Li, Wenming Zheng, Huiru Zheng |
Neural Process. Lett. | 3 |
| 2023 | Learning Transferable Sparse Representations for Cross-Corpus Facial Expression RecognitionabstractAn assumption widely used in traditional facial expression recognition algorithms is that the training and testing are conducted on the same dataset. However, this assumption does not hold in practice, in which the training data and testing data are often from different datasets. In this scenario, directly deploying these algorithms would lead to severe information loss and performance degradation due to the domain shift. To address this challenging problem, in this article, we propose a novel transferable sparse subspace representation method (TSSR) for cross-corpus facial expression recognition. Specifically, in order to reduce the cross-corpus mismatch, inspired by sparse subspace clustering, we advocate reconstructing the source and target samples using the source data points based on$\ell _1-$norm sparse representation. Each data point in source and target corpora can be ideally represented as a combination of a few other source points from its own subspace. Moreover, we take into account the local geometrical information within the cross-corpus data by adopting a graph Laplacian regularizer, which can efficiently preserve the local manifold structure and better transfer knowledge between two corpora. Finally, extensive experiments on several facial expression datasets are conducted to evaluate the recognition performance of TSSR. Experimental results demonstrate the superiority of the proposed method over some state-of-the-art methods. Peng Song 0002, Wenming Zheng |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | GMSS: Graph-Based Multi-Task Self-Supervised Learning for EEG Emotion RecognitionabstractPrevious electroencephalogram (EEG) emotion recognition relies on single-task learning, which may lead to overfitting and learned emotion features lacking generalization. In this paper, a graph-based multi-task self-supervised learning model (GMSS) for EEG emotion recognition is proposed. GMSS has the ability to learn more general representations by integrating multiple self-supervised tasks, including spatial and frequency jigsaw puzzle tasks, and contrastive learning tasks. By learning from multiple tasks simultaneously, GMSS can find a representation that captures all of the tasks thereby decreasing the chance of overfitting on the original task, i.e., emotion recognition task. In particular, the spatial jigsaw puzzle task aims to capture the intrinsic spatial relationships of different brain regions. Considering the importance of frequency information in EEG emotional signals, the goal of the frequency jigsaw puzzle task is to explore the crucial frequency bands for EEG emotion recognition. To further regularize the learned features and encourage the network to learn inherent representations, contrastive learning task is adopted in this work by mapping the transformed data into a common feature space. The performance of the proposed GMSS is compared with several popular unsupervised and supervised methods. Experiments on SEED, SEED-IV, and MPED datasets show that the proposed model has remarkable advantages in learning more discriminative and general features for EEG emotional signals. Yang Li 0019, Fu Li 0002, Boxun Fu, Youshuo Ji, Yijin Zhou, Guangming Shi, Wenming Zheng |
IEEE Trans. Affect. Comput. | 10 |
| 2023 | Variational Instance-Adaptive Graph for EEG Emotion RecognitionabstractThe individual differences and the dynamic uncertain relationships among different electroencephalogram (EEG) regions are essential factors that limit EEG emotion recognition. To address these issues, in this article, we propose a variational instance-adaptive graph method (V-IAG) that simultaneously captures the individual dependencies among different EEG electrodes and estimates the underlying uncertain information. Specifically, we employ two branches, i.e., instance-adaptive branch and variational branch, to construct the graph. Inspired by the attention mechanism, the instance-adaptive branch generates the graph based on the input so as to characterize the individual dependencies among EEG channels. The variational branch generates the probabilistic graph, which quantifies the uncertainties. We combine these two types of graphs to extract more discriminative features. To present more precise graph representation, we propose a new operation named the multi-level and multi-graph convolution operation, which aggregates the features of EEG channels from different frequencies with different graphs. Furthermore, we design the graph coarsening and employ the sparse constraint to obtain more robust features. We conduct extensive experiments on three widely-used EEG emotion recognition databases, i.e., SJTU emotion EEG dataset (SEED), multi-modal physiological emotion recognition dataset (MPED) and DREAMER. The results demonstrate that the proposed model achieves the-state-of-the-art performance. Tengfei Song, Suyuan Liu, Wenming Zheng, Yuan Zong, Zhen Cui 0001, Yang Li 0019 |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | Unsupervised Cross-View Facial Expression Image Generation and RecognitionabstractWe propose an unsupervised cross-view facial expression adaptation network (UCFEAN) to simultaneously generate and recognize cross-view facial expressions in images in an unsupervised manner. The main idea of UCFEAN is to convert the unsupervised domain adaptation between two image spaces with different appearance into semi-supervised learning (SSL) in feature spaces with the same semantic content. The cyclic image generation of cross-view facial expressions based on the generative adversarial network (GAN) is carried out to project unlabelled target images and labelled source images to the corresponding feature spaces with the same semantic content. This helps realize the unsupervised feature learning of the target image. Labels of facial expressions represented in the projected target features can then be learned using the projected source features, because the distributions of the projected features in the two domains are close enough for knowledge transfer by using SSL. Three techniques are developed to train UCFEAN in an effective and stable manner. Extensive experiments are conducted to evaluate the UCFEAN on two multi-view facial expression image databases including RaFD and Multi-PIE. The results show that the proposed method can generate realistic target images of the facial expression and recognize cross-view facial expressions with high precision. Ning Sun 0005, Qingyi Lu, Wenming Zheng, Jixin Liu 0001, Guang Han 0002 |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | FENP: A Database of Neonatal Facial Expression for Pain AnalysisabstractIn this article, we introduce a new neonatal facial expression database for pain analysis. This database, called facial expression of neonatal pain (FENP), contains 11,000 neonatal facial expression images associated with 106 Chinese neonates from two children's hospitals, i.e., the Children's Hospital Affiliated to Nanjing Medical University and Second Affiliated Hospital Affiliated to Nanjing Medical University in China. The facial expression images cover four categories of facial expressions, i.e., severe pain expression, mild pain expression, crying expression and calmness expression, where each category contains 2750 neonatal facial expression images. Based on this database, we also investigate the pain facial expression recognition problem using several state-of-the-art facial expression features and expression recognition methods, such as Gabor+SVM, LBP+SVM, HOG+SVM, LBP+HOG+SVM, and several Convolutional Neural Network (CNN) methods (including AlexNet, VGGNet, GoogLeNet, ResNet and DenseNet). The experimental results indicate that the proposed neonatal pain facial expression database is very suitable for the study of both neonatal pain and facial expression recognition. Moreover, the FENP database is publicly available after signing a license agreement (the users can contact Jingjie Yan ([email protected]), Guanming Lu ([email protected])) or Xiaonan Li ([email protected]). Jingjie Yan, Guanming Lu, Wenming Zheng, Chengwei Huang, Zhen Cui 0001, Yuan Zong, Mengying Chen, Jindu Zhu, Haibo Li 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2023 | Joint Local-Global Discriminative Subspace Transfer Learning for Facial Expression RecognitionabstractTraditional facial expression recognition (FER) has achieved satisfactory results to some extent, and most of the current methods are trained and evaluated on a single database. However, in real applications, the training and testing images are often collected in different scenarios, which will lead to performance degeneration. To tackle this problem, in this paper, we propose a novel transfer learning approach, named joint local-global discriminative subspace transfer learning (LGDSTL), for cross-database FER. In LGDSTL, first, we develop a joint local-global graph as the distance metric, in which we not only consider the local discriminative geometric structure for each database, but also consider a global graph to transfer knowledge. In this way, the discrepancy between the two databases will be significantly reduced. Then, we present a pairwise regression function to guide the discriminative subspace transfer learning. Additionally, a data reconstruction constraint is introduced to preserve the main discriminative information. Finally, comparative studies on six popular benchmarks demonstrate the effectiveness of the proposed approach. Wenjing Zhang 0003, Peng Song 0002, Wenming Zheng |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | SparseDGCNN: Recognizing Emotion From Multichannel EEG SignalsabstractEmotion recognition from EEG signals has attracted much attention in affective computing. Recently, a novel dynamic graph convolutional neural network (DGCNN) model was proposed, which simultaneously optimized the network parameters and a weighted graph$G$characterizing the strength of functional relation between each pair of two electrodes in the EEG recording equipment. In this article, we propose a sparse DGCNN model which modifies DGCNN by imposing a sparseness constraint on$G$and improves the emotion recognition performance. Our work is based on an important observation: the tomography study reveals that different brain regions sampled by EEG electrodes may be related to different functions of the brain and then the functional relations among electrodes are possibly highly localized and sparse. However, introducing sparseness constraint into the graph$G$makes the loss function of sparse DGCNN non-differentiable at some singular points. To ensure that the training process of sparse DGCNN converges, we apply the forward-backward splitting method. To evaluate the performance of sparse DGCNN, we compare it with four representative recognition methods (SVM, DBN, GELM and DGCNN). In addition to comparing different recognition methods, our experiments also compare different features and spectral bands, including EEG features in time-frequency domain (DE, PSD, DASM, RASM, ASM and DCAU on different bands) extracted from four representative EEG datasets (SEED, DEAP, DREAMER, and CMEED). The results show that (1) sparse DGCNN has consistently better accuracy than representative methods and has a good scalability, and (2) DE, PSD, and ASM features on$\gamma$band convey most discriminative emotional information, and fusion of separate features and frequency bands can improve recognition performance. Minjing Yu, Yong-Jin Liu 0001, Guozhen Zhao, Dan Zhang 0014, Wenming Zheng |
IEEE Trans. Affect. Comput. | 6 |
| 2023 | Multi-Source Discriminant Subspace Alignment for Cross-Domain Speech Emotion RecognitionabstractCross-domain speech emotion recognition (SER) is an effective strategy to improve the generalization ability of emotion classification models, which is an important research direction in speech signal processing. However, since the speech signals are non-stationary, it is difficult to train a robust classifier from single-source emotional corpus. To solve this shortcoming, we propose a novel method named multi-source discriminant subspace alignment (MDSA) for cross-domain SER. In MDSA, we first conduct linear discriminant analysis (LDA) in the multi-source domain. Then, the instances in the multi-source discriminant subspace are used to linearly reconstruct the instances in the target subspace. At the same time, the reconstruction contribution of each source discriminant subspace is determined by adaptive weights. Furthermore, the multi-source discriminant subspace is aligned by reducing the loss between projections, which can make our model more robust. In this way, MDSA considers both the alignment of cross-domain data distribution and the structural information of cross-domain instances. Finally, extensive experiments are conducted on five standard emotional corpora, i.e., Berlin, IEMOCAP, CVE, EMOVO, and TESS, and the results demonstrate the proposed MDSA is superior to several state-of-the-art transfer learning algorithms in terms of performance. The codes are available athttps://github.com/shaokai1209/MDSA. Shaokai Li, Peng Song 0002, Wenming Zheng |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Speech Emotion Recognition via an Attentive Time-Frequency Neural NetworkabstractSpectrogram is commonly used as the input feature of deep neural networks to learn the high(er)-level time–frequency pattern of speech signal for speech emotion recognition (SER). Generally, different emotions correspond to specific energy activations both within frequency bands and time frames on spectrogram, which indicates the frequency and time domains are both essential to represent the emotion for SER. However, recent spectrogram-based works mainly focus on modeling the long-term dependency in time domain, which makes these methods suffer from the following issues: 1) neglecting to model the emotion-related correlations within frequency domain during the time–frequency joint learning and 2) ignoring to capture the specific frequency bands associated with emotions. To cope with the issues, we propose an attentive time–frequency neural network (ATFNN) for SER, including a time–frequency neural network (TFNN) and time–frequency attention. Specifically, aiming at the first issue, we design a TFNN with a frequency-domain encoder (F-Encoder) based on the Transformer encoder and a time-domain encoder (T-Encoder) based on the bidirectional long short-term memory (Bi-LSTM). The F-Encoder and T-Encoder model the correlations within frequency bands and time frames, respectively, and they are embedded into a time–frequency joint learning strategy to obtain the time–frequency patterns of speech emotions. Moreover, to handle the second issue, we adopt the time–frequency attention with a frequency-attention network (F-Attention) and a time-attention network (T-Attention) to focus on the emotion-related long-range dependencies between frequency bands and across time frames, which can enhance the emotional discrimination of speech features. Extensive experimental results on three public emotional databases, i.e., IEMOCAP, ABC, and CASIA, show that our proposed ATFNN outperforms the state-of-the-art methods. Cheng Lu 0005, Wenming Zheng, Hailun Lian, Yuan Zong, Chuangao Tang, Sunan Li, Yan Zhao 0037 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2022 | A Novel Micro-Expression Recognition Approach Using Attention-Based Magnification-Adaptive NetworksabstractMicro-Expression recognition (MER) is a challenging task due to the short duration and low intensity of Micro-Expressions. A popular method to tackle this is magnifying MEs so as to enlarge the expression intensity to make recognition easier. However, the single fixed magnification strategy, widely used in existing works of MER, is not appropriate for different subjects, because each subject has specific expression intensity corresponding to different MEs. To cope with this issue, we propose a novel Attention-based Magnification-Adaptive Network (AMAN) to learn adaptive magnification levels for the ME representation. The network consists of two modules: magnification attention (MA module) to adaptively focus on appropriate magnification levels of different MEs, and frame attention (FA module) to focus on discriminative aggregated frames in a ME video. Extensive experiments on three widely used databases manifest that our method yields state-of-art results compared with other methods. Mengting Wei, Wenming Zheng, Yuan Zong, Xingxun Jiang, Cheng Lu 0005, Jiateng Liu |
ICASSP | 2 |
| 2022 | Seeking Salient Facial Regions for Cross-Database Micro-Expression RecognitionabstractCross-Database Micro-Expression Recognition (CD-MER) aims to develop the Micro-Expression Recognition (MER) methods with strong domain adaptability, i.e., the ability to recognize the Micro-Expressions (MEs) of different subjects captured by different imaging devices in different scenes. The development of CDMER is faced with two key problems: 1) the severe feature distribution gap between the source and target databases; 2) the feature representation bottleneck of ME such local and subtle facial expressions. To solve these problems, this paper proposes a novel Transfer Group Sparse Regression method, namely TGSR, which aims to 1) optimize the measurement and better alleviate the difference between the source and target databases, and 2) highlight the valid facial regions to enhance extracted features, by the operation of selecting the group features from the raw face feature, where each region is associated with a group of raw face feature, i.e., the salient facial region selection. Compared with previous transfer group sparse methods, our proposed TGSR has the ability to select the salient facial regions, which is effective in alleviating aforementioned problems for better performance and reducing the computational cost at the same time. We use two public ME databases, i.e., CASME II and SMIC, to evaluate our proposed TGSR method. Experimental results show that our proposed TGSR learns the discriminative and explicable regions, and outperforms most state-of-the-art subspace-learning-based domain-adaptive methods for CDMER. Xingxun Jiang, Yuan Zong, Wenming Zheng, Jiateng Liu, Mengting Wei |
ICPR | 3 |
| 2022 | A Novel Magnification-Robust Network with Sparse Self-Attention for Micro-expression RecognitionabstractExisting works for spontaneous Micro-Expression Recognition (MER) tend to encode Micro-Expression (ME) movements to get more discriminative features. However, MEs’ low intensity makes the capture for motion extremely difficult, and the widely adopted unified-magnification strategy is prone to noise and lacks flexibility. To this end, this paper provides a new insight to encode ME motion and tackle magnification noise. Specifically, we reconstruct a new sequence via magnification techniques to make subtle ME movements more distinguishable. Afterward, Sparse Self-Attention (SSA) rectifies self-attention with Locality Sensitive Hashing (LSH), cutting the space into several hush buckets of related features. Only keys in the same bucket are operated in the attention term for every query feature. The resulting sparsity in the attention matrix prevents the network from attending features stemming from less-informative magnification degrees which could be regarded as noise, while retains the sequence modelling capability of standard self-attention. Extensive experiments on three public MER databases demonstrate our superiority against the state-of-the-art methods. Mengting Wei, Wenming Zheng, Xingxun Jiang, Yuan Zong, Cheng Lu 0005, Jiateng Liu |
ICPR | 2 |
| 2022 | Sample Self-Revised Network for Cross-Dataset Facial Expression RecognitionabstractFacial images with low quality, subjective annotation, severe occlusion, and rare subject identity can lead to the existence of outlier samples in facial expression datasets. These outlier samples are usually far from the center of the dataset in the feature space, resulting in huge differences in feature distribution, which severely restricts the performance of cross-dataset facial expression recognition (FER). To eliminate the influence of outlier samples on cross-dataset FER, we propose an unsupervised domain adaptation (UDA) method called Sample Self-Revised Network (SSRN), which 1) dynamically detects the outlier level of each sample in the source domain to reduce the disturbance of outlier samples to the model training, as well as 2) adaptively revises outlier samples in the source domain to improve transferability of the learned features. Experimental results show that our SSRN outperforms both classic deep UDA methods and state-of-the-art cross-dataset FER results. Wenming Zheng, Yuan Zong, Cheng Lu 0005, Xingxun Jiang |
IJCNN | 2 |
| 2022 | Adaptive Hierarchical Graph Convolutional Network for EEG Emotion RecognitionabstractHuman emotion is closely related to multiple distributed brain regions, and functional connections exist between the regions. However, how to abstract the region-level information to improve electroencephalograph (EEG) emotion recognition performance has not been well considered. To address this problem, we proposed a novel Adaptive Hierarchical Graph Convolutional Network (AHGCN), which includes the basic channel-level graph of EEG channels and the region-level graph of brain regions. Different from previous methods, we propose an adaptive pooling operation to automatically partition brain regions rather than manually define them. To capture the intrinsic functional connections between the brain regions or EEG channels, we design a gated adaptive graph convolution operation. Besides, we develop a graph unpooling operation to integrate the region-level graph and channel-level graph to extract more discrimination features for classification. Experiments on two widely-used datasets show that our proposed method is superior to many state-of-the-art methods on EEG emotion recognition and could find some interesting combinations of EEG channels. Yunlong Xue, Wenming Zheng, Yuan Zong, Hongli Chang, Xingxun Jiang |
IJCNN | 2 |
| 2022 | Coupled Discriminant Subspace Alignment for Cross-database Speech Emotion Recognition
Shaokai Li, Peng Song 0002, Keke Zhao, Wenjing Zhang 0003, Wenming Zheng |
INTERSPEECH | 5 |
| 2022 | Deep Transductive Transfer Regression Network for Cross-Corpus Speech Emotion Recognition
Yan Zhao 0037, Jincen Wang, Ru Ye, Yuan Zong, Wenming Zheng, Li Zhao 0003 |
INTERSPEECH | 5 |
| 2022 | Motion cues guided feature aggregation and enhancement for video object segmentation
Wenming Zheng, Yuan Zong |
Neurocomputing | 2 |
| 2022 | Cross-database micro-expression recognition based on transfer double sparse learning
Jiateng Liu, Yuan Zong, Wenming Zheng |
Multim. Tools Appl. | 3 |
| 2022 | From Regional to Global Brain: A Novel Hierarchical Spatial-Temporal Neural Network Model for EEG Emotion RecognitionabstractIn this paper, we propose a novel Electroencephalograph (EEG) emotion recognition method inspired by neuroscience with respect to the brain response to different emotions. The proposed method, denoted by R2G-STNN, consists of spatial and temporal neural network models with regional to global hierarchical feature learning process to learn discriminative spatial-temporal EEG features. To learn the spatial features, a bidirectional long short term memory (BiLSTM) network is adopted to capture the intrinsic spatial relationships of EEG electrodes within brain region and between brain regions, respectively. Considering that different brain regions play different roles in the EEG emotion recognition, a region-attention layer into the R2G-STNN model is also introduced to learn a set of weights to strengthen or weaken the contributions of brain regions. Based on the spatial feature sequences, BiLSTM is adopted to learn both regional and global spatial-temporal features and the features are fitted into a classifier layer for learning emotion-discriminative features, in which a domain discriminator working corporately with the classifier is used to decrease the domain shift between training and testing data. Finally, to evaluate the proposed method, we conduct both subject-dependent and subject-independent EEG emotion recognition experiments on SEED database, and the experimental results show that the proposed method achieves state-of-the-art performance. Yang Li 0019, Wenming Zheng, Lei Wang 0001, Yuan Zong, Zhen Cui 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2022 | Objective Class-Based Micro-Expression Recognition Under Partial Occlusion Via Region-Inspired Relation Reasoning NetworkabstractMicro-expression recognition (MER) has attracted the attention of many researchers in the past decade. However, occlusion occurs for MER in real-world scenarios. In this paper, a challenging issue in MER that is interesting but unexplored, i.e., occlusion MER, is deeply investigated. First, to research MER under real-world occlusion conditions, synthetic occluded microexpression databases are created by using various community masks. Second, to suppress the influence of occlusion, aRegion-inspiredRelationReasoningNetwork (RRRN) is proposed to model the relations between various facial regions. The RRRN consists of a backbone network, a region-inspired (RI) module and a relation reasoning (RR) module. More specifically, the backbone network aims to extract feature representations from different facial regions, the RI module is designed to compute the adaptive weight from the facial region itself based on the unobstructedness and importance of the region for suppressing the influence of occlusion using an attention mechanism, and the RR module exploits the progressive interactions among these regions by performing graph convolutions. Experiments are conducted on two tasks of MEGC 2018: the holdout-database evaluation task and the composite database evaluation task. Experimental results show that RRRN can be utilized to significantly explore the importance of facial regions and capture the cooperative complementary relationship of facial regions for MER. The results also demonstrate that RRRN outperforms the state-of-the-art approaches, especially with respect to occlusion, where RRRN is more robust. Qirong Mao, Ling Zhou 0005, Wenming Zheng, Xiuyan Shao, Xiaohua Huang 0003 |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | Domain Invariant Feature Learning for Speaker-Independent Speech Emotion RecognitionabstractIn this paper, we propose a novel domain invariant feature learning (DIFL) method to deal with speaker-independent speech emotion recognition (SER). The basic idea of DIFL is to learn the speaker-invariant emotion feature by eliminating domain shifts between the training and testing data caused by different speakers from the perspective of multi-source unsupervised domain adaptation (UDA). Specifically, we embed a hierarchical alignment layer with the strong-weak distribution alignment strategy into the feature extraction block to firstly reduce the discrepancy in feature distributions of speech samples across different speakers as much as possible. Furthermore, multiple discriminators in the discriminator block are utilized to confuse the speaker information of emotion features both inside the training data and between the training and testing data. Through them, a multi-domain invariant representation of emotional speech can be gradually and adaptively achieved by updating network parameters. We conduct extensive experiments on three public datasets, i. e., Emo-DB, eNTERFACE, and CASIA, to evaluate the SER performance of the proposed method, respectively. The experimental results show that the proposed method is superior to the state-of-the-art methods. Cheng Lu 0005, Yuan Zong, Wenming Zheng, Yang Li 0019, Chuangao Tang, Björn W. Schuller |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Cross-Database Micro-Expression Recognition: A BenchmarkabstractCross-database micro-expression recognition (CDMER) is one of recently emerging and interesting problem in micro-expression analysis. CDMER is more challenging than the conventional micro-expression recognition (MER), because the training and testing samples in CDMER come from different micro-expression databases, resulting in inconsistency of the feature distributions between the training and testing sets. In this paper, we contribute to this topic from three aspects. First, we establish a CDMER experimental evaluation protocol aiming to allow the researchers to conveniently work on this topic and evaluate their proposed methods under the same standard. Second, we conduct benchmark experiments by using NINE state-of-the-art domain adaptation (DA) methods and SIX popular spatiotemporal descriptors for investigating CDMER problem from two different perspectives. Third, we propose a novel DA method called region selective transfer regression (RSTR) to deal with the CDMER task. The overall superior performance of RSTR over the state-of-the-art DA methods demonstrates that taking into consideration the facial local region information used in RSTR contributes to developing effective DA methods for dealing with CDMER problem. Tong Zhang 0015, Yuan Zong, Wenming Zheng, C. L. Philip Chen, Xiaopeng Hong, Chuangao Tang, Zhen Cui 0001, Guoying Zhao 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2021 | Uncertain Graph Neural Networks for Facial Action Unit DetectionabstractCapturing the dependencies among different facial action units (AU) is extremely important for the AU detection task. Many studies have employed graph-based deep learning methods to exploit the dependencies among AUs. However, the dependencies among AUs in real world data are often noisy and the uncertainty is essential to be taken into consideration. Rather than employing a deterministic mode, we propose an uncertain graph neural network (UGN) to learn the probabilistic mask that simultaneously captures both the individual dependencies among AUs and the uncertainties. Further, we propose an adaptive weighted loss function based on the epistemic uncertainties to adaptively vary the weights of the training samples during the training process to account for unbalanced data distributions among AUs. We also provide an insightful analysis on how the uncertainties are related to the performance of AU detection. Extensive experiments, conducted on two benchmark datasets, i.e., BP4D and DISFA, demonstrate our method achieves the state-of-the-art performance. Tengfei Song, Lisha Chen, Wenming Zheng |
AAAI | 3 |
| 2021 | Dynamic Probabilistic Graph Convolution for Facial Action Unit Intensity EstimationabstractDeep learning methods have been widely applied to automatic facial action unit (AU) intensity estimation and achieved the state-of-the-art performance. These methods, however, are mostly appearance-based and fail to exploit the underlying structural information among AUs. In this paper, we propose a novel dynamic probabilistic graph convolution (DPG) model to simultaneously exploit AU appearances, AU dynamics, and their semantic structural dependencies for AU intensity estimation. Firstly, we propose to use Bayesian Network to capture the inherent dependencies among AUs. Secondly, we introduce probabilistic graph convolution that allows to perform graph convolution on the distribution of Bayesian Network structure to extract AU structural features. Finally, we introduce a dynamic deep model based on LSTM to simultaneously combine AU appearance features, AU dynamic features, and AU structural features for AU intensity estimation. In experiments, our method achieves comparable and even better performance with the state-of-the-art methods on two benchmark facial AU intensity estimation databases, i.e., FERA 2015 and DISFA. Tengfei Song, Zijun Cui, Yuru Wang, Wenming Zheng |
CVPR | 4 |
| 2021 | Hybrid Message Passing With Performance-Driven Structures for Facial Action Unit DetectionabstractMessage passing neural network has been an effective method to represent dependencies among nodes by propagating messages. However, most of message passing algorithms focus on one structure and messages are estimated by one single approach. For real-world data, like facial action units (AUs), the dependencies may vary in terms of different expressions and individuals. In this paper, we propose a novel hybrid message passing neural network with performance-driven structures (HMP-PS), which combines complementary message passing methods and captures more possible structures in a Bayesian manner. Particularly, a performance-driven Monte Carlo Markov Chain sampling method is proposed for generating high performance graph structures. Besides, hybrid message passing is proposed to combine different types of messages, which provide the complementary information. The contribution of each type of message is adaptively adjusted along with different inputs. The experiments on two widely used benchmark datasets, i.e., BP4D and DISFA, validate that our proposed method can achieve the state-of-the-art performance. Tengfei Song, Zijun Cui, Wenming Zheng |
CVPR | 3 |
| 2021 | Cross-Corpus Speech Emotion Recognition Using Joint Distribution Adaptive RegressionabstractIn this paper, we focus on the research of cross-corpus speech emotion recognition (SER), in which the training and testing speech signals in cross-corpus SER belong to dierent speech corpus. Due to this fact, mismatched feature distributions may exist between the training and testing speech feature sets degrading the performance of most originally well-performing SER methods. To deal with cross-corpus SER, we propose a novel domain adaptation (DA) method called joint distribution adaptive regression (JDAR). The basic idea of JDAR is to learn a regression matrix by jointly considering the marginal and conditional probability distribution between the training and testing speech signals and hence their feature distribution dierence can be alleviated in the subspace spanned by the learned regression matrix. To evaluate the proposed JDAR, we conduct extensive cross-corpus SER experiments on EmoDB, eNTERFACE, and CASIA speech databases. Experimental results show that the proposed JDAR achieves satisfactory performance and outperforms most of state-of-the-art subspace learning based DA methods. Lin Jiang 0007, Yuan Zong, Wenming Zheng, Li Zhao 0003 |
ICASSP | 4 |
| 2021 | Attention-based Spatio-Temporal Graphic LSTM for EEG Emotion RecognitionabstractAutomatic emotion recognition based on electroencephalogram (EEG) is a challenging task in Brain Machine Interfaces (BMI). Since it is still not very clear about the intrinsic connection relationship among the various EEG channels, it is still a challenging task of how to better represent the topology of EEG channels for emotion recognition. On the other hand, the intensity of the emotion may vary along the different time instants, which would affect the recognition accuracy of emotion. To tackle the above issues, in this paper, we propose a novel multichannel EEG emotion recognition method called attention-based spatiotemporal graphic long short-term memory (ASTG-LSTM), in which a Dynamic Structured Learning (DSL) branch that focuses on the most emotion-relevant connectivity of the brain is incorporated to represent inter-channel connections of the EEG signals. In addition, a specific spatio-temporal attention is embedded into the DSL branch to improve the invariance ability against the emotional intensity fluctuation. Extensive experiments on: DEAP and DREAMER are conducted and the experimental results indicate that the proposed ASTG-LSTM model improves the EEG emotion recognition performance compared with many state-of-the-art approaches. Wenming Zheng, Yuan Zong, Hongli Chang, Cheng Lu 0005 |
IJCNN | 2 |
| 2021 | A novel transferability attention neural network model for EEG emotion recognition
Yang Li 0019, Boxun Fu, Fu Li 0002, Guangming Shi, Wenming Zheng |
Neurocomputing | 5 |
| 2021 | Editorial for the special issue of IMAVIS on automatic face analytics for human behavior understanding
Xiaohua Huang 0003, Abhinav Dhall, Guoying Zhao 0001, Wenming Zheng, Matti Pietikäinen |
Image Vis. Comput. | 4 |
| 2021 | A Bi-Hemisphere Domain Adversarial Neural Network Model for EEG Emotion RecognitionabstractIn this paper, we propose a novel neural network model, called bi-hemisphere domain adversarial neural network (BiDANN) model, for electroencephalograph (EEG) emotion recognition. The BiDANN model is inspired by the neuroscience findings that the left and right hemispheres of human's brain are asymmetric to the emotional response. It contains a global and two local domain discriminators that work adversarially with a classifier to learn discriminative emotional features for each hemisphere. At the same time, it tries to reduce the possible domain differences in each hemisphere between the source and target domains so as to improve the generality of the recognition model. In addition, we also propose an improved version of BiDANN, denoted by BiDANN-S, for subject-independent EEG emotion recognition problem by lowering the influences of the personal information of subjects to the EEG emotion recognition. Extensive experiments on the SEED database are conducted to evaluate the performance of both BiDANN and BiDANN-S. The experimental results have shown that the proposed BiDANN and BiDANN models achieve state-of-the-art performance in the EEG emotion recognition. Yang Li 0019, Wenming Zheng, Yuan Zong, Zhen Cui 0001, Tong Zhang 0021 |
IEEE Trans. Affect. Comput. | 2 |
| 2021 | Multi-scale discrepancy adversarial network for crosscorpus speech emotion recognitionabstractOne of the most critical issues in human-computer interaction applications is recognizing human emotions based on speech. In recent years, the challenging problem of cross-corpus speech emotion recognition (SER) has generated extensive research. Nevertheless, the domain discrepancy between training data and testing data remains a major challenge to achieving improved system performance. This paper introduces a novel multi-scale discrepancy adversarial (MSDA) network for conducting multiple timescales domain adaptation for cross-corpus SER, i.e.,integrating domain discriminators of hierarchical levels into the emotion recognition framework to mitigate the gap between the source and target domains. Specifically, we extract two kinds of speech features, i.e., handcraft features and deep features, from three timescales of global, local, and hybrid levels. In each timescale, the domain discriminator and the emotion classifier compete against each other to learn features that minimize the discrepancy between the two domains by fooling the discriminator. Extensive experiments on cross-corpus and cross-language SER were conducted on a combination dataset that combines one Chinese dataset and two English datasets commonly used in SER. The MSDA is affected by the strong discriminate power provided by the adversarial process, where three discriminators are working in tandem with an emotion classifier. Accordingly, the MSDA achieves the best performance over all other baseline methods. The proposed architecture was tested on a combination of one Chinese and two English datasets. The experimental results demonstrate the superiority of our powerful discriminative model for solving cross-corpus SER. Wanlu Zheng, Wenming Zheng, Yuan Zong |
Virtual Real. Intell. Hardw. | 2 |
| 2020 | Instance-Adaptive Graph for EEG Emotion RecognitionabstractTo tackle the individual differences and characterize the dynamic relationships among different EEG regions for EEG emotion recognition, in this paper, we propose a novel instance-adaptive graph method (IAG), which employs a more flexible way to construct graphic connections so as to present different graphic representations determined by different input instances. To fit the different EEG pattern, we employ an additional branch to characterize the intrinsic dynamic relationships between different EEG channels. To give a more precise graphic representation, we design the multi-level and multi-graph convolutional operation and the graph coarsening. Furthermore, we present a type of sparse graphic representation to extract more discriminative features. Experiments on two widely-used EEG emotion recognition datasets are conducted to evaluate the proposed model and the experimental results show that our method achieves the state-of-the-art performance. Tengfei Song, Suyuan Liu, Wenming Zheng, Yuan Zong, Zhen Cui 0001 |
AAAI | 3 |
| 2020 | Variational Pathway Reasoning for EEG Emotion RecognitionabstractResearch on human emotion cognition revealed that connections and pathways exist between spatially-adjacent and functional-related areas during emotion expression (Adolphs 2002a; Bullmore and Sporns 2009). Deeply inspired by this mechanism, we propose a heuristic Variational Pathway Reasoning (VPR) method to deal with EEG-based emotion recognition. We introduce random walk to generate a large number of candidate pathways along electrodes. To encode each pathway, the dynamic sequence model is further used to learn between-electrode dependencies. The encoded pathways around each electrode are aggregated to produce a pseudo maximum-energy pathway, which consists of the most important pair-wise connections. To find those most salient connections, we propose a sparse variational scaling (SVS) module to learn scaling factors of pseudo pathways by using the Bayesian probabilistic process and sparsity constraint, where the former endows good generalization ability while the latter favors adaptive pathway selection. Finally, the salient pathways from those candidates are jointly decided by the pseudo pathways and scaling factors. Extensive experiments on EEG emotion recognition demonstrate that the proposed VPR is superior to those state-of-the-art methods, and could find some interesting pathways w.r.t. different emotions. Tong Zhang 0021, Zhen Cui 0001, Chunyan Xu, Wenming Zheng, Jian Yang 0003 |
AAAI | 4 |
| 2020 | DFEW: A Large-Scale Database for Recognizing Dynamic Facial Expressions in the WildabstractRecently, facial expression recognition (FER) in the wild has gained a lot of researchers' attention because it is a valuable topic to enable the FER techniques to move from the laboratory to the real applications. In this paper, we focus on this challenging but interesting topic and make contributions from three aspects. First, we present a new large-scale 'in-the-wild' dynamic facial expression database, DFEW (Dynamic Facial Expression in the Wild), consisting of over 16,000 video clips from thousands of movies. These video clips contain various challenging interferences in practical scenarios such as extreme illumination, occlusions, and capricious pose changes. Second, we propose a novel method called Expression-Clustered Spatiotemporal Feature Learning (EC-STFL) framework to deal with dynamic FER in the wild. Third, we conduct extensive benchmark experiments on DFEW using a lot of spatiotemporal deep feature learning methods as well as our proposed EC-STFL. Experimental results show that DFEW is a well-designed and challenging database, and the proposed EC-STFL can promisingly improve the performance of existing spatiotemporal deep neural networks in coping with the problem of dynamic FER in the wild. Our DFEW database is publicly available and can be freely downloaded from https://dfew-dataset.github.io/. Xingxun Jiang, Yuan Zong, Wenming Zheng, Chuangao Tang, Wanchuang Xia, Cheng Lu 0005, Jiateng Liu |
ACM Multimedia | 3 |
| 2020 | Feature Selection Based Transfer Subspace Learning for Speech Emotion RecognitionabstractCross-corpus speech emotion recognition has recently received considerable attention due to the widespread existence of various emotional speech. It takes one corpus as the training data aiming to recognize emotions of another corpus, and generally involves two basic problems, i.e., feature matching and feature selection. Many previous works study these two problems independently, or just focus on solving the first problem. In this paper, we propose a novel algorithm, called feature selection based transfer subspace learning (FSTSL), to address these two problems. To deal with the first problem, a latent common subspace is learnt by reducing the difference of different corpora and preserving the important properties. Meanwhile, we adopt the l2,1-norm on the projection matrix to deal with the second problem. Besides, to guarantee the subspace to be robust and discriminative, the geometric information of data is exploited simultaneously in the proposed FSTSL framework. Empirical experiments on cross-corpus speech emotion recognition tasks demonstrate that our proposed method can achieve encouraging results in comparison with state-of-the-art algorithms. Peng Song 0002, Wenming Zheng |
IEEE Trans. Affect. Comput. | 2 |
| 2020 | EEG Emotion Recognition Using Dynamical Graph Convolutional Neural NetworksabstractIn this paper, a multichannel EEG emotion recognition method based on a novel dynamical graph convolutional neural networks (DGCNN) is proposed. The basic idea of the proposed EEG emotion recognition method is to use a graph to model the multichannel EEG features and then perform EEG emotion classification based on this model. Different from the traditional graph convolutional neural networks (GCNN) methods, the proposed DGCNN method can dynamically learn the intrinsic relationship between different electroencephalogram (EEG) channels, represented by an adjacency matrix, via training a neural network so as to benefit for more discriminative EEG feature extraction. Then, the learned adjacency matrix is used to learn more discriminative features for improving the EEG emotion recognition. We conduct extensive experiments on the SJTU emotion EEG dataset (SEED) and DREAMER dataset. The experimental results demonstrate that the proposed method achieves better recognition performance than the state-of-the-art methods, in which the average recognition accuracy of 90.4 percent is achieved for subject dependent experiment while 79.95 percent for subject independent cross-validation one on the SEED database, and the average accuracies of 86.23, 84.54 and 85.02 percent are respectively obtained for valence, arousal and dominance classifications on the DREAMER database. Tengfei Song, Wenming Zheng, Peng Song 0002, Zhen Cui 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2020 | Toward Bridging Microexpressions From Different DomainsabstractRecently, microexpression recognition has attracted a lot of researchers' attention due to its challenges and valuable applications. However, it is noticed that currently most of the existing proposed methods are often evaluated and tested on the single database and, hence, this brings us a question whether these methods are still effective if the training and testing samples belong to different domains, for example, different microexpression databases. In this case, a large feature distribution difference may exist between training (source) and testing (target) samples and, hence, microexpression recognition tasks would become more difficult. To solve this challenging problem, that is, cross-domain microexpression recognition, in this paper, we propose an effective method consisting of an auxiliary set selection model (ASSM) and a transductive transfer regression model (TTRM). In our method, an ASSM is designed to automatically select an optimal set of samples from the target domain to serve as the auxiliary set, which is used for subsequent TTRM training. As for TTRM, it aims at bridging the feature distribution gap between the source and target domains by learning a joint regression model with the source domain samples and the auxiliary set selected from the target domain. We evaluate the proposed TTRM plus ASSM by extensive cross-domain microexpression recognition experiments on SMIC and CASME II databases. Compared with the recent state-of-the-art domain adaptation methods, our proposed method has a more satisfactory performance in dealing with the cross-domain microexpression recognition tasks. Yuan Zong, Wenming Zheng, Zhen Cui 0001, Guoying Zhao 0001, Bin Hu 0001 |
IEEE Trans. Cybern. | 2 |
| 2020 | Deep Manifold-to-Manifold Transforming Network for Skeleton-Based Action RecognitionabstractIn this paper, we will investigate skeleton-based action recognition by employing high-order statistics feature and first-order statistics feature, where the high-order statistics feature is characterized by symmetric positive definite (SPD) matrices. Noting that SPD matrices are theoretically embedded on Riemannian manifolds, we propose an end-to-end deep manifold-to-manifold transforming network (DMT-Net), which can make SPD matrices flow from one Riemannian manifold to another one for facilitating the action recognition task. To learn discriminative SPD features from both spatial and temporal dependencies, we propose a neural network model with three novel layers on manifolds: i.e., (1) the local SPD convolutional layer, (2) the non-linear SPD activation layer, and (3) the Riemannian-preserved recursive layer. The SPD property is preserved through all layers without the singular value decomposition (SVD) operation, which has to be conducted in the existing methods with expensive computation cost. Furthermore, a diagonalizing SPD layer is designed to efficiently calculate the final metric for the classification task. Finally, DMT-Net is further fused with a first order layer to capture temporal evolution information. To evaluate our proposed method, we conduct extensive experiments on the task of action recognition, where the input signals are represented as SPD matrices. The experimental results demonstrate that the proposed method is competitive over state-of-the-art methods. Tong Zhang 0021, Wenming Zheng, Zhen Cui 0001, Yuan Zong, Chaolong Li, Jian Yang 0003 |
IEEE Trans. Multim. | 2 |
| 2020 | Walk-Steered Convolution for Graph ClassificationabstractGraph classification is a fundamental but challenging issue for numerous real-world applications. Despite recent great progress in image/video classification, convolutional neural networks (CNNs) cannot yet cater to graphs well because of graphical non-Euclidean topology. In this article, we propose a walk-steered convolutional (WSC) network to assemble the essential success of standard CNNs, as well as the powerful representation ability of random walk. Instead of deterministic neighbor searching used in previous graphical CNNs, we construct multiscale walk fields (a.k.a. local receptive fields) with random walk paths to depict subgraph structures and advocate graph scalability. To express the internal variations of a walk field, Gaussian mixture models are introduced to encode the principal components of walk paths therein. As an analogy to a standard convolution kernel on image, Gaussian models implicitly coordinate those unordered vertices/nodes and edges in a local receptive field after projecting to the gradient space of Gaussian parameters. We further stack graph coarsening upon Gaussian encoding by using dynamic clustering, such that high-level semantics of graph can be well learned like the conventional pooling on image. The experimental results on several public data sets demonstrate the superiority of our proposed WSC method over many state of the arts for graph classification. Jiatao Jiang, Chunyan Xu, Zhen Cui 0001, Tong Zhang 0021, Wenming Zheng, Jian Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2019 | Bi-modality Fusion for Emotion Recognition in the WildabstractThe emotion recognition in the wild has been a hot research topic in the field of affective computing. Though some progresses have been achieved, the emotion recognition in the wild is still an unsolved problem due to the challenge of head movement, face deformation, illumination variation etc. To deal with these unconstrained challenges, we propose a bi-modality fusion method for video based emotion recognition in the wild. The proposed framework takes advantages of the visual information from facial expression sequences and the speech information from audio. The state-of-the-art CNN based object recognition models are employed to facilitate the facial expression recognition performance. A bi-direction long short term Memory (Bi-LSTM) is employed to capture dynamic information of the learned features. Additionally, to take full advantages of the facial expression information, the VGG16 network is trained on AffectNet dataset to learn a specialized facial expression recognition model. On the other hand, the audio based features, like low level descriptor (LLD) and deep features obtained by spectrogram image, are also developed to improve the emotion recognition performance. The best experimental result shows that the overall accuracy of our algorithm on the Test dataset of the EmotiW challenge is 62.78, which outperforms the best result of EmotiW2018 and ranks 2nd at the EmotiW2019 challenge. Sunan Li, Wenming Zheng, Yuan Zong, Cheng Lu 0005, Chuangao Tang, Xingxun Jiang, Jiateng Liu, Wanchuang Xia |
ICMI | 2 |
| 2019 | Sparse Graphic Attention LSTM for EEG Emotion Recognition
Suyuan Liu, Wenming Zheng, Tengfei Song, Yuan Zong |
ICONIP (4) | 2 |
| 2019 | Cross-Database Micro-Expression Recognition: A BenchmarkabstractCross-database micro-expression recognition (CDMER) is one of recently emerging and interesting problems in micro-expression analysis. CDMER is more challenging than the conventional micro-expression recognition (MER), because the training and testing samples in CDMER come from different micro-expression databases, resulting in inconsistency of the feature distributions between the training and testing sets. In this paper, we contribute to this topic from two aspects. First, we establish a CDMER experimental evaluation protocol and provide a standard platform for evaluating their proposed methods. Second, we conduct extensive benchmark experiments by using NINE state-of-the-art domain adaptation (DA) methods and SIX popular spatiotemporal descriptors for investigating the CDMER problem from two different perspectives and deeply analyze and discuss the experimental results. In addition, all the data and codes involving CDMER in this paper are released on our project website: http://aip.seu.edu.cn/cdmer. Yuan Zong, Wenming Zheng, Xiaopeng Hong, Chuangao Tang, Zhen Cui 0001, Guoying Zhao 0001 |
ICMR | 2 |
| 2019 | EEG Emotion Recognition Based on Graph Regularized Sparse Linear Regression
Yang Li 0019, Wenming Zheng, Zhen Cui 0001, Yuan Zong, Sheng Ge |
Neural Process. Lett. | 2 |
| 2019 | Recurrent Shape RegressionabstractAn end-to-end network architecture, the Recurrent Shape Regression (RSR), is presented to deal with the task of facial shape detection, a crucial step in many computer vision problems. The RSR generalizes the conventional cascaded regression into a recurrent dynamic network through abstracting common latent models with stage-to-stage operations. Instead of invariant regression transformation, we construct shape-dependent dynamic regressors to attain the recurrence of regression action itself. The regressors can be stacked into a high-order regression network to represent more complex shape regression. By further integrating feature learning as well as global shape constraint, the RSR becomes more controllable in entire optimization of shape regression, where the gradient computation can be efficiently back-propagated through time. To handle the possible partial occlusions of shapes, we propose a mimic virtual occlusion strategy by randomly disturbing certain point cliques without the requirement of any annotations of occlusion information or even occluded training data. Extensive experiments on five face datasets demonstrate that the proposed RSR outperforms the recent state-of-the-art cascaded approaches. Zhen Cui 0001, Shengtao Xiao, Zhiheng Niu, Shuicheng Yan, Wenming Zheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2019 | Spatial-Temporal Recurrent Neural Network for Emotion RecognitionabstractIn this paper, we propose a novel deep learning framework, called spatial-temporal recurrent neural network (STRNN), to integrate the feature learning from both spatial and temporal information of signal sources into a unified spatial-temporal dependency model. In STRNN, to capture those spatially co-occurrent variations of human emotions, a multidirectional recurrent neural network (RNN) layer is employed to capture long-range contextual cues by traversing the spatial regions of each temporal slice along different directions. Then a bi-directional temporal RNN layer is further used to learn the discriminative features characterizing the temporal dependencies of the sequences, where sequences are produced from the spatial RNN layer. To further select those salient regions with more discriminative ability for emotion recognition, we impose sparse projection onto those hidden states of spatial and temporal domains to improve the model discriminant ability. Consequently, the proposed two-layer RNN model provides an effective way to make use of both spatial and temporal dependencies of the input signals for emotion recognition. Experimental results on the public emotion datasets of electroencephalogram and facial expression demonstrate the proposed STRNN method is more competitive over those state-of-the-art methods. Tong Zhang 0021, Wenming Zheng, Zhen Cui 0001, Yuan Zong, Yang Li 0019 |
IEEE Trans. Cybern. | 2 |
| 2019 | Spectral Filter TrackingabstractVisual object tracking is a challenging computer vision task with numerous real-world applications. In this paper, we propose a simple but efficient Spectral Filter Tracking (SFT) method from the view of graph, where each candidate image region is modeled as a pixelwise grid graph. Instead of the conventional graph matching, we formulate the tracking as a plain least square regression problem of learning spectral filters on graphs to predict an optimal vertex, which indicates the center of the target. To bypass computationally expensive eigenvalue decomposition on graph Laplacian L, we parameterize spectral graph filters as a polynomial of L to aggregate local graph features according to spectral graph theory, in which Lk exactly encodes a k-hop local neighborhood of each vertex. Thus, different from the holistic regression in those correlation filter based methods, SFT can operate on localized regions around a pixel (i.e., a vertex), which can effectively reduce the influence of local variations and cluttered backgrounds. Furthermore, we observe that the correlation filter tracking may be viewed as a specific case of our proposed spectral filtering method. The implementation of SFT can simply boil down to only a few line codes, but surprisingly it beats the correlation filter based model with the same feature input, and achieves the state-of-the-art performance on OTB-2015 and VOT2016 under the same feature extraction strategy. Zhen Cui 0001, Youyi Cai, Wenming Zheng, Chunyan Xu, Jian Yang 0003 |
IEEE Trans. Image Process. | 3 |
| 2019 | Dynamic Texture Classification Using Unsupervised 3D Filter Learning and Local Binary EncodingabstractLocal binary descriptors, such as local binary pattern (LBP) and its various variants, have been studied extensively in texture and dynamic texture analysis due to their outstanding characteristics, such as grayscale invariance, low computational complexity and good discriminability. Most existing local binary feature extraction methods extract spatio-temporal features from three orthogonal planes of a spatio-temporal volume by viewing a dynamic texture in 3D space. For a given pixel in a video, only a proportion of its surrounding pixels is incorporated in the local binary feature extraction process. We argue that the ignored pixels contain discriminative information that should be explored. To fully utilize the information conveyed by all the pixels in a local neighborhood, we propose extracting local binary features from the spatio-temporal domain with 3D filters that are learned in an unsupervised manner so that the discriminative features along both the spatial and temporal dimensions are captured simultaneously. The proposed approach consists of three components: 1) 3D filtering; 2) binary hashing; and 3) joint histogramming. Densely sampled 3D blocks of a dynamic texture are first normalized to have zero mean and are then filtered by 3D filters that are learned in advance. To preserve more of the structure information, the filter response vectors are decomposed into two complementary components, namely, the signs and the magnitudes, which are further encoded separately into binary codes. The local mean pixels of the 3D blocks are also converted into binary codes. Finally, three types of binary codes are combined via joint or hybrid histograms for the final feature representation. Extensive experiments are conducted on three commonly used dynamic texture databases: 1) UCLA; 2) DynTex; and 3) YUVL. The proposed method provides comparable results to, and even outperforms, many state-of-the-art methods. Xiaochao Zhao, Yaping Lin, Li Liu 0002, Janne Heikkilä, Wenming Zheng |
IEEE Trans. Multim. | 5 |
| 2019 | ℓ1-Norm Heteroscedastic Discriminant Analysis Under Mixture of Gaussian DistributionsabstractFisher’s criterion is one of the most popular discriminant criteria for feature extraction. It is defined as the generalized Rayleigh quotient of the between-class scatter distance to the within-class scatter distance. Consequently, Fisher’s criterion does not take advantage of the discriminant information in the class covariance differences, and hence, its discriminant ability largely depends on the class mean differences. If the class mean distances are relatively large compared with the within-class scatter distance, Fisher’s criterion-based discriminant analysis methods may achieve a good discriminant performance. Otherwise, it may not deliver good results. Moreover, we observe that the between-class distance of Fisher’s criterion is based on the$\ell _{2}$-norm, which would be disadvantageous to separate the classes with smaller class mean distances. To overcome the drawback of Fisher’s criterion, in this paper, we first derive a new discriminant criterion, expressed as amixture of absolute generalized Rayleigh quotients, based on a Bayes error upper bound estimation, where mixture of Gaussians is adopted to approximate the real distribution of data samples. Then, the criterion is further modified by replacing$\ell _{2}$-norm with$\ell _{1}$one to better describe the between-class scatter distance, such that it would be more effective to separate the different classes. Moreover, we propose a novel$\ell _{1}$-norm heteroscedastic discriminant analysis method based on the new discriminant analysis (L1-HDA/GM) for heteroscedastic feature extraction, in which the optimization problem of L1-HDA/GM can be efficiently solved by using the eigenvalue decomposition approach. Finally, we conduct extensive experiments on four real data sets and demonstrate that the proposed method achieves much competitive results compared with the state-of-the-art methods. Wenming Zheng, Cheng Lu 0005, Zhouchen Lin, Tong Zhang 0021, Zhen Cui 0001, Wankou Yang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | Spatio-Temporal Graph Convolution for Skeleton Based Action RecognitionabstractVariations of human body skeletons may be considered as dynamic graphs, which are generic data representation for numerous real-world applications. In this paper, we propose a spatio-temporal graph convolution (STGC) approach for assembling the successes of local convolutional filtering and sequence learning ability of autoregressive moving average. To encode dynamic graphs, the constructed multi-scale local graph convolution filters, consisting of matrices of local receptive fields and signal mappings, are recursively performed on structured graph data of temporal and spatial domain. The proposed model is generic and principled as it can be generalized into other dynamic models. We theoretically prove the stability of STGC and provide an upper-bound of the signal transformation to be learnt. Further, the proposed recursive model can be stacked into a multi-layer architecture. To evaluate our model, we conduct extensive experiments on four benchmark skeleton-based action datasets, including the large-scale challenging NTU RGB+D. The experimental results demonstrate the effectiveness of our proposed model and the improvement over the state-of-the-art. Chaolong Li, Zhen Cui 0001, Wenming Zheng, Chunyan Xu, Jian Yang 0003 |
AAAI | 3 |
| 2018 | Deep Manifold-to-Manifold Transforming NetworkabstractIn this paper, we propose an end-to-end deep manifold-to-manifold transforming network (DMT-Net), which makes SPD matrices flow from one Riemannian manifold to another more discriminative one. For discriminative feature learning, two specific layers on manifolds are developed: (i) the local SPD convolutional layer, (ii) the non-linear SPD activation layer, where positive definiteness is satisfied for both two layers. Further, to relieve computational burden of kernels on relative large-scale data, we design a batch-kernelized layer to favor batchwise kernel optimization of deep networks. Specifically, one reference set dynamically changing with the network training is introduced to break the limitation of memory size. We evaluate our proposed method on action recognition datasets, where input signals are popularly modeled as SPD matrices. The experimental results demonstrate that our DMT-Net is more competitive than state-of-the-art methods. Tong Zhang 0021, Wenming Zheng, Zhen Cui 0001, Chaolong Li |
ICIP | 2 |
| 2018 | Multiple Spatio-temporal Feature Learning for Video-based Emotion Recognition in the WildabstractThe difficulty of emotion recognition in the wild (EmotiW) is how to train a robust model to deal with diverse scenarios and anomalies. The Audio-video Sub-challenge in EmotiW contains audio-video short clips with several emotional labels and the task is to distinguish which label the video belongs to. For the better emotion recognition in videos, we propose a multiple spatio-temporal feature fusion (MSFF) framework, which can more accurately depict emotional information in spatial and temporal dimensions by two mutually complementary sources, including the facial image and audio. The framework is consisted of two parts: the facial image model and the audio model. With respect to the facial image model, three different architectures of spatial-temporal neural networks are employed to extract discriminative features about different emotions in facial expression images. Firstly, the high-level spatial features are obtained by the pre-trained convolutional neural networks (CNN), including VGG-Face and ResNet-50 which are all fed with the images generated by each video. Then, the features of all frames are sequentially input to the Bi-directional Long Short-Term Memory (BLSTM) so as to capture dynamic variations of facial appearance textures in a video. In addition to the structure of CNN-RNN, another spatio-temporal network, namely deep 3-Dimensional Convolutional Neural Networks (3D CNN) by extending the 2D convolution kernel to 3D, is also applied to attain evolving emotional information encoded in multiple adjacent frames. For the audio model, the spectrogram images of speech generated by preprocessing audio, are also modeled in a VGG-BLSTM framework to characterize the affective fluctuation more efficiently. Finally, a fusion strategy with the score matrices of different spatio-temporal networks gained from the above framework is proposed to boost the performance of emotion recognition complementally. Extensive experiments show that the overall accuracy of our proposed MSFF is 60.64%, which achieves a large improvement compared with the baseline and outperform the result of champion team in 2017. Cheng Lu 0005, Wenming Zheng, Chaolong Li, Chuangao Tang, Suyuan Liu, Simeng Yan, Yuan Zong |
ICMI | 2 |
| 2018 | A Novel Neural Network Model based on Cerebral Hemispheric Asymmetry for EEG Emotion RecognitionabstractIn this paper, we propose a novel neural network model, called bi-hemispheres domain adversarial neural network (BiDANN), for EEG emotion recognition. BiDANN is motivated by the neuroscience findings, i.e., the emotional brain's asymmetries between left and right hemispheres. The basic idea of BiDANN is to map the EEG feature data of both left and right hemispheres into discriminative feature spaces separately, in which the data representations can be classified easily. For further precisely predicting the class labels of testing data, we narrow the distribution shift between training and testing data by using a global and two local domain discriminators, which work adversarially to the classifier to encourage domain-invariant data representations to emerge. After that, the learned classifier from labeled training data can be applied to unlabeled testing data naturally. We conduct two experiments to verify the performance of our BiDANN model on SEED database. The experimental results show that the proposed model achieves the state-of-the-art performance. Yang Li 0019, Wenming Zheng, Zhen Cui 0001, Tong Zhang 0021, Yuan Zong |
IJCAI | 2 |
| 2018 | Context-Dependent Diffusion Network for Visual Relationship DetectionabstractVisual relationship detection can bridge the gap between computer vision and natural language for scene understanding of images. Different from pure object recognition tasks, the relation triplets of subject-predicate-object lie on an extreme diversity space, such asperson-behind-person andcar-behind-building, while suffering from the problem of combinatorial explosion. In this paper, we propose a context-dependent diffusion network (CDDN) framework to deal with visual relationship detection. To capture the interactions of different object instances, two types of graphs, word semantic graph and visual scene graph, are constructed to encode global context interdependency. The semantic graph is built through language priors to model semantic correlations across objects, whilst the visual scene graph defines the connections of scene objects so as to utilize the surrounding scene information. For the graph-structured data, we design a diffusion network to adaptively aggregate information from contexts, which can effectively learn latent representations of visual relationships and well cater to visual relationship detection in view of its isomorphic invariance to graphs. Experiments on two widely-used datasets demonstrate that our proposed method is more effective and achieves the state-of-the-art performance. Zhen Cui 0001, Chunyan Xu, Wenming Zheng, Jian Yang 0003 |
ACM Multimedia | 3 |
| 2018 | Face recognition based on recurrent regression neural network
Yang Li 0019, Wenming Zheng, Zhen Cui 0001, Tong Zhang 0021 |
Neurocomputing | 2 |
| 2018 | Multi-cue fusion for emotion recognition in the wild
Jingwei Yan, Wenming Zheng, Zhen Cui 0001, Chuangao Tang, Tong Zhang 0021, Yuan Zong |
Neurocomputing | 2 |
| 2018 | Unsupervised facial expression recognition using domain adaptation based dictionary learning approach
Wenming Zheng, Zhen Cui 0001, Yuan Zong, Tong Zhang 0021, Chuangao Tang |
Neurocomputing | 2 |
| 2018 | Cross-Domain Color Facial Expression Recognition Using Transductive Transfer Subspace LearningabstractFacial expression recognition across domains, e.g., training and testing facial images come from different facial poses, is very challenging due to the different marginal distributions between training and testing facial feature vectors. To deal with such challenging cross-domain facial expression recognition problem, a novel transductive transfer subspace learning method is proposed in this paper. In this method, a labelled facial image set from source domain is combined with an unlabelled auxiliary facial image set from target domain to jointly learn a discriminative subspace and make the class labels prediction of the unlabelled facial images, where a transductive transfer regularized least-squares regression (TTRLSR) model is proposed to this end. Then, based on the auxiliary facial image set, we train a SVM classifier for classifying the expressions of other facial images in the target domain. Moreover, we also investigate the use of color facial features to evaluate the recognition performance of the proposed facial expression recognition method, where color scale invariant feature transform (CSIFT) features associated with 49 landmark facial points are extracted to describe each color facial image. Finally, extensive experiments on BU-3DFE and Multi-PIE multiview color facial expression databases are conducted to evaluate the cross-database & cross-view facial expression recognition performance of the proposed method. Comparisons with state-of-the-art domain adaption methods are also included in the experiments. The experimental results demonstrate that the proposed method achieves much better recognition performance compared with the state-of-the-art methods. Wenming Zheng, Yuan Zong, Minghai Xin |
IEEE Trans. Affect. Comput. | 1 |
| 2018 | Action-Attending Graphic Neural NetworkabstractThe motion analysis of human skeletons is crucial for human action recognition, which is one of the most active topics in computer vision. In this paper, we propose a fully end-to-end action-attending graphic neural network (A2GNN) for skeleton-based action recognition, in which each irregular skeleton is structured as an undirected attribute graph. To extract high-level semantic representation from skeletons, we perform the local spectral graph filtering on the constructed attribute graphs like the standard image convolution operation. Considering not all joints are informative for action analysis, we design an actionattending layer to detect those salient action units (AUs) by adaptively weighting skeletal joints. Herein the filtering responses are parameterized into a weighting function irrelevant to the order of input nodes. To further encode continuous motion variations, the deep features learnt from skeletal graphs are gathered along consecutive temporal slices and then fed into a recurrent gated network. Finally, the spectral graph filtering, action-attending and recurrent temporal encoding are integrated together to jointly train for the sake of robust action recognition as well as the intelligibility of human actions. To evaluate our A2GNN, we conduct extensive experiments on four benchmark skeletonbased action datasets, including the large-scale challenging NTU RGB+D dataset. The experimental results demonstrate that our network achieves the state-of-the-art performances. Chaolong Li, Zhen Cui 0001, Wenming Zheng, Chunyan Xu, Rongrong Ji, Jian Yang 0003 |
IEEE Trans. Image Process. | 3 |
| 2018 | Domain Regeneration for Cross-Database Micro-Expression RecognitionabstractRecently, micro-expression recognition has attracted lots of researchers' attention due to its potential value in many practical applications, e.g., lie detection. In this paper, we investigate an interesting and challenging problem in micro-expression recognition, i.e., cross-database micro-expression recognition, in which the training and testing samples come from different micro-expression databases. Under this problem setting, the consistent feature distribution between the training and testing samples originally existing in conventional micro-expression recognition would be seriously broken and hence the performance of most current well-performing micro-expression recognition methods may sharply drop. In order to overcome it, we propose a simple yet effective framework called Domain Regeneration (DR) in this paper. DR framework aims at learning a domain regenerator to regenerate the micro-expression samples from source and target databases respectively such that they can abide by the same or similar feature distributions. Thus, we are able to use the classifier learned based on the labeled source micro-expression samples to predict the label information of the unlabeled target micro-expression samples. To evaluate the proposed DR framework, we conduct extensive cross-database micro-expression recognition experiments designed based on SMIC and CASME II databases. Experimental results show that compared with recent state-of-the-art cross-database emotion recognition methods, the proposed DR framework has more promising performance. Yuan Zong, Wenming Zheng, Xiaohua Huang 0003, Jingang Shi, Zhen Cui 0001, Guoying Zhao 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Sinusoidal Signal Assisted Multivariate Empirical Mode Decomposition for Brain-Computer InterfacesabstractA brain-computer interface (BCI) is a communication approach that permits cerebral activity to control computers or external devices. Brain electrical activity recorded with electroencephalography (EEG) is most commonly used for BCI. Noise-assisted multivariate empirical mode decomposition (NA-MEMD) is a data-driven time-frequency analysis method that can be applied to nonlinear and nonstationary EEG signals for BCI data processing. However, because white Gaussian noise occupies a broad range of frequencies, some redundant components are introduced. To solve this leakage problem, in this study, we propose using a sinusoidal assisted signal that occupies the same frequency ranges as the original signals to improve MEMD performance. To verify the effectiveness of the proposed sinusoidal signal assisted MEMD (SA-MEMD) method, we compared the decomposition performances of MEMD, NA-MEMD, and the proposed SA-MEMD using synthetic signals and a real-world BCI dataset. The spectral decomposition results indicate that the proposed SA-MEMD can avoid the generation of redundant components and over decomposition, thus, substantially reduce the mode mixing and misalignment that occurs in MEMD and NA-MEMD. Moreover, using SA-MEMD as a signal preprocessing method instead of MEMD or NA-MEMD can significantly improve BCI classification accuracy and reduce calculation time, which indicates that SA-MEMD is a powerful spectral decomposition method for BCI. Sheng Ge, Yanhua Shi, Pan Lin, Junfeng Gao, Gao-Peng Sun, Keiji Iramina, Yuankui Yang, Yue Leng, Haixian Wang, Wenming Zheng |
IEEE J. Biomed. Health Informatics | 11 |
| 2018 | Learning From Hierarchical Spatiotemporal Descriptors for Micro-Expression RecognitionabstractMicro-expression recognition aims to infer genuine emotions that people try to conceal from facial video clips. It is a very challenging task because micro-expressions have a very low intensity and short duration, which makes micro-expressions difficult to observe. Recently, researchers have designed various spatiotemporal descriptors to describe micro-expressions. It is notable that for better capturing the low-intensity facial muscle movement, a fixed spatial division grid, 8× 8 for example, is commonly used to partition the facial images into a few facial blocks before extracting descriptors. However, it is hard to choose an ideal division grid for different micro-expression samples because the division grids affect the discriminative ability of spatiotemporal descriptors to distinguish micro-expressions. To address this problem, in this paper, we design a hierarchical spatial division scheme for spatiotemporal descriptor extraction. By using the proposed scheme, it would not be a problem to determine which division grid is most suitable regarding different micro-expression samples. Furthermore, we propose a kernelized group sparse learning (KGSL) model to process hierarchical scheme based spatiotemporal descriptors such that they are more effective for micro-expression recognition tasks. To evaluate the performance of the proposed micro-expression recognition method consisting of the hierarchical scheme based spatiotemporal descriptors and KGSL, extensive experiments are conducted on two public micro-expression databases: CASME II and SMIC. Compared with many recent state-of-the-art approaches, our method achieves more promising recognition results. Yuan Zong, Xiaohua Huang 0003, Wenming Zheng, Zhen Cui 0001, Guoying Zhao 0001 |
IEEE Trans. Multim. | 3 |
| 2017 | View-Independent Facial Action Unit DetectionabstractAutomatic Facial Action Unit (AU) detection has drawn more and more attention over the past years due to its significance to facial expression analysis. Frontal-view AU detection has been extensively evaluated, but cross-pose AU detection is a less-touched problem due to the scarcity of the related dataset. The challenge of Facial Expression Recognition and Analysis (FERA2017) just released a large-scale videobased AU detection dataset across different facial poses. To deal with this challenging task, we develop a simple and efficient deep learning based system to detect AU occurrence under nine different facial views. In this system, we first crop out facial images by using morphology operations including binary segmentation, connected components labeling and region boundaries extraction, then for each type of AU, we train a corresponding expert network by specifically fine-tuning the VGG-Face network on cross-view facial images, so as to extract more discriminative features for the subsequent binary classification. In the AU detection sub-challenge, our proposed method achieves the mean accuracy of 77.8% (vs. the baseline 56.1%), and promotes the F1 score to 57.4% (vs. the baseline 45.2%). Chuangao Tang, Wenming Zheng, Jingwei Yan, Qiang Li 0044, Yang Li 0019, Tong Zhang 0021, Zhen Cui 0001 |
FG | 2 |
| 2017 | Learning a Target Sample Re-Generator for Cross-Database Micro-Expression RecognitionabstractIn this paper, we investigate the cross-database micro-expression recognition problem, where the training and testing samples are from two different micro-expression databases. Under this setting, the training and testing samples would have different feature distributions and hence the performance of most existing micro-expression recognition methods may decrease greatly. To solve this problem, we propose a simple yet effective method called Target Sample Re-Generator (TSRG) in this paper. By using TSRG, we are able to re-generate the samples from target micro-expression database and the re-generated target samples would share same or similar feature distributions with the original source samples. For this reason, we can then use the classifier learned based on the labeled source samples to accurately predict the micro-expression categories of the unlabeled target samples. To evaluate the performance of the proposed TSRG method, extensive cross-database micro-expression recognition experiments designed based on SMIC and CASME II databases are conducted. Compared with recent state-of-the-art cross-database emotion recognition methods, the proposed TSRG achieves more promising results. Yuan Zong, Xiaohua Huang 0003, Wenming Zheng, Zhen Cui 0001, Guoying Zhao 0001 |
ACM Multimedia | 3 |
| 2017 | Locality-constrained linear coding based bi-layer model for multi-view facial expression recognition
Jianlong Wu, Zhouchen Lin, Wenming Zheng, Hongbin Zha |
Neurocomputing | 3 |
| 2017 | Gender classification using 3D statistical models
Wankou Yang, Changyin Sun 0001, Wenming Zheng, Karl Ricanek |
Multim. Tools Appl. | 3 |
| 2016 | Speech emotion recognition using transfer non-negative matrix factorizationabstractIn practical situations, the emotional speech utterances are often collected from different devices and conditions, which will obviously affect the recognition performance. To address this issue, in this paper, a novel transfer non-negative matrix factorization (TNMF) method is presented for cross-corpus speech emotion recognition. First, the NMF algorithm is adopted to learn a latent common feature space for the source and target datasets. Then, the discrepancies between the feature distributions of different corpora are considered, and the maximum mean discrepancy (MMD) algorithm is used for the similarity measurement. Finally, the TNMF approach, which integrates the NMF and MMD algorithms, is proposed. Experiments are carried out on two popular datasets, and the results verify that the TNMF method can significantly outperform the automatic and competitive methods for cross-corpus speech emotion recognition. Peng Song 0002, Shifeng Ou, Wenming Zheng, Yun Jin, Li Zhao 0003 |
ICASSP | 3 |
| 2016 | Multi-clue fusion for emotion recognition in the wildabstractIn the past three years, Emotion Recognition in the Wild (EmotiW) Grand Challenge has drawn more and more attention due to its huge potential applications. In the fourth challenge, aimed at the task of video based emotion recognition, we propose a multi-clue emotion fusion (MCEF) framework by modeling human emotion from three mutually complementary sources, facial appearance texture, facial action, and audio. To extract high-level emotion features from sequential face images, we employ a CNN-RNN architecture, where face image from each frame is first fed into the fine-tuned VGG-Face network to extract face feature, and then the features of all frames are sequentially traversed in a bidirectional RNN so as to capture dynamic changes of facial textures. To attain more accurate facial actions, a facial landmark trajectory model is proposed to explicitly learn emotion variations of facial components. Further, audio signals are also modeled in a CNN framework by extracting low-level energy features from segmented audio clips and then stacking them as an image-like map. Finally, we fuse the results generated from three clues to boost the performance of emotion recognition. Our proposed MCEF achieves an overall accuracy of 56.66% with a large improvement of 16.19% with respect to the baseline. Jingwei Yan, Wenming Zheng, Zhen Cui 0001, Chuangao Tang, Tong Zhang 0021, Yuan Zong, Ning Sun 0001 |
ICMI | 2 |
| 2016 | A Novel Graph Regularized Sparse Linear Discriminant Analysis Model for EEG Emotion Recognition
Yang Li 0019, Wenming Zheng, Zhen Cui 0001 |
ICONIP (4) | 2 |
| 2016 | Cross-Database Facial Expression Recognition via Unsupervised Domain Adaptive Dictionary Learning
Wenming Zheng, Zhen Cui 0001, Yuan Zong |
ICONIP (2) | 2 |
| 2016 | Spontaneous facial micro-expression analysis using Spatiotemporal Completed Local Quantized Patterns
Xiaohua Huang 0003, Guoying Zhao 0001, Xiaopeng Hong, Wenming Zheng, Matti Pietikäinen |
Neurocomputing | 4 |
| 2016 | A regularized least square based discriminative projections for feature extraction
Wankou Yang, Changyin Sun 0001, Wenming Zheng |
Neurocomputing | 3 |
| 2016 | Cross-corpus speech emotion recognition based on transfer non-negative matrix factorization
Peng Song 0002, Wenming Zheng, Shifeng Ou, Yun Jin, Jinglei Liu, Yanwei Yu |
Speech Commun. | 2 |
| 2016 | Cross-Corpus Speech Emotion Recognition Based on Domain-Adaptive Least-Squares RegressionabstractIn this letter, a novel cross-corpus speech emotion recognition (SER) method using domain-adaptive least-squares regression (DaLSR) model is proposed. In this method, an additional unlabeled data set from target speech corpus is used to serve as an auxiliary data set and combined with the labeled training data set from source speech corpus for jointly training the DaLSR model. In contrast to the traditional least-squares regression (LSR) method, the major novelty of DaLSR is that it is able to handle the mismatch problem between source and target speech corpora. Hence, the proposed DaLSR method is very suitable for coping with cross-corpus SER problem. For evaluating the performance of the proposed method in dealing with the cross-corpus SER problem, we conduct extensive experiments on three emotional speech corpora and compare the results with several state-of-the-art transfer learning methods that are widely used for cross-corpus SER problem. The experimental results show that the proposed method achieves better recognition accuracies than the state-of-the-art methods. Yuan Zong, Wenming Zheng, Tong Zhang 0021, Xiaohua Huang 0003 |
IEEE Signal Process. Lett. | 2 |
| 2016 | Sparse Kernel Reduced-Rank Regression for Bimodal Emotion Recognition From Facial Expression and SpeechabstractA novel bimodal emotion recognition approach from facial expression and speech based on the sparse kernel reduced-rank regression (SKRRR) fusion method is proposed in this paper. In this method, we use the openSMILE feature extractor and the scale invariant feature transform feature descriptor to respectively extract effective features from speech modality and facial expression modality, and then propose the SKRRR fusion approach to fuse the emotion features of two modalities. The proposed SKRRR method is a nonlinear extension of the traditional reduced-rank regression (RRR), where both predictor and response feature vectors in RRR are kernelized by being mapped onto two high-dimensional feature space via two nonlinear mappings, respectively. To solve the SKRRR problem, we propose a sparse representation (SR)-based approach to find the optimal solution of the coefficient matrices of SKRRR, where the introduction of the SR technique aims to fully consider the different contributions of training data samples to the derivation of optimal solution of SKRRR. Finally, we utilize the eNTERFACE '05 and AFEW 4.0 bimodal emotion database to conduct the experiments of monomodal emotion recognition and bimodal emotion recognition, and the results indicate that our presented approach acquires the highest or comparable bimodal emotion recognition rate among some state-of-the-art approaches. Jingjie Yan, Wenming Zheng, Qinyu Xu, Guanming Lu, Haibo Li 0001 |
IEEE Trans. Multim. | 2 |
| 2016 | A Deep Neural Network-Driven Feature Learning Method for Multi-view Facial Expression RecognitionabstractIn this paper, a novel deep neural network (DNN)-driven feature learning method is proposed and applied to multi-view facial expression recognition (FER). In this method, scale invariant feature transform (SIFT) features corresponding to a set of landmark points are first extracted from each facial image. Then, a feature matrix consisting of the extracted SIFT feature vectors is used as input data and sent to a well-designed DNN model for learning optimal discriminative features for expression classification. The proposed DNN model employs several layers to characterize the corresponding relationship between the SIFT feature vectors and their corresponding high-level semantic information. By training the DNN model, we are able to learn a set of optimal features that are well suitable for classifying the facial expressions across different facial views. To evaluate the effectiveness of the proposed method, two nonfrontal facial expression databases, namely BU-3DFE and Multi-PIE, are respectively used to testify our method and the experimental results show that our algorithm outperforms the state-of-the-art methods. Tong Zhang 0021, Wenming Zheng, Zhen Cui 0001, Yuan Zong, Jingwei Yan |
IEEE Trans. Multim. | 2 |
| 2015 | Color facial expression recognition based on color local featuresabstractIn this paper, color facial expression recognition based on color local features is investigated, in which each color facial image is decomposed into three color component images. For each color component image, we extract a set of color local features to represent the color component image, where color local features could be either color local binary patterns (LBP) or color scale-invariant feature transform (SIFT). To cope with the facial expression recognition problem, we use a group sparse least square regression (GSLSR) model to describe the relationship between the color local feature vectors and the associated emotion label vectors and then perform expression recognition based on it. Finally, experiments on the Multi-PIE color facial expression database are conducted to testify the proposed method and compare the results with state-of-the-art methods. Wenming Zheng, Minghai Xin |
ICASSP | 1 |
| 2015 | Cross-pose color facial expression recognition using transductive transfer linear discriminat analysisabstractIn this paper, we propose a novel transductive transfer linear discriminant analysis (TTLDA) approach for cross-pose facial expression recognition (FER), in which training and testing facial images are taken under the two different facial views. The basic idea of the proposed expression recognition method is to choose a set of auxiliary unlabelled facial images from target facial pose and leverage it into the labelled training image set of source facial pose for discriminant analysis, where the labels of the auxiliary images are parameters of TTLDA to be optimized. After learning the class labels of the auxiliary image set, we train a support vector machine (SVM) for classifying the testing facial images based on them. On the other hand, to make full utilize the facial appearance information of color images for improving expression recognition accuracy, we adopt color scale invariant feature transform (SIFT) to describe facial image feature. Finally, we conduct experiments on BU-3DFE and Multi-PIE multiview color facial expression databases to evaluate the proposed cross-pose FER method and compare the results with other methods. Wenming Zheng |
ICIP | 1 |
| 2015 | Transductive Transfer LDA with Riesz-based Volume LBP for Emotion Recognition in The WildabstractIn this paper, we propose the method using Transductive Transfer Linear Discriminant Analysis (TTLDA) and Riesz-based Volume Local Binary Patterns (RVLBP) for image based static facial expression recognition challenge of the Emotion Recognition in the Wild Challenge (EmotiW 2015). The task of this challenge is to assign facial expression labels to frames of some movies containing a face under the real word environment. In our method, we firstly employ a multi-scale image partition scheme to divide each face image into some image blocks and use RVLBP features extracted from each block to describe each facial image. Then, we adopt the TTLDA approach based on RVLBP to cope with the expression recognition task. The experiments on the testing data of SFEW 2.0 database, which is used for image based static facial expression challenge, demonstrate that our method achieves the accuracy of 50%. This result has a 10.87% improvement over the baseline provided by this challenge organizer. Yuan Zong, Wenming Zheng, Xiaohua Huang 0003, Jingwei Yan, Tong Zhang 0021 |
ICMI | 2 |
| 2015 | Dimensionality reduction for speech emotion features by multiscale kernelsabstractTo achieve efficient and compact low-dimensional features for speech emotion recognition, this paper proposes a novel feature reduction method using multiscale kernels in the framework of graph embedding.With Fisher discriminant embedding graph, multiscale Gaussian kernels are used in constructing optimal linear combination of Gram matrices for multiple kernel learning.To evaluate the proposed method, comprehensive experiments, using different public feature sets from the open-source toolbox openSMILE on various corpora, show that the proposed method achieves better performance compared with conventional linear dimensionality reduction methods and singlekernel methods. Xinzhou Xu, Wenming Zheng, Li Zhao 0003, Björn W. Schuller |
INTERSPEECH | 3 |
| 2015 | Online learning 3D context for robust visual tracking
Bineng Zhong 0001, Yingju Shen, Yan Chen 0017, Weibo Xie, Zhen Cui 0001, Hongbo Zhang 0002, Duansheng Chen, Tian Wang 0001, Xin Liu 0011, Shu-Juan Peng, Jin Gou, Jixiang Du, Jing Wang 0049, Wenming Zheng |
Neurocomputing | 14 |
| 2014 | A feature selection and feature fusion combination method for speaker-independent speech emotion recognitionabstractTo enhance the recognition rate of speaker independent speech emotion recognition, a feature selection and feature fusion combination method based on multiple kernel learning is presented. Firstly, multiple kernel learning is used to obtain sparse feature subsets. The features selected at least n times are recombined into another subset named n-subset. The optimal n is determined by 10 cross-validation experiments. Secondly, feature fusion is made at the kernel level. Not only each kind of feature is associated with a kernel, but also the full feature set is associated with a kernel which is not considered in the previous studies. All of the kernels are added together to obtain a combination kernel. The final recognition rate for 7 kinds of emotions on Berlin Database is 83.10%, which outperforms state-of-the-art results and shows the effectiveness of our method. It is also proved that MFCCs play a crucial role in speech emotion recognition. Yun Jin, Peng Song 0002, Wenming Zheng, Li Zhao 0003 |
ICASSP | 3 |
| 2014 | Robust Facial Expression Recognition Using Revised Canonical CorrelationabstractThe poor alignment and large variations in the temporal sale of facial expressions are two crucial problems for facial expression recognition (FER). Canonical correlation (CC) has recently received increasing attention in surveillance face recognition because of its robustness to variations of alignment. But it could not suit well to FER, because the facial expression variations and temporal information are ignored for canonical subspace. This paper proposes the revised canonical correlation method to address the two above described issues for making FER be robust to false detection or mis-alignment. Firstly, this paper presents the local binary pattern to describe the appearance features for enhancing the spatial variations of facial expression. Secondly, this paper proposes the temporal orthogonal locality preserved projection for building a canonical subspace of a video clip, where it mostly captures the motion changes of facial expressions. Then this paper presents the discriminative CC to model the low-dimensional feature space, which increases robustness to imprecise alignment and strengthens discrimination for facial expressions. Extensive experimental results on Extended Cohn-Kanade and MAHNOB-HCI databases demonstrate that the proposed method achieves the best results in recognizing facial expressions and performs robustly with ordinary on general face detection and eye detection. Xiaohua Huang 0003, Guoying Zhao 0001, Matti Pietikäinen, Wenming Zheng |
ICPR | 4 |
| 2014 | Text-independent voice conversion using speaker model alignment method from non-parallel speech
Peng Song 0002, Yun Jin, Wenming Zheng, Li Zhao 0003 |
INTERSPEECH | 3 |
| 2014 | Robust sparsity-preserved learning with application to image visualization
Haixian Wang, Wenming Zheng |
Knowl. Inf. Syst. | 2 |
| 2014 | A Novel Speech Emotion Recognition Method via Incomplete Sparse Least Square RegressionabstractIn this letter, we propose a novel speech emotion recognition method based on least square regression (LSR) model, in which a novel incomplete sparse LSR (ISLSR) model is proposed and utilized to characterize the linear relationship between speech features and the corresponding emotion labels. In training the ISLSR model, both labeled and unlabeled speech data sets are utilized, where the use of unlabeled data set aims to enhance the compatibility of the model such that it is well suitable for the out-of-sample speech data. Another novelty of ISLSR lies in the capability of dealing with feature selection. To evaluate the performance of the proposed method, we conduct experiments on two emotional speech databases. The experimental results on both databases demonstrate that the proposed method achieves better recognition performance in compared with several state-of-the-art methods. Wenming Zheng, Minghai Xin, Xiaolan Wang 0007 |
IEEE Signal Process. Lett. | 1 |
| 2014 | Multi-View Facial Expression Recognition Based on Group Sparse Reduced-Rank RegressionabstractIn this paper, a novel multi-view facial expression recognition method is presented. Different from most of the facial expression methods that use one view of facial feature vectors in the expression recognition, we synthesize multi-view facial feature vectors and combine them to this goal. In the facial feature extraction, we use the grids with multi-scale sizes to partition each facial image into a set of sub regions and carry out the feature extraction in each sub region. To deal with the prediction of expressions, we propose a novel group sparse reduced-rank regression (GSRRR) model to describe the relationship between the multi-view facial feature vectors and the corresponding expression class label vectors. The group sparsity of GSRRR enables us to automatically select the optimal sub regions of a face that contribute most to the expression recognition. To solve the optimization problem of GSRRR, we propose an efficient algorithm using inexact augmented Lagrangian multiplier (ALM) approach. Finally, we conduct extensive experiments on both BU-3DFE and Multi-PIE facial expression databases to evaluate the recognition performance of the proposed method. The experimental results confirm better recognition performance of the proposed method compared with the state of the art methods. Wenming Zheng |
IEEE Trans. Affect. Comput. | 1 |
| 2014 | Fisher Discriminant Analysis With L1-NormabstractFisher linear discriminant analysis (LDA) is a classical subspace learning technique of extracting discriminative features for pattern recognition problems. The formulation of the Fisher criterion is based on the L2-norm, which makes LDA prone to being affected by the presence of outliers. In this paper, we propose a new method, termed LDA-L1, by maximizing the ratio of the between-class dispersion to the within-class dispersion using the L1-norm rather than the L2-norm. LDA-L1 is robust to outliers, and is solved by an iterative algorithm proposed. The algorithm is easy to be implemented and is theoretically shown to arrive at a locally maximal point. LDA-L1 does not suffer from the problems of small sample size and rank limit as existed in the conventional LDA. Experiment results of image recognition confirm the effectiveness of the proposed method. Haixian Wang, Zilan Hu, Wenming Zheng |
IEEE Trans. Cybern. | 4 |
| 2014 | L1-Norm Kernel Discriminant Analysis Via Bayes Error Bound Optimization for Robust Feature ExtractionabstractA novel discriminant analysis criterion is derived in this paper under the theoretical framework of Bayes optimality. In contrast to the conventional Fisher's discriminant criterion, the major novelty of the proposed one is the use of L1 norm rather than L2 norm, which makes it less sensitive to the outliers. With the L1-norm discriminant criterion, we propose a new linear discriminant analysis (L1-LDA) method for linear feature extraction problem. To solve the L1-LDA optimization problem, we propose an efficient iterative algorithm, in which a novel surrogate convex function is introduced such that the optimization problem in each iteration is to simply solve a convex programming problem and a close-form solution is guaranteed to this problem. Moreover, we also generalize the L1-LDA method to deal with the nonlinear robust feature extraction problems via the use of kernel trick, and hereafter proposed the L1-norm kernel discriminant analysis (L1-KDA) method. Extensive experiments on simulated and real data sets are conducted to evaluate the effectiveness of the proposed method in comparing with the state-of-the-art methods. Wenming Zheng, Zhouchen Lin, Haixian Wang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2013 | Non-parallel training for voice conversion based on adaptation methodabstractIn this paper, we propose a simple and efficient non-parallel training scheme for voice conversion (VC). First, the speaker models are adapted from the background model using maximum a posteriori (MAP) technique. Then, by utilizing the parameters of adapted speaker models, the Gaussian normalization and mean transformation methods are proposed for VC, respectively. In addition, to improve the conversion performance of the proposed methods, a combination approach is further presented. Finally, objective and subjective experiments are carried out to evaluate the performance of the proposed scheme, the results demonstrate that our scheme can obtain comparable performance with the traditional GMM method based on parallel corpus. Peng Song 0002, Wenming Zheng, Li Zhao 0003 |
ICASSP | 2 |
| 2013 | Fuzzy two-dimensional local graph embedding discriminant analysis (F2DLGEDA) with its application to face and palm biometrics
Minghua Wan, Wenming Zheng |
Neural Comput. Appl. | 2 |
| 2013 | Complexity-reduced implementations of complete and null-space-based linear discriminant analysis
Gui-Fu Lu, Wenming Zheng |
Neural Networks | 2 |
| 2012 | Speech emotion recognition based on kernel reduced-rank regression
Wenming Zheng |
ICPR | 1 |
| 2012 | Improving CCA via spectral components selection for facial expression recognitionabstractIn this paper, we propose a novel canonical correlation analysis (CCA) algorithm for facial expression recognition. In contrast to the traditional CCA algorithm, the proposed method is capable of selecting the optimal spectral components of the training data matrix in modelling the linear correlation between the facial feature vectors and the corresponding expression class membership vectors. We formulate this spectral selection problem as a sparse optimization problem, where the ℓ1-norm penalty is adopted to this goal. To recognize the emotion category of each facial image, we present a linear regression formula to predict the emotion class membership for each facial image. The experiments on the JAFFE facial expression database confirm the better recognition performance of the proposed method. Wenming Zheng, Minghai Xin |
ISCAS | 2 |
| 2012 | A new discriminant subspace analysis approach for multi-class problems
Wenming Zheng, Zhouchen Lin |
Pattern Recognit. | 1 |
| 2012 | Towards a dynamic expression recognition system under facial occlusion
Xiaohua Huang 0003, Guoying Zhao 0001, Wenming Zheng, Matti Pietikäinen |
Pattern Recognit. Lett. | 3 |
| 2012 | Spatiotemporal Local Monogenic Binary Patterns for Facial Expression RecognitionabstractFeature representation is an important research topic in facial expression recognition from video sequences. In this letter, we propose to use spatiotemporal monogenic binary patterns to describe both appearance and motion information of the dynamic sequences. Firstly, we use monogenic signals analysis to extract the magnitude, the real picture and the imaginary picture of the orientation of each frame, since the magnitude can provide much appearance information and the orientation can provide complementary information. Secondly, the phase-quadrant encoding method and the local bit exclusive operator are utilized to encode the real and imaginary pictures from orientation in three orthogonal planes, and the local binary pattern operator is used to capture the texture and motion information from the magnitude through three orthogonal planes. Finally, both concatenation method and multiple kernel learning method are respectively exploited to handle the feature fusion. The experimental results on the Extended Cohn-Kanade and Oulu-CASIA facial expression databases demonstrate that the proposed methods perform better than the state-of-the-art methods, and are robust to illumination variations. Xiaohua Huang 0003, Guoying Zhao 0001, Wenming Zheng, Matti Pietikäinen |
IEEE Signal Process. Lett. | 3 |
| 2012 | Sparse 2-D Canonical Correlation Analysis via Low Rank Matrix Approximation for Feature ExtractionabstractAlthough 2-D canonical correlation analysis (2DCCA) has been proposed to reduce the computational complexity while reserving local data structure of image, the learned canonical variables of 2DCCA are the linear combination of all the original variables, which makes it hard to interpret the solutions and might have less generality. In this paper, we propose a sparse 2-D canonical correlation analysis (S2DCCA) to solve the drawbacks of the 2DCCA method and apply it to image feature extraction. The basic idea of S2DCCA is to impose two lasso penalties on the objective function of 2DCCA to obtain two sets of sparse projection directions via low rank matrix approximation. We conduct extensive experiments on both FERET and AR databases to evaluate the performance of the proposed method. Jingjie Yan, Wenming Zheng |
IEEE Signal Process. Lett. | 2 |
| 2010 | Dynamic Facial Expression Recognition Using Boosted Component-Based Spatiotemporal Features and Multi-classifier Fusion
Xiaohua Huang 0003, Guoying Zhao 0001, Matti Pietikäinen, Wenming Zheng |
ACIVS (2) | 4 |
| 2010 | Emotion Recognition from Arbitrary View Facial Images
Wenming Zheng, Hao Tang 0001, Zhouchen Lin, Thomas S. Huang |
ECCV (6) | 1 |
| 2010 | A rank-one update algorithm for fast solving kernel Foley-Sammon optimal discriminant vectorsabstractDiscriminant analysis plays an important role in statistical pattern recognition. A popular method is the Foley-Sammon optimal discriminant vectors (FSODVs) method, which aims to find an optimal set of discriminant vectors that maximize the Fisher discriminant criterion under the orthogonal constraint. The FSODVs method outperforms the classic Fisher linear discriminant analysis (FLDA) method in the sense that it can solve more discriminant vectors for recognition. Kernel Foley-Sammon optimal discriminant vectors (KFSODVs) is a nonlinear extension of FSODVs via the kernel trick. However, the current KFSODVs algorithm may suffer from the heavy computation problem since it involves computing the inverse of matrices when solving each discriminant vector, resulting in a cubic complexity for each discriminant vector. This is costly when the number of discriminant vectors to be computed is large. In this paper, we propose a fast algorithm for solving the KFSODVs, which is based on rank-one update (ROU) of the eigensytems. It only requires a square complexity for each discriminant vector. Moreover, we also generalize our method to efficiently solve a family of optimally constrained generalized Rayleigh quotient (OCGRQ) problems which include many existing dimensionality reduction techniques. We conduct extensive experiments on several real data sets to demonstrate the effectiveness of the proposed algorithms. Wenming Zheng, Zhouchen Lin, Xiaoou Tang |
IEEE Trans. Neural Networks | 1 |
| 2009 | A novel approach to expression recognition from non-frontal face imagesabstractNon-frontal view facial expression recognition is important in many scenarios where the frontal view face images may not be available. However, few work on this issue has been done in the past several years because of its technical challenges and the lack of appropriate databases. Recently, a 3D facial expression database (BU-3DFE database) is collected by Yin et al. [10] and has attracted some researchers to study this issue. Based on the BU-3DFE database, in this paper we propose a novel approach to expression recognition from non-frontal view facial images. The novelty of the proposed method lies in recognizing the multi-view expressions under the unified Bayes theoretical framework, where the recognition problem can be formulated as an optimization problem of minimizing an upper bound of Bayes error. We also propose a close-form solution method based on the power iteration approach and rank-one update (ROU) technique to find the optimal solutions of the proposed method. Extensive experiments on BU-3DFE database with 100 subjects and 5 yaw rotation view angles demonstrate the effectiveness of our method. Wenming Zheng, Hao Tang 0001, Zhouchen Lin, Thomas S. Huang |
ICCV | 1 |
| 2009 | Optimizing Multi-Class Spatio-Spectral Filters via Bayes Error Estimation for EEG ClassificationabstractThe method of common spatio-spectral patterns (CSSPs) is an extension of common spatial patterns (CSPs) by utilizing the technique of delay embedding to alleviate the adverse effects of noises and artifacts on the electroencephalogram (EEG) classification. Although the CSSPs method has shown to be more powerful than the CSPs method in the EEG classification, this method is only suitable for two-class EEG classification problems. In this paper, we generalize the two-class CSSPs method to multi-class cases. To this end, we first develop a novel theory of multi-class Bayes error estimation and then present the multi-class CSSPs (MCSSPs) method based on this Bayes error theoretical framework. By minimizing the estimated closed-form Bayes error, we obtain the optimal spatio-spectral filters of MCSSPs. To demonstrate the effectiveness of the proposed method, we conduct extensive experiments on the data set of BCI competition 2005. The experimental results show that our method significantly outperforms the previous multi-class CSPs (MCSPs) methods in the EEG classification. Wenming Zheng, Zhouchen Lin |
NIPS | 1 |
| 2009 | Heteroscedastic Feature Extraction for Texture ClassificationabstractLinear discriminant analysis (LDA) is a well-known feature extraction method in statistical pattern recognition community. The basic idea of LDA is to find a set of optimal discriminant vectors that maximize the Fisher's discriminant criterion. One major problem of the Fisher's criterion is that its discriminant performance largely depends on the class mean differences. Hence, the LDA method may not work well as for the case of the heteroscedastic problem since it can not make use of the discriminant information from the class covariance differences. To this end, in this paper we propose a new discriminant criterion consisting of both class mean and covariance differences to replace the Fisher's criterion. Based on the new discriminant criterion, we propose a heteroscedastic extension method of linear discriminant analysis (namely the HELDA method). We also propose an approximate solution method for HELDA (AHELDA) via matrices joint diagonalization (JD) to reduce the computational complexity. The extensive experiments on texture classification confirm the better classification performance of our method. Wenming Zheng |
IEEE Signal Process. Lett. | 1 |
| 2009 | Fast algorithm for updating the discriminant vectors of dual-space LDAabstractDual-space linear discriminant analysis (DSLDA) is a popular method for discriminant analysis. The basic idea of the DSLDA method is to divide the whole data space into two complementary subspaces, i.e., the range space of the within-class scatter matrix and its complementary space, and then solve the discriminant vectors in each subspace. Hence, the DSLDA method can take full advantage of the discriminant information of the training samples. However, from the computational point of view, the original DSLDA method may not be suitable for online training problems because of its heavy computational cost. To this end, we modify the original DSLDA method and then propose a data order independent incremental algorithm to accurately update the discriminant vectors of the DSLDA method when new samples are inserted into the training data set. We conduct experiments on the AR face database to confirm the better performance of the proposed algorithms in terms of the recognition accuracy and computational efficiency. Wenming Zheng, Xiaoou Tang |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2008 | Locality-Preserved Maximum Information ProjectionabstractDimensionality reduction is usually involved in the domains of artificial intelligence and machine learning. Linear projection of features is of particular interest for dimensionality reduction since it is simple to calculate and analytically analyze. In this paper, we propose an essentially linear projection technique, called locality-preserved maximum information projection (LPMIP), to identify the underlying manifold structure of a data set. LPMIP considers both the within-locality and the between-locality in the processing of manifold learning. Equivalently, the goal of LPMIP is to preserve the local structure while maximize the out-of-locality (global) information of the samples simultaneously. Different from principal component analysis (PCA) that aims to preserve the global information and locality-preserving projections (LPPs) that is in favor of preserving the local structure of the data set, LPMIP seeks a tradeoff between the global and local structures, which is adjusted by a parameter alpha, so as to find a subspace that detects the intrinsic manifold structure for classification tasks. Computationally, by constructing the adjacency matrix, LPMIP is formulated as an eigenvalue problem. LPMIP yields orthogonal basis functions, and completely avoids the singularity problem as it exists in LPP. Further, we develop an efficient and stable LPMIP/QR algorithm for implementing LPMIP, especially, on high-dimensional data set. Theoretical analysis shows that conventional linear projection methods such as (weighted) PCA, maximum margin criterion (MMC), linear discriminant analysis (LDA), and LPP could be derived from the LPMIP framework by setting different graph models and constraints. Extensive experiments on face, digit, and facial expression recognition show the effectiveness of the proposed LPMIP method. Haixian Wang, Sibao Chen 0001, Zilan Hu, Wenming Zheng |
IEEE Trans. Neural Networks | 4 |
| 2007 | Local and Weighted Maximum Margin Discriminant AnalysisabstractIn this paper, we propose a new approach, called local and weighted maximum margin discriminant analysis (LWMMDA), to performing object discrimination. LWMMDA is a subspace learning method that identifies the underlying nonlinear manifold for discrimination. The goal of LWMMDA is to seek a transformation such that data points of different classes are projected as far as possible while points within a same class are as compact as possible. The projections are obtained by maximizing a new discriminant criterion, called local and weighted maximum margin criterion (LWMMC). Different from previous maximum margin criterion (MMC) which seeks only the globally Euclidean structure of data points, LWMMC takes the local property into account, which makes LWMMC more accurate in finding discriminant information. LWMMC has an additional weighted parameter β that further broadens the average margin between different classes. Computationally, LWMMDA completely avoids the singularity problem. Besides, LWMMDA couples the QR-decomposition into its framework, which makes LWMMDA very efficient and stable in implementation. Finally, LWMMDA framework is straightforwardly extended into the reproducing kernel Hilbert space induced by a nonlinear function ϕ. Experiments on digit visualization, face recognition, and facial expression recognition are presented to show the effectiveness of the proposed method. Haixian Wang, Wenming Zheng, Zilan Hu, Sibao Chen 0001 |
CVPR | 2 |
| 2006 | Facial Expression Recognition Based on BoostingTree
Ning Sun 0001, Wenming Zheng, Changyin Sun 0001, Cairong Zou, Li Zhao 0003 |
ISNN (2) | 2 |
| 2006 | Gender Classification Based on Boosting Local Binary Pattern
Ning Sun 0001, Wenming Zheng, Changyin Sun 0001, Cairong Zou, Li Zhao 0003 |
ISNN (2) | 2 |
| 2006 | KDA Plus KPCA for Face Recognition
Wenming Zheng |
ISNN (2) | 1 |
| 2006 | Class-Incremental Generalized Discriminant AnalysisabstractGeneralized discriminant analysis (GDA) is the nonlinear extension of the classical linear discriminant analysis (LDA) via the kernel trick. Mathematically, GDA aims to solve a generalized eigenequation problem, which is always implemented by the use of singular value decomposition (SVD) in the previously proposed GDA algorithms. A major drawback of SVD, however, is the difficulty of designing an incremental solution for the eigenvalue problem. Moreover, there are still numerical problems of computing the eigenvalue problem of large matrices. In this article, we propose another algorithm for solving GDA as for the case of small sample size problem, which applies QR decomposition rather than SVD. A major contribution of the proposed algorithm is that it can incrementally update the discriminant vectors when new classes are inserted into the training set. The other major contribution of this article is the presentation of the modified kernel Gram-Schmidt (MKGS) orthogonalization algorithm for implementing the QR decomposition in the feature space, which is more numerically stable than the kernel Gram-Schmidt (KGS) algorithm. We conduct experiments on both simulated and real data to demonstrate the better performance of the proposed methods. Wenming Zheng |
Neural Comput. | 1 |
| 2006 | Facial expression recognition using kernel canonical correlation analysis (KCCA)abstractIn this correspondence, we address the facial expression recognition problem using kernel canonical correlation analysis (KCCA). Following the method proposed by Lyons et al. and Zhang et al., we manually locate 34 landmark points from each facial image and then convert these geometric points into a labeled graph (LG) vector using the Gabor wavelet transformation method to represent the facial features. On the other hand, for each training facial image, the semantic ratings describing the basic expressions are combined into a six-dimensional semantic expression vector. Learning the correlation between the LG vector and the semantic expression vector is performed by KCCA. According to this correlation, we estimate the associated semantic expression vector of a given test image and then perform the expression classification according to this estimated semantic expression vector. Moreover, we also propose an improved KCCA algorithm to tackle the singularity problem of the Gram matrix. The experimental results on the Japanese female facial expression database and the Ekman's "Pictures of Facial Affect" database illustrate the effectiveness of the proposed method. Wenming Zheng, Cairong Zou, Li Zhao 0003 |
IEEE Trans. Neural Networks | 1 |
| 2005 | Expression Recognition Using Elastic Graph Matching
Yujia Cao, Wenming Zheng, Li Zhao 0003, Cairong Zou |
ACII | 2 |
| 2005 | Discriminative Features Extraction in Minor Component Subspace
Wenming Zheng, Cairong Zou, Li Zhao 0003 |
ACII | 1 |
| 2005 | Weighted maximum margin discriminant analysis with kernels
Wenming Zheng, Cairong Zou, Li Zhao 0003 |
Neurocomputing | 1 |
| 2005 | An Improved Algorithm for Kernel Principal Component Analysis
Wenming Zheng, Cairong Zou, Li Zhao 0003 |
Neural Process. Lett. | 1 |
| 2005 | A note on kernel uncorrelated discriminant analysis
Wenming Zheng |
Pattern Recognit. | 1 |
| 2005 | Foley-Sammon optimal discriminant vectors using kernel approachabstractA new nonlinear feature extraction method called kernel Foley-Sammon optimal discriminant vectors (KFSODVs) is presented in this paper. This new method extends the well-known Foley-Sammon optimal discriminant vectors (FSODVs) from linear domain to a nonlinear domain via the kernel trick that has been used in support vector machine (SVM) and other commonly used kernel-based learning algorithms. The proposed method also provides an effective technique to solve the so-called small sample size (SSS) problem which exists in many classification problems such as face recognition. We give the derivation of KFSODV and conduct experiments on both simulated and real data sets to confirm that the KFSODV method is superior to the previous commonly used kernel-based learning algorithms in terms of the performance of discrimination. Wenming Zheng, Li Zhao 0003, Cairong Zou |
IEEE Trans. Neural Networks | 1 |
| 2004 | Face recognition using two novel nearest neighbor classifiersabstractIn this paper, two novel classifiers, based on locally nearest neighborhood rules, called nearest neighbor line (NNL) and nearest neighbor plane (NNP), are presented for face recognition. The underlying idea of both classifiers is the local linear combination technique that has been previously used in locally linear embedding (LLE) for nonlinear dimension reduction. In comparison to other linear combination based classifiers, such as the nearest feature line (NFL) and the nearest feature plane (NFP), the proposed method has a much lower computation cost. Furthermore, the experimental results on the ORL face database have shown that the performance of both proposed methods are competitive to the NFL and NFP in face classification. Wenming Zheng, Cairong Zou, Li Zhao 0003 |
ICASSP (5) | 1 |
| 2004 | Facial Expression Recognition Using Kernel Discriminant Plane
Wenming Zheng, Cairong Zou, Li Zhao 0003 |
ISNN (1) | 1 |
| 2004 | A Modified Algorithm for Generalized Discriminant AnalysisabstractGeneralized discriminant analysis (GDA) is an extension of the classical linear discriminant analysis (LDA) from linear domain to a nonlinear domain via the kernel trick. However, in the previous algorithm of GDA, the solutions may suffer from the degenerate eigenvalue problem (i.e., several eigenvectors with the same eigenvalue), which makes them not optimal in terms of the discriminant ability. In this letter, we propose a modified algorithm for GDA (MGDA) to solve this problem. The MGDA method aims to remove the degeneracy of GDA and find the optimal discriminant solutions, which maximize the between-class scatter in the subspace spanned by the degenerate eigenvectors of GDA. Theoretical analysis and experimental results on the ORL face database show that the MGDA method achieves better performance than the GDA method. Wenming Zheng, Li Zhao 0003, Cairong Zou |
Neural Comput. | 1 |
| 2004 | An efficient algorithm to solve the small sample size problem for LDA
Wenming Zheng, Li Zhao 0003, Cairong Zou |
Pattern Recognit. | 1 |
| 2004 | Locally nearest neighbor classifiers for pattern classification
Wenming Zheng, Li Zhao 0003, Cairong Zou |
Pattern Recognit. | 1 |