Cheng Cheng 0013

dblp:66/332-13 · DBLP profile ↗
← Back
20ranked-venue papers
11as first author
19since 2021 · last 2026
0000-0002-2138-6286ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 7 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Multi-view multi-scale adapter squeeze excitation network with domain generalization for cross-subject emotion recognition
Cheng Cheng 0013, Ziyu Jia, Weiqi He
Expert Syst. Appl.1
2026 MBDA: A modality-balanced framework with data augmentation and alignment for multimodal emotion recognition
Cheng Cheng 0013, Ruisi Shang, Huazhi Li, Ziyu Jia
Neural Networks1
2026 MS-STGAN: A dual-branch multi-scale spatio-temporal generative adversarial framework for incomplete EEG-based emotion recognition
abstract
Electroencephalography (EEG) enables high-resolution emotion recognition but often suffers from incomplete data in real-world scenarios due to sensor failures or preprocessing errors. To this end, we propose a Multi-Scale Spatio-Temporal Generative Adversarial Network (MS-STGAN). Specifically, we first apply random masking to the EEG channels data to simulate missing data conditions in practical environments. Then, we design a spatio-temporal dual-branch generator to reconstruct complete representations: the spatial branch employs graph convolutional networks (GCNs) to capture robust inter-regional dependencies, while the temporal branch leverages the BiMamba state space model to encode the dynamic evolution of emotions. To further enhance feature learning, multi-scale 2D convolution and deconvolution layers are incorporated before and after both branches, enabling the extraction of diverse spatio-temporal features. Additionally, we introduce a generative adversarial framework, where the generator restores informative features from incomplete inputs and the discriminator enforces the authenticity of reconstructed data. Finally, a fusion module integrates the outputs of both branches for downstream classification. Extensive experiments on the DEAP and SEED-IV datasets validate the effectiveness of each component and demonstrate that MS-STGAN achieves superior performance and strong generalization ability.
Cheng Cheng 0013, Yikang Cheng, Ziyu Jia, Weiqi He
Neural Networks1
2026 MSDA-Net: Multi-source Domain Adaptive Network for Multi-modal Emotion Recognition
abstract
Electroencephalogram (EEG) has shown g reat potential in multi-modal emotion recognition (MER) due to its ability to directly capture emotional states. However, the nonstationarity of EEG signals leads to significant variations across subjects and sessions, posing challenges for subject-independent MER. While previous methods have made significant progress, they often fail to integrate multimodal signals into transfer learning frameworks effectively. To address this limitation, we propose a Multi-source Domain Adaptive Network (MSDA-Net) for MER, designed to mitigate cross-subject and cross-session distribution shifts and enhance recognition performance. Specifically, we first design a feature alignment module to integrate features from different modalities, generating cross-modal feature representations and extracting representative shared features. To further improve generalization, we incorporate domain-specific feature extractors to capture domain-invariant emotional representations. Additionally, we introduce an adapter module to adjust the feature representations between different modalities, aiming to capture inter-individual differences and cross-modal correlations better. Finally, we unify classification loss, discrepancy loss, and maximum mean discrepancy (MMD) loss into a joint optimization framework. Abundant experiments on the SEED and SEED-IV datasets demonstrate the superiority of MSDA-Net, highlighting its effectiveness in improving MER performance.
Cheng Cheng 0013, Xingxing Cai, Hengrui Qi, Wenyun Chen, Yong Zhang 0030
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2025 MPFBL: Modal pairing-based cross-fusion bootstrap learning for multimodal emotion recognition
Yong Zhang 0030, Cheng Cheng 0013, Ziyu Jia
Neurocomputing4
2025 Improving vision-language alignment with graph spiking hybrid networks
Yeming Chen, Heming Zheng, Cheng Cheng 0013
Knowl. Based Syst.6
2025 SASD-MCL: Semi-supervised alignment self-distillation with mixed contrastive learning for cross-subject EEG emotion recognition
Yong Zhang 0030, Wenyun Chen, Xingxing Cai, Cheng Cheng 0013
Neural Networks4
2025 A Cross-Modal Adaptive Masked Autoencoder for Decoding Emotions With Multimodal Data
abstract
Multimodal emotion recognition (MER) has recently gained much attention since it can leverage information over multiple modalities. However, in real life, we often encounter the problem of missing modalities, as well as modeling the heterogeneity and correlation among multimodal data are challenges. To this end, we propose a unified model called cross-modal adaptive masked autoencoder (CMA-MAE) for incomplete multimodal learning. Our CMA-MAE model comprises a cross-modal adaptive fusion encoder (CMAFE) and a multiview adaptive encoder (MVAE) to capture and fuse the heterogeneity and correlation among multimodal features. Additionally, we design a convolutional decoder that progressive upsampling and fusion with the modality-invariant features to generate robust emotional features from partially observable data. To effectively utilize both data with complete and incomplete modalities for feature learning, we adopt an end-to-end approach that simultaneously optimizes classification and reconstruction tasks. Extensive testing on the DEAP and SEED-IV datasets is conducted to assess our model, with the findings demonstrating that our CMA-MAE model outperforms current leading approaches in both incomplete and complete multimodal learning scenarios.
Cheng Cheng 0013, Yong Zhang 0030, Lin Feng 0001, Ziyu Jia
IEEE Trans. Comput. Soc. Syst.1
2025 DISD-Net: A Dynamic Interactive Network With Self-Distillation for Cross-Subject Multi-Modal Emotion Recognition
abstract
Multi-modal Emotion Recognition (MER) has demonstrated competitive performance in affective computing, owing to synthesizing information from diverse modalities. However, many existing approaches still face unresolved challenges, such as: (i) how to learn compact yet representative features from multi-modal data simultaneously and (ii) how to address differences among subjects and enhance the generalization of the emotion recognition model, given the diverse nature of individual biological signals. To this end, we propose a Dynamic Interactive Network with Self-Distillation (DISD-Net) for cross-subject MER. The DISD-Net incorporates a dynamin interactive module to capture the intra- and inter-modal interactions from multi-modal data. Additionally, to enhance compactness in modal representations, we leverage the soft labels generated by the DISD-Net model as supplemental training guidance. This involves incorporating self-distillation, aiming to transfer the knowledge that the DISD-Net model contains hard and soft labels to each modality. Finally, domain adaptation (DA) is seamlessly integrated into the dynamic interactive and self-distillation components, forming a unified framework to extract subject-invariant multi-modal emotional features. Experimental results indicate that the proposed model achieves a mean accuracy of 75.00% with a standard deviation of 7.68% for the DEAP dataset and a mean accuracy of 65.65% with a standard deviation of 5.08% for the SEED-IV dataset.
Cheng Cheng 0013, Xinying Wang 0005, Lin Feng 0001, Ziyu Jia
IEEE Trans. Multim.1
2024 A novel transformer autoencoder for multi-modal emotion recognition with incomplete data
Cheng Cheng 0013, Zhaoxin Fan, Lin Feng 0001, Ziyu Jia
Neural Networks1
2024 Emotion recognition using hierarchical spatial-temporal learning transformer from regional to global brain
Cheng Cheng 0013, Lin Feng 0001, Ziyu Jia
Neural Networks1
2024 Dense Graph Convolutional With Joint Cross-Attention Network for Multimodal Emotion Recognition
abstract
Multimodal emotion recognition (MER) has attracted much attention since it can leverage consistency and complementary relationships across multiple modalities. However, previous studies mostly focused on the complementary information of multimodal signals, neglecting the consistency information of multimodal signals and the topological structure of each modality. To this end, we propose a dense graph convolution network (DGC) equipped with a joint cross attention (JCA), named DG-JCA, for MER. The main advantage of the DG-JCA model is that it simultaneously integrates the spatial topology, consistency, and complementarity of multimodal data into a unified network framework. Meanwhile, DG-JCA extends the graph convolution network (GCN) via a dense connection strategy and introduces cross attention to joint model well-learned features from multiple modalities. Specifically, we first build a topology graph for each modality and then extract neighborhood features of different modalities using DGC driven by dense connections with multiple layers. Next, JCA performs cross-attention fusion in intra- and intermodality based on each modality's characteristics while balancing the contributions of various modalities’ features. Finally, subject-dependent and subject-independent experiments on the DEAP and SEED-IV datasets are conducted to evaluate the proposed method. Abundant experimental results show that the proposed model can effectively extract and fuse multimodal features and achieve outstanding performance in comparison with some state-of-the-art approaches.
Cheng Cheng 0013, Lin Feng 0001, Ziyu Jia
IEEE Trans. Comput. Soc. Syst.1
2024 Hybrid Network Using Dynamic Graph Convolution and Temporal Self-Attention for EEG-Based Emotion Recognition
abstract
The electroencephalogram (EEG) signal has become a highly effective decoding target for emotion recognition and has garnered significant attention from researchers. Its spatial topological and time-dependent characteristics make it crucial to explore both spatial information and temporal information for accurate emotion recognition. However, existing studies often focus on either spatial or temporal aspects of EEG signals, neglecting the joint consideration of both perspectives. To this end, this article proposes a hybrid network consisting of a dynamic graph convolution (DGC) module and temporal self-attention representation (TSAR) module, which concurrently incorporates the representative knowledge of spatial topology and temporal context into the EEG emotion recognition task. Specifically, the DGC module is designed to capture the spatial functional relationships within the brain by dynamically updating the adjacency matrix during the model training process. Simultaneously, the TSAR module is introduced to emphasize more valuable time segments and extract global temporal features from EEG signals. To fully exploit the interactivity between spatial and temporal information, the hierarchical cross-attention fusion (H-CAF) module is incorporated to fuse the complementary information from spatial and temporal features. Extensive experimental results on the DEAP, SEED, and SEED-IV datasets demonstrate that the proposed method outperforms other state-of-the-art methods.
Cheng Cheng 0013, Zikang Yu, Yong Zhang 0030, Lin Feng 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Multi-Modal Network based on Spatio-temporal and Attention for Emotion Recognition
abstract
Multi-modal signals are more powerful for emotion recognition since they can provide richer emotion-related information. However, the heterogeneity and correlation of multimodal inputs in emotion recognition are difficult to explain. Additionally, how to model the temporal-spectral-spatial characteristics of electroencephalogram (EEG) signals with brain electrode dynamics and asymmetry is a challenge. To this end, we propose a multi-modal network that makes use of spatio-temporal features and attention mechanisms. The multi-scale temporal and asymmetric spatial learning module extracts EEG temporalspectral and spatial asymmetry features, while the multi-view attention-enhanced convolutional learning module captures key image features from multiple perspectives and frameworks. Moreover, a cross-modal attention fusion module is designed to integrate the extracted EEG and image features, enabling collaborative emotion recognition. Extensive ablation and comparative experiments on the DEAP and SEED-IV datasets demonstrate the superiority of the proposed model in feature learning and multimodal fusion compared to existing state-of-the-art methods.
Yong Zhang 0030, Wenyun Chen, Cheng Cheng 0013
BIBM3
2023 Action Reinforcement and Indication Module for Single-frame Temporal Action Localization
abstract
Single-frame temporal action localization aims to predict the start time, end time, and categories of action instances in untrimmed videos with only one timestamp label for each action instance. However, recent works have two challenges: incomplete location results caused by various same-class action snippets (or frames) and redundant predictions due to ambiguous background snippets. To tackle the above issues, we design a model consisting of an action reinforcement module (ARM) and an action indication module (AIM). Specifically, the ARM improves the model's generalization ability by augmenting local and global features, thus solving the challenges of incomplete predictions. Simultaneously, the AIM detects the frames that indicate the location of the action instances (called keyframes) in the training period. Then the AIM applies the detected keyframes to filter redundant prediction results in the testing period. Extensive experiments are performed on two benchmark datasets, THUMOS-14 and ActivityNet-v1.3, demonstrating that the model could achieve state-of-the-art performance compared to some competitive approaches.
Zikang Yu, Cheng Cheng 0013, Wujun Wen, Lin Feng 0001
IJCNN2
2023 SS-INR: Spatial-Spectral Implicit Neural Representation Network for Hyperspectral and Multispectral Image Fusion
abstract
Due to the limitation of imaging equipment, it is difficult to acquire hyperspectral images with high spatial resolution directly. Existing approaches improve the resolution of HSIs by fusing multispectral image (MSI) and hyperspectral image (HSI). However, most of them are only feed-forward. They only learn low- to high-resolution feature mappings without considering the ill-posedness of super-resolution tasks, leading to a large solution space of mapping functions and making it difficult to learn a complete mapping function. Moreover, there is a large resolution difference between HSI and MSI, and some up-sampling operations are inevitably employed in the network. Nevertheless, traditional upsampling methods only represent pixel points in a discrete way, failing to adequately restore the continuous spatial and spectral information. To this end, this paper proposes a spatial-spectral implicit neural representation network for hyperspectral and multispectral image fusion (SS-INR). Inspired by the success of implicit neural representation(INR) in continuum reconstruction, we design spatial-INR and spectral-INR for spatial and spectral resolution reconstruction, respectively. SS-INR contains two processes: forward fusion (FF) and back-projection fusion(BPF). In the FF process, the input HSI is first spatially upsampled with Spatial-INR to overcome spatial resolution differences while performing initial fusion with MSI. In the BPF process, we explore the spatial and spectral degradation processes and use them as prior knowledge for error correction. Extensive experiments on five public hyperspectral datasets demonstrate the effectiveness of SS-INR, and SS-INR achieves competitive results compared with existing state-of-the-art fusion methods. The source code for SS-INR will be released at https://github.com/wxy11-27/SS-INR.
Xinying Wang 0005, Cheng Cheng 0013, Shenglan Liu 0001, Ruoxi Song, Xiang-Hai Wang 0001, Lin Feng 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 Multi-Domain Encoding of Spatiotemporal Dynamics in EEG for Emotion Recognition
abstract
The common goal of the studies is to map any emotional states encoded from electroencephalogram (EEG) into 2-dimensional arousal-valance scores. It is still challenging due to each emotion having its specific spatial structure and dynamic dependence over the distinct time segments among EEG signals. This paper aims to model human dynamic emotional behavior by considering the location connectivity and context dependency of brain electrodes. Thus, we designed a hybrid EEG modeling method that mainly adopts the attention mechanism, combining a multi-domain spatial transformer (MST) module and a dynamic temporal transformer (DTT) module, named MSDTTs. Specifically, the MST module extracts single-domain and cross-domain features from different brain regions and fuses them into multi-domain spatial features. Meanwhile, the temporal dynamic excitation (TDE) is inserted into the multi-head convolutional transformer to form the DTT module. These two blocks work together to activate and extract the emotion-related dynamic temporal features within the DTT module. Furthermore, we place the convolutional mapping into the transformer structure to mine the static context features among the keyframes. Overall results show that high classification accuracy of 98.91%/0.14% was obtained by the $\beta$ frequency band of the DEAP dataset, and 97.52%/0.12% and 96.70%/0.26% were obtained by the $\gamma$ frequency band of SEED and SEED-IV datasets. Empirical experiments indicate that our proposed method can achieve remarkable results in comparison with state-of-the-art algorithms.
Cheng Cheng 0013, Yong Zhang 0030, Luyao Liu 0001, Lin Feng 0001
IEEE J. Biomed. Health Informatics1
2022 Multimodal emotion recognition based on manifold learning and convolution neural network
Yong Zhang 0030, Cheng Cheng 0013, YiDie Zhang
Multim. Tools Appl.2
2022 EEG-Based Emotion Recognition Using Spatial-Temporal Graph Convolutional LSTM With Attention Mechanism
abstract
The dynamic uncertain relationship among each brain region is a necessary factor that limits EEG-based emotion recognition. It is a thought-provoking problem to availably employ time-varying spatial and temporal characteristics from multi-channel electroencephalogram (EEG) signals. Although deep learning has made remarkable achievements in emotion recognition, the biological topological information among brain regions does not fully exploit, which is vital for EEG-based emotion recognition. In response to this problem, we design a hybrid model called ST-GCLSTM, which comprises a spatial-graph convolutional network (SGCN) module and an attention-enhanced bi-directional Long Short-Term Memory (LSTM) module. The main advantage of ST-GCLSTM is that it can consider the biological topology information of each brain region to extract representative spatial-temporal features from multiple EEG channels. Specifically, we construct two layers SGCN by introducing adjacency matrices to adaptively learn the intrinsic connection among different EEG channels. Moreover, an attention-enhanced mechanism is placed into a bi-directional LSTM module to extract the crucial spatial-temporal features from sequential EEG data, and then these features serve as the input layer of the classifier to learn discriminative emotion-related features. Extensive experiments on the DEAP, SEED, and SEED-IV datasets demonstrate the effectiveness of the proposed ST-GCLSTM model, revealing that our model had an absolute performance improvement over state-of-the-art strategies.
Lin Feng 0001, Cheng Cheng 0013, Mingyan Zhao, Huiyuan Deng, Yong Zhang 0030
IEEE J. Biomed. Health Informatics2
2019 Multi-Channel Physiological Signal Emotion Recognition Based on ReliefF Feature Selection
abstract
Emotion recognition plays a very important role nowadays. Emotional recognition of electroencephalogram (EEG) signals involves high-dimensional EEG data, which is still a challenging task now. This paper presents an emotion recognition framework based on EEG and peripheral signals, and the ReliefF algorithm is used to select features of EEG and peripheral signals. The DEAP data set is used to process multi-channel physiology signals. In this paper the human emotion is classified into two categories (happy, unhappy) and three categories (happy, neutral, unhappy), respectively. The performance of the proposed method is evaluated in conjunction with ReliefF algorithm. The best average accuracy of K-Nearest Neighbor (KNN) and random forest (RF) are 94.295% and 98.832% in the binary classification tasks, and 79.210% and 98.364% in the triple classification tasks, respectively. At the same time, we also have adopted F1-score as the metric in the binary classification, which are 93.944% and 98.835%. The evaluation on the DEAP data set shows that using the ReliefF algorithm for feature selection can achieve good results on the issue of physiological emotion recognition.
Yong Zhang 0030, Cheng Cheng 0013, Tianzhen Chen
ICPADS2