EDBT 2026 Demo / reviewers in the wild / expert
Wei-Long Zheng
dblp:150/4150
· DBLP profile ↗
78ranked-venue papers
7as first author
47since 2021 · last 2026
0000-0002-9474-6369ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 52 · 4 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 2 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MindCross: Fast New Subject Adaptation with Limited Data for Cross-subject Video Reconstruction from Brain SignalsabstractBrain decoding aims to reconstruct video from brain signals. Existing brain decoding frameworks are primarily built on a subject-dependent paradigm, which requires large amounts of brain data for each subject. However, the expensive cost of collecting brain-video data causes severe data scarcity for brain decoding. Although some cross-subject methods being introduced, they often exhibit an excessive preoccupation with subject-invariant information while neglecting subject-specific information, resulting in slow fine-tune-based adaptation strategy. To achieve fast and data-efficient new subject adaptation, we propose **MindCross**, a novel cross-subject brain decoding framework. MindCross's *N* specific encoders and one shared encoder are designed to extract subject-specific and subject-invariant information, respectively. Additionally, a Top-*K* collaboration module is adopted to enhance new subject decoding with the knowledge learned from previous subjects' encoders. Extensive experiments on fMRI/EEG-to-video benchmarks demonstrate MindCross's efficacy and efficiency of cross-subject decoding and new subject adaptation using only one model. Code of our framework will be released upon publication. Xuan-Hao Liu, Yan-Kai Liu, Bao-Liang Lu, Wei-Long Zheng |
AAAI | 5 |
| 2026 | A Multimodal EEG-Eye Movement Model for Automatic Depression DetectionabstractDepression is a prevalent mental health disorder characterized by persistent sadness and a diminished interest in daily activities, early detection of depression facilitates timely intervention, mitigating its adverse effects. Electroencephalography (EEG) signals and eye movements are emerging as promising biomarkers for depression detection due to their non-invasive nature and cost-effectiveness. Nevertheless, existing studies suffer from methodological constraints, including low specificity, insufficient sample sizes, limited generalizability, and difficulties in large-scale replication, which collectively undermine their clinical utility. To address these challenges, we collected a large-scale depression dataset comprising EEG and eye movements from 1,060 individuals diagnosed with depression and 1,308 healthy controls. To efficiently leverage multimodal data for automatic depression detection, we propose the EEG-Eye Movements Model (E2Mo). E2Mo employs modality-specific encoders to extract discriminative multi-view features from each modality and incorporates a mixture-of-modality-experts architecture with multi pretraining tasks to achieve efficient and robust modality alignment and fusion. Our approach achieves a 70.06% balanced accuracy by leveraging multi-modal data, demonstrating the effectiveness of integrating EEG signals and eye movements for automatic depression detection. Hao-Long Yin, Ren-Jie Dai, Wei-Long Zheng, Qinyu Lv, Zhenghui Yi, Bao-Liang Lu |
AAAI | 4 |
| 2026 | Gram: A Large General EEG Model for Raw Data Classification and RestorationabstractDrawing insights from Large Language Models, researchers have developed several Large Electroencephalogram (EEG) models (LEMs) to learn a generalized representation adaptable to various tasks. However, such LEMs are scarce and neglecting the potential in data restoration tasks. Meanwhile, how to efficiently integrate temporal view and spectral view of EEG data has always been a focal point. In this paper, we propose Gram, a large general EEG model for raw EEG data classification and restoration tasks. Gram consists of two stages. 1) The initial stage quantizes raw EEG patches into base classes rich in temporal information. 2) The second stage features a multi-view layer-fusion masked autoencoder that exploits EEG's complex Temporal and Spectral views through dual training objectives: a spectral mimic target after layer-fusion encoder for visible patches and a base-class classification target after decoder for masked patches. Pretrained on 7000 hours of EEG data, Gram achieves state-of-the-art performance on four cross-subject classification tasks including motor imagery as well as event, emotion, and sleep stage classification. For EEG data restoration, our model significantly improves classification performance by repairing corrupted data in comparison to using noisy data. Ziyi Li 0003, Wei-Long Zheng, Jiwen Xu, Bao-Liang Lu |
IEEE Trans. Affect. Comput. | 2 |
| 2026 | SEED-OLF: A Novel EEG Dataset With Olfactory Stimulation for Emotion Recognition
Wei-Bang Jiang, Wei-Long Zheng, Bao-Liang Lu |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | Multi-to-Single: Reducing Multimodal Dependency in Emotion Recognition Through Contrastive LearningabstractMultimodal emotion recognition is a crucial research area in the field of affective brain-computer interfaces. However, in practical applications, it is often challenging to obtain all modalities simultaneously. To deal with this problem, researchers focus on using cross-modal methods to learn multimodal representations with fewer modalities. However, due to the significant differences in the distribution of different modalities, it is challenging to enable any modality to fully learn multimodal features. To address this limitation, we propose a Multi-to-Single (M2S) emotion recognition model, leveraging contrastive learning and incorporating two innovative modules: 1) a spatial and temporal-sparse (STS) attention mechanism that enhances the encoders' ability to extract features from data; 2) a novel Multi-to-Multi Contrastive Predictive Coding (M2M CPC) that learns and fuses features across different modalities. In the final testing, we only use a single modality for emotion recognition, reducing the dependence on multimodal data. Extensive experiments on five public multimodal emotion datasets demonstrate that our model achieves the state-of-the-art performance in the cross-modal tasks and maintains multimodal performance using only a single modality. Yan-Kai Liu, Jinyu Cai, Bao-Liang Lu, Wei-Long Zheng |
AAAI | 4 |
| 2025 | Decoding Emotions from Missing-Channel EEG Signals via Uncertainty-Aware Modeling with Neural ProcessesabstractElectroencephalography (EEG) signals have been widely utilized in affective brain-computer interface (BCI) applications. However, EEG signals are inherently susceptible to both external environmental noise and internal subject-specific variability, resulting in significant uncertainty. Moreover, in real-world scenarios, excessive external interference may severely degrade the quality of EEG recordings, leading to missing or corrupted channels. Therefore, developing robust emotion decoding methods that can handle incomplete EEG data is critical for enhancing the generalizability and reliability of affective BCIs. In this paper, we propose Fewer-Channel EEG Neural Processes (FC-EEGNP), a novel neural process-based framework designed to learn the uncertain relationships between EEG signals and brain regions, enabling emotion recognition under missing-channel conditions. FC-EEGNP reconstructs high spatial resolution EEG representations from a limited subset of available channels, preserving critical emotional information despite data sparsity. We conduct extensive experiments on three publicly available EEG datasets to evaluate the performance of FC-EEGNP. Results demonstrate that our model consistently achieves state-of-the-art performance across multiple emotion recognition settings, including subject-dependent and cross-subject tasks. Yan-Kai Liu, Xuan-Hao Liu, Yi-Dong Zhao, Bao-Liang Lu, Wei-Long Zheng |
BIBM | 5 |
| 2025 | Attention-Based Graph Net with Mixture of Experts for Emotion Recognition Under Sleep DeprivationabstractThe impact of sleep deprivation on human emotional states has long been a central topic in psychology and neuroscience. However, the lack of high-quality, randomized stimuli materials for eliciting multimodal emotional responses has limited the development of robust datasets in this area. To address this gap, we collected a multimodal dataset comprising four emotional categories-happy, sad, fear, and neutral-from 20 subjects under three sleep conditions: sleep deprivation, sleep recovery, and normal sleep. Each emotional state was elicited using randomized video stimuli specifically designed to provoke the target emotions. Furthermore, we propose a novel emotion classification framework that integrates a Graph Neural Network (GNN) with a MoE (Mixture-of-Experts) architecture. Experimental results demonstrate that our method achieves SOTA performance. Moreover, in all models evaluated, the accuracy of emotional recognition was significantly lower under sleep deprivation compared to sleep recovery and baseline conditions, particularly for positive emotional states. Shi-Heng Tian, Yan-Kai Liu, Bao-Liang Lu, Wei-Long Zheng |
BIBM | 4 |
| 2025 | Self-supervised EEG Representation Learning based on Temporal Prediction and Spatial Reconstruction for Emotion Recognition
Ren-Jie Dai, Keya Hu, Hao-Long Yin, Bao-Liang Lu, Wei-Long Zheng |
CogSci | 5 |
| 2025 | mixEEG: Enhancing EEG Federated Learning for Cross-subject EEG Classification with Tailored mixup
Xuan-Hao Liu, Bao-Liang Lu, Wei-Long Zheng |
CogSci | 3 |
| 2025 | Gram: A Large-Scale General EEG Model for Raw Data Classification and Restoration TasksabstractDrawing insights from Large Language Models, researchers have developed several large-scale Electroencephalogram (EEG) models (LEMs) to learn a generalized representation adaptable to various tasks. However, such LEMs are scarce and neglecting the potential in reconstruction tasks. Meanwhile, how to efficiently integrate temporal view and spectral view of EEG data has always been a focal point. In this paper, we propose Gram, a large general EEG model for raw EEG data classification and reconstruction tasks. Gram consists of two stages. 1) The initial stage quantizes raw EEG patches into base classes rich in temporal information. 2) The second stage features a multi-view layer-fusion masked autoencoder that exploits EEG’s complex Temporal and Spectral views through dual training objectives: a spectral mimic target after layer-fusion encoder for visible patches and a base-class classification target after decoder for masked patches. Pretrained on 7000 hours of EEG data, Gram achieves SOTA performance on 3 cross-subject classification tasks including event, emotion, and sleep stage classification. For EEG data restoration, our model significantly improves classification performance by repairing corrupted data in comparison to using noisy data. The code and pretrained weights are in https://github.com/iiieeeve/Gram. Ziyi Li 0003, Wei-Long Zheng, Bao-Liang Lu |
ICASSP | 2 |
| 2025 | Multi-Source Multi-Target Domain Similarity Network for Cross-Cultural EEG Emotion RecognitionabstractThe significant variations in emotional patterns across different cultures pose a major challenge for cross-cultural electroencephalogram (EEG) emotion recognition. Moreover, this task must address not only differences in feature distributions among different cultures but also among individuals within the same culture. Therefore, we propose a novel domain adaptation approach, the Multi-Source Multi-Target Domain Similarity Network (MSMTDS), which treats each subject as an individual source domain or target domain. Based on the assessment of similarities across both cross-cultural and intra-cultural dimensions, the model dynamically adjusts the weights assigned to each domain, allowing similar domains to play a dominant role in training while reducing the adverse impact of dissimilar domains. Additionally, given the scarcity of EEG data, MSMTDS fully leverages all available data to maximize performance. Extensive experiments on the SEED EEG emotion datasets from three distinct cultures (China, France, and Germany) demonstrate the effectiveness of our approach, achieving state-of-the-art results. Hanwen Shi, Bao-Liang Lu, Wei-Long Zheng |
ICASSP | 4 |
| 2025 | Double Domain Converter Transformer For Improving EEG-Based Emotion Recognition from Video to Game ScenariosabstractEmotion recognition (ER) plays an important role in the field of modern technology and human-computer interaction. Traditional emotion recognition approaches usually utilize videos as stimuli. However, watching video lacks interaction. Recently, more and more game-related stimuli have been used. To explore emotion recognition in a more natural state, this study pioneers the use of Genshin Impact game and gameplay-related videos to evoke positive and neutral emotions. As Electroencephalogram (EEG) feature distributions of video and game scenarios are very different, we propose the Double Domain Converter Transformer network (DDCT) to enhance EEG-based cross-scenario emotion recognition, making full use of information from different domains. Our model has two converters that can convert input data into the source domain and the target domain, diminishing distribution difference in two feature spaces. Experimental results demonstrate that our model achieves a remarkable prediction accuracy of 84.95% in video-game mix-scenario emotion recognition and 74.76% in video to game cross-scenario emotion recognition. Our code is available at https://github.com/AlbertMentat/DDCT. Jun-Yu Pan, Hao-Long Yin, Wei-Long Zheng |
ICASSP | 3 |
| 2025 | STAR: A Spatial-Temporal Autoencoder for EEG Restoration in Emotion RecognitionabstractResearch in emotion recognition using electroencephalography (EEG) has advanced rapidly, and affective EEG-based Brain-computer Interface (aBCI) technology is increasingly moving from lab research to real-world application. Nevertheless, EEG signals are inherently delicate and prone to noise and artifacts, especially in real-world environments where data quality often lags behind laboratory standards. This disparity poses substantial challenges for models trained on high-quality datasets. Conventional methods, such as data interpolation or exclusion, limit model efficacy. To overcome these challenges, we introduce the Spatial-Temporal Autoencoder for EEG Restoration (STAR). STAR leverages dynamic channel and temporal masking to mimic real-world signal degradation and incorporates a spatial-temporal alternating attention mechanism to encapsulate intricate spatiotemporal dynamics within EEG data. Our evaluations on three premium emotion recognition EEG datasets reveal that STAR effectively restores signals across varying corruption levels, significantly bolstering the performance of emotion recognition models in suboptimal conditions. Hao-Long Yin, Wei-Long Zheng, Bao-Liang Lu |
ICASSP | 2 |
| 2025 | Multi-Scale Attention-Based Dense Spatial-Temporal Model for Emotion Induction in Response to Olfactory StimuliabstractAffective Brain-Computer Interfaces (aBCIs) have attracted growing attention due to their potential for decoding human emotional states through electroencephalogram (EEG) signals. However, existing deep learning models often struggle to fully capture both the spatial and temporal dependencies in EEG data, resulting in suboptimal performance in emotion classification. To address these challenges, we propose a novel Multi-Scale Attention-Based Dense Spatial-Temporal Model (MSADM). This model leverages temporal attention, multi-scale dense feature extraction, and attention-based feature fusion to effectively capture and enhance the complicated spatial-temporal dependencies within EEG data, thereby improving the representations. Furthermore, we introduce a new olfactory-based emotion induction paradigm, which effectively mitigates the limitations of traditional visual and auditory stimuli by providing more stable and sustained emotional responses. We also present a novel EEG dataset involving 32 subjects, developed through a pilot study to select appropriate olfactory stimuli. Experimental results demonstrate that our model significantly outperforms existing methods across multiple metrics, demonstrating the effectiveness of both the proposed model and emotion induction paradigm. This study underscores the potential of olfactory stimuli and advanced spatial-temporal modeling techniques for enhancing the robustness and performance of emotion recognition. Wei-Bang Jiang, Wei-Long Zheng, Bao-Liang Lu |
ICASSP | 3 |
| 2025 | EEGMirror: Leveraging EEG Data in the Wild Via Montage-Agnostic Self-Supervision for EEG to Video Decoding
Xuan-Hao Liu, Bao-Liang Lu, Wei-Long Zheng |
ICCV | 3 |
| 2025 | Combining Induction and Transduction for Abstract ReasoningabstractWhen learning an input-output mapping from very few examples, is it better to first infer a latent function that explains the examples, or is it better to directly predict new test outputs, e.g. using a neural network? We study this question on ARC by training neural models for \emph{induction} (inferring latent functions) and \emph{transduction} (directly predicting the test output for a given test input). We train
on synthetically generated variations of Python programs that solve ARC training tasks. We find inductive and transductive models solve different kinds of test problems, despite having the same training problems and sharing the same neural architecture: Inductive program synthesis excels at precise computations, and at composing multiple concepts, while transduction succeeds on fuzzier perceptual concepts. Ensembling them approaches human-level performance on ARC. Wen-Ding Li, Keya Hu, Carter Larsen, Yuqing Wu, Simon Alford, Caleb Woo, Spencer M. Dunn, Hao Tang 0008, Wei-Long Zheng, Yewen Pu, Kevin Ellis |
ICLR | 9 |
| 2025 | Hierarchical Emotion Transformer for Multimodal Joint Emotion Category and Intensity Recognition
Tian-Fang Ma, Wei-Bang Jiang, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (1) | 3 |
| 2025 | Multi-session Meditative EEG Classification with Noise-Robust Self-supervision
Zuxin Song, Mingyu Gou, Tianzhen Chen, Bao-Liang Lu, Wei-Long Zheng |
ICONIP (2) | 7 |
| 2025 | Multimodal Emotion Recognition with Missing Modality via a Unified Multi-task Pre-training FrameworkabstractMultimodal emotion recognition based on physiological signals faces the challenge of missing modality due to issues such as inaccurate signal synchronization and inadequate device contact. Existing methods either require additional generative modules to handle missing modalities, leading to extra computational overhead, or fail to effectively capture both modality-specific and joint representation. In contrast, we propose a Unified Multi-task Pre-training (UMAP) framework based on the mixture of experts structure. Our approach offers two key advantages: (1) It retains a joint structure while flexibly handling both unimodal and multimodal inputs by selecting lightweight modality experts and shared experts, thus preserving both modality-specific features and joint multimodal information. 2) Three pre-training tasks-contrastive learning, modality matching, and modality generation-are integrated into UMAP through different attention masks, enhancing the model's ability to adapt to both complete and incomplete modalities during fine-tuning. Comprehensive experiments conducted on three benchmark datasets demonstrate that UMAP achieves state-of-the-art (SOTA) performance, both in multimodal scenarios and in cases where any modality is missing. The code is available at https://github.com/iiieeeve/UMAP. Ziyi Li 0003, Wei-Long Zheng, Bao-Liang Lu |
ACM Multimedia | 2 |
| 2025 | Human vs AI: How Digital Human News Anchors Affect Our Cognitive Processes?abstractWith the advancement of Artificial Intelligence Generated Content (AIGC) technology, digital human representations are increasingly appearing in multimedia interactions. This trend is particularly prominent in news broadcasting. The uniformity of news anchors' appearances and broadcasting environments has facilitated the widespread adoption of AI-powered news anchors. With high accuracy in news reporting and advancements in technology, AI anchors have been increasingly implemented in various news programs. However, there is currently a lack of objective analysis regarding the cognitive impact of digital human news broadcasting on audiences and its corresponding effects on brain signals. In this work, we investigate the differences in electroencephalography (EEG) responses when subjects watch news broadcasts delivered by digital humans versus real human anchors under various conditions. Our contributions are threefold: 1) We develop a dataset recording EEG signals from 32 subjects while they were watching news broadcasts. According to the presentation format (human/AI anchor) and the level of attention (high/normal), we categorize the dataset into four groups. 2) We utilize EEG signals to analyze the perceptual differences of subjects when watching news presented in different formats and in varying attention states. In addition, we investigate the cognitive differences of the subjects in perceived authenticity and importance of the news under these different conditions. 3) We propose an asymmetric multi-representation learning framework to better utilize and analyze the data. The code and data are available at https://github.com/Arcee-LYK/EEG-News. Yan-Kai Liu, Shunyang Yao, Bao-Liang Lu, Wei-Long Zheng |
ACM Multimedia | 5 |
| 2025 | SEED-MYA: A Novel Myanmar Multimodal Dataset for Enhancing Emotion RecognitionabstractThis paper introduces a novel Myanmar multimodal dataset called SEED-MYA, which is the first culturally and linguistically tailored multimodal emotion recognition dataset for the Burmese speakers. The SEED-MYA dataset consists of EEG and eye movement data collected using Myanmar video stimuli, addressing the underrepresentation of minority cultures in emotion recognition research. To investigate the fundamental characteristics of emotion recognition based on EEG and eye movement data from Myanmar participants, and to validate the quality and effectiveness of the SEED-MYA dataset, we implement the Multimodal Adaptive Emotion Transformer with Cross-Modal Attention (MAET-CMA) as a benchmarking tool. From our experiments on SEED-MYA, we have three main findings: (a) beta and gamma bands play a critical role in distinguishing positive, neutral, and negative emotional states; (b) combining EEG and eye movement data significantly enhances emotion recognition accuracy, with MAET-CMA achieving a maximum accuracy of 91.78%; and (c) EEG signals excel in recognizing negative and positive emotional states, while eye movement data are particularly effective at differentiating neutral emotion in our experimental setup. Our neural activity analysis reveals distinct patterns of activation in temporal, parietal, and prefrontal regions, providing insights into potential culture-related neural responses. We further compare our findings with established benchmarks from Chinese participants (SEED dataset) to explore cultural similarities and differences in emotion recognition. This analysis is structured into three components: band-level EEG comparisons, modality-specific performance analysis, and neural activity differences, providing a multi-level analysis of cultural effects. While some similarities in emotion recognition are observed across two cultures, our cross-cultural performance comparison between SEED and SEED-MYA further indicates that the Chinese dataset generalizes better as a test set. These findings underscore the importance of incorporating culturally diverse datasets in the development of globally applicable emotion recognition systems. Khin Pa Pa Aung, Hao-Long Yin, Tian-Fang Ma, Wei-Long Zheng, Bao-Liang Lu |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | SEED-VII: A Multimodal Dataset of Six Basic Emotions With Continuous Labels for Emotion RecognitionabstractRecognizing emotions from physiological signals is a topic that has garnered widespread interest, and research continues to develop novel techniques for perceiving emotions. However, the emergence of deep learning has highlighted the need for comprehensive and high-quality emotional datasets that enable the accurate decoding of human emotions. To systematically explore human emotions, we develop a multimodal dataset consisting of six basic (happiness, sadness, fear, disgust, surprise, and anger) emotions and the neutral emotion, named SEED-VII. This multimodal dataset includes electroencephalography (EEG) and eye movement signals. The seven emotions in SEED-VII are elicited by 80 different videos and fully investigated with continuous labels that indicate the intensity levels of the corresponding emotions. Additionally, we propose a novel Multimodal Adaptive Emotion Transformer (MAET), that can flexibly process both unimodal and multimodal inputs. Adversarial training is utilized in the MAET to mitigate subject discrepancies, which enhances domain generalization. Our extensive experiments, encompassing both subject-dependent and cross-subject conditions, demonstrate the superior performance of the MAET in terms of handling various inputs. Continuous labels are used to filter the data with high emotional intensity, and this strategy is proven to be effective for attaining improved emotion recognition performance. Furthermore, complementary properties between the EEG signals and eye movements and stable neural patterns of the seven emotions are observed. Wei-Bang Jiang, Xuan-Hao Liu, Wei-Long Zheng, Bao-Liang Lu |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | Investigating the Effects of Sleep Conditions on Emotion Responses with EEG Signals and Eye MovementsabstractExisting studies in psychology and neuroscience have extensively examined the effects of sleep deprivation on emotional responses. More recently, researchers have begun applying deep learning algorithms to further investigate this relationship, emphasizing the importance of accessible and high-quality multimodal datasets across different sleep states. To address this need, we develop SEED-SD, a multimodal dataset comprising data from 40 participants. The dataset includes electroencephalography (EEG) and eye movement signals collected under three sleep conditions: sleep deprivation (SD), sleep recovery (SR), and normal sleep (NS). Each condition contains data corresponding to four basic emotions: happiness, sadness, fear, and neutral state. Additionally, we propose a novelRegion Transformer withLayer-Fusion (ReLF), to conduct comprehensive analyses on the SEED-SD dataset. ReLF incorporates a region- wise self-attention mechanism to extract localized features from EEG and eye movement signals, and supports flexible adaptation to both multimodal and unimodal inputs. Following multimodal generative pre-training, ReLF introduces learnable prompts to replace the missing modality under unimodal settings, thereby enabling effective fine-tuning of the pre-trained model. The experimental results demonstrate that ReLF outperforms the existing models. Our analysis further reveals that SD significantly impacts emotion recognition performance, while SR and NS conditions yield similar results, highlighting the importance of SR in mitigating the adverse effects of SD. Furthermore, we conduct a systematic analysis of multimodal complementarity, critical frequency bands, and neural patterns. Our findings reveal distinct EEG patterns under the SD condition compared to the SR and NS conditions. Notably, the multimodal complementarity and critical frequency bands in both the SD and SR conditions align with those observed in the NS condition. In summary, to the best of our knowledge, SEED-SD is the largest publicly available multimodal dataset for studying the relationship between emotion recognition and sleep states. This dataset lays a crucial foundation for applying deep learning methods in this area. Moreover, through extensive data-driven analysis, this work confirms the inhibitory effect of SD on emotion recognition and the restorative role of SR. The SEED-SD dataset and codes will be public upon paper acceptance. Ziyi Li 0003, Le-Yan Tao, Rui-Xiao Ma, Wei-Long Zheng, Bao-Liang Lu |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | From EEG to Eye Movements: Cross-Modal Emotion Recognition Using Constrained Adversarial Network With Dual AttentionabstractEmotion recognition is a fundamental part of affective computing, obtaining performance gain from multimodal methods. Electroencephalography (EEG) and eye movements are extensively used as they contain complementary information. However, the inconvenient acquisition of EEG is hindering the extensive adoption of multimodal emotion recognition in daily applications while eye movements are more convenient to collect but with lower performance. To tackle this issue, we propose a Constrained Adversarial Network with Dual Attention (CANDA), exploiting the complementary information from multiple modalities during training to improve the test-time performance of single easily acquired modality, i.e., transferring knowledge from a stronger modality to a weaker modality. During training, a common joint space is learned to diminish the distribution discrepancy among different modalities and incorporate the multimodal representations. During test, single modality is converted to the common space achieving comparable performance to multiple modalities. Extensive experiments demonstrate that our model achieves the state-of-the-art performance for cross-modal emotion recognition. Specifically, the mean accuracy increases around 15% on SEED, 15% on SEED-IV, and 2% on SEED-V compared to the latest baseline for emotion recognition. Visualization of features in the joint space illustrates that the distribution of different modalities aligns together with the discriminative ability regarding various emotions. Jia-Wen Liu, Bao-Liang Lu, Wei-Long Zheng |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | Multi-View Self-Supervised Domain Adaptation for EEG-Based Emotion RecognitionabstractResearch on emotion recognition based on EEG signals has made significant progress. Most of the existing studies have focused on supervised learning methods, but real-life data cannot meet the requirement of high quality with labels. In addition, EEG signals have individual variability and instability, which requires transfer learning to enhance the model generalization. In this paper, we propose a multi-view self-supervised domain adaptation model that combines self-supervised learning techniques with domain-adaptive transfer learning algorithms, which can solve the last two problems mentioned above. Specifically, we add a multi-class domain discriminator to construct the adversarial relationship between the sub-networks so that distribution discrepancy of different subjects can be reduced effectively. We conduct both subject-dependent and subject-independent experiments on the SEED and SEED-IV datasets to thoroughly evaluate the performance of our model. The results show that our model achieves outstanding emotion recognition performance even with limited labeled data. In the subject-dependent experiments on both datasets, our model achieves accuracy rates of 85.91% and 87.19% respectively, surpassing the original self-supervised masked autoencoder model by about 3%. In subject-independent experiments, our model demonstrates strong data distribution adaptation capabilities, achieving an accuracy of 69.72% and 62.87%, respectively on the SEED and SEED-IV datasets using only 90 samples for subject-independent experiments. This effectively mitigates the accuracy degradation caused by differences in data distribution across subjects. Furthermore, our model is capable of extracting meaningful features from corrupted EEG data, highlighting its robustness and effectiveness. Lu Zhang 0096, Hanwen Shi, Juan-Zi Li, Wei-Long Zheng, Bao-Liang Lu |
IEEE Trans. Affect. Comput. | 4 |
| 2024 | Addressing Temporal and Auditory Factors in Meditative EEG with Self-Supervised LearningabstractMeditation has been shown to improve mental health by reducing stress and enhancing cognitive functions. To monitor meditation quality, Electroencephalogram (EEG) offers an objective alternative to self-evaluation questionnaires, which suffer from subjectivity and delays. However, previous EEG protocols often inflated classification performance due to temporal correlation and auditory interference. This study introduces a new data acquisition protocol to address these issues. We constructed a dataset from 32 subjects, each participating on three sessions in different days to eliminate temporal interference. During each session, subjects alternated between audio-guided meditation and silent unguided meditation. Using the Multi-view Spectral-Spatial-Temporal Masked Autoencoder (MV-SSTMA) model which was pre-trained on the Emotion EEG dataset SEED, the model achieved superior classification performance in F1 score, accuracy, precision, and recall. These findings offer valuable insights for improving EEG-based meditation monitoring in biomedical applications. Mingyu Gou, Ying-Jie Zhang, Ren-Jie Dai, Hao-Long Yin, Tianzhen Chen, Bao-Liang Lu, Wei-Long Zheng |
BIBM | 9 |
| 2024 | MoGE: Mixture of Graph Experts for Cross-subject Emotion Recognition via Decomposing EEGabstractDecoding emotions of previously unseen subjects from electroencephalography (EEG) signals is challenging due to the inter-subject variability. Domain Generalization (DG) methods aim to mitigate the domain shift among different subjects. Once trained, a DG model can be directly deployed on new subjects without any calibration phase. While existing DG studies on cross-subject emotion recognition mainly focus on the design of loss function for domain alignment or regularization, we introduce Sparse Mixture of Graph Experts (MoGE) model to explore DG issues from a new perspective, i.e. the design of the neural architecture. In the MoGE model, routers allocate each EEG channel to a specialized expert, thereby facilitating the decomposition of the intricate brain into distinct functional areas. Extensive experiments on three public datasets demonstrate that compared to other DG methods, our MoGE model trained with empirical risk minimization (ERM) achieves the state-of-the-art (SOTA) accuracies, 88.0%, 74.3%, and 81.8% on SEED, SEED-IV, and SEED-V datasets, respectively. Our code is available at https://github.com/XuanhaoLiu/MoGE. Xuan-Hao Liu, Wei-Bang Jiang, Wei-Long Zheng, Bao-Liang Lu |
BIBM | 3 |
| 2024 | Emotion Recognition from Eye Movements Using Multi-way Autoregressive ModelabstractThe application of physiological signals in emotion recognition is a popular research topic in human-computer interactions. Eye movement, as an important physiological signal, plays an essential role in medicine, psychology, cognitive science, and other scientific research fields. Previous studies have successfully identified human emotions by combining various eye-related measurements, selecting features, and utilizing machine learning techniques. However, the exploration of eye movement signals in emotion recognition remains insufficient. In this study, we utilize eye tracking heatmap and eye movement trajectory data for emotion recognition for the first time. Based on the two types of eye movement data, we develop a multi-way autoregressive model capable of processing multi-view eye movement data. Compared to traditional deep learning baseline models, our model better adapts to the structure of eye movement data and significantly improves the classification performance. Furthermore, we integrate heatmap and trajectory data with commonly used eye-related measurement features, which further enhance the performance of emotion recognition beyond previous methods. Tian-Fang Ma, Xuan-Hao Liu, Wei-Long Zheng, Bao-Liang Lu |
BIBM | 3 |
| 2024 | Multimodal Multi-View Spectral-Spatial-Temporal Masked Autoencoder for Self-Supervised Emotion RecognitionabstractEmotion recognition is a primary and complex task in emotional intelligence. Due to the complexity of human emotions, utilizing multimodal fusion methods can enhance the performance by leveraging the complementary properties of different modalities. In this paper, we propose a Multimodal Multi-view Spectral-Spatial-Temporal Masked Autoencoder (Multimodal MV-SSTMA) with self-supervised learning to investigate multimodal emotion recognition based on electroencephalogram (EEG) and eye movement signals. Our experimental process comprises three stages: 1) In the pre-training stage, we employ MV-SSTMA to train feature extractors for EEG and eye movement signals; 2) In the fine-tuning stage, the labeled data are input to the feature extractors to fuse and fine-tune the features; 3) In the testing stage, our model is applied to recognize emotions with test data to calculate the accuracies of different methods. Our experimental results demonstrate that the multimodal fusion model outperforms the unimodal model on both SEED-IV and SEED-V datasets. In addition, the proposed model can still effectively recognize emotions with various ratios of missing data. These results underscore the efficiency of multimodal self-supervised learning and data fusion in emotion recognition. Pengxuan Gao, Jia-Wen Liu, Bao-Liang Lu, Wei-Long Zheng |
ICASSP | 5 |
| 2024 | Functional Emotion Transformer for EEG-Assisted Cross-Modal Emotion RecognitionabstractMultimodal emotion recognition based on electroencephalography (EEG) and eye movements has attracted increasing attention due to their high performance and complementary properties. However, there are two challenges that hinder its practical applications: the inconvenient EEG data collection and high-cost data annotation. In contrast, eye movements are convenient to obtain and process in real scenarios. To combine high performance of EEG and easy setups of eye tracking, we propose a novel EEG-assisted Contrastive Learning Framework with a Functional Emotion Transformer (ECO-FET) for cross-modal emotion recognition. ECO-FET leverages both the functional brain connectivity and the spectral-spatial-temporal domain of EEG signals simultaneously, which dramatically benefit the learning of eye movements. The whole process consists of three phases: pre-training, test, and fine-tuning. ECO-FET exploits the complementary information provided by multiple modalities during pre-training in order to improve the performance of unimodal models. In the pre-training phase, unlabeled EEG and eye movement data are fed into the model to contrastively learn the emotional latent representations between the two modalities, while in the test phase, eye movements and few labeled EEG samples are used to predict different emotions. Experimental results on three public datasets demonstrate that ECO-FET surpasses the state-of-the-art dramatically. Wei-Bang Jiang, Ziyi Li 0003, Wei-Long Zheng, Bao-Liang Lu |
ICASSP | 3 |
| 2024 | CEMOAE: A Dynamic Autoencoder with Masked Channel Modeling for Robust EEG-Based Emotion RecognitionabstractEmotion recognition through electroencephalography (EEG) has been an area of active research, but the inherent sensitivity of EEG signals to noise and artifacts poses significant challenges, especially in real-world settings. These complications often necessitate the removal of corrupted channels, making it crucial to develop robust models capable of maintaining performance even when few channels are available. To address this, we propose the Corrupted EMOtion AutoEncoder (CEMOAE), an innovative approach that leverages masked channel modeling to maintain robust performance, achieved through three components: masked autoencoder pretraining for robust representation learning, random masked auxiliary task for implicit modeling of channel corruption, and masked auto-repair to explicitly narrow the data distribution gap between high-quality and corrupted EEG signals. Specifically, we first pretrain a masked autoencoder with the dynamic masking strategy for feature extractor initialization and channel recovery. During the finetuning stage, we mask EEG data using the auxiliary task to mimic real-world EEG corruption. We then employ the pretrained autoencoder to repair these signals and finetune the feature extractor for emotion recognition. Experiments on the SEED dataset demonstrate that CEMOAE achieves SOTA performance for emotion recognition under the random channel corruption simulation, validating the effectiveness of the proposed techniques. Yu-Ting Lan, Wei-Bang Jiang, Wei-Long Zheng, Bao-Liang Lu |
ICASSP | 3 |
| 2024 | Temporal-Spatial Prediction: Pre-Training on Diverse Datasets for EEG ClassificationabstractElectroencephalogram (EEG) classification tasks have received increasing attention because its high application value. Meanwhile, the great success of general pre-training models in language processing areas inspires us to excavate the potential of an EEG pre-trained model. This model is expected to adapt to diverse downstream tasks. However, current studies either ignore the temporal or spatial domain in EEG signals, or only use single datasets in pre-training. The proposed Temporal-Spatial Prediction (TSP) model effectively solve these issues. Specifically, the output of the TSP encoder serves as the input of two tasks: spatial prediction, i.e., masked autoencoder, and temporal prediction, i.e, contrastive predictive coding. In addition, in order to provide more diverse information and thus benefit the downstream fine-tuning, we pre-train TSP on six large EEG datasets with four different numbers of channels. Results on three public downstream datasets SEED, SEED-IV, TUEV demonstrate that TSP achieves the state-of-the-art performance on different EEG classification tasks. In addition, according to the ablation experiments, TSP performs better than the single-domain method, i.e. Temporal Prediction (TP) model and Spatial Prediction (SP) model. Ziyi Li 0003, Li-Ming Zhao, Wei-Long Zheng, Bao-Liang Lu |
ICASSP | 3 |
| 2024 | A Multi-task Emotion Recognition Model Based on Continuously Labeled EEG Signals
Rong-Fei Gu, Yi-Dong Zhao, Li-Ming Zhao, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (5) | 4 |
| 2024 | Detecting Major Depression Disorder with Multiview Eye Movement Features in a Novel Oil Painting ParadigmabstractMajor Depressive Disorder (MDD) is a debilitating condition marked by persistent low mood, reduced interest, cognitive impairments, and vegetative neurological symptoms such as sleep and appetite disturbances. In this paper, we collected eye movement signals from 40 patients diagnosed with MDD and 40 healthy controls to study the relation between eye movements and cognitive processes for depression detection. The eye movement data were captured during a novel emotional cognition task using oil paintings. Subsequently, the data were transformed into multiview eye movement features, including heatmaps, trajectories, and statistical vectors. Rigorous statistical analyses were then conducted on these features to identify significant patterns and correlations between eye movements and depressive symptoms. A multiview invariant & specific eye movement model (MISEYE) was proposed to fuse different eye movement features. The proposed achieved an accuracy rate of 79.88% in depression detection. This performance surpassed not only the outcomes of single-mode approaches and combinations of any two features but also outperformed other fusion methodologies. These findings not only shed light on the intricate relationship between eye movement patterns and MDD but also underscore the potential of eye-tracking technology in psychiatric research. Tian-Fang Ma, Lu-Yu Liu, Li-Ming Zhao, Dan Peng, Wei-Long Zheng, Bao-Liang Lu |
IJCNN | 6 |
| 2024 | EEG2Video: Towards Decoding Dynamic Visual Perception from EEG SignalsabstractOur visual experience in daily life are dominated by dynamic change. Decoding such dynamic information from brain activity can enhance the understanding of the brain’s visual processing system. However, previous studies predominately focus on reconstructing static visual stimuli. In this paper, we explore to decode dynamic visual perception from electroencephalography (EEG), a neuroimaging technique able to record brain activity with high temporal resolution (1000 Hz) for capturing rapid changes in brains. Our contributions are threefold: Firstly, we develop a large dataset recording signals from 20 subjects while they were watching 1400 dynamic video clips of 40 concepts. This dataset fills the gap in the lack of EEG-video pairs. Secondly, we annotate each video clips to investigate the potential for decoding some specific meta information (e.g., color, dynamic, human or not) from EEG. Thirdly, we propose a novel baseline EEG2Video for video reconstruction from EEG signals that better aligns dynamic movements with high temporal resolution brain signals by Seq2Seq architecture. EEG2Video achieves a 2-way accuracy of 79.8% in semantic classification tasks and 0.256 in structural similarity index (SSIM). Overall, our works takes an important step towards decoding dynamic visual perception from EEG signals. Our dataset and code will be released soon. Xuan-Hao Liu, Yan-Kai Liu, Yansen Wang, Kan Ren, Hanwen Shi, Zilong Wang 0006, Dongsheng Li 0002, Bao-Liang Lu, Wei-Long Zheng |
NeurIPS | 9 |
| 2024 | Code Repair with LLMs gives an Exploration-Exploitation TradeoffabstractIteratively improving and repairing source code with large language models (LLMs), known as refinement, has emerged as a popular way of generating programs that would be too complex to construct in one shot. Given a bank of test cases, together with a candidate program, an LLM can improve that program by being prompted with failed test cases. But it remains an open question how to best iteratively refine code, with prior work employing simple greedy or breadth-first strategies. We show here that refinement exposes an explore-exploit tradeoff: exploit by refining the program that passes the most test cases, or explore by refining a lesser considered program. We frame this as an arm-acquiring bandit problem, which we solve with Thompson Sampling. The resulting LLM-based program synthesis algorithm is broadly applicable: Across loop invariant synthesis, visual reasoning puzzles, and competition programming problems, we find that our new method can solve more problems using fewer language model calls. Hao Tang 0008, Keya Hu, Sicheng Zhong, Wei-Long Zheng, Xujie Si, Kevin Ellis |
NeurIPS | 5 |
| 2024 | FuseAnyPart: Diffusion-Driven Facial Parts Swapping via Multiple Reference ImagesabstractFacial parts swapping aims to selectively transfer regions of interest from the source image onto the target image while maintaining the rest of the target image unchanged.
Most studies on face swapping designed specifically for full-face swapping, are either unable or significantly limited when it comes to swapping individual facial parts, which hinders fine-grained and customized character designs.
However, designing such an approach specifically for facial parts swapping is challenged by a reasonable multiple reference feature fusion, which needs to be both efficient and effective.
To overcome this challenge, FuseAnyPart is proposed to facilitate the seamless "fuse-any-part" customization of the face.
In FuseAnyPart, facial parts from different people are assembled into a complete face in latent space within the Mask-based Fusion Module.
Subsequently, the consolidated feature is dispatched to the Addition-based Injection Module for
fusion within the UNet of the diffusion model to create novel characters.
Extensive experiments qualitatively and quantitatively validate the superiority and robustness of FuseAnyPart.
Source codes are available at https://github.com/Thomas-wyh/FuseAnyPart. Siying Cui, Aixi Zhang, Wei-Long Zheng, Senzhang Wang |
NeurIPS | 5 |
| 2023 | Transformer-Based Domain Adaptation for Multi-Modal Emotion Recognition in Response to Game Animation VideosabstractEmotion recognition researches necessitate the strategic selection of stimuli to evoke targeted emotions for robust physiological analyses. This study pioneers the use of Genshin Impact game animation videos to induce positive and neutral emotional states. Notably, it introduces a Transformer-based feature extractor, enhancing Domain-Adversarial Neural Networks (DANN) to advance domain adaptation capabilities. Leveraging the inherent advantages of the Transformer architecture, including parallel processing and handling intricate time-series Electroencephalogram (EEG) and eye movement data, this innovation is reinforced by a discriminator with gradient reversal layers, harmonizing source and target domain distributions. Empirical results demonstrate the effectiveness of the Transformer-based DANN model in cross-subject multimodal emotion recognition, achieving a remarkable prediction accuracy of 83.38% across 59 subjects. Jing-Yi Liu, Jia-Wen Liu, Wei-Long Zheng, Bao-Liang Lu |
BIBM | 3 |
| 2023 | Elastic Graph Transformer Networks for EEG-Based Emotion RecognitionabstractElectroencephalogram (EEG) has been applied in emotion recognition due to excellent temporal resolution with less competitive spatial resolution. This leads to the consequence that the majority of EEG-based emotion recognition models emphasize on exploiting temporal features while ignoring the efficient information provided by spatial resolution. To extract more informative representations, we propose an elastic Graph Transformer network for emotion recognition (EmoGT) inspired by the advantages of Transformer in time-series analysis and the superior performance of graph convolutional networks in topological analysis. Moreover, it is able to be flexibly expanded to cope with multimodal inputs by employing specially designed structures. Experimental results on 3 public datasets demonstrate that our models outperform the state-of-the-art results by 3% on average in both single and multimodal cases, indicating the effectiveness of utilizing temporal and spatial information simultaneously. Wei-Bang Jiang, Xu Yan 0007, Wei-Long Zheng, Bao-Liang Lu |
ICASSP | 3 |
| 2023 | DAformer: Transformer with Domain Adversarial Adaptation for EEG-Based Emotion Recognition with Live-Oil Paintings
Zhong-Wei Jin, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (9) | 3 |
| 2023 | Two-Stream Spectral-Temporal Denoising Network for End-to-End Robust EEG-Based Emotion Recognition
Xuan-Hao Liu, Wei-Bang Jiang, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (3) | 3 |
| 2023 | Naturalistic Emotion Recognition Using EEG and Eye Movements
Ziyi Li 0003, Tian-Fang Ma, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (3) | 6 |
| 2023 | Cross-Subject Decision Confidence Estimation from EEG Signals Using Spectral-Spatial-Temporal Adaptive GCN with Domain AdaptationabstractThe study of the human decision-making process has long been a valuable field for both scientific research and practical application. Towards knowing and taking control of the decision-making process, evaluating the reliability of human decisions objectively plays an important role. Various studies have demonstrated that the confidence level of humans during the decision-making process is an important factor that reflects the correctness of decisions. In literature, several deep learning based methods have been developed to estimate decision confidence using Electroencephalography (EEG). Among these approaches, the spectral-spatial-temporal adaptive graph convolutional neural network (SST-AGCN) stands out. However, SST-AGCN focuses on specific subjects, and may lead to less efficiency in cross-subject situations, which are more common in application scenarios. In this paper, we propose a deep learning model called SST-AGCN with Domain Adaptation (SST-AGCN-DA) for cross-subject decision confidence estimation. To examine the effectiveness of our proposed model, we compare our SST-AGCN-DA with the original SST-AGCN, three typical domain adaption algorithms in the field, and the SST-AGCN with Domain Generalization (SST-AGCN-DG), which is another transfer learning model we developed in this paper. We conduct cross-subject confidence estimation experiments on an EEG dataset collected under a text-based decision-making task. The averaged results of leave-one-out cross-validation come out that the F1-scores of our proposed SST-AGCN-DA and SST-AGCN-DG are 79.45% and 77.04%, respectively, while the original SST-AGCN and the best of the existing domain adaptation algorithms are 74.15% and 74.25%, respectively. Rong-Fei Gu, Rui Li 0048, Wei-Long Zheng, Bao-Liang Lu |
IJCNN | 3 |
| 2023 | Multimodal Adaptive Emotion Transformer with Flexible Modality Inputs on A Novel Dataset with Continuous LabelsabstractEmotion recognition from physiological signals is a topic of widespread interest, and researchers continue to develop novel techniques for perceiving emotions. However, the emergence of deep learning has highlighted the need for high-quality emotional datasets to accurately decode human emotions. In this study, we present a novel multimodal emotion dataset that incorporates electroencephalography (EEG) and eye movement signals to systematically explore human emotions. Seven basic emotions (happy, sad, fear, disgust, surprise, anger, and neutral) are elicited by a large number of 80 videos and fully investigated with continuous labels that indicate the intensity of the corresponding emotions. Additionally, we propose a novel Multimodal Adaptive Emotion Transformer (MAET), that can flexibly process both unimodal and multimodal inputs. Adversarial training is utilized in MAET to mitigate subject discrepancy, which enhances domain generalization. Our extensive experiments, encompassing both subject-dependent and cross-subject conditions, demonstrate MAET's superior performance in handling various inputs. The filtering of data for high emotional evocation using continuous labels proved to be effective in the experiments. Furthermore, the complementary properties between EEG and eye movements are observed. Our code is available at https://github.com/935963004/MAET. Wei-Bang Jiang, Xuan-Hao Liu, Wei-Long Zheng, Bao-Liang Lu |
ACM Multimedia | 3 |
| 2022 | Few-Shot Class-Incremental Learning for EEG-Based Emotion Recognition
Tian-Fang Ma, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (5) | 2 |
| 2022 | A Multi-view Spectral-Spatial-Temporal Masked Autoencoder for Decoding Emotions with Self-supervised LearningabstractAffective Brain-computer Interface has achieved considerable advances that researchers can successfully interpret labeled and flawless EEG data collected in laboratory settings. However, the annotation of EEG data is time-consuming and requires a vast workforce which limits the application in practical scenarios. Furthermore, daily collected EEG data may be partially damaged since EEG signals are sensitive to noise. In this paper, we propose a Multi-view Spectral-Spatial-Temporal Masked Autoencoder (MV-SSTMA) with self-supervised learning to tackle these challenges towards daily applications. The MV-SSTMA is based on a multi-view CNN-Transformer hybrid structure, interpreting the emotion-related knowledge of EEG signals from spectral, spatial, and temporal perspectives. Our model consists of three stages: 1) In the generalized pre-training stage, channels of unlabeled EEG data from all subjects are randomly masked and later reconstructed to learn the generic representations from EEG data; 2) In the personalized calibration stage, only few labeled data from a specific subject are used to calibrate the model; 3) In the personal test stage, our model can decode personal emotions from the sound EEG data as well as damaged ones with missing channels. Extensive experiments on two open emotional EEG datasets demonstrate that our proposed model achieves state-of-the-art performance on emotion recognition. In addition, under the abnormal circumstance of missing channels, the proposed model can still effectively recognize emotions. Rui Li 0048, Wei-Long Zheng, Bao-Liang Lu |
ACM Multimedia | 3 |
| 2022 | Multimodal Vigilance Estimation Using Deep LearningabstractThe phenomenon of increasing accidents caused by reduced vigilance does exist. In the future, the high accuracy of vigilance estimation will play a significant role in public transportation safety. We propose a multimodal regression network that consists of multichannel deep autoencoders with subnetwork neurons (MCDAE$_{sn}$). After we define two thresholds of “0.35” and “0.70” from the percentage of eye closure, the output values are in the continuous range of 0–0.35, 0.36–0.70, and 0.71–1 representing the awake state, the tired state, and the drowsy state, respectively. To verify the efficiency of our strategy, we first applied the proposed approach to a single modality. Then, for the multimodality, since the complementary information between forehead electrooculography and electroencephalography features, we found the performance of the proposed approach using features fusion significantly improved, demonstrating the effectiveness and efficiency of our method. Wei Wu 0022, Wei Sun 0028, Q. M. Jonathan Wu, Yimin Yang 0001, Hui Zhang 0023, Wei-Long Zheng, Bao-Liang Lu |
IEEE Trans. Cybern. | 6 |
| 2020 | Early Action Recognition With Category Exclusion Using Policy-Based Reinforcement LearningabstractThe goal of early action recognition is to predict action label when the sequence is partially observed. The existing methods treat the early action recognition task as sequential classification problems on different observation ratios of an action sequence. Since these models are trained by differentiating positive category from all negative classes, the diverse information of different negative categories is ignored, which we believe can be collected to help improve the recognition performance. In this paper, we step towards to a new direction by introducing category exclusion to early action recognition. We model the exclusion as a mask operation on the classification probability output of a pre-trained early action recognition classifier. Specifically, we use policy-based reinforcement learning to train an agent. The agent generates a series of binary masks to exclude interfering negative categories during action execution and hence help improve the recognition accuracy. The proposed method is evaluated on three benchmark recognition datasets, NTU-RGBD, First-Person Hand Action, as well as UCF-101. The proposed method enhances the recognition accuracy consistently over all different observation ratios on the three datasets, where the accuracy improvements on the early stages are especially significant. Junwu Weng, Xudong Jiang 0001, Wei-Long Zheng, Junsong Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Vigilance Estimation Using a Wearable EOG Device in Real Driving EnvironmentabstractVigilance decrement in driving tasks has been reported to be a major factor in fatal accidents and could severely endanger public transportation safety. However, efficient approaches for estimating vigilance in real driving environment are still lacking. In this paper, we propose a novel approach for implementing continuous vigilance estimation using forehead electrooculograms (EOGs) acquired by wearable dry electrodes in both simulated and real driving environments. To improve the feasibility of this approach for real-world applications, a forehead EOG-based electrode placement with only four electrodes is designed. Flexible dry electrodes and an acquisition board are integrated as a wearable device for recording EOGs. Twenty and ten subjects participated in the simulated and real-world driving environment experiments, respectively. Accurate eye movement parameters from eye-tracking glasses are extracted to calculate the PERCLOS index for vigilance annotation. This is because the vigilance state is a temporally dynamic process, and a continuous conditional random field and a continuous conditional neural field are introduced to construct more accurate vigilance estimation models. To evaluate the efficiency of our system, systematic experiments are performed in real scenarios under various illumination and weather conditions following laboratory simulations as preliminary studies. The experimental results demonstrate that the wearable dry electrode prototype, which has a relatively comfortable forehead setup, can efficiently capture vigilance dynamics. The best mean correlation coefficients achieved by our proposed approach are 71.18% and 66.20% in laboratory simulations and real-world driving environments, respectively. The cross-environment experiments are performed to evaluate the simulated-to-real generalization and a best mean correlation coefficient of 53.96% is achieved. Wei-Long Zheng, Kunpeng Gao, Gang Li 0011, Wei Liu 0078, Chao Liu 0025, Guoxing Wang, Bao-Liang Lu |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2019 | Reducing the Subject Variability of EEG Signals with Adversarial Domain Generalization
Bo-Qun Ma, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (1) | 3 |
| 2019 | Emotion Recognition using Multimodal Residual LSTM NetworkabstractVarious studies have shown that the temporal information captured by conventional long-short-term memory (LSTM) networks is very useful for enhancing multimodal emotion recognition using encephalography (EEG) and other physiological signals. However, the dependency among multiple modalities and high-level temporal-feature learning using deeper LSTM networks is yet to be investigated. Thus, we propose a multimodal residual LSTM (MMResLSTM) network for emotion recognition. The MMResLSTM network shares the weights across the modalities in each LSTM layer to learn the correlation between the EEG and other physiological signals. It contains both the spatial shortcut paths provided by the residual network and temporal shortcut paths provided by LSTM for efficiently learning emotion-related high-level features. The proposed network was evaluated using a publicly available dataset for EEG-based emotion recognition, DEAP. The experimental results indicate that the proposed MMResLSTM network yielded a promising result, with a classification accuracy of 92.87% for arousal and 92.30% for valence. Jia-Xin Ma, Hao Tang 0008, Wei-Long Zheng, Bao-Liang Lu |
ACM Multimedia | 3 |
| 2019 | Identifying Stable Patterns over Time for Emotion Recognition from EEGabstractIn this paper, we investigate stable patterns of electroencephalogram (EEG) over time for emotion recognition using a machine learning approach. Up to now, various findings of activated patterns associated with different emotions have been reported. However, their stability over time has not been fully investigated yet. In this paper, we focus on identifying EEG stability in emotion recognition. We systematically evaluate the performance of various popular feature extraction, feature selection, feature smoothing and pattern classification methods with the DEAP dataset and a newly developed dataset called SEED for this study. Discriminative Graph regularized Extreme Learning Machine with differential entropy features achieves the best average accuracies of 69.67 and 91.07 percent on the DEAP and SEED datasets, respectively. The experimental results indicate that stable patterns exhibit consistency across sessions; the lateral temporal areas activate more for positive emotions than negative emotions in beta and gamma bands; the neural patterns of neutral emotions have higher alpha responses at parietal and occipital sites; and for negative emotions, the neural patterns have significant higher delta responses at parietal and occipital sites and higher gamma responses at prefrontal sites. The performance of our emotion recognition models shows that the neural patterns are relatively stable within and between sessions. Wei-Long Zheng, Jia-Yi Zhu, Bao-Liang Lu |
IEEE Trans. Affect. Comput. | 1 |
| 2019 | EmotionMeter: A Multimodal Framework for Recognizing Human EmotionsabstractIn this paper, we present a multimodal emotion recognition framework called EmotionMeter that combines brain waves and eye movements. To increase the feasibility and wearability of EmotionMeter in real-world applications, we design a six-electrode placement above the ears to collect electroencephalography (EEG) signals. We combine EEG and eye movements for integrating the internal cognitive states and external subconscious behaviors of users to improve the recognition accuracy of EmotionMeter. The experimental results demonstrate that modality fusion with multimodal deep neural networks can significantly enhance the performance compared with a single modality, and the best mean accuracy of 85.11% is achieved for four emotions (happy, sad, fear, and neutral). We explore the complementary characteristics of EEG and eye movements for their representational capacities and identify that EEG has the advantage of classifying happy emotion, whereas eye movements outperform EEG in recognizing fear emotion. To investigate the stability of EmotionMeter over time, each subject performs the experiments three times on different days. EmotionMeter obtains a mean recognition accuracy of 72.39% across sessions with the six-electrode EEG and eye movement features. These experimental results demonstrate the effectiveness of EmotionMeter within and between sessions. Wei-Long Zheng, Wei Liu 0078, Bao-Liang Lu, Andrzej Cichocki |
IEEE Trans. Cybern. | 1 |
| 2018 | Cross-Subject Emotion Recognition Using Deep Adaptation Networks
Wei-Long Zheng, Bao-Liang Lu |
ICONIP (5) | 3 |
| 2018 | WGAN Domain Adaptation for EEG-Based Emotion Recognition
Si-Yang Zhang, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (5) | 3 |
| 2018 | Active Feedback Framework with Scan-Path Clustering for Deep Affective Models
Li-Ming Zhao, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (2) | 3 |
| 2018 | Multimodal Vigilance Estimation with Adversarial Domain Adaptation NetworksabstractRobust vigilance estimation during driving is very crucial in preventing traffic accidents. Many approaches have been proposed for vigilance estimation. However, most of the approaches require collecting subject-specific labeled data for calibration which is high-cost for real-world applications. To solve this problem, domain adaptation methods can be used to align distributions of source subject features (source domain) and new subject features (target domain). By reusing existing data from other subjects, no labeled data of new subjects is required to train models. In this paper, our goal is to apply adversarial domain adaptation networks to cross-subject vigilance estimation. We adopt two kinds of recently proposed adversarial domain adaptation networks and compare their performance with those of several traditional domain adaptation methods and the baseline without domain adaptation. A publicly available dataset, SEED-VIG, is used to evaluate the methods. The dataset includes electroencephalography (EEG) and electrooculography (EOG) signals, as well as the corresponding vigilance level annotations during simulated driving. Compared with the baseline, both adversarial domain adaptation networks achieve improvements over 10% in terms of Pearson's correlation coefficient. In addition, both methods considerably outperform the traditional domain adaptation methods. Wei-Long Zheng, Bao-Liang Lu |
IJCNN | 2 |
| 2018 | Sleep Quality Estimation with Adversarial Domain Adaptation: From Laboratory to Real ScenarioabstractPrevious studies on EEG-based last-night sleep quality estimation mainly focus on evaluation with data from laboratory experiments. However, due to the reality gap constituted with device performance, subject groups, experiment settings and controlled conditions, the models trained solely on laboratory data cannot generalize well to real scenarios. In this work, we investigate the sleep quality estimation for high-speed train drivers as an instance of real-scenario application. Domain adaptation models are adopted to deal with the individual differences across subjects when modeling and testing with the real-scenario data. As it is usually difficult and costly to acquire data and annotate them in real scenarios, the high-quality data in laboratory conditions are used for model trainings. Knowledge from simulation is transferred to reality with domain adaptation methods. A novel approach called Domain Adversarial Neural Network (DANN) is adopted. DANN learns domain independent features through deep networks with an adversarial architecture. The experimental results indicate that DANN outperforms other state-of-the-art methods and achieves 19.55% and 23.50% improvements in terms of accuracy on the cross-subject and cross- scenario tasks, respectively, in comparison with the baseline SVM model. Jia-Jun Tong, Bo-Qun Ma, Wei-Long Zheng, Bao-Liang Lu, Xiao-Qi Song, Shi-Wei Ma |
IJCNN | 4 |
| 2018 | Semi-supervised Deep Generative Modelling of Incomplete Multi-Modality Emotional DataabstractThere are threefold challenges in emotion recognition. First, it is difficult to recognize human's emotional states only considering a single modality. Second, it is expensive to manually annotate the emotional data. Third, emotional data often suffers from missing modalities due to unforeseeable sensor malfunction or configuration issues. In this paper, we address all these problems under a novel multi-view deep generative framework. Specifically, we propose to model the statistical relationships of multi-modality emotional data using multiple modality-specific generative networks with a shared latent space. By imposing a Gaussian mixture assumption on the posterior approximation of the shared latent variables, our framework can learn the joint deep representation from multiple modalities and evaluate the importance of each modality simultaneously. To solve the labeled-data-scarcity problem, we extend our multi-view model to semi-supervised learning scenario by casting the semi-supervised classification problem as a specialized missing data imputation task. To address the missing-modality problem, we further extend our semi-supervised multi-view model to deal with incomplete data, where a missing view is treated as a latent variable and integrated out during inference. This way, the proposed overall framework can utilize all available (both labeled and unlabeled, as well as both complete and incomplete) data to improve its generalization ability. The experiments conducted on two real multi-modal emotion datasets demonstrated the superiority of our framework. Changde Du, Changying Du, Hao Wang 0005, Jinpeng Li 0002, Wei-Long Zheng, Bao-Liang Lu, Huiguang He |
ACM Multimedia | 5 |
| 2017 | Multimodal Emotion Recognition Using Deep Neural Networks
Hao Tang 0008, Wei Liu 0078, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (4) | 3 |
| 2017 | Identifying Gender Differences in Multimodal Emotion Recognition Using Bimodal Deep AutoEncoder
Wei-Long Zheng, Wei Liu 0078, Bao-Liang Lu |
ICONIP (4) | 2 |
| 2017 | Investigating Gender Differences of Brain Areas in Emotion Recognition Using LSTM Neural Network
Wei-Long Zheng, Wei Liu 0078, Bao-Liang Lu |
ICONIP (4) | 2 |
| 2017 | EEG-Based Sleep Quality Evaluation with Deep Transfer Learning
Xing-Zan Zhang, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (4) | 2 |
| 2017 | Emotion Annotation Using Hierarchical Aligned Cluster Analysis
Wei-Ye Zhao, Ting Ji, Qian Ji, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (4) | 5 |
| 2017 | Online Depth Image-Based Object Tracking with Sparse Representation and Object Detection
Wei-Long Zheng, Shan-Chun Shen, Bao-Liang Lu |
Neural Process. Lett. | 1 |
| 2016 | Emotion Recognition Using Multimodal Deep Learning
Wei Liu 0078, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (2) | 2 |
| 2016 | Continuous Vigilance Estimation Using LSTM Neural Networks
Wei-Long Zheng, Wei Liu 0078, Bao-Liang Lu |
ICONIP (2) | 2 |
| 2016 | Personalizing EEG-Based Affective Models with Transfer Learning
Wei-Long Zheng, Bao-Liang Lu |
IJCAI | 1 |
| 2016 | Driving fatigue detection with fusion of EEG and forehead EOGabstractIn this paper, we fuse EEG and forehead EOG to detect drivers' fatigue level by using discriminative graph regularized extreme learning machine (GELM). Twenty-one healthy subjects including twelve men and nine women participate in our driving simulation experiments. Two fusion strategies are adopted: feature level fusion (FLF) and decision level fusion (DLF). PERCLOS (the percentage of eye closure) is calculated by using the eye movement data recorded by eye tracking glasses as the indicator of drivers' fatigue level. The prediction correlation coefficient and root mean square error (RMSE) between the estimated fatigue level and the real fatigue level are both used to evaluate the performance of single modality and fusion modality. A comparative study on modality performance is conducted between GELM and support vector machine (SVM). The experimental results show that fusion modality can improve the performance of driving fatigue detection with a higher prediction correlation coefficient and a lower RMSE value in comparison with solely using EEG or forehead EOG. And FLF achieves better performance than DLF. GELM is more suitable for driving fatigue detection than SVM. Moreover, feature level fusion with GELM achieves the best performance with the prediction correlation coefficient of 0.8080 and the RMSE value of 0.0712 on average. Xue-Qin Huo, Wei-Long Zheng, Bao-Liang Lu |
IJCNN | 2 |
| 2016 | Measuring sleep quality from EEG with machine learning approachesabstractThis study aims at measuring last-night sleep quality from electroencephalography (EEG). We design a sleep experiment to collect waking EEG signals from eight subjects under three different sleep conditions: 8 hours sleep, 6 hours sleep, and 4 hours sleep. We utilize three machine learning approaches, k-Nearest Neighbor (kNN), support vector machine (SVM), and discriminative graph regularized extreme learning machine (GELM), to classify extracted EEG features of power spectral density (PSD). The accuracies of these three classifiers without feature selection are 36.68%, 48.28%, 62.16%, respectively. By using minimal-redundancy-maximal-relevance (MRMR) algorithm and the brain topography, the classification accuracy of GELM with 9 features is improved largely and increased to 83.57% in average. To investigate critical frequency bands for measuring sleep quality, we examine the features of each band and observe their energy changing. The experimental results indicate that Gamma band is more relevant to measuring sleep quality. Wei-Long Zheng, Hai-Wei Ma, Bao-Liang Lu |
IJCNN | 2 |
| 2016 | An unsupervised discriminative extreme learning machine and its applications to data clustering
Yong Peng 0001, Wei-Long Zheng, Bao-Liang Lu |
Neurocomputing | 2 |
| 2015 | Transfer components between subjects for EEG-based emotion recognitionabstractAddressing the structural and functional variability between subjects for robust affective brain-computer interface (aBCI) is challenging but of great importance, since the calibration phase for aBCI is time-consuming. In this paper, we propose a subject transfer framework for electroencephalogram (EEG)-based emotion recognition via component analysis. We compare two state-of-the-art subspace projecting approaches called transfer component analysis (TCA) and kernel principle component analysis (KPCA) for subject transfer. The main idea is to learn a set of transfer components underlying source domain (source subjects) and target domain (target subject). When projected to this subspace, the difference of feature distributions of both domains can be reduced. From the experiments, we show that the two proposed approaches, TCA and KPCA, can achieve an improvement on performance with the best mean accuracies of 71.80% and 79.83%, respectively, in comparison of the baseline of 58.95%. The significant improvement shows the feasibility and efficiency of our approaches for subject transfer emotion recognition from EEG signals. Wei-Long Zheng, Yong-Qi Zhang, Jia-Yi Zhu, Bao-Liang Lu |
ACII | 1 |
| 2015 | Transfer Components Between Subjects for EEG-based Driving Fatigue Detection
Yong-Qi Zhang, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (4) | 2 |
| 2015 | Combining Eye Movements and EEG to Enhance Emotion Recognition
Wei-Long Zheng, Bao-Liang Lu |
IJCAI | 2 |
| 2014 | EEG-based emotion classification using deep belief networksabstractIn recent years, there are many great successes in using deep architectures for unsupervised feature learning from data, especially for images and speech. In this paper, we introduce recent advanced deep learning models to classify two emotional categories (positive and negative) from EEG data. We train a deep belief network (DBN) with differential entropy features extracted from multichannel EEG as input. A hidden markov model (HMM) is integrated to accurately capture a more reliable emotional stage switching. We also compare the performance of the deep models to KNN, SVM and Graph regularized Extreme Learning Machine (GELM). The average accuracies of DBN-HMM, DBN, GELM, SVM, and KNN in our experiments are 87.62%, 86.91%, 85.67%, 84.08%, and 69.66%, respectively. Our experimental results show that the DBN and DBN-HMM models improve the accuracy of EEG-based emotion classification in comparison with the state-of-the-art methods. Wei-Long Zheng, Jia-Yi Zhu, Yong Peng 0001, Bao-Liang Lu |
ICME | 1 |
| 2014 | Online Object Tracking Based on Depth Image with Sparse Coding
Shan-Chun Shen, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (3) | 2 |
| 2014 | EOG-based drowsiness detection using convolutional neural networksabstractThis study provides a new application of convolutional neural networks for drowsiness detection based on electrooculography (EOG) signals. Drowsiness is charged to be one of the major causes of traffic accidents. Such application is helpful to reduce losses of casualty and property. Most attempts at drowsiness detection based on EOG involve a feature extraction step, which is accounted as time-consuming task, and it is difficult to extract effective features. In this paper, an unsupervised learning is proposed to estimate driver fatigue based on EOG. A convolutional neural network with a linear regression layer is applied to EOG signals in order to avoid using of manual features. With a postprocessing step of linear dynamic system (LDS), we are able to capture the physiological status shifting. The performance of the proposed model is evaluated by the correlation coefficients between the final outputs and the local error rates of the subjects. Compared with the results of a manual ad-hoc feature extraction approach, our method is proven to be effective for drowsiness detection. Xuemin Zhu, Wei-Long Zheng, Bao-Liang Lu, Shanguang Chen |
IJCNN | 2 |
| 2014 | EEG-based emotion recognition using discriminative graph regularized extreme learning machineabstractThis study aims at finding the relationship between EEG signals and human emotional states. Movie clips are used as stimuli to evoke positive, neutral and negative emotions of subjects. We introduce a new effective classifier named discriminative graph regularized extreme learning machine (GELM) for EEG-based emotion recognition. The average classification accuracy of GELM using differential entropy (DE) features on the whole five frequency bands is 80.25%, while the accuracy of SVM is 76.62%. These results indicate that GELM is more suitable for emotion recognition than SVM. Additionally, the accuracies of GELM using DE features on Beta and Gamma bands are 79.07%, 79.93% respectively. This suggests that these two bands are more relevant to emotion. The experimental results indicate that the EEG patterns for emotion are generally stable among different experiments and subjects. By using minimal-redundancy-maximal-relevance (MRMR) algorithm and correlation coefficients to select effective features, we get the distribution of top 20 subject-independent features and build a manifold model to monitor the trajectory of emotion changes with time. Jia-Yi Zhu, Wei-Long Zheng, Yong Peng 0001, Ruo-Nan Duan, Bao-Liang Lu |
IJCNN | 2 |