Yi Ding 0012

dblp:89/5503-12 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
18since 2021 · last 2026
0000-0003-2365-3568ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 EEG-DLite: Dataset Distillation for Efficient Large EEG Model Training
abstract
Large-scale EEG foundation models have shown strong generalization across a range of downstream tasks, but their training remains resource-intensive due to the volume and variable quality of EEG data. In this work, we introduce EEG-DLite, a data distillation framework that enables more efficient pre-training by selectively removing noisy and redundant samples from large EEG datasets. EEG-DLite begins by encoding EEG segments into compact latent representations using a self-supervised autoencoder, allowing sample selection to be performed efficiently and with reduced sensitivity to noise. Based on these representations, EEG-DLite filters out outliers and minimizes redundancy, resulting in a smaller yet informative subset that retains the diversity essential for effective foundation model training. Through extensive experiments, we demonstrate that training on only 5 percent of a 2,500-hour dataset curated with EEG-DLite yields performance comparable to, and in some cases better than, training on the full dataset across multiple downstream tasks. To our knowledge, this is the first systematic study of pre-training data distillation in the context of EEG foundation models. EEG-DLite provides a scalable and practical path toward more effective and efficient physiological foundation modeling.
Yuting Tang, Wei-Bang Jiang, Shanglin Li, Yong Li 0032, Xinliang Zhou, Yi Ding 0012, Cuntai Guan
AAAI7
2026 Decoupled Hierarchical Distillation for Multimodal Emotion Recognition
abstract
Human multimodal emotion recognition (MER) seeks to infer human emotions by integrating information from language, visual, and acoustic modalities. Although existing MER approaches have achieved promising results, they still struggle with inherent multimodal heterogeneities and varying contributions from different modalities. To address these challenges, we propose a novel framework, Decoupled Hierarchical Multimodal Distillation (DHMD). DHMD decouples each modality's features into modality-irrelevant (homogeneous) and modality-exclusive (heterogeneous) components using a self-regression mechanism. The framework employs a two-stage knowledge distillation (KD) strategy: (1) coarse-grained KD via a Graph Distillation Unit (GD-Unit) in each decoupled feature space, where a dynamic graph facilitates adaptive distillation among modalities, and (2) fine-grained KD through a cross-modal dictionary matching mechanism, which aligns semantic granularities across modalities to produce more discriminative MER representations. This hierarchical distillation approach enables flexible knowledge transfer and effectively improves cross-modal feature alignment. Experimental results demonstrate that DHMD consistently outperforms state-of-the-art MER methods, achieving 1.3%/2.4% (ACC$_{7}$7), 1.3%/1.9% (ACC$_{2}$2) and 1.9%/1.8% (F1) relative improvement on CMU-MOSI/CMU-MOSEI dataset, respectively. Meanwhile, visualization results reveal that both the graph edges and dictionary activations in DHMD exhibit meaningful distribution patterns across modality-irrelevant/-exclusive feature spaces.
Yong Li 0032, Yuanzhi Wang, Yi Ding 0012, Shiqing Zhang, Ke Lu 0002, Cuntai Guan
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Decoding Covert Speech From EEG by Functional Areas Spatio-Temporal Transformer
abstract
Covert speech involves imagining speaking without audible sound or any movements.Decoding covert speech from electroencephalogram (EEG) is challenging due to a limited understanding of neural pronunciation mapping and the low signal-to-noise ratio of the signal. In this study, we developed a large-scale multi-utterance speech EEG dataset from 57 right-handed native English-speaking subjects, each performing covert and overt speech tasks by repeating the same word in five utterances within a ten-second duration. Given the spatio-temporal nature of the neural activation process during speech pronunciation, we developed a Functional Areas Spatio-temporal Transformer (FAST), an effective framework for converting EEG signals into tokens and utilizing transformer architecture for sequence encoding. Our results reveal distinct and interpretable speech neural features by the visualization of FAST-generated activation maps across frontal and temporal brain regions, with each word being covertly spoken, providing new insights into the discriminative features of the neural representation of covert speech. This is the first report of such a study, which provides interpretable evidence for speech decoding from EEG.
Muyun Jiang, Wei Zhang 0266, Yi Ding 0012, Kok Ann Colin Teo, Laiguan Fong, Shuailei Zhang, Raghavan Bhuvanakantham, Wei Khang Jeremy Sim, Chuan Huat Vince Foo, Rong Hui Jonathan Chua, Parasuraman Padmanabhan, Victoria Leong, Balázs Gulyás, Cuntai Guan
IEEE J. Biomed. Health Informatics3
2025 CAT-Net: A Co-Adaptive Transfer Learning Network for BCI-Assisted Neurorehabilitation
abstract
Brain-computer interfaces (BCIs) hold great potential for motor recovery in post-stroke patients. However, the motor imagery decoding accuracy is limited by the non-stationarity of EEG signals across subjects and sessions. We propose CAT-Net: a Co-Adaptive Transfer learning network to simultaneously address the inter-subject variability and inter-session nonstationarity in EEG data. The proposed method selects a relevant subset of data from all the available subjects’ data to train an initial model, followed by subject-specific transfer learning from the initial model to the target subject to establish a pretrain model. Subsequently, online adaptive training is then applied to incrementally train the pretrain model using the data from previous sessions for the target subject. This proposed network using this unique co-adaptive training method is then evaluated on both upper and lower-limb neurorehabilitation EEG datasets comprising 358 sessions from 33 stroke patients. The results showed significant accuracy improvements, achieving averaged accuracies of 70.6% and 72.3% on the respective datasets, surpassing the state-of-the-art baselines.
Shuailei Zhang, Yi Ding 0012, Muyun Jiang, Effie Chew, Kai Keng Ang, Cuntai Guan
ICASSP2
2025 SelectiveFinetuning: Enhancing Transfer Learning In Sleep Staging Through Selective Domain Alignment
abstract
In practical sleep stage classification, a key challenge is the variability of EEG data across different subjects and environments. Differences in physiology, age, health status, and recording conditions can lead to domain shifts between data. These domain shifts often result in decreased model accuracy and reliability, particularly when the model is applied to new data with characteristics different from those it was originally trained on, which is a typical manifestation of negative transfer. To address this, we propose SelectiveFinetuning in this paper. Our method utilizes a pre-trained Multi-Resolution Convolutional Neural Network (MRCNN) to extract EEG features, capturing the distinctive characteristics of different sleep stages. To mitigate the effect of domain shifts, we introduce a domain aligning mechanism that employs Earth Mover’s Distance (EMD) to evaluate and select source domain data closely matching the target domain. By finetuning the model with selective source data, our SelectiveFinetuning enhances the model’s performance on target domain that exhibits domain shifts compared to the data used for training. Experimental results show that our method outperforms existing baselines, offering greater robustness and adaptability in practical scenarios where data distributions are often unpredictable.
Yi Ding 0012, Xinliang Zhou
ICASSP3
2025 Sera: Separated Coarse-to-fine Representation Alignment for Cross-subject EEG-based Emotion Recognition
abstract
Neuropsychology-inspired models have been utilized in recent advances in EEG emotion recognition, such as convolutional networks for spatial features and Transformers for temporal dependencies. While these methods benefit from domain knowledge like frequency-band features and spatial correlations, most overlook the fundamental fact that EEG signals are complex mixtures of neural source activities recorded at the scalp. EEG signals presenting challenges for emotion recognition, particularly in cross-subject scenarios due to significant inter-subject variance. Inspired by neurophysiological principles, we propose a novel framework, named Sera, for EEG-based emotion recognition that explicitly separates source activities and aligns representations across subjects. Sera introduces two key components: (1) a variational autoencoder (VAE) with multiple multi-stage decoders (M2VAE) designed to disentangle EEG signals into independent sources, mimicking the neural generation process, and (2) a coarse-to-fine representation alignment block (CFRA) to mitigate subject-to-subject variability. The coarse alignment employs adversarial training with a domain discriminator, while the fine-grained alignment matches covariance matrices to capture temporal correlations within EEG segments. Extensive experiments demonstrate that Sera outperforms the state-of-the-art methods with improvements ranging from 1% to 5%, averaging 3.14% and 3.05% on the DEAP and DREAMER datasets, respectively, confirming its effectiveness and neurophysiological grounding. The code is available at: https://github.com/JZH98/Sera-code.
Meiyan Xu, Ziyu Jia, Yong Li 0032, Xinliang Zhou, Junfeng Yao, Yi Ding 0012
ACM Multimedia9
2025 Decoding olfactory response from neurophysiological signal with a multi modal deep learning framework
abstract
The human olfactory system's temporal dynamics are crucial for sensory perception. By learning the temporal dynamics of EEG and utilizing breathing signals, we aim to better understand the neural features of olfactory perception from EEG. To decode the olfactory response effectively, we introduce a new method: the Token Alignment and Cross-Attention Fusion network (TACAF), a multimodal deep learning framework that enhances olfactory EEG decoding using wavelet features for time window selection and spectral analysis for data representation. Spatial features are extracted using spatial learning modules, and temporal dynamics are captured through a multi-head self-attention mechanism. The Temporal Token Semantic Alignment (TTSA) module synchronizes breathing information with EEG data for effective fusion. We collected EEG recordings and breathing signals from 20 subjects to study the decoding responses to pleasant and unpleasant odors. Our evaluation shows that TACAF significantly outperforms existing methods. Further analysis indicates that prolonged odor exposure leads to olfactory adaptation, reducing recognition performance. The findings are visualized through spatial topology maps with saliency mappings, providing insights into the neural mechanisms of olfactory perception.
Chengxuan Tong, Yi Ding 0012, Aung Aung Phyo Wai, Hui Xin Joanna Chua, Xiaorong Wu, Kevin Junliang Lim, Cuntai Guan
Neural Networks2
2025 Beyond Overfitting: Doubly Adaptive Dropout for Generalizable AU Detection
abstract
Facial Action Units (AUs) are essential for conveying psychological states and emotional expressions. While automatic AU detection systems leveraging deep learning have progressed, they often overfit to specific datasets and individual features, limiting their cross-domain applicability. To overcome these limitations, we propose a doubly adaptive dropout approach for cross-domain AU detection, which enhances the robustness of convolutional feature maps and spatial tokens against domain shifts. This approach includes a Channel Drop Unit (CD-Unit) and a Token Drop Unit (TD-Unit), which work together to reduce domain-specific noise at both the channel and token levels. The CD-Unit preserves domain-agnostic local patterns in feature maps, while the TD-Unit helps the model identify AU relationships generalizable across domains. An auxiliary domain classifier, integrated at each layer, guides the selective omission of domain-sensitive features. To prevent excessive feature dropout, a progressive training strategy is used, allowing for selective exclusion of sensitive features at any model layer. Our method consistently outperforms existing techniques in cross-domain AU detection, as demonstrated by extensive experimental evaluations. Visualizations of attention maps also highlight clear and meaningful patterns related to both individual and combined AUs, further validating the approach's effectiveness.
Yong Li 0032, Xuesong Niu, Yi Ding 0012, Xiu-Shen Wei, Cuntai Guan
IEEE Trans. Affect. Comput.4
2025 Decoupled Doubly Contrastive Learning for Cross-Domain Facial Action Unit Detection
abstract
Despite the impressive performance of current vision-based facial action unit (AU) detection approaches, they are heavily susceptible to the variations across different domains and the cross-domain AU detection methods are under-explored. In response to this challenge, we propose a decoupled doubly contrastive adaptation (D2CA) approach to learn a purified AU representation that is semantically aligned for the source and target domains. Specifically, we decompose latent representations into AU-relevant and AU-irrelevant components, with the objective of exclusively facilitating adaptation within the AU-relevant subspace. To achieve the feature decoupling, D2CA is trained to disentangle AU and domain factors by assessing the quality of synthesized faces in cross-domain scenarios when either AU or domain attributes are modified. To further strengthen feature decoupling, particularly in scenarios with limited AU data diversity, D2CA employs a doubly contrastive learning mechanism comprising image and feature-level contrastive learning to ensure the quality of synthesized faces and mitigate feature ambiguities. This new framework leads to an automatically learned, dedicated separation of AU-relevant and domain-relevant factors, and it enables intuitive, scale-specific control of the cross-domain facial image synthesis. Extensive experiments demonstrate the efficacy of D2CA in successfully decoupling AU and domain factors, yielding visually pleasing cross-domain synthesized facial images. Meanwhile, D2CA consistently outperforms state-of-the-art cross-domain AU detection approaches, achieving an average F1 score improvement of 6%-14% across various cross-domain scenarios.
Yong Li 0032, Menglin Liu, Zhen Cui 0001, Yi Ding 0012, Yuan Zong, Wenming Zheng, Shiguang Shan, Cuntai Guan
IEEE Trans. Image Process.4
2025 EEG-Deformer: A Dense Convolutional Transformer for Brain-Computer Interfaces
abstract
Effectively learning the temporal dynamics in electroencephalogram (EEG) signals is challenging yet essential for decoding brain activities using brain-computer interfaces (BCIs). Although Transformers are popular for their long-term sequential learning ability in the BCI field, most methods combining Transformers with convolutional neural networks (CNNs) fail to capture the coarse-to-fine temporal dynamics of EEG signals. To overcome this limitation, we introduce EEG-Deformer, which incorporates two main novel components into a CNN-Transformer: (1) a Hierarchical Coarse-to-Fine Transformer (HCT) block that integrates a Fine-grained Temporal Learning (FTL) branch into Transformers, effectively discerning coarse-to-fine temporal patterns; and (2) a Dense Information Purification (DIP) module, which utilizes multi-level, purified temporal information to enhance decoding accuracy. Comprehensive experiments on three representative cognitive tasksâcognitive attention, driving fatigue, and mental workload detectionâconsistently confirm the generalizability of our proposed EEG-Deformer, demonstrating that it either outperforms or performs comparably to existing state-of-the-art methods. Visualization results show that EEG-Deformer learns from neurophysiologically meaningful brain regions for the corresponding cognitive tasks.
Yi Ding 0012, Yong Li 0032, Rui Liu 0034, Chengxuan Tong, Xinliang Zhou, Cuntai Guan
IEEE J. Biomed. Health Informatics1
2025 REI-Net: A Reference Electrode Standardization Interpolation Technique Based 3D CNN for Motor Imagery Classification
abstract
High-quality scalp EEG datasets are extremely valuable for motor imagery (MI) analysis. However, due to electrode size and montage, different datasets inevitably experience channel information loss, posing a significant challenge for MI decoding. A 2D representation that focuses on the time domain may loss the spatial information in EEG. In contrast, a 3D representation based on topography may suffer from channel loss and introduce noise through different padding methods. In this paper, we propose a framework called Reference Electrode Standardization Interpolation Network (REI-Net). Through an interpolation of 3D representation, REI-Net retains the temporal information in 2D scalp EEG while improving the spatial resolution within a certain montage. Additionally, to overcome the data variability caused by individual differences, transfer learning is employed to enhance the decoding robustness. Our approach achieves promising performance on two widely-recognized MI datasets, with an accuracy of 77.99% on BCI-C IV-2a and an accuracy of 63.94% on Kaya2018. The proposed algorithm outperforms the SOTAs leading to more accurate and robust results.
Meiyan Xu, Jie Jiao, Yi Ding 0012, Jipeng Wu, Peipei Gu, Yijie Pan, Xueping Peng, Naian Xiao, Jiayang Guo
IEEE J. Biomed. Health Informatics4
2025 EmT: A Novel Transformer for Generalized Cross-Subject EEG Emotion Recognition
abstract
Integrating prior knowledge of neurophysiology into neural network architecture enhances the performance of emotion decoding. While numerous techniques emphasize learning spatial and short-term temporal patterns, there has been a limited emphasis on capturing the vital long-term contextual information associated with emotional cognitive processes. In order to address this discrepancy, we introduce a novel transformer model called emotion transformer (EmT). EmT is designed to excel in both generalized cross-subject electroencephalography (EEG) emotion classification and regression tasks. In EmT, EEG signals are transformed into a temporal graph format, creating a sequence of EEG feature graphs using a temporal graph construction (TGC) module. A novel residual multiview pyramid graph convolutional neural network (RMPG) module is then proposed to learn dynamic graph representations for each EEG feature graph within the series, and the learned representations of each graph are fused into one token. Furthermore, we design a temporal contextual transformer (TCT) module with two types of token mixers to learn the temporal contextual information. Finally, the task-specific output (TSO) module generates the desired outputs. Experiments on four publicly available datasets show that EmT achieves higher results than the baseline methods for both EEG emotion classification and regression tasks. The code is available at https://github.com/yi-ding-cs/EmT.
Yi Ding 0012, Chengxuan Tong, Shuailei Zhang, Muyun Jiang, Yong Li 0032, Kevin Junliang Lim, Cuntai Guan
IEEE Trans. Neural Networks Learn. Syst.1
2024 Aggregating intrinsic information to enhance BCI performance through federated learning
abstract
Insufficient data is a long-standing challenge for Brain-Computer Interface (BCI) to build a high-performance deep learning model. Though numerous research groups and institutes collect a multitude of EEG datasets for the same BCI task, sharing EEG data from multiple sites is still challenging due to the heterogeneity of devices. The significance of this challenge cannot be overstated, given the critical role of data diversity in fostering model robustness. However, existing works rarely discuss this issue, predominantly centering their attention on model training within a single dataset, often in the context of inter-subject or inter-session settings. In this work, we propose a hierarchical personalized Federated Learning EEG decoding (FLEEG) framework to surmount this challenge. This innovative framework heralds a new learning paradigm for BCI, enabling datasets with disparate data formats to collaborate in the model training process. Each client is assigned a specific dataset and trains a hierarchical personalized model to manage diverse data formats and facilitate information exchange. Meanwhile, the server coordinates the training procedure to harness knowledge gleaned from all datasets, thus elevating overall performance. The framework has been evaluated in Motor Imagery (MI) classification with nine EEG datasets collected by different devices but implementing the same MI task. Results demonstrate that the proposed framework can boost classification performance up to 8.4% by enabling knowledge sharing between multiple datasets, especially for smaller datasets. Visualization results also indicate that the proposed framework can empower the local models to put a stable focus on task-related areas, yielding better performance. To the best of our knowledge, this is the first end-to-end solution to address this important challenge.
Rui Liu 0034, Yuanyuan Chen 0012, Anran Li 0001, Yi Ding 0012, Han Yu 0001, Cuntai Guan
Neural Networks4
2024 Leveraging temporal dependency for cross-subject-MI BCIs by contrastive learning and self-attention
Yi Ding 0012, Jianzhu Bao, Ke Qin, Chengxuan Tong, Jing Jin 0001, Cuntai Guan
Neural Networks2
2024 MASA-TCN: Multi-Anchor Space-Aware Temporal Convolutional Neural Networks for Continuous and Discrete EEG Emotion Recognition
abstract
Emotion recognition from electroencephalogram (EEG) signals is a critical domain in biomedical research with applications ranging from mental disorder regulation to human-computer interaction. In this paper, we address two fundamental aspects of EEG emotion recognition: continuous regression of emotional states and discrete classification of emotions. While classification methods have garnered significant attention, regression methods remain relatively under-explored. To bridge this gap, we introduce MASA-TCN, a novel unified model that leverages the spatial learning capabilities of Temporal Convolutional Networks (TCNs) for EEG emotion regression and classification tasks. The key innovation lies in the introduction of a space-aware temporal layer, which empowers TCN to capture spatial relationships among EEG electrodes, enhancing its ability to discern nuanced emotional states. Additionally, we design a multi-anchor block with attentive fusion, enabling the model to adaptively learn dynamic temporal dependencies within the EEG signals. Experiments on two publicly available datasets show that MASA-TCN achieves higher results than the state-of-the-art methods for both EEG emotion regression and classification tasks.
Yi Ding 0012, Su Zhang 0004, Chuangao Tang, Cuntai Guan
IEEE J. Biomed. Health Informatics1
2024 LGGNet: Learning From Local-Global-Graph Representations for Brain-Computer Interface
abstract
Neuropsychological studies suggest that co-operative activities among different brain functional areas drive high-level cognitive processes. To learn the brain activities within and among different functional areas of the brain, we propose local-global-graph network (LGGNet), a novel neurologically inspired graph neural network (GNN), to learn local-global-graph (LGG) representations of electroencephalography (EEG) for brain-computer interface (BCI). The input layer of LGGNet comprises a series of temporal convolutions with multiscale 1-D convolutional kernels and kernel-level attentive fusion. It captures temporal dynamics of EEG which then serves as input to the proposed local- and global-graph-filtering layers. Using a defined neurophysiologically meaningful set of local and global graphs, LGGNet models the complex relations within and among functional areas of the brain. Under the robust nested cross-validation settings, the proposed method is evaluated on three publicly available datasets for four types of cognitive classification tasks, namely the attention, fatigue, emotion, and preference classification tasks. LGGNet is compared with state-of-the-art (SOTA) methods, such as DeepConvNet, EEGNet, R2G-STNN, TSception, regularized graph neural network (RGNN), attention-based multiscale convolutional neural network-dynamical graph convolutional network (AMCNN-DGCN), hierarchical recurrent neural network (HRNN), and GraphNet. The results show that LGGNet outperforms these methods, and the improvements are statistically significant ( ) in most cases. The results show that bringing neuroscience prior knowledge into neural network design yields an improvement of classification performance. The source code can be found at https://github.com/yi-ding-cs/LGG.
Yi Ding 0012, Neethu Robinson, Chengxuan Tong, Qiuhao Zeng, Cuntai Guan
IEEE Trans. Neural Networks Learn. Syst.1
2023 TSception: Capturing Temporal Dynamics and Spatial Asymmetry From EEG for Emotion Recognition
abstract
The high temporal resolution and the asymmetric spatial activations are essential attributes of electroencephalogram (EEG) underlying emotional processes in the brain. To learn the temporal dynamics and spatial asymmetry of EEG towards accurate and generalized emotion recognition, we propose TSception, a multi-scale convolutional neural network that can classify emotions from EEG. TSception consists of dynamic temporal, asymmetric spatial, and high-level fusion layers, which learn discriminative representations in the time and channel dimensions simultaneously. The dynamic temporal layer consists of multi-scale 1D convolutional kernels whose lengths are related to the sampling rate of EEG, which learns the dynamic temporal and frequency representations of EEG. The asymmetric spatial layer takes advantage of the asymmetric EEG patterns for emotion, learning the discriminative global and hemisphere representations. The learned spatial representations will be fused by a high-level fusion layer. Using more generalized cross-validation settings, the proposed method is evaluated on two publicly available datasets DEAP and MAHNOB-HCI. The performance of the proposed network is compared with prior reported methods such as SVM, KNN, FBFgMDM, FBTSC, Unsupervised learning, DeepConvNet, ShallowConvNet, and EEGNet. TSception achieves higher classification accuracies and F1 scores than other methods in most of the experiments. The codes are available at: https://github.com/yi-ding-cs/TSception
Yi Ding 0012, Neethu Robinson, Su Zhang 0004, Qiuhao Zeng, Cuntai Guan
IEEE Trans. Affect. Comput.1
2022 TESANet: Self-attention network for olfactory EEG classification
abstract
The olfactory system is known to be associated with emotion during odor stimulation. A well-designed computational model that can correctly recognize preference induced by odor stimulation can be vital in the food and perfume industries. Electroencephalogram (EEG) can be used to study the brain's response to odor stimulation due to its good temporal resolution and low acquisition cost. In this study, we proposed a novel self-attention deep learning framework: Temporal Segment Attention Network (TESANet) to classify the brain state of subjects when they are exposed to pleasant and unpleasant odors. Odor stimulation is a continuous process, the temporal dynamics of the EEG signal should reflect the continuous changes of the brain responses to the given odor, thus we design the model to capture the intercorrelation between time segments of the EEG by utilizing the self-attention mechanism. TESANet consists of a filter-bank layer to extract spectral features, a spatial convolution layer to extract spatial features, a temporal segmentation layer to split the data into overlapping time windows, a Long Short-Term Memory (LSTM) layer to encode the temporal segments, a self-attention layer to decode the temporal dynamics by learning the intercorrelation between time segments, and finally a fully connected layer for classification. Experiments on an olfactory EEG dataset demonstrated that the proposed method outperforms other competing deep learning methods for odor pleasantness classification.
Chengxuan Tong, Yi Ding 0012, Kevin Junliang Lim, Zhuo Zhang 0001, Haihong Zhang, Cuntai Guan
IJCNN2
2020 TSception: A Deep Learning Framework for Emotion Detection Using EEG
abstract
In this paper, we propose a deep learning framework, TSception, for emotion detection from electroencephalogram (EEG). TSception consists of temporal and spatial convolutional layers, which learn discriminative representations in the time and channel domains simultaneously. The temporal learner consists of multi-scale 1D convolutional kernels whose lengths are related to the sampling rate of the EEG signal, which learns multiple temporal and frequency representations. The spatial learner takes advantage of the asymmetry property of emotion responses at the frontal brain area to learn the discriminative representations from the left and right hemispheres of the brain. In our study, a system is designed to study the emotional arousal in an immersive virtual reality (VR) environment. EEG data were collected from 18 healthy subjects using this system to evaluate the performance of the proposed deep learning network for the classification of low and high emotional arousal states. The proposed method is compared with SVM, EEGNet, and LSTM. TSception achieves a high classification accuracy of 86.03%, which outperforms the prior methods significantly (p<; 0.05).
Yi Ding 0012, Neethu Robinson, Qiuhao Zeng, Aung Aung Phyo Wai, Tih Shih Lee, Cuntai Guan
IJCNN1