VLDB 2026 Research / reviewers in the wild / expert
Ziyu Jia
dblp:256/1411
· DBLP profile ↗
55ranked-venue papers
12as first author
50since 2021 · last 2026
0000-0002-8523-1419ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 7 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 7 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DarkFarseer: Robust Spatio-Temporal Kriging Under Graph Sparsity and NoiseabstractThe rapid expansion of the Internet of Things (IoT) has created a growing demand for large-scale sensor deployment. However, the high cost of physical sensors limits the scalability and coverage of sensor networks, making fine-grained sensing difficult. Inductive Spatio-Temporal Kriging (ISK) addresses this challenge by introducing virtual sensors that infer measurements from physical sensors, typically using graph neural networks (GNNs) to model their relationships. Despite its promise, current ISK methods often rely on standard message-passing and generic architectures that fail to effectively capture spatio-temporal features or represent virtual nodes accurately. Additionally, existing graph construction techniques suffer from sparse and noisy connections, further hindering performance. To address these limitations, we propose DarkFarseer, a novel ISK framework with three key innovations. First, the Style-enhanced Temporal-Spatial architecture adopts a temporal-then-spatial processing scheme with a temporal style transfer mechanism to enhance virtual node representations. Second, Regional-semantic Contrastive Learning improves representation learning by aligning virtual nodes with regional component patterns. Third, the Similarity-Based Graph Denoising Strategy mitigates the influence of noisy edges by leveraging temporal similarity and regional structure. Extensive experiments on real-world datasets demonstrate that DarkFarseer significantly outperforms state-of-the-art ISK methods. Zhuoxuan Liang, Wei Wayne Li, Dalin Zhang 0007, Ziyu Jia, Yidan Chen, Moustafa Youssef 0001 |
AAAI | 4 |
| 2026 | Multi-view multi-scale adapter squeeze excitation network with domain generalization for cross-subject emotion recognition
Cheng Cheng 0013, Ziyu Jia, Weiqi He |
Expert Syst. Appl. | 3 |
| 2026 | MBDA: A modality-balanced framework with data augmentation and alignment for multimodal emotion recognition
Cheng Cheng 0013, Ruisi Shang, Huazhi Li, Ziyu Jia |
Neural Networks | 5 |
| 2026 | MS-STGAN: A dual-branch multi-scale spatio-temporal generative adversarial framework for incomplete EEG-based emotion recognitionabstractElectroencephalography (EEG) enables high-resolution emotion recognition but often suffers from incomplete data in real-world scenarios due to sensor failures or preprocessing errors. To this end, we propose a Multi-Scale Spatio-Temporal Generative Adversarial Network (MS-STGAN). Specifically, we first apply random masking to the EEG channels data to simulate missing data conditions in practical environments. Then, we design a spatio-temporal dual-branch generator to reconstruct complete representations: the spatial branch employs graph convolutional networks (GCNs) to capture robust inter-regional dependencies, while the temporal branch leverages the BiMamba state space model to encode the dynamic evolution of emotions. To further enhance feature learning, multi-scale 2D convolution and deconvolution layers are incorporated before and after both branches, enabling the extraction of diverse spatio-temporal features. Additionally, we introduce a generative adversarial framework, where the generator restores informative features from incomplete inputs and the discriminator enforces the authenticity of reconstructed data. Finally, a fusion module integrates the outputs of both branches for downstream classification. Extensive experiments on the DEAP and SEED-IV datasets validate the effectiveness of each component and demonstrate that MS-STGAN achieves superior performance and strong generalization ability. Cheng Cheng 0013, Yikang Cheng, Ziyu Jia, Weiqi He |
Neural Networks | 4 |
| 2026 | Dual-Branch Attention-Based Frequency Domain Network for Cross-Subject SSVEP-BCIsabstractSteady-state visual evoked potential-based brain-computer interfaces (SSVEP-BCIs) hold significant promise for enabling high-speed human-computer interaction in real-world scenarios. However, existing frequency-domain decoding methods treat frequency spectrum features (the real and imaginary spectrum features) as a single feature without considering their unique spatial and spectral characteristics, resulting in insufficient generalizable features and limited classification accuracy in cross-subject scenarios. To address this issue, we propose a Dual-Branch Attention-Based Frequency Domain Network (DB-AFDNet) to independently decode real and imaginary spectral components, aiming to acquire more discriminative and generalizable features for cross-subject applications. Specifically, we construct inter-branch attention similarity constraints to encourage the two branches to have similar attention properties, promoting to learn the consensus characteristics in the dual branches. Furthermore, we propose intra-branch orthogonality constraints to explore branch-specific discriminative features to learn generalizable features. Experimental studies on two public datasets, the Benchmark and Beta datasets, demonstrate that DB-AFDNet outperforms state-of-the-art methods in cross-subject classification, achieving a relative improvement of 1.36$\%$ and 1.45$\%$, respectively. Yi Yang 0067, Ze Wang 0001, Ziyu Jia, Boyu Wang 0004, Shangen Zhang, Chiman Wong, Xiaorong Gao, Tzyy-Ping Jung, Feng Wan 0003 |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | A Domain Adversarial Learning Framework for Major Depression Disorder DiagnosisabstractThe prompt recognition and timely intervention for depression are vital for attaining the most favorable therapeutic results. Notwithstanding its widespread occurrence, depression continues to be poorly understood within both clinical and research frameworks. Recent progress in deep learning methodologies, especially with the application of convolutional neural networks (CNNs) and recurrent neural networks (RNNs), has significantly improved the diagnostic capabilities for Major Depressive Disorder (MDD). However, CNNs and RNNs are limited in extracting essential brain spatial-temporal data during the classification process of MDD. Moreover, cross-subject variability poses an additional challenge in MDD classification. The proposed Domain Adversarial Learning Framework for Major Depression Disorder Diagnosis (MDD-DAL) overcomes the limitations of CNNs and RNNs by incorporating a spatial-temporal transformer architecture that can extract crucial brain spatial-temporal features from electroencephalogram (EEG) signals. Additionally, domain adversarial learning techniques enhance the model’s generalizability across variations in EEG data from different subjects. The model is validated using an MDD diagnosis dataset, achieving outstanding classification performance that showcases its advanced diagnostic capability. Our code is publicly accessible at: https://github.com/shaozheliu/MDD-DAL. Shaozhe Liu, Leike An, Ziyu Jia |
ICASSP | 3 |
| 2025 | SleepSMC: Ubiquitous Sleep Staging via Supervised Multimodal CoordinationabstractSleep staging is critical for assessing sleep quality and tracking health. Polysomnography (PSG) provides comprehensive multimodal sleep-related information, but its complexity and impracticality limit its practical use in daily and ubiquitous monitoring. Conversely, unimodal devices offer more convenience but less accuracy. Existing multimodal learning paradigms typically assume that the data types remain consistent between the training and testing phases. This makes it challenging to leverage information from other modalities in ubiquitous scenarios (e.g., at home) where only one modality is available. To address this issue, we introduce a novel framework for ubiquitous Sleep staging via Supervised Multimodal Coordination, called SleepSMC. To capture category-related consistency and complementarity across modality-level instances, we propose supervised modality-level instance contrastive coordination. Specifically, modality-level instances within the same category are considered positive pairs, while those from different categories are considered negative pairs. To explore the varying reliability of auxiliary modalities, we calculate uncertainty estimates based on the variance in confidence scores for correct predictions during multiple rounds of random masks. These uncertainty estimates are employed to assign adaptive weights to multiple auxiliary modalities during contrastive learning, ensuring that the primary modality learns from high-quality, category-related features. Experimental results on four public datasets, ISRUC-S3, MASS-SS3, Sleep-EDF-78, and ISRUC-S1, show that SleepSMC achieves state-of-the-art cross-subject performance. SleepSMC significantly improves performance when only one modality is present during testing, making it suitable for ubiquitous sleep monitoring. Shuo Ma 0001, Yingwei Zhang 0002, Yiqiang Chen 0001, Hualei Wang, Wei Zhang 0082, Ziyu Jia |
ICLR | 7 |
| 2025 | ST-USleepNet: A Spatial-Temporal Coupling Prominence Network for Multi-Channel Sleep StagingabstractSleep staging is critical to assess sleep quality and diagnose disorders. Despite advancements in artificial intelligence enabling automated sleep staging, significant challenges remain: (1) Simultaneously extracting prominent temporal and spatial sleep features from multi-channel raw signals, including characteristic sleep waveforms and salient spatial brain networks. (2) Capturing the spatial-temporal coupling patterns essential for accurate sleep staging. To address these challenges, we propose a novel framework named ST-USleepNet, comprising a spatial-temporal graph construction module (ST) and a U-shaped sleep network (USleepNet). The ST module converts raw signals into a spatial-temporal graph based on signal similarity, temporal, and spatial relationships to model spatial-temporal coupling patterns. The USleepNet employs a U-shaped structure for both the temporal and spatial streams, mirroring its original use in image segmentation to isolate significant targets. Applied to raw sleep signals and graph data from the ST module, USleepNet effectively segments these inputs, simultaneously extracting prominent temporal and spatial sleep features. Testing on three datasets demonstrates that ST-USleepNet outperforms existing baselines, and model visualizations confirm its efficacy in extracting prominent sleep features and temporal-spatial coupling patterns across various sleep stages. The code is available at https://github.com/Majy-Yuji/ST-USleepNet. Jingying Ma, Qika Lin, Ziyu Jia, Mengling Feng |
IJCAI | 3 |
| 2025 | A Cross-Modal Densely Guided Knowledge Distillation Based on Modality Rebalancing Strategy for Enhanced Unimodal Emotion RecognitionabstractMultimodal emotion recognition has garnered significant attention for its ability to integrate data from multiple modalities to enhance performance. However, physiological signals like electroencephalogram are more challenging to acquire than visual data due to higher collection costs and complexity. This limits the practical application of multimodal networks. To address this issue, this paper proposes a cross-modal knowledge distillation framework for emotion recognition. The framework aims to leverage the strengths of a multimodal teacher network to enhance the performance of a unimodal student network using only the visual modality as input. Specifically, we design a prototype-based modality rebalancing strategy, which dynamically adjusts the convergence rates of different modalities to mitigate modality imbalance issue. It enables the teacher network to better integrate multimodal information. Building upon this, we develop a Cross-Modal Densely Guided Knowledge Distillation (CDGKD) method, which effectively transfers knowledge extracted by the multimodal teacher network to the unimodal student network. Our CDGKD uses multi-level teacher assistant networks to bridge the teacher-student gap and employs dense guidance to reduce error accumulation during knowledge transfer. Experimental results demonstrate that the proposed framework outperforms existing methods on two public emotion datasets, providing an effective solution for emotion recognition in modality-constrained scenarios. Heng Liang, Ziyu Jia |
IJCAI | 5 |
| 2025 | Sera: Separated Coarse-to-fine Representation Alignment for Cross-subject EEG-based Emotion RecognitionabstractNeuropsychology-inspired models have been utilized in recent advances in EEG emotion recognition, such as convolutional networks for spatial features and Transformers for temporal dependencies. While these methods benefit from domain knowledge like frequency-band features and spatial correlations, most overlook the fundamental fact that EEG signals are complex mixtures of neural source activities recorded at the scalp. EEG signals presenting challenges for emotion recognition, particularly in cross-subject scenarios due to significant inter-subject variance. Inspired by neurophysiological principles, we propose a novel framework, named Sera, for EEG-based emotion recognition that explicitly separates source activities and aligns representations across subjects. Sera introduces two key components: (1) a variational autoencoder (VAE) with multiple multi-stage decoders (M2VAE) designed to disentangle EEG signals into independent sources, mimicking the neural generation process, and (2) a coarse-to-fine representation alignment block (CFRA) to mitigate subject-to-subject variability. The coarse alignment employs adversarial training with a domain discriminator, while the fine-grained alignment matches covariance matrices to capture temporal correlations within EEG segments. Extensive experiments demonstrate that Sera outperforms the state-of-the-art methods with improvements ranging from 1% to 5%, averaging 3.14% and 3.05% on the DEAP and DREAMER datasets, respectively, confirming its effectiveness and neurophysiological grounding. The code is available at: https://github.com/JZH98/Sera-code. Meiyan Xu, Ziyu Jia, Yong Li 0032, Xinliang Zhou, Junfeng Yao, Yi Ding 0012 |
ACM Multimedia | 4 |
| 2025 | A Multimodal BiMamba Network with Test-Time Adaptation for Emotion Recognition Based on Physiological SignalsabstractEmotion recognition based on physiological signals plays a vital role in psychological health and human–computer interaction, particularly with the substantial advances in multimodal emotion recognition techniques. However, two key challenges remain unresolved: 1) how to effectively model the intra-modal long-range dependencies and inter-modal correlations in multimodal physiological emotion signals, and 2) how to address the performance limitations resulting from missing multimodal data. In this paper, we propose a multimodal bidirectional Mamba (BiMamba) network with test-time adaptation (TTA) for emotion recognition named BiM-TTA. Specifically, BiM-TTA consists of a multimodal BiMamba network and a multimodal TTA. The former includes intra-modal and inter-modal BiMamba modules, which model long-range dependencies along the time dimension and capture cross-modal correlations along the channel dimension, respectively. The latter (TTA) mitigates the amplified distribution shifts caused by missing multimodal data through two-level entropy-based sample filtering and mutual information sharing across modalities. By addressing these challenges, BiM-TTA achieves state-of-the-art results on two multimodal emotion datasets. Ziyu Jia, Tingyu Du, Zhengyu Tian, Yong Zhang 0030 |
NeurIPS | 1 |
| 2025 | REFED: A Subject Real-time Dynamic Labeled EEG-fNIRS Synchronized Recorded Emotion DatasetabstractAffective brain-computer interfaces (aBCIs) play a crucial role in personalized human–computer interaction and neurofeedback modulation. To develop practical and effective aBCI paradigms and to investigate the spatial-temporal dynamics of brain activity under emotional inducement, portable electroencephalography (EEG) signals have been widely adopted. To further enhance spatial-temporal perception, functional near-infrared spectroscopy (fNIRS) has attracted increasing interest in the aBCI field and has been explored in combination with EEG. However, existing datasets typically provide only static fixation labels, overlooking the dynamic changes in subjects' emotions. Notably, some studies have attempted to collect continuously annotated emotional data, but they have recorded only peripheral physiological signals without directly observing brain activity, limiting insight into underlying neural states under different emotions. To address these challenges, we present the Real-time labeled EEG-fNIRS Dataset (REFED). To the best of our knowledge, this is the first EEG-fNIRS dataset with real-time dynamic emotional annotations. REFED simultaneously records brain signals from both EEG and fNIRS modalities while providing continuous, real-time annotations of valence and arousal. The results of the data analysis demonstrate the effectiveness of emotion inducement and the reliability of real-time annotation. This dataset offers the possibility for studying the neurovascular coupling mechanism under emotional evolution and for developing dynamic, robust affective BCIs. Xiaojun Ning 0001, Jing Wang 0060, Zhiyang Feng, Tianzuo Xin, Shuo Zhang 0015, Shaoqi Zhang, Youfang Lin, Ziyu Jia |
NeurIPS | 10 |
| 2025 | A universal sampling method based on feature and structural comprehensive proximity measure
Jinhui Pang, Cheng Shang, Ziyu Jia, Peng Hao 0003, Xiaoshuai Hao |
Neurocomputing | 3 |
| 2025 | MPFBL: Modal pairing-based cross-fusion bootstrap learning for multimodal emotion recognition
Yong Zhang 0030, Cheng Cheng 0013, Ziyu Jia |
Neurocomputing | 5 |
| 2025 | CCAM: Cross-Channel Association Mining for Ubiquitous Sleep StagingabstractAccurate sleep staging is crucial for wearable sensor-based sleep monitoring and health interventions. Polysomnography (PSG) signals, rich in information from multiple synchronous sensor channels, are frequently utilized in sleep studies due to their high accuracy in sleep staging. However, employing a wearable PSG for sleep monitoring is uncomfortable, complex, expensive, and impractical. The advent of single-channel wearable electroencephalography (EEG) devices enables comfortable sleep monitoring in ubiquitous scenarios (e.g., home). Despite their advantages, such devices lack sufficient information and have severe accuracy limitations. To improve the accuracy of single-channel EEG-based sleep staging, we propose a method called cross-channel association mining (CCAM). Besides the naive feature extraction, CCAM leverages synchronous single-channel EEG and high-density PSG signals to establish pairwise association feature mining models. These association models can transfer information from other auxiliary PSG channels to the target EEG channel. In inference, CCAM exploits association feature mining models to extract effective features from the target EEG channel and improve sleep staging accuracy through directionally multiview information fusion. Experimental results on three public datasets for sleep staging, ISRUC-S1, ISRUC-S3, and sleep heart health study, show that CCAM outperforms state-of-the-art methods, advancing the field of wearable sensor-based sleep monitoring in ubiquitous scenarios. Shuo Ma 0001, Yingwei Zhang 0002, Yiqiang Chen 0001, Weiwen Yang, Jianrong Yang, Ziyu Jia |
IEEE Internet Things J. | 8 |
| 2025 | Two-Stream Dynamic Heterogeneous Graph Recurrent Neural Network for Multi-Label Multi-Modal Emotion RecognitionabstractThe study of the relationship between emotions and physiological signals of subjects under multimedia stimulation is an emerging field, and many important advances are made. However, there are still some challenges: 1) How to effectively utilize the complementarity among spatial-spectral-temporal domain information. 2) How to employ the heterogeneity and the correlation among multi-modal physiological signals simultaneously. 3) How to improve the robustness of the model dealing with missing channels. 4) How to model the dependency among different emotions. In this paper, we propose a novel two-stream Dynamic Heterogeneous Graph Recurrent Neural Network called DHGRNN. Specifically, DHGRNN consists of a spatial-temporal stream, a spatial-spectral stream, a fusion layer, and a multi-label classifier. Each stream is composed of a graph transformer network, evolved graph convolutional neural network, and gated recurrent units. We propose a graph-based two-stream structure to fuse the information of the spatial-spectral-temporal domain simultaneously. Graph transformer network and evolved graph convolutional neural network are used to model the heterogeneity and correlation of multi-modal physiological signals, respectively. To deal with the problem of robustness in the face of missing channel data, we transform it into the problem of dynamic graphs and use a dynamic graph neural network to improve the robustness. In addition, we propose a multi-label classifier to model the dependency among different emotion dimensions. Experiments on three public datasets demonstrate that our proposed model outperforms existing state-of-the-art methods. Jing Wang 0060, Zhiyang Feng, Xiaojun Ning 0001, Youfang Lin, Badong Chen, Ziyu Jia |
IEEE Trans. Affect. Comput. | 6 |
| 2025 | Exploiting the Intrinsic Neighborhood Semantic Structure for Domain Adaptation in EEG-Based Emotion RecognitionabstractDue to the inherent non-stationarity and individual differences present in electroencephalogram (EEG) signals, developing a generalizable model that performs well on new subjects is challenging in EEG-based emotion recognition. Most existing domain adaptation (DA) methods typically mitigate these discrepancies by aligning the marginal distributions of domain feature representations. However, when there is a significant difference in the class-conditional distribution between domain features and labels, the domain-invariant features learned by aligning marginal distributions may have limited discriminative ability for unlabeled target instances or even prove counterproductive. To address this issue, we propose a Neighborhood Semantic Aware Learning-based Dynamic Graph Attention Convolution (NSAL-DGAT) approach that learns target semantic information by considering the inter-domain semantic topological structure, thereby improving classifier adaptation for target instances. Specifically, the proposed NSAL framework is designed to capitalize on the insight that after domain feature alignment, some target samples and their neighboring source samples exhibit similar semantics. By leveraging the neighborhood topological structure, we extract and incorporate semantic target features to train a more transferable classifier. Besides, we implement an entropy weighting mechanism to emphasize representative target semantic information, encouraging target instances to prioritize high-confidence individuals within the source neighborhood. We have conducted extensive experiments on the public SEED dataset and our collected the Hearing-Impaired EEG Dataset (HIED). The experimental results underscore the efficacy of our proposed NSAL-DGAT approach, showcasing state-of-the-art accuracy in subject-dependent as well as subject-independent scenarios. The source code is available at https://github.com/YYingDL/NSAL-DGAT. Yi Yang 0067, Ze Wang 0001, Yu Song 0004, Ziyu Jia, Boyu Wang 0004, Tzyy-Ping Jung, Feng Wan 0003 |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | DistillSleepNet: Heterogeneous Multi-Level Knowledge Distillation via Teacher Assistant for Sleep StagingabstractAccurate sleep staging is crucial for the diagnosis of diseases such as sleep disorders. Existing sleep staging models with excellent performance are usually large and require a lot of computational resources, limiting their application on wearable devices. Therefore, it is a key issue to distil the knowledge embedded in large models into small heterogeneous models for better deployment. In the process of knowledge distillation of heterogeneous models for sleep electroencephalography (EEG) signals, we mainly deal with three major challenges: 1) There are large structural differences between heterogeneous sleep staging models; 2) What kind of knowledge should be conveyed in sleep EEG signals in the knowledge distillation of heterogeneous models; 3) Significant scale differences exist between heterogeneous models. To address these challenges, we design a generic heterogeneous model knowledge distillation framework for sleep staging. Specifically, we first propose a knowledge distillation strategy for heterogeneous models that addresses the large structural differences between heterogeneous models. Then, a multi-level knowledge distillation module is designed to effectively transfer important multi-level feature knowledge. In addition, the teacher assistant module is introduced to ease the scale difference between the heterogeneous models which further enhances the knowledge distillation performance. Experimental results on both Sleep-EDF and ISRUC datasets show that our distillation framework achieves state-of-the-art performance. Ziyu Jia, Heng Liang, Tianzi Jiang |
IEEE Trans. Big Data | 1 |
| 2025 | A Cross-Modal Adaptive Masked Autoencoder for Decoding Emotions With Multimodal DataabstractMultimodal emotion recognition (MER) has recently gained much attention since it can leverage information over multiple modalities. However, in real life, we often encounter the problem of missing modalities, as well as modeling the heterogeneity and correlation among multimodal data are challenges. To this end, we propose a unified model called cross-modal adaptive masked autoencoder (CMA-MAE) for incomplete multimodal learning. Our CMA-MAE model comprises a cross-modal adaptive fusion encoder (CMAFE) and a multiview adaptive encoder (MVAE) to capture and fuse the heterogeneity and correlation among multimodal features. Additionally, we design a convolutional decoder that progressive upsampling and fusion with the modality-invariant features to generate robust emotional features from partially observable data. To effectively utilize both data with complete and incomplete modalities for feature learning, we adopt an end-to-end approach that simultaneously optimizes classification and reconstruction tasks. Extensive testing on the DEAP and SEED-IV datasets is conducted to assess our model, with the findings demonstrating that our CMA-MAE model outperforms current leading approaches in both incomplete and complete multimodal learning scenarios. Cheng Cheng 0013, Yong Zhang 0030, Lin Feng 0001, Ziyu Jia |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2025 | HATNet: EEG-Based Hybrid Attention Transfer Learning Network for Train Driver State DetectionabstractElectroencephalography (EEG) is widely utilized for train driver state detection due to its high accuracy and low latency. However, existing methods for driver status detection rarely use the rich physiological information in EEG to improve detection performance. Moreover, there is currently a lack of EEG datasets for abnormal states of train drivers. To address these gaps, we propose a novel transfer learning model based on a hybrid attention mechanism, named hybrid attention-based transfer learning network (HATNet). We first segment the EEG signals into patches and utilize the hybrid attention module to capture local and global temporal patterns. Then, a channel-wise attention module is introduced to establish spatial representations among EEG channels. Finally, during the training process, we employ a calibration-based transfer learning strategy, which allows for adaptation to the EEG data distribution of new subjects using minimal data. To validate the effectiveness of our proposed model, we conduct a multistimulus oddball experiment to establish a EEG dataset of abnormal states for train drivers. Experimental results on this dataset indicate that: 1) Compared to the state-of-the-art end-to-end models, HATNet achieves the highest classification accuracy in both subject-dependent and subject-independent tasks at 94.26% and 87.03%, respectively, and 2) The proposed hybrid attention module effectively captures the temporal semantic information of EEG data. Shuxiang Lin, Chaojie Fan, Demin Han, Ziyu Jia, Yong Peng 0002, Sam Kwong |
IEEE Trans. Cybern. | 4 |
| 2025 | DynSeizureGAT: Multi-Band Dynamic Graph Attention Network for Interpretable Seizure Detection and Analysis of Drug-Resistant Epilepsy Using SEEGabstractThe dynamic propagation of epileptic discharges complicates Drug-Resistant Epilepsy (DRE) seizure detection using traditional machine learning methods and Stereotactic Electroencephalography (SEEG). Several challenges remain unresolved in prior studies: (1) incomprehensive representations of epileptic brain network features; (2) lacking of flexible and dynamic mechanisms to learn brain network evolving features; and (3) the absence of model mechanisms interpretation corresponds with seizure mechanisms. In response, we propose a novel multi-band dynamic graph attention network, DynSeizureGAT, to detect and analyze DRE seizures with precision and interpretability. Specifically, a seizure network sequence is first constructed by integrating a multi-band directed transfer function matrix and enhanced epileptic index node features. Second, a dynamic graph attention module is integrated to dynamically weigh the contribution of various spatial scales. Third, spatial-spectral-temporal attention mechanisms enhance the model's capacity to better characterize and interpret the ictal and interictal states. Extensive experiments are conducted on the large-scale public clinical SEEG dataset (OpenNeuro). The proposed model demonstrates high seizure detection performance, achieving an average of 94.6% accuracy, 93.4% sensitivity, and 96.4% specificity. In addition, the importance of frequency bands and dynamic abnormal connectivity patterns is successfully quantified and visualized, which contributes most to the explainability. Experimental results indicate that DynSeizureGAT demonstrates strong dynamic propagation feature learning capability, corresponding with seizure propagation mechanisms, and is promising to assist DRE epileptogenic zone localization. Yiping Wang 0002, Jinjie Guo, Ziyu Jia, Gongpeng Cao, Guixia Kang, Jinguo Huang |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Subject-Adaptation Salient Wave Detection Network for Multimodal Sleep Stage ClassificationabstractSleep stage classification is an important step in the diagnosis and treatment of sleep disorders. Despite the high classification performance of previous sleep stage classification work, some challenges remain unresolved: 1) How to effectively capture salient waves in sleep signals to improve sleep stage classification results. 2) How to capture salient waves affected by inter-subject variability. 3) How to adaptively regulate the importance of different modals for different sleep stages. To address these challenges, we propose SleepWaveNet, a multimodal salient wave detection network, which is motivated by the salient object detection task in computer vision. It has a U-Transformer structure to detect salient waves in sleep signals. Meanwhile, the subject-adaptation wave extraction architecture based on transfer learning can adapt to the information of target individuals and extract salient waves with inter-subject variability. In addition, the multimodal attention module can adaptively enhance the importance of specific modal data for sleep stage classification tasks. Experiments on three datasets show that SleepWaveNet has better overall performance than existing baselines. Moreover, visualization experiments show that the model has the ability to capture salient waves with inter-subject variability. Jing Wang 0060, Xiaojun Ning 0001, Youfang Lin, Huy Phan, Ziyu Jia |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | DISD-Net: A Dynamic Interactive Network With Self-Distillation for Cross-Subject Multi-Modal Emotion RecognitionabstractMulti-modal Emotion Recognition (MER) has demonstrated competitive performance in affective computing, owing to synthesizing information from diverse modalities. However, many existing approaches still face unresolved challenges, such as: (i) how to learn compact yet representative features from multi-modal data simultaneously and (ii) how to address differences among subjects and enhance the generalization of the emotion recognition model, given the diverse nature of individual biological signals. To this end, we propose a Dynamic Interactive Network with Self-Distillation (DISD-Net) for cross-subject MER. The DISD-Net incorporates a dynamin interactive module to capture the intra- and inter-modal interactions from multi-modal data. Additionally, to enhance compactness in modal representations, we leverage the soft labels generated by the DISD-Net model as supplemental training guidance. This involves incorporating self-distillation, aiming to transfer the knowledge that the DISD-Net model contains hard and soft labels to each modality. Finally, domain adaptation (DA) is seamlessly integrated into the dynamic interactive and self-distillation components, forming a unified framework to extract subject-invariant multi-modal emotional features. Experimental results indicate that the proposed model achieves a mean accuracy of 75.00% with a standard deviation of 7.68% for the DEAP dataset and a mean accuracy of 65.65% with a standard deviation of 5.08% for the SEED-IV dataset. Cheng Cheng 0013, Xinying Wang 0005, Lin Feng 0001, Ziyu Jia |
IEEE Trans. Multim. | 5 |
| 2024 | A Multimodal Knowledge Distillation Framework for Sleep Physiological Data
Zongting Xie, Heng Liang, Ziyu Jia |
ADMA (4) | 3 |
| 2024 | Spatial-Temporal Mamba Network for EEG-Based Motor Imagery Classification
Xiaoxiao Yang, Ziyu Jia |
ADMA (3) | 2 |
| 2024 | Wavelet Multi-level Multi-scale Neural Network for EEG ClassificationabstractElectroencephalogram (EEG) plays a crucial role in Brain-computer interface (BCI) research by recording brain electrical activity, allowing direct communication with external devices without relying on muscles or peripheral nerves. EEG signals are non-invasive and low-cost, which are commonly used in BCIs. Although existing work has achieved success in many BCI paradigms, there is still a significant challenge: how to automatically extract spectral features from raw EEG signals. To address this, we propose a novel wavelet multi-level multi-scale neural network for EEG-based BCIs, called wavelet-MMNN. Our model uses wavelet decomposition with multi-scale convolutional kernels to automatically extract abundant spectral features from the raw EEG signals. The cross-subject experiments and cross-session experiments are conducted on SSVEP Benchmark dataset, and the results indicate that the proposed model is superior to the state-of-the-art models. Zhenghao Zhou, Yisu Wang, Fengming Zhao, Ziyu Jia |
BIBM | 4 |
| 2024 | VBH-GNN: Variational Bayesian Heterogeneous Graph Neural Networks for Cross-subject Emotion RecognitionabstractThe research on human emotion under electroencephalogram (EEG) is an emerging field in which cross-subject emotion recognition (ER) is a promising but challenging task. Many approaches attempt to find emotionally relevant domain-invariant features using domain adaptation (DA) to improve the accuracy of cross-subject ER. However, two problems still exist with these methods. First, only single-modal data (EEG) is utilized, ignoring the complementarity between multi-modal physiological signals. Second, these methods aim to completely match the signal features between different domains, which is difficult due to the extreme individual differences of EEG. To solve these problems, we introduce the complementarity of multi-modal physiological signals and propose a new method for cross-subject ER that does not align the distribution of signal features but rather the distribution of spatio-temporal relationships between features. We design a Variational Bayesian Heterogeneous Graph Neural Network (VBH-GNN) with Relationship Distribution Adaptation (RDA). The RDA first aligns the domains by expressing the model space as a posterior distribution of a heterogeneous graph for a given source domain. Then, the RDA transforms the heterogeneous graph into an emotion-specific graph to further align the domains for the downstream ER task. Extensive experiments on two public datasets, DEAP and Dreamer, show that our VBH-GNN outperforms state-of-the-art methods in cross-subject scenarios. Xinliang Zhou, Zhengri Zhu, Liming Zhai, Ziyu Jia, Yang Liu 0003 |
ICLR | 5 |
| 2024 | ATTA: Adaptive Test-Time Adaptation for Multi-Modal Sleep Stage Classification
Ziyu Jia, Xihao Yang, Haoyang Deng, Tianzi Jiang |
IJCAI | 1 |
| 2024 | Multi-level Disentangling Network for Cross-Subject Emotion Recognition Based on Multimodal Physiological Signals
Ziyu Jia, Fengming Zhao, Yuzhe Guo, Hairong Chen, Tianzi Jiang |
IJCAI | 1 |
| 2024 | VSGT: Variational Spatial and Gaussian Temporal Graph Models for EEG-based Emotion Recognition
Xinliang Zhou, Jiaping Xiao, Zhengri Zhu, Liming Zhai, Ziyu Jia, Yang Liu 0003 |
IJCAI | 6 |
| 2024 | SDformer: Transformer with Spectral Filter and Dynamic Attention for Multivariate Time Series Long-term Forecasting
Gengyu Lyu, Yiming Huang 0002, Ziyu Jia, Zhen Yang 0004 |
IJCAI | 5 |
| 2024 | Mutual Distillation Extracting Spatial-temporal Knowledge for Lightweight Multi-channel Sleep Stage ClassificationabstractSleep stage classification has important clinical significance for the diagnosis of sleep-related diseases. To pursue more accurate sleep stage classification, multi-channel sleep signals are widely used due to the rich spatial-temporal information contained. However, it leads to a great increment in the size and computational costs, which constrain the application of multi-channel sleep models on hardware devices. Knowledge distillation is an effective way to compress models, yet existing knowledge distillation methods cannot fully extract and transfer the spatial-temporal knowledge in the multi-channel sleep signals. To solve the problem, we propose a general knowledge distillation framework for multi-channel sleep stage classification called spatial-temporal mutual distillation. Based on the spatial relationship of human body and the temporal transition rules of sleep signals, the spatial and temporal modules are designed to extract the spatial-temporal knowledge, thus help the lightweight student model learn the rich spatial-temporal knowledge from large-scale teacher model. The mutual distillation framework transfers the spatial-temporal knowledge mutually. Teacher model and student model can learn from each other, further improving the student model. The results on the ISRUC-III and MASS-SS3 datasets show that our proposed framework compresses the sleep models effectively with minimal performance loss and achieves the state-of-the-art performance compared to the baseline methods. Ziyu Jia, Tianzi Jiang |
KDD | 1 |
| 2024 | SleepMG: Multimodal Generalizable Sleep Staging with Inter-modal Balance of Classification and Domain Discrimination
Shuo Ma 0001, Yingwei Zhang 0002, Yiqiang Chen 0001, Haoran Wang 0012, Ziyu Jia |
ACM Multimedia | 6 |
| 2024 | A novel transformer autoencoder for multi-modal emotion recognition with incomplete data
Cheng Cheng 0013, Zhaoxin Fan, Lin Feng 0001, Ziyu Jia |
Neural Networks | 5 |
| 2024 | Emotion recognition using hierarchical spatial-temporal learning transformer from regional to global brain
Cheng Cheng 0013, Lin Feng 0001, Ziyu Jia |
Neural Networks | 4 |
| 2024 | Multi-source Selective Graph Domain Adaptation Network for cross-subject EEG emotion recognition
Jing Wang 0060, Xiaojun Ning 0001, Yunze Li, Ziyu Jia, Youfang Lin |
Neural Networks | 5 |
| 2024 | Spectral-Spatial Attention Alignment for Multi-Source Domain Adaptation in EEG-Based Emotion RecognitionabstractIn electroencephalographic-based (EEG-based) emotion recognition, high non-stationarity and individual differences in EEG signals could lead to significant discrepancies between sessions/subjects, making generalization to a new session/subject very difficult. Most existing domain adaptation (DA) and multi-source domain adaptation (MSDA) techniques aim to mitigate this discrepancy by aligning feature distributions. However, when confronted with many diverse domain distributions, learning domain-invariant features via aligning pairwise feature distributions between domains can be hard or even counterproductive. To address this issue, this article proposes an attention alignment approach to learning abundant domain-invariant features. The motivation is simple: despite individual differences causing significant differences in feature distributions in EEG-based emotion recognition, shared affective cognitive attributes (attention) of spectral and spatial domains can be observed within the same emotion categories. The proposed spectral-spatial attention alignment multi-source domain adaptation (S2A2-MSDA) constructs domain attention to represent affective cognition attributes in spatial and spectral domains and utilizes domain consistent loss to align them between domains. Furthermore, to facilitate discriminative feature learning on the target classes, S2A2-MSDA learns the conditional semantic information of the target domain using a pseudo-labeling method. This algorithm has been validated on the SEED and SEED-IV datasets in cross-session and cross-subject scenarios, respectively. Experimental results demonstrate that S2A2-MSDA outperforms existing representative DA and MSDA methods, achieving state-of-the-art performance. Yi Yang 0067, Ze Wang 0001, Xucheng Liu, Ziyu Jia, Boyu Wang 0004, Feng Wan 0003 |
IEEE Trans. Affect. Comput. | 5 |
| 2024 | Dense Graph Convolutional With Joint Cross-Attention Network for Multimodal Emotion RecognitionabstractMultimodal emotion recognition (MER) has attracted much attention since it can leverage consistency and complementary relationships across multiple modalities. However, previous studies mostly focused on the complementary information of multimodal signals, neglecting the consistency information of multimodal signals and the topological structure of each modality. To this end, we propose a dense graph convolution network (DGC) equipped with a joint cross attention (JCA), named DG-JCA, for MER. The main advantage of the DG-JCA model is that it simultaneously integrates the spatial topology, consistency, and complementarity of multimodal data into a unified network framework. Meanwhile, DG-JCA extends the graph convolution network (GCN) via a dense connection strategy and introduces cross attention to joint model well-learned features from multiple modalities. Specifically, we first build a topology graph for each modality and then extract neighborhood features of different modalities using DGC driven by dense connections with multiple layers. Next, JCA performs cross-attention fusion in intra- and intermodality based on each modality's characteristics while balancing the contributions of various modalities’ features. Finally, subject-dependent and subject-independent experiments on the DEAP and SEED-IV datasets are conducted to evaluate the proposed method. Abundant experimental results show that the proposed model can effectively extract and fuse multimodal features and achieve outstanding performance in comparison with some state-of-the-art approaches. Cheng Cheng 0013, Lin Feng 0001, Ziyu Jia |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | Exploring Structure Incentive Domain Adversarial Learning for Generalizable Sleep Stage ClassificationabstractSleep stage classification is crucial for sleep state monitoring and health interventions. In accordance with the standards prescribed by the American Academy of Sleep Medicine, a sleep episode follows a specific structure comprising five distinctive sleep stages that collectively form a sleep cycle. Typically, this cycle repeats about five times, providing an insightful portrayal of the subject’s physiological attributes. The progress of deep learning and advanced domain generalization methods allows automatic and even adaptive sleep stage classification. However, applying models trained with visible subject data to invisible subject data remains challenging due to significant individual differences among subjects. Motivated by the periodic category-complete structure of sleep stage classification, we propose a Structure Incentive Domain Adversarial learning (SIDA) method that combines the sleep stage classification method with domain generalization to enable cross-subject sleep stage classification. SIDA includes individual domain discriminators for each sleep stage category to decouple subject dependence differences among different categories and fine-grained learning of domain-invariant features. Furthermore, SIDA directly connects the label classifier and domain discriminators to promote the training process. Experiments on three benchmark sleep stage classification datasets demonstrate that the proposed SIDA method outperforms other state-of-the-art sleep stage classification and domain generalization methods and achieves the best cross-subject sleep stage classification results. Shuo Ma 0001, Yingwei Zhang 0002, Yiqiang Chen 0001, Shuchao Song, Ziyu Jia |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2023 | Exploiting Interactivity and Heterogeneity for Sleep Stage Classification Via Heterogeneous Graph Neural NetworkabstractSleep stage classification based on physiological time-series is essential for sleep quality evaluation and the diagnosis of sleep disorders in clinical practice. Existing machine learning studies have achieved adequate results in sleep stage classification. However, those methods neglect the significance of simultaneously capturing the interactivity and heterogeneity of physiological signals. In this paper, we propose a novel Sleep Heterogeneous Graph Neural Network (SleepHGNN) to employ these essential features. The SleepHGNN is a deep graph network consisting of Heterogeneous Graph Transformer layers, which are composed of a Heterogeneous Message Passing module for capturing the heterogeneity and a Target-Specific Aggregation module for capturing the interactivity of physiological signals. The experiments show that the SleepHGNN outperforms the state-of-the-art models on the sleep stage classification task. The source code of SleepHGNN is available at: https://github.com/zhouyh310/SleepHGNN. Ziyu Jia, Youfang Lin, Xiyang Cai, Jing Wang 0060 |
ICASSP | 1 |
| 2023 | BSTT: A Bayesian Spatial-Temporal Transformer for Sleep Staging
Ziyu Jia |
ICLR | 2 |
| 2023 | Teacher Assistant-Based Knowledge Distillation Extracting Multi-level Features on Single Channel Sleep EEGabstractSleep stage classification is of great significance to the diagnosis of sleep disorders. However, existing sleep stage classification models based on deep learning are usually relatively large in size (wider and deeper), which makes them hard to be deployed on wearable devices. Therefore, it is a challenge to lighten the existing sleep stage classification models. In this paper, we propose a novel general knowledge distillation framework for sleep stage classification tasks called SleepKD. Our SleepKD, composed of the multi-level module, teacher assistant module, and other knowledge distillation modules, aims to lighten large-scale sleep stage classification models. Specifically, the multi-level module is able to transfer the multi-level knowledge extracted from sleep signals by the teacher model (large-scale model) to the student model (lightweight model). Moreover, the teacher assistant module bridges the large gap between the teacher and student network, and further improves the distillation. We evaluate our method on two public sleep datasets (Sleep-EDF and ISRUC-III). Compared to the baseline methods, the results show that our knowledge distillation framework achieves state-of-the-art performance. SleepKD can significantly lighten the sleep model while maintaining its classification performance. The source code is available at https://github.com/HychaoWang/SleepKD. Heng Liang, Ziyu Jia |
IJCAI | 4 |
| 2023 | ASTDF-Net: Attention-Based Spatial-Temporal Dual-Stream Fusion Network for EEG-Based Emotion RecognitionabstractEmotion recognition based on electroencephalography (EEG) has attracted significant attention and achieved considerable advances in the fields of affective computing and human-computer interaction. However, most existing studies ignore the coupling and complementarity of complex spatiotemporal patterns in EEG signals. Moreover, how to exploit and fuse crucial discriminative aspects in high redundancy and low signal-to-noise ratio EEG signals remains a great challenge for emotion recognition. In this paper, we propose a novel attention-based spatial-temporal dual-stream fusion network, named ASTDF-Net, for EEG-based emotion recognition. Specifically, ASTDF-Net comprises three main stages: first, the collaborative embedding module is designed to learn a joint latent subspace to capture the coupling of complicated spatiotemporal information in EEG signals. Second, stacked parallel spatial and temporal attention streams are employed to extract the most essential discriminative features and filter out redundant task-irrelevant factors. Finally, the hybrid attention-based feature fusion module is proposed to integrate significant features discovered from the dual-stream structure to take full advantage of the complementarity of the diverse characteristics. Extensive experiments on two publicly available emotion recognition datasets indicate that our proposed approach consistently outperforms state-of-the-art methods. Peiliang Gong, Ziyu Jia, Pengpai Wang, Yueying Zhou, Daoqiang Zhang |
ACM Multimedia | 2 |
| 2023 | EmotionKD: A Cross-Modal Knowledge Distillation Framework for Emotion Recognition Based on Physiological SignalsabstractEmotion recognition using multi-modal physiological signals is an emerging field in affective computing that significantly improves performance compared to unimodal approaches. The combination of Electroencephalogram(EEG) and Galvanic Skin Response(GSR) signals are particularly effective for objective and complementary emotion recognition. However, the high cost and inconvenience of EEG signal acquisition severely hinder the popularity of multi-modal emotion recognition in real-world scenarios, while GSR signals are easier to obtain. To address this challenge, we propose EmotionKD, a framework for cross-modal knowledge distillation that simultaneously models the heterogeneity and interactivity of GSR and EEG signals under a unified framework. By using knowledge distillation, fully fused multi-modal features can be transferred to an unimodal GSR model to improve performance. Additionally, an adaptive feedback mechanism is proposed to enable the multi-modal model to dynamically adjust according to the performance of the unimodal model during knowledge distillation, which guides the unimodal model to enhance its performance in emotion recognition. Our experiment results demonstrate that the proposed model achieves state-of-the-art performance on two public datasets. Furthermore, our approach has the potential to reduce reliance on multi-modal data with lower sacrificed performance, making emotion recognition more applicable and feasible. The source code is available at https://github.com/YuchengLiu-Alex/EmotionKD Ziyu Jia |
ACM Multimedia | 2 |
| 2023 | A Spatial-Temporal Transformer based on Domain Generalization for Motor Imagery ClassificationabstractMotor imagery (MI) has emerged as a classical paradigm in brain-computer interface (BCI) research. In recent years, advancements in deep learning techniques, such as the application of Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs), have enabled the use of MI classification. Despite their success, CNNs and RNNs are not capable of effectively extracting brain spatial and temporal information necessary for MI classification. Additionally, differences in individual subjects further complicate the classification process. To address these limitations, a novel Spatial-Temporal Transformer based on Domain Generalization (ST-DG) has been proposed for MI classification using EEG signals. This framework utilizes a spatial-temporal transformer architecture to capture essential spatiotemporal characteristics of the brain, while also employing Domain Generalization techniques to account for cross-subject variability and improve the model's generalization performance. Experimental results on two public datasets demonstrate the state-of-the-art classification performance. Shaozhe Liu, Leike An, Chi Zhang 0049, Ziyu Jia |
SMC | 4 |
| 2022 | Multi-Level Spatial-Temporal Adaptation Network for Motor Imagery ClassificationabstractElectroencephalogram (EEG) signals for motor imagery (MI) are easily influenced by the environment and the state of the subject, which exhibit temporal and spatial variance. And this variance is more significant across subjects and sessions, which imposes limitations on the cross-domain MI tasks. To address this problem, we propose a Multi-level Spatial-Temporal Adaptation Network (MSTAN), extracting domain-invariant multi-level spatial-temporal features to overcome domain differences. First, stacked spatial-temporal graph convolution (STGCN) layers and an attention-based readout module are designed to extract spatial-temporal patterns of EEGs at multiple levels. An adaptation scheme is then introduced to narrow domain differences: 1) Individual graph parameters for the source and target domains are designed at each STGCN layer to capture the domain-specific brain region dynamic relationships; 2) The differences of spatial-temporal features between the source and target domain are reduced by minimizing the distribution distance. Experiments are conducted to evaluate the proposed method on a public dataset and the results show that our method achieves state-of-the-art performance in cross-domain motor imagery classification. Jing Wang 0060, Ziyu Jia, Zhiqing Hong, Yunze Li, Youfang Lin |
ICASSP | 3 |
| 2022 | Hybrid spiking neural network for sleep electroencephalogram signals
Ziyu Jia, Junyu Ji, Xinliang Zhou |
Sci. China Inf. Sci. | 1 |
| 2021 | Two-Stream Squeeze-and-Excitation Network for Multi-modal Sleep StagingabstractSleep staging is the basis of sleep medicine for diagnosing psychiatric and neurodegenerative diseases. However, the existing sleep staging methods ignore the fact that multi-modal physiological signals are heterogeneous, and different modalities contribute to sleep staging with distinct impacts on specific stages. Therefore, how to model the heterogeneity of multi-modal signals and adaptively utilize the multi-modal signals for sleep staging remains challenging. To address the above challenges, we design a Two-Stream Squeeze-and-Excitation Network (TSSEN) to capture the features of electroencephalogram (EEG) and electrooculogram (EOG) for sleep staging. The TS-SEN is made up of two independent feature extraction networks for modeling the heterogeneity and a Multi-modal Squeeze-and-Excitation feature fusion module for adaptively utilizing the multi-modal signals. Experiments demonstrate that the TSSEN is superior to the baseline models on the public sleep staging dataset. The implementation code of TS-SEN available at https://github.com/xiyangcai/TS-SEN Xiyang Cai, Ziyu Jia, Zehui Jiao |
BIBM | 2 |
| 2021 | SalientSleepNet: Multimodal Salient Wave Detection Network for Sleep StagingabstractSleep staging is fundamental for sleep assessment and disease diagnosis. Although previous attempts to classify sleep stages have achieved high classification performance, several challenges remain open: 1) How to effectively extract salient waves in multimodal sleep data; 2) How to capture the multi-scale transition rules among sleep stages; 3) How to adaptively seize the key role of specific modality for sleep staging. To address these challenges, we propose SalientSleepNet, a multimodal salient wave detection network for sleep staging. Specifically, SalientSleepNet is a temporal fully convolutional network based on the $U^2$-Net architecture that is originally proposed for salient object detection in computer vision. It is mainly composed of two independent $U^2$-like streams to extract the salient features from multimodal data, respectively. Meanwhile, the multi-scale extraction module is designed to capture multi-scale transition rules among sleep stages. Besides, the multimodal attention module is proposed to adaptively capture valuable information from multimodal data for the specific sleep stage. Experiments on the two datasets demonstrate that SalientSleepNet outperforms the state-of-the-art baselines. It is worth noting that this model has the least amount of parameters compared with the existing deep neural network models. Ziyu Jia, Youfang Lin, Jing Wang 0060, Peiyi Xie, Yingbin Zhang |
IJCAI | 1 |
| 2021 | HetEmotionNet: Two-Stream Heterogeneous Graph Recurrent Neural Network for Multi-modal Emotion RecognitionabstractThe research on human emotion under multimedia stimulation based on physiological signals is an emerging field and important progress has been achieved for emotion recognition based on multi-modal signals. However, it is challenging to make full use of the complementarity among spatial-spectral-temporal domain features for emotion recognition, as well as model the heterogeneity and correlation among multi-modal signals. In this paper, we propose a novel two-stream heterogeneous graph recurrent neural network, named HetEmotionNet, fusing multi-modal physiological signals for emotion recognition. Specifically, HetEmotionNet consists of the spatial-temporal stream and the spatial-spectral stream, which can fuse spatial-spectral-temporal domain features in a unified framework. Each stream is composed of the graph transformer network for modeling the heterogeneity, the graph convolutional network for modeling the correlation, and the gated recurrent unit for capturing the temporal domain or spectral domain dependency. Extensive experiments on two real-world datasets demonstrate that our proposed model achieves better performance than state-of-the-art baselines. Ziyu Jia, Youfang Lin, Jing Wang 0060, Zhiyang Feng, Xiangheng Xie, Caijie Chen |
ACM Multimedia | 1 |
| 2020 | BrainSleepNet: Learning Multivariate EEG Representation for Automatic Sleep StagingabstractSleep is one of the most fundamental physiological activities of human beings. Automatic sleep staging can efficiently assist human experts to diagnose the sleep health of people. However, most of the existing methods only considered one or two kinds of time-domain, frequency-domain, and spatial-domain information from EEG signals. Therefore, how to make full use of the complementarity of different features of EEG signals is still a challenge. To tackle this challenge, in this paper we design BrainSleepNet to capture the comprehensive features of multivariant EEG signals for automatic sleep staging. BrainSleepNet consists of an EEG temporal feature extraction module and an EEG spectral-spatial feature extraction module for the temporal-spectral-spatial representation of EEG signals. To the best of our knowledge, it is the first attempt to integrate EEG temporal-spectral-spatial features simultaneously in a unified model for sleep staging. Experiments on the benchmark dataset MASSSS3 demonstrate that BrainSleepNet outperforms all baseline models. The implementation code of BrainSleepNet is available at https://github.com/ziyujia/sleep. Xiyang Cai, Ziyu Jia, Minfang Tang, Gaoxing Zheng |
BIBM | 2 |
| 2020 | Learning Space-Time-Frequency Representation with Two-Stream Attention Based 3D Network for Motor Imagery ClassificationabstractMotor imagery (MI), as one of the important applications of brain-computer interface (BCI), has lately received great attention. However, current MI researches have not provided satisfactory representations of electroencephalogram (EEG), taking account of the space-time-frequency features for MI classification. Moreover, those models also lack the exploration of attentive spatial, temporal, and spectral dynamics. In this study, we propose TA3D (Two-stream Attention based 3D network), a novel model for MI classification. It mainly consists of two streams: the space-time stream and the space-frequency stream, representing and learning discriminative features in the space-time-frequency dimension. Specifically, each stream contains three key parts: 1) 3D representations of EEG signals depict the spatial information over temporal/spectral distributions; 2) Attention mechanisms adaptively explore attentive dynamics of EEG signals and focus on the most valuable information in separate dimensions; 3) 3D convolutions learn spatial representation, temporal dependence, and spectral dependence. The outputs of the two streams are concatenated for space-time-frequency feature fusion. Extensive experiments implemented on two BCI datasets demonstrate that our model outperforms state-of-the-art MI classification methods. Zhenqi Li, Jing Wang 0060, Ziyu Jia, Youfang Lin |
ICDM | 3 |
| 2020 | GraphSleepNet: Adaptive Spatial-Temporal Graph Convolutional Networks for Sleep Stage ClassificationabstractSleep stage classification is essential for sleep assessment and disease diagnosis. However, how to effectively utilize brain spatial features and transition information among sleep stages continues to be challenging. In particular, owing to the limited knowledge of the human brain, predefining a suitable spatial brain connection structure for sleep stage classification remains an open question. In this paper, we propose a novel deep graph neural network, named GraphSleepNet, for automatic sleep stage classification. The main advantage of the GraphSleepNet is to adaptively learn the intrinsic connection among different electroencephalogram (EEG) channels, represented by an adjacency matrix, thereby best serving the spatial-temporal graph convolution network (ST-GCN) for sleep stage classification. Meanwhile, the ST-GCN consists of graph convolutions for extracting spatial features and temporal convolutions for capturing the transition rules among sleep stages. Experiments on the Montreal Archive of Sleep Studies (MASS) dataset demonstrate that the GraphSleepNet outperforms the state-of-the-art baselines. Ziyu Jia, Youfang Lin, Jing Wang 0060, Ronghao Zhou, Xiaojun Ning 0001, Yuanlai He, Yaoshuai Zhao |
IJCAI | 1 |
| 2020 | SST-EmotionNet: Spatial-Spectral-Temporal based Attention 3D Dense Network for EEG Emotion RecognitionabstractMultimedia stimulation of brain activities has not only become an emerging field for intensive research, but also achieves important progress in the electroencephalogram (EEG) emotion classification based on brain activities. However, how to make full use of different EEG features and the discriminative local patterns among the features for different emotions is challenging. Existing models ignore the complementarity among the spatial-spectral-temporal features and discriminative local patterns in all features, which limits the classification ability of the models to a certain extent. In this paper, we propose a novel spatial-spectral-temporal based attention 3D dense network, named SST-EmotionNet, for EEG emotion recognition. The main advantage of the SST-EmotionNet is the simultaneous integration of spatial-spectral-temporal features in a unified network framework. Meanwhile, a 3D attention mechanism is designed to adaptively explore discriminative local patterns. Extensive experiments on two real-world datasets demonstrate that the SST-EmotionNet outperforms the state-of-the-art baselines. Ziyu Jia, Youfang Lin, Xiyang Cai, Haijun Gou, Jing Wang 0060 |
ACM Multimedia | 1 |
| 2020 | MMCNN: A Multi-branch Multi-scale Convolutional Neural Network for Motor Imagery Classification
Ziyu Jia, Youfang Lin, Jing Wang 0060, Kaixin Yang, Tianhang Liu, Xinwang Zhang |
ECML/PKDD (3) | 1 |