EDBT 2026 Demo / reviewers in the wild / expert
Jiahui Pan 0003
dblp:139/8453-3
· DBLP profile ↗
71ranked-venue papers
6as first author
68since 2021 · last 2026
0000-0002-7576-6743ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 3 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 20 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 4 first-author · 20 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Think Faster Than Words: Efficient LLM Chain-of-Thought Reasoning via Dynamic Shortcut DecodingabstractThis paper proposes shortcut decoding, an efficient framework for accelerating Chain-of-Thought (CoT) reasoning in Large Language Models (LLMs).Existing methods that prune or employ early stopping to reduce latency often compromise reasoning reliability.Motivated by the observation that LLMs frequently converge to correct solutions internally before completing explicit textual reasoning, we propose a dual-signal adaptive controller that integrates lightweight probes over internal hidden states with step-level entropy.This controller detects convergence of reasoning during generation and adaptively selects between a fastexit path and a stability-verified path to remove redundant steps while preserving answer correctness.Experiments across multiple mathematical reasoning benchmarks demonstrate that shortcut decoding reduces token usage by approximately 35%, maintains accuracy comparable to full CoT decoding, and achieves finalanswer accuracy comparable to the full CoT baseline, outperforming existing early-stopping methods without updating the base model.Our code is available at https://github.com/ kuromi9527/shortcut_decoding. Yanhao Wang 0001, Zhikang Chen, Lewei He, Jiahui Pan 0003 |
ACL (1) | 7 |
| 2026 | NaturalGAIA: A Verifiable Benchmark and Hierarchical Framework for Long-Horizon GUI TasksabstractZihan Zheng, Tianle Cui, Taoran Wang, Fengtao Wang, Jiahui Pan, Lewei He, Qianglong Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zihan Zheng, Tianle Cui, Taoran Wang, Fengtao Wang, Jiahui Pan 0003, Lewei He, Qianglong Chen |
ACL (1) | 5 |
| 2026 | Dual-Centroid Alignment with Confidence-Aware Query Optimization for Transductive Few-Shot Learning
Chunjin Ye, Quanlin Chen, Jiahui Pan 0003, Jingcong Li 0001 |
ICIC (1) | 4 |
| 2026 | Subdomain Adaption Method for Cross-Domain Emotion Recognition and Clinical Consciousness DetectionabstractEmotion recognition and consciousness detection are critical in neuroscience. Patients with disorder of consciousness (DOC) often cannot express feelings or respond definitively to stimuli due to impaired functions, posing challenges for accurate annotation and model transfer from healthy individuals to patients with DOC. Existing domain adaptation methods struggle to capture inter-subdomain relationships, resulting in suboptimal cross-domain learning performance. This study proposes a novel Category Maximum Mean Discrepancy (CMMD) subdomain adaptation method for EEG-based emotion recognition and consciousness detection. The CMMD framework aligns subdomain distributions of the same category in target and source domains, capturing subtle EEG features in emotional states and improving cross-domain classification. On the emotion recognition benchmark datasets SEED and SEED-IV, our method outperforms existing baselines, achieving 85.67% and 74.44% accuracy. In clinical experiments on 25 DOC patients, our research explored within-group dynamics among 15 MCS and 10 UWS patients via EEG-based emotion recognition. Experiments revealed higher average emotion recognition accuracy in MCS patients (66.21%) than UWS patients (56.99%), with a positive correlation to CRS-R scores (p<0.01). Our research demonstrated that emotion recognition accuracy in all MCS patients and five UWS patients surpassed chance level. All 15 MCS patients had an average recognition accuracy of 66.21%. Although the accuracy of 5 out of 10 UWS patients ranged from 60.42% to 68.75%, barely exceeding the 59.5-60% significance threshold of chance level, this provides objective neurophysiological evidence for potential residual consciousness, suggesting physicians should avoid overlooking subtle signs of consciousness in UWS patients. Results in clinical application highlight our method's potential in understanding DOC patient subgroups and enhancing consciousness detection. Jiahui Pan 0003, Zhipeng He 0001, Junbiao Zhu, Rongming Liang, Zerong Chen, Qiuyou Xie, Jingcong Li 0001, Yuanqing Li 0001 |
IEEE Internet Things J. | 1 |
| 2026 | PDGCN: A progressive dual-branch graph convolution network for EEG emotion recognition
Lina Qiu, Minjin Wu, You Hu, Baiqiang Long, Tianjian Chen, Jiahui Pan 0003 |
Neural Networks | 6 |
| 2026 | Multi-Scale Dynamic Temporal Network With Graph Matching Domain Adaptation for Cross-Subject EEG Emotion RecognitionabstractElectroencephalography (EEG) is a powerful and objective tool for detecting emotions, with broad applications across various domains. However, EEG-based emotion recognition faces two major challenges: (1) extracting domain-invariant features that maintain emotion-related information across variations induced by individual differences, and (2) aligning the marginal and conditional distributions of data from different individuals in the feature space. To address these challenges, we propose a novel method: the Multi-scale Dynamic Temporal Network with Graph Matching Domain Adaptation, designed for cross-subject EEG emotion recognition. Our method employs a multi-scale dynamic temporal attention module to extract robust domain-invariant features. In addition, it utilizes a domain adaptation network based on Chebyshev graph representations, which reformulates the emotion recognition task as a graph matching problem. This approach effectively aligns cross-subject data distribution. To validate the effectiveness of our method, we conducted extensive experiments on two benchmark databases (SEED and DEAP) under three different cross-validation protocols. The results demonstrate consistent reliability in subjects and sessions. Specifically, in the cross-subject single-session cross-validation task, our method achieved an accuracy of 94.69%$\pm$5.16% on the SEED dataset. On the DEAP dataset, our approach achieved accuracies of 69.45%$\pm$7.26% and 67.52%$\pm$8.28% for the Valence and Arousal dimensions, respectively. These results demonstrate that our method outperforms existing state-of-the-art approaches. Source code is available at:https://github.com/seizeall/MDTN-GMDA. Rongtao Chen, Zhepei Hong, Zhuoyi Huang, Jingcong Li 0001, Jiahui Pan 0003 |
IEEE Trans. Affect. Comput. | 6 |
| 2026 | MSDM: A Lightweight Multi-Scale Dynamic Mamba for Dynamic Facial Expression Recognition in Smart ClassroomsabstractIn smart classroom environments, dynamic facial expression recognition (DFER) is crucial for enhancing teaching quality and improving students' learning experiences. However, existing DFER models face significant challenges in computational efficiency and temporal modeling, which limit their practical application in resource-constrained settings. To address these issues, this paper proposes a novel lightweight DFER framework called Multi-Scale Dynamic Mamba (MSDM). The MSDM model combines a Multi-Scale Attention Fusion Module (MSAFM) to effectively integrate global and local facial features and a Dynamic Temporal Focus (DTF) mechanism to enhance the modeling of long-term facial expression dynamics. These components work together to highlight key facial muscle movements while reducing background interference. Additionally, we introduce Dual-Resolution Bidirectional Mamba (DR Bi-Mamba) blocks that process high- and low-resolution facial images in parallel for coarse-to-fine feature extraction. This bio-inspired strategy enhances robustness by effectively integrating global context and local details. To better align with the practical requirements of smart classroom scenarios, we have developed a dedicated classroom dataset, HM-Class, which addresses the mismatch between existing emotion categories and the high-frequency emotional states observed in educational contexts. Extensive experiments on seven in-the-wild datasets—four DFER datasets, two static facial expression recognition (SFER) datasets, and the HM-Class dataset—show that MSDM outperforms state-of-the-art methods with fewer parameters and lower computational costs. This study offers an efficient solution for affective computing in resource-constrained classroom environments and advances the practical application of DFER technology in educational settings. Jiangyu Cui, Ruixiang Gao, Caiqi Chen, Lei Mo, Jiahui Pan 0003 |
IEEE Trans. Affect. Comput. | 7 |
| 2026 | Watch Where You Move: Region-Aware Dynamic Aggregation and Excitation for Gait RecognitionabstractDeep learning-based gait recognition has achieved great success in various applications. The key to accurate gait recognition lies in considering the unique and diverse behavior patterns indifferent motion regions, especially when covariates affect visual appearance. However, existing methods typically use predefined regions for temporal modeling, with fixed or equivalent temporal scales assigned to different types of regions, which makes it difficult to model motion regions that change dynamically over time and adapt to their specific patterns. To tackle this problem, we introduce a Region-aware Dynamic Aggregation and Excitation framework (GaitRDAE) that automatically searches for motion regions, assigns adaptive temporal scales and applies corresponding attention. Specifically, the framework includes two core modules: the Region-aware Dynamic Aggregation (RDA) module, which dynamically searches the optimal temporal receptive field for each region, and the Region-aware Dynamic Excitation (RDE) module, which emphasizes the learning of motion regions containing more stable behavior patterns while suppressing attention to static regions that are more susceptible to covariates. Experimental results show that GaitRDAE achieves state-of-the-art performance on several benchmark datasets. The source code will be published athttps://github.com/HUAFOR/GaitRDAE. Binyuan Huang, Yongdong Luo, Xianda Guo, Xiawu Zheng, Jiahui Pan 0003, Chengju Zhou |
IEEE Trans. Multim. | 6 |
| 2025 | PlanningArena: A Modular Benchmark for Multidimensional Evaluation of Planning and Tool LearningabstractOne of the research focuses of large language models (LLMs) is the ability to generate action plans.Recent studies have revealed that the performance of LLMs can be significantly improved by integrating external tools.Based on this, we propose a benchmark framework called PlanningArena, which aims to simulate real application scenarios and provide a series of apps and API tools that may be involved in the actual planning process.This framework adopts a modular task structure and combines user portrait analysis to evaluate the ability of LLMs in correctly selecting tools, logical reasoning in complex scenarios, and parsing user information.In addition, we deeply diagnose the task execution effect of LLMs from both macro and micro levels.The experimental results show that even the most outstanding GPT-4o and DeepSeekV3 models only achieved a total score of 56.5% and 41.9% in PlanningArena, respectively, indicating that current LLMs still face challenges in logical reasoning, context memory, and tool calling when dealing with different structures, scenarios, and their complexity.Through this benchmark, we further explore the path to optimize LLMs to perform planning tasks. Zihan Zheng, Tianle Cui, Chuwen Xie, Jiahui Pan 0003, Qianglong Chen, Lewei He |
ACL (1) | 4 |
| 2025 | PR-DA: Prototype Regularization Domain Adaptation for Cross-Subject EEG-Based Emotion RecognitionabstractElectroencephalogram (EEG)-based emotion recognition holds significant potential in healthcare, traffic safety, and entertainment. However, cross-subject emotion recognition remains challenging due to individual differences and the difficulty in extracting domain-invariant features. To address these issues, this paper proposes a novel Prototype Regularization Domain Adaptation (PR-DA) framework. Experimental results on three benchmark datasets (SEED, SEED-IV, and SEED-VII) demonstrate that the proposed PR-DA framework achieves superior performance compared to state-of-the-art methods, with accuracies of 96.30%±2.87%, 86.56%±4.67%, and 50.43%±6.71 %, respectively. The proposed PR-DA framework of-fers a promising approach for cross-subject EEG-based emotion recognition. The source code is available at the following link: https://github.com/seizeall/PR-DA. Rongtao Chen, Zhepei Hong, Qi You, Chuwen Xie, Jiahui Pan 0003 |
BIBM | 6 |
| 2025 | SpinTSA: A Temporal-Spatial-Attentive Network with Multiscale Anchors for Sleep Spindle DetectionabstractAccurate identification of sleep spindles, transient EEG oscillations that are critical for memory consolidation and brain health, is essential for advancing sleep research and clinical applications. Despite considerable efforts, current methods still face significant challenges in capturing the complex dynamics of spindles across different temporal scales, maintaining robustness against inter-subject variability and common EEG artifacts, and generalizing detection across various sleep events without relying on detailed, event-specific rule crafting. To address these challenges, this study introduces SpinTSA, a new deep neural framework that integrates a UNet-BiLSTM backbone for joint spatial-temporal feature extraction, a Progressive Attention Downsampling (PAD) module with squeeze-and-excitation attention for refined feature recalibration, and a Multiscale Temporal Anchor (MTA) mechanism for precise and scalable event localization. Together, these components enable the model to effectively detect spindles with varying morphologies and durations. Evaluations on the public MASS-SS dataset and the DREAMS Sleep Spindles Database show SpinTSA's strong performance, achieving F1-scores of 79.2 % and 61.4 %, respectively, which are significantly higher than existing benchmarks. Furthermore, by using the flexible nature of the MTA module, SpinTSA can be easily extended to detect K-complexes, reaching an F1-score of 61.9% on the MASS-KC dataset. These results highlight SpinTSA's effectiveness and adaptability as a unified framework for automated sleep event detection, bridging the gap between research and clinical applications. Jiahui Pan 0003 |
BIBM | 3 |
| 2025 | RS-EGT: Resting-State EEG Graph Transformer for Large-Scale EEG Signal AnalysisabstractResting-state EEG is a vital modality for investigating fundamental neural dynamics. However, few of the recently emerged large-scale EEG models (LEMs) can effectively decode the diverse and complex nature of resting-state EEG signals, especially those with varying channel configurations. Furthermore, most existing LEMs fail to incorporate training methodologies specifically designed for the unique characteristics of EEG signals. Motivated by the success of large language models (LLMs), we propose the resting-state EEG Graph Transformer (RS-EGT), the first LEM specifically designed for restingstate EEG representation. RS-EGT employs a channel-agnostic architecture that can process EEG signals from various channel configurations, enabling it to learn universal patterns of brain activity. It utilizes a dual self-supervised pretraining framework that combines graph contrastive learning and autoregressive generative learning to capture robust spatial and temporal representations. We pretrain RS-EGT on a large-scale restingstate EEG dataset and validate its generalization capability across downstream tasks. It achieves a TPR of 0.797 on the icare dataset, outperforming previous models. Our work extends the capabilities of LEMs to previously underexplored restingstate scenarios, paving the way for broader applications in neuroscience research. Fei Wang 0026, Jiahui Pan 0003 |
BIBM | 5 |
| 2025 | MSCNN-ADDA: A Cross-Subject P300 EEG Decoding Algorithm Based on a Multi-Scale Convolutional Neural Network and Adversarial Discriminative Domain Adaptation
Wanying He, Yongxi Zhao, Jiahui Pan 0003 |
CogSci | 3 |
| 2025 | Advancing Adolescent Depression Detection through Multi-Task EEG Signals and Biosignal Learning
Aojun Wen, Zijing Guan, Quanlin Chen, Jiahui Pan 0003, Haiyun Huang |
CogSci | 6 |
| 2025 | Meta-MMD Fusion: Enhancing Cross-Subject Motor Imagery ClassificationabstractMotor imagery (MI) is a widely used paradigm in brain-computer interfaces (BCIs). Despite recent advancements, MI classification still faces challenges such as limited data availability and poor performance for new users. In particular, feature alignment based on deep learning in zero-calibration cross-subject frameworks remains inadequately explored. To address these issues, we propose the Meta-MMD, a novel cross-subject MI classification method that integrates meta-learning and maximum mean discrepancy (MMD) strategies. Our method minimizes the discrepancy between support and query distributions within each learning task, enhancing robustness and generalization. Experimental evaluations on two public datasets, BCI Competition IV-2a and BCI Competition IV-2b, were conducted using two backbone networks, EEGNet and DeepConvNet.Based on EEGNet, the accuracies are 69.27% and 79.22%, respectively. Based on DeepConvNet, the accuracies are 67.36% and 78.17%, respectively. Our proposed method outperforms the current state-of-the-art methods. The effectiveness of our method was thus demonstrated. Chao Qu, Jiahui Pan 0003 |
ICASSP | 3 |
| 2025 | A Novel Perspective on Leveraging Hubness in VAE for Eliminating Representative Shift Vectors in Few-Shot LearningabstractFew-Shot Learning (FSL) is a technique aimed at improving a model’s ability to generalize to unseen categories using only a small amount of labeled data. Representative shift vectors and the hubness problem are two common issues that often hinder the performance of FSL. Representative shift vectors refer to the inherent bias caused by the absolute positional differences of various base classes in the feature space. The hubness problem occurs when a class prototype becomes the nearest neighbor for many test instances, regardless of their true class. Inspired by generative models, we find that these two problems can be effectively addressed simultaneously. By leveraging generative models, we eliminate representative shift vectors through learning intra-class diversity and enhance the quality of generated samples by utilizing hubness samples. Experimental results demonstrate that our method improves performance, and our code will be available at https://github.com/cql33/HG-VAE. Quanlin Chen, Chunjin Ye, Jiahui Pan 0003, Jingcong Li 0001 |
ICME | 4 |
| 2025 | Study of Finger Biometrics on Finger Semantic Segmentation and Finger Shape AuthenticationabstractIn hand-based biometrics, fingerprint, finger vein, finger knuckle print, palm print, palm vein, dorsal hand vein, and hand shape are the traits that are getting much attention. However, finger shape (FS), a forgettable trait, has not been studied specifically for identification purposes. In this work, we explore this content as a complement to the hand-based biometrics. Firstly, we annotate the FS on a publicly available finger vein dataset as the ground truth for finger semantic segmentation. Then we explore the finger semantic segmentation task on the annotated data and propose a lightweight network, namely FinSeg-Net (finger segmentation network). Finally, we conduct the FS authentication experiment based on four matching methods; experimental results show that the FS traits can achieve identity authentication. This work is the first study for FS biometrics specifically, and built the first FS dataset, which will be accessed via: https://github.com/SCUT-BIP-Lab/FinSeg. Junduan Huang, Dacan Luo, Weili Yang, Jiahui Pan 0003, Wenxiong Kang |
ICME | 4 |
| 2025 | Time-Frequency Domain Fusion Transformer for Cross-Subject Motor Imagery ClassificationabstractBrain-Computer Interfaces (BCIs) utilizing motor imagery (MI) have become prevalent, yet challenges in deep learning-based MI classification remain, especially regarding domain shift mitigation in low-channel MI and the integration of multimodal features. This study proposes the Time-Frequency Domain Fusion Transformer (TFDFT), a novel multimodal framework designed to overcome these hurdles. The TFDFT employs an agent attention mechanism to seamlessly integrate time and frequency features across dimensions, bolstering the model’s generalization. To counter domain shifts, we have also introduced the Domain Generalization in Conditional Domain Adversarial Network (DG-CDAN) for low-channel MI across subjects. Experimental results demonstrate that TFDFT achieves state-of-the-art performance, surpassing previous methods in cross-subject MI classification with only three channels, achieving 73.30% and 76.37% cross-subject accuracy on the BCIC IV-2a and IV-2b datasets, respectively. The TFDFT significantly surpasses the constraints of single-modal approaches, offering a robust solution for BCI applications. Zijian Xia, Jiahui Pan 0003 |
ICME | 3 |
| 2025 | Multi-soft-label Guided Supervised Contrastive Learning for Gait Emotion RecognitionabstractGait-based emotion recognition has received considerable attention due to its non-invasive capturing manner. However, most existing works learn the gait representations by treating different classes independently, which ignores the inherent class ambiguity in this field. Therefore, we propose a Multi-soft-label Guided Supervised Contrastive Learning (MSL-SCL) framework, which leverages the class correlation information in soft labels to explicitly guide the SCL, thereby alleviating the ambiguous gait representation. Specifically, a Soft-Label SCL (Sof-SCL) module is designed to select positive and negative samples based on their soft-label similarity to the anchors, and the similarity is further incorporated into a novel contrastive loss function for the refinement. Moreover, a Prior-Guided SCL (Prior-SCL) is introduced to capture the subtle changes in gait and employed as soft labels to provide an adaptive supervision for SCL. Extensive experimental results on the Emotion-Gait dataset demonstrate that our method outperforms SOTAs with a mean average precision of 89.8%. Chengju Zhou, Mengxin Xu, Xiaotong Fan, Liangyu Lu, Jiahui Pan 0003, Lewei He |
ICME | 5 |
| 2025 | Adaptive Sleep Staging via Cross-Modal Feature Interaction and Time-Frequency AnalysisabstractAutomated sleep staging is a pivotal element in the field of sleep medicine. Despite promising advancements in deep learning, several challenges remain: (1) how to extract domain-invariant features that preserve rich, emotion-related information under cross-domain conditions; (2) how to align the marginal and conditional distributions of data from different individuals in the feature space. To address these challenges, we propose a novel method, Multi-modal Temporal-Spectral Neural Network for Sleep stage classification (MTSNSleepNet), to enhance the accuracy and reliability of sleep stage classification. MTSNSleepNet incorporates convolutional neural networks (CNNs) and attention-based modules, which are known for their ability to capture complex data patterns, thereby achieving a comprehensive understanding of sleep patterns. It features a dual-stream attention architecture that effectively combines temporal and frequency information, allowing for simultaneous processing and analysis of data in both the time and frequency domains. This provides a more nuanced examination of sleep stage characteristics. Additionally, we designed a multi-scale spatio-temporal analysis mechanism that captures both rapid transitions and slower sleep cycle trends, thereby accurately identifying sleep state transitions. Our experiments on the Sleep-EDF-39 dataset demonstrate that MTSNSleepNet outperforms general models across multiple evaluation metrics. The model achieves an accuracy of 89.05%, an F1 score of 82.61%, and a Cohen Kappa coefficient of 84.8% on the dataset. Furthermore, the model shows particular effectiveness in distinguishing subtle changes between sleep stages. Chuqi Fu, Jiahui Pan 0003 |
IJCNN | 5 |
| 2025 | An Enhanced Autism Detection Model Based on Electroencephalogram SignalsabstractIn recent years, with increasing demand for the diagnosis of autism spectrum disorder (ASD), automated detection methods based on electroencephalogram (EEG) signals have gained significant attention. However, existing approaches still face challenges in terms of accuracy, generalization ability, robustness, and interpretability. We propose an improved ASD detection model, ATCosLNet (Attention-CosCNN-LSTM-Net), which leverages a multidimensional attention mechanism to focus on critical signals, combines Cosine Convolutional Neural Network (CosCNN) for frequency feature extraction, and integrates a tree-structured LSTM to model hierarchical and long-term dependencies, enabling comprehensive extraction of spatiotemporal and frequency domain features from EEG signals. The results of experiments using ten-fold cross-validation demonstrate that ATCosLNet achieves 94.11% classification accuracy on the KAU dataset and 95.80% on the EEG-ASD dataset, significantly outperforming existing detection methods. Moreover, the model exhibits stable performance across different data splits and unseen data, validating its strong generalization ability and robustness, providing an efficient, stable, and interpretable solution for automated ASD detection, advancing the application of EEG signals in ASD diagnosis and offering valuable support for related research and clinical practice. Weihao Huang, Chunhong Jiang, Jiahui Pan 0003 |
IJCNN | 5 |
| 2025 | Attention-aware spatio-temporal learning for multi-view gait-based age estimation and gender classificationabstractAbstract Recently, gait‐based age and gender recognition have attracted considerable attention in the fields of advertisement marketing and surveillance retrieval due to the unique advantage that gaits can be perceived at a long distance. Intuitively, age and gender can be recognised by observing people's static shape (e.g. different hairstyles between males and females) and dynamic motion (e.g. different walking velocities between the elderly and youth). However, most of the existing gait‐based age and gender recognition methods are based on Gait Energy Image (GEI), which loses the capability of explicitly modelling temporal dynamic information and is not robust to the multi‐view recognition that inevitably happens in a real application. Therefore, in this study, an Attention‐aware Spatio‐Temporal Learning (ASTL) framework is proposed, which employs a silhouette sequence as input to learn essential and invariable spatial‐temporal gait representations. More specifically, a Multi‐Scale Temporal Aggregation (MSTA) module provides an effective scheme for dynamic gait description by exploring and aggregating multi‐scale temporal interval information, which is a core supplement to spatial representation. Then, a Multiple Attention Aggregation (MAA) module is designed to help the network focus on the most discriminatory information along temporal, spatial and channel dimensions. Finally, a Multimodal Collaborative Learning (MCL) block gives full play to the advantages of different modal features through a multimodal cooperative learning strategy. The mean absolute error (MAE) for the age estimation and the correct classification rate (CCR) for the gender classification on OU‐MVLP achieve 6.68 years and 97%, respectively, demonstrating the superiority of the method. Ablation experiments and visualisation results also prove the effectiveness of the three individual modules in their framework. Binyuan Huang, Yongdong Luo, Jiahui Xie, Jiahui Pan 0003, Chengju Zhou |
IET Comput. Vis. | 4 |
| 2025 | A robust transductive distribution calibration method for few-shot learning
Jingcong Li 0001, Chunjin Ye, Fei Wang 0026, Jiahui Pan 0003 |
Pattern Recognit. | 4 |
| 2025 | WST-CAA: A Brainprint Recognition Framework Based on Wavelet Scattering Transform and Channel-Axial Attention Dense NetworkabstractBrainprint has emerged as a promising biometrics modality due to its advantages of high security and user subjectivity. Currently, Brainprint recognition faces two fundamental limitations: (1) the inherent susceptibility of EEG signals to physiological artifacts, and (2) the limited capacity of extracting reliable identity-relevant features. To address the above limitations, we propose WST-CAA, a novel Brainprint recognition framework that combines a cascaded preprocessing and wavelet scattering transform pipeline (PWST) with a channel-axial attention dense network (CAADN). The proposed WST-CAA excels in denoising and identity feature extraction, effectively capturing channel-spatial-temporal correlations while suppressing cross-dimensional noise. Extensive experiments on the publicly available datasets, FACED and DEAP, demonstrate that WST-CAA outperforms baseline models, achieving 99.48% accuracy on DEAP and 97.46% on FACED while maintaining state-of-the-art performance across both small and large-scale EEG datasets. The relevant code is available at:https://github.com/jhuangscnu/BrainPrint_WST-CAA. Haoying Zhu, Desen Wang, Jiahui Pan 0003, Junduan Huang |
IEEE Signal Process. Lett. | 4 |
| 2025 | Decoding Musical Neural Activity in Patients With Disorders of Consciousness Through Self-Supervised Contrastive Domain GeneralizationabstractIdentifying the brain responses of patients with disorders of consciousness (DOCs), which include comas, vegetative states (VSs, also called unresponsive wakefulness syndrome (UWS)) and minimally conscious states (MCSs), based on electroencephalography (EEG) has important clinical diagnosis implications. However, due to impaired motor and cognitive abilities, patients with DOCs may not be able to express their feelings and their brain responses to different stimuli, making it difficult to correctly label data. EEG classification algorithms trained with these data cannot make reliable classifications and predictions for clinical diagnosis purposes. To identify the brain responses produced for different types of stimuli in patients with DOCs, we proposed a self-supervised contrastive domain generalization framework (SSCDG) for cross-subject EEG classification. The model was first trained with healthy-subject EEG data induced by different stimuli to learn their corresponding unsupervised representations. Then, we used these representations to train a classifier to predict the emotional states of patients with DOCs under the corresponding stimuli. SSCDG was first evaluated on the SEED dataset, and it achieved an accuracy of 87.6%, which was 1.1% higher than that of the state-of-the-art (SOTA) approaches. Moreover, the SSCDG method was utilized to categorize EEG data acquired from seventeen DOC patients, including eleven in a UWS state and six in an MCS state, with seven patients demonstrating notable accuracy in three-class EEG classification tasks. The SSCDG results indicated that the seven patients with DOCs may have shown classifiable EEG responses to the presented stimuli. Honghua Cai, Jiahui Pan 0003, Qiuyi Xiao, Jiarui Jin, Yuanqing Li 0001, Qiuyou Xie |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | MTADA: A Multi-Task Adversarial Domain Adaptation Network for EEG-Based Cross-Subject Emotion RecognitionabstractIn electroencephalogram (EEG)-based emotion recognition, the applicability of most current models is limited by inter-subject variability and emotion complexity. This study proposes a multi-task adversarial domain adaptation (MTADA) network to enhance cross-subject emotion recognition performance. The model first employs a domain matching strategy to select the source domain that best matches the target domain. Then, adversarial domain adaptation is used to learn the difference between source and target domains, and a fine-grained joint domain discriminator is constructed to align them by incorporating category information. At the same time, a multi-task learning mechanism is utilized to learn the intrinsic relationships between different emotions and predict multiple emotions simultaneously. We conducted comprehensive experiments on two public datasets, DEAP and FACED. On DEAP, the average accuracies for valence, arousal and dominance are 76.39%, 69.74% and 68.26%, respectively. On FACED, the average accuracies for valence and arousal are 78.90% and 77.95%. When using the subject from DEAP as the source domain to predict the subjects in FACED, the accuracies for valence and arousal are 61.07% and 60.82%. These results show that our MTADA model improves cross-subject emotion recognition and outperforms most state-of-the-art methods, which may provide new approach for EEG-based emotion brain-computer interface systems. Lina Qiu, Zuorui Ying, Xianyue Song, Weisen Feng, Chengju Zhou, Jiahui Pan 0003 |
IEEE Trans. Affect. Comput. | 6 |
| 2025 | A Multimodal Consistency-Based Self-Supervised Contrastive Learning Framework for Automated Sleep Staging in Patients With Disorders of ConsciousnessabstractSleep is a fundamental human activity, and automated sleep staging holds considerable investigational potential. Despite numerous deep learning methods proposed for sleep staging that exhibit notable performance, several challenges remain unresolved, including inadequate representation and generalization capabilities, limitations in multimodal feature extraction, the scarcity of labeled data, and the restricted practical application for patients with disorder of consciousness (DOC). This paper proposes MultiConsSleepNet, a multimodal consistency-based sleep staging network. This network comprises a unimodal feature extractor and a multimodal consistency feature extractor, aiming to explore universal representations of electroencephalograms (EEGs) and electrooculograms (EOGs) and extract the consistency of intra- and intermodal features. Additionally, self-supervised contrastive learning strategies are designed for unimodal and multimodal consistency learning to address the current situation in clinical practice where it is difficult to obtain high-quality labeled data but has a huge amount of unlabeled data. It can effectively alleviate the model's dependence on labeled data, and improve the model's generalizability for effective migration to DOC patients. Experimental results on three publicly available datasets demonstrate that MultiConsSleepNet achieves state-of-the-art performance in sleep staging with limited labeled data and effectively utilizes unlabeled data, enhancing its practical applicability. Furthermore, the proposed model yields promising results on a self-collected DOC dataset, offering a novel perspective for sleep staging research in patients with DOC. Jiahui Pan 0003, Yangzuyi Yu, Wanxin Wei, Shuyu Chen 0002, Heyi Zheng, Yanbin He, Yuanqing Li 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | Heterogeneous Domain Adaptation via Correlative and Discriminative Feature LearningabstractHeterogeneous domain adaptation seeks to learn an effective classifier or regression model for unlabeled target samples by using the well-labeled source samples but residing in different feature spaces and lying different distributions. Most recent works have concentrated on learning domain-invariant feature representations to minimize the distribution divergence via target pseudo-labels. However, two critical issues need to be further explored: 1) new feature representations should be not only domain-invariant but also category-correlative and discriminative and 2) alleviating the negative transfer caused by the incorrect pseudo-labeling target samples could boost the adaptation performance during the iterative learning process. To address these issues, in this paper, we put forward a novel heterogeneous domain adaptation method to learn category-correlative and discriminative representations, referred to as correlative and discriminative feature learning (CDFL). Specifically, CDFL aims to learn a feature space where class-specific feature correlations between the source and target domains are maximized, the divergences of marginal and conditional distribution between the source and target domains are minimized, and the distances of inter-class distribution are forced to be maximized to ensure the discriminative ability. Meanwhile, a selective pseudo-labeling procedure based on the correlation coefficient and classifier prediction is introduced to boost class-specific feature correlation and discriminative distribution alignment in an iteration way. Extensive experiments certify that CDFL outperforms the State-of-the-Art algorithms on five standard benchmarks. Yuwu Lu, Dewei Lin, LinLin Shen, Yicong Zhou, Jiahui Pan 0003 |
IEEE Trans. Multim. | 5 |
| 2025 | Publisher Correction: Twinenet: coupling features for synthesizing volume rendered images via convolutional encoder-decoders and multilayer perceptrons
Shengzhou Luo, Jingxing Xu, John Dingliana, Mingqiang Wei, Lewei He, Jiahui Pan 0003 |
Vis. Comput. | 7 |
| 2024 | CariesXrays: Enhancing Caries Detection in Hospital-Scale Panoramic Dental X-rays via Feature Pyramid Contrastive LearningabstractDental caries has been widely recognized as one of the most prevalent chronic diseases in the field of public health. Despite advancements in automated diagnosis across various medical domains, it remains a substantial challenge for dental caries detection due to its inherent variability and intricacies. To bridge this gap, we release a hospital-scale panoramic dental X-ray benchmark, namely “CariesXrays”, to facilitate the advancements in high-precision computer-aided diagnosis for dental caries. It comprises 6,000 panoramic dental X-ray images, with a total of 13,783 instances of dental caries, all meticulously annotated by dental professionals. In this paper, we propose a novel Feature Pyramid Contrastive Learning (FPCL) framework, that jointly incorporates feature pyramid learning and contrastive learning within a unified diagnostic paradigm for automated dental caries detection. Specifically, a robust dual-directional feature pyramid network (D2D-FPN) is designed to adaptively capture rich and informative contextual information from multi-level feature maps, thus enhancing the generalization ability of caries detection across different scales. Furthermore, our model is augmented with an effective proposals-prototype contrastive regularization learning (P2P-CRL) mechanism, which can flexibly bridge the semantic gaps among diverse dental caries with varying appearances, resulting in high-quality dental caries proposals. Extensive experiments on our newly-established CariesXrays benchmark demonstrate the potential of FPCL to make a significant social impact on caries diagnosis. Bingzhi Chen, Sisi Fu, Yishu Liu 0001, Jiahui Pan 0003, Guangming Lu 0002, Zheng Zhang 0006 |
AAAI | 4 |
| 2024 | MC-FAW: A Multi-Scale Convolutional Feature Adaptive Weighting Fusion Network for Detecting Disorders of ConsciousnessabstractThe clinical assessment of patients with disorders of consciousness (DoC) still faces challenges due to the lack of objective and accurate methods. To address this issue, this paper proposes a multi-scale convolutional feature adaptive weighting (MC-FAW) fusion network, which aims to integrate multi-scale electroencephalogram (EEG) functional connectivity features to improve the classification performance of consciousness states. Based on resting-state EEG data from 15 healthy controls (HC), 15 patients in a minimally conscious state (MCS), and 17 patients in a vegetative state (VS), the MC-FAW model effectively captures multi-scale features of brain neural activity and dynamically weights these features to achieve a more comprehensive and accurate assessment of consciousness. The classification results demonstrate that this method achieves classification accuracies of 96.54% for distinguishing HC from MCS, 98.47% for HC from VS, and 85.25% for MCS from VS. These results indicate that our method has the potential to improve the objectivity and accuracy of clinical diagnosis, thereby providing important support for clinical decision-making. Lina Qiu, Xianyue Song, Zuorui Ying, Weisen Feng, Liangquan Zhong, Jiahui Pan 0003 |
BIBM | 6 |
| 2024 | EEG-Based Emotion Recognition via Convolutional Transformer with Class Confusion-Aware Attention
Jiahui Pan 0003, Chenyu Bai |
CogSci | 1 |
| 2024 | A Deep Channel Attention Transformer for Multimodal EEG-EOG-Based Vigilance Estimation
Jiahui Pan 0003, Dehua Lu |
CogSci | 1 |
| 2024 | A Novel Self-Supervised Learning Method for Sleep Staging and its Pilot Study on Patients with Disorder of Consciousness
Jingcong Li 0001, Quanlin Chen, Jiahui Pan 0003, Haiyun Huang |
CogSci | 3 |
| 2024 | An Investigation on EEG-based Prognosis Prediction of Patients with Disorders of Consciousness
Jingcong Li 0001, Jiahui Pan 0003, Fei Wang 0026 |
CogSci | 3 |
| 2024 | EFMLNet: Fusion Model Based on End-to-End Mutual Information Learning for Hybrid EEG-fNIRS Brain-Computer Interface Applications
Lina Qiu, Weisen Feng, Zuorui Ying, Jiahui Pan 0003 |
CogSci | 4 |
| 2024 | Cross-subject EEG Emotion Recognition based on Multitask Adversarial Domain Adaption
Lina Qiu, Zuorui Ying, Weisen Feng, Jiahui Pan 0003 |
CogSci | 4 |
| 2024 | Medical Vision-Language Representation Learning with Cross-Modal Multi-Teacher Contrastive DistillationabstractMedical vision-language representation learning has garnered considerable attention owing to its applicability to extracting generic representations from the image and text modality. However, it still remains challenging to acquire a more comprehensive understanding of intra- and inter-modal semantic knowledge. In this paper, we propose a Cross-Modal Multi-Teacher Contrastive Distillation (CMCD) architecture, which aims to comprehensively learn medical vision-language representation in a unified multi-teacher framework. Specifically, a cross-modal knowledge distillation (CKD) module is designed to refine reconstructed semantics under an additional supervision signal generated by momentum teachers from the other modality, achieving more robust semantic interaction across modalities. To better alleviate the heterogeneity and semantic gaps, the multi-level contrastive learning (MCL) module is conceived to align features of both intra- and inter-modal via contrastive learning from multi-level perspectives. Extensive experiments on two medical downstream tasks, i.e., Med-VQA and Med-ITC, demonstrate that our CMCD consistently outperforms the state-of-the-art methods. Bingzhi Chen, Yishu Liu 0001, Jiahui Pan 0003, Meirong Ding |
ICASSP | 5 |
| 2024 | A Relevant Prototype Domain Gradient Projection Continual Learning Method for Cross-Subject P300 Brain-Computer Interfaces
Zhicong Wu, Honghua Cai, Yuyan Ling, Jiahui Pan 0003 |
ICIC (4) | 4 |
| 2024 | Enhancing Few-Shot Classification without Forgetting Through Multi-level Contrastive ConstraintsabstractMost recent few-shot learning approaches are based on meta-learning with episodic training. However, prior studies encounter two crucial problems: (1) the presence of inductive bias, and (2) the occurrence of catastrophic forgetting. In this paper, we propose a novel Multi-Level Contrastive Constraints (MLCC) framework, that jointly integrates within-episode learning and across-episode learning into a unified interactive learning paradigm to solve these issues. Specifically, we employ a space-aware interaction modeling scheme to explore the correct inductive paradigms for each class between within-episode similarity/dis-similarity distributions. Additionally, with the aim of better utilizing former prior knowledge, a cross-stage distribution adaption strategy is designed to align the across-episode distributions from different time stages, thus reducing the semantic gap between existing and past prediction distribution. Extensive experiments on multiple few-shot datasets demonstrate the consistent superiority of MLCC approach over the existing state-of-the-art baselines. Bingzhi Chen, Haoming Zhou, Yishu Liu 0001, Jiahui Pan 0003, Guangming Lu 0002 |
ICME | 5 |
| 2024 | A Density-driven Iterative Prototype Optimization for Transductive Few-shot Learning
Jingcong Li 0001, Chunjin Ye, Fei Wang 0026, Jiahui Pan 0003 |
IJCAI | 4 |
| 2024 | MFASleepNet: Multi-view fusion attention-based deep neural network for automatic sleep stagingabstractSleep staging is important for both the assessment of sleep quality and the diagnosis of sleep-related disorders. Although previous studies have attempted to automatically classify sleep stages and achieved high classification performance, most automatic staging algorithms currently use only the time or frequency domain information of the data, resulting in limited feature representation capability. To address the above problems, we propose a multi-view fusion attention-based deep neural network (MFASleepNet) to classify sleep stages using single-channel EEG signals. MFASleepNet consists of a multi-scale CNN, a multi-view fusion attention module, a global external attention module and a conditional random field. Specifically, the multi-scale CNN can extract features from different frequency bands of the signal. Then the multi-view fusion attention module can fuse multiple views from different time-frequency spaces to improve the feature representation capability. Subsequently, the global external attention module first captures the global feature representation of each sample through the global context block, and then captures the common feature representation of all samples through the external attention mechanism, and outputs a preliminary sleep stage result. Finally, the conditional random field is used to learn sleep stage transition rules to improve classification performance. We evaluate the performance of our model on the Sleep-EDF dataset, and MFASleepNet achieves an accuracy of 87.2% and an F1 score of 81.1% on the Fpz-Cz channel. The experimental results show that our MFASleepNet can optimise the performance of sleep staging and achieve better staging performance than state-of-the-art methods. Zhoujie Hou, Jiahui Pan 0003, Yuanqing Li 0001 |
IJCNN | 2 |
| 2024 | Stay Focused is All You Need for Adversarial Robustness
Bingzhi Chen, Ruihan Liu, Yishu Liu 0001, Xiaozhao Fang, Jiahui Pan 0003, Guangming Lu 0002, Zheng Zhang 0006 |
ACM Multimedia | 5 |
| 2024 | GaitCTCG: cross-view gait recognition via cascaded residual temporal shift and comprehensive multi-granularity learning
Binyuan Huang, Chengju Zhou, Lewei He, Jiahui Pan 0003 |
Appl. Intell. | 5 |
| 2024 | Multi-degree-of-freedom unmanned aerial vehicle control combining a hybrid brain-computer interface and visual obstacle avoidance
Shanghong Xie, Wei Gao 0038, Qingfu Wu, Nianming Ban, Jiahui Pan 0003 |
Eng. Appl. Artif. Intell. | 8 |
| 2024 | An attention-based adaptive spatial-temporal graph convolutional network for long-video ergonomic risk assessmentabstractErgonomic risk assessment (ERA) is commonly used to identify and analyze postures that are detrimental to the health of workers in industrial workplaces, which is vital to prevent work-related musculoskeletal disorders (WMSDs). Among the automatic approaches, algorithms based on graph convolutional networks (GCNs) have shown promising results in ERA using skeleton sequence as input. However, previous GCN-based methods still have certain limitations. First, the separated modeling of spatial and temporal information and the manually pre-defined topology of graph may restrict the representation diversity of the networks. Additionally, RNN-based temporal modeling often incurs high computational costs and fails to capture long-range temporal dependencies, thereby reducing flexibility in describing long videos. To overcome these challenges, in this study, we propose an attention-based adaptive spatial–temporal graph convolutional network (AAST-GCN), aiming to achieve effective and efficient action representation for ERA in long video. First, we employ an alternate modeling strategy to effectively capture the spatial–temporal information, and propose an improved adaptive adjacency matrix scheme to learn various coordination and relations of body-joints, thus enhancing the flexibility to model diverse postures. Furthermore, we introduce an efficient multi-scale temporal convolutional network as a replacement for RNN-based algorithms, enabling the network to extract various granularities of temporal features. Moreover, to make the network focuses on more valuable information, we employ a spatial–temporal interaction attention (STIA) module. Finally, the aforementioned modules are aggregated within a multi-task learning framework, with the action segmentation serving as the auxiliary task to further improve the accuracy of ERA. We conducted the ergonomic risk assessment on the UW-IOM and TUM Kitchen datasets using our network. Extensive experiments conducted on the most popular datasets UW-IOM and TUM Kitchen demonstrated that our proposed AAST-GCN outperforms other GCN-based methods. Ablation studies and visualization also prove the effectiveness of the individual sub-modules. Chengju Zhou, Jiayu Zeng, Lina Qiu, Shuxi Wang, Pingzhi Liu, Jiahui Pan 0003 |
Eng. Appl. Artif. Intell. | 6 |
| 2024 | Portable vision-based gait assessment for post-stroke rehabilitation using an attention-based lightweight CNN
Chengju Zhou, Daqin Feng, Shuyu Chen 0002, Nianming Ban, Jiahui Pan 0003 |
Expert Syst. Appl. | 5 |
| 2024 | A feature-enhanced hybrid attention network for traffic sign recognition in real scenesabstractAbstract Currently, traffic sign recognition techniques have been brought into the assistive driving of automobiles. However, small traffic sign recognition in real scenes is still a challenging task due to the class imbalance issue and the size limit of the traffic signs. To address the above issues, a feature‐enhanced hybrid attention network is proposed based on YOLOv5s for a small, fast, and accurate traffic sign detector. First, a series of online data augmentation strategies are designed in the preprocessing module for the model training. Second, the hybrid channel and spatial attention module CSAM are integrated into the backbone for a better feature extraction ability. Third, the channel attention module CAM is used in the detection head for a more efficient feature fusion ability. To validate the approach, extensive experiments are conducted based on the Tsinghua‐Tencent 100K dataset. It is found that the novel method achieves state‐of‐the‐art performance with only negligible increases in the model parameter and computational overhead. Specifically, the , parameters, and FLOPs are 85.8%, 7.13 M, and 16.1 G, respectively. Lewei He, Fucai Lan, Chuanzhe Zhou, Yaoguang Ye, Wencong Zhang, Bingzhi Chen, Jiahui Pan 0003 |
IET Image Process. | 7 |
| 2024 | A spatiotemporal network using a local spatial difference stack block for facial micro-expression recognition
Yan Hao, Jiacheng Liao, Zhuoran Deng, Zefeng Zheng, Jiahui Pan 0003 |
Multim. Tools Appl. | 7 |
| 2024 | SFT-SGAT: A semi-supervised fine-tuning self-supervised graph attention network for emotion recognition and consciousness detection
Lina Qiu, Liangquan Zhong, Weisen Feng, Chengju Zhou, Jiahui Pan 0003 |
Neural Networks | 6 |
| 2024 | Sequence-level affective level estimation based on pyramidal facial expression featuresabstractPeople tend to focus on changes in a certain complex human affect in the majority of practical applications of affective computing. Facial expression classification models are unable to represent all human affects through a limited number of expression categories. In this backdrop, this paper studies the Sequence-level affective level estimation (S-ALE), which is more relevant to real scenarios and can depict individual affective level in continuous manner. A spatio-temporal framework applied to S-ALE is proposed, which consists of a Facial Expression Features Pyramid Network (FEFPN) and a Temporal Transformer Encoder (TTE). FEFPN is capable of extracting pyramidal facial expression features, while TTE can effectively capture coarse-grained and fine-grained temporal variations of facial sequences. The proposed model is evaluated on six public datasets across three typical S-ALE tasks (engagement prediction, fatigue detection, and pain assessment), and experimental results show that our method is comparable to or outperforms the state-of-the-art algorithms. Jiacheng Liao, Yan Hao, Zhuoyi Zhou, Jiahui Pan 0003 |
Pattern Recognit. | 4 |
| 2024 | ST-SCGNN: A Spatio-Temporal Self-Constructing Graph Neural Network for Cross-Subject EEG-Based Emotion Recognition and Consciousness DetectionabstractIn this paper, a novel spatio-temporal self-constructing graph neural network (ST-SCGNN) is proposed for cross-subject emotion recognition and consciousness detection. For spatio-temporal feature generation, activation and connection pattern features are first extracted and then combined to leverage their complementary emotion-related information. Next, a self-constructing graph neural network with a spatio-temporal model is presented. Specifically, the graph structure of the neural network is dynamically updated by the self-constructing module of the input signal. Experiments based on the SEED and SEED-IV datasets showed that the model achieved average accuracies of 85.90% and 76.37%, respectively. Both values exceed the state-of-the-art metrics with the same protocol. In clinical besides, patients with disorders of consciousness (DOC) suffer severe brain injuries, and sufficient training data for EEG-based emotion recognition cannot be collected. Our proposed ST-SCGNN method for cross-subject emotion recognition was first attempted in training in ten healthy subjects and testing in eight patients with DOC. We found that two patients obtained accuracies significantly higher than chance level and showed similar neural patterns with healthy subjects. Covert consciousness and emotion-related abilities were thus demonstrated in these two patients. Our proposed ST-SCGNN for cross-subject emotion recognition could be a promising tool for consciousness detection in DOC patients. Jiahui Pan 0003, Rongming Liang, Zhipeng He 0001, Jingcong Li 0001, Yanbin He, Yuanqing Li 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | Twinenet: coupling features for synthesizing volume rendered images via convolutional encoder-decoders and multilayer perceptrons
Shengzhou Luo, Jingxing Xu, John Dingliana, Mingqiang Wei, Lewei He, Jiahui Pan 0003 |
Vis. Comput. | 7 |
| 2023 | Spatio-Temporal Swin Transformer-based 4-D EEG Emotion RecognitionabstractSeveral studies have shown that electroencephalogram (EEG) can be interfered by eye movements and facial muscle movements. These interferences can reduce the accuracy of EEG emotion recognition. In terms of model selection, EEG-based emotion recognition studies mainly use convolutional neural networks and similar methods. These methods rely on global differences to distinguish different emotional states, but overlook the influence of local EEG changes on emotional states. In this paper, we use four-dimensional artificial EEG features as input to a Spatio-Temporal Swin Transformer. The use of artificial EEG features is beneficial for the interpretability of the model. Spatiotemporal attention modules can extract secondary features and useful information from EEG signals while eliminating redundant noise interference. Compared with other emotion recognition methods, the advantage of our study is that the input and output of the spatio-temporal attention modules are consistent. It can adapt to the original input size of the model and can be widely applied as a separable module in other models. We use the public emotion EEG datasets SEED and SEED-IV to evaluate the feasibility and effectiveness of our model. The accuracy of single-subject and cross-subject emotion three-class classification on the SEED dataset is 94.13% ± 2.11% and 89.33% ± 4.37%, respectively. The accuracy of single-subject emotion four-class classification is 82.34% ± 8.17% on the SEED-IV dataset. Our study achieves high accuracy in emotion classification and real-time processing. Zongnan Chen, Jiarui Jin, Jiahui Pan 0003 |
BIBM | 3 |
| 2023 | Automatic Detection of Sleep Spindles and its Application in Patients with Acute Disorders of ConsciousnessabstractSleep spindles play an important role in human sleep and are considered to have great significance in predicting the prognosis of patients with acute disorders of consciousness (ADOC). Although previous studies have achieved high performance in the automatic detection of sleep spindles in normal subjects, the application in ADOC is very limited, and several challenges remain: 1) how to effectively detect patients’ spindles that may decrease in frequency; 2) how to improve the generality of the method to detect more electroencephalogram (EEG) events, such as K-complexes; and 3) how to intuitively reflect the relationship between patients’ spindle density and prognosis. To address the above challenges, we propose SpindleCatcher, a deep learning strategy to detect sleep spindles, and design an experiment to investigate the correlation between spindle density and prognosis in ADOC. SpindleCatcher jointly predicts the locations and durations of spindles in EEG, using a convolutional neural network to extract features from raw EEG signals and two modules for localization and classification tasks. Specifically, a frequency attention module is applied to better focus on signals in the desired frequency ranges to improve the performance of ADOC spindle detection. SpindleCatcher can also detect other EEG events, such as K-complexes. Experiments demonstrate that the proposed method exceeds the baseline methods on spindle detection with an overall recall of 0.817 and F1 score of 0.794 on the publicly available MASS2 dataset and an overall recall of 0.707 and F1 score of 0.681 on the patient dataset. The correlation experiment shows that there may be a strong positive correlation between the sleep spindle density of ADOCs and their outcomes. Zhenglang Yang, Jiahui Pan 0003 |
BIBM | 2 |
| 2023 | Engagement Recognition in Online Learning Based on an Improved Video Vision TransformerabstractOnline learning has gained wide attention and application due to its flexibility and convenience. However, due to the separation of time and space, the level of students' engagement is not easily informed by teachers, which affects the effectiveness of teaching. Automatic detection of students' engagement is an effective way to solve this problem. It can help teachers obtain timely feedback from students and adjust the teaching schedule. In this paper, transformer is first applied in engagement recognition and a novel network based on an improved video vision transformer (ViViT) is proposed to detect student engagement. A new transformer encoder, named Transformer Encoder with Low Complexity (TELC) is proposed. It adopts unit force operated attention (UFO-attention) to eliminate the nonlinearity of the original self-attention in standard ViViT and Patch Merger to fuse the input patches, which allows the network to significantly reduce computational complexity while improving performance. The proposed method is evaluated on the Dataset for Affective States in E-learning Environments (DAiSEE) and achieves an accuracy of 63.91% in the four-level classification task, which is superior to state-of-the-art methods. The experimental results demonstrate the effectiveness of our method, which is more suitable for the practical application of online learning. Zhuoyi Zhou, Jiahui Pan 0003 |
IJCNN | 3 |
| 2023 | A Multiple Command UAV Control System Based on a Hybrid Brain-Computer InterfaceabstractThe difficulty of UAV control in recent years lies in multidimensional movement in 3D space and improving control accuracy. To address this challenge, a UAV control method that incorporates a noninvasive hybrid brain-computer interface and gyroscope is proposed in this paper. We propose an efficient SSVEP deep learning network (CL-NET) based on one-dimensional convolution and a long short-term memory (LSTM) module that enables UAVs to move in the front, back, left and right directions. To improve the performance of CL-NET, an attention module to the network architecture is adopted. The takeoff and landing control of the UAV is realized by a blink state detector based on the electrooculogram (EOG) signal detection algorithm. The UAV was able to fly at a longitudinal tilt and rotate by detecting the current head posture with the help of a gyroscope. The proposed CL-NET model achieves an accuracy of 98.67% for the public dataset and an accuracy of 95.83% for the self-recorded dataset, which are both superior to the state-of-the-art models. In outdoor experiments involving six subjects, the proposed UAV control method reached an average information transfer rate of 44.64 bit/min. The efficiency and stability of our UAV control is thus demonstrated. With the proposed hybrid control strategy, our multidimensional UAV control can maintain excellent performance Qingfu Wu, Shanghong Xie, Jiahui Pan 0003 |
IJCNN | 5 |
| 2023 | Automatic Hemiplegia Gait Assessment for Post-Stroke by an Efficient Hybrid Attention-Based GhostNetabstractVision-based gait analysis provides the possibility to automatically and unobtrusively detect walking pattern alterations caused by stoke. Therefore, it can be used to determine the severity of stroke during stroke rehabilitation outside the hospital, which greatly releases the economic and labor burden on patients and their families. However, state-of-the-art deep learning algorithms for gait analysis usually suffer from high computational complexity and can even lead to overfitting problems on small-scale pathological gait datasets. To realize an efficient and effective system, we constructed a specially designed dataset and proposed a novel lightweight network to lean discriminative gait representation to map the input into one of the stroke severity levels. More specifically, a simulated hemiplegia gait dataset with multiple severity levels is first constructed, including sufficient 2D image sequences collected from 14 subjects. Different from the existing pathological datasets used for coarse classification, which only distinguish different pathological gait types, our proposed dataset is specifically designed for fine classification to assess the severity of hemiplegia that is defined according to medical prior. Second, considering that pathological datasets are usually small-scale, an attention-based lightweight network is proposed. In detail, a lightweight hybrid attention module (LHAM) based on the 1D adaptive convolution for channel attention interaction was developed to enhance the network's ability to integrate and focus on meaningful spatial and channel features. To further lighten the networks, a proposed efficient ghost module (EGM) is used in the bottleneck structure instead of the normal convolutional layer. Extensive experiments on both self-constructed and publicly available datasets demonstrate that the proposed efficient hybrid attention-based GhostNet realizes an effective and efficient gait analysis for stroke rehabilitation. Chengju Zhou, Daqin Feng, Lewei He, Nianming Ban, Shuxi Wang, Jiahui Pan 0003 |
IJCNN | 6 |
| 2023 | Combating Medical Label Noise via Robust Semi-supervised Contrastive Learning
Bingzhi Chen, Zhanhao Ye, Yishu Liu 0001, Zheng Zhang 0006, Jiahui Pan 0003, Guangming Lu 0002 |
MICCAI (1) | 5 |
| 2023 | A cascaded spatiotemporal attention network for dynamic facial expression recognition
Yaoguang Ye, Yongqi Pan, Jiahui Pan 0003 |
Appl. Intell. | 4 |
| 2023 | ICE-GCN: An interactional channel excitation-enhanced graph convolutional network for skeleton-based action recognitionabstractAbstract Thanks to the development of depth sensors and pose estimation algorithms, skeleton-based action recognition has become prevalent in the computer vision community. Most of the existing works are based on spatio-temporal graph convolutional network frameworks, which learn and treat all spatial or temporal features equally, ignoring the interaction with channel dimension to explore different contributions of different spatio-temporal patterns along the channel direction and thus losing the ability to distinguish confusing actions with subtle differences. In this paper, an interactional channel excitation (ICE) module is proposed to explore discriminative spatio-temporal features of actions by adaptively recalibrating channel-wise pattern maps. More specifically, a channel-wise spatial excitation (CSE) is incorporated to capture the crucial body global structure patterns to excite the spatial-sensitive channels. A channel-wise temporal excitation (CTE) is designed to learn temporal inter-frame dynamics information to excite the temporal-sensitive channels. ICE enhances different backbones as a plug-and-play module. Furthermore, we systematically investigate the strategies of graph topology and argue that complementary information is necessary for sophisticated action description. Finally, together equipped with ICE, an interactional channel excited graph convolutional network with complementary topology (ICE-GCN) is proposed and evaluated on three large-scale datasets, NTU RGB+D 60, NTU RGB+D 120, and Kinetics-Skeleton. Extensive experimental results and ablation studies demonstrate that our method outperforms other SOTAs and proves the effectiveness of individual sub-modules. The code will be published at https://github.com/shuxiwang/ICE-GCN . Shuxi Wang, Jiahui Pan 0003, Binyuan Huang, Pingzhi Liu, Zina Li, Chengju Zhou |
Mach. Vis. Appl. | 2 |
| 2023 | A novel semi-supervised meta learning method for subject-transfer brain-computer interfaceabstractThe brain-computer interface (BCI) provides a direct communication pathway between the human brain and external devices. However, the models trained for existing subjects perform poorly on new subjects, which is termed the subject calibration problem. In this paper, we propose a semi-supervised meta learning (SSML) method for subject-transfer calibration. The proposed SSML learns a model-agnostic meta learner with existing subjects and then fine-tunes the meta learner in a semi-supervised learning manner, i.e. using a few labelled samples and many unlabelled samples of the target subject for calibration. It is significant for BCI applications in which labelled data are scarce or expensive while unlabelled data are readily available. Three different BCI paradigms are tested: event-related potential detection, emotion recognition and sleep staging. The SSML achieved classification accuracies of 0.95, 0.89 and 0.83 in the benchmark datasets of three paradigms. The runtime complexity of SSML grows linearly as the number of samples of target subject increases so that is possible to apply it in real-time systems. This study is the first attempt to apply semi-supervised model-agnostic meta learning methodology for subject calibration. The experimental results demonstrated the effectiveness and potential of the SSML method for subject-transfer BCI applications. Jingcong Li 0001, Fei Wang 0026, Haiyun Huang, Feifei Qi, Jiahui Pan 0003 |
Neural Networks | 5 |
| 2023 | AITST - Affective EEG-based person identification via interrelated temporal-spatial transformer
Honghua Cai, Jiarui Jin, Liujiang Li, Yucui Huang, Jiahui Pan 0003 |
Pattern Recognit. Lett. | 6 |
| 2022 | Joint Temporal Convolutional Networks and Adversarial Discriminative Domain Adaptation for EEG-Based Cross-Subject Emotion RecognitionabstractCross-subject emotion recognition is one of the most challenging tasks in electroencephalogram (EEG)-based emotion recognition. To guarantee the constancy of feature representations across domains and to eliminate differences between domains, we explored the feasibility of combining temporal convolutional networks (TCNs) and adversarial discriminative domain adaptation (ADDA) algorithms in solving the problem of domain shift in EEG-based cross-subject emotion recognition. In light of EEG signals that have specific temporal properties, we chose the temporal model TCN as the feature encoder. To verify the validity of the proposed method, we conducted experiments on two public datasets: DEAP and DREAMER. The experimental results show that for the leave-one-subject-out evaluation, average accuracies of 64.33% (valence) and 63.25% (arousal) were obtained on the DEAP dataset, and average accuracies of 66.56% (valence) and 63.69% (arousal) were achieved on the DREAMER dataset. Extensive experiments demonstrate that our method for EEG-based cross-subject emotion recognition is effective. Zhipeng He 0001, Yongshi Zhong, Jiahui Pan 0003 |
ICASSP | 3 |
| 2022 | GaitMSTP: Multi-Granularity Spatio-Temporal Pyramid for Gait Recognition Under Complex Covariation ConditionsabstractGait, with its unique advantage of remote perception without any cooperation from the perceived subject, has become a popular biometric modality for human identity authentication. Diverse spatial representation and temporal modeling are crucial information for gait recognition, especially under covariation conditions. However, most existing algorithms do not fully and explicitly exploit the rich spatial-temporal clues in the gait sequences, leading to a decline in the discriminative ability of gait feature representations. In this paper, we propose a GaitMSTP network for gait recognition under complex covariation conditions, which explicitly models spatio-temporal representations at multi-granularity and multi-sematic levels. More specifically, a Multi-Granularity Temporal Pyramid (MGTP) module is incorporated to extract features at different temporal granularity, which simulates coarse- and fine-grained motion patterns at diverse temporal scales. A Multi-Granularity Spatial Pyramid (MGSP) is designed to capture global and local features at multiple spatial locations and scales. In addition, the multi-granularity spatial-temporal features extracted from shallow to deep semantic levels are further used for supervised learning, aiming to exploit both high-level and low-level of multi-sematic gait characteristics. Extensive experiments on the CASIA-B dataset show that our method outperforms the state-of-the-art algorithms for gait recognition in all scenarios. Binyuan Huang, Chengju Zhou, Jiahui Pan 0003 |
IJCB | 4 |
| 2022 | Emotion-related awareness detection for patients with disorders of consciousness via graph isomorphic networkabstractThe clinical diagnosis of patients with disorders of consciousness (DOC) mainly relies on behavioral scales. However, patients with DOC often have severe dyskinesia, which may lead to misdiagnosis of patients in the minimally conscious state (MCS) as patients in vegetative state (VS). In this paper, we propose an emotion-induced paradigm based on audio-visual stimulation, which can collect electroencephalogram (EEG) signals for consciousness detection without performing behavioral expression tasks. This paradigm exposes patients to emotional videos and stimulates them through video clips, which is more comfortable than event-related potential (ERP), steady-state visual evoked potential (SSVEP) and other paradigms. It effectively reduces the mental burden required by the patients. In terms of algorithms, a graph isomorphic network (GIN) is adopted to automatically classify VS and MCS, using emotional EEG signals from patients with DOC. The accuracy, precision, specificity, and sensitivity of our method are 93.70%, 93.77%, 89.44%, and 95.21%, respectively. Compared with the existing consciousness detection methods, our method has superior performance in consciousness detection for patients with DOC. The experimental results show that our method is feasible in distinguishing MCS and VS, and it is an effective extension of consciousness level detection in patients with DOC. Zhipeng He 0001, Yongshi Zhong, Jiahui Pan 0003 |
SMC | 3 |
| 2021 | Deep facial spatiotemporal network for engagement prediction in online learning
Jiacheng Liao, Jiahui Pan 0003 |
Appl. Intell. | 3 |
| 2021 | An EEG-Based Brain Computer Interface for Emotion Recognition and Its Application in Patients with Disorder of ConsciousnessabstractRecognizing human emotions based on electroencephalogram (EEG) signals has received a great deal of attentions. Most of the existing studies focused on offline analysis, and real-time emotion recognition using a brain computer interface (BCI) approach remains to be further investigated. In this paper, we proposed an EEG-based BCI system for emotion recognition. Specifically, two classes of video clips that represented positive and negative emotions were presented to the subjects one by one, while the EEG data were collected and processed simultaneously, and instant feedback was provided after each clip. Ten healthy subjects participated in the experiment and achieved a high average online accuracy of 91.5$\pm$6.34 percent. The experimental results demonstrated that the subjects emotions had been sufficiently evoked and efficiently recognized by our system. Clinically, patients with disorder of consciousness (DOC), such as coma, vegetative state, minimally conscious state and emergence minimally conscious state, suffer from motor impairment and generally cannot provide adequate emotion expressions. Consequently, doctors have difficulty in detecting the emotional states of these patients. Therefore, we applied our emotion recognition BCI system to patients with DOC. Eight DOC patients participated in our experiment, and three of them achieved significant online accuracy. The experimental results show that the proposed BCI system could be a promising tool to detect the emotional states of patients with DOC. Haiyun Huang, Qiuyou Xie, Jiahui Pan 0003, Yanbin He, Zhenfu Wen, Ronghao Yu, Yuanqing Li 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2020 | Speech emotion recognition using fusion of three multi-task learning-based classifiers: HSF-DNN, MS-CNN and LLD-RNNabstractSpeech emotion recognition plays an increasingly important role in emotional computing and is still a challenging task due to its complexity. In this study, we developed a framework integrating three distinctive classifiers: a deep neural network (DNN), a convolution neural network (CNN), and a recurrent neural network (RNN). The framework was used for categorical recognition of four discrete emotions (i.e., angry, happy, neutral and sad). Frame-level low-level descriptors (LLDs), segment-level mel-spectrograms (MS), and utterance-level outputs of high-level statistical functions (HSFs) on LLDs were passed to RNN, CNN, and DNN, separately. Three individual models of LLD-RNN, MS-CNN, and HSF-DNN were obtained. In the models of MS-CNN and LLD-RNN, the attention mechanism based weighted-pooling method was utilized to aggregate the CNN and RNN outputs. To effectively utilize the interdependencies between the two approaches of emotion description (discrete emotion categories and continuous emotion attributes), a multi-task learning strategy was implemented in these three models to acquire generalized features by simultaneously operating classification of discrete categories and regression of continuous attributes. Finally, a confidence-based fusion strategy was developed to integrate the power of different classifiers in recognizing different emotional states. Three experiments on emotion recognition based on the IEMOCAP corpus were conducted. Our experimental results show that the weighted pooling method based on attention mechanism endowed the neural networks with the capability to focus on emotionally salient parts. The generalized features learned in the multi-task learning helped the neural networks to achieve higher accuracies in the tasks of emotion classification. Furthermore, our proposed fusion system achieved weighted accuracy of 57.1% and unweighted accuracy of 58.3%, which were significantly higher than those of each individual classifier. The effectiveness of the proposed approach based on classifier fusion was thus validated. Zengwei Yao, Weihuang Liu, Yaqian Liu, Jiahui Pan 0003 |
Speech Commun. | 5 |
| 2016 | An EEG-Based brain-computer interface for emotion recognitionabstractIn this paper, an EEG-based brain-computer interface (BCI) system used for emotion recognition is proposed to detect two basic emotional states (happiness and sadness). Selection of frequency bands plays a vital role in distinguishing brain patterns associated with emotions. This paper explores a new method to select suitable subject-specific frequency bands instead of using fixed frequency bands for the emotion recognition. Common spatial pattern and support vector machine were employed to classify two emotional states. Two experiments involving six subjects were conducted to validate our method and BCI system. An average online accuracy of 74.17% for two classes was achieved. The data analysis results demonstrated that the proposed method based on subject-specific frequency bands outperformed the method based on the fixed frequency bands in terms of accuracy. Jiahui Pan 0003, Yuanqing Li 0001, Jun Wang 0002 |
IJCNN | 1 |
| 2016 | Multimodal BCIs: Target Detection, Multidimensional Control, and Awareness Evaluation in Patients With Disorder of ConsciousnessabstractDespite rapid advances in the study of brain–computer interfaces (BCIs) in recent decades, two fundamental challenges, namely, improvement of target detection performance and multidimensional control, continue to be major barriers for further development and applications. In this paper, we review the recent progress in multimodal BCIs (also called hybrid BCIs), which may provide potential solutions for addressing these challenges. In particular, improved target detection can be achieved by developing multimodal BCIs that utilize multiple brain patterns, multimodal signals, or multisensory stimuli. Furthermore, multidimensional object control can be accomplished by generating multiple control signals from different brain patterns or signal modalities. Here, we highlight several representative multimodal BCI systems by analyzing their paradigm designs, detection/control methods, and experimental results. To demonstrate their practicality, we report several initial clinical applications of these multimodal BCI systems, including awareness evaluation/detection in patients with disorder of consciousness (DOC). As an evolving research area, the study of multimodal BCIs is increasingly requiring more synergetic efforts from multiple disciplines for the exploration of the underlying brain mechanisms, the design of new effective paradigms and means of neurofeedback, and the expansion of the clinical applications of these systems. Yuanqing Li 0001, Jiahui Pan 0003, Jinyi Long, Tianyou Yu, Fei Wang 0026, Zhu Liang Yu, Wei Wu 0022 |
Proc. IEEE | 2 |