VLDB 2026 Research / reviewers in the wild / expert
Wei-Bang Jiang
dblp:311/0002 · also Weibang Jiang
· DBLP profile ↗
19ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0003-3759-5100ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EEG-DLite: Dataset Distillation for Efficient Large EEG Model TrainingabstractLarge-scale EEG foundation models have shown strong generalization across a range of downstream tasks, but their training remains resource-intensive due to the volume and variable quality of EEG data. In this work, we introduce EEG-DLite, a data distillation framework that enables more efficient pre-training by selectively removing noisy and redundant samples from large EEG datasets. EEG-DLite begins by encoding EEG segments into compact latent representations using a self-supervised autoencoder, allowing sample selection to be performed efficiently and with reduced sensitivity to noise. Based on these representations, EEG-DLite filters out outliers and minimizes redundancy, resulting in a smaller yet informative subset that retains the diversity essential for effective foundation model training. Through extensive experiments, we demonstrate that training on only 5 percent of a 2,500-hour dataset curated with EEG-DLite yields performance comparable to, and in some cases better than, training on the full dataset across multiple downstream tasks. To our knowledge, this is the first systematic study of pre-training data distillation in the context of EEG foundation models. EEG-DLite provides a scalable and practical path toward more effective and efficient physiological foundation modeling. Yuting Tang, Wei-Bang Jiang, Shanglin Li, Yong Li 0032, Xinliang Zhou, Yi Ding 0012, Cuntai Guan |
AAAI | 2 |
| 2026 | EEG-to-gait decoding via phase-aware representation learningabstractAccurate decoding of lower-limb motion from EEG signals is essential for advancing brain-computer interface (BCI) applications in movement intent recognition and control. This study presents NeuroDyGait, a two-stage, phase-aware EEG-to-gait decoding framework that explicitly models temporal continuity and domain relationships. To address challenges of causal, phase-consistent prediction and cross-subject variability, Stage I learns semantically aligned EEG-motion embeddings via relative contrastive learning with a cross-attention-based metric, while Stage II performs domain relation-aware decoding through dynamic fusion of session-specific heads. Comprehensive experiments on two benchmark datasets (GED and FMD) show substantial gains over baselines, including a recent 2025 model EEG2GAIT. The framework generalizes to unseen subjects and maintains inference latency below 5 ms per window, satisfying real-time BCI requirements. Visualization of learned attention and phase-specific cortical saliency maps further reveals interpretable neural correlates of gait phases. Future extensions will target rehabilitation populations and multimodal integration. Xi Fu, Wei-Bang Jiang, Rui Liu 0034, Gernot R. Müller-Putz, Cuntai Guan |
Neural Networks | 2 |
| 2026 | SEED-OLF: A Novel EEG Dataset With Olfactory Stimulation for Emotion Recognition
Wei-Bang Jiang, Wei-Long Zheng, Bao-Liang Lu |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | Multi-Scale Attention-Based Dense Spatial-Temporal Model for Emotion Induction in Response to Olfactory StimuliabstractAffective Brain-Computer Interfaces (aBCIs) have attracted growing attention due to their potential for decoding human emotional states through electroencephalogram (EEG) signals. However, existing deep learning models often struggle to fully capture both the spatial and temporal dependencies in EEG data, resulting in suboptimal performance in emotion classification. To address these challenges, we propose a novel Multi-Scale Attention-Based Dense Spatial-Temporal Model (MSADM). This model leverages temporal attention, multi-scale dense feature extraction, and attention-based feature fusion to effectively capture and enhance the complicated spatial-temporal dependencies within EEG data, thereby improving the representations. Furthermore, we introduce a new olfactory-based emotion induction paradigm, which effectively mitigates the limitations of traditional visual and auditory stimuli by providing more stable and sustained emotional responses. We also present a novel EEG dataset involving 32 subjects, developed through a pilot study to select appropriate olfactory stimuli. Experimental results demonstrate that our model significantly outperforms existing methods across multiple metrics, demonstrating the effectiveness of both the proposed model and emotion induction paradigm. This study underscores the potential of olfactory stimuli and advanced spatial-temporal modeling techniques for enhancing the robustness and performance of emotion recognition. Wei-Bang Jiang, Wei-Long Zheng, Bao-Liang Lu |
ICASSP | 2 |
| 2025 | NeuroLM: A Universal Multi-task Foundation Model for Bridging the Gap between Language and EEG SignalsabstractRecent advancements for large-scale pre-training with neural signals such as electroencephalogram (EEG) have shown promising results, significantly boosting the development of brain-computer interfaces (BCIs) and healthcare. However, these pre-trained models often require full fine-tuning on each downstream task to achieve substantial improvements, limiting their versatility and usability, and leading to considerable resource wastage. To tackle these challenges, we propose NeuroLM, the first multi-task foundation model that leverages the capabilities of Large Language Models (LLMs) by regarding EEG signals as a foreign language, endowing the model with multi-task learning and inference capabilities. Our approach begins with learning a text-aligned neural tokenizer through vector-quantized temporal-frequency prediction, which encodes EEG signals into discrete neural tokens. These EEG tokens, generated by the frozen vector-quantized (VQ) encoder, are then fed into an LLM that learns causal EEG information via multi-channel autoregression. Consequently, NeuroLM can understand both EEG and language modalities. Finally, multi-task instruction tuning adapts NeuroLM to various downstream tasks. We are the first to demonstrate that, by specific incorporation with LLMs, NeuroLM unifies diverse EEG tasks within a single model through instruction tuning. The largest variant NeuroLM-XL has record-breaking 1.7B parameters for EEG signal processing, and is pre-trained on a large-scale corpus comprising approximately 25,000-hour EEG data. When evaluated on six diverse downstream datasets, NeuroLM showcases the huge potential of this multi-task learning paradigm. Wei-Bang Jiang, Yansen Wang, Bao-Liang Lu, Dongsheng Li 0002 |
ICLR | 1 |
| 2025 | Hierarchical Emotion Transformer for Multimodal Joint Emotion Category and Intensity Recognition
Tian-Fang Ma, Wei-Bang Jiang, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (1) | 2 |
| 2025 | SEED-VII: A Multimodal Dataset of Six Basic Emotions With Continuous Labels for Emotion RecognitionabstractRecognizing emotions from physiological signals is a topic that has garnered widespread interest, and research continues to develop novel techniques for perceiving emotions. However, the emergence of deep learning has highlighted the need for comprehensive and high-quality emotional datasets that enable the accurate decoding of human emotions. To systematically explore human emotions, we develop a multimodal dataset consisting of six basic (happiness, sadness, fear, disgust, surprise, and anger) emotions and the neutral emotion, named SEED-VII. This multimodal dataset includes electroencephalography (EEG) and eye movement signals. The seven emotions in SEED-VII are elicited by 80 different videos and fully investigated with continuous labels that indicate the intensity levels of the corresponding emotions. Additionally, we propose a novel Multimodal Adaptive Emotion Transformer (MAET), that can flexibly process both unimodal and multimodal inputs. Adversarial training is utilized in the MAET to mitigate subject discrepancies, which enhances domain generalization. Our extensive experiments, encompassing both subject-dependent and cross-subject conditions, demonstrate the superior performance of the MAET in terms of handling various inputs. Continuous labels are used to filter the data with high emotional intensity, and this strategy is proven to be effective for attaining improved emotion recognition performance. Furthermore, complementary properties between the EEG signals and eye movements and stable neural patterns of the seven emotions are observed. Wei-Bang Jiang, Xuan-Hao Liu, Wei-Long Zheng, Bao-Liang Lu |
IEEE Trans. Affect. Comput. | 1 |
| 2024 | MoGE: Mixture of Graph Experts for Cross-subject Emotion Recognition via Decomposing EEGabstractDecoding emotions of previously unseen subjects from electroencephalography (EEG) signals is challenging due to the inter-subject variability. Domain Generalization (DG) methods aim to mitigate the domain shift among different subjects. Once trained, a DG model can be directly deployed on new subjects without any calibration phase. While existing DG studies on cross-subject emotion recognition mainly focus on the design of loss function for domain alignment or regularization, we introduce Sparse Mixture of Graph Experts (MoGE) model to explore DG issues from a new perspective, i.e. the design of the neural architecture. In the MoGE model, routers allocate each EEG channel to a specialized expert, thereby facilitating the decomposition of the intricate brain into distinct functional areas. Extensive experiments on three public datasets demonstrate that compared to other DG methods, our MoGE model trained with empirical risk minimization (ERM) achieves the state-of-the-art (SOTA) accuracies, 88.0%, 74.3%, and 81.8% on SEED, SEED-IV, and SEED-V datasets, respectively. Our code is available at https://github.com/XuanhaoLiu/MoGE. Xuan-Hao Liu, Wei-Bang Jiang, Wei-Long Zheng, Bao-Liang Lu |
BIBM | 2 |
| 2024 | Functional Emotion Transformer for EEG-Assisted Cross-Modal Emotion RecognitionabstractMultimodal emotion recognition based on electroencephalography (EEG) and eye movements has attracted increasing attention due to their high performance and complementary properties. However, there are two challenges that hinder its practical applications: the inconvenient EEG data collection and high-cost data annotation. In contrast, eye movements are convenient to obtain and process in real scenarios. To combine high performance of EEG and easy setups of eye tracking, we propose a novel EEG-assisted Contrastive Learning Framework with a Functional Emotion Transformer (ECO-FET) for cross-modal emotion recognition. ECO-FET leverages both the functional brain connectivity and the spectral-spatial-temporal domain of EEG signals simultaneously, which dramatically benefit the learning of eye movements. The whole process consists of three phases: pre-training, test, and fine-tuning. ECO-FET exploits the complementary information provided by multiple modalities during pre-training in order to improve the performance of unimodal models. In the pre-training phase, unlabeled EEG and eye movement data are fed into the model to contrastively learn the emotional latent representations between the two modalities, while in the test phase, eye movements and few labeled EEG samples are used to predict different emotions. Experimental results on three public datasets demonstrate that ECO-FET surpasses the state-of-the-art dramatically. Wei-Bang Jiang, Ziyi Li 0003, Wei-Long Zheng, Bao-Liang Lu |
ICASSP | 1 |
| 2024 | CEMOAE: A Dynamic Autoencoder with Masked Channel Modeling for Robust EEG-Based Emotion RecognitionabstractEmotion recognition through electroencephalography (EEG) has been an area of active research, but the inherent sensitivity of EEG signals to noise and artifacts poses significant challenges, especially in real-world settings. These complications often necessitate the removal of corrupted channels, making it crucial to develop robust models capable of maintaining performance even when few channels are available. To address this, we propose the Corrupted EMOtion AutoEncoder (CEMOAE), an innovative approach that leverages masked channel modeling to maintain robust performance, achieved through three components: masked autoencoder pretraining for robust representation learning, random masked auxiliary task for implicit modeling of channel corruption, and masked auto-repair to explicitly narrow the data distribution gap between high-quality and corrupted EEG signals. Specifically, we first pretrain a masked autoencoder with the dynamic masking strategy for feature extractor initialization and channel recovery. During the finetuning stage, we mask EEG data using the auxiliary task to mimic real-world EEG corruption. We then employ the pretrained autoencoder to repair these signals and finetune the feature extractor for emotion recognition. Experiments on the SEED dataset demonstrate that CEMOAE achieves SOTA performance for emotion recognition under the random channel corruption simulation, validating the effectiveness of the proposed techniques. Yu-Ting Lan, Wei-Bang Jiang, Wei-Long Zheng, Bao-Liang Lu |
ICASSP | 2 |
| 2024 | Large Brain Model for Learning Generic Representations with Tremendous EEG Data in BCIabstractThe current electroencephalogram (EEG) based deep learning models are typically designed for specific datasets and applications in brain-computer interaction (BCI), limiting the scale of the models and thus diminishing their perceptual capabilities and generalizability. Recently, Large Language Models (LLMs) have achieved unprecedented success in text processing, prompting us to explore the capabilities of Large EEG Models (LEMs). We hope that LEMs can break through the limitations of different task types of EEG datasets, and obtain universal perceptual capabilities of EEG signals through unsupervised pre-training. Then the models can be fine-tuned for different downstream tasks. However, compared to text data, the volume of EEG datasets is generally small and the format varies widely. For example, there can be mismatched numbers of electrodes, unequal length data samples, varied task designs, and low signal-to-noise ratio. To overcome these challenges, we propose a unified foundation model for EEG called Large Brain Model (LaBraM). LaBraM enables cross-dataset learning by segmenting the EEG signals into EEG channel patches. Vector-quantized neural spectrum prediction is used to train a semantically rich neural tokenizer that encodes continuous raw EEG channel patches into compact neural codes. We then pre-train neural Transformers by predicting the original neural codes for the masked EEG channel patches. The LaBraMs were pre-trained on about 2,500 hours of various types of EEG signals from around 20 datasets and validated on multiple different types of downstream tasks. Experiments on abnormal detection, event type classification, emotion recognition, and gait prediction show that our LaBraM outperforms all compared SOTA methods in their respective fields. Our code is available at https://github.com/935963004/LaBraM. Wei-Bang Jiang, Li-Ming Zhao, Bao-Liang Lu |
ICLR | 1 |
| 2024 | REmoNet: Reducing Emotional Label Noise via Multi-regularized Self-supervisionabstractEmotion recognition based on electroencephalogram (EEG) has garnered increasing attention in recent years due to non-invasiveness and high reliability of EEG measurements. Despite the promising performance achieved by numerous existing methods, several challenges persist. Firstly, there is the challenge of emotional label noise, stemming from the assumption that emotions remain consistently evoked and stable throughout the entirety of video observation. Such an assumption proves difficult to uphold in practical experimental settings, leading to discrepancies between EEG signals and anticipated emotional states. In addition, there's a need for comprehensive capture of temporal-spatial-spectral characteristics of EEG signals and cope with low signal-to-noise ratio (SNR) issues. To tackle these challenges, we propose a comprehensive pipeline named REmoNet, which leverages novel self-supervised techniques and multi-regularized co-learning. Two self-supervised methods, including masked channel modeling via temporal-spectral transformation and emotion contrastive learning, are introduced to facilitate the comprehensive understanding and extraction of emotion-relevant EEG representations during pre-training. Additionally, fine-tuning with multi-regularized co-learning exploits feature-dependent information through intrinsic similarity, resulting in mitigating emotional label noise. Experimental evaluations on two public datasets demonstrate that our proposed approach, REmoNet, surpasses existing state-of-the-art methods, showcasing its effectiveness in simultaneously addressing raw EEG signals and noisy emotional labels. Wei-Bang Jiang, Yu-Ting Lan, Bao-Liang Lu |
ACM Multimedia | 1 |
| 2024 | Du-IN: Discrete units-guided mask modeling for decoding speech from Intracranial Neural signalsabstractInvasive brain-computer interfaces with Electrocorticography (ECoG) have shown promise for high-performance speech decoding in medical applications, but less damaging methods like intracranial stereo-electroencephalography (sEEG) remain underexplored. With rapid advances in representation learning, leveraging abundant recordings to enhance speech decoding is increasingly attractive. However, popular methods often pre-train temporal models based on brain-level tokens, overlooking that brain activities in different regions are highly desynchronized during tasks. Alternatively, they pre-train spatial-temporal models based on channel-level tokens but fail to evaluate them on challenging tasks like speech decoding, which requires intricate processing in specific language-related areas. To address this issue, we collected a well-annotated Chinese word-reading sEEG dataset targeting language-related brain networks from 12 subjects. Using this benchmark, we developed the Du-IN model, which extracts contextual embeddings based on region-level tokens through discrete codex-guided mask modeling. Our model achieves state-of-the-art performance on the 61-word classification task, surpassing all baselines. Model comparisons and ablation studies reveal that our design choices, including (\romannumeral1) temporal modeling based on region-level tokens by utilizing 1D depthwise convolution to fuse channels in the ventral sensorimotor cortex (vSMC) and superior temporal gyrus (STG) and (\romannumeral2) self-supervision through discrete codex-guided mask modeling, significantly contribute to this performance. Overall, our approach -- inspired by neuroscience findings and capitalizing on region-level representations from specific brain regions -- is suitable for invasive brain modeling and represents a promising neuro-inspired AI approach in brain-computer interfaces. Code and dataset are available at https://github.com/liulab-repository/Du-IN. Haiteng Wang, Wei-Bang Jiang, Zhongtao Chen, Pei-Yang Lin, Peng-Hu Wei, Guo-Guang Zhao, Yun-Zhe Liu |
NeurIPS | 3 |
| 2023 | Elastic Graph Transformer Networks for EEG-Based Emotion RecognitionabstractElectroencephalogram (EEG) has been applied in emotion recognition due to excellent temporal resolution with less competitive spatial resolution. This leads to the consequence that the majority of EEG-based emotion recognition models emphasize on exploiting temporal features while ignoring the efficient information provided by spatial resolution. To extract more informative representations, we propose an elastic Graph Transformer network for emotion recognition (EmoGT) inspired by the advantages of Transformer in time-series analysis and the superior performance of graph convolutional networks in topological analysis. Moreover, it is able to be flexibly expanded to cope with multimodal inputs by employing specially designed structures. Experimental results on 3 public datasets demonstrate that our models outperform the state-of-the-art results by 3% on average in both single and multimodal cases, indicating the effectiveness of utilizing temporal and spatial information simultaneously. Wei-Bang Jiang, Xu Yan 0007, Wei-Long Zheng, Bao-Liang Lu |
ICASSP | 1 |
| 2023 | Two-Stream Spectral-Temporal Denoising Network for End-to-End Robust EEG-Based Emotion Recognition
Xuan-Hao Liu, Wei-Bang Jiang, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (3) | 2 |
| 2023 | Multimodal Adaptive Emotion Transformer with Flexible Modality Inputs on A Novel Dataset with Continuous LabelsabstractEmotion recognition from physiological signals is a topic of widespread interest, and researchers continue to develop novel techniques for perceiving emotions. However, the emergence of deep learning has highlighted the need for high-quality emotional datasets to accurately decode human emotions. In this study, we present a novel multimodal emotion dataset that incorporates electroencephalography (EEG) and eye movement signals to systematically explore human emotions. Seven basic emotions (happy, sad, fear, disgust, surprise, anger, and neutral) are elicited by a large number of 80 videos and fully investigated with continuous labels that indicate the intensity of the corresponding emotions. Additionally, we propose a novel Multimodal Adaptive Emotion Transformer (MAET), that can flexibly process both unimodal and multimodal inputs. Adversarial training is utilized in MAET to mitigate subject discrepancy, which enhances domain generalization. Our extensive experiments, encompassing both subject-dependent and cross-subject conditions, demonstrate MAET's superior performance in handling various inputs. The filtering of data for high emotional evocation using continuous labels proved to be effective in the experiments. Furthermore, the complementary properties between EEG and eye movements are observed. Our code is available at https://github.com/935963004/MAET. Wei-Bang Jiang, Xuan-Hao Liu, Wei-Long Zheng, Bao-Liang Lu |
ACM Multimedia | 1 |
| 2022 | Increasing the Stability of EEG-based Emotion Recognition with a Variant of Neural ProcessesabstractElectroencephalography (EEG) signals provide an incredible promising access to decode human emotions in affective computing. Many approaches have been applied to building EEG-based affective models and much endeavor is made to improve the performance of affective models, where test data has similar quality as training data. However, due to the strong sensitivity of EEG to external factors such as body movement and electromagnetic interference, EEG signals usually have a lot of noise and subjects have to remain as motionless as possible in a quiet environment, which is difficult to be satisfied in real applications and severely influences the user experience. To deal with this problem, a common way is to drop those channels with intensive noise. However, this results in the loss of critical information and neither existing machine learning (ML) nor deep learning (DL) approaches can handle this situation well, especially when too many channels are missed. In this paper, we propose a robust variant of the neural processes model and evaluate the stability of our model under various circumstances to simulate random data corruption in real applications. We conduct two categories of experiments by controlling the number and the places of missing channels separately and compare with classical ML and DL models. The results demonstrate that the performance of our model significantly outperforms the existing models, even using only five channels. We also explore the critical brain regions in the current EEG electrode distribution. Final performance manifests our model has extreme stability in dealing with intractable situations and sheds light on the widespread usage of portable EEG-based affective computing. Yan-Kai Liu, Wei-Bang Jiang, Bao-Liang Lu |
IJCNN | 2 |
| 2021 | Discriminating Surprise and Anger from EEG and Eye Movements with a Graph NetworkabstractEmotion recognition based on EEG and eye movement signals has been studied extensively due to the reliability and stability of signals. The separability of four basic emotions, happy, sad, disgust and fear, has been systematically studied in the existing work. However, there is less research on the emotions of anger and surprise since they are more difficult to be elicited in lab settings. This paper investigates the discrimination ability of EEG and eye movement signals for surprise and anger. To this end, we design a stimulus paradigm that can effectively elicit surprise and anger. We propose a novel Graph Convolutional Network with Channel Attention (GCNCA) to classify three emotions, anger, surprise and neutrality. Experimental results indicate that: a) the proposed GCNCA model has an excellent classification accuracy of 86.47% using EEG and 84.22% using eye movement signals, which are better than other baseline methods; b) EEG and eye movements have a good ability to discriminate surprise and anger, while EEG performs better than eye movements; c) the high-frequency bands of EEG are more distinguishable on classifying surprise and anger than the low-frequency bands; d) there are some differences in neural patterns between surprise and anger, meanwhile critical channels and channel connections of EEG are found. Wei-Bang Jiang, Li-Ming Zhao, Bao-Liang Lu |
BIBM | 1 |
| 2021 | Emotion Transformer Fusion: Complementary Representation Properties of EEG and Eye Movements on Recognizing Anger and SurpriseabstractEmotion recognition plays an important role in di-agnosing and treating many mental disorders as well as affective computing. Among six basic emotions, anger and surprise are relatively hard to be elicited in lab settings, and the complementary representation properties of encephalography (EEG) and eye movement signals on recognizing anger and surprise emotions remain unknown. Although the transformer architecture has the ability of parallelism which avoids many sequential operations as recurrent and convolutional layers, the knowledge of its performance and effectiveness on multimodal emotion recognition from EEG and eye movement signals is limited. To tackle these issues, we elaborately design the experiment and stimuli materials to effectively elicit surprise, anger, and neutral emotions, and propose an Emotion Transformer Fusion (ETF) model based on pure attention mechanism. Results of extensive experiments with multiple models on our dataset indicate that the complementary information of EEG and eye movements significantly improves the performance of discriminating anger, surprise and neutral emotions. Meanwhile, our proposed architecture outperforms baseline models with higher parallelism, which proves the capability of Transformer based architecture on multimodal emotion recognition with EEG and eye movement signals. Wei-Bang Jiang, Rui Li 0048, Bao-Liang Lu |
BIBM | 2 |