Siqi Cai 0002

dblp:241/3543-2 · DBLP profile ↗
← Back
36ranked-venue papers
9as first author
34since 2021 · last 2026
0000-0003-3282-9246ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 4 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Rep Deep & Machine Learning: Exemplar-Free Continual Video Action Recognition via Slow-Fast Collaborative Learning
abstract
In real-world applications, video action recognition models must continuously learn new action categories while retaining previously acquired knowledge. However, most existing approaches rely on storing historical data for replay, which introduces storage burdens and raises data privacy concerns. To address these challenges, we investigate the problem of Exemplar-Free Continual Video Action Recognition (EF-CVAR) and propose a novel framework named Slow-Fast Collaborative Learning (SFCL). SFCL integrates two complementary learning paradigms: a slow branch based on gradient-driven deep learning, which provides strong adaptability to new tasks, and a fast branch based on analytic learning (e.g., Recursive Least Squares), which efficiently preserves old knowledge without requiring access to past samples. To enable effective collaboration between the two branches, we design the Slow-Fast Dynamic Re-parameterization (SFDR) mechanism for adaptive fusion, and the Knowledge Reflection Mechanism (KRM), which mitigates forgetting and task-recency bias via pseudo-feature generation and dual-level knowledge distillation. Extensive experiments on UCF101, HMDB51, and Something-Something V2 demonstrate that SFCL achieves superior performance compared to existing replay-based methods, despite being exemplar-free. Notably, in long-duration continual learning scenarios, SFCL exhibits remarkable robustness, achieving up to a 30.39\% improvement in accuracy over baselines while maintaining a low forgetting rate, highlighting its scalability and effectiveness in real-world video recognition tasks.
Xueyi Zhang 0001, Siqi Cai 0002, Mingrui Lao, Yanming Guo, Huiping Zhuang
AAAI5
2026 Explainable spatial-temporal-spectral meta-learning for subject-independent EEG-based emotion recognition
Siqi Cai 0002, Xueyi Zhang 0001, Xian Tang
Expert Syst. Appl.2
2026 Spiking neural networks for EEG signal analysis: From theory to practice
Siqi Cai 0002, Zheyuan Lin, Wenjie Wei, Shuai Wang 0058, Malu Zhang, Tanja Schultz, Haizhou Li 0001
Neural Networks1
2026 Closed-loop correction reprogramming for fine-grained visual prompting
Xueyi Zhang 0001, Siqi Cai 0002, Mingrui Lao, Haizhou Li 0001
Neural Networks3
2026 The effect of speech representations on EEG-based auditory attention detection
Siqi Cai 0002, Haizhou Li 0001
Pattern Recognit. Lett.3
2025 ATGnet: Adaptive Temporal Graph Network for EEG-enabled Sound Source Tracking in Cocktail Party Scenarios
abstract
Decoding selective auditory attention from electroencephalography (EEG) signals has gained considerable interest. However, few studies have looked into tracking the dynamic trajectory of moving sound source in complex auditory environments, e.g. with multiple moving speakers. We propose a novel model, namely Adaptive Temporal Graph Network (ATGnet), to continuously track the sound source trajectory using spatial-temporal EEG representations. ATGnet incorporates an adaptive graph topology to extract spatial features, and a graph-convolutional long short-term memory (GC-LSTM) network to capture spatial-temporal dependency. We evaluated ATGnet by performing within-subject leave-one-trial-out cross-validation on EEG signals from 10 participants. Experiment results indicate that ATGnet effectively overcomes the variation of signals across trials and subjects. They further confirm that ATGnet robustly tracks both attended and unattended sound sources, and significantly outperforms traditional methods. ATGnet offers a promising solution to continuous sound source tracking in dynamic conditions, with potential applications in neuro-steered hearing devices.
Saurav Pahuja, Gabriel Ivucic, Siqi Cai 0002, Dashanka De Silva, Tanja Schultz, Haizhou Li 0001
ICASSP3
2025 Generalization-Preserved Learning: Closing the Backdoor to Catastrophic Forgetting in Continual Deepfake Detection
Xueyi Zhang 0001, Peiyin Zhu, Zhiyuan Yan 0002, Jikang Cheng, Mingrui Lao, Siqi Cai 0002, Yanming Guo
ICCV7
2025 Decoding Listener's Identity: Person Identification from EEG Signals Using a Lightweight Spiking Transformer
abstract
EEG-based person identification enables applications in security, personalized brain-computer interfaces (BCIs), and cognitive monitoring. However, existing techniques often rely on deep learning architectures at high computational cost, limiting their scope of applications. In this study, we propose a novel EEG person identification approach using spiking neural networks (SNNs) with a lightweight spiking transformer for efficiency and effectiveness. The proposed SNN model is capable of handling the temporal complexities inherent in EEG signals. On the EEG-Music Emotion Recognition Challenge dataset, the proposed model achieves 100% classification accuracy with less than 10% energy consumption of traditional deep neural networks. This study offers a promising direction for energy-efficient and high-performance BCIs. The source code is available at https://github.com/PatrickZLin/Decode-ListenerIdentity.
Zheyuan Lin, Siqi Cai 0002, Haizhou Li 0001
INTERSPEECH2
2025 GTAnet: Geometry-Guided Temporal Attention for EEG-Based Sound Source Tracking in Cocktail Party Scenarios
Saurav Pahuja, Gabriel Ivucic, Siqi Cai 0002, Dashanka De Silva, Haizhou Li 0001, Tanja Schultz
INTERSPEECH3
2025 NeuroSpex+: Dual-Task Training of Neuro-Guided Speaker Extraction with Speech Envelope and Waveform
Dashanka De Silva, Siqi Cai 0002, Saurav Pahuja, Tanja Schultz, Haizhou Li 0001
INTERSPEECH2
2025 Choose Your Expert: Uncertainty-Guided Expert Selection for Continual Deepfake Detection
abstract
The rapid evolution of deepfake techniques presents dual challenges for detection models: adapting to continuously shifting attack distributions while retaining previously learned knowledge. Although recent continual deepfake detection methods have made progress, they often rely on replay-based training, which limits scalability and deployment. Meanwhile, the task structure of deepfake detection offers a unique opportunity that remains under-explored: it is inherently a binary classification problem with a fixed label space, where the main difficulty lies in distributional drift rather than class expansion. This insight enables the modeling of each incremental distribution shift as a dedicated expert, focusing on specific forgery patterns. To this end, we propose a novel analytically driven, replay-free continual detection framework that eliminates the need for iterative gradient updates. In this framework, task-specific experts are constructed via closed-form ridge regression, requiring only a single forward pass and ensuring non-interference with previous tasks. To enhance the model's capacity for fine-grained forgery recognition, we introduce a lightweight Forgery-Aware Residual Enhancer (FARE). At inference, an Uncertainty-Guided Expert Selection module (UGES) dynamically routes each sample to the most confident expert, which does not require prior knowledge of the attack type. The proposed framework achieves a favorable trade-off between efficiency, privacy, and generalization. It achieves state-of-the-art performance across four benchmark datasets, with an average accuracy of 91.82% and only 1.78% forgetting. Notably, it improves cross-forgery generalization by 9.28% on unseen forgery types, demonstrating strong generalization.
Xueyi Zhang 0001, Peiyin Zhu, Jinping Sui, Xiaoda Yang, Mingrui Lao, Siqi Cai 0002, Yanming Guo, Jun Tang 0001
ACM Multimedia7
2025 EventLip: Enhancing Event-Based Lip Reading via Frequency-Aware Spatiotemporal Hypergraph Modeling
abstract
Event cameras, with their microsecond-level temporal resolution and sparse visual encoding, provide a transformative paradigm for automatic lip reading (ALR). However, event data inherently lack explicit spatial structure and exhibit a pronounced frequency-domain bias. The low-frequency components fail to capture crucial lip structural information, which fundamentally impedes the modeling of intra-frame topological dependencies and inter-frame semantic evolution-both of which are critical for robust lip reading. To this end, we propose FAST-HG, a Frequency-Aware SpatioTemporal HyperGraph framework specifically designed for event-based lip reading. First, we apply low-frequency perturbation to improve the model's robustness for capturing discriminative features, and integrate adaptive high-frequency filtering to enhance edge-aware representations. Then, we construct a Spatial Region Hypergraph (SRH) and a Temporal Semantic Hypergraph (TSH). The former captures intra-frame topological dependencies among lip regions, while the latter explicitly models inter-frame structural associations throughout the lip movement process, enabling the model to capture discriminative patterns in lip dynamics. Furthermore, we propose a viseme-aware label smoothing strategy, where a novel viseme-level edit distance is designed to quantify visual similarities between classes and guide the construction of soft labels. FAST-HG achieves 79.85% and 84.03% accuracy on the DVS-Lip and DVS-LRW100 datasets, respectively, significantly outperforming prior methods and establishing a new benchmark for event-based lip reading.
Xueyi Zhang 0001, Jialu Sun, Xianghu Yue, Tianfang Xiao, Siqi Cai 0002, Mingrui Lao, Haizhou Li 0001
ACM Multimedia6
2025 TrustCLIP: Learning from Noisy Labels via Semantic Label Verification and Trust-aligned Gradient Projection
abstract
Prompt learning has emerged as an efficient adaptation paradigm for vision-language models (VLMs), yet it remains highly vulnerable to label noise, which limits its real-world applicability. We propose TrustCLIP, a noise-robust prompt tuning framework that leverages the inherent semantic structure of CLIP through two key components: Semantic Label Verification (SLV) and Trust-aligned Gradient Projection (TGP). SLV defines a semantic trust boundary based on CLIP's zero-shot predictions to identify reliable samples for standard supervised training. For uncertain samples, TGP projects their gradients into a trust-aligned subspace constructed from the gradients of clean samples, thereby preserving semantically aligned learning signals while suppressing noise-induced optimization drift. Unlike prior approaches, TrustCLIP doesn't require additional parameters, loss reweighting, or uncertainty estimation. Extensive experiments on 7 benchmark datasets with both synthetic and real-world noisy labels demonstrate that TrustCLIP consistently outperforms state-of-the-art methods in terms of both robustness and transferability.
Xueyi Zhang 0001, Peiyin Zhu, Mingrui Lao, Siqi Cai 0002, Yanming Guo, Haizhou Li 0001
ACM Multimedia6
2025 Listening to the Brain: Multi-Band sEEG Auditory Reconstruction via Dynamic Spatio-Temporal Hypergraphs
abstract
Speech is a fundamental form of human communication, and speech perception constitutes the initial stage of language comprehension. Although brain-to-speech interface technologies have made significant progress in recent years, most existing studies focus on neural decoding during speech production. Such approaches heavily rely on articulatory motor regions, rendering them unsuitable for individuals with speech motor impairments, such as those with aphasia or locked-in syndrome. To address this limitation, we construct and release NeuroListen, the first publicly available stereo-electroencephalography (sEEG) dataset specifically designed for auditory reconstruction. It contains over 10 hours of neural–speech paired recordings from 5 clinical participants, covering a wide range of semantic categories. Building on this dataset, we propose HyperSpeech, a multi-band neural decoding framework that employs dynamic spatio-temporal hypergraph neural networks to capture high-order dependencies across frequency, spatial, and temporal dimensions. Experimental results demonstrate that HyperSpeech significantly outperforms existing methods across multiple objective speech quality metrics, and achieves superior performance in human subjective evaluations, validating its effectiveness and advancement. This study provides a dedicated dataset and modeling framework for auditory speech decoding, offering foundations for neural language processing and assistive communication systems.
Xueyi Zhang 0001, Ruicong Wang, Jialu Sun, Siqi Cai 0002, Haizhou Li 0001
NeurIPS4
2025 Boosting Discriminability for Robust Multimodal Entity Linking with Visual Modality Missing
abstract
Multimodal Entity Linking (MEL) aims to retrieve ambiguous mentions within multimodal contexts to the referent entities in a multimodal knowledge base, typically based on the assumption of modality completeness. However, when deployed in open-world applications, MEL systems may encounter uncertainly missing of visual modalities from user-proposed mentions. In this paper, we propose a novel setting dubbed MEL-MM to simulate the practical challenge, and reveal that the semantic discriminability is a crucial factor to enhance the anti-missingness resilience. To this end, we introduce an innovative yet efficient approach termed Cross-View Introspective Ranking Distillation (CVIRD), which seeks to sufficiently align the linking similarities between teacher and student models trained from modality-complete and incomplete data. To be specific, as the first concept in CVIRD, Missing-Aware Ranking Distillation (MARD) focuses on modeling the discriminability by formulating the similarity rankings between mention and entities in a missing-sensitive and differentiable manner. Moreover, the second concept of Cross-View Distillation with Introspection (CVDI) aims to improve discriminability extraction in MARD through multi-level distillation, considering both cross-view retrieval and self-consistency. Experiments verify the effectiveness and model-agnostic ability of our method, which achieves superior performance in contrast to competitive missingness-resilient strategies.
Mingrui Lao, Yanming Guo, Xueyi Zhang 0001, Siqi Cai 0002, Zhaoyun Ding, Haizhou Li 0001
SIGIR5
2025 Toward Building Human-Like Sequential Memory Using Brain-Inspired Spiking Neural Models
abstract
The brain is able to acquire and store memories of everyday experiences in real-time. It can also selectively forget information to facilitate memory updating. However, our understanding of the underlying mechanisms and coordination of these processes within the brain remains limited. However, no existing artificial intelligence models have yet matched human-level capabilities in terms of memory storage and retrieval. This study introduces a brain-inspired spiking neural model that integrates the learning and forgetting processes of sequential memory. The proposed model closely mimics the distributed and sparse temporal coding observed in the biological neural system. It employs one-shot online learning for memory formation and uses biologically plausible mechanisms of neural oscillation and phase precession to retrieve memorized sequences reliably. In addition, an active forgetting mechanism is integrated into the spiking neural model, enabling memory removal, flexibility, and updating. The proposed memory model not only enhances our understanding of human memory processes but also provides a robust framework for addressing temporal modeling tasks.
Malu Zhang, Xiaoling Luo 0001, Jibin Wu, Ammar Belatreche, Siqi Cai 0002, Yang Yang 0002, Haizhou Li 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 Robust Decoding of the Auditory Attention from EEG Recordings Through Graph Convolutional Networks
abstract
Auditory attention decoding (AAD) with electroencephalography (EEG) holds great promise in brain-computer interface (BCI). Despite much progress, it remains a research topic on how to effectively evaluate the performance of EEG-based AAD algorithms under an appropriate setting that reflects the use scenarios. It is desired that systems are evaluated under cross-subject and cross-trial settings. However, systems are often reported under same-subject, same-trial settings, where test data are not truly separated from the training data, thus potentially leading to model overfitting due to data leakage. In this paper, we study the robustness of graph convolutional network (GCN), a novel approach to learning the intricate spatial patterns in multi-channel EEG signals, by comparing GCN across cross-subject, cross-trial, and same-subject, same-trial settings. On two publicly available AAD datasets, it is found that GCN exhibits remarkable robustness, outperforming previous conventional convolutional neural network (CNN) solutions. We confirm the superiority of our GCN-based AAD model in terms of generalization and robustness.
Siqi Cai 0002, Haizhou Li 0001
ICASSP1
2024 ASA: An Auditory Spatial Attention Dataset with Multiple Speaking Locations
Zijie Lin, Tianyu He, Siqi Cai 0002, Haizhou Li 0001
INTERSPEECH3
2024 Leveraging Graphic and Convolutional Neural Networks for Auditory Attention Detection with EEG
Saurav Pahuja, Gabriel Ivucic, Pascal Himmelmann, Siqi Cai 0002, Tanja Schultz, Haizhou Li 0001
INTERSPEECH4
2024 Language Without Borders: A Dataset and Benchmark for Code-Switching Lip Reading
abstract
Lip reading aims at transforming the videos of continuous lip movement into textual contents, and has achieved significant progress over the past decade. It serves as a critical yet practical assistance for speech-impaired individuals, with more practicability than speech recognition in noisy environments. With the increasing interpersonal communications in social media owing to globalization, the existing monolingual datasets for lip reading may not be sufficient to meet the exponential proliferation of bilingual and even multilingual users. However, to our best knowledge, research on code-switching is only explored in speech recognition, while the attempts in lip reading are seriously neglected. To bridge this gap, we have collected a bilingual code-switching lip reading benchmark composed of Chinese and English, dubbed CSLR. As the pioneering work, we recruited 62 speakers with proficient foundations in bothspoken Chinese and English to express sentences containing both involved languages. Through rigorous criteria in data selection, CSLR benchmark has accumulated 85,560 video samples with a resolution of 1080x1920, totaling over 71.3 hours of high-quality code-switching lip movement data. To systematically evaluate the technical challenges in CSLR, we implement commonly-used lip reading backbones, as well as competitive solutions in code-switching speech for benchmark testing. Experiments show CSLR to be a challenging and under-explored lip reading task. We hope our proposed benchmark will extend the applicability of code-switching lip reading, and further contribute to the communities of cross-lingual communication and collaboration. Our dataset and benchmark are accessible at https://github.com/cslr-lipreading/CSLR.
Xueyi Zhang 0001, Mingrui Lao, Jun Tang 0001, Yanming Guo, Siqi Cai 0002, Xianghu Yue, Haizhou Li 0001
NeurIPS6
2024 Neurospex: Neuro-Guided Speaker Extraction With Cross-Modal Fusion
abstract
In the study of auditory attention, it has been revealed that there exists a robust correlation between attended speech and elicited neural responses, measurable through electroencephalography (EEG). Therefore, it is possible to use the attention information available within EEG signals to guide the extraction of the target speaker in a cocktail party computationally. In this paper, we present a neuro-guided speaker extraction model, i.e. NeuroSpex, using the EEG response of the listener as the sole auxiliary reference cue to extract attended speech from monaural speech mixtures. We propose a novel EEG signal encoder that captures the attention information. Additionally, we propose a cross-attention (CA) mechanism to enhance the speech feature representations, generating a speaker extraction mask. Experimental results on a publicly available dataset demonstrate that our proposed model outperforms two baseline models across various evaluation metrics.
Dashanka De Silva, Siqi Cai 0002, Saurav Pahuja, Tanja Schultz, Haizhou Li 0001
SLT2
2024 NeuroHeed: Neuro-Steered Speaker Extraction Using EEG Signals
abstract
Humans possess the remarkable ability to selectively attend to a single speaker amidst competing voices and background noise, known asselective auditory attention. Recent studies in auditory neuroscience indicate a strong correlation between the attended speech signal and the corresponding brain's elicited neuronal activities. In this work, we study such brain activities measured using affordable and non-intrusive electroencephalography (EEG) devices. We present NeuroHeed, a speaker extraction model that leverages the listener's synchronized EEG signals to extract the attended speech signal in a cocktail party scenario, in which the extraction process is conditioned on a neuronal attractor encoded from the EEG signal. We propose both an offline and an online NeuroHeed, with the latter designed for real-time inference. In the online NeuroHeed, we additionally propose an autoregressive speaker encoder, which accumulates past extracted speech signals for self-enrollment of the attended speaker information into an auditory attractor, that retains the attentional momentum over time. Online NeuroHeed extracts the current window of the speech signals with guidance from both attractors. Experimental results on KUL dataset two-speaker scenario demonstrate that NeuroHeed effectively extracts brain-attended speech signals with an average scale-invariant signal-to-noise ratio improvement (SI-SDRi) of 14.3 dB and extraction accuracy of 90.8% in offline settings, and SI-SDRi of 11.2 dB and extraction accuracy of 85.1% in online settings.
Zexu Pan, Marvin Borsdorf, Siqi Cai 0002, Tanja Schultz, Haizhou Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2024 A Bio-Inspired Spiking Attentional Neural Network for Attentional Selection in the Listening Brain
abstract
Humans show a remarkable ability in solving the cocktail party problem. Decoding auditory attention from the brain signals is a major step toward the development of bionic ears emulating human capabilities. Electroencephalography (EEG)-based auditory attention detection (AAD) has attracted considerable interest recently. Despite much progress, the performance of traditional AAD decoders remains to be improved, especially in low-latency settings. State-of-the-art AAD decoders based on deep neural networks generally lack the intrinsic temporal coding ability in biological networks. In this study, we first propose a bio-inspired spiking attentional neural network, denoted as BSAnet, for decoding auditory attention. BSAnet is capable of exploiting the temporal dynamics of EEG signals using biologically plausible neurons and an attentional mechanism. Experiments on two publicly available datasets confirm the superior performance of BSAnet over other state-of-the-art systems across various evaluation conditions. Moreover, BSAnet imitates realistic brain-like information processing, through which we show the advantage of brain-inspired computational models.
Siqi Cai 0002, Peiwen Li, Haizhou Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Automatic Detection and Reduction of Compensation in Stroke Patients During Robotic Rehabilitation
abstract
Rehabilitation robots offer promising clinical benefits, however, a critical concern is whether patients can perform exercises correctly without therapist guidance. Patients with stroke, in particular, often resort to compensatory movements without supervision, leading to less-than-optimal outcomes. To address this issue, our work aims to detect and minimize the use of compensation during robotic rehabilitation. Specifically, we developed an artificial neural network (ANN) that uses pressure distribution data to detect compensatory motions. Furthermore, a rehabilitation robot provided real-time force feedback to stroke patients to reduce trunk compensation when necessary. Our experiment involved 8 stroke patients and 12 healthy individuals who completed reaching tasks using an upper limb rehabilitation robot. Results demonstrated the effectiveness of our approach in both offline and online compensation detection. Moreover, we observed a significant reduction in trunk compensation among stroke patients. This study provides a promising solution for unsupervised robotic rehabilitation for stroke patients, paving the way for improved rehabilitation outcomes.
Siqi Cai 0002, Longhan Xie
CoDIT1
2023 Automatic Diagnosis of Pectus Excavatum from CT Images Using a Joint CNN-LSTM Model
abstract
Pectus excavatum (PE) is one of the most common congenital sternal deformities. Accurate preoperative diagnosis of PE is of great significance for subsequent correction and improvement of the patient's quality of life. However, current diagnostic methods rely on the calculation of some PE indices, which is a heavy workload for physical therapists and suffers from measurement errors. To address this issue, we propose an end-to-end automatic assessment of PE, which features a cascaded structure of CNN and LSTM. Specifically, the high-level feature representations of CT images are extracted by the pretrained CNN and then processed through the LSTM layers for classification. In addition, we build up a medical image dataset for PE diagnosis by collecting chest CT images of 42 subjects. Results on this dataset show that the proposed CNN-LSTM framework achieves a relatively high accuracy of 90.20%, which provides a new perspective for the automatic diagnosis of PE in clinics.
Yizhi Liao, Haiyu Zhou, Longhan Xie, Siqi Cai 0002
CoDIT4
2023 Multi-Head Attention and GRU for Improved Match-Mismatch Classification of Speech Stimulus and EEG Response
abstract
This work is based on the participation by the HyperAttention team in the Auditory EEG Decoding Challenge, 2023 (ICASSP 2023 Signal Processing Grand Challenge) task 1, which deals with the match-mismatch classification of speech stimuli and EEG responses of human listeners. We demonstrate the benefits of using mel-spectrograms instead of speech envelopes as input features as well as the effectiveness of Multi-Head Attention and GRU for EEG and speech processing. With a total score of 79.05 %, we reach the second place in the challenge.
Marvin Borsdorf, Saurav Pahuja, Gabriel Ivucic, Siqi Cai 0002, Haizhou Li 0001, Tanja Schultz
ICASSP4
2023 EEG-based Auditory Attention Detection with Spatiotemporal Graph and Graph Convolutional Network
Ruicong Wang, Siqi Cai 0002, Haizhou Li 0001
INTERSPEECH2
2023 Enhancing Subject-Independent EEG-Based Auditory Attention Decoding with WGAN and Pearson Correlation Coefficient
abstract
Electroencephalography (EEG) related research faces a significant challenge of subject independence due to the variation in brain signals and responses among individuals. While deep learning models hold promise in addressing this challenge, their effectiveness depends on large datasets for training and generalization across participants. To overcome this limitation, we propose a solution to the above limitation by increasing the size and quality of training data for subject-independent auditory attention decoding (AAD) using EEG with deep learning. Specifically, our method employs a Wasserstein Generative Adversarial Network (WGAN) to generate synthetic data, with Pearson correlation filtering the most realistic samples. We evaluated this method on a publicly available dataset of selective auditory attention experiments and showed superior performance in subject-independent AAD performance. The mixed training set, consisting of both real and artificial data generated by the WGAN+Pearson Correlation Coefficient, demonstrated approximately 4% improvement in AAD accuracy for a 1-second window. These results demonstrate that deep learning remains a viable approach to overcoming data scarcity in subject-independent AAD tasks based on EEG. Moreover, the proposed method has the potential to improve the generalization and reliability of EEG classification tasks.
Saurav Pahuja, Gabriel Ivucic, Felix Putze, Siqi Cai 0002, Haizhou Li 0001, Tanja Schultz
SMC4
2023 Automatic contour correction of pectus excavatum using computer-aided diagnosis and convolutional neural network
Siqi Cai 0002, Yizhi Liao, Lixuan Lai, Haiyu Zhou, Longhan Xie
Eng. Appl. Artif. Intell.1
2022 A neuroscience-inspired spiking neural network for EEG-based auditory spatial attention detection
Faramarz Faghihi, Siqi Cai 0002, Ahmed A. Moustafa
Neural Networks2
2022 A Biologically Inspired Attention Network for EEG-Based Auditory Attention Detection
abstract
Decoding auditory attention in a cocktail party from neural activities is crucial in the brain-computer interfaces (BCIs). Given that the speech-electroencephalography (EEG) relationships are informative about attentional focus, we propose a novel framework called the biologically inspired attention network (BIAnet) to capture the interactions between EEG and speech. With the neural attention mechanism, the BIAnet can model how each EEG frequency band is related to the subband envelopes of speech by dynamically assigning weights to individual frequency bands at run-time. Results show that the proposed BIAnet outperforms state-of-the-art AAD methods on two publicly available datasets. We also analyze how the BIAnet works and the frequency-specific interactions between EEG and speech signals through data visualization. Overall, the proposed BIAnet provides an accurate, low-latency, and interpretable AAD approach, which has the potential to be extended to general problems in BCIs.
Peiwen Li, Siqi Cai 0002, Enze Su, Longhan Xie
IEEE Signal Process. Lett.2
2022 A Neural-Inspired Architecture for EEG-Based Auditory Attention Detection
abstract
Humans have the ability to focus on one of the sound sources in a noisy scene, which is critical for everyday communication. Auditory attention detection (AAD) seeks to detect selective attention from one’s brain signals. For AAD to be useful in brain–computer interface applications, new approaches with low computational cost, high classification performance, and low latency are required to be developed. In this study, we proposed a novel neural-inspired architecture to mimic the neural computation and coding strategy in the brain for electroencephalography-based AAD. We validated our model through data visualization, and conducted experiments on two publicly available databases. For both KUL and DTU databases, it outperforms both linear and convolutional neural network (CNN) models with consistent improvements from 1 s to 5 s decision windows in terms of detection accuracy. Although the accuracy of the proposed neural-inspired model is inferior to the state-of-the-art spatio-spectral feature (SSF)-CNN model, the computational cost of our model is less than 1% of SSF-CNN’s. Moreover, the neural-inspired decoder is more hardware friendly and energy-efficient due to its biological computing scheme. Overall, the proposed neural-inspired architecture realizes a fast, accurate, and low energy expenditure AAD, which is a big step forward towards practical neuro-steered hearing aids.
Siqi Cai 0002, Peiwen Li, Enze Su, Qi Liu 0005, Longhan Xie
IEEE Trans. Hum. Mach. Syst.1
2022 EEG-Based Auditory Attention Detection via Frequency and Channel Neural Attention
abstract
Humans have the ability to pay attention to one of the sound sources in a multispeaker acoustic environment. Auditory attention detection (AAD) seeks to detect the attended speaker from one’s brain signals that will enable many innovative human–machine systems. However, effective representation learning of electroencephalography (EEG) signals remains a challenge. In this article, we propose a neural attention mechanism that dynamically assigns differentiated weights to the subbands and the channels of EEG signals to derive discriminative representations for AAD. In the nutshell, we would like to build a computational attention mechanism, i.e., neural attention, to model the auditory attention in human brain. We incorporate the proposed neural attention into an AAD system, and validate the neural attention mechanism through comprehensive experiments on two publicly available datasets. The experimental results demonstrate that the proposed system significantly outperforms the state-of-the-art reference baselines.
Siqi Cai 0002, Enze Su, Longhan Xie, Haizhou Li 0001
IEEE Trans. Hum. Mach. Syst.1
2021 Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue
Haizhou Li 0001, Gina-Anne Levow, Chitralekha Gupta, Berrak Sisman, Siqi Cai 0002, David Vandyke, Nina Dethlefs, Yan Wu 0002, Junyi Jessy Li
SIGDIAL6
2020 Low Latency Auditory Attention Detection with Common Spatial Pattern Analysis of EEG Signals
Siqi Cai 0002, Enze Su, Yonghao Song, Longhan Xie, Haizhou Li 0001
INTERSPEECH1
2020 Real-Time Detection of Compensatory Patterns in Patients With Stroke to Reduce Compensation During Robotic Rehabilitation Therapy
abstract
OBJECTIVES: Compensations are commonly employed by patients with stroke during rehabilitation without therapist supervision, leading to suboptimal recovery outcomes. This study investigated the feasibility of the real-time monitoring of compensation in patients with stroke by using pressure distribution data and machine learning algorithms. Whether trunk compensation can be reduced by combining the online detection of compensation and haptic feedback of a rehabilitation robot was also investigated. METHODS: Six patients with stroke did three forms of reaching movements while pressure distribution data were recorded as Dataset1. A support vector machine (SVM) classifier was trained with features extracted from Dataset1. Then, two other patients with stroke performed reaching tasks, and the SVM classifier trained by Dataset1 was employed to classify the compensatory patterns online. Based on the real-time monitoring of compensation, a rehabilitation robot provided an assistive force to patients with stroke to reduce compensations. RESULTS: Good classification performance (F1 score > 0.95) was obtained in both offline and online compensation analysis using the SVM classifier and pressure distribution data of patients with stroke. Based on the real-time detection of compensatory patterns, the angles of trunk rotation, trunk lean-forward and trunk-scapula elevation decreased by 46.95%, 32.35% and 23.75%, respectively. CONCLUSION: High classification accuracies verified the feasibility of detecting compensation in patients with stroke based on pressure distribution data. Since the validity and reliability of the online detection of compensation has been verified, this classifier can be incorporated into a rehabilitation robot to reduce trunk compensations in patients with stroke.
Siqi Cai 0002, Guofeng Li, Enze Su, Xuyang Wei, Shuangyuan Huang, Ke Ma 0007, Haiqing Zheng, Longhan Xie
IEEE J. Biomed. Health Informatics1