Cuntai Guan

dblp:95/7006 · DBLP profile ↗
← Back
155ranked-venue papers
2as first author
63since 2021 · last 2026
0000-0002-0872-3276ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 95 · 1 first-author · 40 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 2 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 29 · 15 since 2021Human-computer interaction and ubiquitous computing · 12 · 1 since 2021Systems, architecture and hardware · 5 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 EEG-DLite: Dataset Distillation for Efficient Large EEG Model Training
abstract
Large-scale EEG foundation models have shown strong generalization across a range of downstream tasks, but their training remains resource-intensive due to the volume and variable quality of EEG data. In this work, we introduce EEG-DLite, a data distillation framework that enables more efficient pre-training by selectively removing noisy and redundant samples from large EEG datasets. EEG-DLite begins by encoding EEG segments into compact latent representations using a self-supervised autoencoder, allowing sample selection to be performed efficiently and with reduced sensitivity to noise. Based on these representations, EEG-DLite filters out outliers and minimizes redundancy, resulting in a smaller yet informative subset that retains the diversity essential for effective foundation model training. Through extensive experiments, we demonstrate that training on only 5 percent of a 2,500-hour dataset curated with EEG-DLite yields performance comparable to, and in some cases better than, training on the full dataset across multiple downstream tasks. To our knowledge, this is the first systematic study of pre-training data distillation in the context of EEG foundation models. EEG-DLite provides a scalable and practical path toward more effective and efficient physiological foundation modeling.
Yuting Tang, Wei-Bang Jiang, Shanglin Li, Yong Li 0032, Xinliang Zhou, Yi Ding 0012, Cuntai Guan
AAAI8
2026 BL-UDA: Towards Unsupervised Domain-Adaptive Surgical Instrument Segmentation with Source Box Labels
abstract
Recent advances in unsupervised domain adaptation (UDA) by adapting the model from one domain to another unseen domain have shown considerable promise in improving surgical instrument segmentation performance across domains. However, existing UDA methods primarily rely on pixel-wise labels, which are always difficult to collect due to the labor-intensive annotation process. In this work, we aim to relax the dependence on pixel-level supervision and investigate a challenging UDA setting - source box annotations, where weak supervision and domain shifts coexist. To achieve this, we introduce a novel unsupervised domain adaptation framework, BL-UDA, which leverages bounding box annotations for surgical instrument segmentation across domains. By utilizing the Segment Anything Model (SAM) for pseudo label generation from box annotations, our method effectively bridges object-level and pixel-level domain adaptation. The proposed BL-UDA framework comprises a teacher-student network with entropy minimization for object detection and an entropy-based label selection strategy for generating box prompts to SAM, facilitating pixel-level domain adaptation. Extensive experiments on the EndoVis 2017 and 2018 datasets demonstrate the superiority of BL-UDA over existing UDA methods, significantly mitigating domain shifts and addressing weak supervision challenges with minimal annotation requirements.
Ziyuan Zhao, Yifang Yin, Yichen Zhang 0002, Xulei Yang, Jun Cheng 0003, Roger Zimmermann, Cuntai Guan, Shaohua Kevin Zhou
ICMR8
2026 Adaptive bottleneck transformer for multimodal EEG, audio, and vision fusion
Sabina Bralina, Adnan Yazici, Cuntai Guan, Min-Ho Lee
Expert Syst. Appl.3
2026 EEG-to-gait decoding via phase-aware representation learning
abstract
Accurate decoding of lower-limb motion from EEG signals is essential for advancing brain-computer interface (BCI) applications in movement intent recognition and control. This study presents NeuroDyGait, a two-stage, phase-aware EEG-to-gait decoding framework that explicitly models temporal continuity and domain relationships. To address challenges of causal, phase-consistent prediction and cross-subject variability, Stage I learns semantically aligned EEG-motion embeddings via relative contrastive learning with a cross-attention-based metric, while Stage II performs domain relation-aware decoding through dynamic fusion of session-specific heads. Comprehensive experiments on two benchmark datasets (GED and FMD) show substantial gains over baselines, including a recent 2025 model EEG2GAIT. The framework generalizes to unseen subjects and maintains inference latency below 5 ms per window, satisfying real-time BCI requirements. Visualization of learned attention and phase-specific cortical saliency maps further reveals interpretable neural correlates of gait phases. Future extensions will target rehabilitation populations and multimodal integration.
Xi Fu, Wei-Bang Jiang, Rui Liu 0034, Gernot R. Müller-Putz, Cuntai Guan
Neural Networks5
2026 Decoupled Hierarchical Distillation for Multimodal Emotion Recognition
abstract
Human multimodal emotion recognition (MER) seeks to infer human emotions by integrating information from language, visual, and acoustic modalities. Although existing MER approaches have achieved promising results, they still struggle with inherent multimodal heterogeneities and varying contributions from different modalities. To address these challenges, we propose a novel framework, Decoupled Hierarchical Multimodal Distillation (DHMD). DHMD decouples each modality's features into modality-irrelevant (homogeneous) and modality-exclusive (heterogeneous) components using a self-regression mechanism. The framework employs a two-stage knowledge distillation (KD) strategy: (1) coarse-grained KD via a Graph Distillation Unit (GD-Unit) in each decoupled feature space, where a dynamic graph facilitates adaptive distillation among modalities, and (2) fine-grained KD through a cross-modal dictionary matching mechanism, which aligns semantic granularities across modalities to produce more discriminative MER representations. This hierarchical distillation approach enables flexible knowledge transfer and effectively improves cross-modal feature alignment. Experimental results demonstrate that DHMD consistently outperforms state-of-the-art MER methods, achieving 1.3%/2.4% (ACC$_{7}$7), 1.3%/1.9% (ACC$_{2}$2) and 1.9%/1.8% (F1) relative improvement on CMU-MOSI/CMU-MOSEI dataset, respectively. Meanwhile, visualization results reveal that both the graph edges and dictionary activations in DHMD exhibit meaningful distribution patterns across modality-irrelevant/-exclusive feature spaces.
Yong Li 0032, Yuanzhi Wang, Yi Ding 0012, Shiqing Zhang, Ke Lu 0002, Cuntai Guan
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 Hierarchical Vision-Language Interaction for Facial Action Unit Detection
abstract
Facial Action Unit (AU) detection seeks to recognize subtle facial muscle activations as defined by the Facial Action Coding System (FACS). A primary challenge w.r.t AU detection is the effective learning of discriminative and generalizable AU representations under conditions of limited annotated data. To address this, we propose a Hierarchical Vision-language Inter action for AU Understanding (HiVA) method, which leverages textual AU descriptions as semantic priors to guide and enhance AU detection. Specifically, HiVA employs a large language model to generate diverse and contextually rich AU descriptions to strengthen language-based representation learning. To capture both fine-grained and holistic vision-language associations, HiVA introduces an AU-aware dynamic graph module that facilitates the learning of AU-specific visual representations. These features are further integrated within a hierarchical cross-modal atten tion architecture comprising two complementary mechanisms: Disentangled Dual Cross-Attention (DDCA), which establishes fine-grained, AU-specific interactions between visual and textual features, and Contextual Dual Cross-Attention (CDCA), which models global inter-AU dependencies. This collaborative, cross modal learning paradigm enables HiVA to leverage multi-grained vision-based AU features in conjunction with refined language based AU details, culminating in robust and semantically en riched AU detection capabilities. Extensive experiments show that HiVA consistently surpasses state-of-the-art approaches. Besides, qualitative analyses reveal that HiVA produces semantically meaningful activation patterns, highlighting its efficacy in learning robust and interpretable cross-modal correspondences for comprehensive facial behavior analysis.
Yong Li 0032, Yizhe Zhang 0001, Tianyi Zhang 0013, Muyun Jiang, Guosen Xie, Cuntai Guan
IEEE Trans. Affect. Comput.8
2026 Decoding Covert Speech From EEG by Functional Areas Spatio-Temporal Transformer
abstract
Covert speech involves imagining speaking without audible sound or any movements.Decoding covert speech from electroencephalogram (EEG) is challenging due to a limited understanding of neural pronunciation mapping and the low signal-to-noise ratio of the signal. In this study, we developed a large-scale multi-utterance speech EEG dataset from 57 right-handed native English-speaking subjects, each performing covert and overt speech tasks by repeating the same word in five utterances within a ten-second duration. Given the spatio-temporal nature of the neural activation process during speech pronunciation, we developed a Functional Areas Spatio-temporal Transformer (FAST), an effective framework for converting EEG signals into tokens and utilizing transformer architecture for sequence encoding. Our results reveal distinct and interpretable speech neural features by the visualization of FAST-generated activation maps across frontal and temporal brain regions, with each word being covertly spoken, providing new insights into the discriminative features of the neural representation of covert speech. This is the first report of such a study, which provides interpretable evidence for speech decoding from EEG.
Muyun Jiang, Wei Zhang 0266, Yi Ding 0012, Kok Ann Colin Teo, Laiguan Fong, Shuailei Zhang, Raghavan Bhuvanakantham, Wei Khang Jeremy Sim, Chuan Huat Vince Foo, Rong Hui Jonathan Chua, Parasuraman Padmanabhan, Victoria Leong, Balázs Gulyás, Cuntai Guan
IEEE J. Biomed. Health Informatics17
2026 Block-Champagne: A Novel Bayesian Framework for Imaging Extended E/MEG Source
abstract
Estimating the extents of E/MEG source activities is crucial for exploring brain dynamics at high spatiotemporal resolution. In this study, we introduce a novel ESI method - Block-Champagne, a Bayesian framework designed to accurately estimate both the locations and extents of extended sources. Our approach leverages a block-sparsity constraint that models each voxel and its neighbors as a single block to account for local homogeneity. The blocks, inherently overlapping in the original source domain, can be adaptively combined to reconstruct sources with arbitrary spatial extents. Furthermore, prior constraints from other neuroimaging modalities with additional spatial information, such as fMRI, can be incorporated to model interactions between distinct sources to further enhance source reconstruction accuracy. The performance of Block-Champagne is quantitatively evaluated through a series of simulation experiments, which demonstrate its overall superiority under various complex scenarios (i.e., SNR, extent size, number of sources, intra-source correlation, & number of EEG channels) compared to benchmark algorithms (including LORETA, EBI-Convex, ts-Cham, L21-Sissy, & BESTIES). Validation results using deep brain stimulation EEG and epilepsy data confirm the practical feasibility of Block-Champagne. Moreover, findings from face processing multimodal data indicate that incorporating relevant and accurate priors significantly enhances source reconstruction accuracy. In conclusion, our study reveals the superiority of the proposed Block-Champagne in accurate reconstruction of extended source, positioning Block-Champagne as a highly promising tool for realistic applications where source locations and extents are of equivalent importance.
Cuntai Guan, Yu Sun 0014
IEEE Trans. Medical Imaging2
2025 Enhancing EEG-based Covert Speech Decoding through Knowledge Transfer
abstract
Covert speech, the imagination of articulation without any actual movement of vocal apparatus, can aid individuals with speech impairments. Recent studies have shown the possibilities of decoding covert speech from non-invasive techniques such as electroencephalogram (EEG). Decoding covert speech from EEG may find broader applications than invasive approaches, but it poses additional challenges due to its less distinct speech-related brain patterns compared with overt speech. In this paper, we propose a novel knowledge transfer framework to build a pretrained model for covert speech from overt speech, including both EEG and audio data. We further integrate acoustic features to enforce a direct neural-to-phonetic mapping from EEG signals to audio representations using data from other subjects. Finally, we fine-tune the pretrained model with covert EEG from the target subjects for decoding. To validate the proposed strategy, we test it on our covert/overt EEG dataset collected from 54 subjects with concurrent EEG and audio recording. Results show that the proposed framework outperforms SOTA methods in terms of classification accuracy of covert speech decoding using EEG.
Muyun Jiang, Min Wu 0008, Balázs Gulyás, Cuntai Guan
ICASSP7
2025 CAT-Net: A Co-Adaptive Transfer Learning Network for BCI-Assisted Neurorehabilitation
abstract
Brain-computer interfaces (BCIs) hold great potential for motor recovery in post-stroke patients. However, the motor imagery decoding accuracy is limited by the non-stationarity of EEG signals across subjects and sessions. We propose CAT-Net: a Co-Adaptive Transfer learning network to simultaneously address the inter-subject variability and inter-session nonstationarity in EEG data. The proposed method selects a relevant subset of data from all the available subjects’ data to train an initial model, followed by subject-specific transfer learning from the initial model to the target subject to establish a pretrain model. Subsequently, online adaptive training is then applied to incrementally train the pretrain model using the data from previous sessions for the target subject. This proposed network using this unique co-adaptive training method is then evaluated on both upper and lower-limb neurorehabilitation EEG datasets comprising 358 sessions from 33 stroke patients. The results showed significant accuracy improvements, achieving averaged accuracies of 70.6% and 72.3% on the respective datasets, surpassing the state-of-the-art baselines.
Shuailei Zhang, Yi Ding 0012, Muyun Jiang, Effie Chew, Kai Keng Ang, Cuntai Guan
ICASSP7
2025 Deep optimal transport for domain adaptation on SPD manifolds
abstract
Recent progress in geometric deep learning has drawn increasing attention from the machine learning community toward domain adaptation on symmetric positive definite (SPD) manifolds—especially for neuroimaging data that often suffer from distribution shifts across sessions. These data, typically represented as covariance matrices of brain signals, inherently lie on SPD manifolds due to their symmetry and positive definiteness. However, conventional domain adaptation methods often overlook this geometric structure when applied directly to covariance matrices, which can result in suboptimal performance. To address this issue, we introduce a new geometric deep learning framework that combines optimal transport theory with the geometry of SPD manifolds. Our approach aligns data distributions while respecting the manifold structure, effectively reducing both marginal and conditional discrepancies. We validate our method on three cross-session brain-computer interface datasets—KU, BNCI2014001, and BNCI2015001—where it consistently outperforms baseline approaches while maintaining the intrinsic geometry of the data. We also provide quantitative results and visualizations to better illustrate the behavior of the learned embeddings.
Ce Ju, Cuntai Guan
Artif. Intell.2
2025 Decoding olfactory response from neurophysiological signal with a multi modal deep learning framework
abstract
The human olfactory system's temporal dynamics are crucial for sensory perception. By learning the temporal dynamics of EEG and utilizing breathing signals, we aim to better understand the neural features of olfactory perception from EEG. To decode the olfactory response effectively, we introduce a new method: the Token Alignment and Cross-Attention Fusion network (TACAF), a multimodal deep learning framework that enhances olfactory EEG decoding using wavelet features for time window selection and spectral analysis for data representation. Spatial features are extracted using spatial learning modules, and temporal dynamics are captured through a multi-head self-attention mechanism. The Temporal Token Semantic Alignment (TTSA) module synchronizes breathing information with EEG data for effective fusion. We collected EEG recordings and breathing signals from 20 subjects to study the decoding responses to pleasant and unpleasant odors. Our evaluation shows that TACAF significantly outperforms existing methods. Further analysis indicates that prolonged odor exposure leads to olfactory adaptation, reducing recognition performance. The findings are visualized through spatial topology maps with saliency mappings, providing insights into the neural mechanisms of olfactory perception.
Chengxuan Tong, Yi Ding 0012, Aung Aung Phyo Wai, Hui Xin Joanna Chua, Xiaorong Wu, Kevin Junliang Lim, Cuntai Guan
Neural Networks7
2025 Self-distillation with beta label smoothing-based cross-subject transfer learning for P300 classification
Shurui Li 0001, Chang Liu 0102, Jing Jin 0001, Cuntai Guan
Pattern Recognit.5
2025 Beyond Overfitting: Doubly Adaptive Dropout for Generalizable AU Detection
abstract
Facial Action Units (AUs) are essential for conveying psychological states and emotional expressions. While automatic AU detection systems leveraging deep learning have progressed, they often overfit to specific datasets and individual features, limiting their cross-domain applicability. To overcome these limitations, we propose a doubly adaptive dropout approach for cross-domain AU detection, which enhances the robustness of convolutional feature maps and spatial tokens against domain shifts. This approach includes a Channel Drop Unit (CD-Unit) and a Token Drop Unit (TD-Unit), which work together to reduce domain-specific noise at both the channel and token levels. The CD-Unit preserves domain-agnostic local patterns in feature maps, while the TD-Unit helps the model identify AU relationships generalizable across domains. An auxiliary domain classifier, integrated at each layer, guides the selective omission of domain-sensitive features. To prevent excessive feature dropout, a progressive training strategy is used, allowing for selective exclusion of sensitive features at any model layer. Our method consistently outperforms existing techniques in cross-domain AU detection, as demonstrated by extensive experimental evaluations. Visualizations of attention maps also highlight clear and meaningful patterns related to both individual and combined AUs, further validating the approach's effectiveness.
Yong Li 0032, Xuesong Niu, Yi Ding 0012, Xiu-Shen Wei, Cuntai Guan
IEEE Trans. Affect. Comput.6
2025 Decoupled Doubly Contrastive Learning for Cross-Domain Facial Action Unit Detection
abstract
Despite the impressive performance of current vision-based facial action unit (AU) detection approaches, they are heavily susceptible to the variations across different domains and the cross-domain AU detection methods are under-explored. In response to this challenge, we propose a decoupled doubly contrastive adaptation (D2CA) approach to learn a purified AU representation that is semantically aligned for the source and target domains. Specifically, we decompose latent representations into AU-relevant and AU-irrelevant components, with the objective of exclusively facilitating adaptation within the AU-relevant subspace. To achieve the feature decoupling, D2CA is trained to disentangle AU and domain factors by assessing the quality of synthesized faces in cross-domain scenarios when either AU or domain attributes are modified. To further strengthen feature decoupling, particularly in scenarios with limited AU data diversity, D2CA employs a doubly contrastive learning mechanism comprising image and feature-level contrastive learning to ensure the quality of synthesized faces and mitigate feature ambiguities. This new framework leads to an automatically learned, dedicated separation of AU-relevant and domain-relevant factors, and it enables intuitive, scale-specific control of the cross-domain facial image synthesis. Extensive experiments demonstrate the efficacy of D2CA in successfully decoupling AU and domain factors, yielding visually pleasing cross-domain synthesized facial images. Meanwhile, D2CA consistently outperforms state-of-the-art cross-domain AU detection approaches, achieving an average F1 score improvement of 6%-14% across various cross-domain scenarios.
Yong Li 0032, Menglin Liu, Zhen Cui 0001, Yi Ding 0012, Yuan Zong, Wenming Zheng, Shiguang Shan, Cuntai Guan
IEEE Trans. Image Process.8
2025 EEG-Deformer: A Dense Convolutional Transformer for Brain-Computer Interfaces
abstract
Effectively learning the temporal dynamics in electroencephalogram (EEG) signals is challenging yet essential for decoding brain activities using brain-computer interfaces (BCIs). Although Transformers are popular for their long-term sequential learning ability in the BCI field, most methods combining Transformers with convolutional neural networks (CNNs) fail to capture the coarse-to-fine temporal dynamics of EEG signals. To overcome this limitation, we introduce EEG-Deformer, which incorporates two main novel components into a CNN-Transformer: (1) a Hierarchical Coarse-to-Fine Transformer (HCT) block that integrates a Fine-grained Temporal Learning (FTL) branch into Transformers, effectively discerning coarse-to-fine temporal patterns; and (2) a Dense Information Purification (DIP) module, which utilizes multi-level, purified temporal information to enhance decoding accuracy. Comprehensive experiments on three representative cognitive tasksâcognitive attention, driving fatigue, and mental workload detectionâconsistently confirm the generalizability of our proposed EEG-Deformer, demonstrating that it either outperforms or performs comparably to existing state-of-the-art methods. Visualization results show that EEG-Deformer learns from neurophysiologically meaningful brain regions for the corresponding cognitive tasks.
Yi Ding 0012, Yong Li 0032, Rui Liu 0034, Chengxuan Tong, Xinliang Zhou, Cuntai Guan
IEEE J. Biomed. Health Informatics8
2025 Explaining E/MEG Source Imaging and Beyond: An Updated Review
abstract
E/MEG source imaging (ESI) provides non-invasive measurements of brain activity with high spatial and temporal resolution. In particular, the wearability and portability of EEG make it an attractive area of research beyond the biomedical communities, especially given the broad application prospects including brain-computer interface (BCI), neuromarketing and neuroergonomics. Although existing reviews offer valuable insights, they often present ESI models in a relatively isolated manner and may not encompass the most recent advancements in the field. In this work, we aim to: 1) provide a timely in-depth review of the widely-explored and state-of-the-art ESI models, including their underlying neurophysiological assumptions and mathematical derivations; 2) list the primary applications of ESI and highlight crucial steps regarding its implementations; 3) discuss current challenges in ESI and propose future research prospects; 4) demonstrate practical usage and implementation details of various representative ESI models. As a rapidly expanding field, ESI is continuously developing and evolving to integrate new technologies. We believe the widespread applications of ESI is happening, and it will dramatically expand our understanding of brain dynamics.
Ioannis Kakkos, George K. Matsopoulos, Cuntai Guan, Yu Sun 0014
IEEE J. Biomed. Health Informatics4
2025 Automated Depression Detection From Text and Audio: A Systematic Review
abstract
Depression is a prevalent mental health disorder that presents significant challenges for timely diagnosis and intervention. Automated Depression Detection (ADD) systems using text and audio offer scalable mental health assessment solutions. This review systematically evaluates 65 studies published between 2018 and 2024, focusing on ADD methods that utilize machine learning models with multimodal data. We examine key methodologies, including data augmentation, multimodal fusion, and feature extraction, along with state-of-the-art ADD systems. The review emphasizes the need for culturally adaptable, high-quality datasets and interpretable models for clinical use. We also identify gaps in longitudinal data and real-world applications. Future research should focus on developing clinically integrated, cross-cultural ADD systems that are interpretable, scalable, and robust. The findings of this review contribute to the research field by providing a comprehensive overview of existing methodologies, identifying gaps in the current literature, and offering insights for future advancements in depression detection using speech and text analysis.
Sinchana Kumbale, Tanmay Surana, Chng Eng Siong, Cuntai Guan
IEEE J. Biomed. Health Informatics6
2025 STARTS: A Self-Adapted Spatio-Temporal Framework for Automatic E/MEG Source Imaging
abstract
To obtain accurate brain source activities, the highly ill-posed source imaging of electro- and magneto-encephalography (E/MEG) requires proficiency in incorporation of biophysiological constraints and signal-processing techniques. Here, we propose a spatio-temporal-constrainted E/MEG source imaging framework-STARTS that can reconstruct the source in a fully automatic way. Specifically, a block-diagonal covariance was adopted to reconstruct the source extents while maintain spatial homogeneity. Temporal basis functions (TBFs) of both sources and noise were estimated and updated in a data-driven fashion to alleviate the influence of noises and further improve source localization accuracy. The performance of the proposed STARTS was quantitatively assessed through a series of simulation experiments, wherein superior results were obtained in comparison with the benchmark ESI algorithms (including LORETA, EBI-Convex, BESTIES & SI-STBF). Additional validations on epileptic and resting-state EEG data further indicate that the STARTS can produce neurophysiologically plausible results. Moreover, a computationally efficient version of STARTS: smooth STARTS was also introduced with an elementary spatial constraint, which exhibited comparable performance and reduced execution cost. In sum, the proposed STARTS, with its advanced spatio-temporal constraints and self-adapted update operation, provides an effective and efficient approach for E/MEG source imaging.
Cuntai Guan, Ruifeng Zheng, Yu Sun 0014
IEEE Trans. Medical Imaging2
2025 EmT: A Novel Transformer for Generalized Cross-Subject EEG Emotion Recognition
abstract
Integrating prior knowledge of neurophysiology into neural network architecture enhances the performance of emotion decoding. While numerous techniques emphasize learning spatial and short-term temporal patterns, there has been a limited emphasis on capturing the vital long-term contextual information associated with emotional cognitive processes. In order to address this discrepancy, we introduce a novel transformer model called emotion transformer (EmT). EmT is designed to excel in both generalized cross-subject electroencephalography (EEG) emotion classification and regression tasks. In EmT, EEG signals are transformed into a temporal graph format, creating a sequence of EEG feature graphs using a temporal graph construction (TGC) module. A novel residual multiview pyramid graph convolutional neural network (RMPG) module is then proposed to learn dynamic graph representations for each EEG feature graph within the series, and the learned representations of each graph are fused into one token. Furthermore, we design a temporal contextual transformer (TCT) module with two types of token mixers to learn the temporal contextual information. Finally, the task-specific output (TSO) module generates the desired outputs. Experiments on four publicly available datasets show that EmT achieves higher results than the baseline methods for both EEG emotion classification and regression tasks. The code is available at https://github.com/yi-ding-cs/EmT.
Yi Ding 0012, Chengxuan Tong, Shuailei Zhang, Muyun Jiang, Yong Li 0032, Kevin Junliang Lim, Cuntai Guan
IEEE Trans. Neural Networks Learn. Syst.7
2025 Federated Graph Neural Networks: Overview, Techniques, and Challenges
abstract
Graph neural networks (GNNs) have attracted extensive research attention in recent years due to their capability to progress with graph data and have been widely used in practical applications. As societies become increasingly concerned with the need for data privacy protection, GNNs face the need to adapt to this new normal. Besides, as clients in federated learning (FL) may have relationships, more powerful tools are required to utilize such implicit information to boost performance. This has led to the rapid development of the emerging research field of federated GNNs (FedGNNs). This promising interdisciplinary field is highly challenging for interested researchers to grasp. The lack of an insightful survey on this topic further exacerbates the entry difficulty. In this article, we bridge this gap by offering a comprehensive survey of this emerging field. We propose a 2-D taxonomy of the FedGNN literature: 1) the main taxonomy provides a clear perspective on the integration of GNNs and FL by analyzing how GNNs enhance FL training as well as how FL assists GNN training and 2) the auxiliary taxonomy provides a view on how FedGNNs deal with heterogeneity across FL clients. Through discussions of key ideas, challenges, and limitations of existing works, we envision future research directions that can help build more robust, explainable, efficient, fair, inductive, and comprehensive FedGNNs.
Rui Liu 0034, Pengwei Xing, Zichao Deng, Anran Li 0001, Cuntai Guan, Han Yu 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 Deep Geodesic Canonical Correlation Analysis for Covariance-Based Neuroimaging Data
abstract
In human neuroimaging, multi-modal imaging techniques are frequently combined to enhance our comprehension of whole-brain dynamics and improve diagnosis in clinical practice. Modalities like electroencephalography and functional magnetic resonance imaging provide distinct views to the brain dynamics due to diametral spatiotemporal sensitivities and underlying neurophysiological coupling mechanisms. These distinct views pose a considerable challenge to learning a shared representation space, especially when dealing with covariance-based data characterized by their geometric structure. To capitalize on the geometric structure, we introduce a measure called geodesic correlation which expands traditional correlation consistency to covariance-based data on the symmetric positive definite (SPD) manifold. This measure is derived from classical canonical correlation analysis and serves to evaluate the consistency of latent representations obtained from paired views. For multi-view, self-supervised learning where one or both latent views are SPD we propose an innovative geometric deep learning framework termed DeepGeoCCA. Its primary objective is to enhance the geodesic correlation of unlabeled, paired data, thereby generating novel representations while retaining the geometric structures. In simulations and experiments with multi-view and multi-modal human neuroimaging data, we find that DeepGeoCCA learns latent representations with high geodesic correlation for unseen data while retaining relevant information for downstream tasks.
Ce Ju, Reinmar J. Kobler, Liyao Tang, Cuntai Guan, Motoaki Kawanabe
ICLR4
2024 Ladder-of-Thought: Using Knowledge as Steps to Elevate Stance Detection
abstract
Stance detection aims to determine the attitude or viewpoint expressed in a document regarding a specific target. Recent advancements in Large Language Models (LLMs), such as Chain-of-Thought (CoT) prompting, have improved the reasoning capabilities of these models by integrating intermediate rationales. However, the efficacy of CoT can be limited by the model’s internal knowledge, resulting in inaccurate rationales that compromise the subsequent stance prediction. This limitation could further lead to hallucinations, where LLMs produce unfaithful responses and erroneous reasoning, affecting the output’s reliability and precision. Moreover, CoT can be challenging to implement on smaller language models with constrained knowledge and reasoning depth, which raises concerns about efficiency. In response to these issues, we propose the Ladder-of-Thought (LoT), a novel framework using knowledge as steps to elevate stance detection. LoT implements a triple-phase Progressive Optimization Framework: 1) External Knowledge Injection, which aims to enrich the model’s intrinsic knowledge base; 2) Intermediate Knowledge Generation, allowing the model to generate more accurate and dependable intermediate knowledge to enhance the downstream prediction; and 3) Downstream Fine-tuning & Prediction, which aims to improve the model’s prediction accuracy. This sequential approach symbolizes ascending a ladder, with each phase representing a progressive step towards achieving optimal reasoning and prediction performance. Our empirical results have demonstrated that LoT achieves state-of-the-art results in zero-shot/few-shot and in-target stance detection, marking a 16% improvement over ChatGPT and a 10% enhancement compared to ChatGPT with CoT on stance detection task.
Kairui Hu, Ming Yan 0007, Wen Haw Chong, Yong Keong Yap, Cuntai Guan, Joey Tianyi Zhou, Ivor W. Tsang
IJCNN5
2024 CASTNet: Cycle-Consistent Attention-based Network for Decoding Open/Close Hand Movement Attempts using EEG
abstract
Electroencephalogram based Brain-Computer Interface (EEG-BCI) employing motor Execution or imagery paradigms generally employs combination of multiple limbs namely right hand, left hand and foot to establish distinct control tasks. The generated neuronal patterns by multiple limb activity are distinct and have been offering promising impact on BCI-based rehabilitation and control applications for decades. However, a huge challenge for present BCI decoding systems lie in decoding finer actions of a single limb such as opening, closing, flexion and extension which is highly difficult for most present-day neural networks, on account of the highly overlapping brain-activations. The inherent high noise-to-signal ratio, inter-subject variability and intra-subject variability associated with the collected EEG signals makes the task challenging for even for deep learning networks which perform efficiently for multiple limb decoding. In this work, a novel network is introduced, named as Cycle-consistent Attention-based Spatio-Temporal Network (CASTNet), for the purposes of classifying open/close attempt based EEG signals of the right hand. The network uses spatio-temporal filters to extract distinguishable features associated with the motor attempts. The network further utilizes transformer layers to capture the long-term dependencies of EEG decoding network. Cycle-consistency is then applied towards the target data against the remaining training set in a subject-dependent setting. The network is validated upon a dataset comprising of 50 subjects performing open close hand movement attempts using right hand. The proposed network shows high efficacy in utilizing data in model adaptation scenarios.
Han Wei Ng, Kavitha P. Thomas, Neethu Robinson, Aung Aung Phyo Wai, Leran Jenny Liang, Nishka Khendry, Aarthy Nagarajan, Cuntai Guan
IJCNN8
2024 Self-Selecting Semi-Supervised Transformer-Attention Convolutional Network for Four Class EEG-Based Motor Imagery Decoding
abstract
Brain-computer interfaces (BCI) serve as an important tool in areas such as neurorehabilitation and constructing prostheses. Electroencephalogram (EEG) motor imagery (MI) signal is a common method used to communicate between the human brain and the computer interface. However, differentiating between multiple motor imagery signals may be challenging due to the presence of high noise-to-signal ratio and small dataset sizes. In this study, we propose a variational autoencoder and transformer-attention based convolutional neural network (SSTACNet) for multi-class EEG-based motor imagery classification. The SSTACNet model leverages upon variational autoencoders’ ability to measure the contrastive distance between two sets of inputs to perform data self-selection. The model further utilizes multi-head self-attention as well as spatial and temporal convolutional filters to achieve superior extraction of signal features. The model additionally utilizes the variational autoencoder’s ability to augment the dataset with feature-informed pseudo-data, achieving stronger classification results. The proposed model outperforms the current state-of-the-art techniques in the BCI Competition IV-2a dataset with an accuracy of 85.52% and 70.56% for the subject-dependent and subject-independent modes, respectively. Codes may be found at: https://github.com/NgHanWei/SSTACNet
Han Wei Ng, Cuntai Guan
IROS2
2024 SHAN: Shape Guided Network for Thyroid Nodule Ultrasound Cross-Domain Segmentation
Wenhuan Lu, Cuntai Guan, Jie Gao 0008, Xi Wei 0002, Xuewei Li 0001
MICCAI (4)3
2024 See, Predict, Plan: Diffusion for Procedure Planning in Robotic Surgical Videos
Ziyuan Zhao, Fen Fang, Xulei Yang, Qianli Xu, Cuntai Guan, Shaohua Kevin Zhou
MICCAI (6)5
2024 Unsupervised Few-Shot Adaptive Re-Learning for EEG-Based Motor Imagery Classification
abstract
To address the effect of both intra- and inter-subject variability in the EEG-based motor-imagery classification, adaptive schemes have been proposed whereby the pre-trained model. However, collection and labelling of additional target data is resource intensive. Furthermore, there may exist significant signal variations across time especially across multiple recording sessions which can result in model performance deterioration following adaptation. This introduces another challenging problem to training classifiers as they are typically unable to automatically determine which data is suitable for fine-tuning purposes, leading to some subjects facing performance drops even after adaptation. To address the data scarcity and variability in adaptive performance, we propose a novel machine-relearning Siamese architecture which utilizes few samples of unlabeled evaluation data to perform data efficient model re-learning. Machine re-learning optimally selects a sub-section of previously known data to update the model parameters. This is implemented via the use of comparative contrastive loss between the Gaussian distributions of target and known data to perform data selection. The highest subject-independent performance achieved an average (N=54) accuracy of 86.63% (±11.79%) without additional data and 88.23% (±10.36%) when utilizing a single unlabeled supplementary data for two-class motor imagery. The previous best accuracy on this dataset is 85.90% (±11.20%) using the best-known method in the literature. Therefore, significantly superior adaptation performance can be achieved while utilizing lesser amount of information from the target subject through reducing EEG feature variability in the training and fine-tuning sets. Codes may be found at: https://github.com/NgHanWei/EEG_Relearning
Han Wei Ng, Cuntai Guan
SMC2
2024 Inter-participant transfer learning with attention based domain adversarial training for P300 detection
Shurui Li 0001, Ian Daly, Cuntai Guan, Andrzej Cichocki, Jing Jin 0001
Neural Networks3
2024 Aggregating intrinsic information to enhance BCI performance through federated learning
abstract
Insufficient data is a long-standing challenge for Brain-Computer Interface (BCI) to build a high-performance deep learning model. Though numerous research groups and institutes collect a multitude of EEG datasets for the same BCI task, sharing EEG data from multiple sites is still challenging due to the heterogeneity of devices. The significance of this challenge cannot be overstated, given the critical role of data diversity in fostering model robustness. However, existing works rarely discuss this issue, predominantly centering their attention on model training within a single dataset, often in the context of inter-subject or inter-session settings. In this work, we propose a hierarchical personalized Federated Learning EEG decoding (FLEEG) framework to surmount this challenge. This innovative framework heralds a new learning paradigm for BCI, enabling datasets with disparate data formats to collaborate in the model training process. Each client is assigned a specific dataset and trains a hierarchical personalized model to manage diverse data formats and facilitate information exchange. Meanwhile, the server coordinates the training procedure to harness knowledge gleaned from all datasets, thus elevating overall performance. The framework has been evaluated in Motor Imagery (MI) classification with nine EEG datasets collected by different devices but implementing the same MI task. Results demonstrate that the proposed framework can boost classification performance up to 8.4% by enabling knowledge sharing between multiple datasets, especially for smaller datasets. Visualization results also indicate that the proposed framework can empower the local models to put a stable focus on task-related areas, yielding better performance. To the best of our knowledge, this is the first end-to-end solution to address this important challenge.
Rui Liu 0034, Yuanyuan Chen 0012, Anran Li 0001, Yi Ding 0012, Han Yu 0001, Cuntai Guan
Neural Networks6
2024 Subject-independent meta-learning framework towards optimal training of EEG-based classifiers
Han Wei Ng, Cuntai Guan
Neural Networks2
2024 Leveraging temporal dependency for cross-subject-MI BCIs by contrastive learning and self-attention
Yi Ding 0012, Jianzhu Bao, Ke Qin, Chengxuan Tong, Jing Jin 0001, Cuntai Guan
Neural Networks7
2024 MASA-TCN: Multi-Anchor Space-Aware Temporal Convolutional Neural Networks for Continuous and Discrete EEG Emotion Recognition
abstract
Emotion recognition from electroencephalogram (EEG) signals is a critical domain in biomedical research with applications ranging from mental disorder regulation to human-computer interaction. In this paper, we address two fundamental aspects of EEG emotion recognition: continuous regression of emotional states and discrete classification of emotions. While classification methods have garnered significant attention, regression methods remain relatively under-explored. To bridge this gap, we introduce MASA-TCN, a novel unified model that leverages the spatial learning capabilities of Temporal Convolutional Networks (TCNs) for EEG emotion regression and classification tasks. The key innovation lies in the introduction of a space-aware temporal layer, which empowers TCN to capture spatial relationships among EEG electrodes, enhancing its ability to discern nuanced emotional states. Additionally, we design a multi-anchor block with attentive fusion, enabling the model to adaptively learn dynamic temporal dependencies within the EEG signals. Experiments on two publicly available datasets show that MASA-TCN achieves higher results than the state-of-the-art methods for both EEG emotion regression and classification tasks.
Yi Ding 0012, Su Zhang 0004, Chuangao Tang, Cuntai Guan
IEEE J. Biomed. Health Informatics4
2024 LGGNet: Learning From Local-Global-Graph Representations for Brain-Computer Interface
abstract
Neuropsychological studies suggest that co-operative activities among different brain functional areas drive high-level cognitive processes. To learn the brain activities within and among different functional areas of the brain, we propose local-global-graph network (LGGNet), a novel neurologically inspired graph neural network (GNN), to learn local-global-graph (LGG) representations of electroencephalography (EEG) for brain-computer interface (BCI). The input layer of LGGNet comprises a series of temporal convolutions with multiscale 1-D convolutional kernels and kernel-level attentive fusion. It captures temporal dynamics of EEG which then serves as input to the proposed local- and global-graph-filtering layers. Using a defined neurophysiologically meaningful set of local and global graphs, LGGNet models the complex relations within and among functional areas of the brain. Under the robust nested cross-validation settings, the proposed method is evaluated on three publicly available datasets for four types of cognitive classification tasks, namely the attention, fatigue, emotion, and preference classification tasks. LGGNet is compared with state-of-the-art (SOTA) methods, such as DeepConvNet, EEGNet, R2G-STNN, TSception, regularized graph neural network (RGNN), attention-based multiscale convolutional neural network-dynamical graph convolutional network (AMCNN-DGCN), hierarchical recurrent neural network (HRNN), and GraphNet. The results show that LGGNet outperforms these methods, and the improvements are statistically significant ( ) in most cases. The results show that bringing neuroscience prior knowledge into neural network design yields an improvement of classification performance. The source code can be found at https://github.com/yi-ding-cs/LGG.
Yi Ding 0012, Neethu Robinson, Chengxuan Tong, Qiuhao Zeng, Cuntai Guan
IEEE Trans. Neural Networks Learn. Syst.5
2024 Graph Neural Networks on SPD Manifolds for Motor Imagery Classification: A Perspective From the Time-Frequency Analysis
abstract
The motor imagery (MI) classification has been a prominent research topic in brain-computer interfaces (BCIs) based on electroencephalography (EEG). Over the past few decades, the performance of MI-EEG classifiers has seen gradual enhancement. In this study, we amplify the geometric deep-learning-based MI-EEG classifiers from the perspective of time-frequency analysis, introducing a new architecture called Graph-CSPNet. We refer to this category of classifiers as Geometric Classifiers, highlighting their foundation in differential geometry stemming from EEG spatial covariance matrices. Graph-CSPNet utilizes novel manifold-valued graph convolutional techniques to capture the EEG features in the time-frequency domain, offering heightened flexibility in signal segmentation for capturing localized fluctuations. To evaluate the effectiveness of Graph-CSPNet, we employ five commonly used publicly available MI-EEG datasets, achieving near-optimal classification accuracies in nine out of 11 scenarios. The Python repository can be found at https://github.com/GeometricBCI/Tensor-CSPNet-and-Graph-CSPNet.
Ce Ju, Cuntai Guan
IEEE Trans. Neural Networks Learn. Syst.2
2023 Efficient Representation Learning for Inner Speech Domain Generalization
Han Wei Ng, Cuntai Guan
CAIP (1)2
2023 SemiGNN-PPI: Self-Ensembling Multi-Graph Neural Network for Efficient and Generalizable Protein-Protein Interaction Prediction
abstract
Protein-protein interactions (PPIs) are crucial in various biological processes and their study has significant implications for drug development and disease diagnosis. Existing deep learning methods suffer from significant performance degradation under complex real-world scenarios due to various factors, e.g., label scarcity and domain shift. In this paper, we propose a self-ensembling multi-graph neural network (SemiGNN-PPI) that can effectively predict PPIs while being both efficient and generalizable. In SemiGNN-PPI, we not only model the protein correlations but explore the label dependencies by constructing and processing multiple graphs from the perspectives of both features and labels in the graph learning process. We further marry GNN with Mean Teacher to effectively leverage unlabeled graph-structured PPI data for self-ensemble graph learning. We also design multiple graph consistency constraints to align the student and teacher graphs in the feature embedding space, enabling the student model to better learn from the teacher model by incorporating more relationships. Extensive experiments on PPI datasets of different scales with different evaluation settings demonstrate that SemiGNN-PPI outperforms state-of-the-art PPI prediction methods, particularly in challenging scenarios such as training with limited annotations and testing on unseen data.
Ziyuan Zhao, Peisheng Qian, Xulei Yang, Zeng Zeng, Cuntai Guan, Tam Wai Leong, Xiaoli Li 0001
IJCAI5
2023 Enhancing the confidence of deep learning classifiers via interpretable saliency maps
abstract
This paper quantifies the quality of heatmap-based eXplainable AI (XAI) methods w.r.t image classification problem. Here, a heatmap is considered desirable if it improves the probability of predicting the correct classes. Different XAI heatmap-based methods are empirically shown to improve classification confidence to different extents depending on the datasets, e.g. Saliency works best on ImageNet and Deconvolution on Chest X-ray Pneumonia dataset. The novelty includes a new gap distribution that shows a stark difference between correct and wrong predictions. Finally, the generative augmentative explanation is introduced, a method to generate heatmaps capable of improving predictive confidence to a high level.
Erico Tjoa, Hong Jing Khok, Tushar Chouhan, Cuntai Guan
Neurocomputing4
2023 A transformer-based deep neural network model for SSVEP classification
Yangsong Zhang 0001, Yudong Pan, Peng Xu 0001, Cuntai Guan
Neural Networks5
2023 Self-Supervised Contrastive Representation Learning for Semi-Supervised Time-Series Classification
abstract
Learning time-series representations when only unlabeled data or few labeled samples are available can be a challenging task. Recently, contrastive self-supervised learning has shown great improvement in extracting useful representations from unlabeled data via contrasting different augmented views of data. In this work, we propose a novel Time-Series representation learning framework via Temporal and Contextual Contrasting (TS-TCC) that learns representations from unlabeled data with contrastive learning. Specifically, we propose time-series-specific weak and strong augmentations and use their views to learn robust temporal relations in the proposed temporal contrasting module, besides learning discriminative representations by our proposed contextual contrasting module. Additionally, we conduct a systematic study of time-series data augmentation selection, which is a key part of contrastive learning. We also extend TS-TCC to the semi-supervised learning settings and propose a Class-Aware TS-TCC (CA-TCC) that benefits from the available few labeled data to further improve representations learned by TS-TCC. Specifically, we leverage the robust pseudo labels produced by TS-TCC to realize a class-aware contrastive loss. Extensive experiments show that the linear evaluation of the features learned by our proposed framework performs comparably with the fully supervised training. Additionally, our framework shows high efficiency in few labeled data and transfer learning scenarios.
Emadeldeen Eldele, Mohamed Ragab 0002, Zhenghua Chen, Min Wu 0008, Chee Keong Kwoh 0001, Xiaoli Li 0001, Cuntai Guan
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 TSception: Capturing Temporal Dynamics and Spatial Asymmetry From EEG for Emotion Recognition
abstract
The high temporal resolution and the asymmetric spatial activations are essential attributes of electroencephalogram (EEG) underlying emotional processes in the brain. To learn the temporal dynamics and spatial asymmetry of EEG towards accurate and generalized emotion recognition, we propose TSception, a multi-scale convolutional neural network that can classify emotions from EEG. TSception consists of dynamic temporal, asymmetric spatial, and high-level fusion layers, which learn discriminative representations in the time and channel dimensions simultaneously. The dynamic temporal layer consists of multi-scale 1D convolutional kernels whose lengths are related to the sampling rate of EEG, which learns the dynamic temporal and frequency representations of EEG. The asymmetric spatial layer takes advantage of the asymmetric EEG patterns for emotion, learning the discriminative global and hemisphere representations. The learned spatial representations will be fused by a high-level fusion layer. Using more generalized cross-validation settings, the proposed method is evaluated on two publicly available datasets DEAP and MAHNOB-HCI. The performance of the proposed network is compared with prior reported methods such as SVM, KNN, FBFgMDM, FBTSC, Unsupervised learning, DeepConvNet, ShallowConvNet, and EEGNet. TSception achieves higher classification accuracies and F1 scores than other methods in most of the experiments. The codes are available at: https://github.com/yi-ding-cs/TSception
Yi Ding 0012, Neethu Robinson, Su Zhang 0004, Qiuhao Zeng, Cuntai Guan
IEEE Trans. Affect. Comput.5
2023 E-Key: An EEG-Based Biometric Authentication and Driving Fatigue Detection System
abstract
Due to the increasing fatal traffic accidents, there are strong desire for more effective and convenient techniques for driving fatigue detection. Here, we propose a unified frameworkE-Keyto simultaneously perform personal identification (PI) and driving fatigue detection using a convolutional attention neural network (CNN-Attention). The performance was assessed using EEG data collected through a wearable dry-sensor system from 31 healthy subjects undergoing a 90-min simulated driving task. In comparison with three widely-used competitive models (including CNN, CNN-LSTM, and Attention), the proposed scheme achieved the best (p < 0.01) performance in both PI (98.5%) and fatigue detection (97.8%). Besides, the spatial-temporal structure of the proposed framework exhibits an optimal balance between classification performance and computational efficiency. Additional validation analyses were conducted to assess the reliability and practicability of the model via re-configuring the kernel size and manipulating the input data, showing that it can achieve a satisfactory performance using a subset of the input data. In sum, these findings would pave the way for further practical implementation of in-vehicle expert system, showing great potential in autonomous driving and car-sharing where currently monitoring of PI and driving fatigue are of particular interest.
Tao Xu 0010, Hongtao Wang 0001, Guanyong Lu, Feng Wan 0003, Mengqi Deng, Peng Qi 0001, Anastasios Bezerianos, Cuntai Guan, Yu Sun 0014
IEEE Trans. Affect. Comput.8
2023 Video Based Cocktail Causal Container for Blood Pressure Classification and Blood Glucose Prediction
abstract
With the development of modern cameras, more physiological signals can be obtained from portable devices like smartphone. Some hemodynamically based non-invasive video processing applications have been applied for blood pressure classification and blood glucose prediction objectives for unobtrusive physiological monitoring at home. However, this approach is still under development with very few publications. In this paper, we propose an end-to-end framework, entitled cocktail causal container, to fuse multiple physiological representations and to reconstruct the correlation between frequency and temporal information during multi-task learning. Cocktail causal container processes hematologic reflex information to classify blood pressure and blood glucose. Since the learning of discriminative features from video physiological representations is quite challenging, we propose a token feature fusion block to fuse the multi-view fine-grained representations to a union discrete frequency space. A causal net is used to analyze the fused higher-order information, so that the framework can be enforced to disentangle the latent factors into the related endogenous association that corresponds to down-stream fusion information to improve the semantic interpretation. Moreover, a pair-wise temporal frequency map is developed to provide valuable insights into extraction of salient photoplethysmograph (PPG) information from fingertip videos obtained by a standard smartphone camera. Extensive comparisons have been implemented for the validation of cocktail causal container using a Clinical dataset and PPG-BP benchmark. The root mean square error of 1.329±0.167 for blood glucose prediction and precision of 0.89±0.03 for blood pressure classification are achieved in Clinical dataset.
Chuanhao Zhang, Emil Jovanov, Hongen Liao, Yuan-Ting Zhang, Benny P. L. Lo, Yuan Zhang 0007, Cuntai Guan
IEEE J. Biomed. Health Informatics7
2023 LE-UDA: Label-Efficient Unsupervised Domain Adaptation for Medical Image Segmentation
abstract
While deep learning methods hitherto have achieved considerable success in medical image segmentation, they are still hampered by two limitations: (i) reliance on large-scale well-labeled datasets, which are difficult to curate due to the expert-driven and time-consuming nature of pixel-level annotations in clinical practices, and (ii) failure to generalize from one domain to another, especially when the target domain is a different modality with severe domain shifts. Recent unsupervised domain adaptation (UDA) techniques leverage abundant labeled source data together with unlabeled target data to reduce the domain gap, but these methods degrade significantly with limited source annotations. In this study, we address this underexplored UDA problem, investigating a challenging but valuable realistic scenario, where the source domain not only exhibits domain shift w.r.t. the target domain but also suffers from label scarcity. In this regard, we propose a novel and generic framework called "Label-Efficient Unsupervised Domain Adaptation" (LE-UDA). In LE-UDA, we construct self-ensembling consistency for knowledge transfer between both domains, as well as a self-ensembling adversarial learning module to achieve better feature alignment for UDA. To assess the effectiveness of our method, we conduct extensive experiments on two different tasks for cross-modality segmentation between MRI and CT images. Experimental results demonstrate that the proposed LE-UDA can efficiently leverage limited source labels to improve cross-domain segmentation performance, outperforming state-of-the-art UDA approaches in the literature.
Ziyuan Zhao, Fangcheng Zhou, Kaixin Xu, Zeng Zeng, Cuntai Guan, Shaohua Kevin Zhou
IEEE Trans. Medical Imaging5
2023 Tensor-CSPNet: A Novel Geometric Deep Learning Framework for Motor Imagery Classification
abstract
Deep learning (DL) has been widely investigated in a vast majority of applications in electroencephalography (EEG)-based brain-computer interfaces (BCIs), especially for motor imagery (MI) classification in the past five years. The mainstream DL methodology for the MI-EEG classification exploits the temporospatial patterns of EEG signals using convolutional neural networks (CNNs), which have been particularly successful in visual images. However, since the statistical characteristics of visual images depart radically from EEG signals, a natural question arises whether an alternative network architecture exists apart from CNNs. To address this question, we propose a novel geometric DL (GDL) framework called Tensor-CSPNet, which characterizes spatial covariance matrices derived from EEG signals on symmetric positive definite (SPD) manifolds and fully captures the temporospatiofrequency patterns using existing deep neural networks on SPD manifolds, integrating with experiences from many successful MI-EEG classifiers to optimize the framework. In the experiments, Tensor-CSPNet attains or slightly outperforms the current state-of-the-art performance on the cross-validation and holdout scenarios in two commonly used MI-EEG datasets. Moreover, the visualization and interpretability analyses also exhibit the validity of Tensor-CSPNet for the MI-EEG classification. To conclude, in this study, we provide a feasible answer to the question by generalizing the DL methodologies on SPD manifolds, which indicates the start of a specific GDL methodology for the MI-EEG classification.
Ce Ju, Cuntai Guan
IEEE Trans. Neural Networks Learn. Syst.2
2022 MMGL: Multi-Scale Multi-View Global-Local Contrastive Learning for Semi-Supervised Cardiac Image Segmentation
abstract
With large-scale well-labeled datasets, deep learning has shown significant success in medical image segmentation. However, it is challenging to acquire abundant annotations in clinical practice due to extensive expertise requirements and costly labeling efforts. Recently, contrastive learning has shown a strong capacity for visual representation learning on unlabeled data, achieving impressive performance rivaling supervised learning in many domains. In this work, we propose a novel multi-scale multi-view global-local contrastive learning (MMGL) framework to thoroughly explore global and local features from different scales and views for robust contrastive learning performance, thereby improving segmentation performance with limited annotations. Extensive experiments on the MM-WHS dataset demonstrate the effectiveness of MMGL framework on semi-supervised cardiac image segmentation, outperforming the state-of-the-art contrastive learning methods by a large margin.
Ziyuan Zhao, Jinxuan Hu, Zeng Zeng, Xulei Yang, Peisheng Qian, Bharadwaj Veeravalli, Cuntai Guan
ICIP7
2022 ACT-NET: Asymmetric Co-Teacher Network for Semi-Supervised Memory-Efficient Medical Image Segmentation
abstract
While deep models have shown promising performance in medical image segmentation, they heavily rely on a large amount of well-annotated data, which is difficult to access, especially in clinical practice. On the other hand, high-accuracy deep models usually come in large model sizes, limiting their employment in real scenarios. In this work, we propose a novel asymmetric co-teacher framework, ACT-Net, to alleviate the burden on both expensive annotations and computational costs for semi-supervised knowledge distillation. We advance teacher-student learning with a co-teacher network to facilitate asymmetric knowledge distillation from large models to small ones by alternating student and teacher roles, obtaining tiny but accurate models for clinical employment. To verify the effectiveness of our ACT-Net, we employ the ACDC dataset for cardiac substructure segmentation in our experiments. Extensive experimental results demonstrate that ACT-Net outperforms other knowledge distillation methods and achieves lossless segmentation performance with 250× fewer parameters.
Ziyuan Zhao, Andong Zhu 0003, Zeng Zeng, Bharadwaj Veeravalli, Cuntai Guan
ICIP5
2022 MATN: Multi-model Attention Network for Gait Prediction from EEG
abstract
Predicting gait trajectories from electroencephalography (EEG) is useful for lower limb rehabilitation. The temporal dynamics of EEG pertaining to individual subjects while they are walking require a personalized treatment inside the deep neural network. Also, the EEG data collected from different sessions or days may exhibit varied distributions, further hindering the model generality. Therefore, we propose the Multi-model ATtention Network (MATN) to address these issues. MATN exploits the self-attention mechanism on the spatio-temporal EEG encoding to learn the subject-specific temporal dynamics. To adapt to varied data distributions, MATN employs knowledge distillation, which includes a teacher and student model. The latter is supervised by the intermediate feature maps from the trained teacher and the labels. Experiments of the regression of bilateral joint angles on the legs (hip, knee, and ankle) are carried out using the MoBI dataset in a subject-specific manner. Our MATN achieves more than 18% improvement on the Pearson's correlation coefficient compared to five other state-of-the-art methods.
Xi Fu, Cuntai Guan
IJCNN3
2022 TESANet: Self-attention network for olfactory EEG classification
abstract
The olfactory system is known to be associated with emotion during odor stimulation. A well-designed computational model that can correctly recognize preference induced by odor stimulation can be vital in the food and perfume industries. Electroencephalogram (EEG) can be used to study the brain's response to odor stimulation due to its good temporal resolution and low acquisition cost. In this study, we proposed a novel self-attention deep learning framework: Temporal Segment Attention Network (TESANet) to classify the brain state of subjects when they are exposed to pleasant and unpleasant odors. Odor stimulation is a continuous process, the temporal dynamics of the EEG signal should reflect the continuous changes of the brain responses to the given odor, thus we design the model to capture the intercorrelation between time segments of the EEG by utilizing the self-attention mechanism. TESANet consists of a filter-bank layer to extract spectral features, a spatial convolution layer to extract spatial features, a temporal segmentation layer to split the data into overlapping time windows, a Long Short-Term Memory (LSTM) layer to encode the temporal segments, a self-attention layer to decode the temporal dynamics by learning the intercorrelation between time segments, and finally a fully connected layer for classification. Experiments on an olfactory EEG dataset demonstrated that the proposed method outperforms other competing deep learning methods for odor pleasantness classification.
Chengxuan Tong, Yi Ding 0012, Kevin Junliang Lim, Zhuo Zhang 0001, Haihong Zhang, Cuntai Guan
IJCNN6
2022 Meta-hallucinator: Towards Few-Shot Cross-Modality Cardiac Image Segmentation
Ziyuan Zhao, Fangcheng Zhou, Zeng Zeng, Cuntai Guan, Shaohua Kevin Zhou
MICCAI (5)4
2022 ME-PLAN: A deep prototypical learning with local attention network for dynamic micro-expression recognition
Sirui Zhao, Huaying Tang, Yangsong Zhang 0001, Hao Wang 0076, Tong Xu 0001, Enhong Chen, Cuntai Guan
Neural Networks8
2022 Visual-to-EEG cross-modal knowledge distillation for continuous emotion recognition
abstract
Visual modality is one of the most dominant modalities for current continuous emotion recognition methods. Compared to which the EEG modality is relatively less sound due to its intrinsic limitation such as subject bias and low spatial resolution. This work attempts to improve the continuous prediction of the EEG modality by using the dark knowledge from the visual modality. The teacher model is built by a cascade convolutional neural network - temporal convolutional network (CNN-TCN) architecture, and the student model is built by TCNs. They are fed by video frames and EEG average band power features, respectively. Two data partitioning schemes are employed, i.e., the trial-level random shuffling (TRS) and the leave-one-subject-out (LOSO). The standalone teacher and student can produce continuous prediction superior to the baseline method, and the employment of the visual-to-EEG cross-modal KD further improves the prediction with statistical significance, i.e., p-value <0.01 for TRS and p-value <0.05 for LOSO partitioning. The saliency maps of the trained student model show that the brain areas associated with the active valence state are not located in precise brain areas. Instead, it results from synchronized activity among various brain areas. And the fast beta and gamma waves, with the frequency of 18−30Hz and 30−45Hz, contribute the most to the human emotion process compared to other bands. The code is available at https://github.com/sucv/Visual_to_EEG_Cross_Modal_KD_for_CER.
Su Zhang 0004, Chuangao Tang, Cuntai Guan
Pattern Recognit.3
2022 Robust Traffic Prediction From Spatial-Temporal Data Based on Conditional Distribution Learning
abstract
Traffic prediction based on massive speed data collected from traffic sensors plays an important role in traffic management. However, it is still challenging to obtain satisfactory performance due to the complex and dynamic spatial-temporal correlations among the data. Recently, many research works have demonstrated the effectiveness of graph neural networks (GNNs) for spatial-temporal modeling. However, such models are restricted by conditional distribution during training, and may not perform well when the target is outside the primary region of interest in the distribution. In this article, we address this problem with a stagewise learning mechanism, in which we redefine speed prediction as a conditional distribution learning followed by speed regression. We first perform a conditional distribution learning for each observed speed class, and then obtain speed prediction by optimizing regression learning, based on the learned conditional distribution. To effectively learn the conditional distribution, we introduce a mean-residue loss, consisting of two parts: 1) a mean loss, which penalizes the differences between the mean of the estimated conditional distribution and the ground truth and 2) a residue loss, which penalizes residue errors of the long tails in the distribution. To optimize the subsequent regression based on distribution information, we combine the mean absolute error (MAE) as another part of the loss function. We also incorporate a GNN-based architecture with our proposed learning mechanism. Mean-residue loss is employed to supervise the hidden speed representation in the network at each time interval, followed by a shared layer to recalibrate the hidden temporal dependencies in the conditional distribution. The experimental results based on three public traffic datasets have demonstrated that the effectiveness of the proposed method outperforms state-of-the-art methods.
Zeng Zeng, Wei Zhao 0035, Peisheng Qian, Yingjie Zhou 0001, Ziyuan Zhao, Cen Chen 0002, Cuntai Guan
IEEE Trans. Cybern.7
2022 Spatio-Spectral Feature Representation for Motor Imagery Classification Using Convolutional Neural Networks
abstract
Convolutional neural networks (CNNs) have recently been applied to electroencephalogram (EEG)-based brain-computer interfaces (BCIs). EEG is a noninvasive neuroimaging technique, which can be used to decode user intentions. Because the feature space of EEG data is highly dimensional and signal patterns are specific to the subject, appropriate methods for feature representation are required to enhance the decoding accuracy of the CNN model. Furthermore, neural changes exhibit high variability between sessions, subjects within a single session, and trials within a single subject, resulting in major issues during the modeling stage. In addition, there are many subject-dependent factors, such as frequency ranges, time intervals, and spatial locations at which the signal occurs, which prevent the derivation of a robust model that can achieve the parameterization of these factors for a wide range of subjects. However, previous studies did not attempt to preserve the multivariate structure and dependencies of the feature space. In this study, we propose a method to generate a spatiospectral feature representation that can preserve the multivariate information of EEG data. Specifically, 3-D feature maps were constructed by combining subject-optimized and subject-independent spectral filters and by stacking the filtered data into tensors. In addition, a layer-wise decomposition model was implemented using our 3-D-CNN framework to secure reliable classification results on a single-trial basis. The average accuracies of the proposed model were 87.15% (±7.31), 75.85% (±12.80), and 70.37% (±17.09) for the BCI competition data sets IV_2a, IV_2b, and OpenBMI data, respectively. These results are better than those obtained by state-of-the-art techniques, and the decomposition model obtained the relevance scores for neurophysiologically plausible electrode channels and frequency domains, confirming the validity of the proposed approach.
Ji-Seon Bang, Min-Ho Lee, Siamac Fazli, Cuntai Guan, Seong-Whan Lee
IEEE Trans. Neural Networks Learn. Syst.4
2021 Time-Series Representation Learning via Temporal and Contextual Contrasting
abstract
Learning decent representations from unlabeled time-series data with temporal dynamics is a very challenging task. In this paper, we propose an unsupervised Time-Series representation learning framework via Temporal and Contextual Contrasting (TS-TCC), to learn time-series representation from unlabeled data. First, the raw time-series data are transformed into two different yet correlated views by using weak and strong augmentations. Second, we propose a novel temporal contrasting module to learn robust temporal representations by designing a tough cross-view prediction task. Last, to further learn discriminative representations, we propose a contextual contrasting module built upon the contexts from the temporal contrasting module. It attempts to maximize the similarity among different contexts of the same sample while minimizing similarity among contexts of different samples. Experiments have been carried out on three real-world time-series datasets. The results manifest that training a linear classifier on top of the features learned by our proposed TS-TCC performs comparably with the supervised training. Additionally, our proposed TS-TCC shows high efficiency in few-labeled data and transfer learning scenarios. The code is publicly available at https://github.com/emadeldeen24/TS-TCC.
Emadeldeen Eldele, Mohamed Ragab 0002, Zhenghua Chen, Min Wu 0008, Chee Keong Kwoh 0001, Xiaoli Li 0001, Cuntai Guan
IJCAI7
2021 MT-UDA: Towards Unsupervised Cross-modality Medical Image Segmentation with Limited Source Labels
Ziyuan Zhao, Kaixin Xu, Shumeng Li, Zeng Zeng, Cuntai Guan
MICCAI (1)5
2021 fMRI-SI-STBF: An fMRI-informed Bayesian electromagnetic spatio-temporal extended source imaging
Ke Liu 0008, Zhu Liang Yu, Wei Wu 0022, Zhenghui Gu, Cuntai Guan
Neurocomputing6
2021 WiFi-Sleep: Sleep Stage Monitoring Using Commodity Wi-Fi Devices
abstract
Sleep monitoring is essential to people's health and wellbeing, which can also assist in the diagnosis and treatment of sleep disorder. Compared with contact-based solutions, contactless sleep monitoring does not attach any device to the human body; hence, it has attracted increasing attention in recent years. Inspired by the recent advances in Wi-Fi-based sensing, this article proposes a low-cost and nonintrusive sleep monitoring system using commodity Wi-Fi devices, namely, WiFi-Sleep. We leverage the fine-grained channel state information from multiple antennas and propose advanced fusion and signal processing methods to extract accurate respiration and body movement information. We introduce a deep learning method combined with clinical sleep medicine prior knowledge to achieve four-stage sleep monitoring with limited data sources (i.e., only respiration and body movement information). We benchmark the performance of WiFi-Sleep with polysomnography, the gold reference standard. Results show that WiFi-Sleep achieves an accuracy of 81.8%, which is comparable to the state-of-the-art sleep stage monitoring using expensive radar devices.
Bohan Yu, Kai Niu 0003, Youwei Zeng, Tao Gu 0001, Leye Wang, Cuntai Guan, Daqing Zhang 0001
IEEE Internet Things J.7
2021 An end-to-end 3D convolutional neural network for decoding attentive mental state
Yangsong Zhang 0001, Huan Cai, Li Nie, Peng Xu 0001, Sirui Zhao, Cuntai Guan
Neural Networks6
2021 Adaptive transfer learning for EEG motor imagery classification with deep Convolutional Neural Network
Kaishuo Zhang, Neethu Robinson, Seong-Whan Lee, Cuntai Guan
Neural Networks4
2021 DSAL: Deeply Supervised Active Learning From Strong and Weak Labelers for Biomedical Image Segmentation
abstract
Image segmentation is one of the most essential biomedical image processing problems for different imaging modalities, including microscopy and X-ray in the Internet-of-Medical-Things (IoMT) domain. However, annotating biomedical images is knowledge-driven, time-consuming, and labor-intensive, making it difficult to obtain abundant labels with limited costs. Active learning strategies come into ease the burden of human annotation, which queries only a subset of training data for annotation. Despite receiving attention, most of active learning methods still require huge computational costs and utilize unlabeled data inefficiently. They also tend to ignore the intermediate knowledge within networks. In this work, we propose a deep active semi-supervised learning framework, DSAL, combining active learning and semi-supervised learning strategies. In DSAL, a new criterion based on deep supervision mechanism is proposed to select informative samples with high uncertainties and low uncertainties for strong labelers and weak labelers respectively. The internal criterion leverages the disagreement of intermediate features within the deep learning network for active sample selection, which subsequently reduces the computational costs. We use the proposed criteria to select samples for strong and weak labelers to produce oracle labels and pseudo labels simultaneously at each active learning iteration in an ensemble learning manner, which can be examined with IoMT Platform. Extensive experiments on multiple medical image datasets demonstrate the superiority of the proposed method over state-of-the-art active learning methods.
Ziyuan Zhao, Zeng Zeng, Kaixin Xu, Cen Chen 0001, Cuntai Guan
IEEE J. Biomed. Health Informatics5
2021 Generative Adversarial Networks-Based Data Augmentation for Brain-Computer Interface
abstract
The performance of a classifier in a brain-computer interface (BCI) system is highly dependent on the quality and quantity of training data. Typically, the training data are collected in a laboratory where the users perform tasks in a controlled environment. However, users' attention may be diverted in real-life BCI applications and this may decrease the performance of the classifier. To improve the robustness of the classifier, additional data can be acquired in such conditions, but it is not practical to record electroencephalogram (EEG) data over several long calibration sessions. A potentially time- and cost-efficient solution is artificial data generation. Hence, in this study, we proposed a framework based on the deep convolutional generative adversarial networks (DCGANs) for generating artificial EEG to augment the training set in order to improve the performance of a BCI classifier. To make a comparative investigation, we designed a motor task experiment with diverted and focused attention conditions. We used an end-to-end deep convolutional neural network for classification between movement intention and rest using the data from 14 subjects. The results from the leave-one subject-out (LOO) classification yielded baseline accuracies of 73.04% for diverted attention and 80.09% for focused attention without data augmentation. Using the proposed DCGANs-based framework for augmentation, the results yielded a significant improvement of 7.32% for diverted attention ( ) and 5.45% for focused attention ( ). In addition, we implemented the method on the data set IVa from BCI competition III to distinguish different motor imagery tasks. The proposed method increased the accuracy by 3.57% ( ). This study shows that using GANs for EEG augmentation can significantly improve BCI performance, especially in real-life applications, whereby users' attention may be diverted.
Fatemeh Fahimi, Strahinja Dosen, Kai Keng Ang, Natalie Mrachacz-Kersting, Cuntai Guan
IEEE Trans. Neural Networks Learn. Syst.5
2021 A Survey on Explainable Artificial Intelligence (XAI): Toward Medical XAI
abstract
Recently, artificial intelligence and machine learning in general have demonstrated remarkable performances in many tasks, from image processing to natural language processing, especially with the advent of deep learning (DL). Along with research progress, they have encroached upon many different fields and disciplines. Some of them require high level of accountability and thus transparency, for example, the medical sector. Explanations for machine decisions and predictions are thus needed to justify their reliability. This requires greater interpretability, which often means we need to understand the mechanism underlying the algorithms. Unfortunately, the blackbox nature of the DL is still unresolved, and many machine decisions are still poorly understood. We provide a review on interpretabilities suggested by different research works and categorize them. The different categories show different dimensions in interpretability research, from approaches that provide "obviously" interpretable information to the studies of complex patterns. By applying the same categorization to interpretability in medical research, it is hoped that: 1) clinicians and practitioners can subsequently approach these methods with caution; 2) insight into interpretability will be born with more considerations for medical practices; and 3) initiatives to push forward data-based, mathematically grounded, and technically grounded medical education are encouraged.
Erico Tjoa, Cuntai Guan
IEEE Trans. Neural Networks Learn. Syst.2
2020 TSception: A Deep Learning Framework for Emotion Detection Using EEG
abstract
In this paper, we propose a deep learning framework, TSception, for emotion detection from electroencephalogram (EEG). TSception consists of temporal and spatial convolutional layers, which learn discriminative representations in the time and channel domains simultaneously. The temporal learner consists of multi-scale 1D convolutional kernels whose lengths are related to the sampling rate of the EEG signal, which learns multiple temporal and frequency representations. The spatial learner takes advantage of the asymmetry property of emotion responses at the frontal brain area to learn the discriminative representations from the left and right hemispheres of the brain. In our study, a system is designed to study the emotional arousal in an immersive virtual reality (VR) environment. EEG data were collected from 18 healthy subjects using this system to evaluate the performance of the proposed deep learning network for the classification of low and high emotional arousal states. The proposed method is compared with SVM, EEGNet, and LSTM. TSception achieves a high classification accuracy of 86.03%, which outperforms the prior methods significantly (p<; 0.05).
Yi Ding 0012, Neethu Robinson, Qiuhao Zeng, Aung Aung Phyo Wai, Tih Shih Lee, Cuntai Guan
IJCNN7
2020 Optimizing Filter-bank Canonical Correlation Analysis for fast response SSVEP Brain-Computer Interface (BCI)
abstract
Steady-State Visual Evoked Potential (SSVEP) BCI brings high accuracy and consistent performance across subjects at the expense of a long stimulus presentation time window. Several recent methods exploited subject-specific features to improve SSVEP recognition performance in a short time window less than 1s. Although the calibration process is tedious and causes inconvenience, small calibration data with short duration resulting in higher performance gains are worth considering. So we propose a method by optimizing Filter-Bank Canonical Correlation Analysis (FBCCA) with subjects' calibrated templates, subject-specific weights and multiple reference types. The proposed method, subject-calibration extended FBCCA (SCEF) leverages independent and distinct discrimination characteristics of multiple references with subject-specific weight-adjusted features to improve SSVEP recognition performance. We tested the proposed method with different parameters compared with FBCCA baseline and state-of-the-art calibration methods on forty targets SSVEP dataset using 0.2s to 4s time windows. Our evaluation results show SCEF with three reference templates and subject-specific weighted features perform significantly better than all FBCCA variants in 0.2 s to 1 s time window (p <; 0.001). SCEF performs marginally, not statistically significant, better than existing methods about 2.69 ± 2.32% mean accuracy across time windows. Including multiple templates and subject-specific weight increases 15.73 ± 5.34% and 8.06 ± 2.06% in mean accuracy resulting the overall performance improvements in short time window. The proposed optimization only requires prior calibration data to create subject-specific templates and weights instead of learning features from calibration data every time. This enables not requiring to repeat the calibration step in every SSVEP session for the same subject while still maintaining accuracy similar to state-of-the-art calibration methods.
Aung Aung Phyo Wai, Ying Chi, Lei Zhang 0006, Xian-Sheng Hua 0001, Cuntai Guan
IJCNN6
2020 Deep Multi-Task Learning for SSVEP Detection and Visual Response Mapping
abstract
Glaucoma is an eye disease that occurs without the onset of symptoms at initial, and late diagnosis results in irreversible degeneration of retinal ganglion cells. Standard automated perimetry is the gold standard for assessing glaucoma; however, the examination is subjective, where responses can fluctuate each time the test is performed, significantly confounding the test's interpretation. In this study, we present our approach that aims to provide a rapid point-of-care diagnostics for glaucoma patients by eliminating the cognitive aspect in existing visual field assessment. Unlike existing methods that mostly report the foveal target detection's accuracy, we employed a multi-task learning architecture that efficiently captures signals simultaneously from the fovea and the neighboring targets in the peripheral vision, generating a visual response map. Furthermore, we designed a multi-task learning module that learns multiple tasks in parallel efficiently. We evaluated our model classification on a 40-classes dataset, with yields 92% and 95% in accuracy and F1 score respectively. Our model is able to perform on a calibration-free user-independent scenario, which is desirable for clinical diagnostics. Our proposed approach could be a stepping stone for an objective assessment of glaucoma patients' visual field.
Hong Jing Khok, Victor Teck Chang Koh, Cuntai Guan
SMC3
2020 A Context-Aware Locality Measure for Inlier Pool Enrichment in Stepwise Image Registration
abstract
We present a feature-based image registration method, the stepwise image registration (SIR), with a closed-form solution. Our SIR creates an inlier pool and a candidate pool as the initialization, and then gradually enriches the inlier pool and refines the transformation. In each step, the enriched correspondence exclusively tunes the transformation coefficient within the confirmed inlier pairs, instead of updating the mapping using the complete putative set. In turn, the refined transformation prunes inconsistent mismatches to alleviate the incoming matching ambiguity. The context-aware locality measure (CALM) is designed for dissimilarity measure. The capability of the CALM can be enhanced by the progressive inlier pool enrichment. Finally, a retrieval process is performed based on the finest CALM and alignment, by which the inlier pool is maximized. Extensive experiments of enrichment evaluation, feature matching, image registration, and image retrieval demonstrate the favorable performance of our SIR against state-of-the-art methods. The code and datasets are available at https://github.com/sucv/SIR.
Su Zhang 0004, Xuying Hao, Yang Yang 0032, Cuntai Guan
IEEE Trans. Image Process.5
2020 Subject-Independent Brain-Computer Interfaces Based on Deep Convolutional Neural Networks
abstract
For a brain-computer interface (BCI) system, a calibration procedure is required for each individual user before he/she can use the BCI. This procedure requires approximately 20-30 min to collect enough data to build a reliable decoder. It is, therefore, an interesting topic to build a calibration-free, or subject-independent, BCI. In this article, we construct a large motor imagery (MI)-based electroencephalography (EEG) database and propose a subject-independent framework based on deep convolutional neural networks (CNNs). The database is composed of 54 subjects performing the left- and right-hand MI on two different days, resulting in 21 600 trials for the MI task. In our framework, we formulated the discriminative feature representation as a combination of the spectral-spatial input embedding the diversity of the EEG signals, as well as a feature representation learned from the CNN through a fusion technique that integrates a variety of discriminative brain signal patterns. To generate spectral-spatial inputs, we first consider the discriminative frequency bands in an information-theoretic observation model that measures the power of the features in two classes. From discriminative frequency bands, spectral-spatial inputs that include the unique characteristics of brain signal patterns are generated and then transformed into a covariance matrix as the input to the CNN. In the process of feature representations, spectral-spatial inputs are individually trained through the CNN and then combined by a concatenation fusion technique. In this article, we demonstrate that the classification accuracy of our subject-independent (or calibration-free) model outperforms that of subject-dependent models using various methods [common spatial pattern (CSP), common spatiospectral pattern (CSSP), filter bank CSP (FBCSP), and Bayesian spatio-spectral filter optimization (BSSFO)].
O-Yeon Kwon, Min-Ho Lee, Cuntai Guan, Seong-Whan Lee
IEEE Trans. Neural Networks Learn. Syst.3
2019 A Comparative Study of Mental States in 2D and 3D Virtual Environments Using EEG
abstract
There is growing evidence that virtual reality (VR) can be used as an optional therapy method for stress relief and treatment of mental disorders such as a variety of anxiety disorders, depression and psychosis. However, more systematic studies to quantity the effects of VR for emotion elicitation and compare the effects of 3D and 2D environment for the same is necessary to design feasible and effective treatments. In this study, we design a cross-over experiment protocol comprising of two emotion eliciting environments relaxation and arousal that is presented to the participant either in a 2D monitor or a 3D head-mounted (HMD) VR display. The EEG data is collected during the experiment and analyzed offline to classify emotion elicited by low and high arousal environments using SVM classification of band power features. A 10-fold cross validation is performed and classification accuracy for each subject is computed for 3D-VR and 2D-screen display. The average classification accuracies over subjects are obtained as 66.88% and 59.27% in 3D-VR and 2D-screen group respectively. The performance difference between 3D-VR and 2D screen is statistically significant $(p\lt 0.05)$, indicating that 3D-VR generates more distinct EEG patterns associated with emotion elicitation. Further, significant differences in band powers from alpha, theta and beta bands are also observed between both groups. The results presented in this paper support the use of VR as an effective tool to study emotion elicitation and regulation, compared to a 2D screen.
Pinar Bilgin, Kat Agres, Neethu Robinson, Aung Aung Phyo Wai, Cuntai Guan
SMC5
2019 EEG Representation in Deep Convolutional Neural Networks for Classification of Motor Imagery
abstract
With deep learning emerging as a powerful machine learning tool to build Brain Computer Interface (BCI) systems, researchers are investigating the use of different type of networks architectures and representations of brain activity to attain superior classification accuracy compared to state-of-the-art machine learning approaches, that rely on processed signal and optimally extracted features. This paper presents a deep learning driven electroencephalography (EEG) -BCI system to perform decoding of hand motor imagery using deep convolution neural network architecture, with spectrally localized time-domain representation of multi-channel EEG as input. A significant increase in decoding performance in terms of accuracy of +6.47% is obtained compared to a wideband EEG representation. We further illustrate the movement class specific feature patterns for both the architectures and demonstrate that higher difference between classes is observed using the proposed architecture. We conclude that the network trained by taking into account the dynamic spatial interactions in distinct frequency bands of EEG, can offer better decoding performance and aid in better interpretation of learned features.
Neethu Robinson, Seong-Whan Lee, Cuntai Guan
SMC3
2018 A Study of SSVEP Responses in Case of Overt and Covert Visual Attention with Different View Angles
abstract
Standard automated perimetry is a common visual field test in clinical practices. But the test effectiveness relies on responses from subject and technician experience in operating the test equipment. Therefore, it calls for a more objective way of measuring visual field; as such we consider SSVEP as potential suitable technique. SSVEP is extensively studied in the context of a brain-computer interface, where successful SSVEP detection relies on fovea vision. But peripheral vision is more critical in assessing the effective field of view in Glaucoma patients. So this study investigates how SSVEP responses exhibit with different view angles and subject's visual attention in peripheral vision. We designed an experiment with single flickering stimulus at three view angles horizontally. Subject performed overt, covert and no visual attention at each stimulus position, while EEG and eye tracking data are recorded simultaneously. We used spectral power amplitude ratio and maximum canonical correlation coefficients to evaluate SSVEP responses. By applying 1-way and 2-way ANOVA tests, there is no statistical significant difference among SSVEP responses when subject paid overt or covert visual attention with different view angles. This might suggest that SSVEP can be a usable approach to measure visual fields when subject oriented with covert visual attention. So reliable SSVEP responses highly depend on the choice of stimulus frequency but do not depend significantly on different visual attention and view angles.
Aung Aung Phyo Wai, Zhiyan Goh, Shi De Foo, Cuntai Guan
SMC4
2018 Learning Temporal Information for Brain-Computer Interface Using Convolutional Neural Networks
abstract
Deep learning (DL) methods and architectures have been the state-of-the-art classification algorithms for computer vision and natural language processing problems. However, the successful application of these methods in motor imagery (MI) brain-computer interfaces (BCIs), in order to boost classification performance, is still limited. In this paper, we propose a classification framework for MI data by introducing a new temporal representation of the data and also utilizing a convolutional neural network (CNN) architecture for classification. The new representation is generated from modifying the filter-bank common spatial patterns method, and the CNN is designed and optimized accordingly for the representation. Our framework outperforms the best classification method in the literature on the BCI competition IV-2a 4-class MI data set by 7% increase in average subject accuracy. Furthermore, by studying the convolutional weights of the trained networks, we gain an insight into the temporal characteristics of EEG.
Siavash Sakhavi, Cuntai Guan, Shuicheng Yan
IEEE Trans. Neural Networks Learn. Syst.2
2017 Effects of transcranial direct current stimulation on the motor-imagery brain-computer interface for stroke recovery: An EEG source-space study
abstract
Recently, noninvasive brain stimulation is gaining significant attention in stroke rehabilitation. In this paper, we investigate the effects of transcranial direct current stimulation (tDCS) on the motor-imagery brain-computer interface (MI-BCI) performance of stroke patients. To this end, we processed the EEG data collected from a randomized control trial (RCT) study of 19 stroke patients grouped into tDCS and sham. An ensemble method for feature extraction is proposed in this study that combines shrinkage regularized Common Spatial Pattern (CSP) features from the sensor-space and the cortical source-space Electroencephalography (EEG) across ten rehabilitation sessions. The classification results of MI vs. Idle state in stroke patients show that the concatenated features from both the sensor- and source space EEG provided an average cross-validation accuracy of 64.5% which is statistically significant (p <; 0.001) compared to either using source or sensor-space features. Further, our findings suggest that the effect of tDCS on the stroke recovery is pronounced in subjects whose delta and alpha band power during the post-tDCS intervention is significantly higher as compared to before intervention. The group-averaged sLORETA activation results showed a significantly higher number of dipoles activated in the tDCS group as compared to sham. In summary, our study paves a new way to analyze the neural correlates of the MI-BCI performance for stroke rehabilitation.
A. Prasad Vinod 0001, Kai Keng Ang, Effie Chew, Cuntai Guan
SMC5
2017 Facilitating motor imagery-based brain-computer interface for stroke patients using passive movement
abstract
Motor imagery-based brain–computer interface (MI-BCI) has been proposed as a rehabilitation tool to facilitate motor recovery in stroke. However, the calibration of a BCI system is a time-consuming and fatiguing process for stroke patients, which leaves reduced time for actual therapeutic interaction. Studies have shown that passive movement (PM) (i.e., the execution of a movement by an external agency without any voluntary motions) and motor imagery (MI) (i.e., the mental rehearsal of a movement without any activation of the muscles) induce similar EEG patterns over the motor cortex. Since performing PM is less fatiguing for the patients, this paper investigates the effectiveness of calibrating MI-BCIs from PM for stroke subjects in terms of classification accuracy. For this purpose, a new adaptive algorithm called filter bank data space adaptation (FB-DSA) is proposed. The FB-DSA algorithm linearly transforms the band-pass-filtered MI data such that the distribution difference between the MI and PM data is minimized. The effectiveness of the proposed algorithm is evaluated by an offline study on data collected from 16 healthy subjects and 6 stroke patients. The results show that the proposed FB-DSA algorithm significantly improved the classification accuracies of the PM and MI calibrated models ( p < 0.05). According to the obtained classification accuracies, the PM calibrated models that were adapted using the proposed FB-DSA algorithm outperformed the MI calibrated models by an average of 2.3 and 4.5 % for the healthy and stroke subjects respectively. In addition, our results suggest that the disparity between MI and PM could be stronger in the stroke patients compared to the healthy subjects, and there would be thus an increased need to use the proposed FB-DSA algorithm in BCI-based stroke rehabilitation calibrated from PM.
Mahnaz Arvaneh, Cuntai Guan, Kai Keng Ang, Tomás Ward, Karen Sui Geok Chua, Christopher Wee Keong Kuah, Gopal Joseph Ephraim Joseph, Koksoon Phua, Chuanchu Wang
Neural Comput. Appl.2
2017 A New Variational Method for Bias Correction and Its Applications to Rodent Brain Extraction
abstract
Brain extraction is an important preprocessing step for further analysis of brain MR images. Significant intensity inhomogeneity can be observed in rodent brain images due to the high-field MRI technique. Unlike most existing brain extraction methods that require bias corrected MRI, we present a high-order and L0regularized variational model for bias correction and brain extraction. The model is composed of a data fitting term, a piecewise constant regularization and a smooth regularization, which is constructed on a 3-D formulation for medical images with anisotropic voxel sizes. We propose an efficient multi-resolution algorithm for fast computation. At each resolution layer, we solve an alternating direction scheme, all subproblems of which have the closed-form solutions. The method is tested on three T2 weighted acquisition configurations comprising a total of 50 rodent brain volumes, which are with the acquisition field strengths of 4.7 Tesla, 9.4 Tesla and 17.6 Tesla, respectively. On one hand, we compare the results of bias correction with N3 and N4 in terms of the coefficient of variations on 20 different tissues of rodent brain. On the other hand, the results of brain extraction are compared against manually segmented gold standards, BET, BSE and 3-D PCNN based on a number of metrics. With the high accuracy and efficiency, our proposed method can facilitate automatic processing of large-scale brain studies.
Huibin Chang, Weimin Huang 0002, Su Huang, Cuntai Guan, Sakthivel Sekar, Kishore Kumar Bhakoo, Yuping Duan
IEEE Trans. Medical Imaging5
2017 A Unified Fisher's Ratio Learning Method for Spatial Filter Optimization
abstract
To detect the mental task of interest, spatial filtering has been widely used to enhance the spatial resolution of electroencephalography (EEG). However, the effectiveness of spatial filtering is undermined due to the significant nonstationarity of EEG. Based on regularization, most of the conventional stationary spatial filter design methods address the nonstationarity at the cost of the interclass discrimination. Moreover, spatial filter optimization is inconsistent with feature extraction when EEG covariance matrices could not be jointly diagonalized due to the regularization. In this paper, we propose a novel framework for a spatial filter design. With Fisher's ratio in feature space directly used as the objective function, the spatial filter optimization is unified with feature extraction. Given its ratio form, the selection of the regularization parameter could be avoided. We evaluate the proposed method on a binary motor imagery data set of 16 subjects, who performed the calibration and test sessions on different days. The experimental results show that the proposed method yields improvement in classification performance for both single broadband and filter bank settings compared with conventional nonunified methods. We also provide a systematic attempt to compare different objective functions in modeling data nonstationarity with simulation studies.To detect the mental task of interest, spatial filtering has been widely used to enhance the spatial resolution of electroencephalography (EEG). However, the effectiveness of spatial filtering is undermined due to the significant nonstationarity of EEG. Based on regularization, most of the conventional stationary spatial filter design methods address the nonstationarity at the cost of the interclass discrimination. Moreover, spatial filter optimization is inconsistent with feature extraction when EEG covariance matrices could not be jointly diagonalized due to the regularization. In this paper, we propose a novel framework for a spatial filter design. With Fisher's ratio in feature space directly used as the objective function, the spatial filter optimization is unified with feature extraction. Given its ratio form, the selection of the regularization parameter could be avoided. We evaluate the proposed method on a binary motor imagery data set of 16 subjects, who performed the calibration and test sessions on different days. The experimental results show that the proposed method yields improvement in classification performance for both single broadband and filter bank settings compared with conventional nonunified methods. We also provide a systematic attempt to compare different objective functions in modeling data nonstationarity with simulation studies.
Cuntai Guan, Haihong Zhang, Kai Keng Ang
IEEE Trans. Neural Networks Learn. Syst.2
2016 Multiway analysis of EEG artifacts based on Block Term Decomposition
abstract
Neural information recorded from electroencephalogram (EEG) provides new possibilities for diagnosis of brain abnormalities, cognitive monitoring, etc. However, many artifacts, such as eye blink and muscle movements, impact and contaminate EEG data. While traditional techniques proposed for artifact removal identified artifact on two-way data, (spatial x temporal), multidimensional nature of EEG data (spatial x temporal x spectral x condition x trial) is overlooked. In this work, we investigate the use of multiway analysis/tensor factorization on the extended EEG tensor (spatial x temporal x spectral), which is constructed from continuous wavelet transform, using Block Term Decomposition (BTD) of rank-(Lr, Lr, 1) for artifact removal. Eight different carefully designed experiments to study artifact typically produced by voluntarily, and sometimes involuntarily, behaviors using a subject were performed and analyzed. After the BTD decomposition, artifacted components are automatically identified removed using spatial and temporal features. The reconstructed signal from proposed method suppresses artifact while retains the signal texture of eight types of artifact investigated.
Sim Kuan Goh, Hussein A. Abbass, Kay Chen Tan, Abdullah Al Mamun 0002, Cuntai Guan, Chuanchu Wang
IJCNN5
2016 A novel supervised locality sensitive Factor analysis to classify voluntary hand movement in multi direction using EEG source space
abstract
Recent advances in EEG-based brain-computer interfaces (BCIs) have shown that brain signals can be used to decode arm movement intention and execution in multiple directions. Conventional approaches use sensor space EEG for classifying movement related tasks. Sensor-space EEG can reveal only limited information about the trivial but complex tasks that involve higher degrees of freedom of the movement. On the contrary, source space analysis is expected to provide more information about the neurophysiological mechanism relevant to the task. To this end, we propose a novel source-space feature extraction technique based on supervised locality sensitive Factor analysis which approximates the neurophysiological functioning of our experimental data in a better way than that of a solely data-driven approach. EEG recordings in the sensor space are transformed into source space using the Weighted Minimum Norm Estimate (wMNE) method. We show that for a multi-class classification problem of classifying the EEG of voluntary arm movement in 4 orthogonal directions, the source space features offer a significant improvement in the classification accuracy compared to sensor space features. One-versus-rest (OVR) approach is used for multiclass classification with Fisher's Linear Discriminant (FLD) as the primary classifier.
A. Prasad Vinod 0001, Cuntai Guan
SMC3
2015 Stress level detection using double-layer subband filter
Tin Lay Nwe, Qianli Xu, Cuntai Guan, Bin Ma 0001
INTERSPEECH3
2015 Cortical Source Localization for Analysing Single-Trial Motor Imagery EEG
abstract
Electroencephalography (EEG) is the most widely used Brain-Computer Interface (BCI) modality to record brain signal. Unlike other neuroimaging modalities like fMRI and PET, EEG is not very effective in localizing the brain sources. However, with the advent of inverse modeling techniques for source localization, it is possible to use EEG as an alternative neuroimaging technique. In this paper, source localization using EEG signal is used to analyze single-trial movement imagination (MI) tasks. Wadsworth physiobank dataset of 109 subjects performing right hand vs left hand movement imagination is considered. Forward modeling based on 3 layered head geometry is co-registered with ICBM 152 template anatomy, which is a non-linear average of fMRI scans of 152 subjects. Inverse modeling is done with the help of Standardized Low Resolution Electromagnetic Tomography (sLORETA). The proposed method presents some preliminary results on how source localization could be used to identify the moment (time instant) of brain source activation even within a single trial.
A. Prasad Vinod 0001, Cuntai Guan
SMC3
2015 Brain-Computer Interface for Neurorehabilitation of Upper Limb After Stroke
abstract
Current rehabilitation therapies for stroke rely on physical practice (PP) by the patients. Motor imagery (MI), the imagination of movements without physical action, presents an alternate neurorehabilitation for stroke patients without relying on residue movements. However, MI is an endogenous mental process that is not physically observable. Recently, advances in brain-computer interface (BCI) technology have enabled the objective detection of MI that spearheaded this alternate neurorehabilitation for stroke. In this review, we present two strategies of using BCI for neurorehabilitation after stroke: detecting MI to trigger a feedback, and detecting MI with a robot to provide concomitant MI and PP. We also present three randomized control trials that employed these two strategies for upper limb rehabilitation. A total of 125 chronic stroke patients were screened over six years. The BCI screening revealed that 103 (82%) patients can use electroencephalogram-based BCI, and 75 (60%) performed well with accuracies above 70%. A total of 67 patients were recruited to complete one of the three RCTs ranging from two to six weeks of which 26 patients, who underwent BCI neurorehabilitation that employed these two strategies, had significant motor improvement of 4.5 measured by Fugl-Meyer Motor Assessment of the upper extremity. Hence, the results demonstrate clinical efficacy of using BCI as an alternate neurorehabilitation for stroke.
Kai Keng Ang, Cuntai Guan
Proc. IEEE2
2015 Cluster-Based Analysis for Personalized Stress Evaluation Using Physiological Signals
abstract
Technology development in wearable sensors and biosignal processing has made it possible to detect human stress from the physiological features. However, the intersubject difference in stress responses presents a major challenge for reliable and accurate stress estimation. This research proposes a novel cluster-based analysis method to measure perceived stress using physiological signals, which accounts for the intersubject differences. The physiological data are collected when human subjects undergo a series of task-rest cycles, incurring varying levels of stress that is indicated by an index of the State Trait Anxiety Inventory. Next, a quantitative measurement of stress is developed by analyzing the physiological features in two steps: 1) a k -means clustering process to divide subjects into different categories (clusters), and 2) cluster-wise stress evaluation using the general regression neural network. Experimental results show a significant improvement in evaluation accuracy as compared to traditional methods without clustering. The proposed method is useful in developing intelligent, personalized products for human stress management.
Qianli Xu, Tin Lay Nwe, Cuntai Guan
IEEE J. Biomed. Health Informatics3
2014 Spatio-temporal variations in hand movement trajectory based brain activation patterns
abstract
The neuro engineering research over the past decades has established Electroencephalography based Brain Computer Interface (EEG-BCI) systems as an efficient means of decoding brain activity. Motor control BCI is a category of BCI that analyzes neural activity recorded over sensory motor area to classify or decode intended motor tasks. For a BCI system, it is desired to have defined and independent output control commands. Decoding movement trajectory parameters such as instantaneous position, speed coordinates from non-invasive brain recordings can thus be a key contribution in motor control BCI applications. In this study, we use Multiple Linear Regressor to estimate the hand movement trajectory from spectrally localized multi-sensor EEG. The algorithm is validated using data collected from subjects as they perform 2-dimensional center-out hand movement towards pre-defined targets at varying speeds. The spatio-temporal variations in motor activity based neural activation patterns using metrics derived from MLR estimator is investigated. The contribution of the predictors to the regression equation and decoding performance at various stages of movement are also studied. An average correlation of 0.63 (p<;0.005) between recorded and estimated trajectory is obtained using the method. The temporally varying involvement of motor, pre-motor and parietal areas; movement task dependent activations and time-varying sensor contribution in reconstruction are further demonstrated.
Neethu Robinson, A. Prasad Vinod 0001, Cuntai Guan
ICARCV3
2014 Spatial filter adaptation based on geodesic-distance for motor EEG classification
abstract
The non-stationarity inherent across sessions recorded on different days poses a major challenge for practical electroencephalography (EEG)-based Brain Computer Interface (BCI) systems. To address this issue, the computational model trained using the training data needs to adapt to the data from the test sessions. In this paper, we propose a novel approach to compute the variations between labelled training data and a batch of unlabelled test data based on the geodesic-distance of the discriminative subspaces of EEG data on the Grassmann manifold. Subsequently, spatial filters can be updated and features that are invariant against such variations can be obtained using a subset of training data that is closer to the test data. Experimental results show that the proposed adaptation method yielded improvements in classification performance.
Cuntai Guan, Kai Keng Ang, Haihong Zhang, Sim Heng Ong
IJCNN2
2014 Evaluation of EEG features during overt visual attention during neurofeedback game
abstract
Brain-Computer Interface (BCI) is an emerging modality for direct communication between brain and computer, bypassing brain's conventional communication pathway of the nerves and muscles. Though BCI investigations have been targeting on the development of assistive devices for paralyzed patients initially, recent BCI research exploits the possibilities of BCI in entertainment and cognitive-skill enhancement through neurofeedback games also. Neurofeedback is an effective tool for boosting cognitive skills of both healthy and attention-deficit people based on real-time feedback and self-regulation of brain signals. This paper investigates the feasibility of employing EEG features related to sustained attention and overt visual attention-shift towards left or right visual periphery from a fixation point in the context of a neurofeedback game. Three healthy subjects have successfully played the proposed neurofeedback game by selecting the targets solely by EEG features related to overt visual attention, offering an average accuracy of 72.22%.
Kavitha P. Thomas, A. Prasad Vinod 0001, Cuntai Guan
SMC3
2014 A BCI speller based on SSVEP using high frequency stimuli design
abstract
We developed and studied a Steady-State Visual Evoked Potential (SSVEP) based BCI system using a high frequency visual stimuli (>25Hz) design for reducing visual fatigue. Existing SSVEP based BCI designs primarily use low frequency visual stimuli (<;20Hz) for eliciting relatively higher SSVEP signal, while the low frequency stimuli can provocate photosensitivity epileptic seizure. On the other hand, high frequency stimuli are visually more comfortable and cause less visual fatigue and seizure. To detect the weak high frequency SSVEP signal, we used multi-channel EEG and introduced canonical correlation analysis to identify the elicited SSVEP frequency. We designed and built a 30-character SSVEP BCI speller system without calibration and evaluated the performance metrics including classification accuracy and subjective fatigue ratings, in both high-frequency and low frequency SSVEP modes. The result indicates that the high frequency stimuli system archieved higher classification accuracy (averaged 80% in the 30-class classification) comparable to that by the low frequency system. Moreover, no subjects rated the visual feeling as unacceptable or uncomfortable with the high frequency system.
Dong-Ok Won, Haihong Zhang, Cuntai Guan, Seong-Whan Lee
SMC3
2014 Quality assessment of EEG signals based on statistics of signal fluctuations
abstract
The quality of the non-invasive EEG signals was always affected by the changes in the contact impedances and the artifacts from eye blinking, eye movements and body movements. An effective quality assessment method is needed to assess the qualities of the EEG signals. This paper proposed a novel method to assess the signal quality of EEG signals based on block-based measurements of the fluctuations of the second-order power amplitudes of the EEG signals. The initial signal quality scores were generated by fusion of the mean power amplitudes and the signal fluctuations of the motor imagery state with reference to background idling state. These scores were subsequently mapped to different quality levels by using fuzzy-c means clustering. Experimental results were conducted on the basis of 3 data sets of 15 healthy subjects performing motor imagery of hand movements and idle, for both gel-based and gel-less electrodes. The results obtained demonstrated that the proposed method was capable of evaluating the quality of the EEG signals, as supported by the clear separation of the assigned quality levels between gel-based and gel-less electrodes. This further validated the assumption that generally the quality of the EEG signals acquired based on the gel-based electrodes was better than that of the gel-less electrodes.
Huijuan Yang, Cuntai Guan, Kai Keng Ang, Koksoon Phua, Chuanchu Wang
SMC2
2014 Mutual information-based optimization of sparse spatio-spectral filters in brain-computer interface
Mahnaz Arvaneh, Cuntai Guan, Kai Keng Ang, Hiok Chai Quek
Neural Comput. Appl.2
2013 Joint spatial-temporal filter design for analysis of motor imagery EEG
abstract
This paper addresses the key issue of discriminative feature extraction of electroencephalogram (EEG) signals in brain-computer interfaces. Recent advances in neuroscience indicate that multiple brain regions can be activated during motor imagery. The signal propagation among the regions can give rise to spurious effects in identifying event-related desynchronization/synchronization for discriminative motor imagery detection in conventional feature extraction methods. Particularly, we propose that computational models which account for both signal propagation and volume conduction effects of the source neuronal activities can more accurately describe EEG during the specific brain activities and lead to more effective feature extraction. To this end, we devise a unified model for joint learning of signal propagation and spatial patterns. The preliminary results obtained with real-world motor imagery EEG data sets confirm that the new methodology can improve classification accuracy with statistical significance.
Haihong Zhang, Cuntai Guan, Sim Heng Ong, Yaozhang Pan, Kai Keng Ang
ICASSP3
2013 Error entropy based adaptive kernel classification for non-stationary EEG analysis
abstract
The performance of Brain-Computer Interface (BCI) applications are sometimes hindered by non-stationarity in the EEG data from sessions on different days. This paper proposes an algorithm for adaptive training of a SVM classifier to address the non-stationarity in EEG by adapting the kernel to data from subsequent sessions. The kernel width parameter of the kernel function of the SVM classifier is adapted using an information theoretic cost function based on minimum error entropy (MEE). An experiment is performed using the proposed method on EEG data collected without feedback from 12 healthy subjects in two sessions on separate days. The results using the proposed method yielded a mean accuracy of 75%, which is significantly better compared to the baseline result of 67% without kernel adaptation (P=0.00029).
Sidath Ravindra Liyanage, Cuntai Guan, Haihong Zhang, Kai Keng Ang, Jianxin Xu 0001, Tong Heng Lee
ICASSP2
2013 Neural decoding of movement targets by unsorted spike trains
abstract
Decoding movement targets from neural activity in motor cortex using invasive brain-computer interface (BCI) has potential application to help disabled patients. Most works employed spike sorting to obtain the single units (SUs) for decoding from the extracellular electrode recordings. However, spike sorting is difficult, computational demanding, and is often limited by the spike waveform variability especially in low SNR and high neuronal density conditions. To address these issues, we proposed a decoding method using unsorted spike trains from recording electrodes based on the maximal likelihood (ML) estimation approach. An experiment was performed to test neuronal data recorded from a rhesus monkey performing the center-out movement task of eight targets. The results showed that the proposed method yielded average correct decoding rate of 98.5% compared to the SU based method that yielded correct decoding rate of 96.3%. The results also showed that the proposed method yielded improved computational efficiency. Thus the proposed method showed potential for real time BCI applications with large scale of neuronal recordings.
Kai Keng Ang, Cuntai Guan
ICASSP3
2013 Maximum dependency and minimum redundancy-based channel selection for motor imagery of walking EEG signal detection
abstract
This paper proposes a novel method to detect motor imagery of walking for the rehabilitation of stroke patients using the Laplacian derivatives (LAD) of power averaged across frequency bands as the feature. We propose to select the most correlated channels by jointly considering the mutual information between the LAD power features of the channels and the class labels, and the redundancy between the LAD power features of the channel with that of the selected channels. Experiments are conducted on the EEG data collected for 11 healthy subjects using proposed method and compared with existing methods. The results show that the proposed method yielded an average classification accuracy of 67.19% by selecting as few as 4 LAD channels. An improved result of 71.45% and 73.23% are achieved by selecting 10 and 22 LAD channels, respectively. Comparison results revealed significantly superior performance of our proposed method compared to that obtained using common spatial pattern and filter bank with power features. Most importantly, our proposed method achieves significant better accuracy for poor BCI performers compared to existing methods. Thus, the results demonstrated the potential of using the proposed method for detecting motor imagery of walking for the rehabilitation of stroke patients.
Huijuan Yang, Cuntai Guan, Chuanchu Wang, Kai Keng Ang
ICASSP2
2013 Hand Movement Trajectory Reconstruction from EEG for Brain-Computer Interface Systems
abstract
Decoding hand movement parameters (for example movement trajectory, speed etc.) from scalp recordings such as Electroencephalography (EEG) is a challenging and less explored area of research in the field of Brain Computer Interface (BCI) systems. By identifying neural features underlying movement parameters, a detailed and well defined control command set can be provided to the BCI output device. A continuous control to the output device is better suited for practical BCI systems, and can be achieved by continuous reconstruction of movement trajectory than discrete brain activity classifications. In this study, we attempt to reconstruct/estimate various parameters of hand movement trajectory from multi channel EEG recordings. The data for analysis is collected by performing an experiment that involved centre-out right hand movement tasks in four different directions at two different speeds in random order. Multiple linear regression (MLR) strategy that fits the recorded movement parameters to a set of spatial, spectral and temporal localized neural data set is adopted. We propose a method to define the predictor set for MLR, using wavelet analysis, to decompose the signal into various sub bands. The correlation between recorded and estimated parameters are calculated and an average correlation coefficient of (0.56 ± 0.16) is obtained over estimating six movement parameters. The promising results achieved using the proposed algorithm, which are better than that of the existing algorithms, indicate the applicability of EEG for continuous motor control.
Neethu Robinson, A. Prasad Vinod 0001, Cuntai Guan
SMC3
2013 eT2FIS: An Evolving Type-2 Neural Fuzzy Inference System
Sau Wai Tung, Hiok Chai Quek, Cuntai Guan
Inf. Sci.3
2013 EEG Data Space Adaptation to Reduce Intersession Nonstationarity in Brain-Computer Interface
abstract
A major challenge in EEG-based brain-computer interfaces (BCIs) is the intersession nonstationarity in the EEG data that often leads to deteriorated BCI performances. To address this issue, this letter proposes a novel data space adaptation technique, EEG data space adaptation (EEG-DSA), to linearly transform the EEG data from the target space (evaluation session), such that the distribution difference to the source space (training session) is minimized. Using the Kullback-Leibler (KL) divergence criterion, we propose two versions of the EEG-DSA algorithm: the supervised version, when labeled data are available in the evaluation session, and the unsupervised version, when labeled data are not available. The performance of the proposed EEG-DSA algorithm is evaluated on the publicly available BCI Competition IV data set IIa and a data set recorded from 16 subjects performing motor imagery tasks on different days. The results show that the proposed EEG-DSA algorithm in both the supervised and unsupervised versions significantly outperforms the results without adaptation in terms of classification accuracy. The results also show that for subjects with poor BCI performances when no adaptation is applied, the proposed EEG-DSA algorithm in both the supervised and unsupervised versions significantly outperforms the unsupervised bias adaptation algorithm (PMean).
Mahnaz Arvaneh, Cuntai Guan, Kai Keng Ang, Hiok Chai Quek
Neural Comput.2
2013 Discriminative Learning of Propagation and Spatial Pattern for Motor Imagery EEG Analysis
abstract
Effective learning and recovery of relevant source brain activity patterns is a major challenge to brain-computer interface using scalp EEG. Various spatial filtering solutions have been developed. Most current methods estimate an instantaneous demixing with the assumption of uncorrelatedness of the source signals. However, recent evidence in neuroscience suggests that multiple brain regions cooperate, especially during motor imagery, a major modality of brain activity for brain-computer interface. In this sense, methods that assume uncorrelatedness of the sources become inaccurate. Therefore, we are promoting a new methodology that considers both volume conduction effect and signal propagation between multiple brain regions. Specifically, we propose a novel discriminative algorithm for joint learning of propagation and spatial pattern with an iterative optimization solution. To validate the new methodology, we conduct experiments involving 16 healthy subjects and perform numerical analysis of the proposed algorithm for EEG classification in motor imagery brain-computer interface. Results from extensive analysis validate the effectiveness of the new methodology with high statistical significance.
Haihong Zhang, Cuntai Guan, Sim Heng Ong, Kai Keng Ang, Yaozhang Pan
Neural Comput.3
2013 Optimizing Spatial Filters by Minimizing Within-Class Dissimilarities in Electroencephalogram-Based Brain-Computer Interface
abstract
A major challenge in electroencephalogram (EEG)-based brain-computer interfaces (BCIs) is the inherent nonstationarities in the EEG data. Variations of the signal properties from intra and inter sessions often lead to deteriorated BCI performances, as features extracted by methods such as common spatial patterns (CSP) are not invariant against the changes. To extract features that are robust and invariant, this paper proposes a novel spatial filtering algorithm called Kullback-Leibler (KL) CSP. The CSP algorithm only considers the discrimination between the means of the classes, but does not consider within-class scatters information. In contrast, the proposed KLCSP algorithm simultaneously maximizes the discrimination between the class means, and minimizes the within-class dissimilarities measured by a loss function based on the KL divergence. The performance of the proposed KLCSP algorithm is compared against two existing algorithms, CSP and stationary CSP (sCSP), using the publicly available BCI competition III dataset IVa and a large dataset from stroke patients performing neuro-rehabilitation. The results show that the proposed KLCSP algorithm significantly outperforms both the CSP and the sCSP algorithms, in terms of classification accuracy, by reducing within-class variations. This results in more compact and separable features.
Mahnaz Arvaneh, Cuntai Guan, Kai Keng Ang, Hiok Chai Quek
IEEE Trans. Neural Networks Learn. Syst.2
2013 Bayesian Learning for Spatial Filtering in an EEG-Based Brain-Computer Interface
abstract
Spatial filtering for EEG feature extraction and classification is an important tool in brain-computer interface. However, there is generally no established theory that links spatial filtering directly to Bayes classification error. To address this issue, this paper proposes and studies a Bayesian analysis theory for spatial filtering in relation to Bayes error. Following the maximum entropy principle, we introduce a gamma probability model for describing single-trial EEG power features. We then formulate and analyze the theoretical relationship between Bayes classification error and the so-called Rayleigh quotient, which is a function of spatial filters and basically measures the ratio in power features between two classes. This paper also reports our extensive study that examines the theory and its use in classification, using three publicly available EEG data sets and state-of-the-art spatial filtering techniques and various classifiers. Specifically, we validate the positive relationship between Bayes error and Rayleigh quotient in real EEG power features. Finally, we demonstrate that the Bayes error can be practically reduced by applying a new spatial filter with lower Rayleigh quotient.
Haihong Zhang, Huijuan Yang, Cuntai Guan
IEEE Trans. Neural Networks Learn. Syst.3
2012 Extracting effective features from high density nirs-based BCI for assessing numerical cognition
abstract
Near-infrared spectroscopy (NIRS)-based Brain-Computer Interface (BCI) was recently proposed to assess level of numerical cognition in subjects. However, existing feature extraction method was only proposed for low density 16 channels NIRS-based BCI. This study investigates the performance of a high density 348 channels NIRS-based BCI on 8 healthy subjects while they solve mental arithmetic problems with two difficulty levels and the rest condition. A novel method of extracting effective features from high density single-trial NIRS data is proposed using common average reference spatial filtering and single-trial baseline reference. The performance of the proposed feature extraction method is presented using 5×5-fold cross-validations on the single-trial NIRS data collected using mutual information-based feature selection and support vector machine classifier. The results yielded an overall average accuracy of 73% and 92% in classifying hard versus easy tasks and hard versus rest tasks respectively using the proposed method, compared to 46% and 62% respectively using existing method. The results demonstrated the effectiveness of using the proposed method in high density NIRS-based BCI for assessing numerical cognition.
Kai Keng Ang, Juanhong Yu, Cuntai Guan
ICASSP3
2012 Multi-frequency band common spatial pattern with sparse optimization in Brain-Computer Interface
abstract
In motor imagery-based Brain Computer Interfaces (BCIs), Common Spatial Pattern (CSP) algorithm is widely used for extracting discriminative patterns from the EEG signals. However, the CSP algorithm is known to be sensitive to noise and artifacts, and its performance greatly depends on the operational frequency band. To address these issues, this paper proposes a novel Sparse Multi-Frequency Band CSP (SMFBCSP) algorithm optimized using a mutual information-based approach. Compared to the use of the cross-validation-based method which finds the regularization parameters by trial and error, the proposed mutual information-based approach directly computes the optimal regularization parameters such that the computational time is substantially reduced. The experimental results on 11 stroke patients showed that the proposed SMFBCSP significantly outperformed three existing algorithms based on CSP, sparse CSP and filter bank CSP in terms of classification accuracy.
Mahnaz Arvaneh, Cuntai Guan, Kai Keng Ang, Hiok Chai Quek
ICASSP2
2012 Fast emotion detection from EEG using asymmetric spatial filtering
abstract
The injection of emotional intelligence in human-computer interfaces is necessary for computer applications to appear intelligent when interacting with people. With the recent development of brain imaging techniques and brain-computer interfaces, computers can actually take a look inside users' head to observe their emotional states. This paper presents an EEG-based emotion detection system which detects emotional states based on short EEG segments of 1s. A novel feature extraction algorithm termed asymmetric spatial filtering is proposed to extract features from high dimensional EEG data. The effectiveness of the proposed method is tested for two types of emotion detection problems on data from five subjects.
Haihong Zhang, Kai Keng Ang, Cuntai Guan, Yaozhang Pan, Chuanchu Wang, Juanhong Yu
ICASSP4
2012 Cluster impurity and forward-backward error maximization-based active learning for EEG signals classification
abstract
This paper investigates how to apply active learning for the classification of motor imagery electroencephalography (EEG) signals to boost the performance for small training size. A new criterion is proposed to select the most representative and informative queries. The candidates are firstly chosen from the samples close to the center of the cluster that has the highest impurity of classes. A predefined number of such candidates and classifiers are forwardly buffered. Subsequently, the query is chosen such that the buffered classifiers can backward maximize the classification errors on labeled data. Experimental results conducted on the BCI competition IV data set IVb show the superior performance of the proposed active learning scheme, which is on average 5.12% higher in accuracy than that of the passive method by choosing the training size from 28 to 112.
Huijuan Yang, Cuntai Guan, Kai Keng Ang, Yaozhang Pan, Haihong Zhang
ICASSP2
2012 Cross-Subject Classification of Speaking Modes Using fNIRS
Christian Herff, Dominic Heger, Felix Putze, Cuntai Guan, Tanja Schultz
ICONIP (2)4
2012 Artifact correction with robust statistics for non-stationary intracranial pressure signal monitoring
Mengling Feng, Liang Yu Loy, Kelvin Sim, Clifton Phua, Cuntai Guan
ICPR6
2012 Iterative clustering and support vectors-based high-confidence query selection for motor imagery EEG signals classification
Huijuan Yang, Cuntai Guan, Kai Keng Ang, Haihong Zhang, Chuanchu Wang
ICPR2
2012 Online ICP forecast for patients with traumatic brain injury
Mengling Feng, Liang Yu Loy, Zhuo Zhang 0001, Cuntai Guan
ICPR5
2012 Extracting and selecting discriminative features from high density NIRS-based BCI for numerical cognition
abstract
Near-Infrared Spectroscopy (NIRS)-based Brain-Computer Interface (BCI) was recently studied for numerical cognition. This study presents a study using high density 348 channels NIRS-based BCI from 8 healthy subjects while solving mental arithmetic problems with two difficulty levels and the rest condition. The existing feature extraction and selection methods on the existing study were presented only for low density 16 channels NIRS-based BCI, and required the specification on the number of features to select to yield desirable performance. This paper presents a method of extracting discriminative features from high density single-trial NIRS data using common average reference spatial filtering and single-trial baseline reference, and a method of automatically selecting a set of discriminative and non-redundant features using the Mutual Information-based Rough Set Reduction (MIRSR) and Supervised Pseudo Self- Evolving Cerebellar (SPSEC) algorithms. The performance of the proposed method is evaluated using 5×5-fold cross-validations on the single-trial NIRS data collected using the support vector machine classifier. The results yielded an overall average accuracy of 71.4% and 91.0% in classifying hard versus easy tasks and hard versus rest tasks respectively using the proposed method, compared to 46.1% and 62.2% respectively using existing methods. The results demonstrated the effectiveness of using the proposed feature extraction and selection method in high density NIRS-based BCI for assessing numerical cognition.
Kai Keng Ang, Juanhong Yu, Cuntai Guan
IJCNN3
2012 Robust EEG channel selection across sessions in brain-computer interface involving stroke patients
abstract
Brain-computer interface (BCI) technology has shown the capability of improving the quality of life for people with severe motor disabilities. To improve the portability and practicability of BCI systems, it is crucial to reduce the number of EEG channels as well as to have a good reliability. However, a relatively neglected issue in the EEG channel selection studies is the robustness of selected channels across sessions. This paper investigates whether the selected channels from first session is also useful for subsequent sessions on other days for a stroke patient. For this purpose, a new robust sparse common spatial pattern (RSCSP) algorithm is proposed for optimal EEG channel selection. Thereafter, the robustness of the proposed algorithm as well as 5 existing channel selection algorithms is investigated across 12 sessions data from 11 stroke patients who performed motor imagery based-BCI rehabilitation. The experimental results show that the proposed RSCSP channel selection algorithm significantly outperforms the other channel selection algorithms, when the 8 channels selected from the first session are evaluated on the 11 subsequent sessions. Moreover, there is no significant difference between the classification results of 8 channels selected by the proposed RSCSP algorithm from the first session and the classification results of 8 optimal channels selected from the same session as the test session.
Mahnaz Arvaneh, Cuntai Guan, Kai Keng Ang, Hiok Chai Quek
IJCNN2
2012 Asymmetric Spatial Pattern for EEG-based emotion detection
abstract
Feature extraction has been a crucial and challenging task for EEG-based BCI applications mainly due to the problems of high-dimensionality and high noise level of EEG signals. In this paper we developed a novel feature extraction algorithm for EEG-based emotion detection problem. The proposed algorithm is derived from viewing EEG signals as the activation/deactivation of sources specific to the brain activities of interest. For binary classification problem, to be more specific, we consider the EEG signals for the two types of brain activities as characterized by the activation/deactivation of two discriminatory sources in the brain, with one source activated and the other one deactivated for one particular type of brain activities. The proposed algorithm, termed Asymmetric Spatial Pattern (ASP), extracts pairs of spatial filters, with each filter corresponding to only one of the two sources. The idea of ASP is neurologically plausible for certain situations. For example, according to the valence hypothesis of emotion, the left hemisphere is more activated in positive emotions and the right hemisphere is more activated in negative emotions. The effectiveness of the proposed algorithm is confirmed by application to real data for two types of EEG-based emotion detection problems: arousal detection (strong v.s. calm), and valence detection (positive v.s. negative). Experimental results on the real data also show that some of the asymmetric spatial patterns by ASP are consistent with the current neurophysiological findings on brain emotion processing.
Cuntai Guan, Kai Keng Ang, Haihong Zhang, Yaozhang Pan
IJCNN2
2012 Dynamically Weighted Classification with Clustering to tackle non-stationarity in Brain computer Interfacing
abstract
This paper addresses an important problem known as EEG non-stationarity in Brain-computer Interfacing. We propose a novel technique called Dynamically Weighted Classification with Clustering (DWCC), which explores hidden states in non-stationary EEG using a modified k-means clustering method by combining cosine distance measure and mutual information criterion. DWCC builds a set of classifiers, one for each pair of clusters from different classes. A dynamically-weighted classifier ensemble network is trained to combine the outputs of the classifiers, where we propose to dynamically assign the weight of a classifier for each test sample based on its distances to the cluster centres associated with the classifier. Experimental results on publicly available BCI Competition IV Dataset 2a yielded a mean accuracy of 81.5% which is statistically significant (t-test p<0.05) compared to the baseline result of 75.9% using a single classifier.
Sidath Ravindra Liyanage, Cuntai Guan, Haihong Zhang, Kai Keng Ang, Jianxin Xu 0001, Tong Heng Lee
IJCNN2
2012 Seizure detection based on spatiotemporal correlation and frequency regularity of scalp EEG
abstract
In this paper, a robust seizure detection system using scalp EEG signal is presented. Two most important and obvious characteristics of seizure EEG, signal variance, and frequency synchronization are carefully chosen as seizure detection indexes. To extract the representation of EEG variance, a spatiotemporal correlation structure is constructed based on space-delay covariance matrices with multi-scale temporal delay. The frequency synchronization of EEG is represented by a regularity index derived from wavelet packet transform. The extracted representations are combined to form a high-dimensional feature vector with redundant information. In order to reduce the redundancy, feature selection is performed using mutual information (MI) based on best individual features. The optimized set of features form a more compact feature vector for each 2-s epoch of multi-channel EEG. Feature vectors are then classified into ictal or interictal class using a linear support vector machine (SVM). To evaluate the proposed seizure detection system, unbiased leave-one-session-out cross-validation using clinical routine EEG from 7 patients are performed in experiments. The proposed method obtains average accuracy of 91.44% and average latency of 6.82 s, which outperforms other 7 commonly used methods. It is also demonstrated that the performance of our method is more robust since the standard deviation of results among patients is smaller than other methods.
Yaozhang Pan, Cuntai Guan, Kai Keng Ang, Koksoon Phua, Huijuan Yang, Shih-Hui Lim
IJCNN2
2012 A modified Wavelet-Common Spatial Pattern method for decoding hand movement directions in brain computer interfaces
abstract
The decoding of hand movement kinematics using non-invasive data acquisition techniques is a recent area of research in Brain Computer Interface (BCI). In this work, we use an Electroencephalography (EEG) based BCI to decode directional information from the brain data collected during an actual hand movement experiment. The objective is to find the discriminative features of movement related potential that can classify any two directions out of the four orthogonal directions in which subject performs right hand movement. The performance using Wavelet-Common Spatial Pattern (W-CSP) algorithm and its variations in terms of spatial regularization is studied and compared. The work further analyzes the involvement of frontal, parietal and motor regions in carrying movement kinematics information with the help of spatial plots given by CSP. The performance variability for different directions in various subjects is another important observation in our results. The work aims to provide a more refined movement control command set for BCIs by developing efficient techniques to decode the direction of movement. © 2012 IEEE.
Neethu Robinson, A. Prasad Vinod 0001, Cuntai Guan, Kai Keng Ang, Keng Peng Tee
IJCNN3
2012 Traffic modeling and identification using a Self-adaptive Fuzzy Inference Network
abstract
Traffic modeling and identification is an important aspect of traffic control today. With an increase in the demands on today's transportation network, an efficient system to model and understand the changes in the network is necessary for policy makers to make timely decisions which affect the overall level of service experienced by commuters. This paper proposes a novel approach to traffic modeling and identification using a Self-adaptive Fuzzy Inference Network (SaFIN). The study is performed on a set of real world traffic data collected along the Pan Island Expressway (PIE) in Singapore. By applying a hybrid fuzzy neural network in the traffic modeling task, SaFIN is able to capitalize on the functionalities of both the fuzzy system and the neural network to (1) provide meaningful and intuitive insights to the traffic data, and (2) demonstrate excellent modeling and identification capabilities for highly nonlinear traffic flow conditions.
Sau Wai Tung, Hiok Chai Quek, Cuntai Guan
IJCNN3
2012 Dynamic initiation and dual-tree complex wavelet feature-based classification of motor imagery of swallow EEG signals
abstract
The use of motor imagery-based brain computer interface has recently been shown to have potential for rehabilitation. This paper proposes a novel scheme to detect motor imagery of swallow from electroencephalography (EEG) signals for dysphagia rehabilitation. The proposed scheme extracts features from the coefficients of dual-tree complex wavelet transform (DT-CWT). A novel sliding window-based peak localization scheme is proposed to dynamically locate the initiation of tongue movement from Electromyography (EMG) signal. Subsequently, effective time segments are extracted from EEG signal for classification based on the detected dynamic initiation location. Comparisons are made between our proposed scheme with that of the three existing approaches. The results based on six healthy subjects show that an increase in averaged accuracy of 9.95% is achieved. Further, an increase in averaged accuracy of 8.02% is resulted comparing our proposed scheme by using and not using the dynamic initiation to extract the time segments. Classification results using EMG data confirm that our results are not due to movements artifacts. Statistical tests with 95% confidence to estimate the accuracy on the respective action at chance level show that five out of six subjects performed above chance level for our proposed dynamic initiation and wavelet feature-based approach.
Huijuan Yang, Cuntai Guan, Kai Keng Ang, Chuanchu Wang, Koksoon Phua, Juanhong Yu
IJCNN2
2012 SoHyFIS-Yager: A self-organizing Yager based Hybrid neural Fuzzy Inference System
Sau Wai Tung, Hiok Chai Quek, Cuntai Guan
Expert Syst. Appl.3
2012 Mutual information-based selection of optimal spatial-temporal patterns for single-trial EEG-based BCIs
Kai Keng Ang, Zhengyang Chin, Haihong Zhang, Cuntai Guan
Pattern Recognit.4
2011 FAPOP: Feature analysis enhanced pseudo outer-product fuzzy rule identification system
abstract
Most existing neural fuzzy systems either overlook the importance of feature analysis; or it is performed as a separate phase prior to the design stage of the systems. This paper proposes a novel neural fuzzy system, named Feature Analysis Enhanced Pseudo Outer-Product Fuzzy Rule Identification System (FAPOP), which integrates its design with feature analysis. The objective is two-folds; namely, (1) to improve the interpretability of the system by identifying features relevant to its computational structure; and (2) to improve the accuracy of the system by identifying features relevant to the application problem. The proposed FAPOP model is subsequently employed in a series of benchmark simulations to demonstrate its efficiency as a neural fuzzy modeling system, and excellent performances have been achieved.
Sau Wai Tung, Hiok Chai Quek, Cuntai Guan
FUZZ-IEEE3
2011 Spatially sparsed Common Spatial Pattern to improve BCI performance
abstract
Common Spatial Pattern (CSP) is widely used in discriminating two classes of EEG in Brain Computer Interface applications. However, the performance of the CSP algorithm is affected by noise and artifacts, and the problem is more pronounced in small training data. To overcome these draw-backs, this paper proposes a new Spatially Sparsed CSP (SS-CSP) algorithm by inducing sparsity in the spatial filters. The proposed algorithm optimizes the spatial filters to emphasize the regions that have high variances between classes, and attenuates the regions with low or irregular variances which can be due to noise or artifacts. The experimental results on 14 subjects from publicly available BCI competition datasets showed that the proposed SSCSP algorithm significantly improved the performance of the subjects with poor CSP accuracy by an average of 11%. The results also showed that the obtained sparse spatial filters are more neurophysilogically relevant.
Mahnaz Arvaneh, Cuntai Guan, Kai Keng Ang, Hiok Chai Quek
ICASSP2
2011 Filter Bank Common Spatial Pattern (FBCSP) algorithm using online adaptive and semi-supervised learning
abstract
The Filter Bank Common Spatial Pattern (FBCSP) algorithm employs multiple spatial filters to automatically select key temporal-spatial discriminative EEG characteristics and the Naïve Bayesian Parzen Window (NBPW) classifier using offline learning in EEG-based Brain-Computer Interfaces (BCI). However, it has yet to address the non-stationarity inherent in the EEG between the initial calibration session and subsequent online sessions. This paper presents the FBCSP that employs the NBPW classifier using online adaptive learning that augments the training data with available labeled data during online sessions. However, employing semi-supervised learning that simply augments the training data with available data using predicted labels can be detrimental to the classification accuracy. Hence, this paper presents the FBCSP using online semi-supervised learning that augments the training data with available data that matches the probabilistic model captured by the NBPW classifier using predicted labels. The performances of FBCSP using online adaptive and semi-supervised learning are evaluated on the BCI Competition IV datasets IIa and IIb and compared to the FBCSP using offline learning. The results showed that the FBCSP using online semi-supervised learning yielded relatively better session-to-session classification results compared against the FBCSP using offline learning. The FBCSP using online adaptive learning on true labels yielded the best results in both datasets, but the FBCSP using online semi-supervised learning on predicted labels is more practical in BCI applications where the true labels are not available.
Kai Keng Ang, Zhengyang Chin, Haihong Zhang, Cuntai Guan
IJCNN4
2011 Filter Bank Feature Combination (FBFC) approach for brain-computer interface
abstract
The Filter Bank Common Spatial Pattern (FBCSP) algorithm constructs and selects subject-specific discriminative CSP features from a filter bank of spatial-temporal filters in a motor imagery brain-computer interface (MI-BCI). However, information from other types of features could be extracted and combined with CSP features to enhance the classification performance. Hence this paper proposes a Filter Bank Feature Combination (FBFC) approach and investigates the use of CSP and Phase Lock Value (PLV) features, where the latter measures the phase synchronization between the EEG electrodes. The performance of the FBFC using CSP and PLV features is evaluated on four-class motor imageries from the publicly available BCI Competition IV Dataset IIa. The experimental results showed that the proposed FBFC using CSP and PLV features yielded a significant improvement in cross-validation accuracies on the training data (p=0.008) and better session-to-session transfer accuracies to the evaluation data compared to the use of CSP features using the FBCSP algorithm. This motivates the research of FBFC using a battery of other features that could possibly benefit EEG-based BCIs and multi-modal BCI systems.
Zhengyang Chin, Kai Keng Ang, Cuntai Guan, Chuanchu Wang, Haihong Zhang
IJCNN3
2011 A linear discriminant analysis method based on mutual information maximization
Haihong Zhang, Cuntai Guan, Yuanqing Li 0001
Pattern Recognit.2
2011 SaFIN: A Self-Adaptive Fuzzy Inference Network
abstract
There are generally two approaches to the design of a neural fuzzy system: 1) design by human experts, and 2) design through a self-organization of the numerical training data. While the former approach is highly subjective, the latter is commonly plagued by one or more of the following major problems: 1) an inconsistent rulebase; 2) the need for prior knowledge such as the number of clusters to be computed; 3) heuristically designed knowledge acquisition methodologies; and 4) the stability-plasticity tradeoff of the system. This paper presents a novel self-organizing neural fuzzy system, named Self-Adaptive Fuzzy Inference Network (SaFIN), to address the aforementioned deficiencies. The proposed SaFIN model employs a new clustering technique referred to as categorical learning-induced partitioning (CLIP), which draws inspiration from the behavioral category learning process demonstrated by humans. By employing the one-pass CLIP, SaFIN is able to incorporate new clusters in each input-output dimension when the existing clusters are not able to give a satisfactory representation of the incoming training data. This not only avoids the need for prior knowledge regarding the number of clusters needed for each input-output dimension, but also allows SaFIN the flexibility to incorporate new knowledge with old knowledge in the system. In addition, the self-automated rule formation mechanism proposed within SaFIN ensures that it obtains a consistent resultant rulebase. Subsequently, the proposed SaFIN model is employed in a series of benchmark simulations to demonstrate its efficiency as a self-organizing neural fuzzy system, and excellent performances have been achieved.
Sau Wai Tung, Hiok Chai Quek, Cuntai Guan
IEEE Trans. Neural Networks3
2011 Optimum Spatio-Spectral Filtering Network for Brain-Computer Interface
abstract
This paper proposes a feature extraction method for motor imagery brain-computer interface (BCI) using electroencephalogram. We consider the primary neurophysiologic phenomenon of motor imagery, termed event-related desynchronization, and formulate the learning task for feature extraction as maximizing the mutual information between the spatio-spectral filtering parameters and the class labels. After introducing a nonparametric estimate of mutual information, a gradient-based learning algorithm is devised to efficiently optimize the spatial filters in conjunction with a band-pass filter. The proposed method is compared with two existing methods on real data: a BCI Competition IV dataset as well as our data collected from seven human subjects. The results indicate the superior performance of the method for motor imagery classification, as it produced higher classification accuracy with statistical significance ( ≥ 95% confidence level) in most cases.
Haihong Zhang, Zhengyang Chin, Kai Keng Ang, Cuntai Guan, Chuanchu Wang
IEEE Trans. Neural Networks4
2010 Learning from other subjects helps reducing Brain-Computer Interface calibration time
abstract
A major limitation of Brain-Computer Interfaces (BCI) is their long calibration time, as much data from the user must be collected in order to tune the BCI for this target user. In this paper, we propose a new method to reduce this calibration time by using data from other subjects. More precisely, we propose an algorithm to regularize the Common Spatial Patterns (CSP) and Linear Discriminant Analysis (LDA) algorithms based on the data from a subset of automatically selected subjects. An evaluation of our approach showed that our method significantly outperformed the standard BCI design especially when the amount of data from the target user is small. Thus, our approach helps in reducing the amount of data needed to achieve a given performance level.
Fabien Lotte, Cuntai Guan
ICASSP2
2010 Towards optimum linear transformation under zero-mean Gaussian mixtures for detection of motor imagery EEG
abstract
Optimum linear transformation under mixture of zero-mean Gaussian conditions is an intriguing problem, especially in learning discriminative spatial components in motor imagery EEG for building brain computer interfaces. However, it is not well addressed in the past. In this paper, we study optimum linear transformation under mixture of zero-mean Gaussian. In particular, we formulate optimum transformation as a Bhattacharyya error bound minimization problem, and derive a numerical solution to estimate the bound from training samples. Based on the solution, we develop an algorithm for selecting optimum linear transformation. The proposed method is evaluated, in comparison with the state-of-the-art methods, using a publicly available data set of motor imagery EEG. The results attest to the superiority of the method for detecting motor imagery.
Haihong Zhang, Cuntai Guan, Chuanchu Wang
ICASSP2
2010 A Brain-Computer Interface for Mental Arithmetic Task from Single-Trial Near-Infrared Spectroscopy Brain Signals
abstract
Near-infrared spectroscopy (NIRS) enables non-invasive recording of cortical hemoglobin oxygenation in human subjects through the intact skull using light in the near-infrared range to determine. Recently, NIRS-based brain-computer interfaces are introduced for discriminating left and right-hand motor imagery. A neuroimaging study has also revealed event-related hemodynamic responses associated with the performance of mental arithmetic tasks. This paper proposes a novel BCI for detecting changes resulting from increases in the magnitude of operands used in a mental arithmetic task, using data from single-trial NIRS brain signals. We measured hemoglobin responses from 20 healthy subjects as they solved mental arithmetic problems with three difficulty levels. Accuracy in recognizing one difficulty level from another is then presented using 5×5-fold cross-validations on the data collected. The results yielded an overall average accuracy of 71.2%, thus demonstrating potential in the proposed NIRS-based BCI in recognizing difficulty of problems encountered by mental arithmetic problem solvers.
Kai Keng Ang, Cuntai Guan, Kerry Lee 0002, Jie Qi Lee, Shoko Nioka, Britton Chance
ICPR2
2010 Spatially Regularized Common Spatial Patterns for EEG Classification
abstract
In this paper, we propose a new algorithm for Brain-Computer Interface (BCI): Spatially Regularized Common Spatial Patterns (SRCSP). SRCSP is an extension of the famous CSP algorithm which includes spatial a priori in the learning process, by adding a regularization term which penalizes spatially non smooth filters. We compared SRCSP and CSP algorithms on data of 14 subjects from BCI competitions. Results suggested that SRCSP can improve performances, around 10% more in classification accuracy, for subjects with poor CSP performances. They also suggested that SRCSP leads to more physiologically relevant filters than CSP.
Fabien Lotte, Cuntai Guan
ICPR2
2010 A Covariate Shift Minimisation Method to Alleviate Non-stationarity Effects for an Adaptive Brain-Computer Interface
abstract
The non-stationary nature of the electroencephalogram (EEG) poses a major challenge for the successful operation of a brain-computer interface (BCI) when deployed over multiple sessions. The changes between the early training measurements and the proceeding multiple sessions can originate as a result of alterations in the subject's brain process, new cortical activities, change of recording conditions and/or change of operation strategies by the subject. These differences and alterations over multiple sessions cause deterioration in BCI system performance if periodic or continuous adaptation to the signal processing is not carried out. In this work, the covariate shift is analyzed over multiple sessions to determine the non-stationarity effects and an unsupervised adaptation approach is employed to account for the degrading effects this might have on performance. To improve the system's online performance, we propose a covariate shift minimization (CSM) method, which takes into account the distribution shift in the feature set domain to reduce the feature set overlap and unbalance for different classes. The analysis and the results demonstrate the importance of CSM, as this method not only improves the accuracy of the system, but also reduces the classification unbalance for different classes by a significant amount.
Abdul Rehman Satti, Cuntai Guan, Damien Coyle, Girijesh Prasad
ICPR2
2010 An Information Theoretic Linear Discriminant Analysis Method
abstract
We propose a novel linear discriminant analysis method and demonstrate its superiority over existing linear methods. Based on information theory, we introduce a non-parametric estimate of mutual information with variable kernel bandwidth. Furthermore, we derive a gradient-based optimization algorithm for learning the optimal linear reduction vectors which maximizes the mutual information estimate. We evaluate the proposed method by running cross-validation on 2 data sets from the UCI repository, together with linear and nonlinear SVMs as classifiers. The result attests to the superority of the method over conventional LDA and its variant, aPAC.
Haihong Zhang, Cuntai Guan, Kai Keng Ang
ICPR2
2010 Application of rough set-based neuro-fuzzy system in NIRS-based BCI for assessing numerical cognition in classroom
abstract
Near-infrared spectroscopy (NIRS) studies have revealed that performing mental arithmetic tasks have associated event-related hemodynamic responses that are detectable. Thus NIRS-based Brain Computer Interface (BCI) has the potential for investigating how to best teach mathematics in a classroom setting. This paper presents a novel computational intelligent method of applying rough set-based neuro-fuzzy system (RNFS) in NIRS-based BCI for assessing numerical cognition. A study is performed on 20 healthy subjects to measure 32 channels of hemoglobin responses in performing three difficulty levels of mental arithmetic. The accuracy is then presented using 5×5-fold cross-validations on the data collected. The results of applying RNFS and its Mutual Information-based Rough Set Reduction (MIRSR) for feature selection is then compared against the Naïve Bayesian Parzen Window classifier and other MI-based feature selection algorithms. The results of applying RNFS yielded significantly better accuracy of 75.7% compared to the other methods, thus demonstrating the potential of RNFS in NIRS-based BCI for assessing numerical cognition.
Kai Keng Ang, Cuntai Guan, Kerry Lee 0002, Jie Qi Lee, Shoko Nioka, Britton Chance
IJCNN2
2010 EEG signal separation for multi-class motor imagery using common spatial patterns based on Joint Approximate Diagonalization
abstract
The design of multiclass BCI is a very challenging task because of the need to extract complex spatial and temporal patterns from noisy multidimensional time series generated from EEG measurements. This paper proposes a Multiclass Common Spatial Pattern (MCSP) based on Joint Approximate Diagonalization (JAD) for multiclass BCIs. The proposed method based on fast Frobenius diagonalization (FFDIAG) is compared with another method based on Jacobi angles on the BCI competition IV dataset 2a. The classification accuracies obtained from 10×10-fold cross-validations on the training dataset are compared using K-Nearest Neighbor, Classification Trees and Support Vector Machine classifiers. The proposed MCSP based on FFDIAG yields an averaged accuracy of 53.6% compared to 32.8% given by the method based on Jacobi angles and 27.8% of the one versus rest CSP methods.
Sidath Ravindra Liyanage, Jianxin Xu 0001, Cuntai Guan, Kai Keng Ang, Tong Heng Lee
IJCNN3
2010 A Study on the impact of spectral variability in brain-computer interface
abstract
The performance of a Brain-Computer Interface (BCI) depends on reliable feature extraction and accurate classification. Motor imagery has been successfully used in BCI for communication and control. During motor imagery, for EEG based BCI, it was known that the discriminative frequency bands are subject-specific. Moreover, such discriminative frequency bands for each subject might vary from time to time. In this paper, we investigate the variability of discriminative spectral ranges and its impact on classification accuracy. It is found that for each subject, his discriminative frequency bands changes significantly from session to session, but keeps almost stable within a session. We then propose a method to adaptively update the discriminative frequency bands using Time-Frequency fisher ratio. From the experimental analysis, it is found that we can reduce the average error rate by 11.50% compared to the case where fixed discriminative frequency bands obtained from calibration session are used.
Kavitha P. Thomas, Cuntai Guan, Chiew Tong Lau, A. Prasad Vinod 0001
ISCAS2
2010 An Evolving Type-2 Neural Fuzzy Inference System
Sau Wai Tung, Hiok Chai Quek, Cuntai Guan
PRICAI3
2009 T2-HyFIS-yager: Type 2 hybrid neural fuzzy inference system realizing yager inference
abstract
The hybrid neural fuzzy inference system (Hy-FIS) is a five layers adaptive neural fuzzy inference system, based on the compositional rule of inference (CRI) scheme, for building and optimizing fuzzy models. To provide the HyFIS architecture with a firmer and more intuitive logical framework that emulates the human reasoning and decision-making mechanism, the fuzzy Yager inference scheme, together with the self-organizing Gaussian discrete incremental clustering (gDIC) technique, were integrated into the HyFIS network to produce the HyFIS-Yager-gDIC . This paper presents T2-HyFIS-Yager, a type-2 hybrid neural fuzzy inference system realizing Yager inference, for learning and reasoning with noise corrupted data. The proposed T2-HyFIS-Yager is used to perform time-series forecasting where a non-stationary time-series is corrupted by additive white noise of known and unknown SNR to demonstrate its superiority as an effective neuro-fuzzy modeling technique.
Sau Wai Tung, Hiok Chai Quek, Cuntai Guan
FUZZ-IEEE3
2009 Voxel selection in fMRI data analysis: A sparse representation method
abstract
This paper proposes an iterative sparse representation-based algorithmfor voxel selection in functionalmagnetic resonance imaging (fMRI) data. The output of the algorithm is a sparse weight vector, of which the magnitude of each entry represents the significance of its corresponding voxel with respect to mental tasks or stimulus. To demonstrate the validity of our algorithm and illustrate its application, we apply this algorithm to the Pittsburgh Brain Activity Interpretation Competition (PBAIC) 2007 fMRI data set for selecting the voxels which are the most relevant to the tasks of the subjects. Compared with three baseline methods, general linear model (GLM)-based statistical parametric mapping (SPM), correlation method and mutual information method, our method shows satisfactory performance for voxel selection.
Yuanqing Li 0001, Zhu Liang Yu, Praneeth Namburi, Cuntai Guan
ICASSP4
2009 Learning EEG-based Spectral-spatial Patterns for Attention Level Measurement
abstract
In our every day life, our brain is constantly processing information and paying attention, reacting accordingly, to all sorts of sensory inputs (auditory, visual, etc.). In some cases, there is a need to accurately measure a person's level of attention to monitor a sportsman performance, to detect Attention Deficit Hyperactivity Disorder (ADHD) in children, to evaluate the effectiveness of neuro-feedback treatment, etc.
Brahim Hamadicharef, Haihong Zhang, Cuntai Guan, Chuanchu Wang, Koksoon Phua, Keng Peng Tee, Kai Keng Ang
ISCAS3
2009 Discriminative FilterBank Selection and EEG Information Fusion for Brain Computer Interface
abstract
Brain computer interface (BCI) provides a direct communication pathway between a human and an external device. In this paper, we propose a new dasiadiscriminative filterbank common spatial pattern (DFBCSP)psila algorithm to select the subject-specific filters automatically during training for a motor imagery based BCI. The subject-specific filters are selected using the fisher ratio values of filtered electroencephalogram (EEG) signal. The channel dasiaC3psila alone could give sufficient information to select the discriminative filterbank for the proposed system. We have also explored the possibility of boosting the system performance by including dynamic temporal features. Fusion of static and dynamic features in the proposed DFBCSP frame work gave an average test accuracy of 92.44%, which is significantly better than conventional filterbank based common spatial pattern algorithms.
Kavitha P. Thomas, Cuntai Guan, Chiew Tong Lau, A. Prasad Vinod 0001
ISCAS2
2008 HyFIS-Yager-gDIC: A Self-organizing Hybrid Neural Fuzzy Inference System Realizing Yager Inference
Sau Wai Tung, Hiok Chai Quek, Cuntai Guan
ICONIP (1)3
2008 Subject-independent brain computer interface through boosting
abstract
This paper presents a subject-independent EEG (Electroencephalogram) classification technique and its application to a P300-based word speller. Due to EEG variations across subjects, a user calibration procedure is usually required to build a subject-specific classification model (SSCM). We remove the user calibration through the boosting of a committee of weak classifiers learned from EEG of a pool of subjects. In particular, we ensemble the weak classifiers based on their confidence that is evaluated according to the classification consistency. Experiments over ten subjects show that the proposed technique greatly outperforms the supervised classification models, hence making P300-based BCIs more convenient for practical uses.
Shijian Lu, Cuntai Guan, Haihong Zhang
ICPR2
2008 Filter Bank Common Spatial Pattern (FBCSP) in Brain-Computer Interface
abstract
In motor imagery-based Brain Computer Interfaces (BCI), discriminative patterns can be extracted from the electroencephalogram (EEG) using the Common Spatial Pattern (CSP) algorithm. However, the performance of this spatial filter depends on the operational frequency band of the EEG. Thus, setting a broad frequency range, or manually selecting a subject-specific frequency range, are commonly used with the CSP algorithm. To address this problem, this paper proposes a novel Filter Bank Common Spatial Pattern (FBCSP) to perform autonomous selection of key temporal-spatial discriminative EEG characteristics. After the EEG measurements have been bandpass-filtered into multiple frequency bands, CSP features are extracted from each of these bands. A feature selection algorithm is then used to automatically select discriminative pairs of frequency bands and corresponding CSP features. A classification algorithm is subsequently used to classify the CSP features. A study is conducted to assess the performance of a selection of feature selection and classification algorithms for use with the FBCSP. Extensive experimental results are presented on a publicly available dataset as well as data collected from healthy subjects and unilaterally paralyzed stroke patients. The results show that FBCSP, using a particular combination feature selection and classification algorithm, yields relatively higher cross-validation accuracies compared to prevailing approaches.
Kai Keng Ang, Zhengyang Chin, Haihong Zhang, Cuntai Guan
IJCNN4
2008 An EEG-based BCI system for 2D cursor control
abstract
In this paper, an electroencephalogram (EEG)-based brain computer interface (BCI) is proposed for two dimensional cursor control. The horizontal and vertical movements of the cursor are controlled by mu/beta rhythm and P300 potential respectively. The main advantages of this system are: (i) two almost independent control signals are produced simultaneously; (ii) the cursor can be moved from a random position to another random position in a screen. These advantages have been demonstrated in our experiment and data analysis.
Yuanqing Li 0001, Chuanchu Wang, Haihong Zhang, Cuntai Guan
IJCNN4
2008 Learning adaptive subject-independent P300 models for EEG-based brain-computer interfaces
abstract
This paper proposes an approach to learn subject-independent P300 models for EEG-based brain-computer interfaces. The P300 models are first learned using a pool of existing subjects and Fisher linear discriminant, and then autonomously adapted to the unlabeled data of a new subject using an unsupervised machine learning technique. In data analysis, we apply this technique to a set of EEG data of 10 subjects performing word spelling in an oddball paradigm. The results are very positive: the adapted models with unlabeled data yield virtually the same classification accuracy as the conventional methods with labeled data. Therefore, it proves the feasibility of P300-based BCIs which can be applied directly to a new subject without training sessions.
Shijian Lu, Cuntai Guan, Haihong Zhang
IJCNN2
2008 Joint feature re-extraction and classification using an iterative semi-supervised support vector machine algorithm
Yuanqing Li 0001, Cuntai Guan
Mach. Learn.2
2008 A self-training semi-supervised SVM algorithm and its application in an EEG-based brain computer interface speller system
Yuanqing Li 0001, Cuntai Guan, Huiqi Li, Zhengyang Chin
Pattern Recognit. Lett.2
2008 Equivalence Probability and Sparsity of Two Sparse Solutions in Sparse Representation
abstract
This paper discusses the estimation and numerical calculation of the probability that the 0-norm and 1-norm solutions of underdetermined linear equations are equivalent in the case of sparse representation. First, we define the sparsity degree of a signal. Two equivalence probability estimates are obtained when the entries of the 0-norm solution have different sparsity degrees. One is for the case in which the basis matrix is given or estimated, and the other is for the case in which the basis matrix is random. However, the computational burden to calculate these probabilities increases exponentially as the number of columns of the basis matrix increases. This computational complexity problem can be avoided through a sampling method. Next, we analyze the sparsity degree of mixtures and establish the relationship between the equivalence probability and the sparsity degree of the mixtures. This relationship can be used to analyze the performance of blind source separation (BSS). Furthermore, we extend the equivalence probability estimates to the small noise case. Finally, we illustrate how to use these theoretical results to guarantee a satisfactory performance in underdetermined BSS.
Yuanqing Li 0001, Andrzej Cichocki, Shun-ichi Amari, Shengli Xie 0001, Cuntai Guan
IEEE Trans. Neural Networks5
2007 Feature Selection Based on Fisher Ratio and Mutual Information Analyses for Robust Brain Computer Interface
abstract
This paper proposes a novel feature selection method based on two-stage analysis of Fisher ratio and mutual information for robust brain computer interface. This method decomposes multichannel brain signals into subbands. The spatial filtering and feature extraction is then processed in each subband. The two-stage analysis of Fisher ratio and mutual information is carried out in the feature domain to reject the noisy feature indexes and select the most informative combination from the remaining. In the approach, we develop two practical solutions, avoiding the difficulties of using high dimensional mutual information in the application, that are the feature indexes clustering using cross mutual information and the latter estimation based on conditional empirical PDF. We test the proposed feature selection method on two BCI data sets and the results are at least comparable to the best results in the literature. The main advantage of proposed method is that the method is free from any time-consuming parameter tweaking and therefore suitable for the BCI system design.
Tran Huy Dat, Cuntai Guan
ICASSP (1)2
2007 A Self-Training Semi-Supervised Support Vector Machine Algorithm and its Applications in Brain Computer Interface
abstract
In this paper, we analyze the convergence of an iterative self-training semi-supervised support vector machine (SVM) algorithm, which is designed for classification in small training data case. This algorithm converges fast and has low computational burden. Its effectiveness is also demonstrated by our data analysis results. Furthermore, we illustrate that this algorithm can be used to significantly reduce training effort and improve adaptability of a brain computer interface (BCI) system, a P300-based speller.
Yuanqing Li 0001, Huiqi Li, Cuntai Guan, Zhengyang Chin
ICASSP (1)3
2007 Brainy Communicator
abstract
The "Brainy Communicator" is a novel state-of-the-art Brain-Computer Interface (BCI) system developed by the Institute for Infocomm Research, Singapore. It allows users to interact with the environment by using just the brain signals. The primary objective of this work is to enhance the quality of life for the people with severe disabilities, by helping them regain some control and communication abilities.
Haihong Zhang, Cuntai Guan, Chuanchu Wang
ICME2
2006 Signal processing for brain-computer interface: enhance feature extraction and classification
abstract
In this paper we present a new scheme for brain signal processing and classification for electroencephalogram based brain-computer interfaces, by emphasizing the extraction of space-time-frequency feature as well as the combination of classifiers. In particular, we use wavelet packets as a time-frequency analysis tool and employ sparse component analysis to recover source components in the brain signals. We subsequently apply multi-class common spatial pattern filters to the signals and thus obtain important space-time-frequency features for discrimination. Furthermore, a Bayesian method is developed to boost the system, by combining multiple support vector machines in a probabilistic way. We have tested the proposed scheme on real multi-class motor imagery signals, and its efficacy has been demonstrated.
Haihong Zhang, Cuntai Guan, Yuanqing Li 0001
ISCAS2
2006 An Extended EM Algorithm for Joint Feature Extraction and Classification in Brain-Computer Interfaces
abstract
For many electroencephalogram (EEG)-based brain-computer interfaces (BCIs), a tedious and time-consuming training process is needed to set parameters. In BCI Competition 2005, reducing the training process was explicitly proposed as a task. Furthermore, an effective BCI system needs to be adaptive to dynamic variations of brain signals; that is, its parameters need to be adjusted online. In this article, we introduce an extended expectation maximization (EM) algorithm, where the extraction and classification of common spatial pattern (CSP) features are performed jointly and iteratively. In each iteration, the training data set is updated using all or part of the test data and the labels predicted in the previous iteration. Based on the updated training data set, the CSP features are reextracted and classified using a standard EM algorithm. Since the training data set is updated frequently, the initial training data set can be small (semi-supervised case) or null (unsupervised case). During the above iterations, the parameters of the Bayes classifier and the CSP transformation matrix are also updated concurrently. In online situations, we can still run the training process to adjust the system parameters using unlabeled data while a subject is using the BCI system. The effectiveness of the algorithm depends on the robustness of CSP feature to noise and iteration convergence, which are discussed in this article. Our proposed approach has been applied to data set IVa of BCI Competition 2005. The data analysis results show that we can obtain satisfying prediction accuracy using our algorithm in the semisupervised and unsupervised cases. The convergence of the algorithm and robustness of CSP feature are also demonstrated in our data analysis.
Yuanqing Li 0001, Cuntai Guan
Neural Comput.2
2006 Probability Estimation for Recoverability Analysis of Blind Source Separation Based on Sparse Representation
abstract
An important application of sparse representation is underdetermined blind source separation (BSS), where the number of sources is greater than the number of observations. Within the stochastic framework, this paper discusses recoverability of underdetermined BSS based on a two-stage sparse representation approach. The two-stage approach is effective when the source matrix is sufficiently sparse. The first stage of the two-stage approach is to estimate the mixing matrix, and the second is to estimate the source matrix by minimizing the 1-norms of the source vectors subject to some constraints. After estimating the mixing matrix and fixing the number of nonzero entries of a source vector, we estimate the recoverability probability (i.e., the probability that the source vector can be recovered). A general case is then considered where the number of nonzero entries of the source vector is fixed and the mixing matrix is drawn from a specific probability distribution. The corresponding probability estimate on recoverability is also obtained. Based on this result, we further estimate the recoverability probability when the sources are also drawn from a distribution (e.g., Laplacian distribution). These probability estimates not only reflect the relationship between the recoverability and sparseness of sources, but also indicate the overall performance and confidence of the two-stage sparse representation approach for solving BSS problems. Several simulation results have demonstrated the validity of the probability estimation approach.
Yuanqing Li 0001, Shun-ichi Amari, Andrzej Cichocki, Cuntai Guan
IEEE Trans. Inf. Theory4
2003 On unit analysis for Cantonese corpus-based TTS
Thomas Choy, Minghui Dong, Cuntai Guan, Haizhou Li 0001
INTERSPEECH4
2002 Multilingual speech recognition with language identification
Bin Ma 0001, Cuntai Guan, Haizhou Li 0001, Chin-Hui Lee 0001
INTERSPEECH2
1997 A space transformation approach for robust speech recognition in noisy environments
Cuntai Guan, Shu Hung Leung, Wing Hong Lau
EUROSPEECH1
1993 Direct modulation on LPC coefficients with application to speech enhancement and improving the performance of speech recognition in noise
Cuntai Guan, Yong Bin Chen, Boxiu Wu
ICASSP (2)1