VLDB 2026 Research / reviewers in the wild / expert
Bao-Liang Lu
dblp:09/3116
· DBLP profile ↗
235ranked-venue papers
9as first author
65since 2021 · last 2026
0000-0001-8359-0058ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 181 · 8 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 40 · 21 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 1 first-author · 14 since 2021Databases, data management, data science and information retrieval · 5Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MindCross: Fast New Subject Adaptation with Limited Data for Cross-subject Video Reconstruction from Brain SignalsabstractBrain decoding aims to reconstruct video from brain signals. Existing brain decoding frameworks are primarily built on a subject-dependent paradigm, which requires large amounts of brain data for each subject. However, the expensive cost of collecting brain-video data causes severe data scarcity for brain decoding. Although some cross-subject methods being introduced, they often exhibit an excessive preoccupation with subject-invariant information while neglecting subject-specific information, resulting in slow fine-tune-based adaptation strategy. To achieve fast and data-efficient new subject adaptation, we propose **MindCross**, a novel cross-subject brain decoding framework. MindCross's *N* specific encoders and one shared encoder are designed to extract subject-specific and subject-invariant information, respectively. Additionally, a Top-*K* collaboration module is adopted to enhance new subject decoding with the knowledge learned from previous subjects' encoders. Extensive experiments on fMRI/EEG-to-video benchmarks demonstrate MindCross's efficacy and efficiency of cross-subject decoding and new subject adaptation using only one model. Code of our framework will be released upon publication. Xuan-Hao Liu, Yan-Kai Liu, Bao-Liang Lu, Wei-Long Zheng |
AAAI | 4 |
| 2026 | A Multimodal EEG-Eye Movement Model for Automatic Depression DetectionabstractDepression is a prevalent mental health disorder characterized by persistent sadness and a diminished interest in daily activities, early detection of depression facilitates timely intervention, mitigating its adverse effects. Electroencephalography (EEG) signals and eye movements are emerging as promising biomarkers for depression detection due to their non-invasive nature and cost-effectiveness. Nevertheless, existing studies suffer from methodological constraints, including low specificity, insufficient sample sizes, limited generalizability, and difficulties in large-scale replication, which collectively undermine their clinical utility. To address these challenges, we collected a large-scale depression dataset comprising EEG and eye movements from 1,060 individuals diagnosed with depression and 1,308 healthy controls. To efficiently leverage multimodal data for automatic depression detection, we propose the EEG-Eye Movements Model (E2Mo). E2Mo employs modality-specific encoders to extract discriminative multi-view features from each modality and incorporates a mixture-of-modality-experts architecture with multi pretraining tasks to achieve efficient and robust modality alignment and fusion. Our approach achieves a 70.06% balanced accuracy by leveraging multi-modal data, demonstrating the effectiveness of integrating EEG signals and eye movements for automatic depression detection. Hao-Long Yin, Ren-Jie Dai, Wei-Long Zheng, Qinyu Lv, Zhenghui Yi, Bao-Liang Lu |
AAAI | 7 |
| 2026 | Online learning with a hedge-based random vector functional link network using multi-forgetting factors
Yimin Wen, Chuangquan Chen, Bao-Liang Lu |
Neurocomputing | 5 |
| 2026 | Gram: A Large General EEG Model for Raw Data Classification and RestorationabstractDrawing insights from Large Language Models, researchers have developed several Large Electroencephalogram (EEG) models (LEMs) to learn a generalized representation adaptable to various tasks. However, such LEMs are scarce and neglecting the potential in data restoration tasks. Meanwhile, how to efficiently integrate temporal view and spectral view of EEG data has always been a focal point. In this paper, we propose Gram, a large general EEG model for raw EEG data classification and restoration tasks. Gram consists of two stages. 1) The initial stage quantizes raw EEG patches into base classes rich in temporal information. 2) The second stage features a multi-view layer-fusion masked autoencoder that exploits EEG's complex Temporal and Spectral views through dual training objectives: a spectral mimic target after layer-fusion encoder for visible patches and a base-class classification target after decoder for masked patches. Pretrained on 7000 hours of EEG data, Gram achieves state-of-the-art performance on four cross-subject classification tasks including motor imagery as well as event, emotion, and sleep stage classification. For EEG data restoration, our model significantly improves classification performance by repairing corrupted data in comparison to using noisy data. Ziyi Li 0003, Wei-Long Zheng, Jiwen Xu, Bao-Liang Lu |
IEEE Trans. Affect. Comput. | 5 |
| 2026 | SEED-OLF: A Novel EEG Dataset With Olfactory Stimulation for Emotion Recognition
Wei-Bang Jiang, Wei-Long Zheng, Bao-Liang Lu |
IEEE Trans. Affect. Comput. | 4 |
| 2026 | Prediction Consistency and Confidence-Based Proxy Domain Construction for Privacy-Preserving in Cross-Subject EEG ClassificationabstractDomainadaptation has proven effective for suppressing the inter-subject variability problem in cross-subject EEG classification tasks in which labeled data is available for source subjects while only unlabeled data is provided for target subjects. Existing domain adaptation methods typically reduced the distribution discrepancy between source and target domains by directly utilizing source domain samples or features. To safeguard the privacy of source domain data, we propose to construct a Proxy Domain by simultaneously considering the prediction Consistency and Confidence (PDCC) of locally trained source models on target EEG samples, serving as the substitute to the source domain. The framework commences with the augmentation and alignment of the source domain data to enhance feature generalizability, after which source models are trained independently on each source subject's data in a decentralized manner. Knowledge transfer from source to target domains is achieved exclusively through accessing to the source domain model, enabling the PDCC-based proxy domain construction that encapsulates the source knowledge. Finally, domain adaptation is performed using the proxy domain and target domain. As a result, PDCC eliminates the need to access source domain data while effectively leveraging source knowledge. Experimental results on four benchmark EEG datasets demonstrate that PDCC consistently outperforms eleven existing methods, including several advanced transfer learning and source-free methods. Especially, the effectiveness of the proxy domain is extensively investigated. Yong Peng 0001, Jiangchuan Liu, Honggang Liu, Natasha M. J. Padfield, Wanzeng Kong, Bao-Liang Lu, Andrzej Cichocki |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Multi-to-Single: Reducing Multimodal Dependency in Emotion Recognition Through Contrastive LearningabstractMultimodal emotion recognition is a crucial research area in the field of affective brain-computer interfaces. However, in practical applications, it is often challenging to obtain all modalities simultaneously. To deal with this problem, researchers focus on using cross-modal methods to learn multimodal representations with fewer modalities. However, due to the significant differences in the distribution of different modalities, it is challenging to enable any modality to fully learn multimodal features. To address this limitation, we propose a Multi-to-Single (M2S) emotion recognition model, leveraging contrastive learning and incorporating two innovative modules: 1) a spatial and temporal-sparse (STS) attention mechanism that enhances the encoders' ability to extract features from data; 2) a novel Multi-to-Multi Contrastive Predictive Coding (M2M CPC) that learns and fuses features across different modalities. In the final testing, we only use a single modality for emotion recognition, reducing the dependence on multimodal data. Extensive experiments on five public multimodal emotion datasets demonstrate that our model achieves the state-of-the-art performance in the cross-modal tasks and maintains multimodal performance using only a single modality. Yan-Kai Liu, Jinyu Cai, Bao-Liang Lu, Wei-Long Zheng |
AAAI | 3 |
| 2025 | Decoding Emotions from Missing-Channel EEG Signals via Uncertainty-Aware Modeling with Neural ProcessesabstractElectroencephalography (EEG) signals have been widely utilized in affective brain-computer interface (BCI) applications. However, EEG signals are inherently susceptible to both external environmental noise and internal subject-specific variability, resulting in significant uncertainty. Moreover, in real-world scenarios, excessive external interference may severely degrade the quality of EEG recordings, leading to missing or corrupted channels. Therefore, developing robust emotion decoding methods that can handle incomplete EEG data is critical for enhancing the generalizability and reliability of affective BCIs. In this paper, we propose Fewer-Channel EEG Neural Processes (FC-EEGNP), a novel neural process-based framework designed to learn the uncertain relationships between EEG signals and brain regions, enabling emotion recognition under missing-channel conditions. FC-EEGNP reconstructs high spatial resolution EEG representations from a limited subset of available channels, preserving critical emotional information despite data sparsity. We conduct extensive experiments on three publicly available EEG datasets to evaluate the performance of FC-EEGNP. Results demonstrate that our model consistently achieves state-of-the-art performance across multiple emotion recognition settings, including subject-dependent and cross-subject tasks. Yan-Kai Liu, Xuan-Hao Liu, Yi-Dong Zhao, Bao-Liang Lu, Wei-Long Zheng |
BIBM | 4 |
| 2025 | Attention-Based Graph Net with Mixture of Experts for Emotion Recognition Under Sleep DeprivationabstractThe impact of sleep deprivation on human emotional states has long been a central topic in psychology and neuroscience. However, the lack of high-quality, randomized stimuli materials for eliciting multimodal emotional responses has limited the development of robust datasets in this area. To address this gap, we collected a multimodal dataset comprising four emotional categories-happy, sad, fear, and neutral-from 20 subjects under three sleep conditions: sleep deprivation, sleep recovery, and normal sleep. Each emotional state was elicited using randomized video stimuli specifically designed to provoke the target emotions. Furthermore, we propose a novel emotion classification framework that integrates a Graph Neural Network (GNN) with a MoE (Mixture-of-Experts) architecture. Experimental results demonstrate that our method achieves SOTA performance. Moreover, in all models evaluated, the accuracy of emotional recognition was significantly lower under sleep deprivation compared to sleep recovery and baseline conditions, particularly for positive emotional states. Shi-Heng Tian, Yan-Kai Liu, Bao-Liang Lu, Wei-Long Zheng |
BIBM | 3 |
| 2025 | Self-supervised EEG Representation Learning based on Temporal Prediction and Spatial Reconstruction for Emotion Recognition
Ren-Jie Dai, Keya Hu, Hao-Long Yin, Bao-Liang Lu, Wei-Long Zheng |
CogSci | 4 |
| 2025 | mixEEG: Enhancing EEG Federated Learning for Cross-subject EEG Classification with Tailored mixup
Xuan-Hao Liu, Bao-Liang Lu, Wei-Long Zheng |
CogSci | 2 |
| 2025 | Gram: A Large-Scale General EEG Model for Raw Data Classification and Restoration TasksabstractDrawing insights from Large Language Models, researchers have developed several large-scale Electroencephalogram (EEG) models (LEMs) to learn a generalized representation adaptable to various tasks. However, such LEMs are scarce and neglecting the potential in reconstruction tasks. Meanwhile, how to efficiently integrate temporal view and spectral view of EEG data has always been a focal point. In this paper, we propose Gram, a large general EEG model for raw EEG data classification and reconstruction tasks. Gram consists of two stages. 1) The initial stage quantizes raw EEG patches into base classes rich in temporal information. 2) The second stage features a multi-view layer-fusion masked autoencoder that exploits EEG’s complex Temporal and Spectral views through dual training objectives: a spectral mimic target after layer-fusion encoder for visible patches and a base-class classification target after decoder for masked patches. Pretrained on 7000 hours of EEG data, Gram achieves SOTA performance on 3 cross-subject classification tasks including event, emotion, and sleep stage classification. For EEG data restoration, our model significantly improves classification performance by repairing corrupted data in comparison to using noisy data. The code and pretrained weights are in https://github.com/iiieeeve/Gram. Ziyi Li 0003, Wei-Long Zheng, Bao-Liang Lu |
ICASSP | 3 |
| 2025 | Multi-Source Multi-Target Domain Similarity Network for Cross-Cultural EEG Emotion RecognitionabstractThe significant variations in emotional patterns across different cultures pose a major challenge for cross-cultural electroencephalogram (EEG) emotion recognition. Moreover, this task must address not only differences in feature distributions among different cultures but also among individuals within the same culture. Therefore, we propose a novel domain adaptation approach, the Multi-Source Multi-Target Domain Similarity Network (MSMTDS), which treats each subject as an individual source domain or target domain. Based on the assessment of similarities across both cross-cultural and intra-cultural dimensions, the model dynamically adjusts the weights assigned to each domain, allowing similar domains to play a dominant role in training while reducing the adverse impact of dissimilar domains. Additionally, given the scarcity of EEG data, MSMTDS fully leverages all available data to maximize performance. Extensive experiments on the SEED EEG emotion datasets from three distinct cultures (China, France, and Germany) demonstrate the effectiveness of our approach, achieving state-of-the-art results. Hanwen Shi, Bao-Liang Lu, Wei-Long Zheng |
ICASSP | 3 |
| 2025 | STAR: A Spatial-Temporal Autoencoder for EEG Restoration in Emotion RecognitionabstractResearch in emotion recognition using electroencephalography (EEG) has advanced rapidly, and affective EEG-based Brain-computer Interface (aBCI) technology is increasingly moving from lab research to real-world application. Nevertheless, EEG signals are inherently delicate and prone to noise and artifacts, especially in real-world environments where data quality often lags behind laboratory standards. This disparity poses substantial challenges for models trained on high-quality datasets. Conventional methods, such as data interpolation or exclusion, limit model efficacy. To overcome these challenges, we introduce the Spatial-Temporal Autoencoder for EEG Restoration (STAR). STAR leverages dynamic channel and temporal masking to mimic real-world signal degradation and incorporates a spatial-temporal alternating attention mechanism to encapsulate intricate spatiotemporal dynamics within EEG data. Our evaluations on three premium emotion recognition EEG datasets reveal that STAR effectively restores signals across varying corruption levels, significantly bolstering the performance of emotion recognition models in suboptimal conditions. Hao-Long Yin, Wei-Long Zheng, Bao-Liang Lu |
ICASSP | 3 |
| 2025 | Multi-Scale Attention-Based Dense Spatial-Temporal Model for Emotion Induction in Response to Olfactory StimuliabstractAffective Brain-Computer Interfaces (aBCIs) have attracted growing attention due to their potential for decoding human emotional states through electroencephalogram (EEG) signals. However, existing deep learning models often struggle to fully capture both the spatial and temporal dependencies in EEG data, resulting in suboptimal performance in emotion classification. To address these challenges, we propose a novel Multi-Scale Attention-Based Dense Spatial-Temporal Model (MSADM). This model leverages temporal attention, multi-scale dense feature extraction, and attention-based feature fusion to effectively capture and enhance the complicated spatial-temporal dependencies within EEG data, thereby improving the representations. Furthermore, we introduce a new olfactory-based emotion induction paradigm, which effectively mitigates the limitations of traditional visual and auditory stimuli by providing more stable and sustained emotional responses. We also present a novel EEG dataset involving 32 subjects, developed through a pilot study to select appropriate olfactory stimuli. Experimental results demonstrate that our model significantly outperforms existing methods across multiple metrics, demonstrating the effectiveness of both the proposed model and emotion induction paradigm. This study underscores the potential of olfactory stimuli and advanced spatial-temporal modeling techniques for enhancing the robustness and performance of emotion recognition. Wei-Bang Jiang, Wei-Long Zheng, Bao-Liang Lu |
ICASSP | 4 |
| 2025 | EEGMirror: Leveraging EEG Data in the Wild Via Montage-Agnostic Self-Supervision for EEG to Video Decoding
Xuan-Hao Liu, Bao-Liang Lu, Wei-Long Zheng |
ICCV | 2 |
| 2025 | NeuroLM: A Universal Multi-task Foundation Model for Bridging the Gap between Language and EEG SignalsabstractRecent advancements for large-scale pre-training with neural signals such as electroencephalogram (EEG) have shown promising results, significantly boosting the development of brain-computer interfaces (BCIs) and healthcare. However, these pre-trained models often require full fine-tuning on each downstream task to achieve substantial improvements, limiting their versatility and usability, and leading to considerable resource wastage. To tackle these challenges, we propose NeuroLM, the first multi-task foundation model that leverages the capabilities of Large Language Models (LLMs) by regarding EEG signals as a foreign language, endowing the model with multi-task learning and inference capabilities. Our approach begins with learning a text-aligned neural tokenizer through vector-quantized temporal-frequency prediction, which encodes EEG signals into discrete neural tokens. These EEG tokens, generated by the frozen vector-quantized (VQ) encoder, are then fed into an LLM that learns causal EEG information via multi-channel autoregression. Consequently, NeuroLM can understand both EEG and language modalities. Finally, multi-task instruction tuning adapts NeuroLM to various downstream tasks. We are the first to demonstrate that, by specific incorporation with LLMs, NeuroLM unifies diverse EEG tasks within a single model through instruction tuning. The largest variant NeuroLM-XL has record-breaking 1.7B parameters for EEG signal processing, and is pre-trained on a large-scale corpus comprising approximately 25,000-hour EEG data. When evaluated on six diverse downstream datasets, NeuroLM showcases the huge potential of this multi-task learning paradigm. Wei-Bang Jiang, Yansen Wang, Bao-Liang Lu, Dongsheng Li 0002 |
ICLR | 3 |
| 2025 | Hierarchical Emotion Transformer for Multimodal Joint Emotion Category and Intensity Recognition
Tian-Fang Ma, Wei-Bang Jiang, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (1) | 4 |
| 2025 | Multi-session Meditative EEG Classification with Noise-Robust Self-supervision
Zuxin Song, Mingyu Gou, Tianzhen Chen, Bao-Liang Lu, Wei-Long Zheng |
ICONIP (2) | 5 |
| 2025 | Multimodal Emotion Recognition with Missing Modality via a Unified Multi-task Pre-training FrameworkabstractMultimodal emotion recognition based on physiological signals faces the challenge of missing modality due to issues such as inaccurate signal synchronization and inadequate device contact. Existing methods either require additional generative modules to handle missing modalities, leading to extra computational overhead, or fail to effectively capture both modality-specific and joint representation. In contrast, we propose a Unified Multi-task Pre-training (UMAP) framework based on the mixture of experts structure. Our approach offers two key advantages: (1) It retains a joint structure while flexibly handling both unimodal and multimodal inputs by selecting lightweight modality experts and shared experts, thus preserving both modality-specific features and joint multimodal information. 2) Three pre-training tasks-contrastive learning, modality matching, and modality generation-are integrated into UMAP through different attention masks, enhancing the model's ability to adapt to both complete and incomplete modalities during fine-tuning. Comprehensive experiments conducted on three benchmark datasets demonstrate that UMAP achieves state-of-the-art (SOTA) performance, both in multimodal scenarios and in cases where any modality is missing. The code is available at https://github.com/iiieeeve/UMAP. Ziyi Li 0003, Wei-Long Zheng, Bao-Liang Lu |
ACM Multimedia | 3 |
| 2025 | Human vs AI: How Digital Human News Anchors Affect Our Cognitive Processes?abstractWith the advancement of Artificial Intelligence Generated Content (AIGC) technology, digital human representations are increasingly appearing in multimedia interactions. This trend is particularly prominent in news broadcasting. The uniformity of news anchors' appearances and broadcasting environments has facilitated the widespread adoption of AI-powered news anchors. With high accuracy in news reporting and advancements in technology, AI anchors have been increasingly implemented in various news programs. However, there is currently a lack of objective analysis regarding the cognitive impact of digital human news broadcasting on audiences and its corresponding effects on brain signals. In this work, we investigate the differences in electroencephalography (EEG) responses when subjects watch news broadcasts delivered by digital humans versus real human anchors under various conditions. Our contributions are threefold: 1) We develop a dataset recording EEG signals from 32 subjects while they were watching news broadcasts. According to the presentation format (human/AI anchor) and the level of attention (high/normal), we categorize the dataset into four groups. 2) We utilize EEG signals to analyze the perceptual differences of subjects when watching news presented in different formats and in varying attention states. In addition, we investigate the cognitive differences of the subjects in perceived authenticity and importance of the news under these different conditions. 3) We propose an asymmetric multi-representation learning framework to better utilize and analyze the data. The code and data are available at https://github.com/Arcee-LYK/EEG-News. Yan-Kai Liu, Shunyang Yao, Bao-Liang Lu, Wei-Long Zheng |
ACM Multimedia | 4 |
| 2025 | SEED-MYA: A Novel Myanmar Multimodal Dataset for Enhancing Emotion RecognitionabstractThis paper introduces a novel Myanmar multimodal dataset called SEED-MYA, which is the first culturally and linguistically tailored multimodal emotion recognition dataset for the Burmese speakers. The SEED-MYA dataset consists of EEG and eye movement data collected using Myanmar video stimuli, addressing the underrepresentation of minority cultures in emotion recognition research. To investigate the fundamental characteristics of emotion recognition based on EEG and eye movement data from Myanmar participants, and to validate the quality and effectiveness of the SEED-MYA dataset, we implement the Multimodal Adaptive Emotion Transformer with Cross-Modal Attention (MAET-CMA) as a benchmarking tool. From our experiments on SEED-MYA, we have three main findings: (a) beta and gamma bands play a critical role in distinguishing positive, neutral, and negative emotional states; (b) combining EEG and eye movement data significantly enhances emotion recognition accuracy, with MAET-CMA achieving a maximum accuracy of 91.78%; and (c) EEG signals excel in recognizing negative and positive emotional states, while eye movement data are particularly effective at differentiating neutral emotion in our experimental setup. Our neural activity analysis reveals distinct patterns of activation in temporal, parietal, and prefrontal regions, providing insights into potential culture-related neural responses. We further compare our findings with established benchmarks from Chinese participants (SEED dataset) to explore cultural similarities and differences in emotion recognition. This analysis is structured into three components: band-level EEG comparisons, modality-specific performance analysis, and neural activity differences, providing a multi-level analysis of cultural effects. While some similarities in emotion recognition are observed across two cultures, our cross-cultural performance comparison between SEED and SEED-MYA further indicates that the Chinese dataset generalizes better as a test set. These findings underscore the importance of incorporating culturally diverse datasets in the development of globally applicable emotion recognition systems. Khin Pa Pa Aung, Hao-Long Yin, Tian-Fang Ma, Wei-Long Zheng, Bao-Liang Lu |
IEEE Trans. Affect. Comput. | 5 |
| 2025 | SEED-VII: A Multimodal Dataset of Six Basic Emotions With Continuous Labels for Emotion RecognitionabstractRecognizing emotions from physiological signals is a topic that has garnered widespread interest, and research continues to develop novel techniques for perceiving emotions. However, the emergence of deep learning has highlighted the need for comprehensive and high-quality emotional datasets that enable the accurate decoding of human emotions. To systematically explore human emotions, we develop a multimodal dataset consisting of six basic (happiness, sadness, fear, disgust, surprise, and anger) emotions and the neutral emotion, named SEED-VII. This multimodal dataset includes electroencephalography (EEG) and eye movement signals. The seven emotions in SEED-VII are elicited by 80 different videos and fully investigated with continuous labels that indicate the intensity levels of the corresponding emotions. Additionally, we propose a novel Multimodal Adaptive Emotion Transformer (MAET), that can flexibly process both unimodal and multimodal inputs. Adversarial training is utilized in the MAET to mitigate subject discrepancies, which enhances domain generalization. Our extensive experiments, encompassing both subject-dependent and cross-subject conditions, demonstrate the superior performance of the MAET in terms of handling various inputs. Continuous labels are used to filter the data with high emotional intensity, and this strategy is proven to be effective for attaining improved emotion recognition performance. Furthermore, complementary properties between the EEG signals and eye movements and stable neural patterns of the seven emotions are observed. Wei-Bang Jiang, Xuan-Hao Liu, Wei-Long Zheng, Bao-Liang Lu |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | Investigating the Effects of Sleep Conditions on Emotion Responses with EEG Signals and Eye MovementsabstractExisting studies in psychology and neuroscience have extensively examined the effects of sleep deprivation on emotional responses. More recently, researchers have begun applying deep learning algorithms to further investigate this relationship, emphasizing the importance of accessible and high-quality multimodal datasets across different sleep states. To address this need, we develop SEED-SD, a multimodal dataset comprising data from 40 participants. The dataset includes electroencephalography (EEG) and eye movement signals collected under three sleep conditions: sleep deprivation (SD), sleep recovery (SR), and normal sleep (NS). Each condition contains data corresponding to four basic emotions: happiness, sadness, fear, and neutral state. Additionally, we propose a novelRegion Transformer withLayer-Fusion (ReLF), to conduct comprehensive analyses on the SEED-SD dataset. ReLF incorporates a region- wise self-attention mechanism to extract localized features from EEG and eye movement signals, and supports flexible adaptation to both multimodal and unimodal inputs. Following multimodal generative pre-training, ReLF introduces learnable prompts to replace the missing modality under unimodal settings, thereby enabling effective fine-tuning of the pre-trained model. The experimental results demonstrate that ReLF outperforms the existing models. Our analysis further reveals that SD significantly impacts emotion recognition performance, while SR and NS conditions yield similar results, highlighting the importance of SR in mitigating the adverse effects of SD. Furthermore, we conduct a systematic analysis of multimodal complementarity, critical frequency bands, and neural patterns. Our findings reveal distinct EEG patterns under the SD condition compared to the SR and NS conditions. Notably, the multimodal complementarity and critical frequency bands in both the SD and SR conditions align with those observed in the NS condition. In summary, to the best of our knowledge, SEED-SD is the largest publicly available multimodal dataset for studying the relationship between emotion recognition and sleep states. This dataset lays a crucial foundation for applying deep learning methods in this area. Moreover, through extensive data-driven analysis, this work confirms the inhibitory effect of SD on emotion recognition and the restorative role of SR. The SEED-SD dataset and codes will be public upon paper acceptance. Ziyi Li 0003, Le-Yan Tao, Rui-Xiao Ma, Wei-Long Zheng, Bao-Liang Lu |
IEEE Trans. Affect. Comput. | 5 |
| 2025 | LibEER: A Comprehensive Benchmark and Algorithm Library for EEG-Based Emotion RecognitionabstractEEG-based emotion recognition (EER) has gained significant attention due to its potential for understanding and analyzing human emotions. While recent advancements in deep learning techniques have substantially improved EER, the field lacks a convincing benchmark and comprehensive open-source libraries. This absence complicates fair comparisons between models and creates reproducibility challenges for practitioners, which collectively hinder progress. To address these issues, we introduce LibEER, a comprehensive benchmark and algorithm library designed to facilitate fair comparisons in EER. LibEER carefully selects popular and powerful baselines, harmonizes key implementation details across methods, and provides a standardized codebase in PyTorch. By offering a consistent evaluation framework with standardized experimental settings, LibEER enables unbiased assessments of seventeen representative deep learning models for EER across the six most widely used datasets. Additionally, we conduct a thorough, reproducible comparison of model performance and efficiency, providing valuable insights to guide researchers in the selection and design of EER models. Moreover, we make observations and in-depth analysis on the experiment results and identify current challenges in this community. We hope that our work will not only lower entry barriers for newcomers to EEG-based emotion recognition but also contribute to the standardization of research in this domain, fostering steady development. The library and source code are publicly available athttps://github.com/XJTU-EEG/LibEER. Huan Liu 0012, Shusen Yang, Yuzhe Zhang 0003, Mengze Wang, Fanyu Gong, Chengxi Xie, Guanjian Liu, Zejun Liu, Yong-Jin Liu 0001, Bao-Liang Lu, Dalin Zhang 0001 |
IEEE Trans. Affect. Comput. | 10 |
| 2025 | From EEG to Eye Movements: Cross-Modal Emotion Recognition Using Constrained Adversarial Network With Dual AttentionabstractEmotion recognition is a fundamental part of affective computing, obtaining performance gain from multimodal methods. Electroencephalography (EEG) and eye movements are extensively used as they contain complementary information. However, the inconvenient acquisition of EEG is hindering the extensive adoption of multimodal emotion recognition in daily applications while eye movements are more convenient to collect but with lower performance. To tackle this issue, we propose a Constrained Adversarial Network with Dual Attention (CANDA), exploiting the complementary information from multiple modalities during training to improve the test-time performance of single easily acquired modality, i.e., transferring knowledge from a stronger modality to a weaker modality. During training, a common joint space is learned to diminish the distribution discrepancy among different modalities and incorporate the multimodal representations. During test, single modality is converted to the common space achieving comparable performance to multiple modalities. Extensive experiments demonstrate that our model achieves the state-of-the-art performance for cross-modal emotion recognition. Specifically, the mean accuracy increases around 15% on SEED, 15% on SEED-IV, and 2% on SEED-V compared to the latest baseline for emotion recognition. Visualization of features in the joint space illustrates that the distribution of different modalities aligns together with the discriminative ability regarding various emotions. Jia-Wen Liu, Bao-Liang Lu, Wei-Long Zheng |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | Multi-View Self-Supervised Domain Adaptation for EEG-Based Emotion RecognitionabstractResearch on emotion recognition based on EEG signals has made significant progress. Most of the existing studies have focused on supervised learning methods, but real-life data cannot meet the requirement of high quality with labels. In addition, EEG signals have individual variability and instability, which requires transfer learning to enhance the model generalization. In this paper, we propose a multi-view self-supervised domain adaptation model that combines self-supervised learning techniques with domain-adaptive transfer learning algorithms, which can solve the last two problems mentioned above. Specifically, we add a multi-class domain discriminator to construct the adversarial relationship between the sub-networks so that distribution discrepancy of different subjects can be reduced effectively. We conduct both subject-dependent and subject-independent experiments on the SEED and SEED-IV datasets to thoroughly evaluate the performance of our model. The results show that our model achieves outstanding emotion recognition performance even with limited labeled data. In the subject-dependent experiments on both datasets, our model achieves accuracy rates of 85.91% and 87.19% respectively, surpassing the original self-supervised masked autoencoder model by about 3%. In subject-independent experiments, our model demonstrates strong data distribution adaptation capabilities, achieving an accuracy of 69.72% and 62.87%, respectively on the SEED and SEED-IV datasets using only 90 samples for subject-independent experiments. This effectively mitigates the accuracy degradation caused by differences in data distribution across subjects. Furthermore, our model is capable of extracting meaningful features from corrupted EEG data, highlighting its robustness and effectiveness. Lu Zhang 0096, Hanwen Shi, Juan-Zi Li, Wei-Long Zheng, Bao-Liang Lu |
IEEE Trans. Affect. Comput. | 5 |
| 2025 | EEG-Based Emotion Monitoring and Regulation System by Learning the Discriminative Brain Network ManifoldabstractEmotion recognition based on electroencephalogram (EEG) is fundamentally associated with human-like intelligence system. However, due to the noise-sensitive characteristics of EEGs and the individual variability of emotions, it is very challenging to extract inherent emotion dependent patterns from emotional EEG signals. In this work, we propose a L1-norm space defined discriminative brain network manifold learning model (L1-SGL), in which the EEG noise outliers can be effectively separated and the pseudolabeled samples caused by subjective feelings can be automatically corrected. Off-line experimental results consistently indicate that the L1-SGL can effectively suppress the influence of noise and achieve an incomparable superiority performance over other existing methods in EEG emotion recognition. Besides, benefiting from the time efficiency of the L1-SGL, an online emotion monitoring and regulation system is further implemented in this work. On-line emotion decoding experimental results (86.30%) of 25 participants prove that the L1-SGL can effectively satisfy the real-time requirements of on-line emotional monitoring applications, and the significant negative emotion regulation experimental results ( $p \lt 0.001$ ) further confirm the feasibility and effectiveness of L1-SGL model in real-time emotion regulation and interactive applications. Overall, the L1-SGL provides a promising solution for the real-time online affective brain-computer interfaces (aBCIs) and the intelligent clinical closed-loop treatments. Cunbo Li, Zehong Cao, Yue Pan 0010, Fali Li, Huafu Chen, Bao-Liang Lu, Feng Wan 0003, Dezhong Yao 0001, Peng Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2024 | Addressing Temporal and Auditory Factors in Meditative EEG with Self-Supervised LearningabstractMeditation has been shown to improve mental health by reducing stress and enhancing cognitive functions. To monitor meditation quality, Electroencephalogram (EEG) offers an objective alternative to self-evaluation questionnaires, which suffer from subjectivity and delays. However, previous EEG protocols often inflated classification performance due to temporal correlation and auditory interference. This study introduces a new data acquisition protocol to address these issues. We constructed a dataset from 32 subjects, each participating on three sessions in different days to eliminate temporal interference. During each session, subjects alternated between audio-guided meditation and silent unguided meditation. Using the Multi-view Spectral-Spatial-Temporal Masked Autoencoder (MV-SSTMA) model which was pre-trained on the Emotion EEG dataset SEED, the model achieved superior classification performance in F1 score, accuracy, precision, and recall. These findings offer valuable insights for improving EEG-based meditation monitoring in biomedical applications. Mingyu Gou, Ying-Jie Zhang, Ren-Jie Dai, Hao-Long Yin, Tianzhen Chen, Bao-Liang Lu, Wei-Long Zheng |
BIBM | 7 |
| 2024 | MoGE: Mixture of Graph Experts for Cross-subject Emotion Recognition via Decomposing EEGabstractDecoding emotions of previously unseen subjects from electroencephalography (EEG) signals is challenging due to the inter-subject variability. Domain Generalization (DG) methods aim to mitigate the domain shift among different subjects. Once trained, a DG model can be directly deployed on new subjects without any calibration phase. While existing DG studies on cross-subject emotion recognition mainly focus on the design of loss function for domain alignment or regularization, we introduce Sparse Mixture of Graph Experts (MoGE) model to explore DG issues from a new perspective, i.e. the design of the neural architecture. In the MoGE model, routers allocate each EEG channel to a specialized expert, thereby facilitating the decomposition of the intricate brain into distinct functional areas. Extensive experiments on three public datasets demonstrate that compared to other DG methods, our MoGE model trained with empirical risk minimization (ERM) achieves the state-of-the-art (SOTA) accuracies, 88.0%, 74.3%, and 81.8% on SEED, SEED-IV, and SEED-V datasets, respectively. Our code is available at https://github.com/XuanhaoLiu/MoGE. Xuan-Hao Liu, Wei-Bang Jiang, Wei-Long Zheng, Bao-Liang Lu |
BIBM | 4 |
| 2024 | Emotion Recognition from Eye Movements Using Multi-way Autoregressive ModelabstractThe application of physiological signals in emotion recognition is a popular research topic in human-computer interactions. Eye movement, as an important physiological signal, plays an essential role in medicine, psychology, cognitive science, and other scientific research fields. Previous studies have successfully identified human emotions by combining various eye-related measurements, selecting features, and utilizing machine learning techniques. However, the exploration of eye movement signals in emotion recognition remains insufficient. In this study, we utilize eye tracking heatmap and eye movement trajectory data for emotion recognition for the first time. Based on the two types of eye movement data, we develop a multi-way autoregressive model capable of processing multi-view eye movement data. Compared to traditional deep learning baseline models, our model better adapts to the structure of eye movement data and significantly improves the classification performance. Furthermore, we integrate heatmap and trajectory data with commonly used eye-related measurement features, which further enhance the performance of emotion recognition beyond previous methods. Tian-Fang Ma, Xuan-Hao Liu, Wei-Long Zheng, Bao-Liang Lu |
BIBM | 4 |
| 2024 | Multimodal Multi-View Spectral-Spatial-Temporal Masked Autoencoder for Self-Supervised Emotion RecognitionabstractEmotion recognition is a primary and complex task in emotional intelligence. Due to the complexity of human emotions, utilizing multimodal fusion methods can enhance the performance by leveraging the complementary properties of different modalities. In this paper, we propose a Multimodal Multi-view Spectral-Spatial-Temporal Masked Autoencoder (Multimodal MV-SSTMA) with self-supervised learning to investigate multimodal emotion recognition based on electroencephalogram (EEG) and eye movement signals. Our experimental process comprises three stages: 1) In the pre-training stage, we employ MV-SSTMA to train feature extractors for EEG and eye movement signals; 2) In the fine-tuning stage, the labeled data are input to the feature extractors to fuse and fine-tune the features; 3) In the testing stage, our model is applied to recognize emotions with test data to calculate the accuracies of different methods. Our experimental results demonstrate that the multimodal fusion model outperforms the unimodal model on both SEED-IV and SEED-V datasets. In addition, the proposed model can still effectively recognize emotions with various ratios of missing data. These results underscore the efficiency of multimodal self-supervised learning and data fusion in emotion recognition. Pengxuan Gao, Jia-Wen Liu, Bao-Liang Lu, Wei-Long Zheng |
ICASSP | 4 |
| 2024 | Functional Emotion Transformer for EEG-Assisted Cross-Modal Emotion RecognitionabstractMultimodal emotion recognition based on electroencephalography (EEG) and eye movements has attracted increasing attention due to their high performance and complementary properties. However, there are two challenges that hinder its practical applications: the inconvenient EEG data collection and high-cost data annotation. In contrast, eye movements are convenient to obtain and process in real scenarios. To combine high performance of EEG and easy setups of eye tracking, we propose a novel EEG-assisted Contrastive Learning Framework with a Functional Emotion Transformer (ECO-FET) for cross-modal emotion recognition. ECO-FET leverages both the functional brain connectivity and the spectral-spatial-temporal domain of EEG signals simultaneously, which dramatically benefit the learning of eye movements. The whole process consists of three phases: pre-training, test, and fine-tuning. ECO-FET exploits the complementary information provided by multiple modalities during pre-training in order to improve the performance of unimodal models. In the pre-training phase, unlabeled EEG and eye movement data are fed into the model to contrastively learn the emotional latent representations between the two modalities, while in the test phase, eye movements and few labeled EEG samples are used to predict different emotions. Experimental results on three public datasets demonstrate that ECO-FET surpasses the state-of-the-art dramatically. Wei-Bang Jiang, Ziyi Li 0003, Wei-Long Zheng, Bao-Liang Lu |
ICASSP | 4 |
| 2024 | CEMOAE: A Dynamic Autoencoder with Masked Channel Modeling for Robust EEG-Based Emotion RecognitionabstractEmotion recognition through electroencephalography (EEG) has been an area of active research, but the inherent sensitivity of EEG signals to noise and artifacts poses significant challenges, especially in real-world settings. These complications often necessitate the removal of corrupted channels, making it crucial to develop robust models capable of maintaining performance even when few channels are available. To address this, we propose the Corrupted EMOtion AutoEncoder (CEMOAE), an innovative approach that leverages masked channel modeling to maintain robust performance, achieved through three components: masked autoencoder pretraining for robust representation learning, random masked auxiliary task for implicit modeling of channel corruption, and masked auto-repair to explicitly narrow the data distribution gap between high-quality and corrupted EEG signals. Specifically, we first pretrain a masked autoencoder with the dynamic masking strategy for feature extractor initialization and channel recovery. During the finetuning stage, we mask EEG data using the auxiliary task to mimic real-world EEG corruption. We then employ the pretrained autoencoder to repair these signals and finetune the feature extractor for emotion recognition. Experiments on the SEED dataset demonstrate that CEMOAE achieves SOTA performance for emotion recognition under the random channel corruption simulation, validating the effectiveness of the proposed techniques. Yu-Ting Lan, Wei-Bang Jiang, Wei-Long Zheng, Bao-Liang Lu |
ICASSP | 4 |
| 2024 | Temporal-Spatial Prediction: Pre-Training on Diverse Datasets for EEG ClassificationabstractElectroencephalogram (EEG) classification tasks have received increasing attention because its high application value. Meanwhile, the great success of general pre-training models in language processing areas inspires us to excavate the potential of an EEG pre-trained model. This model is expected to adapt to diverse downstream tasks. However, current studies either ignore the temporal or spatial domain in EEG signals, or only use single datasets in pre-training. The proposed Temporal-Spatial Prediction (TSP) model effectively solve these issues. Specifically, the output of the TSP encoder serves as the input of two tasks: spatial prediction, i.e., masked autoencoder, and temporal prediction, i.e, contrastive predictive coding. In addition, in order to provide more diverse information and thus benefit the downstream fine-tuning, we pre-train TSP on six large EEG datasets with four different numbers of channels. Results on three public downstream datasets SEED, SEED-IV, TUEV demonstrate that TSP achieves the state-of-the-art performance on different EEG classification tasks. In addition, according to the ablation experiments, TSP performs better than the single-domain method, i.e. Temporal Prediction (TP) model and Spatial Prediction (SP) model. Ziyi Li 0003, Li-Ming Zhao, Wei-Long Zheng, Bao-Liang Lu |
ICASSP | 4 |
| 2024 | Large Brain Model for Learning Generic Representations with Tremendous EEG Data in BCIabstractThe current electroencephalogram (EEG) based deep learning models are typically designed for specific datasets and applications in brain-computer interaction (BCI), limiting the scale of the models and thus diminishing their perceptual capabilities and generalizability. Recently, Large Language Models (LLMs) have achieved unprecedented success in text processing, prompting us to explore the capabilities of Large EEG Models (LEMs). We hope that LEMs can break through the limitations of different task types of EEG datasets, and obtain universal perceptual capabilities of EEG signals through unsupervised pre-training. Then the models can be fine-tuned for different downstream tasks. However, compared to text data, the volume of EEG datasets is generally small and the format varies widely. For example, there can be mismatched numbers of electrodes, unequal length data samples, varied task designs, and low signal-to-noise ratio. To overcome these challenges, we propose a unified foundation model for EEG called Large Brain Model (LaBraM). LaBraM enables cross-dataset learning by segmenting the EEG signals into EEG channel patches. Vector-quantized neural spectrum prediction is used to train a semantically rich neural tokenizer that encodes continuous raw EEG channel patches into compact neural codes. We then pre-train neural Transformers by predicting the original neural codes for the masked EEG channel patches. The LaBraMs were pre-trained on about 2,500 hours of various types of EEG signals from around 20 datasets and validated on multiple different types of downstream tasks. Experiments on abnormal detection, event type classification, emotion recognition, and gait prediction show that our LaBraM outperforms all compared SOTA methods in their respective fields. Our code is available at https://github.com/935963004/LaBraM. Wei-Bang Jiang, Li-Ming Zhao, Bao-Liang Lu |
ICLR | 3 |
| 2024 | A Multi-task Emotion Recognition Model Based on Continuously Labeled EEG Signals
Rong-Fei Gu, Yi-Dong Zhao, Li-Ming Zhao, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (5) | 5 |
| 2024 | Detecting Major Depression Disorder with Multiview Eye Movement Features in a Novel Oil Painting ParadigmabstractMajor Depressive Disorder (MDD) is a debilitating condition marked by persistent low mood, reduced interest, cognitive impairments, and vegetative neurological symptoms such as sleep and appetite disturbances. In this paper, we collected eye movement signals from 40 patients diagnosed with MDD and 40 healthy controls to study the relation between eye movements and cognitive processes for depression detection. The eye movement data were captured during a novel emotional cognition task using oil paintings. Subsequently, the data were transformed into multiview eye movement features, including heatmaps, trajectories, and statistical vectors. Rigorous statistical analyses were then conducted on these features to identify significant patterns and correlations between eye movements and depressive symptoms. A multiview invariant & specific eye movement model (MISEYE) was proposed to fuse different eye movement features. The proposed achieved an accuracy rate of 79.88% in depression detection. This performance surpassed not only the outcomes of single-mode approaches and combinations of any two features but also outperformed other fusion methodologies. These findings not only shed light on the intricate relationship between eye movement patterns and MDD but also underscore the potential of eye-tracking technology in psychiatric research. Tian-Fang Ma, Lu-Yu Liu, Li-Ming Zhao, Dan Peng, Wei-Long Zheng, Bao-Liang Lu |
IJCNN | 7 |
| 2024 | REmoNet: Reducing Emotional Label Noise via Multi-regularized Self-supervisionabstractEmotion recognition based on electroencephalogram (EEG) has garnered increasing attention in recent years due to non-invasiveness and high reliability of EEG measurements. Despite the promising performance achieved by numerous existing methods, several challenges persist. Firstly, there is the challenge of emotional label noise, stemming from the assumption that emotions remain consistently evoked and stable throughout the entirety of video observation. Such an assumption proves difficult to uphold in practical experimental settings, leading to discrepancies between EEG signals and anticipated emotional states. In addition, there's a need for comprehensive capture of temporal-spatial-spectral characteristics of EEG signals and cope with low signal-to-noise ratio (SNR) issues. To tackle these challenges, we propose a comprehensive pipeline named REmoNet, which leverages novel self-supervised techniques and multi-regularized co-learning. Two self-supervised methods, including masked channel modeling via temporal-spectral transformation and emotion contrastive learning, are introduced to facilitate the comprehensive understanding and extraction of emotion-relevant EEG representations during pre-training. Additionally, fine-tuning with multi-regularized co-learning exploits feature-dependent information through intrinsic similarity, resulting in mitigating emotional label noise. Experimental evaluations on two public datasets demonstrate that our proposed approach, REmoNet, surpasses existing state-of-the-art methods, showcasing its effectiveness in simultaneously addressing raw EEG signals and noisy emotional labels. Wei-Bang Jiang, Yu-Ting Lan, Bao-Liang Lu |
ACM Multimedia | 3 |
| 2024 | EEG2Video: Towards Decoding Dynamic Visual Perception from EEG SignalsabstractOur visual experience in daily life are dominated by dynamic change. Decoding such dynamic information from brain activity can enhance the understanding of the brain’s visual processing system. However, previous studies predominately focus on reconstructing static visual stimuli. In this paper, we explore to decode dynamic visual perception from electroencephalography (EEG), a neuroimaging technique able to record brain activity with high temporal resolution (1000 Hz) for capturing rapid changes in brains. Our contributions are threefold: Firstly, we develop a large dataset recording signals from 20 subjects while they were watching 1400 dynamic video clips of 40 concepts. This dataset fills the gap in the lack of EEG-video pairs. Secondly, we annotate each video clips to investigate the potential for decoding some specific meta information (e.g., color, dynamic, human or not) from EEG. Thirdly, we propose a novel baseline EEG2Video for video reconstruction from EEG signals that better aligns dynamic movements with high temporal resolution brain signals by Seq2Seq architecture. EEG2Video achieves a 2-way accuracy of 79.8% in semantic classification tasks and 0.256 in structural similarity index (SSIM). Overall, our works takes an important step towards decoding dynamic visual perception from EEG signals. Our dataset and code will be released soon. Xuan-Hao Liu, Yan-Kai Liu, Yansen Wang, Kan Ren, Hanwen Shi, Zilong Wang 0006, Dongsheng Li 0002, Bao-Liang Lu, Wei-Long Zheng |
NeurIPS | 8 |
| 2024 | Brain Network Manifold Learned by Cognition-Inspired Graph Embedding Model for Emotion RecognitionabstractElectroencephalogram (EEG) brain network embodies the brain’s coordination and interaction mechanism, and the transformations of emotional states are usually accompanied with changes in brain network spatial topologies. To effectively characterize emotions, in this work, we propose a cognition-inspired graph embedding model in the L1-norm space (L1-CGE) to learn an optimal low-dimensional embedded manifold for emotional brain networks. In the L1-CGE, the original brain networks are first encoded in the affinity space with the proposed cognition-inspired metric to construct the latent geometry manifold structure of emotional brain networks, and then the graph learning objective function is defined in the L1-norm space to obtain the optimal low-dimensional representations of brain networks. Essentially, the modularized community structures of emotional brain networks can be effectively emphasized by the L1-CGE to realize an effective depiction for emotions. Compared with existing methods, the L1-CGE model has achieved state-of-the-art performance on three public emotional EEG datasets in off-line conditions. Besides, the robust real-time experimental results have been achieved with the on-line emotion decoding system designed with L1-CGE. Both off- and on-line experimental results consistently demonstrate that the proposed L1-CGE is promising to provide a potential solution for the real-time affective brain-computer interface (aBCI) system. Cunbo Li, Zhaojin Chen, Fali Li, Feng Wan 0003, Zehong Cao, Dezhong Yao 0001, Bao-Liang Lu, Peng Xu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 9 |
| 2023 | Transformer-Based Domain Adaptation for Multi-Modal Emotion Recognition in Response to Game Animation VideosabstractEmotion recognition researches necessitate the strategic selection of stimuli to evoke targeted emotions for robust physiological analyses. This study pioneers the use of Genshin Impact game animation videos to induce positive and neutral emotional states. Notably, it introduces a Transformer-based feature extractor, enhancing Domain-Adversarial Neural Networks (DANN) to advance domain adaptation capabilities. Leveraging the inherent advantages of the Transformer architecture, including parallel processing and handling intricate time-series Electroencephalogram (EEG) and eye movement data, this innovation is reinforced by a discriminator with gradient reversal layers, harmonizing source and target domain distributions. Empirical results demonstrate the effectiveness of the Transformer-based DANN model in cross-subject multimodal emotion recognition, achieving a remarkable prediction accuracy of 83.38% across 59 subjects. Jing-Yi Liu, Jia-Wen Liu, Wei-Long Zheng, Bao-Liang Lu |
BIBM | 4 |
| 2023 | Elastic Graph Transformer Networks for EEG-Based Emotion RecognitionabstractElectroencephalogram (EEG) has been applied in emotion recognition due to excellent temporal resolution with less competitive spatial resolution. This leads to the consequence that the majority of EEG-based emotion recognition models emphasize on exploiting temporal features while ignoring the efficient information provided by spatial resolution. To extract more informative representations, we propose an elastic Graph Transformer network for emotion recognition (EmoGT) inspired by the advantages of Transformer in time-series analysis and the superior performance of graph convolutional networks in topological analysis. Moreover, it is able to be flexibly expanded to cope with multimodal inputs by employing specially designed structures. Experimental results on 3 public datasets demonstrate that our models outperform the state-of-the-art results by 3% on average in both single and multimodal cases, indicating the effectiveness of utilizing temporal and spatial information simultaneously. Wei-Bang Jiang, Xu Yan 0007, Wei-Long Zheng, Bao-Liang Lu |
ICASSP | 4 |
| 2023 | DAformer: Transformer with Domain Adversarial Adaptation for EEG-Based Emotion Recognition with Live-Oil Paintings
Zhong-Wei Jin, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (9) | 4 |
| 2023 | Two-Stream Spectral-Temporal Denoising Network for End-to-End Robust EEG-Based Emotion Recognition
Xuan-Hao Liu, Wei-Bang Jiang, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (3) | 4 |
| 2023 | Naturalistic Emotion Recognition Using EEG and Eye Movements
Ziyi Li 0003, Tian-Fang Ma, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (3) | 7 |
| 2023 | Cross-Subject Decision Confidence Estimation from EEG Signals Using Spectral-Spatial-Temporal Adaptive GCN with Domain AdaptationabstractThe study of the human decision-making process has long been a valuable field for both scientific research and practical application. Towards knowing and taking control of the decision-making process, evaluating the reliability of human decisions objectively plays an important role. Various studies have demonstrated that the confidence level of humans during the decision-making process is an important factor that reflects the correctness of decisions. In literature, several deep learning based methods have been developed to estimate decision confidence using Electroencephalography (EEG). Among these approaches, the spectral-spatial-temporal adaptive graph convolutional neural network (SST-AGCN) stands out. However, SST-AGCN focuses on specific subjects, and may lead to less efficiency in cross-subject situations, which are more common in application scenarios. In this paper, we propose a deep learning model called SST-AGCN with Domain Adaptation (SST-AGCN-DA) for cross-subject decision confidence estimation. To examine the effectiveness of our proposed model, we compare our SST-AGCN-DA with the original SST-AGCN, three typical domain adaption algorithms in the field, and the SST-AGCN with Domain Generalization (SST-AGCN-DG), which is another transfer learning model we developed in this paper. We conduct cross-subject confidence estimation experiments on an EEG dataset collected under a text-based decision-making task. The averaged results of leave-one-out cross-validation come out that the F1-scores of our proposed SST-AGCN-DA and SST-AGCN-DG are 79.45% and 77.04%, respectively, while the original SST-AGCN and the best of the existing domain adaptation algorithms are 74.15% and 74.25%, respectively. Rong-Fei Gu, Rui Li 0048, Wei-Long Zheng, Bao-Liang Lu |
IJCNN | 4 |
| 2023 | Multimodal Adaptive Emotion Transformer with Flexible Modality Inputs on A Novel Dataset with Continuous LabelsabstractEmotion recognition from physiological signals is a topic of widespread interest, and researchers continue to develop novel techniques for perceiving emotions. However, the emergence of deep learning has highlighted the need for high-quality emotional datasets to accurately decode human emotions. In this study, we present a novel multimodal emotion dataset that incorporates electroencephalography (EEG) and eye movement signals to systematically explore human emotions. Seven basic emotions (happy, sad, fear, disgust, surprise, anger, and neutral) are elicited by a large number of 80 videos and fully investigated with continuous labels that indicate the intensity of the corresponding emotions. Additionally, we propose a novel Multimodal Adaptive Emotion Transformer (MAET), that can flexibly process both unimodal and multimodal inputs. Adversarial training is utilized in MAET to mitigate subject discrepancy, which enhances domain generalization. Our extensive experiments, encompassing both subject-dependent and cross-subject conditions, demonstrate MAET's superior performance in handling various inputs. The filtering of data for high emotional evocation using continuous labels proved to be effective in the experiments. Furthermore, the complementary properties between EEG and eye movements are observed. Our code is available at https://github.com/935963004/MAET. Wei-Bang Jiang, Xuan-Hao Liu, Wei-Long Zheng, Bao-Liang Lu |
ACM Multimedia | 4 |
| 2023 | Affective Brain-Computer Interfaces (aBCIs): A TutorialabstractA brain–computer interface (BCI) enables a user to communicate directly with a computer using only the central nervous system. An affective BCI (aBCI) monitors and/or regulates the emotional state of the brain, which could facilitate human cognition, communication, decision-making, and health. The last decade has witnessed rapid progress in aBCI research and applications, but there does not exist a comprehensive and up-to-date tutorial on aBCIs. This tutorial fills the gap. It introduces first the basic concepts of BCIs and then, in detail, the individual components in a closed-loop aBCI system, including signal acquisition, signal processing, feature extraction, emotion recognition, and brain stimulation. Next, it describes three representative applications of aBCIs, i.e., cognitive workload recognition, fatigue estimation, and depression diagnosis and treatment. Several challenges and opportunities in aBCI research and applications, including brain signal acquisition, emotion labeling, diversity and size of aBCI datasets, algorithm comparison, negative transfer in emotion recognition, and privacy protection and security of aBCIs, are also explained. Dongrui Wu, Bao-Liang Lu, Bin Hu 0001, Zhigang Zeng |
Proc. IEEE | 2 |
| 2023 | Joint EEG Feature Transfer and Semisupervised Cross-Subject Emotion RecognitionabstractDue to the weak and nonstationary properties, electroencephalogram (EEG) data present significant individual differences. To align data distributions of different subjects, transfer learning showed promising performance in cross-subject EEG emotion recognition. However, most of the existing models sequentially learned the domain-invariant features and estimated the target domain label information. Such a two-stage strategy breaks the inner connections of both processes, inevitably causing the suboptimality. In this article, we propose a joint EEG feature transfer and semisupervised cross-subject emotion recognition model in which the shared subspace projection matrix and target label are jointly optimized toward the optimum. Extensive experiments are conducted on SEED-IV and SEED, and the results show that the emotion recognition performance is significantly enhanced by the joint learning mode and the spatial-frequency activation patterns of critical EEG frequency bands and brain regions in cross-subject emotion expression are quantitatively identified by analyzing the learned shared subspace. Yong Peng 0001, Honggang Liu, Wanzeng Kong, Feiping Nie 0001, Bao-Liang Lu, Andrzej Cichocki |
IEEE Trans. Ind. Informatics | 5 |
| 2022 | Measuring Decision Confidence Levels from EEG Using a Spectral-Spatial-Temporal Adaptive Graph Convolutional Neural Network
Rui Li 0048, Bao-Liang Lu |
ICONIP (5) | 3 |
| 2022 | Few-Shot Class-Incremental Learning for EEG-Based Emotion Recognition
Tian-Fang Ma, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (5) | 3 |
| 2022 | Increasing the Stability of EEG-based Emotion Recognition with a Variant of Neural ProcessesabstractElectroencephalography (EEG) signals provide an incredible promising access to decode human emotions in affective computing. Many approaches have been applied to building EEG-based affective models and much endeavor is made to improve the performance of affective models, where test data has similar quality as training data. However, due to the strong sensitivity of EEG to external factors such as body movement and electromagnetic interference, EEG signals usually have a lot of noise and subjects have to remain as motionless as possible in a quiet environment, which is difficult to be satisfied in real applications and severely influences the user experience. To deal with this problem, a common way is to drop those channels with intensive noise. However, this results in the loss of critical information and neither existing machine learning (ML) nor deep learning (DL) approaches can handle this situation well, especially when too many channels are missed. In this paper, we propose a robust variant of the neural processes model and evaluate the stability of our model under various circumstances to simulate random data corruption in real applications. We conduct two categories of experiments by controlling the number and the places of missing channels separately and compare with classical ML and DL models. The results demonstrate that the performance of our model significantly outperforms the existing models, even using only five channels. We also explore the critical brain regions in the current EEG electrode distribution. Final performance manifests our model has extreme stability in dealing with intractable situations and sheds light on the widespread usage of portable EEG-based affective computing. Yan-Kai Liu, Wei-Bang Jiang, Bao-Liang Lu |
IJCNN | 3 |
| 2022 | A Multi-view Spectral-Spatial-Temporal Masked Autoencoder for Decoding Emotions with Self-supervised LearningabstractAffective Brain-computer Interface has achieved considerable advances that researchers can successfully interpret labeled and flawless EEG data collected in laboratory settings. However, the annotation of EEG data is time-consuming and requires a vast workforce which limits the application in practical scenarios. Furthermore, daily collected EEG data may be partially damaged since EEG signals are sensitive to noise. In this paper, we propose a Multi-view Spectral-Spatial-Temporal Masked Autoencoder (MV-SSTMA) with self-supervised learning to tackle these challenges towards daily applications. The MV-SSTMA is based on a multi-view CNN-Transformer hybrid structure, interpreting the emotion-related knowledge of EEG signals from spectral, spatial, and temporal perspectives. Our model consists of three stages: 1) In the generalized pre-training stage, channels of unlabeled EEG data from all subjects are randomly masked and later reconstructed to learn the generic representations from EEG data; 2) In the personalized calibration stage, only few labeled data from a specific subject are used to calibrate the model; 3) In the personal test stage, our model can decode personal emotions from the sound EEG data as well as damaged ones with missing channels. Extensive experiments on two open emotional EEG datasets demonstrate that our proposed model achieves state-of-the-art performance on emotion recognition. In addition, under the abnormal circumstance of missing channels, the proposed model can still effectively recognize emotions. Rui Li 0048, Wei-Long Zheng, Bao-Liang Lu |
ACM Multimedia | 4 |
| 2022 | Joint Feature Adaptation and Graph Adaptive Label Propagation for Cross-Subject Emotion Recognition From EEG SignalsabstractThough Electroencephalogram (EEG) could objectively reflect emotional states of our human beings, its weak, non-stationary, and low signal-to-noise properties easily cause the individual differences. To enhance the universality of affective brain-computer interface systems, transfer learning has been widely used to alleviate the data distribution discrepancies among subjects. However, most of existing approaches focused mainly on the domain-invariant feature learning, which is not unified together with the recognition process. In this paper, we propose a joint feature adaptation and graph adaptive label propagation model (JAGP) for cross-subject emotion recognition from EEG signals, which seamlessly unifies the three components of domain-invariant feature learning, emotional state estimation and optimal graph learning together into a single objective. We conduct extensive experiments on two benchmark SEED_IV and SEED_V data sets and the results reveal that 1) the recognition performance is greatly improved, indicating the effectiveness of the triple unification mode; 2) the emotion metric of EEG samples are gradually optimized during model training, showing the necessity of optimal graph learning, and 3) the projection matrix-induced feature importance is obtained based on which the critical frequency bands and brain regions corresponding to subject-invariant features can be automatically identified, demonstrating the superiority of the learned shared subspace. Yong Peng 0001, Wanzeng Kong, Feiping Nie 0001, Bao-Liang Lu, Andrzej Cichocki |
IEEE Trans. Affect. Comput. | 5 |
| 2022 | Tri-training for Dependency Parsing Domain AdaptationabstractIn recent years, the research on dependency parsing focuses on improving the accuracy of the domain-specific (in-domain) test datasets and has made remarkable progress. However, there are innumerable scenarios in the real world that are not covered by the dataset, namely, the out-of-domain dataset. As a result, parsers that perform well on the in-domain data usually suffer from significant performance degradation on the out-of-domain data. Therefore, to adapt the existing in-domain parsers with high performance to a new domain scenario, cross-domain transfer learning methods are essential to solve the domain problem in parsing. This paper examines two scenarios for cross-domain transfer learning: semi-supervised and unsupervised cross-domain transfer learning. Specifically, we adopt a pre-trained language model BERT for training on the source domain (in-domain) data at the subword level and introduce self-training methods varied from tri-training for these two scenarios. The evaluation results on the NLPCC-2019 shared task and universal dependency parsing task indicate the effectiveness of the adopted approaches on cross-domain transfer learning and show the potential of self-learning to cross-lingual transfer learning. Zuchao Li, Hai Zhao 0001, Bao-Liang Lu, Rui Wang 0015 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2022 | Multimodal Vigilance Estimation Using Deep LearningabstractThe phenomenon of increasing accidents caused by reduced vigilance does exist. In the future, the high accuracy of vigilance estimation will play a significant role in public transportation safety. We propose a multimodal regression network that consists of multichannel deep autoencoders with subnetwork neurons (MCDAE$_{sn}$). After we define two thresholds of “0.35” and “0.70” from the percentage of eye closure, the output values are in the continuous range of 0–0.35, 0.36–0.70, and 0.71–1 representing the awake state, the tired state, and the drowsy state, respectively. To verify the efficiency of our strategy, we first applied the proposed approach to a single modality. Then, for the multimodality, since the complementary information between forehead electrooculography and electroencephalography features, we found the performance of the proposed approach using features fusion significantly improved, demonstrating the effectiveness and efficiency of our method. Wei Wu 0022, Wei Sun 0028, Q. M. Jonathan Wu, Yimin Yang 0001, Hui Zhang 0023, Wei-Long Zheng, Bao-Liang Lu |
IEEE Trans. Cybern. | 7 |
| 2021 | Plug-and-Play Domain Adaptation for Cross-Subject EEG-based Emotion RecognitionabstractHuman emotion decoding in affective brain-computer interfaces suffers a major setback due to the inter-subject variability of electroencephalography (EEG) signals. Existing approaches usually require amassing extensive EEG data of each new subject, which is prohibitively time-consuming along with poor user experience. To tackle this issue, we divide EEG representations into private components specific to each subject and shared emotional components that are universal to all subjects. According to this representation partition, we propose a plug-and-play domain adaptation method for dealing with the inter-subject variability. In the training phase, subject-invariant emotional representations and private components of source subjects are separately captured by a shared encoder and private encoders. Furthermore, we build one emotion classifier on the shared partition and subjects' individual classifiers on the combination of these two partitions. In the calibration phase, the model only requires few unlabeled EEG data from incoming target subjects to model their private components. Therefore, besides the shared emotion classifier, we have another pipeline to use the knowledge of source subjects through the similarity of private components. In the test phase, we integrate predictions of the shared emotion classifier with those of individual classifiers ensemble after modulation by similarity weights. Experimental results on the SEED dataset show that our model greatly shortens the calibration time within a minute while maintaining the recognition accuracy, all of which make emotion decoding more generalizable and practicable. Li-Ming Zhao, Xu Yan 0007, Bao-Liang Lu |
AAAI | 3 |
| 2021 | Discriminating Surprise and Anger from EEG and Eye Movements with a Graph NetworkabstractEmotion recognition based on EEG and eye movement signals has been studied extensively due to the reliability and stability of signals. The separability of four basic emotions, happy, sad, disgust and fear, has been systematically studied in the existing work. However, there is less research on the emotions of anger and surprise since they are more difficult to be elicited in lab settings. This paper investigates the discrimination ability of EEG and eye movement signals for surprise and anger. To this end, we design a stimulus paradigm that can effectively elicit surprise and anger. We propose a novel Graph Convolutional Network with Channel Attention (GCNCA) to classify three emotions, anger, surprise and neutrality. Experimental results indicate that: a) the proposed GCNCA model has an excellent classification accuracy of 86.47% using EEG and 84.22% using eye movement signals, which are better than other baseline methods; b) EEG and eye movements have a good ability to discriminate surprise and anger, while EEG performs better than eye movements; c) the high-frequency bands of EEG are more distinguishable on classifying surprise and anger than the low-frequency bands; d) there are some differences in neural patterns between surprise and anger, meanwhile critical channels and channel connections of EEG are found. Wei-Bang Jiang, Li-Ming Zhao, Bao-Liang Lu |
BIBM | 4 |
| 2021 | Wasserstein-Distance-Based Multi-Source Adversarial Domain Adaptation for Emotion Recognition and Vigilance EstimationabstractTo build a subject-independent affective model based on electroencephalography (EEG) is a challenging task due to the domain shift problem caused by individual differences in EEG data. In this paper, we prove a new generalization bound based on Wasserstein distance for multi-source classification and regression problems. Based on our bound, we propose two novel Wasserstein-distance-based multi-source adversarial domain adaptation methods (wMADA) for learning domain invariant and task discriminative domain mappings by dynamically aligning different domain mappings. We evaluate our methods on two typical EEG datasets. The experimental results demonstrate that our wMADA methods successfully handle the multi-source domain shift problem in creating subject-independent affective models and outperform the state-of-the-art domain adaptation methods. Bao-Liang Lu |
BIBM | 2 |
| 2021 | Emotion Transformer Fusion: Complementary Representation Properties of EEG and Eye Movements on Recognizing Anger and SurpriseabstractEmotion recognition plays an important role in di-agnosing and treating many mental disorders as well as affective computing. Among six basic emotions, anger and surprise are relatively hard to be elicited in lab settings, and the complementary representation properties of encephalography (EEG) and eye movement signals on recognizing anger and surprise emotions remain unknown. Although the transformer architecture has the ability of parallelism which avoids many sequential operations as recurrent and convolutional layers, the knowledge of its performance and effectiveness on multimodal emotion recognition from EEG and eye movement signals is limited. To tackle these issues, we elaborately design the experiment and stimuli materials to effectively elicit surprise, anger, and neutral emotions, and propose an Emotion Transformer Fusion (ETF) model based on pure attention mechanism. Results of extensive experiments with multiple models on our dataset indicate that the complementary information of EEG and eye movements significantly improves the performance of discriminating anger, surprise and neutral emotions. Meanwhile, our proposed architecture outperforms baseline models with higher parallelism, which proves the capability of Transformer based architecture on multimodal emotion recognition with EEG and eye movement signals. Wei-Bang Jiang, Rui Li 0048, Bao-Liang Lu |
BIBM | 4 |
| 2021 | EEG-Based Human Decision Confidence Measurement Using Graph Neural Networks
Le-Dian Liu, Rui Li 0048, Yu-Zhong Liu, Hua-Liang Li, Bao-Liang Lu |
ICONIP (6) | 5 |
| 2021 | A Cross-subject and Cross-modal Model for Multimodal Emotion Recognition
Xu Yan 0007, Ziyi Li 0003, Li-Ming Zhao, Yu-Zhong Liu, Hua-Liang Li, Bao-Liang Lu |
ICONIP (6) | 7 |
| 2021 | A Multi-Domain Adaptive Graph Convolutional Network for EEG-based Emotion RecognitionabstractAmong all solutions of emotion recognition tasks, electroencephalogram (EEG) is a very effective tool and has received broad attention from researchers. In addition, information across multimedia in EEG often provides a more complete picture of emotions. However, few of the existing studies concurrently incorporate EEG information from temporal domain, frequency domain and functional brain connectivity. In this paper, we propose a Multi-Domain Adaptive Graph Convolutional Network (MD-AGCN), fusing the knowledge of both the frequency domain and the temporal domain to fully utilize the complementary information of EEG signals. MD-AGCN also considers the topology of EEG channels by combining the inter-channel correlations with the intra-channel information, from which the functional brain connectivity can be learned in an adaptive manner. Extensive experimental results demonstrate that our model exceeds state-of-the-art methods in most experimental settings. At the same time, the results show that MD-AGCN could extract complementary domain information and exploit channel relationships for EEG-based emotion recognition effectively. Rui Li 0048, Bao-Liang Lu |
ACM Multimedia | 3 |
| 2021 | Simplifying Multimodal Emotion Recognition with Single Eye Movement ModalityabstractMultimodal emotion recognition has long been a popular topic in affective computing since it significantly enhances the performance compared with that of a single modality. Among all, the combination of electroencephalography (EEG) and eye movement signals is one of the most attractive practices due to their complementarity and objectivity. However, the high cost and inconvenience of EEG signal acquisition severely hamper the popularization of multimodal emotion recognition in practical scenarios, while eye movement signals are much easier to acquire. To increase the feasibility and the generalization ability of emotion decoding without compromising the performance, we propose a generative adversarial network-based framework. In our model, a single modality of eye movements is used as input and it is capable of mapping the information onto multimodal features. Experimental results on SEED series datasets with different emotion categories demonstrate that our model with multimodal features generated by the single eye movement modality maintains competitive accuracies compared to those with multimodality input and drastically outperforms those single-modal emotion classifiers. This illustrates that the model has the potential to reduce the dependence on multimodalities without sacrificing performance which makes emotion recognition more applicable and practicable. Xu Yan 0007, Li-Ming Zhao, Bao-Liang Lu |
ACM Multimedia | 3 |
| 2020 | A Robust Approach to Estimating Vigilance from EEG with Neural ProcessesabstractRobust vigilance estimation is essential for driving and other tasks that require a high degree of vigilance. Many approaches have been applied to estimating vigilance and much endeavor is made to improve the performance of vigilance estimating models. However, most of the existing approaches require test data to have similar quality as training data, which is difficult to be satisfied. In this paper, we adopt neural processes to deal with noise EEG data problem encountered in real scenarios. A publicly available dataset, SEED-VIG, is used to evaluate the performance and robustness of the proposed neural processes method. The dataset includes electroencephalography (EEG) and the corresponding vigilance level annotations during simulated driving. We compared the neural processes method with the existing regression models. The experimental results demonstrate that the neural processes are far better than all others in robustness, meanwhile maintain high accuracy. Chen-Li Yao, Bao-Liang Lu |
BIBM | 2 |
| 2020 | Joint Semi-Supervised Feature Auto-Weighting and Classification Model for EEG-Based Cross-Subject Sleep Quality EvaluationabstractMeasuring the sleep quality is important or even crucial for people who are engaged in dangerous jobs such as the high-speed train drivers. Since the scalp EEG data are generated by the neural activities of the brain cortex, it is collected from subjects with different hours of sleep time (4 hours, 6 hours and 8 hours) to conduct sleep quality evaluation. To suppress the cross-subject variances of EEG data, in this paper, we propose a joint feature auto-weighting and semi-supervised classification model, termed GRLSR, which is formulated by introducing an auto-weighting variable into the least square regression to adaptively and quantitatively measure the importance of each dimension of the feature. Once the model is solved, besides the measurement results, we can use the auto-weighting variable to 1) analyze the importance of each frequency band in sleep quality expression and 2) identify the capacity of different channels connecting to the sleep effect. Therefore, the proposed GRLSR is a pure data-driven computing model for EEG-based cross-subject sleep quality evaluation. Experimental results show its effectiveness. Yong Peng 0001, Qingxi Li, Wanzeng Kong, Bao-Liang Lu, Andrzej Cichocki |
ICASSP | 5 |
| 2020 | Multimodal Emotion Recognition Using Deep Generalized Canonical Correlation Analysis with an Attention MechanismabstractSince multimodal learning is able to take advantage of the complementarity of multimodal signals, the performance of multimodal emotion recognition usually surpasses that based on a single modality. In this paper, we introduce deep generalized canonical correlation analysis with an attention mechanism (DGCCA-AM) to multimodal emotion recognition. This model extends the conventional canonical correlation analysis (CCA) from two modalities to arbitrarily numerous modalities and implements multimodal adaptive fusion with an attention mechanism. By adjusting the weights matrices to maximize the generalized correlation of different modalities, DGCCA-AM extracts emotion-related information from multiple modalities and discards noises. The attention mechanism allows a neural network to learn adaptive fusion weights for different modalities and produces a more effective multimodal fusion and superior emotion recognition performance. We evaluate DGCCA-AM on a public multimodal dataset, SEED-V. Our experimental results demonstrate that DGCCA-AM achieves a state-of-the-art mean accuracy of 82.11% and standard deviation of 2.76% for five emotion classifications with three modalities. Yu-Ting Lan, Wei Liu 0078, Bao-Liang Lu |
IJCNN | 3 |
| 2020 | Emotion Recognition under Sleep Deprivation Using a Multimodal Residual LSTM NetworkabstractEmotion recognition under sleep deprivation is instructive for the study of mental disorders such as major depressive disorder. Previous studies on emotion recognition under sleep deprivation have been mainly based on psychological research techniques. In this paper, we introduce a multilayer weight-sharing multimodal residual LSTM network for emotion recognition under sleep deprivation. The advantage of our proposed method is that it allows for the combination of three different features: the electroencephalography (EEG) single-channel differential entropy (DE) features, EEG functional strength features with topological correlation connectivity, and eye movement features. The experiments under the conditions of sleep deprivation, sleep recovery and baseline are designed and conducted. The experimental results demonstrate that the proposed method significantly enhances the performance compared with the simple concatenation of the features of different modalities, and the best mean accuracies of 86.86% and 82.03% are achieved for four emotions (happiness, sadness, fear, and neutral) in subject-dependent and cross-subject emotion recognition tasks under 30 hours of sleep deprivation, respectively. The classification accuracy of the happiness emotion is obviously impaired under sleep deprivation, indicating that sleep deprivation impairs the stimulation of the happiness emotion, and one night of sleep recovery can reactivate the elicitation of the happiness emotion to the baseline level. Furthermore, we study the brain neural patterns of the four emotional states. The prefrontal area becomes less activated for the happiness emotion and sadness emotion in the gamma band under sleep deprivation, while the neural pattern of the fear emotion is highly robust with respect to sleep deprivation. Le-Yan Tao, Bao-Liang Lu |
IJCNN | 2 |
| 2020 | Towards Scale-Invariant Graph-related Problem Solving by Iterative Homogeneous GNNsabstractCurrent graph neural networks (GNNs) lack generalizability with respect to scales (graph sizes, graph diameters, edge weights, etc..) when solving many graph analysis problems. Taking the perspective of synthesizing graph theory programs, we propose several extensions to address the issue. First, inspired by the dependency of iteration number of common graph theory algorithms on graph size, we learn to terminate the message passing process in GNNs adaptively according to the computation progress. Second, inspired by the fact that many graph theory algorithms are homogeneous with respect to graph weights, we introduce homogeneous transformation layers that are universal homogeneous function approximators, to convert ordinary GNNs to be homogeneous. Experimentally, we show that our GNN can be trained from small-scale graphs but generalize well to large-scale graphs for a number of basic graph theory problems. It also shows generalizability for applications of multi-body physical simulation and image-based navigation problems. Hao Tang 0008, Zhiao Huang, Jiayuan Gu, Bao-Liang Lu, Hao Su 0001 |
NeurIPS | 4 |
| 2020 | Driver sleepiness detection from EEG and EOG signals using GAN and LSTM networks
Yingying Jiao, Yini Deng, Bao-Liang Lu |
Neurocomputing | 4 |
| 2020 | Vigilance Estimation Using a Wearable EOG Device in Real Driving EnvironmentabstractVigilance decrement in driving tasks has been reported to be a major factor in fatal accidents and could severely endanger public transportation safety. However, efficient approaches for estimating vigilance in real driving environment are still lacking. In this paper, we propose a novel approach for implementing continuous vigilance estimation using forehead electrooculograms (EOGs) acquired by wearable dry electrodes in both simulated and real driving environments. To improve the feasibility of this approach for real-world applications, a forehead EOG-based electrode placement with only four electrodes is designed. Flexible dry electrodes and an acquisition board are integrated as a wearable device for recording EOGs. Twenty and ten subjects participated in the simulated and real-world driving environment experiments, respectively. Accurate eye movement parameters from eye-tracking glasses are extracted to calculate the PERCLOS index for vigilance annotation. This is because the vigilance state is a temporally dynamic process, and a continuous conditional random field and a continuous conditional neural field are introduced to construct more accurate vigilance estimation models. To evaluate the efficiency of our system, systematic experiments are performed in real scenarios under various illumination and weather conditions following laboratory simulations as preliminary studies. The experimental results demonstrate that the wearable dry electrode prototype, which has a relatively comfortable forehead setup, can efficiently capture vigilance dynamics. The best mean correlation coefficients achieved by our proposed approach are 71.18% and 66.20% in laboratory simulations and real-world driving environments, respectively. The cross-environment experiments are performed to evaluate the simulated-to-real generalization and a best mean correlation coefficient of 53.96% is achieved. Wei-Long Zheng, Kunpeng Gao, Gang Li 0011, Wei Liu 0078, Chao Liu 0025, Guoxing Wang, Bao-Liang Lu |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2019 | A Cross-Culture Study on Multimodal Emotion Recognition Using Deep Learning
Wei Liu 0078, Bao-Liang Lu |
ICONIP (4) | 5 |
| 2019 | Reducing the Subject Variability of EEG Signals with Adversarial Domain Generalization
Bo-Qun Ma, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (1) | 4 |
| 2019 | Depersonalized Cross-Subject Vigilance Estimation with Adversarial Domain GeneralizationabstractSubject variability is a major obstacle to vigilance estimation. The conventional subject-specific models fail to perform well on unknown subjects. The existing studies mainly focus on domain adaptation utilizing labeled/unlabeled subject-specific data. However, it is still expensive and inconvenient to collect task-specific data from unknown subjects in some real-world applications. In this paper, we introduce domain generalization methods for building vigilance estimation models without requiring any information from the unknown subjects. We first generalize the structure of Domain Adversarial Neural Network (DANN) into Domain Generalization (DG-DANN), and then propose a novel adversarial structure called Domain Residual Network (DResNet). We compare a popular domain generalization method, Domain-Invariant Component Analysis (DICA), with our proposed approach. In terms of the estimation accuracy and generalization ability, we designed two different settings for evaluation experiments on a public dataset called SEED-VIG. Experimental results indicate that our new model achieves comparable accuracy but more stable performance without using additional information from the unknown subjects in comparison with the state-of-the-art domain adaptation methods. Furthermore, domain generalization models also perform well on the tasks with multiple unknown subjects. Bo-Qun Ma, Bao-Liang Lu |
IJCNN | 4 |
| 2019 | A GAN-Based Data Augmentation Method for Multimodal Emotion Recognition
Lizhen Zhu, Bao-Liang Lu |
ISNN (1) | 3 |
| 2019 | Emotion Recognition using Multimodal Residual LSTM NetworkabstractVarious studies have shown that the temporal information captured by conventional long-short-term memory (LSTM) networks is very useful for enhancing multimodal emotion recognition using encephalography (EEG) and other physiological signals. However, the dependency among multiple modalities and high-level temporal-feature learning using deeper LSTM networks is yet to be investigated. Thus, we propose a multimodal residual LSTM (MMResLSTM) network for emotion recognition. The MMResLSTM network shares the weights across the modalities in each LSTM layer to learn the correlation between the EEG and other physiological signals. It contains both the spatial shortcut paths provided by the residual network and temporal shortcut paths provided by LSTM for efficiently learning emotion-related high-level features. The proposed network was evaluated using a publicly available dataset for EEG-based emotion recognition, DEAP. The experimental results indicate that the proposed MMResLSTM network yielded a promising result, with a classification accuracy of 92.87% for arousal and 92.30% for valence. Jia-Xin Ma, Hao Tang 0008, Wei-Long Zheng, Bao-Liang Lu |
ACM Multimedia | 4 |
| 2019 | Identifying Stable Patterns over Time for Emotion Recognition from EEGabstractIn this paper, we investigate stable patterns of electroencephalogram (EEG) over time for emotion recognition using a machine learning approach. Up to now, various findings of activated patterns associated with different emotions have been reported. However, their stability over time has not been fully investigated yet. In this paper, we focus on identifying EEG stability in emotion recognition. We systematically evaluate the performance of various popular feature extraction, feature selection, feature smoothing and pattern classification methods with the DEAP dataset and a newly developed dataset called SEED for this study. Discriminative Graph regularized Extreme Learning Machine with differential entropy features achieves the best average accuracies of 69.67 and 91.07 percent on the DEAP and SEED datasets, respectively. The experimental results indicate that stable patterns exhibit consistency across sessions; the lateral temporal areas activate more for positive emotions than negative emotions in beta and gamma bands; the neural patterns of neutral emotions have higher alpha responses at parietal and occipital sites; and for negative emotions, the neural patterns have significant higher delta responses at parietal and occipital sites and higher gamma responses at prefrontal sites. The performance of our emotion recognition models shows that the neural patterns are relatively stable within and between sessions. Wei-Long Zheng, Jia-Yi Zhu, Bao-Liang Lu |
IEEE Trans. Affect. Comput. | 3 |
| 2019 | EmotionMeter: A Multimodal Framework for Recognizing Human EmotionsabstractIn this paper, we present a multimodal emotion recognition framework called EmotionMeter that combines brain waves and eye movements. To increase the feasibility and wearability of EmotionMeter in real-world applications, we design a six-electrode placement above the ears to collect electroencephalography (EEG) signals. We combine EEG and eye movements for integrating the internal cognitive states and external subconscious behaviors of users to improve the recognition accuracy of EmotionMeter. The experimental results demonstrate that modality fusion with multimodal deep neural networks can significantly enhance the performance compared with a single modality, and the best mean accuracy of 85.11% is achieved for four emotions (happy, sad, fear, and neutral). We explore the complementary characteristics of EEG and eye movements for their representational capacities and identify that EEG has the advantage of classifying happy emotion, whereas eye movements outperform EEG in recognizing fear emotion. To investigate the stability of EmotionMeter over time, each subject performs the experiments three times on different days. EmotionMeter obtains a mean recognition accuracy of 72.39% across sessions with the six-electrode EEG and eye movement features. These experimental results demonstrate the effectiveness of EmotionMeter within and between sessions. Wei-Long Zheng, Wei Liu 0078, Bao-Liang Lu, Andrzej Cichocki |
IEEE Trans. Cybern. | 4 |
| 2018 | Driver Sleepiness Detection Using LSTM Neural Network
Yini Deng, Yingying Jiao, Bao-Liang Lu |
ICONIP (4) | 3 |
| 2018 | Cross-Subject Emotion Recognition Using Deep Adaptation Networks
Wei-Long Zheng, Bao-Liang Lu |
ICONIP (5) | 4 |
| 2018 | WGAN Domain Adaptation for EEG-Based Emotion Recognition
Si-Yang Zhang, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (5) | 4 |
| 2018 | Multi-view Emotion Recognition Using Deep Canonical Correlation Analysis
Jie-Lin Qiu, Wei Liu 0078, Bao-Liang Lu |
ICONIP (5) | 3 |
| 2018 | Active Feedback Framework with Scan-Path Clustering for Deep Affective Models
Li-Ming Zhao, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (2) | 4 |
| 2018 | Multimodal Vigilance Estimation with Adversarial Domain Adaptation NetworksabstractRobust vigilance estimation during driving is very crucial in preventing traffic accidents. Many approaches have been proposed for vigilance estimation. However, most of the approaches require collecting subject-specific labeled data for calibration which is high-cost for real-world applications. To solve this problem, domain adaptation methods can be used to align distributions of source subject features (source domain) and new subject features (target domain). By reusing existing data from other subjects, no labeled data of new subjects is required to train models. In this paper, our goal is to apply adversarial domain adaptation networks to cross-subject vigilance estimation. We adopt two kinds of recently proposed adversarial domain adaptation networks and compare their performance with those of several traditional domain adaptation methods and the baseline without domain adaptation. A publicly available dataset, SEED-VIG, is used to evaluate the methods. The dataset includes electroencephalography (EEG) and electrooculography (EOG) signals, as well as the corresponding vigilance level annotations during simulated driving. Compared with the baseline, both adversarial domain adaptation networks achieve improvements over 10% in terms of Pearson's correlation coefficient. In addition, both methods considerably outperform the traditional domain adaptation methods. Wei-Long Zheng, Bao-Liang Lu |
IJCNN | 3 |
| 2018 | Sleep Quality Estimation with Adversarial Domain Adaptation: From Laboratory to Real ScenarioabstractPrevious studies on EEG-based last-night sleep quality estimation mainly focus on evaluation with data from laboratory experiments. However, due to the reality gap constituted with device performance, subject groups, experiment settings and controlled conditions, the models trained solely on laboratory data cannot generalize well to real scenarios. In this work, we investigate the sleep quality estimation for high-speed train drivers as an instance of real-scenario application. Domain adaptation models are adopted to deal with the individual differences across subjects when modeling and testing with the real-scenario data. As it is usually difficult and costly to acquire data and annotate them in real scenarios, the high-quality data in laboratory conditions are used for model trainings. Knowledge from simulation is transferred to reality with domain adaptation methods. A novel approach called Domain Adversarial Neural Network (DANN) is adopted. DANN learns domain independent features through deep networks with an adversarial architecture. The experimental results indicate that DANN outperforms other state-of-the-art methods and achieves 19.55% and 23.50% improvements in terms of accuracy on the cross-subject and cross- scenario tasks, respectively, in comparison with the baseline SVM model. Jia-Jun Tong, Bo-Qun Ma, Wei-Long Zheng, Bao-Liang Lu, Xiao-Qi Song, Shi-Wei Ma |
IJCNN | 5 |
| 2018 | Semi-supervised Deep Generative Modelling of Incomplete Multi-Modality Emotional DataabstractThere are threefold challenges in emotion recognition. First, it is difficult to recognize human's emotional states only considering a single modality. Second, it is expensive to manually annotate the emotional data. Third, emotional data often suffers from missing modalities due to unforeseeable sensor malfunction or configuration issues. In this paper, we address all these problems under a novel multi-view deep generative framework. Specifically, we propose to model the statistical relationships of multi-modality emotional data using multiple modality-specific generative networks with a shared latent space. By imposing a Gaussian mixture assumption on the posterior approximation of the shared latent variables, our framework can learn the joint deep representation from multiple modalities and evaluate the importance of each modality simultaneously. To solve the labeled-data-scarcity problem, we extend our multi-view model to semi-supervised learning scenario by casting the semi-supervised classification problem as a specialized missing data imputation task. To address the missing-modality problem, we further extend our semi-supervised multi-view model to deal with incomplete data, where a missing view is treated as a latent variable and integrated out during inference. This way, the proposed overall framework can utilize all available (both labeled and unlabeled, as well as both complete and incomplete) data to improve its generalization ability. The experiments conducted on two real multi-modal emotion datasets demonstrated the superiority of our framework. Changde Du, Changying Du, Hao Wang 0005, Jinpeng Li 0002, Wei-Long Zheng, Bao-Liang Lu, Huiguang He |
ACM Multimedia | 6 |
| 2018 | Graph-Based Bilingual Word Embedding for Statistical Machine TranslationabstractBilingual word embedding has been shown to be helpful for Statistical Machine Translation (SMT). However, most existing methods suffer from two obvious drawbacks. First, they only focus on simple contexts such as an entire document or a fixed-sized sliding window to build word embedding and ignore latent useful information from the selected context. Second, the word sense but not the word should be the minimal semantic unit; however, most existing methods still use word representation. To overcome these drawbacks, this article presents a novel Graph-Based Bilingual Word Embedding (GBWE) method that projects bilingual word senses into a multidimensional semantic space. First, a bilingual word co-occurrence graph is constructed using the co-occurrence and pointwise mutual information between the words. Then, maximum complete subgraphs (cliques), which play the role of a minimal unit for bilingual sense representation, are dynamically extracted according to the contextual information. Consequently, correspondence analysis, principal component analyses, and neural networks are used to summarize the clique-word matrix into lower dimensions to build the embedding model. Without contextual information, the proposed GBWE can be applied to lexical translation. In addition, given contextual information, GBWE is able to give a dynamic solution for bilingual word representations, which can be applied to phrase translation and generation. Empirical results show that GBWE can enhance the performance of lexical translation, as well as Chinese/French-to-English and Chinese-to-Japanese phrase-based SMT tasks (IWSLT, NTCIR, NIST, and WAT). Rui Wang 0015, Hai Zhao 0001, Sabine Ploux, Bao-Liang Lu, Masao Utiyama, Eiichiro Sumita |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2017 | Detecting driver sleepiness from EEG alpha wave during daytime drivingabstractDrowsy driving is the main reason for sleep-related crashes. We have observed that an alpha wave attenuation-disappearance phenomenon and a typical alpha blocking phenomenon commonly exist in the eye closure events during daytime simulated driving experiments. These two alpha-related phenomena prove to respectively represent two different sleepiness levels: the sleep onset and the relaxed wakefulness. Therefore, we propose a novel algorithm for tracking the alpha wave change and detecting the two alpha-related phenomena in real-time for recognizing driver sleepiness. Our proposed algorithm adopts continuous wavelet transform for charactering the signal change and support vector machine for classification. The experimental results indicate that the algorithm is able to detect the start and end points of alpha waves during eye-closed period and distinguish the two types of end points of alpha waves in two alpha-related phenomena with high sensitivity and precision. The eye-closed period detected by our algorithm with alpha waves has a high overlapping rate with that marked by human experts. The main contributions of the proposed algorithm are twofold: to detect alpha waves during the eye-closed period in real-time and serve as an indicator for judging the current sleepiness level as the sleep onset or relaxed wakefulness at the end points of alpha waves. Yingying Jiao, Bao-Liang Lu |
BIBM | 2 |
| 2017 | Multimodal Emotion Recognition Using Deep Neural Networks
Hao Tang 0008, Wei Liu 0078, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (4) | 4 |
| 2017 | Identifying Gender Differences in Multimodal Emotion Recognition Using Bimodal Deep AutoEncoder
Wei-Long Zheng, Wei Liu 0078, Bao-Liang Lu |
ICONIP (4) | 4 |
| 2017 | Investigating Gender Differences of Brain Areas in Emotion Recognition Using LSTM Neural Network
Wei-Long Zheng, Wei Liu 0078, Bao-Liang Lu |
ICONIP (4) | 4 |
| 2017 | EEG-Based Sleep Quality Evaluation with Deep Transfer Learning
Xing-Zan Zhang, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (4) | 3 |
| 2017 | Emotion Annotation Using Hierarchical Aligned Cluster Analysis
Wei-Ye Zhao, Ting Ji, Qian Ji, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (4) | 6 |
| 2017 | Bootstrapping integrative hypothesis test for identifying biomarkers that differentiates lung cancer and chronic obstructive pulmonary disease
Kai-Ming Jiang, Ya-Jing Chen, Jin-Xiong Lv, Bao-Liang Lu, Lei Xu 0001 |
Neurocomputing | 4 |
| 2017 | Discriminative extreme learning machine with supervised sparsity preserving for image classification
Yong Peng 0001, Bao-Liang Lu |
Neurocomputing | 2 |
| 2017 | Robust structured sparse representation via half-quadratic optimization for face recognition
Yong Peng 0001, Bao-Liang Lu |
Multim. Tools Appl. | 2 |
| 2017 | Online Depth Image-Based Object Tracking with Sparse Representation and Object Detection
Wei-Long Zheng, Shan-Chun Shen, Bao-Liang Lu |
Neural Process. Lett. | 3 |
| 2016 | On the Reducibility of Submodular FunctionsabstractThe scalability of submodular optimization methods is critical for their usability in practice. In this paper, we study the reducibility of submodular functions, a property that enables us to reduce the solution space of submodular optimization problems without performance loss. We introduce the concept of reducibility using marginal gains. Then we show that by adding perturbation, we can endow irreducible functions with reducibility, based on which we propose the perturbation-reduction optimization framework. Our theoretical analysis proves that given the perturbation scales, the reducibility gain could be computed, and the performance loss has additive upper bounds. We further conduct empirical studies and the results demonstrate that our proposed framework significantly accelerates existing optimization methods for irreducible submodular functions with a cost of only small performance losses. Jincheng Mei, Bao-Liang Lu |
AISTATS | 3 |
| 2016 | Connecting Phrase based Statistical Machine Translation AdaptationabstractAlthough more additional corpora are now available for Statistical Machine Translation (SMT), only the ones which belong to the same or similar domains of the original corpus can indeed enhance SMT performance directly. A series of SMT adaptation methods have been proposed to select these similar-domain data, and most of them focus on sentence selection. In comparison, phrase is a smaller and more fine grained unit for data selection, therefore we propose a straightforward and efficient connecting phrase based adaptation method, which is applied to both bilingual phrase pair and monolingual n-gram adaptation. The proposed method is evaluated on IWSLT/NIST data sets, and the results show that phrase based SMT performances are significantly improved (up to +1.6 in comparison with phrase based SMT baseline system and +0.9 in comparison with existing methods). Rui Wang 0015, Hai Zhao 0001, Bao-Liang Lu, Masao Utiyama, Eiichiro Sumita |
COLING | 3 |
| 2016 | Emotion Recognition Using Multimodal Deep Learning
Wei Liu 0078, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (2) | 3 |
| 2016 | Continuous Vigilance Estimation Using LSTM Neural Networks
Wei-Long Zheng, Wei Liu 0078, Bao-Liang Lu |
ICONIP (2) | 4 |
| 2016 | A Bilingual Graph-Based Semantic Model for Statistical Machine Translation
Rui Wang 0015, Hai Zhao 0001, Sabine Ploux, Bao-Liang Lu, Masao Utiyama |
IJCAI | 4 |
| 2016 | Personalizing EEG-Based Affective Models with Transfer Learning
Wei-Long Zheng, Bao-Liang Lu |
IJCAI | 2 |
| 2016 | Driving fatigue detection with fusion of EEG and forehead EOGabstractIn this paper, we fuse EEG and forehead EOG to detect drivers' fatigue level by using discriminative graph regularized extreme learning machine (GELM). Twenty-one healthy subjects including twelve men and nine women participate in our driving simulation experiments. Two fusion strategies are adopted: feature level fusion (FLF) and decision level fusion (DLF). PERCLOS (the percentage of eye closure) is calculated by using the eye movement data recorded by eye tracking glasses as the indicator of drivers' fatigue level. The prediction correlation coefficient and root mean square error (RMSE) between the estimated fatigue level and the real fatigue level are both used to evaluate the performance of single modality and fusion modality. A comparative study on modality performance is conducted between GELM and support vector machine (SVM). The experimental results show that fusion modality can improve the performance of driving fatigue detection with a higher prediction correlation coefficient and a lower RMSE value in comparison with solely using EEG or forehead EOG. And FLF achieves better performance than DLF. GELM is more suitable for driving fatigue detection than SVM. Moreover, feature level fusion with GELM achieves the best performance with the prediction correlation coefficient of 0.8080 and the RMSE value of 0.0712 on average. Xue-Qin Huo, Wei-Long Zheng, Bao-Liang Lu |
IJCNN | 3 |
| 2016 | Measuring sleep quality from EEG with machine learning approachesabstractThis study aims at measuring last-night sleep quality from electroencephalography (EEG). We design a sleep experiment to collect waking EEG signals from eight subjects under three different sleep conditions: 8 hours sleep, 6 hours sleep, and 4 hours sleep. We utilize three machine learning approaches, k-Nearest Neighbor (kNN), support vector machine (SVM), and discriminative graph regularized extreme learning machine (GELM), to classify extracted EEG features of power spectral density (PSD). The accuracies of these three classifiers without feature selection are 36.68%, 48.28%, 62.16%, respectively. By using minimal-redundancy-maximal-relevance (MRMR) algorithm and the brain topography, the classification accuracy of GELM with 9 features is improved largely and increased to 83.57% in average. To investigate critical frequency bands for measuring sleep quality, we examine the features of each band and observe their energy changing. The experimental results indicate that Gamma band is more relevant to measuring sleep quality. Wei-Long Zheng, Hai-Wei Ma, Bao-Liang Lu |
IJCNN | 4 |
| 2016 | Discriminative manifold extreme learning machine and applications to image and EEG signal classification
Yong Peng 0001, Bao-Liang Lu |
Neurocomputing | 2 |
| 2016 | An unsupervised discriminative extreme learning machine and its applications to data clustering
Yong Peng 0001, Wei-Long Zheng, Bao-Liang Lu |
Neurocomputing | 3 |
| 2016 | Converting Continuous-Space Language Models into N-gram Language Models with Efficient Bilingual Pruning for Statistical Machine TranslationabstractThe Language Model (LM) is an essential component of Statistical Machine Translation (SMT). In this article, we focus on developing efficient methods for LM construction. Our main contribution is that we propose a Natural N -grams based Converting (NNGC) method for transforming a Continuous-Space Language Model (CSLM) to a Back-off N -gram Language Model (BNLM). Furthermore, a Bilingual LM Pruning (BLMP) approach is developed for enhancing LMs in SMT decoding and speeding up CSLM converting. The proposed pruning and converting methods can convert a large LM efficiently by working jointly. That is, a LM can be effectively pruned before it is converted from CSLM without sacrificing performance, and further improved if an additional corpus contains out-of-domain information. For different SMT tasks, our experimental results indicate that the proposed NNGC and BLMP methods outperform the existing counterpart approaches significantly in BLEU and computational cost. Rui Wang 0015, Masao Utiyama, Isao Goto, Eiichiro Sumita, Hai Zhao 0001, Bao-Liang Lu |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 6 |
| 2015 | On Unconstrained Quasi-Submodular Function OptimizationabstractWith the extensive application of submodularity, its generalizations are constantly being proposed. However, most of them are tailored for special problems. In this paper, we focus on quasi-submodularity, a universal generalization, which satisfies weaker properties than submodularity but still enjoys favorable performance in optimization. Similar to the diminishing return property of submodularity, we first define a corresponding property called the single sub-crossing, then we propose two algorithms for unconstrained quasi-submodular function minimization and maximization, respectively. The proposed algorithms return the reduced lattices in O(n) iterations, and guarantee the objective function values are strictly monotonically increased or decreased after each iteration. Moreover, any local and global optima are definitely contained in the reduced lattices. Experimental results verify the effectiveness and efficiency of the proposed algorithms on lattice reduction. Jincheng Mei, Bao-Liang Lu |
AAAI | 3 |
| 2015 | Transfer components between subjects for EEG-based emotion recognitionabstractAddressing the structural and functional variability between subjects for robust affective brain-computer interface (aBCI) is challenging but of great importance, since the calibration phase for aBCI is time-consuming. In this paper, we propose a subject transfer framework for electroencephalogram (EEG)-based emotion recognition via component analysis. We compare two state-of-the-art subspace projecting approaches called transfer component analysis (TCA) and kernel principle component analysis (KPCA) for subject transfer. The main idea is to learn a set of transfer components underlying source domain (source subjects) and target domain (target subject). When projected to this subspace, the difference of feature distributions of both domains can be reduced. From the experiments, we show that the two proposed approaches, TCA and KPCA, can achieve an improvement on performance with the best mean accuracies of 71.80% and 79.83%, respectively, in comparison of the baseline of 58.95%. The significant improvement shows the feasibility and efficiency of our approaches for subject transfer emotion recognition from EEG signals. Wei-Long Zheng, Yong-Qi Zhang, Jia-Yi Zhu, Bao-Liang Lu |
ACII | 4 |
| 2015 | Graph regularized non-negative local coordinate factorization with pairwise constraints for image representationabstractChen et al. proposed a non-negative local coordinate factorization algorithm for feature extraction (NLCF) [1], which incorporated the local coordinate constraint into non-negative matrix factorization (NMF). However, NLCF is actually a unsupervised method without making use of prior information of problems in hand. In this paper, we propose a novel graph regularized non-negative local coordinate factorization with pairwise constraints algorithm (PCGNLCF) for image representation. PCGNLCF incorporates pairwise constraints and graph Laplacian into NLCF. More specifically, we expect that data points having pairwise must-link constraints will have the similar coordinates as much as possible, while data points with pairwise cannot-link constraints will have distinct coordinates as much as possible. Experimental results show the effectiveness of our proposed method in comparison to the state-of-the-art algorithms on several real-world applications. Yangcheng He, Hongtao Lu 0001, Bao-Liang Lu |
ICME | 3 |
| 2015 | Distance Preserving Marginal Hashing for image retrievalabstractHashing for image retrieval has attracted lots of attentions in recent years due to its fast computational speed and storage efficiency. Many existing hashing methods obtain the hashing functions through mapping neighbor items to similar codes, while ignoring the non-neighbor items. One exception is the Local Linear Spectral Hashing (LLSH), which introduces negative values into the local affinity matrix to map non-neighbor images to non-similar codes. However, setting 10th percentile distance in affinity matrix as a threshold, which is used to judge neighbors and non-neighbors, is not reasonable. In this paper, we propose a novel unsupervised hashing method called Distance Preserving Marginal Hashing (DPMH) which not only makes the average Hamming distance minimized for the intra-cluster pairs and maximized for the inter-cluster pairs, but also preserves the distance of non-neighbor points. Furthermore, we adopt an efficient sequential procedure to learn the hashing functions. The experimental results on two large-scale benchmark datasets demonstrate the effectiveness and efficiency of our method over other state-of-the-art unsupervised methods. Hongtao Lu 0001, Bao-Liang Lu |
ICME | 5 |
| 2015 | Intensity-Depth Face Alignment Using Cascade Shape Regression
Bao-Liang Lu |
ICONIP (4) | 2 |
| 2015 | Transfer Components Between Subjects for EEG-based Driving Fatigue Detection
Yong-Qi Zhang, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (4) | 3 |
| 2015 | Combining Eye Movements and EEG to Enhance Emotion Recognition
Wei-Long Zheng, Bao-Liang Lu |
IJCAI | 4 |
| 2015 | English to Chinese Translation: How Chinese Character Matters
Rui Wang 0015, Hai Zhao 0001, Bao-Liang Lu |
PACLIC | 3 |
| 2015 | Discriminative graph regularized extreme learning machine and its application to face recognition
Yong Peng 0001, Suhang Wang, Xianzhong Long, Bao-Liang Lu |
Neurocomputing | 4 |
| 2015 | Hybrid learning clonal selection algorithm
Yong Peng 0001, Bao-Liang Lu |
Inf. Sci. | 2 |
| 2015 | Enhanced low-rank representation via sparse manifold adaption for semi-supervised learning
Yong Peng 0001, Bao-Liang Lu, Suhang Wang |
Neural Networks | 2 |
| 2015 | Graph Based Semi-Supervised Learning via Structure Preserving Low-Rank Representation
Yong Peng 0001, Xianzhong Long, Bao-Liang Lu |
Neural Process. Lett. | 3 |
| 2015 | Bilingual Continuous-Space Language Model Growing for Statistical Machine TranslationabstractLarger n-gram language models (LMs) perform better in statistical machine translation (SMT). However, the existing approaches have two main drawbacks for constructing larger LMs: 1) it is not convenient to obtain larger corpora in the same domain as the bilingual parallel corpora in SMT; 2) most of the previous studies focus on monolingual information from the target corpora only, and redundant n-grams have not been fully utilized in SMT. Nowadays, continuous-space language model (CSLM), especially neural network language model (NNLM), has been shown great improvement in the estimation accuracies of the probabilities for predicting the target words. However, most of these CSLM and NNLM approaches still consider monolingual information only or require additional corpus. In this paper, we propose a novel neural network based bilingual LM growing method. Compared to the existing approaches, the proposed method enables us to use bilingual parallel corpus for LM growing in SMT. The results show that our new method outperforms the existing approaches on both SMT performance and computational efficiency significantly. Rui Wang 0015, Hai Zhao 0001, Bao-Liang Lu, Masao Utiyama, Eiichiro Sumita |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Large-Scale Nyström Kernel Matrix Approximation Using Randomized SVDabstractThe Nyström method is an efficient technique for the eigenvalue decomposition of large kernel matrices. However, to ensure an accurate approximation, a sufficient number of columns have to be sampled. On very large data sets, the singular value decomposition (SVD) step on the resultant data submatrix can quickly dominate the computations and become prohibitive. In this paper, we propose an accurate and scalable Nyström scheme that first samples a large column subset from the input matrix, but then only performs an approximate SVD on the inner submatrix using the recent randomized low-rank matrix approximation algorithms. Theoretical analysis shows that the proposed algorithm is as accurate as the standard Nyström method that directly performs a large SVD on the inner submatrix. On the other hand, its time complexity is only as low as performing a small SVD. Encouraging results are obtained on a number of large-scale data sets for low-rank approximation. Moreover, as the most computational expensive steps can be easily distributed and there is minimal data transfer among the processors, significant speedup can be further obtained with the use of multiprocessor and multi-GPU systems. Mu Li 0001, Wei Bi, James T. Kwok, Bao-Liang Lu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2014 | Neural Network Based Bilingual Language Model Growing for Statistical Machine TranslationabstractSince larger n-gram Language Model (LM) usually performs better in Statistical Machine Translation (SMT), how to construct efficient large LM is an important topic in SMT.However, most of the existing LM growing methods need an extra monolingual corpus, where additional LM adaption technology is necessary.In this paper, we propose a novel neural network based bilingual LM growing method, only using the bilingual parallel corpus in SMT.The results show that our method can improve both the perplexity score for LM evaluation and BLEU score for SMT, and significantly outperforms the existing LM growing methods without extra corpus. Rui Wang 0015, Hai Zhao 0001, Bao-Liang Lu, Masao Utiyama, Eiichiro Sumita |
EMNLP | 3 |
| 2014 | EEG-based emotion classification using deep belief networksabstractIn recent years, there are many great successes in using deep architectures for unsupervised feature learning from data, especially for images and speech. In this paper, we introduce recent advanced deep learning models to classify two emotional categories (positive and negative) from EEG data. We train a deep belief network (DBN) with differential entropy features extracted from multichannel EEG as input. A hidden markov model (HMM) is integrated to accurately capture a more reliable emotional stage switching. We also compare the performance of the deep models to KNN, SVM and Graph regularized Extreme Learning Machine (GELM). The average accuracies of DBN-HMM, DBN, GELM, SVM, and KNN in our experiments are 87.62%, 86.91%, 85.67%, 84.08%, and 69.66%, respectively. Our experimental results show that the DBN and DBN-HMM models improve the accuracy of EEG-based emotion classification in comparison with the state-of-the-art methods. Wei-Long Zheng, Jia-Yi Zhu, Yong Peng 0001, Bao-Liang Lu |
ICME | 4 |
| 2014 | Saliency Level Set Evolution
Jincheng Mei, Bao-Liang Lu |
ICONIP (2) | 2 |
| 2014 | Online Object Tracking Based on Depth Image with Sparse Coding
Shan-Chun Shen, Wei-Long Zheng, Bao-Liang Lu |
ICONIP (3) | 3 |
| 2014 | Recognizing slow eye movement for driver fatigue detection with machine learning approachabstractSlow eye movement (SEM) regarded as a sign of onset of sleep is very significant for detecting driver fatigue, but its characteristics and detection algorithm have been rarely involved in the study of driver fatigue detection. In this study, some new features were extracted based on wavelet singularity analysis and statistics to detect SEMs. Six subjects participated in this simulated driving experiment, and for each subject, a more than 2 hours electro-oculogram (EOG) session was recorded. Each session was divided into SEM epochs and non-SEM epochs according to the common judgments made by the two of three experts by the visual recognition criteria of SEMs. Regarding the problem of detecting SEMs as an imbalance classification problem, and through the under-sampling and over-sampling methods a 2s horizontal electro-oculogram (HEO) signal could finally be recognized as the category of SEMs or non-SEMs with the classifiers SVM, GELM, and KNN respectively. Results prove that the proposed features was a little better than the wavelet energy features, and through the combination of the wavelet energy features and the new features based on wavelet singularity analysis and statistics, the classification results were improved obviously. Yingying Jiao, Yong Peng 0001, Bao-Liang Lu, Shanguang Chen |
IJCNN | 3 |
| 2014 | EOG-based drowsiness detection using convolutional neural networksabstractThis study provides a new application of convolutional neural networks for drowsiness detection based on electrooculography (EOG) signals. Drowsiness is charged to be one of the major causes of traffic accidents. Such application is helpful to reduce losses of casualty and property. Most attempts at drowsiness detection based on EOG involve a feature extraction step, which is accounted as time-consuming task, and it is difficult to extract effective features. In this paper, an unsupervised learning is proposed to estimate driver fatigue based on EOG. A convolutional neural network with a linear regression layer is applied to EOG signals in order to avoid using of manual features. With a postprocessing step of linear dynamic system (LDS), we are able to capture the physiological status shifting. The performance of the proposed model is evaluated by the correlation coefficients between the final outputs and the local error rates of the subjects. Compared with the results of a manual ad-hoc feature extraction approach, our method is proven to be effective for drowsiness detection. Xuemin Zhu, Wei-Long Zheng, Bao-Liang Lu, Shanguang Chen |
IJCNN | 3 |
| 2014 | EEG-based emotion recognition using discriminative graph regularized extreme learning machineabstractThis study aims at finding the relationship between EEG signals and human emotional states. Movie clips are used as stimuli to evoke positive, neutral and negative emotions of subjects. We introduce a new effective classifier named discriminative graph regularized extreme learning machine (GELM) for EEG-based emotion recognition. The average classification accuracy of GELM using differential entropy (DE) features on the whole five frequency bands is 80.25%, while the accuracy of SVM is 76.62%. These results indicate that GELM is more suitable for emotion recognition than SVM. Additionally, the accuracies of GELM using DE features on Beta and Gamma bands are 79.07%, 79.93% respectively. This suggests that these two bands are more relevant to emotion. The experimental results indicate that the EEG patterns for emotion are generally stable among different experiments and subjects. By using minimal-redundancy-maximal-relevance (MRMR) algorithm and correlation coefficients to select effective features, we get the distribution of top 20 subject-independent features and build a manifold model to monitor the trajectory of emotion changes with time. Jia-Yi Zhu, Wei-Long Zheng, Yong Peng 0001, Ruo-Nan Duan, Bao-Liang Lu |
IJCNN | 5 |
| 2014 | An integrated Gaussian mixture model to estimate vigilance level based on EEG recordings
Jing-Nan Gu, Hongtao Lu 0001, Bao-Liang Lu |
Neurocomputing | 3 |
| 2014 | Parallelized extreme learning machine ensemble based on min-max modular network
Xiaolin Wang 0002, Hai Zhao 0001, Bao-Liang Lu |
Neurocomputing | 4 |
| 2014 | Emotional state classification from EEG data using machine learning approach
Dan Nie, Bao-Liang Lu |
Neurocomputing | 3 |
| 2014 | A Meta-Top-Down Method for Large-Scale Hierarchical ClassificationabstractRecent large-scale hierarchical classification tasks typically have tens of thousands of classes on which the most widely used approach to multiclass classification--one-versus-rest--becomes intractable due to computational complexity. The top-down methods are usually adopted instead, but they are less accurate because of the so-called error-propagation problem in their classifying phase. To address this problem, this paper proposes a meta-top-down method that employs metaclassification to enhance the normal top-down classifying procedure. The proposed method is first analyzed theoretically on complexity and accuracy, and then applied to five real-world large-scale data sets. The experimental results indicate that the classification accuracy is largely improved, while the increased time costs are smaller than most of the existing approaches. Xiaolin Wang 0002, Hai Zhao 0001, Bao-Liang Lu |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2013 | An Empirical Study on Word Segmentation for Chinese Machine Translation
Hai Zhao 0001, Masao Utiyama, Eiichiro Sumita, Bao-Liang Lu |
CICLing (2) | 4 |
| 2013 | Converting Continuous-Space Language Models into N-Gram Language Models for Statistical Machine TranslationabstractNeural network language models, or continuous-space language models (CSLMs), have been shown to improve the performance of statistical machine translation (SMT) when they are used for reranking n-best translations.However, CSLMs have not been used in the first pass decoding of SMT, because using CSLMs in decoding takes a lot of time.In contrast, we propose a method for converting CSLMs into back-off n-gram language models (BNLMs) so that we can use converted CSLMs in decoding.We show that they outperform the original BNLMs and are comparable with the traditional use of CSLMs in reranking. Rui Wang 0015, Masao Utiyama, Isao Goto, Eiichiro Sumita, Hai Zhao 0001, Bao-Liang Lu |
EMNLP | 6 |
| 2013 | Real-Time Head Detection with Kinect for Driving Fatigue Detection
Bao-Liang Lu |
ICONIP (3) | 2 |
| 2013 | Detection of Driving Fatigue Based on Grip Force on Steering Wheel with Wavelet Transformation and Support Vector Machine
Bao-Liang Lu |
ICONIP (3) | 3 |
| 2013 | Marginalized Denoising Autoencoder via Graph Regularization for Domain Adaptation
Yong Peng 0001, Bao-Liang Lu |
ICONIP (2) | 3 |
| 2013 | Structure Preserving Low-Rank Representation for Semi-supervised Face Recognition
Yong Peng 0001, Suhang Wang, Bao-Liang Lu |
ICONIP (2) | 4 |
| 2013 | Labeled Alignment for Recognizing Textual Entailment
Xiaolin Wang 0002, Hai Zhao 0001, Bao-Liang Lu |
IJCNLP | 3 |
| 2013 | EEG-based vigilance estimation using extreme learning machines
Li-Chen Shi, Bao-Liang Lu |
Neurocomputing | 2 |
| 2012 | Online Vigilance Analysis Combining Video and Electrooculography Features
Ruofei Du, Ren-Jie Liu, Tian-Xiang Wu, Bao-Liang Lu |
ICONIP (5) | 4 |
| 2012 | EEG-Based Emotion Recognition in Listening Music by Using Support Vector Machine and Linear Dynamic System
Ruo-Nan Duan, Bao-Liang Lu |
ICONIP (4) | 3 |
| 2012 | EEG-Based Fatigue Classification by Using Parallel Hidden Markov Model and Pattern Classifier Combination
Bao-Liang Lu |
ICONIP (4) | 2 |
| 2012 | Parallel learning of large-scale multi-label classification problems with min-max modular LIBLINEARabstractThe study on pattern classification trends to be towards large-scale, multi-label, and imbalanced problems. The amount of the data which need to be classified is typically dozens of millions and it keeps rapid increasing in recent years. Traditional pattern classification approaches are inefficient and even ineffective in this situation. In our previous work, we proposed a min-max modular (M3) network for dealing with large-scale and imbalanced problems. M3-network is a generalized modular learning framework and includes three main steps: decomposing a large-scale problem into several smaller independent sub-problems, learning these sub-problems in parallel, and combining the results of the sub-problems to generate a solution to the original problem. In this paper, we embed LIBLINEAR into M3-network (M3-liblnear) to deal with large-scale, multi-label, and imbanlanced pattern classification problems. LIBLINEAR is a fast implementation of a linear classifier. M3-Liblinear uses LIBLINEAR as a base classifier to learn each of the sub-problems. We compare M3-Liblinear with Liblinear-cdblock on a large-scale Japanese patent classification problem. Experimental results demonstrate that M3-Liblinear is superior to Liblinear-cdblock in both training time and generalization performance. Bao-Liang Lu, Hai Zhao 0001 |
IJCNN | 2 |
| 2012 | Online vigilance analysis based on electrooculographyabstractThis study provides a highly efficient online method for vigilance analysis and verifies this theory in some experiments. Compared with electroencephalogram (EEG) signals, electrooculography (EOG) signals are easier to collect and faster to process. Some research has proven relations between vigilance and EOG features like blink features and slow eye movement (SEM). This study uses 48 kind of features of eye blinks, SEM and rapid eye movement (REM) from horizontal and vertical channels of EOG signals. It is verified by experiments that the precision of this method is higher than other methods which uses single kind of features like eye blinks. This study also implements an online vigilance analysis method and its precision is close to the offline method after about one minute from the beginning of collecting signals. With the application of dry electrode amplifiers, this algorithm is useful in real-time vigilance estimation in practical environment. This method can be an important part of brain-machine interfaces. Zheng-Ping Wei, Bao-Liang Lu |
IJCNN | 2 |
| 2012 | Spell Checking for Chinese
Hai Zhao 0001, Xiaolin Wang 0002, Bao-Liang Lu |
LREC | 4 |
| 2012 | Towards a Semantic Annotation of English Television News - Building and Evaluating a Constraint Grammar FrameNet
Hai Zhao 0001, Bao-Liang Lu |
PACLIC | 3 |
| 2012 | Gender classification by combining clothing, hair and facial component classifiers
Xiao-Chen Lian, Bao-Liang Lu |
Neurocomputing | 3 |
| 2012 | Multi-view gender classification using symmetry of facial images
Tian-Xiang Wu, Xiao-Chen Lian, Bao-Liang Lu |
Neural Comput. Appl. | 3 |
| 2011 | Time and space efficient spectral clustering via column samplingabstractSpectral clustering is an elegant and powerful approach for clustering. However, the underlying eigen-decomposition takes cubic time and quadratic space w.r.t. the data set size. These can be reduced by the Nyström method which samples only a subset of columns from the matrix. However, the manipulation and storage of these sampled columns can still be expensive when the data set is large. In this paper, we propose a time- and space-efficient spectral clustering algorithm which can scale to very large data sets. A general procedure to orthogonalize the approximated eigenvectors is also proposed. Extensive spectral clustering experiments on a number of data sets, ranging in size from a few thousands to several millions, demonstrate the accuracy and scalability of the proposed approach. We further apply it to the task of image segmentation. For images with more than 10 millions pixels, this algorithm can obtain the eigenvectors in 1 minute on a single machine. Mu Li 0001, Xiao-Chen Lian, James T. Kwok, Bao-Liang Lu |
CVPR | 4 |
| 2011 | Rank-SIFT: Learning to rank repeatable local interest pointsabstractScale-invariant feature transform (SIFT) has been well studied in recent years. Most related research efforts focused on designing and learning effective descriptors to characterize a local interest point. However, how to identify stable local interest points is still a very challenging problem. In this paper, we propose a set of differential features, and based on them we adopt a data-driven approach to learn a ranking function to sort local interest points according to their stabilities across images containing the same visual objects. Compared with the handcrafted rule-based method used by the standard SIFT algorithm, our algorithm substantially improves the stability of detected local interest point on a very challenging benchmark dataset, in which images were generated under very different imaging conditions. Experimental results on the Oxford and PASCAL databases further demonstrate the superior performance of the proposed algorithm on both object image retrieval and category recognition. Rong Xiao 0003, Zhiwei Li 0006, Rui Cai 0002, Bao-Liang Lu, Lei Zhang 0001 |
CVPR | 5 |
| 2011 | An Integrated Hierarchical Gaussian Mixture Model to Estimate Vigilance Level Based on EEG Recordings
Jing-Nan Gu, Hongtao Lu 0001, Bao-Liang Lu |
ICONIP (1) | 4 |
| 2011 | A Sparse Common Spatial Pattern Algorithm for Brain-Computer Interface
Li-Chen Shi, Rui-Hua Sun, Bao-Liang Lu |
ICONIP (1) | 4 |
| 2011 | EEG-Based Emotion Recognition Using Frequency Domain Features and Support Vector Machines
Dan Nie, Bao-Liang Lu |
ICONIP (1) | 3 |
| 2011 | Removing Unrelated Features Based on Linear Dynamical System for Motor-Imagery-Based Brain-Computer Interface
Li-Chen Shi, Bao-Liang Lu |
ICONIP (1) | 3 |
| 2011 | Enhance Top-down method with Meta-Classification for Very Large-scale Hierarchical Classification
Xiaolin Wang 0002, Hai Zhao 0001, Bao-Liang Lu |
IJCNLP | 3 |
| 2011 | A support vector machine classifier with automatic confidence and its application to gender classification
Ji Zheng, Bao-Liang Lu |
Neurocomputing | 2 |
| 2011 | Incorporating cellular sorting structure for better prediction of protein subcellular locationsabstractThis article explores the interdependences between subcellular locations and incorporates them with support vector machines for prediction of protein subcellular localisation. Traditional prediction systems utilise a ‘flat’ structure of classifiers, such as the one-versus-all and one-versus-one schemes, with amino acid compositions to perform the prediction. Apart from those existing studies that ignore the interdependences between subcellular locations, we take advantage of a hierarchical structure to organise the subcellular locations and model their relationships. Here, we propose to use four kinds of hierarchical prediction methods and make comparative studies on three datasets. Experimental results show that three of the hierarchical models outperform the traditional ‘flat’ model in terms of tree loss values. In particular, one hierarchical model outperforms the traditional ‘flat’ model for all evaluation measures. Moreover, we gained some valuable insights into the sorting process by using hierarchical structures. Wen-Yun Yang, Bao-Liang Lu, James T. Kwok |
J. Exp. Theor. Artif. Intell. | 2 |
| 2010 | Online multiple instance learning with no regretabstractMultiple instance (MI) learning is a recent learning paradigm that is more flexible than standard supervised learning algorithms in the handling of label ambiguity. It has been used in a wide range of applications including image classification, object detection and object tracking. Typically, MI algorithms are trained in a batch setting in which the whole training set has to be available before training starts. However, in applications such as tracking, the classifier needs to be trained continuously as new frames arrive. Motivated by the empirical success of a batch MI algorithm called MILES, we propose in this paper an online MI learning algorithm that has an efficient online update procedure and also performs joint feature selection and classification as MILES. Besides, while existing online MI algorithms lack theoretical properties, we prove that the proposed online algorithm has a (cumulative) regret of O(√T), where T is the number of iterations. In other words, the average regret goes to zero asymptotically and it thus achieves the same performance as the best solution in hindsight. Experiments on a number of MI classification and object tracking data sets demonstrate encouraging results. Mu Li 0001, James T. Kwok, Bao-Liang Lu |
CVPR | 3 |
| 2010 | Probabilistic models for supervised dictionary learningabstractDictionary generation is a core technique of the bag-of-visual-words (BOV) models when applied to image categorization. Most of previous approaches generate dictionaries by unsupervised clustering techniques, e.g. k-means. However, the features obtained by such kind of dictionaries may not be optimal for image classification. In this paper, we propose a probabilistic model for supervised dictionary learning (SDLM) which seamlessly combines an unsuper-vised model (a Gaussian Mixture Model) and a supervised model (a logistic regression model) in a probabilistic framework. In the model, image category information directly affects the generation of a dictionary. A dictionary obtained by this approach is a trade-off between minimization of distortions of clusters and maximization of discriminative power of image-wise representations, i.e. histogram representations of images. We further extend the model to incorporate spatial information during the dictionary learning process in a spatial pyramid matching like manner. We extensively evaluated the two models on various benchmark dataset and obtained promising results. Xiao-Chen Lian, Zhiwei Li 0006, Changhu Wang, Bao-Liang Lu, Lei Zhang 0001 |
CVPR | 4 |
| 2010 | Max-Margin Dictionary Learning for Multiclass Image Categorization
Xiao-Chen Lian, Zhiwei Li 0006, Bao-Liang Lu, Lei Zhang 0001 |
ECCV (4) | 3 |
| 2010 | Selecting Optimal Orientations of Gabor Wavelet Filters for Facial Image Analysis
Bao-Liang Lu |
ICISP | 2 |
| 2010 | Making Large-Scale Nyström Approximation Possible
Mu Li 0001, James T. Kwok, Bao-Liang Lu |
ICML | 3 |
| 2010 | Adaptive Ensemble Learning Strategy Using an Assistant Classifier for Large-Scale Imbalanced Patent Categorization
Qi Kong, Hai Zhao 0001, Bao-Liang Lu |
ICONIP (1) | 3 |
| 2010 | Age Classification Combining Contour and Texture Feature
Yan-Ming Tang, Bao-Liang Lu |
ICONIP (2) | 2 |
| 2010 | Multi-view Gender Classification Using Hierarchical Classifiers Structure
Tian-Xiang Wu, Bao-Liang Lu |
ICONIP (2) | 2 |
| 2010 | Pruning Training Samples Using a Supervised Clustering Algorithm
Minzhang Huang, Hai Zhao 0001, Bao-Liang Lu |
ISNN (2) | 3 |
| 2010 | Spectral and Semidefinite Relaxation of the CLUHSIC AlgorithmabstractCLUHSIC is a recent clustering framework that unifies the geometric, spectral and statistical views of clustering.In this paper, we show that the recently proposed discriminative view of clustering, which includes the DIFFRAC and DisKmeans algorithms, can also be unified under the CLUH-SIC framework.Moreover, CLUHSIC involves integer programming and one has to resort to heuristics such as iterative local optimization.In this paper, we propose two relaxations that are much more disciplined.The first one uses spectral techniques while the second one is based on semidefinite programming (SDP).Experimental results on a number of structured clustering tasks show that the proposed method significantly outperforms existing optimization methods for CLUHSIC.Moreover, it can also be used in semi-supervised classification.Experiments on real-world protein subcellular localization data sets clearly demonstrate the ability of CLUHSIC in incorporating structural and evolutionary information. Wen-Yun Yang, James T. Kwok, Bao-Liang Lu |
SDM | 3 |
| 2010 | Protein Subcellular Multi-Localization Prediction Using a Min-Max Modular Support Vector MachineabstractPrediction of protein subcellular localization is an important issue in computational biology because it provides important clues for the characterization of protein functions. Currently, much research has been dedicated to developing automatic prediction tools. Most, however, focus on mono-locational proteins, i.e., they assume that proteins exist in only one location. It should be noted that many proteins bear multi-locational characteristics and carry out crucial functions in biological processes. This work aims to develop a general pattern classifier for predicting multiple subcellular locations of proteins. We use an ensemble classifier, called the min-max modular support vector machine (M(3)-SVM), to solve protein subcellular multi-localization problems; and, propose a module decomposition method based on gene ontology (GO) semantic information for M(3)-SVM. The amino acid composition with secondary structure and solvent accessibility information is adopted to represent features of protein sequences. We apply our method to two multi-locational protein data sets. The M(3)-SVMs show higher accuracy and efficiency than traditional SVMs using the same feature vectors. And the GO decomposition also helps to improve prediction accuracy. Moreover, our method has a much higher rate of accuracy than existing subcellular localization predictors in predicting protein multi-localization. Yang Yang 0030, Bao-Liang Lu |
Int. J. Neural Syst. | 2 |
| 2010 | Modeling and Adaptive Control with Fuzzy Neural Networks - Selected Papers from the 6th International Symposium on Neural Networks
Wen Yu 0001, Bao-Liang Lu |
Neurocomputing | 2 |
| 2010 | A Unified Character-Based Tagging Framework for Chinese Word SegmentationabstractChinese word segmentation is an active area in Chinese language processing though it is suffering from the argument about what precisely is a word in Chinese. Based on corpus-based segmentation standard, we launched this study. In detail, we regard Chinese word segmentation as a character-based tagging problem. We show that there has been a potent trend of using a character-based tagging approach in this field. In particular, learning from segmented corpus with or without additional linguistic resources is treated in a unified way in which the only difference depends on how the feature template set is selected. It differs from existing work in that both feature template selection and tag set selection are considered in our approach, instead of the previous feature template focus only technique. We show that there is a significant performance difference as different tag sets are selected. This is especially applied to a six-tag set, which is good enough for most current segmented corpora. The linguistic meaning of a tag set is also discussed. Our results show that a simple learning system with six n -gram feature templates and a six-tag set can obtain competitive performance in the cases of learning only from a training corpus. In cases when additional linguistic resources are available, an ensemble learning technique, assistant segmenter, is proposed and its effectiveness is verified. Assistant segmenter is also proven to be an effective method as segmentation standard adaptation that outperforms existing ones. Based on the proposed approach, our system provides state-of-the-art performance in all 12 corpora of three international Chinese word segmentation bakeoffs. Hai Zhao 0001, Changning Huang, Mu Li 0001, Bao-Liang Lu |
ACM Trans. Asian Lang. Inf. Process. | 4 |
| 2009 | Gender Classification Based on Support Vector Machine with Automatic Confidence
Zheng Ji, Bao-Liang Lu |
ICONIP (1) | 2 |
| 2009 | Module Combination based on Decision Tree in Min-max Modular Network
Bao-Liang Lu, Zhi-Fei Ye |
IJCCI | 2 |
| 2009 | Incorporating Prior Knowledge into Task Decomposition for Large-Scale Patent Classification
Bao-Liang Lu, Masao Utiyama |
ISNN (2) | 2 |
| 2009 | LogisticLDA: Regularizing Latent Dirichlet Allocation by Logistic Regression
Jia-Cheng Guo, Bao-Liang Lu, Zhiwei Li 0006, Lei Zhang 0001 |
PACLIC | 2 |
| 2009 | Extracting Keyphrases from Chinese News Articles Using TextRank and Query Log Knowledge
Weiming Liang, Changning Huang, Mu Li 0001, Bao-Liang Lu |
PACLIC | 4 |
| 2009 | Computational prediction of novel non-coding RNAs in Arabidopsis thalianaabstractBACKGROUND: Non-coding RNA (ncRNA) genes do not encode proteins but produce functional RNA molecules that play crucial roles in many key biological processes. Recent genome-wide transcriptional profiling studies using tiling arrays in organisms such as human and Arabidopsis have revealed a great number of transcripts, a large portion of which have little or no capability to encode proteins. This unexpected finding suggests that the currently known repertoire of ncRNAs may only represent a small fraction of ncRNAs of the organisms. Thus, efficient and effective prediction of ncRNAs has become an important task in bioinformatics in recent years. Among the available computational methods, the comparative genomic approach seems to be the most powerful to detect ncRNAs. The recent completion of the sequencing of several major plant genomes has made the approach possible for plants. RESULTS: We have developed a pipeline to predict novel ncRNAs in the Arabidopsis (Arabidopsis thaliana) genome. It starts by comparing the expressed intergenic regions of Arabidopsis as provided in two whole-genome high-density oligo-probe arrays from the literature with the intergenic nucleotide sequences of all completely sequenced plant genomes including rice (Oryza sativa), poplar (Populus trichocarpa), grape (Vitis vinifera), and papaya (Carica papaya). By using multiple sequence alignment, a popular ncRNA prediction program (RNAz), wet-bench experimental validation, protein-coding potential analysis, and stringent screening against various ncRNA databases, the pipeline resulted in 16 families of novel ncRNAs (with a total of 21 ncRNAs). CONCLUSION: In this paper, we undertake a genome-wide search for novel ncRNAs in the genome of Arabidopsis by a comparative genomics approach. The identified novel ncRNAs are evolutionarily conserved between Arabidopsis and other recently sequenced plants, and may conduct interesting novel biological functions. Yang Yang 0030, Binglian Zheng, Zhidong Deng, Bao-Liang Lu, Tao Jiang 0001 |
BMC Bioinform. | 6 |
| 2009 | Incorporating prior knowledge into learning by dividing training data
Bao-Liang Lu, Xiaolin Wang 0002, Masao Utiyama |
Frontiers Comput. Sci. China | 1 |
| 2009 | An adaptive image Euclidean distance
Jing Li 0137, Bao-Liang Lu |
Pattern Recognit. | 2 |
| 2009 | Feature selection based on loss-margin of nearest neighbor classification
Yun Li 0009, Bao-Liang Lu |
Pattern Recognit. | 2 |
| 2008 | Semantic Similarity Definition over Gene Ontology by Further Mining of the Information Content
Yuan-Peng Li, Bao-Liang Lu |
APBC | 2 |
| 2008 | String Kernels with Feature Selection for SVM Protein Classification
Wen-Yun Yang, Bao-Liang Lu |
APBC | 2 |
| 2008 | Classification of Protein Sequences Based on Word Segmentation Methods
Yang Yang 0030, Bao-Liang Lu, Wen-Yun Yang |
APBC | 2 |
| 2008 | Gender Classification by Combining Facial and Hair Information
Xiao-Chen Lian, Bao-Liang Lu |
ICONIP (2) | 2 |
| 2008 | Cross Language Text Categorization Using a Bilingual Lexicon
Xiaolin Wang 0002, Bao-Liang Lu |
IJCNLP | 3 |
| 2008 | Large-scale patent classification with min-max modular support vector machinesabstractPatent classification is a large-scale, hierarchical, imbalanced, multi-label problem. The number of samples in a real-world patent classification typically exceeds one million, and this number increases every year. An effective patent classifier must be able to deal with this situation. This paper discusses the use of min-max modular support vector machine (M3-SVM) to deal with large-scale patent classification problems. The method includes three steps: decomposing a large-scale and imbalanced patent classification problem into a group of relatively smaller and more balanced two-class subproblems which are independent of each other, learning these subproblems using support vector machines (SVMs) in parallel, and combining all of the trained SVMs according to the minimization and the maximization rules. M3-SVM has two attractive features which are urgently needed to deal with large-scale patent classification problems. First, it can be realized in a massively parallel form. Second, it can be built up incrementally. Results from experiments using the NTCIR-5 patent data set, which contains more than two million patents, have confirmed these two attractive features, and demonstrate that M3-SVM outperforms conventional SVMs in terms of both training time and generalization performance. Xiao-Lei Chu, Jing Li 0137, Bao-Liang Lu, Masao Utiyama, Hitoshi Isahara |
IJCNN | 4 |
| 2008 | Multi-view gender classification based on local Gabor binary mapping pattern and support vector machinesabstractAbstract—This paper proposes a novel face representation approach, local Gabor binary mapping pattern (LGBMP), for multi-view gender classification. In this approach, a face image is first represented as a series of Gabor magnitude pictures (GMP) by applying multi-scale and multi-orientation Gabor filters. Each GMP is then encoded as a LGBP image where a uniform local binary pattern (LBP) operator is used. After that, each LGBP image is divided into non-overlapping rectangular regions, from which spatial histograms are extracted. Although an LGBP feature vector can be obtained by fitting together the regional histograms, it can not be employed in pattern classification due to its high dimension. We propose that each regional LGBP feature be mapped onto a one-dimensional subspace independently before they are concatenated as a whole feature vector. This is attractive since we reduce the feature dimension and also preserve the spatial information of LGBP image. Two ways have been proposed to map the regional LGBP feature in this paper. One is so-called LGBMP-LDA using linear discriminant analysis (LDA) for dimensionality reduction while the other is to project the regional LGBP feature onto the class center connecting line, namely, LGBMP-CCL. As a result, despite several decades of Gabor filters, the final feature dimension is even less than that of the feature extracted by using LBP directly on gray-scale images. The classification tasks in our work are performed by support vector machines (SVM). The experimental results on the CAS-PEAL face database indicate that the proposed approach achieves higher accuracy than the others such as SVMs+Gray-scale pixel, SVMs+Gabor and SVMs+LBP approach, more particularly, it has the lowest dimension of feature vector. I. He Sun 0003, Bao-Liang Lu |
IJCNN | 3 |
| 2008 | An empirical comparison of min-max-modular k -NN with different voting methods to large-scale text categorization
Bao-Liang Lu, Masao Utiyama, Hitoshi Isahara |
Soft Comput. | 2 |
| 2007 | Person-Specific SIFT Features for Face RecognitionabstractScale invariant feature transform (SIFT) proposed by Lowe has been widely and successfully applied to object detection and recognition. However, the representation ability of SIFT features in face recognition has rarely been investigated systematically. In this paper, we proposed to use the person-specific SIFT features and a simple non-statistical matching strategy combined with local and global similarity on key-points clusters to solve face recognition problems. Large scale experiments on FERET and CAS-PEAL face databases using only one training sample per person have been carried out to compare it with other non person-specific features such as Gabor wavelet feature and local binary pattern feature. The experimental results demonstrate the robustness of SIFT features to expression, accessory and pose variations. Erina Takikawa, Shihong Lao, Masato Kawade, Bao-Liang Lu |
ICASSP (2) | 6 |
| 2007 | A Framework for Multi-view Gender Classification
Jing Li 0137, Bao-Liang Lu |
ICONIP (1) | 2 |
| 2007 | Incorporating Domain Knowledge into a Min-Max Modular Support Vector Machine for Protein Subcellular Localization
Yang Yang 0030, Bao-Liang Lu |
ICONIP (2) | 2 |
| 2007 | Semi-Supervised Clustering for Vigilance Analysis Based on EEGabstractVigilance research is very useful and important to our daily lives. EEG has been proved very effective for measuring vigilance. Up to now, many researches mainly focus on using supervised learning methods to analyze the vigilance. However, the labelled information of vigilance is hard to get and sometimes not reliable. In this paper, we proposed a semi-supervised clustering method for vigilance analysis based on EEG. This method uses the insufficient labeled information to guide the vigilance related feature selection and uses prior knowledge of vigilance state transform to guide the clustering algorithm. The experiment results show that our method can almost correctly distinguish the awake state and the sleeping state by EEG, and can also represent the transform processes of reasonable middle states between the awake state and the sleeping state. Li-Chen Shi, Bao-Liang Lu |
IJCNN | 3 |
| 2007 | Learning Imbalanced Data Sets with a Min-Max Modular Support Vector MachineabstractTo overcome the class imbalance problem in statistical machine learning research area, re-balancing the learning task is one of the most classical and intuitive approach. Besides re-sampling, many researchers consider task decomposition as an alternative method for re-balance. Min-max modular support vector machine combines both intelligent task decomposition methods and the min-max modular network model as classifier ensemble. It overcomes several shortcomings of re-sampling, and could also achieve fast learning and parallel learning. We compare its classification performance with resampling and cost sensitive learning on several imbalanced data sets from different application areas. The experimental results indicate that our method can handle class imbalance problem efficiently. Zhi-Fei Ye, Bao-Liang Lu |
IJCNN | 2 |
| 2007 | Learning Concepts from Large-Scale Data Sets by Pairwise Coupling with Probabilistic OutputsabstractThis paper considers the problems of learning concepts from large-scale data sets. The way we take is completely classification algorithm independent. Firstly, the original problem is decomposed into a series of smaller two-class sub-problems which are easier to be solved. Secondly we present two principles, namely the shrink and expansion principles, to restore the global solution from the intermediate results learned from the sub-problems. In the theoretical analysis, this procedure of integration is described as a statistical inference of a posteriori probability and is degraded as the min-max principles in the special case considering 0-1 outputs. We also propose a revised approach which reduces the computational complexity of the training and testing stage to a linear level . Finally, experiments on both the synthetic and text-classification data are demonstrated. The results indicate that our methods are effective to large scale problems. Bao-Liang Lu |
IJCNN | 2 |
| 2007 | A Confident Majority Voting Strategy for Parallel and Modular Support Vector Machines
Yimin Wen, Bao-Liang Lu |
ISNN (3) | 2 |
| 2007 | A Probabilistic Approach to Feature Selection for Multi-class Text Categorization
Bao-Liang Lu, Masao Uchiyama, Hitoshi Isahara |
ISNN (1) | 2 |
| 2007 | Incremental Learning of Support Vector Machines by Classifier Combining
Yimin Wen, Bao-Liang Lu |
PAKDD | 2 |
| 2007 | Cross-Lingual Document Clustering
Bao-Liang Lu |
PAKDD | 2 |
| 2007 | Multi-View Gender Classification Using Multi-Resolution Local Binary Patterns and Support Vector MachinesabstractIn this paper, we present a novel method for multi-view gender classification considering both shape and texture information to represent facial images. The face area is divided into small regions from which local binary pattern (LBP) histograms are extracted and concatenated into a single vector efficiently representing a facial image. Following the idea of local binary pattern, we propose a new feature extraction approach called multi-resolution LBP, which can retain both fine and coarse local micro-patterns and spatial information of facial images. The classification tasks in this work are performed by support vector machines (SVMs). The experiments clearly show the superiority of the proposed method over both support gray faces and support Gabor faces on the CAS-PEAL face database. A higher correct classification rate of 96.56% and a higher cross validation average accuracy of 95.78% have been obtained. In addition, the simplicity of the proposed method leads to very fast feature extraction, and the regional histograms and fine-to-coarse description of facial images allow for multi-view gender classification. Hui-Cheng Lian, Bao-Liang Lu |
Int. J. Neural Syst. | 2 |
| 2006 | A Comparative Study on Feature Extraction from Protein Sequences for Subcellular Localization PredictionabstractOne of the central problems in computational biology is to identify the protein function in an automated and high-throughput fashion. A key step in this process is to predict subcellular compartment the protein belongs to, since the protein localization closely correlates with its function. A wide variety of methods for protein subcellular localization has been proposed over recent years. They fall into two categories, sequence-based and database-based. The first one is to extract useful features from amino acid sequences and strives to discover the principles behind protein localization process. The second one is more apt to conduct data mining from existing public annotation databases. This paper focuses on the sequence-based approach and exploits the discriminative ability contained in amino acid sequences for protein subcellular localization. By using support vector machines (SVMs) as predictors, we conducted comparisons among amino acid composition approach, amino acid tuple approach, voting scheme, and a new characteristic representation of proteins proposed in this paper. Our experiments are carried out on 7579 eukaryotic protein sequences from 12 subcellular locations. The highest accuracy, 82.8% across 5-fold cross validation, is obtained by voting scheme using five predictors. This is the best performance achieved on this dataset using sequence-based approach. Our experiments demonstrate that there are considerable potentials on improving prediction accuracy by exploiting protein sequences, which have not been fully utilized so far, and more explorations are still needed in this direction Wen-Yun Yang, Bao-Liang Lu, Yang Yang 0030 |
CIBCB | 2 |
| 2006 | Fast Learning for Statistical Face Detection
Zhi-Gang Fan, Bao-Liang Lu |
ICONIP (2) | 2 |
| 2006 | Efficient Classification of Multi-label and Imbalanced Data using Min-Max Modular ClassifiersabstractMany real-world applications, such as text categorization and subcellular localization of protein sequences, involve multi-label classification with imbalanced data. In this paper, we address these problems by using the minmax modular network. The min-max modular network can decompose a multi-label problem into a series of small two-class subproblems, which can then be combined by two simple principles. We also present several decomposition strategies to improve the performance of min-max modular networks. Experimental results on subcellular localization show that our method has better generalization performance than traditional SVMs in solving the multi-label and imbalanced data problems. Moreover, it is also much faster than traditional SVMs. Bao-Liang Lu, James T. Kwok |
IJCNN | 2 |
| 2006 | A New Supervised Clustering Algorithm Based on Min-Max Modular Network with Gaussian-Zero-Crossing FunctionsabstractIn this paper, we show that the shape, size and location of the receptive field around each instance are different and decided by the distribution of training data in a min-max modular network with Gaussian-zero-crossing functions. Based on this property, we propose a new supervised clustering algorithm which has the following features: First, the incremental clustering ability, which means the number of clusters need not to be predefined, it can grow up automatically, also, the training data need not to be processed iteratively; Second, attaching more importance to border instances than non-border instances, which guarantees the good generalization performance and training data reduction ratio; Third, outlier removal ability, which removes noise instances from training data; Last, cluster combination ability, which reduces the number of clusters further. Experiments on an artificial problem and several real-world applications demonstrate these attractive features of our new clustering algorithm. Jing Li 0137, Bao-Liang Lu |
IJCNN | 2 |
| 2006 | Multi-view Gender Classification Using Local Binary Patterns and Support Vector Machines
Hui-Cheng Lian, Bao-Liang Lu |
ISNN (2) | 2 |
| 2006 | Gender Recognition Using a Min-Max Modular Support Vector Machine with Equal Clustering
Bao-Liang Lu |
ISNN (2) | 2 |
| 2006 | Prediction of Protein Subcellular Multi-locations with a Min-Max Modular Support Vector Machine
Yang Yang 0030, Bao-Liang Lu |
ISNN (2) | 2 |
| 2006 | A Modular Reduction Method for k-NN Algorithm with Self-recombination Learning
Hai Zhao 0001, Bao-Liang Lu |
ISNN (1) | 2 |
| 2006 | Effective Tag Set Selection in Chinese Word Segmentation via Conditional Random Field Modeling
Hai Zhao 0001, Changning Huang, Mu Li 0001, Bao-Liang Lu |
PACLIC | 4 |
| 2005 | Extracting Features from Protein Sequences Using Chinese Segmentation Techniques for Subcellular Localization
Yang Yang 0030, Bao-Liang Lu |
CIBCB | 2 |
| 2005 | Fast Recognition of Multi-View Faces with Feature SelectionabstractWe propose a discriminative feature selection method utilizing support vector machines for the challenging task of multiview face recognition. According to the statistical relationship between the two tasks, feature selection and multiclass classification, we integrate the two tasks into a single consistent framework and effectively realize the goal of discriminative feature selection. The classification process can be made faster without degrading the generalization performance through this discriminative feature selection method. On the UMIST multiview face database, our experiments show that this discriminative feature selection method can speed up the multiview face recognition process without degrading the correct rate and outperform the traditional kernel subspace methods. Zhi-Gang Fan, Bao-Liang Lu |
ICCV | 2 |
| 2005 | An algorithm for pruning redundant modules in min-ma modular network [min-ma read min-max]abstractThe min-max modular (M/sup 3/) network is a framework that is capable of solving large-scale pattern classification problems in a parallel way. The M/sup 3/ network has been successfully applied to several large-scale real-world problems. When a complex problem is decomposed into a number of separable problems, however, the M/sup 3/ network suffers from its high redundancy of individual modules. This paper proposes an algorithm, called back-searching (BS) algorithm, to prune these redundant modules. The main idea behind the BS algorithm is to use the actual outputs of the trained M/sup 3/ network associated with training data to find out the redundant modules by means of 'back searching'. In order to ensure the correctness of the algorithm, we prove two propositions theoretically, namely the sufficient proposition and the necessary proposition, and perform simulations on several benchmark and real-world problems. The simulation results indicate that most of all redundant modules can be pruned by our proposed algorithm and the pruned network has the same generalization performance as the original network. Hui-Cheng Lian, Bao-Liang Lu |
IJCNN | 2 |
| 2005 | Fast text categorization with min-max modular support vector machinesabstractThe min-max modular support vector machines (M/sup 3/-SVMs) have been proposed for solving large-scale and complex multiclass classification problems. In this paper, we apply the M/sup 3/-SVMs to multilabel text categorization and introduce a new task decomposition strategy into M/sup 3/-SVMs. A multilabel classification task can be split up into a set of two-class classification tasks. These two-class tasks are to discriminate the C class from non-C class. If these two class tasks are still hard to be learned, we can further divide them into a set of two-class tasks as small as needed and fast training of SVMs on massive multilabel texts can be easily implemented in a massively parallel way. Furthermore, we proposed a new task decomposition strategy called hyperplane task decomposition to improve generalization performance. The experimental results on the RC 1-v2 indicate that the new method has better generalization performance than traditional SVMs and previous M/sup 3/-SVMs using random task decomposition, and is faster than traditional SVMs. Feng-Yao Liu, Hai Zhao 0001, Bao-Liang Lu |
IJCNN | 4 |
| 2005 | On efficient selection of binary classifiers for min-max modular classifierabstractBinary classifiers are fundamental components of multiclass pattern classifiers. How to construct a solution to a multiclass problem by efficiently combining the outputs of binary classifiers is a very important issue in neural network and machine learning research. In this paper, we present three different algorithms for selecting binary classifiers for min-max modular classifier to improve its response performance. We also give a theoretical performance estimation of the proposed algorithms. We prove that quadratic complexity of original min-max combination can be reduced to the level of linear complexity in the number of binary classifiers. The experimental results indicate that our proposed algorithms are efficient and effective. Hai Zhao 0001, Bao-Liang Lu |
IJCNN | 2 |
| 2005 | Typical Sample Selection and Redundancy Reduction for Min-Max Modular Network with GZC Function
Jing Li 0137, Bao-Liang Lu, Michinori Ichikawa |
ISNN (1) | 2 |
| 2005 | Task Decomposition Using Geometric Relation for Min-Max Modular SVMs
Kai-An Wang, Hai Zhao 0001, Bao-Liang Lu |
ISNN (1) | 3 |
| 2005 | A Hierarchical and Parallel Method for Training Support Vector Machines
Yimin Wen, Bao-Liang Lu |
ISNN (1) | 2 |
| 2005 | Structure Pruning Strategies for Min-Max Modular Network
Yang Yang 0030, Bao-Liang Lu |
ISNN (1) | 2 |
| 2005 | Improvement on Response Performance of Min-Max Modular Classifier by Symmetric Module Selection
Hai Zhao 0001, Bao-Liang Lu |
ISNN (2) | 2 |
| 2004 | Feature Selection for Fast Image Classification with Support Vector Machines
Zhi-Gang Fan, Kai-An Wang, Bao-Liang Lu |
ICONIP | 3 |
| 2004 | Fault Diagnosis for Industrial Images Using a Min-Max Modular Neural Network
Bao-Liang Lu |
ICONIP | 2 |
| 2004 | A part-versus-part method for massively parallel training of support vector machinesabstractThis work presents a part-versus-part decomposition method for massively parallel training of multi-class support vector machines (SVMs). By using this method, a massive multi-class classification problem is decomposed into a number of two-class subproblems as small as needed. An important advantage of the part-versus-part method over existing popular pair wise-classification approach is that a large-scale two-class subproblem can be further divided into a number of relatively smaller and balanced two-class subproblems, and fast training of SVMs on massive multi-class classification problems can be easily implemented in a massively parallel way. To demonstrate the effectiveness of the proposed method, we perform simulations on a large-scale text categorization problem. The experimental results show that the proposed method is faster than the existing pairwise-classification approach, better generalization performance can be achieved, and the method scales up to massive, complex multi-class classification problems. Bao-Liang Lu, Kai-An Wang, Masao Utiyama, Hitoshi Isahara |
IJCNN | 1 |
| 2004 | An Adjusted Gaussian Skin-Color Model Based on Principal Component Analysis
Zhi-Gang Fan, Bao-Liang Lu |
ISNN (1) | 2 |
| 2004 | A Cascade Method for Reducing Training Time and the Number of Support Vectors
Yimin Wen, Bao-Liang Lu |
ISNN (1) | 2 |
| 2004 | Analysis of Fault Tolerance of a Combining Classifier
Hai Zhao 0001, Bao-Liang Lu |
ISNN (1) | 2 |
| 2003 | Efficient Part-of-Speech Tagging with a Min-Max Modular Neural-Network Model
Bao-Liang Lu, Michinori Ichikawa, Hitoshi Isahara |
Appl. Intell. | 1 |
| 2003 | Converting general nonlinear programming problems into separable programming problems with feedforward neural networks
Bao-Liang Lu, Koji Ito |
Neural Networks | 1 |
| 2001 | Massively Parallel Classification of EEG Signals Using Min-Max Modular Neural Networks
Bao-Liang Lu, Jonghan Shin, Michinori Ichikawa |
ICANN | 1 |
| 2001 | On-Line Error Detection of Annotated Corpus Using Modular Neural Networks
Bao-Liang Lu, Masaki Murata, Michinori Ichikawa, Hitoshi Isahara |
ICANN | 2 |
| 2001 | Reading auditory discrimination behaviour of freely moving rats from hippocampal EEG
Jonghan Shin, Bao-Liang Lu, Arkadi Talnov, Gen Matsumoto, Jurij Brankack |
Neurocomputing | 2 |
| 2000 | Emergence of Learning: An Approach to Coping with NP-Complete Problems in LearningabstractVarious theoretical results show that learning in conventional feedforward neural networks such as multilayer perceptrons is NP-complete. In this paper we show that learning in min-max modular (M/sup 3/) neural networks is tractable. The key to coping with NP-complete problems in M/sup 3/ networks is to decompose a large-scale problem into a number of manageable, independent subproblems and to make the learning of a large-scale problem emerge from the learning of a number of related small subproblems. Bao-Liang Lu, Michinori Ichikawa |
IJCNN (4) | 1 |
| 1999 | Task decomposition and module combination based on class relations: a modular neural network for pattern classificationabstractIn this paper, we propose a new method for decomposing pattern classification problems based on the class relations among training data. By using this method, we can divide a K-class classification problem into a series of ((2)(K)) two-class problems. These two-class problems are to discriminate class Ci from class Cj for i=1, ..., K and j = i+1, while the existence of the training data belonging to the other K-2 classes is ignored. If the two-class problem of discriminating class Ci from class Cj is still hard to be learned, we can further break down it into a set of two-class subproblems as small as we expect. Since each of the two-class problems can be treated as a completely separate classification problem with the proposed learning framework, all of the two-class problems can be learned in parallel. We also propose two module combination principles which give practical guidelines in integrating individual trained network modules. After learning of each of the two-class problems with a network module, we can easily integrate all of the trained modules into a min-max modular (M3) network according to the module combination principles and obtain a solution to the original problem. Consequently, a large-scale and complex K-class classification problem can be solved effortlessly and efficiently by learning a series of smaller and simpler two-class problems in parallel. Bao-Liang Lu, Masami Ito |
IEEE Trans. Neural Networks | 1 |
| 1999 | Inverting feedforward neural networks using linear and nonlinear programmingabstractThe problem of inverting trained feedforward neural networks is to find the inputs which yield a given output. In general, this problem is an ill-posed problem because the mapping from the output space to the input space is a one-to-many mapping. In this paper, we present a method for dealing with the inverse problem by using mathematical programming techniques. The principal idea behind the method is to formulate the inverse problem as a nonlinear programming (NLP) problem, a separable programming (SP) problem, or a linear programming (LP) problem according to the architectures of networks to be inverted or the types of network inversions to be computed. An important advantage of the method over the existing iterative inversion algorithm is that various designated network inversions of multilayer perceptrons (MLP's) and radial basis function (RBF) neural networks can be obtained by solving the corresponding SP problems, which can be solved by a modified simplex method, a well-developed and efficient method for solving LP problems. We present several examples to demonstrate the proposed method and the applications of network inversions to examining and improving the generalization performance of trained networks. The results show the effectiveness of the proposed method. Bao-Liang Lu, Hajime Kita, Yoshikazu Nishikawa |
IEEE Trans. Neural Networks | 1 |
| 1998 | Decomposition and Parallel Learning of Imbalanced Classification Problems by Min-Max Modular Neural Network
Bao-Liang Lu, Masami Ito |
ICONIP | 1 |