VLDB 2026 Research / reviewers in the wild / expert
Wanzeng Kong
dblp:04/10544
· DBLP profile ↗
91ranked-venue papers
4as first author
68since 2021 · last 2026
0000-0002-0113-6968ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 1 first-author · 29 since 2021Artificial intelligence and machine learning · 33 · 3 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 15 since 2021Human-computer interaction and ubiquitous computing · 6 · 4 since 2021Computer networks · 5 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dense structure distillation towards recurrent scene flow estimation for sparse point cloudabstractAbstract Scene flow, a key representation of motion information in 3D space, plays a critical role in numerous downstream tasks. However, existing point cloud scene flow estimation methods often experience significant performance degradation due to insufficient feature expressiveness or mismatching under sparse point cloud conditions. To alleviate these problems, we propose a novel dense structure distillation towards recurrent scene flow estimation method for sparse point clouds, significantly enhancing performance under low-density point cloud conditions through a dense structure distillation strategy. This module addresses the information loss caused by point cloud sparsity by leveraging high-quality point cloud features to effectively guide the learning of sparse point cloud feature extractors. To further refine the estimation results, a recurrent update strategy is adopted to gradually improve the accuracy and stability of scene flow estimation. Experimental results demonstrate that the proposed method achieves state-of-the-art performance on both the public FlyingThings3D and KITTI datasets, particularly under sparse point cloud conditions, and outperforms existing methods. Weichen Dai 0001, Ziyue Meng, Xiaoyang Weng, Wenhan Su, Wanzeng Kong |
Comput. J. | 5 |
| 2026 | Brain-Eye Collaborative Camouflaged Target DetectionabstractBrain-computer interface (BCI) based on rapid serial visual presentation (RSVP) are widely applied in target detection but suffer from low decoding accuracy. While some methods integrate eye movements to improve localization, they often underutilize EEG’s coarse spatial cues, limiting overall detection effectiveness. We propose a brain-eye collaborative method for camouflaged target detection that integrates EEG and eye movement signals in a coarse-to-fine framework. In the coarse stage, target presence and its image quadrant are determined using brain-eye fusion and contrastive learning on bimodal features. This leverages complementary spatiotemporal information from both modalities to generate a coarse localization region. In the fine stage, eye movement data further refines the target location within this region by identifying high-interest areas. This collaborative approach significantly improves both recognition and localization performance in complex visual search tasks. Our method achieves an F1 score of 83.16% and a balanced accuracy of 85.38%, outperforming existing state-of-the-art methods. Longjie Ma, Weichen Dai 0001, Ziyue Yang 0007, Chenyi Hong, Jianting Cao, Wanzeng Kong |
Int. J. Hum. Comput. Interact. | 7 |
| 2026 | GaborNet: attention based Gabor convolutional networks for contactless palmprint recognition
Xueqin Xiang, Wanzeng Kong, Yong Peng 0001 |
Multim. Tools Appl. | 3 |
| 2026 | Unsupervised multimodal remote sensing image registration via two-stream causal sequential inference
Han Yang 0003, Baoheng Wang, Risheng Huang, Yong Peng 0001, Wanzeng Kong |
Pattern Recognit. | 7 |
| 2026 | Hybrid feature selection for cross-domain few shot learning
Xueqin Xiang, Wanzeng Kong, Jinliang Yao, Haihong Wu |
Pattern Recognit. | 3 |
| 2026 | Multi-View Manifold-Adaptive Kernel Regression for Speech Classification From EEG SignalsabstractDecoding speech intentions from electroencephalogram (EEG) data is the primary task in speech brain-computer interface (BCI) systems, which remains challenging due to the unclear discriminative task-aware features, and underlying nonlinear properties besides the well-known low signal-to-noise ratio of EEG data. Existing approaches typically rely either on single-domain features or performing feature learning by deep neural networks; therefore, they either fail to capture comprehensive signal patterns, or typically require large-sized EEG data to fit the parameter spaces and often have limited interpretability. To address these limitations, we propose a Multi-view Manifold-Adaptive Kernel Regression (MMKR) model for speech recognition from EEG signals in this paper. By treating temporal, spectral, and statistical EEG representations as complementary feature views, view-specific manifold-adaptive kernels are constructed in MMKR to incorporate local graph structure into kernel similarity; besides, a data-driven adaptive view weighting mechanism is used to characterize their contributions. We evaluate MMKR on both overt and imagined speech EEG datasets and the results demonstrate that MMKR achieves superior classification accuracy and robustness compared to some representative single-view, multi-view, and kernel-based baselines. Moreover, analysis on the local manifold-modulated kernel matrix and the learned view contributions are provided. Yong Peng 0001, Wanzeng Kong |
IEEE Signal Process. Lett. | 5 |
| 2026 | High-Fidelity Sonar Waveform Synthesis With Multi-Band Adversarial NetworksabstractRecent advances in underwater sonar detection have highlighted the challenges of acquiring real-world sonar data due to experimental costs and environmental constraints. While deep learning-based methods have shown promise in sonar synthesis, existing approaches often lack fidelity in waveform generation. This manuscript proposes Multi-Band HiFi-GAN, a high-fidelity synthesis method that incorporates a multi-band processing module based on a Pseudo-QMF bank for spectrally-efficient sub-band decomposition, and Multi-Band STFT Discriminator (MBSD) that jointly models time-frequency structures and mitigates aliasing artifacts. Additionally, we design a multi-resolution STFT loss to enhance convergence. Evaluated on DeepShip dataset using Fréchet Audio Distance (FAD), our method demonstrates superior performance, achieving a lower FAD score than existing approaches. Augmenting real training data with synthesized samples improves target classification accuracy to 96.94% at a 100% mixing ratio. The work provides a practical data augmentation solution for underwater acoustic systems, improving synthesis fidelity. Han Yang 0003, Yiwen Shen 0007, Honggang Liu, Wanzeng Kong |
IEEE Signal Process. Lett. | 5 |
| 2026 | Step-Wise Prompting Meets Uncertainty-Aware Dynamic Fusion for Robust EEG-Visual Emotion RecognitionabstractUnderstanding emotional states is fundamental to advancing next-generation AI systems with human-like attributes. Combining the complementary strengths of electroencephalog raphy (EEG) and facial expressions holds great promise for advancing multimodal emotion recognition (MER). EEG provides objective measurements of neural activity, while facial expressions convey rich, externally observable emotional cues. However, existing joint learning frameworks often fall short of fully exploiting the synergy between these modalities. Two key challenges remain unresolved: (1) insufficient cross-modal alignment and interaction prior to fusion which limits the semantic complementarity between modalities; and (2) modality unreliability caused by temporal fluctuations in signal quality and inconsistencies in emotional semantics. To address these limitations, we propose a novel framework that integrates a Step-wise Prompts (SwiP) module with an Uncertainty-Aware Dynamic Fusion (UADF) mechanism. SwiP enables progressive, fine-grained interaction by introducing sequential facial features as visual prompts to guide EEG representation learning, thereby enhancing cross modal complementarity. UADF dynamically adjusts modality contributions through a token- and modality-level uncertainty estimation scheme, enabling the model to selectively emphasize informative inputs and suppress noisy or irrelevant signals. Extensive experiments on benchmark datasets demonstrate that our method achieves state-of-the-art performance, consistently outperforming competitive baselines in both accuracy and stability. These results highlight the potential of our framework as a robust and generalizable solution for real-world affective computing applications. Danyang Hao, Yunyuan Gao, Xiaohui Lou, Wanzeng Kong |
IEEE Trans. Affect. Comput. | 5 |
| 2026 | Brain-Machine Enhanced Intelligence for Semi-Supervised Facial Emotion RecognitionabstractMachine learning, particularly deep learning, typically achieves high facial emotion image recognition accuracy benefiting from multiple labeled data. However, the datasets usually contain insufficient labeled samples and numerous unlabeled data since human labeling is a costly endeavor. For semi-supervised learning of these datasets, self-training procedure solely based on the visual features of images fails to comprehensively understand the intricate high-level semantic features. Since EEG signals contain not only visual information related to the visual stimulus but also emotional information related to brain activity, they are highly suitable as supervisory signals for labeling unlabeled facial emotion images. In this study, we specifically employ EEG signals evoked by visual image stimuli in conjunction with EEGNet3D to learn a discriminative EEG class representation manifold of brain activity. The one-hot class label is replaced with the EEG class representation as the supervisory to train the base model. Then, better pseudo-labeling is achieved using the base model in the EEG class representation manifold. Based on pseudo-labeling results, the utilization of unlabeled data is further improved. Interestingly, our findings reveal that when utilizing EEG class representations as supervisory information for the base model, the base model demonstrates a learning pattern that involves focusing more on the eye area when making judgments about emotions. This behavior closely resembles how the human brain decodes emotions. Experiments show that the performance of the proposed method can be effectively enhanced by combining labeled and pseudo-labeled images. Further experiments demonstrate that our method exhibits strong generalization abilities when applied to new image datasets and other visual networks. Dongjun Liu, Weichen Dai 0001, Hangjie Yi, Honggang Liu, Jianting Cao, Qibin Zhao, Fabio Babiloni, Wanzeng Kong |
IEEE Trans. Affect. Comput. | 8 |
| 2026 | Cognition-Guided Coupled Learning for Camouflaged Object DetectionabstractCamouflaged object detection (COD) aims to identify targets that visually blend into their surroundings, posing significant challenges due to low contrast between targets and backgrounds, diverse target appearances, and limited annotated training samples. A critical yet often overlooked subtask in COD is object presence detection, which determines whether a camouflaged object actually exists in an image. This step is essential as it underpins accurate localization and segmentation and enables timely decisions in practical applications such as disaster rescue and wildlife monitoring. However, most existing methods assume the target is always present and focus mainly on localization, limiting their robustness and reliability in complex scenarios. Inspired by the human brain's remarkable ability to perceive camouflaged objects by integrating edge and contextual cues, we propose a cognitive-visual coupling learning framework to enhance object presence detection. Specifically, electroencephalogram (EEG) signals are incorporated during training to embed human cognitive patterns into the model's visual feature learning through coupled representation learning. Notably, during inference, our framework relies solely on visual input, effectively leveraging cognitive insights acquired during training to improve detection performance. Extensive experiments demonstrate that our approach achieves superior accuracy compared to state-of-the-art methods and exhibits strong generalization across diverse visual backbone architectures and unseen cross-dataset scenarios. Hanru Zhou, Dongjun Liu, Tianyang Qin, Wanzeng Kong |
IEEE Trans. Hum. Mach. Syst. | 6 |
| 2026 | DRFNet: Enhancing Identity Discriminability and Feature Robustness for Cross-Session VEP-Based EEG BiometricsabstractBiometric recognition using visually evoked potentials (VEPs), a type of neural response to visual stimuli recorded via electroencephalography (EEG), has shown great promise. However, the non-stationary nature of EEG signals poses a major challenge in cross-session scenarios, where data collected on different days often leads to performance degradation. To address this, we propose the Discriminative Robust Feature Network (DRFNet) to enhance the robustness and inter-subject discriminability of identity representations across sessions. DRFNet incorporates two key components: (1) A log-power transformation that amplifies inter-individual differences by capturing non-linear energy patterns from VEP features via signal squaring and logarithmic scaling; and (2) A hierarchical normalization strategy with adaptive attention to balance discriminative identity cues with inter-session invariance by stabilizing feature distributions across multiple levels (feature map, batch, and sample). On two public multi-session SSVEP datasets (Dataset A: 30 subjects, 6 s trials; Dataset B: 54 subjects, 4 s trials), our model outperformed state-of-the-art methods, achieving identification accuracies of 92.92% and 86.30%, and equal error rates of 3.92% and 4.09%, respectively. Further analysis demonstrates that filter bank processing and a reduced set of parietal-occipital electrodes can provide more discriminative features while offering a practical path toward system lightweighting. Honggang Liu, Han Yang 0003, Dongjun Liu, Xuanyu Jin, Yong Peng 0001, Wanzeng Kong |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | ASA-STGCN: Adaptive Sparse Awareness-Spatiotemporal Graph Convolutional Network for Multi-Class Motor Imagery EEG ClassificationabstractGraph Convolutional Networks (GCNs) have shown promise in motor imagery electroencephalogram (EEG) signals classification by modeling spatial dynamics and brain connectivity. However, over-smoothing remains a challenge, leading to homogenized node features and reduced discrimination. To address this, we propose an Adaptive Sparse Awareness-Spatiotemporal Graph Convolutional Network (ASA-STGCN) that combines adaptive sparse graph convolution with attention mechanisms. Notably, a Graph Sparse Convolutional Network (GSCN) in the Adaptive Sparse Awareness Spatial Module (ASAM) enhances brain region feature selection, while the Graph Node Neighborhood Awareness Layer (GNNAL) applies self-attention to reinforce critical topological relationships. The Multi-scale Temporal Convolution Module (MTCM) captures both transient and sustained temporal dependencies. Experimental results achieve accuracies of 97.2%±3.4% (binary) and 83.6%±4.9% (four-class) on BCIC-IV-2a, 96.6%±3.1% (binary) on BCIC-III IVa, and 83.41%±4.3 (binary) on OpenBMI. Discussion confirms the model's effectiveness and its potential to support EEG-based neurorehabilitation and clinical brain computer interface applications. Peiqi Yu, Qingshan She, Xugang Xi, Wanzeng Kong |
IEEE J. Biomed. Health Informatics | 5 |
| 2026 | Prediction Consistency and Confidence-Based Proxy Domain Construction for Privacy-Preserving in Cross-Subject EEG ClassificationabstractDomainadaptation has proven effective for suppressing the inter-subject variability problem in cross-subject EEG classification tasks in which labeled data is available for source subjects while only unlabeled data is provided for target subjects. Existing domain adaptation methods typically reduced the distribution discrepancy between source and target domains by directly utilizing source domain samples or features. To safeguard the privacy of source domain data, we propose to construct a Proxy Domain by simultaneously considering the prediction Consistency and Confidence (PDCC) of locally trained source models on target EEG samples, serving as the substitute to the source domain. The framework commences with the augmentation and alignment of the source domain data to enhance feature generalizability, after which source models are trained independently on each source subject's data in a decentralized manner. Knowledge transfer from source to target domains is achieved exclusively through accessing to the source domain model, enabling the PDCC-based proxy domain construction that encapsulates the source knowledge. Finally, domain adaptation is performed using the proxy domain and target domain. As a result, PDCC eliminates the need to access source domain data while effectively leveraging source knowledge. Experimental results on four benchmark EEG datasets demonstrate that PDCC consistently outperforms eleven existing methods, including several advanced transfer learning and source-free methods. Especially, the effectiveness of the proxy domain is extensively investigated. Yong Peng 0001, Jiangchuan Liu, Honggang Liu, Natasha M. J. Padfield, Wanzeng Kong, Bao-Liang Lu, Andrzej Cichocki |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | Adaptive Multimodal Semantic Balancing Framework for Sentiment Analysis
Jiajia Tang, Feiwei Zhou, Xiping Wang, Qibin Zhao, Yu Ding 0001, Wanzeng Kong |
IEEE Trans. Multim. | 7 |
| 2026 | Cognition-driven Adaptive Semantic Decoding Framework for Multimodal Sentiment AnalysisabstractIn real-world scenarios, multimodal sentiment analysis faces significant challenges, particularly in cross-scenario generalization. Existing works fail to effectively deal with the variability in evaluation frameworks and modality combinations, which results in poor transfer performance across different application contexts. In this article, the cognition-driven adaptive semantic decoding framework (CASDF) is proposed to realize an evaluation system and modality-independent multimodal sentiment analysis. Specifically, the adaptive modality association module is proposed to construct the adaptive modality mapping space, which allows us to dynamically adapt to arbitrary modality combinations. This indeed breaks through the limitation of the modality number and effectively deals with the modality gap. Furthermore, similar to the human hierarchical cognition (“perception-concept-decision”), the evaluation system progressive alignment module is presented to establish the unified evaluation system. This consists of the perception, concept, and decision analysis, which contributes to the adaptive cross-task analysis from the discrete sentiment space to the continuous sentiment space. The above joint analysis of the evaluation system and modality number indeed leads to the more flexible and generable multimodal sentiment semantic decoding paradigm. The experiments demonstrate that our sentiment semantic analysis network can achieve state-of-the-art performance. Jiajia Tang, Honggang Liu, Xuanyu Jin, Wanzeng Kong |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2026 | LAE-Net: Large Pretrained Models Assistant Text-Guided Image Editing Adversarial NetworkabstractAutomatic real image editing offers unprecedented freedom to modify the appearance of the image or to edit a few objects through natural language. Recent scalable model families such as diffusion models have showcased remarkable proficiency in editing highly realistic images due to the introduction of vast amounts of training data and large pretrained language models. However, these large diffusion models require iterative evaluation that would significantly hinder the pace of image editing. Moreover, the pioneering work in this field necessitates the learning of a unique textual token that corresponds to each input image, or a group of images containing the same object, leading to the generation of redundant and fragmented models. Given the aforementioned problems, we suggest a novel Large pretrained models Assistant text-guided image Editing adversarial Network (LAE-Net) in this paper. More concretely, we introduce a deep semantic editing network to globally transfer text information among different isolated editing blocks, which would extract features from the source image to differentiate text-required areas from text-irrelevant ones. Furthermore, based on idea that the multi-modal CLIP model, leveraging vision-language alignment, captures comprehensive global semantic cues, whereas the vision-centric DINO model specializes in delivering intricate, fine-grained pixel-level details, the powerful discriminator of LAE-Net is designed by harnessing the visual embeddings derived from both the CLIP and DINO models separately to boost the visual discriminant capability and facilitate training a strong generator for conditioning image generation. Comprehensive experimental evaluations show that our LAE-Net not only delivers outstanding performance but also surpasses several cutting-edge models. Xueqin Xiang, Yong Peng 0001, Wanzeng Kong, Jinliang Yao |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | Dual-Branch EEG Decoding Method for Collaborative Multi-Brain Motor Imagery
Jiaxuan Qin, Li Zhu 0005, Jiangxu Wu, Jinda Liao, Wanzeng Kong |
CogSci | 5 |
| 2025 | RPW-EEG: An Unified Framework for Robust and Practical Watermark of EEG
Tianyang Qin, Hangjie Yi, Jingsheng Qian, Xuanyu Jin, Honggang Liu, Wanzeng Kong |
CogSci | 6 |
| 2025 | Multi-Modal Synergistic Implicit Image Enhancement for Efficient Optical Flow EstimationabstractAs a fundamental visual task, optical flow estimation has widespread applications in computer vision. However, it faces significant challenges under adverse lighting conditions, where low texture and noise make accurate optical flow estimation particularly difficult. In this paper, we propose an optical flow method that employs implicit image enhancement through multi-modal synergistic training. To supplement the scene information missing in the original low-quality image, we utilize a high-low frequency feature enhancement network. The enhancement network is implicitly guided by multi-modal data and the specific subsequent tasks, enabling the model to learn multi-modal knowledge that enhances feature information suitable for optical flow estimation during inference. By using RGBD multi-modal data, the proposed method avoids the reliance on the images captured from the same view, a common limitation in traditional image enhancement methods. During training, the encoded features extracted from the enhanced images are synergistically supervised by features from the RGBD fusion as well as by the optical flow task. Experiments conducted on both synthetic and real datasets demonstrate that the proposed method significantly improves performance on public datasets. Weichen Dai 0001, Hexing Wu, Xiaoyang Weng, Yuhang Ming 0001, Wanzeng Kong |
CVPR | 6 |
| 2025 | Brain Generative Replay for Continual Learning
Jianguo Zhou, Dongjun Liu, Wanzeng Kong |
ICANN (1) | 3 |
| 2025 | Deep Transfer Regression for EEG-based Driving Fatigue DetectionabstractRecently, Electroencephalography (EEG) has been increasingly utilized in driving fatigue detection tasks. However, the inter-subject variabilities in EEG data render models trained on one subject ineffective for being directly applied to others. Transfer learning has been widely used to address this issue, but most existing transfer learning algorithms primarily focused on classification tasks. Therefore, we propose a transfer regression model for EEG-based driving fatigue detection, whose core idea is to learn the weights of models from various source domain data and a base model from target domain training data through an attention network. By assembling models trained on different domain data, predictions are obtained. We conducted experiments on the two subsets of the benchmark SEED-VIG dataset, and the results demonstrate that our transfer regression model effectively enhances the driving fatigue detection performance. The source code is available from https://github.com/SunseaIU/ATR-EEG. Yikai Zhang 0005, Yong Peng 0001, Ziyue Yang 0007, Fei-wei Qin, Wanzeng Kong |
ICASSP | 5 |
| 2025 | PBFE-DAN:Personal Biological Feature Enhanced Domain Adaptation Network for Cross-Session Brainprint RecognitionabstractIn recent years, electroencephalogram (EEG) signal analysis has made significant strides, offering novel solutions for bolstering the security of brain–computer interfaces (BCI) and addressing long-standing vulnerabilities in traditional biometric authentication. Nonetheless, the low signal-to-noise ratio of EEG signals and their susceptibility to environmental influences lead to performance degradation when the model encounters data from subsequent unseen sessions. Targeting the crucial challenge of background noise in cross-session brainprint recognition, this paper presents the Personal Biological Feature Enhanced Domain Adaptation Network (PBFE-DAN). This Transformer-Based framework incorporates depth-ßwise separable convolutions with a Personal Biological Feature Enhancement Unit (PBFEU), which employs a gating mechanism to selectively amplify individual-specific EEG patterns while suppressing background noise. By emphasizing these unique EEG features, PBFE-DAN markedly improves robustness and generalizability across different sessions. Additionally, the self-attention module extracts global, session-invariant features that further enhance the stability and accuracy of cross-session brainprint recognition. Notably, PBFE-DAN can achieve state-of-the-art classification accuracies on both DSIRSVP and SEED-IV datasets, thereby underscoring the effectiveness of the proposed framework in mitigating background noise for cross-session brainprint recognition. Xuanyu Jin, Wanzeng Kong |
IJCNN | 3 |
| 2025 | MPFDAN: Multi-Perspective Feature Dynamic Adaptation Network for Domain Adaptive Object DetectionabstractBased on adversarial training and hierarchical alignment structure, domain adaptive object detection methods have made impressive progress. However, current adaptation methods treats each feature equally, and exerts unchanged alignment strength during alignment process, without sufficiently considering the transferability inconsistency of different features. Such static alignment can only achieve approximate alignment and will bring about negative transfer eventually. To address these issues, we propose the Multi-Perspective Feature Dynamic Adaptation Network (MPFDAN). In this network, the transferability of features is thoroughly considered and utilized from three different perspectives, allowing the alignment process to be dynamically adjusted in different ways. Firstly, regarding local transferability, Shannon entropy is used to adjust the weights of features in different local regions to focus more on regions with higher transferability. Next, from the perspective of global alignment, we dynamically adjust the alignment strength applied during the image-level adaptation process to avoid overfitting. Finally, category information is introduced to achieve category-aware instance-level adaptation, dynamically adjusted based on the differences in category transferability. Experiments on various domain transfer scenarios demonstrate that our MPFDAN outperforms all compared methods, thereby proving the effectiveness of our proposed approach. Wenchao Weng, Weichen Dai 0001, Andrzej Cichocki, Wanzeng Kong |
IJCNN | 6 |
| 2025 | SF-GAN: Semantic fusion generative adversarial networks for text-to-image synthesis
Xueqin Xiang, Wanzeng Kong, Jinliang Yao |
Expert Syst. Appl. | 3 |
| 2025 | Effective feature-sample co-clustering by adaptive feature-sample co-weighting
Yiyan Wang, Mimi Jin, Yong Peng 0001, Ziyue Yang 0007, Feiping Nie 0001, Andrzej Cichocki, Wanzeng Kong |
Inf. Sci. | 8 |
| 2025 | Dual-Brain EEG Decoding for Target Detection via Joint Learning in Shared and Private SpacesabstractHyperscanning enables simultaneous electroencephalography (EEG) recording from multiple individuals, facilitating collaborative brain activity to reduce individual biases and enhance the reliability of decision-making. The decoding of such collaborative paradigm tasks has traditionally relied solely on simple fusion methods based on each individual brain activity, without incorporating cross-brain coupling information. Inspired by social interaction studies on enhanced inter-brain synchrony in collaborative tasks using hyperscanning, we propose a joint learning framework for dual-brain target detection that integrates a shared space construction module and shared feature-guided module. The shared space construction module incorporates brain-to-brain coupling analysis to identify cross-brain synchrony, and further integrates shared and private features through a multi-head fusion mechanism for joint representation learning in shared feature-guided module. Experimental results show an average 10% improvement in balanced accuracy across 12 participant groups compared to traditional single-brain approaches, with some groups achieving up to a 5% gain over state-of-the-art (SOTA) methods. Notably, higher-performing groups exhibit stronger inter-brain coupling and more synchronized target-related responses. These findings advance the development of collaborative brain-computer interface (BCI) systems for more robust and effective target detection. Bingfeng He, Li Zhu 0005, Andrzej Cichocki, Wanzeng Kong |
IEEE Signal Process. Lett. | 5 |
| 2025 | QELDBA: Query-Efficient and Low Distortion Black-Box Attack for Brainprint RecognitionabstractWhile various deep learning techniques for electroencephalogram (EEG)-based brainprint recognition have achieved considerable success, these models remain vulnerable to adversarial attacks. However, existing black-box attack methods suffer from an inherent trade-off between query efficiency and distortion level. To address this challenge and further investigate the security risks of brainprint recognition systems in real-world black-box scenarios, we propose a query-efficient, low-distortion black-box attack method that targets the high-frequency components of EEG signals. Our approach innovatively selects sparse sampling points to estimate more accurate gradient information and leverages historical gradients to guide the prioritization of important points, thereby accelerating the attack process. The perturbations are applied in the high-frequency domain of the EEG signal to enhance stealth and effectiveness. Extensive experiments under black-box settings demonstrate that our method achieves state-of-the-art performance across two datasets and four models. Compared to existing methods, our approach significantly improves attack success rates while reducing the number of queries and minimizing distortion to imperceptible levels, thus achieving a superior balance between query efficiency and perturbation stealth. Jingsheng Qian, Hangjie Yi, Honggang Liu, Xuanyu Jin, Wanzeng Kong |
IEEE Signal Process. Lett. | 5 |
| 2025 | Independent Components Time-Frequency Purification With Channel Consensus Against Adversarial Attack in SSVEP-Based BCIsabstractThe Steady State Visual Evoked Potential (SSVEP) paradigm has been widely employed in various Brain-Computer Interface (BCI) systems. However, recent studies indicate that SSVEP is vulnerable to adversarial attacks, resulting in manipulated results and drastic degradation in recognition performance, which pose inconveniences and even risks to users. Noticing the fact that the adversarial attack on SSVEP is done by adding subtle waveform perturbations into random EEG channels, we propose Independent Components Time-Frequency Purification with Channel Consensus (ICTFP-CC) as a defensive strategy. In particular, we first detect and remove suspicious perturbations with independent component analysis from the time and frequency domain, and then reconstruct the purified EEG signals. Additionally, we introduce a voting mechanism to achieve channel consensus and enhance overall robustness. We conducted experiments on two public datasets and three SSVEP recognition algorithms. The results demonstrate that our method can significantly improve the classification accuracy and information transfer rate of attacked SSVEP signals by a maximum of 46.79 (%) and 62.87 (bits/min). Hangjie Yi, Jingsheng Qian, Yuhang Ming 0001, Wanzeng Kong |
IEEE Signal Process. Lett. | 4 |
| 2025 | Imagined Speech Decoding by Learning Consensus Graph From RKHS-Based Multi-View EEG Features
Zhenye Zhao, Yong Peng 0001, Kenneth P. Camilleri, Wanzeng Kong, Andrzej Cichocki |
IEEE Signal Process. Lett. | 4 |
| 2025 | ID-ProtoFormer: A Dynamic Identity Prototype-Infused Transformer for SSVEP-Based Biometric Recognition
Jiabin Zhu, Xuanyu Jin, Wanzeng Kong |
IEEE Signal Process. Lett. | 3 |
| 2025 | DARN: A Dual Attention Refinement Network for Enhancing Feature Robustness in VEP-Based EEG Biometrics
Honggang Liu, Han Yang 0003, Dongjun Liu, Hangjie Yi, Bingfeng He, Yong Peng 0001, Wanzeng Kong |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2025 | Reinforcement Learning Decoding Method of Multi-User EEG Shared Information Based on Mutual Information MechanismabstractThe multi-user motor imagery brain-computer interface (BCI) is a new approach that uses information from multiple users to improve decision-making and social interaction. Although researchers have shown interest in this field, the current decoding methods are limited to basic approaches like linear averaging or feature integration. They ignored accurately assessing the coupling relationship features, which results in incomplete extraction of multi-source information. To overcome these limitations, we propose a new reinforcement learning electroencephalography (EEG) decoding method based on mutual information mechanisms. Our method enhances the extraction of multi-source common information and uses a dynamic feedback model for inter-brain mutual information reward and punishment mechanisms in the reinforcement learning channel selection module. We feed the single-brain and inter-brain signals after channel selection into deep neural networks, which automatically extract coupled features. Finally, based on the attention indices calculated from EEG signals at prefrontal electrode positions, the output is obtained by voting. Our experimental results show that the average accuracy of dual-brain recognition is improved by 16% compared to single-brain mode. Furthermore, ablation experiments demonstrate that the reinforcement learning module and attention voting module enhance accuracy by 14.5% and 15.7%, respectively. Li Zhu 0005, Wanzeng Kong, Jianting Cao, Andrzej Cichocki |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Pattern-Matching Dynamic Memory Network for Dual-Mode Traffic PredictionabstractIn recent years, deep learning has increasingly gained attention in the field of traffic prediction. Existing traffic prediction models often rely on GCNs or attention mechanisms withO(N2) complexity to dynamically extract traffic node features, which lack efficiency and are not lightweight. Additionally, these models typically only utilize historical data for prediction, without considering the impact of the target information on the prediction. To address these issues, we propose a Pattern-Matching Dynamic Memory Network (PM-DMNet). Unlike traditional attention and graph convolution-based approaches, PM-DMNet employs a novel dynamic memory network that stores the most representative traffic patterns from historical data in a memory matrix through training. It captures traffic pattern features by comparing the similarity between the memory matrix and the current traffic state. This method not only achieves excellent predictive performance but also significantly reduces computational complexity toO(N). The PM-DMNet also introduces two prediction methods: Recursive Multi-step Prediction (RMP) and Parallel Multi-step Prediction (PMP), which leverage the time features of the prediction targets to assist in the prediction process. Furthermore, a transfer attention mechanism is integrated into PMP, transforming historical data features to better align with the predicted target states, thereby capturing trend changes more accurately and reducing errors. Extensive experiments demonstrate the superiority of the proposed model over existing benchmarks. The source codes are available at: https://github.com/wengwenchao123/PM-DMNet Wenchao Weng, Mei Wu 0001, Hanyu Jiang 0001, Wanzeng Kong, Xiangjie Kong 0001, Feng Xia 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | MECG: modality-enhanced convolutional graph for unbalanced multimodal representations
Jiajia Tang, Binbin Ni, Yutao Yang, Yu Ding 0001, Wanzeng Kong |
J. Supercomput. | 5 |
| 2025 | Brain-Machine Cross-Modal Alignment via Sample Relational Learning for Visual ClassificationabstractRecent works on visual classification tasks have leveraged EEG signals to provide additional supervisory information, further improving the performance of the models on natural images. However, previous methods often force machine models to directly match EEG signals, which involves the transfer of modal-specific representations, leading to potentially distorted alignment of modal-shared representations. Moreover, focusing solely on aligning individual sample features neglects the alignment of relationships between samples, making it difficult to capture the potential relational reasoning capabilities in EEG signals. This relational reasoning ability is key to the human brain’s outstanding performance in visual classification tasks. Similarly, for a machine model, the complex relationships between instances are more critical than individual instances. Inspired by this, our idea is to enhance machine visual classification capabilities by imparting human-like relational reasoning, encouraging machine models to focus on the relational structure within EEG signals. To this end, we propose a brain-machine relation alignment method that constructs a cognitive model and a visual model to process EEG signals and visual images, respectively. Instead of forcing the visual model to mimic the output of an individual EEG data sample represented by the cognitive model, we encourage it to learn the mutual relations of EEG data samples. By penalizing the difference in relational structures between EEG signals and visual images, we facilitate the transfer of relational knowledge. Experiments demonstrate that the proposed method significantly improves the classification performance of the visual model. This highlights the potential of relational alignment as a robust mechanism for integrating human relational reasoning into machine learning models. Dongjun Liu, Weichen Dai 0001, Honggang Liu, Hangjie Yi, Wanzeng Kong |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | Fine-grained Semantic Disentanglement Network for Multimodal Sarcasm AnalysisabstractMultimodal sarcasm analysis is one of the most challenging research branch of the sentiment analysis area, due to the presence of cross-modality incongruity. However, existing works mainly attend to the coarse-grained incongruity analysis, and totally ignore the sentiment semantic coupling issue. This indeed limits the discriminate capability and robustness of the sarcasm analysis model. In order to address the above issue, we propose a novel Fine-grained Semantic Disentanglement Network (FSDN). Specifically, the intra-modality semantic disentanglement is performed to investigate the more intrinsic semantic cues of the same modality. Additionally, the inter-modality semantic disentanglement is leveraged to simultaneously facilitate the common and intrinsic semantic cues across modalities. Furthermore, the dual-spatial semantic interaction block is presented to explore the long-range cross-spatial semantic context between the obtained verbal and non-verbal semantic space with the global view. The above semantic disentanglement processes with both local and global views significantly unleash much more robustness even for the sarcasm case consisting of multiple semantic message. Various experiments indicate that the FSDN can receive state-of-the-art or competitive performance. Jiajia Tang, Binbin Ni, Feiwei Zhou, Dongjun Liu, Yu Ding 0001, Yong Peng 0001, Andrzej Cichocki, Qibin Zhao, Wanzeng Kong |
ACM Trans. Multim. Comput. Commun. Appl. | 9 |
| 2025 | Hybrid Feature Integrated Transformer for 3D Hand Reconstruction from a Single RGB ImageabstractReconstructing a 3D hand from a single RGB image is a very challenging task. Most of the existing Transformer-based 3D hand reconstructing methods do not fully consider the local spatial information from low-level image features, which would be crucial for capturing fine details and accurate shapes of the hand. Consequently, this oversight often leads to reconstructed hands that lack the precision and realism necessary for many applications, such as augmented reality, and hand gesture recognition. To address this limitation, in this paper, we propose a novel and efficient method named HybridMETRO to both utilize low-level and high-level image features for accurate reconstructing 3D hand pose and mesh vertices from a single RGB image. Specifically, we introduce the deformable attention into the encoder of Transformer, making it no longer limited by the length of the image feature sequence. Based on the above mechanism, we further propose an interleaved updating multi-scale feature encoder to fuse low-level and high-level features. Moreover, we incorporate the Graph Convolutional Residual (GCR) module to build a novel decoder to capture explicit semantic connections between grid vertices and thus improve spatial locality of extracted features. Experimental results demonstrate that, when compared with state-of-the-art methods, our proposed HybridMETRO could achieve better performance with significantly smaller model parameters that are about half of METRO’s and a quarter of HandOccNet’s. Xueqin Xiang, Wanzeng Kong, Jinliang Yao |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Implicit guidance for enhancing low-light optical flow estimation via channel attention networks
Weichen Dai 0001, Hexing Wu, Xiaoyang Weng, Wanzeng Kong |
Vis. Comput. | 4 |
| 2024 | The Impact of Dynamic Icons on Mobile APP Interfaces: Evidence from EEG and Eye-tracking SignalsabstractThis study investigates the cognitive impact of dynamic icons in mobile interfaces by integrating electroencephalography (EEG) and eye-tracking technologies. Traditional research on mobile app interface design has relied mainly on questionnaire and eye-tracking methods for behavioral analysis. This research adds a new dimension by examining the neural mechanisms associated with dynamic icons. We employed EEG to analyze channel-wise power spectrum density (PSD), focusing on the alpha and theta frequency bands related to attention and working memory. Concurrently, eye-tracking data were analyzed through Areas of Interest (AOIs) and fixation metrics to assess visual attention patterns. The results indicate that dynamic icons significantly enhance neural activity, with a 15% increase in alpha band power and a 20% increase in theta band power compared to static icons. Additionally, eye-tracking data show a 30% increase in total fixation duration on AOIs containing dynamic icons, particularly in the left-top quarter of the mobile interface. This effect was observed without changes in the first fixation duration, suggesting that dynamic icons have a stronger impact on sustained attention rather than on initial capture. These findings highlight that dynamic icons not only attract and maintain visual attention more effectively but also enhance cognitive processing efficiency. This study provides valuable insights for optimizing mobile app interface design, emphasizing the benefits of incorporating dynamic elements to improve user engagement and interface effectiveness. Ruizhe Yang, Jiaxuan Qin, Letao Fang, Haojie Tao, Li Zhu 0005, Xuanyu Jin, Wanzeng Kong |
BIBM | 8 |
| 2024 | Comparisons on Perception Mechanism of Mental Rotation Between Health and Stroke Groups with EEG IndicatorsabstractClinically, motor rehabilitation and mechanism research have garnered increasing attention. However, cognitive impairments often accompany stroke patients. Mental rotation is a crucial task for cognitive evaluation and training. In our study, we proposed a method to compare mental rotation perception mechanisms between healthy individuals and stroke patients using EEG indicators. The experiment was designed to accommodate stroke patients. In the data analysis module, we utilized power spectral density (PSD) to explore single-channel frequency domain indicators and phase-locked value (PLV) to analyze connectivity between channels. Mechanism analysis based on EEG indicators considered three main conditions: counter-clockwise and clockwise rotation, lesion area versus functional area across strokes, and mental rotation perception sub-stages. Our experimental results show significant differences between the stroke group and health group in the θ and α frequency bands across counter-clock and clock-wise perception and lesion connectivity conditions. The stroke group reveals a compensatory effect in the lesion areas, with higher PLV values (averaging 0.2101) compared to those in common cognitive functional areas of the health group and the connectivity within lesion areas was significantly lower (averaging 0.0310) than that outside the lesion areas. These findings could assist in stroke treatment and monitoring rehabilitation progress. Lingmin Zhou, Li Zhu 0005, Haibin Xia, Guifen Yang, Xuanyu Jin, Wanzeng Kong |
BIBM | 7 |
| 2024 | An Automated Sleep Staging Method with EEG-based Sleep Structure Computation
Ruixiang Liao, Li Zhu 0005, Wanzeng Kong |
CogSci | 3 |
| 2024 | Enhanced Coherence-Aware Network with Hierarchical Disentanglement for Aspect-Category Sentiment AnalysisabstractAspect-category-based sentiment analysis (ACSA), which aims to identify aspect categories and predict their sentiments has been intensively studied due to its wide range of NLP applications. Most approaches mainly utilize intrasentential features. However, a review often includes multiple different aspect categories, and some of them do not explicitly appear in the review. Even in a sentence, there is more than one aspect category with its sentiments, and they are entangled intra-sentence, which makes the model fail to discriminately preserve all sentiment characteristics. In this paper, we propose an enhanced coherence-aware network with hierarchical disentanglement (ECAN) for ACSA tasks. Specifically, we explore coherence modeling to capture the contexts across the whole review and to help the implicit aspect and sentiment identification. To address the issue of multiple aspect categories and sentiment entanglement, we propose a hierarchical disentanglement module to extract distinct categories and sentiment features. Extensive experimental and visualization results show that our ECAN effectively decouples multiple categories and sentiments entangled in the coherence representations and achieves state-of-the-art (SOTA) performance. Our codes and data are available online: https://github.com/cuijin-23/ECAN. Jin Cui 0005, Fumiyo Fukumoto, Xinfeng Wang, Yoshimi Suzuki, Jiyi Li, Noriko Tomuro, Wanzeng Kong |
LREC/COLING | 7 |
| 2024 | An EEG-based Decoding Method for Motor Imagery Intentions in Mixed-Subject Settings with Adversarial DisentanglementabstractSmall samples and significant inter-subject variability are the two main challenges in current electroencephalogram (EEG) based Motor Imagery (MI) Brain-Computer Interface (BCI) decoding methods. To overcome these challenges, we proposed an EEG-based decoding method for MI intentions in mixed-subject settings with adversarial disentanglement, which utilize the EEG decoding focusing on MI related task information and decrease the inference of inter-subject variability under mixed-subject setting. The method includes three main modules: data augmentation module, dual-label training module, disentanglement training module. It first uses a mixed-subject settings approach, which involves shuffling the data from all subjects to augment the data for a single model. We then create a dual-label dataset using motor imagery labels and identity labels. Finally, a disentanglement training strategy is employed to optimize the negative entropy loss, measuring the inter-subject variability. Our experiment results show that our method achieves higher accuracy compared to traditional one-to-one model training methods and lower variance with the mixed-subject settings. It has achieved a $\mathbf{7 5. 9 3 \%}$ average classification accuracy across four classes on the BCIC-IV-2a dataset with the best classification accuracy reaches $\mathbf{9 0. 2 8 \%}$, indicating that our model has the capability to disentangle identity-related information during the feature extraction and has more stable performance across different subjects. Li Zhu 0005, Jiazheng Zhang, Chengrui Chen, Andrzej Cichocki, Jianghan Yan, Wanzeng Kong |
CW | 9 |
| 2024 | Exploration of Common Cognitive Perception Biomarker for Human: An EEG Empirical Analysis StudyabstractCognitive perception is a fundamental human physiological function, closely interrelated with decision-making, belief formation, and problem-solving. Decoding the mechanisms of cognitive perception is crucial. Advances in brain signal processing and imaging have made the study of human cognitive perception more precise and convenient. In this paper, we explore potential common biomarkers in terms of spatial rotation (SR) and working memory (WM) on basis of their similarities. We propose a power spectrum density (PSD)-based analysis pipeline that includes statistical analysis, ratio definition, impact measurement of brain state, and inter-stage analysis. Our experimental results show that the average distribution of PSD is similar across different frequency bands, with corresponding specialization observed during inter-stage analysis. Li Zhu 0005, Lingmin Zhou, Haibin Xia, Jianghan Yan, Guifen Yang, Wanzeng Kong |
CW | 7 |
| 2024 | AEGIS-Net: Attention-Guided Multi-Level Feature Aggregation for Indoor Place RecognitionabstractWe present AEGIS-Net, a novel indoor place recognition model that takes in RGB point clouds and generates global place descriptors by aggregating lower-level color, geometry features and higher-level implicit semantic features. However, rather than simple feature concatenation, self-attention modules are employed to select the most important local features that best describe an indoor place. Our AEGIS-Net is made of a semantic encoder, a semantic decoder and an attention-guided feature embedding. The model is trained in a 2-stage process with the first stage focusing on an auxiliary semantic segmentation task and the second one on the place recognition task. We evaluate our AEGIS-Net on the ScanNetPR dataset and compare its performance with a pre-deep-learning feature-based method and five state-of-the-art deep-learning-based methods. Our AEGIS-Net achieves exceptional performance and outperforms all six methods. Yuhang Ming 0001, Jian Ma 0001, Xingrui Yang 0001, Weichen Dai 0001, Yong Peng 0001, Wanzeng Kong |
ICASSP | 6 |
| 2024 | Label Rectified and Graph Adaptive Semi-Supervised Regression for Electrode Shifted Gesture RecognitionabstractSurface electromyography (sEMG) noninvasively records muscle activities. It provides valuable information about muscle contractions and enables real-time decoding into hand gestures. Recently many studies have successfully demonstrated this capability. However, the accuracy of gesture recognition decreases significantly due to electrode shifts. Without increasing the density of electrodes which may cause the curse of dimensionality and result in higher costs, we propose a label rectified and graph adaptive semi-supervised regression (LRGASR) model for electrode shifted gesture recognition. LRGASR on one hand learns an optimal graph to characterize the underlying semantic connectionship of both non-shifted and shifted sEMG samples and takes advantage of label rectification to reduce the feature-label inconsistency of shifted ones. Experimental results show that LRGASR achieved the average recognition accuracies 78.20% and 87.28% on the SeNic and ISRMyo sEMG data sets, which outperforms six existing models. Chengxi Zhu, Yong Peng 0001, Yinfeng Fang, Wanzeng Kong |
ICASSP | 4 |
| 2024 | Time-Frequency Jointed Imperceptible Adversarial Attack to Brainprint Recognition with Deep Learning ModelsabstractEEG-based brainprint recognition with deep learning models has garnered much attention in biometric identification. Yet, studies have indicated vulnerability to adversarial attacks in deep learning models with EEG inputs. In this paper, we introduce a novel adversarial attack method that jointly attacks time-domain and frequency-domain EEG signals by employing wavelet transform. Different from most existing methods which only target time-domain EEG signals, our method not only takes advantage of the time-domain attack’s potent adversarial strength but also benefits from the imperceptibility inherent in frequency-domain attack, achieving a better balance between attack performance and imperceptibility. Extensive experiments are conducted in both white- and grey-box scenarios and the results demonstrate that our attack method achieves state-of-the-art attack performance on three datasets and three deep-learning models. In the meanwhile, the perturbations in the signals attacked by our method are barely perceptible to the human visual system. Hangjie Yi, Yuhang Ming 0001, Dongjun Liu, Wanzeng Kong |
ICME | 4 |
| 2024 | Fine-grained Dual-space Context Analysis Network for Multimodal Emotion RecognitionabstractMultimodal emotion recognition has received widespread attention in variety of domains, which can utilize multiple modalities of emotional information to improve the performance of emotion recognition. However, existing coarse-grained works mainly attend to the temporal domain context and totally ignore the temporal-spatial domain context, which results in the significant deterioration of the emotion analysis performance. In this work, the fine-grained dual-space context analysis network (FDCAN) is proposed to fully investigate the fine-grained emotion context among the joint temporal-spatial emotion representative space. Specifically, the depthwise separable convolution based operation is leveraged to exploit the temporal and spatial space from EEG and EOG modality. Furthermore, the correlation analysis based technique is introduced to investigate cross-modality correlation messages from the above obtained spaces, leading to the coupled temporal-spatial representative space. Additionally, the attention mechanism based procedure is performed to deal with the comprehensive and sophisticated multi-modality emotion context from the coupled representative space. Note that, the above carefully designed hierarchical dual-space emotion context procedure indeed provide a new detection and bears the strong potential to facilitate the emotion analysis. We validated the effectiveness of the proposed framework on the popular and public multimodal emotion analysis benchmark DEAP. The experimental results demonstrated that our model achieved better performance of 97.4% and 96.9% for binary and four-class emotion classification task. Jiajia Tang, Wanzeng Kong, Yanyan Ying |
IJCNN | 3 |
| 2024 | A Privacy-Preserving Brainprint Recognition System Based on Feature Homomorphic EncryptionabstractIn recent years, the extensive use of electroencephalogram (EEG) biometric recognition technology has raised significant privacy concerns, particularly when the biometric computation process is carried out on an untrusted server. This paper discusses the practical scenario of performing EEG biometric recognition computation on an untrusted server within EEG biometric recognition systems. To address the privacy issues in previous systems, we propose, for the first time, a privacy-preserving EEG biometric recognition systems using homomorphic encryption. This approach enables the transmission and computation of EEG data without exposing the EEG information. Additionally, due to the large volume of raw EEG data, applying homomorphic encryption to the raw EEG data would significantly increase the time and space overhead of the homomorphic computation. This paper further proposes the application of homomorphic encryption to brainprint feature, as the data volume of brainprint feature is small, resulting in a significant reduction in computational time and space overhead. We conducted a quantitative evaluation of the system’s classification accuracy, time overhead, and space overhead, and compared it with the unencrypted system. Experimental results demonstrate that our approach maintains classification accuracy, meets the real-time requirements of brainprint recognition systems in terms of time overhead, and keeps space overhead within an acceptable range, while preserving the privacy and security of user information. Hangjie Yi, Wanzeng Kong |
IJCNN | 3 |
| 2024 | Temporal-channel cascaded transformer for imagined handwriting character recognition
Wenhui Zhou 0001, Liangyan Mo, Wanzeng Kong, Guojun Dai |
Neurocomputing | 6 |
| 2024 | DSFE: Decoding EEG-Based Finger Motor Imagery Using Feature-Dependent Frequency, Feature Fusion and Ensemble LearningabstractAccurate decoding finger motor imagery is essential for fine motor control using EEG signals. However, decoding finger motor imagery is particularly challenging compared with ordinary motor imagery. This paper proposed a novel EEG decoding method of feature-dependent frequency band selection, feature fusion, and ensemble learning (DSFE) for finger motor imagery. First, a feature-dependent frequency band selection method based on correlation coefficient (FDCC) was proposed to select feature-specific effective bands. Second, a feature fusion method was proposed to fuse different types of candidate features to produce multiple refined sets of decoding features. Finally, an ensemble model using the weighted voting strategy was proposed to make full use of these diverse sets of final features. The results on a public EEG dataset of five fingers motor imagery showed that the DSFE method is effective and achieves the highest decoding accuracy of 50.64%, which is 7.64% higher than existing studies using exactly the same data. The experiments further revealed that both the effective frequency bands of different subjects and the effective frequency bands of different types of features are different in finger motor imagery. Furthermore, compared with two-hand motor imagery, the effective decoding information of finger motor imagery is transferred to the lower frequency. The idea and findings in this paper provide a valuable perspective for understanding fine motor imagery in-depth. Kun Yang 0002, Ruochen Li 0006, Jing Xu 0004, Li Zhu 0005, Wanzeng Kong |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | A General DNA-Like Hybrid Symbiosis Framework: An EEG Cognitive Recognition MethodabstractIn electroencephalogram (EEG) cognitive recognition research, the combined use of artificial neural networks (ANNs) and spiking neural networks (SNNs) plays an important role to realize different categories of recognition tasks. However, most of the existing studies focus on the unidirectional interaction between an ANN and a SNN, which may be overly dependent on the performance of ANNs or SNNs. Inspired by the symbiosis phenomenon in nature, in this study, we propose a general DNA-like Hybrid Symbiosis (DNA-HS) framework, which enables mutual learning between the ANN and the SNN generated by this ANN through parametric genetic algorithm and bidirectional interaction mechanism to enhance the optimization ability of the model parameters, resulting in a significant improvement of the performance of the DNA-HS framework in all aspects. By comparing with seven typical EEG cognitive recognition models, the performance of the seven hybrid network frameworks constructed using this method on different EEG-based cognitive recognition tasks are all improved to different degrees, verifying the effectiveness of the proposed method. This unified hybrid network framework similar to the DNA structure is expected to open up a new approach and form a new research paradigm for EEG-based cognitive recognition task. Hong Zeng 0002, Yue Zhao 0030, Fabio Babiloni, Ming Tao 0004, Wanzeng Kong, Guojun Dai |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | DMF-GAN: Deep Multimodal Fusion Generative Adversarial Networks for Text-to-Image SynthesisabstractText-to-image synthesis aims to generate highquality realistic images conditioned on text description. The great challenge of this task depends on deeply and seamlessly integrating image and text information. Thus, in this paper, we propose a deep multimodal fusion generative adversarial networks (DMF-GAN) that allows effective semantic interactions for finegrained text-to-image generation. Specifically, through a novel recurrent semantic fusion network, DMF-GAN could consistently manipulate global assignment of text information among isolated fusion blocks. With the assistance of a multi-head attention module, DMF-GAN could model word information from different perspectives and further improve the semantic consistency. In addition, a word-level discriminator is proposed to provide the generator with fine-grained feedback related to each word. Compared with current state-of-the-art methods, our proposed DMFGAN could efficiently synthesize realistic and text-alignment images and achieve better performance on challenging benchmarks. The code link:https://github.com/xueqinxiang/DMF-GAN Xueqin Xiang, Wanzeng Kong, Yong Peng 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Unbiased Semantic Representation Learning Based on Causal Disentanglement for Domain GeneralizationabstractDomain generalization primarily mitigates domain shift among multiple source domains, generalizing the trained model to an unseen target domain. However, the spurious correlation usually caused by context prior (e.g., background) makes it challenging to get rid of the domain shift. Therefore, it is critical to model the intrinsic causal mechanism. The existing domain generalization methods only attend to disentangle the semantic and context-related features by modeling the causation between input and labels, which totally ignores the unidentifiable but important confounders. In this article, a Causal Disentangled Intervention Model (CDIM) is proposed for the first time, to the best of our knowledge, to construct confounders via causal intervention. Specifically, a generative model is employed to disentangle the semantic and context-related features. The contextual information of each domain from generative model can be considered as a confounder layer, and the center of all context-related features is utilized for fine-grained hierarchical modeling of confounders. Then the semantic and confounding features from each layer are combined to train an unbiased classifier, which exhibits both transferability and robustness across an unknown distribution domain. CDIM is evaluated on three widely recognized benchmark datasets, namely, Digit-DG, PACS, and NICO, through extensive ablation studies. The experimental results clearly demonstrate that the proposed model achieves state-of-the-art performance. Xuanyu Jin, Wanzeng Kong, Jiajia Tang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | EEG-based emotion recognition with cascaded convolutional recurrent neural networks
Yu Zhang 0009, Yuliang Ma 0002, Yunyuan Gao, Wanzeng Kong |
Pattern Anal. Appl. | 5 |
| 2023 | Brain-Machine Coupled Learning Method for Facial Emotion RecognitionabstractNeural network models of machine learning have shown promising prospects for visual tasks, such as facial emotion recognition (FER). However, the generalization of the model trained from a dataset with a few samples is limited. Unlike the machine, the human brain can effectively realize the required information from a few samples to complete the visual tasks. To learn the generalization ability of the brain, in this article, we propose a novel brain-machine coupled learning method for facial emotion recognition to let the neural network learn the visual knowledge of the machine and cognitive knowledge of the brain simultaneously. The proposed method utilizes visual images and electroencephalogram (EEG) signals to couple training the models in the visual and cognitive domains. Each domain model consists of two types of interactive channels, common and private. Since the EEG signals can reflect brain activity, the cognitive process of the brain is decoded by a model following reverse engineering. Decoding the EEG signals induced by the facial emotion images, the common channel in the visual domain can approach the cognitive process in the cognitive domain. Moreover, the knowledge specific to each domain is found in each private channel using an adversarial strategy. After learning, without the participation of the EEG signals, only the concatenation of both channels in the visual domain is used to classify facial emotion images based on the visual knowledge of the machine and the cognitive knowledge learned from the brain. Experiments demonstrate that the proposed method can produce excellent performance on several public datasets. Further experiments show that the proposed method trained from the EEG signals has good generalization ability on new datasets and can be applied to other network models, illustrating the potential for practical applications. Dongjun Liu, Weichen Dai 0001, Hangkui Zhang, Xuanyu Jin, Jianting Cao, Wanzeng Kong |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | BAFN: Bi-Direction Attention Based Fusion Network for Multimodal Sentiment AnalysisabstractAttention-based networks currently identify their effectiveness in multimodal sentiment analysis. However, existing methods ignore the redundancy of auxiliary modalities. More importantly, existing methods only attend to top-down attention (static process) or down-top attention (implicit process), leading to the coarse-grained multimodal sentiment context. In this paper, during the preprocessing period, we first propose the multimodal dynamic enhanced block to capture the intra-modality sentiment context. This can effectively decrease the intra-modality redundancy of auxiliary modalities. Furthermore, the bi-direction attention block is proposed to capture fine-grained multimodal sentiment context via the novel bi-direction multimodal dynamic routing mechanism. Specifically, the bi-direction attention block first highlights the explicit and low-level multimodal sentiment context. Then, the low-level multimodal context is transmitted to a carefully designed bi-direction multimodal dynamic routing procedure. This allows us to dynamically update and investigate high-level and much more fine-grained multimodal sentiment contexts. The experiments demonstrate that our fusion network can achieve state-of-the-art performance. Notably, our model outperforms the best baseline on the metric ‘Acc-7’ with an improvement of 6.9%. Jiajia Tang, Dongjun Liu, Xuanyu Jin, Yong Peng 0001, Qibin Zhao, Yu Ding 0001, Wanzeng Kong |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2023 | Joint EEG Feature Transfer and Semisupervised Cross-Subject Emotion RecognitionabstractDue to the weak and nonstationary properties, electroencephalogram (EEG) data present significant individual differences. To align data distributions of different subjects, transfer learning showed promising performance in cross-subject EEG emotion recognition. However, most of the existing models sequentially learned the domain-invariant features and estimated the target domain label information. Such a two-stage strategy breaks the inner connections of both processes, inevitably causing the suboptimality. In this article, we propose a joint EEG feature transfer and semisupervised cross-subject emotion recognition model in which the shared subspace projection matrix and target label are jointly optimized toward the optimum. Extensive experiments are conducted on SEED-IV and SEED, and the results show that the emotion recognition performance is significantly enhanced by the joint learning mode and the spatial-frequency activation patterns of critical EEG frequency bands and brain regions in cross-subject emotion expression are quantitatively identified by analyzing the learned shared subspace. Yong Peng 0001, Honggang Liu, Wanzeng Kong, Feiping Nie 0001, Bao-Liang Lu, Andrzej Cichocki |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | MMT: Multi-way Multi-modal Transformer for Multimodal LearningabstractThe heart of multimodal learning research lies the challenge of effectively exploiting fusion representations among multiple modalities.However, existing two-way cross-modality unidirectional attention could only exploit the intermodal interactions from one source to one target modality. This indeed fails to unleash the complete expressive power of multimodal fusion with restricted number of modalities and fixed interactive direction.In this work, the multiway multimodal transformer (MMT) is proposed to simultaneously explore multiway multimodal intercorrelations for each modality via single block rather than multiple stacked cross-modality blocks. The core idea of MMT is the multiway multimodal attention, where the multiple modalities are leveraged to compute the multiway attention tensor. This naturally benefits us to exploit comprehensive many-to-many multimodal interactive paths. Specifically, the multiway tensor is comprised of multiple interconnected modality-aware core tensors that consist of the intramodal interactions. Additionally, the tensor contraction operation is utilized to investigate intermodal dependencies between distinct core tensors.Essentially, our tensor-based multiway structure allows for easily extending MMT to the case associated with an arbitrary number of modalities. Taking MMT as the basis, the hierarchical network is further established to recursively transmit the low-level multiway multimodal interactions to high-level ones. The experiments demonstrate that MMT can achieve state-of-the-art or comparable performance. Jiajia Tang, Xuanyu Jin, Wanzeng Kong, Yu Ding 0001, Qibin Zhao |
IJCAI | 5 |
| 2022 | Dynamically Adjust Word Representations Using Unaligned Multimodal InformationabstractMultimodal Sentiment Analysis is a promising research area for modeling multiple heterogeneous modalities. Two major challenges that exist in this area are a) multimodal data is unaligned in nature due to the different sampling rates of each modality, and b) long-range dependencies between elements across modalities. These challenges increase the difficulty of conducting efficient multimodal fusion. In this work, we propose a novel end-to-end network named Cross Hyper-modality Fusion Network (CHFN). The CHFN is an interpretable Transformer-based neural model that provides an efficient framework for fusing unaligned multimodal sequences. The heart of our model is to dynamically adjust word representations in different non-verbal contexts using unaligned multimodal sequences. It is concerned with the influence of non-verbal behavioral information at the scale of the entire utterances and then integrates this influence into verbal expression. We conducted experiments on both publicly available multimodal sentiment analysis datasets CMU-MOSI and CMU-MOSEI. The experiment results demonstrate that our model surpasses state-of-the-art models. In addition, we visualize the learned interactions between language modality and non-verbal behavior information and explore the underlying dynamics of multimodal language data. Jiwei Guo, Jiajia Tang, Weichen Dai 0001, Yu Ding 0001, Wanzeng Kong |
ACM Multimedia | 5 |
| 2022 | A new patterns of self-organization activity of brain: Neural energy codingabstractAccording to the basic principles and methods of information theory, the operation way of neural coding is studied and analyzed by using the minimum mutual information and the maximum entropy principle. This paper describes how the principles of minimum mutual information and maximum entropy are used to evaluate the amount of information in neural responses. Its main contribution is as follows: (1) that the expression of neural information is closely related to the utilization of neural energy, and it is found that the highly evolved nervous system strictly follows the two basic principles of economy and efficiency in energy consumption and utilization; (2) In order to verify the relationship between neural information processing and energy utilization, this paper uses the concept of energy-efficiency ratio to measure the economy and high efficiency of the nervous system in term of energy utilization by using the maximum entropy principle; (3) The numerical results show that the energy consumed by the nervous system reflects not only the internal law of neural information conduction and processing, but also the self-organization structure of neural information coding. The results suggest that energy neural coding, a novel neural information processing method, can be used to understand how brain activity works. Such a coding pattern can not only be extended to research the large-scale neuroscience field, but also unify brain models at all levels by use of the energy theory. This will provide a scientific theoretical basis for the exploration of how the brain works and the computational principles of brain-like artificial intelligence. Jinchao Zheng, Rubin Wang, Wanzeng Kong |
Inf. Sci. | 3 |
| 2022 | Adaptive multi-task learning using lagrange multiplier for automatic art analysis
Xueqin Xiang, Wanzeng Kong, Yong Peng 0001, Jinliang Yao |
Multim. Tools Appl. | 3 |
| 2022 | Joint Feature Adaptation and Graph Adaptive Label Propagation for Cross-Subject Emotion Recognition From EEG SignalsabstractThough Electroencephalogram (EEG) could objectively reflect emotional states of our human beings, its weak, non-stationary, and low signal-to-noise properties easily cause the individual differences. To enhance the universality of affective brain-computer interface systems, transfer learning has been widely used to alleviate the data distribution discrepancies among subjects. However, most of existing approaches focused mainly on the domain-invariant feature learning, which is not unified together with the recognition process. In this paper, we propose a joint feature adaptation and graph adaptive label propagation model (JAGP) for cross-subject emotion recognition from EEG signals, which seamlessly unifies the three components of domain-invariant feature learning, emotional state estimation and optimal graph learning together into a single objective. We conduct extensive experiments on two benchmark SEED_IV and SEED_V data sets and the results reveal that 1) the recognition performance is greatly improved, indicating the effectiveness of the triple unification mode; 2) the emotion metric of EEG samples are gradually optimized during model training, showing the necessity of optimal graph learning, and 3) the projection matrix-induced feature importance is obtained based on which the critical frequency bands and brain regions corresponding to subject-invariant features can be automatically identified, demonstrating the superiority of the learned shared subspace. Yong Peng 0001, Wanzeng Kong, Feiping Nie 0001, Bao-Liang Lu, Andrzej Cichocki |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | Analysis of Functional Corticomuscular Coupling Based on Multiscale Transfer Spectral EntropyabstractFunctional corticomuscular coupling (FCMC) between the cerebral motor cortex and muscle activity reflects multi-layer and nonlinear interactions in the sensorimotor system. Considering the inherent multiscale characteristics of physiological signals, we proposed multiscale transfer spectral entropy (MSTSE) and introduced the unidirectionally coupled Hénon maps model to verify the effectiveness of MSTSE. We recorded electroencephalogram (EEG) and surface electromyography (sEMG) in steady-state grip tasks of 29 healthy participants and 27 patients. Then, we used MSTSE to analyze the FCMC base on EEG of the bilateral motor areas and the sEMG of the flexor digitorum superficialis (FDS). The results show that MSTSE is superior to transfer spectral entropy (TSE) method in restraining the spurious coupling and detecting the coupling more accurately. The coupling strength was higher in the β1, β2, and γ2 bands, among which, it was highest in the β1 band, and reached its maximum at the 22-30 scale. On the directional characteristics of FCMC, the coupling strength of EEG→sEMG is superior to the opposite direction in most cases. In addition, the coupling strength of the stroke-affected side was lower than that of healthy controls' right hand in the β1 and β2 bands and the stroke-unaffected side in the β1 band. The coupling strength of the stroke-affected side was higher than that of the stroke-unaffected side and the right hand of healthy controls in the sEMG→EEG direction of γ2 band. This study provides a new perspective and lays a foundation for analyzing FCMC and motor dysfunction. Xugang Xi, Jinsuo Ding, Yun-Bo Zhao, Ting Wang 0021, Wanzeng Kong |
IEEE J. Biomed. Health Informatics | 6 |
| 2021 | CTFN: Hierarchical Learning for Multimodal Sentiment Analysis Using Coupled-Translation Fusion NetworkabstractJiajia Tang, Kang Li, Xuanyu Jin, Andrzej Cichocki, Qibin Zhao, Wanzeng Kong. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jiajia Tang, Xuanyu Jin, Andrzej Cichocki, Qibin Zhao, Wanzeng Kong |
ACL/IJCNLP (1) | 6 |
| 2021 | Fuzzy graph clustering
Yong Peng 0001, Feiping Nie 0001, Wanzeng Kong |
Inf. Sci. | 4 |
| 2021 | Emotion recognition based on sparse representation of phase synchronization features
Wanzeng Kong, Xulin Song |
Multim. Tools Appl. | 1 |
| 2021 | Transfer of semi-supervised broad learning system in electroencephalography signal classification
Yukai Zhou, Qingshan She, Yuliang Ma 0002, Wanzeng Kong, Yingchun Zhang |
Neural Comput. Appl. | 4 |
| 2020 | Joint Semi-Supervised Feature Auto-Weighting and Classification Model for EEG-Based Cross-Subject Sleep Quality EvaluationabstractMeasuring the sleep quality is important or even crucial for people who are engaged in dangerous jobs such as the high-speed train drivers. Since the scalp EEG data are generated by the neural activities of the brain cortex, it is collected from subjects with different hours of sleep time (4 hours, 6 hours and 8 hours) to conduct sleep quality evaluation. To suppress the cross-subject variances of EEG data, in this paper, we propose a joint feature auto-weighting and semi-supervised classification model, termed GRLSR, which is formulated by introducing an auto-weighting variable into the least square regression to adaptively and quantitatively measure the importance of each dimension of the feature. Once the model is solved, besides the measurement results, we can use the auto-weighting variable to 1) analyze the importance of each frequency band in sleep quality expression and 2) identify the capacity of different channels connecting to the sleep effect. Therefore, the proposed GRLSR is a pure data-driven computing model for EEG-based cross-subject sleep quality evaluation. Experimental results show its effectiveness. Yong Peng 0001, Qingxi Li, Wanzeng Kong, Bao-Liang Lu, Andrzej Cichocki |
ICASSP | 3 |
| 2020 | A Factorized Extreme Learning Machine and Its Applications in EEG-Based Emotion Recognition
Yong Peng 0001, Rixin Tang, Wanzeng Kong, Feiping Nie 0001 |
ICONIP (5) | 3 |
| 2020 | Joint low-rank representation and spectral regression for robust subspace learning
Yong Peng 0001, Leijie Zhang, Wanzeng Kong, Fei-wei Qin |
Knowl. Based Syst. | 3 |
| 2019 | A CNN-based Approach for three-class classification of motor imagery EEG data including 'rest state' in hybrid multi-user BCIabstractMulti-user BCI (brain computer interface) refers to a kind of BCI in which there are several users participating in one task. Differently from P300 and SSVEP (steady-state visual evoked potential), MI (motor imagery)-BCI does not rely on external stimulus, which is more widely used in the assistance of disabled people. The main problem of MI-BCI is to achieve asynchronous control, which needs to improve the rest state recognition accuracy. We proposed a CNN-based approach for three-class classification of motor imagery EEG data including rest state in hybrid multi-user BCI. Firstly, we designed a two-user hybrid MI-BCI experimental paradigm and moreover proposed a CNN-based processing framework method which contains several strategies for inter-brain phase locked value (PLV) features fusion. Results show that the alpha band shows the significantly better performance than other bands and the classification performance of multi-user MI is better than single-user MI. Multi-user BCI is a potential way to enhance the performance of asynchronous MI-BCI. Chongwei Su, Dariusz Zapala, Li Zhu 0005, Gaochao Cui, Wanzeng Kong |
BIBM | 6 |
| 2019 | Gender recognition in emotion perception using EEG featuresabstractGender recognition is widely studied in different areas and especially, many physiological theories have shown that there are many differences in emotional processing between different genders. However, there are many observations need to be verified. In our paper, we focus on the gender recognition in emotion perception using diverse EEG (electroencephalogram) features. The time-frequency and phase locked value (PLV) features are extracted and different feature selection methods are compared. We performed these methods on DEAP dataset. Results show that 1) the gender recognition accuracy using PLV feature is higher than using traditional time-frequency feature. 2) extreme learning machine (ELM) classifier has the best performance. 3) the gender recognition accuracy of theta and gamma bands is higher than other bands, which indicates that theta and gamma bands contain more discriminate gender information. 4) arousal has a greater impact than valence. And recognition accuracy of `calm' emotion is lower than other emotions 5) feature selection methods can improve the accuracy. Li Zhu 0005, Wanzeng Kong |
BIBM | 4 |
| 2019 | Idle-State Detection in Multi-user Motor Imagery Brain Computer Interface with Cross-Brain CSP and Hyper-Brain-NetworkabstractMotor imagery (MI) is a kind of spontaneous controlled brain computer interface (BCI) paradigm, which is more likely to the concept of 'mind control'. The idle state detection is an important problem to construct a robust MI-BCI system since it needs to tell whether the subject is in MI task and the idle state contains much diverse cases. Herein, EEG-based multi-user BCI refers to two or more subjects engage in a coordinate task while their EEG are simultaneously recorded. The objective of this paper is to explore how the multi-user MI-BCI performance in idle detection based on CSP (common spatial pattern) and brain-network features. We proposed several strategies for cross-brain feature fusion. Results show that 1) Through CSP features, the classification accuracy of cross-brain outperforms the single brain CSP feature across different strategies. 2) Through brain-network features, the classification accuracy of concatenated with the paired subjects outperforms the single brain-network, while the inter-brain-network is lower than single subject 3) alpha frequency band shows better performance than other bands. Multi-user MI-BCI would be a potential way to improve the idle state detection accuracy. Li Zhu 0005, Chongwei Su, Gaochao Cui, Changle Zhou, Wanzeng Kong |
CW | 6 |
| 2019 | Flexible Non-negative Matrix Factorization with Adaptively Learned Graph RegularizationabstractNon-negative matrix factorization (NMF) is an efficient model in learning parts-based data representation. Since the local geometrical structure can be effectively modeled by a nearest neighbor graph, the graph regularized NMF (GNMF) was proposed to make the learned representation more faithfully and better characterize the intrinsic structure of data. However, GNMF shares a similar paradigm with most of existing graph-based learning models which perform learning tasks on a fixed input graph. In this paper, we propose a new Flexible NMF model with adaptively learned Graph regularization (FNMFG) in which the graph is jointly learned with simultaneous performing the matrix factorization. An efficient iterative method with guaranteed convergence and relative low complexity is developed to optimize the FNMFG objective. Experiments compare FNMFG method with state-of-the-art algorithms and demonstrate its improved performance. Yong Peng 0001, Yanfang Long, Fei-wei Qin, Wanzeng Kong, Feiping Nie 0001, Andrzej Cichocki |
ICASSP | 4 |
| 2019 | Joint Structured Graph Learning and Clustering Based on Concept FactorizationabstractAs one of the matrix factorization models, concept factorization (CF) achieved promising performance in learning data representation in both original feature space and reproducible kernel Hilbert space (RKHS). Based on the consensuses that 1) learning performance of models can be enhanced by exploiting the geometrical structure of data and 2) jointly performing structured graph learning and clustering can avoid the suboptimal solutions caused by the two-stage strategy in graph-based learning, we developed a new CF model with self-expression. Our model has a combined coefficient matrix which is able to learn more efficiently. In other words, we propose a CF-based joint structured graph learning and clustering model (JSGCF). A new efficient iterative method is developed to optimize the JSGCF objective function. Experimental results on representative data sets demonstrate the effectiveness of our new JSGCF algorithm. Yong Peng 0001, Rixin Tang, Wanzeng Kong, Feiping Nie 0001, Andrzej Cichocki |
ICASSP | 3 |
| 2019 | Joint Structured Graph Learning and Unsupervised Feature SelectionabstractThe central task in graph-based unsupervised feature selection (GUFS) depends on two folds, one is to accurately characterize the geometrical structure of the original feature space with a graph and the other is to make the selected features well preserve such intrinsic structure. Currently, most of the existing GUFS methods use a two-stage strategy which constructs graph first and then perform feature selection on this fixed graph. Since the performance of feature selection severely depends on the quality of graph, the selection results will be unsatisfactory if the given graph is of low-quality. To this end, we propose a joint graph learning and unsupervised feature selection (JGUFS) model in which the graph can be adjusted to adapt the feature selection process. The JGUFS objective function is optimized by an efficient iterative algorithm whose convergence and complexity are analyzed in detail. Experimental results on representative benchmark data sets demonstrate the improved performance of JGUFS in comparison with state-of-the-art methods and therefore we conclude that it is promising of allowing the feature selection process to change the data graph. Yong Peng 0001, Leijie Zhang, Wanzeng Kong, Feiping Nie 0001, Andrzej Cichocki |
ICASSP | 3 |
| 2019 | Deep Multimodal Multilinear Fusion with High-order Polynomial PoolingabstractTensor-based multimodal fusion techniques have exhibited great predictive performance. However, one limitation is that existing approaches only consider bilinear or trilinear pooling, which fails to unleash the complete expressive power of multilinear fusion with restricted orders of interactions. More importantly, simply fusing features all at once ignores the complex local intercorrelations, leading to the deterioration of prediction. In this work, we first propose a polynomial tensor pooling (PTP) block for integrating multimodal features by considering high-order moments, followed by a tensorized fully connected layer. Treating PTP as a building block, we further establish a hierarchical polynomial fusion network (HPFN) to recursively transmit local correlations into global ones. By stacking multiple PTPs, the expressivity capacity of HPFN enjoys an exponential growth w.r.t. the number of layers, which is shown by the equivalence to a very deep convolutional arithmetic circuits. Various experiments demonstrate that it can achieve the state-of-the-art performance. Jiajia Tang, Wanzeng Kong, Qibin Zhao |
NeurIPS | 4 |
| 2019 | Large-scale trip planning for bike-sharing systems
Zhi Li 0052, Jiayu Gan, Pengqian Lu, Wanzeng Kong |
Pervasive Mob. Comput. | 6 |
| 2018 | Task-Independent EEG Identification via Low-Rank Matrix Decomposition
Xianghao Kong, Wanzeng Kong, Qiaonan Fan, Qibin Zhao, Andrzej Cichocki |
BIBM | 2 |
| 2018 | A New Method for Brain Death Diagnosis Based on Phase Synchronization Analysis With EEG
Jianting Cao, Wanzeng Kong, Jiajia Tang, Yong Peng 0001 |
BIBM | 4 |
| 2018 | Emotional-state brain network analysis revealed by minimum spanning tree using EEG signals
Shaokai Zhao, Jiajia Tang, Tao Zhang 0062, Yong Peng 0001, Wanzeng Kong |
BIBM | 7 |
| 2018 | Parallel Vector Field Regularized Non-Negative Matrix Factorization for Image RepresentationabstractNon-negative Matrix Factorization (NMF) is a popular model in machine learning, which can learn parts-based representation by seeking for two non-negative matrices whose product can best approximate the original matrix. However, the manifold structure is not considered by NMF and many of the existing work use the graph Laplacian to ensure the smoothness of the learned representation coefficients on the data manifold. Further, beyond smoothness, it is suggested by recent theoretical work that we should ensure second order smoothness for the NMF mapping, which measures the linearity of the NMF mapping along the data manifold. Based on the equivalence between the gradient field of a linear function and a parallel vector field, we propose to find the NMF mapping which minimizes the approximation error, and simultaneously requires its gradient field to be as parallel as possible. The continuous objective function on the manifold can be discretized and optimized under the general NMF framework. Extensive experimental results suggest that the proposed parallel field regularized NMF provides a better data representation and achieves higher accuracy in image clustering. Yong Peng 0001, Rixin Tang, Wanzeng Kong, Fei-wei Qin, Feiping Nie 0001 |
ICASSP | 3 |
| 2017 | Task-Free Brainprint Recognition Based on Degree of Brain Networks
Wanzeng Kong, Qiaonan Fan, Luyun Wang, Bei Jiang, Yong Peng 0001 |
ICONIP (2) | 1 |
| 2017 | Assessment of driving fatigue based on intra/inter-region phase synchronization
Wanzeng Kong, Zhanpeng Zhou, Bei Jiang, Fabio Babiloni, Gianluca Borghini |
Neurocomputing | 1 |
| 2017 | Orthogonal extreme learning machine for image classification
Yong Peng 0001, Wanzeng Kong |
Neurocomputing | 2 |
| 2016 | Shortcomings/Limitations of Blockwise Granger Causality and Advances of Blockwise New CausalityabstractMultivariate blockwise Granger causality (BGC) is used to reflect causal interactions among blocks of multivariate time series. In particular, spectral BGC and conditional spectral BGC are used to disclose blockwise causal flow among different brain areas in various frequencies. In this paper, we demonstrate that: 1) BGC in time domain may not necessarily disclose true causality and 2) due to the use of the transfer function or its inverse matrix and partial information of the multivariate linear regression model, both of spectral BGC and conditional spectral BGC have shortcomings and/or limitations, which may inevitably lead to misinterpretation. We then, in time and frequency domains, develop two new multivariate blockwise causality methods for the linear regression model called blockwise new causality (BNC) and spectral BNC, respectively. By several examples, we confirm that BNC measures are more reasonable and sensitive to reflect true causality or trend of true causality than BGC or conditional BGC. Finally, for electroencephalograph data from an epilepsy patient, we analyze event-related potential causality and demonstrate that both of the BGC and BNC methods show significant causality flow in frequency domain, but the spectral BNC method yields satisfactory and convincing results, which are consistent with an event-related time-frequency power spectrum activity. The spectral BGC method is shown to generate misleading results. Thus, we deeply believe that our new blockwise causality definitions as well as our previous NC definitions may have wide applications to reflect true causality among two blocks of time series or two univariate time series in economics, neuroscience, and engineering. Sanqing Hu, Xinxin Jia, Wanzeng Kong, Yu Cao 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2016 | Comparison Analysis: Granger Causality and New Causality and Their Applications to Motor ImageryabstractIn this paper we first point out a fatal drawback that the widely used Granger causality (GC) needs to estimate the autoregressive model, which is equivalent to taking a series of backward recursive operations which are infeasible in many irreversible chemical reaction models. Thus, new causality (NC) proposed by Hu et al. (2011) is theoretically shown to be more sensitive to reveal true causality than GC. We then apply GC and NC to motor imagery (MI) which is an important mental process in cognitive neuroscience and psychology and has received growing attention for a long time. We study causality flow during MI using scalp electroencephalograms from nine subjects in Brain-computer interface competition IV held in 2008. We are interested in three regions: Cz (central area of the cerebral cortex), C3 (left area of the cerebral cortex), and C4 (right area of the cerebral cortex) which are considered to be optimal locations for recognizing MI states in the literature. Our results show that: 1) there is strong directional connectivity from Cz to C3/C4 during left- and right-hand MIs based on GC and NC; 2) during left-hand MI, there is directional connectivity from C4 to C3 based on GC and NC; 3) during right-hand MI, there is strong directional connectivity from C3 to C4 which is much clearly revealed by NC than by GC, i.e., NC largely improves the classification rate; and 4) NC is demonstrated to be much more sensitive to reveal causal influence between different brain regions than GC. Sanqing Hu, Wanzeng Kong, Yu Cao 0002, Robert Kozma 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2014 | Causality from Cz to C3/C4 or between C3 and C4 revealed by granger causality and new causality during motor imageryabstractInteraction between different brain regions has received wide attention recently. Granger causality (GC) is one of the most popular methods to explore causality relationship between different brain regions. In 2011, Hu et. al [1] pointed out shortcomings and/or limitations of GC by using a large of number of illustrative examples and showed that GC is only a causality definition in the sense of Granger and does not reflect real causality at all, and meanwhile proposed a new causality (NC) which is shown to be more reasonable and understandable than GC by those examples. Motor imagery (MI) is an important mental process in cognitive neuroscience and cognitive psychology and has received growing attention for a long time. However, there is few work about causality flow so far during MI based on scalp EEG. In this paper, we use scalp EEG to study causality flow during MI. The scalp EEGs are from 9 subjects in BCI competition IV held in 2008 [2] and provided by Graz University of Technology. We are interested in three regions: Cz (the centre of cerebral cortex), C3 (the left of cerebral cortex) and C4 (the right of cerebral cortex) which are considered to be optimal locations for recognizing MI states in literature. We apply GC and NC to scalp EEG and find that i) there is strong directional connectivity from Cz to C3/C4 during left hand and right hand MI based on GC and NC. ii) During left hand MI, there is directional connectivity from C4 to C3 based on GC and NC. iii) During right hand MI, there is strong directional connectivity from C3 to C4 which is much clearly revealed by NC method than by GC method. iv) Our results suggest that NC method in time and frequency domains is demonstrated to be much better to reveal causal influence between different brain regions than GC method. Thus, we deeply believe that NC method will shed new light on causality analysis in economics and neuroscience. Sanqing Hu, Wanzeng Kong, Yu Cao 0002 |
IJCNN | 4 |
| 2013 | Motor imagery classification based on joint regression model and spectral power
Sanqing Hu, Qiangqiang Tian, Yu Cao 0002, Wanzeng Kong |
Neural Comput. Appl. | 5 |
| 2013 | Robust and smart spectral clustering from normalized cut
Wanzeng Kong, Sanqing Hu, Guojun Dai |
Neural Comput. Appl. | 1 |