EDBT 2026 Demo / reviewers in the wild / expert
Huiguang He
dblp:84/5905
· DBLP profile ↗
67ranked-venue papers
0as first author
44since 2021 · last 2026
0000-0002-0684-1711ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 7 since 2021Human-computer interaction and ubiquitous computing · 2Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Disentangled multimodal domain generalization network for zero-calibration vigilance estimation
Kangning Wang 0005, Wei Wei 0046, Weibo Yi, Huiguang He, Minpeng Xu, Shuang Qiu 0002, Dong Ming |
Knowl. Based Syst. | 5 |
| 2026 | Quantization and disentanglement for cross-modal alignment in neural speech reconstruction from brain activityabstractUnderstanding and reconstructing speech from brain activity is central to advancing neuroscience and brain–computer interface (BCI) research. Despite recent progress, current brain-to-speech approaches remain limited by entangled neural representations, overly fine-grained decoding targets, and poor generalization. To address these challenges, we propose a novel framework that integrates quantization with representation disentanglement. Instead of learning brain-specific discrete units, we directly align magnetoencephalography (MEG) signals with the codebook of a pretrained neural audio codec. To balance representational capacity and decoding complexity, we introduce a discrete contrastive loss combining codebook clustering and temporal pooling. Furthermore, we factorize speech into prosody, content, and timbre, and decode these subspaces separately, simplifying learning and enhancing decoding precision. Experiments on two public MEG datasets demonstrate strong performance in both intra- and cross-subject settings, yielding substantial improvements in prosody-related and perceptual metrics. Our findings highlight a principled pathway towards more reliable brain-to-speech decoding, with potential applications in assistive communication and BCI technologies. Changde Du, Huiguang He |
Pattern Recognit. | 4 |
| 2025 | BP-GPT: Auditory Neural Decoding Using fMRI-prompted LLMabstractDecoding language information from brain signals represents a vital research area within brain-computer interfaces, particularly in the context of deciphering the semantic information from the fMRI signal. Although existing work uses LLM to achieve this goal, their method does not use an end-to-end approach and avoids the LLM in the mapping of fMRI-to-text, leaving space for the exploration of the LLM in auditory decoding. In this paper, we introduce a novel method, the Brain Prompt GPT (BP-GPT). By using the brain representation that is extracted from the fMRI as a prompt, our method can utilize GPT-2 to decode fMRI signals into stimulus text. Further, we introduce the text prompt and align the fMRI prompt to it. By introducing the text prompt, our BP-GPT can extract a more robust brain prompt and promote the decoding of pre-trained LLM. We evaluate our BP-GPT on the open-source auditory semantic decoding dataset and achieve a significant improvement up to 4.61% on METEOR and 2.43% on BERTScore across all the subjects compared to the state-of-the-art method. The experimental results demonstrate that using brain representation as a prompt to further drive LLM for auditory neural decoding is feasible and effective. The code is available at https://github.com/1994cxy/BP-GPT. Changde Du, Huiguang He |
ICASSP | 5 |
| 2025 | ThicknessVAE: Learning a Lateral Prior for Clothed Human Body ReconstructionabstractSandwich-like structures have shown remarkable efficacy in clothed human reconstruction. However, these approaches often generate unrealistic side geometries due to inadequate handling of lateral regions. This paper addresses this limitation by incorporating the side geometry of clothed humans as a prior. We propose ThicknessVAE, a novel two-stage method that makes two key contributions: (1) We learn a prototype from point clouds for the lateral regions of clothed humans to extract common and detailed geometric features. (2) We utilize this prototype as a prior to transform geometric features into a thickness map associated with clothed human images, enabling refined normal integration for sandwich-like reconstruction methods. By seamlessly integrating our model into the sandwich-like reconstruction pipeline, we achieve highly realistic side views. Both qualitative and quantitative experiments demonstrate that our approach is comparable to state-of-the-art methods in terms of side-view realism. Xiaotao Wu, Zhaoxin Fan, Huiguang He, Dinggang Shen |
ICASSP | 3 |
| 2025 | Lesion Localization Prior-Driven Few-Shot Learning for Branch Atheromatous Disease Diagnosis
Kaijun Zhang, Shengde Li, Shengpei Wang, Shangyi Shi, Huiguang He |
ICIG (1) | 8 |
| 2025 | Animate Your Thoughts: Reconstruction of Dynamic Natural Vision from Human Brain ActivityabstractReconstructing human dynamic vision from brain activity is a challenging task with great scientific significance. Although prior video reconstruction methods have made substantial progress, they still suffer from several limitations, including: (1) difficulty in simultaneously reconciling semantic (e.g. categorical descriptions), structure (e.g. size and color), and consistent motion information (e.g. order of frames); (2) low temporal resolution of fMRI, which poses a challenge in decoding multiple frames of video dynamics from a single fMRI frame; (3) reliance on video generation models, which introduces ambiguity regarding whether the dynamics observed in the reconstructed videos are genuinely derived from fMRI data or are hallucinations from generative model. To overcome these limitations, we propose a two-stage model named Mind-Animator. During the fMRI-to-feature stage, we decouple semantic, structure, and motion features from fMRI. Specifically, we employ fMRI-vision-language tri-modal contrastive learning to decode semantic feature from fMRI and design a sparse causal attention mechanism for decoding multi-frame video motion features through a next-frame-prediction task. In the feature-to-video stage, these features are integrated into videos using an inflated Stable Diffusion, effectively eliminating external video data interference. Extensive experiments on multiple video-fMRI datasets demonstrate that our model achieves state-of-the-art performance. Comprehensive visualization analyses further elucidate the interpretability of our model from a neurobiological perspective. Project page: https://mind-animator-design.github.io/. Yizhuo Lu, Changde Du, Xuanliu Zhu, Liuyun Jiang, Xujin Li, Huiguang He |
ICLR | 7 |
| 2025 | EmoGrowth: Incremental Multi-label Emotion Decoding with Augmented Emotional Relation GraphabstractEmotion recognition systems face significant challenges in real-world applications, where novel emotion categories continually emerge and multiple emotions often co-occur. This paper introduces multi-label fine-grained class incremental emotion decoding, which aims to develop models capable of incrementally learning new emotion categories while maintaining the ability to recognize multiple concurrent emotions. We propose an Augmented Emotional Semantics Learning (AESL) framework to address two critical challenges: past- and future-missing partial label problems. AESL incorporates an augmented Emotional Relation Graph (ERG) for reliable soft label generation and affective dimension-based knowledge distillation for future-aware feature learning. We evaluate our approach on three datasets spanning brain activity and multimedia domains, demonstrating its effectiveness in decoding up to 28 fine-grained emotion categories. Results show that AESL significantly outperforms existing methods while effectively mitigating catastrophic forgetting. Our code is available at https://github.com/ChangdeDu/EmoGrowth. Kaicheng Fu, Changde Du, Shuangchen Zhao, Huiguang He |
ICML | 7 |
| 2025 | A temporal-spectral fusion transformer with subject-specific adapter for enhancing RSVP-BCI decoding
Xujin Li, Wei Wei 0046, Shuang Qiu 0002, Huiguang He |
Neural Networks | 4 |
| 2025 | Adaptive Domain Alignment Neural Networks for Cross-Domain EEG Emotion RecognitionabstractElectroencephalography (EEG) - based Emotion recognition is now facing great challenge of the intra- and inter-subject variability of EEG signal. Researchers attempted to handle this challenge by using transfer learning methods which usually share two main limitations: most of these methods align marginal distributions instead of conditional distributions of source and target data, making the alignment process classwise ambiguous; also, they prefer to use Multi-Layer Perceptron (MLP) with redundant parameters as classifiers, which is shown by recent research that could result serious over-fitting towards labeled data and prevent the model to draw a proper representation space. In our work, we propose a novel domain alignment method: Adaptive Domain Alignment Neural Networks (ADANN). Our method directly model conditional distributions of source and target domains by two sets of label-wise prototypes, representing the density maximum of each class, while the normalized correspond similarity naturally represents the conditional probability. The predicted label for a sample is given by the argument maxima of similarities and therefore the MLP classifier is not required. Using context-instance contrastive learning to align two sets of prototypes, their corresponding conditional distributions are being learned simultaneously. Exhaustive cross-domain experiments have been conducted under protocols that are strongly related to practical application scenarios and our proposed method achieves better or similar performance compared with recent state-of-the-art methods. Xuezhu Hong, Changde Du, Huiguang He |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | Recognizing Natural Images From EEG With Language-Guided Contrastive LearningabstractElectroencephalography (EEG), known for its convenient noninvasive acquisition but moderate signal-to-noise ratio, has recently gained much attention due to the potential to decode image information. However, previous works have not delivered sufficient evidence of this task, primarily limited by performance and biological plausibility. In this work, we first introduce a self-supervised framework to demonstrate the feasibility of recognizing images from EEG signals. Contrastive learning is leveraged to align the representations of EEG responses with image stimuli. Then, language descriptions of the stimuli generated by large language models (LLMs) help guide learning core semantic information. With the framework, we attain significantly above-chance results on the THINGS-EEG2 dataset, achieving a top-1 accuracy of 19.7% and a top-5 accuracy of 51.5% in challenging 200-way zero-shot tasks. Furthermore, we conduct thorough experiments to resolve the human visual responses with EEG from temporal, spatial, spectral, and semantic perspectives. These results provide evidence of feasibility and plausibility regarding EEG-based image recognition, substantiated by comparative studies with the THINGS-Magnetoencephalography (MEG) dataset. The findings offer valuable insights for neural decoding and real-world applications of brain-computer interfaces (BCIs), such as health care and robot control. The code is available at https://github.com/eeyhsong/NICE-LLM. Yonghao Song, Yijun Wang 0001, Huiguang He, Xiaorong Gao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Enhancing SSVEP-Based BCI Performance via Consensus Information Transfer Among SubjectsabstractThe brain-computer interface (BCI) based on steady-state visual evoked potential (SSVEP) has received considerable attention for its high communication speed. While large datasets provide an important opportunity to enhance decoding accuracies, the key challenge lies in the exploration of existing data to extract valuable information based on the distinctive characteristics of brain responses. In this study, we introduce ConsenNet, a framework designed to enhance SSVEP classification performance by leveraging information from the diverse perspectives of existing subjects. First, this study exploits the diversity of existing subjects to generate new samples, which retain both task-related components and variability. This effectively enhances the network generalization capability on new subjects. Second, the structured knowledge that encapsulates the interrelationships between categories has been constructed and then transferred from the teacher network to the student network, guiding the student network to extract invariant features across subjects. Finally, our model incorporates a small amount of new subject data for model calibration in the final stage. Offline experiments conducted on three public datasets demonstrate the superiority of ConsenNet over 19 methods compared in this study, while online experiments validate its feasibility for real-world applications. Wei Wei 0046, Shuang Qiu 0002, Xujin Li, Yijun Wang 0001, Huiguang He |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | CLIP-MUSED: CLIP-Guided Multi-Subject Visual Neural Information Semantic DecodingabstractThe study of decoding visual neural information faces challenges in generalizing single-subject decoding models to multiple subjects, due to individual differences. Moreover, the limited availability of data from a single subject has a constraining impact on model performance. Although prior multi-subject decoding methods have made significant progress, they still suffer from several limitations, including difficulty in extracting global neural response features, linear scaling of model parameters with the number of subjects, and inadequate characterization of the relationship between neural responses of different subjects to various stimuli.
To overcome these limitations, we propose a CLIP-guided Multi-sUbject visual neural information SEmantic Decoding (CLIP-MUSED) method. Our method consists of a Transformer-based feature extractor to effectively model global neural representations. It also incorporates learnable subject-specific tokens that facilitates the aggregation of multi-subject data without a linear increase of parameters. Additionally, we employ representational similarity analysis (RSA) to guide token representation learning based on the topological relationship of visual stimuli in the representation space of CLIP, enabling full characterization of the relationship between neural responses of different subjects under different stimuli. Finally, token representations are used for multi-subject semantic decoding. Our proposed method outperforms single-subject decoding methods and achieves state-of-the-art performance among the existing multi-subject methods on two fMRI datasets. Visualization results provide insights into the effectiveness of our proposed method. Code is available at https://github.com/CLIP-MUSED/CLIP-MUSED. Qiongyi Zhou, Changde Du, Shengpei Wang, Huiguang He |
ICLR | 4 |
| 2024 | Contrastive fine-grained domain adaptation network for EEG-based vigilance estimation
Kangning Wang 0005, Wei Wei 0046, Weibo Yi, Shuang Qiu 0002, Huiguang He, Minpeng Xu, Dong Ming |
Neural Networks | 5 |
| 2024 | Growing Like a Tree: Finding Trunks From Graph Skeleton TreesabstractThe message-passing paradigm has served as the foundation of graph neural networks (GNNs) for years, making them achieve great success in a wide range of applications. Despite its elegance, this paradigm presents several unexpected challenges for graph-level tasks, such as the long-range problem, information bottleneck, over-squashing phenomenon, and limited expressivity. In this study, we aim to overcome these major challenges and break the conventional "node- and edge-centric" mindset in graph-level tasks. To this end, we provide an in-depth theoretical analysis of the causes of the information bottleneck from the perspective of information influence. Building on the theoretical results, we offer unique insights to break this bottleneck and suggest extracting a skeleton tree from the original graph, followed by propagating information in a distinctive manner on this tree. Drawing inspiration from natural trees, we further propose to find trunks from graph skeleton trees to create powerful graph representations and develop the corresponding framework for graph-level tasks. Extensive experiments on multiple real-world datasets demonstrate the superiority of our model. Comprehensive experimental analyses further highlight its capability of capturing long-range dependencies and alleviating the over-squashing problem, thereby providing novel insights into graph-level tasks. Zhongyu Huang, Yingheng Wang, Chaozhuo Li, Huiguang He |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Improved Video Emotion Recognition With Alignment of CNN and Human Brain RepresentationsabstractThe ability to perceive emotions is an important criterion for judging whether a machine is intelligent. To this end, a large number of emotion recognition algorithms have been developed especially for visual information such as video. Most previous studies are based on hand-crafted features or CNN, in which the former fails to extract expressive features and the latter still faces the undesired affective gap. This drives us to think about what if we could incorporate the human emotional perception capability into CNN. In this paper, we attempt to address this question by exploring alignment between representations of neural networks and human brain activity. In particular, we employ a visually evoked emotional brain activity dataset to conduct a jointly training strategy for CNN. In the training phase, we introduce the representation similarity analysis (RSA) to align the CNN with human brain to obtain more brain-like features. Specifically, representation similarity matrices (RSMs) of multiple convolutional layers are averaged with learnable weights and related to the RSM of human brain. In order to obtain emotion-related brain activity, we conduct voxel selection and denoising with a banded ridge model before computing the RSM. Sufficient experiments on two challenging video emotion recognition datasets and multiple popular CNN architectures suggest that human brain activity is promising to provide an inductive bias for CNN towards better performance of emotion recognition. Our source code is available inhttps://osf.io/ucx57. Kaicheng Fu, Changde Du, Shengpei Wang, Huiguang He |
IEEE Trans. Affect. Comput. | 4 |
| 2024 | PGCN: Pyramidal Graph Convolutional Network for EEG Emotion RecognitionabstractEmotion recognition is essential in the diagnosis and rehabilitation of various mental diseases. In the last decade, electroencephalogram (EEG)-based emotion recognition has been intensively investigated due to its prominative accuracy and reliability, and graph convolutional network (GCN) has become a mainstream model to decode emotions from EEG signals. However, the electrode relationship, especially long-range electrode dependencies across the scalp, may be underutilized by GCNs, although such relationships have been proven to be important in emotion recognition. The small receptive field makes shallow GCNs only aggregate local nodes. On the other hand, stacking too many layers leads to over-smoothing. To solve these problems, we propose the pyramidal graph convolutional network (PGCN), which aggregates features at three levels: local, mesoscopic, and global. First, we construct a vanilla GCN based on the 3D topological relationships of electrodes, which is used to integrate two-order local features; Second, we construct several mesoscopic brain regions based on priori knowledge and employ mesoscopic attention to sequentially calculate the virtual mesoscopic centers to focus on the functional connections of mesoscopic brain regions; Finally, we fuse the node features and their 3D positions to construct a numerical relationship adjacency matrix to integrate structural and functional connections from the global perspective. Experimental results on four public datasets indicate that PGCN enhances the relationship modelling across the scalp and achieves stateof-the-art performance in both subject-dependent and subjectindependent scenarios. Meanwhile, PGCN makes an effective trade-off between enhancing network depth and receptive fields while suppressing the ensuing over-smoothing. Our codes are publicly accessible athttps://github.com/Jinminbox/PGCN. Ming Jin 0006, Changde Du, Huiguang He, Ting Cai 0001, Jinpeng Li 0002 |
IEEE Trans. Multim. | 3 |
| 2024 | Multi-View Multi-Label Fine-Grained Emotion Decoding From Human Brain ActivityabstractDecoding emotional states from human brain activity play an important role in the brain-computer interfaces. Existing emotion decoding methods still have two main limitations: one is only decoding a single emotion category from a brain activity pattern and the decoded emotion categories are coarse-grained, which is inconsistent with the complex emotional expression of humans; the other is ignoring the discrepancy of emotion expression between the left and right hemispheres of the human brain. In this article, we propose a novel multi-view multi-label hybrid model for fine-grained emotion decoding (up to 80 emotion categories) which can learn the expressive neural representations and predict multiple emotional states simultaneously. Specifically, the generative component of our hybrid model is parameterized by a multi-view variational autoencoder, in which we regard the brain activity of left and right hemispheres and their difference as three distinct views and use the product of expert mechanism in its inference network. The discriminative component of our hybrid model is implemented by a multi-label classification network with an asymmetric focal loss. For more accurate emotion decoding, we first adopt a label-aware module for emotion-specific neural representation learning and then model the dependency of emotional states by a masked self-attention mechanism. Extensive experiments on two visually evoked emotional datasets show the superiority of our method. Kaicheng Fu, Changde Du, Shengpei Wang, Huiguang He |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | RoBrain: Towards Robust Brain-to-Image Reconstruction via Cross-Domain Contrastive Learning
Changde Du, Huiguang He |
ICONIP (3) | 3 |
| 2023 | A Fine-Grained Domain Adaptation Method for Cross-Session Vigilance Estimation in SSVEP-Based BCI
Kangning Wang 0005, Shuang Qiu 0002, Wei Wei 0046, Huiguang He, Minpeng Xu, Dong Ming |
ICONIP (3) | 5 |
| 2023 | Mammo-Net: Integrating Gaze Supervision and Interactive Information in Multi-view Mammogram Classification
Changkai Ji, Changde Du, Sheng Wang 0014, Chong Ma 0004, Jiaming Xie, Huiguang He, Dinggang Shen |
MICCAI (7) | 8 |
| 2023 | Auditory Attention Decoding with Task-Related Multi-View Contrastive LearningabstractThe human brain can easily focus on one speaker and suppress others in scenarios such as a cocktail party. Recently, researchers found that auditory attention can be decoded from the electroencephalogram (EEG) data. However, most existing deep learning methods are difficult to use prior knowledge of different views (that is attended speech and EEG are task-related views) and extract an unsatisfactory representation. Inspired by Broadbent's filter model, we decode auditory attention in a multi-view paradigm and extract the most relevant and important information utilizing the missing view. Specifically, we propose an auditory attention decoding (AAD) method based on multi-view VAE with task-related multi-view contrastive (TMC) learning. Employing TMC learning in multi-view VAE can utilize the missing view to accumulate prior knowledge of different views into the fusion of representation, and extract the approximate task-related representation. We examine our method on two popular AAD datasets, and demonstrate the superiority of our method by comparing it to the state-of-the-art method. Changde Du, Qiongyi Zhou, Huiguang He |
ACM Multimedia | 4 |
| 2023 | MindDiffuser: Controlled Image Reconstruction from Human Brain Activity with Semantic and Structural DiffusionabstractReconstructing visual stimuli from brain recordings has been a meaningful and challenging task. Especially, the achievement of precise and controllable image reconstruction bears great significance in propelling the progress and utilization of brain-computer interfaces. Despite the advancements in complex image reconstruction techniques, the challenge persists in achieving a cohesive alignment of both semantic (concepts and objects) and structure (position, orientation, and size) with the image stimuli. To address the aforementioned issue, we propose a two-stage image reconstruction model called MindDiffuser1. In Stage 1, the VQ-VAE latent representations and the CLIP text embeddings decoded from fMRI are put into Stable Diffusion, which yields a preliminary image that contains semantic information. In Stage 2, we utilize the CLIP visual feature decoded from fMRI as supervisory information, and continually adjust the two feature vectors decoded in Stage 1 through backpropagation to align the structural information. The results of both qualitative and quantitative analyses demonstrate that our model has surpassed the current state-of-the-art models on Natural Scenes Dataset (NSD). The subsequent experimental findings corroborate the neurobiological plausibility of the model, as evidenced by the interpretability of the multimodal feature employed, which align with the corresponding brain responses. Yizhuo Lu, Changde Du, Qiongyi Zhou, Dianpeng Wang, Huiguang He |
ACM Multimedia | 5 |
| 2023 | A multimodal approach to estimating vigilance in SSVEP-based BCI
Kangning Wang 0005, Shuang Qiu 0002, Wei Wei 0046, Shengpei Wang, Huiguang He, Minpeng Xu, Tzyy-Ping Jung, Dong Ming |
Expert Syst. Appl. | 6 |
| 2023 | Cross-modal guiding and reweighting network for multi-modal RSVP-based target detection
Jiayu Mao, Shuang Qiu 0002, Wei Wei 0046, Huiguang He |
Neural Networks | 4 |
| 2023 | Decoding Visual Neural Representations by Multimodal Learning of Brain-Visual-Linguistic FeaturesabstractDecoding human visual neural representations is a challenging task with great scientific significance in revealing vision-processing mechanisms and developing brain-like intelligent machines. Most existing methods are difficult to generalize to novel categories that have no corresponding neural data for training. The two main reasons are 1) the under-exploitation of the multimodal semantic knowledge underlying the neural data and 2) the small number of paired (stimuli-responses) training data. To overcome these limitations, this paper presents a generic neural decoding method called BraVL that uses multimodal learning of brain-visual-linguistic features. We focus on modeling the relationships between brain, visual and linguistic features via multimodal deep generative models. Specifically, we leverage the mixture-of-product-of-experts formulation to infer a latent code that enables a coherent joint generation of all three modalities. To learn a more consistent joint representation and improve the data efficiency in the case of limited brain activity data, we exploit both intra- and inter-modality mutual information maximization regularization terms. In particular, our BraVL model can be trained under various semi-supervised scenarios to incorporate the visual and textual features obtained from the extra categories. Finally, we construct three trimodal matching datasets, and the extensive experiments lead to some interesting conclusions and cognitive insights: 1) decoding novel visual categories from human brain activity is practically possible with good accuracy; 2) decoding models using the combination of visual and linguistic features perform much better than those using either of them alone; 3) visual perception may be accompanied by linguistic influences to represent the semantics of visual stimuli. Changde Du, Kaicheng Fu, Jinpeng Li 0002, Huiguang He |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Graph-Enhanced Emotion Neural DecodingabstractBrain signal-based emotion recognition has recently attracted considerable attention since it has powerful potential to be applied in human-computer interaction. To realize the emotional interaction of intelligent systems with humans, researchers have made efforts to decode human emotions from brain imaging data. The majority of current efforts use emotion similarities (e.g., emotion graphs) or brain region similarities (e.g., brain networks) to learn emotion and brain representations. However, the relationships between emotions and brain regions are not explicitly incorporated into the representation learning process. As a result, the learned representations may not be informative enough to benefit specific tasks, e.g., emotion decoding. In this work, we propose a novel idea of graph-enhanced emotion neural decoding, which takes advantage of a bipartite graph structure to integrate the relationships between emotions and brain regions into the neural decoding process, thus helping learn better representations. Theoretical analyses conclude that the suggested emotion-brain bipartite graph inherits and generalizes the conventional emotion graphs and brain networks. Comprehensive experiments on visually evoked emotion datasets demonstrate the effectiveness and superiority of our approach. Zhongyu Huang, Changde Du, Yingheng Wang, Kaicheng Fu, Huiguang He |
IEEE Trans. Medical Imaging | 5 |
| 2022 | Going Deeper into Permutation-Sensitive Graph Neural NetworksabstractThe invariance to permutations of the adjacency matrix, i.e., graph isomorphism, is an overarching requirement for Graph Neural Networks (GNNs). Conventionally, this prerequisite can be satisfied by the invariant operations over node permutations when aggregating messages. However, such an invariant manner may ignore the relationships among neighboring nodes, thereby hindering the expressivity of GNNs. In this work, we devise an efficient permutation-sensitive aggregation mechanism via permutation groups, capturing pairwise correlations between neighboring nodes. We prove that our approach is strictly more powerful than the 2-dimensional Weisfeiler-Lehman (2-WL) graph isomorphism test and not less powerful than the 3-WL test. Moreover, we prove that our approach achieves the linear sampling complexity. Comprehensive experiments on multiple synthetic and real-world datasets demonstrate the superiority of our model. Zhongyu Huang, Yingheng Wang, Chaozhuo Li, Huiguang He |
ICML | 4 |
| 2022 | Graph Emotion Decoding from Visually Evoked Neural Responses
Zhongyu Huang, Changde Du, Yingheng Wang, Huiguang He |
MICCAI (8) | 4 |
| 2022 | VigilanceNet: Decouple Intra- and Inter-Modality Learning for Multimodal Vigilance Estimation in RSVP-Based BCIabstractRecently, brain-computer interface (BCI) technology has made impressive progress and has been developed for many applications. Thereinto, the BCI system based on rapid serial visual presentation (RSVP) is a promising information detection technology. However, the use of RSVP is closely related to the user's performance, which can be influenced by their vigilance levels. Therefore it is crucial to detect vigilance levels in RSVP-based BCI. In this paper, we conducted a long-term RSVP target detection experiment to collect electroencephalography (EEG) and electrooculogram (EOG) data at different vigilance levels. In addition, to estimate vigilance levels in RSVP-based BCI, we propose a multimodal method named VigilanceNet using EEG and EOG. Firstly, we define the multiplicative relationships in conventional EOG features that can better describe the relationships between EOG features, and design an outer product embedding module to extract the multiplicative relationships. Secondly, we propose to decouple the learning of intra- and inter-modality to improve multimodal learning. Specifically, for intra-modality, we introduce an intra-modality representation learning (intra-RL) method to obtain effective representations of each modality by letting each modality independently predict vigilance levels during the multimodal training process. For inter-modality, we employ the cross-modal Transformer based on cross-attention to capture the complementary information between EEG and EOG, which only pays attention to the inter-modality relations. Extensive experiments and ablation studies are conducted on the RSVP and SEED-VIG public datasets. The results demonstrate the effectiveness of the method in terms of regression error and correlation. Wei Wei 0046, Changde Du, Shuang Qiu 0002, Sanli Tian, Huiguang He |
ACM Multimedia | 7 |
| 2022 | TFF-Former: Temporal-Frequency Fusion Transformer for Zero-training Decoding of Two BCI TasksabstractBrain-computer interface (BCI) systems provide a direct connection between the human brain and external devices. Visual evoked BCI systems including Event-related Potential (ERP) and Steady-state Visual Evoked Potential (SSVEP) have attracted extensive attention because of their strong brain responses and wide applications. Previous studies have made some breakthroughs in within-subject decoding algorithms for specific tasks. However, there are two challenges in current decoding algorithms in BCI systems. Firstly, current decoding algorithms cannot accurately classify EEG signals without the data of the new subject, but the calibration procedure is time-consuming. Secondly, algorithms are tailored to extract features for one specific task, which limits their applications across tasks. In this study, we proposed a Temporal-Frequency Fusion Transformer (TFF-Former) for zero-training decoding across two BCI tasks. EEG data were organized into temporal-spatial and frequency-spatial forms, which can be considered as two views. In the TFF-Former framework, two symmetrical Transformer streams were designed to extract view-specific features. The cross-view module based on the cross-attention mechanism was proposed to guide each stream to strengthen common representations of features across EEG views. Additionally, an attention-based fusion module was built to fuse the representations from the two views effectively. The mean mask mechanism was applied to adaptively decrease redundant EEG tokens aggregation for the integration of common representations. We validated our method on the self-collected RSVP dataset and benchmark SSVEP dataset. Experimental results demonstrated that our TFF-Former model achieved competitive performance compared with models in each of the above paradigms. It can further promote the application of visual evoked EEG-based BCI system. Xujin Li, Wei Wei 0046, Shuang Qiu 0002, Huiguang He |
ACM Multimedia | 4 |
| 2022 | A Zero-Training Method for RSVP-Based Brain Computer Interface
Xujin Li, Shuang Qiu 0002, Wei Wei 0046, Huiguang He |
PRCV (2) | 4 |
| 2022 | Dynamic Domain Adaptation for Class-Aware Cross-Subject and Cross-Session EEG Emotion RecognitionabstractIt is vital to develop general models that can be shared across subjects and sessions in the real-world deployment of electroencephalogram (EEG) emotion recognition systems. Many prior studies have exploited domain adaptation algorithms to alleviate the inter-subject and inter-session discrepancies of EEG distributions. However, these methods only aligned the global domain divergence, but overlooked the local domain divergence with respect to each emotion category. This degenerates the emotion-discriminating ability of the domain invariant features. In this paper, we argue that aligning the EEG data within the same emotion categories is important for generalizable and discriminative features. Hence, we propose the dynamic domain adaptation (DDA) algorithm where the global and local divergences are disposed by minimizing the global domain discrepancy and local subdomain discrepancy, respectively. To tackle the absence of emotion labels in the target domain, we introduce a dynamic training strategy where the model focuses on optimizing the global domain discrepancy in the early training steps, and then gradually switches to the local subdomain discrepancy. The DDA algorithm is formally implemented as an unsupervised version and a semi-supervised version for different experimental settings. Based on the coarse-to-fine alignment, our model achieves the average peak accuracy of 91.08%, 92.89% on SEED, and 81.58%, 80.82% on SEED-IV in the cross-subject and cross-session scenarios, respectively. Zhunan Li, Enwei Zhu, Ming Jin 0006, Cunhang Fan, Huiguang He, Ting Cai 0001, Jinpeng Li 0002 |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | Deep Modality Assistance Co-Training Network for Semi-Supervised Multi-Label Semantic DecodingabstractMulti-label semantic decoding is a challenging task with great scientific significance and application value. The existing methods mainly focus on label learning and ignore the amount of information contained in the sample itself, especially non-image sample, which may limit their performance. To address these issues, we propose a novel semi-supervised modality assistance co-training network, which utilizes image modality to assist non-image modality for multi-label learning. In real application, there are two thorny issues: (i) non-image modality tends to be missing owing to the difficulty in obtaining them; (ii) although the image modality is easy to obtain from the Internet, image label annotation is still time-consuming and expensive. Therefore, the proposed method utilizes a small number of paired & labeled images and non-image modalities, and a large number of unpaired & unlabeled images from web sources to improve results. It consists of the modality-specific feature generators, the feature translators and the label relationship network. Specifically, the modality-specific feature generators are used to generate different features (views) for each modality. Semantic translators are employed to capture the relationship between the paired modalities and impute the missing modality feature by using unpaired & unlabeled images. Label relation network is a graph convolution network (GCN) aiming to capture the correlation between labels. To mine the information in unlabeled features, the co-training mechanism is considered. With this mechanism, we introduce a multi-view orthogonality constraint and a multi-label co-regularization constraint. Extensive experiments on three computer vision and neuroscience datasets demonstrate the effectiveness of the proposed method. Changde Du, Haibao Wang, Qiongyi Zhou, Huiguang He |
IEEE Trans. Multim. | 5 |
| 2022 | Structured Neural Decoding With Multitask Transfer Learning of Deep Neural Network RepresentationsabstractThe reconstruction of visual information from human brain activity is a very important research topic in brain decoding. Existing methods ignore the structural information underlying the brain activities and the visual features, which severely limits their performance and interpretability. Here, we propose a hierarchically structured neural decoding framework by using multitask transfer learning of deep neural network (DNN) representations and a matrix-variate Gaussian prior. Our framework consists of two stages, Voxel2Unit and Unit2Pixel. In Voxel2Unit, we decode the functional magnetic resonance imaging (fMRI) data to the intermediate features of a pretrained convolutional neural network (CNN). In Unit2Pixel, we further invert the predicted CNN features back to the visual images. Matrix-variate Gaussian prior allows us to take into account the structures between feature dimensions and between regression tasks, which are useful for improving decoding effectiveness and interpretability. This is in contrast with the existing single-output regression models that usually ignore these structures. We conduct extensive experiments on two real-world fMRI data sets, and the results show that our method can predict CNN features more accurately and reconstruct the perceived natural images and faces with higher quality. Changde Du, Changying Du, Haibao Wang, Huiguang He |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2021 | Enhancing Detection of SSVEPs for High-Speed Brain-Computer Interface with a Siamese ArchitectureabstractBrain-Computer Interface (BCI) is a direct communication medium between brain and outside world. This study focuses on a Steady-State Visually Evoked Potential (SSVEP)based BCI due to its large number of instruction set. However, it is still challenging to decode multi-class SSVEPs. To improve target identification accuracy, we propose a Siamese Correlation Analysis model (SiamCA), which involves two feature extractors with tied parameters and a top decision network. We consider two datasets for benchmarking the performance of the proposed model and compare it with FBCCA, TRCA, ensemble-TRCA and a deep learning method named ConvCA. The proposed method realize a significantly higher average classification accuracy than the compared method at different data length (0.2–1.0s) on two datasets. This suggests that the proposed SiamCA model is a promising methodology for target identification of SSVEPs and could further improve the performance of SSVEP-based BCI system. Shuang Qiu 0002, Minghao Geng, Huiguang He |
BIBM | 4 |
| 2021 | A Cross-Modal Guiding and Fusion Method for Multi-Modal RSVP-based Image RetrievalabstractRapid Serial Visual Presentation (RSVP) is an important paradigm in Brain-Computer Interface (BCI). It can be used in speller, image retrieval, anomaly detection, etc. RSVP paradigm uses a small number of target pictures in a high speed presented picture sequence to induce specific event-related potential (ERP) components. However, the application of RSVP based BCI is challenged by the accuracy of ERP detection. Thus, the goal of this study is to introduce other related modalities to the traditional EEG-based BCI to make robust predictions and improve the detection performance. First, we introduce the eye movement modality into the RSVP-based BCI and collect a multimodality RSVP-based dataset simultaneously during the image retrieval task. Second, we design a simple but efficient CNN-based network with two modality fusion modules to fully utilize the multi-modality data in two stages. In the feature extraction stage, we propose a Cross-modality-Guided Feature Calibration (cm-GFC) module to enable the EEG modality feature to modify the eye movement modality feature, and the aim is to make eye movement modality features and EEG modality features are more complementary. In the feature fusion stage, we propose a Dynamic Gated Fusion (DGF) module, which applies modality-specific gates to retain the complementary information of the two modalities and reduce redundant information from the two modalities. To evaluate our method, we conduct extensive experiments on the dataset with EEG and eye movement data are from 20 subjects. The proposed method achieves a high balanced accuracy of 87.83 ± 2.31% of classification, which outperforms a series of single modality and multi-modality approaches. Jiayu Mao, Shuang Qiu 0002, Wei Wei 0046, Huiguang He |
IJCNN | 5 |
| 2021 | Filter Bank Adversarial Domain Adaptation For Motor Imagery Brain Computer InterfaceabstractMotor imagery (MI) based Brain-computer interface (BCI) is a promising BCI paradigm that can help neuromuscular injury patients to recover or replace their motor abilities. However, electroencephalography (EEG) based MI-BCI suffers from its long calibration time and low classification accuracy, which restrict its application. Thus, it is important to reduce the calibration time of MI-BCI and enhance its prediction accuracy. In this study, we propose a filter bank Wasserstein adversarial domain adaptation framework (FBWADA) that uses a short amount of training data from a new target subject, and all collected data from an existing subject. A Convolutional Neural Networks (CNN) based feature extractor is designed to extract feature from EEG data. Filter bank strategy is employed to extract feature from multiple sub bands and integrate predictions from all sub bands. Wasserstein Generative Adversarial Networks (WGAN) based domain adaptation network aligns the marginal and conditional distribution of target and source. We evaluate our method on Data set 2a of BCI competition IV. Experiment results show that our method achieves the best performance among compared methods under different amount of training data. Performance of our method trained with certain blocks of data is similar to or better than the best comparing method trained with one more block. This indicates that our method could reduce the need for training data for at least one block. Shuang Qiu 0002, Wei Wei 0046, Xuelin Ma, Huiguang He |
IJCNN | 5 |
| 2021 | Multi-subject data augmentation for target subject semantic decoding with deep multi-view adversarial learning
Changde Du, Shengpei Wang, Haibao Wang, Huiguang He |
Inf. Sci. | 5 |
| 2021 | SNAP: Shaping neural architectures progressively via information density criterion
Zhiqiang Chen 0002, Ting-Bing Xu, Weijian Liao, Zhengcheng Li, Jinpeng Li 0002, Cheng-Lin Liu 0001, Huiguang He |
Pattern Recognit. | 7 |
| 2021 | Multi-task contrastive learning for automatic CT and X-ray diagnosis of COVID-19
Jinpeng Li 0002, Gangming Zhao, Yaling Tao, Penghua Zhai, Hao Chen 0081, Huiguang He, Ting Cai 0001 |
Pattern Recognit. | 6 |
| 2021 | A prototype-based SPD matrix network for domain adaptation EEG emotion recognition
Shuang Qiu 0002, Xuelin Ma, Huiguang He |
Pattern Recognit. | 4 |
| 2021 | Boundary Aware U-Net for Retinal Layers Segmentation in Optical Coherence Tomography ImagesabstractRetinal layers segmentation in optical coherence tomography (OCT) images is a critical step in the diagnosis of numerous ocular diseases. Automatic layers segmentation requires separating each individual layer instance with accurate boundary detection, but remains a challenging task since it suffers from speckle noise, intensity inhomogeneity, and the low contrast around boundary. In this work, we proposed a boundary aware U-Net (BAU-Net) for retinal layers segmentation by detecting accurate boundary. Based on encoder-decoder architecture, we design a dual tasks framework with low-level outputs for boundary detection and high-level outputs for layers segmentation. Specifically, we first use the multi-scale input strategy to enrich the spatial information in the deep features of encoder. For low-level features from encoder, we design an edge aware (EA) module in skip connection to extract the pure edge features. Then, a U-structure feature enhanced (UFE) module is designed in all skip connections to enlarge the features receptive fields from the encoder. Besides, a canny edge fusion (CEF) module is introduced to aforementioned architecture, which can fuse the priory edge information from segmentation task to boundary detection branch for a better predication. Furthermore, we model each boundary as a vertical coordinates distribution for boundary detection. Based on this distribution, a topology guarantee loss with combined A-scan regression loss and structure loss is proposed to make an accurate and guaranteed topological boundary set. The method is evaluated on two public datasets and the results demonstrate that the BAU-Net achieves promising performance than other state-of-the-art methods. Bo Wang 0168, Wei Wei 0046, Shuang Qiu 0002, Shengpei Wang, Huiguang He |
IEEE J. Biomed. Health Informatics | 6 |
| 2021 | CSU-Net: A Context Spatial U-Net for Accurate Blood Vessel Segmentation in Fundus ImagesabstractBlood vessel segmentation in fundus images is a critical procedure in the diagnosis of ophthalmic diseases. Recent deep learning methods achieve high accuracy in vessel segmentation but still face the challenge to segment the microvascular and detect the vessel boundary. This is due to the fact that common Convolutional Neural Networks (CNN) are unable to preserve rich spatial information and a large receptive field simultaneously. Besides, CNN models for vessel segmentation usually are trained by equal pixel level cross-entropy loss, which tend to miss fine vessel structures. In this paper, we propose a novel Context Spatial U-Net (CSU-Net) for blood vessel segmentation. Compared with the other U-Net based models, we design a two-channel encoder: a context channel with multi-scale convolution to capture more receptive field and a spatial channel with large kernel to retain spatial information. Also, to combine and strengthen the features extracted from two paths, we introduce a feature fusion module (FFM) and an attention skip module (ASM). Furthermore, we propose a structure loss, which adds a spatial weight to cross-entropy loss and guide the network to focus more on the thin vessels and boundaries. We evaluated this model on three public datasets: DRIVE, CHASE-DB1 and STARE. The results show that the CSU-Net achieves higher segmentation accuracy than the current state-of-the-art methods. Bo Wang 0168, Shengpei Wang, Shuang Qiu 0002, Wei Wei 0046, Haibao Wang, Huiguang He |
IEEE J. Biomed. Health Informatics | 6 |
| 2021 | Dynamical Channel Pruning by Conditional Accuracy Change for Deep Neural NetworksabstractChannel pruning is an effective technique that has been widely applied to deep neural network compression. However, many existing methods prune from a pretrained model, thus resulting in repetitious pruning and fine-tuning processes. In this article, we propose a dynamical channel pruning method, which prunes unimportant channels at the early stage of training. Rather than utilizing some indirect criteria (e.g., weight norm, absolute weight sum, and reconstruction error) to guide connection or channel pruning, we design criteria directly related to the final accuracy of a network to evaluate the importance of each channel. Specifically, a channelwise gate is designed to randomly enable or disable each channel so that the conditional accuracy changes (CACs) can be estimated under the condition of each channel disabled. Practically, we construct two effective and efficient criteria to dynamically estimate CAC at each iteration of training; thus, unimportant channels can be gradually pruned during the training process. Finally, extensive experiments on multiple data sets (i.e., ImageNet, CIFAR, and MNIST) with various networks (i.e., ResNet, VGG, and MLP) demonstrate that the proposed method effectively reduces the parameters and computations of baseline network while yielding the higher or competitive accuracy. Interestingly, if we Double the initial Channels and then Prune Half (DCPH) of them to baseline's counterpart, it can enjoy a remarkable performance improvement by shaping a more desirable structure. Zhiqiang Chen 0002, Ting-Bing Xu, Changde Du, Cheng-Lin Liu 0001, Huiguang He |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2020 | Conditional Generative Neural Decoding with Structured CNN Feature PredictionabstractDecoding visual contents from human brain activity is a challenging task with great scientific value. Two main facts that hinder existing methods from producing satisfactory results are 1) typically small paired training data; 2) under-exploitation of the structural information underlying the data. In this paper, we present a novel conditional deep generative neural decoding approach with structured intermediate feature prediction. Specifically, our approach first decodes the brain activity to the multilayer intermediate features of a pretrained convolutional neural network (CNN) with a structured multi-output regression (SMR) model, and then inverts the decoded CNN features to the visual images with an introspective conditional generation (ICG) model. The proposed SMR model can simultaneously leverage the covariance structures underlying the brain activities, the CNN features and the prediction tasks to improve the decoding accuracy and interpretability. Further, our ICG model can 1) leverage abundant unpaired images to augment the training data; 2) self-evaluate the quality of its conditionally generated images; and 3) adversarially improve itself without extra discriminator. Experimental results show that our approach yields state-of-the-art visual reconstructions from brain activities. Changde Du, Changying Du, Huiguang He |
AAAI | 4 |
| 2020 | Simultaneous Neural Spike Encoding and Decoding Based on Cross-modal Dual Deep Generative ModelabstractNeural encoding and decoding of retinal ganglion cells (RGCs) have been attached great importance in the research work of brain-machine interfaces. Much effort has been invested to mimic RGC and get insight into RGC signals to reconstruct stimuli. However, there remain two challenges. On the one hand, complex nonlinear processes in retinal neural circuits hinder encoding models from enhancing their ability to fit the natural stimuli and modelling RGCs accurately. On the other hand, current research of the decoding process is separate from that of the encoding process, in which the liaison of mutual promotion between them is neglected. In order to alleviate the above problems, we propose a cross-modal dual deep generative model (CDDG) in this paper. CDDG treats the RGC spike signals and the stimuli as two modalities, which learns a shared latent representation for the concatenated modality and two modal-specific latent representations. Then, it imposes distribution consistency restriction on different latent space, cross-consistency and cycle-consistency constraints on the generated variables. Thus, our model ensures cross-modal generation from RGC spike signals to stimuli and vice versa. In our framework, the generation from stimuli to RGC spike signals is equivalent to neural encoding while the inverse process is equivalent to neural decoding. Hence, the proposed method integrates neural encoding and decoding and exploits the reciprocity between them. The experimental results demonstrate that our proposed method can achieve excellent encoding and decoding performance compared with the state-of-the-art methods on three salamander RGC spike datasets with natural stimuli. Qiongyi Zhou, Changde Du, Haibao Wang, Jian K. Liu, Huiguang He |
IJCNN | 6 |
| 2020 | A CNN-based comparing network for the detection of steady-state visual evoked potential responses
Jiezhen Xing, Shuang Qiu 0002, Xuelin Ma, Chenyao Wu, Jinpeng Li 0002, Shengpei Wang, Huiguang He |
Neurocomputing | 7 |
| 2020 | Semi-supervised cross-modal image generation with generative adversarial networks
Changde Du, Huiguang He |
Pattern Recognit. | 3 |
| 2020 | Multisource Transfer Learning for Cross-Subject EEG Emotion RecognitionabstractElectroencephalogram (EEG) has been widely used in emotion recognition due to its high temporal resolution and reliability. Since the individual differences of EEG are large, the emotion recognition models could not be shared across persons, and we need to collect new labeled data to train personal models for new users. In some applications, we hope to acquire models for new persons as fast as possible, and reduce the demand for the labeled data amount. To achieve this goal, we propose a multisource transfer learning method, where existing persons are sources, and the new person is the target. The target data are divided into calibration sessions for training and subsequent sessions for test. The first stage of the method is source selection aimed at locating appropriate sources. The second is style transfer mapping, which reduces the EEG differences between the target and each source. We use few labeled data in the calibration sessions to conduct source selection and style transfer. Finally, we integrate the source models to recognize emotions in the subsequent sessions. The experimental results show that the three-category classification accuracy on benchmark SEED improves by 12.72% comparing with the nontransfer method. Our method facilitates the fast deployment of emotion recognition models by reducing the reliance on the labeled data amount, which has practical significance especially in fast-deployment scenarios. Jinpeng Li 0002, Shuang Qiu 0002, Yuan-Yuan Shen, Cheng-Lin Liu 0001, Huiguang He |
IEEE Trans. Cybern. | 5 |
| 2019 | Doubly Semi-Supervised Multimodal Adversarial Learning for Classification, Generation and RetrievalabstractLearning over incomplete multi-modality data is a challenging problem with strong practical applications. Most existing multi-modal data imputation approaches have two limitations: (1) they are unable to accurately control the semantics of imputed modalities; and (2) without a shared low-dimensional latent space, they do not scale well with multiple modalities. To overcome the limitations, we propose a novel doubly semi-supervised multi-modal learning framework (DSML) with a modality-shared latent space and modality-specific generators, encoders and classifiers. We design novel softmax-based discriminators to train all modules adversarially. As a unified framework, DSML can be applied in multi-modal semi-supervised classification, missing modality imputation and fast cross-modality retrieval tasks simultaneously. Experiments on multiple datasets demonstrate its advantages. Changde Du, Changying Du, Huiguang He |
ICME | 3 |
| 2019 | Learning "What" and "Where": An Interpretable Neural Encoding ModelabstractNeural encoding modeling aims to reveal how brain processes perceived information by establishing a quantitative relationship between stimuli and evoked brain activities. In the field of visual neuroscience, many studies have been dedicated to building the neural encoding model for primary visual cortex and demonstrate that the population receptive field (pRF) models can be used to explain how neurons in primary visual cortex work. However, these models rely on either the inflexible prior assumptions imposed on the spatial characteristics of pRF or the clumsy parameter estimation methods which requiring too much manual adjustment. Suffering from these issues, current methods yield dissatisfactory performance on mimicking brain activity. In this paper, we address the problems under a novel "what and where" neural encoding framework. Basing on deep neural network (DNN) and the separability of the spatial ("where") and visual feature ("what") dimensions, the proposed method is not only powerful in extracting nonlinear features from images, but also rich in interpretability. Owing to two forms of regularization: sparsity and smoothness, receptive fields are estimated automatically for each voxel without prior assumptions on shape, which gets rid of the shortcomings of previous methods. Extensive empirical evaluations on publicly available fMRI dataset show that the proposed method has superior performance gains over several existing methods. Haibao Wang, Changde Du, Huiguang He |
IJCNN | 4 |
| 2019 | Dual Encoding U-Net for Retinal Vessel Segmentation
Bo Wang 0168, Shuang Qiu 0002, Huiguang He |
MICCAI (1) | 3 |
| 2019 | Automatic brain labeling via multi-atlas guided fully convolutional networks
Longwei Fang, Lichi Zhang, Dong Nie, Xiaohuan Cao, Islem Rekik, Seong-Whan Lee, Huiguang He, Dinggang Shen |
Medical Image Anal. | 7 |
| 2019 | Reconstructing Perceived Images From Human Brain Activities With Bayesian Deep Multiview LearningabstractNeural decoding, which aims to predict external visual stimuli information from evoked brain activities, plays an important role in understanding human visual system. Many existing methods are based on linear models, and most of them only focus on either the brain activity pattern classification or visual stimuli identification. Accurate reconstruction of the perceived images from the measured human brain activities still remains challenging. In this paper, we propose a novel deep generative multiview model for the accurate visual image reconstruction from the human brain activities measured by functional magnetic resonance imaging (fMRI). Specifically, we model the statistical relationships between the two views (i.e., the visual stimuli and the evoked fMRI) by using two view-specific generators with a shared latent space. On the one hand, we adopt a deep neural network architecture for visual image generation, which mimics the stages of human visual processing. On the other hand, we design a sparse Bayesian linear model for fMRI activity generation, which can effectively capture voxel correlations, suppress data noise, and avoid overfitting. Furthermore, we devise an efficient mean-field variational inference method to train the proposed model. The proposed method can accurately reconstruct visual images via Bayesian inference. In particular, we exploit a posterior regularization technique in the Bayesian inference to regularize the model posterior. The quantitative and qualitative evaluations conducted on multiple fMRI data sets demonstrate the proposed method can reconstruct visual images more accurately than the state of the art. Changde Du, Changying Du, Huiguang He |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2018 | Improving Image Classification Performance with Automatically Hierarchical Label ClusteringabstractImage classification is a common and foundational problem in computer vision. In traditional image classification, a category is assigned with single label, which is difficult for networks to learn better features. On the contrary, hierarchical labels can depict the structure of categories better, which helps network to learn more hierarchical features and improve the classification performance. Though many datasets contain images with multi-labels, the labels in these datasets usually lack of hierarchy. To overcome this problem, we propose a new method to improve image classification performance with Automatically Hierarchical Label Clustering (AHLC). Firstly, AHLC calculates the similarity between each pair of original categories by how easily they are misclassified with a pre-trained classifier. Secondly, AHLC obtains hierarchical labels by merging similar categories using hierarchical clustering. Finally, AHLC trains a new classifier with hierarchical labels to improve the original classification performance. We evaluate our method on MNIST and CIFAR-100 datasets and the results demonstrate the superiority of our method. The main contribution of this work is that we can simply improve an existing classification network by AHLC without extra information or heavy architecture redesign. Zhiqiang Chen 0002, Changde Du, Huiguang He |
ICPR | 5 |
| 2018 | Multi-label Semantic Decoding from Human Brain ActivityabstractIt is meaningful to decode the semantic information from functional magnetic resonance imaging (fMRI) brain signals evoked by natural images. Semantic decoding can be viewed as a classification problem. Since a natural image may contain many semantic information of different objects, the single label classification model is not appropriate to cope with semantic decoding problem, which motivates the multi-label classification model. However, most multi-label models always treat each label equally. Actually, if dataset is associated with a large number of semantic labels, it will be difficult to get an accurate prediction of semantic label when the label appears with a low frequency in this dataset. So we should increase the relative importance degree to the labels that associate with little instances. In order to improve multi-label prediction performance, in this paper, we firstly propose a multinomial label distribution to estimate the importance degree of each associated label for an instance by using conditional probability, and then establish a deep neural network (DNN) based model which contains both multinomial label distribution and label co-occurrence information to realize the multi-label classification of semantic information in fMRI brain signals. Experiments on three fMRI recording datasets demonstrate that our approach performs better than the state-of-the-art methods on semantic information prediction. Changde Du, Zhiqiang Chen 0002, Huiguang He |
ICPR | 5 |
| 2018 | Semi-supervised Deep Generative Modelling of Incomplete Multi-Modality Emotional DataabstractThere are threefold challenges in emotion recognition. First, it is difficult to recognize human's emotional states only considering a single modality. Second, it is expensive to manually annotate the emotional data. Third, emotional data often suffers from missing modalities due to unforeseeable sensor malfunction or configuration issues. In this paper, we address all these problems under a novel multi-view deep generative framework. Specifically, we propose to model the statistical relationships of multi-modality emotional data using multiple modality-specific generative networks with a shared latent space. By imposing a Gaussian mixture assumption on the posterior approximation of the shared latent variables, our framework can learn the joint deep representation from multiple modalities and evaluate the importance of each modality simultaneously. To solve the labeled-data-scarcity problem, we extend our multi-view model to semi-supervised learning scenario by casting the semi-supervised classification problem as a specialized missing data imputation task. To address the missing-modality problem, we further extend our semi-supervised multi-view model to deal with incomplete data, where a missing view is treated as a latent variable and integrated out during inference. This way, the proposed overall framework can utilize all available (both labeled and unlabeled, as well as both complete and incomplete) data to improve its generalization ability. The experiments conducted on two real multi-modal emotion datasets demonstrated the superiority of our framework. Changde Du, Changying Du, Hao Wang 0005, Jinpeng Li 0002, Wei-Long Zheng, Bao-Liang Lu, Huiguang He |
ACM Multimedia | 7 |
| 2018 | Predicting Epileptic Seizures from Intracranial EEG Using LSTM-Based Multi-task Learning
Xuelin Ma, Shuang Qiu 0002, Xiaoqin Lian, Huiguang He |
PRCV (2) | 5 |
| 2017 | Sharing deep generative representation for perceived image reconstruction from human brain activityabstractDecoding human brain activities via functional magnetic resonance imaging (fMRI) has gained increasing attention in recent years. While encouraging results have been reported in brain states classification tasks, reconstructing the details of human visual experience still remains difficult. Two main challenges that hinder the development of effective models are the perplexing fMRI measurement noise and the high dimensionality of limited data instances. Existing methods generally suffer from one or both of these issues and yield dissatisfactory results. In this paper, we tackle this problem by casting the reconstruction of visual stimulus as the Bayesian inference of missing view in a multiview latent variable model. Sharing a common latent representation, our joint generative model of external stimulus and brain response is not only “deep” in extracting nonlinear features from visual images, but also powerful in capturing correlations among voxel activities of fMRI recordings. The nonlinearity and deep structure endow our model with strong representation ability, while the correlations of voxel activities are critical for suppressing noise and improving prediction. We devise an efficient variational Bayesian method to infer the latent variables and the model parameters. To further improve the reconstruction accuracy, the latent representations of testing instances are enforced to be close to that of their neighbours from the training set via posterior regularization. Experiments on three fMRI recording datasets demonstrate that our approach can more accurately reconstruct visual stimuli. Changde Du, Changying Du, Huiguang He |
IJCNN | 3 |
| 2017 | Multi-modal multiple kernel learning for accurate identification of Tourette syndrome children
Hongwei Wen, Islem Rekik, Shengpei Wang, Zhiqiang Chen 0002, Jishui Zhang, Yun Peng 0005, Huiguang He |
Pattern Recognit. | 9 |
| 2015 | Spherical volume-preserving Demons registration
Xuejiao Chen, Jiaxi Hu, Huiguang He, Jing Hua 0001 |
Comput. Aided Des. | 3 |
| 2013 | Ricci flow-based spherical parameterization and surface registration
Huiguang He, Guangyu Zou, Xiaopeng Zhang 0001, Xianfeng Gu, Jing Hua 0001 |
Comput. Vis. Image Underst. | 2 |
| 2013 | Accurate prediction of AD patients using cortical thickness networks
Dai Dai, Huiguang He, Joshua T. Vogelstein, Zeng-Guang Hou |
Mach. Vis. Appl. | 2 |
| 2010 | Comparative Analysis of Quasi-Conformal Deformations in Shape Space
Vahid Taimouri, Huiguang He, Jing Hua 0001 |
MICCAI (3) | 2 |
| 2006 | CWME: A Framework of Group Support System for Emergency Responses
Yaodong Li, Huiguang He, Baihua Xiao, Chunheng Wang, Fei-Yue Wang 0001 |
ISI | 2 |
| 2005 | The design and implementation of a novel platform for medical data visualizationabstractAs medical imaging applications become more complex, the design of software platforms for medical imaging is getting greater priority. With this demand, we have designed and implemented a novel software platform for medical data visualization in traditional object-oriented fashion with some common design patterns. This platform integrates the mainstream medical data reconstruction and visualization algorithms and 3D human-computer interaction based on 3D widgets into a consistent framework. The design goals, the overall framework and the implementation of some key technologies are addressed in details, and some application examples are also given to demonstrate the abilities of this platform. The ultimate objective is to provide a flexible, reliable and extensible 3D medical data visualization platform for the medical imaging society. Jian Xue 0002, Jie Tian 0001, Mingchang Zhao, Huiguang He |
CSCWD (2) | 4 |
| 2002 | A New 3D Surface Reconstruction Method and its Application in Industrial CTabstractA fast surface reconstruction algorithm is proposed in this paper for processing large scale industrial CT images with high resolution. Through the following main steps: surface tracking in the single layer, a data caching mechanism and triangle strip generation, the algorithm can extract and represent surfaces efficiently on current mainstream PCs. The experimental results tested in real industrial CT datasets are also reported. Mingchang Zhao, Jie Tian 0001, Huiguang He |
CSCWD | 3 |