Changde Du

dblp:178/4485 · DBLP profile ↗
← Back
39ranked-venue papers
8as first author
23since 2021 · last 2026
0000-0002-0084-433XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021
YearPublicationVenuePosition
2026 Prompt-guided dual-channel attention model predicts brain activation from functional and structural profiles
Wei Huang 0016, Hengjiang Li, Sizhuo Wang, Changde Du, Kaiwen Cheng, Huafu Chen
Pattern Recognit.7
2026 Quantization and disentanglement for cross-modal alignment in neural speech reconstruction from brain activity
abstract
Understanding and reconstructing speech from brain activity is central to advancing neuroscience and brain–computer interface (BCI) research. Despite recent progress, current brain-to-speech approaches remain limited by entangled neural representations, overly fine-grained decoding targets, and poor generalization. To address these challenges, we propose a novel framework that integrates quantization with representation disentanglement. Instead of learning brain-specific discrete units, we directly align magnetoencephalography (MEG) signals with the codebook of a pretrained neural audio codec. To balance representational capacity and decoding complexity, we introduce a discrete contrastive loss combining codebook clustering and temporal pooling. Furthermore, we factorize speech into prosody, content, and timbre, and decode these subspaces separately, simplifying learning and enhancing decoding precision. Experiments on two public MEG datasets demonstrate strong performance in both intra- and cross-subject settings, yielding substantial improvements in prosody-related and perceptual metrics. Our findings highlight a principled pathway towards more reliable brain-to-speech decoding, with potential applications in assistive communication and BCI technologies.
Changde Du, Huiguang He
Pattern Recognit.2
2025 BP-GPT: Auditory Neural Decoding Using fMRI-prompted LLM
abstract
Decoding language information from brain signals represents a vital research area within brain-computer interfaces, particularly in the context of deciphering the semantic information from the fMRI signal. Although existing work uses LLM to achieve this goal, their method does not use an end-to-end approach and avoids the LLM in the mapping of fMRI-to-text, leaving space for the exploration of the LLM in auditory decoding. In this paper, we introduce a novel method, the Brain Prompt GPT (BP-GPT). By using the brain representation that is extracted from the fMRI as a prompt, our method can utilize GPT-2 to decode fMRI signals into stimulus text. Further, we introduce the text prompt and align the fMRI prompt to it. By introducing the text prompt, our BP-GPT can extract a more robust brain prompt and promote the decoding of pre-trained LLM. We evaluate our BP-GPT on the open-source auditory semantic decoding dataset and achieve a significant improvement up to 4.61% on METEOR and 2.43% on BERTScore across all the subjects compared to the state-of-the-art method. The experimental results demonstrate that using brain representation as a prompt to further drive LLM for auditory neural decoding is feasible and effective. The code is available at https://github.com/1994cxy/BP-GPT.
Changde Du, Huiguang He
ICASSP2
2025 Animate Your Thoughts: Reconstruction of Dynamic Natural Vision from Human Brain Activity
abstract
Reconstructing human dynamic vision from brain activity is a challenging task with great scientific significance. Although prior video reconstruction methods have made substantial progress, they still suffer from several limitations, including: (1) difficulty in simultaneously reconciling semantic (e.g. categorical descriptions), structure (e.g. size and color), and consistent motion information (e.g. order of frames); (2) low temporal resolution of fMRI, which poses a challenge in decoding multiple frames of video dynamics from a single fMRI frame; (3) reliance on video generation models, which introduces ambiguity regarding whether the dynamics observed in the reconstructed videos are genuinely derived from fMRI data or are hallucinations from generative model. To overcome these limitations, we propose a two-stage model named Mind-Animator. During the fMRI-to-feature stage, we decouple semantic, structure, and motion features from fMRI. Specifically, we employ fMRI-vision-language tri-modal contrastive learning to decode semantic feature from fMRI and design a sparse causal attention mechanism for decoding multi-frame video motion features through a next-frame-prediction task. In the feature-to-video stage, these features are integrated into videos using an inflated Stable Diffusion, effectively eliminating external video data interference. Extensive experiments on multiple video-fMRI datasets demonstrate that our model achieves state-of-the-art performance. Comprehensive visualization analyses further elucidate the interpretability of our model from a neurobiological perspective. Project page: https://mind-animator-design.github.io/.
Yizhuo Lu, Changde Du, Xuanliu Zhu, Liuyun Jiang, Xujin Li, Huiguang He
ICLR2
2025 EmoGrowth: Incremental Multi-label Emotion Decoding with Augmented Emotional Relation Graph
abstract
Emotion recognition systems face significant challenges in real-world applications, where novel emotion categories continually emerge and multiple emotions often co-occur. This paper introduces multi-label fine-grained class incremental emotion decoding, which aims to develop models capable of incrementally learning new emotion categories while maintaining the ability to recognize multiple concurrent emotions. We propose an Augmented Emotional Semantics Learning (AESL) framework to address two critical challenges: past- and future-missing partial label problems. AESL incorporates an augmented Emotional Relation Graph (ERG) for reliable soft label generation and affective dimension-based knowledge distillation for future-aware feature learning. We evaluate our approach on three datasets spanning brain activity and multimedia domains, demonstrating its effectiveness in decoding up to 28 fine-grained emotion categories. Results show that AESL significantly outperforms existing methods while effectively mitigating catastrophic forgetting. Our code is available at https://github.com/ChangdeDu/EmoGrowth.
Kaicheng Fu, Changde Du, Shuangchen Zhao, Huiguang He
ICML2
2025 Adaptive Domain Alignment Neural Networks for Cross-Domain EEG Emotion Recognition
abstract
Electroencephalography (EEG) - based Emotion recognition is now facing great challenge of the intra- and inter-subject variability of EEG signal. Researchers attempted to handle this challenge by using transfer learning methods which usually share two main limitations: most of these methods align marginal distributions instead of conditional distributions of source and target data, making the alignment process classwise ambiguous; also, they prefer to use Multi-Layer Perceptron (MLP) with redundant parameters as classifiers, which is shown by recent research that could result serious over-fitting towards labeled data and prevent the model to draw a proper representation space. In our work, we propose a novel domain alignment method: Adaptive Domain Alignment Neural Networks (ADANN). Our method directly model conditional distributions of source and target domains by two sets of label-wise prototypes, representing the density maximum of each class, while the normalized correspond similarity naturally represents the conditional probability. The predicted label for a sample is given by the argument maxima of similarities and therefore the MLP classifier is not required. Using context-instance contrastive learning to align two sets of prototypes, their corresponding conditional distributions are being learned simultaneously. Exhaustive cross-domain experiments have been conducted under protocols that are strongly related to practical application scenarios and our proposed method achieves better or similar performance compared with recent state-of-the-art methods.
Xuezhu Hong, Changde Du, Huiguang He
IEEE Trans. Affect. Comput.2
2024 CLIP-MUSED: CLIP-Guided Multi-Subject Visual Neural Information Semantic Decoding
abstract
The study of decoding visual neural information faces challenges in generalizing single-subject decoding models to multiple subjects, due to individual differences. Moreover, the limited availability of data from a single subject has a constraining impact on model performance. Although prior multi-subject decoding methods have made significant progress, they still suffer from several limitations, including difficulty in extracting global neural response features, linear scaling of model parameters with the number of subjects, and inadequate characterization of the relationship between neural responses of different subjects to various stimuli. To overcome these limitations, we propose a CLIP-guided Multi-sUbject visual neural information SEmantic Decoding (CLIP-MUSED) method. Our method consists of a Transformer-based feature extractor to effectively model global neural representations. It also incorporates learnable subject-specific tokens that facilitates the aggregation of multi-subject data without a linear increase of parameters. Additionally, we employ representational similarity analysis (RSA) to guide token representation learning based on the topological relationship of visual stimuli in the representation space of CLIP, enabling full characterization of the relationship between neural responses of different subjects under different stimuli. Finally, token representations are used for multi-subject semantic decoding. Our proposed method outperforms single-subject decoding methods and achieves state-of-the-art performance among the existing multi-subject methods on two fMRI datasets. Visualization results provide insights into the effectiveness of our proposed method. Code is available at https://github.com/CLIP-MUSED/CLIP-MUSED.
Qiongyi Zhou, Changde Du, Shengpei Wang, Huiguang He
ICLR2
2024 Improved Video Emotion Recognition With Alignment of CNN and Human Brain Representations
abstract
The ability to perceive emotions is an important criterion for judging whether a machine is intelligent. To this end, a large number of emotion recognition algorithms have been developed especially for visual information such as video. Most previous studies are based on hand-crafted features or CNN, in which the former fails to extract expressive features and the latter still faces the undesired affective gap. This drives us to think about what if we could incorporate the human emotional perception capability into CNN. In this paper, we attempt to address this question by exploring alignment between representations of neural networks and human brain activity. In particular, we employ a visually evoked emotional brain activity dataset to conduct a jointly training strategy for CNN. In the training phase, we introduce the representation similarity analysis (RSA) to align the CNN with human brain to obtain more brain-like features. Specifically, representation similarity matrices (RSMs) of multiple convolutional layers are averaged with learnable weights and related to the RSM of human brain. In order to obtain emotion-related brain activity, we conduct voxel selection and denoising with a banded ridge model before computing the RSM. Sufficient experiments on two challenging video emotion recognition datasets and multiple popular CNN architectures suggest that human brain activity is promising to provide an inductive bias for CNN towards better performance of emotion recognition. Our source code is available inhttps://osf.io/ucx57.
Kaicheng Fu, Changde Du, Shengpei Wang, Huiguang He
IEEE Trans. Affect. Comput.2
2024 PGCN: Pyramidal Graph Convolutional Network for EEG Emotion Recognition
abstract
Emotion recognition is essential in the diagnosis and rehabilitation of various mental diseases. In the last decade, electroencephalogram (EEG)-based emotion recognition has been intensively investigated due to its prominative accuracy and reliability, and graph convolutional network (GCN) has become a mainstream model to decode emotions from EEG signals. However, the electrode relationship, especially long-range electrode dependencies across the scalp, may be underutilized by GCNs, although such relationships have been proven to be important in emotion recognition. The small receptive field makes shallow GCNs only aggregate local nodes. On the other hand, stacking too many layers leads to over-smoothing. To solve these problems, we propose the pyramidal graph convolutional network (PGCN), which aggregates features at three levels: local, mesoscopic, and global. First, we construct a vanilla GCN based on the 3D topological relationships of electrodes, which is used to integrate two-order local features; Second, we construct several mesoscopic brain regions based on priori knowledge and employ mesoscopic attention to sequentially calculate the virtual mesoscopic centers to focus on the functional connections of mesoscopic brain regions; Finally, we fuse the node features and their 3D positions to construct a numerical relationship adjacency matrix to integrate structural and functional connections from the global perspective. Experimental results on four public datasets indicate that PGCN enhances the relationship modelling across the scalp and achieves stateof-the-art performance in both subject-dependent and subjectindependent scenarios. Meanwhile, PGCN makes an effective trade-off between enhancing network depth and receptive fields while suppressing the ensuing over-smoothing. Our codes are publicly accessible athttps://github.com/Jinminbox/PGCN.
Ming Jin 0006, Changde Du, Huiguang He, Ting Cai 0001, Jinpeng Li 0002
IEEE Trans. Multim.2
2024 Multi-View Multi-Label Fine-Grained Emotion Decoding From Human Brain Activity
abstract
Decoding emotional states from human brain activity play an important role in the brain-computer interfaces. Existing emotion decoding methods still have two main limitations: one is only decoding a single emotion category from a brain activity pattern and the decoded emotion categories are coarse-grained, which is inconsistent with the complex emotional expression of humans; the other is ignoring the discrepancy of emotion expression between the left and right hemispheres of the human brain. In this article, we propose a novel multi-view multi-label hybrid model for fine-grained emotion decoding (up to 80 emotion categories) which can learn the expressive neural representations and predict multiple emotional states simultaneously. Specifically, the generative component of our hybrid model is parameterized by a multi-view variational autoencoder, in which we regard the brain activity of left and right hemispheres and their difference as three distinct views and use the product of expert mechanism in its inference network. The discriminative component of our hybrid model is implemented by a multi-label classification network with an asymmetric focal loss. For more accurate emotion decoding, we first adopt a label-aware module for emotion-specific neural representation learning and then model the dependency of emotional states by a masked self-attention mechanism. Extensive experiments on two visually evoked emotional datasets show the superiority of our method.
Kaicheng Fu, Changde Du, Shengpei Wang, Huiguang He
IEEE Trans. Neural Networks Learn. Syst.2
2023 RoBrain: Towards Robust Brain-to-Image Reconstruction via Cross-Domain Contrastive Learning
Changde Du, Huiguang He
ICONIP (3)2
2023 Mammo-Net: Integrating Gaze Supervision and Interactive Information in Multi-view Mammogram Classification
Changkai Ji, Changde Du, Sheng Wang 0014, Chong Ma 0004, Jiaming Xie, Huiguang He, Dinggang Shen
MICCAI (7)2
2023 Auditory Attention Decoding with Task-Related Multi-View Contrastive Learning
abstract
The human brain can easily focus on one speaker and suppress others in scenarios such as a cocktail party. Recently, researchers found that auditory attention can be decoded from the electroencephalogram (EEG) data. However, most existing deep learning methods are difficult to use prior knowledge of different views (that is attended speech and EEG are task-related views) and extract an unsatisfactory representation. Inspired by Broadbent's filter model, we decode auditory attention in a multi-view paradigm and extract the most relevant and important information utilizing the missing view. Specifically, we propose an auditory attention decoding (AAD) method based on multi-view VAE with task-related multi-view contrastive (TMC) learning. Employing TMC learning in multi-view VAE can utilize the missing view to accumulate prior knowledge of different views into the fusion of representation, and extract the approximate task-related representation. We examine our method on two popular AAD datasets, and demonstrate the superiority of our method by comparing it to the state-of-the-art method.
Changde Du, Qiongyi Zhou, Huiguang He
ACM Multimedia2
2023 MindDiffuser: Controlled Image Reconstruction from Human Brain Activity with Semantic and Structural Diffusion
abstract
Reconstructing visual stimuli from brain recordings has been a meaningful and challenging task. Especially, the achievement of precise and controllable image reconstruction bears great significance in propelling the progress and utilization of brain-computer interfaces. Despite the advancements in complex image reconstruction techniques, the challenge persists in achieving a cohesive alignment of both semantic (concepts and objects) and structure (position, orientation, and size) with the image stimuli. To address the aforementioned issue, we propose a two-stage image reconstruction model called MindDiffuser1. In Stage 1, the VQ-VAE latent representations and the CLIP text embeddings decoded from fMRI are put into Stable Diffusion, which yields a preliminary image that contains semantic information. In Stage 2, we utilize the CLIP visual feature decoded from fMRI as supervisory information, and continually adjust the two feature vectors decoded in Stage 1 through backpropagation to align the structural information. The results of both qualitative and quantitative analyses demonstrate that our model has surpassed the current state-of-the-art models on Natural Scenes Dataset (NSD). The subsequent experimental findings corroborate the neurobiological plausibility of the model, as evidenced by the interpretability of the multimodal feature employed, which align with the corresponding brain responses.
Yizhuo Lu, Changde Du, Qiongyi Zhou, Dianpeng Wang, Huiguang He
ACM Multimedia2
2023 Decoding Visual Neural Representations by Multimodal Learning of Brain-Visual-Linguistic Features
abstract
Decoding human visual neural representations is a challenging task with great scientific significance in revealing vision-processing mechanisms and developing brain-like intelligent machines. Most existing methods are difficult to generalize to novel categories that have no corresponding neural data for training. The two main reasons are 1) the under-exploitation of the multimodal semantic knowledge underlying the neural data and 2) the small number of paired (stimuli-responses) training data. To overcome these limitations, this paper presents a generic neural decoding method called BraVL that uses multimodal learning of brain-visual-linguistic features. We focus on modeling the relationships between brain, visual and linguistic features via multimodal deep generative models. Specifically, we leverage the mixture-of-product-of-experts formulation to infer a latent code that enables a coherent joint generation of all three modalities. To learn a more consistent joint representation and improve the data efficiency in the case of limited brain activity data, we exploit both intra- and inter-modality mutual information maximization regularization terms. In particular, our BraVL model can be trained under various semi-supervised scenarios to incorporate the visual and textual features obtained from the extra categories. Finally, we construct three trimodal matching datasets, and the extensive experiments lead to some interesting conclusions and cognitive insights: 1) decoding novel visual categories from human brain activity is practically possible with good accuracy; 2) decoding models using the combination of visual and linguistic features perform much better than those using either of them alone; 3) visual perception may be accompanied by linguistic influences to represent the semantics of visual stimuli.
Changde Du, Kaicheng Fu, Jinpeng Li 0002, Huiguang He
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Graph-Enhanced Emotion Neural Decoding
abstract
Brain signal-based emotion recognition has recently attracted considerable attention since it has powerful potential to be applied in human-computer interaction. To realize the emotional interaction of intelligent systems with humans, researchers have made efforts to decode human emotions from brain imaging data. The majority of current efforts use emotion similarities (e.g., emotion graphs) or brain region similarities (e.g., brain networks) to learn emotion and brain representations. However, the relationships between emotions and brain regions are not explicitly incorporated into the representation learning process. As a result, the learned representations may not be informative enough to benefit specific tasks, e.g., emotion decoding. In this work, we propose a novel idea of graph-enhanced emotion neural decoding, which takes advantage of a bipartite graph structure to integrate the relationships between emotions and brain regions into the neural decoding process, thus helping learn better representations. Theoretical analyses conclude that the suggested emotion-brain bipartite graph inherits and generalizes the conventional emotion graphs and brain networks. Comprehensive experiments on visually evoked emotion datasets demonstrate the effectiveness and superiority of our approach.
Zhongyu Huang, Changde Du, Yingheng Wang, Kaicheng Fu, Huiguang He
IEEE Trans. Medical Imaging2
2022 Graph Emotion Decoding from Visually Evoked Neural Responses
Zhongyu Huang, Changde Du, Yingheng Wang, Huiguang He
MICCAI (8)2
2022 VigilanceNet: Decouple Intra- and Inter-Modality Learning for Multimodal Vigilance Estimation in RSVP-Based BCI
abstract
Recently, brain-computer interface (BCI) technology has made impressive progress and has been developed for many applications. Thereinto, the BCI system based on rapid serial visual presentation (RSVP) is a promising information detection technology. However, the use of RSVP is closely related to the user's performance, which can be influenced by their vigilance levels. Therefore it is crucial to detect vigilance levels in RSVP-based BCI. In this paper, we conducted a long-term RSVP target detection experiment to collect electroencephalography (EEG) and electrooculogram (EOG) data at different vigilance levels. In addition, to estimate vigilance levels in RSVP-based BCI, we propose a multimodal method named VigilanceNet using EEG and EOG. Firstly, we define the multiplicative relationships in conventional EOG features that can better describe the relationships between EOG features, and design an outer product embedding module to extract the multiplicative relationships. Secondly, we propose to decouple the learning of intra- and inter-modality to improve multimodal learning. Specifically, for intra-modality, we introduce an intra-modality representation learning (intra-RL) method to obtain effective representations of each modality by letting each modality independently predict vigilance levels during the multimodal training process. For inter-modality, we employ the cross-modal Transformer based on cross-attention to capture the complementary information between EEG and EOG, which only pays attention to the inter-modality relations. Extensive experiments and ablation studies are conducted on the RSVP and SEED-VIG public datasets. The results demonstrate the effectiveness of the method in terms of regression error and correlation.
Wei Wei 0046, Changde Du, Shuang Qiu 0002, Sanli Tian, Huiguang He
ACM Multimedia3
2022 GREN: Graph-Regularized Embedding Network for Weakly-Supervised Disease Localization in X-Ray Images
abstract
Locating diseases in chest X-ray images with few careful annotations saves large human effort. Recent works approached this task with innovative weakly-supervised algorithms such as multi-instance learning (MIL) and class activation maps (CAM), however, these methods often yield inaccurate or incomplete regions. One of the reasons is the neglection of the pathological implications hidden in the relationship across anatomical regions within each image and the relationship across images. In this paper, we argue that the cross-region and cross-image relationship, as contextual and compensating information, is vital to obtain more consistent and integral regions. To model the relationship, we propose the Graph Regularized Embedding Network (GREN), which leverages the intra-image and inter-image information to locate diseases on chest X-ray images. GREN uses a pre-trained U-Net to segment the lung lobes, and then models the intra-image relationship between the lung lobes using an intra-image graph to compare different regions. Meanwhile, the relationship between in-batch images is modeled by an inter-image graph to compare multiple images. This process mimics the training and decision-making process of a radiologist: comparing multiple regions and images for diagnosis. In order for the deep embedding layers of the neural network to retain structural information (important in the localization task), we use the Hash coding and Hamming distance to compute the graphs, which are used as regularizers to facilitate training. By means of this, our approach achieves the state-of-the-art result on NIH chest X-ray dataset for weakly-supervised disease localization. Our codes are accessible online.
Baolian Qi, Gangming Zhao, Changde Du, Chengwei Pan, Yizhou Yu, Jinpeng Li 0002
IEEE J. Biomed. Health Informatics4
2022 Deep Modality Assistance Co-Training Network for Semi-Supervised Multi-Label Semantic Decoding
abstract
Multi-label semantic decoding is a challenging task with great scientific significance and application value. The existing methods mainly focus on label learning and ignore the amount of information contained in the sample itself, especially non-image sample, which may limit their performance. To address these issues, we propose a novel semi-supervised modality assistance co-training network, which utilizes image modality to assist non-image modality for multi-label learning. In real application, there are two thorny issues: (i) non-image modality tends to be missing owing to the difficulty in obtaining them; (ii) although the image modality is easy to obtain from the Internet, image label annotation is still time-consuming and expensive. Therefore, the proposed method utilizes a small number of paired & labeled images and non-image modalities, and a large number of unpaired & unlabeled images from web sources to improve results. It consists of the modality-specific feature generators, the feature translators and the label relationship network. Specifically, the modality-specific feature generators are used to generate different features (views) for each modality. Semantic translators are employed to capture the relationship between the paired modalities and impute the missing modality feature by using unpaired & unlabeled images. Label relation network is a graph convolution network (GCN) aiming to capture the correlation between labels. To mine the information in unlabeled features, the co-training mechanism is considered. With this mechanism, we introduce a multi-view orthogonality constraint and a multi-label co-regularization constraint. Extensive experiments on three computer vision and neuroscience datasets demonstrate the effectiveness of the proposed method.
Changde Du, Haibao Wang, Qiongyi Zhou, Huiguang He
IEEE Trans. Multim.2
2022 Structured Neural Decoding With Multitask Transfer Learning of Deep Neural Network Representations
abstract
The reconstruction of visual information from human brain activity is a very important research topic in brain decoding. Existing methods ignore the structural information underlying the brain activities and the visual features, which severely limits their performance and interpretability. Here, we propose a hierarchically structured neural decoding framework by using multitask transfer learning of deep neural network (DNN) representations and a matrix-variate Gaussian prior. Our framework consists of two stages, Voxel2Unit and Unit2Pixel. In Voxel2Unit, we decode the functional magnetic resonance imaging (fMRI) data to the intermediate features of a pretrained convolutional neural network (CNN). In Unit2Pixel, we further invert the predicted CNN features back to the visual images. Matrix-variate Gaussian prior allows us to take into account the structures between feature dimensions and between regression tasks, which are useful for improving decoding effectiveness and interpretability. This is in contrast with the existing single-output regression models that usually ignore these structures. We conduct extensive experiments on two real-world fMRI data sets, and the results show that our method can predict CNN features more accurately and reconstruct the perceived natural images and faces with higher quality.
Changde Du, Changying Du, Haibao Wang, Huiguang He
IEEE Trans. Neural Networks Learn. Syst.1
2021 Multi-subject data augmentation for target subject semantic decoding with deep multi-view adversarial learning
Changde Du, Shengpei Wang, Haibao Wang, Huiguang He
Inf. Sci.2
2021 Dynamical Channel Pruning by Conditional Accuracy Change for Deep Neural Networks
abstract
Channel pruning is an effective technique that has been widely applied to deep neural network compression. However, many existing methods prune from a pretrained model, thus resulting in repetitious pruning and fine-tuning processes. In this article, we propose a dynamical channel pruning method, which prunes unimportant channels at the early stage of training. Rather than utilizing some indirect criteria (e.g., weight norm, absolute weight sum, and reconstruction error) to guide connection or channel pruning, we design criteria directly related to the final accuracy of a network to evaluate the importance of each channel. Specifically, a channelwise gate is designed to randomly enable or disable each channel so that the conditional accuracy changes (CACs) can be estimated under the condition of each channel disabled. Practically, we construct two effective and efficient criteria to dynamically estimate CAC at each iteration of training; thus, unimportant channels can be gradually pruned during the training process. Finally, extensive experiments on multiple data sets (i.e., ImageNet, CIFAR, and MNIST) with various networks (i.e., ResNet, VGG, and MLP) demonstrate that the proposed method effectively reduces the parameters and computations of baseline network while yielding the higher or competitive accuracy. Interestingly, if we Double the initial Channels and then Prune Half (DCPH) of them to baseline's counterpart, it can enjoy a remarkable performance improvement by shaping a more desirable structure.
Zhiqiang Chen 0002, Ting-Bing Xu, Changde Du, Cheng-Lin Liu 0001, Huiguang He
IEEE Trans. Neural Networks Learn. Syst.3
2020 Conditional Generative Neural Decoding with Structured CNN Feature Prediction
abstract
Decoding visual contents from human brain activity is a challenging task with great scientific value. Two main facts that hinder existing methods from producing satisfactory results are 1) typically small paired training data; 2) under-exploitation of the structural information underlying the data. In this paper, we present a novel conditional deep generative neural decoding approach with structured intermediate feature prediction. Specifically, our approach first decodes the brain activity to the multilayer intermediate features of a pretrained convolutional neural network (CNN) with a structured multi-output regression (SMR) model, and then inverts the decoded CNN features to the visual images with an introspective conditional generation (ICG) model. The proposed SMR model can simultaneously leverage the covariance structures underlying the brain activities, the CNN features and the prediction tasks to improve the decoding accuracy and interpretability. Further, our ICG model can 1) leverage abundant unpaired images to augment the training data; 2) self-evaluate the quality of its conditionally generated images; and 3) adversarially improve itself without extra discriminator. Experimental results show that our approach yields state-of-the-art visual reconstructions from brain activities.
Changde Du, Changying Du, Huiguang He
AAAI1
2020 Simultaneous Neural Spike Encoding and Decoding Based on Cross-modal Dual Deep Generative Model
abstract
Neural encoding and decoding of retinal ganglion cells (RGCs) have been attached great importance in the research work of brain-machine interfaces. Much effort has been invested to mimic RGC and get insight into RGC signals to reconstruct stimuli. However, there remain two challenges. On the one hand, complex nonlinear processes in retinal neural circuits hinder encoding models from enhancing their ability to fit the natural stimuli and modelling RGCs accurately. On the other hand, current research of the decoding process is separate from that of the encoding process, in which the liaison of mutual promotion between them is neglected. In order to alleviate the above problems, we propose a cross-modal dual deep generative model (CDDG) in this paper. CDDG treats the RGC spike signals and the stimuli as two modalities, which learns a shared latent representation for the concatenated modality and two modal-specific latent representations. Then, it imposes distribution consistency restriction on different latent space, cross-consistency and cycle-consistency constraints on the generated variables. Thus, our model ensures cross-modal generation from RGC spike signals to stimuli and vice versa. In our framework, the generation from stimuli to RGC spike signals is equivalent to neural encoding while the inverse process is equivalent to neural decoding. Hence, the proposed method integrates neural encoding and decoding and exploits the reciprocity between them. The experimental results demonstrate that our proposed method can achieve excellent encoding and decoding performance compared with the state-of-the-art methods on three salamander RGC spike datasets with natural stimuli.
Qiongyi Zhou, Changde Du, Haibao Wang, Jian K. Liu, Huiguang He
IJCNN2
2020 Semi-supervised cross-modal image generation with generative adversarial networks
Changde Du, Huiguang He
Pattern Recognit.2
2019 Doubly Semi-Supervised Multimodal Adversarial Learning for Classification, Generation and Retrieval
abstract
Learning over incomplete multi-modality data is a challenging problem with strong practical applications. Most existing multi-modal data imputation approaches have two limitations: (1) they are unable to accurately control the semantics of imputed modalities; and (2) without a shared low-dimensional latent space, they do not scale well with multiple modalities. To overcome the limitations, we propose a novel doubly semi-supervised multi-modal learning framework (DSML) with a modality-shared latent space and modality-specific generators, encoders and classifiers. We design novel softmax-based discriminators to train all modules adversarially. As a unified framework, DSML can be applied in multi-modal semi-supervised classification, missing modality imputation and fast cross-modality retrieval tasks simultaneously. Experiments on multiple datasets demonstrate its advantages.
Changde Du, Changying Du, Huiguang He
ICME1
2019 Learning "What" and "Where": An Interpretable Neural Encoding Model
abstract
Neural encoding modeling aims to reveal how brain processes perceived information by establishing a quantitative relationship between stimuli and evoked brain activities. In the field of visual neuroscience, many studies have been dedicated to building the neural encoding model for primary visual cortex and demonstrate that the population receptive field (pRF) models can be used to explain how neurons in primary visual cortex work. However, these models rely on either the inflexible prior assumptions imposed on the spatial characteristics of pRF or the clumsy parameter estimation methods which requiring too much manual adjustment. Suffering from these issues, current methods yield dissatisfactory performance on mimicking brain activity. In this paper, we address the problems under a novel "what and where" neural encoding framework. Basing on deep neural network (DNN) and the separability of the spatial ("where") and visual feature ("what") dimensions, the proposed method is not only powerful in extracting nonlinear features from images, but also rich in interpretability. Owing to two forms of regularization: sparsity and smoothness, receptive fields are estimated automatically for each voxel without prior assumptions on shape, which gets rid of the shortcomings of previous methods. Extensive empirical evaluations on publicly available fMRI dataset show that the proposed method has superior performance gains over several existing methods.
Haibao Wang, Changde Du, Huiguang He
IJCNN3
2019 Reconstructing Perceived Images From Human Brain Activities With Bayesian Deep Multiview Learning
abstract
Neural decoding, which aims to predict external visual stimuli information from evoked brain activities, plays an important role in understanding human visual system. Many existing methods are based on linear models, and most of them only focus on either the brain activity pattern classification or visual stimuli identification. Accurate reconstruction of the perceived images from the measured human brain activities still remains challenging. In this paper, we propose a novel deep generative multiview model for the accurate visual image reconstruction from the human brain activities measured by functional magnetic resonance imaging (fMRI). Specifically, we model the statistical relationships between the two views (i.e., the visual stimuli and the evoked fMRI) by using two view-specific generators with a shared latent space. On the one hand, we adopt a deep neural network architecture for visual image generation, which mimics the stages of human visual processing. On the other hand, we design a sparse Bayesian linear model for fMRI activity generation, which can effectively capture voxel correlations, suppress data noise, and avoid overfitting. Furthermore, we devise an efficient mean-field variational inference method to train the proposed model. The proposed method can accurately reconstruct visual images via Bayesian inference. In particular, we exploit a posterior regularization technique in the Bayesian inference to regularize the model posterior. The quantitative and qualitative evaluations conducted on multiple fMRI data sets demonstrate the proposed method can reconstruct visual images more accurately than the state of the art.
Changde Du, Changying Du, Huiguang He
IEEE Trans. Neural Networks Learn. Syst.1
2018 Improving Image Classification Performance with Automatically Hierarchical Label Clustering
abstract
Image classification is a common and foundational problem in computer vision. In traditional image classification, a category is assigned with single label, which is difficult for networks to learn better features. On the contrary, hierarchical labels can depict the structure of categories better, which helps network to learn more hierarchical features and improve the classification performance. Though many datasets contain images with multi-labels, the labels in these datasets usually lack of hierarchy. To overcome this problem, we propose a new method to improve image classification performance with Automatically Hierarchical Label Clustering (AHLC). Firstly, AHLC calculates the similarity between each pair of original categories by how easily they are misclassified with a pre-trained classifier. Secondly, AHLC obtains hierarchical labels by merging similar categories using hierarchical clustering. Finally, AHLC trains a new classifier with hierarchical labels to improve the original classification performance. We evaluate our method on MNIST and CIFAR-100 datasets and the results demonstrate the superiority of our method. The main contribution of this work is that we can simply improve an existing classification network by AHLC without extra information or heavy architecture redesign.
Zhiqiang Chen 0002, Changde Du, Huiguang He
ICPR2
2018 Multi-label Semantic Decoding from Human Brain Activity
abstract
It is meaningful to decode the semantic information from functional magnetic resonance imaging (fMRI) brain signals evoked by natural images. Semantic decoding can be viewed as a classification problem. Since a natural image may contain many semantic information of different objects, the single label classification model is not appropriate to cope with semantic decoding problem, which motivates the multi-label classification model. However, most multi-label models always treat each label equally. Actually, if dataset is associated with a large number of semantic labels, it will be difficult to get an accurate prediction of semantic label when the label appears with a low frequency in this dataset. So we should increase the relative importance degree to the labels that associate with little instances. In order to improve multi-label prediction performance, in this paper, we firstly propose a multinomial label distribution to estimate the importance degree of each associated label for an instance by using conditional probability, and then establish a deep neural network (DNN) based model which contains both multinomial label distribution and label co-occurrence information to realize the multi-label classification of semantic information in fMRI brain signals. Experiments on three fMRI recording datasets demonstrate that our approach performs better than the state-of-the-art methods on semantic information prediction.
Changde Du, Zhiqiang Chen 0002, Huiguang He
ICPR2
2018 Redundancy-resistant Generative Hashing for Image Retrieval
abstract
By optimizing probability distributions over discrete latent codes, Stochastic Generative Hashing (SGH) bypasses the critical and intractable binary constraints on hash codes. While encouraging results were reported, SGH still suffers from the deficient usage of latent codes, i.e., there often exist many uninformative latent dimensions in the code space, a disadvantage inherited from its auto-encoding variational framework. Motivated by the fact that code redundancy usually is severer when more complex decoder network is used, in this paper, we propose a constrained deep generative architecture to simplify the decoder for data reconstruction. Specifically, our new framework forces the latent hashing codes to not only reconstruct data through the generative network but also retain minimal squared L2 difference to the last real-valued network hidden layer. Furthermore, during posterior inference, we propose to regularize the standard auto-encoding objective with an additional term that explicitly accounts for the negative redundancy degree of latent code dimensions. We interpret such modifications as Bayesian posterior regularization and design an adversarial strategy to optimize the generative, the variational, and the redundancy-resistanting parameters. Empirical results show that our new method can significantly boost the quality of learned codes and achieve state-of-the-art performance for image retrieval.
Changying Du, Xingyu Xie, Changde Du, Hao Wang 0005
IJCAI3
2018 Multi-view Adversarially Learned Inference for Cross-domain Joint Distribution Matching
abstract
Many important data mining problems can be modeled as learning a (bidirectional) multidimensional mapping between two data domains. Based on the generative adversarial networks (GANs), particularly conditional ones, cross-domain joint distribution matching is an increasingly popular kind of methods addressing such problems. Though significant advances have been achieved, there are still two main disadvantages of existing models, i.e., the requirement of large amount of paired training samples and the notorious instability of training. In this paper, we propose a multi-view adversarially learned inference (ALI) model, termed as MALI, to address these issues. Unlike the common practice of learning direct domain mappings, our model relies on shared latent representations of both domains and can generate arbitrary number of paired faking samples, benefiting from which usually very few paired samples (together with sufficient unpaired ones) is enough for learning good mappings. Extending the vanilla ALI model, we design novel discriminators to judge the quality of generated samples (both paired and unpaired), and provide theoretical analysis of our new formulation. Experiments on image-to-image translation, image-to-attribute generation (multi-label classification), attribute-to-image generation tasks demonstrate that our semi-supervised learning framework yields significant performance improvements over existing ones. Results on cross-modality retrieval show that our latent space based method can achieve competitive similarity search performance in relative fast speed, compared to those methods that compute similarities in the high-dimensional data space.
Changying Du, Changde Du, Xingyu Xie, Chen Zhang 0003, Hao Wang 0005
KDD2
2018 Semi-supervised Deep Generative Modelling of Incomplete Multi-Modality Emotional Data
abstract
There are threefold challenges in emotion recognition. First, it is difficult to recognize human's emotional states only considering a single modality. Second, it is expensive to manually annotate the emotional data. Third, emotional data often suffers from missing modalities due to unforeseeable sensor malfunction or configuration issues. In this paper, we address all these problems under a novel multi-view deep generative framework. Specifically, we propose to model the statistical relationships of multi-modality emotional data using multiple modality-specific generative networks with a shared latent space. By imposing a Gaussian mixture assumption on the posterior approximation of the shared latent variables, our framework can learn the joint deep representation from multiple modalities and evaluate the importance of each modality simultaneously. To solve the labeled-data-scarcity problem, we extend our multi-view model to semi-supervised learning scenario by casting the semi-supervised classification problem as a specialized missing data imputation task. To address the missing-modality problem, we further extend our semi-supervised multi-view model to deal with incomplete data, where a missing view is treated as a latent variable and integrated out during inference. This way, the proposed overall framework can utilize all available (both labeled and unlabeled, as well as both complete and incomplete) data to improve its generalization ability. The experiments conducted on two real multi-modal emotion datasets demonstrated the superiority of our framework.
Changde Du, Changying Du, Hao Wang 0005, Jinpeng Li 0002, Wei-Long Zheng, Bao-Liang Lu, Huiguang He
ACM Multimedia1
2017 Nonlinear Maximum Margin Multi-View Learning with Adaptive Kernel
abstract
Existing multi-view learning methods based on kernel function either require the user to select and tune a single predefined kernel or have to compute and store many Gram matrices to perform multiple kernel learning. Apart from the huge consumption of manpower, computation and memory resources, most of these models seek point estimation of their parameters, and are prone to overfitting to small training data. This paper presents an adaptive kernel nonlinear max-margin multi-view learning model under the Bayesian framework. Specifically, we regularize the posterior of an efficient multi-view latent variable model by explicitly mapping the latent representations extracted from multiple data views to a random Fourier feature space where max-margin classification constraints are imposed. Assuming these random features are drawn from Dirichlet process Gaussian mixtures, we can adaptively learn shift-invariant kernels from data according to Bochners theorem. For inference, we employ the data augmentation idea for hinge loss, and design an efficient gradient-based MCMC sampler in the augmented space. Having no need to compute the Gram matrix, our algorithm scales linearly with the size of training set. Extensive experiments on real-world datasets demonstrate that our method has superior performance.
Jia He 0001, Changying Du, Changde Du, Fuzhen Zhuang, Qing He 0003, Guoping Long
IJCAI3
2017 Sharing deep generative representation for perceived image reconstruction from human brain activity
abstract
Decoding human brain activities via functional magnetic resonance imaging (fMRI) has gained increasing attention in recent years. While encouraging results have been reported in brain states classification tasks, reconstructing the details of human visual experience still remains difficult. Two main challenges that hinder the development of effective models are the perplexing fMRI measurement noise and the high dimensionality of limited data instances. Existing methods generally suffer from one or both of these issues and yield dissatisfactory results. In this paper, we tackle this problem by casting the reconstruction of visual stimulus as the Bayesian inference of missing view in a multiview latent variable model. Sharing a common latent representation, our joint generative model of external stimulus and brain response is not only “deep” in extracting nonlinear features from visual images, but also powerful in capturing correlations among voxel activities of fMRI recordings. The nonlinearity and deep structure endow our model with strong representation ability, while the correlations of voxel activities are critical for suppressing noise and improving prediction. We devise an efficient variational Bayesian method to infer the latent variables and the model parameters. To further improve the reconstruction accuracy, the latent representations of testing instances are enforced to be close to that of their neighbours from the training set via posterior regularization. Experiments on three fMRI recording datasets demonstrate that our approach can more accurately reconstruct visual stimuli.
Changde Du, Changying Du, Huiguang He
IJCNN1
2016 Bayesian Group Feature Selection for Support Vector Learning Machines
Changde Du, Changying Du, Shandian Zhe, A-Li Luo, Qing He 0003, Guoping Long
PAKDD (1)1
2016 Efficient Bayesian Maximum Margin Multiple Kernel Learning
Changying Du, Changde Du, Guoping Long, Xin Jin 0004, Yucheng Li 0002
ECML/PKDD (1)2
2016 Online Bayesian Multiple Kernel Bipartite Ranking
Changying Du, Changde Du, Guoping Long, Qing He 0003, Yucheng Li 0002
UAI2