Jiangrong Shen

dblp:208/3564 · DBLP profile ↗
← Back
36ranked-venue papers
9as first author
32since 2021 · last 2026
0000-0003-3683-3779ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 8 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Key-value pair-free continual learner via task-specific prompt-prototype
abstract
Continual learning aims to enable models to acquire new knowledge while retaining previously learned information. Prompt-based methods have shown remarkable performance in this domain; however, they typically rely on key-value pairing, which can introduce inter-task interference and hinder scalability. To overcome these limitations, we propose a novel approach employing task-specific Prompt-Prototype (ProP), thereby eliminating the need for key-value pairs. In our method, task-specific prompts facilitate more effective feature learning for the current task, while corresponding prototypes capture the representative features of the input. During inference, predictions are generated by binding each task-specific prompt with its associated prototype. Additionally, we introduce regularization constraints during prompt initialization to penalize excessively large values, thereby enhancing stability. Experiments on several widely used datasets demonstrate the effectiveness of the proposed method. In contrast to mainstream prompt-based approaches, our framework removes the dependency on key-value pairs, offering a fresh perspective for future continual learning research.
Haihua Luo, Xuming Ran, Zhengji Li, Huiyan Xue, Jiangrong Shen, Tommi Kärkkäinen, Qi Xu 0008, Fengyu Cong
Neural Networks6
2026 Generic-to-Personalised Learning for Multimodal Image Synthesis With Bidirectional Variational GAN
abstract
Multimodal image synthesis, which predicts target-modality images from source-modality images, has garnered considerable attention in the field of clinical diagnosis. Both unidirectional and bidirectional multimodal image synthesis methods have been explored in the medical domain, however, unidirectional models heavily rely on paired images, while current bidirectional models typically overlook local image details due to their unsupervised training patterns. In this work, we propose a Bidirectional Variational Generative Adversarial Network (BVGAN) for multimodal image synthesis, which achieves high-quality bidirectional translations between any two modalities using only a limited number paired images. Firstly, BVGAN's generator incorporates a variational structure (VAS) to regularise the latent space for noise reduction. This regularisation imposes smoothness to the latent space, enabling BVGAN to produce high-quality, noise-free images. Secondly, a novel generic-to-personalised (GTP) learning strategy is introduced to train BVGAN and reduce its reliance on a large sets of paired images. GTP initially leverages an unsupervised learning model to capture the global mapping between two modalities using unpaired images from generic patients. It then applies a supervised learning model to refine the mapping for individual patient, enhancing image details. Finally, the GTP learning strategy along with VAS enables BVGAN to achieve state-of-the-art performance on two multi-modality medical datasets: Brain CTMRI and BRATS.
Long Chen 0019, Xirui Dong, Jiangrong Shen, Lu Zhang 0053, Qi Xu 0008, Gang Pan 0001, Qiang Zhang 0008
IEEE Trans. Multim.3
2025 SpikingYOLOX: Improved YOLOX Object Detection with Fast Fourier Convolution and Spiking Neural Networks
abstract
In recent years, with the advancements in brain science, spiking neural networks (SNNs) have garnered significant attention. SNNs can generate spikes that mimic the function of neurons transmission in humans brain, thereby significantly reducing computational costs by the event-driven nature during training. While deep SNNs have shown impressive performance on classification tasks, they still face challenges in more complex tasks such as object detection. In this paper, we propose SpikingYOLOX, extending the structure of the original YOLOX by introducing signed spiking neurons and fast Fourier convolution (FFC). The designed ternary signed spiking neurons could generate three kinds of spikes to obtain more robust features in the deep layer of the backbone. Meanwhile, we integrate FFC with SNN modules to enhance object detection performance, because its global receptive field is beneficial to the object detection task. Extensive experiments demonstrate that the proposed SpikingYOLOX achieves state-of-the-art performance among other SNN-based object detection methods.
Wei Miao 0006, Jiangrong Shen, Qi Xu 0008, Timo Hämäläinen 0002, Yi Xu 0008, Fengyu Cong
AAAI2
2025 ALADE-SNN: Adaptive Logit Alignment in Dynamically Expandable Spiking Neural Networks for Class Incremental Learning
abstract
Inspired by the human brain's ability to adapt to new tasks without erasing prior knowledge, we develop spiking neural networks (SNNs) with dynamic structures for Class Incremental Learning (CIL). Our analytical experiments reveal that limited datasets introduce biases in logits distributions among tasks. Fixed features from frozen past-task extractors can cause overfitting and hinder the learning of new tasks. To address these challenges, we propose the ALADE-SNN framework, which includes adaptive logit alignment for balanced feature representation and OtoN suppression to manage weights mapping frozen old features to new classes during training, releasing them during fine-tuning. This approach dynamically adjusts the network architecture based on analytical observations, improving feature extraction and balancing performance between new and old tasks. Experiment results show that ALADE-SNN achieves an average incremental accuracy of 75.42 ± 0.74% on the CIFAR100-B0 dataset over 10 incremental steps. ALADE-SNN not only matches the performance of DNN-based methods but also surpasses state-of-the-art SNN-based continual learning algorithms. This advancement enhances continual learning in neuromorphic computing, offering a brain-inspired, energy-efficient solution for real-time data processing.
Wenyao Ni, Jiangrong Shen, Qi Xu 0008, Huajin Tang
AAAI2
2025 Improving the Sparse Structure Learning of Spiking Neural Networks from the View of Compression Efficiency
abstract
The human brain utilizes spikes for information transmission and dynamically reorganizes its network structure to boost energy efficiency and cognitive capabilities throughout its lifespan. Drawing inspiration from this spike-based computation, Spiking Neural Networks (SNNs) have been developed to construct event-driven models that emulate this efficiency. Despite these advances, deep SNNs continue to suffer from over-parameterization during training and inference, a stark contrast to the brain’s ability to self-organize. Furthermore, existing sparse SNNs are challenged by maintaining optimal pruning levels due to a static pruning ratio, resulting in either under or over-pruning. In this paper, we propose a novel two-stage dynamic structure learning approach for deep SNNs, aimed at maintaining effective sparse training from scratch while optimizing compression efficiency. The first stage evaluates the compressibility of existing sparse subnetworks within SNNs using the PQ index, which facilitates an adaptive determination of the rewiring ratio for synaptic connections based on data compression insights. In the second stage, this rewiring ratio critically informs the dynamic synaptic connection rewiring process, including both pruning and regrowth. This approach significantly improves the exploration of sparse structures training in deep SNNs, adapting sparsity dynamically from the point view of compression efficiency. Our experiments demonstrate that this sparse training approach not only aligns with the performance of current deep SNNs models but also significantly improves the efficiency of compressing sparse SNNs. Crucially, it preserves the advantages of initiating training with sparse models and offers a promising solution for implementing Edge AI on neuromorphic hardware.
Jiangrong Shen, Qi Xu 0008, Gang Pan 0001, Badong Chen
ICLR1
2025 Hybrid Spiking Vision Transformer for Object Detection with Event Cameras
abstract
Event-based object detection has attracted increasing attention for its high temporal resolution, wide dynamic range, and asynchronous address-event representation. Leveraging these advantages, spiking neural networks (SNNs) have emerged as a promising approach, offering low energy consumption and rich spatiotemporal dynamics. To further enhance the performance of event-based object detection, this study proposes a novel hybrid spike vision Transformer (HsVT) model. The HsVT model integrates a spatial feature extraction module to capture local and global features, and a temporal feature extraction module to model time dependencies and long-term patterns in event sequences. This combination enables HsVT to capture spatiotemporal features, improving its capability in handling complex event-based object detection tasks. To support research in this area, we developed the Fall Detection dataset as a benchmark for event-based object detection tasks. The Fall DVS detection dataset protects facial privacy and reduces memory usage thanks to its event-based representation. Experimental results demonstrate that HsVT outperforms existing SNN methods and achieves competitive performance compared to ANN-based models, with fewer parameters and lower energy consumption.
Qi Xu 0008, Jiangrong Shen, Biwu Chen, Huajin Tang, Gang Pan 0001
ICML3
2025 Self-cross Feature based Spiking Neural Networks for Efficient Few-shot Learning
abstract
Deep neural networks (DNNs) excel in computer vision tasks, especially, few-shot learning (FSL), which is increasingly important for generalizing from limited examples. However, DNNs are computationally expensive with scalability issues in real world. Spiking Neural Networks (SNNs), with their event-driven nature and low energy consumption, are particularly efficient in processing sparse and dynamic data, though they still encounter difficulties in capturing complex spatiotemporal features and performing accurate cross-class comparisons. To further enhance the performance and efficiency of SNNs in few-shot learning, we propose a few-shot learning framework based on SNNs, which combines a self-feature extractor module and a cross-feature contrastive module to refine feature representation and reduce power consumption. We apply the combination of temporal efficient training loss and InfoNCE loss to optimize the temporal dynamics of spike trains and enhance the discriminative power. Experimental results show that the proposed FSL-SNN significantly improves the classification performance on the neuromorphic dataset N-Omniglot, and also achieves competitive performance to ANNs on static datasets such as CUB and miniImageNet with low power consumption.
Qi Xu 0008, Junyang Zhu, Dongdong Zhou, Jiangrong Shen, Qiang Zhang 0008
ICML6
2025 Efficient ANN-SNN Conversion with Error Compensation Learning
abstract
Artificial neural networks (ANNs) have demonstrated outstanding performance in numerous tasks, but deployment in resource-constrained environments remains a challenge due to their high computational and memory requirements. Spiking neural networks (SNNs) operate through discrete spike events and offer superior energy efficiency, providing a bio-inspired alternative. However, current ANN-to-SNN conversion often results in significant accuracy loss and increased inference time due to conversion errors such as clipping, quantization, and uneven activation. This paper proposes a novel ANN-to-SNN conversion framework based on error compensation learning. We introduce a learnable threshold clipping function, dual-threshold neurons, and an optimized membrane potential initialization strategy to mitigate the conversion error. Together, these techniques address the clipping error through adaptive thresholds, dynamically reduce the quantization error through dual-threshold neurons, and minimize the non-uniformity error by effectively managing the membrane potential. Experimental results on CIFAR-10, CIFAR-100, ImageNet datasets show that our method achieves high-precision and ultra-low latency among existing conversion methods. Using only two time steps, our method significantly reduces the inference time while maintains competitive accuracy of 94.75% on CIFAR-10 dataset under ResNet-18 structure. This research promotes the practical application of SNNs on low-power hardware, making efficient real-time processing possible.
Chang Liu 0030, Jiangrong Shen, Xuming Ran, Mingkun Xu, Qi Xu 0008, Yi Xu 0008, Gang Pan 0001
ICML2
2025 Advanced SpikingYOLOX: Extending Spiking Neural Network on Object Detection with Spike-based Partial Self-Attention and 2D-Spiking Transformer
abstract
Brain-inspired Spiking Neural Networks (SNNs) have garnered significant attention due to their bio-plausibility and low power consumption advantages compared to Artificial Neural Networks (ANNs). However, the application of SNN in computer vision remains limited, primarily due to their inferior performance. In this work, we aim to bridge the performance gap between ANNs and SNNs in object detection by our Advanced SpikingYOLOX. The proposed approach extends the SpikingYOLOX with two key innovations: PSA-SNN and 2D-Spiking Transformer, both designed to enhance object detection performance. PSA-SNN extends spike-based self-attention by incorporating high-speed partial self-attention with an SNN-based 2D-Spiking Transformer in the deepest layer of the backbone, significantly improving feature extraction. The 2D-Spiking Transformer redefines the role of spiking neurons in Transformer sequences (Key, Query, Value), demonstrating that applying an additional spiking layer solely to the Value sequence yields the best performance while maintaining computational efficiency in spike-driven Transformers. We conduct extensive experiments on static images and the Advanced SpikingYOLOX achieves state-of-the-art performance among other SNN-based object detection methods. This work paves the way for more advanced SNN applications in object detection and broader computer vision tasks.
Wei Miao 0006, Jiangrong Shen, Hongming Xu 0002, Tommi Kärkkäinen, Qi Xu 0008, Yi Xu 0008, Fengyu Cong
ACM Multimedia2
2025 Spiking Neural Networks with Temporal Attention-Guided Adaptive Fusion for imbalanced Multi-modal Learning
Jiangrong Shen, Yulin Xie, Qi Xu 0008, Gang Pan 0001, Huajin Tang, Badong Chen
ACM Multimedia1
2025 Local-Global Coupling Spiking Graph Transformer for Brain Disorders Diagnosis from Two Perspectives
abstract
Brain disorders have been consistently associated with abnormalities in specific brain regions or neural circuits. Identifying key brain regional activities and functional connectivity patterns is essential for discovering more precise neurobiological biomarkers. However, previous studies have primarily emphasized alterations in functional connectivity while overlooking abnormal neuronal population activity within brain regions. To bridge this gap, we propose a novel Local-Global Coupling Spiking Graph Transformer (LGC-SGT) that jointly models both inter-regional connectivity differences and deviations in neuronal population firing rates within brain regions, enabling a dual-perspective neuropathological analysis. The global pathway leverages spike-based computation in LGC-SGT to model biologically plausible aberrant neural firing dynamics, while the local pathway adaptively captures abnormal graph-based representations of brain connectivity learned by local plasticity in the liquid state machine module. Furthermore, we design a shortcut-enhanced output strategy in LGC-SGT with the hybrid loss function to suppress outlier interference caused by inter-individual and inter-center variability, enabling a more robust decision boundary. Extensive experiments on three brain disorder datasets demonstrate that our model consistently outperforms state-of-the-art graph methods in brain disorder diagnosis. Moreover, it facilitates the extraction of interpretable neurobiological biomarkers by jointly analyzing regional neural activity and functional connectivity, offering a more comprehensive framework for brain disorder understanding and diagnosis.
Jiangrong Shen, Kaizhong Zheng, Liangjun Chen, Badong Chen
NeurIPS2
2025 Context gating in spiking neural networks: Achieving lifelong learning through integration of local and global plasticity
Jiangrong Shen, Wenyao Ni, Qi Xu 0008, Gang Pan 0001, Huajin Tang
Knowl. Based Syst.1
2025 Temporal spiking generative adversarial networks for heading direction decoding
Jiangrong Shen, Jian K. Liu, Qi Xu 0008, Gang Pan 0001, Xiaodong Chen 0005, Huajin Tang
Neural Networks1
2025 Robust Sensory Information Reconstruction and Classification With Augmented Spikes
abstract
Sensory information recognition is primarily processed through the ventral and dorsal visual pathways in the primate brain visual system, which exhibits layered feature representations bearing a strong resemblance to convolutional neural networks (CNNs), encompassing reconstruction and classification. However, existing studies often treat these pathways as distinct entities, focusing individually on pattern reconstruction or classification tasks, overlooking a key feature of biological neurons, the fundamental units for neural computation of visual sensory information. Addressing these limitations, we introduce a unified framework for sensory information recognition with augmented spikes. By integrating pattern reconstruction and classification within a single framework, our approach not only accurately reconstructs multimodal sensory information but also provides precise classification through definitive labeling. Experimental evaluations conducted on various datasets including video scenes, static images, dynamic auditory scenes, and functional magnetic resonance imaging (fMRI) brain activities demonstrate that our framework delivers state-of-the-art pattern reconstruction quality and classification accuracy. The proposed framework enhances the biological realism of multimodal pattern recognition models, offering insights into how the primate brain visual system effectively accomplishes the reconstruction and classification tasks through the integration of ventral and dorsal pathways.
Qi Xu 0008, Sibo Liu, Xuming Ran, Jiangrong Shen, Huajin Tang, Jian K. Liu, Gang Pan 0001, Qiang Zhang 0008
IEEE Trans. Neural Networks Learn. Syst.5
2024 Efficient Spiking Neural Networks with Sparse Selective Activation for Continual Learning
abstract
The next generation of machine intelligence requires the capability of continual learning to acquire new knowledge without forgetting the old one while conserving limited computing resources. Spiking neural networks (SNNs), compared to artificial neural networks (ANNs), have more characteristics that align with biological neurons, which may be helpful as a potential gating function for knowledge maintenance in neural networks. Inspired by the selective sparse activation principle of context gating in biological systems, we present a novel SNN model with selective activation to achieve continual learning. The trace-based K-Winner-Take-All (K-WTA) and variable threshold components are designed to form the sparsity in selective activation in spatial and temporal dimensions of spiking neurons, which promotes the subpopulation of neuron activation to perform specific tasks. As a result, continual learning can be maintained by routing different tasks via different populations of neurons in the network. The experiments are conducted on MNIST and CIFAR10 datasets under the class incremental setting. The results show that the proposed SNN model achieves competitive performance similar to and even surpasses the other regularization-based methods deployed under traditional ANNs.
Jiangrong Shen, Wenyao Ni, Qi Xu 0008, Huajin Tang
AAAI1
2024 Adaptive deep spiking neural network with global-local learning via balanced excitatory and inhibitory mechanism
abstract
The training method of Spiking Neural Networks (SNNs) is an essential problem, and how to integrate local and global learning is a worthy research interest. However, the current integration methods do not consider the network conditions suitable for local and global learning, and thus fail to balance their advantages. In this paper, we propose an Excitation-Inhibition Mechanism-assisted Hybrid Learning(EIHL) algorithm that adjusts the network connectivity by using the excitation-inhibition mechanism and then switches between local and global learning according to the network connectivity. The experimental results on CIFAR10/100 and DVS-CIFAR10 demonstrate that the EIHL not only has better accuracy performance than other methods but also has excellent sparsity advantage. Especially, the Spiking VGG11 is trained by EIHL, STBP, and STDP on DVS_CIFAR10, respectively. The accuracy of the Spiking VGG11 model on EIHL is 62.45%, which is 4.35% higher than STBP and 11.40% higher than STDP, and the sparsity is 18.74%, which is 18.74% higher than the other two methods. Moreover, the excitation-inhibition mechanism used in our method also offers a new perspective on the field of SNN learning.
Qi Xu 0008, Xuming Ran, Jiangrong Shen, Pan Lv, Qiang Zhang 0008, Gang Pan 0001
ICLR4
2024 The Balanced Multi-Modal Spiking Neural Networks with Online Loss Adjustment and Time Alignment
abstract
Optimizing multi-modal learning of SNNs has the advantages of energy efficiency and performance improvements. However, modality imbalance in multi-modal SNNs results in performance decline due to heterogeneity of modalities and temporal inconsistencies across different SNNs branches. In this paper, we propose the Balanced Multi-modal SNNs (BM-SNNs) model, equipped with a novel online loss adjustment (LA) algorithm and time alignment (TA) modules, ultimately achieving balanced training across multiple modalities. LA supervises the learning of uni-modal feature extractors by adding unimodal loss components without additional classifier. Moreover, the modulation factors enable the adaptive adjustment of unimodal learning rate. Furthermore, the proposed TA adopts the optimal timestep for different modalities to avoid information redundancy. Experimental results reveal that BM-SNNs model improves the performance of both multi-modal model and unimodal branches through adaptively exploiting intra-modal and cross-modal information.
Jianing Han, Jiangrong Shen, Qi Xu 0008, Jian K. Liu, Huajin Tang
ICME2
2024 Towards efficient deep spiking neural networks construction with spiking activity based pruning
abstract
The emergence of deep and large-scale spiking neural networks (SNNs) exhibiting high performance across diverse complex datasets has led to a need for compressing network models due to the presence of a significant number of redundant structural units, aiming to more effectively leverage their low-power consumption and biological interpretability advantages. Currently, most model compression techniques for SNNs are based on unstructured pruning of individual connections, which requires specific hardware support. Hence, we propose a structured pruning approach based on the activity levels of convolutional kernels named Spiking Channel Activity-based (SCA) network pruning framework. Inspired by synaptic plasticity mechanisms, our method dynamically adjusts the network’s structure by pruning and regenerating convolutional kernels during training, enhancing the model’s adaptation to the current target task. While maintaining model performance, this approach refines the network architecture, ultimately reducing computational load and accelerating the inference process. This indicates that structured dynamic sparse learning methods can better facilitate the application of deep SNNs in low-power and high-efficiency scenarios.
Qi Xu 0008, Jiangrong Shen, Hongming Xu 0002, Long Chen 0019, Gang Pan 0001
ICML3
2024 RSNN: Recurrent Spiking Neural Networks for Dynamic Spatial-Temporal Information Processing
Qi Xu 0008, Xuanye Fang, Jiangrong Shen, De Ma, Yi Xu 0008, Gang Pan 0001
ACM Multimedia4
2024 Reversing Structural Pattern Learning with Biologically Inspired Knowledge Distillation for Spiking Neural Networks
abstract
Spiking neural networks (SNNs) have superb characteristics in sensory information recognition tasks due to their biological plausibility. However, the performance of some current spiking-based models is limited by their structures which means either fully connected or too-deep structures bring too much redundancy. This redundancy from both connection and neurons is one of the key factors hindering the practical application of SNNs. Although Some pruning methods were proposed to tackle this problem, they normally ignored the fact the neural topology in the human brain could be adjusted dynamically. Inspired by this, this paper proposed an evolutionary-based structure construction method for constructing more reasonable SNNs. By integrating the knowledge distillation and connection pruning method, the synaptic connections in SNNs can be optimized dynamically to reach an optimal state. As a result, the structure of SNNs could not only absorb knowledge from the teacher model but also search for deep but sparse network topology. Experimental results on CIFAR100, Tiny-imagenet and DVS-Gesture show that the proposed structure learning method can get pretty well performance while reducing the connection redundancy. The proposed method explores a novel dynamical way for structure learning from scratch in SNNs which could build a bridge to close the gap between deep learning and bio-inspired neural dynamics.
Qi Xu 0008, Xuanye Fang, Jiangrong Shen, Qiang Zhang 0008, Gang Pan 0001
ACM Multimedia4
2024 SpikingMiniLM: energy-efficient spiking transformer for natural language understanding
Jiangrong Shen, Zeke Wang, Qinghai Guo, Rui Yan 0005, Gang Pan 0001, Huajin Tang
Sci. China Inf. Sci.2
2024 Hierarchical Spiking-Based Model for Efficient Image Classification With Enhanced Feature Extraction and Encoding
abstract
Thanks to their event-driven nature, spiking neural networks (SNNs) are surmised to be great computation-efficient models. The spiking neurons encode beneficial temporal facts and possess excessive anti-noise properties. However, the high-quality encoding of spatio-temporal complexity and also its training optimization of SNNs are restricted by means of the contemporary problem, this article proposes a novel hierarchical event-driven visual device to explore how information transmits and signifies in the retina the usage of biologically manageable mechanisms. This cognitive model is an augmented spiking-based framework consisting of the function learning capacity of convolutional neural networks (CNNs) with the cognition capability of SNNs. Furthermore, this visual device is modeled in a biological realism way with unsupervised learning rules and advanced spike firing rate encoding methods. We train and test them on some image datasets (Modified National Institute of Standards and Technology (MNIST), Canadian Institute for Advanced Research (CIFAR)10, and its noisy versions) to show that our mannequin can process greater vital data than present cognitive models. This article also proposes a novel quantization approach to make the proposed spiking-based model more efficient for neuromorphic hardware implementation. The outcomes show this joint CNN-SNN model can reap excessive focus accuracy and get more effective generalization ability.
Qi Xu 0008, Jiangrong Shen, Jian K. Liu, Huajin Tang, Gang Pan 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 ESL-SNNs: An Evolutionary Structure Learning Strategy for Spiking Neural Networks
abstract
Spiking neural networks (SNNs) have manifested remarkable advantages in power consumption and event-driven property during the inference process. To take full advantage of low power consumption and improve the efficiency of these models further, the pruning methods have been explored to find sparse SNNs without redundancy connections after training. However, parameter redundancy still hinders the efficiency of SNNs during training. In the human brain, the rewiring process of neural networks is highly dynamic, while synaptic connections maintain relatively sparse during brain development. Inspired by this, here we propose an efficient evolutionary structure learning (ESL) framework for SNNs, named ESL-SNNs, to implement the sparse SNN training from scratch. The pruning and regeneration of synaptic connections in SNNs evolve dynamically during learning, yet keep the structural sparsity at a certain level. As a result, the ESL-SNNs can search for optimal sparse connectivity by exploring all possible parameters across time. Our experiments show that the proposed ESL-SNNs framework is able to learn SNNs with sparse structures effectively while reducing the limited accuracy. The ESL-SNNs achieve merely 0.28% accuracy loss with 10% connection density on the DVS-Cifar10 dataset. Our work presents a brand-new approach for sparse training of SNNs from scratch with biologically plausible evolutionary mechanisms, closing the gap in the expressibility between sparse training and dense training. Hence, it has great potential for SNN lightweight training and inference with low power consumption and small memory usage.
Jiangrong Shen, Qi Xu 0008, Jian K. Liu, Yueming Wang 0001, Gang Pan 0001, Huajin Tang
AAAI1
2023 Constructing Deep Spiking Neural Networks from Artificial Neural Networks with Knowledge Distillation
abstract
Spiking neural networks (SNNs) are well-known as brain-inspired models with high computing efficiency, due to a key component that they utilize spikes as information units, close to the biological neural systems. Although spiking based models are energy efficient by taking advantage of discrete spike signals, their performance is limited by current network structures and their training methods. As discrete signals, typical SNNs cannot apply the gradient descent rules directly into parameter adjustment as artificial neural networks (ANNs). Aiming at this limitation, here we propose a novel method of constructing deep SNN models with knowledge distillation (KD) that uses ANN as the teacher model and SNN as the student model. Through the ANN-SNN joint training algorithm, the student SNN model can learn rich feature information from the teacher ANN model through the KD method, yet it avoids training SNN from scratch when communicating with non-differentiable spikes. Our method can not only build a more efficient deep spiking structure feasibly and reasonably but use few time steps to train the whole model compared to direct training or ANN to SNN methods. More importantly, it has a superb ability of noise immunity for various types of artificial noises and natural signals. The proposed novel method provides efficient ways to improve the performance of SNN through constructing deeper structures in a high-throughput fashion, with potential usage for light and efficient brain-inspired computing of practical scenarios.
Qi Xu 0008, Jiangrong Shen, Jian K. Liu, Huajin Tang, Gang Pan 0001
CVPR3
2023 Learnable Surrogate Gradient for Direct Training Spiking Neural Networks
abstract
Spiking neural networks (SNNs) have increasingly drawn massive research attention due to biological interpretability and efficient computation. Recent achievements are devoted to utilizing the surrogate gradient (SG) method to avoid the dilemma of non-differentiability of spiking activity to directly train SNNs by backpropagation. However, the fixed width of the SG leads to gradient vanishing and mismatch problems, thus limiting the performance of directly trained SNNs. In this work, we propose a novel perspective to unlock the width limitation of SG, called the learnable surrogate gradient (LSG) method. The LSG method modulates the width of SG according to the change of the distribution of the membrane potentials, which is identified to be related to the decay factors based on our theoretical analysis. Then we introduce the trainable decay factors to implement the LSG method, which can optimize the width of SG automatically during training to avoid the gradient vanishing and mismatch problems caused by the limited width of SG. We evaluate the proposed LSG method on both image and neuromorphic datasets. Experimental results show that the LSG method can effectively alleviate the blocking of gradient propagation caused by the limited width of SG when training deep SNNs directly. Meanwhile, the LSG method can help SNNs achieve competitive performance on both latency and accuracy.
Shuang Lian, Jiangrong Shen, Qianhui Liu, Rui Yan 0005, Huajin Tang
IJCAI2
2023 Bipolar Population Threshold Encoding for Audio Recognition with Deep Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) have been in-creasingly investigated for audio recognition due to the low power consumption on neuromorphic hardware by mimicking biological neural systems. Since the SNNs are learned from spikes, a critical step lies in the efficient neural encoding of real-valued sound signals to represent complex temporal patterns in speech and environmental sounds. In this paper, we propose a novel Bipolar Population Threshold (BPT) encoding model that effectively captures the trajectory information of time-series speech data by combining temporal and spatial dimensions. The bipolar encoding technique uses positive and negative neurons to capture the dynamic changes in the audio signal, while the threshold intervals allow for a sparse representation that focuses on encoding significant changes, resulting in an efficient and simplified recognition process. Extensively experimenting on three benchmark datasets including the TIDIGITS with speeches, RWCP with sounds, and MedleyDB with music, the numeric results show the superiority of the proposed method by consistently outperforming the state-of-the-art approaches while with fewer spikes, especially in capturing the complex spatio-temporal patterns of audio signals.
Xiaocui Lin, Jiangrong Shen, Huajin Tang
IJCNN2
2023 EICIL: Joint Excitatory Inhibitory Cycle Iteration Learning for Deep Spiking Neural Networks
abstract
Spiking neural networks (SNNs) have undergone continuous development and extensive study for decades, leading to increased biological plausibility and optimal energy efficiency. However, traditional training methods for deep SNNs have some limitations, as they rely on strategies such as pre-training and fine-tuning, indirect coding and reconstruction, and approximate gradients. These strategies lack a complete training model and require gradient approximation. To overcome these limitations, we propose a novel learning method named Joint Excitatory Inhibitory Cycle Iteration learning for Deep Spiking Neural Networks (EICIL) that integrates both excitatory and inhibitory behaviors inspired by the signal transmission of biological neurons.By organically embedding these two behavior patterns into one framework, the proposed EICIL significantly improves the bio-mimicry and adaptability of spiking neuron models, as well as expands the representation space of spiking neurons. Extensive experiments based on EICIL and traditional learning methods demonstrate that EICIL outperforms traditional methods on various datasets, such as CIFAR10 and CIFAR100, revealing the crucial role of the learning approach that integrates both behaviors during training.
Zihang Shao, Xuanye Fang, Chaoran Feng 0001, Jiangrong Shen, Qi Xu 0008
NeurIPS5
2023 Enhancing Adaptive History Reserving by Spiking Convolutional Block Attention Module in Recurrent Neural Networks
abstract
Spiking neural networks (SNNs) serve as one type of efficient model to process spatio-temporal patterns in time series, such as the Address-Event Representation data collected from Dynamic Vision Sensor (DVS). Although convolutional SNNs have achieved remarkable performance on these AER datasets, benefiting from the predominant spatial feature extraction ability of convolutional structure, they ignore temporal features related to sequential time points. In this paper, we develop a recurrent spiking neural network (RSNN) model embedded with an advanced spiking convolutional block attention module (SCBAM) component to combine both spatial and temporal features of spatio-temporal patterns. It invokes the history information in spatial and temporal channels adaptively through SCBAM, which brings the advantages of efficient memory calling and history redundancy elimination. The performance of our model was evaluated in DVS128-Gesture dataset and other time-series datasets. The experimental results show that the proposed SRNN-SCBAM model makes better use of the history information in spatial and temporal dimensions with less memory space, and achieves higher accuracy compared to other models.
Qi Xu 0008, Yuyuan Gao, Jiangrong Shen, Xuming Ran, Huajin Tang, Gang Pan 0001
NeurIPS3
2023 HybridSNN: Combining Bio-Machine Strengths by Boosting Adaptive Spiking Neural Networks
abstract
Spiking neural networks (SNNs), inspired by the neuronal network in the brain, provide biologically relevant and low-power consuming models for information processing. Existing studies either mimic the learning mechanism of brain neural networks as closely as possible, for example, the temporally local learning rule of spike-timing-dependent plasticity (STDP), or apply the gradient descent rule to optimize a multilayer SNN with fixed structure. However, the learning rule used in the former is local and how the real brain might do the global-scale credit assignment is still not clear, which means that those shallow SNNs are robust but deep SNNs are difficult to be trained globally and could not work so well. For the latter, the nondifferentiable problem caused by the discrete spike trains leads to inaccuracy in gradient computing and difficulties in effective deep SNNs. Hence, a hybrid solution is interesting to combine shallow SNNs with an appropriate machine learning (ML) technique not requiring the gradient computing, which is able to provide both energy-saving and high-performance advantages. In this article, we propose a HybridSNN, a deep and strong SNN composed of multiple simple SNNs, in which data-driven greedy optimization is used to build powerful classifiers, avoiding the derivative problem in gradient descent. During the training process, the output features (spikes) of selected weak classifiers are fed back to the pool for the subsequent weak SNN training and selection. This guarantees HybridSNN not only represents the linear combination of simple SNNs, as what regular AdaBoost algorithm generates, but also contains neuron connection information, thus closely resembling the neural networks of a brain. HybridSNN has the benefits of both low power consumption in weak units and overall data-driven optimizing strength. The network structure in HybridSNN is learned from training samples, which is more flexible and effective compared with existing fixed multilayer SNNs. Moreover, the topological tree of HybridSNN resembles the neural system in the brain, where pyramidal neurons receive thousands of synaptic input signals through their dendrites. Experimental results show that the proposed HybridSNN is highly competitive among the state-of-the-art SNNs.
Jiangrong Shen, Jian K. Liu, Yueming Wang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2022 Convolutional Neural Network Based Sleep Stage Classification with Class Imbalance
abstract
Accurate sleep stage classification is vital to assess sleep quality and diagnose sleep disorders. Numerous deep learning based models have been designed for accomplishing this labor automatically. However, the class imbalance problem existing in polysomnography (PSG) datasets has been barely investigated in previous studies, which is one of the most challenging obstacles for the real-world sleep staging application. To address this issue, this paper proposes novel methods with signal-driven and image-driven ways of noise addition to balance the imbalanced relationship in the training dataset samples. We evaluate the effectiveness of the proposed methods which are integrated into a convolutional neural network (CNN) based model. Experimental results evaluated on Sleep-EDF-V1, Sleep-EDF and CCSHS databases demonstrate that the proposed balancing approaches with specific tensity Gaussian white noise could enhance the overall or stage N1 recognition to some degree, especially the combination of two types of Data augmentation (DA) strategies shows the superiority of overall accuracy improvement.
Qi Xu 0008, Dongdong Zhou, Jian Wang 0112, Jiangrong Shen, Lauri Kettunen, Fengyu Cong
IJCNN4
2022 Robust Transcoding Sensory Information With Neural Spikes
abstract
Neural coding, including encoding and decoding, is one of the key problems in neuroscience for understanding how the brain uses neural signals to relate sensory perception and motor behaviors with neural systems. However, most of the existed studies only aim at dealing with the continuous signal of neural systems, while lacking a unique feature of biological neurons, termed spike, which is the fundamental information unit for neural computation as well as a building block for brain-machine interface. Aiming at these limitations, we propose a transcoding framework to encode multi-modal sensory information into neural spikes and then reconstruct stimuli from spikes. Sensory information can be compressed into 10% in terms of neural spikes, yet re-extract 100% of information by reconstruction. Our framework can not only feasibly and accurately reconstruct dynamical visual and auditory scenes, but also rebuild the stimulus patterns from functional magnetic resonance imaging (fMRI) brain activities. More importantly, it has a superb ability of noise immunity for various types of artificial noises and background signals. The proposed framework provides efficient ways to perform multimodal feature representation and reconstruction in a high-throughput fashion, with potential usage for efficient neuromorphic computing in a noisy environment.
Qi Xu 0008, Jiangrong Shen, Xuming Ran, Huajin Tang, Gang Pan 0001, Jian K. Liu
IEEE Trans. Neural Networks Learn. Syst.2
2021 Dynamic Spatiotemporal Pattern Recognition With Recurrent Spiking Neural Network
abstract
Our real-time actions in everyday life reflect a range of spatiotemporal dynamic brain activity patterns, the consequence of neuronal computation with spikes in the brain. Most existing models with spiking neurons aim at solving static pattern recognition tasks such as image classification. Compared with static features, spatiotemporal patterns are more complex due to their dynamics in both space and time domains. Spatiotemporal pattern recognition based on learning algorithms with spiking neurons therefore remains challenging. We propose an end-to-end recurrent spiking neural network model trained with an algorithm based on spike latency and temporal difference backpropagation. Our model is a cascaded network with three layers of spiking neurons where the input and output layers are the encoder and decoder, respectively. In the hidden layer, the recurrently connected neurons with transmission delays carry out high-dimensional computation to incorporate the spatiotemporal dynamics of the inputs. The test results based on the data sets of spiking activities of the retinal neurons show that the proposed framework can recognize dynamic spatiotemporal patterns much better than using spike counts. Moreover, for 3D trajectories of a human action data set, the proposed framework achieves a test accuracy of 83.6% on average. Rapid recognition is achieved through the learning methodology-based on spike latency and the decoding process using the first spike of the output neurons. Taken together, these results highlight a new model to extract information from activity patterns of neural computation in the brain and provide a novel approach for spike-based neuromorphic computing.
Jiangrong Shen, Jian K. Liu, Yueming Wang 0001
Neural Comput.1
2020 Recognizing Scoring in Basketball Game from AER Sequence by Spiking Neural Networks
abstract
The automatic score detection and recognition in basketball game has important application potentials, for examples, basketball technique analysis and 24 second control in the game. Although existing studies have been conducted on broadcast videos, most of them usually learned a machine learning algorithm on long videos recorded by traditional cameras. Address Event Representation (AER) sensor provides a possibility to deal with the problem by a human sensing manner. It represents the visual information as a series of spike-based events and records event sequences. Compared to traditional videos, AER events can fully utilize their addresses and timestamp information, forming precise spatio-temporal features with significantly less storage cost. More importantly, it issues spikes which can be naturally processed by human-style spiking neural networks (SNNs). In this paper, we propose to recognize scoring in basketball game from AER sequences. A new model is designed to extract dynamic features and discriminate different event streams using SNN. To handle the imbalance problem between positive and negative samples, we use an imbalanced Tempotron algorithm in our SNN model. Meanwhile, an AER sequence dataset of basketball games is collected. The experimental results demonstrate that our method achieves better performance compared with existing models.
Jiangrong Shen, Jian K. Liu, Yueming Wang 0001
IJCNN1
2020 Deep CovDenseSNN: A hierarchical event-driven dynamic framework with spiking neurons in noisy environment
Qi Xu 0008, Jianxin Peng, Jiangrong Shen, Huajin Tang, Gang Pan 0001
Neural Networks3
2018 Jointly Learning Network Connections and Link Weights in Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are considered to be biologically plausible and power-efficient on neuromorphic hardware. However, unlike the brain mechanisms, most existing SNN algorithms have fixed network topologies and connection relationships. This paper proposes a method to jointly learn network connections and link weights simultaneously. The connection structures are optimized by the spike-timing-dependent plasticity (STDP) rule with timing information, and the link weights are optimized by a supervised algorithm. The connection structures and the weights are learned alternately until a termination condition is satisfied. Experiments are carried out using four benchmark datasets. Our approach outperforms classical learning methods such as STDP, Tempotron, SpikeProp, and a state-of-the-art supervised algorithm. In addition, the learned structures effectively reduce the number of connections by about 24%, thus facilitate the computational efficiency of the network.
Jiangrong Shen, Yueming Wang 0001, Huajin Tang, Hang Yu 0010, Zhaohui Wu 0001, Gang Pan 0001
IJCAI2
2018 CSNN: An Augmented Spiking based Framework with Perceptron-Inception
abstract
Spiking Neural Networks (SNNs) represent and transmit information in spikes, which is considered more biologically realistic and computationally powerful than the traditional Artificial Neural Networks. The spiking neurons encode useful temporal information and possess highly anti-noise property. The feature extraction ability of typical SNNs is limited by shallow structures. This paper focuses on improving the feature extraction ability of SNNs in virtue of powerful feature extraction ability of Convolutional Neural Networks (CNNs). CNNs can extract abstract features resorting to the structure of the convolutional feature maps. We propose a CNN-SNN (CSNN) model to combine feature learning ability of CNNs with cognition ability of SNNs. The CSNN model learns the encoded spatial temporal representations of images in an event-driven way. We evaluate the CSNN model on the handwritten digits images dataset MNIST and its variational databases. In the presented experimental results, the proposed CSNN model is evaluated regarding learning capabilities, encoding mechanisms, robustness to noisy stimuli and its classification performance. The results show that CSNN behaves well compared to other cognitive models with significantly fewer neurons and training samples. Our work brings more biological realism into modern image classification models, with the hope that these models can inform how the brain performs this high-level vision task.
Qi Xu 0008, Hang Yu 0010, Jiangrong Shen, Huajin Tang, Gang Pan 0001
IJCAI4