EDBT 2026 Demo / reviewers in the wild / expert
Huajin Tang
dblp:18/434
· DBLP profile ↗
159ranked-venue papers
12as first author
84since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 142 · 12 first-author · 70 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 27 since 2021Systems, architecture and hardware · 8 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | S3Net: Spatiotemporally Separated Sparse Network for Neuromorphic Vision ProcessingabstractDynamic Vision Sensor (DVS) asynchronously records sparse events triggered by changes in pixel intensity, offering high temporal resolution and low latency. Existing frame-based methods process event data densely, violating its inherent sparsity and introducing computational redundancy. While asynchronous models preserve the event stream's native format, they often neglect spatial information, compromising their adaptability and efficiency. To address these limitations, we propose a Spatiotemporally Separated Sparse Network (S3Net) for efficient event stream encoding and learning. Specifically, we employ a learnable sparse encoding scheme to construct a voxel-structured representation that effectively extracts spatiotemporal relationships among event data. After that, we propose a dual-branch architecture to capture localized spatial dependencies and dynamic temporal patterns of event data. By explicitly decoupling spatial and temporal modeling, S3Net enables end-to-end asynchronous processing of variable-length event sequences, achieving both strong representational capacity and high computational efficiency. Experimental results on six event-based datasets demonstrate that S3Net achieves state-of-the-art performance. Compared to frame-based methods, it significantly reduces computational overhead and model complexity, while also outperforming existing asynchronous approaches in inference speed without compromising accuracy. Extensive experiments across six event-based datasets show that S3Net establishes new state-of-the-art performance. Our method reduces computational costs by 35% and model parameters by 27% compared to frame-based approaches, while delivering 1.58× faster inference than existing point-based methods at comparable accuracy levels. Rong Xiao 0001, Wanying Xu, Chenwei Tang, Shudong Huang, Huajin Tang |
AAAI | 6 |
| 2026 | MPD-SGR: Robust Spiking Neural Networks with Membrane Potential Distribution-Driven Surrogate Gradient RegularizationabstractThe surrogate gradient (SG) method has shown significant promise in enhancing the performance of deep spiking neural networks (SNNs), but it also introduces vulnerabilities to adversarial attacks. Although spike coding strategies and neural dynamics parameters have been extensively studied for their impact on robustness, the critical role of gradient magnitude, which reflects the model's sensitivity to input perturbations, remains underexplored. In SNNs, the gradient magnitude is primarily determined by the interaction between the membrane potential distribution (MPD) and the SG function. In this study, we investigate the relationship between the MPD and SG and their implications for improving the robustness of SNNs. Our theoretical analysis reveals that reducing the proportion of membrane potentials lying within the gradient-available range of the SG function effectively mitigates the sensitivity of SNNs to input perturbations. Building upon this insight, we propose a novel MPD-driven surrogate gradient regularization (MPD-SGR) method, which enhances robustness by explicitly regularizing the MPD based on its interaction with the SG function. Extensive experiments across multiple image classification benchmarks and diverse network architectures confirm that the MPD-SGR method significantly enhances the resilience of SNNs to adversarial perturbations and exhibits strong generalizability across diverse network configurations, SG functions, and spike encoding schemes. Runhao Jiang, Chengzhi Jiang, Rui Yan 0005, Huajin Tang |
AAAI | 4 |
| 2026 | S³: Spiking Neurons as an Isolating Segmenter for Brain Signal DecodingabstractRecent brain decoding studies have primarily emphasized the development of brain decoders, while largely neglecting the segmentation step. Existing methods typically adopt fixed-length segmentation, which might overlook subject- or task-level variability and disrupt temporal patterns within brain signals. To address this gap, we propose S3, which leverages spiking neurons as an isolating segmenter for brain signal decoding. S3 segments brain signals adaptively, considering subject- and task-level variability while preserving intrinsic temporal patterns of brain signals. It exploits the unique reset mechanism of spiking neurons to isolate previous irrelevant temporal patterns during the generation of each segmentation point. To optimize S3 for enhancing task performance in the absence of segmentation labels, we develop an optimization method where segmentation pseudo-labels are created with a stochastic-greedy algorithm to optimize them, while circumventing gradient blockade between S3 and task performance. Experiments on 10 downstream tasks across 13 public datasets demonstrate that S3 consistently outperforms existing methods, validating its effectiveness, generalizability and interpretability. Sha Zhao, Shi Gu, De Ma, Huajin Tang, Gang Pan 0001 |
AAAI | 7 |
| 2026 | SPEAK: Spiking Neurons as an Entropy-Aware Tokenizer for Large Language ModelsabstractMing Chen, Wenyao Li, Chao Liang, Shi Gu, Peng Lin, De Ma, Huajin Tang, Qian Zheng, Gang Pan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shi Gu, De Ma, Huajin Tang, Gang Pan 0001 |
ACL (1) | 7 |
| 2026 | MAS-SNN: Adaptive Sampling and Representation via Memory-Augmented Recurrent Spiking Neural Networks for Event-Based Object DetectionabstractEvent-based object detection aims to directly recognize and localize objects from asynchronous event streams, facilitating reliable perception in scenarios with fast motion and complex lighting conditions. Its performance largely depends on how temporal information is organized during event sampling. In such dynamic scenarios, fixed sampling strategies struggle to adapt to temporal variations in event streams, whereas recent learnable approaches often exhibit limitations in either their coupling with downstream detection objectives or their ability to fully exploit temporal cues. In this paper, we propose MAS-SNN, a memory-augmented Spiking Neural Network (SNN) framework for event-based object detection, which formulates event sampling as a task-driven, learnable process tightly integrated with representation learning. MAS-SNN introduces a Spiking Memory Attention Embedding (SMAE), composed of a Spiking Memory Buffer (SMB) and a Spiking Memory Attention (SMA) mechanism, to preserve and selectively reuse historical spiking states, thereby fully leveraging the intrinsic temporal dynamics of spiking neurons while maintaining the inherent sparsity and asynchronicity of event-driven vision. MAS-SNN achieves superior performance on N-Caltech 101 and Gen1, demonstrating the effectiveness of memory augmented spiking sampling for event-based object detection. Qingfeng Shi, Huaning Li, Huajin Tang, Rui Yan 0005 |
ICIC | 4 |
| 2026 | SpikeLoRA: Power-efficient low-rank adaptation based on spiking neural network
Qiugang Zhan, Fangyi Ding, Guisong Liu, Xiurui Xie, Huajin Tang |
Neurocomputing | 6 |
| 2026 | MSER: Multi-scale event representation model for enhanced spatio-temporal feature extraction
Wanying Xu, Rong Xiao 0001, Chenwei Tang, Jiancheng Lv 0001, Huajin Tang |
Neurocomputing | 7 |
| 2026 | SFedCA: Credit Assignment-Based Active Client Selection Strategy for Spiking Federated LearningabstractThe spiking federated learning (FL) is an emerging distributed learning paradigm that allows resource-constrained devices to train collaboratively at low power consumption without exchanging local data. It takes advantage of both the privacy computation property in FL and the energy efficiency in spiking neural networks (SNNs). However, existing spiking FL methods employ a random selection approach for client aggregation, assuming unbiased client participation. This neglect of statistical heterogeneity significantly affects the convergence and precision of the global model. In this work, we propose a credit assignment-based active client selection strategy for spiking federated learning, the SFedCA, to aggregate clients contributing to the global sample distribution balance judiciously. Specifically, the client credits are assigned by the firing intensity state before and after local model training, which reflects the difference in local data distribution from the global model. The comprehensive experiments are conducted on various non-identical and independent distribution (non-IID) scenarios. The experimental results demonstrate that the SFedCA outperforms the existing state-of-the-art spiking FL methods and requires fewer communication rounds. Qiugang Zhan, Jinbo Cao, Xiurui Xie, Huajin Tang, Malu Zhang, Shantian Yang, Guisong Liu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | EvHDR-GS: Event-guided HDR Video Reconstruction with 3D Gaussian SplattingabstractHigh Dynamic Range (HDR) video reconstruction seeks to accurately restore the extensive dynamic range present in real-world scenes and is widely employed in downstream applications. Existing methods typically operate on one or a small number of consecutive frames, which often leads to inconsistent brightness across the video due to their limited perspective on the video sequence. Moreover, supervised learning-based approaches are susceptible to data bias, resulting in reduced effectiveness when confronted with test inputs exhibiting a domain gap relative to the training data. To address these limitations, we present an event-guided HDR video reconstruction method through building 3D Gaussian Splatting (3DGS), to ensure consistent brightness imposed by 3D consistency. We introduce HDR 3D Gaussians capable of simultaneously representing HDR and low-dynamic-range (LDR) colors. Furthermore, we incorporate a learnable HDR-to-LDR transformation optimized by input event streams and LDR frames to eliminate the data bias. Experimental results on both synthetic and real-world datasets demonstrate that the proposed method achieves state-of-the-art performance. Zhan Lu, De Ma, Huajin Tang, Xudong Jiang 0001, Gang Pan 0001 |
AAAI | 4 |
| 2025 | EvHDR-NeRF: Building High Dynamic Range Radiance Fields with Single Exposure Images and EventsabstractWe present EvHDR-NeRF to recover a High Dynamic Range (HDR) radiance field from event streams and a set of Low Dynamic Range (LDR) views with single exposures. Using the EvHDR-NeRF, we can generate both novel HDR views and novel LDR views under different exposures. The key to our method is to model the new relationship between events streams and LDR images, which considers both the Camera Response Function (CRF) and exposure time. Based on this relationship, we categorize events into inter-frame events and intra-exposure. The former is utilized for building HDR radiance field and the latter is used to deblur potentially blurred images. Compared to existing methods, this method can effectively reconstruct the HDR radiance field even when the input images are degraded. Experimental results demonstrate that our method achieves state-of-the-art HDR reconstruction, providing a more adaptable and accurate solution for complex imaging applications. Zhanfeng Liao, De Ma, Huajin Tang, Gang Pan 0001 |
AAAI | 4 |
| 2025 | ALADE-SNN: Adaptive Logit Alignment in Dynamically Expandable Spiking Neural Networks for Class Incremental LearningabstractInspired by the human brain's ability to adapt to new tasks without erasing prior knowledge, we develop spiking neural networks (SNNs) with dynamic structures for Class Incremental Learning (CIL). Our analytical experiments reveal that limited datasets introduce biases in logits distributions among tasks. Fixed features from frozen past-task extractors can cause overfitting and hinder the learning of new tasks. To address these challenges, we propose the ALADE-SNN framework, which includes adaptive logit alignment for balanced feature representation and OtoN suppression to manage weights mapping frozen old features to new classes during training, releasing them during fine-tuning. This approach dynamically adjusts the network architecture based on analytical observations, improving feature extraction and balancing performance between new and old tasks. Experiment results show that ALADE-SNN achieves an average incremental accuracy of 75.42 ± 0.74% on the CIFAR100-B0 dataset over 10 incremental steps. ALADE-SNN not only matches the performance of DNN-based methods but also surpasses state-of-the-art SNN-based continual learning algorithms. This advancement enhances continual learning in neuromorphic computing, offering a brain-inspired, energy-efficient solution for real-time data processing. Wenyao Ni, Jiangrong Shen, Qi Xu 0008, Huajin Tang |
AAAI | 4 |
| 2025 | GRSN: Gated Recurrent Spiking Neurons for POMDPs and MARLabstractSpiking neural networks (SNNs) are widely applied in various fields due to their energy-efficient and fast-inference capabilities. Applying SNNs to reinforcement learning (RL) can significantly reduce the computational resource requirements for agents and improve the algorithm's performance under resource-constrained conditions. However, in current spiking reinforcement learning (SRL) algorithms, the simulation results of multiple time steps can only correspond to a single-step decision in RL. This is quite different from the real temporal dynamics in the brain and also fails to fully exploit the capacity of SNNs to process temporal data. In order to address this temporal mismatch issue and further take advantage of the inherent temporal dynamics of spiking neurons, we propose a novel temporal alignment paradigm (TAP) that leverages the single-step update of spiking neurons to accumulate historical state information in RL and introduces gated units to enhance the memory capacity of spiking neurons. Experimental results show that our method can solve partially observable Markov decision processes (POMDPs) and multi-agent cooperation problems with similar performance as recurrent neural networks (RNNs) but with about 50\% power consumption. Runhao Jiang, Rui Yan 0005, Huajin Tang |
AAAI | 5 |
| 2025 | EvSTVSR: Event Guided Space-Time Video Super-ResolutionabstractIn the domain of space-time video super-resolution, it is typically challenging to handle complex motions (including large and nonlinear motions) and varying illumination scenes due to the lack of inter-frame information. Leveraging the dense temporal information provided by event signals offers a promising solution. Traditional event-based methods typically rely on multiple images, using motion estimation and compensation, which can introduce errors. Accumulated errors from multiple frames often lead to artifacts and blurriness in the output. To mitigate these issues, we propose EvSTVSR, a method that uses fewer adjacent frames and integrates dense temporal information from events to guide alignment. Additionally, we introduce a coordinate-based feature fusion upsampling module to achieve spatial super-resolution. Experimental results demonstrate that our method not only outperforms existing RGB-based approaches but also excels in handling large motion scenarios. Haojie Yan, Zhan Lu, De Ma, Huajin Tang, Gang Pan 0001 |
AAAI | 5 |
| 2025 | Two-Stream Spiking Neural Network for Event-based Action RecognitionabstractSpiking neural networks (SNNs) are increasingly applied to event-based data generated by event cameras due to their asynchronous and sparse properties. Event cameras can inherently respond to the changes in the scene, which is a quite desirable property for action recognition tasks. However, existing works of SNNs for event-based action recognition are still limited. To capture the rich dynamics embedded in event streams, we propose the two-stream SNN that consists of spatial spiking stream and motion spiking stream to address event-based action recognition. To effectively build the two-stream SNN, we present a motion feature aggregation strategy and an attention-based two-stream fusion method. The motion feature aggregation strategy accumulates motion information and groups it into distinct channels for input into the SNN, which can alleviate the dilemma of information loss caused by compact representation. The attention-based two-stream fusion method can fuse the spatial and motion features effectively using the channel-wise attention mechanism, which helps our network to achieve better integration of two-stream information. Extensive experimental results on three event-based action recognition datasets show our proposed two-stream SNN achieves competitive performance with much fewer trainable parameters, which demonstrates the effectiveness of our work in event-based action recognition tasks. Shuang Lian, Qianhui Liu, Ziling Wang, Zhibin Zuo, Rui Yan 0005, Huajin Tang |
ICASSP | 8 |
| 2025 | Adaptive Gradient-Based Timesurface for Event-based DetectionabstractThe advantages of high temporal resolution and high dynamic range provided by event cameras are particularly suitable for moving object detection, especially in scenarios with motion blur and extreme lighting conditions. Current popular methods predominantly focus on designing powerful network architectures to extract event features, often neglecting the rationality of event representation design which has been proven to impact significantly on downstream tasks. In particular, current event representations typically rely on fixed hyperparameters, without considering variations in relative motion speed, a key factor in motion-rich scenes captured by event cameras. To tackle this challenge, we propose a gradient-based scaled Timesurface (STS) motivated by the observation of the relationship between motion speeds and the gradient strength, which adaptively rescales the decay factor at different spatial positions. Additionally, we propose a dataset called RotateDigit, which is the first event dataset featuring clear motion level annotations to our best knowledge. Proposed STS method is verified using Spiking Neural Network (SNNs) due to the sharing asynchronous and sparse properties with event camera. Experimental results on RotateDigit and Gen1 show the performance improvement achieved by STS, which validates the rationality and effectiveness of our work. Ziling Wang, Shuang Lian, Rui Yan 0005, Huajin Tang |
ICASSP | 5 |
| 2025 | E-NeMF: Event-based Neural Motion Field for Novel Space-time View Synthesis of Dynamic Scenes
Haojie Yan, De Ma, Huajin Tang, Gang Pan 0001 |
ICCV | 5 |
| 2025 | CAM: Asynchronous GPU-Initiated, CPU-Managed SSD Management for Batching Storage AccessabstractWith the wide adoption of GPU and the explosion in data volumes, existing accelerator-centric systems require massive storage access. They adopt high-performance storage devices like NVMe SSDs to scale up single-node systems cost-effectively and leverage the CPU to manage these SSDs. However, they suffer from performance bottlenecks because of the high CPU OS kernel overhead and the CPU memory intermediated data transfer. To address this issue, GPU-initiated and GPU-managed SSD management is proposed to allow the GPU to fully manipulate SSDs: 1) direct data transfer from SSD to GPU memory (data plane) and 2) GPU-managed SSD control (control plane). This can potentially enable these GPU systems to fully leverage the SSD bandwidth. However, we still identify two severe issues. First, the GPU-management SSD control leads to low GPU Streaming Multiprocessor utilization. Second, it leads to the serial execution of SSD accesses with GPU computation, which slows down the overall computing task. To this end, we propose CAM, the first asynchronous GPU-initialized, CPU-managed SSD management for batching storage access. It 1) offloads the SSD control plane from GPU to CPU, thus maximizing GPU streaming multiprocessor utilization, and 2) adopts asynchronous user-friendly APIs that allow programmers to easily overlap GPU computation and SSD I/O operations while keeping a synchronous programming experience. As such, CAM enables us to achieve the best of two worlds: high performance and high programmability. The experimental results show that CAM can perform GNN model training, mergesort, and GEMM up to$\mathbf{1.84}\times, \mathbf{1.5}\times$, and$\mathbf{1.84}\times \mathbf{faster}$compared to the existing state-of-the-art GPU systems, while keeping high programmability. Ziyu Song, Jie Zhang 0081, Jie Sun 0017, Mo Sun 0001, Zihan Yang 0004, Xuzheng Chen, Fei Wu 0001, Huajin Tang, Zeke Wang |
ICDE | 9 |
| 2025 | Hybrid Spiking Vision Transformer for Object Detection with Event CamerasabstractEvent-based object detection has attracted increasing attention for its high temporal resolution, wide dynamic range, and asynchronous address-event representation. Leveraging these advantages, spiking neural networks (SNNs) have emerged as a promising approach, offering low energy consumption and rich spatiotemporal dynamics. To further enhance the performance of event-based object detection, this study proposes a novel hybrid spike vision Transformer (HsVT) model. The HsVT model integrates a spatial feature extraction module to capture local and global features, and a temporal feature extraction module to model time dependencies and long-term patterns in event sequences. This combination enables HsVT to capture spatiotemporal features, improving its capability in handling complex event-based object detection tasks. To support research in this area, we developed the Fall Detection dataset as a benchmark for event-based object detection tasks. The Fall DVS detection dataset protects facial privacy and reduces memory usage thanks to its event-based representation. Experimental results demonstrate that HsVT outperforms existing SNN methods and achieves competitive performance compared to ANN-based models, with fewer parameters and lower energy consumption. Qi Xu 0008, Jiangrong Shen, Biwu Chen, Huajin Tang, Gang Pan 0001 |
ICML | 5 |
| 2025 | Training High Performance Spiking Neural Network by Temporal Model CalibrationabstractSpiking Neural Networks (SNNs) are considered promising energy-efficient models due to their dynamic capability to process spatial-temporal spike information. Existing work has demonstrated that SNNs exhibit temporal heterogeneity, which leads to diverse outputs of SNNs at different time steps and has the potential to enhance their performance. Although SNNs obtained by direct training methods achieve state-of-the-art performance, current methods introduce limited temporal heterogeneity through the dynamics of spiking neurons or network structures. They lack the improvement of temporal heterogeneity through the lens of the gradient. In this paper, we first conclude that the diversity of the temporal logit gradients in current methods is limited. This leads to insufficient temporal heterogeneity and results in temporally miscalibrated SNNs with degraded performance. Based on the above analysis, we propose a Temporal Model Calibration (TMC) method, which can be seen as a logit gradient rescaling mechanism across time steps. Experimental results show that our method can improve the temporal logit gradient diversity and generate temporally calibrated SNNs with enhanced performance. In particular, our method achieves state-of-the-art accuracy on ImageNet, DVSCIFAR10, and N-Caltech101. Codes are available at https://github.com/zju-bmi-lab/TMC. Changping Wang, De Ma, Huajin Tang, Gang Pan 0001 |
ICML | 4 |
| 2025 | Brain-Inspired Spatial Continuous State Encoding for Efficient Spiking-Based NavigationabstractSpiking neural networks (SNNs) show great potential in mapless navigation tasks due to their low power consumption, but the continuous representation of spatial information poses a challenge to SNN training. Neuroscience findings reveal that spatial cognition cells encode spatial information through population spike patterns. Inspired by this, we propose a navigation method based on SNNs, leveraging spatial cognition cells, which include grid cells (GCs), head direction cells (HDCs), and boundary vector cells (BVCs). Our method integrates spike-based information to achieve precise navigation goal encoding and egocentric environment perception, significantly improving SNN navigation capabilities in complex environments. Simulation and real-world experiments demonstrate that our method achieves significant improvements in navigation success rate and energy efficiency, showcasing superior adaptability across environments. Our work provides a novel approach to developing efficient brain-inspired navigation systems. Qingao Chai, Jiashuo Wang, Runhao Jiang, Huajin Tang |
ICRA | 6 |
| 2025 | HSRL: A Hierarchical Control System Based on Spiking Deep Reinforcement Learning for Robot NavigationabstractReinforcement Learning (RL) has shown promise in robotic navigation tasks, yet applying it to real-world environments remains challenging due to dynamic complexities and the need for dynamically feasible actions. We propose a hierarchical control framework based on Spiking Deep Reinforcement Learning (SDRL) for robust robot navigation in real environments. Our approach utilizes a two-layer architecture: a high-level decision layer powered by a Spiking GRU network for handling partially observable environments, and a low-level executive layer employing Continuous Attractor Neural Networks (CANNs) to ensure precise and continuous actions. This hierarchical structure allows real-time decisionmaking that respects the physical constraints of the robot. Experimental results show that our method adapts effectively to new environments without fine-tuning and surpasses existing methods in performance. We also explore the implementation on the Darwin3 chip, paving the way for biologically inspired motion control in future robotic applications. Shibo Zhou, Chaohui Lin, Qingao Chai, Rui Yan 0005, De Ma, Gang Pan 0001, Huajin Tang |
ICRA | 8 |
| 2025 | EDyGS: Event Enhanced Dynamic 3D Radiance Fields from Blurry Monocular VideoabstractThe task of generating novel views in dynamic scenes plays a critical role in the 3D vision domain. Neural Radiance Fields (NeRFs) and 3D Gaussian Splatting (3DGS) have shown great promise in this domain but struggle with motion blur, which often arises in real-world scenarios due to camera or object motion. Existing methods address camera motion blur but fall short in dynamic scenes, where the coupling of camera and object motion complicates multi-view consistency and temporal coherence. In this work, we propose EDyGS, a model designed to reconstruct sharp novel views from event streams and monocular videos of dynamic scenes with motion blur. Our approach introduces a motion-mask 3D Gaussian model that assigns each Gaussian an additional attribute to distinguish between static and dynamic regions. By leveraging this motion mask field, we separate and optimize the static and dynamic regions independently. A progressive learning strategy is adopted, where static regions are reconstructed by jointly optimizing camera poses and learnable 3D Gaussians, while dynamic regions are modeled using an implicit deformation field alongside learnable 3D Gaussians. We conduct both quantitative and qualitative experiments on synthetic and real-world data. Experimental results demonstrate that EDyGS effectively handles blurry inputs in dynamic scenes. Mengxu Lu, De Ma, Huajin Tang, Gang Pan 0001 |
IJCAI | 5 |
| 2025 | Deep Spiking Neural Network with Adaptive Temporal Feature Fusion for Energy-efficient Event-based Person Re-IdentificationabstractVideo surveillance is widely used in various public spaces, and person re-identification (ReId) based on video frames has become a research hotspot in the field of computer vision. However, most video-based methods struggle with issues such as poor lighting or motion blur, which result in blurry textures and hinder the extraction of discriminative features for specific identities. Additionally, frame-based methods may introduce privacy concerns and suffer from redundant storage between frames. In contrast, bio-inspired sensors—such as event cameras—have asynchronous properties and higher temporal resolution. They record events when changes in light intensity occur, offering a better response to lighting variations in specific scenes. Spiking Neural Networks (SNNs), which transmit information through sparse spikes, are naturally suited to process such asynchronous and sparse event-stream inputs. In this work, we explore the possibility of using only SNN for person ReId based on event-stream data generated by event cameras for the first time. While SNNs for person ReId suffer from issues such as blurry matching and poor feature distinction in spike-based feature matching, we propose an adaptive temporal feature representation (ATFR) method based on membrane potential to address these issues. We evaluate our method on two event-based datasets (Event-ReId and Event-PRID-2011). Compared to existing methods, our model achieves competitive performance while improving energy efficiency by at least 6 times. Qingfeng Shi, Jiaqiang Jiang, Huajin Tang, Rui Yan 0005 |
IJCNN | 3 |
| 2025 | Spiking Neural Networks with Temporal Attention-Guided Adaptive Fusion for imbalanced Multi-modal Learning
Jiangrong Shen, Yulin Xie, Qi Xu 0008, Gang Pan 0001, Huajin Tang, Badong Chen |
ACM Multimedia | 5 |
| 2025 | The roles of internal dynamics and proprioceptive feedback in motor cortex during movement execution
Hongru Jiang, Xiangdong Bu, Zhiyan Zheng, Huajin Tang, Xiaochuan Pan, Yao Chen 0004 |
Neurocomputing | 4 |
| 2025 | NeuroSimWorm: A multisensory framework for modeling and simulating neural circuits of Caenorhabditis elegansabstractBiological behaviors emerge from the dynamic interplay among the inner neurodynamic system, embodied mechanical structure, and external environmental inputs. Nonetheless, existing approaches simply consider the static brain model that cannot fully exploit the potential of continuous interaction and feedback from the body and the environment. To address these problems, we introduce NeuroSimWorm , a multisensory closed-loop neural circuit simulation approach of the widely studied organism, Caenorhabditis elegans ( C. elegans ). The full closed-loop simulation platform integrates four key subcomponents including Environment, Neural Computing, Biomechanical Model, and Visualization modules. We initially define multiple sensory environments for chemical, mechanical, and thermal stimuli. Subsequently, we construct four types of neural circuits including Locomotion, Chemosensation , Thermosensation, and Mechanosensation . Fitness functions based on multisensory optimization are proposed, enabling the virtual nematode to achieve complex intelligent behaviors such as autonomous locomotion, foraging, thermosensory regulation, and tactile avoidance. Thus, NeuroSimWorm presents a feasible way to understand the mechanism of biological intelligence by modeling the connectome and simulating behaviors in the physical surroundings. Mengxiao Zhang 0003, Lijun Kang, Gang Pan 0001, Huajin Tang |
Neurocomputing | 7 |
| 2025 | Context gating in spiking neural networks: Achieving lifelong learning through integration of local and global plasticity
Jiangrong Shen, Wenyao Ni, Qi Xu 0008, Gang Pan 0001, Huajin Tang |
Knowl. Based Syst. | 5 |
| 2025 | Temporal spiking generative adversarial networks for heading direction decoding
Jiangrong Shen, Jian K. Liu, Qi Xu 0008, Gang Pan 0001, Xiaodong Chen 0005, Huajin Tang |
Neural Networks | 8 |
| 2025 | STSF: Spiking Time Sparse Feedback Learning for Spiking Neural NetworksabstractSpiking neural networks (SNNs) are biologically plausible models known for their computational efficiency. A significant advantage of SNNs lies in the binary information transmission through spike trains, eliminating the need for multiplication operations. However, due to the spatio-temporal nature of SNNs, direct application of traditional backpropagation (BP) training still results in significant computational costs. Meanwhile, learning methods based on unsupervised synaptic plasticity provide an alternative for training SNNs but often yield suboptimal results. Thus, efficiently training high-accuracy SNNs remains a challenge. In this article, we propose a highly efficient and biologically plausible spiking time sparse feedback (STSF) learning method. This algorithm modifies synaptic weights by incorporating a neuromodulator for global supervised learning using sparse direct feedback alignment (DFA) and local homeostasis learning with vanilla spike-timing-dependent plasticity (STDP). Such neuromorphic global-local learning focuses on instantaneous synaptic activity, enabling independent and simultaneous optimization of each network layer, thereby improving biological plausibility, enhancing parallelism, and reducing storage overhead. Incorporating sparse fixed random feedback connections for global error modulation, which uses selection operations instead of multiplication operations, further improves computational efficiency. Experimental results demonstrate that the proposed algorithm markedly reduces the computational cost with significantly higher accuracy comparable to current state-of-the-art algorithms across a wide range of classification tasks. Our implementation codes are available at https://github.com/hppeace/STSF. Rong Xiao 0001, Chenwei Tang, Shudong Huang, Jiancheng Lv 0001, Huajin Tang |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Spiking Neural Network for Ultralow-Latency and High-Accurate Object DetectionabstractSpiking Neural Networks (SNNs) have attracted significant attention for their energy-efficient and brain-inspired event-driven properties. Recent advancements, notably Spiking-YOLO, have enabled SNNs to undertake advanced object detection tasks. Nevertheless, these methods often suffer from increased latency and diminished detection accuracy, rendering them less suitable for latency-sensitive mobile platforms. Additionally, the conversion of artificial neural networks (ANNs) to SNNs frequently compromises the integrity of the ANNs' structure, resulting in poor feature representation and heightened conversion errors. To address the issues of high latency and low detection accuracy, we introduce two solutions: timestep compression and spike-time-dependent integrated (STDI) coding. Timestep compression effectively reduces the number of timesteps required in the ANN-to-SNN conversion by condensing information. The STDI coding employs a time-varying threshold to augment information capacity. Furthermore, we have developed an SNN-based spatial pyramid pooling (SPP) structure, optimized to preserve the network's structural efficacy during conversion. Utilizing these approaches, we present the ultralow latency and highly accurate object detection model, SUHD. SUHD exhibits exceptional performance on challenging datasets like PASCAL VOC and MS COCO, achieving a remarkable reduction of approximately 750 times in timesteps and a 30% enhancement in mean average precision (mAP) compared to Spiking-YOLO on MS COCO. To the best of our knowledge, SUHD is currently the deepest spike-based object detection model, achieving ultralow timesteps for lossless conversion. Jinye Qu, Tielin Zhang, Huajin Tang, Hong Qiao |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Toward High-Accuracy and Low-Latency Spiking Neural Networks With Two-Stage OptimizationabstractSpiking neural networks (SNNs) operating with asynchronous discrete events show higher energy efficiency with sparse computation. A popular approach for implementing deep SNNs is artificial neural network (ANN)-SNN conversion combining both efficient training of ANNs and efficient inference of SNNs. However, the accuracy loss is usually nonnegligible, especially under few time steps, which restricts the applications of SNN on latency-sensitive edge devices greatly. In this article, we first identify that such performance degradation stems from the misrepresentation of the negative or overflow residual membrane potential in SNNs. Inspired by this, we decompose the conversion error into three parts: quantization error, clipping error, and residual membrane potential representation error. With such insights, we propose a two-stage conversion algorithm to minimize those errors, respectively. In addition, we show that each stage achieves significant performance gains in a complementary manner. By evaluating on challenging datasets including CIFAR- 10, CIFAR- 100, and ImageNet, the proposed method demonstrates the state-of-the-art performance in terms of accuracy, latency, and energy preservation. Furthermore, our method is evaluated using a more challenging object detection task, revealing notable gains in regression performance under ultralow latency, when compared with existing spike-based detection algorithms. Codes will be available at: https://github.com/Windere/snn-cvt-dual-phase. Yuhao Zhang 0007, Shuang Lian, Xiaoxin Cui, Rui Yan 0005, Huajin Tang |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Robust Sensory Information Reconstruction and Classification With Augmented SpikesabstractSensory information recognition is primarily processed through the ventral and dorsal visual pathways in the primate brain visual system, which exhibits layered feature representations bearing a strong resemblance to convolutional neural networks (CNNs), encompassing reconstruction and classification. However, existing studies often treat these pathways as distinct entities, focusing individually on pattern reconstruction or classification tasks, overlooking a key feature of biological neurons, the fundamental units for neural computation of visual sensory information. Addressing these limitations, we introduce a unified framework for sensory information recognition with augmented spikes. By integrating pattern reconstruction and classification within a single framework, our approach not only accurately reconstructs multimodal sensory information but also provides precise classification through definitive labeling. Experimental evaluations conducted on various datasets including video scenes, static images, dynamic auditory scenes, and functional magnetic resonance imaging (fMRI) brain activities demonstrate that our framework delivers state-of-the-art pattern reconstruction quality and classification accuracy. The proposed framework enhances the biological realism of multimodal pattern recognition models, offering insights into how the primate brain visual system effectively accomplishes the reconstruction and classification tasks through the integration of ventral and dorsal pathways. Qi Xu 0008, Sibo Liu, Xuming Ran, Jiangrong Shen, Huajin Tang, Jian K. Liu, Gang Pan 0001, Qiang Zhang 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Successive POI Recommendation via Brain-Inspired Spatiotemporal Aware RepresentationabstractExisting approaches usually perform spatiotemporal representation in the spatial and temporal dimensions, respectively, which isolates the spatial and temporal natures of the target and leads to sub-optimal embeddings. Neuroscience research has shown that the mammalian brain entorhinal-hippocampal system provides efficient graph representations for general knowledge. Moreover, entorhinal grid cells present concise spatial representations, while hippocampal place cells represent perception conjunctions effectively. Thus, the entorhinal-hippocampal system provides a novel angle for spatiotemporal representation, which inspires us to propose the SpatioTemporal aware Embedding framework (STE) and apply it to POIs (STEP). STEP considers two types of POI-specific representations: sequential representation and spatiotemporal conjunctive representation, learned using sparse unlabeled data based on the proposed graph-building policies. Notably, STEP jointly represents the spatiotemporal natures of POIs using both observations and contextual information from integrated spatiotemporal dimensions by constructing a spatiotemporal context graph. Furthermore, we introduce a successive POI recommendation method using STEP, which achieves state-of-the-art performance on two benchmarks. In addition, we demonstrate the excellent performance of the STE representation approach in other spatiotemporal representation-centered tasks through a case study of the traffic flow prediction problem. Therefore, this work provides a novel solution to spatiotemporal representation and paves a new way for spatiotemporal modeling-related tasks. Gehua Ma, Rui Yan 0005, Huajin Tang |
AAAI | 5 |
| 2024 | Efficient Spiking Neural Networks with Sparse Selective Activation for Continual LearningabstractThe next generation of machine intelligence requires the capability of continual learning to acquire new knowledge without forgetting the old one while conserving limited computing resources. Spiking neural networks (SNNs), compared to artificial neural networks (ANNs), have more characteristics that align with biological neurons, which may be helpful as a potential gating function for knowledge maintenance in neural networks. Inspired by the selective sparse activation principle of context gating in biological systems, we present a novel SNN model with selective activation to achieve continual learning. The trace-based K-Winner-Take-All (K-WTA) and variable threshold components are designed to form the sparsity in selective activation in spatial and temporal dimensions of spiking neurons, which promotes the subpopulation of neuron activation to perform specific tasks. As a result, continual learning can be maintained by routing different tasks via different populations of neurons in the network. The experiments are conducted on MNIST and CIFAR10 datasets under the class incremental setting. The results show that the proposed SNN model achieves competitive performance similar to and even surpasses the other regularization-based methods deployed under traditional ANNs. Jiangrong Shen, Wenyao Ni, Qi Xu 0008, Huajin Tang |
AAAI | 4 |
| 2024 | EAS-SNN: End-to-End Adaptive Sampling and Representation for Event-Based Detection with Recurrent Spiking Neural Networks
Ziling Wang, Huaning Li, Runhao Jiang, De Ma, Huajin Tang |
ECCV (60) | 7 |
| 2024 | The Balanced Multi-Modal Spiking Neural Networks with Online Loss Adjustment and Time AlignmentabstractOptimizing multi-modal learning of SNNs has the advantages of energy efficiency and performance improvements. However, modality imbalance in multi-modal SNNs results in performance decline due to heterogeneity of modalities and temporal inconsistencies across different SNNs branches. In this paper, we propose the Balanced Multi-modal SNNs (BM-SNNs) model, equipped with a novel online loss adjustment (LA) algorithm and time alignment (TA) modules, ultimately achieving balanced training across multiple modalities. LA supervises the learning of uni-modal feature extractors by adding unimodal loss components without additional classifier. Moreover, the modulation factors enable the adaptive adjustment of unimodal learning rate. Furthermore, the proposed TA adopts the optimal timestep for different modalities to avoid information redundancy. Experimental results reveal that BM-SNNs model improves the performance of both multi-modal model and unimodal branches through adaptively exploiting intra-modal and cross-modal information. Jianing Han, Jiangrong Shen, Qi Xu 0008, Jian K. Liu, Huajin Tang |
ICME | 5 |
| 2024 | An Event-based Feature Representation Method for Event Stream Classification using Deep Spiking Neural NetworksabstractEvent streams output by event cameras have low data redundancy and retain accurate temporal information in the form of Address Event Representation (AER) which are different from the outputs of traditional frame-based cameras. Spiking Neural Networks (SNNs) are considered an effective tool for handling event-based scenarios due to their inherent temporal properties. However, most existing SNNs directly convert an event stream to several static frames with temporal relationships by channel-wise accumulation of events. These serial frames lose temporal characteristics in some extent and potentially affect the capacity of the SNNs to learn and recognize event streams. In this work, we proposed a novel event-based feature descriptor called time interval correlation time-surface (TICTS) for SNNs and introduced this event-based feature extraction method into SNNs. The TICTS can capture more precise temporal correlation from event streams, thereby facilitating SNNs to learn temporal information more effectively. The experimental results show that SNNs with the proposed TICTS exhibit superior performance and increased stability across datasets with varying speeds. In addition, shallow network using TICTS can achieve competitive accuracy compared to deep networks, underscoring the effectiveness of TICTS in reducing the redundant network size of SNNs. Limei Liang, Runhao Jiang, Huajin Tang, Rui Yan 0005 |
IJCNN | 3 |
| 2024 | Adaptive Multi-Level Firing for Direct Training Deep Spiking Neural NetworksabstractSpiking neural networks (SNNs) with bio-inspired spatio-temporal dynamics, have increasingly manifested their superiority in energy efficiency. However, the non-differential spiking activity hinders the implementation of the efficient error backpropagation in training SNNs. The existing solution circumvents this problem by a relaxation on the gradient calculation using a continuous function, which is referred to as surrogate gradient learning. Nevertheless, such a solution leads to a prominent gradient mismatch problem due to the low precision of spikes, which limits the performance of directly trained SNNs on deeper architectures. To tackle this issue, we propose the adaptive multilevel firing (AMLF) method incorporating the spiking dormant-suppressed residual network (spiking DS-ResNet) to get well-behaved SNNs. The AMLF method can not only adaptively implement incremental expression ability of spiking neurons to alleviate gradient mismatch issue, but also enable more efficient gradient propagation during training by enlarging the gradient-available area. We conduct a series of experiments on static image datasets and neuromorphic datasets. Results show that the proposed AMLF method can help SNNs achieve competitive performance in terms of both latency and accuracy. Haosong Qi, Shuang Lian, Huajin Tang |
IJCNN | 4 |
| 2024 | Multi-scale Harmonic Mean Time Surfaces for Event-based Object ClassificationabstractEvent cameras have attracted increasing attention in the field of computer vision due to their advantages in terms of high temporal resolution, high dynamic range and low power consumption. However, the output of event cameras is a sparse and discrete event stream, with each individual event in the event streams containing only little information. Therefore, extracting more effective features from the available information in the event streams is currently a major challenge in event-based object classification. In this paper, we propose a novel event-based feature representation to encode the spatiotemporal features of event streams. The Harmonic Mean Time Surfaces (HMTS) representation makes efficient use of information about past events, which enhances the spatiotemporal relationship between events, thus establishing robust representation. Furthermore, we propose a multi-scale feature extraction model that can construct broader event-region correlations. In the classification stage, the extracted spatiotemporal features are classified using a spiking neural network with an event-driven Tempotron rule. To demonstrate the effectiveness of the proposed method, we conduct a series of experiments on the N-MNIST, MNIST-DVS, DVS128 Gesture and DailyAction-DVS datasets. Experimental results show that our model achieves superior classification performance and exhibits a high level of robustness to noise. Huajin Tang, Rui Yan 0005 |
IJCNN | 3 |
| 2024 | CASRL: Collision Avoidance with Spiking Reinforcement Learning Among Dynamic, Decision-Making AgentsabstractDeveloping an efficient collision avoidance policy with Spiking Reinforcement Learning for dynamic, decision-making agents remains challenging. Moreover, the implementation of energy-efficient collision avoidance is important for mobile robots that operate with limited on-board computing resources. Most existing energy-efficient methods via spiking reinforcement learning are predominately concerned with the navigational capabilities of a single agent, and are unable to handle a large, and possibly varying number of agents. To overcome these limitations, we propose a model called collision avoidance with spiking reinforcement learning (CASRL), based on proximal policy optimization algorithms. This proposed model consists of an actor with spiking neural networks (SNNs) and a critic with deep neural networks (DNNs). Our spiking reinforcement learning algorithm is advantageous to handle an arbitrary number of other agents by virtue of a spiking-gated transformer (SpikeGTr) architecture and an accumulate-to-fire (ATF) module. Extensive experimental results demonstrate that CASRL obtains a competitive success rate of navigation and exhibits higher time-efficiency for navigation in crowded scenarios compared to traditional DNN-based methods. Ka-Wa Yip, Mengwen Yuan, Rui Yan 0005, Huajin Tang |
IROS | 7 |
| 2024 | Event-ID: Intrinsic Decomposition Using an Event CameraabstractReconstructing 3D scenes from multi-view images is challenging, especially under extreme scenarios. We propose Event-ID, an event-based intrinsic decomposition framework that leverages events and images for stable decomposition under extreme scenarios. Our method is based on two observations: event cameras maintain good imaging quality under blurry or poorly exposed scenarios, and event signals from different viewpoints exhibit similarity in diffuse regions while varying in specular regions. We establish an event-based reflectance model and introduce an event-based warping method to extract specular clues. Our two-stage framework constructs a radiance field and decomposes the scene into normal, material, and lighting. Experimental results demonstrate superior performance compared to state-of-the-art methods. Our project can be found at https://zehaoc.github.io/EventID.github.io/ Zhan Lu, De Ma, Huajin Tang, Xudong Jiang 0001, Gang Pan 0001 |
ACM Multimedia | 4 |
| 2024 | FEEL-SNN: Robust Spiking Neural Networks with Frequency Encoding and Evolutionary Leak FactorabstractCurrently, researchers think that the inherent robustness of spiking neural networks (SNNs) stems from their biologically plausible spiking neurons, and are dedicated to developing more bio-inspired models to defend attacks. However, most work relies solely on experimental analysis and lacks theoretical support, and the direct-encoding method and fixed membrane potential leak factor they used in spiking neurons are simplified simulations of those in the biological nervous system, which makes it difficult to ensure generalizability across all datasets and networks. Contrarily, the biological nervous system can stay reliable even in a highly complex noise environment, one of the reasons is selective visual attention and non-fixed membrane potential leaks in biological neurons. This biological finding has inspired us to design a highly robust SNN model that closely mimics the biological nervous system. In our study, we first present a unified theoretical framework for SNN robustness constraint, which suggests that improving the encoding method and evolution of the membrane potential leak factor in spiking neurons can improve SNN robustness. Subsequently, we propose a robust SNN (FEEL-SNN) with Frequency Encoding (FE) and Evolutionary Leak factor (EL) to defend against different noises, mimicking the selective visual attention mechanism and non-fixed leak observed in biological systems. Experimental results confirm the efficacy of both our FE, EL, and FEEL methods, either in isolation or in conjunction with established robust enhancement algorithms, for enhancing the robustness of SNNs. Mengting Xu, De Ma, Huajin Tang, Gang Pan 0001 |
NeurIPS | 3 |
| 2024 | SpikingMiniLM: energy-efficient spiking transformer for natural language understanding
Jiangrong Shen, Zeke Wang, Qinghai Guo, Rui Yan 0005, Gang Pan 0001, Huajin Tang |
Sci. China Inf. Sci. | 7 |
| 2024 | Reliable object tracking by multimodal hybrid feature extraction and transformer-based fusion
Hongze Sun, Rui Liu 0047, Wuque Cai, Jun Wang 0031, Huajin Tang, Yan Cui 0004, Dezhong Yao 0001, Daqing Guo |
Neural Networks | 6 |
| 2024 | Trainable Spiking-YOLO for low-latency and high-performance object detection
Mengwen Yuan, Huixiang Liu, Gang Pan 0001, Huajin Tang |
Neural Networks | 6 |
| 2024 | Enhancing SNN-based spatio-temporal learning: A benchmark dataset and Cross-Modality Attention model
Shibo Zhou, Mengwen Yuan, Runhao Jiang, Rui Yan 0005, Gang Pan 0001, Huajin Tang |
Neural Networks | 7 |
| 2024 | Brain-Inspired Computing: A Systematic Survey and Future TrendsabstractBrain-inspired computing (BIC) is an emerging research field that aims to build fundamental theories, models, hardware architectures, and application systems toward more general artificial intelligence (AI) by learning from the information processing mechanisms or structures/functions of biological nervous systems. It is regarded as one of the most promising research directions for future intelligent computing in the post-Moore era. In the past few years, various new schemes in this field have sprung up to explore more general AI. These works are quite divergent in the aspects of modeling/algorithm, software tool, hardware platform, and benchmark data since BIC is an interdisciplinary field that consists of many different domains, including computational neuroscience, AI, computer science, statistical physics, material science, and microelectronics. This situation greatly impedes researchers from obtaining a clear picture and getting started in the right way. Hence, there is an urgent requirement to do a comprehensive survey in this field to help correctly recognize and analyze such bewildering methodologies. What are the key issues to enhance the development of BIC? What roles do the current mainstream technologies play in the general framework of BIC? Which techniques are truly useful in real-world applications? These questions largely remain open. To address the above issues, in this survey, we first clarify the biggest challenge of BIC: how can AI models benefit from the recent advancements in computational neuroscience? With this challenge in mind, we will focus on discussing the concept of BIC and summarize four components of BIC infrastructure development: 1) modeling/algorithm; 2) hardware platform; 3) software tool; and 4) benchmark data. For each component, we will summarize its recent progress, main challenges to resolve, and future trends. Based on these studies, we present a general framework for the real-world applications of BIC systems, which is promising to benefit both AI and brain science. Finally, we claim that it is extremely important to build a research ecology to promote prosperity continuously in this field. Guoqi Li 0002, Lei Deng 0003, Huajin Tang, Gang Pan 0001, Yonghong Tian 0001, Kaushik Roy 0001, Wolfgang Maass 0001 |
Proc. IEEE | 3 |
| 2024 | Corrections to "Brain-Inspired Computing: A Systematic Survey and Future Trends"abstractPresents corrections to the paper, (Corrections to “Brain-Inspired Computing: A Systematic Survey and Future Trends”). Guoqi Li 0002, Lei Deng 0003, Huajin Tang, Gang Pan 0001, Yonghong Tian 0001, Kaushik Roy 0001, Wolfgang Maass 0001 |
Proc. IEEE | 3 |
| 2024 | Spiking Transfer Learning From RGB Image to Neuromorphic Event StreamabstractRecent advances in bio-inspired vision with event cameras and associated spiking neural networks (SNNs) have provided promising solutions for low-power consumption neuromorphic tasks. However, as the research of event cameras is still in its infancy, the amount of labeled event stream data is much less than that of the RGB database. The traditional method of converting static images into event streams by simulation to increase the sample size cannot simulate the characteristics of event cameras such as high temporal resolution. To take advantage of both the rich knowledge in labeled RGB images and the features of the event camera, we propose a transfer learning method from the RGB to the event domain in this paper. Specifically, we first introduce a transfer learning framework named R2ETL (RGB to Event Transfer Learning), including a novel encoding alignment module and a feature alignment module. Then, we introduce the temporal centered kernel alignment (TCKA) loss function to improve the efficiency of transfer learning. It aligns the distribution of temporal neuron states by adding a temporal learning constraint. Finally, we theoretically analyze the amount of data required by the deep neuromorphic model to prove the necessity of our method. Numerous experiments demonstrate that our proposed framework outperforms the state-of-the-art SNN and artificial neural network (ANN) models trained on event streams, including N-MNIST, CIFAR10-DVS and N-Caltech101. This indicates that the R2ETL framework is able to leverage the knowledge of labeled RGB images to help the training of SNN on event streams. Qiugang Zhan, Guisong Liu, Xiurui Xie, Malu Zhang, Huajin Tang |
IEEE Trans. Image Process. | 6 |
| 2024 | Attention-Based Deep Spiking Neural Networks for Temporal Credit Assignment ProblemsabstractThe temporal credit assignment (TCA) problem, which aims to detect predictive features hidden in distracting background streams, remains a core challenge in biological and machine learning. Aggregate-label (AL) learning is proposed by researchers to resolve this problem by matching spikes with delayed feedback. However, the existing AL learning algorithms only consider the information of a single timestep, which is inconsistent with the real situation. Meanwhile, there is no quantitative evaluation method for TCA problems. To address these limitations, we propose a novel attention-based TCA (ATCA) algorithm and a minimum editing distance (MED)-based quantitative evaluation method. Specifically, we define a loss function based on the attention mechanism to deal with the information contained within the spike clusters and use MED to evaluate the similarity between the spike train and the target clue flow. Experimental results on musical instrument recognition (MedleyDB), speech recognition (TIDIGITS), and gesture recognition (DVS128-Gesture) show that the ATCA algorithm can reach the state-of-the-art (SOTA) level compared with other AL learning algorithms. Rui Yan 0005, Huajin Tang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Event-Driven Spiking Learning Algorithm Using Aggregated LabelsabstractTraditional spiking learning algorithm aims to train neurons to spike at a specific time or on a particular frequency, which requires precise time and frequency labels in the training process. While in reality, usually only aggregated labels of sequential patterns are provided. The aggregate-label (AL) learning is proposed to discover these predictive features in distracting background streams only by aggregated spikes. It has achieved much success recently, but it is still computationally intensive and has limited use in deep networks. To address these issues, we propose an event-driven spiking aggregate learning algorithm (SALA) in this article. Specifically, to reduce the computational complexity, we improve the conventional spike-threshold-surface (STS) calculation in AL learning by analytical calculating voltage peak values in spiking neurons. Then we derive the algorithm to multilayers by event-driven strategy using aggregated spikes. We conduct comprehensive experiments on various tasks including temporal clue recognition, segmented and continuous speech recognition, and neuromorphic image classification. The experimental results demonstrate that the new STS method improves the efficiency of AL learning significantly, and the proposed algorithm outperforms the conventional spiking algorithm in various temporal clue recognition tasks. Xiurui Xie, Yansong Chua, Guisong Liu, Malu Zhang, Guangchun Luo, Huajin Tang |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Effective Active Learning Method for Spiking Neural NetworksabstractA large quantity of labeled data is required to train high-performance deep spiking neural networks (SNNs), but obtaining labeled data is expensive. Active learning is proposed to reduce the quantity of labeled data required by deep learning models. However, conventional active learning methods in SNNs are not as effective as that in conventional artificial neural networks (ANNs) because of the difference in feature representation and information transmission. To address this issue, we propose an effective active learning method for a deep SNN model in this article. Specifically, a loss prediction module ActiveLossNet is proposed to extract features and select valuable samples for deep SNNs. Then, we derive the corresponding active learning algorithm for deep SNN models. Comprehensive experiments are conducted on CIFAR-10, MNIST, Fashion-MNIST, and SVHN on different SNN frameworks, including seven-layer CIFARNet and 20-layer ResNet-18. The comparison results demonstrate that the proposed active learning algorithm outperforms random selection and conventional ANN active learning methods. In addition, our method converges faster than conventional active learning methods. Xiurui Xie, Guisong Liu, Qiugang Zhan, Huajin Tang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Hierarchical Spiking-Based Model for Efficient Image Classification With Enhanced Feature Extraction and EncodingabstractThanks to their event-driven nature, spiking neural networks (SNNs) are surmised to be great computation-efficient models. The spiking neurons encode beneficial temporal facts and possess excessive anti-noise properties. However, the high-quality encoding of spatio-temporal complexity and also its training optimization of SNNs are restricted by means of the contemporary problem, this article proposes a novel hierarchical event-driven visual device to explore how information transmits and signifies in the retina the usage of biologically manageable mechanisms. This cognitive model is an augmented spiking-based framework consisting of the function learning capacity of convolutional neural networks (CNNs) with the cognition capability of SNNs. Furthermore, this visual device is modeled in a biological realism way with unsupervised learning rules and advanced spike firing rate encoding methods. We train and test them on some image datasets (Modified National Institute of Standards and Technology (MNIST), Canadian Institute for Advanced Research (CIFAR)10, and its noisy versions) to show that our mannequin can process greater vital data than present cognitive models. This article also proposes a novel quantization approach to make the proposed spiking-based model more efficient for neuromorphic hardware implementation. The outcomes show this joint CNN-SNN model can reap excessive focus accuracy and get more effective generalization ability. Qi Xu 0008, Jiangrong Shen, Jian K. Liu, Huajin Tang, Gang Pan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | ESL-SNNs: An Evolutionary Structure Learning Strategy for Spiking Neural NetworksabstractSpiking neural networks (SNNs) have manifested remarkable advantages in power consumption and event-driven property during the inference process. To take full advantage of low power consumption and improve the efficiency of these models further, the pruning methods have been explored to find sparse SNNs without redundancy connections after training. However, parameter redundancy still hinders the efficiency of SNNs during training. In the human brain, the rewiring process of neural networks is highly dynamic, while synaptic connections maintain relatively sparse during brain development. Inspired by this, here we propose an efficient evolutionary structure learning (ESL) framework for SNNs, named ESL-SNNs, to implement the sparse SNN training from scratch. The pruning and regeneration of synaptic connections in SNNs evolve dynamically during learning, yet keep the structural sparsity at a certain level. As a result, the ESL-SNNs can search for optimal sparse connectivity by exploring all possible parameters across time. Our experiments show that the proposed ESL-SNNs framework is able to learn SNNs with sparse structures effectively while reducing the limited accuracy. The ESL-SNNs achieve merely 0.28% accuracy loss with 10% connection density on the DVS-Cifar10 dataset. Our work presents a brand-new approach for sparse training of SNNs from scratch with biologically plausible evolutionary mechanisms, closing the gap in the expressibility between sparse training and dense training. Hence, it has great potential for SNN lightweight training and inference with low power consumption and small memory usage. Jiangrong Shen, Qi Xu 0008, Jian K. Liu, Yueming Wang 0001, Gang Pan 0001, Huajin Tang |
AAAI | 6 |
| 2023 | Constructing Deep Spiking Neural Networks from Artificial Neural Networks with Knowledge DistillationabstractSpiking neural networks (SNNs) are well-known as brain-inspired models with high computing efficiency, due to a key component that they utilize spikes as information units, close to the biological neural systems. Although spiking based models are energy efficient by taking advantage of discrete spike signals, their performance is limited by current network structures and their training methods. As discrete signals, typical SNNs cannot apply the gradient descent rules directly into parameter adjustment as artificial neural networks (ANNs). Aiming at this limitation, here we propose a novel method of constructing deep SNN models with knowledge distillation (KD) that uses ANN as the teacher model and SNN as the student model. Through the ANN-SNN joint training algorithm, the student SNN model can learn rich feature information from the teacher ANN model through the KD method, yet it avoids training SNN from scratch when communicating with non-differentiable spikes. Our method can not only build a more efficient deep spiking structure feasibly and reasonably but use few time steps to train the whole model compared to direct training or ANN to SNN methods. More importantly, it has a superb ability of noise immunity for various types of artificial noises and natural signals. The proposed novel method provides efficient ways to improve the performance of SNN through constructing deeper structures in a high-throughput fashion, with potential usage for light and efficient brain-inspired computing of practical scenarios. Qi Xu 0008, Jiangrong Shen, Jian K. Liu, Huajin Tang, Gang Pan 0001 |
CVPR | 5 |
| 2023 | Adaptive Smoothing Gradient Learning for Spiking Neural NetworksabstractSpiking neural networks (SNNs) with biologically inspired spatio-temporal dynamics demonstrate superior energy efficiency on neuromorphic architectures. Error backpropagation in SNNs is prohibited by the all-or-none nature of spikes. The existing solution circumvents this problem by a relaxation on the gradient calculation using a continuous function with a constant relaxation de- gree, so-called surrogate gradient learning. Nevertheless, such a solution introduces additional smoothing error on spike firing which leads to the gradients being estimated inaccurately. Thus, how to adaptively adjust the relaxation degree and eliminate smoothing error progressively is crucial. Here, we propose a methodology such that training a prototype neural network will evolve into training an SNN gradually by fusing the learnable relaxation degree into the network with random spike noise. In this way, the network learns adaptively the accurate gradients of loss landscape in SNNs. The theoretical analysis further shows optimization on such a noisy network could be evolved into optimization on the embedded SNN with shared weights progressively. Moreover, The experiments on static images, dynamic event streams, speech, and instrumental sounds show the proposed method achieves state-of-the-art performance across all the datasets with remarkable robustness on different relaxation degrees. Runhao Jiang, Shuang Lian, Rui Yan 0005, Huajin Tang |
ICML | 5 |
| 2023 | CMCI: A Robust Multimodal Fusion Method for Spiking Neural Networks
Runhao Jiang, Jianing Han, Yingying Xue, Huajin Tang |
ICONIP (3) | 5 |
| 2023 | Learnable Surrogate Gradient for Direct Training Spiking Neural NetworksabstractSpiking neural networks (SNNs) have increasingly drawn massive research attention due to biological interpretability and efficient computation. Recent achievements are devoted to utilizing the surrogate gradient (SG) method to avoid the dilemma of non-differentiability of spiking activity to directly train SNNs by backpropagation. However, the fixed width of the SG leads to gradient vanishing and mismatch problems, thus limiting the performance of directly trained SNNs. In this work, we propose a novel perspective to unlock the width limitation of SG, called the learnable surrogate gradient (LSG) method. The LSG method modulates the width of SG according to the change of the distribution of the membrane potentials, which is identified to be related to the decay factors based on our theoretical analysis. Then we introduce the trainable decay factors to implement the LSG method, which can optimize the width of SG automatically during training to avoid the gradient vanishing and mismatch problems caused by the limited width of SG. We evaluate the proposed LSG method on both image and neuromorphic datasets. Experimental results show that the LSG method can effectively alleviate the blocking of gradient propagation caused by the limited width of SG when training deep SNNs directly. Meanwhile, the LSG method can help SNNs achieve competitive performance on both latency and accuracy. Shuang Lian, Jiangrong Shen, Qianhui Liu, Rui Yan 0005, Huajin Tang |
IJCAI | 6 |
| 2023 | A Low Latency Adaptive Coding Spike Framework for Deep Reinforcement LearningabstractIn recent years, spiking neural networks (SNNs) have been used in reinforcement learning (RL) due to their low power consumption and event-driven features. However, spiking reinforcement learning (SRL), which suffers from fixed coding methods, still faces the problems of high latency and poor versatility. In this paper, we use learnable matrix multiplication to encode and decode spikes, improving the flexibility of the coders and thus reducing latency. Meanwhile, we train the SNNs using the direct training method and use two different structures for online and offline RL algorithms, which gives our model a wider range of applications. Extensive experiments have revealed that our method achieves optimal performance with ultra-low latency (as low as 0.8% of other SRL methods) and excellent energy efficiency (up to 5X the DNNs) in different algorithms and different environments. Rui Yan 0005, Huajin Tang |
IJCAI | 3 |
| 2023 | Bipolar Population Threshold Encoding for Audio Recognition with Deep Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) have been in-creasingly investigated for audio recognition due to the low power consumption on neuromorphic hardware by mimicking biological neural systems. Since the SNNs are learned from spikes, a critical step lies in the efficient neural encoding of real-valued sound signals to represent complex temporal patterns in speech and environmental sounds. In this paper, we propose a novel Bipolar Population Threshold (BPT) encoding model that effectively captures the trajectory information of time-series speech data by combining temporal and spatial dimensions. The bipolar encoding technique uses positive and negative neurons to capture the dynamic changes in the audio signal, while the threshold intervals allow for a sparse representation that focuses on encoding significant changes, resulting in an efficient and simplified recognition process. Extensively experimenting on three benchmark datasets including the TIDIGITS with speeches, RWCP with sounds, and MedleyDB with music, the numeric results show the superiority of the proposed method by consistently outperforming the state-of-the-art approaches while with fewer spikes, especially in capturing the complex spatio-temporal patterns of audio signals. Xiaocui Lin, Jiangrong Shen, Huajin Tang |
IJCNN | 4 |
| 2023 | Spiking Reinforcement Learning with Memory Ability for Mapless NavigationabstractOur study focuses on mapless navigation in robotics, which involves navigating without an established obstacle map of the environment. Spiking Neural Networks (SNNs) have recently been applied to this task using Deep Reinforcement Learning (DRL), but face challenges in dynamic and partially observable environments, as well as inaccuracies in transmitted data. To overcome these issues, we propose a Multi-Critic DDPG with Spiking Memory (MC-DDPGSM) framework. Our approach introduces a spiking Gate Recurrent Unit layer (Spiking-GRU) to provide memory function and evaluates the state-action value with multi-critic networks. The experimental results demonstrate that our method achieves better performance (success rate, navigation distance, navigation time spent, and power consumption) in complex navigation tasks compared to the state-of-the-art approaches. Furthermore, our model can be transferred to unseen environments without the need for fine-tuning. Mengwen Yuan, Chaofei Hong, Gang Pan 0001, Huajin Tang |
IROS | 6 |
| 2023 | Temporal Conditioning Spiking Latent Variable Models of the Neural Response to Natural Visual ScenesabstractDeveloping computational models of neural response is crucial for understanding sensory processing and neural computations. Current state-of-the-art neural network methods use temporal filters to handle temporal dependencies, resulting in an **unrealistic and inflexible processing paradigm**. Meanwhile, these methods target **trial-averaged firing rates** and fail to capture important features in spike trains. This work presents the temporal conditioning spiking latent variable models (***TeCoS-LVM***) to simulate the neural response to natural visual stimuli. We use spiking neurons to produce spike outputs that directly match the recorded trains. This approach helps to avoid losing information embedded in the original spike trains. We exclude the temporal dimension from the model parameter space and introduce a temporal conditioning operation to allow the model to adaptively explore and exploit temporal dependencies in stimuli sequences in a **natural paradigm**. We show that TeCoS-LVM models can produce more realistic spike activities and accurately fit spike statistics than powerful alternatives. Additionally, learned TeCoS-LVM models can generalize well to longer time scales. Overall, while remaining computationally tractable, our model effectively captures key features of neural coding systems. It thus provides a useful tool for building accurate predictive computational accounts for various sensory perception circuits. Gehua Ma, Runhao Jiang, Rui Yan 0005, Huajin Tang |
NeurIPS | 4 |
| 2023 | Enhancing Adaptive History Reserving by Spiking Convolutional Block Attention Module in Recurrent Neural NetworksabstractSpiking neural networks (SNNs) serve as one type of efficient model to process spatio-temporal patterns in time series, such as the Address-Event Representation data collected from Dynamic Vision Sensor (DVS). Although convolutional SNNs have achieved remarkable performance on these AER datasets, benefiting from the predominant spatial feature extraction ability of convolutional structure, they ignore temporal features related to sequential time points. In this paper, we develop a recurrent spiking neural network (RSNN) model embedded with an advanced spiking convolutional block attention module (SCBAM) component to combine both spatial and temporal features of spatio-temporal patterns. It invokes the history information in spatial and temporal channels adaptively through SCBAM, which brings the advantages of efficient memory calling and history redundancy elimination. The performance of our model was evaluated in DVS128-Gesture dataset and other time-series datasets. The experimental results show that the proposed SRNN-SCBAM model makes better use of the history information in spatial and temporal dimensions with less memory space, and achieves higher accuracy compared to other models. Qi Xu 0008, Yuyuan Gao, Jiangrong Shen, Xuming Ran, Huajin Tang, Gang Pan 0001 |
NeurIPS | 6 |
| 2023 | Grid cell modeling with mapping representation of self-motion for path integration
Jiru Wang, Rui Yan 0005, Huajin Tang |
Neural Comput. Appl. | 3 |
| 2023 | Dual memory model for experience-once task-incremental lifelong learning
Gehua Ma, Runhao Jiang, Huajin Tang |
Neural Networks | 4 |
| 2023 | Towards Energy-Preserving Natural Language Understanding With Spiking Neural NetworksabstractArtificial neural networks have shown promising results in a variety of natural language understanding (NLU) tasks. Despite their successes, conventional neural-based NLU models are criticized for high energy consumption, making them laborious to be widely applied in low-power electronics, such as smartphones and intelligent terminals. In this paper, we introduce a potential direction to alleviate this bottleneck by proposing a spiking encoder. The core of our model is bi-directional spiking neural network (SNN) which transforms numeric values into discrete spiking signals and replaces massive multiplications with much cheaper additive operations. We examine our model on sentiment classification and machine translation tasks. Experimental results reveal that our model achieves comparable classification and translation accuracy to advancedTransformerbaseline, whereas significantly reduces the required computational energy to 0.82%. Rong Xiao 0001, Yu Wan 0004, Baosong Yang, Haibo Zhang 0013, Huajin Tang, Derek F. Wong, Boxing Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | Human-Level Control Through Directly Trained Deep Spiking Q-NetworksabstractAs the third-generation neural networks, spiking neural networks (SNNs) have great potential on neuromorphic hardware because of their high energy efficiency. However, deep spiking reinforcement learning (DSRL), that is, the reinforcement learning (RL) based on SNNs, is still in its preliminary stage due to the binary output and the nondifferentiable property of the spiking function. To address these issues, we propose a deep spiking Q -network (DSQN) in this article. Specifically, we propose a directly trained DSRL architecture based on the leaky integrate-and-fire (LIF) neurons and deep Q -network (DQN). Then, we adapt a direct spiking learning algorithm for the DSQN. We further demonstrate the advantages of using LIF neurons in DSQN theoretically. Comprehensive experiments have been conducted on 17 top-performing Atari games to compare our method with the state-of-the-art conversion method. The experimental results demonstrate the superiority of our method in terms of performance, stability, generalization and energy efficiency. To the best of our knowledge, our work is the first one to achieve state-of-the-art performance on multiple Atari games with the directly trained SNN. Guisong Liu, Wenjie Deng, Xiurui Xie, Li Huang 0002, Huajin Tang |
IEEE Trans. Cybern. | 5 |
| 2023 | A Cell-Based Fast Memetic Algorithm for Automated Convolutional Neural Architecture DesignabstractNeural architecture search (NAS) has attracted much attention in recent years. It automates the neural network construction for different tasks, which is traditionally addressed manually. In the literature, evolutionary optimization (EO) has been proposed for NAS due to its strong global search capability. However, despite the success enjoyed by EO, it is worth noting that existing EO algorithms for NAS are often very computationally expensive, which makes these algorithms unpractical in reality. Keeping this in mind, in this article, we propose an efficient memetic algorithm (MA) for automated convolutional neural network (CNN) architecture search. In contrast to existing EO algorithms for CNN architecture design, a new cell-based architecture search space, and new global and local search operators are proposed for CNN architecture search. To further improve the efficiency of our proposed algorithm, we develop a one-epoch-based performance estimation strategy without any pretrained models to evaluate each found architecture on the training datasets. To investigate the performance of the proposed method, comprehensive empirical studies are conducted against 34 state-of-the-art peer algorithms, including manual algorithms, reinforcement learning (RL) algorithms, gradient-based algorithms, and evolutionary algorithms (EAs), on widely used CIFAR10 and CIFAR100 datasets. The obtained results confirmed the efficacy of the proposed approach for automated CNN architecture design. Junwei Dong, Boyu Hou, Liang Feng 0001, Huajin Tang, Kay Chen Tan, Yew-Soon Ong |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Spiking Deep Residual NetworksabstractSpiking neural networks (SNNs) have received significant attention for their biological plausibility. SNNs theoretically have at least the same computational power as traditional artificial neural networks (ANNs). They possess the potential of achieving energy-efficient machine intelligence while keeping comparable performance to ANNs. However, it is still a big challenge to train a very deep SNN. In this brief, we propose an efficient approach to build deep SNNs. Residual network (ResNet) is considered a state-of-the-art and fundamental model among convolutional neural networks (CNNs). We employ the idea of converting a trained ResNet to a network of spiking neurons named spiking ResNet (S-ResNet). We propose a residual conversion model that appropriately scales continuous-valued activations in ANNs to match the firing rates in SNNs and a compensation mechanism to reduce the error caused by discretization. Experimental results demonstrate that our proposed method achieves state-of-the-art performance on CIFAR-10, CIFAR-100, and ImageNet 2012 with low latency. This work is the first time to build an asynchronous SNN deeper than 100 layers, with comparable performance to its original ANN. Yangfan Hu, Huajin Tang, Gang Pan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Event-Based Multimodal Spiking Neural Network with Attention MechanismabstractHuman brain can effectively integrate visual and auditory information. Dynamic Vision Sensor (DVS) and Dynamic Audio Sensor (DAS) are event-based sensors imitating the mechanism of human retina and cochlea. Since the sensors record the visual and auditory input as asynchronous discrete events, they are inherently suitable to cooperate with the spiking neural network (SNN). Existing works of SNNs for processing events mainly focus on unimodality, however, audiovisual multimodal SNNs are still limited. In this paper, we propose an end-to-end event-based multimodal spiking neural network. The network consists of visual and auditory unimodal subnetworks and a novel attention-based cross-modal subnetwork for fusion. The attention mechanism measures the significance of each modality and allocates the weights to two modalities. We evaluate our proposed multimodal network on an event-based audiovisual joint dataset (MNIST-DVS and N-TIDIGITS datasets). Experimental results show the performance improvement of this multimodal network and the effectiveness of our proposed attention mechanism. Qianhui Liu, Dong Xing, Lang Feng 0002, Huajin Tang, Gang Pan 0001 |
ICASSP | 4 |
| 2022 | Learning Local Event-based Descriptor for Patch-based Stereo MatchingabstractStereo matching is an indispensable function that enables machine vision system to obtain depth information of its environment. However, most of existing algorithms rely on conventional camera, which follows the frame-based scheme and has several shortcomings: low dynamic range, low temporal resolution and high power consumption. To address these issues, we propose two novel patch-based stereo matching methods that exploit the output from a pair of neuromorphic vision sensors. Compared to frame-based camera, neuromorphic vision sensor has independent pixels that generates events at the time intensity changes occur. Based on this unique output, we first construct event representations and present a novel encoding method, which integrates with attention mechanism to encode rich spatial-temporal information of event streams. Then, we design efficient and accuracy networks and propose corresponding loss to train them, which are used to extract event-based descriptors from representations. Finally, the disparity maps are calculated based on local features and refined by two simple smoothing methods. Extensive experiments on the Multi Vehicle Stereo Event Camera Dataset demonstrate the effectiveness of our methods. Peigen Liu, Guang Chen 0001, Zhijun Li 0001, Huajin Tang, Alois C. Knoll |
ICRA | 4 |
| 2022 | Multi-Level Firing with Spiking DS-ResNet: Enabling Better and Deeper Directly-Trained Spiking Neural NetworksabstractSpiking neural networks (SNNs) are bio-inspired neural networks with asynchronous discrete and sparse characteristics, which have increasingly manifested their superiority in low energy consumption. Recent research is devoted to utilizing spatio-temporal information to directly train SNNs by backpropagation. However, the binary and non-differentiable properties of spike activities force directly trained SNNs to suffer from serious gradient vanishing and network degradation, which greatly limits the performance of directly trained SNNs and prevents them from going deeper. In this paper, we propose a multi-level firing (MLF) method based on the existing spatio-temporal back propagation (STBP) method, and spiking dormant-suppressed residual network (spiking DS-ResNet). MLF enables more efficient gradient propagation and the incremental expression ability of the neurons. Spiking DS-ResNet can efficiently perform identity mapping of discrete spikes, as well as provide a more suitable connection for gradient propagation in deep SNNs. With the proposed method, our model achieves superior performances on a non-neuromorphic dataset and two neuromorphic datasets with much fewer trainable parameters and demonstrates the great ability to combat the gradient vanishing and degradation problem in deep SNNs. Lang Feng 0002, Qianhui Liu, Huajin Tang, De Ma, Gang Pan 0001 |
IJCAI | 3 |
| 2022 | Training Deep Convolutional Spiking Neural Networks With Spike Probabilistic Global PoolingabstractRecent work on spiking neural networks (SNNs) has focused on achieving deep architectures. They commonly use backpropagation (BP) to train SNNs directly, which allows SNNs to go deeper and achieve higher performance. However, the BP training procedure is computing intensive and complicated by many trainable parameters. Inspired by global pooling in convolutional neural networks (CNNs), we present the spike probabilistic global pooling (SPGP) method based on a probability function for training deep convolutional SNNs. It aims to remove the difficulty of too many trainable parameters brought by multiple layers in the training process, which can reduce the risk of overfitting and get better performance for deep SNNs (DSNNs). We use the discrete leaky-integrate-fire model and the spatiotemporal BP algorithm for training DSNNs directly. As a result, our model trained with the SPGP method achieves competitive performance compared to the existing DSNNs on image and neuromorphic data sets while minimizing the number of trainable parameters. In addition, the proposed SPGP method shows its effectiveness in performance improvement, convergence, and generalization ability. Shuang Lian, Qianhui Liu, Rui Yan 0005, Gang Pan 0001, Huajin Tang |
Neural Comput. | 5 |
| 2022 | Event stream learning using spatio-temporal event surface
Junfei Dong, Runhao Jiang, Rong Xiao 0001, Rui Yan 0005, Huajin Tang |
Neural Networks | 5 |
| 2022 | Toward Efficient Processing and Learning With Spikes: New Approaches for Multispike LearningabstractSpikes are the currency in central nervous systems for information transmission and processing. They are also believed to play an essential role in low-power consumption of the biological systems, whose efficiency attracts increasing attentions to the field of neuromorphic computing. However, efficient processing and learning of discrete spikes still remain a challenging problem. In this article, we make our contributions toward this direction. A simplified spiking neuron model is first introduced with the effects of both synaptic input and firing output on the membrane potential being modeled with an impulse function. An event-driven scheme is then presented to further improve the processing efficiency. Based on the neuron model, we propose two new multispike learning rules which demonstrate better performance over other baselines on various tasks, including association, classification, and feature detection. In addition to efficiency, our learning rules demonstrate high robustness against the strong noise of different types. They can also be generalized to different spike coding schemes for the classification task, and notably, the single neuron is capable of solving multicategory classifications with our learning rules. In the feature detection task, we re-examine the ability of unsupervised spike-timing-dependent plasticity with its limitations being presented, and find a new phenomenon of losing selectivity. In contrast, our proposed learning rules can reliably solve the task over a wide range of conditions without specific constraints being applied. Moreover, our rules cannot only detect features but also discriminate them. The improved performance of our methods would contribute to neuromorphic computing as a preferable choice. Qiang Yu 0005, Shenglan Li, Huajin Tang, Longbiao Wang, Jianwu Dang 0001, Kay Chen Tan |
IEEE Trans. Cybern. | 3 |
| 2022 | Effective Transfer Learning Algorithm in Spiking Neural NetworksabstractAs the third generation of neural networks, spiking neural networks (SNNs) have gained much attention recently because of their high energy efficiency on neuromorphic hardware. However, training deep SNNs requires many labeled data that are expensive to obtain in real-world applications, as traditional artificial neural networks (ANNs). In order to address this issue, transfer learning has been proposed and widely used in traditional ANNs, but it has limited use in SNNs. In this article, we propose an effective transfer learning framework for deep SNNs based on the domain in-variance representation. Specifically, we analyze the rationality of centered kernel alignment (CKA) as a domain distance measurement relative to maximum mean discrepancy (MMD) in deep SNNs. In addition, we study the feature transferability across different layers by testing on the Office-31, Office-Caltech-10, and PACS datasets. The experimental results demonstrate the transferability of SNNs and show the effectiveness of the proposed transfer learning framework by using CKA in SNNs. Qiugang Zhan, Guisong Liu, Xiurui Xie, Guolin Sun, Huajin Tang |
IEEE Trans. Cybern. | 5 |
| 2022 | Robust Transcoding Sensory Information With Neural SpikesabstractNeural coding, including encoding and decoding, is one of the key problems in neuroscience for understanding how the brain uses neural signals to relate sensory perception and motor behaviors with neural systems. However, most of the existed studies only aim at dealing with the continuous signal of neural systems, while lacking a unique feature of biological neurons, termed spike, which is the fundamental information unit for neural computation as well as a building block for brain-machine interface. Aiming at these limitations, we propose a transcoding framework to encode multi-modal sensory information into neural spikes and then reconstruct stimuli from spikes. Sensory information can be compressed into 10% in terms of neural spikes, yet re-extract 100% of information by reconstruction. Our framework can not only feasibly and accurately reconstruct dynamical visual and auditory scenes, but also rebuild the stimulus patterns from functional magnetic resonance imaging (fMRI) brain activities. More importantly, it has a superb ability of noise immunity for various types of artificial noises and background signals. The proposed framework provides efficient ways to perform multimodal feature representation and reconstruction in a high-throughput fashion, with potential usage for efficient neuromorphic computing in a noisy environment. Qi Xu 0008, Jiangrong Shen, Xuming Ran, Huajin Tang, Gang Pan 0001, Jian K. Liu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Indoor Lighting Estimation Using an Event CameraabstractImage-based methods for indoor lighting estimation suffer from the problem of intensity-distance ambiguity. This paper introduces a novel setup to help alleviate the ambiguity based on the event camera. We further demonstrate that estimating the distance of a light source becomes a well-posed problem under this setup, based on which an optimization-based method and a learning-based method are proposed. Our experimental results validate that our approaches not only achieve superior performance for indoor lighting estimation (especially for the close light) but also significantly alleviate the intensity-distance ambiguity. Peisong Niu, Huajin Tang, Gang Pan 0001 |
CVPR | 4 |
| 2021 | Event-based Action Recognition Using Motion Information and Spiking Neural NetworksabstractEvent-based cameras have attracted increasing attention due to their advantages of biologically inspired paradigm and low power consumption. Since event-based cameras record the visual input as asynchronous discrete events, they are inherently suitable to cooperate with the spiking neural network (SNN). Existing works of SNNs for processing events mainly focus on the task of object recognition. However, events from the event-based camera are triggered by dynamic changes, which makes it an ideal choice to capture actions in the visual scene. Inspired by the dorsal stream in visual cortex, we propose a hierarchical SNN architecture for event-based action recognition using motion information. Motion features are extracted and utilized from events to local and finally to global perception for action recognition. To the best of the authors’ knowledge, it is the first attempt of SNN to apply motion information to event-based action recognition. We evaluate our proposed SNN on three event-based action recognition datasets, including our newly published DailyAction-DVS dataset comprising 12 actions collected under diverse recording conditions. Extensive experimental results show the effectiveness of motion information and our proposed SNN architecture for event-based action recognition. Qianhui Liu, Dong Xing, Huajin Tang, De Ma, Gang Pan 0001 |
IJCAI | 3 |
| 2021 | Few-Shot Learning in Spiking Neural Networks by Multi-Timescale OptimizationabstractLearning new concepts rapidly from a few examples is an open issue in spike-based machine learning. This few-shot learning imposes substantial challenges to the current learning methodologies of spiking neuron networks (SNNs) due to the lack of task-related priori knowledge. The recent learning-to-learn (L2L) approach allows SNNs to acquire priori knowledge through example-level learning and task-level optimization. However, existing L2L-based frameworks do not target the neural dynamics (i.e., neuronal and synaptic parameter changes) on different timescales. This diversity of temporal dynamics is an important attribute in spike-based learning, which facilitates the networks to rapidly acquire knowledge from very few examples and gradually integrate this knowledge. In this work, we consider the neural dynamics on various timescales and provide a multi-timescale optimization (MTSO) framework for SNNs. This framework introduces an adaptive-gated LSTM to accommodate two different timescales of neural dynamics: short-term learning and long-term evolution. Short-term learning is a fast knowledge acquisition process achieved by a novel surrogate gradient online learning (SGOL) algorithm, where the LSTM guides gradient updating of SNN on a short timescale through an adaptive learning rate and weight decay gating. The long-term evolution aims to slowly integrate acquired knowledge and form a priori, which can be achieved by optimizing the LSTM guidance process to tune SNN parameters on a long timescale. Experimental results demonstrate that the collaborative optimization of multi-timescale neural dynamics can make SNNs achieve promising performance for the few-shot learning tasks. Runhao Jiang, Jie Zhang 0012, Rui Yan 0005, Huajin Tang |
Neural Comput. | 4 |
| 2021 | Why grid cells function as a metric for space
Suogui Dang, Yining Wu, Rui Yan 0005, Huajin Tang |
Neural Networks | 4 |
| 2021 | A Novel Illumination-Robust Hand Gesture Recognition System With Event-Based Neuromorphic Vision SensorabstractThe hand gesture recognition system is a noncontact and intuitive communication approach, which, in turn, allows for natural and efficient interaction. This work focuses on developing a novel and robust gesture recognition system, which is insensitive to environmental illumination and background variation. In the field of gesture recognition, standard vision sensors, such as CMOS cameras, are widely used as the sensing devices in state-of-the-art hand gesture recognition systems. However, such cameras depend on environmental constraints, such as lighting variability and the cluttered background, which significantly deteriorates their performances. In this work, we propose an event-based gesture recognition system to overcome the detriment constraints and enhance the robustness of the recognition performance. Our system relies on a biologically inspired neuromorphic vision sensor that has microsecond temporal resolution, high dynamic range, and low latency. The sensor output is a sequence of asynchronous events instead of discrete frames. To interpret the visual data, we utilize a wearable glove as an interaction device with five high-frequency (>100 Hz) active LED markers (ALMs), representing fingers and palm, which are tracked precisely in the temporal domain using a restricted spatiotemporal particle filter algorithm. The latency of the sensing pipeline is negligible compared with the dynamics of the environment as the sensor's temporal resolution allows us to distinguish high frequencies precisely. We design an encoding process to extract features and adopt a lightweight network to classify the hand gestures. The recognition accuracy of our system is comparable to the state-of-the-art methods. To study the robustness of the system, experiments considering illumination and background variations are performed, and the results show that our system is more robust than the state-of-the-art deep learning-based gesture recognition systems. Note to Practitioners-This article addresses the robustness of the hand gesture recognition system that is important for gesture recognition-based applications. Existing methods rely on either the large-volume data to train a deep learning model or to restrict the applied environments (e.g., an ideal environment without dynamic background). However, a vision-based deep learning model requires large computational resources, while the ideal environment limits the practicality of the system. In this work, we introduce a biologically inspired neuromorphic vision sensor and an ALM glove and build a novel gesture recognition system to tackle the above issue. The neuromorphic vision sensor has a microsecond temporal resolution and a high dynamic range. With these properties, the sensing system of our prototype operates in a very low-latency space, which, in turn, ensures that our gesture recognition system is robust to illumination variance and dynamic background. Thus, this work is valuable to the research of illumination-robust gesture recognition systems. Preliminary experiments suggest that our system prototype is feasible, but it has not yet been incorporated into an online gesture recognition system nor tested with complex gestures. In future work, we will concentrate on the improvement of the signal processing methods that advance the current system to complex and practical applications. Guang Chen 0001, Zhongcong Xu, Zhijun Li 0001, Huajin Tang, Sanqing Qu, Kejia Ren, Alois C. Knoll |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2021 | NeuroAED: Towards Efficient Abnormal Event Detection in Visual Surveillance With Neuromorphic Vision SensorabstractAbnormal event detection is an important task in research and industrial applications, which has received considerable attention in recent years. Existing methods usually rely on standard frame-based cameras to record the data and process them with computer vision technologies. In contrast, this paper presents a novel neuromorphic vision based abnormal event detection system. Compared to the frame-based camera, neuromorphic vision sensors, such as Dynamic Vision Sensor (DVS), do not acquire full images at a fixed frame rate but rather have independent pixels that output intensity changes (called events) asynchronously at the time they occur. Thus, it avoids the design of the encryption scheme. Since events are triggered by moving edges on the scene, DVS is a natural motion detector for the abnormal objects and automatically filters out any temporally-redundant information. Based on this unique output, we first propose a highly efficient method based on the event density to select activated event cuboids and locate the foreground. We design a novel event-based multiscale spatio-temporal descriptor to extract features from the activated event cuboids for the abnormal event detection. Additionally, we build the NeuroAED dataset, the first public dataset dedicated to abnormal event detection with neuromorphic vision sensor. The NeuroAED dataset consists of four sub-datasets: Walking, Campus, Square, and Stair dataset. Experiments are conducted based on these datasets and demonstrate the high efficiency and accuracy of our method. Guang Chen 0001, Peigen Liu, Zhengfa Liu, Huajin Tang, Lin Hong, Jinhu Dong, Jörg Conradt, Alois C. Knoll |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2021 | Robust Environmental Sound Recognition With Sparse Key-Point Encoding and Efficient Multispike LearningabstractThe capability for environmental sound recognition (ESR) can determine the fitness of individuals in a way to avoid dangers or pursue opportunities when critical sound events occur. It still remains mysterious about the fundamental principles of biological systems that result in such a remarkable ability. Additionally, the practical importance of ESR has attracted an increasing amount of research attention, but the chaotic and nonstationary difficulties continue to make it a challenging task. In this article, we propose a spike-based framework from a more brain-like perspective for the ESR task. Our framework is a unifying system with consistent integration of three major functional parts which are sparse encoding, efficient learning, and robust readout. We first introduce a simple sparse encoding, where key points are used for feature representation, and demonstrate its generalization to both spike- and nonspike-based systems. Then, we evaluate the learning properties of different learning rules in detail with our contributions being added for improvements. Our results highlight the advantages of multispike learning, providing a selection reference for various spike-based developments. Finally, we combine the multispike readout with the other parts to form a system for ESR. Experimental results show that our framework performs the best as compared to other baseline approaches. In addition, we show that our spike-based framework has several advantageous characteristics including early decision making, small dataset acquiring, and ongoing dynamic processing. Our framework is the first attempt to apply the multispike characteristic of nervous neurons to ESR. The outstanding performance of our approach would potentially contribute to draw more research efforts to push the boundaries of spike-based paradigm to a new horizon. Qiang Yu 0005, Yanli Yao, Longbiao Wang, Huajin Tang, Jianwu Dang 0001, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Effective AER Object Classification Using Segmented Probability-Maximization Learning in Spiking Neural NetworksabstractAddress event representation (AER) cameras have recently attracted more attention due to the advantages of high temporal resolution and low power consumption, compared with traditional frame-based cameras. Since AER cameras record the visual input as asynchronous discrete events, they are inherently suitable to coordinate with the spiking neural network (SNN), which is biologically plausible and energy-efficient on neuromorphic hardware. However, using SNN to perform the AER object classification is still challenging, due to the lack of effective learning algorithms for this new representation. To tackle this issue, we propose an AER object classification model using a novel segmented probability-maximization (SPA) learning algorithm. Technically, 1) the SPA learning algorithm iteratively maximizes the probability of the classes that samples belong to, in order to improve the reliability of neuron responses and effectiveness of learning; 2) a peak detection (PD) mechanism is introduced in SPA to locate informative time points segment by segment, based on which information within the whole event stream can be fully utilized by the learning. Extensive experimental results show that, compared to state-of-the-art methods, not only our model is more effective, but also it requires less information to reach a certain level of accuracy. Qianhui Liu, Haibo Ruan, Dong Xing, Huajin Tang, Gang Pan 0001 |
AAAI | 4 |
| 2020 | A Novel Mathematic Entorhinal-Hippocampal System Building Cognitive Map
Jianxin Peng, Suogui Dang, Rui Yan 0005, Huajin Tang |
ICONIP (2) | 4 |
| 2020 | Deep Spiking Neural Network Using Spatio-temporal Backpropagation with Variable ResistanceabstractIn recent years, the learning of deep spiking neural networks(SNN) has attracted increasing researchers' interest, and has also made important progresses in theories and applications. It is desired to choose a neuron model with biological features and suitable for SNN training. Currently, Leaky Integrate-and-Fire(LIF) model is mainly used in deep SNN and some factors that can express the spatio-temporal information are ignored in the model. In this work, inspired by the Hodgkin-Huxley(H-H) model, we propose an improved LIF neuron model, which is an iterative current-based LIF model with voltage-based variable resistance. The improved neuron model is closer to the characteristics of the biological neuron model, which can make use of the spatio-temporal information. We further construct a new SNN learning algorithm that uses spatio-temporal back propagation by defining a loss function. We evaluated the proposed methods on single-label and multi-label data sets. The experimental results show that the variable resistance of the neuron model will affect the performance of the model. Choosing the appropriate relationship between the variable resistance and the membrane voltage can effectively improve the recognition accuracy. Xianglan Wen, Pengjie Gu, Rui Yan 0005, Huajin Tang |
IJCNN | 4 |
| 2020 | An FPGA Implementation of Deep Spiking Neural Networks for Low-Power and Fast ClassificationabstractA spiking neural network (SNN) is a type of biological plausibility model that performs information processing based on spikes. Training a deep SNN effectively is challenging due to the nondifferention of spike signals. Recent advances have shown that high-performance SNNs can be obtained by converting convolutional neural networks (CNNs). However, the large-scale SNNs are poorly served by conventional architectures due to the dynamic nature of spiking neurons. In this letter, we propose a hardware architecture to enable efficient implementation of SNNs. All layers in the network are mapped on one chip so that the computation of different time steps can be done in parallel to reduce latency. We propose new spiking max-pooling method to reduce computation complexity. In addition, we apply approaches based on shift register and coarsely grained parallels to accelerate convolution operation. We also investigate the effect of different encoding methods on SNN accuracy. Finally, we validate the hardware architecture on the Xilinx Zynq ZCU102. The experimental results on the MNIST data set show that it can achieve an accuracy of 98.94% with eight-bit quantized weights. Furthermore, it achieves 164 frames per second (FPS) under 150 MHz clock frequency and obtains 41[Formula: see text] speed-up compared to CPU implementation and 22 times lower power than GPU implementation. Xiping Ju, Biao Fang, Rui Yan 0005, Huajin Tang |
Neural Comput. | 5 |
| 2020 | Deep CovDenseSNN: A hierarchical event-driven dynamic framework with spiking neurons in noisy environment
Qi Xu 0008, Jianxin Peng, Jiangrong Shen, Huajin Tang, Gang Pan 0001 |
Neural Networks | 4 |
| 2020 | Unsupervised AER Object Recognition Based on Multiscale Spatio-Temporal Features and Spiking NeuronsabstractThis article proposes an unsupervised address event representation (AER) object recognition approach. The proposed approach consists of a novel multiscale spatio-temporal feature (MuST) representation of input AER events and a spiking neural network (SNN) using spike-timing-dependent plasticity (STDP) for object recognition with MuST. MuST extracts the features contained in both the spatial and temporal information of AER event flow, and forms an informative and compact feature spike representation. We show not only how MuST exploits spikes to convey information more effectively, but also how it benefits the recognition using SNN. The recognition process is performed in an unsupervised manner, which does not need to specify the desired status of every single neuron of SNN, and thus can be flexibly applied in real-world recognition tasks. The experiments are performed on five AER datasets including a new one named GESTURE-DVS. Extensive experimental results show the effectiveness and advantages of the proposed approach. Qianhui Liu, Gang Pan 0001, Haibo Ruan, Dong Xing, Qi Xu 0008, Huajin Tang |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2020 | An Event-Driven Categorization Model for AER Image Sensors Using Multispike Encoding and LearningabstractIn this article, we present a systematic computational model to explore brain-based computation for object recognition. The model extracts temporal features embedded in address-event representation (AER) data and discriminates different objects by using spiking neural networks (SNNs). We use multispike encoding to extract temporal features contained in the AER data. These temporal patterns are then learned through the tempotron learning rule. The presented model is consistently implemented in a temporal learning framework, where the precise timing of spikes is considered in the feature-encoding and learning process. A noise-reduction method is also proposed by calculating the correlation of an event with the surrounding spatial neighborhood based on the recently proposed time-surface technique. The model evaluated on wide spectrum data sets (MNIST, N-MNIST, MNIST-DVS, AER Posture, and Poker Card) demonstrates its superior recognition performance, especially for the events with noise. Rong Xiao 0001, Huajin Tang, Rui Yan 0005, Garrick Orchard |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | A Multi-spike Approach for Robust Sound RecognitionabstractThe extraordinary performance of the brain on various cognitive tasks motivates the design of a biologically plausible system for the challenging task of environmental sound recognition. In this paper, we propose a novel approach based on multi-spike learning and key-point encoding. Our encoding extracts local temporal and spectral information from the sound and converts it into spatiotemporal spike pattern, which is further learned by the following spiking neural networks. Our experiments demonstrate the robustness and effectiveness of our approach across a variety of noise conditions, outperforming other conventional baseline methods in both mismatched and multi-condition scenarios. Qiang Yu 0005, Yanli Yao, Longbiao Wang, Huajin Tang, Jianwu Dang 0001 |
ICASSP | 4 |
| 2019 | Dance to Music Expressively: A Brain-Inspired System Based on Audio-Semantic Model for Cognitive Development of Robots
Dengju Li, Rui Yan 0005, Huajin Tang |
ICONIP (4) | 4 |
| 2019 | STCA: Spatio-Temporal Credit Assignment with Delayed Feedback in Deep Spiking Neural NetworksabstractThe temporal credit assignment problem, which aims to discover the predictive features hidden in distracting background streams with delayed feedback, remains a core challenge in biological and machine learning. To address this issue, we propose a novel spatio-temporal credit assignment algorithm called STCA for training deep spiking neural networks (DSNNs). We present a new spatiotemporal error backpropagation policy by defining a temporal based loss function, which is able to credit the network losses to spatial and temporal domains simultaneously. Experimental results on MNIST dataset and a music dataset (MedleyDB) demonstrate that STCA can achieve comparable performance with other state-of-the-art algorithms with simpler architectures. Furthermore, STCA successfully discovers predictive sensory features and shows the highest performance in the unsegmented sensory event detection tasks. Pengjie Gu, Rong Xiao 0001, Gang Pan 0001, Huajin Tang |
IJCAI | 4 |
| 2019 | Fast and Accurate Classification with a Multi-Spike Learning Algorithm for Spiking NeuronsabstractThe formulation of efficient supervised learning algorithms for spiking neurons is complicated and remains challenging. Most existing learning methods with the precisely firing times of spikes often result in relatively low efficiency and poor robustness to noise. To address these limitations, we propose a simple and effective multi-spike learning rule to train neurons to match their output spike number with a desired one. The proposed method will quickly find a local maximum value (directly related to the embedded feature) as the relevant signal for synaptic updates based on membrane potential trace of a neuron, and constructs an error function defined as the difference between the local maximum membrane potential and the firing threshold. With the presented rule, a single neuron can be trained to learn multi-category tasks, and can successfully mitigate the impact of the input noise and discover embedded features. Experimental results show the proposed algorithm has higher precision, lower computation cost, and better noise robustness than current state-of-the-art learning methods under a wide range of learning tasks. Rong Xiao 0001, Qiang Yu 0005, Rui Yan 0005, Huajin Tang |
IJCAI | 4 |
| 2019 | A temporal encoding method based on expansion representationabstractTemporal encoding of visual stimulus based on the spiking neural networks is important and challenging. Inspired by the representation of neuron population in the sensory pathway, we propose a new hierarchical encoding method, which amplifies the difference among spike trains through the synaptic projection matrix and leaves the top neurons keeping active by a winner-take-all scheme. Finally, the neurons are assigned to positive and negative subthreshold membrane oscillation to fire new spike trains. In the optical character recognition task, we compare the readout performance of the proposed encoding method and the phase encoding method. When the image is severely destroyed, for example, at 25% noise, our method still maintains 88% recognition accuracy, which is 30% higher than the accuracy of phase encoding. The robustness of readout neurons to perform recognition tasks is enhanced through the proposed encoding method. We performed the benchmark two-class classification experiment on Caltech 101 dataset, and the accuracy can reach 99.4%, which is 17% higher than the phase encoding. Compared with the other spiking deep neural network, our method also shows great improvement and strong scalability. Yan Dai 0008, Mengwen Yuan, Huajin Tang, Rui Yan 0005 |
IJCNN | 3 |
| 2019 | An Unsupervised Spiking Deep Neural Network for Object Recognition
Zeyang Song, Mengwen Yuan, Huajin Tang |
ISNN (2) | 4 |
| 2019 | Reinforcement Learning in Spiking Neural Networks with Stochastic and Deterministic SynapsesabstractThough succeeding in solving various learning tasks, most existing reinforcement learning (RL) models have failed to take into account the complexity of synaptic plasticity in the neural system. Models implementing reinforcement learning with spiking neurons involve only a single plasticity mechanism. Here, we propose a neural realistic reinforcement learning model that coordinates the plasticities of two types of synapses: stochastic and deterministic. The plasticity of the stochastic synapse is achieved by the hedonistic rule through modulating the release probability of synaptic neurotransmitter, while the plasticity of the deterministic synapse is achieved by a variant of a reward-modulated spike-timing-dependent plasticity rule through modulating the synaptic strengths. We evaluate the proposed learning model on two benchmark tasks: learning a logic gate function and the 19-state random walk problem. Experimental results show that the coordination of diverse synaptic plasticities can make the RL model learn in a rapid and stable form. Mengwen Yuan, Rui Yan 0005, Huajin Tang |
Neural Comput. | 4 |
| 2019 | A structure-time parallel implementation of spike-based deep learning
Huajin Tang, Rui Yan 0005 |
Neural Networks | 3 |
| 2019 | Guest Editorial Special Section on Emerging Information Sharing and Design Technologies on Robotics and Mechatronics Systems for Intelligent ManufacturingabstractThe ten papers in this special section aim gather the latest research and development works on design, sensing, and intelligent control of robotics and mechatronics systems resulting from the emerging information sharing and design technologies. Guilin Yang, I-Ming Chen 0001, Chin-Yin Chen, Huajin Tang, Chi Zhang 0014 |
IEEE Trans. Ind. Informatics | 4 |
| 2018 | A Gesture Recognition Method Based on Spiking Neural Networks for Cognition Development
Dong Niu, Dengju Li, Rui Yan 0005, Huajin Tang |
ICONIP (1) | 4 |
| 2018 | Jointly Learning Network Connections and Link Weights in Spiking Neural NetworksabstractSpiking neural networks (SNNs) are considered to be biologically plausible and power-efficient on neuromorphic hardware. However, unlike the brain mechanisms, most existing SNN algorithms have fixed network topologies and connection relationships. This paper proposes a method to jointly learn network connections and link weights simultaneously. The connection structures are optimized by the spike-timing-dependent plasticity (STDP) rule with timing information, and the link weights are optimized by a supervised algorithm. The connection structures and the weights are learned alternately until a termination condition is satisfied. Experiments are carried out using four benchmark datasets. Our approach outperforms classical learning methods such as STDP, Tempotron, SpikeProp, and a state-of-the-art supervised algorithm. In addition, the learned structures effectively reduce the number of connections by about 24%, thus facilitate the computational efficiency of the network. Jiangrong Shen, Yueming Wang 0001, Huajin Tang, Hang Yu 0010, Zhaohui Wu 0001, Gang Pan 0001 |
IJCAI | 4 |
| 2018 | CSNN: An Augmented Spiking based Framework with Perceptron-InceptionabstractSpiking Neural Networks (SNNs) represent and transmit information in spikes, which is considered more biologically realistic and computationally powerful than the traditional Artificial Neural Networks. The spiking neurons encode useful temporal information and possess highly anti-noise property. The feature extraction ability of typical SNNs is limited by shallow structures. This paper focuses on improving the feature extraction ability of SNNs in virtue of powerful feature extraction ability of Convolutional Neural Networks (CNNs). CNNs can extract abstract features resorting to the structure of the convolutional feature maps. We propose a CNN-SNN (CSNN) model to combine feature learning ability of CNNs with cognition ability of SNNs. The CSNN model learns the encoded spatial temporal representations of images in an event-driven way. We evaluate the CSNN model on the handwritten digits images dataset MNIST and its variational databases. In the presented experimental results, the proposed CSNN model is evaluated regarding learning capabilities, encoding mechanisms, robustness to noisy stimuli and its classification performance. The results show that CSNN behaves well compared to other cognitive models with significantly fewer neurons and training samples. Our work brings more biological realism into modern image classification models, with the hope that these models can inform how the brain performs this high-level vision task. Qi Xu 0008, Hang Yu 0010, Jiangrong Shen, Huajin Tang, Gang Pan 0001 |
IJCAI | 5 |
| 2018 | A Supervised Multi-Spike Learning Algorithm for Spiking Neural NetworksabstractThe formulation of efficient supervised learning algorithms for Spiking Neural Network (SNN) is difficult and remains challenging. This paper presents a supervised multispike learning algorithm, which is used to train neurons to output spike train with a target firing rate. The proposed algorithm simplifies the expression of the membrane potential by assuming a special condition of the threshold, thus allows the application of a gradient descent to optimize the synaptic weights. Additionally, in the presented experimental results, the proposed algorithm is evaluated regarding its initial setups, its classification performance for rate-based and timing-based patterns and its capability to sound recognition. The results also demonstrate that the proposed algorithm can achieve a competitive accuracy in temporal pattern classification and sound recognition. Huajin Tang, Gang Pan 0001 |
IJCNN | 2 |
| 2018 | Spike-based encoding and learning of spectrum features for robust sound recognition
Rong Xiao 0001, Huajin Tang, Pengjie Gu |
Neurocomputing | 2 |
| 2018 | Connections Between Nuclear-Norm and Frobenius-Norm-Based RepresentationsabstractA lot of works have shown that frobenius-norm-based representation (FNR) is competitive to sparse representation and nuclear-norm-based representation (NNR) in numerous tasks such as subspace clustering. Despite the success of FNR in experimental studies, less theoretical analysis is provided to understand its working mechanism. In this brief, we fill this gap by building the theoretical connections between FNR and NNR. More specially, we prove that: 1) when the dictionary can provide enough representative capacity, FNR is exactly NNR even though the data set contains the Gaussian noise, Laplacian noise, or sample-specified corruption and 2) otherwise, FNR and NNR are two solutions on the column space of the dictionary. Xi Peng 0001, Canyi Lu, Zhang Yi 0001, Huajin Tang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2018 | Sparse Temporal Encoding of Visual Features for Robust Object Recognition by Spiking NeuronsabstractRobust object recognition in spiking neural systems remains a challenging in neuromorphic computing area as it needs to solve both the effective encoding of sensory information and also its integration with downstream learning neurons. We target this problem by developing a spiking neural system consisting of sparse temporal encoding and temporal classifier. We propose a sparse temporal encoding algorithm which exploits both spatial and temporal information derived from an spike-timing-dependent plasticity-based HMAX feature extraction process. The temporal feature representation, thus, becomes more appropriate to be integrated with a temporal classifier based on spiking neurons rather than with nontemporal classifier. The algorithm has been validated on two benchmark data sets and the results show the temporal feature encoding and learning-based method achieves high recognition accuracy. The proposed model provides an efficient approach to perform feature representation and recognition in a consistent temporal learning framework, which is easily adapted to neuromorphic implementations. Yajing Zheng, Rui Yan 0005, Huajin Tang, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2017 | An Event-Driven Computational System with Spiking Neurons for Object Recognition
Rong Xiao 0001, Huajin Tang |
ICONIP (6) | 3 |
| 2017 | Cognitive memory and mapping in a brain-like system for robotic navigation
Huajin Tang, Aditya Narayanamoorthy, Rui Yan 0005 |
Neural Networks | 1 |
| 2017 | Constructing the L2-Graph for Robust Subspace Learning and Subspace ClusteringabstractUnder the framework of graph-based learning, the key to robust subspace clustering and subspace learning is to obtain a good similarity graph that eliminates the effects of errors and retains only connections between the data points from the same subspace (i.e., intrasubspace data points). Recent works achieve good performance by modeling errors into their objective functions to remove the errors from the inputs. However, these approaches face the limitations that the structure of errors should be known prior and a complex convex problem must be solved. In this paper, we present a novel method to eliminate the effects of the errors from the projection space (representation) rather than from the input space. We first prove that ℓ1-, ℓ2-, ℓ∞-, and nuclear-norm-based linear projection spaces share the property of intrasubspace projection dominance, i.e., the coefficients over intrasubspace data points are larger than those over intersubspace data points. Based on this property, we introduce a method to construct a sparse similarity graph, called L2-graph. The subspace clustering and subspace learning algorithms are developed upon L2-graph. We conduct comprehensive experiment on subspace learning, image clustering, and motion segmentation and consider several quantitative benchmarks classification/clustering accuracy, normalized mutual information, and running time. Results show that L2-graph outperforms many state-of-the-art methods in our experiments, including L1-graph, low rank representation (LRR), and latent LRR, least square regression, sparse subspace clustering, and locally linear representation. Xi Peng 0001, Zhiding Yu, Zhang Yi 0001, Huajin Tang |
IEEE Trans. Cybern. | 4 |
| 2017 | Bag of Events: An Efficient Probability-Based Feature Extraction Method for AER Image SensorsabstractAddress event representation (AER) image sensors represent the visual information as a sequence of events that denotes the luminance changes of the scene. In this paper, we introduce a feature extraction method for AER image sensors based on the probability theory, namely, bag of events (BOE). The proposed approach represents each object as the joint probability distribution of the concurrent events, and each event corresponds to a unique activated pixel of the AER sensor. The advantages of BOE include: 1) it is a statistical learning method and has a good interpretability in mathematics; 2) BOE can significantly reduce the effort to tune parameters for different data sets, because it only has one hyperparameter and is robust to the value of the parameter; 3) BOE is an online learning algorithm, which does not require the training data to be collected in advance; 4) BOE can achieve competitive results in real time for feature extraction (>275 frames/s and >120,000 events/s); and 5) the implementation complexity of BOE only involves some basic operations, e.g., addition and multiplication. This guarantees the hardware friendliness of our method. The experimental results on three popular AER databases (i.e., MNIST-dynamic vision sensor, Poker Card, and Posture) show that our method is remarkably faster than two recently proposed AER categorization systems while preserving a good classification accuracy. Xi Peng 0001, Bo Zhao 0018, Rui Yan 0005, Huajin Tang, Zhang Yi 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2017 | Guest Editorial Learning in Neuromorphic Systems and Cyborg IntelligenceabstractNeuromorphic computing has become an important emerging research area in recent years. By emulating computational principles and architecture found in neural systems, neuromorphic computing has led to the development of neuromorphic sensors, processors, and sensory motor systems for robotic agents. It has also led to rapid progress in related areas covering computational theories of sensory coding, synaptic computing, learning, and signal processing algorithms, circuit designs, and implementations. The work in these areas shows neuromorphic approaches with appealing computational advantages over conventional approaches, but at the same time, neuromorphic systems still pose many research challenges. Neuromorphic computing overlaps with another area called cyborg intelligence which is dedicated to integrating artificial intelligence (AI) with biological intelligence closely and deeply by connecting computer systems and biological beings. Cyborg intelligence aims to compensate for the weaknesses of both systems by combining the computational power of machines with the perceptive and cognitive abilities of biological systems. Recently, many of the advances in cyborg intelligence methods, systems, and applications have demonstrated the trend of the rapid integration of cyborg intelligence with neuromorphic computing in both breadth and depth. These areas pose innumerable interesting and significant questions for AI and could fundamentally change the landscape of AI research. Zhaohui Wu 0001, Ryad Benosman, Huajin Tang, Shih-Chii Liu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2016 | Hebbian learning analysis of a grid cell based cognitive mapping systemabstractIt is believed that animals sense the environment by encoding the spatial information into an internal representation. Grid cells in the rat entorhinal cortex are found with periodic firing fields and have been hypothesized as the basis for the ‘cognitive map’. This paper presents a grid-cell based computational model for cognitive map building together with an analysis of the learning from grid cells to place cells. Experiment results demonstrate that the learning from grid cells to place cells plays an important role affecting the performance of map building and that the proposed computational model provides an alternative approach for robot mapping. Miaolong Yuan, Huajin Tang, Weiyun Yau |
CEC | 3 |
| 2016 | Robot-to-human handover with obstacle avoidance via continuous time Recurrent Neural NetworkabstractParallel with the development of service robots, it is vital for the robots to carry out handovers autonomously. Robot-to-human handover is a coordination in time and space for a robot to deliver an object to human. A good robot-to-human handover should consider human safety and preference, natural motion planning that mimics human and adaptability to the changes of the environment. Conventional handover motion mostly rely on sampling-based algorithms that emphasizes on kinematic and dynamic analysis. This kind of motion planning could become complicated and slow in response if the handover motion is implemented in a dynamic environment where real time motion planning is required. To simplify the implementation of robot-to-human handover, a motion learning and generation framework that based on Continuous Time Recurrent Neural Network(CTRNN) is proposed. The proposed framework is equipped with the capabilities of object recognition, motion generation based on past learning experience and obstacle adaptation. As compared with conventional method, the proposed framework could be easily extended to handover motion with high dimensional configuration spaces as the motion can be generated from the learnt experience. In the proposed framework, the handover behaviour can be learnt via human-guided motion teaching which provides an intuitive and visible solution for motion planning. The proposed framework has been experimentally evaluated on a customized design robot via robotto-human handover testing. Based on the testing, the feasibility of the proposed framework had been justified. Huajin Tang, Boon Hwa Tan, Rui Yan 0005 |
CEC | 1 |
| 2016 | A Unified Framework for Representation-Based Subspace Clustering of Out-of-Sample and Large-Scale Dataabstract-norm-based representation, and have achieved the state-of-the-art performance. However, these methods have suffered from the following two limitations. First, the time complexities of these methods are at least proportional to the cube of the data size, which make those methods inefficient for solving the large-scale problems. Second, they cannot cope with the out-of-sample data that are not used to construct the similarity graph. To cluster each out-of-sample datum, the methods have to recalculate the similarity graph and the cluster membership of the whole data set. In this paper, we propose a unified framework that makes the representation-based subspace clustering algorithms feasible to cluster both the out-of-sample and the large-scale data. Under our framework, the large-scale problem is tackled by converting it as the out-of-sample problem in the manner of sampling, clustering, coding, and classifying. Furthermore, we give an estimation for the error bounds by treating each subspace as a point in a hyperspace. Extensive experimental results on various benchmark data sets show that our methods outperform several recently proposed scalable methods in clustering a large-scale data set. Xi Peng 0001, Huajin Tang, Lei Zhang 0005, Zhang Yi 0001, Shijie Xiao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2016 | A Spiking Neural Network System for Robust Sequence RecognitionabstractThis paper proposes a biologically plausible network architecture with spiking neurons for sequence recognition. This architecture is a unified and consistent system with functional parts of sensory encoding, learning, and decoding. This is the first systematic model attempting to reveal the neural mechanisms considering both the upstream and the downstream neurons together. The whole system is a consistent temporal framework, where the precise timing of spikes is employed for information processing and cognitive computing. Experimental results show that the system is competent to perform the sequence recognition, being robust to noisy sensory inputs and invariant to changes in the intervals between input stimuli within a certain range. The classification ability of the temporal learning rule used in the system is investigated through two benchmark tasks that outperform the other two widely used learning rules for classification. The results also demonstrate the computational power of spiking neurons over perceptrons for processing spatiotemporal patterns. In summary, the system provides a general way with spiking neurons to encode external stimuli into spatiotemporal spikes, to learn the encoded spike patterns with temporal learning rules, and to decode the sequence order with downstream neurons. The system structure would be beneficial for developments in both hardware and software. Qiang Yu 0005, Rui Yan 0005, Huajin Tang, Kay Chen Tan, Haizhou Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | Robust Subspace Clustering via Thresholding Ridge RegressionabstractGiven a data set from a union of multiple linear subspaces, a robust subspace clustering algorithm fits each group of data points with a low-dimensional subspace and then clusters these data even though they are grossly corrupted or sampled from the union of dependent subspaces. Under the framework of spectral clustering, recent works using sparse representation, low rank representation and their extensions achieve robust clustering results by formulating the errors (e.g., corruptions) into their objective functions so that the errors can be removed from the inputs. However, these approaches have suffered from the limitation that the structure of the errors should be known as the prior knowledge. In this paper, we present a new method of robust subspace clustering by eliminating the effect of the errors from the projection space (representation) rather than from the input space. We firstly prove that ell_1-, ell_2-, and ell_infty-norm-based linear projection spaces share the property of intra-subspace projection dominance, i.e., the coefficients over intra-subspace data points are larger than those over inter-subspace data points. Based on this property, we propose a robust and efficient subspace clustering algorithm, called Thresholding Ridge Regression (TRR). TRR calculates the ell2-norm-based coefficients of a given data set and performs a hard thresholding operator; and then the coefficients are used to build a similarity graph for clustering. Experimental studies show that TRR outperforms the state-of-the-art methods with respect to clustering quality, robustness, and time-saving. Xi Peng 0001, Zhang Yi 0001, Huajin Tang |
AAAI | 3 |
| 2015 | An Entorhinal-Hippocampal Model for Simultaneous Cognitive Map BuildingabstractHippocampal place cells and entorhinal grid cells have been hypothesized to be able to form map-like spatial representation of the environment, namely cognitive map. In most prior approaches, either neural network methods or only hippocampal models are used for building cognitive maps, lacking biological fidelity to the entorhinal-hippocampal system. This paper presents a novel computational model to build cognitive maps of real environments using both place cells and grid cells. The proposed model includes two major components: (1) A competitive Hebbian learning algorithm is used to select velocity-coupled grid cell population activities, which path-integrate self-motion signals to determine computation of place cell population activities; (2) Visual cues of environments are used to correct the accumulative errors intrinsically associated with the path integration process. Experiments performed on a mobile robot show that cognitive maps of the real environment can be efficiently built. The proposed model would provide an alternative neuro-inspired approach for robotic mapping, navigation and localization. Miaolong Yuan, Vui Ann Shim, Huajin Tang, Haizhou Li 0001 |
AAAI | 4 |
| 2015 | Fast low rank representation based spatial pyramid matching for image classification
Xi Peng 0001, Rui Yan 0005, Bo Zhao 0018, Huajin Tang, Zhang Yi 0001 |
Knowl. Based Syst. | 4 |
| 2015 | Adaptive Memetic Computing for Evolutionary Multiobjective OptimizationabstractInspired by biological evolution, a plethora of algorithms with evolutionary features have been proposed. These algorithms have strengths in certain aspects, thus yielding better optimization performance in a particular problem. However, in a wide range of problems, none of them are superior to one another. Synergetic combination of these algorithms is one of the potential ways to ameliorate their search ability. Based on this idea, this paper proposes an adaptive memetic computing as the synergy of a genetic algorithm, differential evolution, and estimation of distribution algorithm. The ratio of the number of fitter solutions produced by the algorithms in a generation defines their adaptability features in the next generation. Subsequently, a subset of solutions undergoes local search using the evolutionary gradient search algorithm. This memetic technique is then implemented in two prominent frameworks of multiobjective optimization: the domination- and decomposition-based frameworks. The performance of the adaptive memetic algorithms is validated in a wide range of test problems with different characteristics and difficulties. Vui Ann Shim, Kay Chen Tan, Huajin Tang |
IEEE Trans. Cybern. | 3 |
| 2015 | Feedforward Categorization on AER Motion Events Using Cortex-Like Features in a Spiking Neural NetworkabstractThis paper introduces an event-driven feedforward categorization system, which takes data from a temporal contrast address event representation (AER) sensor. The proposed system extracts bio-inspired cortex-like features and discriminates different patterns using an AER based tempotron classifier (a network of leaky integrate-and-fire spiking neurons). One of the system's most appealing characteristics is its event-driven processing, with both input and features taking the form of address events (spikes). The system was evaluated on an AER posture dataset and compared with two recently developed bio-inspired models. Experimental results have shown that it consumes much less simulation time while still maintaining comparable performance. In addition, experiments on the Mixed National Institute of Standards and Technology (MNIST) image dataset have demonstrated that the proposed system can work not only on raw AER data but also on images (with a preprocessing step to convert images into AER events) and that it can maintain competitive accuracy even when noise is added. The system was further evaluated on the MNIST dynamic vision sensor dataset (in which data is recorded using an AER dynamic vision sensor), with testing accuracy of 88.14%. Bo Zhao 0018, Ruoxi Ding, Shoushun Chen, Bernabé Linares-Barranco, Huajin Tang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2014 | A new learning rule for classification of spatiotemporal spike patternsabstractIn this paper, we present a new learning rule for classification of spatiotemporal spike patterns. This rule is derived from the common Widrow-Hoff rule, and it can be used for both the association and the classification. We mainly focus on investigating its classification ability in this paper. Through experimental simulations, it can be seen that this rule can successfully train the neuron to reproduce the desired spikes. In the classification task, the neuron is capable to classify different categories with the learning rule. We have proposed two decision-making schemes which are the absolute confidence and the relative confidence criteria. The classification performance is largely improved by the relative confidence criterion. The performance of this rule on classification of spatiotemporal spike patterns is also investigated and benchmarked by the tempotron rule. Qiang Yu 0005, Huajin Tang, Kay Chen Tan |
IJCNN | 2 |
| 2014 | Bio-inspired categorization using event-driven feature extraction and spike-based learningabstractThis paper presents a fully event-driven feedforward architecture that accounts for rapid categorization. The proposed algorithm processes the address event data generated either from an image or from Address-Event-Representation (AER) temporal contrast vision sensor. Bio-inspired, cortex-like, spike-based features are obtained through event-driven convolution and neural competition. The extracted spike feature patterns are then classified by a network of leaky integrate-and-fire (LIE) spiking neurons, in which the weights are trained using tempotron learning rule. One appealing characteristic of our system is the fully event-driven processing. The input, the features, and the classification are all based on address events (spikes). Experimental results on three datasets have proved the efficacy of the proposed algorithm. Bo Zhao 0018, Shoushun Chen, Huajin Tang |
IJCNN | 3 |
| 2014 | Direction-driven navigation using cognitive map for mobile robotsabstractBuilding an integrated mobile robot that can navigate freely and deliberately in an indoor environment is a challenging task. In this paper, we propose a direction-driven navigation system, in which a mobile robot is instructed to follow the given directions instead of a global path. This is similar to humans when travelling to a target destination by following the directional guidance from someone else. This system consists of two main modules including a grid-based direction planner and a multilayered asymmetrical local navigation module. The directions of movement are computed from a brain-inspired spatial cognitive map using the grid-based direction planner, while the motion of the robot is controlled by the local navigation module. In order to improve the mapping performance using RGB-D input, an improved template comparison method is suggested. The proposed system is implemented in a mobile robot that performs simultaneous localization, mapping, and navigation in a typical office environment. The experimental results indicate that the robot can efficiently navigate to target destinations. Vui Ann Shim, Miaolong Yuan, Huajin Tang, Haizhou Li 0001 |
IROS | 4 |
| 2014 | Flexible and robust robotic arm design and skill learning by using recurrent neural networksabstractIt is undeniable that the ability to grasp and handle an object is vital for service robots. From object recognition to object grasping motion, the motion execution should be as fast as possible. Due to the possible position variation of the target object to be grasped, online planning of grasping motion should be done. In order to achieve flexible grasping motion, recurrent neural network could be implemented as an alternative to conventional manipulation method which is based on kinematic and dynamic analysis. However, the application of recurrent neural network model requires good and easily obtainable training data. Hence, a novel robotic arm design with high flexibility is proposed to facilitate the training and implementation of the recurrent neural network model. The feasibility of the proposed robotic arm design is evaluated via the training, learning and testing of stochastic continuous time recurrent neural network (S-CTRNN) model with grasping a box motion. Boon Hwa Tan, Huajin Tang, Rui Yan 0005, Jun Tani |
IROS | 2 |
| 2014 | Vision enhanced neuro-cognitive structure for robotic spatial cognition
Huajin Tang |
Neurocomputing | 2 |
| 2014 | Guest editorial: Special issue on brain inspired models of cognitive memory
Huajin Tang, Kiruthika Ramanathan, Ning Ning 0001 |
Neurocomputing | 1 |
| 2014 | A brain-inspired spiking neural network model with temporal encoding and learning
Qiang Yu 0005, Huajin Tang, Kay Chen Tan, Haoyong Yu |
Neurocomputing | 2 |
| 2014 | Real-Time Keypoint Recognition Using Restricted Boltzmann MachineabstractFeature point recognition is a key component in many vision-based applications, such as vision-based robot navigation, object recognition and classification, image-based modeling, and augmented reality. Real-time performance and high recognition rates are of crucial importance to these applications. In this brief, we propose a novel method for real-time keypoint recognition using restricted Boltzmann machine (RBM). RBMs are generative models that can learn probability distributions of many different types of data including labeled and unlabeled data sets. Due to the inherent noise of the training data sets, we use an RBM to model statistical distributions of the training data. Furthermore, the learned RBM can be used as a competitive classifier to recognize the keypoints in real-time during the tracking stage, thus making it advantageous to be employed in applications that require real-time performance. Experiments have been conducted under a variety of conditions to demonstrate the effectiveness and generalization of the proposed approach. Miaolong Yuan, Huajin Tang, Haizhou Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2013 | Temporal coding of local spectrogram features for robust sound recognitionabstractThere is much evidence to suggest that the human auditory system uses localised time-frequency information for the robust recognition of sounds. Despite this, conventional systems typically rely on features extracted from short windowed frames over time, covering the whole frequency spectrum. Such approaches are not inherently robust to noise, as each frame will contain a mixture of the spectral information from noise and signal. Here, we propose a novel approach based on the temporal coding of Local Spectrogram Features (LSFs), which generate spikes that are used to train a Spiking Neural Network (SNN) with temporal learning. LSFs represent robust location information in the spectrogram surrounding keypoints, which are detected in a signal-driven manner such that the effect of noise on the temporal coding is reduced. Our experiments demonstrate the robust performance of our approach across a variety of noise conditions, such that it is able to outperform the conventional frame-based baseline methods. Jonathan William Dennis, Qiang Yu 0005, Huajin Tang, Tran Huy Dat, Haizhou Li 0001 |
ICASSP | 3 |
| 2013 | A Simplified Cerebellum-Based Model for Motor Control in Brain Based Devices
Vui Ann Shim, Chris Stephen Naveen Ranjit, Huajin Tang |
ICONIP (1) | 4 |
| 2013 | A hierarchical organized memory model using spiking neuronsabstractThe recent identification of neural cliques, which are network-level memory coding units in the hippocampus, enables population codes to be the neuronal representation of memory. It has been discovered that the timing of spikes plays an important role in the neural computation and information processing in the brain. Moreover, these memory-coding units have been observed organizing in a hierarchical manner in the brain. Inspired by these exciting findings, we present a hierarchically organized memory model with spiking neurons, which can store both associative memory and episodic memory with temporal population codes. The basic structure of the hierarchical model is composed of three layers with different functions and can be extended to more complicated networks by duplicating and connecting the basic three-layer network. With a spike-timing based learning algorithm, the spiking neural network with theta and gamma oscillations is able to store spatiotemporal memory items within gamma cycles, and links these memories into a sequence. The spiking-timing-dependent plasticity (STDP) contributes to the formation of both associative memory and episodic memory via fast and slow N-methyl-D-aspartate (NMDA) channels, respectively. Huajin Tang, Kay Chen Tan |
IJCNN | 2 |
| 2013 | A computationally efficient associative memory model of hippocampus CA3 by spiking neuronsabstractThe hippocampus is involved with the storage and retrieval of short-term associative memories. In this paper, we propose a computationally efficient associative memory model of the hippocampus CA3 region by spiking neurons, and explores the storage of auto-associative memory. The spiking neural network encodes different associative memories by different subsets of the principal neurons. These memory items are activated in different gamma subcycles, and auto-associative memory is maintained by the synaptic modifications of recurrent collaterals by N-methyl-D-aspartate (NMDA) channels. Accurate formation of auto-associative memory is achievable in single presentation of memory items when synaptic modifications depend on fast NMDA channels having a deactivation time within the duration of a gamma subcycle. Simulation results also show that spike response model (SRM) improves computational efficiency over the integrate-and-fire (I&F) neuron model. Chin Hiong Tan, Huajin Tang, Eng Yeow Cheu |
IJCNN | 2 |
| 2013 | RGB-D based cognitive map building and navigationabstractThis paper describes a cognitive map building and navigation system using an RGB-D sensor for mobile robots. A brain-inspired simultaneously localization and mapping (SLAM) system, that requires raw odometry data and RGB-D information, is used to construct a spatial cognitive map of an office environment. The cognitive map contains a set of spatial coordinates that the robot has traveled. A global path is extracted from the built cognitive map and subsequently used by a local planner to instruct the robot to navigate. The global path is a subset of the path that builds up the cognitive map. This is different from other path planning mechanisms that construct a path based on a ground-truth map. Experiment results show that the employment of the RGB-D sensor significantly improves the mapping results. Vui Ann Shim, Miaolong Yuan, Chithra Srinivasan, Huajin Tang, Haizhou Li 0001 |
IROS | 5 |
| 2013 | Dynamical properties of continuous attractor neural network with background tuning
Huajin Tang, Haizhou Li 0001, Luping Shi |
Neurocomputing | 2 |
| 2013 | Continuous attractors of discrete-time recurrent neural networks
Huajin Tang, Haizhou Li 0001 |
Neural Comput. Appl. | 2 |
| 2013 | A Spike-Timing-Based Integrated Model for Pattern RecognitionabstractDuring the past few decades, remarkable progress has been made in solving pattern recognition problems using networks of spiking neurons. However, the issue of pattern recognition involving computational process from sensory encoding to synaptic learning remains underexplored, as most existing models or algorithms target only part of the computational process. Furthermore, many learning algorithms proposed in the literature neglect or pay little attention to sensory information encoding, which makes them incompatible with neural-realistic sensory signals encoded from real-world stimuli. By treating sensory coding and learning as a systematic process, we attempt to build an integrated model based on spiking neural networks (SNNs), which performs sensory neural encoding and supervised learning with precisely timed sequences of spikes. With emerging evidence of precise spike-timing neural activities, the view that information is represented by explicit firing times of action potentials rather than mean firing rates has been receiving increasing attention. The external sensory stimulation is first converted into spatiotemporal patterns using a latency-phase encoding method and subsequently transmitted to the consecutive network for learning. Spiking neurons are trained to reproduce target signals encoded with precisely timed spikes. We show that when a supervised spike-timing-based learning is used, different spatiotemporal patterns are recognized by different spike patterns with a high time precision in milliseconds. Huajin Tang, Kay Chen Tan, Haizhou Li 0001, Luping Shi |
Neural Comput. | 2 |
| 2013 | Dynamics Analysis of a Population Decoding ModelabstractInformation processing in the nervous system involves the activity of large populations of neurons. It is difficult to extract information from these population codes because of the noise inherent in neuronal responses. We propose a divisive normalization model to read the population codes. The dynamics of the model are analyzed by continuous attractor theory. Under certain conditions, the model possesses continuous attractors. Moreover, the explicit expressions of the continuous attractors are provided. Simulations are employed to illustrate the theory. Huajin Tang, Haizhou Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2013 | Rapid Feedforward Computation by Temporal Encoding and Learning With Spiking NeuronsabstractPrimates perform remarkably well in cognitive tasks such as pattern recognition. Motivated by recent findings in biological systems, a unified and consistent feedforward system network with a proper encoding scheme and supervised temporal rules is built for solving the pattern recognition task. The temporal rules used for processing precise spiking patterns have recently emerged as ways of emulating the brain's computation from its anatomy and physiology. Most of these rules could be used for recognizing different spatiotemporal patterns. However, there arises the question of whether these temporal rules could be used to recognize real-world stimuli such as images. Furthermore, how the information is represented in the brain still remains unclear. To tackle these problems, a proper encoding method and a unified computational model with consistent and efficient learning rule are proposed. Through encoding, external stimuli are converted into sparse representations, which also have properties of invariance. These temporal patterns are then learned through biologically derived algorithms in the learning layer, followed by the final decision presented through the readout layer. The performance of the model with images of digits from the MNIST database is presented. The results show that the proposed model is capable of recognizing images correctly with a performance comparable to that of current benchmark algorithms. The results also suggest a plausibility proof for a class of feedforward models of rapid and robust recognition in the brain. Qiang Yu 0005, Huajin Tang, Kay Chen Tan, Haizhou Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2012 | Learning real-world stimuli by single-spike coding and tempotron ruleabstract10.1109/IJCNN.2012.6252369 Huajin Tang, Qiang Yu 0005, Kay Chen Tan |
IJCNN | 1 |
| 2012 | Pattern recognition computation in a spiking neural network with temporal encoding and learningabstractMany conventional methods have been widely studied to solve the pattern recognition task, but most of them lack the biological plausibility. This paper presents a spiking neural network of integrate-and-fire neurons to perform pattern recognition. A biologically plausible supervised synaptic learning rule is used so that neurons can efficiently make a decision. The whole system contains encoding, learning and readout. It can classify complex patterns of activities stored in a vector, as well as the real-world stimuli. We test the performance of the network with digital images from the MNIST and images of alphabetic letters. It turns out to be able to classify these patterns correctly. In addition, the synaptic dynamics is shown to be compatible with many experimental observations on induction of long-term modifications, like spike-timing-dependent plasticity (STDP). Qiang Yu 0005, Kay Chen Tan, Huajin Tang |
IJCNN | 3 |
| 2012 | Neural Networks and Learning Systems Come TogetherabstractThis issue marks the beginning of the IEEE Transactions on Neural Networks and Learning Systems (TNNLS). By adding "Learning Systems" to the title, we now state explicitly the scope of the Transactions to include neural networks as well as related learning systems. This issue marks a new era in the history of our Transactions. The Transactions is now ready to face the challenges of the next 10-20 years. With the evolution of the fields of neural networks in particular and computational intelligence in general, the IEEE Transactions on Neural Networks and Learning Systems will continue to grow and to succeed in this ever-changing world. Also included are a few comments about the review process of TNN manuscripts and the introduction of 14 new TNNLS Associate Editors. Short biographies are included for the new Associate Editors. Bart Baesens, Pantelis Bouboulis, Sergio Cruces, Carlotta Domeniconi, Shiro Ikeda, Xuelong Li 0001, Patricia Melin, Vadrevu Sree Hari Rao, Björn W. Schuller, Huajin Tang, Cong Wang 0033, Jian Yang 0003, Derong Zhao, Derong Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 11 |
| 2011 | A Neuro-cognitive Robot for Spatial Navigation
Huajin Tang, Chin Hiong Tan |
ICONIP (1) | 2 |
| 2011 | Associative Memory Model of Hippocampus CA3 Using Spike Response Neurons
Chin Hiong Tan, Eng Yeow Cheu, Qiang Yu 0005, Huajin Tang |
ICONIP (1) | 5 |
| 2011 | Brain Inspired Cognitive System for Learning and Memory
Huajin Tang |
ICONIP (1) | 1 |
| 2010 | Restricted Boltzmann machine based algorithm for multi-objective optimizationabstractRestricted Boltzmann machine is an energy-based stochastic neural network with unsupervised learning. This network consists of a layer of hidden unit and visible unit in an undirected generative network. In this paper, restricted Boltzmann machine is modeled as estimation of distribution algorithm in the context of multi-objective optimization. The probabilities of the joint configuration over the visible and hidden units in the network are trained until the distribution over the global state reach a certain degree of thermal equilibrium. Subsequently, the probabilistic model is constructed using the energy function of the network. Moreover, the proposed algorithm incorporates clustering in phenotype space and other canonical operators. The effects on the stability of the trained network and clustering in optimization are rigorously examined. Experimental investigations are conducted to analyze the performance of the algorithm in scalable problems with high numbers of objective functions and decision variables. Huajin Tang, Vui Ann Shim, Kay Chen Tan, Jun Yong Chia |
IEEE Congress on Evolutionary Computation | 1 |
| 2010 | Memory Dynamics in Attractor Networks with Saliency WeightsabstractMemory is a fundamental part of computational systems like the human brain. Theoretical models identify memories as attractors of neural network activity patterns based on the theory that attractor (recurrent) neural networks are able to capture some crucial characteristics of memory, such as encoding, storage, retrieval, and long-term and working memory. In such networks, long-term storage of the memory patterns is enabled by synaptic strengths that are adjusted according to some activity-dependent plasticity mechanisms (of which the most widely recognized is the Hebbian rule) such that the attractors of the network dynamics represent the stored memories. Most of previous studies on associative memory are focused on Hopfield-like binary networks, and the learned patterns are often assumed to be uncorrelated in a way that minimal interactions between memories are facilitated. In this letter, we restrict our attention to a more biological plausible attractor network model and study the neuronal representations of correlated patterns. We have examined the role of saliency weights in memory dynamics. Our results demonstrate that the retrieval process of the memorized patterns is characterized by the saliency distribution, which affects the landscape of the attractors. We have established the conditions that the network state converges to unique memory and multiple memories. The analytical result also holds for other cases for variable coding levels and nonbinary levels, indicating a general property emerging from correlated memories. Our results confirmed the advantage of computing with graded-response neurons over binary neurons (i.e., reducing of spurious states). It was also found that the nonuniform saliency distribution can contribute to disappearance of spurious states when they exit. Huajin Tang, Haizhou Li 0001, Rui Yan 0005 |
Neural Comput. | 1 |
| 2010 | A discrete-time neural network for optimization problems with hybrid constraintsabstractRecurrent neural networks have become a prominent tool for optimizations including linear or nonlinear variational inequalities and programming, due to its regular mathematical properties and well-defined parallel structure. This brief presents a general discrete-time recurrent network for linear variational inequalities and related optimization problems with hybrid constraints. In contrary to the existing discrete-time networks, this general model can operate not only on bound constraints, but also on hybrid constraints comprised of inequality, equality and bound constraints. The model has dynamical properties of global convergence, asymptotical and exponential convergences under some weaker conditions. Numerical examples demonstrate its efficacy and performance. Huajin Tang, Haizhou Li 0001, Zhang Yi 0001 |
IEEE Trans. Neural Networks | 1 |
| 2009 | Analysis of Continuous Attractors for 2-D Linear Threshold Neural NetworksabstractThis brief investigates continuous attractors of the well-developed model in visual cortex, i.e., the linear threshold (LT) neural networks, based on a parameterized 2-D model. On the basis of existing results on nondegenerate equilibria in mathematics, we further discuss degenerate equilibria for such networks and present properties and distributions of the equilibria, which enables us to draw the coexistence conditions of nondegenerate and degenerate equilibria (e.g., singular lines). Our theoretical results provide a useful framework for precise tuning on the network parameters, e.g., the feedbacks and visual inputs. Simulations are also presented to illustrate the theoretical findings. Lan Zou, Huajin Tang, Kay Chen Tan, Weinian Zhang |
IEEE Trans. Neural Networks | 2 |
| 2009 | Nontrivial Global Attractors in 2-D Multistable Attractor Neural NetworksabstractAttractor dynamics is a crucial problem for attractor neural networks, as it is the underling computational mechanism for memory storage and retrieval in neural systems. This brief studies a class of attractor network consisting of linearized threshold neurons, and analyzes global attractors based on a parameterized 2-D model. On the basis of previous results on nondegenerate and degenerate equilibria in mathematics, we further elucidate all possible nontrivial global attractors. Our theoretical result provides precise descriptions on how the changes of network parameters affect the attractors' distribution and landscape, and it may give a feasible solution towards specifying attractors by specifying weights. Simulations are presented to illustrate the theoretical results. Lan Zou, Huajin Tang, Kay Chen Tan, Weinian Zhang |
IEEE Trans. Neural Networks | 2 |
| 2008 | An asynchronous recurrent linear threshold network approach to solving the traveling salesman problem
Eu Jin Teoh, Kay Chen Tan, Huajin Tang, Cheng Xiang 0001, Chi Keong Goh |
Neurocomputing | 3 |
| 2006 | A Columnar Competitive Model with Simulated Annealing for Solving Combinatorial Optimization ProblemsabstractOne of the major drawbacks of the Hopfield network is that when it is applied to certain polytopes of combinatorial problems, such as the traveling salesman problem (TSP), the obtained solutions are often invalid, requiring numerous trial-and-error setting of the network parameters thus resulting in low-computation efficiency. With this in mind, this article presents a columnar competitive model (CCM) which incorporates a winner-takes-all (WTA) learning rule for solving the TSP. Theoretical analysis for the convergence of the CCM shows that the competitive computational neural network guarantees the convergence of the network to valid states and avoids the tedious procedure of determining the penalty parameters. In addition, its intrinsic competitive learning mechanism enables a fast and effective evolving of the network. Simulation results illustrate that the competitive model offers more and better valid solutions as compared to the original Hopfield network. Eu Jin Teoh, Huajin Tang, Kay Chen Tan |
IJCNN | 2 |
| 2006 | An Improvement on Competitive Neural Networks Applied to Image Segmentation
Rui Yan 0005, Meng Joo Er, Huajin Tang |
ISNN (2) | 3 |
| 2006 | Dynamics analysis and analog associative memory of networks with LT neuronsabstractThe additive recurrent network structure of linear threshold neurons represents a class of biologically-motivated models, where nonsaturating transfer functions are necessary for representing neuronal activities, such as that of cortical neurons. This paper extends the existing results of dynamics analysis of such linear threshold networks by establishing new and milder conditions for boundedness and asymptotical stability, while allowing for multistability. As a condition for asymptotical stability, it is found that boundedness does not require a deterministic matrix to be symmetric or possess positive off-diagonal entries. The conditions put forward an explicit way to design and analyze such networks. Based on the established theory, an alternate approach to study such networks is through permitted and forbidden sets. An application of the linear threshold (LT) network is analog associative memory, for which a simple design method describing the associative memory is suggested in this paper. The proposed design method is similar to a generalized Hebbian approach, but with distinctions of additional network parameters for normalization, excitation and inhibition, both on a global and local scale. The computational abilities of the network are dependent on its nonlinear dynamics, which in turn is reliant upon the sparsity of the memory vectors. Huajin Tang, Kay Chen Tan, Eu Jin Teoh |
IEEE Trans. Neural Networks | 1 |
| 2005 | Analysis of Cyclic Dynamics for Networks of Linear Threshold NeuronsabstractThe network of neurons with linear threshold (LT) transfer functions is a prominent model to emulate the behavior of cortical neurons. The analysis of dynamic properties for LT networks has attracted growing interest, such as multistability and boundedness. However, not much is known about how the connection strength and external inputs are related to oscillatory behaviors. Periodic oscillation is an important characteristic that relates to nondivergence, which shows that the network is still bounded although unstable modes exist. By concentrating on a general parameterized two-cell network, theoretical results for geometrical properties and existence of periodic orbits are presented. Although it is restricted to two-dimensional systems, the analysis can provide a useful contribution to analyze cyclic dynamics of some specific LT networks of high dimension. As an application, it is extended to an important class of biologically motivated networks of large scale: the winner-take-all model using local excitation and global inhibition. Huajin Tang, Kay Chen Tan, Weinian Zhang |
Neural Comput. | 1 |
| 2004 | Dynamical optimal learning for FNN and its applicationsabstractThis work presents a new dynamical optimal learning (DOL) algorithm for three-layer linear neural networks and investigates its generalization ability. The optimal learning rates can be fully determined during the training process. The mean squared error is guaranteed to be stably decreased and the learning is less sensitive to initial parameter settings. The simulation results illustrate that the proposed DOL algorithm gives better generalization performance and faster convergence as compared to standard error back propagation algorithm. Huajin Tang, Kay Chen Tan, Tong Heng Lee |
FUZZ-IEEE | 1 |
| 2004 | Global exponential stability of discrete-time neural networks for constrained quadratic optimization
Kay Chen Tan, Huajin Tang, Z. Yi |
Neurocomputing | 2 |
| 2004 | New dynamical optimal learning for linear multilayer FNNabstractThis letter presents a new dynamical optimal learning (DOL) algorithm for three-layer linear neural networks and investigates its generalization ability. The optimal learning rates can be fully determined during the training process. The mean squared error (mse) is guaranteed to be stably decreased and the learning is less sensitive to initial parameter settings. The simulation results illustrate that the proposed DOL algorithm gives better generalization performance and faster convergence as compared to standard error back propagation algorithm. Kay Chen Tan, Huajin Tang |
IEEE Trans. Neural Networks | 2 |
| 2004 | A columnar competitive model for solving combinatorial optimization problemsabstractThe major drawbacks of the Hopfield network when it is applied to some combinatorial problems, e.g., the traveling salesman problem (TSP), are invalidity of the obtained solutions, trial-and-error setting value process of the network parameters and low-computation efficiency. This letter presents a columnar competitive model (CCM) which incorporates winner-takes-all (WTA) learning rule for solving the TSP. Theoretical analysis for the convergence of the CCM shows that the competitive computational neural network guarantees the convergence to valid states and avoids the onerous procedures of determining the penalty parameters. In addition, its intrinsic competitive learning mechanism enables a fast and effective evolving of the network. The simulation results illustrate that the competitive model offers more and better valid solutions as compared to the original Hopfield network. Huajin Tang, Kay Chen Tan, Zhang Yi 0001 |
IEEE Trans. Neural Networks | 1 |