Malu Zhang

dblp:156/7882 · DBLP profile ↗
← Back
94ranked-venue papers
7as first author
79since 2021 · last 2026
0000-0002-2345-0974ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 81 · 7 first-author · 66 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 1 first-author · 31 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards Training-Free and Accurate ANN-to-SNN Conversion via Activation-Aware Redistribution
abstract
Conversion represents an effective approach for obtaining low-power models by transforming Artificial Neural Networks (ANNs) into event-driven Spiking Neural Networks (SNNs) without additional training. However, existing training-free conversion methods often incur substantial conversion errors. Here, we first reveal that these conversion errors primarily arise from a distributional mismatch, as the activation distributions of ANNs exhibit channel-wise shifts and scaling, whereas spike rates lack corresponding channel-specific characteristics. To address this limitation, we propose Adaptive Integrate-and-Fire (AIF) neurons with channel-specific thresholds and membrane-potential offsets that dynamically adjust spike rates. These parameters are optimized to jointly minimize conversion errors and maximize information entropy, enabling AIF neurons to capture the activation distribution characteristics of the original ANN. Moreover, AIF neurons can be seamlessly integrated into Transformer architectures with only negligible additional computational cost. Our method achieves state-of-the-art results on multiple vision and natural language processing benchmarks, in particular attaining a notable top-1 accuracy of 85.52% on ImageNet-1K.
Honglin Cao, Shuai Wang 0058, Zijian Zhou 0005, Ammar Belatreche, Wenjie Wei, Malu Zhang, Haizhou Li 0001
AAAI9
2026 HardF-SNN: Hardware-Friendly Quantization for Spiking Neural Networks with Efficient Integer-Arithmetic-Only Inference
abstract
Spiking Neural Networks (SNNs) are emerging as a promising energy-efficient alternative to Artificial Neural Networks (ANNs) due to their event-driven computation paradigm. However, recent advances toward large-scale high-performance SNNs inevitably lead to substantial memory and computational overhead. While quantization offers a potential way, many quantization approaches fail to deliver verifiable efficiency gains on resource-constrained hardware platforms. In this paper, we propose a lightweight and hardware-friendly SNN, termed HardF-SNN. Specifically, we first build a baseline model using shared-scale quantization and BN folding to simulate integer-only inference, as this has not been thoroughly discussed in prior SNN works. Then, through empirical and theoretical analysis, we identify that the baseline suffers from accuracy degradation and may cause training failure. To mitigate these issues, we propose proportional shared-scale quantization for enhanced dynamic range and integer-only BN using bit-shifting to stabilize training. Extensive experiments show that HardF-SNN achieves an optimal balance between performance and efficiency with excellent hardware compatibility. To demonstrate its effectiveness on resource-limited platforms, HardF-SNN is deployed on a dedicated FPGA-based hardware accelerator. Evaluation results indicate that our implementation achieves significant performance improvements over several existing hardware accelerators.
Jieyuan Zhang, Yimeng Shan, Jibin Wu, Wenyu Chen 0001, Malu Zhang
AAAI7
2026 Training-Free ANN-to-SNN Conversion for High-Performance Spiking Transformers
abstract
Leveraging the event-driven paradigm, Spiking Neural Networks (SNNs) offer a promising approach for constructing energy-efficient Transformer architectures. Compared to directly trained Spiking Transformers, ANN-to-SNN conversion methods bypass the high training costs. However, existing methods still suffer from notable limitations, failing to effectively handle nonlinear operations in Transformer architectures and requiring additional fine-tuning processes for pre-trained ANNs. To address these issues, we propose a high-performance and training-free ANN-to-SNN conversion framework tailored for Transformer architectures. Specifically, we introduce a Multi-basis Exponential Decay (MBE) neuron, which employs an exponential decay strategy and multi-basis encoding method to efficiently approximate various nonlinear operations. It removes the requirement for weight modifications in pre-trained ANNs. Extensive experiments across diverse tasks (CV, NLU, NLG) and mainstream Transformer architectures (ViT, RoBERTa, GPT-2) demonstrate that our method achieves near-lossless conversion accuracy with significantly lower latency. This provides a promising pathway for the efficient and scalable deployment of Spiking Transformers in real-world applications.
Wenjie Wei, Dehao Zhang, Shuai Wang 0058, Qian Sun 0014, Jieyuan Zhang, Malu Zhang
AAAI10
2026 GASE: Generalized adaptive static enhancement for temporal sentence grounding
Ran Ran 0001, Kaiwen Shen, Jiwei Wei, Ruikun Chai, Shiyuan He, Zeyu Ma 0002, Malu Zhang, Yang Yang 0002
Knowl. Based Syst.8
2026 Spiking neural networks for EEG signal analysis: From theory to practice
Siqi Cai 0002, Zheyuan Lin, Wenjie Wei, Shuai Wang 0058, Malu Zhang, Tanja Schultz, Haizhou Li 0001
Neural Networks6
2026 Lightweight and Personalized Single-Eye Emotion Recognition via CNN-SNN Spatiotemporal Learning and Memory-Inferred Event Features
abstract
Emotion recognition is essential for improving user experience and interaction quality in human-centered applications. While recent studies have leveraged both event and traditional cameras to enhance eye-based emotion recognition, their practical deployment is hindered by the scarcity of event cameras and the complexity of dual-modality frameworks. Personalization, which is critical for handling individual differences in emotional expression, is also affected by these factors, resulting in reduced performance and adaptation efficiency. To address these challenges, we propose a lightweight and personalized single-eye emotion recognition network, called LPSEER. LPSEER introduces a novel hybrid neural architecture that integrates a convolutional neural network (CNN) and a spiking neural network (SNN) to capture spatiotemporal features from video frames and events, respectively. Additionally, we design a memorybased event feature inference (MEFI) module that recalls event features from video frames, eliminating the reliance on event cameras during inference and personalization while retaining the discriminative advantages of event-based representations. Experimental results demonstrate that LPSEER achieves state-of- the-art recognition accuracy while maintaining the smallest model size and lowest computational cost. Further experiments confirm the strong generalization capabilities and the ability to achieve faster, more accurate personalization. These advantages collectively enable lightweight, accurate, and efficient emotion recognition for real-world human-centered applications.
Qianhui Liu, Jiqing Zhang, Yang Wang 0106, Malu Zhang, Xin Yang 0011, Gang Pan 0001, Haizhou Li 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Spike-Driven Lightweight Large Language Model With Evolutionary Computation
abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks, but their deployment in resource-constrained environments remains challenging due to substantial memory and computational requirements. Benefiting from the sparse event-driven computation paradigm of Spiking Neural Networks (SNNs), some research has focused on designing spike-based language models. However, existing spike-based language models achieve only partial computational efficiency gains and fail to address memory constraints comprehensively. In this paper, we propose an evolved and quantized spike-driven language model (EQ-SpikeLM) to address identified challenges. This model incorporates two primary innovations. First, inspired by the artificial bee colony algorithm in evolutionary computation, we propose an architecture evolution method, namely ABC-Arc. This method optimizes network topology by systematically removing redundant neural pathways. Second, a dynamic post-training quantization (DynPTQ) strategy is developed for the evolved SpikeLM, facilitating the conversion of floating-point parameters to lower-bit precision without requiring model retraining. By combining these two methods, EQ-SpikeLM significantly reduces storage and computational demands while preserving model performance. Experimental evaluation on the GLUE benchmark demonstrates EQ-SpikeLM’s ability to maintain performance equivalent to its uncompressed counterpart, with a substantial reduction in both model size and power consumption. These results position EQ-SpikeLM as a viable approach for deploying large language models in resource-constrained edge computing scenarios.
Malu Zhang, Wenjie Wei, Zijian Zhou 0005, Wanlong Liu, Jie Zhang 0118, Ammar Belatreche, Yang Yang 0002
IEEE Trans. Evol. Comput.1
2026 SNN-FT: Temporal-Coded Spiking Neural Networks for Fourier Transform
abstract
The Fourier transform (FT) stands as a fundamental tool in modern signal processing with widespread applications across various scientific and engineering fields. Therefore, there remains a need for continued research efforts to devise energy-efficient implementations of the FT. Due to their inherent energy efficiency, biologically plausible spiking neural networks (SNNs) emerge as a promising alternative solution. However, current SNN implementations of the FT suffer from two key shortcomings, namely, high latency and reduced accuracy. In this article, we analyze the underlying causes of these limitations and highlight deficiencies in the existing spike-based encoding mechanisms and spiking neuron models. We then propose a new SNN-based FT (SNN-FT) based on a logarithmically polarized time-to-first-spike (TTFS) encoding method (called LP-TTFS) along with a novel piecewise spiking neuron (PTSN) model based on ternary spikes (referred to as PTSN). The resulting SNN-FT is mathematically equivalent to the conventional FT and demonstrates superior performance in accuracy as well as reduced latency. We assess the performance of the proposed SNN-FT alternative through extensive experiments on FT-based applications, such as radar and audio signal processing, and the obtained results demonstrate the efficacy of SNN-FT and its superiority over the existing approaches. This study unveils a novel energy-efficient neuromorphic computing technique with great potential for FT applications across diverse scientific and engineering domains.
Shuai Wang 0058, Haorui Zheng, Ammar Belatreche, Guoqing Wang 0001, Yeying Jin, Jibin Wu, Malu Zhang, Yang Yang 0002, Haizhou Li 0001
IEEE Trans. Neural Networks Learn. Syst.8
2026 SFedCA: Credit Assignment-Based Active Client Selection Strategy for Spiking Federated Learning
abstract
The spiking federated learning (FL) is an emerging distributed learning paradigm that allows resource-constrained devices to train collaboratively at low power consumption without exchanging local data. It takes advantage of both the privacy computation property in FL and the energy efficiency in spiking neural networks (SNNs). However, existing spiking FL methods employ a random selection approach for client aggregation, assuming unbiased client participation. This neglect of statistical heterogeneity significantly affects the convergence and precision of the global model. In this work, we propose a credit assignment-based active client selection strategy for spiking federated learning, the SFedCA, to aggregate clients contributing to the global sample distribution balance judiciously. Specifically, the client credits are assigned by the firing intensity state before and after local model training, which reflects the difference in local data distribution from the global model. The comprehensive experiments are conducted on various non-identical and independent distribution (non-IID) scenarios. The experimental results demonstrate that the SFedCA outperforms the existing state-of-the-art spiking FL methods and requires fewer communication rounds.
Qiugang Zhan, Jinbo Cao, Xiurui Xie, Huajin Tang, Malu Zhang, Shantian Yang, Guisong Liu
IEEE Trans. Neural Networks Learn. Syst.5
2026 Inhibiting Error Exacerbation in Offline Reinforcement Learning With Data Sparsity
abstract
Offline reinforcement learning (RL) aims to learn effective agents from previously collected datasets, facilitating the safety and efficiency of RL by avoiding real-time interaction. However, in practical applications, the approximation error of the out-of-distribution (OOD) state-actions can cause considerable overestimation due to error exacerbation during training, finally degrading the performance. In contrast to prior works that merely addressed the OOD state-actions, we discover that all data introduces estimation error whose magnitude is directly related to data sparsity. Consequently, the impact of data sparsity is inevitable and vital when inhibiting the error exacerbation. In this article, we propose an offline RL approach to inhibit error exacerbation with data sparsity (IEEDS), which includes a novel value estimation method to consider the impact of data sparsity on the training of agents. Specifically, the value estimation phase includes two innovations: 1) replace Q-net with V-net, a smaller and denser state space makes data more concentrated, contributing to more accurate value estimation and 2) introduce state sparsity to the training by design state-aware-sparsity Markov decision process (MDP), further lessening the impact of sparse states. We theoretically prove the convergence of IEEDS under state-aware-sparsity MDP. Extensive experiments on offline RL benchmarks reveal that IEEDS's superior performance.
Fan Zhang 0068, Malu Zhang, Wenyu Chen 0001, Siying Wang 0002, Yang Yang 0002
IEEE Trans. Neural Networks Learn. Syst.2
2025 Towards Accurate Binary Spiking Neural Networks: Learning with Adaptive Gradient Modulation Mechanism
abstract
Binary Spiking Neural Networks (BSNNs) inherit the event-driven paradigm of SNNs, while also adopting the reduced storage burden of binarization techniques. These distinct advantages grant BSNNs lightweight and energy-efficient characteristics, rendering them ideal for deployment on resource-constrained edge devices. However, due to the binary synaptic weights and non-differentiable spike function, effectively training BSNNs remains an open question. In this paper, we conduct an in-depth analysis of the challenge for BSNN learning, namely the frequent weight sign flipping problem. To mitigate this issue, we propose an Adaptive Gradient Modulation Mechanism (AGMM), which is designed to reduce the frequency of weight sign flipping by adaptively adjusting the gradients during the learning process. The proposed AGMM can enable BSNNs to achieve faster convergence speed and higher accuracy, effectively narrowing the gap between BSNNs and their full-precision equivalents. We validate AGMM on both static and neuromorphic datasets, and results indicate that it achieves state-of-the-art results among BSNNs. This work substantially reduces storage demands and enhances SNNs' inherent energy efficiency, making them highly feasible for resource-constrained environments.
Wenjie Wei, Ammar Belatreche, Honglin Cao, Zijian Zhou 0005, Shuai Wang 0058, Malu Zhang, Yang Yang 0002
AAAI7
2025 Advancing Spiking Neural Networks Towards Multiscale Spatiotemporal Interaction Learning
abstract
Recent advancements in neuroscience research have propelled the development of Spiking Neural Networks (SNNs), which not only have the potential to further advance neuroscience research but also serve as an energy-efficient alternative to Artificial Neural Networks (ANNs) due to their spike-driven characteristics. However, previous studies often overlooked the multiscale information and its spatiotemporal correlation between event data, leading SNN models to approximate each frame of input events as static images. We hypothesize that this oversimplification significantly contributes to the performance gap between SNNs and traditional ANNs. To address this issue, we have designed a Spiking Multiscale Attention (SMA) module that captures multiscale spatiotemporal interaction information. Furthermore, we developed a regularization method named Attention ZoneOut (AZO), which utilizes spatiotemporal attention weights to reduce the model's generalization error through pseudo-ensemble training. Our approach has achieved state-of-the-art results on mainstream neuromorphic datasets. Additionally, we have reached a performance of 77.1\% on the Imagenet-1K dataset using a 104-layer ResNet architecture enhanced with SMA and AZO. This achievement confirms the state-of-the-art performance of SNNs with non-transformer architectures and underscores the effectiveness of our method in bridging the performance gap between SNN models and traditional ANN models.
Yimeng Shan, Malu Zhang, Rui-Jie Zhu 0003, Xuerui Qiu, Jason Kamran Eshraghian, Haicheng Qu
AAAI2
2025 Leveraging Asynchronous Spiking Neural Networks for Ultra Efficient Event-Based Visual Processing
abstract
Event cameras encode visual information by generating asynchronous and sparse event streams, which hold great potential for low latency and low power consumption. Despite many successful implementations of event camera-based applications, most of them accumulate the events into frames and then utilize conventional frame-based computer vision algorithms. These frame-based methods, though typically effective, diminish the inherent advantages of the event camera's low latency and low power consumption. To solve the above problems, we propose ASGCN, which efficiently processes data on an event-by-event basis and dynamically evolves into a corresponding dynamic representation, enabling low latency and high sparsity of data representation. The sparsity computation is further improved by introducing brain-inspired spiking neural networks, resulting in low power consumption for ASGCN. Extensive and diverse experiments demonstrate the energy efficiency and low latency advantages of our processing pipeline. Especially on real-world event camera datasets, our pipeline consumes more than 10,000 times less energy and achieves similar performance compared to current frame-based methods.
Dingyi Zeng, Honglin Cao, Wanlong Liu, Yichen Xiao, Chengzhuo Lu, Wenyu Chen 0001, Malu Zhang, Guoqing Wang 0001, Yang Yang 0002
AAAI8
2025 A Compressive Memory-based Retrieval Approach for Event Argument Extraction
abstract
Recent works have demonstrated the effectiveness of retrieval augmentation in the Event Argument Extraction (EAE) task. However, existing retrieval-based EAE methods have two main limitations: (1) input length constraints and (2) the gap between the retriever and the inference model. These issues limit the diversity and quality of the retrieved information. In this paper, we propose a Compressive Memory-based Retrieval (CMR) mechanism for EAE, which addresses the two limitations mentioned above. Our compressive memory, designed as a dynamic matrix that effectively caches retrieved information and supports continuous updates, overcomes the limitations of input length. Additionally, after pre-loading all candidate demonstrations into the compressive memory, the model further retrieves and filters relevant information from the memory based on the input query, bridging the gap between the retriever and the inference model. Extensive experiments show that our method achieves new state-of-the-art performance on three public datasets (RAMS, WikiEvents, ACE05), significantly outperforming existing retrieval-based EAE methods.
Wanlong Liu, Enqi Zhang, Shaohuan Cheng, Dingyi Zeng, Li Zhou 0010, Chen Zhang 0020, Malu Zhang, Wenyu Chen 0001
COLING7
2025 Rethinking Spiking Self-Attention Mechanism: Implementing a-XNOR Similarity Calculation in Spiking Transformers
abstract
Transformers significantly raise the performance limits across various tasks, spurring research into integrating them into spiking neural networks. However, a notable performance gap remains between existing spiking Transformers and their artificial neural network counterparts. Here, we first analyze the cause of this gap and attribute it to the dot product’s ineffectiveness in measuring similarity between spiking queries and keys, due to numerous non-spiking events. To address this, we propose a novel α-XNOR similarity measure tailored for spike trains. It redefines the correlation between non-spike pairs as a specific value α, effectively overcoming the limitations of dot-product similarity. Furthermore, considering the sparse nature of spike trains where spikes carry more information than non-spikes, the α-XNOR similarity correspondingly highlights the distinct importance of spikes over non-spikes. Extensive experiments demonstrate that α-XNOR similarity significantly improves performance across different spiking Transformer architectures on various static and neuromorphic datasets, further revealing the potential of spiking Transformers.
Yichen Xiao, Shuai Wang 0058, Dehao Zhang, Wenjie Wei, Yimeng Shan, Yulin Jiang, Malu Zhang
CVPR8
2025 Mixed-Precision Graph Neural Quantization for Low Bit Large Language Models
abstract
Post-Training Quantization (PTQ) is pivotal for deploying large language models (LLMs) within resource-limited settings by significantly reducing resource demands. However, existing PTQ strategies underperform at low bit levels (< 3 bits) due to the significant difference between the quantized and original weights. To enhance the quantization performance at low bit widths, we introduce a Mixed-precision Graph Neural PTQ (MG-PTQ) approach, employing a graph neural network (GNN) module to capture dependencies among weights and adaptively assign quantization bit-widths. Through the information propagation of the GNN module, our method more effectively captures dependencies among target weights, leading to a more accurate assessment of weight importance and optimized allocation of quantization strategies. Extensive experiments on the WikiText2 and C4 datasets demonstrate that our MG-PTQ method outperforms previous state-of-the-art PTQ method GPTQ, setting new benchmarks for quantization performance under low-bit (< 3 bits) conditions.
Wanlong Liu, Yichen Xiao, Dingyi Zeng, Hongyang Zhao, Wenyu Chen 0001, Malu Zhang
ICASSP6
2025 Enhancing Document-Level Relation Extraction through Entity-Pair-Level Interaction Modeling
abstract
Document-level relation extraction aims at extracting relational facts between two entities in a document. Existing approaches mainly focus on target entities, utilizing techniques such as graph neural networks to enhance their representations. However, they ignore the rich semantic correlations among entity pairs which provide wider and multifaceted information at a higher level. In this paper, we propose the Relation-based Entity-pair-level Inference (REI) model, which facilitates information interaction at the entity-pair level, enhancing logical reasoning among entities and capturing semantic correlations among entity pairs. Our REI model comprises two modules: Relation-based Information Aggregation (RIA) and Entity-pair-level Information Interaction (EII). The RIA module builds and integrates relation representations to filter out distractions from unrelated entity pairs, while the EII module models entity-pair-level information interaction through multi-head attentions. Extensive experiments on the DocRED, DWIE, CDR, and GDA datasets demonstrate the superiority of the proposed REI model, outperforming previous state-of-the-art approaches. Furthermore, we provide detailed experimental analyses based on the performance gains and illustrate the interpretability.
Wanlong Liu, Dingyi Zeng, Li Zhou 0010, Yichen Xiao, Malu Zhang, Wenyu Chen 0001
ICASSP5
2025 Decoupled Feature Matching for Few-shot Counting and Localization
abstract
Few-shot counting (FSC) aims to train a generalized visual counting model that can count any novel category given a small number of support samples. Current prevalent approaches treat FSC as a feature-matching task, leveraging attention to aggregate information from all other query patches or supports for each query patch. However, we notice that this operation blends target features with non-target features, making it difficult for the model to differentiate between targets and non-targets, thereby impacting counting accuracy. To tackle this issue, we develop a Decoupled Feature Matching Module (DFMM), which decouples target and non-target regions and conducts self-aggregation within respective regions. Furthermore, we design a Consistency Alignment Loss (CAL) to facilitate discriminative ability between targets and non-targets across multiple scales. Besides, we adopt a localization paradigm for counting and propose an Anchor-based Assignment Strategy to stabilize the optimization process and improve counting accuracy. Experiments on FSC147 and CARPK demonstrate that our method can achieve performance on par with state-of-the-art methods. Qualitative and quantitative experiments both confirm the efficacy of our proposed components.
Fan Zhang 0068, Wenyu Chen 0001, Malu Zhang, Xuanting Xie
ICASSP4
2025 Memory-Free and Parallel Computation for Quantized Spiking Neural Networks
abstract
Quantized Spiking Neural Networks (QSNNs) offer superior energy efficiency and are well-suited for deployment on resource-limited edge devices. However, limited bit-width weight and membrane potential result in a notable performance decline. In this study, we first identify a new underlying cause for this decline: the loss of historical information due to the quantized membrane potential. To tackle this issue, we introduce a memory-free quantization method that captures all historical information without directly storing membrane potentials, resulting in better performance with less memory requirements. To further improve the computational efficiency, we propose a parallel training and asynchronous inference framework that greatly increases training speed and energy efficiency. We combine the proposed memory-free quantization and parallel computation methods to develop a high-performance and efficient QSNN, named MFP-QSNN. Extensive experiments show that our MFP-QSNN achieves state-of-the-art performance on various static and neuromorphic image datasets, requiring less memory and faster training speeds. The efficiency and efficacy of the MFP-QSNN highlight its potential for energy-efficient neuromorphic computing.
Dehao Zhang, Shuai Wang 0058, Yichen Xiao, Wenjie Wei, Yimeng Shan, Malu Zhang, Yang Yang 0002
ICASSP6
2025 Quantized Spike-driven Transformer
abstract
Spiking neural networks (SNNs) are emerging as a promising energy-efficient alternative to traditional artificial neural networks (ANNs) due to their spike-driven paradigm. However, recent research in the SNN domain has mainly focused on enhancing accuracy by designing large-scale Transformer structures, which typically rely on substantial computational resources, limiting their deployment on resource-constrained devices. To overcome this challenge, we propose a quantized spike-driven Transformer baseline (QSD-Transformer), which achieves reduced resource demands by utilizing a low bit-width parameter. Regrettably, the QSD-Transformer often suffers from severe performance degradation. In this paper, we first conduct empirical analysis and find that the bimodal distribution of quantized spike-driven self-attention (Q-SDSA) leads to spike information distortion (SID) during quantization, causing significant performance degradation. To mitigate this issue, we take inspiration from mutual information entropy and propose a bi-level optimization strategy to rectify the information distribution in Q-SDSA. Specifically, at the lower level, we introduce an information-enhanced LIF to rectify the information distribution in Q-SDSA. At the upper level, we propose a fine-grained distillation scheme for the QSD-Transformer to align the distribution in Q-SDSA with that in the counterpart ANN. By integrating the bi-level optimization strategy, the QSD-Transformer can attain enhanced energy efficiency without sacrificing its high-performance advantage. We validate the QSD-Transformer on various visual tasks, and experimental results indicate that our method achieves state-of-the-art results in the SNN domain. For instance, when compared to the prior SNN benchmark on ImageNet, the QSD-Transformer achieves 80.3\% top-1 accuracy, accompanied by significant reductions of 6.0$\times$ and 8.1$\times$ in power consumption and model size, respectively. Code is available at https://github.com/bollossom/QSD-Transformer.
Xuerui Qiu, Malu Zhang, Jieyuan Zhang, Wenjie Wei, Honglin Cao, Junsheng Guo, Rui-Jie Zhu 0003, Yimeng Shan, Yang Yang 0002, Haizhou Li 0001
ICLR2
2025 Spiking Vision Transformer with Saccadic Attention
abstract
The combination of Spiking Neural Networks (SNNs) and Vision Transformers (ViTs) holds potential for achieving both energy efficiency and high performance, particularly suitable for edge vision applications. However, a significant performance gap still exists between SNN-based ViTs and their ANN counterparts. Here, we first analyze why SNN-based ViTs suffer from limited performance and identify a mismatch between the vanilla self-attention mechanism and spatio-temporal spike trains. This mismatch results in degraded spatial relevance and limited temporal interactions. To address these issues, we draw inspiration from biological saccadic attention mechanisms and introduce an innovative Saccadic Spike Self-Attention (SSSA) method. Specifically, in the spatial domain, SSSA employs a novel spike distribution-based method to effectively assess the relevance between Query and Key pairs in SNN-based ViTs. Temporally, SSSA employs a saccadic interaction module that dynamically focuses on selected visual areas at each timestep and significantly enhances whole scene understanding through temporal interactions. Building on the SSSA mechanism, we develop a SNN-based Vision Transformer (SNN-ViT). Extensive experiments across various visual tasks demonstrate that SNN-ViT achieves state-of-the-art performance with linear computational complexity. The effectiveness and efficiency of the SNN-ViT highlight its potential for power-critical edge vision applications.
Shuai Wang 0058, Malu Zhang, Dehao Zhang, Ammar Belatreche, Yichen Xiao, Yimeng Shan, Qian Sun 0014, Enqi Zhang, Yang Yang 0002
ICLR2
2025 QP-SNN: Quantized and Pruned Spiking Neural Networks
abstract
Brain-inspired Spiking Neural Networks (SNNs) leverage sparse spikes to encode information and operate in an asynchronous event-driven manner, offering a highly energy-efficient paradigm for machine intelligence. However, the current SNN community focuses primarily on performance improvement by developing large-scale models, which limits the applicability of SNNs in resource-limited edge devices. In this paper, we propose a hardware-friendly and lightweight SNN, aimed at effectively deploying high-performance SNN in resource-limited scenarios. Specifically, we first develop a baseline model that integrates uniform quantization and structured pruning, called QP-SNN baseline. While this baseline significantly reduces storage demands and computational costs, it suffers from performance decline. To address this, we conduct an in-depth analysis of the challenges in quantization and pruning that lead to performance degradation and propose solutions to enhance the baseline's performance. For weight quantization, we propose a weight rescaling strategy that utilizes bit width more effectively to enhance the model's representation capability. For structured pruning, we propose a novel pruning criterion using the singular value of spatiotemporal spike activities to enable more accurate removal of redundant kernels. Extensive experiments demonstrate that integrating two proposed methods into the baseline allows QP-SNN to achieve state-of-the-art performance and efficiency, underscoring its potential for enhancing SNN deployment in edge intelligence computing.
Wenjie Wei, Malu Zhang, Zijian Zhou 0005, Ammar Belatreche, Yimeng Shan, Honglin Cao, Jieyuan Zhang, Yang Yang 0002
ICLR2
2025 BSO: Binary Spiking Online Optimization Algorithm
abstract
Binary Spiking Neural Networks (BSNNs) offer promising efficiency advantages for resource-constrained computing. However, their training algorithms often require substantial memory overhead due to latent weights storage and temporal processing requirements. To address this issue, we propose Binary Spiking Online (BSO) optimization algorithm, a novel online training algorithm that significantly reduces training memory. BSO directly updates weights through flip signals under the online training framework. These signals are triggered when the product of gradient momentum and weights exceeds a threshold, eliminating the need for latent weights during training. To enhance performance, we propose T-BSO, a temporal-aware variant that leverages the inherent temporal dynamics of BSNNs by capturing gradient information across time steps for adaptive threshold adjustment. Theoretical analysis establishes convergence guarantees for both BSO and T-BSO, with formal regret bounds characterizing their convergence rates. Extensive experiments demonstrate that both BSO and T-BSO achieve superior optimization performance compared to existing training methods for BSNNs. The codes are available at https://github.com/hamingsi/BSO.
Wenjie Wei, Ammar Belatreche, Shuai Wang 0058, Malu Zhang, Yang Yang 0002
ICML6
2025 Binary Event-Driven Spiking Transformer
abstract
Transformer-based Spiking Neural Networks (SNNs) introduce a novel event-driven self-attention paradigm that combines the high performance of Transformers with the energy efficiency of SNNs. However, the larger model size and increased computational demands of the Transformer structure limit their practicality in resource-constrained scenarios. In this paper, we integrate binarization techniques into Transformer-based SNNs and propose the Binary Event-Driven Spiking Transformer, i.e. BESTformer. The proposed BESTformer can significantly reduce storage and computational demands by representing weights and attention maps with a mere 1-bit. However, BESTformer suffers from a severe performance drop from its full-precision counterpart due to the limited representation capability of binarization. To address this issue, we propose a Coupled Information Enhancement (CIE) method, which consists of a reversible framework and information enhancement distillation. By maximizing the mutual information between the binary model and its full-precision counterpart, the CIE method effectively mitigates the performance degradation of the BESTformer. Extensive experiments on static and neuromorphic datasets demonstrate that our method achieves superior performance to other binary SNNs, showcasing its potential as a compact yet high-performance model for resource-limited edge devices. The repository of this paper is available at https://github.com/CaoHLin/BESTFormer.
Honglin Cao, Zijian Zhou 0005, Wenjie Wei, Ammar Belatreche, Dehao Zhang, Malu Zhang, Yang Yang 0002, Haizhou Li 0001
IJCAI7
2025 Temporal-coded Spiking Transformer
abstract
Spiking Neural Networks (SNNs) have garnered significant attention due to their biological plausibility and low power consumption. While spiking transformers enhance performance by combining SNNs with transformer architecture, most rely on rate coding, limiting energy efficiency. Temporal coding methods, such as Time-To-First-Spike (TTFS) coding, offer a more efficient alternative by encoding information based on the timing of a single spike. However, integrating TTFS with transformer architecture faces challenges due to incompatibility with batch normalization (BN) and residual connections (RC), which disrupt the precise spike firing times. In this paper, we propose temporal-coded BN (tBN) and temporal-coded RC (tRC) to address these issues. Building on tBN and tRC, we develop temporal-coded spiking attention (TSA) and temporal-coded spiking transformer (T-SpikeFormer), the first to combine TTFS coding with transformer architecture. Experimental results show our model achieves state-of-the-art performance for temporal-coded SNNs and comparable results to rate-coded SNNs while significantly reducing power consumption.
Qian Sun 0014, Chengzhuo Lu, Wenyu Chen 0001, Wenjie Wei, Jieyuan Zhang, Yalan Ye, Yang Yang 0002, Malu Zhang
ACM Multimedia10
2025 Bipolar Self-attention for Spiking Transformers
abstract
Harnessing the event-driven characteristic, Spiking Neural Networks (SNNs) present a promising avenue toward energy-efficient Transformer architectures. However, existing Spiking Transformers still suffer significant performance gaps compared to their Artificial Neural Network counterparts. Through comprehensive analysis, we attribute this gap to these two factors. First, the binary nature of spike trains limits Spiking Self-attention (SSA)’s capacity to capture negative–negative and positive–negative membrane potential interactions on Querys and Keys. Second, SSA typically omits Softmax functions to avoid energy-intensive multiply-accumulate operations, thereby failing to maintain row-stochasticity constraints on attention scores. To address these issues, we propose a Bipolar Self-attention (BSA) paradigm, effectively modeling multi-polar membrane potential interactions with a fully spike-driven characteristic. Specifically, we demonstrate that ternary matrix multiplication provides a closer approximation to real-valued computation on both distribution and local correlation, enabling clear differentiation between homopolar and heteropolar interactions. Moreover, we propose a shift-based Softmax approximation named Shiftmax, which efficiently achieves low-entropy activation and partly maintains row-stochasticity without non-linear operation, enabling precise attention allocation. Extensive experiments show that BSA achieves substantial performance improvements across various tasks, including image classification, semantic segmentation, and event-based tracking. These results establish its potential as a fundamental building block for energy-efficient Spiking Transformers.
Shuai Wang 0058, Malu Zhang, Dehao Zhang, Yimeng Shan, Jieyuan Zhang, Yichen Xiao, Honglin Cao, Zeyu Ma 0002, Yang Yang 0002, Haizhou Li 0001
NeurIPS2
2025 S2NN: Sub-bit Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) offer an energy-efficient paradigm for machine intelligence, but their continued scaling poses challenges for resource-limited deployment. Despite recent advances in binary SNNs, the storage and computational demands remain substantial for large-scale networks. To further explore the compression and acceleration potential of SNNs, we propose Sub-bit Spiking Neural Networks (S$^2$NNs) that represent weights with less than one bit. Specifically, we first establish an S$^2$NN baseline by leveraging the clustering patterns of kernels in well-trained binary SNNs. This baseline is highly efficient but suffers from \textit{outlier-induced codeword selection bias} during training. To mitigate this issue, we propose an \textit{outlier-aware sub-bit weight quantization} (OS-Quant) method, which optimizes codeword selection by identifying and adaptively scaling outliers. Furthermore, we propose a \textit{membrane potential-based feature distillation} (MPFD) method, improving the performance of highly compressed S$^2$NN via more precise guidance from a teacher model. Extensive results on vision reveal that S$^2$NN outperforms existing quantized SNNs in both performance and efficiency, making it promising for edge computing applications.
Wenjie Wei, Malu Zhang, Jieyuan Zhang, Ammar Belatreche, Shuai Wang 0058, Yimeng Shan, Honglin Cao, Guoqing Wang 0001, Yang Yang 0002, Haizhou Li 0001
NeurIPS2
2025 Unveiling the Spatial-temporal Effective Receptive Fields of Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) demonstrate significant potential for energy-efficient neuromorphic computing through an event-driven paradigm. While training methods and computational models have greatly advanced, SNNs struggle to achieve competitive performance in visual long-sequence modeling tasks. In artificial neural networks, the effective receptive field (ERF) serves as a valuable tool for analyzing feature extraction capabilities in visual long-sequence modeling. Inspired by this, we introduce the Spatio-Temporal Effective Receptive Field (ST-ERF) to analyze the ERF distributions across various Transformer-based SNNs. Based on the proposed ST-ERF, we reveal that these models suffer from establishing a robust global ST-ERF, thereby limiting their visual feature modeling capabilities. To overcome this issue, we propose two novel channel-mixer architectures: \underline{m}ulti-\underline{l}ayer-\underline{p}erceptron-based m\underline{ixer} (MLPixer) and \underline{s}plash-and-\underline{r}econstruct \underline{b}lock (SRB). These architectures enhance global spatial ERF through all timesteps in early network stages of Transformer-based SNNs, improving performance on challenging visual long-sequence modeling tasks. Extensive experiments conducted on the Meta-SDT variants and across object detection and semantic segmentation tasks further validate the effectiveness of our proposed method. Beyond these specific applications, we believe the proposed ST-ERF framework can provide valuable insights for designing and optimizing SNN architectures across a broader range of tasks. The code is available at \href{https://github.com/EricZhang1412/Spatial-temporal-ERF}{\faGithub~EricZhang1412/Spatial-temporal-ERF}.
Jieyuan Zhang, Shuai Wang 0058, Wenjie Wei, Qian Sun 0014, Malu Zhang, Yang Yang 0002, Haizhou Li 0001
NeurIPS7
2025 Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence Modeling
abstract
The explosive growth in sequence length has intensified the demand for effective and efficient long sequence modeling. Benefiting from intrinsic oscillatory membrane dynamics, Resonate-and-Fire (RF) neurons can efficiently extract frequency components from input signals and encode them into spatiotemporal spike trains, making them well-suited for long sequence modeling. However, RF neurons exhibit limited effective memory capacity and a trade-off between energy efficiency and training speed on complex temporal tasks. Inspired by the dendritic structure of biological neurons, we propose a Dendritic Resonate-and-Fire (D-RF) model, which explicitly incorporates a multi-dendritic and soma architecture. Each dendritic branch encodes specific frequency bands by utilizing the intrinsic oscillatory dynamics of RF neurons, thereby collectively achieving comprehensive frequency representation. Furthermore, we introduce an adaptive threshold mechanism into the soma structure. This mechanism adjusts the firing threshold according to historical spiking activity, thereby reducing redundant spikes while maintaining training efficiency in long-sequence tasks. Extensive experiments demonstrate that our method maintains competitive accuracy while substantially ensuring sparse spikes without compromising computational efficiency during training. These results underscore its potential as an effective and efficient solution for long sequence modeling on edge platforms.
Dehao Zhang, Malu Zhang, Shuai Wang 0058, Wenjie Wei, Zeyu Ma 0002, Guoqing Wang 0001, Yang Yang 0002, Haizhou Li 0001
NeurIPS2
2025 Document-level relation extraction with structural encoding and entity-pair-level information interaction
Wanlong Liu, Yichen Xiao, Shaohuan Cheng, Dingyi Zeng, Li Zhou 0010, Weishan Kong, Malu Zhang, Wenyu Chen 0001
Expert Syst. Appl.7
2025 Electroencephalography Decoding with Conditional Identification Generator
abstract
Decoding Electroencephalography (EEG) signals are extremely useful for advancing and understanding human-artificial intelligence (AI) interaction systems. Recent advancements in deep neural networks (DNNs) have demonstrated significant promise in this respect due to their ability to model complex nonlinear relationships. However, DNNs face persistent challenges in addressing the inter-person variability inherent in EEG signals, which limits their generalizability. To tackle this limitation, we propose a novel framework that integrates conditional identification information, leveraging the interaction between EEG signals and individual traits to enhance the model's internal representation and improve decoding accuracy. Building on this foundation, we further introduce a privacy-preserving conditional information generator - a generative model that derives embedding knowledge directly from raw EEG signals. This approach eliminates the need for personal identification via individual tests, ensuring both efficiency and privacy. Experimental evaluations conducted on WithMe dataset confirm that this framework outperforms baseline network architectures. Notably, our approach achieves substantial improvements in decoding accuracy for both familiar and unseen subjects, paving the way for efficient, robust, and privacy-conscious human-computer interface systems.
Pengfei Sun 0003, Jorg De Winne, Malu Zhang, Paul Devos, Dick Botteldooren
Int. J. Neural Syst.3
2025 Efficient Automatic Modulation Classification in Nonterrestrial Networks With SNN-Based Transformer
abstract
With the development of informatization of IoT devices, nonterrestrial networks (NTNs) are becoming more and more important. NTN, including air and space networks, face challenges, such as high-computational complexity, bandwidth requirements, and memory constraints. An intelligent automatic modulation classification (AMC) mechanism based on neural networks plays a pivotal role in enhancing spectrum efficiency, throughput, and link reliability. Past work in AMC has evolved from likelihood-based and feature-based methods to traditional machine learning techniques and, more recently, to deep neural networks (DNNs). However, existing DNN architectures pose challenges for NTN due to high-computational complexity, bandwidth requirements, and memory consumption. Addressing this problems, we proposes a spiking transformer-based model for AMC, exploiting temporal dynamics for enhanced performance. Biologically inspired spiking neural networks enable us to exploit the sparse and binarized activation properties of spiking neurons, allowing us to build AMC models with high-energy efficiency and high availability that can be used in NTN systems. Furthermore, we introduce a weight binarization method to reduce the model size, which also further reduces the bandwidth and memory requirements of AMC in NTN edge deployment. Experimental results demonstrate the superiority of our approach over state-of-the-art methods, with the binarized model achieving comparable accuracy at a fraction of the size.
Dingyi Zeng, Yichen Xiao, Wanlong Liu, Huilin Du, Enqi Zhang, Dehao Zhang, Malu Zhang, Wenyu Chen 0001
IEEE Internet Things J.8
2025 ESTSformer: Efficient spatio-temporal spiking transformer
Chengzhuo Lu, Huilin Du, Wenjie Wei, Qian Sun 0014, Dingyi Zeng, Wenyu Chen 0001, Malu Zhang, Yang Yang 0002
Neural Networks8
2025 Delayed knowledge transfer: Cross-modal knowledge transfer from delayed stimulus to EEG for continuous attention detection based on spike-represented EEG signals
Pengfei Sun 0003, Jorg De Winne, Malu Zhang, Paul Devos, Dick Botteldooren
Neural Networks3
2025 Ternary spike-based neuromorphic signal processing system
Shuai Wang 0058, Dehao Zhang, Ammar Belatreche, Yichen Xiao, Hongyu Qing, Wenjie Wei, Malu Zhang, Yang Yang 0002
Neural Networks7
2025 Spiking Neural Networks With Adaptive Membrane Time Constant for Event-Based Tracking
abstract
The brain-inspired Spiking Neural Networks (SNNs) work in an event-driven manner and have an implicit recurrence in neuronal membrane potential to memorize information over time, which are inherently suitable to handle temporal event-based streams. Despite their temporal nature and recent approaches advancements, these methods have predominantly been assessed on event-based classification tasks. In this paper, we explore the utility of SNNs for event-based tracking tasks. Specifically, we propose a brain-inspired adaptive Leaky Integrate-and-Fire neuron (BA-LIF) that can adaptively adjust the membrane time constant according to the inputs, thereby accelerating the leakage of meaningless noise features and reducing the decay of valuable information. SNNs composed of our proposed BA-LIF neurons can achieve high performance without a careful and time-consuming trial-by-error initialization on the membrane time constant. The adaptive capability of our network is further improved by introducing an extra temporal feature aggregator (TFA) that assigns attention weights over the temporal dimension. Extensive experiments on various event-based tracking datasets validate the effectiveness of our proposed method. We further validate the generalization capability of our method by applying it to other event-classification tasks.
Jiqing Zhang, Malu Zhang, Yuanchen Wang, Qianhui Liu, Haizhou Li 0001, Xin Yang 0011
IEEE Trans. Image Process.2
2025 Delayed Memory Unit: Modeling Temporal Dependency Through Delay Gate
abstract
Recurrent neural networks (RNNs) are widely recognized for their proficiency in modeling temporal dependencies, making them highly prevalent in sequential data processing applications. Nevertheless, vanilla RNNs are confronted with the well-known issue of gradient vanishing and exploding, posing a significant challenge for learning and establishing long-range dependencies. Additionally, gated RNNs tend to be over-parameterized, resulting in poor computational efficiency and network generalization. To address these challenges, this article proposes a novel delayed memory unit (DMU). The DMU incorporates a delay line structure along with delay gates into vanilla RNN, thereby enhancing temporal interaction and facilitating temporal credit assignment. Specifically, the DMU is designed to directly distribute the input information to the optimal time instant in the future, rather than aggregating and redistributing it over time through intricate network dynamics. Our proposed DMU demonstrates superior temporal modeling capabilities across a broad range of sequential modeling tasks, utilizing considerably fewer parameters than other state-of-the-art gated RNN models in applications such as speech recognition, radar gesture recognition, ECG waveform segmentation, and permuted sequential (PS) image classification.
Pengfei Sun 0003, Jibin Wu, Malu Zhang, Paul Devos, Dick Botteldooren
IEEE Trans. Neural Networks Learn. Syst.3
2025 EMWQ: An Efficient Mixed Precision Weight Quantization Method for Large Language Models
abstract
Large language models (LLMs) have gained a lot of attention and achievements recently because of their significant comprehension and generative abilities. However, the large-scale parameters of LLMs require considerable computational resources in the training and inference process, which restricts their wide application. To overcome this challenge, we propose an efficient mixed precision weight quantization (EMWQ) method for LLMs in this article. Specifically, we introduce a new outlier detection method by analyzing the weight distribution instead of the conventional weight magnitude. Then, we propose a dual-quantization strategy that quantizes both the outlier critical columns and the residual matrices with different precision. Besides, we introduce two effective EMWQ-based application frameworks, the EMWQ-R and EMWQ-O in our study. Comprehensive experiments are conducted on the Penn Treebank (PTB), C4, ARC-Easy datasets, and MMLU benchmark across various tasks. The comparison results demonstrate that the proposed EMWQ achieves state-of-the-art performance in mixed precision quantization and further reduces computational memory cost. Besides, it has higher generalizability compared with conventional methods.
Xiurui Xie, Guowei Peng, Malu Zhang, Guangchun Luo, Yang Yang 0002, Guisong Liu
IEEE Trans. Neural Networks Learn. Syst.4
2025 Toward Building Human-Like Sequential Memory Using Brain-Inspired Spiking Neural Models
abstract
The brain is able to acquire and store memories of everyday experiences in real-time. It can also selectively forget information to facilitate memory updating. However, our understanding of the underlying mechanisms and coordination of these processes within the brain remains limited. However, no existing artificial intelligence models have yet matched human-level capabilities in terms of memory storage and retrieval. This study introduces a brain-inspired spiking neural model that integrates the learning and forgetting processes of sequential memory. The proposed model closely mimics the distributed and sparse temporal coding observed in the biological neural system. It employs one-shot online learning for memory formation and uses biologically plausible mechanisms of neural oscillation and phase precession to retrieve memorized sequences reliably. In addition, an active forgetting mechanism is integrated into the spiking neural model, enabling memory removal, flexibility, and updating. The proposed memory model not only enhances our understanding of human memory processes but also provides a robust framework for addressing temporal modeling tasks.
Malu Zhang, Xiaoling Luo 0001, Jibin Wu, Ammar Belatreche, Siqi Cai 0002, Yang Yang 0002, Haizhou Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2025 TCJA-SNN: Temporal-Channel Joint Attention for Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are attracting widespread interest due to their biological plausibility, energy efficiency, and powerful spatiotemporal information representation ability. Given the critical role of attention mechanisms in enhancing neural network performance, the integration of SNNs and attention mechanisms exhibits tremendous potential to deliver energy-efficient and high-performance computing paradigms. In this article, we present a novel temporal-channel joint attention mechanism for SNNs, referred to as TCJA-SNN. The proposed TCJA-SNN framework can effectively assess the significance of spike sequence from both spatial and temporal dimensions. More specifically, our essential technical contribution lies on: 1) we employ the squeeze operation to compress the spike stream into an average matrix. Then, we leverage two local attention mechanisms based on efficient 1-D convolutions to facilitate comprehensive feature extraction at the temporal and channel levels independently and 2) we introduce the cross-convolutional fusion (CCF) layer as a novel approach to model the interdependencies between the temporal and channel scopes. This layer effectively breaks the independence of these two dimensions and enables the interaction between features. Experimental results demonstrate that the proposed TCJA-SNN outperforms the state-of-the-art (SOTA) on all standard static and neuromorphic datasets, including Fashion-MNIST, CIFAR10, CIFAR100, CIFAR10-DVS, N-Caltech 101, and DVS128 Gesture. Furthermore, we effectively apply the TCJA-SNN framework to image generation tasks by leveraging a variation autoencoder. To the best of our knowledge, this study is the first instance where the SNN-attention mechanism has been employed for high-level classification and low-level generation tasks. Our implementation codes are available at https://github.com/ridgerchu/TCJA.
Rui-Jie Zhu 0003, Malu Zhang, Qihang Zhao, Yule Duan 0001, Liang-Jian Deng
IEEE Trans. Neural Networks Learn. Syst.2
2024 Restoring Speaking Lips from Occlusion for Audio-Visual Speech Recognition
abstract
Prior studies on audio-visual speech recognition typically assume the visibility of speaking lips, ignoring the fact that visual occlusion occurs in real-world videos, thus adversely affecting recognition performance. To address this issue, we propose a framework that restores occluded lips in a video by utilizing both the video itself and the corresponding noisy audio. Specifically, the framework aims to achieve these three tasks: detecting occluded frames, masking occluded areas, and reconstruction of masked regions. We tackle the first two issues by utilizing the Class Activation Map (CAM) obtained from occluded frame detection to facilitate the masking of occluded areas. Additionally, we introduce a novel synthesis-matching strategy for the reconstruction to ensure the compatibility of audio features with different levels of occlusion. Our framework is evaluated in terms of Word Error Rate (WER) on the original videos, the videos corrupted by concealed lips, and the videos restored using the framework with several existing state-of-the-art audio-visual speech recognition methods. Experimental results substantiate that our framework significantly mitigates performance degradation resulting from lip occlusion. Under -5dB noise conditions, AV-Hubert's WER increases from 10.62% to 13.87% due to lip occlusion, but rebounds to 11.87% in conjunction with the proposed framework. Furthermore, the framework also demonstrates its capacity to produce natural synthesized images in qualitative assessments.
Zexu Pan, Malu Zhang, Robby T. Tan, Haizhou Li 0001
AAAI3
2024 A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue Evaluators
abstract
Automatic evaluation is an integral aspect of dialogue system research. The traditional reference-based NLG metrics are generally found to be unsuitable for dialogue assessment. Consequently, recent studies have suggested various unique, reference-free neural metrics that better align with human evaluations. Notably among them, large language models (LLMs), particularly the instruction-tuned variants like ChatGPT, are shown to be promising substitutes for human judges. Yet, existing works on utilizing LLMs for automatic dialogue evaluation are limited in their scope in terms of the number of meta-evaluation datasets, mode of evaluation, coverage of LLMs, etc. Hence, it remains inconclusive how effective these LLMs are. To this end, we conduct a comprehensive study on the application of LLMs for automatic dialogue evaluation. Specifically, we analyze the multi-dimensional evaluation capability of 30 recently emerged LLMs at both turn and dialogue levels, using a comprehensive set of 12 meta-evaluation datasets. Additionally, we probe the robustness of the LLMs in handling various adversarial perturbations at both turn and dialogue levels. Finally, we explore how model-level and dimension-level ensembles impact the evaluation performance. All resources are available at https://github.com/e0397123/comp-analysis.
Chen Zhang 0020, Luis Fernando D'Haro, Yiming Chen 0010, Malu Zhang, Haizhou Li 0001
AAAI4
2024 Raindrop Clarity: A Dual-Focused Dataset for Day and Night Raindrop Removal
Yeying Jin, Xin Li 0082, Yan Zhang 0004, Malu Zhang
ECCV (6)5
2024 Spiking-Leaf: A Learnable Auditory Front-End for Spiking Neural Networks
abstract
Brain-inspired spiking neural networks (SNNs) have demonstrated great potential for temporal signal processing. However, their performance in speech processing remains limited due to the lack of an effective auditory front-end. To address this limitation, we introduce Spiking-LEAF, a learnable auditory front-end meticulously designed for SNN-based speech processing. Spiking-LEAF combines a learnable filter bank with a novel two-compartment spiking neuron model called IHC-LIF. The IHC-LIF neurons draw inspiration from the structure of inner hair cells (IHC) and they leverage segregated dendritic and somatic compartments to effectively capture multi-scale temporal dynamics of speech signals. Additionally, the IHC-LIF neurons incorporate the lateral feedback mechanism along with spike regularization loss to enhance spike encoding efficiency. On keyword spotting and speaker identification tasks, the proposed Spiking-LEAF outperforms both SOTA spiking auditory front-ends and conventional real-valued acoustic features in terms of classification accuracy, noise robustness, and encoding efficiency.
Zeyang Song, Jibin Wu, Malu Zhang, Zheng Shou 0001, Haizhou Li 0001
ICASSP3
2024 LitE-SNN: Designing Lightweight and Efficient Spiking Neural Network through Spatial-Temporal Compressive Network Search and Joint Optimization
Qianhui Liu, Malu Zhang, Gang Pan 0001, Haizhou Li 0001
IJCAI3
2024 Global-Local Convolution with Spiking Neural Networks for Energy-efficient Keyword Spotting
abstract
25th Annual Conference of the International Speech Communication Association, Interspeech 2024, Kos, Greece, September 1-5, 2024. ISCA 2024
Shuai Wang 0058, Dehao Zhang, Wenjie Wei, Jibin Wu, Malu Zhang
INTERSPEECH7
2024 Q-SNNs: Quantized Spiking Neural Networks
abstract
Brain-inspired Spiking Neural Networks (SNNs) leverage sparse spikes to represent information and process them in an asynchronous event-driven manner, offering an energy-efficient paradigm for the next generation of machine intelligence. However, the current focus within the SNN community prioritizes accuracy optimization through the development of large-scale models, limiting their viability in resource-constrained and low-power edge devices. To address this challenge, we introduce a lightweight and hardware-friendly Quantized SNN (Q-SNN) that applies quantization to both synaptic weights and membrane potentials. By significantly compressing these two key elements, the proposed Q-SNNs substantially reduce both memory usage and computational complexity. Moreover, to prevent the performance degradation caused by this compression, we present a new Weight-Spike Dual Regulation (WS-DR) method inspired by information entropy theory. Experimental evaluations on various datasets, including static and neuromorphic, demonstrate that our Q-SNNs outperform existing methods in terms of both model size and accuracy. These state-of-the-art results in efficiency and efficacy suggest that the proposed method can significantly improve edge intelligent computing.
Wenjie Wei, Ammar Belatreche, Yichen Xiao, Honglin Cao, Zhenbang Ren, Guoqing Wang 0003, Malu Zhang, Yang Yang 0002
ACM Multimedia8
2024 JoReS-Diff: Joint Retinex and Semantic Priors in Diffusion Model for Low-light Image Enhancement
abstract
Low-light image enhancement (LLIE) has achieved promising performance by employing conditional diffusion models. Despite the success of some conditional methods, previous methods may neglect the importance of a sufficient formulation of task-specific condition strategy, resulting in suboptimal visual outcomes. In this study, we propose JoReS-Diff, a novel approach that incorporates Retinex- and semantic-based priors as the additional pre-processing condition to regulate the generating capabilities of the diffusion model. We first leverage pre-trained decomposition network to generate the Retinex prior, which is updated with better quality by an adjustment network and integrated into a refinement network to implement Retinex-based conditional generation at both feature- and image-levels. Moreover, the semantic prior is extracted from the input image with an off-the-shelf semantic segmentation model and incorporated through semantic attention layers. By treating Retinex- and semantic-based priors as the condition, JoReS-Diff presents a unique perspective for establishing an diffusion model for LLIE and similar image enhancement tasks. Extensive experiments validate the rationality and superiority of our approach.
Yuhui Wu 0001, Guoqing Wang 0001, Zhiwen Wang 0004, Yang Yang 0002, Tianyu Li 0003, Malu Zhang, Chongyi Li, Heng Tao Shen
ACM Multimedia6
2024 Spike-based Neuromorphic Model for Sound Source Localization
abstract
Biological systems possess remarkable sound source localization (SSL) capabilities that are critical for survival in complex environments. This ability arises from the collaboration between the auditory periphery, which encodes sound as precisely timed spikes, and the auditory cortex, which performs spike-based computations. Inspired by these biological mechanisms, we propose a novel neuromorphic SSL framework that integrates spike-based neural encoding and computation. The framework employs Resonate-and-Fire (RF) neurons with a phase-locking coding (RF-PLC) method to achieve energy-efficient audio processing. The RF-PLC method leverages the resonance properties of RF neurons to efficiently convert audio signals to time-frequency representation and encode interaural time difference (ITD) cues into discriminative spike patterns. In addition, biological adaptations like frequency band selectivity and short-term memory effectively filter out many environmental noises, enhancing SSL capabilities in real-world settings. Inspired by these adaptations, we propose a spike-driven multi-auditory attention (MAA) module that significantly improves both the accuracy and robustness of the proposed SSL framework. Extensive experimentation demonstrates that our SSL framework achieves state-of-the-art accuracy in SSL tasks. Furthermore, it shows exceptional noise robustness and maintains high accuracy even at very low signal-to-noise ratios. By mimicking biological hearing, this neuromorphic approach contributes to the development of high-performance and explainable artificial intelligence systems capable of superior performance in real-world environments.
Dehao Zhang, Shuai Wang 0058, Ammar Belatreche, Wenjie Wei, Yichen Xiao, Haorui Zheng, Zijian Zhou 0005, Malu Zhang, Yang Yang 0002
NeurIPS8
2024 Tensor decomposition based attention module for spiking neural networks
Rui-Jie Zhu 0003, Xuerui Qiu, Yule Duan 0001, Malu Zhang, Liang-Jian Deng
Knowl. Based Syst.5
2024 A two-stage spiking meta-learning method for few-shot classification
Qiugang Zhan, Bingchao Wang, Anning Jiang, Xiurui Xie, Malu Zhang, Guisong Liu
Knowl. Based Syst.5
2024 QDAP: Downsizing adaptive policy for cooperative multi-agent reinforcement learning
Zhitong Zhao, Siying Wang 0002, Fan Zhang 0068, Malu Zhang, Wenyu Chen 0001
Knowl. Based Syst.5
2024 Delay learning based on temporal coding in Spiking Neural Networks
Pengfei Sun 0003, Jibin Wu, Malu Zhang, Paul Devos, Dick Botteldooren
Neural Networks3
2024 A universal ANN-to-SNN framework for achieving high accuracy and low latency deep Spiking Neural Networks
Malu Zhang, Xiaoling Luo 0001, Hong Qu 0002
Neural Networks3
2024 Efficient spiking neural network design via neural architecture search
Qianhui Liu, Malu Zhang, Lang Feng 0002, De Ma, Haizhou Li 0001, Gang Pan 0001
Neural Networks3
2024 Spiking Transfer Learning From RGB Image to Neuromorphic Event Stream
abstract
Recent advances in bio-inspired vision with event cameras and associated spiking neural networks (SNNs) have provided promising solutions for low-power consumption neuromorphic tasks. However, as the research of event cameras is still in its infancy, the amount of labeled event stream data is much less than that of the RGB database. The traditional method of converting static images into event streams by simulation to increase the sample size cannot simulate the characteristics of event cameras such as high temporal resolution. To take advantage of both the rich knowledge in labeled RGB images and the features of the event camera, we propose a transfer learning method from the RGB to the event domain in this paper. Specifically, we first introduce a transfer learning framework named R2ETL (RGB to Event Transfer Learning), including a novel encoding alignment module and a feature alignment module. Then, we introduce the temporal centered kernel alignment (TCKA) loss function to improve the efficiency of transfer learning. It aligns the distribution of temporal neuron states by adding a temporal learning constraint. Finally, we theoretically analyze the amount of data required by the deep neuromorphic model to prove the necessity of our method. Numerous experiments demonstrate that our proposed framework outperforms the state-of-the-art SNN and artificial neural network (ANN) models trained on event streams, including N-MNIST, CIFAR10-DVS and N-Caltech101. This indicates that the R2ETL framework is able to leverage the knowledge of labeled RGB images to help the training of SNN on event streams.
Qiugang Zhan, Guisong Liu, Xiurui Xie, Malu Zhang, Huajin Tang
IEEE Trans. Image Process.5
2024 Biologically Plausible Sparse Temporal Word Representations
abstract
Word representations, usually derived from a large corpus and endowed with rich semantic information, have been widely applied to natural language tasks. Traditional deep language models, on the basis of dense word representations, requires large memory space and computing resource. The brain-inspired neuromorphic computing systems, with the advantages of better biological interpretability and less energy consumption, still have major difficulties in the representation of words in terms of neuronal activities, which has restricted their further application in more complicated downstream language tasks. Comprehensively exploring the diverse neuronal dynamics of both integration and resonance, we probe into three spiking neuron models to post-process the original dense word embeddings, and test the generated sparse temporal codes on several tasks concerning both word-level and sentence-level semantics. The experimental results show that our sparse binary word representations could perform on par with or even better than original word embeddings in capturing semantic information, while requiring less storage. Our methods provide a robust representation foundation of language in terms of neuronal activities, which could potentially be applied to future downstream natural language tasks under neuromorphic computing systems.
Yuguo Liu, Wenyu Chen 0001, Malu Zhang, Hong Qu 0002
IEEE Trans. Neural Networks Learn. Syst.5
2024 Event-Driven Spiking Learning Algorithm Using Aggregated Labels
abstract
Traditional spiking learning algorithm aims to train neurons to spike at a specific time or on a particular frequency, which requires precise time and frequency labels in the training process. While in reality, usually only aggregated labels of sequential patterns are provided. The aggregate-label (AL) learning is proposed to discover these predictive features in distracting background streams only by aggregated spikes. It has achieved much success recently, but it is still computationally intensive and has limited use in deep networks. To address these issues, we propose an event-driven spiking aggregate learning algorithm (SALA) in this article. Specifically, to reduce the computational complexity, we improve the conventional spike-threshold-surface (STS) calculation in AL learning by analytical calculating voltage peak values in spiking neurons. Then we derive the algorithm to multilayers by event-driven strategy using aggregated spikes. We conduct comprehensive experiments on various tasks including temporal clue recognition, segmented and continuous speech recognition, and neuromorphic image classification. The experimental results demonstrate that the new STS method improves the efficiency of AL learning significantly, and the proposed algorithm outperforms the conventional spiking algorithm in various temporal clue recognition tasks.
Xiurui Xie, Yansong Chua, Guisong Liu, Malu Zhang, Guangchun Luo, Huajin Tang
IEEE Trans. Neural Networks Learn. Syst.4
2024 Minicolumn-Based Episodic Memory Model With Spiking Neurons, Dendrites and Delays
abstract
Episodic memory is fundamental to the brain's cognitive function, but how neuronal activity is temporally organized during its encoding and retrieval is still unknown. In this article, combining hippocampus structure with a spiking neural network (SNN), a new bionic spiking temporal memory (BSTM) model is proposed to explore the encoding, formation, and retrieval of episodic memory. For encoding episodic memory, the spike-timing-dependent-plasticity (STDP) learning algorithm and a proposed minicolumn selection algorithm are used to encode each input item into several active minicolumns. For the formation of episodic memory, a sequential memory algorithm is proposed to store the contexts between items. For retrieval of episodic memory, the local retrieval algorithm and the global retrieval algorithm are proposed to retrieve sequence information, achieving multisentence prediction and multitime step prediction. All functions of BSTM are based on bionic spiking neurons, which have biological characteristics including columnar and dendritic structures, firing and receiving spikes, and delaying transmission. To test the performance of the BSTM model, the Children's Book Test (CBT) data set was used to conduct a series of experiments under different settings, including changing the number of minicolumns, neurons and sequences, modifying sequence items, etc. Compared to other sequence memory algorithms, the experimental results show that the proposed BSTM achieves higher accuracy and better robustness.
Yi Chen 0034, Jilun Zhang, Xiaoling Luo 0001, Malu Zhang, Hong Qu 0002, Zhang Yi 0001
IEEE Trans. Neural Networks Learn. Syst.5
2023 Substructure Aware Graph Neural Networks
abstract
Despite the great achievements of Graph Neural Networks (GNNs) in graph learning, conventional GNNs struggle to break through the upper limit of the expressiveness of first-order Weisfeiler-Leman graph isomorphism test algorithm (1-WL) due to the consistency of the propagation paradigm of GNNs with the 1-WL.Based on the fact that it is easier to distinguish the original graph through subgraphs, we propose a novel framework neural network framework called Substructure Aware Graph Neural Networks (SAGNN) to address these issues. We first propose a Cut subgraph which can be obtained from the original graph by continuously and selectively removing edges. Then we extend the random walk encoding paradigm to the return probability of the rooted node on the subgraph to capture the structural information and use it as a node feature to improve the expressiveness of GNNs. We theoretically prove that our framework is more powerful than 1-WL, and is superior in structure perception. Our extensive experiments demonstrate the effectiveness of our framework, achieving state-of-the-art performance on a variety of well-proven graph tasks, and GNNs equipped with our framework perform flawlessly even in 3-WL failed graphs. Specifically, our framework achieves a maximum performance improvement of 83% compared to the base models and 32% compared to the previous state-of-the-art methods.
Dingyi Zeng, Wanlong Liu, Wenyu Chen 0001, Li Zhou 0010, Malu Zhang, Hong Qu 0002
AAAI5
2023 Seeing What You Said: Talking Face Generation Guided by a Lip Reading Expert
abstract
Talking face generation, also known as speech-to-lip generation, reconstructs facial motions concerning lips given coherent speech input. The previous studies revealed the importance of lip-speech synchronization and visual quality. Despite much progress, they hardly focus on the content of lip movements i.e., the visual intelligibility of the spoken words, which is an important aspect of generation quality. To address the problem, we propose using a lipreading expert to improve the intelligibility of the generated lip regions by penalizing the incorrect generation results. Moreover, to compensate for data scarcity, we train the lip-reading expert in an audio-visual self-supervised manner. With a lip-reading expert, we propose a novel contrastive learning to enhance lip-speech synchronization, and a transformer to encode audio synchronically with video, while considering global temporal dependency of audio. For evaluation, we propose a new strategy with two different lip-reading experts to measure intelligibility of the generated videos. Rigorous experiments show that our proposal is superior to other State-of-the-art (SOTA) methods, such as Wav2Lip, in reading intelligibility i.e., over 38% Word Error Rate (WER) on LRS2 dataset and 27.8% accuracy on LRW dataset. We also achieve the SOTA performance in lip-speech synchronization and comparable performances in visual quality.
Xinyuan Qian 0001, Malu Zhang, Robby T. Tan, Haizhou Li 0001
CVPR3
2023 Temporal-Coded Spiking Neural Networks with Dynamic Firing Threshold: Learning with Event-Driven Backpropagation
abstract
Spiking Neural Networks (SNNs) offer a highly promising computing paradigm due to their biological plausibility, exceptional spatiotemporal information processing capability and low power consumption. As a temporal encoding scheme for SNNs, Time-To-First-Spike (TTFS) encodes information using the timing of a single spike, which allows spiking neurons to transmit information through sparse spike trains and results in lower power consumption and higher computational efficiency compared to traditional rate-based encoding counterparts. However, despite the advantages of the TTFS encoding scheme, the effective and efficient training of TTFS-based deep SNNs remains a significant and open research problem. In this work, we first examine the factors underlying the limitations of applying existing TTFS-based learning algorithms to deep SNNs. Specifically, we investigate issues related to over-sparsity of spikes and the complexity of finding the ‘causal set'. We then propose a simple yet efficient dynamic firing threshold (DFT) mechanism for spiking neurons to address these issues. Building upon the proposed DFT mechanism, we further introduce a novel direct training algorithm for TTFS-based deep SNNs, called DTA-TTFS. This method utilizes event-driven processing and spike timing to enable efficient learning of deep SNNs. The proposed training method was validated on the image classification task and experimental results clearly demonstrate that our proposed method achieves state-of-the-art accuracy in comparison to existing TTFS-based learning algorithms, while maintaining high levels of sparsity and energy efficiency on neuromorphic inference accelerator.
Wenjie Wei, Malu Zhang, Hong Qu 0002, Ammar Belatreche, Jian Zhang 0020, Hong Chen 0002
ICCV2
2023 Spatial-Temporal Self-Attention for Asynchronous Spiking Neural Networks
abstract
The brain-inspired spiking neural networks (SNNs) are receiving increasing attention due to their asynchronous event-driven characteristics and low power consumption. As attention mechanisms recently become an indispensable part of sequence dependence modeling, the combination of SNNs and attention mechanisms holds great potential for energy-efficient and high-performance computing paradigms. However, the existing works cannot benefit from both temporal-wise attention and the asynchronous characteristic of SNNs. To fully leverage the advantages of both SNNs and attention mechanisms, we propose an SNNs-based spatial-temporal self-attention (STSA) mechanism, which calculates the feature dependence across the time and space domains without destroying the asynchronous transmission properties of SNNs. To further improve the performance, we also propose a spatial-temporal relative position bias (STRPB) for STSA to consider the spatiotemporal position of spikes. Based on the STSA and STRPB, we construct a spatial-temporal spiking Transformer framework, named STS-Transformer, which is powerful and enables SNNs to work in an asynchronous event-driven manner. Extensive experiments are conducted on popular neuromorphic datasets and speech datasets, including DVS128 Gesture, CIFAR10-DVS, and Google Speech Commands, and our experimental results can outperform other state-of-the-art models.
Chengzhuo Lu, Yuguo Liu, Malu Zhang, Hong Qu 0002
IJCAI5
2023 A Low Power and Low Latency FPGA-Based Spiking Neural Network Accelerator
abstract
Spiking Neural Networks (SNNs), known as the third generation of the neural network, are famous for their biological plausibility and brain-like characteristics. Recent efforts further demonstrate the potential of SNNs in high-speed inference by designing accelerators with the parallelism of temporal or spatial dimensions. However, with the limitation of hardware resources, the accelerator designs must utilize off-chip memory to store many intermediate data, which leads to both high power consumption and long latency. In this paper, we focus on the data flow between layers to improve arithmetic efficiency. Based on the spike discrete property, we design a convolution-pooling(CONVP) unit that fuses the processing of the convolutional layer and pooling layer to reduce latency and resource utilization. Furthermore, for the fully-connected layer, we apply intra-output parallelism and inter-output parallelism to accelerate network inference. We demonstrate the effectiveness of our proposed hardware architecture by implementing different SNN models with the different datasets on a Zynq XA7Z020 FPGA. The experiments show that our accelerator can achieve about x28 inference speed up with a competitive power compared with FPGA implementation on MNIST dataset and a x15 inference speed up with low power compared with ASIC design on DVSGesture dataset.
Yi Chen 0034, Zihang Zeng, Malu Zhang, Hong Qu 0002
IJCNN4
2023 Bio-inspired Active Learning method in spiking neural network
Qiugang Zhan, Guisong Liu, Xiurui Xie, Malu Zhang, Guolin Sun
Knowl. Based Syst.4
2023 DPGNN: Dual-perception graph neural network for representation learning
Li Zhou 0010, Wenyu Chen 0001, Dingyi Zeng, Shaohuan Cheng, Wanlong Liu, Malu Zhang, Hong Qu 0002
Knowl. Based Syst.6
2023 Supervised Learning in Multilayer Spiking Neural Networks With Spike Temporal Error Backpropagation
abstract
The brain-inspired spiking neural networks (SNNs) hold the advantages of lower power consumption and powerful computing capability. However, the lack of effective learning algorithms has obstructed the theoretical advance and applications of SNNs. The majority of the existing learning algorithms for SNNs are based on the synaptic weight adjustment. However, neuroscience findings confirm that synaptic delays can also be modulated to play an important role in the learning process. Here, we propose a gradient descent-based learning algorithm for synaptic delays to enhance the sequential learning performance of single spiking neuron. Moreover, we extend the proposed method to multilayer SNNs with spike temporal-based error backpropagation. In the proposed multilayer learning algorithm, information is encoded in the relative timing of individual neuronal spikes, and learning is performed based on the exact derivatives of the postsynaptic spike times with respect to presynaptic spike times. Experimental results on both synthetic and realistic datasets show significant improvements in learning efficiency and accuracy over the existing spike temporal-based learning algorithms. We also evaluate the proposed learning method in an SNN-based multimodal computational model for audiovisual pattern recognition, and it achieves better performance compared with its counterparts.
Xiaoling Luo 0001, Hong Qu 0002, Zhang Yi 0001, Jilun Zhang, Malu Zhang
IEEE Trans. Neural Networks Learn. Syst.6
2023 A Tandem Learning Rule for Effective Training and Rapid Inference of Deep Spiking Neural Networks
abstract
Spiking neural networks (SNNs) represent the most prominent biologically inspired computing model for neuromorphic computing (NC) architectures. However, due to the nondifferentiable nature of spiking neuronal functions, the standard error backpropagation algorithm is not directly applicable to SNNs. In this work, we propose a tandem learning framework that consists of an SNN and an artificial neural network (ANN) coupled through weight sharing. The ANN is an auxiliary structure that facilitates the error backpropagation for the training of the SNN at the spike-train level. To this end, we consider the spike count as the discrete neural representation in the SNN and design an ANN neuronal activation function that can effectively approximate the spike count of the coupled SNN. The proposed tandem learning rule demonstrates competitive pattern recognition and regression capabilities on both the conventional frame- and event-based vision datasets, with at least an order of magnitude reduced inference time and total synaptic operations over other state-of-the-art SNN implementations. Therefore, the proposed tandem learning rule offers a novel solution to training efficient, low latency, and high-accuracy deep SNNs with low computing resources.
Jibin Wu, Yansong Chua, Malu Zhang, Guoqi Li 0002, Haizhou Li 0001, Kay Chen Tan
IEEE Trans. Neural Networks Learn. Syst.3
2022 A Hybrid Learning Framework for Deep Spiking Neural Networks with One-Spike Temporal Coding
abstract
Bio-inspired spiking neural networks (SNNs) are compelling candidates for spatio-temporal information processing on ultra-low power neuromorphic computing chips. However, the existing SNN training methods have not fully exploited the temporal information of spikes that plays a critical role in sparse information representation and communication. Hereby, we present a hybrid learning framework for deep SNNs with one-spike temporal coding to make full utilization of the spike timing. We first propose a novel ANN-to-SNN conversion method based on forward propagation mechanisms of ANNs and SNNs to offer a good initialization for SNNs. The performance of the converted SNN is further improved by training with a timing-based backpropagation (BP) method. Experimental results demonstrate that the proposed hybrid learning framework can achieve competitive accuracies on both visual and audio recognition tasks with significantly improved training efficiency over direct SNN BP methods.
Jibin Wu, Malu Zhang, Qi Liu 0005, Haizhou Li 0001
ICASSP3
2022 Temporal-Sequential Learning with Columnar-Structured Spiking Neural Networks
Xiaoling Luo 0001, Yi Chen 0034, Malu Zhang, Hong Qu 0002
ICONIP (4)4
2022 Signed Neuron with Memory: Towards Simple, Accurate and High-Efficient ANN-SNN Conversion
abstract
Spiking Neural Networks (SNNs) are receiving increasing attention due to their biological plausibility and the potential for ultra-low-power event-driven neuromorphic hardware implementation. Due to the complex temporal dynamics and discontinuity of spikes, training SNNs directly usually suffers from high computing resources and a long training time. As an alternative, SNN can be converted from a pre-trained artificial neural network (ANN) to bypass the difficulty in SNNs learning. However, the existing ANN-to-SNN methods neglect the inconsistency of information transmission between synchronous ANNs and asynchronous SNNs. In this work, we first analyze how the asynchronous spikes in SNNs may cause conversion errors between ANN and SNN. To address this problem, we propose a signed neuron with memory function, which enables almost no accuracy loss during the conversion process, and maintains the properties of asynchronous transmission in the converted SNNs. We further propose a new normalization method, named neuron-wise normalization, to significantly shorten the inference latency in the converted SNNs. We conduct experiments on challenging datasets including CIFAR10 (95.44% top-1), CIFAR100 (78.3% top-1) and ImageNet (73.16% top-1). Experimental results demonstrate that the proposed method outperforms the state-of-the-art works in terms of accuracy and inference time. The code is available at https://github.com/ppppps/ANN2SNNConversion_SNM_NeuronNorm.
Malu Zhang, Yi Chen 0034, Hong Qu 0002
IJCAI2
2022 Training Spiking Neural Networks with Local Tandem Learning
abstract
Spiking neural networks (SNNs) are shown to be more biologically plausible and energy efficient over their predecessors. However, there is a lack of an efficient and generalized training method for deep SNNs, especially for deployment on analog computing substrates. In this paper, we put forward a generalized learning rule, termed Local Tandem Learning (LTL). The LTL rule follows the teacher-student learning approach by mimicking the intermediate feature representations of a pre-trained ANN. By decoupling the learning of network layers and leveraging highly informative supervisor signals, we demonstrate rapid network convergence within five training epochs on the CIFAR-10 dataset while having low computational complexity. Our experimental results have also shown that the SNNs thus trained can achieve comparable accuracies to their teacher ANNs on CIFAR-10, CIFAR-100, and Tiny ImageNet datasets. Moreover, the proposed LTL rule is hardware friendly. It can be easily implemented on-chip to perform fast parameter calibration and provide robustness against the notorious device non-ideality issues. It, therefore, opens up a myriad of opportunities for training and deployment of SNN on ultra-low-power mixed-signal neuromorphic computing chips.
Qu Yang, Jibin Wu, Malu Zhang, Yansong Chua, Xinchao Wang, Haizhou Li 0001
NeurIPS3
2022 Progressive Tandem Learning for Pattern Recognition With Deep Spiking Neural Networks
abstract
Spiking neural networks (SNNs) have shown clear advantages over traditional artificial neural networks (ANNs) for low latency and high computational efficiency, due to their event-driven nature and sparse communication. However, the training of deep SNNs is not straightforward. In this paper, we propose a novel ANN-to-SNN conversion and layer-wise learning framework for rapid and efficient pattern recognition, which is referred to as progressive tandem learning. By studying the equivalence between ANNs and SNNs in the discrete representation space, a primitive network conversion method is introduced that takes full advantage of spike count to approximate the activation value of ANN neurons. To compensate for the approximation errors arising from the primitive network conversion, we further introduce a layer-wise learning method with an adaptive training scheduler to fine-tune the network weights. The progressive tandem learning framework also allows hardware constraints, such as limited weight precision and fan-in connections, to be progressively imposed during training. The SNNs thus trained have demonstrated remarkable classification and regression capabilities on large-scale object recognition, image reconstruction, and speech separation tasks, while requiring at least an order of magnitude reduced inference time and synaptic operations than other state-of-the-art SNN implementations. It, therefore, opens up a myriad of opportunities for pervasive mobile and embedded devices with a limited power budget.
Jibin Wu, Chenglin Xu, Daquan Zhou, Malu Zhang, Haizhou Li 0001, Kay Chen Tan
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Rectified Linear Postsynaptic Potential Function for Backpropagation in Deep Spiking Neural Networks
abstract
Spiking neural networks (SNNs) use spatiotemporal spike patterns to represent and transmit information, which are not only biologically realistic but also suitable for ultralow-power event-driven neuromorphic implementation. Just like other deep learning techniques, deep SNNs (DeepSNNs) benefit from the deep architecture. However, the training of DeepSNNs is not straightforward because the well-studied error backpropagation (BP) algorithm is not directly applicable. In this article, we first establish an understanding as to why error BP does not work well in DeepSNNs. We then propose a simple yet efficient rectified linear postsynaptic potential function (ReL-PSP) for spiking neurons and a spike-timing-dependent BP (STDBP) learning algorithm for DeepSNNs where the timing of individual spikes is used to convey information (temporal coding), and learning (BP) is performed based on spike timing in an event-driven manner. We show that DeepSNNs trained with the proposed single spike time-based learning algorithm can achieve the state-of-the-art classification accuracy. Furthermore, by utilizing the trained model parameters obtained from the proposed STDBP learning algorithm, we demonstrate ultralow-power inference operations on a recently proposed neuromorphic inference accelerator. The experimental results also show that the neuromorphic hardware consumes 0.751 mW of the total power consumption and achieves a low latency of 47.71 ms to classify an image from the Modified National Institute of Standards and Technology (MNIST) dataset. Overall, this work investigates the contribution of spike timing dynamics for information encoding, synaptic plasticity, and decision-making, providing a new perspective to the design of future DeepSNNs and neuromorphic hardware.
Malu Zhang, Jibin Wu, Ammar Belatreche, Burin Amornpaisannon, Venkata Pavan Kumar Miriyala, Hong Qu 0002, Yansong Chua, Trevor E. Carlson, Haizhou Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2021 Deep Spiking Neural Network with Neural Oscillation and Spike-Phase Information
abstract
Deep spiking neural network (DSNN) is a promising computational model towards artificial intelligence. It benefits from both the DNNs and SNNs through a hierarchy structure to extract multiple levels of abstraction and the event-driven computational manner to provide ultra-low-power neuromorphic implementation, respectively. However, how to efficiently train the DSNNs remains an open question because of the non-differentiable spike function that prevents the traditional back-propagation (BP) learning algorithm directly applied to DSNNs. Here, inspired by the findings from the biological neural networks, we address the above-mentioned problem by introducing neural oscillation and spike-phase information to DSNNs. Specifically, we propose an Oscillation Postsynaptic Potential (Os-PSP) and phase-locking active function, and further put forward a new spiking neuron model, namely Resonate Spiking Neuron (RSN). Based on the RSN, we propose a Spike-Level-Dependent Back-Propagation (SLDBP) learning algorithm for DSNNs. Experimental results show that the proposed learning algorithm resolves the problems caused by the incompatibility between the BP learning algorithm and SNNs, and achieves state-of-the-art performance in single spike-based learning algorithms. This work investigates the contribution of introducing biologically inspired mechanisms, such as neural oscillation and spike-phase information to DSNNs and providing a new perspective to design future DSNNs.
Yi Chen 0034, Hong Qu 0002, Malu Zhang
AAAI3
2021 GCC-PHAT with Speech-oriented Attention for Robotic Sound Source Localization
abstract
Robotic audition is a basic sense that helps robots perceive the surroundings and interact with humans. Sound Source Localization (SSL) is an essential module for a robotic system. However, the performance of most sound source localization techniques degrades in noisy and reverberant environments due to inaccurate Time Difference of Arrival (TDoA) estimation. In robotic sound source localization, we are more interested in detecting the arrival of human speech than other sound sources. Ideally, we expect an effective TDoA estimation to respond only to speech signals, while masking off other interferences. In this paper, we propose a novel technique that learns to attend to speech fundamental frequency and harmonics while suppressing noise interference and reverberation. The novel TDoA feature is referred to as Generalized Cross Correlation with Phase Transform and Speech Mask (GCC-PHAT-SM). We perform sound source localization experiments on real-world data captured from a robotic platform. Experiments show that GCC-PHAT-SM feature significantly outperforms traditional Generalized Cross Correlation (GCC) feature in noisy and reverberant acoustic environments.
Xinyuan Qian 0001, Zihan Pan, Malu Zhang, Haizhou Li 0001
ICRA4
2021 HuRAI: A brain-inspired computational model for human-robot auditory interface
Jibin Wu, Qi Liu 0005, Malu Zhang, Zihan Pan, Haizhou Li 0001, Kay Chen Tan
Neurocomputing3
2021 A new recursive least squares-based learning algorithm for spiking neurons
Hong Qu 0002, Xiaoling Luo 0001, Yi Chen 0034, Malu Zhang, Zefang Li
Neural Networks6
2021 Multi-Tone Phase Coding of Interaural Time Difference for Sound Source Localization With Spiking Neural Networks
abstract
Mammals exhibit remarkable capability of detecting and localizing sound sources in complex acoustic environments by using binaural cues in the spiking manner. Emulating the auditory process for sound source localization (SSL) by mammals, we propose a computational model for accurate and robust SSL under the neuromorphic spiking neural network (SNN) framework. The center of this model is a Multi-Tone Phase Coding (MTPC) scheme, which encodes the interaural time difference (ITD) between binaural pure tones into discriminative spike patterns that can be directly classified by SNNs. As such, SSL can be implemented as an event-driven task on highly efficient, neuromorphic parallel processors. We evaluate the proposed computational model on a directional audio dataset recorded from a microphone array in a realistic acoustic environment with background noise, obstruction, reflection, and other interferences. We report superior localization capability with a mean absolute error (MAE) of 1.02°or 100% classification accuracy with an angle resolution of 5°, which surpasses other SNN-based biologically plausible neuromorphic approaches by a relatively large margin and on par with human performance in similar tasks. This study opens up many application opportunities in human-robot interaction where energy efficiency is crucial. As a case study, we successfully deploy the proposed SSL system in a robotic platform to track the speaker and orient the robot's attention.
Zihan Pan, Malu Zhang, Jibin Wu, Haizhou Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Supervised learning in spiking neural networks with synaptic delay-weight plasticity
Malu Zhang, Jibin Wu, Ammar Belatreche, Zihan Pan, Xiurui Xie, Yansong Chua, Guoqi Li 0002, Hong Qu 0002, Haizhou Li 0001
Neurocomputing1
2020 An end-to-end functional spiking model for sequential feature learning
Xiurui Xie, Guisong Liu, Guolin Sun, Malu Zhang, Hong Qu 0002
Knowl. Based Syst.5
2019 MPD-AL: An Efficient Membrane Potential Driven Aggregate-Label Learning Algorithm for Spiking Neurons
abstract
One of the long-standing questions in biology and machine learning is how neural networks may learn important features from the input activities with a delayed feedback, commonly known as the temporal credit-assignment problem. The aggregate-label learning is proposed to resolve this problem by matching the spike count of a neuron with the magnitude of a feedback signal. However, the existing threshold-driven aggregate-label learning algorithms are computationally intensive, resulting in relatively low learning efficiency hence limiting their usability in practical applications. In order to address these limitations, we propose a novel membrane-potential driven aggregate-label learning algorithm, namely MPD-AL. With this algorithm, the easiest modifiable time instant is identified from membrane potential traces of the neuron, and guild the synaptic adaptation based on the presynaptic neurons’ contribution at this time instant. The experimental results demonstrate that the proposed algorithm enables the neurons to generate the desired number of spikes, and to detect useful clues embedded within unrelated spiking activities and background noise with a better learning efficiency over the state-of-the-art TDP1 and Multi-Spike Tempotron algorithms. Furthermore, we propose a data-driven dynamic decoding scheme for practical classification tasks, of which the aggregate labels are hard to define. This scheme effectively improves the classification accuracy of the aggregate-label learning algorithms as demonstrated on a speech recognition task.
Malu Zhang, Jibin Wu, Yansong Chua, Xiaoling Luo 0001, Zihan Pan, Haizhou Li 0001
AAAI1
2019 Neural Population Coding for Effective Temporal Classification
abstract
Neural encoding plays an important role in faithfully describing the temporally rich patterns, whose instances include human speech and environmental sounds. To classify such spatio-temporal patterns with the Spiking Neural Networks (SNNs), how these patterns are encoded has a direct impact on the complexity of the task. In this paper, we study several existing temporal and population coding schemes in speech and audio recognition. We show that, with population neural coding, the encoded patterns are linearly separable using the Support Vector Machine (SVM). We note that the population neural coding effectively project the temporal information into the spatial domain, thus improving linear separability of the patterns. We achieve an accuracy of 95% and 100% on TIDIGITS and RWCP datasets respectively with SVM classifier. We further implement the Tempotron as an SNN-based classifier on the same datasets and achieve similar results. The study suggests that an effective neural coding scheme is just as important as the classifier.
Zihan Pan, Jibin Wu, Malu Zhang, Haizhou Li 0001, Yansong Chua
IJCNN3
2019 Deep Spiking Neural Network with Spike Count based Learning Rule
abstract
Deep spiking neural networks (SNNs) support asynchronous event-driven computation, massive parallelism and demonstrate great potential to improve the energy efficiency of its synchronous analog counterpart. However, insufficient attention has been paid to neural encoding when designing SNN learning rules. Remarkably, the temporal credit assignment has been performed on rate-coded spiking inputs, leading to poor learning efficiency. In this paper, we introduce a novel spike-based learning rule for rate-coded deep SNNs, whereby the spike count of each neuron is used as a surrogate for gradient backpropagation. We evaluate the proposed learning rule by training deep spiking multi-layer perceptron (MLP) and spiking convolutional neural network (CNN) on the UCI machine learning and MNIST handwritten digit datasets. We show that the proposed learning rule achieves state-of-the-art accuracies on all benchmark datasets. The proposed learning rule allows introducing latency, spike rate and hardware constraints into the SNN learning, which is superior to the indirect approach in which conventional artificial neural networks are first trained and then converted to SNNs. Hence, it allows direct deployment to the neuromorphic hardware and supports efficient inference. Notably, a test accuracy of 98.40% was achieved on the MNIST dataset in our experiments with only 10 simulation time steps, when the same latency constraint is imposed during training.
Jibin Wu, Yansong Chua, Malu Zhang, Qu Yang, Guoqi Li 0002, Haizhou Li 0001
IJCNN3
2019 Competitive STDP-based Feature Representation Learning for Sound Event Classification
abstract
Humans are good at discriminating environmental sounds and associating them with opportunities or dangers. While the deep learning approach to sound event classification (SEC) is achieving human parity, unsolved problems remain, instances include high computational cost, requirement of massive labeled training data, and question of biological plausibility. Motivated by the human auditory system, we propose a biologically plausible SEC system, which integrates the auditory front-end, population coding, competitive spike-timing-dependent plasticity (STDP) based feature representation learning and supervised temporal classification into a unified spiking neural network (SNN) system. The proposed SEC system achieves a classification accuracy on the RWCP database that is on par with other competitive baseline systems. Furthermore, the STDP-based feature representation learning shows low intra-class variability and high inter-class variability in our experiments, which is highly desirable for pattern classification tasks.
Jibin Wu, Malu Zhang, Haizhou Li 0001, Yansong Chua
IJCNN2
2019 Robust Sound Recognition: A Neuromorphic Approach
Jibin Wu, Zihan Pan, Malu Zhang, Rohan Kumar Das, Yansong Chua, Haizhou Li 0001
INTERSPEECH3
2019 The maximum points-based supervised learning rule for spiking neural networks
Xiurui Xie, Guisong Liu, Hong Qu 0002, Malu Zhang
Soft Comput.5
2019 A Highly Effective and Robust Membrane Potential-Driven Supervised Learning Method for Spiking Neurons
abstract
Spiking neurons are becoming increasingly popular owing to their biological plausibility and promising computational properties. Unlike traditional rate-based neural models, spiking neurons encode information in the temporal patterns of the transmitted spike trains, which makes them more suitable for processing spatiotemporal information. One of the fundamental computations of spiking neurons is to transform streams of input spike trains into precisely timed firing activity. However, the existing learning methods, used to realize such computation, often result in relatively low accuracy performance and poor robustness to noise. In order to address these limitations, we propose a novel highly effective and robust membrane potential-driven supervised learning (MemPo-Learn) method, which enables the trained neurons to generate desired spike trains with higher precision, higher efficiency, and better noise robustness than the current state-of-the-art spiking neuron learning methods. While the traditional spike-driven learning methods use an error function based on the difference between the actual and desired output spike trains, the proposed MemPo-Learn method employs an error function based on the difference between the output neuron membrane potential and its firing threshold. The efficiency of the proposed learning method is further improved through the introduction of an adaptive strategy, called skip scan training strategy, that selectively identifies the time steps when to apply weight adjustment. The proposed strategy enables the MemPo-Learn method to effectively and efficiently learn the desired output spike train even when much smaller time steps are used. In addition, the learning rule of MemPo-Learn is improved further to help mitigate the impact of the input noise on the timing accuracy and reliability of the neuron firing dynamics. The proposed learning method is thoroughly evaluated on synthetic data and is further demonstrated on real-world classification tasks. Experimental results show that the proposed method can achieve high learning accuracy with a significant improvement in learning time and better robustness to different types of noise.
Malu Zhang, Hong Qu 0002, Ammar Belatreche, Yi Chen 0034, Zhang Yi 0001
IEEE Trans. Neural Networks Learn. Syst.1
2018 Botnet Detection based on Fuzzy Association Rules
abstract
Difficult to be detected in complex network environments, botnets have been huge threats to network security. As the circumscriptions of normal traffics and botnet traffics are blurring, the commonly used botnet detection methods based on traffic analysis often result in high false positive rates. To overcome this issue, we propose an effective botnet detection method based on fuzzy association rules. The proposed method can calculate the features of botnet traffic accurately, which can be used to recognize the normal traffic and botnet. We first collect the data in the laboratory by setting different botnets in the controlled experiment. The botnet traffic features, association rules support, trust and membership are calculated by the proposed method, which are further used to distinguish the type of botnet. When our method is compared with other methods in our data set, we find the former performs better. For the generality, we also test our method on the public data set and also find the higher accuracy rates, which demonstrates the proposed method is effective in detecting the botnets.
Fengmao Lv, Quanhui Liu, Malu Zhang, Xiaosong Zhang 0001
ICPR4
2017 A Fast Precise-Spike and Weight-Comparison Based Learning Approach for Evolving Spiking Neural Networks
Lin Zuo, Hong Qu 0002, Malu Zhang
ICONIP (3)4
2017 A Dynamic Region Generation Algorithm for Image Segmentation Based on Spiking Neural Network
Lin Zuo, Linyao Ma, Yanqing Xiao, Malu Zhang, Hong Qu 0002
ICONIP (3)4
2017 Efficient training of supervised spiking neural networks via the normalized perceptron based learning rule
Xiurui Xie, Hong Qu 0002, Guisong Liu, Malu Zhang
Neurocomputing4
2017 Supervised learning in spiking neural networks with noise-threshold
Malu Zhang, Hong Qu 0002, Xiurui Xie, Jürgen Kurths
Neurocomputing1
2015 Improved perception-based spiking neuron learning rule for real-time user authentication
Hong Qu 0002, Xiurui Xie, Yongshuai Liu, Malu Zhang, Li Lu 0001
Neurocomputing4