Dehao Zhang

dblp:194/7033 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Computer networks · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Training-Free ANN-to-SNN Conversion for High-Performance Spiking Transformers
abstract
Leveraging the event-driven paradigm, Spiking Neural Networks (SNNs) offer a promising approach for constructing energy-efficient Transformer architectures. Compared to directly trained Spiking Transformers, ANN-to-SNN conversion methods bypass the high training costs. However, existing methods still suffer from notable limitations, failing to effectively handle nonlinear operations in Transformer architectures and requiring additional fine-tuning processes for pre-trained ANNs. To address these issues, we propose a high-performance and training-free ANN-to-SNN conversion framework tailored for Transformer architectures. Specifically, we introduce a Multi-basis Exponential Decay (MBE) neuron, which employs an exponential decay strategy and multi-basis encoding method to efficiently approximate various nonlinear operations. It removes the requirement for weight modifications in pre-trained ANNs. Extensive experiments across diverse tasks (CV, NLU, NLG) and mainstream Transformer architectures (ViT, RoBERTa, GPT-2) demonstrate that our method achieves near-lossless conversion accuracy with significantly lower latency. This provides a promising pathway for the efficient and scalable deployment of Spiking Transformers in real-world applications.
Wenjie Wei, Dehao Zhang, Shuai Wang 0058, Qian Sun 0014, Jieyuan Zhang, Malu Zhang
AAAI4
2025 Rethinking Spiking Self-Attention Mechanism: Implementing a-XNOR Similarity Calculation in Spiking Transformers
abstract
Transformers significantly raise the performance limits across various tasks, spurring research into integrating them into spiking neural networks. However, a notable performance gap remains between existing spiking Transformers and their artificial neural network counterparts. Here, we first analyze the cause of this gap and attribute it to the dot product’s ineffectiveness in measuring similarity between spiking queries and keys, due to numerous non-spiking events. To address this, we propose a novel α-XNOR similarity measure tailored for spike trains. It redefines the correlation between non-spike pairs as a specific value α, effectively overcoming the limitations of dot-product similarity. Furthermore, considering the sparse nature of spike trains where spikes carry more information than non-spikes, the α-XNOR similarity correspondingly highlights the distinct importance of spikes over non-spikes. Extensive experiments demonstrate that α-XNOR similarity significantly improves performance across different spiking Transformer architectures on various static and neuromorphic datasets, further revealing the potential of spiking Transformers.
Yichen Xiao, Shuai Wang 0058, Dehao Zhang, Wenjie Wei, Yimeng Shan, Yulin Jiang, Malu Zhang
CVPR3
2025 Memory-Free and Parallel Computation for Quantized Spiking Neural Networks
abstract
Quantized Spiking Neural Networks (QSNNs) offer superior energy efficiency and are well-suited for deployment on resource-limited edge devices. However, limited bit-width weight and membrane potential result in a notable performance decline. In this study, we first identify a new underlying cause for this decline: the loss of historical information due to the quantized membrane potential. To tackle this issue, we introduce a memory-free quantization method that captures all historical information without directly storing membrane potentials, resulting in better performance with less memory requirements. To further improve the computational efficiency, we propose a parallel training and asynchronous inference framework that greatly increases training speed and energy efficiency. We combine the proposed memory-free quantization and parallel computation methods to develop a high-performance and efficient QSNN, named MFP-QSNN. Extensive experiments show that our MFP-QSNN achieves state-of-the-art performance on various static and neuromorphic image datasets, requiring less memory and faster training speeds. The efficiency and efficacy of the MFP-QSNN highlight its potential for energy-efficient neuromorphic computing.
Dehao Zhang, Shuai Wang 0058, Yichen Xiao, Wenjie Wei, Yimeng Shan, Malu Zhang, Yang Yang 0002
ICASSP1
2025 Spiking Vision Transformer with Saccadic Attention
abstract
The combination of Spiking Neural Networks (SNNs) and Vision Transformers (ViTs) holds potential for achieving both energy efficiency and high performance, particularly suitable for edge vision applications. However, a significant performance gap still exists between SNN-based ViTs and their ANN counterparts. Here, we first analyze why SNN-based ViTs suffer from limited performance and identify a mismatch between the vanilla self-attention mechanism and spatio-temporal spike trains. This mismatch results in degraded spatial relevance and limited temporal interactions. To address these issues, we draw inspiration from biological saccadic attention mechanisms and introduce an innovative Saccadic Spike Self-Attention (SSSA) method. Specifically, in the spatial domain, SSSA employs a novel spike distribution-based method to effectively assess the relevance between Query and Key pairs in SNN-based ViTs. Temporally, SSSA employs a saccadic interaction module that dynamically focuses on selected visual areas at each timestep and significantly enhances whole scene understanding through temporal interactions. Building on the SSSA mechanism, we develop a SNN-based Vision Transformer (SNN-ViT). Extensive experiments across various visual tasks demonstrate that SNN-ViT achieves state-of-the-art performance with linear computational complexity. The effectiveness and efficiency of the SNN-ViT highlight its potential for power-critical edge vision applications.
Shuai Wang 0058, Malu Zhang, Dehao Zhang, Ammar Belatreche, Yichen Xiao, Yimeng Shan, Qian Sun 0014, Enqi Zhang, Yang Yang 0002
ICLR3
2025 Binary Event-Driven Spiking Transformer
abstract
Transformer-based Spiking Neural Networks (SNNs) introduce a novel event-driven self-attention paradigm that combines the high performance of Transformers with the energy efficiency of SNNs. However, the larger model size and increased computational demands of the Transformer structure limit their practicality in resource-constrained scenarios. In this paper, we integrate binarization techniques into Transformer-based SNNs and propose the Binary Event-Driven Spiking Transformer, i.e. BESTformer. The proposed BESTformer can significantly reduce storage and computational demands by representing weights and attention maps with a mere 1-bit. However, BESTformer suffers from a severe performance drop from its full-precision counterpart due to the limited representation capability of binarization. To address this issue, we propose a Coupled Information Enhancement (CIE) method, which consists of a reversible framework and information enhancement distillation. By maximizing the mutual information between the binary model and its full-precision counterpart, the CIE method effectively mitigates the performance degradation of the BESTformer. Extensive experiments on static and neuromorphic datasets demonstrate that our method achieves superior performance to other binary SNNs, showcasing its potential as a compact yet high-performance model for resource-limited edge devices. The repository of this paper is available at https://github.com/CaoHLin/BESTFormer.
Honglin Cao, Zijian Zhou 0005, Wenjie Wei, Ammar Belatreche, Dehao Zhang, Malu Zhang, Yang Yang 0002, Haizhou Li 0001
IJCAI6
2025 Bipolar Self-attention for Spiking Transformers
abstract
Harnessing the event-driven characteristic, Spiking Neural Networks (SNNs) present a promising avenue toward energy-efficient Transformer architectures. However, existing Spiking Transformers still suffer significant performance gaps compared to their Artificial Neural Network counterparts. Through comprehensive analysis, we attribute this gap to these two factors. First, the binary nature of spike trains limits Spiking Self-attention (SSA)’s capacity to capture negative–negative and positive–negative membrane potential interactions on Querys and Keys. Second, SSA typically omits Softmax functions to avoid energy-intensive multiply-accumulate operations, thereby failing to maintain row-stochasticity constraints on attention scores. To address these issues, we propose a Bipolar Self-attention (BSA) paradigm, effectively modeling multi-polar membrane potential interactions with a fully spike-driven characteristic. Specifically, we demonstrate that ternary matrix multiplication provides a closer approximation to real-valued computation on both distribution and local correlation, enabling clear differentiation between homopolar and heteropolar interactions. Moreover, we propose a shift-based Softmax approximation named Shiftmax, which efficiently achieves low-entropy activation and partly maintains row-stochasticity without non-linear operation, enabling precise attention allocation. Extensive experiments show that BSA achieves substantial performance improvements across various tasks, including image classification, semantic segmentation, and event-based tracking. These results establish its potential as a fundamental building block for energy-efficient Spiking Transformers.
Shuai Wang 0058, Malu Zhang, Dehao Zhang, Yimeng Shan, Jieyuan Zhang, Yichen Xiao, Honglin Cao, Zeyu Ma 0002, Yang Yang 0002, Haizhou Li 0001
NeurIPS4
2025 Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence Modeling
abstract
The explosive growth in sequence length has intensified the demand for effective and efficient long sequence modeling. Benefiting from intrinsic oscillatory membrane dynamics, Resonate-and-Fire (RF) neurons can efficiently extract frequency components from input signals and encode them into spatiotemporal spike trains, making them well-suited for long sequence modeling. However, RF neurons exhibit limited effective memory capacity and a trade-off between energy efficiency and training speed on complex temporal tasks. Inspired by the dendritic structure of biological neurons, we propose a Dendritic Resonate-and-Fire (D-RF) model, which explicitly incorporates a multi-dendritic and soma architecture. Each dendritic branch encodes specific frequency bands by utilizing the intrinsic oscillatory dynamics of RF neurons, thereby collectively achieving comprehensive frequency representation. Furthermore, we introduce an adaptive threshold mechanism into the soma structure. This mechanism adjusts the firing threshold according to historical spiking activity, thereby reducing redundant spikes while maintaining training efficiency in long-sequence tasks. Extensive experiments demonstrate that our method maintains competitive accuracy while substantially ensuring sparse spikes without compromising computational efficiency during training. These results underscore its potential as an effective and efficient solution for long sequence modeling on edge platforms.
Dehao Zhang, Malu Zhang, Shuai Wang 0058, Wenjie Wei, Zeyu Ma 0002, Guoqing Wang 0001, Yang Yang 0002, Haizhou Li 0001
NeurIPS1
2025 Efficient blockchain synchronization mechanism over NDN based on directed Interest forwarding
Dehao Zhang, Jiapeng Xiu, Zhengqiu Yang, Shao-Yong Guo 0001
Comput. Commun.1
2025 Efficient Automatic Modulation Classification in Nonterrestrial Networks With SNN-Based Transformer
abstract
With the development of informatization of IoT devices, nonterrestrial networks (NTNs) are becoming more and more important. NTN, including air and space networks, face challenges, such as high-computational complexity, bandwidth requirements, and memory constraints. An intelligent automatic modulation classification (AMC) mechanism based on neural networks plays a pivotal role in enhancing spectrum efficiency, throughput, and link reliability. Past work in AMC has evolved from likelihood-based and feature-based methods to traditional machine learning techniques and, more recently, to deep neural networks (DNNs). However, existing DNN architectures pose challenges for NTN due to high-computational complexity, bandwidth requirements, and memory consumption. Addressing this problems, we proposes a spiking transformer-based model for AMC, exploiting temporal dynamics for enhanced performance. Biologically inspired spiking neural networks enable us to exploit the sparse and binarized activation properties of spiking neurons, allowing us to build AMC models with high-energy efficiency and high availability that can be used in NTN systems. Furthermore, we introduce a weight binarization method to reduce the model size, which also further reduces the bandwidth and memory requirements of AMC in NTN edge deployment. Experimental results demonstrate the superiority of our approach over state-of-the-art methods, with the binarized model achieving comparable accuracy at a fraction of the size.
Dingyi Zeng, Yichen Xiao, Wanlong Liu, Huilin Du, Enqi Zhang, Dehao Zhang, Malu Zhang, Wenyu Chen 0001
IEEE Internet Things J.6
2025 Ternary spike-based neuromorphic signal processing system
Shuai Wang 0058, Dehao Zhang, Ammar Belatreche, Yichen Xiao, Hongyu Qing, Wenjie Wei, Malu Zhang, Yang Yang 0002
Neural Networks2
2024 Global-Local Convolution with Spiking Neural Networks for Energy-efficient Keyword Spotting
abstract
25th Annual Conference of the International Speech Communication Association, Interspeech 2024, Kos, Greece, September 1-5, 2024. ISCA 2024
Shuai Wang 0058, Dehao Zhang, Wenjie Wei, Jibin Wu, Malu Zhang
INTERSPEECH2
2024 Spike-based Neuromorphic Model for Sound Source Localization
abstract
Biological systems possess remarkable sound source localization (SSL) capabilities that are critical for survival in complex environments. This ability arises from the collaboration between the auditory periphery, which encodes sound as precisely timed spikes, and the auditory cortex, which performs spike-based computations. Inspired by these biological mechanisms, we propose a novel neuromorphic SSL framework that integrates spike-based neural encoding and computation. The framework employs Resonate-and-Fire (RF) neurons with a phase-locking coding (RF-PLC) method to achieve energy-efficient audio processing. The RF-PLC method leverages the resonance properties of RF neurons to efficiently convert audio signals to time-frequency representation and encode interaural time difference (ITD) cues into discriminative spike patterns. In addition, biological adaptations like frequency band selectivity and short-term memory effectively filter out many environmental noises, enhancing SSL capabilities in real-world settings. Inspired by these adaptations, we propose a spike-driven multi-auditory attention (MAA) module that significantly improves both the accuracy and robustness of the proposed SSL framework. Extensive experimentation demonstrates that our SSL framework achieves state-of-the-art accuracy in SSL tasks. Furthermore, it shows exceptional noise robustness and maintains high accuracy even at very low signal-to-noise ratios. By mimicking biological hearing, this neuromorphic approach contributes to the development of high-performance and explainable artificial intelligence systems capable of superior performance in real-world environments.
Dehao Zhang, Shuai Wang 0058, Ammar Belatreche, Wenjie Wei, Yichen Xiao, Haorui Zheng, Zijian Zhou 0005, Malu Zhang, Yang Yang 0002
NeurIPS1
2022 Data Association Between Event Streams and Intensity Frames Under Diverse Baselines
Dehao Zhang, Qiankun Ding, Peiqi Duan 0002, Chu Zhou, Boxin Shi
ECCV (7)1
2016 Enhancing Traffic Engineering Performance and Flow Manageability in Hybrid SDN
abstract
Hybrid Software-Defined Networking (HSDN) is a transitional networking form of SDN where SDN elements are partially deployed in traditional networks. Previous researches show that redirecting every flow of source-destination pair through at least one SDN switch can obtain flow manageability, e.g., access control and traffic measurement. Intuitively, the selection of SDN switch as the waypoint for every flow has a significant effect on the Traffic Engineering (TE) performance, such as maximum link utilization and routing efficiency. And it is worth noting that SDN switch can split traffic to the outgoing links to exactly profit the TE performance. In this paper, from the perspective of TE performance, we propose a flow routing and splitting (FRS) algorithm whereby we jointly determine an appropriate SDN switch for every flow as the waypoint, as well as optimizing the traffic splitting fractions for every SDN switch among its outgoing links to minimize the maximum link utilization. We conduct simulations with different SDN deployment rate. The results indicate that, when 20% of the SDN switches are deployed, the proposed FRS algorithm can obtain a lower maximum link utilization compared with other state-of-art works. Not only that, FRS algorithm can also generate a little longer paths for every flow on the average, which has a limited influence on the routing efficiency.
Cheng Ren, Sheng Wang 0006, Jing Ren 0002, Xiong Wang 0001, Tongyu Song, Dehao Zhang
GLOBECOM6