EDBT 2026 Demo / reviewers in the wild / expert
Zhaokun Zhou
dblp:241/5641
· DBLP profile ↗
23ranked-venue papers
5as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 3 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Computer networks · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spikingformer: A Key Foundation Model for Spiking Neural NetworksabstractSpiking neural networks (SNNs) offer a promising energy-efficient alternative to artificial neural networks, due to their event-driven spiking computation. However, some foundation SNN backbones (including Spikformer and SEW ResNet) suffer from non-spike computations (integer-float multiplications) caused by the structure of their residual connections. These non-spike computations increase SNNs' power consumption and make them unsuitable for deployment on mainstream neuromorphic hardware. In this paper, we analyze the spike-driven behavior of the residual connection methods in SNNs. We then present Spikingformer, a novel spiking transformer backbone that merges the MS Residual connection with Self-Attention in a biologically plausible way to address the non-spike computation challenge in Spikformer while maintaining global modeling capabilities. We evaluate Spikingformer across 13 datasets spanning large static images, neuromorphic data, and natural language tasks, and demonstrate the effectiveness and universality of Spikingformer, setting a vital benchmark for spiking neural networks. In addition, with the spike-driven features and global modeling capabilities, Spikingformer is expected to become a more efficient general-purpose SNN backbone towards energy-efficient artificial intelligence. Chenlin Zhou, Liutao Yu, Zhaokun Zhou, Han Zhang 0035, Jiaqi Wang 0003, Zhengyu Ma, Yonghong Tian 0001 |
AAAI | 3 |
| 2026 | Rethinking Neuromorphic Object Detection with Hybrid Dynamic Interaction Transformers
Dianze Li, Jianing Li 0001, Xu Liu 0006, Zhaokun Zhou, Xiaopeng Fan 0001, Yonghong Tian 0001 |
Int. J. Comput. Vis. | 4 |
| 2026 | Spatially-enhanced Spiking neural network for efficient point cloud analysis
Yijie Lu, Zhiyi Pan 0001, Renrui Zhang, Yanhao Jia, Kaiwei Che, Zhaokun Zhou |
Neural Networks | 6 |
| 2026 | SAH-NeRF: Enhancing NeRF on Novel View Synthesis With an SNN-ANN Hybrid FrameworkabstractNeural Radiance Field (NeRF) utilizes Artificial Neural Networks (ANNs) to map 3D points and their corresponding 2D viewing directions to colors and densities. However, this approach encounters several challenges due to the inherent characteristics of ANNs. Firstly, ANNs tend to extract smooth features across sampled points, which makes it difficult to accurately represent the non-smooth variations between object surfaces and the surrounding air. Secondly, ANNs compute densities and colors independently for each point, failing to account for the sequential dependencies among points along the same ray. To tackle these issues, Spiking Neural Networks (SNNs) are introduced, which are better at processing sequential information and non-smooth representations. In this work, we propose a novel hybrid NeRF framework called SAH-NeRF that combines ANNs with SNNs. By harnessing ANNs’ robust representation capabilities alongside SNNs’ strengths in handling non-smooth distributions, our method could significantly improve the performance of three ANN-based NeRFs, surpassing state-of-the-art methods including 3D Gaussian Splatting. Notably, our SAH-NeRF could meanwhile enhance novel view synthesis and reduce energy consumption. Yiqian Chang, Peixi Peng, Zhaokun Zhou, Xuan Wang 0002, Yonghong Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Identity-Preserving Face Privacy Enhancement via Diffusion Models with Cognitive-Aware Obfuscation
Haoxuan Tai, Zhaokun Zhou, Yuesheng Zhu |
CogSci | 3 |
| 2025 | SpikingPoint: Rethinking Point as Spike for Efficient 3D Point Cloud AnalysisabstractSpiking Neural Networks (SNNs), due to their unique spike-based inference mechanism, offer low power consumption and biological plausibility. As a fundamental technology for various real-world applications, 3D point cloud analysis faces significant challenges related to high computational overhead and energy-intensive. In fact, each point can be viewed as a specialized spike data containing positional information in 3D space. Therefore, utilizing the spiking features of SNNs to represent point clouds holds significant potential. However, this exploration faces two main challenges: SNN-compatible architecture design and spike-based 3D spatial modeling. In this work, we introduce the SpikingPoint, a pure Spiking Multi-layer Perceptron (MLP) Architecture that leverages the low power consumption of SNNs and the computational efficiency of linear layers. Furthermore, we propose the spiking 3D Position Embedding (SPE) that effectively models the 3D spatial feature into spike-form features. SpikingPoint achieves competitive performance with a small number of parameters while reducing energy consumption. For example, when achieving similar accuracy, the theoretical energy consumption of our method is reduced by 97.1% compared to the mainstream ANN-based KPConv. With fewer parameters, SpikingPoint surpasses the state-of-the-art SNN-based P2SResLNet on ModelNet40 by 1.42% in accuracy. Zhaokun Zhou, Yijie Lu, Jiaqiyu Zhan, Guibo Luo, Yuesheng Zhu |
ICASSP | 1 |
| 2025 | Spiking Transformer with Spatial-Temporal Spiking Self-AttentionabstractSpiking Neural Networks are celebrated for energy efficiency and biological plausibility. Building on Spiking Self-Attention (SSA), Spiking Transformers are extensively studied due to their exceptional performance. However, SSA focuses solely on spatial dimension at each time step, overlooking the crucial features across temporal dimension. To address this, we propose the Spatial-Temporal Spiking Self-Attention (STSSA), a spike-driven mechanism that leverages both spatial and temporal information with negligible additional computational overhead. Specifically, we extract the Representative Spiking Temporal Tokens (RSTT) and apply temporal window masking to the RSTT. These tokens are inserted between the Query and Key to integrate temporal features. Furthermore, we design a Multi-dimensional Learnable Scaling Factor (MLSF) to adapt to STSSA. Our results consistently demonstrate that STSSA outperforms SSA across extensive experiments on sequential, neuromorphic, and static datasets. Notably, STSSA achieves performance improvements of 5.7% and 2.9% over SSA on Sequential CIFAR-100 and CIFAR-10DVS, respectively. STSSA provides a powerful alternative within the family of Spiking Self-Attention mechanisms. Zhaokun Zhou, Li Yuan 0007, Yuesheng Zhu |
ICASSP | 1 |
| 2025 | Multiplication-Free Parallelizable Spiking Neurons with Efficient Spatio-Temporal DynamicsabstractSpiking Neural Networks (SNNs) are distinguished from Artificial Neural Networks (ANNs) for their complex neuronal dynamics and sparse binary activations (spikes) inspired by the biological neural system. Traditional neuron models use iterative step-by-step dynamics, resulting in serial computation and slow training speed of SNNs. Recently, parallelizable spiking neuron models have been proposed to fully utilize the massive parallel computing ability of graphics processing units to accelerate the training of SNNs. However, existing parallelizable spiking neuron models involve dense floating operations and can only achieve high long-term dependencies learning ability with a large order at the cost of huge computational and memory costs. To solve the dilemma of performance and costs, we propose the mul-free channel-wise Parallel Spiking Neuron, which is hardware-friendly and suitable for SNNs’ resource-restricted application scenarios. The proposed neuron imports the channel-wise convolution to enhance the learning ability, induces the sawtooth dilations to reduce the neuron order, and employs the bit-shift operation to avoid multiplications. The algorithm for the design and implementation of acceleration methods is discussed extensively. Our methods are validated in neuromorphic Spiking Heidelberg Digits voices, sequential CIFAR images, and neuromorphic DVS-Lip vision datasets, achieving superior performance over SOTA spiking neurons. Training speed results demonstrate the effectiveness of our acceleration methods, providing a practical reference for future research. Our code is available at Github. Wei Fang 0006, Zhengyu Ma, Zihan Huang, Zhaokun Zhou, Yonghong Tian 0001, Timothée Masquelier |
NeurIPS | 5 |
| 2024 | A Multi-modal Spiking Meta-learner with Brain-Inspired Task-Aware Modulation Scheme
Zhaokun Zhou, Kaiwei Che, Li Yuan 0007 |
ICANN (10) | 2 |
| 2024 | Spike-driven Transformer V2: Meta Spiking Neural Network Architecture Inspiring the Design of Next-generation Neuromorphic ChipsabstractNeuromorphic computing, which exploits Spiking Neural Networks (SNNs) on neuromorphic chips, is a promising energy-efficient alternative to traditional AI. CNN-based SNNs are the current mainstream of neuromorphic computing. By contrast, no neuromorphic chips are designed especially for Transformer-based SNNs, which have just emerged, and their performance is only on par with CNN-based SNNs, offering no distinct advantage. In this work, we propose a general Transformer-based SNN architecture, termed as ``Meta-SpikeFormer", whose goals are: (1) *Lower-power*, supports the spike-driven paradigm that there is only sparse addition in the network; (2) *Versatility*, handles various vision tasks; (3) *High-performance*, shows overwhelming performance advantages over CNN-based SNNs; (4) *Meta-architecture*, provides inspiration for future next-generation Transformer-based neuromorphic chip designs. Specifically, we extend the Spike-driven Transformer in \citet{yao2023spike} into a meta architecture, and explore the impact of structure, spike-driven self-attention, and skip connection on its performance. On ImageNet-1K, Meta-SpikeFormer achieves 80.0\% top-1 accuracy (55M), surpassing the current state-of-the-art (SOTA) SNN baselines (66M) by 3.7\%. This is the first direct training SNN backbone that can simultaneously supports classification, detection, and segmentation, obtaining SOTA results in SNNs. Finally, we discuss the inspiration of the meta SNN architecture for neuromorphic chip design. Man Yao, Tianxiang Hu, Zhaokun Zhou, Yonghong Tian 0001, Bo Xu 0002, Guoqi Li 0002 |
ICLR | 5 |
| 2024 | Spiking Transformer with Experts MixtureabstractSpiking Neural Networks (SNNs) provide a sparse spike-driven mechanism which is believed to be critical for energy-efficient deep learning.
Mixture-of-Experts (MoE), on the other side, aligns with the brain mechanism of distributed and sparse processing, resulting in an efficient way of enhancing model capacity and conditional computation.
In this work, we consider how to incorporate SNNs’ spike-driven and MoE’s conditional computation into a unified framework.
However, MoE uses softmax to get the dense conditional weights for each expert and TopK to hard-sparsify the network, which does not fit the properties of SNNs.
To address this issue, we reformulate MoE in SNNs and introduce the Spiking Experts Mixture Mechanism (SEMM) from the perspective of sparse spiking activation.
Both the experts and the router output spiking sequences, and their element-wise operation makes SEMM computation spike-driven and dynamic sparse-conditional.
By developing SEMM into Spiking Transformer, the Experts Mixture Spiking Attention (EMSA) and the Experts Mixture Spiking Perceptron (EMSP) are proposed, which performs routing allocation for head-wise and channel-wise spiking experts, respectively. Experiments show that SEMM realizes sparse conditional computation and obtains a stable improvement on neuromorphic and static datasets with approximate computational overhead based on the Spiking Transformer baselines. Zhaokun Zhou, Yijie Lu, Yanhao Jia, Kaiwei Che, Liwei Huang, Yuesheng Zhu, Guoqi Li 0002, Zhaofei Yu, Li Yuan 0007 |
NeurIPS | 1 |
| 2024 | QKFormer: Hierarchical Spiking Transformer using Q-K AttentionabstractSpiking Transformers, which integrate Spiking Neural Networks (SNNs) with Transformer architectures, have attracted significant attention due to their potential for low energy consumption and high performance. However, there remains a substantial gap in performance between SNNs and Artificial Neural Networks (ANNs). To narrow this gap, we have developed QKFormer, a direct training spiking transformer with the following features: i) _Linear complexity and high energy efficiency_, the novel spike-form Q-K attention module efficiently models the token or channel attention through binary vectors and enables the construction of larger models. ii) _Multi-scale spiking representation_, achieved by a hierarchical structure with the different numbers of tokens across blocks. iii) _Spiking Patch Embedding with Deformed Shortcut (SPEDS)_, enhances spiking information transmission and integration, thus improving overall performance. It is shown that QKFormer achieves significantly superior performance over existing state-of-the-art SNN models on various mainstream datasets. Notably, with comparable size to Spikformer (66.34 M, 74.81\%), QKFormer (64.96 M) achieves a groundbreaking top-1 accuracy of **85.65\%** on ImageNet-1k, substantially outperforming Spikformer by **10.84\%**. To our best knowledge, this is the first time that directly training SNNs have exceeded 85\% accuracy on ImageNet-1K. Chenlin Zhou, Han Zhang 0035, Zhaokun Zhou, Liutao Yu, Liwei Huang, Xiaopeng Fan 0001, Li Yuan 0007, Zhengyu Ma, Yonghong Tian 0001 |
NeurIPS | 3 |
| 2023 | Imbalanced Conditional Conv-Transformer for Mathematical Expression Recognition
Shuaijian Ji, Zhaokun Zhou, Yuqing Wang 0006, Baishan Duan, Zhenyu Weng, Yuesheng Zhu |
ICANN (6) | 2 |
| 2023 | Spikformer: When Spiking Neural Network Meets Transformer
Zhaokun Zhou, Yuesheng Zhu, Yaowei Wang 0001, Shuicheng Yan, Yonghong Tian 0001, Li Yuan 0007 |
ICLR | 1 |
| 2023 | TRMER: Transformer-Based End to End Printed Mathematical Expression RecognitionabstractAs a fundamental task of transcribing formula images into structural mathematical expressions, Printed Mathematical Expression Recognition (PMER) is wildly used in many fields. However, there is still a lack of an end-to-end approach toward fully exploring the spatial structure and semantic information in the formula to achieve high recognition accuracy. In this work, a Transformer-based Mathematical Expression Recognition (TRMER) model, is proposed to enhance the recognition accuracy. A Dual-Branch Encoder (DBE) is developed to extract multi-scaled feature maps from a formula image so that the spatial and semantic information can be obtained synchronously, and the different feature maps are fused with a Fusion Enhancement Module (FEM) by merging and reinforcing the spatial-semantic information. A standard transformer-based decoder is developed to decode the rich spatial-semantic information of the image and output a recognized mathematical expression in LaTex sequence. The experimental results have illustrated that the TRMER has achieved state-of-the-art recognition performance. Zhaokun Zhou, Shuaijian Ji, Yuqing Wang 0006, Zhenyu Weng, Yuesheng Zhu |
IJCNN | 1 |
| 2023 | Parallel Spiking Neurons with High Efficiency and Ability to Learn Long-term DependenciesabstractVanilla spiking neurons in Spiking Neural Networks (SNNs) use charge-fire-reset neuronal dynamics, which can only be simulated serially and can hardly learn long-time dependencies. We find that when removing reset, the neuronal dynamics can be reformulated in a non-iterative form and parallelized. By rewriting neuronal dynamics without reset to a general formulation, we propose the Parallel Spiking Neuron (PSN), which generates hidden states that are independent of their predecessors, resulting in parallelizable neuronal dynamics and extremely high simulation speed. The weights of inputs in the PSN are fully connected, which maximizes the utilization of temporal information. To avoid the use of future inputs for step-by-step inference, the weights of the PSN can be masked, resulting in the masked PSN. By sharing weights across time-steps based on the masked PSN, the sliding PSN is proposed to handle sequences of varying lengths. We evaluate the PSN family on simulation speed and temporal/static data classification, and the results show the overwhelming advantage of the PSN family in efficiency and accuracy. To the best of our knowledge, this is the first study about parallelizing spiking neurons and can be a cornerstone for the spiking deep learning research. Our codes are available at https://github.com/fangwei123456/Parallel-Spiking-Neuron. Wei Fang 0006, Zhaofei Yu, Zhaokun Zhou, Yanqi Chen, Zhengyu Ma, Timothée Masquelier, Yonghong Tian 0001 |
NeurIPS | 3 |
| 2023 | Spike-driven TransformerabstractSpiking Neural Networks (SNNs) provide an energy-efficient deep learning option due to their unique spike-based event-driven (i.e., spike-driven) paradigm. In this paper, we incorporate the spike-driven paradigm into Transformer by the proposed Spike-driven Transformer with four unique properties: (1) Event-driven, no calculation is triggered when the input of Transformer is zero; (2) Binary spike communication, all matrix multiplications associated with the spike matrix can be transformed into sparse additions; (3) Self-attention with linear complexity at both token and channel dimensions; (4) The operations between spike-form Query, Key, and Value are mask and addition. Together, there are only sparse addition operations in the Spike-driven Transformer. To this end, we design a novel Spike-Driven Self-Attention (SDSA), which exploits only mask and addition operations without any multiplication, and thus having up to $87.2\times$ lower computation energy than vanilla self-attention. Especially in SDSA, the matrix multiplication between Query, Key, and Value is designed as the mask operation. In addition, we rearrange all residual connections in the vanilla Transformer before the activation functions to ensure that all neurons transmit binary spike signals. It is shown that the Spike-driven Transformer can achieve 77.1\% top-1 accuracy on ImageNet-1K, which is the state-of-the-art result in the SNN field. Man Yao, Zhaokun Zhou, Li Yuan 0007, Yonghong Tian 0001, Bo Xu 0002, Guoqi Li 0002 |
NeurIPS | 3 |
| 2022 | Dual Branch Network Towards Accurate Printed Mathematical Expression Recognition
Yuqing Wang 0006, Zhenyu Weng, Zhaokun Zhou, Shuaijian Ji, Zhongjie Ye, Yuesheng Zhu |
ICANN (4) | 3 |
| 2022 | PRPN: Progressive region prediction network for natural scene text detection
Yuanhong Zhong, Tao Chen 0004, Jing Zhang 0037, Zhaokun Zhou |
Knowl. Based Syst. | 5 |
| 2021 | Recovery of image and video based on compressive sensing via tensor approximation and Spatio-temporal correlation
Yuanhong Zhong, Jing Zhang 0037, Zhaokun Zhou |
Multim. Tools Appl. | 3 |
| 2021 | Code Caching-Assisted Computation Offloading and Resource Allocation for Multi-User Mobile Edge ComputingabstractUtilizing the data caching technology to reduce data transmission is a promising technique for improving the performance of mobile edge computing (MEC), because the delay and energy consumption produced by data transmission constitute the dominant cost of task execution in MEC. Besides, computation tasks generally consist of input parameters, executive codes, and computation results. The executive codes are fixed and can output difference computation results under different input parameters. Motivated by this, we consider to proactively cache executive codes of tasks at the MEC server to reduce the weighted sum of task execution delay and users’ energy consumption. Aiming at establishing optimal system design, we formulate the problem as a non-linear programming problem which involves jointly optimizing the executive code caching strategy, computation offloading decision, wireless resource allocation, and computing resource allocation. We propose to find the optimal solution by employing an alternating optimization framework. The optimal wireless resource and computing resource allocation problem are firstly addressed by utilizing convex optimization technology. Then, a dynamic programming-based algorithm has been developed to achieve the optimal executive code caching and computation offloading strategies. Extensive simulation results show that the proposed scheme operates well and can substantially reduce the system cost over other benchmark schemes. Zhixiong Chen 0003, Zhaokun Zhou, Chen Chen 0037 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2020 | Dynamic Task Caching and Computation Offloading for Mobile Edge ComputingabstractMobile edge computing (MEC) provides information technology and cloud-computing capabilities within the network edge close to mobile users, thereby addressing computing demand for users. However, the energy consumption of data uploading from users to the MEC server makes it hard to meet users' demand in some specific applications, i.e., interactive gaming and augmented reality. Motivated by this, we integrate a task caching mechanism into computation offloading technique. Specifically, it allows the MEC server to proactively cache some tasks and users to offload their tasks to the MEC server. Since the limited storage capacity and task demands for users are changing dynamically, which tasks are cached has to be judiciously decided to maximize the MEC system performance. The objective of this paper is to minimize the system cost, which is defined as the average total user energy consumption of all time slots. By formulating the problem as an integer-programming, we propose to find the optimal solution with two steps. Through which we have obtained the optimal online computation offloading and task cache update strategy. Simulation results show that in comparison with the other two baselines, the proposed scheme can effectively reduce the system cost. Zhixiong Chen 0003, Zhaokun Zhou |
GLOBECOM | 2 |
| 2020 | Specific category region proposal network for text detection in natural sceneabstractNatural scene text usually carries considerable abstract semantic information, which is closely related to the surrounding environment. Thus, natural scene text detection plays a vital role in image content retrieval and understanding. In this study, the authors propose a novel specific category region proposal network (SCRPN) based on maximally stable extremal regions (MSER) and fully convolutional network (FCN) for natural scene text detection. First, FCN for pixel‐level recognition is utilised to obtain the text saliency map and MSER is used to obtain oversegmented regions. Then, the multiple features of oversegmented regions and text saliency map are used for region aggregation. Next, single‐linkage clustering method is adopted to cluster the segmentation regions to obtain a hierarchical structure of text region proposals. Finally, for the top‐ranking region proposals, SCRPN built an end‐to‐end pipeline for scene text detection directly. Experiments on street view text and international conference on document analysis and recognition (ICDAR) 2013 have demonstrated the effectiveness of SCRPN for generating the text proposals. SCRPN could work with various two‐stage text detection networks; thus, faster region convolutional neural network was used as the text detection framework to evaluate the performance of SCRPN in the ICDAR 2015 and MSRA‐TD500 benchmarks. The experimental results confirmed that SCRPN makes text detection more robust in complex scenarios. Yuanhong Zhong, Zhaokun Zhou, Jing Zhang 0037 |
IET Image Process. | 3 |