Shibo Zhou

dblp:227/8778 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Deep learning architectures and training · 26% Reinforcement learning · 15% Robot navigation and mapping · 12%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Emerging computing paradigms · 100%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 21 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
spiking neural network
3.142026
Dynamic Token Masking in Spiking Neural Network · Int. J. Comput. Vis. 2026
Efficient 3D Recognition with Event-driven Spike Sparse Convolution · AAAI 2025
Spiking Neural Network as Adaptive Event Stream Slicer · NeurIPS 2024
Emerging computing paradigms
neuromorphic computing
2.232025
Spiking Neural Networks Need High-Frequency Information · NeurIPS 2025
Efficient 3D Recognition with Event-driven Spike Sparse Convolution · AAAI 2025
Temporal-Coded Deep Spiking Neural Network with Easy Training and Robust Performance · AAAI 2021
Machine learning › Efficient and distributed learning
model compression
1.012026
Dynamic Token Masking in Spiking Neural Network · Int. J. Comput. Vis. 2026
Natural language and speech › Language models and text generation
token masking
1.012026
Dynamic Token Masking in Spiking Neural Network · Int. J. Comput. Vis. 2026
Machine learning › Reinforcement learning
deep reinforcement learning
0.912025
HSRL: A Hierarchical Control System Based on Spiking Deep Reinforcement Learning for Robot Navigation · ICRA 2025
Robotics › Robot navigation and mapping
mobile robot navigation
0.912025
HSRL: A Hierarchical Control System Based on Spiking Deep Reinforcement Learning for Robot Navigation · ICRA 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
Table2LaTeX-RL: High-Fidelity LaTeX Code Generation from Table Images via Reinforced Multimodal Language Models · NeurIPS 2025
Computer vision › 3D vision › 3d object recognition
point cloud recognition
0.912025
Efficient 3D Recognition with Event-driven Spike Sparse Convolution · AAAI 2025
Robotics › Robot navigation and mapping › learning-based navigation
reinforcement-learning-based navigation
0.912025
HSRL: A Hierarchical Control System Based on Spiking Deep Reinforcement Learning for Robot Navigation · ICRA 2025
Machine learning › Reinforcement learning › deep reinforcement learning
spiking reinforcement learning
0.912025
HSRL: A Hierarchical Control System Based on Spiking Deep Reinforcement Learning for Robot Navigation · ICRA 2025
Machine learning › Deep learning architectures and training › spiking neural network
spiking transformer
0.912025
Spiking Neural Networks Need High-Frequency Information · NeurIPS 2025
Natural language and speech › Information extraction and text analysis › document understanding › table recognition
table structure recognition
0.912025
Table2LaTeX-RL: High-Fidelity LaTeX Code Generation from Table Images via Reinforced Multimodal Language Models · NeurIPS 2025
Program synthesis and code generation › code generation with language models
image-to-code generation
0.912025
Table2LaTeX-RL: High-Fidelity LaTeX Code Generation from Table Images via Reinforced Multimodal Language Models · NeurIPS 2025
Emerging computing paradigms › neuromorphic computing
spiking neural network
0.912025
Spiking Neural Networks Need High-Frequency Information · NeurIPS 2025
Computer vision › Video understanding and tracking › object tracking
event-based tracking
0.812024
Spiking Neural Network as Adaptive Event Stream Slicer · NeurIPS 2024
Computer vision › Video understanding and tracking
event recognition
0.812024
Spiking Neural Network as Adaptive Event Stream Slicer · NeurIPS 2024
Machine learning › Reinforcement learning › policy optimization
group relative policy optimization
0.312025
Table2LaTeX-RL: High-Fidelity LaTeX Code Generation from Table Images via Reinforced Multimodal Language Models · NeurIPS 2025
Robotics › Motion planning and robot control › robot control
hierarchical control
0.312025
HSRL: A Hierarchical Control System Based on Spiking Deep Reinforcement Learning for Robot Navigation · ICRA 2025
Computer vision › Image recognition and object detection
image classification
0.312025
Spiking Neural Networks Need High-Frequency Information · NeurIPS 2025
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.312025
Table2LaTeX-RL: High-Fidelity LaTeX Code Generation from Table Images via Reinforced Multimodal Language Models · NeurIPS 2025
Robotics › Motion planning and robot control
robot control
0.312025
HSRL: A Hierarchical Control System Based on Spiking Deep Reinforcement Learning for Robot Navigation · ICRA 2025

Methods — techniques the papers use, named apart from their topics

spike voxel coding · 1.7spike sparse convolution · 1.7event-driven learning · 1.7dual-reward strategy · 1.7depthwise convolution · 1.7dynamic token masking · 1.0spiking neural network · 0.9reinforcement learning · 0.9max-pooling · 0.9max pooling · 0.9continuous attractor neural network · 0.9GRU · 0.9GRPO · 0.9weight quantization · 0.5temporal coding · 0.5
YearPublicationVenuePosition
2026 Dynamic Token Masking in Spiking Neural Network
Yuetong Fang, Deming Zhou, Shibo Zhou, Renjing Xu
Int. J. Comput. Vis.5
2025 Efficient 3D Recognition with Event-driven Spike Sparse Convolution
abstract
Spiking Neural Networks (SNNs) provide an energy-efficient way to extract 3D spatio-temporal features. Point clouds are sparse 3D spatial data, which suggests that SNNs should be well-suited for processing them. However, when applying SNNs to point clouds, they often exhibit limited performance and fewer application scenarios. We attribute this to inappropriate preprocessing and feature extraction methods. To address this issue, we first introduce the Spike Voxel Coding (SVC) scheme, which encodes the 3D point clouds into a sparse spike train space, reducing the storage requirements and saving time on point cloud preprocessing. Then, we propose a Spike Sparse Convolution (SSC) model for efficiently extracting 3D sparse point cloud features. Combining SVC and SSC, we design an efficient 3D SNN backbone (E-3DSNN), which is friendly with neuromorphic hardware. For instance, SSC can be implemented on neuromorphic chips with only minor modifications to the addressing function of vanilla spike convolution. Experiments on ModelNet40, KITTI, and Semantic KITTI datasets demonstrate that E-3DSNN achieves state-of-the-art (SOTA) results with remarkable efficiency. Notably, our E-3DSNN (1.87M) obtained 91.7% top-1 accuracy on ModelNet40, surpassing the current best SNN baselines (14.3M) by 3.0%. To our best knowledge, it is the first direct training 3D SNN backbone that can simultaneously handle various 3D computer vision tasks (e.g., classification, detection, and segmentation) with an event-driven nature.
Xuerui Qiu, Man Yao, Jieyuan Zhang, Yuhong Chou, Shibo Zhou, Bo Xu 0002, Guoqi Li 0002
AAAI6
2025 HSRL: A Hierarchical Control System Based on Spiking Deep Reinforcement Learning for Robot Navigation
abstract
Reinforcement Learning (RL) has shown promise in robotic navigation tasks, yet applying it to real-world environments remains challenging due to dynamic complexities and the need for dynamically feasible actions. We propose a hierarchical control framework based on Spiking Deep Reinforcement Learning (SDRL) for robust robot navigation in real environments. Our approach utilizes a two-layer architecture: a high-level decision layer powered by a Spiking GRU network for handling partially observable environments, and a low-level executive layer employing Continuous Attractor Neural Networks (CANNs) to ensure precise and continuous actions. This hierarchical structure allows real-time decisionmaking that respects the physical constraints of the robot. Experimental results show that our method adapts effectively to new environments without fine-tuning and surpasses existing methods in performance. We also explore the implementation on the Darwin3 chip, paving the way for biologically inspired motion control in future robotic applications.
Shibo Zhou, Chaohui Lin, Qingao Chai, Rui Yan 0005, De Ma, Gang Pan 0001, Huajin Tang
ICRA2
2025 Spiking Neural Networks Need High-Frequency Information
abstract
Spiking Neural Networks promise brain-inspired and energy-efficient computation by transmitting information through binary (0/1) spikes. Yet, their performance still lags behind that of artificial neural networks, often assumed to result from information loss caused by sparse and binary activations. In this work, we challenge this long-standing assumption and reveal a previously overlooked frequency bias: **spiking neurons inherently suppress high-frequency components and preferentially propagate low-frequency information.** This frequency-domain imbalance, we argue, is the root cause of degraded feature representation in SNNs. Empirically, on Spiking Transformers, adopting Avg-Pooling (low-pass) for token mixing lowers performance to 76.73% on Cifar-100, whereas replacing it with Max-Pool (high-pass) pushes the top-1 accuracy to 79.12%. Accordingly, we introduce **Max-Former** that restores high-frequency signals through two frequency-enhancing operators: (1) extra Max-Pool in patch embedding, and (2) Depth-Wise Convolution in place of self-attention. Notably, **Max-Former** attains 82.39% top-1 accuracy on ImageNet using only 63.99M parameters, surpassing Spikformer (74.81%, 66.34M) by +7.58%. Extending our insight beyond transformers, our **Max-ResNet-18** achieves state-of-the-art performance on convolution-based benchmarks: 97.17% on CIFAR-10 and 83.06% on CIFAR-100. We hope this simple yet effective solution inspires future research to explore the distinctive nature of spiking neural networks. Code is available: https://github.com/bic-L/MaxFormer.
Yuetong Fang, Deming Zhou, ZeCui Zeng, Lusong Li, Shibo Zhou, Renjing Xu
NeurIPS7
2025 Table2LaTeX-RL: High-Fidelity LaTeX Code Generation from Table Images via Reinforced Multimodal Language Models
abstract
In this work, we address the task of table image to LaTeX code generation, with the goal of automating the reconstruction of high-quality, publication-ready tables from visual inputs. A central challenge of this task lies in accurately handling complex tables—those with large sizes, deeply nested structures, and semantically rich or irregular cell content—where existing methods often fail. We begin with a comprehensive analysis, identifying key challenges and highlighting the limitations of current evaluation protocols. To overcome these issues, we propose a reinforced multimodal large language model (MLLM) framework, where a pre-trained MLLM is fine-tuned on a large-scale table-to-LaTeX dataset. To further improve generation quality, we introduce a dual-reward reinforcement learning strategy based on Group Relative Policy Optimization (GRPO). Unlike standard approaches that optimize purely over text outputs, our method incorporates both a structure-level reward on LaTeX code and a visual fidelity reward computed from rendered outputs, enabling direct optimization of the visual output quality. We adopt a hybrid evaluation protocol combining TEDS-Structure and CW-SSIM, and show that our method achieves state-of-the-art performance, particularly on structurally complex tables, demonstrating the effectiveness and robustness of our approach.
Jun Ling, Yao Qi, Shibo Zhou, Yanqin Huang, Yang Yang 0002, Heng Tao Shen, Peng Wang 0023
NeurIPS4
2024 Spiking Neural Network as Adaptive Event Stream Slicer
abstract
Event-based cameras are attracting significant interest as they provide rich edge information, high dynamic range, and high temporal resolution. Many state-of-the-art event-based algorithms rely on splitting the events into fixed groups, resulting in the omission of crucial temporal information, particularly when dealing with diverse motion scenarios (e.g., high/low speed). In this work, we propose SpikeSlicer, a novel-designed event processing framework capable of splitting events stream adaptively. SpikeSlicer utilizes a low-energy spiking neural network (SNN) to trigger event slicing. To guide the SNN to fire spikes at optimal time steps, we propose the Spiking Position-aware Loss (SPA-Loss) to modulate the neuron's state. Additionally, we develop a Feedback-Update training strategy that refines the slicing decisions using feedback from the downstream artificial neural network (ANN). Extensive experiments demonstrate that our method yields significant performance improvements in event-based object tracking and recognition. Notably, SpikeSlicer provides a brand-new SNN-ANN cooperation paradigm, where the SNN acts as an efficient, low-energy data processor to assist the ANN in improving downstream performance, injecting new perspectives and potential avenues of exploration.
Jiahang Cao, Hao Cheng 0015, Qiang Zhang 0029, Shibo Zhou, Renjing Xu
NeurIPS6
2024 Enhancing SNN-based spatio-temporal learning: A benchmark dataset and Cross-Modality Attention model
Shibo Zhou, Mengwen Yuan, Runhao Jiang, Rui Yan 0005, Gang Pan 0001, Huajin Tang
Neural Networks1
2021 Temporal-Coded Deep Spiking Neural Network with Easy Training and Robust Performance
abstract
Spiking neural network (SNN) is promising but the development has fallen far behind conventional deep neural networks (DNNs) because of difficult training. To resolve the training problem, we analyze the closed-form input-output response of spiking neurons and use the response expression to build abstract SNN models for training. This avoids calculating membrane potential during training and makes the direct training of SNN as efficient as DNN. We show that the nonleaky integrate-and-fire neuron with single-spike temporal-coding is the best choice for direct-train deep SNNs. We develop an energy-efficient phase-domain signal processing circuit for the neuron and propose a direct-train deep SNN framework. Thanks to easy training, we train deep SNNs under weight quantizations to study their robustness over low-cost neuromorphic hardware. Experiments show that our direct-train deep SNNs have the highest CIFAR-10 classification accuracy among SNNs, achieve ImageNet classification accuracy within 1% of the DNN of equivalent architecture, and are robust to weight quantization and noise perturbation.
Shibo Zhou, Xiaohua Li 0003, Sanjeev Tannirkulam Chandrasekaran, Arindam Sanyal
AAAI1
2020 Temporal Pulses Driven Spiking Neural Network for Time and Power Efficient Object Recognition in Autonomous Driving
abstract
Accurate real-time object recognition from sensory data has long been a crucial and challenging task for autonomous driving. Even though deep neural networks (DNNs) have been widely applied in this area, their considerable processing latency, power consumption, as well as computational complexity have been challenging issues for real-time autonomous driving applications. In this paper, we propose an approach to address the real-time object recognition problem utilizing spiking neural networks (SNNs). The proposed SNN model works directly with raw LiDAR temporal pulses without the pulse-to-point cloud preprocessing procedure, which can significantly reduce delay and power consumption. Being evaluated on various datasets derived from LiDAR and dynamic vision sensor (DVS), including Sim LiDAR, KITTI, and DVS-barrel, our proposed model has shown remarkable time and power efficiency, while achieving comparable recognition performance as the state-of-the-art methods. This paper highlights the SNN's great potentials in autonomous driving and related applications. To the best of our knowledge, this is the first attempt to use SNN to perform time and energy efficient object recognition directly on LiDAR temporal pulses in the setting of autonomous driving.
Wei Wang 0196, Shibo Zhou, Jingxi Li, Xiaohua Li 0003, Junsong Yuan 0001, Zhanpeng Jin
ICPR2
2020 Spiking Neural Networks with Single-Spike Temporal-Coded Neurons for Network Intrusion Detection
abstract
Spiking neural network (SNN) is interesting due to its strong bio-plausibility and high energy efficiency. However, its performance is falling far behind conventional deep neural networks (DNNs). In this paper, considering a general class of single-spike temporal-coded integrate-and-fire neurons, we analyze the input-output expressions of both leaky and nonleaky neurons. We show that SNNs built with leaky neurons suffer from the overly-nonlinear and overly-complex input-output response, which is the major reason for their difficult training and low performance. This reason is more fundamental than the commonly believed problem of nondifferentiable spikes. To support this claim, we show that SNNs built with nonleaky neurons can have a less-complex and less-nonlinear input-output response. They can be easily trained and can have superior performance, which is demonstrated by experimenting with the SNNs over two popular network intrusion detection datasets, i.e., the NSL-KDD and the AWID datasets. Our experiment results show that the proposed SNNs outperform a comprehensive list of DNN models and classic machine learning models. This paper demonstrates that SNNs can be promising and competitive in contrast to common beliefs.
Shibo Zhou, Xiaohua Li 0003
ICPR1