Kairong Yu

dblp:326/8953 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Deep learning architectures and training · 70% Efficient and distributed learning · 30%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Emerging computing paradigms · 100%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
spiking neural network
2.632025
Enhanced Self-Distillation Framework for Efficient Spiking Neural Network Training · NeurIPS 2025
TS-SNN: Temporal Shift Module for Spiking Neural Networks · ICML 2025
STAA-SNN: Spatial-Temporal Attention Aggregator for Spiking Neural Networks · CVPR 2025
Emerging computing paradigms
neuromorphic computing
2.632025
TS-SNN: Temporal Shift Module for Spiking Neural Networks · ICML 2025
Temporal Separation with Entropy Regularization for Knowledge Distillation in Spiking Neural Networks · CVPR 2025
FSTA-SNN: Frequency-Based Spatial-Temporal Attention Module for Spiking Neural Networks · AAAI 2025
Emerging computing paradigms › neuromorphic computing
spiking neural network
2.632025
TS-SNN: Temporal Shift Module for Spiking Neural Networks · ICML 2025
Temporal Separation with Entropy Regularization for Knowledge Distillation in Spiking Neural Networks · CVPR 2025
FSTA-SNN: Frequency-Based Spatial-Temporal Attention Module for Spiking Neural Networks · AAAI 2025
Machine learning › Deep learning architectures and training
attention mechanism
1.722025
STAA-SNN: Spatial-Temporal Attention Aggregator for Spiking Neural Networks · CVPR 2025
FSTA-SNN: Frequency-Based Spatial-Temporal Attention Module for Spiking Neural Networks · AAAI 2025
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
1.722025
Enhanced Self-Distillation Framework for Efficient Spiking Neural Network Training · NeurIPS 2025
Temporal Separation with Entropy Regularization for Knowledge Distillation in Spiking Neural Networks · CVPR 2025
Machine learning › Deep learning architectures and training › efficient deep learning
efficient neural network architecture
0.912025
TS-SNN: Temporal Shift Module for Spiking Neural Networks · ICML 2025
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
model distillation
0.912025
Temporal Separation with Entropy Regularization for Knowledge Distillation in Spiking Neural Networks · CVPR 2025
Machine learning › Deep learning architectures and training › attention mechanism
self-attention
0.912025
STAA-SNN: Spatial-Temporal Attention Aggregator for Spiking Neural Networks · CVPR 2025
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
self-distillation
0.912025
Enhanced Self-Distillation Framework for Efficient Spiking Neural Network Training · NeurIPS 2025
Machine learning › Deep learning architectures and training › attention mechanism › multi-dimensional attention
spatio-temporal attention
0.912025
STAA-SNN: Spatial-Temporal Attention Aggregator for Spiking Neural Networks · CVPR 2025
Machine learning › Deep learning architectures and training › attention mechanism › self-attention
spiking self-attention
0.912025
STAA-SNN: Spatial-Temporal Attention Aggregator for Spiking Neural Networks · CVPR 2025
Machine learning › Deep learning architectures and training › spiking neural network
surrogate gradient
0.912025
Enhanced Self-Distillation Framework for Efficient Spiking Neural Network Training · NeurIPS 2025
Emerging computing paradigms
knowledge distillation
0.912025
Temporal Separation with Entropy Regularization for Knowledge Distillation in Spiking Neural Networks · CVPR 2025
Machine learning › Efficient and distributed learning
model compression
0.312025
Temporal Separation with Entropy Regularization for Knowledge Distillation in Spiking Neural Networks · CVPR 2025
Machine learning › Deep learning architectures and training
positional encoding
0.312025
STAA-SNN: Spatial-Temporal Attention Aggregator for Spiking Neural Networks · CVPR 2025
Emerging computing paradigms
neuromorphic hardware
0.312025
Enhanced Self-Distillation Framework for Efficient Spiking Neural Network Training · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

temporal shift · 1.7temporal separation · 1.7spatial-temporal attention · 1.7residual combination · 1.7frequency analysis · 1.7entropy regularization · 1.7time-step random dropout · 0.9step attention · 0.9spike-driven self-attention · 0.9rate-based backpropagation · 0.9positional encoding · 0.9knowledge distillation · 0.9
YearPublicationVenuePosition
2025 FSTA-SNN: Frequency-Based Spatial-Temporal Attention Module for Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) are emerging as a promising alternative to Artificial Neural Networks (ANNs) due to their inherent energy efficiency. Owing to the inherent sparsity in spike generation within SNNs, the in-depth analysis and optimization of intermediate output spikes are often neglected. This oversight significantly restricts the inherent energy efficiency of SNNs and diminishes their advantages in spatiotemporal feature extraction, resulting in a lack of accuracy and unnecessary energy expenditure. In this work, we analyze the inherent spiking characteristics of SNNs from both temporal and spatial perspectives. In terms of spatial analysis, we find that shallow layers tend to focus on learning vertical variations, while deeper layers gradually learn horizontal variations of features. Regarding temporal analysis, we observe that there is not a significant difference in feature learning across different time steps. This suggests that increasing the time steps has limited effect on feature learning. Based on the insights derived from these analyses, we propose a Frequency-based Spatial-Temporal Attention (FSTA) module to enhance feature learning in SNNs. This module aims to improve the feature learning capabilities by suppressing redundant spike features. The experimental results indicate that the introduction of the FSTA module significantly reduces the spike firing rate of SNNs, demonstrating superior performance compared to state-of-the-art baselines across multiple datasets.
Kairong Yu, Tianqing Zhang, Hongwei Wang 0001, Qi Xu 0008
AAAI1
2025 Temporal Separation with Entropy Regularization for Knowledge Distillation in Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs), inspired by the human brain, offer significant computational efficiency through discrete spike-based information transfer. Despite their potential to reduce inference energy consumption, a performance gap persists between SNNs and Artificial Neural Networks (ANNs), primarily due to current training methods and inherent model limitations. While recent research has aimed to enhance SNN learning by employing knowledge distillation (KD) from ANN teacher networks, traditional distillation techniques often overlook the distinctive spatiotemporal properties of SNNs, thus failing to fully leverage their advantages. To overcome these challenge, we propose a novel logit distillation method characterized by temporal separation and entropy regularization. This approach improves existing SNN distillation techniques by performing distillation learning on logits across different time steps, rather than merely on aggregated output features. Furthermore, the integration of entropy regularization stabilizes model optimization and further boosts the performance. Extensive experimental results indicate that our method surpasses prior SNN distillation strategies, whether based on logit distillation, feature distillation, or a combination of both. Our project is available at https://github.com/yukairong/TSER.
Kairong Yu, Chengting Yu, Tianqing Zhang, Xiaochen Zhao, Hongwei Wang 0001, Qiang Zhang 0008, Qi Xu 0008
CVPR1
2025 STAA-SNN: Spatial-Temporal Attention Aggregator for Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) have gained significant attention due to their biological plausibility and energy efficiency, making them promising alternatives to Artificial Neural Networks (ANNs). However, the performance gap between SNNs and ANNs remains a substantial challenge hindering the widespread adoption of SNNs. In this paper, we propose a Spatial-Temporal Attention Aggregator SNN (STAA-SNN) framework, which dynamically focuses on and captures both spatial and temporal dependencies. First, we introduce a spike-driven self-attention mechanism specifically designed for SNNs. Additionally, we pioneeringly incorporate position encoding to integrate latent temporal relationships into the incoming features. For spatial-temporal information aggregation, we employ step attention to selectively amplify relevant features to variant steps. Finally, we implement a time-step random dropout strategy to avoid local optima. The framework demonstrates exceptional performance across diverse datasets and exhibits strong generalization capabilities. Notably, STAA-SNN achieves state-of-the-art results on neuromorphic datasets CIFAR10-DVS of 82.10% and with performances of 97.14%, 82.05% and 70.40% on the static datasets CIFAR-10, CIFAR-100 and ImageNet, respectively. Furthermore, this model exhibits improved performance ranging from 0.33% to 2.80% with fewer time steps.
Tianqing Zhang, Kairong Yu, Xian Zhong, Hongwei Wang 0001, Qi Xu 0008, Qiang Zhang 0008
CVPR2
2025 DA-LIF: Dual Adaptive Leaky Integrate-and-Fire Model for Deep Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) are valued for their ability to process spatio-temporal information efficiently, offering biological plausibility, low energy consumption, and compatibility with neuromorphic hardware. However, the commonly used Leaky Integrate-and-Fire (LIF) model overlooks neuron heterogeneity and independently processes spatial and temporal information, limiting the expressive power of SNNs. In this paper, we propose the Dual Adaptive Leaky Integrate- and-Fire (DA-LIF) model, which introduces spatial and temporal tuning with independently learnable decays. Evaluations on both static (CIFAR10/100, ImageNet) and neuromorphic datasets (CIFAR10-DVS, DVS128 Gesture) demonstrate superior accuracy with fewer timesteps compared to state-of-the-art methods. Importantly, DA-LIF achieves these improvements with minimal additional parameters, maintaining low energy consumption. Extensive ablation studies further highlight the robustness and effectiveness of the DA-LIF model.
Tianqing Zhang, Kairong Yu, Jian Zhang 0083, Hongwei Wang 0001
ICASSP2
2025 TS-SNN: Temporal Shift Module for Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) are increasingly recognized for their biological plausibility and energy efficiency, positioning them as strong alternatives to Artificial Neural Networks (ANNs) in neuromorphic computing applications. SNNs inherently process temporal information by leveraging the precise timing of spikes, but balancing temporal feature utilization with low energy consumption remains a challenge. In this work, we introduce Temporal Shift module for Spiking Neural Networks (TS-SNN), which incorporates a novel Temporal Shift (TS) module to integrate past, present, and future spike features within a single timestep via a simple yet effective shift operation. A residual combination method prevents information loss by integrating shifted and original features. The TS module is lightweight, requiring only one additional learnable parameter, and can be seamlessly integrated into existing architectures with minimal additional computational cost. TS-SNN achieves state-of-the-art performance on benchmarks like CIFAR-10 (96.72%), CIFAR-100 (80.28%), and ImageNet (70.61%) with fewer timesteps, while maintaining low energy consumption. This work marks a significant step forward in developing efficient and accurate SNN architectures.
Kairong Yu, Tianqing Zhang, Qi Xu 0008, Gang Pan 0001, Hongwei Wang 0001
ICML1
2025 Head-Tail-Aware KL Divergence in Knowledge Distillation for Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) have emerged as a promising approach for energy-efficient and biologically plausible computation. However, due to limitations in existing training methods and inherent model constraints, SNNs often exhibit a performance gap when compared to Artificial Neural Networks (ANNs). Knowledge distillation (KD) has been explored as a technique to transfer knowledge from ANN teacher models to SNN student models to mitigate this gap. Traditional KD methods typically use Kullback-Leibler (KL) divergence to align output distributions. However, conventional KL-based approaches fail to fully exploit the unique characteristics of SNNs, as they tend to overemphasize high-probability predictions while neglecting low-probability ones, leading to suboptimal generalization. To address this, we propose Head-Tail Aware Kullback-Leibler (HTA-KL) divergence, a novel KD method for SNNs. HTA-KL introduces a cumulative probability-based mask to dynamically distinguish between high- and low-probability regions. It assigns adaptive weights to ensure balanced knowledge transfer, enhancing the overall performance. By integrating forward KL (FKL) and reverse KL (RKL) divergence, our method effectively align both head and tail regions of the distribution. We evaluate our methods on CIFAR-10, CIFAR-100 and Tiny ImageNet datasets. Our method outperforms existing methods on most datasets with fewer timesteps.
Tianqing Zhang, Zixin Zhu, Kairong Yu, Hongwei Wang 0001
IJCNN3
2025 Enhanced Self-Distillation Framework for Efficient Spiking Neural Network Training
abstract
Spiking Neural Networks (SNNs) exhibit exceptional energy efficiency on neuromorphic hardware due to their sparse activation patterns. However, conventional training methods based on surrogate gradients and Backpropagation Through Time (BPTT) not only lag behind Artificial Neural Networks (ANNs) in performance, but also incur significant computational and memory overheads that grow linearly with the temporal dimension. To enable high-performance SNN training under limited computational resources, we propose an enhanced self-distillation framework, jointly optimized with rate-based backpropagation. Specifically, the firing rates of intermediate SNN layers are projected onto lightweight ANN branches, and high-quality knowledge generated by the model itself is used to optimize substructures through the ANN pathways. Unlike traditional self-distillation paradigms, we observe that low-quality self-generated knowledge may hinder convergence. To address this, we decouple the teacher signal into reliable and unreliable components, ensuring that only reliable knowledge is used to guide the optimization of the model. Extensive experiments on CIFAR-10, CIFAR-100, CIFAR10-DVS, and ImageNet demonstrate that our method reduces training complexity while achieving high-performance SNN training. Our code is available at https://github.com/Intelli-Chip-Lab/enhanced-self-distillation-framework-for-snn.
Xiaochen Zhao, Chengting Yu, Kairong Yu, Aili Wang 0002
NeurIPS3