EDBT 2026 Demo / reviewers in the wild / expert
Li Lun
dblp:366/0014
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2026
0009-0000-1780-1646ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 7 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DyNeuro: A Hybrid Neuromorphic Accelerator with Dynamic Spatio-Temporal Variation Adaptations
Youming Yang 0002, Yi Zhong 0002, Li Lun, Tao Zhang 0140, Xiaoxin Cui, Yuan Wang 0001 |
ISCAS | 3 |
| 2025 | Towards Effective and Sparse Adversarial Attack on Spiking Neural Networks via Breaking Invisible Surrogate GradientsabstractSpiking neural networks (SNNs) have shown their competence in handling spatial-temporal event-based data with low energy consumption. Similar to conventional artificial neural networks (ANNs), SNNs are also vulnerable to gradient-based adversarial attacks, wherein gradients are calculated by spatial-temporal back-propagation (STBP) and surrogate gradients (SGs). However, the SGs may be invisible for an inference-only model as they do not influence the inference results, and current gradient-based attacks are ineffective for binary dynamic images captured by the dynamic vision sensor (DVS). While some approaches addressed the issue of invisible SGs through universal SGs, their SGs lack a correlation with the victim model, resulting in sub-optimal performance. Moreover, the imperceptibility of existing SNN-based binary attacks is still insufficient. In this paper, we introduce an innovative potential-dependent surrogate gradient (PDSG) method to establish a robust connection between the SG and the model, thereby enhancing the adaptability of adversarial attacks across various models with invisible SGs. Additionally, we propose the sparse dynamic attack (SDA) to effectively attack binary dynamic images. Utilizing a generation-reduction paradigm, SDA can fully optimize the sparsity of adversarial perturbations. Experimental results demonstrate that our PDSG and SDA outperform state-of-the-art SNN-based attacks across various models and datasets. Specifically, our PDSG achieves 100% attack success rate on ImageNet, and our SDA obtains 82% attack success rate by modifying only 0.24% of the pixels on CIFAR10DVS. The code is available at https://github.com/ryime/PDSG-SDA. Li Lun, Kunyu Feng, Qinglong Ni, Ling Liang 0003, Yuan Wang 0001, Ying Li 0056, Dunshan Yu, Xiaoxin Cui |
CVPR | 1 |
| 2025 | CROSSCUT: A Multi-Core Neuromorphic Accelerator Improving Resource-UtilizationabstractNeuromorphic computing is attracting significant attention due to its bio-mimetic characteristics. Consequently, neuromorphic hardware platforms have emerged as innovative computing architectures for acceleration. However, the fixed nature of data flow and resources leads to considerable inefficiencies in storage and computation, thereby limiting both utilization efficiency and overall performance. This severely hinders the deployment of edge artificial intelligence (AI) models. To address these issues, we present a multi-core neuromorphic accelerator named CROSSCUT. This crossbar-based system supports both spiking neural network (SNN) and artificial neural network (ANN) paradigms and has a capacity of 256K neurons and 288M synapses. By leveraging the Neuron Package Mechanism (NPM) and Synapse Compress Mechanism (SCM), CROSSCUT can increase input data scale by 64 times and reduce wasted resources and computations by 46.7%, ensuring high compatibility with diverse network structures in machine learning models. Additionally, a Tree-Mesh hybrid network on chip (NoC) is constructed for inter-core communication. Implemented on Xilinx XCVU9P FPGA, CROSSCUT can achieve a peak performance of 431.9 GSOPS/s and 121.13 GSOPS/W energy efficiency. The inference accuracy on MNIST is 98.2%. Youming Yang 0002, Yi Zhong 0002, Zilin Wang 0001, Tao Zhang 0140, Li Lun, Yingying Cui, Xiaoxin Cui, Song Jia, Yuan Wang 0001 |
ISCAS | 5 |
| 2025 | A Reconfigurable Digital Compute-In-Memory Heterogeneous Macro for Differential Frame Convolution and Spiking Neural NetworkabstractThe application of artificial neural network (ANN) in video processing encounters significant challenges, including large data volumes, numerous linear operations, and high power consumption. The fusion of convolutional neural network (CNN) and spiking neural network (SNN) provides a dual benefit of achieving high accuracy while maintaining low power consumption. However, ongoing challenges remain in minimizing multiply-accumulate (MAC) operations and optimizing data movement. In this work, we propose a reconfigurable digital compute-in-memory (RDCIM) heterogeneous macro without the sense amplifier, tailored for the diverse computational demands of CNN and SNN. To improve energy efficiency, differential frame convolution (DFC) is adopted to mitigate the computational overhead. In addition, computational resources are functionally reused to accommodate four data flow types, supporting both DFC and SNN operations. Implemented by TSMC 28nm technology, the proposed RDCIM heterogeneous macro achieves the peak energy efficiency of 29.13 TOPS/W for DFC and 0.56 pJ/SOP for SNN, operating at a frequency of 284 MHz. Li Lun, Zhenhui Dai, Yingying Cui, Xiaoxin Cui |
ISCAS | 2 |
| 2025 | HyNITA: A Neuromorphic Inference and Training Accelerator for Hybrid ANN-SNN Fusion ModelsabstractIn order to achieve the brain-like advantages over conservative computers, previous neuromorphic researchers have stretched the hardware explorations of the hybrid artificial neural network (ANN) and spiking neural network (SNN) inference approaches, as well as the efficient bio-plausible and gradient-based SNN training mechanisms. However, a versatile accelerator for both ANN-SNN inference and training is little addressed. In this work, we introduce HyNITA, a neuromorphic processor that supports accelerating both inference and training tasks of hybrid ANN and SNN models. Regarding the similarity and distinction, a pair of working stages are distinguished and distributed to multiple simple cores. The accelerator optimizes the interchange dataflow in a scalable chip design, following a reconfigurable design methodology to integrate the involved equation calculations in the dynamic process of neurons. The evaluation results show it achieves an accuracy of 99.65% and 99.34% on training ANN MNIST and SNN N-MNIST datasets. Yi Zhong 0002, Li Lun, Zilin Wang 0001, Jinhao Ruan, Yipeng Gao, Xiaoxin Cui, Xing Zhang 0002, Yuan Wang 0001 |
ISCAS | 2 |
| 2024 | SPAT: FPGA-based Sparsity-Optimized Spiking Neural Network Training Accelerator with Temporal Parallel DataflowabstractSpiking neural networks (SNNs), as biologically inspired computational models, possess significant advantages in energy efficiency due to their event-driven operations. However, challenges remain in attaining high computational efficiency for SNN training. In this work, we propose a novel SNN training accelerator employing temporal parallelism and sparsity optimizations to achieve superior efficiency. A temporal parallel dataflow is designed to concurrently integrate spikes across multiple time steps, enhancing throughput and data reuse. To reduce latency and improve energy efficiency, we leverage the sparsity of SNNs and employ methods such as zero gating and zero skipping. Implemented on a field-programmable gate array (FPGA), the proposed training accelerator demonstrates 2.3-fold speedup and 15.7-fold energy reduction compared to NVIDIA A100 GPU on N-MNIST dataset. Li Lun, Mingqi Yin, Zhenhui Dai, Xiaole Cui, Xiaoxin Cui |
ISCAS | 2 |
| 2024 | A 16.41 TOPS/W CNN Accelerator with Event-Based Layer Fusion for Real-Time InferenceabstractThis paper proposes a convolutional neural network (CNN) accelerator architecture for real-time tasks in edge devices. An event-based layer fusion technique is adopted to eliminate on-chip storage requirements and off-chip data movement caused by features. Cross-layer pipeline is elaborated during layer fusion to obtain high throughput and low latency. An adaptive fully unrolling event-driven core is designed and a cyclic storage method is exploited to reduce the storage space for partial sum in the core. Modified LeNet is accelerated with the proposed architecture. The accelerator can reach an energy efficiency of 16.41 TOPS/W and a latency of 0.85μs under TSMC 28nm technology, and a frame rate of 369.4K FPS under FPGA. Li Lun, Zhenhui Dai, Xiaoxin Cui |
ISCAS | 2 |
| 2024 | NeuroREC: A 28-nm Efficient Neuromorphic Processor for Radar Emitter ClassificationabstractRadar emitter classification (REC) plays an important role in modern warfare. Traditional REC methods have difficulty identifying complex radar signals in the present day. Inspired by biology, spiking neural networks (SNNs) have gradually gained widespread attention due to their low power characteristics. Compared with convolutional neural networks (CNNs), SNNs are more suitable for application in the field of REC. The reason is that SNN can not only maintain higher accuracy in the presence of noise interference, but also reduce the power consumption of mobile devices. However, it is challenging to make full use of the input sparsity of radar emitter signals and the weight sparsity of pruned SNN models. In this paper, a 28-nm neuromorphic processor for REC named NeuroREC is proposed. It uses matrix compression algorithms to store sparse weights on chip, and designs corresponding spike detection circuits for this purpose. As a single-core design, we propose a ping-pong running mechanism to alleviate the imbalance between IO throughput and peak performance. Two SNN models for classifying RadioML2016.b and RadioML2018.a datasets are deployed on the chip, achieving competitive accuracy with only 8 timesteps, and demonstrating better robustness than CNN. Fabricated in 28-nm CMOS process, NeuroREC runs at frequencies ranging from 22.5MHz to 744MHz. Under specific sparsity conditions, it can reach an energy efficiency of 7.22TSOP/W for 8-bit weight. Zilin Wang 0001, Zehong Ou, Yi Zhong 0002, Youming Yang 0002, Li Lun, Hufei Li, Jian Cao 0002, Xiaoxin Cui, Song Jia, Yuan Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |