EDBT 2026 Demo / reviewers in the wild / expert
Wei Liu 0118
dblp:49/3283-118
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0001-8906-6636ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spiking-NeRF: Neural Graphics Acceleration With Spiking Feature Encoding for Edge 3D RenderingabstractNeural Radiance Fields (NeRF) have demonstrated remarkable potential for high-fidelity 3D scene reconstruction and rendering. However, achieving real-time performance remains a major challenge due to two critical bottlenecks: the high memory demand of multi-resolution hash encoding and the considerable computational cost of floating-point interpolation. To address these limitations, we propose Spiking-NeRF, a braininspired algorithm-hardware co-design framework. On the algorithm side, we introduce a spiking feature encoding scheme based on Integrate-and-Fire (IF) neurons, which transforms continuous voxel features into sparse spikes, reducing hash storage overhead by $\mathbf{7 5 \%}$. We further propose a global importance-based pruning strategy that compresses hash tables by $\mathbf{7 1. 3 \%}$ by removing lowaccessed entries. To reduce interpolation complexity, we design a hard-threshold weight discretization method that eliminates floating-point operations in favor of bitwise logic. On the hardware side, we accelerate critical stages of the NeRF pipeline by integrating a spike-skipping mechanism that dynamically bypasses hash entries, reducing memory traffic by 32.46%. We also co-optimize on-chip storage by leveraging access locality patterns across different resolution levels of the hash structure. Experimental results demonstrate that Spiking-NeRF achieves real-time rendering performance while maintaining high visual fidelity. Compared to edge GPUs, our design improves throughput by $111.2 \times$ and reduces power consumption by $41.67 \times$. Against SOTA NeRF accelerators, Spiking-NeRF achieves up to $2.48 \times$ higher throughput and $5.45 \times$ lower energy usage, underscoring the potential of spike-based computing for next-generation lowpower neural graphics systems. Jianzhen Gao, Wei Liu 0118, Hengyi Zhou, Zhiyi Yu, Shanlin Xiao |
ASP-DAC | 2 |
| 2026 | An Algorithm-Hardware Co-Design for Efficient and Robust Spiking Neural Networks via SparsityabstractSNN deployment faces a dilemma: rate codes are power-hungry while temporal codes lack noise resilience. This paper proposes an SNN algorithm-hardware co-design, which uses sparse coding and a zero-skipping accelerator to alleviate the rate-temporal trade-off. The design reduces network spike count by 88% compared to rate coding while enhancing fault tolerance. Benchmarked against state-of-the-art rate-coding and temporal-coding accelerators, the prototype saves 88% and 89% energy, achieves $4.5 \times$ and $26.8 \times$ higher throughput, and uses 82% fewer LUTs, enabling efficient and robust edge inference. Wei Liu 0118, Yinsheng Chen, Jilong Luo, Yusa Wang, Zhiyi Yu, Shanlin Xiao |
ASP-DAC | 1 |
| 2026 | SpikeVPR: Energy-efficient visual place recognition via multi-scale spiking transformers
Hengyi Zhou, Yinsheng Chen, Jilong Luo, Jianzheng Gao, Wei Liu 0118, Shanlin Xiao |
Neurocomputing | 5 |
| 2025 | SCSC: Leveraging Sparsity and Fault-Tolerance for Energy-Efficient Spiking Neural NetworksabstractSpiking neural networks (SNNs) are more energy-efficient for processing sparse spike signals and demonstrate better fault tolerance compared to artificial neural networks (ANNs). In neuromorphic chips, synaptic weight access and neuron computation operations constitute 75%-95% of the chip's energy consumption. Therefore, our primary strategy to achieve highly energy-efficient SNNs is to enhance network sparsity while leveraging SNNs' high fault tolerance to reduce both weight access and neuron computation energy. The coding module is an essential component of SNNs, responsible for encoding non-spiking inputs into spike trains. However, previous coding schemes often exhibit poor sparsity or fault tolerance performance. Thus, we propose a novel coding scheme for SNNs: spiking convolutional sparse coding (SCSC). SCSC utilizes convolutional kernels as dictionaries and achieves sparsity through neural layers. Additionally, dynamic firing thresholds in neural layers balance sparsity with network performance and fault tolerance. The experimental results indicate that SCSC can increase network sparsity by 10%-20% and achieves higher accuracy than baseline networks when dealing with disturbances. Furthermore, we utilize approximate DRAM to store synaptic weights and selectively deactivate specific neuronal computing modules. With only a 1% decrease in accuracy, SCSC can reduce synaptic weight access energy by 29% and neuronal computing energy by 49%. Wei Liu 0118, Jinghai Wang, Zhiyi Yu, Shanlin Xiao |
ASP-DAC | 3 |
| 2025 | Towards In-Situ Neuromorphic Computing Architecture for Event Stream Super-ResolutionabstractEvent-based cameras, with their unique event stream representation, effectively mitigate motion blur in highspeed, high-exposure scenarios but suffer from low spatial resolution. To address this, we propose a super-resolution hardware accelerator for event streams based on Spiking Neural Networks (SNNs). In terms of network architecture, we incorporate hardware-friendly algorithmic designs by simplifying neuron models and optimizing convolution operations. On the hardware side, the design adopts a hierarchical structure featuring a highly parallel computational array. Additionally, by proposing a Kernel-Channel-Timestamp-Row (KCTR) dataflow and dual-pipeline structure, the design achieves in-situ computing, eliminating intermediate storage within layers and significantly reducing inter-layer spike storage. Evaluations on the N-MNIST and ASL-DVS datasets demonstrate root mean square errors (RMSE) of $\mathbf{1. 2 9 6}$ and $\mathbf{0. 1 2 1}$ for reconstructed super-resolution event streams. In downstream applications, the classification accuracies reach 98.84% and 99.73%, respectively. The proposed accelerator, designed using a 28 nm CMOS process, improves reconstruction speed by 95.6% compared to a GPU, operates at 500 MHz, and consumes only 0.546 pJ per synaptic operation. Yihe Yu, Wei Liu 0118, Jinghai Wang, Zhiyi Yu, Shanlin Xiao |
DAC | 4 |
| 2025 | A Hybrid Stochastic-Binary Computing Batch Normalization Engine for Low-Power On-Chip Learning Spiking Neural NetworksabstractBatch normalization (BN) has proven to be a critical component in speeding up the training of deep spiking neural networks in deep learning. However, conventional BN implementations face significant challenges in terms of excessive off-chip memory bandwidth requirements and complex circuit designs, hindering their applicability for on-chip training in spiking neural networks (SNNs). This article introduces a novel hybrid stochastic-binary computing BN engine (HBN) that strikes an optimal balance between computational efficiency and hardware resource utilization, enabling efficient on-chip learning for SNNs. While conventional binary-mode BN engines offer temporal efficiency, they demand substantial hardware resources. In contrast, stochastic computing (SC)-based BN approaches reduce hardware overhead but introduce latency penalties and necessitate additional random number generation (RNG) circuits. To overcome these limitations, we propose a hybrid architecture that seamlessly integrates binary and stochastic computing (SC) paradigms. Our co-designed methodology effectively balances computational latency and hardware footprint. This is achieved by a rounding-free SC multiplier unified with binary-circuit map ping, which eliminates latency and RNG overheads. Extensive validation across both static image datasets and neuromorphic datasets demonstrates that HBN maintains algorithmic fidelity while achieving unprecedented computational efficiency. Simulation results reveal 98.7% reduction in floating-point operations (FLOPs), 98.5% latency improvement, and 98.2% energy consumption reduction compared with conventional BN implementations. FPGA implementation on the ZCU102 platform demonstrates practical hardware advantages, including 74.9% reduction in look-up table (LUT) utilization, 83.6% decrease in flip-flop (FF) count, and 13.7% reduction in block RAM (BRAM) allocation. Notably, the design achieves 63.7% power reduction compared with state-of-the-art implementations while maintaining complete DSP-free operation. Wei Liu 0118, Zhiyi Yu, Shanlin Xiao |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |