EDBT 2026 Demo / reviewers in the wild / expert
Lichen Feng
dblp:182/3823
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0002-7685-2141ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust Compression Circuit for Real-time Read-out of 320x240 SPAD Array Based on FPGA
Yaoqi Bao, Zhuang Yan, Dong Li 0046, Lichen Feng, Rui Ma 0007, Zhangming Zhu |
ISCAS | 5 |
| 2026 | Improving MAC Accuracy of Pre-Aligned Floating-Point CIM Macros for Compound AI with Statistical Group Features
Yunlong Liu 0006, Lichen Feng, Dong Li 0046, Zhangming Zhu |
ISCAS | 2 |
| 2026 | An Asynchronous Analog-Computing Spiking Neural Network With Improved Tolerance to Nonidealities for Always-On Near-Sensor AIabstractSpiking Neural Networks (SNN) is well-suited for always-on near-sensor intelligence, due to its spike-driven nature; however, its IC realization is complicated by the temporal dimension. Asynchronous low-power SNN chips employing analog computing-in-memory (CIM) techniques have been demonstrated to enable real-time, energy-efficient inference. However, their tolerance to nonidealities remains to be improved, and their peripherals for multiphase or multilevel signal control are complex. This paper proposes a general-purpose, spike-driven SNN chip designed with efficient analog-computing circuits, featuring two key contributions: 1) A compact direct current-add CIM synapse design that significantly simplifies control peripherals, thereby reducing latency and enhancing energy efficiency. 2) A PVT-aware multi-network learning method underpinned by detailed analyses and modeling, coupled with a label-normalization-based synapse-strengthening method, to mitigate the impact of nonidealities. Fabricated in 65nm CMOS process, the proposed design achieves a state-of-the-art latency of$10\mu $s and an energy efficiency of 0.40pJ/spike. The generalization ability of the design is verified through its successful application to two distinct tasks: voice activity detection (VAD) and ECG anomaly detection. The tolerance to nonidealities is validated by the VAD task. The ten chips maintain over 90% detection accuracy across signal-to-noise ratios (SNRs) of$4\sim 16$dB, ±10% supply voltage variation, and a temperature range of$- 25\sim 55^{\circ }$C. Lichen Feng, Hongwei Shan, Libo Qian, Zhangming Zhu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2025 | Cascade Pre-attention: Regulating Neuronal Activation Distributions in MetaFormer-Based Spiking Neural Networks
Yukun Xue, Lichen Feng |
ICANN (1) | 2 |
| 2024 | An AER-based spiking convolution neural network system for image classification with low latency and high energy efficiency
Lichen Feng, Hongwei Shan, Liying Yang 0001, Zhangming Zhu |
Neurocomputing | 2 |
| 2024 | A Review of Sub- μ W CMOS Analog Computing Circuits for Instant 1-Dimensional Audio Signal Processing in Always-On Edge DevicesabstractReducing the power consumption of intelligence edge devices that require long standby periods is highly significant. The use of CMOS analog computing circuits for instant feature extraction and wake-up signal generation from the sensed signal has emerged as an efficient signal processing paradigm for data volume reduction and power minimization. This review highlights the progress of ultra-low-power (sub-$\mu $W) analog computing circuits for instant 1-dimensional audio signal processing in edge devices in recent years, at architecture and circuit levels. A comparative analysis of existing works is presented, offering choices that cater to various performance requirements. The low power consumption in the range of tens of nanowatts has been achieved. Finally, we provide insights into the future potential of these analog computing circuits. Zhangming Zhu, Lichen Feng |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | A 109-GOPs/W FPGA-Based Vision Transformer Accelerator With Weight-Loop Dataflow Featuring Data Reusing and Resource SavingabstractThe Vision Transformer (ViT) models have demonstrated excellent performance in computer vision tasks, but a large amount of computation and memory access for massive matrix multiplications lead to degraded hardware performance compared to convolutional neural network (CNN). In this paper, we propose a ViT accelerator with a novel “Weight-Loop” dataflow and its computing unit, for efficient matrix multiplication computation. By data partitioning and rearrangement, the number of memory accesses and the number of registers are greatly reduced, and the adder trees are eliminated. A computation pipeline with the proposed dataflow scheduling method is constructed to maintain a high utilization rate through zero bubble switching. Moreover, a novel accurate dual INT8 multiply-accumulate (DI8MAC) method for DSP optimization is introduced to eliminate the additional correction circuits by weight encoding. Verified in the Xilinx XCZU9EG FPGA, the proposed ViT accelerator achieves the lowest inference latencies of 3.91 ms and 13.98 ms for ViT-S and ViT-B, respectively. The throughput of the accelerator can reach up to 2330.2 GOPs with an energy efficiency of 109 GOPs/W, showing a significant improvement compared to the state-of-the-art works. Lichen Feng, Hongwei Shan, Zhangming Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Memory-Efficient Deformable Convolution Based Joint Denoising and Demosaicing for UHD ImagesabstractThis paper introduces deformable convolution in deep learning based joint denoising and demosaicing (JDD), which yields more adaptable representation and larger receptive fields in features extraction for a superior restoration performance. However, the deformable convolution generally leads to considerable computational load and irregular memory access bottleneck, limiting its extensive deployment on edge devices. To address this issue, we develop grouping strategy and assign independent offsets to each kernel group to reduce the computation latency while keeping the accuracy. Motivated by the exploration for aggregate distribution characteristics of deformable offsets, we present the offset sharing methodology to simplify the memory access complexity of deformable convolution. As for hardware acceleration, we specially design a novel deformable matrix multiplication workflow incorporated with a deformable memory mapping unit to boost the computational throughput. The verification experiments on FPGA demonstrate that the proposed deformable convolution based JDD can restore 4K Ultra High Definition (UHD) images at 70FPS and yields significant promotion in visual effect and objective quality assessment. Juntao Guan, Yangang Li, Huanan Li, Lichen Feng, Yintang Yang, Lin Gu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |