Zongpei Fu

dblp:372/2278 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0008-0166-4104ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Memory-Efficient In-Sensor Event Denoising with a Lightweight Point-Cloud Network
Zongpei Fu, Xiaojin Zhao, Wen Bin Ye 0001
ISCAS1
2026 A Low-Cost Energy-Efficient FFT-CNN Processor for Versatile mmWave Radar in IoT Applications
abstract
The millimeter-wave (mmWave) radar Internet of Things (IoT) applications based on Convolutional Neural Network (CNN) algorithms are widely employed across various industries. However, they face challenges in deployment on low-cost and resource-constrained edge devices. Current CNN accelerators can efficiently handle CNN computations but cannot effectively accelerate the essential fast Fourier Transform (FFT) computations in mmWave radar signal processing, thus failing to achieve end-to-end high-efficiency radar signal processing. This work presents a low-cost energy-efficient FFT-CNN processor with a unified architecture that combines CNN and FFT to enhance performance for mmWave radar IoT applications. The proposed FFT-CNN processor employs a reconfigurable dual-mode processing element to maximize hardware resource sharing, supporting both butterfly operations for FFT and multiply-accumulate computations for CNN acceleration. This sharing technique reduces hardware resource consumption by 22.1% in LUTs, 22.7% in FFs, 9.5% in BRAMs, and 16.7% in DSPs, leading to a total power reduction of 34.9%. Compared to state-of-the-art processors optimized for radar applications, the proposed processor achieves the highest performance-to-resource ratio and energy efficiency, making it ideal for a wide range of low-cost radar IoT applications.
Juhua Chen, Zongpei Fu, Wen Bin Ye 0001
IEEE Trans. Circuits Syst. I Regul. Pap.2
2026 A 95.3% 12-Class, 108-nJ/Inference Keyword Spotting Chip With Hybrid FFT-BFNet Architecture and Exponent-Aware Nonuniform Quantization in 65-nm CMOS
abstract
This article presents a 65-nm keyword spotting (KWS) chip that achieves 95.3% accuracy on 12-class tasks with 108.04-nJ/inference efficiency through cross-domain hardware-algorithm innovations. The unified fast Fourier transform (FFT)-butterfly-structured neural network (BFNet) accelerator fundamentally rethinks computational reuse: by replacing dense pointwise convolutions with butterfly-based sparse operations mirroring FFT’s dataflow, it slashes$6.3\times $multiply-accumulate (MAC) operations and halves parameter counts while preserving model expressivity. A 6-bit exponent-aware nonuniform quantization (EANUQ) scheme compresses weights, achieving a 25% reduction in storage while maintaining an accuracy loss of less than 0.01% with lightweight on-chip decoders. Hardware resource sharing extends beyond computation: Mel-filter-banks reuse fully-connected (FC) layer multipliers through decomposed 8-bit arithmetic, and FFT output buffers double as convolutional neural network (CNN) feature map memory. Measured at 0.65 V/600 kHz, the 0.58-$\text {mm}^{2}$core demonstrates$1.9\times $–$15.5\times $better energy efficiency than prior 65/28-nm implementations, with 14.82-ms end-to-end latency.
Zongpei Fu, Kaixiang Lin, Xiaojin Zhao, Wen Bin Ye 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2024 A 593nJ/Inference DVS Hand Gesture Recognition Processor Embedded With Reconfigurable Multiple Constant Multiplication Technique
abstract
Hand gesture recognition (HGR) is a popular technique for edge-based human-computer interaction. Dynamic vision sensors (DVS) are often used in HGR systems due to their low latency, high dynamic range, low energy consumption, and asynchronous event triggering. While Spiking Neural Networks (SNNs) are commonly thought to consume less energy than Convolutional Neural Networks (CNNs) in DVS-based HGR systems, this work demonstrates that a DVS-based HGR system on chip (SoC) incorporating CNN can achieve lower power consumption through algorithm and hardware co-design. The proposed edge-side processor for DVS-based HGR integrates a median filter processing core and an AI accelerator core for preprocessing and CNN inference on the DVS output data. To reduce hardware costs without sacrificing accuracy, the median filtering core uses a simplified median filtering function tailored to the specific application scenario. The paper suggests using reconfigurable multiple constant multiplication (RMCM) techniques for the AI accelerator core to share computational resources among processing element (PE) arrays, thereby reducing computational costs and power consumption. The entire DVS gesture processor was implemented in a 65nm technology, achieving an energy requirement of 593.4nJ per inference on-chip with a guaranteed accuracy of 92.4%.
Zongpei Fu, Wen Bin Ye 0001
IEEE Trans. Circuits Syst. I Regul. Pap.1