Jiapei Zheng

dblp:328/8755 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2026
0009-0001-8917-735XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 4 first-author · 10 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 NS-FPS: Accelerating Farthest Point Sampling via Neighbor Search in Large-Scale Point Clouds
Jiapei Zheng, Shuan Yang, Siqi He, Qi Liu 0010, Chixiao Chen
ISCA1
2026 PSVS: A Parallel-Series Voltage-Sensing 1T4R RRAM Macro Achieving 15-Level Storage and 17.88 Mb/mm2 Density for LLM Inference
Kaijun Zhang, Wencong Wu, Jiapei Zheng, Jinru Lai, Xiaoxin Xu, Qi Liu 0010, Chixiao Chen
ISCAS4
2026 COPIC: A Codesigned Accelerator for Efficient Octree-Based Learned Point Cloud Compression at Edge
abstract
With the growing adoption of light detection and ranging (LiDAR) technology and the increasing resolution of captured data, efficient point cloud compression (PCC) has become a critical challenge. While PCC is essential for reducing bandwidth and storage demands, existing octree-based learned PCC methods face two fundamental bottlenecks: 1) octree construction and serialization that involve irregular, pointwise memory accesses that are difficult to parallelize and 2) autoregressive entropy model inference that requires heavy computation. These issues lead to high latency and energy consumption, making real-time deployment on edge devices impractical. This article introduces COPIC, a hardware accelerator designed for real-time PCC on edge devices. COPIC integrates two key components. The first is OctCAM, a content-addressable-memory (CAM)-based module for rapid octree construction and serialization. By enabling hardware-level parallel search, OctCAM significantly reduces the latency and energy consumption caused by irregular memory access patterns in octree processing. The second is PCC-Engine, a dedicated neural network inference module. Combined with a lightweight PCC-TinyAttention network and a window-based key-value caching approach, it further minimizes execution latency and power usage. Experimental results show that COPIC achieves a$25.9\times $speedup compared to GPU implementations, with a total processing latency of 40.1 ms and an average energy consumption of 0.23 J per frame, while delivering a compression rate 12.4% higher than standard proposals.
Jiapei Zheng, Yutong Su, Siqi He, Haozhe Zhu, Qi Liu 0010, Chixiao Chen
IEEE Trans. Very Large Scale Integr. Syst.1
2025 Hydra: Harnessing Expert Popularity for Efficient Mixture-of-Expert Inference on Chiplet System
abstract
The rapid growth of model sizes in advanced artificial intelligence algorithms, particularly in Transformerbased large language models (LLMs), has led to significant computational overhead. Mixture-of-Expert (MoE) models offer a solution through their sparsely gating mechanism but introduce new challenges of extensive all-to-all communication and model computational inefficiencies. This paper presents Hydra, a software/hardware co-design aimed at accelerating MoE inference on chiplet-based architectures. In software, Hydra employs a popularity-aware expert mapping strategy to optimize interchiplet communication. In hardware, it incorporates Content Addressable Memory (CAM) to eliminate expensive explicit token (un)-permutation based on sparse matrix multiplications and a redundant-calculation-skipping softmax engine to bypass unnecessary division and exponential operations. Evaluated in 22 nm technology, Hydra achieves latency reductions of $14.2 \times$ and $3.5 \times$ and power reductions of $169.1 \times$ and $18.9 \times$ over GPU and state-of-the-art MoE accelerator, respectively, thereby offering a scalable and efficient solution for MoE model deployment.
Siqi He, Haozhe Zhu, Jiapei Zheng, Lizhou Wu, Bo Jiao 0003, Qi Liu 0010, Xiaoyang Zeng, Chixiao Chen
DAC3
2025 A 0.22 pJ/bit Processing-in-Controller GEMV Macro with Weight Prefetch for Efficient Near-Memory Computing
abstract
3D-stacked DRAM is a key technology enabling the development of large language models (LLMs). However, the intensive computational demands lead to substantial energy consumption resulting from large-scale data movement. Integrating processing-in-memory (PIM) within DRAM has been employed to reduce data movement, but it elevates manufacturing costs and incurs additional area and power consumption. On the other hand, implementing processing-near-memory (PNM) outside the DRAM controller fails to eliminate the high energy consumption associated with interconnects during data transfer. To address these challenges, we propose a processing-in-controller (PIC) architecture aimed at 3D-stacked DRAM for efficient data movement. A size-scalable General Matrix-Vector Multiplication (GEMV) structure is proposed, supporting configurations ranging from 16 × 16 to 128 × 128, thereby enhancing its adaptability for a variety of applications. To address the low throughput bottleneck caused by DDR read latency, a ping-pong buffer with weight prefetching is introduced in the PNM. Based on 28 nm CMOS technology, experimental results demonstrate that the energy consumption of this PIC is 0.22 pJ/bit, while data movement efficiency improves by nearly 92% compared to traditional DRAM operations.
Jiapei Zheng, Siqi He, Lizhou Wu, Chen Mu, Haozhe Zhu, Liyu Lin, Qi Liu 0010, Chixiao Chen
ISCAS2
2024 CAMPER: Exploring the Potential of Content Addressable Memory for 3D Point Cloud Efficient Range Search
abstract
The use of Light Detection and Ranging (LiDAR) for sensing has continuously improved the precision and performance of autonomous driving. At the same time, the large number of high-precision point clouds generated by LiDAR require real-time processing, and the range search is the key part of the processing pipeline. Content-Addressable Memory (CAM) has proven its efficiency for search tasks on switches and routers, but so far, there is still a lack of exploration on its application in point cloud range search. In this work, we propose CAMPER, a CAM-centered accelerator, aiming to explore the potential of CAM for point cloud range search. We developed a ripple comparison 13T (RC-13T) CAM cell for distance comparison, designed a spatial approximation search algorithm based on Chebyshev distance, and discussed the flexibility and scalability of the architecture. The results show that in the range search task of 64k@64k points, CAMPER achieves a latency of 0.83ms and a power consumption of 114.6mW. Compared with GPU, the throughput is increased by 10.4×; compared with SOTA accelerator, the energy efficiency is increased by about 228×.
Jiapei Zheng, Lizhou Wu, Yutong Su, Zhangcheng Huang 0001, Chixiao Chen, Qi Liu 0010
DAC1
2024 GauSPU: 3D Gaussian Splatting Processor for Real-Time SLAM Systems
abstract
3D Gaussian Splatting (3DGS) has recently emerged as a promising technique in the realms of 3D vision and robotics. Its capacity for rapid rendering and high-fidelity reconstruction makes it an attractive candidate for integration into Simultaneous Localization and Mapping (SLAM) systems. However, existing 3DGS-based SLAM systems still suffer from inadequate tracking throughput due to tremendous recursion in volume rendering and irregular memory access for gradient backpropagation. To address these challenges, this paper proposes GauSPU, an algorithm-hardware co-designed accelerator for supporting real-time 3DGS-based SLAM. On the algorithm side, we present a sparse-tile-sampling (STS) method for efficient pose tracking. The STS focuses on informative image regions, discarding the rest to alleviate computational workload while maintaining accuracy. At the hardware level, we make twofold efforts. Firstly, we design a sparsity-adaptive ray recursion unit (SA-RRU) to accelerate volume rendering by leveraging irregular spatial sparsity. The SA-RRU introduces a sub-tile-wise execution pattern and a Morton-based thread allocation scheme to optimize sparsity utilization. Additionally, a sparsity-aware task dispatcher ensures efficient fine-grained task scheduling. Secondly, we propose a memory-access-relaxed backpropagation engine (MAR-BE) for efficient gradient aggregation. It comprises a gradient buffer unit (GBU) for coalescing partial gradients and a pose backward unit (PBU) for pipeline-fused backpropagation, collaboratively eliminating the costly atomic operations. Sufficient experiments demonstrate that, through the integration of GauSPU and GPU, the system achieves a throughput of 33.6 FPS for real-time pose tracking in 3DGS-SLAM, presenting a significant$63.9\times$improvement in energy efficiency compared to the RTX3090 baseline.
Lizhou Wu, Haozhe Zhu, Siqi He, Jiapei Zheng, Chixiao Chen, Xiaoyang Zeng
MICRO4
2024 Hi-NeRF: A Multicore NeRF Accelerator With Hierarchical Empty Space Skipping for Edge 3-D Rendering
abstract
Neural radiance field (NeRF) has proved to be promising in augmented/virtual-reality applications. However, the deployment of NeRF on edge devices suffers from inadequate throughput due to redundant ray sampling and congested memory access. To address these challenges, this article proposes Hi-NeRF, a multirendering-core accelerator for efficient edge NeRF rendering. On the architecture level, a hierarchical empty space skipping (HESS) scheme is adopted, which efficiently locates the effective samples with fewer skipping steps and thus accelerates the ray marching process. Furthermore, to alleviate the memory access bottleneck, a vertex-interleaved mapping (VIM) method that eliminates memory bank conflicts is also proposed. On the hardware level, ineffective sample filters (ISFs) and voxel access filters (VCFs) are introduced to further exploit spatial sparsity and data locality at run-time. The experimental results show that our work achieves$2.67\times $rendering throughput and$11.2\times $energy efficiency compared to a SOTA NeRF rendering accelerator. The energy efficiency can be improved by$561\times $compared to a commercial GPU.
Lizhou Wu, Haozhe Zhu, Jiapei Zheng, Yinuo Cheng, Qi Liu 0010, Xiaoyang Zeng, Chixiao Chen
IEEE Trans. Very Large Scale Integr. Syst.3
2023 TiPU: A Spatial-Locality-Aware Near-Memory Tile Processing Unit for 3D Point Cloud Neural Network
abstract
Energy-efficient 3D point cloud neural network accelerators are desired for autonomous driving and AR/VR applications. This paper proposes TiPU, a spatial-locality-aware near-memory tile processing unit where the point clouds are partitioned into tiles to process spatial features locally. Intra-tile farthest point sampling and cross-tile neighbor search are employed to avoid unnecessary distance computing. To efficiently facilitate the tile operations, TiPU architecture consists of a tile-based unified distance computing unit, a near-CAM feature extractor, and a near-SRAM-computing MLP engine. The experimental results show that, compared to GPU implementation, TiPU achieves 15.7× processing speed and reduces 7308× energy consumption.
Jiapei Zheng, Hao Jiang 0024, Xinkai Nie, Zhangcheng Huang 0001, Chixiao Chen, Qi Liu 0010
DAC1
2023 A $2.53 \mu \mathrm{W}/\text{channel}$ Event-Driven Neural Spike Sorting Processor with Sparsity-Aware Computing-In-Memory Macros
abstract
Spike sorting processors with high energy efficiency are widely used in large-scale neural signal processing tasks to monitor the activity of neurons in brains. This paper presents a low-power processor for high-accuracy spike sorting and on-chip incremental learning using an algorithm-hardware co-design approach. The processor introduces an event-driven mechanism with adaptive-threshold detection to conditionally activate the system in order to reduce power consumption. Sparsity-aware computing-in-memory (CIM) macros are also developed in our design to store templates and perform complicated computations efficiently. The prototype is designed using 28nm technology with an area of 0.018 mm2/channel and an overall power efficiency of$\mathbf{2.53} \mu \mathbf{W}/\mathbf{channel}$and 84nW/(channel.cluster) at the voltage of 0.72V. Moreover, the accuracy of the whole design can reach 94.5% in a 32-channel scenario.
Hao Jiang 0024, Jiapei Zheng, Yunzhengmao Wang, Jinshan Zhang 0006, Haozhe Zhu, Liangjian Lyu, Yingping Chen, Chixiao Chen, Qi Liu 0010
ISCAS2