Yipu Zhang 0002

dblp:118/5586-2 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0007-9465-5807ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HERO: Hardware-Efficient RL-based Optimization Framework for NeRF Quantization
Yipu Zhang 0002, Chaofang Ma, Jinming Ge, Jiang Xu 0001, Wei Zhang 0012
ASP-DAC1
2026 FLICKER: A Fine-Grained Contribution-Aware Accelerator for Real-Time 3D Gaussian Splatting
abstract
Recently, 3D Gaussian Splatting (3DGS) has become a mainstream rendering technique for its photorealistic quality and low latency. However, the need to process massive noncontributing Gaussian points makes it struggle on resource-limited edge computing platforms and limits its use in next-gen AR/VR devices. A contribution-based prior skipping strategy is effective in alleviating this inefficiency, but the associated contribution-testing workload becomes prohibitive when it is further applied to the edge. In this paper, we present FLICKER, a contribution-aware 3DGS accelerator that leverages a hardware–software co-design framework, including adaptive leader pixels, pixel-rectangle grouping, hierarchical Gaussian testing, and mixed-precision architecture, to achieve near-pixel-level, contribution-driven rendering with minimal overhead. Experimental results show that our design achieves up to 1.5× speedup, 2.6× energy efficiency improvement, and 14% area reduction over a state-of-the-art accelerator. Meanwhile, it also achieves 19.8× speedup and 26.7× energy efficiency compared with a common edge GPU.
Wenhui Ou, Zhuoyu Wu, Yipu Zhang 0002, Dongjun Wu, Frederick Ziyang Hong, C. Patrick Yue
DATE3
2026 DRACO: A Hardware-Efficient Robot Rigid Body Dynamics Accelerator with Precision-Aware Quantization Framework
abstract
Rigid Body Dynamics (RBD) computation is a critical component of robotic control, often dominating system runtime due to its algorithmic complexity and high parallelism demands. CPUs suffer from limited parallelism and cache-unfriendly access patterns, while GPUs incur prohibitive memory-access latency and per-task response time, making them unsuitable for real-time control. Both platforms also consume excessive power for edge deployment. FPGAs offer superior latency, energy efficiency, and customizable hardware-level parallelism, emerging as promising targets for RBD acceleration. However, existing FPGA designs still face critical limitations. First, the intensive use of multiply-accumulate operations leads to high Digital Signal Processing (DSP) slices consumptionespecially for high degrees-of-freedom (DOF) robots-resulting in limited scalability. Second, RBD functions include mass matrix inversion function, which is inefficient on FPGA due to reciprocal operations falling on the longest latency path, severely limiting performance. Third, mismatched processing rates across modules introduce idle cycles, resulting in poor DSP utilization. To address these issues, we propose DRACO, a hardwareefficient and high-performance RBD accelerator based on FPGA, introducing three key innovations. First, we propose a precisionaware quantization framework that reduces DSP demand by up to$4 \times$while preserving motion accuracy. This is also the first study to systematically evaluate quantization impact on robot control and motion for hardware acceleration. Second, we leverage a hardware-efficient division deferring optimization in mass matrix inversion algorithm, which decouples reciprocal operations from the longest latency path to improve the performance. Finally, we present an inter-module DSP reuse methodology to improve DSP utilization and save DSP usage. Experiment results show that DRACO achieves up to$8 \times$throughput improvement and$7.4 \times$latency reduction over state-of-the-art (SOTA) RBD accelerators across various robot types, demonstrating its effectiveness and scalability for high-DOF robotic systems.
Yipu Zhang 0002, Linfeng Du, Chaofang Ma, Jiang Xu 0001, Wei Zhang 0012
HPCA3
2026 Guiding multimodal LLMs for efficient visual place recognition
Zhijian He, Jintao Cheng, Yipu Zhang 0002, Chi-Man Vong, Jin Wu 0002, Xieyuanli Chen
Pattern Recognit. Lett.4
2025 SpNeRF: Memory Efficient Sparse Volumetric Neural Rendering Accelerator for Edge Devices
abstract
Neural rendering has gained prominence for its high-quality output, which is crucial for AR/VR applications. However, its large voxel grid data size and irregular access patterns challenge real-time processing on edge devices. While previous works have focused on improving data locality, they have not adequately addressed the issue of large voxel grid sizes, which necessitate frequent off-chip memory access and substantial on-chip memory. This paper introduces SpNeRF, a software-hardware co-design solution tailored for sparse volumetric neural rendering. We first identify memory-bound rendering inefficiencies and analyze the inherent sparsity in the voxel grid data of neural rendering. To enhance efficiency, we propose novel preprocessing and online decoding steps, reducing the memory size for voxel grid. The preprocessing step employs hash mapping to support irregular data access while maintaining a minimal memory size. The online decoding step enables efficient on-chip sparse voxel grid processing, incorporating bitmap masking to mitigate PSNR loss caused by hash collisions. To further optimize performance, we design a dedicated hardware architecture supporting our sparse voxel grid processing technique. Experimental results demonstrate that SpNeRF achieves an average 21.07× reduction in memory size while maintaining comparable PSNR levels. When benchmarked against Jetson XNX, Jetson ONX, RT-NeRF. Edge and NeuRex. Edge, our design achieves speedups of 95.1×, 63.5×, 1.5× and 10.3×, and improves energy efficiency by 625.6×, 529.1×, 4×, and 4.4×, respectively.
Yipu Zhang 0002, Jiang Xu 0001, Wei Zhang 0012
DATE1
2025 FLEX: Leveraging FPGA-CPU Synergy for Mixed-Cell-Height Legalization Acceleration
abstract
Legalization is a critical yet time-consuming step in very large-scale integration (VLSI) design, tasked with iteratively relocating standard cells to eliminate overlaps while resolving design rule violations. This process is repeatedly invoked during VLSI physical design. However, increasing spatial constraints and complex design rules impose significant challenges on existing CPU- and GPU-based legalizers, including suboptimal task assignment, inefficient algorithm, and long hardware idle time caused by processing tasks with irregular computational patterns in parallel.
Linfeng Du, Yipu Zhang 0002, Chaofang Ma, Hanwei Fan, Jiang Xu 0001, Wei Zhang 0012
ICPP4
2024 Data-Pattern-Based Predictive On-Chip Power Meter in DNN Accelerator
abstract
Advanced power management techniques, such as voltage drop mitigation and fast power management, can greatly enhance energy efficiency in contemporary hardware design. Nevertheless, the implementation of these innovative techniques necessitates accurate and fine-grained power modeling, as well as timely responses for effective coordination with the power management unit. Additionally, existing performance-counter-based and RTL-based on-chip power meters have difficulty in providing sufficient response time for fast power and voltage management scenarios. In this article, we propose PROPHET, a data-pattern-based power modeling method for multiply-accumulate-based (MACC) deep neural network (DNN) accelerators. Our proposed power model extracts the predefined data patterns during memory access and then a pretrained power model can predict the dynamic power of the DNN accelerators. Thus, PROPHET can predict dynamic power and provide sufficient responding time for power management units. In the experiments, we evaluate our predictive power model in four DNN accelerators with different dataflows and data types. In power model training and verification, our proposed data-patterns-based power model can realize the 2-cycle temporal resolution with$R^{2} \gt 0.9$, normalized mean absolute error <7%, and the area and power overhead lower than 4.5%.
Tingyuan Liang, Jingbo Jiang, Yipu Zhang 0002, Zhe Lin 0007, Zhiyao Xie, Wei Zhang 0012
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4