EDBT 2026 Demo / reviewers in the wild / expert
Beining Zhao 0001
dblp:360/7861-1
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0006-4550-7974ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Packetized Pipelined Pillar Feature Net Accelerator for LiDAR 3D Object DetectionabstractImplementing LiDAR-based 3D object detection algorithms in practical autonomous driving situations presents a significant challenge. In current research algorithms, the inherent sparsity and randomness of point cloud data necessitate significant memory usage and frequent data read/write operations during preprocessing. Such demands are not well-suited for terminal devices with stringent real-time requirements and constrained resources. In this paper, we present a packetized processing Pillar Feature Net accelerator for LiDAR 3D object detection. By integrating voxelization and feature extraction into a pipelined architecture, the proposed accelerator significantly reduces the storage requirements for point cloud data and enhances the speed of feature extraction and pseudo-image generation. Experimental results indicate that the proposed method improves the computational throughput from point cloud data to pseudo-image generation by 1.2 times and eliminates the need for off-chip memory access during preprocessing. Qingyu Deng, Xinyu Chen 0007, Wei Zhang 0001, Beining Zhao 0001, Yuhang Gu, Shan Cao 0001, Zhiyuan Jiang |
ISCAS | 4 |
| 2025 | MPQA: Mixed-Precision Quantization Accelerator for CNN InferenceabstractMixed-precision quantization CNNs have become crucial for edge vision algorithms. While model quantization techniques have improved inference efficiency, existing approaches have not fully addressed data coherence in coarse-grained parallelism or adaptation to various 2D data sizes. This study presents a novel CNN accelerator equipped with mixed-precision multiplier and feature linking mechanisms to enhance mixed-precision inference on edge NPUs. The accelerator incorporates reconfigurable multi-precision multiplication support in MAC units, enabling 8-bit/4-bit mixed-precision quantization. A feature linking mechanism is developed to overcome feature map size limitations, achieving scalable processing dimensions. The accelerator design, implemented on a Xilinx Zynq UltraScale+ MPSoC FPGA platform, demonstrates improved hardware performance and enhanced inference speeds for edge device CNN models such as Tiny-YOLOv3 and ResNet18. This work provides a significant advancement in deploying mixed-quantization CNNs on edge NPUs, enhancing the adaptability and performance of edge computing devices for complex neural network models. Beining Zhao 0001, Yu Li 0051, Jiahao Zuo, Wei Zhang 0001, Xinyu Chen 0007, Shan Cao 0001, Zhiyuan Jiang |
ISCAS | 1 |
| 2025 | Near-Sensor LiDAR and Visual Feature Extraction and Communication for Low-Latency Roadside Cooperative PerceptionabstractAutonomous driving technologies are swiftly evolving, characterized by two main strategies: Single-Vehicle Autonomous Driving (SVAD) and Vehicle-Infrastructure Cooperative Autonomous Driving (VICAD). SVAD depends entirely on the vehicle’s internal sensors and processing capabilities, whereas VICAD benefits from a synergistic network combining roadside infrastructure, connected vehicles, and cloud services to boost safety and efficiency. Nevertheless, VICAD encounters challenges with high-bandwidth data transmission and perception latency. To mitigate these concerns, we introduce an innovative intelligent roadside unit (I-RSU) platform integrating perception, computing, and communication into one cohesive system. The platform features dual neural processing units (NPUs) for the effective extraction of images and LiDAR features, alongside a C-V2X communication module, all realized on a Field-Programmable Gate Array (FPGA). This setup minimizes latency and expenses by enabling computation near the sensors and facilitating selective data transmission. Our system also supports multi-modal fusion, enhancing overall perception and safety. Through extensive real-world trials and simulations, our system demonstrates a substantial reduction in end-to-end latency, providing a scalable solution for VICAD scenarios. Wei Zhang 0388, Yuhang Gu, Beining Zhao 0001, Qingyu Deng, Xinyu Chen 0007, Yi Shi 0004, Limin Jiang, Shan Cao 0001, Zhiyuan Jiang, Ruiqing Mao, Sheng Zhou 0001 |
IEEE Internet Things J. | 3 |
| 2025 | A Heterogeneous CNN Compilation Framework for RISC-V CPU and NPU Integration Based on ONNX-MLIRabstractWith the continuous advancement of convolutional neural networks (CNNs), many neural network processing units (NPU) have emerged in recent years. NPUs offer advantages such as improved energy efficiency and low latency compared to traditional processors. However, NPUs often struggle to keep pace with the growing complexity of algorithmic models, which limits the implementation of certain AI applications. Central processing units (CPU) are known for their versatility, but are computationally inefficient. Combining the strengths of both architectures could present a viable solution to these challenges. Despite this potential, there is currently insufficient academic focus on CPU/NPU co-computing, and few architectures or toolchains have been developed to support such integration. In this paper, we propose a RISC-V CPU/NPU co-computing-based compilation framework to solve the problem of mapping from algorithms to heterogeneous architectures. To bridge the computational gap between heterogeneous platforms, within ONNX-MLIR, we developed three phases to support operator packing, algorithm adaption, and data coordination. To integrate heterogeneous compilation platforms, we introduce four additional functions to handle task scheduling, weight quantization, weight rearrangement, and memory management. In particular: (1) to optimize the transition overhead, an operator packing technique is introduced by extending the ONNX-MLIR framework with custom operators, which enhances NPU efficiency by reducing the number of transitions between operations. (2) At the aspect of task scheduling, a three-mode task scheduler is developed to allow users to customize task distribution for correctness verification, performance optimization, or detailed analysis on heterogeneous platforms, which demonstrates the versatility of our framework. (3) In terms of memory management, an optimized scheme for CPU/NPU architectures is proposed to ensure accurate data access. This memory management employs a dual-end approach with self-checks and a release-after-use mechanism to further improve efficiency. Experimental results demonstrate that the framework effectively generates instructions for various CNNs, significantly enhancing the efficiency and performance of heterogeneous architectures. It achieves up to a 7.06× reduction in code density and up to a 5.58× improvement in memory usage, narrowing the computational gap between CPUs and NPUs. Shan Cao 0001, Meiling Yang, Yintao Liu 0001, Yu Li 0051, Beining Zhao 0001, Xinyu Chen 0007, Zhiyuan Jiang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | Hybrid-Grained Pruning and Hardware Acceleration for Convolutional Neural NetworksabstractThroughout various convolutional neural network (CNN) models, the sparsity increases as the network deepens, which poses significant potential to model compression and hardware acceleration. In this paper, a dual-factor hybrid-grained pruning method is introduced to make a good balance between model compression and accuracy preservation. The pro-posed pruning method combines hardware-friendly unstructured vector-level pruning with structured filter-level pruning to explore multiple grains of sparsity in CNNs. The architecture of the corresponding hardware accelerator is then proposed based on the row-based convolution dataflow, which could fully utilize the hybrid sparsity to accelerate CNN processing. Experimental results demonstrate that the proposed method increases the compression rate by 1.08× while causing 0.21% accuracy loss compared to the state-of-the-art filter pruning method in VGG16, and 2.39% hardware resource increase compared to the accelerator without sparsity optimization. Yu Li 0051, Shan Cao 0001, Beining Zhao 0001, Wei Zhang 0001, Zhiyuan Jiang |
ISCAS | 3 |