EDBT 2026 Demo / reviewers in the wild / expert
Guocheng Zhao
dblp:377/5204
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0001-5477-3020ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LSAF: A load-balancing SpGEMM acceleration framework with dynamic package and static partition for multi-core systolic arrays
Yongxiang Cao, Guocheng Zhao, Dongcheng Shi, Runhua Zhang 0002 |
Parallel Comput. | 3 |
| 2025 | HMSA: High-Performance Heterogeneous Mixed-Precision CNN Systolic Array Accelerator on FPGAabstractIn power-constrained and real-time-demanding embedded scenarios, Field-Programmable Gate Arrays (FPGAs) emerge as ideal options for accelerating neural network inference, owing to the reconfigurability, high reliability, and flexibility of FPGAs. Mixed precision quantization technology significantly reduces computational complexity and bandwidth requirements while preserving model accuracy. However, existing FPGA accelerators fail to fully leverage the parallel advantages of mixed precision, which leads to the actual inference speedup being markedly lower than the theoretical prediction. Reviewing existing methods, we found three main drawbacks in enhancing practical computational performance. Firstly, FPGAs primarily rely on DSP slices to achieve high-performance parallel multiplication and accumulation (MAC) in neural network inference. However, the current DSP PE design and data packing methods are not compatible with mixed-precision models. Secondly, the remaining logic resources are not fully utilized to accelerate computations. Thirdly, the mixed-precision quantization bit-width selection method without hardware-guided guidance leads to additional model accuracy loss. To address these challenges, we propose a high-performance heterogeneous mixed-precision systolic array (SA) accelerator, HMSA. It aims to leverage mixed-precision quantization fully, enhancing the practical inference efficiency of neural networks on embedded FPGAs. We propose an optimized DSP data packing method guided by a resource-performance cost model, which enhances the parallel computing performance of accelerators. We propose a heterogeneous convolutional acceleration architecture based on SA architecture with high scalability, enabling efficient utilization of FPGA’s heterogeneous computing resources. In terms of the algorithm, we propose an optimization method for bit-width selection based on FPGA hardware architecture, aiming to avoid accuracy degradation without improving inference speed. Experiments confirm that HMSA on the Xilinx XC7VX690T FPGA reaches a peak throughput of 6.385 TOP/s at W1A8 precision. When inferring mixed precision neural networks, HMSA achieves 3.53×, 5.46×, and 1.58× improvements in actual throughput/DSP compared to state-of-the-art MPA, MSD, and MP-OPU. Compared with the state-of-the-art MBFQuant and Edge-MPQ mixed quantization algorithms, the proposed optimized bit-width selection method effectively reduces the model accuracy loss. Yongxiang Cao, Huiyong Li 0005, Dongcheng Shi, Guocheng Zhao |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2024 | Highly Efficient Load-Balanced Dataflow for SpGEMMs on Systolic ArraysabstractTo enhance the efficiency of sparse neural network models, compression methods are commonly employed to store the non-zero elements in a sparse storage format. Sparse General Matrix Multiplication (SpGEMM) is a critical computation in deep neural networks. However, when utilizing systolic arrays for SpGEMM computations, a challenge arises due to the irregular flow of compressed, non-zero element activation data. This irregularity leads to varying lengths of activation data streams entering the systolic array per batch, potentially resulting in the underutilization of processing units. Our research focuses on repackaging compressed data streams using hardware-software co-design to minimize software pre-processing time. We also package unevenly sized sparse matrix rows post-compression into multiple groups of activation data streams with approximately equal lengths. This approach aims to evenly distribute the workload across the fixed output of the systolic array, thereby improving the utilization rate of Processing Elements (PE). Our evaluation demonstrates that our method achieves a 2.01x acceleration compared to uncompressed sparse data streams and a 3.63x average acceleration relative to TPUs. Furthermore, compared to the state-of-the-art SpGEMM accelerator SADD, our approach achieves an average of 2.01x acceleration. Guocheng Zhao, Yongxiang Cao |
ACM Great Lakes Symposium on VLSI | 1 |
| 2024 | Compression Format and Systolic Array Structure Co-design for Accelerating Sparse Matrix Multiplication in DNNs
Yongxiang Cao, Jixiang Jiang, Guocheng Zhao, Yanfei Song |
ICA3PP (4) | 3 |