EDBT 2026 Demo / reviewers in the wild / expert
Yongxiang Cao
dblp:331/7320
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
0009-0000-2020-5116ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HCTA: A heterogeneous conv-transformer networks accelerator overcoming nonlinear operator bottlenecks
Yongxiang Cao, Yanfei Song |
Neurocomputing | 1 |
| 2026 | LSAF: A load-balancing SpGEMM acceleration framework with dynamic package and static partition for multi-core systolic arrays
Yongxiang Cao, Guocheng Zhao, Dongcheng Shi, Runhua Zhang 0002 |
Parallel Comput. | 1 |
| 2025 | SparDR: Accelerating Unstructured Sparse DNN Inference via Dataflow OptimizationabstractUnstructured sparsity is becoming a key dimension in exploring the inference efficiency of neural networks. However, its data layout presents irregularity, making it difficult to match the parallel computing mode of hardware, resulting in low computational and memory access efficiency. We have studied this issue and found that the main reason is that existing sparse acceleration libraries and compilers perform sparse matrix multiplication optimization exploration through the splitting and reconstruction of sparse patterns, thus ignoring the acceleration of sparse convolution operations centered on data streams, which may miss some optimization opportunities for sparse operations. In this article, we propose SparDR, a general sparse convolution operation acceleration method centered around data streams. Through novel feature map data stream reconstruction and convolutional kernel data representation, redundant zero value calculations are effectively avoided, addressing efficiency is improved, and memory overhead is reduced. SparDR is based on TVM and allows for automatic scheduling across different hardware configurations. Compared with the current mainstream five methods on four types of hardware, the inference delay acceleration reaches 1.1-12 x and the memory usage decreases by 20%. Runhua Zhang 0002, Yongxiang Cao, Yaochen Han |
DATE | 4 |
| 2025 | SAM-Lightning: Segment Anything Model for Efficient Inference and Reduced Memory Footprint
Yanfei Song, Bangzheng Pu, Yongxiang Cao, Runhua Zhang 0002, Yiqing Shen 0003 |
PRICAI (5) | 6 |
| 2025 | HMSA: High-Performance Heterogeneous Mixed-Precision CNN Systolic Array Accelerator on FPGAabstractIn power-constrained and real-time-demanding embedded scenarios, Field-Programmable Gate Arrays (FPGAs) emerge as ideal options for accelerating neural network inference, owing to the reconfigurability, high reliability, and flexibility of FPGAs. Mixed precision quantization technology significantly reduces computational complexity and bandwidth requirements while preserving model accuracy. However, existing FPGA accelerators fail to fully leverage the parallel advantages of mixed precision, which leads to the actual inference speedup being markedly lower than the theoretical prediction. Reviewing existing methods, we found three main drawbacks in enhancing practical computational performance. Firstly, FPGAs primarily rely on DSP slices to achieve high-performance parallel multiplication and accumulation (MAC) in neural network inference. However, the current DSP PE design and data packing methods are not compatible with mixed-precision models. Secondly, the remaining logic resources are not fully utilized to accelerate computations. Thirdly, the mixed-precision quantization bit-width selection method without hardware-guided guidance leads to additional model accuracy loss. To address these challenges, we propose a high-performance heterogeneous mixed-precision systolic array (SA) accelerator, HMSA. It aims to leverage mixed-precision quantization fully, enhancing the practical inference efficiency of neural networks on embedded FPGAs. We propose an optimized DSP data packing method guided by a resource-performance cost model, which enhances the parallel computing performance of accelerators. We propose a heterogeneous convolutional acceleration architecture based on SA architecture with high scalability, enabling efficient utilization of FPGA’s heterogeneous computing resources. In terms of the algorithm, we propose an optimization method for bit-width selection based on FPGA hardware architecture, aiming to avoid accuracy degradation without improving inference speed. Experiments confirm that HMSA on the Xilinx XC7VX690T FPGA reaches a peak throughput of 6.385 TOP/s at W1A8 precision. When inferring mixed precision neural networks, HMSA achieves 3.53×, 5.46×, and 1.58× improvements in actual throughput/DSP compared to state-of-the-art MPA, MSD, and MP-OPU. Compared with the state-of-the-art MBFQuant and Edge-MPQ mixed quantization algorithms, the proposed optimized bit-width selection method effectively reduces the model accuracy loss. Yongxiang Cao, Huiyong Li 0005, Dongcheng Shi, Guocheng Zhao |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2024 | Highly Efficient Load-Balanced Dataflow for SpGEMMs on Systolic ArraysabstractTo enhance the efficiency of sparse neural network models, compression methods are commonly employed to store the non-zero elements in a sparse storage format. Sparse General Matrix Multiplication (SpGEMM) is a critical computation in deep neural networks. However, when utilizing systolic arrays for SpGEMM computations, a challenge arises due to the irregular flow of compressed, non-zero element activation data. This irregularity leads to varying lengths of activation data streams entering the systolic array per batch, potentially resulting in the underutilization of processing units. Our research focuses on repackaging compressed data streams using hardware-software co-design to minimize software pre-processing time. We also package unevenly sized sparse matrix rows post-compression into multiple groups of activation data streams with approximately equal lengths. This approach aims to evenly distribute the workload across the fixed output of the systolic array, thereby improving the utilization rate of Processing Elements (PE). Our evaluation demonstrates that our method achieves a 2.01x acceleration compared to uncompressed sparse data streams and a 3.63x average acceleration relative to TPUs. Furthermore, compared to the state-of-the-art SpGEMM accelerator SADD, our approach achieves an average of 2.01x acceleration. Guocheng Zhao, Yongxiang Cao |
ACM Great Lakes Symposium on VLSI | 2 |
| 2024 | Compression Format and Systolic Array Structure Co-design for Accelerating Sparse Matrix Multiplication in DNNs
Yongxiang Cao, Jixiang Jiang, Guocheng Zhao, Yanfei Song |
ICA3PP (4) | 1 |
| 2023 | LOCP: Latency-optimized channel pruning for CNN inference acceleration on GPUs
Runhua Zhang 0002, Yongxiang Cao, Chenhui Zhu |
J. Supercomput. | 5 |
| 2022 | An Efficient Sparse CNNs Accelerator on FPGAabstractConvolutional Neural Networks (CNNs) have achieved remarkable performance at a huge computational cost. By improving the model sparsity, it can effectively reduce the complexity. However, with deepening of sparsity, the problems of unbalanced workloads, computing fragmentation and mapping access conflict caused by irregular sparsity have become more and more remarkable. These problems pose great challenges for efficient computation of sparse CNN s. In order to make full use of two side of sparsity introduced by activations and weights, and overcome the above problems, this paper proposes an efficient sparse CNN s accelerator on FPGA to achieve the inference acceleration. We designed and implemented the accelerator on the Zynq UltraScale+ MPSoC ZCU102 evaluation board. By running AlexNet, VGG16 and ResNet50 networks on the accelerator to evaluated the peeformance. Experimental results show that the method proposed in this paper can achieve more than 97% reduction in collision rate and 2.35x improvement in computing performance and 9.37x improvement in energy efficiency. Haojie Wang 0004, Dong Dong 0001, Yongxiang Cao |
CLUSTER | 6 |