Yasong Cao

dblp:321/8481 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
0009-0003-7683-989XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 first-author · 8 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Precision boundary modeling for area-efficient Block Floating Point accumulation
Xin Ju 0005, Yasong Cao, Zhongdi Luo, Jianchao Yang, Jingkui Yang, Dong Chen 0015, Mei Wen
J. Syst. Archit.4
2025 WinAcc: Window-based Acceleration of Neural Networks Using Block Floating Point
abstract
Deep Neural Networks (DNNs) impose significant computational demands, necessitating optimizations for computational and energy efficiencies. Per-vector scaling, which applies a scaling factor to blocks of elements using narrow integer types, effectively reduces storage and computational overhead. However, the frequent occurrence of floating-point accumulations between vectors limits further improvements in energy efficiency. State-of-the-art accelerators address this challenge by grouping and summing vector products based on their exponent differences, thereby reducing the overhead associated with intra-group shifting and accumulation. Nevertheless, this approach increases the complexity of register usage and grouping logic, leading to limited energy benefits and hardware efficiency. In this context, we introduce WinAcc, a novel algorithm and architecture co-designed solution that utilizes a low-cost accumu-lator to handle the majority of data in DNNs, offering low area overhead and high energy efficiency gains. Our key insight is that the data of DNNs follows a Laplace-like distribution, which enables the use of a customized data format with a narrow dynamic range to encode most of the data. This allows for the design of a low-cost accumulator with narrow shifters and adders, significantly reducing reliance on floating-point accumulator and consequently improving energy efficiency. Compared with state-of-the-art architecture Bucket, WinAcc achieves 33.95% energy reduction across seven representative DNNs and reduces area by 9.5% while maintaining superior model performance.
Xin Ju 0005, Mei Wen, Yasong Cao, Junzhong Shen, Zhaoyun Chen, Yang Shi 0008
DATE5
2025 SparSynergy: Unlocking Flexible and Efficient DNN Acceleration Through Multi-Level Sparsity
abstract
To more effectively address the computational and memory requirements of deep neural networks (DNNs), leveraging multi-level sparsity-including value-level and bit-level sparsity-has emerged as a pivotal strategy. While substantial research has been dedicated to exploring value-level and bit-level sparsity individually, the combination of both has largely been overlooked until now. In this paper, we propose SparSynergy, which-to the best of our knowledge-is the first accelerator that synergistically integrates multi-level sparsity into a unified framework, maximizing computational efficiency and minimizing memory usage. However, jointly considering multi-level sparsity is non-trivial, as it presents several challenges: (1) increased hardware overhead due to the complexity of incorporating multiple sparsity levels, (2) bandwidth-intensive data transmission during multiplexing, and (3) decreased throughput and scalability caused by bottlenecks in bit-serial computation. Our proposed SparSynergy addresses these challenges by introducing a unified sparsity format and a cooptimized hardware design. Experimental results demonstrate that SparSynergy achieves a 5.38 x geometric mean improvement in the energy-delay product (EDP) when compared with the tensor core, across workloads with varying degrees of sparsity. Furthermore, SparSynergy significantly improves accuracy retention compared to state-of-the-art accelerators for representative DNNs.
Jingkui Yang, Mei Wen, Junzhong Shen, Jianchao Yang, Yasong Cao, Minjin Tang, Zhaoyun Chen, Yang Shi 0008
DATE5
2024 BitShare: An Efficient Precision-Scalable Accelerator with Combining-Like-Terms GEMM
abstract
Narrow-precision fixed-point (INT) computation is a significant approach for reducing memory requirements and enhancing the performance of accelerators for Deep Neural Networks (DNNs). Different DNNs, as well as different layers within the DNNs, may exhibit varying numerical distributions, necessitating INT formats with different minimum bit-widths. Therefore, DNN accelerators need to support multi-precision INT computations to strike a better balance between DNN inference accuracy and performance. However, existing precision-scalable accelerators face challenges such as low bandwidth utilization, insufficient utilization of computing resources across different precision modes, and complex circuit structures with associated overhead. In this paper, we propose (1) a hardware-friendly Combining-Like-Terms GEMM (CLT-GEMM) scheme that supports multiple computing modes of 2/4/8 bits and their combinations to align with the various bit-width settings of DNNs; (2) and subsequently design an efficient systolic accelerator with scalable precision, named BitShare, which features DataMap module and Multi-mode adder-tree-based accumulators. Compared to the state-of-the-art precision-scalable design, BitBlade, our accelerator achieves a 57.25% reduction in bandwidth requirement and exhibits an improvement of$1.14\times$and$1.12\times$in area and power efficiency$(2\mathbf{b}\times 2\mathbf{b})$, respectively.
Yasong Cao, Mei Wen, Junzhong Shen, Zhongxing Li
ASAP1
2024 ABS: Accumulation Bit-Width Scaling Method for Designing Low-Precision Tensor Core
abstract
A big gap exists between deep neural network (DNN) applications’ computational demand and the computing power of DNN accelerators. Low-precision floating-point (LP-FP) computation is one of the important means to improve the performance of DNN training and inference. However, the high-precision accumulators are typically applied to summating the dot products during general matrix multiplication (GEMM) in tensor cores (TCs). As the precision of data decreases, the accumulator becomes the main consumer of multiply-accumulate’s (MAC’s) area and power. Reducing the accumulators’ bit-width is of significant importance for improving the area- and energy-efficiency of TCs. There are two main challenges: 1) theoretical support on the floating-point (FP) formats with the lowest bit-width of TC’s accumulators and 2) how to integrate the LP-FP TC in the framework of DNN training and inference to evaluate its benefits. In this article, we propose accumulation bit-width scaling (ABS), a novel ABS method, to guide the design of LP-FP TCs. We 1) implement this method by constructing a novel variance retention ratio (VRR) model to predict the FP format with the minimum bit-width for TC’s accumulator; 2) provide a generator of DNN accelerator based on a systolic-array (SA) TC, supporting many low-precision configurations; and 3) design an LP-FP DNN executing framework that supports software-simulation mode and hardware-accelerator mode to run LP-FP DNN tasks. The experimental results show that the LP-FP TC guided by our ABS method has a maximum reduction of 76.47% and 75.60% in area and power consumption, respectively, compared with the advanced TCs.
Yasong Cao, Mei Wen, Zhongdi Luo, Xin Ju 0005, Haolan Huang, Junzhong Shen
IEEE Trans. Very Large Scale Integr. Syst.1
2022 BP-Im2col: Implicit Im2col Supporting AI Backpropagation on Systolic Arrays
abstract
State-of-the-art systolic array-based accelerators adopt the traditional im2col algorithm to accelerate the inference of convolutional layers. However, traditional im2col cannot efficiently support AI backpropagation. Backpropagation in convolutional layers involves performing transposed convolution and dilated convolution, which usually introduces plenty of zero-spaces into the feature map or kernel. The zero-space data reorganization interfere with the continuity of training and incur additional and non-negligible overhead in terms of off- and on-chip storage, access and performance. Since countermeasures for backpropagation are rarely proposed, we propose BP-im2col, a novel im2col algorithm for AI backpropagation, and implement it in RTL on a TPU-like accelerator. Experiments on TPU-like accelerator indicate that BP-im2col reduces the backpropagation runtime by 34.9% on average, and reduces the bandwidth of off-chip memory and on-chip buffers by at least 22.7% and 70.6% respectively, over a baseline accelerator adopting the traditional im2col. It further reduces the additional storage overhead in the backpropagation process by at least 74.78%.
Jianchao Yang, Mei Wen, Junzhong Shen, Yasong Cao, Minjin Tang, Renyu Yang, Jiawei Fei, Chunyuan Zhang
ICCD4
2022 Mentha: Enabling Sparse-Packing Computation on Systolic Arrays
abstract
Generalized Sparse Matrix-Matrix Multiplication (SpGEMM) is a critical kernel in domains like graph analytic and scientific computation. As a kind of classical special-purpose architecture, systolic arrays were first used for complex computing problems, e.g., matrix multiplication. However, classical systolic arrays are not efficient enough when handling sparse matrices due to the fact that the PEs containing zero-valued entries perform unnecessary operations that do not contribute to the result. Accordingly, in this paper, we propose Mentha, a framework that enables systolic arrays to accelerate sparse matrix computation by employing a sparse-packing algorithm suitable for various dataflow of systolic array. Firstly, Mentha supports both online and offline methods. By packing the rows or columns of the sparse matrix, the zero-valued items in the matrix are significantly reduced and the density of the matrix is improved. In addition, acceleration benefits can be obtained by the adaptation scheme even with limited resources. Moreover, we reconfigure PEs in systolic arrays at a low cost (1.28x in area, 1.21x in power) and find that our method outperforms TPU-like systolic arrays by 1.2~3.3x in terms of SpMM and 1.3~4.4x in terms of SpGEMM when dealing with moderately sparse matrices (sparsity < 0.9), while its performance is at least 9.7x better than cuSPARSE. Furthermore, experimental results show a FLOPs reduction of roughly 3.4x in the neural network.
Minjin Tang, Mei Wen, Yasong Cao, Junzhong Shen, Jianchao Yang, Jiawei Fei, Yang Guo 0003, Sheng Liu 0001
ICPP3
2022 S-SIM: A Simulator for Systolic Array-based DNN Accelerators with Tile Access Awareness
abstract
As NN accelerators emerging, many analytical models are presented to help designers to carry out hardware design space exploration. However, these models cannot accurately simulate the systolic array-based NN accelerator due to their pervasiveness or abstraction. In this paper, we propose a compute-centric simulator driven by the execution of events from the tiles of the mapping matrix, which can accurately model the systolic array-based accelerator. The simulator focuses on the conflicts when the tile is used for data access, or various interruptions caused by hardware resource limitations. Experimental results show that the proposed simulator achieves more than 95% accuracy compared to the real scenes.
Mei Wen, Renyu Yang, Junzhong Shen, Yasong Cao
ISCAS5
2022 TILE-SIM: A Systematic Approach to Systolic Array-based Accelerator Evaluation
abstract
The systolic array provides extremely high efficiency for running matrix multiplication, and is one of the mainstream architectures of today’s deep learning accelerators. In order to develop efficient accelerators, people usually employ simulators to make design trade-offs. However, current simulators suffer from coarse-grained modeling methods and ideal assumptions, which limits their ability of describing structural characteristics of systolic arrays. In addition, they do not support the exploration of microarchitecture. This paper presents TILE-SIM, a computing-centric systematic method for evaluating systolic array accelerators by using an event-driven method. TILE-SIM can obtain accurate results and provide the best mapping scheme for different workload due to its fine-grained modeling technique and deny of ideal assumption. Experimental results show that TILE-SIM plays a significant role in design trade-offs and outperforms state-of-the-art simulators, with an accuracy of more than 95%.
Mei Wen, Jiawei Fei, Junzhong Shen, Yasong Cao
ISPASS5