Junzhong Shen

dblp:171/0935 · DBLP profile ↗
← Back
35ranked-venue papers
6as first author
22since 2021 · last 2026
0000-0001-6233-6800ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 33 · 6 first-author · 20 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Uni-STC: Unified Sparse Tensor Core
abstract
Modern processors are increasingly adopting tensor cores as key computational units. Compared to existing designs for dense and structured sparsity, recent dual-side sparse tensor cores have evolved to support general sparsity. However, existing methods still face limitations on generality (incomplete sparse kernel support prevents broad applicability) and performance (outer-product/row-row schemes yield unsatisfactory hardware utilisation, data reuse, and energy efficiency). In this paper, we propose Uni-STC, a unified sparse tensor core that delivers high-performance dataflows for four key sparse kernels: sparse matrix-vector multiplication (SpMV), sparse matrixsparse vector multiplication (SpMSpV), sparse matrix-multiple vector multiplication (SpMM), and sparse general matrix-matrix multiplication (SpGEMM). To efficiently support these diverse sparse workloads, we first introduce BBC, a unified sparse format co-designed with Uni-STC's dataflow. We then design UniSTC's architecture supporting (1) fine-grained task partitioning to improve resource utilisation, (2) parallel sparse-tile processing to enhance data reuse, and (3) a dynamic network to reduce intermediate data movement and energy consumption. Evaluated across 2893 SuiteSparse and 302 DLMC matrices, Uni-STC demonstrates significant improvements, outperforming the state-of-the-art RM-STC with a$2.21 \times$geomean speedup and$2.96 \times$higher energy efficiency.
Haocheng Lian, Meichen Dong, Yijie Nie, Junzhong Shen, Chun Huang 0006, Bingcai Sui, Weifeng Liu 0002
HPCA7
2026 C-CIM: A Multi-Mode Convolution-Capable SRAM-CIM
abstract
SRAM is widely used in computing-in-memory (CIM) neural network accelerators because of its relatively mature technology and good compatibility with complementary metal oxide semiconductor logic process. Digital SRAM-CIM is favored by researchers because of its stability and accuracy. However, the current digital SRAM-CIM macro only supports the weight-stationary dataflows, which means the repeated movement of graph data. Some special deep neural network layers, such as depth-wise, make the utilization of computing resources inside CIM low. To overcome these problems, we propose C-CIM, which can switch between input-stationary and weight-stationary dataflows and support matrix multiplication as well as convolution operations with multiple mainstream convolution kernel sizes (1×1, 3×3, 5×5 and 7×7). The C-CIM achieves an average performance of 27.31TOPS/W@8b at a frequency of 1GHz. Experimental results show that our proposed SRAM-CIM successfully outperforms baseline in terms of performance optimization, achieving up to 7.6× performance speedup and up to 86.84% reduction in activation relocation.
Renyu Yang, Xin Ju 0005, Mei Wen, Jinjin Deng, Junzhong Shen, Tianyu Wang 0003, Zhaoyan Shen, Zili Shao
ACM Trans. Design Autom. Electr. Syst.6
2025 WinAcc: Window-based Acceleration of Neural Networks Using Block Floating Point
abstract
Deep Neural Networks (DNNs) impose significant computational demands, necessitating optimizations for computational and energy efficiencies. Per-vector scaling, which applies a scaling factor to blocks of elements using narrow integer types, effectively reduces storage and computational overhead. However, the frequent occurrence of floating-point accumulations between vectors limits further improvements in energy efficiency. State-of-the-art accelerators address this challenge by grouping and summing vector products based on their exponent differences, thereby reducing the overhead associated with intra-group shifting and accumulation. Nevertheless, this approach increases the complexity of register usage and grouping logic, leading to limited energy benefits and hardware efficiency. In this context, we introduce WinAcc, a novel algorithm and architecture co-designed solution that utilizes a low-cost accumu-lator to handle the majority of data in DNNs, offering low area overhead and high energy efficiency gains. Our key insight is that the data of DNNs follows a Laplace-like distribution, which enables the use of a customized data format with a narrow dynamic range to encode most of the data. This allows for the design of a low-cost accumulator with narrow shifters and adders, significantly reducing reliance on floating-point accumulator and consequently improving energy efficiency. Compared with state-of-the-art architecture Bucket, WinAcc achieves 33.95% energy reduction across seven representative DNNs and reduces area by 9.5% while maintaining superior model performance.
Xin Ju 0005, Mei Wen, Yasong Cao, Junzhong Shen, Zhaoyun Chen, Yang Shi 0008
DATE6
2025 SparSynergy: Unlocking Flexible and Efficient DNN Acceleration Through Multi-Level Sparsity
abstract
To more effectively address the computational and memory requirements of deep neural networks (DNNs), leveraging multi-level sparsity-including value-level and bit-level sparsity-has emerged as a pivotal strategy. While substantial research has been dedicated to exploring value-level and bit-level sparsity individually, the combination of both has largely been overlooked until now. In this paper, we propose SparSynergy, which-to the best of our knowledge-is the first accelerator that synergistically integrates multi-level sparsity into a unified framework, maximizing computational efficiency and minimizing memory usage. However, jointly considering multi-level sparsity is non-trivial, as it presents several challenges: (1) increased hardware overhead due to the complexity of incorporating multiple sparsity levels, (2) bandwidth-intensive data transmission during multiplexing, and (3) decreased throughput and scalability caused by bottlenecks in bit-serial computation. Our proposed SparSynergy addresses these challenges by introducing a unified sparsity format and a cooptimized hardware design. Experimental results demonstrate that SparSynergy achieves a 5.38 x geometric mean improvement in the energy-delay product (EDP) when compared with the tensor core, across workloads with varying degrees of sparsity. Furthermore, SparSynergy significantly improves accuracy retention compared to state-of-the-art accelerators for representative DNNs.
Jingkui Yang, Mei Wen, Junzhong Shen, Jianchao Yang, Yasong Cao, Minjin Tang, Zhaoyun Chen, Yang Shi 0008
DATE3
2025 Initial-Key Cache: An Efficient KV Cache Strategy Focusing on Initial and Key Tokens for LLMs
abstract
Large language models (LLMs) have developed rapidly in recent years and have demonstrated excellent performance in various application fields. Despite their outstanding performance, LLMs also introduce significant challenges in the practical inference process, mainly because of their computational and memory-intensive characteristics. Due to the autoregressive nature of the attention mechanism, KV caching can effectively accelerate the inference of LLMs by substituting quadratic-complexity computation with linear-complexity memory accesses. However, during the calculation, it is necessary to transfer the KV cache value to the computing unit, which not only requires a larger storage capacity but also places higher demands on storage bandwidth. In the process of inference and text generation for long texts, storage has become a bottleneck that limits the performance of the inference. In this paper, we first noticed the attention sink phenomenon that high attention scores are allocated towards initial tokens as "attention sink" even if they are not semantically significant in short input sentences and further observed that the attention sink tends to diminish as the length of the input sentence increases. Based on the above preliminary empirical results, we proposed Initial-key Cache, an efficient KV cache selection algorithm that not only effectively addresses the disappearance of the sink phenomenon in long inputs but also comprehensively considers the importance of middle-sentence tokens. We conducted a series of experiments on four baseline models, namely LLaMA, Qwen, Pythia, and OPT, to assess the performance of the Initial-Key Cache. The experiment results indicate that the Initial-Key Cache saves KV cache memory usage, while almost not losing the model’s accuracy.
Zhongyi Tang, Zejiang He, Junzhong Shen, Yiyue Hu, Luchen Zhou, Yongzhang Nie, Yongwen Wang
IJCNN3
2025 CAMO: A High-Performance CIM-based Lightweight CNN Accelerator for Mobile Devices
abstract
Digital Compute-in-Memory (CIM) macros revolutionize the Von Neumann architecture by significantly reducing data movements between CPU and memory. However, when dealing with lightweight CNNs with various convolution types, existing GEMM (general matrix multiplication)-oriented solutions suffer from underutilization and large activation traffic, leading to unsatisfied energy and area efficiencies. In this context, we propose CAMO, in which the key contributions are: (1) A novel convolution mapping mechanism suitable for CIM macros, that maximizes data reuse and reduces activation traffic. (2) A convolution-capable CIM macro, that also supports small-scale GEMM. (3) A CIM-based architecture that supports multiple computing modes. The experimental results show that CAMO achieves up to 31.49× performance speedup and 77.5% activation traffic reduction compared to the baseline architecture.
Xin Ju 0005, Renyu Yang, Mei Wen, Junzhong Shen, Tianyu Wang 0009, Zhaoyan Shen, Zili Shao
ISCAS4
2025 MAP-SIM: A DNN-Specific Mapping Optimization Framework for Shared-Memory CPU-Systolic Array Architectures
abstract
As performance demands continue to rise, Shared-Memory Heterogeneous Systems (SMHSs) have been widely adopted for their ability to enable efficient communication and data sharing between different heterogeneous cores. However, existing SMHS face challenges in uneven workload distribution among heterogeneous cores and suboptimal mapping schemes, preventing them from fully leveraging their architectural advantages. To address these issues, this paper proposes a mapping-aware framework for modeling SMHSs called MAP-SIM. By performing performance modeling for CPUs and Systolic Arrays (SAs), and considering rational schemes for the partition and mapping of computational tasks, MAP-SIM aims to evaluate and optimize the computational performance of heterogeneous multicore architectures. The experimental results show that compared to previous work, MAP-SIM can increase simulation speed by 14 to 67 times and can also enhance the computational performance of SMHS by 1.4 to 4.4 times.
Mei Wen, Junzhong Shen, Zhaoyun Chen, Yang Shi 0008, Tianyu Wang 0009, Zili Shao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2025 SPSA: Exploring Sparse-Packing Computation on Systolic Arrays From Scratch
abstract
Sparse matrix-matrix multiplication (SpMM) and Generalized SpMM (SpGEMM) are essential computational kernels in domains, such as graph analytics and scientific computation. While systolic arrays have traditionally been employed as specialized architectures for complex computing problems like matrix multiplication, they exhibit inefficiency when dealing with sparse matrices. This inefficiency arises from the unnecessary operations performed by processing elements (PEs) that contain zero-valued entries, which do not contribute to the final result. To address this issue, we propose SPSA, a framework that leverages a sparse-packing algorithm suitable for systolic arrays to accelerate sparse matrix computations. Our approach achieves significant reduction of zero-valued items and improves matrix density by packing the rows or columns of the sparse matrix. Furthermore, we have introduced for the first time a data representation format tailored to systolic arrays, called CSXD, which further enhances storage and computational efficiency. Importantly, our adaptation scheme enables acceleration benefits even with limited resources. Through sparse packing, SPSA achieved a$5.2\times $performance improvement compared to the dense baseline, and further reached a$6.4\times $enhancement via CSXD. Simultaneously, CSXD realized an average storage efficiency improvement of$15.0\times $. Through extensive evaluations, SPSA outperforms previous designs on CPU, GPU, and ASIC platforms. Finally, in end-to-end evaluations, SPSA achieved a performance improvement of 3.9 times across the workloads of BERT, VGG19, and ResNet50.
Minjin Tang, Mei Wen, Jianchao Yang, Zeyu Xue, Junzhong Shen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 ISOAcc: In-situ Shift Operation-based Accelerator For Efficient in-SRAM Multiplication
abstract
Digital SRAM-based CIM architectures must balance three critical factors: quantized neural network bitwidth, accuracy loss, and computational efficiency, each crucial to optimizing performance and efficiency. In Domain Specific Accelerators (DSAs), flexible and specific hardware design, when incorporated with tailored Power-of-2 (P-2) quantization schemes, addresses this issue. However, in CIMs, the absence of flexible and specific hardware to support dynamic switching between general and tailored quantization schemes hinders the adoption of efficient quantization methods. In this article, we propose the I n-situ S hift O peration based Acc elerator ( ISOAcc ) for efficient SRAM-based multiplication. The key idea is to introduce transmission gates near the SRAM array to enable the selection of bits from either the same or the neighbor line when data flows from one row to another. This functionally equals a shift operation. By configuring the transmission gates array in a cascade manner, ISOAcc can support 0 to 15-bit shift with a negligible overhead. The ISOAcc can directly leverage P-2 quantization schemes in hardware, thereby greatly reducing multiplication cycles. We have chosen five well-known neural networks to evaluate ISOAcc. The evaluations show that ISOAcc achieves an average performance improvement of 3.24× and an energy reduction of 75%, compared with the state-of-the-art (SOTA) SRAM-based CIM design, Bit-Parallel.
Gaoyang Zhao, Junzhong Shen, Rongzhen Lin
ACM Trans. Design Autom. Electr. Syst.2
2025 SSpMM: Efficiently Scalable SpMM Kernels Across Multiple Generations of Tensor Cores
abstract
Sparse-Dense Matrix-Matrix Multiplication (SpMM) has emerged as a foundational primitive in HPC and AI. Recent advancements have aimed to accelerate SpMM by harnessing the powerful Tensor Cores found in modern GPUs. However, despite these efforts, existing methods frequently encounter performance degradation when ported across different Tensor Core architectures. Recognizing that scalable SpMM across multiple generations of Tensor Cores relies on the effective use of general-purpose instructions, we have meticulously developed a SpMM library named SSpMM. However, a significant conflict exists between granularity and performance in current Tensor Core instructions. To resolve this, we introduce the innovative Transpose Mapping Scheme, which elegantly implements fine-grained kernels using coarse-grained instructions. Additionally, we propose the Register Shuffle Method to further enhance performance. Finally, we introduce Sparse Vector Compression, a technique that ensures our kernels are scalable with both structured and unstructured sparsity. Our experimental results, conducted on four generations of Tensor Core GPUs using over 3,000 sparse matrices from well established matrix collections, demonstrate that SSpMM achieves an average speedup of 2.04×, 2.81×, 2.07×, and 1.87×, respectively, over the state-of-the-art SpMM solution. Furthermore, we have integrated SSpMM into PyTorch, achieving a 1.81× speedup in end-to-end Transformer inference compared to cuDNN.
Zeyu Xue, Mei Wen, Jianchao Yang, Minjin Tang, Zhongdi Luo, Yang Shi 0008, Zhaoyun Chen, Junzhong Shen, Johannes Langguth
IEEE Trans. Parallel Distributed Syst.9
2025 FAMS: A FrAmework of Memory-Centric Mapping for DNNs on Systolic Array Accelerators
abstract
In recent years, deep neural networks (DNNs) have experienced rapid development. These DNNs demonstrate significant variations in architecture and scale, creating a substantial demand for domain-specific accelerators that are optimized for both high performance and low energy consumption. Systolic array accelerators, due to their efficient dataflow and parallel processing capabilities, offer significant advantages when performing computations for DNNs. Existing studies frequently overlook various hardware constraints in systolic array accelerators when representing mapping strategies. This oversight includes ignoring the differences in delays between communication and computation operations, as well as overlooking the capacities of multilevel memory hierarchies. Such omissions can lead to inaccuracies in predicting accelerator performance and inefficiencies in system design. We propose the FAMS framework, which introduces a memory-centric notation capable of fully representing the mapping of DNN operations on systolic array accelerators. Memory-centric notation moves away from the idealized assumptions of previous notations and considers various hardware constraints, thereby expanding the effective design and mapping spaces. The FAMS framework also includes a cycle-accurate simulator, which takes the hardware configurations, task descriptions, and mapping strategy represented by memory-centric notation as inputs, providing various metrics such as latency and energy consumption. The experimental results demonstrate that our proposed FAMS framework reduces latency by up to 29.7% and increases throughput by 42.4% compared to the state-of-the-art TENET framework. Additionally, under hardware configurations with a MAC delay of 2 and 3 clock cycles, the FAMS framework enhances performance by 12.0% and 25.4%, respectively.
Hao Sun 0023, Junzhong Shen, Zhongyi Tang, Changwu Zhang, Yang Shi 0008, Hengzhu Liu
IEEE Trans. Very Large Scale Integr. Syst.2
2024 BitShare: An Efficient Precision-Scalable Accelerator with Combining-Like-Terms GEMM
abstract
Narrow-precision fixed-point (INT) computation is a significant approach for reducing memory requirements and enhancing the performance of accelerators for Deep Neural Networks (DNNs). Different DNNs, as well as different layers within the DNNs, may exhibit varying numerical distributions, necessitating INT formats with different minimum bit-widths. Therefore, DNN accelerators need to support multi-precision INT computations to strike a better balance between DNN inference accuracy and performance. However, existing precision-scalable accelerators face challenges such as low bandwidth utilization, insufficient utilization of computing resources across different precision modes, and complex circuit structures with associated overhead. In this paper, we propose (1) a hardware-friendly Combining-Like-Terms GEMM (CLT-GEMM) scheme that supports multiple computing modes of 2/4/8 bits and their combinations to align with the various bit-width settings of DNNs; (2) and subsequently design an efficient systolic accelerator with scalable precision, named BitShare, which features DataMap module and Multi-mode adder-tree-based accumulators. Compared to the state-of-the-art precision-scalable design, BitBlade, our accelerator achieves a 57.25% reduction in bandwidth requirement and exhibits an improvement of$1.14\times$and$1.12\times$in area and power efficiency$(2\mathbf{b}\times 2\mathbf{b})$, respectively.
Yasong Cao, Mei Wen, Junzhong Shen, Zhongxing Li
ASAP3
2024 MACO: Exploring GEMM Acceleration on a Loosely-Coupled Multi-Core Processor
abstract
General-purpose processor vendors have integrated customized accelerator in their products due to the widespread use of General Matrix-Matrix Multiplication (GEMM) kernels. However, it remains a challenge to further improve the flexibility and scalability of these GEMM-enhanced processors to cater to the emerging large-scale GEMM workloads. In this paper we propose MACO, a novel loosely-coupled multi-core general-purpose archi-tecture optimized for GEMM-related applications. To enhance the programmability and flexibility of MACO, the paper introduces a tile-based instruction set architecture. Additionally, the paper presents techniques such as hardware-assisted data prefetching and locking, and predictive address translation to further enhance the computational efficiency of MACO for GEMM workloads. The experimental results demonstrate that MACO exhibits good scalability, achieving an average computational efficiency of 90 % across multiple cores. Furthermore, evaluations on state-of-the-art deep neural networks show that MACO can achieve up to 1.1 TFLOPS with 88 % computational efficiency, indicating its adaptivity to deep learning workloads.
Bingcai Sui, Junzhong Shen, Caixia Sun
DATE2
2024 MAP-SIM: A Performance Model for Shared-Memory Heterogeneous Systems with Mapping Awareness
Mei Wen, Junzhong Shen, Zhaoyun Chen, Yang Shi 0008
ICA3PP (1)3
2024 MSA2: An Efficient Sparsity-Aware Accelerator for Matrix Multiplication with Multi-core Systolic Arrays
Minjin Tang, Mei Wen, Junzhong Shen, Jingkui Yang, Zeyu Xue, Zili Shao
ICA3PP (3)3
2024 Enhancing the PE Utilization for Multi-Precision Systolic Array via Optimizing Computation Latency
abstract
Systolic array (SA) architectures are widely recognized as the optimal choice for Convolutional Neural Networks (CNNs). However, existing SAs suffer from reduced computational efficiency when confronted with an inadequate workload scale. Furthermore, this issue becomes even more pronounced in the accelerators that support multiple precisions. In this paper, by analyzing the under-utilization of processing element (PE), we propose a SA accelerator that optimizes computation latency for multi-precision scenarios. Considering dynamic changes in data precision, we incorporate a switching strategy to further enhance computational efficiency. Experimental results demonstrate that our proposed design only incurs a 1.203% increase in area compared to the classic approach, while achieving an average performance improvement of 20% on small-scale CNN models.
Mei Wen, Xin Ju 0005, Junzhong Shen, Yang Guo 0003
ISCAS4
2024 HyFiSS: A Hybrid Fidelity Stall-Aware Simulator for GPGPUs
abstract
The widespread adoption of GPUs has driven the development of GPU simulators, which, in turn, lead advancements in both GPU architectures and software optimization. Trace-driven cycle-accurate Cycle-accurate simulators, which provide detailed microarchitectural models and clock-level precision, come at the cost of extended simulation times and require high computational resources. Their scalability has become a bottleneck. A growing trend is the adoption of cycle-approximate simulators, which introduce mathematical modeling of partial hardware units and utilize sampling to accelerate simulation. However, this approach faces challenges regarding the accuracy of performance predictions. To address these limitations, we introduce HyFiSS, a hybrid fidelity stall-aware GPU simulator. HyFiSS features fine-grained stall events tracking and attribution by constructing a detailed execution pipeline model for various stall events on Streaming Multiprocessors (SMs). It accurately emulates the thread block scheduler behavior using real-time scheduling logs and utilizes sampling based on thread block sets to minimize the precision loss due to fine-grained sampling points on the microarchitectural state. We achieve a balance between reliability, speed, and the level of simulation detail, especially regarding bottlenecks. By evaluating a diverse set of benchmarks, HyFiSS achieves a mean absolute percentage error in predicting active cycles that is comparable to the state-of-the-art cycle-accurate simulator Accel-Sim. Moreover, HyFiSS achieves a substantial 12.8 × speedup in the simulation efficiency compared to Accel-Sim. HyFiSS also requires at least 3.2 × less disk storage than both Accel-Sim and another state-of-the-art cycle-approximate simulator PPT-GPU due to its efficient SASS (Streaming Assembler) traces compression. With precise, per-cycle stall events statistics, HyFiSS can provide accurate GPU performance metrics and stall cause reporting. This significantly simplifies performance analysis, bottleneck identification, and performance optimization tasks for researchers, making it easier to enhance GPU performance effectively.
Jianchao Yang, Mei Wen, Dong Chen 0015, Zhaoyun Chen, Zeyu Xue, Junzhong Shen, Yang Shi 0008
MICRO7
2024 ABS: Accumulation Bit-Width Scaling Method for Designing Low-Precision Tensor Core
abstract
A big gap exists between deep neural network (DNN) applications’ computational demand and the computing power of DNN accelerators. Low-precision floating-point (LP-FP) computation is one of the important means to improve the performance of DNN training and inference. However, the high-precision accumulators are typically applied to summating the dot products during general matrix multiplication (GEMM) in tensor cores (TCs). As the precision of data decreases, the accumulator becomes the main consumer of multiply-accumulate’s (MAC’s) area and power. Reducing the accumulators’ bit-width is of significant importance for improving the area- and energy-efficiency of TCs. There are two main challenges: 1) theoretical support on the floating-point (FP) formats with the lowest bit-width of TC’s accumulators and 2) how to integrate the LP-FP TC in the framework of DNN training and inference to evaluate its benefits. In this article, we propose accumulation bit-width scaling (ABS), a novel ABS method, to guide the design of LP-FP TCs. We 1) implement this method by constructing a novel variance retention ratio (VRR) model to predict the FP format with the minimum bit-width for TC’s accumulator; 2) provide a generator of DNN accelerator based on a systolic-array (SA) TC, supporting many low-precision configurations; and 3) design an LP-FP DNN executing framework that supports software-simulation mode and hardware-accelerator mode to run LP-FP DNN tasks. The experimental results show that the LP-FP TC guided by our ABS method has a maximum reduction of 76.47% and 75.60% in area and power consumption, respectively, compared with the advanced TCs.
Yasong Cao, Mei Wen, Zhongdi Luo, Xin Ju 0005, Haolan Huang, Junzhong Shen
IEEE Trans. Very Large Scale Integr. Syst.6
2022 BP-Im2col: Implicit Im2col Supporting AI Backpropagation on Systolic Arrays
abstract
State-of-the-art systolic array-based accelerators adopt the traditional im2col algorithm to accelerate the inference of convolutional layers. However, traditional im2col cannot efficiently support AI backpropagation. Backpropagation in convolutional layers involves performing transposed convolution and dilated convolution, which usually introduces plenty of zero-spaces into the feature map or kernel. The zero-space data reorganization interfere with the continuity of training and incur additional and non-negligible overhead in terms of off- and on-chip storage, access and performance. Since countermeasures for backpropagation are rarely proposed, we propose BP-im2col, a novel im2col algorithm for AI backpropagation, and implement it in RTL on a TPU-like accelerator. Experiments on TPU-like accelerator indicate that BP-im2col reduces the backpropagation runtime by 34.9% on average, and reduces the bandwidth of off-chip memory and on-chip buffers by at least 22.7% and 70.6% respectively, over a baseline accelerator adopting the traditional im2col. It further reduces the additional storage overhead in the backpropagation process by at least 74.78%.
Jianchao Yang, Mei Wen, Junzhong Shen, Yasong Cao, Minjin Tang, Renyu Yang, Jiawei Fei, Chunyuan Zhang
ICCD3
2022 Mentha: Enabling Sparse-Packing Computation on Systolic Arrays
abstract
Generalized Sparse Matrix-Matrix Multiplication (SpGEMM) is a critical kernel in domains like graph analytic and scientific computation. As a kind of classical special-purpose architecture, systolic arrays were first used for complex computing problems, e.g., matrix multiplication. However, classical systolic arrays are not efficient enough when handling sparse matrices due to the fact that the PEs containing zero-valued entries perform unnecessary operations that do not contribute to the result. Accordingly, in this paper, we propose Mentha, a framework that enables systolic arrays to accelerate sparse matrix computation by employing a sparse-packing algorithm suitable for various dataflow of systolic array. Firstly, Mentha supports both online and offline methods. By packing the rows or columns of the sparse matrix, the zero-valued items in the matrix are significantly reduced and the density of the matrix is improved. In addition, acceleration benefits can be obtained by the adaptation scheme even with limited resources. Moreover, we reconfigure PEs in systolic arrays at a low cost (1.28x in area, 1.21x in power) and find that our method outperforms TPU-like systolic arrays by 1.2~3.3x in terms of SpMM and 1.3~4.4x in terms of SpGEMM when dealing with moderately sparse matrices (sparsity < 0.9), while its performance is at least 9.7x better than cuSPARSE. Furthermore, experimental results show a FLOPs reduction of roughly 3.4x in the neural network.
Minjin Tang, Mei Wen, Yasong Cao, Junzhong Shen, Jianchao Yang, Jiawei Fei, Yang Guo 0003, Sheng Liu 0001
ICPP4
2022 S-SIM: A Simulator for Systolic Array-based DNN Accelerators with Tile Access Awareness
abstract
As NN accelerators emerging, many analytical models are presented to help designers to carry out hardware design space exploration. However, these models cannot accurately simulate the systolic array-based NN accelerator due to their pervasiveness or abstraction. In this paper, we propose a compute-centric simulator driven by the execution of events from the tiles of the mapping matrix, which can accurately model the systolic array-based accelerator. The simulator focuses on the conflicts when the tile is used for data access, or various interruptions caused by hardware resource limitations. Experimental results show that the proposed simulator achieves more than 95% accuracy compared to the real scenes.
Mei Wen, Renyu Yang, Junzhong Shen, Yasong Cao
ISCAS4
2022 TILE-SIM: A Systematic Approach to Systolic Array-based Accelerator Evaluation
abstract
The systolic array provides extremely high efficiency for running matrix multiplication, and is one of the mainstream architectures of today’s deep learning accelerators. In order to develop efficient accelerators, people usually employ simulators to make design trade-offs. However, current simulators suffer from coarse-grained modeling methods and ideal assumptions, which limits their ability of describing structural characteristics of systolic arrays. In addition, they do not support the exploration of microarchitecture. This paper presents TILE-SIM, a computing-centric systematic method for evaluating systolic array accelerators by using an event-driven method. TILE-SIM can obtain accurate results and provide the best mapping scheme for different workload due to its fine-grained modeling technique and deny of ideal assumption. Experimental results show that TILE-SIM plays a significant role in design trade-offs and outperforms state-of-the-art simulators, with an accuracy of more than 95%.
Mei Wen, Jiawei Fei, Junzhong Shen, Yasong Cao
ISPASS4
2020 Towards Memory-Efficient Streaming Processing with Counter-Cascading Sketching on FPGA
abstract
Obtaining item frequencies in data streams with limited space is a well-recognized and challenging problem in a wide range of applications. Sketch-based solutions have been widely used to address this challenge due to their ability to accurately record the data streams at a low memory cost. However, most sketches suffer from low memory utilization due to the adoption of a fixed counter size. Accordingly, in this work, we propose a counter-cascading scheduling algorithm to maximize the memory utilization of sketches without incurring any accuracy loss. In addition, we propose an FPGA-based system design that supports sketch parameter learning, counter-cascading record and online query. We implement our designs on Xilinx VCU118, and conduct evaluations on real-world traces, thereby demonstrating that our design can achieve higher accuracy with lower storage; the performance achieved is 10× ~ 20× better than that of state-of-the-art sketches.
Minjin Tang, Mei Wen, Junzhong Shen, Chunyuan Zhang
DAC3
2020 Scalable FPGA-based Architecture for High-Performance Per-Flow Traffic Measurement
abstract
Per-flow traffic measurement has emerged as a critical but challenging task in data center in recent years in the face of massive network traffic. Many approximate methods have been proposed to resolve the existing resource-accuracy trade-off in per-flow traffic measurement, one of which is the sketch-based method. However, sketches are affected by their high computational cost and low throughput; moreover, their measurement accuracy is hard to guarantee under the conditions of changing network bandwidth or flow size distribution. Recently, FPGA platforms have been widely deployed in data centers, as they demonstrate a good fit for high-speed network processing. In this work, we propose a scalable pipelined architecture for high high-throughput per-flow traffic measurement on FPGA. We adopts memory-friendly D-left hashing in our design, which guarantees high space utilization that successfully addressing the challenge of tracking high speed data stream under limit memory resource on FPGA. Comparisons with state-of-the-art sketch-based solutions show that our design outperforms state-of-the-art sketch-based methods in terms of throughput by over 80x.
Junzhong Shen, Mei Wen, Minjin Tang, Chunyuan Zhang
FPGA1
2020 Towards a Deep-Pipelined Architecture for Accelerating Deep GCN on a Multi-FPGA Platform
Qixuan Cheng, Mei Wen, Junzhong Shen, Chunyuan Zhang
ICA3PP (1)3
2020 Toward an Efficient Deep Pipelined Template-Based Architecture for Accelerating the Entire 2-D and 3-D CNNs on FPGA
abstract
3-D convolutional neural networks (3-D CNNs) are used efficiently in many computer vision applications. Most previous work in this area has concentrated only on design and optimization of accelerators for 2-D CNNs, with few attempts having been made to accelerate 3-D CNNs on FPGA. We find the acceleration of 3-D CNNs on FPGA to be challenging due to their high computational complexity and storage demands. More importantly, although the computational patterns of 2-D and 3-D CNNs are analogous, the conventional approaches that have been adopted for acceleration of 2-D CNNs may be unfit for 3-D CNN acceleration. In this paper, in order to accelerate 2-D and 3-D CNNs using a uniform framework, we first propose a uniform template-based architecture that uses templates based on the Winograd algorithm to ensure the rapid development of 2-D and 3-D CNN accelerators. Then, with the aim of efficiently mapping all layers of 2-D/3-D CNNs onto a pipelined accelerator, techniques are developed to improve the throughput and computational efficiency of the accelerator, including layer fusion, layer clustering, and workload-balancing scheme. Finally, we demonstrate the effectiveness of the deep pipelined architecture by accelerating real-life 2-D and 3-D CNNs on the state-of-the-art FPGA platform. On VCU118, we achieve 3.7 TOPS for VGG-16, which outperforms state-of-the-art FPGA-based CNN accelerators. Comparisons with CPU and GPU solutions demonstrate that our implementation of 3-D CNN achieves gains of up to 17.8× and 64.2× in performance and energy relative to a CPU solution, and a 5.0× energy efficiency gain over a GPU solution.
Junzhong Shen, You Huang, Mei Wen, Chunyuan Zhang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 Scale-out Acceleration for 3D CNN-based Lung Nodule Segmentation on a Multi-FPGA System
abstract
Three-dimensional convolutional neural networks (3D CNNs) have become a promising method in lung nodule segmentation. The high computational complexity and memory requirements of 3D CNNs make it challenging to accelerate 3D CNNs on a single FPGA. In this work, we focus on accelerating the 3D CNN-based lung nodule segmentation on a multi-FPGA platform by proposing an efficient mapping scheme that takes advantage of the massive parallelism provided by the platform, as well as maximizing the computational efficiency of the accelerators. Experimental results show that our system integrating with four Xilinx VCU118 can achieve state-of-the-art performance of 14.5 TOPS, in addition with a 29.4x performance gain over CPU and 10.5x more energy efficiency over GPU.
Junzhong Shen, You Huang, Mei Wen, Chunyuan Zhang
DAC1
2019 Accelerating 3D CNN-based Lung Nodule Segmentation on a Multi-FPGA System
abstract
Lung nodule segmentation is one of the most significant steps in many Computer Aided Detection (CAD) systems used for lung nodule identification and classification. Three-dimensional convolutional neural networks (3D CNNs) have become a promising method in lung nodule segmentation, as this method can achieve higher detection accuracy than conventional methods. It has been proven that FPGAs can provide the most energy-efficient solution for CNN acceleration. However, the high computational complexity and memory requirements of 3D CNNs make it challenging to accelerate 3D CNNs on a single FPGA, as this will further bottleneck the performance of a 3D CNN-based CAD system. Accordingly, in this work, we focus on accelerating the 3D CNN-based lung nodule segmentation on a multi-FPGA platform by proposing an efficient mapping scheme that takes advantage of the massive parallelism provided by the platform, as well as maximizing the computational efficiency of the accelerators. Experimental results show that our system is able to achieve high computational efficiency and thereby a state-of-the-art performance of 14.5 TOPS at 200 MHz. Comparisons with CPU and GPU solutions demonstrate that our system achieves a 29.4x performance gain over CPU and a 10.5x energy efficiency improvement over GPU.
Junzhong Shen, You Huang, Mei Wen, Chunyuan Zhang
FPGA1
2019 An Efficient Design Flow for Accelerating Complicated-connected CNNs on a Multi-FPGA Platform
abstract
Convolutional Neural Networks (CNNs) have achieved impressive performance on various computer vision tasks. To facilitate better performance, some complicated-connected CNN models (e.g., GoogLeNet and DenseNet) have recently been proposed, and have achieved state-of-the-art performance in the fields of image classification and segmentation. However, CNNs are computation- and memory-intensive. Thus, it is significant to develop hardware accelerators in order to accelerate the inference and training processes of CNNs. Due to the high-performance, reconfigurable and energy-efficient nature of Field-Programmable Gate Arrays (FPGAs), many FPGA-based accelerators have been proposed to implement CNNs and have achieved higher throughput and energy efficiency. However, the large number of parameters involved in complicated-connected CNN models have exceeded the limited hardware resources of single FPGA board, which are unable to meet the memory and computation resource demands associated with mapping entire CNN models. Accordingly, in this paper, we propose a complete design flow to accelerate the inference of complicated-connected CNNs on a multi-FPGA platform, including DAG abstraction, mapping scheme generation and design space exploration. In addition, a multi-FPGA system with flexible inter-FPGA communications is proposed to efficiently support our design flow. Experimental results on representative models illustrate that the proposed multi-FPGA system design can achieve a throughput acceleration of up to 145.2× and 2.5× compared to CPU and GPU solutions, as well as an energy efficiency improvement of up to 139.1× and 4.8× compared to multi-core CPU and GPU solutions.
Junzhong Shen, Mei Wen, Chunyuan Zhang
ICPP2
2019 Towards a Uniform Architecture for the Efficient Implementation of 2D and 3D Deconvolutional Neural Networks on FPGAs
abstract
Three-dimensional deconvolution is widely used in many computer vision applications. However, most previous works have only focused on accelerating 2D deconvolutional neural networks (DCNNs) on FPGAs, while the acceleration of 3D DCNNs has not been studied in depth as they have higher computational complexity and sparsity than 2D DCNNs. In this paper, we focus on the acceleration of both 2D and 3D DCNNs on FPGAs by proposing efficient schemes for mapping 2D and 3D DCNNs on a uniform architecture. By implementing our design on the Xilinx VC709 platform for four real-life 2D and 3D DCNNs, we can achieve up to 3.0 TOPS with high hardware efficiency. Comparisons with CPU and GPU solutions demonstrate that we can achieve an improvement of up to 63.3 × in throughput relative to a CPU solution and an improvement of up to 8.3 × in energy efficiency compared to a GPU solution.
Junzhong Shen, Mei Wen, Chunyuan Zhang
ISCAS2
2018 Towards a Uniform Template-based Architecture for Accelerating 2D and 3D CNNs on FPGA
abstract
Three-dimensional convolutional neural networks (3D CNNs) are used efficiently in many computer vision applications. Most previous work in this area has concentrated only on designing and optimizing accelerators for 2D CNN, with few attempts made to accelerate 3D CNN on FPGA. We find accelerating 3D CNNs on FPGA to be challenge due to their high computational complexity and storage demands. More importantly, although the computation patterns of 2D and 3D CNNs are analogous, the conventional approaches adopted for accelerating 2D CNNs may be unfit for 3D CNN acceleration. In this paper, in order to accelerate 2D and 3D CNNs using a uniform framework, we propose a uniform template-based architecture that uses templates based on the Winograd algorithm to ensure fast development of 2D and 3D CNN accelerators. Furthermore, we also develop a uniform analytical model to facilitate efficient design space explorations of 2D and 3D CNN accelerators based on our architecture. Finally, we demonstrate the effectiveness of the template-based architecture by implementing accelerators for real-life 2D and 3D CNNs (VGG16 and C3D) on multiple FPGA platforms. On S2C VUS440, we achieve up to 1.13 TOPS and 1.11 TOPS under low resource utilization for VGG16 and C3D, respectively. End-to-end comparisons with CPU and GPU solutions demonstrate that our implementation of C3D achieves gains of up to 13x and 60x in performance and energy relative to a CPU solution, and a 6.4x energy efficiency gain over a GPU solution.
Junzhong Shen, You Huang, Yuran Qiao, Mei Wen, Chunyuan Zhang
FPGA1
2018 Towards a Multi-array Architecture for Accelerating Large-scale Matrix Multiplication on FPGAs
abstract
Large-scale floating-point matrix multiplication is a fundamental kernel in many scientific and engineering applications. Most existing work only focus on accelerating matrix multiplication on FPGA by adopting a linear systolic array. This paper towards the extension of this architecture by proposing a scalable and highly configurable multi-array architecture. In addition, we propose a work-stealing scheme to ensure the equality in the workload partition among multiple linear arrays. Furthermore, an analytical model is developed to determine the optimal design parameters. Experiments on a real-life convolutional neural network (CNN) show that we can obtain the optimal extension of the linear array architecture.
Junzhong Shen, Yuran Qiao, You Huang, Mei Wen, Chunyuan Zhang
ISCAS1
2017 Optimizing OpenCL Implementation of Deep Convolutional Neural Network on FPGA
Yuran Qiao, Junzhong Shen, Dafei Huang, Qianming Yang, Mei Wen, Chunyuan Zhang
NPC2
2017 FPGA-accelerated deep convolutional neural networks for high throughput and energy efficiency
abstract
Summary Recent breakthroughs in the deep convolutional neural networks (CNNs) have led to great improvements in the accuracy of both vision and auditory systems. Characterized by their deep structures and large numbers of parameters, deep CNNs challenge the computational performance of today. Hardware specialization in the form of field‐programmable gate array offers a promising path towards major leaps in computational performance while achieving high‐energy efficiency. In this paper, we focus on accelerating deep CNNs using the Xilinx Zynq‐zq7045 FPGA SoC. As most of the computational workload can be converted to matrix multiplications, we adopt a matrix multiplier‐based accelerator architecture. Dedicated units are designed to eliminate the conversion overhead. We also design a customized memory system according to the memory access pattern of CNNs. To make the accelerator easily usable by application developers, our accelerator supports Caffe, which is a widely used software framework of deep CNN. Different CNN models can be adopted by our accelerator, with good performance portability. The experimental results show that for a typical application of CNN, image classification, an average throughout of 77.8 GFLOPS is achieved, while the energy efficiency is 4.7× better than an Nvidia K20 GPGPU. © 2016 The Authors. Concurrency and Computation: Practice and Experience Published by John Wiley & Sons Ltd
Yuran Qiao, Junzhong Shen, Qianming Yang, Mei Wen, Chunyuan Zhang
Concurr. Comput. Pract. Exp.2
2015 Unified Virtual Memory Support for Deep CNN Accelerator on SoC FPGA
Yuran Qiao, Junzhong Shen, Qianming Yang, Mei Wen
ICA3PP (1)3