Zhengkun Gu

dblp:348/7574 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
0009-0005-3064-4469ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 HRAMTran: A Hybrid-RAM Transformer Accelerator With Dynamic Sparsity Floating-Point CIM and Written-Back Transpose Array
abstract
Transformer model performs outstandingly in various tasks involving artificial intelligence. In this work, we propose a hybrid-RAM Transformer accelerator (HRAMTran) utilizing computing-in-memory (CIM) based on spin-orbit torque magnetic random access memory (SOT-MRAM) and static RAM (SRAM), which supports dynamic sparsity in floating-point (FP) matrix multiplication (MM) and written-back transpose, thereby realizing efficient attention mechanism. First, a dynamic sparsity-based MM scheme is proposed, which dynamically ignores low-impact elements during vector multiplication, thereby effectively reducing the latency and energy consumption of MM. Second, a data-reuse multiply-and-accumulate (MAC) scheme for mantissa is designed to further optimize MM, which shares partial operation result to reduce redundant computation. The SOT-MRAM and SRAM based CIM architectures with dynamic sparsity and data-reuse schemes are constructed to perform weight (Query (Q), Key (K), and Value (V)) and dynamic MM, respectively. This hybrid-RAM CIM method can realize the optimization of energy and latency during attention mechanism computation. Moreover, written-back transpose SRAM array that can write multiple bits into a column simultaneously is designed to significantly reduce write-back cycles for KT. Finally, the HRAMTran accelerator is built to evaluate the performance of transformer implementation through performing machine translation for the WMT14 dataset. Results show that this accelerator realizes 3.6 µJ/Token and 68.77 TFLOPS/W, achieving 4.33× and 2.39× improvement compared with the state-of-the-art transformer accelerator.
Xianan Zhu, Zhengkun Gu, Zhizhong Zhang 0004, Kun Zhang 0030, Weisheng Zhao 0001, Yue Zhang 0010
ICCAD5
2025 A Heterogeneous System With Computing in Memory Processing Elements to Accelerate CNN Inference
abstract
Computing in memory (CIM) is one of the promising solutions to improve computing performance by integrating logic in memory. This work presents an efficient heterogeneous system based on the ultrafast CIM architecture (HS-CIM) to accelerate convolutional neural network (CNN) inference. First, an ultrafast CIM architecture is proposed based on the static random access memory (SRAM) by utilizing the novel total input and full digital scheme to implement multiply-and-accumulate (MAC) operation, which effectively addresses the high delay issue caused by high-precision computing in CIM architecture. Second, a heterogeneous system based on the proposed CIM architecture (HS-CIM) has been constructed with an aligned global cache and an adaptive pruning scheme to eliminate performance degradation and accuracy loss caused by input data bandwidth limitations. Meanwhile, efficient input data and weight data mapping schemes are proposed to minimize the delay and energy caused by input data transmission from the cache to the CIM architecture, thus realizing efficient CNN inference in the HS-CIM system. Finally, we analyze the performance of HS-CIM at the layout level by implementing the LeNet-5 and VGG models of CNN to recognize the image of MNIST and CIFAR-10 datasets, respectively. Results show that the energy efficiency of the proposed CIM architecture achieves 58.32 TOPS/W with 8-bit input/weight precision. Meanwhile, the inference accuracy of the HS-CIM system for MNIST and CIFAR-10 is 99.25% and 92.34%, respectively.
Youxiang Chen, Zhengkun Gu, Haiming Qiu, Kun Zhang 0030, Weisheng Zhao 0001, Yue Zhang 0010
IEEE Trans. Circuits Syst. I Regul. Pap.4
2024 FRM-CIM: Full-Digital Recursive MAC Computing in Memory System Based on MRAM for Neural Network Applications
abstract
Computing in memory (CIM) realizes energy-efficient neural network algorithms by implementing highly parallel multiply-and-accumulate (MAC) operation. However, the MAC delay of CIM will sharply increase with the improvement of computing precision, which restricts its development. In this work, we propose a full-digital recursive MAC (FRM) operation based on spin-transfer-torque magnetic random access memory (STT-MRAM) CIM system to enable fast and energy-efficient image recognition application. First, the fast FRM scheme is proposed by utilizing the recursive operations of read and addition in segmented bit-line array, which effectively reduces the delay of MAC operations to 3.5ns and 4ns for 8-bit and 16-bit input and weight precision, respectively. Second, we design an image recognition system using FRM-CIM architecture as the processing element (PE), where the adaptive pruning method for layers is proposed to improve the compatibility of it with the neural network. By performing image recognition for the MNIST and CIFAR-10 datasets, results show that the throughput and energy efficiency of the FRM-CIM system are 58.51TOPS/mm2 and 11.3--56.72 TOPS/W under 8--16-bit precision, which are improved by 4.3 times and 2.6 times compared with the state-of-the-art works. Finally, the recognition accuracy can reach 96.65% and 82.7% on MNIST and CIFAR-10, respectively.
Zhengkun Gu, Youxiang Chen, Weisheng Zhao 0001, Yue Zhang 0010
DAC4
2024 RSACIM: Resistance Summation Analog Computing in Memory With Accuracy Optimization Scheme Based on MRAM
abstract
Computing in memory (CIM) has become a promising candidate to address the Von Neumann bottleneck in processors designed for data-intensive applications. In this article, we propose a resistance summation analog computing in memory (RSACIM) with accuracy optimization scheme in spin transfer torque magnetic random access memory (STT-MRAM), in order to realize energy-efficient and highly reliable analog multiply-and-accumulation (MAC) operation. Firstly, we construct a resistance summation array by serial magnetic tunnel junctions (MTJs) to perform analog MAC operation utilizing time domain technology. Secondly, in order to reduce the impact of position-dependent error caused by resistance summation mechanism, we propose an accuracy optimization scheme to maximize the sensing margin (SM) and computation accuracy. Finally, we design a power-gated reconfigurability control scheme to implement power saving corresponding to different precisions for both input and weight. Evaluation on a 2 Kb RSACIM architecture shows an energy efficiency of 92.9 TOPS/W. System level simulation shows that comparing to existing CIMs based on MRAM, RSACIM architecture saves the inference energy by 4.2 times with 8.4 times lower latency in CIFAR10 image classification task.
Zhengkun Gu, Youxiang Chen, Kun Zhang 0030, Youguang Zhang, Yue Zhang 0010
IEEE Trans. Circuits Syst. I Regul. Pap.2
2023 TAM: A Computing in Memory based on Tandem Array within STT-MRAM for Energy-Efficient Analog MAC Operation
abstract
Computing in memory (CIM) has been demonstrated promising for energy efficient computing. However, the dramatic growth of the data scale in neural network processors has aroused a demand for CIM architecture of higher bit density, for which the spin transfer torque magnetic RAM (STT-MRAM) with high bit density and performance arises as an up-and-coming candidate solution. In this work, we propose an analog CIM scheme based on tandem array within STT-MRAM (TAM) to further improve energy efficiency while achieving high bit density. First, the resistance summation based analog MAC operation minimizes the effect of low tunnel magnetoresistance (TMR) by the serial magnetic tunnel junctions (MTJs) structure in the proposed tandem array with smaller area overhead. Moreover, a read scheme of resistive-to-binary is designed to achieve the MAC results accurately and reliably. Besides, the data-dependent error caused by MTJs in series has been eliminated with a proposed dynamic selection circuit. Simulation results of a 2Kb TAM architecture show 113.2 TOPS/W and 63.7 TOPS/W for 4-bit and 8-bit input/weight precision, respectively, and reduction by 39.3% for bit-cell area compared with existing array of MTJs in series.
Zhengkun Gu, Zuolei Hao, Weisheng Zhao 0001, Yue Zhang 0010
DATE2