EDBT 2026 Demo / reviewers in the wild / expert
Kun Zhang 0030
dblp:96/3115-30
· DBLP profile ↗
10ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0001-7215-7953ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HRAMTran: A Hybrid-RAM Transformer Accelerator With Dynamic Sparsity Floating-Point CIM and Written-Back Transpose ArrayabstractTransformer model performs outstandingly in various tasks involving artificial intelligence. In this work, we propose a hybrid-RAM Transformer accelerator (HRAMTran) utilizing computing-in-memory (CIM) based on spin-orbit torque magnetic random access memory (SOT-MRAM) and static RAM (SRAM), which supports dynamic sparsity in floating-point (FP) matrix multiplication (MM) and written-back transpose, thereby realizing efficient attention mechanism. First, a dynamic sparsity-based MM scheme is proposed, which dynamically ignores low-impact elements during vector multiplication, thereby effectively reducing the latency and energy consumption of MM. Second, a data-reuse multiply-and-accumulate (MAC) scheme for mantissa is designed to further optimize MM, which shares partial operation result to reduce redundant computation. The SOT-MRAM and SRAM based CIM architectures with dynamic sparsity and data-reuse schemes are constructed to perform weight (Query (Q), Key (K), and Value (V)) and dynamic MM, respectively. This hybrid-RAM CIM method can realize the optimization of energy and latency during attention mechanism computation. Moreover, written-back transpose SRAM array that can write multiple bits into a column simultaneously is designed to significantly reduce write-back cycles for KT. Finally, the HRAMTran accelerator is built to evaluate the performance of transformer implementation through performing machine translation for the WMT14 dataset. Results show that this accelerator realizes 3.6 µJ/Token and 68.77 TFLOPS/W, achieving 4.33× and 2.39× improvement compared with the state-of-the-art transformer accelerator. Xianan Zhu, Zhengkun Gu, Zhizhong Zhang 0004, Kun Zhang 0030, Weisheng Zhao 0001, Yue Zhang 0010 |
ICCAD | 8 |
| 2025 | A Heterogeneous System With Computing in Memory Processing Elements to Accelerate CNN InferenceabstractComputing in memory (CIM) is one of the promising solutions to improve computing performance by integrating logic in memory. This work presents an efficient heterogeneous system based on the ultrafast CIM architecture (HS-CIM) to accelerate convolutional neural network (CNN) inference. First, an ultrafast CIM architecture is proposed based on the static random access memory (SRAM) by utilizing the novel total input and full digital scheme to implement multiply-and-accumulate (MAC) operation, which effectively addresses the high delay issue caused by high-precision computing in CIM architecture. Second, a heterogeneous system based on the proposed CIM architecture (HS-CIM) has been constructed with an aligned global cache and an adaptive pruning scheme to eliminate performance degradation and accuracy loss caused by input data bandwidth limitations. Meanwhile, efficient input data and weight data mapping schemes are proposed to minimize the delay and energy caused by input data transmission from the cache to the CIM architecture, thus realizing efficient CNN inference in the HS-CIM system. Finally, we analyze the performance of HS-CIM at the layout level by implementing the LeNet-5 and VGG models of CNN to recognize the image of MNIST and CIFAR-10 datasets, respectively. Results show that the energy efficiency of the proposed CIM architecture achieves 58.32 TOPS/W with 8-bit input/weight precision. Meanwhile, the inference accuracy of the HS-CIM system for MNIST and CIFAR-10 is 99.25% and 92.34%, respectively. Youxiang Chen, Zhengkun Gu, Haiming Qiu, Kun Zhang 0030, Weisheng Zhao 0001, Yue Zhang 0010 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2024 | RSACIM: Resistance Summation Analog Computing in Memory With Accuracy Optimization Scheme Based on MRAMabstractComputing in memory (CIM) has become a promising candidate to address the Von Neumann bottleneck in processors designed for data-intensive applications. In this article, we propose a resistance summation analog computing in memory (RSACIM) with accuracy optimization scheme in spin transfer torque magnetic random access memory (STT-MRAM), in order to realize energy-efficient and highly reliable analog multiply-and-accumulation (MAC) operation. Firstly, we construct a resistance summation array by serial magnetic tunnel junctions (MTJs) to perform analog MAC operation utilizing time domain technology. Secondly, in order to reduce the impact of position-dependent error caused by resistance summation mechanism, we propose an accuracy optimization scheme to maximize the sensing margin (SM) and computation accuracy. Finally, we design a power-gated reconfigurability control scheme to implement power saving corresponding to different precisions for both input and weight. Evaluation on a 2 Kb RSACIM architecture shows an energy efficiency of 92.9 TOPS/W. System level simulation shows that comparing to existing CIMs based on MRAM, RSACIM architecture saves the inference energy by 4.2 times with 8.4 times lower latency in CIFAR10 image classification task. Zhengkun Gu, Youxiang Chen, Kun Zhang 0030, Youguang Zhang, Yue Zhang 0010 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2023 | Magnetic coupling governed pinning directions in magnetic tunnel junctions under magnetic field annealing with zero magnetic field cooling
Shaohua Yan, Shiyang Lu, Xiaonan Zhao, Runrun Hao, Zitong Zhou, Kun Zhang 0030, Shishen Yan, Qunwen Leng |
Sci. China Inf. Sci. | 9 |
| 2023 | Implementation of 16 Boolean logic operations based on one basic cell of spin-transfer-torque magnetic random access memory
Kaihua Cao, Kun Zhang 0030, Kewen Shi, Zuolei Hao, Wenlong Cai, Ao Du, Jialiang Yin, Jianfeng Gao 0005, Weisheng Zhao 0001 |
Sci. China Inf. Sci. | 3 |
| 2022 | Reconfigurable Bit-Serial Operation Using Toggle SOT-MRAM for High-Performance Computing in Memory ArchitectureabstractComputing in memory (CIM) is a promising candidate for high throughput and energy-efficient data-driven applications, which mitigates the well-known memory bottleneck in Von Neumann architecture. In this paper, we present a reconfigurable bit-serial operation using toggle spin-orbit torque magnetic random access memory (TSOT-MRAM) to perform the computation completely in the bit-cell array instead of in a peripheral circuit. This bit-serial CIM (BSCIM) scheme achieves higher throughput and energy efficiency in CIM. First, basic Boolean logic operations are realized by utilizing the feature of TSOT device. A bit-cell array that implements the bit-serial operation is then built to provide the communication between column and row necessary for arithmetic operations, such as the carry propagation of addition and multiplication. Finally, we analyze the reliability of BSCIM scheme and demonstrate the performance advantage by performing convolution operations for$28\times 28$handwritten digit images in a BSCIM architecture. The results show that the delay and energy of BSCIM architecture are respectively reduced by 1.16-5.49 times and 1.12-1.43 times compared with the existing digital CIM architectures. Besides, its throughput and energy efficiency are also enhanced to 51.2 GOPS and 9.9 TOPS/W respectively. Yining Bai, Zuolei Hao, Guanda Wang, Kun Zhang 0030, Youguang Zhang, Weifeng Lv, Yue Zhang 0010 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2021 | Time-Domain Computing in Memory Using Spintronics for Energy-Efficient Convolutional Neural NetworkabstractThe data transfer bottleneck in Von Neumann architecture owing to the separation between processor and memory hinders the development of high-performance computing. The computing in memory (CIM) concept is widely considered as a promising solution for overcoming this issue. In this article, we present a time-domain CIM (TD-CIM) scheme using spintronics, which can be applied to construct the energy-efficient convolutional neural network (CNN). Basic Boolean logic operations are implemented through recording the bit-line output at different moments. A multi-addend addition mechanism is then introduced based on the TD-CIM circuit, which can eliminate the cascaded full adders. To further optimize the compatibility of TD-CIM circuit for CNN, we also propose a quantization method that transforms floating-point parameters of pre-trained CNN models into fixed-point parameters. Finally, we build a TD-CIM architecture integrating with a highly reconfigurable array of field-free spin-orbit torque magnetic random access memory (SOT-MRAM) and evaluate its benefits for the quantized CNN. By performing digit recognition with the MNIST dataset, we find that the delay and energy are respectively reduced by 1.22.7 times and 2.4×103-1.1×104times compared with STT-CIM and CRAM based on spintronic memory. Finally, the recognition accuracy can reach 98.65% and 91.11% on MNIST and CIFAR10, respectively. Yue Zhang 0010, Chenyu Lian, Yining Bai, Guanda Wang, Zhizhong Zhang 0004, Zhenyi Zheng, Kun Zhang 0030, Georgios Ch. Sirakoulis, Youguang Zhang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 9 |
| 2020 | A Novel In-memory Computing Scheme Based on Toggle Spin Torque MRAMabstractThis paper proposes a novel in-memory computing (IMC) scheme based on toggle spin torque magnetic random access memory (TST-MRAM), called TST-IMC, which makes full use of the unique TST writing mechanism. In this scheme, all of the computing results are directly written in bit-cells without transferring data out of the memory array. Varied Boolean logic operations, such as, NAND, NOR and XOR, can be achieved by specially configuring decision cells. We can also implement three-input majority logic through replacing a decision cell with a datum cell, which can further be used to realize the carry of full-adder. By using 28 nm CMOS technology node and 50 nm-diameter TST-MRAM, we perform mixed simulations to validate the functionality of the proposed TST-IMC scheme. Simulation results show that XOR logic operation can be carried out within 4 ns at 1.8 V supply voltage while the other basic logic operations can be faster, i.e. within 2 ns. In addition, TST-IMC 33% less time and 44% energy saved comparing with existing IMC schemes. Yining Bai, Yue Zhang 0010, Guanda Wang, Zhizhong Zhang 0004, Zhenyi Zheng, Kun Zhang 0030, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 7 |
| 2020 | An In-memory Highly Reconfigurable Logic Circuit Based on Diode-assisted Enhanced Magnetoresistance DeviceabstractIn the post-Moore era, in order to solve the problem of von Neumann bottleneck and memory wall caused by separation of memory and processor, in-memory-processing (IMP) technique has aroused great attention. Novel non-volatile memory (NVM) based on spintronic devices shows promise for satisfying the needs of low-power consumption and high speed for IMP. However, most spintronic memories based on magnetic tunnel junctions (MTJs) can only implement simple and specific logic functions due to the limits of single device and circuit structure. Otherwise, performing logic functions in memory generates vast dynamic power consumption during frequent reading and writing processes because of the high resistance of miniaturized MTJ. In this paper, we propose an in-memory highly reconfigurable logic circuit based on diode-assisted enhanced magnetoresistance (DEMR) device. Our circuit can realize 16 different logic functions with extremely limited circuit area benefiting from the special structure of DEMR device. With appropriate adjustment of control bit and current, the proposed circuit can further implement complex functions like full adder. The proposed reconfigurable circuit can flexibly meet the performance requirements in different scenarios and will contribute a lot for future in-memory chip design. Yue Zhang 0010, Kun Zhang 0030, Zhizhong Zhang 0004, Youguang Zhang, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2020 | Efficient Time-Domain In-Memory Computing Based on TST-MRAMabstractIn-memory computing is highly promising to address the processor-memory data transfer bottleneck in current computational paradigm. We firstly propose a timedomain in-memory computing (TIMC) scheme based on highspeed low-power toggle spin torque random access memory (TST-MRAM). The difference of voltage drops of bitline caused by simultaneously-activated bit-cells is reflected to time domain. Reconfigurable logic operations can be performed by utilizing D flip-flops (DFFs) to record the outputs at different moments. In order to demonstrate the advantages of this scheme in terms of speed and energy consumption, an efficient multi-digit addition circuit has been designed and analyzed. Compared with existing IMC schemes, such as spin-transfer torque computing-in-memory (STT-CiM) structure, up to 67% energy saving and 10 times delay improvement can be achieved in the case of four-digit addition by using TIMC scheme. Yue Zhang 0010, Chenyu Lian, Yining Bai, Guanda Wang, Kun Zhang 0030, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 7 |