Zhenyi Zheng

dblp:207/3975 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
3since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 since 2021
YearPublicationVenuePosition
2026 Sparsh: Breaking the Communication Bottleneck in Sequence Parallel Video Diffusion Inference with Predictive Sparse Communication
Zhenyi Zheng, Jiangsu Du
Euro-Par (1)3
2025 TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
abstract
As the model size continuously increases, pipeline parallelism shows great promise in throughput-oriented LLM inference due to its low demand on communications. However, imbalanced pipeline workloads and complex data dependencies in the prefill and decode phases result in massive pipeline bubbles and further severe performance reduction.
Hongbin Zhang 0006, Taosheng Wei, Zhenyi Zheng, Jiangsu Du, Zhiguang Chen 0001, Yutong Lu
ICPP3
2021 Time-Domain Computing in Memory Using Spintronics for Energy-Efficient Convolutional Neural Network
abstract
The data transfer bottleneck in Von Neumann architecture owing to the separation between processor and memory hinders the development of high-performance computing. The computing in memory (CIM) concept is widely considered as a promising solution for overcoming this issue. In this article, we present a time-domain CIM (TD-CIM) scheme using spintronics, which can be applied to construct the energy-efficient convolutional neural network (CNN). Basic Boolean logic operations are implemented through recording the bit-line output at different moments. A multi-addend addition mechanism is then introduced based on the TD-CIM circuit, which can eliminate the cascaded full adders. To further optimize the compatibility of TD-CIM circuit for CNN, we also propose a quantization method that transforms floating-point parameters of pre-trained CNN models into fixed-point parameters. Finally, we build a TD-CIM architecture integrating with a highly reconfigurable array of field-free spin-orbit torque magnetic random access memory (SOT-MRAM) and evaluate its benefits for the quantized CNN. By performing digit recognition with the MNIST dataset, we find that the delay and energy are respectively reduced by 1.22.7 times and 2.4×103-1.1×104times compared with STT-CIM and CRAM based on spintronic memory. Finally, the recognition accuracy can reach 98.65% and 91.11% on MNIST and CIFAR10, respectively.
Yue Zhang 0010, Chenyu Lian, Yining Bai, Guanda Wang, Zhizhong Zhang 0004, Zhenyi Zheng, Kun Zhang 0030, Georgios Ch. Sirakoulis, Youguang Zhang
IEEE Trans. Circuits Syst. I Regul. Pap.7
2020 A Novel In-memory Computing Scheme Based on Toggle Spin Torque MRAM
abstract
This paper proposes a novel in-memory computing (IMC) scheme based on toggle spin torque magnetic random access memory (TST-MRAM), called TST-IMC, which makes full use of the unique TST writing mechanism. In this scheme, all of the computing results are directly written in bit-cells without transferring data out of the memory array. Varied Boolean logic operations, such as, NAND, NOR and XOR, can be achieved by specially configuring decision cells. We can also implement three-input majority logic through replacing a decision cell with a datum cell, which can further be used to realize the carry of full-adder. By using 28 nm CMOS technology node and 50 nm-diameter TST-MRAM, we perform mixed simulations to validate the functionality of the proposed TST-IMC scheme. Simulation results show that XOR logic operation can be carried out within 4 ns at 1.8 V supply voltage while the other basic logic operations can be faster, i.e. within 2 ns. In addition, TST-IMC 33% less time and 44% energy saved comparing with existing IMC schemes.
Yining Bai, Yue Zhang 0010, Guanda Wang, Zhizhong Zhang 0004, Zhenyi Zheng, Kun Zhang 0030, Weisheng Zhao 0001
ACM Great Lakes Symposium on VLSI6