Xianzeng Guo

dblp:405/1689 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Input Sparsity Aware In-Memory Computing Macro Based on SOT-MRAM Multi-Level Cell for Efficient Deep Neural Network Acceleration
abstract
Deep neural network (DNN) technology has gained widespread applications, but its high energy demands continue to drive the advancement of low-power computing architectures, particularly in in-memory computing (IMC) architectures based on non-volatile memory. Among these, spin-transfer torque magnetic random-access memory (STT-MRAM)-based IMC architectures have achieved some progress, but their performance remains constrained by limited resistance and binary characteristics. By contrast, the next-generation spin-orbit torque MRAM (SOT-MRAM) offers superior magnetic tunnel junction (MTJ) resistance and more flexible cell structures, presenting significant potential for energy-efficient IMC implementation. In this work, leveraging the ultra-high MTJ resistance and the separation of read/write paths in SOT-MRAM, we propose a multi-level cell (MLC) structure-based high energy-efficiency IMC architecture (MLC-SOT-IMC), which performs standard multiplication operations by optimizing the conductance mapping paradigm. The proposed architecture not only maintains high inference accuracy but also significantly enhances integration density and reduces the overhead per bit. Additionally, a self-terminating time-to-digital converter (TDC) readout circuit, which is dependent on input sparsity, is introduced to eliminate the excess power consumption associated with ineffective pulses after readout completion. Ultimately, the proposed MLC-SOT-IMC architecture achieves an inference energy efficiency of 6388.98 1-bit TOPS/W under an input sparsity of 50%, with the peak energy efficiency reaching 8426.19 1-bit TOPS/W at an input sparsity of 90%.
Chao Wang 0094, Qihang Gao, Xianzeng Guo, Zhongzhen Tong, Zhaohao Wang, Weisheng Zhao 0001
DATE3
2026 High-Performance and High-Density NAND-Like SOT-MRAM for FinFET Technology Nodes
abstract
This paper proposes a comprehensive optimization framework for NAND-like spintronics memory (NAND-SPIN) in advanced FinFET technology nodes. At bit-cell structure level, we propose a NAND-SPIN-GND design which is configured with a grounded bit line (BL) to minimize the parasitic resistance in both read and write paths, thereby decreasing read latency by 35.0% and write energy by 27.3%. At device and layout level, a short-circuiting bottom electrode (SBE) design is proposed, which shorts non-contributing spin-orbit torque (SOT) segments by the BEs, reducing read latency by up to 41.4% and write energy by 55.8%. In addition, a compact capacitance symmetric source line (SL)-type reference scheme is introduced to address the inherent capacitance asymmetry in conventional SL-type reference scheme, resulting in a 48.6% reduction in read latency compared to the conventional word line (WL)-type reference scheme.
Chao Wang 0094, Xianzeng Guo, Luman Xiang, Zhaohao Wang, Weisheng Zhao 0001
DATE2
2025 High-Speed VGSOT-MRAM Design for Non-Volatile Cache Memories
abstract
Spin-orbit torque magnetic random-access memory (SOT-MRAM) is a promising candidate for next-generation memory systems, particularly for cache applications, owing to its ultra-fast write speed and high endurance. However, SOTMRAM confronts challenges of large bit-cell area and high write current. Voltage-gated-SOT-MRAM (VGSOT-MRAM) mitigates these issues through the voltage controlled magnetic anisotropy (VCMA) mechanism, reducing the write current and enabling high-density device structure, but at the cost of slower read speed due to high device resistance. To address the read speed issue, we propose the local read bit-line (LRBL) scheme, which decreases the load capacitance of the read path and can reduce the read latency by 55.0% with minimal area overhead. Additionally, an efficient parallel-discharge-serial-sensing (PDSS) scheme is proposed to optimize the sequential read operations in cache, achieving up to 84.8% latency reduction during cache pre-fetch operations. Furthermore, implementing an appropriate error checking and correction (ECC) algorithm can further diminish the total read latency by 26.4%.
Xianzeng Guo, Chao Wang 0094, Luman Xiang, Zhaohao Wang
ISCAS1
2025 Technically Feasible Robust Complementary SOT-MRAM Design for Improving the Area and Energy Efficiency
abstract
Spin-orbit torque magnetic random-access memory (SOT-MRAM), which exhibits sub-nanosecond write speed and high endurance, is a promising candidate for the future high-level cache. Nevertheless, SOT-MRAM faces challenge in meeting the high read performance requirements of cache applications due to the limited ON/OFF ratio. Consequently, extensive investigation has been conducted into robust complementary bit-cell (CBC) designs based on SOT-MRAM. However, previous designs suffer from significant technology feasibility, area and performance issues. In this paper, the feasibility and performance of the existing complementary write schemes are analyzed, and optimized U-type and toggle spin torque (TST) schemes with practicality and conciseness are presented. The previous CBC designs are evaluated and optimized in terms of circuit and layout, while the 1-word-line-3-bit-line (1WL3BL) CBC designs with both U-type and TST schemes are proposed, which can reduce the bit-cell area by 24.64%-27.54% and improve the write and read performance. In comparison to the conventional CBC design, the proposed 1WL3BL CBC design can reduce the write energy and read latency by up to 36.91% and 21.93%, respectively. Furthermore, the proposed low-voltage read scheme demonstrates the capability to enhance the read performance and conserve the read energy under the aggressive read-related process parameters.
Chao Wang 0094, Zhongkui Zhang, Xianzeng Guo, Qihang Gao, Zhaohao Wang, Weisheng Zhao 0001
IEEE Trans. Circuits Syst. I Regul. Pap.4