VLDB 2026 Research / reviewers in the wild / expert
Chao Wang 0094
dblp:188/7759-94
· DBLP profile ↗
17ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0003-4836-7648ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 7 first-author · 16 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Input Sparsity Aware In-Memory Computing Macro Based on SOT-MRAM Multi-Level Cell for Efficient Deep Neural Network AccelerationabstractDeep neural network (DNN) technology has gained widespread applications, but its high energy demands continue to drive the advancement of low-power computing architectures, particularly in in-memory computing (IMC) architectures based on non-volatile memory. Among these, spin-transfer torque magnetic random-access memory (STT-MRAM)-based IMC architectures have achieved some progress, but their performance remains constrained by limited resistance and binary characteristics. By contrast, the next-generation spin-orbit torque MRAM (SOT-MRAM) offers superior magnetic tunnel junction (MTJ) resistance and more flexible cell structures, presenting significant potential for energy-efficient IMC implementation. In this work, leveraging the ultra-high MTJ resistance and the separation of read/write paths in SOT-MRAM, we propose a multi-level cell (MLC) structure-based high energy-efficiency IMC architecture (MLC-SOT-IMC), which performs standard multiplication operations by optimizing the conductance mapping paradigm. The proposed architecture not only maintains high inference accuracy but also significantly enhances integration density and reduces the overhead per bit. Additionally, a self-terminating time-to-digital converter (TDC) readout circuit, which is dependent on input sparsity, is introduced to eliminate the excess power consumption associated with ineffective pulses after readout completion. Ultimately, the proposed MLC-SOT-IMC architecture achieves an inference energy efficiency of 6388.98 1-bit TOPS/W under an input sparsity of 50%, with the peak energy efficiency reaching 8426.19 1-bit TOPS/W at an input sparsity of 90%. Chao Wang 0094, Qihang Gao, Xianzeng Guo, Zhongzhen Tong, Zhaohao Wang, Weisheng Zhao 0001 |
DATE | 1 |
| 2026 | High-Performance and High-Density NAND-Like SOT-MRAM for FinFET Technology NodesabstractThis paper proposes a comprehensive optimization framework for NAND-like spintronics memory (NAND-SPIN) in advanced FinFET technology nodes. At bit-cell structure level, we propose a NAND-SPIN-GND design which is configured with a grounded bit line (BL) to minimize the parasitic resistance in both read and write paths, thereby decreasing read latency by 35.0% and write energy by 27.3%. At device and layout level, a short-circuiting bottom electrode (SBE) design is proposed, which shorts non-contributing spin-orbit torque (SOT) segments by the BEs, reducing read latency by up to 41.4% and write energy by 55.8%. In addition, a compact capacitance symmetric source line (SL)-type reference scheme is introduced to address the inherent capacitance asymmetry in conventional SL-type reference scheme, resulting in a 48.6% reduction in read latency compared to the conventional word line (WL)-type reference scheme. Chao Wang 0094, Xianzeng Guo, Luman Xiang, Zhaohao Wang, Weisheng Zhao 0001 |
DATE | 1 |
| 2026 | A Low Power and High Reliability Nonvolatile SRAM Using In-Plane VGSOT-MRAM with Pre-Charge Restore SchemeabstractConventional magnetic nonvolatile-static random access memory (MNV-SRAM) suffers from large store current, which leads to low area and energy efficiency, severely limiting their development and application. This paper proposes a 10T-2MTJ NV-SRAM cell based on the in-plane voltage-gated spin-orbit torque magnetic tunnel junction (VGSOT-MTJ), which enables field-free deterministic magnetization switching and reduces the required store current by leveraging the voltage-controlled magnetic anisotropy (VCMA) effect to assist the store operation. Thereby, the proposed design achieves the smallest SRAM cell area compared to prior works, due to the relaxed transistor drive strength requirement. On the other hand, existing NV-SRAM restore schemes exhibit a substantial deterioration in restore error ratio (RSER) with increasing MTJ resistance. Targeting the high resistance characteristics of VGSOT-MTJ, we innovatively propose a pre-charge restore scheme with sensitive transistor isolation. Simulation results demonstrate that the proposed design achieves the lowest read and write energy in SRAM mode, with store energy 1.58× to 2.48× lower than other in-plane MTJ-based designs. And the proposed restore scheme significantly improves restore reliability with over 98.7% RSER enhancement, and shows superior robustness across different MTJ resistances and tunnel magnetoresistance ratio (TMR) conditions. Zhongzhen Tong, Mingche Li, Weimeng Zhao, Zhongkui Zhang, Yaling Wang, Chao Wang 0094, Zhaohao Wang |
DATE | 8 |
| 2026 | Highly Energy-Efficient In-Memory Computing Architecture Based on VGSOT-MRAM for Reconfigurable BNN/TNN Acceleration
Qihang Gao, Chao Wang 0094, Chenghang Li, Zhongzhen Tong, Zhaohao Wang |
ISCAS | 2 |
| 2026 | MEPUF: A Lightweight and ML-Resistant Strong PUF Integrating Dual-Mode MRAM and Configurable AES for Reliable UAVs
Bi Wang 0002, Luyao Shi, Chao Wang 0094, Zhaohao Wang |
ISCAS | 5 |
| 2026 | Multi-Retention and Bit-Level Approximate STT-MRAM for High-Efficiency AI Applications
Yulong Qiu, Chao Wang 0094, Weimeng Zhao, Zhongzhen Tong, Zhaohao Wang |
ISCAS | 2 |
| 2026 | BaM-CIM: A High Throughput Booth Algorithm-Based In-MRAM Computing Macro Using Hybrid VGSOT-MTJ/GAA-CNTFETabstractAs artificial intelligence (AI) and computational models grow in scale, the demand for computational power and storage has significantly increased. The computing-in-memory (CIM) architecture addresses this challenge by performing computations directly within the memory array, reducing data transfer between the processor and memory. This paper introduces a Booth algorithm-based In-MRAM computing architecture (BaM-CIM) using a hybrid voltage-gated spin-orbit torque MTJ (VGSOT-MTJ) and gate-all-around carbon nanotube field-effect transistors (GAA-CNTFETs) for efficient multiply-and-accumulate (MAC) computing. The key contributions of BaM-CIM are as follows: 1) A Voltage divider reference (VDR) cell is proposed, which enables read operations using only a 2T1M cell structure. Compared to complementary read cells, the VDR reduces the area by half and achieves robust data sensing without requiring precharge/discharge operations. 2) The BaM-CIM circuit is proposed to complete 8b-W/8b-IN/21b-OUT computations in only two cycles (1.6 ns), reducing the number of cycles by 75% compared to single-bit input serial operations and by 50% compared to two-bit serial operations. 3) A three-input 8b Booth computing adder (BCA), along with Modified computing shift adder (MCSA) and Modified computing post adder (MCPA), which can achieve higher energy efficiency. BaM-CIM with 128 Kb is simulated, achieving throughput and energy efficiency of 0.93 TOPS and 258.4 TOPS/W, respectively, at a 0.6 V supply voltage and 1.28 TOPS and 169.5 TOPS/W, respectively, at a 0.8 V supply voltage with 8b-IN, 8b-W, and 21b-OUT. Chenghang Li, Zhongzhen Tong, Yulong Qiu, Jiye Yao, Chao Wang 0094, Zhaohao Wang, Xiaoyang Lin, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | High-Speed VGSOT-MRAM Design for Non-Volatile Cache MemoriesabstractSpin-orbit torque magnetic random-access memory (SOT-MRAM) is a promising candidate for next-generation memory systems, particularly for cache applications, owing to its ultra-fast write speed and high endurance. However, SOTMRAM confronts challenges of large bit-cell area and high write current. Voltage-gated-SOT-MRAM (VGSOT-MRAM) mitigates these issues through the voltage controlled magnetic anisotropy (VCMA) mechanism, reducing the write current and enabling high-density device structure, but at the cost of slower read speed due to high device resistance. To address the read speed issue, we propose the local read bit-line (LRBL) scheme, which decreases the load capacitance of the read path and can reduce the read latency by 55.0% with minimal area overhead. Additionally, an efficient parallel-discharge-serial-sensing (PDSS) scheme is proposed to optimize the sequential read operations in cache, achieving up to 84.8% latency reduction during cache pre-fetch operations. Furthermore, implementing an appropriate error checking and correction (ECC) algorithm can further diminish the total read latency by 26.4%. Xianzeng Guo, Chao Wang 0094, Luman Xiang, Zhaohao Wang |
ISCAS | 2 |
| 2025 | Approximate SOT-MRAM for Neural Network Acceleration with Superior Read PerformanceabstractMagnetoresistive random-access memory (MRAM) has been demonstrated to be a suitable memory technology for neural network (NN) acceleration due to its non-volatility, high density, and fast access speed. However, compared to the widely used static random-access memory (SRAM), MRAM still exhibits a notable disparity in speed and energy. In this paper, we propose a read-related approximation computation (RAC) strategy based on spin-orbit torque MRAM (SOT-MRAM) to enhance the computational speed of NN, and then we introduce a reference reconfigurable array (RRA) architecture to further decrease the read latency and energy consumption, significantly improving the speed and energy efficiency of weight retrieval during computations. Furthermore, we propose an algorithm to verify and optimize NN model performance. The proposed architecture is evaluated using a 28 nm process combined with a SPICE model of the SOT-MRAM. Simulation results indicate that the read speed increases by 4.71X, the read energy consumption is reduced by 70.7%, while the model accuracy loss remains below 1%. Yulong Qiu, Chao Wang 0094, Zhongzhen Tong, Siyuan Cheng 0021, Zhaohao Wang |
ISCAS | 2 |
| 2025 | A Self-Decryption Pass Transistor Logic-Based In-MRAM Computing Macro Using Hybrid VGSOT-MTJ/GAA-CNTFETabstractSpintronic devices and gate-all-around carbon nanotube field-effect-transistors (GAA-CNTFETs)-based computing in-memory architecture are competitive candidates for applications in battery-powered tiny artificial intelligence (AI) edge devices. Meanwhile, data encryption and decryption are also necessary to protect AI model weights and the customized data used to guarantee neural network (NN) inference accuracy. In this study, we propose a self-decryption pass transistor logic (PTL)-based in-MRAM computing macro (SP-CIM) that utilizes hybrid voltage-gated spin-orbit torque magnetic tunnel junctions (VGSOT-MTJ)/GAA-CNTFET. The proposed SP-CIM macro enables simultaneous data access, decryption, and full-accuracy multiply-and-accumulate (MAC) operations using the newly introduced voltage-divider self-decryption cell, without the need for additional decryption logic. Compared to existing in-memory decryption strategies, this design reduces energy consumption by 45.7% and decreases decryption delay by 87.2%. To enhance area efficiency and reduce computing latency, we propose a PTL-based multiplication cell that achieves full-accuracy local 2b-IN TEXPRESERVE0 2b-W operations with only 20 transistors (20T). Additionally, novel PTL-based full-swing output half adders (10T-HA) and full adders (14T-FA) are proposed to construct the local adder tree, achieving reductions of 31.8%, 76.4%, and 41.4% in energy, delay, and area, respectively, compared to conventional adder trees in CIM macros. Simulations of the 288 kb SP-CIM macro demonstrated throughput and energy efficiency of 2.25 TOPS and 226.6 TOPS/W, respectively, at a 0.6 V supply voltage, and 2.97 TOPS and 154.1 TOPS/W, respectively, at a 0.8 V supply voltage, with 8b-IN, 8b-W, and 24b-OUT. Zhongzhen Tong, Sifan Sun, Chenghang Li, Jiye Yao, Yulong Qiu, Chao Wang 0094, Zhaohao Wang, Amara Amara, Xiaoyang Lin, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2025 | Technically Feasible Robust Complementary SOT-MRAM Design for Improving the Area and Energy EfficiencyabstractSpin-orbit torque magnetic random-access memory (SOT-MRAM), which exhibits sub-nanosecond write speed and high endurance, is a promising candidate for the future high-level cache. Nevertheless, SOT-MRAM faces challenge in meeting the high read performance requirements of cache applications due to the limited ON/OFF ratio. Consequently, extensive investigation has been conducted into robust complementary bit-cell (CBC) designs based on SOT-MRAM. However, previous designs suffer from significant technology feasibility, area and performance issues. In this paper, the feasibility and performance of the existing complementary write schemes are analyzed, and optimized U-type and toggle spin torque (TST) schemes with practicality and conciseness are presented. The previous CBC designs are evaluated and optimized in terms of circuit and layout, while the 1-word-line-3-bit-line (1WL3BL) CBC designs with both U-type and TST schemes are proposed, which can reduce the bit-cell area by 24.64%-27.54% and improve the write and read performance. In comparison to the conventional CBC design, the proposed 1WL3BL CBC design can reduce the write energy and read latency by up to 36.91% and 21.93%, respectively. Furthermore, the proposed low-voltage read scheme demonstrates the capability to enhance the read performance and conserve the read energy under the aggressive read-related process parameters. Chao Wang 0094, Zhongkui Zhang, Xianzeng Guo, Qihang Gao, Zhaohao Wang, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2024 | BSTCIM: A Balanced Symmetry Ternary Fully Digital In-MRAM Computing Macro for Energy Efficiency Neural NetworkabstractSilicon-based traditional binary computing in-memory (TBCIM) architectures are approaching their energy efficiency and throughput limits owing to challenges facing Moore’s Law. Thus, it is essential to explore architecture based on novel devices and computing paradigms to fulfill data-centric applications, such as artificial intelligence. In this paper, we propose a balanced symmetry ternary (BST) fully digital in-MRAM computing macro (BSTCIM) using hybrid voltage-gated spin-orbit torque magnetic tunnel junctions (VGSOT-MTJ) and gate-all-around carbon nanotube field-effect-transistors (GAA-CNTFET) technology. The overall computing is based on the highest efficiency multi-bit ternary system. BSTCIM includes a ternary dot product (TDP) unit with 4 GAA-CNTFETs and 2 VGSOT-MTJs achieving TDP operation without complex logic circuits. The multi-bit ternary multiply-and-accumulate (MAC) operation is realized through the proposed ternary adder tree and ternary post adder which accumulate TDP results within the digital domain enabling high accuracy neural network inference. Furthermore, due to the advantages of BST, ternary signed MAC is more easily performed compared to TBCIM macros that adapt 2’s complement or separate signed bit calculations. BSTCIM with 288 kb is simulated, achieving throughput and energy efficiency of 0.72 TOPS and 54.5 TOPS/W, respectively, at a 0.6 V supply voltage and 1.15 TOPS and 33.7 TOPS/W, respectively at a 0.8 V supply voltage with 8b-IN, 8b-W, and 20b-OUT. Moreover, the figure-of-merit for BSTCIM is 1.13–33.6 times higher than that of existing CIM macros. Zhongzhen Tong, Chenghang Li, Chao Wang 0094, Suteng Zhao, Qianyong Peng, Daming Zhou, Zhaohao Wang, Xiaoyang Lin, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | Variation Aware Evaluation Approach and Design Methodology for SOT-MRAMabstractSpin-orbit torque magnetic random access memory (SOT-MRAM), which exhibits sub-nanosecond write speed and high reliability, is a promising candidate for the future high-level cache. However, SOT-MRAM faces the problem of large bit-cell layout area due to its structural characteristics and write performance requirements, therefore it is necessary to explore the bit-cell design with optimal overall performance under the unified bit-cell area. In this paper, we propose a comprehensive variation aware evaluation approach for the area, latency, and energy of SOT-MRAM under the uniform yield standard. Based on this, the mainstream SOT-MRAM bit-cell designs with high-density method and multi-finger configuration are evaluated, meanwhile bit-cell designs with excellent write performance and their optimum area ranges are identified. Moreover, the source line read (SLR) mode with higher robustness against transistor variation is proposed to improve the read performance, and the dual SL (DSL) method is proposed to further reduce the read latency and write energy. With the DSL method, the read latency and write energy of 2-word-line (WL)-type bit-cells can be reduced by up to 36.5% and 12.6%, respectively. In addition, the DSL method can solve the shunt current issue of 1WL-type bit-cells and reduce the read latency and write energy by up to 43.6% and 17.4%, respectively. Chao Wang 0094, Zhaohao Wang, Shixing Li, Zhongkui Zhang, Youguang Zhang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | Layout Aware Optimization Methodology for SOT-MRAM Based on Technically Feasible Top-Pinned Magnetic Tunnel Junction ProcessabstractThe emerging spin-orbit torque magnetic random-access memory (SOT-MRAM) shows promising prospects in high-level cache applications due to its subnanosecond switching speed and high reliability. However, SOT-MRAM faces the issue of large bit-cell layout area, which is currently the focus of attention. Although many design and evaluation works have emerged, the lack of a unified standard for realistic SOT process has hindered the development of relevant research toward practicality. In this article, the bit-cell area of the SOT-MRAM will be evaluated and optimized based on the technically feasible process. First of all, based on the state-of-the-art top-pinned SOT nanopillar process, the SOT-MRAM design rules are proposed. On this basis, this article systematically summarizes four basic device layout modes and provides optimized layout suggestions for conventional SOT bit-cells with different types and sizes of devices. In addition, a series of area-efficient SOT bit-cell designs based on the common area (CA) and dual common (DC) solutions are proposed, which can reduce the layout area of SOT bit-cells by up to 38.4% with reasonable write latency and energy overhead. Chao Wang 0094, Zhaohao Wang, Zhongkui Zhang, Jiagao Feng, Youguang Zhang, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | Reconfigurable and Dynamically Transformable In-Cache-MPUF System With True Randomness Based on the SOT-MRAMabstractIn this paper, we present a reconfigurable Physically Unclonable Functions (PUF) based on the Spin-Orbit-Torque Magnetic Random-Access Memory (SOT-MRAM), which exploits thermal noise as the true dynamic entropy source. Therefore, the MRAM cells could be configured to random final states with stochastic switching mechanism. The proposed PUF is constructed and reconfigured by combining the small-capacity true random number generator (TRNG) and high-reliability secure hash algorithm (SHA-512), realizing the dynamic transformation between SOT-MRAM based last level cache and PUF (In-Cache-MPUF). Thanks to the full reconfigurability and the high endurance of SOT-MRAM, the proposed In-Cache-MPUF can achieve$10^{\textbf {14}}$maximum PUF bits per cell, which has greatly motivated the implementations compared with the traditional weak PUFs utilizing the static entropy source of process variations. The Monte-Carlo simulation results using 40 nm technology and a compact MTJ model show that the proposed PUF has desirable randomness as the digitized bit streams passing all the NIST tests, achieving 50.0428% uniqueness as well as 49.9236% uniformity. It also shows comparable reliability to the state-of-the-art works: a maximum bit error rate of 0.14% and 0.12% at 100 °C and 0.9 V, respectively. In addition, the system level performance is tested and validated by gem5. Zhengyi Hou, Zhaohao Wang, Chao Wang 0094, Min Wang 0033, You Wang 0002, Cenlin Duan, Jianlei Yang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | Computing-in-Memory Paradigm Based on STT-MRAM with Synergetic Read/Write-Like ModesabstractWith the surge in demand for data storage and processing in emerging applications, the traditional CMOS-based Von-Neumann architecture is facing challenges such as memory wall and static power consumption. In order to conquer the above-mentioned bottlenecks in computing systems, computing in-memory (CiM) architectures based on non-volatile memory (NVM) have been widely researched. In this paper, we propose a CiM paradigm based on spin-transfer torque magnetic random access memory (STT-MRAM), which combines common read-like mode (RLM) and write-like mode (WLM). On the basis of realizing the basic functions AND/OR/NAND/NOR, our design coordinates the high speed of RLM and the integrity of WLM to perform complex operations like full-adder (FA) and XOR/XNOR. In addition, the high speed and low power consumption of the proposed CiM paradigm are established by circuit-level simulation with a 40 nm design kit. Chao Wang 0094, Zhaohao Wang, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 1 |
| 2020 | Computing-in-Memory Architecture Based on Field-Free SOT-MRAM with Self-Reference MethodabstractOn the current computing platforms, the memory wall between processor and memory has become the toughest challenge for the traditional Von-Neumann computer architecture. Computing-in-Memory (CIM) is taken as a promising approach to solving the above bottleneck in computing systems. In this paper, we propose a CIM platform with field-free spinorbit torque magnetic random access memory (SOT-MRAM). The self-reference (SelfRef) method is designed to enhance the read reliability and directly obtain logic results through memory-like read operations without adding logic cells. Memory read/write and logic operations, including NOT, AND/NAND and OR/NOR, can be implemented in the same SOT-MRAM chip. The speed and power penalties caused by SelfRef scheme are acceptable thanks to the ultrafast switching of the SOT. The read reliability and logic correctness of the proposed CIM are demonstrated by hybrid simulation on a 40 nm technology node. Chao Wang 0094, Zhaohao Wang, Jianlei Yang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 1 |