VLDB 2026 Research / reviewers in the wild / expert
Zhongzhen Tong
dblp:312/7462
· DBLP profile ↗
16ranked-venue papers
5as first author
16since 2021 · last 2026
0000-0001-8907-939XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 5 first-author · 16 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Input Sparsity Aware In-Memory Computing Macro Based on SOT-MRAM Multi-Level Cell for Efficient Deep Neural Network AccelerationabstractDeep neural network (DNN) technology has gained widespread applications, but its high energy demands continue to drive the advancement of low-power computing architectures, particularly in in-memory computing (IMC) architectures based on non-volatile memory. Among these, spin-transfer torque magnetic random-access memory (STT-MRAM)-based IMC architectures have achieved some progress, but their performance remains constrained by limited resistance and binary characteristics. By contrast, the next-generation spin-orbit torque MRAM (SOT-MRAM) offers superior magnetic tunnel junction (MTJ) resistance and more flexible cell structures, presenting significant potential for energy-efficient IMC implementation. In this work, leveraging the ultra-high MTJ resistance and the separation of read/write paths in SOT-MRAM, we propose a multi-level cell (MLC) structure-based high energy-efficiency IMC architecture (MLC-SOT-IMC), which performs standard multiplication operations by optimizing the conductance mapping paradigm. The proposed architecture not only maintains high inference accuracy but also significantly enhances integration density and reduces the overhead per bit. Additionally, a self-terminating time-to-digital converter (TDC) readout circuit, which is dependent on input sparsity, is introduced to eliminate the excess power consumption associated with ineffective pulses after readout completion. Ultimately, the proposed MLC-SOT-IMC architecture achieves an inference energy efficiency of 6388.98 1-bit TOPS/W under an input sparsity of 50%, with the peak energy efficiency reaching 8426.19 1-bit TOPS/W at an input sparsity of 90%. Chao Wang 0094, Qihang Gao, Xianzeng Guo, Zhongzhen Tong, Zhaohao Wang, Weisheng Zhao 0001 |
DATE | 4 |
| 2026 | A Low Power and High Reliability Nonvolatile SRAM Using In-Plane VGSOT-MRAM with Pre-Charge Restore SchemeabstractConventional magnetic nonvolatile-static random access memory (MNV-SRAM) suffers from large store current, which leads to low area and energy efficiency, severely limiting their development and application. This paper proposes a 10T-2MTJ NV-SRAM cell based on the in-plane voltage-gated spin-orbit torque magnetic tunnel junction (VGSOT-MTJ), which enables field-free deterministic magnetization switching and reduces the required store current by leveraging the voltage-controlled magnetic anisotropy (VCMA) effect to assist the store operation. Thereby, the proposed design achieves the smallest SRAM cell area compared to prior works, due to the relaxed transistor drive strength requirement. On the other hand, existing NV-SRAM restore schemes exhibit a substantial deterioration in restore error ratio (RSER) with increasing MTJ resistance. Targeting the high resistance characteristics of VGSOT-MTJ, we innovatively propose a pre-charge restore scheme with sensitive transistor isolation. Simulation results demonstrate that the proposed design achieves the lowest read and write energy in SRAM mode, with store energy 1.58× to 2.48× lower than other in-plane MTJ-based designs. And the proposed restore scheme significantly improves restore reliability with over 98.7% RSER enhancement, and shows superior robustness across different MTJ resistances and tunnel magnetoresistance ratio (TMR) conditions. Zhongzhen Tong, Mingche Li, Weimeng Zhao, Zhongkui Zhang, Yaling Wang, Chao Wang 0094, Zhaohao Wang |
DATE | 3 |
| 2026 | Highly Energy-Efficient In-Memory Computing Architecture Based on VGSOT-MRAM for Reconfigurable BNN/TNN Acceleration
Qihang Gao, Chao Wang 0094, Chenghang Li, Zhongzhen Tong, Zhaohao Wang |
ISCAS | 4 |
| 2026 | Multi-Retention and Bit-Level Approximate STT-MRAM for High-Efficiency AI Applications
Yulong Qiu, Chao Wang 0094, Weimeng Zhao, Zhongzhen Tong, Zhaohao Wang |
ISCAS | 4 |
| 2026 | A Fully-Parallel Digital MRAM Computing-in-Memory Macro Featuring a High-Efficient Dynamic Adder Tree and Bit-Splitting MAC
Zhongzhen Tong, Jiye Yao, Shaohui Ma, Yulong Qiu, Zhaohao Wang, Amara Amara, Xiaoyang Lin |
ISCAS | 1 |
| 2026 | BaM-CIM: A High Throughput Booth Algorithm-Based In-MRAM Computing Macro Using Hybrid VGSOT-MTJ/GAA-CNTFETabstractAs artificial intelligence (AI) and computational models grow in scale, the demand for computational power and storage has significantly increased. The computing-in-memory (CIM) architecture addresses this challenge by performing computations directly within the memory array, reducing data transfer between the processor and memory. This paper introduces a Booth algorithm-based In-MRAM computing architecture (BaM-CIM) using a hybrid voltage-gated spin-orbit torque MTJ (VGSOT-MTJ) and gate-all-around carbon nanotube field-effect transistors (GAA-CNTFETs) for efficient multiply-and-accumulate (MAC) computing. The key contributions of BaM-CIM are as follows: 1) A Voltage divider reference (VDR) cell is proposed, which enables read operations using only a 2T1M cell structure. Compared to complementary read cells, the VDR reduces the area by half and achieves robust data sensing without requiring precharge/discharge operations. 2) The BaM-CIM circuit is proposed to complete 8b-W/8b-IN/21b-OUT computations in only two cycles (1.6 ns), reducing the number of cycles by 75% compared to single-bit input serial operations and by 50% compared to two-bit serial operations. 3) A three-input 8b Booth computing adder (BCA), along with Modified computing shift adder (MCSA) and Modified computing post adder (MCPA), which can achieve higher energy efficiency. BaM-CIM with 128 Kb is simulated, achieving throughput and energy efficiency of 0.93 TOPS and 258.4 TOPS/W, respectively, at a 0.6 V supply voltage and 1.28 TOPS and 169.5 TOPS/W, respectively, at a 0.8 V supply voltage with 8b-IN, 8b-W, and 21b-OUT. Chenghang Li, Zhongzhen Tong, Yulong Qiu, Jiye Yao, Chao Wang 0094, Zhaohao Wang, Xiaoyang Lin, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2026 | An FD-SOI-Based Compact In-Pixel Computing Architecture Enabling Real-Time Feature ExtractionabstractTo empower resource-limited edge devices in artificial intelligence (AI) and Internet of Things (IoT) applications, it is essential to overcome challenges posed by restricted area resources and the high latency demands of transmitting and processing substantial sensory data. In-pixel computing addresses these challenges effectively, and the Fully Depleted Silicon-On-Insulator (FD-SOI)-based pixel, which relies on an FD-SOI transistor whose current is made photosensitive to light by applying a negative back-gate voltage, shows significant potential with its compact structure and in-situ computation capability. In this paper, for the first time, we present an FD-SOI-based chip-level architecture for in-pixel computing. Our design implements programmable, massively parallel convolution with low latency using a pulse-width modulation (PWM) input encoding scheme. Furthermore, the proposed compact 1P1T (1 Phototransistor 1 Transistor) pixel design, integrated with an improved single-slope analog-to-digital converter (SS ADC), greatly enhances area efficiency. Validated through simulation in a 22nm FD-SOI process, the design achieves 990 frames/s under typical outdoor illumination conditions, with the figure of merit (FoM) of 16.86pJ/pixel/frame. In addition, the proposed architecture has been evaluated on hand gesture recognition (5697 training and 633 validation images across six categories), achieving an accuracy of 97.48%. The results demonstrate that, compared to state-of-the-art designs, our approach achieves a$6\times $improvement in in-pixel convolution speed and a$7.3\times $reduction in area overhead. Yijiao Wang, Jiayao Wu, Zhongzhen Tong, Xinrui Duan, Yiming Shi, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2026 | A CIM Macro Embedded With Sign Operations for Parallel Signed Multibit Multiplication-and-Accumulation Using Hybrid Cell Array
Jin Zhang 0036, Zhongzhen Tong, Qiang Zhao 0007, Chunyu Peng, Wenjuan Lu, Zhi-Ting Lin, Xiulong Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2025 | Approximate SOT-MRAM for Neural Network Acceleration with Superior Read PerformanceabstractMagnetoresistive random-access memory (MRAM) has been demonstrated to be a suitable memory technology for neural network (NN) acceleration due to its non-volatility, high density, and fast access speed. However, compared to the widely used static random-access memory (SRAM), MRAM still exhibits a notable disparity in speed and energy. In this paper, we propose a read-related approximation computation (RAC) strategy based on spin-orbit torque MRAM (SOT-MRAM) to enhance the computational speed of NN, and then we introduce a reference reconfigurable array (RRA) architecture to further decrease the read latency and energy consumption, significantly improving the speed and energy efficiency of weight retrieval during computations. Furthermore, we propose an algorithm to verify and optimize NN model performance. The proposed architecture is evaluated using a 28 nm process combined with a SPICE model of the SOT-MRAM. Simulation results indicate that the read speed increases by 4.71X, the read energy consumption is reduced by 70.7%, while the model accuracy loss remains below 1%. Yulong Qiu, Chao Wang 0094, Zhongzhen Tong, Siyuan Cheng 0021, Zhaohao Wang |
ISCAS | 3 |
| 2025 | A Self-Decryption Pass Transistor Logic-Based In-MRAM Computing Macro Using Hybrid VGSOT-MTJ/GAA-CNTFETabstractSpintronic devices and gate-all-around carbon nanotube field-effect-transistors (GAA-CNTFETs)-based computing in-memory architecture are competitive candidates for applications in battery-powered tiny artificial intelligence (AI) edge devices. Meanwhile, data encryption and decryption are also necessary to protect AI model weights and the customized data used to guarantee neural network (NN) inference accuracy. In this study, we propose a self-decryption pass transistor logic (PTL)-based in-MRAM computing macro (SP-CIM) that utilizes hybrid voltage-gated spin-orbit torque magnetic tunnel junctions (VGSOT-MTJ)/GAA-CNTFET. The proposed SP-CIM macro enables simultaneous data access, decryption, and full-accuracy multiply-and-accumulate (MAC) operations using the newly introduced voltage-divider self-decryption cell, without the need for additional decryption logic. Compared to existing in-memory decryption strategies, this design reduces energy consumption by 45.7% and decreases decryption delay by 87.2%. To enhance area efficiency and reduce computing latency, we propose a PTL-based multiplication cell that achieves full-accuracy local 2b-IN TEXPRESERVE0 2b-W operations with only 20 transistors (20T). Additionally, novel PTL-based full-swing output half adders (10T-HA) and full adders (14T-FA) are proposed to construct the local adder tree, achieving reductions of 31.8%, 76.4%, and 41.4% in energy, delay, and area, respectively, compared to conventional adder trees in CIM macros. Simulations of the 288 kb SP-CIM macro demonstrated throughput and energy efficiency of 2.25 TOPS and 226.6 TOPS/W, respectively, at a 0.6 V supply voltage, and 2.97 TOPS and 154.1 TOPS/W, respectively, at a 0.8 V supply voltage, with 8b-IN, 8b-W, and 24b-OUT. Zhongzhen Tong, Sifan Sun, Chenghang Li, Jiye Yao, Yulong Qiu, Chao Wang 0094, Zhaohao Wang, Amara Amara, Xiaoyang Lin, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2024 | BSTCIM: A Balanced Symmetry Ternary Fully Digital In-MRAM Computing Macro for Energy Efficiency Neural NetworkabstractSilicon-based traditional binary computing in-memory (TBCIM) architectures are approaching their energy efficiency and throughput limits owing to challenges facing Moore’s Law. Thus, it is essential to explore architecture based on novel devices and computing paradigms to fulfill data-centric applications, such as artificial intelligence. In this paper, we propose a balanced symmetry ternary (BST) fully digital in-MRAM computing macro (BSTCIM) using hybrid voltage-gated spin-orbit torque magnetic tunnel junctions (VGSOT-MTJ) and gate-all-around carbon nanotube field-effect-transistors (GAA-CNTFET) technology. The overall computing is based on the highest efficiency multi-bit ternary system. BSTCIM includes a ternary dot product (TDP) unit with 4 GAA-CNTFETs and 2 VGSOT-MTJs achieving TDP operation without complex logic circuits. The multi-bit ternary multiply-and-accumulate (MAC) operation is realized through the proposed ternary adder tree and ternary post adder which accumulate TDP results within the digital domain enabling high accuracy neural network inference. Furthermore, due to the advantages of BST, ternary signed MAC is more easily performed compared to TBCIM macros that adapt 2’s complement or separate signed bit calculations. BSTCIM with 288 kb is simulated, achieving throughput and energy efficiency of 0.72 TOPS and 54.5 TOPS/W, respectively, at a 0.6 V supply voltage and 1.15 TOPS and 33.7 TOPS/W, respectively at a 0.8 V supply voltage with 8b-IN, 8b-W, and 20b-OUT. Moreover, the figure-of-merit for BSTCIM is 1.13–33.6 times higher than that of existing CIM macros. Zhongzhen Tong, Chenghang Li, Chao Wang 0094, Suteng Zhao, Qianyong Peng, Daming Zhou, Zhaohao Wang, Xiaoyang Lin, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2024 | A High Throughput In-MRAM-Computing Scheme Using Hybrid p-SOT-MTJ/GAA-CNTFETabstractSilicon-based semiconductor transistors are approaching their physical limits due to shrinking feature sizes. Simultaneously, traditional silicon-based von Neumann architectures exhibit significant latency and power consumption issues in data-centric applications, such as the Internet of Things and artificial intelligence. To tackle these challenges, this study introduces a novel approach: Magnetoresistance Random Access Memory (MRAM) computing in-memory (CIM) using gate-all-around carbon nanotube field-effect transistors (GAA-CNTFET). The proposed MRAM array comprised three transistors and one perpendicular magnetic anisotropy spin-orbit torque magnetic tunnel junction (p-SOT-MTJ) (3T1M) cell and achieves full-array Boolean logic operations and half/full-adder operations. The calculated results can be stored in-situ during the computing phase without requiring additional peripheral circuits. A 16 Kb MRAM was simulated in both GAA-CNTFET/p-SOT-MTJ and 14-nm FinFET/p-SOT-MTJ technologies to examine the effectiveness of the proposed design. Compared to its 14-nm FinFET/p-SOT-MTJ counterparts, the write and computing latencies of the GAA-CNTFET/p-SOT-MTJ CIM macro were reduced by approximately 21% and 20.6%, respectively, while the read and computing energy consumption by approximately 45.3% and 24.7%, respectively. Moreover, the proposed in-memory Boolean logic throughput was 8192 GOPS, which was approximately 160–250 times higher than that of existing CIM solutions, in which only two rows of word lines can be activated. Zhongzhen Tong, Yunlong Liu 0006, Xinrui Duan, Suteng Zhao, Chenghang Li, Zhi-Ting Lin, Xiulong Wu, Zhaohao Wang, Xiaoyang Lin |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2024 | A Computing In-Memory Multibit Multiplication Based on Decoupling and In-Array StoringabstractMultiplications are basic operations of neural networks. Therefore, multiplication results are crucial in analyzing the operating process of neural networks. However, the multiplication strategies are generally based on analog-domain circuits, and the results are in a multiply-and-accumulate (MAC) form. The result of each multiplication in MAC cannot be distinguished accurately using these strategies. Therefore, we proposed an in-memory multibit multiplication based on the decoupling and in-array storage strategy to overcome this problem, and the core module is the 10T1C SRAM cell. Multibit multiplications are decoupled by a series of logical operations. Therefore, in the analysis mode, multiplication results can be saved and outputted in the normal read mode without requiring additional storage. When executing the neural network, the operation results are stored in the cells. Hence, the operands stored in the array are retained. Accumulation operations are completed based on the charge-sharing technology; thus, the linearity of accumulation is high. We simulated and analyzed the performance of the proposed circuit in a 28 nm CMOS process. The absolute value of integral nonlinearity is at most 0.29. Further, due to high data operation parallelism, the throughputs of the logical operation and MAC are up to 6307.8 and 802.8 GOPS, respectively. Jin Zhang 0036, Zhongzhen Tong, Hao Wang 0239, Qiang Zhao 0007, Jiaqun Wang, Zhi-Ting Lin, Xiulong Wu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | A Fully Digital SRAM-Based Four-Layer In-Memory Computing Unit Achieving Multiplication Operations and Results StoreabstractThe separation of memory and arithmetic logic unit (ALU) in the von Neumann computing architecture hinders the development of big data and high-performance computing. In-memory computing (IMC) as a new computation method significantly reduces the latency and power consumption of data processing. In this study, we propose a fully digital static random access memory (SRAM)-based IMC architecture, which has the following advantages: 1) it simplifies multiplication to multicycle addition operations, reuses logic cells, and reduces hardware overhead; 2) by adding a pair of nMOS transistors to achieve internal write-back, the computational efficiency is improved, and at the same time, the final result of the multiplication can be stored locally, eliminating the need to read the computational result immediately; and 3) this scheme can be easily expanded to multiplication operations with different bit widths, which provides good scalability. A 4-kb SRAM-IMC macro chip is manufactured using the SMIC 55-nm technology to realize 4-bit multiplication, with an energy efficiency of 51.4 TOPS/W (0.9 V) and a throughput of 234.3 GOPS/mm2. The proposed multiplication–accumulation architecture is applied to a neural network, which achieves 98.7% accuracy with the Mixed National Institute of Standards and Technology database (MNIST) dataset. Zhi-Ting Lin, Shaoying Zhang, Jianping Xia, Yunwei Liu, Kefeng Yu, Zhongzhen Tong, Xiulong Wu, Wenjuan Lu, Chunyu Peng, Qiang Zhao 0007 |
IEEE Trans. Very Large Scale Integr. Syst. | 11 |
| 2023 | In-Memory Transposable Multibit Multiplication Based on Diagonal Symmetry Weight BlockabstractA possible approach to overcome the von Neumann bottleneck and meet the increasing demand for better computing performance is to computing in-memory (CIM). The results of the in-memory calculations are primarily reflected in the vertical bitline (BL) analog voltage. However, the nonlinearity of the BL discharge deteriorates with the increase in discharge voltage. In this study, we propose a diagonal symmetry weight block (DSWB) based on an eight-transistor (8T) static random access memory (SRAM) that can achieve multibit transposable operations. In addition, to guarantee linearity and complete multibit multiplication operations, we propose a cascode current mirror (CCM)-based multiplier. To achieve low-overhead and more efficient quantification, our proposed CIM macro uses a counter-type quantization circuit to read out the analog calculation results. We simulated the performance of the proposed 8T SRAM in a 28-nm complementary metal–oxide–semiconductor process. The integral nonlinearity (INL) of the proposed CCM-based CIM decreased by approximately 54.4% compared with the traditional CIM. Furthermore, the proposed in-memory multibit multiplication throughput density was 6.74 GOPS/kb; this throughput density improvement is approximately 3.3–10.5 times higher than the existing CIM works. Zhongzhen Tong, Yue Zhao 0029, Jin Zhang 0036, Zhi-Ting Lin, Xiaoyang Lin, Xiulong Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2022 | Configurable Memory With a Multilevel Shared Structure Enabling In-Memory ComputingabstractFrequent to-and-from data transfers in the von Neumann architecture limit the overall throughput. One of the promising approaches used to overcome von Neumann bottleneck is in-memory computing (IMC) that aims to embed computing in memory to reduce the transfer of memory-processor data. This study proposes a configurable 6-transistor (6T) static random access memory (SRAM) array with a multilevel shared structure for IMC. A multilevel shared structure can effectively improve the utilization rate of the module. In addition to the conventional SRAM operation, the configurable structure can also perform the sum of absolute differences (SAD) and Hamming distance (HD) calculations. To quickly identify the minimum value among multiple calculation results, a four-input sense amplifier (SA) is proposed. The performance of the proposed memory is simulated in a 65-nm CMOS process. The post-layout simulation results show good linearity of the multirow read in the SAD and HD modes. The mean time required by the four-input SA to obtain the result is 190 ps. The SAD and HD calculations yield consumptions of 67.44 fJ/byte and 0.64 fJ/bit, respectively, at 0.8 V. Furthermore, a single column-sharing comparator consumes 2.78 and 3.41 pJ at 0.8 V in the SAD and HD modes, respectively. Yue Zhao 0029, Zhi-Ting Lin, Xiulong Wu, Qiang Zhao 0007, Wenjuan Lu, Chunyu Peng, Zhongzhen Tong, Junning Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |