VLDB 2026 Research / reviewers in the wild / expert
Zhaohao Wang
dblp:89/9136
· DBLP profile ↗
61ranked-venue papers
11as first author
33since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 46 · 1 first-author · 27 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Theory of computation · 3 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Input Sparsity Aware In-Memory Computing Macro Based on SOT-MRAM Multi-Level Cell for Efficient Deep Neural Network AccelerationabstractDeep neural network (DNN) technology has gained widespread applications, but its high energy demands continue to drive the advancement of low-power computing architectures, particularly in in-memory computing (IMC) architectures based on non-volatile memory. Among these, spin-transfer torque magnetic random-access memory (STT-MRAM)-based IMC architectures have achieved some progress, but their performance remains constrained by limited resistance and binary characteristics. By contrast, the next-generation spin-orbit torque MRAM (SOT-MRAM) offers superior magnetic tunnel junction (MTJ) resistance and more flexible cell structures, presenting significant potential for energy-efficient IMC implementation. In this work, leveraging the ultra-high MTJ resistance and the separation of read/write paths in SOT-MRAM, we propose a multi-level cell (MLC) structure-based high energy-efficiency IMC architecture (MLC-SOT-IMC), which performs standard multiplication operations by optimizing the conductance mapping paradigm. The proposed architecture not only maintains high inference accuracy but also significantly enhances integration density and reduces the overhead per bit. Additionally, a self-terminating time-to-digital converter (TDC) readout circuit, which is dependent on input sparsity, is introduced to eliminate the excess power consumption associated with ineffective pulses after readout completion. Ultimately, the proposed MLC-SOT-IMC architecture achieves an inference energy efficiency of 6388.98 1-bit TOPS/W under an input sparsity of 50%, with the peak energy efficiency reaching 8426.19 1-bit TOPS/W at an input sparsity of 90%. Chao Wang 0094, Qihang Gao, Xianzeng Guo, Zhongzhen Tong, Zhaohao Wang, Weisheng Zhao 0001 |
DATE | 5 |
| 2026 | High-Performance and High-Density NAND-Like SOT-MRAM for FinFET Technology NodesabstractThis paper proposes a comprehensive optimization framework for NAND-like spintronics memory (NAND-SPIN) in advanced FinFET technology nodes. At bit-cell structure level, we propose a NAND-SPIN-GND design which is configured with a grounded bit line (BL) to minimize the parasitic resistance in both read and write paths, thereby decreasing read latency by 35.0% and write energy by 27.3%. At device and layout level, a short-circuiting bottom electrode (SBE) design is proposed, which shorts non-contributing spin-orbit torque (SOT) segments by the BEs, reducing read latency by up to 41.4% and write energy by 55.8%. In addition, a compact capacitance symmetric source line (SL)-type reference scheme is introduced to address the inherent capacitance asymmetry in conventional SL-type reference scheme, resulting in a 48.6% reduction in read latency compared to the conventional word line (WL)-type reference scheme. Chao Wang 0094, Xianzeng Guo, Luman Xiang, Zhaohao Wang, Weisheng Zhao 0001 |
DATE | 4 |
| 2026 | A Low Power and High Reliability Nonvolatile SRAM Using In-Plane VGSOT-MRAM with Pre-Charge Restore SchemeabstractConventional magnetic nonvolatile-static random access memory (MNV-SRAM) suffers from large store current, which leads to low area and energy efficiency, severely limiting their development and application. This paper proposes a 10T-2MTJ NV-SRAM cell based on the in-plane voltage-gated spin-orbit torque magnetic tunnel junction (VGSOT-MTJ), which enables field-free deterministic magnetization switching and reduces the required store current by leveraging the voltage-controlled magnetic anisotropy (VCMA) effect to assist the store operation. Thereby, the proposed design achieves the smallest SRAM cell area compared to prior works, due to the relaxed transistor drive strength requirement. On the other hand, existing NV-SRAM restore schemes exhibit a substantial deterioration in restore error ratio (RSER) with increasing MTJ resistance. Targeting the high resistance characteristics of VGSOT-MTJ, we innovatively propose a pre-charge restore scheme with sensitive transistor isolation. Simulation results demonstrate that the proposed design achieves the lowest read and write energy in SRAM mode, with store energy 1.58× to 2.48× lower than other in-plane MTJ-based designs. And the proposed restore scheme significantly improves restore reliability with over 98.7% RSER enhancement, and shows superior robustness across different MTJ resistances and tunnel magnetoresistance ratio (TMR) conditions. Zhongzhen Tong, Mingche Li, Weimeng Zhao, Zhongkui Zhang, Yaling Wang, Chao Wang 0094, Zhaohao Wang |
DATE | 9 |
| 2026 | Highly Energy-Efficient In-Memory Computing Architecture Based on VGSOT-MRAM for Reconfigurable BNN/TNN Acceleration
Qihang Gao, Chao Wang 0094, Chenghang Li, Zhongzhen Tong, Zhaohao Wang |
ISCAS | 5 |
| 2026 | A Post-Fabrication Tunable TDC-PUF Using VG-SOT MRAM for Robust Hardware Security
Tianrui Guo, Zhaohao Wang |
ISCAS | 4 |
| 2026 | A 6.86Tb/s Bandwidth SOT-MRAM Sensing Scheme with Configurable Full-Column Over Frequency Technique for Near Memory Computing
Xinpeng Jiang, Hanting Chen, Zhaohao Wang, He Zhang 0011, Weisheng Zhao 0001 |
ISCAS | 3 |
| 2026 | MEPUF: A Lightweight and ML-Resistant Strong PUF Integrating Dual-Mode MRAM and Configurable AES for Reliable UAVs
Bi Wang 0002, Luyao Shi, Chao Wang 0094, Zhaohao Wang |
ISCAS | 6 |
| 2026 | Multi-Retention and Bit-Level Approximate STT-MRAM for High-Efficiency AI Applications
Yulong Qiu, Chao Wang 0094, Weimeng Zhao, Zhongzhen Tong, Zhaohao Wang |
ISCAS | 5 |
| 2026 | A Fully-Parallel Digital MRAM Computing-in-Memory Macro Featuring a High-Efficient Dynamic Adder Tree and Bit-Splitting MAC
Zhongzhen Tong, Jiye Yao, Shaohui Ma, Yulong Qiu, Zhaohao Wang, Amara Amara, Xiaoyang Lin |
ISCAS | 5 |
| 2026 | Design and Implementation of a Scalable 64 p-bits Ising Computing Chip with Integrated SOT-MTJs for Efficient Computing
Jinhao Li 0007, Jialiang Yin, Linying Liu, Chengyuan Sun, Hong-Xi Liu, Kaihua Cao, Zhaohao Wang, Wenlong Cai, He Zhang 0011 |
ISCAS | 11 |
| 2026 | Constructing binary relations in interval-valued information tables using optimized methods
Zhuyun Dong, Zhaohao Wang |
Appl. Intell. | 2 |
| 2026 | A quality-configurable approximate cache design based on NAND-like SOT MRAM with high energy efficiency
Zhengyi Hou, Luyao Shi, Bi Wang 0002, Bi Wu 0002, Lirida A. B. Naviner, Zhaohao Wang |
Integr. | 7 |
| 2026 | BaM-CIM: A High Throughput Booth Algorithm-Based In-MRAM Computing Macro Using Hybrid VGSOT-MTJ/GAA-CNTFETabstractAs artificial intelligence (AI) and computational models grow in scale, the demand for computational power and storage has significantly increased. The computing-in-memory (CIM) architecture addresses this challenge by performing computations directly within the memory array, reducing data transfer between the processor and memory. This paper introduces a Booth algorithm-based In-MRAM computing architecture (BaM-CIM) using a hybrid voltage-gated spin-orbit torque MTJ (VGSOT-MTJ) and gate-all-around carbon nanotube field-effect transistors (GAA-CNTFETs) for efficient multiply-and-accumulate (MAC) computing. The key contributions of BaM-CIM are as follows: 1) A Voltage divider reference (VDR) cell is proposed, which enables read operations using only a 2T1M cell structure. Compared to complementary read cells, the VDR reduces the area by half and achieves robust data sensing without requiring precharge/discharge operations. 2) The BaM-CIM circuit is proposed to complete 8b-W/8b-IN/21b-OUT computations in only two cycles (1.6 ns), reducing the number of cycles by 75% compared to single-bit input serial operations and by 50% compared to two-bit serial operations. 3) A three-input 8b Booth computing adder (BCA), along with Modified computing shift adder (MCSA) and Modified computing post adder (MCPA), which can achieve higher energy efficiency. BaM-CIM with 128 Kb is simulated, achieving throughput and energy efficiency of 0.93 TOPS and 258.4 TOPS/W, respectively, at a 0.6 V supply voltage and 1.28 TOPS and 169.5 TOPS/W, respectively, at a 0.8 V supply voltage with 8b-IN, 8b-W, and 21b-OUT. Chenghang Li, Zhongzhen Tong, Yulong Qiu, Jiye Yao, Chao Wang 0094, Zhaohao Wang, Xiaoyang Lin, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2025 | An Adaptive Sparse Matrix Compression CIM Accelerator based on 256Kb SOT-MRAM for Downlink Massive MIMO CommunicationsabstractDownlink precoding in massive multiple input multiple output (MIMO) systems involves high-dimensional sparse matrix calculations, which poses challenges to existing architectures. Computing-in-memory (CIM) has significant advantages in handling large-scale parallel operations, but sparse computing for wireless communication remains underexplored. In this paper, we propose a novel CIM accelerator based on magnetic random access memory (MRAM) leveraging adaptive multi-sparse mode technology for optimized sparse matrix multiplication in MIMO communication systems. This architecture represents the first application of CIM technology for processing sparse matrices in MIMO precoding tasks, minimizing storage requirements and enhancing parallel processing speed. Experimental results demonstrate that, for a 32×256×8 MIMO downlink precoding task with 90% sparsity, the symbol error rate is reduced to 0.1% at a signal-to-noise ratio of 20dB, achieving 8.35× reduction in storage overhead, 39.4× power saving and 9.85× speedup. These results position our accelerator as a promising candidate for processing sparse data in 5G massive MIMO systems. Liangchen Li, Changyu Li, Anyang Yu, Junda Zhao, Zhaohao Wang, Chengyuan Sun, Kaihua Cao, Wang Kang 0001, He Zhang 0011, Weisheng Zhao 0001 |
ICCAD | 7 |
| 2025 | High-Speed VGSOT-MRAM Design for Non-Volatile Cache MemoriesabstractSpin-orbit torque magnetic random-access memory (SOT-MRAM) is a promising candidate for next-generation memory systems, particularly for cache applications, owing to its ultra-fast write speed and high endurance. However, SOTMRAM confronts challenges of large bit-cell area and high write current. Voltage-gated-SOT-MRAM (VGSOT-MRAM) mitigates these issues through the voltage controlled magnetic anisotropy (VCMA) mechanism, reducing the write current and enabling high-density device structure, but at the cost of slower read speed due to high device resistance. To address the read speed issue, we propose the local read bit-line (LRBL) scheme, which decreases the load capacitance of the read path and can reduce the read latency by 55.0% with minimal area overhead. Additionally, an efficient parallel-discharge-serial-sensing (PDSS) scheme is proposed to optimize the sequential read operations in cache, achieving up to 84.8% latency reduction during cache pre-fetch operations. Furthermore, implementing an appropriate error checking and correction (ECC) algorithm can further diminish the total read latency by 26.4%. Xianzeng Guo, Chao Wang 0094, Luman Xiang, Zhaohao Wang |
ISCAS | 4 |
| 2025 | A High-Gain Three-Stage Auto-Zeroing Residual Amplifier for High-Precision Pipelined SAR ADCabstractThis paper proposes a residual amplifier (RA) with high open-loop gain, wide output swing, and integrated auto-zeroing (AZ) functionality. To satisfy the stringent relative gain error requirements of high-precision successive-approximation-register (SAR) analog-to-digital converters (ADCs), the RA employs a three-stage architecture to achieve enhanced open-loop gain while maintaining a broad output voltage range. Stability in the multi-stage design is ensured through the strategic placement of the second and third poles using the complex-pole method. The input stage incorporates a current-reuse technique to double the effective trans-conductance, reducing thermal noise and extending bandwidth without increasing static power consumption. Additionally, an AZ technique is integrated to suppress offset voltage and mitigate low-frequency 1/f noise. Simulated in a 40nm CMOS process, the opamp achieves an open-loop gain exceeding 130dB, and a loop gain of over 101 dB when configured as 61.5× switched-capacitor (SC) RA. The design delivers a bandwidth of 30 MHz, a phase margin greater than 60°, and an input-reference total noise of 14.8 μVrms, with a power consumption of 5.6 mW. These results demonstrate the RA’s capability to meet the demands of high-resolution pipelined SAR ADCs, combining precision, dynamic performance, and power efficiency. Renjie Fu, Yiqin Chen, Hongjie Ye, Bi Wang 0002, Zhaohao Wang |
ISCAS | 6 |
| 2025 | Approximate SOT-MRAM for Neural Network Acceleration with Superior Read PerformanceabstractMagnetoresistive random-access memory (MRAM) has been demonstrated to be a suitable memory technology for neural network (NN) acceleration due to its non-volatility, high density, and fast access speed. However, compared to the widely used static random-access memory (SRAM), MRAM still exhibits a notable disparity in speed and energy. In this paper, we propose a read-related approximation computation (RAC) strategy based on spin-orbit torque MRAM (SOT-MRAM) to enhance the computational speed of NN, and then we introduce a reference reconfigurable array (RRA) architecture to further decrease the read latency and energy consumption, significantly improving the speed and energy efficiency of weight retrieval during computations. Furthermore, we propose an algorithm to verify and optimize NN model performance. The proposed architecture is evaluated using a 28 nm process combined with a SPICE model of the SOT-MRAM. Simulation results indicate that the read speed increases by 4.71X, the read energy consumption is reduced by 70.7%, while the model accuracy loss remains below 1%. Yulong Qiu, Chao Wang 0094, Zhongzhen Tong, Siyuan Cheng 0021, Zhaohao Wang |
ISCAS | 6 |
| 2025 | Critical current for field-free switching of the in-plane magnetization in the three-terminal magnetic tunnel junction
Hongjie Ye, Zhengjie Yan, Zhaohao Wang |
Sci. China Inf. Sci. | 4 |
| 2025 | A Self-Decryption Pass Transistor Logic-Based In-MRAM Computing Macro Using Hybrid VGSOT-MTJ/GAA-CNTFETabstractSpintronic devices and gate-all-around carbon nanotube field-effect-transistors (GAA-CNTFETs)-based computing in-memory architecture are competitive candidates for applications in battery-powered tiny artificial intelligence (AI) edge devices. Meanwhile, data encryption and decryption are also necessary to protect AI model weights and the customized data used to guarantee neural network (NN) inference accuracy. In this study, we propose a self-decryption pass transistor logic (PTL)-based in-MRAM computing macro (SP-CIM) that utilizes hybrid voltage-gated spin-orbit torque magnetic tunnel junctions (VGSOT-MTJ)/GAA-CNTFET. The proposed SP-CIM macro enables simultaneous data access, decryption, and full-accuracy multiply-and-accumulate (MAC) operations using the newly introduced voltage-divider self-decryption cell, without the need for additional decryption logic. Compared to existing in-memory decryption strategies, this design reduces energy consumption by 45.7% and decreases decryption delay by 87.2%. To enhance area efficiency and reduce computing latency, we propose a PTL-based multiplication cell that achieves full-accuracy local 2b-IN TEXPRESERVE0 2b-W operations with only 20 transistors (20T). Additionally, novel PTL-based full-swing output half adders (10T-HA) and full adders (14T-FA) are proposed to construct the local adder tree, achieving reductions of 31.8%, 76.4%, and 41.4% in energy, delay, and area, respectively, compared to conventional adder trees in CIM macros. Simulations of the 288 kb SP-CIM macro demonstrated throughput and energy efficiency of 2.25 TOPS and 226.6 TOPS/W, respectively, at a 0.6 V supply voltage, and 2.97 TOPS and 154.1 TOPS/W, respectively, at a 0.8 V supply voltage, with 8b-IN, 8b-W, and 24b-OUT. Zhongzhen Tong, Sifan Sun, Chenghang Li, Jiye Yao, Yulong Qiu, Chao Wang 0094, Zhaohao Wang, Amara Amara, Xiaoyang Lin, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2025 | MS-SCIM: A Mixed-Signal Stochastic Computing-in-Memory Paradigm for Information SecurityabstractStochastic Computing (SC), an emerging paradigm with advantages in hardware cost and fault tolerance, is well-suited for applications in image processing and information security. However, the existing SC paradigms suffer from significant hardware costs due to the conversion between binary and stochastic sequences, and the long latency caused by low computational parallelism. In this work, we demonstrate a Mixed-Signal Stochastic Computing-In-Memory (MS-SCIM) paradigm utilizing spin orbit torque magnetic random access memory (SOT-MRAM) arrays, in order to realize a energy-efficient and conversion-less SC method for the first time. The main contributions include: 1) The inherent stochastic switching behaviors of spintronic devices are exploited to enable the SOT-MRAM array to serve both as a parallel true random number generator (TRNG) and a CIM cell. 2) A high-parallelism mixed signal stochastic CIM paradigm is proposed to accelerate SC-based edge detection algorithm. The whole process achieves binary outputs without the conversion circuits and the energy efficiency achieves 446 Tops/W. 3) Based on the results of MS-SCIM, a novel image steganography method using stochastic bit streams is leveraged for information security, which enables lossless embedding and extraction of secret image information in a$156\times 156$size, enhancing both capacity and undetectability. Pengxu Wang, Yijiao Wang, Jialiang Yin, Jiayao Wu, Xinrui Duan, Zhaohao Wang, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2025 | Technically Feasible Robust Complementary SOT-MRAM Design for Improving the Area and Energy EfficiencyabstractSpin-orbit torque magnetic random-access memory (SOT-MRAM), which exhibits sub-nanosecond write speed and high endurance, is a promising candidate for the future high-level cache. Nevertheless, SOT-MRAM faces challenge in meeting the high read performance requirements of cache applications due to the limited ON/OFF ratio. Consequently, extensive investigation has been conducted into robust complementary bit-cell (CBC) designs based on SOT-MRAM. However, previous designs suffer from significant technology feasibility, area and performance issues. In this paper, the feasibility and performance of the existing complementary write schemes are analyzed, and optimized U-type and toggle spin torque (TST) schemes with practicality and conciseness are presented. The previous CBC designs are evaluated and optimized in terms of circuit and layout, while the 1-word-line-3-bit-line (1WL3BL) CBC designs with both U-type and TST schemes are proposed, which can reduce the bit-cell area by 24.64%-27.54% and improve the write and read performance. In comparison to the conventional CBC design, the proposed 1WL3BL CBC design can reduce the write energy and read latency by up to 36.91% and 21.93%, respectively. Furthermore, the proposed low-voltage read scheme demonstrates the capability to enhance the read performance and conserve the read energy under the aggressive read-related process parameters. Chao Wang 0094, Zhongkui Zhang, Xianzeng Guo, Qihang Gao, Zhaohao Wang, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2024 | Series-Parallel Hybrid SOT-MRAM Computing-in-Memory Macro with Multi-Method Modulation for High Area and Energy EfficiencyabstractComputing-in-memory (CIM) shows its superiority in lots of applications like neural network inference. Recently, there are lots of exploration of the application of Magnetic Random-Access Memory (MRAM) in CIM. This paper aims to investigate the potential of Spin-Orbit-Torque-MRAM (SOT-MRAM) in CIM and proposes a high area and energy efficiency SOT-MRAM CIM macro based on a 6T-4J weight group. The bit-cell array adopts series-parallel hybrid architecture, which combines both serial and parallel configurations of Magnetic Tunnel Junction (MTJ) to solve the problem of high energy cost and low flexibility caused by MRAM-series and MRAM-parallel architecture, respectively. Additionally, the proposed SOT-MRAM CIM macro incorporates a multi-method modulation scheme, ranging from input unit to array, which meanwhile allows for configurable input precision (2/4/6/8-bit). The SOT-MRAM CIM macro is designed and verified in both 180nm and 28nm nodes, based on the verified electrical performance of the SOT-MRAM array in a 200-nm wafer pre-fabricated. The simulation results in 28nm show that this macro can achieve energy efficiency of 23.7~29.6 Tops/W at 8-bit input and output precision. Weiliang Huang, Jinyu Bai, Wang Kang 0001, Zhaohao Wang, Kaihua Cao, He Zhang 0011, Weisheng Zhao 0001 |
DAC | 4 |
| 2024 | Quasi-atomic relations based rough set model and convex geometry
Zhaohao Wang |
Appl. Intell. | 1 |
| 2024 | BSTCIM: A Balanced Symmetry Ternary Fully Digital In-MRAM Computing Macro for Energy Efficiency Neural NetworkabstractSilicon-based traditional binary computing in-memory (TBCIM) architectures are approaching their energy efficiency and throughput limits owing to challenges facing Moore’s Law. Thus, it is essential to explore architecture based on novel devices and computing paradigms to fulfill data-centric applications, such as artificial intelligence. In this paper, we propose a balanced symmetry ternary (BST) fully digital in-MRAM computing macro (BSTCIM) using hybrid voltage-gated spin-orbit torque magnetic tunnel junctions (VGSOT-MTJ) and gate-all-around carbon nanotube field-effect-transistors (GAA-CNTFET) technology. The overall computing is based on the highest efficiency multi-bit ternary system. BSTCIM includes a ternary dot product (TDP) unit with 4 GAA-CNTFETs and 2 VGSOT-MTJs achieving TDP operation without complex logic circuits. The multi-bit ternary multiply-and-accumulate (MAC) operation is realized through the proposed ternary adder tree and ternary post adder which accumulate TDP results within the digital domain enabling high accuracy neural network inference. Furthermore, due to the advantages of BST, ternary signed MAC is more easily performed compared to TBCIM macros that adapt 2’s complement or separate signed bit calculations. BSTCIM with 288 kb is simulated, achieving throughput and energy efficiency of 0.72 TOPS and 54.5 TOPS/W, respectively, at a 0.6 V supply voltage and 1.15 TOPS and 33.7 TOPS/W, respectively at a 0.8 V supply voltage with 8b-IN, 8b-W, and 20b-OUT. Moreover, the figure-of-merit for BSTCIM is 1.13–33.6 times higher than that of existing CIM macros. Zhongzhen Tong, Chenghang Li, Chao Wang 0094, Suteng Zhao, Qianyong Peng, Daming Zhou, Zhaohao Wang, Xiaoyang Lin, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 9 |
| 2024 | A High Throughput In-MRAM-Computing Scheme Using Hybrid p-SOT-MTJ/GAA-CNTFETabstractSilicon-based semiconductor transistors are approaching their physical limits due to shrinking feature sizes. Simultaneously, traditional silicon-based von Neumann architectures exhibit significant latency and power consumption issues in data-centric applications, such as the Internet of Things and artificial intelligence. To tackle these challenges, this study introduces a novel approach: Magnetoresistance Random Access Memory (MRAM) computing in-memory (CIM) using gate-all-around carbon nanotube field-effect transistors (GAA-CNTFET). The proposed MRAM array comprised three transistors and one perpendicular magnetic anisotropy spin-orbit torque magnetic tunnel junction (p-SOT-MTJ) (3T1M) cell and achieves full-array Boolean logic operations and half/full-adder operations. The calculated results can be stored in-situ during the computing phase without requiring additional peripheral circuits. A 16 Kb MRAM was simulated in both GAA-CNTFET/p-SOT-MTJ and 14-nm FinFET/p-SOT-MTJ technologies to examine the effectiveness of the proposed design. Compared to its 14-nm FinFET/p-SOT-MTJ counterparts, the write and computing latencies of the GAA-CNTFET/p-SOT-MTJ CIM macro were reduced by approximately 21% and 20.6%, respectively, while the read and computing energy consumption by approximately 45.3% and 24.7%, respectively. Moreover, the proposed in-memory Boolean logic throughput was 8192 GOPS, which was approximately 160–250 times higher than that of existing CIM solutions, in which only two rows of word lines can be activated. Zhongzhen Tong, Yunlong Liu 0006, Xinrui Duan, Suteng Zhao, Chenghang Li, Zhi-Ting Lin, Xiulong Wu, Zhaohao Wang, Xiaoyang Lin |
IEEE Trans. Circuits Syst. I Regul. Pap. | 10 |
| 2024 | Variation Aware Evaluation Approach and Design Methodology for SOT-MRAMabstractSpin-orbit torque magnetic random access memory (SOT-MRAM), which exhibits sub-nanosecond write speed and high reliability, is a promising candidate for the future high-level cache. However, SOT-MRAM faces the problem of large bit-cell layout area due to its structural characteristics and write performance requirements, therefore it is necessary to explore the bit-cell design with optimal overall performance under the unified bit-cell area. In this paper, we propose a comprehensive variation aware evaluation approach for the area, latency, and energy of SOT-MRAM under the uniform yield standard. Based on this, the mainstream SOT-MRAM bit-cell designs with high-density method and multi-finger configuration are evaluated, meanwhile bit-cell designs with excellent write performance and their optimum area ranges are identified. Moreover, the source line read (SLR) mode with higher robustness against transistor variation is proposed to improve the read performance, and the dual SL (DSL) method is proposed to further reduce the read latency and write energy. With the DSL method, the read latency and write energy of 2-word-line (WL)-type bit-cells can be reduced by up to 36.5% and 12.6%, respectively. In addition, the DSL method can solve the shunt current issue of 1WL-type bit-cells and reduce the read latency and write energy by up to 43.6% and 17.4%, respectively. Chao Wang 0094, Zhaohao Wang, Shixing Li, Zhongkui Zhang, Youguang Zhang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | The granulation attribute reduction of multi-label data
Zhaohao Wang |
Appl. Intell. | 1 |
| 2023 | NAND-SPIN-based processing-in-MRAM architecture for convolutional neural network acceleration
Yinglin Zhao, Jianlei Yang 0001, Bing Li 0017, Xingzhou Cheng, Xucheng Ye, Xiaotao Jia, Zhaohao Wang, Youguang Zhang, Weisheng Zhao 0001 |
Sci. China Inf. Sci. | 8 |
| 2023 | Layout Aware Optimization Methodology for SOT-MRAM Based on Technically Feasible Top-Pinned Magnetic Tunnel Junction ProcessabstractThe emerging spin-orbit torque magnetic random-access memory (SOT-MRAM) shows promising prospects in high-level cache applications due to its subnanosecond switching speed and high reliability. However, SOT-MRAM faces the issue of large bit-cell layout area, which is currently the focus of attention. Although many design and evaluation works have emerged, the lack of a unified standard for realistic SOT process has hindered the development of relevant research toward practicality. In this article, the bit-cell area of the SOT-MRAM will be evaluated and optimized based on the technically feasible process. First of all, based on the state-of-the-art top-pinned SOT nanopillar process, the SOT-MRAM design rules are proposed. On this basis, this article systematically summarizes four basic device layout modes and provides optimized layout suggestions for conventional SOT bit-cells with different types and sizes of devices. In addition, a series of area-efficient SOT bit-cell designs based on the common area (CA) and dual common (DC) solutions are proposed, which can reduce the layout area of SOT bit-cells by up to 38.4% with reasonable write latency and energy overhead. Chao Wang 0094, Zhaohao Wang, Zhongkui Zhang, Jiagao Feng, Youguang Zhang, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | IMGA: Efficient In-Memory Graph Convolution Network Aggregation With Data Flow OptimizationsabstractAggregating features from neighbor vertices is a fundamental operation in graph convolution network (GCN). However, the sparsity in graph data creates poor spatial and temporal locality, causing dynamic and irregular memory access patterns and limiting the performance of aggregation on the Von Neumann architecture. The emerging processing-in-memory (PIM) architecture is based on emerging nonvolatile memory (NVM), like spin-orbit torque magnetic RAM (SOT-MRAM), and demonstrates promising prospects in alleviating the Von Neumann bottleneck. However, the limited memory capacity of PIM medium still incurs non-negligible data movements between PIM architecture and external memory. To solve this challenge, we propose an SOT-MRAM-based in-memory computing architecture, called IMGA, for efficient in-situ graph aggregation. Specifically, we design adaptive data flow management strategies that reuse vertex data in MRAM when processing graphs of different scales and adopt edge data as the control signal source to utilize the graph’s structural information. A reordering optimization strategy leveraging hardware–software co-design principle is proposed to further reduce the costly data movement. Experimental results demonstrate that IMGA achieves an average$2523\times $and$21\times $speedup, and 1.03E+6 and 1.04E+3 energy efficiency compared with CPU and GPU, respectively. Yuntao Wei, Shangtong Zhang, Jianlei Yang 0001, Xiaotao Jia, Zhaohao Wang, Gang Qu 0001, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2022 | Reconfigurable and Dynamically Transformable In-Cache-MPUF System With True Randomness Based on the SOT-MRAMabstractIn this paper, we present a reconfigurable Physically Unclonable Functions (PUF) based on the Spin-Orbit-Torque Magnetic Random-Access Memory (SOT-MRAM), which exploits thermal noise as the true dynamic entropy source. Therefore, the MRAM cells could be configured to random final states with stochastic switching mechanism. The proposed PUF is constructed and reconfigured by combining the small-capacity true random number generator (TRNG) and high-reliability secure hash algorithm (SHA-512), realizing the dynamic transformation between SOT-MRAM based last level cache and PUF (In-Cache-MPUF). Thanks to the full reconfigurability and the high endurance of SOT-MRAM, the proposed In-Cache-MPUF can achieve$10^{\textbf {14}}$maximum PUF bits per cell, which has greatly motivated the implementations compared with the traditional weak PUFs utilizing the static entropy source of process variations. The Monte-Carlo simulation results using 40 nm technology and a compact MTJ model show that the proposed PUF has desirable randomness as the digitized bit streams passing all the NIST tests, achieving 50.0428% uniqueness as well as 49.9236% uniformity. It also shows comparable reliability to the state-of-the-art works: a maximum bit error rate of 0.14% and 0.12% at 100 °C and 0.9 V, respectively. In addition, the system level performance is tested and validated by gem5. Zhengyi Hou, Zhaohao Wang, Chao Wang 0094, Min Wang 0033, You Wang 0002, Cenlin Duan, Jianlei Yang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | Computing-in-Memory Paradigm Based on STT-MRAM with Synergetic Read/Write-Like ModesabstractWith the surge in demand for data storage and processing in emerging applications, the traditional CMOS-based Von-Neumann architecture is facing challenges such as memory wall and static power consumption. In order to conquer the above-mentioned bottlenecks in computing systems, computing in-memory (CiM) architectures based on non-volatile memory (NVM) have been widely researched. In this paper, we propose a CiM paradigm based on spin-transfer torque magnetic random access memory (STT-MRAM), which combines common read-like mode (RLM) and write-like mode (WLM). On the basis of realizing the basic functions AND/OR/NAND/NOR, our design coordinates the high speed of RLM and the integrity of WLM to perform complex operations like full-adder (FA) and XOR/XNOR. In addition, the high speed and low power consumption of the proposed CiM paradigm are established by circuit-level simulation with a 40 nm design kit. Chao Wang 0094, Zhaohao Wang, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 2 |
| 2021 | A New Description of Transversal Matroids Through Rough Set ApproachabstractMatroid theory is a useful tool for the combinatorial optimization issue in data mining, machine learning and knowledge discovery. Recently, combining matroid theory with rough sets is becoming interesting. In this paper, rough set approaches are used to investigate an important class of matroids, transversal matroids. We first extend the concept of upper approximation number functions in rough set theory and propose the notion of generalized upper approximation number functions on a set system. By means of the new notion, we give some necessary and sufficient conditions for a subset to be a partial transversal of a set system. Furthermore, we obtain a new description of a transversal matroid by the generalized upper approximation number function. We show that a transversal matroid can be induced by the generalized upper approximation number function which can be decomposed into the sum of some elementary generalized upper approximation number functions. Conversely, we also prove that a generalized upper approximation number function can induce a transversal matroid. Finally, we apply the generalized upper approximation number function to study the relationship among transversal matroids. Zhaohao Wang |
Fundam. Informaticae | 1 |
| 2020 | PRISM: Energy-Efficient Polymorphic Operation Based on Spin-Orbit Torque Memory for Reconfigurable ComputingabstractEmerging Non-Volatile Memories (NVMs) including resistive RAM (ReRAM), phase-change memory (PCM), and magnetic RAM (MRAM), have opened up new pathways for the NVM-based reconfigurable computing. Those NVMs technologies can achieve significant energy-efficient computational operations with only minor modification of the peripheral circuits. However, the supported operations are limited by the array structure and low energy-efficiency of implementing the computation using the memory array. In this paper, the Spin Orbit torque-MRAM based polymorphic circuits are proposed to support the reconfigurable computation for reducing the power consumption and improving the functionalities of the single memory array. With the high speed and energy-efficiency write operation, the proposed memory array support both read-out and write-in reconfigurable operations. Liang Chang 0002, Zhaohao Wang, Youguang Zhang, Weisheng Zhao 0001, Jun Zhou 0017 |
ISCAS | 2 |
| 2020 | Computing-in-Memory Architecture Based on Field-Free SOT-MRAM with Self-Reference MethodabstractOn the current computing platforms, the memory wall between processor and memory has become the toughest challenge for the traditional Von-Neumann computer architecture. Computing-in-Memory (CIM) is taken as a promising approach to solving the above bottleneck in computing systems. In this paper, we propose a CIM platform with field-free spinorbit torque magnetic random access memory (SOT-MRAM). The self-reference (SelfRef) method is designed to enhance the read reliability and directly obtain logic results through memory-like read operations without adding logic cells. Memory read/write and logic operations, including NOT, AND/NAND and OR/NOR, can be implemented in the same SOT-MRAM chip. The speed and power penalties caused by SelfRef scheme are acceptable thanks to the ultrafast switching of the SOT. The read reliability and logic correctness of the proposed CIM are demonstrated by hybrid simulation on a 40 nm technology node. Chao Wang 0094, Zhaohao Wang, Jianlei Yang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 2 |
| 2020 | A Comparative Cross-layer Study on Racetrack Memories: Domain Wall vs SkyrmionabstractRacetrack memory (RM), a new storage scheme in which information flows along a nanotrack, has been considered as a potential candidate for future high-density storage device instead of hard disk drive (HDD). The first RM technology, which was proposed in 2008 by IBM, relies on a train of opposite magnetic domains separated by domain walls (DWs), named DW-RM. After 10 years of intensive research, a variety of fundamental advancements has been achieved; unfortunately, no product has been available until now. With increasing effort and resources dedicated to the development of DW-RM, it is likely that new materials and mechanisms will soon be discovered for practical applications. However, new concepts might also be on the horizon. Recently, an alternative information carrier, magnetic skyrmion, which was experimentally discovered in 2009, has been regarded as a promising replacement of DW for RM, named skyrmion-based RM (SK-RM). Intensive effort has been involved and amazing advances have been made in observing, writing, manipulating, and deleting individual skyrmions. So, what is the relationship between DW and skyrmion? What are the key differences between DW and skyrmion, or between DW-RM and SK-RM? What benefits could SK-RM bring and what challenges need to be addressed before application? In this review article, we intend to answer these questions through a comparative cross-layer study between DW-RM and SK-RM. This work will provide guidelines, especially for circuit and architecture researchers on RM. Wang Kang 0001, Bi Wu 0002, Xing Chen 0012, Daoqian Zhu, Zhaohao Wang, Xichao Zhang, Youguang Zhang, Weisheng Zhao 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 5 |
| 2020 | The uncertainty measures for covering rough set models
Zhaohao Wang, Jianping Deng |
Soft Comput. | 1 |
| 2020 | A Novel High Performance and Energy Efficient NUCA Architecture for STT-MRAM LLCs With Thermal ConsiderationabstractAs the speed gap of the modern processor and the off-chip main memory enlarges, on-chip cache capacity increases to sustain the performance scaling. As a result, the cache power occupies a large portion of the total power budget. Spin transfer torque magnetic memory (STT-MRAM) is proposed as a promising solution for the low power cache design due to its high integration density and ultralow leakage power. Nevertheless, the high write power and latency of STT-MRAM become new barriers for the commercialization of this emerging technology. In this paper, we investigate the thermal effect on the access performance of STT-MRAM, and observe that the temperature can affect the write delay and energy significantly. Then, we explore the nonuniform cache access (NUCA) design of the chip-multiprocessors with STT-MRAM-based last level cache (LLC). A thermal aware data migration policy, called “Thermosiphon,” which takes advantage of the thermal property of STT-MRAM, is proposed to reduce the LLC write energy. This policy splits the LLC into different regions dynamically based on the thermal distribution monitored by thermal sensors available on-chip, and adaptively migrates write intensive data among different thermal regions considering the thermal gradient. Compared to the conventional NUCA design, our proposed design can save 41.2% write energy at most and 13.01% on average with negligible hardware overhead. Bi Wu 0002, Pengcheng Dai, Yuanqing Cheng, Ying Wang 0001, Jianlei Yang 0001, Zhaohao Wang, Dijun Liu, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2019 | CORN: In-Buffer Computing for Binary Neural NetworkabstractBinary Neural Networks (BNNs) have obtained great attention since they reduce memory usage and power consumption as well as achieve a satisfying recognition accuracy on Image Classification. In particular to the computation of BNNs, the multiply-accumulate operations of convolution-layer are replaced with the bit-wise operations (XNOR and pop-count). Such bit-wise operations are well suited for the hardware accelerator such as in-memory computing (IMC). However, an additional digital processing unit (DPU) is required for the pop-count operation, which induces considerable data movement between the Process Engines (PEs) and data buffers reducing the efficiency of the IMC. In this paper, we present a BNN computing accelerator, namely CORN, which consists of a Spin-Orbit-Torque Magnetic RAM (SOT-MRAM) based data buffer to perform the majority operation (to replace the pop-count process) with the SOT-MRAM-based IMC to accelerate the computing of BNNs. CORN can naturally implement the XNOR operation in the NVM memory array, and feed results to the computing data buffer for the majority write operation. Such a design removes the pop-counter implemented by the DPU and reduces data movement between the data buffer and the memory array. Based on the evaluation results, CORN achieves 61% and 14% power saving with 1.74× and 2.12× speedup, compared to the FPGA and DPU based IMC architecture, respectively. Liang Chang 0002, Zhaohao Wang, Youguang Zhang, Weisheng Zhao 0001, Yuan Xie 0001 |
DATE | 3 |
| 2019 | An Uncertainty Measure Based on Lower and Upper Approximations for Generalized Rough set ModelsabstractUncertainty measures are an important tool for analyzing data. There is the uncertainty of a rough set caused by its boundary region in rough set models. Thus the uncertainty measurement issue is also an important topic for rough set theory. Shannon entropy has been introduced into rough set theory. However, there are relatively few studies on the uncertainty measure in generalized rough set models. We know that the boundary region of a rough set is closely related to the upper and lower approximations in rough set models. In this paper, from the viewpoint of the upper and lower approximations, we propose new uncertainty measures, the upper rough entropy and the lower rough entropy, in generalized rough set models. Then we focus on the investigations of the upper rough entropy, and give the concepts of the upper joint entropy, the upper conditional entropy and the mutual information with respect to a general binary relation. Some important properties of these measures are obtained. The connections among these measures are given. Furthermore, comparing with the existing uncertainty measures, the upper rough entropy has high distinguishing degree. Theoretical analysis and experimental results show that the proposed entropy is better effective than some existing measures. Zhaohao Wang, Huifang Yue, Jianping Deng |
Fundam. Informaticae | 1 |
| 2019 | The lattice and matroid representations of definable sets in generalized rough sets based on relations
Zhaohao Wang, Qinrong Feng |
Inf. Sci. | 1 |
| 2019 | Approximation via a double-matroid structure
Huangjian Yi, Zhaohao Wang |
Soft Comput. | 3 |
| 2019 | The structures and the connections on four types of covering rough sets
Zhaohao Wang, Qinrong Feng |
Soft Comput. | 1 |
| 2019 | Exploiting Spin-Orbit Torque Devices As Reconfigurable Logic for Circuit ObfuscationabstractCircuit obfuscation is a frequently used approach to conceal logic functionalities in order to prevent reverse engineering attacks on fabricated chips. Efficient obfuscation implementations are expected with lower design complexity and overhead but higher attack difficulties. In this paper, an emerging obfuscation approach is proposed by leveraging spin-orbit torque (SOT) devices-based look-up-tables as reconfigurable logic to replace the carefully selected gates. It is essentially impossible to identify the obfuscated gate with SOTs inside according to the physical geometry characteristics because the configured functionalities are represented by magnetization states. Such an obfuscation approach makes the circuit security further improved with high exponential attack complexities. Experiments on MCNC and ISCAS 85/89 benchmark suits show that the proposed approach could reduce the area overheads due to obfuscation by 10% averagely. Jianlei Yang 0001, Qiang Zhou 0001, Zhaohao Wang, Hai Li 0001, Yiran Chen 0001, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | DASM: Data-Streaming-Based Computing in Nonvolatile Memory Architecture for Embedded SystemabstractEmerging nonvolatile memories (NVMs), including resistive RAM (RRAM), phase-change memory (PCM), and magnetic RAM (MRAM), have opened up new pathways for Computing-In-Memory (CIM). Those NVM technologies can achieve energy-efficient computational operations with only minor modification of the peripheral circuits. Despite many advantages provided by computational NVMs, parallelism is not sufficiently explored in such CIM designs. To break through this limitation on performance gain, we propose a data-streaming design for the NVM-based CIM (e.g., DASM) by leveraging the underlying parallelism in the hardware. DASM benefits from the massive parallelism of data-streaming computing, reduction in data movement of the CIM, and the nonvolatility of memory arrays. Specifically, data streaming operations can be implemented with CIM bitwise operations in both read-out and write-in procedures. In addition, we use the multilevel power gating for the memory array and connections to further boost the performance. Finally, we study a case of inference process for the quantized deep-neural-network-based on the DASM design. DASM architecture achieves 47.8×, 5.1×, 2.1× speedup compared to the NVIDIA Jetson TK1 embedded GPU board, Intel Xeon E5-2640 CPU, the state-of-the-art field-programmable gate array (FPGA) design, with much lower power consumption. Liang Chang 0002, Zhaohao Wang, Youguang Zhang, Yufei Ding 0001, Weisheng Zhao 0001, Yuan Xie 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | PXNOR-BNN: In/With Spin-Orbit Torque MRAM Preset-XNOR Operation-Based Binary Neural NetworksabstractConvolution neural networks (CNNs) have demonstrated superior capability in computer vision, speech recognition, autonomous driving, and so forth, which are opening up an artificial intelligence (AI) era. However, conventional CNNs require significant matrix computation and memory usage leading to power and memory issues for mobile deployment and embedded chips. On the algorithm side, the emerging binary neural networks (BNNs) promise portable intelligence by replacing the costly massive floating-point compute-andaccumulate operations with lightweight bit-wise XNOR and popcount operations. On the hardware side, the computingin-memory (CIM) architectures developed by the non-volatile memory (NVM) present outstanding performance regarding high speed and good power efficiency. In this paper, we propose an NVM-based CIM architecture employing a Preset-XNOR operation in/with the spin-orbit torque magnetic random access memory (SOT-MRAM) to accelerate the computation of BNNs (PXNOR-BNN). PXNOR-BNN performs the XNOR operation of BNNs inside the computing-buffer array with only slight modifications of the peripheral circuits. Based on the layer evaluation results, PXNOR-BNN can achieve similar performance compared with the read-based SOT-MRAM counterpart. Finally, the end-to-end estimation demonstrates 12.3× speedup compared with the baseline with 96.6-image/s/W throughput efficiency. Liang Chang 0002, Zhaohao Wang, Youguang Zhang, Yuan Xie 0001, Weisheng Zhao 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Spintronics based stochastic computing for efficient Bayesian inference systemabstractBayesian inference is an effective approach for solving statistical learning problems especially with uncertainty and incompleteness. However, inference efficiencies are physically limited by the bottlenecks of conventional computing platforms. In this paper, an emerging Bayesian inference system is proposed by exploiting spintronics based stochastic computing. A stochastic bitstream generator is realized as the kernel components by leveraging the inherent randomness of spintronics devices. The proposed system is evaluated by typical applications of data fusion and Bayesian belief networks. Simulation results indicate that the proposed approach could achieve significant improvement on inference efficiencies in terms of power consumption and inference speed. Xiaotao Jia, Jianlei Yang 0001, Zhaohao Wang, Yiran Chen 0001, Hai Li 0001, Weisheng Zhao 0001 |
ASP-DAC | 3 |
| 2018 | Progresses and challenges of spin orbit torque driven magnetization switching and application (Invited)abstractSpin orbit torque (SOT) has been proposed as a potential alternative mechanism to the conventional spin transfer torque (STT) for the magnetization switching. Recently, theoretical and experimental works revealed the novel factors influencing the SOT-driven magnetization switching. Emerging SOT-based spintronics memories and circuits were explored to implement fast and energy-efficient write operation. However, the perspective of the SOT mechanism is still challenged by some serious shortcomings, such as area penalty, relatively large switching current density and undesirable use of external magnetic field. Here, we review the progresses in the SOT mechanism involving the magnetization dynamics, device design and circuit development. Key issues to be addressed in optimizing the SOT devices are pointed out. In particular, we discuss the potential solutions to develop high-density SOT-based memories and circuits. Zhaohao Wang, Zuwei Li, Liang Chang 0002, Wang Kang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 1 |
| 2018 | Radiation hardening design for spin-orbit torque magnetic random access memoryabstractAlthough the magnetic tunnel junction (MTJ) is intrinsically immune to radiation, the read/write operations of magnetic random access memory (MRAM) may be vulnerable to radiation-induced current. In this paper, we investigate the radiation hardening design for spin orbit torque based MRAM (SOT-MRAM). The hardening technique is firstly studied at the device level by optimizing the dimension and magnetic parameters. Then we propose radiation hardening read and write circuits addressing the influence of single event upset (SEU). Based on a physics-based SOT-MTJ compact model and a 65nm CMOS design kit, simulation results show that the proposed MOS-stacked read sensing amplifier and write circuits of six PMOS transistors as a feed-back structure to charge/discharge sensitive nodes can correct soft errors. Bi Wang 0002, Zhaohao Wang, Kaihua Cao, Youguang Zhang, Yuanfu Zhao, Weisheng Zhao 0001 |
ISCAS | 2 |
| 2018 | Multi-bit nonvolatile flip-flop based on NAND-like spin transfer torque MRAMabstractNonvolatile flip-flops (NVFFs) integrating emerging spintronics devices such as magnetic tunnel junction (MTJ) are under intensive investigation. They allow computing systems to be powered-off during the standby state, hence high static power issue of conventional CMOS technology can be addressed. MTJ based on spin transfer torque (STT) effect provide non-volatility, good endurance and 3D integration with CMOS based circuits. However, it suffers from relative long switching delay, high switching power and asymmetric switching issues. In this work, we first present a multi-bit NVFF using NAND-like spintronics (NANS-SPIN) devices which are written by STT and spin orbit torque (SOT) currents. It shows advantages in terms of power consumption, area overhead and write voltage. Then, functionality and performance of the proposed NVFF will be simulated and validated. Erya Deng, Zhaohao Wang, Wang Kang 0001, Shaoqian Wei, Weisheng Zhao 0001 |
VLSI-SoC | 2 |
| 2017 | Advanced Low Power Spintronic Memories beyond STT-MRAMabstractUntil now, spin transfer torque magnetic random access memory (STT-MRAM) has drawn considerable R&D interest worldwide. A number of companies and universities are currently involved in this promising technology. In 2016, Everspin released the first 256M STT-MRAM chip, indicating the commercialization and application of STT-MRAM. Nevertheless, STT-MRAM still has some intrinsic limitations, such as dynamic write power and speed, compared with CMOS-based memory technologies. Following the technical evolution process from toggle-MRAM to STT-MRAM, the continuous pursuit of high performance, high density, low power and scalability, drives the intensive R&D of new memory technologies. In this paper, we will show the recent progress in advanced spintronic memories beyond STT-MRAM, such as the spin Hall effect (SHE)-driven and voltage-driven MRAMs. These advanced MRAM technologies do have some unique advantages compared with STT-MRAM, but they also suffer from new design and fabrication challenges. In addition, we will present the latest research in emerging spintronic devices, e.g., magnetic skyrmions, which are potential as information carriers in future spintronic memories, e.g., racetrack memory. Wang Kang 0001, Zhaohao Wang, He Zhang 0011, Youguang Zhang, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2017 | PRESCOTT: Preset-based cross-point architecture for spin-orbit-torque magnetic random access memoryabstractDue to nearly zero leakage power consumption, non-volatile magnetoresistive random access memory (MRAM) is becoming one of the promising candidates for replacing conventional volatile memories (e.g. SRAM and DRAM). In particular, emerging spin-orbit torque (SOT) MRAM is considered to outperform spin-transfer torque (STT) MRAM due to its fast switching, separate read/write paths, and lower energy dissipation. However, the SOT-MRAM technology is still in its infancy; one key design challenge is that the control of SOT-MRAM, which involves three terminals, is more complicated compared with STT-MRAM. In this paper, we propose a novel MRAM write scheme called PRESCOTT1, where the “1” and “0” data values can be written into memory cells through the SOT and STT, respectively. As a result, the write current is unidirectional rather than bi-directional, which addresses the control complexity. Using this unidirectional write scheme, we design a PreSET-based cross-point (CP) MRAM to improve programing speed, write energy dissipation and storage density compared to conventional MRAM. Circuit simulation results demonstrate that our PreSET-based CP MRAM can achieve around 67.14% average write energy reduction and 50.86% improvement in programming speed, compared with CP STT-MRAM. Liang Chang 0002, Zhaohao Wang, Alvin Oliver Glova, Jishen Zhao, Youguang Zhang, Yuan Xie 0001, Weisheng Zhao 0001 |
ICCAD | 2 |
| 2017 | Pseudo-Differential Sensing Framework for STT-MRAM: A Cross-Layer PerspectiveabstractWith the rapid increase of leakage currents, non-volatile memories have become competitive candidates in the next-generation computer architecture. Among them, STT-MRAM shows great promise in working memory with high density, high speed and tremendous endurance, etc. However, based on our investigations, the dynamic write power and read reliability are two critical challenges of STT-MRAM. In this work, we propose a synergistic pseudo-differential sensing (PDS) framework that employs device, circuit and architectural techniques to address these challenges. In specific, three design techniques, including cell cluster, asymmetric sensing amplifier and self-error-detection-correction, are proposed to implement the PDS framework. We show that the holistic device-circuit-architecture cross-layer co-design enables STT-MRAM to be utilized in the cache memory, benefiting from the improved density, reliability and energy-efficiency. Our experimental results show that the proposed PDS scheme improves the read margin by ~35.6 percent, reduces the area, read latency, read energy, write latency and write power by ~46.7, ~9.8, ~30.3, ~2.3 and ~31.1 percent respectively, compared with the typical 1T1MTJ cell structure for the cache capacity of 8 MB. In addition, the proposed PDS scheme reduces the dynamic energy by ~32.9 percent and leakage energy by ~830 percent, improves the IPC by ~1.3 percent and miss rate by ~36.9 percent respectively, compared with conventional SRAM based cache. Wang Kang 0001, Liang Chang 0002, Zhaohao Wang, Weifeng Lv, Guangyu Sun 0003, Weisheng Zhao 0001 |
IEEE Trans. Computers | 3 |
| 2016 | A rough set approach to the characterization of transversal matroids
Guoye Xu, Zhaohao Wang |
Int. J. Approx. Reason. | 2 |
| 2015 | The approximation number function and the characterization of covering approximation space
Zhaohao Wang, Qinrong Feng, Lan Shu |
Inf. Sci. | 1 |
| 2015 | Spintronics: Emerging Ultra-Low-Power Circuits and Systems beyond MOS TechnologyabstractConventional MOS integrated circuits and systems suffer serve power and scalability challenges as technology nodes scale into ultra-deep-micron technology nodes (e.g., below 40nm). Both static and dynamic power dissipations are increasing, caused mainly by the intrinsic leakage currents and large data traffic. Alternative approaches beyond charge-only-based electronics, and in particular, spin-based devices, show promising potential to overcome these issues by adding the spin freedom of electrons to electronic circuits. Spintronics provides data non-volatility, fast data access, and low-power operation, and has now become a hot topic in both academia and industry for achieving ultra-low-power circuits and systems. The ITRS report on emerging research devices identified themagnetic tunnel junction(MTJ) nanopillar (one of the Spintronics nanodevices) as one of the most promising technologies to be part of future micro-electronic circuits. In this review we will give an overview of the status and prospects of spin-based devices and circuits that are currently under intense investigation and development across the world, and address particularly their merits and challenges for practical applications. We will also show that, with a rapid development of Spintronics, some novel computing architectures and paradigms beyond classic Von-Neumann architecture have recently been emerging for next-generation ultra-low-power circuits and systems. Wang Kang 0001, Yue Zhang 0010, Zhaohao Wang, Jacques-Olivier Klein, Claude Chappert, Dafine Ravelosona, Gefei Wang, Youguang Zhang, Weisheng Zhao 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2014 | An overview of spin-based integrated circuitsabstractConventional CMOS integrated circuits suffer from serve power and scalability challenges as technology node scales into ultra-deep-micron technology nodes. Alternative approaches beyond charge-only based circuits. In particular, spin-based devices or integrated circuits show promising merits to overcome these issues by adding the spin freedom of electrons to the electronic circuits. Spintronics has now become a hot topic in both academics and industrials. This paper overviews the status and prospects of spin-based integrated circuits under intense investigation and address particularly their merits and challenges for practical applications. Wang Kang 0001, Weisheng Zhao 0001, Zhaohao Wang, Jacques-Olivier Klein, Yue Zhang 0010, Djaafar Chabi, Youguang Zhang, Dafine Ravelosona, Claude Chappert |
ASP-DAC | 3 |
| 2014 | Ferroelectric tunnel memristor-based neuromorphic network with 1T1R crossbar architectureabstractEmerging ferroelectric tunnel memristors show large OFF/ON resistance ratio (>100) and high operation speed (~10ns), promising to be widely applied in the future synapse-like systems. In this paper we propose a neuromorphic network with ferroelectric tunnel memristor. This network is arranged with classical crossbar topology, in which each crosspoint forms a synapse consisting of a MOS transistor and a memristor. Based on this architecture, we design a spike-timing dependent plasticity (STDP) scheme and a parallel supervised learning circuit. Using a compact model of ferroelectric tunnel memristor and CMOS 40nm design kit, we perform transient simulation to validate the functionality of the proposed STDP and learning circuit. Simulation results show the potential of our neuromorphic network in low power (~100nA or ~1μA) and high speed (μs or ~100ns) computing system. Zhaohao Wang, Weisheng Zhao 0001, Wang Kang 0001, Youguang Zhang, Jacques-Olivier Klein, Claude Chappert |
IJCNN | 1 |
| 2014 | Design and analysis of crossbar architecture based on complementary resistive switching non-volatile memory cells
Weisheng Zhao 0001, Jean-Michel Portal, Wang Kang 0001, Mathieu Moreau, Yue Zhang 0010, Hassen Aziza, Jacques-Olivier Klein, Zhaohao Wang, Damien Querlioz, Damien Deleruyelle, Marc Bocquet, Dafine Ravelosona, Christophe Muller, Claude Chappert |
J. Parallel Distributed Comput. | 8 |
| 2013 | Spin-electronics based logic fabricsabstractAdvanced computing ICs in ultra deep-micron technology nodes (e.g. 40 nm) suffer from high power issues, which become one of the major bottlenecks for the future performance progress. Both static and dynamic power dissipation are increasing, caused mainly by the intrinsic leakage currents and large data traffic. Alternative approaches beyond charge-based logic circuits become hot research topics to overcome these issues definitively. By integrating the spin freedom of electrons to electronic devices, spin-electronics is promising for ultra-low power computing as it can provide non-volatility, fast data control and high logic density etc. Today, most of large microelectronics industries investigate this emerging field. In this invited paper for the special session “Nanoscale logic fabrics”, we overview spin-electronics based logic fabrics under intense investigation and address particularly the impact of this technology on logic architectures and new computing paradigms. Weisheng Zhao 0001, Jacques-Olivier Klein, Zhaohao Wang, Yue Zhang 0010, Nesrine Ben Romdhane, Damien Querlioz, Dafine Ravelosona, Claude Chappert |
VLSI-SoC | 3 |
| 2013 | Minimal Description and Maximal Description in Covering-based Rough SetsabstractRough set theory is an important technique in knowledge discovery in databases. In covering-based rough sets, seven types of rough set models were established in recent years. This paper defines the concept of maximal description of an element, and further explores the properties and structures of several types by means of the concepts of maximal description and minimal description. Finally, we study the relationship between covering-based rough sets and the generalized rough sets based on binary relation. Zhaohao Wang, Lan Shu, Xiuyong Ding |
Fundam. Informaticae | 1 |