He Zhang 0011

dblp:24/2058-11 · DBLP profile ↗
← Back
27ranked-venue papers
3as first author
23since 2021 · last 2026
0000-0001-9262-3106ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 3 first-author · 21 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 An Operator-Circuit Co-design Digital SOT-MRAM Computing-in-Memory Accelerator with Double Bit Density and Full-Utilized Bandwidth/Throughput
abstract
Computing-in-Memory (CIM) demonstrates exceptional performance on edge AI applications, owing to its in-situ computation capability with minimal data transfer consumption. However, volatile CIMs suffer from inevitable data retention power overhead, while non-volatile MRAM-CIMs still necessitate periodic weight updates constrained by limited memory space, diminishing the intrinsic advantage of CIMs. In this work, we propose a digital SOT-MRAM CIM accelerator with circuit-architecture-operator cross-layer design, achieving double bit density and full utilization of both data transmission bandwidth and computing throughput, thereby satisfying the stringent hardware demands for edge AI applications. Firstly, we propose a refined 2T-1MTJ non-complementary memory cell with an XOR-integrated pre-charged sense amplifier (X-SA), which significantly promotes the storage density and consumes only 6.284 fJ per read-based XOR operation. Then, we devise a channel-flatten data mapping (CFDM) scheme and an operator-aware residual fusion (OARF) structure to full utilize the storage and computing resources. Furthermore, an operator fusion method towards non-linear layers is proposed, achieving an 89.84% size reduction in non-binary parameters. System-level simulations at 40nm demonstrate that our work achieves 284.25 TOPS/W energy efficiency and 5.41 TOPS/mm2area efficiency with an accuracy of 98.72% (87.78%) on MNIST (CIFAR-10) dataset.
Tianshuo Bai, Jingcheng Gu, Lehao Tan, Wente Yi, Haolin Ge, Zhenyu Xue, He Zhang 0011, Na Lei, Biao Pan
DATE8
2026 An Effective SNN Macro with Real-Time STDP and Dynamic LIF Model Based on Thermally Interplayed Spin-Orbit Torque MTJ
abstract
Spiking neural networks (SNNs) have emerged as a promising paradigm for effective event-driven computation. However, CMOS-based SNN designs are limited by power consumption and complexity, while nonvolatile memory (NVM)-based SNN designs often lack biological characteristics and require active capacitive circuits to emulate neuronal dynamics. In this paper, we propose a thermally interplayed spin-orbit torque magnetic tunnel junction (TI-MTJ) macro that integrates core SNN functionalities. Our neuron array autonomously achieves leaky integrate-and-fire (LIF) model within the TI-MTJ device, thus improving power efficiency and simplifying circuit structure. Additionally, the proposed synaptic array provides adaptive in-situ responses based on a simplified spike-timing-dependent plasticity (STDP) rule. To enhance biological plausibility, our macro incorporates real-time spike monitoring and inhibition mechanisms. A comprehensive device-circuit-algorithm co-optimization framework validates the high performance of the TI-MTJ macro, achieving a synaptic energy consumption of 6.07fJ per spike, an inference accuracy of 97.76% on the MNIST dataset, and an energy efficiency of 22.8TOPS/W.
Changyu Li, Linjun Jiang, Liangchen Li, Dehang Zhu, Junda Zhao, Wang Kang 0001, Wenlong Cai, He Zhang 0011, Weisheng Zhao 0001
DATE9
2026 A 6.86Tb/s Bandwidth SOT-MRAM Sensing Scheme with Configurable Full-Column Over Frequency Technique for Near Memory Computing
Xinpeng Jiang, Hanting Chen, Zhaohao Wang, He Zhang 0011, Weisheng Zhao 0001
ISCAS4
2026 Design and Implementation of a Scalable 64 p-bits Ising Computing Chip with Integrated SOT-MTJs for Efficient Computing
Jinhao Li 0007, Jialiang Yin, Linying Liu, Chengyuan Sun, Hong-Xi Liu, Kaihua Cao, Zhaohao Wang, Wenlong Cai, He Zhang 0011
ISCAS13
2026 A 4/8b High-Precision Fully-Parallel In-Sensor Computing Chip with Subthreshold Digital Pixel and Hybrid Pulse Modulation
Junda Zhao, Yimo Du, Taoyi Wang, Junzhan Liu, He Zhang 0011, Wang Kang 0001
ISCAS6
2026 High-performance true random number generator based on SOT-MTJ spin relaxation
Jialiang Yin, Xiuye Zhang, Wenlong Cai, Ao Du, Binchao Tang, Shijian Bao, Daoqian Zhu, Kewen Shi, Lang Zeng, He Zhang 0011, Kaihua Cao, Weisheng Zhao 0001
Sci. China Inf. Sci.15
2026 High-Efficiency and Low-Deviation Analog-Digital Hybrid Compute-in-Memory Architecture With Dynamic Weight Division
abstract
Compute-in-memory (CIM) reduces data movement but suffers from an accuracy–efficiency trade-off: Analog CIM (ACIM) is energy-efficient but loses accuracy and incurs higher cost at large bit-widths, while digital CIM (DCIM) supports high precision but is inefficient for low-precision tasks. To overcome these challenges, we propose an analog–digital hybrid CIM (HCIM) architecture to address this trade-off, including 1) an analog–digital hybrid 10T SRAM cell without additional transistors and a dual-capacitor-based multicycle weighting module to reduce area; 2) a successive-approximation-register (SAR) ADC with a pseudo C-2C capacitor array that can be reconfigured from an 8-bit ADC into two parallel 4-bit ADCs to improve configurability; 3) configurable weight division and computing resource allocation strategies. Simulations in a 28-nm process show that HCIM achieves 15.56 TOPS/W at 12-bit ($8+4$) with$1.33\times $and$2.35\times $efficiency improvement over DCIM and ACIM and$16\times $lower error. It achieves 27.87 TOPS/W at 8-bit and 78.13 TOPS/W at 4-bit, demonstrating superior energy efficiency, computational accuracy, and flexibility.
Linjun Jiang, Sifan Sun, Wente Yi, Dengwen Li, Wang Kang 0001, He Zhang 0011, Weisheng Zhao 0001
IEEE Trans. Circuits Syst. I Regul. Pap.9
2026 A 40 nm Buffer-Free 7T-SRAM Analog Charge-Domain CIM Macro With Merging Timing Based On Time-Row Division Strategy
abstract
Computing-in-memory (CIM) macros based on static random access memory (SRAM) are meant to increase capacity while improving energy efficiency and reducing computing latency. However, traditional analog designs still face several key challenges, including long computing latency from separated computing phases, negative voltage fluctuations from massive parallel computing, and low bitcell density from additional transistors and capacitors for multiplication. On the other hand, only time-aligned inputs are supported in the works. To overcome the above challenges, this work proposes a buffer-free 7T-SRAM charge-domain CIM macro. It has four key features: 1) a compact 7T SRAM bitcell structure for high-energy efficiency; 2) a configurable input unit to support different sizes of input activations; 3) a time-row division (RD) strategy to support real-time processing and alleviate negative voltage fluctuations; and 4) a merging timing to conceal the input phase for high throughput. The fabricated 512-Kb SRAM-CIM macro in 40 nm achieves 79.3–290.4 Tops/W at 4-bit precision.
Linjun Jiang, Sifan Sun, Changyu Li, Wang Kang 0001, He Zhang 0011
IEEE Trans. Very Large Scale Integr. Syst.6
2026 Self-Calibrating Analog Circuitry for Softmax-Scaled Function With Analog Computing-In-Memory
abstract
Analog computing-in-memory (ACIM) has garnered widespread attention due to its advantage of high energy efficiency. However, it faces large power and hardware costs to handle sophisticated nonlinear functions, such as the softmax, due to costly exponentiation and division. Existing digital-domain approaches often rely on dedicated modules to carry out these operations, leading to a cost expensive area and high-power consumption. To address the issues, we propose a self-calibrating analog circuitry for a softmax-scaled function with ACIM. By exploiting transistor subthreshold properties, the work eliminates expensive digital operations while mapping exponentiation and division to successive analog circuits. A self-calibration module further mitigates partial mismatch-induced deviations by dynamically tuning bias voltages, improving overall fitting accuracy and system robustness. The proposed softmax-enabled ACIM work achieves energy efficiency of 55.06–60.08TOPS/W and 684.15 GOPS/mm2at 4-bit precision. In comparison with the state-of-the-art ACIMs with softmax implications, our proposed work shows higher energy efficiency and area efficiency.
Linjun Jiang, He Zhang 0011, Wang Kang 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2025 A 24.65 TOPS/W@INT8 Hybrid Analog-Digital Multi-core SRAM CIM Macro with Optimal Weight Dividing and Resource Allocation Strategies
abstract
Compute-in-memory (CIM) technology integrates memory and computation to reduce memory bottlenecks in modern systems. However, current CIM architectures face challenges in balancing accuracy and energy efficiency. Analog-CIM (ACIM) is energy-efficient but less accurate, while Digital-CIM (DCIM) is accurate but consumes more energy. In this paper, we propose a novel multi-core hybrid analog-digital CIM macro that effectively addresses this trade-off. Our approach intelligently allocates computation tasks to ACIM and DCIM cores based on their accuracy requirements, achieving a balance of accuracy and efficiency. Additionally, we developed an optimization framework to determine the optimal weight divide ratio and computing resource allocation for the hybrid CIM. Experimental results demonstrate the efficacy of our approach. The proposed hybrid CIM achieves an outstanding energy efficiency of 24.65 TOPS/W at 8-bit precision, surpassing DCIM by a factor of 1.33 while maintaining a low error rate of only 0.4%, which is 30 times better than ACIM at the same precision.
Wente Yi, Sifan Sun, Wenjia Wang 0011, Jinyu Bai, He Zhang 0011, Wang Kang 0001
ASP-DAC6
2025 An Adaptive Sparse Matrix Compression CIM Accelerator based on 256Kb SOT-MRAM for Downlink Massive MIMO Communications
abstract
Downlink precoding in massive multiple input multiple output (MIMO) systems involves high-dimensional sparse matrix calculations, which poses challenges to existing architectures. Computing-in-memory (CIM) has significant advantages in handling large-scale parallel operations, but sparse computing for wireless communication remains underexplored. In this paper, we propose a novel CIM accelerator based on magnetic random access memory (MRAM) leveraging adaptive multi-sparse mode technology for optimized sparse matrix multiplication in MIMO communication systems. This architecture represents the first application of CIM technology for processing sparse matrices in MIMO precoding tasks, minimizing storage requirements and enhancing parallel processing speed. Experimental results demonstrate that, for a 32×256×8 MIMO downlink precoding task with 90% sparsity, the symbol error rate is reduced to 0.1% at a signal-to-noise ratio of 20dB, achieving 8.35× reduction in storage overhead, 39.4× power saving and 9.85× speedup. These results position our accelerator as a promising candidate for processing sparse data in 5G massive MIMO systems.
Liangchen Li, Changyu Li, Anyang Yu, Junda Zhao, Zhaohao Wang, Chengyuan Sun, Kaihua Cao, Wang Kang 0001, He Zhang 0011, Weisheng Zhao 0001
ICCAD12
2025 Model quantization for computing-in-memory: a survey
Sifan Sun, Jinyu Bai, Hanting Chen, Kaiwen Deng, Zhiwei Xie 0009, He Zhang 0011, Wang Kang 0001, Weisheng Zhao 0001
Sci. China Inf. Sci.8
2025 A 0.88 e‾rms 8-Mpixel 3D-Stacked Low Temporal-Noise CMOS Image Sensor With Auto-Zero Single-Slope ADC, Fast Correlated Multi-Sampling, Row-Wise Noise Reduction, and Dark Current Non-Uniformity Calibration Techniques
abstract
This paper presents a low temporal noise, low-power, 8-Mpixel, rolling-shutter (RS)-type, back-illuminated CMOS image sensor (CIS) employing through silicon via (TSV) 3D-stack technology. To achieve temporal noise less than 1erms-, we explored auto-zero (AZ) column single-slope (SS) ADC and fast correlated multi-sampling (CMS) techniques. The pixel signal was sampled two times by the readout circuits using a 9-bits ADC, resulting in a 10-bits digital output. To enhance image quality in low light conditions, we adopted a parity column counter (PCC) for power supply stabilization and H-banding elimination, and employed row-wise noise reduction (RWNR) and dark-current non-uniformity calibration (DCNUC) techniques for reducing row-wise noise and improving image uniformity. Our CIS chip was fabricated using a 55nm 1P4M (pixel substrate) and a 55nm 1P5M (logic substrate) CIS 3D stacked process. The die area is ~3.99*3.45 mm2with 1.008-μm pixel pitch and the total energy consumption is 170mW under a 2.8V analog-VDD and a 1.2V digital-VDD. The chip achieves a temporal noise of only ~0.88erms-, fixed pattern noise (FPN) of ~25.08μVrms, row-wise noise of ~5.5μVrmsand an energy efficiency figure-of-merit (FoM) of ~0.6erms-*nJ/step at a frame rate of 60 frames per second (FPS).
Wang Kang 0001, Jing Kou, Liangchen Li, He Zhang 0011, Weisheng Zhao 0001
IEEE Trans. Circuits Syst. I Regul. Pap.7
2024 Series-Parallel Hybrid SOT-MRAM Computing-in-Memory Macro with Multi-Method Modulation for High Area and Energy Efficiency
abstract
Computing-in-memory (CIM) shows its superiority in lots of applications like neural network inference. Recently, there are lots of exploration of the application of Magnetic Random-Access Memory (MRAM) in CIM. This paper aims to investigate the potential of Spin-Orbit-Torque-MRAM (SOT-MRAM) in CIM and proposes a high area and energy efficiency SOT-MRAM CIM macro based on a 6T-4J weight group. The bit-cell array adopts series-parallel hybrid architecture, which combines both serial and parallel configurations of Magnetic Tunnel Junction (MTJ) to solve the problem of high energy cost and low flexibility caused by MRAM-series and MRAM-parallel architecture, respectively. Additionally, the proposed SOT-MRAM CIM macro incorporates a multi-method modulation scheme, ranging from input unit to array, which meanwhile allows for configurable input precision (2/4/6/8-bit). The SOT-MRAM CIM macro is designed and verified in both 180nm and 28nm nodes, based on the verified electrical performance of the SOT-MRAM array in a 200-nm wafer pre-fabricated. The simulation results in 28nm show that this macro can achieve energy efficiency of 23.7~29.6 Tops/W at 8-bit input and output precision.
Weiliang Huang, Jinyu Bai, Wang Kang 0001, Zhaohao Wang, Kaihua Cao, He Zhang 0011, Weisheng Zhao 0001
DAC7
2024 MixMixQ: Quantization with Mixed Bit-Sparsity and Mixed Bit-Width for CIM Accelerators
abstract
Quantization is vital for deploying neural networks on Computing-In-Memory (CIM) based accelerators due to inherent limitations in memory devices and data interfaces’ representational capacities. However, traditional quantization algorithms often overlook CIM’s unique computing paradigm, leading to suboptimal performance. To address this, we introduce MixMixQ, a novel quantization algorithm specifically designed for CIM accelerators that strategically integrates mixed bit-sparsity and mixed bit-width, enhancing overall hardware efficiency while preserving high accuracy. Notably, our method can enhance hardware efficiency by up to 294% compared to traditional quantization methods, with only a minimal 0.13% decrease in accuracy compared to a full-precision network.
Jinyu Bai, He Zhang 0011, Wang Kang 0001
ACM Great Lakes Symposium on VLSI2
2024 PipeCIM: A High-Throughput Computing-In-Memory Microprocessor With Nested Pipeline and RISC-V Extended Instructions
abstract
The large number of multiply accumulate (MAC) operations in Convolutional Neural Network (CNN) leads to substantial data migration and computation. Although computing-in-memory (CIM) proves to be a promising paradigm for MAC operations, high throughput CNN accelerator still confronts bottlenecks from: the low MAC utilization and the uncessary off-chip memory access. In this paper, we propose a high throughput CIM-based CNN accelerator PipeCIM with three hierarchies of pipelines: Intra-Macro, Near-Memory and Tile-Level. The Intra-Macro Pipeline parallelly executes data transfer and in-memory-computing (IMC) operations. The Near-Memory Pipeline alleviates memory access for pooling and data reshaping. The Tile-Level Pipeline establishes a layer-wise pipeline to further improve the throughput while reducing control complexity. PipeCIM introduces the nested scheme and a Unidirectional Divergent Connection Protocol (UDTCP) to simplify the control of data flow with the help of customized RISC-V instructions. To validate our design, PipeCIM was prototyped in 55 nm process node, achieving energy efficiency of 133.8 TOPS/W and peak throughput of 819 GOPS with a 16KB CIM array, which can accelerate VGG-16 to 128.56$\times$or Inception to 19.754$\times$compared to the baseline.
Tingran Chen, Wenjia Wang 0011, Haotian Fu, Wente Yi, Bojun Cheng, He Zhang 0011, Biao Pan
IEEE Trans. Circuits Syst. I Regul. Pap.7
2024 CiTST-AdderNets: Computing in Toggle Spin Torques MRAM for Energy-Efficient AdderNets
abstract
Recently, Adder Neural Networks (AdderNets) have gained widespread attention as an alternative to traditional Convolutional Neural Networks (CNNs) for deep learning tasks. AdderNets use lightweight addition operations to replace multiplication and accumulation (MAC) operations, but can keep almost the same accuracy compared to other CNNs. Nevertheless, challenges still exist with regards to hardware resources, power consumption, and communication bandwidth, primarily due to the ‘Von-Neumann bottlenecks’. However, computing-in-memory (CIM) architecture based on magnetic random-access memory (MRAM) has great potential for edge DNN implementation. In this paper, we propose a novel CIM paradigm using a novel Toggle-Spin-Torques (TST) driven MRAM for energy-efficient AdderNets (called CiTST_AdderNets). In CiTST_AdderNets, MRAM is driven by the interplay of the field-free spin orbit torque (SOT) effect and the spin transfer torque (STT) effect, which offers a fascinating prospect for energy efficiency and speed. Furthermore, a novel CIM paradigm is proposed to implement the dominating subtraction and sum operations in AdderNets, reducing data transfer and the related energy. Meanwhile, a highly parallel array structure integrating computation and storage is designed to support CiTST_AdderNets. In addition, a mapping strategy is proposed to efficiently map the convolution layer on the array. Fully connected layers can also be efficiently computed. The CiTST-AdderNets macro is designed by using a 65-nm CMOS process. Results show that our CiTST-AdderNets consumes about 1.65 mJ, 9.29 mJ, and 42.46 mJ for running VGG8, ResNet-50, and ResNet-18 respectively at 8-bit fixed-point precision. Compared to state-of-the-art platforms, our macro achieves an energy efficiency improvement of 1.45 x to 66.78 x.
Lichuan Luo, Erya Deng, Dijun Liu, Zhen Wang 0070, Weiliang Huang, He Zhang 0011, Xiao Liu 0051, Jinyu Bai, Junzhan Liu, Youguang Zhang, Wang Kang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.6
2024 Toward Energy-efficient STT-MRAM-based Near Memory Computing Architecture for Embedded Systems
abstract
Convolutional Neural Networks (CNNs) have significantly impacted embedded system applications across various domains. However, this exacerbates the real-time processing and hardware resource-constrained challenges of embedded systems. To tackle these issues, we propose spin-transfer torque magnetic random-access memory (STT-MRAM)-based near memory computing (NMC) design for embedded systems. We optimize this design from three aspects: Fast-pipelined STT-MRAM readout scheme provides higher memory bandwidth for NMC design, enhancing real-time processing capability with a non-trivial area overhead. Direct index compression format in conjunction with digital sparse matrix-vector multiplication (SpMV) accelerator supports various matrices of practical applications that alleviate computing resource requirements. Custom NMC instructions and stream converter for NMC systems dynamically adjust available hardware resources for better utilization. Experimental results demonstrate that the memory bandwidth of STT-MRAM achieves 26.7 GB/s. Energy consumption and latency improvement of digital SpMV accelerator are up to 64× and 1,120× across sparsity matrices spanning from 10% to 99.8%. Single-precision and double-precision elements transmission increased up to 8× and 9.6×, respectively. Furthermore, our design achieves a throughput of up to 15.9× over state-of-the-art designs.
Yueting Li 0001, He Zhang 0011, Biao Pan, Keni Qiu, Wang Kang 0001, Jun Wang 0041, Weisheng Zhao 0001
ACM Trans. Embed. Comput. Syst.3
2023 Toward Energy-Efficient Sparse Matrix-Vector Multiplication with near STT-MRAM Computing Architecture
abstract
Sparse Matrix-Vector Multiplication (SpMV) is one of the vital computational primitives used in modern workloads. SpMV performs memory access, leading to unnecessary data transmission, massive data access, and redundant multiplicative accumulators. Therefore, we propose the near spin-transfer torque magnetic random access memory (STT-MRAM) processing architecture from three optimization perspectives. These optimizations include (1) the NMP controller receives the instruction through the AXI4 bus to implement the SpMV operation in the following steps, identifies valid data, and encodes the index depending on the kernel size, (2) the NMP controller uses high-level synthesis dataflow in the shared buffer for achieving better performance throughput while do not consume bus bandwidth, and (3) the configurable MACs are implemented in the NMP core without matching step entirely during the multiplication. Using these optimizations, the NMP architecture can access the pipelined STT-MRAM (read bandwidth is 26.7GB/s). The experimental simulation results show that this design achieves up to 66x and 28x speedup compared with state-of-the-art ones and 69x speedup without sparse optimization.
Yueting Li 0001, He Zhang 0011, Hao Cai 0001, Shuqin Lv, Renguang Liu, Weisheng Zhao 0001
ASP-DAC2
2022 CP-SRAM: charge-pulsation SRAM marco for ultra-high energy-efficiency computing-in-memory
abstract
SRAM-based computing-in-memory (SRAM-CIM) provides fast speed and good scalability with advanced process technology. However, the energy efficiency of the state-of-the-art current-domain SRAM-CIM bit-cell structure is limited and the peripheral circuitry (e.g., DAC/ADC) for high-precision is expensive. This paper proposes a charge-pulsation SRAM (CP-SRAM) structure to achieve ultra-high energy-efficiency thanks to its charge-domain mechanism. Furthermore, our proposed CP-SRAM CIM supports configurable precision (2/4/6-bit). The CP-SRAM CIM macro was designed in 180nm (with silicon verification) and 40nm (simulation) nodes. The simulation results in 40nm show that our macro can achieve energy efficiency of ~2950Tops/W at 2-bit precision, ~576.4 Tops/W at 4-bit precision and ~111.7 Tops/W at 6-bit precision, respectively.
He Zhang 0011, Linjun Jiang, Tingran Chen, Junzhan Liu, Wang Kang 0001, Weisheng Zhao 0001
DAC1
2022 SpinCIM: spin orbit torque memory for ternary neural networks based on the computing-in-memory architecture
Lichuan Luo, Dijun Liu, He Zhang 0011, Youguang Zhang, Jinyu Bai, Wang Kang 0001
CCF Trans. High Perform. Comput.3
2022 HD-CIM: Hybrid-Device Computing-In-Memory Structure Based on MRAM and SRAM to Reduce Weight Loading Energy of Neural Networks
abstract
SRAM based computing-in-memory (SRAM-CIM) techniques have been widely studied for neural networks (NNs) to solve the “Von Neumann bottleneck”. However, as the scale of the NN model increasingly expands, the weight cannot be fully stored on-chip owing to the big device size (limited capacity) of SRAM. In this case, the NN weight data have to be frequently loaded from external memories, such as DRAM and Flash memory, which results in high energy consumption and low efficiency. In this paper, we propose a hybrid-device computing-in-memory (HD-CIM) architecture based on SRAM and MRAM (magnetic random-access memory). In our HD-CIM, the NN weight data are stored in on-chip MRAM and are loaded into SRAM-CIM core, significantly reducing energy and latency. Besides, in order to improve the data transfer efficiency between MRAM and SRAM, a high-speed pipelined MRAM readout structure is proposed to reduce the BL charging time. Our results show that the NN weight data loading energy in our design is only 0.242 pJ/bit, which is 289$\times $less in comparison with that from off-chip DRAM. Moreover, the energy breakdown and efficiency are analyzed based on different NN models, such as VGG19, ResNet18 and MobileNetV1. Our design can improve$\mathbf {58\times \,\,to\,\,124\times }$energy efficiency.
He Zhang 0011, Junzhan Liu, Jinyu Bai, Lichuan Luo, Shaoqian Wei, Wang Kang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.1
2021 SpinLiM: Spin Orbit Torque Memory for Ternary Neural Networks Based on the Logic-in-Memory Architecture
abstract
Logic-in-memory architecture based on spintronic memories shows fascinating prospects in neural networks (NNs) for its high energy efficiency and good endurance. In this work, we leveraged two magnetic tunnel junctions (MTJs), which are driven by the interplay of field-free spin orbit torque (SOT) and spin transfer torque (STT) effects, to achieve a novel statefullogic-in-memory paradigm for ternary multiplication operations. Based on this paradigm, we further proposed a highly parallel array structure to serve for ternary neural networks (TNNs). Our results demonstrate the advantage of our design in power consumption compared with CPU, GPU and other state-of-the-art works.
Lichuan Luo, He Zhang 0011, Jinyu Bai, Youguang Zhang, Wang Kang 0001, Weisheng Zhao 0001
DATE2
2020 High-Density, Low-Power Voltage-Control Spin Orbit Torque Memory with Synchronous Two-Step Write and Symmetric Read Techniques
abstract
Voltage-control spin orbit torque (VC-SOT) magnetic tunnel junction (MTJ) has the potential to achieve high-speed and low-power spintronic memory, owing to the adaptive voltage modulated energy barrier of the MTJ. However, the three-terminal device structure needs two access transistors (one for write operation and the other one for read operation) and thus occupies larger bit-cell area compared to two terminal MTJs. A feasible method to reduce area overhead is to stack multiple VC-SOT MTJs on a common antiferromagnetic strip to share the write access transistors. In this structure, high density can be achieved. However, write and read operations face problems and the design space is not sure given a strip length. In this paper, we propose a synchronous two-step multi-bit write and symmetric read method by exploiting the selective VC-SOT driven MTJ switching mechanism. Then hybrid circuits are designed and evaluated based a physics-based VC-SOT MTJ model and a 40nm CMOS design-kit to show the feasibility and performance of our method. Our work enables high-density, low-power, high-speed voltage-control SOT memory.
Wang Kang 0001, Liuyang Zhang, He Zhang 0011, Brajesh Kumar Kaushik, Weisheng Zhao 0001
DATE4
2020 Deep Neural Network accelerator with Spintronic Memory
abstract
Utilizing emerging nonvolatile memories to accelerate deep neural network (DNN) has been considered as one of the promising approaches to solve the bottleneck of data transfer during the multiplication and accumulation (MAC). Among them, spintronic memories show tempting prospect due to their low access power, fast access speed, high density, and relatively mature process. As shown in fig.1, according to the principle to achieve DNN computing, it can be mainly divided into three different technical routes. The first one is an "analog" method [1, 2], as shown in fig.1(a). By transforming the digital input signals into multi-level voltage signals, and applying them to different columns of the memory array, the MAC results can be obtained in different columns with current integrator and analog to digital converter (ADC). Besides, the WL drivers can control the pulse width of different rows, to achieve the effect of multi-bit weights. This method can theoretically achieve high energy efficiency and computing speed. However, the variation of magnetic tunnel junction (MTJ) may have influence on the computing accuracy. Besides, the power consumption and area overhead of the ADC are also challenging. The other two methods are in a "digital" way, and they realize MAC computing through row-by-row read/write operation. Fig.1(b) shows the second reading-based method [3]. The weights of the neural network are stored in the memory cell. By putting the input signal to the modified sensing amplifier (SA), it can also achieve XOR function, which is the core of binary NN, with the content stored in the memory cell. Nevertheless, the modification to the SA is usually to add extra transistors in the read path, which will increase the bit error rate. Fig.1(c) shows the diagram of the last one, which is based on the "stateful logic" [4]. The input data is sent to the modified write driver when the WL receiving weight signals from outside I/O. Based on a unique logic paradigm, it can realize XOR function for BNN within 1 or several memory cells during a write cycle. In this talk, we will review the main research status of DNN accelerators based on spintronic memories. Particularly, our recent work on DNN accelerating will be introduced, which can be implemented with different spintronic memories.
He Zhang 0011, Wang Kang 0001, Youguang Zhang, Weisheng Zhao 0001
ACM Great Lakes Symposium on VLSI1
2017 Advanced Low Power Spintronic Memories beyond STT-MRAM
abstract
Until now, spin transfer torque magnetic random access memory (STT-MRAM) has drawn considerable R&D interest worldwide. A number of companies and universities are currently involved in this promising technology. In 2016, Everspin released the first 256M STT-MRAM chip, indicating the commercialization and application of STT-MRAM. Nevertheless, STT-MRAM still has some intrinsic limitations, such as dynamic write power and speed, compared with CMOS-based memory technologies. Following the technical evolution process from toggle-MRAM to STT-MRAM, the continuous pursuit of high performance, high density, low power and scalability, drives the intensive R&D of new memory technologies. In this paper, we will show the recent progress in advanced spintronic memories beyond STT-MRAM, such as the spin Hall effect (SHE)-driven and voltage-driven MRAMs. These advanced MRAM technologies do have some unique advantages compared with STT-MRAM, but they also suffer from new design and fabrication challenges. In addition, we will present the latest research in emerging spintronic devices, e.g., magnetic skyrmions, which are potential as information carriers in future spintronic memories, e.g., racetrack memory.
Wang Kang 0001, Zhaohao Wang, He Zhang 0011, Youguang Zhang, Weisheng Zhao 0001
ACM Great Lakes Symposium on VLSI3
2017 Programmable Stateful In-Memory Computing Paradigm via a Single Resistive Device
abstract
Data transfer bandwidth and the related energy consumption has become two of the most critical bottlenecks in conventional von-Newman architecture, owing to the separation of the processor and memory units and the performance mismatch between the two. Realization of the unity of logic computing and data storage in the same die has opened up a promising research direction of in-memory computing (IMC). Meanwhile nonvolatile memory (NVM) based programmable (or reconfigurable) logic architecture has always been a hot topic in the circuit and system societies. To date, lots of interest has been attracted and amazing advance has been made in the two fields, yet none can fully exploit the advantages of both. This paper takes a major step forward by introducing a novel nonvolatile programmable stateful IMC architecture via a single resistive device, which is a completely different design paradigm from previous studies. Each memory cell can perform different Boolean logic functions by dynamically programming the input signals. The computing output result is insitu stored in the memory cell itself and can be readout with a memory-like operation. We will first give a brief review on this filed and then introduce our recent work. We will illustrate how the programmable stateful IMC operations can be implemented via a single resistive device and how the logic computing and data storage can be united within a memory chip.
Wang Kang 0001, He Zhang 0011, Youguang Zhang, Weisheng Zhao 0001
ICCD2