EDBT 2026 Demo / reviewers in the wild / expert
Wang Kang 0001
dblp:138/9373-1 · also Kang Wang 0001
· DBLP profile ↗
71ranked-venue papers
10as first author
41since 2021 · last 2026
0000-0002-3169-6034ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 68 · 9 first-author · 40 since 2021Software engineering, systems software and programming languages · 10 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Noise-Aware Adaptive Sampling for Robust Diffusion Models on Analog Compute-in-MemoryabstractDiffusion models achieve state-of-the-art image generation but impose heavy computational burdens on digital computers. Compute-in-memory (CIM) architectures offer promising acceleration, but inherent noise causes severe performance degradation through weight perturbations. We find that reducing sampling steps improves robustness but limits generation versatility, and that noise at earlier steps causes more severe degradation due to error accumulation. Based on these insights, we propose EtaMix, a novel noise-aware sampling strategy that interpolates between stochastic and deterministic sampling without requiring training or hardware modifications. EtaMix applies more stochastic sampling initially to offset weight perturbations, then gradually transitions to deterministic sampling. Experimental results show EtaMix achieves up to 2.01× and 5.12× FID improvements under different noise conditions for DDPM and DDIM, respectively. Yuannuo Feng, Wenyong Zhou, Yuexi Lv, Guangyao Wang, Zhengwu Liu, Ngai Wong 0001, Wang Kang 0001 |
DATE | 8 |
| 2026 | An Effective SNN Macro with Real-Time STDP and Dynamic LIF Model Based on Thermally Interplayed Spin-Orbit Torque MTJabstractSpiking neural networks (SNNs) have emerged as a promising paradigm for effective event-driven computation. However, CMOS-based SNN designs are limited by power consumption and complexity, while nonvolatile memory (NVM)-based SNN designs often lack biological characteristics and require active capacitive circuits to emulate neuronal dynamics. In this paper, we propose a thermally interplayed spin-orbit torque magnetic tunnel junction (TI-MTJ) macro that integrates core SNN functionalities. Our neuron array autonomously achieves leaky integrate-and-fire (LIF) model within the TI-MTJ device, thus improving power efficiency and simplifying circuit structure. Additionally, the proposed synaptic array provides adaptive in-situ responses based on a simplified spike-timing-dependent plasticity (STDP) rule. To enhance biological plausibility, our macro incorporates real-time spike monitoring and inhibition mechanisms. A comprehensive device-circuit-algorithm co-optimization framework validates the high performance of the TI-MTJ macro, achieving a synaptic energy consumption of 6.07fJ per spike, an inference accuracy of 97.76% on the MNIST dataset, and an energy efficiency of 22.8TOPS/W. Changyu Li, Linjun Jiang, Liangchen Li, Dehang Zhu, Junda Zhao, Wang Kang 0001, Wenlong Cai, He Zhang 0011, Weisheng Zhao 0001 |
DATE | 7 |
| 2026 | FALCON: A Fast and Low-Power Current-Mode Near-Sensor-Computing Architecture for Real-Time Edge Visual Processing
Jing Kou, Jinyao Mi, Junda Zhao, Junzhan Liu, Wang Kang 0001 |
DATE | 7 |
| 2026 | InFuzz: An Efficient and Lightweight In-Memory-Computing Cryptographic Fuzzy Extractor for IoT Security
Jing Kou, Wang Kang 0001 |
ISCAS | 7 |
| 2026 | FABS-CIM: Unlocking A/D Conversion Bottlenecks of Bit-Serial Computing-In-Memory with Analog Shift-and-Addition and In-Situ Batch Normalization
Junda Zhao, Jing Kou, Junzhan Liu, Wang Kang 0001 |
ISCAS | 6 |
| 2026 | A 4/8b High-Precision Fully-Parallel In-Sensor Computing Chip with Subthreshold Digital Pixel and Hybrid Pulse Modulation
Junda Zhao, Yimo Du, Taoyi Wang, Junzhan Liu, He Zhang 0011, Wang Kang 0001 |
ISCAS | 7 |
| 2026 | NoiseGuard: A Comprehensive Framework With Noise Modeling, Noise-Aware Training, and Noise Compensation for In-Memory Computing SoC
Guangyao Wang, Yizhe Chen, Yuexi Lv, Yuannuo Feng, Jenny Ma, Saiya Wang, Guilin Zhao, Yong Pei, Minghua Tang, Wang Kang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 13 |
| 2026 | High-Efficiency and Low-Deviation Analog-Digital Hybrid Compute-in-Memory Architecture With Dynamic Weight DivisionabstractCompute-in-memory (CIM) reduces data movement but suffers from an accuracy–efficiency trade-off: Analog CIM (ACIM) is energy-efficient but loses accuracy and incurs higher cost at large bit-widths, while digital CIM (DCIM) supports high precision but is inefficient for low-precision tasks. To overcome these challenges, we propose an analog–digital hybrid CIM (HCIM) architecture to address this trade-off, including 1) an analog–digital hybrid 10T SRAM cell without additional transistors and a dual-capacitor-based multicycle weighting module to reduce area; 2) a successive-approximation-register (SAR) ADC with a pseudo C-2C capacitor array that can be reconfigured from an 8-bit ADC into two parallel 4-bit ADCs to improve configurability; 3) configurable weight division and computing resource allocation strategies. Simulations in a 28-nm process show that HCIM achieves 15.56 TOPS/W at 12-bit ($8+4$) with$1.33\times $and$2.35\times $efficiency improvement over DCIM and ACIM and$16\times $lower error. It achieves 27.87 TOPS/W at 8-bit and 78.13 TOPS/W at 4-bit, demonstrating superior energy efficiency, computational accuracy, and flexibility. Linjun Jiang, Sifan Sun, Wente Yi, Dengwen Li, Wang Kang 0001, He Zhang 0011, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2026 | A 40 nm Buffer-Free 7T-SRAM Analog Charge-Domain CIM Macro With Merging Timing Based On Time-Row Division StrategyabstractComputing-in-memory (CIM) macros based on static random access memory (SRAM) are meant to increase capacity while improving energy efficiency and reducing computing latency. However, traditional analog designs still face several key challenges, including long computing latency from separated computing phases, negative voltage fluctuations from massive parallel computing, and low bitcell density from additional transistors and capacitors for multiplication. On the other hand, only time-aligned inputs are supported in the works. To overcome the above challenges, this work proposes a buffer-free 7T-SRAM charge-domain CIM macro. It has four key features: 1) a compact 7T SRAM bitcell structure for high-energy efficiency; 2) a configurable input unit to support different sizes of input activations; 3) a time-row division (RD) strategy to support real-time processing and alleviate negative voltage fluctuations; and 4) a merging timing to conceal the input phase for high throughput. The fabricated 512-Kb SRAM-CIM macro in 40 nm achieves 79.3–290.4 Tops/W at 4-bit precision. Linjun Jiang, Sifan Sun, Changyu Li, Wang Kang 0001, He Zhang 0011 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2026 | Self-Calibrating Analog Circuitry for Softmax-Scaled Function With Analog Computing-In-MemoryabstractAnalog computing-in-memory (ACIM) has garnered widespread attention due to its advantage of high energy efficiency. However, it faces large power and hardware costs to handle sophisticated nonlinear functions, such as the softmax, due to costly exponentiation and division. Existing digital-domain approaches often rely on dedicated modules to carry out these operations, leading to a cost expensive area and high-power consumption. To address the issues, we propose a self-calibrating analog circuitry for a softmax-scaled function with ACIM. By exploiting transistor subthreshold properties, the work eliminates expensive digital operations while mapping exponentiation and division to successive analog circuits. A self-calibration module further mitigates partial mismatch-induced deviations by dynamically tuning bias voltages, improving overall fitting accuracy and system robustness. The proposed softmax-enabled ACIM work achieves energy efficiency of 55.06–60.08TOPS/W and 684.15 GOPS/mm2at 4-bit precision. In comparison with the state-of-the-art ACIMs with softmax implications, our proposed work shows higher energy efficiency and area efficiency. Linjun Jiang, He Zhang 0011, Wang Kang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | A 24.65 TOPS/W@INT8 Hybrid Analog-Digital Multi-core SRAM CIM Macro with Optimal Weight Dividing and Resource Allocation StrategiesabstractCompute-in-memory (CIM) technology integrates memory and computation to reduce memory bottlenecks in modern systems. However, current CIM architectures face challenges in balancing accuracy and energy efficiency. Analog-CIM (ACIM) is energy-efficient but less accurate, while Digital-CIM (DCIM) is accurate but consumes more energy. In this paper, we propose a novel multi-core hybrid analog-digital CIM macro that effectively addresses this trade-off. Our approach intelligently allocates computation tasks to ACIM and DCIM cores based on their accuracy requirements, achieving a balance of accuracy and efficiency. Additionally, we developed an optimization framework to determine the optimal weight divide ratio and computing resource allocation for the hybrid CIM. Experimental results demonstrate the efficacy of our approach. The proposed hybrid CIM achieves an outstanding energy efficiency of 24.65 TOPS/W at 8-bit precision, surpassing DCIM by a factor of 1.33 while maintaining a low error rate of only 0.4%, which is 30 times better than ACIM at the same precision. Wente Yi, Sifan Sun, Wenjia Wang 0011, Jinyu Bai, He Zhang 0011, Wang Kang 0001 |
ASP-DAC | 7 |
| 2025 | Accuracy Is Not Always We Need: Precision-Aware Bayesian Yield OptimizationabstractIntegrated circuit yield optimization plays a vital role in ensuring reliable semiconductor manufacturing, directly impacting both product quality and production costs. Current approaches to yield optimization face two fundamental challenges that limit their practical effectiveness. First, yield estimation requires intensive computational resources. Second, traditional black-box optimization methods inefficiently allocate these resources across design candidates. Most existing approaches compound these issues by performing detailed yield estimations uniformly across all candidates, regardless of their potential quality. To address these limitations, we introduce a novel precision-aware yield optimization framework that intelligently adapts computational resource allocation based on each design candidate’s predicted performance. Our approach moves beyond simple simulation counting by incorporating a Figure of Merit (FoM) as a continuous quality metric. By combining a Continuous AutoRegression model to characterize the relationship between true yield and precision levels with a sophisticated multi-fidelity acquisition strategy, our framework achieves optimal resource distribution. Experimental validation on four industry-standard benchmark circuits demonstrates that our method converges with fewer than 1,000 simulations, reducing simulation costs by over $10 \times$ while achieving better final designs and robustness than state-of-the-art high-fidelity approaches. Jing Kou, Zidong Chen, Haiyan Qin, Wang Kang 0001, Wei W. Xing |
DAC | 5 |
| 2025 | Efficient Weight Mapping and Resource Scheduling on Crossbar-based Multi-core CIM SystemsabstractCrossbar-based computing-in-memory (CIM) systems facilitate large-scale parallel multiply-and-accumulate (MAC) operations, while a domain-specific compiler (DSC) plays a pivotal role in optimizing the deployment of neural network algorithms on such systems. With the development of multi-core and large-core architectures, some key compiler problems such as high parallel processing, resource utilization, and crossbar array assignment methods have not been solved. For low-latency application scenarios, we have designed a resource scheduling strategy for our hardware system based on stream data processing to reduce the latency caused by intra-core and intercore communication. Additionally, a weight mapping strategy has been developed to maximize the potential of crossbar arrays in convolutional neural networks (CNNs) deployment. Experimental results on our multi-core eFlash-based CIM system-on-chip (SoC) demonstrate that these two technologies help CNNs achieve a 76% reduction in latency, a 30% improvement in resource utilization, and the use rate of crossbar array that can reach up to 94.7%. Sifan Sun, Aifei Zhang, Haiyan Qin, Minhao Gu, Shihang Fu, Shuaikai Liu, Baosen Liu, Wang Kang 0001 |
DAC | 10 |
| 2025 | Multi-Agent Yield Analysis For Circuit DesignabstractSemiconductor yield estimation presents a critical challenge in modern manufacturing, directly impacting production costs and market competitiveness. Traditional estimation methods, particularly Monte Carlo simulation, while reliable, become computationally prohibitive for complex modern circuits. Contemporary approaches, including importance sampling and machine learning techniques, face fundamental limitations in consistency across circuit topologies and practical validation. This work introduces YieldAgent, a novel Large Language Model (LLM)-powered framework that revolutionizes yield estimation through dynamic integration of multiple analytical strategies. YieldAgent employs a three-layer agent architecture to analyze circuit characteristics and historical data, optimizing estimation methods while balancing computational efficiency and precision. The framework incorporates Retrieval-Augmented Generation for domain knowledge integration and Tree-structured Parzen Estimators for dynamic hyperparameter optimization. Experimental validation across 12nm and 40nm technology nodes demonstrates that YieldAgent reduces computational overhead by up to $2.9 \times$ while maintaining or exceeding state-of-the-art accuracy. The system’s ability to adapt across different circuit topologies and technology nodes establishes a new paradigm for scalable, intelligent yield estimation in electronic design automation. Haiyan Qin, Jing Kou, Wang Kang 0001, Wei W. Xing |
DAC | 4 |
| 2025 | HyIMC: Analog-Digital Hybrid In-Memory Computing SoC for High-Quality Low-Latency Speech EnhancementabstractIn-memory computing (IMC) holds significant promise for accelerating deep learning-based speech enhancement (DL-SE). However, existing IMC architectures face challenges in simultaneously achieving high precision, energy efficiency, and the necessary parallelism for DL-SE's inherent temporal dependencies. This paper introduces HyIMC, a novel hybrid analog-digital IMC architecture designed to address these limitations. HyIMC features: 1) a hybrid analog-digital design optimized for DL-SE algorithms; 2) a schedule controller that efficiently manages recurrent dataflow within skip connections; and 3) non-key dimension shrinkage, a model compression technique that preserves accuracy. Implemented on a 40nm eFlash-based IMC SoC prototype, HyIMC achieves 160 TOPS/W energy efficiency, compresses the DL-SE model size by ~600%, improves the feature of merit by ~1200%, and enhances perceptual evaluation of speech quality by ~120%. Wanru Mao, Guangyao Wang, Tianshuo Bai, Jingcheng Gu, Xitong Yang, Aifei Zhang, Xiaohang Wei, Wang Kang 0001 |
DATE | 11 |
| 2025 | Towards Accurate Characterization of In-Memory Computing Non-Idealities: A Physics & Data Co-Driven Generative FrameworkabstractAnalog in-memory computing (IMC) promises unprecedented energy efficiency for deep learning acceleration, but suffers from non-idealities that severely degrade inference accuracy in fabricated chips. Therefore, accurate modeling of these non-idealities becomes significant. In this work, we present PDGM-IMC, the first physics and data co-driven generative framework for IMC non-idealities characterizing. Unlike traditional physical models that fail to model complex non-ideality behaviors, or black-box neural networks that lack interpretability and generalization, PDGM-IMC leverages normalizing flows with custom transformations directly derived from device physics principles. This novel approach enables explicit modeling of complex probability distributions, spatial correlations, and die-to-die variations that previous methods could not capture. Validated on multiple dies of a commercial eFlash-based IMC SoC, PDGM-IMC improves modeling accuracy by 4.6× for the input circuit and IMC array and by 2.0× for the output circuit, significantly outperforming existing approaches. By extracting the statistical signature of fabricated chips, PDGM-IMC enables accurate pre-silicon prediction of post-silicon behavior, fundamentally transforming hardware-aware neural network optimization for analog accelerators. The source code and the pre-trained models are publicly available at https://github.com/BUAA-BASIC-Lab/PDGM-IMC. Jing Kou, Guangyao Wang, Saiya Wang, Yuexi Lv, Xinghao Cui, Wei. W. Xing, Wang Kang 0001 |
ICCAD | 10 |
| 2025 | An Adaptive Sparse Matrix Compression CIM Accelerator based on 256Kb SOT-MRAM for Downlink Massive MIMO CommunicationsabstractDownlink precoding in massive multiple input multiple output (MIMO) systems involves high-dimensional sparse matrix calculations, which poses challenges to existing architectures. Computing-in-memory (CIM) has significant advantages in handling large-scale parallel operations, but sparse computing for wireless communication remains underexplored. In this paper, we propose a novel CIM accelerator based on magnetic random access memory (MRAM) leveraging adaptive multi-sparse mode technology for optimized sparse matrix multiplication in MIMO communication systems. This architecture represents the first application of CIM technology for processing sparse matrices in MIMO precoding tasks, minimizing storage requirements and enhancing parallel processing speed. Experimental results demonstrate that, for a 32×256×8 MIMO downlink precoding task with 90% sparsity, the symbol error rate is reduced to 0.1% at a signal-to-noise ratio of 20dB, achieving 8.35× reduction in storage overhead, 39.4× power saving and 9.85× speedup. These results position our accelerator as a promising candidate for processing sparse data in 5G massive MIMO systems. Liangchen Li, Changyu Li, Anyang Yu, Junda Zhao, Zhaohao Wang, Chengyuan Sun, Kaihua Cao, Wang Kang 0001, He Zhang 0011, Weisheng Zhao 0001 |
ICCAD | 11 |
| 2025 | PAR-CIM: A Precise/Approximate Reconfigurable Digital CIM Macro with 0.35-4b Fractional Mixed-Bitwidth QuantizationabstractDigital computing-in-memory (DCIM) enables efficient deep neural networks (DNNs) acceleration but faces limitations in resource overhead, energy efficiency, and architectural flexibility. Existing approximate or reconfigurable DCIM solutions tackle these issues partially without achieving a holistic balance. To address this, we propose PAR-CIM, a highly energy-efficient reconfigurable CIM macro that integrates precise and approximate paradigm. First, we introduce layer/gate-level approximate computation (LGAC) into the adder tree (AT) of the DCIM core, achieving full operation with only 0.35× the area of traditional implementations. Then, we develop a 0.35-4b fractional mixed-bitwidth quantization (FMBQ) algorithm, combining second-order Taylor sensitivity analysis with DoReFa-Net. This is complemented by a high-precision low-approximation (HPLA) mapping scheme to enhance energy efficiency. Additionally, a multi-bit reconfigurable computation mode (MBRM) strategy further improves architectural flexibility and enables the implementation of the proposed design. Under 40nm technology, PAR-CIM achieves 3048 TOPS/W at 1b/1b operations. With FMBQ, ResNet18 and our custom V-FuseMBA trained on CIFAR-10 achieve over 86.61% compression with accuracy loss under 0.74%, reaching classification accuracies of 93.67% and 92.86%, respectively. Zhenyu Xue, Wente Yi, Tianshuo Bai, Lehao Tan, Jingcheng Gu, Weijie Ding, Wang Kang 0001, Biao Pan |
ICCAD | 8 |
| 2025 | ACSNN: A 61.25 TOPS/W, 1.65 ns delay SNN Processor that combines CIM-inspired Synapse and Asynchronous ArchitectureabstractSpiking Neural Networks (SNNs) offer their biological plausibility and dynamic sensitivity which have gained significant attention in real-time systems. However, designing SNN processors with high throughput and real-time response remains challenging due to high power consumption, high area costs and routing competition. In this paper, we present ACSNN, a processor that integrates a CIM-inspired synapse array, Leaky Integrate-and-Fire (LIF) neurons and asynchronous architecture to achieve high parallelism, high energy and area efficiency while declining routing competition. Experimental results demonstrate an impressive power efficiency of 61.25 TOPS/W and a high peak throughput of 2,415 GOPS with 1.65 ns minimum compute delay, highlighting superior performance and power efficiency compared to state-of-the-art designs. Tingran Chen, Yuxuan Ran, Yueting Li 0001, Wang Kang 0001, Biao Pan |
ISCAS | 5 |
| 2025 | Model quantization for computing-in-memory: a survey
Sifan Sun, Jinyu Bai, Hanting Chen, Kaiwen Deng, Zhiwei Xie 0009, He Zhang 0011, Wang Kang 0001, Weisheng Zhao 0001 |
Sci. China Inf. Sci. | 9 |
| 2025 | A 0.88 e‾rms 8-Mpixel 3D-Stacked Low Temporal-Noise CMOS Image Sensor With Auto-Zero Single-Slope ADC, Fast Correlated Multi-Sampling, Row-Wise Noise Reduction, and Dark Current Non-Uniformity Calibration TechniquesabstractThis paper presents a low temporal noise, low-power, 8-Mpixel, rolling-shutter (RS)-type, back-illuminated CMOS image sensor (CIS) employing through silicon via (TSV) 3D-stack technology. To achieve temporal noise less than 1erms-, we explored auto-zero (AZ) column single-slope (SS) ADC and fast correlated multi-sampling (CMS) techniques. The pixel signal was sampled two times by the readout circuits using a 9-bits ADC, resulting in a 10-bits digital output. To enhance image quality in low light conditions, we adopted a parity column counter (PCC) for power supply stabilization and H-banding elimination, and employed row-wise noise reduction (RWNR) and dark-current non-uniformity calibration (DCNUC) techniques for reducing row-wise noise and improving image uniformity. Our CIS chip was fabricated using a 55nm 1P4M (pixel substrate) and a 55nm 1P5M (logic substrate) CIS 3D stacked process. The die area is ~3.99*3.45 mm2with 1.008-μm pixel pitch and the total energy consumption is 170mW under a 2.8V analog-VDD and a 1.2V digital-VDD. The chip achieves a temporal noise of only ~0.88erms-, fixed pattern noise (FPN) of ~25.08μVrms, row-wise noise of ~5.5μVrmsand an energy efficiency figure-of-merit (FoM) of ~0.6erms-*nJ/step at a frame rate of 60 frames per second (FPS). Wang Kang 0001, Jing Kou, Liangchen Li, He Zhang 0011, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | Series-Parallel Hybrid SOT-MRAM Computing-in-Memory Macro with Multi-Method Modulation for High Area and Energy EfficiencyabstractComputing-in-memory (CIM) shows its superiority in lots of applications like neural network inference. Recently, there are lots of exploration of the application of Magnetic Random-Access Memory (MRAM) in CIM. This paper aims to investigate the potential of Spin-Orbit-Torque-MRAM (SOT-MRAM) in CIM and proposes a high area and energy efficiency SOT-MRAM CIM macro based on a 6T-4J weight group. The bit-cell array adopts series-parallel hybrid architecture, which combines both serial and parallel configurations of Magnetic Tunnel Junction (MTJ) to solve the problem of high energy cost and low flexibility caused by MRAM-series and MRAM-parallel architecture, respectively. Additionally, the proposed SOT-MRAM CIM macro incorporates a multi-method modulation scheme, ranging from input unit to array, which meanwhile allows for configurable input precision (2/4/6/8-bit). The SOT-MRAM CIM macro is designed and verified in both 180nm and 28nm nodes, based on the verified electrical performance of the SOT-MRAM array in a 200-nm wafer pre-fabricated. The simulation results in 28nm show that this macro can achieve energy efficiency of 23.7~29.6 Tops/W at 8-bit input and output precision. Weiliang Huang, Jinyu Bai, Wang Kang 0001, Zhaohao Wang, Kaihua Cao, He Zhang 0011, Weisheng Zhao 0001 |
DAC | 3 |
| 2024 | MixMixQ: Quantization with Mixed Bit-Sparsity and Mixed Bit-Width for CIM AcceleratorsabstractQuantization is vital for deploying neural networks on Computing-In-Memory (CIM) based accelerators due to inherent limitations in memory devices and data interfaces’ representational capacities. However, traditional quantization algorithms often overlook CIM’s unique computing paradigm, leading to suboptimal performance. To address this, we introduce MixMixQ, a novel quantization algorithm specifically designed for CIM accelerators that strategically integrates mixed bit-sparsity and mixed bit-width, enhancing overall hardware efficiency while preserving high accuracy. Notably, our method can enhance hardware efficiency by up to 294% compared to traditional quantization methods, with only a minimal 0.13% decrease in accuracy compared to a full-precision network. Jinyu Bai, He Zhang 0011, Wang Kang 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2024 | An End-to-End In-Memory Computing System Based on a 40-nm eFlash-Based IMC SoC: Circuits, Toolchains, and Systems Co-Design FrameworkabstractDespite its promising potential for Artificial Intelligence (AI) applications, current In-Memory Computing (IMC) technology faces a variety of challenges before mass production. One of the major challenges we face is the absence of efficient toolchains for deploying canonical networks on IMC chips. To address this issue, we propose a co-designed framework that integrates circuit, toolchain, and system elements specifically for IMC. More specifically, our framework consists of several key techniques to improve the key performance including (a) an 8-bit hardware-friendly Quantization-Aware Training (QAT) approach to quantify the deep learning network from floating-point data to fixed-point data, (b) a novel operator optimization technique to increase the computing precision when running the algorithm models on the IMC chips, and (c) an efficient mapping strategy based on the Integer Linear Programming (ILP) approach to improve the computation resource utilization of the IMC array. We assess our method on our 40nm eFlash-based IMC SoC chip with voice recognition, speech noise reduction, and person detection tasks. Our experimental results show an accuracy over 94.60% in a quiet environment and 87.27% in a white noise environment and a false recognition rate below 1 time per 24 hours for voice recognition, a 21.53% improvement for the Perceptual Evaluation of Speech Quality (PESQ) for noise reduction, and a 97.80% accuracy in person detection. Tianshuo Bai, Wanru Mao, Guangyao Wang, Aifei Zhang, Shihang Fu, Shuaikai Liu, Jianchao Hu, Xitong Yang, Biao Pan, Wei W. Xing, Wang Kang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 12 |
| 2024 | CIMQ: A Hardware-Efficient Quantization Framework for Computing-In-Memory-Based Neural Network AcceleratorsabstractThe novel computing-in-memory (CIM) technology has demonstrated significant potential in enhancing the performance and efficiency of convolutional neural networks (CNNs). However, due to the low precision of memory devices and data interfaces, an additional quantization step is necessary. Conventional NN quantization methods fail to account for the hardware characteristics of CIM, resulting in inferior system performance and efficiency. This article proposes CIMQ, a hardware-efficient quantization framework designed to improve the efficiency of CIM-based NN accelerators. The holistic framework focuses on the fundamental computing elements in CIM hardware: inputs, weights, and outputs (or activations, weights, and partial sums in NNs) with four innovative techniques. First, bit-level sparsity induced activation quantization is introduced to decrease dynamic computation energy. Second, inspired by the unique computation paradigm of CIM, an innovative arraywise quantization granularity is proposed for weight quantization. Third, partial sums are quantized with a reparametrized clipping function to reduce the required resolution of analog-to-digital converters (ADCs). Finally, to improve the accuracy of quantized neural networks (QNNs), the post-training quantization (PTQ) is enhanced with a random quantization dropping strategy. The effectiveness of the proposed framework has been demonstrated through experimental results on various NNs and datasets (CIFAR10, CIFAR100, and ImageNet). In typical cases, the hardware efficiency can be improved up to 222% with a 58.97% improvement in accuracy compared to conventional quantization methods. Jinyu Bai, Sifan Sun, Weisheng Zhao 0001, Wang Kang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | CIM²PQ: An Arraywise and Hardware-Friendly Mixed Precision Quantization Method for Analog Computing-In-MemoryabstractComputing-in-memory (CIM) architecture is a promising convolutional neural network (CNN) accelerator known for its highly efficient matrix-vector multiplications (MVMs). However, due to the low-precision computation and limited size of CIM memory arrays, it is necessary to decompose the huge MVMs into smaller subsets. Conventional NN quantization methods overlook the characteristics of CIM hardware, resulting in diminished system performance and efficiency. This paper proposes a mixed precision quantization (MPQ) method based on evolutionary algorithm for CIM-based accelerators, while considering the hardware characteristics of CIM, called CIMPQ, which can automatically generate quantization strategies for NN model to improve the efficiency of CIM systems. Firstly, inspired by the CIM computing paradigm, an array-wise quantization granularity is introduced in the MPQ search space, which can jointly quantize the inputs, weights, and partial sums. Secondly, a production procedure containing fine-grained crossover and progressive adaptive mutation is proposed, which can efficiently explore the search space and speed up the search process. Thirdly, we propose a fast and efficient strategy evaluation method to obtain the performance of quantization strategy on the CIM platform, saving the evaluation time significantly without requiring fine-tuning. Finally, to protect CIM-friendly strategies with lower bit-widths but worse algorithm performance, we propose a strategy selection method based on multi-objective optimization, named qNSGA-III. The effectiveness of the proposed method has been demonstrated through experimental results of various NNs and datasets. For ResNet-18, the hardware efficiency and accuracy can be improved to 117% with 7.05%, 113% with 3.37%, and 119% with 5.78%, on CIFAR-10, CIFAR-100 and ImageNet, respectively, compared to the baseline MPQ method. Sifan Sun, Jinyu Bai, Zhaoyu Shi, Weisheng Zhao 0001, Wang Kang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | Toward Write Optimization for Skyrmion Racetrack Memory by Skyrmion RepermutationabstractSkyrmion racetrack memory (Sky-RM), as an emerging non-volatile memory technology, provides both high-density data storage and low-access latency, accompanying with a novel feature of using current to shift data along a racetrack. However, it is time-comsuing and energy-hungry to write/inject new particle-like Skyrimons on Sky-RM compared with other manipulations. Thus, the goal of this work is to presents a novel strategy to optimize write performance in Sky-RM, Permutation-Write (PW). Particuarly, PW reduces the number of Skyrmion injections when writing data, by “re-permuting" existing Skyrmion particles within a racetrack. Moreover, based on the flexible concept of re-permutation, this work further proposes PW+ strategy, which is an optimized PW strategy integrated with injection and shift optimizations to achieve better performance. Our evaluation results justify that PW+ can greatly reduce the amount of Skyrmion injections, which are considered highly expensive in Sky-RM, by at least 50%, compared with other state-of-the-art write-optimized strategies. Furthermore, PW+ only incurs 1.79 injections, averaged for every 64-bit-word write. We show that PW+ strategy can bring significant benefits over other state-of-the-art write-optimized strategies by about 16-50% for every 64-bit-word write, for both performance and energy efficiency. Tsun-Yu Yang, Xiangjun Peng, Wang Kang 0001, Ming-Chang Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | CiTST-AdderNets: Computing in Toggle Spin Torques MRAM for Energy-Efficient AdderNetsabstractRecently, Adder Neural Networks (AdderNets) have gained widespread attention as an alternative to traditional Convolutional Neural Networks (CNNs) for deep learning tasks. AdderNets use lightweight addition operations to replace multiplication and accumulation (MAC) operations, but can keep almost the same accuracy compared to other CNNs. Nevertheless, challenges still exist with regards to hardware resources, power consumption, and communication bandwidth, primarily due to the ‘Von-Neumann bottlenecks’. However, computing-in-memory (CIM) architecture based on magnetic random-access memory (MRAM) has great potential for edge DNN implementation. In this paper, we propose a novel CIM paradigm using a novel Toggle-Spin-Torques (TST) driven MRAM for energy-efficient AdderNets (called CiTST_AdderNets). In CiTST_AdderNets, MRAM is driven by the interplay of the field-free spin orbit torque (SOT) effect and the spin transfer torque (STT) effect, which offers a fascinating prospect for energy efficiency and speed. Furthermore, a novel CIM paradigm is proposed to implement the dominating subtraction and sum operations in AdderNets, reducing data transfer and the related energy. Meanwhile, a highly parallel array structure integrating computation and storage is designed to support CiTST_AdderNets. In addition, a mapping strategy is proposed to efficiently map the convolution layer on the array. Fully connected layers can also be efficiently computed. The CiTST-AdderNets macro is designed by using a 65-nm CMOS process. Results show that our CiTST-AdderNets consumes about 1.65 mJ, 9.29 mJ, and 42.46 mJ for running VGG8, ResNet-50, and ResNet-18 respectively at 8-bit fixed-point precision. Compared to state-of-the-art platforms, our macro achieves an energy efficiency improvement of 1.45 x to 66.78 x. Lichuan Luo, Erya Deng, Dijun Liu, Zhen Wang 0070, Weiliang Huang, He Zhang 0011, Xiao Liu 0051, Jinyu Bai, Junzhan Liu, Youguang Zhang, Wang Kang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 11 |
| 2024 | Toward Energy-efficient STT-MRAM-based Near Memory Computing Architecture for Embedded SystemsabstractConvolutional Neural Networks (CNNs) have significantly impacted embedded system applications across various domains. However, this exacerbates the real-time processing and hardware resource-constrained challenges of embedded systems. To tackle these issues, we propose spin-transfer torque magnetic random-access memory (STT-MRAM)-based near memory computing (NMC) design for embedded systems. We optimize this design from three aspects: Fast-pipelined STT-MRAM readout scheme provides higher memory bandwidth for NMC design, enhancing real-time processing capability with a non-trivial area overhead. Direct index compression format in conjunction with digital sparse matrix-vector multiplication (SpMV) accelerator supports various matrices of practical applications that alleviate computing resource requirements. Custom NMC instructions and stream converter for NMC systems dynamically adjust available hardware resources for better utilization. Experimental results demonstrate that the memory bandwidth of STT-MRAM achieves 26.7 GB/s. Energy consumption and latency improvement of digital SpMV accelerator are up to 64× and 1,120× across sparsity matrices spanning from 10% to 99.8%. Single-precision and double-precision elements transmission increased up to 8× and 9.6×, respectively. Furthermore, our design achieves a throughput of up to 15.9× over state-of-the-art designs. Yueting Li 0001, He Zhang 0011, Biao Pan, Keni Qiu, Wang Kang 0001, Jun Wang 0041, Weisheng Zhao 0001 |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2023 | Skyrmion Vault: Maximizing Skyrmion Lifespan for Enabling Low-Power Skyrmion Racetrack MemoryabstractSkyrmion racetrack memory (SK-RM) has demonstrated great potential as a high-density and low-cost nonvolatile memory. Nevertheless, even though random data accesses are supported on SK-RM, data accesses can not be carried out on individual data bit directly. Instead, special skyrmion manipulations, such as injecting and shifting, are required to support random information update and deletion. With such special manipulations, the latency and energy consumption of skyrmion manipulations could quickly accumulate and induce additional overhead on the data read/write path of SK-RM. Meanwhile, injection operation consumes more energy and has higher latency than any other manipulations. Although prior arts have tried to alleviate the overhead of skyrmion manipulations, the possibility of minimizing injections through buffering skyrmions for future reuse and energy conservation receives much less attention. Such observation motivates us to propose the concept of skyrmion vault to effectively utilize the skyrmion buffer track structure for energy conservation through maximizing the lifespan of injected skyrmions and minimizing the number of skyrmion injections. Experimental results have shown promising improvements in both energy consumption and skyrmions' lifespan. Syue-Wei Lu, Shuo-Han Chen, Yu-Pei Liang, Yuan-Hao Chang 0001, Wang Kang 0001, Tseng-Yi Chen, Wei-Kuan Shih |
ASP-DAC | 5 |
| 2023 | Hierarchical Non-Structured Pruning for Computing-In-Memory Accelerators with Reduced ADC Resolution RequirementabstractThe crossbar architecture, which is comprised of novel nano-devices, enables high-speed and energy-efficient computing-in-memory (CIM) for neural networks. However, the overhead from analog-to-digital converters (ADCs) substantially degrades the energy efficiency of CIM accelerators. In this paper, we introduce a hierarchical non-structured pruning strategy where value-level and bit-level pruning are performed jointly on neural networks to reduce the resolution of ADCs by using the famous alternating direction method of multipliers (ADMM). To verify the effectiveness, we deployed the proposed method to a variety of state-of-the-art convolutional neural networks on two image classification benchmark datasets: CIFAR10, and ImageNet. The results show that our pruning method can reduce the required resolution of ADCs to 2 or 3 bits with only slight accuracy loss (~0.25 %), and thus can improve the hardware efficiency by 180%. Wenlu Xue, Jinyu Bai, Sifan Sun, Wang Kang 0001 |
DATE | 4 |
| 2023 | OPT: Optimal Proposal Transfer for Efficient Yield Optimization for Analog and SRAM CircuitsabstractYield optimization is one of the central challenges in submicrometer integrated circuit manufacture. However, yield optimization is computationally expensive due to intensive yield estimation and intractable optimization processes. In this work, we first reinvent the state-of-the-art all sensitivity adversarial importance sampling (ASAIS) yield optimization from a Laplace approximation perspective, which also reveals its limitations and suggests improvements. We then generalize it with infinite components and discover the key ingredient in yield optimization to be an effective proposal distribution transfer (OPT) procedure, which is captured using conditional normalizing flow (CNF). To deliver a reliable yield optimization pipeline that accounts for the uncertainty due to the lack of data, we propose sequential ensemble, the first empirical uncertainty estimation that enables tractable Bayesian yield optimization without introducing an extra surrogate for the first time. We conduct extensive experiments against five state-of-the-art baselines and show that the proposed method delivers superior performance: a speedup of 1.01x-11.94x (5.57x on average) with higher yield designs, and most importantly, excellent robustness and consistency in all our experiments on analog and SRAM circuits. Guohao Dai 0002, Yuanqing Cheng, Wang Kang 0001, Wei W. Xing |
ICCAD | 4 |
| 2023 | Experimental Demonstration of STT-MRAM-based Nonvolatile Instantly On/Off System for IoT Applications: Case StudiesabstractEnergy consumption has been a big challenge for electronic devices, particularly for battery-powered Internet of Things (IoT) equipment. To address such a challenge, on the one hand, low-power electronic design methodologies and novel power management techniques have been proposed, such as nonvolatile memories and instantly on/off systems; on the other hand, the energy harvesting technology by collecting signals from human activity or the environment has attracted widespread attention in the IoT area. However, the system with self-powered energy harvesting may suffer frequent energy failures or fluctuating energy conditions, which degrade system reliability and user experience. Therefore, how to make the system under unreliable power inputs operate correctly and efficiently is one of the most critical issues for energy harvesting technology. In this article, we built an instantly on/off system based on nonvolatile STT-MRAM for IoT applications, which can instantly power on/off under different conditions of the harvested energy. The system powers on and operates normally when the harvested energy is enough (over the preset threshold); otherwise, the system powers off and stores the operational data back to the nonvolatile STT-MRAM. We described implementations of the hardware/software co-designed architecture (with image acquisition as an example) based on the commercialized 32 MB STT-MRAM, and we experimentally demonstrated the system functionality and efficiency under five typical energy harvesting scenarios, including radio frequency, thermal, solar, piezoelectric, and WIFI. Our experimental results show that the power consumption and data restore time were reduced by 15.1% and 714 times, respectively, in comparison with the DRAM-based counterpart. Yueting Li 0001, Wang Kang 0001, Kunyu Zhou, Keni Qiu, Weisheng Zhao 0001 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2022 | SMART: on simultaneously marching racetracks to improve the performance of racetrack-based main memoryabstractRaceTrack Memory (RTM) is a promising media for modern Main Memory subsystems. However, the "shift-before-access" principle, as the nature of RTM, introduces considerable overheads to the access latency. To obtain more insights for the mitigation of shift overheads, this work characterizes and observes that the access patterns, exhibited by the state-of-the-art RTM-based Main Memory, mismatches with the granularity of shift commands (i.e., a group of RaceTracks called Domain Block Cluster (DBC)). Based on the characterization, we propose a novel mechanism called SMART, which simultaneously and proactively marches all DBCs within a subarray, so that subsequent accesses to other DBCs can be served without additional shift commands. Evaluation results show that, averaged across 15 real-world workloads, SMART significantly outperforms other state-of-the-art proposals of RTM-based Main Memory by at least 1.53X in terms of the total execution time, on two different generations of RTM technologies. Xiangjun Peng, Ming-Chang Yang, Ho Ming Tsui, Chi Ngai Leung, Wang Kang 0001 |
DAC | 5 |
| 2022 | Efficient bayesian yield analysis and optimization with active learningabstractYield optimization for circuit design is computationally intensive due to the expensive yield estimation based on Monte Carlo methods and the difficult optimization process. In this work, a uniform framework to solve these problems simultaneously is proposed. Firstly, a novel efficient Bayesian yield analysis framework, BYA, is proposed by deriving a Bayesian estimation for the yield and introducing active learning based on reductions of integral entropy. A tractable convolutional entropy infill technique is then proposed to efficiently solve the entropy reduction problem. Lastly, we extend BYA for yield optimization by transforming knowledge across the design space and variational space. Experimental results based on SRAM and adder circuits show that BYA is 410x faster (in terms of the number of simulations) than standard MC and averagely 10x (up to 10000x) more accurate than the state-of-the-art method for yield estimation, and is about 5x faster than the SOTA yield optimization methods. Xiang Jin, Linxu Shi, Wang Kang 0001, Wei W. Xing |
DAC | 4 |
| 2022 | CP-SRAM: charge-pulsation SRAM marco for ultra-high energy-efficiency computing-in-memoryabstractSRAM-based computing-in-memory (SRAM-CIM) provides fast speed and good scalability with advanced process technology. However, the energy efficiency of the state-of-the-art current-domain SRAM-CIM bit-cell structure is limited and the peripheral circuitry (e.g., DAC/ADC) for high-precision is expensive. This paper proposes a charge-pulsation SRAM (CP-SRAM) structure to achieve ultra-high energy-efficiency thanks to its charge-domain mechanism. Furthermore, our proposed CP-SRAM CIM supports configurable precision (2/4/6-bit). The CP-SRAM CIM macro was designed in 180nm (with silicon verification) and 40nm (simulation) nodes. The simulation results in 40nm show that our macro can achieve energy efficiency of ~2950Tops/W at 2-bit precision, ~576.4 Tops/W at 4-bit precision and ~111.7 Tops/W at 6-bit precision, respectively. He Zhang 0011, Linjun Jiang, Tingran Chen, Junzhan Liu, Wang Kang 0001, Weisheng Zhao 0001 |
DAC | 6 |
| 2022 | SpinCIM: spin orbit torque memory for ternary neural networks based on the computing-in-memory architecture
Lichuan Luo, Dijun Liu, He Zhang 0011, Youguang Zhang, Jinyu Bai, Wang Kang 0001 |
CCF Trans. High Perform. Comput. | 6 |
| 2022 | HD-CIM: Hybrid-Device Computing-In-Memory Structure Based on MRAM and SRAM to Reduce Weight Loading Energy of Neural NetworksabstractSRAM based computing-in-memory (SRAM-CIM) techniques have been widely studied for neural networks (NNs) to solve the “Von Neumann bottleneck”. However, as the scale of the NN model increasingly expands, the weight cannot be fully stored on-chip owing to the big device size (limited capacity) of SRAM. In this case, the NN weight data have to be frequently loaded from external memories, such as DRAM and Flash memory, which results in high energy consumption and low efficiency. In this paper, we propose a hybrid-device computing-in-memory (HD-CIM) architecture based on SRAM and MRAM (magnetic random-access memory). In our HD-CIM, the NN weight data are stored in on-chip MRAM and are loaded into SRAM-CIM core, significantly reducing energy and latency. Besides, in order to improve the data transfer efficiency between MRAM and SRAM, a high-speed pipelined MRAM readout structure is proposed to reduce the BL charging time. Our results show that the NN weight data loading energy in our design is only 0.242 pJ/bit, which is 289$\times $less in comparison with that from off-chip DRAM. Moreover, the energy breakdown and efficiency are analyzed based on different NN models, such as VGG19, ResNet18 and MobileNetV1. Our design can improve$\mathbf {58\times \,\,to\,\,124\times }$energy efficiency. He Zhang 0011, Junzhan Liu, Jinyu Bai, Lichuan Luo, Shaoqian Wei, Wang Kang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2021 | SpinLiM: Spin Orbit Torque Memory for Ternary Neural Networks Based on the Logic-in-Memory ArchitectureabstractLogic-in-memory architecture based on spintronic memories shows fascinating prospects in neural networks (NNs) for its high energy efficiency and good endurance. In this work, we leveraged two magnetic tunnel junctions (MTJs), which are driven by the interplay of field-free spin orbit torque (SOT) and spin transfer torque (STT) effects, to achieve a novel statefullogic-in-memory paradigm for ternary multiplication operations. Based on this paradigm, we further proposed a highly parallel array structure to serve for ternary neural networks (TNNs). Our results demonstrate the advantage of our design in power consumption compared with CPU, GPU and other state-of-the-art works. Lichuan Luo, He Zhang 0011, Jinyu Bai, Youguang Zhang, Wang Kang 0001, Weisheng Zhao 0001 |
DATE | 5 |
| 2021 | Spin-Orbit Torque Nonvolatile Flip-Flop DesignsabstractFlip-flops (FFs) are basic units in electronic circuits. Recently, nonvolatile FFs (NVFFs) have attracted great interests for power-gating applications and a variety of NVFFs have been proposed by integrating nonvolatile memory devices. Among them, magnetic tunnel junction (MTJ) based NVFFs show considerable potential in terms of zero static power consumption and high endurance. Nevertheless, the mainstream spin transfer torque (STT) effect based MTJ switching approach for data storing still consumes much dynamic power and long delay, limiting the system performance and data reliability. The spin-orbit torque (SOT) effect provides an alternative approach for high-speed and low-power MTJ switching, therefore rather promising for NVFF design. In this work, we propose four NVFF designs based on the FF architectures (either DFF or SRFF) and perpendicular MTJ (pMTJ). The circuit structures and operations are investigated, and the performance is evaluated and compared at the 40 nm process technology node. Simulation results show that the proposed NVFFs can achieve high read speed (<; 200 ps), low read power consumption (<; 10 fJ) and area efficiency. Erya Deng, Wang Kang 0001, Weisheng Zhao 0001, Shaoqian Wei, You Wang 0002, Deming Zhang |
ISCAS | 2 |
| 2021 | HSC: A Hybrid Spin/CMOS Logic Based In-Memory Engine with Area-Efficient Mapping StrategyabstractRecent advances in deep learning have shown that binary neural networks (BNNs) can provide a satisfying accuracy on various tasks with significant reduction in computation power and memory cost. Theoretically, the multiply-and-accumulate (MAC) operations of BNNs can be replaced by in-memory XNOR operations, thereby avoiding frequent data transfer between the buffer and the processor. However, devices supporting in-memory implementation of XNOR operations together with efficient weight-matrix mapping strategy is still an open research area. In this paper, a hybrid spin/CMOS cell (HSC) structure is proposed in which the XNOR operation can be simply realized in an in-memory computing manner by the non-volatile data from the spin component and the volatile data from the CMOS component. Given the time/spatial trade-off, a novel weight mapping method to break the large memory array and unroll the 3D kernel into 2D weight matrix is designed to cooperate with the proposed HSC structure in a time-division way. System-level simulation results show that the proposed BNN processor can achieve a 3.32* speedup and 11.9* improvement in throughput and energy efficiency, which could be attributed to the device and mapping method co-design. Erya Deng, Jinyu Bai, Wang Kang 0001, Biao Pan |
ISCAS | 5 |
| 2020 | Permutation-Write: Optimizing Write Performance and Energy for Skyrmion Racetrack MemoryabstractSkyrmion racetrack memory (Sky-RM) is an emerging non-volatile memory technology that offers not only high-density data storage but also the feature of shifting data along a racetrack by current. However, compared to other shifting based manipulations, writing/injecting new particle-like Skyrmions is much more time-consuming and energy-hungry. To optimize the write performance and energy for Sky-RM, this paper introduces a new write strategy, called Permutation-Write. Specifically, this strategy circumvents new Skyrmion injections during data writes by "re-permuting" existing Skyrmion particles in the racetrack. Evaluation results show that Permutation-Write strategy can significantly reduce the number of expensive Skyrmion injections by at least 50% compared to other state-of-the-art write reduction strategies, and merely introduce 2.25 injections on average for every 64-bit-word write under realistic workloads. Tsun-Yu Yang, Ming-Chang Yang, Wang Kang 0001 |
DAC | 4 |
| 2020 | High-Density, Low-Power Voltage-Control Spin Orbit Torque Memory with Synchronous Two-Step Write and Symmetric Read TechniquesabstractVoltage-control spin orbit torque (VC-SOT) magnetic tunnel junction (MTJ) has the potential to achieve high-speed and low-power spintronic memory, owing to the adaptive voltage modulated energy barrier of the MTJ. However, the three-terminal device structure needs two access transistors (one for write operation and the other one for read operation) and thus occupies larger bit-cell area compared to two terminal MTJs. A feasible method to reduce area overhead is to stack multiple VC-SOT MTJs on a common antiferromagnetic strip to share the write access transistors. In this structure, high density can be achieved. However, write and read operations face problems and the design space is not sure given a strip length. In this paper, we propose a synchronous two-step multi-bit write and symmetric read method by exploiting the selective VC-SOT driven MTJ switching mechanism. Then hybrid circuits are designed and evaluated based a physics-based VC-SOT MTJ model and a 40nm CMOS design-kit to show the feasibility and performance of our method. Our work enables high-density, low-power, high-speed voltage-control SOT memory. Wang Kang 0001, Liuyang Zhang, He Zhang 0011, Brajesh Kumar Kaushik, Weisheng Zhao 0001 |
DATE | 2 |
| 2020 | Deep Neural Network accelerator with Spintronic MemoryabstractUtilizing emerging nonvolatile memories to accelerate deep neural network (DNN) has been considered as one of the promising approaches to solve the bottleneck of data transfer during the multiplication and accumulation (MAC). Among them, spintronic memories show tempting prospect due to their low access power, fast access speed, high density, and relatively mature process. As shown in fig.1, according to the principle to achieve DNN computing, it can be mainly divided into three different technical routes. The first one is an "analog" method [1, 2], as shown in fig.1(a). By transforming the digital input signals into multi-level voltage signals, and applying them to different columns of the memory array, the MAC results can be obtained in different columns with current integrator and analog to digital converter (ADC). Besides, the WL drivers can control the pulse width of different rows, to achieve the effect of multi-bit weights. This method can theoretically achieve high energy efficiency and computing speed. However, the variation of magnetic tunnel junction (MTJ) may have influence on the computing accuracy. Besides, the power consumption and area overhead of the ADC are also challenging. The other two methods are in a "digital" way, and they realize MAC computing through row-by-row read/write operation. Fig.1(b) shows the second reading-based method [3]. The weights of the neural network are stored in the memory cell. By putting the input signal to the modified sensing amplifier (SA), it can also achieve XOR function, which is the core of binary NN, with the content stored in the memory cell. Nevertheless, the modification to the SA is usually to add extra transistors in the read path, which will increase the bit error rate. Fig.1(c) shows the diagram of the last one, which is based on the "stateful logic" [4]. The input data is sent to the modified write driver when the WL receiving weight signals from outside I/O. Based on a unique logic paradigm, it can realize XOR function for BNN within 1 or several memory cells during a write cycle. In this talk, we will review the main research status of DNN accelerators based on spintronic memories. Particularly, our recent work on DNN accelerating will be introduced, which can be implemented with different spintronic memories. He Zhang 0011, Wang Kang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2020 | A Comparative Cross-layer Study on Racetrack Memories: Domain Wall vs SkyrmionabstractRacetrack memory (RM), a new storage scheme in which information flows along a nanotrack, has been considered as a potential candidate for future high-density storage device instead of hard disk drive (HDD). The first RM technology, which was proposed in 2008 by IBM, relies on a train of opposite magnetic domains separated by domain walls (DWs), named DW-RM. After 10 years of intensive research, a variety of fundamental advancements has been achieved; unfortunately, no product has been available until now. With increasing effort and resources dedicated to the development of DW-RM, it is likely that new materials and mechanisms will soon be discovered for practical applications. However, new concepts might also be on the horizon. Recently, an alternative information carrier, magnetic skyrmion, which was experimentally discovered in 2009, has been regarded as a promising replacement of DW for RM, named skyrmion-based RM (SK-RM). Intensive effort has been involved and amazing advances have been made in observing, writing, manipulating, and deleting individual skyrmions. So, what is the relationship between DW and skyrmion? What are the key differences between DW and skyrmion, or between DW-RM and SK-RM? What benefits could SK-RM bring and what challenges need to be addressed before application? In this review article, we intend to answer these questions through a comparative cross-layer study between DW-RM and SK-RM. This work will provide guidelines, especially for circuit and architecture researchers on RM. Wang Kang 0001, Bi Wu 0002, Xing Chen 0012, Daoqian Zhu, Zhaohao Wang, Xichao Zhang, Youguang Zhang, Weisheng Zhao 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2020 | Shift-Limited Sort: Optimizing Sorting Performance on Skyrmion Memory-Based SystemsabstractModern nonvolatile memories (NVMs) are widely recognized as energy-efficient replacements of classical memory/storage media, such as SRAM, DRAM, and mechanical hard disk. Among the popular NVMs, the skyrmion racetrack memory (SK-RM) is well known for its high storage density and unique supports of insert/delete operations. However, the existing algorithms designed for classical media might experience serious performance degradation when working on the SK-RM, due to the distinct characteristics of SK-RM. Thus, the existing algorithms should be redesigned to adapt to the brand-new memory model based on the SK-RM, so as to fully reveal the potentials of SK-RM. In particular, many existing algorithms tend to access the in-memory data in a random-hopping fashion, which generates many time-consuming shift operations of SK-RM. It is therefore crucial for the existing algorithms to eliminate unnecessary shift operations of SK-RM to boost the performance of the algorithms. In many modern applications, such as multimedia and data analysis, it is a common operation to process two or more arrays/vectors of data to perform certain computation tasks. In the arrays/vectors, an appropriate data placement strategy is critical for avoiding unnecessary shift operations of SK-RM. The observation thus motivates this work in proposing a recursive back-to-back data placement manner to effectively reduces the shift operations of SK-RM. To demonstrate the back-to-back data placement, we take sorting algorithms as a case study, and propose a novel shift-limited sorting algorithm for SK-RM. Analytical studies show that the shift-limited sort effectively enhances the time complexity of classical merge sort from O(dn lg n) to O(n lg n), where d is the bit distance between adjacent access ports on the nanotracks of the SK-RM. After that, the efficacy of the proposed shift-limited sort is then verified by experimental studies, where the results are encouraging. Yun-Shan Hsieh, Po-Chun Huang, Ping-Xiang Chen, Yuan-Hao Chang 0001, Wang Kang 0001, Ming-Chang Yang, Wei-Kuan Shih |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2019 | ZUMA: Enabling Direct Insertion/Deletion Operations with Emerging Skyrmion Racetrack MemoryabstractData insertion and deletion are common operations exist in various applications. However, traditional memory architecture can only perform an indirect insertion/deletion with multiple data read and write operations, which is significantly time and energy consuming. To mitigate this problem, we propose to leverage the unique capability of emerging skyrmion racetrack memory technology that it can naturally support direct insertion/deletion operations inside a racetrack. In this work, we first present a circuit level model for skyrmion racetrack memory. Then, we further propose a novel memory architecture to enable an efficient large size data insertion/deletion. With the help of the model and the architecture, we study several potential applications to leverage the insertion and deletion operations. Experimental results demonstrate that the efficiency of these operations can be substantially improved. Zheng Liang 0003, Guangyu Sun 0003, Wang Kang 0001, Xing Chen 0012, Weisheng Zhao 0001 |
DAC | 3 |
| 2019 | Voltage-Controlled Magnetoelectric Memory Bit-cell Design With Assisted Body-bias in FD-SOIabstractVoltage-controlled magnetic anisotropy (VCMA)-magnetic tunnel junction (MTJ) is incorporated into FD-SOI CMOS technology. The design space of 1 transistor-1 MTJ (1T-1M) bit-cell is explored through varied VCMA pulse duration/amplitude and scaling down transistor dimensions. The design point with 1.1 V VCMA pulse amplitude, 0.44 ns pulse duration and W/L = 400 nm/30 nm access transistor shows the ultra low write energy in VCMA-MTJ based bit-cell. It achieves a minimum 3.18 fJ/bit switching energy with 28-nm FD-SOI process. Access transistor sizing is studied, while the ultra low power implementation may lead to MTJ switching failure. Voltage assisted techniques for failure mitigation are proposed based on body-bias generator (BBG). The BBG not only provides VCMA pulse signal to control MTJ barrier, but also generates body-bias to boost the transistor performance. In the presence of forward body-bias (FBB) and increased VCMA pulse level, the proposed strategy is effective in switching failure compensation as well as writing delay improvement. Hao Cai 0001, Menglin Han, Weiwei Shan, Jun Yang 0006, You Wang 0002, Wang Kang 0001, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 6 |
| 2019 | Magnetic Skyrmion-Based Neural Recording System Design for Brain Machine InterfaceabstractNext-generation brain machine interface demand a high-channel-count neural recording system to wirelessly monitor activities of thousands of neurons. In order to achieve high-density neural recording, further development of single recording channel comprised of a neural amplifier front-end (AFE) and an analog-to-digit converter (ADC) is critical. Despite the great progress made in CMOS implementation of custom-designed neural recording system, hybrid limitations of increasing area and power consumption in line with Moore's law drove great demand for post-CMOS substitutes. Magnetic skyrmion with nano particle-like and non-volatile properties are of both fundamental and applied interests for future bio-inspired electronics. In this work, we propose a compact model including both AFE and ADC based on current-induced skyrmion motion. The proposed system achieved a power consumption of 0.63 pJ/channel with an area overhead of 0.14 μm2. The purpose of this work is to explore the feasibility of magnetic skyrmion for building large-scale, dense neuronal recording system which could pave a new way for future brain machine interface application. Biao Pan, Wang Kang 0001, Xing Chen 0012, Jinyu Bai, Jianlei Yang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 2 |
| 2019 | SR-WTA: Skyrmion Racing Winner-Takes-All Module for Spiking Neural ComputingabstractSpiking neural network (SNN) has emerged as one of the popular architectures in complex pattern recognition and classification tasks. However, hardware implementation of such algorithms using conventional CMOS based neuron consume resources and power that are orders of magnitude higher than that in human brain. This can be attributed to the mismatch of the computational architecture between biological brain and the current Boolean logic computing platform. Magnetic skyrmions have been intensively studied as a prospective information carrier in neuromorphic computing hardware design. In this work, a compact time-domain skyrmion-racing winner-takes-all (SR-WTA) leaky-integrate-fire (LIF) spiking neuron network is presented for the first time. The skyrmion motion dynamics in the LIF neuron and the behaviors of the neuron network was investigated comprehensively. Both SPICE and micromagnetic simulations are performed to evaluate the functionality and performance of the proposed SR-WTA based SNN. Biao Pan, Wang Kang 0001, Xing Chen 0012, Jinyu Bai, Jianlei Yang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 2 |
| 2019 | An STT-MRAM Based in Memory Architecture for Low Power Integral ComputingabstractThe integral histogram image plays an important role in accelerating the feature computation in vision algorithms. However, the computational process of the integral histogram, called integral computation, has high computational complexity and numerous memory access operations, which limit its wide application. This brief proposes an in-memory computational architecture based on Spin Transfer Torque Magnetic Random Access Memory (STT-MRAM) to solve these problems. The architecture can work in two different modes depending on the requirements: the integral computation mode and the memory mode. The architecture can figure out the integral histogram when in the integral computation mode, and just store the data directly when in the memory mode. Utilizing the non-volatile, high density and low power characteristics of STT-MRAM, we integrate the computational units into the memory array to achieve parallel computation. Reduced number of data transmission between storage units and computation units contributes to cut down the latency and energy consumption. The evaluation results show that, comparing with the state-of-the-art work, our architecture provides$1.1\times \sim 9\times$performance improvements and reduces 87.4$\sim$97.3 percent energy consumption for$64\times 64\sim 512\times 512$size images, just with a 8 percent area overhead. Yinglin Zhao, Wang Kang 0001, Shouyi Yin, Youguang Zhang, Shaojun Wei, Weisheng Zhao 0001 |
IEEE Trans. Computers | 3 |
| 2018 | Process variation aware data management for magnetic skyrmions racetrack memoryabstractSkyrmions racetrack memory (SKM) has been identified as a promising candidate for future on-chip cache. Similar to many other nanoscale technologies, process variations also adversely impact the reliability and performance of SKM cache. In this work, we propose the first holistic solution for employing SKM as last-level caches. We first present a novel SKM cache architecture and a physical-to-logic mapping scheme based on our comprehensive analysis on working mechanism of SKM. We then model the impact of process variations on SKM cache performance. By leveraging the developed model, we propose a process variation aware data management technique to minimize the performance degradation of SKM cache incurred by process variations. Experimental results show that the proposed SKM cache can achieve a geometric mean of 1.28x IPC improvement, 2x density increase, and 23% energy reduction compared to Domain Wall racetrack memory (DWM) under the same area constraint across 15 workloads. In addition, our dynamic data management technique can further improve the system IPC by 25% w.r.t. the worst-case design. Fan Chen 0001, Wang Kang 0001, Weisheng Zhao 0001, Hai Li 0001, Yiran Chen 0001 |
ASP-DAC | 3 |
| 2018 | Magnetic skyrmions for future potential memory and logic applications: Alternative information carriersabstractMagnetic skyrmions are swirling topological configurations, which are mostly induced by chiral interactions between atomic spins in non-centrosymmetric magnetic bulks or in thin films with broken inversion symmetry. They hold promise as information carriers in future ultra-dense, low-power memory and logic devices owing to the nanocale size and extremely low spin-polarized currents needed to move them. To date, an intense research effort has led to the identification, creation/annihilation, motion and manipulation of skyrmions at room temperature. Meanwhile, a rich variety of skyrmion-based device concepts and prototypes have been proposed, indicating the considerable potential of magnetic skyrmions in future electronic applications. However, current studies mainly focus on physical or principle investigations, whereas the electrical design methodology, implementation and evaluations are still lacking. In this paper, we will bring the readers in the “design, automation and test (DAT) society” the current status and outlook of skyrmions in relation to future potential racetrack memory and neuromorphic computing applications. Most importantly, we also want to evoke the effort from the DAT society to address the challenges, e.g., all-electrical manipulation of skyrmions at room temperature, for the research and development of practical skyrmion-based electronics. Wang Kang 0001, Xing Chen 0012, Daoqian Zhu, Yangqi Huang, Youguang Zhang, Weisheng Zhao 0001 |
DATE | 1 |
| 2018 | Enabling Resilient Voltage-Controlled MeRAM Using Write Assist TechniquesabstractReliability concerns arise in nonvolatile magnetoelectric random access memory (MeRAM) due to continuously nanotechnology scaling down and CMOS-magnetic hybrid integration. The primary objective of this work is to investigate failure mitigation in voltage-controlled magnetic anisotropy-magnetic tunnel junction (VCMA-MTJ) based 1T-1MTJ MeRAM bit-cell, by using MTJ compact model and 28nm fully depleted silicon on insulator (FD-SOI) process design-kit. A comprehensive reliability study is performed considering process variation and aging degradations, including hot carrier injection (HCI), bias temperature instability (BTI), soft breakdown (SBD) and radiation effect. Write assist techniques are proposed to ensure failure resilient MeRAM design. Bit line (BL) boost and negative source line (SL) methods show high efficiency in writing latency improvement and failure mitigation. Hao Cai 0001, You Wang 0002, Wang Kang 0001, Lirida A. B. Naviner, Weiwei Shan, Jun Yang 0006, Weisheng Zhao 0001 |
ISCAS | 3 |
| 2018 | Design and Data Management for Magnetic Racetrack MemoryabstractBenefiting from its ultra-high storage density, high energy efficiency, and non-volatility, racetrack memory demonstrates great potential in replacing conventional SRAM as large on-chip memory. Integrating the tape-like racetrack memory, however, faces unique design challenges from cell structure to architecture design. This paper reviews some cross-layer design methodologies for racetrack memory as on-chip cache hierarchy. Research studies show that with proper architectural design and data management, racetrack memory can achieve significant area reduction, system performance enhancement, and energy saving compared to state-of-the-art memory technologies. Bing Li 0017, Fan Chen 0001, Wang Kang 0001, Weisheng Zhao 0001, Yiran Chen 0001, Hai Li 0001 |
ISCAS | 3 |
| 2018 | Progresses and challenges of spin orbit torque driven magnetization switching and application (Invited)abstractSpin orbit torque (SOT) has been proposed as a potential alternative mechanism to the conventional spin transfer torque (STT) for the magnetization switching. Recently, theoretical and experimental works revealed the novel factors influencing the SOT-driven magnetization switching. Emerging SOT-based spintronics memories and circuits were explored to implement fast and energy-efficient write operation. However, the perspective of the SOT mechanism is still challenged by some serious shortcomings, such as area penalty, relatively large switching current density and undesirable use of external magnetic field. Here, we review the progresses in the SOT mechanism involving the magnetization dynamics, device design and circuit development. Key issues to be addressed in optimizing the SOT devices are pointed out. In particular, we discuss the potential solutions to develop high-density SOT-based memories and circuits. Zhaohao Wang, Zuwei Li, Liang Chang 0002, Wang Kang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 6 |
| 2018 | Multi-bit nonvolatile flip-flop based on NAND-like spin transfer torque MRAMabstractNonvolatile flip-flops (NVFFs) integrating emerging spintronics devices such as magnetic tunnel junction (MTJ) are under intensive investigation. They allow computing systems to be powered-off during the standby state, hence high static power issue of conventional CMOS technology can be addressed. MTJ based on spin transfer torque (STT) effect provide non-volatility, good endurance and 3D integration with CMOS based circuits. However, it suffers from relative long switching delay, high switching power and asymmetric switching issues. In this work, we first present a multi-bit NVFF using NAND-like spintronics (NANS-SPIN) devices which are written by STT and spin orbit torque (SOT) currents. It shows advantages in terms of power consumption, area overhead and write voltage. Then, functionality and performance of the proposed NVFF will be simulated and validated. Erya Deng, Zhaohao Wang, Wang Kang 0001, Shaoqian Wei, Weisheng Zhao 0001 |
VLSI-SoC | 3 |
| 2017 | Voltage-controlled MRAM for working memory: Perspectives and challengesabstractMagnetic random access memory (MRAM) has been widely studied for future nonvolatile working memory candidate. However, the mainstream current (spin transfer torque, STT or spin Hall effect, SHE) driven MRAMs (STT-MRAM or SHE-MRAM) face intrinsic problems in terms of high write power and long latency, significantly limiting the applications for low-power and high-speed working memories. The recently-developed new-generation MRAM, named VCMA-MRAM, which exploits the voltage-controlled magnetic anisotropy (VCMA) effect to write (or assist to write) data information into magnetic tunnel junctions (MTJs), holds the promise to efficiently overcome these problems. Despite the impressive possibility of improving write power and speed, this technology, however, is currently under intensive research and development (R&D), and some challenges still await answers. In this paper, we investigate the perspectives and challenges of VCMA-MRAM for working memories from a cross-layer (device/circuit/architecture) design point of view. We demonstrate that VCMA-MRAM outperforms STT-MRAM and SHE-MRAM in terms of area, speed, energy consumption and instruction-per-cycle (IPC) performance, benefiting from the low-power and high-speed VCMA-driven data writing mechanism. On the other hand, challenges in terms of device fabrication and circuit design should be efficiently addressed before practical applications. Wang Kang 0001, Liang Chang 0002, Youguang Zhang, Weisheng Zhao 0001 |
DATE | 1 |
| 2017 | Energy Efficient Magnetic Tunnel Junction Based Hybrid LSI Using Multi-Threshold UTBB-FD-SOI DeviceabstractThe energy scalability of ultra-low power nonvolatile (NV) large-scale integration (LSI) is explored in this paper. Multi-threshold computing (super/near/sub-$V_t$) in hybrid CMOS/ magnetic tunnel junction (MTJ) circuits are investigated based on SPICE-compatible MTJ model and fully depleted silicon on insulator (FD-SOI) devices. Ultra-low supply voltage operation bottlenecks associated with performance loss, parametric variations and function failure are studied in differential pair-based sensing circuit, MTJ writing/control circuit and other building blocks. A case study is performed with three typical NV-flip-flops (NV-FF), which are implemented with 28nm FD-SOI low $V_t$ (LVT) device and forward back-bias. Results show that MTJ writing/control circuit must operate at nominal supply (super-$V_t$) region to guarantee MTJ switching; sensing circuit is configured with near-$V_t$ operation (0.6V) with robustness consideration, whereas other parts could be implemented with near/sub-$V_t$ computing to achieve ultra-low power consumption and energy efficient operations. Hao Cai 0001, You Wang 0002, Lirida A. B. Naviner, Wang Kang 0001, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2017 | Advanced Low Power Spintronic Memories beyond STT-MRAMabstractUntil now, spin transfer torque magnetic random access memory (STT-MRAM) has drawn considerable R&D interest worldwide. A number of companies and universities are currently involved in this promising technology. In 2016, Everspin released the first 256M STT-MRAM chip, indicating the commercialization and application of STT-MRAM. Nevertheless, STT-MRAM still has some intrinsic limitations, such as dynamic write power and speed, compared with CMOS-based memory technologies. Following the technical evolution process from toggle-MRAM to STT-MRAM, the continuous pursuit of high performance, high density, low power and scalability, drives the intensive R&D of new memory technologies. In this paper, we will show the recent progress in advanced spintronic memories beyond STT-MRAM, such as the spin Hall effect (SHE)-driven and voltage-driven MRAMs. These advanced MRAM technologies do have some unique advantages compared with STT-MRAM, but they also suffer from new design and fabrication challenges. In addition, we will present the latest research in emerging spintronic devices, e.g., magnetic skyrmions, which are potential as information carriers in future spintronic memories, e.g., racetrack memory. Wang Kang 0001, Zhaohao Wang, He Zhang 0011, Youguang Zhang, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2017 | Programmable Stateful In-Memory Computing Paradigm via a Single Resistive DeviceabstractData transfer bandwidth and the related energy consumption has become two of the most critical bottlenecks in conventional von-Newman architecture, owing to the separation of the processor and memory units and the performance mismatch between the two. Realization of the unity of logic computing and data storage in the same die has opened up a promising research direction of in-memory computing (IMC). Meanwhile nonvolatile memory (NVM) based programmable (or reconfigurable) logic architecture has always been a hot topic in the circuit and system societies. To date, lots of interest has been attracted and amazing advance has been made in the two fields, yet none can fully exploit the advantages of both. This paper takes a major step forward by introducing a novel nonvolatile programmable stateful IMC architecture via a single resistive device, which is a completely different design paradigm from previous studies. Each memory cell can perform different Boolean logic functions by dynamically programming the input signals. The computing output result is insitu stored in the memory cell itself and can be readout with a memory-like operation. We will first give a brief review on this filed and then introduce our recent work. We will illustrate how the programmable stateful IMC operations can be implemented via a single resistive device and how the logic computing and data storage can be united within a memory chip. Wang Kang 0001, He Zhang 0011, Youguang Zhang, Weisheng Zhao 0001 |
ICCD | 1 |
| 2017 | Pseudo-Differential Sensing Framework for STT-MRAM: A Cross-Layer PerspectiveabstractWith the rapid increase of leakage currents, non-volatile memories have become competitive candidates in the next-generation computer architecture. Among them, STT-MRAM shows great promise in working memory with high density, high speed and tremendous endurance, etc. However, based on our investigations, the dynamic write power and read reliability are two critical challenges of STT-MRAM. In this work, we propose a synergistic pseudo-differential sensing (PDS) framework that employs device, circuit and architectural techniques to address these challenges. In specific, three design techniques, including cell cluster, asymmetric sensing amplifier and self-error-detection-correction, are proposed to implement the PDS framework. We show that the holistic device-circuit-architecture cross-layer co-design enables STT-MRAM to be utilized in the cache memory, benefiting from the improved density, reliability and energy-efficiency. Our experimental results show that the proposed PDS scheme improves the read margin by ~35.6 percent, reduces the area, read latency, read energy, write latency and write power by ~46.7, ~9.8, ~30.3, ~2.3 and ~31.1 percent respectively, compared with the typical 1T1MTJ cell structure for the cache capacity of 8 MB. In addition, the proposed PDS scheme reduces the dynamic energy by ~32.9 percent and leakage energy by ~830 percent, improves the IPC by ~1.3 percent and miss rate by ~36.9 percent respectively, compared with conventional SRAM based cache. Wang Kang 0001, Liang Chang 0002, Zhaohao Wang, Weifeng Lv, Guangyu Sun 0003, Weisheng Zhao 0001 |
IEEE Trans. Computers | 1 |
| 2016 | PDS: pseudo-differential sensing scheme for STT-MRAMabstractSTT-MRAM has been considered as one of the most promising nonvolatile memory candidates in the next-generation of computer architecture. However, the read reliability and dynamic write power concerns greatly hinder its practical application. In this paper, we propose a synergistic solution, namely pseudo-differential sensing (PDS), to jointly address these two concerns. Three techniques, including cell cluster, asymmetric sensing amplifier (ASA) and self-error-detection-correction (SEDC), are proposed to implement the PDS concept. Our experimental results show that the PDS scheme with the 3T3MTJ cell cluster can reduce the area (~21.7%) and write power (~25.6%) of the differential sensing (DS) scheme while improve the read reliability (read margin, ~35.6%) of the typical sensing (TS) scheme for a 16 Mbit cache. Furthermore, the PDS scheme with the 1T3MTJ cell cluster can outperform both the TS and DS schemes in terms of area (~40.0%, ~66.1%), read latency (~16.6%, ~32.1%), read power (~16.7%, ~37.1%), write latency (~5.4%, 16.3%) and write power (~18.6%, ~43.4%). Wang Kang 0001, Tingting Pang, Bi Wu 0002, Weifeng Lv, Youguang Zhang, Guangyu Sun 0003, Weisheng Zhao 0001 |
DAC | 1 |
| 2016 | Quantitative evaluation of reliability and performance for STT-MRAMabstractDue to its non-volatility, high access speed, ultra low power consumption and unlimited writing/reading cycles, STT-MRAM (Spin Transfer Torque Magnetic Random Access Memory) has emerged as the most promising candidate for the next generation universal memory. However, the process of commercialization of STT-MRAM is hampered by its poor reliability. Generally, these reliability issues are caused by the PVT (Process Variations, Voltage, and Temperature) of both MTJ (Magnetic Tunneling Junction) and transistor. Mitigation and alleviating the impacts of the intrinsic properties and PVT on STT-MRAM is a challenging work. This paper discusses the errors occurring in STT-MRAM resulting from its poor reliability, and analyzes the causes of such errors. To obtain a quantitative assessment of PVT impact on STT-MRAM reliability, we investigate three aspects: writing/reading operation error rate, power consumption and access delay of a single cell. This study is carried out on Cadence platform for 45 nm technology node and the PMA (Perpendicular Magnetic Anisotropy) MTJ model used in the investigation comes from SP INLIB. These quantitative information would be helpful for designing reliability enhancing strategies of STT-MRAM. Liuyang Zhang, Aida Todri, Wang Kang 0001, Youguang Zhang, Lionel Torres, Yuanqing Cheng, Weisheng Zhao 0001 |
ISCAS | 3 |
| 2016 | Read disturbance issue and design techniques for nanoscale STT-MRAM
Yi Ran, Wang Kang 0001, Youguang Zhang, Jacques-Olivier Klein, Weisheng Zhao 0001 |
J. Syst. Archit. | 2 |
| 2016 | Skyrmion-Electronics: An Overview and OutlookabstractThe well-known empirical phenomenon known as Moore's Law has held true for the past half century. However, it is beginning to break down, owing to limitations arising from leakage currents caused by the quantum effect. As a result, the search for alternatives or complementary technologies that can aid the downscaling of complementary metal-oxide-semiconductor (CMOS) technology has been accelerated in the field of electronics. Among various potential candidates, spintronic technology has attracted considerable interest and attention, especially for the topological spin textures known as magnetic skyrmions. Magnetic skyrmions are expected to have topologically protected stability and nanoscale size, and require a very low driving current density, therefore they are considered as potential building blocks for future spintronic devices and integrated circuits. Furthermore, recent experimental demonstrations of the control of individual nanometer-scale skyrmions, including their creation, detection, transportation, and manipulation at room temperature, further highlight their potential for future electronic applications. In this paper, we review the current status and outlook of skyrmions from the viewpoint of electronic applications. First, the fundamental and elementary functionality of skyrmions, such as electric write-in, read-out, transmission, and manipulation, are introduced. Then, potential electronic applications of skyrmions for nonvolatile memory and logic circuits are described with case studies. Finally, we conclude with an analysis of current challenges, limitations, and future trends of skyrmion research. Wang Kang 0001, Yangqi Huang, Xichao Zhang, Weisheng Zhao 0001 |
Proc. IEEE | 1 |
| 2015 | Spintronics: Emerging Ultra-Low-Power Circuits and Systems beyond MOS TechnologyabstractConventional MOS integrated circuits and systems suffer serve power and scalability challenges as technology nodes scale into ultra-deep-micron technology nodes (e.g., below 40nm). Both static and dynamic power dissipations are increasing, caused mainly by the intrinsic leakage currents and large data traffic. Alternative approaches beyond charge-only-based electronics, and in particular, spin-based devices, show promising potential to overcome these issues by adding the spin freedom of electrons to electronic circuits. Spintronics provides data non-volatility, fast data access, and low-power operation, and has now become a hot topic in both academia and industry for achieving ultra-low-power circuits and systems. The ITRS report on emerging research devices identified themagnetic tunnel junction(MTJ) nanopillar (one of the Spintronics nanodevices) as one of the most promising technologies to be part of future micro-electronic circuits. In this review we will give an overview of the status and prospects of spin-based devices and circuits that are currently under intense investigation and development across the world, and address particularly their merits and challenges for practical applications. We will also show that, with a rapid development of Spintronics, some novel computing architectures and paradigms beyond classic Von-Neumann architecture have recently been emerging for next-generation ultra-low-power circuits and systems. Wang Kang 0001, Yue Zhang 0010, Zhaohao Wang, Jacques-Olivier Klein, Claude Chappert, Dafine Ravelosona, Gefei Wang, Youguang Zhang, Weisheng Zhao 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2014 | An overview of spin-based integrated circuitsabstractConventional CMOS integrated circuits suffer from serve power and scalability challenges as technology node scales into ultra-deep-micron technology nodes. Alternative approaches beyond charge-only based circuits. In particular, spin-based devices or integrated circuits show promising merits to overcome these issues by adding the spin freedom of electrons to the electronic circuits. Spintronics has now become a hot topic in both academics and industrials. This paper overviews the status and prospects of spin-based integrated circuits under intense investigation and address particularly their merits and challenges for practical applications. Wang Kang 0001, Weisheng Zhao 0001, Zhaohao Wang, Jacques-Olivier Klein, Yue Zhang 0010, Djaafar Chabi, Youguang Zhang, Dafine Ravelosona, Claude Chappert |
ASP-DAC | 1 |
| 2014 | Spintronics for low-power computingabstractMicroelectronics has been following Moore's law for almost 40 years. However this trend tends to run out of steam in recent technology nodes. The continuous improvements in the size of the transistors and in the operating frequencies result in serious power consumption, heat dissipation and reliability issues. Spintronics (Nobel Prize of Physics, 2007 awarded to Prof. Fert from Univ. Paris-Sud and Peter Grünberg from Forschungszentrum Jülich) nanodevices can reduce significantly the power, improve the reliability or allow new functionalities. The 2010 ITRS report on emerging research devices identified Magnetic Tunnel Junction (MTJ) nanopillar (the preeminent spintronics nanodevice) as one of the most promising technologies to be part of the future microelectronics circuits. It provides data non-volatility, hardness to radiations, fast data access and low-power operations. Magnetic memories become the most promising candidate for both low power logic computing and the data storage. This tutorial paper presents multi-discipline questions (Device, Circuit, Architecture, System and CAD) related to this topic to share the most recent results and discuss the future challenges. Yue Zhang 0010, Weisheng Zhao 0001, Jacques-Olivier Klein, Wang Kang 0001, Damien Querlioz, Youguang Zhang, Dafine Ravelosona, Claude Chappert |
DATE | 4 |
| 2014 | Ferroelectric tunnel memristor-based neuromorphic network with 1T1R crossbar architectureabstractEmerging ferroelectric tunnel memristors show large OFF/ON resistance ratio (>100) and high operation speed (~10ns), promising to be widely applied in the future synapse-like systems. In this paper we propose a neuromorphic network with ferroelectric tunnel memristor. This network is arranged with classical crossbar topology, in which each crosspoint forms a synapse consisting of a MOS transistor and a memristor. Based on this architecture, we design a spike-timing dependent plasticity (STDP) scheme and a parallel supervised learning circuit. Using a compact model of ferroelectric tunnel memristor and CMOS 40nm design kit, we perform transient simulation to validate the functionality of the proposed STDP and learning circuit. Simulation results show the potential of our neuromorphic network in low power (~100nA or ~1μA) and high speed (μs or ~100ns) computing system. Zhaohao Wang, Weisheng Zhao 0001, Wang Kang 0001, Youguang Zhang, Jacques-Olivier Klein, Claude Chappert |
IJCNN | 3 |
| 2014 | Design and analysis of crossbar architecture based on complementary resistive switching non-volatile memory cells
Weisheng Zhao 0001, Jean-Michel Portal, Wang Kang 0001, Mathieu Moreau, Yue Zhang 0010, Hassen Aziza, Jacques-Olivier Klein, Zhaohao Wang, Damien Querlioz, Damien Deleruyelle, Marc Bocquet, Dafine Ravelosona, Christophe Muller, Claude Chappert |
J. Parallel Distributed Comput. | 3 |