EDBT 2026 Demo / reviewers in the wild / expert
Weisheng Zhao 0001
dblp:48/5196-1
· DBLP profile ↗
169ranked-venue papers
6as first author
83since 2021 · last 2026
0000-0001-8088-0404ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 152 · 6 first-author · 71 since 2021Software engineering, systems software and programming languages · 16 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 9 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Effective SNN Macro with Real-Time STDP and Dynamic LIF Model Based on Thermally Interplayed Spin-Orbit Torque MTJabstractSpiking neural networks (SNNs) have emerged as a promising paradigm for effective event-driven computation. However, CMOS-based SNN designs are limited by power consumption and complexity, while nonvolatile memory (NVM)-based SNN designs often lack biological characteristics and require active capacitive circuits to emulate neuronal dynamics. In this paper, we propose a thermally interplayed spin-orbit torque magnetic tunnel junction (TI-MTJ) macro that integrates core SNN functionalities. Our neuron array autonomously achieves leaky integrate-and-fire (LIF) model within the TI-MTJ device, thus improving power efficiency and simplifying circuit structure. Additionally, the proposed synaptic array provides adaptive in-situ responses based on a simplified spike-timing-dependent plasticity (STDP) rule. To enhance biological plausibility, our macro incorporates real-time spike monitoring and inhibition mechanisms. A comprehensive device-circuit-algorithm co-optimization framework validates the high performance of the TI-MTJ macro, achieving a synaptic energy consumption of 6.07fJ per spike, an inference accuracy of 97.76% on the MNIST dataset, and an energy efficiency of 22.8TOPS/W. Changyu Li, Linjun Jiang, Liangchen Li, Dehang Zhu, Junda Zhao, Wang Kang 0001, Wenlong Cai, He Zhang 0011, Weisheng Zhao 0001 |
DATE | 10 |
| 2026 | Late Breaking Results: Algorithm-Hardware Co-Design of a Sparsity-Aware Dense-Sparse Scheme for DNN AcceleratorsabstractDeep neural networks (DNNs) in modern applications have increased the demand for energy-efficient DNN inference solutions, especially on resource-constrained platforms. However, the growing model capacity of DNNs incurs significant memory traffic and energy consumption. To address these challenges, we propose a novel solution that presents an algorithm-hardware co-design for reconfigurable DNN acceleration. This design exploits value- and bit-level sparsity to minimize memory footprint and enhance computational efficiency. To achieve this, the proposed algorithm leverages a static dense-sparse storage format, along with a dynamic bit-processing scheme that removes non-contributing bits. Building on this algorithm, a flexible processing element array is designed to perform LUT-based shift-accumulate operations, with fine-grained per-layer configurability. Experimental results show that this design yields 13–24% storage savings across the evaluated DNN models, while delivering up to 8.4× effective sparsity. Based on post-implementation FPGA results (from our RTL design), the proposed accelerator delivers 1.41× lower LUT usage than state-of-the-art design at similar throughput. Yueting Li 0001, Terry Tao Ye, Weisheng Zhao 0001 |
DATE | 4 |
| 2026 | Input Sparsity Aware In-Memory Computing Macro Based on SOT-MRAM Multi-Level Cell for Efficient Deep Neural Network AccelerationabstractDeep neural network (DNN) technology has gained widespread applications, but its high energy demands continue to drive the advancement of low-power computing architectures, particularly in in-memory computing (IMC) architectures based on non-volatile memory. Among these, spin-transfer torque magnetic random-access memory (STT-MRAM)-based IMC architectures have achieved some progress, but their performance remains constrained by limited resistance and binary characteristics. By contrast, the next-generation spin-orbit torque MRAM (SOT-MRAM) offers superior magnetic tunnel junction (MTJ) resistance and more flexible cell structures, presenting significant potential for energy-efficient IMC implementation. In this work, leveraging the ultra-high MTJ resistance and the separation of read/write paths in SOT-MRAM, we propose a multi-level cell (MLC) structure-based high energy-efficiency IMC architecture (MLC-SOT-IMC), which performs standard multiplication operations by optimizing the conductance mapping paradigm. The proposed architecture not only maintains high inference accuracy but also significantly enhances integration density and reduces the overhead per bit. Additionally, a self-terminating time-to-digital converter (TDC) readout circuit, which is dependent on input sparsity, is introduced to eliminate the excess power consumption associated with ineffective pulses after readout completion. Ultimately, the proposed MLC-SOT-IMC architecture achieves an inference energy efficiency of 6388.98 1-bit TOPS/W under an input sparsity of 50%, with the peak energy efficiency reaching 8426.19 1-bit TOPS/W at an input sparsity of 90%. Chao Wang 0094, Qihang Gao, Xianzeng Guo, Zhongzhen Tong, Zhaohao Wang, Weisheng Zhao 0001 |
DATE | 6 |
| 2026 | High-Performance and High-Density NAND-Like SOT-MRAM for FinFET Technology NodesabstractThis paper proposes a comprehensive optimization framework for NAND-like spintronics memory (NAND-SPIN) in advanced FinFET technology nodes. At bit-cell structure level, we propose a NAND-SPIN-GND design which is configured with a grounded bit line (BL) to minimize the parasitic resistance in both read and write paths, thereby decreasing read latency by 35.0% and write energy by 27.3%. At device and layout level, a short-circuiting bottom electrode (SBE) design is proposed, which shorts non-contributing spin-orbit torque (SOT) segments by the BEs, reducing read latency by up to 41.4% and write energy by 55.8%. In addition, a compact capacitance symmetric source line (SL)-type reference scheme is introduced to address the inherent capacitance asymmetry in conventional SL-type reference scheme, resulting in a 48.6% reduction in read latency compared to the conventional word line (WL)-type reference scheme. Chao Wang 0094, Xianzeng Guo, Luman Xiang, Zhaohao Wang, Weisheng Zhao 0001 |
DATE | 5 |
| 2026 | A 6.86Tb/s Bandwidth SOT-MRAM Sensing Scheme with Configurable Full-Column Over Frequency Technique for Near Memory Computing
Xinpeng Jiang, Hanting Chen, Zhaohao Wang, He Zhang 0011, Weisheng Zhao 0001 |
ISCAS | 5 |
| 2026 | High-performance true random number generator based on SOT-MTJ spin relaxation
Jialiang Yin, Xiuye Zhang, Wenlong Cai, Ao Du, Binchao Tang, Shijian Bao, Daoqian Zhu, Kewen Shi, Lang Zeng, He Zhang 0011, Kaihua Cao, Weisheng Zhao 0001 |
Sci. China Inf. Sci. | 17 |
| 2026 | ReNN-RV: Run-Time PE Reconfiguration for DNN Inference Acceleration With Custom RISC-V ISAabstractDeep neural network (DNN) accelerators integrated with RISC-V Instruction Set Architecture (ISA) extensions have enabled efficient computing on resource-constrained platforms. However, their specialization in regular compute patterns limits effectiveness on irregular workloads, making it challenging to achieve high throughput and energy efficiency. To tackle these challenges, we present ReNN-RV, which integrates a computation-aware RISC-V ISA extension with an instructiondriven processing pipeline to efficiently accelerate run-time reconfigurable processing elements (RePEs). The computation-aware ISA employs configurable opcodes and custom encodings to support fine-grained task scheduling, while an instructiondriven pipeline implements it with minimal control complexity. Moreover, theRePEaccelerator provides seamless switching between multiply-accumulate (MAC) and non-MAC operations by configuring a path multiplexer to realize multiple operators at run time. Experimental results demonstrate that ReNN-RV achieves average reductions of 14.6× in cycle count and 15.3× in execution time across representative DNN workloads compared with the baseline RISC-V design. On average, ReNN-RV outperforms state-of-the-art designs by 10.1× for energy efficiency and 10.3× for computational throughput. Yueting Li 0001, Terry Tao Ye, Ngai Wong 0001, Zhenhua Zhu 0002, Yongfu Li 0002, Weisheng Zhao 0001 |
IEEE Trans. Computers | 6 |
| 2026 | CIMinus: Empowering Sparse DNN Workloads Modeling and Exploration on SRAM-Based CIM ArchitecturesabstractCompute-in-memory (CIM) has emerged as a pivotal direction for accelerating workloads in the field of machine learning, such as Deep Neural Networks (DNNs). However, the effectively exploitation of sparsity in CIM systems presents numerous challenges, due to the inherent limitations in their rigid array structures. Designing sparse DNN dataflows and developing efficient mapping strategies also become more complex when accounting for diverse sparsity patterns and the flexibility of a multi-macro CIM structure. Despite these complexities, there is still an absence of a unified systematic view and modeling approach for diverse sparse DNN workloads in CIM systems. In this paper, we propose CIMinus, a framework dedicated to cost modeling for sparse DNN workloads on CIM architectures. It provides an in-depth energy consumption analysis at the level of individual components and an assessment of the overall workload latency. We validate CIMinus against contemporary CIM architectures and demonstrate its applicability in two use-cases. These cases provide valuable insights into both the impact of sparsity patterns and the effectiveness of mapping strategies, bridging the gap between theoretical design and practical implementation. Yingjie Qi, Jianlei Yang 0001, Rubing Yang, Cenlin Duan, Xiaolin He, Ziyan He, Weitao Pan, Weisheng Zhao 0001 |
IEEE Trans. Computers | 8 |
| 2026 | GCoDE: Efficient Device-Edge Co-Inference for GNNs via Architecture-Mapping Co-SearchabstractGraph Neural Networks (GNNs) have emerged as the state-of-the-art graph learning method. However, achieving efficient GNN inference on edge devices poses significant challenges, limiting their application in real-world edge scenarios. This is due to the high computational cost of GNNs and limited hardware resources on edge devices, which prevent GNN inference from meeting real-time and energy requirements. As an emerging paradigm, device-edge co-inference shows potential for improving inference efficiency and reducing energy consumption on edge devices. Despite its potential, research on GNN device-edge co-inference remains scarce, and our findings show that traditional model partitioning methods are ineffective for GNNs. To address this, we propose GCoDE, the first automatic framework forGNN architecture-mappingCo-design and deployment onDevice-Edge hierarchies. By abstracting the device communication process into an explicit operation, GCoDE fuses the architecture and mapping scheme in a unified design space for joint optimization. Additionally, GCoDE’s system performance awareness enables effective evaluation of architecture efficiency across diverse heterogeneous systems. By analyzing the energy consumption of various GNN operations, GCoDE introduces an energy prediction method that improves energy assessment accuracy and identifies energy-efficient solutions. Using a constraint-based random search strategy, GCoDE identifies the optimal solution in 1.5 hours, balancing accuracy and efficiency. Moreover, the integrated co-inference engine in GCoDE enables efficient deployment and execution of GNN co-inference. Experimental results show that GCoDE can achieve up to 44.9× speedup and 98.2% energy reduction compared to existing approaches across diverse applications and system configurations. Jianlei Yang 0001, Yingjie Qi, Zhi Yang 0001, Weisheng Zhao 0001, Chunming Hu |
IEEE Trans. Computers | 6 |
| 2026 | Efficient SRAM-PIM Co-Design by Joint Exploration of Value-Level and Bit-Level SparsityabstractProcessing-in-memory (PIM) architectures mitigate the Von Neumann bottleneck by integrating computation units into memory arrays. Among PIM architectures, digital SRAMPIM has become a prominent approach, directly integrating digital logic within the SRAM array. However, the rigid crossbar architecture and full array activation pose challenges in efficiently utilizing value-level sparsity. Moreover, neural network models exhibit a high proportion of zero bits within non-zero values, which remain underutilized due to architectural constraints. To overcome these limitations, we present Dyadic Block PIM (DB-PIM), a groundbreaking algorithm-architecture co-design framework to harness both value-level and bit-level sparsity. At the algorithm level, our hybrid-grained pruning technique, combined with a novel sparsity pattern, enables effective sparsity management. Architecturally, DB-PIM incorporates a sparse network and customized digital SRAM-PIM macros, including input pre-processing unit (IPU), dyadic block multiply units (DBMUs), and Canonical Signed Digit (CSD)-based adder trees. It circumvents structured zero values in weights and bypasses unstructured zero bits within non-zero weights and block-wise all-zero bit columns in input features. As a result, the DBPIM framework skips a majority of unnecessary computations, thereby driving significant gains in computational efficiency. Experimental results demonstrate that our DB-PIM framework achieves up to 8.01× speedup and 85.28% energy savings, significantly boosting computational efficiency in digital SRAMPIM systems. Cenlin Duan, Jianlei Yang 0001, Yiou Wang, Yingjie Qi, Xiaolin He, Bonan Yan, Xiaotao Jia, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 10 |
| 2026 | ACE-GNN: Adaptive GNN Co-Inference With System-Aware Scheduling in Dynamic Edge EnvironmentsabstractThe device-edge co-inference paradigm effectively bridges the gap between the high resource demands of Graph Neural Networks (GNNs) and limited device resources, making it a promising solution for advancing edge GNN applications. Existing research enhances GNN co-inference by leveraging offline model splitting and pipeline parallelism (PP), which enables more efficient computation and resource utilization during inference. However, the performance of these static deployment methods is significantly affected by environmental dynamics such as network fluctuations and multi-device access, which remain unaddressed. We present ACE-GNN, the first Adaptive GNN Co-inference framework tailored for dynamic Edge environments, to boost system performance and stability. ACE-GNN achieves performance awareness for complex multi-device access edge systems via system-level abstraction and two novel prediction methods, enabling rapid runtime scheme optimization. Moreover, we introduce a data parallelism (DP) mechanism in the runtime optimization space, enabling adaptive scheduling between PP and DP to leverage their distinct advantages and maintain stable system performance. Also, an efficient batch inference strategy and specialized communication middleware are implemented to further improve performance. Extensive experiments across diverse applications and edge settings demonstrate that ACE-GNN achieves a speedup of up to 12.7× and an energy savings of 82.3% compared to GCoDE, as well as 11.7× better energy efficiency than Fograph. Jianlei Yang 0001, Yingjie Qi, Xinming Wei, Cenlin Duan, Weisheng Zhao 0001, Chunming Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2026 | NFGen: Normalizing Flow-Based Joint Generative Model for Variability-Aware Design Technology Co-OptimizationabstractAs transistor sizes continue shrinking, impacts of variability has become ever more paramount in circuit design and manufacturing. Their accurate representations in model cards help save design margins and provide appropriate guidelines in design technology co-optimization (DTCO). To address such a challenge, we propose a novel machine learning framework, Normalizing Flow-Based Joint Generative Model (NFGen), which generates a comprehensive model library from a limited number of model cards. Unlike traditional generative methods that focus on the marginal distribution of model card parameters, NFGen is the first model to approximate their joint distribution, which includes information on their correlation and thus enables closer representation of variability effects. In addition, we introduce two similarity metrics to rigorously evaluate the quality of generated model cards. Experimental results show that NFGen reduces overall error by 2x to 8x compared to state-of-the-art methods, validating its superiority in variability-aware DTCO. Zhenxing Dou, Yijiao Wang, Peng Wang 0022, Runsheng Wang, Weisheng Zhao 0001, A. Asenov |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2026 | BaM-CIM: A High Throughput Booth Algorithm-Based In-MRAM Computing Macro Using Hybrid VGSOT-MTJ/GAA-CNTFETabstractAs artificial intelligence (AI) and computational models grow in scale, the demand for computational power and storage has significantly increased. The computing-in-memory (CIM) architecture addresses this challenge by performing computations directly within the memory array, reducing data transfer between the processor and memory. This paper introduces a Booth algorithm-based In-MRAM computing architecture (BaM-CIM) using a hybrid voltage-gated spin-orbit torque MTJ (VGSOT-MTJ) and gate-all-around carbon nanotube field-effect transistors (GAA-CNTFETs) for efficient multiply-and-accumulate (MAC) computing. The key contributions of BaM-CIM are as follows: 1) A Voltage divider reference (VDR) cell is proposed, which enables read operations using only a 2T1M cell structure. Compared to complementary read cells, the VDR reduces the area by half and achieves robust data sensing without requiring precharge/discharge operations. 2) The BaM-CIM circuit is proposed to complete 8b-W/8b-IN/21b-OUT computations in only two cycles (1.6 ns), reducing the number of cycles by 75% compared to single-bit input serial operations and by 50% compared to two-bit serial operations. 3) A three-input 8b Booth computing adder (BCA), along with Modified computing shift adder (MCSA) and Modified computing post adder (MCPA), which can achieve higher energy efficiency. BaM-CIM with 128 Kb is simulated, achieving throughput and energy efficiency of 0.93 TOPS and 258.4 TOPS/W, respectively, at a 0.6 V supply voltage and 1.28 TOPS and 169.5 TOPS/W, respectively, at a 0.8 V supply voltage with 8b-IN, 8b-W, and 21b-OUT. Chenghang Li, Zhongzhen Tong, Yulong Qiu, Jiye Yao, Chao Wang 0094, Zhaohao Wang, Xiaoyang Lin, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2026 | An FD-SOI-Based Compact In-Pixel Computing Architecture Enabling Real-Time Feature ExtractionabstractTo empower resource-limited edge devices in artificial intelligence (AI) and Internet of Things (IoT) applications, it is essential to overcome challenges posed by restricted area resources and the high latency demands of transmitting and processing substantial sensory data. In-pixel computing addresses these challenges effectively, and the Fully Depleted Silicon-On-Insulator (FD-SOI)-based pixel, which relies on an FD-SOI transistor whose current is made photosensitive to light by applying a negative back-gate voltage, shows significant potential with its compact structure and in-situ computation capability. In this paper, for the first time, we present an FD-SOI-based chip-level architecture for in-pixel computing. Our design implements programmable, massively parallel convolution with low latency using a pulse-width modulation (PWM) input encoding scheme. Furthermore, the proposed compact 1P1T (1 Phototransistor 1 Transistor) pixel design, integrated with an improved single-slope analog-to-digital converter (SS ADC), greatly enhances area efficiency. Validated through simulation in a 22nm FD-SOI process, the design achieves 990 frames/s under typical outdoor illumination conditions, with the figure of merit (FoM) of 16.86pJ/pixel/frame. In addition, the proposed architecture has been evaluated on hand gesture recognition (5697 training and 633 validation images across six categories), achieving an accuracy of 97.48%. The results demonstrate that, compared to state-of-the-art designs, our approach achieves a$6\times $improvement in in-pixel convolution speed and a$7.3\times $reduction in area overhead. Yijiao Wang, Jiayao Wu, Zhongzhen Tong, Xinrui Duan, Yiming Shi, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2026 | High-Efficiency and Low-Deviation Analog-Digital Hybrid Compute-in-Memory Architecture With Dynamic Weight DivisionabstractCompute-in-memory (CIM) reduces data movement but suffers from an accuracy–efficiency trade-off: Analog CIM (ACIM) is energy-efficient but loses accuracy and incurs higher cost at large bit-widths, while digital CIM (DCIM) supports high precision but is inefficient for low-precision tasks. To overcome these challenges, we propose an analog–digital hybrid CIM (HCIM) architecture to address this trade-off, including 1) an analog–digital hybrid 10T SRAM cell without additional transistors and a dual-capacitor-based multicycle weighting module to reduce area; 2) a successive-approximation-register (SAR) ADC with a pseudo C-2C capacitor array that can be reconfigured from an 8-bit ADC into two parallel 4-bit ADCs to improve configurability; 3) configurable weight division and computing resource allocation strategies. Simulations in a 28-nm process show that HCIM achieves 15.56 TOPS/W at 12-bit ($8+4$) with$1.33\times $and$2.35\times $efficiency improvement over DCIM and ACIM and$16\times $lower error. It achieves 27.87 TOPS/W at 8-bit and 78.13 TOPS/W at 4-bit, demonstrating superior energy efficiency, computational accuracy, and flexibility. Linjun Jiang, Sifan Sun, Wente Yi, Dengwen Li, Wang Kang 0001, He Zhang 0011, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 10 |
| 2026 | TinyFormer: Efficient Sparse Transformer Design and Deployment on Tiny DevicesabstractDeveloping deep learning models on tiny devices (e.g. Microcontroller units, MCUs) has attracted much attention in various embedded IoT applications. However, it is challenging to efficiently design and deploy recent advanced models (e.g. transformers) on tiny devices due to their severe hardware resource constraints. In this work, we proposeTinyFormer, a framework specifically designed to develop and deploy resource-efficient transformer models on MCUs. TinyFormer consists ofSuperNAS,SparseNAS, andSparseEngine. Separately, SuperNAS aims to search for an appropriate supernet from a vast search space. SparseNAS evaluates the best sparse single-path transformer model from the identified supernet. Finally, SparseEngine efficiently deploys the searched sparse models onto MCUs. To the best of our knowledge, SparseEngine is the first deployment framework capable of performing inference of sparse transformer models on MCUs. Evaluation results on the CIFAR-10 dataset demonstrate that TinyFormer can design efficient transformers with an accuracy of 96.1% while adhering to hardware constraints of 1MB storage and 320KB memory. Additionally, TinyFormer achieves significant speedups in sparse inference, up to$12.2\times $comparing to the CMSIS-NN library. TinyFormer is believed to bring powerful transformers into TinyML scenarios and to greatly expand the scope of deep learning applications. Jianlei Yang 0001, Jiacheng Liao, Fanding Lei, Meichen Liu, Lingkun Long, Han Wan, Bei Yu 0001, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 9 |
| 2025 | A Custom RISC-V ISA with Scalable Processing Units for Efficient Neural Network InferenceabstractA customized RISC-V ISA with integrated digital accelerators offers a promising solution to improve energy efficiency in neural network inference.However, it often requires multiple instructions per accelerator operation, which limits computational efficiency during deep neural network inference.To overcome the instruction overhead, this design introduces a dedicated instruction set that enables scalable and fine-grained accelerator control.By incorporating the pattern-driven instruction mode, this design exploits the neural layer regularity to support efficient instruction iteration.Furthermore, this digital accelerator leverages hardware reuse for logic operations, forming a fusion-style architecture that integrates reconfigurable components.Experimental results demonstrate that the custom RISC-V ISA achieves an average runtime speedup of 8.26× and reduces the instruction count by 14.71×.This design also yields an average 8.73× reduction in cycles per instruction across MobileNetV2, ResNet50, VGG19, EfficientNet, and DenseNet-BC, validating its effectiveness across representative benchmarks.Additionally, it improves average energy efficiency by 1.74×, outperforming state-of-the-art designs. Yueting Li 0001, Wanshuang Lin, Wendong Xu, Ngai Wong 0001, Weisheng Zhao 0001 |
CF | 5 |
| 2025 | MIRACLE: Multimodal Information Retrieval via a Combined In-Memory Processing and Content Addressable Memory ApproachabstractThe rapid advancement of information technology has brought multimodal information retrieval into the research spotlight. Neural networks, particularly Transformers, have emerged as the dominant solution for extracting multimodal feature vectors. While neural network acceleration has been extensively explored, the subsequent retrieval stage in multimodal scenarios remains under-optimized. Conventional retrieval approaches, such as cosine similarity sorting on von Neumann architectures, suffer from significant data migration and computational inefficiencies. Hashing methods enhance storage and computation efficiency but encounter challenges in energy-efficient implementation and mitigating accuracy losses due to modal heterogeneity. This paper presents a hybrid architecture that integrates in-memory processing (PIM) and content-addressable memory (CAM) to address these challenges. Transformer-extracted features are processed via in-memory random hashing leveraging device-intrinsic properties, with CAM facilitating parallel search space reduction. A final cosine similarity reranking stage refines the results while balancing accuracy with energy efficiency. Experimental evaluations validate that the proposed method, when compared to the baseline traditional CPU-based cosine similarity retrieval, 1) achieves almost identical level of accuracy, dramatically outperforming other pure CAMbased Hamming distance retrieval approaches; and 2) reduces latency by $9.45 \times$ and energy consumption by $30.20 \times$. Xuehui Liu, Tianyang Yu, Shuo Ran, Bi Wu 0002, Xiaotao Jia, Weiqiang Liu 0001, Gang Qu 0001, Weisheng Zhao 0001 |
DAC | 10 |
| 2025 | CIMFlow: An Integrated Framework for Systematic Design and Evaluation of Digital CIM ArchitecturesabstractDigital Compute-in-Memory (CIM) architectures have shown great promise in Deep Neural Network (DNN) acceleration by effectively addressing the “memory wall” bottleneck. However, the development and optimization of digital CIM accelerators are hindered by the lack of comprehensive tools that encompass both software and hardware design spaces. Moreover, existing design and evaluation frameworks often lack support for the capacity constraints inherent in digital CIM architectures. In this paper, we present CIMFlow, an integrated framework that provides an out-of-the-box workflow for implementing and evaluating DNN workloads on digital CIM architectures. CIMFlow bridges the compilation and simulation infrastructures with a flexible instruction set architecture (ISA) design, and addresses the constraints of digital CIM through advanced partitioning and parallelism strategies in the compilation flow. Our evaluation demonstrates that CIMFlow enables systematic prototyping and optimization of digital CIM architectures across diverse configurations, providing researchers and designers with an accessible platform for extensive design space exploration. Yingjie Qi, Jianlei Yang 0001, Yiou Wang, Dayu Wang, Cenlin Duan, Xiaolin He, Weisheng Zhao 0001 |
DAC | 9 |
| 2025 | An Adaptive Sparse Matrix Compression CIM Accelerator based on 256Kb SOT-MRAM for Downlink Massive MIMO CommunicationsabstractDownlink precoding in massive multiple input multiple output (MIMO) systems involves high-dimensional sparse matrix calculations, which poses challenges to existing architectures. Computing-in-memory (CIM) has significant advantages in handling large-scale parallel operations, but sparse computing for wireless communication remains underexplored. In this paper, we propose a novel CIM accelerator based on magnetic random access memory (MRAM) leveraging adaptive multi-sparse mode technology for optimized sparse matrix multiplication in MIMO communication systems. This architecture represents the first application of CIM technology for processing sparse matrices in MIMO precoding tasks, minimizing storage requirements and enhancing parallel processing speed. Experimental results demonstrate that, for a 32×256×8 MIMO downlink precoding task with 90% sparsity, the symbol error rate is reduced to 0.1% at a signal-to-noise ratio of 20dB, achieving 8.35× reduction in storage overhead, 39.4× power saving and 9.85× speedup. These results position our accelerator as a promising candidate for processing sparse data in 5G massive MIMO systems. Liangchen Li, Changyu Li, Anyang Yu, Junda Zhao, Zhaohao Wang, Chengyuan Sun, Kaihua Cao, Wang Kang 0001, He Zhang 0011, Weisheng Zhao 0001 |
ICCAD | 13 |
| 2025 | Ultra Energy-Efficient Butterfly Counting in Bipartite Networks via Algorithm-Architecture Co-OptimizationabstractButterfly counting (BFC) problem, which counts the number of butterfly structure in a graph, is fundamental in bipartite network analysis. Recently, considerable efforts toward accelerations of BFC on both CPU and GPU platforms have been reported. However, the underlying BFC algorithms require repetitive vertex traversal and suffer from substantial latency and energy consumption concerns because data in large graphs has very limited reusability. In this paper, we introduce a hardware-software co-optimization approach to tackle these issues. A key innovation behind our approach is an algorithm that employs iterative lightweight arithmetical operations and facilitates highly parallel and pipelined processing. We further develop optimized data compression and pruning strategies to improve the efficiency of processing sparse data. These pivotal advancements are seamlessly integrated with a purpose-built hardware architecture to augment the overall implementation efficiency. Our proposed strategies are thoroughly evaluated on Zynq UltraScale+ FPGA platform. Compared with the state-of-the-art CPU (with 512 GB DRAM) and CPU+GPU (with 128 GB DRAM) implementations, our design achieve speedups of 15.84×and 1.35×, respectively, with only 4 GB DRAM. Meanwhile, our design’s energy efficiency is 50.14× over the CPU+GPU accelerator. Jianlei Yang 0001, Xiaotao Jia, Gang Qu 0001, Weisheng Zhao 0001 |
ICCAD | 6 |
| 2025 | HRAMTran: A Hybrid-RAM Transformer Accelerator With Dynamic Sparsity Floating-Point CIM and Written-Back Transpose ArrayabstractTransformer model performs outstandingly in various tasks involving artificial intelligence. In this work, we propose a hybrid-RAM Transformer accelerator (HRAMTran) utilizing computing-in-memory (CIM) based on spin-orbit torque magnetic random access memory (SOT-MRAM) and static RAM (SRAM), which supports dynamic sparsity in floating-point (FP) matrix multiplication (MM) and written-back transpose, thereby realizing efficient attention mechanism. First, a dynamic sparsity-based MM scheme is proposed, which dynamically ignores low-impact elements during vector multiplication, thereby effectively reducing the latency and energy consumption of MM. Second, a data-reuse multiply-and-accumulate (MAC) scheme for mantissa is designed to further optimize MM, which shares partial operation result to reduce redundant computation. The SOT-MRAM and SRAM based CIM architectures with dynamic sparsity and data-reuse schemes are constructed to perform weight (Query (Q), Key (K), and Value (V)) and dynamic MM, respectively. This hybrid-RAM CIM method can realize the optimization of energy and latency during attention mechanism computation. Moreover, written-back transpose SRAM array that can write multiple bits into a column simultaneously is designed to significantly reduce write-back cycles for KT. Finally, the HRAMTran accelerator is built to evaluate the performance of transformer implementation through performing machine translation for the WMT14 dataset. Results show that this accelerator realizes 3.6 µJ/Token and 68.77 TFLOPS/W, achieving 4.33× and 2.39× improvement compared with the state-of-the-art transformer accelerator. Xianan Zhu, Zhengkun Gu, Zhizhong Zhang 0004, Kun Zhang 0030, Weisheng Zhao 0001, Yue Zhang 0010 |
ICCAD | 9 |
| 2025 | Finesse: An Agile Design Framework for Pairing-based Cryptography via Software/Hardware Co-DesignabstractPairing-based cryptography (PBC) is crucial in modern cryptographic applications.With the rapid advancement of adversarial research and the growing diversity of application requirements, PBC accelerators need regular updates in algorithms, parameter configurations, and hardware design.However, traditional design methodologies face significant challenges, including prolonged design cycles, difficulties in balancing performance and flexibility, and insufficient support for potential architectural exploration.To address these challenges, we introduce Finesse, an agile design framework based on co-design methodology.Finesse leverages a co-optimization cycle driven by a specialized compiler and a multi-granularity hardware simulator, enabling both optimized performance metrics and effective design space exploration.Furthermore, Finesse adopts a modular design flow to significantly shorten design cycles, while its versatile abstraction ensures flexibility across various curve families and hardware architectures.Finesse offers flexibility, efficiency, and rapid prototyping, comparing with previous frameworks.With compilation times reduced to minutes, Finesse enables faster iteration cycles and streamlined hardware-software co-design.Experiments on popular curves * Both authors contributed equally to this research. Tianwei Pan, Tianao Dai, Jianlei Yang 0001, Hongbin Jing, Zeyu Hao, Xiaotao Jia, Chunming Hu, Weisheng Zhao 0001 |
ISCA | 9 |
| 2025 | Model quantization for computing-in-memory: a survey
Sifan Sun, Jinyu Bai, Hanting Chen, Kaiwen Deng, Zhiwei Xie 0009, He Zhang 0011, Wang Kang 0001, Weisheng Zhao 0001 |
Sci. China Inf. Sci. | 10 |
| 2025 | Orbitronics for energy-efficient magnetization switching
Daoqian Zhu, Qingtao Xia, Jianing Liang, Zhiyang Peng, Chen Xiao, Renyou Xu, Xiantao Shang, Shiyang Lu, Dapeng Zhu, Kaihua Cao, Weisheng Zhao 0001 |
Sci. China Inf. Sci. | 16 |
| 2025 | AM-CIM: Approximate Memory Based Near Sensor Compute-in-Memory Architecture for Keyword SpottingabstractCompute-In-Memory (CIM) has emerged as a promising solution to address the von-Neumann bottleneck, making it a key technology for intelligent computing in edge IoT devices, particularly for real-time applications like keyword spotting (KWS). However, traditional CIM architectures face challenges such as high resource consumption, especially in data conversion, which can significantly impact chip area and energy efficiency. To address these challenges, this work proposes a computational CIM architecture utilizing multilevel analog memory, named AM-CIM, tailored for near-sensor (NS) computation of real-time KWS applications. Additionally, approximate memory technology is integrated into the AM-CIM architecture, employing data resilience scheduling for analog memory which contributes to significant reductions in hardware overhead. This integration facilitates a hardware-software co-design approach. To deploy KWS tasks in AM-CIM, a gated recurrent unit (GRU) network, referred to as MAC-GRU, is implemented. By employing Mel-energy as the input feature at the near-sensor end, the system achieves a 93.13% reduction in feature extraction power consumption. Evaluation results based on TSMC 180-nm technology demonstrate that the AM-CIM architecture achieves an accuracy of 88.51% for 10-keyword classification with a power consumption of$546~\mu W$, while reducing analog memory area by 43.32%. Xiaotao Jia, Guangcai Yuan, Jianyi Yu, Cong Shi 0003, Qi Wei 0001, Youguang Zhang, Weisheng Zhao 0001, Fei Qiao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2025 | A Self-Decryption Pass Transistor Logic-Based In-MRAM Computing Macro Using Hybrid VGSOT-MTJ/GAA-CNTFETabstractSpintronic devices and gate-all-around carbon nanotube field-effect-transistors (GAA-CNTFETs)-based computing in-memory architecture are competitive candidates for applications in battery-powered tiny artificial intelligence (AI) edge devices. Meanwhile, data encryption and decryption are also necessary to protect AI model weights and the customized data used to guarantee neural network (NN) inference accuracy. In this study, we propose a self-decryption pass transistor logic (PTL)-based in-MRAM computing macro (SP-CIM) that utilizes hybrid voltage-gated spin-orbit torque magnetic tunnel junctions (VGSOT-MTJ)/GAA-CNTFET. The proposed SP-CIM macro enables simultaneous data access, decryption, and full-accuracy multiply-and-accumulate (MAC) operations using the newly introduced voltage-divider self-decryption cell, without the need for additional decryption logic. Compared to existing in-memory decryption strategies, this design reduces energy consumption by 45.7% and decreases decryption delay by 87.2%. To enhance area efficiency and reduce computing latency, we propose a PTL-based multiplication cell that achieves full-accuracy local 2b-IN TEXPRESERVE0 2b-W operations with only 20 transistors (20T). Additionally, novel PTL-based full-swing output half adders (10T-HA) and full adders (14T-FA) are proposed to construct the local adder tree, achieving reductions of 31.8%, 76.4%, and 41.4% in energy, delay, and area, respectively, compared to conventional adder trees in CIM macros. Simulations of the 288 kb SP-CIM macro demonstrated throughput and energy efficiency of 2.25 TOPS and 226.6 TOPS/W, respectively, at a 0.6 V supply voltage, and 2.97 TOPS and 154.1 TOPS/W, respectively, at a 0.8 V supply voltage, with 8b-IN, 8b-W, and 24b-OUT. Zhongzhen Tong, Sifan Sun, Chenghang Li, Jiye Yao, Yulong Qiu, Chao Wang 0094, Zhaohao Wang, Amara Amara, Xiaoyang Lin, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 12 |
| 2025 | A Heterogeneous System With Computing in Memory Processing Elements to Accelerate CNN InferenceabstractComputing in memory (CIM) is one of the promising solutions to improve computing performance by integrating logic in memory. This work presents an efficient heterogeneous system based on the ultrafast CIM architecture (HS-CIM) to accelerate convolutional neural network (CNN) inference. First, an ultrafast CIM architecture is proposed based on the static random access memory (SRAM) by utilizing the novel total input and full digital scheme to implement multiply-and-accumulate (MAC) operation, which effectively addresses the high delay issue caused by high-precision computing in CIM architecture. Second, a heterogeneous system based on the proposed CIM architecture (HS-CIM) has been constructed with an aligned global cache and an adaptive pruning scheme to eliminate performance degradation and accuracy loss caused by input data bandwidth limitations. Meanwhile, efficient input data and weight data mapping schemes are proposed to minimize the delay and energy caused by input data transmission from the cache to the CIM architecture, thus realizing efficient CNN inference in the HS-CIM system. Finally, we analyze the performance of HS-CIM at the layout level by implementing the LeNet-5 and VGG models of CNN to recognize the image of MNIST and CIFAR-10 datasets, respectively. Results show that the energy efficiency of the proposed CIM architecture achieves 58.32 TOPS/W with 8-bit input/weight precision. Meanwhile, the inference accuracy of the HS-CIM system for MNIST and CIFAR-10 is 99.25% and 92.34%, respectively. Youxiang Chen, Zhengkun Gu, Haiming Qiu, Kun Zhang 0030, Weisheng Zhao 0001, Yue Zhang 0010 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2025 | A Time and Energy-Efficient Asynchronous Hybrid-Searching Auto Frequency Calibration for a 3.2 GHz Phase-Locked LoopabstractWide-band wireless system-on-chip demands phase-locked loops(PLL) designed with multi-band voltage controlled oscillators (VCO), which requires auto frequency calibration (AFC) for frequency presetting. This paper proposes a high-speed energy-efficient integrated AFC with asynchronous hybrid searching technique. The asynchronous architecture breaks the minimum limit of search time, while the hybrid method overcomes nonmonotonic variation of frequency errors in binary search. A true single-phase clock (TSPC)-based RF digital counter further accelerates AFC by directly quantizing the VCO output frequency. In this paper, a 3.2GHz PLL is presented utilizing the high-speed AFC to achieve fast locking performance. Implemented in 28nm CMOS technology, the proposed AFC for a 7-bit VCO achieves a calibration time of 0.88-$4.74\boldsymbol {\mu }$s across available tuning range, while the time of each step reaches 120ns level. With the aid of AFC, the prototype PLL reaches settle in less than$11.8\boldsymbol {\mu }$s, while achieving 318.2fs integrated jitter and -64.3dBc reference spur. The Figure-of-Merit (FoM) of the 3.2GHz PLL achieves -239.66dB for$\text {FoM}_{\text {jitter}}$, and -186.16dB for$\text {FoM}_{\text {T}_{\text {S}}}$. Fanxun Cai, Lianbo Wu, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | MS-SCIM: A Mixed-Signal Stochastic Computing-in-Memory Paradigm for Information SecurityabstractStochastic Computing (SC), an emerging paradigm with advantages in hardware cost and fault tolerance, is well-suited for applications in image processing and information security. However, the existing SC paradigms suffer from significant hardware costs due to the conversion between binary and stochastic sequences, and the long latency caused by low computational parallelism. In this work, we demonstrate a Mixed-Signal Stochastic Computing-In-Memory (MS-SCIM) paradigm utilizing spin orbit torque magnetic random access memory (SOT-MRAM) arrays, in order to realize a energy-efficient and conversion-less SC method for the first time. The main contributions include: 1) The inherent stochastic switching behaviors of spintronic devices are exploited to enable the SOT-MRAM array to serve both as a parallel true random number generator (TRNG) and a CIM cell. 2) A high-parallelism mixed signal stochastic CIM paradigm is proposed to accelerate SC-based edge detection algorithm. The whole process achieves binary outputs without the conversion circuits and the energy efficiency achieves 446 Tops/W. 3) Based on the results of MS-SCIM, a novel image steganography method using stochastic bit streams is leveraged for information security, which enables lossless embedding and extraction of secret image information in a$156\times 156$size, enhancing both capacity and undetectability. Pengxu Wang, Yijiao Wang, Jialiang Yin, Jiayao Wu, Xinrui Duan, Zhaohao Wang, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2025 | Technically Feasible Robust Complementary SOT-MRAM Design for Improving the Area and Energy EfficiencyabstractSpin-orbit torque magnetic random-access memory (SOT-MRAM), which exhibits sub-nanosecond write speed and high endurance, is a promising candidate for the future high-level cache. Nevertheless, SOT-MRAM faces challenge in meeting the high read performance requirements of cache applications due to the limited ON/OFF ratio. Consequently, extensive investigation has been conducted into robust complementary bit-cell (CBC) designs based on SOT-MRAM. However, previous designs suffer from significant technology feasibility, area and performance issues. In this paper, the feasibility and performance of the existing complementary write schemes are analyzed, and optimized U-type and toggle spin torque (TST) schemes with practicality and conciseness are presented. The previous CBC designs are evaluated and optimized in terms of circuit and layout, while the 1-word-line-3-bit-line (1WL3BL) CBC designs with both U-type and TST schemes are proposed, which can reduce the bit-cell area by 24.64%-27.54% and improve the write and read performance. In comparison to the conventional CBC design, the proposed 1WL3BL CBC design can reduce the write energy and read latency by up to 36.91% and 21.93%, respectively. Furthermore, the proposed low-voltage read scheme demonstrates the capability to enhance the read performance and conserve the read energy under the aggressive read-related process parameters. Chao Wang 0094, Zhongkui Zhang, Xianzeng Guo, Qihang Gao, Zhaohao Wang, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2025 | A 0.88 e‾rms 8-Mpixel 3D-Stacked Low Temporal-Noise CMOS Image Sensor With Auto-Zero Single-Slope ADC, Fast Correlated Multi-Sampling, Row-Wise Noise Reduction, and Dark Current Non-Uniformity Calibration TechniquesabstractThis paper presents a low temporal noise, low-power, 8-Mpixel, rolling-shutter (RS)-type, back-illuminated CMOS image sensor (CIS) employing through silicon via (TSV) 3D-stack technology. To achieve temporal noise less than 1erms-, we explored auto-zero (AZ) column single-slope (SS) ADC and fast correlated multi-sampling (CMS) techniques. The pixel signal was sampled two times by the readout circuits using a 9-bits ADC, resulting in a 10-bits digital output. To enhance image quality in low light conditions, we adopted a parity column counter (PCC) for power supply stabilization and H-banding elimination, and employed row-wise noise reduction (RWNR) and dark-current non-uniformity calibration (DCNUC) techniques for reducing row-wise noise and improving image uniformity. Our CIS chip was fabricated using a 55nm 1P4M (pixel substrate) and a 55nm 1P5M (logic substrate) CIS 3D stacked process. The die area is ~3.99*3.45 mm2with 1.008-μm pixel pitch and the total energy consumption is 170mW under a 2.8V analog-VDD and a 1.2V digital-VDD. The chip achieves a temporal noise of only ~0.88erms-, fixed pattern noise (FPN) of ~25.08μVrms, row-wise noise of ~5.5μVrmsand an energy efficiency figure-of-merit (FoM) of ~0.6erms-*nJ/step at a frame rate of 60 frames per second (FPS). Wang Kang 0001, Jing Kou, Liangchen Li, He Zhang 0011, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2025 | A High-Speed, Low-Power, High-Reliability and Fully Single Event Double Node Upset Tolerant Design for Magnetic Random Access MemoryabstractMagnetic Random Access Memory (MRAM) has enormous application potential in the aerospace field due to its nonvolatile, high speed, low power, and inherent radiation resistance characteristics. Due to its high sensing reliability, pre-charge differential sense amplifier (PCDSA) has been proposed and widely used in MRAM products. However, such PCDSA is based on traditional CMOS technology, and as the size of CMOS technology continues to shrink, its sensing result is easily affected by single event upset (SEU) or even the single event double node upset (SEDU). Recently, a TSC-PCDSA has been proposed to fully tolerate SEDU. However, it still suffers from slow speed, high power consumption and low reliability during normal sense operation. To address these issues, this paper proposes a novel PCDSA circuit that uses 6 three-input approximate C-elements (TACs) and 2 three-input standard C-elements (TSCs) to provide SEDU-tolerance. By reducing the number of transistors on the discharge path and increasing the difference in discharge current, the proposed PCDSA can achieve high speed, low power and high reliability. By using a physics-based STT-MTJ compact model and a commercial CMOS 40 nm design kit, hybrid simulations have been performed to demonstrate its functionality and evaluate its performance. Simulation results show that when the TMR is 150%, the width of N1-N12 is 480 nm and the$\text {V}_{\text {DD}}$is 1.1 V, the proposed PCDSA sensing error rate (SER) is close to 0% during normal sense operation, achieving a high sense speed of 123.6 ps and a low sense energy of 1.6533 fJ. Compared with the previously proposed TSC-PCDSA, the sense reliability is greatly improved, and the sense time and sense energy are reduced by 1.84 times and 1.27 times, respectively. Moreover, the proposed PCDSA can fully tolerate SEDU by optimizing the layout design. In the worst case where deposited charge$Q_{\text {inj}}$is 2 pC, it can achieve a shorter recover time of 1.28244 ns and a lower recover energy dissipation of 2.1604 pJ than the previously proposed TSC-PCDSA. Shixuan Wang, Yue Zhang 0010, Weisheng Zhao 0001, Lang Zeng, Deming Zhang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2025 | A 32 kb 55 nm Radiation-Hardened SRAM Chip With SEU ≤1.1 E-11 Upsets/Bit-Day, SEL >107.1 MeV ⋅ cm²/mg, and TID >100 Krad(Si) for Space ApplicationsabstractIn this paper, a 32kb radiation-hardened (RH) static random access memory (SRAM) chip, named BH55RHSRAM32K, is proposed and fabricated for space applications. The chip is hardened from the view of the circuit level, layout level, and system level and is fabricated using a 55 nm CMOS process design kit with an RH cell library. At the circuit level, the proposed RH-14T SRAM cell and radiation-hardened pre-charged sense amplifier (RH-PCSA) cell adopt a polarity hardening method, making them fully tolerant of single event upset (SEU). At the layout level, the sensitive nodes in the proposed RH-14T SRAM cell and RH-PCSA cell layouts are isolated. Furthermore, the proposed RH-14T SRAM array adopts a bit-interleaved design, effectively reducing single event double upsets (SEDU). At the system level, an error correction coding (ECC) circuit is implemented to enhance SEU tolerance. Experimental results show that the proposed 32kb RH-SRAM chip can not only obtains superior radiation tolerance, i.e., the SEU ≤ 1.1E-11 upsets/bit-day, the SEL > 107.1 MeV⋅cm2/mg, and the TID > 100 Krad(Si), but also a faster access speed of < 10 ns and a lower write power consumption of 14.664 mW in comparison with the related products. Deming Zhang, Dingyi Luo, Lang Zeng, Bi Wang 0002, Yue Zhang 0010, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2024 | Towards Efficient SRAM-PIM Architecture Design by Exploiting Unstructured Bit-Level SparsityabstractBit-level sparsity in neural network models harbors immense untapped potential. Eliminating redundant calculations of randomly distributed zero-bits significantly boosts computational efficiency. Yet, traditional digital SRAM-PIM architecture, limited by rigid crossbar architecture, struggles to effectively exploit this unstructured sparsity. To address this challenge, we propose Dyadic Block PIM (DB-PIM), a groundbreaking algorithm-architecture co-design framework. First, we propose an algorithm coupled with a distinctive sparsity pattern, termed a dyadic block (DB), that preserves the random distribution of non-zero bits to maintain accuracy while restricting the number of these bits in each weight to improve regularity. Architecturally, we develop a custom PIM macro that includes dyadic block multiplication units (DBMUs) and Canonical Signed Digit (CSD)-based adder trees, specifically tailored for Multiply-Accumulate (MAC) operations. An input pre-processing unit (IPU) further refines performance and efficiency by capitalizing on block-wise input sparsity. Results show that our proposed co-design framework achieves a remarkable speedup of up to 7.69× and energy savings of 83.43%. Cenlin Duan, Jianlei Yang 0001, Yiou Wang, Yingjie Qi, Xiaolin He, Bonan Yan, Xiaotao Jia, Weisheng Zhao 0001 |
DAC | 10 |
| 2024 | Series-Parallel Hybrid SOT-MRAM Computing-in-Memory Macro with Multi-Method Modulation for High Area and Energy EfficiencyabstractComputing-in-memory (CIM) shows its superiority in lots of applications like neural network inference. Recently, there are lots of exploration of the application of Magnetic Random-Access Memory (MRAM) in CIM. This paper aims to investigate the potential of Spin-Orbit-Torque-MRAM (SOT-MRAM) in CIM and proposes a high area and energy efficiency SOT-MRAM CIM macro based on a 6T-4J weight group. The bit-cell array adopts series-parallel hybrid architecture, which combines both serial and parallel configurations of Magnetic Tunnel Junction (MTJ) to solve the problem of high energy cost and low flexibility caused by MRAM-series and MRAM-parallel architecture, respectively. Additionally, the proposed SOT-MRAM CIM macro incorporates a multi-method modulation scheme, ranging from input unit to array, which meanwhile allows for configurable input precision (2/4/6/8-bit). The SOT-MRAM CIM macro is designed and verified in both 180nm and 28nm nodes, based on the verified electrical performance of the SOT-MRAM array in a 200-nm wafer pre-fabricated. The simulation results in 28nm show that this macro can achieve energy efficiency of 23.7~29.6 Tops/W at 8-bit input and output precision. Weiliang Huang, Jinyu Bai, Wang Kang 0001, Zhaohao Wang, Kaihua Cao, He Zhang 0011, Weisheng Zhao 0001 |
DAC | 8 |
| 2024 | A Combined Content Addressable Memory and In-Memory Processing Approach for k-Clique Counting Accelerationabstractk-Clique counting problem plays an important role in graph mining which has seen a growing number of applications. However, current k-Clique counting accelerators cannot meet the performance requirement mainly because they struggle with high data transfer issue incurred by the intensive set intersection operations and the inability of load balancing. In this paper, we propose to solve this problem with a hybrid framework of content addressable memory (CAM) and in-memory processing (PIM). Specifically, we first utilize CAM for binary induced subgraph generation in order to reduce the search space, then we use PIM to implement in-place parallel k-Clique counting through iterative Boolean logic "AND" like operation. To take full advantage of this combined CAM and PIM framework, we develop dynamic task scheduling strategies that can achieve near optimal load balancing among the PIM arrays. Experimental results demonstrate that, compared with state-of-the-art CPU and GPU platforms, our approach achieves speedups of 167.5× and 28.8×, respectively. Meanwhile, the energy efficiency is improved by 788.3× over the GPU baseline. Xidi Ma, Tianyang Yu, Bi Wu 0002, Gang Qu 0001, Weisheng Zhao 0001 |
DAC | 7 |
| 2024 | GNNavigator: Towards Adaptive Training of Graph Neural Networks via Automatic Guideline ExplorationabstractGraph Neural Networks (GNNs) succeed significantly in many applications recently. However, balancing GNNs training runtime cost, memory consumption, and attainable accuracy for various applications is non-trivial. Previous training methodologies suffer from inferior adaptability and lack a unified training optimization solution. To address the problem, this work proposes GNNavigator, an adaptive GNN training configuration optimization framework. GN-Navigator meets diverse GNN application requirements due to our unified software-hardware co-abstraction, proposed GNNs training performance model, and practical design space exploration solution. Experimental results show that GNNavigator can achieve up to 3.1× speedup and 44.9% peak memory reduction with comparable accuracy to state-of-the-art approaches. Jianlei Yang 0001, Yingjie Qi, Bei Yu 0001, Weisheng Zhao 0001, Chunming Hu |
DAC | 7 |
| 2024 | FRM-CIM: Full-Digital Recursive MAC Computing in Memory System Based on MRAM for Neural Network ApplicationsabstractComputing in memory (CIM) realizes energy-efficient neural network algorithms by implementing highly parallel multiply-and-accumulate (MAC) operation. However, the MAC delay of CIM will sharply increase with the improvement of computing precision, which restricts its development. In this work, we propose a full-digital recursive MAC (FRM) operation based on spin-transfer-torque magnetic random access memory (STT-MRAM) CIM system to enable fast and energy-efficient image recognition application. First, the fast FRM scheme is proposed by utilizing the recursive operations of read and addition in segmented bit-line array, which effectively reduces the delay of MAC operations to 3.5ns and 4ns for 8-bit and 16-bit input and weight precision, respectively. Second, we design an image recognition system using FRM-CIM architecture as the processing element (PE), where the adaptive pruning method for layers is proposed to improve the compatibility of it with the neural network. By performing image recognition for the MNIST and CIFAR-10 datasets, results show that the throughput and energy efficiency of the FRM-CIM system are 58.51TOPS/mm2 and 11.3--56.72 TOPS/W under 8--16-bit precision, which are improved by 4.3 times and 2.6 times compared with the state-of-the-art works. Finally, the recognition accuracy can reach 96.65% and 82.7% on MNIST and CIFAR-10, respectively. Zhengkun Gu, Youxiang Chen, Weisheng Zhao 0001, Yue Zhang 0010 |
DAC | 6 |
| 2024 | PPGNN: Fast and Accurate Privacy-Preserving Graph Neural Network Inference via Parallel and Pipelined Arithmetic-and-Logic FHE AcceleratorabstractGraph Neural Networks (GNNs) are increasingly used in fields like social media and bioinformatics, promoting the prosperity of cloud-based GNN inference services. Nevertheless, data privacy becomes a critical issue when handling sensitive information. Fully Homomorphic Encryption (FHE) enables computations on encrypted data, while privacy-preserving GNN inference generally necessitates ensuring graph structure data confidentiality and maintaining computation precision, both of which are computationally expensive in FHE. Existing schemes of GNNs inference with FHE are deterred by either computational overhead, accuracy degradation, or incomplete data protection. This paper presents PPGNN to address these challenges all at once. We first propose a novel privacy-preserving GNN inference algorithm utilizing a high-accuracy arithmetic-and-logic FHE approach, meanwhile only need much smaller parameters, substantially reducing computational complexity and facilitating parallel processing. Correspondingly, a dedicated hardware architecture has been designed to implement these innovations, with featured specialized units for arithmetic and logic FHE operations in a pipelined manner. Collectively, PPGNN achieves 2.7× and 1.5× speedup over state-of-the-art Arithmetic FHE and Logic FHE accelerators while ensuring high accuracy, simultaneously with about 18× energy reduction on average. Yuntao Wei, Song Bian 0001, Weisheng Zhao 0001, Yier Jin |
DAC | 5 |
| 2024 | Graph Neural Networks Automated Design and Deployment on Device-Edge Co-Inference SystemsabstractThe key to device-edge co-inference paradigm is to partition models into computation-friendly and computation-intensive parts across the device and the edge, respectively. However, for Graph Neural Networks (GNNs), we find that simply partitioning without altering their structures can hardly achieve the full potential of the co-inference paradigm due to various computational-communication overheads of GNN operations over heterogeneous devices. We present GCoDE, the first automatic framework for GNN that innovatively Co-designs the architecture search and the mapping of each operation on Device-Edge hierarchies. GCoDE abstracts the device communication process into an explicit operation and fuses the search of architecture and the operations mapping in a unified space for joint-optimization. Also, the performance-awareness approach, utilized in the constraint-based search process of GCoDE, enables effective evaluation of architecture efficiency in diverse heterogeneous systems. We implement the co-inference engine and runtime dispatcher in GCoDE to enhance the deployment efficiency. Experimental results show that GCoDE can achieve up to 44.9× speedup and 98.2% energy reduction compared to existing approaches across various applications and system configurations. Jianlei Yang 0001, Yingjie Qi, Zhi Yang 0001, Weisheng Zhao 0001, Chunming Hu |
DAC | 6 |
| 2024 | LLP-ECCA: A Low-Latency and Programmable Framework for Elliptic Curve Cryptography AcceleratorsabstractElliptic curve cryptography (ECC) plays a pivotal role in safeguarding data integrity and authentication in contemporary communication contexts, particularly within the domain of Intelligent Transport Systems (ITS). In the realm of ITS, vehicles communicate via the V2X (vehicle-to-everything) protocol, necessitating low-latency responses and minimal power consumption. Given the evolving nature of V2X protocol standards across the globe, programmability becomes a rigid requirement. However, existing strategies cannot meet all these vehicular equipment demands. This paper introduces a novel framework tailored for ECC acceleration to address the issues. Specifically, we propose the design of an Application Specific Instruction Set Processor (ASIP), augmented by pipeline and dual-issue techniques. Furthermore, the envisioned ASIP integrates a hybrid control framework founded on Finite State Machines (FSM), facilitating agile and effective management. Notably, a general GF(p256) Barrett modular multiplier is specially devised to optimize latency and area utilization. Experimental results on Xilinx Kintex Ultrscale+ FPGA demonstrate that the proposed ECC accelerator generates a signature within 131us and verifies a message within 181us, and the performance meets the requirements of today’s V2X standard. Tianao Dai, Jianlei Yang 0001, Zhaojun Lu, Xiaotao Jia, Gang Qu 0001, Weisheng Zhao 0001 |
ITC-Asia | 8 |
| 2024 | CRISP: Triangle Counting Acceleration via Content Addressable Memory-Integrated 3D-Stacked MemoryabstractTriangle Counting is a fundamental problem in graph analysis, which usually needs to traverse the graph and perform set-intersections of neighbor sets. However, existing approaches suffer from heavy off-chip memory access and set-intersection overhead, which are both memory-bound and computation-bound. Fortunately, the emerging 3D-stacked computation-in-memory (CIM) architecture can reduce off-chip memory access, and the content addressable memory (CAM) can achieve parallel comparison. However, existing solutions have not effectively combined the high bandwidth of 3D-stacked memory with the high computational capabilities of CAM arrays. Besides, there exist many fruitless searches in the triangle counting process. Thus, we propose CRISP, a software-hardware co-design architecture to address these issues. At the level of software design, a new storage format named Two-Pointer CSR is proposed to eliminate fruitless searches during the set-intersection process. At the level of hardware design, CRISP integrates a novel Presence-Bits based Content Addressable Memory (PB-CAM) near the memory bank of 3D-stacked memory to fully exploit the high internal bandwidth. Through the presence bits comparison, the PB-CAM can effectively reduce both the off-chip memory access and set-intersection operations. Experimental results show that compared with previous state-of-the-art near-DIMM and HBM-PIM triangle counting accelerators, CRISP achieves speedups of 5.7× and 1.8 respectively. Shangtong Zhang, Weisheng Zhao 0001, Yier Jin |
ITC-Asia | 3 |
| 2024 | Spin-orbit torque efficiency enhancement to tungsten-based SOT-MTJs by interface modification with an ultrathin MgO
Shiyang Lu, Xiaobai Ning, Sixi Zhen, Xiaofei Fan, Danrong Xiong, Dapeng Zhu, Gefei Wang, Kaihua Cao, Weisheng Zhao 0001 |
Sci. China Inf. Sci. | 11 |
| 2024 | HGNAS: Hardware-Aware Graph Neural Architecture Search for Edge DevicesabstractGraph Neural Networks (GNNs) are becoming increasingly popular for graph-based learning tasks such as point cloud processing due to their state-of-the-art (SOTA) performance. Nevertheless, the research community has primarily focused on improving model expressiveness, lacking consideration of how to design efficient GNN models for edge scenarios with real-time requirements and limited resources. Examining existing GNN models reveals varied execution across platforms and frequent Out-Of-Memory (OOM) problems, highlighting the need for hardware-aware GNN design. To address this challenge, this work proposes a novel hardware-aware graph neural architecture search framework tailored for resource constraint edge devices, namely HGNAS. To achieve hardware awareness, HGNAS integrates an efficient GNN hardware performance predictor that evaluates the latency and peak memory usage of GNNs in milliseconds. Meanwhile, we study GNN memory usage during inference and offer a peak memory estimation method, enhancing the robustness of architecture evaluations when combined with predictor outcomes. Furthermore, HGNAS constructs a fine-grained design space to enable the exploration of extreme performance architectures by decoupling the GNN paradigm. In addition, the multi-stage hierarchical search strategy is leveraged to facilitate the navigation of huge candidates, which can reduce the single search time to a few GPU hours. To the best of our knowledge, HGNAS is the first automated GNN design framework for edge devices, and also the first work to achieve hardware awareness of GNNs across different platforms. Extensive experiments across various applications and edge devices have proven the superiority of HGNAS. It can achieve up to a$10.6\boldsymbol{\times}$speedup and an$82.5\%$peak memory reduction with negligible accuracy loss compared to DGCNN on ModelNet40. Jianlei Yang 0001, Yingjie Qi, Yumeng Shi, Cenlin Duan, Weisheng Zhao 0001, Chunming Hu |
IEEE Trans. Computers | 7 |
| 2024 | CIMQ: A Hardware-Efficient Quantization Framework for Computing-In-Memory-Based Neural Network AcceleratorsabstractThe novel computing-in-memory (CIM) technology has demonstrated significant potential in enhancing the performance and efficiency of convolutional neural networks (CNNs). However, due to the low precision of memory devices and data interfaces, an additional quantization step is necessary. Conventional NN quantization methods fail to account for the hardware characteristics of CIM, resulting in inferior system performance and efficiency. This article proposes CIMQ, a hardware-efficient quantization framework designed to improve the efficiency of CIM-based NN accelerators. The holistic framework focuses on the fundamental computing elements in CIM hardware: inputs, weights, and outputs (or activations, weights, and partial sums in NNs) with four innovative techniques. First, bit-level sparsity induced activation quantization is introduced to decrease dynamic computation energy. Second, inspired by the unique computation paradigm of CIM, an innovative arraywise quantization granularity is proposed for weight quantization. Third, partial sums are quantized with a reparametrized clipping function to reduce the required resolution of analog-to-digital converters (ADCs). Finally, to improve the accuracy of quantized neural networks (QNNs), the post-training quantization (PTQ) is enhanced with a random quantization dropping strategy. The effectiveness of the proposed framework has been demonstrated through experimental results on various NNs and datasets (CIFAR10, CIFAR100, and ImageNet). In typical cases, the hardware efficiency can be improved up to 222% with a 58.97% improvement in accuracy compared to conventional quantization methods. Jinyu Bai, Sifan Sun, Weisheng Zhao 0001, Wang Kang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | DDC-PIM: Efficient Algorithm/Architecture Co-Design for Doubling Data Capacity of SRAM-Based Processing-in-MemoryabstractProcessing-in-memory (PIM), as a novel computing paradigm, provides significant performance benefits from the aspect of effective data movement reduction. SRAM-based PIM has been demonstrated as one of the most promising candidates due to its endurance and compatibility. However, the integration density of SRAM-based PIM is much lower than other nonvolatile memory-based ones, due to its inherent 6T structure for storing a single bit. Within comparable area constraints, SRAM-based PIM exhibits notably lower capacity. Thus, aiming to unleash its capacity potential, we propose DDC-PIM, an efficient algorithm/architecture co-design methodology that effectively doubles the equivalent data capacity. At the algorithmic level, we propose a filter-wise complementary correlation (FCC) algorithm to obtain a bitwise complementary pair. At the architecture level, we exploit the intrinsic cross-coupled structure of 6T SRAM to store the bitwise complementary pair in their complementary states$(Q/\overline {Q})$, thereby maximizing the data capacity of each SRAM cell. The dual-broadcast input structure and reconfigurable unit support both depthwise and pointwise convolution, adhering to the requirements of various neural networks. Evaluation results show that DDC-PIM yields about$2.84\times $speedup on MobileNetV2 and$2.69\times $on EfficientNet-B0 with negligible accuracy loss compared with PIM baseline implementation. Compared with state-of-the-art SRAM-based PIM macros, DDC-PIM achieves up to$8.41\times $and$2.75\times $improvement in weight density and area efficiency, respectively. Cenlin Duan, Jianlei Yang 0001, Xiaolin He, Yingjie Qi, Yiou Wang, Ziyan He, Bonan Yan, Xiaotao Jia, Weitao Pan, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 12 |
| 2024 | APIM: An Antiferromagnetic MRAM-Based Processing-In-Memory System for Efficient Bit-Level Operations of Quantized Convolutional Neural NetworksabstractQuantized Convolutional Neural Network (QCNN) is an attractive approach that reduces hardware overheads, especially for energy-constrained systems. However, existing QCNNs still require non-trivial hardware resources and memory capacity in order not to compromise model accuracy. To address this issue, we propose an antiferromagnetic magnetic random-access memory (ARAM)-based processing-in-memory (PIM) system, leveraging bit-level sparsity. Three optimization techniques are proposed to optimize hardware resource utilization while preserving CNN accuracy. Firstly, the ARAM-based memory subsystem allows dynamic adaptation of variable bit-width across CNN layers. Secondly, the bit-level accelerator employs the bit-fusion format engineered for processing data from the ARAM subsystem. Thirdly, a customized data path within the RISC-V core guarantees efficient instruction processing to the ARAM-based memory subsystem and bit-level accelerator, enabling optimal bit-level data transmission and computation. Experimental results demonstrate that this design remarkably reduces data movement by 50%-83% across existing CNNs. Compared to state-of-the-art designs, it enhances throughput and latency by an average of 5x and 10x, respectively. In addition, this design achieves speedups between 1.63x and 2.96x, outstripping other designs in AlexNet, VGG16, and ResNet18 benchmarks. Yueting Li 0001, Daoqian Zhu, Jinhao Li 0007, Ao Du, Yue Zhang 0010, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2024 | CIM²PQ: An Arraywise and Hardware-Friendly Mixed Precision Quantization Method for Analog Computing-In-MemoryabstractComputing-in-memory (CIM) architecture is a promising convolutional neural network (CNN) accelerator known for its highly efficient matrix-vector multiplications (MVMs). However, due to the low-precision computation and limited size of CIM memory arrays, it is necessary to decompose the huge MVMs into smaller subsets. Conventional NN quantization methods overlook the characteristics of CIM hardware, resulting in diminished system performance and efficiency. This paper proposes a mixed precision quantization (MPQ) method based on evolutionary algorithm for CIM-based accelerators, while considering the hardware characteristics of CIM, called CIMPQ, which can automatically generate quantization strategies for NN model to improve the efficiency of CIM systems. Firstly, inspired by the CIM computing paradigm, an array-wise quantization granularity is introduced in the MPQ search space, which can jointly quantize the inputs, weights, and partial sums. Secondly, a production procedure containing fine-grained crossover and progressive adaptive mutation is proposed, which can efficiently explore the search space and speed up the search process. Thirdly, we propose a fast and efficient strategy evaluation method to obtain the performance of quantization strategy on the CIM platform, saving the evaluation time significantly without requiring fine-tuning. Finally, to protect CIM-friendly strategies with lower bit-widths but worse algorithm performance, we propose a strategy selection method based on multi-objective optimization, named qNSGA-III. The effectiveness of the proposed method has been demonstrated through experimental results of various NNs and datasets. For ResNet-18, the hardware efficiency and accuracy can be improved to 117% with 7.05%, 113% with 3.37%, and 119% with 5.78%, on CIFAR-10, CIFAR-100 and ImageNet, respectively, compared to the baseline MPQ method. Sifan Sun, Jinyu Bai, Zhaoyu Shi, Weisheng Zhao 0001, Wang Kang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | Multicorner Timing Analysis Acceleration for Iterative Physical Design of ICsabstractWe propose a multi-corner multi-stage timing analysis prediction framework using a generalized linear model with latent features. We then further improve such methods using kernel trick extension, transfer learning with knowledge from previous designs, and multi-output feature engineering to deliver state-of-the-art (SOTA) prediction accuracy with very limited training data. Most importantly, our method is equipped with a Bayesian decision strategy to deliver reliable predictions with accuracy close to 100%, pushing the frontier of the machine-learning-based STA for practical implementation in the industry environment, where reliability is highly desired. Experimental results show that the accuracy of our proposed method outperforms the SOTA competitors by up to 4x and can improve prediction accuracy to 100% with little extra STA executions. Wei W. Xing, Longze Wang, Zhelong Wang, Zhaoyu Shi, Ning Xu 0006, Yuanqing Cheng, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2024 | BSTCIM: A Balanced Symmetry Ternary Fully Digital In-MRAM Computing Macro for Energy Efficiency Neural NetworkabstractSilicon-based traditional binary computing in-memory (TBCIM) architectures are approaching their energy efficiency and throughput limits owing to challenges facing Moore’s Law. Thus, it is essential to explore architecture based on novel devices and computing paradigms to fulfill data-centric applications, such as artificial intelligence. In this paper, we propose a balanced symmetry ternary (BST) fully digital in-MRAM computing macro (BSTCIM) using hybrid voltage-gated spin-orbit torque magnetic tunnel junctions (VGSOT-MTJ) and gate-all-around carbon nanotube field-effect-transistors (GAA-CNTFET) technology. The overall computing is based on the highest efficiency multi-bit ternary system. BSTCIM includes a ternary dot product (TDP) unit with 4 GAA-CNTFETs and 2 VGSOT-MTJs achieving TDP operation without complex logic circuits. The multi-bit ternary multiply-and-accumulate (MAC) operation is realized through the proposed ternary adder tree and ternary post adder which accumulate TDP results within the digital domain enabling high accuracy neural network inference. Furthermore, due to the advantages of BST, ternary signed MAC is more easily performed compared to TBCIM macros that adapt 2’s complement or separate signed bit calculations. BSTCIM with 288 kb is simulated, achieving throughput and energy efficiency of 0.72 TOPS and 54.5 TOPS/W, respectively, at a 0.6 V supply voltage and 1.15 TOPS and 33.7 TOPS/W, respectively at a 0.8 V supply voltage with 8b-IN, 8b-W, and 20b-OUT. Moreover, the figure-of-merit for BSTCIM is 1.13–33.6 times higher than that of existing CIM macros. Zhongzhen Tong, Chenghang Li, Chao Wang 0094, Suteng Zhao, Qianyong Peng, Daming Zhou, Zhaohao Wang, Xiaoyang Lin, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 11 |
| 2024 | Toward Energy-efficient STT-MRAM-based Near Memory Computing Architecture for Embedded SystemsabstractConvolutional Neural Networks (CNNs) have significantly impacted embedded system applications across various domains. However, this exacerbates the real-time processing and hardware resource-constrained challenges of embedded systems. To tackle these issues, we propose spin-transfer torque magnetic random-access memory (STT-MRAM)-based near memory computing (NMC) design for embedded systems. We optimize this design from three aspects: Fast-pipelined STT-MRAM readout scheme provides higher memory bandwidth for NMC design, enhancing real-time processing capability with a non-trivial area overhead. Direct index compression format in conjunction with digital sparse matrix-vector multiplication (SpMV) accelerator supports various matrices of practical applications that alleviate computing resource requirements. Custom NMC instructions and stream converter for NMC systems dynamically adjust available hardware resources for better utilization. Experimental results demonstrate that the memory bandwidth of STT-MRAM achieves 26.7 GB/s. Energy consumption and latency improvement of digital SpMV accelerator are up to 64× and 1,120× across sparsity matrices spanning from 10% to 99.8%. Single-precision and double-precision elements transmission increased up to 8× and 9.6×, respectively. Furthermore, our design achieves a throughput of up to 15.9× over state-of-the-art designs. Yueting Li 0001, He Zhang 0011, Biao Pan, Keni Qiu, Wang Kang 0001, Jun Wang 0041, Weisheng Zhao 0001 |
ACM Trans. Embed. Comput. Syst. | 8 |
| 2024 | An Energy-Efficient Bayesian Neural Network Implementation Using Stochastic Computing MethodabstractThe robustness of Bayesian neural networks (BNNs) to real-world uncertainties and incompleteness has led to their application in some safety-critical fields. However, evaluating uncertainty during BNN inference requires repeated sampling and feed-forward computing, making them challenging to deploy in low-power or embedded devices. This article proposes the use of stochastic computing (SC) to optimize the hardware performance of BNN inference in terms of energy consumption and hardware utilization. The proposed approach adopts bitstream to represent Gaussian random number and applies it in the inference phase. This allows for the omission of complex transformation computations in the central limit theorem-based Gaussian random number generating (CLT-based GRNG) method and the simplification of multipliers as AND operations. Furthermore, an asynchronous parallel pipeline calculation technique is proposed in computing block to enhance operation speed. Compared with conventional binary radix-based BNN, SC-based BNN (StocBNN) realized by FPGA with 128-bit bitstream consumes much less energy consumption and hardware resources with less than 0.1% accuracy decrease when dealing with MNIST/Fashion-MNIST datasets. Xiaotao Jia, Huiyi Gu, Jianlei Yang 0001, Weitao Pan, Youguang Zhang, Sorin Cotofana, Weisheng Zhao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 9 |
| 2023 | Toward Energy-Efficient Sparse Matrix-Vector Multiplication with near STT-MRAM Computing ArchitectureabstractSparse Matrix-Vector Multiplication (SpMV) is one of the vital computational primitives used in modern workloads. SpMV performs memory access, leading to unnecessary data transmission, massive data access, and redundant multiplicative accumulators. Therefore, we propose the near spin-transfer torque magnetic random access memory (STT-MRAM) processing architecture from three optimization perspectives. These optimizations include (1) the NMP controller receives the instruction through the AXI4 bus to implement the SpMV operation in the following steps, identifies valid data, and encodes the index depending on the kernel size, (2) the NMP controller uses high-level synthesis dataflow in the shared buffer for achieving better performance throughput while do not consume bus bandwidth, and (3) the configurable MACs are implemented in the NMP core without matching step entirely during the multiplication. Using these optimizations, the NMP architecture can access the pipelined STT-MRAM (read bandwidth is 26.7GB/s). The experimental simulation results show that this design achieves up to 66x and 28x speedup compared with state-of-the-art ones and 69x speedup without sparse optimization. Yueting Li 0001, He Zhang 0011, Hao Cai 0001, Shuqin Lv, Renguang Liu, Weisheng Zhao 0001 |
ASP-DAC | 8 |
| 2023 | TOTAL: Multi-Corners Timing Optimization Based on Transfer and Active LearningabstractIn modern advanced integrated circuit design, a design normally needs to be progressively optimized until the static timing analysis (STA) of full process corners meets the timing constraints. To improve efficiency, using machine learning to predict the path timings directly in order to reduce the extensive time-consuming SPICE simulations has become a promising technique to approach fast design closure. However, current methods lack both flexibility and reliability to be used in a practical industrial environment. To resolve these challenges, we propose TOTAL, which is constructed using a generalized linear model with latent features to effectively capture knowledge transferred from previous designs and delivers state-of-the-art (SOTA) prediction accuracy that is up to 6.6x improvement over the competitors in terms of mean absolute error (MAE). Most importantly, TOTAL is equipped with a Bayesian decision strategy to actively update uncertain predictions and deliver reliable predictions with accuracy close to 100%, pushing the frontier of the machine-learning-based STA for practical implementation. Wei W. Xing, Rongqi Lu, Zhelong Wang, Ning Xu 0006, Yuanqing Cheng, Weisheng Zhao 0001 |
DAC | 7 |
| 2023 | Hardware-Aware Graph Neural Network Automated Design for Edge Computing PlatformsabstractGraph neural networks (GNNs) have emerged as a popular strategy for handling non-Euclidean data due to their state-of-the-art performance. However, most of the current GNN model designs mainly focus on task accuracy, lacking in considering hardware resources limitation and real-time requirements of edge application scenarios. Comprehensive profiling of typical GNN models indicates that their execution characteristics are significantly affected across different computing platforms, which demands hardware awareness for efficient GNN designs. In this work, HGNAS is proposed as the first Hardware-aware Graph Neural Architecture Search framework targeting resource constraint edge devices. By decoupling the GNN paradigm, HGNAS constructs a fine-grained design space and leverages an efficient multi-stage search strategy to explore optimal architectures within a few GPU hours. Moreover, HGNAS achieves hardware awareness during the GNN architecture design by leveraging a hardware performance predictor, which could balance the GNN model accuracy and efficiency corresponding to the characteristics of targeted devices. Experimental results show that HGNAS can achieve about 10.6× speedup and 88.2% peak memory reduction with a negligible accuracy loss compared to DGCNN on various edge devices, including Nvidia RTX3080, Jetson TX2, Intel i7-8700K and Raspberry Pi 3B+. Jianlei Yang 0001, Yingjie Qi, Yumeng Shi, Weisheng Zhao 0001, Chunming Hu |
DAC | 6 |
| 2023 | TAM: A Computing in Memory based on Tandem Array within STT-MRAM for Energy-Efficient Analog MAC OperationabstractComputing in memory (CIM) has been demonstrated promising for energy efficient computing. However, the dramatic growth of the data scale in neural network processors has aroused a demand for CIM architecture of higher bit density, for which the spin transfer torque magnetic RAM (STT-MRAM) with high bit density and performance arises as an up-and-coming candidate solution. In this work, we propose an analog CIM scheme based on tandem array within STT-MRAM (TAM) to further improve energy efficiency while achieving high bit density. First, the resistance summation based analog MAC operation minimizes the effect of low tunnel magnetoresistance (TMR) by the serial magnetic tunnel junctions (MTJs) structure in the proposed tandem array with smaller area overhead. Moreover, a read scheme of resistive-to-binary is designed to achieve the MAC results accurately and reliably. Besides, the data-dependent error caused by MTJs in series has been eliminated with a proposed dynamic selection circuit. Simulation results of a 2Kb TAM architecture show 113.2 TOPS/W and 63.7 TOPS/W for 4-bit and 8-bit input/weight precision, respectively, and reduction by 39.3% for bit-cell area compared with existing array of MTJs in series. Zhengkun Gu, Zuolei Hao, Weisheng Zhao 0001, Yue Zhang 0010 |
DATE | 6 |
| 2023 | THE-V: Verifiable Privacy-Preserving Neural Network via Trusted Homomorphic ExecutionabstractPrivacy-preserving machine learning (PPML) schemes aim at protecting client-side data privacy in two-party secure computing tasks such as private deep neural network (DNN) inference. While fully homomorphic encryption (FHE) can provide provable security for client data privacy, efficiently verifying that such homomorphic DNN inference protocol is honestly executed on the server presents to be challenging. In this work, we propose THE-V, a novel DNN inference framework that combines FHE and Trusted Execution Environment (TEE) to achieve data privacy, verifiable execution and efficient computation all at once. We first point out that, while the trivial solution of executing FHE entirely within TEE can ensure both private and verifiable computing, the limited resource within TEE becomes a severe computational bottleneck. To solve such dilemma, we devise a new strategy of securely outsourcing computation-heavy tasks in TEE to untrusted environments. By rigorous experiments, we show that we can achieve verifiable and private DNN inference with up to$15\times$speedup compared with the state-of-the-art solution. Yuntao Wei, Song Bian 0001, Weisheng Zhao 0001, Yier Jin |
ICCAD | 4 |
| 2023 | Implementation of 16 Boolean logic operations based on one basic cell of spin-transfer-torque magnetic random access memory
Kaihua Cao, Kun Zhang 0030, Kewen Shi, Zuolei Hao, Wenlong Cai, Ao Du, Jialiang Yin, Jianfeng Gao 0005, Weisheng Zhao 0001 |
Sci. China Inf. Sci. | 14 |
| 2023 | NAND-SPIN-based processing-in-MRAM architecture for convolutional neural network acceleration
Yinglin Zhao, Jianlei Yang 0001, Bing Li 0017, Xingzhou Cheng, Xucheng Ye, Xiaotao Jia, Zhaohao Wang, Youguang Zhang, Weisheng Zhao 0001 |
Sci. China Inf. Sci. | 10 |
| 2023 | Layout Aware Optimization Methodology for SOT-MRAM Based on Technically Feasible Top-Pinned Magnetic Tunnel Junction ProcessabstractThe emerging spin-orbit torque magnetic random-access memory (SOT-MRAM) shows promising prospects in high-level cache applications due to its subnanosecond switching speed and high reliability. However, SOT-MRAM faces the issue of large bit-cell layout area, which is currently the focus of attention. Although many design and evaluation works have emerged, the lack of a unified standard for realistic SOT process has hindered the development of relevant research toward practicality. In this article, the bit-cell area of the SOT-MRAM will be evaluated and optimized based on the technically feasible process. First of all, based on the state-of-the-art top-pinned SOT nanopillar process, the SOT-MRAM design rules are proposed. On this basis, this article systematically summarizes four basic device layout modes and provides optimized layout suggestions for conventional SOT bit-cells with different types and sizes of devices. In addition, a series of area-efficient SOT bit-cell designs based on the common area (CA) and dual common (DC) solutions are proposed, which can reduce the layout area of SOT bit-cells by up to 38.4% with reasonable write latency and energy overhead. Chao Wang 0094, Zhaohao Wang, Zhongkui Zhang, Jiagao Feng, Youguang Zhang, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | IMGA: Efficient In-Memory Graph Convolution Network Aggregation With Data Flow OptimizationsabstractAggregating features from neighbor vertices is a fundamental operation in graph convolution network (GCN). However, the sparsity in graph data creates poor spatial and temporal locality, causing dynamic and irregular memory access patterns and limiting the performance of aggregation on the Von Neumann architecture. The emerging processing-in-memory (PIM) architecture is based on emerging nonvolatile memory (NVM), like spin-orbit torque magnetic RAM (SOT-MRAM), and demonstrates promising prospects in alleviating the Von Neumann bottleneck. However, the limited memory capacity of PIM medium still incurs non-negligible data movements between PIM architecture and external memory. To solve this challenge, we propose an SOT-MRAM-based in-memory computing architecture, called IMGA, for efficient in-situ graph aggregation. Specifically, we design adaptive data flow management strategies that reuse vertex data in MRAM when processing graphs of different scales and adopt edge data as the control signal source to utilize the graph’s structural information. A reordering optimization strategy leveraging hardware–software co-design principle is proposed to further reduce the costly data movement. Experimental results demonstrate that IMGA achieves an average$2523\times $and$21\times $speedup, and 1.03E+6 and 1.04E+3 energy efficiency compared with CPU and GPU, respectively. Yuntao Wei, Shangtong Zhang, Jianlei Yang 0001, Xiaotao Jia, Zhaohao Wang, Gang Qu 0001, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2023 | Outgoing EditorialabstractDear TCAS-I Readers, Weisheng Zhao 0001, Hai Li 0001, Domenico Zito |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | Experimental Demonstration of STT-MRAM-based Nonvolatile Instantly On/Off System for IoT Applications: Case StudiesabstractEnergy consumption has been a big challenge for electronic devices, particularly for battery-powered Internet of Things (IoT) equipment. To address such a challenge, on the one hand, low-power electronic design methodologies and novel power management techniques have been proposed, such as nonvolatile memories and instantly on/off systems; on the other hand, the energy harvesting technology by collecting signals from human activity or the environment has attracted widespread attention in the IoT area. However, the system with self-powered energy harvesting may suffer frequent energy failures or fluctuating energy conditions, which degrade system reliability and user experience. Therefore, how to make the system under unreliable power inputs operate correctly and efficiently is one of the most critical issues for energy harvesting technology. In this article, we built an instantly on/off system based on nonvolatile STT-MRAM for IoT applications, which can instantly power on/off under different conditions of the harvested energy. The system powers on and operates normally when the harvested energy is enough (over the preset threshold); otherwise, the system powers off and stores the operational data back to the nonvolatile STT-MRAM. We described implementations of the hardware/software co-designed architecture (with image acquisition as an example) based on the commercialized 32 MB STT-MRAM, and we experimentally demonstrated the system functionality and efficiency under five typical energy harvesting scenarios, including radio frequency, thermal, solar, piezoelectric, and WIFI. Our experimental results show that the power consumption and data restore time were reduced by 15.1% and 714 times, respectively, in comparison with the DRAM-based counterpart. Yueting Li 0001, Wang Kang 0001, Kunyu Zhou, Keni Qiu, Weisheng Zhao 0001 |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2023 | BoA-PTA: A Bayesian Optimization Accelerated PTA Solver for SPICE SimulationabstractOne of the greatest challenges in integrated circuit design is the repeated executions of computationally expensive SPICE simulations, particularly when highly complex chip testing/verification is involved. Recently, pseudo-transient analysis (PTA) has shown to be one of the most promising continuation SPICE solvers. However, the PTA efficiency is highly influenced by the inserted pseudo-parameters. In this work, we proposed BoA-PTA, a Bayesian optimization accelerated PTA that can substantially accelerate simulations and improve convergence performance without introducing extra errors. Furthermore, our method does not require any pre-computation data or offline training. The acceleration framework can either speed up ongoing, repeated simulations (e.g., Monte-Carlo simulations) immediately or improve new simulations of completely different circuits. BoA-PTA is equipped with cutting-edge machine learning techniques, such as deep learning, Gaussian process, Bayesian optimization, non-stationary monotonic transformation, and variational inference via reparameterization. We assess BoA-PTA in 43 benchmark circuits and real industrial circuits against other SOTA methods and demonstrate an average of 1.5x (maximum 3.5x) for the benchmark circuits and up to 250x speedup for the industrial circuit designs over the original CEPTA without sacrificing any accuracy. Wei W. Xing, Xiang Jin, Tian Feng 0002, Dan Niu, Weisheng Zhao 0001, Zhou Jin 0001 |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2022 | Work-in-Progress: Toward Energy-efficient Near STT-MRAM Processing Architecture for Neural NetworksabstractThe size of parameters in artificial neural network (NN) applications grows quickly from a handful to the GB-level. The data transmission poses a key challenge for NN, and either neuron is removed or data compression reduces pressure on memory access but cannot successfully decrease data traffic. Therefore, we propose the near spin-transfer-torque magnetic random processing architecture for developing energy-efficient NNs. Our approach provides system architects with a preliminary scheme to obtain real-time transmission that near memory controller directly compresses non-zero elements, and encodes the corresponding index depending on the kernel size. Furthermore, it adjusts the number of multiplication accumulators and avoids unnecessary hardware overheads during computation. The preliminary experimental results demonstrated this design verified with weights that currently achieve up to 3.05x speedup and 29.6% power compared with the unoptimized one. Yueting Li 0001, Bingluo Zhao, Jun Wang 0041, Weisheng Zhao 0001 |
CODES+ISSS | 6 |
| 2022 | Eventor: an efficient event-based monocular multi-view stereo accelerator on FPGA platformabstractEvent cameras are bio-inspired vision sensors that asynchronously represent pixel-level brightness changes as event streams. Event-based monocular multi-view stereo (EMVS) is a technique that exploits the event streams to estimate semi-dense 3D structure with known trajectory. It is a critical task for event-based monocular SLAM. However, the required intensive computation workloads make it challenging for real-time deployment on embedded platforms. In this paper, Eventor is proposed as a fast and efficient EMVS accelerator by realizing the most critical and time-consuming stages including event back-projection and volumetric ray-counting on FPGA. Highly paralleled and fully pipelined processing elements are specially designed via FPGA and integrated with the embedded ARM as a heterogeneous system to improve the throughput and reduce the memory footprint. Meanwhile, the EMVS algorithm is reformulated to a more hardware-friendly manner by rescheduling, approximate computing and hybrid data quantization. Evaluation results on DAVIS dataset show that Eventor achieves up to 24X improvement in energy efficiency compared with Intel i5 CPU platform. Jianlei Yang 0001, Yingjie Qi, Meng Dong, Yuhao Yang 0008, Runze Liu 0001, Weitao Pan, Bei Yu 0001, Weisheng Zhao 0001 |
DAC | 9 |
| 2022 | CP-SRAM: charge-pulsation SRAM marco for ultra-high energy-efficiency computing-in-memoryabstractSRAM-based computing-in-memory (SRAM-CIM) provides fast speed and good scalability with advanced process technology. However, the energy efficiency of the state-of-the-art current-domain SRAM-CIM bit-cell structure is limited and the peripheral circuitry (e.g., DAC/ADC) for high-precision is expensive. This paper proposes a charge-pulsation SRAM (CP-SRAM) structure to achieve ultra-high energy-efficiency thanks to its charge-domain mechanism. Furthermore, our proposed CP-SRAM CIM supports configurable precision (2/4/6-bit). The CP-SRAM CIM macro was designed in 180nm (with silicon verification) and 40nm (simulation) nodes. The simulation results in 40nm show that our macro can achieve energy efficiency of ~2950Tops/W at 2-bit precision, ~576.4 Tops/W at 4-bit precision and ~111.7 Tops/W at 6-bit precision, respectively. He Zhang 0011, Linjun Jiang, Tingran Chen, Junzhan Liu, Wang Kang 0001, Weisheng Zhao 0001 |
DAC | 7 |
| 2022 | Stateful implication logic based on perpendicular magnetic tunnel junctions
Wenlong Cai, Mengxing Wang 0001, Kaihua Cao, Huaiwen Yang, Shouzhong Peng, Huisong Li, Weisheng Zhao 0001 |
Sci. China Inf. Sci. | 7 |
| 2022 | Femtosecond laser-assisted switching in perpendicular magnetic tunnel junctions with double-interface free layer
Luding Wang, Wenlong Cai, Kaihua Cao, Kewen Shi, Bert Koopmans, Weisheng Zhao 0001 |
Sci. China Inf. Sci. | 6 |
| 2022 | Triangle Counting Accelerations: From Algorithm to In-Memory Computing ArchitectureabstractTriangles are the basic substructure of networks and triangle counting (TC) has been a fundamental graph computing problem in numerous fields such as social network analysis. Nevertheless, like other graph computing problems, due to the high memory-computation ratio and random memory access pattern, TC involves a large amount of data transfers thus suffers from the bandwidth bottleneck in the traditional Von-Neumann architecture. To overcome this challenge, in this paper, we propose to accelerate TC with the emerging processing-in-memory (PIM) architecture through an algorithm-architecture co-optimization manner. To enable the efficient in-memory implementations, we come up to reformulate TC with bitwise logic operations (such as AND), and develop customized graph compression and mapping techniques for efficient data flow management. With the emerging computational Spin-Transfer Torque Magnetic RAM (STT-MRAM) array, which is one of the most promising PIM enabling techniques, the device-to-architecture co-simulation results demonstrate that the proposed TC in-memory accelerator outperforms the state-of-the-art GPU and FPGA accelerations by 12.2x and 31.8x, respectively, and achieves a 34x energy efficiency improvement over the FPGA accelerator. Jianlei Yang 0001, Yinglin Zhao, Xiaotao Jia, Rong Yin 0001, Xuhang Chen 0001, Gang Qu 0001, Weisheng Zhao 0001 |
IEEE Trans. Computers | 8 |
| 2022 | S2 Engine: A Novel Systolic Architecture for Sparse Convolutional Neural NetworksabstractConvolutional neural networks (CNNs) have achieved great success in performing cognitive tasks. However, execution of CNNs requires a large amount of computing resources and generates heavy memory traffic, which impose a severe challenge on computing system design. Through optimizing parallel executions and data reuse in convolution, systolic architecture demonstrates great advantages in accelerating CNN computations. However, regular internal data transmission path in traditional systolic architecture prevents the systolic architecture from completely leveraging the benefits introduced by neural network sparsity.Deployment of fine-grained sparsity on the existing systolic architectures is greatly hindered by the incurred computational overheads.In this work, we propose S2Engine a novel systolic architecture that can fully exploit the sparsity in CNNs with maximized data reuse. S2Engine transmits compressed data internally and allows each processing element to dynamically select an aligned data from the compressed dataflow in convolution. Compared to the naive systolic array, S2Engine achieves about 3.2 and about 3.0 improvements on speed and power efficiency, respectively. Jianlei Yang 0001, Wenzhi Fu, Xingzhou Cheng, Xucheng Ye, Pengcheng Dai, Weisheng Zhao 0001 |
IEEE Trans. Computers | 6 |
| 2022 | Accelerating Graph-Connected Component Computation With Emerging Processing-In-Memory ArchitectureabstractComputing the connected component (CC) of a graph is a basic graph computing problem, which has numerous applications like graph partitioning and pattern recognition. Existing methods for computing CC suffer from memory wall problems because of the frequent data transmission between CPU and memory. To overcome this challenge, in this article, we propose to accelerate CC computation with the emerging processing-in-memory (PIM) architecture through an algorithm–architecture co-design manner. The innovation lies in computing CC with bitwise logical operations (such as AND and OR), and the customized data flow management methods to accelerate computation and reduce energy consumption. As a proof of concept, experimental results with computational spin-transfer torque magnetic RAM (STT-MRAM) arrays demonstrate on average$19.8\times $and$12.4\times $speedups compared with the CPU and GPU implementations, and a$35.4 \times $energy efficiency improvement over the CPU implementation. Moreover, we investigate the potential associations between graph computing and bitwise Boolean logic, which could help design more general in-memory graph computing accelerators in the future. Xuhang Chen 0001, Xiaotao Jia, Jianlei Yang 0001, Gang Qu 0001, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2021 | FedSkel: Efficient Federated Learning on Heterogeneous Systems with Skeleton Gradients UpdateabstractFederated learning aims to protect users' privacy while performing data analysis from different participants. However, it is challenging to guarantee the training efficiency on heterogeneous systems due to the various computational capabilities and communication bottlenecks. In this work, we propose FedSkel to enable computation-efficient and communication-efficient federated learning on edge devices by only updating the model's essential parts, named skeleton networks. FedSkel is evaluated on real edge devices with imbalanced datasets. Experimental results show that it could achieve up to 5.52x speedups for CONV layers' back-propagation, 1.82x speedups for the whole training process, and reduce 64.8% communication cost, with negligible accuracy loss. Junyu Luo 0002, Jianlei Yang 0001, Xucheng Ye, Xin Guo 0008, Weisheng Zhao 0001 |
CIKM | 5 |
| 2021 | SpinLiM: Spin Orbit Torque Memory for Ternary Neural Networks Based on the Logic-in-Memory ArchitectureabstractLogic-in-memory architecture based on spintronic memories shows fascinating prospects in neural networks (NNs) for its high energy efficiency and good endurance. In this work, we leveraged two magnetic tunnel junctions (MTJs), which are driven by the interplay of field-free spin orbit torque (SOT) and spin transfer torque (STT) effects, to achieve a novel statefullogic-in-memory paradigm for ternary multiplication operations. Based on this paradigm, we further proposed a highly parallel array structure to serve for ternary neural networks (TNNs). Our results demonstrate the advantage of our design in power consumption compared with CPU, GPU and other state-of-the-art works. Lichuan Luo, He Zhang 0011, Jinyu Bai, Youguang Zhang, Wang Kang 0001, Weisheng Zhao 0001 |
DATE | 6 |
| 2021 | A Reconfigurable Arbiter PUF Based on STT-MRAMabstractWith the rapid development of the Internet of Things (IoT) infrastructure, electronic devices are becoming ubiquitous, in which authentication and secure communication are required. As a result, novel hardware security primitives have been developed to overcome the deficiencies of conventional security methods and address the growing security issues. Physical unclonable function (PUF) is an emerging hardware security primitive that plays an important role in authenticity and reliability of integrated circuits (ICs). Spin-transfer torque magne- toresistive random access memory (STT-MRAM) is a promising technology that is dense, fast, non-volatile, highly endurant and energy-efficient. STT-MRAM is considered a promising primitive as it has several intrinsic randomness sources, such as stochastic switching, process variations and statistical read/write failures. This paper proposes a novel hybrid STT-MRAM/complementary metal-oxide semiconductor (CMOS) based reconfigurable arbiter PUF. The functionality of the design is validated by a 28nm CMOS technology and a compact magnetic tunnel junction (MTJ) model. Simulation results show that the proposed PUF has a mean intra-hamming distance (HD) of 0.24%, a mean inter-HD of 51.1% and passes the National Institute of Standards and Technology (NIST) statistical tests. You Wang 0002, Zhengyi Hou, Deming Zhang, Erya Deng, Weisheng Zhao 0001 |
ISCAS | 7 |
| 2021 | Spin-Orbit Torque Nonvolatile Flip-Flop DesignsabstractFlip-flops (FFs) are basic units in electronic circuits. Recently, nonvolatile FFs (NVFFs) have attracted great interests for power-gating applications and a variety of NVFFs have been proposed by integrating nonvolatile memory devices. Among them, magnetic tunnel junction (MTJ) based NVFFs show considerable potential in terms of zero static power consumption and high endurance. Nevertheless, the mainstream spin transfer torque (STT) effect based MTJ switching approach for data storing still consumes much dynamic power and long delay, limiting the system performance and data reliability. The spin-orbit torque (SOT) effect provides an alternative approach for high-speed and low-power MTJ switching, therefore rather promising for NVFF design. In this work, we propose four NVFF designs based on the FF architectures (either DFF or SRFF) and perpendicular MTJ (pMTJ). The circuit structures and operations are investigated, and the performance is evaluated and compared at the 40 nm process technology node. Simulation results show that the proposed NVFFs can achieve high read speed (<; 200 ps), low read power consumption (<; 10 fJ) and area efficiency. Erya Deng, Wang Kang 0001, Weisheng Zhao 0001, Shaoqian Wei, You Wang 0002, Deming Zhang |
ISCAS | 3 |
| 2021 | SpinSim: A Computer Architecture-Level Variation Aware STT-MRAM Performance Evaluation FrameworkabstractWith low power consumption, fast access speed, high scalability and infinite endurance, spin-transfer torque magnetoresistive random access memory (STT-MRAM) is considered as one of the most promising alternatives to SRAM. However, The performance of STT-MRAM is significantly influenced by several reliability issues, such as process variations and stochastic switching. Most of the reliability analysis of relative circuits are performed at bit-cell and memory level, while that at computer-system level is missing. This paper proposes an efficient framework for performance evaluation of STT-MRAM on computer architecture-level implemented by GEM5+NVMain co-simulator in consideration of the reliability issues. The results show that the overall average latency and energy of STT-MRAM can be up to 5.996% and 20.65% larger than that of the nominal cases in a computer system-level memory architecture taking reliability issues into account. Because reliability issues are considered during the design phase, our framework can provide more accurate performance evaluation and contribute to a higher yield of STT-MRAM based computer systems. You Wang 0002, Zhengyi Hou, Deming Zhang, Erya Deng, Gefei Wang, Weisheng Zhao 0001 |
ISCAS | 8 |
| 2021 | Computing-in-Memory Paradigm Based on STT-MRAM with Synergetic Read/Write-Like ModesabstractWith the surge in demand for data storage and processing in emerging applications, the traditional CMOS-based Von-Neumann architecture is facing challenges such as memory wall and static power consumption. In order to conquer the above-mentioned bottlenecks in computing systems, computing in-memory (CiM) architectures based on non-volatile memory (NVM) have been widely researched. In this paper, we propose a CiM paradigm based on spin-transfer torque magnetic random access memory (STT-MRAM), which combines common read-like mode (RLM) and write-like mode (WLM). On the basis of realizing the basic functions AND/OR/NAND/NOR, our design coordinates the high speed of RLM and the integrity of WLM to perform complex operations like full-adder (FA) and XOR/XNOR. In addition, the high speed and low power consumption of the proposed CiM paradigm are established by circuit-level simulation with a 40 nm design kit. Chao Wang 0094, Zhaohao Wang, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 4 |
| 2021 | Fully Single Event Double Node Upset Tolerant Design for Magnetic Random Access MemoryabstractBenefitting from its non-volatility, high speed, low power and inherent radiation hardened characteristic, magnetic random access memory (MRAM) has been used in aerospace and avionic electronics. Owing to its high sensing reliability, precharge differential sense amplifier (PCDSA) has been proposed and widely used in MRAM products. However, such PCDSA is based on the conventional CMOS technology and its sensing result is prone to be affected by the single event upset (SEU) and even the single event double node upset (SEDU) when the CMOS technology node shrinks into the nanometer scale. In this paper, we propose a novel PCDSA to tolerate the SEDU, in which the special three-input C-element that behaves as an inverter when its inputs have the same logic value and holds its previous value when its inputs have the different logic values is employed. By using a physics-based STT-MTJ compact model and a commercial CMOS 40 nm design kit, hybrid simulations have been performed to demonstrate its functionality and evaluate its performance. Simulation results show that it can fully tolerate the SEDU when the amount of the deposited charge (Qinj) reaches up to 2 pC. In the worst case where the Qinjis 2 pC, it can achieve a small recover time of 1.3368 ns and low recover energy dissipation of 1.967 pJ with the optimized VDDof 1 V. Deming Zhang, Lang Zeng, You Wang 0002, Bi Wang 0002, Erya Deng, Chuanjie Wang, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 11 |
| 2021 | Brief Industry Paper: optimizing Memory Efficiency of Graph Neural Networks on Edge Computing PlatformsabstractGraph neural networks (GNN) have achieved state-of-the-art performance on various industrial tasks. However, the poor efficiency of GNN inference and frequent Out-of-Memory (OOM) problem limit the successful application of GNN on edge computing platforms. To tackle these problems, a feature decomposition approach is proposed for memory efficiency optimization of GNN inference. The proposed approach could achieve outstanding optimization on various GNN models, covering a wide range of datasets, which speeds up the inference by up to 3×. Furthermore, the proposed feature decomposition could significantly reduce the peak memory usage (up to 5× in memory efficiency improvement) and mitigate OOM problems during GNN inference. Jianlei Yang 0001, Yeqi Gao, Yingjie Qi, Yunli Chen, Pengcheng Dai, Weisheng Zhao 0001, Chunming Hu |
RTAS | 9 |
| 2021 | Spintronics for Energy- Efficient Computing: An Overview and OutlookabstractFrom the discovery of giant magnetoresistance (GMR) to tunnel magnetoresistance (TMR), their subsequent application in large capacity hard disk drives (HDDs) greatly speeded up the information era over the past decades. However, the growing demand for big-data storage and processing is limited by the von-Neumann architecture due to the memory bottleneck and power dissipation. Taking advantage of nonvolatility, high speed, and low power, magnetic random access memory (MRAM) becomes a promising candidate to overcome this limitation through processing-in-memory (PIM) architectures. In this article, we provide an overview of existing technology and give a roadmap of spintronic devices for future energy-efficient computing and its relevant integration architectures. We begin with the fundamentals of Toggle-MRAM and spin-transfer torque (STT)-MRAM, which already have commercial applications. We then introduce spin-orbit torque (SOT), a critical mechanism to realize low-power data manipulation in the next generation of MRAM and summarize the recent experimental breakthroughs of field-free SOT switching schemes. Finally, we present MRAM-based PIM architectures and novel spintronic devices, provide an application outlook, and deliver the future development potential of energy-efficient computing systems. Zongxia Guo, Jialiang Yin, Daoqian Zhu, Kewen Shi, Gefei Wang, Kaihua Cao, Weisheng Zhao 0001 |
Proc. IEEE | 8 |
| 2021 | Efficient Computation Reduction in Bayesian Neural Networks Through Feature Decomposition and MemorizationabstractThe Bayesian method is capable of capturing real-world uncertainties/incompleteness and properly addressing the overfitting issue faced by deep neural networks. In recent years, Bayesian neural networks (BNNs) have drawn tremendous attention to artificial intelligence (AI) researchers and proved to be successful in many applications. However, the required high computation complexity makes BNNs difficult to be deployed in computing systems with a limited power budget. In this article, an efficient BNN inference flow is proposed to reduce the computation cost and then is evaluated using both software and hardware implementations. A feature decomposition and memorization (DM) strategy is utilized to reform the BNN inference flow in a reduced manner. About half of the computations could be eliminated compared with the traditional approach that has been proved by theoretical analysis and software validations. Subsequently, in order to resolve the hardware resource limitations, a memory-friendly computing framework is further deployed to reduce the memory overhead introduced by the DM strategy. Finally, we implement our approach in Verilog and synthesize it with a 45-nm FreePDK technology. Hardware simulation results on multilayer BNNs demonstrate that, when compared with the traditional BNN inference method, it provides an energy consumption reduction of 73% and a 4× speedup at the expense of 14% area overhead. Xiaotao Jia, Jianlei Yang 0001, Runze Liu 0001, Sorin Cotofana, Weisheng Zhao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2020 | SparseTrain: Exploiting Dataflow Sparsity for Efficient Convolutional Neural Networks TrainingabstractTraining Convolutional Neural Networks (CNNs) usually requires a large number of computational resources. In this paper, SparseTrain is proposed to accelerate CNN training by fully exploiting the sparsity. It mainly involves three levels of innovations: activation gradients pruning algorithm, sparse training dataflow, and accelerator architecture. By applying a stochastic pruning algorithm on each layer, the sparsity of back-propagation gradients can be increased dramatically without degrading training accuracy and convergence rate. Moreover, to utilize both natural sparsity (resulted from ReLU or Pooling layers) and artificial sparsity (brought by pruning algorithm), a sparse-aware architecture is proposed for training acceleration. This architecture supports forward and back-propagation of CNN by adopting 1-Dimensional convolution dataflow. We have built a cycle-accurate architecture simulator to evaluate the performance and efficiency based on the synthesized design with 14nm FinFET technologies. Evaluation results on AlexNet/ResNet show that SparseTrain could achieve about 2.7× speedup and 2.2× energy efficiency improvement on average compared with the original training process. Pengcheng Dai, Jianlei Yang 0001, Xucheng Ye, Xingzhou Cheng, Junyu Luo 0002, Linghao Song, Yiran Chen 0001, Weisheng Zhao 0001 |
DAC | 8 |
| 2020 | TCIM: Triangle Counting Acceleration With Processing-In-MRAM ArchitectureabstractTriangle counting (TC) is a fundamental problem in graph analysis and has found numerous applications, which motivates many TC acceleration solutions in the traditional computing platforms like GPU and FPGA. However, these approaches suffer from the bandwidth bottleneck because TC calculation involves a large amount of data transfers. In this paper, we propose to overcome this challenge by designing a TC accelerator utilizing the emerging processing-in-MRAM (PIM) architecture. The true innovation behind our approach is a novel method to perform TC with bitwise logic operations (such as AND), instead of the traditional approaches such as matrix computations. This enables the efficient in-memory implementations of TC computation, which we demonstrate in this paper with computational Spin-Transfer Torque Magnetic RAM (STT-MRAM) arrays. Furthermore, we develop customized graph slicing and mapping techniques to speed up the computation and reduce the energy consumption. We use a device-to-architecture co-simulation framework to validate our proposed TC accelerator. The results show that our data mapping strategy could reduce 99.99% of the computation and 72% of the memory WRITE operations. Compared with the existing GPU or FPGA accelerators, our in-memory accelerator achieves speedups of 9× and 23.4×, respectively, and a 20.6× energy efficiency improvement over the FPGA accelerator. Jianlei Yang 0001, Yinglin Zhao, Yingjie Qi, Meichen Liu, Xingzhou Cheng, Xiaotao Jia, Gang Qu 0001, Weisheng Zhao 0001 |
DAC | 10 |
| 2020 | High-Density, Low-Power Voltage-Control Spin Orbit Torque Memory with Synchronous Two-Step Write and Symmetric Read TechniquesabstractVoltage-control spin orbit torque (VC-SOT) magnetic tunnel junction (MTJ) has the potential to achieve high-speed and low-power spintronic memory, owing to the adaptive voltage modulated energy barrier of the MTJ. However, the three-terminal device structure needs two access transistors (one for write operation and the other one for read operation) and thus occupies larger bit-cell area compared to two terminal MTJs. A feasible method to reduce area overhead is to stack multiple VC-SOT MTJs on a common antiferromagnetic strip to share the write access transistors. In this structure, high density can be achieved. However, write and read operations face problems and the design space is not sure given a strip length. In this paper, we propose a synchronous two-step multi-bit write and symmetric read method by exploiting the selective VC-SOT driven MTJ switching mechanism. Then hybrid circuits are designed and evaluated based a physics-based VC-SOT MTJ model and a 40nm CMOS design-kit to show the feasibility and performance of our method. Our work enables high-density, low-power, high-speed voltage-control SOT memory. Wang Kang 0001, Liuyang Zhang, He Zhang 0011, Brajesh Kumar Kaushik, Weisheng Zhao 0001 |
DATE | 6 |
| 2020 | Towards Systems Education for Artificial Intelligence: A Course Practice in Intelligent Computing ArchitecturesabstractWith the rapid development of artificial intelligence (AI) community, education in AI is receiving more and more attentions. There have been many AI related courses in the respects of algorithms and applications, while not many courses in system level are seriously taken into considerations. In order to bridge the gap between AI and computing systems, we are trying to explore how to conduct AI education from the perspective of computing systems. In this paper, a course practice in intelligent computing architectures are provided to demonstrate the system education in AI era. The motivation for this course practice is first introduced as well as the learning orientations. The main goal of this course aims to teach students for designing AI accelerators on FPGA platforms. The elaborated course contents include lecture notes and related technical materials. Especially several practical labs and projects are detailed illustrated. Finally, some teaching experiences and effects are discussed as well as some potential improvements in the future. Jianlei Yang 0001, Xiaopeng Gao, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2020 | A Novel In-memory Computing Scheme Based on Toggle Spin Torque MRAMabstractThis paper proposes a novel in-memory computing (IMC) scheme based on toggle spin torque magnetic random access memory (TST-MRAM), called TST-IMC, which makes full use of the unique TST writing mechanism. In this scheme, all of the computing results are directly written in bit-cells without transferring data out of the memory array. Varied Boolean logic operations, such as, NAND, NOR and XOR, can be achieved by specially configuring decision cells. We can also implement three-input majority logic through replacing a decision cell with a datum cell, which can further be used to realize the carry of full-adder. By using 28 nm CMOS technology node and 50 nm-diameter TST-MRAM, we perform mixed simulations to validate the functionality of the proposed TST-IMC scheme. Simulation results show that XOR logic operation can be carried out within 4 ns at 1.8 V supply voltage while the other basic logic operations can be faster, i.e. within 2 ns. In addition, TST-IMC 33% less time and 44% energy saved comparing with existing IMC schemes. Yining Bai, Yue Zhang 0010, Guanda Wang, Zhizhong Zhang 0004, Zhenyi Zheng, Kun Zhang 0030, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 8 |
| 2020 | An In-memory Highly Reconfigurable Logic Circuit Based on Diode-assisted Enhanced Magnetoresistance DeviceabstractIn the post-Moore era, in order to solve the problem of von Neumann bottleneck and memory wall caused by separation of memory and processor, in-memory-processing (IMP) technique has aroused great attention. Novel non-volatile memory (NVM) based on spintronic devices shows promise for satisfying the needs of low-power consumption and high speed for IMP. However, most spintronic memories based on magnetic tunnel junctions (MTJs) can only implement simple and specific logic functions due to the limits of single device and circuit structure. Otherwise, performing logic functions in memory generates vast dynamic power consumption during frequent reading and writing processes because of the high resistance of miniaturized MTJ. In this paper, we propose an in-memory highly reconfigurable logic circuit based on diode-assisted enhanced magnetoresistance (DEMR) device. Our circuit can realize 16 different logic functions with extremely limited circuit area benefiting from the special structure of DEMR device. With appropriate adjustment of control bit and current, the proposed circuit can further implement complex functions like full adder. The proposed reconfigurable circuit can flexibly meet the performance requirements in different scenarios and will contribute a lot for future in-memory chip design. Yue Zhang 0010, Kun Zhang 0030, Zhizhong Zhang 0004, Youguang Zhang, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 7 |
| 2020 | Deep Neural Network accelerator with Spintronic MemoryabstractUtilizing emerging nonvolatile memories to accelerate deep neural network (DNN) has been considered as one of the promising approaches to solve the bottleneck of data transfer during the multiplication and accumulation (MAC). Among them, spintronic memories show tempting prospect due to their low access power, fast access speed, high density, and relatively mature process. As shown in fig.1, according to the principle to achieve DNN computing, it can be mainly divided into three different technical routes. The first one is an "analog" method [1, 2], as shown in fig.1(a). By transforming the digital input signals into multi-level voltage signals, and applying them to different columns of the memory array, the MAC results can be obtained in different columns with current integrator and analog to digital converter (ADC). Besides, the WL drivers can control the pulse width of different rows, to achieve the effect of multi-bit weights. This method can theoretically achieve high energy efficiency and computing speed. However, the variation of magnetic tunnel junction (MTJ) may have influence on the computing accuracy. Besides, the power consumption and area overhead of the ADC are also challenging. The other two methods are in a "digital" way, and they realize MAC computing through row-by-row read/write operation. Fig.1(b) shows the second reading-based method [3]. The weights of the neural network are stored in the memory cell. By putting the input signal to the modified sensing amplifier (SA), it can also achieve XOR function, which is the core of binary NN, with the content stored in the memory cell. Nevertheless, the modification to the SA is usually to add extra transistors in the read path, which will increase the bit error rate. Fig.1(c) shows the diagram of the last one, which is based on the "stateful logic" [4]. The input data is sent to the modified write driver when the WL receiving weight signals from outside I/O. Based on a unique logic paradigm, it can realize XOR function for BNN within 1 or several memory cells during a write cycle. In this talk, we will review the main research status of DNN accelerators based on spintronic memories. Particularly, our recent work on DNN accelerating will be introduced, which can be implemented with different spintronic memories. He Zhang 0011, Wang Kang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2020 | PRISM: Energy-Efficient Polymorphic Operation Based on Spin-Orbit Torque Memory for Reconfigurable ComputingabstractEmerging Non-Volatile Memories (NVMs) including resistive RAM (ReRAM), phase-change memory (PCM), and magnetic RAM (MRAM), have opened up new pathways for the NVM-based reconfigurable computing. Those NVMs technologies can achieve significant energy-efficient computational operations with only minor modification of the peripheral circuits. However, the supported operations are limited by the array structure and low energy-efficiency of implementing the computation using the memory array. In this paper, the Spin Orbit torque-MRAM based polymorphic circuits are proposed to support the reconfigurable computation for reducing the power consumption and improving the functionalities of the single memory array. With the high speed and energy-efficiency write operation, the proposed memory array support both read-out and write-in reconfigurable operations. Liang Chang 0002, Zhaohao Wang, Youguang Zhang, Weisheng Zhao 0001, Jun Zhou 0017 |
ISCAS | 5 |
| 2020 | Efficient Time-Domain In-Memory Computing Based on TST-MRAMabstractIn-memory computing is highly promising to address the processor-memory data transfer bottleneck in current computational paradigm. We firstly propose a timedomain in-memory computing (TIMC) scheme based on highspeed low-power toggle spin torque random access memory (TST-MRAM). The difference of voltage drops of bitline caused by simultaneously-activated bit-cells is reflected to time domain. Reconfigurable logic operations can be performed by utilizing D flip-flops (DFFs) to record the outputs at different moments. In order to demonstrate the advantages of this scheme in terms of speed and energy consumption, an efficient multi-digit addition circuit has been designed and analyzed. Compared with existing IMC schemes, such as spin-transfer torque computing-in-memory (STT-CiM) structure, up to 67% energy saving and 10 times delay improvement can be achieved in the case of four-digit addition by using TIMC scheme. Yue Zhang 0010, Chenyu Lian, Yining Bai, Guanda Wang, Kun Zhang 0030, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 9 |
| 2020 | Computing-in-Memory Architecture Based on Field-Free SOT-MRAM with Self-Reference MethodabstractOn the current computing platforms, the memory wall between processor and memory has become the toughest challenge for the traditional Von-Neumann computer architecture. Computing-in-Memory (CIM) is taken as a promising approach to solving the above bottleneck in computing systems. In this paper, we propose a CIM platform with field-free spinorbit torque magnetic random access memory (SOT-MRAM). The self-reference (SelfRef) method is designed to enhance the read reliability and directly obtain logic results through memory-like read operations without adding logic cells. Memory read/write and logic operations, including NOT, AND/NAND and OR/NOR, can be implemented in the same SOT-MRAM chip. The speed and power penalties caused by SelfRef scheme are acceptable thanks to the ultrafast switching of the SOT. The read reliability and logic correctness of the proposed CIM are demonstrated by hybrid simulation on a 40 nm technology node. Chao Wang 0094, Zhaohao Wang, Jianlei Yang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 6 |
| 2020 | Voltage-Gated Spin-Hall Effect Based Magnetic Non-Volatile Flip-Flop for High Speed, Low Power and Compact Cell AreaabstractIn this paper, we present a novel magnetic nonvolatile flip-flop (MNV-FF) for fast and low-power backup operation with a compact cell area. It employs perpendicular magnetic tunnel junctions (p-MTJs) as its non-volatile data backup storage units and exploits the voltage-gated spin-hall effect (VGSHE) for data backup operation. Benefitting from the assistance of the voltage-controlled magnetic anisotropy (VCMA) effect, the critical write current for 1-ns backup operation can be reduced to 3 μA or even lower, thus resulting in high speed and low power consumption. Moreover, such small write current allows to be driven by the cross-coupled inverters in the master latch, instead of a dedicated write driver, leading to a low cell area overhead. Additionally, by using an antiferromagnetic (AFM) metal that can provide both an exchange bias and the SHE instead of the heavy metal, no external magnetic field is required, making it suitable for practical applications. Our simulation results show that our proposed VGSHE-based MNV-FF can achieve 58.2× less backup energy, 1.85× less backup delay and 1.625× less cell area overhead than the previous SHE-based MNV-FF. Deming Zhang, Lang Zeng, Weisheng Zhao 0001 |
ISCAS | 5 |
| 2020 | An STT-MRAM based reconfigurable computing-in-memory architecture for general purpose computing
Xiaotao Jia, Jianlei Yang 0001, Weisheng Zhao 0001 |
CCF Trans. High Perform. Comput. | 7 |
| 2020 | Prototyping federated learning on edge computing systems
Jianlei Yang 0001, Yixiao Duan, Huanyu Zhou, Jingyuan Wang 0001, Weisheng Zhao 0001 |
Frontiers Comput. Sci. | 6 |
| 2020 | A Comparative Cross-layer Study on Racetrack Memories: Domain Wall vs SkyrmionabstractRacetrack memory (RM), a new storage scheme in which information flows along a nanotrack, has been considered as a potential candidate for future high-density storage device instead of hard disk drive (HDD). The first RM technology, which was proposed in 2008 by IBM, relies on a train of opposite magnetic domains separated by domain walls (DWs), named DW-RM. After 10 years of intensive research, a variety of fundamental advancements has been achieved; unfortunately, no product has been available until now. With increasing effort and resources dedicated to the development of DW-RM, it is likely that new materials and mechanisms will soon be discovered for practical applications. However, new concepts might also be on the horizon. Recently, an alternative information carrier, magnetic skyrmion, which was experimentally discovered in 2009, has been regarded as a promising replacement of DW for RM, named skyrmion-based RM (SK-RM). Intensive effort has been involved and amazing advances have been made in observing, writing, manipulating, and deleting individual skyrmions. So, what is the relationship between DW and skyrmion? What are the key differences between DW and skyrmion, or between DW-RM and SK-RM? What benefits could SK-RM bring and what challenges need to be addressed before application? In this review article, we intend to answer these questions through a comparative cross-layer study between DW-RM and SK-RM. This work will provide guidelines, especially for circuit and architecture researchers on RM. Wang Kang 0001, Bi Wu 0002, Xing Chen 0012, Daoqian Zhu, Zhaohao Wang, Xichao Zhang, Youguang Zhang, Weisheng Zhao 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 9 |
| 2020 | Write Back Energy Optimization for STT-MRAM-based Last-level Cache with Data Pattern CharacterizationabstractTraditional memory technologies face severe challenges in meeting the ever-increasing power and memory bandwidth requirements for high-performance computing and big-data analyses. Several emerging memory technologies are promising as the replacements of SRAM or DRAM. Among them, STT-MRAM can be used to replace SRAM as the last-level cache (LLC). However, it suffers from high write energy and latency. In this article, we investigate data patterns written from SRAM-based upper-level cache to STT-MRAM-based LLC to explore the write energy reduction potential. Depending on the data layout within a cache line, redundant bits can be identified and eliminated from write back operations to save STT-MRAM write energy. We also propose a dynamic profiling method to accommodate different application characteristics. The extensive simulation results show that write energy can be saved by 37.05% ∼ 38.89% for static profiling and 19.76% ∼ 34.29% for dynamic profiling. Keren Liu, Bi Wu 0002, Weisheng Zhao 0001, Yuanqing Cheng, Ying Wang 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2020 | Hardware Security in Spin-based Computing-in-memory: Analysis, Exploits, and Mitigation TechniquesabstractComputing-in-memory (CIM) is proposed to alleviate the processor-memory data transfer bottleneck in traditional von Neumann architectures, and spintronics-based magnetic memory has demonstrated many facilitation in implementing CIM paradigm. Since hardware security has become one of the major concerns in circuit designs, this article, for the first time, investigates spin-based computing-in-memory (SpinCIM) from a security perspective. We focus on two fundamental questions: (1) How can the new SpinCIM computing paradigm be exploited to enhance hardware security?; (2) What security concerns has this new SpinCIM computing paradigm incurred? Jianlei Yang 0001, Yinglin Zhao, Xiaotao Jia, Gang Qu 0001, Weisheng Zhao 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 6 |
| 2020 | SPINBIS: Spintronics-Based Bayesian Inference System With Stochastic ComputingabstractBayesian inference is an effective approach for solving statistical learning problems, especially with uncertainty and incompleteness. However, Bayesian inference is a computing-intensive task whose efficiency is physically limited by the bottlenecks of conventional computing platforms. In this paper, a spintronics-based stochastic computing (SC) approach is proposed for efficient Bayesian inference. The inherent stochastic switching behaviors of spintronic devices are exploited to build a stochastic bitstream generator (SBG) for SC with hybrid CMOS/magnetic tunnel junction (MTJ) circuits design. Aiming to improve the inference efficiency, an SBG sharing strategy is leveraged to reduce the required SBG array scale by integrating a switch network between SBG array and SC logic. A device-to-architecture level framework is proposed to evaluate the performance of spintronics-based Bayesian inference system (SPINBIS). Experimental results on data fusion applications have shown that SPINBIS could improve the energy efficiency about 12× than MTJ-based approach with 45% design area overhead and about 26× than FPGA-based approach. Xiaotao Jia, Jianlei Yang 0001, Pengcheng Dai, Runze Liu 0001, Yiran Chen 0001, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2020 | A Novel High Performance and Energy Efficient NUCA Architecture for STT-MRAM LLCs With Thermal ConsiderationabstractAs the speed gap of the modern processor and the off-chip main memory enlarges, on-chip cache capacity increases to sustain the performance scaling. As a result, the cache power occupies a large portion of the total power budget. Spin transfer torque magnetic memory (STT-MRAM) is proposed as a promising solution for the low power cache design due to its high integration density and ultralow leakage power. Nevertheless, the high write power and latency of STT-MRAM become new barriers for the commercialization of this emerging technology. In this paper, we investigate the thermal effect on the access performance of STT-MRAM, and observe that the temperature can affect the write delay and energy significantly. Then, we explore the nonuniform cache access (NUCA) design of the chip-multiprocessors with STT-MRAM-based last level cache (LLC). A thermal aware data migration policy, called “Thermosiphon,” which takes advantage of the thermal property of STT-MRAM, is proposed to reduce the LLC write energy. This policy splits the LLC into different regions dynamically based on the thermal distribution monitored by thermal sensors available on-chip, and adaptively migrates write intensive data among different thermal regions considering the thermal gradient. Compared to the conventional NUCA design, our proposed design can save 41.2% write energy at most and 13.01% on average with negligible hardware overhead. Bi Wu 0002, Pengcheng Dai, Yuanqing Cheng, Ying Wang 0001, Jianlei Yang 0001, Zhaohao Wang, Dijun Liu, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2019 | ZUMA: Enabling Direct Insertion/Deletion Operations with Emerging Skyrmion Racetrack MemoryabstractData insertion and deletion are common operations exist in various applications. However, traditional memory architecture can only perform an indirect insertion/deletion with multiple data read and write operations, which is significantly time and energy consuming. To mitigate this problem, we propose to leverage the unique capability of emerging skyrmion racetrack memory technology that it can naturally support direct insertion/deletion operations inside a racetrack. In this work, we first present a circuit level model for skyrmion racetrack memory. Then, we further propose a novel memory architecture to enable an efficient large size data insertion/deletion. With the help of the model and the architecture, we study several potential applications to leverage the insertion and deletion operations. Experimental results demonstrate that the efficiency of these operations can be substantially improved. Zheng Liang 0003, Guangyu Sun 0003, Wang Kang 0001, Xing Chen 0012, Weisheng Zhao 0001 |
DAC | 5 |
| 2019 | eSLAM: An Energy-Efficient Accelerator for Real-Time ORB-SLAM on FPGA PlatformabstractSimultaneous Localization and Mapping (SLAM) is a critical task for autonomous navigation. However, due to the computational complexity of SLAM algorithms, it is very difficult to achieve real-time implementation on low-power platforms. We propose an energy-efficient architecture for real-time ORB (Oriented-FAST and Rotated-BRIEF) based visual SLAM system by accelerating the most time-consuming stages of feature extraction and matching on FPGA platform. Moreover, the original ORB descriptor pattern is reformed as a rotational symmetric manner which is much more hardware friendly. Optimizations including rescheduling and parallelizing are further utilized to improve the throughput and reduce the memory footprint. Compared with Intel i7 and ARM Cortex-A9 CPUs on TUM dataset, our FPGA realization achieves up to 3× and 31× frame rate improvement, as well as up to 71× and 25× energy efficiency improvement, respectively. Runze Liu 0001, Jianlei Yang 0001, Yiran Chen 0001, Weisheng Zhao 0001 |
DAC | 4 |
| 2019 | CORN: In-Buffer Computing for Binary Neural NetworkabstractBinary Neural Networks (BNNs) have obtained great attention since they reduce memory usage and power consumption as well as achieve a satisfying recognition accuracy on Image Classification. In particular to the computation of BNNs, the multiply-accumulate operations of convolution-layer are replaced with the bit-wise operations (XNOR and pop-count). Such bit-wise operations are well suited for the hardware accelerator such as in-memory computing (IMC). However, an additional digital processing unit (DPU) is required for the pop-count operation, which induces considerable data movement between the Process Engines (PEs) and data buffers reducing the efficiency of the IMC. In this paper, we present a BNN computing accelerator, namely CORN, which consists of a Spin-Orbit-Torque Magnetic RAM (SOT-MRAM) based data buffer to perform the majority operation (to replace the pop-count process) with the SOT-MRAM-based IMC to accelerate the computing of BNNs. CORN can naturally implement the XNOR operation in the NVM memory array, and feed results to the computing data buffer for the majority write operation. Such a design removes the pop-counter implemented by the DPU and reduces data movement between the data buffer and the memory array. Based on the evaluation results, CORN achieves 61% and 14% power saving with 1.74× and 2.12× speedup, compared to the FPGA and DPU based IMC architecture, respectively. Liang Chang 0002, Zhaohao Wang, Youguang Zhang, Weisheng Zhao 0001, Yuan Xie 0001 |
DATE | 5 |
| 2019 | Voltage-Controlled Magnetoelectric Memory Bit-cell Design With Assisted Body-bias in FD-SOIabstractVoltage-controlled magnetic anisotropy (VCMA)-magnetic tunnel junction (MTJ) is incorporated into FD-SOI CMOS technology. The design space of 1 transistor-1 MTJ (1T-1M) bit-cell is explored through varied VCMA pulse duration/amplitude and scaling down transistor dimensions. The design point with 1.1 V VCMA pulse amplitude, 0.44 ns pulse duration and W/L = 400 nm/30 nm access transistor shows the ultra low write energy in VCMA-MTJ based bit-cell. It achieves a minimum 3.18 fJ/bit switching energy with 28-nm FD-SOI process. Access transistor sizing is studied, while the ultra low power implementation may lead to MTJ switching failure. Voltage assisted techniques for failure mitigation are proposed based on body-bias generator (BBG). The BBG not only provides VCMA pulse signal to control MTJ barrier, but also generates body-bias to boost the transistor performance. In the presence of forward body-bias (FBB) and increased VCMA pulse level, the proposed strategy is effective in switching failure compensation as well as writing delay improvement. Hao Cai 0001, Menglin Han, Weiwei Shan, Jun Yang 0006, You Wang 0002, Wang Kang 0001, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 7 |
| 2019 | A Skyrmion Racetrack Memory based Computing In-memory Architecture for Binary Neural Convolutional NetworkabstractA Skyrmion Racetrack Memory (SRM) based Computing In-Memory Architecture (SRM-CIM) was proposed in this paper. Both data and computing operation can be achieved in SRM-CIM. SRM-CIM is used to support convolutional computing in Binary Convolutional Neural Network (BCNN). Experimental results show that SRM-CIM achieves 98.7% and 82% energy reduction when compared with RRAM and SOT-MRAM based counterparts. Yinglin Zhao, Shouyi Yin, Youguang Zhang, Shaojun Wei, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 7 |
| 2019 | Magnetic Skyrmion-Based Neural Recording System Design for Brain Machine InterfaceabstractNext-generation brain machine interface demand a high-channel-count neural recording system to wirelessly monitor activities of thousands of neurons. In order to achieve high-density neural recording, further development of single recording channel comprised of a neural amplifier front-end (AFE) and an analog-to-digit converter (ADC) is critical. Despite the great progress made in CMOS implementation of custom-designed neural recording system, hybrid limitations of increasing area and power consumption in line with Moore's law drove great demand for post-CMOS substitutes. Magnetic skyrmion with nano particle-like and non-volatile properties are of both fundamental and applied interests for future bio-inspired electronics. In this work, we propose a compact model including both AFE and ADC based on current-induced skyrmion motion. The proposed system achieved a power consumption of 0.63 pJ/channel with an area overhead of 0.14 μm2. The purpose of this work is to explore the feasibility of magnetic skyrmion for building large-scale, dense neuronal recording system which could pave a new way for future brain machine interface application. Biao Pan, Wang Kang 0001, Xing Chen 0012, Jinyu Bai, Jianlei Yang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 8 |
| 2019 | SR-WTA: Skyrmion Racing Winner-Takes-All Module for Spiking Neural ComputingabstractSpiking neural network (SNN) has emerged as one of the popular architectures in complex pattern recognition and classification tasks. However, hardware implementation of such algorithms using conventional CMOS based neuron consume resources and power that are orders of magnitude higher than that in human brain. This can be attributed to the mismatch of the computational architecture between biological brain and the current Boolean logic computing platform. Magnetic skyrmions have been intensively studied as a prospective information carrier in neuromorphic computing hardware design. In this work, a compact time-domain skyrmion-racing winner-takes-all (SR-WTA) leaky-integrate-fire (LIF) spiking neuron network is presented for the first time. The skyrmion motion dynamics in the LIF neuron and the behaviors of the neuron network was investigated comprehensively. Both SPICE and micromagnetic simulations are performed to evaluate the functionality and performance of the proposed SR-WTA based SNN. Biao Pan, Wang Kang 0001, Xing Chen 0012, Jinyu Bai, Jianlei Yang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 7 |
| 2019 | Modulation and Demodulation of Digital Frequency Shift Keying System Based on Spin Torque Nano Oscillator with Voltage Controlled Magnetic Anisotropy EffectabstractIn this work, a spin torque nano oscillator (STNO) device whose frequency can be tuned by Voltage Controlled Magnetic Anisotropy effect (VCMA) is proposed. The requirement of magnetic bias field in previous STNO devices is eliminated by the introduction of VCMA effect. Based on VCMA-STNO, a novel architecture is proposed which can compose of a modulation/demodulation digital frequency shift keying (DFSK) communication system. The proposed architecture utilizes VCMA-STNO as core devices and is much simpler comparing with its CMOS counterpart. The proposed VCMA-STNO modulation/demodulation architecture will help to design next generation spintronics DFSK communication system. Lang Zeng, Zuodong Zhang, Haoxuan Chen, Tianqi Gao, Deming Zhang, Mingzhi Long, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 8 |
| 2019 | An STT-MRAM Based in Memory Architecture for Low Power Integral ComputingabstractThe integral histogram image plays an important role in accelerating the feature computation in vision algorithms. However, the computational process of the integral histogram, called integral computation, has high computational complexity and numerous memory access operations, which limit its wide application. This brief proposes an in-memory computational architecture based on Spin Transfer Torque Magnetic Random Access Memory (STT-MRAM) to solve these problems. The architecture can work in two different modes depending on the requirements: the integral computation mode and the memory mode. The architecture can figure out the integral histogram when in the integral computation mode, and just store the data directly when in the memory mode. Utilizing the non-volatile, high density and low power characteristics of STT-MRAM, we integrate the computational units into the memory array to achieve parallel computation. Reduced number of data transmission between storage units and computation units contributes to cut down the latency and energy consumption. The evaluation results show that, comparing with the state-of-the-art work, our architecture provides$1.1\times \sim 9\times$performance improvements and reduces 87.4$\sim$97.3 percent energy consumption for$64\times 64\sim 512\times 512$size images, just with a 8 percent area overhead. Yinglin Zhao, Wang Kang 0001, Shouyi Yin, Youguang Zhang, Shaojun Wei, Weisheng Zhao 0001 |
IEEE Trans. Computers | 7 |
| 2019 | Exploiting Spin-Orbit Torque Devices As Reconfigurable Logic for Circuit ObfuscationabstractCircuit obfuscation is a frequently used approach to conceal logic functionalities in order to prevent reverse engineering attacks on fabricated chips. Efficient obfuscation implementations are expected with lower design complexity and overhead but higher attack difficulties. In this paper, an emerging obfuscation approach is proposed by leveraging spin-orbit torque (SOT) devices-based look-up-tables as reconfigurable logic to replace the carefully selected gates. It is essentially impossible to identify the obfuscated gate with SOTs inside according to the physical geometry characteristics because the configured functionalities are represented by magnetization states. Such an obfuscation approach makes the circuit security further improved with high exponential attack complexities. Experiments on MCNC and ISCAS 85/89 benchmark suits show that the proposed approach could reduce the area overheads due to obfuscation by 10% averagely. Jianlei Yang 0001, Qiang Zhou 0001, Zhaohao Wang, Hai Li 0001, Yiran Chen 0001, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2019 | DASM: Data-Streaming-Based Computing in Nonvolatile Memory Architecture for Embedded SystemabstractEmerging nonvolatile memories (NVMs), including resistive RAM (RRAM), phase-change memory (PCM), and magnetic RAM (MRAM), have opened up new pathways for Computing-In-Memory (CIM). Those NVM technologies can achieve energy-efficient computational operations with only minor modification of the peripheral circuits. Despite many advantages provided by computational NVMs, parallelism is not sufficiently explored in such CIM designs. To break through this limitation on performance gain, we propose a data-streaming design for the NVM-based CIM (e.g., DASM) by leveraging the underlying parallelism in the hardware. DASM benefits from the massive parallelism of data-streaming computing, reduction in data movement of the CIM, and the nonvolatility of memory arrays. Specifically, data streaming operations can be implemented with CIM bitwise operations in both read-out and write-in procedures. In addition, we use the multilevel power gating for the memory array and connections to further boost the performance. Finally, we study a case of inference process for the quantized deep-neural-network-based on the DASM design. DASM architecture achieves 47.8×, 5.1×, 2.1× speedup compared to the NVIDIA Jetson TK1 embedded GPU board, Intel Xeon E5-2640 CPU, the state-of-the-art field-programmable gate array (FPGA) design, with much lower power consumption. Liang Chang 0002, Zhaohao Wang, Youguang Zhang, Yufei Ding 0001, Weisheng Zhao 0001, Yuan Xie 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2019 | PXNOR-BNN: In/With Spin-Orbit Torque MRAM Preset-XNOR Operation-Based Binary Neural NetworksabstractConvolution neural networks (CNNs) have demonstrated superior capability in computer vision, speech recognition, autonomous driving, and so forth, which are opening up an artificial intelligence (AI) era. However, conventional CNNs require significant matrix computation and memory usage leading to power and memory issues for mobile deployment and embedded chips. On the algorithm side, the emerging binary neural networks (BNNs) promise portable intelligence by replacing the costly massive floating-point compute-andaccumulate operations with lightweight bit-wise XNOR and popcount operations. On the hardware side, the computingin-memory (CIM) architectures developed by the non-volatile memory (NVM) present outstanding performance regarding high speed and good power efficiency. In this paper, we propose an NVM-based CIM architecture employing a Preset-XNOR operation in/with the spin-orbit torque magnetic random access memory (SOT-MRAM) to accelerate the computation of BNNs (PXNOR-BNN). PXNOR-BNN performs the XNOR operation of BNNs inside the computing-buffer array with only slight modifications of the peripheral circuits. Based on the layer evaluation results, PXNOR-BNN can achieve similar performance compared with the read-based SOT-MRAM counterpart. Finally, the end-to-end estimation demonstrates 12.3× speedup compared with the baseline with 96.6-image/s/W throughput efficiency. Liang Chang 0002, Zhaohao Wang, Youguang Zhang, Yuan Xie 0001, Weisheng Zhao 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2019 | An Adaptive Thermal-Aware ECC Scheme for Reliable STT-MRAM LLC DesignabstractConsidering the insatiable demand for high-performance computing, on-chip cache capacity increases rapidly. Spin-transfer-torque magnetoresistive random-access memory (STT-MRAM) is a promising cache candidate due to ultralow standby power, high-access speed, and integration density. Unfortunately, when the feature size of magnetic tunnel junction (MTJ) scales down to 1 Xnm, read current approaches write current closely, which may result in read disturbance threatening the reliability of STT-MRAM. Furthermore, the elevating on-chip temperature reduces the thermal stability of STT-MRAM remarkably and aggravates the read disturbance. Error correction code (ECC) is an effective technique to enhance memory reliability. In this paper, we take advantage of the thermal dependence of STT-MRAM and propose a thermally adaptive ECC design, called “Chameleon,” that can adjust the ECC protection strength dynamically to reduce the ECC storage overhead and improve the cache access performance and energy efficiency. Experimental results show that compared to the conservative nonadaptive ECC scheme, our design can improve both cache performance and energy consumption effectively. Bi Wu 0002, Yuanqing Cheng, Ying Wang 0001, Dijun Liu, Weisheng Zhao 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2018 | Process variation aware data management for magnetic skyrmions racetrack memoryabstractSkyrmions racetrack memory (SKM) has been identified as a promising candidate for future on-chip cache. Similar to many other nanoscale technologies, process variations also adversely impact the reliability and performance of SKM cache. In this work, we propose the first holistic solution for employing SKM as last-level caches. We first present a novel SKM cache architecture and a physical-to-logic mapping scheme based on our comprehensive analysis on working mechanism of SKM. We then model the impact of process variations on SKM cache performance. By leveraging the developed model, we propose a process variation aware data management technique to minimize the performance degradation of SKM cache incurred by process variations. Experimental results show that the proposed SKM cache can achieve a geometric mean of 1.28x IPC improvement, 2x density increase, and 23% energy reduction compared to Domain Wall racetrack memory (DWM) under the same area constraint across 15 workloads. In addition, our dynamic data management technique can further improve the system IPC by 25% w.r.t. the worst-case design. Fan Chen 0001, Wang Kang 0001, Weisheng Zhao 0001, Hai Li 0001, Yiran Chen 0001 |
ASP-DAC | 4 |
| 2018 | Spintronics based stochastic computing for efficient Bayesian inference systemabstractBayesian inference is an effective approach for solving statistical learning problems especially with uncertainty and incompleteness. However, inference efficiencies are physically limited by the bottlenecks of conventional computing platforms. In this paper, an emerging Bayesian inference system is proposed by exploiting spintronics based stochastic computing. A stochastic bitstream generator is realized as the kernel components by leveraging the inherent randomness of spintronics devices. The proposed system is evaluated by typical applications of data fusion and Bayesian belief networks. Simulation results indicate that the proposed approach could achieve significant improvement on inference efficiencies in terms of power consumption and inference speed. Xiaotao Jia, Jianlei Yang 0001, Zhaohao Wang, Yiran Chen 0001, Hai Li 0001, Weisheng Zhao 0001 |
ASP-DAC | 6 |
| 2018 | Magnetic skyrmions for future potential memory and logic applications: Alternative information carriersabstractMagnetic skyrmions are swirling topological configurations, which are mostly induced by chiral interactions between atomic spins in non-centrosymmetric magnetic bulks or in thin films with broken inversion symmetry. They hold promise as information carriers in future ultra-dense, low-power memory and logic devices owing to the nanocale size and extremely low spin-polarized currents needed to move them. To date, an intense research effort has led to the identification, creation/annihilation, motion and manipulation of skyrmions at room temperature. Meanwhile, a rich variety of skyrmion-based device concepts and prototypes have been proposed, indicating the considerable potential of magnetic skyrmions in future electronic applications. However, current studies mainly focus on physical or principle investigations, whereas the electrical design methodology, implementation and evaluations are still lacking. In this paper, we will bring the readers in the “design, automation and test (DAT) society” the current status and outlook of skyrmions in relation to future potential racetrack memory and neuromorphic computing applications. Most importantly, we also want to evoke the effort from the DAT society to address the challenges, e.g., all-electrical manipulation of skyrmions at room temperature, for the research and development of practical skyrmion-based electronics. Wang Kang 0001, Xing Chen 0012, Daoqian Zhu, Yangqi Huang, Youguang Zhang, Weisheng Zhao 0001 |
DATE | 7 |
| 2018 | A Scalable Pipelined Dataflow Accelerator for Object Region Proposals on FPGA PlatformabstractRegion proposal is critical for object detection while it usually poses a bottleneck in improving the computation efficiency on traditional control-flow architectures. We have observed region proposal tasks are potentially suitable for performing pipelined parallelism by exploiting dataflow driven acceleration. In this paper, a scalable pipelined dataflow accelerator is proposed for efficient region proposals on FPGA platform. The accelerator processes image data by a streaming manner with three sequential stages: resizing, kernel computing and sorting. First, Ping-Pong cache strategy is adopted for rotation loading in resize module to guarantee continuous output streaming. Then, a multiple pipelines architecture with tiered memory is utilized in kernel computing module to complete the main computation tasks. Finally, a bubble-pushing heap sort method is exploited in sorting module to find the top-k largest candidates efficiently. Our design is implemented with high level synthesis on FPGA platforms, and experimental results on VOC2007 datasets show that it could achieve about 3.67X speedups than traditional desktop CPU platform and >250X energy efficiency improvement than embedded ARM platform. Wenzhi Fu, Jianlei Yang 0001, Pengcheng Dai, Yiran Chen 0001, Weisheng Zhao 0001 |
FPT | 5 |
| 2018 | Design Space Exploration of Magnetic Tunnel Junction based Stochastic Computing in Deep LearningabstractMagnetic tunnel junction (MTJ) is considered as a promising memory candidate in the more than Moore era because of high power efficiency, fast access speed, nearly infinite endurance and easy 3D integration. The nondeterministic switching behavior has been profited to exploit new directions for computing methods, such as stochastic computing. In this paper, the application of stochastic switching behavior in stochastic computing is explored for deep neural network (DNN). Stochastic computing method features low logic complexity, low energy consumption and fine-grained parallelism, boosting the performance of DNN system by combining MTJ. As a key block of stochastic computing, MTJ based true random number generator design is presented in details. The functionality has been validated by combining the hardware design and post-processing in software. Simulation results are demonstrated visibly by handwritten digits recognition test to show the accuracy. Furthermore, the performance is investigated in terms of accuracy, energy consumption and memory occupation to find more efficient techniques. You Wang 0002, Yue Zhang 0010, Youguang Zhang, Weisheng Zhao 0001, Hao Cai 0001, Lirida A. B. Naviner |
ACM Great Lakes Symposium on VLSI | 4 |
| 2018 | Enabling Resilient Voltage-Controlled MeRAM Using Write Assist TechniquesabstractReliability concerns arise in nonvolatile magnetoelectric random access memory (MeRAM) due to continuously nanotechnology scaling down and CMOS-magnetic hybrid integration. The primary objective of this work is to investigate failure mitigation in voltage-controlled magnetic anisotropy-magnetic tunnel junction (VCMA-MTJ) based 1T-1MTJ MeRAM bit-cell, by using MTJ compact model and 28nm fully depleted silicon on insulator (FD-SOI) process design-kit. A comprehensive reliability study is performed considering process variation and aging degradations, including hot carrier injection (HCI), bias temperature instability (BTI), soft breakdown (SBD) and radiation effect. Write assist techniques are proposed to ensure failure resilient MeRAM design. Bit line (BL) boost and negative source line (SL) methods show high efficiency in writing latency improvement and failure mitigation. Hao Cai 0001, You Wang 0002, Wang Kang 0001, Lirida A. B. Naviner, Weiwei Shan, Jun Yang 0006, Weisheng Zhao 0001 |
ISCAS | 7 |
| 2018 | Design and Data Management for Magnetic Racetrack MemoryabstractBenefiting from its ultra-high storage density, high energy efficiency, and non-volatility, racetrack memory demonstrates great potential in replacing conventional SRAM as large on-chip memory. Integrating the tape-like racetrack memory, however, faces unique design challenges from cell structure to architecture design. This paper reviews some cross-layer design methodologies for racetrack memory as on-chip cache hierarchy. Research studies show that with proper architectural design and data management, racetrack memory can achieve significant area reduction, system performance enhancement, and energy saving compared to state-of-the-art memory technologies. Bing Li 0017, Fan Chen 0001, Wang Kang 0001, Weisheng Zhao 0001, Yiran Chen 0001, Hai Li 0001 |
ISCAS | 4 |
| 2018 | NEAR: A Novel Energy Aware Replacement Policy for STT-MRAM LLCsabstractAs the technology node shrinks, leakage power becomes a bottleneck for processor performance and memory capacity scalings. Spin Torque Transfer Magnetic Random Access Memory (STT-MRAM) has negligible leakage power, fast access speed, high integration density and non-volatility. Therefore, it is a promising candidate for the last level cache design. However, it suffers from high write energy and slow write speed. In the paper, we observe that the traditional cache replacement policy is not optimal when applied to STT-MRAM from the energy consumption perspective. So we propose a novel write energy aware cache replacement policy, which utilizes a MinHash function to identify the similarities between the cache line to be written back and candidates for the replacement. The cache line with the highest similarity is chosen as the victim. In addition, we propose a new metric for cache replacement considering both performance and write energy to improve the replacement policy further. The experimental results show that our proposed policy can reduce write energy by 33.6% on average compared to the state-of-the-art Least Recently Used (LRU) replacement policy with only 0.5% performance penalty and negligible hardware overhead. Yuanqing Cheng, Ying Wang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 5 |
| 2018 | Progresses and challenges of spin orbit torque driven magnetization switching and application (Invited)abstractSpin orbit torque (SOT) has been proposed as a potential alternative mechanism to the conventional spin transfer torque (STT) for the magnetization switching. Recently, theoretical and experimental works revealed the novel factors influencing the SOT-driven magnetization switching. Emerging SOT-based spintronics memories and circuits were explored to implement fast and energy-efficient write operation. However, the perspective of the SOT mechanism is still challenged by some serious shortcomings, such as area penalty, relatively large switching current density and undesirable use of external magnetic field. Here, we review the progresses in the SOT mechanism involving the magnetization dynamics, device design and circuit development. Key issues to be addressed in optimizing the SOT devices are pointed out. In particular, we discuss the potential solutions to develop high-density SOT-based memories and circuits. Zhaohao Wang, Zuwei Li, Liang Chang 0002, Wang Kang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 8 |
| 2018 | Radiation hardening design for spin-orbit torque magnetic random access memoryabstractAlthough the magnetic tunnel junction (MTJ) is intrinsically immune to radiation, the read/write operations of magnetic random access memory (MRAM) may be vulnerable to radiation-induced current. In this paper, we investigate the radiation hardening design for spin orbit torque based MRAM (SOT-MRAM). The hardening technique is firstly studied at the device level by optimizing the dimension and magnetic parameters. Then we propose radiation hardening read and write circuits addressing the influence of single event upset (SEU). Based on a physics-based SOT-MTJ compact model and a 65nm CMOS design kit, simulation results show that the proposed MOS-stacked read sensing amplifier and write circuits of six PMOS transistors as a feed-back structure to charge/discharge sensitive nodes can correct soft errors. Bi Wang 0002, Zhaohao Wang, Kaihua Cao, Youguang Zhang, Yuanfu Zhao, Weisheng Zhao 0001 |
ISCAS | 6 |
| 2018 | Multi-bit nonvolatile flip-flop based on NAND-like spin transfer torque MRAMabstractNonvolatile flip-flops (NVFFs) integrating emerging spintronics devices such as magnetic tunnel junction (MTJ) are under intensive investigation. They allow computing systems to be powered-off during the standby state, hence high static power issue of conventional CMOS technology can be addressed. MTJ based on spin transfer torque (STT) effect provide non-volatility, good endurance and 3D integration with CMOS based circuits. However, it suffers from relative long switching delay, high switching power and asymmetric switching issues. In this work, we first present a multi-bit NVFF using NAND-like spintronics (NANS-SPIN) devices which are written by STT and spin orbit torque (SOT) currents. It shows advantages in terms of power consumption, area overhead and write voltage. Then, functionality and performance of the proposed NVFF will be simulated and validated. Erya Deng, Zhaohao Wang, Wang Kang 0001, Shaoqian Wei, Weisheng Zhao 0001 |
VLSI-SoC | 5 |
| 2018 | Power Supply Noise Aware Task Scheduling on Homogeneous 3D MPSoCs Considering the Thermal Constraint
Yinglin Zhao, Jianlei Yang 0001, Weisheng Zhao 0001, Aida Todri, Yuanqing Cheng |
J. Comput. Sci. Technol. | 3 |
| 2017 | Voltage-controlled MRAM for working memory: Perspectives and challengesabstractMagnetic random access memory (MRAM) has been widely studied for future nonvolatile working memory candidate. However, the mainstream current (spin transfer torque, STT or spin Hall effect, SHE) driven MRAMs (STT-MRAM or SHE-MRAM) face intrinsic problems in terms of high write power and long latency, significantly limiting the applications for low-power and high-speed working memories. The recently-developed new-generation MRAM, named VCMA-MRAM, which exploits the voltage-controlled magnetic anisotropy (VCMA) effect to write (or assist to write) data information into magnetic tunnel junctions (MTJs), holds the promise to efficiently overcome these problems. Despite the impressive possibility of improving write power and speed, this technology, however, is currently under intensive research and development (R&D), and some challenges still await answers. In this paper, we investigate the perspectives and challenges of VCMA-MRAM for working memories from a cross-layer (device/circuit/architecture) design point of view. We demonstrate that VCMA-MRAM outperforms STT-MRAM and SHE-MRAM in terms of area, speed, energy consumption and instruction-per-cycle (IPC) performance, benefiting from the low-power and high-speed VCMA-driven data writing mechanism. On the other hand, challenges in terms of device fabrication and circuit design should be efficiently addressed before practical applications. Wang Kang 0001, Liang Chang 0002, Youguang Zhang, Weisheng Zhao 0001 |
DATE | 4 |
| 2017 | A true random number generator based on parallel STT-MTJsabstractRandom number generators are an essential part of cryptographic systems. For the highest level of security, true random number generators (TRNG) are needed instead of pseudorandom number generators. In this paper, the stochastic behavior of the spin transfer torque magnetic tunnel junction (STT-MTJ) is utilized to produce a TRNG design. A parallel structure with multiple MTJs is proposed that minimizes device variation effects. The design is validated in a 28-nm CMOS process with Monte Carlo simulation using a compact model of the MTJ. The National Institute of Standards and Technology (NIST) statistical test suite is used to verify the randomness quality when generating encryption keys for the Transport Layer Security or Secure Sockets Layer (TLS/SSL) cryptographic protocol. This design has a generation speed of 177.8 Mbit/s, and an energy of 0.64 pJ is consumed to set up the state in one MTJ. Yuanzhuo Qu, Jie Han 0001, Bruce F. Cockburn, Witold Pedrycz, Yue Zhang 0010, Weisheng Zhao 0001 |
DATE | 6 |
| 2017 | Energy Efficient Magnetic Tunnel Junction Based Hybrid LSI Using Multi-Threshold UTBB-FD-SOI DeviceabstractThe energy scalability of ultra-low power nonvolatile (NV) large-scale integration (LSI) is explored in this paper. Multi-threshold computing (super/near/sub-$V_t$) in hybrid CMOS/ magnetic tunnel junction (MTJ) circuits are investigated based on SPICE-compatible MTJ model and fully depleted silicon on insulator (FD-SOI) devices. Ultra-low supply voltage operation bottlenecks associated with performance loss, parametric variations and function failure are studied in differential pair-based sensing circuit, MTJ writing/control circuit and other building blocks. A case study is performed with three typical NV-flip-flops (NV-FF), which are implemented with 28nm FD-SOI low $V_t$ (LVT) device and forward back-bias. Results show that MTJ writing/control circuit must operate at nominal supply (super-$V_t$) region to guarantee MTJ switching; sensing circuit is configured with near-$V_t$ operation (0.6V) with robustness consideration, whereas other parts could be implemented with near/sub-$V_t$ computing to achieve ultra-low power consumption and energy efficient operations. Hao Cai 0001, You Wang 0002, Lirida A. B. Naviner, Wang Kang 0001, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2017 | Advanced Low Power Spintronic Memories beyond STT-MRAMabstractUntil now, spin transfer torque magnetic random access memory (STT-MRAM) has drawn considerable R&D interest worldwide. A number of companies and universities are currently involved in this promising technology. In 2016, Everspin released the first 256M STT-MRAM chip, indicating the commercialization and application of STT-MRAM. Nevertheless, STT-MRAM still has some intrinsic limitations, such as dynamic write power and speed, compared with CMOS-based memory technologies. Following the technical evolution process from toggle-MRAM to STT-MRAM, the continuous pursuit of high performance, high density, low power and scalability, drives the intensive R&D of new memory technologies. In this paper, we will show the recent progress in advanced spintronic memories beyond STT-MRAM, such as the spin Hall effect (SHE)-driven and voltage-driven MRAMs. These advanced MRAM technologies do have some unique advantages compared with STT-MRAM, but they also suffer from new design and fabrication challenges. In addition, we will present the latest research in emerging spintronic devices, e.g., magnetic skyrmions, which are potential as information carriers in future spintronic memories, e.g., racetrack memory. Wang Kang 0001, Zhaohao Wang, He Zhang 0011, Youguang Zhang, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 6 |
| 2017 | PRESCOTT: Preset-based cross-point architecture for spin-orbit-torque magnetic random access memoryabstractDue to nearly zero leakage power consumption, non-volatile magnetoresistive random access memory (MRAM) is becoming one of the promising candidates for replacing conventional volatile memories (e.g. SRAM and DRAM). In particular, emerging spin-orbit torque (SOT) MRAM is considered to outperform spin-transfer torque (STT) MRAM due to its fast switching, separate read/write paths, and lower energy dissipation. However, the SOT-MRAM technology is still in its infancy; one key design challenge is that the control of SOT-MRAM, which involves three terminals, is more complicated compared with STT-MRAM. In this paper, we propose a novel MRAM write scheme called PRESCOTT1, where the “1” and “0” data values can be written into memory cells through the SOT and STT, respectively. As a result, the write current is unidirectional rather than bi-directional, which addresses the control complexity. Using this unidirectional write scheme, we design a PreSET-based cross-point (CP) MRAM to improve programing speed, write energy dissipation and storage density compared to conventional MRAM. Circuit simulation results demonstrate that our PreSET-based CP MRAM can achieve around 67.14% average write energy reduction and 50.86% improvement in programming speed, compared with CP STT-MRAM. Liang Chang 0002, Zhaohao Wang, Alvin Oliver Glova, Jishen Zhao, Youguang Zhang, Yuan Xie 0001, Weisheng Zhao 0001 |
ICCAD | 7 |
| 2017 | Thermosiphon: A thermal aware NUCA architecture for write energy reduction of the STT-MRAM based LLCsabstractAs the speed gap of the modern processor and the off-chip main memory enlarges, on-chip cache capacity increases to sustain the performance scaling. As a result, the cache power occupies a large portion of the total power budget. STT-MRAM (Spin Transfer Torque Magnetic Memory) is proposed as a promising solution for the low power cache design due to its high integration density and ultra-low leakage. Nevertheless, the high write power and latency of STT-MRAM become new barriers for the commercialization of this emerging technology. In this paper, we investigate the thermal effect on the access performance of STT-MRAM and observe that the temperature can affect the write delay and energy significantly. Then, we explore the NUCA (Non-Uniform Cache Access) design of the CMPs (Chip-Multi-Processors)with STT-MRAM based LLC (Last Level Cache). A thermal aware data migration policy, called “Thermosiphon”, which takes advantage of the thermal property of STT-MRAM, is proposed to reduce the LLC write energy. This policy splits the LLC into different regions based on the thermal distribution and adaptively migrate write intensive data considering the temperature gradient among different thermal regions. Compared to the conventional NUCA design, our proposed design can save 22.5% write energy with negligible hardware overhead. Bi Wu 0002, Yuanqing Cheng, Pengcheng Dai, Jianlei Yang 0001, Youguang Zhang, Dijun Liu, Ying Wang 0001, Weisheng Zhao 0001 |
ICCAD | 8 |
| 2017 | Programmable Stateful In-Memory Computing Paradigm via a Single Resistive DeviceabstractData transfer bandwidth and the related energy consumption has become two of the most critical bottlenecks in conventional von-Newman architecture, owing to the separation of the processor and memory units and the performance mismatch between the two. Realization of the unity of logic computing and data storage in the same die has opened up a promising research direction of in-memory computing (IMC). Meanwhile nonvolatile memory (NVM) based programmable (or reconfigurable) logic architecture has always been a hot topic in the circuit and system societies. To date, lots of interest has been attracted and amazing advance has been made in the two fields, yet none can fully exploit the advantages of both. This paper takes a major step forward by introducing a novel nonvolatile programmable stateful IMC architecture via a single resistive device, which is a completely different design paradigm from previous studies. Each memory cell can perform different Boolean logic functions by dynamically programming the input signals. The computing output result is insitu stored in the memory cell itself and can be readout with a memory-like operation. We will first give a brief review on this filed and then introduce our recent work. We will illustrate how the programmable stateful IMC operations can be implemented via a single resistive device and how the logic computing and data storage can be united within a memory chip. Wang Kang 0001, He Zhang 0011, Youguang Zhang, Weisheng Zhao 0001 |
ICCD | 5 |
| 2017 | Pseudo-Differential Sensing Framework for STT-MRAM: A Cross-Layer PerspectiveabstractWith the rapid increase of leakage currents, non-volatile memories have become competitive candidates in the next-generation computer architecture. Among them, STT-MRAM shows great promise in working memory with high density, high speed and tremendous endurance, etc. However, based on our investigations, the dynamic write power and read reliability are two critical challenges of STT-MRAM. In this work, we propose a synergistic pseudo-differential sensing (PDS) framework that employs device, circuit and architectural techniques to address these challenges. In specific, three design techniques, including cell cluster, asymmetric sensing amplifier and self-error-detection-correction, are proposed to implement the PDS framework. We show that the holistic device-circuit-architecture cross-layer co-design enables STT-MRAM to be utilized in the cache memory, benefiting from the improved density, reliability and energy-efficiency. Our experimental results show that the proposed PDS scheme improves the read margin by ~35.6 percent, reduces the area, read latency, read energy, write latency and write power by ~46.7, ~9.8, ~30.3, ~2.3 and ~31.1 percent respectively, compared with the typical 1T1MTJ cell structure for the cache capacity of 8 MB. In addition, the proposed PDS scheme reduces the dynamic energy by ~32.9 percent and leakage energy by ~830 percent, improves the IPC by ~1.3 percent and miss rate by ~36.9 percent respectively, compared with conventional SRAM based cache. Wang Kang 0001, Liang Chang 0002, Zhaohao Wang, Weifeng Lv, Guangyu Sun 0003, Weisheng Zhao 0001 |
IEEE Trans. Computers | 6 |
| 2016 | PDS: pseudo-differential sensing scheme for STT-MRAMabstractSTT-MRAM has been considered as one of the most promising nonvolatile memory candidates in the next-generation of computer architecture. However, the read reliability and dynamic write power concerns greatly hinder its practical application. In this paper, we propose a synergistic solution, namely pseudo-differential sensing (PDS), to jointly address these two concerns. Three techniques, including cell cluster, asymmetric sensing amplifier (ASA) and self-error-detection-correction (SEDC), are proposed to implement the PDS concept. Our experimental results show that the PDS scheme with the 3T3MTJ cell cluster can reduce the area (~21.7%) and write power (~25.6%) of the differential sensing (DS) scheme while improve the read reliability (read margin, ~35.6%) of the typical sensing (TS) scheme for a 16 Mbit cache. Furthermore, the PDS scheme with the 1T3MTJ cell cluster can outperform both the TS and DS schemes in terms of area (~40.0%, ~66.1%), read latency (~16.6%, ~32.1%), read power (~16.7%, ~37.1%), write latency (~5.4%, 16.3%) and write power (~18.6%, ~43.4%). Wang Kang 0001, Tingting Pang, Bi Wu 0002, Weifeng Lv, Youguang Zhang, Guangyu Sun 0003, Weisheng Zhao 0001 |
DAC | 7 |
| 2016 | Spin wave based synapse and neuron for ultra low power neuromorphic computation systemabstractIn this work, we have proposed that the neural synapses and neurons can be realized by utilizing spin waves (SWs) as information carrier. The SWs is excited by spin torque nano-oscillator (STNO), and detected with several different physical mechanisms: 1) tunneling magnetic-resistance 2) spin pumping and 3) inverse spin hall effect. The proposed SWs based synapses and neurons can be further combined together to form a neuromorphic computation system with crossbar structure. Possible ultra low power consumption and ultra high speed are the advantage of our proposed SWs based synapses and neurons. Lang Zeng, Deming Zhang, Youguang Zhang, Fanghui Gong, Tianqi Gao, Sa Tu, Haiming Yu, Weisheng Zhao 0001 |
ISCAS | 8 |
| 2016 | Quantitative evaluation of reliability and performance for STT-MRAMabstractDue to its non-volatility, high access speed, ultra low power consumption and unlimited writing/reading cycles, STT-MRAM (Spin Transfer Torque Magnetic Random Access Memory) has emerged as the most promising candidate for the next generation universal memory. However, the process of commercialization of STT-MRAM is hampered by its poor reliability. Generally, these reliability issues are caused by the PVT (Process Variations, Voltage, and Temperature) of both MTJ (Magnetic Tunneling Junction) and transistor. Mitigation and alleviating the impacts of the intrinsic properties and PVT on STT-MRAM is a challenging work. This paper discusses the errors occurring in STT-MRAM resulting from its poor reliability, and analyzes the causes of such errors. To obtain a quantitative assessment of PVT impact on STT-MRAM reliability, we investigate three aspects: writing/reading operation error rate, power consumption and access delay of a single cell. This study is carried out on Cadence platform for 45 nm technology node and the PMA (Perpendicular Magnetic Anisotropy) MTJ model used in the investigation comes from SP INLIB. These quantitative information would be helpful for designing reliability enhancing strategies of STT-MRAM. Liuyang Zhang, Aida Todri, Wang Kang 0001, Youguang Zhang, Lionel Torres, Yuanqing Cheng, Weisheng Zhao 0001 |
ISCAS | 7 |
| 2016 | Read disturbance issue and design techniques for nanoscale STT-MRAM
Yi Ran, Wang Kang 0001, Youguang Zhang, Jacques-Olivier Klein, Weisheng Zhao 0001 |
J. Syst. Archit. | 5 |
| 2016 | Skyrmion-Electronics: An Overview and OutlookabstractThe well-known empirical phenomenon known as Moore's Law has held true for the past half century. However, it is beginning to break down, owing to limitations arising from leakage currents caused by the quantum effect. As a result, the search for alternatives or complementary technologies that can aid the downscaling of complementary metal-oxide-semiconductor (CMOS) technology has been accelerated in the field of electronics. Among various potential candidates, spintronic technology has attracted considerable interest and attention, especially for the topological spin textures known as magnetic skyrmions. Magnetic skyrmions are expected to have topologically protected stability and nanoscale size, and require a very low driving current density, therefore they are considered as potential building blocks for future spintronic devices and integrated circuits. Furthermore, recent experimental demonstrations of the control of individual nanometer-scale skyrmions, including their creation, detection, transportation, and manipulation at room temperature, further highlight their potential for future electronic applications. In this paper, we review the current status and outlook of skyrmions from the viewpoint of electronic applications. First, the fundamental and elementary functionality of skyrmions, such as electric write-in, read-out, transmission, and manipulation, are introduced. Then, potential electronic applications of skyrmions for nonvolatile memory and logic circuits are described with case studies. Finally, we conclude with an analysis of current challenges, limitations, and future trends of skyrmion research. Wang Kang 0001, Yangqi Huang, Xichao Zhang, Weisheng Zhao 0001 |
Proc. IEEE | 5 |
| 2016 | Radiation-Induced Soft Error Analysis of STT-MRAM: A Device to Circuit ApproachabstractSpin-transfer torque magnetic random access memory (STT-MRAM) is a promising emerging memory technology due to its various advantageous features such as scalability, nonvolatility, density, endurance, and fast speed. However, the reliability of STT-MRAM is severely impacted by environmental disturbances because radiation strike on the access transistor could introduce potential write and read failures for 1T1MTJ cells. In this paper, a comprehensive approach is proposed to evaluate the radiation-induced soft errors spanning from device modeling to circuit level analysis. The simulation based on 3-D metal-oxide-semiconductor transistor modeling is first performed to capture the radiation-induced transient current pulse. Then a compact switching model of magnetic tunneling junction (MTJ) is developed to analyze the various mechanisms of STT-MRAM write failures. The probability of failure of 1T1MTJ is characterized and built as look-up-tables. This approach enables designers to consider the effect of different factors such as radiation strength, write current magnitude and duration time on soft error rate of STT-MRAM memory arrays. Meanwhile, comprehensive write and sense circuits are evaluated for bit error rate analysis under random radiation effects and transistors process variation, which is critical for performance optimization of practical STT-MRAM read and sense circuits. Jianlei Yang 0001, Peiyuan Wang, Yaojun Zhang, Yuanqing Cheng, Weisheng Zhao 0001, Yiran Chen 0001, Hai Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2016 | Temperature Impact Analysis and Access Reliability Enhancement for 1T1MTJ STT-RAMabstractSpin-transfer torque magnetic random access memory (STT-RAM) is a promising and emerging technology due to its many advantageous features such as scalability, nonvolatility, density, endurance, and fast access speed. However, the operation of STT-RAM is severely affected by environmental factors such as process variations and temperature. As the temperature rockets up in modern computing systems, it is highly desirable to understand thermal impact on STT-RAM operations and reliability. In this paper, a thermal-aware MTJ model, calibrated and validated by experimental measurements, is proposed as the basis for thoroughly thermal aware analysis of a 1T1MTJ STT-RAM cell structure. Using this model, we investigate temperature effect on memory cell access behavior in terms of access latency, energy, and reliability on a 45-nm technology node. Thermal impact on a more advanced 11-nm technology node is also evaluated in the paper. Additionally, we propose a thermal-aware design for STT-RAM sensing circuit using a body-biasing technique, which can enlarge read margin dramatically to enhance read reliability under temperature variations. Moreover, our proposed technique can suppress read disturbance effectively as well. Experimental results show that our proposed sensing circuit can enlarge read margin by 2.47× when reading “0” and 3.15× when reading “1,” and reduce read disturbance error rate by 55.6% on average. Bi Wu 0002, Yuanqing Cheng, Jianlei Yang 0001, Aida Todri, Weisheng Zhao 0001 |
IEEE Trans. Reliab. | 5 |
| 2016 | Alleviating Through-Silicon-Via Electromigration for 3-D Integrated Circuits Taking Advantage of Self-Healing EffectabstractThree-dimensional integration is considered to be a promising technology to tackle the global interconnect scaling problem for terascale integrated circuits (ICs). Three-dimensional ICs typically employ through-silicon-vias (TSVs) to vertically connect planar circuits. Due to its immature fabrication process, several defects, such as void, misalignment, and dust contamination, may be introduced. These defects can significantly increase current densities within TSVs and cause severe electromigration (EM) effects, which can degrade the reliability of 3-D ICs considerably. In this paper, we propose an effective framework to mitigate EM effect of the defective TSV. At first, we analyze various possible TSV defects and their impacts on EM reliability. Based on the observation that EM can be significantly alleviated by self-healing effect, we design an EM mitigation module to protect defective TSVs from EM. To guarantee EM mitigation efficiency, we propose two defective TSV protection schemes, i.e., neighbor sharing and global sharing. Experimental results show that the global-sharing scheme performs the best and can improve the EM mean time to failure by more than 70× on average with only 0.7% area overhead and less than 0.5% performance degradation compared with naked design without any EM protection. Yuanqing Cheng, Aida Todri, Jianlei Yang 0001, Weisheng Zhao 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2015 | Quantitative modeling of racetrack memory, a tradeoff among area, performance, and powerabstractRecently, an emerging non-volatile memory called Racetrack Memory (RM) becomes promising to satisfy the requirement of increasing on-chip memory capacity. RM can achieve ultra-high storage density by integrating many bits in a tape-like racetrack, and also provide comparable read/write speed with SRAM. However, the lack of circuit-level modeling has limited the design exploration of RM, especially in the system-level. To overcome this limitation, we develop an RM circuit-level model, with careful study of device configurations and circuit layouts. This model introduces Macro Unit (MU) as the building block of RM, and analyzes the interaction of its attributes. Moreover, we integrate the model into NVsim to enable the automatic exploration of its huge design space. Our case study of RM cache demonstrates significant variance under different optimization targets, in respect of area, performance, and energy. In addition, we show that the cross-layer optimization is critical for adoption of RM as on-chip memory. Chao Zhang 0007, Guangyu Sun 0003, Fan Mi, Hai Li 0001, Weisheng Zhao 0001 |
ASP-DAC | 6 |
| 2015 | Spintronic devices as key elements for energy-efficient neuroinspired architectures
Nicolas Locatelli, Adrien F. Vincent, Alice Mizrahi, Joseph S. Friedman, Damir Vodenicarevic, Joo-Von Kim, Jacques-Olivier Klein, Weisheng Zhao 0001, Julie Grollier, Damien Querlioz |
DATE | 8 |
| 2015 | From device to system: cross-layer design exploration of racetrack memory
Guangyu Sun 0003, Chao Zhang 0007, Hehe Li, Yue Zhang 0010, Yizi Gu, Jacques-Olivier Klein, Dafine Ravelosona, Yongpan Liu, Weisheng Zhao 0001, Huazhong Yang |
DATE | 11 |
| 2015 | A High-Speed Robust NVM-TCAM Design Using Body Bias FeedbackabstractAs manufacture process scales down rapidly, the design of ternary content-addressable memory (TCAM) requiring high storage density, fast access speed and low power consumption becomes very challenging. In recent years, many novel TCAM designs have been inspired by the research on emerging nonvolatile memory technologies, such as magnetic tunneling junction (MTJ), phase change memory (PCM), and memristor. These designs store a data as the resistive variable of a nonvolatile device, which usually results in limited sensing margin and therefore constrains the searching speed of TCAM architecture severely. To further enhance the performance and robustness of TCAMs, we proposed two novel cell designs that utilize MTJs as data storage units - the symmetrical dual-N structure and the asymmetrical P-N scheme. In both designs, a body bias feedback circuit is integrated to enlarge the sensing margins. Compared with an existing MTJ-based TCAM structure, the tolerance in gate voltage variation of the symmetrical dua-N (asymmetrical P-N) scheme can significantly improve 59.5% (21.2%). The latency and the dynamic energy consumption in one searching operation at the word length of 256 bits are merely 590.35ps (97.89ps) and 65.05fJ/bit (36.85fJ/bit), not even mentioning that the use of nonvolatile MTJ devices avoids unnecessary leakage power consumption. Bonan Yan, Yaojun Zhang, Jianlei Yang 0001, Hai Li 0001, Weisheng Zhao 0001, Pierre Chor-Fung Chia |
ACM Great Lakes Symposium on VLSI | 6 |
| 2015 | Hi-fi playback: tolerating position errors in shift operations of racetrack memoryabstractRacetrack memory is an emerging non-volatile memory based on spintronic domain wall technology. It can achieve ultra-high storage density. Also, its read/write speed is comparable to that of SRAM. Due to the tape-like structure of its storage cell, a "shift" operation is introduced to access racetrack memory. Thus, prior research mainly focused on minimizing shift latency/energy of racetrack memory while leveraging its ultra-high storage density. Yet the reliability issue of a shift operation, however, is not well addressed. In fact, racetrack memory suffers from unsuccessful shift due to domain misalignment. Such a problem is called "position error" in this work. It can significantly reduce mean-time-to-failure (MTTF) of racetrack memory to an intolerable level. Even worse, conventional error correction codes (ECCs), which are designed for "bit errors", cannot protect racetrack memory from the position errors. Chao Zhang 0007, Guangyu Sun 0003, Xian Zhang 0001, Weisheng Zhao 0001, Tao Wang 0004, Yun Liang 0001, Yongpan Liu, Yu Wang 0002, Jiwu Shu |
ISCA | 5 |
| 2015 | A new self-reference sensing scheme for TLC MRAMabstractDensity is one of the major design factors of magnetic random access memory (MRAM). Very recently, a tri-level cell (TLC) structure was proposed to enhance the storage density of MRAM. In this work, we propose a new self-reference sensing scheme for the TLC MRAM cell based on its unique property called state ordering. Simulation results show that compared to conventional design, our proposed self-reference scheme achieves on average 61% saving on sensing delay while also demonstrating significantly enhanced resilience to device parametric variations. Bonan Yan, Lun Yang, Weisheng Zhao 0001, Yiran Chen 0001, Hai Li 0001 |
ISCAS | 4 |
| 2015 | Vortex-based spin transfer oscillator compact model for IC designabstractSpintronic oscillators are nanodevices that are serious candidates for CMOS integration due to their compactness and easy frequency tunability. Among them vortex-based oscillators appear as one of the most promising technology because of their lower power supply and higher quality factors. To assess their potential in circuits and systems, compact models describing their behavior are necessary. In this work, we propose an implementation of a spintronic nano-oscillator (STNO) model for integrated circuit (IC) architectures design. The modeled device is a vortex-based magnetic oscillator demonstrating self-sustained magnetization oscillations under current bias, inducing alternating voltage across the device. This model describes the coupled electrical and magnetic behavior of the device, taking into account phase and amplitude noises associated with thermal fluctuations. Compatibility with commercial CMOS design kits is demonstrated, and an implementation in a CMOS circuit is proposed for AC signal generation. These results will allow to develop and evaluate innovative hybrid STNO/CMOS systems and their potential to efficiently complement existing full-CMOS technologies. Nicolas Locatelli, Damir Vodenicarevic, Weisheng Zhao 0001, Jacques-Olivier Klein, Julie Grollier, Damien Querlioz |
ISCAS | 3 |
| 2015 | A body-biasing of readout circuit for STT-RAM with improved thermal reliabilityabstractAs the integration density rockets up for contemporary VLSI circuits, power consumption limits the scalability of technology advancement of CMOS. Spin transfer torque-magnetic random access memory (STT-MRAM), as one of the emerging non-CMOS technologies, has the promising prospect of low standby power, fast access speed and compatibility with the CMOS fabrication process. However, with the technology node scaling down, typical 1 Transistor-1 Magnetic Tunnel Junction (1T-1MTJ) STT-RAM cell suffers from severe reliability challenges, especially for read operation under temperature fluctuation. In this paper, we quantitatively analyze the temperature effect on read reliability of STT-RAM cell and propose a novel body-biasing feedback readout circuit design to improve the read sensing margin under different temperatures. The experiments based on 40nm CMOS technology and MTJ compact model validate the effectiveness of the proposed method. The improved sensing margin also permits a smaller sensing current for reading such that higher read energy efficiency can be achieved. Lun Yang, Yuanqing Cheng, Yuhao Wang 0002, Hao Yu 0001, Weisheng Zhao 0001, Aida Todri |
ISCAS | 5 |
| 2015 | Perspectives of racetrack memory based on current-induced domain wall motion: From device to systemabstractCurrent-induced domain wall motion (CIDWM) is regarded as a promising way towards achieving emerging high-density, high-speed and low-power non-volatile devices. Racetrack memory is an attractive concept based on this phenomenon, which can store and transfer a series of data along a magnetic nanowire. Although the first prototype has been successfully fabricated, its advancement is relatively arduous caused by certain technique and material limitations. Particularly, the storage capacity issue is one of the most serious bottlenecks hindering its application for practical systems. In this paper, we present two alternative solutions to improve the capacity of racetrack memory: magnetic field assistance and chiral domain wall (DW) motion. The former one can lower the current density for DW shifting; the latter one can utilize materials with low resistivity. Both of them are able to increase the nanowire length and allow higher feasibility of large-capacity racetrack memory. Furthermore, system level simulation shows that a racetrack memory based cache can improve system performance by about 15.8% and significantly reduces the energy consumption, compared to the SRAM counterpart. Yue Zhang 0010, Chao Zhang 0007, Jacques-Olivier Klein, Dafine Ravelosona, Guangyu Sun 0003, Weisheng Zhao 0001 |
ISCAS | 6 |
| 2015 | Energy-efficient neuromorphic computation based on compound spin synapse with stochastic learningabstractRecently, magnetic tunnel junction with in-plane magnetization (i-MTJ) has been exploited to behave as a binary stochastic synapse. However, it suffers from its limited level of synaptic weight, resulting in an inaccurate learning. In this work, a compound synapse that employs multiple perpendicular MTJs (p-MTJs) in series is proposed. It possesses an analog-like synaptic weight under weak programming conditions, which leads to a stochastic learning rule and low power consumption per synaptic event. By performing system-level simulations on the MNIST database, it has been demonstrated that such compound spin synapses can realize stochastic neuromorphic computation with high accuracy and low energy consumption. Deming Zhang, Lang Zeng, Yuanzhuo Qu, Youguang Zhang, Mengxing Wang 0001, Weisheng Zhao 0001, Tianqi Tang 0001, Yu Wang 0002 |
ISCAS | 6 |
| 2015 | On-Chip Universal Supervised Learning Methods for Neuro-Inspired Block of Memristive NanodevicesabstractScaling down beyond CMOS transistors requires the combination of new computing paradigms and novel devices. In this context, neuromorphic architecture is developed to achieve robust and ultra-low power computing systems. Memristive nanodevices are often associated with this architecture to implement efficiently synapses for ultra-high density. In this article, we investigate the design of a neuro-inspired logic block (NLB) dedicated to on-chip function learning and propose learning strategy. It is composed of an array of memristive nanodevices as synapses associated to neuronal circuits. Supervised learning methods are proposed for different type of memristive nanodevices and simulations are performed to demonstrate the ability to learn logic functions with memristive nanodevices. Benefiting from a compact implementation of neuron circuits and the optimization of learning process, this architecture requires small number of nanodevices and moderate power consumption. Djaafar Chabi, Weisheng Zhao 0001, Damien Querlioz, Jacques-Olivier Klein |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2015 | Spintronics: Emerging Ultra-Low-Power Circuits and Systems beyond MOS TechnologyabstractConventional MOS integrated circuits and systems suffer serve power and scalability challenges as technology nodes scale into ultra-deep-micron technology nodes (e.g., below 40nm). Both static and dynamic power dissipations are increasing, caused mainly by the intrinsic leakage currents and large data traffic. Alternative approaches beyond charge-only-based electronics, and in particular, spin-based devices, show promising potential to overcome these issues by adding the spin freedom of electrons to electronic circuits. Spintronics provides data non-volatility, fast data access, and low-power operation, and has now become a hot topic in both academia and industry for achieving ultra-low-power circuits and systems. The ITRS report on emerging research devices identified themagnetic tunnel junction(MTJ) nanopillar (one of the Spintronics nanodevices) as one of the most promising technologies to be part of future micro-electronic circuits. In this review we will give an overview of the status and prospects of spin-based devices and circuits that are currently under intense investigation and development across the world, and address particularly their merits and challenges for practical applications. We will also show that, with a rapid development of Spintronics, some novel computing architectures and paradigms beyond classic Von-Neumann architecture have recently been emerging for next-generation ultra-low-power circuits and systems. Wang Kang 0001, Yue Zhang 0010, Zhaohao Wang, Jacques-Olivier Klein, Claude Chappert, Dafine Ravelosona, Gefei Wang, Youguang Zhang, Weisheng Zhao 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 9 |
| 2014 | An overview of spin-based integrated circuitsabstractConventional CMOS integrated circuits suffer from serve power and scalability challenges as technology node scales into ultra-deep-micron technology nodes. Alternative approaches beyond charge-only based circuits. In particular, spin-based devices or integrated circuits show promising merits to overcome these issues by adding the spin freedom of electrons to the electronic circuits. Spintronics has now become a hot topic in both academics and industrials. This paper overviews the status and prospects of spin-based integrated circuits under intense investigation and address particularly their merits and challenges for practical applications. Wang Kang 0001, Weisheng Zhao 0001, Zhaohao Wang, Jacques-Olivier Klein, Yue Zhang 0010, Djaafar Chabi, Youguang Zhang, Dafine Ravelosona, Claude Chappert |
ASP-DAC | 2 |
| 2014 | Spintronics for low-power computingabstractMicroelectronics has been following Moore's law for almost 40 years. However this trend tends to run out of steam in recent technology nodes. The continuous improvements in the size of the transistors and in the operating frequencies result in serious power consumption, heat dissipation and reliability issues. Spintronics (Nobel Prize of Physics, 2007 awarded to Prof. Fert from Univ. Paris-Sud and Peter Grünberg from Forschungszentrum Jülich) nanodevices can reduce significantly the power, improve the reliability or allow new functionalities. The 2010 ITRS report on emerging research devices identified Magnetic Tunnel Junction (MTJ) nanopillar (the preeminent spintronics nanodevice) as one of the most promising technologies to be part of the future microelectronics circuits. It provides data non-volatility, hardness to radiations, fast data access and low-power operations. Magnetic memories become the most promising candidate for both low power logic computing and the data storage. This tutorial paper presents multi-discipline questions (Device, Circuit, Architecture, System and CAD) related to this topic to share the most recent results and discuss the future challenges. Yue Zhang 0010, Weisheng Zhao 0001, Jacques-Olivier Klein, Wang Kang 0001, Damien Querlioz, Youguang Zhang, Dafine Ravelosona, Claude Chappert |
DATE | 2 |
| 2014 | Ferroelectric tunnel memristor-based neuromorphic network with 1T1R crossbar architectureabstractEmerging ferroelectric tunnel memristors show large OFF/ON resistance ratio (>100) and high operation speed (~10ns), promising to be widely applied in the future synapse-like systems. In this paper we propose a neuromorphic network with ferroelectric tunnel memristor. This network is arranged with classical crossbar topology, in which each crosspoint forms a synapse consisting of a MOS transistor and a memristor. Based on this architecture, we design a spike-timing dependent plasticity (STDP) scheme and a parallel supervised learning circuit. Using a compact model of ferroelectric tunnel memristor and CMOS 40nm design kit, we perform transient simulation to validate the functionality of the proposed STDP and learning circuit. Simulation results show the potential of our neuromorphic network in low power (~100nA or ~1μA) and high speed (μs or ~100ns) computing system. Zhaohao Wang, Weisheng Zhao 0001, Wang Kang 0001, Youguang Zhang, Jacques-Olivier Klein, Claude Chappert |
IJCNN | 2 |
| 2014 | Spin-transfer torque magnetic memory as a stochastic memristive synapseabstractSpin-transfer torque magnetic memory (STT-MRAM) is currently under intense academic and industrial development, since it features nonvolatility, high write and read speed and high endurance. In this work, we show that when used in an original regime, it can additionally act as a stochastic memristive device, appropriate to implement a “synaptic” function. We introduce basic concepts relating to STT-MRAM cell behavior and its possible use to implement learning-capable synapses. System-level simulations on a problem of car counting highlight the potential of the technology for learning systems. Monte Carlo simulations show its robustness to device variations. These results open the way for unexplored applications of STT-MRAM in robust, low power, cognitive-type systems. Adrien F. Vincent, Jerome Larroque, Weisheng Zhao 0001, Nesrine Ben Romdhane, Olivier Bichler, Christian Gamrat, Jacques-Olivier Klein, Sylvie Galdin-Retailleau, Damien Querlioz |
ISCAS | 3 |
| 2014 | Robust learning approach for neuro-inspired nanoscale crossbar architectureabstractScaling beyond CMOS require a new combination of computing paradigm and new devices. In this context, memristor are often considered as best candidate to implement efficiently synapses in hardware neural networks. In this article, we analyze the impact of memristor parameter variability. We build an analytical model of the global reliability at the crossbar level. It is based on a supervised learning method with multilayer and redundancy extensions. Comparisons with Monte Carlo simulations of small neural network validate our analytical model. It can be used to extrapolate directly the reliability of large-scale neural system. Our extrapolations show that high defect rate and important parameter variability can be handle efficiency with a moderate amount of redundancy. Djaafar Chabi, Damien Querlioz, Weisheng Zhao 0001, Jacques-Olivier Klein |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2014 | Design and analysis of crossbar architecture based on complementary resistive switching non-volatile memory cells
Weisheng Zhao 0001, Jean-Michel Portal, Wang Kang 0001, Mathieu Moreau, Yue Zhang 0010, Hassen Aziza, Jacques-Olivier Klein, Zhaohao Wang, Damien Querlioz, Damien Deleruyelle, Marc Bocquet, Dafine Ravelosona, Christophe Muller, Claude Chappert |
J. Parallel Distributed Comput. | 1 |
| 2013 | Spin-electronics based logic fabricsabstractAdvanced computing ICs in ultra deep-micron technology nodes (e.g. 40 nm) suffer from high power issues, which become one of the major bottlenecks for the future performance progress. Both static and dynamic power dissipation are increasing, caused mainly by the intrinsic leakage currents and large data traffic. Alternative approaches beyond charge-based logic circuits become hot research topics to overcome these issues definitively. By integrating the spin freedom of electrons to electronic devices, spin-electronics is promising for ultra-low power computing as it can provide non-volatility, fast data control and high logic density etc. Today, most of large microelectronics industries investigate this emerging field. In this invited paper for the special session “Nanoscale logic fabrics”, we overview spin-electronics based logic fabrics under intense investigation and address particularly the impact of this technology on logic architectures and new computing paradigms. Weisheng Zhao 0001, Jacques-Olivier Klein, Zhaohao Wang, Yue Zhang 0010, Nesrine Ben Romdhane, Damien Querlioz, Dafine Ravelosona, Claude Chappert |
VLSI-SoC | 1 |
| 2012 | MRAM crossbar based configurable logic blockabstractSpintronics-based non-volatile storage devices promise great potential to be integrated in reconfigurable circuits to overcome the major hurdles related to conventional flash and SRAM memories, such as low logic density, high standby power and long (re) boot latency. In this paper, we describe a compact design of configurable logic block based on Magnetic RAM (MRAM) crossbar architecture. The logic density can be increased greatly (~5 times) compared to conventional designs; the standby power can be nearly zero thanks to the non-volatility of MRAM. Its high speed and power efficiency (~10.4 Tera-OPS/Watt in computing mode) are also demonstrated through mixed CMOS/Magnetic spice simulations. Fully dynamic reconfiguration through context switching is also studied, which could be achieved with low area overhead. Yahya Lakys, Weisheng Zhao 0001, Jacques-Olivier Klein, Claude Chappert |
ISCAS | 2 |
| 2012 | Nanodevice-based novel computing paradigms and the neuromorphic approachabstractDeep submicron (<;90nm) Integrated Circuits (IC) suffer from both high static and dynamic power consumption, which are caused respectively by the growing leakage currents and large capacitance bus traffic. Nanodevice based novel computing paradigms are currently under intense investigation to overcome these issues and build up the next generation ICs performing with higher power efficiency and operating performance. In this paper, an overview and current status of this field is first presented, and then we focus on the memristive nanodevices based neuromorphic approach, which is considered as one of the most promising computing paradigms for power reduction and process variation or defect tolerance. Weisheng Zhao 0001, Damien Querlioz, Jacques-Olivier Klein, Djaafar Chabi, Claude Chappert |
ISCAS | 1 |
| 2011 | Magnetic memory (MRAM), a new area for 2D and 3D SoC/SiP designabstractNo abstract available. Lionel Torres, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2011 | Design of MRAM based logic circuits and its applicationsabstractAs the fabrication technology node shrinks down to 90nm or below, high standby power becomes one of the major critical issues for CMOS logic circuits due to the high leakage currents. A number of non-volatile storage technologies such as FRAM, MRAM, PCRAM and RRAM and so on, are under investigation to bring the non-volatility into the logic circuits and then eliminate completely the standby power issue. Thanks to its infinite endurance, high switching/sensing speed and easy 3D integration after CMOS process, MRAM is considered as the most promising one. Numerous logic circuits based on MRAM technology have been proposed and prototyped in the last years. In this paper, we present an overview and current status of these logic circuits and their potential applications in the future. Weisheng Zhao 0001, Lionel Torres, Yoann Guillemenet, Vitorio Cargnini, Yahya Lakys, Jacques-Olivier Klein, Dafine Ravelosona, Gilles Sassatelli, Claude Chappert |
ACM Great Lakes Symposium on VLSI | 1 |
| 2011 | Embedded MRAM for high-speed computingabstractAs the fabrication technology node shrinks down to 90nm or below, high standby power becomes one of the major critical issues for CMOS high-speed computing circuits (e.g. logic and cache memory) due to the high leakage currents. A number of non-volatile storage technologies such as FeRAM, MRAM, PCRAM and RRAM and so on, are under investigation to bring the non-volatility into the logic circuits and then eliminate completely the standby power issue. Thanks to its infinite endurance, high switching/sensing speed and easy 3D integration after CMOS process, MRAM is considered as the most promising one. Numerous logic circuits based on MRAM technology have been proposed and prototyped in the last years. In this paper, we present an overview and current status of these logic circuits and discuss their potential applications in the future from both the physics and architecture points of view. Weisheng Zhao 0001, Yue Zhang 0010, Yahya Lakys, Jacques-Olivier Klein, Daniel Etiemble, D. Revelosona, Claude Chappert, Lionel Torres, Vitorio Cargnini, Raphael Martins Brum, Yoann Guillemenet, Gilles Sassatelli |
VLSI-SoC | 1 |
| 2010 | High Density Asynchronous LUT Based on Non-volatile MRAM TechnologyabstractIn this article, we present the architecture design of high-performance Asynchronous Look Up Table (LUT) embedded with a non-volatile Magnetic RAM (MRAM) as the configuration memory, called MALUT. It promises a number of advantages over the traditional FPGA circuits such as “free” standby power, high operating frequency and instant on/off etc. Thanks to the 3D integration and high density of MRAM, fine-grain run-time reconfiguration and multi-context configuration can be achieved. An automatic design flow has been developed for the design of complex hybrid CMOS/MRAM circuits. Based on CMOS 130nm and MRAM 120nm technology, mixed simulations and layout implementation have been done to demonstrate the expected operation and configuration performances. At last, we discuss and conclude. Sumanta Chaudhuri, Weisheng Zhao 0001, Jacques-Olivier Klein, Claude Chappert, Pascale Mazoyer |
FPL | 2 |
| 2010 | Design of embedded MRAM macros for memory-in-logic applicationsabstractInternational audience Sumanta Chaudhuri, Weisheng Zhao 0001, Jacques-Olivier Klein, Claude Chappert, Pascale Mazoyer |
ACM Great Lakes Symposium on VLSI | 2 |
| 2010 | Carbon nanotube-based programmable devices for adaptive architecturesabstractWe show that optic ally-gated carbon nanotube field effect transistors can be used as 2-terminal like devices with light sensitivity and memory capabilities. In particular, their channel resistivity can be adjusted precisely, within a large range and memorized. These devices can thus be used as synapses in neural network type of circuits. We demonstrate experimentally these properties, build a device model and propose a circuit architecture, which allows very efficient parallel learning. Guillaume Agnus, Arianna Filoramo, Jean-Philippe Bourgoin, Vincent Derycke, Weisheng Zhao 0001 |
ISCAS | 5 |