Hao Cai 0001

dblp:08/3328-1 · DBLP profile ↗
← Back
53ranked-venue papers
7as first author
42since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 49 · 6 first-author · 38 since 2021Software engineering, systems software and programming languages · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Soft-Error Resilient MRAM-OTP BCAM for DDR4 STT-MRAM Redundancy Management
abstract
Memory systems operating in high-radiation environments require robust protection against single-event effects (SEE) damage. While STT-MRAM offers inherent advantages due to its spin-based storage, conventional redundancy repair architectures remain vulnerable due to separated storage and configuration circuits, slow boot performance, and radiation-induced errors in long signal paths. This paper proposes a novel radiation-hardened one-time programmable (OTP) content-addressable memory (CAM) based on magnetic tunnel junctions (MTJs) for efficient column redundancy in STT-MRAM macros. The design incorporates a radiation-hardened-by-design (RHBD) CAM array with built-in self-repair (BISR), featuring complementary OTP MTJ bitcells enabling parallel programming and disturbance-free matching, a soft-error resilient array with dual-node hardened latches and dual match-line sensing, and a DDR4-compatible repair mechanism supporting TMR-Latch-based fast initialization and energy-efficient search with inter-loop termination. The proposed system significantly improves wake-up speed to less than two clock cycles, and reduces power consumption to less than 12fJ, offering a viable solution for 37MeV radiation-tolerant memory systems.
Zhenghan Fang, Hao Cai 0001
DATE5
2026 Non-volatile Spintronic Flip-Flops with Checkpoint Preservation Supported in RISC-V Platform
abstract
Due to ambient energy’s inherent instability, intermittent computing is essential for task completion. This work comprehensively explores the spintronic flip-flop implementation in the open-source RISC-V platform. Magnetic tunnel junction (MTJ) has great potential for non-volatile flip-flop (NV-FF) implementation because of its high density, low read and write energy consumption, and compatibility with CMOS process. To the best of the authors’ knowledge, the checkpoint preservation is firstly supported in this work. The proposed non-volatile differential sampling latch (NV-DSL) achieves 7.39 fJ/bit data transfer energy consumption. The phased write strategy reduces write energy by 24.3%. A generalized NV-FF design methodology is further established, achieving a 68.88% area reduction. The power consumption of proposed non-volatile RISC-V processor is reduced by nearly 75%. When performing atomic tasks, the energy consumption and latency are reduced by 61.4% and 43.87%, respectively, compared with the cache scheme.
Jiongzhe Su, Mingtao Chen, Zhanpeng Qiu, Bo Liu 0019, Hao Cai 0001
DATE5
2026 Equivalent-0ns-Replacement Self-Aware-Access LLC on Dual-Port SOT-MRAM by Sense-While-Replace
abstract
Last-Level Cache (LLC) is increasingly required to be energy-efficient and area-saving. Emerging Non-volatile memory (NVM), such as Magnetic-resistive Random Access Memory (MRAM), present potential solutions for LLC as its ultra-low leakage and area. However, the high replacement-latency caused by its high write-latency and power consumption hinders MRAM in LLC applications. Thus, this paper proposes a novel Sense-While-Replace (SWR) strategy for dual-port SOT-MRAM, which liberates the conflict between reading and writing to conceal the impact of high write-latency on system performance. Furthermore, Self-aware access circuits are proposed, which accelerate reading and obtain utmost writing-energy saving. Under 40-nm CMOS technology, the 4Kb Macro achieves <3ns@32bits read and <75% energy-saving. Most crucially, SWR supports CPU continuously read whereas preserved from replacement latency, which improves performance by up to 8% even compared to SRAM.
Keyang Zhang, Quanhai Zhu, Zhenghan Fang, Hao Cai 0001
DATE5
2026 Hierarchical Fault Mitigation Architecture in Radiation-Hardened STT-MRAM
Yantong Di, Huanghui Wang, Jingchao Zhang, Bo Liu 0019, Hao Cai 0001
ISCAS9
2026 Hybrid-Domain STT-MRAM Compute-in-Memory Scheme for Quantized Deep Neural Network
Tingxuan Shi, Huanghui Wang, Zeying Ding, Yantong Di, Bo Liu 0019, Hao Cai 0001
ISCAS9
2026 3D-TANoC: Thermal-Aware 3D LLM Accelerator with Hierarchical NoC and Operator-Aware Dataflow Mapping Strategy
Xingyu Xu 0008, Dengke Liu, Zihan Zou, Xilong Kang, Hui Kou, Hao Cai 0001, Bo Liu 0019
ISCAS7
2026 A Highly-Scalable and Full-Connected SOT P-Bit Ising Annealer for Combinatorial Optimization
Huanghui Wang, Yantong Di, Zeying Ding, Jiongzhe Su, Jingchao Zhang, Xin Si, Bo Liu 0019, Hao Cai 0001
ISCAS11
2026 Low Bit-Width LLM Acceleration via Symmetric Lookup Format and Compute-in-Decoding Paradigm
Zihan Zou, Jiaming Lin, Xinming Yan, Shikuang Chen, Chen Zhang 0025, Xilong Kang, Hao Cai 0001, Bo Liu 0019
IEEE Trans. Computers9
2026 Endurance-Oriented STT-MRAM Implementation in Normally-off Microcontroller Unit
abstract
With the rising demand for machine learning applications in edge devices, emerging edge devices are required to achieve higher endurance under wide temperatures for continuous weight updates under training. It becomes critical for the extensive utilization of non-volatile memory (NVM) devices in battery-powered embedded applications to achieve better endurance than$10^{6}$, provided by eFlash. In this paper, we proposed an embedded spin transfer torque magnetic random access memory (STT-MRAM) macro design with a power-insensitive write driver, a near-driver cell protection write scheme, and a temperature self-adaptive power module, which emphasized MRAMs’ high endurance and low power without compromising device-level speed and compactness characteristics. The proposed design managed to operate under a wide temperature range from -40C to 125C while maintaining over$10^{9}$times higher endurance than the traditional approach. With its high endurance, high density, and low leakage power, the eMRAM is later verified with Microcontroller Unit (MCU).
Ting-Xuan Shi, Hao Cai 0001
IEEE Trans. Circuits Syst. I Regul. Pap.5
2026 Elaborated Dual-Path SOT-MRAM Achieving 500-MHz Read and 100-MHz Write for Energy-Constraint Applications
abstract
This brief proposes a novel read-write scheme for energy-efficient SOT-MRAM. First, a dual-path-based smart write scheme is introduced, which utilizes a low-leakage half Schmitt Trigger (LLH-ST) to enable both early judge termination and write complete termination in a dual write path. Based on the pulse-width-dependent critical switching current characteristics of SOT devices, we propose a dynamic gradient-ascent (DGA) write driver to mitigate write energy consumption under device level variations. For write operation, a feedback voltage-controlled negative differential resistance (FV-NDR) structure is applied. It enables voltage latching by comparing and feeding back the stored data, allowing rapid readout during the discharge phase and achieving low power consumption. Based on a 40nm CMOS technology, the proposed dual-path-based smart write scheme with DGA write driver achieves power savings of 86.27% at early judge termination and 67.64% at write complete termination, with no additional timing overhead compared to other write termination schemes. The proposed FV-NDR structure achieves 75.35% read power savings compared to traditional readout, offering a read window of 257.1 mV and enabling 1.9ns readout.
Zhenghan Fang, Hao Cai 0001
IEEE Trans. Circuits Syst. I Regul. Pap.5
2026 An SOT-MRAM-Based Δ-Bias-Σ Computing Paradigm for High Energy Efficiency Computer Vision Applications
abstract
Spin–orbit torque (SOT)-MRAM offers a promising solution for computer vision (CV) applications through its superior throughput and bit-cell energy efficiency. Meanwhile, redundant features lead to dominant energy consumption in memory-access peripherals and analog-to-digital converters (ADCs). Existing delta-sigma in-memory computing ($\Delta \Sigma $IMC) only reduces logic switching power, which has limited energy benefit for non-volatile memory (NVM)-based computing cores. This paper proposes an SOT-MRAM-based$\Delta $-bias-$\Sigma $computing macro to address the critical energy bottlenecks in CV applications. Voltage-delta-sensitive multiplier (VDM) realizes in-SOT-MRAM multiplication of incremental values, while eliminating redundant memory-access power caused by repetitive features. Delta-sensitive analog biasing (DS-biasing) scheme reduces the required ADC dynamic range, while delta-input-adaptive (DA) variable resolution ADC eliminates unnecessary quantization energy from repetitive inputs. Simulations based on 32kb SOT-MRAM computing macro demonstrate that the VDM achieves$6.45\times $energy reduction over conventional MRAM compute unit, while the DS-biasing combined with DA-ADC achieves$1.8\times $A/D conversion energy reduction compared with$\Delta \Sigma $IMC. The CV computing system achieves 546TOPS/W/b on MNIST image classification task with 97.9% accuracy, while improving 34% and 5% energy efficiency in video edge detection task and MobileViT inference task, respectively.
Zhenghan Fang, Bo Liu 0019, Hao Cai 0001
IEEE Trans. Circuits Syst. I Regul. Pap.7
2026 A Versatile One-Time-Programmable STT-MRAM for Security-Aware Scenario
abstract
Recently, one-time-programmable (OTP) memory has been widely used in micro controller unit (MCU) due to its storage reliability and tamper-proof. Based on the analysis of the breakdown mechanism of magnetic tunnel junction and the measurement results used for modeling, this article demonstrates a 96-Kb versatile MRAM-OTP macro integrated with a 6-Kb OTP-based physical unclonable function (PUF) in 55-nm fully depleted silicon on insulator process. The MRAM-OTP macro realize minimum 3.0 V program voltage and 100 ns for 32-bit program speed. The on-chip power supply method trims with PVT variations and completes voltage switching between two modes. The multi-bit programming design saves about 97% time and 5% energy using self-termination. The discernible dual-mode sense amplifier saves 57.4% power in OTP reading with no lacks of sensing yield. We also presented the application scheme of the proposed MRAM-OTP macro within the trimming information storage and security MCU design, with the PUF helped to complete the authentication process together with OTP and the error correcting code assisted encrypting and correcting process. The trimming information can help to enhance the reliability of STT-MRAM main area. With the reconfigurable bit-cell design, the Inter-HD and Intra-HD of the OTP-based PUF can reach 49.77% and 0.08%, respectively.
Jiongzhe Su, Mingtao Chen, Keyang Zhang, Quanhai Zhu, Bo Liu 0019, Hao Cai 0001
IEEE Trans. Reliab.8
2025 SUArch: Accelerating Layer-wise N: M Sparse Pattern with a Unified Architecture for Deep-learning Edge Device
abstract
Deep neural networks are of the essence for user applications on edge devices. However, the computation and memory-intensive nature of deep neural networks conflicts with the resource-constrained devices. Moreover, the heterogeneity across different models imposes new challenges on deployment on edge devices. To boost the capabilities of edge devices, we propose SUArch, which innovates on three fronts: 1) a layer-wise N:M sparsity aware training approach to strike a balance between accuracy and training cost; 2) a sparsity alignment unit based on the butterfly network to maximize hardware utilization and eliminate extra overhead; 3) a mode-heterogenous processing element array to effectively accomplish the unified support for Convolution Neural Network and Transformer. The experimental results demonstrate that when running convolution-based and attention-based models under an industrial 28-nm process, the proposed SUArch realizes an energy efficiency of 52.1 TOPS/W. Compared to state-of-the-art architecture, SUArch achieves an energy efficiency improvement of 2.07× while accuracy loss is within 0.7%.
Xilong Kang, Qingwen Wei, Ningyuan Li 0004, Xingyu Xu 0008, Hao Cai 0001, Bo Liu 0019
ASP-DAC5
2025 AmPEC: Approximate MRAM with Partial Error Correction for Fine-grained Energy-quality Trade-off
abstract
Spin Transfer Torque-Magnetic Random Access Memory (STT-MRAM) is a promising nonvolatile memory technology for future on-chip storage. However, its energy consumption during read and write operations poses a challenge to its overall energy efficiency. To address this, approximate storage techniques are explored to enhance energy saving in STT-MRAM while maintaining a low error probability. This work proposes a fine-grained, quality-tunable approximate STT-MRAM that allows for an energy-quality trade-off. Compared with uniform approach that approximate all the bits to the same quality, our approach leverages bit-level approximate methods and data mapping to minimize quality loss. The reuse of circuits for different modes eliminates area overhead, and the modified read and write schemes are achieved through control signal manipulation. Partial Error Correction Code (ECC) is employed to check the most significant bits (MSBs) and ensure the minimum precision when quality of MSBs cannot be guaranteed. The evaluation demonstrates a 49.5% reduction in energy consumption with negligible area overhead and image quality loss. Partial error correction further enhances image quality without requiring additional column area.
Lan-yang Sun, Yaoru Hou, Hao Cai 0001
ASP-DAC3
2025 TWDP: A Vision Transformer Accelerator with Token-Weight Dual-Pruning Strategy for Edge Device Deployment
abstract
Vision Transformers (ViTs) have attracted significant attention due to their superior accuracy compared to convolutional neural networks (CNNs) in various computer vision tasks. However, their substantial computational load and significant memory footprint lead to excessive delay and considerable data storage overhead, posing challenges for resource-limited edge device deployment. To address these issues, we present TWDP, a vision transformer accelerator employing a Token-Weight Dual-Pruning strategy to enhance the efficiency of the inference process. Firstly, we propose a parameter-free self-adaptive token pruning method to skip redundant computations in an image-dependent manner. Secondly, we apply a Hessian-aware layer-wise N:M weight pruning approach to minimize storage overhead, memory access, and computational power consumption. Additionally, to manage the complex computing patterns in ViTs, an overlapping dataflow is utilized to further reduce temporal storage and inference latency. Implemented and evaluated under an industrial 28nm technology, the proposed TWDP framework reduces 66.1% weight storage requirements and achieves an energy efficiency of 2070.9 FPS/W. Compared to state-of-the-art architectures, TWDP obtains a 1.6× energy efficiency improvement with negligible accuracy loss, demonstrating the superiority of TWDP in edge device deployment scenarios.
Guang Yang 0036, Xinming Yan, Hui Kou, Zihan Zou, Qingwen Wei, Hao Cai 0001, Bo Liu 0019
ASP-DAC6
2025 OutlierCIM: Outlier-Aware Digital CIM-Based LLM Accelerator with Hybrid-Strategy Quantization and Unified FP-INT Computation
abstract
Activation outliers in Large Language Models (LLMs), which exhibit large magnitudes but small quantities, significantly affect model performance and pose challenges for the acceleration of LLMs. To address this bottleneck, researchers have proposed several co-design frameworks with outlier-aware algorithms and dedicated hardware. However, they face challenges balancing model accuracy with hardware efficiency when accelerating LLMs in a low bit-width manner. To this end, we propose OutlierCIM, the first algorithm and hardware codesign framework for the compute-in-memory (CIM) accelerator with outlier-aware quantization algorithm. The key contributions of OutlierCIM are 1) an outlier-clustered tiling strategy that regulates memory access and reduces inefficient workloads which are both introduced by outliers, 2) a hybrid-strategy quantization and a reconfigurable double-bit CIM macro array that overcome the low storage utilization and high latency of outlier-based LLM quantization, and 3) a quantization factor post-processing strategy and a dedicated quantizer that efficiently unify the multiplication and accumulation of outlier-caused FP-INT workloads. Implemented in a 28 nm CMOS technology, OutlierCIM occupies an area of $2.25 \mathrm{~mm}^{2}$. When evaluated at comprehensive benchmarks, OutlierCIM achieves up to $4.54 \times$ energy efficiency improvement and $3.91 \times$ speedup compared to the state-of-the-art outlier-aware accelerators.
Zihan Zou, Shikuang Chen, Chen Zhang 0001, Xin Si, Hao Cai 0001, Bo Liu 0019
DAC8
2025 A Hardware Prototype of an MRAM-based Stochastic Computing System
Zhengkun Yu, Tianhan Fei, Hao Cai 0001, Heng Shi 0003, Siting Liu 0001
ACM Great Lakes Symposium on VLSI3
2025 H3D-LLM: Heterogeneous 3D Chiplet Design for LLM Inference with Dynamic Task Scheduling and Memory-Aware Orchestration
abstract
The exponential growth of Large Language Model (LLM) intensifies hardware demands for energy-efficient, low-latency architectures with scalable memory bandwidth. While 3D chiplet integration addresses conventional systems’ memory wall limitations, three critical challenges persist: asymmetric compression constraints from divergent sparsity-precision requirements across attention/projection layers, tier-level load imbalance from static resource allocation in dynamic computation patterns, and coupling-induced signal degradation in high-density TSV networks, especially under LLM-phase-specific traffic with spatiotemporal burstiness. To address these, we present H3D-LLM, a vertically heterogeneous architecture combining analog/digital Computing-in-Memory (CIM) and Neural Processing Unit (NPU) chiplets through three innovations. First, a Sparse-Aware Dynamic Execution Framework (SADEF) with Precision-Adaptive Quantization Mechanism (PAQM) enables hardware-aware compression via layer-wise unstructured sparsity detection and INT4/8-FP/BF16 mixed precision. Second, a 3D Spatio-Temporal Interleaved Parallelism (3D-STIP) with Semantic-Aware Tiered Storage (SATS) eliminates resource imbalance and improves memory efficiency through dynamic sub-batch partitioning and Key-Value (KV) cache aware management. Third, a Phase-Adaptive TSV Management (PATM) scheme with dynamic encoding and cluster-based allocation enhances inter-connect efficiency and signal integrity through runtime-aware partitioning and phase-specific dataflow scheduling. Evaluations on Llama-7B demonstrate that H3D-LLM achieves 12.3× higher energy efficiency and 8.4× faster inference than the A800 GPU, while its TSV strategy increases eye height by 12% and reduces bit error rate by up to 60× compared to naïve 3D accelerators.
Hui Kou, Chenjie Xia, Liyi Li 0004, Hao Cai 0001, Xin Si, Bo Liu 0019
ICCAD5
2025 S-DMA: Sparse Diffusion Models Acceleration via Spatiality-Aware Prediction and Dimension-Adaptive Dataflow
abstract
Diffusion Models (DMs) have demonstrated remarkable performance in a variety of image generation tasks.However, their complex architectures and intensive computations result in significant overhead and latency, posing challenges for hardware deployment.To address these issues, researchers have explored the sparsity in DMs to reduce computational workloads, including semantic sparsity in image generation and spatial sparsity in local editing.Unfortunately, existing sparsity prediction methods face critical limitations in deployment: 1) additional prediction overheads offset the benefits of sparsity; 2) convolution and general matrix multiplication (GEMM) exhibit distinct sparsity patterns, which current co-design frameworks struggle to process.In this paper, we introduce S-DMA, a software-hardware co-design framework that unifies efficient sparsity prediction while supporting various sparse operators.First, we propose a spatiality-aware similarity computation method that leverages the local similarity of images, reducing the computational complexity of sparsity prediction from O(𝑁 2 ) to O(N ).Second, we implement NAND-based similarity for sparsity prediction, which minimizes the computational overheads and ensures adaptability to different sparsity schemes.Finally, a dedicated hardware architecture is designed to efficiently leverage the algorithm optimizations.A NAND-based sparsity prediction processing unit is designed to adaptively handle the sparsity patterns.Additionally, a sparsity-aware reduction network and a dimension-adaptive Bo Liu is the corresponding author.
Zihan Zou, Xinming Yan, Guang Yang 0036, Hao Cai 0001, Bo Liu 0019
MICRO6
2025 PF²A-ViT: Parameter-Free and Feature-Aware Dynamic Token Pruning Accelerator With Complementary Quantization-Encoding for Vision Transformer
abstract
Vision Transformers (ViTs) have achieved outstanding performance in visual applications. However, ViT’s increasing parameters and computation overhead limit its deployment on hardware. Previous ViT accelerators focus on optimizing the core attention mechanism of ViTs due to the high overheads of language transformer-based neural networks. Nevertheless, linear layers are the actual bottleneck of ViT inference, accounting for larger than 90% FLOPs on numerous ViTs because of the short and fixed token length of ViTs. To this end, we propose PF2A-ViT, an algorithm and accelerator co-design framework, comprehensively accelerating both core attention and linear layers. At the algorithm level, a parameter-free and feature-aware dynamic token pruning (PF2ATP) strategy is proposed to reduce the dimension of feature maps and dynamically adjust the pruning ratio according to the complexity of the feature without complex execution of subnetworks. Meanwhile, a mixed-precision quantization strategy combines Hessian trace and parameter-aware signal-to-quantization-noise ratio to boost the deployment efficiency of PF2ATP. In addition, a cluster-regroup-based weight encoding strategy is proposed to compensate for the bit-wise redundant information of the quantization strategy. At the hardware level, a token pruning module based on bitonic sorters is designed to fully leverage PF2ATP. Simultaneously, a 3D-PE array with reconfigurable 4/8-bit processing elements (PE) is designed to implement the mixed-precision quantization strategy and equipped with greedy bit-wise compensation decoders to exploit the encoding strategy. Extensive experiments on multiple ViTs demonstrate the achievements of PF2A-ViT: (1) Maximally realize 3.89× speedup, 5.54× energy efficiency compared to state-of-the-art ViT accelerators. (2) Occupying a 2.25 mm area and consuming 76 mW power in 28-nm technology.
Zihan Zou, Xinming Yan, Chen Zhang 0025, Shikuang Chen, Guang Yang 0036, Han Yan 0014, Hao Cai 0001, Bo Liu 0019
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2024 FDCA: Fine-grained Digital-CIM based CNN Accelerator with Hybrid Quantization and Weight-Stationary Dataflow
abstract
Digital-Compute-in-memory (DCIM) has demonstrated significant energy and area efficiency in convolutional neural network (CNN) accelerators, particularly for high precision applications. However, to mitigate parasitic effects on word and bit lines, most DCIMs employ fine-grained multiply-accumulate operations, which introduces new challenges and opportunities but has not been widely explored. This paper proposes FDCA: a fine-grained digital-CIM based CNN accelerator with hybrid quantization and weight-stationary dataflow, in which the key contributions are :1) a hybrid quantization approach for CNNs leveraging hessian trace and approximation is utilized. This method incorporates the ratio of computation time and storage time into quantization, achieving high efficiency while maintaining accuracy; 2) a Cartesian Genetic Programming based approximate shift and accumulate with error compensation is proposed, where an approximate adder tree is generated to compensate for errors introduced by DCIM; 3) an optimized weight-stationary dataflow is used to improve the utilization of CIM and eliminate dataflow stalls. The experimental results demonstrate that under 28-nm process, when running VGG16 and ResNet50 on CIFAR100, the proposed FDCA achieves 17.1TOPS/W and 18.79TOPS/W with only a slight decrease in accuracy by 0.71% and 0.98%, respectively. Compared to previous works, this work achieves 1.76× and 1.28× better in energy efficiency with less accuracy loss.
Bo Liu 0019, Qingwen Wei, Yang Zhang 0132, Xingyu Xu 0008, Zihan Zou, Xinxiang Huang, Xin Si, Hao Cai 0001
DAC8
2024 Small-Footprint Automatic Speech Recognition System using Two-Stage Transfer Learning based Symmetrized Ternary Weight Network
abstract
Traditional automatic speech recognition (ASR) models face challenges when deployed on edge devices due to their high computational requirements and storage demands. To address this issue, we present a novel ASR system specifically designed for edge applications, encompassing both keyword spotting (KWS) and speaker verification (SV) functionalities with on chip learning for speaker registration. Our proposed system employs a compact model trained using a two-stage transfer learning method for on-the-fly small-sample speaker registration. In the proposed model, sparsity-controllable weights are symmetrically ternary-quantized to further exploit data reuse. Additionally, we introduce a Huffman-coding based weight lossy compression method to achieve efficient storage compaction. Moreover, we propose a specialized classifier taking the signal-to-noise ratio into account to enhance the accuracy of SV. The proposed ASR system has been successfully deployed on a 1.65mm2custom chip fabricated under 28-nm technology, with only 10.84KB of on-chip memory. This compact system effectively handles KWS and SV tasks, as well as on-chip speaker registration.
Xuanhao Zhang, Hui Kou, Chenjie Xia, Hao Cai 0001, Bo Liu 0019
ICASSP4
2024 Live Demonstration: A Target-Separable BWN Inspired Speech Recognition Processor with Low-power Precision-adaptive Approximate Computing
abstract
In this live demonstration, a speech recognition system supporting both keyword spotting (KWS) and speaker verification (SV) is presented. The live demonstration is composed of three parts: the speech recognition system, a display screen, and a PC. The system can be spotted by the speaker (pre-trained in the PC), and the recognition results of the KWS and the SV can be shown on the display screen under different background noises.
Chenjie Xia, Xuanhao Zhang, Zihan Zou, Hao Cai 0001, Bo Liu 0019
ISCAS4
2024 Complementary Series-connected STT-MTJ for Time-based Computing-in-Memory
abstract
Computing-in-memory (CIM) based on spin transfer torque magnetic random access memory (STT-MRAM) is promised to be an effective way to overcome the "memory wall" bottleneck. In this work, we proposed a novel complementary series-connected magnetic tunnel junction (STT-MTJ) structure for time-based Computing-in-Memory (CST-CIM). The bit-cell with four transistors and one MTJ is utilized to establish a series-connected structure to improve the limited resistance of MTJ, which can be applied for high-linearity and sufficient-margin multiply-and-accumulate (MAC) operation. In addition, for peripheral computing circuit, a customized successive-approximation-register time-to-digital converter (SAR-TDC) is used for high energy efficiency and low latency. To optimize the multi-bit MAC operation, we proposed a novel hardware friendly signed binary weight mapping strategy, which can provide the computing flexibility with 1-8bit quantization. Simulation result shows the proposed CST-CIM architecture can achieve low computation latency of 5ns and peak energy efficiency of 106.7 TOPS/W.
Bo Liu 0019, Xin Si, Hao Cai 0001
ISCAS4
2024 Timing Error Tolerant CNN Accelerator With Layerwise Approximate Multiplication
abstract
Exploiting the error tolerance in computation, approximate circuits become an emerging computing paradigm to increase the energy efficiency in digital systems, which is crucial in high-performance and low-power systems for the edge Internet-of-Things (EIoT) devices. Inspired by the state-of-the-art high-efficiency NN accelerators, three techniques are proposed for effectively integrating the approximate computing unit into CNN accelerator to achieve a dynamic energy-accuracy trade-off: (1) An approximate multiplier that can be configured to three precision modes is proposed. A weight pre-encoding method is used to save hardware overhead. (2) For hybrid-accuracy layer-wise mapping, the hessian-aware layer-wise accuracy scaling is proposed, which concerns inference accuracy and hardware overhead simultaneously. A progressive re-training approach is proposed to enable an aggressive approximation configuration and higher power reduction. (3) A tensor multiplication unit (TMU) with timing error detection and correction (TEDC) approach is proposed, enabling an aggressive voltage scaling and a 41.5% power reduction is obtained. An energy-efficient CNN accelerator is proposed and shows how deep learning can be brought to EIoT devices by running each layer at its appropriate computational accuracy. Implemented under 28-nm CMOS technology, the CNN accelerator achieves the energy efficiency of 14.4 TOPS/W. The proposed accelerator and method are conducted on the applications of keyword spotting of GSCD, CIFAR10 and CIFAR100, 44.5%~46.7% multiplication energy is saved while reducing the accuracy by less than 0.6%.
Bo Liu 0019, Na Xie, Qingwen Wei, Guang Yang 0036, Chonghang Xie, Weiqiang Liu 0001, Hao Cai 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2024 Low-Overhead Triple-Node-Upset Self-Recoverable Latch Design for Ultra-Dynamic Voltage Scaling Application
abstract
Ultra-dynamic voltage scaling (UDVS) is a popular trade-off technique between delay and power performance. However, voltage scaling will degrade the radiation-aware reliability of traditional latch obviously. In addition, although the shrinkage of feature sizes results in the reduction of latch area, the occurrence possibility of double node upset (DNU) and triple node upset (TNU) events are increasing. Achieving a good balance among delay, power, area and reliability performance is becoming an important issue in the design of radiation-hardened latches, especially considering the coming commercial aerospace applications. Therefore, this paper proposes a TNU self-recoverable latch with wide voltage range (TRLW), which is low overhead and very suitable for UDVS technique. The TRLW latch is mainly composed of two completely interlocking triangle structures, and is able to self-recover from any possible TNU event. Clock-gated isolated cells are skillfully utilized to avoid current conflict. Meanwhile, six transmission gates are carefully integrated into TRLW latch to reduce the propagation delay$\textit{t}_{d2q}$and critical path delay$\textit{t}_{crit}$. Accordingly, the overall performance of TRLW latch is always excellent from normal voltage to near-threshold voltage (NTV). Simulation results based on 28nm CMOS process show that TRLW latch can achieve complete SNU, DNU and TNU self-recovery in all possible cases, and the soft error rate of TRLW latch only raises by 4.6$\%$when the supply voltage is decreased from 0.9 V to 0.3 V. Moreover, compared with the other reported TNU self-recovery latches, TRLW latch consistently achieves the minimum delay, power, area and delay-power-area product (DPAP) under different process, voltage and temperature (PVT) conditions, and obtains average reductions of 3.43$\times$, 3.03$\times$, 2.66$\times$, 1.40$\times$and 10.83$\times$for$\textit{t}_{d2q}$,$\textit{t}_{crit}$, power, area and DPAP when operating from 0.5 V to 1.0 V.
Xin Chen 0039, Hao Cai 0001, Congyi Zhu, Ying Zhang 0068, Weiqiang Liu 0001
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 Layer-Wise Mixed-Modes CNN Processing Architecture With Double-Stationary Dataflow and Dimension-Reshape Strategy
abstract
With the development of convolutional neural networks (CNN) across various domains, the growth in network structure complexity and computational load has increasingly become a research focus in the deployment of neural networks. The key to current research on neural network accelerators lies in striking a balance between computational accuracy and energy efficiency. This paper proposes a software-hardware co-design to strike the balance for CNN edge applications. On the hardware side, a 3-dimensional tensor engine (3D-TE), achieved with reconfigurable Tensor Processing Units (TPUs), is introduced for efficient convolution computation. We optimize the CNN dataflow on 3D-TE using a dimension reshaping method for feature maps rearrangement, and a double stationary dataflow scheduling to reduce memory access. This paper adopts a configurable approximate multiplier design based on Boolean Matrix Factorization (BMF) based logic synthesis applied in the architecture of TPU. The proposed 3D-TE, characterized by its configurable precision, enables the TPUs to dynamically adapt the bitwidth of features and weights in response to varying precision requirements. On the software side, a hessian-guided layer precision mapping is adopted to reduce unnecessary computational overhead, and a progressive re-training approach is proposed to enable a better approximation configuration and higher power reduction. Fabricated on 28-nm CMOS, this work achieves an optimized energy efficiency of 14.9 TOPS/W and 12.1 TOPS/W for ResNet56 and MobileNetV2 respectively, with 0.6V supply voltage and 150MHz clock frequency, representing an improvement of$1.33\times \sim 8.28\times $over the state-of-the-art works.
Bo Liu 0019, Xinxiang Huang, Yang Zhang 0132, Guang Yang 0036, Han Yan 0014, Chen Zhang 0025, Zejv Li, Yuanhao Wang 0009, Hao Cai 0001
IEEE Trans. Circuits Syst. I Regul. Pap.9
2024 A CFMB STT-MRAM-Based Computing-in-Memory Proposal With Cascade Computing Unit for Edge AI Devices
abstract
The application of non-volatile memory technology is increasingly attractive for Computing-in-memory (CIM) owing to high integration density and negligible standby power consumption. This study proposes an spin-transfer-torque (STT) magnetic random access memory (MRAM) based CIM macro which incorporates following innovative features: 1) cross-feedback margin-boost (CFMB) scheme to enable robust and fast reading operations against process variation and limited Tunneling Magnetoresistance Ratio (TMR); 2) cascade computing units (CCU) and related design method for efficient and stable multi-bit multiply-and-accumulate (MAC) operation; and 3) dual computing mode scheme and resolution adjustable quantization module to optimize energy efficiency and operating speed. The post-simulations are performed under 28nm CMOS&MTJ technology. The results demonstrate the achievement in energy efficiency of 36.4 TOPS/W while performing MAC operations with up to 16-bit weights, 4-bit inputs, and 22-bit outputs.
Yongliang Zhou, Chenghu Dai, Licai Hao, Chunyu Peng, Hao Cai 0001, Xiulong Wu
IEEE Trans. Circuits Syst. I Regul. Pap.9
2024 Layer-Sensitive Neural Processing Architecture for Error-Tolerant Applications
abstract
Neural network (NN) operation has high requirements for storage resources and parallel computing, which bring huge challenges to the deployment of NNs in Internet-of-Things (IoT) devices. Consequently, this work proposed a low-power NN architecture, comprising an energy-efficient NN processor and a Cortex-M3 host processor to achieve state-of-the-art (SOTA) end-to-end inference at the edge. The innovations of this article are as follows: 1) to minimize the bit width of the weight while keeping the loss of accuracy within a small range, cross-layer error tolerance has been analyzed, and mixed precision quantization has been adopted for cross-layer mapping; 2) dynamic reconfigurable tensor processing unit (DR-TPU) with approximate computing has been proposed, which brings$1.45\times $computing energy reduction within 0.46% accurate loss in ResNet-50; and 3) a customized input feature map (IFM) reuse and over-writeback strategy has been adopted, eliminating the recurrent fetching from the on-chip and off-chip memories. The times of on-chip storage access can be reduced by 25%–60%, and the capacity of on-chip memory can be reduced to half of the original. The processor has been implemented at 28-nm CMOS technology. Combining the above work, the proposed architecture can achieve a 53.1% reduction of power and 17.2-TOPS/W energy efficiency.
Zeju Li, Qinfan Wang, Zihan Zou, Qiao Shen 0001, Na Xie, Hao Cai 0001, Hao Zhang 0111, Bo Liu 0019
IEEE Trans. Very Large Scale Integr. Syst.6
2023 Toward Energy-Efficient Sparse Matrix-Vector Multiplication with near STT-MRAM Computing Architecture
abstract
Sparse Matrix-Vector Multiplication (SpMV) is one of the vital computational primitives used in modern workloads. SpMV performs memory access, leading to unnecessary data transmission, massive data access, and redundant multiplicative accumulators. Therefore, we propose the near spin-transfer torque magnetic random access memory (STT-MRAM) processing architecture from three optimization perspectives. These optimizations include (1) the NMP controller receives the instruction through the AXI4 bus to implement the SpMV operation in the following steps, identifies valid data, and encodes the index depending on the kernel size, (2) the NMP controller uses high-level synthesis dataflow in the shared buffer for achieving better performance throughput while do not consume bus bandwidth, and (3) the configurable MACs are implemented in the NMP core without matching step entirely during the multiplication. Using these optimizations, the NMP architecture can access the pipelined STT-MRAM (read bandwidth is 26.7GB/s). The experimental simulation results show that this design achieves up to 66x and 28x speedup compared with state-of-the-art ones and 69x speedup without sparse optimization.
Yueting Li 0001, He Zhang 0011, Hao Cai 0001, Shuqin Lv, Renguang Liu, Weisheng Zhao 0001
ASP-DAC4
2023 Work-in-Process: Error-Compensation-Based Energy-Efficient MAC Unit for CNNs
abstract
Approximate circuits sacrifice accuracy in exchange for energy efficiency and have been widely used in hardware deployment of neural networks (NNs). Since convolution accounts for most of the power consumption in NNs, it is necessary to design an approximate multiplication and accumulation (MAC) unit which improve the energy efficiency of hardware with ignorable accuarcy loss. In this work, an error-compensation-based energy-efficient MAC unit is proposed in which approximate multipliers are designed by Boolean matrix factorization and approximate adders are generated by Cartesian genetic programming. The proposed MAC unit is conducted on CIFAR10 using ResNet-18, where PDP is reduced by 58.8% with an accuracy loss of 0.81%.
Xingyu Xu 0008, Qingwen Wei, Yang Zhang 0132, Hao Cai 0001, Bo Liu 0019
CASES4
2022 Triple-Skipping Near-MRAM Computing Framework for AIoT Era
abstract
Near memory computing (NMC) paradigm shows great significance in non-von Neumann architecture to reduce data movement. The normally-off and instance-on characteristics of spin-transfer torque magnetic random access memory (STT-MRAM) promise energy-efficient storage in the AIoT era. To avoid unnecessary memory-related processing, we propose a novel write-read-calculation triple-skipping (TS) NMC for multiply-accumulate (MAC) operation with minimally modified peripheral circuits. The proposed TS-NMC is evaluated with a custom micro control unit (MCU) in 28-nm high-K metal gate (HKMG) CMOS process and foundry announced universal two-transistor two-magnetic tunnel junction (2T-2MTJ) MRAM cell. The framework consists of a sparse flag which is defined in extra STT-MRAM columns with only 0.73% area overhead, and a calculation block for NMC logic with 9.9% overhead. The TS-NMC can successfully work at 0.6-V supply voltage under 20MHz. This Near-MRAM framework can offer up to ~9S.6 % energy saving compared to commercial SRAM refer to ultra-low-power benchmark (ULP-Benchmark). Classification task on MNIST takes 13nJ/pattern. The energy access of memory, calculation, and the total can be reduced by$52.49\times, 2.7\times$, and 11.3 × respectively from the TS scheme.
Juntong Chen, Hao Cai 0001, Bo Liu 0019, Jun Yang 0006
DATE2
2022 A Target-Separable BWN Inspired Speech Recognition Processor with Low-power Precision-adaptive Approximate Computing
abstract
This paper proposes a speech recognition processor based on a target-separable binarized weight network (BWN), capable of performing both speaker verification (SV) and keyword spotting (KWS). In traditional speech recognition system, the SV based on traditional model and the KWS based on neural networks (NN) model are two independent hardware modules. In this work, both SV and KWS are processed by the proposed BWN with unified training and optimization framework which can be performed for various application scenarios. By the system-architecture co-design, SV and KWS share most of the network parameters, and the classification part is calculated separately according to different targets. An energy-efficient NN accelerator which can be dynamically reconfigured to process different layers of the BWN with splitting calculation of frequency domain convolution is proposed. SV and KWS can be achieved with only one time calculation of each input speech frame, which greatly improves the computing energy efficiency. The computing units of the NN accelerator are optimized using precision-adaptive approximate computing method with Dual-VDD to further reduce the energy cost. Compared to state-of-the-arts, this work can achieve about 4 × reduction in power consumption while maintaining high system adaptability and accuracy.
Bo Liu 0019, Hao Cai 0001, Haige Wu, Anfeng Xue, Zhen Wang 0019, Jun Yang 0006
DATE2
2022 A Low Power DNN-based Speech Recognition Processor with Precision Recoverable Approximate Computing
abstract
This paper proposes a low power speech recognition processor based on an optimized DNN with precision recoverable approximate computing. In order to accelerate and improve energy utilization of DNN, an approximate multiplier based on cartesian genetic programming with weight pre-classification and mismatch compensation is proposed. A partial retraining scheme based on approximate noise is proposed to recover the accuracy loss caused by approximate computing. Experimental results show that the proposed approximate multiplier reduces power consumption by 42.9%, and the partial retraining scheme can recover accuracy of 3.03%~4.34%. Implemented under 22nm, the proposed processor can support the recognition of 10 keywords under different noise types and signal-to-noise ratios (5dB~clean), while the recognition accuracy is 83.36% ~89.82% and power consumption is 8.6μ W.
Bo Liu 0019, Xuetao Wang, Anfeng Xue, Haige Wu, Hao Cai 0001
ISCAS7
2022 Self-compensation tensor multiplication unit for adaptive approximate computing in low-power CNN processing
Bo Liu 0019, Hao Cai 0001, Reyuan Zhang, Zhen Wang 0019, Jun Yang 0006
Sci. China Inf. Sci.3
2022 Bit-error-rate aware sensing-error correction interaction in spintronic MRAM
Hao Cai 0001, Xinfang Tong, Xinning Liu, Bo Liu 0019
J. Syst. Archit.1
2022 An Efficient BCNN Deployment Method Using Quality-Aware Approximate Computing
abstract
As the artificial intelligence and Internet of Things (AIoT) develop rapidly, the deployment of artificial neural networks in edge computing is becoming significant with great challenge. The binarized convolutional neural network (BCNN) is one of the most widely adopted light-weight ANNs in AIoT, which can achieve the balance of system accuracy and hardware resource consumption, compared to others. To achieve high power and area efficiency in BCNN deployment, many approximate computing (AxC) techniques are integrated to make full use of the resilience of BCNN. As the research focused on the integration of AxC in circuit design, the design of AxC itself is not fully considered when applied to specific applications or domains. Based on circuit-architecture-system co-design, this article proposes an efficient BCNN deployment method, including a quality-circuit co-design method for approximate adder generation, a quality-aware intercompensation approach for addition tree, and a computing quality involved retraining approach for BCNN deployment. Experimental results show that the proposed quality model can achieve 86.43% in average accuracy while evaluating nine types of typical approximate adders. The proposed method is conducted on the applications of keyword spotting of GSCD, MNIST, and CIFAR-10, and we can further rise the approximation degree by 50%–75%, while reducing the accuracy by less than 1%.
Bo Liu 0019, Xuetao Wang, Anfeng Xue, Qiao Shen 0001, Na Xie, Yu Gong 0002, Zhen Wang 0019, Jun Yang 0006, Hao Cai 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.11
2022 Proposal of Analog In-Memory Computing With Magnified Tunnel Magnetoresistance Ratio and Universal STT-MRAM Cell
abstract
In-memory computing (IMC) is an effective solution for energy-efficient artificial intelligence applications. Analog IMC amortizes the power consumption of multiple sensing amplifiers with an analog-to-digital converter (ADC) and simultaneously completes the calculation of multi-line data with a high parallelism degree. Based on a universal one-transistor one-magnetic tunnel junction (MTJ) spin transfer torque magnetic RAM (STT-MRAM) cell, this paper demonstrates a novel tunneling magnetoresistance (TMR) ratio magnifying method to realize analog IMC. Previous concerns including low TMR ratio and analog calculation nonlinearity are addressed using device-circuit interaction. The TMR is magnified$7500\times $using a latch structure in combination with the device. Peripheral circuits are minimally modified to enable in-memory matrix-vector multiplication. A current mirror with a feedback structure is implemented to enhance analog computing linearity and calculation accuracy. The proposed design maximumly supports 1024 2-bit input and 1-bit weight multiply-and-accumulate (MAC) computations simultaneously. The proposal is simulated using the 28-nm CMOS process and MTJ compact model. The integral nonlinearity is reduced by 57.6% compared with the conventional structure. 9.47-25.4 TOPS/W is realized with 2-bit input, 1-bit weight, and 4-bit output convolution neural network (CNN).
Hao Cai 0001, Yanan Guo 0004, Bo Liu 0019, Mingyang Zhou 0002, Juntong Chen, Xinning Liu, Jun Yang 0006
IEEE Trans. Circuits Syst. I Regul. Pap.1
2022 Quality Driven Systematic Approximation for Binary-Weight Neural Network Deployment
abstract
Neural networks (NNs) with large scales of artificial neurons are increasingly used in recognition and classification tasks. In power-constrained scenarios, the tradeoff between performance and hardware consumptions must be carefully evaluated before silicon tape-out. In this paper, we proposed a systematic approach to design ultra-low power NN system. This work is motivated by the facts that NNs are resilient to approximation in many of the computations and NNs are outputting statistical tensors which are acceptable to less-than-perfect results. We resort to the front-back end approach with a twofold aim: (1) a fast and accurate design approach is proposed by estimating the computing quality of low-power approximate adder arrays, and it is adopted to evaluate the neural network system; (2) a quality configurable engine with different approximation degrees while processing NNs is implemented. The proposed work is demonstrated with a comprehensive keyword spotting (KWS) system as an ultra-low power NN engine. The experimental environment is setup with ten keywords from the google speech command dataset (GSCD) using an industrial 22-nm ultra-low-leakage (ULL) process. Comparing to the state-of-the-art KWS processors, the proposed approximate NN engine can demonstrate over 60% improvement in power efficiency and$1.1\times $area efficiency while achieving similar recognition accuracy.
Yu Gong 0002, Hao Cai 0001, Haige Wu, Hao Yan 0002, Zhen Wang 0019, Longxing Shi, Bo Liu 0019
IEEE Trans. Circuits Syst. I Regul. Pap.2
2022 More is Less: Domain-Specific Speech Recognition Microprocessor Using One-Dimensional Convolutional Recurrent Neural Network
abstract
Low-power keywords recognition has been a focus of acoustic signal processing for several decades. This work investigates the domain-specific speech recognition microprocessor based on optimized one-dimensional convolutional recurrent neural network (1D-CRNN). Compared to previous DNN based frameworks, the proposed 1D-CRNN can process both the feature extraction and keywords classification, and achieve high recognition accuracy with reduced computation operations under wide range background noise SNRs. An energy-efficient 1D-CRNN accelerator is implemented to dynamically reconfigure and process the different layers. This accelerator has the characteristics of “More is Less” in three aspects: 1) the hybrid network with more complex layers is much more compact and requires less computation; 2) although the weight width quantized to 8 bits requires more memory size and multiplication energy cost, the required network neurons can be reduced and hardware utilization can be improved; 3) an energy-aware self-compensation tensor multiplication unit with dual power supply based on approximation design method can be utilized for 1D-CRNN computing. Compared to the state-of-the-art architectures, the novel more-is-less architecture can achieve a much lower power consumption of$1.4~\mu \text{W}\sim 2.1~\mu \text{W}$(over 80% reduced) under an industry 22nm technology, while maintaining higher system adaptability (support SNRs: −5dB~Clean) for 1~5 real-time keywords recognition.
Bo Liu 0019, Hao Cai 0001, Xiaoling Ding, Yu Gong 0002, Weiqiang Liu 0001, Jinjiang Yang, Zhen Wang 0019, Jun Yang 0006
IEEE Trans. Circuits Syst. I Regul. Pap.2
2021 A 1D-CRNN Inspired Reconfigurable Processor for Noise-robust Low-power Keywords Recognition
abstract
A low-power high-accuracy reconfigurable processor is proposed for noise-robust keywords recognition and evaluated in 22nm technology, which is based on an optimized one-dimensional convolutional recurrent neural network (1D-CRNN). In traditional DNN-based keywords recognition system, the speech feature extraction based on traditional algorithms and the DNN based keywords classification are two independent modules. Compared to the traditional architecture, both the feature extraction and keywords classification are processed by the proposed 1D-CRNN with weight/data bit width quantized to 8/8 bits. Therefore unified training and optimization framework can be performed for various application scenarios and input loads. The proposed 1D-CRNN based keywords recognition system can achieve a higher recognition accuracy with reduced computation operations. Based on system-architecture co-design, an energy-efficient DNN accelerator which can be dynamically reconfigured to process the 1D-CRNN with different configurations is proposed. The processing circuits of the accelerator are optimized to further improve the energy efficiency using a fine-grained precision reconfigurable approximate multiplier. Compared to the state-of-the-art architectures, this work can support 1~5 real-time keywords recognition with lower power consumption, while maintaining higher system capability and adaptability.
Bo Liu 0019, Zeyu Shen 0003, Lepeng Huang, Yu Gong 0002, Hao Cai 0001
DATE6
2021 A survey of in-spin transfer torque MRAM computing
Hao Cai 0001, Bo Liu 0019, Juntong Chen, Lirida A. B. Naviner, Yongliang Zhou, Zhen Wang 0019, Jun Yang 0006
Sci. China Inf. Sci.1
2020 An Ultra-low Power Keyword-Spotting Accelerator Using Circuit-Architecture-System Co-design and Self-adaptive Approximate Computing Based BWN
abstract
This paper proposed an ultra-low power keyword-spotting (KWS) accelerator using circuit-architecture-system co-design and precision self-adaptive approximate computing based binarized weight network (BWN). To reduce the power consumption while maintaining the system recognition accuracy for different background noise, we first proposed a bit-by-bit layer-by-layer quantization method to quantize the deep neural network (DNN) to BWN. Then, we proposed a precision self-adaptive approximate addition unit to further reduce the BWN energy consumption. Evaluated under TSMC22nm ULL process technology, this work can support up to 10 keywords real time recognition under different background noise types and SNRs (from 5dB to near microphone) with power consumption of 13.6uW.
Bo Liu 0019, Hao Cai 0001, Zeyu Shen 0003, Yu Gong 0002, Lepeng Huang, Zhen Wang 0019
ACM Great Lakes Symposium on VLSI3
2020 A Learning-Based Timing Prediction Framework for Wide Supply Voltage Design
abstract
Wide voltage design provides the tremendous benefits for state-of-the-art circuit design in terms of power consumption reduction and energy efficiency enhancement. The traditional design and verification flow depends on the standard cell libraries, which are only available from foundries for limited PVT (Process-Voltage-Temperature) corners near the nominal voltages, leading to remarkable characterization effort and storage overhead. In this paper, a learning-based framework is proposed to predict circuit path delays across multiple voltages and process corners without the requirement of cell library for each PVT corner, which consists of dilated-CNN (Conventional Neural Network) based feature engineering and ensemble model. The proposed method was verified with the supply voltages ranging from 0.5V to 0.9V under FF, SS and TT corners. Experimental results demonstrate that the prediction error is limited by 4.9% and 7.9% respectively within and across process corners for various working temperatures, which achieves significant precision enhancement compared with related learning-based methods.
Peng Cao 0002, Hao Cai 0001, Aiguo Bu
ACM Great Lakes Symposium on VLSI3
2020 A Modeling Attack Resilient Physical Unclonable Function Based on STT-MRAM
abstract
Physical unclonable function (PUF) is considered as a promising hardware security primitive for a variety of applications. Recently, with the rapid development of integrated circuit (IC), the requirement for low complexity, high power efficiency and high performance PUFs become urgent. Moreover, a variety of powerful attack approaches have been carried out to counterfeit PUFs. This paper proposes a novel PUF design by utilizing the spin transfer torque magnetic random-access memory (STT-MRAM). The intrinsic process variation of STT-MRAM is exploited as an entropy source for generating PUF response. The primary performance metrics in terms of reliability, uniformity, uniqueness, and diffuseness of our proposed PUF have been verified, which validate its functionality. In addition, machine learning based modeling attacks are employed to evaluate the security level of proposed STT-MRAM based PUF (MPUF). The statistical results show that MPUF is much more immune to modeling attacks compared with the traditional Arbiter PUF.
Zhengyi Hou, You Wang 0002, Deming Zhang, Hao Cai 0001
ACM Great Lakes Symposium on VLSI5
2020 A Background Noise Self-adaptive VAD Using SNR Prediction Based Precision Dynamic Reconfigurable Approximate Computing
abstract
This paper proposed a background-noise self-adaptive voice activity detection (VAD) accelerator using SNR prediction based precision dynamic reconfigurable approximate computing. To improve the energy efficiency while maintaining high recognition accuracy for different background noises, two optimization techniques are proposed. Firstly, we proposed a SNR prediction module to analyze and pre-classify the back-ground noise into different levels, and a binarized weight network (BWN) accelerator with reconfigurable data bit width to implement the feature classification of VAD. Then, we proposed an approximate computing architecture with precision self-adaptive approximate addition unit to further reduce the energy consumption of BWN accelerator. Evaluated under 28nm process technology, this work can achieve high recognition accuracy (speech/none-speech hit rate: 95%/92% @10dB, 90%/87% @5dB, and 85%/80% @-5dB) under different background noise (SNR-5dB) with a low power consumption of 2 ~ 8uW.
Bo Liu 0019, Yan Li 0056, Lepeng Huang, Hao Cai 0001, Shisheng Guo, Yu Gong 0002, Zhen Wang 0019
ACM Great Lakes Symposium on VLSI4
2020 Binarized Weight Neural-Network Inspired Ultra-Low Power Speech Recognition Processor with Time-Domain Based Digital-Analog Mixed Approximate Computing
abstract
In this paper, an ultra-low power speech recognition processor is implemented based on an optimized binarized weight neural-network (BWN). To accelerate the BWN and make it energy efficient, we proposed an approximate computing architecture for the quantized BWN based on time-domain digital-analog mixed addition unit and precision optimization with fault-tolerant training method. Experimental results show that the proposed digital-analog mixed approximate computing architecture can significantly reduce the power consumption while maintaining the recognition accuracy. Implemented under TSMC 28nm, the proposed processor can support 10 keywords real time recognition under different noise types and SNRs, while the power consumption is 56μW.
Bo Liu 0019, Hao Cai 0001, Yu Gong 0002, Yan Li 0056, Zhen Wang 0019
ISCAS2
2020 Interplay Bitwise Operation in Emerging MRAM for Efficient In-memory Computing
Hao Cai 0001, Honglan Jiang, Yongliang Zhou, Menglin Han, Bo Liu 0019
CCF Trans. High Perform. Comput.1
2020 Towards an automated design flow for memristor based VLSI circuits
Hao Cai 0001, Chao Wang 0068, Jun Yang 0006
Integr.2
2019 Voltage-Controlled Magnetoelectric Memory Bit-cell Design With Assisted Body-bias in FD-SOI
abstract
Voltage-controlled magnetic anisotropy (VCMA)-magnetic tunnel junction (MTJ) is incorporated into FD-SOI CMOS technology. The design space of 1 transistor-1 MTJ (1T-1M) bit-cell is explored through varied VCMA pulse duration/amplitude and scaling down transistor dimensions. The design point with 1.1 V VCMA pulse amplitude, 0.44 ns pulse duration and W/L = 400 nm/30 nm access transistor shows the ultra low write energy in VCMA-MTJ based bit-cell. It achieves a minimum 3.18 fJ/bit switching energy with 28-nm FD-SOI process. Access transistor sizing is studied, while the ultra low power implementation may lead to MTJ switching failure. Voltage assisted techniques for failure mitigation are proposed based on body-bias generator (BBG). The BBG not only provides VCMA pulse signal to control MTJ barrier, but also generates body-bias to boost the transistor performance. In the presence of forward body-bias (FBB) and increased VCMA pulse level, the proposed strategy is effective in switching failure compensation as well as writing delay improvement.
Hao Cai 0001, Menglin Han, Weiwei Shan, Jun Yang 0006, You Wang 0002, Wang Kang 0001, Weisheng Zhao 0001
ACM Great Lakes Symposium on VLSI1
2018 Design Space Exploration of Magnetic Tunnel Junction based Stochastic Computing in Deep Learning
abstract
Magnetic tunnel junction (MTJ) is considered as a promising memory candidate in the more than Moore era because of high power efficiency, fast access speed, nearly infinite endurance and easy 3D integration. The nondeterministic switching behavior has been profited to exploit new directions for computing methods, such as stochastic computing. In this paper, the application of stochastic switching behavior in stochastic computing is explored for deep neural network (DNN). Stochastic computing method features low logic complexity, low energy consumption and fine-grained parallelism, boosting the performance of DNN system by combining MTJ. As a key block of stochastic computing, MTJ based true random number generator design is presented in details. The functionality has been validated by combining the hardware design and post-processing in software. Simulation results are demonstrated visibly by handwritten digits recognition test to show the accuracy. Furthermore, the performance is investigated in terms of accuracy, energy consumption and memory occupation to find more efficient techniques.
You Wang 0002, Yue Zhang 0010, Youguang Zhang, Weisheng Zhao 0001, Hao Cai 0001, Lirida A. B. Naviner
ACM Great Lakes Symposium on VLSI5
2018 Enabling Resilient Voltage-Controlled MeRAM Using Write Assist Techniques
abstract
Reliability concerns arise in nonvolatile magnetoelectric random access memory (MeRAM) due to continuously nanotechnology scaling down and CMOS-magnetic hybrid integration. The primary objective of this work is to investigate failure mitigation in voltage-controlled magnetic anisotropy-magnetic tunnel junction (VCMA-MTJ) based 1T-1MTJ MeRAM bit-cell, by using MTJ compact model and 28nm fully depleted silicon on insulator (FD-SOI) process design-kit. A comprehensive reliability study is performed considering process variation and aging degradations, including hot carrier injection (HCI), bias temperature instability (BTI), soft breakdown (SBD) and radiation effect. Write assist techniques are proposed to ensure failure resilient MeRAM design. Bit line (BL) boost and negative source line (SL) methods show high efficiency in writing latency improvement and failure mitigation.
Hao Cai 0001, You Wang 0002, Wang Kang 0001, Lirida A. B. Naviner, Weiwei Shan, Jun Yang 0006, Weisheng Zhao 0001
ISCAS1
2017 Energy Efficient Magnetic Tunnel Junction Based Hybrid LSI Using Multi-Threshold UTBB-FD-SOI Device
abstract
The energy scalability of ultra-low power nonvolatile (NV) large-scale integration (LSI) is explored in this paper. Multi-threshold computing (super/near/sub-$V_t$) in hybrid CMOS/ magnetic tunnel junction (MTJ) circuits are investigated based on SPICE-compatible MTJ model and fully depleted silicon on insulator (FD-SOI) devices. Ultra-low supply voltage operation bottlenecks associated with performance loss, parametric variations and function failure are studied in differential pair-based sensing circuit, MTJ writing/control circuit and other building blocks. A case study is performed with three typical NV-flip-flops (NV-FF), which are implemented with 28nm FD-SOI low $V_t$ (LVT) device and forward back-bias. Results show that MTJ writing/control circuit must operate at nominal supply (super-$V_t$) region to guarantee MTJ switching; sensing circuit is configured with near-$V_t$ operation (0.6V) with robustness consideration, whereas other parts could be implemented with near/sub-$V_t$ computing to achieve ultra-low power consumption and energy efficient operations.
Hao Cai 0001, You Wang 0002, Lirida A. B. Naviner, Wang Kang 0001, Weisheng Zhao 0001
ACM Great Lakes Symposium on VLSI1