EDBT 2026 Demo / reviewers in the wild / expert
Liang Chang 0002
dblp:72/6746-2
· DBLP profile ↗
41ranked-venue papers
13as first author
32since 2021 · last 2026
0000-0002-6685-5576ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 39 · 12 first-author · 30 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ERSRP: A 55nm 46.28 MPixels/(s·mm2) 104 FPS Efficient Real-Time Super-Resolution Processor with Layer-Fused Lightweight EngineabstractThis paper presents ERSRP, a 55nm Edge Real-Time Super-Resolution Processor that achieves 104 FPS at FHD resolution with a peak area efficiency of 46.28 MPixels/(s·mm2), 8.86× higher than the state-of-the-art. The processor is designed through a software-hardware co-optimization approach, addressing the low utilization, high memory demand, and workload imbalance challenges inherent in lightweight SR networks. At the algorithmic level, an Ultra-Lightweight Super-Resolution (ULSR) model is proposed that integrates depth-wise and point-wise separable blocks with a pixel-shuffle mechanism to achieve high-quality reconstruction (37.18 dB PSNR and 0.9581 SSIM on Set5) with only 5.62K parameters. At the hardware level, the ERSRP introduces a Lightweight Accelerated Engine (LAE) sup-porting a Layer Parallel Computing Scheme (LPCS) to improve lightweight operator throughput by 48.9%. A Point-wise Layer Fused Scheme (PLFS) further enhances utilization by 4.95× through inter-core and intra-core fusion without intermediate memory. Fabricated in a 55nm UMC CMOS process, the ERSRP achieves a throughput of 215.7 MPixels/s and supports 104 FPS real-time SR at FHD. Gaoxiang Wu, Liang Chang 0002, Jingke Wang, Zhicheng Hu, Xin Zhao 0044, Fengbin Tu, Jun Zhou 0017 |
ISCAS | 2 |
| 2026 | IDEA: A Real-Time Unified Accelerator for Dual-Task Image Restoration: Dehazing and Illumination Enhancement
Gaoxiang Wu, Jingke Wang, Liang Chang 0002, Jun Zhou 0017 |
ISCAS | 3 |
| 2026 | QSAP: Energy and Area-Efficient Query-Based Sparsity-Aware Accelerator for Voxel-Based Point Cloud Neural NetworksabstractVoxel-based neural networks have been widely applied to the processing of large-scale outdoor point cloud data, which first convert points into voxels and then extract features using several sparse convolution and normal convolution layers. The hardware implementation of these networks suffers from complex rulebook generation, irregular memory access, and low hardware utilization. Meanwhile, these networks still have much data sparsity. In this paper, we propose an energy and area-efficient query-based sparsity-aware accelerator for voxel-based point cloud neural networks, namely QSAP. Specifically, a dedicated unit is used to improve the efficiency of rulebook generation. A query-based input feature-writing method is proposed to enhance parallel reading potential. An efficient weight-mapping method is introduced to store unpruned weights in on-chip buffers. A novel input-feature-reading method with consecutive queries is proposed to enable out-of-order execution of convolution operations, thereby improving hardware utilization. The hardware utilization is further enhanced by a proposed pop strategy that minimizes the total number of empty FIFOs. The MAC related to zero value is also skipped in this process. As a result, QSAP achieves superior performance on 22 nm technology with the throughput, energy efficiency, area efficiency, frame rate, and frame energy of 1074 GOPS, 7.79 TOPS/W, 590 GOPS/mm$\mathbf {^{2}}$, 35.2 FPS, and 3.91 mJ/Frame, respectively, better than state-of-the-art works. Licheng Wu, Ting Yue, Xin Zhao 0044, Donghui Xue, Jinxi Huang, Liang Chang 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 9 |
| 2025 | Athena: Accelerating Quantized Convolutional Neural Networks under Fully Homomorphic EncryptionabstractDeep learning under FHE is difficult due to two aspects: (1) formidable amount of ciphertext computations like convolutions, so frequent bootstrapping is inevitable which in turn exacerbates the problem; (2) lack of the support to various non-linear functions in terms of the diversity and accuracy.Previous work primarily used the CKKS-based approach, which requires large parameters and places a heavy burden on the hardware.In this paper, we propose Athena, including a novel framework targeting quantized convolutional neural networks under FHE, and a specialized accelerator to release the maximum potential of the framework.Unlike the classic CKKS-based approach, Athena only requires much smaller parameters, i.e., 2 15 degree and approximately 5 MB ciphertext size.Athena uses a uniform representation, functional bootstrapping, to accurately support any type of activation functions, and is not limited to polynomial approximate fitted functions such as ReLU and sigmoid.We highlight the following results: (1) the accuracy varies by +0.01%/-0.24% compared with the plaintext quantized CNN;(2) the inference performance on the Athena accelerator achieves a speedup of 1.5× to 2.3×, an EDAP improvement of 3.8× to 9.9×, compared with state-of-the-art FHE accelerators. Yinghao Yang 0001, Xicheng Xu, Liang Chang 0002, Xiaowei Li 0001 |
MICRO | 3 |
| 2025 | Trident: The Acceleration Architecture for High-Performance Private Set IntersectionabstractPrivate Set Intersection (PSI) is imperative in discovering the properties of the same data owned by two competitive parties, without revealing anything else of their respective data asset. Existing PSI solutions such as APSI and ORI-PSI suffer from severe communication and computation overhead due to inefficient communication and FHE polynomial evaluation, which hinders their deployment in practice. This issue is evident in both the upper-level protocol and the lower-level hardware platform. In this paper, we propose a novel software/hardware co-design acceleration architecture for PSI, termed as “Trident”, which includes two tightly coupled segments: from the protocol perspective, we investigate existing bottlenecks and propose a new PSI protocol with significantly less communication and computation under the security guarantee; besides, we re-architect the hardware platform by designing a PSI-specific accelerator, implemented with both FPGA and ASIC, targeting the key operations in the proposed protocol. We build a real-world experimental environment with two instantiated parties to verify the acceleration architecture, and highlight the following results: (1) up to 130$\boldsymbol{\times}$/145$\boldsymbol{\times}$speedup for the computation ofreceiverandsenderparties; (2) up to 37$\boldsymbol{\times}$reduction of communication overhead. (3) up to 93,651$\boldsymbol{\times}$and 74,326$\boldsymbol{\times}$higher energy efficiency over the CPU-based ORI-PSI and APSI, respectively. Jinkai Zhang, Yinghao Yang 0001, Zhe Zhou 0003, Zhicheng Hu, Xin Zhao 0044, Liang Chang 0002, Xiaowei Li 0001 |
IEEE Trans. Computers | 6 |
| 2025 | Exploiting the Memory-Compute-Coupling Feature for CIM Accelerator Design OptimizationabstractSRAM computing-in-memory (CIM) accelerators have evolved as a promising solution to the memory wall problem in neural network (NN) models. By integrating memory and compute resources in each macro, CIM accelerators offer massive in-situ computing parallelism and large memory capacity, enabling spatial mapping with layer fusion and potentially keeping layers stationary in CIM. However, CIM’s memory-compute coupling (MCC) feature poses challenges in designing CIM accelerators. From an architecture aspect, designers must balance CIM’s memory and compute resources by optimizing the macro’s memory-compute ratio (MCR) configuration across diverse scenarios. From a mapping aspect, conventional mappings, which allocate each macro exclusively to each layer, face two major problems: a layer-fusion dilemma (the accelerator suffers from excessive memory access due to layer replications or performance degradation due to load imbalance) and a layer-eviction issue (storing layers stationary in CIM is usually infeasible due to limited CIM capacity). To address these challenges, this paper introduces MCC-DSE, an MCC-aware Design Space Exploration framework for architecture-mapping co-optimization of CIM accelerators. We also propose a three-axis CIM division mapping, which interleaves multiple layers in each macro to concurrently optimize memory access and performance during layer fusion as well as reserves a part of CIM memory in each macro for layer pinning. Compared to baseline architecture and mapping, MCC-DSE shows a 1.4x 8.3x EDP reduction across various workloads and chip areas. Moreover, MCC-DSE provides insights into CIM accelerator optimization, such as selecting optimal MCR and configuring CIM dynamically for different scenarios. Yongkun Wu, Jia Chen 0032, Zhenhua Zhu 0002, Jingyu He, Pingcheng Dong, Yonghao Tan, Xin Zhao 0044, Liang Chang 0002, Yu Wang 0002, Fengbin Tu, Chi-Ying Tsui, Kwang-Ting Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2025 | An Ultra-High Performance and Scalable Optical Flow Hardware Accelerator Based on FPGA for Autonomous DrivingabstractOptical flow plays an extremely important role in the field of computer vision and extremely high real-time performance is required especially in autonomous driving. Traditional optical flow methods generally improve the accuracy of optical flow through the image pyramid technique. Nevertheless, the incorporation of the image pyramid elevates the computational complexity. Moreover, the data relationships between pyramid layers result in strong data dependencies, which renders it difficult to accelerate via parallel processing and makes it challenging to fulfill real-time demands in practical scenarios. To address this issue, in this paper, we propose an ultra-high performance and scalable optical flow hardware accelerator based on FPGA with several techniques, including an adaptive optical flow computation technique based on dynamic direction prediction to reduce computation without accuracy degradation, a highly scalable computing architecture with configurable numbers of PEs to improve the flexibility and hardware utilization under different hardware resource constraints, and a reconfigurable pyramid-layer pipeline technique to improve performance and reduce memory size. The proposed hardware accelerator was implemented and evaluated on a Xilinx FPGA ZCU104 achieving ultra-high performance (405 FPS) while maintaining high accuracy (AEE 1.02) compared with SOTA hardware accelerators. Ye Liu 0011, Shuang Hao 0005, Xiuyuan Qi, Zili Huang, Liang Chang 0002, Yu Long 0005, Jun Zhou 0017 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2025 | PIPECIM: Energy-Efficient Pipelined Computing-in-Memory Computation Engine With Sparsity-Aware TechniqueabstractComputing-in-memory (CIM) architecture has become a promising solution to improve the parallelism of the multiply-and-accumulation (MAC) operation for artificial intelligence (AI) processors. Recently, revived CIM engine partly relieves the memory wall issue by integrating computation in/with the memory. However, current CIM solutions still require large data movements with the increase of the practical neural network model and massive input data. Previous CIM works only considered computation without concern for the memory attribute, leading to a low memory computing ratio. This article presents a static-random access-memory (SRAM)-based digital CIM macro supporting pipeline mode and computation-memory-aware technique to improve the memory computing ratio. We develop a novel weight driver with fine-grained ping-pong operation, avoiding the computation stall caused by weight update. Based on our evaluation, the peak energy efficiency is 19.78 TOPS/W at the 22-nm technology node, 8-bit width, and 50% sparsity of the input feature map. Liang Chang 0002, Jingke Wang, Xin Zhao 0044, Wuyang Hao, Haining Tan, Yinhe Han 0001, Jun Zhou 0017 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | A RRAM-based High Energy-efficient Accelerator Supporting Multimodal Tasks for Virtual Reality Wearable DevicesabstractVirtual reality (VR) wearable devices can achieve immersive entertainment by fusing multi-modal tasks from various senses. However, constrained by the short battery life and limited hardware resources of the VR devices, running multiple tasks simultaneously with different modals is difficult. In this paper, we propose an energy-efficient accelerator that supports Multi-modal Tasks for VR devices, namely MTVR. We present a multi-task computing solution based on the flexible multi-task computing core design and efficient computing unit allocation strategy, which simultaneously achieves efficient work of multi-modal tasks. We design an early exit detector to skip invalid calculations, greatly saving energy. In addition, a fine-grained tiny value skip method at multiplier and adder levels is proposed to save energy further. We provide a hybrid RRAM and SRAM memory access scheme, reducing the external memory access (EMA). Through experimental evaluation, the multitask computing core achieves an average computational utilization of 95%. When the invalid input ratio is 90%, energy saving brought by the early exit detector can reach 88%. The tiny value skip method further achieved 13% energy saving. Hybrid memory access scheme obtains 98.9% EMA reduction. We deployed the MTVR accelerator in FPGA and self-designed RRAM, achieving energy efficiency of 3.6 TOPS/W, higher than other single-task accelerators. Xin Zhao 0044, Zhicheng Hu, Zilong Guo, Haodong Fan, Liang Chang 0002 |
DAC | 7 |
| 2024 | An Ultra-Low Power Time-Domain based SNN Processor for ECG ClassificationabstractWearable devices for ECG arrhythmia detection based on artificial neural networks (ANN) are very popular. However, the energy consumption of electrocardiogram (ECG) processing in ANN has become one of the most critical factors. One solution is using a spiking neural network (SNN), effectively reducing power consumption and improving energy efficiency. Nevertheless, the inevitable membrane potential storage and accumulation of SNN result in significant energy and area overheads. This paper proposes a time domain (TD) based SNN processor for ECG classification. We propose a novel memory delay unit (MDU), part of the memory delay line (MDL), to store and accumulate membrane potential. With this method, power consumption can be significantly reduced. Also, we propose a wave generator that works with MDL to maximize computing efficiency. Compared with digital neurons, our proposed TD neurons reduce power consumption by 32.5% and achieve a classification accuracy of 96.8%. It is very suitable for arrhythmia detection wearable devices. Haodong Fan, Liang Chang 0002, Junlu Zhou, Shuisheng Lin, Jun Zhou 0017 |
ISCAS | 2 |
| 2024 | SuperHCA: A Super-Resolution Accelerator with Sparsity-Aware Heterogeneous Core ArchitectureabstractDeep learning-based super-resolution (SR) models have emerged as a potential approach to achieving high-quality images. The large SR networks can achieve a high peak signal-noise ratio (PSNR), a metric to evaluate the quality of the image. However, the SR networks typically contain large amounts of parameters, inducing high computation capacity and memory bandwidth requirements, which are difficult to deploy on embedded hardware. In this work, we develop the Anchor-Based Shuffle Net (ABSN) oriented to develop a hardware accelerator with a dynamic-scale fixed-point (DSFP) quantization method. In addition, we implement the dynamic quantization adaption in hardware. We design a Super-resolution Heterogeneous Accelerator, namely SuperHCA, employing a sparsity-aware heterogeneous architecture to distinguish between dense and sparse workloads to improve inference efficiency. Furthermore, we provide Slice Layer Fusion (SLF) computation in the heterogeneous cores to reduce external memory access and on-chip buffer sizes. The SuperHCA achieves 91 FPS with the lowest area overhead compared to the state-of-the-art works. Zhicheng Hu, Xin Zhao 0044, Liang Chang 0002 |
ISCAS | 5 |
| 2024 | USR-LUT: A High-Efficient Universal Super Resolution Accelerator with Lookup TableabstractSuper-resolution (SR) can promote medical diagnosis efficiency by enriching the details of captured images, such as gastroscopy and colonoscopy. However, the wireless capsule detector used for diagnosis is constrained by the camera’s low resolution and limited computing resources, making it difficult to deploy computation- and memory access-intensive SR models. In this paper, we propose an efficient universal SR accelerator based on lookup tables, namely USR-LUT, which can support various SR algorithms. We design a LUT-based computing unit (LCU) with higher efficiency and lower area overhead. By utilizing the sparsity of deconvolution, we propose an efficient data mapping scheme that can flexibly support convolution and deconvolution with different kernel sizes, achieving a 3.24× acceleration for deconvolution. Tile-based computing is adopted to reduce memory resources and external memory access (EMA) overhead. Through experimental evaluation, compared with LUT-based SR algorithms, the USR-LUT achieves the least LUT storage entries of 82k and the smallest LUT resource overhead of 0.078MB, respectively. The USR-LUT achieves the highest area efficiency of 175.7GOPS/mm2and throughput area ratio (TAR) of 39.9fps/mm2under the 8-bit precision compared with the state-of-the-art works. To the best of our knowledge, this is the first work of SR accelerators adopting LUT-based computing, which is suitable for tiny mobile devices. Xin Zhao 0044, Zhicheng Hu, Liang Chang 0002 |
ISCAS | 3 |
| 2024 | General Purpose Deep Learning Accelerator Based on Bit InterleavingabstractAlong with the rapid evolution of deep neural networks, the ever-increasing complexity imposes formidable computation intensity on the hardware accelerator. In this paper, we propose a novel computing philosophy called “bit interleaving” and the associate accelerator couple called “Bitlet” and Bitlet-X to maximally exploit the bit-level sparsity. Apart from the existing bit-serial/parallel accelerators, Bitlet leverages the abundant “sparsity parallelism” in the parameters to enforce the inference acceleration. Bitlet is versatile by supporting diverse precisions on a single platform, including floating-point 32 and fixed-point from 1b to 24b. The versatility enables Bitlet feasible for both efficient inference and training. Besides, by updating the key compute engine in the accelerator, Bitlet-X could furthermore improve the peak power consumption and efficiency for the inference-only scenario, with competitive accuracy. Empirical studies on 12 domain-specific deep learning applications highlight the following results: (1) up to 81×/21× energy efficiency improvement for training/inference over recent high-performance GPUs; (2) up to 15×/8× higher speedup/efficiency over state-of-the-art fixed-point accelerators; (3) 1.5mm2 area and scalable power consumption from 570mW (fp32) to 432mW (16b) and 365mW (8b) @28nm TSMC; (4) 1.3× improvement of the peak power efficiency for the Bitlet-X over Bitlet; (5) highly configurable justified by the ablation and sensitivity studies. Liang Chang 0002, Xin Zhao 0044, Zhicheng Hu, Jun Zhou 0017, Xiaowei Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | SuperHCA: An Efficient Deep-Learning Edge Super-Resolution Accelerator With Sparsity-Aware Heterogeneous Core ArchitectureabstractDeep learning-based super-resolution (SR) generative models have recently emerged as a promising approach for generating high-quality images. While large SR networks can achieve a high peak signal-to-noise ratio (PSNR) to assess image quality, they often come with a high number of parameters, leading to increased computational and memory requirements that can be challenging to deploy on embedded hardware. In this study, we introduce the Anchor-Based Shuffle Net (ABSN), which is designed to create a hardware accelerator using a dynamic-scale fixed-point (DSFP) quantization method. Additionally, we incorporate dynamic quantization adaptation in the hardware design. Our Super-resolution Heterogeneous Accelerator, SuperHCA, utilizes a sparsity-aware heterogeneous architecture to optimize inference efficiency by distinguishing between dense and sparse workloads. We also propose Slice Layer Fusion (SLF) dataflow and feature-sharing bit interleaving (FSBI) methods in the heterogeneous cores to reduce on-chip buffer sizes. The SuperHCA achieves a frame rate of 91 fps at a target resolution of FHD, with the highest throughput area ratio (TAR) of 22.75 fps/mm2 compared to existing state-of-the-art works. Zhicheng Hu, Xin Zhao 0044, Liang Chang 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | An Efficient GCN Accelerator Based on Workload Reorganization and Feature ReductionabstractThe irregular adjacency matrix and the mismatched computation patterns of Aggregation and Combination phases make Graph Neural Networks (GNNs) challenging to compute efficiently. This paper proposes a software and hardware co-design system to reduce computational latency and memory access based on workload reorganization and feature reduction. In software, the adjacency matrix is preprocessed, and the workload in both feature and node dimensions is concentrated to optimize memory access and hardware utilization. The interlayer nodes are analyzed using Principal Component Analysis (PCA) to explore the minimum feature vector length based on information redundancy, and a unique weight initialization is utilized for retraining to trim the feature vector to the minimum length. In hardware, an efficient GCN accelerator is designed to fully support the reorganized workload by reconfigurable output node computation. The hardware accelerator is implemented using 28-nm CMOS technology. It achieves 3.3 TOPS peak throughput and 2.6 TOPS/W energy efficiency. Compared with HyGCN, this result shows that the proposed method can improve the overall performance by$5\times $with a negligible accuracy loss of less than 0.5%. Chenjia Xie, Zihan Ning, Liang Chang 0002, Yuan Du |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | HDSuper: High-Quality and High Computational Utilization Edge Super-Resolution Accelerator With Hardware-Algorithm Co-Design TechniquesabstractSuper-resolution (SR) techniques have been employed to construct high-definition images from low-quality images. Various neural networks have demonstrated excellent image-reconstruction quality in SR accelerators. However, deploying SR networks on edge devices is limited by resources and power consumption induced by significant algorithm parameters, computation complexity, and external memory accesses. This work explores the hardware algorithm co-design techniques to provide an end-to-end platform with a lightweight super-resolution network (LSR) and an efficient, high-quality SR accelerator HDSuper. For algorithm design, the improved depth-wise separable convolution and pixelshuffle layers are developed to reduce network size and computation complexity by considering the hardware constraints. Also, the improved channel attention (CA) blocks enhance the image reconstruction quality. For hardware accelerator design, we design a unified computing core (UCC) combined with an efficient flattening-and-allocation (F-A) mapping strategy to support various operators with high computational utilization. In addition, we design the patch computing scheme to reduce the external memory access of the hardware architecture. Based on the evaluation, the proposed algorithm achieves high-quality image reconstruction with$37.44dB$PSNR. Finally, the FPGA demonstration and ASIC layout under UMC 55nm are achieved with low power consumption ($2.08 W$and$152 mW$) under the lowest hardware resources compared to the state-of-the-art works. Xin Zhao 0044, Liang Chang 0002, Dongqi Fan, Zhicheng Hu, Ting Yue, Fengbin Tu, Jun Zhou 0017 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | IPOCIM: Artificial Intelligent Architecture Design Space Exploration With Scalable Ping-Pong Computing-in-Memory MacroabstractComputing-in-memory (CIM) architecture has become a possible solution to designing an energy-efficient artificial intelligent processor. Various CIM demonstrators indicated the computing efficiency of CIM macro and CIM-based processors. However, previous studies mainly focus on macro optimization and low CIM capacity without considering the weight update strategy of CIM architecture. The artificial intelligence (AI) processor with a CIM engine practically induces issues, including updating memory data and supporting different operators. For instance, AI-oriented applications usually contain various weight parameters. The weight stored in the CIM architecture should be reloaded for the considerable gap between the capacity of CIM and growing weight parameters. The computation efficiency of the CIM architecture is reduced by the weight updating and waiting. In addition, the natural parallelism of CIM leads to the mismatch of various convolution kernel sizes in different networks and layers, which reduces hardware utilization efficiency. In this work, we develop a CIM engine with a ping-pong computing strategy as an alternative to typical CIM macro and weight buffer, hiding the data update latency and improving the data reuse ratio. Based on the ping-pong engine, we propose a flexible CIM architecture adapting to different sizes of neural networks, namely, intelligent pong computing-in memory (IPOCIM), with a fine-grained data flow mapping strategy. Based on the evaluation, IPOCIM can achieve a 1.27–$6.27\times $performance and 2.34–$5.30\times $energy efficiency improvement compared to the state-of-the-art works. Liang Chang 0002, Xin Zhao 0044, Ting Yue, Shuisheng Lin, Jun Zhou 0017 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2023 | HDSuper: Algorithm-Hardware Co-design for Light-weight High-quality Super-Resolution AcceleratorabstractSuper-resolution (SR) networks have been gradually applied to embedded devices with good-quality image reconstruction. However, the hardware performance and power efficiency are limited by a large number of algorithm parameters, computation complexity, and hardware resources, obstructing the development of a high-quality SR accelerator. This paper proposes an end-to-end platform with a lightweight super-resolution network (LSR) and an efficient, high-quality super-resolution architecture HDSuper, to perform algorithm-hardware co-design for the SR accelerator. For algorithm design, we employ depth-wise separable convolution and pixelshuffle to reduce network size and computation complexity by considering the hardware constraints. For hardware design, we provide a unified computing core (UCC) combined with an efficient flattening-and-allocation (F-A) mapping strategy to support various operators with high computational utilization. We adopt the patch training method to reduce the external memory access of the hardware architecture. Based on the evaluation, the proposed algorithm achieves high-quality image reconstruction with 37.44dB PSNR. Finally, we implement the image reconstruction in FPGA demonstration, achieving high-quality image reconstruction with 2.08W power consumption under the lowest hardware resources compared to the state-of-the-art works. Liang Chang 0002, Xin Zhao 0044, Dongqi Fan, Zhicheng Hu, Jun Zhou 0017 |
DAC | 1 |
| 2023 | A FPGA-Based Iterative 6DoF Pose Refinement Processing Unit for Fast and Energy-Efficient Pose Estimation in Picking RobotsabstractFast and energy-efficient 6D pose estimation is essential for robotic applications, especially for picking robots in industrial scene. Introducing iterative pose refinement in the final stage of pose estimation pipeline can effectively improve the precision. However, this procedure can be very time consuming due to iterative convolutional neural network (CNN) inference on resource and power constrained platforms. In this paper, we propose a FPGA-based iterative pose refinement processing unit that achieves fast and energy-efficient pose estimation for picking robots. The design and implementation are based on a Xilinx Zynq UltraScale+ MPSoC. Our experimental results demonstrate that the proposed FPGA-Based processing unit is 21.78 times faster and 23.89 times more energy-efficient compared with the baseline, significantly enhances the speed and energy efficiency of pose refinement. The evaluation result on the datasets shows little accuracy drop compared to the baseline implementation. Le Jin, Guoshun Zhou, Liang Chang 0002, Jun Zhou 0017 |
IECON | 9 |
| 2023 | TDPRO: Time-Domain-Based Computing-in Memory Engine for Ultra-Low Power ECG ProcessorabstractFor the wearable biomedical signal detection, both high accuracy and low-power consumption are critical requirements. Various works have employed the neural network to improve the detecting accuracy and develop the biomedical processor. However, the biomedical processor with neural network engine contains massive data movements and large data buffers. One solution is the computing-in memory (CIM) architecture, which locates more data near the computing engine to reduce data movements. In traditional CIM-based solution, the detecting accuracy and power consumption is difficult to be optimized simultaneously, where the accuracy should be satisfied for the detection. To date, the time-domain computing engine have been developed to employ both digital and time domain computation. In this work, we present a high-precision time-domain engine to perform 8-bit multiplication and addition operation for the biomedical signal detection. With the high precision time-domain engine, we develop a CIM-based neural-network processor, namely TDPRO, to perform the detection of arrhythmia. In addition, we develop TD-zero-jumping (TDJ) and idle-shutdown (ISD) techniques according to signal features and data mapping strategy, further optimizing the power consumption. Based on our evaluation, the TD-based 8-bit mulitply-accumulation operation is robust, without declining the accuracy of biomedical signal detection. We design a ECG processor with the proposed TDPRO architecture, which obtains 98.60% high accuracy and 75.7% power saving compared to the recent the state-of-the-art study. Liang Chang 0002, Siqi Yang 0002, Zhiyuan Chang, Haodong Fan, Junlu Zhou, Jun Zhou 0017 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | A High Accuracy and Low Power CNN-Based Environmental Sound Classification ProcessorabstractThe environmental sound classification (ESC) has attracted increasing attention as the environmental sound contains a wealth of information that can be used to detect particular events. However, so far, most of the existing work in ESC still remains in the stage of algorithm design and the design of ESC processor has not been thoroughly investigated. The existing ESC processor designs have issues in meeting low power consumption and high accuracy simultaneously due to the lack of joint-optimization between algorithm and hardware, and very few work has demonstrated a complete ESC system containing all the necessary modules. In this work, a high accuracy and low power CNN-based ESC processor has been proposed, featuring: 1) a big-small CNN-based reconfigurable ESC processing hardware architecture to reduce the power consumption and hardware overhead while maintaining high classification accuracy. 2) a Mel feature adaptation engine reusing the neural network processing unit to further reduce the power consumption. 3) an event-driven ESC processing technique to reduce the inference time and the power consumption. The design has been implemented on a Kintex-7 FPGA and achieves low power consumption of 0.313W with high accuracy of 84.5% for the ESC-50 dataset, outperforming other state-of-the-art ESC processors. Lujie Peng, Junyu Yang, Longke Yan, Xiben Jiao, Jianbiao Xiao, Liang Chang 0002, Yu Long 0005, Jun Zhou 0017 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2023 | ADAS: A High Computational Utilization Dynamic Reconfigurable Hardware Accelerator for Super ResolutionabstractSuper-resolution (SR) based on deep learning has obtained superior performance in image reconstruction. Recently, various algorithm efforts have been committed to improving image reconstruction quality and speed. However, the inference of SR contains huge amounts of computation and data access, leading to low hardware implementation efficiency. For instance, the up-sampling with the deconvolution process requires considerable computation resources. In addition, the sizes of output feature maps of several middle layers are extraordinarily large, which is challenging to optimize, causing serious data access issues. In this work, we present an all-on-chip hardware architecture based on the deconvolution scheme and feature map segmentation strategy, namely ADAS, where all the generated data by the middle layers are buffered on-chip to avoid large data movements between on- and off-chip. In ADAS, we develop a hardware-friendly and efficient deconvolution scheme to accelerate the computation. Also, the dynamic reconfigurable process element (PE) combined with efficient mapping is proposed to enhance PE utilization up to nearly 100% and support multiple scaling factors. Based on our experimental results, ADAS demonstrates real-time image SR and better image reconstruction quality with PSNR (37.15 dB ) and SSIM (0.9587). Compared to baseline and validated with the FPGA platform, ADAS can support scaling factors of 2, 3, and 4, achieving 2.68 ×, 5.02 ×, and 8.28 × speedup. Liang Chang 0002, Xin Zhao 0044, Jun Zhou 0017 |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2022 | An energy-efficient seizure detection processor using event-driven multi-stage CNN classification and segmented data processing with adaptive channel selectionabstractRecently wearable EEG monitoring devices with seizure detection processor using convolutional neural network (CNN) have been proposed to detect the seizure onset of patients in real time for alert or stimulation purpose. High energy efficiency and accuracy are required for the seizure detection processor due to the tight energy constraint of wearable devices. However, the use of CNN and multi-channel processing nature of seizure detection result in significant energy consumption. In this work, an energy-efficient seizure detection processor is proposed, featuring multi-stage CNN classification, segmented data processing and adaptive channel selection to reduce the energy consumption while achieving high accuracy. The design has been fabricated and tested using a 55nm process technology. Compared with several state-of-the-art designs, the proposed design achieves the lowest energy per classification (0.32 μJ) with high sensitivity (97.78%) and low false positive rate per hour (0.5). Jiahao Liu 0006, Zirui Zhong, Hui Qiu, Jianbiao Xiao, Jiajing Fan, Zhaomin Zhang, Sixu Li, Siqi Yang 0002, Weiwei Shan, Shuisheng Lin, Liang Chang 0002, Jun Zhou 0017 |
DAC | 13 |
| 2022 | TDPRO: Ultra-low Power ECG Processor with High-Precision Time-Domain Computing EngineabstractIn wearable biomedical signal detection, the low-power consumption is a critical requirement. However, the process of biomedical signal detection with traditional neural-network processor is uneconomical for large data movements. A typical solution is the near memory computing (NMC) method, locating more data near the computing engine to save energy, where the detecting accuracy and power consumption is difficult to be optimized simultaneously. In addition, a suitable computing engine is needed to match both power and computation budget. In this work, we combine the NMC-based ECG processor equipped with a high-precision time-domain engine to perform the detection of arrhythmia, namely TDPRO. The proposed TDPRO supports high precision multiplication and addition operation with 8-bit input and weight parameters. Also, we propose TD-zero-jumping and idle-shutdown technique to further reduce 63%$\sim$ 91% power consumption of the time-domain engine. The error rate of 8-bit MAC operation in the TDPRO is 1.18%, which is suitable for the ECG detection. Liang Chang 0002, Siqi Yang 0002, Huinan Wang, Jianbo Xiao, Xin Zhao 0044, Shuisheng Lin, Jun Zhou 0017 |
ISCAS | 1 |
| 2022 | ReverSearch: Search-based energy-efficient Processing-in-Memory ArchitectureabstractRecent development of the processing-in-memory (PIM) architecture has demonstrated high efficiency by reducing data movements. However, the performance of the conventional PIM architecture is limited by several issues, including frequent bit-line operations, complicated control of data flow, and massive inter-macro data movements. In addition, both analog- and digital-PIM solutions have obstacles to meet requirement of high-precision computation. In this work, we explore the tradeoff between data movement and energy efficiency of PIM architecture. We develop a PIM architecture, namely ReverSearch, to accelerate multiple-and-accumulate operation, equipped with reverse searching engine and look up table operations. Also, the corresponding data mapping and data flow methods are provided to improve the performance of the ReverSearch architecture. Based on our evaluation, ReverSearch improves the energy efficiency by 17.26 × and 3.68 ×, compared to the baseline of LUT-Cache [1] and LAcc [2]. Weihang Li, Liang Chang 0002, Jiajing Fan, Xin Zhao 0044, Hengtan Zhang, Shuisheng Lin, Jun Zhou 0017 |
ISCAS | 2 |
| 2022 | MobileSP: An FPGA-Based Real-Time Keypoint Extraction Hardware Accelerator for Mobile VSLAMabstractKeypoint extraction is a key technique for Visual Simultaneous Localization and Mapping (VSLAM). Recently, Convolutional Neural Network (CNN) has been used in the keypoint extraction for improving the accuracy. As one of the state-of-the-art CNN based keypoint extraction techniques, the SuperPoint ranked top in the CVPR2020 image matching challenge. However, the use of complex CNN makes it difficult to meet the real-time performance on a mobile platform with limited resource such as mobile robots and wearable Augmented Reality (AR) devices. In this work, based on the SuperPoint, we proposed an FPGA-based real-time keypoint extraction hardware accelerator through algorithm-hardware co-design for mobile VSLAM applications, which is named as MobileSP. Several algorithm and hardware level design techniques have been proposed to reduce the computation and improve the processing speed while maintaining high accuracy, including a partially shared detection & description encoding architecture, a pre-sorting based Non-Maximum Suppression (NMS) engine and a software-hardware hybrid pipeline computing technique. The design has been implemented and evaluated on a ZCU104 FPGA board. It achieves real-time performance of 42 fps with low Absolute Trajectory Error (ATE) of 1.82 cm simultaneously, outperforming several state-of-the-art designs. Ye Liu 0011, Xiuyuan Qi, Liang Chang 0002, Yu Long 0005, Jun Zhou 0017 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2022 | ULECGNet: An Ultra-Lightweight End-to-End ECG Classification Neural NetworkabstractECG classification is a key technology in intelligent electrocardiogram (ECG) monitoring. In the past, traditional machine learning methods such as support vector machine (SVM) and K-nearest neighbor (KNN) have been used for ECG classification, but with limited classification accuracy. Recently, the end-to-end neural network has been used for ECG classification and shows high classification accuracy. However, the end-to-end neural network has large computational complexity including a large number of parameters and operations. Although dedicated hardware such as field-programmable gate array (FPGA) and application-specific integrated circuit (ASIC) can be developed to accelerate the neural network, they result in large power consumption, large design cost, or limited flexibility. In this work, we have proposed an ultra-lightweight end-to-end ECG classification neural network that has extremely low computational complexity (∼8.2k parameters & ∼227k multiplication/addition operations) and can be squeezed into a low-cost microcontroller (MCU) such as MSP432 while achieving 99.1% overall classification accuracy. This outperforms the state-of-the-art ECG classification neural network. Implemented on MSP432, the proposed design consumes only 0.4 mJ and 3.1 mJ per heartbeat classification for normal and abnormal heartbeats respectively for real-time ECG classification. Jianbiao Xiao, Jiahao Liu 0006, Huanqi Yang, Ning Wang 0070, Zhen Zhu 0005, Yu Long 0005, Liang Chang 0002, Jun Zhou 0017 |
IEEE J. Biomed. Health Informatics | 9 |
| 2021 | BitX: Empower Versatile Inference with Hardware Runtime PruningabstractClassic DNN pruning mostly leverages software-based methodologies to tackle the accuracy/speed tradeoff, which involves complicated procedures like critical parameter searching, fine-tuning and sparse training to find the best plan. In this paper, we explore the opportunities of hardware runtime pruning and propose a hardware runtime pruning methodology, termed as “BitX” to empower versatile DNN inference. It targets the abundant useless bits in the parameters, pinpoints and prunes these bits on-the-fly in the proposed BitX accelerator. The versatility of BitX lies in: (1) software effortless; (2) orthogonal to the software-based pruning; and (3) multi-precision support (including both floating point and fixed point). Empirical studies on image classification and object detection models highlight the following results: (1) up to 4.82x speedup over the original non-pruned DNN and 14.76x speedup collaborated with the software-pruned DNN; (2) up to 0.07% and 0.9% higher accuracy for the floating-point and fixed-point DNN, respectively; (3) 2.00x and 3.79x performance improvement over the state-of-the-art accelerators, with 0.039 mm2 and 68.62 mW (floating-point 32), 36.41 mW(16-bit fixed point) power consumption under TSMC 28 nm technology library. Mingzhe Zhang 0005, Liang Chang 0002, Xiaowei Li 0001 |
ICPP | 7 |
| 2021 | Energy-Efficient Spin-Orbit Torque MRAM Operations for Neural Network ProcessorabstractEmerging energy-efficient neural network processor is a promising hardware design to accelerate neural network algorithms with high performance and low power consumption. Typically, static random-access memory (SRAM) is employed to develop large buffers using in the processor. The bit cell of SRAM contains six transistors, leading to low density and large leakage current. In particular, several AI processors need multiple port and transfer-based SRAMs, which decrease the density and increase the power consumption. Recently, emerging spin-orbit torque magnetic random-access memory (SOT-MRAM) becomes a possible solution to replace the SRAM as working memory. However, more operations should be supported by the SOT- MRAM to provide sufficient functions, such as multiple-port memory, transpose memory, data-streaming operations. In this paper, we develop the working memory of neural network processor with SOT-MRAM to build the design library including the transpose operations, multiple-port memory, and data-streaming based buffer arrays. Equiped with those operations provided by SOT-MRAM, we can build high performance and energy-efficient neural network processors. Liang Chang 0002, Zixuan Zhu 0001, Zhen Zhu 0005, Siqi Yang 0002, Weihang Li, Jun Zhou 0017 |
ISCAS | 1 |
| 2021 | Distilling Bit-level Sparsity Parallelism for General Purpose Deep Learning AccelerationabstractAlong with the rapid evolution of deep neural networks, the ever-increasing complexity imposes formidable computation intensity to the hardware accelerator. In this paper, we propose a novel computing philosophy called “bit interleaving” and the associate accelerator design called “Bitlet” to maximally exploit the bit-level sparsity. Apart from existing bit-serial/parallel accelerators, Bitlet leverages the abundant “sparsity parallelism” in the parameters to enforce the inference acceleration. Bitlet is versatile by supporting diverse precisions on a single platform, including floating-point 32 and fixed-point from 1b to 24b. The versatility enables Bitlet feasible for both efficient inference and training. Empirical studies on 12 domain-specific deep learning applications highlight the following results: (1) up to 81 × /21 × energy efficiency improvement for training/inference over recent high performance GPUs; (2) up to 15 × /8 × higher speedup/efficiency over state-of-the-art fixed-point accelerators; (3) 1.5mm2 area and scalable power consumption from 570mW (float32) to 432mW (16b) and 365mW (8b) @28nm TSMC; (4) highly configurable justified by ablation and sensitivity studies. Liang Chang 0002, Zixuan Zhu 0001, Shengjian Lu, Yanhuan Liu, Mingzhe Zhang 0005 |
MICRO | 2 |
| 2021 | Energy-efficient computing-in-memory architecture for AI processor: device, circuit, architecture perspective
Liang Chang 0002, Zhaomin Zhang, Jianbiao Xiao, Zhen Zhu 0005, Weihang Li, Zixuan Zhu 0001, Siqi Yang 0002, Jun Zhou 0017 |
Sci. China Inf. Sci. | 1 |
| 2021 | A Fast and Energy-Efficient SNN Processor With Adaptive Clock/Event-Driven Computation Scheme and Online LearningabstractIn the recent years, the spiking neural network (SNN) has attracted increasing attention due to its low energy consumption and online learning potential. However, the design of SNN processor has not been thoroughly investigated in the past, resulting in limited performance and energy consumption. In this work, a fast and energy-efficient SNN processor with adaptive clock/event-driven computation scheme and online learning capability has been proposed. Several techniques have been proposed to reduce the computation time and energy consumption, including Adaptive Clock- and Event-Driven Computing Scheme, Neighboring PE Borrowing Technique, Compressed Spike Routing Technique and Reconfigurable PE for Inference and Learning. Implemented on a Virtex-7 FPGA, the proposed design achieves computation time of 3.15 ms/image, inference energy consumption of$0.028~\mu $J/synapse/image and online learning energy consumption of$0.297~\mu $J/synapse/image for the MNIST 10-class dataset, which outperform several state-of-the-art SNN processors. The proposed SNN processor is suitable for real-time and energy-constrained applications. Sixu Li, Zhaomin Zhang, Ruixin Mao, Jianbiao Xiao, Liang Chang 0002, Jun Zhou 0017 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2020 | PRISM: Energy-Efficient Polymorphic Operation Based on Spin-Orbit Torque Memory for Reconfigurable ComputingabstractEmerging Non-Volatile Memories (NVMs) including resistive RAM (ReRAM), phase-change memory (PCM), and magnetic RAM (MRAM), have opened up new pathways for the NVM-based reconfigurable computing. Those NVMs technologies can achieve significant energy-efficient computational operations with only minor modification of the peripheral circuits. However, the supported operations are limited by the array structure and low energy-efficiency of implementing the computation using the memory array. In this paper, the Spin Orbit torque-MRAM based polymorphic circuits are proposed to support the reconfigurable computation for reducing the power consumption and improving the functionalities of the single memory array. With the high speed and energy-efficiency write operation, the proposed memory array support both read-out and write-in reconfigurable operations. Liang Chang 0002, Zhaohao Wang, Youguang Zhang, Weisheng Zhao 0001, Jun Zhou 0017 |
ISCAS | 1 |
| 2020 | SemiMap: A Semi-Folded Convolution Mapping for Speed-Overhead Balance on CrossbarsabstractCrossbar architecture has been widely used in neural network (NN) accelerators, involving conventional and emerging devices. It performs well on the fully connected layer through efficient vector-matrix multiplication. Whereas, the advantages degrade on the convolutional layer with huge data reuse, since the execution speed and resource overhead are imbalanced when using existing fully unfolded or fully folded mapping strategy. To address this issue, we propose a novel semi-folded mapping (SemiMap) framework for implementing the convolution on crossbars. It simultaneously folds the physical resources along the row dimension of feature maps (FMs) and unfolds them along the column dimension. The former reduces the resource overhead, and the latter maintains the parallelism. An FM slicing scheme is further proposed to enable the processing of large-size image. Via our mapping framework, a row-by-row streaming pipeline for intraimage dataflow and periodical pipeline for interimage dataflow are easy to be obtained. To validate the idea, we build a many-crossbar architecture with several designs to guarantee the overall functionality and performance. Based on the measurement data of a fabricated chip, a mapping compiler and a cycle-accurate simulator are developed for the hardware simulation of large-scale networks. We evaluate the proposed SemiMap on various convolutional NNs across different network scale. ${>} 35 {\times }$ resource saving and several hundred times cycle reduction are demonstrated compared to the existing fully unfolded and fully folded strategies, respectively. This paper jumps out of the current extreme mapping schemes, and provides a balanced solution on how to efficiently deploy the computational graphs with data reuse on many-crossbar architecture. Lei Deng 0003, Yuan Xie 0001, Ling Liang 0003, Guanrui Wang, Liang Chang 0002, Xing Hu 0001, Liu Liu 0017, Jing Pei, Guoqi Li 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2019 | CORN: In-Buffer Computing for Binary Neural NetworkabstractBinary Neural Networks (BNNs) have obtained great attention since they reduce memory usage and power consumption as well as achieve a satisfying recognition accuracy on Image Classification. In particular to the computation of BNNs, the multiply-accumulate operations of convolution-layer are replaced with the bit-wise operations (XNOR and pop-count). Such bit-wise operations are well suited for the hardware accelerator such as in-memory computing (IMC). However, an additional digital processing unit (DPU) is required for the pop-count operation, which induces considerable data movement between the Process Engines (PEs) and data buffers reducing the efficiency of the IMC. In this paper, we present a BNN computing accelerator, namely CORN, which consists of a Spin-Orbit-Torque Magnetic RAM (SOT-MRAM) based data buffer to perform the majority operation (to replace the pop-count process) with the SOT-MRAM-based IMC to accelerate the computing of BNNs. CORN can naturally implement the XNOR operation in the NVM memory array, and feed results to the computing data buffer for the majority write operation. Such a design removes the pop-counter implemented by the DPU and reduces data movement between the data buffer and the memory array. Based on the evaluation results, CORN achieves 61% and 14% power saving with 1.74× and 2.12× speedup, compared to the FPGA and DPU based IMC architecture, respectively. Liang Chang 0002, Zhaohao Wang, Youguang Zhang, Weisheng Zhao 0001, Yuan Xie 0001 |
DATE | 1 |
| 2019 | DASM: Data-Streaming-Based Computing in Nonvolatile Memory Architecture for Embedded SystemabstractEmerging nonvolatile memories (NVMs), including resistive RAM (RRAM), phase-change memory (PCM), and magnetic RAM (MRAM), have opened up new pathways for Computing-In-Memory (CIM). Those NVM technologies can achieve energy-efficient computational operations with only minor modification of the peripheral circuits. Despite many advantages provided by computational NVMs, parallelism is not sufficiently explored in such CIM designs. To break through this limitation on performance gain, we propose a data-streaming design for the NVM-based CIM (e.g., DASM) by leveraging the underlying parallelism in the hardware. DASM benefits from the massive parallelism of data-streaming computing, reduction in data movement of the CIM, and the nonvolatility of memory arrays. Specifically, data streaming operations can be implemented with CIM bitwise operations in both read-out and write-in procedures. In addition, we use the multilevel power gating for the memory array and connections to further boost the performance. Finally, we study a case of inference process for the quantized deep-neural-network-based on the DASM design. DASM architecture achieves 47.8×, 5.1×, 2.1× speedup compared to the NVIDIA Jetson TK1 embedded GPU board, Intel Xeon E5-2640 CPU, the state-of-the-art field-programmable gate array (FPGA) design, with much lower power consumption. Liang Chang 0002, Zhaohao Wang, Youguang Zhang, Yufei Ding 0001, Weisheng Zhao 0001, Yuan Xie 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2019 | PXNOR-BNN: In/With Spin-Orbit Torque MRAM Preset-XNOR Operation-Based Binary Neural NetworksabstractConvolution neural networks (CNNs) have demonstrated superior capability in computer vision, speech recognition, autonomous driving, and so forth, which are opening up an artificial intelligence (AI) era. However, conventional CNNs require significant matrix computation and memory usage leading to power and memory issues for mobile deployment and embedded chips. On the algorithm side, the emerging binary neural networks (BNNs) promise portable intelligence by replacing the costly massive floating-point compute-andaccumulate operations with lightweight bit-wise XNOR and popcount operations. On the hardware side, the computingin-memory (CIM) architectures developed by the non-volatile memory (NVM) present outstanding performance regarding high speed and good power efficiency. In this paper, we propose an NVM-based CIM architecture employing a Preset-XNOR operation in/with the spin-orbit torque magnetic random access memory (SOT-MRAM) to accelerate the computation of BNNs (PXNOR-BNN). PXNOR-BNN performs the XNOR operation of BNNs inside the computing-buffer array with only slight modifications of the peripheral circuits. Based on the layer evaluation results, PXNOR-BNN can achieve similar performance compared with the read-based SOT-MRAM counterpart. Finally, the end-to-end estimation demonstrates 12.3× speedup compared with the baseline with 96.6-image/s/W throughput efficiency. Liang Chang 0002, Zhaohao Wang, Youguang Zhang, Yuan Xie 0001, Weisheng Zhao 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2018 | Progresses and challenges of spin orbit torque driven magnetization switching and application (Invited)abstractSpin orbit torque (SOT) has been proposed as a potential alternative mechanism to the conventional spin transfer torque (STT) for the magnetization switching. Recently, theoretical and experimental works revealed the novel factors influencing the SOT-driven magnetization switching. Emerging SOT-based spintronics memories and circuits were explored to implement fast and energy-efficient write operation. However, the perspective of the SOT mechanism is still challenged by some serious shortcomings, such as area penalty, relatively large switching current density and undesirable use of external magnetic field. Here, we review the progresses in the SOT mechanism involving the magnetization dynamics, device design and circuit development. Key issues to be addressed in optimizing the SOT devices are pointed out. In particular, we discuss the potential solutions to develop high-density SOT-based memories and circuits. Zhaohao Wang, Zuwei Li, Liang Chang 0002, Wang Kang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 5 |
| 2017 | Voltage-controlled MRAM for working memory: Perspectives and challengesabstractMagnetic random access memory (MRAM) has been widely studied for future nonvolatile working memory candidate. However, the mainstream current (spin transfer torque, STT or spin Hall effect, SHE) driven MRAMs (STT-MRAM or SHE-MRAM) face intrinsic problems in terms of high write power and long latency, significantly limiting the applications for low-power and high-speed working memories. The recently-developed new-generation MRAM, named VCMA-MRAM, which exploits the voltage-controlled magnetic anisotropy (VCMA) effect to write (or assist to write) data information into magnetic tunnel junctions (MTJs), holds the promise to efficiently overcome these problems. Despite the impressive possibility of improving write power and speed, this technology, however, is currently under intensive research and development (R&D), and some challenges still await answers. In this paper, we investigate the perspectives and challenges of VCMA-MRAM for working memories from a cross-layer (device/circuit/architecture) design point of view. We demonstrate that VCMA-MRAM outperforms STT-MRAM and SHE-MRAM in terms of area, speed, energy consumption and instruction-per-cycle (IPC) performance, benefiting from the low-power and high-speed VCMA-driven data writing mechanism. On the other hand, challenges in terms of device fabrication and circuit design should be efficiently addressed before practical applications. Wang Kang 0001, Liang Chang 0002, Youguang Zhang, Weisheng Zhao 0001 |
DATE | 2 |
| 2017 | PRESCOTT: Preset-based cross-point architecture for spin-orbit-torque magnetic random access memoryabstractDue to nearly zero leakage power consumption, non-volatile magnetoresistive random access memory (MRAM) is becoming one of the promising candidates for replacing conventional volatile memories (e.g. SRAM and DRAM). In particular, emerging spin-orbit torque (SOT) MRAM is considered to outperform spin-transfer torque (STT) MRAM due to its fast switching, separate read/write paths, and lower energy dissipation. However, the SOT-MRAM technology is still in its infancy; one key design challenge is that the control of SOT-MRAM, which involves three terminals, is more complicated compared with STT-MRAM. In this paper, we propose a novel MRAM write scheme called PRESCOTT1, where the “1” and “0” data values can be written into memory cells through the SOT and STT, respectively. As a result, the write current is unidirectional rather than bi-directional, which addresses the control complexity. Using this unidirectional write scheme, we design a PreSET-based cross-point (CP) MRAM to improve programing speed, write energy dissipation and storage density compared to conventional MRAM. Circuit simulation results demonstrate that our PreSET-based CP MRAM can achieve around 67.14% average write energy reduction and 50.86% improvement in programming speed, compared with CP STT-MRAM. Liang Chang 0002, Zhaohao Wang, Alvin Oliver Glova, Jishen Zhao, Youguang Zhang, Yuan Xie 0001, Weisheng Zhao 0001 |
ICCAD | 1 |
| 2017 | Pseudo-Differential Sensing Framework for STT-MRAM: A Cross-Layer PerspectiveabstractWith the rapid increase of leakage currents, non-volatile memories have become competitive candidates in the next-generation computer architecture. Among them, STT-MRAM shows great promise in working memory with high density, high speed and tremendous endurance, etc. However, based on our investigations, the dynamic write power and read reliability are two critical challenges of STT-MRAM. In this work, we propose a synergistic pseudo-differential sensing (PDS) framework that employs device, circuit and architectural techniques to address these challenges. In specific, three design techniques, including cell cluster, asymmetric sensing amplifier and self-error-detection-correction, are proposed to implement the PDS framework. We show that the holistic device-circuit-architecture cross-layer co-design enables STT-MRAM to be utilized in the cache memory, benefiting from the improved density, reliability and energy-efficiency. Our experimental results show that the proposed PDS scheme improves the read margin by ~35.6 percent, reduces the area, read latency, read energy, write latency and write power by ~46.7, ~9.8, ~30.3, ~2.3 and ~31.1 percent respectively, compared with the typical 1T1MTJ cell structure for the cache capacity of 8 MB. In addition, the proposed PDS scheme reduces the dynamic energy by ~32.9 percent and leakage energy by ~830 percent, improves the IPC by ~1.3 percent and miss rate by ~36.9 percent respectively, compared with conventional SRAM based cache. Wang Kang 0001, Liang Chang 0002, Zhaohao Wang, Weifeng Lv, Guangyu Sun 0003, Weisheng Zhao 0001 |
IEEE Trans. Computers | 2 |