Jinshan Yue

dblp:180/3832 · DBLP profile ↗
← Back
33ranked-venue papers
1as first author
25since 2021 · last 2026
0000-0001-8234-7400ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 30 · 1 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021
YearPublicationVenuePosition
2026 An area/energy-efficient RRAM computing-in-memory macro with fully-charge-domain multi-bit computation
Shengzhe Yan, Zhuoyu Dai, Zeyu Guo 0002, Zhaori Cong, Zhihang Qian, Xiangqu Fu, Chunmeng Dou, Dashan Shang, Jinshan Yue
Sci. China Inf. Sci.12
2025 Efficient Edge Vision Transformer Accelerator with Decoupled Chunk Attention and Hybrid Computing-In-Memory
abstract
Vision Transformers (ViTs) are new foundation models for vision applications. Edge-deploying ViTs to realize energy-saving, low-latency, and high-performance dense predictions have wide applications, such as autonomous driving and surveillance image analysis. However, the quadratic complexity of the self-attention mechanism renders ViTs slow and resource-intensive, particularly for pixel-level dense predictions that involve long contexts. Additionally, the pyramid-like architecture of modern ViT variants leads to an unbalanced workload, further reducing hardware utilization and decreasing the throughput of conventional edge devices. To this end, we propose an algorithm-hardware co-optimized edge ViT accelerator tailored for efficient dense predictions. At the algorithm level, we propose a decoupled chunk attention (DCA) mechanism implemented in a pipelined manner to reduce off-chip memory access, thereby enabling efficient dense predictions within limited on-chip memory. At the architecture level, we introduce a hybrid architecture that combines SRAM-based computing-in-memory (CIM) and nonvolatile RRAM storage to eliminate extensive off-chip memory access, with a fusion scheduling to balance workloads and minimize intermediate on-chip memory access. At the circuit level, a bit/element two-way-reconfigurable CIM macro is proposed to improve hardware utilization across pyramidal ViT blocks with varied matrix sizes. The experimental results on object detection, semantic segmentation, and depth estimation tasks demonstrate that our design can efficiently process patch lengths up to 16384 with a speedup of 18.5×-217.1×, a reduction in memory accesses of 1.7×-7.4×, and an improvement in energy efficiency of 1.8×, under less than 1% performance degradation.
Yi Li 0049, Zijian Ye, Xiangqu Fu, Songqi Wang, Shucheng Du, Ning Lin, Dashan Shang, Jinshan Yue, Xiaojuan Qi 0001, Feng Zhang 0014
DAC8
2025 An Energy-Efficient High-Utilization Hardware Architecture for Attention Mechanism in Transformer using Balanced Systolic Array and Multi-Row Interleaved Operation Ordering
abstract
Transformer-based neural networks have achieved remarkable performance. Designing energy-efficient and high-speed accelerators for the attention mechanism, which dominates the energy and latency in Transformers, has become increasingly significant. Existing attention accelerators commonly use algorithm-hardware co-design to achieve higher energy efficiency and speed. However, deeply customized algorithms make these accelerators dependent on a particular application. Therefore, optimizing hardware architecture is crucial for achieving general-purpose acceleration. We observe two limitations in the hardware architecture of existing attention accelerators. First, the widely used input stationary, weight stationary, and output stationary systolic arrays (SAs) can’t balance data reuse, register saving, and utilization, which hinders to build more energy-efficient and faster SA-based accelerators. Second, layer-by-layer operation ordering introduces high SRAM access overhead of intermediate results. To address the first limitation, we propose the “Balanced Systolic Array”, which improves energy efficiency by 40% compared to conventional systolic arrays and achieves a utilization rate of 99.5%. To address the second limitation, we propose “Multi-Row Interleaved” operation ordering, which reduces the SRAM energy by 31.7% By integrating two techniques, the proposed attention accelerator achieves a 39% improvement in energy efficiency and a 38% enhancement in throughput×energy efficiency compared to previous works.
Haiyang Zhou, Hongyang Hu, Jinshan Yue, Hanghang Gao, Yuanlu Xie, Xiaoxin Xu, Chunmeng Dou, Ming Liu 0022
DAC3
2025 A monolithic 3D IGZO-RRAM-SRAM-integrated architecture for robust and efficient compute-in-memory enabling equivalent-ideal device metrics
Shengzhe Yan, Zhaori Cong, Zhuoyu Dai, Zeyu Guo 0002, Zhihang Qian, Xufan Li, Chuanke Chen, Nianduan Lu, Chunmeng Dou, Guanhua Yang, Xiaoxin Xu, Di Geng, Jinshan Yue, Ling Li 0013, Ming Liu 0022
Sci. China Inf. Sci.15
2025 SHMT: An SRAM and HBM Hybrid Computing-in-Memory Architecture With Optimized KV Cache for Multimodal Transformer
abstract
Multimodal Transformer (MMT) algorithms have become the state-of-the-art for multimodal tasks such as image captioning. The Encoder-Decoder (E-D) structure, consisting of Encoder, Decoder-causal, and Decoder-cross components, provides a flexible and effective framework for multimodal tasks. However, previous accelerators mainly focus on the dataflow and hardware optimization of the Encoder, which fails to accelerate the entire E-D structure efficiently. There remain three challenges: 1) the lack of pipeline and multicore optimization at the module, layer, and E-D level; 2) the Decoder-causal and Decoder-cross computations have lower arithmetic intensity compared to the Encoder, requiring a better solution for the varying arithmetic intensities; and 3) the autoregressive algorithm in Decoder-causal leads to redundant KV Cache accesses and considerable idle power. In this paper, SHMT, an SRAM and HBM hybrid computing-in-memory (CIM) architecture, is designed to efficiently support multimodal Transformers with three key contributions: 1) a multi-level pipelined multicore scheme, including pipeline optimization across E-D layer-head-module levels and a multicore network-on-chip (NoC) architecture, to reduce inference latency and off-chip accesses; 2) a heterogeneous SRAM-HBM architecture, utilizing high-density HBM-CIM for low-arithmetic-intensity (LAI) parts and high-performance SRAM-CIM for high-arithmetic-intensity (HAI) parts; and 3) by integrating KV Cache with zero-padding in SRAM-CIM, SHMT eliminates redundant read-write operations in KV Cache, reducing idle power consumption. Experiment results show that SHMT achieves 212× speedup, reduces energy consumption by 208×~2000× per token, and achieves 13.3× higher energy efficiency compared to NVIDIA A100 GPU.
Xiangqu Fu, Jinshan Yue, Muhammad Faizan, Zhi Li 0062, Qiang Huo, Feng Zhang 0014
IEEE Trans. Circuits Syst. I Regul. Pap.2
2025 An RRAM-Based Computing-in-Memory Macro With Low-Power Readout/Hold Circuits and Activation Differential Strategy for AdderNet
abstract
AdderNet is an innovative neural network (NN) structure that substitutes multiplications with additions in convolutional operations, while computing-in-memory (CIM) is an efficient architecture that tackles the memory bottleneck for von Neumann architectures. Previous work has explored the SRAM-based CIM AdderNet circuits and demonstrates high energy efficiency. However, it still suffers low storage density, repetitive readout, and redundant comparisons. In this brief, an RRAM-based CIM macro is proposed for efficient AdderNet with the following innovations. First, RRAM cells are adopted to replace SRAM for high-density weight storage. A low-power readout and hold circuit is proposed to save redundant read power of weight data held for multiple cycles. Second, an 8-bit comparator with an early-stop strategy is proposed to compare 8-bit activations and weights in one cycle. Third, an activation (ACT) differential strategy is proposed to reduce redundant comparisons. The proposed 28-nm RRAM CIM macro achieves 12.8-TOPS/mm2peak area efficiency and 126-TOPS/W peak energy efficiency, which is$3.0\times $and$1.2\times $compared with the state-of-the-art AdderNet CIM macro.
Zhihang Qian, Shengzhe Yan, Zhuoyu Dai, Zeyu Guo 0002, Zhaori Cong, Yifan He 0003, Chunmeng Dou, Feng Zhang 0014, Jinshan Yue, Yongpan Liu
IEEE Trans. Very Large Scale Integr. Syst.9
2025 A High-Density Energy-Efficient CNM Macro Using Hybrid RRAM and SRAM for Memory-Bound Applications
abstract
The big data era has facilitated various memory-centric algorithms, such as the Transformer decoder, neural network, stochastic computing (SC), and genetic sequence matching, which impose high demands on memory capacity, bandwidth, and access power consumption. The emerging nonvolatile memory devices and compute-near-memory (CNM) architecture offer a promising solution for memory-bound tasks. This work proposes a hybrid resistive random access memory (RRAM) and static random access memory (SRAM) CNM architecture. The main contributions include: 1) proposing an energy-efficient and high-density CNM architecture based on the hybrid integration of RRAM and SRAM arrays; 2) designing low-power CNM circuits using the logic gates and dynamic-logic adder with configurable datapath; and 3) proposing a broadcast mechanism with output-stationary workflow to reduce memory access. The proposed RRAM-SRAM CNM architecture and dataflow tailored for four distinct applications are evaluated at a 28-nm technology, achieving 4.62-TOPS$/$W energy efficiency and 1.20-Mb$/$mm2memory density, which shows$11.35\times $–$25.81\times $and$1.44\times $–$4.92\times $improvement compared to previous works, respectively.
Shengzhe Yan, Xiangqu Fu, Zhihang Qian, Zhi Li 0062, Zeyu Guo 0002, Zhuoyu Dai, Zhaori Cong, Chunmeng Dou, Feng Zhang 0014, Jinshan Yue, Dashan Shang
IEEE Trans. Very Large Scale Integr. Syst.11
2025 An RRAM Digital Computing-in-Memory Macro With Dual-Mode Multiplication and Maximum Value Rounding Adder Tree
abstract
Implementing digital computing-in-memory (DCIM) based on resistive memory (RRAM) faces several critical challenges due to the small signal margin, large device variations, and large energy- and area-overhead induced by the digital adder tree (AT). To address these issues, we propose an RRAM DCIM macro based on the standard foundry one-transistor-one-resistor (1T1R) cell array featuring: 1) dual-mode MAC operation for efficiency- or accuracy-oriented optimization; 2) margin-enhanced digitized unit (MEDU) to amplify the signal ratio; and 3) maximum value rounding AT (MVR-AT) to reduce its power- and area-overhead. A test chip is demonstrated using a 180 nm CMOS process to verify the concept. It achieves a peak energy efficiency (EF) of 63.08 TOPS/W in the efficiency-oriented mode and a minimum error rate of 1.58% in the accuracy-oriented mode. Their combination can meet the requirements of different workloads in AI computing tasks to optimize the overall power consumption with negligible accuracy loss.
Wang Ye, Hanghang Gao, Zhidao Zhou, Linfang Wang, Weizeng Li, Zhi Li 0062, Jinshan Yue, Xiaoxin Xu, Hongyang Hu, Chunmeng Dou
IEEE Trans. Very Large Scale Integr. Syst.7
2024 IG-CRM: Area/Energy-Efficient IGZO-Based Circuits and Architecture Design for Reconfigurable CIM/CAM Applications
abstract
Artificial intelligence is evolving with various algorithms such as deep neural network (DNN), Transformer, recommendation system (RecSys) and graph convolutional network (GCN). Correspondingly, multiply-accumulate (MAC) and content search are two main operations, which can be efficiently executed on the emerging computing-in-memory (CIM) and content-addressable-memory (CAM) paradigms. Recently, the emerging Indium-Gallium-Zine-Oxide (IGZO) transistor becomes a promising candidate for both CIM/CAM circuits, featuring ultra-low leakage with >300s data retention time and high-density BEOL fabrication. This paper proposes IG-CRM, the first IGZO-based circuits and architecture design for Reconfigurable CIM/CAM applications. The main contributions include: 1) at cell level, propose IGZO-based 3T0C/4T0C cell design that enables both CIM and CAM functionalities while matching IGZO/CMOS voltage; 2) at circuit level, utilize the BEOL IGZO transistor to reduce digital adder tree area in CIM circuits; 3) at architecture level, propose a reconfigurable CIM/CAM architecture with four macro structures based on 3T0C/4T0C cells. The proposed IG-CRM architecture shows high area/energy efficiency on various applications including DNN, Transformer, RecSys and GCN. Experiment results show that IG-CRM achieves 8.09X area saving compared with the SRAM-based non-reconfigurable CIM/CAM baseline, and 1.53×103X/51.9X speedup and 1.63×104X/7.62×103X energy efficiency improvement compared with CPU and GPU on average.
Zeyu Guo 0002, Jinshan Yue, Shengzhe Yan, Zhuoyu Dai, Xiangqu Fu, Zhaori Cong, Zening Niu, Lihua Xu, Guanhua Yang, Di Geng, Ling Li 0013
DAC2
2024 A 2T P-Channel Logic Flash Cell for Reconfigurable Interconnection in Chiplet-Based Computing-In-Memory Accelerators
abstract
In this work, we propose a two-transistor (2T) p-type channel (p-channel) logic-compatible flash cell. Compared to the previous designs, the proposed structure features reduced area-cost and enhanced ability to pass through the logic ‘1’. Due to these advantages, we explore its application as the reconfigurable interconnections in the chiplet-based system. By integrating them into the silicon interposer, the 2T p-channel flash cells can potentially lead to the dense and flexible interconnection between multiple computing-in-memory (CIM) chiplets, resulting in highly reconfigurable and scalable chiplet-based CIM accelerators. A 180nm 1Kb 2T p-channel flash cell array is fabricated and characterized. The characterization results show the 2T p-channel flash cells exhibit a signal ratio >103over 1000 program/erase (P/E) cycles and the device-to-device variations are less than 21.07%. Their typical behaviors as routers are also confirmed by circuit simulations.
Weizeng Li, Linfang Wang, Zhi Li 0062, Wang Ye, Zhidao Zhou, Haiyang Zhou, Hanghang Gao, Jinshan Yue, Hongyang Hu, Fengman Liu, Chunmeng Dou
ISCAS8
2024 A Multichiplet Computing-in-Memory Architecture Exploration Framework Based on Various CIM Devices
abstract
Computing-in-memory (CIM) architectures based on various devices, such as resistive random access memory, SRAM, DRAM, etc., have demonstrated promising energy efficiency. Single-device-based CIM chips show different advantages on performance, power, or area metrics under different workload/operators sizes and application requirements. Some nonidealities, such as the write endurance of some nonvolatile devices, also influence the design choices. Motivated by the emerging 2.5-D/3-D chiplet integration, this work aims to combine the advantages of CIM/storage chips based on different devices, and proposes a design exploration framework to combine the advantages of CIM chips based on these devices in a 3-D-stack architecture. This work proposes: 1) an evaluation method for the power, performance, and area metrics of the multichiplet CIM architecture; 2) an abstraction for the single-device-based CIM chiplets and artificial intelligence algorithm operators; and 3) a mapping and optimization strategy to explore the 2.5-D/3-D CIM chiplet set. The effectiveness of the mapping strategy is verified with a small-scale brute-force search. The proposed design exploration framework can help to find a better-multichiplet CIM architecture. Under a simple design case, the proposed 3-D CIM architecture shows$4.68\times $–$53.32\times $energy efficiency compared with the single CIM chip baselines. The abstracted chiplet library is open-source available in the open-sourcehttps://github.com/dai0dai/3D_CIM_Chiplet_Architecture_Exploration.
Zhuoyu Dai, Feibin Xiang, Xiangqu Fu, Yifan He 0003, Wenyu Sun, Yongpan Liu, Guanhua Yang, Feng Zhang 0014, Jinshan Yue, Ling Li 0013
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.9
2024 A 28-nm Computing-in-Memory-Based Super-Resolution Accelerator Incorporating Macro-Level Pipeline and Texture/Algebraic Sparsity
abstract
Super-resolution (SR) task using the convolutional neural network is a crucial task in improving image and video quality. The introduction of the residual block (RB) raises the depth of the algorithm to perform better reconstruction. The processing of the RB leads to a decrease in hardware utilization and frequent off-chip communications. It is hard to apply such algorithms on edge devices with limited performance. Computing-in-memory (CiM) is one promising method to reduce high power caused by massive data movement in multiply-accumulation computation. The algebraic sparsity (AS) is the structured sparsity (SS) optimization for imaging computing. However, it is an unsolved problem to simultaneously realize the texture sparsity (TS) of the image and the SS of the algorithm in the CiM scheme while maintaining high hardware utilization. Thus, we propose a CiM-based SR task accelerator. There are three key contributions: first, a texture-aware workflow and a dynamic grouping CiM engine can concurrently support TS coupling with AS. Second, a macro-level pipeline scheme together with two custom-sized CiM macros and a high reuse-rate Hadamard transformation circuit reaches 91% hardware utilization. Third, a novel weight update strategy is devised to reduce the performance loss induced by the weight updating. The accelerator prototype is fabricated in a 28-nm CMOS. It scores a 22.8-44.3-TOPS/W peak energy efficiency at the voltage supply of 0.54-1.1 V and the operating frequency of 50-200 MHz, indicating 1.8-6.8x higher compared to the state-of-the-art CiM processors.
Hao Wu 0084, Yong Chen 0005, Yiyang Yuan, Jinshan Yue, Xiangqu Fu, Qirui Ren, Pui-In Mak, Xinghua Wang 0005, Feng Zhang 0014
IEEE Trans. Circuits Syst. I Regul. Pap.4
2024 Write-Verify-Free MLC RRAM Using Nonbinary Encoding for AI Weight Storage at the Edge
abstract
High-density and reliable multilevel-cell (MLC) resistive random access memory (RRAM) is expected to meet the ever-increasing demand for on-chip weight storages in the intelligent edge devices. However, due to the device variations, many write-and-verify (WAV) iterations are usually required to program the RRAM cell, which causes high power consumption, long latency, and degradation on the memory lifetime. To address this issue, we propose a write–verify-free MLC RRAM macro for weight storage with 1) a cascode-current-mirror multibit write (CCM-MW) driver and 2) a nonbinary programming scheme (NB-PS) with a radix not greater than 2. A 180-nm 400-Kb RRAM test chip is demonstrated in silicon. For 2-bit-per-cell MLC storage, the value error rates can be reduced by 24.13% after introducing two redundant bits (RBDs). In addition, compared to the single-level cell (SLC) storage scheme, a 37.50% reduction in the number of cells can be achieved to store the ResNet-8 model with a 0.79% loss in inference accuracy without the need for WAV iterations.
Junjie An, Zhidao Zhou, Linfang Wang, Wang Ye, Weizeng Li, Hanghang Gao, Zhi Li 0062, Jinghui Tian, Hongyang Hu, Jinshan Yue, Lingyan Fan, Shibing Long, Qi Liu 0010, Chunmeng Dou
IEEE Trans. Very Large Scale Integr. Syst.11
2023 A User-Friendly Fast and Accurate Simulation Framework for Non-Ideal Factors in Computing-in-Memory Architecture
abstract
Computing-in-memory (CIM) architecture utilizing emerging non-volatile devices is promising for energy-efficient neural network (NN) applications. However, the non-ideal factors of non-volatile devices and analog circuits may incur severe accuracy loss, which cuts the algorithm and hardware design apart. The algorithm/hardware designers are skilled in either macro-scope NN models or detailed circuit/device errors, while sophisticated research of the joint effect on accuracy loss is urgently needed. In this paper, we propose a user-friendly, fast, and accurate simulation framework (CIMUFAS) to explore the impact of various non-ideal devices/circuits on algorithm accuracy. Based on multiple architecture-level CIM mapping/scheduling workflow, sophisticated non-ideal factors with different error models are established. The CIMUFAS also provides easy-to-use interfaces to flexibly support user-defined models/parameters for specified devices/circuits. Besides, the CIMUFAS framework achieves reasonable simulation time. Compared with MNSIM 2.0, the simulation time is reduced by 45% even after adding a more realistic hardware configuration. This CIMUFAS framework is verified with two fabricated CIM chips with <0.04% accuracy mismatch. The source code of CIMUFAS is publicly available at https://github.com/Hlal/CIMUFAS.
Jinshan Yue, Chaojie He, Zhuoyu Dai, Feibin Xiang, Zhaori Cong, Yifan He 0003, Xiaoyu Feng, Yongpan Liu
ISCAS2
2023 Recent progress in InGaZnO FETs for high-density 2T0C DRAM applications
Shengzhe Yan, Zhaori Cong, Nianduan Lu, Jinshan Yue
Sci. China Inf. Sci.4
2023 SAMBA: Single-ADC Multi-Bit Accumulation Compute-in-Memory Using Nonlinearity- Compensated Fully Parallel Analog Adder Tree
abstract
Performing data-intensive tasks in the von Neumann architecture is challenging to achieve both high performance and energy efficiency due to the memory wall bottleneck. Compute-in-memory (CiM) is a promising mitigation approach by enabling parallel and in-situ multiply-accumulate (MAC) operations within the memory array. Thanks to the good matching of capacitors, SRAM-based charge-domain CiM (Q-CiM) has shown its potential for higher row-wise parallelism. However, the peripheral circuits of Q-CiM, such as the input drivers and analog-digital converters (ADCs), limit further improvement of throughput and area efficiency. This paper proposes a single-ADC multi-bit accumulation CiM macro architecture SAMBA, which can perform multi-bit MAC operation with ReLU of two vectors in one CiM cycle by only a single A/D conversion to mitigate the ADC overhead. In addition, post-correction methods are proposed to compensate the non-linearity of sensitive circuit modules in SAMBA to recover the accuracy drop due to the capacitor mismatch. A proof-of-concept macro is fabricated in a 65nm process and achieves 51.2GOPS throughput and 10.3TOPS/W energy efficiency, while showing 88.6% accuracy on CIFAR-10 and 64.8% accuracy on the CIFAR-100 with VGG-8 model.
Guodong Yin, Mufeng Zhou, Mingyen Lee, Xirui Du, Jinshan Yue, Jiaxin Liu 0001, Huazhong Yang, Yongpan Liu, Xueqing Li 0002
IEEE Trans. Circuits Syst. I Regul. Pap.8
2023 P3 ViT: A CIM-Based High-Utilization Architecture With Dynamic Pruning and Two-Way Ping-Pong Macro for Vision Transformer
abstract
Transformers have made remarkable contributions to natural language processing (NLP) and many other fields. Recently, transformer-based models have achieved state-of-the-art (SOTA) performance on computer vision tasks compared with traditional convolutional neural networks (CNNs). Unfortunately, existing CNN accelerators cannot efficiently support transformer due to the high computational overhead and redundant data accesses associated with the ‘KQV’ matrix operations in the transformer models. If the recently-developed NLP transformer accelerators are applied to the vision transformer (ViT) models, their efficiency would decrease due to three challenges. 1) Redundant data storage and access still exist in ViT data flow scheduling. 2) For matrix transposition in transformer models, the previous transpose-operation schemes lack flexibility, resulting in extra area overhead. 3) The sparse acceleration schemes for NLP in prior transformer accelerators cannot efficiently accelerate ViT with relatively fewer tokens. To overcome these challenges, we propose$P^{3}$ViT, a computing-in-memory (CIM)-based architecture, to efficiently accelerate ViT, achieving high utilization on data flow scheduling. There are three key contributions: 1) P3ViT architecture supports three ping-pong pipeline scheduling modes, involving inter-core parallel and intra-core ping-pong pipeline mode (IEP-IAP3), inter-core pipeline and parallel mode (IEP2), and full parallel mode, to eliminate redundant memory accesses. 2) A two-way ping-pong CIM macro is proposed, which can be configured to regular calculation mode and transpose calculation mode to adapt to both$\text{Q}\times \text{K}^{\mathrm {T}}$and$\text{A}\times \text{V}$tasks. 3) P3ViT also runs a small prediction network. It prunes redundant tokens to be a standard number hierarchically and dynamically, enabling high-throughput and high-utilization attention computation. Measurements show that P3ViT achieves$1.13\times $higher energy efficiency than the state-of-the-art transformer accelerator and achieves$30.8\times $and$14.6\times $speedup compared to CPU and GPU.
Xiangqu Fu, Qirui Ren, Hao Wu 0084, Feibin Xiang, Jinshan Yue, Yong Chen 0005, Feng Zhang 0014
IEEE Trans. Circuits Syst. I Regul. Pap.6
2023 A 40-nm SONOS Digital CIM Using Simplified LUT Multiplier and Continuous Sample-Hold Sense Amplifier for AI Edge Inference
abstract
Digital computing in memory (CIM) exhibits high precision as well as high energy efficiency (EE) yet still lacks discussion in nonvolatile memory (NVM). In this article, we propose a 40-nm silicon-oxide-nitride-oxide-silicon (SONOS)-based digital NVM CIM macro (DNV-CIM) featuring: 1) a simplified lookup table multiplier (SLUTM) combined with a lookup table (LUT) mapping scheme to improve area and EE and 2) a continuous sample-hold sense amplifier (CSH-SA) with an optimized voltage clamper and comparator for continuous read to reduce overall energy and time consumption for deep neural network (DNN) inference tasks. Performance evaluations indicate that the proposed DNV-CIM can achieve 93.04% accuracy and an EE up to 39.9 TOPS/W when running a 4-bit quantized ResNet18 trained on the CIFAR-10 dataset. This work presents a highly efficient digital CIM solution that can be readily implemented with commodity NVM.
Hongyang Hu, Haiyang Zhou, Danian Dong, Jinshan Yue, Wan Pang, Xiaoxin Xu, Chunmeng Dou
IEEE Trans. Very Large Scale Integr. Syst.6
2022 Sparsity-Aware Non-Volatile Computing-In-Memory Macro with Analog Switch Array and Low-Resolution Current-Mode ADC
abstract
Non-volatile computing-in-memory (nvCIM) is a novel architecture used for deep neural networks (DNNs) because it can reduce the movement of data between computing units and memory units. As sparsity has made great progress in DNNs, the existing nvCIM architecture is only optimized for structured sparsity but little for unstructured sparsity. To solve this problem, the sparsity-aware nvCIM macro is proposed to improve the computing performance and network classification accuracy, and to support both structured and unstructured sparsity. First, the analog switch array is used to take advantage of the structured sparsity and to improve the computing parallelism. Second, the low-resolution current-mode analog-to-digital converter (CMADC) is designed to optimize the unstructured sparsity. Experimental results show that the peak equivalent energy efficiency of the proposed nvCIM macro is 9.1 TOPS/W (A8W8, 8-bit activations and 8-bit weights) with only 0.51% accuracy loss, and 584.9 TOPS/W (A1W1), which is 4.8 -$7.5\times$compared to the state-of-the-art nvCIM macros.
Yifan He 0003, Jinshan Yue, Wenyu Sun, Huazhong Yang, Yongpan Liu
ASP-DAC3
2022 Toward Low-Bit Neural Network Training Accelerator by Dynamic Group Accumulation
abstract
Low-bit quantization is a big challenge for neural network training. Conventional training hardware adopts FP32 to accumulate the partial-sum result, which seriously degrades energy efficiency. In this paper, a technology called dynamic group accumulation (DGA) is proposed to reduce the accumulation error. First, we model the proposed group accumulation method and give the optimal DGA algorithm. Second, we design a training architecture and implement a hardware-efficient DGA unit. Third, we make a comprehensive analysis of the DGA algorithm and training architecture. The proposed method is evaluated on CIFAR and ImageNet datasets, and results show that DGA can reduce accumulation bit-width by 6 bits while achieving the same precision as the static group method. With the FP12 DGA, the CNN algorithm only loses 0.11% accuracy in ImageNet training, and our architecture saves 32% of power consumption compared to the FP32 baseline.
Yixiong Yang, Ruoyang Liu, Wenyu Sun, Jinshan Yue, Huazhong Yang, Yongpan Liu
ASP-DAC4
2022 C-RRAM: A Fully Input Parallel Charge-Domain RRAM-based Computing-in-Memory Design with High Tolerance for RRAM Variations
abstract
Previous RRAM-based computing-in-memory works mainly focus on the current-domain approach. However, the performance and accuracy of current computation are limited by large read currents of RRAM cells and their variations. This work presents a novel RRAM-based charge-domain design, C-RRAM, to resolve these limitations. A 3TlRlC cell is proposed to execute MAC operations by capacitor discharging. The resistance variations can be tolerated with reasonable discharging time, which is accelerated by a positive feedback loop. Also, the output from each cell is accumulated by charge sharing instead of summing currents to eliminate the static current path in readout circuits. In this way, robust and efficient RRAM-based CIM operation is enabled with fully input parallelism. A $512\times 514$ RRAM array is implemented to evaluate the benefits of the proposed charge-domain approach. The experiment results show that C-RRAM can suppress the output variations by $41\times$ and incur negligible accuracy loss for ResNet-18 on Cifar10 dataset. Compared to previous ITIR current-domain RRAM designs, it achieves $1.2\times$ energy efficiency and $127\times$ area efficiency due to improved parallelism.
Yifan He 0003, Jinshan Yue, Wenyu Sun, Lu Zhang 0074, Yongpan Liu
ISCAS3
2022 PACA: A Pattern Pruning Algorithm and Channel-Fused High PE Utilization Accelerator for CNNs
abstract
In recent years, convolutional neural networks (CNNs) have achieved significant advancements in various fields. However, the computation and storage overheads of CNNs are overwhelming for Internet-of-Things devices. Both network pruning algorithms and hardware accelerators have been introduced to empower CNN inference at the edge. Network pruning algorithms reduce the size and computational cost of CNNs by regularizing unimportant weights to zeros. However, existing works lack intrakernel structured types to tradeoff between sparsity and hardware efficiency, and the index storage for irregularly pruned networks is significant. Hardware accelerators leverage the sparsity of pruned CNNs to improve energy efficiency. However, their process element (PE) utilization rate is low because of uneven sparsity among input convolutional kernels. To overcome these problems, we propose PACA: a Pattern pruning Algorithm and Channel-fused high PE utilization Accelerator for CNNs. It includes three parts: a pattern pruning algorithm to explore the intrakernel sparsity type and reduce the index storage, a channel-fused hardware architecture to reduce the PEs’ idle rate and improve the performance, and a heuristic and taboo search-based smart fusion scheduler to analyze the idle PE problem and schedule the channel fusion in hardware. To demonstrate the effectiveness of PACA, we have implemented the software parts by Python and the hardware architecture by RTL codes. Experimental results on various datasets show that compared with an existing work, PACA can reduce the index storage overhead by$3.47\times $–$5.63\times $with 3.85–9.12 average patterns, and it can improve the hardware performance by$2.01\times $–$5.53\times $because of PEs’ idle rate reduction.
Jingyu Wang 0004, Songming Yu, Zhuqing Yuan, Jinshan Yue, Ruoyang Liu, Yanzhi Wang 0001, Huazhong Yang, Xueqing Li 0002, Yongpan Liu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 Accuracy Optimization With the Framework of Non-Volatile Computing-In-Memory Systems
abstract
Computing-in-memory (CIM) is a new architecture which is more energy-efficient than the Von Neumann architecture due to the fact that it performs calculation in the memory units which can reduce a large amount of data movement. Nowadays, CIM with non-volatile memory (nvCIM), such as resistive random access memory (RRAM), has become a research frontier to further improve computing performance. Recent works mainly explored how to improve the computing performance of nvCIM, but seldom paid attention to the problem of accuracy loss. In this paper, we propose the nvCIM framework, which can systematically analyze the relationship between the classification accuracy of network models and main analog factors. Based on the nvCIM framework, we further provide detailed optimization methods, including RRAM features, array properties, and ADC parameters. The adaptive voltage-controlled SET and pulse-controlled RESET (VSPR) program-verify scheme is proposed to achieve high-resolution RRAM. And the margin enhancement based current-mode sense amplifier (MECSA) and offset reduction based analog-to-digital converter (ORADC) are proposed to improve the accuracy of analog part computing. Experimental results show that the macro-level and system-level energy efficiency is 112.1 TOPS/W and 9.86 TOPS/W respectively with less than 3.51% loss in contrast to the ideal accuracy, which is$2.9\times $-$25.9\times $compared to the energy efficiency of existing RRAM based nvCIM accelerators.
Yifan He 0003, Jinshan Yue, Huazhong Yang, Yongpan Liu
IEEE Trans. Circuits Syst. I Regul. Pap.3
2021 Block-Circulant Neural Network Accelerator Featuring Fine-Grained Frequency-Domain Quantization and Reconfigurable FFT Modules
abstract
Block-circulant based compression is a popular technique to accelerate neural network inference. Though storage and computing costs can be reduced by transforming weights into block-circulant matrices, this method incurs uneven data distribution in the frequency domain and imbalanced workload. In this paper, we propose RAB: a Reconfigurable Architecture Block-Circulant Neural Network Accelerator to solve the problems via two techniques. First, a fine-grained frequency-domain quantization is proposed to accelerate MAC operations. Second, a reconfigurable architecture is designed to transform FFT/IFFT modules into MAC modules, which alleviates the imbalanced workload and further improves efficiency. Experimental results show that RAB can achieve 1.9x/1.8x area/energy efficiency improvement compared with the state-of-the-art block-circulant compression based accelerator.
Yifan He 0003, Jinshan Yue, Yongpan Liu, Huazhong Yang
ASP-DAC2
2021 A Non-Volatile Computing-In-Memory Framework With Margin Enhancement Based CSA and Offset Reduction Based ADC
abstract
Nowadays, deep neural network (DNN) has played an important role in machine learning. Non-volatile computingin-memory (nvCIM) for DNN has become a new architecture to optimize hardware performance and energy efficiency. However, the existing nvCIM accelerators focus on system-level performance but ignore analog factors. In this paper, the sense margin and offset are considered in the proposed nvCIM framework. The margin enhancement based current-mode sense amplifier (MECSA) and the offset reduction based analog-to-digital converter (ORADC) are proposed to improve the accuracy of the ADC. Based on the above methods, the nvCIM framework is displayed and the experiment results show that the proposed framework has an improvement on area, power, and latency with the high accuracy of network models, and the energy efficiency is 2.3 - 20.4x compared to the existing RRAM based nvCIM accelerators.
Yifan He 0003, Jinshan Yue, Huazhong Yang, Yongpan Liu
ASP-DAC3
2020 High PE Utilization CNN Accelerator with Channel Fusion Supporting Pattern-Compressed Sparse Neural Networks
abstract
Recently CNN-based methods have made remarkable progress in broad fields. Both network pruning algorithms and hardware accelerators have been introduced to accelerate CNN. However, existing pruning algorithms have not fully studied the pattern pruning method, and current index storage scheme of sparse CNN is not efficient. Furthermore, the performance of existing accelerators suffers from no-load PEs on sparse networks. This work proposes a software-hardware co-design to address these problems. The software includes an ADMM-based method which compresses the patterns of convolution kernels with acceptable accuracy loss, and a Huffman encoding method which reduces index storage overhead. The hardware is a fusion-enabled systolic architecture, which can reduce PEs' no-load rate and improve performance by supporting the channel fusion. On CIFAR-10, this work achieves 5.63× index storage reduction with 2-7 patterns among different layers with 0.87% top-1 accuracy loss. Compared with the state-of-art accelerator, this work achieves 1.54×-1.79× performance and 25%-34% reduction of no-load rate with reasonable area and power overheads.
Jingyu Wang 0004, Songming Yu, Jinshan Yue, Zhuqing Yuan, Huazhong Yang, Xueqing Li 0002, Yongpan Liu
DAC3
2020 RL Based Network Accelerator Compiler for Joint Compression Hyper-Parameter Search
abstract
Although compression techniques like pruning or quantization are beneficial for accelerators' energy efficiency, the large search space makes finding the appropriate compression scheme difficult. Besides, most existing works ignore the combination of both pruning and quantization. In this paper, we propose a reinforcement learning (RL) based joint compression framework to find the appropriate pruning ratio and quantization bit-width for accelerators. By interacting with the energy model of the target accelerator, the RL agent can learn the effect of compression scheme on both accuracy and energy efficiency. Through a long trial-and-error process, the agent can finally reach an optimal trade-off between accuracy and energy efficiency. Compared with control groups whose compression hyper-parameters are not jointly optimized, the proposed framework can achieve at least 25% energy reduction with higher accuracy or much higher accuracy with small disadvantages on energy. Compared with 8-bit quantized baseline, the framework can achieve 90% and 85% energy reduction on Cifar10 and Cifar100 respectively.
Xiaoyu Feng, Jinshan Yue, Huazhong Yang, Yongpan Liu
ISCAS2
2019 AERIS: area/energy-efficient 1T2R ReRAM based processing-in-memory neural network system-on-a-chip
abstract
ReRAM-based processing-in-memory (PIM) architecture is a promising solution for deep neural networks (NN), due to its high energy efficiency and small footprint. However, traditional PIM architecture has to use a separate crossbar array to store either positive or negative (P/N) weights, which limits both energy efficiency and area efficiency. Even worse, imbalance running time of different layers and idle ADCs/DACs even lower down the whole system efficiency. This paper proposes AERIS, an Area/Energy-efficient 1T2R ReRAM based processing-In-memory NN System-on-a-chip to enhance both energy and area efficiency. We propose an area-efficient 1T2R ReRAM structure to represent both P/N weights in a single array, and a reference current cancelling scheme (RCS) is also presented for better accuracy. Moreover, a layer-balance scheduling strategy, as well as the power gating technique for interface circuits, such as ADCs/DACs, is adopted for higher energy efficiency. Experiment results show that compared with state-of-the-art ReRAM-based architectures, AERIS achieves 8.5x/1.3x peak energy/area efficiency improvements in total, due to layer-balance scheduling for different layers, power gating of interface circuits, and 1T2R ReRAM circuits. Furthermore, we demonstrate that the proposed RCS compensates the non-ideal factors of ReRAM and improves NN accuracy by 5.2% in the XNOR net on CIFAR-10 dataset.
Jinshan Yue, Yongpan Liu, Fang Su, Shuangchen Li, Zhibo Wang 0004, Wenyu Sun, Xueqing Li 0002, Huazhong Yang
ASP-DAC1
2018 PATH: Performance-Aware Task Scheduling for Energy-Harvesting Nonvolatile Processors
Jinyang Li 0002, Yongpan Liu, Hehe Li, Chenchen Fu, Jinshan Yue, Xiaoyu Feng, Chun Jason Xue, Jingtong Hu, Huazhong Yang
IEEE Trans. Very Large Scale Integr. Syst.6
2017 CNN-based pattern recognition on nonvolatile IoT platform for smart ultraviolet monitoring: (Invited paper)
abstract
Intelligent computing and maintenance-free powering are two desirable characteristics of wearable IoT devices. Energy harvesting nonvolatile intelligent processor (NIP) with neural network computation capability has the potential to advance these goals. Individual ultraviolet (UV) exposure monitoring progressively becomes one conspicuous application of wearable devices. In resource constrained wearable sensor nodes, we can alleviate the data transmission burden via convolutional neural networks (CNNs) based pattern recognition. Nevertheless, in spite of the substantially improved computing capability of NIP, typically computational and memory intensive CNNs are still too bulky for on-node implementation. We develop an CNN-based pattern recognition system for nonvolatile IoT platform for smart UV monitoring, and propose a optimization method to achieve extremely tiny and efficient CNNs. Experimental results show that the offline-trained CNN can recognize individual UV exposure patterns with accuracy of 85%, and the simplified on-node CNN can achieve 93.2% parameters reduction with only 5% accuracy loss.
Jinyang Li 0002, Qingwei Guo, Fang Su, Jinshan Yue, Jingtong Hu, Huazhong Yang, Yongpan Liu
ICCAD5
2017 CORAL: Coarse-grained reconfigurable architecture for Convolutional Neural Networks
abstract
Convolutional Neural Network (CNN) has become one of the most successful technologies for visual classification and other applications. As CNN models continue to evolve and adopt different kernel sizes in various applications, it is necessary for the hardware architecture to support reconfigurability. Previous FPGAs and programmable ASICs are fine-grained reconfigurable but with energy efficiency compromise. Considering specific features of CNNs, this paper presents an energy efficient coarse-grained reconfigurable architecture, denoted as CORAL. An application-specific configuration neural block is proposed for convolution operations with reconfigurable data quantization to reduce both energy consumption and on-chip memory requirements. An optimal data loading strategy is presented for CORAL to achieve the best energy efficiency. Experimental results show that CORAL improves 80.0% energy efficiency while reduces 78.9% chip area and 81.0% reconfiguration time compared with the best up-to-date programmable ASIC solution.
Yongpan Liu, Jinshan Yue, Jinyang Li 0002, Huazhong Yang
ISLPED3
2017 Data Backup Optimization for Nonvolatile SRAM in Energy Harvesting Sensor Nodes
abstract
Nonvolatile static random access memory (nvSRAM) has been widely investigated as a promising on-chip memory architecture in energy harvesting sensor nodes, due to zero standby power, resilience to power failures, and fast read/write operations. However, conventional approaches back up all data from static random access memory into nonvolatile memory when power failures happen. It leads to significant energy overhead and peak inrush current, which has a negative impact on the system performance and circuit reliability. This paper proposes a holistic data backup optimization to mitigate these problems in nvSRAM, consisting of a partial backup algorithm and a run-time adaptive write policy. A statistic dead-block predictor is employed to achieve dead block identification with trivial hardware overhead. An adaptive policy is used to switch between write-back and write-through strategy to reduce the rollback induced by backup failures. Experimental results show that the proposed scheme improves the performance by 4.6% on average while the backup power consumption and the inrush current are reduced by 38.1% and 54% on average compared to the full backup scheme. What is more, the backup capacitor size for energy buffer can be reduced by 40% on average under the same performance constraint.
Yongpan Liu, Jinshan Yue, Hehe Li, Qinghang Zhao, Mengying Zhao, Chun Jason Xue, Guangyu Sun 0003, Meng-Fan Chang, Huazhong Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2016 Performance-aware task scheduling for energy harvesting nonvolatile processors considering power switching overhead
abstract
Nonvolatile processors have manifested strong vitality in battery-less energy harvesting sensor nodes due to their characteristics of zero standby power, resilience to power failures and fast read/write operations. However, I/O and sensing operations cannot store their system states after power off, hence they are sensitive to power failures and high power switching overhead is induced during power oscillation, which significantly degrades the system performance. In this paper, we propose a novel performance-aware task scheduling technique considering power switching overhead for energy harvesting nonvolatile processors. We first give the analysis of the power switching overhead on energy harvesting sensor nodes. Then, the scheduling problem is formulated by MILP (Mixed Integer Linear Programming). Furthermore, a task splitting strategy is adopted to improve the performance and an heuristic scheduling algorithm is proposed to reduce the problem complexity. Experimental results show that the proposed scheduling approach can improve the performance by 14% on average compared to the state-of-the-art scheduling strategy. With the employment of the task splitting approach, the execution time can be further reduced by 10.6%.
Hehe Li, Yongpan Liu, Chenchen Fu, Chun Jason Xue, Donglai Xiang, Jinshan Yue, Jinyang Li 0002, Jingtong Hu, Huazhong Yang
DAC6