Kunming Shao

dblp:357/2753 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0002-9459-0779ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 DS-CIM: Digital Stochastic Computing-In-Memory Featuring Accurate OR-Accumulation via Sample Region Remapping for Edge AI Models
abstract
Stochastic computing (SC) offers hardware simplicity but suffers from low throughput, while high-throughput Digital Computing-in-Memory (DCIM) is bottlenecked by costly adder logic for matrix-vector multiplication (MVM). To address this trade-off, this paper introduces a digital stochastic CIM (DS-CIM) architecture that achieves both high accuracy and efficiency. We implement signed multiply-accumulation (MAC) in a compact, unsigned OR-based circuit by modifying the data representation. Throughput is enhanced by replicating this low-cost circuit 64 times with only a 1× area increase. Our core strategy, a shared Pseudo Random Number Generator (PRNG) with 2D partitioning, enables single-cycle mutually exclusive activation to eliminate OR-gate collisions. We also resolve the 1s saturation issue via stochastic process analysis and data remapping, significantly improving accuracy and resilience to input sparsity. Our high-accuracy DS-CIM1 variant achieves 94.45% accuracy for INT8 ResNet18 on CIFAR-10 with a root-mean-squared error (RMSE) of just 0.74%. Meanwhile, our high-efficiency DS-CIM2 variant attains an energy efficiency of 3566.1 TOPS/W and an area efficiency of 363.7 TOPS/mm2, while maintaining a low RMSE of 3.81%. The DS-CIM capability with larger models is further demonstrated through experiments with INT8 ResNet50 on ImageNet and the FP8 LLaMA-7B model.
Kunming Shao, Jiangnan Yu, Zhipeng Liao, Yi Zou 0001, Kwang-Ting Cheng, Chi-Ying Tsui
DATE1
2026 Configurable Dataflow and Adaptive Mapping Optimization for Hybrid ReRAM and SRAM Compute-in-Memory Accelerator
abstract
Hybrid compute-in-memory (CIM) designs have been proposed recently to facilitate the storing of large number of weights of a neural network on-chip. Notably, ReSCIM wang2024res pairs an SRAM cell with a dedicated ReRAM crossbar, allowing ReRAM to serve as the local storage, significantly enhancing the storage capacity of the SRAM-CIM. The SRAM is custom-designed not only to serve as a storage element for CIM but also to function as a sense amplifier to retrieve the data from the ReRAM, which enables super high bandwidth of weight data loading into the CIM engine. However, existing mapping tools for CIM are inadequate for ReSCIM since they do not fully exploit the unique hardware characteristics and advantages of this novel architecture. In this work, we propose an analytical energy and latency model, which incorporates four key factors: hardware, workload, dataflow, and mapping (HWDM), for executing inference of neural network on the ReSCIM accelerator. Specifically, we first characterize the ReSCIM accelerator hardware specifications and the neural network layers. Next, we introduce three dataflows for ReSCIM, leveraging the high weight-loading bandwidth to reduce memory access for various workloads and layer types. Finally, we develop an algorithm to generate optimal mapping and dataflow strategies aimed at minimizing latency or energy consumption. Using our HWDM model, we design a tile-based ReSCIM accelerator and conduct extensive simulations to obtain the cycle-accurate latency and gate-level energy consumption metrics for inference across different neural networks. We conduct design space exploration (DSE) using the HWDM model on a comprehensive set of benchmarks to minimize inference energy or latency. Experimental results show that our optimal ReSCIM accelerator achieves a 44% reduction in EDP reduction compared to the weight-stationary and fixed mapping baseline for SEResNet50. Moreover, our design exhibits 1.74× higher energy efficiency than the state-of-the-art hybrid TL-nvSRAM wang2023tl accelerator on ResNet 18.
Jingyu He, Kunming Shao, Kwang-Ting Cheng, Chi-Ying Tsui
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2026 Balancing FP8 Computation Accuracy and Efficiency on Digital CIM via Shift-Aware On-the-Fly Aligned-Mantissa Bitwidth Prediction
abstract
FP8 low-precision formats have gained significant adoption in transformer inference and training. However, existing digital compute-in-memory (DCIM) architectures face challenges in supporting variable FP8 aligned-mantissa bitwidths, as unified alignment strategies and fixed-precision multiply accumulate (MAC) units struggle to handle input data with diverse distributions. This work presents a flexible FP8 DCIM accelerator with three innovations: 1) a dynamic shift-aware bitwidth prediction (DSBP) with on-the-fly input prediction that adaptively adjusts weight (2/4/6/8b) and input ($2\sim 12$b) aligned-mantissa precision; 2) a FIFO-based input alignment unit (FIAU) replacing complex barrel shifters with pointer-based control; and 3) a precision-scalable INT MAC array achieving flexible weight precision with minimal overhead. Implemented in 28-nm CMOS with a$64~\times ~96$CIM array, the design achieves 20.4 TFLOPS/W for fixed E5M7, demonstrating$2.8\times $higher FP8 efficiency than previous work while supporting all FP8 formats. Results on Llama-7b show that the DSBP achieves higher efficiency than fixed bitwidth mode at the same accuracy level on both BoolQ and Winogrande datasets, with configurable parameters enabling flexible accuracy–efficiency tradeoffs.
Kunming Shao, Zhipeng Liao, Xijie Huang, Kwang-Ting Cheng, Chi-Ying Tsui, Yi Zou 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2025 SynDCIM: A Performance-Aware Digital Computing-in-Memory Compiler with Multi-Spec-Oriented Subcircuit Synthesis
abstract
Digital Computing-in-Memory (DCIM) is an innovative technology that integrates multiply-accumulation (MAC) logic directly into memory arrays to enhance the performance of modern AI computing. However, the need for customized memory cells and logic components currently necessitates significant manual effort in DCIM design. Existing tools for facilitating DCIM macro designs struggle to optimize subcircuit synthesis to meet user-defined performance criteria, thereby limiting the potential system-level acceleration that DCIM can offer. To address these challenges and enable the agile design of DCIM macros with optimal architectures, we present SynDCIM - a performance-aware DCIM compiler that employs multi-spec-oriented subcircuit synthesis. SynDCIM features an automated performance-to-layout generation process that aligns with user-defined performance expectations. This is supported by a scalable subcircuit library and a multi-spec-oriented searching algorithm for effective subcircuit synthesis. The effectiveness of SynDCIM is demonstrated through extensive experiments and validated with a test chip fabricated in a 40nm CMOS process. Testing results reveal that designs generated by SynDCIM exhibit competitive performance when compared to state-of-the-art manually designed DCIM macros.
Kunming Shao, Fengshi Tian, Jiakun Zheng, Jia Chen 0032, Jingyu He, Hui Wu 0010, Jinbo Chen 0002, Xihao Guan, Fengbin Tu, Jie Yang 0033, Mohamad Sawan, Kwang-Ting Cheng, Chi-Ying Tsui
DATE1
2025 A Flexible Precision Scaling Deep Neural Network Accelerator with Efficient Weight Combination
abstract
Deploying mixed-precision neural networks on edge devices is friendly to hardware resources and power consumption. To support fully mixed-precision neural network inference, it is necessary to design flexible hardware accelerators for continuous varying precision operations. However, the previous works have issues on hardware utilization and overhead of reconfigurable logic. In this paper, we propose an efficient accelerator for 2 ∼ 8-bit precision scaling with serial activation input and parallel weight preloaded. First, we set two loading modes for the weight operands and decompose the weight into the corresponding bitwidths, which extends the weight precision support efficiently. Then, to improve hardware utilization of low-precision operations, we design the architecture that performs bit-serial MAC operation with systolic dataflow, and the partial sums are combined spatially. Furthermore, we designed an efficient carry save adder tree supporting both signed and unsigned number summation across rows. The experiment result shows that the proposed accelerator, synthesized with TSMC 28nm CMOS technology, achieves peak throughput of 4.09TOPS and peak energy efficiency of 68.94TOPS/W at 2/2-bit operations.
Kunming Shao, Fengshi Tian, Kwang-Ting Cheng, Chi-Ying Tsui, Yi Zou 0001
ISCAS2
2025 DIRC-RAG: Accelerating Edge RAG with Robust High-Density and High-Loading-Bandwidth Digital In-ReRAM Computation
abstract
Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by integrating external knowledge retrieval but faces challenges on edge devices due to high storage, energy, and latency demands. Computing-in-Memory (CIM) offers a promising solution by storing document embeddings in CIM macros and enabling in-situ parallel retrievals but is constrained by either low memory density or limited computational accuracy. To address these challenges, we present DIRC-RAG, a novel edge RAG acceleration architecture leveraging Digital In-ReRAM Computation (DIRC). DIRC integrates a high-density multi-level ReRAM subarray with an SRAM cell, utilizing SRAM and differential sensing for robust ReRAM readout and digital multiply-accumulate (MAC) operations. By storing all document embeddings within the CIM macro, DIRC achieves ultra-low-power, single-cycle data loading, substantially reducing both energy consumption and latency compared to off-chip DRAM. A query-stationary (QS) dataflow is supported for RAG tasks, minimizing on-chip data movement and reducing SRAM buffer requirements. We introduce error optimization for the DIRC ReRAM-SRAM cell by extracting the bit-wise spatial error distribution of the ReRAM subarray and applying targeted bit-wise data remapping. An error detection circuit is also implemented to enhance readout resilience against device-and circuit-level variations.Simulation results demonstrate that DIRC-RAG under TSMC 40nm process achieves an on-chip non-volatile memory density of 5.18Mb/mm2and a throughput of 131 TOPS. It delivers a 4MB retrieval latency of 5.6μs/query and an energy consumption of 0.956μJ/query, while maintaining the retrieval precision.
Kunming Shao, Zhipeng Liao, Jiangnan Yu, Xijie Huang, Jingyu He, Fengshi Tian, Yi Zou 0001, Kwang-Ting Cheng, Chi-Ying Tsui
ISLPED1
2024 ReSCIM: Variation-Resilient High Weight-Loading Bandwidth In-Memory Computation Based on Fine-Grained Hybrid Integration of Multi-Level ReRAM and SRAM Cells
abstract
SRAM-CIM is a promising approach to implement efficient accelerator architecture as it enables accurate, energy-efficient AI computing, supporting both analog and digital computation. However, it has low area efficiency. On the other hand, Resistive RAM (ReRAM) provides dense on-chip storage, especially with multi-level cells (MLC), but ReRAM-CIM may introduce inaccuracies due to device variation and only supports analog computation. To leverage the strengths of both technologies, a hybrid architecture that combines them at a fine granularity is desirable. Previous hybrid designs incorporate ReRAM resistors into SRAM to improve storage density. However, they face scalability limitations and restricted signal margins for multi-level RRAM readout, leading to degraded computation accuracy. In this work, we propose ReSCIM, a hybrid compute-in-memory (CIM) architecture that seamlessly integrates multi-level ReRAM into SRAM cells at a fine-grained level. By incorporating a compact ReRAM crossbar in each SRAM cell, a dense CIM marco using SRAM-based computation is achieved. We develop an energy-efficient differential sensing scheme that enables parallel weight loading from local ReRAM crossbars to SRAM cells. This scheme allows multi-bit ReRAM data readout using a single SRAM cell and offers resilience to device variations. Furthermore, We designed a ReSCIM accelerator architecture for efficient AI acceleration, fully utilizing the highly scalable storage and exceptional weight-loading bandwidth. We employ a folded weight-mapping approach for MLC ReRAM cells to guarantee accurate classification even under substantial ReRAM device variations. Experimental results show that ReSCIM accelerators based on both analog and digital-based CIM achieve 60% energy savings and 98% latency savings, and 59× higher area efficiency compared to state-of-the-art all-weights-on-chip AI accelerators on AlexNet.
Jingyu He, Kunming Shao, Jiakun Zheng, Fengshi Tian, Kwang-Ting Cheng, Chi-Ying Tsui
ICCAD3
2023 AutoDCIM: An Automated Digital CIM Compiler
abstract
Digital Computing-in-Memory (DCIM) is an emerging architecture that integrates digital logic into memory for efficient AI computing. However, current DCIM designs heavily rely on manual efforts. This increases DCIM design time and limits the optimization space, making it challenging to satisfy the user specifications of diverse AI applications. This paper presents AutoDCIM, the first automated DCIM compiler. Au-toDCIM takes the user specifications as inputs and generates a DCIM macro architecture with an optimized layout. AutoDCIM’s template-based generation balances handcrafted cell design and agile macro development. AutoDCIM’s layout exploration loop analyzes diverse DCIM array partitioning schemes to satisfy user specifications. The auto-generated DCIM macros present competitive efficiency results in comparison with state-of-the-art silicon-verified DCIM macros.
Jia Chen 0032, Fengbin Tu, Kunming Shao, Fengshi Tian, Xiao Huo, Chi-Ying Tsui, Kwang-Ting Cheng
DAC3