Yiqi Wang 0005

dblp:178/8638-5 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0002-9657-3617ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2026 HR-DCIM: High-Reliability Floating-Point Digital CIM Architecture With Unified Low-Cost Iterative Error Correction
abstract
Digital computing-in-memory (CIM) is a promising computing paradigm for the neural network (NN) acceleration. However, during the actual deployment process of digital CIM chips, we find that existing digital CIM designs face severe computing reliability issues, which are crucial for real product development but remain underexplored. Therefore, this work pioneers a systematic computing reliability analysis for digital CIM across off-memory and in-memory levels. We find that both the off-memory floating-point (FP) exponent alignment and the in-memory random cell bit-flip errors impair digital CIM's computing reliability, causing significant truncation and bit-flip accuracy loss. Critically, existing reliability solutions are incompatible with the unique multi-row accumulation structure of digital CIM, which either severely damage digital CIM's performance or result in prohibitive overhead. To address above challenges, we propose HR-DCIM: a highreliability FP digital CIM architecture featuring unified lowcost iterative error correction. Specifically, for the off-memory reliability, we propose an exponent-mantissa joint-alignment mechanism to repurpose inherent invalid bits of aligned mantissas as compensation bits to reduce alignment truncation loss, without damaging digital CIM's performance. Then, for the in-memory reliability, we propose a remainder aliasing-based unified multiply-accumulation (MAC) error correction mechanism to correct possible MAC errors caused by various cell error cases with low-cost iteration. Experimental results show that the proposed techniques enable digital CIM to maintain high performance and efficiency across various operating voltage conditions without significant accuracy loss.
Yiqi Wang 0005, Zhiheng Yue, Zihan Wu 0006, Huiming Han, Shaojun Wei, Yang Hu 0001, Fengbin Tu, Shouyi Yin
HPCA2
2026 MRCIM: A Many-Core Reconfigurable Computing-in-Memory Processor Combining CPU and Tensor Modes for NN Acceleration
abstract
Many-core architecture is a promising architecture to accelerate increasingly larger neural networks (NNs). Most many-core architectures couple a standalone CPU core and a tensor core together as a compute node. However, the existing architectures suffer from inefficiency at the architecture, data flow, and control flow levels: The standalone scalar CPU core with deep out-of-order pipeline and low data parallelism per instruction incurs high hardware overhead and low throughput; Fixed proportions of CPU and tensor cores execute computations alternately in each cluster, leading to core under-utilization under diverse workloads; The MIMD parallelism strategy causes redundant instruction cache (I-Cache) accesses, which increases power consumption. To tackle the above limitations, we propose MRCIM, a many-core reconfigurable computing-in-memory (CIM) processor with reconfigurable cores featuring both CPU and tensor modes. 1) We design a reconfigurable CPU core by reusing the CIM-based tensor core’s inherent memory and computing logic to simplify the pipeline logic and improve the data parallelism of conventional CPU. 2) We propose interleaved workload execution (IWE) and adaptive workload mapping (AWM) scheduling strategies, which dynamically adjust the proportion of CPU core and tensor core in a cluster, making them work in parallel with high utilization. 3) We propose a hybrid MIMD/SIMD control flow to bypass unnecessary I-Cache accesses by instruction forwarding and sharing, thereby reducing power consumption. Experimental results show MRCIM achieves 166.48x~446.67x speedup and 96.76x~309.01x energy saving over Intel i9-13900k CPU, 12.62x~27.62x speedup and 5.49x~17.82x energy saving over NVIDIA RTX 4090 GPU. Compared with state-of-the-art NN processor architectures, our MRCIM achieves average 6.84x, 7.51x, and 3.66x speedup and average 4.57x, 3.03x, and 3.11x energy saving over Simba, LUT-ICC, and MAICC.
Yiqi Wang 0005, Zihan Wu 0006, Huiming Han, Shaojun Wei, Yang Hu 0001, Chao Li 0009, Fengbin Tu, Shouyi Yin
IEEE Trans. Circuits Syst. I Regul. Pap.2
2025 ER-DCIM: Error-Resilient Digital CIM Architecture with Run-Time MAC-Cell Error Correction
abstract
Digital computing-in-memory (CIM) is an emerging solution to break through the limitations of memory wall by integrating digital logic into SRAM, which is able to achieve high area and energy efficiency with no accuracy loss. Digital CIM’s SRAM cells are still prone to errors like conventional SRAM due to noise and variation, especially under low-voltage operation for high energy efficiency. The SRAM cell errors cause multiplyaccumulation (MAC) result errors, which may seriously damage the neural network inference accuracy. However, traditional SRAM’s error correcting code (ECC) that corrects errors in one read-out row is incompatible with digital CIM, which reads out multiple rows simultaneously for computation. Detecting and correcting computational MAC errors and SRAM cell errors (i.e., MAC-cell errors) in digital CIM remain largely unexplored.To address digital CIM’s unique MAC-cell error resilience needs, we propose ER-DCIM, an error-resilient digital CIM with run-time MAC-cell error correction to guarantee computation correctness. The proposed residue code-based MAC error correction mechanism is the first to correct additive errors in the MAC result in real time during DCIM computation. Then, we propose a progressive cell error correction mechanism to correct underlying cell error in a timely manner, avoiding performance loss due to stalling computation. Further, we design a mode switcher to repurpose redundant error-resilient logic reserved for low-voltage mode to improve performance in high-voltage mode. Experimental results show that the proposed techniques enable digital CIM to maintain high throughput and energy efficiency without accuracy loss in both low-voltage and high-voltage modes.
Yiqi Wang 0005, Zihan Wu 0006, Shaojun Wei, Yang Hu 0001, Fengbin Tu, Shouyi Yin
HPCA2
2023 SDP: Co-Designing Algorithm, Dataflow, and Architecture for In-SRAM Sparse NN Acceleration
abstract
Processing-in-memory (PIM) is a promising architecture for neural network (NN) acceleration. Most previous PIMs are based on analog computing, so their accuracy and memory cell array utilization are limited by analog deviation and ADC overhead. Digital PIM is an emerging type of PIM architecture that integrates digital logic in memory cells, which can make full utilization of the cell array without accuracy loss. However, digital PIM’s rigid crossbar architecture and full array activation raise new challenges in sparse NN acceleration. Conventional unstructured or structured sparsity cannot perform well on both the weight and input side of digital PIM. We take the opportunities from digital PIM’s bit-serial processing and in-memory customization, to tackle the above challenges by the co-designing sparse algorithm, multiplication dataflow, and PIM architecture. At the algorithm level, we propose double-broadcast hybrid-grained pruning to exploit weight sparsity with better accuracy and efficiency balance. At the dataflow level, we propose a bit-serial Booth in-SRAM multiplication dataflow for stable acceleration from the input side. At the architecture level, we design a sparse digital PIM (SDP) accelerator with customized SRAM-PIM macros to support the proposed techniques. SDP achieves$3.59\times $,$8.15\times $,$3.11\times $area efficiency, and$6.95\times $,$29.44\times $,$39.40\times $energy savings, over state-of-the-art sparse NN architectures SIGMA, SRE, and Bit Prudent.
Fengbin Tu, Yiqi Wang 0005, Ling Liang 0003, Yufei Ding 0001, Leibo Liu, Shaojun Wei, Shouyi Yin, Yuan Xie 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 SPCIM: Sparsity-Balanced Practical CIM Accelerator With Optimized Spatial-Temporal Multi-Macro Utilization
abstract
Compute-in-memory (CIM) is a promising technique that reduces data movement in neural network (NN) acceleration. To achieve higher efficiency, some recent CIM accelerators exploit NN sparsity based on CIM’s small-grained operation unit (OU) feature. However, new problems arise in a practical multi-macro accelerator: The mismatch between workload parallelism and CIM macro organization causes spatial under-utilization; The multiple macros’ different computation time leads to temporal under-utilization. To solve the under-utilization problems, we propose a Sparsity-balanced Practical CIM accelerator (SPCIM), including optimized dataflow and hardware architecture design. For the CIM dataflow design, we first propose a reconfigurable cluster topology for CIM macro organization. Then we regularize weight sparsity in the OU-height pattern and reorder the weight matrix based on the sparsity ratio. The cluster topology can be reshaped to match workload parallelism for higher spatial utilization. Each CIM cluster’s workload is dynamically rebalanced for higher temporal utilization. Our hardware architecture supports the proposed dataflow with a spatial input dispatcher and a temporal workload allocator. Experimental results show that, compared with the baseline sparse CIM accelerator that suffers from spatial and temporal under-utilization, SPCIM achieves$2.94\times $speedup and$2.86\times $energy saving. The proposed sparsity-balanced dataflow and architecture are generic and scalable, which can be applied to other CIM accelerators. We strengthen two state-of-the-art CIM accelerators with the SPCIM techniques, improving their energy efficiency by$1.92\times $and$5.59\times $, respectively.
Yiqi Wang 0005, Fengbin Tu, Leibo Liu, Shaojun Wei, Yuan Xie 0001, Shouyi Yin
IEEE Trans. Circuits Syst. I Regul. Pap.1
2022 GQNA: Generic Quantized DNN Accelerator With Weight-Repetition-Aware Activation Aggregating
abstract
Quantization is a prominent approach to compress model sizes of deep neural networks (DNNs), which clusters high-precision weights into a smaller set of quantization levels and represents high-precision weights by low-precision indexes. To achieve the same accuracy, nonuniform quantized DNNs (NUQ-DNNs) with unequal quantization intervals need lower index precision than uniform quantized DNNs (UQ-DNNs) with equal intervals, achieving smaller model sizes. Hence, deploying NUQ-DNNs on accelerators costs less on- and off-chip memory accesses than UQ-DNNs, which are more valuable for edge devices. However, accelerating NUQ-DNNs is nontrivial, since weight indexes cannot be directly used for computations. Previous NUQ-DNN accelerators adopt standard convolutions by decoding weight indexes into actual-weights multiplied with activations, causing abundant look-up overhead and redundant computations. In this work, we propose a weight-repetition-aware activation aggregating (WPAA) convolution approach to accelerate inference of variable-precision NUQ- and UQ-DNNs. By merging convolutions of multiple kernels, WPAA requires no look-up operation and removes redundant computations. Based on WPAA, we design a generic quantized DNN accelerator (GQNA). Furthermore, we propose a layer-adaptive kernel-reordering merging scheme to off-line adjust merging order of kernels for minimizing energy consumption of GQNA. Implemented under TSMC 28-nm technology, GQNA achieves 31.9 and 32.6 TOPS/W energy efficiency for 1-b UQ- and NUQ-VGG-16, respectively.
Jianxun Yang, Fengbin Tu, Yiqi Wang 0005, Leibo Liu, Shaojun Wei, Shouyi Yin
IEEE Trans. Circuits Syst. I Regul. Pap.4