EDBT 2026 Demo / reviewers in the wild / expert
Chenyang Zhao 0008
dblp:22/8876-8
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-5054-8141ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Light-CIM: A Lightweight ADC/DAC-Fewer RRAM CIM DNN Accelerator With Fully Analog Tiles and Nonideality-Aware Algorithm for Consumer ElectronicsabstractNeuromorphic computing has emerged as a revolutionary technology in consumer electronics, with computing-in-memory (CIM) attracting considerable attention for its potential to minimize data transfer. However, most CIM accelerators necessitate numerous digital-to-analog converters (DACs) and analog-to-digital converters (ADCs) for mixed-signal data processing, resulting in substantial area and energy overheads. This study introduces a lightweight CIM accelerator, Light-CIM, which operates with fully analog tiles (FANTs) and employs a nonideality-aware algorithm. A FANT consists of one-transistor-one-resistor (1T1R) arrays based on resistive random access memory (RRAM) and customized analog peripheral circuits for data processing. The intratile data computation, transfer, and buffering are all in analog voltage, current, or RRAM resistance, thus eliminating costly DACs and ADCs for intermediate data conversions in conventional CIM accelerators. The fully analog approach significantly reduces power consumption attributed to ADCs, accounting for only 2.5% of the total power consumption. Additionally, a nonideality-aware training algorithm is employed to enhance the robustness of the hardware system. It models and incorporates nonidealities of circuits in software training, including read nonlinearities, mismatches, variations, and noises in the hardware analog data flow. Experimental results demonstrate that Light-CIM achieves accuracy close to software performance in various NN models. Light-CIM accomplishes a compute density of 3.91 TOPS/${\mathrm { mm}}^{2}$and an energy efficiency of 3.08 TOPS/W, both highly competitive compared to state-of-the-art works. Chenyang Zhao 0008, Jinbei Fang, Xiaoyong Xue, Xiaoyang Zeng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | A Heuristic and Greedy Weight Remapping Scheme with Hardware Optimization for Irregular Sparse Neural Networks Implemented on CIM Accelerator in Edge AI ApplicationsabstractComputing-in-memory (CIM) is a promising technique for hardware acceleration of neural networks (NNs) with high performance and efficiency. However, conventional dense mapping scheme cannot well support the compression and optimization of irregular sparse NNs. In this paper, we propose a heuristic and greedy weight remapping scheme for irregular sparse neural networks implemented on CIM accelerator in edge AI applications. The genetic algorithm (GA) is proposed for the first time to be utilized in the column shuffle for sparse weight remapping. Combined with the granularity exploration of the CIM, the proportion of the compressible all-zero rows increase remarkably. A greedy algorithm is then employed to planarize the unevenly compressed units, thus to improve the storage utilization of the crossbar. For hardware optimization, the pipeline is customized with a zero-skipping circuit to leverage the bit-level activation sparsity at runtime. Our results show that the proposed remapping scheme achieves 70%-94% utilization rate of the sparsity, and an average of $1.3 \times$ increment compared with the naive compression. The cooptimized CIM achieves $3-7.6 \times$ speedup and $2.1- 4.8 \times$ energy efficiency, compared with the baseline for dense NNs. Lizhou Wu, Chenyang Zhao 0008, Xueru Yu, Shoumian Chen, Jun Han 0003, Xiaoyong Xue, Xiaoyang Zeng |
ASPDAC | 2 |
| 2023 | ARBiS: A Hardware-Efficient SRAM CIM CNN Accelerator With Cyclic-Shift Weight Duplication and Parasitic-Capacitance Charge Sharing for AI Edge ApplicationabstractComputing-in-memory (CIM) relieves the Von Neumann bottleneck by storing the weights of neural networks in memory arrays. However, two challenges still exist, hindering the efficient acceleration of convolutional neural networks (CNN) in artificial intelligence (AI) edge devices. Firstly, the activations for sliding window (SW) operations in CNN still bring high memory access pressure. This can be alleviated by increasing the SW parallelism, but simple array replication suffers from poor array utilization and large peripheral circuits overhead. Secondly, the partial sums from individual CIM arrays, which are usually accumulated to obtain the final sum, introduce large latency due to enormous shift-and-add operations. Moreover, high-resolution ADCs are also needed to reduce the quantization error of partial sums, further increasing the hardware costs. In this paper, a hardware-efficient CIM accelerator, ARBiS, is proposed with improved activation reusability and bit-scalable matrix-vector-multiplication (MVM) for CNN acceleration in AI edge applications. The cyclic-shift weight duplication exploits a third dimension of receptive field (RF) depth for SW weight mapping to reduce the memory accesses of activations, improving the array utilization. The parasitic-capacitance charge sharing is employed to realize high-precision analog MVM in order to reduce the ADC cost. Compared with conventional architectures, ARBiS with parallel processing of 9 SW operations achieves 56.6%~58.8% alleviation of memory access pressure. Meanwhile, ARBiS configured with 8-bit ADCs saves 92.53%~94.53% ADC energy consumption. An ARBiS accelerator is evaluated to realize a computational efficiency (CE) of 10.28 (10.43) TOPS/mm2, an energy efficiency (EE) of 91.19 (112.36) TOPS/W with 8-bit (4-bit) ADCs, achieving$11.4\sim 11.7\times $($11.6\sim 11.8\times $),$1.1\sim 3.3\times $($1.4\sim 4\times $) improvements over state-of-the-art works, respectively. Chenyang Zhao 0008, Jinbei Fang, Xiaoyong Xue, Xiaoyang Zeng |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |