EDBT 2026 Demo / reviewers in the wild / expert
Heng Zhang 0024
dblp:55/826-24
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2025
0009-0001-7827-7849ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Corrections to "An Efficient Two-Stage Pipelined Compute-in-Memory Macro for Accelerating Transformer Feed-Forward Networks"abstractIn the above article [1], the die photograph on the right side of original Fig. 9 was inadvertently mirrored horizontally, as shown in Fig. 1. This occurred during the annotation process, where the image used had already been flipped without our awareness. As a result, the internal layout labeling (e.g., CIMA1, CIMA2, and ADC) appeared in reverse orientation relative to the actual die.Fig. 1.Difference clarification between the original Fig. 9 of our published article and the revised Fig. 9. Fig. 9.Die photograph and measure setup for the proposed chip. Heng Zhang 0024, Wenhe Yin, Sunan He, Yuan Du |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2024 | A Compilation Framework for SRAM Computing-in-Memory Systems With Optimized Weight Mapping and Error CorrectionabstractDeploying convolution-based algorithms into SRAM computing-in-memory (CIM) systems faces various challenges, such as operator incompatibility and intrinsic non-ideal error. This paper proposes a compilation framework to address this issue. Efficient weight mapping strategies are introduced to improve the utilization of SRAM-CIM macro. The intrinsic non-ideal errors of SRAM-CIM macro are also taken into consideration, and two efficient error correction schemes are proposed, which include calibration of computation voltage linear error (CCVLE) and the mitigation of analog-to-digital quantization error (MAQE). In addition, bit-width flexibility and signed-unsigned reconfigurability are also supported to facilitate the deployment of various convolution-based algorithms. ResNet18, finite impulse response (FIR) filtering, and Gaussian image filtering are deployed into a multi-macro SRAM-CIM system. These algorithms serve as deployment representatives of convolutional neural network (CNN), digital signal processing (DSP), and digital image processing (DIP), respectively. The results show that the introduced weight mapping strategies improve the macro utilization by 63.29% and 21.10% for two types of frequently used convolution layers compared to the traditional strategy. Moreover, the proposed error correction schemes achieve similar algorithm accuracy to the floating-point results, and the deployment result of ResNet18 achieves 66.3%~70.1% top-1 classification accuracy evaluated on the ImageNet dataset with different throughput tradeoffs. Yichuan Bai, Yaqing Li, Heng Zhang 0024, Aojie Jiang, Yuan Du |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | An 11T1C Bit-Level-Sparsity-Aware Computing- in-Memory Macro With Adaptive Conversion Time and Computation VoltageabstractA static random-access memory (SRAM)-based computing-in-memory (CiM) is a promising architecture for efficiently performing high-precision integer (INT) multiplication and accumulation (MAC) operations. In this work, we propose a charge-domain bit-level-sparsity-aware analog CiM (ACiM) macro for an area-energy-efficient convolutional neural network (CNN). An 11T1C ACiM bit-cell is proposed to dynamically remove the computation capacitors during the accumulation phase based on the weight value (W) for improving the partial sums and analog computing accuracy margin (ACAM). The computation voltage is dynamically adjusted according to column-wise sparsity by the bit-level-sparsity-aware controller to improve energy efficiency. To digitize the MAC computing results, a 2-8bit column-parallel time-interleaved hybrid analog-to-digital converter (ADC) is designed by sharing the voltage reference generator, which achieves a low unit pitch size. A$256\times 64~11$T1C ACiM macro prototype with hybrid ADCs is implemented using 55nm CMOS process. The silicon measurement results show that the proposed ACiM achieves a throughput of 51.2-153.6 GOPS, core area efficiency reaching 112-336GOPS/mm2, and energy efficiency ranging from 17 to 111 TOPS/W with 8bit weights and 8bit inputs. Yuandong Li, Heng Zhang 0024, Jingjing Lv, Anying Jiang, Yuan Du |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | An Efficient Two-Stage Pipelined Compute-in-Memory Macro for Accelerating Transformer Feed-Forward NetworksabstractTransformer architectures have achieved state-of-the-art performance in various applications. However, deploying transformer models on resource-constrained platforms is still challenging due to its dynamic workloads, intensive computations, and substantial memory access. In this article, we propose a two-stage pipelined compute-in-memory (CIM) macro for effectively deploying and accelerating the feed-forward network (FFN) layers of transformer models. Two independent CIM arrays are designed to execute the two distinct linear projections in FFN layers, which are interconnected by co-designed analog rectified linear unit (ReLU) circuits to realize the nonlinear activation function. The analog multiply-and-add (MAC) results from the first CIM array are streamed directly to the analog ReLU circuits, and subsequently to the next CIM array for performing another linear projection. This architecture eliminates the need for analog-to-digital converters (ADCs) and digital-to-analog converters (DACs) for internal results’ staging, thereby enhancing overall macro efficiency and reducing computing latency. A proof-of-concept macro is fabricated using TSMC 65-nm process and achieves 4.096 TOPS peak throughput, 4.39 TOPS/mm2 area efficiency, and 49.83 TOPS/W energy efficiency. To map transformer models onto the proposed macro, we quantize the FFN layers of BERTMINI model under per-token granularity for activations and per-tensor granularity for weights using quantization-aware training (QAT), which exhibits excellent accuracy across multiple benchmarks. Heng Zhang 0024, Wenhe Yin, Sunan He, Yuan Du |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2023 | SSM-CIM: An Efficient CIM Macro Featuring Single-Step Multi-bit MAC Computation for CNN Edge InferenceabstractCompute-in-memory (CIM) is a promising approach to solving the memory-wall problem existing in traditional computing architectures. In this paper, we introduce SSM-CIM, a charge-domain, static random-access memory (SRAM)-based CIM macro designed for area-energy-efficient convolutional neural network (CNN) inference. SSM-CIM utilizes an original sign-magnitude data encoding method for both inputs and weights. By codesigning four adjacent SRAM computing cells and employing a 3-bit digital-to-analog converter (DAC), SSM-CIM performs accurate 4-bit multiply-and-accumulate (MAC) computation in a single step, eliminating the peripheral digital shift-and-add circuits. To digitize the MAC computing results, a dedicated multi-reference assisted SAR ADC is designed by reusing the reference voltages from the DAC, which offers significant power and area savings. In addition, analog computing errors and quantization errors are analyzed to ensure the multi-bit computing accuracy of SSM-CIM. SSM-CIM is implemented and evaluated using 28-nm global foundry process. The post-layout simulation results validate the excellent computing linearity and accuracy of SSM-CIM. Benefitting from the compact layout design and fully parallel computing flow, the$144\times 256$macro achieves a peak throughput of 2.3 TOPS, an area efficiency of 10.2 TOPS/mm2, and an energy efficiency of 205.4 TOPS/W with 4-bit weights and 4-bit inputs. Heng Zhang 0024, Sunan He, Xinjie Guo, Shaodi Wang, Yuan Du |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |