EDBT 2026 Demo / reviewers in the wild / expert
Zhi Li 0062
dblp:43/3166-62
· DBLP profile ↗
7ranked-venue papers
0as first author
7since 2021 · last 2026
0009-0000-9649-7338ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A BEOL Ferroelectric FET-based Computing Unit for Digital Computing-in-Memory
Junyu Zhu, Zexue Bian, Weizeng Li, Junzhe Shen, Hanghang Gao, Zhidao Zhou, Zhongze Han, Zhi Li 0062, Hongyang Hu, Chunmeng Dou |
ISCAS | 8 |
| 2025 | An energy-efficient FeFET-based computing-in-memory macro using BEOL-integrated HZO ferroelectric capacitors
Weizeng Li, Zhidao Zhou, Linfang Wang, Junyu Zhu, Junzhe Shen, Hongyang Hu, Baihan Wang, Zhi Li 0062, Wang Ye, Zhongze Han, Hanghang Gao, Chunmeng Dou |
Sci. China Inf. Sci. | 8 |
| 2025 | SHMT: An SRAM and HBM Hybrid Computing-in-Memory Architecture With Optimized KV Cache for Multimodal TransformerabstractMultimodal Transformer (MMT) algorithms have become the state-of-the-art for multimodal tasks such as image captioning. The Encoder-Decoder (E-D) structure, consisting of Encoder, Decoder-causal, and Decoder-cross components, provides a flexible and effective framework for multimodal tasks. However, previous accelerators mainly focus on the dataflow and hardware optimization of the Encoder, which fails to accelerate the entire E-D structure efficiently. There remain three challenges: 1) the lack of pipeline and multicore optimization at the module, layer, and E-D level; 2) the Decoder-causal and Decoder-cross computations have lower arithmetic intensity compared to the Encoder, requiring a better solution for the varying arithmetic intensities; and 3) the autoregressive algorithm in Decoder-causal leads to redundant KV Cache accesses and considerable idle power. In this paper, SHMT, an SRAM and HBM hybrid computing-in-memory (CIM) architecture, is designed to efficiently support multimodal Transformers with three key contributions: 1) a multi-level pipelined multicore scheme, including pipeline optimization across E-D layer-head-module levels and a multicore network-on-chip (NoC) architecture, to reduce inference latency and off-chip accesses; 2) a heterogeneous SRAM-HBM architecture, utilizing high-density HBM-CIM for low-arithmetic-intensity (LAI) parts and high-performance SRAM-CIM for high-arithmetic-intensity (HAI) parts; and 3) by integrating KV Cache with zero-padding in SRAM-CIM, SHMT eliminates redundant read-write operations in KV Cache, reducing idle power consumption. Experiment results show that SHMT achieves 212× speedup, reduces energy consumption by 208×~2000× per token, and achieves 13.3× higher energy efficiency compared to NVIDIA A100 GPU. Xiangqu Fu, Jinshan Yue, Muhammad Faizan, Zhi Li 0062, Qiang Huo, Feng Zhang 0014 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | A High-Density Energy-Efficient CNM Macro Using Hybrid RRAM and SRAM for Memory-Bound ApplicationsabstractThe big data era has facilitated various memory-centric algorithms, such as the Transformer decoder, neural network, stochastic computing (SC), and genetic sequence matching, which impose high demands on memory capacity, bandwidth, and access power consumption. The emerging nonvolatile memory devices and compute-near-memory (CNM) architecture offer a promising solution for memory-bound tasks. This work proposes a hybrid resistive random access memory (RRAM) and static random access memory (SRAM) CNM architecture. The main contributions include: 1) proposing an energy-efficient and high-density CNM architecture based on the hybrid integration of RRAM and SRAM arrays; 2) designing low-power CNM circuits using the logic gates and dynamic-logic adder with configurable datapath; and 3) proposing a broadcast mechanism with output-stationary workflow to reduce memory access. The proposed RRAM-SRAM CNM architecture and dataflow tailored for four distinct applications are evaluated at a 28-nm technology, achieving 4.62-TOPS$/$W energy efficiency and 1.20-Mb$/$mm2memory density, which shows$11.35\times $–$25.81\times $and$1.44\times $–$4.92\times $improvement compared to previous works, respectively. Shengzhe Yan, Xiangqu Fu, Zhihang Qian, Zhi Li 0062, Zeyu Guo 0002, Zhuoyu Dai, Zhaori Cong, Chunmeng Dou, Feng Zhang 0014, Jinshan Yue, Dashan Shang |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | An RRAM Digital Computing-in-Memory Macro With Dual-Mode Multiplication and Maximum Value Rounding Adder TreeabstractImplementing digital computing-in-memory (DCIM) based on resistive memory (RRAM) faces several critical challenges due to the small signal margin, large device variations, and large energy- and area-overhead induced by the digital adder tree (AT). To address these issues, we propose an RRAM DCIM macro based on the standard foundry one-transistor-one-resistor (1T1R) cell array featuring: 1) dual-mode MAC operation for efficiency- or accuracy-oriented optimization; 2) margin-enhanced digitized unit (MEDU) to amplify the signal ratio; and 3) maximum value rounding AT (MVR-AT) to reduce its power- and area-overhead. A test chip is demonstrated using a 180 nm CMOS process to verify the concept. It achieves a peak energy efficiency (EF) of 63.08 TOPS/W in the efficiency-oriented mode and a minimum error rate of 1.58% in the accuracy-oriented mode. Their combination can meet the requirements of different workloads in AI computing tasks to optimize the overall power consumption with negligible accuracy loss. Wang Ye, Hanghang Gao, Zhidao Zhou, Linfang Wang, Weizeng Li, Zhi Li 0062, Jinshan Yue, Xiaoxin Xu, Hongyang Hu, Chunmeng Dou |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2024 | A 2T P-Channel Logic Flash Cell for Reconfigurable Interconnection in Chiplet-Based Computing-In-Memory AcceleratorsabstractIn this work, we propose a two-transistor (2T) p-type channel (p-channel) logic-compatible flash cell. Compared to the previous designs, the proposed structure features reduced area-cost and enhanced ability to pass through the logic ‘1’. Due to these advantages, we explore its application as the reconfigurable interconnections in the chiplet-based system. By integrating them into the silicon interposer, the 2T p-channel flash cells can potentially lead to the dense and flexible interconnection between multiple computing-in-memory (CIM) chiplets, resulting in highly reconfigurable and scalable chiplet-based CIM accelerators. A 180nm 1Kb 2T p-channel flash cell array is fabricated and characterized. The characterization results show the 2T p-channel flash cells exhibit a signal ratio >103over 1000 program/erase (P/E) cycles and the device-to-device variations are less than 21.07%. Their typical behaviors as routers are also confirmed by circuit simulations. Weizeng Li, Linfang Wang, Zhi Li 0062, Wang Ye, Zhidao Zhou, Haiyang Zhou, Hanghang Gao, Jinshan Yue, Hongyang Hu, Fengman Liu, Chunmeng Dou |
ISCAS | 3 |
| 2024 | Write-Verify-Free MLC RRAM Using Nonbinary Encoding for AI Weight Storage at the EdgeabstractHigh-density and reliable multilevel-cell (MLC) resistive random access memory (RRAM) is expected to meet the ever-increasing demand for on-chip weight storages in the intelligent edge devices. However, due to the device variations, many write-and-verify (WAV) iterations are usually required to program the RRAM cell, which causes high power consumption, long latency, and degradation on the memory lifetime. To address this issue, we propose a write–verify-free MLC RRAM macro for weight storage with 1) a cascode-current-mirror multibit write (CCM-MW) driver and 2) a nonbinary programming scheme (NB-PS) with a radix not greater than 2. A 180-nm 400-Kb RRAM test chip is demonstrated in silicon. For 2-bit-per-cell MLC storage, the value error rates can be reduced by 24.13% after introducing two redundant bits (RBDs). In addition, compared to the single-level cell (SLC) storage scheme, a 37.50% reduction in the number of cells can be achieved to store the ResNet-8 model with a 0.79% loss in inference accuracy without the need for WAV iterations. Junjie An, Zhidao Zhou, Linfang Wang, Wang Ye, Weizeng Li, Hanghang Gao, Zhi Li 0062, Jinghui Tian, Hongyang Hu, Jinshan Yue, Lingyan Fan, Shibing Long, Qi Liu 0010, Chunmeng Dou |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |