Zhuoyu Dai

dblp:352/9585 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2026
0009-0008-6777-9051ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 An area/energy-efficient RRAM computing-in-memory macro with fully-charge-domain multi-bit computation
Shengzhe Yan, Zhuoyu Dai, Zeyu Guo 0002, Zhaori Cong, Zhihang Qian, Xiangqu Fu, Chunmeng Dou, Dashan Shang, Jinshan Yue
Sci. China Inf. Sci.4
2025 A monolithic 3D IGZO-RRAM-SRAM-integrated architecture for robust and efficient compute-in-memory enabling equivalent-ideal device metrics
Shengzhe Yan, Zhaori Cong, Zhuoyu Dai, Zeyu Guo 0002, Zhihang Qian, Xufan Li, Chuanke Chen, Nianduan Lu, Chunmeng Dou, Guanhua Yang, Xiaoxin Xu, Di Geng, Jinshan Yue, Ling Li 0013, Ming Liu 0022
Sci. China Inf. Sci.4
2025 An RRAM-Based Computing-in-Memory Macro With Low-Power Readout/Hold Circuits and Activation Differential Strategy for AdderNet
abstract
AdderNet is an innovative neural network (NN) structure that substitutes multiplications with additions in convolutional operations, while computing-in-memory (CIM) is an efficient architecture that tackles the memory bottleneck for von Neumann architectures. Previous work has explored the SRAM-based CIM AdderNet circuits and demonstrates high energy efficiency. However, it still suffers low storage density, repetitive readout, and redundant comparisons. In this brief, an RRAM-based CIM macro is proposed for efficient AdderNet with the following innovations. First, RRAM cells are adopted to replace SRAM for high-density weight storage. A low-power readout and hold circuit is proposed to save redundant read power of weight data held for multiple cycles. Second, an 8-bit comparator with an early-stop strategy is proposed to compare 8-bit activations and weights in one cycle. Third, an activation (ACT) differential strategy is proposed to reduce redundant comparisons. The proposed 28-nm RRAM CIM macro achieves 12.8-TOPS/mm2peak area efficiency and 126-TOPS/W peak energy efficiency, which is$3.0\times $and$1.2\times $compared with the state-of-the-art AdderNet CIM macro.
Zhihang Qian, Shengzhe Yan, Zhuoyu Dai, Zeyu Guo 0002, Zhaori Cong, Yifan He 0003, Chunmeng Dou, Feng Zhang 0014, Jinshan Yue, Yongpan Liu
IEEE Trans. Very Large Scale Integr. Syst.3
2025 A High-Density Energy-Efficient CNM Macro Using Hybrid RRAM and SRAM for Memory-Bound Applications
abstract
The big data era has facilitated various memory-centric algorithms, such as the Transformer decoder, neural network, stochastic computing (SC), and genetic sequence matching, which impose high demands on memory capacity, bandwidth, and access power consumption. The emerging nonvolatile memory devices and compute-near-memory (CNM) architecture offer a promising solution for memory-bound tasks. This work proposes a hybrid resistive random access memory (RRAM) and static random access memory (SRAM) CNM architecture. The main contributions include: 1) proposing an energy-efficient and high-density CNM architecture based on the hybrid integration of RRAM and SRAM arrays; 2) designing low-power CNM circuits using the logic gates and dynamic-logic adder with configurable datapath; and 3) proposing a broadcast mechanism with output-stationary workflow to reduce memory access. The proposed RRAM-SRAM CNM architecture and dataflow tailored for four distinct applications are evaluated at a 28-nm technology, achieving 4.62-TOPS$/$W energy efficiency and 1.20-Mb$/$mm2memory density, which shows$11.35\times $–$25.81\times $and$1.44\times $–$4.92\times $improvement compared to previous works, respectively.
Shengzhe Yan, Xiangqu Fu, Zhihang Qian, Zhi Li 0062, Zeyu Guo 0002, Zhuoyu Dai, Zhaori Cong, Chunmeng Dou, Feng Zhang 0014, Jinshan Yue, Dashan Shang
IEEE Trans. Very Large Scale Integr. Syst.7
2024 IG-CRM: Area/Energy-Efficient IGZO-Based Circuits and Architecture Design for Reconfigurable CIM/CAM Applications
abstract
Artificial intelligence is evolving with various algorithms such as deep neural network (DNN), Transformer, recommendation system (RecSys) and graph convolutional network (GCN). Correspondingly, multiply-accumulate (MAC) and content search are two main operations, which can be efficiently executed on the emerging computing-in-memory (CIM) and content-addressable-memory (CAM) paradigms. Recently, the emerging Indium-Gallium-Zine-Oxide (IGZO) transistor becomes a promising candidate for both CIM/CAM circuits, featuring ultra-low leakage with >300s data retention time and high-density BEOL fabrication. This paper proposes IG-CRM, the first IGZO-based circuits and architecture design for Reconfigurable CIM/CAM applications. The main contributions include: 1) at cell level, propose IGZO-based 3T0C/4T0C cell design that enables both CIM and CAM functionalities while matching IGZO/CMOS voltage; 2) at circuit level, utilize the BEOL IGZO transistor to reduce digital adder tree area in CIM circuits; 3) at architecture level, propose a reconfigurable CIM/CAM architecture with four macro structures based on 3T0C/4T0C cells. The proposed IG-CRM architecture shows high area/energy efficiency on various applications including DNN, Transformer, RecSys and GCN. Experiment results show that IG-CRM achieves 8.09X area saving compared with the SRAM-based non-reconfigurable CIM/CAM baseline, and 1.53×103X/51.9X speedup and 1.63×104X/7.62×103X energy efficiency improvement compared with CPU and GPU on average.
Zeyu Guo 0002, Jinshan Yue, Shengzhe Yan, Zhuoyu Dai, Xiangqu Fu, Zhaori Cong, Zening Niu, Lihua Xu, Guanhua Yang, Di Geng, Ling Li 0013
DAC4
2024 A Multichiplet Computing-in-Memory Architecture Exploration Framework Based on Various CIM Devices
abstract
Computing-in-memory (CIM) architectures based on various devices, such as resistive random access memory, SRAM, DRAM, etc., have demonstrated promising energy efficiency. Single-device-based CIM chips show different advantages on performance, power, or area metrics under different workload/operators sizes and application requirements. Some nonidealities, such as the write endurance of some nonvolatile devices, also influence the design choices. Motivated by the emerging 2.5-D/3-D chiplet integration, this work aims to combine the advantages of CIM/storage chips based on different devices, and proposes a design exploration framework to combine the advantages of CIM chips based on these devices in a 3-D-stack architecture. This work proposes: 1) an evaluation method for the power, performance, and area metrics of the multichiplet CIM architecture; 2) an abstraction for the single-device-based CIM chiplets and artificial intelligence algorithm operators; and 3) a mapping and optimization strategy to explore the 2.5-D/3-D CIM chiplet set. The effectiveness of the mapping strategy is verified with a small-scale brute-force search. The proposed design exploration framework can help to find a better-multichiplet CIM architecture. Under a simple design case, the proposed 3-D CIM architecture shows$4.68\times $–$53.32\times $energy efficiency compared with the single CIM chip baselines. The abstracted chiplet library is open-source available in the open-sourcehttps://github.com/dai0dai/3D_CIM_Chiplet_Architecture_Exploration.
Zhuoyu Dai, Feibin Xiang, Xiangqu Fu, Yifan He 0003, Wenyu Sun, Yongpan Liu, Guanhua Yang, Feng Zhang 0014, Jinshan Yue, Ling Li 0013
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 A User-Friendly Fast and Accurate Simulation Framework for Non-Ideal Factors in Computing-in-Memory Architecture
abstract
Computing-in-memory (CIM) architecture utilizing emerging non-volatile devices is promising for energy-efficient neural network (NN) applications. However, the non-ideal factors of non-volatile devices and analog circuits may incur severe accuracy loss, which cuts the algorithm and hardware design apart. The algorithm/hardware designers are skilled in either macro-scope NN models or detailed circuit/device errors, while sophisticated research of the joint effect on accuracy loss is urgently needed. In this paper, we propose a user-friendly, fast, and accurate simulation framework (CIMUFAS) to explore the impact of various non-ideal devices/circuits on algorithm accuracy. Based on multiple architecture-level CIM mapping/scheduling workflow, sophisticated non-ideal factors with different error models are established. The CIMUFAS also provides easy-to-use interfaces to flexibly support user-defined models/parameters for specified devices/circuits. Besides, the CIMUFAS framework achieves reasonable simulation time. Compared with MNSIM 2.0, the simulation time is reduced by 45% even after adding a more realistic hardware configuration. This CIMUFAS framework is verified with two fabricated CIM chips with <0.04% accuracy mismatch. The source code of CIMUFAS is publicly available at https://github.com/Hlal/CIMUFAS.
Jinshan Yue, Chaojie He, Zhuoyu Dai, Feibin Xiang, Zhaori Cong, Yifan He 0003, Xiaoyu Feng, Yongpan Liu
ISCAS4