Kaining Zhou

dblp:298/1240 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 61% Memory systems · 30% Hardware accelerators and domain-specific architectures · 9%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation › simulation › architectural simulation
full-system simulation
0.912025
A Full-system, Programmable, and Extensible In-Memory Computing Simulation Framework for Deep Learning · DAC 2025
Memory systems
processing-in-memory
0.912025
A Full-system, Programmable, and Extensible In-Memory Computing Simulation Framework for Deep Learning · DAC 2025
Performance modeling and evaluation › simulation › simulation software
simulation framework
0.912025
A Full-system, Programmable, and Extensible In-Memory Computing Simulation Framework for Deep Learning · DAC 2025
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.312025
A Full-system, Programmable, and Extensible In-Memory Computing Simulation Framework for Deep Learning · DAC 2025

Methods — techniques the papers use, named apart from their topics

tensor operator mapping · 0.9device-circuit modeling · 0.9ISA extension · 0.9
YearPublicationVenuePosition
2025 A Full-system, Programmable, and Extensible In-Memory Computing Simulation Framework for Deep Learning
abstract
In-memory computing (IMC) has established itself as an attractive alternative to hardware accelerators in addressing the memory wall problem for artificial intelligence (AI) workloads. However, designing programmable IMC-based computing platforms for today’s large generative AI models, such as large language models (LLMs) and diffusion transformers (DiTs), is hindered by the absence of a simulator that is able to address the associated scalability challenges while simultaneously incorporating device and circuit-level behaviors intrinsic to IMCs. To address this challenge, we present IMCsim, a versatile fullsystem IMC simulation framework. IMCsim integrates software runtime libraries for AI models, introduces a new set of ISA extensions to express common tensor operators, and provides flexibility in mapping these operators to various IMC architectures. As such, IMCsim enables designers to explore trade-offs between performance, energy, area, and computational accuracy for various IMC design choices. To demonstrate the functionality, efficiency, and versatility of IMCsim, we model three types of IMCs: (1) embedded non-volatile memory (eNVM)-based, (2) SRAM-based, and (3) digital IMCs. We validate IMCsim using measured data from two laboratory-tested IMC prototype ICs—a 22 nm MRAM-based IMC and a 28 nm SRAM-based IMC—and a digital IMC design in 28 nm. Next, we demonstrate the utility of IMCsim by exploring the architectural design space to obtain insights for maximizing utilization of IMC-based processors for diverse workloads—ResNet-18, Llama, and a DiT—using the three IMC types. Finally, we employ IMCsim as a design tool to obtain an efficient chip architecture and layout in 28 nm for a lightweight DiT.
Kaining Zhou, Jian Huang 0006, Nam Sung Kim, Naresh Shanbhag
DAC1