Jinsen Zhu

dblp:437/0278 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0003-4370-6273ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 62% Memory systems · 38%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
dataflow optimization
1.012026
PIMapping: A Tile-Level Dataflow Optimization Framework for PIM Architecture · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2026
Memory systems
processing-in-memory
1.012026
PIMapping: A Tile-Level Dataflow Optimization Framework for PIM Architecture · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2026
Hardware accelerators and domain-specific architectures
matrix-vector multiplication
0.312026
PIMapping: A Tile-Level Dataflow Optimization Framework for PIM Architecture · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2026
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network acceleration
0.312026
PIMapping: A Tile-Level Dataflow Optimization Framework for PIM Architecture · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2026

Methods — techniques the papers use, named apart from their topics

dataflow analysis · 1.0congestion-aware scheduling · 1.0
YearPublicationVenuePosition
2026 PIMapping: A Tile-Level Dataflow Optimization Framework for PIM Architecture
abstract
Process-in-memory (PIM) accelerators demonstrate outstanding performance in accelerating matrix-vector multiplication (MVM) tasks in neural networks. To achieve better acceleration performance, extensive research has focused on the design of hierarchical tile-based PIM architectures, presenting challenges for the hardware deployment of algorithms. Multilayer parallelism in tile-based architectures requires the support of dataflow optimization techniques. However, existing research primarily performs dataflow analysis using computational models and performance metrics that are not suitable for tile-level dataflows. In this paper, we propose an analytical framework, PIMapping, for tile-based PIM architectures that supports multi-layer parallelism mapping and tile-level dataflow optimization. In this work, we first establish a general dataflow representation for tile-level dataflow, which serves as an intermediate representation (IR) for multi-layer DNN mapping and scheduling tasks at tile level. Next, based on the proposed representation, we introduce a data proximity-based mapping method aimed at minimizing the inter-tile communication overhead. Furthermore, we propose a congestion-aware scheduling algorithm to minimize inter-tile communication conflicts. Experimental case-studies are conducted to map common DNN algorithms onto different PIM architectures. The results demonstrate significant improvements over the state-of-the-art mapping framework in terms of communication latency, required inter-tile bandwidth, and pipeline efficiency.
Ziqian Zhu, Jinsen Zhu, Hongbing Pan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3