Junyi Zhu 0016

dblp:414/4136 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Processor architecture and microarchitecture · 56% Memory systems · 44%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › processing-in-memory › computing-in-memory
in-SRAM computing
1.012026
Enabling Efficient Vector Processing: A Heterogeneous Vector Architecture with in-SRAM Computing · ACM Trans. Archit. Code Optim. 2026
Memory systems
processing-in-memory
1.012026
Enabling Efficient Vector Processing: A Heterogeneous Vector Architecture with in-SRAM Computing · ACM Trans. Archit. Code Optim. 2026
Processor architecture and microarchitecture › vector processor
vector processing unit
1.012026
Enabling Efficient Vector Processing: A Heterogeneous Vector Architecture with in-SRAM Computing · ACM Trans. Archit. Code Optim. 2026
Processor architecture and microarchitecture
vector processor
1.012026
Enabling Efficient Vector Processing: A Heterogeneous Vector Architecture with in-SRAM Computing · ACM Trans. Archit. Code Optim. 2026
Processor architecture and microarchitecture
instruction set architecture
0.312026
Enabling Efficient Vector Processing: A Heterogeneous Vector Architecture with in-SRAM Computing · ACM Trans. Archit. Code Optim. 2026
Processor architecture and microarchitecture › instruction set architecture
ISA extension
0.312026
Enabling Efficient Vector Processing: A Heterogeneous Vector Architecture with in-SRAM Computing · ACM Trans. Archit. Code Optim. 2026

Methods — techniques the papers use, named apart from their topics

vector batch pipelining · 1.0architectural co-design · 1.0
YearPublicationVenuePosition
2026 Enabling Efficient Vector Processing: A Heterogeneous Vector Architecture with in-SRAM Computing
abstract
The growing demand for high-performance, energy-efficient execution of data-parallel workloads has driven the resurgence of vector processors, yet their expanding instruction sets exacerbate the area-performance tradeoff of vector processing units (VPUs). Processing-in-memory (PIM) technique offers a promising path to mitigate this tradeoff by offloading vector operations near data. However, integrating the computing-capable SRAM (C-SRAM) with conventional VPUs introduces significant architectural challenges, including inefficient coordination between heterogeneous devices, the lack of a unified hardware/software interface, and underutilized parallelism within the C-SRAM arrays due to unoptimized data handling. To address these challenges, this article proposes a heterogeneous VPU (HVPU) that seamlessly integrates a standard VPU inside a vector processor with a C-SRAM for more efficient vector processing. HVPU introduces a standardized interface between the processor frontend and the C-SRAM, enabling instruction dispatch, dynamic hazard resolution, and concurrent execution. Furthermore, it employs multiple independent PIM blocks coupled with dual controlling pipelines inside the C-SRAM for further performance improvement. This design facilitates pipelined data loading and computing, effectively hiding memory latency and fully unlocking the parallel potential of C-SRAM. The system is supported by a user-friendly and generic programming model featuring a two-layer extended ISA system and a vector batch pipelining mechanism. Experimental results show that our HVPU-enhanced processor achieves significant speedups of 5.11× to 39.0× over the Xuantie-910 baseline on vector benchmarks, respectively, while reducing energy consumption by 75% on average. Meanwhile, it outperforms state-of-the-art PIM accelerators by 1.33× to 1.97× on the same benchmarks, with minimal area overhead. This demonstrates that our architectural co-design effectively alleviates the area-performance tradeoff in vector processors and offers a scalable heterogeneous architecture template for efficient vector processing.
Dunbo Zhang, Shangshang Yao, Qingjie Lang, Junyi Zhu 0016, Li Shen 0007
ACM Trans. Archit. Code Optim.5