Chenhao Ma 0006

dblp:400/4804 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 53% Memory systems · 47%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
neural processing unit
0.912025
NVR: Vector Runahead on NPUs for Sparse Memory Access · DAC 2025
Memory systems
cache
0.312025
NVR: Vector Runahead on NPUs for Sparse Memory Access · DAC 2025
Memory systems › cache
cache miss reduction
0.312025
NVR: Vector Runahead on NPUs for Sparse Memory Access · DAC 2025
Memory systems › cache
prefetching
0.312025
NVR: Vector Runahead on NPUs for Sparse Memory Access · DAC 2025

Methods — techniques the papers use, named apart from their topics

speculative execution · 0.9runahead execution · 0.9
YearPublicationVenuePosition
2026 Joint learning video segmentation with different prior guidance
Chenhao Ma 0006, Xiuhui Deng, Jason Junwei Zeng, Jianning Zhang, Zhe Jiang 0004, Ying Huo
J. Syst. Archit.2
2025 NVR: Vector Runahead on NPUs for Sparse Memory Access
abstract
Deep Neural Networks are increasingly leveraging sparsity to reduce the scaling up of model parameter size. However, reducing wall-clock time through sparsity and pruning remains challenging due to irregular memory access patterns, leading to frequent cache misses. In this paper, we present NPU Vector Runahead (NVR), a prefetching mechanism tailored for NPUs to address cache miss problems in sparse DNN workloads. Rather than optimising memory patterns with high overhead and poor portability, NVR adapts runahead execution to the unique architecture of NPUs. NVR provides a general micro-architectural solution for sparse DNN workloads without requiring compiler or algorithmic support, operating as a decoupled, speculative, lightweight hardware sub-thread alongside the NPU, with minimal hardware overhead (under 5%). NVR achieves an average 90% reduction in cache misses compared to SOTA prefetching in general-purpose processors, delivering 4 x average speedup on sparse workloads versus NPUs without prefetching. Moreover, we investigate the advantages of incorporating a small cache (16 KB) into the NPU combined with NVR. Our evaluation shows that expanding this modest cache delivers 5x higher performance benefits than increasing the $\mathbf{L 2}$ cache size by the same amount.
Hui Wang 0166, Zhengpeng Zhao, Jing Wang 0113, Yushu Du, Chenhao Ma 0006, Xiaomeng Han, Dean You, Jiapeng Guan, Zhe Jiang 0004
DAC8