Runhua Zhang 0002

dblp:249/7886-2 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0003-3487-5289ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LSAF: A load-balancing SpGEMM acceleration framework with dynamic package and static partition for multi-core systolic arrays
Yongxiang Cao, Guocheng Zhao, Dongcheng Shi, Runhua Zhang 0002
Parallel Comput.5
2025 FlexPie: Accelerate Distributed Inference on Edge Devices with Flexible Combinatorial Optimization
Runhua Zhang 0002, Jinkun Geng, Chenhui Zhu
DASFAA (1)1
2025 SparDR: Accelerating Unstructured Sparse DNN Inference via Dataflow Optimization
abstract
Unstructured sparsity is becoming a key dimension in exploring the inference efficiency of neural networks. However, its data layout presents irregularity, making it difficult to match the parallel computing mode of hardware, resulting in low computational and memory access efficiency. We have studied this issue and found that the main reason is that existing sparse acceleration libraries and compilers perform sparse matrix multiplication optimization exploration through the splitting and reconstruction of sparse patterns, thus ignoring the acceleration of sparse convolution operations centered on data streams, which may miss some optimization opportunities for sparse operations. In this article, we propose SparDR, a general sparse convolution operation acceleration method centered around data streams. Through novel feature map data stream reconstruction and convolutional kernel data representation, redundant zero value calculations are effectively avoided, addressing efficiency is improved, and memory overhead is reduced. SparDR is based on TVM and allows for automatic scheduling across different hardware configurations. Compared with the current mainstream five methods on four types of hardware, the inference delay acceleration reaches 1.1-12 x and the memory usage decreases by 20%.
Runhua Zhang 0002, Yongxiang Cao, Yaochen Han
DATE3
2025 SAM-Lightning: Segment Anything Model for Efficient Inference and Reduced Memory Footprint
Yanfei Song, Bangzheng Pu, Yongxiang Cao, Runhua Zhang 0002, Yiqing Shen 0003
PRICAI (5)7
2025 JOVS: Joint Optimization of Vectorization and Scheduling for DNN on AI DSPs
abstract
Recent embedded devices have integrated digital signal processors (DSPs) to balance performance and power when executing complex Deep Neural Network (DNN) workloads. With modern AI DSPs providing specialized tensor computation vector instructions and limited on-chip memory, fully releasing the potential of these DSPs remains a significant challenge. The performance of AI DSPs relies heavily on vendor-provided libraries and compilers. In practice, vendor-provided libraries are inflexible and prevent further optimization. State-of-the-art compilers usually focus on a single optimization (vectorization or scheduling), which is insufficient to address this challenge.
Yaochen Han, Runhua Zhang 0002, Rui She 0001
SPAA3
2024 A high-performance dataflow-centric optimization framework for deep learning inference on the edge
Runhua Zhang 0002, Jinkun Geng
J. Syst. Archit.1
2023 Xenos : Dataflow-Centric Optimization to Accelerate Model Inference on Edge Devices
Runhua Zhang 0002, Jinkun Geng, Chenhui Zhu
DASFAA (1)1
2023 LOCP: Latency-optimized channel pruning for CNN inference acceleration on GPUs
Runhua Zhang 0002, Yongxiang Cao, Chenhui Zhu
J. Supercomput.4
2021 SRQ: Self-Reference quantization scheme for lightweight neural network
Shuangxi Huang, Runhua Zhang 0002
DCC5
2021 Robustness-aware 2-bit quantization with real-time performance for neural network
Runhua Zhang 0002, Shuangxi Huang, Donghuan Xu
Neurocomputing3
2019 HiPower: A High-Performance RDMA Acceleration Solution for Distributed Transaction Processing
Runhua Zhang 0002, Jinkun Geng, Shuai Wang 0028, Kaihui Gao, Guowei Shen
NPC1