Kunxiong Zhu

dblp:322/5248 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
0009-0007-3050-0152ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 30% GPUs and heterogeneous computing · 23% Hardware accelerators and domain-specific architectures · 23%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
1.012026
FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations · ASPLOS (2) 2026
Memory systems › memory hierarchy
memory hierarchy optimization
1.012026
FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations · ASPLOS (2) 2026
GPUs and heterogeneous computing › embedded GPU
mobile GPU
1.012026
FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations · ASPLOS (2) 2026
Embedded and real-time systems › on-device inference
mobile inference
1.012026
FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations · ASPLOS (2) 2026
Memory systems › memory access optimization
memory streaming
0.312026
FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations · ASPLOS (2) 2026

Methods — techniques the papers use, named apart from their topics

static scheduling · 1.02.5d texture memory · 1.0
YearPublicationVenuePosition
2026 FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations
abstract
The increasing size and complexity of modern deep neural networks (DNNs) pose significant challenges for on-device inference on mobile GPUs, with limited memory and computational resources. Existing DNN acceleration frameworks primarily deploy a weight preloading strategy, where all model parameters are loaded into memory before execution on mobile GPUs. We posit that this approach is not adequate for modern DNN workloads that comprise very large model(s) and possibly execution of several distinct models in succession. In this work, we introduce FlashMem, a memory streaming framework designed to efficiently execute large-scale modern DNNs and multi-DNN workloads while minimizing memory consumption and reducing inference latency. Instead of fully preloading weights, FlashMem statically determines model loading schedules and dynamically streams them on demand, leveraging 2.5D texture memory to minimize data transformations and improve execution efficiency. Experimental results on 11 models demonstrate that FlashMem achieves 2.0× to 8.4× memory reduction and 1.7× to 75.0× speedup compared to existing frameworks, enabling efficient execution of large-scale models and multi-DNN support on resource-constrained mobile GPUs.
Zhihao Shu, Md. Musfiqur Rahman Sanim, Hangyu Zheng, Kunxiong Zhu, Miao Yin, Gagan Agrawal, Wei Niu 0002
ASPLOS (2)4
2024 Efficient Point Cloud Analytics on Edge Devices
abstract
Point clouds are crucial for 3D geometry representation, and vital in applications like autonomous driving and augmented reality. Despite advancements in deep learning-based analytics, their high computational cost limits deployment on edge devices with constrained resources. To this end, we analyze PointNet++, a leading point cloud analytics framework, identifying two major bottlenecks: 1) GPU is underutilized due to limited parallelism and excessive kernel launches in the sampling stage and voting stage, and 2) irregular memory accesses in the grouping stage. To address these, we propose parallel sampling and voting to enhance GPU utilization and fuse subroutines in grouping to improve memory efficiency. Experimental results demonstrate that our optimizations result in significant speedup (up to $5.0 \times, 3.2 \times$ on average) across various point cloud workloads on edge devices.
Kunxiong Zhu, Zhenlin Wu 0001, Hongyuan Liu 0002
ICPADS1