VLDB 2026 Research / reviewers in the wild / expert
Kunxiong Zhu
dblp:322/5248
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2026
0009-0007-3050-0152ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Memory systems · 30% GPUs and heterogeneous computing · 23% Hardware accelerators and domain-specific architectures · 23% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
machine learning accelerator |
1.0 | 1 | 2026 | FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations · ASPLOS (2) 2026 |
Memory systems › memory hierarchy
memory hierarchy optimization |
1.0 | 1 | 2026 | FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations · ASPLOS (2) 2026 |
GPUs and heterogeneous computing › embedded GPU
mobile GPU |
1.0 | 1 | 2026 | FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations · ASPLOS (2) 2026 |
Embedded and real-time systems › on-device inference
mobile inference |
1.0 | 1 | 2026 | FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations · ASPLOS (2) 2026 |
Memory systems › memory access optimization
memory streaming |
0.3 | 1 | 2026 | FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations · ASPLOS (2) 2026 |
Methods — techniques the papers use, named apart from their topics
static scheduling · 1.02.5d texture memory · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy OptimizationsabstractThe increasing size and complexity of modern deep neural networks (DNNs) pose significant challenges for on-device inference on mobile GPUs, with limited memory and computational resources. Existing DNN acceleration frameworks primarily deploy a weight preloading strategy, where all model parameters are loaded into memory before execution on mobile GPUs. We posit that this approach is not adequate for modern DNN workloads that comprise very large model(s) and possibly execution of several distinct models in succession. In this work, we introduce FlashMem, a memory streaming framework designed to efficiently execute large-scale modern DNNs and multi-DNN workloads while minimizing memory consumption and reducing inference latency. Instead of fully preloading weights, FlashMem statically determines model loading schedules and dynamically streams them on demand, leveraging 2.5D texture memory to minimize data transformations and improve execution efficiency. Experimental results on 11 models demonstrate that FlashMem achieves 2.0× to 8.4× memory reduction and 1.7× to 75.0× speedup compared to existing frameworks, enabling efficient execution of large-scale models and multi-DNN support on resource-constrained mobile GPUs. Zhihao Shu, Md. Musfiqur Rahman Sanim, Hangyu Zheng, Kunxiong Zhu, Miao Yin, Gagan Agrawal, Wei Niu 0002 |
ASPLOS (2) | 4 |
| 2024 | Efficient Point Cloud Analytics on Edge DevicesabstractPoint clouds are crucial for 3D geometry representation, and vital in applications like autonomous driving and augmented reality. Despite advancements in deep learning-based analytics, their high computational cost limits deployment on edge devices with constrained resources. To this end, we analyze PointNet++, a leading point cloud analytics framework, identifying two major bottlenecks: 1) GPU is underutilized due to limited parallelism and excessive kernel launches in the sampling stage and voting stage, and 2) irregular memory accesses in the grouping stage. To address these, we propose parallel sampling and voting to enhance GPU utilization and fuse subroutines in grouping to improve memory efficiency. Experimental results demonstrate that our optimizations result in significant speedup (up to $5.0 \times, 3.2 \times$ on average) across various point cloud workloads on edge devices. Kunxiong Zhu, Zhenlin Wu 0001, Hongyuan Liu 0002 |
ICPADS | 1 |