Shaoxun Zeng

dblp:351/5891 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
0009-0004-6493-1768ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 GPU Checkpoint/Restore Made Fast and Lightweight
Shaoxun Zeng, Tingxu Ren, Jiwu Shu, Youyou Lu
FAST1
2025 Frugal: Efficient and Economic Embedding Model Training with Commodity GPUs
abstract
Embedding models show superiority in learning representations of massive ID-type features in sparse learning scenarios such as recommendation systems (e.g., user/item IDs) and graph learning (e.g., node/edge IDs). Commodity GPUs are highly favored for their cost-efficient computing power, which is ideally suited for the low computing demand of memory-intensive embedding models. However, directly running embedding model training on commodity GPUs yields poor performance because of their deficient communication resources (including low communication bandwidth and no PCIe P2P support).
Minhui Xie, Shaoxun Zeng, Youyou Lu
ASPLOS (1)2
2025 Medusa: Accelerating Serverless LLM Inference with Materialization
abstract
Serverless is a promising paradigm to provide scalable, cost-efficient, and easy-to-use model inference services. However, the cold start of model inference functions requires loading models to the devices, which incurs high latencies and undermines the benefits of serverless computing. In LLMs, things get even worse since two extra stages are introduced: a KV cache initialization stage that profiles and anticipates memory reservation for KV cache, and a capturing stage which dynamically constructs CUDA graphs for different batch sizes. Both stages are paramount to the inference performance, but become the main culprit of cold start latency.
Shaoxun Zeng, Minhui Xie, Youmin Chen, Youyou Lu
ASPLOS (1)1
2025 Weaver: Efficient Multi-LLM Serving with Attention Offloading
Qing Wang 0031, Shaoxun Zeng, Youyou Lu, Jiwu Shu
USENIX ATC3
2024 Volley: Accelerating Write-Read Orders in Disaggregated Storage
abstract
Modern data centers deploy disaggregated storage systems (e.g., NVMe over Fabrics, NVMe-oF) for fine-grained resource elasticity and high resource utilization. A client-side writeback cache is used to absorb writes and buffer frequently accessed data, thereby eliminating unnecessary remote storage accesses and improving performance. Yet, a cache miss on the full cache triggers an evict-and-fetch operation which evicts the old entries before new data blocks are fetched. Existing systems perform the evict-and-fetch operation by sequentially executing write and read I/O operations, which reduces the concurrency and makes it challenging to fully utilize the fast network and storage devices.
Shaoxun Zeng, Xiaojian Liao, Youyou Lu
EuroSys1
2023 SingularFS: A Billion-Scale Distributed File System Using a Single Metadata Server
Youyou Lu, Wenhao Lv, Xiaojian Liao, Shaoxun Zeng, Jiwu Shu
USENIX ATC5