EDBT 2026 Demo / reviewers in the wild / expert
Shaoxun Zeng
dblp:351/5891
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2026
0009-0004-6493-1768ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GPU Checkpoint/Restore Made Fast and Lightweight
Shaoxun Zeng, Tingxu Ren, Jiwu Shu, Youyou Lu |
FAST | 1 |
| 2025 | Frugal: Efficient and Economic Embedding Model Training with Commodity GPUsabstractEmbedding models show superiority in learning representations of massive ID-type features in sparse learning scenarios such as recommendation systems (e.g., user/item IDs) and graph learning (e.g., node/edge IDs). Commodity GPUs are highly favored for their cost-efficient computing power, which is ideally suited for the low computing demand of memory-intensive embedding models. However, directly running embedding model training on commodity GPUs yields poor performance because of their deficient communication resources (including low communication bandwidth and no PCIe P2P support). Minhui Xie, Shaoxun Zeng, Youyou Lu |
ASPLOS (1) | 2 |
| 2025 | Medusa: Accelerating Serverless LLM Inference with MaterializationabstractServerless is a promising paradigm to provide scalable, cost-efficient, and easy-to-use model inference services. However, the cold start of model inference functions requires loading models to the devices, which incurs high latencies and undermines the benefits of serverless computing. In LLMs, things get even worse since two extra stages are introduced: a KV cache initialization stage that profiles and anticipates memory reservation for KV cache, and a capturing stage which dynamically constructs CUDA graphs for different batch sizes. Both stages are paramount to the inference performance, but become the main culprit of cold start latency. Shaoxun Zeng, Minhui Xie, Youmin Chen, Youyou Lu |
ASPLOS (1) | 1 |
| 2025 | Weaver: Efficient Multi-LLM Serving with Attention Offloading
Qing Wang 0031, Shaoxun Zeng, Youyou Lu, Jiwu Shu |
USENIX ATC | 3 |
| 2024 | Volley: Accelerating Write-Read Orders in Disaggregated StorageabstractModern data centers deploy disaggregated storage systems (e.g., NVMe over Fabrics, NVMe-oF) for fine-grained resource elasticity and high resource utilization. A client-side writeback cache is used to absorb writes and buffer frequently accessed data, thereby eliminating unnecessary remote storage accesses and improving performance. Yet, a cache miss on the full cache triggers an evict-and-fetch operation which evicts the old entries before new data blocks are fetched. Existing systems perform the evict-and-fetch operation by sequentially executing write and read I/O operations, which reduces the concurrency and makes it challenging to fully utilize the fast network and storage devices. Shaoxun Zeng, Xiaojian Liao, Youyou Lu |
EuroSys | 1 |
| 2023 | SingularFS: A Billion-Scale Distributed File System Using a Single Metadata Server
Youyou Lu, Wenhao Lv, Xiaojian Liao, Shaoxun Zeng, Jiwu Shu |
USENIX ATC | 5 |