Zuocheng Shi

dblp:348/6420 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0004-7835-6522ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
GPUs and heterogeneous computing · 43% Storage systems · 25% Electronic design automation · 12%
Artificial intelligence
4 papers
Graph learning · 79% Efficient and distributed learning · 21%

Topics — the 11 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning
graph neural network training
1.522025
Hyperion: Co-Optimizing SSD Access and GPU Computation for Cost-Efficient GNN Training · ICDE 2025
Legion: Automatically Pushing the Envelope of Multi-GPU System for Billion-Scale GNN Training · USENIX ATC 2023
Machine learning › Efficient and distributed learning
distributed training
0.912025
Understanding Stragglers in Large Model Training Using What-if Analysis · OSDI 2025
Machine learning › Graph learning › graph neural network
graph neural network inference
0.912025
Helios: Efficient Distributed Dynamic Graph Sampling for Online GNN Inference · PPoPP 2025
Machine learning › Graph learning › graph neural network training
out-of-core GNN training
0.912025
Hyperion: Co-Optimizing SSD Access and GPU Computation for Cost-Efficient GNN Training · ICDE 2025
Storage systems
data placement
0.912025
Moment: Co-optimizing Physical Communication Topology and Data Placement for Multi-GPU Out-of-core GNN Training · SC 2025
GPUs and heterogeneous computing › multi-GPU computing
multi-GPU training
0.912025
Moment: Co-optimizing Physical Communication Topology and Data Placement for Multi-GPU Out-of-core GNN Training · SC 2025
Electronic design automation › design optimization
topology optimization
0.912025
Moment: Co-optimizing Physical Communication Topology and Data Placement for Multi-GPU Out-of-core GNN Training · SC 2025
GPUs and heterogeneous computing
graph neural network training
0.712023
Legion: Automatically Pushing the Envelope of Multi-GPU System for Billion-Scale GNN Training · USENIX ATC 2023
GPUs and heterogeneous computing
multi-GPU computing
0.712023
Legion: Automatically Pushing the Envelope of Multi-GPU System for Billion-Scale GNN Training · USENIX ATC 2023
Distributed systems
distributed graph processing
0.312025
Helios: Efficient Distributed Dynamic Graph Sampling for Online GNN Inference · PPoPP 2025
High-performance computing
large-scale training
0.312025
Understanding Stragglers in Large Model Training Using What-if Analysis · OSDI 2025

Methods — techniques the papers use, named apart from their topics

what-if analysis · 1.7unified cache · 1.7query-aware sample cache · 1.7pre-sampling · 1.7max-flow · 0.9knapsack algorithms · 0.9asynchronous i/o · 0.9asynchronous IO · 0.9
YearPublicationVenuePosition
2025 Hyperion: Co-Optimizing SSD Access and GPU Computation for Cost-Efficient GNN Training
abstract
SSDs are traditionally regarded as a cheap but slow way to scale up GNN training. Several GNN systems explore cheap single-machine single-GPU out-of-core training but fall short in terms of TPC (throughput per monetary cost). The underlying reason is that the existing systems 1) overly focus on minimizing the number of SSD accesses, which results in substantial unnecessary overhead on the CPU side, or 2) exhaust all GPU parallelism to saturate SSD but fail to overlap SSD accesses with GNN computation. In this work, we present Hyperion, a cost-efficient system for terabyte-scale GNN training. We argue that co-optimizing GPU-initiated asynchronous SSD access and GNN computation pipeline enables us to only add cheap NVMe SSDs, rather than expensive GPU servers, to achieve in- memory-like throughput and thus maximal TPC of GNN training. However, this is non-trivial due to imbalanced workloads and interference among IO submission, IO completion, and cache lookup. To tackle the challenges, Hyperion proposes three key designs. First, Hyperion proposes the first GPU-initiated pipeline- friendly asynchronous disk IO stack, which only requires about 1% GPU cores to saturate SSD throughput and wastes no GPU cores between IO submission and completion to fully overlap disk IO and computation. Second, we propose a new GPU-managed, disaggregated, and unified cache that disaggregates cache lookup from disk IO and fully utilizes CPU/GPU memory hierarchy by a unified static cache policy. Third, we propose a GNN-aware general TPC-analytical model that precisely predicts TPC under diverse hardware settings and GNN models and provide a hint to guide users to select hardware, e.g., number of SSDs, under a limited budget to maximize TPC. Experiments demonstrate that Hyperion can improve the TPC by over 3.1x on terabyte-scale graphs compared to SOTA out-of-core baselines and improve 60 x TPC compared to distributed in-memory baselines.
Jie Sun 0017, Mo Sun 0001, Zuocheng Shi, Zihan Yang 0004, Jie Zhang 0081, Zeke Wang, Fei Wu 0001
ICDE4
2025 Understanding Stragglers in Large Model Training Using What-if Analysis
Jinkun Lin, Ziheng Jiang, Zuquan Song, Sida Zhao, Menghan Yu, Zhanghan Wang, Zuocheng Shi, Zherui Liu, Shuguang Wang, Haibin Lin, Xin Liu 0086, Aurojit Panda, Jinyang Li 0001
OSDI8
2025 Helios: Efficient Distributed Dynamic Graph Sampling for Online GNN Inference
abstract
Online GNN inference has been widely explored by applications such as online recommendation and financial fraud detection systems, where even minor delays can result in significant financial impact. Real-time dynamic graph sampling enables online GNN inference to reflect the latest graph updates in real-world graphs. However, online GNN inference typically demands millisecond-level latency Service Level Objectives (SLOs) as its performance guarantees, which poses great challenges for existing dynamic graph sampling approaches based on graph databases. The issues mainly arise from two aspects: long tail latency due to imbalanced data-dependent sampling and large communication overhead incurred by distributed sampling. To address these issues, we propose Helios, an efficient distributed dynamic graph sampling service to meet the stringent latency SLOs. The key ideas of Helios are 1) pre-sampling the dynamic graph in an event-driven approach, and 2) maintaining a query-aware sample cache to build the complete K-hop sampling results locally for inference requests. Experiments on multiple datasets show that Helios achieves up to 67× higher serving throughput and up to 32× lower P99 query latency compared to baselines.
Jie Sun 0017, Zuocheng Shi, Li Su 0005, Wenting Shen, Zeke Wang, Yong Li 0045, Wenyuan Yu, Wei Lin 0016, Fei Wu 0001, Bingsheng He, Jingren Zhou 0001
PPoPP2
2025 Moment: Co-optimizing Physical Communication Topology and Data Placement for Multi-GPU Out-of-core GNN Training
abstract
Graph Neural Networks (GNNs) are widely employed in applications like recommendation systems, social network analysis, and fraud detection, but training large-scale GNNs is challenging due to its memory limitations. Existing systems face a trade-off between throughput and monetary cost: Distributed systems require expensive memory scaling, while single-machine out-of-core systems are limited by GPU/PCIe throughput. To this end, we propose Moment, a physical communication topology and data placement co-optimizer to enable high-throughput and low-cost GNN training in a single multi-GPU machine. Moment addresses communication contention and GPU load imbalance issues by modeling the physical topology as capacity-constrained directed graphs and formulating communication scheduling as a max-flow problem. It also introduces a data-distribution-aware knapsack algorithm for optimized data placement. Experimental results show that Moment outperforms out-of-core systems by up to 6.51 × and distributed systems by up to 3.02 ×, with only 50% monetary cost.
Zuocheng Shi, Jie Sun 0017, Ziyu Song, Mo Sun 0001, Fei Wu 0001, Zeke Wang
SC1
2023 Legion: Automatically Pushing the Envelope of Multi-GPU System for Billion-Scale GNN Training
Jie Sun 0017, Li Su 0005, Zuocheng Shi, Wenting Shen, Zeke Wang, Lei Wang 0004, Jie Zhang 0081, Yong Li 0020, Wenyuan Yu, Jingren Zhou 0001, Fei Wu 0001
USENIX ATC3