Chenxuan Yao

dblp:351/0897 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0002-1143-052XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2026 SpotCC: Facilitating Coded Computation for Prediction Serving Systems on Spot Instances
abstract
The growing adoption of prediction serving systems (PSSes) has made cost-saving deployment on preemptible spot instances crucial, yet frequent preemptions severely harm availability. While coded computation (CC) can keep availability cost-effectively by encoding original jobs into parity ones, its direct application to spot instances incurs prohibitive decoding overhead and tail latency under frequent preemptions. We identify two findings for optimization: (i) decoding asymmetry (only original job failures require decoding); (ii) preemption unevenness (variation in preemption rates across cloud regions). Leveraging these findings, we propose SpotCC, a new CC framework that strategically dispatches parity jobs to high-preemption (volatile) regions and original jobs to low-preemption (stable) regions. SpotCC designs locality-based and fine-grained volatility identification to reduce decoding operations and mitigate job congestion, respectively, and adaptively tunes configurations for decoding minimization. Experiments show that SpotCC improves P99 latency by 83.9% over state-of-the-arts, while maintaining ultra-low monetary costs.
Yuchong Hu, Ziling Duan, Chenxuan Yao, Xiaolu Li 0002, Leihua Qin, Dan Feng 0001
HPCA5
2026 COAR: Computation-Aware Erasure-Coded Repair in Storage Clusters for Distributed Data Analytics
Yuchong Hu, Chenxuan Yao, Patrick P. C. Lee, Dan Feng 0001
ICDCS4
2025 Saving Memory via Residual Reduction for DNN Training with Compressed Communication
Xinjue Zheng, Zhangqiang Ming, Yuchong Hu, Chenxuan Yao, Wenxiang Zhou, Dan Feng 0001
Euro-Par (2)4
2025 LowDiff: Efficient Frequent Checkpointing via Low-Cost Differential for High-Performance Distributed Training Systems
abstract
Distributed training of large deep-learning models often leads to failures, so checkpointing is commonly employed for recovery. State-of-the-art studies focus on frequent checkpointing for fast recovery from failures. However, it generates numerous checkpoints, incurring substantial costs and thus degrading training performance. Recently, differential checkpointing has been proposed to reduce costs, but it is limited to recommendation systems, so its application to general distributed training systems remains unexplored.
Chenxuan Yao, Yuchong Hu, Xinjue Zheng, Wenxiang Zhou
SC1
2024 ADTopk: All-Dimension Top-k Compression for High-Performance Data-Parallel DNN Training
abstract
Data-parallel deep neural networks (DNN) training systems deployed across nodes have been widely used in various domains, while the system performance is often bottlenecked by the communication overhead among workers for synchronizing gradients. Top-k sparsification compression is the de facto approach to alleviate the communication bottleneck, which truncates the gradient to its largest k elements before sending it to other nodes.
Zhangqiang Ming, Yuchong Hu, Wenxiang Zhou, Xinjue Zheng, Chenxuan Yao, Dan Feng 0001
HPDC5