Resilience-Aware Elastic Scaling for Cloud-Native Online DL Training on Multi-Tenant GPU Clusters

vldb26-988 · Regular Research · Qianhao Wu, Jiazhi Jiang, Guihui Ling, Yue Pang
Abstract

Online deep learning (DL) training has become pivotal in powering real-time applications. Yet tidal workload fluctuations leave GPU clusters significantly underutilized during off-peak periods. This not only wastes GPU capacity but also exacerbates scarcity for other GPU-intensive jobs on cloud-native GPU clusters. Cluster-wide resource leasing across different tenants enabled by elastic scaling offers a promising opportunity to enhance GPU utilization for cloud-native online DL training on multi-tenant GPU clusters. Existing solutions do not address the unique challenges of maintaining system stability during elastic scaling, including prolonged disruptions due to job reconstruction, failures arising from triggering dependency-unaware operations, and unreliable reclamation of loaned GPU resources. In this paper, we introduce WeFlex, a resilience-aware elastic scaling solution engineered for cloud-native online deep learning jobs in multi-tenant GPU clusters. WeFlex allows GPUs from online training jobs to be leased to other GPU-intensive jobs during low-demand periods while ensuring rapid reclamation as demand surges. It significantly reduces the duration of training disruptions by constructing an interruption mitigation pipeline, prevents dependency-unaware operation failures via topology-aware pod orchestration, and ensures reclamation of GPU resources through right-of-return GPU leasing. Evaluations on production GPU clusters at a 10,000-plus scale demonstrate that WeFlex enhances GPU utilization of online training by about 25\% while reliably maintaining continuous training performance.

Assigned reviewers

No reviewers assigned yet.

Candidates from the panel ranked by taxonomy affinity

#ReviewerMatchLoadWhy