EDBT 2026 Demo / reviewers in the wild / expert
Wei Gao 0064
dblp:28/2073-64
· DBLP profile ↗
12ranked-venue papers
8as first author
12since 2021 · last 2026
0000-0002-7048-1722ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ResiHP: Taming LLM Training Failures with Dynamic Hybrid ParallelismabstractHybrid parallelism underpins large-scale LLM training across tens of thousands of GPUs. At such scale, hardware failures on individual devices lead to performance skew across devices, diminishing overall training efficiency. Existing resilient systems overlook sequence length variability in datasets and device performance skew under hybrid parallelism. As a result, (1) iteration time fluctuations induced by sequence length variability can trigger spurious fail-slow detections, and (2) failures are mitigated through individual adaptations in hybrid parallelism, leading to unnecessary detection overhead and inefficient resilient training. Tenghui Ma, Jihu Guo, Wei Gao 0064, Sitian Lu, Zhisheng Ye 0002, Hanjing Wang, Dahua Lin |
HPDC | 3 |
| 2026 | SPPO: Making Million-Token LLM Training Practical on Modest GPU ClustersabstractIn recent years, Large Language Models (LLMs) have exhibited remarkable capabilities, driving advancements in real-world applications. However, training LLMs on increasingly long input sequences imposes significant challenges due to high GPU memory and computational demands. Existing solutions face two key limitations: (1) memory reduction techniques, such as activation recomputation and CPU offloading, compromise training efficiency; (2) distributed parallelism strategies require excessive GPU resources, limiting the scalability of input sequence length. Qiaoling chen, Shenggui Li, Wei Gao 0064, Peng Sun 0006, Yonggang Wen 0001, Tianwei Zhang 0004 |
ICS | 3 |
| 2025 | IceFrog: A Layer-Elastic Scheduling System for Deep Learning Training in GPU ClustersabstractThe high resource demand of deep learning training (DLT) workloads necessitates the design of efficient schedulers. While most existing schedulers expedite DLT workloads by considering GPU sharing and elastic training, they neglectlayer elasticity, which dynamically freezes certain layers of a network. This technique has been shown to significantly speed up individual workloads. In this paper, we explore how to incorporatelayer elasticityinto DLT scheduler designs to achieve higher cluster-wide efficiency. A key factor that hinders the application of layer elasticity in GPU clusters is the potential loss in model accuracy, making users reluctant to enable layer elasticity for their workloads. It is necessary to have an efficient layer-elastic system, which can well balance training accuracy and speed for layer elasticity. We introduceIceFrog, the first scheduling system that utilizes layer elasticity to improve the efficiency of DLT workloads in GPU clusters. It achieves this goal with superior algorithmic designs and intelligent resource management. In particular, (1) we model the frozen penalty and layer-aware throughput to measure the effective progress metric of layer-elastic workloads. (2) We design a novel scheduler to further improve the efficiency of layer elasticity. We implement and deployIceFrogin a physical cluster of 48 GPUs. Extensive evaluations and large-scale simulations show thatIceFrogreduces average job completion times by 36-48% relative to state-of-the-art DL schedulers. Wei Gao 0064, Zhuoyuan Ouyang, Peng Sun 0006, Tianwei Zhang 0004, Yonggang Wen 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2024 | AutoSched: An Adaptive Self-configured Framework for Scheduling Deep Learning Training WorkloadsabstractModern Deep Learning Training (DLT) schedulers in GPU datacenters are designed to be very sophisticated with many configurations. These configurations need to be adjusted delicately as they can significantly affect the scheduling performance. Existing schedulers require the datacenter operator to tune the configurations only once before they are deployed, based on the historical workload traces. Unfortunately, workloads in a datacenter would experience dynamic changes and deviate a lot from the historical ones over time, making the pre-determined configurations less effective. Wei Gao 0064, Shangwei Guo, Peng Sun 0006, Yonggang Wen 0001, Tianwei Zhang 0004 |
ICS | 1 |
| 2024 | Ymir: A Scheduler for Foundation Model Fine-tuning Workloads in DatacentersabstractThe breakthrough of foundation models makes foundation model fine-tuning (FMF) workloads prevalent in modern GPU datacenters. However, existing schedulers tailored for model training do not consider the unique characteristics of FMs, making them inefficient in handling FMF workloads. To bridge the gap, we propose Ymir, a scheduler to improve the efficiency of FMF workloads in GPU datacenters. Ymir leverages the shared FM backbone architecture to expedite FMF workloads from two aspects: (1) Ymir investigates the task transferability among different FMF workloads and automatically merges FMF workloads with the same FM into one to improve the cluster-wide efficiency via transfer learning. (2) Ymir reuses the fine-tuning runtime of FMF workloads to reduce the significant context switch overhead. We conduct 32-GPU physical experiments and 240-GPU trace-driven simulations to validate the effectiveness of Ymir. Ymir can reduce the average job completion time by up to 4.3 × compared with existing state-of-the-art schedulers. It also promotes scheduling fairness by fully exploiting the task transferability. More supplementary materials can be found on our project website https://sites.google.com/view/ymir-project. Wei Gao 0064, Weiming Zhuang, Minghao Li 0005, Peng Sun 0006, Yonggang Wen 0001, Tianwei Zhang 0004 |
ICS | 1 |
| 2024 | UniSched: A Unified Scheduler for Deep Learning Training Jobs With Different User DemandsabstractThe growth of deep learning training (DLT) jobs in modern GPU clusters calls for efficient deep learning (DL) scheduler designs. Due to the extensive applications of DL technology, developers may have different demands for their DLT jobs. It is important for a GPU cluster to support all these demands and efficiently execute those DLT jobs. Unfortunately, existing DL schedulers mainly focus on part of those demands, and cannot provide comprehensive scheduling services.In this work, we present UniSched, a unified scheduler to optimize different types of scheduling objectives (e.g., guaranteeing the deadlines of SLO jobs, minimizing the latency of best-effort jobs). Meanwhile, UniSchedsupports different job stopping criteria (e.g., iteration-based, performance-based). UniSched includes two key components: Estimator for estimating the job duration, and Selector for selecting jobs and allocating resources. We perform large-scale simulations over the job traces from the production clusters. Compared to state-of-the-art schedulers, UniSchedcan significantly decrease the deadline miss rate of SLO jobs by up to 6.84×, and the latency of best-effort jobs by up to 4.02×, To demonstrate the practicality of UniSched, we implement and deploy a prototype on Kubernetes in a physical cluster consisting of 64 GPUs. Wei Gao 0064, Zhisheng Ye 0002, Peng Sun 0006, Tianwei Zhang 0004, Yonggang Wen 0001 |
IEEE Trans. Computers | 1 |
| 2023 | Automatic Transformation Search Against Deep Leakage From GradientsabstractCollaborative learning has gained great popularity due to its benefit of data privacy protection: participants can jointly train a Deep Learning model without sharing their training sets. However, recent works discovered that an adversary can fully recover the sensitive training samples from the shared gradients. Such reconstruction attacks pose severe threats to collaborative learning. Hence, effective mitigation solutions are urgently desired. In this paper, we systematically analyze existing reconstruction attacks and propose to leverage data augmentation to defeat these attacks: by preprocessing sensitive images with carefully-selected transformation policies, it becomes infeasible for the adversary to extract training samples from the corresponding gradients. We first design two new metrics to quantify the impacts of transformations on data privacy and model usability. With the two metrics, we design a novel search method to automatically discover qualified policies from a given data augmentation library. Our defense method can be further combined with existing collaborative training systems without modifying the training protocols. We conduct comprehensive experiments on various system settings. Evaluation results demonstrate that the policies discovered by our method can defeat state-of-the-art reconstruction attacks in collaborative learning, with high efficiency and negligible impact on the model performance. Wei Gao 0064, Shangwei Guo, Tianwei Zhang 0004, Tao Xiang 0001, Han Qiu 0001, Yonggang Wen 0001, Yang Liu 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Titan: a scheduler for foundation model fine-tuning workloadsabstractThe recent breakthrough of foundation model (FM) research raises a new trend to acquire efficient DL models by fine-tuning FMs with low-resource datasets. Current GPU clusters are mainly established to develop DL models by training from scratch. How to tailor a GPU cluster scheduler for FM fine-tuning workloads is still not explored. Wei Gao 0064, Peng Sun 0006, Yonggang Wen 0001, Tianwei Zhang 0004 |
SoCC | 1 |
| 2022 | Astraea: A Fair Deep Learning Scheduler for Multi-Tenant GPU ClustersabstractModern GPU clusters are designed to support distributed Deep Learning jobs from multiple tenants concurrently. Each tenant may have varied and dynamic resource demands. Unfortunately, existing GPU schedulers fail to thoroughly consider the fairness among the tenants and jobs, which can result in unbalanced resource allocation and unfair user experience. In this article, we present an efficient solution to provide strong fairness while maintaining high scheduling effectiveness in multi-tenant GPU clusters. First, we introduce a novel Long-Term GPU-time Fairness metric, which can comprehensively evaluate the fairness at both the tenant and job levels, based on both the temporal and spatial impacts of resource allocation. Second, we design a new and practical GPU scheduler,Astraea, to enforce the desired fairness among tenants and jobs. Large-scale evaluations show thatAstraeacan improve tenant fairness by up to 9.42× compared to state-of-the-art schedulers, without sacrificing the average job completion time. Zhisheng Ye 0002, Peng Sun 0006, Wei Gao 0064, Tianwei Zhang 0004, Xiaolin Wang 0001, Shengen Yan, Yingwei Luo |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | Chronus: A Novel Deadline-aware Scheduler for Deep Learning Training JobsabstractModern GPU clusters support Deep Learning training (DLT) jobs in a distributed manner. Job scheduling is the key to improve the training performance, resource utilization and fairness across users. Different training jobs may require various objectives and demands in terms of completion time. How to efficiently satisfy all these requirements is not extensively studied. Wei Gao 0064, Zhisheng Ye 0002, Peng Sun 0006, Yonggang Wen 0001, Tianwei Zhang 0004 |
SoCC | 1 |
| 2021 | Privacy-Preserving Collaborative Learning With Automatic Transformation SearchabstractCollaborative learning has gained great popularity due to its benefit of data privacy protection: participants can jointly train a Deep Learning model without sharing their training sets. However, recent works discovered that an adversary can fully recover the sensitive training samples from the shared gradients. Such reconstruction attacks pose severe threats to collaborative learning. Hence, effective mitigation solutions are urgently desired.In this paper, we propose to leverage data augmentation to defeat reconstruction attacks: by preprocessing sensitive images with carefully-selected transformation policies, it becomes infeasible for the adversary to extract any useful information from the corresponding gradients. We design a novel search method to automatically discover qualified policies. We adopt two new metrics to quantify the impacts of transformations on data privacy and model usability, which can significantly accelerate the search speed. Comprehensive evaluations demonstrate that the policies discovered by our method can defeat existing reconstruction attacks in collaborative learning, with high efficiency and negligible impact on the model performance. Wei Gao 0064, Shangwei Guo, Tianwei Zhang 0004, Han Qiu 0001, Yonggang Wen 0001, Yang Liu 0003 |
CVPR | 1 |
| 2021 | Improving Transferability of Adversarial Patches on Face Recognition With Generative ModelsabstractFace recognition is greatly improved by deep convolutional neural networks (CNNs). Recently, these face recognition models have been used for identity authentication in security sensitive applications. However, deep CNNs are vulnerable to adversarial patches, which are physically realizable and stealthy, raising new security concerns on the real-world applications of these models. In this paper, we evaluate the robustness of face recognition models using adversarial patches based on transferability, where the attacker has limited accessibility to the target models. First, we extend the existing transfer-based attack techniques to generate transferable adversarial patches. However, we observe that the transferability is sensitive to initialization and degrades when the perturbation magnitude is large, indicating the overfitting to the substitute models. Second, we propose to regularize the adversarial patches on the low dimensional data manifold. The manifold is represented by generative models pre-trained on legitimate human face images. Using face-like features as adversarial perturbations through optimization on the manifold, we show that the gaps between the responses of substitute models and the target models dramatically decrease, exhibiting a better transferability. Extensive digital world experiments are conducted to demonstrate the superiority of the proposed method in the black-box setting. We apply the proposed method in the physical world as well. Zihao Xiao 0002, Xianfeng Gao, Chilin Fu, Yinpeng Dong, Wei Gao 0064, Jun Zhou 0011, Jun Zhu 0001 |
CVPR | 5 |