Abeda Sultana

dblp:222/4141 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
3since 2021 · last 2026
0000-0003-4973-3507ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Resource Heterogeneity-Aware and Utilization-Enhanced Scheduling for Deep Learning Clusters
abstract
Scheduling deep learning (DL) models to train on powerful clusters with accelerators like GPUs and TPUs, presently falls short, either lacking fine-grained heterogeneity awareness or leaving resources substantially under-utilized. To fill this gap, we propose a novel task-level heterogeneity-aware scheduler for DL clusters, <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Hadar</i>, based on an optimization framework able to boost cluster resource utilization. <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Hadar</i> leverages the performance traits of DL jobs on a heterogeneous DL cluster to make scheduling decisions across both spatial and temporal dimensions. It characterizes the task-level performance heterogeneity for optimization and involves the primal-dual framework employing a dual subroutine, to solve the optimization problem and guide the scheduling design. Our trace-driven simulation with representative DL model training workloads demonstrates that <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Hadar</i> accelerates the total training time duration by 1.20× when compared with its state-of-the-art heterogeneity-aware counterpart, Gavel. Further, our <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Hadar</i> scheduler is enhanced to <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Hadar</i>E by forking each job into multiple copies to let a job train concurrently on heterogeneous GPUs resided on separate available cluster nodes (i.e., machines or servers) for resource utilization enhancement. <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Hadar</i>E is evaluated extensively on physical DL clusters for comparison with <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Hadar</i> and Gavel. With substantial enhancement in cluster resource utilization (by 1.45×), <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Hadar</i>E exhibits considerable speed-ups in DL model training, reducing the total training time duration by 50% (or 80%) on an Amazon’s AWS (or our lab) cluster, while producing trained DL models with consistently better inference quality than those trained by <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Hadar</i>.
Abeda Sultana, Nabin Pakka, Fei Xu 0009, Xu Yuan 0001, Li Chen 0019, Nian-Feng Tzeng
IEEE Trans. Computers1
2024 Hadar: Heterogeneity-Aware Optimization-Based Online Scheduling for Deep Learning Cluster
abstract
With the wide adoption of deep neural network (DNN) models for various applications, enterprises, and cloud providers have built deep learning clusters and increasingly deployed specialized accelerators, such as GPUs and TPUs, for DNN training jobs. To arbitrate cluster resources among multi-user jobs, existing schedulers fall short, either lacking fine-grained heterogeneity awareness or hardly generalizable to various scheduling policies. To fill this gap, we propose a novel design of a task-level heterogeneity-aware scheduler, Hadar, based on an online optimization framework that can express other scheduling algorithms. Hadar leverages the performance traits of DNN jobs on a heterogeneous cluster, characterizes the task-level performance heterogeneity in the optimization problem, and makes scheduling decisions across both spatial and temporal dimensions. The primal-dual framework is employed, with our design of a dual subroutine, to solve the optimization problem and guide the scheduling design. Extensive trace-driven simulations with representative DNN models have been conducted to demonstrate that Hadar improves the average job completion time (JCT) by 3× over an Apache YARN-based resource manager used in production. Moreover, Hadar outperforms Gavel [1], the state-of-the-art heterogeneity-aware scheduler, by 2.5× for the average JCT, shortens the queuing delay by 13%, and improves FTF (Finish-Time-Fairness) by 1.5%.
Abeda Sultana, Fei Xu 0009, Xu Yuan 0001, Li Chen 0019, Nian-Feng Tzeng
IPDPS1
2022 Eiffel: Efficient and Fair Scheduling in Adaptive Federated Learning
abstract
Emerging machine learning (ML) technologies, in combination with the increasing computational power of mobile devices, lead to the extensive adoption of ML-based applications. Different from conventional model training that needs to collect all the user data in centralized cloud servers, federated learning (FL) has recently drawn increasing research attention as it enables privacy-preserving model training. With FL, decentralized edge devices in participation, train their model copies locally over their siloed datasets, and periodically synchronize the model parameters. However, model training is computationally extensive which easily drains the battery of mobile devices. In addition, due to the uneven distribution of siloed datasets, the shared model may become biased. To address theefficiencyandfairnessconcerns in a resource-constrained federated learning setting, in this paper, we proposeEiffelto judiciously select mobile devices to participate in the global model aggregation, and adaptively adjust the frequency of local and global model updates.Eiffelaims to make scheduling and coordination for the federated learning towards both resource efficiency and model fairness. We have conducted theoretical analysis ofEiffelfrom the perspectives of fairness and convergence. Extensive experiments with a wide variety of real-world datasets and models, both on a networked prototype system and in a larger-scale simulated environment, have demonstrated that while maintaining similar accuracy performance,Eiffeloutperforms existing baselines with respect to reducing communication overhead by up to 6× for higher efficiency and improving the fairness metric by up to 57% compared to the state-of-the-art algorithms.
Abeda Sultana, Md. Mainul Haque, Li Chen 0019, Fei Xu 0009, Xu Yuan 0001
IEEE Trans. Parallel Distributed Syst.1
2020 E-LAS: Design and Analysis of Completion-Time Agnostic Scheduling for Distributed Deep Learning Cluster
abstract
With the prosperity of deep learning, enterprises, and large platform providers, such as Microsoft, Amazon, and Google, have built and provided GPU clusters to facilitate distributed deep learning training. As deep learning training workloads are heterogeneous, with a diverse range of characteristics and resource requirements, it becomes increasingly crucial to design an efficient and optimal scheduler for distributed deep learning jobs in the GPU cluster. This paper aims to propose a simple and yet effective scheduler, called E-LAS, with the objective of reducing the averaged training completion time of deep learning jobs. Without relying on the estimation or prior knowledge of the job running time, E-LAS leverages the real-time epoch progress rate, unique for distributed deep learning training jobs, as well as the attained services from temporal and spatial domains, to guide the scheduling decisions. The theoretical analysis for E-LAS is conducted to offer a deeper understanding on the components of scheduling criteria. Furthermore, we present a placement algorithm to achieve better resource utilization without involving much implementation overhead, complementary to the scheduling algorithm. Extensive simulations have been conducted, demonstrating that E-LAS improves the averaged job completion time (JCT) by 10 × over an Apache YARN-based resource manager used in production. Moreover, E-LAS outperforms Tiresias, the state-of-the-art scheduling algorithm customized for deep learning jobs, by almost 1.5 × for the average JCT as well as queuing time.
Abeda Sultana, Li Chen 0019, Fei Xu 0009, Xu Yuan 0001
ICPP1