Menglu Yu

dblp:293/8005 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
3since 2021 · last 2022
0000-0002-3653-3453ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 3 · 3 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 53% Distributed systems · 47%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
1.122022
GADGET: Online Resource Optimization for Scheduling Ring-All-Reduce Learning Jobs · INFOCOM 2022
A Sum-of-Ratios Multi-Dimensional-Knapsack Decomposition for DNN Resource Scheduling · INFOCOM 2021
Distributed systems › distributed machine learning
distributed deep learning
1.122022
GADGET: Online Resource Optimization for Scheduling Ring-All-Reduce Learning Jobs · INFOCOM 2022
A Sum-of-Ratios Multi-Dimensional-Knapsack Decomposition for DNN Resource Scheduling · INFOCOM 2021
Cloud and datacenter computing
job scheduling
1.122022
GADGET: Online Resource Optimization for Scheduling Ring-All-Reduce Learning Jobs · INFOCOM 2022
A Sum-of-Ratios Multi-Dimensional-Knapsack Decomposition for DNN Resource Scheduling · INFOCOM 2021
Distributed systems › distributed machine learning
parameter server
0.512021
A Sum-of-Ratios Multi-Dimensional-Knapsack Decomposition for DNN Resource Scheduling · INFOCOM 2021
Distributed systems › distributed machine learning
distributed training
0.322022
GADGET: Online Resource Optimization for Scheduling Ring-All-Reduce Learning Jobs · INFOCOM 2022
A Sum-of-Ratios Multi-Dimensional-Knapsack Decomposition for DNN Resource Scheduling · INFOCOM 2021

Methods — techniques the papers use, named apart from their topics

analytical modeling · 1.1trace-driven experiments · 0.6greedy algorithm · 0.6sum-of-ratios multi-dimensional-knapsack decomposition · 0.5
YearPublicationVenuePosition
2022 GADGET: Online Resource Optimization for Scheduling Ring-All-Reduce Learning Jobs
abstract
Fueled by advances in distributed deep learning (DDL), recent years have witnessed a rapidly growing demand for resource-intensive distributed/parallel computing to process DDL computing jobs. To resolve network communication bottleneck and load balancing issues in distributed computing, the so-called "ring-all-reduce" decentralized architecture has been increasingly adopted to remove the need for dedicated parameter servers. To date, however, there remains a lack of theoretical understanding on how to design resource optimization algorithms for efficiently scheduling ring-all-reduce DDL jobs in computing clusters. This motivates us to fill this gap by proposing a series of new resource scheduling designs for ring-all-reduce DDL jobs. Our contributions in this paper are threefold: i) We propose a new resource scheduling analytical model for ring-all-reduce deep learning, which covers a wide range of objectives in DDL performance optimization (e.g., excessive training avoidance, energy efficiency, fairness); ii) Based on the proposed performance analytical model, we develop an efficient resource scheduling algorithm called GADGET (greedy ring-all-reduce distributed graph embedding technique), which enjoys a provable strong performance guarantee; iii) We conduct extensive trace-driven experiments to demonstrate the effectiveness of the GADGET approach and its superiority over the state of the art.
Menglu Yu, Bo Ji 0001, Chuan Wu 0001, Hridesh Rajan, Jia Liu 0002
INFOCOM1
2022 On scheduling ring-all-reduce learning jobs in multi-tenant GPU clusters with communication contention
abstract
Powered by advances in deep learning (DL) techniques, machine learning and artificial intelligence have achieved astonishing successes. However, the rapidly growing needs for DL also led to communication- and resource-intensive distributed training jobs for large-scale DL training, which are typically deployed over GPU clusters. To sustain the ever-increasing demand for DL training, the so-called "ring-all-reduce" (RAR) technologies have recently emerged as a favorable computing architecture to efficiently process network communication and computation load in GPU clusters. The most salient feature of RAR is that it removes the need for dedicated parameter servers, thus alleviating the potential communication bottleneck. However, when multiple RAR-based DL training jobs are deployed over GPU clusters, communication bottlenecks could still occur due to contentions between DL training jobs. So far, there remains a lack of theoretical understanding on how to design contention-aware resource scheduling algorithms for RAR-based DL training jobs, which motivates us to fill this gap in this work. Our main contributions are three-fold: i) We develop a new analytical model that characterizes both communication overhead related to the worker distribution of the job and communication contention related to the co-location of different jobs; ii) Based on the proposed analytical model, we formulate the problem as a non-convex integer program to minimize the makespan of all RAR-based DL training jobs. To address the unique structure in this problem that is not amenable for optimization algorithm design, we reformulate the problem into an integer linear program that enables provable approximation algorithm design called SJF-BCO (Smallest Job First with Balanced Contention and Overhead); and iii) We conduct extensive experiments to show the superiority of SJF-BCO over existing schedulers. Collectively, our results contribute to the state-of-the-art of distributed GPU system optimization and algorithm design.
Menglu Yu, Bo Ji 0001, Hridesh Rajan, Jia Liu 0002
MobiHoc1
2021 A Sum-of-Ratios Multi-Dimensional-Knapsack Decomposition for DNN Resource Scheduling
abstract
In recent years, to sustain the resource-intensive computational needs for training deep neural networks (DNNs), it is widely accepted that exploiting the parallelism in large-scale computing clusters is critical for the efficient deployments of DNN training jobs. However, existing resource schedulers for traditional computing clusters are not well suited for DNN training, which results in unsatisfactory job completion time performance. The limitations of these resource scheduling schemes motivate us to propose a new computing cluster resource scheduling framework that is able to leverage the special layered structure of DNN jobs and significantly improve their job completion times. Our contributions in this paper are three-fold: i) We develop a new resource scheduling analytical model by considering DNN's layered structure, which enables us to analytically formulate the resource scheduling optimization problem for DNN training in computing clusters; ii) Based on the proposed performance analytical model, we then develop an efficient resource scheduling algorithm based on the widely adopted parameter-server architecture using a sum-of-ratios multi-dimensional-knapsack decomposition (SMD) method to offer strong performance guarantee; iii) We conduct extensive numerical experiments to demonstrate the effectiveness of the proposed schedule algorithm and its superior performance over the state of the art.
Menglu Yu, Chuan Wu 0001, Bo Ji 0001, Jia Liu 0002
INFOCOM1