Yixin Bao

dblp:179/2382 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0002-6921-2154ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Systems, architecture and hardware · 3 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
9 papers
Cloud and datacenter computing · 42% Distributed systems · 25% Parallel and multicore computing · 18%
Artificial intelligence
7 papers
Learning paradigms · 43% Efficient and distributed learning · 40% Learning theory · 14%

Topics — the 28 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems › distributed machine learning
distributed training
1.732024
Optimizing Task Placement and Online Scheduling for Distributed GNN Training Acceleration in Heterogeneous Systems · IEEE/ACM Trans. Netw. 2024
Optimizing Task Placement and Online Scheduling for Distributed GNN Training Acceleration · INFOCOM 2022
A generic communication scheduler for distributed DNN training acceleration · SOSP 2019
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
1.742023
Deep Learning-Based Job Placement in Distributed Machine Learning Clusters With Heterogeneous Workloads · IEEE/ACM Trans. Netw. 2023
Deep Learning-based Job Placement in Distributed Machine Learning Clusters · INFOCOM 2019
Online Job Scheduling in Distributed Machine Learning Clusters · INFOCOM 2018
Embedded and real-time systems › multicore real-time systems
task mapping and scheduling
1.322024
Optimizing Task Placement and Online Scheduling for Distributed GNN Training Acceleration in Heterogeneous Systems · IEEE/ACM Trans. Netw. 2024
Optimizing Task Placement and Online Scheduling for Distributed GNN Training Acceleration · INFOCOM 2022
Machine learning › Efficient and distributed learning
distributed training
1.242020
Preemptive All-reduce Scheduling for Expediting Distributed DNN Training · INFOCOM 2020
Deep Learning-based Job Placement in Distributed Machine Learning Clusters · INFOCOM 2019
Online Job Scheduling in Distributed Machine Learning Clusters · INFOCOM 2018
Cloud and datacenter computing › job scheduling
job placement
1.022023
Deep Learning-Based Job Placement in Distributed Machine Learning Clusters With Heterogeneous Workloads · IEEE/ACM Trans. Netw. 2023
Deep Learning-based Job Placement in Distributed Machine Learning Clusters · INFOCOM 2019
Machine learning › Learning paradigms › continual learning
catastrophic forgetting
0.912025
Understanding the Forgetting of (Replay-based) Continual Learning via Feature Learning: Angle Matters · ICML 2025
Machine learning › Learning paradigms
continual learning
0.912025
Understanding the Forgetting of (Replay-based) Continual Learning via Feature Learning: Angle Matters · ICML 2025
Machine learning › Learning theory › neural network theory
feature learning theory
0.912025
Understanding the Forgetting of (Replay-based) Continual Learning via Feature Learning: Angle Matters · ICML 2025
Machine learning › Learning paradigms › continual learning
rehearsal-based continual learning
0.912025
Understanding the Forgetting of (Replay-based) Continual Learning via Feature Learning: Angle Matters · ICML 2025
Cloud and datacenter computing › cluster resource management and scheduling › cluster scheduling
deep learning cluster scheduling
0.822021
DL2: A Deep Learning-Driven Scheduler for Deep Learning Clusters · IEEE Trans. Parallel Distributed Syst. 2021
Optimus: an efficient dynamic resource scheduler for deep learning clusters · EuroSys 2018
Parallel and multicore computing › parallel scheduling
communication scheduling
0.822020
Preemptive All-reduce Scheduling for Expediting Distributed DNN Training · INFOCOM 2020
A generic communication scheduler for distributed DNN training acceleration · SOSP 2019
Distributed systems › distributed machine learning
distributed graph neural network training
0.812024
Optimizing Task Placement and Online Scheduling for Distributed GNN Training Acceleration in Heterogeneous Systems · IEEE/ACM Trans. Netw. 2024
Parallel and multicore computing › parallel scheduling
heterogeneous multiprocessor scheduling
0.812024
Optimizing Task Placement and Online Scheduling for Distributed GNN Training Acceleration in Heterogeneous Systems · IEEE/ACM Trans. Netw. 2024
Parallel and multicore computing › parallel scheduling › resource-aware scheduling
interference-aware scheduling
0.712023
Deep Learning-Based Job Placement in Distributed Machine Learning Clusters With Heterogeneous Workloads · IEEE/ACM Trans. Netw. 2023
Distributed systems › distributed machine learning › distributed training
distributed GNN training
0.612022
Optimizing Task Placement and Online Scheduling for Distributed GNN Training Acceleration · INFOCOM 2022
Cloud and datacenter computing
cluster resource management and scheduling
0.512021
DL2: A Deep Learning-Driven Scheduler for Deep Learning Clusters · IEEE Trans. Parallel Distributed Syst. 2021
Cloud and datacenter computing
resource allocation
0.512021
DL2: A Deep Learning-Driven Scheduler for Deep Learning Clusters · IEEE Trans. Parallel Distributed Syst. 2021
Machine learning › Efficient and distributed learning › distributed training › communication-efficient training
communication optimization
0.412020
Preemptive All-reduce Scheduling for Expediting Distributed DNN Training · INFOCOM 2020
High-performance computing › collective communication
all-reduce
0.412020
Preemptive All-reduce Scheduling for Expediting Distributed DNN Training · INFOCOM 2020
Machine learning › Efficient and distributed learning › distributed training
gradient aggregation
0.412019
A generic communication scheduler for distributed DNN training acceleration · SOSP 2019
Machine learning › Efficient and distributed learning › distributed training › distributed training systems
parameter server
0.312018
Online Job Scheduling in Distributed Machine Learning Clusters · INFOCOM 2018
Cloud and datacenter computing › cluster resource management and scheduling
cluster scheduling
0.312018
Optimus: an efficient dynamic resource scheduler for deep learning clusters · EuroSys 2018
Cloud and datacenter computing › resource allocation
dynamic resource allocation
0.312018
Optimus: an efficient dynamic resource scheduler for deep learning clusters · EuroSys 2018
Cloud and datacenter computing
job scheduling
0.312018
Online Job Scheduling in Distributed Machine Learning Clusters · INFOCOM 2018
Distributed systems
distributed machine learning
0.212023
Deep Learning-Based Job Placement in Distributed Machine Learning Clusters With Heterogeneous Workloads · IEEE/ACM Trans. Netw. 2023
Parallel and multicore computing › parallel scheduling
training job scheduling
0.212023
Deep Learning-Based Job Placement in Distributed Machine Learning Clusters With Heterogeneous Workloads · IEEE/ACM Trans. Netw. 2023
Machine learning › Graph learning
graph neural network
0.212022
Optimizing Task Placement and Online Scheduling for Distributed GNN Training Acceleration · INFOCOM 2022
Interconnection networks and networks-on-chip
remote direct memory access
0.112019
A generic communication scheduler for distributed DNN training acceleration · SOSP 2019

Methods — techniques the papers use, named apart from their topics

theoretical analysis · 1.9simulation · 1.9online scheduling · 1.9deep reinforcement learning · 1.4actor-critic · 1.4experience replay · 1.0two-layer CNN · 0.9signal-noise data model · 0.9polynomial ReLU · 0.9testbed experiments · 0.8sequence-to-sequence model · 0.7attention · 0.7neural network · 0.5preemptive scheduling · 0.4dynamic programming · 0.4
YearPublicationVenuePosition
2025 Understanding the Forgetting of (Replay-based) Continual Learning via Feature Learning: Angle Matters
abstract
Continual learning (CL) is crucial for advancing human-level intelligence, but its theoretical understanding, especially regarding factors influencing forgetting, is still relatively limited. This work aims to build a unified theoretical framework for understanding CL using feature learning theory. Different from most existing studies that analyze forgetting under linear regression model or lazy training, we focus on a more practical two-layer convolutional neural network (CNN) with polynomial ReLU activation for sequential tasks within a signal-noise data model. Specifically, we theoretically reveal how the angle between task signal vectors influences forgetting that: acute or small obtuse angles lead to benign forgetting, whereas larger obtuse angles result in harmful forgetting. Furthermore, we demonstrate that the replay method alleviates forgetting by expanding the range of angles corresponding to benign forgetting. Our theoretical results suggest that mid-angle sampling, which selects examples with moderate angles to the prototype, can enhance the replay method’s ability to mitigate forgetting. Experiments on synthetic and real-world datasets confirm our theoretical results and highlight the effectiveness of our mid-angle sampling strategy.
Shiyuan Ren, Wei Huang 0034, Miao Zhang 0001, Xiang Deng 0002, Yixin Bao, Liqiang Nie
ICML6
2024 Optimizing Task Placement and Online Scheduling for Distributed GNN Training Acceleration in Heterogeneous Systems
abstract
Training Graph Neural Networks (GNNs) on large graphs is resource-intensive and time-consuming, mainly due to the large graph data that cannot be fit into the memory of a single machine, but have to be fetched from distributed graph storage and processed on the go. Unlike distributed deep neural network (DNN) training, the bottleneck in distributed GNN training lies largely in large graph data transmission for constructing mini-batches of training samples. Existing solutions often advocate data-computation colocation, and do not work well with limited resources and heterogeneous training devices in heterogeneous clusters. The potentials of strategical task placement and optimal scheduling of data transmission and task execution have not been well explored. This paper designs an efficient algorithm framework for task placement and execution scheduling of distributed GNN training in heterogeneous systems, to better resource utilization, improve execution pipelining, and expedite training completion. Our framework consists of two modules: (i) an online scheduling algorithm that schedules the execution of training tasks, and the data transmission plan; and (ii) an exploratory task placement scheme that decides the placement of each training task. We conduct thorough theoretical analysis, testbed experiments and simulation studies, and observe up to 48% training speed-up with our algorithm as compared to representative baselines in our testbed settings.
Ziyue Luo, Yixin Bao, Chuan Wu 0001
IEEE/ACM Trans. Netw.2
2023 Deep Learning-Based Job Placement in Distributed Machine Learning Clusters With Heterogeneous Workloads
abstract
Nowadays, most leading IT companies host a variety of distributed machine learning (ML) workloads in ML clusters to support AI-driven services, such as speech recognition, machine translation, and image processing. While multiple jobs are executed concurrently in a shared cluster to improve resource utilization, interference among co-located ML jobs can lead to significant performance downgrade. Existing cluster schedulers, such as YARN and Mesos, are interference-agnostic in their job placement, leading to suboptimal resource efficiency and usage. Some literature has studied interference-aware job placement policy, but relies on detailed workload profiling and interference modeling, which is not a general solution. In this work, we present Harmony, a deep learning-driven ML cluster scheduler that places heterogeneous training jobs (either with parameter server architecture or all-reduce architecture) in a manner that minimizes interference and maximizes performance (i.e., training completion time minimization). The design of Harmony is based on a carefully designed deep reinforcement learning (DRL) framework enhanced with reward modeling. The DRL integrates a dynamic sequence-to-sequence model with the state-of-the-art techniques to stabilize training and improve convergence, including actor-critic algorithm, job-aware action space exploration, multi-head attention, and experience replay. In view of a common lack of reward samples corresponding to different placement decisions, we build an auxiliary sequence-to-sequence reward prediction model, which is trained with historical samples and used for producing reward for unseen placement. Experiments using real ML workloads in a Kubernetes cluster of 6 GPU servers show that Harmony outperforms representative schedulers by 16%–42% in terms of average job completion time.
Yixin Bao, Yanghua Peng, Chuan Wu 0001
IEEE/ACM Trans. Netw.1
2022 Optimizing Task Placement and Online Scheduling for Distributed GNN Training Acceleration
abstract
Training Graph Neural Networks (GNN) on large graphs is resource-intensive and time-consuming, mainly due to the large graph data that cannot be fit into the memory of a single machine, but have to be fetched from distributed graph storage and processed on the go. Unlike distributed deep neural network (DNN) training, the bottleneck in distributed GNN training lies largely in large graph data transmission for constructing mini-batches of training samples. Existing solutions often advocate data-computation colocation, and do not work well with limited resources where the colocation is infeasible. The potentials of strategical task placement and optimal scheduling of data transmission and task execution have not been well explored. This paper designs an efficient algorithm framework for task placement and execution scheduling of distributed GNN training, to better resource utilization, improve execution pipelining, and expediting training completion. Our framework consists of two modules: (i) an online scheduling algorithm that schedules the execution of training tasks, and the data transmission plan; and (ii) an exploratory task placement scheme that decides the placement of each training task. We conduct thorough theoretical analysis, testbed experiments and simulation studies, and observe up to 67% training speed-up with our algorithm as compared to representative baselines.
Ziyue Luo, Yixin Bao, Chuan Wu 0001
INFOCOM2
2021 DL2: A Deep Learning-Driven Scheduler for Deep Learning Clusters
abstract
Efficient resource scheduling is essential for maximal utilization of expensive deep learning (DL) clusters. Existing cluster schedulers either are agnostic to machine learning (ML) workload characteristics, or use scheduling heuristics based on operators' understanding of particular ML framework and workload, which are less efficient or not general enough. In this article, we show that DL techniques can be adopted to design a generic and efficient scheduler. Specifically, we propose DL2, a DL-driven scheduler for DL clusters, targeting global training job expedition by dynamically resizing resources allocated to jobs. DL2 advocates a joint supervised learning and reinforcement learning approach: a neural network is warmed up via offline supervised learning based on job traces produced by the existing cluster scheduler; then the neural network is plugged into the live DL cluster, fine-tuned by reinforcement learning carried out throughout the training progress of the DL jobs, and used for deciding job resource allocation in an online fashion. We implement DL2 on Kubernetes and enable dynamic resource scaling in DL jobs on MXNet. Extensive evaluation shows that DL2 outperforms fairness scheduler (i.e., DRF) by 44.1 percent and expert heuristic scheduler (i.e., Optimus) by 17.5 percent in terms of average job completion time.
Yanghua Peng, Yixin Bao, Yangrui Chen, Chuan Wu 0001, Wei Lin 0016
IEEE Trans. Parallel Distributed Syst.2
2020 Elastic parameter server load distribution in deep learning clusters
abstract
In distributed DNN training, parameter servers (PS) can become performance bottlenecks due to PS stragglers, caused by imbalanced parameter distribution, bandwidth contention, or computation interference. Few existing studies have investigated efficient parameter (aka load) distribution among PSs. We observe significant training inefficiency with the current parameter assignment in representative machine learning frameworks (e.g., MXNet, TensorFlow), and big potential for training acceleration with better PS load distribution. We design PSLD, a dynamic parameter server load distribution scheme, to mitigate PS straggler issues and accelerate distributed model training in the PS architecture. An exploitation-exploration method is carefully designed to scale in and out parameter servers and adjust parameter distribution among PSs on the go. We also design an elastic PS scaling module to carry out our scheme with little interruption to the training process. We implement our module on top of open-source PS architectures, including MXNet and BytePS. Testbed experiments show up to 2.86x speed-up in model training with PSLD, for different ML models under various straggler settings.
Yangrui Chen, Yanghua Peng, Yixin Bao, Chuan Wu 0001, Yibo Zhu 0001, Chuanxiong Guo
SoCC3
2020 Preemptive All-reduce Scheduling for Expediting Distributed DNN Training
abstract
Data-parallel training is widely used for scaling DNN training over large datasets, using the parameter server or all-reduce architecture. Communication scheduling has been promising to accelerate distributed DNN training, which aims to overlap communication with computation by scheduling the order of communication operations. We identify two limitations of previous communication scheduling work. First, layer-wise computation graph has been a common assumption, while modern machine learning frameworks (e.g., TensorFlow) use a sophisticated directed acyclic graph (DAG) representation as the execution model. Second, the default sizes of tensors are often less than optimal for transmission scheduling and bandwidth utilization. We propose PACE, a communication scheduler that preemptively schedules (potentially fused) all-reduce tensors based on the DAG of DNN training, guaranteeing maximal overlapping of communication with computation and high bandwidth utilization. The scheduler contains two integrated modules: given a DAG, we identify the best tensor-preemptive communication schedule that minimizes the training time; exploiting the optimal communication scheduling as an oracle, a dynamic programming approach is developed for generating a good DAG, which merges small communication tensors for efficient bandwidth utilization. Experiments in a GPU testbed show that PACE accelerates training with representative system configurations, achieving up to 36% speed-up compared with state-of-the-art solutions.
Yixin Bao, Yanghua Peng, Yangrui Chen, Chuan Wu 0001
INFOCOM1
2019 Deep Learning-based Job Placement in Distributed Machine Learning Clusters
abstract
Production machine learning (ML) clusters commonly host a variety of distributed ML workloads, e.g., speech recognition, machine translation. While server sharing among jobs improves resource utilization, interference among co-located ML jobs can lead to significant performance downgrade. Existing cluster schedulers (e.g., Mesos) are interference-oblivious in their job placement, causing suboptimal resource efficiency. Interference-aware job placement has been studied in the literature, but was treated using detailed workload profiling and interference modeling, which is not a general solution. This paper presents Harmony, a deep learning-driven ML cluster scheduler that places training jobs in a manner that minimizes interference and maximizes performance (i.e., training completion time). Harmony is based on a carefully designed deep reinforcement learning (DRL) framework augmented with reward modeling. The DRL employs state-of-the-art techniques to stabilize training and improve convergence, including actor-critic algorithm, job-aware action space exploration and experience replay. In view of a common lack of reward samples corresponding to different placement decisions, we build an auxiliary reward prediction model, which is trained using historical samples and used for producing reward for unseen placement. Experiments using real ML workloads in a Kubernetes cluster of 6 GPU servers show that Harmony outperforms representative schedulers by 25% in terms of average job completion time.
Yixin Bao, Yanghua Peng, Chuan Wu 0001
INFOCOM1
2019 A generic communication scheduler for distributed DNN training acceleration
abstract
We present ByteScheduler, a generic communication scheduler for distributed DNN training acceleration. ByteScheduler is based on our principled analysis that partitioning and rearranging the tensor transmissions can result in optimal results in theory and good performance in real-world even with scheduling overhead. To make ByteScheduler work generally for various DNN training frameworks, we introduce a unified abstraction and a Dependency Proxy mechanism to enable communication scheduling without breaking the original dependencies in framework engines. We further introduce a Bayesian Optimization approach to auto-tune tensor partition size and other parameters for different training models under various networking conditions. ByteScheduler now supports TensorFlow, PyTorch, and MXNet without modifying their source code, and works well with both Parameter Server (PS) and all-reduce architectures for gradient synchronization, using either TCP or RDMA. Our experiments show that ByteScheduler accelerates training with all experimented system configurations and DNN models, by up to 196% (or 2.96X of original speed).
Yanghua Peng, Yibo Zhu 0001, Yangrui Chen, Yixin Bao, Bairen Yi, Chang Lan, Chuan Wu 0001, Chuanxiong Guo
SOSP4
2018 AutoHighlight : Automatic Highlights Detection and Segmentation in Soccer Matches
abstract
More and more football fans prefer to watch football highlights because it takes too much time to watch the entire football game. At the same time, many TV stations and sports companies need to store and retrieve fragments of different events in football matches. The traditional way is to analyze and cut the football match manually. This method is time consuming and the process largely depends on personal choice, which results in unsecured accuracy. Therefore, it’s necessary to develop a fast and accurate system for automatic summarization and analysis of football matches video. Automatic event detection and segmentation in football matches is about detecting specific event and extracting the related moments for football viewers and companies. This paper presents a deep learning based event detection and segmentation system for emphasizing specific events during football matches. The proposed system use live text of football matches as auxiliary information to segment the important event segmentation. We firstly use the deep learning model to tag each live text. Second, we merge successive live texts with the same label. Then we go to the football match to segment corresponding segmentation to the merged live text. Finally, if necessary, we run pure video based event detection model to these segmentation for finer grained segmentation. Experiments on real football matches show encouraging results. The proposed system greatly improves the accuracy and speed of segmentation.
Kaiyu Tang, Yixin Bao, Yining Lin
IEEE BigData2
2018 Optimus: an efficient dynamic resource scheduler for deep learning clusters
abstract
Deep learning workloads are common in today's production clusters due to the proliferation of deep learning driven AI services (e.g., speech recognition, machine translation). A deep learning training job is resource-intensive and time-consuming. Efficient resource scheduling is the key to the maximal performance of a deep learning cluster. Existing cluster schedulers are largely not tailored to deep learning jobs, and typically specifying a fixed amount of resources for each job, prohibiting high resource efficiency and job performance. This paper proposes Optimus, a customized job scheduler for deep learning clusters, which minimizes job training time based on online resource-performance models. Optimus uses online fitting to predict model convergence during training, and sets up performance models to accurately estimate training speed as a function of allocated resources in each job. Based on the models, a simple yet effective method is designed and used for dynamically allocating resources and placing deep learning tasks to minimize job completion time. We implement Optimus on top of Kubernetes, a cluster manager for container orchestration, and experiment on a deep learning cluster with 7 CPU servers and 6 GPU servers, running 9 training jobs using the MXNet framework. Results show that Optimus outperforms representative cluster schedulers by about 139% and 63% in terms of job completion time and makespan, respectively.
Yanghua Peng, Yixin Bao, Yangrui Chen, Chuan Wu 0001, Chuanxiong Guo
EuroSys2
2018 Online Job Scheduling in Distributed Machine Learning Clusters
abstract
Nowadays large-scale distributed machine learning systems have been deployed to support various analytics and intelligence services in IT firms. To train a large dataset and derive the prediction/inference model, e.g., a deep neural network, multiple workers are run in parallel to train partitions of the input dataset, and update shared model parameters. In a shared cluster handling multiple training jobs, a fundamental issue is how to efficiently schedule jobs and set the number of concurrent workers to run for each job, such that server resources are maximally utilized and model training can be completed in time. Targeting a distributed machine learning system using the parameter server framework, w e design an online algorithm for scheduling the arriving jobs and deciding the adjusted numbers of concurrent workers and parameter servers for each job over its course, to maximize overall utility of all jobs, contingent on their completion times. Our online algorithm design utilizes a primal-dual framework coupled with efficient dual subroutines, achieving good long-term performance guarantees with polynomial time complexity. Practical effectiveness of the online algorithm is evaluated using trace-driven simulation and testbed experiments, which demonstrate its outperformance as compared to commonly adopted scheduling algorithms in today's cloud systems.
Yixin Bao, Yanghua Peng, Chuan Wu 0001, Zongpeng Li
INFOCOM1
2016 Online influence maximization in non-stationary Social Networks
abstract
Social networks have been popular platforms for information propagation. An important use case is viral marketing: given a promotion budget, an advertiser can choose some influential users as the seed set and provide them free or discounted sample products; in this way, the advertiser hopes to increase the popularity of the product in the users' friend circles by the world-of-mouth effect, and thus maximizes the number of users that information of the production can reach. There has been a body of literature studying the influence maximization problem. Nevertheless, the existing studies mostly investigate the problem on a one-off basis, assuming fixed known influence probabilities among users, or the knowledge of the exact social network topology. In practice, the social network topology and the influence probabilities are typically unknown to the advertiser, which can be varying over time, i.e., in cases of newly established, strengthened or weakened social ties. In this paper, we focus on a dynamic non-stationary social network and design a randomized algorithm, RSB, based on multi-armed bandit optimization, to maximize influence propagation over time. The algorithm produces a sequence of online decisions and calibrates its explore-exploit strategy utilizing outcomes of previous decisions. It is rigorously proven to achieve an upper-bounded regret in reward and applicable to large-scale social networks. Practical effectiveness of the algorithm is evaluated using real-world datasets, which demonstrates that our algorithm outperforms previous stationary methods under non-stationary conditions.
Yixin Bao, Zhi Wang 0001, Chuan Wu 0001, Francis C. M. Lau 0001
IWQoS1
2016 Face database generation based on text-video correlation
Dan Zeng 0001, Yixin Bao, Fan Zhao 0004, Qi Tian 0001
Neurocomputing2