Yidi Wu 0001

dblp:213/8964-1 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
9since 2021 · last 2023
0000-0002-3996-9316ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2023 FEC: Efficient Deep Recommendation Model Training with Flexible Embedding Communication
abstract
Embedding-based deep recommendation models (EDRMs), which contain small dense models and large embedding tables, are widely used in industry. Embedding communication constitutes the main cost for the distributed training of EDRMs, and thus we propose two strategies to improve its efficiency, i.e.,embedding tiering andpre-fetching. In particular, embedding tiering uses AllReduce to communicate popular embeddings that are accessed frequently. This is counter-intuitive as embeddings belong to the sparse embedding tables, but reasonable because the access pattern of popular embeddings resembles dense models. Pre-fetching starts communication early for embeddings that receive no updates such that they are removed from the critical path of training. We implement embedding tiering and pre-fetching in a system called FEC and compare it with the state-of-the-art systems on real datasets. The results show that FEC consistently outperforms the existing methods on all datasets, and its speed can be up to 6.65x and 2.42x in terms of embedding communication time and training throughput compared with the best performing baseline.
Kaihao Ma, Xiao Yan 0002, Zhenkun Cai, Yidi Wu 0001, James Cheng
Proc. ACM Manag. Data5
2022 HGL: Accelerating Heterogeneous GNN Training with Holistic Representation and Optimization
abstract
Graph neural networks (GNNs) have shown to significantly improve graph analytics. Existing systems for GNN training are primarily designed for homogeneous graphs. In industry, however, most graphs are actually heterogeneous in nature (i.e., having multiple types of nodes and edges). Existing systems train a heterogeneous GNN (HetGNN) as a composition of homogeneous GNN (HomoGNN) and thus suffer from critical limitations such as lack of memory optimization and limited operator parallelism. To address these limitations, we propose HGL - a heterogeneity-aware system for GNN training. At the core of HGL is an intermediate representation, called HIR, which provides a holistic representation for GNNs and enables cross-relation optimization in HetGNN training. We devise tailored optimizations on HIR, including graph stitching, operator fusion and operator bundling. Compared with DGL and PyG, HGL achieves a speedup from 7 to 22 times for training HetGNNs.
Yuntao Gui, Yidi Wu 0001, Han Yang 0002, Tatiana Jin, Boyang Li 0016, Qihui Zhou, James Cheng, Fan Yu 0004
SC2
2022 TensorOpt: Exploring the Tradeoffs in Distributed DNN Training With Auto-Parallelism
abstract
Effective parallelization strategies are crucial for the performance of distributed deep neural network (DNN) training. Recently, several methods have been proposed to search parallelization strategies but they all optimize a single objective (e.g., execution time, memory consumption) and produce only one strategy. We proposeFrontier Tracking(FT), an efficient algorithm that findsa set of Pareto-optimal parallelization strategiesto explore the best trade-off among different objectives. FT can minimize the memory consumption when the number of devices is limited and fully utilize additional resources to reduce the execution time. Based onFT, we develop a user-friendly system, calledTensorOpt, which allows users to run their distributed DNN training jobs without caring the details about searching and coding parallelization strategies. Experimental results show that TensorOpt is more flexible in adapting to resource availability compared with existing frameworks.
Zhenkun Cai, Xiao Yan 0002, Kaihao Ma, Yidi Wu 0001, James Cheng, Teng Su, Fan Yu 0004
IEEE Trans. Parallel Distributed Syst.4
2022 Elastic Deep Learning in Multi-Tenant GPU Clusters
abstract
We study how to support elasticity, that is, the ability to dynamically adjust the parallelism (i.e., the number of GPUs), for deep neural network (DNN) training in a GPU cluster. Elasticity can benefit multi-tenant GPU cluster management in many ways, for example, achieving various scheduling objectives (e.g., job throughput, job completion time, GPU efficiency) according to cluster load variations, utilizing transient idle resources, and supporting performance profiling, job migration, and straggler mitigation. We propose EDL, which enables elastic deep learning with a simple API and can be easily integrated with existing deep learning frameworks such as TensorFlow and PyTorch. EDL also incorporates techniques that are necessary to reduce the overhead of parallelism adjustments, such as stop-free scaling and dynamic data pipeline. We demonstrate with experiments that EDL can indeed bring significant benefits to the above-listed applications in GPU cluster management.
Yidi Wu 0001, Kaihao Ma, Xiao Yan 0002, Zhi Liu 0002, Zhenkun Cai, James Cheng, Fan Yu 0004
IEEE Trans. Parallel Distributed Syst.1
2021 DGCL: an efficient communication library for distributed GNN training
abstract
Graph neural networks (GNNs) have gained increasing popularity in many areas such as e-commerce, social networks and bio-informatics. Distributed GNN training is essential for handling large graphs and reducing the execution time. However, for distributed GNN training, a peer-to-peer communication strategy suffers from high communication overheads. Also, different GPUs require different remote vertex embeddings, which leads to an irregular communication pattern and renders existing communication planning solutions unsuitable. We propose the distributed graph communication library (DGCL) for efficient GNN training on multiple GPUs. At the heart of DGCL is a communication planning algorithm tailored for GNN training, which jointly considers fully utilizing fast links, fusing communication, avoiding contention and balancing loads on different links. DGCL can be easily adopted to extend existing single-GPU GNN systems to distributed training. We conducted extensive experiments on different datasets and network configurations to compare DGCL with alternative communication schemes. In our experiments, DGCL reduces the communication time of the peer-to-peer communication by 77.5% on average and the training time for an epoch by up to 47%.
Zhenkun Cai, Xiao Yan 0002, Yidi Wu 0001, Kaihao Ma, James Cheng, Fan Yu 0004
EuroSys3
2021 Seastar: vertex-centric programming for graph neural networks
abstract
Graph neural networks (GNNs) have achieved breakthrough performance in graph analytics such as node classification, link prediction and graph clustering. Many GNN training frameworks have been developed, but they are usually designed as a set of manually written, GNN-specific operators plugged into existing deep learning systems, which incurs high memory consumption, poor data locality, and large semantic gap between algorithm design and implementation. This paper proposes the Seastar system, which presents a vertex-centric programming model for GNN training on GPU and provides idiomatic python constructs to enable easy development of novel homogeneous and heterogeneous GNN models. We also propose novel optimizations to produce highly efficient fused GPU kernels for forward and backward passes in GNN training. Compared with the state-of-the art GNN systems, DGL and PyG, Seastar achieves better usability, up to 2 and 8 times less memory consumption, and 14 and 3 times faster execution, respectively.
Yidi Wu 0001, Kaihao Ma, Zhenkun Cai, Tatiana Jin, Boyang Li 0016, Chengguang Zheng, James Cheng, Fan Yu 0004
EuroSys1
2021 Vertex-Centric Visual Programming for Graph Neural Networks
abstract
Graph neural networks (GNNs) have achieved remarkable performance in many graph analytics tasks such as node classification, link prediction and graph clustering. Existing GNN systems (e.g., PyG and DGL) adopt a tensor-centric programming model and train GNNs with manually written operators. Such design results in poor usability due to the large semantic gap between the API and the GNN models, and suffers from inferior efficiency because of high memory consumption and massive data movement. We demonstrateSeastar, a novel GNN training framework that adopts avertex-centric programming paradigm and supportsautomatic kernel generation, to simplify model development and improve training efficiency. We will (i) show how to express GNN models succinctly using a visual "drag-and-drop'' interface or Seastar's vertex-centric python API; (ii) demonstrate the performance advantage of Seastar over existing GNN systems in convergence speed, training throughput and memory consumption; and (iii) illustrate how Seastar's optimizations (e.g., operator fusion and constant folding) improve training efficiency by profiling the run-time performance.
Yidi Wu 0001, Yuntao Gui, Tatiana Jin, James Cheng, Xiao Yan 0002, Peiqi Yin, Yufei Cai, Bo Tang 0016, Fan Yu 0004
SIGMOD Conference1
2021 Scaling Large Production Clusters with Partitioned Synchronization
Yihui Feng, Zhi Liu 0002, Yunjian Zhao, Tatiana Jin, Yidi Wu 0001, James Cheng, Chao Li 0009
USENIX ATC5
2021 Timestamped State Sharing for Stream Analytics
abstract
State access in existing distributed stream processing systems is restricted locally within each operator. However, in advanced stream analytics such as online learning and dynamic graph analytics, enabling state sharing across different operators makes application development easier and stream processing more efficient. In addition, when stream records are timestamped, proper time semantics should be defined for both state updates and fetches. We propose a new state abstraction to address the limitations of existing systems and develop a distributed stream processing system, Nova, with native support for timestamped state sharing. We validate the expressiveness and efficiency of Nova with extensive experiments.
Yunjian Zhao, Zhi Liu 0002, Yidi Wu 0001, Guanxian Jiang, James Cheng, Kunlong Liu, Xiao Yan 0002
IEEE Trans. Parallel Distributed Syst.3
2018 FlexPS: Flexible Parallelism Control in Parameter Server Architecture
abstract
As a general abstraction for coordinating the distributed storage and access of model parameters, the parameter server (PS) architecture enables distributed machine learning to handle large datasets and high dimensional models. Many systems, such as Parameter Server and Petuum, have been developed based on the PS architecture and widely used in practice. However, none of these systems supports changing parallelism during runtime, which is crucial for the efficient execution of machine learning tasks with dynamic workloads. We propose a new system, called FlexPS, which introduces a novel multi-stage abstraction to support flexible parallelism control. With the multi-stage abstraction, a machine learning task can be mapped to a series of stages and the parallelism for a stage can be set according to its workload. Optimizations such as stage scheduler, stage-aware consistency controller, and direct model transfer are proposed for the efficiency of multi-stage machine learning in FlexPS. As a general and complete PS systems, FlexPS also incorporates many optimizations that are not limited to multi-stage machine learning. We conduct extensive experiments using a variety of machine learning workloads, showing that FlexPS achieves significant speedups and resource saving compared with the state-of-the-art PS systems such as Petuum and Multiverso.
Tatiana Jin, Yidi Wu 0001, Zhenkun Cai, Xiao Yan 0002, Fan Yang 0091, Yuying Guo, James Cheng
Proc. VLDB Endow.3