EDBT 2026 Demo / reviewers in the wild / expert
Yemao Xu
dblp:184/9221
· DBLP profile ↗
8ranked-venue papers
3as first author
5since 2021 · last 2023
0000-0003-3639-5449ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | SSD-SGD: Communication Sparsification for Distributed Deep Learning TrainingabstractIntensive communication and synchronization cost for gradients and parameters is the well-known bottleneck of distributed deep learning training. Based on the observations that Synchronous SGD (SSGD) obtains good convergence accuracy while asynchronous SGD (ASGD) delivers a faster raw training speed, we propose Several Steps Delay SGD (SSD-SGD) to combine their merits, aiming at tackling the communication bottleneck via communication sparsification. SSD-SGD explores both global synchronous updates in the parameter servers and asynchronous local updates in the workers in each periodic iteration. The periodic and flexible synchronization makes SSD-SGD achieve good convergence accuracy and fast training speed. To the best of our knowledge, we strike the new balance between synchronization quality and communication sparsification, and improve the tradeoff between accuracy and training speed. Specifically, the core components of SSD-SGD include proper warm-up stage, steps delay stage, and the novel algorithm of global gradient for local update (GLU). GLU is critical for local update operations by using global gradient information to effectively compensate for the delayed local weights. Furthermore, we implement SSD-SGD on MXNet framework and comprehensively evaluate its performance with CIFAR-10 and ImageNet datasets. Experimental results show that SSD-SGD can accelerate distributed training speed under different experimental configurations, by up to 110% (or 2.1× of the original speed), while achieving good convergence accuracy. Yemao Xu, Dezun Dong, Dongsheng Wang 0004, Enda Yu, Weixia Xu 0001, Xiangke Liao |
ACM Trans. Archit. Code Optim. | 1 |
| 2022 | CP-SGD: Distributed stochastic gradient descent with compression and periodic compensation
Enda Yu, Dezun Dong, Yemao Xu, Shuo Ouyang, Xiangke Liao |
J. Parallel Distributed Comput. | 3 |
| 2021 | CD-SGD: Distributed Stochastic Gradient Descent with Compression and Delay CompensationabstractCommunication overhead is the key challenge for distributed training. Gradient compression is a widely used approach to reduce communication traffic. When combining with a parallel communication mechanism method like pipeline, gradient compression technique can greatly alleviate the impact of communication overhead. However, there exist two problems of gradient compression technique to be solved. Firstly, gradient compression brings in extra computation cost, which will delay the next training iteration. Secondly, gradient compression usually leads to a decrease in convergence accuracy. In this paper, we combine parallel mechanism with gradient quantization and delayed full-gradient compensation, and propose a new distributed optimization method named CD-SGD, which can hide the overhead of gradient compression, overlap part of the communication and obtain high convergence accuracy. The local update operation in CD-SGD allows the next iteration to be launched quickly without waiting for the completion of gradient compression and the current communication process. Besides, the accuracy loss caused by gradient compression is solved by k-step correction method introduced in CD-SGD. We prove that CD-SGD has convergence guarantee and it achieves at least convergence rate. We conduct extensive experiments on MXNet to verify the convergence properties and scaling performance of CD-SGD. Experimental results on a 16-GPU cluster show that convergence accuracy of CD-SGD is close to or even slightly better than that of S-SGD, and its end-to-end time is 30 less than 2-bit gradient compression under a 56Gbps bandwidth environment. Enda Yu, Dezun Dong, Yemao Xu, Shuo Ouyang, Xiangke Liao |
ICPP | 3 |
| 2021 | FastHorovod: Expediting Parallel Message-Passing Schedule for Distributed DNN TrainingabstractLarge-scale deep neural networks training have been widely deployed on dense-GPU public cloud clusters. Intensive communication and synchronization cost for gradients and parameters is becoming the bottleneck of distributed deep learning training. Horovod is one of the most popular distributed communication frameworks to address the scale-out issue of deep learning training on GPU clusters. Existing public-cloud GPU datacenters, such as Amazon EC2 and Alibaba GPU cloud, are usually equipped with commodity high-speed Ethernet and TCP networking. In current vanilla Horovod, however, we observe that one GPU device is merely associated with at most one proxy communication process. The proxy process is responsible for dealing with all the communication operations of parameter all-reduce for one or multiple GPUs. Such configuration makes communication interface based on TCP protocols suffer from limited network goodput and incur training performance penalties. In this paper, we make the first attempt to improve the message passing interface of Horovod and address the mismatching between the computation and communication capability when deploying Horovod in TCP-based public-cloud GPU clusters. We propose FastHorovod to exploit more cost-efficient auxiliary communication processes on CPU to expedite parallel message-passing schedule for GPU. We conduct extensive experiments against state-of-the-art Horovod. The experiment results show that our design can significantly accelerate the distributed training communication on TCP-based public-cloud GPU clusters, and FastHorovod improves the training speed of AlexNet and VGG16 models by 64.5% and 72.6% respectively. Yanghai Wang, Dezun Dong, Yemao Xu, Shuo Ouyang, Xiangke Liao |
ISCC | 3 |
| 2021 | Communication optimization strategies for distributed deep neural network training: A survey
Shuo Ouyang, Dezun Dong, Yemao Xu, Liquan Xiao |
J. Parallel Distributed Comput. | 3 |
| 2020 | OD-SGD: One-Step Delay Stochastic Gradient Descent for Distributed TrainingabstractThe training of modern deep learning neural network calls for large amounts of computation, which is often provided by GPUs or other specific accelerators. To scale out to achieve faster training speed, two update algorithms are mainly applied in the distributed training process, i.e., the Synchronous SGD algorithm (SSGD) and Asynchronous SGD algorithm (ASGD). SSGD obtains good convergence point while the training speed is slowed down by the synchronous barrier. ASGD has faster training speed but the convergence point is lower when compared to SSGD. To sufficiently utilize the advantages of SSGD and ASGD, we propose a novel technology named One-step Delay SGD (OD-SGD) to combine their strengths in the training process. Therefore, we can achieve similar convergence point and training speed as SSGD and ASGD separately. To the best of our knowledge, we make the first attempt to combine the features of SSGD and ASGD to improve distributed training performance. Each iteration of OD-SGD contains a global update in the parameter server node and local updates in the worker nodes, the local update is introduced to update and compensate the delayed local weights. We evaluate our proposed algorithm on MNIST, CIFAR-10, and ImageNet datasets. Experimental results show that OD-SGD can obtain similar or even slightly better accuracy than SSGD, while its training speed is much faster, which even exceeds the training speed of ASGD. Yemao Xu, Dezun Dong, Weixia Xu 0001, Xiangke Liao |
ACM Trans. Archit. Code Optim. | 1 |
| 2019 | SketchDLC: A Sketch on Distributed Deep Learning Communication via Trace CapturingabstractWith the fast development of deep learning (DL), the communication is increasingly a bottleneck for distributed workloads, and a series of optimization works have been done to scale out successfully. Nevertheless, the network behavior has not been investigated much yet. We intend to analyze the network behavior and then carry out some research through network simulation. Under this circumstance, an accurate communication measurement is necessary, as it is an effective way to study the network behavior and the basis for accurate simulation. Therefore, we propose to capture the deep learning communication (DLC) trace to achieve the measurement. To the best of our knowledge, we make the first attempt to capture the communication trace for DL training. In this article, we first provide detailed analyses about the communication mechanism of MXNet, which is a representative framework for distributed DL. Secondly, we define the DLC trace format to describe and record the communication behaviors. Third, we present the implementation of method for trace capturing. Finally, we make some statistics and analyses about the distributed DL training, including communication pattern, overlap ratio between computation and communication, computation overhead, synchronization overhead, update overhead, and so forth. Both the statistics and analyses are based on the trace files captured in a cluster with six machines. On the one hand, our trace files provide a sketch on the DLC, which contributes to understanding the communication details. On the other hand, the captured trace files can be used for figuring out various overheads, as they record the communication behaviors of each node. Yemao Xu, Dezun Dong, Weixia Xu 0001, Xiangke Liao |
ACM Trans. Archit. Code Optim. | 1 |
| 2016 | Dynamic Power-Performance Adjustment on Clustered Multi-Threading ProcessorsabstractDynamic optimization techniques such as dynamic voltage and frequency scaling (DVFS), power gating (PG) and thread migration are widely used in current multi-core processor platforms to boost performance, lower power and improve energy efficiency. To obtain the best optimization results, an important issue is to predict the performance of each thread quantitatively. In this paper, a separated and clustered performance predictor, called SCP, is proposed oriented to clustered multi- threading (CMT) processors. SCP model is implemented in AMD's FX-8320 processor and its accuracy is about 6% for SPEC CPU 2006 benchmark suite. To illustrate the application and effectiveness of SCP model furthermore, we propose PPEP-SCP model for power capping, which is a combination of SCP model and PPEP model. Compared to power capping strategy based on PPEP model, strategy based on PPEP-SCP performs better in terms of the performance and the difference between the actual and the target power consumption, because PPEP-SCP based strategy can perform optimization through both DVFS and PG techniques. Li Shen 0007, Zhiying Wang 0003, Yemao Xu |
NAS | 5 |