EDBT 2026 Demo / reviewers in the wild / expert
Shuo Ouyang
dblp:197/3973
· DBLP profile ↗
6ranked-venue papers
1as first author
5since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | CP-SGD: Distributed stochastic gradient descent with compression and periodic compensation
Enda Yu, Dezun Dong, Yemao Xu, Shuo Ouyang, Xiangke Liao |
J. Parallel Distributed Comput. | 4 |
| 2021 | CD-SGD: Distributed Stochastic Gradient Descent with Compression and Delay CompensationabstractCommunication overhead is the key challenge for distributed training. Gradient compression is a widely used approach to reduce communication traffic. When combining with a parallel communication mechanism method like pipeline, gradient compression technique can greatly alleviate the impact of communication overhead. However, there exist two problems of gradient compression technique to be solved. Firstly, gradient compression brings in extra computation cost, which will delay the next training iteration. Secondly, gradient compression usually leads to a decrease in convergence accuracy. In this paper, we combine parallel mechanism with gradient quantization and delayed full-gradient compensation, and propose a new distributed optimization method named CD-SGD, which can hide the overhead of gradient compression, overlap part of the communication and obtain high convergence accuracy. The local update operation in CD-SGD allows the next iteration to be launched quickly without waiting for the completion of gradient compression and the current communication process. Besides, the accuracy loss caused by gradient compression is solved by k-step correction method introduced in CD-SGD. We prove that CD-SGD has convergence guarantee and it achieves at least convergence rate. We conduct extensive experiments on MXNet to verify the convergence properties and scaling performance of CD-SGD. Experimental results on a 16-GPU cluster show that convergence accuracy of CD-SGD is close to or even slightly better than that of S-SGD, and its end-to-end time is 30 less than 2-bit gradient compression under a 56Gbps bandwidth environment. Enda Yu, Dezun Dong, Yemao Xu, Shuo Ouyang, Xiangke Liao |
ICPP | 4 |
| 2021 | FastHorovod: Expediting Parallel Message-Passing Schedule for Distributed DNN TrainingabstractLarge-scale deep neural networks training have been widely deployed on dense-GPU public cloud clusters. Intensive communication and synchronization cost for gradients and parameters is becoming the bottleneck of distributed deep learning training. Horovod is one of the most popular distributed communication frameworks to address the scale-out issue of deep learning training on GPU clusters. Existing public-cloud GPU datacenters, such as Amazon EC2 and Alibaba GPU cloud, are usually equipped with commodity high-speed Ethernet and TCP networking. In current vanilla Horovod, however, we observe that one GPU device is merely associated with at most one proxy communication process. The proxy process is responsible for dealing with all the communication operations of parameter all-reduce for one or multiple GPUs. Such configuration makes communication interface based on TCP protocols suffer from limited network goodput and incur training performance penalties. In this paper, we make the first attempt to improve the message passing interface of Horovod and address the mismatching between the computation and communication capability when deploying Horovod in TCP-based public-cloud GPU clusters. We propose FastHorovod to exploit more cost-efficient auxiliary communication processes on CPU to expedite parallel message-passing schedule for GPU. We conduct extensive experiments against state-of-the-art Horovod. The experiment results show that our design can significantly accelerate the distributed training communication on TCP-based public-cloud GPU clusters, and FastHorovod improves the training speed of AlexNet and VGG16 models by 64.5% and 72.6% respectively. Yanghai Wang, Dezun Dong, Yemao Xu, Shuo Ouyang, Xiangke Liao |
ISCC | 4 |
| 2021 | vSketchDLC: A Sketch on Distributed Deep Learning Communication via Fine-grained Tracing Visualization
Yanghai Wang, Shuo Ouyang, Dezun Dong, Enda Yu, Xiangke Liao |
NPC | 2 |
| 2021 | Communication optimization strategies for distributed deep neural network training: A survey
Shuo Ouyang, Dezun Dong, Yemao Xu, Liquan Xiao |
J. Parallel Distributed Comput. | 1 |
| 2019 | A region search evolutionary algorithm for many-objective optimization
Yongqi Liu 0003, Hui Qin, Liqiang Yao, Shuo Ouyang, Jie Li 0038 |
Inf. Sci. | 7 |