VLDB 2026 Research / reviewers in the wild / expert
Enda Yu
dblp:295/8923
· DBLP profile ↗
12ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0003-2661-0889ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 5 first-author · 9 since 2021Computer networks · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy ServersabstractMixture-of-Experts (MoE) models face memory and PCIe latency bottlenecks when deployed on commodity hardware. Offloading expert weights to CPU memory results in PCIe transfer latency that exceeds GPU computation by several folds. We present PreScope, a prediction-driven expert scheduling system that addresses three key challenges: inaccurate activation prediction, PCIe bandwidth competition, and cross-device scheduling complexity. Our solution includes: 1) Learnable Layer-Aware Predictor (LLaPor) that captures layer-specific expert activation patterns; 2) Prefetch-Aware Cross-Layer Scheduling (PreSched) that generates globally optimal plans balancing prefetching costs and loading overhead; 3) Asynchronous I/O Optimizer (AsyncIO) that decouples I/O from computation, eliminating waiting bubbles. PreScope achieves 141% higher throughput and 74.6% lower latency than state-of-the-art solutions. Enda Yu, Dezun Dong, Zhaoning Zhang 0001, Zhe Bai, Weiling Yang, Haojie Wang 0004, Dongsheng Li 0001, Yongwei Wu 0001, Xiangke Liao |
ICS | 1 |
| 2025 | DSL-SGD: Distributed Local Stochastic Gradient Descent with Delayed Synchronization
Enda Yu, Zhe Bai, Dezun Dong |
APPT | 1 |
| 2025 | LLMEmu: Execution-Driven Emulator for High-Fidelity Distributed LLM TrainingabstractTransformer-based large models, with trillions of parameters and massive datasets, have driven breakthroughs in NLP, vision, and multimodal tasks. However, their rapid growth poses substantial challenges for training within limited GPU resources, making distributed training indispensable. Pingjing Lu, Enda Yu, Dezun Dong |
IWQoS | 3 |
| 2025 | LLMEmu: A lightweight performance emulator for high-fidelity distributed LLM training
Enda Yu, Pingjing Lu, Dezun Dong |
Perform. Evaluation | 2 |
| 2024 | Enhancing Gradient Compression for Distributed Deep Learning
Zhe Bai, Enda Yu, Dezun Dong, Pingjing Lu |
APNet | 2 |
| 2023 | In-network aggregation for data center networks: A survey
Aoxiang Feng, Dezun Dong, Enda Yu |
Comput. Commun. | 5 |
| 2023 | SSD-SGD: Communication Sparsification for Distributed Deep Learning TrainingabstractIntensive communication and synchronization cost for gradients and parameters is the well-known bottleneck of distributed deep learning training. Based on the observations that Synchronous SGD (SSGD) obtains good convergence accuracy while asynchronous SGD (ASGD) delivers a faster raw training speed, we propose Several Steps Delay SGD (SSD-SGD) to combine their merits, aiming at tackling the communication bottleneck via communication sparsification. SSD-SGD explores both global synchronous updates in the parameter servers and asynchronous local updates in the workers in each periodic iteration. The periodic and flexible synchronization makes SSD-SGD achieve good convergence accuracy and fast training speed. To the best of our knowledge, we strike the new balance between synchronization quality and communication sparsification, and improve the tradeoff between accuracy and training speed. Specifically, the core components of SSD-SGD include proper warm-up stage, steps delay stage, and the novel algorithm of global gradient for local update (GLU). GLU is critical for local update operations by using global gradient information to effectively compensate for the delayed local weights. Furthermore, we implement SSD-SGD on MXNet framework and comprehensively evaluate its performance with CIFAR-10 and ImageNet datasets. Experimental results show that SSD-SGD can accelerate distributed training speed under different experimental configurations, by up to 110% (or 2.1× of the original speed), while achieving good convergence accuracy. Yemao Xu, Dezun Dong, Dongsheng Wang 0004, Enda Yu, Weixia Xu 0001, Xiangke Liao |
ACM Trans. Archit. Code Optim. | 5 |
| 2023 | Communication Optimization Algorithms for Distributed Deep Learning Systems: A SurveyabstractDeep learning's widespread adoption in various fields has made distributed training across multiple computing nodes essential. However, frequent communication between nodes can significantly slow down training speed, creating a bottleneck in distributed training. To address this issue, researchers are focusing on communication optimization algorithms for distributed deep learning systems. In this paper, we propose a standard that systematically classifies all communication optimization algorithms based on mathematical modeling, which is not achieved by existing surveys in the field. We categorize existing works into four categories based on the optimization strategies of communication: communication masking, communication compression, communication frequency reduction, and hybrid optimization. Finally, we discuss potential future challenges and research directions in the field of communication optimization algorithms for distributed deep learning systems. Enda Yu, Dezun Dong, Xiangke Liao |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | DNNEmu: A Lightweight Performance Emulator for Distributed DNN Training
Enda Yu, Dezun Dong, Zhengbin Pang |
ICA3PP | 2 |
| 2022 | CP-SGD: Distributed stochastic gradient descent with compression and periodic compensation
Enda Yu, Dezun Dong, Yemao Xu, Shuo Ouyang, Xiangke Liao |
J. Parallel Distributed Comput. | 1 |
| 2021 | CD-SGD: Distributed Stochastic Gradient Descent with Compression and Delay CompensationabstractCommunication overhead is the key challenge for distributed training. Gradient compression is a widely used approach to reduce communication traffic. When combining with a parallel communication mechanism method like pipeline, gradient compression technique can greatly alleviate the impact of communication overhead. However, there exist two problems of gradient compression technique to be solved. Firstly, gradient compression brings in extra computation cost, which will delay the next training iteration. Secondly, gradient compression usually leads to a decrease in convergence accuracy. In this paper, we combine parallel mechanism with gradient quantization and delayed full-gradient compensation, and propose a new distributed optimization method named CD-SGD, which can hide the overhead of gradient compression, overlap part of the communication and obtain high convergence accuracy. The local update operation in CD-SGD allows the next iteration to be launched quickly without waiting for the completion of gradient compression and the current communication process. Besides, the accuracy loss caused by gradient compression is solved by k-step correction method introduced in CD-SGD. We prove that CD-SGD has convergence guarantee and it achieves at least convergence rate. We conduct extensive experiments on MXNet to verify the convergence properties and scaling performance of CD-SGD. Experimental results on a 16-GPU cluster show that convergence accuracy of CD-SGD is close to or even slightly better than that of S-SGD, and its end-to-end time is 30 less than 2-bit gradient compression under a 56Gbps bandwidth environment. Enda Yu, Dezun Dong, Yemao Xu, Shuo Ouyang, Xiangke Liao |
ICPP | 1 |
| 2021 | vSketchDLC: A Sketch on Distributed Deep Learning Communication via Fine-grained Tracing Visualization
Yanghai Wang, Shuo Ouyang, Dezun Dong, Enda Yu, Xiangke Liao |
NPC | 4 |