EDBT 2026 Demo / reviewers in the wild / expert
Yerui Sun
dblp:355/5601
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0003-5452-001XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Distributed systems · 44% Parallel and multicore computing · 44% Hardware accelerators and domain-specific architectures · 13% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing › parallel programming models
automatic parallelization |
1.0 | 1 | 2026 | HetAuto: Cross-Cluster Auto-Parallelism for Heterogeneous Distributed Training · EuroSys 2026 |
Distributed systems › distributed machine learning
distributed training |
1.0 | 1 | 2026 | HetAuto: Cross-Cluster Auto-Parallelism for Heterogeneous Distributed Training · EuroSys 2026 |
Hardware accelerators and domain-specific architectures › accelerator architecture
heterogeneous accelerator |
0.3 | 1 | 2026 | HetAuto: Cross-Cluster Auto-Parallelism for Heterogeneous Distributed Training · EuroSys 2026 |
Methods — techniques the papers use, named apart from their topics
random forest · 1.0monte carlo tree search · 1.0cost model · 1.0communication overlap · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HetAuto: Cross-Cluster Auto-Parallelism for Heterogeneous Distributed TrainingabstractAs large neural network models (e.g., LLMs) grow in scale, single-cluster resources become insufficient, making cross-cluster distributed training essential. Cross-cluster training is challenging: hardware heterogeneity complicates load balancing and parallelization strategy and introduces hardware compatibility issues in implementation; cross-cluster communication bottlenecks severely impact training throughput. We present HetAuto, an automatic parallelization system for efficient cross-cluster heterogeneous large model training. HetAuto contributes three key innovations: (1) a principle-guided MCTS algorithm with a random forest-enhanced cost model that efficiently searches parallelization strategies and quickly evaluates their performance under heterogeneous configurations; (2) cross-cluster communication optimizations including Virtual-1F1B scheduling that overlaps communication with computation and an optimized resharding strategy for inter-stage communication; and (3) a unified API enabling seamless integration of diverse accelerators. We evaluate HetAuto across 4 different clusters with up to 736 heterogeneous devices. The evaluation results show that HetAuto achieves up to 1.57× training throughput improvement over representative baselines, and strikes an efficient balance between solution quality and search overhead. Guicheng Qi, Junwei Su, Liqi Yang, Tingwen Xie, Yerui Sun, Chuan Wu 0001 |
EuroSys | 6 |