Guicheng Qi

dblp:418/3188 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2026
0009-0005-0606-5109ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 44% Parallel and multicore computing · 44% Hardware accelerators and domain-specific architectures · 13%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › parallel programming models
automatic parallelization
1.012026
HetAuto: Cross-Cluster Auto-Parallelism for Heterogeneous Distributed Training · EuroSys 2026
Distributed systems › distributed machine learning
distributed training
1.012026
HetAuto: Cross-Cluster Auto-Parallelism for Heterogeneous Distributed Training · EuroSys 2026
Hardware accelerators and domain-specific architectures › accelerator architecture
heterogeneous accelerator
0.312026
HetAuto: Cross-Cluster Auto-Parallelism for Heterogeneous Distributed Training · EuroSys 2026

Methods — techniques the papers use, named apart from their topics

random forest · 1.0monte carlo tree search · 1.0cost model · 1.0communication overlap · 1.0
YearPublicationVenuePosition
2026 HetAuto: Cross-Cluster Auto-Parallelism for Heterogeneous Distributed Training
abstract
As large neural network models (e.g., LLMs) grow in scale, single-cluster resources become insufficient, making cross-cluster distributed training essential. Cross-cluster training is challenging: hardware heterogeneity complicates load balancing and parallelization strategy and introduces hardware compatibility issues in implementation; cross-cluster communication bottlenecks severely impact training throughput. We present HetAuto, an automatic parallelization system for efficient cross-cluster heterogeneous large model training. HetAuto contributes three key innovations: (1) a principle-guided MCTS algorithm with a random forest-enhanced cost model that efficiently searches parallelization strategies and quickly evaluates their performance under heterogeneous configurations; (2) cross-cluster communication optimizations including Virtual-1F1B scheduling that overlaps communication with computation and an optimized resharding strategy for inter-stage communication; and (3) a unified API enabling seamless integration of diverse accelerators. We evaluate HetAuto across 4 different clusters with up to 736 heterogeneous devices. The evaluation results show that HetAuto achieves up to 1.57× training throughput improvement over representative baselines, and strikes an efficient balance between solution quality and search overhead.
Guicheng Qi, Junwei Su, Liqi Yang, Tingwen Xie, Yerui Sun, Chuan Wu 0001
EuroSys1
2025 ECCheck: Enhancing In-Memory Checkpoint with Erasure Coding in Distributed DNN Training
abstract
Distributed large model training is intensively time and resource consuming. Failures during the long training period are often inevitable, and can incur substantial recovery costs. Checkpointing has been the standard fault tolerance approach, which periodically stores the latest model states at remote persistent storage. This process can be time-consuming due to limited network bandwidth, and adversely affects training throughput. In-memory checkpointing addresses this issue by saving checkpoint data into host memory instead of remote storage. However, host memory is non-persistent, and may not provide sufficient resilience in case of machine failure. We propose ECCheck, a novel in-memory checkpoint system that employs erasure coding to enhance fault tolerance in distributed deep neural network training. ECCheck advocates serialization-free encoding and decoding in model checkpointing. Several techniques are proposed to minimize computation and communication overhead incurred by erasure coding. Extensive experiments demonstrate that ECCheck achieves superior fault tolerance compared to state-of-the-art solutions, while maintaining high checkpointing frequency, low checkpointing stalls, and fast recovery from failures.
Guicheng Qi, Zongpeng Li, Chuan Wu 0001, Zhuwei Peng, Yi Zheng 0007
ICDCS1