Zuquan Song

dblp:379/6912 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
0009-0008-7576-6162ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 3 since 2021Computer networks · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Distributed systems · 68% High-performance computing · 32%
Artificial intelligence
4 papers
Efficient and distributed learning · 73% Language models and text generation · 14% Trustworthy machine learning · 14%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
fault tolerance
2.632025
Robust LLM Training Infrastructure at ByteDance · SOSP 2025
Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training · SOSP 2025
Minder: Faulty Machine Detection for Large-scale Distributed Model Training · NSDI 2025
Machine learning › Efficient and distributed learning
distributed training
1.122025
Understanding Stragglers in Large Model Training Using What-if Analysis · OSDI 2025
Robust LLM Training Infrastructure at ByteDance · SOSP 2025
Distributed systems › fault tolerance
checkpointing
0.912025
ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development · NSDI 2025
High-performance computing
collective communication
0.912025
Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training · SOSP 2025
Distributed systems › distributed machine learning
distributed training
0.912025
Minder: Faulty Machine Detection for Large-scale Distributed Model Training · NSDI 2025
Distributed systems › fault tolerance
failure recovery
0.912025
Robust LLM Training Infrastructure at ByteDance · SOSP 2025
High-performance computing › large-scale training
large language model training
0.912025
Robust LLM Training Infrastructure at ByteDance · SOSP 2025
Distributed systems
root cause analysis
0.912025
Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training · SOSP 2025
Machine learning › Efficient and distributed learning › distributed training
distributed training systems
0.312025
ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development · NSDI 2025
Machine learning › Trustworthy machine learning › robustness
fault tolerance
0.312025
Robust LLM Training Infrastructure at ByteDance · SOSP 2025
Natural language and speech › Language models and text generation
large language model training
0.312025
Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training · SOSP 2025
High-performance computing
large-scale training
0.312025
Understanding Stragglers in Large Model Training Using What-if Analysis · OSDI 2025

Methods — techniques the papers use, named apart from their topics

what-if analysis · 1.7tracing · 1.7fault diagnosis · 1.7failure tolerance · 1.7data-driven localization · 1.7
YearPublicationVenuePosition
2025 Minder: Faulty Machine Detection for Large-scale Distributed Model Training
Yangtao Deng, Zhuo Jiang, Xingjian Zhang 0009, Zhang Zhang 0003, Zuquan Song, Gaohong Liu, Fuliang Li, Shuguang Wang, Haibin Lin, Jianxi Ye, Minlan Yu
NSDI8
2025 ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development
Borui Wan, Mingji Han, Yiyao Sheng, Yanghua Peng, Haibin Lin, Mofan Zhang, Zhichao Lai, Menghan Yu, Junda Zhang, Zuquan Song, Xin Liu 0086, Chuan Wu 0001
NSDI10
2025 Understanding Stragglers in Large Model Training Using What-if Analysis
Jinkun Lin, Ziheng Jiang, Zuquan Song, Sida Zhao, Menghan Yu, Zhanghan Wang, Zuocheng Shi, Zherui Liu, Shuguang Wang, Haibin Lin, Xin Liu 0086, Aurojit Panda, Jinyang Li 0001
OSDI3
2025 Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training
abstract
Reliability is essential for ensuring efficiency in LLM training. However, many real-world reliability issues remain difficult to resolve, resulting in wasted resources and degraded model performance. Unfortunately, today's collective communication libraries operate as black boxes, hiding critical information needed for effective root cause analysis.
Yangtao Deng, Qinlong Wang, Xiaoyun Zhi, Zhuo Jiang, Haohan Xu, Zuquan Song, Gaohong Liu, Shuguang Wang, Wencong Xiao, Jianxi Ye, Minlan Yu, Hong Xu 0001
SOSP9
2025 Robust LLM Training Infrastructure at ByteDance
abstract
The training scale of large language models (LLMs) has reached tens of thousands of GPUs and is still continuously expanding, enabling faster learning of larger models. Accompanying the expansion of the resource scale is the prevalence of failures (CUDA error, NaN values, job hang, etc.), which poses significant challenges to training stability. Any large-scale LLM training infrastructure should strive for minimal training interruption, efficient fault diagnosis, and effective failure tolerance to enable highly efficient continuous training. This paper presents ByteRobust, a large-scale GPU infrastructure management system tailored for robust and stable training of LLMs. It exploits the uniqueness of LLM training process and gives top priorities to detecting and recovering failures in a routine manner. Leveraging parallelisms and characteristics of LLM training, ByteRobust enables high-capacity fault tolerance, prompt fault demarcation, and localization with an effective data-driven approach, comprehensively ensuring continuous and efficient training of LLM tasks. ByteRobust is deployed on a production GPU platform with over 200,000 GPUs and advances the state of the art in training robustness by achieving 97% ETTR for a three-month training job on 9,600 GPUs.
Borui Wan, Gaohong Liu, Zuquan Song, Jun Wang 0039, Guangming Sheng, Shuguang Wang, Houmin Wei, Weiqiang Lou, Mofan Zhang, Kaihua Jiang, Cheng Ren, Xiaoyun Zhi, Menghan Yu, Zhe Nan, Zhuolin Zheng, Baoquan Zhong, Qinlong Wang, Jinxin Chi, Wang Zhang 0017, Zixian Du, Sida Zhao, Jingzhe Tang, Zherui Liu, Chuan Wu 0001, Yanghua Peng, Haibin Lin, Wencong Xiao, Xin Liu 0086
SOSP3