EDBT 2026 Demo / reviewers in the wild / expert
Gaohong Liu
dblp:327/4574
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2025
0009-0005-2551-7879ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Distributed systems · 67% High-performance computing · 33% | |
| Artificial intelligence
3 papers |
Reinforcement learning · 69% Language models and text generation · 10% Efficient and distributed learning · 10% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems
fault tolerance |
2.6 | 3 | 2025 | Robust LLM Training Infrastructure at ByteDance · SOSP 2025 Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training · SOSP 2025 Minder: Faulty Machine Detection for Large-scale Distributed Model Training · NSDI 2025 |
Machine learning › Reinforcement learning
policy optimization |
0.9 | 1 | 2025 | DAPO: An Open-Source LLM Reinforcement Learning System at Scale · NeurIPS 2025 |
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement learning for language models |
0.9 | 1 | 2025 | DAPO: An Open-Source LLM Reinforcement Learning System at Scale · NeurIPS 2025 |
High-performance computing
collective communication |
0.9 | 1 | 2025 | Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training · SOSP 2025 |
Distributed systems › distributed machine learning
distributed training |
0.9 | 1 | 2025 | Minder: Faulty Machine Detection for Large-scale Distributed Model Training · NSDI 2025 |
Distributed systems › fault tolerance
failure recovery |
0.9 | 1 | 2025 | Robust LLM Training Infrastructure at ByteDance · SOSP 2025 |
High-performance computing › large-scale training
large language model training |
0.9 | 1 | 2025 | Robust LLM Training Infrastructure at ByteDance · SOSP 2025 |
Distributed systems
root cause analysis |
0.9 | 1 | 2025 | Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training · SOSP 2025 |
Machine learning › Efficient and distributed learning
distributed training |
0.3 | 1 | 2025 | Robust LLM Training Infrastructure at ByteDance · SOSP 2025 |
Machine learning › Trustworthy machine learning › robustness
fault tolerance |
0.3 | 1 | 2025 | Robust LLM Training Infrastructure at ByteDance · SOSP 2025 |
Natural language and speech › Language models and text generation
large language model training |
0.3 | 1 | 2025 | Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training · SOSP 2025 |
Methods — techniques the papers use, named apart from their topics
tracing · 1.7fault diagnosis · 1.7failure tolerance · 1.7data-driven localization · 1.7reinforcement learning at scale · 0.9policy optimization · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DAPO: An Open-Source LLM Reinforcement Learning System at ScaleabstractInference scaling empowers LLMs with unprecedented reasoning ability, with reinforcement learning as the core technique to elicit complex reasoning. However, key technical details of state-of-the-art reasoning LLMs are concealed (such as in OpenAI o1 blog and DeepSeek R1 technical report), thus the community still struggles to reproduce their RL training results. We propose the **D**ecoupled Clip and **D**ynamic s**A**mpling **P**olicy **O**ptimization (**DAPO**) algorithm, and fully open-source a state-of-the-art large-scale RL system that achieves 50 points on AIME 2024 using Qwen2.5-32B base model. Unlike previous works that withhold training details, we introduce four key techniques of our algorithm that make large-scale LLM RL a success. In addition, we open-source our training code, which is built on the verl framework, along with a carefully curated and processed dataset. These components of our open-source system enhance reproducibility and support future research in large-scale LLM RL. Qiying Yu, Zheng Zhang 0001, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Juncai Liu, Lingjun Liu, Xin Liu 0039, Haibin Lin, Bole Ma, Guangming Sheng, Yuxuan Tong, Chi Zhang 0022, Mofan Zhang, Ru Zhang 0006, Wang Zhang 0017, Jiaze Chen, Jiangjie Chen, Hongli Yu, Yuxuan Song 0002, Xiangpeng Wei, Hao Zhou 0012, Wei-Ying Ma, Ya-Qin Zhang, Mingxuan Wang |
NeurIPS | 9 |
| 2025 | Minder: Faulty Machine Detection for Large-scale Distributed Model Training
Yangtao Deng, Zhuo Jiang, Xingjian Zhang 0009, Zhang Zhang 0003, Zuquan Song, Gaohong Liu, Fuliang Li, Shuguang Wang, Haibin Lin, Jianxi Ye, Minlan Yu |
NSDI | 10 |
| 2025 | Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM TrainingabstractReliability is essential for ensuring efficiency in LLM training. However, many real-world reliability issues remain difficult to resolve, resulting in wasted resources and degraded model performance. Unfortunately, today's collective communication libraries operate as black boxes, hiding critical information needed for effective root cause analysis. Yangtao Deng, Qinlong Wang, Xiaoyun Zhi, Zhuo Jiang, Haohan Xu, Zuquan Song, Gaohong Liu, Shuguang Wang, Wencong Xiao, Jianxi Ye, Minlan Yu, Hong Xu 0001 |
SOSP | 10 |
| 2025 | Robust LLM Training Infrastructure at ByteDanceabstractThe training scale of large language models (LLMs) has reached tens of thousands of GPUs and is still continuously expanding, enabling faster learning of larger models. Accompanying the expansion of the resource scale is the prevalence of failures (CUDA error, NaN values, job hang, etc.), which poses significant challenges to training stability. Any large-scale LLM training infrastructure should strive for minimal training interruption, efficient fault diagnosis, and effective failure tolerance to enable highly efficient continuous training. This paper presents ByteRobust, a large-scale GPU infrastructure management system tailored for robust and stable training of LLMs. It exploits the uniqueness of LLM training process and gives top priorities to detecting and recovering failures in a routine manner. Leveraging parallelisms and characteristics of LLM training, ByteRobust enables high-capacity fault tolerance, prompt fault demarcation, and localization with an effective data-driven approach, comprehensively ensuring continuous and efficient training of LLM tasks. ByteRobust is deployed on a production GPU platform with over 200,000 GPUs and advances the state of the art in training robustness by achieving 97% ETTR for a three-month training job on 9,600 GPUs. Borui Wan, Gaohong Liu, Zuquan Song, Jun Wang 0039, Guangming Sheng, Shuguang Wang, Houmin Wei, Weiqiang Lou, Mofan Zhang, Kaihua Jiang, Cheng Ren, Xiaoyun Zhi, Menghan Yu, Zhe Nan, Zhuolin Zheng, Baoquan Zhong, Qinlong Wang, Jinxin Chi, Wang Zhang 0017, Zixian Du, Sida Zhao, Jingzhe Tang, Zherui Liu, Chuan Wu 0001, Yanghua Peng, Haibin Lin, Wencong Xiao, Xin Liu 0086 |
SOSP | 2 |