EDBT 2026 Demo / reviewers in the wild / expert
Chris Cai
dblp:359/6979
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
0009-0002-3416-8939ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Efficient and distributed learning · 87% Language models and text generation · 13% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
High-performance computing · 50% Parallel and multicore computing · 50% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
distributed training |
1.7 | 2 | 2025 | WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model Training · OSDI 2025 Scaling Llama 3 Training with Efficient Parallelism Strategies · ISCA 2025 |
High-performance computing
large-scale training |
0.9 | 1 | 2025 | Scaling Llama 3 Training with Efficient Parallelism Strategies · ISCA 2025 |
Parallel and multicore computing › parallel computing › parallel machine learning
parallel training |
0.9 | 1 | 2025 | WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model Training · OSDI 2025 |
Natural language and speech › Language models and text generation › large language model training › language model pretraining
large language model pretraining |
0.3 | 1 | 2025 | Scaling Llama 3 Training with Efficient Parallelism Strategies · ISCA 2025 |
Methods — techniques the papers use, named apart from their topics
tensor parallelism · 1.7pipeline parallelism · 1.7fully sharded data parallelism · 1.7context parallelism · 1.74d parallelism · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scaling Llama 3 Training with Efficient Parallelism StrategiesabstractLlama is a widely used open-source large language model.This paper presents the design and implementation of the parallelism techniques used in Llama 3 pre-training.To achieve efficient training on tens of thousands of GPUs, Llama 3 employs a combination of four-dimensional parallelism: fully sharded data parallelism, tensor parallelism, pipeline parallelism, and context parallelism.Beyond achieving efficiency through parallelism and model co-design, we Weiwei Chu, Xinfeng Xie, Jiecao Yu, Jie Wang 0022, Amar Phanishayee, Chunqiang Tang, Yuchen Hao, Muhammet Mustafa Ozdal, Vedanuj Goswami, Naman Goyal 0001, Abhishek Kadian, Andrew Gu, Chris Cai, Xiaodong Wang 0020, Min Si, Pavan Balaji, Ching-Hsiang Chu, Jongsoo Park |
ISCA | 15 |
| 2025 | WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model Training
Zheng Wang 0075, Anna Cai, Xinfeng Xie, Zaifeng Pan, Yue Guan 0003, Weiwei Chu, Jie Wang 0022, Shikai Li, Chris Cai, Yuchen Hao, Yufei Ding 0001 |
OSDI | 10 |
| 2023 | Towards GPU Memory Efficiency for Distributed Training at ScaleabstractThe scale of deep learning models has grown tremendously in recent years. State-of-the-art models have reached billions of parameters and terabyte-scale model sizes. Training of these models demands memory bandwidth and capacity that can only be accommodated distributively over hundreds to thousands of GPUs. However, large-scale distributed training suffers from GPU memory inefficiency, such as memory under-utilization and out-of-memory events (OOMs). There is a lack of understanding of actual GPU memory behavior of distributed training on terabyte-size models, which hinders the development of effective solutions to such inefficiency. In this paper, we present a systematic analysis of GPU memory behavior of large-scale distributed training jobs in production at Meta. Our analysis is based on offline training jobs of multi-terabyte Deep Learning Recommendation Models from one of Meta's largest production clusters. We measure GPU memory inefficiency; characterize GPU memory utilization, and provide fine-grained GPU memory usage analysis. We further show how to build on the understanding to develop a practical GPU provisioning system in production. Runxiang Cheng, Chris Cai, Selman Yilmaz, Malay Bag, Mrinmoy Ghosh, Tianyin Xu |
SoCC | 2 |