Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Chris Cai

dblp:359/6979 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
0009-0002-3416-8939ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Efficient and distributed learning · 87% Language models and text generation · 13%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 50% Parallel and multicore computing · 50%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
1.722025
WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model Training · OSDI 2025
Scaling Llama 3 Training with Efficient Parallelism Strategies · ISCA 2025
High-performance computing
large-scale training
0.912025
Scaling Llama 3 Training with Efficient Parallelism Strategies · ISCA 2025
Parallel and multicore computing › parallel computing › parallel machine learning
parallel training
0.912025
WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model Training · OSDI 2025
Natural language and speech › Language models and text generation › large language model training › language model pretraining
large language model pretraining
0.312025
Scaling Llama 3 Training with Efficient Parallelism Strategies · ISCA 2025

Methods — techniques the papers use, named apart from their topics

tensor parallelism · 1.7pipeline parallelism · 1.7fully sharded data parallelism · 1.7context parallelism · 1.74d parallelism · 1.7
YearPublicationVenuePosition
2025 Scaling Llama 3 Training with Efficient Parallelism Strategies
abstract
Llama is a widely used open-source large language model.This paper presents the design and implementation of the parallelism techniques used in Llama 3 pre-training.To achieve efficient training on tens of thousands of GPUs, Llama 3 employs a combination of four-dimensional parallelism: fully sharded data parallelism, tensor parallelism, pipeline parallelism, and context parallelism.Beyond achieving efficiency through parallelism and model co-design, we
Weiwei Chu, Xinfeng Xie, Jiecao Yu, Jie Wang 0022, Amar Phanishayee, Chunqiang Tang, Yuchen Hao, Muhammet Mustafa Ozdal, Vedanuj Goswami, Naman Goyal 0001, Abhishek Kadian, Andrew Gu, Chris Cai, Xiaodong Wang 0020, Min Si, Pavan Balaji, Ching-Hsiang Chu, Jongsoo Park
ISCA15
2025 WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model Training
Zheng Wang 0075, Anna Cai, Xinfeng Xie, Zaifeng Pan, Yue Guan 0003, Weiwei Chu, Jie Wang 0022, Shikai Li, Chris Cai, Yuchen Hao, Yufei Ding 0001
OSDI10
2023 Towards GPU Memory Efficiency for Distributed Training at Scale
abstract
The scale of deep learning models has grown tremendously in recent years. State-of-the-art models have reached billions of parameters and terabyte-scale model sizes. Training of these models demands memory bandwidth and capacity that can only be accommodated distributively over hundreds to thousands of GPUs. However, large-scale distributed training suffers from GPU memory inefficiency, such as memory under-utilization and out-of-memory events (OOMs). There is a lack of understanding of actual GPU memory behavior of distributed training on terabyte-size models, which hinders the development of effective solutions to such inefficiency. In this paper, we present a systematic analysis of GPU memory behavior of large-scale distributed training jobs in production at Meta. Our analysis is based on offline training jobs of multi-terabyte Deep Learning Recommendation Models from one of Meta's largest production clusters. We measure GPU memory inefficiency; characterize GPU memory utilization, and provide fine-grained GPU memory usage analysis. We further show how to build on the understanding to develop a practical GPU provisioning system in production.
Runxiang Cheng, Chris Cai, Selman Yilmaz, Malay Bag, Mrinmoy Ghosh, Tianyin Xu
SoCC2