Andrew Gu

dblp:287/4942 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0001-1236-9484ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 93% Language models and text generation · 7%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
High-performance computing · 60% Distributed systems · 30% Cloud and datacenter computing · 10%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
2.432025
Scaling Llama 3 Training with Efficient Parallelism Strategies · ISCA 2025
TorchTitan: One-stop PyTorch native solution for production ready LLM pretraining · ICLR 2025
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel · Proc. VLDB Endow. 2023
High-performance computing
large-scale training
1.722025
Scaling Llama 3 Training with Efficient Parallelism Strategies · ISCA 2025
TorchTitan: One-stop PyTorch native solution for production ready LLM pretraining · ICLR 2025
Recommender systems › advertising
advertising recommendation
1.012026
Meta Lattice: Model Space Redesign for Cost-Effective Industry-Scale Ads Recommendations · KDD (1) 2026
Distributed systems › distributed machine learning
distributed training systems
0.912025
TorchTitan: One-stop PyTorch native solution for production ready LLM pretraining · ICLR 2025
Machine learning › Efficient and distributed learning › distributed training › data parallel training
fully sharded data parallel
0.712023
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel · Proc. VLDB Endow. 2023
Machine learning › Efficient and distributed learning › distributed training
large model training
0.712023
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel · Proc. VLDB Endow. 2023
Cloud and datacenter computing
cluster resource management and scheduling
0.312026
Meta Lattice: Model Space Redesign for Cost-Effective Industry-Scale Ads Recommendations · KDD (1) 2026
Natural language and speech › Language models and text generation › large language model training › language model pretraining
large language model pretraining
0.312025
Scaling Llama 3 Training with Efficient Parallelism Strategies · ISCA 2025
Operating systems › resource management
memory management
0.212023
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel · Proc. VLDB Endow. 2023

Methods — techniques the papers use, named apart from their topics

tensor parallelism · 1.7pipeline parallelism · 1.7hardware-software co-design · 1.7fully sharded data parallelism · 1.7float8 training · 1.7elastic scaling · 1.7context parallelism · 1.7sharding · 1.3data-parallel training · 0.7data parallel training · 0.7
YearPublicationVenuePosition
2026 Meta Lattice: Model Space Redesign for Cost-Effective Industry-Scale Ads Recommendations
abstract
The rapidly evolving landscape of products, surfaces, policies, and regulations poses significant challenges for deploying state-of-the-art recommendation models at industry scale, primarily due to data fragmentation across domains and escalating infrastructure costs that hinder sustained quality improvements.
Yuxin Chen 0001, Mengyue Hang, Andrew Gu, Buyun Zhang, Fan Yang 0094, Feifan Gu, Jade Nie, Jiayi Xu 0001, Jiyan Yang, Jongsoo Park, Laming Chen, Longhao Jin, Qin Huang 0006, Shali Jiang 0003, Shiwen Shen, Shuaiwen Wang, Siyang Yuan, Tongyi Tang, Weilin Zhang, Xi Liu 0011, Xiaohan Wei, Yuchen Hao, Xiaozhen Xia, Yasmine Badr, Zeliang Chen, Chengze Fan, Qianru Li 0002, Sihan Zeng, Yinbin Ma, Maxim Naumov, Yantao Yao, Ellie Wen
KDD (1)5
2025 TorchTitan: One-stop PyTorch native solution for production ready LLM pretraining
abstract
The development of large language models (LLMs) has been instrumental in advancing state-of-the-art natural language processing applications. Training LLMs with billions of parameters and trillions of tokens requires sophisticated distributed systems that enable composing and comparing several state-of-the-art techniques in order to efficiently scale across thousands of accelerators. However, existing solutions are complex, scattered across multiple libraries/repositories, lack interoperability, and are cumbersome to maintain. Thus, curating and empirically comparing training recipes requires non-trivial engineering effort. This paper introduces **TORCHTITAN**$^1$, a PyTorch-native distributed training system that unifies and advances state-of-the-art techniques, streamlining integration and reducing engineering overhead. TORCHTITAN enables seamless application of 4D parallelism in a modular and composable manner, while featuring elastic scaling to adapt to changing computational requirements. The system provides comprehensive logging, efficient checkpointing, and debugging tools, ensuring production-ready training. Moreover, TORCHTITAN incorporates innovative hardware-software co-designed solutions, leveraging cutting-edge features like Float8 training and SymmetricMemory to maximize hardware utilization. As a flexible experimental test bed, TORCHTITAN facilitates the curation and comparison of custom recipes for diverse training contexts. By leveraging TORCHTITAN, we developed optimized training recipes for the Llama 3.1 family and provide actionable guidance on selecting and combining distributed training techniques to maximize training efficiency, based on our hands-on experiences. We thoroughly assess TORCHTITAN on the Llama 3.1 family of LLMs, spanning 8 billion to 405 billion parameters, and showcase its exceptional performance, modular composability, and elastic scalability. By stacking training optimizations, we demonstrate accelerations ranging from 65.08% on Llama 3.1 8B at 128 GPU scale (1D), 12.59% on Llama 3.1 70B at 256 GPU scale (2D), to 30% on Llama 3.1 405B at 512 GPU scale (3D) on NVIDIA H100 GPUs over optimized baselines. We also demonstrate the effectiveness of 4D parallelism in enabling long context training. $^1$ GitHub: [https://github.com/pytorch/torchtitan](https://github.com/pytorch/torchtitan)
Wanchao Liang, Less Wright, Will Constable, Andrew Gu, Chien-Chin Huang, Iris Zhang, Howard Huang, Sanket Purandare, Gokul Nadathur, Stratos Idreos
ICLR5
2025 Scaling Llama 3 Training with Efficient Parallelism Strategies
abstract
Llama is a widely used open-source large language model.This paper presents the design and implementation of the parallelism techniques used in Llama 3 pre-training.To achieve efficient training on tens of thousands of GPUs, Llama 3 employs a combination of four-dimensional parallelism: fully sharded data parallelism, tensor parallelism, pipeline parallelism, and context parallelism.Beyond achieving efficiency through parallelism and model co-design, we
Weiwei Chu, Xinfeng Xie, Jiecao Yu, Jie Wang 0022, Amar Phanishayee, Chunqiang Tang, Yuchen Hao, Muhammet Mustafa Ozdal, Vedanuj Goswami, Naman Goyal 0001, Abhishek Kadian, Andrew Gu, Chris Cai, Xiaodong Wang 0020, Min Si, Pavan Balaji, Ching-Hsiang Chu, Jongsoo Park
ISCA14
2023 PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
abstract
It is widely acknowledged that large models have the potential to deliver superior performance across a broad range of domains. Despite the remarkable progress made in the field of machine learning systems research, which has enabled the development and exploration of large models, such abilities remain confined to a small group of advanced users and industry leaders, resulting in an implicit technical barrier for the wider community to access and leverage these technologies. In this paper, we introduce PyTorch Fully Sharded Data Parallel (FSDP) as an industry-grade solution for large model training. FSDP has been closely co-designed with several key PyTorch core components including Tensor implementation, dispatcher system, and CUDA memory caching allocator, to provide non-intrusive user experiences and high training efficiency. Additionally, FSDP natively incorporates a range of techniques and settings to optimize resource utilization across a variety of hardware configurations. The experimental results demonstrate that FSDP is capable of achieving comparable performance to Distributed Data Parallel while providing support for significantly larger models with near-linear scalability in terms of TFLOPS.
Yanli Zhao, Andrew Gu, Rohan Varma, Chien-Chin Huang, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, Alban Desmaison, Can Balioglu, Pritam Damania, Bernard Nguyen, Geeta Chauhan, Yuchen Hao, Ajit Mathews
Proc. VLDB Endow.2