VLDB 2026 Research / reviewers in the wild / expert
Andrew Gu
dblp:287/4942
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0001-1236-9484ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Efficient and distributed learning · 93% Language models and text generation · 7% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
High-performance computing · 60% Distributed systems · 30% Cloud and datacenter computing · 10% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
distributed training |
2.4 | 3 | 2025 | Scaling Llama 3 Training with Efficient Parallelism Strategies · ISCA 2025 TorchTitan: One-stop PyTorch native solution for production ready LLM pretraining · ICLR 2025 PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel · Proc. VLDB Endow. 2023 |
High-performance computing
large-scale training |
1.7 | 2 | 2025 | Scaling Llama 3 Training with Efficient Parallelism Strategies · ISCA 2025 TorchTitan: One-stop PyTorch native solution for production ready LLM pretraining · ICLR 2025 |
Recommender systems › advertising
advertising recommendation |
1.0 | 1 | 2026 | Meta Lattice: Model Space Redesign for Cost-Effective Industry-Scale Ads Recommendations · KDD (1) 2026 |
Distributed systems › distributed machine learning
distributed training systems |
0.9 | 1 | 2025 | TorchTitan: One-stop PyTorch native solution for production ready LLM pretraining · ICLR 2025 |
Machine learning › Efficient and distributed learning › distributed training › data parallel training
fully sharded data parallel |
0.7 | 1 | 2023 | PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel · Proc. VLDB Endow. 2023 |
Machine learning › Efficient and distributed learning › distributed training
large model training |
0.7 | 1 | 2023 | PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel · Proc. VLDB Endow. 2023 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.3 | 1 | 2026 | Meta Lattice: Model Space Redesign for Cost-Effective Industry-Scale Ads Recommendations · KDD (1) 2026 |
Natural language and speech › Language models and text generation › large language model training › language model pretraining
large language model pretraining |
0.3 | 1 | 2025 | Scaling Llama 3 Training with Efficient Parallelism Strategies · ISCA 2025 |
Operating systems › resource management
memory management |
0.2 | 1 | 2023 | PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel · Proc. VLDB Endow. 2023 |
Methods — techniques the papers use, named apart from their topics
tensor parallelism · 1.7pipeline parallelism · 1.7hardware-software co-design · 1.7fully sharded data parallelism · 1.7float8 training · 1.7elastic scaling · 1.7context parallelism · 1.7sharding · 1.3data-parallel training · 0.7data parallel training · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Meta Lattice: Model Space Redesign for Cost-Effective Industry-Scale Ads RecommendationsabstractThe rapidly evolving landscape of products, surfaces, policies, and regulations poses significant challenges for deploying state-of-the-art recommendation models at industry scale, primarily due to data fragmentation across domains and escalating infrastructure costs that hinder sustained quality improvements. Yuxin Chen 0001, Mengyue Hang, Andrew Gu, Buyun Zhang, Fan Yang 0094, Feifan Gu, Jade Nie, Jiayi Xu 0001, Jiyan Yang, Jongsoo Park, Laming Chen, Longhao Jin, Qin Huang 0006, Shali Jiang 0003, Shiwen Shen, Shuaiwen Wang, Siyang Yuan, Tongyi Tang, Weilin Zhang, Xi Liu 0011, Xiaohan Wei, Yuchen Hao, Xiaozhen Xia, Yasmine Badr, Zeliang Chen, Chengze Fan, Qianru Li 0002, Sihan Zeng, Yinbin Ma, Maxim Naumov, Yantao Yao, Ellie Wen |
KDD (1) | 5 |
| 2025 | TorchTitan: One-stop PyTorch native solution for production ready LLM pretrainingabstractThe development of large language models (LLMs) has been instrumental in advancing state-of-the-art natural language processing applications. Training LLMs with billions of parameters and trillions of tokens requires sophisticated distributed systems that enable composing and comparing several state-of-the-art techniques in order to efficiently scale across thousands of accelerators. However, existing solutions are complex, scattered across multiple libraries/repositories, lack interoperability, and are cumbersome to maintain. Thus, curating and empirically comparing training recipes requires non-trivial engineering effort.
This paper introduces **TORCHTITAN**$^1$, a PyTorch-native distributed training system that unifies and advances state-of-the-art techniques, streamlining integration and reducing engineering overhead. TORCHTITAN enables seamless application of 4D parallelism in a modular and composable manner, while featuring elastic scaling to adapt to changing computational requirements. The system provides comprehensive logging, efficient checkpointing, and debugging tools, ensuring production-ready training. Moreover, TORCHTITAN incorporates innovative hardware-software co-designed solutions, leveraging cutting-edge features like Float8 training and SymmetricMemory to maximize hardware utilization.
As a flexible experimental test bed, TORCHTITAN facilitates the curation and comparison of custom recipes for diverse training contexts. By leveraging TORCHTITAN, we developed optimized training recipes for the Llama 3.1 family and provide actionable guidance on selecting and combining distributed training techniques to maximize training efficiency, based on our hands-on experiences.
We thoroughly assess TORCHTITAN on the Llama 3.1 family of LLMs, spanning 8 billion to 405 billion parameters, and showcase its exceptional performance, modular composability, and elastic scalability. By stacking training optimizations, we demonstrate accelerations ranging from 65.08% on Llama 3.1 8B at 128 GPU scale (1D), 12.59% on Llama 3.1 70B at 256 GPU scale (2D), to 30% on Llama 3.1 405B at 512 GPU scale (3D) on NVIDIA H100 GPUs over optimized baselines. We also demonstrate the effectiveness of 4D parallelism in enabling long context training.
$^1$ GitHub: [https://github.com/pytorch/torchtitan](https://github.com/pytorch/torchtitan) Wanchao Liang, Less Wright, Will Constable, Andrew Gu, Chien-Chin Huang, Iris Zhang, Howard Huang, Sanket Purandare, Gokul Nadathur, Stratos Idreos |
ICLR | 5 |
| 2025 | Scaling Llama 3 Training with Efficient Parallelism StrategiesabstractLlama is a widely used open-source large language model.This paper presents the design and implementation of the parallelism techniques used in Llama 3 pre-training.To achieve efficient training on tens of thousands of GPUs, Llama 3 employs a combination of four-dimensional parallelism: fully sharded data parallelism, tensor parallelism, pipeline parallelism, and context parallelism.Beyond achieving efficiency through parallelism and model co-design, we Weiwei Chu, Xinfeng Xie, Jiecao Yu, Jie Wang 0022, Amar Phanishayee, Chunqiang Tang, Yuchen Hao, Muhammet Mustafa Ozdal, Vedanuj Goswami, Naman Goyal 0001, Abhishek Kadian, Andrew Gu, Chris Cai, Xiaodong Wang 0020, Min Si, Pavan Balaji, Ching-Hsiang Chu, Jongsoo Park |
ISCA | 14 |
| 2023 | PyTorch FSDP: Experiences on Scaling Fully Sharded Data ParallelabstractIt is widely acknowledged that large models have the potential to deliver superior performance across a broad range of domains. Despite the remarkable progress made in the field of machine learning systems research, which has enabled the development and exploration of large models, such abilities remain confined to a small group of advanced users and industry leaders, resulting in an implicit technical barrier for the wider community to access and leverage these technologies. In this paper, we introduce PyTorch Fully Sharded Data Parallel (FSDP) as an industry-grade solution for large model training. FSDP has been closely co-designed with several key PyTorch core components including Tensor implementation, dispatcher system, and CUDA memory caching allocator, to provide non-intrusive user experiences and high training efficiency. Additionally, FSDP natively incorporates a range of techniques and settings to optimize resource utilization across a variety of hardware configurations. The experimental results demonstrate that FSDP is capable of achieving comparable performance to Distributed Data Parallel while providing support for significantly larger models with near-linear scalability in terms of TFLOPS. Yanli Zhao, Andrew Gu, Rohan Varma, Chien-Chin Huang, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, Alban Desmaison, Can Balioglu, Pritam Damania, Bernard Nguyen, Geeta Chauhan, Yuchen Hao, Ajit Mathews |
Proc. VLDB Endow. | 2 |