Geeta Chauhan

dblp:272/9960 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
2since 2021 · last 2024
0009-0003-0830-7330ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 1 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 71% Trustworthy machine learning · 16% Deep learning architectures and training · 13%
Software engineering, system software, and programming languages
2 papers
Runtime systems and virtual machines · 44% Compilers and program optimization · 44% Operating systems · 12%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%
Network and information security
1 paper
Privacy and data protection · 100%

Topics — the 7 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Runtime systems and virtual machines › dynamic compilation
just-in-time compilation
0.812024
PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation · ASPLOS (2) 2024
Machine learning › Efficient and distributed learning
distributed training
0.712023
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel · Proc. VLDB Endow. 2023
Machine learning › Efficient and distributed learning › distributed training › data parallel training
fully sharded data parallel
0.712023
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel · Proc. VLDB Endow. 2023
Machine learning › Efficient and distributed learning › distributed training
large model training
0.712023
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel · Proc. VLDB Endow. 2023
Machine learning › Trustworthy machine learning
interpretability
0.412020
Building Recommender Systems with PyTorch · KDD 2020
Machine learning › Deep learning architectures and training › deep learning systems
deep learning framework
0.212024
PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation · ASPLOS (2) 2024
Operating systems › resource management
memory management
0.212023
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel · Proc. VLDB Endow. 2023

Methods — techniques the papers use, named apart from their topics

just-in-time compilation · 1.5graph compilation · 1.5sharding · 1.3pytorch · 1.3deep learning · 1.3data-parallel training · 0.7data parallel training · 0.7
YearPublicationVenuePosition
2024 PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation
abstract
This paper introduces two extensions to the popular PyTorch machine learning framework, TorchDynamo and TorchInductor, which implement the torch.compile feature released in PyTorch 2. TorchDynamo is a Python-level just-in-time (JIT) compiler that enables graph compilation in PyTorch programs without sacrificing the flexibility of Python. It achieves this by dynamically modifying Python bytecode before execution and extracting sequences of PyTorch operations into an FX graph, which is then JIT compiled using one of many extensible backends. TorchInductor is the default compiler backend for TorchDynamo, which translates PyTorch programs into OpenAI's Triton for GPUs and C++ for CPUs. Results show that TorchDynamo is able to capture graphs more robustly than prior approaches while adding minimal overhead, and TorchInductor is able to provide a 2.27× inference and 1.41× training geometric mean speedup on an NVIDIA A100 GPU across 180+ real-world models, which outperforms six other compilers. These extensions provide a new way to apply optimizations through compilers in eager mode frameworks like PyTorch.
Jason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein, Animesh Jain, Michael Voznesensky, Bin Bao, Peter Bell 0008, David Berard, Evgeni Burovski, Geeta Chauhan, Anjali Chourdia, Will Constable, Alban Desmaison, Zach DeVito, Elias Ellison, Will Feng, Jiong Gong, Michael Gschwind, Brian Hirsh, Sherlock Huang, Kshiteej Kalambarkar, Laurent Kirsch, Michael Lazos, Mario Lezcano Casado, Yanbo Liang, Jason Liang, Yinghai Lu, C. K. Luk, Bert Maher, Yunjie Pan, Christian Puhrsch, Matthias Reso, Mark Saroufim, Marcos Yukio Siraichi, Helen Suk, Shunting Zhang, Michael Suo, Phil Tillet, Xu Zhao 0004, Eikan Wang, Keren Zhou 0001, Richard Zou, Ajit Mathews, Xiaoquan Wen, Gregory Chanan, Peng Wu 0001, Soumith Chintala
ASPLOS (2)11
2023 PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
abstract
It is widely acknowledged that large models have the potential to deliver superior performance across a broad range of domains. Despite the remarkable progress made in the field of machine learning systems research, which has enabled the development and exploration of large models, such abilities remain confined to a small group of advanced users and industry leaders, resulting in an implicit technical barrier for the wider community to access and leverage these technologies. In this paper, we introduce PyTorch Fully Sharded Data Parallel (FSDP) as an industry-grade solution for large model training. FSDP has been closely co-designed with several key PyTorch core components including Tensor implementation, dispatcher system, and CUDA memory caching allocator, to provide non-intrusive user experiences and high training efficiency. Additionally, FSDP natively incorporates a range of techniques and settings to optimize resource utilization across a variety of hardware configurations. The experimental results demonstrate that FSDP is capable of achieving comparable performance to Distributed Data Parallel while providing support for significantly larger models with near-linear scalability in terms of TFLOPS.
Yanli Zhao, Andrew Gu, Rohan Varma, Chien-Chin Huang, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, Alban Desmaison, Can Balioglu, Pritam Damania, Bernard Nguyen, Geeta Chauhan, Yuchen Hao, Ajit Mathews
Proc. VLDB Endow.15
2020 Building Recommender Systems with PyTorch
abstract
In this tutorial we show how to build deep learning recommendation systems and resolve the associated interpretability, integrity and privacy challenges. We start with an overview of the PyTorch framework, features that it offers and a brief review of the evolution of recommendation models. We delineate their typical components and build a proxy deep learning recommendation model (DLRM) in PyTorch. Then, we discuss how to interpret recommendation system results as well as how to address the corresponding integrity and quality challenges.
Dheevatsa Mudigere, Maxim Naumov, Joe Spisak, Geeta Chauhan, Narine Kokhlikyan, Amanpreet Singh, Vedanuj Goswami
KDD4