Can Balioglu

dblp:266/5860 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
1since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Efficient and distributed learning · 87% Optimization for machine learning · 13%
Software engineering, system software, and programming languages
1 paper
Operating systems · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
1.122023
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel · Proc. VLDB Endow. 2023
Elastic Machine Learning Algorithms in Amazon SageMaker · SIGMOD Conference 2020
Machine learning › Efficient and distributed learning › distributed training › data parallel training
fully sharded data parallel
0.712023
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel · Proc. VLDB Endow. 2023
Machine learning › Efficient and distributed learning › distributed training
large model training
0.712023
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel · Proc. VLDB Endow. 2023
Machine learning › Efficient and distributed learning › distributed training › distributed training systems
elastic training
0.412020
Elastic Machine Learning Algorithms in Amazon SageMaker · SIGMOD Conference 2020
Machine learning › Optimization for machine learning
hyperparameter optimization
0.412020
Elastic Machine Learning Algorithms in Amazon SageMaker · SIGMOD Conference 2020
Operating systems › resource management
memory management
0.212023
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel · Proc. VLDB Endow. 2023

Methods — techniques the papers use, named apart from their topics

sharding · 1.3resumable training · 0.9incremental training · 0.9data-parallel training · 0.7data parallel training · 0.7
YearPublicationVenuePosition
2023 PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
abstract
It is widely acknowledged that large models have the potential to deliver superior performance across a broad range of domains. Despite the remarkable progress made in the field of machine learning systems research, which has enabled the development and exploration of large models, such abilities remain confined to a small group of advanced users and industry leaders, resulting in an implicit technical barrier for the wider community to access and leverage these technologies. In this paper, we introduce PyTorch Fully Sharded Data Parallel (FSDP) as an industry-grade solution for large model training. FSDP has been closely co-designed with several key PyTorch core components including Tensor implementation, dispatcher system, and CUDA memory caching allocator, to provide non-intrusive user experiences and high training efficiency. Additionally, FSDP natively incorporates a range of techniques and settings to optimize resource utilization across a variety of hardware configurations. The experimental results demonstrate that FSDP is capable of achieving comparable performance to Distributed Data Parallel while providing support for significantly larger models with near-linear scalability in terms of TFLOPS.
Yanli Zhao, Andrew Gu, Rohan Varma, Chien-Chin Huang, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, Alban Desmaison, Can Balioglu, Pritam Damania, Bernard Nguyen, Geeta Chauhan, Yuchen Hao, Ajit Mathews
Proc. VLDB Endow.12
2020 Elastic Machine Learning Algorithms in Amazon SageMaker
abstract
There is a large body of research on scalable machine learning (ML). Nevertheless, training ML models on large, continuously evolving datasets is still a difficult and costly undertaking for many companies and institutions. We discuss such challenges and derive requirements for an industrial-scale ML platform. Next, we describe the computational model behind Amazon SageMaker, which is designed to meet such challenges. SageMaker is an ML platform provided as part of Amazon Web Services (AWS), and supports incremental training, resumable and elastic learning as well as automatic hyperparameter optimization. We detail how to adapt several popular ML algorithms to its computational model. Finally, we present an experimental evaluation on large datasets, comparing SageMaker to several scalable, JVM-based implementations of ML algorithms, which we significantly outperform with regard to computation time and cost.
Edo Liberty, Zohar S. Karnin, Bing Xiang, Laurence Rouesnel, Baris Coskun, Ramesh Nallapati, Julio Delgado, Amir Sadoughi, Yury Astashonok, Piali Das, Can Balioglu, Saswata Chakravarty, Madhav Jha, Philip Gautier, David Arpin, Tim Januschowski, Valentin Flunkert, Yuyang Wang 0001, Jan Gasthaus, Lorenzo Stella, Syama Sundar Rangapuram, David Salinas, Sebastian Schelter, Alexander J. Smola
SIGMOD Conference11