Ryan Chard

dblp:124/2122 · DBLP profile ↗
← Back
2ranked-venue papers in the field
0as first author
2since 2021 · last 2024
0000-0002-6781-7432ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 2
YearPublicationVenuePosition
2024 Model and Data Management for Machine Learning (M2ML): Integrating Instruments, Edge and HPC for Accelerated Machine Learning
abstract
The use of data produced by scientific instruments, such as the Advanced Photon Source Upgrade (APS-U), to train and fine-tune machine learning models is becoming increasingly challenging due to high data production rates, large data volumes, and the growing complexity of machine learning models. To address these challenges, researchers have developed frameworks like fairDMS to efficiently organize vast amounts of data and models for rapid querying when model degradation is detected. However, the complexity of these frameworks and the physically distributed nature of experimental facilities complicate their deployment.Here we introduce a high-performance model and data management framework for machine learning, M2ML. In contrast to previous frameworks, M2ML abstracts the tasks into three key elements that can be easily called and accessed by users. M2ML is capable of utilizing a variety of computational resources, that are distributed across scientific facilities, to accelerate machine learning tasks. For example, it can automatically transfer data from an experimental facility (such as APS-U) to a high performance computing (HPC) facility (such as the Argonne Leadership Computing Facility (ALCF)), train machine learning models at the HPC facility, and deploy the trained models on edge computing devices back at the experimental facility for inferencing. M2ML provides a unified interface for (on-the-fly) model (re)training, storage, evaluation, fine-tuning, and inferencing using heterogeneous resources that can be geographically distributed. M2ML uses Globus services such as Globus Transfer and Globus Compute (formerly FuncX). We evaluate M2ML using a high energy diffraction microscopy (HEDM) workflow that employs BraggNN to predict the diffraction peak locations. Results show that, although the BraggNN model is small, M2ML can significantly accelerate the workflow through selective assignment of tasks to different computing resources.
Weijian Zheng, Hemant Sharma, Ryan Chard, Peter Kenesei, Jun-Sang Park, Nicholas Schwarz, Antonino Miceli, Ian T. Foster, Rajkumar Kettimuthu
IEEE Big Data3
2023 Trillion Parameter AI Serving Infrastructure for Scientific Discovery: A Survey and Vision
abstract
Deep learning methods are transforming research, enabling new techniques, and ultimately leading to new discoveries. As the demand for more capable AI models continues to grow, we are now entering an era of Trillion Parameter Models (TPM), or models with more than a trillion parameters---such as Huawei's PanGu-Σ. We describe a vision for the ecosystem of TPM users and providers that caters to the specific needs of the scientific community. We then outline the significant technical challenges and open problems in system design for serving TPMs to enable scientific research and discovery. Specifically, we describe the requirements of a comprehensive software stack and interfaces to support the diverse and flexible requirements of researchers.
Nathaniel Hudson 0001, J. Gregory Pauloski, Matt Baughman, Alok Kamatar, Mansi Sakarvadia, Logan T. Ward, Ryan Chard, André Bauer 0001, Maksim Levental, Will Engler, Owen Price Skelly, Ben Blaiszik, Rick L. Stevens, Kyle Chard, Ian T. Foster
BDCAT7