Sara Kokkila Schumacher

dblp:252/2928 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0002-2338-4815ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Lowering the Barrier: A Science Gateway for Scalable Machine Learning
abstract
We present a modular science gateway that simplifies the deployment of machine learning workflows across heterogeneous computing environments. Designed for domain researchers without DevOps expertise, our system lowers the barrier to scalable ML by integrating Clowder, an open-source data management platform, with Ray and Kubernetes for distributed execution. The platform supports data ingestion, visualization, and metadata management, with ML-specific UI enhancements. Workflows are executed via containerized extractors that interact with a shared Ray cluster. To demonstrate real-world utility, we apply our system to the detection of ice wedge polygons from Arctic satellite imagery. Researchers can fine-tune and run inference workflows through a simple web interface, without managing infrastructure. This approach enhances accessibility and reproducibility, and promotes the reuse of data and models across research communities.
Vismayak Mohanarajan, Luigi Marini, Kenton McHenry, Sara Kokkila Schumacher, Amal Perera, Anna Liljedahl
eScience4
2025 Efficient and Cost-Effective HPC on the Cloud
abstract
HPC applications are increasingly utilizing cloud resources due to their cost-effectiveness. Among these resources, spot compute instances present an opportunity to run applications at deep discounts compared to on-demand instances. However, they present unique challenges for tightly-coupled HPC applications due to potential interruptions. Traditional parallel programming models like MPI are not inherently fault-tolerant, and existing methods to handle these interruptions are inefficient and require significant programmer effort. In this paper, we present Charm++ as an alternative solution that natively supports fault tolerance, dynamic load balancing, and resource rescaling. We present a tool to run Charm++ applications with a mix of on-demand and spot instances which can detect and efficiently handle spot interruptions without a shared filesystem. We show that using spot instances can result in up to 60% cost savings for our benchmark application.
Aditya Bhosale, Laxmikant V. Kalé, Sara Kokkila Schumacher
HPDC3
2025 The Cloud, Like Building and Running Go Binaries
abstract
We document the UX challenges of targeting distributed batch-processing code to Cloud resources. The challenges largely stem from the need for every team member to know everything about the mechanics of achieving scale. Thus, both running and writing code become overwhelming tasks. We present Lunchpail, a tool designed to address these challenges.We present two case studies, one of a team running AI/ML workloads and one of the code they wrote to make it happen. We quantify the challenges with two novel UX metrics: multiplicity and divergence. We show that the code base manifests a multitude of concerns, including distribution, packaging, and automation; 64–98% of the team’s code diverges from the main goal of the application. The story is paralleled when running workloads. Users switch between 3–7 types of tasks on a daily basis (high multiplicity). The nature of these tasks differ greatly from the users’ core competencies (high divergence). In particular, we show that all users assume the daily burdens of cluster operators.We demonstrate that four angles of attack, combined, can yield significant reductions in complexity: 1) Adopt a Serverless approach, allowing code to focus on that core "2%". 2) Treat application packaging like building a Golang binary via go build. This binary embeds source, configuration, deployment logic, and a lightweight runtime that channels data to workers with fan-out and queuing. 3) Treat running distributed applications pipelines against Cloud resources like launching said binaries, with simple bash "|" syntax; 4) When possible, avoid multi-tenancy, and instead target Cloud virtual machines directly.We present a large experimental study to quantify the viability of obtaining a dedicated "burst" of cloud resources for every job run. We show VMs can be ready in well under a minute, which is 10-20x faster than scaling a Kubernetes cluster.We embody this approach in Lunchpail. Lunchpail itself is small, weighing in at 12k lines of code (10% of the size of Kubeflow, 2.5% of Ray, 1% of Kueue). We validate Lunchpail against AI/ML code, legacy chip design workloads, and show that it adds little overhead on top of acquiring Cloud VMs.
Diana Arroyo, Paul Castro, Thuan Doan, Nick Mitchell, Sara Kokkila Schumacher, Ed Seabolt, Aleksander Slominski, Ansu Varghese, Lionel Villard, Cora Coleman
ICDCS5
2024 Automated Data Management and Learning-Based Scheduling for Ray-Based Hybrid HPC-Cloud Systems
Tingkai Liu, Huili Tao, Yicheng Lu, Zhongbo Zhu, Marquita Ellis, Sara Kokkila Schumacher, Volodymyr V. Kindratenko
Euro-Par (1)6
2021 Generalizable coordination of large multiscale workflows: challenges and learnings at scale
abstract
The advancement of machine learning techniques and the heterogeneous architectures of most current supercomputers are propelling the demand for large multiscale simulations that can automatically and autonomously couple diverse components and map them to relevant resources to solve complex problems at multiple scales. Nevertheless, despite the recent progress in workflow technologies, current capabilities are limited to coupling two scales. In the first-ever demonstration of using three scales of resolution, we present a scalable and generalizable framework that couples pairs of models using machine learning and in situ feedback. We expand upon the massively parallel Multiscale Machine-Learned Modeling Infrastructure (MuMMI), a recent, award-winning workflow, and generalize the framework beyond its original design. We discuss the challenges and learnings in executing a massive multiscale simulation campaign that utilized over 600,000 node hours on Summit and achieved more than 98% GPU occupancy for more than 83% of the time. We present innovations to enable several orders of magnitude scaling, including simultaneously coordinating 24,000 jobs, and managing several TBs of new data per day and over a billion files in total. Finally, we describe the generalizability of our framework and, with an upcoming open-source release, discuss how the presented framework may be used for new applications.
Harsh Bhatia, Francesco Di Natale, Joseph Y. Moon, Joseph R. Chavez, Fikret Aydin, Christopher B. Stanley, Tomas Oppelstrup, Chris Neale, Sara Kokkila Schumacher, Dong H. Ahn, Stephen Herbein, Timothy S. Carpenter, Sandrasegaram Gnanakaran, Peer-Timo Bremer, James N. Glosli, Felice C. Lightstone, Helgi I. Ingólfsson
SC10
2019 Preparation and optimization of a diverse workload for a large-scale heterogeneous system
abstract
Productivity from day one on supercomputers that leverage new technologies requires significant preparation. An institution that procures a novel system architecture often lacks sufficient institutional knowledge and skills to prepare for it. Thus, the "Center of Excellence" (CoE) concept has emerged to prepare for systems such as Summit and Sierra, currently the top two systems in the Top 500. This paper documents CoE experiences that prepared a workload of diverse applications and math libraries for a heterogeneous system. We describe our approach to this preparation, including our management and execution strategies, and detail our experiences with and reasons for using different programming approaches. Our early science and performance results show that the project enabled significant early seismic science with up to a l4X throughput increase over Cori. In addition to our successes, we discuss our challenges and failures so others may benefit from our experience.
Ian Karlin, Yoonho Park, Bronis R. de Supinski, Bert Still, D. A. Beckingsale, Robert Blake, Tong Chen 0001, Guojing Cong, Carlos H. A. Costa, Johann Dahm, Giacomo Domeniconi, Thomas Epperly, Aaron Fisher, Sara Kokkila Schumacher, Steve H. Langer, Hai Le, Naoya Maruyama, Xinyu Que, David F. Richards, Björn Sjögreen, Jonathan Wong, Carol S. Woodward, Ulrike Meier Yang, Bob Anderson, David Appelhans, Levi Barnes, Peter D. Barnes Jr., Sorin Bastea, David Böhme, Jamie A. Bramwell, James M. Brase, José R. Brunheroto, Barry Chen, Charway R. Cooper, Tony Degroot, Robert D. Falgout, Todd Gamblin, David J. Gardner, James N. Glosli, John A. Gunnels, Max P. Katz, Tzanio V. Kolev, I-Feng W. Kuo, Matthew P. LeGendre, Pei-Hung Lin, Shelby Lockhart, Kathleen McCandless, Claudia Misale, Jaime H. Moreno, Rob Neely, Jarom Nelson, Rao Nimmakayala, Kathryn M. O'Brien, Kevin O'Brien, Ramesh Pankajakshan, Roger A. Pearce, Slaven Peles, Phil Regier, Steven C. Rennich, Martin Schulz 0001, Howard Scott, James C. Sexton, Kathleen Shoga, Shiv Sundram, Guillaume Thomas-Collignon, Brian Van Essen, Alexey Voronin, Bob Walkup, Chris Ward, Hui-Fang Wen, Daniel A. White, Christopher Young, Cyril Zeller, Edward Zywicz
SC15
2019 A massively parallel infrastructure for adaptive multiscale simulations: modeling RAS initiation pathway for cancer
abstract
Computational models can define the functional dynamics of complex systems in exceptional detail. However, many modeling studies face seemingly incommensurate requirements: to gain meaningful insights into some phenomena requires models with high resolution (microscopic) detail that must nevertheless evolve over large (macroscopic) length- and time-scales. Multiscale modeling has become increasingly important to bridge this gap. Executing complex multiscale models on current petascale computers with high levels of parallelism and heterogeneous architectures is challenging. Many distinct types of resources need to be simultaneously managed, such as GPUs and CPUs, memory size and latencies, communication bottlenecks, and filesystem bandwidth. In addition, robustness to failure of compute nodes, network, and filesystems is critical.
Francesco Di Natale, Harsh Bhatia, Timothy S. Carpenter, Chris Neale, Sara Kokkila Schumacher, Tomas Oppelstrup, Liam Stanton, Shiv Sundram, Thomas Scogland, Gautham Dharuman, Michael P. Surh, Yue Yang 0034, Claudia Misale, Lars Schneidenbach, Carlos H. A. Costa, Changhoan Kim, Bruce D'Amora, Sandrasegaram Gnanakaran, Dwight V. Nissley, Frederick H. Streitz, Felice C. Lightstone, Peer-Timo Bremer, James N. Glosli, Helgi I. Ingólfsson
SC5