Jim Garlick

dblp:19/6025 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Cloud and datacenter computing · 43% Electronic design automation · 38% Memory systems · 6%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
job scheduling
2.132026
Flux Fiction: Hopping Toward Storage Graph Scheduling With El Capitan's Rabbits · HPDC 2026
Flux Emulator: First Insights into Optimizing Scheduling for Exascale HPC · HPDC 2025
Scalable I/O-Aware Job Scheduling for Burst Buffer Enabled HPC Clusters · HPDC 2016
Electronic design automation › hardware verification and test › functional verification › emulation
full-system emulation
1.922026
Flux Fiction: Hopping Toward Storage Graph Scheduling With El Capitan's Rabbits · HPDC 2026
Flux Emulator: First Insights into Optimizing Scheduling for Exascale HPC · HPDC 2025
Memory systems
non-volatile memory
0.312026
Flux Fiction: Hopping Toward Storage Graph Scheduling With El Capitan's Rabbits · HPDC 2026
Parallel and multicore computing › parallel scheduling › resource-aware scheduling
i/o-aware scheduling
0.212016
Scalable I/O-Aware Job Scheduling for Burst Buffer Enabled HPC Clusters · HPDC 2016
Storage systems › storage performance
i/o interference
0.112016
Scalable I/O-Aware Job Scheduling for Burst Buffer Enabled HPC Clusters · HPDC 2016
Storage systems › file systems › distributed file system
parallel file system
0.112016
Scalable I/O-Aware Job Scheduling for Burst Buffer Enabled HPC Clusters · HPDC 2016

Methods — techniques the papers use, named apart from their topics

queueing policy evaluation · 1.0job trace replay · 1.0graph-based scheduling · 0.9conservative backfilling · 0.9bandwidth modeling · 0.2EASY backfilling · 0.2
YearPublicationVenuePosition
2026 Flux Fiction: Hopping Toward Storage Graph Scheduling With El Capitan's Rabbits
abstract
Modern HPC systems are placing increasing demands on job schedulers due to their scale and novel hardware. El Capitan’s Rabbit nodes exemplify this challenge: unlike traditional systems where storage is remote and shared, Rabbit nodes wire local NVMe SSDs directly to compute nodes via PCIe, forcing schedulers to actively track storage topology, capacity, and cross-job persistence, concerns they were never designed to handle. We introduce Flux Fiction, a fully plugin-based HPC system emulator built on top of Flux that replays historical job traces to evaluate scheduling policies in Flux. We validate Flux Fiction against the LLNL Tuolumne cluster using two workloads across four queueing policies, achieving a P99-bounded slowdown error below 1 in 7 of 8 experiments and a maximum utilization error of 1.2%. We then use Flux Fiction to explore Rabbit storage scheduling, demonstrating its ability to explore novel scheduling scenarios.
Walter J. Ashworth, Ian Lumsden, Jim Garlick, Mark Grondona, Olga Pearce, Stephanie Brink, Daniel Milroy, Tapasya Patki, Thomas Scogland, Michela Taufer
HPDC3
2025 Flux Emulator: First Insights into Optimizing Scheduling for Exascale HPC
abstract
El Capitan, currently the world's largest supercomputer at 1.742 Ex-aflop/s, introduces challenges in scheduling due to its scale and innovative rabbit nodes, which traditional schedulers cannot efficiently handle. Flux, a resource and job management system, handles dynamic resource allocation tailored for exascale systems through its graph-based scheduler, Fluxion. This work introduces the Flux Emulator, a tool designed to test scheduling policies in Fluxion without impacting production systems. The emulator plugs into the real components of Flux and Fluxion to mimic job execution, emulate resource usage, and collect information on how the job behaves. Preliminary tests show negligible overhead introduced by the emulator and demonstrate its effectiveness in evaluating scheduli ng policies, like conservative backfilling, in a fraction of the time required with a real system.
Walter J. Ashworth, Ian Lumsden, Jim Garlick, Mark Grondona, Olga Pearce, Stephanie Brink, Dewi Yokelson, Daniel Milroy, Tapasya Patki, Thomas Scogland, Michela Taufer
HPDC3
2020 Flux: Overcoming scheduling challenges for exascale workflows
Dong H. Ahn, Ned Bass, Albert Chu, Jim Garlick, Mark Grondona, Stephen Herbein, Helgi I. Ingólfsson, Joe Koning, Tapasya Patki, Thomas Scogland, Becky Springmeyer, Michela Taufer
Future Gener. Comput. Syst.4
2016 Scalable I/O-Aware Job Scheduling for Burst Buffer Enabled HPC Clusters
abstract
The economics of flash vs. disk storage is driving HPC centers to incorporate faster solid-state burst buffers into the storage hierarchy in exchange for smaller parallel file system (PFS) bandwidth. In systems with an underprovisioned PFS, avoiding I/O contention at the PFS level will become crucial to achieving high computational efficiency. In this paper, we propose novel batch job scheduling techniques that reduce such contention by integrating I/O awareness into scheduling policies such as EASY backfilling. We model the available bandwidth of links between each level of the storage hierarchy (i.e., burst buffers, I/O network, and PFS), and our I/O-aware schedulers use this model to avoid contention at any level in the hierarchy. We integrate our approach into Flux, a next-generation resource and job management framework, and evaluate the effectiveness and computational costs of our I/O-aware scheduling. Our results show that by reducing I/O contention for underprovisioned PFSes, our solution reduces job performance variability by up to 33% and decreases I/O-related utilization losses by up to 21%, which ultimately increases the amount of science performed by scientific workloads.
Stephen Herbein, Dong H. Ahn, Don Lipari, Thomas Scogland, Marc Stearman, Mark Grondona, Jim Garlick, Becky Springmeyer, Michela Taufer
HPDC7
2006 Data-Preservation in Scientific Workflow Middleware
abstract
This paper investigates data-preservation, a feature of scientific workflow middleware (SWM) useful for supporting data provenance and "smart recomputation." We observe that in order for an SWM supporting data preservation to achieve decent performance, it should execute on top of copy-on-write file systems. Unfortunately, most file systems in-use at scientific computing facilities were designed without copy-on-write semantics. In response, we design, implement and evaluate a middleware-level solution that is based on user-provided hints and parallelization. The solution can be deployed on top of current file systems and is able to scale almost arbitrarily. Our validation is based on real use-cases from astrophysics and experiments on a cluster with 4 file systems
David T. Liu, Michael J. Franklin, Ghaleb Abdulla, Jim Garlick, Marcus Miller
SSDBM4