EDBT 2026 Demo / reviewers in the wild / expert
Daniel Milroy
dblp:182/0265 · also Daniel J. Milroy
· DBLP profile ↗
8ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0001-6500-3227ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Flux Fiction: Hopping Toward Storage Graph Scheduling With El Capitan's RabbitsabstractModern HPC systems are placing increasing demands on job schedulers due to their scale and novel hardware. El Capitan’s Rabbit nodes exemplify this challenge: unlike traditional systems where storage is remote and shared, Rabbit nodes wire local NVMe SSDs directly to compute nodes via PCIe, forcing schedulers to actively track storage topology, capacity, and cross-job persistence, concerns they were never designed to handle. We introduce Flux Fiction, a fully plugin-based HPC system emulator built on top of Flux that replays historical job traces to evaluate scheduling policies in Flux. We validate Flux Fiction against the LLNL Tuolumne cluster using two workloads across four queueing policies, achieving a P99-bounded slowdown error below 1 in 7 of 8 experiments and a maximum utilization error of 1.2%. We then use Flux Fiction to explore Rabbit storage scheduling, demonstrating its ability to explore novel scheduling scenarios. Walter J. Ashworth, Ian Lumsden, Jim Garlick, Mark Grondona, Olga Pearce, Stephanie Brink, Daniel Milroy, Tapasya Patki, Thomas Scogland, Michela Taufer |
HPDC | 7 |
| 2026 | IFlux: Intent-Aware Storage Tiers & Software Scheduling for HPC Systems
Hariharan Devarajan, Vanessa Sochat, Daniel Milroy, Thomas Scogland |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2025 | Predictive Execution of Workflows in a HPC+Cloud EnvironmentabstractMeeting deadlines for data-intensive workflows on HPC systems is challenging as jobs experience varying wait times before resources become available. This impact is significant in hybrid HPC+Cloud scheduling, which can lead to resource idleness, deadline violations, and higher costs. To address these issues, we propose scheduling data-intensive workflows over a combined HPC+Cloud hybrid environment in a deterministic manner by scavenging unused HPC resources. We predict resource availability (RA) of HPC systems, and exploit this prediction to dynamically split resource allocation between HPC's unused and Cloud's on-demand resources to complete a workflow by a given deadline. The deterministic resource allocation allows for preloading input data for workflow tasks, avoiding execution delays. Further, we develop an adaptive scaling algorithm that effectively backs up the targeted HPC allocation on Cloud facilities to avoid workflow execution delays in the event of incorrect RA estimation. Experiments show that our scheduling technique imposes minimal impact on HPC production jobs, saves cost for$>75 \%$workflow runs, suggests accurate budgets with a mean 7.11 % to 14.75 % cost estimation error, and finishes a mean 98 % to 99.4 % of tasks before deadlines. Subhendu Behera, Jae-Seung Yeom, Daniel Milroy, Marc Niethammer, Frank Mueller 0001 |
HiPC | 3 |
| 2025 | Flux Emulator: First Insights into Optimizing Scheduling for Exascale HPCabstractEl Capitan, currently the world's largest supercomputer at 1.742 Ex-aflop/s, introduces challenges in scheduling due to its scale and innovative rabbit nodes, which traditional schedulers cannot efficiently handle. Flux, a resource and job management system, handles dynamic resource allocation tailored for exascale systems through its graph-based scheduler, Fluxion. This work introduces the Flux Emulator, a tool designed to test scheduling policies in Fluxion without impacting production systems. The emulator plugs into the real components of Flux and Fluxion to mimic job execution, emulate resource usage, and collect information on how the job behaves. Preliminary tests show negligible overhead introduced by the emulator and demonstrate its effectiveness in evaluating scheduli ng policies, like conservative backfilling, in a fraction of the time required with a real system. Walter J. Ashworth, Ian Lumsden, Jim Garlick, Mark Grondona, Olga Pearce, Stephanie Brink, Dewi Yokelson, Daniel Milroy, Tapasya Patki, Thomas Scogland, Michela Taufer |
HPDC | 8 |
| 2024 | Enabling Workload-Driven Elasticity in MPI-based EnsemblesabstractInterdisciplinary workflows are evolving to accom-modate the growing resource diversity and parallelism in modern computing systems. The necessity of integrating various components, including multi-scale simulations and Artificial Intelligence and Machine Learning (AI/ML) with ensemble methods, has made the workflows increasingly complex and challenging to manage using traditional High performance computing (HPC) infrastructure. Cloud computing provides capabilities such as container orchestration, automation, and elasticity to manage the growing heterogeneity, scale, and complexity of HPC systems and workflows. Converged computing, a growing movement that integrates HPC and cloud technologies into a seamless environ-ment, can provide a means to bridge the gap between needs and capabilities. In particular, ensemble-based HPC workflows can benefit from the potential efficiency improvements afforded by these capabilities. While MPI - based (Message Passing Interface) workflows have been demonstrated to scale using cloud-native orchestration in Kubernetes, there is little work on understanding the combined impact of autoscaling and elasticity on MPI-based workflows. In this study, we explore the cost and performance of elasticity applied to ensembles of MPI-based HPC simulations using the Flux Operator. We propose a workload-driven autoscaling algorithm that outperforms CPU utilization-based autoscaling for MPI-based ensembles. The efficiency gains afforded by elastic, autoscaled approaches for MPI - based ensembles are described. We demonstrate that the workload-driven algorithm can reduce ensemble completion time by up to$4.7\times$in comparison with CPU utilization-based autoscaling. Md Rajib Hossen, Vanessa Sochat, Abhik Sarkar, Mohammad A. Islam 0001, Daniel Milroy |
CLUSTER | 5 |
| 2024 | Predicting Cross-Architecture Performance of Parallel ProgramsabstractA variety of hardware architectures, both CPUs and GPUs, are used today to build supercomputers and parallel clusters. Often times, users can choose which hardware platform they want to run on. Modern scientific workflows have multiple computational tasks, and each task may be better suited for a different architecture in terms of performance. Deciding where to run an application or workflow task is not straightforward because of the complexity of applications, and hardware architectures, which makes performance predictions challenging. Hence, modeling the performance of scientific applications across a variety of architectures is important for achieving the best performance. In this paper, we present a machine learning based methodology to model the relative performance of applications across multiple architectures using hardware performance counters. Our machine learning model can predict the relative performance of an application with a mean absolute error of 0.11, and can be used effectively to make performance-aware and multi-architecture scheduling decisions, reducing makespan by up to 20%. Daniel Nichols, Alexander Movsesyan, Jae-Seung Yeom, Abhik Sarkar, Daniel Milroy, Tapasya Patki, Abhinav Bhatele |
IPDPS | 5 |
| 2022 | Scalable Composition and Analysis Techniques for Massive Scientific WorkflowsabstractComposite science workflows are gaining traction to manage the combined effects of (1) extreme hardware heterogeneity in new High Performance Computing (HPC) systems and (2) growing software complexity – effects necessitated by the convergence of traditional HPC with data sciences. Composing, analyzing, and optimizing a composite workflow remains highly challenging as the component technologies are generally developed in isolation and often feature widely varying levels of performance, scalability, and interoperability. In this paper, we propose novel workflow composition and analysis techniques to create and optimize a scalable and effective composite workflow for heterogeneous HPC centers, and define the performance space of variables that impact composite workflow performance. We present PerfFlowAspect, an Aspect Oriented Programming (AOP)-based tool to perform cross-cutting performance analysis of composite workflows and better understand the impact of key performance variables on workflows. Our solution directly addresses AOP concerns that can affect workflow performance and covers the full software lifecycle, ranging from the workflow's initial composition through performance analysis and optimization. We use our science workflow composition techniques to implement the American Heart Association Molecule Screening (AHA MoleS) workflow. Through experimentation, we demonstrate that tuning a single performance variable can improve AHA MoleS workflow performance by a factor of up to 2.45x. Our evaluation suggests that our techniques can significantly enhance the ability of a multi-disciplinary research and development team to create a high performance composite workflow. Dong H. Ahn, Jeffrey Mast, Stephen Herbein, Francesco Di Natale, Daniel A. Kirshner, Sam Ade Jacobs, Ian Karlin, Daniel Milroy, Bronis R. de Supinski, Brian Van Essen, Jonathan E. Allen, Felice C. Lightstone |
e-Science | 9 |
| 2019 | Making Root Cause Analysis Feasible for Large Code Bases: A Solution Approach for a Climate ModelabstractLarge-scale simulation codes that model complicated science and engineering applications typically have huge and complex code bases. For such simulation codes, where bit-for-bit comparisons are too restrictive, finding the source of statistically significant discrepancies (e.g., from a previous version, alternative hardware or supporting software stack) in output is non-trivial at best. Although there are many tools for program comprehension through debugging or slicing, few (if any) scale to a model as large as the Community Earth System Model (CESM#8482;), which consists of more than 1.5 million lines of Fortran code. Currently for the CESM, we can easily determine whether a discrepancy exists in the output using a by now well-established statistical consistency testing tool. However, this tool provides no information as to the possible cause of the detected discrepancy, leaving developers in a seemingly impossible (and frustrating) situation. Therefore, our aim in this work is to provide the tools to enable developers to trace a problem detected through the CESM output to its source. To this end, our strategy is to reduce the search space for the root cause(s) to a tractable size via a series of techniques that include creating a directed graph of internal CESM variables, extracting a subgraph (using a form of hybrid program slicing), partitioning into communities, and ranking nodes by centrality. Runtime variable sampling then becomes feasible in this reduced search space. We demonstrate the utility of this process on multiple examples of CESM simulation output by illustrating how sampling can be performed as part of an efficient parallel iterative refinement procedure to locate error sources, including sensitivity to CPU instructions. By providing CESM developers with tools to identify and understand the reason for statistically distinct output, we have positively impacted the CESM software development cycle and, in particular, its focus on quality assurance. Daniel Milroy, Allison H. Baker, Dorit Hammerling, Youngsung Kim, Elizabeth R. Jessup, Thomas Hauser |
HPDC | 1 |