VLDB 2026 Research / reviewers in the wild / expert
Stephen Herbein
dblp:143/6409
· DBLP profile ↗
14ranked-venue papers
2as first author
4since 2021 · last 2022
0000-0003-0141-0653ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Scalable Composition and Analysis Techniques for Massive Scientific WorkflowsabstractComposite science workflows are gaining traction to manage the combined effects of (1) extreme hardware heterogeneity in new High Performance Computing (HPC) systems and (2) growing software complexity – effects necessitated by the convergence of traditional HPC with data sciences. Composing, analyzing, and optimizing a composite workflow remains highly challenging as the component technologies are generally developed in isolation and often feature widely varying levels of performance, scalability, and interoperability. In this paper, we propose novel workflow composition and analysis techniques to create and optimize a scalable and effective composite workflow for heterogeneous HPC centers, and define the performance space of variables that impact composite workflow performance. We present PerfFlowAspect, an Aspect Oriented Programming (AOP)-based tool to perform cross-cutting performance analysis of composite workflows and better understand the impact of key performance variables on workflows. Our solution directly addresses AOP concerns that can affect workflow performance and covers the full software lifecycle, ranging from the workflow's initial composition through performance analysis and optimization. We use our science workflow composition techniques to implement the American Heart Association Molecule Screening (AHA MoleS) workflow. Through experimentation, we demonstrate that tuning a single performance variable can improve AHA MoleS workflow performance by a factor of up to 2.45x. Our evaluation suggests that our techniques can significantly enhance the ability of a multi-disciplinary research and development team to create a high performance composite workflow. Dong H. Ahn, Jeffrey Mast, Stephen Herbein, Francesco Di Natale, Daniel A. Kirshner, Sam Ade Jacobs, Ian Karlin, Daniel Milroy, Bronis R. de Supinski, Brian Van Essen, Jonathan E. Allen, Felice C. Lightstone |
e-Science | 4 |
| 2022 | Reproducing and Extending Analytical Performance Models of Generalized Hierarchical SchedulingabstractWorkflows in High-Performance Computing (HPC) are rapidly changing towards more complex and large-scale workflows. In particular, high-throughput and ensemble workflows are becoming increasingly common. These workflows impose significant burden on current HPC scheduling systems which typically use slow, centralized schedulers. Generalized hierarchical scheduling (GHS) is a potential solution to face modern workflows but is not widely adopted in HPC yet. One difficulty hindering widespread adoption is the lack of performance models to configure and fit application requirements. The few existing models are often built on stick assumptions that can substantially reduce the analysis realism. In this paper, we reproduce the analysis and improve the realism of a state-of-the-art model presented in “An Analytical Performance Model of Generalized Hierarchical Scheduling” [1] by Herbein and co-authors. Specifically, we first reproduce four key analysis studies in the original paper and then expand the model by removing key assumptions, one at a time. In doing so, we extend the realism of the original model. We empirically validate our extended model using three different scenarios and discuss the observed accuracy. Jakob Lüttgau, Silvina Caíno-Lores, Kae Suarez, Dong H. Ahn, Stephen Herbein, Michela Taufer |
e-Science | 5 |
| 2022 | LuxIO: Intelligent Resource Provisioning and Auto-Configuration for Storage ServicesabstractStorage in HPC is typically a single Remote and Static Storage (RSS) resource. However, applications demonstrate diverse I/O requirements that can be better served by a multi-storage approach. Current practice employs ephemeral storage systems running on either node-local or shared storage resources. Yet, the burden of provisioning and configuring intermediate storage falls solely on the users, while global job schedulers offer little to no support for custom deployments. This lack of support often leads to over- or under-provisioning of resources and poorly configured storage systems. To mitigate this, we present LuxIO, an intelligent storage resource provisioning and auto-configuration service. LuxIO constructs storage deployments configured to best match I/O requirements. LuxIO-tuned storage services show performance improvements up to 2× across common applications and benchmarks, while introducing minimal overhead of 93.40 ms on top of existing job scheduling pipelines. LuxIO improves resource utilization by up to 25% in select workflows. Keith Bateman, Neeraj Rajesh, Jaime Cernuda, Luke Logan, Stephen Herbein, Antonios Kougkas, Xian-He Sun |
HIPC | 6 |
| 2021 | Generalizable coordination of large multiscale workflows: challenges and learnings at scaleabstractThe advancement of machine learning techniques and the heterogeneous architectures of most current supercomputers are propelling the demand for large multiscale simulations that can automatically and autonomously couple diverse components and map them to relevant resources to solve complex problems at multiple scales. Nevertheless, despite the recent progress in workflow technologies, current capabilities are limited to coupling two scales. In the first-ever demonstration of using three scales of resolution, we present a scalable and generalizable framework that couples pairs of models using machine learning and in situ feedback. We expand upon the massively parallel Multiscale Machine-Learned Modeling Infrastructure (MuMMI), a recent, award-winning workflow, and generalize the framework beyond its original design. We discuss the challenges and learnings in executing a massive multiscale simulation campaign that utilized over 600,000 node hours on Summit and achieved more than 98% GPU occupancy for more than 83% of the time. We present innovations to enable several orders of magnitude scaling, including simultaneously coordinating 24,000 jobs, and managing several TBs of new data per day and over a billion files in total. Finally, we describe the generalizability of our framework and, with an upcoming open-source release, discuss how the presented framework may be used for new applications. Harsh Bhatia, Francesco Di Natale, Joseph Y. Moon, Joseph R. Chavez, Fikret Aydin, Christopher B. Stanley, Tomas Oppelstrup, Chris Neale, Sara Kokkila Schumacher, Dong H. Ahn, Stephen Herbein, Timothy S. Carpenter, Sandrasegaram Gnanakaran, Peer-Timo Bremer, James N. Glosli, Felice C. Lightstone, Helgi I. Ingólfsson |
SC | 12 |
| 2020 | CanarIO: Sounding the Alarm on IO-Related Performance DegradationabstractUsers interact with High Performance Computing (HPC) machines through batch systems, which take user job submissions and allocate them to computing resources. While some resource managers have a generalized resource model, in nearly all modern systems, nodes are the only resource managed. Other resources, such as parallel file systems, are also necessary for jobs to make progress, but schedulers are blind to these resources. Facility staff can manually detect critical problems and manually hold jobs that need particular file systems, but this requires manual monitoring. Without human intervention, modern schedulers will happily run jobs whose required resources are not available. As a result, resources are wasted when IO-intensive jobs are scheduled on file systems with degraded performance.We introduce CanarIO, a tool for predicting the IO-sensitivity of HPC jobs and detecting IO-related performance degradation on HPC systems. CanarIO uses a set of "canary" IO probes run at regular intervals on the system. Using performance measurements from these jobs, CanarIO builds classifiers that can determine which jobs are IO-sensitive and when file system performance is degraded. We demonstrate the accuracy of our tool with a simulation of system execution using real HPC data. Specifically, we detect 37.5% of IO degradation events and correctly identify >90% of IO-sensitive jobs. We show that with CanarIO predictions we recover >1,500 node-hours in 10 days, with a potential maximum of nearly 10,000 node-hours. CanarIO is the first step necessary for augmenting schedulers to be resource-aware. Michael R. Wyatt II, Stephen Herbein, Kathleen Shoga, Todd Gamblin, Michela Taufer |
IPDPS | 2 |
| 2020 | Flux: Overcoming scheduling challenges for exascale workflows
Dong H. Ahn, Ned Bass, Albert Chu, Jim Garlick, Mark Grondona, Stephen Herbein, Helgi I. Ingólfsson, Joe Koning, Tapasya Patki, Thomas Scogland, Becky Springmeyer, Michela Taufer |
Future Gener. Comput. Syst. | 6 |
| 2018 | PRIONN: Predicting Runtime and IO using Neural NetworksabstractFor job allocation decision, current batch schedulers have access to and use only information on the number of nodes and runtime because it is readily available at submission time from user job scripts. User-provided runtimes are typically inaccurate because users overestimate or lack understanding of job resource requirements. Beyond the number of nodes and runtime, other system resources, including IO and network, are not available but play a key role in system performance. There is the need for automatic, general, and scalable tools that provide accurate resource usage information to schedulers so that, by becoming resource-aware, they can better manage system resources. Michael R. Wyatt II, Stephen Herbein, Todd Gamblin, Adam Moody, Dong H. Ahn, Michela Taufer |
ICPP | 2 |
| 2018 | Evaluation of an interference-free node allocation policy on fat-tree clusters
Samuel Pollard, Stephen Herbein, Abhinav Bhatele |
SC | 3 |
| 2016 | Machine Learning Predictions of Runtime and IO Traffic on High-End ClustersabstractWe use supervised machine learning algorithms (i.e., Decision Trees, Random Forest, and K-nearest Neighbors) to predict performance characteristics such as runtime and IO traffic of batch jobs on high-end clusters, using only user job scripts as input. We show that decision trees outperform other algorithms and accurately predict the runtime of 73% of jobs within a error tolerance of 10 minutes, which is a 51% improvement over the user requested runtime. Ryan McKenna, Stephen Herbein, Adam Moody, Todd Gamblin, Michela Taufer |
CLUSTER | 2 |
| 2016 | Scalable I/O-Aware Job Scheduling for Burst Buffer Enabled HPC ClustersabstractThe economics of flash vs. disk storage is driving HPC centers to incorporate faster solid-state burst buffers into the storage hierarchy in exchange for smaller parallel file system (PFS) bandwidth. In systems with an underprovisioned PFS, avoiding I/O contention at the PFS level will become crucial to achieving high computational efficiency. In this paper, we propose novel batch job scheduling techniques that reduce such contention by integrating I/O awareness into scheduling policies such as EASY backfilling. We model the available bandwidth of links between each level of the storage hierarchy (i.e., burst buffers, I/O network, and PFS), and our I/O-aware schedulers use this model to avoid contention at any level in the hierarchy. We integrate our approach into Flux, a next-generation resource and job management framework, and evaluate the effectiveness and computational costs of our I/O-aware scheduling. Our results show that by reducing I/O contention for underprovisioned PFSes, our solution reduces job performance variability by up to 33% and decreases I/O-related utilization losses by up to 21%, which ultimately increases the amount of science performed by scientific workloads. Stephen Herbein, Dong H. Ahn, Don Lipari, Thomas Scogland, Marc Stearman, Mark Grondona, Jim Garlick, Becky Springmeyer, Michela Taufer |
HPDC | 1 |
| 2016 | Performance characterization of irregular I/O at the extreme scale
Stephen Herbein, Sean McDaniel, Norbert Podhorszki, Jeremy Logan, Scott Klasky, Michela Taufer |
Parallel Comput. | 1 |
| 2015 | A Two-Tiered Approach to I/O Quality of Service in Docker ContainersabstractLinux containers allow applications to run in complete isolation from one another without the extra overhead of running entirely separate operating systems. This approach eliminates memory overheads associated with virtualization and virtual machines and helps businesses run their day-today applications. Unfortunately, multiple applications sharing the same resources can result in substantial resource contention among the applications in the containers and substantial performance loss. One way to mitigate this loss in performance is by ensuring quality of service (QoS) guaranteeing that the application of interest meets the performance requirements. Existing work targets ways of managing CPU, network, and memory contention, however, no solutions exist for managing contention associated with I/O. To address the I/O contention challenge in containers, we propose a two-tiered approach (i.e., at both the cluster and node levels) that extends Docker and Docker Swarm, making both capable of monitoring and controlling the I/O of Dockers containers. We demonstrate how our two-tiered approach has the potential for higher resource utilization without the effects of contention. Sean McDaniel, Stephen Herbein, Michela Taufer |
CLUSTER | 2 |
| 2014 | Using surrogate-based modeling to predict optimal I/O parameters of applications at the extreme scaleabstractOn petascale systems, the selection of optimal values for I/O parameters without taking into account the I/O size and pattern can cause the I/O time to dominate the simulation time, compromising the application's scalability. In this paper, we adopt and adapt an engineering method called surrogate-based modeling to efficiently search for the optimal I/O parameter values and accurately predict the associated I/O times at the extreme scale. Our approach allows us to address both the search and prediction in a short time, even when the application's I/O is large and exhibits irregular patterns. Michael Matheny, Stephen Herbein, Norbert Podhorszki, Scott Klasky, Michela Taufer |
ICPADS | 2 |
| 2013 | Efficient SDS Simulations on Multi-GPU Nodes of XSEDE High-End ClustersabstractEfficiently studying Sodium Dodecyl Sulfate (SDS) molecules' formations in the presence of different molar concentrations on high-end GPU clusters whose nodes share accelerators exposes us to several challenges, including the need to dynamically adapt the job lengths. Neither virtualization nor lightweight OS solutions can easily support generality, portability, and maintainability in concert. Our solution complements rather than rewrites existing workflow and resource managers with a companion module that complements functions of the workflow manager and a wrapper module that extends functions of the resource managers. Results on the Keene land cluster show how, by using our modules, accelerated SDS simulations more efficiently use the cluster's GPUs while leading to relevant scientific observations. Samuel Schlachter, Stephen Herbein, Michela Taufer, Shuching Ou, Sandeep Patel, Jeremy Logan |
e-Science | 2 |