Jay F. Lofstead

dblp:34/6823 · also Gerald F. Lofstead II, Gerald Fredrick Lofstead · DBLP profile ↗
← Back
48ranked-venue papers
14as first author
20since 2021 · last 2026
0000-0002-4697-2919ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 34 · 12 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Aligning Storage Benchmark Metrics with Application-Level Performance
abstract
The current state of practice in HPC is that performance metrics reported by storage benchmarks are disconnected from those obtained through application-level performance analysis tools, making it difficult for users to determine whether performance tuning efforts are effective or whether observed performance indicates underutilization of the system. In this work, we propose an approach to align IO500 storage benchmark results with application-level performance by deconstructing benchmark components and recalculating their metrics. The results show that certain universal metrics, such as bandwidth, can be meaningfully aligned with application performance, enabling more consistent and interpretable evaluation. Our findings also identify metrics that remain missing or cannot be reconciled, highlighting the need for standardization to align metrics produced by benchmarks and performance analysis tools.
Radita Liem, Julian M. Kunkel, Jay F. Lofstead, Sarah Neuwirth
SSDBM4
2025 Maximizing Insights, Minimizing Data: I/O Time Prediction Using Transfer Learning
Adrian Voß, Radita Liem, Julian M. Kunkel, Jay F. Lofstead, Philip H. Carns
HiPC4
2025 XIO: Toward eXplainable I/O for HPC Systems
Sarah Neuwirth, Hariharan Devarajan, Chen Wang 0004, Jay F. Lofstead
SSDBM4
2025 Guest Editorial:Special Section on SC22 Student Cluster Competition
abstract
Since 2015, as part of a Reproducibility Initiative, the International Conference for High Performance Computing, Networking, Storage, and Analysis (SC) has encouraged, and more recently required, all technical papers to include an Artifact Description (AD) Appendix. Each year, the conference selects one accepted paper to form the basis of a Reproducibility Challenge in the Student Cluster Competition (SCC) - a mainstay in the conference program - for the following year. In the SCC, teams of undergraduate students partner with their educational institution and vendors to build, benchmark and operate a small HPC system during the conference week. The Reproducibility Challenge component has them attempt to reproduce parts of the selected work on their own cluster. At SC22, teams worked to reproduce the SC21 paper “Productivity, Portability, Performance: Data-Centric Python” by Ziogas et al. This special section includes a selection of reports from student teams, along with an extended version of that paper, discussing the students’ results...
Omer F. Rana, Josef Spillner, Stephen Leak, Jay F. Lofstead, Rafael Tolosana-Calasanz
IEEE Trans. Parallel Distributed Syst.4
2024 PerSSD: Persistent, Shared, and Scalable Data with Node-Local Storage for Scientific Workflows in Cloud Infrastructure
abstract
Computational workflows need to retain data from both intermediate stages and final results to ensure the reproducibility and trustworthiness of scientific discoveries. While cloud infrastructure offers advantages like elasticity and automation, it compromises the persistence of intermediate data to ensure performance and reduce costs. Utilizing node-local storage can enhance performance but requires manual data transfers to persistent storage, making the technique impractical. To address these challenges, we propose a software architecture called Persistent, Shared, and Scalable Data (PerSSD) that integrates cloud operators and a Network File System (NFS) to make node-local data persistent and shareable across cloud nodes while ensuring performance. PerSSD outperforms traditional cloud object storage, achieving 35% reduction in the overall execution time of an earth science workflow, all while ensuring data persistence and shareability.
Paula Olaya, Sophia Wen, Jay F. Lofstead, Michela Taufer
IEEE Big Data3
2024 Hades: A Context-Aware Active Storage Framework for Accelerating Large-Scale Data Analysis
abstract
Modern simulation workflows generate and analyze massive amounts of data using I/O libraries like Adios2 and NetCDF. Although extensive work has optimized the I/O processes during the simulation phase, executing analytical queries—which often require iterative traversals of large files for insights—is cumbersome and usually constrained by low I/O performance. Instead of waiting for the analysis phase to process queries, quantities can be derived asynchronously during data production and cached, speeding up future queries. In this work, we introduce a context-aware I/O layer named ’Hades.’ It is designed to efficiently derive insights from selected quantities without compromising overall workflow performance. Hades actively and asynchronously computes and stores these quantities while the data is in transit. Hades leverages a hierarchical buffering system with data access-aware prefetching to ensure quick and timely access to relevant data. It offers a flexible query interface empowering users to easily define derived quantities and provide control over data placement decisions. Hades is implemented using an Adios2 plugin engine and the Hermes buffering platform, enabling transparent use by any Adios-powered application or workflow. Experimental results demonstrate performance improvements by up to 3-4x for tested real-world scientific producer-consumer workflows.
Jaime Cernuda, Luke Logan, Ana Gainaru, Scott Klasky, Jay F. Lofstead, Antonios Kougkas, Xian-He Sun
CCGrid5
2024 To Derive or Not to Derive: I/O Libraries Take Charge of Derived Quantities Computation
abstract
The ever-increasing volume of data produced by HPC simulations necessitates scalable methods for data exploration and knowledge extraction. Scientific data analysis often involves complex queries across distributed datasets, requiring manipulation of multiple primary variables and generating derived data that needs to be handled efficiently, creating challenges for applications that need to parse many large datasets. Relying on individual applications to handle all intermediate data generally leads to redundant computations across studies and unnecessary data transfers. In this paper, we investigate the performance of different approaches where applications define derived variables as quantities of interest (QoIs) and offload the computation and transfer of these QoIs to the I/O library. This significantly reduces redundancy and optimizes data movement across the distributed storage and processing infrastructure by allowing control over when and where derived variables are computed. We present a detailed analysis of the performance-storage trade-offs associated with different solutions and showcase results for our study on two large-scale datasets created from climate and combustion simulations.
Ana Gainaru, Norbert Podhorszki, Liz Dulac, Qian Gong, Scott Klasky, Greg Eisenhauer, Antonios Kougkas, Xian-He Sun, Jay F. Lofstead
SBAC-PAD9
2024 High-Quality I/O Bandwidth Prediction with Minimal Data via Transfer Learning Workflow
abstract
Providing a high-quality performance prediction has the potential to enhance various aspects of a cluster, such as devising scheduling and provisioning policies, guiding procurement decisions, suggesting candidate applications for tuning, and identifying probable scaling and porting challenges. Creating such a prediction for the I/O metrics is still challenging, however, due to the intricate interplay of multiple cluster components, making this an ideal case for machine learning. Nevertheless, achieving the required accuracy level with machine learning calls for a substantial amount of high-quality data, which is often a difficult challenge for most HPC clusters. In this work we explore the use of transfer learning to predict the applications’ I/O bandwidth based on a public dataset. As a result, our experiment can provide an I/O bandwidth prediction for a different cluster comparable to the current state-of-the-art result while employing 100 times less data than needed to construct the base model. Furthermore, we evaluate potential future improvements of the proposed workflow.
Dmytro Povaliaiev, Radita Liem, Julian M. Kunkel, Jay F. Lofstead, Philip H. Carns
SBAC-PAD4
2023 Enabling Scalability in the Cloud for Scientific Workflows: An Earth Science Use Case
abstract
Scientific discovery increasingly relies on interoperable, multimodular workflows generating intermediate data. The complexity of managing intermediate data may cause performance losses or unexpected costs. This paper defines an approach to composing these scientific workflows on cloud services, focusing on workflow data orchestration, management, and scalability. We demonstrate the effectiveness of our approach with the SOMOSPIE scientific workflow that deploys machine learning (ML) models to predict high-resolution soil moisture using an HPC service (LSF) and an open-source cloud-native service (K8s) and object storage. Our approach enables scientists to scale from coarse-grained to fine-grained resolution and from a small to a larger region of interest. Using our empirical observations, we generate a cost model for the execution of workflows with hidden intermediate data on cloud services.
Paula Olaya, Jakob Lüttgau, Camila Roa, Ricardo M. Llamas, Rodrigo Vargas, Sophia Wen, I-Hsin Chung, Seetharami R. Seelam, Yoonho Park, Jay F. Lofstead, Michela Taufer
CLOUD10
2023 Message from the Program Committee Chairs
abstract
Welcome to the 35th IEEE International Symposium on Computer Architecture and High Performance Computing, SBAC-PAD 2023, in Porto Alegre, Brazil. The conference is sponsored by the Brazilian Computer Society (SBC) and the IEEE Computer Society through two of its technical committees: the Technical Committee on Computer Architecture (TCCA), and the Technical Committee on Parallel Processing (TCPP).
Jay F. Lofstead, Vinod E. F. Rebello
SBAC-PAD1
2023 Building Trust in Earth Science Findings through Data Traceability and Results Explainability
abstract
To trust findings in computational science, scientists need workflows that trace the data provenance and support results explainability. As workflows become more complex, tracing data provenance and explaining results become harder to achieve. In this paper, we propose a computational environment that automatically creates a workflow execution's record trail and invisibly attaches it to the workflow's output, enabling data traceability and results explainability. Our solution transforms existing container technology, includes tools for automatically annotating provenance metadata, and allows effective movement of data and metadata across the workflow execution. We demonstrate the capabilities of our environment with the study of SOMOSPIE, an earth science workflow. Through a suite of machine learning modeling techniques, this workflow predicts soil moisture values from the 27 km resolution satellite data down to higher resolutions necessary for policy making and precision agriculture. By running the workflow in our environment, we can identify the causes of different accuracy measurements for predicted soil moisture values in different resolutions of the input data and link different results to different machine learning methods used during the soil moisture downscaling, all without requiring scientists to know aspects of workflow design and implementation.
Paula Olaya, Dominic Kennedy, Ricardo M. Llamas, Leobardo Valera, Rodrigo Vargas, Jay F. Lofstead, Michela Taufer
IEEE Trans. Parallel Distributed Syst.6
2022 Exploring Spatial Indexing for Accelerated Feature Retrieval in HPC
abstract
Despite the critical role that range queries play in analysis and visualization for HPC applications, there has been no comprehensive analysis of indices that are designed to accelerate range queries and the extent to which they are viable in HPC. In this paper we present the first such evaluation, examining 20 open-source C and C++ libraries that support range queries. Contributions of this paper include answering the following questions: which of the implementations are viable in HPC, how do these libraries compare in terms of build time, query time, memory usage, and scalability, what are other trade-offs between these implementations, is there a single overall best solution, and when does a brute force solution offer the best performance? We also share key insights learned during this process that can assist both HPC application scientists and spatial index developers. While we find that there is no single best solution, three libraries, Boost, CGAL and R-tree, offer some of the best performance, scalability, memory overheads, and support for different mesh types. We find several areas where the spatial indices could be substantially improved: better performance when there are a large number of query matches, reduced memory overheads, and better support for GPUs or other accelerators.
Margaret Lawson, William Gropp, Jay F. Lofstead
CCGRID3
2022 Failure Sources in Machine Learning for Medicine - A Study
abstract
Machine learning (ML) inherently suffers from at least a small amount of inaccuracy. Typically, these errors are acceptable in trade for either speed to an answer or the ability to find an answer at all. For high consequence domains, such as medicine where a wrong diagnosis can mean the difference between catching a disease early or not or prescribing debilitating treatment when it may not be needed, certain kinds and types errors are less acceptable. In a study attempting to reproduce ML for medicine research, many difficulties are encountered. These difficulties highlight both the need for higher standards to achieve reproducible ML in general and especially when it comes to high-stakes domains. This paper explores some of those difficulties with a focus on the error sources and discussions about how they may be addressed.
Hana Ahmed, Roselyne Tchoua, Jay F. Lofstead
e-Science3
2022 Augmenting Singularity to Generate Fine-grained Workflows, Record Trails, and Data Provenance
abstract
The use of containerization technology in high performance computing (HPC) workflows has substantially increased recently because it makes workflows much easier to develop and deploy. Although many HPC workflows include multiple data and multiple applications, they have traditionally all been bundled together into one monolithic container. This hinders the ability to trace the thread of execution, thus preventing scientists from establishing data provenance, or having workflow reproducibility. To provide a solution to this problem we extend the functionality of a popular HPC container runtime, Singularity. We implement both the ability to compose fine-grained containerized workflows and execute these workflows within the Singularity runtime with automatic metadata collection. Specifically, the new functionality collects a record trail of execution and creates data provenance. The use of our augmented Singularity is demonstrated with an earth science workflow, SOMOSPIE. The workflow is composed via our augmented Singularity which creates fine-grained containers and collects the metadata to trace, explain, and reproduce the prediction of soil moisture at a fine resolution.
Dominic Kennedy, Paula Olaya, Jay F. Lofstead, Rodrigo Vargas, Michela Taufer
e-Science3
2022 P-RECS'22: 5th International Workshop on Practical Reproducible Evaluation of Systems
abstract
The P-RECS workshop focuses heavily on practical, actionable aspects of reproducibility in broad areas of computational science and data exploration, with special emphasis on issues in which community collaboration can be essential for adopting novel methodologies, techniques and frameworks aimed at addressing some of the challenges we face today. The workshop brings together researchers and experts to share experiences and advance the state of the art in the reproducible evaluation of computer systems, featuring contributed papers and invited talks.
Jay F. Lofstead, Carlos Maltzahn, Ivo Jimenez
HPDC1
2022 NSDF-FUSE: A Testbed for Studying Object Storage via FUSE File Systems
abstract
This work presents NSDF-FUSE, a testbed for evaluating settings and performance of FUSE-based file systems on top of S3-compatible object storage; the testbed is part of a suite of services from the National Science Data Fabric (NSDF) project (an NSF-funded project that is delivering cyberinfrastructures for data scientists). We demonstrate how NSDF-FUSE can be deployed to evaluate eight different mapping packages that mount S3-compatible object storage to a file system, as well as six data patterns representing different I/O operations on two cloud platforms. NSDF-FUSE is open-source and can be easily extended to run with other software mapping packages and different cloud platforms.
Paula Olaya, Jakob Lüttgau, Naweiluo Zhou, Jay F. Lofstead, Giorgio Scorzelli, Valerio Pascucci, Michela Taufer
HPDC4
2022 LabStor: A Modular and Extensible Platform for Developing High-Performance, Customized I/O Stacks in Userspace
abstract
Traditionally, I/O systems have been developed within the confines of a centralized OS kernel. This led to monolithic and rigid storage systems that are limited by low development speed, expressiveness, and performance. Various assumptions are imposed including reliance on the UNIX-file abstraction, the POSIX standard, and a narrow set of I/O policies. However, this monolithic design philosophy makes it difficult to develop and deploy new I/O approaches to satisfy the rapidly-evolving I/O requirements of modern scientific applications. To this end, we propose LabStor: a modular and extensible platform for developing high-performance, customized I/O stacks. Single-purpose I/O modules (e.g, I/O schedulers) can be developed in the comfort of userspace and released as plug-ins, while end-users can compose these modules to form workload- and hardware-specific I/O stacks. Evaluations show that by switching to a fully modular design, tailored I/O stacks can yield performance improvements of up to 60% in various applications.
Luke Logan, Jaime Cernuda, Jay F. Lofstead, Xian-He Sun, Antonios Kougkas
SC3
2022 EMPRESS: Accelerating Scientific Discovery through Descriptive Metadata Management
abstract
High-performance computing scientists are producing unprecedented volumes of data that take a long time to load for analysis. However, many analyses only require loading in the data containing particular features of interest and scientists have many approaches for identifying these features. Therefore, if scientists store information (descriptive metadata) about these identified features, then for subsequent analyses they can use this information to only read in the data containing these features. This can greatly reduce the amount of data that scientists have to read in, thereby accelerating analysis. Despite the potential benefits of descriptive metadata management, no prior work has created a descriptive metadata system that can help scientists working with a wide range of applications and analyses to restrict their reads to data containing features of interest. In this article, we present EMPRESS, the first such solution. EMPRESS offers all of the features needed to help accelerate discovery: It can accelerate analysis by up to 300 ×, supports a wide range of applications and analyses, is high-performing, is highly scalable, and requires minimal storage space. In addition, EMPRESS offers features required for a production-oriented system: scalable metadata consistency techniques, flexible system configurations, fault tolerance as a service, and portability.
Margaret Lawson, William Gropp, Jay F. Lofstead
ACM Trans. Storage3
2021 pMEMCPY: a simple, lightweight, and portable I/O library for storing data in persistent memory
abstract
Persistent memory (PMEM) devices can achieve comparable performance to DRAM while providing significantly more capacity. This has made the technology compelling as an expansion to main memory. Rethinking PMEM as storage devices can offer a high performance buffering layer for HPC applications to temporarily, but safely store data. However, modern parallel I/O libraries, such as HDF5 and pNetCDF, are complicated and introduce significant software and metadata overheads when persisting data to these storage devices, wasting much of their potential. In this work, we explore the potential of PMEM as storage through pMEMCPY: a simple, lightweight, and portable I/O library for storing data in persistent memory. We demonstrate that our approach is up to 2x faster than other popular parallel I/O libraries under real workloads.
Luke Logan, Jay F. Lofstead, Scott Levy, Patrick M. Widener, Xian-He Sun, Antonios Kougkas
CLUSTER2
2021 Interpreting Write Performance of Supercomputer I/O Systems with Regression Models
abstract
This work seeks to advance the state of the art in HPC I/O performance analysis and interpretation. In particular, we demonstrate effective techniques to: (1) model output performance in the presence of I/O interference from production loads; (2) build features from write patterns and key parameters of the system architecture and configurations; (3) employ suitable machine learning algorithms to improve model accuracy. We train models with five popular regression algorithms and conduct experiments on two distinct production HPC platforms. We find that the lasso and random forest models predict output performance with high accuracy on both of the target systems. We also explore use of the models to guide adaptation in I/O middleware systems, and show potential for improvements of at least 15% from model-guided adaptation on 70% of samples, and improvements up to 10 x on some samples for both of the target systems.
Zilong Tan, Philip H. Carns, Jeffrey S. Chase, Kevin Harms, Jay F. Lofstead, Sarp Oral, Sudharshan S. Vazhkudai, Feiyi Wang
IPDPS6
2020 Stitch It Up: Using Progressive Data Storage to Scale Science
abstract
Generally, scientific simulations load the entire simulation domain into memory because most, if not all, of the data changes with each time step. This has driven application structures that have, in turn, affected the design of popular IO libraries, such as HDF-5, ADIOS, and NetCDF. This assumption makes sense for many cases, but there is also a significant collection of simulations where this approach results in vast swaths of unchanged data written each time step. This paper explores a new IO approach that is capable of stitching together a coherent global view of the total simulation space at any given time. This benefit is achieved with no performance penalty compared to running with the full data set in memory, at a radically smaller process requirement, and results in radical data reduction with no fidelity loss. Additionally, the structures employed enable online simulation monitoring.
Jay F. Lofstead, John Mitchell 0001, Enze Chen
IPDPS1
2020 Characterizing Output Bottlenecks of a Production Supercomputer: Analysis and Implications
abstract
This article studies the I/O write behaviors of the Titan supercomputer and its Lustre parallel file stores under production load. The results can inform the design, deployment, and configuration of file systems along with the design of I/O software in the application, operating system, and adaptive I/O libraries. We propose a statistical benchmarking methodology to measure write performance across I/O configurations, hardware settings, and system conditions. Moreover, we introduce two relative measures to quantify the write-performance behaviors of hardware components under production load. In addition to designing experiments and benchmarking on Titan, we verify the experimental results on one real application and one real application I/O kernel, XGC and HACC IO, respectively. These two are representative and widely used to address the typical I/O behaviors of applications. In summary, we find that Titan’s I/O system is variable across the machine at fine time scales. This variability has two major implications. First, stragglers lessen the benefit of coupled I/O parallelism (striping). Peak median output bandwidths are obtained with parallel writes to many independent files, with no striping or write sharing of files across clients (compute nodes). I/O parallelism is most effective when the application—or its I/O libraries—distributes the I/O load so that each target stores files for multiple clients and each client writes files on multiple targets in a balanced way with minimal contention. Second, our results suggest that the potential benefit of dynamic adaptation is limited. In particular, it is not fruitful to attempt to identify “good locations” in the machine or in the file system: component performance is driven by transient load conditions and past performance is not a useful predictor of future performance. For example, we do not observe diurnal load patterns that are predictable.
Sarp Oral, Christopher Zimmer 0001, Jong Choi 0001, David Dillow, Scott Klasky, Jay F. Lofstead, Norbert Podhorszki, Jeffrey S. Chase
ACM Trans. Storage7
2019 LABIOS: A Distributed Label-Based I/O System
abstract
In the era of data-intensive computing, large-scale applications, in both scientific and the BigData communities, demonstrate unique I/O requirements leading to a proliferation of different storage devices and software stacks, many of which have conflicting requirements. In this paper, we investigate how to support a wide variety of conflicting I/O workloads under a single storage system. We introduce the idea of a Label, a new data representation, and, we present LABIOS: a new, distributed, Label- based I/O system. LABIOS boosts I/O performance by up to 17x via asynchronous I/O, supports heterogeneous storage resources, offers storage elasticity, and promotes in-situ analytics via data provisioning. LABIOS demonstrates the effectiveness of storage bridging to support the convergence of HPC and BigData workloads on a single platform.
Antonios Kougkas, Hariharan Devarajan, Jay F. Lofstead, Xian-He Sun
HPDC3
2019 Session details: Cloud Systems
abstract
No abstract available.
Jay F. Lofstead
HPDC1
2019 Creating repeatable, reusable experimentation pipelines with popper: tutorial
abstract
Popper is an experimentation protocol for conducting scientific explorations and writing academic articles following a DevOps approach. The Popper CLI tool helps researchers automate the execution and validation of an experimentation pipeline. In this tutorial we give an introduction to the concepts and CLI tool, and go over hands-on exercises that help.
Ivo Jimenez, Jay F. Lofstead, Carlos Maltzahn
PPoPP2
2018 Harmonia: An Interference-Aware Dynamic I/O Scheduler for Shared Non-volatile Burst Buffers
abstract
Modern HPC systems employ burst buffer installations to reduce the peak I/O requirements for external storage and deal with the burstiness of I/O in modern scientific applications. These I/O buffering resources are shared between multiple applications that run concurrently. This leads to severe performance degradation due to contention, a phenomenon called cross-application I/O interference. In this paper, we first explore the negative effects of interference at the burst buffer layer and we present two new metrics that can quantitatively describe the slowdown applications experience due to interference. We introduce Harmonia, a new dynamic I/O scheduler that is aware of interference, adapts to the underlying system, implements a new 2-way decision-making process and employs several scheduling policies to maximize the system efficiency and applications' performance. Our evaluation shows that Harmonia, through better I/O scheduling, can outperform by 3× existing state-of-the-art buffering management solutions and can lead to better resource utilization.
Antonios Kougkas, Hariharan Devarajan, Xian-He Sun, Jay F. Lofstead
CLUSTER4
2018 Tintenfisch: File System Namespace Schemas and Generators
Michael Sevilla, Reza Nasirigerdeh, Carlos Maltzahn, Jeff LeFevre, Noah Watkins, Peter Alvaro, Margaret Lawson, Jay F. Lofstead, James Pivarski
HotStorage8
2018 quiho: Automated Performance Regression Testing Using Inferred Resource Utilization Profiles
abstract
We introduce quiho, a framework for profiling application performance that can be used in automated performance regression tests. quiho profiles an application by applying sensitivity analysis, in particular statistical regression analysis (SRA), using application-independent performance feature vectors that characterize the performance of machines. The result of the SRA, feature importance specifically, is used as a proxy to identify hardware and low-level system software behavior. The relative importance of these features serve as a performance profile of an application (termed inferred resource utilization profile or IRUP), which is used to automatically validate performance behavior across multiple revisions of an application»s code base without having to instrument code or obtain performance counters. We demonstrate that quiho can successfully discover performance regressions by showing its effectiveness in profiling application performance for synthetically introduced regressions as well as those found in real-world applications.
Ivo Jimenez, Noah Watkins, Michael Sevilla, Jay F. Lofstead, Carlos Maltzahn
ICPE4
2017 Predicting Output Performance of a Petascale Supercomputer
abstract
In this paper, we develop a predictive model useful for output performance prediction of supercomputer file systems under production load. Our target environment is Titan---the 3rd fastest supercomputer in the world---and its Lustre-based multi-stage write path. We observe from Titan that although output performance is highly variable at small time scales, the mean performance is stable and consistent over typical application run times. Moreover, we find that output performance is non-linearly related to its correlated parameters due to interference and saturation on individual stages on the path. These observations enable us to build a predictive model of expected write times of output patterns and I/O configurations, using feature transformations to capture non-linear relationships. We identify the candidate features based on the structure of the Lustre/Titan write path, and use feature transformation functions to produce a model space with 135,000 candidate models. By searching for the minimal mean square error in this space we identify a good model and show that it is effective.
Yezhou Huang, Jeffrey S. Chase, Jong Choi 0001, Scott Klasky, Jay F. Lofstead, Sarp Oral
HPDC6
2016 SuperGlue: Standardizing Glue Components for HPC Workflows
abstract
This poster describes our work on SuperGlue, a set of generic, reusable components for composing scientific workflows. These are distributed data analysis and manipulation tools that can be chained together to form a variety of real-time workflows providing analytical results during the execution of the primary scientific code. Unlike existing components used in IAWs, SuperGlue components do not have a fixed data type. This one change enables using these components on completely different kinds of simulations that share nothing in their output format. Key to making this work is using a typed transport mechanism between different components. Many options exist for these transports and the particular mechanism selected is not critical.
Jay F. Lofstead, Alexis Champsaur, Jai Dayal, Matthew Wolf, Greg Eisenhauer
CLUSTER1
2016 Managing I/O Interference in a Shared Burst Buffer System
abstract
In this work, we investigate the problem of inter-application interference in a shared Burst Buffer (BB) system. A BB is a new storage technology for HPC architectures that acts as an intermediate layer between performance-hungry HPC applications and the slow parallel file system. While the BB is meant to alleviate the problem of slow I/O in HPC systems, it is itself prone to performance degradation under interference. We observe that the magnitude of interference effects can reach a level that matters to the HPC system and the jobs that run on it. We investigate I/O scheduling techniques as a mechanism to mitigate BB I/O interference. With our results, we show that scheduling techniques tuned to BBs can control interference and significant performance benefits can be achieved.
Sagar Thapaliya, Purushotham V. Bangalore, Jay F. Lofstead, Kathryn Mohror, Adam Moody
ICPP3
2016 DAOS and friends: a proposal for an exascale storage system
abstract
The DOE Extreme-Scale Technology Acceleration Fast Forward Storage and IO Stack project is going to have significant impact on storage systems design within and beyond the HPC community. With phase two of the project starting, it is an excellent opportunity to explore the complete design and how it will address the needs of extreme scale platforms. This paper examines each layer of the proposed stack in some detail along with cross-cutting topics, such as transactions and metadata management. This paper not only provides a timely summary of important aspects of the design specifications but also captures the underlying reasoning that is not available elsewhere. We encourage the broader community to understand the design, intent, and future directions to foster discussion guiding phase two and the ultimate production storage stack based on this work. An initial performance evaluation of the early prototype implementation is also provided to validate the presented design.
Jay F. Lofstead, Ivo Jimenez, Carlos Maltzahn, Quincey Koziol, John Bent, Eric Barton
SC1
2015 SODA: Science-Driven Orchestration of Data Analytics
abstract
As scientific simulation applications evolve on the path towards exascale, a new model of scientific inquiry is required where concurrently with the running simulation, online analytics operate on the data it produces. By avoiding offline data storage except when absoluately necessary, it enables speeding up the scientific discovery process by providing rapid insights into the simulated science phenomena and affording more frequent, detailed data analytics than is possible with the traditional purely offline approach of using disk for intermediate data storage. However, a challenge for online analytics is to respond to behavior dynamics caused by changing simulation outputs and by unforeseen events on the underlying hardware/software platforms. This paper presents SODA, a set of run-time abstractions for online orchestration of data analytics, realized by embedding analytics tasks into workstations that monitor component behavior and enable responses to run-time changes in their resource demands and in the platform's resource availability. For high end simulations running on a leadership class machine, experimental evaluations show SODA can invoke efficient orchestration operations responding to a diverse set of run-time dynamics at different granularities to meet end-user and analysis specific requirements.
Jai Dayal, Jay F. Lofstead, Greg Eisenhauer, Karsten Schwan, Matthew Wolf, Hasan Abbasi, Scott Klasky
e-Science2
2015 The Role of Container Technology in Reproducible Computer Systems Research
abstract
Evaluating experimental results in the field of computer systems is a challenging task, mainly due to the many changes in software and hardware that computational environments go through. In this position paper, we analyze salient features of container technology that, if leveraged correctly, can help reduce the complexity of reproducing experiments in systems research. We present a use case in the area of distributed storage systems to illustrate the extensions that we envision, mainly in terms of container management infrastructure. We also discuss the benefits and limitations of using containers as a way of reproducing research in other areas of experimental systems research.
Ivo Jimenez, Carlos Maltzahn, Adam Moody, Kathryn Mohror, Jay F. Lofstead, Remzi H. Arpaci-Dusseau, Andrea C. Arpaci-Dusseau
IC2E5
2014 An innovative storage stack addressing extreme scale platforms and Big Data applications
abstract
Current production HPC IO stack design is unlikely to offer sufficient features and performance to adequately serve extreme scale science platform requirements as well as Big Data problems. A joint effort between the US Department of Energy's Office of Advanced Simulation and Computing and Advanced Scientific Computing Research commissioned a project to develop a design and prototype for an IO stack suitable for the extreme scale environment. It will be referred to as the Fast Forward Storage and IO (FFSIO) project. This is a joint effort led by Lawrence Livermore National Laboratory, with the DOE Data Management Nexus leads Rob Ross and Gary Grider as coordinators and contract lead Mark Gary.
Jay F. Lofstead, Ivo Jimenez, Carlos Maltzahn, Quincey Koziol, John Bent, Eric Barton
CLUSTER1
2014 Hello ADIOS: the challenges and lessons of developing leadership class I/O frameworks
abstract
SUMMARY Applications running on leadership platforms are more and more bottlenecked by storage input/output (I/O). In an effort to combat the increasing disparity between I/O throughput and compute capability, we created Adaptable IO System (ADIOS) in 2005. Focusing on putting users first with a service oriented architecture, we combined cutting edge research into new I/O techniques with a design effort to create near optimal I/O methods. As a result, ADIOS provides the highest level of synchronous I/O performance for a number of mission critical applications at various Department of Energy Leadership Computing Facilities. Meanwhile ADIOS is leading the push for next generation techniques including staging and data processing pipelines. In this paper, we describe the startling observations we have made in the last half decade of I/O research and development, and elaborate the lessons we have learned along this journey. We also detail some of the challenges that remain as we look toward the coming Exascale era. Copyright © 2013 John Wiley & Sons, Ltd.
Qing Liu 0002, Jeremy Logan, Yuan Tian 0004, Hasan Abbasi, Norbert Podhorszki, Jong Choi 0001, Scott Klasky, Roselyne Tchoua, Jay F. Lofstead, Ron A. Oldfield, Manish Parashar, Nagiza F. Samatova, Karsten Schwan, Arie Shoshani, Matthew Wolf, Kesheng Wu, Weikuan Yu
Concurr. Comput. Pract. Exp.9
2013 Experiences Applying Data Staging Technology in Unconventional Ways
abstract
Several efforts have shown the potential of using additional compute-area resources to enhance the IO path to storage. Efforts like data staging, IO forwarding, and similar techniques can accelerate IO performance and reduce the impact of IO time to a compute application. Hybrid staging enhanced this path by adding processing functionality to locations along the data path to storage. While these efforts have been effective, they have taken a somewhat limited view of the potential benefits using some additional compute resources can offer both to enhance a compute application as well as to offering a way to exploit HPC-style resources for non-traditional tasks. Over the last few years, we have been experimenting with the potential for other sorts of activities using a staging style approach to add or enable new functionality. The efforts in this area have yielded a collection of small projects that yield some insights into both the potential and limitations of this approach for both achieving exascale computing and for enabling alternative uses for HPC resources.
Jay F. Lofstead, Ron A. Oldfield, Todd Kordenbrock
CCGRID1
2013 A case of system-wide power management for scientific applications
abstract
The advance of high-performance computing systems towards exascale will be constrained by the systems' energy consumption levels. Large numbers of processing components, memory, interconnects, and storage components must all be considered to achieve exascale performance within a targeted energy bound. While application-aware power allocation schemes for computing resources are well studied, a portable and scalable budget-constrained power management scheme for scientific applications on exascale systems is still required. Execution activities within scientific applications can be categorized as CPU-bound, I/O-bound and communication-bound. Such activities tend to be clustered into ‘phases’, offering opportunities to manage their power consumption separately. Our experiments have demonstrated that their performance and energy consumption are affected differently by CPU frequency, an opportunity to fine tune CPU frequency for a minimal impact on the total execution time but significant savings on the energy consumption. By exploiting this opportunity, we present a phase-aware hierarchical power management framework that can opportunistically deliver good tradeoffs between system power consumption and application performance under a power budget. Our hierarchical power management framework consists of two main techniques: Phase-Aware CPU Frequency Scaling (PAFS) and opportunistic provisioning for power-constrained performance optimization. We have performed a systematic evaluation using both simulations and representative scientific applications on real systems. Our results show that our techniques can achieve 4.3%–17% better energy efficiency for large-scale scientific applications.
Jay F. Lofstead, Teng Wang 0001, Weikuan Yu
CLUSTER2
2013 Insights for exascale IO APIs from building a petascale IO API
abstract
Near the dawn of the petascale era, IO libraries had reached a stability in their function and data layout with only incremental changes being incorporated. The shift in technology, particularly the scale of parallel file systems and the number of compute processes, prompted revisiting best practices for optimal IO performance.
Jay F. Lofstead, Robert B. Ross
SC1
2012 D2T: Doubly Distributed Transactions for High Performance and Distributed Computing
abstract
Current exascale computing projections suggest rather than a monolithic simulation running for the majority of the machine, a collection of components comprising the scientific discovery process will be employed in an online workflow. This move to an online workflow scenario requires knowledge that inter-step operations are completed and correct before the next phase begins. Further, dynamic load balancing or fault tolerance techniques may dynamically deploy or redeploy resources for optimal use of computing resources. These newly configured resources should only be used if they are successfully deployed. Our D2T system offers a mechanism to support these kinds of operations by providing database-like transactions with distributed servers and clients. Ultimately, with adequate hardware support, full ACID compliance is possible for the transactions. To prove the viability of this approach, we show that the D2T protocol has less than 1.2 seconds of overhead using 4096 clients and 32 servers with good scaling characteristics using this initial prototype implementation.
Jay F. Lofstead, Jai Dayal, Karsten Schwan, Ron A. Oldfield
CLUSTER1
2012 Extending MPI to Better Support Multi-application Interaction
Jay F. Lofstead, Jai Dayal
EuroMPI1
2011 EDO: Improving Read Performance for Scientific Applications through Elastic Data Organization
abstract
Large scale scientific applications are often bottlenecked due to the writing of checkpoint-restart data. Much work has been focused on improving their write performance. With the mounting needs of scientific discovery from these datasets, it is also important to provide good read performance for many common access patterns, which requires effective data organization. To address this issue, we introduce Elastic Data Organization (EDO), which can transparently enable different data organization strategies for scientific applications. Through its flexible data ordering algorithms, EDO harmonizes different access patterns with the underlying file system. Two levels of data ordering are introduced in EDO. One works at the level of data groups (a.k.a process groups). It uses Hilbert Space Filling Curves (SFC) to balance the distribution of data groups across storage targets. Another governs the ordering of data elements within a data group. It divides a data group into sub chunks and strikes a good balance between the size of sub chunks and the number of seek operations. Our experimental results demonstrate that EDO is able to achieve balanced data distribution across all dimensions and improve the read performance of multidimensional datasets in scientific applications.
Yuan Tian 0004, Scott Klasky, Hasan Abbasi, Jay F. Lofstead, Ray W. Grout, Norbert Podhorszki, Qing Liu 0002, Yandong Wang 0001, Weikuan Yu
CLUSTER4
2011 Six degrees of scientific data: reading patterns for extreme scale science IO
abstract
Petascale science simulations generate 10s of TBs of application data per day, much of it devoted to their checkpoint/restart fault tolerance mechanisms. Previous work demonstrated the importance of carefully managing such output to prevent application slowdown due to IO blocking, resource contention negatively impacting simulation performance and to fully exploit the IO bandwidth available to the petascale machine. This paper takes a further step in understanding and managing extreme-scale IO. Specifically, its evaluations seek to understand how to efficiently read data for subsequent data analysis, visualization, checkpoint restart after a failure, and other read-intensive operations. In their entirety, these actions support the 'end-to-end' needs of scientists enabling the scientific processes being undertaken. Contributions include the following. First, working with application scientists, we define 'read' benchmarks that capture the common read patterns used by analysis codes. Second, these read patterns are used to evaluate different IO techniques at scale to understand the effects of alternative data sizes and organizations in relation to the performance seen by end users. Third, defining the novel notion of a 'data district' to characterize how data is organized for reads, we experimentally compare the read performance seen with the ADIOS middleware's log-based BP format to that seen by the logically contiguous NetCDF or HDF5 formats commonly used by analysis tools. Measurements assess the performance seen across patterns and with different data sizes, organizations, and read process counts. Outcomes demonstrate that high end-to-end IO performance requires data organizations that offer flexibility in data layout and placement on parallel storage targets, including in ways that can make tradeoffs in the performance of data writes vs. reads.
Jay F. Lofstead, Milo Polte, Garth A. Gibson, Scott Klasky, Karsten Schwan, Ron A. Oldfield, Matthew Wolf, Qing Liu 0002
HPDC1
2010 PreDatA - preparatory data analytics on peta-scale machines
abstract
Peta-scale scientific applications running on High End Computing (HEC) platforms can generate large volumes of data. For high performance storage and in order to be useful to science end users, such data must be organized in its layout, indexed, sorted, and otherwise manipulated for subsequent data presentation, visualization, and detailed analysis. In addition, scientists desire to gain insights into selected data characteristics `hidden' or `latent' in these massive datasets while data is being produced by simulations. PreDatA, short for Preparatory Data Analytics, is an approach to preparing and characterizing data while it is being produced by the large scale simulations running on peta-scale machines. By dedicating additional compute nodes on the machine as `staging' nodes and by staging simulations' output data through these nodes, PreDatA can exploit their computational power to perform select data manipulations with lower latency than attainable by first moving data into file systems and storage. Such intransit manipulations are supported by the PreDatA middleware through asynchronous data movement to reduce write latency, application-specific operations on streaming data that are able to discover latent data characteristics, and appropriate data reorganization and metadata annotation to speed up subsequent data access. PreDatA enhances the scalability and flexibility of the current I/O stack on HEC platforms and is useful for data pre-processing, runtime data analysis and inspection, as well as for data exchange between concurrently running simulations.
Fang Zheng 0003, Hasan Abbasi, Ciprian Docan, Jay F. Lofstead, Qing Liu 0002, Scott Klasky, Manish Parashar, Norbert Podhorszki, Karsten Schwan, Matthew Wolf
IPDPS4
2010 EFFIS: An End-to-end Framework for Fusion Integrated Simulation
abstract
The purpose of the Fusion Simulation Project is to develop a predictive capability for integrated modeling of magnetically confined burning plasmas. In support of this mission, the Center for Plasma Edge Simulation has developed an End-to-end Framework for Fusion Integrated Simulation (EFFIS) that combines critical computer science technologies in an effective manner to support leadership class computing and the coupling of complex plasma physics models. We describe here the main components of EFFIS and how they are being utilized to address our goal of integrated predictive plasma edge simulation.
Julian C. Cummings, Jay F. Lofstead, Karsten Schwan, Alex Sim, Arie Shoshani, Ciprian Docan, Manish Parashar, Scott Klasky, Norbert Podhorszki, Roselyne Tchoua
PDP2
2010 Managing Variability in the IO Performance of Petascale Storage Systems
abstract
Significant challenges exist for achieving peak or even consistent levels of performance when using IO systems at scale. They stem from sharing IO system resources across the processes of single largescale applications and/or multiple simultaneous programs causing internal and external interference, which in turn, causes substantial reductions in IO performance. This paper presents interference effects measurements for two different file systems at multiple supercomputing sites. These measurements motivate developing a 'managed' IO approach using adaptive algorithms varying the IO system workload based on current levels and use areas. An implementation of these methods deployed for the shared, general scratch storage system on Oak Ridge National Laboratory machines achieves higher overall performance and less variability in both a typical usage environment and with artificially introduced levels of 'noise'. The latter serving to clearly delineate and illustrate potential problems arising from shared system usage and the advantages derived from actively managing it.
Jay F. Lofstead, Fang Zheng 0003, Qing Liu 0002, Scott Klasky, Ron A. Oldfield, Todd Kordenbrock, Karsten Schwan, Matthew Wolf
SC1
2009 Extending I/O through high performance data services
abstract
The complexity of HPC systems has increased the burden on the developer as applications scale to hundreds of thousands of processing cores. Moreover, additional efforts are required to achieve acceptable I/O performance, where it is important how I/O is performed, which resources are used, and where I/O functionality is deployed. Specifically, by scheduling I/O data movement and by effectively placing operators affecting data volumes or information about the data, tremendous gains can be achieved both in the performance of simulation output and in the usability of output data. Previous studies have shown the value of using asynchronous I/O, of employing a staging area, and of performing select operations on data before it is written to disk. Leveraging such insights, this paper develops and experiments with higher level I/O abstractions, termed ldquodata servicesrdquo, that manage output data from `source to sink': where/when it is captured, transported towards storage, and filtered or manipulated by service functions to improve its information content. Useful services include data reduction, data indexing, and those that manage how I/O is performed, i.e., the control aspects of data movement. Our data services implementation distinguishes control aspects - the control plane - from data movement - the data plane, so that both may be changed separably. This results in runtime flexibility not only in which services to employ, but also in where to deploy them and how they use I/O resources. The outcome is consistently high levels of I/O performance at large scale, without requiring application change.
Hasan Abbasi, Jay F. Lofstead, Fang Zheng 0003, Karsten Schwan, Matthew Wolf, Scott Klasky
CLUSTER2
2009 Adaptable, metadata rich IO methods for portable high performance IO
abstract
Since IO performance on HPC machines strongly depends on machine characteristics and configuration, it is important to carefully tune IO libraries and make good use of appropriate library APIs. For instance, on current petascale machines, independent IO tends to outperform collective IO, in part due to bottlenecks at the metadata server. The problem is exacerbated by scaling issues, since each IO library scales differently on each machine, and typically, operates efficiently to different levels of scaling on different machines. With scientific codes being run on a variety of HPC resources, efficient code execution requires us to address three important issues: (1) end users should be able to select the most efficient IO methods for their codes, with minimal effort in terms of code updates or alterations; (2) such performance-driven choices should not prevent data from being stored in the desired file formats, since those are crucial for later data analysis; and (3) it is important to have efficient ways of identifying and selecting certain data for analysis, to help end users cope with the flood of data produced by high end codes. This paper employs ADIOS, the adaptable IO system, as an IO API to address (1)-(3) above. Concerning (1), ADIOS makes it possible to independently select the IO methods being used by each grouping of data in an application, so that end users can use those IO methods that exhibit best performance based on both IO patterns and the underlying hardware. In this paper, we also use this facility of ADIOS to experimentally evaluate on petascale machines alternative methods for high performance IO. Specific examples studied include methods that use strong file consistency vs. delayed parallel data consistency, as that provided by MPI-IO or POSIX IO. Concerning (2), to avoid linking IO methods to specific file formats and attain high IO performance, ADIOS introduces an efficient intermediate file format, termed BP, which can be converted, at small cost, to the standard file formats used by analysis tools, such as NetCDF and HDF-5. Concerning (3), associated with BP are efficient methods for data characterization, which compute attributes that can be used to identify data sets without having to inspect or analyze the entire data contents of large files.
Jay F. Lofstead, Fang Zheng 0003, Scott Klasky, Karsten Schwan
IPDPS1