Stephen Lien Harrell

dblp:175/6841 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
5since 2021 · last 2023
0000-0001-5327-525XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2023 Taming Metadata-intensive HPC Jobs Through Dynamic, Application-agnostic QoS Control
abstract
Modern I/O applications that run on HPC infrastructures are increasingly becoming read and metadata intensive. However, having multiple applications submitting large amounts of metadata operations can easily saturate the shared parallel file system's metadata resources, leading to overall performance degradation and I/O unfairness. We present PADLL, an application and file system agnostic storage middleware that enables QoS control of data and metadata workflows in HPC storage systems. It adopts ideas from Software-Defined Storage, building data plane stages that mediate and rate limit POSIX requests submitted to the shared file system, and a control plane that holistically coordinates how all I/O workflows are handled. We demonstrate its performance and feasibility under multiple QoS policies using synthetic benchmarks, real-world applications, and traces collected from a production file system. Results show that PADLL can enforce complex storage QoS policies over concurrent metadata-aggressive jobs, ensuring fairness and prioritization.
Ricardo Macedo, Mariana Miranda, Yusuke Tanimura, Jason H. Haga, Amit Ruhela, Stephen Lien Harrell, R. Todd Evans, José Pereira 0001, João Paulo 0001
CCGrid6
2022 Protecting Metadata Servers From Harm Through Application-level I/O Control
abstract
Modern large-scale I/O applications that run on HPC infrastructures are increasingly becoming metadata-intensive. Unfortunately, having multiple concurrent applications submitting massive amounts of metadata operations can easily saturate the shared parallel file system's metadata resources, leading to unresponsiveness of the storage backend and overall performance degradation. To address these challenges, we present Padll, a storage middleware that enables system administrators to proactively control and ensure QoS over metadata workflows in HPC storage systems. We demonstrate its performance and feasibility by controlling the rate of both synthetic and realistic I/O workloads. Results show that Padll can dynamically control metadata-aggressive workloads, prevent I/O burstiness, and ensure I/O fairness and prioritization.
Ricardo Macedo, Mariana Miranda, Yusuke Tanimura, Jason H. Haga, Amit Ruhela, Stephen Lien Harrell, R. Todd Evans, João Paulo 0001
CLUSTER6
2022 VPIC 2.0: Next Generation Particle-in-Cell Simulations
abstract
VPIC is a general purpose particle-in-cell simulation code for modeling plasma phenomena such as magnetic reconnection, fusion, solar weather, and laser-plasma interaction in three dimensions using large numbers of particles. VPIC's capacity in both fidelity and scale makes it particularly well-suited for plasma research on pre-exascale and exascale platforms. In this article, we demonstrate the unique challenges involved in preparing the VPIC code for operation at exascale, outlining important optimizations to make VPIC efficient on accelerators. Specifically, we show the work undertaken in adapting VPIC to exploit the portability-enabling framework Kokkos and highlight the enhancements to VPIC's modeling capabilities to achieve performance at exascale. We assess the achieved performance-portability trade-off through a suite of studies on nine different varieties of modern pre-exascale hardware. Our performance-portability study includes weak-scaling runs on three of the top ten TOP500 supercomputers, as well as a comparison of low-level system performance of hardware from four different vendors.
Robert F. Bird, Nigel Tan, Scott V. Luedtke, Stephen Lien Harrell, Michela Taufer, Brian J. Albright
IEEE Trans. Parallel Distributed Syst.4
2022 Advancing Adoption of Reproducibility in HPC: A Preface to the Special Section
abstract
In this special section we bring you a practice and experience effort in reproducibility for large-scale computational science at SC20. This section includes nine critiques, each by a student team that reproduced results from a paper published at SC19, during the following year’s Student Cluster Competition. The paper is also included in this section and has been expanded upon, now including an analysis of the outcomes of the students’ reproducibility experiments. Lastly, this special section encapsulates a variety of advances in reproducibility in the SC conference series technical program.
Stephen Lien Harrell, Scott Michael, Carlos Maltzahn
IEEE Trans. Parallel Distributed Syst.1
2021 Transparency and Reproducibility Practice in Large-Scale Computational Science: A Preface to the Special Section
abstract
With this special section we bring you a practice and experience effort in transparency and reproducibility for large-scale computational science. A unique section, it consists of a research work plus six critques, each by a student team that reproduced the work. The original research work has been expanded in its science and also in its contribution to open science with a discussion of the student effort. Our letter contemplates implications as well.
Beth Plale, Stephen Lien Harrell
IEEE Trans. Parallel Distributed Syst.2
2018 Hybrid HPC Cloud Strategies from the Student Cluster Competition
abstract
The value of running high performance and throughput applications on cloud or on-premises resources is a topic of increasing importance. At the annual Student Cluster Competition at the SC conference, teams race High Performance Compute (HPC) clusters for 46 hours straight to find solutions to scientific and engineering application workloads. In 2017, a cloud component was added, and students were tasked with deciding how to best to use the cloud. They had to decide which competition applications should run in the cloud versus their on-premises clusters and efficiently manage their computing budget, both in terms of on-premises power and their monetary cloud budget. Using these teams' experiences as a case study, we have a unique opportunity to examine and glean insights from the strategies, outcomes and attitudes toward using a hybrid model for HPC.
Stephen Lien Harrell
IEEE CLOUD1
2018 Special Issue on SCC'17 Reproducibility Initiative
C. Kristopher Garrett, Stephen Lien Harrell, Michael A. Heroux
Parallel Comput.2
2010 Cost-Effective HPC: The Community or the Cloud?
abstract
The increasing availability of commercial cloud computing resources in recent years has caught the attention of the high-performance computing (HPC) and scientific computing community. Many researchers have subsequently examined the relative computational performance of commercially available cloud computing offerings across a number of HPC application bench-marks and scientific workflows, but the analogous cost comparisons-i.e., comparisons between the cost of doing scientific computation in traditional HPC environments vs. cloud computing environments-are less frequently discussed and are difficult to make in meaning-ful ways. Such comparisons are of interest to traditional HPC resource providers as well as to members of the scientific research community who need access to HPC resources on a routine basis. This paper is a case study of costs incurred by faculty end-users of Purdue University's HPC “community cluster” program. We develop and present a per node-hour cloud computing equivalent cost that is based upon actual usage patterns of the community cluster participants and is suitable for direct comparison to hourly costs charged by one commercial cloud computing provider. We find that the majority of community cluster participants incur substantially lower out-of-pocket costs in this community cluster program than in purchasing cloud computing HPC products.
Adam G. Carlyle, Stephen Lien Harrell, Preston M. Smith
CloudCom2