Gary Grider

dblp:31/6753 · also Gary A. Grider · DBLP profile ↗
← Back
24ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0001-5749-3377ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Lessons from Profiling and Optimizing Placement in AMR Codes
abstract
Block-structured Adaptive Mesh Refinement (AMR), while essential for improving efficiency in large-scale irregular and dynamic simulations, poses unique optimization challenges. Previous work has identified load imbalance and synchronization overhead as key obstacles to performance, but the deep understanding of complex runtime behavior needed to systematically address them remains elusive. In this paper, we integrate telemetry collection, analysis, and intervention to bridge this understanding gap. Establishing reliable, actionable telemetry required systematic tuning to eliminate cross-stack performance anomalies. Leveraging this foundation we design CPLX, a tunable placement policy balancing compute load and communication locality, improving runtime by up to$\mathbf{2 1. 6 \%}$over optimized baselines. Our experience highlights the empirical nature of placement optimization, requiring theoretical models to be grounded in observed runtime behavior.
Ankush Jain, Chuck Cranor, Qing Zheng, Dominic Manno, George Amvrosiadis, Gary Grider
CLUSTER6
2024 CARP: Range Query-Optimized Indexing for Streaming Data
abstract
Ingestion of data generated by high-performance scientific applications continues to stress available storage resources. Efficient range-based analyses on this data can be enabled by reordering it on attributes of interest, but require expensive post-processing sorts to realize the query benefits of reordering. In-situ indexing techniques, while write-efficient, are orders of magnitude slower at range queries than sorted indices. Range queries are necessary for analyzing continuous physical attributes and tracking phenomena such as energy bands and wave fronts. We present CARP, a scalable data partitioner for range queries that reorders data in-situ as it is streamed to storage during application I/O. Motivated by our findings that real application distributions tend to be highly skewed and dynamic, CARP dynamically discovers and adapts its data partitions to track these characteristics. As a result, CARP can approximate the query performance of a sort without any ingestion overhead, making it $5 \times$ faster than prior work.
Ankush Jain, Chuck Cranor, Qing Zheng, Bradley W. Settlemyer, George Amvrosiadis, Gary Grider
SC6
2023 KV-CSD: A Hardware-Accelerated Key-Value Store for Data-Intensive Applications
abstract
Popular software key-value stores such as LevelDB and RocksDB are often tailored for efficient writing. Yet, they tend to also perform well on read operations. This is because while data is initially stored in a format that favors writes, it is later transformed by the DB in the background into a format that better accommodates reads. Write-optimized key-value stores can still block writes. This happens when those background workers cannot keep up with the foreground insertion workload.This paper advocates for a hardware-accelerated key-value store, enabling performance-critical operations, like background data reorganization and queries, to execute directly on storage instead of a host as existing key-value stores do. This better hides background work latency, prevents it from blocking foreground writes, and improves overall I/O efficiency. Our prototype, called KV-CSD, is a key-value based computational storage device consisting of an NVMe SSD and a System-on-a-Chip (SoC) that implements an ordered key-value store atop the SSD. Through offloaded processing, KV-CSD streamlines data insertion, reduces host-device data movement for both background data reorganization and query processing, and shows up to 10.6× lower write times and up to 7.4× faster queries compared to the current state-of-the-art software key-value stores on a real scientific dataset.
Inhyuk Park, Qing Zheng, Dominic Manno, Soonyeal Yang, Jason Lee 0004, David Bonnie, Bradley W. Settlemyer, Youngjae Kim 0001, Woosuk Chung, Gary Grider
CLUSTER10
2022 GUFI: Fast, Secure File System Metadata Search for Both Privileged and Unprivileged Users
abstract
Modern High-Performance Computing (HPC) data centers routinely store massive data sets resulting in millions of directories and billions of files. To efficiently search and sift through these files and directories we present the Grand Unified File Index (GUFI), a novel file system metadata index that enables both privileged and regular users to rapidly locate and characterize data sets of interest. GUFI uses a hierarchical index that preserves file access permissions such that the index can be securely accessed by users while still enabling efficient, advanced analysis of storage system usage by cluster administrators. Compared with the current state-of-the-art indexing for file system metadata, GUFI is able to provide speedups of 1.5× to 230× for queries executed by administrators on a real production file system namespace. Queries executed by users, which typically cannot rely on cluster-wide indexing, see even greater speedups using GUFI.
Dominic Manno, Jason Lee 0004, Prajwal Challa, Qing Zheng, David Bonnie, Gary Grider, Bradley W. Settlemyer
SC6
2021 DeltaFS: a scalable no-ground-truth filesystem for massively-parallel computing
abstract
High-Performance Computing (HPC) is known for its use of massive concurrency. But it can be challenging for a parallel filesystem's control plane to utilize cores when every client process must globally synchronize and serialize its metadata mutations with those of other clients. We present DeltaFS, a new paradigm for distributed filesystem metadata. DeltaFS allows jobs to self-commit their namespace changes to logs, avoiding the cost of global synchronization. Followup jobs selectively merge logs produced by previous jobs as needed, a principle we term No Ground Truth which allows for efficient data sharing. By avoiding unnecessary synchronization of metadata operations, DeltaFS improves metadata operation throughput up to 98X leveraging parallelism on the nodes where job processes run. This speedup grows as job size increases. DeltaFS enables efficient inter-job communication, reducing overall workflow runtime by significantly improving client metadata operation latency up to 49X and resource usage up to 52X.
Qing Zheng, Chuck Cranor, Gregory R. Ganger, Garth A. Gibson, George Amvrosiadis, Bradley W. Settlemyer, Gary Grider
SC7
2020 Streaming Data Reorganization at Scale with DeltaFS Indexed Massive Directories
abstract
Complex storage stacks providing data compression, indexing, and analytics help leverage the massive amounts of data generated today to derive insights. It is challenging to perform this computation, however, while fully utilizing the underlying storage media. This is because, while storage servers with large core counts are widely available, single-core performance and memory bandwidth per core grow slower than the core count per die. Computational storage offers a promising solution to this problem by utilizing dedicated compute resources along the storage processing path. We present DeltaFS Indexed Massive Directories (IMDs), a new approach to computational storage. DeltaFS IMDs harvest available (i.e., not dedicated) compute, memory, and network resources on the compute nodes of an application to perform computation on data. We demonstrate the efficiency of DeltaFS IMDs by using them to dynamically reorganize the output of a real-world simulation application across 131,072 CPU cores. DeltaFS IMDs speed up reads by 1,740× while only slightly slowing down the writing of data during simulation I/O for in situ data processing.
Qing Zheng, Chuck Cranor, Ankush Jain, Gregory R. Ganger, Garth A. Gibson, George Amvrosiadis, Bradley W. Settlemyer, Gary Grider
ACM Trans. Storage8
2019 Compact Filters for Fast Online Data Partitioning
abstract
We are approaching a point in time when it will be infeasible to catalog and query data after it has been generated. This trend has fueled research on in-situ data processing (i.e. operating on data as it is streamed to storage). One important example of this approach is in-situ data indexing. Prior work has shown the feasibility of indexing at scale as a two-step process. First, one partitions data by key across the CPU cores of a parallel job. Then each core indexes its subset as data is persisted. Online partitioning requires transferring data over the network so that it can be indexed and stored by the core responsible for the data. This approach is becoming increasingly costly as new computing platforms emphasize parallelism instead of individual core performance that is crucial for communication libraries and systems software in general. In addition to indexing, scalable online data partitioning is also useful in other contexts such as load balancing and efficient compression. We present FilterKV, an efficient data management scheme for fast online data partitioning of key-value (KV) pairs. FilterKV reduces the total amount of data sent over the network and to storage. We achieve this by: (a) partitioning pointers to KV pairs instead of the KV pairs themselves and (b) using a compact format to represent and store KV pointers. Results from LANL show that FilterKV can reduce total write slowdown (including partitioning overhead) by up to 3x across 4096 CPU cores.
Qing Zheng, Chuck Cranor, Ankush Jain, Gregory R. Ganger, Garth A. Gibson, George Amvrosiadis, Bradley W. Settlemyer, Gary Grider
CLUSTER8
2018 Fail-Slow at Scale: Evidence of Hardware Performance Faults in Large Production Systems
Haryadi S. Gunawi, Riza O. Suminto, Russell Sears, Casey Golliher, Swaminathan Sundararaman, Tim Emami, Weiguang Sheng, Nematollah Bidokhti, Caitie McCaffrey, Gary Grider, Parks M. Fields, Kevin Harms, Robert B. Ross, Andree Jacobson, Robert Ricci, Kirk Webb, Peter Alvaro, H. Birali Runesha, Mingzhe Hao, Huaicheng Li
FAST11
2018 Scaling embedded in-situ indexing with deltaFS
Qing Zheng, Chuck Cranor, Danhao Guo, Gregory R. Ganger, George Amvrosiadis, Garth A. Gibson, Bradley W. Settlemyer, Gary Grider
SC8
2018 Fail-Slow at Scale: Evidence of Hardware Performance Faults in Large Production Systems
abstract
Fail-slow hardware is an under-studied failure mode. We present a study of 114 reports of fail-slow hardware incidents, collected from large-scale cluster deployments in 14 institutions. We show that all hardware types such as disk, SSD, CPU, memory, and network components can exhibit performance faults. We made several important observations such as faults convert from one form to another, the cascading root causes and impacts can be long, and fail-slow faults can have varying symptoms. From this study, we make suggestions to vendors, operators, and systems designers.
Haryadi S. Gunawi, Riza O. Suminto, Russell Sears, Casey Golliher, Swaminathan Sundararaman, Tim Emami, Weiguang Sheng, Nematollah Bidokhti, Caitie McCaffrey, Deepthi Srinivasan, Biswaranjan Panda, Andrew Baptist, Gary Grider, Parks M. Fields, Kevin Harms, Robert B. Ross, Andree Jacobson, Robert Ricci, Kirk Webb, Peter Alvaro, H. Birali Runesha, Mingzhe Hao, Huaicheng Li
ACM Trans. Storage14
2017 An Early Functional and Performance Experiment of the MarFS Hybrid Storage EcoSystem
abstract
Many computing sites, LANL being one of them, have a requirement for long-term retention of mostly cold data. Although the main function of this storage tier is capacity, it does also have a bandwidth requirement. For many years, tape was the best economic solution for this requirement. However, over time, data sets have grown larger more quickly than tape bandwidth has improved. We have now entered a regime in which disk is the more economically efficient medium for this storage tier. Also more and more, data dominates the computing world. There is a "sea" of data out there in many different formats such as file, object, and Key-value that needs to be efficiently managed and effectively used. In this paper, we introduce a new hybrid storage system named MarFS. MarFS is a Near-POSIX File System using scale-out commercial/cloud for data and many POSIX file systems for metadata services. MarFS is an approach to support a data lake for HPC that sits on industry based commodity storage hardware and is a software layer that provides a global namespace and near POSIX semantics. MarFS provides the capability to serve as an umbrella over a variety of underling storage layers. In this paper, we present the system architecture of the proposed MarFS near-POISX file system, we conduct early functional performance testing cases on MarFS's software components, and finally we address current deployment status and future development works of the MarFS.
Hsing-bung Chen, Gary Grider, David Richard Montoya
IC2E2
2016 Power usage of production supercomputers and production workloads
abstract
Summary Power is becoming an increasingly important concern for large supercomputer centers. However, to date, there have been a dearth of studies of power usage ‘in the wild’—on production supercomputers running production workloads. In this paper, we present the initial results of a project to characterize the power usage of the three Top500 supercomputers at Los Alamos National Laboratory: Cielo, Roadrunner, and Luna (#15, #19, and #47, respectively, on the June 2012 Top500 list). Power measurements taken both at the switchboard level and within the compute racks are presented and discussed. Some noteworthy results of this study are that (1) variability in power consumption differs across architectures, even when running a similar workload and (2) Los Alamos National Laboratory's scientific workload draws, on average, only 70–75% of LINPACK power and only 40–55% of nameplate power, implying that power capping may enable a substantial reduction in power and cooling infrastructure while impacting comparatively few applications. Copyright © 2013 John Wiley & Sons, Ltd.
Scott Pakin, Curtis B. Storlie, Michael Lang 0003, Bob Fields, Eloy E. Romero Jr., Craig Idler, Sarah Ellen Michalak, Hugh Greenberg, Josip Loncaric, Randal Rheinheimer, Gary Grider, Joanne Wendelberger
Concurr. Comput. Pract. Exp.11
2015 The Impact of Vectorization on Erasure Code Computing in Cloud Storages - A Performance and Power Consumption Study
abstract
Erasure code storage systems are becoming popular choices for cloud storage systems due to cost-effective storage space saving schemes and higher fault-resilience capabilities. Both erasure code encoding and decoding procedures are involving heavy array, matrix, and table-lookup compute intensive operations. Multi-core, many-core, and streaming SIMD extension are implemented in modern CPU designs. In this paper, we study the power consumption and energy efficiency of erasure code computing using traditional Intel x86 platform and Intel Streaming SIMD extension platform. We use a breakdown power consumption analysis approach and conduct power studies of erasure code encoding process on various storage devices. We present the impact of various storage devices on erasure code based storage systems in terms of processing time, power utilization, and energy cost. Finally we conclude our studies and demonstrate the Intel x86's Streaming SIMD extensions computing is a cost-effective and favorable choice for future power efficient HPC cloud storage systems.
Hsing-bung Chen, Gary Grider, Jeff Inman, Parks Fields, Jeffery A. Kuehn
CLOUD2
2015 MDHIM: A Parallel Key/Value Framework for HPC
Hugh Greenberg, John Bent, Gary Grider
HotStorage3
2015 An empirical study of performance, power consumption, and energy cost of erasure code computing for HPC cloud storage systems
abstract
Erasure code storage systems are becoming popular choices for cloud storage systems due to cost-effective storage space saving schemes and higher fault-resilience capabilities. Both erasure code encoding and decoding procedures are involving heavy array, matrix, and table-lookup compute intensive operations. Multi-core, many-core, and streaming SIMD extension are implemented in modern CPU designs. In this paper, we study the power consumption and energy efficiency of erasure code computing using traditional Intel x86 platform and Intel Streaming SIMD extension platform. We use a breakdown power consumption analysis approach and conduct power studies of erasure code encoding process on various storage devices. We present the impact of various storage devices on erasure code based storage systems in terms of processing time, power utilization, and energy cost. Finally we conclude our studies and demonstrate the Intel x86's Streaming SIMD extensions computing is a cost-effective and favorable choice for future power efficient HPC cloud storage systems.
Hsing-bung Chen, Gary Grider, Jeff Inman, Parks Fields, Jeff Alan Kuehn
NAS2
2014 Cost of Tape versus Disk for Archival Storage
abstract
For archiving large datasets in high-performance computing facilities, tape technology has a long history of providing inexpensive capacity. However, as the memory-size of supercomputers continues to grow geometrically, the cost of tape bandwidth is becoming more important. The projected costs for tape-drives, robotics, and maintenance, are creating challenges for tape-based archives. The advent of erasure-coded object storage, driven by the "cloud storage" industry, might make it practical to implement archives using disks, or hybrid disk-and-tape systems. We used linear optimization techniques to investigate when and how this transition might best be made, taking into consideration our significant investment in tape technology. Our models introduce a technique to systematically relax constraints on the relationship between tape-capacity and tape-bandwidth, which governs a trade-off between cost and performance. We ran parameter studies that support some preliminary conclusions about paths forward for archive infrastructure at LANL.
Jeff Inman, Gary Grider, Hsing-bung Chen
IEEE CLOUD2
2013 I/O acceleration with pattern detection
John Bent, Aaron Torres, Gary Grider, Garth A. Gibson, Carlos Maltzahn, Xian-He Sun
HPDC4
2012 Jitter-free co-processing on a prototype exascale storage stack
abstract
In the petascale era, the storage stack used by the extreme scale high performance computing community is fairly homogeneous across sites. On the compute edge of the stack, file system clients or IO forwarding services direct IO over an interconnect network to a relatively small set of IO nodes. These nodes forward the requests over a secondary storage network to a spindle-based parallel file system. Unfortunately, this architecture will become unviable in the exascale era. As the density growth of disks continues to outpace increases in their rotational speeds, disks are becoming increasingly cost-effective for capacity but decreasingly so for bandwidth. Fortunately, new storage media such as solid state devices are filling this gap; although not cost-effective for capacity, they are so for performance. This suggests that the storage stack at exascale will incorporate solid state storage between the compute nodes and the parallel file systems. There are three natural places into which to position this new storage layer: within the compute nodes, the IO nodes, or the parallel file system. In this paper, we argue that the IO nodes are the appropriate location for HPC workloads and show results from a prototype system that we have built accordingly. Running a pipeline of computational simulation and visualization, we show that our prototype system reduces total time to completion by up to 30%.
John Bent, Sorin Faibish, James P. Ahrens, Gary Grider, John Patchett, Percy Tzelnic, Jonathan Woodring
MSST4
2012 Storage challenges at Los Alamos National Lab
abstract
There yet exist no truly parallel file systems. Those that make the claim fall short when it comes to providing adequate concurrent write performance at large scale. This limitation causes large usability headaches in HPC. Users need two major capabilities missing from current parallel file systems. One, they need low latency interactivity. Two, they need high bandwidth for large parallel IO; this capability must be resistant to IO patterns and should not require tuning. There are no existing parallel file systems which provide these features. Frighteningly, exascale renders these features even less attainable from currently available parallel file systems. Fortunately, there is a path forward.
John Bent, Gary Grider, Brett Kettering, Adam Manzanares, Meghan McClelland, Aaron Torres, Alfred Torrez
MSST2
2012 On the role of burst buffers in leadership-class storage systems
abstract
The largest-scale high-performance (HPC) systems are stretching parallel file systems to their limits in terms of aggregate bandwidth and numbers of clients. To further sustain the scalability of these file systems, researchers and HPC storage architects are exploring various storage system designs. One proposed storage system design integrates a tier of solid-state burst buffers into the storage system to absorb application I/O requests. In this paper, we simulate and explore this storage system design for use by large-scale HPC systems. First, we examine application I/O patterns on an existing large-scale HPC system to identify common burst patterns. Next, we describe enhancements to the CODES storage system simulator to enable our burst buffer simulations. These enhancements include the integration of a burst buffer model into the I/O forwarding layer of the simulator, the development of an I/O kernel description language and interpreter, the development of a suite of I/O kernels that are derived from observed I/O patterns, and fidelity improvements to the CODES models. We evaluate the I/O performance for a set of multiapplication I/O workloads and burst buffer configurations. We show that burst buffers can accelerate the application perceived throughput to the external storage system and can reduce the amount of external storage bandwidth required to meet a desired application perceived throughput goal.
Ning Liu 0008, Jason Cope, Philip H. Carns, Christopher D. Carothers, Robert B. Ross, Gary Grider, Adam Crume, Carlos Maltzahn
MSST6
2010 Integration Experiences and Performance Studies of A COTS Parallel Archive System
abstract
Present and future Archive Storage Systems have been challenged to (a) scale to very high bandwidths, (b) scale in metadata performance, (c) support policy-based hierarchical storage management capability, (d) scale in supporting changing needs of very large data sets, (e) support standard interface, and (f) utilize commercial-off-the-shelf (COTS) hardware. Parallel file systems have also been demanded to perform the same manner but at one or more orders of magnitude faster in performance. Archive systems continue to improve substantially comparable to file systems in their design due to the need for speed and bandwidth, especially metadata searching speeds such as more caching and less robust semantics. Currently, the number of extreme highly scalable parallel archive solutions is very limited especially for moving a single large striped parallel disk file onto many tapes in parallel. We believe that a hybrid storage approach of using COTS components and an innovative software technology can bring new capabilities into a production environment for the HPC community. This solution is much faster than the approach of creating and maintaining a complete end-to-end unique parallel archive software solution. We relay our experience of integrating a global parallel file system and a standard backup/archive product with an innovative parallel software code to construct a scalable and parallel archive storage system. Our solution has a high degree of overlap with current parallel archive products including (a) doing parallel movement to/from tape for a single large parallel file, (b) hierarchical storage management, (c) ILM features, (d) high volume (non-single parallel file) archives for backup/archive/content management, and (e) leveraging all free file movement tools in Linux such as copy, move, ls, tar, etc. We have successfully applied our working COTS Parallel Archive System to the current world's first petaflop/s computing system, LANL's Roadrunner machine, and demonstrated its capability to address requirements of future archival storage systems. Now this new Parallel Archive System is used on the LANL's Turquoise Network.
Hsing-bung Chen, Gary Grider, Cody Scott, Milton Turley, Aaron Torres, Kathy Sanchez, John Bremer
CLUSTER2
2009 PLFS: a checkpoint filesystem for parallel applications
abstract
Parallel applications running across thousands of processors must protect themselves from inevitable system failures. Many applications insulate themselves from failures by checkpointing. For many applications, checkpointing into a shared single file is most convenient. With such an approach, the size of writes are often small and not aligned with file system boundaries. Unfortunately for these applications, this preferred data layout results in pathologically poor performance from the underlying file system which is optimized for large, aligned writes to non-shared files. To address this fundamental mismatch, we have developed a virtual parallel log structured file system, PLFS. PLFS remaps an application's preferred data layout into one which is optimized for the underlying file system. Through testing on PanFS, Lustre, and GPFS, we have seen that this layer of indirection and reorganization can reduce checkpoint time by an order of magnitude for several important benchmarks and real applications without any application modification.
John Bent, Garth A. Gibson, Gary Grider, Ben McClelland, Paul Nowoczynski, James Nunez, Milo Polte, Meghan Wingate
SC3
2007 A Cost-Effective, High Bandwidth Server I/O network Architecture for Cluster Systems
abstract
In this paper we present a cost-effective, high bandwidth server I/O network architecture, named PaScal (Parallel and Scalable). We use the PaScal server I/O network to support data-intensive scientific applications running on very large-scale Linux clusters. PaScal server I/O network architecture provides (1) bi-level data transfer network by combining high speed interconnects for computing inter-process communication (IPC) requirements and low-cost gigabit Ethernet interconnect for global IP based storage/file access, (2) bandwidth on demand I/O network architecture without re-wiring and reconfiguring the system, (3) multi-path routing scheme, (4) reliability improvement through reducing large number of network components in server I/O network, and (5) global storage/file systems support in heterogeneous multi-cluster and grids environments. We have compared the PaScal server I/O network architecture with the federated server I/O network architecture (FESIO). Concurrent MPI-I/O performance testing results and deployment cost comparison demonstrate that the PaScal server I/O network architecture can outperform the FESIO network architecture in many categories: cost-effectiveness, scalability, and manageability and ease of large-scale I/O network.
Hsing-bung Chen, Gary Grider, Parks Fields
IPDPS2
2006 PaScal - a new parallel and scalable server IO networking infrastructure for supporting global storage/file systems in large-size Linux clusters
abstract
This paper presents the design and implementation of a new I/O networking infrastructure, named PaScal (parallel and scalable I/O networking framework). PaScal is used to support high data bandwidth IP based global storage systems for large scale Linux clusters. PaScal has several unique properties. It employs (1) Multi-level switch-fabric interconnection network by combining high speed interconnects for computing inter-process communication (IPC) requirements and low-cost Gigabit Ethernet interconnect for global IP based storage/file access, (2) A bandwidth on demand scaling I/O networking architecture, (3) open-standard IP networks (routing and switching), (4) multipath routing for load balancing and failover, (5) open shortest path first (OSPF) routing software, and (6) Supporting a global file system in multi-cluster and multi-platform environments. We describe both the hardware and software components of our proposed PaScal. We have implemented the PaScal I/O infrastructure on several large-size Linux clusters at LANL. We have conducted a sequence of parallel MPI-IO assessment benchmarks on LANL's Pink 1024 node Linux cluster and the Panasas global parallel file system. Performance results from our parallel MPI-IO benchmarks on the Pink cluster demonstrate that the PaScal I/O Infrastructure is robust and capable of scaling in bandwidth on large-size Linux clusters
Gary Grider, Hsing-bung Chen, James Nunez, Stephen W. Poole, Rosie Wacha, Parks Fields, Robert Martinez, Paul Martinez, Satsangat Khalsa, Abbie Matthews, Garth A. Gibson
IPCCC1