Galen M. Shipman

dblp:23/3104 · DBLP profile ↗
← Back
32ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0001-6297-2145ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 22 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 6 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 6Applied, interdisciplinary, general and emerging computing · 6 · 1 first-authorArtificial intelligence and machine learning · 3Computer networks · 1
YearPublicationVenuePosition
2026 Machine Learning-Driven Early Performance Prediction Framework for Accelerated Microarchitecture Simulation
abstract
Rapid and accurate performance estimation is critical in evaluating novel microarchitectures, as it enables efficient exploration of architectural trade-offs. Unfortunately, traditional simulation techniques, while precise in predicting performance and power, incur tremendous slowdowns versus real machines. Despite prior works having explored machine learning–based performance prediction, the area remains far from sufficiently studied with existing approaches typically requiring large comprehensive datasets, frequent retraining, and heavy memory footprints with limited accuracy. Here, we introduce a new, fast and accurate, early-stage preview framework that uses partial simulation data, and leverages a smaller, faster tree-based machine learning (ML) model to forecast performance metrics such as IPC and Power. By training on a diverse set of configurations, our framework dynamically captures relationships between microarchitectural parameters in large OoO cores versus overall performance and other metrics. Collecting data from as few as 10 sample points taken during warmup, representing only 25 million instructions, our models achieve mean absolute percentage errors of 3-4%, preserving a majority of the model’s predictive accuracy while achieving a 25× speedup (96% reduction in simulation time). By comparison, linear regression techniques from the same point in simulation show an error of 50%. In cache DSE, we improve ranking accuracy by 25× compared to state-of-the-art prediction methods. Our results also show the proposed framework can accurately predict the performance of unseen (untrained) microarchitectural components including new prefetchers and branch predictors.
Aiden Stickney, Osvaldo Castro, Aaron Chan, Paul Gratz, Jiang Hu 0001, Aakash Tyagi, Jered Dominguez-Trujillo, Galen M. Shipman, Kevin Sheridan
DATE8
2025 DX100: Programmable Data Access Accelerator for Indirection
abstract
Indirect memory accesses frequently appear in applications where memory bandwidth is a critical bottleneck.Prior indirect memory access proposals, such as indirect prefetchers, runahead execution, fetchers, and decoupled access/execute architectures, primarily focus on improving memory access latency by loading data ahead of computation but still rely on the DRAM controllers to reorder memory requests and enhance memory bandwidth utilization.DRAM controllers have limited visibility to future memory accesses due to the small capacity of request buffers and the restricted memorylevel parallelism of conventional core and memory systems.We introduce DX100, a programmable data access accelerator for indirect memory accesses.DX100 is shared across cores to offload bulk indirect memory accesses and associated address calculation operations.DX100 reorders, interleaves, and coalesces memory requests to improve DRAM row-buffer hit rate and memory bandwidth utilization.DX100 provides a general-purpose ISA to support diverse access types, loop patterns, conditional accesses
Alireza Khadem, Kamalakkannan Kamalavasan, Zhenyan Zhu, Akash Poptani, Yufeng Gu, Jered Dominguez-Trujillo, Nishil Talati, Daichi Fujiki, Scott A. Mahlke, Galen M. Shipman, Reetuparna Das
ISCA10
2025 Performance Analysis of Open MPI on AMR Applications over Slingshot-11
Maxim Moraru, Howard Pritchard, Derek Schafer, Galen M. Shipman, Patrick G. Bridges
EuroMPI4
2024 Optimizing Neighbor Collectives with Topology Objects
abstract
Many HPC applications implement non-cartesian neighbor data exchanges using MPI point-to-point operations rather than utilizing native MPI neighbor collective methods. Each application must therefore implement their own commu-nication optimizations, rather than leveraging any optimizations that could be provided by MPI. While an interface for such optimizations is provided within MPI through neighborhood collectives, applications avoid these methods due to the lack of performance optimizations within them along with large costs associated with graph communicator formation. This paper presents a novel approach for creating local, non-cartesian topol-ogy objects that provides finer control over the aforementioned setup costs. Any additional setup costs, such as initializing per-iteration optimizations, can then be deferred until additional information is available, such as within persistent initialization calls. This paper describes our implementation within an MPI extension library and demonstrates the effectiveness of our approach in simple benchmarks and real-world applications.
Gerald Collom, Derek Schafer, Amanda Bienz, Patrick G. Bridges, Galen M. Shipman
CLUSTER5
2021 Scaling implicit parallelism via dynamic control replication
abstract
We present dynamic control replication, a run-time program analysis that enables scalable execution of implicitly parallel programs on large machines through a distributed and efficient dynamic dependence analysis. Dynamic control replication distributes dependence analysis by executing multiple copies of an implicitly parallel program while ensuring that they still collectively behave as a single execution. By distributing and parallelizing the dependence analysis, dynamic control replication supports efficient, on-the-fly computation of dependences for programs with arbitrary control flow at scale. We describe an asymptotically scalable algorithm for implementing dynamic control replication that maintains the sequential semantics of implicitly parallel programs.
Michael Bauer 0001, Wonchan Lee, Elliott Slaughter, Mario Di Renzo, Manolis Papadakis, Galen M. Shipman, Patrick S. McCormick, Michael Garland, Alex Aiken
PPoPP7
2020 Mochi: Composing Data Services for High-Performance Computing Environments
Robert B. Ross, George Amvrosiadis, Philip H. Carns, Chuck Cranor, Matthieu Dorier, Kevin Harms, Gregory R. Ganger, Garth A. Gibson, Samuel K. Gutierrez, Robert Latham, Robert W. Robey, Dana Robinson, Bradley W. Settlemyer, Galen M. Shipman, Shane Snyder, Jérome Soumagne, Qing Zheng
J. Comput. Sci. Technol.14
2018 Programmable Caches with a Data Management Language and Policy Engine
abstract
Our analysis of the key-value activity generated by the ParSplice molecular dynamics simulation demonstrates the need for more complex cache management strategies. Baseline measurements show clear key access patterns and hot spots that offer significant opportunity for optimization. We use the data management language and policy engine from the Mantle system to dynamically explore a variety of techniques, ranging from basic algorithms and heuristics to statistical models, calculus, and machine learning. While Mantle was originally designed for distributed file systems, we show how the collection of abstractions effectively decomposes the problem into manageable policies for a different application and storage system. Our exploration of this space results in a dynamically sized cache policy that does not sacrifice any performance while using 32-66% less memory than the default ParSplice configuration.
Michael Sevilla, Carlos Maltzahn, Peter Alvaro, Reza Nasirigerdeh, Bradley W. Settlemyer, Danny Perez, David Rich, Galen M. Shipman
CCGrid8
2018 Isometry: A Path-Based Distributed Data Transfer System
abstract
Data transfers in parallel systems have a significant impact on the performance of applications. Most existing systems generally support only data transfers between memories with a direct hardware connection and have limited facilities for handling transformations to the data's layout in memory. As a result, to move data between memories that are not directly connected, higher levels of the software stack must explicitly divide a multi-hop transfer into a sequence of single-hop transfers and decide how and where to perform data layout conversions if needed. This approach results in inefficiencies, as the higher levels lack enough information to plan transfers as a whole, while the lower level that does the transfer sees only the individual single-hop requests.
Sean Treichler, Galen M. Shipman, Patrick S. McCormick, Alex Aiken
ICS3
2017 Integrating External Resources with a Task-Based Programming Model
abstract
Accessing external resources (e.g., loading input data, checkpointing snapshots, and out-of-core processing) can have a significant impact on the performance of supercomputer applications. However, no existing programming systems for high-performance computing directly manage and optimize these external accesses. As a result, users must explicitly manage external accesses alongside their computation at the application level, which can result in both correctness and performance issues. We address this limitation by introducing Iris, a task-based programming model with semantics for external resources. Iris allows applications to describe their access requirements to external resources and the relationship of those accesses to the computation. Iris incorporates external I/O into a deferred execution model, reschedules external I/O to overlap I/O with computation, and reduces external I/O when possible. We evaluate Iris on three microbenchmarks representative of important workloads in HPC and a full combustion simulation, S3D. We demonstrate that the Iris implementation of S3D reduces the external I/O overhead by up to 20×, compared to the Legion and the Fortran implementations.
Sean Treichler, Galen M. Shipman, Michael Bauer 0001, Noah Watkins, Carlos Maltzahn, Patrick S. McCormick, Alex Aiken
HiPC3
2017 Control replication: compiling implicit parallelism to efficient SPMD with logical regions
abstract
We present control replication, a technique for generating high-performance and scalable SPMD code from implicitly parallel programs. In contrast to traditional parallel programming models that require the programmer to explicitly manage threads and the communication and synchronization between them, implicitly parallel programs have sequential execution semantics and naturally avoid the pitfalls of explicitly parallel code. However, without optimizations to distribute control overhead, scalability is often poor.
Elliott Slaughter, Wonchan Lee, Sean Treichler, Michael Bauer 0001, Galen M. Shipman, Patrick S. McCormick, Alex Aiken
SC6
2017 A Distributed Multi-GPU System for Fast Graph Processing
abstract
We present Lux, a distributed multi-GPU system that achieves fast graph processing by exploiting the aggregate memory bandwidth of multiple GPUs and taking advantage of locality in the memory hierarchy of multi-GPU clusters. Lux provides two execution models that optimize algorithmic efficiency and enable important GPU optimizations, respectively. Lux also uses a novel dynamic load balancing strategy that is cheap and achieves good load balance across GPUs. In addition, we present a performance model that quantitatively predicts the execution times and automatically selects the runtime configurations for Lux applications. Experiments show that Lux achieves up to 20X speedup over state-of-the-art shared memory systems and up to two orders of magnitude speedup over distributed systems.
Yongkee Kwon, Galen M. Shipman, Patrick S. McCormick, Mattan Erez, Alex Aiken
Proc. VLDB Endow.3
2017 Optimizing End-to-End Big Data Transfers over Terabits Network Infrastructure
abstract
While future terabit networks hold the promise of significantly improving big-data motion among geographically distributed data centers, significant challenges must be overcome even on today's 100 gigabit networks to realize end-to-end performance. Multiple bottlenecks exist along the end-to-end path from source to sink, for instance, the data storage infrastructure at both the source and sink and its interplay with the wide-area network are increasingly the bottleneck to achieving high performance. In this paper, we identify the issues that lead to congestion on the path of an end-to-end data transfer in the terabit network environment, and we present a new bulk data movement framework for terabit networks, called LADS. LADS exploits the underlying storage layout at each endpoint to maximize throughput without negatively impacting the performance of shared storage resources for other users. LADS also uses the Common Communication Interface (CCI) in lieu of the sockets interface to benefit from hardware-level zero-copy, and operating system bypass capabilities when available. It can further improve data transfer performance under congestion on the end systems using buffering at the source using flash storage. With our evaluations, we show that LADS can avoid congested storage elements within the shared storage resource, improving input/output bandwidth, and data transfer rates across the high speed networks. We also investigate the performance degradation problems of LADS due to I/O contention on the parallel file system (PFS), when multiple LADS tools share the PFS. We design and evaluate a meta-scheduler to coordinate multiple I/O streams while sharing the PFS, to minimize the I/O contention on the PFS. With our evaluations, we observe that LADS with meta-scheduling can further improve the performance by up to 14 percent relative to LADS without meta-scheduling.
Youngjae Kim 0001, Scott Atchley, Geoffroy Vallée, Sankeun Lee 0001, Galen M. Shipman
IEEE Trans. Parallel Distributed Syst.5
2015 LADS: Optimizing Data Transfers Using Layout-Aware Data Scheduling
Youngjae Kim 0001, Scott Atchley, Geoffroy Vallée, Galen M. Shipman
FAST4
2014 Web-based visual analytics for extreme scale climate science
abstract
In this paper, we introduce a Web-based visual analytics framework for democratizing advanced visualization and analysis capabilities pertinent to large-scale earth system simulations. We address significant limitations of present climate data analysis tools such as tightly coupled dependencies, inefficient data movements, complex user interfaces, and static visualizations. Our Web-based visual analytics framework removes critical barriers to the widespread accessibility and adoption of advanced scientific techniques. Using distributed connections to back-end diagnostics, we minimize data movements and leverage HPC platforms. We also mitigate system dependency issues by employing a RESTful interface. Our framework embraces the visual analytics paradigm via new visual navigation techniques for hierarchical parameter spaces, multi-scale representations, and interactive spatio-temporal data mining methods that retain details. Although generalizable to other science domains, the current work focuses on improving exploratory analysis of large-scale Community Land Model (CLM) and Community Atmosphere Model (CAM) simulations.
Chad A. Steed, Katherine J. Evans, John Harney, Brian C. Jewell, Galen M. Shipman, Brian E. Smith, Peter E. Thornton, Dean N. Williams
IEEE BigData5
2014 Department of energy strategic roadmap for Earth system science data integration
abstract
The U.S. Department of Energy (DOE) Office of Biological and Environmental Research (BER) Climate and Environmental Sciences Division (CESD) produces a diversity of data, information, software, and model codes across its research and informatics programs and facilities. This information includes raw and reduced observational and instrumentation data, model codes, model-generated results, and integrated data products. Currently, most of these data and information are prepared and shared for program specific activities, corresponding to CESD organization research. A major challenge facing BER CESD is how best to inventory, integrate, and deliver these vast and diverse resources for the purpose of accelerating Earth system science research. This paper provides a concept for a CESD Integrated Data Ecosystem and an initial roadmap for its implementation to address this integration challenge in the “Big Data” domain.
Dean N. Williams, Giriprakash Palanisamy, Galen M. Shipman, Thomas A. Boden, Jimmy W. Voyles
IEEE BigData3
2014 Accelerating Data Acquisition, Reduction, and Analysis at the Spallation Neutron Source
abstract
ORNL operates the world's brightest neutron source, the Spallation Neutron Source (SNS). Funded by the US DOE Office of Basic Energy Science, this national user facility hosts hundreds of scientists from around the world, providing a platform to enable break-through research in materials science, sustainable energy, and basic science. While the SNS provides scientists with advanced experimental instruments, the deluge of data generated from these instruments represents both a big data challenge and a big data opportunity. For example, instruments at the SNS can now generate multiple millions of neutron events per second providing unprecedented experiment fidelity but leaving the user with a dataset that cannot be processed and analyzed in a timely fashion using legacy techniques. To address this big data challenge, ORNL has developed a near real-time streaming data reduction and analysis infrastructure. The Accelerating Data Acquisition, Reduction, and Analysis (ADARA) system provides a live streaming data infrastructure based on a high-performance publish subscribe system, in situ data reduction, visualization, and analysis tools, and integration with a high-performance computing and data storage infrastructure. ADARA allows users of the SNS instruments to analyze their experiment as it is run and make changes to the experiment in real-time and visualize the results of these changes. In this paper we describe ADARA, provide a high-level architectural overview of the system, and present a set of use-cases and real-world demonstrations of the technology.
Galen M. Shipman, Stuart I. Campbell, David Dillow, Mathieu Doucet, Jim Kohl, Garrett E. Granroth, Ross G. Miller, Dale Stansberry, Thomas Proffen, Russel Taylor
eScience1
2014 Best Practices and Lessons Learned from Deploying and Operating Large-Scale Data-Centric Parallel File Systems
abstract
The Oak Ridge Leadership Computing Facility (OLCF) has deployed multiple large-scale parallel file systems (PFS) to support its operations. During this process, OLCF acquired significant expertise in large-scale storage system design, file system software development, technology evaluation, benchmarking, procurement, deployment, and operational practices. Based on the lessons learned from each new PFS deployment, OLCF improved its operating procedures, and strategies. This paper provides an account of our experience and lessons learned in acquiring, deploying, and operating large-scale parallel file systems. We believe that these lessons will be useful to the wider HPC community.
Sarp Oral, James Simmons, Jason Hill, Dustin Leverman, Feiyi Wang, Matthew Ezell, Ross G. Miller, Douglas Fuller, Raghul Gunasekaran, Youngjae Kim 0001, Saurabh Gupta 0002, Devesh Tiwari, Sudharshan S. Vazhkudai, James H. Rogers, David Dillow, Galen M. Shipman, Arthur S. Bland
SC16
2014 The Earth System Grid Federation: An open infrastructure for access to distributed geospatial data
Luca Cinquini, Daniel J. Crichton, Chris Mattmann, John Harney, Galen M. Shipman, Feiyi Wang, Rachana Ananthakrishnan, Neill Miller, Sebastien Denvil, Mark Morgan, Zed Pobre, Gavin M. Bell, Charles M. Doutriaux, Bob Drach, Dean N. Williams, Philip Kershaw, Stephen Pascoe, Estanislao Gonzalez, Sandro Fiore, Roland Schweitzer
Future Gener. Comput. Syst.5
2014 Coordinating Garbage Collectionfor Arrays of Solid-State Drives
abstract
Although solid-state drives (SSDs) offer significant performance improvements over hard disk drives (HDDs) for a number of workloads, they can exhibit substantial variance in request latency and throughput as a result of garbage collection (GC). When GC conflicts with an I/O stream, the stream can make no forward progress until the GC cycle completes. GC cycles are scheduled by logic internal to the SSD based on several factors such as the pattern, frequency, and volume of write requests. When SSDs are used in a RAID with currently available technology, the lack of coordination of the SSD-local GC cycles amplifies this performance variance. We propose a global garbage collection (GGC) mechanism to improve response times and reduce performance variability for a RAID of SSDs. We include a high-level design of SSD-aware RAID controller and GGC-capable SSD devices and algorithms to coordinate the GGC cycles. We develop reactive and proactive GC coordination algorithms and evaluate their I/O performance and block erase counts for various workloads. Our simulations show that GC coordination by a reactive scheme improves average response time and reduces performance variability for a wide variety of enterprise workloads. For bursty, write-dominated workloads, response time was improved by 69 percent and performance variability was reduced by 71 percent. We show that a proactive GC coordination algorithm can further improve the I/O response times by up to 9 percent and the performance variability by up to 15 percent. We also observe that it could increase the lifetimes of SSDs with some workloads (e.g., Financial) by reducing the number of block erase counts by up to 79 percent relative to a reactive algorithm for write-dominant enterprise workloads.
Youngjae Kim 0001, Junghee Lee 0004, Sarp Oral, David Dillow, Feiyi Wang, Galen M. Shipman
IEEE Trans. Computers6
2013 Layout-aware I/O Scheduling for terabits data movement
abstract
Many science facilities, such as the Department of Energy's Leadership Computing Facilities and experimental facilities including the Spallation Neutron Source, Stanford Linear Accelerator Center, and Advanced Photon Source, produce massive amounts of experimental and simulation data. These data are often shared among the facilities and with collaborating institutions. Moving large datasets over the wide-area network (WAN) is a major problem inhibiting collaboration. Next-generation, terabit-networks will help alleviate the problem, however, the parallel storage systems on the endsystem hosts at these institutions can become a bottleneck for terabit data movement. The parallel storage system (PFS) is shared by simulation systems, experimental systems, analysis and visualization clusters, in addition to wide-area data movers. These competing uses often induce temporary, but significant, I/O load imbalances on the storage system, which impact the performance of all the users. The problem is a serious concern because some resources are more expensive (e.g. super computers) or have time-critical deadlines (e.g. experimental data from a light source), but parallel file systems handle all requests fairly even if some storage servers are under heavy load. This paper investigates the problem of competing workloads accessing the parallel file system and how the performance of wide-area data movement can be improved in these environments. First, we study the I/O load imbalance problems using actual I/O performance data collected from the Spider storage system at the Oak Ridge Leadership Computing Facility. Second, we present I/O optimization solutions with layout-awareness on end-system hosts for bulk data movement. With our evaluation, we show that our I/O optimization techniques can avoid the I/O congested disk groups, improving storage I/O times on parallel storage systems for terabit data movement.
Youngjae Kim 0001, Scott Atchley, Geoffroy Vallée, Galen M. Shipman
IEEE BigData4
2013 Preemptible I/O Scheduling of Garbage Collection for Solid State Drives
abstract
Unlike hard disks, flash devices use out-of-place updates operations and require a garbage collection (GC) process to reclaim invalid pages to create free blocks. This GC process is a major cause of performance degradation when running concurrently with other I/O operations as internal bandwidth is consumed to reclaim these invalid pages. The invocation of the GC process is generally governed by a low watermark on free blocks and other internal device metrics that different workloads meet at different intervals. This results in an I/O performance that is highly dependent on workload characteristics. In this paper, we examine the GC process and propose a semipreemptible GC (PGC) scheme that allows GC processing to be preempted while pending I/O requests in the queue are serviced. Moreover, we further enhance flash performance by pipelining internal GC operations and merge them with pending I/O requests whenever possible. Our experimental evaluation of this semi-PGC scheme with realistic workloads demonstrates both improved performance and reduced performance variability. Write-dominant workloads show up to a 66.56% improvement in average response time with a 83.30% reduced variance in response time compared to the non-PGC scheme. In addition, we explore opportunities of a new NAND flash device that supports suspend/resume commands for read, write, and erase operations for fully PGC (F-PGC). Our experiments with an F-PGC enabled flash device show that request response time can be improved by up to 14.57% compared to semi-PGC.
Junghee Lee 0004, Youngjae Kim 0001, Galen M. Shipman, Sarp Oral, Jongman Kim
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2012 The Earth System Grid Federation: An open infrastructure for access to distributed geospatial data
abstract
The Earth System Grid Federation (ESGF) is a multi-agency, international collaboration that aims at developing the software infrastructure needed to facilitate and empower the study of climate change on a global scale. The ESGF's architecture employs a system of geographically distributed peer nodes, which are independently administered yet united by the adoption of common federation protocols and application programming interfaces (APIs). The cornerstones of its interoperability are the peer-to-peer messaging that is continuously exchanged among all nodes in the federation; a shared architecture and API for search and discovery; and a security infrastructure based on industry standards (OpenID, SSL, GSI and SAML). The ESGF software is developed collaboratively across institutional boundaries and made available to the community as open source. It has now been adopted by multiple Earth science projects and allows access to petabytes of geophysical data, including the entire model output used for the next international assessment report on climate change (IPCC-AR5) and a suite of satellite observations (obs4MIPs) and reanalysis data sets (ANA4MIPs).
Luca Cinquini, Daniel J. Crichton, Chris Mattmann, John Harney, Galen M. Shipman, Feiyi Wang, Rachana Ananthakrishnan, Neill Miller, Sebastien Denvil, Mark Morgan, Zed Pobre, Gavin M. Bell, Bob Drach, Dean N. Williams, Philip Kershaw, Stephen Pascoe, Estanislao Gonzalez, Sandro Fiore, Roland Schweitzer
eScience5
2012 Active Flash: Out-of-core data analytics on flash storage
abstract
Next generation science will increasingly come to rely on the ability to perform efficient, on-the-fly analytics of data generated by high-performance computing (HPC) simulations, modeling complex physical phenomena. Scientific computing workflows are stymied by the traditional chaining of simulation and data analysis, creating multiple rounds of redundant reads and writes to the storage system, which grows in cost with the ever-increasing gap between compute and storage speeds in HPC clusters. Recent HPC acquisitions have introduced compute node-local flash storage as a means to alleviate this I/O bottleneck. We propose a novel approach, Active Flash, to expedite data analysis pipelines by migrating to the location of the data, the flash device itself. We argue that Active Flash has the potential to enable true out-of-core data analytics by freeing up both the compute core and the associated main memory. By performing analysis locally, dependence on limited bandwidth to a central storage system is reduced, while allowing this analysis to proceed in parallel with the main application. In addition, offloading work from the host to the more power-efficient controller reduces peak system power usage, which is already in the megawatt range and poses a major barrier to HPC system scalability. We propose an architecture for Active Flash, explore energy and performance trade-offs in moving computation from host to storage, demonstrate the ability of appropriate embedded controllers to perform data analysis and reduction tasks at speeds sufficient for this application, and present a simulation study of Active Flash scheduling policies. These results show the viability of the Active Flash model, and its capability to potentially have a transformative impact on scientific data analysis.
Simona Boboila, Youngjae Kim 0001, Sudharshan S. Vazhkudai, Peter Desnoyers, Galen M. Shipman
MSST5
2012 D-factor: a quantitative model of application slow-down in multi-resource shared systems
abstract
Scheduling multiple jobs onto a platform enhances system utilization by sharing resources. The benefits from higher resource utilization include reduced cost to construct, operate, and maintain a system, which often include energy consumption. Maximizing these benefits, while satisfying performance limits, comes at a price -- resource contention among jobs increases job completion time. In this paper, we analyze slow-downs of jobs due to contention for multiple resources in a system; referred to as dilation factor. We observe that multiple-resource contention creates non-linear dilation factors of jobs. From this observation, we establish a general quantitative model for dilation factors of jobs in multi-resource systems. A job is characterized by a vector-valued loading statistics and dilation factors of a job set are given by a quadratic function of their loading vectors. We demonstrate how to systematically characterize a job, maintain the data structure to calculate the dilation factor (loading matrix), and calculate the dilation factor of each job. We validated the accuracy of the model with multiple processes running on a native Linux server, virtualized servers, and with multiple MapReduce workloads co-scheduled in a cluster. Evaluation with measured data shows that the D-factor model has an error margin of less than 16%. We also show that the model can be integrated with an existing on-line scheduler to minimize the makespan of workloads.
Seung-Hwan Lim, Jae-Seok Huh, Youngjae Kim 0001, Galen M. Shipman, Chita R. Das
SIGMETRICS4
2011 Enhancing I/O throughput via efficient routing and placement for large-scale parallel file systems
abstract
As storage systems get larger to meet the demands of petascale systems, careful planning must be applied to avoid congestion points and extract the maximum performance. In addition, the large data sets generated by such systems makes it desirable for all compute resources to have common access to this data without needing to copy it to each machine. This paper describes a method of placing I/O close to the storage nodes to minimize contention on Cray's SeaStar2+ network, and extends it to a routed Lustre configuration to gain the same benefits when running against a center-wide file system. Our experiments using half of the resources of Spider - the center-wide file system at the Oak Ridge Leadership Computing Facility - show that I/O write bandwidth can be improved by up to 45% (from 71.9 to 104 GB/s) for a direct-attached configuration and by 137% (47.6 GB/s to 115 GB/s) for a routed configuration. We demonstrated up to 20.7% reduction in run-time for production scientific applications. With the full Spider system, we demonstrated over 240 GB/s of aggregate bandwidth using our techniques.
David Dillow, Galen M. Shipman, Sarp Oral, Youngjae Kim 0001
IPCCC2
2011 A semi-preemptive garbage collector for solid state drives
abstract
NAND flash memory is a preferred storage media for various platforms ranging from embedded systems to enterprise-scale systems. Flash devices do not have any mechanical moving parts and provide low-latency access. They also require less power compared to rotating media. Unlike hard disks, flash devices use out-of-update operations and they require a garbage collection (GC) process to reclaim invalid pages to create free blocks. This GC process is a major cause of performance degradation when running concurrently with other I/O operations as internal bandwidth is consumed to reclaim these invalid pages. The invocation of the GC process is generally governed by a low watermark on free blocks and other internal device metrics that different workloads meet at different intervals. This results in I/O performance that is highly dependent on workload characteristics. In this paper, we examine the GC process and propose a semi-preemptive GC scheme that can preempt on-going GC processing and service pending I/O requests in the queue. Moreover, we further enhance flash performance by pipelining internal GC operations and merge them with pending I/O requests whenever possible. Our experimental evaluation of this semi-preemptive GC sheme with realistic workloads demonstrate both improved performance and reduced performance variability. Write-dominant workloads show up to a 66.56% improvement in average response time with a 83.30% reduced variance in response time compared to the non-preemptive GC scheme.
Junghee Lee 0004, Youngjae Kim 0001, Galen M. Shipman, Sarp Oral, Feiyi Wang, Jongman Kim
ISPASS3
2011 Harmonia: A globally coordinated garbage collector for arrays of Solid-State Drives
abstract
Solid-State Drives (SSDs) offer significant performance improvements over hard disk drives (HDD) on a number of workloads. The frequency of garbage collection (GC) activity is directly correlated with the pattern, frequency, and volume of write requests, and scheduling of GC is controlled by logic internal to the SSD. SSDs can exhibit significant performance degradations when garbage collection (GC) conflicts with an ongoing I/O request stream. When using SSDs in a RAID array, the lack of coordination of the local GC processes amplifies these performance degradations. No RAID controller or SSD available today has the technology to overcome this limitation. This paper presents Harmonia, a Global Garbage Collection (GGC) mechanism to improve response times and reduce performance variability for a RAID array of SSDs. Our proposal includes a high-level design of SSD-aware RAID controller and GGC-capable SSD devices, as well as algorithms to coordinate the global GC cycles. Our simulations show that this design improves response time and reduces performance variability for a wide variety of enterprise workloads. For bursty, write dominant workloads response time was improved by 69% while performance variability was reduced by 71%.
Youngjae Kim 0001, Sarp Oral, Galen M. Shipman, Junghee Lee 0004, David Dillow, Feiyi Wang
MSST3
2010 Efficient Object Storage Journaling in a Distributed Parallel File System
Sarp Oral, Feiyi Wang, David Dillow, Galen M. Shipman, Ross G. Miller, Oleg Drokin
FAST4
2010 Functional Partitioning to Optimize End-to-End Performance on Many-core Architectures
abstract
Scaling computations on emerging massive-core supercomputers is a daunting task, which coupled with the significantly lagging system I/O capabilities exacerbates applications' end-to-end performance. The I/O bottleneck often negates potential performance benefits of assigning additional compute cores to an application. In this paper, we address this issue via a novel functional partitioning (FP) runtime environment that allocates cores to specific application tasks - checkpointing, de-duplication, and scientific data format transformation - so that the deluge of cores can be brought to bear on the entire gamut of application activities. The focus is on utilizing the extra cores to support HPC application I/O activities and also leverage solid-state disks in this context. For example, our evaluation shows that dedicating 1 core on an oct-core machine for checkpointing and its assist tasks using FP can improve overall execution time of a FLASH benchmark on 80 and 160 cores by 43.95% and 41.34%, respectively.
Sudharshan S. Vazhkudai, Ali Raza Butt, Xiaosong Ma, Youngjae Kim 0001, Christian Engelmann, Galen M. Shipman
SC8
2007 Network Fault Tolerance in Open MPI
Galen M. Shipman, Richard L. Graham, George Bosilca
Euro-Par1
2006 Open MPI: A High-Performance, Heterogeneous MPI
abstract
The growth in the number of generally available, distributed, heterogeneous computing systems places increasing importance on the development of user-friendly tools that enable application developers to efficiently use these resources. Open MPI provides support for several aspects of heterogeneity within a single, open-source MPI implementation. Through careful abstractions, heterogeneous support maintains efficient use of uniform computational platforms. We describe Open MPI's architecture for heterogeneous network and processor support. A key design features of this implementation is the transparency to the application developer while maintaining very high levels of performance. This is demonstrated with the results of several numerical experiments
Richard L. Graham, Galen M. Shipman, Brian W. Barrett, Ralph H. Castain, George Bosilca, Andrew Lumsdaine
CLUSTER2
2006 Infiniband scalability in Open MPI
abstract
Infiniband is becoming an important interconnect technology in high performance computing. Efforts in large scale Infiniband deployments are raising scalability questions in the HPC community. Open MPI, a new open source implementation of the MPI standard targeted for production computing, provides several mechanisms to enhance Infiniband scalability. Initial comparisons with MVAPICH, the most widely used Infiniband MPI implementation, show similar performance but with much better scalability characteristics. Specifically, small message latency is improved by up to 10% in medium/large jobs and memory usage per host is reduced by as much as 300%. In addition, Open MPI provides predictable latency that is close to optimal without sacrificing bandwidth performance.
Galen M. Shipman, Timothy S. Woodall, Richard L. Graham, Arthur B. Maccabe, Patrick G. Bridges
IPDPS1