Christopher S. Daley

dblp:15/9309 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
3since 2021 · last 2024
0000-0003-3105-0804ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2024 Evaluating the potential of disaggregated memory systems for HPC applications
abstract
Summary Disaggregated memory is a promising approach that addresses the limitations of traditional memory architectures by enabling memory to be decoupled from compute nodes and shared across a data center. Cloud platforms have deployed such systems to improve overall system memory utilization, but performance can vary across workloads. High‐performance computing (HPC) is crucial in scientific and engineering applications, where HPC machines also face the issue of underutilized memory. As a result, improving system memory utilization while understanding workload performance is essential for HPC operators. Therefore, learning the potential of a disaggregated memory system before deployment is a critical step. This paper proposes a methodology for exploring the design space of a disaggregated memory system. It incorporates key metrics that affect performance on disaggregated memory systems: memory capacity, local and remote memory access ratio, injection bandwidth, and bisection bandwidth, providing an intuitive approach to guide machine configurations based on technology trends and workload characteristics. We apply our methodology to analyze thirteen diverse workloads, including AI training, data analysis, genomics, protein, fusion, atomic nuclei, and traditional HPC bookends. Our methodology demonstrates the ability to comprehend the potential and pitfalls of a disaggregated memory system and provides motivation for machine configurations. Our results show that eleven of our thirteen applications can leverage injection bandwidth disaggregated memory without affecting performance, while one pays a rack bisection bandwidth penalty and two pay the system‐wide bisection bandwidth penalty. In addition, we also show that intra‐rack memory disaggregation would meet the application's memory requirement and provide enough remote memory bandwidth.
Nan Ding 0006, Pieter Maris, Hai Ah Nam, Taylor L. Groves, Muaaz Gul Awan, LeAnn Lindsey, Christopher S. Daley, Oguz Selvitopi, Leonid Oliker, Nicholas J. Wright, Samuel Williams 0001
Concurr. Comput. Pract. Exp.7
2022 A Portable Sparse Solver Framework for Large Matrices on Heterogeneous Architectures
abstract
Programming applications on heterogeneous systems with hardware accelerators is challenging due to the disjoint address spaces between the host (CPU) and the device (GPU). The limited device memory further exacerbates the challenges as most data-intensive applications will not fit in the limited device memory. CUDA Unified Memory (UM) was introduced to mitigate such challenges. UM improves GPU programmability by supporting oversubscription, on-demand paging, and migration. However, when the working set of an application exceeds the device memory capacity, the resulting data movement can cause significant performance losses. We propose a tiling-based task-parallel framework, named DeepSparseGPU, to accelerate sparse eigensolvers on GPUs by minimizing data movement between the host and device. To this end, we tile all operations in a sparse solver and express the entire computation as a directed acyclic graph (DAG). We design and develop a memory manager (MM) to execute larger inputs that do not fit into GPU memory. MM keeps track of the data on CPU and GPU, and automatically moves data between them as needed. We use OpenMP target offload in our implementation to achieve portability beyond NVIDIA hardware. Performance evaluations show that DeepSparseGPU transfers 1.39x-2.18x less host to device (H2D) and device to host (D2H) data, while executing up to 2.93x faster than the UM-based baseline version.
Fazlay Rabbi, Christopher S. Daley, Ümit V. Çatalyürek, Hasan Metin Aktulga
HIPC2
2021 Non-recurring engineering (NRE) best practices: a case study with the NERSC/NVIDIA OpenMP contract
abstract
The NERSC supercomputer, Perlmutter, consists of AMD CPUs and NVIDIA GPUs. NERSC users expect to be able to use OpenMP to take advantage of the highly capable GPUs. This paper describes how NERSC/NVIDIA constructed a Non-Recurring Engineering (NRE) contract to add OpenMP GPU-offload support to the NVIDIA HPC compilers. The paper describes how the contract incorporated the strengths of both parties and encouraged collaboration to improve the quality of the final deliverable. We include our best practices and how this particular contract took into account emerging OpenMP specifications, NERSC workload requirements, and how to use OpenMP most efficiently on GPU hardware. This paper includes OpenMP application performance results obtained with the NVIDIA compilers distributed in the NVIDIA HPC SDK.
Christopher S. Daley, Annemarie Southwell, Rahulkumar Gayatri, Scott Biersdorfff, Craig Toepfer, Güray Özen, Nicholas J. Wright
SC1
2020 Experiences in porting mini-applications to OpenACC and OpenMP on heterogeneous systems
abstract
Summary This article studies mini‐applications—Minisweep, GenASiS, GPP, and FF—that use computational methods commonly encountered in HPC. We have ported these applications to develop OpenACC and OpenMP versions, and evaluated their performance on Titan (Cray XK7 with K20x GPUs), Cori (Cray XC40 with Intel KNL), Summit (IBM AC922 with Volta GPUs), and Cori‐GPU (Cray CS‐Storm 500NX with Intel Skylake and Volta GPUs). Our goals are for these new ports to be useful to both application and compiler developers, to document and describe the lessons learned and the methodology to create optimized OpenMP and OpenACC versions, and to provide a description of possible migration paths between the two specifications. Cases where specific directives or code patterns result in improved performance for a given architecture are highlighted. We also include discussions of the functionality and maturity of the latest compilers available on the above platforms with respect to OpenACC or OpenMP implementations.
Verónica G. Vergara Larrea, Reuben D. Budiardja, Rahulkumar Gayatri, Christopher S. Daley, Oscar R. Hernandez, Wayne Joubert
Concurr. Comput. Pract. Exp.4
2020 Performance characterization of scientific workflows for the optimal use of Burst Buffers
Christopher S. Daley, Devarshi Ghoshal, Glenn K. Lockwood, Sudip S. Dosanjh, Lavanya Ramakrishnan, Nicholas J. Wright
Future Gener. Comput. Syst.1
2015 Ongoing verification of a multiphysics community code: FLASH
abstract
SUMMARY When developing a complex, multi‐authored code, daily testing on multiple platforms and under a variety of conditions is essential. It is therefore necessary to have a regression test suite that is easily administered and configured, as well as a way to easily view and interpret the test suite results. We describe the methodology for verification of FLASH, a highly capable multiphysics scientific application code with a wide user base. The methodology uses a combination of unit and regression tests and an in‐house testing software that is optimized for operation under limited resources. Although our practical implementations do not always comply with theoretical regression‐testing research, our methodology provides a comprehensive verification of a large scientific code under resource constraints.Copyright © 2013 John Wiley & Sons, Ltd.
Anshu Dubey, Klaus Weide, Dongwook Lee 0004, John Bachan, Christopher S. Daley, Samuel Olofin, Noel T. Taylor, Paul M. Rich, Lynn B. Reid
Softw. Pract. Exp.5
2013 Parallel Algorithms for Using Lagrangian Markers in Immersed Boundary Method with Adaptive Mesh Refinement in FLASH
abstract
Computational fluid dynamics (CFD) are at the forefront of computational mechanics in requiring large-scale computational resources associated with high performance computing (HPC). Many flows of practical interest also include moving and deforming boundaries. High fidelity computations of fluid-structure interactions (FSI) are amongst the most challenging problems in computational mechanics. Additionally, many FSI applications have different resolution requirements in different parts of the domain and therefore requirement adaptive mesh refinement (AMR) for computational efficiency. FLASH is a well established AMR code with an existing Lagrangian framework which could be augmented and exploited to implement an immersed boundary method for simulating fluid-structure interactions atop an existing infrastructure. This paper describes the augmentations to the Lagrangian framework, and the new parallel algorithms added to the FLASH infrastructure that enabled the implementation of immersed boundary method in FLASH. The paper also presents scaling behavior and performance analysis of the implementations.
Prateeti Mohapatra, Anshu Dubey, Christopher S. Daley, Marcos Vanella, Elias Balaras
SBAC-PAD3
2012 Optimization of multigrid based elliptic solver for large scale simulations in the FLASH code
abstract
SUMMARY FLASH is a multiphysics multiscale adaptive mesh refinement (AMR) code originally designed for simulation of reactive flows often found in Astrophysics. With its wide user base and flexible applications configuration capability, FLASH has a dual task of maintaining scalability and portability in all its solvers. The scalability of fully explicit solvers in the code is tied very closely to that of the underlying mesh. Others such as the Poisson solver based on a multigrid method have more complex scaling behavior. Multigrid methods suffer from processor starvation and dominating communication costs at coarser grids with increase in the number of processors. In this paper, we propose a combination of uniform grid mesh with AMR mesh, and the merger of two different sets of solvers to overcome the scalability limitation of the Poisson solver in FLASH. The principal challenge in the proposed merger is the efficiency of the communication algorithm to map the mesh back and forth between uniform grid and AMR. We present two different parallel mapping algorithms and also discuss results from performance studies of the two implementations. Copyright © 2012 John Wiley & Sons, Ltd.
Christopher S. Daley, Marcos Vanella, Anshu Dubey, Klaus Weide, Elias Balaras
Concurr. Comput. Pract. Exp.1
2011 Parallel algorithms for moving Lagrangian data on block structured Eulerian meshes
Anshu Dubey, Katie Antypas, Christopher S. Daley
Parallel Comput.3