Ryan E. Grant

dblp:51/6006 · also Ryan Eric Grant · DBLP profile ↗
← Back
48ranked-venue papers
10as first author
13since 2021 · last 2025
0000-0002-0163-3892ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 39 · 9 first-author · 9 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2025 NAV: A Comparative Analysis Tool for Nsight Systems GPU Traces
abstract
High-performance computing (HPC) and data centers increasingly rely on Graphics Processing Units (GPUs) in large supercomputers. Yet, this reliance poses challenges in quickly understanding the performance impacts of code changes. This paper introduces NAV, a versatile tool designed to rapidly analyze and compare GPU performance traces. Built upon NVIDIA’s NsightTMSystems (NSYS), NAV enhances NSYS’s capabilities by accelerating visualization for large traces, increasing the variety of data representation formats, and adding comparative analysis capabilities. Our tool, NAV, provides complimentary functions on top of NSYS to quickly access performance data and perform comparative analysis. NAV offers detailed visual and written representations of traces at various granularity levels and efficiently handles large trace files through parallelization. It extracts trace data 1.15 to 3.5 times faster than comparable NSYS recipes for typical developer trace sizes, automating the generation of valuable data representations and streamlining workload analysis. This paper outlines NAV’s key features and functionalities, demonstrating its effectiveness through use cases that highlight its benefits for rapid application analysis and assessing the impact of code changes.
Ethan Shama, Ryan E. Grant
PDP2
2025 Utilizing Network Hardware Parallelism for MPI Partitioned Collective Communication
abstract
Parallel distributed applications running on large-scale high-performance computing systems depend on effective point-to-point and collective communication to meet performance goals. Beginning with version 4.0, the Message Passing Interface (MPI) introduced the partitioned communication API, providing tools for addressing communication bottlenecks raised by hybrid communication models. This API allows individual actors (CPU threads, GPU threads, etc.) to initiate communication on portions of complete buffers, enabling additional communication/computation overlap. Intuitively, the utility of partitioned communication could benefit from network-level support: If there are multiple paths between endpoints, an MPI-aware network could disperse partitions across these paths, avoiding the data serialization entailed by a dependency on a single path. The Cerio Rockport Ethernet Fabric has the ability to expose this capability to communication middleware. In this work we develop this capability to allow for user-level path selection for MPI partitioned communication and explore how this capability impacts point-to-point performance, collective design, and Allreduce efficiency in a Large Language Model task
Yiltan Hassan Temuçin, Amirreza Barati Sedeh, Whit Schonbein, Ryan E. Grant, Ahmad Afsahi
PDP4
2024 Smart Network Traffic Prediction for Scientific Applications
abstract
Network traffic in HPC systems can impact application performance by inducing costly re-transmissions or consuming memory bandwidth. Emerging SmartNIC technologies present new opportunities for addressing these issues and optimizing network performance by providing a platform for the intelligent utilization of network resources through machine learning models of network traffic. However, SmartNICs also present challenges for deploying such models insofar as they offer relatively limited computational and memory resources, and these resources must be shared with other services. Based on an analysis of traffic data collected from eight scientific applications and proxies, we explore lightweight approaches to modeling network traffic using dynamic linear regression. Depending on the application, normalized root mean squared error for static regression may be less than 1 %, and dynamic regression can reduce this error by an order of magnitude. We further refine the dynamic approach by adding an additional classifier that categorizes predictions generated by the model as reliable or unreliable, showing that the technique can achieve good precision and recall. Finally, we evaluate the performance of dynamic regression and classification on NVIDIA BlueField-2 and BlueField-3 SmartNICs, demonstrating these computationally lightweight techniques are feasible on contemporary SmartNIC platforms.
Whit Schonbein, Tinotenda Matsika, Ryan E. Grant
PDP3
2024 Improved MPI Collectives for 3D-FFT
Yuang Yan, Natasha Kuk, Ryan E. Grant
EuroMPI3
2023 A Dynamic Network-Native MPI Partitioned Aggregation Over InfiniBand Verbs
abstract
Modern HPC systems require efficient hybrid programming model to utilize their hardware resources effectively. The Message Passing Interface (MPI) has accommodated next-generation hardware by providing new APIs such as the MPI Partitioned interface. This API provides a user with fine-grain communication without the overhead of traditional MPI point-to-point communication in multi-threaded workloads.To the best of our knowledge, we present the first work on detailed low-level design for an MPI Partitioned implementation. We guide readers through a method to map the MPI Partitioned interface to the InfiniBand Verbs API. Alongside implementation details, we also study the aggregation of user partitions and how we can efficiently send them over the network. We study a brute force approach and using the Partitioned LogGP (PLogGP) model to predict ideal aggregation. We observe that using the PLogGP model provides comparable performance without exhausting computing resources to search the entire solution space. The PLogGP design was further optimized by considering how the partition arrival pattern can be used to dynamically modify our aggregation scheme. We profiled our micro-benchmarks to provide analysis on how and why this additional optimization is beneficial to our results and how we can fine-tune this mechanism. Finally, we evaluated our PLogGP and Timer-based PLogGP designs with a commonly used communication pattern in HPC (communication sweep) to observe the impact when communicating with multiple processes in an application-like scenario at 1024 cores.
Yiltan Hassan Temuçin, Scott Levy, Whit Schonbein, Ryan E. Grant, Ahmad Afsahi
CLUSTER4
2023 Modeling and Benchmarking the Potential Benefit of Early-Bird Transmission in Fine-Grained Communication
abstract
Traditional point-to-point communication sends data only after the entirety of the data is available. This includes situations where multiple actors (e.g., threads) contribute to the send buffer. As a result, cases where the completion times of these actors are widely distributed may be lost opportunities for optimization because data ready to be sent is waiting to be transmitted. Fine-grained communication exposes these opportunities by allowing buffers to be divided into element s that can then be sent independently (see e.g., Partitioned Communication in Message Passing Interface v4.0). While some research has been directed at exploring the utility of such ‘early-bird’ transmission, the overall search space for finding the best performing actor completion timings and element counts is large. In this work, we present an abstract model of fine-grained communication based on the LogGP model and a complementary benchmark. We use the model to explore actor completion timing scenarios and identify trends in communication behavior based on factors such as overall message size and delay between actor completions. We evaluate the benchmarks on three systems utilizing distinct network technologies and show that: (i) smaller numbers of element s are able to exploit most of the benefit of early-bird communication, (ii) performance benefit will depend non-trivially on application behavior, and (iii) benefits are highly network-dependent.
Whit Schonbein, Scott Levy, Matthew G. F. Dosanjh, W. Pepper Marts, Elizabeth Reid 0002, Ryan E. Grant
ICPP6
2023 Enabling power measurement and control on Astra: The first petascale Arm supercomputer
abstract
Summary Astra, deployed in 2018, was the first petascale supercomputer to utilize processors based on the ARM instruction set. The system was also the first under Sandia's Vanguard program which seeks to provide an evaluation vehicle for novel technologies that with refinement could be utilized in demanding, large‐scale HPC environments. In addition to ARM, several other important first‐of‐a‐kind developments were used in the machine, including new approaches to cooling the datacenter and machine. This article documents our experiences building a power measurement and control infrastructure for Astra. While this is often beyond the control of users today, the accurate measurement, cataloging, and evaluation of power, as our experiences show, is critical to the successful deployment of a large‐scale platform. While such systems exist in part for other architectures, Astra required new development to support the novel Marvell ThunderX2 processor used in compute nodes. In addition to documenting the measurement of power during system bring up and for subsequent on‐going routine use, we present results associated with controlling the power usage of the processor, an area which is becoming of progressively greater interest as data centers and supercomputing sites look to improve compute/energy efficiency and find additional sources for full system optimization.
Ryan E. Grant, Simon D. Hammond, James H. Laros III, Michael J. Levenhagen, Stephen Olivier, Kevin T. Pedretti, Lee Ward, Andrew J. Younge
Concurr. Comput. Pract. Exp.1
2023 Design of a portable implementation of partitioned point-to-point communication primitives
abstract
Abstract The Message Passing Interface (MPI) has been the dominant message passing solution for scientific computing for decades. MPI point‐to‐point communications are highly efficient mechanisms for process‐to‐process communication. However, MPI performance when processes utilize multiple threads is slowed by concurrency protections in the MPI library. MPI's current thread level interface imposes these overheads throughout the library when thread safety is needed. While much work has been done to reduce multithreading overheads in MPI, a solution is needed that reduces the number of messages exchanged in a threaded environment. Partitioned communication is included in the MPI 4.0 standard as an alternative that addresses the challenges of multithreaded communication in MPI today. Partitioned communication reduces overall message volume by creating a buffer‐sharing mechanism between threads such that they can indicate when portions of a communication buffer are available to be sent. Separation of the control and data planes in MPI is enabled by allowing persistent initialization and single occurrence message buffer matching from the indication that the data is ready to be sent. This enables the usage of underlying hardware primitives like triggered operations, where commands (destination, size, etc.) can be set up prior to data buffer readiness and readiness triggered with a simple doorbell/counter later. This approach is useful for future development of MPI operations in environments where traditional networking commands can have performance challenges, like accelerators (GPUs, FPGAs). In this paper, we detail the design and implementation of a layered library (built on top of MPI‐3.1) and an integrated Open MPI solution that supports the new, MPI‐4.0 partitioned communication feature set. The library will enable applications to use currently released MPI implementations and older legacy libraries to provide partitioned communication support while also enabling further exploration of this new communication model in new applications and use cases. We will compare the designs of the library and native Open MPI support, provide performance results and comparisons between the two approaches, and lessons learned from the implementation of partitioned communication in both library and native forms. We find that the native implementation and library have similar performance with a percentage difference under 0.94% in microbenchmarks and performance within 5% for a partitioned communication enabled proxy application.
W. Pepper Marts, Andrew Worley, Prema Soundarajan, Derek Schafer, Matthew G. F. Dosanjh, Ryan E. Grant, Purushotham V. Bangalore, Anthony Skjellum, Sheikh K. Ghafoor
Concurr. Comput. Pract. Exp.6
2022 Micro-Benchmarking MPI Partitioned Point-to-Point Communication
abstract
Modern High-Performance Computing (HPC) architectures have developed the need for scalable hybrid programming models. The latest Message Passing Interface (MPI) 4.0 standard has introduced a new communication model: MPI Partitioned Point-to-Point communication. This new model allows for the contribution of data from multiple threads with lower overheads than with traditional MPI point-to-point communication. In this paper, we design the first publicly available micro-benchmark suite for MPI Partitioned to measure various metrics that can give insight into the benefits of using this new model and scenarios where MPI point-to-point is better suited. Suggestions are provided to application developers on how to choose partition size for their application based on compute and message size. We evaluate MPI Partitioned communication with both a hot and cold CPU cache, system noise with different probability distributions, point-to-point communication directly, and with commonly used MPI communication patterns such as a halo exchange and Sweep3D.
Yiltan Hassan Temuçin, Ryan E. Grant, Ahmad Afsahi
ICPP2
2022 "Smarter" NICs for faster molecular dynamics: a case study
abstract
This work evaluates the benefits of using a “smart” network interface card (SmartNIC) as a compute accelerator for the example of the MiniMD molecular dynamics proxy application. The accelerator is NVIDIA's BlueField-2 card, which includes an 8-core Arm processor along with a small amount of DRAM and storage. We test the networking and data movement performance of these cards compared to a standard Intel server host using microbenchmarks and MiniMD. In MiniMD, we identify two distinct classes of computation, namely core computation and maintenance computation, which are executed in sequence. We restructure the algorithm and code to weaken this dependence and increase task parallelism, thereby making it possible to increase utilization of the BlueField-2 concurrently with the host. We evaluate our implementation on a cluster consisting of 16 dual-socket Intel Broadwell host nodes with one BlueField-2 per host-node. Our results show that while the overall compute performance of BlueField-2 is limited, using them with a modified MiniMD algorithm allows for up to 20% speedup over the host CPU baseline with no loss in simulation accuracy.
Sara Karamati, Clay Hughes, Karl S. Hemmert, Ryan E. Grant, Whit Schonbein, Scott Levy, Thomas M. Conte, Jeffrey Young 0001, Richard W. Vuduc
IPDPS4
2021 MiniMod: A Modular Miniapplication Benchmarking Framework for HPC
abstract
The HPC application community has proposed many new application communication structures, middleware interfaces, and communication models to improve HPC application performance. Modifying proxy applications is the standard practice for the evaluation of these novel methodologies. Currently, this requires the creation of a new version of the proxy application for each combination of the approach being tested. In this article, we present a modular proxy-application framework, MiniMod, that enables evaluation of a combination of independently written computation kernels, data transfer logic, communication access, and threading libraries. MiniMod is designed to allow rapid development of individual modules which can be combined at runtime. Through MiniMod, developers only need a single implementation to evaluate application impact under a variety of scenarios.We demonstrate the flexibility of MiniMod’s design by using it to implement versions of a heat diffusion kernel and the miniFE finite element proxy application, along with a variety of communication, granularity, and threading modules. We examine how changing communication libraries, communication granularities, and threading approaches impact these applications on an HPC system. These experiments demonstrate that MiniMod can rapidly improve the ability to assess new middleware techniques for scientific computing applications and next-generation hardware platforms.
W. Pepper Marts, Matthew G. F. Dosanjh, Scott Levy, Whit Schonbein, Ryan E. Grant, Patrick G. Bridges
CLUSTER5
2021 RVMA: Remote Virtual Memory Access
abstract
Remote Direct Memory Access (RDMA) capabilities have been provided by high-end networks for many years, but the network environments surrounding RDMA are evolving. RDMA performance has historically relied on using strict ordering guarantees to determine when data transfers complete, but modern adaptively-routed networks no longer provide those guarantees. RDMA also exposes low-level details about memory buffers: either all clients are required to coordinate access using a single shared buffer, or exclusive resources must be allocatable per-client for an unbounded amount of time. This makes RDMA unattractive for use in many-to-one communication models such as those found in public internet client-server situations. Remote Virtual Memory Access (RVMA) is a novel approach to data transfer which adapts and builds upon RDMA to provide better usability, resource management, and fault tolerance. RVMA provides a lightweight completion notification mechanism which addresses RDMA performance penalties imposed by adaptively-routed networks, enabling high-performance data transfer regardless of message ordering. RVMA also provides receiver-side resource management, abstracting away previously-exposed details from the sender-side and removing the RDMA requirement for exclusive/coordinated resources. RVMA requires only small hardware modifications from current designs, provides performance comparable or superior to traditional RDMA networks, and offers many new features. In this paper, we describe RVMA's receiver-managed resource approach and how it enables a variety of new data-transfer approaches on high-end networks. In particular, we demonstrate how an RVMA NIC could implement the first hardware-based fault tolerant RDMA-like solution. We present the design and validation of an RVMA simulation model in a popular simulation suite and use it to evaluate the advantages of RVMA at large scale. In addition to support for adaptive routing and easy programmability, RVMA can outperform RDMA on a 3D sweep application by 4.4X.
Ryan E. Grant, Michael J. Levenhagen, Matthew G. F. Dosanjh, Patrick M. Widener
IPDPS1
2021 Implementation and evaluation of MPI 4.0 partitioned communication libraries
Matthew G. F. Dosanjh, Andrew Worley, Derek Schafer, Prema Soundararajan, Sheikh K. Ghafoor, Anthony Skjellum, Purushotham V. Bangalore, Ryan E. Grant
Parallel Comput.8
2020 A survey of MPI usage in the US exascale computing project
abstract
Summary The Exascale Computing Project (ECP) is currently the primary effort in the United States focused on developing “exascale” levels of computing capabilities, including hardware, software, and applications. In order to obtain a more thorough understanding of how the software projects under the ECP are using, and planning to use the Message Passing Interface (MPI), and help guide the work of our own project within the ECP, we created a survey. Of the 97 ECP projects active at the time the survey was distributed, we received 77 responses, 56 of which reported that their projects were using MPI. This paper reports the results of that survey for the benefit of the broader community of MPI developers.
David E. Bernholdt, Swen Böhm, George Bosilca, Manjunath Gorentla Venkata, Ryan E. Grant, Thomas J. Naughton, Howard Pritchard, Martin Schulz 0001, Geoffroy Vallée
Concurr. Comput. Pract. Exp.5
2020 Tail queues: A multi-threaded matching architecture
abstract
Summary As we approach exascale, computational parallelism will have to drastically increase in order to meet throughput targets. Many‐core architectures have exacerbated this problem by trading reduced clock speeds, core complexity, and computation throughput for increasing parallelism. This presents two major challenges for communication libraries such as MPI: the library must leverage the performance advantages of thread level parallelism and avoid the scalability problems associated with increasing the number of processes to that scale. Hybrid programming models, such as MPI+X, have been proposed to address these challenges. MPI THREAD MULTIPLE is MPI's thread safe mode. While there has been work to optimize it, it largely remains non‐performant in most implementations. While current applications avoid MPI multithreading due to performance concerns, it is expected to be utilized in future applications. One of the major synchronous data structures required by MPI is the matching engine. In this paper, we present a parallel matching algorithm that can improve MPI matching for multithreaded applications. We then perform a feasibility study to demonstrate the performance benefit of the technique.
Matthew G. F. Dosanjh, Ryan E. Grant, Whit Schonbein, Patrick G. Bridges
Concurr. Comput. Pract. Exp.2
2020 Hardware MPI message matching: Insights into MPI matching behavior to inform design
abstract
Summary This paper explores key differences of MPI match lists for several important United States Department of Energy (DOE) applications and proxy applications. This understanding is critical in determining the most promising hardware matching design for any given high‐speed network. The results of MPI match list studies for the major open‐source MPI implementations, MPICH and Open MPI, are presented, and we modify an MPI simulator, LogGOPSim, to provide match list statistics. These results are discussed in the context of several different potential design approaches to MPI matching–capable hardware. The data illustrate the requirements for different hardware designs in terms of performance and memory capacity. This paper's contributions are the collection and analysis of data to help inform hardware designers of common MPI requirements and highlight the difficulties in determining these requirements by only examining a single MPI implementation.
Kurt B. Ferreira, Ryan E. Grant, Michael J. Levenhagen, Scott Levy, Taylor L. Groves
Concurr. Comput. Pract. Exp.2
2020 Foreword to the Special Issue of the Workshop on Exascale MPI (ExaMPI 2017)
abstract
The aim of the Workshop on Exascale MPI (ExaMPI 2017), held in conjunction with SC17: The International Conference for High Performance Computing, Networking, Storage and Analysis, was to bring together researchers and developers to present and discuss innovative algorithms and concepts in the Message Passing programming model and to create a forum for open and potentially controversial discussions on the future of MPI in the Exascale era. This special issue includes selected papers from this workshop that include innovative algorithms for collective operations, extensions to MPI, including datacentric models, scheduling/routing to avoid network congestion, fault-tolerant communication, interoperability of MPI and PGAS models, and use of MPI in large-scale simulations. The first paper, titled “A survey of MPI usage in the US Exascale Computing Project,” provides an analysis of the survey that was conducted to understand how MPI is currently used and intended to be used by different applications that are part of the Exascale Computing Project (ECP).1 The results of analysis provide specific recommendations for MPI implementors, tool developers, and the MPI Forum. The next three papers focus on the issue of message matching in MPI and provide different options to address message matching at exascale. The paper titled “Tail Queues: A Multi-threaded Matching Architecture” introduces a novel parallel matching architecture and prototype implementation based on MPICH to improve the performance of message matching.2 The paper titled “Communication-Aware Message Matching in MPI” uses a novel message queue architecture that allocates dedicated message queues based on the frequency of communication between various processes to reduce the queue search time and also reduces memory consumption.3 These performance improvements result in a speedup of 5 times on the FDS application. The next paper, titled “Hardware MPI Message Matching: Insights into MPI Matching Behavior to Inform Design,” explores what hardware features are needed to support efficient message matching through the evaluation of message matching characteristics of major MPI implementations.4 The next two papers consider support for fault tolerance. The first paper, titled “EReinit: Scalable and Efficient Fault-Tolerance for Bulk-Synchronous MPI Applications,” describes a global-restart model to improve the recovery time of applications when dealing with faults.5 The second paper, titled “The Unexpected Virtue of Almost: Exploiting MPI Collective Operations to Approximately Coordinate Checkpoints,” describes an uncoordinated checkpointing mechanism that makes use of collective operations already used in an application to force checkpoints.6 The paper titled “Optimizing Point-to-Point Communication between Adaptive MPI Endpoints in Shared Memory” describes an approach to optimize point-to-point communication in a shared memory environment and hence improve MPI multithreading support.7 The paper titled “On the Memory Attribution Problem: A Solution and Case Study Using MPI” describes a solution to capture and analyze memory usage by an MPI application and the MPI library.8 Lastly, the paper titled “Twister2: Design of a Big Data Toolkit” describes an architecture to support different types of data-intensive applications in a unified framework.9 Overall, these nine papers contribute to the knowledge base and advancement of the Message Passing Interface in diverse and useful ways. While illustrating the staying power of MPI after a quarter century, they also point to areas of opportunity for enhancement at Exascale as well as to new potential application areas.
Anthony Skjellum, Purushotham V. Bangalore, Ryan E. Grant
Concurr. Comput. Pract. Exp.3
2019 Fuzzy Matching: Hardware Accelerated MPI Communication Middleware
abstract
Contemporary parallel scientific codes often rely on message passing for inter-process communication. However, inefficient coding practices or multithreading (e.g., via MPI_THREAD_MULTIPLE) can severely stress the underlying message processing infrastructure, resulting in potentially un-acceptable impacts on application performance. In this article, we propose and evaluate a novel method for addressing this issue: 'Fuzzy Matching'. This approach has two components. First, it exploits the fact most server-class CPUs include vector operations to parallelize message matching. Second, based on a survey of point-to-point communication patterns in representative scientific applications, the method further increases parallelization by allowing matches based on 'partial truth', i.e., by identifying probable rather than exact matches. We evaluate the impact of this approach on memory usage and performance on Knight's Landing and Skylake processors. At scale (262,144 Intel Xeon Phi cores), the method shows up to 1.13 GiB of memory savings per node in the MPI library, and improvement in matching time of 95.9%; smaller-scale runs show run-time improvements of up to 31.0% for full applications, and up to 6.1% for optimized proxy applications.
Matthew G. F. Dosanjh, Whit Schonbein, Ryan E. Grant, Patrick G. Bridges, S. Mahdieh Ghazimirsaeed, Ahmad Afsahi
CCGRID3
2019 MPI tag matching performance on ConnectX and ARM
abstract
As we approach Exascale, message matching has increasingly become a significant factor in HPC application performance. To address this, network vendors have placed higher precedence on improving MPI message matching performance. ConnectX-5, Mellanox's new network interface card, has both hardware and software matching layers. The performance characteristics of these layers have yet to be studied under real world circumstances. In this work we offer an initial evaluation of ConnectX-5 message matching performance. To analyze this new hardware we executed a series of micro-benchmarks and applications on Astra, an ARM-based ConnectX-5 HPC system, while varying hardware and software matching parameters. The benchmark results show the ConnectX-5 is sensitive to queue depths, and that hardware message matching increases performance for applications that send messages between 1KiB and 16KiB. Furthermore, the hardware matching system was capable of matching wildcard receives without negatively impacting performance. Finally, for some applications, a significant improvement can be observed when leveraging the ConnectX-5's hardware matching.
W. Pepper Marts, Matthew G. F. Dosanjh, Whit Schonbein, Ryan E. Grant, Patrick G. Bridges
EuroMPI4
2019 INCA: in-network compute assistance
abstract
Current proposals for in-network data processing operate on data as it streams through a network switch or endpoint. Since compute resources must be available when data arrives, these approaches provide deadline-based models of execution. This paper introduces a deadline-free general compute model for network endpoints called INCA: In-Network Compute Assistance. INCA builds upon contemporary NIC offload capabilities to provide on-NIC, deadline-free, general-purpose compute capacities that can be utilized when the network is inactive. We demonstrate INCA is Turing complete, and provide a detailed design for extending existing hardware to support this model. We evaluate runtimes for a selection of kernels, including several optimizations, and show INCA can provide up to a 11% speedup for applications with minimal code modifications and between 25% to 37% when applications are optimized for INCA.
Whit Schonbein, Ryan E. Grant, Matthew G. F. Dosanjh, Dorian C. Arnold
SC2
2019 A dynamic, unified design for dedicated message matching engines for collective and point-to-point communications
S. Mahdieh Ghazimirsaeed, Ryan E. Grant, Ahmad Afsahi
Parallel Comput.2
2019 Using simulation to examine the effect of MPI message matching costs on application performance
Scott Levy, Kurt B. Ferreira, Whit Schonbein, Ryan E. Grant, Matthew G. F. Dosanjh
Parallel Comput.4
2018 Measuring Multithreaded Message Matching Misery
Whit Schonbein, Matthew G. F. Dosanjh, Ryan E. Grant, Patrick G. Bridges
Euro-Par3
2018 The Case for Semi-Permanent Cache Occupancy: Understanding the Impact of Data Locality on Network Processing
abstract
The performance critical path for MPI implementations relies on fast receive side operation, which in turn requires fast list traversal. The performance of list traversal is dependent on data-locality; whether the data is currently contained in a close-to-core cache due to its temporal locality or if its spacial locality allows for predictable pre-fetching. In this paper, we explore the effects of data locality on the MPI matching problem by examining both forms of locality. First, we explore spacial locality, by combining multiple entries into a single linked list element, we can control and modify this form of locality. Secondly, we explore temporal locality by utilizing a new technique called "hot caching", a process that creates a thread to periodically access certain data, increasing its temporal locality. In this paper, we show that by increasing data locality, we can improve MPI performance on a variety of architectures up to 4x for micro-benchmarks and up to 2x for an application.
Matthew G. F. Dosanjh, S. Mahdieh Ghazimirsaeed, Ryan E. Grant, Whit Schonbein, Michael J. Levenhagen, Patrick G. Bridges, Ahmad Afsahi
ICPP3
2018 Improving MPI Multi-threaded RMA Communication Performance
abstract
One-sided communication is crucial to enabling communication concurrency. As core counts have increased, particularly with many-core architectures, one-sided (RMA) communication has been proposed to address the ever increasing contention at the network interface. The difficulty in using one-sided (RMA) communication with MPI is that the performance of MPI implementations using RMA with multiple concurrent threads is not well understood. Past studies have been done using MPI RMA in combination with multi-threading (RMA-MT) but they have been performed on older MPI implementations lacking RMA-MT optimizations. In addition prior work has only been done at smaller scale (<=512 cores).
Nathan T. Hjelm, Matthew G. F. Dosanjh, Ryan E. Grant, Taylor L. Groves, Patrick G. Bridges, Dorian C. Arnold
ICPP3
2018 Characterizing MPI matching via trace-based simulation
Kurt B. Ferreira, Scott Levy, Kevin T. Pedretti, Ryan E. Grant
Parallel Comput.4
2018 Unraveling Network-Induced Memory Contention: Deeper Insights with Machine Learning
abstract
Remote Direct Memory Access (RDMA) is expected to be an integral communication mechanism for future exascale systems-enabling asynchronous data transfers, so that applications may fully utilize CPU resources while simultaneously sharing data amongst remote nodes. In this work we examine Network-induced Memory Contention (NiMC) on Infiniband networks. We expose the interactions between RDMA, main-memory and cache, when applications and out-of-band services compete for memory resources. We then explore NiMC's resulting impact on application-level performance. For a range of hardware technologies and HPC workloads, we quantify NiMC and show that NiMC's impact grows with scale resulting in up to 3X performance degradation at scales as small as 8K processes even in applications that previously have been shown to be performance resilient in the presence of noise. Additionally, this work examines the problem of predicting NiMC's impact on applications by leveraging machine learning and easily accessible performance counters. This approach provides additional insights about the root cause of NiMC and facilitates dynamic selection of potential solutions. Lastly, we evaluated three potential techniques to reduce NiMC's impact, namely hardware offloading, core reservation and software-based network throttling.
Taylor L. Groves, Ryan E. Grant, Aaron Gonzales, Dorian C. Arnold
IEEE Trans. Parallel Distributed Syst.2
2017 A Tale of Two Systems: Using Containers to Deploy HPC Applications on Supercomputers and Clouds
abstract
Containerization, or OS-level virtualization has taken root within the computing industry. However, container utilization and its impact on performance and functionality within High Performance Computing (HPC) is still relatively undefined. This paper investigates the use of containers with advanced supercomputing and HPC system software. With this, we define a model for parallel MPI application DevOps and deployment using containers to enhance development effort and provide container portability from laptop to clouds or supercomputers. In this endeavor, we extend the use of Sin- gularity containers to a Cray XC-series supercomputer. We use the HPCG and IMB benchmarks to investigate potential points of overhead and scalability with containers on a Cray XC30 testbed system. Furthermore, we also deploy the same containers with Docker on Amazon's Elastic Compute Cloud (EC2), and compare against our Cray supercomputer testbed. Our results indicate that Singularity containers operate at native performance when dynamically linking Cray's MPI libraries on a Cray supercomputer testbed, and that while Amazon EC2 may be useful for initial DevOps and testing, scaling HPC applications better fits supercomputing resources like a Cray.
Andrew J. Younge, Kevin T. Pedretti, Ryan E. Grant, Ron Brightwell
CloudCom3
2017 Enabling Diverse Software Stacks on Supercomputers Using High Performance Virtual Clusters
abstract
While large-scale simulations have been the hallmark of the High Performance Computing (HPC) community for decades, Large Scale Data Analytics (LSDA) workloads are gaining attention within the scientific community not only as a processing component to large HPC simulations, but also as standalone scientific tools for knowledge discovery. With the path towards Exascale, new HPC runtime systems are also emerging in a way that differs from classical distributed computing models. However, system software for such capabilities on the latest extreme-scale DOE supercomputing needs to be enhanced to more appropriately support these types of emerging software ecosystems. In this paper, we propose the use of Virtual Clusters on advanced supercomputing resources to enable systems to support not only HPC workloads, but also emerging big data stacks. Specifically, we have deployed the KVM hypervisor within Cray's Compute Node Linux on a XC-series supercomputer testbed. We also use libvirt and QEMU to manage and provision VMs directly on compute nodes, leveraging Ethernet-over-Aries network emulation. To our knowledge, this is the first known use of KVM on a true MPP supercomputer. We investigate the overhead our solution using HPC benchmarks, both evaluating single-node performance as well as weak scaling of a 32-node virtual cluster. Overall, we find single node performance of our solution using KVM on a Cray is very efficient with near-native performance. However overhead increases by up to 20% as virtual cluster size increases, due to limitations of the Ethernet-over-Aries bridged network. Furthermore, we deploy Apache Spark with large data analysis workloads in a Virtual Cluster, effectively demonstrating how diverse software ecosystems can be supported by High Performance Virtual Clusters.
Andrew J. Younge, Kevin T. Pedretti, Ryan E. Grant, Brian L. Gaines, Ron Brightwell
CLUSTER3
2017 sPIN: high-performance streaming processing in the network
abstract
Optimizing communication performance is imperative for large-scale computing because communication overheads limit the strong scalability of parallel applications. Today's network cards contain rather powerful processors optimized for data movement. However, these devices are limited to fixed functions, such as remote direct memory access. We develop sPIN, a portable programming model to offload simple packet processing functions to the network card. To demonstrate the potential of the model, we design a cycle-accurate simulation environment by combining the network simulator Log-GOPSim and the CPU simulator gem5. We implement offloaded message matching, datatype processing, and collective communications and demonstrate transparent full-application speedups. Furthermore, we show how sPIN can be used to accelerate redundant in-memory filesystems and several other use cases. Our work investigates a portable packet-processing network acceleration model similar to compute acceleration with CUDA or OpenCL. We show how such network acceleration enables an eco-system that can significantly speed up applications and system services.
Torsten Hoefler, Salvatore Di Girolamo, Konstantin Taranov, Ryan E. Grant, Ron Brightwell
SC4
2016 RMA-MT: A Benchmark Suite for Assessing MPI Multi-threaded RMA Performance
abstract
Reaching Exascale will require leveraging massive parallelism while potentially leveraging asynchronous communication to help achieve scalability at such large levels of concurrency. MPI is a good candidate for providing the mechanisms to support communication at such large scales. Two existing MPI mechanisms are particularly relevant to Exascale: multi-threading, to support massive concurrency, and Remote Memory Access (RMA), to support asynchronous communication. Unfortunately, multi-threaded MPI RMA code has not been extensively studied. Part of the reason for this is that no public benchmarks or proxy applications exist to assess its performance. The contributions of this paper are the design and demonstration of the first available proxy applications and micro-benchmark suite for multi-threaded RMA in MPI, a study of multi-threaded RMA performance of different MPI implementations, and an evaluation of how these benchmarks can be used to test development for both performance and correctness.
Matthew G. F. Dosanjh, Taylor L. Groves, Ryan E. Grant, Ron Brightwell, Patrick G. Bridges
CCGrid3
2016 (SAI) Stalled, Active and Idle: Characterizing Power and Performance of Large-Scale Dragonfly Networks
abstract
Exascale networks are expected to comprise a significant part of the total monetary cost and 10-20% of the power budget allocated to exascale systems. Yet, our understanding of current and emerging workloads on these networks is limited. Left ignored, this knowledge gap likely will translate into missed opportunities for (1) improved application performance and (2) decreased power and monetary costs in next generation systems. This work targets a detailed understanding and analysis of the performance and utilization of the dragonfly network topology. Using the Structural Simulation Toolkit (SST) and a range of relevant workloads on a dragonfly topology of 110,592 nodes, we examine network design tradeoffs amongst execution time, power, bandwidth, and the number of global links. Our simulations report stalled, active and idle time on a per-port level of the fabric, in order to provide a detailed picture of future networks. The results of this work show potential savings of 3-10% of the exascale power budget and provide valuableinsights to researchers looking for new opportunities to improve performance and increase power efficiency of next generation HPC systems.
Taylor L. Groves, Ryan E. Grant, Karl S. Hemmert, Simon D. Hammond, Michael J. Levenhagen, Dorian C. Arnold
CLUSTER2
2016 NiMC: Characterizing and Eliminating Network-Induced Memory Contention
abstract
Remote Direct Memory Access (RDMA) is expected to be an integral communication mechanism for future exascale systems -- enabling asynchronous data transfers, so that applications may fully utilize all CPU resources while simultaneously sharing data amongst remote nodes. We examined this network-induced memory contention (NiMC), the interactions between RDMA and the memory subsystem when applications and out-of-band services compete for memory resources, and NiMC's resulting impact on application-level performance. For a range of hardware technologies and HPC workloads, we quantified NiMC and show that NiMC's impact grows with scale resulting in up to 3X performance degradation at scales as small as 8K processes even in applications that previously have been shown to be performance resilient in the presence of noise. We also evaluated three potential techniques to reduce NiMC's performance impact, namely hardware offloading, core reservation and software-based network throttling. While all three of these solutions show promise, we provide guidelines that help select the best solution for a given environment.
Taylor L. Groves, Ryan E. Grant, Dorian C. Arnold
IPDPS2
2016 MPI Sessions: Leveraging Runtime Infrastructure to Increase Scalability of Applications at Exascale
abstract
MPI includes all processes in MPI_COMM_WORLD; this is untenable for reasons of scale, resiliency, and overhead. This paper offers a new approach, extending MPI with a new concept called Sessions, which makes two key contributions: a tighter integration with the underlying runtime system; and a scalable route to communication groups. This is a fundamental change in how we organise and address MPI processes that removes well-known scalability barriers by no longer requiring the global communicator MPI_COMM_WORLD.
Daniel J. Holmes, Kathryn Mohror, Ryan E. Grant, Anthony Skjellum, Martin Schulz 0001, Wesley Bland, Jeffrey M. Squyres
EuroMPI3
2016 Program optimizations: The interplay between power, performance, and energy
Edgar A. León, Ian Karlin, Ryan E. Grant, Matthew G. F. Dosanjh
Parallel Comput.3
2015 Re-evaluating Network Onload vs. Offload for the Many-Core Era
abstract
This paper explores the trade-offs between on-loaded versus offloaded network stack processing for systems with varying CPU frequencies. This study explores the differences of onload and offload using experiments run at different DVFS settings to change the frequency, while measuring performance and power. This allows for a quantitative comparison of the the performance and power and trade-offs between onload and offload cards, with a wide range of CPU performances. The results show that there is often a significant performance increase in using offloaded cards especially at lower CPU frequencies, with only a small increase in power usage. This study also uses MPI profiling to analyze why some applications see a larger benefit than others. This paper's contributions are an analytical, quantitative analysis of the trade-offs between onload and offload. While there has been debate to this question, this is the first, to the authors' knowledge, analytical evaluation of the performance difference. The range of frequencies analyzed give insight on how this MPI might perform on different architectures, such as the low frequency, many-core CPUs. Finally, the power measurements allow for the study to provide further depth in the analysis.
Matthew G. F. Dosanjh, Ryan E. Grant, Patrick G. Bridges, Ron Brightwell
CLUSTER2
2015 Optimizing Explicit Hydrodynamics for Power, Energy, and Performance
abstract
Practical considerations for future supercomputer designs will impose limits on both instantaneous power consumption and total energy consumption. Working within these constraints while providing the maximum possible performance, application developers will need to optimize their code for speed alongside power and energy concerns. This paper analyzes the effectiveness of several code optimizations including loop fusion, data structure transformations, and global allocations. A per component measurement and analysis of different architectures is performed, enabling the examination of code optimizations on different compute subsystems. Using an explicit hydrodynamics proxy application from the U.S. Department of Energy, LULESH, we show how code optimizations impact different computational phases of the simulation. This provides insight for simulation developers into the best optimizations to use during particular simulation compute phases when optimizing code for future supercomputing platforms. We examine and contrast both x86 and Blue Gene architectures with respect to these optimizations.
Edgar A. León, Ian Karlin, Ryan E. Grant
CLUSTER3
2015 Scalable connectionless RDMA over unreliable datagrams
Ryan E. Grant, Mohammad J. Rashti, Pavan Balaji, Ahmad Afsahi
Parallel Comput.1
2014 Energy Consumption of Resilience Mechanisms in Large Scale Systems
abstract
As HPC systems continue to grow to meet the requirements of tomorrow's exascale-class systems, two of the biggest challenges are power consumption and system resilience. On current systems, the dominant resilience technique is checkpoint/restart. It is believed, however, that this technique alone will not scale to the level necessary to support future systems. Therefore, alternative methods have been suggested to augment checkpoint/restart -- for example process replication. In this paper we address both resilience and power together, this is in contrast to much of the competed work which does so independently. Using an analytical model that accounts for both power consumption and failures, we study the performance of checkpoint and replication-based techniques on current and future systems and use power measurements from current systems to validate our findings. Lastly, in an attempt to optimize power consumption for replication, we introduce a new protocol termed shadow replication which not only reduces energy consumption but also produces faster response times than checkpoint/restart and traditional replication when operating under system power constraints.
Bryan N. Mills, Taieb Znati, Rami G. Melhem, Kurt B. Ferreira, Ryan E. Grant
PDP5
2013 Protocols for Fully Offloaded Collective Operations on Accelerated Network Adapters
abstract
With each successive generation, network adapters for high-performance networks are becoming more powerful and feature rich. High-performance NICs can now provide support for performing complex group communication operations on the NIC without any host CPU involvement. Several "offloading interfaces" have been designed with the collective communications goal being the complete offloading of arbitrary communication patterns. In this work, we analyze the offloading model offered in the Portals 4 specification in detail. We perform a theoretical analysis based on abstract communication graphs and show several protocols for implementing offloaded communication schedules. Based on our analysis, we propose and implement an extension to the Portals 4 specification that enables offloading any communication pattern completely to the NIC. Our measurements with several advanced communication algorithms confirm that the enhancements provide good overlap and asynchronous progress in practical settings. Altogether, we demonstrate a complete and simple scheme for implementing arbitrary offloaded communication algorithms and hardware. Our protocols can act as a blueprint for the development of communication hardware and middleware while optimizing the whole communication stack.
Timo Schneider, Torsten Hoefler, Ryan E. Grant, Brian W. Barrett, Ron Brightwell
ICPP3
2011 RDMA Capable iWARP over Datagrams
abstract
iWARP is a state of the art high-speed connection-based RDMA networking technology for Ethernet networks to provide InfiniBand-like zero-copy and one-sided communication capabilities over Ethernet. Despite the benefits offered by iWARP, many data center and web-based applications, such as stock-market trading and media-streaming applications, that rely on data gram-based semantics (mostly through UDP/IP) cannot take advantage of it because the iWARP standard is only defined over reliable, connection-oriented transports. This paper presents an RDMA model that functions over reliable and unreliable data grams. The ability to use data grams significantly expands the application space serviced by iWARP and can bring the scalability advantages of a connectionless transport to iWARP. In our previous work, we had developed an iWARP data gram solution using send/receive semantics showing excellent memory scalability and performance benefits over the current TCP-based iWARP. In this paper, we demonstrate an improved iWARP design that provides true RDMA semantics over data grams. Specifically, because traditional RDMA semantics do not map well to unreliable communication, we propose RDMA Write-Record, the first and the only method capable of supporting RDMA Write over both unreliable and reliable data grams. We demonstrate through a proof-of-concept software implementation that data gram-iWARP is feasible for real-world applications. Our proposed RDMA Write-Record method has been designed with data loss in mind and can provide superior performance under conditions of packet loss. It is shown through micro-benchmarks that by using RDMA capable data gram-iWARP a maximum of 256% increase in large message bandwidth and a maximum of 24.4\% improvement in small message latency can be achieved over traditional iWARP. For application results we focus on streaming applications, showing a 24% improvement in memory usage and up to a 74% improvement in performance, although the proposed approach is also applicable to the HPC domain.
Ryan E. Grant, Mohammad J. Rashti, Ahmad Afsahi, Pavan Balaji
IPDPS1
2010 iWARP redefined: Scalable connectionless communication over high-speed Ethernet
abstract
iWARP represents the leading edge of high performance Ethernet technologies. By utilizing an asynchronous communication model, iWARP brings the advantages of OS bypass and RDMA technology to Ethernet. The current specification of iWARP is only defined over connection-oriented transports such as TCP. The memory requirements of many connections along with TCP's flow and reliability controls lead to scalability and performance issues for large-scale HPC and datacenter applications. In this research, we propose guidelines to extend iWARP over datagrams to provide better scalability and performance. While the proposed extension is designed for use in both HPC and datacenters, the emphasis of this paper is on HPC applications. We present our software implementation of datagram-iWARP over UDP and MPI over datagram-iWARP. Our microbenchmark and MPI application results show performance and memory usage benefits for MPI applications, promoting the use of datagram-iWARP for large-scale HPC applications.
Mohammad J. Rashti, Ryan E. Grant, Ahmad Afsahi, Pavan Balaji
HiPC2
2010 A study of hardware assisted IP over InfiniBand and its impact on enterprise data center performance
abstract
High-performance sockets implementations such as the Sockets Direct Protocol (SDP) have traditionally showed major performance advantages compared to the TCP/IP stack over InfiniBand (IPoIB). These stacks bypass the kernel-based TCP/IP and take advantage of network hardware features, providing enhanced performance. SDP has excellent performance but limited utility as only applications relying on the TCP/IP sockets API can use it and other IP stack uses (IPSec, UDP, SCTP) or TCP layer modifications (iSCSI) cannot benefit from it. Recently, newer generations of InfiniBand adapters, such as ConnectX from Mellanox, have provided hardware support for the IP stack itself, such as Large Send Offload and Large Receive Offload. As such high performance socket networks are likely to be deployed or converged with existing Ethernet networking solutions, the performance of such technologies is important to assess. In this paper we take a first look at the performance advantages provided by these offload techniques and compare them to SDP. Our micro-benchmarks and enterprise data-center experiments show that hardware assisted IPoIB can provide competitive performance with SDP and even outperform it in some cases.
Ryan E. Grant, Pavan Balaji, Ahmad Afsahi
ISPASS1
2009 Evaluation of ConnectX Virtual Protocol Interconnect for Data Centers
abstract
With the emergence of new technologies such as Virtual Protocol Interconnect (VPI) for the modern data center, the separation between commodity networking technology and high-performance interconnects is shrinking. With VPI, a single network adapter on a data center server can easily be configured to use one port to interface with Ethernet traffic and another port to interface with high-bandwidth, low-latency InfiniBand technology. In this paper, we evaluate ConnectX VPI using microbenchmarks as well as real traces from a three-tier data center architecture. We find that with VPI each network segment in the data center can use the most optimal configuration (whether InfiniBand or Ethernet) without having to fall back to the lowest common denominator, as is currently the case. Our results show a maximum 26.7% increase in bandwidth, a 54.5% reduction in latency, and a 5% increase in real data center throughput.
Ryan E. Grant, Ahmad Afsahi, Pavan Balaji
ICPADS1
2009 Improving energy efficiency of asymmetric chip multithreaded multiprocessors through reduced OS noise scheduling
abstract
Abstract The performance of the emerging chip multithreaded symmetric multiprocessors (SMPs) is of great importance to the high performance computing community. However, the growing power consumption of such systems is of increasing concern, and techniques that can be used to increase the overall system power efficiency while sustaining the performance are very desirable. Operating system (OS) noise can have a dramatic effect on the system performance. Effectively handling the smaller OS tasks while simultaneously preserving application thread synchronicity leads to gains in the overall system efficiency. Recently, under a fixed power budget, asymmetric multiprocessors (AMP) have been proposed to improve the performance of multithreaded applications. An AMP in this context is a multiprocessor system in which its processors are not operating at the same frequency. This paper proposes two simple scheduling methods that reduce the impact of OS noise, while simultaneously taking advantage of an opportunity to increase the overall machine energy efficiency on AMP servers. Prototyping AMPs on a commercial 2‐way dual‐core Hyper‐Threaded (HT) Intel Xeon SMP server, using real power measurements across six SPEC OpenMP applications, indicates that the first proposed scheduler performs better on average for HT‐enabled systems, whereas the second scheduler is superior on average for HT‐disabled systems. Copyright © 2009 John Wiley & Sons, Ltd.
Ryan E. Grant, Ahmad Afsahi
Concurr. Comput. Pract. Exp.1
2007 Improving system efficiency through scheduling and power management
abstract
The performance of the emerging commercial chip multithreaded multiprocessors is of great importance to the high performance computing community. However, the growing power consumption of such systems is of increasing concern, and techniques that could be effectively used to increase overall system power efficiency while sustaining performance are very desirable.
Ryan E. Grant, Ahmad Afsahi
CLUSTER1
2007 A Comprehensive Analysis of OpenMP Applications on Dual-Core Intel Xeon SMPs
abstract
Hybrid chip multithreaded SMPs present new challenges as well as new opportunities to maximize performance. Our intention is to discover the optimal operating configuration of such systems for scientific applications and to identify the shared resources that might become a bottleneck to performance under the different hardware configurations. This knowledge will be useful to the research community in developing software techniques to improve the performance of shared memory programs on modern multi-core multiprocessors. In this paper, we study a two-way dual-core Hyper-Threaded (HT) Intel Xeon SMP server under single program and multi-program multithreaded workloads using the NAS OpenMP benchmark suite. Our performance results indicate that in the single-program case, the CMP-based SMP and CMT-based SMP configurations have the highest average speedup across all of the applications. The most efficient architecture is a single HT-enabled dual-core processor that is almost comparable to the performance of a 2-way dual-core HT-disabled system.
Ryan E. Grant, Ahmad Afsahi
IPDPS1
2006 Power-performance efficiency of asymmetric multiprocessors for multi-threaded scientific applications
abstract
Recently, under a fixed power budget, asymmetric multiprocessors (AMP) have been proposed to improve the performance of multi-threaded applications compared to symmetric multiprocessors. An AMP is a multiprocessor system in which its processors are not operating at the same frequency. Power consumption has become an important design constraint in servers and high-performance server clusters. This paper explores the power-performance efficiency of hyper-threaded (HT) AMP servers, and proposes a new scheduling algorithm that can be used to reduce the overall power consumption of a server while maintaining a high level of performance. Prototyping AMPs on a commercial 4-way SMP server, we show that on average 15.6% energy savings and 6.1% slowdown for the HT-disabled case, and 7.1% energy savings and 4.8% slowdown for the HT-enabled case can be achieved across NAS and SPEC OpenMP applications.
Ryan E. Grant, Ahmad Afsahi
IPDPS1