Derek Schafer

dblp:249/5683 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0001-8438-5144ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Performance Analysis of Open MPI on AMR Applications over Slingshot-11
Maxim Moraru, Howard Pritchard, Derek Schafer, Galen M. Shipman, Patrick G. Bridges
EuroMPI3
2024 Quantifying and Modeling Irregular MPI Communication
abstract
Many modern scientific applications have communication patterns where both the number of communication partners and amount of data transmitted between process pairs vary significantly and change over time. This work describes an approach to measure and model the behavior of these irregular, dynamic MPI communication patterns on modern high-performance computing systems. Specifically, this approach quantifies communication behavior using a small number of stochastic random variables that capture key features of irregular communication patterns, and estimates the distributions of these variables either parameterically or empirically. This work then demonstrates that the collected parameters and their distributions can be used to measure and model the communication performance of several MPI applications. It also presents a synthetic benchmark that uses these distributions to recreate statistically similar communication patterns. This approach provides a lightweight method to reproduce communication patterns of a variety of applications with minimal overhead while gaining additional insights into the performance and characteristics of various irregular communication patterns.
Carson Woods, Derek Schafer, Patrick G. Bridges, Anthony Skjellum
CCGrid2
2024 Optimizing Neighbor Collectives with Topology Objects
abstract
Many HPC applications implement non-cartesian neighbor data exchanges using MPI point-to-point operations rather than utilizing native MPI neighbor collective methods. Each application must therefore implement their own commu-nication optimizations, rather than leveraging any optimizations that could be provided by MPI. While an interface for such optimizations is provided within MPI through neighborhood collectives, applications avoid these methods due to the lack of performance optimizations within them along with large costs associated with graph communicator formation. This paper presents a novel approach for creating local, non-cartesian topol-ogy objects that provides finer control over the aforementioned setup costs. Any additional setup costs, such as initializing per-iteration optimizations, can then be deferred until additional information is available, such as within persistent initialization calls. This paper describes our implementation within an MPI extension library and demonstrates the effectiveness of our approach in simple benchmarks and real-world applications.
Gerald Collom, Derek Schafer, Amanda Bienz, Patrick G. Bridges, Galen M. Shipman
CLUSTER2
2024 A More Scalable Sparse Dynamic Data Exchange
abstract
Parallel architectures are continually increasing in performance and scale while underlying algorithmic infrastruc-ture often fails to take full advantage of available compute power. Within the context of MPI, irregular communication patterns create bottlenecks in parallel applications. One common bottleneck is the sparse dynamic data exchange, often required when forming communication patterns within applications. There is a large variety of approaches for these dynamic exchanges, with optimizations implemented directly in parallel applications. This paper proposes a novel API within an MPI eXtension library, allowing applications to utilize the variety of provided optimizations for sparse dynamic data exchange methods. Fur-ther, the paper presents novel locality-aware sparse dynamic data exchange algorithms. Finally, performance results show locality-aware approaches achieve up to 128x over existing approaches when exchanging only pattern of communication, and up to 54x when exchanging data to be communicated as well.
Andrew Geyko, Gerald Collom, Derek Schafer, Patrick G. Bridges, Amanda Bienz
HiPC3
2024 Understanding GPU Triggering APIs for MPI+X Communication
Patrick G. Bridges, Anthony Skjellum, Evan Drake Suggs, Derek Schafer, Purushotham V. Bangalore
EuroMPI4
2023 Design of a portable implementation of partitioned point-to-point communication primitives
abstract
Abstract The Message Passing Interface (MPI) has been the dominant message passing solution for scientific computing for decades. MPI point‐to‐point communications are highly efficient mechanisms for process‐to‐process communication. However, MPI performance when processes utilize multiple threads is slowed by concurrency protections in the MPI library. MPI's current thread level interface imposes these overheads throughout the library when thread safety is needed. While much work has been done to reduce multithreading overheads in MPI, a solution is needed that reduces the number of messages exchanged in a threaded environment. Partitioned communication is included in the MPI 4.0 standard as an alternative that addresses the challenges of multithreaded communication in MPI today. Partitioned communication reduces overall message volume by creating a buffer‐sharing mechanism between threads such that they can indicate when portions of a communication buffer are available to be sent. Separation of the control and data planes in MPI is enabled by allowing persistent initialization and single occurrence message buffer matching from the indication that the data is ready to be sent. This enables the usage of underlying hardware primitives like triggered operations, where commands (destination, size, etc.) can be set up prior to data buffer readiness and readiness triggered with a simple doorbell/counter later. This approach is useful for future development of MPI operations in environments where traditional networking commands can have performance challenges, like accelerators (GPUs, FPGAs). In this paper, we detail the design and implementation of a layered library (built on top of MPI‐3.1) and an integrated Open MPI solution that supports the new, MPI‐4.0 partitioned communication feature set. The library will enable applications to use currently released MPI implementations and older legacy libraries to provide partitioned communication support while also enabling further exploration of this new communication model in new applications and use cases. We will compare the designs of the library and native Open MPI support, provide performance results and comparisons between the two approaches, and lessons learned from the implementation of partitioned communication in both library and native forms. We find that the native implementation and library have similar performance with a percentage difference under 0.94% in microbenchmarks and performance within 5% for a partitioned communication enabled proxy application.
W. Pepper Marts, Andrew Worley, Prema Soundarajan, Derek Schafer, Matthew G. F. Dosanjh, Ryan E. Grant, Purushotham V. Bangalore, Anthony Skjellum, Sheikh K. Ghafoor
Concurr. Comput. Pract. Exp.4
2022 Reconfigurable switches for high performance and flexible MPI collectives
abstract
Abstract There has been much effort in offloading MPI collective operations into hardware. But while NIC‐based collective acceleration is well‐studied, offloading their processing into the switching fabric, despite numerous advantages, has been much more limited. A major problem with fixed logic implementations is that either only a fraction of the possible collective communication is accelerated or that logic is wasted in the applications that do not need a particular capability. Using reconfigurable logic has numerous advantages: exactly the required operations can be implemented; the level of desired performance can be specified; and new, possibly complex, operations can be defined and implemented. We have designed an in‐switch collective accelerator,MPI‐FPGA, and demonstrated its use with seven MPI collectives and over a set of benchmarks and proxy applications (MiniApps). The accelerator uses a novel two‐level switch design containing fully pipelined vectorized aggregation logic units. Essential to this work is providing support for sub‐communicator collectives that enables communicators of arbitrary shape, and that is scalable to large systems. A streaming interface improves the performance for long messages. While this reconfigurable design is generally applicable, we prototype it with an FPGA‐centric cluster. A sampleMPI‐FPGAdesign in a direct network achieves considerable speedups over conventional clusters in the most likely scenarios. We also present results for indirect networks with reconfigurable high‐radix switches and show that this approach is competitive withSHArPtechnology for the subset of operations thatSHArPsupports.MPI‐FPGAis fully integrated into MPICH and is transparent to MPI applications.
Pouya Haghi, Anqi Guo, Qingqing Xiong, Chen Yang 0010, Tong Geng, Justin T. Broaddus, Ryan J. Marshall, Derek Schafer, Anthony Skjellum, Martin C. Herbordt
Concurr. Comput. Pract. Exp.8
2021 Implementation and evaluation of MPI 4.0 partitioned communication libraries
Matthew G. F. Dosanjh, Andrew Worley, Derek Schafer, Prema Soundararajan, Sheikh K. Ghafoor, Anthony Skjellum, Purushotham V. Bangalore, Ryan E. Grant
Parallel Comput.3
2020 Why is MPI (perceived to be) so complex?: Part 1 - Does strong progress simplify MPI?
abstract
Strong progress is optional in MPI. MPI allows implementations where progress (for example, updating the message-transport state machines or interaction with network devices) is only made during certain MPI procedure calls. Generally speaking, strong progress implies the ability to achieve progress (to transport data through the network from senders to receivers and exchange protocol messages) without explicit calls from user processes to MPI procedures. For instance, data given to a send procedure that matches a pre-posted receive on the receiving process is moved from source to destination in due course regardless of how often (including zero times) the sender or receiver processes call MPI in the meantime. Further, nonblocking operations and persistent collective operations work ‘in the background’ of user processes once all processes in the communicator’s group have performed the starting step for the operation. Overall, strong progress is meant to enhance the potential for overlap of communication and computation and improve predictability of procedure execution times by eliminating progress effort from user threads. This paper posits that strong progress is desirable as an MPI implementation property and examines whether strong progress: This paper explores such possibilities and sets forth principles that underpin MPI and interactions with normal and fault modes of operation. The key contribution of this paper is the conclusion that, whether measured by absolute performance, by performance portability, or by interface simplicity, strong progress in MPI is no worse than weak progress and, in most scenarios, has more potential to fulfil the aforementioned desirable attributes.
Daniel J. Holmes, Anthony Skjellum, Derek Schafer
EuroMPI3
2019 User-Level Scheduled Communications for MPI
abstract
Composability is one of seven reasons for the long-standing and continuing success of MPI. Extending MPI by composing its operations with user-level operations provides useful integration with the progress engine and completion notification methods of MPI. However, the existing extensibility mechanism in MPI (generalized requests) is not widely utilized and has significant drawbacks. MPI can be generalized via scheduled communication primitives, for example, by utilizing implementation techniques from existing MPI-3 nonblocking collectives and from forthcoming MPI-4 persistent and partitioned APIs. Non-trivial schedules are used internally in some MPI libraries; but, they are not accessible to end-users. Message-based communication patterns can be built as libraries on top of MPI. Such libraries can have comparable implementation maturity and potentially higher performance than MPI library code, but do not require intimate knowledge of the MPI implementation. Libraries can provide performance-portable interfaces that cross MPI implementation boundaries. The ability to compose additional user-defined operations using the same progress engine benefits all kinds of general purpose HPC libraries. We propose a definition for MPI schedules: a user-level programming model suitable for creating persistent collective communication composed with new application-specific sequences of user-defined operations managed by MPI and fully integrated with MPI progress and completion notification. The API proposed offers a path to standardization for extensible communication schedules involving user-defined operations. Our approach has the potential to introduce event-driven programming into MPI (beyond the tools interface), although connecting schedules with events comprises future work. Early performance results described here are promising and indicate strong overlap potential.
Derek Schafer, Sheikh K. Ghafoor, Daniel J. Holmes, Martin Ruefenacht, Anthony Skjellum
HiPC1