W. Pepper Marts

dblp:247/5660 · also William Pepper Marts · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
5since 2021 · last 2025
0009-0009-3764-1905ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Measuring Thread Timing to Assess the Feasibility of Early-Bird Message Delivery Across Systems and Scales
abstract
ABSTRACT Early‐bird communication is a communication/computation overlap technique that leverages fine‐grained communication to improve application run‐time. Communication is divided such that each individual thread can initiate transmission of its portion of the data upon completion rather than waiting for a dedicated communication phase. The benefit of early‐bird communication depends on the completion timing of the individual threads: On the one hand, if all threads are complete at nearly the same time, the overheads of sending multiple messages will accumulate, leading to performance that is worse than if a single message had been sent. On the other hand, if thread completions are spread out in time, those that complete earlier can send data while others continue working, leading to performance that is better than if a single message had been sent. The challenge is that the completion times are currently unknown and can vary based on application, problem size, system software, and underlying hardware. In this paper, we address this lacuna by measuring and evaluating the potential overlap afforded by early‐bird communication for a selection of proxy applications. These measurements help us understand whether a given application could benefit from early‐bird communication. We present our technique for gathering this data and evaluate data collected from three proxy applications: MiniFE, MiniMD, and MiniQMC. Each application is run on three systems with distinct CPU architectures and strong scales across three run sizes. To characterize the behavior of these workloads, we study the trends of thread timings at both a macro level, across all threads across all runs of an application, and a micro level, that is, within a single process of a single run. We observe that our tested applications exhibit significantly different thread arrival distributions. The machine used had a significant impact, with the window of potential overlap varying by as much as an order of magnitude.
W. Pepper Marts, Matthew G. F. Dosanjh, Whit Schonbein, Scott Levy, Patrick G. Bridges
Concurr. Comput. Pract. Exp.1
2024 CMB: A Configurable Messaging Benchmark to Explore Fine-Grained Communication
abstract
Modern communication APIs provide increased ability to specify when, where, and how to send data between processes. One recent innovation is fine-grained communication, where processes are able to send subsets of data as it is ready rather than waiting for the entirety of the data to be completed. Allowing data to be sent when it is ready increases opportunities for overlapping communication and computation. However, with multiple fine-grained, thread-safe interfaces, the task of optimizing an application’s peer-to-peer fine-grained communication is complex. In this paper, we present the Configurable Messaging Benchmark (CMB), a tool for evaluating the application impact of fine-grained communication. Using the CMB we perform a case study to measure the impact of different fine-grained implementations on a variety of realistic application profiles. Initial results reveal a large optimization space ranging from potential speedups as high as 52.97% to slowdowns as high as 289.55% relative to bulk-synchronous MPI message passing.
W. Pepper Marts, Donald A. Kruse, Matthew G. F. Dosanjh, Whit Schonbein, Scott Levy, Patrick G. Bridges
CCGrid1
2023 Modeling and Benchmarking the Potential Benefit of Early-Bird Transmission in Fine-Grained Communication
abstract
Traditional point-to-point communication sends data only after the entirety of the data is available. This includes situations where multiple actors (e.g., threads) contribute to the send buffer. As a result, cases where the completion times of these actors are widely distributed may be lost opportunities for optimization because data ready to be sent is waiting to be transmitted. Fine-grained communication exposes these opportunities by allowing buffers to be divided into element s that can then be sent independently (see e.g., Partitioned Communication in Message Passing Interface v4.0). While some research has been directed at exploring the utility of such ‘early-bird’ transmission, the overall search space for finding the best performing actor completion timings and element counts is large. In this work, we present an abstract model of fine-grained communication based on the LogGP model and a complementary benchmark. We use the model to explore actor completion timing scenarios and identify trends in communication behavior based on factors such as overall message size and delay between actor completions. We evaluate the benchmarks on three systems utilizing distinct network technologies and show that: (i) smaller numbers of element s are able to exploit most of the benefit of early-bird communication, (ii) performance benefit will depend non-trivially on application behavior, and (iii) benefits are highly network-dependent.
Whit Schonbein, Scott Levy, Matthew G. F. Dosanjh, W. Pepper Marts, Elizabeth Reid 0002, Ryan E. Grant
ICPP4
2023 Design of a portable implementation of partitioned point-to-point communication primitives
abstract
Abstract The Message Passing Interface (MPI) has been the dominant message passing solution for scientific computing for decades. MPI point‐to‐point communications are highly efficient mechanisms for process‐to‐process communication. However, MPI performance when processes utilize multiple threads is slowed by concurrency protections in the MPI library. MPI's current thread level interface imposes these overheads throughout the library when thread safety is needed. While much work has been done to reduce multithreading overheads in MPI, a solution is needed that reduces the number of messages exchanged in a threaded environment. Partitioned communication is included in the MPI 4.0 standard as an alternative that addresses the challenges of multithreaded communication in MPI today. Partitioned communication reduces overall message volume by creating a buffer‐sharing mechanism between threads such that they can indicate when portions of a communication buffer are available to be sent. Separation of the control and data planes in MPI is enabled by allowing persistent initialization and single occurrence message buffer matching from the indication that the data is ready to be sent. This enables the usage of underlying hardware primitives like triggered operations, where commands (destination, size, etc.) can be set up prior to data buffer readiness and readiness triggered with a simple doorbell/counter later. This approach is useful for future development of MPI operations in environments where traditional networking commands can have performance challenges, like accelerators (GPUs, FPGAs). In this paper, we detail the design and implementation of a layered library (built on top of MPI‐3.1) and an integrated Open MPI solution that supports the new, MPI‐4.0 partitioned communication feature set. The library will enable applications to use currently released MPI implementations and older legacy libraries to provide partitioned communication support while also enabling further exploration of this new communication model in new applications and use cases. We will compare the designs of the library and native Open MPI support, provide performance results and comparisons between the two approaches, and lessons learned from the implementation of partitioned communication in both library and native forms. We find that the native implementation and library have similar performance with a percentage difference under 0.94% in microbenchmarks and performance within 5% for a partitioned communication enabled proxy application.
W. Pepper Marts, Andrew Worley, Prema Soundarajan, Derek Schafer, Matthew G. F. Dosanjh, Ryan E. Grant, Purushotham V. Bangalore, Anthony Skjellum, Sheikh K. Ghafoor
Concurr. Comput. Pract. Exp.1
2021 MiniMod: A Modular Miniapplication Benchmarking Framework for HPC
abstract
The HPC application community has proposed many new application communication structures, middleware interfaces, and communication models to improve HPC application performance. Modifying proxy applications is the standard practice for the evaluation of these novel methodologies. Currently, this requires the creation of a new version of the proxy application for each combination of the approach being tested. In this article, we present a modular proxy-application framework, MiniMod, that enables evaluation of a combination of independently written computation kernels, data transfer logic, communication access, and threading libraries. MiniMod is designed to allow rapid development of individual modules which can be combined at runtime. Through MiniMod, developers only need a single implementation to evaluate application impact under a variety of scenarios.We demonstrate the flexibility of MiniMod’s design by using it to implement versions of a heat diffusion kernel and the miniFE finite element proxy application, along with a variety of communication, granularity, and threading modules. We examine how changing communication libraries, communication granularities, and threading approaches impact these applications on an HPC system. These experiments demonstrate that MiniMod can rapidly improve the ability to assess new middleware techniques for scientific computing applications and next-generation hardware platforms.
W. Pepper Marts, Matthew G. F. Dosanjh, Scott Levy, Whit Schonbein, Ryan E. Grant, Patrick G. Bridges
CLUSTER1
2019 MPI tag matching performance on ConnectX and ARM
abstract
As we approach Exascale, message matching has increasingly become a significant factor in HPC application performance. To address this, network vendors have placed higher precedence on improving MPI message matching performance. ConnectX-5, Mellanox's new network interface card, has both hardware and software matching layers. The performance characteristics of these layers have yet to be studied under real world circumstances. In this work we offer an initial evaluation of ConnectX-5 message matching performance. To analyze this new hardware we executed a series of micro-benchmarks and applications on Astra, an ARM-based ConnectX-5 HPC system, while varying hardware and software matching parameters. The benchmark results show the ConnectX-5 is sensitive to queue depths, and that hardware message matching increases performance for applications that send messages between 1KiB and 16KiB. Furthermore, the hardware matching system was capable of matching wildcard receives without negatively impacting performance. Finally, for some applications, a significant improvement can be observed when leveraging the ConnectX-5's hardware matching.
W. Pepper Marts, Matthew G. F. Dosanjh, Whit Schonbein, Ryan E. Grant, Patrick G. Bridges
EuroMPI1