Whit Schonbein

dblp:117/1671 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0003-4955-2984ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
YearPublicationVenuePosition
2025 Utilizing Network Hardware Parallelism for MPI Partitioned Collective Communication
abstract
Parallel distributed applications running on large-scale high-performance computing systems depend on effective point-to-point and collective communication to meet performance goals. Beginning with version 4.0, the Message Passing Interface (MPI) introduced the partitioned communication API, providing tools for addressing communication bottlenecks raised by hybrid communication models. This API allows individual actors (CPU threads, GPU threads, etc.) to initiate communication on portions of complete buffers, enabling additional communication/computation overlap. Intuitively, the utility of partitioned communication could benefit from network-level support: If there are multiple paths between endpoints, an MPI-aware network could disperse partitions across these paths, avoiding the data serialization entailed by a dependency on a single path. The Cerio Rockport Ethernet Fabric has the ability to expose this capability to communication middleware. In this work we develop this capability to allow for user-level path selection for MPI partitioned communication and explore how this capability impacts point-to-point performance, collective design, and Allreduce efficiency in a Large Language Model task
Yiltan Hassan Temuçin, Amirreza Barati Sedeh, Whit Schonbein, Ryan E. Grant, Ahmad Afsahi
PDP3
2025 Measuring Thread Timing to Assess the Feasibility of Early-Bird Message Delivery Across Systems and Scales
abstract
ABSTRACT Early‐bird communication is a communication/computation overlap technique that leverages fine‐grained communication to improve application run‐time. Communication is divided such that each individual thread can initiate transmission of its portion of the data upon completion rather than waiting for a dedicated communication phase. The benefit of early‐bird communication depends on the completion timing of the individual threads: On the one hand, if all threads are complete at nearly the same time, the overheads of sending multiple messages will accumulate, leading to performance that is worse than if a single message had been sent. On the other hand, if thread completions are spread out in time, those that complete earlier can send data while others continue working, leading to performance that is better than if a single message had been sent. The challenge is that the completion times are currently unknown and can vary based on application, problem size, system software, and underlying hardware. In this paper, we address this lacuna by measuring and evaluating the potential overlap afforded by early‐bird communication for a selection of proxy applications. These measurements help us understand whether a given application could benefit from early‐bird communication. We present our technique for gathering this data and evaluate data collected from three proxy applications: MiniFE, MiniMD, and MiniQMC. Each application is run on three systems with distinct CPU architectures and strong scales across three run sizes. To characterize the behavior of these workloads, we study the trends of thread timings at both a macro level, across all threads across all runs of an application, and a micro level, that is, within a single process of a single run. We observe that our tested applications exhibit significantly different thread arrival distributions. The machine used had a significant impact, with the window of potential overlap varying by as much as an order of magnitude.
W. Pepper Marts, Matthew G. F. Dosanjh, Whit Schonbein, Scott Levy, Patrick G. Bridges
Concurr. Comput. Pract. Exp.3
2024 CMB: A Configurable Messaging Benchmark to Explore Fine-Grained Communication
abstract
Modern communication APIs provide increased ability to specify when, where, and how to send data between processes. One recent innovation is fine-grained communication, where processes are able to send subsets of data as it is ready rather than waiting for the entirety of the data to be completed. Allowing data to be sent when it is ready increases opportunities for overlapping communication and computation. However, with multiple fine-grained, thread-safe interfaces, the task of optimizing an application’s peer-to-peer fine-grained communication is complex. In this paper, we present the Configurable Messaging Benchmark (CMB), a tool for evaluating the application impact of fine-grained communication. Using the CMB we perform a case study to measure the impact of different fine-grained implementations on a variety of realistic application profiles. Initial results reveal a large optimization space ranging from potential speedups as high as 52.97% to slowdowns as high as 289.55% relative to bulk-synchronous MPI message passing.
W. Pepper Marts, Donald A. Kruse, Matthew G. F. Dosanjh, Whit Schonbein, Scott Levy, Patrick G. Bridges
CCGrid4
2024 Leveraging High-Performance Data Transfer to Offload Data Management Tasks to SmartNICs
abstract
Network interface controllers (NICs) with general-purpose compute capabilities (‘SmartNICs’) present an opportunity for reducing host application overheads by offloading non-critical tasks to the NIC. In addition to moving computation, offloading requires that associated data is also transferred to the NIC. To meet this need, we introduce a high-performance, general-purpose data movement service that facilitates the of-floading of tasks to SmartNICs: The SmartNIC Data Movement Service (SDMS). SDMS provides near-line-rate transfer band-widths between the host and NIC. Moreover, SDMS's In-transit Data Placement (IDP) feature can reduce (or even eliminate) the cost of serializing data on the NIC by performing the necessary data formatting during the transfer. To illustrate these capabilities, we provide an in-depth case study using SDMS to offload data management operations related to Apache Arrow, a popular data format standard. For single-column tables, SDMS can achieve more than 87% of baseline throughput for data buffers that are 128 KiB or larger (and more than 95% of baseline throughput for buffers that are 1 MiB or larger) while also nearly eliminating the host and SmartNIC overhead associated with Arrow operations.
Scott Levy, Whit Schonbein, Craig D. Ulmer
CLUSTER2
2024 Smart Network Traffic Prediction for Scientific Applications
abstract
Network traffic in HPC systems can impact application performance by inducing costly re-transmissions or consuming memory bandwidth. Emerging SmartNIC technologies present new opportunities for addressing these issues and optimizing network performance by providing a platform for the intelligent utilization of network resources through machine learning models of network traffic. However, SmartNICs also present challenges for deploying such models insofar as they offer relatively limited computational and memory resources, and these resources must be shared with other services. Based on an analysis of traffic data collected from eight scientific applications and proxies, we explore lightweight approaches to modeling network traffic using dynamic linear regression. Depending on the application, normalized root mean squared error for static regression may be less than 1 %, and dynamic regression can reduce this error by an order of magnitude. We further refine the dynamic approach by adding an additional classifier that categorizes predictions generated by the model as reliable or unreliable, showing that the technique can achieve good precision and recall. Finally, we evaluate the performance of dynamic regression and classification on NVIDIA BlueField-2 and BlueField-3 SmartNICs, demonstrating these computationally lightweight techniques are feasible on contemporary SmartNIC platforms.
Whit Schonbein, Tinotenda Matsika, Ryan E. Grant
PDP1
2023 A Dynamic Network-Native MPI Partitioned Aggregation Over InfiniBand Verbs
abstract
Modern HPC systems require efficient hybrid programming model to utilize their hardware resources effectively. The Message Passing Interface (MPI) has accommodated next-generation hardware by providing new APIs such as the MPI Partitioned interface. This API provides a user with fine-grain communication without the overhead of traditional MPI point-to-point communication in multi-threaded workloads.To the best of our knowledge, we present the first work on detailed low-level design for an MPI Partitioned implementation. We guide readers through a method to map the MPI Partitioned interface to the InfiniBand Verbs API. Alongside implementation details, we also study the aggregation of user partitions and how we can efficiently send them over the network. We study a brute force approach and using the Partitioned LogGP (PLogGP) model to predict ideal aggregation. We observe that using the PLogGP model provides comparable performance without exhausting computing resources to search the entire solution space. The PLogGP design was further optimized by considering how the partition arrival pattern can be used to dynamically modify our aggregation scheme. We profiled our micro-benchmarks to provide analysis on how and why this additional optimization is beneficial to our results and how we can fine-tune this mechanism. Finally, we evaluated our PLogGP and Timer-based PLogGP designs with a commonly used communication pattern in HPC (communication sweep) to observe the impact when communicating with multiple processes in an application-like scenario at 1024 cores.
Yiltan Hassan Temuçin, Scott Levy, Whit Schonbein, Ryan E. Grant, Ahmad Afsahi
CLUSTER3
2023 Modeling and Benchmarking the Potential Benefit of Early-Bird Transmission in Fine-Grained Communication
abstract
Traditional point-to-point communication sends data only after the entirety of the data is available. This includes situations where multiple actors (e.g., threads) contribute to the send buffer. As a result, cases where the completion times of these actors are widely distributed may be lost opportunities for optimization because data ready to be sent is waiting to be transmitted. Fine-grained communication exposes these opportunities by allowing buffers to be divided into element s that can then be sent independently (see e.g., Partitioned Communication in Message Passing Interface v4.0). While some research has been directed at exploring the utility of such ‘early-bird’ transmission, the overall search space for finding the best performing actor completion timings and element counts is large. In this work, we present an abstract model of fine-grained communication based on the LogGP model and a complementary benchmark. We use the model to explore actor completion timing scenarios and identify trends in communication behavior based on factors such as overall message size and delay between actor completions. We evaluate the benchmarks on three systems utilizing distinct network technologies and show that: (i) smaller numbers of element s are able to exploit most of the benefit of early-bird communication, (ii) performance benefit will depend non-trivially on application behavior, and (iii) benefits are highly network-dependent.
Whit Schonbein, Scott Levy, Matthew G. F. Dosanjh, W. Pepper Marts, Elizabeth Reid 0002, Ryan E. Grant
ICPP1
2022 "Smarter" NICs for faster molecular dynamics: a case study
abstract
This work evaluates the benefits of using a “smart” network interface card (SmartNIC) as a compute accelerator for the example of the MiniMD molecular dynamics proxy application. The accelerator is NVIDIA's BlueField-2 card, which includes an 8-core Arm processor along with a small amount of DRAM and storage. We test the networking and data movement performance of these cards compared to a standard Intel server host using microbenchmarks and MiniMD. In MiniMD, we identify two distinct classes of computation, namely core computation and maintenance computation, which are executed in sequence. We restructure the algorithm and code to weaken this dependence and increase task parallelism, thereby making it possible to increase utilization of the BlueField-2 concurrently with the host. We evaluate our implementation on a cluster consisting of 16 dual-socket Intel Broadwell host nodes with one BlueField-2 per host-node. Our results show that while the overall compute performance of BlueField-2 is limited, using them with a modified MiniMD algorithm allows for up to 20% speedup over the host CPU baseline with no loss in simulation accuracy.
Sara Karamati, Clay Hughes, Karl S. Hemmert, Ryan E. Grant, Whit Schonbein, Scott Levy, Thomas M. Conte, Jeffrey Young 0001, Richard W. Vuduc
IPDPS5
2021 MiniMod: A Modular Miniapplication Benchmarking Framework for HPC
abstract
The HPC application community has proposed many new application communication structures, middleware interfaces, and communication models to improve HPC application performance. Modifying proxy applications is the standard practice for the evaluation of these novel methodologies. Currently, this requires the creation of a new version of the proxy application for each combination of the approach being tested. In this article, we present a modular proxy-application framework, MiniMod, that enables evaluation of a combination of independently written computation kernels, data transfer logic, communication access, and threading libraries. MiniMod is designed to allow rapid development of individual modules which can be combined at runtime. Through MiniMod, developers only need a single implementation to evaluate application impact under a variety of scenarios.We demonstrate the flexibility of MiniMod’s design by using it to implement versions of a heat diffusion kernel and the miniFE finite element proxy application, along with a variety of communication, granularity, and threading modules. We examine how changing communication libraries, communication granularities, and threading approaches impact these applications on an HPC system. These experiments demonstrate that MiniMod can rapidly improve the ability to assess new middleware techniques for scientific computing applications and next-generation hardware platforms.
W. Pepper Marts, Matthew G. F. Dosanjh, Scott Levy, Whit Schonbein, Ryan E. Grant, Patrick G. Bridges
CLUSTER4
2020 Tail queues: A multi-threaded matching architecture
abstract
Summary As we approach exascale, computational parallelism will have to drastically increase in order to meet throughput targets. Many‐core architectures have exacerbated this problem by trading reduced clock speeds, core complexity, and computation throughput for increasing parallelism. This presents two major challenges for communication libraries such as MPI: the library must leverage the performance advantages of thread level parallelism and avoid the scalability problems associated with increasing the number of processes to that scale. Hybrid programming models, such as MPI+X, have been proposed to address these challenges. MPI THREAD MULTIPLE is MPI's thread safe mode. While there has been work to optimize it, it largely remains non‐performant in most implementations. While current applications avoid MPI multithreading due to performance concerns, it is expected to be utilized in future applications. One of the major synchronous data structures required by MPI is the matching engine. In this paper, we present a parallel matching algorithm that can improve MPI matching for multithreaded applications. We then perform a feasibility study to demonstrate the performance benefit of the technique.
Matthew G. F. Dosanjh, Ryan E. Grant, Whit Schonbein, Patrick G. Bridges
Concurr. Comput. Pract. Exp.3
2019 Fuzzy Matching: Hardware Accelerated MPI Communication Middleware
abstract
Contemporary parallel scientific codes often rely on message passing for inter-process communication. However, inefficient coding practices or multithreading (e.g., via MPI_THREAD_MULTIPLE) can severely stress the underlying message processing infrastructure, resulting in potentially un-acceptable impacts on application performance. In this article, we propose and evaluate a novel method for addressing this issue: 'Fuzzy Matching'. This approach has two components. First, it exploits the fact most server-class CPUs include vector operations to parallelize message matching. Second, based on a survey of point-to-point communication patterns in representative scientific applications, the method further increases parallelization by allowing matches based on 'partial truth', i.e., by identifying probable rather than exact matches. We evaluate the impact of this approach on memory usage and performance on Knight's Landing and Skylake processors. At scale (262,144 Intel Xeon Phi cores), the method shows up to 1.13 GiB of memory savings per node in the MPI library, and improvement in matching time of 95.9%; smaller-scale runs show run-time improvements of up to 31.0% for full applications, and up to 6.1% for optimized proxy applications.
Matthew G. F. Dosanjh, Whit Schonbein, Ryan E. Grant, Patrick G. Bridges, S. Mahdieh Ghazimirsaeed, Ahmad Afsahi
CCGRID2
2019 MPI tag matching performance on ConnectX and ARM
abstract
As we approach Exascale, message matching has increasingly become a significant factor in HPC application performance. To address this, network vendors have placed higher precedence on improving MPI message matching performance. ConnectX-5, Mellanox's new network interface card, has both hardware and software matching layers. The performance characteristics of these layers have yet to be studied under real world circumstances. In this work we offer an initial evaluation of ConnectX-5 message matching performance. To analyze this new hardware we executed a series of micro-benchmarks and applications on Astra, an ARM-based ConnectX-5 HPC system, while varying hardware and software matching parameters. The benchmark results show the ConnectX-5 is sensitive to queue depths, and that hardware message matching increases performance for applications that send messages between 1KiB and 16KiB. Furthermore, the hardware matching system was capable of matching wildcard receives without negatively impacting performance. Finally, for some applications, a significant improvement can be observed when leveraging the ConnectX-5's hardware matching.
W. Pepper Marts, Matthew G. F. Dosanjh, Whit Schonbein, Ryan E. Grant, Patrick G. Bridges
EuroMPI3
2019 INCA: in-network compute assistance
abstract
Current proposals for in-network data processing operate on data as it streams through a network switch or endpoint. Since compute resources must be available when data arrives, these approaches provide deadline-based models of execution. This paper introduces a deadline-free general compute model for network endpoints called INCA: In-Network Compute Assistance. INCA builds upon contemporary NIC offload capabilities to provide on-NIC, deadline-free, general-purpose compute capacities that can be utilized when the network is inactive. We demonstrate INCA is Turing complete, and provide a detailed design for extending existing hardware to support this model. We evaluate runtimes for a selection of kernels, including several optimizations, and show INCA can provide up to a 11% speedup for applications with minimal code modifications and between 25% to 37% when applications are optimized for INCA.
Whit Schonbein, Ryan E. Grant, Matthew G. F. Dosanjh, Dorian C. Arnold
SC1
2019 Using simulation to examine the effect of MPI message matching costs on application performance
Scott Levy, Kurt B. Ferreira, Whit Schonbein, Ryan E. Grant, Matthew G. F. Dosanjh
Parallel Comput.3
2018 Measuring Multithreaded Message Matching Misery
Whit Schonbein, Matthew G. F. Dosanjh, Ryan E. Grant, Patrick G. Bridges
Euro-Par1
2018 The Case for Semi-Permanent Cache Occupancy: Understanding the Impact of Data Locality on Network Processing
abstract
The performance critical path for MPI implementations relies on fast receive side operation, which in turn requires fast list traversal. The performance of list traversal is dependent on data-locality; whether the data is currently contained in a close-to-core cache due to its temporal locality or if its spacial locality allows for predictable pre-fetching. In this paper, we explore the effects of data locality on the MPI matching problem by examining both forms of locality. First, we explore spacial locality, by combining multiple entries into a single linked list element, we can control and modify this form of locality. Secondly, we explore temporal locality by utilizing a new technique called "hot caching", a process that creates a thread to periodically access certain data, increasing its temporal locality. In this paper, we show that by increasing data locality, we can improve MPI performance on a variety of architectures up to 4x for micro-benchmarks and up to 2x for an application.
Matthew G. F. Dosanjh, S. Mahdieh Ghazimirsaeed, Ryan E. Grant, Whit Schonbein, Michael J. Levenhagen, Patrick G. Bridges, Ahmad Afsahi
ICPP4
2012 Inspirational anchors: minimal computational models in cognitive science
abstract
In model-based science, a minimal computational model (MCM) is a computational model developed without the guidance of significant empirical evidence about the mechanism being modelled. Despite their historical and contemporary prominence in cognitive science, MCMs face serious challenges: (1) because of the lack of empirical grounding, it is hard to see how we are justified in making inferences from a model to its target and (2) if they say nothing about their targets, it seems that their utility is limited to the articulation of mere logical possibilities. In this article, I scrutinise this challenge by surveying and rejecting some alternative accounts of the epistemological role of MCMs. I argue that these models are best viewed as fulcra upon which additional research is leveraged. In support, I draw connections between cognitive and economic modelling, and sketch how a prominent account of the latter can be extended to cover the former.
Whit Schonbein
J. Exp. Theor. Artif. Intell.1