VLDB 2026 Research / reviewers in the wild / expert
Patrick G. Bridges
dblp:97/3677
· DBLP profile ↗
56ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0003-4801-0390ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 42 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 5 · 3 first-author · 1 since 2021Computer networks · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scaling Laws for the Workload Throughput of Emerging Heterogeneous ClustersabstractNext-generation HPC clusters are evolving into highly heterogeneous systems that integrate traditional computing resources with emerging accelerator technologies such as quantum processors, neuromorphic units, dataflow architectures, and specialized AI accelerators within a unified infrastructure. These advanced systems enable workloads to dynamically utilize different accelerators during various computation phases, creating complex execution patterns. The performance of the workloads can therefore be impacted by many factors, including how the accelerators are shared, their utilization, and their placement within the system. Moreover, effects such as the system and network state due to the overall system load can significantly impact the job completion rate. Understanding, identifying, and quantifying the impact of the most critical factors (e.g., the number of allocated accelerators) will help decide the investment decisions for accelerator acquisition and deployment that can improve the overall system throughput. This paper extensively studies these complex interactions among advanced accelerators within an HPC cluster and various workloads. We introduce a novel analytical model which predicts the speedup of a workload given an accelerator/system configuration. This model can be used to quantify the effect of augmenting additional accelerators on job performance running on an HPC cluster. We validate the model using both simulated and real environments. Akhil Alasandagutti, Joshua Suetterlein, Jesun Sahariar Firoz, Stephen J. Young, Joseph B. Manzano, Jason R. Stewart, Patrick G. Bridges, Trilce Estrada, Kevin J. Barker |
CCGrid | 7 |
| 2025 | Grey-Box Machine Learning Prediction of Parallel Application ScalingabstractAccurate prediction of parallel application performance in HPC systems is essential for efficient resource allocation and system design. Classical performance models estimate of speedup based on theoretical assumptions, but their applicability is limited by parameter estimation, data acquisition, and real-world system issues such as latency and network congestion. This paper describes performance prediction using classical performance models boosted by a trainable machine learning framework. Domain-informed machine-learning models estimate the overhead of an application for a given problem size and resource configuration as a coefficient of the estimated speedup provided by performance laws. We evaluate this approach on two HPC mini-applications and two full applications with varying patterns of computation and communication and also evaluate the prediction accuracy on runs with varying processors-per-node configurations. Our results show that this method significantly improves the accuracy of performance predictions over standard analytical models and black-box regressors, while remaining robust even with limited training data. Akhil Alasandagutti, Patrick G. Bridges, Trilce Estrada |
HiPC | 2 |
| 2025 | Performance Analysis of Open MPI on AMR Applications over Slingshot-11
Maxim Moraru, Howard Pritchard, Derek Schafer, Galen M. Shipman, Patrick G. Bridges |
EuroMPI | 5 |
| 2025 | Measuring Thread Timing to Assess the Feasibility of Early-Bird Message Delivery Across Systems and ScalesabstractABSTRACT Early‐bird communication is a communication/computation overlap technique that leverages fine‐grained communication to improve application run‐time. Communication is divided such that each individual thread can initiate transmission of its portion of the data upon completion rather than waiting for a dedicated communication phase. The benefit of early‐bird communication depends on the completion timing of the individual threads: On the one hand, if all threads are complete at nearly the same time, the overheads of sending multiple messages will accumulate, leading to performance that is worse than if a single message had been sent. On the other hand, if thread completions are spread out in time, those that complete earlier can send data while others continue working, leading to performance that is better than if a single message had been sent. The challenge is that the completion times are currently unknown and can vary based on application, problem size, system software, and underlying hardware. In this paper, we address this lacuna by measuring and evaluating the potential overlap afforded by early‐bird communication for a selection of proxy applications. These measurements help us understand whether a given application could benefit from early‐bird communication. We present our technique for gathering this data and evaluate data collected from three proxy applications: MiniFE, MiniMD, and MiniQMC. Each application is run on three systems with distinct CPU architectures and strong scales across three run sizes. To characterize the behavior of these workloads, we study the trends of thread timings at both a macro level, across all threads across all runs of an application, and a micro level, that is, within a single process of a single run. We observe that our tested applications exhibit significantly different thread arrival distributions. The machine used had a significant impact, with the window of potential overlap varying by as much as an order of magnitude. W. Pepper Marts, Matthew G. F. Dosanjh, Whit Schonbein, Scott Levy, Patrick G. Bridges |
Concurr. Comput. Pract. Exp. | 5 |
| 2024 | CMB: A Configurable Messaging Benchmark to Explore Fine-Grained CommunicationabstractModern communication APIs provide increased ability to specify when, where, and how to send data between processes. One recent innovation is fine-grained communication, where processes are able to send subsets of data as it is ready rather than waiting for the entirety of the data to be completed. Allowing data to be sent when it is ready increases opportunities for overlapping communication and computation. However, with multiple fine-grained, thread-safe interfaces, the task of optimizing an application’s peer-to-peer fine-grained communication is complex. In this paper, we present the Configurable Messaging Benchmark (CMB), a tool for evaluating the application impact of fine-grained communication. Using the CMB we perform a case study to measure the impact of different fine-grained implementations on a variety of realistic application profiles. Initial results reveal a large optimization space ranging from potential speedups as high as 52.97% to slowdowns as high as 289.55% relative to bulk-synchronous MPI message passing. W. Pepper Marts, Donald A. Kruse, Matthew G. F. Dosanjh, Whit Schonbein, Scott Levy, Patrick G. Bridges |
CCGrid | 6 |
| 2024 | Quantifying and Modeling Irregular MPI CommunicationabstractMany modern scientific applications have communication patterns where both the number of communication partners and amount of data transmitted between process pairs vary significantly and change over time. This work describes an approach to measure and model the behavior of these irregular, dynamic MPI communication patterns on modern high-performance computing systems. Specifically, this approach quantifies communication behavior using a small number of stochastic random variables that capture key features of irregular communication patterns, and estimates the distributions of these variables either parameterically or empirically. This work then demonstrates that the collected parameters and their distributions can be used to measure and model the communication performance of several MPI applications. It also presents a synthetic benchmark that uses these distributions to recreate statistically similar communication patterns. This approach provides a lightweight method to reproduce communication patterns of a variety of applications with minimal overhead while gaining additional insights into the performance and characteristics of various irregular communication patterns. Carson Woods, Derek Schafer, Patrick G. Bridges, Anthony Skjellum |
CCGrid | 3 |
| 2024 | Optimizing Neighbor Collectives with Topology ObjectsabstractMany HPC applications implement non-cartesian neighbor data exchanges using MPI point-to-point operations rather than utilizing native MPI neighbor collective methods. Each application must therefore implement their own commu-nication optimizations, rather than leveraging any optimizations that could be provided by MPI. While an interface for such optimizations is provided within MPI through neighborhood collectives, applications avoid these methods due to the lack of performance optimizations within them along with large costs associated with graph communicator formation. This paper presents a novel approach for creating local, non-cartesian topol-ogy objects that provides finer control over the aforementioned setup costs. Any additional setup costs, such as initializing per-iteration optimizations, can then be deferred until additional information is available, such as within persistent initialization calls. This paper describes our implementation within an MPI extension library and demonstrates the effectiveness of our approach in simple benchmarks and real-world applications. Gerald Collom, Derek Schafer, Amanda Bienz, Patrick G. Bridges, Galen M. Shipman |
CLUSTER | 4 |
| 2024 | A More Scalable Sparse Dynamic Data ExchangeabstractParallel architectures are continually increasing in performance and scale while underlying algorithmic infrastruc-ture often fails to take full advantage of available compute power. Within the context of MPI, irregular communication patterns create bottlenecks in parallel applications. One common bottleneck is the sparse dynamic data exchange, often required when forming communication patterns within applications. There is a large variety of approaches for these dynamic exchanges, with optimizations implemented directly in parallel applications. This paper proposes a novel API within an MPI eXtension library, allowing applications to utilize the variety of provided optimizations for sparse dynamic data exchange methods. Fur-ther, the paper presents novel locality-aware sparse dynamic data exchange algorithms. Finally, performance results show locality-aware approaches achieve up to 128x over existing approaches when exchanging only pattern of communication, and up to 54x when exchanging data to be communicated as well. Andrew Geyko, Gerald Collom, Derek Schafer, Patrick G. Bridges, Amanda Bienz |
HiPC | 4 |
| 2024 | Understanding GPU Triggering APIs for MPI+X Communication
Patrick G. Bridges, Anthony Skjellum, Evan Drake Suggs, Derek Schafer, Purushotham V. Bangalore |
EuroMPI | 1 |
| 2023 | Evaluating the Viability of LogGP for Modeling MPI Performance with Non-contiguous Datatypes on Modern ArchitecturesabstractModern architectures and communication systems software include complex hardware, communication abstractions, and optimizations that make their performance difficult to measure, model, and understand. This paper examines the ability of modified versions of the existing Netgauge communication performance measurement tool and LogGOPS performance model to accurately characterize communication behavior of modern hardware, MPI abstractions, and implementations. This includes analyzing their ability to model both GPU-aware communication in different MPI implementations and quantifying the performance characteristics of different approaches to non-contiguous data communication on modern GPU systems. This paper also applies these techniques to quantify the performance of different implementations and optimization approaches to non-contiguous data communication on a variety of systems, demonstrating that modern communication system design approaches can result in widely-varying and difficult-to-predict performance variation, even within the same hardware/communication software combination. Nicholas H. Bacon, Patrick G. Bridges, Scott Levy, Kurt B. Ferreira, Amanda Bienz |
EuroMPI | 2 |
| 2021 | MiniMod: A Modular Miniapplication Benchmarking Framework for HPCabstractThe HPC application community has proposed many new application communication structures, middleware interfaces, and communication models to improve HPC application performance. Modifying proxy applications is the standard practice for the evaluation of these novel methodologies. Currently, this requires the creation of a new version of the proxy application for each combination of the approach being tested. In this article, we present a modular proxy-application framework, MiniMod, that enables evaluation of a combination of independently written computation kernels, data transfer logic, communication access, and threading libraries. MiniMod is designed to allow rapid development of individual modules which can be combined at runtime. Through MiniMod, developers only need a single implementation to evaluate application impact under a variety of scenarios.We demonstrate the flexibility of MiniMod’s design by using it to implement versions of a heat diffusion kernel and the miniFE finite element proxy application, along with a variety of communication, granularity, and threading modules. We examine how changing communication libraries, communication granularities, and threading approaches impact these applications on an HPC system. These experiments demonstrate that MiniMod can rapidly improve the ability to assess new middleware techniques for scientific computing applications and next-generation hardware platforms. W. Pepper Marts, Matthew G. F. Dosanjh, Scott Levy, Whit Schonbein, Ryan E. Grant, Patrick G. Bridges |
CLUSTER | 6 |
| 2021 | SAMPRA: Scalable Analysis, Management, Protection of Research ArtifactsabstractThis paper describes SAMPRA, a framework for supporting effective research on sensitive data being deployed at the University of New Mexico. SAMPRA and its associated implementation are designed to support the needs of a diverse set of use-cases from researchers across different disciplines at UNM, including from clinical neurosciences, forensic anthropology, and community health. From these use-cases, we identified a set of common unaddressed demands when handling data with privacy/protection requirements, particularly collaborative research projects, interfacing with scientific instruments, and full-lifecycle management of sensitive data. To properly address and accelerate research projects with these needs, SAMPRA a) integrates privacy-preserving storage and data transfer systems with data-centric virtual environments, and b) supports effective researcher use of the system through active collaboration between local IT personnel, campus enterprise IT service providers, and campus data librarians by defining clear roles with associated personnel. By doing so, SAMPRA seeks to meet the needs of research on sensitive data across the entire data lifecycle and avoid the pitfalls of generic “one-size-fits-all” services. Patrick G. Bridges, Zeinab Akhavan, Jonathan Wheeler, Hussein Al-Azzawi, Orlando Albillar, Grace Faustino |
e-Science | 1 |
| 2020 | Tail queues: A multi-threaded matching architectureabstractSummary As we approach exascale, computational parallelism will have to drastically increase in order to meet throughput targets. Many‐core architectures have exacerbated this problem by trading reduced clock speeds, core complexity, and computation throughput for increasing parallelism. This presents two major challenges for communication libraries such as MPI: the library must leverage the performance advantages of thread level parallelism and avoid the scalability problems associated with increasing the number of processes to that scale. Hybrid programming models, such as MPI+X, have been proposed to address these challenges. MPI THREAD MULTIPLE is MPI's thread safe mode. While there has been work to optimize it, it largely remains non‐performant in most implementations. While current applications avoid MPI multithreading due to performance concerns, it is expected to be utilized in future applications. One of the major synchronous data structures required by MPI is the matching engine. In this paper, we present a parallel matching algorithm that can improve MPI matching for multithreaded applications. We then perform a feasibility study to demonstrate the performance benefit of the technique. Matthew G. F. Dosanjh, Ryan E. Grant, Whit Schonbein, Patrick G. Bridges |
Concurr. Comput. Pract. Exp. | 4 |
| 2019 | Fuzzy Matching: Hardware Accelerated MPI Communication MiddlewareabstractContemporary parallel scientific codes often rely on message passing for inter-process communication. However, inefficient coding practices or multithreading (e.g., via MPI_THREAD_MULTIPLE) can severely stress the underlying message processing infrastructure, resulting in potentially un-acceptable impacts on application performance. In this article, we propose and evaluate a novel method for addressing this issue: 'Fuzzy Matching'. This approach has two components. First, it exploits the fact most server-class CPUs include vector operations to parallelize message matching. Second, based on a survey of point-to-point communication patterns in representative scientific applications, the method further increases parallelization by allowing matches based on 'partial truth', i.e., by identifying probable rather than exact matches. We evaluate the impact of this approach on memory usage and performance on Knight's Landing and Skylake processors. At scale (262,144 Intel Xeon Phi cores), the method shows up to 1.13 GiB of memory savings per node in the MPI library, and improvement in matching time of 95.9%; smaller-scale runs show run-time improvements of up to 31.0% for full applications, and up to 6.1% for optimized proxy applications. Matthew G. F. Dosanjh, Whit Schonbein, Ryan E. Grant, Patrick G. Bridges, S. Mahdieh Ghazimirsaeed, Ahmad Afsahi |
CCGRID | 4 |
| 2019 | Workflows for Performance Predictable and Reproducible HPC ApplicationsabstractThis poster presents an HPC application workflow system whose goal is to provide verifiably-reproducible HPC application performance. This system combines existing container, experiment, and data management techniques with HPC performance models, allowing it to both maximize performance reproducibility and inform users when application performance deviates from what should be expected even when running at scales or for lengths of time at which the application had never run. Keira Haskins, Quincy Wofford, Patrick G. Bridges |
CLUSTER | 3 |
| 2019 | MPI tag matching performance on ConnectX and ARMabstractAs we approach Exascale, message matching has increasingly become a significant factor in HPC application performance. To address this, network vendors have placed higher precedence on improving MPI message matching performance. ConnectX-5, Mellanox's new network interface card, has both hardware and software matching layers. The performance characteristics of these layers have yet to be studied under real world circumstances. In this work we offer an initial evaluation of ConnectX-5 message matching performance. To analyze this new hardware we executed a series of micro-benchmarks and applications on Astra, an ARM-based ConnectX-5 HPC system, while varying hardware and software matching parameters. The benchmark results show the ConnectX-5 is sensitive to queue depths, and that hardware message matching increases performance for applications that send messages between 1KiB and 16KiB. Furthermore, the hardware matching system was capable of matching wildcard receives without negatively impacting performance. Finally, for some applications, a significant improvement can be observed when leveraging the ConnectX-5's hardware matching. W. Pepper Marts, Matthew G. F. Dosanjh, Whit Schonbein, Ryan E. Grant, Patrick G. Bridges |
EuroMPI | 5 |
| 2018 | Measuring Multithreaded Message Matching Misery
Whit Schonbein, Matthew G. F. Dosanjh, Ryan E. Grant, Patrick G. Bridges |
Euro-Par | 4 |
| 2018 | The Case for Semi-Permanent Cache Occupancy: Understanding the Impact of Data Locality on Network ProcessingabstractThe performance critical path for MPI implementations relies on fast receive side operation, which in turn requires fast list traversal. The performance of list traversal is dependent on data-locality; whether the data is currently contained in a close-to-core cache due to its temporal locality or if its spacial locality allows for predictable pre-fetching. In this paper, we explore the effects of data locality on the MPI matching problem by examining both forms of locality. First, we explore spacial locality, by combining multiple entries into a single linked list element, we can control and modify this form of locality. Secondly, we explore temporal locality by utilizing a new technique called "hot caching", a process that creates a thread to periodically access certain data, increasing its temporal locality. In this paper, we show that by increasing data locality, we can improve MPI performance on a variety of architectures up to 4x for micro-benchmarks and up to 2x for an application. Matthew G. F. Dosanjh, S. Mahdieh Ghazimirsaeed, Ryan E. Grant, Whit Schonbein, Michael J. Levenhagen, Patrick G. Bridges, Ahmad Afsahi |
ICPP | 6 |
| 2018 | Improving MPI Multi-threaded RMA Communication PerformanceabstractOne-sided communication is crucial to enabling communication concurrency. As core counts have increased, particularly with many-core architectures, one-sided (RMA) communication has been proposed to address the ever increasing contention at the network interface. The difficulty in using one-sided (RMA) communication with MPI is that the performance of MPI implementations using RMA with multiple concurrent threads is not well understood. Past studies have been done using MPI RMA in combination with multi-threading (RMA-MT) but they have been performed on older MPI implementations lacking RMA-MT optimizations. In addition prior work has only been done at smaller scale (<=512 cores). Nathan T. Hjelm, Matthew G. F. Dosanjh, Ryan E. Grant, Taylor L. Groves, Patrick G. Bridges, Dorian C. Arnold |
ICPP | 5 |
| 2018 | An evaluation of the state of time synchronization on leadership class supercomputersabstractSummary We present a detailed examination of time agreement characteristics for nodes within extreme‐scale parallel computers. Using a software tool we introduce in this paper, we quantify attributes of clock skew among nodes in three representative high‐performance computers sited at three national laboratories. Our measurements detail the statistical properties of time agreement among nodes and how time agreement drifts over typical application execution durations. We discuss the implications of our measurements, why the current state of the field is inadequate, and propose strategies to address observed shortcomings. Terry R. Jones, George Ostrouchov, Gregory A. Koenig, Oscar H. Mondragon, Patrick G. Bridges |
Concurr. Comput. Pract. Exp. | 5 |
| 2017 | Evaluating the Viability of Using Compression to Mitigate Silent Corruption of Read-Mostly Application DataabstractAggregating millions of hardware components to construct an exascale computing platform will pose significant resilience challenges. In addition to slowdowns associated with detected errors, silent errors are likely to further degrade application performance. Moreover, silent data corruption (SDC) has the potential to undermine the integrity of the results produced by important scientific applications. In this paper, we propose an application-independent mechanism to efficiently detect and correct SDC in read-mostly memory, where SDC may be most likely to occur. We use memory protection mechanisms to maintain compressed backups of application memory. We detect SDC by identifying changes in memory contents that occur without explicit write operations. We demonstrate that, for several applications, our approach can potentially protect a significant fraction of application memory pages from SDC with modest overheads. Moreover, our proposed technique can be straightforwardly combined with many other approaches to provide a significant bulwark against SDC. Scott Levy, Kurt B. Ferreira, Patrick G. Bridges |
CLUSTER | 3 |
| 2016 | RMA-MT: A Benchmark Suite for Assessing MPI Multi-threaded RMA PerformanceabstractReaching Exascale will require leveraging massive parallelism while potentially leveraging asynchronous communication to help achieve scalability at such large levels of concurrency. MPI is a good candidate for providing the mechanisms to support communication at such large scales. Two existing MPI mechanisms are particularly relevant to Exascale: multi-threading, to support massive concurrency, and Remote Memory Access (RMA), to support asynchronous communication. Unfortunately, multi-threaded MPI RMA code has not been extensively studied. Part of the reason for this is that no public benchmarks or proxy applications exist to assess its performance. The contributions of this paper are the design and demonstration of the first available proxy applications and micro-benchmark suite for multi-threaded RMA in MPI, a study of multi-threaded RMA performance of different MPI implementations, and an evaluation of how these benchmarks can be used to test development for both performance and correctness. Matthew G. F. Dosanjh, Taylor L. Groves, Ryan E. Grant, Ron Brightwell, Patrick G. Bridges |
CCGrid | 5 |
| 2016 | Scheduling In-Situ Analytics in Next-Generation ApplicationsabstractNext-generation applications increasingly rely on in situ analytics to guide computation, reduce the amount of I/O performed, and perform other important tasks. Scheduling where and when to run analytics is challenging, however. This paper quantifies the costs and benefits of different approaches to scheduling applications and analytics on nodes in large-scale applications, including space sharing, uncoordinated time sharing, and gang scheduled time sharing. Oscar H. Mondragon, Patrick G. Bridges, Scott Levy, Kurt B. Ferreira, Patrick M. Widener |
CCGrid | 2 |
| 2016 | How I Learned to Stop Worrying and Love In Situ Analytics: Leveraging Latent Synchronization in MPI Collective AlgorithmsabstractScientific workloads running on current extreme-scale systems routinely generate tremendous volumes of data for postprocessing. This data movement has become a serious issue due to its energy cost and the fact that I/O bandwidths have not kept pace with data generation rates. In situ analytics is an increasingly popular alternative in which post-simulation processing is embedded into an application, running as part of the same MPI job. This can reduce data movement costs but introduces a new potential source of interference for the application. Using a validated simulation-based approach, we investigate how best to mitigate the interference from time-shared in situ tasks for a number of key extreme-scale workloads. This paper makes a number of contributions. First, we show that the independent scheduling of in situ analytics tasks can significantly degradation application performance, with slowdowns exceeding 1000%. Second, we demonstrate that the degree of synchronization found in many modern collective algorithms is sufficient to significantly reduce the overheads of this interference to less than 10% in most cases. Finally, we show that many applications already frequently invoke collective operations that use these synchronizing MPI algorithms. Therefore, the syncronization introduced by these MPI collective algorithms can be leveraged to efficiently schedule analytics tasks with minimal changes to existing applications. This paper provides critical analysis and guidance for MPI users and developers on the importance of scheduling in situ analytics tasks. It shows the degree of synchronization needed to mitigate the performance impacts of these time-shared coupled codes and demonstrates how that synchronization can be realized in an extreme-scale environment using modern collective algorithms. Scott Levy, Kurt B. Ferreira, Patrick M. Widener, Patrick G. Bridges, Oscar H. Mondragon |
EuroMPI | 4 |
| 2016 | Improving application resilience to memory errors with lightweight compressionabstractIn next-generation extreme-scale systems, application performance will be limited by memory performance characteristics. The first exascale system is projected to contain many petabytes of memory. In addition to the sheer volume of the memory required, device trends, such as shrinking feature sizes and reduced supply voltages, have the potential to increase the frequency of memory errors. As a result, resilience to memory errors is a key challenge. In this paper, we evaluate the viability of using memory compression to repair detectable uncorrectable errors (DUEs) in memory. We develop a software library, evaluate its performance and demonstrate that it is able to significantly compress memory of HPC applications. Further, we show that exploiting compressed memory pages to correct memory errors can significantly improve application performance on next-generation systems. Scott Levy, Kurt B. Ferreira, Patrick G. Bridges |
SC | 3 |
| 2016 | Understanding performance interference in next-generation HPC systemsabstractNext-generation systems face a wide range of new potential sources of application interference, including resilience actions, system software adaptation, and in situ analytics programs. In this paper, we present a new model for analyzing the performance of bulk-synchronous HPC applications based on the use of extreme value theory. After validating this model against both synthetic and real applications, the paper then uses both simulation and modeling techniques to profile next-generation interference sources and characterize their behavior and performance impact on a selection of HPC benchmarks, mini-applications, and applications. Lastly, this work shows how the model can be used to understand how current interference mitigation techniques in multi-processors work. Oscar H. Mondragon, Patrick G. Bridges, Scott Levy, Kurt B. Ferreira, Patrick M. Widener |
SC | 2 |
| 2015 | Re-evaluating Network Onload vs. Offload for the Many-Core EraabstractThis paper explores the trade-offs between on-loaded versus offloaded network stack processing for systems with varying CPU frequencies. This study explores the differences of onload and offload using experiments run at different DVFS settings to change the frequency, while measuring performance and power. This allows for a quantitative comparison of the the performance and power and trade-offs between onload and offload cards, with a wide range of CPU performances. The results show that there is often a significant performance increase in using offloaded cards especially at lower CPU frequencies, with only a small increase in power usage. This study also uses MPI profiling to analyze why some applications see a larger benefit than others. This paper's contributions are an analytical, quantitative analysis of the trade-offs between onload and offload. While there has been debate to this question, this is the first, to the authors' knowledge, analytical evaluation of the performance difference. The range of frequencies analyzed give insight on how this MPI might perform on different architectures, such as the low frequency, many-core CPUs. Finally, the power measurements allow for the study to provide further depth in the analysis. Matthew G. F. Dosanjh, Ryan E. Grant, Patrick G. Bridges, Ron Brightwell |
CLUSTER | 3 |
| 2014 | Characterizing the Impact of Rollback Avoidance at Extreme-Scale: A Modeling ApproachabstractResilience to failure is a key concern for next-generation high-performance computing systems. The dominant fault tolerance mechanism, coordinated checkpoint/restart, is projected to no longer be a viable option on these systems due to its predicted overheads. Rollback avoidance has the potential to prolong the viability of coordinated checkpoint/restart by allowing an application to make meaningful forward progress, perhaps with degraded performance, despite the occurrence or imminence of a failure. In this paper, we present two general analytic models for the performance of rollback avoidance techniques and validate these models against the performance of existing rollback avoidance techniques. We then use these models to evaluate the applicability of rollback avoidance for next-generation exascale systems. This includes analysis of exascale system design questions such as: (1) how effective must an application-specific rollback avoidance technique be to usefully augment checkpointing in an exascale system? (2) when is rollback avoidance on its own a viable alternative to coordinated checkpointing? and (3) how do rollback avoidance techniques and system characteristics interact to influence application performance? Scott Levy, Kurt B. Ferreira, Patrick G. Bridges |
ICPP | 3 |
| 2014 | Accelerating incremental checkpointing for extreme-scale computing
Kurt B. Ferreira, Rolf Riesen, Patrick G. Bridges, Dorian C. Arnold, Ron Brightwell |
Future Gener. Comput. Syst. | 3 |
| 2013 | Virtual TCP offload: optimizing ethernet overlay performance on advanced interconnects
Patrick G. Bridges, Jack Lange, Peter A. Dinda |
HPDC | 2 |
| 2012 | VNET/P: bridging the cloud and high performance computing through fast overlay networkingabstractIt is now possible to allow VMs hosting HPC applications to seamlessly bridge distributed cloud resources and tightly-coupled supercomputing and cluster resources. However, to achieve the application performance that the tightly-coupled resources are capable of, it is important that the overlay network not introduce significant overhead relative to the native hardware, which is not the case for current user-level tools, including our own existing VNET/U system. In response, we describe the design, implementation, and evaluation of a layer 2 virtual networking system that has negligible latency and bandwidth overheads in 1--10 Gbps networks. Our system, VNET/P, is directly embedded into our publicly available Palacios virtual machine monitor (VMM). VNET/P achieves native performance on 1 Gbps Ethernet networks and very high performance on 10 Gbps Ethernet networks and InfiniBand. The NAS benchmarks generally achieve over 95% of their native performance on both 1 and 10 Gbps. These results suggest it is feasible to extend a software-based overlay network designed for computing at wide-area scales into tightly-coupled environments. Lei Xia 0001, Jack Lange, Peter A. Dinda, Patrick G. Bridges |
HPDC | 6 |
| 2012 | On the Viability of Compression for Reducing the Overheads of Checkpoint/Restart-Based Fault ToleranceabstractThe increasing size and complexity of high performance computing (HPC) systems have led to major concerns over fault frequencies and the mechanisms necessary to tolerate these faults. Previous studies have shown that state-of-the-field checkpoint/restart mechanisms will not scale sufficiently for future generation systems. Therefore, optimizations that reduce checkpoint overheads are necessary to keep checkpoint/restart mechanisms effective. In this work, we demonstrate that checkpoint data compression is a feasible mechanism for reducing checkpoint commit latencies and storage overheads. Leveraging a simple model for checkpoint compression viability, we show: (1) checkpoint data compression is feasible for many types of scientific applications expected to run on extreme scale systems, (2) checkpoint compression viability scales with checkpoint size, (3) user-level versus system-level checkpoints bears little impact on checkpoint compression viability, and (4) checkpoint compression viability scales with application process count. Lastly, we describe the impact that checkpoint compression might have on future generation extreme scale systems. Dewan Ibtesham, Dorian C. Arnold, Patrick G. Bridges, Kurt B. Ferreira, Ron Brightwell |
ICPP | 3 |
| 2012 | Optimizing overlay-based virtual networking through optimistic interrupts and cut-through forwardingabstractOverlay-based virtual networking provides a powerful model for realizing virtual distributed and parallel computing systems with strong isolation, portability, and recoverability properties. However, in extremely high throughput and low latency networks, such overlays can suffer from bandwidth and latency limitations, which is of particular concern if we want to apply the model in HPC environments. Through careful study of an existing very high performance overlay-based virtual network system, we have identified two core issues limiting performance: delayed and/or excessive virtual interrupt delivery into guests, and copies between host and guest data buffers done during encapsulation. We respond with two novel optimizations: optimistic, timer-free virtual interrupt injection, and zero-copy cut-through data forwarding. These optimizations improve the latency and bandwidth of the overlay network on 10 Gbps interconnects, resulting in near-native performance for a wide range of microbenchmarks and MPI application benchmarks. Lei Xia 0001, Patrick G. Bridges, Peter A. Dinda, Jack Lange |
SC | 3 |
| 2012 | Alleviating scalability issues of checkpointing protocolsabstractCurrent fault tolerance protocols are not sufficiently scalable for the exascale era. The most-widely used method, coordinated checkpointing, places enormous demands on the I/O subsystem and imposes frequent synchronizations. Uncoordinated protocols use message logging which introduces message rate limitations or undesired memory and storage requirements to hold payload and event logs. In this paper we propose a combination of several techniques, namely coordinated checkpointing, optimistic message logging, and a protocol that glues them together. This combination eliminates some of the drawbacks of each individual approach and proves to be an alternative for many types of exascale applications. We evaluate performance and scaling characteristics of this combination using simulation and a partial implementation. While not a universal solution, the combined protocol is suitable for a large range of existing and future applications that use coordinated checkpointing and enhances their scalability. Rolf Riesen, Kurt B. Ferreira, Dilma Da Silva, Pierre Lemarinier, Dorian C. Arnold, Patrick G. Bridges |
SC | 6 |
| 2011 | Exploiting MISD Performance Opportunities in Multi-core Systems
Patrick G. Bridges, Donour Sizemore, Scott Levy |
HotOS | 1 |
| 2011 | libhashckpt: Hash-Based Incremental Checkpointing Using GPU's
Kurt B. Ferreira, Rolf Riesen, Ron Brightwell, Patrick G. Bridges, Dorian C. Arnold |
EuroMPI | 4 |
| 2011 | Evaluating the viability of process replication reliability for exascale systemsabstractAs high-end computing machines continue to grow in size, issues such as fault tolerance and reliability limit application scalability. Current techniques to ensure progress across faults, like checkpoint-restart, are increasingly problematic at these scales due to excessive overheads predicted to more than double an application's time to solution. Replicated computing techniques, particularly state machine replication, long used in distributed and mission critical systems, have been suggested as an alternative to checkpoint-restart. In this paper, we evaluate the viability of using state machine replication as the primary fault tolerance mechanism for upcoming exascale systems. We use a combination of modeling, empirical analysis, and simulation to study the costs and benefits of this approach in comparison to checkpoint/restart on a wide range of system parameters. These results, which cover different failure distributions, hardware mean time to failures, and I/O bandwidths, show that state machine replication is a potentially useful technique for meeting the fault tolerance demands of HPC applications on future exascale platforms. Kurt B. Ferreira, Jon Stearley, James H. Laros III, Ron A. Oldfield, Kevin T. Pedretti, Ron Brightwell, Rolf Riesen, Patrick G. Bridges, Dorian C. Arnold |
SC | 8 |
| 2011 | Minimal-overhead virtualization of a large scale supercomputerabstractVirtualization has the potential to dramatically increase the usability and reliability of high performance computing (HPC) systems. However, this potential will remain unrealized unless overheads can be minimized. This is particularly challenging on large scale machines that run carefully crafted HPC OSes supporting tightly-coupled, parallel applications. In this paper, we show how careful use of hardware and VMM features enables the virtualization of a large-scale HPC system, specifically a Cray XT4 machine, with < = 5% overhead on key HPC applications, microbenchmarks, and guests at scales of up to 4096 nodes. We describe three techniques essential for achieving such low overhead: passthrough I/O, workload-sensitive selection of paging mechanisms, and carefully controlled preemption. These techniques are forms of symbiotic virtualization, an approach on which we elaborate. Jack Lange, Kevin T. Pedretti, Peter A. Dinda, Patrick G. Bridges, Chang Bae, Philip Soltero, Alex Merritt |
VEE | 4 |
| 2011 | Inferring users' online activities through traffic analysisabstractTraffic analysis may threaten user privacy, even if the traffic is encrypted. In this paper, we use IEEE 802.11 wireless local area networks (WLANs) as an example to show that inferring users' online activities accurately by traffic analysis without the administrator's privilege is possible during very short periods (e.g., a few seconds). The online activities we investigated include web browsing, chatting, online gaming, downloading, uploading and video watching, etc. We implement a hierarchical classification system based on machine learning algorithms to discover what a user is doing on his/her computer. Furthermore, we conduct experiments in different network environments (e.g., at home, on university campus, and in public areas) with different application scenarios to evaluate the performance of the classification system. Results show that our system can distinguish different online applications on the accuracy of about 80% in 5 seconds and over 90% accuracy if the eavesdropping lasts for 1 minute. Fan Zhang 0019, Wenbo He 0003, Xue (Steve) Liu, Patrick G. Bridges |
WISEC | 4 |
| 2010 | The Impact of System Design Parameters on Application Noise SensitivityabstractOperating system noise, or “jitter,” is a key limiter of application scalability in high end computing systems. Several studies have attempted to quantify the sources and effects of system interference, though few of these studies show the influence that architectural and system characteristics have on the impact of OS noise at scale. In this paper, we examine the impact of three such system properties: platform balance, “noisy” node distribution, and non-blocking collective operations. Using a previouslydeveloped noise injection tool, we explore how the impact of noise varies with these platform characteristics. We provide detailed performance results that indicate that a system with relatively less network bandwidth is able to absorb more noise than a system with more network bandwidth. Our results also show that application performance can be significantly degraded by only a subset of noisy nodes. Furthermore, the placement of the noisy nodes is also important, especially for applications that make substantial use of collective communication operations that are tree-based. Lastly, performance results indicate that nonblocking collective operations have the ability to greatly mitigate the impact of OS interference. Combined, these results show that the impact of OS noise is not solely a property of application communication behavior, but is also influenced by other properties of the system architecture and system software environment. Kurt B. Ferreira, Patrick G. Bridges, Ron Brightwell, Kevin T. Pedretti |
CLUSTER | 2 |
| 2010 | Palacios and Kitten: New high performance operating systems for scalable virtualized and native supercomputingabstractPalacios is a new open-source VMM under development at Northwestern University and the University of New Mexico that enables applications executing in a virtualized environment to achieve scalable high performance on large machines. Palacios functions as a modularized extension to Kitten, a high performance operating system being developed at Sandia National Laboratories to support large-scale supercomputing applications. Together, Palacios and Kitten provide a thin layer over the hardware to support full-featured virtualized environments alongside Kitten's lightweight native environment. Palacios supports existing, unmodified applications and operating systems by using the hardware virtualization technologies in recent AMD and Intel processors. Additionally, Palacios leverages Kitten's simple memory management scheme to enable low-overhead pass-through of native devices to a virtualized environment. We describe the design, implementation, and integration of Palacios and Kitten. Our benchmarks show that Palacios provides near native (within 5%), scalable performance for virtualized environments running important parallel applications. This new architecture provides an incremental path for applications to use supercomputers, running specialized lightweight host operating systems, that is not significantly performance-compromised. Jack Lange, Kevin T. Pedretti, Trammell Hudson, Peter A. Dinda, Lei Xia 0001, Patrick G. Bridges, Andy Gocke, Steven Jaconette, Michael J. Levenhagen, Ron Brightwell |
IPDPS | 7 |
| 2009 | Using application communication characteristics to drive dynamic MPI reconfigurationabstractModern HPC applications, for example adaptive mesh refinement and multi-physics codes, have dynamic communication characteristics which result in poor performance on current MPI implementations. Current MPI implementations do not change transport protocols or allocate resources based on the application characteristics, resulting in degraded application performance. In this paper, we describe PRO-MPI, a protocol reconfiguration and optimization system for MPI that we are developing to meet the needs of dynamic modern HPC applications. PRO-MPI uses profiles of past application communication characteristics to dynamically reconfigure MPI protocol choices. We show that such dynamic reconfiguration can improve the performance of important MPI applications significantly when exact communication profiles are known. We also present preliminary data showing that profiles from past application runs with different (but related) inputs can be used to optimize the performance of later application runs. Manjunath Gorentla Venkata, Patrick G. Bridges, Patrick M. Widener |
IPDPS | 2 |
| 2009 | Instruction-level simulation of a cluster at scaleabstractInstruction-level simulation is necessary to evaluate new architectures. However, single-node simulation cannot predict the behavior of a parallel application on a supercomputer. We present a scalable simulator that couples a cycle-accurate node simulator with a supercomputer network model. Our simulator executes individual instances of IBM's Mambo PowerPC simulator on hundreds of cores. We integrated a NIC emulator into Mambo and model the network instead of fully simulating it. This decouples the individual node simulators and makes our design scalable. Edgar A. León, Rolf Riesen, Arthur B. Maccabe, Patrick G. Bridges |
SC | 4 |
| 2009 | Designing and implementing lightweight kernels for capability computingabstractAbstract In the early 1990s, researchers at Sandia National Laboratories and the University of New Mexico began development of customized system software for massively parallel ‘capability’ computing platforms. These lightweight kernels have proven to be essential for delivering the full power of the underlying hardware to applications. This claim is underscored by the success of several supercomputers, including the Intel Paragon, Intel Accelerated Strategic Computing Initiative Red, and the Cray XT series of systems, each having established a new standard for high‐performance computing upon introduction. In this paper, we describe our approach to lightweight compute node kernel design and discuss the design principles that have guided several generations of implementation and deployment. A broad strategy of operating system specialization has led to a focus on user‐level resource management, deterministic behavior, and scalable system services. The relative importance of each of these areas has changed over the years in response to changes in applications and hardware and system architecture. We detail our approach and the associated principles, describe how our application of these principles has changed over time, and provide design and performance comparisons to contemporaneous supercomputing operating systems. Copyright © 2008 John Wiley & Sons, Ltd. Rolf Riesen, Ron Brightwell, Patrick G. Bridges, Trammell Hudson, Arthur B. Maccabe, Patrick M. Widener, Kurt B. Ferreira |
Concurr. Comput. Pract. Exp. | 3 |
| 2009 | Cholla: A Framework for Composing and Coordinating Adaptations in Networked SystemsabstractThe ability of networked system software to adapt in a controlled manner to changes in the environment and requirements is crucial, but difficult to realize in complex systems with multiple interacting software layers/components. Typically, many components in a networked system implement adaptive behaviors and encapsulate their own adaptation logic (policies), making the whole system's adaptive behavior hard to analyze, coordinate, and test. This paper describes Cholla, a software architecture that separates the policy decisions of how and when adaptive components in networked systems react to their environment into separate centralized controllers that are constructed from composable rule sets. Centralizing policy decisions into controllers in Cholla facilitates the analysis, coordination, and testing of these policies, while the composable nature of these controllers allows them to be customized to changing user, application, and hardware demands. In addition to describing the architecture of Cholla, this paper also presents a Linux-based prototype implementation of this architecture that controls and coordinates adaptation policy decisions inside network protocols and multimedia applications. An experimental evaluation of this prototype demonstrates that Cholla's controller architecture enables component-based construction and customization of adaptation policies in networked systems, and that these policies can effectively control and coordinate adaptation. Patrick G. Bridges, Matti A. Hiltunen, Richard D. Schlichting |
IEEE Trans. Computers | 1 |
| 2009 | Lightweight Online Performance Monitoring and Tuning with Embedded GossipabstractUnderstanding and tuning the performance of large-scale long-running applications is difficult, with both standard trace-based and statistical methods having substantial shortcomings that limit their usefulness. This paper describes a new performance monitoring approach called Embedded Gossip (EG) designed to enable lightweight online performance monitoring and tuning. EG works by piggybacking performance information on existing messages and performing information correlation online, giving each process in a parallel application a weakly consistent global view of the behavior of the entire application. To demonstrate the viability of EG, this paper presents the design and experimental evaluation of two different online monitoring systems and an online global adaptation system driven by Embedded Gossiping. In addition, we present a metric system for evaluating the suitability of an application to EG-based monitoring and adaptation, a general architecture for implementing EG-based monitoring systems, and a modified global commit algorithm appropriate for use in EG-based global adaptation systems. Together, these results demonstrate that EG is an efficient low-overhead approach for addressing a wide range of parallel performance monitoring tasks and that results from these systems can effectively drive online global adaptation. Patrick G. Bridges, Arthur B. Maccabe |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2008 | Exploiting Latent I/O Asynchrony in Petascale Science ApplicationsabstractCurrent and emerging large-scale HPC applications face daunting I/O challenges. In existing codes, problems arise both from large data volumes and from the need to perform complex online data manipulations, including data staging, reorganization, and transformation. We describe three related techniques for enabling, encouraging, and exploiting latent I/O asynchrony in HPC applications: data taps, IOgraphs, and Metabots. Mary Payne, Patrick M. Widener, Matthew Wolf, Hasan Abbasi, Scott McManus, Patrick G. Bridges, Karsten Schwan |
eScience | 6 |
| 2008 | Characterizing application sensitivity to OS interference using kernel-level noise injectionabstractOperating system noise has been shown to be a key limiter of application scalability in high-end systems. While several studies have attempted to quantify the sources and effects of system interference using user-level mechanisms, there are few published studies on the effect of different kinds of kernel-generated noise on application performance at scale. In this paper, we examine the sensitivity of real-world, large-scale applications to a range of OS noise patterns using a kernel-based noise injection mechanism implemented in the Catamount lightweight kernel. Our results demonstrate the importance of how noise is generated, in terms of frequency and duration, and how this impact changes with application scale. For example, our results show that 2.5% net processor noise at 10,000 nodes can have no impact or can result in over a factor of 20 slowdown for the same application, depending solely on how the noise is generated. We also discuss how the characteristics of the applications we studied, for example computation/communication ratios, collective communication sizes, and other characteristics, related to their tendency to amplify or absorb noise. Finally, we discuss the implications of our findings on the design of new operating systems, middleware, and other system services for high-end parallel systems. Kurt B. Ferreira, Patrick G. Bridges, Ron Brightwell |
SC | 2 |
| 2007 | Embedded Gossip: Lightweight Online Measurement for Large-Scale ApplicationsabstractFor large-scale parallel applications, lightweight online monitoring can enable a wide range of online adaptations, including load balancing, power management, and progress monitoring. The processing and monitoring overhead of centralized global tracing techniques make them unsuitable for such tasks. Purely local tools, on the other hand, fail to provide the global information necessary for many desirable online adaptations of large-scale applications. In this paper, we describe a novel distributed online measurement method for large-scale applications called Embedded Gossip (EG). EG works by piggybacking performance information about application behavior on existing application messages and merging received information with previously known data in a fashion customized to the needs of a particular monitoring task. EG thus provides each process with both local and global views of application behavior with low overhead. To illustrate the capabilities of Embedded Gossip, we also show that it disseminates global information in a timely fashion for a wide range of monitoring tasks, including critical path profiling, workload imbalance monitoring, and progress monitoring. This global information has a wide range of potential uses, including imbalance detection for load balancing and energy management tools, progress monitoring for batch schedulers, and a wide range of other performance debugging and optimization techniques. Patrick G. Bridges, Arthur B. Maccabe |
ICDCS | 2 |
| 2007 | A configurable and extensible transport protocol
Patrick G. Bridges, Gary T. Wong, Matti A. Hiltunen, Richard D. Schlichting, Matthew J. Barrick |
IEEE/ACM Trans. Netw. | 1 |
| 2006 | Infiniband scalability in Open MPIabstractInfiniband is becoming an important interconnect technology in high performance computing. Efforts in large scale Infiniband deployments are raising scalability questions in the HPC community. Open MPI, a new open source implementation of the MPI standard targeted for production computing, provides several mechanisms to enhance Infiniband scalability. Initial comparisons with MVAPICH, the most widely used Infiniband MPI implementation, show similar performance but with much better scalability characteristics. Specifically, small message latency is improved by up to 10% in medium/large jobs and memory usage per host is reduced by as much as 300%. In addition, Open MPI provides predictable latency that is close to optimal without sacrificing bandwidth performance. Galen M. Shipman, Timothy S. Woodall, Richard L. Graham, Arthur B. Maccabe, Patrick G. Bridges |
IPDPS | 5 |
| 2005 | Online Critical Path Profiling for Parallel ApplicationsabstractOnline monitoring of parallel applications is increasingly important for techniques such as load balancing, protocol adaptation, and online anomaly detection. Unfortunately, existing online monitoring techniques only monitor individual hosts in a distributed-memory parallel application. In this paper, we show how a new monitoring technique, message-centric monitoring, can be used for online monitoring of the complete critical path in distributed-memory parallel applications. Results from an MPI-based message-centric monitoring prototype called IMPuLSE show that it has less than 3% runtime overhead, accurately measures whole-system performance as the application runs, and captures data that can be used by nodes to detect unusual system behaviors at runtime Patrick G. Bridges, Arthur B. Maccabe |
CLUSTER | 2 |
| 2003 | A Performance Comparison of Linux and a Lightweight KernelabstractIn this paper, we compare running the Linux operating system on the compute nodes of ASCI Red hardware to running a specialized, highly-optimized lightweight kernel (LWK) operating system. We have ported Linux to the compute and service nodes of the ASCI Red supercomputer, and have run several benchmarks. We present performance and scalability results for Linux compared with the LWK environment. To our knowledge, this is the first direct comparison on identical hardware of Linux and an operating system designed specifically for large-scale supercomputers. In addition to presenting these results, we discuss the limitations of both operating systems, in terms of the empirical evidence as well as other important factors. Ron Brightwell, Rolf Riesen, Keith D. Underwood, Trammell Hudson, Patrick G. Bridges, Arthur B. Maccabe |
CLUSTER | 5 |
| 2001 | Supporting Coordinated Adaption in Networked SystemsabstractSummary form only given. Our position is that the true potential of adaptation can only be realized if support is provided for more general solutions, including adaptations that span multiple hosts and multiple system components, and algorithmic adaptations that involve changing the underlying algorithms used by the system at runtime. Such a general solution must, however, address the difficult issues related to these types of adaptations. Adaptation by multiple related components, for example, must be coordinated so that these adaptations work together to implement consistent adaptation policies. Likewise, large-scale algorithmic adaptations need to be coordinated using graceful adaptation strategies in which as much normal processing as possible continues during the changeover. Here, we summarize our approach to addressing these problems in Cactus, a system for constructing highly-configurable distributed services and protocols. Patrick G. Bridges, Wen-Ke Chen, Matti A. Hiltunen, Richard D. Schlichting |
HotOS | 1 |
| 2000 | Experiences building a communication-oriented JavaOSabstractMobile code makes it easier to maintain, debug, update, and customize a system. Active networks are one of the more interesting applications of mobile code: code is injected into the nodes of a network to customize the network's functionality, such as routing, and to add new features, such as special-purpose congestion control and filtering algorithms. The challenge is to develop a communication-oriented platform for such systems. We refer to mobile code targeted at low-level, communication-oriented systems like active networks as liquid software, the key distinction being that liquid software is focused on the efficient transfer of data, not high-performance computation. To this end, we have designed and implemented Joust, which consists of a complete re-implementation of the Java virtual machine (including both the runtime system and a just-in-time compiler), running on the Scout operating system (a configurable, communication-oriented OS). The result is a configurable, high-performance platform for running liquid software. We present the results of implementing two different applications of liquid software on Joust, including a prototype architecture for active networks. Copyright © 2000 John Wiley & Sons, Ltd. John H. Hartman, Larry L. Peterson, Andy C. Bavier, Peter A. Bigot, Patrick G. Bridges, Allen Brady Montz, Rob Piltz, Todd A. Proebsting, Oliver Spatscheck |
Softw. Pract. Exp. | 5 |
| 1996 | Analysis of Techniques to Improve Protocol Processing LatencyabstractThis paper describes several techniques designed to improve protocol latency, and reports on their effectiveness when measured on a modern RISC machine employing the DEC Alpha processor. We found that the memory system---which has long been known to dominate network throughput---is also a key factor in protocol latency. As a result, improving instruction cache effectiveness can greatly reduce protocol processing overheads. An important metric in this context is the memory cycles per instructions (mCPI), which is the average number of cycles that an instruction stalls waiting for a memory access to complete. The techniques presented in this paper reduce the mCPI by a factor of 1.35 to 5.8. In analyzing the effectiveness of the techniques, we also present a detailed study of the protocol processing behavior of two protocol stacks---TCP/IP and RPC---on a modern RISC processor. David Mosberger, Larry L. Peterson, Patrick G. Bridges, Sean W. O'Malley |
SIGCOMM | 3 |