EDBT 2026 Demo / reviewers in the wild / expert
Taylor L. Groves
dblp:139/3334
· DBLP profile ↗
16ranked-venue papers
6as first author
3since 2021 · last 2024
0000-0002-7020-8881ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 6 first-author · 3 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Evaluating the potential of disaggregated memory systems for HPC applicationsabstractSummary Disaggregated memory is a promising approach that addresses the limitations of traditional memory architectures by enabling memory to be decoupled from compute nodes and shared across a data center. Cloud platforms have deployed such systems to improve overall system memory utilization, but performance can vary across workloads. High‐performance computing (HPC) is crucial in scientific and engineering applications, where HPC machines also face the issue of underutilized memory. As a result, improving system memory utilization while understanding workload performance is essential for HPC operators. Therefore, learning the potential of a disaggregated memory system before deployment is a critical step. This paper proposes a methodology for exploring the design space of a disaggregated memory system. It incorporates key metrics that affect performance on disaggregated memory systems: memory capacity, local and remote memory access ratio, injection bandwidth, and bisection bandwidth, providing an intuitive approach to guide machine configurations based on technology trends and workload characteristics. We apply our methodology to analyze thirteen diverse workloads, including AI training, data analysis, genomics, protein, fusion, atomic nuclei, and traditional HPC bookends. Our methodology demonstrates the ability to comprehend the potential and pitfalls of a disaggregated memory system and provides motivation for machine configurations. Our results show that eleven of our thirteen applications can leverage injection bandwidth disaggregated memory without affecting performance, while one pays a rack bisection bandwidth penalty and two pay the system‐wide bisection bandwidth penalty. In addition, we also show that intra‐rack memory disaggregation would meet the application's memory requirement and provide enough remote memory bandwidth. Nan Ding 0006, Pieter Maris, Hai Ah Nam, Taylor L. Groves, Muaaz Gul Awan, LeAnn Lindsey, Christopher S. Daley, Oguz Selvitopi, Leonid Oliker, Nicholas J. Wright, Samuel Williams 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2023 | Not all applications have boring communication patterns: Profiling message matching with BMMabstractSummary Message matching within MPI is an important performance consideration for applications that utilize two‐sided semantics. In this work, we present an instrumentation of the CrayMPI library that allows the collection of detailed message‐matching statistics as well as an implementation of hashed matching in software. We use this functionality to profile key DOE applications with complex communication patterns to determine under what circumstances an application might benefit from hardware offload capabilities within the NIC to accelerate message matching. We find that there are several applications and libraries that exhibit sufficiently long match list lengths to motivate a Binned Message Matching approach. Taylor L. Groves, Naveen Ravichandrasekaran, Brandon Cook 0001, Noel Keen, David Trebotich, Nicholas J. Wright, Robert Alverson, Duncan Roweth, Keith D. Underwood |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | Performance Evaluation of Adaptive Routing on Dragonfly-based Production SystemsabstractPerformance of applications in production environments can be sensitive to network congestion. Cray Aries supports adaptively routing each network packet independently based on the load or congestion encountered as a packet traverses the network. Software can dictate different routing policies, adjusting between minimal and non-minimal bias, for each posted message. We have extensively evaluated the sensitivity of the routing bias selection on application performance as well as whole system performance in both production and controlled conditions. We show that the default routing bias used in Aries-based systems is often sub-optimal and that using a higher bias towards minimal routes will not only reduce the congestion effects on the application but also will decrease the overall congestion on the network. This routing scheme results in not only improved mean performance (by up to 12%) of most production applications but also reduced run-to-run variability. Our study prompted the two supercomputing facilities (ALCF and NERSC) to change the default routing mode on their Aries-based systems. We present the substantial improvement measured in the overall congestion management and interconnect performance in production after making this change. Sudheer Chunduri, Kevin Harms, Taylor L. Groves, Peter Mendygral, Justs Zarins, Michèle Weiland, Yasaman Ghadar |
IPDPS | 3 |
| 2020 | Quantifying the impact of network congestion on application performance and network metricsabstractIn modern high-performance computing (HPC) systems, network congestion is an important factor that contributes to performance degradation. However, how network congestion impacts application performance is not fully understood. As Aries network, a recent HPC network architecture featuring a dragonfly topology, is equipped with network counters measuring packet transmission statistics on each router, these network metrics can potentially be utilized to understand network performance. In this work, by experiments on a large HPC system, we quantify the impact of network congestion on various applications' performance in terms of execution time, and we correlate application performance with network metrics. Our results demonstrate diverse impacts of network congestion: while applications with intensive MPI operations (such as HACC and MILC) suffer from more than 40% extension in their execution times under network congestion, applications with less intensive MPI operations (such as Graph500 and HPCG) are mostly not affected. We also demonstrate that a stall-to-flit ratio metric derived from Aries network counters is positively correlated with performance degradation and, thus, this metric can serve as an indicator of network congestion in HPC systems. Yijia Zhang 0002, Taylor L. Groves, Brandon Cook 0001, Nicholas J. Wright, Ayse K. Coskun |
CLUSTER | 2 |
| 2020 | The Case of Performance Variability on Dragonfly-based SystemsabstractPerformance of a parallel code running on a large supercomputer can vary significantly from one run to another even when the executable and its input parameters are left unchanged. Such variability can occur due to perturbation of the computation and/or communication in the code. In this paper, we investigate the case of performance variability arising due to network effects on supercomputers that use a dragonfly topology - specifically, Cray XC systems equipped with the Aries interconnect. We perform post-mortem analysis of network hardware counters, profiling output, job queue logs, and placement information, all gathered from periodic representative application runs. We investigate the causes of performance variability using deviation prediction and recursive feature elimination. Additionally, using time-stepped performance data of individual applications, we train machine learning models that can forecast the execution time of future time steps. Abhinav Bhatele, Jayaraman J. Thiagarajan, Taylor L. Groves, Rushil Anirudh, Staci A. Smith, Brandon Cook 0001, David K. Lowenthal |
IPDPS | 3 |
| 2020 | Hardware MPI message matching: Insights into MPI matching behavior to inform designabstractSummary This paper explores key differences of MPI match lists for several important United States Department of Energy (DOE) applications and proxy applications. This understanding is critical in determining the most promising hardware matching design for any given high‐speed network. The results of MPI match list studies for the major open‐source MPI implementations, MPICH and Open MPI, are presented, and we modify an MPI simulator, LogGOPSim, to provide match list statistics. These results are discussed in the context of several different potential design approaches to MPI matching–capable hardware. The data illustrate the requirements for different hardware designs in terms of performance and memory capacity. This paper's contributions are the collection and analysis of data to help inform hardware designers of common MPI requirements and highlight the difficulties in determining these requirements by only examining a single MPI implementation. Kurt B. Ferreira, Ryan E. Grant, Michael J. Levenhagen, Scott Levy, Taylor L. Groves |
Concurr. Comput. Pract. Exp. | 5 |
| 2019 | GPCNeT: designing a benchmark suite for inducing and measuring contention in HPC networksabstractNetwork congestion is one of the biggest problems facing HPC systems today, affecting system throughput, performance, user experience, and reproducibility. Congestion manifests as run-to-run variability due to contention for shared resources (e.g., filesystems) or routes between compute endpoints. Despite its significance, current network benchmarks fail to proxy the real-world network utilization seen on congested systems. We propose a new open-source benchmark suite called the Global Performance and Congestion Network Tests (GPCNeT) to advance the state of the practice in this area. The guiding principles used in designing GPCNeT are described and the methodology employed to maximize its utility is presented. The capabilities of GPCNeT are evaluated by analyzing results from several world's largest HPC systems, including an evaluation of congestion management on a next-generation network. The results show that systems of all technologies and scales are susceptible to congestion and this work motivates the need for congestion control in next-generation networks. Sudheer Chunduri, Taylor L. Groves, Peter Mendygral, Brian Austin, Jacob Balma, Krishna Kandalla, Kalyan Kumaran, Glenn K. Lockwood, Scott Parker, Steven Warren, Nathan Wichmann, Nicholas J. Wright |
SC | 2 |
| 2019 | Bandwidth steering in HPC using silicon nanophotonicsabstractAs bytes-per-FLOP ratios continue to decline, communication is becoming a bottleneck for performance scaling. This paper describes bandwidth steering in HPC using emerging reconfigurable silicon photonic switches. We demonstrate that placing photonics in the lower layers of a hierarchical topology efficiently changes the connectivity and consequently allows operators to recover from system fragmentation that is otherwise hard to mitigate using common task placement strategies. Bandwidth steering enables efficient utilization of the higher layers of the topology and reduces cost with no performance penalties. In our simulations with a few thousand network endpoints, bandwidth steering reduces static power consumption per unit throughput by 36% and dynamic power consumption by 14% compared to a reference fat tree topology. Such improvements magnify as we taper the bandwidth of the upper network layer. In our hardware testbed, bandwidth steering improves total application execution time by 69%, unaffected by bandwidth tapering. George Michelogiannakis, Yiwen Shen 0002, Min Yee Teh, Xiang Meng 0003, Benjamin Aivazi, Taylor L. Groves, John Shalf, Madeleine Glick, Manya Ghobadi, Larry Dennison, Keren Bergman |
SC | 6 |
| 2018 | Improving MPI Multi-threaded RMA Communication PerformanceabstractOne-sided communication is crucial to enabling communication concurrency. As core counts have increased, particularly with many-core architectures, one-sided (RMA) communication has been proposed to address the ever increasing contention at the network interface. The difficulty in using one-sided (RMA) communication with MPI is that the performance of MPI implementations using RMA with multiple concurrent threads is not well understood. Past studies have been done using MPI RMA in combination with multi-threading (RMA-MT) but they have been performed on older MPI implementations lacking RMA-MT optimizations. In addition prior work has only been done at smaller scale (<=512 cores). Nathan T. Hjelm, Matthew G. F. Dosanjh, Ryan E. Grant, Taylor L. Groves, Patrick G. Bridges, Dorian C. Arnold |
ICPP | 4 |
| 2018 | Unraveling Network-Induced Memory Contention: Deeper Insights with Machine LearningabstractRemote Direct Memory Access (RDMA) is expected to be an integral communication mechanism for future exascale systems-enabling asynchronous data transfers, so that applications may fully utilize CPU resources while simultaneously sharing data amongst remote nodes. In this work we examine Network-induced Memory Contention (NiMC) on Infiniband networks. We expose the interactions between RDMA, main-memory and cache, when applications and out-of-band services compete for memory resources. We then explore NiMC's resulting impact on application-level performance. For a range of hardware technologies and HPC workloads, we quantify NiMC and show that NiMC's impact grows with scale resulting in up to 3X performance degradation at scales as small as 8K processes even in applications that previously have been shown to be performance resilient in the presence of noise. Additionally, this work examines the problem of predicting NiMC's impact on applications by leveraging machine learning and easily accessible performance counters. This approach provides additional insights about the root cause of NiMC and facilitates dynamic selection of potential solutions. Lastly, we evaluated three potential techniques to reduce NiMC's impact, namely hardware offloading, core reservation and software-based network throttling. Taylor L. Groves, Ryan E. Grant, Aaron Gonzales, Dorian C. Arnold |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | Understanding Performance Variability on the Aries Dragonfly NetworkabstractThis work evaluates performance variability in the Cray Aries dragonfly network and characterizes its impact on MPI Allreduce. The execution time of Allreduce is limited by the performance of the slowest participating process, which can vary by more than an order of magnitude. We utilize counters from the network routers to provide a better understanding of how competing workloads can influence performance. Specifically, we examine the relationships between message size, process counts, Aries counters and the Allreduce communication-time. Our results suggest that competing traffic from other jobs can significantly impact performance on the Aries Dragonfly Network. Furthermore, we show that Aries network counters are a valuable tool, explaining up to 70% of the performance variability for our experiments on a large-scale production system. Taylor L. Groves, Yizi Gu, Nicholas J. Wright |
CLUSTER | 1 |
| 2016 | RMA-MT: A Benchmark Suite for Assessing MPI Multi-threaded RMA PerformanceabstractReaching Exascale will require leveraging massive parallelism while potentially leveraging asynchronous communication to help achieve scalability at such large levels of concurrency. MPI is a good candidate for providing the mechanisms to support communication at such large scales. Two existing MPI mechanisms are particularly relevant to Exascale: multi-threading, to support massive concurrency, and Remote Memory Access (RMA), to support asynchronous communication. Unfortunately, multi-threaded MPI RMA code has not been extensively studied. Part of the reason for this is that no public benchmarks or proxy applications exist to assess its performance. The contributions of this paper are the design and demonstration of the first available proxy applications and micro-benchmark suite for multi-threaded RMA in MPI, a study of multi-threaded RMA performance of different MPI implementations, and an evaluation of how these benchmarks can be used to test development for both performance and correctness. Matthew G. F. Dosanjh, Taylor L. Groves, Ryan E. Grant, Ron Brightwell, Patrick G. Bridges |
CCGrid | 2 |
| 2016 | (SAI) Stalled, Active and Idle: Characterizing Power and Performance of Large-Scale Dragonfly NetworksabstractExascale networks are expected to comprise a significant part of the total monetary cost and 10-20% of the power budget allocated to exascale systems. Yet, our understanding of current and emerging workloads on these networks is limited. Left ignored, this knowledge gap likely will translate into missed opportunities for (1) improved application performance and (2) decreased power and monetary costs in next generation systems. This work targets a detailed understanding and analysis of the performance and utilization of the dragonfly network topology. Using the Structural Simulation Toolkit (SST) and a range of relevant workloads on a dragonfly topology of 110,592 nodes, we examine network design tradeoffs amongst execution time, power, bandwidth, and the number of global links. Our simulations report stalled, active and idle time on a per-port level of the fabric, in order to provide a detailed picture of future networks. The results of this work show potential savings of 3-10% of the exascale power budget and provide valuableinsights to researchers looking for new opportunities to improve performance and increase power efficiency of next generation HPC systems. Taylor L. Groves, Ryan E. Grant, Karl S. Hemmert, Simon D. Hammond, Michael J. Levenhagen, Dorian C. Arnold |
CLUSTER | 1 |
| 2016 | NiMC: Characterizing and Eliminating Network-Induced Memory ContentionabstractRemote Direct Memory Access (RDMA) is expected to be an integral communication mechanism for future exascale systems -- enabling asynchronous data transfers, so that applications may fully utilize all CPU resources while simultaneously sharing data amongst remote nodes. We examined this network-induced memory contention (NiMC), the interactions between RDMA and the memory subsystem when applications and out-of-band services compete for memory resources, and NiMC's resulting impact on application-level performance. For a range of hardware technologies and HPC workloads, we quantified NiMC and show that NiMC's impact grows with scale resulting in up to 3X performance degradation at scales as small as 8K processes even in applications that previously have been shown to be performance resilient in the presence of noise. We also evaluated three potential techniques to reduce NiMC's performance impact, namely hardware offloading, core reservation and software-based network throttling. While all three of these solutions show promise, we provide guidelines that help select the best solution for a given environment. Taylor L. Groves, Ryan E. Grant, Dorian C. Arnold |
IPDPS | 1 |
| 2015 | A LogP Extension for Modeling Tree Aggregation NetworksabstractAs high-performance systems continue to expand in power and size, scalable communication and data transfer is necessary to facilitate next generation monitoring and analysis. Many popular frameworks such as MapReduce, MPI and MRNet utilize scalable reduction operations to fulfill the performance requirements of a large distributed system. The structures to handle these aggregations may simply consist of a single level with children reporting directly to the parent node, or it may be layered to create a large tree with varying breadth and height. Despite their common-place, the techniques for modeling these Tree Aggregation Networks (TANs) are lacking. This paper addresses this need by introducing a novel extension of the LogP framework for Tree Aggregation Networks. Our TAN model adheres to the simplicity of the LogP model, but utilizes structural insights to provide a simple yet precise performance estimate. Additionally, our model makes no assumptions of the underlying NIC transfer mechanisms or uniformity of tree breadth, making it suitable for a wide range of environments. To evaluate our TAN model, we compare it against the traditional LogP model for predicting the performance of the Multicast Reduction Network (MRNet) framework. Taylor L. Groves, Samuel K. Gutierrez, Dorian C. Arnold |
CLUSTER | 1 |
| 2009 | Probability Delegation Forwarding in Delay Tolerant NetworksabstractDelay tolerant networks are a type of wireless mobile networks that do not guarantee the existence of a path between a source and a destination at any time. In such a network, one of the critical issues is to reliably deliver data with a low latency. Naive forwarding approaches, such as flooding and its derivatives, make the routing cost (here defined as the number of copies duplicated for a message) very high. Many efforts have been made to reduce the cost while maintaining performance. Recently, an approach called delegation forwarding (DF) caught significant attention in the research community because of its simplicity and good performance. In a network with N nodes, it reduces the cost to O(radic(N)) which is better than O(N) in other methods. In this paper, we extend the DF algorithm by putting forward a new scheme called probability delegation forwarding (PDF) that can further reduce the cost to O(Nlog2+2p(1+p)), p isin (0, 1). Simulation results show that PDF can achieve similar delivery ratio, which is the most important metric in DTNs, as the DF scheme at a lower cost if p is not too small. In addition, we propose the threshold probability delegation forwarding (TPDF) scheme to close the latency gap between the DF and PDF schemes. Xiao Chen 0001, Jian Shen 0005, Taylor L. Groves, Jie Wu 0001 |
ICCCN | 3 |