EDBT 2026 Demo / reviewers in the wild / expert
Amanda Bienz
dblp:173/5133
· DBLP profile ↗
8ranked-venue papers
3as first author
6since 2021 · last 2024
0000-0002-8891-934XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Optimizing Neighbor Collectives with Topology ObjectsabstractMany HPC applications implement non-cartesian neighbor data exchanges using MPI point-to-point operations rather than utilizing native MPI neighbor collective methods. Each application must therefore implement their own commu-nication optimizations, rather than leveraging any optimizations that could be provided by MPI. While an interface for such optimizations is provided within MPI through neighborhood collectives, applications avoid these methods due to the lack of performance optimizations within them along with large costs associated with graph communicator formation. This paper presents a novel approach for creating local, non-cartesian topol-ogy objects that provides finer control over the aforementioned setup costs. Any additional setup costs, such as initializing per-iteration optimizations, can then be deferred until additional information is available, such as within persistent initialization calls. This paper describes our implementation within an MPI extension library and demonstrates the effectiveness of our approach in simple benchmarks and real-world applications. Gerald Collom, Derek Schafer, Amanda Bienz, Patrick G. Bridges, Galen M. Shipman |
CLUSTER | 3 |
| 2024 | A More Scalable Sparse Dynamic Data ExchangeabstractParallel architectures are continually increasing in performance and scale while underlying algorithmic infrastruc-ture often fails to take full advantage of available compute power. Within the context of MPI, irregular communication patterns create bottlenecks in parallel applications. One common bottleneck is the sparse dynamic data exchange, often required when forming communication patterns within applications. There is a large variety of approaches for these dynamic exchanges, with optimizations implemented directly in parallel applications. This paper proposes a novel API within an MPI eXtension library, allowing applications to utilize the variety of provided optimizations for sparse dynamic data exchange methods. Fur-ther, the paper presents novel locality-aware sparse dynamic data exchange algorithms. Finally, performance results show locality-aware approaches achieve up to 128x over existing approaches when exchanging only pattern of communication, and up to 54x when exchanging data to be communicated as well. Andrew Geyko, Gerald Collom, Derek Schafer, Patrick G. Bridges, Amanda Bienz |
HiPC | 5 |
| 2023 | Evaluating the Viability of LogGP for Modeling MPI Performance with Non-contiguous Datatypes on Modern ArchitecturesabstractModern architectures and communication systems software include complex hardware, communication abstractions, and optimizations that make their performance difficult to measure, model, and understand. This paper examines the ability of modified versions of the existing Netgauge communication performance measurement tool and LogGOPS performance model to accurately characterize communication behavior of modern hardware, MPI abstractions, and implementations. This includes analyzing their ability to model both GPU-aware communication in different MPI implementations and quantifying the performance characteristics of different approaches to non-contiguous data communication on modern GPU systems. This paper also applies these techniques to quantify the performance of different implementations and optimization approaches to non-contiguous data communication on a variety of systems, demonstrating that modern communication system design approaches can result in widely-varying and difficult-to-predict performance variation, even within the same hardware/communication software combination. Nicholas H. Bacon, Patrick G. Bridges, Scott Levy, Kurt B. Ferreira, Amanda Bienz |
EuroMPI | 5 |
| 2023 | Characterizing the performance of node-aware strategies for irregular point-to-point communication on heterogeneous architectures
Shelby Lockhart, Amanda Bienz, William Gropp, Luke N. Olson |
Parallel Comput. | 2 |
| 2022 | A Locality-Aware Bruck AllgatherabstractCollective algorithms are an essential part of MPI, allowing application programmers to utilize underlying optimizations of common distributed operations. The MPI_Allgather gathers data, which is originally distributed across all processes, so that all data is available to each process. For small data sizes, the Bruck algorithm is commonly implemented to minimize the maximum number of messages communicated by any process. However, the cost of each step of communication is dependent upon the relative locations of source and destination processes, with non-local messages, such as inter-node, significantly more costly than local messages, such as intra-node. This paper optimizes the Bruck algorithm with locality-awareness, minimizing the number and size of non-local messages to improve performance and scalability of the allgather operation. Amanda Bienz, Shreeman Gautam, Amun Kharel |
EuroMPI | 1 |
| 2022 | Tausch: A halo exchange library for large heterogeneous computing systems using MPI, OpenCL, and CUDA
Lukas Spies, Amanda Bienz, J. David Moulton, Luke N. Olson, Andrew Reisner |
Parallel Comput. | 2 |
| 2019 | Node aware sparse matrix-vector multiplication
Amanda Bienz, William Gropp, Luke N. Olson |
J. Parallel Distributed Comput. | 1 |
| 2018 | Improving Performance Models for Irregular Point-to-Point CommunicationabstractParallel applications are often unable to take full advantage of emerging parallel architectures due to scaling limitations, which arise due to inter-process communication. Performance models are used to analyze the sources of communication costs. However, traditional models for point-to-point communication fail to capture the full cost of many irregular operations, such as sparse matrix methods. In this paper, a node-aware based model is presented. Furthermore, the model is extended to include communication queue search time as well as an additional parameter estimating network contention. The resulting model is applied to a variety of irregular communication patterns throughout matrix operations, displaying improved accuracy over traditional models. Amanda Bienz, William Gropp, Luke N. Olson |
EuroMPI | 1 |