Aaron Welch

dblp:42/8808 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
0000-0002-8988-3027ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 56% High-performance computing · 28% Interconnection networks and networks-on-chip · 17%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › memory interference
cache contention
1.012026
Performance Analysis of Conveyors: Memory Dominates? · HPDC 2026
Memory systems › memory interference
memory contention
1.012026
Performance Analysis of Conveyors: Memory Dominates? · HPDC 2026
High-performance computing
message aggregation
1.012026
Performance Analysis of Conveyors: Memory Dominates? · HPDC 2026
Interconnection networks and networks-on-chip
network bandwidth
0.312026
Performance Analysis of Conveyors: Memory Dominates? · HPDC 2026
Interconnection networks and networks-on-chip › high-speed networks
supercomputer interconnect
0.312026
Performance Analysis of Conveyors: Memory Dominates? · HPDC 2026

Methods — techniques the papers use, named apart from their topics

profiling · 1.0performance measurement · 1.0
YearPublicationVenuePosition
2026 Performance Analysis of Conveyors: Memory Dominates?
abstract
Small-message aggregation is critical for scaling irregular, communication intensive applications in high-performance computing. In this paper, contrary to conventional wisdom, we present the first systematic study showing that memory contention, not network bandwidth, is the dominant bottleneck in message aggregation runtimes. Using the state-of-the-art conveyors library as our reference implementation, we conducted extensive experiments on HPC systems featuring Slingshot 11 and InfiniBand interconnects, scaling to 16k cores (256 nodes) and processing 10s–100s GB of data. Our measurements reveal that interference between user data and aggregation buffers drives LLC miss rates to 77%, inflating memory costs by 2–3× over the algorithmic baseline. Consequently, we advocate for dedicated near-memory subsystems to improve the scalability and performance of message aggregation runtimes. This paper also demonstrates up to an order of magnitude higher latency for conveyor termination compared to a traditional HPC barrier, and it examines the impact of communication context isolation and the critical challenge of programmability.
Shubhendra Pal Singhal, Aaron Welch, Oscar R. Hernandez, Stephen W. Poole, Akihiro Hayashi, Vivek Sarkar
HPDC2
2023 Extending OpenSHMEM with Aggregation Support for Improved Message Rate Performance
Aaron Welch, Oscar R. Hernandez, Stephen W. Poole
Euro-Par1