Michael McKinsey

dblp:354/5129 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0001-7201-5733ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Parallel sorting algorithm classification: is manual instrumentation necessary?
Michael McKinsey, Dewi Yokelson, Stephanie Brink, Thomas Scogland, Olga Pearce
Future Gener. Comput. Syst.1
2025 Thicket Workflow for Classifying Parallel Sorting Algorithms
abstract
Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We develop an approach to learn parallel sorting algorithm classes from performance data directly in order to classify parallel sorting algorithms without using the source code. In this paper, we focus on the workflow and interfaces we developed for collecting, processing, and modeling performance data using Caliper, Thicket, PyTorch, and Scikit-learn. Our workflow results in classification accuracy of our machine learning models up to 95.3% across five different algorithm classes.
Michael McKinsey, Stephanie Brink, Stephanie Lam, Dewi Yokelson, Olga Pearce
HPDC1
2025 Cross-Architecture Performance Analysis Using the RAJA Performance Suite
abstract
Modern supercomputer architectures are diverse and becoming increasingly complex. Scientists are constantly porting code and re-optimizing it for the new architecture, but achieving good performance is challenging. Performance portability programming models such as RAJA, Kokkos, and OpenMP enable codes to maintain a single-source code rather than rewriting for each target architecture. However, portability models alone will not result in optimal performance as hardware has varying specifications (e.g., cache sizes and speeds) and parallel algorithms may use varying amounts of memory and compute resources. We present a systematic analysis of application behaviors across a diverse set of CPU and GPU hardware. We leverage the RAJA Performance Suite, which contains a curated set of kernels commonly found in HPC applications, to perform an in-depth GPU and memory analysis as well as a quantitative performance portability evaluation across different compute platforms. In analyzing the performance portability scores, we identify gaps and opportunities to achieve consistent performance across platforms. We provide a comprehensive analysis across seven architectures, including the most recent GPU systems with new physical memory layouts, where kernels demonstrate a runtime speedup of up to 44 ×. Although the speedup highlights the baseline improvements of newer hardware, the performance portability scores calculated, ranging from 0% to 92%, showcase where opportunities remain for scientists to increase utilization of the newer systems.
Dewi Yokelson, Stephanie Brink, Jason Burmark, Michael McKinsey, Befikir Bogale, Ian Lumsden, Michela Taufer, Thomas Scogland, Olga Pearce
ICPP4
2023 Thicket: Seeing the Performance Experiment Forest for the Individual Run Trees
abstract
Thicket is an open-source Python toolkit for Exploratory Data Analysis (EDA) of multi-run performance experiments. It enables an understanding of optimal performance configuration for large-scale application codes. Most performance tools focus on a single execution (e.g., single platform, single measurement tool, single scale). Thicket bridges the gap to convenient analysis in multi-dimensional, multi-scale, multi-architecture, and multi-tool performance datasets by providing an interface for interacting with the performance data. Thicket has a modular structure composed of three components. The first component is a data structure for multi-dimensional performance data, which is composed automatically on the portable basis of call trees, and accommodates any subset of dimensions present in the dataset. The second is the metadata, enabling distinction and sub-selection of dimensions in performance data. The third is a dimensionality reduction mechanism, enabling analysis such as computing aggregated statistics on a given data dimension. Extensible mechanisms are available for applying analyses (e.g., top-down on Intel CPUs), data science techniques (e.g., K-means clustering from scikit-learn), modeling performance (e.g., Extra-P), and interactive visualization. We demonstrate the power and flexibility of Thicket through two case studies, first with the open-source RAJA Performance Suite on CPU and GPU clusters and another with a large physics simulation run on both a traditional HPC cluster and an AWS Parallel Cluster instance.
Stephanie Brink, Michael McKinsey, David Böhme, Connor Scully-Allison, Ian Lumsden, W. Daryl Hawkins, Treece Burgess, Vanessa Lama, Jakob Lüttgau, Katherine E. Isaacs, Michela Taufer, Olga Pearce
HPDC2