Stephanie Brink

dblp:252/4605 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0002-1458-8453ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 2 first-author · 9 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Flux Fiction: Hopping Toward Storage Graph Scheduling With El Capitan's Rabbits
abstract
Modern HPC systems are placing increasing demands on job schedulers due to their scale and novel hardware. El Capitan’s Rabbit nodes exemplify this challenge: unlike traditional systems where storage is remote and shared, Rabbit nodes wire local NVMe SSDs directly to compute nodes via PCIe, forcing schedulers to actively track storage topology, capacity, and cross-job persistence, concerns they were never designed to handle. We introduce Flux Fiction, a fully plugin-based HPC system emulator built on top of Flux that replays historical job traces to evaluate scheduling policies in Flux. We validate Flux Fiction against the LLNL Tuolumne cluster using two workloads across four queueing policies, achieving a P99-bounded slowdown error below 1 in 7 of 8 experiments and a maximum utilization error of 1.2%. We then use Flux Fiction to explore Rabbit storage scheduling, demonstrating its ability to explore novel scheduling scenarios.
Walter J. Ashworth, Ian Lumsden, Jim Garlick, Mark Grondona, Olga Pearce, Stephanie Brink, Daniel Milroy, Tapasya Patki, Thomas Scogland, Michela Taufer
HPDC6
2026 Parallel sorting algorithm classification: is manual instrumentation necessary?
Michael McKinsey, Dewi Yokelson, Stephanie Brink, Thomas Scogland, Olga Pearce
Future Gener. Comput. Syst.3
2025 Performance Optimization of an Exascale Implicit Kinetic Plasma Simulation on El Capitan
abstract
We present performance scaling and optimization of iPIC3D - an exascale-class, GPU-enabled implicit particle-in-cell code for planetary-scale magnetosphere modeling and plasma simulation - on El Capitan. Our strong and weak scaling studies demonstrate near-linear scaling up to 8,000 nodes (32,000 APUs) with a parallel efficiency of nearly 100%. We optimize iPIC3D to leverage AMD’s MI300A APUs, the Merced Lustre filesystem, and Rabbit nodes for high-bandwidth I/O. Optimizations reduce memory usage by 97% and runtime by 74%, enabling simulations that are 1.8 times larger than before. Rabbit further improve checkpointing bandwidth by 2 times, ensuring scalable fault- tolerant simulations on exascale architectures.
Ian Lumsden, Stefano Markidis, Andong Hu, Ivy Bo Peng, Luca Pennati, Dewi Yokelson, Stephanie Brink, Olga Pearce, Thomas Scogland, Hariharan Devarajan, Bronis R. de Supinski, Gian Luca Delzanno, Michela Taufer
eScience7
2025 Flux Emulator: First Insights into Optimizing Scheduling for Exascale HPC
abstract
El Capitan, currently the world's largest supercomputer at 1.742 Ex-aflop/s, introduces challenges in scheduling due to its scale and innovative rabbit nodes, which traditional schedulers cannot efficiently handle. Flux, a resource and job management system, handles dynamic resource allocation tailored for exascale systems through its graph-based scheduler, Fluxion. This work introduces the Flux Emulator, a tool designed to test scheduling policies in Fluxion without impacting production systems. The emulator plugs into the real components of Flux and Fluxion to mimic job execution, emulate resource usage, and collect information on how the job behaves. Preliminary tests show negligible overhead introduced by the emulator and demonstrate its effectiveness in evaluating scheduli ng policies, like conservative backfilling, in a fraction of the time required with a real system.
Walter J. Ashworth, Ian Lumsden, Jim Garlick, Mark Grondona, Olga Pearce, Stephanie Brink, Dewi Yokelson, Daniel Milroy, Tapasya Patki, Thomas Scogland, Michela Taufer
HPDC6
2025 Thicket Workflow for Classifying Parallel Sorting Algorithms
abstract
Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We develop an approach to learn parallel sorting algorithm classes from performance data directly in order to classify parallel sorting algorithms without using the source code. In this paper, we focus on the workflow and interfaces we developed for collecting, processing, and modeling performance data using Caliper, Thicket, PyTorch, and Scikit-learn. Our workflow results in classification accuracy of our machine learning models up to 95.3% across five different algorithm classes.
Michael McKinsey, Stephanie Brink, Stephanie Lam, Dewi Yokelson, Olga Pearce
HPDC2
2025 Cross-Architecture Performance Analysis Using the RAJA Performance Suite
abstract
Modern supercomputer architectures are diverse and becoming increasingly complex. Scientists are constantly porting code and re-optimizing it for the new architecture, but achieving good performance is challenging. Performance portability programming models such as RAJA, Kokkos, and OpenMP enable codes to maintain a single-source code rather than rewriting for each target architecture. However, portability models alone will not result in optimal performance as hardware has varying specifications (e.g., cache sizes and speeds) and parallel algorithms may use varying amounts of memory and compute resources. We present a systematic analysis of application behaviors across a diverse set of CPU and GPU hardware. We leverage the RAJA Performance Suite, which contains a curated set of kernels commonly found in HPC applications, to perform an in-depth GPU and memory analysis as well as a quantitative performance portability evaluation across different compute platforms. In analyzing the performance portability scores, we identify gaps and opportunities to achieve consistent performance across platforms. We provide a comprehensive analysis across seven architectures, including the most recent GPU systems with new physical memory layouts, where kernels demonstrate a runtime speedup of up to 44 ×. Although the speedup highlights the baseline improvements of newer hardware, the performance portability scores calculated, ranging from 0% to 92%, showcase where opportunities remain for scientists to increase utilization of the newer systems.
Dewi Yokelson, Stephanie Brink, Jason Burmark, Michael McKinsey, Befikir Bogale, Ian Lumsden, Michela Taufer, Thomas Scogland, Olga Pearce
ICPP2
2025 A Global Perspective on Supercomputer Power Provisioning: Case Studies from United States and Europe
abstract
Electrical provisioning in high performance computing is transitioning from simple nameplate Thermal Design Power (TDP) models to more nuanced approaches based on expected electrical load.This paper captures current power
Tapasya Patki, Barry Rountree, Torsten Wilde, Andrea Bartolini, Stephanie Brink, Esa Heiskanen, Sachin Idgunji, Matthias Maiterth, James H. Rogers, Ermal Rrapaj, Ralf Schneider, Woong Shin, Kathleen Shoga, Christian Simmendinger, Nicholas J. Wright, Zhengji Zhao
ICS5
2025 Enabling Lightweight Performance Analysis of Complex Scientific Workflows with PerfFlowAspect
Aliza Lisan, Tapasya Patki, Stephanie Brink, Konstantinos Parasyris, Brian Gunnarson, Giorgis Georgakoudis, Hank Childs
SSDBM3
2024 Design Concerns for Integrated Scripting and Interactive Visualization in Notebook Environments
abstract
Interactive visualization can support fluid exploration but is often limited to predetermined tasks. Scripting can support a vast range of queries but may be more cumbersome for free-form exploration. Embedding interactive visualization in scripting environments, such as computational notebooks, provides an opportunity to leverage the strengths of both direct manipulation and scripting. We investigate interactive visualization design methodology, choices, and strategies under this paradigm through a design study of calling context trees used in performance analysis, a field which exemplifies typical exploratory data analysis workflows with Big Data and hard to define problems. We first produce a formal task analysis assigning tasks to graphical or scripting contexts based on their specificity, frequency, and suitability. We then design a notebook-embedded interactive visualization and validate it with intended users. In a follow-up study, we present participants with multiple graphical and scripting interaction modes to elicit feedback about notebook-embedded visualization design, finding consensus in support of the interaction model. We report and reflect on observations regarding the process and design implications for combining visualization and scripting in notebooks.
Connor Scully-Allison, Ian Lumsden, Katy Williams, Jesse Bartels, Michela Taufer, Stephanie Brink, Abhinav Bhatele, Olga Pearce, Katherine E. Isaacs
IEEE Trans. Vis. Comput. Graph.6
2023 Thicket: Seeing the Performance Experiment Forest for the Individual Run Trees
abstract
Thicket is an open-source Python toolkit for Exploratory Data Analysis (EDA) of multi-run performance experiments. It enables an understanding of optimal performance configuration for large-scale application codes. Most performance tools focus on a single execution (e.g., single platform, single measurement tool, single scale). Thicket bridges the gap to convenient analysis in multi-dimensional, multi-scale, multi-architecture, and multi-tool performance datasets by providing an interface for interacting with the performance data. Thicket has a modular structure composed of three components. The first component is a data structure for multi-dimensional performance data, which is composed automatically on the portable basis of call trees, and accommodates any subset of dimensions present in the dataset. The second is the metadata, enabling distinction and sub-selection of dimensions in performance data. The third is a dimensionality reduction mechanism, enabling analysis such as computing aggregated statistics on a given data dimension. Extensible mechanisms are available for applying analyses (e.g., top-down on Intel CPUs), data science techniques (e.g., K-means clustering from scikit-learn), modeling performance (e.g., Extra-P), and interactive visualization. We demonstrate the power and flexibility of Thicket through two case studies, first with the open-source RAJA Performance Suite on CPU and GPU clusters and another with a large physics simulation run on both a traditional HPC cluster and an AWS Parallel Cluster instance.
Stephanie Brink, Michael McKinsey, David Böhme, Connor Scully-Allison, Ian Lumsden, W. Daryl Hawkins, Treece Burgess, Vanessa Lama, Jakob Lüttgau, Katherine E. Isaacs, Michela Taufer, Olga Pearce
HPDC1
2023 Scalable Comparative Visualization of Ensembles of Call Graphs
abstract
Optimizing the performance of large-scale parallel codes is critical for efficient utilization of computing resources. Code developers often explore various execution parameters, such as hardware configurations, system software choices, and application parameters, and are interested in detecting and understanding bottlenecks in different executions. They often collect hierarchical performance profiles represented as call graphs, which combine performance metrics with their execution contexts. The crucial task of exploring multiple call graphs together is tedious and challenging because of the many structural differences in the execution contexts and significant variability in the collected performance metrics (e.g., execution runtime). In this paper, we present Ensemble CallFlow to support the exploration of ensembles of call graphs using new types of visualizations, analysis, graph operations, and features. We introduce ensemble-Sankey, a new visual design that combines the strengths of resource-flow (Sankey) and box-plot visualization techniques. Whereas the resource-flow visualization can easily and intuitively describe the graphical nature of the call graph, the box plots overlaid on the nodes of Sankey convey the performance variability within the ensemble. Our interactive visual interface provides linked views to help explore ensembles of call graphs, e.g., by facilitating the analysis of structural differences, and identifying similar or distinct call graphs. We demonstrate the effectiveness and usefulness of our design through case studies on large-scale parallel codes.
Suraj P. Kesavan, Harsh Bhatia, Abhinav Bhatele, Stephanie Brink, Olga Pearce, Todd Gamblin, Peer-Timo Bremer, Kwan-Liu Ma
IEEE Trans. Vis. Comput. Graph.4
2022 Enabling Call Path Querying in Hatchet to Identify Performance Bottlenecks in Scientific Applications
abstract
As computational science applications benefit from larger-scale, more heterogeneous high performance computing (HPC) systems, the process of studying their performance becomes increasingly complex. The performance data analysis library Hatchet provides some insights into this complexity, but is currently limited in its analysis capabilities. Missing capabilities include the handling of relational caller-callee data captured by HPC profilers. To address this shortcoming, we augment Hatchet with a Call Path Query Language that leverages relational data in the performance analysis of scientific applications. Specifically, our Query Language enables data reduction using call path pattern matching. We demonstrate the effectiveness of our Query Language in identifying performance bottlenecks and enhancing Hatchet's analysis capabilities through three case studies. In the first case study, we compare the performance of sequential and multi-threaded versions of the graph alignment application Fido. In doing so, we identify the existence of large memory inefficiencies in both versions. In the second case study, we examine the performance of MPI calls in the linear algebra mini-application AMG2013 when using MVAPICH and Spectrum-MPI. In doing so, we identify hidden performance losses in specific MPI functions. In the third case study, we illustrate the use of our Query Language in Hatchet's interactive visualization. In doing so, we show that our Query Language enables a simple and intuitive way to massively reduce profiling data.
Ian Lumsden, Jakob Lüttgau, Vanessa Lama, Connor Scully-Allison, Stephanie Brink, Katherine E. Isaacs, Olga Pearce, Michela Taufer
e-Science5
2021 Introducing Application Awareness Into a Unified Power Management Stack
abstract
Effective power management in a data center is critical to ensure that power delivery constraints are met while maximizing the performance of users' workloads. Power limiting is needed in order to respond to greater-than-expected power demand. HPC sites have generally tackled this by adopting one of two approaches: (1) a system-level power management approach that is aware of the facility or site-level power requirements, but is agnostic to the application demands; OR (2) a job-level power management solution that is aware of the application design patterns and requirements, but is agnostic to the site-level power constraints. Simultaneously incorporating solutions from both domains often leads to conflicts in power management mechanisms. This, in turn, affects system stability and leads to irreproducibility of performance. To avoid this irreproducibility, HPC sites have to choose between one of the two approaches, thereby leading to missed opportunities for efficiency gains.This paper demonstrates the need for the HPC community to collaborate towards seamless integration of system-aware and application-aware power management approaches. This is achieved by proposing a new dynamic policy that inherits the benefits of both approaches from tight integration of a resource manager and a performance-aware job runtime environment. An empirical comparison of this integrated management approach against state-of-the-art solutions exposes the benefits of investing in end-to-end solutions to optimize for system-wide performance or efficiency objectives. With our proposed system-application integrated policy, we observed up to 7% reduction in system time dedicated to jobs and up to 11% savings in compute energy, compared to a baseline that is agnostic to system power and application design constraints.
Daniel C. Wilson, Siddhartha Jana, Aniruddha Marathe, Stephanie Brink, Christopher Cantalupo, Diana R. Guttman, Brad Geltz, Lowren H. Lawson, Asma Al-Rawi, Ali Mohammad, Fuat Keceli, Federico Ardanaz, Jonathan Eastep, Ayse K. Coskun
IPDPS4
2021 Evaluating adaptive and predictive power management strategies for optimizing visualization performance on supercomputers
Stephanie Brink, Matthew Larsen, Hank Childs, Barry Rountree
Parallel Comput.1
2019 Hatchet: pruning the overgrowth in parallel profiles
abstract
Performance analysis is critical for eliminating scalability bottlenecks in parallel codes. There are many profiling tools that can instrument codes and gather performance data. However, analytics and visualization tools that are general, easy to use, and programmable are limited. In this paper, we focus on the analytics of structured profiling data, such as that obtained from calling context trees or nested region timers in code. We present a set of techniques and operations that build on the pandas data analysis library to enable analysis of parallel profiles. We have implemented these techniques in a Python-based library called Hatchet that allows structured data to be filtered, aggregated, and pruned. Using performance datasets obtained from profiling parallel codes, we demonstrate performing common performance analysis tasks reproducibly with a few lines of Hatchet code. Hatchet brings the power of modern data science tools to bear on performance analysis.
Abhinav Bhatele, Stephanie Brink, Todd Gamblin
SC2