Dewi Yokelson

dblp:271/4437 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0003-1453-5906ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Parallel sorting algorithm classification: is manual instrumentation necessary?
Michael McKinsey, Dewi Yokelson, Stephanie Brink, Thomas Scogland, Olga Pearce
Future Gener. Comput. Syst.2
2025 Performance Optimization of an Exascale Implicit Kinetic Plasma Simulation on El Capitan
abstract
We present performance scaling and optimization of iPIC3D - an exascale-class, GPU-enabled implicit particle-in-cell code for planetary-scale magnetosphere modeling and plasma simulation - on El Capitan. Our strong and weak scaling studies demonstrate near-linear scaling up to 8,000 nodes (32,000 APUs) with a parallel efficiency of nearly 100%. We optimize iPIC3D to leverage AMD’s MI300A APUs, the Merced Lustre filesystem, and Rabbit nodes for high-bandwidth I/O. Optimizations reduce memory usage by 97% and runtime by 74%, enabling simulations that are 1.8 times larger than before. Rabbit further improve checkpointing bandwidth by 2 times, ensuring scalable fault- tolerant simulations on exascale architectures.
Ian Lumsden, Stefano Markidis, Andong Hu, Ivy Bo Peng, Luca Pennati, Dewi Yokelson, Stephanie Brink, Olga Pearce, Thomas Scogland, Hariharan Devarajan, Bronis R. de Supinski, Gian Luca Delzanno, Michela Taufer
eScience6
2025 Flux Emulator: First Insights into Optimizing Scheduling for Exascale HPC
abstract
El Capitan, currently the world's largest supercomputer at 1.742 Ex-aflop/s, introduces challenges in scheduling due to its scale and innovative rabbit nodes, which traditional schedulers cannot efficiently handle. Flux, a resource and job management system, handles dynamic resource allocation tailored for exascale systems through its graph-based scheduler, Fluxion. This work introduces the Flux Emulator, a tool designed to test scheduling policies in Fluxion without impacting production systems. The emulator plugs into the real components of Flux and Fluxion to mimic job execution, emulate resource usage, and collect information on how the job behaves. Preliminary tests show negligible overhead introduced by the emulator and demonstrate its effectiveness in evaluating scheduli ng policies, like conservative backfilling, in a fraction of the time required with a real system.
Walter J. Ashworth, Ian Lumsden, Jim Garlick, Mark Grondona, Olga Pearce, Stephanie Brink, Dewi Yokelson, Daniel Milroy, Tapasya Patki, Thomas Scogland, Michela Taufer
HPDC7
2025 Thicket Workflow for Classifying Parallel Sorting Algorithms
abstract
Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We develop an approach to learn parallel sorting algorithm classes from performance data directly in order to classify parallel sorting algorithms without using the source code. In this paper, we focus on the workflow and interfaces we developed for collecting, processing, and modeling performance data using Caliper, Thicket, PyTorch, and Scikit-learn. Our workflow results in classification accuracy of our machine learning models up to 95.3% across five different algorithm classes.
Michael McKinsey, Stephanie Brink, Stephanie Lam, Dewi Yokelson, Olga Pearce
HPDC4
2025 Cross-Architecture Performance Analysis Using the RAJA Performance Suite
abstract
Modern supercomputer architectures are diverse and becoming increasingly complex. Scientists are constantly porting code and re-optimizing it for the new architecture, but achieving good performance is challenging. Performance portability programming models such as RAJA, Kokkos, and OpenMP enable codes to maintain a single-source code rather than rewriting for each target architecture. However, portability models alone will not result in optimal performance as hardware has varying specifications (e.g., cache sizes and speeds) and parallel algorithms may use varying amounts of memory and compute resources. We present a systematic analysis of application behaviors across a diverse set of CPU and GPU hardware. We leverage the RAJA Performance Suite, which contains a curated set of kernels commonly found in HPC applications, to perform an in-depth GPU and memory analysis as well as a quantitative performance portability evaluation across different compute platforms. In analyzing the performance portability scores, we identify gaps and opportunities to achieve consistent performance across platforms. We provide a comprehensive analysis across seven architectures, including the most recent GPU systems with new physical memory layouts, where kernels demonstrate a runtime speedup of up to 44 ×. Although the speedup highlights the baseline improvements of newer hardware, the performance portability scores calculated, ranging from 0% to 92%, showcase where opportunities remain for scientists to increase utilization of the newer systems.
Dewi Yokelson, Stephanie Brink, Jason Burmark, Michael McKinsey, Befikir Bogale, Ian Lumsden, Michela Taufer, Thomas Scogland, Olga Pearce
ICPP1
2024 Enabling Performance Observability for Heterogeneous HPC Workflows with SOMA
abstract
Heterogeneous workflows represent a promising approach for overcoming traditional application performance limitations and to accelerate scientific insight on high-performance computing (HPC) platforms. As HPC platforms grow in size and complexity, managing and optimizing workflow resources while maximizing scientific output assumes vital importance. Optimal workflow resource allocation requires high-quality and timely information about the state of the hardware resources, the status of the pending tasks, the performance of the tasks that have already been executed, and the current status of the workflow itself. A robust performance observability framework that captures and delivers this information can fundamentally improve the quality of decision-making within the workflow system, setting the stage for the adaptive execution of workflow tasks. We propose the use of SOMA, a service-based performance observability framework for such HPC workflows. With the RADICAL-Pilot runtime system as a development vehicle, SOMA demonstrates that service-based architectures coupled with an appropriate data model can serve the performance monitoring needs of large-scale ensemble workflows in a low-overhead fashion. Effective observability of workflow performance requires exporting, storing, and analyzing several types of performance data from across the application and workflow software stacks. Our study finds significant benefits in integrating observability frameworks as first-class citizens within an HPC workflow software stack. In this paper, we demonstrate how SOMA can simultaneously observe the performance states of the individual tasks, system hardware, and the workflow as a whole. Such information can then be employed to calculate better resource allocation and task configuration.
Dewi Yokelson, Mikhail Titov, Srinivasan Ramesh, Ozgur O. Kilic, Matteo Turilli, Shantenu Jha, Allen D. Malony
ICPP1
2024 SOMA: Observability, monitoring, and in situ analytics for exascale applications
abstract
Summary With the rise of exascale systems and large, data‐centric workflows, the need to observe and analyze high performance computing (HPC) applications during their execution is becoming increasingly important. HPC applications are typically not designed with online monitoring in mind, therefore, the observability challenge lies in being able to access and analyze interesting events with low overhead while seamlessly integrating such capabilities into existing and new applications. We explore how our service‐based observation, monitoring, and analytics (SOMA) approach to collecting and aggregating both application‐specific diagnostic data and performance data addresses these needs. We present our SOMA framework and demonstrate its viability with LULESH, a hydrodynamics proxy application. Then we focus on Astaroth, a multi‐GPU library for stencil computations, highlighting the integration of the TAU and APEX performance tools and SOMA for application and performance data monitoring.
Dewi Yokelson, Oskar Lappi, Srinivasan Ramesh, Miikka S. Väisälä, Kevin A. Huck, Touko Puro, Boyana Norris, Maarit J. Korpi-Lagg, Keijo Heljanko, Allen D. Malony
Concurr. Comput. Pract. Exp.1