Sarah Neuwirth

dblp:160/0630 · DBLP profile ↗
← Back
19ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0001-7409-153XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 6 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Metis: Agentic Knowledge Synthesis for Explainable I/O Performance in HPC Systems
abstract
I/O performance explainability in HPC requires contextual characterization across the full software and system stack. Contextual characterization identifies the semantics and runtime role of I/O functions. State-of-the-art contextual characterization is still largely manual, but remains highly valuable for explaining bottlenecks and guiding optimization. However, manual contextual characterization is difficult to scale and hard to reproduce as I/O libraries and cross-layer interactions grow in complexity. We present Metis, a framework for systematic characterization of HPC I/O functions that uses agentic LLMs to integrate heterogeneous evidence sources and quantify agents agreement. Across evaluation, Metis improves held-out-category generalization over an MCP Tool baseline (0.90 vs. 0.35), reduces runtime (27.7 s vs. 84.5 s) while increasing throughput (41.8 vs. 14.28 functions/min), and sustains high verifier throughput under federated scaling (330K–1.18M functions/s). These results demonstrate that Metis is an effective and practical approach for explainable characterization of complex HDF5 behavior, enabling more trustworthy and reproducible HPC I/O analysis.
Karim Youssef, Sarah Neuwirth, Neeraj Rajesh, Hariharan Devarajan
HPDC2
2026 Aligning Storage Benchmark Metrics with Application-Level Performance
abstract
The current state of practice in HPC is that performance metrics reported by storage benchmarks are disconnected from those obtained through application-level performance analysis tools, making it difficult for users to determine whether performance tuning efforts are effective or whether observed performance indicates underutilization of the system. In this work, we propose an approach to align IO500 storage benchmark results with application-level performance by deconstructing benchmark components and recalculating their metrics. The results show that certain universal metrics, such as bandwidth, can be meaningfully aligned with application performance, enabling more consistent and interpretable evaluation. Our findings also identify metrics that remain missing or cannot be reconciled, highlighting the need for standardization to align metrics produced by benchmarks and performance analysis tools.
Radita Liem, Julian M. Kunkel, Jay F. Lofstead, Sarah Neuwirth
SSDBM5
2025 VerifyIO: Verifying Adherence to Parallel I/O Consistency Semantics
abstract
High-performance computing (HPC) applications generate and consume substantial amounts of data, typically managed by parallel file systems. These applications access file systems either through the POSIX interface or by using highlevel I/O libraries. While the POSIX consistency model remains dominant in HPC, emerging file systems and popular I/O libraries increasingly adopt alternative consistency models that relax semantics in various ways, creating significant challenges for correctness and portability. This paper addresses these challenges by proposing a trace-driven I/O consistency verification workflow, implemented in our open-source tool, VerifyIO, which collects execution traces, detects data conflicts, and verifies proper synchronization against specified consistency models. Our extensive evaluation of 91 test case executions across three widely used I/O libraries with four I/O consistency models reveals critical consistency issues at both application and implementation levels.
Chen Wang 0004, Zhaobin Zhu, Kathryn Mohror, Sarah Neuwirth, Marc Snir
IPDPS4
2025 Comprehensive Performance Analysis of Portals4 Communication Primitives on BXI Hardware
abstract
This paper presents a comprehensive performance analysis of the BullSequana eXascale Interconnect (BXI) using the Portals4 programming model. The main contributions include: (1) the design and implementation of PtlBench, a Portals4 microbenchmark suite that evaluates low-level features such as bandwidth, latency, cache effects, and triggered operations; (2) a detailed comparison of Portals4-compatible MPI implementations (OpenMPI, ParaStationMPI) and our custom Portals4 device for the PGAS library GPI-2, covering point-topoint, one-sided, and collective operations; and (3) applicationlevel analysis using the Himeno and SSCA1 benchmarks to assess the impact on different communication patterns. These results provide valuable information of BXI’s capabilities and limitations for real-world HPC workloads and communication models.
Niklas Bartelheimer, Sarah Neuwirth
MASCOTS2
2025 XIO: Toward eXplainable I/O for HPC Systems
Sarah Neuwirth, Hariharan Devarajan, Chen Wang 0004, Jay F. Lofstead
SSDBM1
2025 Advancing HPC Performance Modeling with an Interactive, Automated and Tool-Agnostic ML-Driven Workflow
Zhaobin Zhu, Chen Wang 0004, Sarah Neuwirth
SSDBM3
2024 Tarazu: An Adaptive End-to-end I/O Load-balancing Framework for Large-scale Parallel File Systems
abstract
The imbalanced I/O load on large parallel file systems affects the parallel I/O performance of high-performance computing (HPC) applications. One of the main reasons for I/O imbalances is the lack of a global view of system-wide resource consumption. While approaches to address the problem already exist, the diversity of HPC workloads combined with different file striping patterns prevents widespread adoption of these approaches. In addition, load-balancing techniques should be transparent to client applications. To address these issues, we propose Tarazu , an end-to-end control plane where clients transparently and adaptively write to a set of selected I/O servers to achieve balanced data placement. Our control plane leverages real-time load statistics for global data placement on distributed storage servers, while our design model employs trace-based optimization techniques to minimize latency for I/O load requests between clients and servers and to handle multiple striping patterns in files. We evaluate our proposed system on an experimental cluster for two common use cases: the synthetic I/O benchmark IOR and the scientific application I/O kernel HACC-I/O. We also use a discrete-time simulator with real HPC application traces from emerging workloads running on the Summit supercomputer to validate the effectiveness and scalability of Tarazu in large-scale storage environments. The results show improvements in load balancing and read performance of up to 33% and 43%, respectively, compared to the state-of-the-art.
Arnab Kumar Paul, Sarah Neuwirth, Bharti Wadhwa, Feiyi Wang, Sarp Oral, Ali Raza Butt
ACM Trans. Storage2
2023 Toward Reproducible Benchmarking of PGAS and MPI Communication Schemes
abstract
With the forthcoming age of exascale computing, the efficient support of different programming models has become a crucial performance factor for high-performance computing systems. When adopting novel programming paradigms, the performance assessment through reproducible and comparable benchmarks plays an essential role in both the effective use of the heterogeneous system hardware and the application performance tuning. Alternatives to MPI such as the partitioned global address space (PGAS) model have become increasingly popular. One such PGAS API is the Global Address Space Programming Interface (GASPI). This paper introduces the GASPI Benchmark Suite (GBS), which combines a comprehensive set of microbenchmarks with application kernels. The microbenchmarks target common GASPI communication patterns, including one-sided, collective, passive, and global atomics, while the application kernels stress communication schemes commonly found in real HPC applications. The effectiveness of GBS is demonstrated by evaluating the GASPI communication performance for the networking communication standard InfiniBand.
Niklas Bartelheimer, Sarah Neuwirth
ICPADS2
2023 Characterization of Large-scale HPC Workloads with non-naïve I/O Roofline Modeling and Scoring
abstract
This paper introduces a novel approach to characterize system and application performance in high performance computing (HPC) systems. Traditional metrics such as computations and memory accesses alone are no longer sufficient to evaluate the performance of such systems. To address this challenge, an empirical I/O Roofline model and corresponding workload analysis workflow are proposed that can be adapted in the future to enable multidimensional evaluation for application performance characterization across different HPC systems. The model focuses on commonly used performance metrics such as I/O operations per second (IOPS) and I/O bandwidth, and leverages the well-known Roofline modeling technique to intuitively characterize I/O performance and identify performance bottlenecks without requiring deep knowledge of the I/O stack. Furthermore, based on the I/O Roofline model, a scoring approach is described that provides a unified and comprehensible method for evaluating the performance of different systems and applications. The effectiveness of the approach is demonstrated by evaluating the performance of various application I/O kernels using the empirical Roofline model.
Zhaobin Zhu, Sarah Neuwirth
ICPADS2
2023 Modeling the Impact of System-Level Parameters on I/O Performance of HPC Applications
abstract
Modern High Performance Computing (HPC) workloads prioritize data, challenging storage systems to meet rising I/O demands. The network's role in inter-node communication and data transfer significantly impacts overall application performance. Our study systematically examines bandwidth and latency variations in CephFS due to diverse I/O patterns. We formalize latency, bandwidth, I/O size, and block-to-file size ratio relationships and train a model for predicting bandwidth and estimating latency in similar I/O workloads, offering valuable insights into optimizing HPC storage.
Debasmita Biswas, Arnab Kumar Paul, Sarah Neuwirth, Ali Raza Butt
MASCOTS3
2022 Assessment of the I/O and Storage Subsystem in Modular Supercomputing Architectures
abstract
The Modular Supercomputing Architecture (MSA) plays a key role in Europe's Exascale computing strategy. Developed in the European project series DEEP, MSA breaks with traditional HPC architectures by integrating heterogeneous computing resources in a modular way at the system level, organizing them in compute modules with different hardware and performance characteristics. Heterogeneous applications and workflows can be run on exactly matching computing resources, therefore improving the time to solution and energy use. While mapping different code parts on the best-suited compute modules of the DEEP-EST prototype has been evaluated yet, the parallel I/O and storage environment has been neglected. Therefore, this position paper makes a first attempt to provide a complete overview and analysis of the complex I/O subsystem of the DEEP-EST prototype and IO-SEA, the main EU project addressing I/O and data management in MSA systems.
Sarah Neuwirth
CLUSTER1
2022 A Comprehensive I/O Knowledge Cycle for Modular and Automated HPC Workload Analysis
abstract
On the way to the exascale era, millions of parallel processing elements are required. Accordingly, one major chal-lenge is the ever-widening gap between computational power and underlying I/O systems. To bridge this gap, I/O resources must be used efficiently, thus a profound I/O knowledge is required. In this work, we analyze state-of-the-art approaches that can be applied to improve the general I/O understanding and performance. Based on our analysis, we present an automated, modular, tool-agnostic I/Oanalysis workflow and a prototype implementation that can be used to generate, extract, store, analyze, and use I/O knowledge in a structured and reproducible way.
Zhaobin Zhu, Sarah Neuwirth, Thomas Lippert
CLUSTER2
2021 Toward a Comprehensive Benchmark Suite for Evaluating GASPI in HPC Environments
abstract
With the forthcoming age of exascale computing, the efficient support of different programming models has become a crucial performance factor for high-performance computing systems. When adopting novel programming paradigms, the performance assessment through standardized and comparable benchmarks plays an essential role in both the effective use of the heterogeneous system hardware and the application performance tuning. Alternatives to MPI such as the partitioned global address space (PGAS) model have become increasingly popular. One such PGAS API is the Global Address Space Programming Interface (GASPI). This paper introduces the GASPI Benchmark Suite (GBS), which combines a comprehensive, standardized set of microbenchmarks with application kernels. The microbenchmarks target common GASPI communication patterns, including onesided, collective, passive, and global atomics, while the application kernels stress communication schemes commonly found in real HPC applications. The effectiveness of GBS is demonstrated by evaluating the GASPI communication performance for the two networking communication standards Ethernet and InfiniBand.
Sarah Neuwirth
CLUSTER1
2021 Parallel I/O Evaluation Techniques and Emerging HPC Workloads: A Perspective
abstract
Emerging workloads such as artificial intelligence, big data analytics and complex multi-step workflows alongside future exascale applications are anticipated future HPC workloads, which will result in a more diverse I/O system workload and even less predictable I/O behavior and access patterns. Along with the ever increasing gap between the compute and storage performance capabilities, the in-depth understanding of extreme-scale I/O behavior and the I/O performance modeling and prediction are essential tools of the large-scale I/O evaluation process for addressing the needs of extreme-scale hybrid workloads. In this survey article, we focus on the state-of-the-art of the I/O behavior and performance analysis process for HPC systems in a 5-year time window and identify future research challenges.
Sarah Neuwirth, Arnab Kumar Paul
CLUSTER1
2019 iez: Resource Contention Aware Load Balancing for Large-Scale Parallel File Systems
abstract
Parallel I/O performance is crucial to sustaining scientific applications on large-scale High-Performance Computing (HPC) systems. However, I/O load imbalance in the underlying distributed and shared storage systems can significantly reduce overall application performance. There are two conflicting challenges to mitigate this load imbalance: (i) optimizing systemwide data placement to maximize the bandwidth advantages of distributed storage servers, i.e., allocating I/O resources efficiently across applications and job runs; and (ii) optimizing client-centric data movement to minimize I/O load request latency between clients and servers, i.e., allocating I/O resources efficiently in service to a single application and job run. Moreover, existing approaches that require application changes limit wide-spread adoption in commercial or proprietary deployments. We propose iez, an “end-to-end control plane” where clients transparently and adaptively write to a set of selected I/O servers to achieve balanced data placement. Our control plane leverages realtime load information for distributed storage server global data placement while our design model leverages trace-based optimization techniques to minimize I/O load request latency between clients and servers. We evaluate our proposed system on an experimental cluster for two common use cases: synthetic I/O benchmark IOR for large sequential writes and a scientific application I/O kernel, HACC-I/O. Results show read and write performance improvements of up to 34% and 32%, respectively, compared to the state of the art.
Bharti Wadhwa, Arnab Kumar Paul, Sarah Neuwirth, Feiyi Wang, Sarp Oral, Ali Raza Butt, Jon Bernard, Kirk W. Cameron
IPDPS3
2017 Automatic and Transparent Resource Contention Mitigation for Improving Large-Scale Parallel File System Performance
abstract
Proportional to the scale increases in HPC systems, many scientific applications are becoming increasingly data intensive, and parallel I/O has become one of the dominant factors impacting the large-scale HPC application performance. On a typical large-scale HPC system, we have observed that the lack of a global workload coordination coupled with the shared nature of storage systems cause load imbalance and resource contention over the end-to-end I/O paths resulting in severe performance degradation. I/O load imbalance on HPC systems is generally a self-inflicted wound and mostly occurs between the I/O paths and resources consumed by each individual job. In this paper, we introduce TAPP-IO, a dynamic, shared load balancing framework for mitigating resource contention. TAPP-IO extends our previous work and solves two major limitations: First, it transparently intercepts file creation calls during runtime to balance the workload over all available storage targets. The usage of TAPP-IO requires no application source code modifications and is independent from any I/O middleware. The framework can be applied to almost any HPC platform and is suitable for systems that lack a centralized file system resource manager. Second, the framework proposes a new placement strategy to support not only file-per-process I/O, but also single shared file I/O. This opens the door to a new class of scientific applications that can leverage the placement library for improved performance. We demonstrate the effectiveness of our integration on the Titan system at the Oak Ridge National Laboratory. Our experiments with a synthetic benchmark and real-world HPC workload show that, even in a noisy production environment, TAPP-IO can improve large-scale application performance significantly.
Sarah Neuwirth, Feiyi Wang, Sarp Oral, Ulrich Brüning 0001
ICPADS1
2016 Using Balanced Data Placement to Address I/O Contention in Production Environments
abstract
Designed for capacity and capability, HPC I/O systems are inherently complex and shared among multiple, concurrent jobs competing for resources. Lack of centralized coordination and control often render the end-to-end I/O paths vulnerable to load imbalance and contention. With the emergence of data-intensive HPC applications, storage systems are further contended for performance and scalability. This paper proposes to unify two key approaches to tackle the imbalanced use of I/O resources and to achieve an end-to-end I/O performance improvement in the most transparent way. First, it utilizes a topology-aware, Balanced Placement I/O method (BPIO) for mitigating resource contention. Second, it takes advantage of the platform-neutral ADIOS middleware, which provides a flexible I/O mechanism for scientific applications. By integrating BPIO with ADIOS, referred to as Aequilibro, we obtain an end-to-end and per job I/O performance improvement for ADIOS-enabled HPC applications without requiring any code changes. Aequilibro can be applied to almost any HPC platform and is mostly suitable for systems that lack a centralized file system resource manager. We demonstrate the effectiveness of our integration on the Titan system at the Oak Ridge National Laboratory. Our experiments with a synthetic benchmark and real-world HPC workload show that, even in a noisy production environment, Aequilibro can improve large-scale application performance significantly.
Sarah Neuwirth, Feiyi Wang, Sarp Oral, Sudharshan S. Vazhkudai, James H. Rogers, Ulrich Brüning 0001
SBAC-PAD1
2015 Scalable communication architecture for network-attached accelerators
abstract
On the road to Exascale computing, novel communication architectures are required to overcome the limitations of host-centric accelerators. Typically, accelerator devices require a local host CPU to configure and operate them. This limits the number of accelerators per host system. Network-attached accelerators are a new architectural approach for scaling the number of accelerators and host CPUs independently. In this paper, the communication architecture for network-attached accelerators is described which enables remote initialization and control of the accelerator devices. Furthermore, an operative prototype implementation is presented. The prototype accelerator node consists of an Intel Xeon Phi coprocessor and an EXTOLL NIC. The EXTOLL interconnect provides new features to enable accelerator-to-accelerator direct communication without a local host. Workloads can be dynamically assigned to CPUs and accelerators at run-time in an N to M ratio. The latency, bandwidth, and performance of the low-level implementation and MPI communication layer are presented. The LAMMPS molecular dynamics simulator is used to evaluate the communication architecture. The internode communication time is improved by up to 47%.
Sarah Neuwirth, Dirk Frey, Mondrian Nüssle, Ulrich Brüning 0001
HPCA1
2015 Communication Models for Distributed Intel Xeon Phi Coprocessors
abstract
The emergence of accelerator technology in current supercomputing systems is changing the landscape of supercom-puting architectures. Accelerators like GPGPUs and coprocessors are optimized for parallel computation while being more energy efficient. Their computational power per watt plays a crucial role in developing exaflop systems. Today's accelerators come with some limitations. They require a local host to configure and operate them. In addition, the number of host CPUs and accelerators does not scale independently. Another problem is the unbalanced communication between distributed accelerators. New communication frameworks are developed to optimize the internode communication. In this paper, four communication models using the Intel Xeon Phi coprocessor technology are compared. The Intel Xeon Phi coprocessor is based on the Intel Many Integrated Cores technology. It is an attractive accelerator due to its embedded Linux operating system, up to 1 TFLOPS of performance on a single chip, and its x86 64 compatibility. DCFA-MPI, MVAPICH2-MIC, and HAM-Offload are compared against the communication architecture for network-attached accelerators (NAA). Each communication model optimizes a different layer of the MIC communication architecture. The NAA approach makes the accelerator device independent from a local host system. Furthermore, it enables the accelerator to source and sink network traffic. Workloads can be dynamically assigned during run-time in an N to M ratio between CPUs and accelerators. The latency, bandwidth, and performance of the MPI communication layer of a prototype implementation are evaluated.
Sarah Neuwirth, Dirk Frey, Ulrich Brüning 0001
ICPADS1