Holger Brunst

dblp:68/142 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
2since 2021 · last 2023
0000-0003-2224-0630ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 3 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 56% Storage systems · 35% High-performance computing · 9%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › data compression
file compression
0.712023
Rapidgzip: Parallel Decompression and Seeking in Gzip Files Using Cache Prefetching · HPDC 2023
Storage systems
file systems
0.712023
Rapidgzip: Parallel Decompression and Seeking in Gzip Files Using Cache Prefetching · HPDC 2023
Memory systems › cache
prefetching
0.712023
Rapidgzip: Parallel Decompression and Seeking in Gzip Files Using Cache Prefetching · HPDC 2023
Memory systems
memory-bound computation
0.412019
An early evaluation of Intel's optane DC persistent memory module and its impact on high-performance scientific applications · SC 2019
Memory systems
non-volatile memory
0.412019
An early evaluation of Intel's optane DC persistent memory module and its impact on high-performance scientific applications · SC 2019
Memory systems › non-volatile memory › persistent memory
optane persistent memory
0.412019
An early evaluation of Intel's optane DC persistent memory module and its impact on high-performance scientific applications · SC 2019
Memory systems › non-volatile memory
persistent memory
0.412019
An early evaluation of Intel's optane DC persistent memory module and its impact on high-performance scientific applications · SC 2019
High-performance computing
scientific computing systems
0.412019
An early evaluation of Intel's optane DC persistent memory module and its impact on high-performance scientific applications · SC 2019
Storage systems › object storage
distributed object store
0.112019
An early evaluation of Intel's optane DC persistent memory module and its impact on high-performance scientific applications · SC 2019
Memory systems › non-volatile memory
NVRAM
0.112019
An early evaluation of Intel's optane DC persistent memory module and its impact on high-performance scientific applications · SC 2019

Methods — techniques the papers use, named apart from their topics

performance evaluation · 0.4STREAM benchmark · 0.4
YearPublicationVenuePosition
2023 Rapidgzip: Parallel Decompression and Seeking in Gzip Files Using Cache Prefetching
abstract
Gzip is a file compression format, which is ubiquitously used. Although a multitude of gzip implementations exist, only pugz can fully utilize current multi-core processor architectures for decompression. Yet, pugz cannot decompress arbitrary gzip files. It requires the decompressed stream to only contain byte values 9-126. In this work, we present a generalization of the parallelization scheme used by pugz that can be reliably applied to arbitrary gzip-compressed data without compromising performance. We show that the requirements on the file contents posed by pugz can be dropped by implementing an architecture based on a cache and a parallelized prefetcher. This architecture can safely handle faulty decompression results, which can appear when threads start decompressing in the middle of a gzip file by using trial and error. Using 128 cores, our implementation reaches 8.7 GB/s decompression bandwidth for gzip-compressed base64-encoded data, a speedup of 55 over the single-threaded GNU gzip, and 5.6 GB/s for the Silesia corpus, a speedup of 33 over GNU gzip.
Maximilian Knespel, Holger Brunst
HPDC2
2022 First Experiences in Performance Benchmarking with the New SPEChpc 2021 Suites
abstract
Modern High Performance Computing (HPC) sys-tems are built with innovative system architectures and novel programming models to further push the speed limit of computing. The increased complexity poses challenges for performance portability and performance evaluation. The Standard Perfor-mance Evaluation Corporation (SPEC) has a long history of producing industry-standard benchmarks for modern computer systems. SPEC's newly released SPEChpc 2021 benchmark suites, developed by the High Performance Group, are a bold attempt to provide a fair and objective benchmarking tool designed for state-of-the-art HPC systems. With the support of multiple host and accelerator programming models, the suites are portable across both homogeneous and heterogeneous architectures. Different workloads are developed to fit system sizes ranging from a few compute nodes to a few hundred compute nodes. In this work we present our first experiences in performance benchmarking the new SPEChpc2021 suites and evaluate their portability and basic performance characteristics on various popular and emerging HPC architectures, including x86 CPU, NVIDIA GPU, and AMD GPU. This study provides a first-hand experience of executing the SPEChpc 2021 suites at scale on production HPC systems, discusses real-world use cases, and serves as an initial guideline for using the benchmark suites.
Holger Brunst, Sunita Chandrasekaran, Florina M. Ciorba, Nick Hagerty, Robert Henschel, Guido Juckeland, Junjie Li 0003, Verónica G. Vergara Larrea, Sandra Wienke, Miguel Zavala
CCGRID1
2019 An early evaluation of Intel's optane DC persistent memory module and its impact on high-performance scientific applications
abstract
Memory and I/O performance bottlenecks in supercomputing simulations are two key challenges that must be addressed on the road to Exascale. The new byte-addressable persistent non-volatile memory technology from Intel, DCPMM, promises to be an exciting opportunity to break with the status quo, with unprecedented levels of capacity at near-DRAM speeds. Here, we explore the potential of DCPMM in the context of two high-performance scientific applications in terms of outright performance, efficiency and usability for both its Memory and App Direct modes. In Memory mode, we show equivalent performance and better efficiency for a CASTEP simulation that is limited by memory capacity on conventional DRAM-only systems without any changes to the application. For IFS, we demonstrate that a distributed object-store over NVRAM reduces the data contention created in weather forecasting data producer-consumer workflows. In addition, we also present the achievable memory bandwidth performance using STREAM.
Michèle Weiland, Holger Brunst, Tiago Quintino, Nick Johnson, Olivier Iffrig, Simon D. Smart, Christian Herold, Antonino Bonanni, Adrian Jackson, Mark Parsons 0001
SC2
2017 Using adaptive runtime filtering to support an event-based performance analysis
abstract
Summary Event‐based performance monitoring and analysis are effective means when tuning parallel applications for optimal resource usage. In this article, we address the data capacity challenge that arises when applying the tracing methodology to large‐scale parallel applications and long execution times. Existing approaches use static, pre‐defined event filters to reduce the performance data to a manageable size. In contrast, we propose self‐guided filters that automatically adapt to an application's runtime behaviour and therefore, do not require any previous knowledge or application executions. Our contribution consists of four adaptive runtime filters, which target a specific type of data redundancy each. The filters focus on detecting identical events in loop iterations, constant events with no variation in time, and very short, highly frequent, typically not very meaningful events, having a severe impact on the total data volume. We evaluate our prototype implementation with five real‐world applications and achieve a data reduction of two orders of magnitude while increasing execution time less than 1%. Likewise, we show that the qualitative impact of our filters on performance analysis in state‐of‐the‐art analysis tools can be reduced by adding feedback methods and statistical information to the filtered traces. Copyright © 2017 John Wiley & Sons, Ltd.
Jonas Stolle, Michael Wagner 0003, Jens Doleschal, Felix Schmitt 0004, Holger Brunst
Concurr. Comput. Pract. Exp.5
2016 Structural Clustering: A New Approach to Support Performance Analysis at Scale
abstract
The increasing complexity of high performance computing systems creates high demands on performance tools and human analysts due to an unmanageable volume of data gathered for performance analysis. A promising approach for reducing data volume is classification of data from multiple processes into groups of similar behavior to aid in analyzing application performance and identifying hot spots. However, existing approaches for structural and temporal classification of performance data suffer from lack of scalability or produce misleading results. To address this problem, we present a novel and effective structural similarity measure to efficiently classify data from parallel processes and introduce a method for efficient storage of the classified data. Using four examples, we show how existing performance analysis techniques benefit from our structural classification. Finally, we present a case study with 15 applications on up to 65,536 parallel processes that demonstrates the generality and scalability of our classification approach.
Matthias Weber 0002, Ronny Brendel, Tobias Hilbrich, Kathryn Mohror, Martin Schulz 0001, Holger Brunst
IPDPS6
2015 Event-Action Mappings for Parallel Tools Infrastructures
Tobias Hilbrich, Martin Schulz 0001, Holger Brunst, Joachim Jenke, Bronis R. de Supinski, Matthias S. Müller
Euro-Par3
2013 Alignment-Based Metrics for Trace Comparison
Matthias Weber 0002, Kathryn Mohror, Martin Schulz 0001, Bronis R. de Supinski, Holger Brunst, Wolfgang E. Nagel
Euro-Par5
2012 Trace File Comparison with a Hierarchical Sequence Alignment Algorithm
abstract
Performance optimization, especially in the field of HPC, is an integral part of today's software development process. One powerful way of optimizing applications is to analyze their event traces. Yet, the comparison of traces of multiple application runs is cumbersome. The impact of optimizations in the source code or the usage of different compiler flags has to be tracked manually. The challenge is to automatically identify exactly those areas that changed in the large amount of trace data. We propose a novel solution that combines sequence alignment algorithms with call graph analysis to compare and highlight traces event-wise. Our approach is able to automatically detect differences by aligning event traces. Fine-grained execution time differences can be extracted and displayed in performance charts. The results of our implementation are presented and discussed.
Matthias Weber 0002, Ronny Brendel, Holger Brunst
ISPA3
2012 Performance analysis of multi-level parallelism: inter-node, intra-node and hardware accelerators
abstract
SUMMARY The advent of multi‐core processors has made parallel computing techniques mandatory on mainstream systems. With the recent rise in hardware accelerators, hybrid parallelism adds yet another dimension of complexity to the process of software development. The inner workings of a parallel program are usually difficult to understand and verify. This paper presents a tool for graphical program flow analysis of hardware accelerated parallel programs. It monitors the hybrid program execution to record and visualize many performance relevant events along the way. Representative real‐world applications written for both IBM's Cell processor and NVIDIA's CUDA API are studied exemplarily. With our combined monitoring and visualization approach for hardware accelerated multi‐core and multi‐node systems we take the next step in tool evolution towards a highly improved level of detail, precision, and completeness. The contents of this paper is of interest to developers of hardware accelerated applications as well as performance tool architects. Copyright © 2011 John Wiley & Sons, Ltd.
Daniel Hackenberg, Guido Juckeland, Holger Brunst
Concurr. Comput. Pract. Exp.3
2010 High Resolution Program Flow Visualization of Hardware Accelerated Hybrid Multi-core Applications
abstract
The advent of multi-core processors has made parallel computing techniques mandatory on main stream systems. With the recent rise of hardware accelerators, hybrid parallelism adds yet another dimension of complexity to the process of software development. This article presents a tool for graphical program flow analysis of hardware accelerated parallel programs. It monitors the hybrid program execution to record and visualize many performance relevant events along the way. Representative real-world applications written for both IBM's Cell processor and NVIDIA's CUDA API are studied exemplarily. To the best of our knowledge, this approach is the first that visualizes the parallelism in hybrid multi-core systems at the presented level of detail.
Daniel Hackenberg, Guido Juckeland, Holger Brunst
CCGRID3
2008 Event Tracing and Visualization for Cell Broadband Engine Systems
Daniel Hackenberg, Holger Brunst, Wolfgang E. Nagel
Euro-Par2
2005 Monitoring cache behavior on parallel SMP architectures and related programming tools
Thomas Brandes, Helmut Schwamborn, Michael Gerndt, Jürgen Jeitner, Edmond Kereku, Martin Schulz 0001, Holger Brunst, Wolfgang E. Nagel, Reinhard Neumann, Ralph Müller-Pfefferkorn, Bernd Trenkler, Wolfgang Karl, Jie Tao 0001, Hans-Christian Hoppe
Future Gener. Comput. Syst.7
2003 A Distributed Performance Analysis Architecture for Clusters
abstract
The use of a cluster for distributed performance analysis of parallel trace data is discussed. We propose an analysis architecture that uses multiple cluster nodes as a server to execute analysis operations in parallel and communicate to remote clients where performance visualization and user interactions occur. The client-server system developed, VNG, is highly configurable and is shown to perform well for traces of large size, when compared to leading trace visualization systems.
Holger Brunst, Wolfgang E. Nagel, Allen D. Malony
CLUSTER1
2001 Group-Based Performance Analysis for Multithreaded SMP Cluster Applications
Holger Brunst, Wolfgang E. Nagel, Hans-Christian Hoppe
Euro-Par1