Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Andreas Knüpfer

dblp:32/6648 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
0since 2021 · last 2020
0000-0003-3591-397XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Storage systems · 61% Performance modeling and evaluation · 13% Parallel and multicore computing · 13%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › i/o optimization
i/o forwarding
0.112012
Enabling event tracing at leadership-class scale through I/O forwarding middleware · HPDC 2012
Storage systems › file systems › distributed file system
parallel file system
0.112012
Enabling event tracing at leadership-class scale through I/O forwarding middleware · HPDC 2012
Parallel and multicore computing › parallel computing
parallel program analysis
0.112006
M09 - Program analysis tools for massively parallel applications: how to achieve highest performance · SC 2006
Performance modeling and evaluation
performance analysis tools
0.112006
M09 - Program analysis tools for massively parallel applications: how to achieve highest performance · SC 2006
High-performance computing › supercomputing
leadership-class systems
0.012012
Enabling event tracing at leadership-class scale through I/O forwarding middleware · HPDC 2012
High-performance computing
performance optimization
0.012006
M09 - Program analysis tools for massively parallel applications: how to achieve highest performance · SC 2006

Methods — techniques the papers use, named apart from their topics

write buffering · 0.1log aggregation · 0.1profiling · 0.1
YearPublicationVenuePosition
2020 PIKA: Center-Wide and Job-Aware Cluster Monitoring
abstract
Nowadays, performance optimization is more or less an established procedure in high-performance computing (HPC) centers. To sustainably increase compute efficiency of such systems, we need to increase the awareness of efficiency on both the operator's and the users' side. Therefore, we propose an infrastructure for continuous monitoring and analysis, which automatically characterizes HPC jobs and provides a systematic approach to identify underperforming compute jobs with optimization potential. The recorded metadata and time-series data can be visualized live at runtime or post-mortem and are eventually stored for long-term analysis. The monitoring has a negligible overhead on the compute nodes and neither influences nor limits the user applications.
Robert Dietrich, Andreas Knüpfer, Wolfgang E. Nagel
CLUSTER3
2020 From stirring to mixing: artificial intelligence in the process industry
abstract
The introduction of AI methods in production or production-related environments meets with resistance from operators due to their lack of relevant experience and their responsibility for plant safety. To overcome these inhibitions one requires prototypical implementations, which offer considerable benefits, meet the highest requirements for reliability, are accepted by the operating personnel, and get support from those in charge. As a result, AI technologies must be embedded into the complex IT/OT infrastructure of the companies. Traceability, maintainability and longevity must also be guaranteed. As a first step towards this, we present a concept of the demonstrator and its first results, which should make AI comprehensible by visualizing challenges and exploring possibilities in the process industry.
Valentin Khaydarov, Sebastian Heinze, Markus Graube, Andreas Knüpfer, Maximilian Knespel, Silke Merkelbach, Leon Urbas
ETFA4
2017 Automatic Adaption of the Sampling Frequency for Detailed Performance Analysis
abstract
One of the most urgent challenges in event based performance analysis is the enormous amount of collected data. Combining event tracing and periodic sampling has been a successful approach to allow a detailed event-based recording of MPI communication and a coarse recording of the remaining application with periodic sampling. In this paper, we present a novel approach to automatically adapt the sampling frequency during runtime to the given amount of buffer space, releasing users to find an appropriate sampling frequency themselves. This way, the entire measurement can be kept within a single memory buffer, which avoids disruptive intermediate memory buffer flushes, excessive data volumes, and measurement delays due to slow file system interaction. We describe our approach to sort and store samples based on their order of occurrence in an hierarchical array based on powers of two. Furthermore, we evaluate the feasibility as well as the overhead of the approach with the prototype implementation OTFX based on the Open Trace Format 2, a state-of-the-art Open Source event trace library used by the performance analysis tools Vampir, Scalasca, and Tau.
Michael Wagner 0003, Andreas Knüpfer
CCGrid2
2015 MPI-focused Tracing with OTFX: An MPI-aware In-memory Event Tracing Extension to the Open Trace Format 2
abstract
Performance analysis tools are more than ever inevitable to develop applications that utilize the enormous computing resources of high performance computing (HPC) systems. In event-based performance analysis the amount of collected data is one of the most urgent challenges. The resulting measurement bias caused by uncoordinated intermediate memory buffer flushes in the monitoring tool can render a meaningful analysis of the parallel behavior impossible. In this paper we address the impact of intermediate memory buffer flushes and present a method to avoid file interaction in the monitoring tool entirely. We propose an MPI-focused tracing approach that provides the complete MPI communication behavior and adapts the remaining application events to an amount that fits into a single memory buffer. We demonstrate the capabilities of our method with an MPI-focused prototype implementation of OTFX, based on the Open Trace Format 2, a state-of-the-art Open Source event tracing library used by the performance analysis tools Vampir, Scalasca, and Tau. In a comparison to OTF2 based on seven applications from different scientific domains, our prototype introduces in average 5.1% less overhead and reduces the trace size up to three orders of magnitude.
Michael Wagner 0003, Jens Doleschal, Andreas Knüpfer
EuroMPI3
2013 Hierarchical Memory Buffering Techniques for an In-Memory Event Tracing Extension to the Open Trace Format 2
abstract
One of the most urgent challenges in event based performance analysis is the enormous amount of collected data. A real-time event reduction is crucial to enable a complete in-memory event tracing workflow, which circumvents the limitations of current parallel file systems to support event tracing on large scale systems. However, a traditional single flat memory buffer fails to support real-time event reduction. To address this issue, we present a hierarchical memory buffer, which is capable to support event reduction. We show that this hierarchical memory buffer does not introduce additional overhead, regardless of the grade of reduction. In addition, we evaluate its main parameter: the size of the internal memory bins. The hierarchical memory buffer is based on the Open Trace Format 2, a state-of-the-art Open Source event trace library used by the performance analysis tools VAMPIR, SCALASCA, and TAU.
Michael Wagner 0003, Andreas Knüpfer, Wolfgang E. Nagel
ICPP2
2013 Runtime message uniquification for accurate communication analysis on incomplete MPI event traces
abstract
Communication analysis of parallel applications based on event traces depends on correct matching of associated MPI send and receive events. Selective monitoring techniques, however, may result in incomplete MPI event traces and, in that case, current matching strategies fail. In this paper we introduce an additional unique identifier for each message to make MPI events distinguishable from others. Therefore, it is possible to identify missing MPI events and match all remaining MPI events correctly. An overhead study with a real-life application and a benchmark suite demonstrates the applicability and benefits of this approach.
Michael Wagner 0003, Jens Doleschal, Wolfgang E. Nagel, Andreas Knüpfer
EuroMPI4
2012 Enabling event tracing at leadership-class scale through I/O forwarding middleware
abstract
Event tracing is an important tool for understanding the performance of parallel applications. As concurrency increases in leadership-class computing systems, the quantity of performance log data can overload the parallel file system, perturbing the application being observed. In this work we present a solution for event tracing at leadership scales. We enhance the I/O forwarding system software to aggregate and reorganize log data prior to writing to the storage system, significantly reducing the burden on the underlying file system for this type of traffic. Furthermore, we augment the I/O forwarding system with a write buffering capability to limit the impact of artificial perturbations from log data accesses on traced applications. To validate the approach, we modify the Vampir tracing toolset to take advantage of this new capability and show that the approach increases the maximum traced application size by a factor of 5x to more than 200,000 processes.
Thomas Ilsche, Joseph Schuchart, Jason Cope, Dries Kimpe, Terry R. Jones, Andreas Knüpfer, Kamil Iskra, Robert B. Ross, Wolfgang E. Nagel, Stephen W. Poole
HPDC6
2012 Holistic Debugging of MPI Derived Datatypes
abstract
The Message Passing Interface (MPI) specifies an API that allows programmers to create efficient and scalable parallel applications. The standard defines multiple constraints for each function parameter. For performance reasons, no MPI implementation checks all of these constraints at runtime. Derived data types are an important concept of MPI and allow users to describe an application's data structures for efficient and convenient communication. Using existing infrastructure we present scalable algorithms to detect usage errors of basic and derived MPI data types. We detect errors that include constraints for construction and usage of derived data types, matching their type signatures in communication, and detecting erroneous overlaps of communication buffers. We implement these checks in the MUST runtime error detection framework. We provide a novel representation of error locations to highlight usage errors. Further, approaches to buffer overlap checking can cause unacceptable overheads for non-contiguous data types. We present an algorithm that uses patterns in derived MPI data types to avoid these overheads without losing precision. Application results for the benchmark suites SPEC MPI2007 and NAS Parallel Benchmarks for up to 2048 cores show that our approach applies to a broad range of applications and that our extended overlap check improves performance by two orders of magnitude. Finally, we augment our runtime error detection component with a debugger extension to support in-depth analysis of the errors that we find as well as semantic errors. This extension to gdb provides information about MPI data type handles and enables gdb -- and other debuggers based on gdb -- to display the content of a buffer as used in MPI communications.
Joachim Jenke, Tobias Hilbrich, Andreas Knüpfer, Bronis R. de Supinski, Matthias S. Müller
IPDPS3
2010 Special section: Tools for program development and analysis in computational science
Jie Tao 0001, Arndt Bode, Andreas Knüpfer, Dieter Kranzlmüller, Jens Volkert, Roland Wismüller
Future Gener. Comput. Syst.3
2009 Pattern Matching and I/O Replay for POSIX I/O in Parallel Programs
Michael Kluge, Andreas Knüpfer, Matthias S. Müller, Wolfgang E. Nagel
Euro-Par2
2006 M09 - Program analysis tools for massively parallel applications: how to achieve highest performance
abstract
Today's HPC environments are increasingly complex in order to achieve highest performance. Hardware platforms introduce features like out-of-order execution, multi-level caches, multi-cores, non-uniform memory access etc. Application software combines OpenMP, MPI, optimized libraries and various types of compiler optimization to exploit potential performance.To reach a reasonable percentage of the theoretical peak performance, three fundamental steps need to be accomplished. First, correctness must be guaranteed especially during the course of optimization. Second, the actual performance achieved needs to be determined. In particular the contributions/limitations of all sub-systems involved (CPU, memory, network, I/O) have to be identified. Third, actual optimization can only be successful with the previously obtained knowledge.Those steps are by no means trivial. There are sophisticated tools beyond simple profiling to support the HPC user. The tutorial introduces a variety of such tools: it shows how they play together and how they scale with long-running massively parallel cases.
Andreas Knüpfer, Dieter Kranzlmüller, Bernd Mohr, Wolfgang E. Nagel
SC1
2006 Compressible memory data structures for event-based trace analysis
Andreas Knüpfer, Wolfgang E. Nagel
Future Gener. Comput. Syst.1
2005 Knowledge Based Automatic Scalability Analysis and Extrapolation for MPI Programs
Michael Kluge, Andreas Knüpfer, Wolfgang E. Nagel
Euro-Par2
2005 Construction and Compression of Complete Call Graphs for Post-Mortem Program Trace Analysis
abstract
Compressed complete call graphs (cCCGs) are a newly developed memory data structure for event based program traces. The most important advantage over linear lists or arrays traditionally used is the ability to apply lossy or lossless data compression. The compression scheme is completely transparent with respect to read access decompression is not required. This approach is a new way to cope with todays challenges when analyzing enormous amounts of trace data. The article focuses on CCG construction and compression, querying and evaluation are briefly covered.
Andreas Knüpfer, Wolfgang E. Nagel
ICPP1