EDBT 2026 Demo / reviewers in the wild / expert
Andreas Knüpfer
dblp:32/6648
· DBLP profile ↗
14ranked-venue papers
3as first author
0since 2021 · last 2020
0000-0003-3591-397XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 3 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Storage systems · 61% Performance modeling and evaluation · 13% Parallel and multicore computing · 13% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems › i/o optimization
i/o forwarding |
0.1 | 1 | 2012 | Enabling event tracing at leadership-class scale through I/O forwarding middleware · HPDC 2012 |
Storage systems › file systems › distributed file system
parallel file system |
0.1 | 1 | 2012 | Enabling event tracing at leadership-class scale through I/O forwarding middleware · HPDC 2012 |
Parallel and multicore computing › parallel computing
parallel program analysis |
0.1 | 1 | 2006 | M09 - Program analysis tools for massively parallel applications: how to achieve highest performance · SC 2006 |
Performance modeling and evaluation
performance analysis tools |
0.1 | 1 | 2006 | M09 - Program analysis tools for massively parallel applications: how to achieve highest performance · SC 2006 |
High-performance computing › supercomputing
leadership-class systems |
0.0 | 1 | 2012 | Enabling event tracing at leadership-class scale through I/O forwarding middleware · HPDC 2012 |
High-performance computing
performance optimization |
0.0 | 1 | 2006 | M09 - Program analysis tools for massively parallel applications: how to achieve highest performance · SC 2006 |
Methods — techniques the papers use, named apart from their topics
write buffering · 0.1log aggregation · 0.1profiling · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | PIKA: Center-Wide and Job-Aware Cluster MonitoringabstractNowadays, performance optimization is more or less an established procedure in high-performance computing (HPC) centers. To sustainably increase compute efficiency of such systems, we need to increase the awareness of efficiency on both the operator's and the users' side. Therefore, we propose an infrastructure for continuous monitoring and analysis, which automatically characterizes HPC jobs and provides a systematic approach to identify underperforming compute jobs with optimization potential. The recorded metadata and time-series data can be visualized live at runtime or post-mortem and are eventually stored for long-term analysis. The monitoring has a negligible overhead on the compute nodes and neither influences nor limits the user applications. Robert Dietrich, Andreas Knüpfer, Wolfgang E. Nagel |
CLUSTER | 3 |
| 2020 | From stirring to mixing: artificial intelligence in the process industryabstractThe introduction of AI methods in production or production-related environments meets with resistance from operators due to their lack of relevant experience and their responsibility for plant safety. To overcome these inhibitions one requires prototypical implementations, which offer considerable benefits, meet the highest requirements for reliability, are accepted by the operating personnel, and get support from those in charge. As a result, AI technologies must be embedded into the complex IT/OT infrastructure of the companies. Traceability, maintainability and longevity must also be guaranteed. As a first step towards this, we present a concept of the demonstrator and its first results, which should make AI comprehensible by visualizing challenges and exploring possibilities in the process industry. Valentin Khaydarov, Sebastian Heinze, Markus Graube, Andreas Knüpfer, Maximilian Knespel, Silke Merkelbach, Leon Urbas |
ETFA | 4 |
| 2017 | Automatic Adaption of the Sampling Frequency for Detailed Performance AnalysisabstractOne of the most urgent challenges in event based performance analysis is the enormous amount of collected data. Combining event tracing and periodic sampling has been a successful approach to allow a detailed event-based recording of MPI communication and a coarse recording of the remaining application with periodic sampling. In this paper, we present a novel approach to automatically adapt the sampling frequency during runtime to the given amount of buffer space, releasing users to find an appropriate sampling frequency themselves. This way, the entire measurement can be kept within a single memory buffer, which avoids disruptive intermediate memory buffer flushes, excessive data volumes, and measurement delays due to slow file system interaction. We describe our approach to sort and store samples based on their order of occurrence in an hierarchical array based on powers of two. Furthermore, we evaluate the feasibility as well as the overhead of the approach with the prototype implementation OTFX based on the Open Trace Format 2, a state-of-the-art Open Source event trace library used by the performance analysis tools Vampir, Scalasca, and Tau. Michael Wagner 0003, Andreas Knüpfer |
CCGrid | 2 |
| 2015 | MPI-focused Tracing with OTFX: An MPI-aware In-memory Event Tracing Extension to the Open Trace Format 2abstractPerformance analysis tools are more than ever inevitable to develop applications that utilize the enormous computing resources of high performance computing (HPC) systems. In event-based performance analysis the amount of collected data is one of the most urgent challenges. The resulting measurement bias caused by uncoordinated intermediate memory buffer flushes in the monitoring tool can render a meaningful analysis of the parallel behavior impossible. In this paper we address the impact of intermediate memory buffer flushes and present a method to avoid file interaction in the monitoring tool entirely. We propose an MPI-focused tracing approach that provides the complete MPI communication behavior and adapts the remaining application events to an amount that fits into a single memory buffer. We demonstrate the capabilities of our method with an MPI-focused prototype implementation of OTFX, based on the Open Trace Format 2, a state-of-the-art Open Source event tracing library used by the performance analysis tools Vampir, Scalasca, and Tau. In a comparison to OTF2 based on seven applications from different scientific domains, our prototype introduces in average 5.1% less overhead and reduces the trace size up to three orders of magnitude. Michael Wagner 0003, Jens Doleschal, Andreas Knüpfer |
EuroMPI | 3 |
| 2013 | Hierarchical Memory Buffering Techniques for an In-Memory Event Tracing Extension to the Open Trace Format 2abstractOne of the most urgent challenges in event based performance analysis is the enormous amount of collected data. A real-time event reduction is crucial to enable a complete in-memory event tracing workflow, which circumvents the limitations of current parallel file systems to support event tracing on large scale systems. However, a traditional single flat memory buffer fails to support real-time event reduction. To address this issue, we present a hierarchical memory buffer, which is capable to support event reduction. We show that this hierarchical memory buffer does not introduce additional overhead, regardless of the grade of reduction. In addition, we evaluate its main parameter: the size of the internal memory bins. The hierarchical memory buffer is based on the Open Trace Format 2, a state-of-the-art Open Source event trace library used by the performance analysis tools VAMPIR, SCALASCA, and TAU. Michael Wagner 0003, Andreas Knüpfer, Wolfgang E. Nagel |
ICPP | 2 |
| 2013 | Runtime message uniquification for accurate communication analysis on incomplete MPI event tracesabstractCommunication analysis of parallel applications based on event traces depends on correct matching of associated MPI send and receive events. Selective monitoring techniques, however, may result in incomplete MPI event traces and, in that case, current matching strategies fail. In this paper we introduce an additional unique identifier for each message to make MPI events distinguishable from others. Therefore, it is possible to identify missing MPI events and match all remaining MPI events correctly. An overhead study with a real-life application and a benchmark suite demonstrates the applicability and benefits of this approach. Michael Wagner 0003, Jens Doleschal, Wolfgang E. Nagel, Andreas Knüpfer |
EuroMPI | 4 |
| 2012 | Enabling event tracing at leadership-class scale through I/O forwarding middlewareabstractEvent tracing is an important tool for understanding the performance of parallel applications. As concurrency increases in leadership-class computing systems, the quantity of performance log data can overload the parallel file system, perturbing the application being observed. In this work we present a solution for event tracing at leadership scales. We enhance the I/O forwarding system software to aggregate and reorganize log data prior to writing to the storage system, significantly reducing the burden on the underlying file system for this type of traffic. Furthermore, we augment the I/O forwarding system with a write buffering capability to limit the impact of artificial perturbations from log data accesses on traced applications. To validate the approach, we modify the Vampir tracing toolset to take advantage of this new capability and show that the approach increases the maximum traced application size by a factor of 5x to more than 200,000 processes. Thomas Ilsche, Joseph Schuchart, Jason Cope, Dries Kimpe, Terry R. Jones, Andreas Knüpfer, Kamil Iskra, Robert B. Ross, Wolfgang E. Nagel, Stephen W. Poole |
HPDC | 6 |
| 2012 | Holistic Debugging of MPI Derived DatatypesabstractThe Message Passing Interface (MPI) specifies an API that allows programmers to create efficient and scalable parallel applications. The standard defines multiple constraints for each function parameter. For performance reasons, no MPI implementation checks all of these constraints at runtime. Derived data types are an important concept of MPI and allow users to describe an application's data structures for efficient and convenient communication. Using existing infrastructure we present scalable algorithms to detect usage errors of basic and derived MPI data types. We detect errors that include constraints for construction and usage of derived data types, matching their type signatures in communication, and detecting erroneous overlaps of communication buffers. We implement these checks in the MUST runtime error detection framework. We provide a novel representation of error locations to highlight usage errors. Further, approaches to buffer overlap checking can cause unacceptable overheads for non-contiguous data types. We present an algorithm that uses patterns in derived MPI data types to avoid these overheads without losing precision. Application results for the benchmark suites SPEC MPI2007 and NAS Parallel Benchmarks for up to 2048 cores show that our approach applies to a broad range of applications and that our extended overlap check improves performance by two orders of magnitude. Finally, we augment our runtime error detection component with a debugger extension to support in-depth analysis of the errors that we find as well as semantic errors. This extension to gdb provides information about MPI data type handles and enables gdb -- and other debuggers based on gdb -- to display the content of a buffer as used in MPI communications. Joachim Jenke, Tobias Hilbrich, Andreas Knüpfer, Bronis R. de Supinski, Matthias S. Müller |
IPDPS | 3 |
| 2010 | Special section: Tools for program development and analysis in computational science
Jie Tao 0001, Arndt Bode, Andreas Knüpfer, Dieter Kranzlmüller, Jens Volkert, Roland Wismüller |
Future Gener. Comput. Syst. | 3 |
| 2009 | Pattern Matching and I/O Replay for POSIX I/O in Parallel Programs
Michael Kluge, Andreas Knüpfer, Matthias S. Müller, Wolfgang E. Nagel |
Euro-Par | 2 |
| 2006 | M09 - Program analysis tools for massively parallel applications: how to achieve highest performanceabstractToday's HPC environments are increasingly complex in order to achieve highest performance. Hardware platforms introduce features like out-of-order execution, multi-level caches, multi-cores, non-uniform memory access etc. Application software combines OpenMP, MPI, optimized libraries and various types of compiler optimization to exploit potential performance.To reach a reasonable percentage of the theoretical peak performance, three fundamental steps need to be accomplished. First, correctness must be guaranteed especially during the course of optimization. Second, the actual performance achieved needs to be determined. In particular the contributions/limitations of all sub-systems involved (CPU, memory, network, I/O) have to be identified. Third, actual optimization can only be successful with the previously obtained knowledge.Those steps are by no means trivial. There are sophisticated tools beyond simple profiling to support the HPC user. The tutorial introduces a variety of such tools: it shows how they play together and how they scale with long-running massively parallel cases. Andreas Knüpfer, Dieter Kranzlmüller, Bernd Mohr, Wolfgang E. Nagel |
SC | 1 |
| 2006 | Compressible memory data structures for event-based trace analysis
Andreas Knüpfer, Wolfgang E. Nagel |
Future Gener. Comput. Syst. | 1 |
| 2005 | Knowledge Based Automatic Scalability Analysis and Extrapolation for MPI Programs
Michael Kluge, Andreas Knüpfer, Wolfgang E. Nagel |
Euro-Par | 2 |
| 2005 | Construction and Compression of Complete Call Graphs for Post-Mortem Program Trace AnalysisabstractCompressed complete call graphs (cCCGs) are a newly developed memory data structure for event based program traces. The most important advantage over linear lists or arrays traditionally used is the ability to apply lossy or lossless data compression. The compression scheme is completely transparent with respect to read access decompression is not required. This approach is a new way to cope with todays challenges when analyzing enormous amounts of trace data. The article focuses on CCG construction and compression, querying and evaluation are briefly covered. Andreas Knüpfer, Wolfgang E. Nagel |
ICPP | 1 |