EDBT 2026 Demo / reviewers in the wild / expert
Joachim Jenke
dblp:118/5320 · also Joachim Protze
· DBLP profile ↗
20ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0003-0640-8966ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Review of MPI Continuations and Their Integration into PMPI Tools
Alexander Optenhöfel, Joachim Jenke, Ben Thärigen, Joseph Schuchart |
EuroMPI | 2 |
| 2025 | Verifying MPI API Usage Requirements with Contracts
Yussur Mustafa Oraji, Simon Schwitanski, Alexander Hück, Joachim Jenke, Sebastian Kreutzer, Christian H. Bischof |
EuroMPI | 4 |
| 2025 | MPI Finally Needs to Deal with Threads
Joseph Schuchart, Joachim Jenke, Simon Schwitanski |
EuroMPI | 2 |
| 2024 | RMASanitizer: Generalized Runtime Detection of Data Races in Remote Memory Access ApplicationsabstractRemote Memory Access (RMA) programming models enable processes running on a distributed-memory computer to access and manipulate the memory of other processes directly. Such one-sided communication has the benefit that the receiving process is not actively involved in the communication compared to the classical two-sided message-passing model. The three programming models MPI RMA, OpenSHMEM, and GASPI provide such a communication scheme. However, RMA models require the developer to synchronize the accesses with corresponding API calls correctly. Concurrent modifications of the same (remote) memory location due to wrong or missing synchronization lead to data races. Such data races are undefined behavior and may result in non-deterministic failures of the program execution. This paper presents RMASanitizer, an on-the-fly race detector for MPI RMA, OpenSHMEM, and GASPI applications. It relies on a generalized race detection model independent of the concrete RMA programming model. RMASanitizer combines a dynamic on-the-fly analysis with a static analysis at compile-time that detects and instruments only relevant memory accesses. It is implemented as part of the MPI correctness checking framework MUST which we extended with support for OpenSHMEM and GASPI. We show that RMASanitizer can detect races in MPI RMA, OpenSHMEM, and GASPI applications with an accuracy of over 95 percent by running it on the data race benchmark suite RMARaceBench. On proxy applications, the slowdown for the execution with up to 700 processes ranges from 1.1x to 30x, depending on the application, showing that our tool is applicable in practice. Simon Schwitanski, Yussur Mustafa Oraji, Cornelius Pätzold, Joachim Jenke, Felix Tomski, Matthias S. Müller |
ICPP | 4 |
| 2024 | MPI-BugBench: A Framework for Assessing MPI Correctness Tools
Tim Jammer, Emmanuelle Saillard, Simon Schwitanski, Joachim Jenke, Radjasouria Vinayagame, Alexander Hück, Christian H. Bischof |
EuroMPI | 4 |
| 2023 | Investigating the Usage of MPI at Argument-Granularity in HPC CodesabstractThis study focuses on gaining insights into the usage of the Message-Passing Interface (MPI) in a large set of High-Performance Computing (HPC) codes by analyzing MPI function calls and their argument usage patterns. Previous work has focused on analyzing MPI feature usage by statically matching function calls. However, this approach does not reveal common argument-specific call patterns or cross-interactions between MPI functions. In particular, MPI exposes its internal data structures using handles, and users pass these handles to MPI constructor functions, e.g., to create custom communicators. Tracking the relevant MPI arguments of these constructors and cross-referencing them with other MPI calls in a target code can reveal common user interactions. These insights can be used to optimize, e.g., datatype construction at a library level or to extend MPI correctness debugging tools to verify correct construction of these data structures. To that end, we statically analyze codes to extract MPI function calls and their arguments, cross-reference them with other MPI calls, and provide statistics on common argument patterns and cross-use of MPI functions. We believe that these insights can guide further development within the MPI community to ultimately benefit users. Alexander Hück, Tim Jammer, Joachim Jenke, Christian H. Bischof |
EuroMPI | 3 |
| 2023 | A Shim Layer for Transparently Adding Meta Data to MPI HandlesabstractMPI tool or abstraction libraries often have the need to bind meta data information to different kinds of MPI opaque handles. Communicator, window and datatype handles allow to associate key-value pairs as a mean to store meta information. The more short-living request handles, however, do not provide such functionality. As a result, several tool libraries use map-like data structures to bind internal meta information to MPI opaque handles and track those handles across their lifetime, using the handle value as a key. This results in several challenges. In this paper we make the case that request handles associated with different concurrent communication operations are not guaranteed to be unique when returned from the MPI library, so that a simple map may result in conflicts. Furthermore, MPI handles are not guaranteed to be constant over their lifetime, making the use of a map even more questionable. In this work, we present a shim layer wrapping MPI opaque handles that is transparent to the application and that allows tool and abstraction libraries to uniquely distinguish semantically different handles. At the same time, the handle shim layer allows to store the meta information in the wrapped handle avoiding the map-like data structures and also sand-boxes the MPI handle so that changes over its lifetime do not cause any harm. We provide a thread-safe proof-of-concept implementation for most relevant MPI-4 functions that can be used with multi-threaded MPI applications. The implementation transparently supports the different underlying base-language types for handles chosen by the MPI implementations. We evaluate the integration into several tools and abstraction libraries. In an overhead evaluation on synthetic benchmarks, our handle shim layer reduces the runtime overhead compared to a map-like data structure by about in most single-threaded cases. Joachim Jenke, Michael Knobloch, Marc-André Hermanns, Simon Schwitanski |
EuroMPI | 1 |
| 2022 | On-the-Fly Calculation of Model Factors for Multi-paradigm Applications
Joachim Jenke, Fabian Orland, Kingshuk Haldar, Thore Koritzius, Christian Terboven |
Euro-Par | 1 |
| 2022 | Towards a Hybrid MPI Correctness Benchmark SuiteabstractHigh-performance computing codes often combine the Message-Passing Interface (MPI) with a shared-memory programming model, e.g., OpenMP, for efficient computations. These so-called hybrid models may issue MPI calls concurrently from different threads at the highest level of MPI thread support. The correct use of either MPI or OpenMP can be complex and error-prone. The hybrid model increases this complexity even further. While correctness analysis tools exist for both programming paradigms, for hybrid models, a new set of potential errors exist, whose detection requires combining knowledge of MPI and OpenMP primitives. Unfortunately, correctness tools do not fully support the hybrid model yet, and their current capabilities are also hard to assess. In previous work, to enable structured comparisons of correctness tools and improve their coverage, we proposed the MPI-CorrBench test suite for MPI. Likewise, others proposed the DataRaceBench test suite for OpenMP. However, for the particular error classes of the hybrid model, no such test suite exists. Hence, we propose a hybrid MPI-OpenMP test suite to (1) facilitate the correctness tool development in this area and, subsequently, (2) further encourage the use of the hybrid model at the highest level of MPI thread support. To that end, we discuss issues with this hybrid model and the knowledge correctness tools need to combine w.r.t. MPI and OpenMP to detect these. In our evaluation of two state-of-the-art correctness tools, we see that for most cases of concurrent and conflicting MPI operations, these tools can cope with the added complexity of OpenMP. However, more intricate errors, where user code interferes with MPI, e.g., a data race on a buffer, still evade tool analysis. Tim Jammer, Alexander Hück, Jan-Patrick Lehr, Joachim Jenke, Simon Schwitanski, Christian H. Bischof |
EuroMPI | 4 |
| 2022 | MPI detach - Towards automatic asynchronous local completion
Joachim Jenke, Marc-André Hermanns, Matthias S. Müller, Van Man Nguyen, Julien Jaeger, Emmanuelle Saillard, Patrick Carribault, Denis Barthou |
Parallel Comput. | 1 |
| 2021 | ARBALEST: Dynamic Detection of Data Mapping Issues in Heterogeneous OpenMP ApplicationsabstractFrom OpenMP 4.0 onwards, programmers can offload code regions to accelerators by using the target offloading feature. However, incorrect usage of target offloading constructs may incur data mapping issues. A data mapping issue occurs when the host fails to observe updates on the accelerator or vice versa. It may further lead to multiple memory issues such as use of uninitialized memory, use of stale data, and data race. To the best of our knowledge, currently there is no prior work on dynamic detection of data mapping issues in heterogeneous OpenMP applications.In this paper, we identify possible root causes of data mapping issues in OpenMP's standard memory model and the unified memory model. We find that data mapping issues primarily result from incorrect settings of map and nowait clauses in target offloading constructs. Further, the novel unified memory model introduced in OpenMP 5.0 cannot avoid the occurrence of data mapping issues. To mitigate the difficulty of detecting data mapping issues, we propose ARBALEST, an on-the-fly data mapping issue detector for OpenMP applications. For each variable mapped to the accelerator, ARBALEST's detection algorithm leverages a state machine to track the last write's visibility. ARBALEST requires constant storage space for each memory location and takes amortized constant time per memory access. To demonstrate ARBALEST's effectiveness, an experimental comparison with four other dynamic analysis tools (Valgrind, Archer, AddressSanitizer, MemorySanitizer) has been carried out on a number of open-source benchmark suites. The evaluation results show that ARBALEST delivers demonstrably better precision than the other four tools, and its execution time overhead is comparable to that of state-of-the-art dynamic analysis tools. Lechen Yu, Joachim Jenke, Oscar R. Hernandez, Vivek Sarkar |
IPDPS | 2 |
| 2020 | MPI Detach - Asynchronous Local CompletionabstractWhen aiming for large scale parallel computing, waiting time due to network latency, synchronization, and load imbalance are the primary opponents of high parallel efficiency. A common approach to hide latency with computation is the use of non-blocking communication. In the presence of a consistent load imbalance, synchronization cost is just the visible symptom of the load imbalance. Tasking approaches as in OpenMP, TBB, OmpSs, or C++20 coroutines promise to expose a higher degree of concurrency, which can be distributed on available execution units and significantly increase load balance. Available MPI non-blocking functionality does not integrate seamlessly into such tasking parallelization. In this work, we present a slim extension of the MPI interface to allow seamless integration of non-blocking communication with available concepts of asynchronous execution in OpenMP and C++. Joachim Jenke, Marc-André Hermanns, Ali C. Demiralp, Matthias S. Müller, Torsten W. Kuhlen |
EuroMPI | 1 |
| 2018 | Thread-local concurrency: a technique to handle data race detection at programming model abstractionabstractWith greater adoption of various high-level parallel programming models to harness on-node parallelism, accurate data race detection has become more crucial than ever. However, existing tools have great difficulty spotting data races through these high-level models, as they primarily target low-level concurrent execution models (e.g., concurrency expressed at the level of POSIX threads). In this paper, we propose a novel technique to accurately detect those data races that can occur at higher levels of concurrent execution. The core idea of our technique is to introduce the general concept of Thread-Local Concurrency (TLC) as a new way to translate the concurrency expressed by a high-level programming paradigm into the low execution level understood by the existing tools. Specifically, we extend the definition of vector clocks to allow the existing state-of-the-art race detectors to recognize those races that occur at the higher level of concurrency with minor modifications to these tools. Our evaluation with our prototype implemented within ThreadSanitizer shows that TLC can allow the existing tool to detect these races accurately with only small additional analysis overheads. Joachim Jenke, Martin Schulz 0001, Dong H. Ahn, Matthias S. Müller |
HPDC | 1 |
| 2016 | ARCHER: Effectively Spotting Data Races in Large OpenMP ApplicationsabstractOpenMP plays a growing role as a portable programming model to harness on-node parallelism, yet, existing data race checkers for OpenMP have high overheads and generate many false positives. In this paper, we propose the first OpenMP data race checker, ARCHER, that achieves high accuracy, low overheads on large applications, and portability. ARCHER incorporates scalable happens-before tracking, exploits structured parallelism via combined static and dynamic analysis, and modularly interfaces with OpenMP runtimes. ARCHER significantly outperforms TSan and Intel® Inspector XE, while providing the same or better precision. It has helped detect critical data races in the Hypre library that is central to many projects at Lawrence Livermore National Laboratory and elsewhere. Simone Atzeni, Ganesh Gopalakrishnan, Zvonimir Rakamaric, Dong H. Ahn, Ignacio Laguna, Martin Schulz 0001, Gregory L. Lee, Joachim Jenke, Matthias S. Müller |
IPDPS | 8 |
| 2016 | Runtime Correctness Analysis of MPI-3 Nonblocking CollectivesabstractThe Message Passing Interface (MPI) includes nonblocking collective operations that support additional overlap between computation and communication. These new operations enable complex data movement between large numbers of processes. However, their asynchronous behavior hides and complicates the detection of defects in their use. We highlight a lack of correctness tool support for these operations and extend the MUST runtime MPI correctness tool to alleviate this complexity. We introduce a classification to summarize the types of correctness analyses that are applicable to MPI's nonblocking collectives. We identify complex wait-for dependencies in deadlock situations and incorrect use of communication buffers as the most challenging types of usage errors. We devise, demonstrate, and evaluate the applicability of correctness analyses for these errors. A scalable analysis mechanism allows our runtime approach to scale with the application. Benchmark measurements highlight the scalability and applicability of our approach at up to 4,096 application processes and with low overhead. Tobias Hilbrich, Matthias Weber 0002, Joachim Jenke, Bronis R. de Supinski, Wolfgang E. Nagel |
EuroMPI | 3 |
| 2015 | Event-Action Mappings for Parallel Tools Infrastructures
Tobias Hilbrich, Martin Schulz 0001, Holger Brunst, Joachim Jenke, Bronis R. de Supinski, Matthias S. Müller |
Euro-Par | 4 |
| 2013 | Intralayer Communication for Tree-Based Overlay NetworksabstractWhile various HPC tools use Tree-Based Overlay Networks (TBONs) to increase their scalability, some use cases do not map well to a tree-based hierarchy. We provide the concept of intralayer communication to improve this situation, where nodes in a specific hierarchy layer may exchange messages directly with each other. This concept targets data preprocessing that allows tool developers to avoid load imbalances in higher hierarchy levels. We implement intralayer communication within the Generic Tools Infrastructure (GTI) that provides TBON services, as well as a high-level abstraction to ease the creation of scalable runtime tools. An extension of GTI's abstractions allows simple and efficient use of intralayer communication. We demonstrate this capability with a runtime message matching tool for MPI's point-to-point communication, which we evaluate in an application study with up to 16,384 processes. Low overheads for two benchmark suites show the applicability of our approach, while a stress test demonstrates close to constant overheads across scales. The stress test measurements demonstrate that intralayer communication reduces application slowdown by two orders of magnitude at 2,048 processes, compared to a previous TBON-based implementation. Tobias Hilbrich, Joachim Jenke, Bronis R. de Supinski, Martin Schulz 0001, Matthias S. Müller, Wolfgang E. Nagel |
ICPP | 2 |
| 2013 | Distributed wait state tracking for runtime MPI deadlock detectionabstractThe widely used Message Passing Interface (MPI) with its multitude of communication functions is prone to usage errors. Runtime error detection tools aid in the removal of these errors. We develop MUST as one such tool that provides a wide variety of automatic correctness checks. Its correctness checks can be run in a distributed mode, except for its deadlock detection. This limitation applies to a wide range of tools that either use centralized detection algorithms or a timeout approach. In order to provide scalable and distributed deadlock detection with detailed insight into deadlock situations, we propose a model for MPI blocking conditions that we use to formulate a distributed algorithm. This algorithm implements scalable MPI deadlock detection in MUST. Stress tests at up to 4,096 processes demonstrate the scalability of our approach. Finally, overhead results for a complex benchmark suite demonstrate an average runtime increase of 34% at 2,048 processes. Tobias Hilbrich, Bronis R. de Supinski, Wolfgang E. Nagel, Joachim Jenke, Christel Baier, Matthias S. Müller |
SC | 4 |
| 2012 | Holistic Debugging of MPI Derived DatatypesabstractThe Message Passing Interface (MPI) specifies an API that allows programmers to create efficient and scalable parallel applications. The standard defines multiple constraints for each function parameter. For performance reasons, no MPI implementation checks all of these constraints at runtime. Derived data types are an important concept of MPI and allow users to describe an application's data structures for efficient and convenient communication. Using existing infrastructure we present scalable algorithms to detect usage errors of basic and derived MPI data types. We detect errors that include constraints for construction and usage of derived data types, matching their type signatures in communication, and detecting erroneous overlaps of communication buffers. We implement these checks in the MUST runtime error detection framework. We provide a novel representation of error locations to highlight usage errors. Further, approaches to buffer overlap checking can cause unacceptable overheads for non-contiguous data types. We present an algorithm that uses patterns in derived MPI data types to avoid these overheads without losing precision. Application results for the benchmark suites SPEC MPI2007 and NAS Parallel Benchmarks for up to 2048 cores show that our approach applies to a broad range of applications and that our extended overlap check improves performance by two orders of magnitude. Finally, we augment our runtime error detection component with a debugger extension to support in-depth analysis of the errors that we find as well as semantic errors. This extension to gdb provides information about MPI data type handles and enables gdb -- and other debuggers based on gdb -- to display the content of a buffer as used in MPI communications. Joachim Jenke, Tobias Hilbrich, Andreas Knüpfer, Bronis R. de Supinski, Matthias S. Müller |
IPDPS | 1 |
| 2012 | MPI runtime error detection with MUST: advances in deadlock detectionabstractThe widely used Message Passing Interface (MPI) is complex and rich. As a result, application developers require automated tools to avoid and to detect MPI programming errors. We present the Marmot Umpire Scalable Tool (MUST) that detects such errors with significantly increased scalability. We present improvements to our graph-based deadlock detection approach for MPI, which cover future MPI extensions. Our enhancements also check complex MPI constructs that no previous graph-based detection approach handled correctly. Finally, we present optimizations for the processing of MPI operations that reduce runtime deadlock detection overheads. Existing approaches often require O(p) analysis time per MPI operation, for p processes. We empirically observe that our improvements lead to sub-linear or better analysis time per operation for a wide range of real world applications. Tobias Hilbrich, Joachim Jenke, Martin Schulz 0001, Bronis R. de Supinski, Matthias S. Müller |
SC | 2 |