Josef Weidendorfer

dblp:10/1152 · DBLP profile ↗
← Back
18ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0001-7159-1432ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 6 since 2021Software engineering, systems software and programming languages · 2Security and privacy · 1
YearPublicationVenuePosition
2026 Simulating MPI Collectives on Tofino Smart Switches in SimGrid
abstract
Programmable smart switches enable In-Network Computing, e.g., to accelerate HPC workloads by offloading collective operations from host CPUs. However, evaluating the benefits of these devices remains challenging due to the cost and complexity of deployment on real hardware. In this paper, we address this limitation by simulating Intel Tofino smart switches using the SimGrid framework. We take advantage of the components of the S4U and SMPI modules and introduce a new network component that represents smart switches that reproduce the latency and computational capabilities of the Tofino architecture. We validate this model on a physical testbed and present a performance evaluation of MPI collective operations offloading. Although we focus on simulating Tofino-class switches, our approach can be adapted to other smart switch architectures. Our preliminary results indicate that small-scale simulations achieve latency comparable to the real hardware when offloading MPI_Allreduce. This study lays the groundwork for future assessments of Tofino smart switches at scale.
Ahmad Moh'd Saleh A. Belbeisi, Majid Salimi Beni, Thomas Erbesdobler, Ehab Saleh, Matthew Tovey, Amir Raoofy, Josef Weidendorfer
CF7
2026 A Dedicated CPU Core for MPI Progress: Towards Improved Overlap in Non-Blocking Two-Sided Communication
Ehab Saleh, Amir Raoofy, Robert Mijakovic, Ahmad Moh'd Saleh A. Belbeisi, Josef Weidendorfer
ISPDC5
2025 POSTER: Performance Comparison of GPU Programming Models Using HeCBench Benchmarks
abstract
GPUs play an important role in High-Performance Computing.The choice of GPU programming models plays a crucial role in achieving portability and performance.High-level programming models, such as SYCL and OpenMP offloading, have emerged, offering unified abstractions that enable developers to target multiple architectures with a single, maintainable codebase.However, achieving consistent performance across different models remains a significant challenge due to variations in abstraction levels, compiler optimizations, and runtime behavior.We present a profiling-based methodology for systematically comparing GPU programming models on NVIDIA and AMD GPUs.We apply our methodology to over 150 benchmarks of HeCBench, demonstrating its effectiveness in identifying performance issues in OpenMP, SYCL, HIP and CUDA implementations for AMD and NVIDIA GPUs. CCS Concepts• General and reference
Jakob Schäffeler, Bengisu Elis, Amir Raoofy, Josef Weidendorfer, Martin Schulz 0001
CF4
2024 A Portable Tool to Compare Performance Profiles from GPU Offloading Programming Models
abstract
GPUs are growingly dominating the High-Performance Computing ecosystem, and therefore, the ease of their programming is getting increasingly important. Standard and high-level offloading methods, like OpenMP offloading and OpenACC, facilitate portable and efficient offloading across different GPU platforms. However, pinpointing and troubleshooting performance variations among different models, implementations, or architectures poses a challenge due to varying abstraction levels and profilers employed. Therefore, to tackle this problem and to unwind the performance issues related to various offloading abstractions and models that are entangled together in practice, in this work, we introduce a portable tool to enable the comparison of performance profiles acquired from various offloading models and GPU platforms. For this, the tool first processes the collected profiles by different profilers to extract key performance indicatory metrics. For ease of comparison, the tool utilizes plots depicting the metrics of all target variants for relative comparison. Moreover, we demonstrate the tool's capabilities by discussing specific issues discovered by using the tool when comparing OpenMP offloading and CUDA implementations of Babelstream.
Jakob Schäffeler, Bengisu Elis, Amir Raoofy, Josef Weidendorfer, Martin Schulz 0001
CF4
2023 Phase-aware System-Side Sampling for HPC
abstract
HPC compute centers always benefit from better insights into the application mix running on their systems. We present a low-overhead statistical sampling tool running in the background on the system side, which can capture application compute phases. Our tool leverages eBPF (extended Berkeley Packet Filter) from modern Linux kernels to extract phase information provided by developers via instrumentation or instruction pointers. We outline how this tool can be integrated into the monitoring framework DCDB, and we show resulting performance insights.
Julian Scheipl, Amir Raoofy, Michael Ott 0001, Josef Weidendorfer
CF4
2023 From reactive to proactive load balancing for task-based parallel applications in distributed memory machines
abstract
Summary Load balancing is often a challenge in task‐parallel applications. The balancing problems are divided into static and dynamic. “Static” means that we have some prior knowledge about load information and perform balancing before execution, while “dynamic” must rely on partial information of the execution status to balance the load at runtime. Conventionally, work stealing is a practical approach used in almost all shared memory systems. In distributed memory systems, the communication overhead can make stealing tasks too late. To improve, people have proposed a reactive approach to relax communication in balancing load. The approach leaves one dedicated thread per process to monitor the queue status and offload tasks reactively from a slow to a fast process. However, reactive decisions might be mistaken in high imbalance cases. First, this article proposes a performance model to analyze reactive balancing behaviors and understand the bound leading to incorrect decisions. Second, we introduce a proactive approach to improve further balancing tasks at runtime. The approach exploits task‐based programming models with a dedicated thread as well, namely . Nevertheless, the main idea is to force not only to monitor load; it will characterize tasks and train load prediction models by online learning. “Proactive” indicates offloading tasks before each execution phase proactively with an appropriate number of tasks at once to a potential victim (denoted by an underloaded/fast process). The experimental results confirm speedup improvements from to in important use cases compared to the previous solutions. Furthermore, this approach can support co‐scheduling tasks across multiple applications.
Minh Thanh Chung, Josef Weidendorfer, Karl Fürlinger, Dieter Kranzlmüller
Concurr. Comput. Pract. Exp.2
2022 Always-on instrumentation for application introspection in HPC
abstract
Obtaining insights into the dynamic behavior of user code is crucial for supercomputing centers to support both better operation and co-design of future systems. To this end, always-on instrumentation is the key: enabling all running code to dynamically forward metadata such as compute phase changes to the system would provide important information for those goals. To keep the overhead low, the system must be able to deactivate instrumentation points with high trigger frequency on demand. In this poster, we present a simple always-on instrumentation method for C/C++ which can be easily used by developers, copying a single source file into their code base. Our evaluations show that the overhead in the deactivated state is low enough for the manual instrumentation to stay compiled in, all the time.
Amir Raoofy, Josef Weidendorfer, Michael Ott 0001
CF2
2022 A Profiling-Based Approach to Cache Partitioning of Program Data
Sergej Breiter, Josef Weidendorfer, Minh Thanh Chung, Karl Fürlinger
PDCAT2
2017 Dynamic Co-Scheduling Driven by Main Memory Bandwidth Utilization
abstract
Most applications running on supercomputers achieve only a fraction of a system's peak performance. It has been demonstrated that the co-scheduling of applications can improve the overall system utilization. However, following this approach, applications need to fulfill certain criteria such that the mutual slowdown is kept at a minimum. In this paper, we present an HPC scheduler that applies co-scheduling and utilizes virtual machine migration for a re-orchestration of applications at runtime based on their main memory bandwidth requirements. Given a job queue consisting of main memory-bound applications and compute-bound applications, we can see a throughput increase of up to 35% while at the same time reducing energy consumption by around 30%.
Jens Breitbart, Simon Pickartz, Stefan Lankes, Josef Weidendorfer, Antonello Monti
CLUSTER4
2016 Automatic Co-scheduling Based on Main Memory Bandwidth Usage
Jens Breitbart, Josef Weidendorfer, Carsten Trinitis
JSSPP2
2014 A Novel Variable Ordering Heuristic for BDD-based K-Terminal Reliability
abstract
Modern exact methods solving the NP-hard k-terminal reliability problem are based on Binary Decision Diagrams (BDDs). The system redundancy structure represented by the input graph is converted into a BDD whose size highly depends on the predetermined variable ordering. As finding the optimal available ordering has exponential complexity, a heuristic must be used. Currently, the breadth-first-search is considered to be state-of-the-art. Based on Hardy's decomposition approach, we derive a novel static heuristic which yields significantly smaller BDD sizes for a wide variety of network structures, especially irregular ones. As a result, runtime and memory requirements can be drastically reduced for BDD-based reliability methods. Applying the decomposition method with the new heuristic to three medium-sized irregular networks from the literature, an average speedup of around 9,400 is gained and the memory consumption drops to less than 0.1 percent.
Minh Lê, Josef Weidendorfer, Max Walter
DSN2
2013 Real Asynchronous MPI Communication in Hybrid Codes through OpenMP Communication Tasks
abstract
With the number of cores growing faster than memory per node, hybrid programming models (mixing message passing with shared memory paradigms) become a requirement for efficient use of HPC systems. For this scenario, achieving efficient communication is challenging. This is true even when using asynchronous communication, as most MPI implementations can only advance communication inside library calls. In this paper we propose to move communication into a new type of OpenMP task, which gets scheduled as part of the regular OpenMP work-pool. We show for compute intensive iterative stencil algorithms, that this provides real asynchronous communication. Without complicating the programming interface, our results show an excellent performance independent of the communication to computation ratio.
David Büttner, Jean-Thomas Acquaviva, Josef Weidendorfer
ICPADS3
2013 Expression Tree Evaluation by Dynamic Code Generation - Are Accelerators Up for the Task?
abstract
Dynamic code generation techniques are useful if the benefit of code specialized to values only known at runtime outweighs generation time. Such techniques are increasingly employed for HPC applications to tune their runtime behavior. The simulation software investigated in this paper is a typical example: It spends a significant portion of computing time evaluating symbolic formulas which are set up dynamically from model data. However, any software tuning has to match the hardware. Due to the so-called power wall, HPC systems are increasingly equipped with throughput-oriented accelerator components, to allow for rising performance as known from Top500 history. To best exploit such systems, it is important to understand how well applications map to heterogeneous components. While dynamic code generation can work well for standard multi-core systems, in this paper, we research the benefit of accelerators for this scenario. For our application we show that - while the generated code runs well on the accelerator - the generation itself has serious issues, and much better maps to standard multi-cores. Therefore, we see the need that coming HPC systems still have to be equipped with a significant portion of latency-oriented, thus complex general-purpose hardware.
Thomas Müller 0001, Josef Weidendorfer, Andreas Blaszczyk
ICPP2
2012 An integrated simulation framework for invasive computing
Michael Gerndt, Frank Hannig, Andreas Herkersdorf, Andreas Hollmann, Marcel Meyer, Sascha Roloff, Josef Weidendorfer, Thomas Wild, Aurang Zaib
FDL7
2012 Invasive computing with iOMP
Michael Gerndt, Andreas Hollmann, Marcel Meyer, Martin Schreiber 0001, Josef Weidendorfer
FDL5
2011 Compact data structure and scalable algorithms for the sparse grid technique
abstract
The sparse grid discretization technique enables a compressed representation of higher-dimensional functions. In its original form, it relies heavily on recursion and complex data structures, thus being far from well-suited for GPUs. In this paper, we describe optimizations that enable us to implement compression and decompression, the crucial sparse grid algorithms for our application, on Nvidia GPUs. The main idea consists of a bijective mapping between the set of points in a multi-dimensional sparse grid and a set of consecutive natural numbers. The resulting data structure consumes a minimum amount of memory. For a 10-dimensional sparse grid with approximately 127 million points, it consumes up to 30 times less memory than trees or hash tables which are typically used. Compared to a sequential CPU implementation, the speedups achieved on GPU are up to 17 for compression and up to 70 for decompression, respectively. We show that the optimizations are also applicable to multicore CPUs.
Alin Florindor Murarasu, Josef Weidendorfer, Gerrit Buse, Daniel Butnaru, Dirk Pflüger
PPoPP2
2011 Sparse matrix operations on several multi-core architectures
Carsten Trinitis, Tilman Küstner, Josef Weidendorfer, Jasmin Smajic
J. Supercomput.3
2004 A Data Structure Oriented Monitoring Environment for Fortran OpenMP Programs
Edmond Kereku, Tianchao Li 0001, Michael Gerndt, Josef Weidendorfer
Euro-Par4