Tobias Weinzierl

dblp:68/6473 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
3since 2021 · last 2026
0000-0002-6208-1841ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 1 first-author · 2 since 2021Theory of computation · 4 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 An Extension of C++ with Memory-Centric Specifications for HPC to Reduce Memory Footprints and Streamline MPI Development
abstract
C++ leans towards a memory-inefficient storage of structs: The compiler inserts padding bits, while it is not able to exploit knowledge about the range of integers, enums or bitsets. Furthermore, the language provides no support for arbitrary floating-point precisions. We propose a language extension based upon attributes through which developers can guide the compiler what memory arrangements would be beneficial: Can multiple booleans or integers with limited range be squeezed into one bit field, do floating-point numbers hold fewer significant bits than in the IEEE standard, and is a programmer willing to trade attribute ordering guarantees for a more compact object representation? The extension offers the opportunity to fall back to normal alignment and native C++ floating-point representations via plain C++ assignments, no dependencies upon external libraries are introduced, and the resulting code remains (syntactically) standard C++. As MPI remains the de facto standard for distributed memory calculations in C++, we furthermore propose additional attributes which streamline the MPI datatype modelling in combination with our memory optimisation extensions. Our work implements the language annotations within LLVM and demonstrates their potential impact through smoothed particle hydrodynamics benchmarks. They uncover the potential gains in terms of performance and development productivity.
Pawel K. Radtke, Cristian G. Barrera-Hinojosa, Mladen Ivkovic, Tobias Weinzierl
ACM Trans. Math. Softw.4
2025 Annotation-Guided AoS-to-SoA Conversions and GPU Offloading With Data Views in C++
abstract
ABSTRACT The C++ programming language provides classes and structs as fundamental modeling entities. Consequently, C++ code tends to favor array‐of‐structs (AoS) for encoding data sequences, even though structure‐of‐arrays (SoA) yields better performance for some calculations. We propose a C++ language extension based on attributes that allows developers to guide the compiler in selecting memory arrangements, that is, to select the optimal choice between AoS and SoA dynamically depending on both the execution context and algorithm step. The compiler can then automatically convert data into the preferred format prior to the calculations and convert results back afterward. The compiler handles all the complexity of determining which data to convert and how to manage data transformations. Our implementation realizes the compiler‐extension for the new annotations in Clang and demonstrates their effectiveness through a smoothed particle hydrodynamics (SPH) code, which we evaluate on an Intel CPU, an ARM CPU, and a Grace‐Hopper GPU. While the separation of concerns between data structure and operators is elegant and provides performance improvements, the new annotations do not eliminate the need for performance engineering. Instead, they challenge conventional performance wisdom and necessitate rethinking approaches how to write efficient implementations.
Pawel K. Radtke, Tobias Weinzierl
Concurr. Comput. Pract. Exp.2
2021 Delayed approximate matrix assembly in multigrid with dynamic precisions
abstract
Summary The accurate assembly of the system matrix is an important step in any code that solves partial differential equations on a mesh. We either explicitly set up a matrix, or we work in a matrix‐free environment where we have to be able to quickly return matrix entries upon demand. Either way, the construction can become costly due to nontrivial material parameters entering the equations, multigrid codes requiring cascades of matrices that depend upon each other, or dynamic adaptive mesh refinement that necessitates the recomputation of matrix entries or the whole equation system throughout the solve. We propose that these constructions can be performed concurrently with the multigrid cycles. Initial geometric matrices and low accuracy integrations kickstart the multigrid iterations, while improved assembly data is fed to the solver as and when it becomes available. The time to solution is improved as we eliminate an expensive preparation phase traditionally delaying the actual computation. We eliminate algorithmic latency. Furthermore, we desynchronize the assembly from the solution process. This anarchic increase in the concurrency level improves the scalability. Assembly routines are notoriously memory‐ and bandwidth‐demanding. As we work with iteratively improving operator accuracies, we finally propose the use of a hierarchical, lossy compression scheme such that the memory footprint is brought down aggressively where the system matrix entries carry little information or are not yet available with high accuracy.
Charles D. Murray, Tobias Weinzierl
Concurr. Comput. Pract. Exp.2
2020 Lightweight task offloading exploiting MPI wait times for parallel adaptive mesh refinement
abstract
Summary Balancing the workload of sophisticated simulations is inherently difficult, since we have to balance both computational workload and memory footprint over meshes that can change any time or yield unpredictable cost per mesh entity, while modern supercomputers and their interconnects start to exhibit fluctuating performance. We propose a novel lightweight balancing technique for MPI+X to accompany traditional, prediction‐based load balancing. It is a reactive diffusion approach that uses online measurements of MPI idle time to migrate tasks temporarily from overloaded to underemployed ranks. Tasks are deployed to ranks which otherwise would wait, processed with high priority, and made available to the overloaded ranks again. This migration is nonpersistent. Our approach hijacks idle time to do meaningful work and is totally nonblocking, asynchronous and distributed without a global data view. Tests with a seismic simulation code developed in the ExaHyPE engine uncover the method's potential. We found speed‐ups of up to 2‐3 for ill‐balanced scenarios without logical modifications of the code base and show that the strategy is capable to react quickly to temporarily changing workload or node performance.
Philipp Samfass, Tobias Weinzierl, Dominic Etienne Charrier, Michael Bader
Concurr. Comput. Pract. Exp.2
2019 A multi-core ready discrete element method with triangles using dynamically adaptive multiscale grids
abstract
The simulation of vast numbers of rigid bodies of non-analytical shapes and of tremendously different sizes that collide with each other is computationally challenging. A bottleneck is the identification of all particle contact points per time step. We propose a tree-based multilevel meta data structure to administer the particles. The data structure plus a purpose-made tree traversal identifying the contact points introduce concurrency to the particle comparisons, whilst they keep the absolute number of particle-to-particle comparisons low. Furthermore, a novel adaptivity criterion allows explicit time stepping to work with comparably large time steps. It optimises both toward low algorithmic complexity per time step and low numbers of time steps. We study three different parallelisation strategies exploiting our traversal's concurrency. The fusion of two of them yields promising speedups once we rely on maximally asynchronous task-based realisations. Our work shows that new computer architecture can push the boundary of rigid particle computability, yet if and only if the right data structures and data processing schemes are chosen.
Konstantinos Krestenitis, Tobias Weinzierl
Concurr. Comput. Pract. Exp.2
2019 The Peano Software - Parallel, Automaton-based, Dynamically Adaptive Grid Traversals
abstract
We discuss the design decisions, design alternatives, and rationale behind the third generation of Peano, a framework for dynamically adaptive Cartesian meshes derived from spacetrees. Peano ties the mesh traversal to the mesh storage and supports only one element-wise traversal order resulting from space-filling curves. The user is not free to choose a traversal order herself. The traversal can exploit regular grid subregions and shared memory as well as distributed memory systems with almost no modifications to a serial application code. We formalize the software design by means of two interacting automata—one automaton for the multiscale grid traversal and one for the application-specific algorithmic steps. This yields a callback-based programming paradigm. We further sketch the supported application types and the two data storage schemes realized before we detail high-performance computing aspects and lessons learned. Special emphasis is put on observations regarding the used programming idioms and algorithmic concepts. This transforms our report from a “one way to implement things” code description into a generic discussion and summary of some alternatives, rationale, and design decisions to be made for any tree-based adaptive mesh refinement software.
Tobias Weinzierl
ACM Trans. Math. Softw.1
2018 Quasi-matrix-free Hybrid Multigrid on Dynamically Adaptive Cartesian Grids
abstract
We present a family of spacetree-based multigrid realizations using the tree’s multiscale nature to derive coarse grids. They align with matrix-free geometric multigrid solvers as they never assemble the system matrices, which is cumbersome for dynamically adaptive grids and full multigrid. The most sophisticated realizations use BoxMG to construct operator-dependent prolongation and restriction in combination with Galerkin/Petrov-Galerkin coarse-grid operators. This yields robust solvers for nontrivial elliptic problems. We embed the algebraic, problem-dependent, and grid-dependent multigrid operators as stencils into the grid and evaluate all matrix-vector productsin situthroughout the grid traversals. Such an approach is not literally matrix-free as the grid carries the matrix. We propose to switch to a hierarchical representation of all operators. Only differences of algebraic operators to their geometric counterparts are held. These hierarchical differences can be stored and exchanged with small memory footprint. Our realizations support arbitrary dynamically adaptive grids while they vertically integrate the multilevel operations through spacetree linearization. This yields good memory access characteristics, while standard colouring of mesh entities with domain decomposition allows us to use parallel many-core clusters. All realization ingredients are detailed such that they can be used by other codes.
Marion Weinzierl, Tobias Weinzierl
ACM Trans. Math. Softw.2
2017 Complex Additive Geometric Multilevel Solvers for Helmholtz Equations on Spacetrees
abstract
We introduce a family of implementations of low-order, additive, geometric multilevel solvers for systems of Helmholtz equations arising from Schrödinger equations. Both grid spacing and arithmetics may comprise complex numbers, and we thus can apply complex scaling to the indefinite Helmholtz operator. Our implementations are based on the notion of a spacetree and work exclusively with a finite number of precomputed local element matrices. They are globally matrix-free. Combining various relaxation factors with two grid transfer operators allows us to switch from additive multigrid over a hierarchical basis method into a Bramble-Pasciak-Xu (BPX)-type solver, with several multiscale smoothing variants within one code base. Pipelining allows us to realize full approximation storage (FAS) within the additive environment where, amortized, each grid vertex carrying degrees of freedom is read/written only once per iteration. The codes realize a single-touch policy. Among the features facilitated by matrix-free FAS is arbitrary dynamic mesh refinement (AMR) for all solver variants. AMR as an enabler for full multigrid (FMG) cycling—the grid unfolds throughout the computation—allows us to reduce the cost per unknown. The present work primary contributes toward software realization and design questions. Our experiments show that the consolidation of single-touch FAS, dynamic AMR, and vectorization-friendly, complex scaled, matrix-free FMG cycles delivers a mature implementation blueprint for solvers of Helmholtz equations in general. For this blueprint, we put particular emphasis on a strict implementation formalism as well as some implementation correctness proofs.
Bram Reps, Tobias Weinzierl
ACM Trans. Math. Softw.2
2016 Two particle-in-grid realisations on spacetrees
Tobias Weinzierl, B. Verleye, Pierre Henri, Dirk Roose
Parallel Comput.1
2013 Cluster Optimization and Parallelization of Simulations with Dynamically Adaptive Grids
Martin Schreiber 0001, Tobias Weinzierl, Hans-Joachim Bungartz
Euro-Par2
2010 A precompiler to reduce the memory footprint of multiscale PDE solvers in C++
Hans-Joachim Bungartz, Wolfgang Eckhardt, Tobias Weinzierl, Christoph Zenger 0001
Future Gener. Comput. Syst.3
2010 An efficient parallel implementation of the MSPAI preconditioner
Thomas Huckle, Alexander Kallischko, Andreas Roy, Matous Sedlacek, Tobias Weinzierl
Parallel Comput.5
2006 A Parallel Adaptive Cartesian PDE Solver Using Space-Filling Curves
Hans-Joachim Bungartz, Miriam Mehl, Tobias Weinzierl
Euro-Par3