Harald Köstler

dblp:41/5505 · DBLP profile ↗
← Back
22ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0002-6992-2690ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 6 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Software engineering, systems software and programming languages · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 The European master for HPC curriculum
abstract
International audience
Pascal Bouvry, Mats Brorsson, Ramon Canal, Aryan Eftekhari, Siegfried Höfinger, Didier Smets, Harald Köstler, Tomás Kozubek, Ezhilmathi Krishnasamy, Josep Llosa, Alexandra Lukas-Rother, Xavier Martorell, Dirk Pleiter, Ana Proykova, Maria-Ribera Sancho, Olaf Schenk, Cristina Silvano
J. Parallel Distributed Comput.7
2025 Analyzing performance portability for a SYCL implementation of the 2D shallow water equations
abstract
Abstract SYCL is an open standard for targeting heterogeneous hardware from C++. In this work, we evaluate a SYCL implementation for a discontinuous Galerkin discretization of the 2D shallow water equations targeting CPUs, GPUs, and also FPGAs. The discretization uses polynomial orders zero to two on unstructured triangular meshes. Separating memory accesses from the numerical code allow us to optimize data accesses for the target architecture. A performance analysis shows good portability across x86 and ARM CPUs, GPUs from different vendors, and even two variants of Intel Stratix 10 FPGAs. Measuring the energy to solution shows that GPUs yield an up to 10x higher energy efficiency in terms of degrees of freedom per joule compared to CPUs. With custom designed caches, FPGAs offer a meaningful complement to the other architectures with particularly good computational performance on smaller meshes. FPGAs with High Bandwidth Memory are less affected by bandwidth issues and have similar energy efficiency as latest generation CPUs.
Markus Büttner, Christoph Alt, Tobias Kenter, Harald Köstler, Christian Plessl, Vadym Aizinger
J. Supercomput.4
2024 Code Generation for Octree-Based Multigrid Solvers with Fused Higher-Order Interpolation and Communication
Richard Angersbach, Sebastian Kuckuk, Harald Köstler
Euro-Par (3)3
2024 waLBerla-wind: A lattice-Boltzmann-based high-performance flow solver for wind energy applications
abstract
Summary This article presents the development of a new wind turbine simulation software to study wake flow physics. To this end, the design and development of waLBerla‐wind , a new simulator based on the lattice‐Boltzmann method that is known for its excellent performance and scaling properties, will be presented. Here it will be used for large eddy simulations (LES) coupled with actuator wind turbine models. Due to its modular software design, waLBerla‐wind is flexible and extensible with regard to turbine configurations. Additionally it is performance portable across different hardware architectures, another critical design goal. The new solver is validated by presenting force distributions and velocity profiles and comparing them with experimental data and a vortex solver. Furthermore, waLBerla‐wind 's performance is compared to a theoretical peak performance, and analyzed with weak and strong scaling benchmarks on CPU and GPU systems. This analysis demonstrates the suitability for large‐scale applications and future cost‐effective full wind farm simulations.
Helen Schottenhamml, Ani Anciaux-Sedrakian, Frédéric Blondel, Harald Köstler, Ulrich Rüde
Concurr. Comput. Pract. Exp.4
2023 MD-Bench: A performance-focused prototyping harness for state-of-the-art short-range molecular dynamics algorithms
Rafael Ravedutti L. Machado, Jan Eitzinger, Jan Laukemann, Georg Hager, Harald Köstler, Gerhard Wellein
Future Gener. Comput. Syst.5
2022 Evolving generalizable multigrid-based helmholtz preconditioners with grammar-guided genetic programming
abstract
Solving the indefinite Helmholtz equation is not only crucial for the understanding of many physical phenomena but also represents an outstandingly-difficult benchmark problem for the successful application of numerical methods. Here we introduce a new approach for evolving eficient preconditioned iterative solvers for Helmholtz problems with multi-objective grammar-guided genetic programming. Our approach is based on a novel context-free grammar, which enables the construction of multigrid preconditioners that employ a tailored sequence of operations on each discretization level. To find solvers that generalize well over the given domain, we propose a custom method of successive problem difficulty adaption, in which we evaluate a preconditioner's efficiency on increasingly ill-conditioned problem instances. We demonstrate our approach's effectiveness by evolving multigrid-based preconditioners for a two-dimensional indefinite Helmholtz problem that outperform several human-designed methods for different wavenumbers up to systems of linear equations with more than a million unknowns.
Jonas Schmitt, Harald Köstler
GECCO2
2021 On revisiting energy and performance in microservices applications: A cloud elasticity-driven approach
Igor Fontana De Nardin, Rodrigo da Rosa Righi, Thiago Roberto Lima Lopes, Cristiano André da Costa, Heon Young Yeom, Harald Köstler
Parallel Comput.6
2020 Constructing efficient multigrid solvers with genetic programming
abstract
For many linear and nonlinear systems that arise from the discretization of partial differential equations the construction of an efficient multigrid solver is a challenging task. Here we present a novel approach for the optimization of geometric multigrid methods that is based on evolutionary computation, a generic program optimization technique inspired by the principle of natural evolution. A multigrid solver is represented as a tree of mathematical expressions which we generate based on a tailored grammar. The quality of each solver is evaluated in terms of convergence and compute performance using automated local Fourier analysis (LFA) and roofline performance modeling, respectively. Based on these objectives a multi-objective optimization is performed using grammar-guided genetic programming with a non-dominated sorting based selection. To evaluate the model-based prediction and to target concrete applications, scalable implementations of an evolved solver can be automatically generated with the ExaStencils framework. We demonstrate the effectiveness of our approach by constructing multigrid solvers for a linear elastic boundary value problem that are competitive with common V- and W-cycles.
Jonas Schmitt, Sebastian Kuckuk, Harald Köstler
GECCO3
2019 Code generation for massively parallel phase-field simulations
abstract
This article describes the development of automatic program generation technology to create scalable phase-field methods for material science applications. To simulate the formation of microstructures in metal alloys, we employ an advanced, thermodynamically consistent phase-field method. A state-of-the-art large-scale implementation of this model requires extensive, time-consuming, manual code optimization to achieve unprecedented fine mesh resolution. Our new approach starts with an abstract description based on free-energy functionals which is formally transformed into a continuous PDE and discretized automatically to obtain a stencil-based time-stepping scheme. Subsequently, an automatized performance engineering process generates highly optimized, performance-portable code for CPUs and GPUs. We demonstrate the efficiency for real-world simulations on large-scale GPU-based (PizDaint) and CPU-based (SuperMUC-NG) supercomputers. Our technique simplifies program development and optimization for a wide class of models.
Martin Bauer 0003, Johannes Hötzer, Dominik Ernst, Julian Hammer, Marco Seiz, Henrik Hierl, Jan Hönig, Harald Köstler, Gerhard Wellein, Britta Nestler, Ulrich Rüde
SC8
2018 Whole Program Generation of Massively Parallel Shallow Water Equation Solvers
abstract
The study of ocean currents has been an active area of research for decades. As a model close to the water surface, the shallow water equations (SWE) can be used. For realistic simulations, efficient numerical solvers are necessary that exhibit a good node-level performance while still maintaining scalability. When comparing the discretized model and the actual implementation, one often finds that they differ vastly. This gap makes it hard for domain experts to implement their models and high performance computing (HPC) experts are required to ensure an optimal implementation. Using domain-specific languages (DSLs) and code generation techniques can be a useful tool to bridge this gap. In recent years, ExaStencils and its DSL ExaSlang have proven to provide a suitable platform for this. We present an extension from up to now elliptic to hyperbolic partial differential equations (PDEs) in this work, namely the SWE. After setting up a suitable discretization, we demonstrate how it can be mapped to ExaSlang code. This code is still quite similar to the original, mathematically motivated specification and can be easily written by domain experts. Still, solvers generated from this abstract representation can be run on large-scale clusters. We demonstrate this by giving performance and scalability results on the state-of-the-art GPU cluster Piz Daint where we solve for close to a trillion unknowns on 2048 GPUs. From there, we discuss the performance impact of different optimizations such as overlapping computation and communication, or switching to a hybrid CPU-GPU parallelization scheme.
Sebastian Kuckuk, Harald Köstler
CLUSTER2
2018 Unified Code Generation for the Parallel Computation of Pairwise Interactions Using Partial Evaluation
abstract
The evaluation of pairwise interactions is fundamental for the simulation of most molecular processes. Their efficient computation is therefore crucial for the overall performance of these simulations on modern computer architectures. We show how code for the computation of pairwise interactions on parallel and heterogeneous platforms can be generated from a unified base through partial evaluation of higher-order functions. For this purpose we introduce a complete implementation of the neighbor list algorithm based on the AnyDSL framework, from which we are able to generate executables for both CPU and GPU through compile-time specialization. Furthermore, we discuss the advantages and disadvantages of our approach and compare it with the miniMD simulation package from the Mantevo project, which is implemented in the C++ programming language and uses a similar computational core as the widely used molecular dynamics package LAMMPS. Finally, we assess the performance of our implementation in a number of test cases on modern CPU and GPU hardware.
Jonas Schmitt, Harald Köstler, Jan Eitzinger, Richard Membarth
ISPDC2
2018 Knowledge Amalgamation for Computational Science and Engineering
Theresa Pollinger, Michael Kohlhase, Harald Köstler
CICM3
2017 Towards generating efficient flow solvers with the ExaStencils approach
abstract
Summary ExaStencils aims at providing intuitive interfaces for the specification of numerical problems and resulting solvers, particularly those from the class of (geometric) multigrid methods. It envisions a multi‐layered domain‐specific language and a sophisticated code generation framework ultimately emitting source code in a chosen target language. We present our recent advances in fully generating solvers applied to 3D fluid mechanics for nonisothermal/non‐Newtonian flows. In detail, a system of time‐dependent, nonlinear partial differential equations is discretized on a cubic, nonuniform, and staggered grid using finite volumes. We examine the contained problem of coupled Navier‐Stokes and temperature equations, which are linearized and solved using the SIMPLE algorithm and geometric multigrid solvers, as well as the incorporation of non‐Newtonian properties. Furthermore, we provide details on necessary extensions to our domain‐specific language and code generation framework, in particular, those concerning the handling of boundary conditions, support for nonequidistant staggered grids, and supplying specialized functions to express operations reoccurring in the scope of finite volume discretizations. Many of these enhancements are generalizable and thus suitable for utilization in similar projects. Lastly, we demonstrate the applicability of our code generation approach by providing convincing performance results for fully generated and automatically parallelized solvers.
Sebastian Kuckuk, Gundolf Haase, Diego A. Vasco, Harald Köstler
Concurr. Comput. Pract. Exp.4
2015 Massively parallel phase-field simulations for ternary eutectic directional solidification
abstract
Microstructures forming during ternary eutectic directional solidification processes have significant influence on the macroscopic mechanical properties of metal alloys. For a realistic simulation, we use the well established thermodynamically consistent phase-field method and improve it with a new grand potential formulation to couple the concentration evolution. This extension is very compute intensive due to a temperature dependent diffusive concentration. We significantly extend previous simulations that have used simpler phase-field models or were performed on smaller domain sizes. The new method has been implemented within the massively parallel HPC framework waLBerla that is designed to exploit current supercomputers efficiently. We apply various optimization techniques, including buffering techniques, explicit SIMD kernel vectorization, and communication hiding. Simulations utilizing up to 262,144 cores have been run on three different supercomputing architectures and weak scalability results are shown. Additionally, a hierarchical, mesh-based data reduction strategy is developed to keep the I/O problem manageable at scale.
Martin Bauer 0003, Johannes Hötzer, Marcus Jainta, Philipp Steinmetz, Marco Berghoff, Florian Schornbaum, Christian Godenschwager, Harald Köstler, Britta Nestler, Ulrich Rüde
SC8
2015 Tsunami and Storm Surge Simulation Using Low Power Architectures - Concept and Evaluation
abstract
Performing a tsunami or storm surge simulation in real time is a highly challenging research topic that calls for a collaboration between mathematicians and computer scientists. One must combine mathematical models with numerical methods and rely on computational performance and code parallelization to produce accurate simulation results as fast as possible. The traditional modeling approaches require a lot of computing power and significant amounts of electrical energy; they are also highly dependent on uninterrupted access to a reliable power supply. This paper presents a concept how to develop suitable low power hardware architectures for tsunami and storm surge simulations based on cooperative software and hardware simulation. The main goal is to be able - if necessary - to perform simulations in-situ and battery-powered. For flood warning systems installed in regions with weak or unreliable power and computing infrastructure, this would significantly decrease the risk of failure at the most critical moments.
Dominik Schoenwetter, Alexander Ditter, Bruno Kleinert, Arne Hendricks, Vadym Aizinger, Harald Köstler, Dietmar Fey
SIMULTECH6
2015 Performance modeling and analysis of heterogeneous lattice Boltzmann simulations on CPU-GPU clusters
Christian Feichtinger, Johannes Habich, Harald Köstler, Ulrich Rüde, Takayuki Aoki
Parallel Comput.3
2015 A Gauss-Seidel Iteration Scheme for Reference-Free 3-D Histological Image Reconstruction
abstract
Three-dimensional (3-D) reconstruction of histological slice sequences offers great benefits in the investigation of different morphologies. It features very high-resolution which is still unmatched by in vivo 3-D imaging modalities, and tissue staining further enhances visibility and contrast. One important step during reconstruction is the reversal of slice deformations introduced during histological slice preparation, a process also called image unwarping. Most methods use an external reference, or rely on conservative stopping criteria during the unwarping optimization to prevent straightening of naturally curved morphology. Our approach shows that the problem of unwarping is based on the superposition of low-frequency anatomy and high-frequency errors. We present an iterative scheme that transfers the ideas of the Gauss-Seidel method to image stacks to separate the anatomy from the deformation. In particular, the scheme is universally applicable without restriction to a specific unwarping method, and uses no external reference. The deformation artifacts are effectively reduced in the resulting histology volumes, while the natural curvature of the anatomy is preserved. The validity of our method is shown on synthetic data, simulated histology data using a CT data set and real histology data. In the case of the simulated histology where the ground truth was known, the mean Target Registration Error (TRE) between the unwarped and original volume could be reduced to less than 1 pixel on average after six iterations of our proposed method.
Simone Gaffling, Volker Daum, Stefan Steidl, Andreas K. Maier, Harald Köstler, Joachim Hornegger
IEEE Trans. Medical Imaging5
2014 Parallel multigrid on hierarchical hybrid grids: a performance study on current high performance computing clusters
abstract
SUMMARY This article studies the performance and scalability of a geometric multigrid solver implemented within the hierarchical hybrid grids (HHG) software package on current high performance computing clusters up to nearly 300,000 cores. HHG is based on unstructured tetrahedral finite elements that are regularly refined to obtain a block‐structured computational grid. One challenge is the parallel mesh generation from an unstructured input grid that roughly approximates a human head within a 3D magnetic resonance imaging data set. This grid is then regularly refined to create the HHG grid hierarchy. As test platforms, a BlueGene/P cluster located at Jülich supercomputing center and an Intel Xeon 5650 cluster located at the local computing center in Erlangen are chosen. To estimate the quality of our implementation and to predict runtime for the multigrid solver, a detailed performance and communication model is developed and used to evaluate the measured single node performance, as well as weak and strong scaling experiments on both clusters. Thus, for a given problem size, one can predict the number of compute nodes that minimize the overall runtime of the multigrid solver. Overall, HHG scales up to the full machines, where the biggest linear system solved on Jugene had more than one trillion unknowns. Copyright © 2012 John Wiley & Sons, Ltd.
Björn Gmeiner, Harald Köstler, Markus Stürmer, Ulrich Rüde
Concurr. Comput. Pract. Exp.2
2014 Towards a performance-portable description of geometric multigrid algorithms using a domain-specific language
Richard Membarth, Oliver Reiche, Christian Schmitt 0003, Frank Hannig, Jürgen Teich, Markus Stürmer, Harald Köstler
J. Parallel Distributed Comput.7
2013 A framework for hybrid parallel flow simulations with a trillion cells in complex geometries
abstract
waLBerla is a massively parallel software framework for simulating complex flows with the lattice Boltzmann method (LBM). Performance and scalability results are presented for SuperMUC, the world's fastest x86-based supercomputer ranked number 6 on the Top500 list, and JUQUEEN, a Blue Gene/Q system ranked as number 5.
Christian Godenschwager, Florian Schornbaum, Martin Bauer 0003, Harald Köstler, Ulrich Rüde
SC4
2011 A flexible Patch-based lattice Boltzmann parallelization approach for heterogeneous GPU-CPU clusters
Christian Feichtinger, Johannes Habich, Harald Köstler, Georg Hager, Ulrich Rüde, Gerhard Wellein
Parallel Comput.3
2007 3D optical flow computation using a parallel variational multigrid scheme with application to cardiac C-arm CT motion
El Mostafa Kalmoun, Harald Köstler, Ulrich Rüde
Image Vis. Comput.2