Ulrich Rüde

dblp:75/2928 · DBLP profile ↗
← Back
28ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0001-8796-8599ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Complexity analysis and scalability of a matrix-free extrapolated geometric multigrid solver for curvilinear coordinates representations from fusion plasma applications
abstract
Tokamak fusion reactors are promising alternatives for future energy production. Gyrokinetic simulations are important tools to understand physical processes inside tokamaks and to improve the design of future plants. In gyrokinetic codes such as Gysela, these simulations involve at each time step the solution of a gyrokinetic Poisson equation defined on disk-like cross sections. The authors of [14] , [15] proposed to discretize a simplified differential equation using symmetric finite differences derived from the resulting energy functional and to use an implicitly extrapolated geometric multigrid scheme tailored to problems in curvilinear coordinates. In this article, we extend the discretization to a more realistic partial differential equation and demonstrate the optimal linear complexity of the proposed solver, in terms of computation and memory. We provide a general framework to analyze floating point operations and memory usage of matrix-free approaches for stencil-based operators. Finally, we give an efficient matrix-free implementation for the considered solver exploiting a task-based multithreaded parallelism which takes advantage of the disk-shaped geometry of the problem. We demonstrate the parallel efficiency for the solution of problems of size up to 50 million unknowns.
Philippe Leleux, Christina Schwarz, Martin Joachim Kühn, Carola Kruse, Ulrich Rüde
J. Parallel Distributed Comput.5
2024 waLBerla-wind: A lattice-Boltzmann-based high-performance flow solver for wind energy applications
abstract
Summary This article presents the development of a new wind turbine simulation software to study wake flow physics. To this end, the design and development of waLBerla‐wind , a new simulator based on the lattice‐Boltzmann method that is known for its excellent performance and scaling properties, will be presented. Here it will be used for large eddy simulations (LES) coupled with actuator wind turbine models. Due to its modular software design, waLBerla‐wind is flexible and extensible with regard to turbine configurations. Additionally it is performance portable across different hardware architectures, another critical design goal. The new solver is validated by presenting force distributions and velocity profiles and comparing them with experimental data and a vortex solver. Furthermore, waLBerla‐wind 's performance is compared to a theoretical peak performance, and analyzed with weak and strong scaling benchmarks on CPU and GPU systems. This analysis demonstrates the suitability for large‐scale applications and future cost‐effective full wind farm simulations.
Helen Schottenhamml, Ani Anciaux-Sedrakian, Frédéric Blondel, Harald Köstler, Ulrich Rüde
Concurr. Comput. Pract. Exp.5
2023 Scalable Flow Simulations with the Lattice Boltzmann Method
abstract
The primary goal of the EuroHPC JU project SCALABLE is to develop an industrial Lattice Boltzmann Method (LBM)-based computational fluid dynamics (CFD) solver capable of exploiting current and future extreme scale architectures, expanding current capabilities of existing industrial LBM solvers by at least two orders of magnitude in terms of processor cores and lattice cells, while preserving its accessibility from both the end-user and software developer's point of view. This is accomplished by transferring technology and knowledge between an academic code (waLBerla) and an industrial code (LaBS). This paper briefly introduces the characteristics and main features of both software packages involved in the process. We also highlight some of the performance achievements in scales of up to tens of thousand of cores presented on one academic and one industrial benchmark case.
Markus Holzer 0005, Gabriel Staffelbach, Ilan Rocchi, Jayesh Badwaik, Andreas Herten, Radim Vavrík, Ondrej Vysocky, Lubomir Riha, Romain Cuidard, Ulrich Rüde
CF10
2021 Parallel solution of saddle point systems with nested iterative solvers based on the Golub-Kahan Bidiagonalization
abstract
Summary The Golub‐Kahan bidiagonalization is widely used in the singular value decomposition of rectangular matrices and has been generalized to an iterative solver for symmetric indefinite linear systems with a two‐by‐two block structure. In this work, we present a scalability study of this generalized solver as implemented in a recent release of the parallel numerical library PETSc (Portable, Extensible Toolkit for Scientific Computation). We present an improved solver performance for the two‐dimensional (2D) Stokes equations as compared to previous work. Furthermore, we investigate the performance of different parallel inner solvers in the outer Golub‐Kahan iteration for a three‐dimensional Stokes problem. The study includes parallel sparse direct solvers and multigrid methods. When increasing the number of cores for a fixed total problem size, the solver exhibits good speedups of up to 50% at the 1024 core count. For the tests in which the total problem size grows while the workload in each core stays constant, the parallel performance of the solver scales almost linearly with the increase in the core counts. In particular, the computation time increases only by about 15% when the number of cores increases from 80 to 1024 for a 2D test case.
Carola Kruse, Masha Sosonkina, Mario Arioli, Nicolas Tardieu, Ulrich Rüde
Concurr. Comput. Pract. Exp.5
2019 Scalable GPU Communication with Code Generation on Stencil Applications
abstract
Clusters with GPUs are mainstream in HPC as shown by the last edition of the Top500 list, increasing the demand for GPU capable scientific computing software. Programming large scale GPU systems in an efficient and future-proof way present numerous challenges, such as optimizations for a variety of GPUs and interconnect hardware, hiding communication overhead with computation and efficient domain partitioning. We present an improvement to the CUDA-based communication of stencil applications in the WALBERLA framework, achieving scalability while supporting different GPUs and communication infrastructures. We utilize the lattice Boltzmann Method for fluid flows as a representative of stencil-based scientific computing and implement a communication hiding strategy that is capable of adjusting to a system's computing and communication capabilities. We compare the use of CUDAMemCopy with the use of customized pack/unpack kernels and show that packing achieves almost linear weak scaling behavior in the Santos Dumont supercomputer with up to 128 GPUs. We also show that the proposed approach is not sensitive to the direction of the domain partitioning, one of the biggest challenges when communicating 3D domains in GPU-based stencil simulations.
João Victor Tozatti Risso, Martin Bauer 0003, Paulo Roberto de Carvalho, Ulrich Rüde, Daniel Weingaertner
SBAC-PAD4
2019 Code generation for massively parallel phase-field simulations
abstract
This article describes the development of automatic program generation technology to create scalable phase-field methods for material science applications. To simulate the formation of microstructures in metal alloys, we employ an advanced, thermodynamically consistent phase-field method. A state-of-the-art large-scale implementation of this model requires extensive, time-consuming, manual code optimization to achieve unprecedented fine mesh resolution. Our new approach starts with an abstract description based on free-energy functionals which is formally transformed into a continuous PDE and discretized automatically to obtain a stencil-based time-stepping scheme. Subsequently, an automatized performance engineering process generates highly optimized, performance-portable code for CPUs and GPUs. We demonstrate the efficiency for real-world simulations on large-scale GPU-based (PizDaint) and CPU-based (SuperMUC-NG) supercomputers. Our technique simplifies program development and optimization for a wide class of models.
Martin Bauer 0003, Johannes Hötzer, Dominik Ernst, Julian Hammer, Marco Seiz, Henrik Hierl, Jan Hönig, Harald Köstler, Gerhard Wellein, Britta Nestler, Ulrich Rüde
SC11
2018 A Local Parallel Communication Algorithm for Polydisperse Rigid Body Dynamics
Sebastian Eibl, Ulrich Rüde
Parallel Comput.2
2015 Massively parallel phase-field simulations for ternary eutectic directional solidification
abstract
Microstructures forming during ternary eutectic directional solidification processes have significant influence on the macroscopic mechanical properties of metal alloys. For a realistic simulation, we use the well established thermodynamically consistent phase-field method and improve it with a new grand potential formulation to couple the concentration evolution. This extension is very compute intensive due to a temperature dependent diffusive concentration. We significantly extend previous simulations that have used simpler phase-field models or were performed on smaller domain sizes. The new method has been implemented within the massively parallel HPC framework waLBerla that is designed to exploit current supercomputers efficiently. We apply various optimization techniques, including buffering techniques, explicit SIMD kernel vectorization, and communication hiding. Simulations utilizing up to 262,144 cores have been run on three different supercomputing architectures and weak scalability results are shown. Additionally, a hierarchical, mesh-based data reduction strategy is developed to keep the I/O problem manageable at scale.
Martin Bauer 0003, Johannes Hötzer, Marcus Jainta, Philipp Steinmetz, Marco Berghoff, Florian Schornbaum, Christian Godenschwager, Harald Köstler, Britta Nestler, Ulrich Rüde
SC10
2015 Performance modeling and analysis of heterogeneous lattice Boltzmann simulations on CPU-GPU clusters
Christian Feichtinger, Johannes Habich, Harald Köstler, Ulrich Rüde, Takayuki Aoki
Parallel Comput.4
2014 Parallel multigrid on hierarchical hybrid grids: a performance study on current high performance computing clusters
abstract
SUMMARY This article studies the performance and scalability of a geometric multigrid solver implemented within the hierarchical hybrid grids (HHG) software package on current high performance computing clusters up to nearly 300,000 cores. HHG is based on unstructured tetrahedral finite elements that are regularly refined to obtain a block‐structured computational grid. One challenge is the parallel mesh generation from an unstructured input grid that roughly approximates a human head within a 3D magnetic resonance imaging data set. This grid is then regularly refined to create the HHG grid hierarchy. As test platforms, a BlueGene/P cluster located at Jülich supercomputing center and an Intel Xeon 5650 cluster located at the local computing center in Erlangen are chosen. To estimate the quality of our implementation and to predict runtime for the multigrid solver, a detailed performance and communication model is developed and used to evaluate the measured single node performance, as well as weak and strong scaling experiments on both clusters. Thus, for a given problem size, one can predict the number of compute nodes that minimize the overall runtime of the multigrid solver. Overall, HHG scales up to the full machines, where the biggest linear system solved on Jugene had more than one trillion unknowns. Copyright © 2012 John Wiley & Sons, Ltd.
Björn Gmeiner, Harald Köstler, Markus Stürmer, Ulrich Rüde
Concurr. Comput. Pract. Exp.4
2013 A framework for hybrid parallel flow simulations with a trillion cells in complex geometries
abstract
waLBerla is a massively parallel software framework for simulating complex flows with the lattice Boltzmann method (LBM). Performance and scalability results are presented for SuperMUC, the world's fastest x86-based supercomputer ranked number 6 on the Top500 list, and JUQUEEN, a Blue Gene/Q system ranked as number 5.
Christian Godenschwager, Florian Schornbaum, Martin Bauer 0003, Harald Köstler, Ulrich Rüde
SC5
2012 Hierarchical Hybrid Grids for Mantle Convection: A First Study
abstract
In this article we consider the application of the Hierarchical Hybrid Grid Framework (HHG) to the geodynamical problem of simulating mantle convection. We describe the generation of a refined icosahedral grid and a further subdivision of the resulting prisms into tetrahedral elements. Based on this mesh, we present performance results for HHG and compare these to the also Finite Element program TERRA, which is a well-known code for mantle convection using a matrix-free representation of the stiffness matrix. In our analysis we consider the most time consuming part of TERRA's solution algorithm and evaluate it in a strong scaling setup. Finally we present strong and weak scaling results for HHG to verify its parallel concepts, algorithms and grid flexibility on Jugene.
Björn Gmeiner, Marcus Mohr 0001, Ulrich Rüde
ISPDC3
2011 Implementation of Multigrid on QPACE
abstract
We developed and optimized a multigrid method on the QPACE cluster. The QPACE cluster is an acclerator-based cluster using the Power Cell 8i CPU that is built by the special research field SFB TR 55 for Lattice Quantum Chromo dynamics computations. The cluster uses a custom 3D to rus network build using FPGAs. Our goal was to evaluate the QPACE architecture for a type of algorithm that uses a communication pattern not limited to nearest neighbor communication. We provide a model of the communication network taking into account the specific characteristics of the network and the network processor. For the implementation we chose to use an accelerator-centric programming model by using the SPUs, only.
Matthias Bolten, Daniel Brinkers, Ulrich Rüde, Markus Stürmer
CLUSTER3
2011 A flexible Patch-based lattice Boltzmann parallelization approach for heterogeneous GPU-CPU clusters
Christian Feichtinger, Johannes Habich, Harald Köstler, Georg Hager, Ulrich Rüde, Gerhard Wellein
Parallel Comput.5
2010 Direct Numerical Simulation of Particulate Flows on 294912 Processor Cores
abstract
This paper describes computational models for particle-laden flows based on a fully resolved fluid-structure interaction. The flow simulation uses the Lattice Boltzmann method, while the particles are handled by a rigid body dynamics algorithm. The particles can have individual non-spherical shapes, creating the need for a non-trivial collision detection and special contact models. An explicit coupling algorithm transfers momenta from the fluid to the particles in each time step, while the particles impose moving boundaries for the flow solver. All algorithms and their interaction are fully parallelized. Scaling experiments and a careful performance analysis are presented for up to 294912 processor cores of the Blue Gene at the Jülich Supercomputing center. The largest simulations involve 264 million particles that are coupled to a fluid which is simultaneously resolved by 150 billion cells for the Lattice Boltzmann method. The paper will conclude with a computational experiment for the segregation of suspensions of particles of different density, as an example of the many industrial applications that are enabled by this new methodology.
Jan Götz, Klaus Iglberger, Markus Stürmer, Ulrich Rüde
SC4
2010 Coupling multibody dynamics and computational fluid dynamics on 8192 processor cores
Jan Götz, Klaus Iglberger, Christian Feichtinger, Stefan Donath, Ulrich Rüde
Parallel Comput.5
2009 Localized Parallel Algorithm for Bubble Coalescence in Free Surface Lattice-Boltzmann Method
Stefan Donath, Christian Feichtinger, Thomas Pohl, Jan Götz, Ulrich Rüde
Euro-Par5
2009 A Parallel Rigid Body Dynamics Algorithm
Klaus Iglberger, Ulrich Rüde
Euro-Par2
2009 Detail-preserving fluid control
Nils Thürey, Richard Keiser, Mark Pauly, Ulrich Rüde
Graph. Model.4
2007 3D optical flow computation using a parallel variational multigrid scheme with application to cardiac C-arm CT motion
El Mostafa Kalmoun, Harald Köstler, Ulrich Rüde
Image Vis. Comput.3
2005 Is 1.7 x 10^10 Unknowns the Largest Finite Element System that Can Be Solved Today?
abstract
The hierarchical hybrid Grids (HHG) framework attempts to remove limitations on the size of problem that can be solved using a finite element discretization of a partial differential equation (PDE) by using a process of regular refinement, of an unstructured input grid, to generate a nested hierarchy of patch-wise structured grids that is suitable for use with geometric multigrid. The regularity of the resulting grids may be exploited in such a way that it is no longer necessary to explicitly assemble the global discretization matrix. In particular, given an appropriate input grid, the discretization matrix may be defined implicitly using stencils that are constant for each structured patch. This drastically reduces the amonnt of memory required for the discretization, thus allowing for a much larger problem to be solved. Here we present a brief description of the HHG framework: detailing the principles that led to solving a finite element system with 1.7 x 10^10 unknowns, on an SGI Altix supercomputer, using 1024 nodes, with an overall performance of 0.96 TFLOP/s, on a logically unstructured grid, using geometric mmiltigrid as a solver.
Ben Bergen 0002, Frank Hülsemann, Ulrich Rüde
SC3
2004 Performance Evaluation of Parallel Large-Scale Lattice Boltzmann Applications on Three Supercomputing Architectures
abstract
Computationally intensive programs with moderate communication requirements such as CFD codes suffer from the standard slow interconnects of commodity "off the shelf" (COTS) hardware. We will introduce different large-scale applications of the Lattice Boltzmann Method (LBM) in fluid dynamics, material science, and chemical engineering and present results of the parallel performance on different architectures. It will be shown that a high speed communication network in combination with an efficient CPU is mandatory in order to achieve the required performance. An estimation of the necessary CPU count to meet the performance of 1 TFlop/s will be given as well as a prediction as to which architecture is the most suitable for LBM. Finally, ratios of costs to application performance for tailored HPC systems and COTS architectures will be presented.
Thomas Pohl, Frank Deserno, Nils Thürey, Ulrich Rüde, Peter Lammers, Gerhard Wellein, Thomas Zeiser
SC4
2003 Hierarchical Hybrid Grids as Basis for Parallel Numerical Solution of PDE
Frank Hülsemann, Ben Bergen 0002, Ulrich Rüde
Euro-Par3
2003 Cache Performance Optimizations for Parallel Lattice Boltzmann Codes
Jens Wilke, Thomas Pohl, Markus Kowarschik, Ulrich Rüde
Euro-Par4
2003 Editorial
Rosemary A. Renaut, Ulrich Rüde
Future Gener. Comput. Syst.2
2000 Numerical Algorithms for Linear and Nonlinear Algebra
Ulrich Rüde, Hans-Joachim Bungartz
Euro-Par1
1999 Memory Characteristics of Iterative Methods
abstract
Conventional implementations of iterative numerical algorithms, especially multigrid methods, merely reach a disappointing small percentage of the theoretically available CPU performance when applied to representative large problems.One of the most important reasons for this phenomenon is that the current DRAM technology cannot provide the data fast enough to keep the CPU busy.Although the fundamentals of cache optimizations are quite simple, current compilers cannot optimize even elementary iterative schemes.In this paper, we analyze the memory and cache behavior of iterative methods with extensive profiling and describe program transformation techniques to improve the cache performance of two-and three-dimensional multigrid algorithms.This project is partially funded by DFG Ru 422/7-1,2.1 All benchmarks in the article were compiled with native FORTRAN77 compilers and aggressive optimizations enabled.On the Intel platform we used egcs (V2.91.60).The platforms include an Intel PentiumII Xeon PC (450 MHz, 450 MFLOPS), a SUN Ultra 60 (296 MHz, 592 MFLOPS), a HP SPP2200 Convex Exemplar Node (200 MHz, 800 MFLOPS), a Compaq PWS 500au (500 MHz, 1 GFLOPS), and a Compaq XP1000 (500 MHz, 1 GFLOPS).1
Christian Weiß 0001, Wolfgang Karl, Markus Kowarschik, Ulrich Rüde
SC4
1997 Iterative Algorithms on High Performance Architectures
Ulrich Rüde
Euro-Par1