Michael Bader

dblp:14/1333 · DBLP profile ↗
← Back
20ranked-venue papers
2as first author
5since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 2 first-author · 4 since 2021Theory of computation · 2
YearPublicationVenuePosition
2024 Fused GEMMs towards an efficient GPU implementation of the ADER-DG method in SeisSol
abstract
Summary This study shows how GPU performance of the ADER discontinuous Galerkin method in SeisSol (an earthquake simulation software) can be further improved while preserving its original design that ensures high CPU performance. We introduce a new code generator (“ChainForge”) that fuses subsequent batched matrix multiplications (“GEMMs”) into a single GPU kernel, holding intermediate results in shared memory as long as necessary. The generator operates as an external module linked against SeisSol's domain specific language YATeTo and, as a result, the original SeisSol source code remains mainly unchanged. In this paper, we discuss several challenges related to automatic fusion of GPU kernels and provide solutions to them. By and large, we gain 60% in performance of SeisSol's wave propagation solver using Fused‐GEMMs compared to the original GPU implementation. We demonstrated this on benchmarks as well as on a real production scenario simulating the Northridge 1994 earthquake.
Ravil Dorozhinskii, Gonzalo Brito Gadeschi, Michael Bader
Concurr. Comput. Pract. Exp.3
2023 The EU Center of Excellence for Exascale in Solid Earth (ChEESE): Implementation, results, and roadmap for the second phase
abstract
The EU Center of Excellence for Exascale in Solid Earth (ChEESE) develops exascale transition capabilities in the domain of Solid Earth, an area of geophysics rich in computational challenges embracing different approaches to exascale (capability, capacity, and urgent computing). The first implementation phase of the project (ChEESE-1P; 2018–2022) addressed scientific and technical computational challenges in seismology, tsunami science, volcanology, and magnetohydrodynamics, in order to understand the phenomena, anticipate the impact of natural disasters, and contribute to risk management. The project initiated the optimisation of 10 community flagship codes for the upcoming exascale systems and implemented 12 Pilot Demonstrators that combine the flagship codes with dedicated workflows in order to address the underlying capability and capacity computational challenges. Pilot Demonstrators reaching more mature Technology Readiness Levels (TRLs) were further enabled in operational service environments on critical aspects of geohazards such as long-term and short-term probabilistic hazard assessment, urgent computing, and early warning and probabilistic forecasting. Partnership and service co-design with members of the project Industry and User Board (IUB) leveraged the uptake of results across multiple research institutions, academia, industry, and public governance bodies (e.g. civil protection agencies). This article summarises the implementation strategy and the results from ChEESE-1P, outlining also the underpinning concepts and the roadmap for the on-going second project implementation phase (ChEESE-2P; 2023–2026).
Arnau Folch, Claudia Abril, Michael Afanasiev, Giorgio Amati, Michael Bader, Rosa M. Badia, Hafize B. Bayraktar, Sara Barsotti, Roberto Basili 0002, Fabrizio Bernardi, Christian Boehm, Beatriz Brizuela, Federico Brogi, Eduardo Cabrera, Emanuele Casarotti, Manuel Jesús Castro Díaz, Matteo Cerminara, Antonella Cirella, Alexey Cheptsov, Javier Conejero, Antonio Costa 0002, Marc de la Asunción, Josep de la Puente, Marco Djuric, Ravil Dorozhinskii, Gabriela Espinosa, Tomaso Esposti Ongaro, Joan Farnós, Nathalie Favretto-Cristini, Andreas Fichtner, Alexandre Fournier, Alice-Agnes Gabriel, Jean-Matthieu Gallard, Steven J. Gibbons, Sylfest Glimsdal, José Manuel González-Vida, José Gracia, Rose Gregorio, Natalia Gutiérrez, Benedikt Halldorsson, Okba Hamitou, Guillaume Houzeaux, Stephan Jaure, Mouloud Kessar, Lukas Krenz, Lion Krischer, Soline Laforet, Piero Lanucara, Bo Li 0147, Maria Concetta Lorenzino, Stefano Lorito, Finn Løvholt, Giovanni Macedonio, Jorge Macías Sánchez, Guillermo Marin, Beatriz Martínez Montesinos, Leonardo Mingari, Geneviève Moguilny, Vadim Montellier, Marisol Monterrubio Velasco, Georges-Emmanuel Moulard, Masaru Nagaso, Massimo Nazaria, Christoph Niethammer, Federica Pardini, Marta Pienkowska, Luca Pizzimenti, Natalia Poiata, Leonhard Rannabauer, Otilio Rojas, Juan Esteban Rodriguez, Fabrizio Romano, Oleksandr Rudyy, Vittorio Ruggiero, Philipp Samfass, Carlos Sánchez-Linares, Sabrina Sanchez, Laura Sandri, Antonio Scala, Nathanaël Schaeffer, Joseph Schuchart, Jacopo Selva, Amadine Sergeant, Angela Stallone, Matteo Taroni, Solvi Thrastarson, Manuel Titos, Nadia Tonelllo, Roberto Tonini, Thomas Ulrich, Jean-Pierre Vilotte, Malte Vöge, Manuela Volpe, Sara Aniko Wirp, Uwe Wössner
Future Gener. Comput. Syst.5
2021 SeisSol on Distributed Multi-GPU Systems: CUDA Code Generation for the Modal Discontinuous Galerkin Method
abstract
We present a GPU implementation of the high order Discontinuous Galerkin (DG) scheme in SeisSol, a software package for simulating seismic waves and earthquake dynamics. Our particular focus is on providing a performance portable solution for heterogeneous distributed multi-GPU systems. We therefore redesigned SeisSol’s code generation cascade for GPU programming models. This includes CUDA source code generation for the performance-critical small batched matrix multiplications kernels. The parallelisation extends the existing MPI+X scheme and supports SeisSol’s cluster-wise Local Time Stepping (LTS) algorithm for ADER time integration.
Ravil Dorozhinskii, Michael Bader
HPC Asia2
2021 3D acoustic-elastic coupling with gravity: the dynamics of the 2018 palu, sulawesi earthquake and tsunami
abstract
We present a highly scalable 3D fully-coupled Earth & ocean model of earthquake rupture and tsunami generation and perform the first fully coupled simulation of an actual earthquake-tsunami event and a 3D benchmark problem of tsunami generation by a megathrust dynamic earthquake rupture. Multi-petascale simulations, with excellent performance demonstrated on three different platforms, allow high-resolution forward modeling. Our largest mesh has ≈261 billion degrees of freedom, resolving at least 15 Hz of the acoustic wave field. We self-consistently model seismic, acoustic and surface gravity wave propagation in elastic (Earth) and acoustic (ocean) materials sourced by physics-based non-linear earthquake dynamic rupture, thereby gaining insight into the tsunami generation process without relying on approximations that have previously been applied to permit solution of this challenging problem. Complicated geometries, including high-resolution bathymetry, coastlines and segmented earthquake faults are discretized by adaptive unstructured tetrahedral meshes. This inevitably leads to large differences in element sizes and wave speeds which can be mitigated by ADER local time-stepping and a Discontinuous Galerkin discretization yielding high-order accuracy in time and space.
Lukas Krenz, Carsten Uphoff, Thomas Ulrich, Alice-Agnes Gabriel, Lauren S. Abrahams, Eric M. Dunham, Michael Bader
SC7
2021 High performance uncertainty quantification with parallelized multilevel Markov chain Monte Carlo
abstract
Numerical models of complex real-world phenomena often necessitate High Performance Computing (HPC). Uncertainties increase problem dimensionality further and pose even greater challenges.
Linus Seelinger, Anne Reinarz, Leonhard Rannabauer, Michael Bader, Peter Bastian, Robert Scheichl
SC4
2020 Lightweight task offloading exploiting MPI wait times for parallel adaptive mesh refinement
abstract
Summary Balancing the workload of sophisticated simulations is inherently difficult, since we have to balance both computational workload and memory footprint over meshes that can change any time or yield unpredictable cost per mesh entity, while modern supercomputers and their interconnects start to exhibit fluctuating performance. We propose a novel lightweight balancing technique for MPI+X to accompany traditional, prediction‐based load balancing. It is a reactive diffusion approach that uses online measurements of MPI idle time to migrate tasks temporarily from overloaded to underemployed ranks. Tasks are deployed to ranks which otherwise would wait, processed with high priority, and made available to the overloaded ranks again. This migration is nonpersistent. Our approach hijacks idle time to do meaningful work and is totally nonblocking, asynchronous and distributed without a global data view. Tests with a seismic simulation code developed in the ExaHyPE engine uncover the method's potential. We found speed‐ups of up to 2‐3 for ill‐balanced scenarios without logical modifications of the code base and show that the strategy is capable to react quickly to temporarily changing workload or node performance.
Philipp Samfass, Tobias Weinzierl, Dominic Etienne Charrier, Michael Bader
Concurr. Comput. Pract. Exp.4
2020 CHAMELEON: Reactive Load Balancing for Hybrid MPI+OpenMP Task-Parallel Applications
abstract
Many applications in high performance computing are designed based on underlying performance and execution models. While these models could successfully be employed in the past for balancing load within and between compute nodes, modern software and hardware increasingly make performance predictability difficult if not impossible. Consequently, balancing computational load becomes much more difficult. Aiming to tackle these challenges in search for a general solution, we present a novel library for fine-granular task-based reactive load balancing in distributed memory based on MPI and OpenMP. With our approach, individual migratable tasks can be executed on any MPI rank. The actual executing rank is determined at run time based on online performance data. We evaluate our approach under an enforced power cap and under enforced clock frequency changes for a synthetic benchmark and show its robustness for work-induced imbalances for a realistic application. Our experiments demonstrate speedups of up to 1.31X.
Jannis Klinkenberg, Philipp Samfass, Michael Bader, Christian Terboven, Matthias S. Müller
J. Parallel Distributed Comput.3
2020 Yet Another Tensor Toolbox for Discontinuous Galerkin Methods and Other Applications
abstract
The numerical solution of partial differential equations is at the heart of many grand challenges in supercomputing. Solvers based on high-order discontinuous Galerkin (DG) discretisation have been shown to scale on large supercomputers with excellent performance and efficiency if the implementation exploits all levels of parallelism and is tailored to the specific architecture. However, every year new supercomputers emerge and the list of hardware-specific considerations grows simultaneously with the list of desired features in a DG code. Thus, we believe that a sustainable DG code needs an abstraction layer to implement the numerical scheme in a suitable language. We explore the possibility to abstract the numerical scheme as small tensor operations, describe them in a domain-specific language (DSL) resembling the Einstein notation, and to map them to small General Matrix-Matrix Multiplication routines. The compiler for our DSL implements classic optimisations that are used for large tensor contractions, and we present novel optimisation techniques such as equivalent sparsity patterns and optimal index permutations for temporary tensors. Our application examples, which include the earthquake simulation software SeisSol, show that the generated kernels achieve over 50% peak performance of a recent 48-core Skylake system while the DSL considerably simplifies the implementation.
Carsten Uphoff, Michael Bader
ACM Trans. Math. Softw.2
2018 Hybrid MPI+OpenMP Reactive Work Stealing in Distributed Memory in the PDE Framework sam(oa)^2
abstract
"Equal work results in equal execution time" is an assumption that has fundamentally driven design and implementation of parallel applications for decades. However, increasing hardware variability on current architectures (e.g., through Turbo Boost, dynamic voltage and frequency scaling or thermal effects) necessitate a revision of this assumption. Expecting an increase of these effects on future (exascale-)systems, in this paper, we present reactive work stealing across nodes on distributed memory machines using only MPI and OpenMP. We develop a novel distributed work stealing concept that - based on on-line performance monitoring - selectively steals and remotely executes tasks across MPI boundaries. This concept has been implemented in the parallel adaptive mesh refinement (AMR) framework sam(oa)2for OpenMP tasks of traversing a grid section. Corresponding performance measurements in the presence of enforced CPU clock frequency imbalances demonstrate that a state-of-the-art cost-based (chains-on-chains partitioning) load balancing mechanism is insufficient and can even degrade performance while distributed work stealing successfully mitigates the frequency-induced imbalances. Furthermore, our results indicate that our approach is also suitable for load balancing work-induced imbalances in a realistic AMR test case.
Philipp Samfass, Jannis Klinkenberg, Michael Bader
CLUSTER3
2017 Extreme scale multi-physics simulations of the tsunamigenic 2004 sumatra megathrust earthquake
abstract
We present a high-resolution simulation of the 2004 Sumatra-Andaman earthquake, including non-linear frictional failure on a megathrustsplay fault system. Our method exploits unstructured meshes capturing the complicated geometries in subduction zones that are crucial to understand large earthquakes and tsunami generation. These up-to-date largest and longest dynamic rupture simulations enable analysis of dynamic source effects on the seafloor displacements.
Carsten Uphoff, Sebastian Rettenberger, Michael Bader, Elizabeth H. Madden, Thomas Ulrich, Stephanie Wollherr, Alice-Agnes Gabriel
SC3
2017 Parallel Memory-Efficient Adaptive Mesh Refinement on Structured Triangular Meshes with Billions of Grid Cells
abstract
We present sam(oa)2, a software package for a dynamically adaptive, parallel solution of 2D partial differential equations on triangular grids created via newest vertex bisection. An element order imposed by the Sierpinski space-filling curve provides an algorithm for grid generation, refinement, and traversal that is inherently memory efficient. Based purely on stack and stream data structures, it completely avoids random memory access. Using an element-oriented data view suitable for local operators, concrete simulation scenarios are implemented based on control loops and event hooks, which hide the complexity of the underlying traversal scheme. Two case studies are presented: two-phase flow in heterogeneous porous media and tsunami wave propagation, demonstrated on the Tohoku tsunami 2011 in Japan. sam(oa)2features hybrid MPI+OpenMP parallelization based on the Sierpinski order induced on the elements. Sections defined by contiguous grid cells define atomic tasks for OpenMP work sharing and stealing, as well as for migration of grid cells between MPI processes. Using optimized communication and load balancing algorithms, sam(oa)2achieves 88% strong scaling efficiency from 16 to 512 cores and 92% efficiency in a weak scaling test on 8,192 cores with 10 billion elements—all tests including adaptive mesh refinement and load balancing in each time step.
Oliver Meister, Kaveh Rahnema, Michael Bader
ACM Trans. Math. Softw.3
2016 Petascale Local Time Stepping for the ADER-DG Finite Element Method
abstract
In this work we present a clustered local time stepping (LTS) scheme for the arbitrary high-order derivatives discontinuous Galerkin finite element scheme. By clustering elements of similar time step, our scheme meets regularity requirements of modern hardware through the design of the numerical discretization. We present a detailed description of our clustered local time stepping scheme for the seismic simulation package SeisSol. Our scheme is able to capture homogeneous and heterogeneous time step variations in the computational domain and maintains a large fraction of the theoretical speedup offered by LTS. From an engineering standpoint, our scheme addresses all important performance characteristics of state-of-the-art supercomputers. The combined algorithmic and computational performance results for SeisSol show that we are able to leverage the large potential of local time stepping by reducing time-to-solution by several factors (2.3 - 4.1), sustaining more than 53% of SuperMUC-II's HPL performance, what corresponds to more than 1.5 PFLOPS performance on 86,016 cores.
Alexander Breuer, Alexander Heinecke, Michael Bader
IPDPS3
2015 Optimizing I/O for Petascale Seismic Simulations on Unstructured Meshes
abstract
SeisSol simulates earthquake dynamics by coupling seismic wave propagation and dynamic rupture simulations with high order accuracy on fully adaptive, unstructured meshes. In this paper we present an optimization of SeisSol's I/O implementations to establish a workflow that supports petascale simulations on large unstructured datasets. Our implementations can handle meshes with more than 1 billion cells and 660 billion degrees of reedom. The results show that SeisSol can initialize the mesh structure within 35 seconds on 2048 SuperMUC nodes from our new optimized mesh format. For the wave field output we implemented carefully tuned I/O routines based on HDF5 and MPI-IO. With an aggregation strategy we are able to increase the write bandwidth from 832 MiB/s to 6.7 GiB/s on 2048 SuperMUC nodes.
Sebastian Rettenberger, Michael Bader
CLUSTER2
2014 Petascale High Order Dynamic Rupture Earthquake Simulations on Heterogeneous Supercomputers
abstract
We present an end-to-end optimization of the innovative Arbitrary high-order DERivative Discontinuous Galerkin (ADER-DG) software SeisSol targeting Intel® Xeon Phi coprocessor platforms, achieving unprecedented earthquake model complexity through coupled simulation of full frictional sliding and seismic wave propagation. SeisSol exploits unstructured meshes to flexibly adapt for complicated geometries in realistic geological models. Seismic wave propagation is solved simultaneously with earthquake faulting in a multiphysical manner leading to a heterogeneous solver structure. Our architecture aware optimizations deliver up to 50% of peak performance, and introduce an efficient compute-communication overlapping scheme shadowing the multiphysics computations. SeisSol delivers near-optimal weak scaling, reaching 8.6 DP-PFLOPS on 8,192 nodes of the Tianhe-2 supercomputer. Our performance model projects reaching 18 -- 20 DP-PFLOPS on the full Tianhe-2 machine. Of special relevance to modern civil engineering needs, our pioneering simulation of the 1992 Landers earthquake shows highly detailed rupture evolution and ground motion at frequencies up to 10 Hz.
Alexander Heinecke, Alexander Breuer, Sebastian Rettenberger, Michael Bader, Alice-Agnes Gabriel, Christian Pelties, Arndt Bode, William L. Barth, Xiangke Liao, Karthikeyan Vaidyanathan, Mikhail Smelyanskiy, Pradeep Dubey
SC4
2013 Improving kinetic energy storage for vehicles through the combination of rolling element and active magnetic bearings
abstract
The demand for short term energy storage providing high power for electric and hybrid-electric vehicles is increasing dramatically. Stationary flywheel energy storage systems (FESS) are established as uninterruptible power supply (UPS) and represent an emerging market. In contrast, mobile FESS are currently only used in few applications such as motor sports. To enable a wider use in personal and public transportation the lifespan of the flywheel's bearings needs to be increased significantly. This paper presents an alternative approach to extend the lifespan of the flywheel's bearings by using a combination of rolling element and active magnetic bearings (CREAMB).
Manes Recheis, Armin Buchroithner, Ivan Andrasec, Thomas Gallien, Bernhard Schweighofer, Michael Bader, Hannes Wegleiter
IECON6
2012 Shared memory parallelization of fully-adaptive simulations using a dynamic tree-split and -join approach
abstract
In this work we present an approach for the parallelization of hyperbolic simulations on shared-memory architectures running on fully-adaptive grids. We tackle the parallelization problem with a dynamic sub-tree split- and join-approach by running computations on those split sub-trees in parallel using lightweight tasks. The traversal of sub-trees created by tree-splittings is built upon an inherently cache efficient approach for solving hyperbolic PDEs on dynamically adaptive triangular grids using a Sierpiński space filling curve. Our communication scheme among sub-trees stores the exchange-data to/from adjacent sub-trees in a consecutive memory area which is further utilized for an improved run-length-encoded data exchange. To give results for a concrete scenario, we implemented a solver for the shallow water equations which demands for fully-adaptive grid refinement and coarsening after each time-step. Our results give detailed statistics about optimization of the split size, parallelization overhead and also strong scalability results for a simulation running on multi-socket Intel and AMD architectures.
Martin Schreiber 0001, Hans-Joachim Bungartz, Michael Bader
HiPC3
2012 Teaching Parallel Programming Models on a Shallow-Water Code
abstract
We present a software package that supports teaching different parallel programming models in a computational science and engineering context. It implements a Finite Volume solver for the shallow water equations, with application to tsunami simulation in mind. The numerical model is kept simple, using patches of Cartesian grids as computational domain, which can be connected via ghost layers. The Finite Volume method is restricted to piecewise constant approximation in each grid cell, but the computation of fluxes between cells can be based on the simple Lax-Friedrichs method, as well as on versatile approximate Riemann solvers, which allows realistic simulations. We present how this code can be used to study parallelization with CUDA, MPI, OpenMP, and hybrid approaches - and is useful for both introductory lectures in parallel computing and more advanced courses.
Alexander Breuer, Michael Bader
ISPDC2
2010 Matrix exponentials and parallel prefix computation in a quantum control problem
Thomas Auckenthaler, Michael Bader, Thomas Huckle, A. Spörl, Konrad Waldherr
Parallel Comput.2
2008 Exploiting the Locality Properties of Peano Curves for Parallel Matrix Multiplication
Michael Bader
Euro-Par1
2000 A Fast Solver for Convection Diffusion Equations Based on Nested Dissection with Incomplete Elimination
Michael Bader, Christoph Zenger 0001
Euro-Par1