Katharina Kormann

dblp:85/8754 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
1since 2021 · last 2021
0000-0003-1956-2073ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Theory of computation · 2 · 1 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2021 hyper.deal: An Efficient, Matrix-free Finite-element Library for High-dimensional Partial Differential Equations
abstract
This work presents the efficient, matrix-free finite-element library hyper.deal for solving partial differential equations in two to six dimensions with high-order discontinuous Galerkin methods. It builds upon the low-dimensional finite-element library deal.II to create complex low-dimensional meshes and to operate on them individually. These meshes are combined via a tensor product on the fly and the library provides new special-purpose highly optimized matrixfree functions exploiting domain decomposition as well as shared memory via MPI-3.0 features. Both node-level performance analyses and strong/weak-scaling studies on up to 147,456 CPU cores confirm the efficiency of the implementation. Results of the library hyper.deal are reported for high-dimensional advection problems and for the solution of the Vlasov-Poisson equation in up to 6D phase space. Copyright © 2020, arXiv, All rights reserved.
Peter Munch, Katharina Kormann, Martin Kronbichler 0002
ACM Trans. Math. Softw.2
2020 Evaluation of performance portability frameworks for the implementation of a particle-in-cell code
abstract
Summary This paper reports on an in‐depth evaluation of the performance portability frameworks Kokkos and RAJA with respect to their suitability for the implementation of complex particle‐in‐cell (PIC) simulation codes, extending previous studies based on codes from other domains. At the example of a particle‐in‐cell model, we implemented the hotspot of the code in C++ and parallelized it using OpenMP, OpenACC, CUDA, Kokkos, and RAJA, targeting multi‐core (CPU) and graphics (GPU) processors. Both Kokkos and RAJA appear mature, are usable for complex codes, and keep their promise to provide performance portability across different architectures. Comparing the obtainable performance on state‐of‐the art hardware, but also considering aspects such as code complexity, feature availability, and overall productivity, we finally draw the conclusion that the Kokkos framework would be suited best to tackle the massively parallel implementation of the full PIC model.
Victor Artigues, Katharina Kormann, Markus Rampp, Klaus Reuter
Concurr. Comput. Pract. Exp.2
2019 Fast Matrix-Free Evaluation of Discontinuous Galerkin Finite Element Operators
abstract
We present an algorithmic framework for matrix-free evaluation of discontinuous Galerkin finite element operators. It relies on fast quadrature with sum factorization on quadrilateral and hexahedral meshes, targeting general weak forms of linear and nonlinear partial differential equations. Different algorithms and data structures are compared in an in-depth performance analysis. The implementations of the local integrals are optimized by vectorization over several cells and faces and an even-odd decomposition of the one-dimensional interpolations. Up to 60% of the arithmetic peak on Intel Haswell, Broadwell, and Knights Landing processors is reached when running from caches and up to 40% of peak when also considering the access to vectors from main memory. On 2×14 Broadwell cores, the throughput is up to 2.2 billion unknowns per second for the 3D Laplacian and up to 4 billion unknowns per second for the 3D advection on affine geometries, close to a simple copy operation at 4.7 billion unknowns per second. Our experiments show that MPI ghost exchange has a considerable impact on performance and we present strategies to mitigate this effect. Finally, various options for evaluating geometry terms and their performance are discussed. Our implementations are publicly available through the deal.II finite element library.
Martin Kronbichler 0002, Katharina Kormann
ACM Trans. Math. Softw.2
2011 Parallel Finite Element Operator Application: Graph Partitioning and Coloring
abstract
We present an efficient implementation of parallel finite element operator application for hexahedral elements. The implementation is tailored to data structures for adaptively refined meshes and exploits parallelism on modern computer systems. The evaluation of local shape functions and gradients is performed with sum-factorization that makes use of the tensor-product form. For shared memory parallelization, we propose a novel two-level partitioning/coloring approach that avoids race conditions when writing into the result vector. We give evidence for the good performance of our implementation. We employ the optimized operator implementation on a problem in quantum dynamics described by the time-dependent Schroedinger equation. We obtain a speedup of more than a factor four over conventional solvers based on sparse matrices for a moderate polynomial order of four in three dimensions.
Katharina Kormann, Martin Kronbichler 0002
eScience1