EDBT 2026 Demo / reviewers in the wild / expert
Matthew G. Knepley
dblp:30/5730
· DBLP profile ↗
8ranked-venue papers
1as first author
4since 2021 · last 2022
0000-0002-2292-0735ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 since 2021Theory of computation · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Landau collision operator in the CUDA programming model applied to thermal quench plasmasabstractCollisional processes are critical in the understanding of non-Maxwellian plasmas. The Landau form of the Fokker-Planck equation is the gold standard for modeling collisions in most plasmas, however$\mathcal{O}(N^{2})$work complexity inhibits its widespread use. We show that with advanced numerical methods and GPU hardware this cost can be effectively mitigated. This paper extends previous work on a conservative, high order accurate, finite element discretization with adaptive mesh refinement of the Landau operator, with extensions to GPU hardware and implementations in both the CUDA and Kokkos programming languages. This work focuses on the Landau kernels and on NVIDIA hardware, however preliminary results on AMD and Fujitsu/ARM hardware, as well as end-to-end performance of a velocity space model of a plasma thermal quench, are also presented. Both the fully implicit Landau time integrator and the plasma thermal quench model are publicly available in PETSc (Portable, Extensible, Toolkit for Scientific computing). Mark F. Adams, Dylan P. Brennan, Matthew G. Knepley |
IPDPS | 3 |
| 2022 | The PetscSF Scalable Communication LayerabstractPetscSF, the communication component of the Portable, Extensible Toolkit for Scientific Computation (PETSc), is designed to provide PETSc's communication infrastructure suitable for exascale computers that utilize GPUs and other accelerators. PetscSF provides a simple application programming interface (API) for managing common communication patterns in scientific computations by using a star-forest graph representation. PetscSF supports several implementations based on MPI and NVSHMEM, whose selection is based on the characteristics of the application or the target architecture. An efficient and portable model for network and intra-node communication is essential for implementing large-scale applications. The Message Passing Interface, which has been the de facto standard for distributed memory systems, has developed into a large complex API that does not yet provide high performance on the emerging heterogeneous CPU-GPU-based exascale systems. In this article, we discuss the design of PetscSF, how it can overcome some difficulties of working directly with MPI on GPUs, and we demonstrate its performance, scalability, and novel features. Junchao Zhang 0002, Jed Brown, Satish Balay, Jacob Faibussowitsch, Matthew G. Knepley, Oana Marin, Richard Tran Mills, Todd S. Munson, Barry Smith 0002, Stefano Zampini |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2021 | Toward performance-portable PETSc for GPU-based exascale systems
Richard Tran Mills, Mark F. Adams, Satish Balay, Jed Brown, Alp Dener, Matthew G. Knepley, Scott Kruger, Hannah Morgan, Todd S. Munson, Karl Rupp, Barry Smith 0002, Stefano Zampini, Hong Zhang 0006, Junchao Zhang 0002 |
Parallel Comput. | 6 |
| 2021 | PCPATCH: Software for the Topological Construction of Multigrid Relaxation MethodsabstractEffective relaxation methods are necessary for good multigrid convergence. For many equations, standard Jacobi and Gauß–Seidel are inadequate, and more sophisticated space decompositions are required; examples include problems with semidefinite terms or saddle point structure. In this article, we present a unifying software abstraction, PCPATCH, for the topological construction of space decompositions for multigrid relaxation methods. Space decompositions are specified by collecting topological entities in a mesh (such as all vertices or faces) and applying a construction rule (such as taking all degrees of freedom in the cells around each entity). The software is implemented in PETSc and facilitates the elegant expression of a wide range of schemes merely by varying solver options at runtime. In turn, this allows for the very rapid development of fast solvers for difficult problems. Patrick E. Farrell, Matthew G. Knepley, Lawrence Mitchell, Florian Wechsung |
ACM Trans. Math. Softw. | 2 |
| 2018 | A performance spectrum for parallel computational frameworks that solve PDEsabstractSummary Important computational physics problems are often large‐scale in nature, and it is highly desirable to have robust and high performing computational frameworks that can quickly address these problems. However, it is no trivial task to determine whether a computational framework is performing efficiently or is scalable. The aim of this paper is to present various strategies for better understanding the performance of any parallel computational frameworks for solving PDEs. Important performance issues that negatively impact time‐to‐solution are discussed, and we propose a performance spectrum analysis that can enhance one's understanding of critical aforementioned performance issues. As proof of concept, we examine commonly used finite element simulation packages and software and apply the performance spectrum to quickly analyze the performance and scalability across various hardware platforms, software implementations, and numerical discretizations. It is shown that the proposed performance spectrum is a versatile performance model that is not only extendable to more complex PDEs such as hydrostatic ice sheet flow equations but also useful for understanding hardware performance in a massively parallel computing environment. Potential applications and future extensions of this work are also discussed. Justin Chang 0001, K. B. Nakshatrala, Matthew G. Knepley, S. Lennart Johnsson |
Concurr. Comput. Pract. Exp. | 3 |
| 2016 | A stochastic performance model for pipelined Krylov methodsabstractSummary Pipelined Krylov methods seek to ameliorate the latency due to inner products necessary for projection by overlapping it with the computation associated with sparse matrix‐vector multiplication. We clarify a folk theorem that this can only result in a speedup of 2× over the naive implementation. Examining many repeated runs, we show that stochastic noise also contributes to the latency, and we model this using an analytical probability distribution. Our analysis shows that speedups greater than 2× are possible with these algorithms. Copyright © 2016 John Wiley & Sons, Ltd. Hannah Morgan, Matthew G. Knepley, Patrick Sanan, L. Ridgway Scott |
Concurr. Comput. Pract. Exp. | 2 |
| 2013 | Finite Element Integration on GPUsabstractWe present a novel finite element integration method for low-order elements on GPUs. We achieve more than 100GF for element integration on first order discretizations of both the Laplacian and Elasticity operators on an NVIDIA GTX285, which has a nominal single precision peak flop rate of 1 TF/s and bandwidth of 159 GB/s, corresponding to a bandwidth limited peak of 40 GF/s. Matthew G. Knepley, Andy R. Terrel |
ACM Trans. Math. Softw. | 1 |
| 2012 | Composable Linear Solvers for MultiphysicsabstractThe Portable, Extensible Toolkit for Scientific computing (PETSc), which focuses on the scalable solution of problems based on partial differential equations, now incorporates new components that allow full compos ability of solvers for multiphysics and multilevel methods. Through strong encapsulation, we achieve arbitrary, dynamic composition of hierarchical methods for coupled problems and allow customization of all components in composite solvers. For example, we support block decompositions with nested multigrid as well as multigrid on the fully coupled system with block-decomposed smoothers. This paper provides an overview of PETSc's new multiphysics capabilities, which have been used in parallel applications including lithosphere dynamics, subduction and mantle convection, ice sheet dynamics, subsurface reactive flow, fusion, mesoscale materials modeling, and power networks. Jed Brown, Matthew G. Knepley, David A. May, Lois C. McInnes, Barry Smith 0002 |
ISPDC | 2 |