Paul F. Fischer

dblp:71/2990 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
4since 2021 · last 2023
0000-0002-6506-4502ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 1 first-author · 4 since 2021Theory of computation · 2 · 2 first-author
YearPublicationVenuePosition
2023 Exascale Multiphysics Nuclear Reactor Simulations for Advanced Designs
abstract
ENRICO is a coupled application developed under the U.S. Department of Energy's Exascale Computing Project (ECP) targeting the modeling of advanced nuclear reactors. It couples radiation transport with heat and fluid simulation, including the high-fidelity, highresolution Monte-Carlo code Shift and the Computational fluid dynamics code NekRS. NekRS is a highly-performant open-source code for simulation of incompressible and low-Mach fluid flow, heat transfer, and combustion with a particular focus on turbulent flows in complex domains. It is based on rapidly convergent high-order spectral element discretizations that feature minimal numerical dissipation and dispersion. State-of-the-art multilevel preconditioners, efficient high-order time-splitting methods, and runtime-adaptive communication strategies are built on a fast OCCA-based kernel library, libParanumal, to provide scalability and portability across the spectrum of current and future high-performance computing platforms. On Frontier, Nek5000/RS has recently achieved an unprecedented milestone in breaching over 1 billion spectral elements and 350 billion degrees of freedom. Shift has demonstrated the capability to transport upwards of 1 billion particles per second in full core nuclear reactor simulations featuring complete temperature-dependent, continuous-energy physics on Frontier. Shift achieved a weak-scaling efficiency of 97.8% on 8192 nodes of Frontier and calculated 6 reactions in 214,896 fuel pin regions below 1% statistical error yielding first-of-a-kind resolution for a Monte Carlo transport application.
Elia Merzari, Steven P. Hamilton, Thomas M. Evans 0001, Misun Min, Paul F. Fischer, Stefan Kerkemeier, Jun Fang 0005, Paul K. Romano, Yu-Hsiang Lan, Malachi Phillips, Elliott Biondo, Katherine Royston, Timothy C. Warburton, Noel Chalmers, Thilina Ratnayaka
SC5
2022 Optimization of Full-Core Reactor Simulations on Summit
abstract
Nek5000/RS, a highly-performant open-source spectral element code, has recently achieved an unprecedented milestone in the simulation of nuclear reactors: the first full core computational fluid dynamics simulations of reactor cores, including pebble beds with 352,625 pebbles and 98M spectral elements (51 billion gridpoints), advanced in less than 0.25 seconds per Navier-Stokes timestep. The authors present performance and optimization considerations necessary to achieve this milestone when running on all of Summit. These optimizations led to a fourfold reduction in time-to-solution, making it possible to perform high-fidelity simulations of a single flow-through time in less than six hours for a full reactor core under prototypical conditions.
Misun Min, Yu-Hsiang Lan, Paul F. Fischer, Elia Merzari, Stefan Kerkemeier, Malachi Phillips, Thilina Ratnayaka, April Novak, Derek Gaston, Noel Chalmers, Timothy C. Warburton
SC3
2022 NekRS, a GPU-accelerated spectral element Navier-Stokes solver
Paul F. Fischer, Stefan Kerkemeier, Misun Min, Yu-Hsiang Lan, Malachi Phillips, Thilina Ratnayaka, Elia Merzari, Ananias Tomboulides, Ali Karakus, Noel Chalmers, Timothy C. Warburton
Parallel Comput.1
2021 GPU algorithms for Efficient Exascale Discretizations
Ahmad Abdelfattah, Valeria Barra, Natalie N. Beams, Ryan Bleile, Jed Brown, Sylvain Camier, Robert Carson, Noel Chalmers, Veselin Dobrev, Yohann Dudouit, Paul F. Fischer, Ali Karakus, Stefan Kerkemeier, Tzanio V. Kolev, Yu-Hsiang Lan, Elia Merzari, Misun Min, Malachi Phillips, Thilina Ratnayaka, Robert N. Rieben, Thomas Stitt, Ananias Tomboulides, Stanimire Tomov, Vladimir Z. Tomov, Arturo Vargas, Timothy C. Warburton, Kenneth Weiss 0001
Parallel Comput.11
2019 OpenACC acceleration for the PN-PN-2 algorithm in Nek5000
Evelyn Otero, Misun Min, Paul F. Fischer, Philipp Schlatter, Erwin Laure
J. Parallel Distributed Comput.4
2017 Why is MPI so slow?: analyzing the fundamental limits in implementing MPI-3.1
abstract
This paper provides an in-depth analysis of the software overheads in the MPI performance-critical path and exposes mandatory performance overheads that are unavoidable based on the MPI-3.1 specification. We first present a highly optimized implementation of the MPI-3.1 standard in which the communication stack---all the way from the application to the low-level network communication API---takes only a few tens of instructions. We carefully study these instructions and analyze the root cause of the overheads based on specific requirements from the MPI standard that are unavoidable under the current MPI standard. We recommend potential changes to the MPI standard that can minimize these overheads. Our experimental results on a variety of network architectures and applications demonstrate significant benefits from our proposed changes.
Kenneth Raffenetti, Abdelhalim Amer, Lena Oden, Charles Archer, Wesley Bland, Hajime Fujita 0002, Yanfei Guo, Tomislav Janjusic, Dmitry Durnov, Michael Blocksome, Min Si, Akhil Langer, Gengbin Zheng, Masamichi Takagi, Paul K. Coffman, Sayantan Sur, Alexander Sannikov, Sergey Oblomov, Michael Chuvelev, Masayuki Hatanaka, Paul F. Fischer, Thilina Ratnayaka, Matthew Otten, Misun Min, Pavan Balaji
SC24
2016 Nekbone performance on GPUs with OpenACC and CUDA Fortran implementations
Stefano Markidis, Erwin Laure, Matthew Otten, Paul F. Fischer, Misun Min
J. Supercomput.5
2015 Evaluation of Parallel Communication Models in Nekbone, a Nek5000 Mini-Application
abstract
Nekbone is a proxy application of Nek5000, a scalable Computational Fluid Dynamics (CFD) code used for modelling incompressible flows. The Nekbone mini-application is used by several international co-design centers to explore new concepts in computer science and to evaluate their performance. We present the design and implementation of a new communication kernel in the Nekbone mini-application with the goal of studying the performance of different parallel communication models. First, a new MPI blocking communication kernel has been developed to solve Nekbone problems in a three-dimensional Cartesian mesh and process topology. The new MPI implementation delivers a 13% performance improvement compared to the original implementation. The new MPI communication kernel consists of approximately 500 lines of code against the original 7,000 lines of code, allowing experimentation with new approaches in Nekbone parallel communication. Second, the MPI blocking communication in the new kernel was changed to the MPI non-blocking communication. Third, we developed a new Partitioned Global Address Space (PGAS) communication kernel, based on the GPI-2 library. This approach reduces the synchronization among neighbor processes and is on average 3% faster than the new MPI-based, non-blocking, approach. In our tests on 8,192 processes, the GPI-2 communication kernel is 3% faster than the new MPI non-blocking communication kernel. In addition, we have used the OpenMP in all the versions of the new communication kernel. Finally, we highlight the future steps for using the new communication kernel in the parent application Nek5000.
Ilya Ivanov, Dana Akhmetova, Ivy Bo Peng, Stefano Markidis, Erwin Laure, Mirko Rahn, Valeria Bartsch, Alistair Hart, Paul F. Fischer
CLUSTER11
2010 Speeding up Nek5000 with autotuning and specialization
abstract
Autotuning technology has emerged recently as a systematic process for evaluating alternative implementations of a computation, in order to select the best-performing solution for a particular architecture. Specialization optimizes code customized to a particular class of input data set. In this paper, we demonstrate how compiler-based autotuning that incorporates specialization for expected data set sizes of key computations can be used to speed up Nek5000, a spectral-element code. Nek5000 makes heavy use of what are effectively Basic Linear Algebra Subroutine (BLAS) calls, but for very small matrices. Through autotuning and specialization, we can achieve significant performance gains over hand-tuned libraries (e.g., Goto, ATLAS, and ACML BLAS). Additional performance gains are obtained from using higher-level compiler optimizations that aggregate multiple BLAS calls. We demonstrate more than 2.2X performance gains on an Opteron over the original manually tuned implementation, and speedups of up to 1.26X on the entire application running on 256 nodes of the Cray XT5 Jaguar system at Oak Ridge.
Mary W. Hall, Jacqueline Chame, Chun Chen 0002, Paul F. Fischer, Paul D. Hovland
ICS5
2001 Fast Parallel Direct Solvers for Coarse Grid Problems
Henry M. Tufo, Paul F. Fischer
J. Parallel Distributed Comput.2
2001 Generalized scans and tridiagonal systems
Paul F. Fischer, Franco P. Preparata, John E. Savage
Theor. Comput. Sci.1
1999 Terascale Spectral Element Algorithms and Implementations
abstract
We describe the development and implementation of an efficient spectral element code for multimillion gridpoint simulations of incompressible flows in general two- and three-dimensional domains. We review basic and recently developed algorithmic underpinnings that have resulted in good parallel and vector performance on a broad range of architectures, including the terascale computing systems now coming online at the DOE labs. Sustained performance of 219 GFLOPS has been recently achieved on 2048 nodes of the Intel ASCI-Red machine at Sandia.
Henry M. Tufo, Paul F. Fischer
SC2
1999 Numerical Simulation and Immersive Visualization of Hairpin Vortices
abstract
To better understand the vortex dynamics of coherent structures in turbulent and transitional boundary layers, we consider direct numerical simulation of the interaction between a flat-plate-boundary-layer flow and an isolated hemispherical roughness element. Of principal interest is the evolution of hairpin vortices that form an interlacing pattern in the wake of the hemisphere, lift away from the wall, and are stretched by the shearing action of the boundary layer. Using animations of unsteady three-dimensional representations of this flow, produced by the vtk toolkit and enhanced to operate in a CAVE virtual environment, we identify and study several key features in the evolution of this complex vortex topology not previously observed in other visualization formats. 1 Introduction We present visualization results of numerically generated hairpin vortex formation and evolution in incompressible boundary layer flows. At moderate flow speeds, hairpin vortices provide an exampl...
Henry M. Tufo, Paul F. Fischer, Michael E. Papka, Kristopher J. Blom
SC2
1995 Generalized Scans and Tri-Diagonal Systems
Paul F. Fischer, Franco P. Preparata, John E. Savage
STACS1
1991 Numerical simulation of incompressible fluid flows
abstract
Abstract In this paper we discuss temporal, spatial, and architectural aspects of the numerical simulation of time‐dependent incompressible fluid flows. In particular, we consider: high‐order operator‐integration‐factor time‐splitting methods; sliding‐mesh spectral mortar‐element spatial discretizations; and data‐parallel distributed‐memory medium‐grained parallel solution techniques. Numerous flow examples are presented.
George Anagnostou, Yvon Maday, Anthony T. Patera, Paul F. Fischer, Einar M. Rønquist
Concurr. Pract. Exp.4