VLDB 2026 Research / reviewers in the wild / expert
Barry Smith 0002
dblp:s/BarryFSmith · also Barry F. Smith
· DBLP profile ↗
19ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0001-5955-8111ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 2 since 2021Theory of computation · 4 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Seamless Interoperability of MPI-OpenMP ApplicationsabstractA chasm exists between mathematical software libraries written for MPI-based applications and those written for OpenMP applications. Recently, however, PETSc enables the simple use of its MPI-based linear solvers from OpenMP applications. Separately, the MPICH MPI development team has started a new project to allow almost seamless MPI use in OpenMP applications. Both proposed approaches would result in a similar user experience. We discuss the reasons for these projects and their potential for providing more numerical library choices for OpenMP applications, including the unlimited assortment of linear solvers available in PETSc. In addition, we present the performance of an application using the first approach, demonstrating its efficacy. Barry Smith 0002, Marsha J. Berger, Junchao Zhang 0002, Hui Zhou 0012 |
ACM Trans. Math. Softw. | 1 |
| 2022 | The PetscSF Scalable Communication LayerabstractPetscSF, the communication component of the Portable, Extensible Toolkit for Scientific Computation (PETSc), is designed to provide PETSc's communication infrastructure suitable for exascale computers that utilize GPUs and other accelerators. PetscSF provides a simple application programming interface (API) for managing common communication patterns in scientific computations by using a star-forest graph representation. PetscSF supports several implementations based on MPI and NVSHMEM, whose selection is based on the characteristics of the application or the target architecture. An efficient and portable model for network and intra-node communication is essential for implementing large-scale applications. The Message Passing Interface, which has been the de facto standard for distributed memory systems, has developed into a large complex API that does not yet provide high performance on the emerging heterogeneous CPU-GPU-based exascale systems. In this article, we discuss the design of PetscSF, how it can overcome some difficulties of working directly with MPI on GPUs, and we demonstrate its performance, scalability, and novel features. Junchao Zhang 0002, Jed Brown, Satish Balay, Jacob Faibussowitsch, Matthew G. Knepley, Oana Marin, Richard Tran Mills, Todd S. Munson, Barry Smith 0002, Stefano Zampini |
IEEE Trans. Parallel Distributed Syst. | 9 |
| 2021 | Toward performance-portable PETSc for GPU-based exascale systems
Richard Tran Mills, Mark F. Adams, Satish Balay, Jed Brown, Alp Dener, Matthew G. Knepley, Scott Kruger, Hannah Morgan, Todd S. Munson, Karl Rupp, Barry Smith 0002, Stefano Zampini, Hong Zhang 0006, Junchao Zhang 0002 |
Parallel Comput. | 11 |
| 2020 | PETSc DMNetwork: A Library for Scalable Network PDE-Based Multiphysics SimulationsabstractWe present DMNetwork, a high-level package included in the PETSc library for the simulation of multiphysics phenomena over large-scale networked systems. The library aims at applications that have networked structures such as those in electrical, gas, and water distribution systems. DMNetwork provides data and topology management, parallelization for multiphysics systems over a network, and hierarchical and composable solvers to exploit the problem structure. DMNetwork eases the simulation development cycle by providing the necessary infrastructure through simple abstractions to define and query the network components. This article presents the design of DMNetwork, illustrates its user interface, and demonstrates its ability to solve multiphysics systems, such as an electric circuit, a network of power grid and water subnetworks, and transient hydraulic systems over large networks with more than 2 billion variables on extreme-scale computers using up to 30,000 processors. Shrirang Abhyankar, Getnet D. Betrie, Daniel A. Maldonado, Lois C. McInnes, Barry Smith 0002, Hong Zhang 0006 |
ACM Trans. Math. Softw. | 5 |
| 2018 | Vectorized Parallel Sparse Matrix-Vector Multiplication in PETSc Using AVX-512abstractEmerging many-core CPU architectures with high degrees of single-instruction, multiple data (SIMD) parallelism promise to enable increasingly ambitious simulations based on partial differential equations (PDEs) via extreme-scale computing. However, such architectures present several challenges to their efficient use. Here, we explore the efficient implementation of sparse matrix-vector (SpMV) multiplications---a critical kernel for the workhorse iterative linear solvers used in most PDE-based simulations---on recent CPU architectures from Intel as well as the second-generation Knights Landing Intel Xeon Phi, which features many CPU cores, wide SIMD lanes, and on-package high-bandwidth memory. Traditional SpMV algorithms use compressed sparse row storage format, which is a hindrance to exploiting wide SIMD lanes. We study alternative matrix formats and present an efficient optimized SpMV kernel, based on a sliced ELLPACK representation, which we have implemented in the PETSc library. In addition, we demonstrate the benefit of using this representation to accelerate preconditioned iterative solvers in realistic PDE-based simulations in parallel. Hong Zhang 0006, Richard Tran Mills, Karl Rupp, Barry Smith 0002 |
ICPP | 4 |
| 2014 | Hierarchical Krylov and nested Krylov methods for extreme-scale computing
Lois C. McInnes, Barry Smith 0002, Hong Zhang 0006, Richard Tran Mills |
Parallel Comput. | 2 |
| 2012 | Composable Linear Solvers for MultiphysicsabstractThe Portable, Extensible Toolkit for Scientific computing (PETSc), which focuses on the scalable solution of problems based on partial differential equations, now incorporates new components that allow full compos ability of solvers for multiphysics and multilevel methods. Through strong encapsulation, we achieve arbitrary, dynamic composition of hierarchical methods for coupled problems and allow customization of all components in composite solvers. For example, we support block decompositions with nested multigrid as well as multigrid on the fully coupled system with block-decomposed smoothers. This paper provides an overview of PETSc's new multiphysics capabilities, which have been used in parallel applications including lithosphere dynamics, subduction and mantle convection, ice sheet dynamics, subsurface reactive flow, fusion, mesoscale materials modeling, and power networks. Jed Brown, Matthew G. Knepley, David A. May, Lois C. McInnes, Barry Smith 0002 |
ISPDC | 5 |
| 2009 | Enabling high-fidelity neutron transport simulations on petascale architecturesabstractThe UNIC code is being developed as part of the DOE's Nuclear Energy Advanced Modeling and Simulation (NEAMS) program. UNIC is an unstructured, deterministic neutron transport code that allows a highly detailed description of a nuclear reactor. The primary goal of our simulation efforts is to reduce the uncertainties and biases in reactor design calculations by progressively replacing existing multilevel averaging (homogenization) techniques with more direct solution methods based on first principles. Since the neutron transport equation is seven dimensional (three in space, two in angle, one in energy, and one in time), these simulations are among the most memory and computationally intensive in all of computational science. In order to model the complex physics of a reactor core, billions of spatial elements, hundreds of angles, and thousands of energy groups are necessary, leading to problem sizes with petascale degrees of freedom. Therefore, these calculations exhaust memory resources on current and even next-generation architectures. In this paper, we present UNIC simulation results for two important representative problems in reactor design and analysis---PHENIX and ZPR-6. In each case, UNIC shows good weak scalability on up to 163,840 cores of Blue Gene/P (Argonne) and 122,800 cores of XT5 (Oak Ridge). While our current per processor performance is less than ideal, we demonstrate a clear ability to effectively utilize the leadership computing platforms. Over the coming months, we aim to improve the per processor performance while maintaining the high parallel efficiency by employing better algorithms such as spatial p- and h-multigrid preconditioners, optimized matrix-tensor operations, and weighted partitioning for better load balancing. Combining these additional algorithmic improvements with the availability of larger parallel machines should allow us to realize our long-term goal of explicit geometry coupled multiphysics reactor simulations. In the long run, these high-fidelity simulations will be able to replace expensive mockup experiments and reduce the uncertainty in crucial reactor design and operational parameters. Dinesh K. Kaushik, Micheal Smith, Allan B. Wollaber, Barry Smith 0002, Andrew R. Siegel, Won Sik Yang |
SC | 4 |
| 2008 | Improving the Performance of Tensor Matrix Vector Multiplication in Cumulative Reaction Probability Based Quantum Chemistry Codes
Dinesh K. Kaushik, William Gropp, Michael Minkoff, Barry Smith 0002 |
HiPC | 4 |
| 2007 | Nonuniformly Communicating Noncontiguous Data: A Case Study with PETSc and MPIabstractDue to the complexity associated with developing parallel applications, scientists and engineers rely on high-level software libraries such as PETSc, ScaLAPACK and PESSL to ease this task. Such libraries assist developers by providing abstractions for mathematical operations, data representation and management of parallel layouts of the data, while internally using communication libraries such as MPI and PVM. With high-level libraries managing data layout and communication internally, it can be expected that they organize application data suitably for performing the library operations optimally. However, this places additional overhead on the underlying communication library by making the data layout noncontiguous in memory and communication volumes (data transferred by a process to each of the other processes) nonuniform. In this paper, we analyze the overheads associated with these two aspects (noncontiguous data layouts and nonuniform communication volumes) in the context of the PETSc software toolkit over the MPI communication library. We describe the issues with the current approaches used by MPICH2 (an implementation of MPI), propose different approaches to handle these issues and evaluate these approaches with micro-benchmarks as well as an application over the PETSc software library. Our experimental results demonstrate close to an order of magnitude improvement in the performance of a 3-D Laplacian multi-grid solver application when evaluated on a 128 processor cluster. Pavan Balaji, Darius Buntinas, Satish Balay, Barry Smith 0002, Rajeev Thakur, William Gropp |
IPDPS | 4 |
| 2007 | SIPs: Shift-and-invert parallel spectral transformationsabstractSIPs is a new efficient and robust software package implementing multiple shift-and-invert spectral transformations on parallel computers. Built on top of SLEPc and PETSc, it can compute very large numbers of eigenpairs for sparse symmetric generalized eigenvalue problems. The development of SIPs is motivated by applications in nanoscale materials modeling, in which the growing size of the matrices and the pathological eigenvalue distribution challenge the efficiency and robustness of the solver. In this article, we present a parallel eigenvalue algorithm based on distributed spectrum slicing. We describe the object-oriented design and implementation techniques in SIPs, and demonstrate its numerical performance on an advanced distributed computer. Hong Zhang 0006, Barry Smith 0002, Michael G. Sternberg, Peter Zapol |
ACM Trans. Math. Softw. | 2 |
| 2005 | Making automatic differentiation truly automatic: coupling PETSc with ADIC
Paul D. Hovland, Boyana Norris, Barry Smith 0002 |
Future Gener. Comput. Syst. | 3 |
| 2002 | Parallel components for PDEs and optimization: some issues and experiences
Boyana Norris, Satish Balay, Steven Benson, Lori A. Diachin, Paul D. Hovland, Lois C. McInnes, Barry Smith 0002 |
Parallel Comput. | 7 |
| 2001 | High-performance parallel implicit CFD
William Gropp, Dinesh K. Kaushik, David E. Keyes, Barry Smith 0002 |
Parallel Comput. | 4 |
| 2000 | Analyzing the Parallel Scalability of an Implicit Unstructured Mesh CFD Code
William Gropp, Dinesh K. Kaushik, Barry Smith 0002, David E. Keyes |
HiPC | 3 |
| 2000 | Performance Modeling and Tuning of an Unstructured Mesh CFD ApplicationabstractThis paper describes performance tuning experiences with a three-dimensional unstructured grid Euler flow code from NASA, which we have reimplemented in the PETSc framework and ported to several large-scale machines, including the ASCI Red and Blue Pacific machines, the SGI Origin, the Cray T3E, and Beowulf clusters. The code achieves a respectable level of performance for sparse problems, typical of scientific and engineering codes based on partial differential equations, and scales well up to thousands of processors. Since the gap between CPU speed and memory access rate is widening, the code is analyzed from a memory-centric perspective (in contrast to traditional flop-orientation) to understand its sequential and parallel performance. Performance tuning is approached on three fronts: data layouts to enhance locality of reference, algorithmic parameters, and parallel programming model. This effort was guided partly by some simple performance models developed for the sparse matrix-vector product operation. William Gropp, Dinesh K. Kaushik, David E. Keyes, Barry Smith 0002 |
SC | 4 |
| 1999 | Achieving High Sustained Performance in an Unstructured Mesh CFD ApplicationabstractThis paper highlights a three-year project by an interdisciplinary team on a legacy F77 computational fluid dynamics code, with the aim of demonstrating that implicit unstructured grid simulations can execute at rates not far from those of explicit structured grid codes, provided attention is paid to data motion complexity and the reuse of data positioned at the levels of the memory hierarchy closest to the processor, in addition to traditional operation count complexity. The demonstration code is from NASA and the enabling parallel hardware and (freely available) software toolkit are from DOE, but the resulting methodology should be broadly applicable, and the hardware limitations exposed should allow programmers and vendors of parallel platforms to focus with greater encouragement on sparse codes with indirect addressing. This snapshot of ongoing work shows a performance of 15 microseconds per degree of freedom to steady-state convergence of Euler flow on a mesh with 2.8 million vertices using 3072 dual-processor nodes of ASCI Red, corresponding to a sustained floating-point rate of 0.227 Tflop/s. W. K. Anderson, William Gropp, Dinesh K. Kaushik, David E. Keyes, Barry Smith 0002 |
SC | 5 |
| 1984 | The Linked Inference Principle, II: The User's Viewpoint
Larry Wos, Robert Veroff, Barry Smith 0002, William McCune |
CADE | 3 |
| 1984 | A New Use of an Automated Reasoning Assistant: Open Questions in Equivalential Calculus and the Study of Infinite Domains
Larry Wos, Steven K. Winker, Barry Smith 0002, Robert Veroff, Lawrence J. Henschen |
Artif. Intell. | 3 |