VLDB 2026 Research / reviewers in the wild / expert
Martin Berzins
dblp:b/MartinBerzins
· DBLP profile ↗
25ranked-venue papers
4as first author
3since 2021 · last 2024
0000-0002-5419-0634ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 1 first-author · 1 since 2021Theory of computation · 4 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Monotone Operator Theory-Inspired Message Passing for Learning Long-Range Interaction on Graphs
Justin M. Baker, Martin Berzins, Thomas Strohmer, Bao Wang 0001 |
AISTATS | 3 |
| 2024 | Algorithm 1041: HiPPIS - A High-order Positivity-preserving Mapping Software for Structured MeshesabstractPolynomial interpolation is an important component of many computational problems. In several of these computational problems, failure to preserve positivity when using polynomials to approximate or map data values between meshes can lead to negative unphysical quantities. Currently, most polynomial-based methods for enforcing positivity are based on splines and polynomial rescaling. The spline-based approaches build interpolants that are positive over the intervals in which they are defined and may require solving a minimization problem and/or system of equations. The linear polynomial rescaling methods allow for high-degree polynomials but enforce positivity only at limited locations (e.g., quadrature nodes). This work introduces open-source software (HiPPIS) for high-order data-bounded interpolation (DBI) and positivity-preserving interpolation (PPI) that addresses the limitations of both the spline and polynomial rescaling methods. HiPPIS is suitable for approximating and mapping physical quantities such as mass, density, and concentration between meshes while preserving positivity. This work provides Fortran and Matlab implementations of the DBI and PPI methods, presents an analysis of the mapping error in the context of PDEs, and uses several 1D and 2D numerical examples to demonstrate the benefits and limitations of HiPPIS. Timbwaoga A. J. Ouermi, Robert M. Kirby, Martin Berzins |
ACM Trans. Math. Softw. | 3 |
| 2021 | Logically Parallel Communication for Fast MPI+Threads ApplicationsabstractSupercomputing applications are increasingly adopting the MPI+threads programming model over the traditional “MPI everywhere” approach to better handle the disproportionate increase in the number of cores compared with other on-node resources. In practice, however, most applications observe a slower performance with MPI+threads primarily because of poor communication performance. Recent research efforts on MPI libraries address this bottleneck by mapping logically parallel communication, that is, operations that are not subject to MPI's ordering constraints to the underlying network parallelism. Domain scientists, however, typically do not expose such communication independence information because the existing MPI-3.1 standard's semantics can be limiting. Researchers had initially proposed user-visible endpoints to combat this issue, but such a solution requires intrusive changes to the standard (new APIs). The upcoming MPI-4.0 standard, on the other hand, allows applications to relax unneeded semantics and provides them with many opportunities to express logical communication parallelism. In this article, we show how MPI+threads applications can achieve high performance with logically parallel communication. Through application case studies, we compare the capabilities of the new MPI-4.0 standard with those of the existing one and user-visible endpoints (upper bound). Logical communication parallelism can boost the overall performance of an application by over 2×. Rohit Zambre, Damodar Sahasrabudhe, Hui Zhou 0012, Martin Berzins, Aparna Chandramowlishwaran, Pavan Balaji |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2019 | An Evaluation of An Asynchronous Task Based Dataflow Approach For UintahabstractThe challenge of running complex physics code on the largest computers available has led to dataflow paradigms being explored. While such approaches are often applied at smaller scales, the challenge of extreme-scale data flow computing remains. The Uintah dataflow framework has consistently used dataflow computing at the largest scales on complex physics applications. At present Uintah contains two main dataflow models. Both are based upon asynchronous communication. One uses a static graph-based approach with asynchronous communication and the other uses a more dynamic approach that was introduced almost a decade ago. Subsequent changes within the Uintah runtime system combined with many more large scale experiments, has necessitated a reevaluation of these two approaches, comparing them in the context of large scale problems. While the static approach has worked well for some large-scale simulations, the dynamic approach is seen to offer performance improvements over the static case for a challenging fluid-structure interaction problem at large scale that involves fluid flow and a moving solid represented using particle method on an adaptive mesh. Alan Humphrey, Martin Berzins |
COMPSAC (2) | 2 |
| 2019 | Node failure resiliency for Uintah without checkpointingabstractSummary The frequency of failures in upcoming exascale supercomputers may well be greater than at present due to many‐core architectures if component failure rates remain unchanged. This potential increase in failure frequency coupled with I/O challenges at exascale may prove problematic for current resiliency approaches such as checkpoint restarting, although the use of fast intermediate memory may help. Algorithm‐based fault tolerance (ABFT) using adaptive mesh refinement (AMR) is one resiliency approach used to address these challenges. For adaptive mesh codes, a coarse mesh version of the solution may be used to restore the fine mesh solution. This paper addresses the implementation of the ABFT approach within the Uintah software framework: both at a software level within Uintah and in the data reconstruction method used for the recovery of lost data. This method has two problems: inaccuracies introduced during the reconstruction propagate forward in time, and the physical consistency of variables, such as positivity or boundedness, may be violated during interpolation. These challenges can be addressed by the combination of two techniques: (1) a fault‐tolerant message passing interface (MPI) implementation to recover from runtime node failures, and (2) high‐order interpolation schemes to preserve the physical solution and reconstruct lost data. The approach considered here uses a “limited essentially nonoscillatory” (LENO) scheme along with AMR to rebuild the lost data without checkpointing using Uintah. Experiments were carried out using a fault‐tolerant MPI‐user‐level failure mitigation to recover from runtime failure and LENO to recover data on patches belonging to failed ranks, while the simulation was continued to the end. Results show that this ABFT approach is up to 10× faster than the traditional checkpointing method. The new interpolation approach is more accurate than linear interpolation and not subject to the overshoots found in other interpolation methods. Damodar Sahasrabudhe, Martin Berzins, John A. Schmidt |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | Parallel Breadth First Search on GPU clustersabstractFast, scalable, low-cost, and low-power execution of parallel graph algorithms is important for a wide variety of commercial and public sector applications. Breadth First Search (BFS) imposes an extreme burden on memory bandwidth and network communications and has been proposed as a benchmark that may be used to evaluate current and future parallel computers. Hardware trends and manufacturing limits strongly imply that many-core devices, such as NVIDIA® GPUs and the Intel® Xeon Phi®, will become central components of such future systems. GPUs are well known to deliver the highest FLOPS/watt and enjoy a very significant memory bandwidth advantage over CPU architectures. Recent work has demonstrated that GPUs can deliver high performance for parallel graph algorithms and, further, that it is possible to encapsulate that capability in a manner that hides the low level details of the GPU architecture and the CUDA language but preserves the high throughput of the GPU. We extend previous research on GPUs and on scalable graph processing on supercomputers and demonstrate that a high-performance parallel graph machine can be created using commodity GPUs and networking hardware. Zhisong Fu, Harish Kumar Dasari, Bradley R. Bebee, Martin Berzins, Bryan Thompson 0001 |
IEEE BigData | 4 |
| 2014 | Scalable large-scale fluid-structure interaction solvers in the Uintah framework via hybrid task-based parallelism algorithmsabstractSUMMARY Uintah is a software framework that provides an environment for solving fluid–structure interaction problems on structured adaptive grids for large‐scale science and engineering problems involving the solution of partial differential equations. Uintah uses a combination of fluid flow solvers and particle‐based methods for solids, together with adaptive meshing and a novel asynchronous task‐based approach with fully automated load balancing. When applying Uintah to fluid–structure interaction problems, the combination of adaptive meshing and the movement of structures through space present a formidable challenge in terms of achieving scalability on large‐scale parallel computers. The Uintah approach to the growth of the number of core counts per socket together with the prospect of less memory per core is to adopt a model that uses MPI to communicate between nodes and a shared memory model on‐node so as to achieve scalability on large‐scale systems. For this approach to be successful, it is necessary to design data structures that large numbers of cores can simultaneously access without contention. This scalability challenge is addressed here for Uintah, by the development of new hybrid runtime and scheduling algorithms combined with novel lock‐free data structures, making it possible for Uintah to achieve excellent scalability for a challenging fluid–structure problem with mesh refinement on as many as 260K cores. Copyright © 2013 John Wiley & Sons, Ltd. Martin Berzins |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | A survey of high level frameworks in block-structured adaptive mesh refinement packages
Anshu Dubey, Ann S. Almgren, John B. Bell, Martin Berzins, Steven R. Brandt, Greg Bryan, Phillip Colella, Daniel T. Graves, Michael Lijewski, Frank Löffler 0001, Brian W. O'Shea, Erik Schnetter, Brian van Straalen, Klaus Weide |
J. Parallel Distributed Comput. | 4 |
| 2013 | Large Scale Parallel Solution of Incompressible Flow Problems Using Uintah and HypreabstractThe Uintah Software framework was developed to provide an environment for solving fluid-structure interaction problems on structured adaptive grids on large-scale, long-running, data-intensive problems. Uintah uses a combination of fluid-flow solvers and particle-based methods for solids together with a novel asynchronous task-based approach with fully automated load balancing. As Uintah is often used to solve incompressible flow problems in combustion applications it is important to have a scalable linear solver. While there are many such solvers available, the scalability of those codes varies greatly. The hypre software offers a range of solvers and pre-conditioners for different types of grids. The weak scalability of Uintah and hypre is addressed for particular examples of both packages when applied to a number of incompressible flow problems. After careful software engineering to reduce startup costs, much better than expected weak scalability is seen for up to 100K cores on NSFs Kraken architecture and up to260K cpu cores, on DOEs new Titan machine. The scalability is found to depend in a crtitical way on the choice of algorithm used by hypre for a realistic application problem. John A. Schmidt, Martin Berzins, Jeremy Thornock, Tony Saad, James C. Sutherland |
CCGRID | 2 |
| 2013 | Investigating applications portability with the Uintah DAG-based runtime system on PetaScale supercomputersabstractPresent trends in high performance computing present formidable challenges for applications code using multicore nodes possibly with accelerators and/or co-processors and reduced memory while still attaining scalability. Software frameworks that execute machine-independent applications code using a runtime system that shields users from architectural complexities offer a possible solution. The Uintah framework for example, solves a broad class of large-scale problems on structured adaptive grids using fluid-flow solvers coupled with particle-based solids methods. Uintah executes directed acyclic graphs of computational tasks with a scalable asynchronous and dynamic runtime system for CPU cores and/or accelerators/co-processors on a node. Uintah's clear separation between application and runtime code has led to scalability increases of 1000x without significant changes to application code. This methodology is tested on three leading Top500 machines; OLCF Titan, TACC Stampede and ALCF Mira using three diverse and challenging applications problems. This investigation of scalability with regard to the different processors and communications performance leads to the overall conclusion that the adaptive DAG-based approach provides a very powerful abstraction for solving challenging multi-scale multi-physics engineering problems on some of the largest and most powerful computers available today. Alan Humphrey, John A. Schmidt, Martin Berzins |
SC | 4 |
| 2011 | Introduction
Martin Berzins, Daniela di Serafino, Martin J. Gander, Luc Giraud |
Euro-Par (2) | 1 |
| 2011 | Scalable parallel regridding algorithms for block-structured adaptive mesh refinementabstractAbstract Block‐structured adaptive mesh refinement (BSAMR) is widely used within simulation software because it improves the utilization of computing resources by refining the mesh only where necessary. For BSAMR to scale onto existing petascale and eventually exascale computers all portions of the simulation need to weak scale ideally. Any portions of the simulation that do not will become a bottleneck at larger numbers of cores. The challenge is to design algorithms that will make it possible to avoid these bottlenecks on exascale computers. One step of existing BSAMR algorithms involves determining where to create new patches of refinement. The Berger–Rigoutsos algorithm is commonly used to perform this task. This paper provides a detailed analysis of the performance of two existing parallel implementations of the Berger–Rigoutsos algorithm and develops a new parallel implementation of the Berger–Rigoutsos algorithm and a tiled algorithm that exhibits ideal scalability. The analysis and computational results up to 98 304 cores are used to design performance models which are then used to predict how these algorithms will perform on 100 M cores. Copyright © 2011 John Wiley & Sons, Ltd. Justin Luitjens, Martin Berzins |
Concurr. Comput. Pract. Exp. | 2 |
| 2010 | Improving the performance of Uintah: A large-scale adaptive meshing computational frameworkabstractUintah is a highly parallel and adaptive multi-physics framework created by the Center for Simulation of Accidental Fires and Explosions in Utah. Uintah, which is built upon the Common Component Architecture, has facilitated the simulation of a wide variety of fluid-structure interaction problems using both adaptive structured meshes for the fluid and particles to model solids. Uintah was originally designed for, and has performed well on, about a thousand processors. The evolution of Uintah to use tens of thousands processors has required improvements in memory usage, data structure design, load balancing algorithms and cost estimation in order to improve strong and weak scalability up to 98,304 cores for situations in which the mesh used varies adaptively and also cases in which particles that represent the solids move from mesh cell to mesh cell. Justin Luitjens, Martin Berzins |
IPDPS | 2 |
| 2007 | Parallelization and scalability issues of a multilevel elastohydrodynamic lubrication solverabstractAbstract The computation of numerical solutions to elastohydrodynamic lubrication problems is only possible on fine meshes by using a combination of multigrid and multilevel techniques. In this paper, we show how the parallelization of both multigrid and multilevel multi‐integration for these problems may be accomplished and discuss the scalability of the resulting code. A performance model of the solver is constructed and used to perform an analysis of the results obtained. Results are shown with good speed‐ups and excellent scalability for distributed memory architectures and in agreement with the model. Copyright © 2006 John Wiley & Sons, Ltd. Christopher E. Goodyer, Martin Berzins |
Concurr. Comput. Pract. Exp. | 2 |
| 2007 | Parallelization and scalability of a spectral element channel flow solver for incompressible Navier-Stokes equationsabstractAbstract Direct numerical simulation (DNS) of turbulent flows is widely recognized to demand fine spatial meshes, small timesteps, and very long runtimes to properly resolve the flow field. To overcome these limitations, most DNS is performed on supercomputing machines. With the rapid development of terascale (and, eventually, petascale) computing on thousands of processors, it has become imperative to consider the development of DNS algorithms and parallelization methods that are capable of fully exploiting these massively parallel machines. A highly parallelizable algorithm for the simulation of turbulent channel flow that allows for efficient scaling on several thousand processors is presented. A model that accurately predicts the performance of the algorithm is developed and compared with experimental data. The results demonstrate that the proposed numerical algorithm is capable of scaling well on petascale computing machines and thus will allow for the development and analysis of high Reynolds number channel flows. Copyright © 2007 John Wiley & Sons, Ltd. C. W. Hamman, Robert M. Kirby, Martin Berzins |
Concurr. Comput. Pract. Exp. | 3 |
| 2007 | Parallel space-filling curve generation through sortingabstractAbstract In this paper we consider the scalability of parallel space‐filling curve generation as implemented through parallel sorting algorithms. Multiple sorting algorithms are studied and results show that space‐filling curves can be generated quickly in parallel on thousands of processors. In addition, performance models are presented that are consistent with measured performance and offer insight into performance on still larger numbers of processors. At large numbers of processors, the scalability of adaptive mesh refined codes depends on the individual components of the adaptive solver. One such component is the dynamic load balancer. In adaptive mesh refined codes, the mesh is constantly changing resulting in load imbalance among the processors requiring a load‐balancing phase. The load balancing may occur often, requiring the load balancer to perform quickly. One common method for dynamic load balancing is to use space‐filling curves. Space‐filling curves, in particular the Hilbert curve, generate good partitions quickly in serial. However, at tens and hundreds of thousands of processors serial generation of space‐filling curves will hinder scalability. In order to avoid this issue we have developed a method that generates space‐filling curves quickly in parallel by reducing the generation to integer sorting. Copyright © 2007 John Wiley & Sons, Ltd. Justin Luitjens, Martin Berzins, Tom Henderson |
Concurr. Comput. Pract. Exp. | 2 |
| 2001 | Review of The Computational Beauty of Nature by Gary William Flake
Martin Berzins |
Artif. Intell. | 1 |
| 2000 | A comparison of some dynamic load-balancing algorithms for a parallel adaptive flow solver
Nasir Touheed, Paul M. Selwood, Peter K. Jimack, Martin Berzins |
Parallel Comput. | 4 |
| 1999 | A Structured SADT Approach to the Support of a Parallel Adaptive 3D CFD Code
Jonathan M. Nash, Martin Berzins, Paul M. Selwood |
Euro-Par | 2 |
| 1999 | Parallel unstructured tetrahedral mesh adaptation: algorithms, implementation and scalabilityabstractThe use of unstructured adaptive tetrahedral meshes in the solution of transient flows poses a challenge for parallel computing due to the irregular and frequently changing nature of the data and its distribution. A parallel mesh adaptation algorithm, PTETRAD, for unstructured tetrahedral meshes (based on the serial code TETRAD) is described and analysed. The portable implementation of the parallel code in C with MPI is described and discussed. The scalability of the code is considered, analysed and illustrated by numerical experiments using a shock wave diffraction problem. Copyright © 1999 John Wiley & Sons, Ltd. Paul M. Selwood, Martin Berzins |
Concurr. Pract. Exp. | 2 |
| 1998 | SPRINT2D: Adaptive Software for PDEsabstractSPRINT2D is a set of software tools for solving both steady an unsteady partial differential equations in two-space variables. The software consists of a set of coupled modules for mesh generation, spatial discretization, time integration, nonlinear equations, linear algebra, spatial adaptivity, and visualization. The software uses unstructured triangular meshes and adaptive local error control in both space and time. The class of problems solved includes systems of parabolic, elliptic, and hyperbolic equations; for the latter by use of Riemann-solver-based methods. This article describes the software and shows how the adaptive techniques may be used to increase the reliability of the solution for a Burgers' equations problem, an electrostatics problem from elastohydrodynamic lubrication, and a challenging gas jet problem. Martin Berzins, R. Fairlie, S. V. Pennington, J. M. Ware, L. E. Scales |
ACM Trans. Math. Softw. | 1 |
| 1996 | Shock preserving quadratic interpolation for visualization on triangular meshes
Paul Pratt, Martin Berzins |
Comput. Graph. | 2 |
| 1995 | Dynamic load-balancing for PDE solvers on adaptive unstructured meshesabstractAbstract Modern PDE solvers written for time‐dependent problems increasingly employ adaptive unstructured meshes (Flaherty et al. , 1989) in order to both increase efficiency and control the numerical error. If a distributed memory parallel computer is to be used, there arises the significant problem of dividing the domain equally amongst the processors whilst minimising the inter‐subdomain dependencies. A number of graph‐based algorithms have recently been proposed for steady‐state calculations. The paper considers an extension to such methods which renders them more suitable for time‐dependent problems in which the mesh may be changed frequently. Chris Walshaw, Martin Berzins |
Concurr. Pract. Exp. | 2 |
| 1994 | New NAG library software for first-order partial differential equationsabstractNew NAG Fortran Library routines are described for the solution of systems of nonlinear, first-order, time-dependent partial differential equations in one space dimension, with scope for coupled ordinary differential or algebraic equations. The method-of-lines is used with spatial discretization by either the central-difference Keller box scheme or an upwind scheme for hyperbolic systems of conservation laws. The new routines have the same structure as existing library routines for the solution of second-order partial differential equations, and much of the existing library software is reused. Results are presented for several computational examples to show that the software provides physically realistic numerical solutions to a challenging class of problems. S. V. Pennington, Martin Berzins |
ACM Trans. Math. Softw. | 2 |
| 1991 | Algorithm 690: Chebyshev polynomial software for elliptic-parabolic systems of PDEsabstractPDECHEB is a FORTRAN 77 software package that semidiscretizes a wide range of time-dependent partial differential equations in one space variable. The software implements a family of spacial discretization formulas, based on piecewise Chebyshev polynomial expansions with C 0 continuity. The package has been designed to be used in conjunction with a general integrator for initial value problems to provide a powerful software tool for the solution of parabolic-elliptic PDEs with coupled differential algebraic equations. Examples are provided to illustrate the use of the package with the DASSL d.a.e integrator of Petzold [18]. Martin Berzins, Peter M. Dew |
ACM Trans. Math. Softw. | 1 |