VLDB 2026 Research / reviewers in the wild / expert
Marsha J. Berger
dblp:10/2868 · also Marsha Berger
· DBLP profile ↗
8ranked-venue papers
4as first author
2since 2021 · last 2025
0000-0002-7584-895XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Seamless Interoperability of MPI-OpenMP ApplicationsabstractA chasm exists between mathematical software libraries written for MPI-based applications and those written for OpenMP applications. Recently, however, PETSc enables the simple use of its MPI-based linear solvers from OpenMP applications. Separately, the MPICH MPI development team has started a new project to allow almost seamless MPI use in OpenMP applications. Both proposed approaches would result in a similar user experience. We discuss the reasons for these projects and their potential for providing more numerical library choices for OpenMP applications, including the unlimited assortment of linear solvers available in PETSc. In addition, we present the performance of an application using the first approach, demonstrating its efficacy. Barry Smith 0002, Marsha J. Berger, Junchao Zhang 0002, Hui Zhou 0012 |
ACM Trans. Math. Softw. | 2 |
| 2024 | The Well: a Large-Scale Collection of Diverse Physics Simulations for Machine LearningabstractMachine learning based surrogate models offer researchers powerful tools for accelerating simulation-based workflows. However, as standard datasets in this space often cover small classes of physical behavior, it can be difficult to evaluate the efficacy of new approaches. To address this gap, we introduce the Well: a large-scale collection of datasets containing numerical simulations of a wide variety of spatiotemporal physical systems. The Well draws from domain experts and numerical software developers to provide 15TB of data across 16 datasets covering diverse domains such as biological systems, fluid dynamics, acoustic scattering, as well as magneto-hydrodynamic simulations of extra-galactic fluids or supernova explosions. These datasets can be used individually or as part of a broader benchmark suite. To facilitate usage of the Well, we provide a unified PyTorch interface for training and evaluating models. We demonstrate the function of this library by introducing example baselines that highlight the new challenges posed by the complex dynamics of the Well. The code and data is available at https://github.com/PolymathicAI/the_well. Ruben Ohana, Michael McCabe, Lucas Meyer, Rudy Morel, Fruzsina Julia Agocs, Miguel Beneitez, Marsha J. Berger, Blakesley Burkhart, Stuart B. Dalziel, Drummond B. Fielding, Daniel Fortunato, Jared A. Goldberg, Keiya Hirashima, Yan-Fei Jiang, Rich R. Kerswell, Suryanarayana Maddu, Jonah Miller, Payel Mukhopadhyay, Stefan S. Nixon, Jeff Shen, Romain Watteaux, Bruno Régaldo-Saint Blancard, François Rozet, Liam Holden Parker, Miles D. Cranmer, Shirley Ho |
NeurIPS | 7 |
| 2017 | Toucan - A Translator for Communication Tolerant MPI ApplicationsabstractWe discuss early results with Toucan, a source-to-source translator that automatically restructures C/C++ MPI applications tooverlap communication with computation. We co-designed the translator and runtime system to enable dynamic, dependence-driven execution of MPI applications, and require only a modest amount of programmer annotation. Co-design was essential to realizing overlap through dynamic code block reordering and avoiding the limitations of static code relocation and inlining. We demonstrate that Toucan hides significant communication in four representative applications running on up to 24Kcores of NERSC's Edison platform. Using Toucan, we have hidden from 33% to 85% of the communication overhead, with performance meeting or exceeding that of painstakingly hand-written overlap variants. Sergio M. Martin, Marsha J. Berger, Scott B. Baden |
IPDPS | 2 |
| 2005 | High Resolution Aerospace Applications using the NASA Columbia SupercomputerabstractThis paper focuses on the parallel performance of two high-performance aerodynamic simulation packages on the newly installed NASA Columbia supercomputer. These packages include both a high-fidelity, unstructured, Reynolds-averaged Navier-Stokes solver, and a fully-automated inviscid flow package for cut-cell Cartesian grids. The complementary combination of these two simulation codes enables high-fidelity characterization of aerospace vehicle design performance over the entire flight envelope through extensive parametric analysis and detailed simulation of critical regions of the flight envelope. Both packages are industrial-level codes designed for complex geometry and incorporate customized multigrid solution algorithms. The performance of these codes on Columbia is examined using both MPI and OpenMP and using both the NUMAlink and InfiniBand interconnect fabrics. Numerical results demonstrate good scalability on up to 2016 cpus using the NUMAlink4 interconnect, with measured computational rates in the vicinity of 3 TFLOP/s, while InfiniBand showed some performance degradation at high CPU counts, particularly with multigrid. Nonetheless, the results are encouraging enough to indicate that larger test cases using combined MPI/OpenMP communication should scale well on even more processors. Dimitri J. Mavriplis, Michael J. Aftosmis, Marsha J. Berger |
SC | 3 |
| 2005 | Performance of a new CFD flow solver using a hybrid programming paradigm
Marsha J. Berger, Michael J. Aftosmis, D. D. Marshall, Scott M. Murman |
J. Parallel Distributed Comput. | 1 |
| 1991 | An algorithm for point clustering and grid generationabstractA special-purpose point clustering algorithm is described, and its application to automatic grid generation, a technique used to solve partial differential equations, is considered. Extensions of techniques common in computer vision and pattern recognition literature are used to partition points into a set of enclosing rectangles. Examples from 2-D calculations are shown, but the algorithm generalizes readily to three dimensions.> Marsha J. Berger, Isidore Rigoutsos |
IEEE Trans. Syst. Man Cybern. | 1 |
| 1987 | A Partitioning Strategy for Nonuniform Problems on MultiprocessorsabstractWe consider the partitioning of a problem on a domain with unequal work estimates in different subdomains in a way that balances the workload across multiple processors. Such a problem arises for example in solving partial differential equations using an adaptive method that places extra grid points in certain subregions of the domain. We use a binary decomposition of the domain to partition it into rectangles requiring equal computational effort. We then study the communication costs of mapping this partitioning onto different multiprocessors: a mesh- connected array, a tree machine, and a hypercube. The communication cost expressions can be used to determine the optimal depth of the above partitioning. Marsha J. Berger, Shahid H. Bokhari |
IEEE Trans. Computers | 1 |
| 1985 | A Partitioning Strategy for PDEs Across Multiprocessors
Marsha J. Berger, Shahid H. Bokhari |
ICPP | 1 |