Marsha J. Berger

dblp:10/2868 · also Marsha Berger · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
2since 2021 · last 2025
0000-0002-7584-895XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Towards Seamless Interoperability of MPI-OpenMP Applications
abstract
A chasm exists between mathematical software libraries written for MPI-based applications and those written for OpenMP applications. Recently, however, PETSc enables the simple use of its MPI-based linear solvers from OpenMP applications. Separately, the MPICH MPI development team has started a new project to allow almost seamless MPI use in OpenMP applications. Both proposed approaches would result in a similar user experience. We discuss the reasons for these projects and their potential for providing more numerical library choices for OpenMP applications, including the unlimited assortment of linear solvers available in PETSc. In addition, we present the performance of an application using the first approach, demonstrating its efficacy.
Barry Smith 0002, Marsha J. Berger, Junchao Zhang 0002, Hui Zhou 0012
ACM Trans. Math. Softw.2
2024 The Well: a Large-Scale Collection of Diverse Physics Simulations for Machine Learning
abstract
Machine learning based surrogate models offer researchers powerful tools for accelerating simulation-based workflows. However, as standard datasets in this space often cover small classes of physical behavior, it can be difficult to evaluate the efficacy of new approaches. To address this gap, we introduce the Well: a large-scale collection of datasets containing numerical simulations of a wide variety of spatiotemporal physical systems. The Well draws from domain experts and numerical software developers to provide 15TB of data across 16 datasets covering diverse domains such as biological systems, fluid dynamics, acoustic scattering, as well as magneto-hydrodynamic simulations of extra-galactic fluids or supernova explosions. These datasets can be used individually or as part of a broader benchmark suite. To facilitate usage of the Well, we provide a unified PyTorch interface for training and evaluating models. We demonstrate the function of this library by introducing example baselines that highlight the new challenges posed by the complex dynamics of the Well. The code and data is available at https://github.com/PolymathicAI/the_well.
Ruben Ohana, Michael McCabe, Lucas Meyer, Rudy Morel, Fruzsina Julia Agocs, Miguel Beneitez, Marsha J. Berger, Blakesley Burkhart, Stuart B. Dalziel, Drummond B. Fielding, Daniel Fortunato, Jared A. Goldberg, Keiya Hirashima, Yan-Fei Jiang, Rich R. Kerswell, Suryanarayana Maddu, Jonah Miller, Payel Mukhopadhyay, Stefan S. Nixon, Jeff Shen, Romain Watteaux, Bruno Régaldo-Saint Blancard, François Rozet, Liam Holden Parker, Miles D. Cranmer, Shirley Ho
NeurIPS7
2017 Toucan - A Translator for Communication Tolerant MPI Applications
abstract
We discuss early results with Toucan, a source-to-source translator that automatically restructures C/C++ MPI applications tooverlap communication with computation. We co-designed the translator and runtime system to enable dynamic, dependence-driven execution of MPI applications, and require only a modest amount of programmer annotation. Co-design was essential to realizing overlap through dynamic code block reordering and avoiding the limitations of static code relocation and inlining. We demonstrate that Toucan hides significant communication in four representative applications running on up to 24Kcores of NERSC's Edison platform. Using Toucan, we have hidden from 33% to 85% of the communication overhead, with performance meeting or exceeding that of painstakingly hand-written overlap variants.
Sergio M. Martin, Marsha J. Berger, Scott B. Baden
IPDPS2
2005 High Resolution Aerospace Applications using the NASA Columbia Supercomputer
abstract
This paper focuses on the parallel performance of two high-performance aerodynamic simulation packages on the newly installed NASA Columbia supercomputer. These packages include both a high-fidelity, unstructured, Reynolds-averaged Navier-Stokes solver, and a fully-automated inviscid flow package for cut-cell Cartesian grids. The complementary combination of these two simulation codes enables high-fidelity characterization of aerospace vehicle design performance over the entire flight envelope through extensive parametric analysis and detailed simulation of critical regions of the flight envelope. Both packages are industrial-level codes designed for complex geometry and incorporate customized multigrid solution algorithms. The performance of these codes on Columbia is examined using both MPI and OpenMP and using both the NUMAlink and InfiniBand interconnect fabrics. Numerical results demonstrate good scalability on up to 2016 cpus using the NUMAlink4 interconnect, with measured computational rates in the vicinity of 3 TFLOP/s, while InfiniBand showed some performance degradation at high CPU counts, particularly with multigrid. Nonetheless, the results are encouraging enough to indicate that larger test cases using combined MPI/OpenMP communication should scale well on even more processors.
Dimitri J. Mavriplis, Michael J. Aftosmis, Marsha J. Berger
SC3
2005 Performance of a new CFD flow solver using a hybrid programming paradigm
Marsha J. Berger, Michael J. Aftosmis, D. D. Marshall, Scott M. Murman
J. Parallel Distributed Comput.1
1991 An algorithm for point clustering and grid generation
abstract
A special-purpose point clustering algorithm is described, and its application to automatic grid generation, a technique used to solve partial differential equations, is considered. Extensions of techniques common in computer vision and pattern recognition literature are used to partition points into a set of enclosing rectangles. Examples from 2-D calculations are shown, but the algorithm generalizes readily to three dimensions.>
Marsha J. Berger, Isidore Rigoutsos
IEEE Trans. Syst. Man Cybern.1
1987 A Partitioning Strategy for Nonuniform Problems on Multiprocessors
abstract
We consider the partitioning of a problem on a domain with unequal work estimates in different subdomains in a way that balances the workload across multiple processors. Such a problem arises for example in solving partial differential equations using an adaptive method that places extra grid points in certain subregions of the domain. We use a binary decomposition of the domain to partition it into rectangles requiring equal computational effort. We then study the communication costs of mapping this partitioning onto different multiprocessors: a mesh- connected array, a tree machine, and a hypercube. The communication cost expressions can be used to determine the optimal depth of the above partitioning.
Marsha J. Berger, Shahid H. Bokhari
IEEE Trans. Computers1
1985 A Partitioning Strategy for PDEs Across Multiprocessors
Marsha J. Berger, Shahid H. Bokhari
ICPP1