VLDB 2026 Research / reviewers in the wild / expert
Ann S. Almgren
dblp:64/7586
· DBLP profile ↗
9ranked-venue papers
0as first author
1since 2021 · last 2021
0000-0003-2103-312XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
High-performance computing · 85% Parallel and multicore computing · 8% Processor architecture and microarchitecture · 5% |
Topics — the 9 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing › scientific computing systems
adaptive mesh refinement |
0.6 | 2 | 2018 | Phase asynchronous AMR execution for productive and performant astrophysical flows · SC 2018 Perilla: metadata-based optimizations of an asynchronous runtime for adaptive mesh refinement · SC 2016 |
High-performance computing › supercomputing
exascale computing |
0.4 | 1 | 2020 | Preparing nuclear astrophysics for exascale · SC 2020 |
High-performance computing › performance engineering
performance portability |
0.4 | 1 | 2020 | Preparing nuclear astrophysics for exascale · SC 2020 |
High-performance computing
scientific computing |
0.4 | 1 | 2020 | Preparing nuclear astrophysics for exascale · SC 2020 |
High-performance computing › scientific computing
scientific computing application |
0.3 | 1 | 2018 | Phase asynchronous AMR execution for productive and performant astrophysical flows · SC 2018 |
High-performance computing › numerical linear algebra › linear solver
iterative linear solvers |
0.1 | 1 | 2012 | Optimization of geometric multigrid for emerging multi- and manycore processors · SC 2012 |
High-performance computing › performance optimization
many-core processor optimization |
0.1 | 1 | 2012 | Optimization of geometric multigrid for emerging multi- and manycore processors · SC 2012 |
High-performance computing › numerical linear algebra › linear solver › iterative linear solvers
multigrid method |
0.1 | 1 | 2012 | Optimization of geometric multigrid for emerging multi- and manycore processors · SC 2012 |
Processor architecture and microarchitecture
SIMD |
0.1 | 1 | 2012 | Optimization of geometric multigrid for emerging multi- and manycore processors · SC 2012 |
Methods — techniques the papers use, named apart from their topics
metadata-based optimization · 0.2threaded wavefront · 0.1operator fusion · 0.1dynamic threading · 0.1communication aggregation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Porting WarpX to GPU-accelerated platformsabstractWarpX is a general purpose electromagnetic particle-in-cell code that was originally designed to run on many-core CPU architectures. We describe the strategy, based on the AMReX library, followed to allow WarpX to use the GPU-accelerated nodes on OLCF’s Summit supercomputer, a strategy we believe will extend to the upcoming machines Frontier and Aurora. We summarize the challenges encountered, lessons learned, and give current performance results on a series of relevant benchmark problems. Andrew Myers 0001, Ann S. Almgren, Ligia Diana Amorim, John B. Bell, Luca Fedeli, Lixin Ge, Kevin Gott, David P. Grote, Mark J. Hogan, Axel Huebl, Revathi Jambunathan, Rémi Lehe, Cho-Kuen Ng, Michael E. Rowan, Olga Shapoval, Maxence Thévenet, Jean-Luc Vay, Henri Vincenti, Eloise Yang, Neïl Zaïm, Weiqun Zhang, Yinjian Zhao, Edoardo Zoni |
Parallel Comput. | 2 |
| 2020 | Preparing nuclear astrophysics for exascaleabstractAstrophysical explosions such as supernovae are fascinating events that require sophisticated algorithms and substantial computational power to model. Castro and MAESTROeX are nuclear astrophysics codes that simulate thermonuclear fusion in the context of supernovae and X-ray bursts. Examining these nuclear burning processes using high resolution simulations is critical for understanding how these astrophysical explosions occur. In this paper we describe the changes that have been made to these codes to transform them from standard MPI + OpenMP codes targeted at petascale CPU-based systems into a form compatible with the pre-exascale systems now online and the exascale systems coming soon. We then discuss what new science is possible to run on systems such as Summit and Perlmutter that could not have been achieved on the previous generation of supercomputers. Max P. Katz, Ann S. Almgren, Maria Barrios Sazo, Kiran Eiden, Kevin Gott, Alice Harpole, Jean M. Sexton, Donald E. Willcox, Weiqun Zhang, Michael Zingale |
SC | 2 |
| 2018 | Phase asynchronous AMR execution for productive and performant astrophysical flows
Muhammed Nufail Farooqi, Tan Nguyen 0001, Weiqun Zhang, Ann S. Almgren, John Shalf, Didem Unat |
SC | 4 |
| 2017 | Nonintrusive AMR Asynchrony for Communication Optimization
Muhammed Nufail Farooqi, Didem Unat, Tan Nguyen 0001, Weiqun Zhang, Ann S. Almgren, John Shalf |
Euro-Par | 5 |
| 2017 | Overlapping Data Transfers with Computation on GPU with TilesabstractGPUs are employed to accelerate scientific applications however they require much more programming effort from the programmers particularly because of the disjoint address spaces between the host and the device. OpenACC and OpenMP 4.0 provide directive based programming solutions to alleviate the programming burden however synchronous data movement can create a performance bottleneck in fully taking advantage of GPUs. We propose a tiling based programming model and its library that simplifies the development of GPU programs and overlaps the data movement with computation. The programming model decomposes the data and computation into tiles and treats them as the main data transfer and execution units, which enables pipelining the transfers to hide the transfer latency. Moreover, partitioning application data into tiles allows the programmer to still take advantage of GPU even though application data cannot fit into the device memory. The library leverages C++ lambda functions, OpenACC directives, CUDA streams and tiling API from TiDA to support both productivity and performance. We show the performance of the library on a data transfer-intensive and a compute-intensive kernels and compare its speedup against OpenACC and CUDA. The results indicate that the library can hide the transfer latency, handle the cases where there is no sufficient device memory, and achieves reasonable performance. Burak Bastem, Didem Unat, Weiqun Zhang, Ann S. Almgren, John Shalf |
ICPP | 4 |
| 2016 | Perilla: metadata-based optimizations of an asynchronous runtime for adaptive mesh refinementabstractHardware architecture is increasingly complex, urging the development of asynchronous runtime systems with advance resource and locality management supports. However, these supports may come at the cost of complicating the user interface while programming remains one of the major constraints to wide adoption of asynchronous runtimes in practice. In this paper, we propose a solution that leverages application metadata to enable challenging optimizations as well as to facilitate the task of transforming legacy code to an asynchronous representation. We develop Perilla, a task graph-based runtime system that requires only modest programming effort. Perilla utilizes metadata of an AMR software framework to enable various optimizations at the communication layer without complicating its API. Experimental results with different applications on up to 24K processor cores show that Perilla can realize up to 1.44x speedup over the synchronous code variant. The metadata enabled optimizations account for 25% to 100% of the performance improvement. Tan Nguyen 0001, Didem Unat, Weiqun Zhang, Ann S. Almgren, Muhammed Nufail Farooqi, John Shalf |
SC | 4 |
| 2014 | s-Step Krylov Subspace Methods as Bottom Solvers for Geometric MultigridabstractGeometric multigrid solvers within adaptive mesh refinement (AMR) applications often reach a point where further coarsening of the grid becomes impractical as individual sub domain sizes approach unity. At this point the most common solution is to use a bottom solver, such as BiCGStab, to reduce the residual by a fixed factor at the coarsest level. Each iteration of BiCGStab requires multiple global reductions (MPI collectives). As the number of BiCGStab iterations required for convergence grows with problem size, and the time for each collective operation increases with machine scale, bottom solves in large-scale applications can constitute a significant fraction of the overall multigrid solve time. In this paper, we implement, evaluate, and optimize a communication-avoiding s-step formulation of BiCGStab (CABiCGStab for short) as a high-performance, distributed-memory bottom solver for geometric multigrid solvers. This is the first time s-step Krylov subspace methods have been leveraged to improve multigrid bottom solver performance. We use a synthetic benchmark for detailed analysis and integrate the best implementation into BoxLib in order to evaluate the benefit of a s-step Krylov subspace method on the multigrid solves found in the applications LMC and Nyx on up to 32,768 cores on the Cray XE6 at NERSC. Overall, we see bottom solver improvements of up to 4.2x on synthetic problems and up to 2.7x in real applications. This results in as much as a 1.5x improvement in solver performance in real applications. Samuel Williams 0001, Michael Lijewski, Ann S. Almgren, Brian van Straalen, Erin Carson, Nicholas Knight, James Demmel |
IPDPS | 3 |
| 2014 | A survey of high level frameworks in block-structured adaptive mesh refinement packages
Anshu Dubey, Ann S. Almgren, John B. Bell, Martin Berzins, Steven R. Brandt, Greg Bryan, Phillip Colella, Daniel T. Graves, Michael Lijewski, Frank Löffler 0001, Brian W. O'Shea, Erik Schnetter, Brian van Straalen, Klaus Weide |
J. Parallel Distributed Comput. | 2 |
| 2012 | Optimization of geometric multigrid for emerging multi- and manycore processorsabstractMultigrid methods are widely used to accelerate the convergence of iterative solvers for linear systems used in a number of different application areas. In this paper, we explore optimization techniques for geometric multigrid on existing and emerging multicore systems including the Opteron-based Cray XE6, Intel® Xeon® E5-2670 and X5550 processor-based Infiniband clusters, as well as the new Intel® Xeon Phi coprocessor (Knights Corner). Our work examines a variety of novel techniques including communication-aggregation, threaded wavefront-based DRAM communication-avoiding, dynamic threading decisions, SIMDization, and fusion of operators. We quantify performance through each phase of the V-cycle for both single-node and distributed-memory experiments and provide detailed analysis for each class of optimization. Results show our optimizations yield significant speedups across a variety of subdomain sizes while simultaneously demonstrating the potential of multi- and manycore processors to dramatically accelerate single-node performance. However, our analysis also indicates that improvements in networks and communication will be essential to reap the potential of manycore processors in large-scale multigrid calculations. Samuel Williams 0001, Dhiraj D. Kalamkar, Amik Singh, Anand M. Deshpande, Brian van Straalen, Mikhail Smelyanskiy, Ann S. Almgren, Pradeep Dubey, John Shalf, Leonid Oliker |
SC | 7 |