Erik W. Draeger

dblp:01/7047 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
6since 2021 · last 2024
0000-0003-4063-0253ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2024 Designing a GPU-Accelerated Communication Layer for Efficient Fluid-Structure Interaction Computations on Heterogeneous Systems
abstract
As biological research demands simulations with increasingly larger cell counts, optimizing these models for largescale deployment on heterogeneous supercomputing resources becomes crucial. This requires the redesign of fluid-structure interaction tasks written around distributed data structures built for CPU-based systems, where design flexibility and overall memory footprint are key considerations, to instead be performant on CPU-GPU machines. This paper describes the trade-offs of offloading communication tasks to the GPUs and the corresponding changes to the underlying data structures required, along with new algorithms that significantly reduce time-to-solution. At scale performance of our GPU implementation is evaluated on the Polaris and Frontier leadership systems. Real-world workloads involving millions of deformable cells are evaluated. We analyze the competing factors that come into play when designing a communication layer for a fluid-structure interaction code, including code efficiency, complexity, and GPU memory demands, and offer advice to other high performance computing applications facing similar decisions.
Aristotle X. Martin, Bálint Joó, Runxin Wu, Mohammed Shihab Kabir, Erik W. Draeger, Amanda Randles
SC6
2023 Optimizing Cloud Computing Resource Usage for Hemodynamic Simulation
abstract
Cloud computing resources are becoming an increasingly attractive option for simulation workflows but require users to assess a wider variety of hardware options and associated costs than required by traditional in-house hardware or fixed allocations at leadership computing facilities. The pay-as-you-go model used by cloud providers gives users the opportunity to make more nuanced cost-benefit decisions at runtime by choosing hardware that best matches a given workload, but creates the risk of suboptimal allocation strategies or inadvertent cost overruns. In this work, we propose the use of an iteratively-refined performance model to optimize cloud simulation campaigns against overall cost, throughput, or maximum time to solution. Hemodynamic simulations represent an excellent use case for these assessments, as the relative costs and dominant terms in the performance model can vary widely with hardware, numerical parameters and physics models. Performance and scaling behavior of hemodynamic simulations on multiple cloud services as well as a traditional compute cluster are collected and evaluated, and an initial performance model is proposed along with a strategy for dynamically refining it with additional experimental data.
William Ladd, Christopher Jensen, Madhurima Vardhan, Jeff Ames, Jeff R. Hammond, Erik W. Draeger, Amanda Randles
IPDPS6
2023 Enhancing Adaptive Physics Refinement Simulations Through the Addition of Realistic Red Blood Cell Counts
abstract
Simulations of cancer cell transport require accurately modeling mm-scale and longer trajectories through a circulatory system containing trillions of deformable red blood cells, whose intercellular interactions require submicron fidelity. Using a hybrid CPU-GPU approach, we extend the advanced physics refinement (APR) method to couple a finely-resolved region of explicitly-modeled red blood cells to a coarsely-resolved bulk fluid domain. We further develop algorithms that: capture the dynamics at the interface of differing viscosities, maintain hematocrit within the cell-filled volume, and move the finely-resolved region and encapsulated cells while tracking an individual cancer cell. Comparison to a fully-resolved fluid-structure interaction model is presented for verification. Finally, we use the advanced APR method to simulate cancer cell transport over a mm-scale distance while maintaining a local region of RBCs, using a fraction of the computational power required to run a fully-resolved model.
Sayan Roychowdhury, Samreen T. Mahmud, Aristotle X. Martin, Peter Balogh, Daniel F. Puleri, John Gounley, Erik W. Draeger, Amanda Randles
SC7
2023 Cloud Computing to Enable Wearable-Driven Longitudinal Hemodynamic Maps
abstract
Tracking hemodynamic responses to treatment and stimuli over long periods remains a grand challenge. Moving from established single-heartbeat technology to longitudinal profiles would require continuous data describing how the patient's state evolves, new methods to extend the temporal domain over which flow is sampled, and high-throughput computing resources. While personalized digital twins can accurately measure 3D hemodynamics over several heartbeats, state-of-the-art methods would require hundreds of years of wallclock time on leadership scale systems to simulate one day of activity. To address these challenges, we propose a cloud-based, parallel-in-time framework leveraging continuous data from wearable devices to capture the first 3D patient-specific, longitudinal hemodynamic maps. We demonstrate the validity of our method by establishing ground truth data for 750 beats and comparing the results. Our cloud-based framework is based on an initial fixed set of simulations to enable the wearable-informed creation of personalized longitudinal hemodynamic maps.
Cyrus Tanade, Emily Rakestraw, William Ladd, Erik W. Draeger, Amanda Randles
SC4
2022 High Performance Adaptive Physics Refinement to Enable Large-Scale Tracking of Cancer Cell Trajectory
abstract
The ability to track simulated cancer cells through the circulatory system, important for developing a mechanistic understanding of metastatic spread, pushes the limits of today's supercomputers by requiring the simulation of large fluid volumes at cellular-scale resolution. To overcome this challenge, we introduce a new adaptive physics refinement (APR) method that captures cellular-scale interaction across large domains and leverages a hybrid CPU-GPU approach to maximize performance. Through algorithmic advances that integrate multi-physics and multi-resolution models, we establish a finely resolved window with explicitly modeled cells coupled to a coarsely resolved bulk fluid domain. In this work we present multiple validations of the APR framework by comparing against fully resolved fluid-structure interaction methods and employ techniques, such as latency hiding and maximizing memory bandwidth, to effectively utilize heterogeneous node architectures. Collectively, these computational developments and performance optimizations provide a robust and scalable framework to enable system-level simulations of cancer cell transport.
Daniel F. Puleri, Sayan Roychowdhury, Peter Balogh, John Gounley, Erik W. Draeger, Jeff Ames, Adebayo Adebiyi, Simbarashe Chidyagwai, Benjamín Hernández, Seyong Lee, Shirley V. Moore, Jeffrey S. Vetter, Amanda Randles
CLUSTER5
2022 Propagation Pattern for Moment Representation of the Lattice Boltzmann Method
abstract
A propagation pattern for the moment representation of the regularized lattice Boltzmann method (LBM) in three dimensions is presented. Using effectively lossless compression, the simulation state is stored as a set of moments of the lattice Boltzmann distribution function, instead of the distribution function itself. An efficient cache-aware propagation pattern for this moment representation has the effect of substantially reducing both the storage and memory bandwidth required for LBM simulations. This paper extends recent work with the moment representation by expanding the performance analysis on central processing unit (CPU) architectures, considering how boundary conditions are implemented, and demonstrating the effectiveness of the moment representation on a graphics processing unit (GPU) architecture.
John Gounley, Madhurima Vardhan, Erik W. Draeger, Pedro Valero-Lara, Shirley V. Moore, Amanda Randles
IEEE Trans. Parallel Distributed Syst.3
2019 Multi-physics simulations of particle tracking in arterial geometries with a scalable moving window algorithm
abstract
In arterial systems, cancer cell trajectories determine metastatic cancer locations; similarly, particle trajectories determine drug delivery distribution. Predicting trajectories is challenging, as the dynamics are affected by local interactions with red blood cells, complex hemodynamic flow structure, and downstream factors such as stenoses or blockages. Direct simulation is not possible, as a single simulation of a large arterial domain with explicit red blood cells is currently intractable on even the largest supercomputers. To overcome this limitation, we present a multi-physics adaptive window algorithm, in which individual red blood cells are explicitly modeled in a small region of interest moving through a coupled arterial fluid domain. We describe the coupling between the window and fluid domains, including automatic insertion and deletion of explicit cells and dynamic tracking of cells of interest by the window. We show that this algorithm scales efficiently on heterogeneous architectures and enables us to perform large, highly-resolved particle-tracking simulations that would otherwise be intractable.
Gregory Herschlag, John Gounley, Sayan Roychowdhury, Erik W. Draeger, Amanda Randles
CLUSTER4
2019 Moment representation in the lattice Boltzmann method on massively parallel hardware
abstract
The widely-used lattice Boltzmann method (LBM) for computational fluid dynamics is highly scalable, but also significantly memory bandwidth-bound on current architectures. This paper presents a new regularized LBM implementation that reduces the memory footprint by only storing macroscopic, moment-based data. We show that the amount of data that must be stored in memory during a simulation is reduced by up to 47%. We also present a technique for cache-aware data re-utilization and show that optimizing cache utilization to limit data motion results in a similar improvement in time to solution. These new algorithms are implemented in the hemodynamics solver HARVEY and demonstrated using both idealized and realistic biological geometries. We develop a performance model for the moment representation algorithm and evaluate the performance on Summit.
Madhurima Vardhan, John Gounley, Luiz Hegele, Erik W. Draeger, Amanda Randles
SC4
2017 Massively parallel first-principles simulation of electron dynamics in materials
Erik W. Draeger, Xavier Andrade, John A. Gunnels, Abhinav Bhatele, André Schleife, Alfredo A. Correa
J. Parallel Distributed Comput.1
2016 Interactive exploration of atomic trajectories through relative-angle distribution and associated uncertainties
abstract
Exploration of atomic trajectories is fundamental to understanding and characterizing complex chemical systems important in many applications. For instance, any new insight into the mechanisms of ionic migration in catalytic materials could lead to a substantial increase in battery performance. A new statistical measure, called the relative-angle distribution, has been proposed to understand complex motion - whether Brownian, ballistic, or diffusive. The relative-angle distribution can be represented as a collection of 1D histograms, but is currently created in a slow, offline process, making any parameter exploration a tedious and time-consuming task. Furthermore, the resulting plot can hide uncertainty in both the data and the visualization. As a result, once rastered or printed at a fixed resolution, these histograms can be misleading. We present a new analysis tool for the exploration of atomic trajectories that combines an interactive histogram visualization with uncertainty information for both data and plotting errors, and is also linked to an interactive 3D display of trajectories. Our tool enables a holistic exploration of trajectories previously not feasible, with the potential for significant scientific impact. In collaboration with domain experts, we have deployed our tool ta analyze molecular dynamics simulations of lithium-ion diffusion. Users have found that the tool significantly accelerates the exploration process and have used it to validate a number of previously unconfirmed hypotheses.
Harsh Bhatia, Attila Gyulassy, Valerio Pascucci, Martina Bremer, Mitchell T. Ong, Vincenzo Lordi, Erik W. Draeger, John E. Pask, Peer-Timo Bremer
PacificVis7
2016 Massively Parallel First-Principles Simulation of Electron Dynamics in Materials
abstract
We present a highly scalable, parallel implementation of first-principles electron dynamics coupled with molecular dynamics (MD). By using optimized kernels, network topology aware communication, and by fully distributing all terms in the time-dependent Kohn-Sham equation, we demonstrate unprecedented time to solution for disordered aluminum systems of 2,000 atoms (22,000 electrons) and 5,400 atoms (59,400 electrons), with wall clock time as low as 7.5 seconds per MD time step. Despite a significant amount of non-local communication required in every iteration, we achieved excellent strong scaling and sustained performance on the Sequoia Blue Gene/Q supercomputer at LLNL. We obtained up to 59% of the theoretical sustained peak performance on 16,384 nodes and performance of 8.75 Petaflop/s (43% of theoretical peak) on the full 98,304 node machine (1,572,864 cores). Scalable explicit electron dynamics allows for the study of phenomena beyond the reach of standard first principles MD, in particular, materials subject to strong or rapid perturbations, such as pulsed electromagnetic radiation, particle irradiation, or strong electric currents.
Erik W. Draeger, Xavier Andrade, John A. Gunnels, Abhinav Bhatele, André Schleife, Alfredo A. Correa
IPDPS1
2016 Modeling dilute solutions using first-principles molecular dynamics: computing more than a million atoms with over a million cores
abstract
First-Principles Molecular Dynamics (FPMD) methods, although powerful, are notoriously expensive computationally due to the quantum modeling of electrons. Traditional FPMD approaches have typically been limited to a few thousand atoms at most, due to O(N3) or worse solver complexity and the large amount of communication required for highly parallel implementations. Attempts to lower the complexity have often introduced uncontrolled approximations or systematic errors. Using a robust new algorithm, we have developed an O(N) complexity solver for electronic structure problems with fully controllable numerical error. Its minimal use of global communications yields excellent scalability, allowing for very accurate FPMD simulations of more than a million atoms on over a million cores. At these scales, this approach provides multiple orders of magnitude speedup compared to the standard plane-wave approach typically used in condensed matter applications, without sacrificing accuracy. This will open up entire new classes of FPMD simulations, e.g. dilute aqueous solutions.
Jean-Luc Fattebert, Daniel Osei-Kuffuor, Erik W. Draeger, Tadashi Ogitsu, William D. Krauss
SC3
2015 Massively parallel models of the human circulatory system
abstract
The potential impact of blood flow simulations on the diagnosis and treatment of patients suffering from vascular disease is tremendous. Empowering models of the full arterial tree can provide insight into diseases such as arterial hypertension and enables the study of the influence of local factors on global hemodynamics. We present a new, highly scalable implementation of the lattice Boltzmann method which addresses key challenges such as multiscale coupling, limited memory capacity and bandwidth, and robust load balancing in complex geometries. We demonstrate the strong scaling of a three-dimensional, high-resolution simulation of hemodynamics in the systemic arterial tree on 1,572,864 cores of Blue Gene/Q. Faster calculation of flow in full arterial networks enables unprecedented risk stratification on a perpatient basis. In pursuit of this goal, we have introduced computational advances that significantly reduce time-to-solution for biofluidic simulations.
Amanda Randles, Erik W. Draeger, Tomas Oppelstrup, Liam Krauss, John A. Gunnels
SC2
2012 Mapping applications with collectives over sub-communicators on torus networks
abstract
The placement of tasks in a parallel application on specific nodes of a supercomputer can significantly impact performance. Traditionally, this task mapping has focused on reducing the distance between communicating tasks on the physical network. This minimizes the number of hops that point-to-point messages travel and thus reduces link sharing between messages and contention. However, for applications that use collectives over sub-communicators, this heuristic may not be optimal. Many collectives can benefit from an increase in bandwidth even at the cost of an increase in hop count, especially when sending large messages. For example, placing communicating tasks in a cube configuration rather than a plane or a line on a torus network increases the number of possible paths messages might take. This increases the available bandwidth which can lead to significant performance gains. We have developed Rubik, a tool that provides a simple and intuitive interface to create a wide variety of mappings for structured communication patterns. Rubik supports a number of elementary operations such as splits, tilts, or shifts, that can be combined into a large number of unique patterns. Each operation can be applied to disjoint groups of processes involved in collectives to increase the effective bandwidth. We demonstrate the use of Rubik for improving performance of two parallel codes, pF3D and Qbox, which use collectives over sub-communicators.
Abhinav Bhatele, Todd Gamblin, Steve H. Langer, Peer-Timo Bremer, Erik W. Draeger, Bernd Hamann, Katherine E. Isaacs, Aaditya G. Landge, Joshua A. Levine, Valerio Pascucci, Martin Schulz 0001, Charles H. Still
SC5
2012 Toward real-time modeling of human heart ventricles at cellular resolution: simulation of drug-induced arrhythmias
abstract
We have developed a highly efficient and scalable cardiac electrophysiology simulation capability that supports groundbreaking resolution and detail to elucidate the mechanisms of sudden cardiac death from arrhythmia. We can simulate thousands of heartbeats at a resolution of 0.1 mm, comparable to the size of cardiac cells, thereby enabling scientific inquiry not previously possible. Based on scaling results from the partially deployed Sequoia IBM Blue Gene/Q machine at Lawrence Livermore National Laboratory and planned optimizations, we estimate that by SC12 we will simulate 8 -- 10 heartbeats per minute -- a time-to-solution 400 -- 500 times faster than the state-of-the-art. Performance between 8 and 11 PFlop/s on the full 1,572,864 cores is anticipated, representing 40 -- 55 percent of peak. The power of the model is demonstrated by illuminating the subtle arrhythmogenic mechanisms of anti-arrhythmic drugs that paradoxically increase arrhythmias in some patient populations.
Arthur A. Mirin, David F. Richards, James N. Glosli, Erik W. Draeger, Bor Chan, Jean-Luc Fattebert, William D. Krauss, Tomas Oppelstrup, John Jeremy Rice, John A. Gunnels, Viatcheslav Gurev, Changhoan Kim, John Magerlein, Matthias Reumann, Hui-Fang Wen
SC4
2009 Beyond homogeneous decomposition: scaling long-range forces on Massively Parallel Systems
abstract
With supercomputers anticipated to expand from thousands to millions of cores, one of the challenges facing scientists is how to effectively utilize this ever-increasing number. We report here an approach that creates a heterogeneous decomposition by partitioning effort according to the scaling properties of the component algorithms. We demonstrate our strategy by developing a capability to model hot dense plasma. We have performed benchmark calculations ranging from millions to billions of charged particles, including a 2.8 billion particle simulation that achieved 259.9 TFlop/s (26% of peak performance) on the 294,912 cpu JUGENE computer at the Jülich Supercomputing Centre in Germany. With this unprecedented simulation capability we have begun an investigation of plasma fusion physics under conditions where both theory and experiment are lacking--in the strongly-coupled regime as the plasma begins to burn.
David F. Richards, James N. Glosli, Bor Chan, Milo R. Dorr, Erik W. Draeger, Jean-Luc Fattebert, William D. Krauss, Thomas E. Spelce, Frederick H. Streitz, Michael P. Surh, John A. Gunnels
SC5
2006 Gordon Bell finalists I - Large-scale electronic structure calculations of high-Z metals on the BlueGene/L platform
abstract
First-principles simulations of high-Z metallic systems using the Qbox code on the BlueGene/L supercomputer demonstrate unprecedented performance and scaling for a quantum simulation code. Specifically designed to take advantage of massively-parallel systems like BlueGene/L, Qbox demonstrates excellent parallel efficiency and peak performance. A sustained peak performance of 207.3 TFlop/s was measured on 65,536 nodes, corresponding to 56.5% of the theoretical full machine peak using all 128k CPUs.
François Gygi, Erik W. Draeger, Martin Schulz 0001, Bronis R. de Supinski, John A. Gunnels, Vernon Austel, James C. Sexton, Franz Franchetti, Stefan Kral, Christoph W. Ueberhuber, Juergen Lorenz
SC2
2005 Large-Scale First-Principles Molecular Dynamics simulations on the BlueGene/L Platform using the Qbox code
abstract
We demonstrate that the Qbox code supports unprecedented large-scale First-Principles Molecular Dynamics (FPMD) applications on the BlueGene/L supercomputer. Qbox is an FPMD implementation specifically designed for large-scale parallel platforms such as BlueGene/L. Strong scaling tests for a Materials Science application show an 86% scaling efficiency between 1024 and 32,768 CPUs. Measurements of performance by means of hardware counters show that 36% of the peak FPU performance can be attained.
François Gygi, Robert K. Yates, Juergen Lorenz, Erik W. Draeger, Franz Franchetti, Christoph W. Ueberhuber, Bronis R. de Supinski, Stefan Kral, John A. Gunnels, James C. Sexton
SC4