EDBT 2026 Demo / reviewers in the wild / expert
Amanda Randles
dblp:128/1814 · also Amanda E. Peters, Amanda Peters Randles
· DBLP profile ↗
25ranked-venue papers
6as first author
9since 2021 · last 2024
0000-0001-6318-3885ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Designing a GPU-Accelerated Communication Layer for Efficient Fluid-Structure Interaction Computations on Heterogeneous SystemsabstractAs biological research demands simulations with increasingly larger cell counts, optimizing these models for largescale deployment on heterogeneous supercomputing resources becomes crucial. This requires the redesign of fluid-structure interaction tasks written around distributed data structures built for CPU-based systems, where design flexibility and overall memory footprint are key considerations, to instead be performant on CPU-GPU machines. This paper describes the trade-offs of offloading communication tasks to the GPUs and the corresponding changes to the underlying data structures required, along with new algorithms that significantly reduce time-to-solution. At scale performance of our GPU implementation is evaluated on the Polaris and Frontier leadership systems. Real-world workloads involving millions of deformable cells are evaluated. We analyze the competing factors that come into play when designing a communication layer for a fluid-structure interaction code, including code efficiency, complexity, and GPU memory demands, and offer advice to other high performance computing applications facing similar decisions. Aristotle X. Martin, Bálint Joó, Runxin Wu, Mohammed Shihab Kabir, Erik W. Draeger, Amanda Randles |
SC | 7 |
| 2023 | Optimizing Cloud Computing Resource Usage for Hemodynamic SimulationabstractCloud computing resources are becoming an increasingly attractive option for simulation workflows but require users to assess a wider variety of hardware options and associated costs than required by traditional in-house hardware or fixed allocations at leadership computing facilities. The pay-as-you-go model used by cloud providers gives users the opportunity to make more nuanced cost-benefit decisions at runtime by choosing hardware that best matches a given workload, but creates the risk of suboptimal allocation strategies or inadvertent cost overruns. In this work, we propose the use of an iteratively-refined performance model to optimize cloud simulation campaigns against overall cost, throughput, or maximum time to solution. Hemodynamic simulations represent an excellent use case for these assessments, as the relative costs and dominant terms in the performance model can vary widely with hardware, numerical parameters and physics models. Performance and scaling behavior of hemodynamic simulations on multiple cloud services as well as a traditional compute cluster are collected and evaluated, and an initial performance model is proposed along with a strategy for dynamically refining it with additional experimental data. William Ladd, Christopher Jensen, Madhurima Vardhan, Jeff Ames, Jeff R. Hammond, Erik W. Draeger, Amanda Randles |
IPDPS | 7 |
| 2023 | Enhancing Adaptive Physics Refinement Simulations Through the Addition of Realistic Red Blood Cell CountsabstractSimulations of cancer cell transport require accurately modeling mm-scale and longer trajectories through a circulatory system containing trillions of deformable red blood cells, whose intercellular interactions require submicron fidelity. Using a hybrid CPU-GPU approach, we extend the advanced physics refinement (APR) method to couple a finely-resolved region of explicitly-modeled red blood cells to a coarsely-resolved bulk fluid domain. We further develop algorithms that: capture the dynamics at the interface of differing viscosities, maintain hematocrit within the cell-filled volume, and move the finely-resolved region and encapsulated cells while tracking an individual cancer cell. Comparison to a fully-resolved fluid-structure interaction model is presented for verification. Finally, we use the advanced APR method to simulate cancer cell transport over a mm-scale distance while maintaining a local region of RBCs, using a fraction of the computational power required to run a fully-resolved model. Sayan Roychowdhury, Samreen T. Mahmud, Aristotle X. Martin, Peter Balogh, Daniel F. Puleri, John Gounley, Erik W. Draeger, Amanda Randles |
SC | 8 |
| 2023 | Cloud Computing to Enable Wearable-Driven Longitudinal Hemodynamic MapsabstractTracking hemodynamic responses to treatment and stimuli over long periods remains a grand challenge. Moving from established single-heartbeat technology to longitudinal profiles would require continuous data describing how the patient's state evolves, new methods to extend the temporal domain over which flow is sampled, and high-throughput computing resources. While personalized digital twins can accurately measure 3D hemodynamics over several heartbeats, state-of-the-art methods would require hundreds of years of wallclock time on leadership scale systems to simulate one day of activity. To address these challenges, we propose a cloud-based, parallel-in-time framework leveraging continuous data from wearable devices to capture the first 3D patient-specific, longitudinal hemodynamic maps. We demonstrate the validity of our method by establishing ground truth data for 750 beats and comparing the results. Our cloud-based framework is based on an initial fixed set of simulations to enable the wearable-informed creation of personalized longitudinal hemodynamic maps. Cyrus Tanade, Emily Rakestraw, William Ladd, Erik W. Draeger, Amanda Randles |
SC | 5 |
| 2022 | High Performance Adaptive Physics Refinement to Enable Large-Scale Tracking of Cancer Cell TrajectoryabstractThe ability to track simulated cancer cells through the circulatory system, important for developing a mechanistic understanding of metastatic spread, pushes the limits of today's supercomputers by requiring the simulation of large fluid volumes at cellular-scale resolution. To overcome this challenge, we introduce a new adaptive physics refinement (APR) method that captures cellular-scale interaction across large domains and leverages a hybrid CPU-GPU approach to maximize performance. Through algorithmic advances that integrate multi-physics and multi-resolution models, we establish a finely resolved window with explicitly modeled cells coupled to a coarsely resolved bulk fluid domain. In this work we present multiple validations of the APR framework by comparing against fully resolved fluid-structure interaction methods and employ techniques, such as latency hiding and maximizing memory bandwidth, to effectively utilize heterogeneous node architectures. Collectively, these computational developments and performance optimizations provide a robust and scalable framework to enable system-level simulations of cancer cell transport. Daniel F. Puleri, Sayan Roychowdhury, Peter Balogh, John Gounley, Erik W. Draeger, Jeff Ames, Adebayo Adebiyi, Simbarashe Chidyagwai, Benjamín Hernández, Seyong Lee, Shirley V. Moore, Jeffrey S. Vetter, Amanda Randles |
CLUSTER | 13 |
| 2022 | Distributed Acceleration of Adhesive Dynamics SimulationsabstractCell adhesion plays a critical role in processes ranging from leukocyte migration to cancer cell transport during metastasis. Adhesive cell interactions can occur over large distances in microvessel networks with cells traveling over distances much greater than the length scale of their own diameter. Therefore, biologically relevant investigations necessitate efficient modeling of large field-of-view domains, but current models are limited by simulating such geometries at the sub-micron scale required to model adhesive interactions which greatly increases the computational requirements for even small domain sizes. In this study we introduce a hybrid scheme reliant on both on-node and distributed parallelism to accelerate a fully deformable adhesive dynamics cell model. This scheme leads to performant system usage of modern supercomputers which use a many-core per-node architecture. On-node acceleration is augmented by a combination of spatial data structures and algorithmic changes to lessen the need for atomic operations. This deformable adhesive cell model accelerated with hybrid parallelization allows us to bridge the gap between high-resolution cell models which can capture the sub-micron adhesive interactions between the cell and its microenvironment, and large-scale fluid-structure interaction (FSI) models which can track cells over considerable distances. By integrating the sub-micron simulation environment into a distributed FSI simulation we enable the study of previously unfeasible research questions involving numerous adhesive cells in microvessel networks such as cancer cell transport through the microcirculation. Daniel F. Puleri, Aristotle X. Martin, Amanda Randles |
EuroMPI | 3 |
| 2022 | Evaluation of U-Net Based Architectures for Automatic Aortic Dissection SegmentationabstractSegmentation and reconstruction of arteries is important for a variety of medical and engineering fields, such as surgical planning and physiological modeling. However, manual methods can be laborious and subject to a high degree of human variability. In this work, we developed various convolutional neural network ( CNN ) architectures to segment Stanford type B aortic dissections ( TBADs ), characterized by a tear in the descending aortic wall creating a normal channel of blood flow called a true lumen and a pathologic channel within the wall called a false lumen. We introduced several variations to the two-dimensional ( 2D ) and three-dimensional (3 D ) U-Net, where small stacks of slices were inputted into the networks instead of individual slices or whole geometries. We compared these variations with a variety of CNN segmentation architectures and found that stacking the input data slices in the upward direction with 2D U-Net improved segmentation accuracy, as measured by the Dice similarity coefficient ( DC ) and point-by-point average distance ( AVD ), by more than 15\% . Our optimal architecture produced DC scores of 0.94, 0.88, and 0.90 and AVD values of 0.074, 0.22, and 0.11 in the whole aorta, true lumen, and false lumen, respectively. Altogether, the predicted reconstructions closely matched manual reconstructions. Bradley Feiger, Erick Lorenzana, Colin L. V. Cooke, Roarke Horstmeyer, Muath Bishawi, Julie Doberne, G. Chad Hughes, David Ranney, Soraya Voigt, Amanda Randles |
ACM Trans. Comput. Heal. | 10 |
| 2022 | Propagation Pattern for Moment Representation of the Lattice Boltzmann MethodabstractA propagation pattern for the moment representation of the regularized lattice Boltzmann method (LBM) in three dimensions is presented. Using effectively lossless compression, the simulation state is stored as a set of moments of the lattice Boltzmann distribution function, instead of the distribution function itself. An efficient cache-aware propagation pattern for this moment representation has the effect of substantially reducing both the storage and memory bandwidth required for LBM simulations. This paper extends recent work with the moment representation by expanding the performance analysis on central processing unit (CPU) architectures, considering how boundary conditions are implemented, and demonstrating the effectiveness of the moment representation on a graphics processing unit (GPU) architecture. John Gounley, Madhurima Vardhan, Erik W. Draeger, Pedro Valero-Lara, Shirley V. Moore, Amanda Randles |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2021 | Analysis of GPU Data Access Patterns on Complex Geometries for the D3Q19 Lattice Boltzmann AlgorithmabstractGPU performance of the lattice Boltzmann method (LBM) depends heavily on memory access patterns. When implemented with GPUs on complex domains, typically, geometric data is accessed indirectly and lattice data is accessed lexicographically. Although there are a variety of other options, no study has examined the relative efficacy between them. Here, we examine a suite of memory access schemes via empirical testing and performance modeling. We find strong evidence that semi-direct is often better suited than the more common indirect addressing, providing increased computational speed and reducing memory consumption. For the layout, we find that the Collected Structure of Arrays (CSoA) and bundling layouts outperform the common Structure of Array layout; on V100 and P100 devices, CSoA consistently outperforms bundling, however the relationship is more complicated on K40 devices. When compared to state-of-the-art practices, our recommendations lead to speedups of 10-40 percent and reduce memory consumption up to 17 percent. Using performance modeling and computational experimentation, we determine the mechanisms behind the accelerations. We demonstrate that our results hold across multiple GPUs on two leadership class systems, and present the first near-optimal strong results for LBM with arterial geometries run on GPUs. Gregory Herschlag, Seyong Lee, Jeffrey S. Vetter, Amanda Randles |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2019 | Multi-physics simulations of particle tracking in arterial geometries with a scalable moving window algorithmabstractIn arterial systems, cancer cell trajectories determine metastatic cancer locations; similarly, particle trajectories determine drug delivery distribution. Predicting trajectories is challenging, as the dynamics are affected by local interactions with red blood cells, complex hemodynamic flow structure, and downstream factors such as stenoses or blockages. Direct simulation is not possible, as a single simulation of a large arterial domain with explicit red blood cells is currently intractable on even the largest supercomputers. To overcome this limitation, we present a multi-physics adaptive window algorithm, in which individual red blood cells are explicitly modeled in a small region of interest moving through a coupled arterial fluid domain. We describe the coupling between the window and fluid domains, including automatic insertion and deletion of explicit cells and dynamic tracking of cells of interest by the window. We show that this algorithm scales efficiently on heterogeneous architectures and enables us to perform large, highly-resolved particle-tracking simulations that would otherwise be intractable. Gregory Herschlag, John Gounley, Sayan Roychowdhury, Erik W. Draeger, Amanda Randles |
CLUSTER | 5 |
| 2019 | Investigating the Role of VR in a Simulation-Based Medical Planning System for Coronary Interventions
Madhurima Vardhan, Harvey Shi, John Gounley, S. James Chen, Andrew Kahn, Jane A. Leopold, Amanda Randles |
MICCAI (5) | 7 |
| 2019 | Moment representation in the lattice Boltzmann method on massively parallel hardwareabstractThe widely-used lattice Boltzmann method (LBM) for computational fluid dynamics is highly scalable, but also significantly memory bandwidth-bound on current architectures. This paper presents a new regularized LBM implementation that reduces the memory footprint by only storing macroscopic, moment-based data. We show that the amount of data that must be stored in memory during a simulation is reduced by up to 47%. We also present a technique for cache-aware data re-utilization and show that optimizing cache utilization to limit data motion results in a similar improvement in time to solution. These new algorithms are implemented in the hemodynamics solver HARVEY and demonstrated using both idealized and realistic biological geometries. We develop a performance model for the moment representation algorithm and evaluate the performance on Summit. Madhurima Vardhan, John Gounley, Luiz Hegele, Erik W. Draeger, Amanda Randles |
SC | 5 |
| 2019 | Performance portability study for massively parallel computational fluid dynamics application on scalable heterogeneous architectures
Seyong Lee, John Gounley, Amanda Randles, Jeffrey S. Vetter |
J. Parallel Distributed Comput. | 3 |
| 2018 | GPU Data Access on Complex Geometries for D3Q19 Lattice Boltzmann MethodabstractGPU performance of the lattice Boltzmann method (LBM) depends heavily on memory access patterns. When LBM is advanced with GPUs on complex computational domains, geometric data is typically accessed indirectly, and lattice data is typically accessed lexicographically in the Structure of Array (SoA) layout. Although there are a variety of existing access patterns beyond the typical choices, no study has yet examined the relative efficacy between them. Here, we compare a suite of memory access schemes via empirical testing and performance modeling. We find strong evidence that semi-direct addressing is the superior addressing scheme for the majority of cases examined: Semi-direct addressing increases computational speed and often reduces memory consumption. For lattice layout, we find that the Collected Structure of Arrays (CSoA) layout outperforms the SoA layout. When compared to state-of-the-art practices, our recommended addressing modifications lead to performance gains between 10-40% across different complex geometries, fluid volume fractions, and resolutions. The modifications also lead to a decrease in memory consumption by as much as 17%. Having discovered these improvements, we examine a highly resolved arterial geometry on a leadership class system. On this system we present the first near-optimal strong results for LBM with arterial geometries run on GPUs. We also demonstrate that the above recommendations remain valid for large scale, many device simulations, which leads to an increased computational speed and average memory usage reductions. To understand these observations, we employ performance modeling which reveals that semi-direct methods outperform indirect methods due to a reduced number of total loads/stores in memory, and that CSoA outperforms SoA and bundling due to improved caching behavior. Gregory Herschlag, Seyong Lee, Jeffrey S. Vetter, Amanda Randles |
IPDPS | 4 |
| 2015 | Massively parallel models of the human circulatory systemabstractThe potential impact of blood flow simulations on the diagnosis and treatment of patients suffering from vascular disease is tremendous. Empowering models of the full arterial tree can provide insight into diseases such as arterial hypertension and enables the study of the influence of local factors on global hemodynamics. We present a new, highly scalable implementation of the lattice Boltzmann method which addresses key challenges such as multiscale coupling, limited memory capacity and bandwidth, and robust load balancing in complex geometries. We demonstrate the strong scaling of a three-dimensional, high-resolution simulation of hemodynamics in the systemic arterial tree on 1,572,864 cores of Blue Gene/Q. Faster calculation of flow in full arterial networks enables unprecedented risk stratification on a perpatient basis. In pursuit of this goal, we have introduced computational advances that significantly reduce time-to-solution for biofluidic simulations. Amanda Randles, Erik W. Draeger, Tomas Oppelstrup, Liam Krauss, John A. Gunnels |
SC | 1 |
| 2015 | Scaling Support Vector Machines on modern HPC platforms
Yang You 0001, Haohuan Fu, Shuaiwen Song, Amanda Randles, Darren J. Kerbyson, Andrés Márquez 0001, Guangwen Yang 0002, Adolfy Hoisie |
J. Parallel Distributed Comput. | 4 |
| 2014 | A Spatio-temporal Coupling Method to Reduce the Time-to-Solution of Cardiovascular SimulationsabstractWe present a new parallel-in-time method designed to reduce the overall time-to-solution of a patient-specific cardiovascular flow simulation. Using a modified Para real algorithm, our approach extends strong scalability beyond spatial parallelism with fully controllable accuracy and no decrease in stability. We discuss the coupling of spatial and temporal domain decompositions used in our implementation, and showcase the use of the method on a study of blood flow through the aorta. We observe an additional 40% reduction in overall wall clock time with no significant loss of accuracy, in agreement with a predictive performance model. Amanda Randles, Efthimios Kaxiras |
IPDPS | 1 |
| 2014 | MIC-SVM: Designing a Highly Efficient Support Vector Machine for Advanced Modern Multi-core and Many-Core ArchitecturesabstractSupport Vector Machine (SVM) has been widely used in data-mining and Big Data applications as modern commercial databases start to attach an increasing importance to the analytic capabilities. In recent years, SVM was adapted to the field of High Performance Computing for power/performance prediction, auto-tuning, and runtime scheduling. However, even at the risk of losing prediction accuracy due to insufficient runtime information, researchers can only afford to apply offline model training to avoid significant runtime training overhead. Advanced multi- and many-core architectures offer massive parallelism with complex memory hierarchies which can make runtime training possible, but form a barrier to efficient parallel SVM design. To address the challenges above, we designed and implemented MIC-SVM, a highly efficient parallel SVM for x86 based multi-core and many-core architectures, such as the Intel Ivy Bridge CPUs and Intel Xeon Phi co-processor (MIC). We propose various novel analysis methods and optimization techniques to fully utilize the multilevel parallelism provided by these architectures and serve as general optimization methods for other machine learning tools. MIC-SVM achieves 4.4-84x and 18-47x speedups against the popular LIBSVM, on MIC and Ivy Bridge CPUs respectively, for several real-world data-mining datasets. Even compared with GPUSVM, run on a top of the line NVIDIA k20x GPU, the performance of our MIC-SVM is competitive. We also conduct a cross-platform performance comparison analysis, focusing on Ivy Bridge CPUs, MIC and GPUs, and provide insights on how to select the most suitable advanced architectures for specific algorithms and input data patterns. Yang You 0001, Shuaiwen Song, Haohuan Fu, Andrés Márquez 0001, Maryam Mehri Dehnavi, Kevin J. Barker, Kirk W. Cameron, Amanda Randles, Guangwen Yang 0002 |
IPDPS | 8 |
| 2013 | Performance Analysis of the Lattice Boltzmann Model Beyond Navier-StokesabstractThe lattice Boltzmann method is increasingly important in facilitating large-scale fluid dynamics simulations. To date, these simulations have been built on discretized velocity models of up to 27 neighbors. Recent work has shown that higher order approximations of the continuum Boltzmann equation enable not only recovery of the Navier-Stokes hydrodynamics, but also simulations for a wider range of Knudsen numbers, which is especially important in micro- and nanoscale flows. These higher-order models have significant impact on both the communication and computational complexity of the application. We present a performance study of the higher-order models as compared to the traditional ones, on both the IBM Blue Gene/P and Blue Gene/Q architectures. We study the tradeoffs of many optimizations methods such as the use of deep halo level ghost cells that, alongside hybrid programming models, reduce the impact of extended models and enable efficient modeling of extreme regimes of computational fluid dynamics. Amanda Randles, Vivek Kale, Jeff R. Hammond, William Gropp, Efthimios Kaxiras |
IPDPS | 1 |
| 2013 | Massively Parallel Model of Extended Memory Use in Evolutionary Game DynamicsabstractTo study the emergence of cooperative behavior, we have developed a scalable parallel framework for evolutionary game dynamics. This is a critical computational tool enabling large-scale agent simulation research. An important aspect is the amount of history, or memory steps, that each agent can keep. When six memory steps are taken into account, the strategy space spans 24096 potential strategies, requiring large populations of agents. We introduce a multi-level decomposition method that allows us to exploit both multi-node and thread-level parallel scaling while minimizing communication overhead. We present the results of a production run modeling up to six memory steps for populations consisting of up to 1018 agents, making this study one of the largest yet undertaken. The high rate of mutation within the population results in a non-trivial parallel implementation. The strong and weak scaling studies provide insight into parallel scalability and programmability trade-offs for large-scale simulations, while exhibiting near perfect weak and strong scaling on 16,384 tasks on Blue Gene/Q. We further show 99% weak scaling up to 294,912 processors 82% strong scaling efficiency up to 262,144 processors of Blue Gene/P. Our framework marks an important step in the study of game dynamics with potential applications in fields ranging from biology to economics and sociology. Amanda Randles, David G. Rand, J. Gregory Morrisett, Jayanta Sircar, Martin A. Nowak, Hanspeter Pfister |
IPDPS | 1 |
| 2011 | Evaluation of Artery Visualizations for Heart Disease DiagnosisabstractHeart disease is the number one killer in the United States, and finding indicators of the disease at an early stage is critical for treatment and prevention. In this paper we evaluate visualization techniques that enable the diagnosis of coronary artery disease. A key physical quantity of medical interest is endothelial shear stress (ESS). Low ESS has been associated with sites of lesion formation and rapid progression of disease in the coronary arteries. Having effective visualizations of a patient's ESS data is vital for the quick and thorough non-invasive evaluation by a cardiologist. We present a task taxonomy for hemodynamics based on a formative user study with domain experts. Based on the results of this study we developed HemoVis, an interactive visualization application for heart disease diagnosis that uses a novel 2D tree diagram representation of coronary artery trees. We present the results of a formal quantitative user study with domain experts that evaluates the effect of 2D versus 3D artery representations and of color maps on identifying regions of low ESS. We show statistically significant results demonstrating that our 2D visualizations are more accurate and efficient than 3D representations, and that a perceptually appropriate color map leads to fewer diagnostic mistakes than a rainbow color map. Michelle Borkin, Krzysztof Z. Gajos, Amanda Randles, Dimitrios Mitsouras, Simone Melchionna, Frank J. Rybicki, Charles L. Feldman, Hanspeter Pfister |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2010 | Multiscale Simulation of Cardiovascular flows on the IBM Bluegene/P: Full Heart-Circulation System at Red-Blood Cell ResolutionabstractWe present the first large-scale simulation of blood flow in the coronary artieries and other vessels supplying blood to the heart muscle, with a realistic description of human arterial geometry at spatial resolutions from centimeters down to 10 microns (near the size of red blood cells). This multiscale simulation resolves the fluid into a billion volume units, embedded in a bounding space of 300 billion voxels, coupled with the concurrent motion of 300 million red blood cells, which interact with one another and with the surrounding fluid. The level of detail is sufficient to describe phenomena of potential physiological and clinical significance, such as the development of atherosclerotic plaques. The simulation achieves excellent scalability on up to 294, 912 Blue Gene/P computational cores. Amanda Randles, Simone Melchionna, Efthimios Kaxiras, Jonas Lätt, Joy K. Sircar, Massimo Bernaschi, Mauro Bisson, Sauro Succi |
SC | 1 |
| 2008 | Asynchronous task dispatch for high throughput computing for the eServer IBM Blue Gene® SupercomputerabstractHigh Throughput Computing (HTC) environments strive "to provide large amounts of processing capacity to customers over long periods of time by exploiting existing resources on the network" according to Basney and Livny [1]. A single Blue Gene/L rack can provide thousands of CPU resources into HTC environments. This paper discusses the implementation of an asynchronous task dispatch system that exploits a recently released feature of the Blue Gene/L control system - called HTC mode - and presents data on experimental runs consisting of the asynchronous submission of multiple batches of thousands of tasks for financial workloads. The methodology developed here demonstrates how systems with very large processor counts and light-weight kernels can be configured to deliver capacity computing at the individual processor level in future petascale computing systems. Amanda Randles, Alan King, Tom Budnik, Paul McCarthy, Pat Michaud, Mike Mundy, Jim Sexton, Greg Stewart |
IPDPS | 1 |
| 2008 | An Efficient Parallel Implementation of the Hidden Markov Methods for Genomic Sequence-Search on a Massively Parallel SystemabstractBioinformatics databases used for sequence comparison and sequence alignment are growing exponentially. This has popularized programs that carry out database searches. Current implementations of sequence alignment methods based on hidden Markov models (HMM) have proven to be computationally intensive and, hence, amenable to architectures with multiple processors. In this paper, we describe a modified version of the original parallel implementation of HMMs on a massively parallel system. This is part of the HMMER bioinformatics code. HMMER 2.3.2 uses profile HMMs for sensitive database searching based on statistical descriptions of a sequence family's consensus (Durbin et al., 1998), Two of the nine programs were further parallelized to take advantage of the large number of processors, namely, hmmsearch and hmmpfam. For our study, we start by porting the parallel virtual machine (PVM) versions of these two programs currently available as part of the HMMER suite of programs. We report the performance of these nonoptimized versions as baselines. Our work also includes the introduction of an alternate sequence file indexing, multiple-master configuration, dynamic data collection and, finally, load balancing via the indexed sequence files. This set of optimizations constitutes our modified version for massively parallel systems. Our results show parallel performance improvements of more than one order of magnitude (16 times) for hmmsearch and hmmpfam. Karl Jiang, Oystein Thorsen, Amanda Randles, Brian E. Smith, Carlos P. Sosa |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2006 | Poster reception - Optimizing EUDOC for the IBM eServer Blue Gene supercomputerabstractThe EUDOC application code was ported to and optimized for the IBM eServer Blue Gene (BG/L) supercomputer. EUDOC is a molecular docking program that has shown success predicting drug-bound protein complexes and identifying new drug leads. Single node performance was optimized to obtain a 4X improvement by handtuning critical sections of code. Many of the techniques used are applicable to other applications running on BG/L. Load balancing schemes were studied to maximize processor utilization and scalability of the application. A static load balance scheme achieves 40% processor utilization, while a head-tailtail scheme that takes better advantage of the high-speed interconnect on BG/L peaked above 98%. Performance results are shown using 512 and 2048 processors on two different drug targets with a database of 65536 "drug-like" molecules. The results suggest linear scaling. We will show how EUDOC was optimized to obtain this speedup, and present the load balancing results. Yuan-Ping Pang, Brent A. Swartz, Brian E. Smith, Timothy J. Mullins, Amanda Randles, Roy G. Musselman |
SC | 5 |