EDBT 2026 Demo / reviewers in the wild / expert
Sadaf R. Alam
dblp:62/3773
· DBLP profile ↗
28ranked-venue papers
19as first author
1since 2021 · last 2022
0000-0002-2534-5078ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 17 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
High-performance computing · 29% Performance modeling and evaluation · 29% Cloud and datacenter computing · 18% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 11 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.3 | 1 | 2018 | RM-replay: a high-fidelity tuning, optimization and exploration tool for resource management · SC 2018 |
Hardware accelerators and domain-specific architectures › accelerator integration
accelerator interconnect |
0.2 | 1 | 2016 | A PCIe congestion-aware performance model for densely populated accelerator servers · SC 2016 |
Performance modeling and evaluation
benchmarking |
0.2 | 3 | 2008 | Early evaluation of IBM BlueGene/P · SC 2008 Cray XT4: an early evaluation for petascale scientific simulation · SC 2007 Performance evaluation of the cray XT3 configured with dual core opteron processors · PPoPP 2007 |
High-performance computing › supercomputing
supercomputing systems |
0.2 | 2 | 2008 | Early evaluation of IBM BlueGene/P · SC 2008 Performance evaluation of the cray XT3 configured with dual core opteron processors · PPoPP 2007 |
High-performance computing › scientific computing systems
biomolecular simulation |
0.1 | 1 | 2010 | Optimal Utilization of Heterogeneous Resources for Biomolecular Simulations · SC 2010 |
GPUs and heterogeneous computing
multi-GPU computing |
0.1 | 1 | 2010 | Optimal Utilization of Heterogeneous Resources for Biomolecular Simulations · SC 2010 |
Performance modeling and evaluation › simulation
simulation-based evaluation |
0.1 | 1 | 2018 | RM-replay: a high-fidelity tuning, optimization and exploration tool for resource management · SC 2018 |
High-performance computing › scientific computing systems
molecular dynamics simulation |
0.1 | 1 | 2006 | Performance characterization of molecular dynamics techniques for biomolecular simulations · PPoPP 2006 |
High-performance computing
scientific computing |
0.1 | 1 | 2006 | Performance characterization of molecular dynamics techniques for biomolecular simulations · PPoPP 2006 |
Distributed systems › communication optimization
communication-computation overlap |
0.0 | 1 | 2010 | Optimal Utilization of Heterogeneous Resources for Biomolecular Simulations · SC 2010 |
Bioinformatics and computational biology › molecular informatics › molecular modeling
biomolecular simulation |
0.0 | 1 | 2006 | Performance characterization of molecular dynamics techniques for biomolecular simulations · PPoPP 2006 |
Methods — techniques the papers use, named apart from their topics
replay-based tuning · 0.3optimization · 0.3congestion-aware modeling · 0.2congestion graph · 0.2microbenchmarks · 0.2application benchmarks · 0.2pipelining · 0.1parametric study · 0.1overlapping · 0.1performance evaluation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | IEEE Special Issue on Innovative R&D Toward the Exascale EraabstractThis special issue on Innovative research and development toward the exascale era explores new foundational and translational research toward enabling exascale computing for emerging scientific and societal challenges. Exascale computing is defined as the capability to perform 1018 operations per second. Productively harnessing such a scale of processing, storage, and networking capabilities for diverse domains— including high-performance computing (HPC) simulations, artificial intelligence (AI), and extreme data-driven computing— relies on not only revitalizing existing parallel and distributed computing technologies but also innovating new solutions. Papers in this special issue explore diverse topics in research encompassing parallel, distributed, and heterogeneous systems for exascale, including advances in applications, programming environments, runtimes, libraries, innovative algorithms, domain-specific frameworks, systems architecture, performance analysis, data processing, and networking technologies. Sadaf R. Alam, Lois C. McInnes, Kengo Nakajima |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | RM-replay: a high-fidelity tuning, optimization and exploration tool for resource management
Maxime Martinasso, Miguel Gila, Mauro Bianco, Sadaf R. Alam, Colin McMurtrie, Thomas C. Schulthess |
SC | 4 |
| 2016 | A PCIe congestion-aware performance model for densely populated accelerator serversabstractMeteoSwiss, the Swiss national weather forecast institute, has selected densely populated accelerator servers as their primary system to compute weather forecast simulation. Servers with multiple accelerator devices that are primarily connected by a PCI-Express (PCIe) network achieve a significantly higher energy efficiency. Memory transfers between accelerators in such a system are subjected to PCIe arbitration policies. In this paper, we study the impact of PCIe topology and develop a congestion-aware performance model for PCIe communication. We present an algorithm for computing congestion factors of every communication in a congestion graph that characterizes the dynamic usage of network resources by an application. Our model applies to any PCIe tree topology. Our validation results on two different topologies of 8 GPU devices demonstrate that our model achieves an accuracy of over 97% within the PCIe network. We demonstrate the model on a weather forecast application to identify the best algorithms for its communication patterns among GPUs. Maxime Martinasso, Grzegorz Kwasniewski, Sadaf R. Alam, Thomas C. Schulthess, Torsten Hoefler |
SC | 3 |
| 2013 | Evaluation of Inter- and Intra-node Data Transfer Efficiencies between GPU Devices and their Impact on Scalable ApplicationsabstractData movement is of high relevance for GPU Computing. Communication and performance efficiencies of applications and systems with GPU accelerators depend on on- and off-node data paths, thereby making tuning and optimization an increasingly complex task. In this paper we conduct an in-depth study to establish the parameters that influence performance of data transfers between on-node GPU devices, and located on separate nodes (off-node). We compare the most recent version of MVAPICH2 featuring seamless remote GPU transfers with our own low-level benchmarks, and discuss the bottlenecks that may arise. Data path performance and bottlenecks between GPU devices are analyzed and compared for two substantially different systems: an IBM datable relying on an InfiniBand QDR fabric with two on-node GPU devices, and a Cray XK6, featuring a single GPU per node, and connected through a Gemini interconnect. Finally, we adapt LAMMPS, a GPU-accelerated application, to benefit from efficient inter-GPU data transfers, and validate our findings. Antonio J. Peña, Sadaf R. Alam |
CCGRID | 2 |
| 2013 | Performance modeling of microsecond scale biological molecular dynamics simulations on heterogeneous architecturesabstractSUMMARY Performance improvements in biomolecular simulations based on molecular dynamics (MD) codes are widely desired. Unfortunately, the factors, which allowed past performance improvements, particularly the microprocessor clock frequencies, are no longer increasing. Hence, novel software and hardware solutions are being explored for accelerating performance of widely used MD codes. In this paper, we describe our efforts on porting, optimizing and tuning of Large‐scale Atomic/Molecular Massively Parallel Simulator, a popular MD framework, on heterogeneous architectures: multi‐core processors with graphical processing unit (GPU) accelerators. Our implementation is based on accelerating the most computationally expensive non‐bonded interaction terms on the GPUs and overlapping the computation on the CPU and GPUs. This functionality is built on top of message passing interface that allows multi‐level parallelism to be extracted even at the workstation level with the multi‐core CPUs and allows extension of the implementation on GPU‐enabled clusters. We hypothesize that the optimal benefit of heterogeneous architectures for applications will come by utilizing all possible resources (for example, CPU‐cores and GPU devices on GPU‐enabled clusters). Benchmarks for a range of biomolecular system sizes are provided, and an analysis is performed on four generations of NVIDIA's GPU devices. On GPU‐enabled Linux clusters, by overlapping and pipelining computation and communication, we observe up to 10‐folds application acceleration in multi‐core and multi‐GPU environments illustrating significant performance improvements. Detailed analysis of the implementation is presented that allows identification of bottlenecks in algorithm, indicating that code optimization and improvements on GPUs could allow microsecond scale simulation throughput on workstations and inexpensive GPU clusters, putting widely desired biologically relevant simulation time‐scales within reach of a large user community. In order to systematically optimize simulation throughput and to enable performance prediction, we have developed a parameterized performance model that will allow developers and users to explore the performance potential of future heterogeneous systems for biological simulations. Copyright © 2012 John Wiley & Sons, Ltd. Pratul K. Agarwal, Scott S. Hampton, Jeffrey D. Poznanovic, Arvind Ramanathan, Sadaf R. Alam, Paul S. Crozier |
Concurr. Comput. Pract. Exp. | 5 |
| 2010 | Optimal Utilization of Heterogeneous Resources for Biomolecular SimulationsabstractBiomolecular simulations have traditionally benefited from increases in the processor clock speed and coarse-grain inter-node parallelism on large-scale clusters. With stagnating clock frequencies, the evolutionary path for performance of microprocessors is maintained by virtue of core multiplication. Graphical processing units (GPUs) offer revolutionary performance potential at the cost of increased programming complexity. Furthermore, it has been extremely challenging to effectively utilize heterogeneous resources (host processor and GPU cores) for scientific simulations, as underlying systems, programming models and tools are continually evolving. In this paper, we present a parametric study demonstrating approaches to exploit resources of heterogeneous systems to reduce time-to-solution of a production-level application for biological simulations. By overlapping and pipelining computation and communication, we observe up to 10-fold application acceleration in multi-core and multi-GPU environments illustrating significant performance improvements over code acceleration approaches, where the host-to-accelerator ratio is static, and is constrained by a given algorithmic implementation. Scott S. Hampton, Sadaf R. Alam, Paul S. Crozier, Pratul K. Agarwal |
SC | 2 |
| 2009 | Impact of Quad-Core Cray XT4 System and Software Stack on Scientific Computation
Sadaf R. Alam, Richard F. Barrett, Heike Jagode, Jeffery A. Kuehn, Stephen W. Poole, Ramanan Sankaran |
Euro-Par | 1 |
| 2009 | Performance Characterization of a Hierarchical MPI Implementation on Large-scale Distributed-memory PlatformsabstractThe building blocks of emerging Petascale massively parallel processing (MPP) systems are multi-core processors with four or more cores as a single processing element and a customized network interface. The resulting memory and communication hierarchy of these platforms are now exposed to application developers and end users by creating a hierarchical or multi-core aware message-passing (MPI) programming interface and by providing a handful of runtime, tunable parameters that allows mapping and control of MPI tasks and message handling. We characterize performance of MPI communication patterns and present strategies for optimizing applications performance on Cray XT series systems that are composed of contemporary AMD processors and a proprietary network infrastructure. We highlight dependencies in its memory and network subsystems, which could influence production-level applications performance. We demonstrate that MPI micro-benchmarks could mislead an application developer or end user since these benchmarks often do not expose the interplay between memory allocation and usage in the user space, which depends on the number of tasks or cores and workload characteristics. Our studies show performance improvements compared to the default options for our target scientific benchmarks and production-level applications. Sadaf R. Alam, Richard F. Barrett, Jeffery A. Kuehn, Stephen W. Poole |
ICPP | 1 |
| 2009 | Performance analysis and projections for Petascale applications on Cray XT series systemsabstractThe Petascale Cray XT5 system at the Oak Ridge National Laboratory (ORNL) Leadership Computing Facility (LCF) shares a number of system and software features with its predecessor, the Cray XT4 system including the quad-core AMD processor and a multi-core aware MPI library. We analyze performance of scalable scientific applications on the quad-core Cray XT4 system as part of the early system access using a combination of micro-benchmarks and Petascale ready applications. Particularly, we evaluate impact of key changes that occurred during the dual-core to quad-core processor upgrade on applications behavior and provide projections for the next-generation massively-parallel platforms with multi-core processors, specifically for proposed Petascale Cray XT5 system. We compare and contrast the quad-core XT4 system features with the upcoming XT5 system and discuss strategies for improving scaling and performance for our target applications. Sadaf R. Alam, Richard F. Barrett, Jeffery A. Kuehn, Stephen W. Poole |
IPDPS | 1 |
| 2008 | Experimental Evaluation of Molecular Dynamics Simulations on Multi-core Systems
Sadaf R. Alam, Pratul K. Agarwal, Scott S. Hampton, Hong Ong |
HiPC | 1 |
| 2008 | Impact of multicores on large-scale molecular dynamics simulationsabstractProcessing nodes of the Cray XT and IBM Blue Gene Massively Parallel Processing (MPP) systems are composed of multiple execution units, sharing memory and network subsystems. These multicore processors offer greater computational power, but may be hindered by resource contention. In order to understand and avoid such situations, we investigate the impact of resource contention on three scalable molecular dynamics suites: AMBER (PMEMD module), LAMMPS, and NAMD. The results reveal the factors that can inhibit scaling and performance efficiency on emerging multicore processors. Sadaf R. Alam, Pratul K. Agarwal, Scott S. Hampton, Hong Ong, Jeffrey S. Vetter |
IPDPS | 1 |
| 2008 | A Methodology for Developing High Fidelity Communication Models for Large-Scale Applications Targeted on Multicore SystemsabstractResource sharing and implementation of software stack for emerging multicore processors introduce performance and scaling challenges for large-scale scientific applications, particularly on systems with thousands of processing elements. Traditional performance optimization, tuning and modeling techniques that rely on uniform representation of computation and communication requirements are only partially useful due to the complexity of applications and underlying systems and software architecture. In this paper, we propose a workload modeling methodology that allows application developers to capture and represent hierarchical decomposition and distribution of their applications thereby allowing them to explore and identify optimal mapping of a workload on a target system. We demonstrate the proposed methodology on a Teraflopsscale fusion application that is developed using message-passing (MPI) programming paradigm. Using our analysis and projection results, we obtain insight into the performance characteristics of the application on a quad-core system and also identify optimal mapping on a Teraflops-scale platform. 1. Charles W. Lively, Valerie Taylor 0001, Sadaf R. Alam, Jeffrey S. Vetter |
SBAC-PAD | 3 |
| 2008 | Early evaluation of IBM BlueGene/PabstractBlueGene/P (BG/P) is the second generation BlueGene architecture from IBM, succeeding BlueGene/L (BG/L). BG/P is a system-on-a-chip (SoC) design that uses four PowerPC 450 cores operating at 850 MHz with a double precision, dual pipe floating point unit per core. These chips are connected with multiple interconnection networks including a 3-D torus, a global collective network, and a global barrier network. The design is intended to provide a highly scalable, physically dense system with relatively low power requirements per flop. In this paper, we report on our examination of BG/P, presented in the context of a set of important scientific applications, and as compared to other major large scale supercomputers in use today. Our investigation confirms that BG/P has good scalability with an expected lower performance per processor when compared to the Cray XT4's Opteron. We also find that BG/P uses very low power per floating point operation for certain kernels, yet it has less of a power advantage when considering science driven metrics for mission applications. Sadaf R. Alam, Richard F. Barrett, M. Bast, Mark R. Fahey, Jeffery A. Kuehn, Collin McCurdy, James H. Rogers, Philip C. Roth, Ramanan Sankaran, Jeffrey S. Vetter, Patrick H. Worley, Weikuan Yu |
SC | 1 |
| 2008 | Performance characteristics of biomolecular simulations on high-end systems with multi-core processors
Sadaf R. Alam, Pratul K. Agarwal, Jeffrey S. Vetter |
Parallel Comput. | 1 |
| 2007 | An Application Specific Memory Characterization Technique for Co-processor AcceleratorsabstractCommodity accelerator technologies including reconfigurable devices provide an order of magnitude performance improvement compared to mainstream microprocessor systems. A number of compute-intensive scientific applications, therefore, can potentially benefit from commodity computing devices available in the form of co-processor accelerators. However, there has been little progress in accelerating production-level scientific applications using these technologies due to several programming and performance challenges. One of the key perfomance challenges is performance sustainability. While computation is often accelerated substantially by accelerator devices, the achievable performance is significantly lower once the data transfer costs and overheads are incorporated. We present an application-specific memory characterization technique for an FPGA-accelerated system that enabled us to reduce data transfer overhead by a factor of five for a production-scale scientific application. Our proposed technique extends to applications that exhibit similar memory behavior and to co-processor accelerator systems that support data streaming, pipelining, and overlapped execution. Sadaf R. Alam, Jeffrey S. Vetter, Melissa C. Smith |
ASAP | 1 |
| 2007 | Performance Evaluation of a Scalable Molecular Dynamics Simulation Framework on a Massively-Parallel SystemabstractThe successors of distributed-memory, massively-parallel processing (MPP) systems that are based on multi-core processor technologies and high-bandwidth communication networks are expected to deliver Petascale computing power for scientific communities in near future. This report presents preliminary performance evaluation and benchmarking results of a scalable biomolecular simulation framework on the Cray XT4 MPP system that contains multi-core Opteron processors. We identify not only the performance enhancing features but also the bottlenecks for biomolecular simulation test cases on this system using a combination of application and vendor specific performance tools. Our results show that unprecedented performance has been achieved for large-scale test cases on the system; however, the critical challenges remain for longer time scale simulations on MPP systems. Sadaf R. Alam, Pratul K. Agarwal, Jeffery A. Kuehn |
BIBE | 1 |
| 2007 | Sensitivity Analysis of Biomolecular Simulations using Symbolic ModelsabstractPerformance and scaling of biomolecular simulations frameworks largely depends on not only the workload characteristics of the simulations but also the design of underlying processor architecture and interconnection networks. Because construction of Teraflops and Petaflops scale prototype systems for evaluation alone is impractical and cost-prohibitive, architects use analytical models of workloads and architecture simulators to guide their design decisions and tradeoffs. To address the problem of providing scalable yet precise input for network simulators, we have developed a technique to model symbolically the communication patterns of production-level scientific applications to study workload growth rates and to carry out sensitivity analysis. We apply our symbolic modeling scheme to the particle mesh ewald (PME) implementation in the sander package of the AMBER framework and demonstrate how the increase in computation, memory and communication requirements impact the performance and scaling of the PME method on the next-generation massively-parallel systems. Sadaf R. Alam, Nikhil Bhatia, Jeffrey S. Vetter |
BIBE | 1 |
| 2007 | Balancing productivity and performance on the cell broadband engineabstractThe cell broadband engine (BE) is a heterogeneous multicore processor, combining a general-purpose POWER architecture core with eight independent single-instruction-multiple-data (SIMD) cores. Each core is capable of very high performance; however, users must explicitly manage data movement, scheduling, and synchronization. While these attributes provide some of the cell processorpsilas greatest performance strengths, they also form its greatest weaknesses in terms of developer productivity, code portability, and initial performance efficiencies. In this paper, we evaluate productivity and relative performance improvements of a cell BE system for a diverse set of kernels and applications. Our experimental workload includes algorithms from scientific, cognitive, and imaging problem domains. Our results demonstrate that the cell processor could be several times faster than a SSE-enabled, contemporary dual-core processor, and could sustain a high performance-to-productivity ratio. We outline strategies for transforming applications to exploit the cellpsilas architectural features, and measure productivity by comparing programming effort in terms of lines of code and performance. For instance, our measurements revealed that a covariance matrix creation routine - a common routine in hyperspectral imaging - ran over eight times faster than a 2.66 GHz Intel Woodcrest processor while sustaining a productivity metric of over two by parallelizing across the heterogeneous cores, unrolling loops, and improving instruction level parallelism with SIMD instructions in a high-level language. Sadaf R. Alam, Jeremy S. Meredith, Jeffrey S. Vetter |
CLUSTER | 1 |
| 2007 | An Exploration of Performance Attributes for Symbolic Modeling of Emerging Processing Devices
Sadaf R. Alam, Nikhil Bhatia, Jeffrey S. Vetter |
HPCC | 1 |
| 2007 | On the Path to Enable Multi-scale Biomolecular Simulations on PetaFLOPS Supercomputer with Multi-core ProcessorsabstractBiological processes occurring inside cell involve multiple scales of time and length; many popular theoretical and computational multi-scale techniques utilize biomolecular simulations based on molecular dynamics. Till recently, the computing power required for simulating the relevant scales was even beyond the reach of fastest supercomputers. The availability of petaFLOPS-scale computing power in near future holds great promise. Unfortunately, the bio-simulations software technology has not kept up with the changes in hardware. In particular, with the introduction of multi-core processing technologies in systems with tens of thousands of processing cores, it is unclear whether the existing biomolecular simulation frameworks will be able to scale and to utilize these resources effectively. While the multi-core processing systems provide higher processing capabilities, their memory and IO subsystems are posing new challenges to application and system software developers. In this preliminary study, we attempt to characterize computation, communication and memory efficiencies of bio-molecular simulations on a Cray XT3 system, which has recently been upgraded to dual-core Opteron processors. We identify that the application efficiencies using the multi-core processors reduce with the increase of the simulated system size. Further, we measure the communication overhead of using both cores in the processor simultaneously and identify that the MPI communication performance can be as low as 50% as compared to the single-core execution times. We conclude that not only the biomolecular simulations need to be aware of the underlying multi-core hardware in order to achieve maximum performance but also the system software needs to provide processor and memory placement features in the high-end systems. Our results on a stand-alone dual-core AMD system confirm that combinations of processor and memory affinity schemes can result in over 12% performance gains. Sadaf R. Alam, Pratul K. Agarwal |
IPDPS | 1 |
| 2007 | Analysis of a Computational Biology Simulation Technique on Emerging Processing ArchitecturesabstractMulti-paradigm, multi-threaded and multi-core computing devices available today provide several orders of magnitude performance improvement over mainstream microprocessors. These devices include the STI Cell Broadband Engine, graphical processing units (GPU) and the Cray massively-multithreaded processors - available in desktop computing systems as well as proposed for supercomputing platforms. The main challenge in utilizing these powerful devices is their unique programming paradigms. GPUs and the Cell systems require code developers to manage code and data explicitly, while the Cray multithreaded architecture requires them to generate a very large number of threads or independent tasks concurrently. In this paper, we explain strategies for optimizing a molecular dynamics (MD) calculation that is used in biomolecular simulations on three devices: Cell, GPU and MTA-2. We show that the Cray MTA-2 system requires minimal code modification and does not outperform the microprocessor runs; but it demonstrates an improved workload scaling behavior over the microprocessor implementation. On the other hand, substantial porting and optimization efforts on the Cell and the GPU systems result in a 5times to 6times improvement, respectively, over a 2.2 GHz Opteron system. Jeremy S. Meredith, Sadaf R. Alam, Jeffrey S. Vetter |
IPDPS | 2 |
| 2007 | Performance evaluation of the cray XT3 configured with dual core opteron processorsabstractNo abstract available. Richard F. Barrett, Sadaf R. Alam, Jeffrey S. Vetter |
PPoPP | 2 |
| 2007 | Cray XT4: an early evaluation for petascale scientific simulationabstractThe scientific simulation capabilities of next generation high-end computing technology will depend on striking a balance among memory, processor, I/O, and local and global network performance across the breadth of the scientific simulation space. The Cray XT4 combines commodity AMD dual core Opteron processor technology with the second generation of Cray's custom communication accelerator in a system design whose balance is claimed to be driven by the demands of scientific simulation. This paper presents an evaluation of the Cray XT4 using micro-benchmarks to develop a controlled understanding of individual system components, providing the context for analyzing and comprehending the performance of several petascale-ready applications. Results gathered from several strategic application domains are compared with observations on the previous generation Cray XT3 and other high-end computing systems, demonstrating performance improvements across a wide variety of application benchmark problems. Sadaf R. Alam, Jeffery A. Kuehn, Richard F. Barrett, Jeffrey M. Larkin, Mark R. Fahey, Ramanan Sankaran, Patrick H. Worley |
SC | 1 |
| 2006 | Hierarchical Model Validation of Symbolic Performance Models of Scientific Kernels
Sadaf R. Alam, Jeffrey S. Vetter |
Euro-Par | 1 |
| 2006 | An Analysis of System Balance Requirements for Scientific ApplicationsabstractScientific applications are diverse in terms of the resource requirements, and tend to vary significantly from commercial applications. In order to provide sustained performance, a target high performance computing (HPC) platform must offer a balance between CPU performance to memory, interconnect and I/O subsystems performance. We characterize the system balance requirements for two large-scale Office of Science applications, GYRO (fusion simulation) and POP (climate modeling), and develop platform-independent parameterized requirement models. We measure the parallel efficiencies for GYRO and POP on three multiprocessor systems: an SMP cluster (IBM p690), a shared-memory system (SGI Altix) and a vector supercomputer (Cray XI). The higher computational intensity and interconnect bandwidth requirements of GYRO result in higher performance efficiencies on the vector platform. At the same time, small message sizes in POP benefit from low MPI latencies of the shared-memory platform. Overall results confirm system balance requirements that are generated by the requirement models Sadaf R. Alam, Jeffrey S. Vetter |
ICPP | 1 |
| 2006 | A framework to develop symbolic performance models of parallel applicationsabstractPerformance and workload modeling has numerous uses at every stage of the high-end computing lifecycle: design, integration, procurement, installation and tuning. Despite the tremendous usefulness of performance models, their construction remains largely a manual, complex, and time-consuming exercise. We propose a new approach to the model construction, called modeling assertions (MA), which borrows advantages from both the empirical and analytical modeling techniques. This strategy has many advantages over traditional methods: incremental construction of realistic performance models, straightforward model validation against empirical data, and intuitive error bounding on individual model terms. We demonstrate this new technique on the NAS parallel CG and SP benchmarks by constructing high fidelity models for the floating-point operation cost, memory requirements, and MPI message volume. These models are driven by a small number of key input parameters thereby allowing efficient design space exploration of future problem sizes and architectures Sadaf R. Alam, Jeffrey S. Vetter |
IPDPS | 1 |
| 2006 | Early evaluation of the Cray XT3abstractOak Ridge National Laboratory recently received delivery of a 5,294 processor Cray XT3. The XT3 is Cray's third-generation massively parallel processing system. The system builds on a single processor node - built around the AMD Opteron - and uses a custom chip - called SeaStar - to provide interprocess or communication. In addition, the system uses a lightweight operating system on the compute nodes. This paper describes our initial experiences with the system, including micro-benchmark, kernel, and application benchmark results. In particular, we provide performance results for strategic Department of Energy applications areas including climate and fusion. We demonstrate experiments on the installed system, scaling applications up to 4,096 processors. Jeffrey S. Vetter, Sadaf R. Alam, Thomas H. Dunigan, Mark R. Fahey, Philip C. Roth, Patrick H. Worley |
IPDPS | 2 |
| 2006 | Performance characterization of molecular dynamics techniques for biomolecular simulationsabstractLarge-scale simulations and computational modeling using molecular dynamics (MD) continues to make significant impacts in the field of biology. It is well known that simulations of biological events at native time and length scales requires computing power several orders of magnitude beyond today's commonly available systems. Supercomputers, such as IBM Blue Gene/L and Cray XT3, will soon make tens to hundreds of teraFLOP/s of computing power available by utilizing thousands of processors. The popular algorithms and MD applications, however, were not initially designed to run on thousands of processors. In this paper, we present detailed investigations of the performance issues, which are crucial for improving the scalability of the MD-related algorithms and applications on massively parallel processing (MPP) architectures. Due to the varying characteristics of biological input problems, we study two prototypical biological complexes that use the MD algorithm: an explicit solvent and an implicit solvent. In particular, we study the AMBER application, which supports a variety of these types of input problems. For the explicit solvent problem, we focused on the particle mesh Ewald (PME) method for calculating the electrostatic energy, and for the implicit solvent model, we targeted the Generalized Born (GB) calculation. We uncovered and subsequently modified a limitation in AMBER that restricted the scaling beyond 128 processors. We collected performance data for experiments on up to 2048 Blue Gene/L and XT3 processors and subsequently identified that the scaling is largely limited by the underlying algorithmic characteristics and also by the implementation of the algorithms. Furthermore, we found that the input problem size of biological system is constrained by memory available per node. In conclusion, our results indicate that MD codes can significantly benefit from the current generation architectures with relatively modest optimization efforts. Nevertheless, the key for enabling scientific breakthroughs lies in exploiting the full potential of these new architectures. Sadaf R. Alam, Jeffrey S. Vetter, Pratul K. Agarwal, Al Geist |
PPoPP | 1 |