Sadaf R. Alam

dblp:62/3773 · DBLP profile ↗
← Back
28ranked-venue papers
19as first author
1since 2021 · last 2022
0000-0002-2534-5078ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 17 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
High-performance computing · 29% Performance modeling and evaluation · 29% Cloud and datacenter computing · 18%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 11 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.312018
RM-replay: a high-fidelity tuning, optimization and exploration tool for resource management · SC 2018
Hardware accelerators and domain-specific architectures › accelerator integration
accelerator interconnect
0.212016
A PCIe congestion-aware performance model for densely populated accelerator servers · SC 2016
Performance modeling and evaluation
benchmarking
0.232008
Early evaluation of IBM BlueGene/P · SC 2008
Cray XT4: an early evaluation for petascale scientific simulation · SC 2007
Performance evaluation of the cray XT3 configured with dual core opteron processors · PPoPP 2007
High-performance computing › supercomputing
supercomputing systems
0.222008
Early evaluation of IBM BlueGene/P · SC 2008
Performance evaluation of the cray XT3 configured with dual core opteron processors · PPoPP 2007
High-performance computing › scientific computing systems
biomolecular simulation
0.112010
Optimal Utilization of Heterogeneous Resources for Biomolecular Simulations · SC 2010
GPUs and heterogeneous computing
multi-GPU computing
0.112010
Optimal Utilization of Heterogeneous Resources for Biomolecular Simulations · SC 2010
Performance modeling and evaluation › simulation
simulation-based evaluation
0.112018
RM-replay: a high-fidelity tuning, optimization and exploration tool for resource management · SC 2018
High-performance computing › scientific computing systems
molecular dynamics simulation
0.112006
Performance characterization of molecular dynamics techniques for biomolecular simulations · PPoPP 2006
High-performance computing
scientific computing
0.112006
Performance characterization of molecular dynamics techniques for biomolecular simulations · PPoPP 2006
Distributed systems › communication optimization
communication-computation overlap
0.012010
Optimal Utilization of Heterogeneous Resources for Biomolecular Simulations · SC 2010
Bioinformatics and computational biology › molecular informatics › molecular modeling
biomolecular simulation
0.012006
Performance characterization of molecular dynamics techniques for biomolecular simulations · PPoPP 2006

Methods — techniques the papers use, named apart from their topics

replay-based tuning · 0.3optimization · 0.3congestion-aware modeling · 0.2congestion graph · 0.2microbenchmarks · 0.2application benchmarks · 0.2pipelining · 0.1parametric study · 0.1overlapping · 0.1performance evaluation · 0.1
YearPublicationVenuePosition
2022 IEEE Special Issue on Innovative R&D Toward the Exascale Era
abstract
This special issue on Innovative research and development toward the exascale era explores new foundational and translational research toward enabling exascale computing for emerging scientific and societal challenges. Exascale computing is defined as the capability to perform 1018 operations per second. Productively harnessing such a scale of processing, storage, and networking capabilities for diverse domains— including high-performance computing (HPC) simulations, artificial intelligence (AI), and extreme data-driven computing— relies on not only revitalizing existing parallel and distributed computing technologies but also innovating new solutions. Papers in this special issue explore diverse topics in research encompassing parallel, distributed, and heterogeneous systems for exascale, including advances in applications, programming environments, runtimes, libraries, innovative algorithms, domain-specific frameworks, systems architecture, performance analysis, data processing, and networking technologies.
Sadaf R. Alam, Lois C. McInnes, Kengo Nakajima
IEEE Trans. Parallel Distributed Syst.1
2018 RM-replay: a high-fidelity tuning, optimization and exploration tool for resource management
Maxime Martinasso, Miguel Gila, Mauro Bianco, Sadaf R. Alam, Colin McMurtrie, Thomas C. Schulthess
SC4
2016 A PCIe congestion-aware performance model for densely populated accelerator servers
abstract
MeteoSwiss, the Swiss national weather forecast institute, has selected densely populated accelerator servers as their primary system to compute weather forecast simulation. Servers with multiple accelerator devices that are primarily connected by a PCI-Express (PCIe) network achieve a significantly higher energy efficiency. Memory transfers between accelerators in such a system are subjected to PCIe arbitration policies. In this paper, we study the impact of PCIe topology and develop a congestion-aware performance model for PCIe communication. We present an algorithm for computing congestion factors of every communication in a congestion graph that characterizes the dynamic usage of network resources by an application. Our model applies to any PCIe tree topology. Our validation results on two different topologies of 8 GPU devices demonstrate that our model achieves an accuracy of over 97% within the PCIe network. We demonstrate the model on a weather forecast application to identify the best algorithms for its communication patterns among GPUs.
Maxime Martinasso, Grzegorz Kwasniewski, Sadaf R. Alam, Thomas C. Schulthess, Torsten Hoefler
SC3
2013 Evaluation of Inter- and Intra-node Data Transfer Efficiencies between GPU Devices and their Impact on Scalable Applications
abstract
Data movement is of high relevance for GPU Computing. Communication and performance efficiencies of applications and systems with GPU accelerators depend on on- and off-node data paths, thereby making tuning and optimization an increasingly complex task. In this paper we conduct an in-depth study to establish the parameters that influence performance of data transfers between on-node GPU devices, and located on separate nodes (off-node). We compare the most recent version of MVAPICH2 featuring seamless remote GPU transfers with our own low-level benchmarks, and discuss the bottlenecks that may arise. Data path performance and bottlenecks between GPU devices are analyzed and compared for two substantially different systems: an IBM datable relying on an InfiniBand QDR fabric with two on-node GPU devices, and a Cray XK6, featuring a single GPU per node, and connected through a Gemini interconnect. Finally, we adapt LAMMPS, a GPU-accelerated application, to benefit from efficient inter-GPU data transfers, and validate our findings.
Antonio J. Peña, Sadaf R. Alam
CCGRID2
2013 Performance modeling of microsecond scale biological molecular dynamics simulations on heterogeneous architectures
abstract
SUMMARY Performance improvements in biomolecular simulations based on molecular dynamics (MD) codes are widely desired. Unfortunately, the factors, which allowed past performance improvements, particularly the microprocessor clock frequencies, are no longer increasing. Hence, novel software and hardware solutions are being explored for accelerating performance of widely used MD codes. In this paper, we describe our efforts on porting, optimizing and tuning of Large‐scale Atomic/Molecular Massively Parallel Simulator, a popular MD framework, on heterogeneous architectures: multi‐core processors with graphical processing unit (GPU) accelerators. Our implementation is based on accelerating the most computationally expensive non‐bonded interaction terms on the GPUs and overlapping the computation on the CPU and GPUs. This functionality is built on top of message passing interface that allows multi‐level parallelism to be extracted even at the workstation level with the multi‐core CPUs and allows extension of the implementation on GPU‐enabled clusters. We hypothesize that the optimal benefit of heterogeneous architectures for applications will come by utilizing all possible resources (for example, CPU‐cores and GPU devices on GPU‐enabled clusters). Benchmarks for a range of biomolecular system sizes are provided, and an analysis is performed on four generations of NVIDIA's GPU devices. On GPU‐enabled Linux clusters, by overlapping and pipelining computation and communication, we observe up to 10‐folds application acceleration in multi‐core and multi‐GPU environments illustrating significant performance improvements. Detailed analysis of the implementation is presented that allows identification of bottlenecks in algorithm, indicating that code optimization and improvements on GPUs could allow microsecond scale simulation throughput on workstations and inexpensive GPU clusters, putting widely desired biologically relevant simulation time‐scales within reach of a large user community. In order to systematically optimize simulation throughput and to enable performance prediction, we have developed a parameterized performance model that will allow developers and users to explore the performance potential of future heterogeneous systems for biological simulations. Copyright © 2012 John Wiley & Sons, Ltd.
Pratul K. Agarwal, Scott S. Hampton, Jeffrey D. Poznanovic, Arvind Ramanathan, Sadaf R. Alam, Paul S. Crozier
Concurr. Comput. Pract. Exp.5
2010 Optimal Utilization of Heterogeneous Resources for Biomolecular Simulations
abstract
Biomolecular simulations have traditionally benefited from increases in the processor clock speed and coarse-grain inter-node parallelism on large-scale clusters. With stagnating clock frequencies, the evolutionary path for performance of microprocessors is maintained by virtue of core multiplication. Graphical processing units (GPUs) offer revolutionary performance potential at the cost of increased programming complexity. Furthermore, it has been extremely challenging to effectively utilize heterogeneous resources (host processor and GPU cores) for scientific simulations, as underlying systems, programming models and tools are continually evolving. In this paper, we present a parametric study demonstrating approaches to exploit resources of heterogeneous systems to reduce time-to-solution of a production-level application for biological simulations. By overlapping and pipelining computation and communication, we observe up to 10-fold application acceleration in multi-core and multi-GPU environments illustrating significant performance improvements over code acceleration approaches, where the host-to-accelerator ratio is static, and is constrained by a given algorithmic implementation.
Scott S. Hampton, Sadaf R. Alam, Paul S. Crozier, Pratul K. Agarwal
SC2
2009 Impact of Quad-Core Cray XT4 System and Software Stack on Scientific Computation
Sadaf R. Alam, Richard F. Barrett, Heike Jagode, Jeffery A. Kuehn, Stephen W. Poole, Ramanan Sankaran
Euro-Par1
2009 Performance Characterization of a Hierarchical MPI Implementation on Large-scale Distributed-memory Platforms
abstract
The building blocks of emerging Petascale massively parallel processing (MPP) systems are multi-core processors with four or more cores as a single processing element and a customized network interface. The resulting memory and communication hierarchy of these platforms are now exposed to application developers and end users by creating a hierarchical or multi-core aware message-passing (MPI) programming interface and by providing a handful of runtime, tunable parameters that allows mapping and control of MPI tasks and message handling. We characterize performance of MPI communication patterns and present strategies for optimizing applications performance on Cray XT series systems that are composed of contemporary AMD processors and a proprietary network infrastructure. We highlight dependencies in its memory and network subsystems, which could influence production-level applications performance. We demonstrate that MPI micro-benchmarks could mislead an application developer or end user since these benchmarks often do not expose the interplay between memory allocation and usage in the user space, which depends on the number of tasks or cores and workload characteristics. Our studies show performance improvements compared to the default options for our target scientific benchmarks and production-level applications.
Sadaf R. Alam, Richard F. Barrett, Jeffery A. Kuehn, Stephen W. Poole
ICPP1
2009 Performance analysis and projections for Petascale applications on Cray XT series systems
abstract
The Petascale Cray XT5 system at the Oak Ridge National Laboratory (ORNL) Leadership Computing Facility (LCF) shares a number of system and software features with its predecessor, the Cray XT4 system including the quad-core AMD processor and a multi-core aware MPI library. We analyze performance of scalable scientific applications on the quad-core Cray XT4 system as part of the early system access using a combination of micro-benchmarks and Petascale ready applications. Particularly, we evaluate impact of key changes that occurred during the dual-core to quad-core processor upgrade on applications behavior and provide projections for the next-generation massively-parallel platforms with multi-core processors, specifically for proposed Petascale Cray XT5 system. We compare and contrast the quad-core XT4 system features with the upcoming XT5 system and discuss strategies for improving scaling and performance for our target applications.
Sadaf R. Alam, Richard F. Barrett, Jeffery A. Kuehn, Stephen W. Poole
IPDPS1
2008 Experimental Evaluation of Molecular Dynamics Simulations on Multi-core Systems
Sadaf R. Alam, Pratul K. Agarwal, Scott S. Hampton, Hong Ong
HiPC1
2008 Impact of multicores on large-scale molecular dynamics simulations
abstract
Processing nodes of the Cray XT and IBM Blue Gene Massively Parallel Processing (MPP) systems are composed of multiple execution units, sharing memory and network subsystems. These multicore processors offer greater computational power, but may be hindered by resource contention. In order to understand and avoid such situations, we investigate the impact of resource contention on three scalable molecular dynamics suites: AMBER (PMEMD module), LAMMPS, and NAMD. The results reveal the factors that can inhibit scaling and performance efficiency on emerging multicore processors.
Sadaf R. Alam, Pratul K. Agarwal, Scott S. Hampton, Hong Ong, Jeffrey S. Vetter
IPDPS1
2008 A Methodology for Developing High Fidelity Communication Models for Large-Scale Applications Targeted on Multicore Systems
abstract
Resource sharing and implementation of software stack for emerging multicore processors introduce performance and scaling challenges for large-scale scientific applications, particularly on systems with thousands of processing elements. Traditional performance optimization, tuning and modeling techniques that rely on uniform representation of computation and communication requirements are only partially useful due to the complexity of applications and underlying systems and software architecture. In this paper, we propose a workload modeling methodology that allows application developers to capture and represent hierarchical decomposition and distribution of their applications thereby allowing them to explore and identify optimal mapping of a workload on a target system. We demonstrate the proposed methodology on a Teraflopsscale fusion application that is developed using message-passing (MPI) programming paradigm. Using our analysis and projection results, we obtain insight into the performance characteristics of the application on a quad-core system and also identify optimal mapping on a Teraflops-scale platform. 1.
Charles W. Lively, Valerie Taylor 0001, Sadaf R. Alam, Jeffrey S. Vetter
SBAC-PAD3
2008 Early evaluation of IBM BlueGene/P
abstract
BlueGene/P (BG/P) is the second generation BlueGene architecture from IBM, succeeding BlueGene/L (BG/L). BG/P is a system-on-a-chip (SoC) design that uses four PowerPC 450 cores operating at 850 MHz with a double precision, dual pipe floating point unit per core. These chips are connected with multiple interconnection networks including a 3-D torus, a global collective network, and a global barrier network. The design is intended to provide a highly scalable, physically dense system with relatively low power requirements per flop. In this paper, we report on our examination of BG/P, presented in the context of a set of important scientific applications, and as compared to other major large scale supercomputers in use today. Our investigation confirms that BG/P has good scalability with an expected lower performance per processor when compared to the Cray XT4's Opteron. We also find that BG/P uses very low power per floating point operation for certain kernels, yet it has less of a power advantage when considering science driven metrics for mission applications.
Sadaf R. Alam, Richard F. Barrett, M. Bast, Mark R. Fahey, Jeffery A. Kuehn, Collin McCurdy, James H. Rogers, Philip C. Roth, Ramanan Sankaran, Jeffrey S. Vetter, Patrick H. Worley, Weikuan Yu
SC1
2008 Performance characteristics of biomolecular simulations on high-end systems with multi-core processors
Sadaf R. Alam, Pratul K. Agarwal, Jeffrey S. Vetter
Parallel Comput.1
2007 An Application Specific Memory Characterization Technique for Co-processor Accelerators
abstract
Commodity accelerator technologies including reconfigurable devices provide an order of magnitude performance improvement compared to mainstream microprocessor systems. A number of compute-intensive scientific applications, therefore, can potentially benefit from commodity computing devices available in the form of co-processor accelerators. However, there has been little progress in accelerating production-level scientific applications using these technologies due to several programming and performance challenges. One of the key perfomance challenges is performance sustainability. While computation is often accelerated substantially by accelerator devices, the achievable performance is significantly lower once the data transfer costs and overheads are incorporated. We present an application-specific memory characterization technique for an FPGA-accelerated system that enabled us to reduce data transfer overhead by a factor of five for a production-scale scientific application. Our proposed technique extends to applications that exhibit similar memory behavior and to co-processor accelerator systems that support data streaming, pipelining, and overlapped execution.
Sadaf R. Alam, Jeffrey S. Vetter, Melissa C. Smith
ASAP1
2007 Performance Evaluation of a Scalable Molecular Dynamics Simulation Framework on a Massively-Parallel System
abstract
The successors of distributed-memory, massively-parallel processing (MPP) systems that are based on multi-core processor technologies and high-bandwidth communication networks are expected to deliver Petascale computing power for scientific communities in near future. This report presents preliminary performance evaluation and benchmarking results of a scalable biomolecular simulation framework on the Cray XT4 MPP system that contains multi-core Opteron processors. We identify not only the performance enhancing features but also the bottlenecks for biomolecular simulation test cases on this system using a combination of application and vendor specific performance tools. Our results show that unprecedented performance has been achieved for large-scale test cases on the system; however, the critical challenges remain for longer time scale simulations on MPP systems.
Sadaf R. Alam, Pratul K. Agarwal, Jeffery A. Kuehn
BIBE1
2007 Sensitivity Analysis of Biomolecular Simulations using Symbolic Models
abstract
Performance and scaling of biomolecular simulations frameworks largely depends on not only the workload characteristics of the simulations but also the design of underlying processor architecture and interconnection networks. Because construction of Teraflops and Petaflops scale prototype systems for evaluation alone is impractical and cost-prohibitive, architects use analytical models of workloads and architecture simulators to guide their design decisions and tradeoffs. To address the problem of providing scalable yet precise input for network simulators, we have developed a technique to model symbolically the communication patterns of production-level scientific applications to study workload growth rates and to carry out sensitivity analysis. We apply our symbolic modeling scheme to the particle mesh ewald (PME) implementation in the sander package of the AMBER framework and demonstrate how the increase in computation, memory and communication requirements impact the performance and scaling of the PME method on the next-generation massively-parallel systems.
Sadaf R. Alam, Nikhil Bhatia, Jeffrey S. Vetter
BIBE1
2007 Balancing productivity and performance on the cell broadband engine
abstract
The cell broadband engine (BE) is a heterogeneous multicore processor, combining a general-purpose POWER architecture core with eight independent single-instruction-multiple-data (SIMD) cores. Each core is capable of very high performance; however, users must explicitly manage data movement, scheduling, and synchronization. While these attributes provide some of the cell processorpsilas greatest performance strengths, they also form its greatest weaknesses in terms of developer productivity, code portability, and initial performance efficiencies. In this paper, we evaluate productivity and relative performance improvements of a cell BE system for a diverse set of kernels and applications. Our experimental workload includes algorithms from scientific, cognitive, and imaging problem domains. Our results demonstrate that the cell processor could be several times faster than a SSE-enabled, contemporary dual-core processor, and could sustain a high performance-to-productivity ratio. We outline strategies for transforming applications to exploit the cellpsilas architectural features, and measure productivity by comparing programming effort in terms of lines of code and performance. For instance, our measurements revealed that a covariance matrix creation routine - a common routine in hyperspectral imaging - ran over eight times faster than a 2.66 GHz Intel Woodcrest processor while sustaining a productivity metric of over two by parallelizing across the heterogeneous cores, unrolling loops, and improving instruction level parallelism with SIMD instructions in a high-level language.
Sadaf R. Alam, Jeremy S. Meredith, Jeffrey S. Vetter
CLUSTER1
2007 An Exploration of Performance Attributes for Symbolic Modeling of Emerging Processing Devices
Sadaf R. Alam, Nikhil Bhatia, Jeffrey S. Vetter
HPCC1
2007 On the Path to Enable Multi-scale Biomolecular Simulations on PetaFLOPS Supercomputer with Multi-core Processors
abstract
Biological processes occurring inside cell involve multiple scales of time and length; many popular theoretical and computational multi-scale techniques utilize biomolecular simulations based on molecular dynamics. Till recently, the computing power required for simulating the relevant scales was even beyond the reach of fastest supercomputers. The availability of petaFLOPS-scale computing power in near future holds great promise. Unfortunately, the bio-simulations software technology has not kept up with the changes in hardware. In particular, with the introduction of multi-core processing technologies in systems with tens of thousands of processing cores, it is unclear whether the existing biomolecular simulation frameworks will be able to scale and to utilize these resources effectively. While the multi-core processing systems provide higher processing capabilities, their memory and IO subsystems are posing new challenges to application and system software developers. In this preliminary study, we attempt to characterize computation, communication and memory efficiencies of bio-molecular simulations on a Cray XT3 system, which has recently been upgraded to dual-core Opteron processors. We identify that the application efficiencies using the multi-core processors reduce with the increase of the simulated system size. Further, we measure the communication overhead of using both cores in the processor simultaneously and identify that the MPI communication performance can be as low as 50% as compared to the single-core execution times. We conclude that not only the biomolecular simulations need to be aware of the underlying multi-core hardware in order to achieve maximum performance but also the system software needs to provide processor and memory placement features in the high-end systems. Our results on a stand-alone dual-core AMD system confirm that combinations of processor and memory affinity schemes can result in over 12% performance gains.
Sadaf R. Alam, Pratul K. Agarwal
IPDPS1
2007 Analysis of a Computational Biology Simulation Technique on Emerging Processing Architectures
abstract
Multi-paradigm, multi-threaded and multi-core computing devices available today provide several orders of magnitude performance improvement over mainstream microprocessors. These devices include the STI Cell Broadband Engine, graphical processing units (GPU) and the Cray massively-multithreaded processors - available in desktop computing systems as well as proposed for supercomputing platforms. The main challenge in utilizing these powerful devices is their unique programming paradigms. GPUs and the Cell systems require code developers to manage code and data explicitly, while the Cray multithreaded architecture requires them to generate a very large number of threads or independent tasks concurrently. In this paper, we explain strategies for optimizing a molecular dynamics (MD) calculation that is used in biomolecular simulations on three devices: Cell, GPU and MTA-2. We show that the Cray MTA-2 system requires minimal code modification and does not outperform the microprocessor runs; but it demonstrates an improved workload scaling behavior over the microprocessor implementation. On the other hand, substantial porting and optimization efforts on the Cell and the GPU systems result in a 5times to 6times improvement, respectively, over a 2.2 GHz Opteron system.
Jeremy S. Meredith, Sadaf R. Alam, Jeffrey S. Vetter
IPDPS2
2007 Performance evaluation of the cray XT3 configured with dual core opteron processors
abstract
No abstract available.
Richard F. Barrett, Sadaf R. Alam, Jeffrey S. Vetter
PPoPP2
2007 Cray XT4: an early evaluation for petascale scientific simulation
abstract
The scientific simulation capabilities of next generation high-end computing technology will depend on striking a balance among memory, processor, I/O, and local and global network performance across the breadth of the scientific simulation space. The Cray XT4 combines commodity AMD dual core Opteron processor technology with the second generation of Cray's custom communication accelerator in a system design whose balance is claimed to be driven by the demands of scientific simulation. This paper presents an evaluation of the Cray XT4 using micro-benchmarks to develop a controlled understanding of individual system components, providing the context for analyzing and comprehending the performance of several petascale-ready applications. Results gathered from several strategic application domains are compared with observations on the previous generation Cray XT3 and other high-end computing systems, demonstrating performance improvements across a wide variety of application benchmark problems.
Sadaf R. Alam, Jeffery A. Kuehn, Richard F. Barrett, Jeffrey M. Larkin, Mark R. Fahey, Ramanan Sankaran, Patrick H. Worley
SC1
2006 Hierarchical Model Validation of Symbolic Performance Models of Scientific Kernels
Sadaf R. Alam, Jeffrey S. Vetter
Euro-Par1
2006 An Analysis of System Balance Requirements for Scientific Applications
abstract
Scientific applications are diverse in terms of the resource requirements, and tend to vary significantly from commercial applications. In order to provide sustained performance, a target high performance computing (HPC) platform must offer a balance between CPU performance to memory, interconnect and I/O subsystems performance. We characterize the system balance requirements for two large-scale Office of Science applications, GYRO (fusion simulation) and POP (climate modeling), and develop platform-independent parameterized requirement models. We measure the parallel efficiencies for GYRO and POP on three multiprocessor systems: an SMP cluster (IBM p690), a shared-memory system (SGI Altix) and a vector supercomputer (Cray XI). The higher computational intensity and interconnect bandwidth requirements of GYRO result in higher performance efficiencies on the vector platform. At the same time, small message sizes in POP benefit from low MPI latencies of the shared-memory platform. Overall results confirm system balance requirements that are generated by the requirement models
Sadaf R. Alam, Jeffrey S. Vetter
ICPP1
2006 A framework to develop symbolic performance models of parallel applications
abstract
Performance and workload modeling has numerous uses at every stage of the high-end computing lifecycle: design, integration, procurement, installation and tuning. Despite the tremendous usefulness of performance models, their construction remains largely a manual, complex, and time-consuming exercise. We propose a new approach to the model construction, called modeling assertions (MA), which borrows advantages from both the empirical and analytical modeling techniques. This strategy has many advantages over traditional methods: incremental construction of realistic performance models, straightforward model validation against empirical data, and intuitive error bounding on individual model terms. We demonstrate this new technique on the NAS parallel CG and SP benchmarks by constructing high fidelity models for the floating-point operation cost, memory requirements, and MPI message volume. These models are driven by a small number of key input parameters thereby allowing efficient design space exploration of future problem sizes and architectures
Sadaf R. Alam, Jeffrey S. Vetter
IPDPS1
2006 Early evaluation of the Cray XT3
abstract
Oak Ridge National Laboratory recently received delivery of a 5,294 processor Cray XT3. The XT3 is Cray's third-generation massively parallel processing system. The system builds on a single processor node - built around the AMD Opteron - and uses a custom chip - called SeaStar - to provide interprocess or communication. In addition, the system uses a lightweight operating system on the compute nodes. This paper describes our initial experiences with the system, including micro-benchmark, kernel, and application benchmark results. In particular, we provide performance results for strategic Department of Energy applications areas including climate and fusion. We demonstrate experiments on the installed system, scaling applications up to 4,096 processors.
Jeffrey S. Vetter, Sadaf R. Alam, Thomas H. Dunigan, Mark R. Fahey, Philip C. Roth, Patrick H. Worley
IPDPS2
2006 Performance characterization of molecular dynamics techniques for biomolecular simulations
abstract
Large-scale simulations and computational modeling using molecular dynamics (MD) continues to make significant impacts in the field of biology. It is well known that simulations of biological events at native time and length scales requires computing power several orders of magnitude beyond today's commonly available systems. Supercomputers, such as IBM Blue Gene/L and Cray XT3, will soon make tens to hundreds of teraFLOP/s of computing power available by utilizing thousands of processors. The popular algorithms and MD applications, however, were not initially designed to run on thousands of processors. In this paper, we present detailed investigations of the performance issues, which are crucial for improving the scalability of the MD-related algorithms and applications on massively parallel processing (MPP) architectures. Due to the varying characteristics of biological input problems, we study two prototypical biological complexes that use the MD algorithm: an explicit solvent and an implicit solvent. In particular, we study the AMBER application, which supports a variety of these types of input problems. For the explicit solvent problem, we focused on the particle mesh Ewald (PME) method for calculating the electrostatic energy, and for the implicit solvent model, we targeted the Generalized Born (GB) calculation. We uncovered and subsequently modified a limitation in AMBER that restricted the scaling beyond 128 processors. We collected performance data for experiments on up to 2048 Blue Gene/L and XT3 processors and subsequently identified that the scaling is largely limited by the underlying algorithmic characteristics and also by the implementation of the algorithms. Furthermore, we found that the input problem size of biological system is constrained by memory available per node. In conclusion, our results indicate that MD codes can significantly benefit from the current generation architectures with relatively modest optimization efforts. Nevertheless, the key for enabling scientific breakthroughs lies in exploiting the full potential of these new architectures.
Sadaf R. Alam, Jeffrey S. Vetter, Pratul K. Agarwal, Al Geist
PPoPP1