VLDB 2026 Research / reviewers in the wild / expert
Alistair P. Rendell
dblp:54/3648
· DBLP profile ↗
24ranked-venue papers
1as first author
2since 2021 · last 2022
0000-0002-9445-0146ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 1 first-author · 2 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
High-performance computing · 63% GPUs and heterogeneous computing · 28% Parallel and multicore computing · 9% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing
scientific computing systems |
1.6 | 4 | 2022 | Scaling Correlated Fragment Molecular Orbital Calculations on Summit · SC 2022 Enabling large-scale correlated electronic structure calculations: scaling the RI-MP2 method on summit · SC 2021 Scaling the hartree-fock matrix build on summit · SC 2020 |
High-performance computing › scientific computing systems
quantum chemistry simulation |
1.5 | 3 | 2022 | Scaling Correlated Fragment Molecular Orbital Calculations on Summit · SC 2022 Enabling large-scale correlated electronic structure calculations: scaling the RI-MP2 method on summit · SC 2021 Scaling the hartree-fock matrix build on summit · SC 2020 |
GPUs and heterogeneous computing › multi-GPU computing
distributed GPU computing |
1.1 | 2 | 2022 | Scaling Correlated Fragment Molecular Orbital Calculations on Summit · SC 2022 Enabling large-scale correlated electronic structure calculations: scaling the RI-MP2 method on summit · SC 2021 |
GPUs and heterogeneous computing › GPU resource management
GPU load balancing |
0.4 | 1 | 2020 | Scaling the hartree-fock matrix build on summit · SC 2020 |
Parallel and multicore computing › load balancing
dynamic load balancing |
0.1 | 1 | 2020 | Scaling the hartree-fock matrix build on summit · SC 2020 |
Computational science and engineering
computational chemistry |
0.1 | 1 | 2009 | Liquid water: obtaining the right answer for the right reasons · SC 2009 |
High-performance computing › quantum chemistry
coupled-cluster method |
0.1 | 1 | 2009 | Liquid water: obtaining the right answer for the right reasons · SC 2009 |
High-performance computing › supercomputing
petascale computing |
0.1 | 1 | 2009 | Liquid water: obtaining the right answer for the right reasons · SC 2009 |
High-performance computing › scientific computing systems
computational chemistry |
0.0 | 1 | 2003 | Enabling the Efficient Use of SMP Clusters: The GAMESS/DDI Model · SC 2003 |
High-performance computing › scientific computing systems
electronic structure calculation |
0.0 | 1 | 2003 | Enabling the Efficient Use of SMP Clusters: The GAMESS/DDI Model · SC 2003 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 2003 | Enabling the Efficient Use of SMP Clusters: The GAMESS/DDI Model · SC 2003 |
High-performance computing
cluster computing |
0.0 | 1 | 2003 | Enabling the Efficient Use of SMP Clusters: The GAMESS/DDI Model · SC 2003 |
High-performance computing › cluster computing
SMP cluster |
0.0 | 1 | 2003 | Enabling the Efficient Use of SMP Clusters: The GAMESS/DDI Model · SC 2003 |
Methods — techniques the papers use, named apart from their topics
RI-MP2 · 1.1distributed many-GPU algorithm · 0.6molecular fragmentation · 0.5many-GPU algorithm · 0.5fragmentation-based algorithm · 0.4fock digestion · 0.4dynamic load balancing · 0.4coupled-cluster theory · 0.3basis set convergence · 0.3message passing · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Scaling Correlated Fragment Molecular Orbital Calculations on SummitabstractCorrelated electronic structure calculations enable an accurate prediction of the physicochemical properties of complex molecular systems; however, the scale of these calculations is limited by their extremely high computational cost. The Fragment Molecular Orbital (FMO) method is arguably one of the most effective ways to lower this computational cost while retaining predictive accuracy. In this paper, a novel distributed many-GPU algorithm and implementation of the FMO method are presented. When applied in tandem with the Hartree-Fock and RI-MP2 methods, the new implementation enables correlated calculations on 623,016 electrons and 146,592 atoms in less than 45 minutes using 99.8% of the Summit supercomputer (27,600 GPUs). The implementation demonstrates remarkable speedups with respect to other current GPU and CPU codes, and excellent strong scalability on Summit achieving 94.6 % parallel efficiency on 4600 nodes. This work makes feasible correlated quantum chemistry calculations on significantly larger molecular systems than before and with higher accuracy. Giuseppe M. J. Barca, Calum Snowdon, Jorge L. Galvez Vallejo, Fazeleh S. Kazemian, Alistair P. Rendell, Mark S. Gordon |
SC | 5 |
| 2021 | Enabling large-scale correlated electronic structure calculations: scaling the RI-MP2 method on summitabstractSecond-order Møller-Plesset perturbation theory using the Resolution-of-the-Identity approximation (RI-MP2) is a state-of-the-art approach to accurately estimate many-body electronic correlation effects. This is critical for predicting the physicochemical properties of complex molecular systems; however, the scale of these calculations is limited by their extremely high computational cost. In this paper, a novel many-GPU algorithm and implementation of a molecular-fragmentation-based RI-MP2 method are presented that enable correlated calculations on over 180,000 electrons and 45,000 atoms using up to the entire Summit supercomputer in 12 minutes. The implementation demonstrates remarkable speedups with respect to other current GPU and CPU codes, excellent strong scalability on Summit achieving 89.1% parallel efficiency on 4600 nodes, and shows nearly-ideal weak scaling up to 612 nodes. This work makes feasible ab initio correlated quantum chemistry calculations on significantly larger molecular scales than before on both large supercomputing systems and on commodity clusters, with a potential for major impact on progress in chemical, physical, biological and engineering sciences. Giuseppe M. J. Barca, Jorge L. Galvez Vallejo, David Poole 0001, Melisa Alkan, Ryan Stocks, Alistair P. Rendell, Mark S. Gordon |
SC | 6 |
| 2020 | Scaling the hartree-fock matrix build on summitabstractUsage of Graphics Processing Units (GPU) has become strategic for simulating the chemistry of large molecular systems, with the majority of top supercomputers utilizing GPUs as their main source of computational horsepower. In this paper, a new fragmentation-based Hartree-Fock matrix build algorithm designed for scaling on many-GPU architectures is presented. The new algorithm uses a novel dynamic load balancing scheme based on a binned shell-pair container to distribute batches of significant shell quartets with the same code path to different GPUs. This maximizes computational throughput and load balancing, and eliminates GPU thread divergence due to integral screening. Additionally, the code uses a novel Fock digestion algorithm to contract electron repulsion integrals into the Fock matrix, which exploits all forms of permutational symmetry and eliminates thread synchronization requirements. The implementation demonstrates excellent scalability on the Summit computer, achieving good strong scaling performance up to 4096 nodes, and linear weak scaling up to 612 nodes. Giuseppe M. J. Barca, David Poole 0001, Jorge L. Galvez Vallejo, Melisa Alkan, Colleen Bertoni, Alistair P. Rendell, Mark S. Gordon |
SC | 6 |
| 2018 | Development and Application of a Hybrid Programming Environment on an ARM/DSP System for High Performance ComputingabstractThe nCore Brown-Dwarf system has a unique architecture where each node is comprised of two different low-power System-on-Chip (LPSoC) processors from Texas Instruments; the ARM/DSP Keystone II SoC and the DSP based Keystone I SoC. These LPSoC processors have, through use of the C66x multi-core DSP, been shown to be capable of running floating-point intensive HPC application codes. However, it is non-trivial to run such codes across all processing elements of a node simultaneously. This paper demonstrates a hybrid programming environment that combines OpenMP, OpenCL and MPI to enable application execution across multiple Brown-Dwarf nodes. This environment is evaluated using two diverse application codes. The first is Level-3 BLAS matrix multiplication (GEMM), which is a standard HPC floating-point intensive benchmark. The second is a unique real-world scientific code for biostructure based drug design developed by the Southwest Research Institute called Rhodium. Performance and energy-efficiency of Rhodium is presented alongside comparisons with conventional x86 based HPC systems with attached accelerators. Results indicate that the Brown-Dwarf system remains competitive with contemporary systems for memory-bound computations. Gaurav Mitra, Jonathan Bohmann, Ian Lintault, Alistair P. Rendell |
IPDPS | 4 |
| 2014 | PGAS-FMM: Implementing a distributed fast multipole method using the X10 programming languageabstractSUMMARY The fast multipole method (FMM) is a complex, multi‐stage algorithm over a distributed tree data structure, with multiple levels of parallelism and inherent data locality. X10 is a modern partitioned global address space language with support for asynchronous activities. The parallel tasks comprising FMM may be expressed in X10 by using a scalable pattern of activities. This paper demonstrates the use of X10 to implement FMM for simulation of electrostatic interactions between ions in a cyclotron resonance mass spectrometer. X10's task‐parallel model is used to express parallelism by using a pattern of activities mapping directly onto the tree. X10's work stealing runtime handles load balancing fine‐grained parallel activities, avoiding the need for explicit work sharing. The use of global references and active messages to create and synchronize parallel activities over a distributed tree structure is also demonstrated. In contrast to previous simulations of ion trajectories in cyclotron resonance mass spectrometers, our code enables both simulation of realistic particle numbers and guaranteed error bounds. Single‐node performance is comparable with the fastest published FMM implementations, and critical expansion operators are faster for high accuracy calculations. A comparison of parallel and sequential codes shows the overhead of activity management and work stealing in this application is low. Scalability is evaluated for 8k cores on a Blue Gene/Q system and 512 cores on a Nehalem/InfiniBand cluster. Copyright © 2013 John Wiley & Sons, Ltd. Josh Milthorpe, Alistair P. Rendell |
Concurr. Comput. Pract. Exp. | 2 |
| 2013 | Deterministic global optimization in ab-initio quantum chemistry
Pete P. Janes, Alistair P. Rendell |
J. Glob. Optim. | 2 |
| 2012 | Efficient update of ghost regions using active messagesabstractThe use of ghost regions is a common feature of many distributed grid applications. A ghost region holds local read-only copies of remotely-held boundary data which are exchanged and cached many times over the course of a computation. X10 is a modern parallel programming language intended to support productive development of distributed applications. X10 supports the “active message” paradigm, which combines data transfer and computation in one-sided communications. A central feature of X10 is the distributed array, which distributes array data across multiple places, providing standard read and write operations as well as powerful high-level operations. We used active messages to implement ghost region updates for X10 distributed arrays using two different update algorithms. Our implementation exploits multiple levels of parallelism and avoids global synchronization; it also supports split-phase ghost updates, which allows for overlapping computation and communication. We compare the performance of these algorithms on two platforms: an Intel x86-64 cluster over QDR InfiniBand, and a Blue Gene/P system, using both stand-alone benchmarks and an example computational chemistry application code. Our results suggest that on a dynamically threaded architecture, a ghost region update using only pairwise synchronization exhibits superior scaling to an update that uses global collective synchronization. Josh Milthorpe, Alistair P. Rendell |
HiPC | 2 |
| 2012 | Implementation of 3D FFTs Across Multiple GPUs in Shared Memory EnvironmentsabstractIn this paper, a novel implementation of the distributed 3D Fast Fourier Transform (FFT) on a multi-GPU platform using CUDA is presented. The 3D FFT is the core of many simulation methods, thus its fast calculation is critical. The main bottleneck of the distributed 3D FFT is the global data exchange which must be performed. The latest version of CUDA introduces direct GPU-to-GPU transfers using a Unified Virtual Address space (UVA) that provides new possibilities for optimising the communication part of the FFT. Here, we propose different implementations of the distributed 3D FFT, investigate their behaviour, and compare their performance with the single GPU CUFFT and CPU-based FFTW libraries. In particular, we demonstrate the advantage of direct GPU-to-GPU transfers over data exchanges via host main memory. Our preliminary results show that running the distributed 3D FFT with four GPUs can bring a 12% speedup over the single node (CUFFT) while also enabling the calculation of 3D FFTs of larger datasets. Replacing the global data exchange via shared memory with direct GPU-to-GPU transfers reduces the execution time by up to 49%. This clearly shows that direct GPU-to-GPU transfers are the key factor in obtaining good performance on multi-GPU systems. Nimalan Nandapalan, Jirí Jaros, Alistair P. Rendell, Bradley E. Treeby |
PDCAT | 3 |
| 2012 | Generating optimal CUDA sparse matrix-vector product implementations for evolving GPU hardwareabstractSUMMARY The CUDA model for graphics processing units (GPUs) presents the programmer with a plethora of different programming options. These includes different memory types, different memory access methods and different data types. Identifying which options to use and when is a non‐trivial exercise. This paper explores the effect of these different options on the performance of a routine that evaluates sparse matrix–vector products (SpMV) across three different generations of NVIDIA GPU hardware. A process for analysing performance and selecting the subset of implementations that perform best is proposed. The potential for mapping sparse matrix attributes to optimal CUDA SpMV implementations is discussed. Copyright © 2011 John Wiley & Sons, Ltd. Ahmed H. El Zein, Alistair P. Rendell |
Concurr. Comput. Pract. Exp. | 2 |
| 2011 | X10 as a Parallel Language for Scientific Computation: Practice and ExperienceabstractX10 is an emerging Partitioned Global Address Space (PGAS) language intended to increase significantly the productivity of developing scalable HPC applications. The language has now matured to a point where it is meaningful to consider writing large scale scientific application codes in X10. This paper reports our experiences writing three codes from the chemistry/material science domain: Fast Multipole Method (FMM), Particle Mesh Ewald (PME) and Hartree-Fock (HF), entirely in X10. Performance results are presented for up to 256 places on a Blue Gene/P system. During the course of this work our experiences have been shared with the X10 development team, so that application requirements could inform language design discussions as the language capabilities influenced algorithm design. This resulted in improvements in the language implementation and standard class libraries, including the design of the array API and support for complex math. Data constructs in X10 such as \emph{places} and \emph{distributed arrays}, and parallel constructs such as \emph{finish} and \emph{async}, simplify implementation of the applications in comparison with MPI. However, current implementation limitations in X10 2.1.2 make it difficult to achieve scalable performance using the most natural expressions of the algorithms. The most serious limitation is the use of point-to-point communication patterns, rather than collectives, to implement parallel constructs and array operations. This issue will be addressed in future releases of X10. Josh Milthorpe, V. Ganesh, Alistair P. Rendell, David Grove |
IPDPS | 3 |
| 2011 | Profiling Directed NUMA Optimization on Linux Systems: A Case Study of the Gaussian Computational Chemistry CodeabstractThe parallel performance of applications running on Non-Uniform Memory Access (NUMA) platforms is strongly influenced by the relative placement of memory pages to the threads that access them. As a consequence there are Linux application programmer interfaces (APIs) to control this. For large parallel codes it can, however, be difficult to determine how and when to use these APIs. In this paper we introduce the NUMAgrind profiling tool which can be used to simplify this process. It extends the Val grind binary translation framework to include a model which incorporates cache coherency, memory locality domains and interconnect traffic for arbitrary NUMA topologies. Using NUMAgrind, cache misses can be mapped to memory locality domains, page access modes determined, and pages that are referenced by multiple threads quickly determined. We show how the NUMAgrind tool can be used to guide the use of Linux memory and thread placement APIs in the Gaussian computational chemistry code. The performance of the code before and after use of these APIs is also presented for three different commodity NUMA platforms. Rui Yang 0005, Joseph Antony, Alistair P. Rendell, Danny Robson, Peter E. Strazdins |
IPDPS | 3 |
| 2010 | Region-Based Prefetch Techniques for Software Distributed Shared Memory SystemsabstractAlthough shared memory programming models show good programmability compared to message passing programming models, their implementation by page-based software distributed shared memory systems usually suffers from high memory consistency costs. The major part of these costs is inter-node data transfer for keeping virtual shared memory consistent. A good prefetch strategy can reduce this cost. We develop two prefetch techniques, TReP and HReP, which are based on the execution history of each parallel region. These techniques are evaluated using offline simulations with the NAS Parallel Benchmarks and the LINPACK benchmark. On average, TReP achieves an efficiency (ratio of pages prefetched that were subsequently accessed) of 96% and a coverage (ratio of access faults avoided by prefetches) of 65%. HReP achieves an efficiency of 91% but has a coverage of 79%. Treating the cost of an incorrectly prefetched page to be equivalent to that of a miss, these techniques have an effective page miss rate of 63% and 71% respectively. Additionally, these two techniques are compared with two well-known software distributed shared memory (sDSM) prefetch techniques, Adaptive++ and TODFCM. TReP effectively reduces page miss rate by 53% and 34% more, and HReP effectively reduces page miss rate by 62% and 43% more, compared to Adaptive++ and TODFCM respectively. As for Adaptive++, these techniques also permit bulk prefetching for pages predicted using temporal locality, amortizing network communication costs and permitting bandwidth improvement from multi-rail network interfaces. Peter E. Strazdins, Alistair P. Rendell |
CCGRID | 3 |
| 2010 | From Sparse Matrix to Optimal GPU CUDA Sparse Matrix Vector Product ImplementationabstractThe CUDA model for GPUs presents the programmer with a plethora of different programming options. These includes different memory types, different memory access methods, and different data types. Identifying which options to use and when is a non-trivial exercise. This paper explores the effect of these different options on the performance of a routine that evaluates sparse matrix vector products. A process for analysing performance and selecting the subset of implementations that perform best is proposed. The potential for mapping sparse matrix attributes to optimal CUDA sparse matrix vector product implementation is discussed. Ahmed H. El Zein, Alistair P. Rendell |
CCGRID | 2 |
| 2009 | Integrating software distributed shared memory and message passing programmingabstractSoftware Distributed Shared Memory (SDSM) systems provide programmers with a shared memory programming environment across distributed memory architectures. In contrast to the message passing programming environment, the SDSM can resolve data dependencies within the application without the programmer having to explicitly specify communication. However, this service is provided at a cost to performance. Thus it makes sense to use message passing directly when data dependencies are easy to solve using message passing. For example, it is not complicated to specify data transfer for large contiguous regions of memory. This paper outlines how the Danui SDSM library has been extended to include support for message passing. Four different message passing transfers are identified depending on whether the data being sent/received resides in private or globally shared buffers. Transfers between globally shared buffers are further categorized as symmetrical or asymmetrical depending on whether they correspond to the same region of shared memory. The implication of each transfer type on the memory consistency of the global address space is discussed. Central to the Danui SDSM extension is the use of information provided and implied by message passing operations. The overhead of the implementation is analyzed. H'sien J. Wong, Alistair P. Rendell |
CLUSTER | 2 |
| 2009 | A Simple Performance Model for Multithreaded Applications Executing on Non-uniform Memory Access ComputersabstractIn this work, we extend and evaluate a simple performance model to account for NUMA and bandwidth effects for single and multi-threaded calculations within the Gaussian 03 computational chemistry code on a contemporary multi-core, NUMA platform. By using the thread and memory placement APIs in Solaris, we present results for a set of calculations from which we analyze on-chip interconnect and intra-core bandwidth contention and show the importance of load-balancing between threads. The extended model predicts single threaded performance to within 1% errors and most multi-threaded experiments within 15% errors. Our results and modeling shows that accounting for bandwidth constraints within user-space code is beneficial. Rui Yang 0005, Joseph Antony, Alistair P. Rendell |
HPCC | 3 |
| 2009 | Non-threaded and Threaded Approaches to MultiRail Communication with uDAPLabstractuDAPL is portable and platform independent communication library, which provides RDMA as well as send/recv operations. Some well known software has attempted to take advantage of uDAPL's portability, such as Open MPI, MVAPICH2, Intel MPI, and Cluster OpenMP. However, network performance is still the bottleneck for those software. Engaging "multirail" network is a method to by-pass it. In this paper, we have designed a non-threaded and a threaded approaches to improve performance of uDAPL over multirail configured clusters. The two approaches will be evaluated on different InfiniBand multirail configured clusters. The results shows that threaded approach improves 33% and 148% of the uni-directional bandwidth on the multi-port and the multi-HCA configured network respectively, and the non-threaded approach improves ~90% of the uni-directional bandwidth on the multi-HCA configured network. A similar improvements have been achieved for the bi-directional bandwidth. Alistair P. Rendell, Peter E. Strazdins |
NPC | 2 |
| 2009 | Liquid water: obtaining the right answer for the right reasonsabstractWater is ubiquitous on our planet and plays an essential role in several key chemical and biological processes. Accurate models for water are crucial in understanding, controlling and predicting the physical and chemical properties of complex aqueous systems. Over the last few years we have been developing a molecular-level based approach for a macroscopic model for water that is based on the explicit description of the underlying intermolecular interactions between molecules in water clusters. In the absence of detailed experimental data for small water clusters, highly-accurate theoretical results are required to validate and parameterize model potentials. As an example of the benchmarks needed for the development of accurate models for the interaction between water molecules, for the most stable structure of (H2O)20 we ran a coupled-cluster calculation on the ORNL's Jaguar petaflop computer that used over 100 TB of memory for a sustained performance of 487 TFLOP/s (double precision) on 96,000 processors, lasting for 2 hours. By this summer we will have studied multiple structures of both (H2O)20 and (H2O)24 and completed basis set and other convergence studies and anticipate the sustained performance rising close to 1 PFLOP/s. Edoardo Aprà, Alistair P. Rendell, Robert J. Harrison, Vinod Tipparaju, Bert de Jong, Sotiris S. Xantheas |
SC | 2 |
| 2008 | Reinforcement learning for automated performance tuning: Initial evaluation for sparse matrix format selectionabstractThe field of reinforcement learning has developed techniques for choosing beneficial actions within a dynamic environment. Such techniques learn from experience and do not require teaching. This paper explores how reinforcement learning techniques might be used to determine efficient storage formats for sparse matrices. Three different storage formats are considered: coordinate, compressed sparse row, and blocked compressed sparse row. Which format performs best depends heavily on the nature of the matrix and the computer system being used. To test the above a program has been written to generate a series of sparse matrices, where any given matrix performs optimally using one of the three different storage types. For each matrix several sparse matrix vector products are performed. The goal of the learning agent is to predict the optimal sparse matrix storage format for that matrix. The proposed agent uses five attributes of the sparse matrix: the number of rows, the number of columns, the number of non-zero elements, the standard deviation of non-zeroes per row and the mean number of neighbours. The agent is characterized by two parameters: an exploration rate and a parameter that determines how the state space is partitioned. The ability of the agent to successfully predict the optimal storage format is analyzed for a series of 1,000 automatically generated test matrices. Warren Armstrong, Alistair P. Rendell |
CLUSTER | 2 |
| 2007 | The design of MPI based distributed shared memory systems to support OpenMP on clustersabstractOpenMP can be supported in cluster environments by using distributed shared memory (DSM) systems. A portable approach for building DSM systems is to layer it on MPI. With these goals in mind, this paper makes two contributions. The first is a discussion about two software DSM systems that we have implemented using MPI. One uses background polling threads while the other uses processes that are driven only by incoming MPI messages. Comparisons of the two approaches show the latter to be a more scalable architecture that is better suited for the multi-core processors that are becoming commonplace. The second contribution recognizes that a common workaround for sub-team synchronizations in OpenMP is to use the flush directive on shared variables within busy-wait loops. In such a situation, only the flush in the last iteration of the busy-wait loop will result in the conditions necessary for exiting the loop. Thus transfer of the shared value need only be done if there were changes. We implement in our DSM a flush mechanism that eliminates the unnecessary data transfers entirely without any additional support or hints from the programmer. H'sien J. Wong, Alistair P. Rendell |
CLUSTER | 2 |
| 2007 | On the Use of Incomplete LU Decomposition as a Preconditioning Technique for Density Fitting in Electronic Structure Computations
Rui Yang 0005, Alistair P. Rendell, Michael J. Frisch |
ICCSA (1) | 2 |
| 2006 | Exploring Thread and Memory Placement on NUMA Architectures: Solaris and Linux, UltraSPARC/FirePlane and Opteron/HyperTransport
Joseph Antony, Pete P. Janes, Alistair P. Rendell |
HiPC | 3 |
| 2003 | Enabling the Efficient Use of SMP Clusters: The GAMESS/DDI ModelabstractAn important advance in cluster computing is the evolution from single processor clusters to multi-processor SMP clusters. Due to the increased complexity in the memory model on SMP clusters, new approaches are needed for applications that make use of distributed-memory paradigms. This paper presents new communications software developments that are designed to take advantage of SMP cluster hardware. Although the specific focus is on the central field of computational chemistry and materials science, as embodied in the popular electronic structure package GAMESS (General Atomic and Molecular Electronic Structure System), the impact of these new developments will be far broader in scope. Following a summary of the essential features of the distributed data interface (DDI) in the current implementation of GAMESS, the new developments for SMP clusters are described. The advantages of these new features are illustrated using timing benchmarks on several hardware platforms, using a typical computational chemistry application. Ryan M. Olson, Michael W. Schmidt, Mark S. Gordon, Alistair P. Rendell |
SC | 4 |
| 2000 | Computational chemistry on Fujitsu vector-parallel processors: Hardware and programming environment
Ross H. Nobes, Alistair P. Rendell, Jarek Nieplocha |
Parallel Comput. | 2 |
| 2000 | Computational chemistry on Fujitsu vector-parallel processors: Development and performance of applications software
Alistair P. Rendell, Andrey A. Bliznyuk, Ross H. Nobes, Elena V. Akhmatskaya, Herbert A. Früchtl, Paul W.-C. Kung, Victor Milman, Han Lung |
Parallel Comput. | 1 |