EDBT 2026 Demo / reviewers in the wild / expert
Hongzhang Shan
dblp:37/3718
· DBLP profile ↗
23ranked-venue papers
13as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 13 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
10 papers |
High-performance computing · 41% Performance modeling and evaluation · 18% Parallel and multicore computing · 15% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% |
Topics — the 30 heaviest of 32, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing
scientific computing systems |
0.3 | 3 | 2015 | Parallel implementation and performance optimization of the configuration-interaction method · SC 2015 Linearly scaling 3D fragment method for large-scale electronic structure calculations · SC 2008 A Comparison of Three Programming Models for Adaptive Applications on the Origin2000 · SC 2000 |
Parallel and multicore computing
load balancing |
0.2 | 1 | 2015 | Parallel implementation and performance optimization of the configuration-interaction method · SC 2015 |
Emerging computing paradigms › quantum computing › quantum simulation
quantum many-body simulation |
0.2 | 1 | 2015 | Parallel implementation and performance optimization of the configuration-interaction method · SC 2015 |
Performance modeling and evaluation
benchmarking |
0.2 | 3 | 2008 | Characterizing and predicting the I/O performance of HPC applications using a parameterized synthetic benchmark · SC 2008 Apex-Map: A Global Data Access Benchmark to Analyze HPC Systems and Parallel Programming Paradigms · SC 2005 Investigation of leading HPC I/O performance using a scientific-application derived benchmark · SC 2007 |
High-performance computing
performance optimization at scale |
0.1 | 2 | 2015 | Linearly scaling 3D fragment method for large-scale electronic structure calculations · SC 2008 Parallel implementation and performance optimization of the configuration-interaction method · SC 2015 |
Performance modeling and evaluation › benchmarking
i/o benchmarking |
0.1 | 2 | 2008 | Characterizing and predicting the I/O performance of HPC applications using a parameterized synthetic benchmark · SC 2008 Investigation of leading HPC I/O performance using a scientific-application derived benchmark · SC 2007 |
High-performance computing › scientific computing systems
electronic structure calculation |
0.1 | 1 | 2008 | Linearly scaling 3D fragment method for large-scale electronic structure calculations · SC 2008 |
High-performance computing
parallel i/o |
0.1 | 1 | 2008 | Characterizing and predicting the I/O performance of HPC applications using a parameterized synthetic benchmark · SC 2008 |
Storage systems › file systems › distributed file system
parallel file system |
0.1 | 1 | 2007 | Investigation of leading HPC I/O performance using a scientific-application derived benchmark · SC 2007 |
High-performance computing › parallel i/o
scientific application i/o |
0.1 | 1 | 2007 | Investigation of leading HPC I/O performance using a scientific-application derived benchmark · SC 2007 |
High-performance computing › sparse linear algebra
sparse matrix computation |
0.1 | 1 | 2015 | Parallel implementation and performance optimization of the configuration-interaction method · SC 2015 |
High-performance computing
collective communication |
0.1 | 1 | 2006 | Particles and contiuum - Performance modeling and optimization of a high energy colliding beam simulation code · SC 2006 |
Performance modeling and evaluation › communication modeling
communication performance modeling |
0.1 | 1 | 2006 | Particles and contiuum - Performance modeling and optimization of a high energy colliding beam simulation code · SC 2006 |
Memory systems
data locality |
0.1 | 1 | 2005 | Apex-Map: A Global Data Access Benchmark to Analyze HPC Systems and Parallel Programming Paradigms · SC 2005 |
Performance modeling and evaluation
workload characterization |
0.1 | 1 | 2005 | Apex-Map: A Global Data Access Benchmark to Analyze HPC Systems and Parallel Programming Paradigms · SC 2005 |
Distributed systems
grid computing |
0.0 | 1 | 2003 | Job Superscheduler Architecture and Performance in Computational Grid Environments · SC 2003 |
Cloud and datacenter computing
job scheduling |
0.0 | 1 | 2003 | Job Superscheduler Architecture and Performance in Computational Grid Environments · SC 2003 |
Cloud and datacenter computing
resource management |
0.0 | 1 | 2003 | Job Superscheduler Architecture and Performance in Computational Grid Environments · SC 2003 |
Parallel and multicore computing
programming models |
0.0 | 2 | 2000 | A Comparison of Three Programming Models for Adaptive Applications on the Origin2000 · SC 2000 Parallel Sorting on Cache-coherent DSM Multiprocessors · SC 1999 |
Computational science and engineering › computational chemistry › electronic structure calculation
density functional theory |
0.0 | 1 | 2008 | Linearly scaling 3D fragment method for large-scale electronic structure calculations · SC 2008 |
Computational science and engineering › materials science
materials science simulation |
0.0 | 1 | 2008 | Linearly scaling 3D fragment method for large-scale electronic structure calculations · SC 2008 |
High-performance computing
i/o intensive applications |
0.0 | 1 | 2008 | Characterizing and predicting the I/O performance of HPC applications using a parameterized synthetic benchmark · SC 2008 |
Memory systems
cache coherence |
0.0 | 1 | 1999 | Parallel Sorting on Cache-coherent DSM Multiprocessors · SC 1999 |
Memory systems › cache coherence
cache-coherent shared memory |
0.0 | 1 | 1999 | Parallel Sorting on Cache-coherent DSM Multiprocessors · SC 1999 |
Parallel and multicore computing
parallel algorithms |
0.0 | 1 | 1999 | Parallel Sorting on Cache-coherent DSM Multiprocessors · SC 1999 |
Parallel and multicore computing › parallel algorithms › sorting
parallel sorting |
0.0 | 1 | 1999 | Parallel Sorting on Cache-coherent DSM Multiprocessors · SC 1999 |
Interconnection networks and networks-on-chip
network topology |
0.0 | 1 | 2006 | Particles and contiuum - Performance modeling and optimization of a high energy colliding beam simulation code · SC 2006 |
Interconnection networks and networks-on-chip › network topology
torus network |
0.0 | 1 | 2006 | Particles and contiuum - Performance modeling and optimization of a high energy colliding beam simulation code · SC 2006 |
High-performance computing › performance engineering
performance portability |
0.0 | 1 | 1997 | Application Restructuring and Performance Portability on Shared Virtual Memory and Hardware-Coherent Multiprocessors · PPoPP 1997 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 2005 | Apex-Map: A Global Data Access Benchmark to Analyze HPC Systems and Parallel Programming Paradigms · SC 2005 |
Methods — techniques the papers use, named apart from their topics
matrix-vector multiplication · 0.2lanczos reorthogonalization · 0.2patching scheme · 0.2fragment method · 0.2divide-and-conquer · 0.2benchmarking · 0.1asynchronous i/o · 0.1performance modeling · 0.1microbenchmarking · 0.1synthetic performance probing · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | MPI usage at NERSC: Present and FutureabstractIn this poster, we describe how MPI is used at the National Energy Research Scientific Computing Center (NERSC) NERSC is the production high-performance computing center for the US Department of Energy, with more than 5000 users and 800 distinct projects. Through a variety of tools (e.g., User Survey, application team collaborations, etc.), we determine how MPI is used on our latest systems, with a particular focus on advanced features and how early applications intend to use MPI on NERSC's upcoming Intel Knights Landing (KNL) many-core system1 - one of the first to be deployed. In the poster, we also compare the usage of MPI to exascale developmental programming models such as UPC++ and HPX, with an eye on what features and extensions to MPI are plausible and useful for NERSC users. We also discuss perceived shortcomings of MPI, and why certain groups use other parallel programming models on the systems. In addition to a broad survey of the NERSC HPC population, we follow the evolution of a few key application codes2 that are being highly optimized for the KNL architecture using advanced OpenMP techniques. We study how these highly optimized on-node proxy apps and full applications start to make the transition to using full hybrid MPI+OpenMP implementations on the self-hosted KNL system. Alice E. Koniges, Brandon Cook 0001, Jack Deslippe, Thorsten Kurth, Hongzhang Shan |
EuroMPI | 5 |
| 2015 | Parallel implementation and performance optimization of the configuration-interaction methodabstractThe configuration-interaction (CI) method, long a popular approach to describe quantum many-body systems, is cast as a very large sparse matrix eigenpair problem with matrices whose dimension can exceed one billion. Such formulations place high demands on memory capacity and memory bandwidth --- two quantities at a premium today. In this paper, we describe an efficient, scalable implementation, BIGSTICK, which, by factorizing both the basis and the interaction into two levels, can reconstruct the nonzero matrix elements on the fly, reduce the memory requirements by one or two orders of magnitude, and enable researchers to trade reduced resources for increased computational time. We optimize BIGSTICK on two leading HPC platforms --- the Cray XC30 and the IBM Blue Gene/Q. Specifically, we not only develop an empirically-driven load balancing strategy that can evenly distribute the matrix-vector multiplication across 256K threads, we also developed techniques that improve the performance of the Lanczos reorthogonalization. Combined, these optimizations improved performance by 1.3-8× depending on platform and configuration. Hongzhang Shan, Samuel Williams 0001, Calvin W. Johnson, Kenneth S. McElvain, W. Erich Ormand |
SC | 1 |
| 2014 | UPC++: A PGAS Extension for C++abstractPartitioned Global Address Space (PGAS) languages are convenient for expressing algorithms with large, random-access data, and they have proven to provide high performance and scalability through lightweight one-sided communication and locality control. While very convenient for moving data around the system, PGAS languages have taken different views on the model of computation, with the static Single Program Multiple Data (SPMD) model providing the best scalability. In this paper we present UPC++, a PGAS extension for C++ that has three main objectives: 1) to provide an object-oriented PGAS programming model in the context of the popular C++ language, 2) to add useful parallel programming idioms unavailable in UPC, such as asynchronous remote function invocation and multidimensional arrays, to support complex scientific applications, 3) to offer an easy on-ramp to PGAS programming through interoperability with other existing parallel programming systems (e.g., MPI, OpenMP, CUDA). We implement UPC++ with a "compiler-free" approach using C++ templates and runtime libraries. We borrow heavily from previous PGAS languages and describe the design decisions that led to this particular set of language features, providing significantly more expressiveness than UPC with very similar performance characteristics. We evaluate the programmability and performance of UPC++ using five benchmarks on two representative supercomputers, demonstrating that UPC++ can deliver excellent performance at large scale up to 32K cores while offering PGAS productivity features to C++ applications. Yili Zheng, Amir Kamil, Michael B. Driscoll, Hongzhang Shan, Katherine A. Yelick |
IPDPS | 4 |
| 2010 | Developing a Parameterized Performance Proxy for Sequential Scientific KernelsabstractA simple, synthetic performance proxy for scientific applications is of great interest to the scientific computing community for the development of new products, procurements, and performance related questions in general. To develop such a performance proxy, we enhance the capability of the memory performance benchmark, Apex-MAP, by adding new concepts to capture the effects of computational details and programming styles. We test the fidelity of using Apex-MAP as a performance proxy with five sequential kernels on three different platforms with five inputs each. The relative performance difference between the kernels and Apex-MAP configured with corresponding parameters is generally within 10%. The quality of prediction measured by the coefficient of determination R^2 is over 98% for most cases. We also discuss experiences we gained during this study about how to improve the current version of Apex-MAP without affecting its basic concepts and designs so that it can reliably be used across platforms. Hongzhang Shan, Erich Strohmaier |
HPCC | 1 |
| 2009 | HPC global file system performance analysis using a scientific-application derived benchmark
Julian Borrill, Leonid Oliker, John Shalf, Hongzhang Shan, Andrew Uselton |
Parallel Comput. | 4 |
| 2008 | Characterizing and predicting the I/O performance of HPC applications using a parameterized synthetic benchmarkabstractThe unprecedented parallelism of new supercomputing platforms poses tremendous challenges to achieving scalable performance for I/O intensive applications. Performance assessments using traditional I/O system and component benchmarks are difficult to relate back to application I/O requirements. However, the complexity of full applications motivates development of simpler synthetic I/O benchmarks as proxies to the full application. In this paper we examine the I/O requirements of a range of HPC applications and describe how the LLNL IOR synthetic benchmark was chosen as suitable proxy for the diverse workload. We show a procedure for selecting IOR parameters to match the I/O patterns of the selected applications and show it can accurately predict the I/O performance of the full applications. We conclude that IOR is an effective replacement for full-application I/O benchmarks and can bridge the gap of understanding that typically exists between stand-alone benchmarks and the full applications they intend to model. Hongzhang Shan, Katie Antypas, John Shalf |
SC | 1 |
| 2008 | Linearly scaling 3D fragment method for large-scale electronic structure calculationsabstractWe present a new linearly scaling three-dimensional fragment (LS3DF) method for large scale ab initio electronic structure calculations. LS3DF is based on a divide-and-conquer approach, which incorporates a novel patching scheme that effectively cancels out the artificial boundary effects due to the subdivision of the system. As a consequence, the LS3DF program yields essentially the same results as direct density functional theory (DFT) calculations. The fragments of the LS3DF algorithm can be calculated separately with different groups of processors. This leads to almost perfect parallelization on over one hundred thousand processors. After code optimization, we were able to achieve 60.3 Tflop/s, which is 23.4% of the theoretical peak speed on 30,720 Cray XT4 processor cores. In a separate run on a BlueGene/P system, we achieved 107.5 Tflop/s on 131,072 cores, or 24.2% of peak. Our 13,824-atom ZnTeO alloy calculation runs 400 times faster than a direct DFT calculation, even presuming that the direct DFT calculation can scale well up to 17,280 processor cores. These results demonstrate the applicability of the LS3DF method to material simulations, the advantage of using linearly scaling algorithms over conventional O(N3) methods, and the potential for petascale computation using the LS3DF method. Lin-Wang Wang, Byounghak Lee, Hongzhang Shan, Zhengji Zhao, Juan C. Meza, Erich Strohmaier, David H. Bailey |
SC | 3 |
| 2007 | Scientific Application Performance on Candidate PetaScale PlatformsabstractAfter a decade where HEC (high-end computing) capability was dominated by the rapid pace of improvements to CPU clock frequency, the performance of next-generation supercomputers is increasingly differentiated by varying interconnect designs and levels of integration. Understanding the tradeoffs of these system designs, in the context of high-end numerical simulations, is a key step towards making effective petascale computing a reality. This work represents one of the most comprehensive performance evaluation studies to date on modern NEC systems, including the IBM Power5, AMD Opteron, IBM BG/L, and Cray X1E. A novel aspect of our study is the emphasis on full applications, with real input data at the scale desired by computational scientists in their unique domain. We examine six candidate ultra-scale applications, representing a broad range of algorithms and computational structures. Our work includes the highest concurrency experiments to date on five of our six applications, including 32K processor scalability for two of our codes and describe several successful optimizations strategies on BG/L, as well as improved X1E vectorization. Overall results indicate that our evaluated codes have the potential to effectively utilize petascale resources; however, several applications would require reengineering to incorporate the additional levels of parallelism necessary to achieve the vast concurrency of upcoming ultra-scale systems. Leonid Oliker, Andrew Canning, Jonathan Carter 0002, Costin Iancu, Michael Lijewski, Shoaib Kamil 0001, John Shalf, Hongzhang Shan, Erich Strohmaier, Stéphane Ethier, Tom Goodale |
IPDPS | 8 |
| 2007 | Investigation of leading HPC I/O performance using a scientific-application derived benchmarkabstractWith the exponential growth of high-fidelity sensor and simulated data, the scientific community is increasingly reliant on ultrascale HPC resources to handle their data analysis requirements. However, to utilize such extreme computing power effectively, the I/O components must be designed in a balanced fashion, as any architectural bottleneck will quickly render the platform intolerably inefficient. To understand I/O performance of data-intensive applications in realistic computational settings, we develop a lightweight, portable benchmark called MADbench2, which is derived directly from a large-scale Cosmic Microwave Background (CMB) data analysis package. Our study represents one of the most comprehensive I/O analyses of modern parallel filesystems, examining a broad range of system architectures and configurations, including Lustre on the Cray XT3 and Intel Itanium2 cluster; GPFS on IBM Power5 and AMD Opteron platforms; two BlueGene/L installations utilizing GPFS and PVFS2 filesystems; and CXFS on the SGI Altix3700. We present extensive synchronous I/O performance data comparing a number of key parameters including concurrency, POSIX- versus MPI-IO, and unique- versus shared-file accesses, using both the default environment as well as highly-tuned I/O parameters. Finally, we explore the potential of asynchronous I/O and quantify the volume of computation required to hide a given volume of I/O. Overall our study quantifies the vast differences in performance and functionality of parallel filesystems across state-of-the-art platforms, while providing system designers and computational scientists a lightweight tool for conducting further analyses. Julian Borrill, Leonid Oliker, John Shalf, Hongzhang Shan |
SC | 4 |
| 2007 | APEX-Map: a parameterized scalable memory access probe for high-performance computing systemsabstractAbstract The memory wall between the peak performance of microprocessors and their memory performance has become the prominent performance bottleneck for many scientific application codes. New benchmarks measuring data access speeds locally and globally in a variety of different ways are needed to explore the ever increasing diversity of architectures for high‐performance computing. In this paper, we introduce a novel benchmark, APEX‐Map, which focuses on global data movement and measures how fast global data can be fed into computational units. APEX‐Map is a parameterized, synthetic performance probe and integrates concepts for temporal and spatial locality into its design. Our first parallel implementation in MPI and various results obtained with it are discussed in detail. By measuring the APEX‐Map performance with parameter sweeps for a whole range of temporal and spatial localities performance surfaces can be generated. These surfaces are ideally suited to study the characteristics of the computational platforms and are useful for performance comparison. Results on a global‐memory vector platform and distributed‐memory superscalar platforms clearly reflect the design differences between these different architectures. Published in 2007 by John Wiley & Sons, Ltd. Erich Strohmaier, Hongzhang Shan |
Concurr. Comput. Pract. Exp. | 2 |
| 2006 | Performance Analysis of a High Energy Colliding Beam Simulation Code on Four HPC ArchitecturesabstractThe high energy colliders are essential to study the inner structure of nuclear and elementary particles. A parallel particle simulation code, BeamBeam3D, has been developed and actively used to model the beam dynamics and to optimize the performance of these colliders. In this paper, we analyzed the performance characteristics of BeamBeam3D on four leading high performance computing architectures, including a massive parallel system, a commodity-based cluster, an advanced vector platform, and a novel architecture focused on low power consumption and high density. We examine how to partition the workload among the processors to effectively use the computing resources, whether these platforms exhibit similar performance bottlenecks and how to address them, whether some platforms perform substantially better than others, and finally, the implications of BeamBeam3D for the design of the next generation supercomputer architectures Hongzhang Shan, Ji Qiang, Erich Strohmaier, Katherine A. Yelick |
ICPP | 1 |
| 2006 | Particles and contiuum - Performance modeling and optimization of a high energy colliding beam simulation codeabstractAn accurate modeling of the beam-beam interaction is essential to maximizing the luminosity in existing and future colliders. BeamBeam3D was the first parallel code that can be used to study this interaction fully self-consistently on high-performance computing platforms. Various all-to-all personalized communication (AAPC) algorithms dominate its communication patterns, for which we developed a sequence of performance models using a series of micro-benchmarks. We find that for SMP based systems the most important performance constraint is node-adapter contention, while for 3D-Torus topologies good performance models are not possible without considering link contention. The best average model prediction error is very low on SMP based systems with of 3% to 7%. On torus based systems errors of 29% are higher but optimized performance can again be predicted within 8% in some cases. These excellent results across five different systems indicate that this methodology for performance modeling can be applied to a large class of algorithms.1 Hongzhang Shan, Erich Strohmaier, Ji Qiang, David H. Bailey, Katherine A. Yelick |
SC | 1 |
| 2005 | Apex-Map: A Synthetic Scalable Benchmark Probe to Explore Data Access Performance on Highly Parallel Systems
Erich Strohmaier, Hongzhang Shan |
Euro-Par | 2 |
| 2005 | Apex-Map: A Global Data Access Benchmark to Analyze HPC Systems and Parallel Programming ParadigmsabstractThe memory wall and global data movement have become the dominant performance bottleneck for many scientific applications. New characterizations of data access streams and related benchmarks to measure their performances are therefore needed to compare HPC systems, software, and programming paradigms effectively. In this paper, we introduce a novel global data access benchmark, Apex-Map. It is a parameterized synthetic performance probe and integrates concepts for temporal and spatial locality into its design. We measured Apex-Map performance for a whole range of temporal and spatial localities on several advanced processors and parallel computing platforms and use the generated performance surfaces forperformance comparisons and to study the characteristics of these different architectures. We demonstrate that the results of Apex-Map clearly reflect many specific characteristics of the used systems. We also show the utility of Apex-Map for analyzing the performance effects of three leading parallel programming models and demonstrate their relative merits. Erich Strohmaier, Hongzhang Shan |
SC | 2 |
| 2004 | Performance characteristics of the Cray X1 and their implications for application performance tuningabstractDuring the last decade the scientific computing community has optimized many applications for execution on superscalar computing platforms. The recent arrival of the Japanese Earth Simulator has revived interest in vector architectures especially in the US. It is important to examine how to port our current scientific applications to the new vector platforms and how to achieve high performance. The success of porting these applications will also influence the acceptance of new vector architectures. In this paper, we first investigate the memory performance characteristics of the Cray X1, a recently released vector platform, and determine the most influential performance factors. Then, we examine how to optimize applications tuned on superscalar platforms for the Cray X1 using its performance characteristics as guidelines. Finally, we evaluate the different types of optimizations used, the effort for their implementations, and whether they provide any performance benefits when ported back to superscalar platforms. Hongzhang Shan, Erich Strohmaier |
ICS | 1 |
| 2003 | Job Superscheduler Architecture and Performance in Computational Grid EnvironmentsabstractComputational grids hold great promise in utilizing geographically separated heterogeneous resources to solve large-scale complex scientific problems. However, a number of major technical hurdles, including distributed resource management and effective job scheduling, stand in the way of realizing these gains. In this paper, we propose a novel grid superscheduler architecture and three distributed job migration algorithms. We also model the critical interaction between the superscheduler and autonomous local schedulers. Extensive performance comparisons with ideal, central, and local schemes using real workloads from leading computational centers are conducted in a simulation environment. Additionally, synthetic workloads are used to perform a detailed sensitivity analysis of our superscheduler. Several key metrics demonstrate that substantial performance gains can be achieved via smart superscheduling in distributed computational grids. Hongzhang Shan, Leonid Oliker, Rupak Biswas |
SC | 1 |
| 2003 | Message passing and shared address space parallelism on an SMP cluster
Hongzhang Shan, Jaswinder Pal Singh, Leonid Oliker, Rupak Biswas |
Parallel Comput. | 1 |
| 2002 | A Comparison of Three Programming Models for Adaptive Applications on the Origin2000
Hongzhang Shan, Jaswinder Pal Singh, Leonid Oliker, Rupak Biswas |
J. Parallel Distributed Comput. | 1 |
| 2001 | Message Passing Vs. Shared Address Space on a Clusters of SMPsabstractThe emergence of scalable computer architectures using clusters of PCs (or PC-SMPs) with commodity networking has made them attractive platforms for high-end scientific computing. Currently, message passing (MP) and shared address space (SAS) are the two leading programming paradigms for these systems. MP has been standardized with MPI, and is the most common and mature parallel programming approach. However, MP code development can be extremely difficult, especially for irregularly structured computations. SAS offers substantial ease of programming, but may suffer from performance limitations due to poor spatial locality and high protocol overhead. In this paper, we compare the performance of and programming effort required for six applications under both programming models on a 32-CPU PC-SMP cluster. Our application suite consists of codes that typically do not exhibit scalable performance under shared-memory programming due to their high communication-to-computation ratios and complex communication patterns. Results indicate that SAS can achieve about half the parallel efficiency of MPI for most of our applications; however on certain classes of problems, SAS performance is competitive with MPI. Hongzhang Shan, Jaswinder Pal Singh, Leonid Oliker, Rupak Biswas |
IPDPS | 1 |
| 2000 | A Comparison of Three Programming Models for Adaptive Applications on the Origin2000
Hongzhang Shan, Jaswinder Pal Singh, Leonid Oliker, Rupak Biswas |
SC | 1 |
| 1999 | A comparison of MPI, SHMEM and cache-coherent shared address space programming models on the SGI Origin2000abstractWe compare the performance of three major programming models -- a load-store cache-coherent shared address space (CC-SAS), message passing (MP) and the segmented SHMEM model -- on a modern, 64-processor hardware cache-coherent machine, one of the two major types of platforms upon which high-performance computing is converging. We focus on applications that are either regular and predictable or at least do not require fine-grained dynamic replication of irregularly accessed data. Within this class, we use programs with a range of important communication patterns. We examine whether the basic parallel algorithm and communication structuring approaches needed for best performance are similar or different among the models, whether some models have substantial performance advantages over others as problem size and number of processors change, what the sources of these performance differences are, where the programs spend their time, and whether substantial improvements can be obtained by mo... Hongzhang Shan, Jaswinder Pal Singh |
International Conference on Supercomputing | 1 |
| 1999 | Parallel Sorting on Cache-coherent DSM MultiprocessorsabstractThe performance of parallel sorting is not well understood on hardware cachecoherent shared address space (CC-SAS) multiprocessors, which increasingly dominate the market for tightly-coupled multiprocessing.We study two high-performance parallel sorting algorithms, radix and sample sorting, under three major programming models-a load-store CC-SAS, message passing, and the segmented SHMEM model-on a 64-processor SGI Origin2000.We observe surprisingly good speedups on this demanding application.The performance of radix sort is greatly affected by the programming model and particular implementation used.Sample sort exhibits more uniform performance across programming models on this platform, but it is usually not so good as that of the best radix sort for larger data sets if each is allowed to use the best programming model for itself.The best combination of algorithm and programming model is radix sorting under the SHMEM model for larger data sets and sample sorting under CC-SAS for smaller data sets.Recently, a new type of platform has begun to dominate tightly-coupled multiprocessing, and together with less tightly-coupled commodity clusters constitutes the stateof-the-art going forward in multiprocessing.This type of platform supports a shared address space with implicit coherent replication of data in hardware, even though memory is physically distributed.It is very different from the traditional message-passing 1 Hongzhang Shan, Jaswinder Pal Singh |
SC | 1 |
| 1997 | Application Restructuring and Performance Portability on Shared Virtual Memory and Hardware-Coherent MultiprocessorsabstractThe performance portability of parallel programs across a wide range of emerging coherent shared address space systems is not well understood. Programs that run well on efficient, hardware cache-coherent systems often do not perform well on less optimal or more commodity-based communication architectures. This paper studies this issue of performance portability, with the commodity communication architecture of interest being page-grained shared virtual memory. We begin with applications that perform well on moderat scale hardware cache-coherent systems, and find that they do not do so well on SVM systems. Then, we examine whether and how the applications can be improved for SVM systems --- through data structuring or algorithmic enhancements---and the nature and difficulty of the optimization. Finally, we examine the impact of the successful optimizations on hardware-coherent platforms themselves, to see whether they are helpful, harmful or neutral on those platforms. We develop a systematic methodology to explore optimizations in different structured classes. The results, and the difficulty of the optimizations, lead insight not only into performance portability but also into the viability of SVM as a platform for these types of applications. Dongming Jiang, Hongzhang Shan, Jaswinder Pal Singh |
PPoPP | 2 |