EDBT 2026 Demo / reviewers in the wild / expert
Harvey J. Wasserman
dblp:41/4312
· DBLP profile ↗
14ranked-venue papers
3as first author
0since 2021 · last 2010
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 3 first-authorSoftware engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Performance modeling and evaluation · 56% High-performance computing · 18% Processor architecture and microarchitecture · 12% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation
analytical modeling |
0.0 | 1 | 2001 | Predictive performance and scalability modeling of a large-scale application · SC 2001 |
Performance modeling and evaluation
performance prediction |
0.0 | 1 | 2001 | Predictive performance and scalability modeling of a large-scale application · SC 2001 |
Performance modeling and evaluation › parallel system performance
scalability modeling |
0.0 | 1 | 2001 | Predictive performance and scalability modeling of a large-scale application · SC 2001 |
Memory systems
memory hierarchy |
0.0 | 1 | 1997 | Performance Evaluation of the SGI Origin2000: A Memory-Centric Characterization of LANL ASCI Applications · SC 1997 |
High-performance computing
scientific computing systems |
0.0 | 2 | 2001 | Predictive performance and scalability modeling of a large-scale application · SC 2001 Vectorization on Monte Carlo particle transport: an architectural study using the LANL benchmark "GAMTEB" · SC 1989 |
Performance modeling and evaluation
benchmarking |
0.0 | 2 | 1992 | The Performance Realities of Massively Parallel Processors: A Case Study · SC 1992 A performance comparison of three supercomputers: Fujitsu VP-2600, NEC SX-3, and CRAY Y-MP · SC 1991 |
Processor architecture and microarchitecture
vector processing |
0.0 | 3 | 1992 | Vectorization on Monte Carlo particle transport: an architectural study using the LANL benchmark "GAMTEB" · SC 1989 The Performance Realities of Massively Parallel Processors: A Case Study · SC 1992 Performance evaluation of the IBM RISC System/6000: comparison of an optimized scalar processor with two vector processors · SC 1990 |
Parallel and multicore computing › parallel architecture
massively parallel processor |
0.0 | 1 | 1992 | The Performance Realities of Massively Parallel Processors: A Case Study · SC 1992 |
Processor architecture and microarchitecture
SIMD |
0.0 | 1 | 1992 | The Performance Realities of Massively Parallel Processors: A Case Study · SC 1992 |
High-performance computing
supercomputing |
0.0 | 2 | 1997 | Performance Evaluation of the SGI Origin2000: A Memory-Centric Characterization of LANL ASCI Applications · SC 1997 The Performance Realities of Massively Parallel Processors: A Case Study · SC 1992 |
High-performance computing › supercomputing
supercomputer performance evaluation |
0.0 | 1 | 1991 | A performance comparison of three supercomputers: Fujitsu VP-2600, NEC SX-3, and CRAY Y-MP · SC 1991 |
High-performance computing › code optimization
vectorization |
0.0 | 1 | 1989 | Vectorization on Monte Carlo particle transport: an architectural study using the LANL benchmark "GAMTEB" · SC 1989 |
Performance modeling and evaluation › parallel system performance › speedup modeling
amdahl's law |
0.0 | 1 | 1989 | Vectorization on Monte Carlo particle transport: an architectural study using the LANL benchmark "GAMTEB" · SC 1989 |
Parallel and multicore computing › parallel computing
parallel computing environments |
0.0 | 1 | 1988 | Performance comparison of the Cray-2 and Cray X-MP/416 supercomputers · SC 1988 |
Methods — techniques the papers use, named apart from their topics
parametric modeling · 0.0analytical modeling · 0.0performance model for hierarchical memory systems · 0.0performance modeling · 0.0code porting · 0.0performance analysis · 0.0benchmarking · 0.0standard-fortran benchmark codes · 0.0vectorization · 0.0benchmark suite · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2010 | Performance Analysis of High Performance Computing Applications on the Amazon Web Services CloudabstractCloud computing has seen tremendous growth, particularly for commercial web applications. The on-demand, pay-as-you-go model creates a flexible and cost-effective means to access compute resources. For these reasons, the scientific computing community has shown increasing interest in exploring cloud computing. However, the underlying implementation and performance of clouds are very different from those at traditional supercomputing centers. It is therefore critical to evaluate the performance of HPC applications in today's cloud environments to understand the tradeoffs inherent in migrating to the cloud. This work represents the most comprehensive evaluation to date comparing conventional HPC platforms to Amazon EC2, using real applications representative of the workload at a typical supercomputing center. Overall results indicate that EC2 is six times slower than a typical mid-range Linux cluster, and twenty times slower than a modern HPC system. The interconnect on the EC2 cloud platform severely limits performance and causes significant variability. Keith R. Jackson, Lavanya Ramakrishnan, Krishna Muriki, Shane Canon, Shreyas Cholia, John Shalf, Harvey J. Wasserman, Nicholas J. Wright |
CloudCom | 7 |
| 2005 | A performance comparison between the Earth Simulator and other terascale systems on a characteristic ASCI workloadabstractAbstract This work gives a detailed analysis of the relative performance of the recently installed Earth Simulator and the next top four systems in the Top500 list using predictive performance models. The Earth Simulator uses vector processing nodes interconnected using a single‐stage, cross‐bar network, whereas the next top four systems are built using commodity based superscalar microprocessors and interconnection networks. The performance that can be achieved results from an interplay of system characteristics, application requirements and scalability behavior. Detailed performance models are used here to predict the performance of two codes representative of the ASCI workload, namely SAGE and Sweep3D. The performance models encapsulate fully the behavior of these codes and have been previously validated on many large‐scale systems. One result of this analysis is to size systems, built from the same nodes and networks as those in the top five, that will have the same performance as the Earth Simulator. In particular, the largest ASCI machine, ASCI Q, is expected to achieve a similar performance to the Earth Simulator on the representative workload. Published in 2005 by John Wiley & Sons, Ltd. Darren J. Kerbyson, Adolfy Hoisie, Harvey J. Wasserman |
Concurr. Pract. Exp. | 3 |
| 2001 | Predictive performance and scalability modeling of a large-scale applicationabstractIn this work we present a predictive analytical model that encompasses the performance and scaling characteristics of an important ASCI application. SAGE (SAIC's Adaptive Grid Eulerian hydrocode) is a multidimensional hydrodynamics code with adaptive mesh refinement. The model is validated against measurements on several systems including ASCI Blue Mountain, ASCI White, and a Compaq Alphaserver ES45 system showing high accuracy. It is parametric --- basic machine performance numbers (latency, MFLOPS rate, bandwidth) and application characteristics (problem size, decomposition method, etc.) serve as input. The model is applied to add insight into the performance of current systems, to reveal bottlenecks, and to illustrate where tuning efforts can be effective. We also use the model to predict performance on future systems. Darren J. Kerbyson, Henry J. Alme, Adolfy Hoisie, Fabrizio Petrini, Harvey J. Wasserman, Michael L. Gittings |
SC | 5 |
| 2000 | A General Predictive Performance Model for Wavefront Algorithms on Clusters of SMPsabstractWe propose and validate a closed-end, analytical, general, predictive performance model for applications based on wavefront algorithms on clusters of SMPs. Wavefront algorithms are ubiquitous in parallel computing, since they represent a means of enabling parallelism in computations that contain recurrences. Our particular interest in wavefront algorithms derives from their use in discrete ordinates neutral particle transport computations representative of ASCI, but other important uses are well known. The proposed model captures the tradeoff between processor utilization and communication requirements characteristics of wavefront algorithms. The general model can predict the performance of this class of applications on distributed architectures with a network of lower dimensionality compared to that of an MPP, of which clusters of SMPs are one example. We validate the model using a compact-application from the ASCI workload on a large-scale cluster of SGI Origin 2000s in existence at the Los Alamos National Laboratory. The proposed model validates well on all clusters configurations utilized. Adolfy Hoisie, Olaf M. Lubeck, Harvey J. Wasserman, Fabrizio Petrini, Hank Alme |
ICPP | 3 |
| 1997 | Performance Evaluation of the SGI Origin2000: A Memory-Centric Characterization of LANL ASCI ApplicationsabstractIn this paper the authors compare single processor performance of the SGI Origin and PowerChallenge and utilize a previously reported performance model for hierarchical memory systems to explain the results. Both the Origin and PowerChallenge use the same microprocessor (MIPS R10000) but have significant differences in their memory subsystems. Their memory model includes the effect of overlap between CPU and memory operations and allows them to infer the individual contributions of all three improvements in the Origin`s memory architecture and relate the effectiveness of each improvement to application characteristics. Harvey J. Wasserman, Olaf M. Lubeck, Federico Bassetti |
SC | 1 |
| 1996 | Benchmark Tests on the Digital Equipment Corporation Alpha AXP 21164-based AlphaServer 8400, Including a Comparison of Optimized Vector and Superscalar ProcessingabstractThis paper reports single-processor performance of the DEC 8400 system, a multi-cpu mainframe based on the DEC 21164 microprocessor.Performance is compared with single processors of the CRAY J90, the IBM RISC System/6000 Model 39H, and the SGI Power Onyx (75 MHz MIPS R8000).Benchmark codes representing Los Alamos applications with a range of computational characteristics were used.An important part of the comparison uses a particle transport application with two implementations, one that is highly vectorized and one that is unvectorizable with memory access patterns more suitable for superscalar processors.The results suggest that the best architecture/implementation match is the vectorizable code running on the vector processor. Harvey J. Wasserman |
International Conference on Supercomputing | 1 |
| 1992 | The Performance Realities of Massively Parallel Processors: A Case StudyabstractThe authors present the results of an architectural comparison of SIMD (single-instruction multiple-data) massive parallelism, as implemented in the Thinking Machines Corp. CM-2, and vector or concurrent-vector processing, as implemented in the Cray Research Inc., Y-MP/8. The comparison is based primarily upon three application codes taken from the LANL (Los Alamos National Laboratory) CM-2 workload. Tests were run by porting CM Fortran codes to the Y-MP, so that nearly the same level of optimization was obtained on both machines. The results for fully configured systems, using measured data rather than scaled data from smaller configurations, show that the Y-MP/8 is faster than the 64 k CM-2 for all three codes. A simple model that accounts for the relative characteristic computational speeds of the two machines, and reduction in overall CM-2 performance due to communication or SIMD conditional execution, accurately predicts the performance of two of the three codes. The authors show the similarity of the CM-2 and Y-MP programming models and comment on selected future massively parallel processor designs.> Olaf M. Lubeck, Margaret L. Simmons, Harvey J. Wasserman |
SC | 3 |
| 1991 | A performance comparison of three supercomputers: Fujitsu VP-2600, NEC SX-3, and CRAY Y-MPabstractThe performance of two second-generation supercomputers, the NEC SX-3 and the Fujitsu VP2600, is analyzed using the Standard Los Alamos Benchmark Set, the Mendez Fluid Dynamics Codes, and some highly vectorizable production-type codes from Los Alamos.For comparison, data are also given for a single processor of the CRAY Y-MP8/264.Factors affecting performance such as memory bandwidth, vector register organization, and the effects of multiple vector pipelines are examined.On a highly vectorizable code that can take advantage of multiple vector pipes, the SX-3 and VP2600 are faster than the single CRAY Y-MP processor by factors of seven to eight. Margaret L. Simmons, Harvey J. Wasserman, Olaf M. Lubeck, Christopher Eoyang, Raul Mendez, Hiroo Harada, Misako Ishiguro |
SC | 2 |
| 1990 | Performance evaluation of the IBM RISC System/6000: comparison of an optimized scalar processor with two vector processorsabstractThe authors report the performance of the 6000-series computers as measured using a set of portable, standard-Fortran, computationally intensive benchmark codes that represent the scientific workload at the Los Alamos National Laboratory. On all but three of the benchmark codes, the 40-ns RISC (reduced instruction set computer) system was able to perform as well as a single Convex C-240 processor, a vector processor that also has a 40-ns clock cycle, and, on these same codes, it performed as well as the FPS-500, a vector processor with a 30-ns clock cycle.> Margaret L. Simmons, Harvey J. Wasserman |
SC | 2 |
| 1990 | Performance comparison of the CRAY-2 and CRAY X-MP/416 supercomputers
Margaret L. Simmons, Harvey J. Wasserman |
J. Supercomput. | 2 |
| 1989 | Vectorization on Monte Carlo particle transport: an architectural study using the LANL benchmark "GAMTEB"abstractFully vectorized versions of the Los Alamos National Laboratory benchmark code Gamteb, a Monte Carlo photon transport algorithm, were developed for the Cyber 205/ETA-10 and Cray X-MP/Y-MP architectures. Single-processor performance measurements of the vector and scalar implementations were modeled in a modified Amdahl's Law that accounts for additional data motion in the vector code. The performance and implementation strategy of the vector codes are related to architectural features of each machine. Speedups between fifteen and eighteen for Cyber 205/ETA-10 architectures, and about nine for CRAY X-MP/Y-MP architectures are observed. The best single processor execution time for the problem was 0.33 seconds on the ETA-10G, and 0.42 seconds on the CRAY Y-MP. Patrick J. Burns 0002, Mark Christon, Roland Schweitzer, Olaf M. Lubeck, Harvey J. Wasserman |
SC | 5 |
| 1988 | Performance comparison of the Cray-2 and Cray X-MP/416 supercomputersabstractThe serial and parallel performance of the Cray-2 is analyzed using the standard Los Alamos benchmark set plus codes adopted for parallel processing. For comparison, architectural and performance data are given for the Cray-X-MP/416. Factors affecting performance, such as memory bandwidth, size and access speed of memory, and software exploitation of hardware, are examined. The parallel-processing environments of both machines are evaluated, and speedup measurements for the parallel codes are given.> Margaret L. Simmons, Harvey J. Wasserman |
SC | 2 |
| 1988 | The performance of minisupercomputers: Alliant FX/8, Convex C-1, and SCS-40
Harvey J. Wasserman, Margaret L. Simmons, Olaf M. Lubeck |
Parallel Comput. | 1 |
| 1988 | A Debugger for Parallel ProcessesabstractAbstract A system for analysing and debugging parallel Fortran codes is in use on the Sun workstation. The system is composed of a parallel processing simulatormtsim, a window‐ and mouse‐based debugging toolmtdbx, and a set of real‐time display routines. The simulatormtsim, which is called from Fortran by a set of routines having the same syntax as the CRAY X‐MP multitasking library, causes several concurrently active user tasks to be executed. The debuggermtdbxis based on the Sundbxtooldebugging facility. It has an enhanced command interface with functional control of parallel processes in multiple windows. The display routines offer two real‐time views of multitasking synchronization primitives as they are used during execution. These three components of the debugging system afford the opportunity to analyse the behaviour of parallel processes by dynamic interaction at run‐time. James H. Griffin, Harvey J. Wasserman, Lauren P. McGavran |
Softw. Pract. Exp. | 2 |