EDBT 2026 Demo / reviewers in the wild / expert
Erich Strohmaier
dblp:63/4265
· DBLP profile ↗
22ranked-venue papers
7as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 7 first-authorArtificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
8 papers |
High-performance computing · 33% Performance modeling and evaluation · 29% Memory systems · 16% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% |
Topics — the 25 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing
scientific computing systems |
0.2 | 2 | 2009 | Memory-efficient optimization of Gyrokinetic particle-to-grid interpolation for multicore processors · SC 2009 Linearly scaling 3D fragment method for large-scale electronic structure calculations · SC 2008 |
Performance modeling and evaluation
benchmarking |
0.2 | 4 | 2007 | A genetic algorithms approach to modeling the performance of memory-bound computations · SC 2007 Apex-Map: A Global Data Access Benchmark to Analyze HPC Systems and Parallel Programming Paradigms · SC 2005 TOP500 - TOP500 supercomputer · SC 2006 |
Performance modeling and evaluation
workload characterization |
0.1 | 3 | 2009 | Quantifying Locality In The Memory Access Patterns of HPC Applications · SC 2005 Apex-Map: A Global Data Access Benchmark to Analyze HPC Systems and Parallel Programming Paradigms · SC 2005 Memory-efficient optimization of Gyrokinetic particle-to-grid interpolation for multicore processors · SC 2009 |
High-performance computing
performance optimization at scale |
0.1 | 2 | 2008 | Linearly scaling 3D fragment method for large-scale electronic structure calculations · SC 2008 A genetic algorithms approach to modeling the performance of memory-bound computations · SC 2007 |
Parallel and multicore computing
parallel algorithms |
0.1 | 1 | 2009 | Memory-efficient optimization of Gyrokinetic particle-to-grid interpolation for multicore processors · SC 2009 |
High-performance computing › scientific computing systems
particle-in-cell simulation |
0.1 | 1 | 2009 | Memory-efficient optimization of Gyrokinetic particle-to-grid interpolation for multicore processors · SC 2009 |
High-performance computing › scientific computing systems
electronic structure calculation |
0.1 | 1 | 2008 | Linearly scaling 3D fragment method for large-scale electronic structure calculations · SC 2008 |
Distributed systems › communication optimization
communication-computation overlap |
0.1 | 1 | 2007 | Optimizing communication overlap for high-speed networks · PPoPP 2007 |
Distributed systems
communication optimization |
0.1 | 1 | 2007 | Optimizing communication overlap for high-speed networks · PPoPP 2007 |
Memory systems
memory-bound computation |
0.1 | 1 | 2007 | A genetic algorithms approach to modeling the performance of memory-bound computations · SC 2007 |
Embedded and real-time systems › real-time communication
message scheduling |
0.1 | 1 | 2007 | Optimizing communication overlap for high-speed networks · PPoPP 2007 |
Performance modeling and evaluation
performance prediction |
0.1 | 1 | 2007 | A genetic algorithms approach to modeling the performance of memory-bound computations · SC 2007 |
High-performance computing
collective communication |
0.1 | 1 | 2006 | Particles and contiuum - Performance modeling and optimization of a high energy colliding beam simulation code · SC 2006 |
Performance modeling and evaluation › communication modeling
communication performance modeling |
0.1 | 1 | 2006 | Particles and contiuum - Performance modeling and optimization of a high energy colliding beam simulation code · SC 2006 |
Memory systems › cache
cache performance |
0.1 | 1 | 2005 | Quantifying Locality In The Memory Access Patterns of HPC Applications · SC 2005 |
Memory systems
data locality |
0.1 | 1 | 2005 | Apex-Map: A Global Data Access Benchmark to Analyze HPC Systems and Parallel Programming Paradigms · SC 2005 |
Performance modeling and evaluation › workload characterization
locality analysis |
0.1 | 1 | 2005 | Quantifying Locality In The Memory Access Patterns of HPC Applications · SC 2005 |
Memory systems
memory access patterns |
0.1 | 1 | 2005 | Quantifying Locality In The Memory Access Patterns of HPC Applications · SC 2005 |
Memory systems › data locality
spatial and temporal locality |
0.1 | 1 | 2005 | Quantifying Locality In The Memory Access Patterns of HPC Applications · SC 2005 |
Computational science and engineering › computational chemistry › electronic structure calculation
density functional theory |
0.0 | 1 | 2008 | Linearly scaling 3D fragment method for large-scale electronic structure calculations · SC 2008 |
Computational science and engineering › materials science
materials science simulation |
0.0 | 1 | 2008 | Linearly scaling 3D fragment method for large-scale electronic structure calculations · SC 2008 |
Performance modeling and evaluation › benchmarking › parallel benchmark suites
HPC benchmark suite |
0.0 | 1 | 2006 | TOP500 - TOP500 supercomputer · SC 2006 |
Interconnection networks and networks-on-chip
network topology |
0.0 | 1 | 2006 | Particles and contiuum - Performance modeling and optimization of a high energy colliding beam simulation code · SC 2006 |
Interconnection networks and networks-on-chip › network topology
torus network |
0.0 | 1 | 2006 | Particles and contiuum - Performance modeling and optimization of a high energy colliding beam simulation code · SC 2006 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 2005 | Apex-Map: A Global Data Access Benchmark to Analyze HPC Systems and Parallel Programming Paradigms · SC 2005 |
Methods — techniques the papers use, named apart from their topics
patching scheme · 0.2fragment method · 0.2divide-and-conquer · 0.2load balancing · 0.1charge deposition kernel parallelization · 0.1heuristic search · 0.1genetic algorithm · 0.1analytic performance model · 0.1MultiMAPS · 0.1Apex-MAPS · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Analysis and Prediction of Data Transfer Throughput for Data-Intensive WorkloadsabstractScientific workflows are increasingly transferring large amounts of data between high performance computing (HPC) systems. Even though these HPC systems are connected via high-speed dedicated networks and use dedicated data transfer nodes (DTNs), it is still difficult to predict the data transfer throughput because of variations in data transfer protocols, host configurations, performance of file systems, and overlapping workloads. In order to provide reliable performance prediction for better resource management and job scheduling, we need models for predicting data transfer throughput under real-world conditions. In this paper, we explore different machine learning approaches for building data-driven models to improve performance and prediction of large-scale data transfer throughput. In addition to the variables already collected by the network monitoring system, we also develop heuristics to derive additional metrics for improving the prediction accuracy. We use the prediction results to identify the importance of different network parameters in predicting the throughput for large-scale data transfers. Through extensive tests, we identify key network parameters, discover interesting variations among different HPC sites, and show that we can predict throughput with high accuracy. We also analyze our models and results to provide recommendations for improving the performance of big data transfers. Devarshi Ghoshal, Kesheng Wu, Eric Pouyoul, Erich Strohmaier |
IEEE BigData | 4 |
| 2014 | A power-measurement methodology for large-scale, high-performance computingabstractImprovement in the energy efficiency of supercomputers can be accelerated by improving the quality and comparability of efficiency measurements. The ability to generate accurate measurements at extreme scale are just now emerging. The realization of system-level measurement capabilities can be accelerated with a commonly adopted and high quality measurement methodology for use while running a workload, typically a benchmark. This paper describes a methodology that has been developed collaboratively through the Energy Efficient HPC Working Group to support architectural analysis and comparative measurements for rankings, such as the Top500 and Green500. To support measurements with varying amounts of effort and equipment required we present three distinct levels of measurement, which provide increasing levels of accuracy. Level 1 is similar to the Green500 run rules today, a single average power measurement extrapolated from a subset of a machine. Level 2 is more comprehensive, but still widely achievable. Level 3 is the most rigorous of the three methodologies but is only possible at a few sites. However, the Level 3 methodology generates a high quality result that exposes details that the other methodologies may miss. In addition, we present case studies from the Leibniz Supercomputing Centre (LRZ), Argonne National Laboratory (ANL) and Calcul Québec Université Laval that explore the benefits and difficulties of gathering high quality, system-level measurements on large-scale machines. Thomas Scogland, Craig P. Steffen, Torsten Wilde, Florent Parent, Susan Coghlan, Natalie J. Bates, Wu-chun Feng, Erich Strohmaier |
ICPE | 8 |
| 2010 | Developing a Parameterized Performance Proxy for Sequential Scientific KernelsabstractA simple, synthetic performance proxy for scientific applications is of great interest to the scientific computing community for the development of new products, procurements, and performance related questions in general. To develop such a performance proxy, we enhance the capability of the memory performance benchmark, Apex-MAP, by adding new concepts to capture the effects of computational details and programming styles. We test the fidelity of using Apex-MAP as a performance proxy with five sequential kernels on three different platforms with five inputs each. The relative performance difference between the kernels and Apex-MAP configured with corresponding parameters is generally within 10%. The quality of prediction measured by the coefficient of determination R^2 is over 98% for most cases. We also discuss experiences we gained during this study about how to improve the current version of Apex-MAP without affecting its basic concepts and designs so that it can reliably be used across platforms. Hongzhang Shan, Erich Strohmaier |
HPCC | 2 |
| 2010 | Characterizing the Relation Between Apex-Map Synthetic Probes and Reuse Distance DistributionsabstractCharacterizing a memory reference stream using reuse distance distribution can enable predicting the performance on a given architecture. Benchmarks can subject an architecture to a limited set of reuse distance distributions, but it cannot exhaustively test it. In contrast, Apex-Map, a synthetic memory probe with parameterized locality, can provide a better coverage of the machine use scenarios. Unfortunately, it requires a lot of expertise to relate an application memory behavior to an Apex-Map parameter set. In this work we present a mathematical formulation that describes the relation between Apex-Map and reuse distance distributions. We also introduce a process through which we can automate the estimation of Apex-Map locality parameters for a given application. This process finds the best parameters for Apex-Map probes that generate a reuse distance distribution similar to that of the original application. We tested this scheme on benchmarks from Scalable Synthetic Compact Applications and Unbalanced Tree Search, and we show that this scheme provides an accurate Apex-Map parameterization with a small percentage of mismatch in reuse distance distributions, about 3% in average and less than 8% in the worst case, on the tested applications. Khaled Z. Ibrahim, Erich Strohmaier |
ICPP | 2 |
| 2009 | Memory-efficient optimization of Gyrokinetic particle-to-grid interpolation for multicore processorsabstractWe present multicore parallelization strategies for the particle-to-grid interpolation step in the Gyrokinetic Toroidal Code (GTC), a 3D particle-in-cell (PIC) application to study turbulent transport in magnetic-confinement fusion devices. Particle-grid interpolation is a known performance bottleneck in several PIC applications. In GTC, this step involves particles depositing charges to a 3D toroidal mesh, and multiple particles may contribute to the charge at a grid point. We design new parallel algorithms for the GTC charge deposition kernel, and analyze their performance on three leading multicore platforms. We implement thirteen different variants for this kernel and identify the best-performing ones given typical PIC parameters such as the grid size, number of particles per cell, and the GTC-specific particle Larmor radius variation. We find that our best strategies can be 2x faster than the reference optimized MPI implementation, and our analysis provides insight into desirable architectural features for high-performance PIC simulation codes. Kamesh Madduri, Samuel Williams 0001, Stéphane Ethier, Leonid Oliker, John Shalf, Erich Strohmaier, Katherine A. Yelick |
SC | 6 |
| 2008 | Power efficiency in high performance computingabstractAfter 15 years of exponential improvement in microprocessor clock rates, the physical principles allowing for Dennard scaling, which enabled performance improvements without a commensurate increase in power consumption, have all but ended. Until now, most HPC systems have not focused on power efficiency. However, as the cost of power reaches parity with capital costs, it is increasingly important to compare systems with metrics based on the sustained performance per watt. Therefore we need to establish practical methods to measure power consumption of such systems in- situ in order to support such metrics. Our study provides power measurements for various computational loads on the largest scale HPC systems ever involved in such an assessment. This study demonstrates clearly that, contrary to conventional wisdom, the power consumed while running the high performance Linpack (HPL) benchmark is very close to the power consumed by any subset of a typical compute-intensive scientific workload. Therefore, HPL, which in most cases cannot serve as a suitable workload for performance measurements, can be used for the purposes of power measurement. Furthermore, we show through measurements on a large scale system that the power consumed by smaller subsets of the system can be projected straightforwardly and accurately to estimate the power consumption of the full system. This allows a less invasive approach for determining the power consumption of large-scale systems. Shoaib Kamil 0001, John Shalf, Erich Strohmaier |
IPDPS | 3 |
| 2008 | Linearly scaling 3D fragment method for large-scale electronic structure calculationsabstractWe present a new linearly scaling three-dimensional fragment (LS3DF) method for large scale ab initio electronic structure calculations. LS3DF is based on a divide-and-conquer approach, which incorporates a novel patching scheme that effectively cancels out the artificial boundary effects due to the subdivision of the system. As a consequence, the LS3DF program yields essentially the same results as direct density functional theory (DFT) calculations. The fragments of the LS3DF algorithm can be calculated separately with different groups of processors. This leads to almost perfect parallelization on over one hundred thousand processors. After code optimization, we were able to achieve 60.3 Tflop/s, which is 23.4% of the theoretical peak speed on 30,720 Cray XT4 processor cores. In a separate run on a BlueGene/P system, we achieved 107.5 Tflop/s on 131,072 cores, or 24.2% of peak. Our 13,824-atom ZnTeO alloy calculation runs 400 times faster than a direct DFT calculation, even presuming that the direct DFT calculation can scale well up to 17,280 processor cores. These results demonstrate the applicability of the LS3DF method to material simulations, the advantage of using linearly scaling algorithms over conventional O(N3) methods, and the potential for petascale computation using the LS3DF method. Lin-Wang Wang, Byounghak Lee, Hongzhang Shan, Zhengji Zhao, Juan C. Meza, Erich Strohmaier, David H. Bailey |
SC | 6 |
| 2007 | Scientific Application Performance on Candidate PetaScale PlatformsabstractAfter a decade where HEC (high-end computing) capability was dominated by the rapid pace of improvements to CPU clock frequency, the performance of next-generation supercomputers is increasingly differentiated by varying interconnect designs and levels of integration. Understanding the tradeoffs of these system designs, in the context of high-end numerical simulations, is a key step towards making effective petascale computing a reality. This work represents one of the most comprehensive performance evaluation studies to date on modern NEC systems, including the IBM Power5, AMD Opteron, IBM BG/L, and Cray X1E. A novel aspect of our study is the emphasis on full applications, with real input data at the scale desired by computational scientists in their unique domain. We examine six candidate ultra-scale applications, representing a broad range of algorithms and computational structures. Our work includes the highest concurrency experiments to date on five of our six applications, including 32K processor scalability for two of our codes and describe several successful optimizations strategies on BG/L, as well as improved X1E vectorization. Overall results indicate that our evaluated codes have the potential to effectively utilize petascale resources; however, several applications would require reengineering to incorporate the additional levels of parallelism necessary to achieve the vast concurrency of upcoming ultra-scale systems. Leonid Oliker, Andrew Canning, Jonathan Carter 0002, Costin Iancu, Michael Lijewski, Shoaib Kamil 0001, John Shalf, Hongzhang Shan, Erich Strohmaier, Stéphane Ethier, Tom Goodale |
IPDPS | 9 |
| 2007 | Optimizing communication overlap for high-speed networksabstractModern networking hardware supports true non-blocking communicationand effective exploitation of this feature can lead to significantapplication performance improvements. We believe that algorithm design and optimization techniques that hide latency by taking advantage of communication overlap will facilitate obtaining good parallel efficiency and performance on the highly concurrent contemporary systems. Finding an optimal, performance portable implementation when using non-blocking communication primitives is non-trivial and intimidating to many application developers. In this paper we present a methodology for discovering optimal message sizes and schedules for a variety of application scenarios. This is achieved by combining an analytic model that takes into account the variability of performance parameters with system scale and load with heuristics designed to avoid network congestion. We perform experiments to understand network behavior in the presence of overlap and purge the optimization space for any system based on either resource or implementation constraints. Our approach isable to choose optimal or nearly optimal implementation parameters fora variety of highly non-trivial scenarios and networks with different performance characteristics. Implementations based on parameters chosen by the models are able to hide over 90% of communicationoverhead in all cases. Costin Iancu, Erich Strohmaier |
PPoPP | 2 |
| 2007 | A genetic algorithms approach to modeling the performance of memory-bound computationsabstractBenchmarks that measure memory bandwidth, such as STREAM, Apex-MAPS and MultiMAPS, are increasingly popular due to the "Von Neumann" bottleneck of modern processors which causes many calculations to be memory-bound. We present a scheme for predicting the performance of HPC applications based on the results of such benchmarks. A Genetic Algorithm approach is used to "learn" bandwidth as a function of cache hit rates per machine with MultiMAPS as the fitness test. The specific results are 56 individual performance predictions including 3 full-scale parallel applications run on 5 different modern HPC architectures, with various CPU counts and inputs, predicted within 10 % average difference with respect to independently verified runtimes. Mustafa M. Tikir, Laura Carrington, Erich Strohmaier, Allan Snavely |
SC | 3 |
| 2007 | APEX-Map: a parameterized scalable memory access probe for high-performance computing systemsabstractAbstract The memory wall between the peak performance of microprocessors and their memory performance has become the prominent performance bottleneck for many scientific application codes. New benchmarks measuring data access speeds locally and globally in a variety of different ways are needed to explore the ever increasing diversity of architectures for high‐performance computing. In this paper, we introduce a novel benchmark, APEX‐Map, which focuses on global data movement and measures how fast global data can be fed into computational units. APEX‐Map is a parameterized, synthetic performance probe and integrates concepts for temporal and spatial locality into its design. Our first parallel implementation in MPI and various results obtained with it are discussed in detail. By measuring the APEX‐Map performance with parameter sweeps for a whole range of temporal and spatial localities performance surfaces can be generated. These surfaces are ideally suited to study the characteristics of the computational platforms and are useful for performance comparison. Results on a global‐memory vector platform and distributed‐memory superscalar platforms clearly reflect the design differences between these different architectures. Published in 2007 by John Wiley & Sons, Ltd. Erich Strohmaier, Hongzhang Shan |
Concurr. Comput. Pract. Exp. | 1 |
| 2006 | Performance Analysis of a High Energy Colliding Beam Simulation Code on Four HPC ArchitecturesabstractThe high energy colliders are essential to study the inner structure of nuclear and elementary particles. A parallel particle simulation code, BeamBeam3D, has been developed and actively used to model the beam dynamics and to optimize the performance of these colliders. In this paper, we analyzed the performance characteristics of BeamBeam3D on four leading high performance computing architectures, including a massive parallel system, a commodity-based cluster, an advanced vector platform, and a novel architecture focused on low power consumption and high density. We examine how to partition the workload among the processors to effectively use the computing resources, whether these platforms exhibit similar performance bottlenecks and how to address them, whether some platforms perform substantially better than others, and finally, the implications of BeamBeam3D for the design of the next generation supercomputer architectures Hongzhang Shan, Ji Qiang, Erich Strohmaier, Katherine A. Yelick |
ICPP | 3 |
| 2006 | Particles and contiuum - Performance modeling and optimization of a high energy colliding beam simulation codeabstractAn accurate modeling of the beam-beam interaction is essential to maximizing the luminosity in existing and future colliders. BeamBeam3D was the first parallel code that can be used to study this interaction fully self-consistently on high-performance computing platforms. Various all-to-all personalized communication (AAPC) algorithms dominate its communication patterns, for which we developed a sequence of performance models using a series of micro-benchmarks. We find that for SMP based systems the most important performance constraint is node-adapter contention, while for 3D-Torus topologies good performance models are not possible without considering link contention. The best average model prediction error is very low on SMP based systems with of 3% to 7%. On torus based systems errors of 29% are higher but optimized performance can again be predicted within 8% in some cases. These excellent results across five different systems indicate that this methodology for performance modeling can be applied to a large class of algorithms.1 Hongzhang Shan, Erich Strohmaier, Ji Qiang, David H. Bailey, Katherine A. Yelick |
SC | 2 |
| 2006 | TOP500 - TOP500 supercomputerabstractNow in its 14th year, the TOP500 list of supercomputers serves as a "Who's Who" in the field of High Performance Computing (HPC). The TOP500 list was started in 1993 as a project to compile a list of the most powerful supercomputers in the world. It has evolved from a simple ranking system to a major source of information to analyze trends in HPC. The 28th TOP500 list will be published in November 2006 just in time for SC06.This BoF will present detailed analyses of the TOP500 and discuss the changes in the HPC marketplace during the past years. This includes the presentation of a new metric to track power consumption and updates on various benchmark initiatives such as the HPC Challenge benchmarks and the APEX-Map project. The BoF is meant as an open forum for discussion and feedback between the TOP500 authors and the user community. Erich Strohmaier |
SC | 1 |
| 2005 | Apex-Map: A Synthetic Scalable Benchmark Probe to Explore Data Access Performance on Highly Parallel Systems
Erich Strohmaier, Hongzhang Shan |
Euro-Par | 1 |
| 2005 | Apex-Map: A Global Data Access Benchmark to Analyze HPC Systems and Parallel Programming ParadigmsabstractThe memory wall and global data movement have become the dominant performance bottleneck for many scientific applications. New characterizations of data access streams and related benchmarks to measure their performances are therefore needed to compare HPC systems, software, and programming paradigms effectively. In this paper, we introduce a novel global data access benchmark, Apex-Map. It is a parameterized synthetic performance probe and integrates concepts for temporal and spatial locality into its design. We measured Apex-Map performance for a whole range of temporal and spatial localities on several advanced processors and parallel computing platforms and use the generated performance surfaces forperformance comparisons and to study the characteristics of these different architectures. We demonstrate that the results of Apex-Map clearly reflect many specific characteristics of the used systems. We also show the utility of Apex-Map for analyzing the performance effects of three leading parallel programming models and demonstrate their relative merits. Erich Strohmaier, Hongzhang Shan |
SC | 1 |
| 2005 | Quantifying Locality In The Memory Access Patterns of HPC ApplicationsabstractSeveral benchmarks for measuring the memory performance of HPC systems along dimensions of spatial and temporal memory locality have recently been proposed. However, little is understood about the relationships of these benchmarks to real applications and to each other. We propose a methodology for producing architecture-neutral characterizations of the spatial and temporal locality exhibited by the memory access patterns of applications. We demonstrate that the results track intuitive notions of locality on several synthetic and application benchmarks. We employ the methodology to analyze the memory performance components of the HPC Challenge Benchmarks, the Apex-MAP benchmark, and their relationships to each other and other benchmarks and applications. We show that this analysis can be used to both increase understanding of the benchmarks and enhance their usefulness by mapping them, along with applications, to a 2-D space along axes of spatial and temporal locality. Jonathan Weinberg, Michael O. McCracken, Erich Strohmaier, Allan Snavely |
SC | 3 |
| 2005 | Recent trends in the marketplace of high performance computing
Erich Strohmaier, Jack J. Dongarra, Hans Werner Meuer, Horst D. Simon |
Parallel Comput. | 1 |
| 2004 | Performance characteristics of the Cray X1 and their implications for application performance tuningabstractDuring the last decade the scientific computing community has optimized many applications for execution on superscalar computing platforms. The recent arrival of the Japanese Earth Simulator has revived interest in vector architectures especially in the US. It is important to examine how to port our current scientific applications to the new vector platforms and how to achieve high performance. The success of porting these applications will also influence the acceptance of new vector architectures. In this paper, we first investigate the memory performance characteristics of the Cray X1, a recently released vector platform, and determine the most influential performance factors. Then, we examine how to optimize applications tuned on superscalar platforms for the Cray X1 using its performance characteristics as guidelines. Finally, we evaluate the different types of optimizations used, the effort for their implementations, and whether they provide any performance benefits when ported back to superscalar platforms. Hongzhang Shan, Erich Strohmaier |
ICS | 2 |
| 1999 | The marketplace of high-performance computing
Erich Strohmaier, Jack J. Dongarra, Hans Werner Meuer, Horst D. Simon |
Parallel Comput. | 1 |
| 1997 | Statistical Performance Modeling: Case Study of the NPB 2.1 Results
Erich Strohmaier |
Euro-Par | 1 |
| 1997 | Changing technologies of HPC
Jack J. Dongarra, Hans Werner Meuer, Horst D. Simon, Erich Strohmaier |
Future Gener. Comput. Syst. | 4 |