EDBT 2026 Demo / reviewers in the wild / expert
David H. Bailey
dblp:74/2925
· DBLP profile ↗
30ranked-venue papers
19as first author
0since 2021 · last 2016
0000-0002-7574-8342ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 14 first-authorTheory of computation · 7 · 5 first-authorSoftware engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
16 papers |
High-performance computing · 48% Performance modeling and evaluation · 40% Memory systems · 6% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 78% Program analysis · 17% Software testing · 5% |
Topics — the 30 heaviest of 41, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization › approximate computing › precision tuning
floating-point precision tuning |
0.4 | 2 | 2016 | Floating-point precision tuning using blame analysis · ICSE 2016 Precimonious: tuning assistant for floating-point precision · SC 2013 |
Performance modeling and evaluation
benchmarking |
0.2 | 7 | 2009 | Misleading performance claims in parallel computations · DAC 2009 S12 - The HPC Challenge (HPCC) benchmark suite · SC 2006 ESP: A System Utilization Benchmark · SC 2000 |
High-performance computing
performance optimization at scale |
0.2 | 3 | 2016 | Linearly scaling 3D fragment method for large-scale electronic structure calculations · SC 2008 Floating-point precision tuning using blame analysis · ICSE 2016 ESP: A System Utilization Benchmark · SC 2000 |
Program analysis
dynamic analysis |
0.2 | 1 | 2013 | Precimonious: tuning assistant for floating-point precision · SC 2013 |
Compilers and program optimization › approximate computing › precision tuning
mixed-precision tuning |
0.2 | 1 | 2013 | Precimonious: tuning assistant for floating-point precision · SC 2013 |
Compilers and program optimization › approximate computing
precision tuning |
0.2 | 1 | 2013 | Precimonious: tuning assistant for floating-point precision · SC 2013 |
High-performance computing
scientific computing systems |
0.1 | 5 | 2008 | Linearly scaling 3D fragment method for large-scale electronic structure calculations · SC 2008 High performance computing meets experimental mathematics · SC 2002 Performance results for two of the NAS parallel benchmarks · SC 1991 |
High-performance computing › scientific computing systems
electronic structure calculation |
0.1 | 1 | 2008 | Linearly scaling 3D fragment method for large-scale electronic structure calculations · SC 2008 |
High-performance computing
collective communication |
0.1 | 1 | 2006 | Particles and contiuum - Performance modeling and optimization of a high energy colliding beam simulation code · SC 2006 |
Performance modeling and evaluation › communication modeling
communication performance modeling |
0.1 | 1 | 2006 | Particles and contiuum - Performance modeling and optimization of a high energy colliding beam simulation code · SC 2006 |
Performance modeling and evaluation › benchmarking › parallel benchmark suites
HPC benchmark suite |
0.1 | 1 | 2006 | S12 - The HPC Challenge (HPCC) benchmark suite · SC 2006 |
Memory systems
memory access patterns |
0.1 | 1 | 2006 | S12 - The HPC Challenge (HPCC) benchmark suite · SC 2006 |
Performance modeling and evaluation
workload characterization |
0.1 | 1 | 2006 | S12 - The HPC Challenge (HPCC) benchmark suite · SC 2006 |
High-performance computing
parallel numerical algorithms |
0.0 | 2 | 2002 | High performance computing meets experimental mathematics · SC 2002 A Strassen-Newton algorithm for high-speed parallelizable matrix inversion · SC 1988 |
High-performance computing
supercomputing |
0.0 | 2 | 2006 | S12 - The HPC Challenge (HPCC) benchmark suite · SC 2006 Massively parallel vs. parallel vector supercomputers: a user's perspective (panel) · SC 1993 |
Performance modeling and evaluation
parallel performance evaluation |
0.0 | 1 | 2009 | Misleading performance claims in parallel computations · DAC 2009 |
Computational science and engineering › computational chemistry › electronic structure calculation
density functional theory |
0.0 | 1 | 2008 | Linearly scaling 3D fragment method for large-scale electronic structure calculations · SC 2008 |
Computational science and engineering › materials science
materials science simulation |
0.0 | 1 | 2008 | Linearly scaling 3D fragment method for large-scale electronic structure calculations · SC 2008 |
Performance modeling and evaluation › benchmarking › parallel benchmark suites
NAS parallel benchmarks |
0.0 | 3 | 1992 | NAS Parallel Benchmark Results · SC 1992 Performance results for two of the NAS parallel benchmarks · SC 1991 The NAS parallel benchmarks - summary and preliminary results · SC 1991 |
Performance modeling and evaluation › benchmarking
parallel benchmark suites |
0.0 | 3 | 1992 | NAS Parallel Benchmark Results · SC 1992 Performance results for two of the NAS parallel benchmarks · SC 1991 The NAS parallel benchmarks - summary and preliminary results · SC 1991 |
Interconnection networks and networks-on-chip
network topology |
0.0 | 1 | 2006 | Particles and contiuum - Performance modeling and optimization of a high energy colliding beam simulation code · SC 2006 |
Interconnection networks and networks-on-chip › network topology
torus network |
0.0 | 1 | 2006 | Particles and contiuum - Performance modeling and optimization of a high energy colliding beam simulation code · SC 2006 |
High-performance computing
fast fourier transform |
0.0 | 2 | 1991 | Performance results for two of the NAS parallel benchmarks · SC 1991 FFTs in external of hierarchical memory · SC 1989 |
High-performance computing › supercomputer architecture
vector supercomputer |
0.0 | 2 | 1993 | Massively parallel vs. parallel vector supercomputers: a user's perspective (panel) · SC 1993 Vector Computer Memory Bank Contention · IEEE Trans. Computers 1987 |
Algorithms and data structures
number-theoretic algorithms |
0.0 | 1 | 2002 | High performance computing meets experimental mathematics · SC 2002 |
Processor architecture and microarchitecture
instruction set architecture |
0.0 | 1 | 1993 | RISC microprocessors and scientific computing · SC 1993 |
Processor architecture and microarchitecture › instruction set architecture › RISC
RISC processor |
0.0 | 1 | 1993 | RISC microprocessors and scientific computing · SC 1993 |
Performance modeling and evaluation › simulation
monte carlo methods |
0.0 | 1 | 1991 | Performance results for two of the NAS parallel benchmarks · SC 1991 |
High-performance computing › numerical linear algebra › linear solver
poisson solver |
0.0 | 1 | 1991 | Performance results for two of the NAS parallel benchmarks · SC 1991 |
High-performance computing › supercomputing
supercomputer performance evaluation |
0.0 | 3 | 1992 | NAS Parallel Benchmark Results · SC 1992 Misleading Performance in the Supercomputing Field · SC 1992 The NAS parallel benchmarks - summary and preliminary results · SC 1991 |
Methods — techniques the papers use, named apart from their topics
search-based optimization · 0.5dynamic analysis · 0.5blame analysis · 0.5search-based type tuning · 0.2patching scheme · 0.2fragment method · 0.2divide-and-conquer · 0.2arbitrary precision arithmetic · 0.1performance modeling · 0.1microbenchmarking · 0.1HPCC kernels · 0.1PSLQ integer relation algorithm · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Floating-point precision tuning using blame analysisabstractWhile tremendously useful, automated techniques for tuning the precision of floating-point programs face important scalability challenges. We present Blame Analysis, a novel dynamic approach that speeds up precision tuning. Blame Analysis performs floating-point instructions using different levels of accuracy for their operands. The analysis determines the precision of all operands such that a given precision is achieved in the final result of the program. Our evaluation on ten scientific programs shows that Blame Analysis is successful in lowering operand precision. As it executes the program only once, the analysis is particularly useful when targeting reductions in execution time. In such case, the analysis needs to be combined with search-based tools such as Precimonious. Our experiments show that combining Blame Analysis with Precimonious leads to obtaining better results with significant reduction in analysis time: the optimized programs execute faster (in three cases, we observe as high as 39.9% program speedup) and the combined analysis time is 9x faster on average, and up to 38x faster than Precimonious alone. Cindy Rubio-González, Cuong Nguyen 0001, Benjamin Mehne, Koushik Sen, James Demmel, William Kahan, Costin Iancu, Wim T. L. P. Lavrijsen, David H. Bailey, David Hough 0001 |
ICSE | 9 |
| 2014 | Automated simplification of large symbolic expressions
David H. Bailey, Jonathan M. Borwein, Alexander D. Kaiser |
J. Symb. Comput. | 1 |
| 2013 | High-precision computation: Applications and challenges [Keynote I]abstractSummary form only given, as follows. High-precision floating-point arithmetic software, ranging from "double-double" or "quad" precision to arbitrarily high-precision (hundreds or thousands of digits), has been available for years. Such facilities are standard features of Mathematica and Maple, and software packages such as MPFR, QD and ARPREC are available on the Internet. Some of these packages include high-level language interface modules that make conversion of standard-precision programs a relatively simple task. However, until recently such facilities were widely considered as novelty items - why would anyone need such exalted levels of numeric precision in "practical" research or engineering? In fact, during the past decade or two, numerous applications have arisen for high-precision floatingpoint arithmetic. This presentation will briefly describe some of these applications, which mostly arise in mathematical physics, applied physics and mathematics. Many heretofore unknown identities and relationships have been discovered, and features have been identified in computed data that were not "visible" with ordinary 64-bit precision. Applications of double-double (31 digits) or quad-double precision (62 digits) are particularly common, but there are also some interesting applications for as high as 50,000 digits. The speaker will also outline what is needed in improved facilities for high-precision computation to address challenges that lie ahead. David H. Bailey |
IEEE Symposium on Computer Arithmetic | 1 |
| 2013 | Extending Summation Precision for Network Reduction OperationsabstractDouble precision summation is at the core of numerous important algorithms such as Newton-Krylov methods and other operations involving inner products, but the effectiveness of summation is limited by the accumulation of rounding errors, which are an increasing problem with the scaling of modern HPC systems and data sets. To reduce the impact of precision loss, researchers have proposed increased- and arbitrary-precision libraries that provide reproducible error or even bounded error accumulation for large sums, but do not guarantee an exact result. Such libraries can also increase computation time significantly. We propose big integer (BigInt) expansions of double precision variables that enable arbitrarily large summations without error and provide exact and reproducible results. This is feasible with performance comparable to that of double-precision floating point summation, by the inclusion of simple and inexpensive logic into modern NICs to accelerate performance on large-scale systems. George Michelogiannakis, Xiaoye S. Li, David H. Bailey, John Shalf |
SBAC-PAD | 3 |
| 2013 | Precimonious: tuning assistant for floating-point precisionabstractGiven the variety of numerical errors that can occur, floating-point programs are difficult to write, test and debug. One common practice employed by developers without an advanced background in numerical analysis is using the highest available precision. While more robust, this can degrade program performance significantly. In this paper we present Precimonious, a dynamic program analysis tool to assist developers in tuning the precision of floating-point programs. Precimonious performs a search on the types of the floating-point program variables trying to lower their precision subject to accuracy constraints and performance goals. Our tool recommends a type instantiation that uses lower precision while producing an accurate enough answer without causing exceptions. We evaluate Precimonious on several widely used functions from the GNU Scientific Library, two NAS Parallel Benchmarks, and three other numerical programs. For most of the programs analyzed, Precimonious reduces precision, which results in performance improvements as high as 41%. Cindy Rubio-González, Cuong Nguyen 0001, Hong Diep Nguyen, James Demmel, William Kahan, Koushik Sen, David H. Bailey, Costin Iancu, David Hough 0001 |
SC | 7 |
| 2011 | High-precision numerical integration: Progress and challenges
David H. Bailey, Jonathan M. Borwein |
J. Symb. Comput. | 1 |
| 2009 | Misleading performance claims in parallel computationsabstractIn a previous humorous note entitled "Twelve Ways to Fool the Masses ...," I outlined twelve common ways in which performance figures for technical computer systems can be distorted. In this paper and accompanying conference talk, I give a reprise of these twelve "methods" and give some actual examples that have appeared in peer-reviewed literature in years past. I then propose guidelines for reporting performance, the adoption of which would raise the level of professionalism and reduce the level of confusion, not only in the world of device simulation but also in the larger arena of technical computing. David H. Bailey |
DAC | 1 |
| 2008 | Linearly scaling 3D fragment method for large-scale electronic structure calculationsabstractWe present a new linearly scaling three-dimensional fragment (LS3DF) method for large scale ab initio electronic structure calculations. LS3DF is based on a divide-and-conquer approach, which incorporates a novel patching scheme that effectively cancels out the artificial boundary effects due to the subdivision of the system. As a consequence, the LS3DF program yields essentially the same results as direct density functional theory (DFT) calculations. The fragments of the LS3DF algorithm can be calculated separately with different groups of processors. This leads to almost perfect parallelization on over one hundred thousand processors. After code optimization, we were able to achieve 60.3 Tflop/s, which is 23.4% of the theoretical peak speed on 30,720 Cray XT4 processor cores. In a separate run on a BlueGene/P system, we achieved 107.5 Tflop/s on 131,072 cores, or 24.2% of peak. Our 13,824-atom ZnTeO alloy calculation runs 400 times faster than a direct DFT calculation, even presuming that the direct DFT calculation can scale well up to 17,280 processor cores. These results demonstrate the applicability of the LS3DF method to material simulations, the advantage of using linearly scaling algorithms over conventional O(N3) methods, and the potential for petascale computation using the LS3DF method. Lin-Wang Wang, Byounghak Lee, Hongzhang Shan, Zhengji Zhao, Juan C. Meza, Erich Strohmaier, David H. Bailey |
SC | 7 |
| 2006 | S12 - The HPC Challenge (HPCC) benchmark suiteabstractIn 2003, the DARPA's High Productivity Computing Systems released the HPCC suite. It examines the performance of HPC architectures using kernels with various memory access patterns of well known computational kernels. Consequently, HPCC results bound the performance of real applications as a function of memory access characteristics and define performance boundaries of HPC architectures. The suite was intended to augment the TOP500 list and by now the results are publicly available for 6 out of 10 of the world's fastest computers. Implementations exist in most of the major high-end programming languages and environments, accompanied by countless optimization efforts. The increased publicity enjoyed by HPCC doesn't necessarily translate into deeper understanding of the performance issues that HPCC benchmarks. And so this tutorial will introduce attendees to HPCC, provide tools to examine differences in HPC architectures, and give hands-on training that will hopefully lead to better understanding of parallel environments. Piotr Luszczek, David H. Bailey, Jack J. Dongarra, Jeremy Kepner, Robert F. Lucas, Rolf Rabenseifner, Daisuke Takahashi |
SC | 2 |
| 2006 | Particles and contiuum - Performance modeling and optimization of a high energy colliding beam simulation codeabstractAn accurate modeling of the beam-beam interaction is essential to maximizing the luminosity in existing and future colliders. BeamBeam3D was the first parallel code that can be used to study this interaction fully self-consistently on high-performance computing platforms. Various all-to-all personalized communication (AAPC) algorithms dominate its communication patterns, for which we developed a sequence of performance models using a series of micro-benchmarks. We find that for SMP based systems the most important performance constraint is node-adapter contention, while for 3D-Torus topologies good performance models are not possible without considering link contention. The best average model prediction error is very low on SMP based systems with of 3% to 7%. On torus based systems errors of 29% are higher but optimized performance can again be predicted within 8% in some cases. These excellent results across five different systems indicate that this methodology for performance modeling can be applied to a large class of algorithms.1 Hongzhang Shan, Erich Strohmaier, Ji Qiang, David H. Bailey, Katherine A. Yelick |
SC | 4 |
| 2005 | Performance Modeling: Understanding the Past and Predicting the Future
David H. Bailey, Allan Snavely |
Euro-Par | 1 |
| 2002 | High performance computing meets experimental mathematicsabstractIn this paper we describe some novel applications of high performance computing in a discipline now known as "experimental mathematics." The paper reviews some recent published work, and then presents some new results that have not yet appeared in the literature. A key technique inovlved in this research is the PSLQ integer relation algorithm (recently named one of ten "algorithms of the century" by Computing in Science and Engineering). This algorithm permits one to recognize a numeric constant in terms of the formula that it satisfies. We present a variant of PSLQ that is well-suited for parallel computation, and give several examples of new mathematical results that we have found using it. Two of these computations were performed on highly parallel computers, since they are not feasible on conventional systems. We also describe a new software package for performing arbitrary precision arithmetic, which is required in this research. David H. Bailey, David John Broadhurst, Yozo Hida, Xiaoye S. Li, Brandon Thompson |
SC | 1 |
| 2002 | Design, implementation and testing of extended and mixed precision BLASabstractThis article describes the design rationale, a C implementation, and conformance testing of a subset of the new Standard for the BLAS (Basic Linear Algebra Subroutines): Extended and Mixed Precision BLAS. Permitting higher internal precision and mixed input/output types and precisions allows us to implement some algorithms that are simpler, more accurate, and sometimes faster than possible without these features. The new BLAS are challenging to implement and test because there are many more subroutines than in the existing Standard, and because we must be able to assess whether a higher precision is used for internal computations than is used for either input or output variables. We have therefore developed an automated process of generating and systematically testing these routines. Our methodology is applicable to languages besides C. In particular, our algorithms used in the testing code will be valuable to all other BLAS implementors. Our extra precision routines achieve excellent performance---close to half of the machine peak Megaflop rate even for the Level 2 BLAS, when the data access is stride one. Xiaoye S. Li, James Demmel, David H. Bailey, Greg Henry, Yozo Hida, Jimmy Iskandar, William Kahan, Suh Y. Kang, Anil Kapur, Michael C. Martin, Brandon Thompson, Teresa Tung, Daniel J. Yoo |
ACM Trans. Math. Softw. | 3 |
| 2001 | Algorithms for Quad-Double Precision Floating Point ArithmeticabstractA quad-double number is an unevaluated sum of four IEEE double precision numbers, capable of representing at least 212 bits of significand. We present the algorithms for various arithmetic operations (including the four basic operations and various algebraic and transcendental operations) on quad-double numbers. The performance of the algorithms, implemented in C++, is also presented. Yozo Hida, Xiaoye S. Li, David H. Bailey |
IEEE Symposium on Computer Arithmetic | 3 |
| 2000 | System Utilization Benchmark on the Cray T3E and IBM SP
Adrian T. Wong, Leonid Oliker, William T. Kramer, Teresa L. Kaltz, David H. Bailey |
JSSPP | 5 |
| 2000 | ESP: A System Utilization BenchmarkabstractThis article describes a new benchmark, called the Effective System Performance (ESP) test, which is designed to measure system-level performance, including such factors as job scheduling efficiency, handling of large jobs and shutdown-reboot times. In particular, this test can be used to study the effects of various scheduling policies and parameters. We present here some results that we have obtained so far on the Cray T3E and IBM SP systems, together with insights obtained from simulations. Adrian T. Wong, Leonid Oliker, William T. Kramer, Teresa L. Kaltz, David H. Bailey |
SC | 5 |
| 1995 | A Fortran-90 Based Multiprecision SystemabstractA new version of a Fortran multiprecision computation system, based on the Fortran 90 language, is described. With this new approach, a translator program is not required—translation of Fortran code for multiprecision is accomplished by merely utilizing advanced features of Fortran 90, such as derived data types and operator extensions. This approach results in more-reliable translation and permits programmers of multiprecision applications to utilize the full power of Fortran 90. Three multiprecision data types are supported in this system: multiprecision integer, real, and complex. All the usual Fortran conventions for mixed-mode operations are supported, and many of the Fortran intrinsics, such as SIN, EXP, and MOD, are supported with multiprecision arguments. An interesting application of this software, wherein new number-theoretic identities have been discovered by means of multiprecision computations, is included also. David H. Bailey |
ACM Trans. Math. Softw. | 1 |
| 1993 | RISC microprocessors and scientific computingabstractThis paper discusses design features in currently available RISC microprocessors that result in less-than-optimal sustained performance on large-scale scientific calculations. Recommendations for future designs are suggested. The author is with the Numerical Aerodynamic Simulation (NAS) Systems Division at NASA Ames Research Center, Moffett Field, CA 94035. 1. Introduction Scientists accustomed to running large-scale, computationally intensive applications have traditionally utilized conventional vector supercomputers, such as those manufactured by Cray Research, Inc., Fujitsu or NEC. However, with the recent popularization of RISC workstations, many of these same scientists are using their workstations not just to edit their source codes and display their results, but to perform their computations as well. Another avenue from which scientific computer users have been introduced to RISC processors, an avenue which is potentially very significant for the future of scientific computin... David H. Bailey |
SC | 1 |
| 1993 | Massively parallel vs. parallel vector supercomputers: a user's perspective (panel)abstractNo abstract available. Gary Mountry, David H. Bailey, Eugene D. Brooks III, David W. Forslund, Robert J. Harrison, Don Eric Heller, Tom Kraay |
SC | 2 |
| 1993 | Algorithm 719: Multiprecision translation and execution of FORTRAN programsabstractThis paper describes two Fortran utilities for multiprecision computation. The first is a package of Fortran subroutines that perform a variety of arithmetic operations and transcendental functions on floating point numbers of arbitrarily high precision. This package is in some cases over 200 times faster than that of certain other packages that have been developed for this purpose. The second utility is a translator program, which facilitates the conversion of ordinary Fortran programs to use this package. By means of source directives (special comments) in the original Fortran program, the user declares the precision level and specifies which variables in each subprogram are to be treated as multiprecision. The translator program reads this source program and outputs a program with the appropriate multiprecision subroutine calls. This translator supports multiprecision integer, real, and complex datatypes. The required array space for multiprecision data types is automatically allocated. In the evaluation of computational expressions, all of the usual conventions for operator precedence and mixed mode operations are upheld. Furthermore, most of the Fortran-77 intrinsics, such as ABS, MOD, NINT, COS, EXP are supported and produce true multiprecision values. David H. Bailey |
ACM Trans. Math. Softw. | 1 |
| 1992 | Misleading Performance in the Supercomputing FieldabstractThe problems of misleading performance reporting and the evident lack of careful refereeing in the supercomputing field are discussed in detail. Included are some examples that have appeared in recently published scientific papers. Some guidelines for reporting performance are presented, the adoption of which would raise the level of professionalism and reduce the level of confusion in the field of supercomputing.> David H. Bailey |
SC | 1 |
| 1992 | NAS Parallel Benchmark ResultsabstractThe NAS (Numerical Aerodynamic Simulation) parallel benchmarks have been developed at NASA Ames Research Center to study the performance of parallel supercomputer. The eight benchmark problems are specified in a 'pencil and paper' fashion. The performance results of various systems using the NAS parallel benchmarks are presented. These results represent the best results that have been reported to the authors for the specific systems listed. They represent implementation efforts performed by personnel in both the NAS Applied Research Branch of NASA Ames Research Center and in other organizations.> David H. Bailey, Leonardo Dagum, Eric Barszcz, Horst D. Simon |
SC | 1 |
| 1991 | The NAS parallel benchmarks - summary and preliminary resultsabstractArticle Free Access Share on The NAS parallel benchmarks—summary and preliminary results Authors: D. H. Bailey Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , E. Barszcz Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , J. T. Barton Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , D. S. Browning Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , R. L. Carter Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , L. Dagum Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , R. A. Fatoohi Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , P. O. Frederickson Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , T. A. Lasinski Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , R. S. Schreiber Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , H. D. Simon Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , V. Venkatakrishnan Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile , S. K. Weeratunga Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CA Numerical Aerodynamic Simulation (NAS) Systems Division, NASA Ames Research Center, Mail Stop T045-1, Moffett Field, CAView Profile Authors Info & Claims Supercomputing '91: Proceedings of the 1991 ACM/IEEE conference on SupercomputingAugust 1991 Pages 158–165https://doi.org/10.1145/125826.125925Published:01 August 1991Publication History 405citation1,162DownloadsMetricsTotal Citations405Total Downloads1,162Last 12 Months161Last 6 weeks25 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF David H. Bailey, Eric Barszcz, John T. Barton, D. S. Browning, Robert L. Carter, Leonardo Dagum, Rod A. Fatoohi, Paul O. Frederickson, T. A. Lasinski, Robert Schreiber, Horst D. Simon, V. Venkatakrishnan, Sisira Weeratunga |
SC | 1 |
| 1991 | Performance results for two of the NAS parallel benchmarksabstractTwo problems from the recently published NAS Parallel Benchmarks have been implemented on three advanced computer systems. These two benchmarks are the following: (1) an embarrassingly parallel Monte Carlo statistical calculation and (2) a Poisson partial differential equation solver based on three-dimensiona l fast Fourier transforms. The first requires virtually no interprocessor communication, while the second has a substantial communication re quirement. This paper briefly describes the two problems stud ied, discusses the implementation schemes employed, and gives performance results on the Cray Y-MP, the Intel iPSC/860 and the Connection Machine-2. David H. Bailey, Paul O. Frederickson |
SC | 1 |
| 1991 | Using Strassen's algorithm to accelerate the solution of linear systems
David H. Bailey, King Lee, Horst D. Simon |
J. Supercomput. | 1 |
| 1990 | FFTs in external or hierarchical memory
David H. Bailey |
J. Supercomput. | 1 |
| 1989 | FFTs in external of hierarchical memoryabstractConventional algorithms for computing large one-dimensional fast Fourier transforms (FFTs), even those algorithms recently developed for vector and parallel computers, are largely unsuitable for systems with external or hierarchical memory. The principal reason for this is the fact that most FFT algorithms require at least m complete passes through the data set to compute a 2m-point FFT. David H. Bailey |
SC | 1 |
| 1988 | A Strassen-Newton algorithm for high-speed parallelizable matrix inversionabstractTechniques are described for computing matrix inverses by algorithms that are highly suited to massively parallel computation. The techniques are based on an algorithm suggested by V. Strassen (1969). Variations of this scheme use matrix Newton iterations and other methods to improve the numerical stability while at the same time preserving a very high level of parallelism. One-processor Cray-2 implementations of these schemes range from one that is up to 55% faster than a conventional library routine to one that is slower than a library routine but achieves excellent numerical stability. The problem of computing the solution to a single set of linear equations is discussed, and it is shown that this problem can also be solved efficiently using these techniques.> David H. Bailey, Helaman R. P. Gerguson |
SC | 1 |
| 1987 | Vector Computer Memory Bank ContentionabstractA number of recent vector supercomputer designs have featured main memories with very large capacities, and presumably even larger memories are planned for future generations. While the memory chips used in these computers can store much larger amounts of data than before, their operation speeds are rather slow when compared to the significantly faster CPU (central processing unit) circuitry in new supercomputer designs. A consequence of this speed disparity between CPU's and main memory is that memory access times and memory bank reservation times (as measured in CPU ticks) are sharply increased from previous generations. David H. Bailey |
IEEE Trans. Computers | 1 |
| 1987 | A high-performance fast Fourier transform algorithm for the Cray-2
David H. Bailey |
J. Supercomput. | 1 |