VLDB 2026 Research / reviewers in the wild / expert
Michael A. Heroux
dblp:21/1024 · also Mike Heroux
· DBLP profile ↗
30ranked-venue papers
6as first author
1since 2021 · last 2026
0000-0002-5893-0273ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 1 first-authorTheory of computation · 8 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Trilinos: Enabling Scientific Computing across Diverse Hardware Architectures at ScaleabstractTrilinos is a community-developed, open-source software framework that facilitates building large-scale, complex, multiscale, multiphysics simulation code bases for scientific and engineering problems. Since the Trilinos framework has undergone substantial changes to support new applications and new hardware architectures, this document is an update to “An Overview of the Trilinos project” by Heroux et al. (ACM Transactions on Mathematical Software, 31(3):397–423, 2005). It describes the design of Trilinos, introduces its new organization in product areas, and highlights established and new features available in Trilinos. Particular focus is put on the modernized software stack based on the Kokkos ecosystem to deliver performance portability across heterogeneous hardware architectures. This article also outlines the organization of the Trilinos community and the contribution model to help onboard interested users and contributors. Matthias Mayr, Alexander Heinlein, Christian A. Glusa, Sivasankaran Rajamanickam, Maarten Arnst, Roscoe A. Bartlett, Luc Berger-Vergiat, Erik G. Bowman, Karen D. Devine, Graham Harper, Michael A. Heroux, Mark Hoemmen, Jonathan J. Hu, Brian Michael Kelley, Kyungjoo Kim, Drew P. Kouri, Paul Kuberry, Kim Liegeois, Curtis C. Ober, Roger P. Pawlowski, Carl Pearson, Mauro Perego, Eric T. Phipps, Denis Ridzal, Nathan V. Roberts, Christopher M. Siefert, Heidi Thornquist, Romin Tomasetti, Christian Trott, Ray S. Tuminaro, James M. Willenbring, Michael M. Wolf, Ichitaro Yamazaki |
ACM Trans. Math. Softw. | 11 |
| 2018 | Special Issue on SCC'17 Reproducibility Initiative
C. Kristopher Garrett, Stephen Lien Harrell, Michael A. Heroux |
Parallel Comput. | 3 |
| 2017 | Special Issue on SC16 Student Cluster Competition Reproducibility Initiative
Michael A. Heroux, C. Kristopher Garrett |
Parallel Comput. | 1 |
| 2017 | Modeling and Simulating Multiple Failure Masking Enabled by Local Recovery for Stencil-Based Applications at Extreme ScalesabstractObtaining multi-process hard failure resilience at the application level is a key challenge that must be overcome before the promise of exascale can be fully realized. Previous work has shown that online global recovery can dramatically reduce the overhead of failures when compared to the more traditional approach of terminating the job and restarting it from the last stored checkpoint. If online recovery is performed in a local manner further scalability is enabled, not only due to the intrinsic lower costs of recovering locally, but also due to derived effects when using some application types. In this paper we model one such effect, namely multiple failure masking, that manifests when running Stencil parallel computations on an environment when failures are recovered locally. First, the delay propagation shape of one or multiple failures recovered locally is modeled to enable several analyses of the probability of different levels of failure masking under certain Stencil application behaviors. Our results indicate that failure masking is an extremely desirable effect at scale which manifestation is more evident and beneficial as the machine size or the failure rate increase. Marc Gamell, Keita Teranishi, Jackson R. Mayo, Hemanth Kolla, Michael A. Heroux, Jacqueline Chen, Manish Parashar |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2016 | Parallel subdomain solver strategies for the algebraic additive Schwarz preconditioner
Radu Popescu, Michael A. Heroux, Simone Deparis |
Parallel Comput. | 2 |
| 2015 | Exploring Failure Recovery for Stencil-based Applications at Extreme ScalesabstractApplication resilience is a key challenge that must be addressed in order to realize the exascale vision. Previous work has shown that online recovery, even when done in a global manner (i.e., involving all processes), can dramatically reduce the overhead of failures when compared to the more traditional approach of terminating the job and restarting it from the last stored checkpoint. In this paper we suggest going one step further, and explore how local recovery can be used for certain classes of applications to reduce the overheads due to failures. Specifically we study the feasibility of local recovery for stencil-based parallel applications and we show how multiple independent failures can be masked to effectively reduce the impact on the total time to solution. Marc Gamell, Keita Teranishi, Michael A. Heroux, Jackson R. Mayo, Hemanth Kolla, Jacqueline Chen, Manish Parashar |
HPDC | 3 |
| 2015 | Local recovery and failure masking for stencil-based applications at extreme scalesabstractApplication resilience is a key challenge that has to be addressed to realize the exascale vision. Online recovery, even when it involves all processes, can dramatically reduce the overhead of failures as compared to the more traditional approach where the job is terminated and restarted from the last checkpoint. In this paper we explore how local recovery can be used for certain classes of applications to further reduce overheads due to resilience. Specifically we develop programming support and scalable runtime mechanisms to enable online and transparent local recovery for stencil-based parallel applications on current leadership class systems. We also show how multiple independent failures can be masked to effectively reduce the impact on the total time to solution. We integrate these mechanisms with the S3D combustion simulation, and experimentally demonstrate (using the Titan Cray-XK7 system at ORNL) the ability to tolerate high failure rates (i.e., node failures every 5 seconds) with low overhead while sustaining performance, at scales up to 262144 cores. Marc Gamell, Keita Teranishi, Michael A. Heroux, Jackson R. Mayo, Hemanth Kolla, Jacqueline Chen, Manish Parashar |
SC | 3 |
| 2015 | Assessing a mini-application as a performance proxy for a finite element method engineering applicationabstractSummary The performance of a large‐scale, production‐quality science and engineering application (‘app’) is often dominated by a small subset of the code. Even within that subset, computational and data access patterns are often repeated, so that an even smaller portion can represent the performance‐impacting features. If application developers, parallel computing experts, and computer architects can together identify this representative subset and then develop a small mini‐application (‘miniapp’) that can capture these primary performance characteristics, then this miniapp can be used to both improve the performance of the app as well as provide a tool for co‐design for the high‐performance computing community. However, a critical question is whether a miniapp can effectively capture key performance behavior of an app. This study provides a comparison of an implicit finite element semiconductor device modeling app on unstructured meshes with an implicit finite element miniapp on unstructured meshes. The goal is to assess whether the miniapp is predictive of the performance of the app. Single compute node performance will be compared, as well as scaling up to 16,000 cores. Results indicate that the miniapp can be reasonably predictive of the performance characteristics of the app for a single iteration of the solver on a single compute node. Published 2015. This article is a U.S. Government work and is in the public domain in the USA. Paul T. Lin, Michael A. Heroux, Richard F. Barrett, Alan B. Williams |
Concurr. Comput. Pract. Exp. | 2 |
| 2015 | Assessing the role of mini-applications in predicting key performance characteristics of scientific and engineering applications
Richard F. Barrett, Paul S. Crozier, Douglas Doerfler, Michael A. Heroux, Paul T. Lin, Heidi Thornquist, Timothy G. Trucano, Courtenay T. Vaughan |
J. Parallel Distributed Comput. | 4 |
| 2015 | Editorial: ACM TOMS Replicated Computational Results InitiativeabstractThe scientific community relies on the peer review process for assuring the quality of published material, the goal of which is to build a body of work we can trust. Computational journals such as the ACM Transactions on Mathematical Software (TOMS) use this process for rigorously promoting the clarity and completeness of content, and citation of prior work. At the same time, it is unusual to independently confirm computational results. ACM TOMS has established a Replicated Computational Results (RCR) review process as part of the manuscript peer review process. The purpose is to provide independent confirmation that results contained in a manuscript are replicable. Successful completion of the RCR process awards a manuscript with the Replicated Computational Results Designation. This issue of ACM TOMS contains the first [Van Zee and van de Geijn 2015] of what we anticipate to be a growing number of articles to receive the RCR designation, and the related RCR reviewer report [Willenbring 2015]. We hope that the TOMS RCR process will serve as a model for other publications and increase the confidence in and value of computational results in TOMS articles. Michael A. Heroux |
ACM Trans. Math. Softw. | 1 |
| 2014 | Domain Decomposition Preconditioners for Communication-Avoiding Krylov Methods on a Hybrid CPU/GPU ClusterabstractKrylov subspace projection methods are widely used iterative methods for solving large-scale linear systems of equations. Researchers have demonstrated that communication avoiding (CA) techniques can improve Krylov methods' performance on modern computers, where communication is becoming increasingly expensive compared to arithmetic operations. In this paper, we extend these studies by two major contributions. First, we present our implementation of a CA variant of the Generalized Minimum Residual (GMRES) method, called CAGMRES, for solving no symmetric linear systems of equations on a hybrid CPU/GPU cluster. Our performance results on up to 120 GPUs show that CA-GMRES gives a speedup of up to 2.5x in total solution time over standard GMRES on a hybrid cluster with twelve Intel Xeon CPUs and three Nvidia Fermi GPUs on each node. We then outline a domain decomposition framework to introduce a family of preconditioners that are suitable for CA Krylov methods. Our preconditioners do not incur any additional communication and allow the easy reuse of existing algorithms and software for the sub domain solves. Experimental results on the hybrid CPU/GPU cluster demonstrate that CA-GMRES with preconditioning achieve a speedup of up to 7.4x over CAGMRES without preconditioning, and speedup of up to 1.7x over GMRES with preconditioning in total solution time. These results confirm the potential of our framework to develop a practical and effective preconditioned CA Krylov method. Ichitaro Yamazaki, Sivasankaran Rajamanickam, Erik G. Boman, Mark Hoemmen, Michael A. Heroux, Stanimire Tomov |
SC | 5 |
| 2014 | Exascale design space exploration and co-design
Sudip S. Dosanjh, Richard F. Barrett, Douglas Doerfler, Simon D. Hammond, Karl S. Hemmert, Michael A. Heroux, Paul T. Lin, Kevin T. Pedretti, Arun Rodrigues, Timothy G. Trucano, Justin Luitjens |
Future Gener. Comput. Syst. | 6 |
| 2012 | Overview of the TriBITS lifecycle model: A Lean/Agile software lifecycle model for research-based computational science and engineering softwareabstractSoftware lifecycles are becoming an increasingly important issue for computational science & engineering (CSE) software. The process by which a piece of CSE software begins life as a set of research requirements and then matures into a trusted high-quality capability is both commonplace and extremely challenging. Although an implicit lifecycle is obviously being used in any effort, the challenges of this process-respecting the competing needs of research vs. production-cannot be overstated. Here we describe a proposal for a well-defined software life-cycle process based on modern Lean/Agile software engineering principles. What we propose is appropriate for many CSE software projects that are initially heavily focused on research but also are expected to eventually produce usable high-quality capabilities. The model is related to TriBITS, a build, integration and testing system, which serves as a strong foundation for this lifecycle model, and aspects of this lifecycle model are ingrained in the TriBITS system. Indeed this lifecycle process, if followed, will enable large-scale sustainable integration of many complex CSE software efforts across several institutions. Roscoe A. Bartlett, Michael A. Heroux, James M. Willenbring |
eScience | 2 |
| 2012 | Toward codesign in high performance computing systemsabstractPreparations for exascale computing have led to the realization that computing environments will be significantly different from those that provide petascale capabilities. This change is driven by energy constraints, which has compelled hardware architects to design systems that will require a significant re-thinking of how application algorithms are selected and implemented. The "codesign" principle may offer a common basis for application and system developers as well as architects to work synergistically towards achieving exascale computing. This paper aims to introduce to the embedded system design community the unique challenges and opportunities as well as exciting developments in exascale HPC system codesign. Given the success of adopting codesign practices in the embedded system design area, this effort should be mutually beneficial to both communities. Richard F. Barrett, Xiaobo Sharon Hu, Sudip S. Dosanjh, Steven G. Parker, Michael A. Heroux, John Shalf |
ICCAD | 5 |
| 2012 | ShyLU: A Hybrid-Hybrid Solver for Multicore PlatformsabstractWith the ubiquity of multicore processors, it is crucial that solvers adapt to the hierarchical structure of modern architectures. We present ShyLU, a “hybrid-hybrid” solver for general sparse linear systems that is hybrid in two ways: First, it combines direct and iterative methods. The iterative part is based on approximate Schur complements where we compute the approximate Schur complement using a value-based dropping strategy or structure-based probing strategy. Second, the solver uses two levels of parallelism via hybrid programming (MPI+threads). ShyLU is useful both in shared-memory environments and on large parallel computers with distributed memory. In the latter case, it should be used as a subdomain solver. We argue that with the increasing complexity of compute nodes, it is important to exploit multiple levels of parallelism even within a single compute node. We show the robustness of ShyLU against other algebraic preconditioners. ShyLU scales well up to 384 cores for a given problem size. We also study the MPI-only performance of ShyLU against a hybrid implementation and conclude that on present multicore nodes MPI-only implementation is better. However, for future multicore machines (96 or more cores) hybrid/ hierarchical algorithms and implementations are important for sustained performance. Sivasankaran Rajamanickam, Erik G. Boman, Michael A. Heroux |
IPDPS | 3 |
| 2011 | Achieving Exascale Computing through Hardware/Software Co-design
Sudip S. Dosanjh, Richard F. Barrett, Michael A. Heroux, Arun Rodrigues |
EuroMPI | 3 |
| 2011 | Self-similarity of parallel machines
Robert W. Numrich, Michael A. Heroux |
Parallel Comput. | 2 |
| 2010 | A Light-weight API for Portable Multicore ProgrammingabstractMulticore nodes have become ubiquitous in just a few years. At the same time, writing portable parallel software for multicore nodes is extremely challenging. Widely available programming models such as OpenMP and Pthreads are not useful for devices such as graphics cards, and more flexible programming models such as RapidMind are only available commercially. OpenCL represents the first truly portable standard, but its availability is limited. In the presence of such transition, we have developed a minimal application programming interface (API) for multicore nodes that allows us to write portable parallel linear algebra software that can use any of the aforementioned programming models and any future standard models. We utilize C++ template meta-programming to enable users to write parallel kernels that can be executed on a variety of node types, including Cell, GPUs and multicore CPUs. The support for a parallel node is provided by implementing a Node object, according to the requirements specified by the API. This ability to provide custom support for particular node types gives developers a level of control not allowed by the current slate of proprietary parallel programming APIs. We demonstrate implementations of the API for a simple vector dot-product on sequential CPU, multicore CPU and GPU nodes. Christopher G. Baker, Michael A. Heroux, H. Carter Edwards, Alan B. Williams |
PDP | 2 |
| 2009 | Parallel Phase Model: A Programming Model for High-end Parallel Machines with ManycoresabstractThis paper presents a parallel programming model, Parallel Phase Model (PPM), for next-generation high-end parallel machines based on a distributed memory architecture consisting of a networked cluster of nodes with a large number of cores on each node. PPM has a unified high-level programming abstraction that facilitates the design and implementation of parallel algorithms to exploit both the parallelism of the many cores and the parallelism at the cluster level. The programming abstraction will be suitable for expressing both fine-grained and coarse-grained parallelism. It includes a few high-level parallel programming language constructs that can be added as an extension to an existing (sequential or parallel) programming language such as C; and the implementation of PPM also includes a light-weight runtime library that runs on top of an existing network communication software layer (e.g. MPI). Design philosophy of PPM and details of the programming abstraction are also presented. Several unstructured applications that inherently require high-volume random fine-grained data accesses have been implemented in PPM with very promising results. Ron Brightwell, Michael A. Heroux, Zhaofang Wen |
ICPP | 2 |
| 2008 | Initial Experiences with the BEC Parallel Programming EnvironmentabstractBundle-exchange-compute (BEC) is a new virtual shared memory parallel programming environment for distributed-memory machines. Different from and complementary to other global address space (GAS) programming model research efforts, BEC has built-in efficient support for unstructured applications that inherently require high-volume random fine-grained communication, such as parallel graph algorithms, sparse-matrices, and large-scale physics simulations. In BEC, the global view of shared data structures enables ease of algorithm design and programming; and for good application performance, fine-grained (random) accesses to shared data are automatically and dynamically bundled together for coarse-grained message-passing. BEC frees the users from explicit management of data distribution, locality, and communication. Therefore, BEC is much easier to program than MPI, while achieving comparable application performance. This paper presents some initial BEC applications, which show that simple BEC programs can match very complex and highly optimized MPI codes. Michael A. Heroux, Zhaofang Wen, Yuesheng Xu |
ISPDC | 1 |
| 2008 | PyTrilinos: High-performance distributed-memory solvers for PythonabstractPyTrilinos is a collection of Python modules that are useful for serial and parallel scientific computing. This collection contains modules that cover serial and parallel dense linear algebra, serial and parallel sparse linear algebra, direct and iterative linear solution techniques, domain decomposition and multilevel preconditioners, nonlinear solvers, and continuation algorithms. Also included are a variety of related utility functions and classes, including distributed I/O, coloring algorithms, and matrix generation. PyTrilinos vector objects are integrated with the popular NumPy Python module, gathering together a variety of high-level distributed computing operations with serial vector operations. PyTrilinos is a set of interfaces to existing, compiled libraries. This hybrid framework uses Python as front-end, and efficient precompiled libraries for all computationally expensive tasks. Thus, we take advantage of both the flexibility and ease of use of Python, and the efficiency of the underlying C++, C, and FORTRAN numerical kernels. Out numerical results show that, for many important problem classes, the overhead required by the Python interpreter is negligible. To run in parallel, PyTrilinos simply requires a standard Python interpreter. The fundamental MPI calls are encapsulated under an abstract layer that manages all interprocessor communications. This makes serial and parallel scripts using PyTrilinos virtually identical. Marzio Sala, William F. Spotz, Michael A. Heroux |
ACM Trans. Math. Softw. | 3 |
| 2008 | On the design of interfaces to sparse direct solversabstractWe discuss the design of general, flexible, consistent, reusable, and efficient interfaces to software libraries for the direct solution of systems of linear equations on both serial and distributed memory architectures. We introduce a set of abstract classes to access the linear system matrix elements and their distribution, access vector elements, and control the solution of the linear system. We describe a concrete implementation of the proposed interfaces, and report examples of applications and numerical results showing that the overhead induced by the object-oriented design is negligible under typical conditions of usage. We include examples of applications, and we comment on the advantages and limitations of the design. Marzio Sala, Kendall S. Stanley, Michael A. Heroux |
ACM Trans. Math. Softw. | 3 |
| 2007 | Optimal Kernels to Optimal Solutions: Algorithm and Software Issues in Solver DevelopmentabstractSummary form only given. Computer modeling and simulation efforts are most useful when they can provide an optimal solution to a problem. Toward this goal we need a vertical hierarchy of algorithmic capabilities and software tools, all of which must work together to make optimal solutions possible. In this presentation we discuss efforts in the Trilinos project to enable optimal solutions for scientific and engineering applications. We discuss how the goal of optimal solutions drives requirements not only for good optimization algorithms but also ever faster and more robust forward solvers and scalable kernels, as well as an abstraction layer to couple solver components. Michael A. Heroux |
PDP | 1 |
| 2007 | Improving the Development Process for CSE SoftwareabstractScientific and engineering programming has been around since the beginning of computing, often being the driving force for new system development and innovation. At the same time a continual focus on new modeling capabilities, and some apparent cultural issues, find software processes for many computational science and engineering (CSE) software projects lacking. Certainly there are notable exceptions, but our experience has been that CSE software projects, although committed to writing high-quality software, have few if any formal software processes and tools in place, and are often unaware of formal software quality assurance (SQA) concepts. Presently, increasing complexity of applications and a broad push to certify computations are dictating a higher standard for CSE software quality; it is no longer sufficient to claim to write high quality software. However, traditional software development models can be impractical for CSE projects to implement. Despite this, CSE software teams can benefit by implementing valuable SQA processes and tools. In this paper we outline some the processes and tools that are successfully used by the Trilinos Project. These tools and processes have been useful not only in increasing verifiable software quality, but also have improved overall software quality, and the development experience in general Michael A. Heroux, James M. Willenbring, Michael N. Phenow |
PDP | 1 |
| 2006 | An envolutionary path towards virtual shared memory with random accessabstractNo abstract available. Jonathan Leighton Brown, Sue Goudy, Michael A. Heroux, Shan Shan Huang, Zhaofang Wen |
SPAA | 3 |
| 2005 | An overview of the Trilinos projectabstractThe Trilinos Project is an effort to facilitate the design, development, integration, and ongoing support of mathematical software libraries within an object-oriented framework for the solution of large-scale, complex multiphysics engineering and scientific problems. Trilinos addresses two fundamental issues of developing software for these problems: (i) providing a streamlined process and set of tools for development of new algorithmic implementations and (ii) promoting interoperability of independently developed software.Trilinos uses a two-level software structure designed around collections of packages . A Trilinos package is an integral unit usually developed by a small team of experts in a particular algorithms area such as algebraic preconditioners, nonlinear solvers, etc. Packages exist underneath the Trilinos top level, which provides a common look-and-feel, including configuration, documentation, licensing, and bug-tracking.Here we present the overall Trilinos design, describing our use of abstract interfaces and default concrete implementations. We discuss the services that Trilinos provides to a prospective package and how these services are used by various packages. We also illustrate how packages can be combined to rapidly develop new algorithms. Finally, we discuss how Trilinos facilitates high-quality software engineering practices that are increasingly required from simulation software. Michael A. Heroux, Roscoe A. Bartlett, Victoria E. Howle, Robert J. Hoekstra, Jonathan J. Hu, Tamara G. Kolda, Richard B. Lehoucq, Kevin R. Long, Roger P. Pawlowski, Eric T. Phipps, Andrew G. Salinger, Heidi Thornquist, Ray S. Tuminaro, James M. Willenbring, Alan B. Williams, Kendall S. Stanley |
ACM Trans. Math. Softw. | 1 |
| 2004 | Vector reduction/transformation operatorsabstractDevelopment of flexible linear algebra interfaces is an increasingly critical issue. Efficient and expressive interfaces are well established for some linear algebra abstractions, but not for vectors. Vectors differ from other abstractions in the diversity of necessary operations, sometimes requiring dozens for a given algorithm (e.g. interior-point methods for optimization). We discuss a new approach based on operator objects that are transported to the underlying data by the linear algebra library implementation, allowing developers of abstract numerical algorithms to easily extend the functionality regardless of computer architecture, application or data locality/organization. Numerical experiments demonstrate efficient implementation. Roscoe A. Bartlett, Bart G. van Bloemen Waanders, Michael A. Heroux |
ACM Trans. Math. Softw. | 3 |
| 2002 | An overview of the sparse basic linear algebra subprograms: The new standard from the BLAS technical forumabstractWe discuss the interface design for the Sparse Basic Linear Algebra Subprograms (BLAS), the kernels in the recent standard from the BLAS Technical Forum that are concerned with unstructured sparse matrices. The motivation for such a standard is to encourage portable programming while allowing for library-specific optimizations. In particular, we show how this interface can shield one from concern over the specific storage scheme for the sparse matrix. This design makes it easy to add further functionality to the sparse BLAS in the future.We illustrate the use of the Sparse BLAS with examples in the three supported programming languages, Fortran 95, Fortran 77, and C. Iain S. Duff, Michael A. Heroux, Roldan Pozo |
ACM Trans. Math. Softw. | 2 |
| 1999 | Massively parallel computing: A Sandia perspective
David E. Womble, Sudip S. Dosanjh, Bruce Hendrickson, Michael A. Heroux, Steven J. Plimpton, James L. Tomkins, David S. Greenberg |
Parallel Comput. | 4 |
| 1998 | An Object-Oriented Framework for Block PreconditioningabstractGeneral software for preconditioning the iterative solution of linear systems is greatly lagging behind the literature. This is partly because specific problems and specific matrix and preconditioner data structures in order to be solved efficiently, i.e., multiple implementations of a preconditioner with specialized data structures are required. This article presents a framework to support preconditioning with various, possibly user-defined, data structures for matrices that are partitioned into blocks. The main idea is to define data structures for the blocks, and an upper layer of software which uses these blocks transparently of their data structure. This transparency can be accomplished by using an object-oriented language. Thus, various preconditioners, such as block relaxations and block-incomplete factorizations, only need to be defined once and will work with any block type. In addition, it is possible to transparently interchange various approximate or exact techniques for inverting pivot blocks, or solving systems whose coefficient matrices are diagonal blocks. This leads to a rich variety of preconditioners that can be selected. Operations with the blocks are performed with optimized libraries or fundamental data types. Comparisons with an optimized Fortran 77 code on both workstations and Cray supercomputers show that this framework can approach the efficiency of Fortran 77, as long as suitable block sized and block types are chosen. Edmond Chow, Michael A. Heroux |
ACM Trans. Math. Softw. | 2 |