Sven Hammarling

dblp:35/6648 · DBLP profile ↗
← Back
15ranked-venue papers
1as first author
1since 2021 · last 2021
0000-0003-3865-4897ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Theory of computation · 10 · 1 first-author · 1 since 2021Systems, architecture and hardware · 4Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2021 A Set of Batched Basic Linear Algebra Subprograms and LAPACK Routines
abstract
This article describes a standard API for a set of Batched Basic Linear Algebra Subprograms (Batched BLAS or BBLAS). The focus is on many independent BLAS operations on small matrices that are grouped together and processed by a single routine, called a Batched BLAS routine. The matrices are grouped together in uniformly sized groups, with just one group if all the matrices are of equal size. The aim is to provide more efficient, but portable, implementations of algorithms on high-performance many-core platforms. These include multicore and many-core CPU processors, GPUs and coprocessors, and other hardware accelerators with floating-point compute facility. As well as the standard types of single and double precision, we also include half and quadruple precision in the standard. In particular, half precision is used in many very large scale applications, such as those associated with machine learning.
Ahmad Abdelfattah, Timothy B. Costa, Jack J. Dongarra, Mark Gates, Azzam Haidar, Sven Hammarling, Nicholas J. Higham, Jakub Kurzak, Piotr Luszczek, Stanimire Tomov, Mawussi Zounon
ACM Trans. Math. Softw.6
2019 PLASMA: Parallel Linear Algebra Software for Multicore Using OpenMP
abstract
The recent version of the Parallel Linear Algebra Software for Multicore Architectures (PLASMA) library is based on tasks with dependencies from the OpenMP standard. The main functionality of the library is presented. Extensive benchmarks are targeted on three recent multicore and manycore architectures, namely, an Intel Xeon, Intel Xeon Phi, and IBM POWER 8 processors.
Jack J. Dongarra, Mark Gates, Azzam Haidar, Jakub Kurzak, Piotr Luszczek, Panruo Wu, Ichitaro Yamazaki, Asim YarKhan, Maksims Abalenkovs, Negin Bagherpour, Sven Hammarling, Jakub Sístek, David Stevens, Mawussi Zounon, Samuel D. Relton
ACM Trans. Math. Softw.11
2017 Optimized Batched Linear Algebra for Modern Architectures
Jack J. Dongarra, Sven Hammarling, Nicholas J. Higham, Samuel D. Relton, Mawussi Zounon
Euro-Par2
2013 An algorithm for the complete solution of quadratic eigenvalue problems
abstract
We develop a new algorithm for the computation of all the eigenvalues and optionally the right and left eigenvectors of dense quadratic matrix polynomials. It incorporates scaling of the problem parameters prior to the computation of eigenvalues, a choice of linearization with favorable conditioning and backward stability properties, and a preprocessing step that reveals and deflates the zero and infinite eigenvalues contributed by singular leading and trailing matrix coefficients. The algorithm is backward-stable for quadratics that are not too heavily damped. Numerical experiments show that our MATLAB implementation of the algorithm, quadeig, outperforms the MATLAB function polyeig in terms of both stability and efficiency.
Sven Hammarling, Christopher J. Munro, Françoise Tisseur
ACM Trans. Math. Softw.1
2008 Cache efficient bidiagonalization using BLAS 2.5 operators
abstract
On cache based computer architectures using current standard algorithms, Householder bidiagonalization requires a significant portion of the execution time for computing matrix singular values and vectors. In this paper we reorganize the sequence of operations for Householder bidiagonalization of a general m × n matrix, so that two (_GEMV) vector-matrix multiplications can be done with one pass of the unreduced trailing part of the matrix through cache. Two new BLAS operations approximately cut in half the transfer of data from main memory to cache, reducing execution times by up to 25 per cent. We give detailed algorithm descriptions and compare timings with the current LAPACK bidiagonalization algorithm.
Gary W. Howell, James Demmel, Charles T. Fulton, Sven Hammarling, Karen Marmol
ACM Trans. Math. Softw.4
1997 Key Concepts for Parallel Out-of-Core LU Factorization
Jack J. Dongarra, Sven Hammarling, David W. Walker
Parallel Comput.2
1997 Practical Experience in the Numerical Dangers of Heterogeneous Computing
abstract
Special challenges exist in writing reliable numerical library software for heterogeneous computing environments. Although a lot of software for distributed-memory parallel computers has been written, porting this software to a network of workstations requires careful consideration. The symptoms of heterogeneous computing failures can range from erroneous results without warning to deadlock. Some of the problems are straightforward to solve, but for others the solutions are not so obvious, or incur an unacceptable overhead. Making software robust on heterogeneous systems often requires additional communication. We describe and illustrate the problems encountered during the development of ScaLAPACK and the NAG Numerical PVM Library. Where possible, we suggest ways to avoid potential pitfalls, or if that is not possible, we recommend that the software not be used on heterogeneous networks.
L. Susan Blackford, Andrew J. Cleary, Antoine Petitet, R. Clint Whaley, James Demmel, Inderjit S. Dhillon, H. Ren, Ken Stanley, Jack J. Dongarra, Sven Hammarling
ACM Trans. Math. Softw.10
1996 ScaLAPACK: A Portable Linear Algebra Library for Distributed Memory Computers - Design Issues and Performance
abstract
This paper outlines the content and performance of ScaLAPACK, a collection of mathematical software for linear algebra computations on distributed memory computers. The importance of developing standards for computational and message passing interfaces is discussed. We present the different components and building blocks of ScaLAPACK, and indicate the difficulties inherent in producing correct codes for networks of heterogeneous processors. Finally, this paper briefly describes future directions for the ScaLAPACK library and concludes by suggesting alternative approaches to mathematical libraries, explaining how ScaLAPACK could be integrated into efficient and user-friendly distributed systems.
L. Susan Blackford, Andrew J. Cleary, James Demmel, Inderjit S. Dhillon, Jack J. Dongarra, Sven Hammarling, Greg Henry, Antoine Petitet, Ken Stanley, David W. Walker, R. Clint Whaley
SC7
1990 LAPACK: a portable linear algebra library for high-performance computers
abstract
The goal of the LAPACK project is to design and implement a portable linear algebra library for efficient use on a variety of high-performance computers. The library is based on the widely used LINPACK and EISPACK packages for solving linear equations, eigenvalue problems, and linear least-squares problems, but extends their functionality in a number of ways. The major methodology for making the algorithms run faster is to restructure them to perform block matrix operations (e.g., matrix-matrix multiplication) in their inner loops. These block operations may be optimized to exploit the memory hierarchy of a specific architecture. The LAPACK project is also working on new algorithms that yield higher relative accuracy for a variety of linear algebra problems.>
Edward C. Anderson, Zhaojun Bai, Jack J. Dongarra, Anne Greenbaum, A. McKenney, Jeremy Du Croz, Sven Hammarling, James Demmel, Christian H. Bischof, Danny C. Sorensen
SC7
1990 A set of level 3 basic linear algebra subprograms
abstract
This paper describes an extension to the set of Basic Linear Algebra Subprograms. The extensions are targeted at matrix-vector operations that should provide for efficient and portable implementations of algorithms for high-performance computers
Jack J. Dongarra, Jeremy Du Croz, Sven Hammarling, Iain S. Duff
ACM Trans. Math. Softw.3
1990 Algorithm 679; a set of level 3 basic linear algebra subprograms: model implementation and test programs
abstract
This paper describes a model implementation and test software for the Level 3 Basic Linear Algebra Subprograms (Level3 BLAS). The Level3 BLAS are targeted at matrix-matrix operations with the aim of providing more efficient, but portable, implementations of algorithms on high-performance computers. The model implementation provides a portable set of Fortran 77 Level 3 BLAS for machines where specialized implementations do not exist or are not required. The test software aims to verify that specialized implementations meet the specification of the Level 3 BLAS and that implementations are correctly installed.
Jack J. Dongarra, Jeremy Du Croz, Sven Hammarling, Iain S. Duff
ACM Trans. Math. Softw.3
1989 Comments, with reply, on 'Solving the generalized eigenvalue problem with singular forms by M.D. Zoltowski
abstract
The above-titled letter (ibid., vol.75, no.11, p.1546-8, Nov. 1987) proposed two numerical algorithms for the solution of generalized eigenvalue problem Ax= lambda , where it is assumed that the matrix B is Hermitian. It was claimed by the author that standard algorithms are not available if the matrix B is singular. The commenters draw attention to the literature in this field based on the QZ algorithm. The author thanks the commenters for supplying the information but points out to prospective readers certain caveats with regard to use of the QZ algorithm.>
K. Vince Fernando, Sven Hammarling
Proc. IEEE2
1988 An extended set of FORTRAN basic linear algebra subprograms
abstract
This paper describes an extension to the set of Basic Linear Algebra Subprograms. The extensions are targeted at matrix-vector operations that should provide for efficient and portable implementations of algorithms for high-performance computers.
Jack J. Dongarra, Jeremy Du Croz, Sven Hammarling, Richard J. Hanson
ACM Trans. Math. Softw.3
1988 Algorithm 656: an extended set of basic linear algebra subprograms: model implementation and test programs
abstract
This paper describes a model implementation and test software for the Level 2 Basic Linear Algebra Subprograms (Level 2 BLAS). Level 2 BLAS are targeted at matrix-vector operations with the aim of providing more efficient, but portable, implementations of algorithms on high-performance computers. The model implementation provides a portable set of FORTRAN 77 Level 2 BLAS for machines where specialized implementations do not exist or are not required. The test software aims to verify that specialized implementations meet the specification of Level 2 BLAS and that implementations are correctly installed.
Jack J. Dongarra, Jeremy Du Croz, Sven Hammarling, Richard J. Hanson
ACM Trans. Math. Softw.3
1988 Corrigenda: "An Extended Set of FORTRAN Basic Linear Algebra Subprograms"
abstract
No abstract available.
Jack J. Dongarra, Jeremy Du Croz, Sven Hammarling, Richard J. Hanson
ACM Trans. Math. Softw.3