VLDB 2026 Research / reviewers in the wild / expert
Hartwig Anzt
dblp:69/8652
· DBLP profile ↗
37ranked-venue papers
13as first author
16since 2021 · last 2025
0000-0003-2177-952XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 10 first-author · 13 since 2021Theory of computation · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | pyGinkgo: A Sparse Linear Algebra Operator Framework for PythonabstractSparse linear algebra is a cornerstone of many scientific computing and machine learning applications. Python has become a popular choice for these applications due to its simplicity and ease of use. Yet high-performance sparse kernels in Python remain limited in functionality, especially on modern CPU and GPU architectures. We present pyGinkgo, a lightweight and Pythonic interface to the Ginkgo library, offering high-performance sparse linear algebra support with platform portability across CUDA, HIP, and OpenMP backends. pyGinkgo bridges the gap between high-performance C++ backends and Python usability by exposing Ginkgo’s capabilities via Pybind11 and a NumPy and PyTorch compatible interface. We benchmark pyGinkgo’s performance against state-of-the-art Python libraries including SciPy, CuPy, PyTorch and TensorFlow. Results across hardware from different vendors demonstrate that pyGinkgo consistently outperforms existing Python tools in both Sparse Matrix Vector (SpMV) product and iterative solver performance, while maintaining performance parity with native Ginkgo C++ code. Our work positions pyGinkgo as a compelling backend for sparse machine learning models and scientific workflows. Keshvi Tuteja, Gregor Olenik, Roman Mishchuk, Yu-Hsiang Tsai, Markus Götz, Achim Streit, Hartwig Anzt, Charlotte Debus |
ICPP | 7 |
| 2025 | Multifacets of lossy compression for scientific data in the Joint-Laboratory of Extreme Scale Computing
Franck Cappello, Mario C. Acosta, Emmanuel Agullo, Hartwig Anzt, Jon Calhoun 0001, Sheng Di, Luc Giraud, Thomas Grützmacher, Sian Jin, Kentaro Sano, Kento Sato, Amarjit Singh, Dingwen Tao, Jiannan Tian, Tomohiro Ueno, Robert Underwood, Frédéric Vivien, Xavier Yepes, Kazutomo Yoshii, Boyuan Zhang 0002 |
Future Gener. Comput. Syst. | 4 |
| 2025 | JuMonC: A RESTful tool for enabling monitoring and control of simulations at scaleabstractAs systems and simulations grow in size and complexity, it is challenging to maintain efficient use of resources and avoid failures. In this scenario, monitoring becomes even more important and mandatory. This paper describes and discusses the benefits of the advanced monitoring and control tool JuMonC, which runs under user control alongside HPC simulations and provides valuable metrics via REST-API. In addition, plugin extensibility allows JuMonC to go a step further and provide computational steering of the simulation itself. To demonstrate the benefits and usability of JuMonC for large-scale simulations, two use cases are described employing nekRS and ICON on JURECA-DC, a supercomputer located at the Jülich Supercomputing Centre (JSC). Furthermore, a large-scale use case with nekRS on JSC’s flagship system JUWELS Booster is described. Finally, the interplay between JuMonC and LLview (a standard monitoring tool for HPC systems) is presented using a simple and secure JuMonC-LLview plugin, which collects performance metrics and enables their analysis in LLview. Overall, the portability and usefulness of JuMonC, together with its low performance impact, make it an important application for both current and future generations of exascale HPC systems. • Description of JuMonC, a novel tool for monitoring and control for simulations at scale. • Applications with nekRS and ICON. • Performance measurements up to 3200 GPUs. • Elaboration of the integration into LLview. • Discussion of the benefit for exascale workflows. Christian Witzler, Filipe Souza Mendes Guimarães, Daniel Mira, Hartwig Anzt, Jens Henrik Göbbert, Wolfgang Frings, Mathis Bode |
Future Gener. Comput. Syst. | 4 |
| 2024 | Batched sparse and mixed-precision linear algebra interface for efficient use of GPU hardware accelerators in scientific applications
Piotr Luszczek, Ahmad Abdelfattah, Hartwig Anzt, Stanimire Tomov |
Future Gener. Comput. Syst. | 3 |
| 2023 | PAQR: Pivoting Avoiding QR factorizationabstractThe solution of linear least-squares problems is at the heart of many scientific and engineering applications. While any method able to minimize the backward error of such problems is considered numerically stable, the theory states that the forward error depends on the condition number of the matrix in the system of equations. On the one hand, the QR factorization is an efficient method to solve such problems, but the solutions it produces may have large forward errors when the matrix is rank deficient. On the other hand, rank-revealing QR (RRQR) is able to produce smaller forward errors on rank deficient matrices, but its cost is prohibitive compared to QR due to memory-inefficient operations. The aim of this paper is to propose PAQR for the solution of rank-deficient linear least-squares problems as an alternative solution method. It has the same (or smaller) cost as QR and is as accurate as QR with column pivoting in many practical cases. In addition to presenting the algorithm and its implementations on different hardware architectures, we compare its accuracy and performance results on a variety of application-derived problems. Wissam M. Sid-Lakhdar, Sébastien Cayrols, Daniel Bielich, Ahmad Abdelfattah, Piotr Luszczek, Mark Gates, Stanimire Tomov, Hans Johansen, David B. Williams-Young, Timothy A. Davis 0001, Jack J. Dongarra, Hartwig Anzt |
IPDPS | 12 |
| 2023 | Sparse matrix-vector and matrix-multivector products for the truncated SVD on graphics processorsabstractSummary Many practical algorithms for numerical rank computations implement an iterative procedure that involves repeated multiplications of a vector, or a collection of vectors, with both a sparse matrix and its transpose. Unfortunately, the realization of these sparse products on current high performance libraries often deliver much lower arithmetic throughput when the matrix involved in the product is transposed. In this work, we propose a hybrid sparse matrix layout, named CSRC, that combines the flexibility of some well‐known sparse formats to offer a number of appealing properties: (1) CSRC can be obtained at low cost from the popular CSR (compressed sparse row) format; (2) CSRC has similar storage requirements as CSR; and especially, (3) the implementation of the sparse product kernels delivers high performance for both the direct product and its transposed variant on modern graphics accelerators thanks to a significant reduction of atomic operations compared to a conventional implementation based on CSR. This solution thus renders considerably higher performance when integrated into an iterative algorithm for the truncated singular value decomposition (SVD), such as the randomized SVD or, as demonstrated in the experimental results, the block Golub–Kahan–Lanczos algorithm. José Ignacio Aliaga, Hartwig Anzt, Enrique S. Quintana-Ortí, Andrés Tomás |
Concurr. Comput. Pract. Exp. | 2 |
| 2023 | Providing performance portable numerics for Intel GPUsabstractSummary With discrete Intel GPUs entering the high‐performance computing landscape, there is an urgent need for production‐ready software stacks for these platforms. In this article, we report how we enable the Ginkgo math library to execute on Intel GPUs by developing a kernel backed based on the DPC++ programming environment. We discuss conceptual differences between the CUDA and DPC++ programming models and describe workflows for simplified code conversion. We evaluate the performance of basic and advanced sparse linear algebra routines available in Ginkgo's DPC++ backend in the hardware‐specific performance bounds and compare against routines providing the same functionality that ship with Intel's oneMKL vendor library. Yu-Hsiang Tsai, Terry Cojean, Hartwig Anzt |
Concurr. Comput. Pract. Exp. | 3 |
| 2023 | Three-precision algebraic multigrid on GPUs
Yu-Hsiang Tsai, Natalie N. Beams, Hartwig Anzt |
Future Gener. Comput. Syst. | 3 |
| 2023 | Integrating batched sparse iterative solvers for the collision operator in fusion plasma simulations on GPUs
Aditya Kashi, Pratik Nayak, Dhruva Kulkarni, Aaron Scheinberg, Paul Lin, Hartwig Anzt |
J. Parallel Distributed Comput. | 6 |
| 2023 | Using Ginkgo's memory accessor for improving the accuracy of memory-bound low precision BLASabstractAbstract The roofline model not only provides a powerful tool to relate an application's performance with the specific constraints imposed by the target hardware but also offers a graphic representation of the balance between memory access cost and compute throughput. In this work, we present a strategy to break up the tight coupling between the precision format used for arithmetic operations and the storage format employed for memory operations. (At a high level, this idea is equivalent to compressing/decompressing the data in registers before/after invoking store/load memory operations.) In practice, we demonstrate that a “memory accessor” that hides the data compression behind the memory access, can virtually push the bandwidth‐induced roofline, yielding higher performance for memory‐bound applications using high precision arithmetic that can handle the numerical effects associated with lossy compression. We also demonstrate that memory‐bound applications operating on low precision data can increase the accuracy by relying on the memory accessor to perform all arithmetic operations in high precision. In particular, we demonstrate that memory‐bound BLAS operations (including the sparse matrix‐vector product) can be re‐engineered with the memory accessor and that the resulting accessor‐enabled BLAS routines achieve lower rounding errors while delivering the same performance as the fast low precision BLAS. Thomas Grützmacher, Hartwig Anzt, Enrique S. Quintana-Ortí |
Softw. Pract. Exp. | 2 |
| 2022 | Batched sparse iterative solvers on GPU for the collision operator for fusion plasma simulationsabstractBatched linear solvers, which solve many small related but independent problems, are important in several applications. This is increasingly the case for highly parallel processors such as graphics processing units (GPUs), which need a substantial amount of work to keep them operating efficiently and solving smaller problems one-by-one is not an option. Because of the small size of each problem, the task of coming up with a parallel partitioning scheme and mapping the problem to hardware is not trivial. In recent history, significant attention has been given to batched dense linear algebra. However, there is also an interest in utilizing sparse iterative solvers in a batched form, and this presents further challenges. An example use case is found in a gyrokinetic Particle-In-Cell (PIC) code used for modeling magnetically confined fusion plasma devices. The collision operator has been identified as a bottleneck, and a proxy app has been created for facilitating optimizations and porting to GPUs. The current collision kernel linear solver does not run on the GPU-a major bottleneck. As these matrices are well-conditioned, batched iterative sparse solvers are an attractive option. A batched sparse iterative solver capability has recently been developed in the Ginkgo library. In this paper, we describe how the software architecture can be used to develop an efficient solution for the XGC collision proxy app. Comparisons for the solve times on NVIDIA V100 and A100 GPUs and AMD MI100 GPUs with one dual-socket Intel Xeon Skylake CPU node with 40 OpenMP threads are presented for matrices representative of those required in the collision kernel of XGC. The results suggest that GINKGO's batched sparse iterative solvers are well suited for efficient utilization of the GPU for this problem, and the performance portability of Ginkgo in conjunction with Kokkos (used within XGC as the heterogeneous programming model) allows seamless execution for exascale oriented heterogeneous architectures at the various leadership supercomputing facilities. Aditya Kashi, Pratik Nayak, Dhruva Kulkarni, Aaron Scheinberg, Paul Lin, Hartwig Anzt |
IPDPS | 6 |
| 2022 | Compression and load balancing for efficient sparse matrix-vector product on multicore processors and graphics processing unitsabstractSummary We contribute to the optimization of the sparse matrix‐vector product by introducing a variant of the coordinate sparse matrix format that balances the workload distribution and compresses both the indexing arrays and the numerical information. Our approach is multi‐platform, in the sense that the realizations for (general‐purpose) multicore processors as well as graphics accelerators (GPUs) are built upon common principles, but differ in the implementation details, which are adapted to avoid thread divergence in the GPU case or maximize compression element‐wise (i.e., for each matrix entry) for multicore architectures. Our evaluation on the two last generations of NVIDIA GPUs as well as Intel and AMD processors demonstrate the benefits of the new kernels when compared with the optimized implementations of the sparse matrix‐vector product in NVIDIA's cuSPARSE and Intel's MKL, respectively. José Ignacio Aliaga, Hartwig Anzt, Thomas Grützmacher, Enrique S. Quintana-Ortí, Andrés Tomás |
Concurr. Comput. Pract. Exp. | 2 |
| 2022 | Ginkgo - A math library designed for platform portability
Terry Cojean, Yu-Hsiang Tsai, Hartwig Anzt |
Parallel Comput. | 3 |
| 2022 | Ginkgo: A Modern Linear Operator Algebra Framework for High Performance ComputingabstractIn this article, we present Ginkgo , a modern C++ math library for scientific high performance computing. While classical linear algebra libraries act on matrix and vector objects, Ginkgo ’s design principle abstracts all functionality as “linear operators,” motivating the notation of a “linear operator algebra library.” Ginkgo ’s current focus is oriented toward providing sparse linear algebra functionality for high performance graphics processing unit (GPU) architectures, but given the library design, this focus can be easily extended to accommodate other algorithms and hardware architectures. We introduce this sophisticated software architecture that separates core algorithms from architecture-specific backends and provide details on extensibility and sustainability measures. We also demonstrate Ginkgo ’s usability by providing examples on how to use its functionality inside the MFEM and deal.ii finite element ecosystems. Finally, we offer a practical demonstration of Ginkgo ’s high performance on state-of-the-art GPU architectures. Hartwig Anzt, Terry Cojean, Goran Flegar, Fritz Göbel, Thomas Grützmacher, Pratik Nayak, Tobias Ribizel, Yu-Hsiang Tsai, Enrique S. Quintana-Ortí |
ACM Trans. Math. Softw. | 1 |
| 2021 | Mixed Precision Incomplete and Factorized Sparse Approximate Inverse Preconditioning on GPUs
Fritz Göbel, Thomas Grützmacher, Tobias Ribizel, Hartwig Anzt |
Euro-Par | 4 |
| 2021 | Adaptive Precision Block-Jacobi for High Performance Preconditioning in the Ginkgo Linear Algebra SoftwareabstractThe use of mixed precision in numerical algorithms is a promising strategy for accelerating scientific applications. In particular, the adoption of specialized hardware and data formats for low-precision arithmetic in high-end GPUs (graphics processing units) has motivated numerous efforts aiming at carefully reducing the working precision in order to speed up the computations. For algorithms whose performance is bound by the memory bandwidth, the idea of compressing its data before (and after) memory accesses has received considerable attention. One idea is to store an approximate operator–like a preconditioner–in lower than working precision hopefully without impacting the algorithm output. We realize the first high-performance implementation of an adaptive precision block-Jacobi preconditioner which selects the precision format used to store the preconditioner data on-the-fly, taking into account the numerical properties of the individual preconditioner blocks. We implement the adaptive block-Jacobi preconditioner as production-ready functionality in the Ginkgo linear algebra library, considering not only the precision formats that are part of the IEEE standard, but also customized formats which optimize the length of the exponent and significand to the characteristics of the preconditioner blocks. Experiments run on a state-of-the-art GPU accelerator show that our implementation offers attractive runtime savings. Goran Flegar, Hartwig Anzt, Terry Cojean, Enrique S. Quintana-Ortí |
ACM Trans. Math. Softw. | 2 |
| 2020 | Multiprecision Block-Jacobi for Iterative Triangular Solves
Fritz Göbel, Hartwig Anzt, Terry Cojean, Goran Flegar, Enrique S. Quintana-Ortí |
Euro-Par | 2 |
| 2020 | A customized precision format based on mantissa segmentation for accelerating sparse linear algebraabstractSummary In this work, we pursue the idea of radically decoupling the floating point format used for arithmetic operations from the format used to store the data in memory. We complement this idea with a customized precision memory format derived by splitting the mantissa (significand) of standard IEEE formats into segments, such that values can be accessed faster if lower accuracy is acceptable. Combined with precision‐aware algorithms that dynamically adapt the data access accuracy to the numerical requirements, the customized precision memory format can render attractive runtime savings without impacting the memory footprint of the data or the accuracy of the final result. In an experimental analysis using the adaptive precision Jacobi method on diagonalizable test problems, we assess the benefits of the mantissa‐segmenting customized precision format on recent multi‐ and manycore architectures. Thomas Grützmacher, Terry Cojean, Goran Flegar, Fritz Göbel, Hartwig Anzt |
Concurr. Comput. Pract. Exp. | 5 |
| 2020 | Parallel selection on GPUs
Tobias Ribizel, Hartwig Anzt |
Parallel Comput. | 2 |
| 2019 | ParILUT - A Parallel Threshold ILU for GPUsabstractIn this paper, we present the first algorithm for computing threshold ILU factorizations on GPU architectures. The proposed ParILUT-GPU algorithm is based on interleaving parallel fixed-point iterations that approximate the incomplete factors for an existing nonzero pattern with a strategy that dynamically adapts the nonzero pattern to the problem characteristics. This requires the efficient selection of thresholds that separate the values to be dropped from the incomplete factors, and we design a novel selection algorithm tailored towards GPUs. All components of the ParILUT-GPU algorithm make heavy use of the features available in the latest NVIDIA GPU generations, and outperform existing multithreaded CPU implementations. Hartwig Anzt, Tobias Ribizel, Goran Flegar, Edmond Chow, Jack J. Dongarra |
IPDPS | 1 |
| 2019 | Adaptive precision in block-Jacobi preconditioning for iterative sparse linear system solversabstractSummary We propose an adaptive scheme to reduce communication overhead caused by data movement by selectively storing the diagonal blocks of a block‐Jacobi preconditioner in different precision formats (half, single, or double). This specialized preconditioner can then be combined with any Krylov subspace method for the solution of sparse linear systems to perform all arithmetic in double precision. We assess the effects of the adaptive precision preconditioner on the iteration count and data transfer cost of a preconditioned conjugate gradient solver. A preconditioned conjugate gradient method is, in general, a memory bandwidth‐bound algorithm, and therefore its execution time and energy consumption are largely dominated by the costs of accessing the problem's data in memory. Given this observation, we propose a model that quantifies the time and energy savings of our approach based on the assumption that these two costs depend linearly on the bit length of a floating point number. Furthermore, we use a number of test problems from the SuiteSparse matrix collection to estimate the potential benefits of the adaptive block‐Jacobi preconditioning scheme. Hartwig Anzt, Jack J. Dongarra, Goran Flegar, Nicholas J. Higham, Enrique S. Quintana-Ortí |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | Variable-size batched Gauss-Jordan elimination for block-Jacobi preconditioning on graphics processors
Hartwig Anzt, Jack J. Dongarra, Goran Flegar, Enrique S. Quintana-Ortí |
Parallel Comput. | 1 |
| 2018 | A Jaccard Weights Kernel Leveraging Independent Thread Scheduling on GPUsabstractJaccard weights are a popular metric for identifying communities in social network analytics. In this paper we present a kernel for efficiently computing the Jaccard weight matrix on G PU s. The kernel design is guided by fine-grained parallelism and the independent thread scheduling supported by NVIDIA's Volta architecture. This technology makes it possible to interleave the execution of divergent branches for enhanced data reuse and a higher instruction per cycle rate for memory-bound algorithms. In a performance evaluation using a set of publicly available social networks, we report the kernel execution time and analyze the built-in hardware counters on different GPU architectures. The findings have implications beyond the specific algorithm and suggest a reformulation of other data-sparse algorithms. Hartwig Anzt, Jack J. Dongarra |
SBAC-PAD | 1 |
| 2018 | Variable-Size Batched Condition Number Calculation on GPUsabstractWe present a kernel that is designed to quickly compute the condition number of a large collection of tiny matrices on a graphics processing unit (GPU). The matrices can differ in size and the process integrates the use of pivoting to ensure a numerically-stable matrix inversion. The performance assessment reveals that, in double precision arithmetic, the new GPU kernel achieves up to 550 GFLOPs (billions of floating-point operations per second) and 800 GFLOPs on NVIDIA's P100 and V100 GPUs, respectively. The results also demonstrate a considerable speed-up with respect to a workflow that computes the condition number via launching a set of four batched kernels. In addition, we present a variable-size batched kernel for the computation of the matrix infinity norm. We show that this memory-bound kernel achieves up to 90% of the sustainable peak bandwidth. Hartwig Anzt, Jack J. Dongarra, Goran Flegar, Thomas Grützmacher |
SBAC-PAD | 1 |
| 2018 | Using Jacobi iterations and blocking for solving sparse triangular systems in incomplete factorization preconditioning
Edmond Chow, Hartwig Anzt, Jennifer A. Scott, Jack J. Dongarra |
J. Parallel Distributed Comput. | 2 |
| 2018 | Incomplete Sparse Approximate Inverses for Parallel Preconditioning
Hartwig Anzt, Thomas Huckle, Jürgen Bräckle, Jack J. Dongarra |
Parallel Comput. | 1 |
| 2017 | Variable-Size Batched LU for Small Matrices and Its Integration into Block-Jacobi PreconditioningabstractWe present a set of new batched CUDA kernels for the LU factorization of a large collection of independent problems of different size, and the subsequent triangular solves. All kernels heavily exploit the registers of the graphics processing unit (GPU) in order to deliver high performance for small problems. The development of these kernels is motivated by the need for tackling this embarrasingly-parallel scenario in the context of block-Jacobi preconditioning that is relevant for the iterative solution of sparse linear systems. Hartwig Anzt, Jack J. Dongarra, Goran Flegar, Enrique S. Quintana-Ortí |
ICPP | 1 |
| 2017 | Preconditioned Krylov solvers on GPUs
Hartwig Anzt, Mark Gates, Jack J. Dongarra, Moritz Kreutzer, Gerhard Wellein, Martin Koehler |
Parallel Comput. | 1 |
| 2016 | Implementation and Tuning of Batched Cholesky Factorization and Solve for NVIDIA GPUsabstractMany problems in engineering and scientific computing require the solution of a large number of small systems of linear equations. Due to their high processing power, Graphics Processing Units became an attractive target for this class of problems, and routines based on the LU and the QR factorization have been provided by NVIDIA in the cuBLAS library. This work addresses the situation where the systems of equations are symmetric positive definite. The paper describes the implementation and tuning of the kernels for the Cholesky factorization and the forward and backward substitution. Targeted workloads involve the solution of thousands of linear systems of the same size, where the focus is on matrix dimensions from 5 by 5 to 100 by 100. Due to the lack of a cuBLAS Cholesky factorization, execution rates of cuBLAS LU and cuBLAS QR are used for comparison against the proposed Cholesky factorization in this work. Execution rates of forward and backward substitution routines are compared to equivalent cuBLAS routines. Comparisons against optimized multicore implementations are also presented. Superior performance is reached in all cases. Jakub Kurzak, Hartwig Anzt, Mark Gates, Jack J. Dongarra |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2015 | Accelerating collaborative filtering using concepts from high performance computingabstractIn this paper we accelerate the Alternating Least Squares (ALS) algorithm used for generating product recommendations on the basis of implicit feedback datasets. We approach the algorithm with concepts proven to be successful in High Performance Computing. This includes the formulation of the algorithm as a mix of cache-optimized algorithm-specific kernels and standard BLAS routines, acceleration via graphics processing units (GPUs), use of parallel batched kernels, and autotuning to identify performance winners. For benchmark datasets, the multi-threaded CPU implementation we propose achieves more than a 10 times speedup over the implementations available in the GraphLab and Spark MLlib software packages. For the GPU implementation, the parameters of an algorithm-specific kernel were optimized using a comprehensive autotuning sweep. This results in an additional 2 times speedup over our CPU implementation. Mark Gates, Hartwig Anzt, Jakub Kurzak, Jack J. Dongarra |
IEEE BigData | 2 |
| 2015 | Iterative Sparse Triangular Solves for Preconditioning
Hartwig Anzt, Edmond Chow, Jack J. Dongarra |
Euro-Par | 1 |
| 2015 | Unveiling the performance-energy trade-off in iterative linear system solvers for multithreaded processorsabstractSummary In this paper, we analyze the interactions occurring in the triangle performance‐power‐energy for the execution of a pivotal numerical algorithm, the iterative conjugate gradient (CG) method, on a diverse collection of parallel multithreaded architectures. This analysis is especially timely in a decade where the power wall has arisen as a major obstacle to build faster processors. Moreover, the CG method has recently been proposed as a complement to the LINPACK benchmark, as this iterative method is argued to be more archetypical of the performance of today's scientific and engineering applications. To gain insights about the benefits of hands‐on optimizations we include runtime and energy efficiency results for both out‐of‐the‐box usage relying exclusively on compiler optimizations, and implementations manually optimized for target architectures, that range from general‐purpose and digital signal multicore processors to manycore graphics processing units, all representative of current multithreaded systems. Copyright © 2014 John Wiley & Sons, Ltd. José Ignacio Aliaga, Hartwig Anzt, María Isabel Castillo, Juan Carlos Fernández 0002, German Leon, Enrique S. Quintana-Ortí |
Concurr. Comput. Pract. Exp. | 2 |
| 2015 | Experiences in autotuning matrix multiplication for energy minimization on GPUsabstractSummary In this paper, we report extensive results and analysis of autotuning the computationally intensive graphics processing units kernel for dense matrix–matrix multiplication in double precision. In contrast to traditional autotuning and/or optimization for runtime performance only, we also take the energy efficiency into account. For kernels achieving equal performance, we show significant differences in their energy balance. We also identify the memory throughput as the most influential metric that trades off performance and energy efficiency. As a result, the performance optimal case ends up not being the most efficient kernel in overall resource use. Copyright © 2015 John Wiley & Sons, Ltd. Hartwig Anzt, Blake Haugen, Jakub Kurzak, Piotr Luszczek, Jack J. Dongarra |
Concurr. Comput. Pract. Exp. | 1 |
| 2014 | Improving the Performance of CA-GMRES on Multicores with Multiple GPUsabstractThe Generalized Minimum Residual (GMRES) method is one of the most widely-used iterative methods for solving nonsymmetric linear systems of equations. In recent years, techniques to avoid communication in GMRES have gained attention because in comparison to floating-point operations, communication is becoming increasingly expensive on modern computers. Since graphics processing units (GPUs) are now becoming crucial component in computing, we investigate the effectiveness of these techniques on multicore CPUs with multiple GPUs. While we present the detailed performance studies of a matrix powers kernel on multiple GPUs, we particularly focus on orthogonalization strategies that have a great impact on both the numerical stability and performance of GMRES, especially as the matrix becomes sparser or ill-conditioned. We present the experimental results on two eight-core Intel Sandy Bridge CPUs with three NDIVIA Fermi GPUs and demonstrate that significant speedups can be obtained by avoiding communication, either on a GPU or between the GPUs. As part of our study, we investigate several optimization techniques for the GPU kernels that can also be used in other iterative solvers besides GMRES. Hence, our studies not only emphasize the importance of avoiding communication on GPUs, but they also provide insight about the effects of these optimization techniques on the performance of the sparse solvers, and may have greater impact beyond GMRES. Ichitaro Yamazaki, Hartwig Anzt, Stanimire Tomov, Mark Hoemmen, Jack J. Dongarra |
IPDPS | 2 |
| 2013 | Reformulated Conjugate Gradient for the Energy-Aware Solution of Linear Systems on GPUsabstractIn this paper we introduce a redesign of the conjugate gradient method for the iterative solution of sparse linear systems on heterogeneous systems accelerated by graphics processing units (GPUs). Reshaping the GPU kernels induced by the classical formulation of the CG method into algorithm-specific routines results in a slight increase of performance and, more importantly, enables the efficient exploitation of power-saving techniques implicit in the hardware, like the processor C-states, that produce remarkable energy savings. Numerical experiments using data matrices from a popular sparse matrix collection show that the time overhead naturally associated with the application of these energy-aware techniques is no longer crucial to the overall runtime performance. José Ignacio Aliaga, Enrique S. Quintana-Ortí, Hartwig Anzt |
ICPP | 4 |
| 2013 | A block-asynchronous relaxation method for graphics processing units
Hartwig Anzt, Stanimire Tomov, Jack J. Dongarra, Vincent Heuveline |
J. Parallel Distributed Comput. | 1 |
| 2012 | GPU-Accelerated Asynchronous Error Correction for Mixed Precision Iterative Refinement
Hartwig Anzt, Piotr Luszczek, Jack J. Dongarra, Vincent Heuveline |
Euro-Par | 1 |