Ricard Borrell

dblp:68/9750 · also Ricard Borrell Pol, Rick Borrell · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
3since 2021 · last 2024
0000-0003-3747-7104ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2024 Extending Sparse Patterns to Improve Inverse Preconditioning on GPU Architectures
abstract
Graphic Processing Units (GPUs) have become a key component of high-end computing infrastructures due to their massively parallel architecture, which delivers large floating-point operations per cycle rates. Many scientific workloads benefit from GPUs and, in particular, numerical methods solving linear systems of equations Ax = b typically run on GPUs. Among them, the Conjugate Gradient (CG) method, which targets linear systems with Symmetric and Positive Definite (SPD) matrices, runs on GPUs using its preconditioned form. However, state-of-the-art preconditioning techniques like the Factorized Sparse Approximate Inverse (FSAI) preconditioner ignore the benefits of data coalescence and locality on GPU architectures and leave substantial performance on the table. These approaches are exclusively based on numerical criteria.
Sergi Laut, Ricard Borrell, Marc Casas
HPDC2
2022 Communication-aware Sparse Patterns for the Factorized Approximate Inverse Preconditioner
abstract
The Conjugate Gradient (CG) method is an iterative solver targeting linear systems of equations Ax=b where A is a symmetric and positive definite matrix. CG convergence properties improve when preconditioning is applied to reduce the condition number of matrix A. While many different options can be found in the literature, the Factorized Sparse Approximate Inverse (FSAI) preconditioner constitutes a highly parallel option based on approximating A-1. This paper proposes the Communication-aware Factorized Sparse Approximate Inverse preconditioner (FSAIE-Comm), a method to generate extensions of the FSAI sparse pattern that are not only cache friendly, but also avoid increasing communication costs in distributed memory systems. We also propose a filtering strategy to reduce inter-process imbalance. We evaluate FSAIE-Comm on a heterogeneous set of 39 matrices achieving an average solution time decrease of 17.98%, 26.44% and 16.74% on three different architectures, respectively, Intel Skylake, Fujitsu A64FX and AMD Zen 2 with respect to FSAI. In addition, we consider a set of 8 large matrices running on up to 32,768 CPU cores, and we achieve an average solution time decrease of 12.59%.
Sergi Laut, Marc Casas, Ricard Borrell
HPDC3
2021 Cache-aware Sparse Patterns for the Factorized Sparse Approximate Inverse Preconditioner
abstract
Conjugate Gradient is a widely used iterative method to solve linear systems Ax=b with matrix A being symmetric and positive definite. Part of its effectiveness relies on finding a suitable preconditioner that accelerates its convergence. Factorized Sparse Approximate Inverse (FSAI) preconditioners are a prominent and easily parallelizable option. An essential element of a FSAI preconditioner is the definition of its sparse pattern, which constraints the approximation of the inverse A-1. This definition is generally based on numerical criteria. In this paper we introduce complementary architecture-aware criteria to increase the numerical effectiveness of the preconditioner without incurring in significant performance costs. In particular, we define cache-aware pattern extensions that do not trigger additional cache misses when accessing vector x in the y=Ax Sparse Matrix-Vector (SpMV) kernel. As a result, we obtain very significant reductions in terms of average solution time ranging between 12.94% and 22.85% on three different architectures - Intel Skylake, POWER9 and A64FX - over a set of 72 test matrices.
Sergi Laut, Ricard Borrell, Marc Casas
HPDC2
2020 Heterogeneous CPU/GPU co-execution of CFD simulations on the POWER9 architecture: Application to airplane aerodynamics
Ricard Borrell, Damien Dosimont, Marta Garcia-Gasulla, Guillaume Houzeaux, Oriol Lehmkuhl, V. Mehta, Herbert Owen, Mariano Vázquez, Guillermo Oyarzun
Future Gener. Comput. Syst.1
2018 Efficient CFD code implementation for the ARM-based Mont-Blanc architecture
abstract
Since 2011, the European project Mont-Blanc has been focused on enabling ARM-based technology for HPC, developing both hardware platforms and system software. The latest Mont-Blanc prototypes use system-on-chip (SoC) devices that combine a CPU and a GPU sharing a common main memory. Specific developments of parallel computing software and well-suited implementation approaches are of crucial importance for such a heterogeneous architecture in order to efficiently exploit its potential. This paper is devoted to the optimizations carried out in the TermoFluids CFD code to efficiently run it on the Mont-Blanc system. The underlying numerical method is based on an unstructured finite-volume discretization of the Navier–Stokes equations for the numerical simulation of incompressible turbulent flows. It is implemented using a portable and modular operational approach based on a minimal set of linear algebra operations. An architecture-specific heterogeneous multilevel MPI+OpenMP+OpenCL implementation of such kernels is proposed. It includes optimizations of the storage formats, dynamic load balancing between the CPU and GPU devices and hiding of communication overheads by overlapping computations and data transfers. A detailed performance study shows time reductions of up to 2.1× on the kernels’ execution with the new heterogeneous implementation, its scalability on up to 128 Mont-Blanc nodes and the energy savings (around 40%) achieved with the Mont-Blanc system versus the high-end hybrid supercomputer MinoTauro.
Guillermo Oyarzun, Ricard Borrell, Andrey V. Gorobets, Filippo Mantovani, Assensi Oliva
Future Gener. Comput. Syst.2