VLDB 2026 Research / reviewers in the wild / expert
Pablo Ezzatti
dblp:71/8239
· DBLP profile ↗
52ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-2368-8907ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 12 · 3 since 2021Software engineering, systems software and programming languages · 10 · 3 since 2021Databases, data management, data science and information retrieval · 10 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optimizing SpMV Kernel for Banded Matrices on FPGAs Using High-Level Synthesis
Federico Favaro, Ernesto Dufrechu, Juan P. Oliver, Pablo Ezzatti |
Euro-Par (1) | 4 |
| 2026 | A LightGBM Framework for Operational Tide Forecasting and Predictive Imputation in the Río De La Plata
Diego Silva Piedra, Mónica Fossati, Pablo Ezzatti |
ICCSA (2) | 3 |
| 2025 | Load-balanced SpMV kernel for the bmSparse matrix formatabstractThe design of sparse matrix storage formats is essential to achieve high-performance sparse kernels in modern parallel architectures. The bitmap-based bmSparse format was designed with the SPGEMM operation as its main focus, but it shows potential for other operations as well when the sparse matrix has a convenient structure. In this paper, we propose a new SPMV kernel that greatly improves the load balance of previous open-source implementations and utilizes the GPU resources more efficiently. The results show speedups of up to $100 \times$ regarding other SPMV kernels for bmSparse. Finally, we leverage a hybrid implementation between the new and existing approach, selecting the kernel which will probably result in a better performance depending on each matrix characteristics. Gonzalo Berger, Ernesto Dufrechu, Pablo Ezzatti |
PDP | 3 |
| 2025 | A synchronization-free incomplete LU factorization for GPUs with level-set analysisabstractIncomplete factorization methods are powerful algebraic preconditioners widely used to accelerate the convergence of linear solvers. The parallelization of ILU methods has been extensively studied, particularly for GPUs, which are ubiquitous parallel computing devices. In recent years, synchronization-free methods have become the mainstream approach for solving sparse triangular linear systems.Although the sparse triangular solver and ILU factorization are closely related, the application of synchronization-free strategies to ILU factorization has not been explored in the literature to the same extent as the triangular solver. In this work, we present synchronization-free implementations of the ILU-0 preconditioner on GPUs. Specifically, we propose three implementations that vary in how row updates are handled after each coefficient elimination, as well as an additional approach that leverages a prior level-set analysis to optimize the execution schedule. Manuel Freire 0002, Ernesto Dufrechu, Pablo Ezzatti |
PDP | 3 |
| 2024 | A new level-set analysis and sparse storage format for the SPTRSV in GPUsabstractDue to its relevant role in many numerical methods, the solution of sparse triangular linear systems (SpTRSV) in parallel platforms is continuously studied to extract as much performance as possible from the latest hardware architectures. In the case of GPUs, the latest solvers use the synchronization-free paradigm. When the problem involves several system solutions for the same matrix, they often pre-process it through a levelset analysis to improve the equation solution scheduling in the solution phase. In addition, other optimizations address the load balancing issues and irregular memory access of the SpTRSV. In this work, we modify the classical approach to compute the level sets used in the parallel SpTRSV computation, and we show that the new strategy generally reduces the computation time of the solver. Furthermore, we design an internal matrix representation that can significantly accelerate the solution stage at the cost of increasing the memory storage requirements of the algorithm. The experimental evaluation shows that the proposed modifications can improve the performance of a recent levelset and synchronization-free solver by up to 70%, significantly outperforming other state-of-the-art solvers, especially when several linear systems must be solved for each analysis phase. Manuel Freire 0002, Ernesto Dufrechu, Pablo Ezzatti |
SBAC-PAD | 3 |
| 2023 | Evaluation of architecture-aware optimization techniques for Convolutional Neural NetworksabstractThe growing need to perform Neural network inference with low latency is giving place to a broad spectrum of heterogeneous devices with deep learning capabilities. Therefore, obtaining the best performance from each device and choosing the most suitable platform for a given problem has become challenging. This paper evaluates multiple inference platforms using architecture-aware optimizations for convolutional neural networks. Specifically, we use TensorRT and OpenVINO frameworks for hardware optimizations on top of the platform-aware NetAdapt algorithm. The experimental evaluation shows that on MobileNet and AlexNet, using NetAdapt with TensorRT or Open-VINO can improve latency up to 10 x and 5.3 x, respectively. Moreover, a throughput test using different batch sizes showed variable performance improvement on the different devices. Discussing the experimental results can guide the selection of devices and optimizations for different AI solutions. Raúl Marichal, Guillermo Toyos, Ernesto Dufrechu, Pablo Ezzatti |
PDP | 4 |
| 2023 | Advancing on an efficient sparse matrix multiplication kernel for modern GPUsabstractSummary The sparse matrix multiplication (SpGeMM) increased its importance in the last years due to its data science and machine learning applications. Consequently, considerable research has focused on accelerating this kernel in GPUs. Designing massively‐parallel algorithms for the SpGeMM is a challenging task since the computation pattern is highly irregular, and the required memory and operations depend on the interaction between the nonzero layout of the inputs. One strategy to attack this kernel consists of proposing new sparse matrix storage formats that contribute to mitigating this irregularity. In previous work, we commenced a study of the recently proposed bmSparse matrix format, suggesting several modifications to the SpGeMM algorithm. This work integrates the previous extensions and proposes new improvements to unleash bmSparse's full potential before comparing it to more consolidated options. In particular, we enhance one of the most computationally demanding stages with an adaptive technique, apply optimizations to achieve more efficient data accesses, and analyze the effect of using Tensor Cores to accelerate the multiplication stage of the algorithm. The experimental results on a set of real‐world sparse matrices show that the optimized implementation largely outperforms vendor implementations such as NVIDIA cuSparse Intel MKL‐CSR variant, while being competitive with MKL's‐BSR. Gonzalo Berger, Manuel Freire 0002, Renzo Marini, Ernesto Dufrechu, Pablo Ezzatti |
Concurr. Comput. Pract. Exp. | 5 |
| 2023 | A GPU method for the analysis stage of the SPTRSV kernel
Manuel Freire 0002, Juan Ferrand, Franco Seveso, Ernesto Dufrechu, Pablo Ezzatti |
J. Supercomput. | 5 |
| 2021 | Improving the performance of graph database queries using linear algebra operationsabstractThe application of graph databases to different domains is gaining momentum. The Resource Description Framework (RDF) is one of the data models supported by graph databases, and SPARQL is the standard query language for RDF graphs. These databases are also known as RDF triplestores. Many triplestores are implemented over the relational data model, using tables to store graphs and translating SPARQL queries into SQL queries, and this approach can lead to unnecessary overheads. On the other hand, in the context of High- Performance Computing (HPC), implementations over hybrid hardware platforms using Numerical Linear Algebra (NLA) operations have become an effective and efficient computing strategy in the last decade. In particular, Graphics Processing Units (GPUs) have been adopted to perform general-purpose computations due to their high performance, reasonable prices, and an attractive relationship between computing capacity and energy consumption. In the context described above, this paper presents an initial study on the efficient implementation of a set of SPARQL queries in terms of NLA operations. Additionally, we evaluate the performance of implementing these operations on GPUs. Bruno Amaral, Juan Manuel San Martin, Lorena Etcheverry, Pablo Ezzatti |
CLEI | 4 |
| 2021 | Proximity tracing applications for COVID-19: data privacy and securityabstractSince the beginning of 2020, COVID-19 has had a strong impact on the health of the world population. Tracing the contacts of infected people is one of the main strategies for controlling the pandemic. Given the high rates of contagion, which makes difficult an effective manual tracing, multiple initiatives arose for developing digital proximity tracing technologies. In this paper, we discuss in depth the security and personal data protection requirements that these technologies must satisfy, and we present an exhaustive and detailed list of the various applications that have been deployed globally. In particular, we identify potential threats that could undermine the satisfaction of the analyzed requirements, violating hegemonic personal data protection regulations. Gustavo Betarte, Juan Diego Campo, Andrea Delgado 0001, Pablo Ezzatti, Laura González 0001, Alvaro Martín, Rodrigo Martínez, Bárbara Muracciole |
CLEI | 4 |
| 2021 | Assessing the solution of one sparse triangular linear system on multi-many core platformsabstractThe solution of sparse triangular linear systems is an important building block for a large number of numerical methods used in science and engineering. It is then crucial to count with implementations of this operation that can execute efficiently in the most recent hardware platforms. In the case of GPUs, several methods have been proposed in the last years. These methods belong to two main categories. On the one hand, there are the methods that rely on a previous analysis of the sparse matrix to determine a better execution schedule and, on the other hand, there are methods that decide this scheduling dynamically. The experimental results in the literature are not conclussive in favour of any of these strategies. However, the experimental evaluations usually focus on the use case where many systems have to be solved with the same sparse matrix, where the analysis phase needs to be performed only once and its cost is not important in relation with the total runtime. In this work we are interested in determining which is the best strategy, according to the degree of parallelism of the problem, when only one sytem is to be solved. The experimental evaluation performed on NVIDIA P100 accelerators shows that the self-scheduled routines present important advantages when the degree of parallelism of the problem allows it. Raúl Marichal, Ernesto Dufrechu, Pablo Ezzatti |
CLEI | 3 |
| 2021 | Machine learning for optimal selection of sparse triangular system solvers on GPUs
Ernesto Dufrechu, Pablo Ezzatti, Manuel Freire 0002, Enrique S. Quintana-Ortí |
J. Parallel Distributed Comput. | 2 |
| 2021 | Factorized solution of generalized stable Sylvester equations using many-core GPU accelerators
Peter Benner, Ernesto Dufrechu, Pablo Ezzatti, Rodrigo Gallardo, Enrique S. Quintana-Ortí |
J. Supercomput. | 3 |
| 2020 | An asynchronous computation architecture for enhancing the performance of the Weather Research and Forecasting modelabstractSummary The Weather Research and Forecasting (WRF) model does not usually show an adequate scalability for small‐sized domains in large computing platforms. With this in mind, we propose a novel asynchronous software architecture for the WRF model that decouples the calculation of the solar radiation from the rest of the model. The experimental evaluation was conducted on real‐world data for forecasting the photovoltaic energy fed into the electrical grid of Uruguay by twelve solar power plants. The experimental results show that the proposed architecture can make a better use of the hardware resources available on the computing platform, reducing the total runtime of the WRF while maintaining the numerical accuracy. Rodrigo Baya, Martín Pedemonte, Alejandro Gutiérrez Arce, Pablo Ezzatti |
Concurr. Comput. Pract. Exp. | 4 |
| 2020 | Using analysis information in the synchronization-free GPU solution of sparse triangular systemsabstractSummary The solution of sparse triangular linear systems is one of the most important building blocks for a large number of science and engineering problems. For these reasons, it has been studied steadily for several decades, principally in order to take advantage of emerging parallel platforms. In the context of massively parallel platforms such as GPUs, the standard strategy of parallel solution is based on performing a level‐set analysis of the sparse matrix, and the kernel included in the nVidia cuSparse library is the most prominent example of this approach. However, a weak spot of this implementation is the costly analysis phase and the constant synchronizations with the CPU during the solution stage. In previous work, we presented a self‐scheduled and synchronization‐free GPU algorithm that avoided the analysis phase and the synchronizations of the standard approach. Here, we extend this proposal and show how the level‐set information can be leveraged to improve its performance. In particular, we present new GPU solution routines that attack some of the weak spots of the self‐scheduled solver, such as the under‐utilization of the GPU resources in the case of highly sparse matrices. The experimental evaluation reveals a sensible runtime reduction over cuSparse and the state‐of‐the‐art synchronization‐free method. Ernesto Dufrechu, Pablo Ezzatti |
Concurr. Comput. Pract. Exp. | 2 |
| 2019 | Automatic Selection of Sparse Triangular Linear System Solvers on GPUs through Machine Learning TechniquesabstractThe solution of sparse triangular linear systems is often the most time-consuming stage of preconditioned iterative methods to solve general sparse linear systems, where it has to be applied several times for the same sparse matrix. For this reason, its computational performance has a strong impact on a wide range of scientific and engineering applications, which has motivated the study of its efficient execution on massively parallel platforms. In this sense, several methods have been proposed to tackle this operation on graphics processing units (GPUs), which can be classified under either the level-set or the self-scheduling paradigms. The results obtained from the experimental evaluation of the different methods suggest that both paradigms perform well for certain problems but poorly for others. Additionally, the relation between the properties of the linear systems and the performance of the different solvers is not evident a-priori. In this context, techniques that allow to predict inexpensively which is be the best solver for a particular linear system can lead to important runtime reductions. Our approach leverages machine learning techniques to select the best sparse triangular solver for a given linear system, with focus on the case where a small number of triangular systems has to be solved for the same matrix. We study the performance of several methods using different features derived from the sparse matrices, obtaining models with more than 80% of accuracy and acceptable prediction speed. These results are an important advance towards the automatic selection of the best GPU solver for a given sparse triangular linear system, and the characterization of the performance of these kernels. Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí |
SBAC-PAD | 2 |
| 2019 | Avoiding Synchronization to Accelerate a CFD Solver in GPUabstractThe caffa3d.MBRi is an open source, GPU-aware, general purpose incompressible flow solver, aimed at providing a useful tool for numerical simulation of real world fluid flow problems that require both geometrical flexibility and parallel computation capabilities to afford tens and hundreds million cells simulations. At the core of this tool there are a number of linear solvers that can be selected according to the characteristics of the problem to solve. For band matrices, the most efficient linear solver included in caffa3d.MBRi is the Strongly Implicit Procedure (SIP) solver. The parallelization of this solver follows the hyper-planes strategy, where the computations in one hyper-plane bare no dependencies and can be executed in parallel, while the hyper-planes have to be processed sequentially. In this work, we analyze this strategy to reach an efficient GPU implementation of the SIP solver for the caffa3d.MBRi. In particular, we design and implement a self-scheduling procedure to avoid the overhead of CPU-GPU synchronization implied by the hyper-planes strategy, outperforming the standard GPU implementation of the SIP by approximately 2×. Ernesto Dufrechu, Pablo Ezzatti, Gabriel Usera |
SBAC-PAD | 2 |
| 2019 | A GPU-aware mixed-precision solver for low-rank algebraic Riccati equationsabstractSummary We investigate different alternatives for the solution of algebraic Riccati equations on hybrid hardware platforms (ie, CPUs+GPUs). We evaluate a mixed‐precision approach that uses single‐precision arithmetic to obtain an approximation to the solution and later improve it to the desired precision, applying some steps of an economic iterative refinement. This method exploits the higher performance of the hardware to accelerate the solver when single‐precision arithmetic is employed and simultaneously obtains a high‐accuracy solution with the iterative refinement. We extend this approach to exploit the low‐rank property of the equation, when possible, to further improve its efficiency. The experimental evaluation shows that the mixed‐precision approach reports time and energy savings and provides similar or even more accurate solutions than well‐known methods like the sign function iteration or the structure‐preserving doubling algorithm. Peter Benner, Ernesto Dufrechu, Pablo Ezzatti, Alfredo Remón, Jens Saak |
Concurr. Comput. Pract. Exp. | 3 |
| 2019 | Power-aware computingabstractPower- Pablo Ezzatti, Enrique S. Quintana-Ortí, Alfredo Remón, Jens Saak |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | Accelerating the task/data-parallel version of ILUPACK's BiCG in multi-CPU/GPU configurations
José Ignacio Aliaga, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí |
Parallel Comput. | 3 |
| 2019 | An efficient GPU version of the preconditioned GMRES method
José Ignacio Aliaga, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí |
J. Supercomput. | 3 |
| 2018 | Extending ILUPACK with a GPU Version of the BiCGStab MethodabstractThe solution of sparse linear systems of large dimension is a important stage in problems that span a diverse kind of applications. For this reason, a number of iterative solvers have been developed, among which ILUPACK integrates an inverse-based multilevel ILU preconditioner with appealing numerical properties. In this work we extend the iterative methods available in ILUPACK. Concretely, we develop a data-parallel implementation of the BiCGStab method for GPUs hardware platforms that completes the functionality of ILUPACK-preconditioned solvers for general linear systems. The experimental evaluation carried out in a hybrid hardware platform, including a multicore CPU and a Nvidia GPU, shows that our novel proposal reaches speedups values between 5 and 10× when is compared with the CPU counterpart and values of up to 8.2× runtime reduction over other GPU solvers. José Ignacio Aliaga, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí |
CLEI | 3 |
| 2018 | A New GPU Algorithm to Compute a Level Set-Based Analysis for the Parallel Solution of Sparse Triangular SystemsabstractA myriad of problems in science and engineering, involve the solution of sparse triangular linear systems. They arise frequently as part of direct and iterative solvers for linear systems and eigenvalue problems, and hence can be considered as a key building block of sparse numerical linear algebra. This is why, since the early days, their parallel solution has been exhaustively studied, and efficient implementations of this kernel can be found for almost every hardware platform. In the GPU context, the most widespread implementation of this kernel is the one distributed in NVIDIA CUSPARSE library, which relies on a preprocessing stage to aggregate the unknowns of the triangular system into level sets. This determines an execution schedule for the solution of the system, where the level sets have to be processed sequentially while the unknowns that belong to one level set can be solved in parallel. One of the disadvantages of the CUSPARSE implementation is that this preprocessing stage is often extremely slow in comparison to the runtime of the solving phase. In this work, we present a parallel GPU algorithm that is able to compute the same level sets as CU S PARSE but takes significantly less runtime. Our experiments on a set of matrices from the SuiteSparse collection show acceleration factors of up to 44×. Additionally, we provide a routine capable of solving a triangular linear system on the same pass used to calculate the level sets, yielding important performance benefits. Ernesto Dufrechu, Pablo Ezzatti |
IPDPS | 2 |
| 2018 | Task Parallelism in the WRF Model Through Computation Offloading to Many-Core DevicesabstractIn the last decade the use of hybrid hardware (e.g., multicore processors + coprocessors) has been growing on the HPC field. However, this evolution in the HPC hardware has not been fully exploited by the WRF model since it shows limitations in the scalability when a large number of computing units are used. In a previous work, we proposed an asynchronous architecture for the WRF that overlaps the radiation computation with the execution of the rest of the model. In this work, we extend this idea with the aim of exploiting the computational power offered by hybrid hardware platforms. Specifically, we implement an OpenMP version of the asynchronous architecture and include the use of two types of coprocessors, a Xeon Phi and a GPU. The experimental evaluation performed shows that our proposal is able to adequately exploit these secondary computation devices, reaching interesting runtime reductions when solving tests cases from real scenarios. Rodrigo Baya, Claudio Porrini, Martín Pedemonte, Pablo Ezzatti |
PDP | 4 |
| 2018 | Solving Sparse Triangular Linear Systems in Modern GPUs: A Synchronization-Free AlgorithmabstractSparse triangular linear systems are ubiquitous in a wide range of science and engineering fields, and represent one of the most important building blocks of Sparse Numerical Lineal Algebra methods. For this reason, their parallel solution has been subject of exhaustive study, and efficient implementations of this kernel can be found for almost every hardware platform. However, the strong data dependencies that serialize a great deal of the execution and the load imbalance inherent to the triangular structure poses serious difficulties for its parallel performance, specially in the context of massively- parallel processors such as GPUs. To this day, the most widespread GPU implementation of this kernel is the one distributed in NVIDIA CUSPARSE library, which relies on a preprocessing stage to determine the parallel execution schedule. Although the solution phase is highly efficient, this strategy pays the cost of constant synchronizations with the CPU. In this work, we present a synchronization-free GPU al- gorithm to solve sparse triangular linear systems for the CSR format. The experimental evaluation shows performance improvements over CUSPARSE and a recently proposed synchronization-free method for the CSC matrix format. Ernesto Dufrechu, Pablo Ezzatti |
PDP | 2 |
| 2017 | A VNS with Parallel Evaluation of Solutions for the Inverse Lighting Problem
Ignacio Decia, Rodrigo Leira, Martín Pedemonte, Eduardo Fernández 0003, Pablo Ezzatti |
EvoApplications (1) | 5 |
| 2017 | Solving Sparse Differential Riccati Equations on Hybrid CPU-GPU Platforms
Peter Benner, Ernesto Dufrechu, Pablo Ezzatti, Hermann Mena, Enrique S. Quintana-Ortí, Alfredo Remón |
ICCSA (1) | 3 |
| 2017 | Overcoming Memory-Capacity Constraints in the Use of ILUPACK on Graphics ProcessorsabstractAn important number of scientific and engineering problems currently require the solution of large and sparse linear systems of equations. In previous work, we applied a GPU accelerator to the solution of sparse linear systems of moderate dimension via ILUPACK, showing important reductions in the execution time while maintaining the quality of the solution. Unfortunately, the use of GPUs attached to only one compute node strongly limits the memory available to solve the systems, and thus the size of the problems that can be tackled with this approach. In this work we introduce a distributed-parallel version of ILUPACK that overcomes these limitations. The results of the evaluation show that the inclusion of multiple GPUs, located on distinct nodes of a cluster, yields relevant reductions in the execution time for large problems and, more importantly, allows to increase the dimension of the problems, showing interesting scaling properties. José Ignacio Aliaga, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí |
SBAC-PAD | 3 |
| 2017 | Extending the Gauss-Huard method for the solution of Lyapunov matrix equations and matrix inversionabstractSummary The solution of linear systems is a recurrent operation in scientific and engineering applications, traditionally addressed via the LU factorization. The Gauss–Huard (GH) algorithm has been introduced as an efficient alternative in modern platforms equipped with accelerators, although this approach presented some functional constraints. In particular, it was not possible to reuse part of the computations in the solution of delayed linear systems or in the inversion of the matrix. Here, we adapt GH to overcome these two deficiencies of GH, yielding new algorithms that exhibit the same computational cost as their corresponding counterparts based on the LU factorization of the matrix. We evaluate the novel GH extensions on the solution of Lyapunov matrix equations via the LRCF‐ADI method, validating our approach via experiments with three benchmarks from model order reduction. Copyright © 2017 John Wiley & Sons, Ltd. Peter Benner, Pablo Ezzatti, Enrique S. Quintana-Ortí, Alfredo Remón |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | A comparison of various schemes for solving the transport equation in many-core platforms
Marcelo Bondarenco, Pablo Gamazo, Pablo Ezzatti |
J. Supercomput. | 3 |
| 2016 | Taking advantage of HPC techniques in the operational forecast of the Río de la PlataabstractIn this paper we address the use of high performance computing techniques with the aim of accelerating the runtime of a numerical model for the South Atlantic Ocean. This numerical model is a component of a larger system, that includes higher precision models to compute the hidrodynamics of the Río de la Plata river. Our work includes, in a first stage, a thorough study of the use of traditional HPC techniques on the CPU, and in a second stage we perform a preliminar study about the use of GPUs to accelerate the optimized version of the model. Specifically, for the first stage we made a profiling of the original model to identify the routines with larger computational cost. After that, we design and implement some variants for the three most expensive ones. Finally, we validate the proposals with an experimental evaluation. The results obtained show important acceleration values for the routines (up to 11 x), and these accelerations impact in the whole model with a runtime reduction of more than three times. Finally, we migrated one of the most costly routines of the optimized version to the GPU as a proof of concept. Rodrigo Baya, Ernesto Dufrechu, Pablo Ezzatti, Michelle Jackson, Mónica Fossati |
CLEI | 3 |
| 2016 | Assessing the explicit finite difference method on a massive parallel platformabstractThis work addresses the resolution of the transport (advection-diffusion) equation in 3D using an explicit scheme for the finite d ifference method. Our initative is motivated by the advantages offered by this scheme for parallel processing. We propose three implementations, a sequential code (in C) and two parallel versions (C-CUDA and C with OpenMP). The experimental comparison is focused on the performance of each implementation using different grid sizes, and in the case of the OpenMP implementation, several number of threads. Additionally, we measured the accurancy of this scheme when the detail of the discretization grows. The results show that the parallel implementations reach significant speed up compared with the sequential counterpart. In addition, the GPU variant offers an further runtime reduction of up to 10x. Marcelo Bondarenco, Pablo Gamazo, Pablo Ezzatti |
CLEI | 3 |
| 2016 | Overview of HPC benchmarks in hybrid hardware platforms (CPUs+GPUs)abstractThis work studies the use of architectures that include GPUs to accelerate the most popular benchmarks in the high performance computing field, namely HPL, HPCG and the one used for the Graph500 ranking. Specifically, the installation and configuration of implementations of these benchmarks that are able to exploit the computer power of different massively parallel hardware platforms are discussed in order to provide a helpful insight on these topics to other researchers. The results obtained in the experimental evaluation show the benefits of the use of this kind of hardware architectures in computation-bounded and memory-bounded algorithms. Danilo Espino, Gerardo Ares, Martín Pedemonte, Pablo Ezzatti |
CLEI | 4 |
| 2016 | The Impact of Panel Factorization on the Gauss-Huard Algorithm for the Solution of Linear Systems on Modern Architectures
Sandra Catalán, Pablo Ezzatti, Enrique S. Quintana-Ortí, Alfredo Remón |
ICA3PP | 2 |
| 2016 | Exploiting task and data parallelism in ILUPACK's preconditioned CG solver on NUMA architectures and many-core accelerators
José Ignacio Aliaga, Rosa M. Badia, Maria Barreda, Matthias Bollhöfer, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí |
Parallel Comput. | 6 |
| 2015 | Solving dense linear systems with hybrid ARM+GPU platformsabstractThe necessity of reducing the energy consumption while improving the computational performance has encouraged the development of new hardware platforms. In this line, hybrid architectures that integrate ARM processors with graphics accelerators offer a positive balance between computing capabilities and energy requirements. However, in order to make an efficient use of this hardware, it is necessary to develop new methods and computational kernels, as well as to adapt existing ones. The solution of linear systems of equations is a basic operation in the solution of different problems. Its relevance and computational cost has motivated an important amount of work, and in consequence, it is possible to find high performance solvers for most hardware platforms. In this work we study the solution of dense linear systems of equations in an NVIDIA Jetson TK1 device via the Gauss-Huard method. The experimental evaluation shows that the new solvers outperform the ones available in the MAGMA library for systems of dimesion n ≤ 6,000. Juan Pablo Silva, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí, Peter Benner, Alfredo Remón |
CLEI | 3 |
| 2015 | Extending lyapack for the solution of band Lyapunov equations on hybrid CPU-GPU platforms
Peter Benner, Alfredo Remón, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí |
J. Supercomput. | 4 |
| 2014 | Accelerating the general band matrix multiplication using graphics processorsabstractIn this paper, we leverage the intrinsic data-parallelism of the band matrix-matrix product to accelerate this operation on Graphics Processing Units (GPUs). In particular, we propose a Level-3 BLAS style algorithm to tackle the band matrix-matrix product and implement two GPU-based versions that off-load the most expensive computations - i.e., general dense matrix-matrix multiplication, triangular matrixmatrix multiplication and matrix addition - to the hardware accelerator. Results collected using GPUs for the two most recent generations of NVIDIA (“Fermi” and “Kepler”) and a complete set of benchmark cases (which differ in the matrix dimensions and bandwidth) show that the GPU-enabled implementations deliver a notable reduction of the execution time. Peter Benner, Alfredo Remón, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí |
CLEI | 4 |
| 2014 | Accelerating Band Linear Algebra Operations on GPUs with Application in Model Reduction
Peter Benner, Ernesto Dufrechu, Pablo Ezzatti, Pablo Igounet, Enrique S. Quintana-Ortí, Alfredo Remón |
ICCSA (6) | 3 |
| 2014 | Leveraging Data-Parallelism in ILUPACK using Graphics ProcessorsabstractIn this paper, we address the exploitation of data parallelism for the solution of sparse symmetric positive definite linear systems via iterative methods on Graphics Processing Units (GPUs). In particular, we accelerate the preconditioned CG-based iterative solver underlying the incomplete LU decomposition package (ILUPACK) by off-loading the most expensive computations i.e., The solution of sparse triangular systems and sparse matrix-vector products-to the hardware accelerator. The results collected using GPUs from the two most recent generations from NVIDIA ("Fermi" and "Kepler") and a benchmark test bed of sparse linear systems show that the GPU-enabled implementations deliver a notable reduction of the execution time, while maintaining the convergence rate and numerical properties of the original ILUPACK solver. José Ignacio Aliaga, Matthias Bollhöfer, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí |
ISPDC | 4 |
| 2014 | Another step to the full GPU implementation of the weather research and forecasting model
Juan Pablo Silva, José Hagopian, Marcel Burdiat, Ernesto Dufrechu, Martín Pedemonte, Alejandro Gutiérrez Arce, Gabriel Cazes, Pablo Ezzatti |
J. Supercomput. | 8 |
| 2013 | On the Impact of Optimization on the Time-Power-Energy Balance of Dense Linear Algebra Factorizations
Peter Benner, Pablo Ezzatti, Enrique S. Quintana-Ortí, Alfredo Remón |
ICA3PP (2) | 2 |
| 2013 | Matrix inversion on CPU-GPU platforms with applications in control theoryabstractSUMMARY In this paper, we tackle the inversion of large‐scale dense matrices via conventional matrix factorizations (LU, Cholesky, andLDLT) and the Gauss–Jordan method on hybrid platforms consisting of a multicore CPU and a many‐core graphics processor (GPU). Specifically, we introduce the different matrix inversion algorithms by using a unified framework based on the notation from the FLAME project; we develop hybrid implementations for those matrix operations underlying the algorithms, alternative to those in existing libraries for single GPU systems; and we perform an extensive experimental study on a platform equipped with state‐of‐the‐art general‐purpose architectures from Intel (Santa Clara, CA, USA) and a ‘Fermi’ GPU from NVIDIA (Santa Clara, CA, USA) that exposes the efficiency of the different inversion approaches. Our study and experimental results show the simplicity and performance advantage of the Gauss–Jordan elimination‐based inversion methods and the difficulties associated with the symmetric indefinite case. Copyright © 2012 John Wiley & Sons, Ltd. Peter Benner, Pablo Ezzatti, Enrique S. Quintana-Ortí, Alfredo Remón |
Concurr. Comput. Pract. Exp. | 2 |
| 2013 | Accelerating the Lyapack library using GPUs
Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí, Alfredo Remón |
J. Supercomput. | 2 |
| 2012 | Accelerating radiative heat transfer calculations on modern hardwareabstractWidely used for the resolution of radiative heat transfer calculations, the radiosity method entails the use of view factors which may require significant calculation efforts in complex geometries. In this work, we study the heat transfer of the filament of an incandescent light bulb using the radiosity method. Due to the high computational cost of the Monte Carlo method used for computing the view factors, two high-performance computing (HPC) techniques, namely, a parallel multi-core approach based on OpenMP and a massively parallel implementation over two different graphics processing units (GPUs) in CUDA were study. The use of such techniques enabled to reduce the calculation time up to 10× in two Quad Core INTEL Xeon processors, 112× in an NVIDIA Tesla C1060 GPU and 199× in an NVIDIA Tesla C2070. Ernesto Dufrechu, Federico Favre, Martín Pedemonte, Pedro Curto, Pablo Ezzatti |
CLEI | 5 |
| 2012 | GPU Acceleration of the caffa3d.MB Model
Pablo Igounet, Pablo Alfaro, Gabriel Usera, Pablo Ezzatti |
ICCSA (4) | 4 |
| 2012 | High Performance Implementations of the BST Method on Hybrid CPU-GPU PlatformsabstractModel order reduction is necessary in many complex scientific and engineering applications. Among the methods for model reduction, those based on the SVD are well-known for their beneficial theoretical properties, though they require O(n3) floating-point arithmetic operations, with n being in the range of 103- 104for many practical applications. In this paper we propose several high performance implementations of the Balanced Stochastic Truncation method for model reduction. The new routines carefully distribute the computations among the computational resources of a hybrid platform composed of one (or more) multicore CPU(s) and a manycore GPU. Our results show that model reduction of a large-scale problem with 9,669 state variables, which previously required the use of a cluster of computers, can now be carried out in the target platform in less than 25 minutes. Peter Benner, Pablo Ezzatti, Enrique S. Quintana-Ortí, Alfredo Remón |
ISPA | 2 |
| 2011 | Efficient Model Order Reduction of Large-Scale Systems on Multi-core Platforms
Pablo Ezzatti, Enrique S. Quintana-Ortí, Alfredo Remón |
ICCSA (5) | 1 |
| 2011 | High Performance Matrix Inversion on a Multi-core Platform with Several GPUsabstractInversion of large-scale matrices appears in a few scientific applications like model reduction or optimal control. Matrix inversion requires an important computational effort and, therefore, the application of high performance computing techniques and architectures for matrices with dimension in the order of thousands. Following the recent uprise of graphics processors (GPUs), we present and evaluate high performance codes for matrix inversion, based on Gauss-Jordan elimination with partial pivoting, which off-load the main computational kernels to one or more GPUs while performing fine-grain operations on the general-purpose processor. The target architecture consists of a multi-core processor connected to several GPUs. Parallelism is extracted from parallel implementations of BLAS and from the concurrent execution of operations in the available computational units. Numerical experiments on a system with two Intel QuadCore processors and four NVIDIA cl060 GPUs illustrate the efficiency and the scalability of the different implementations, which deliver over 1.2 x 1012floating point operations per second. Pablo Ezzatti, Enrique S. Quintana-Ortí, Alfredo Remón |
PDP | 1 |
| 2011 | A mixed-precision algorithm for the solution of Lyapunov equations on hybrid CPU-GPU platforms
Peter Benner, Pablo Ezzatti, Daniel Kressner, Enrique S. Quintana-Ortí, Alfredo Remón |
Parallel Comput. | 2 |
| 2011 | Using graphics processors to accelerate the computation of the matrix inverse
Pablo Ezzatti, Enrique S. Quintana-Ortí, Alfredo Remón |
J. Supercomput. | 1 |
| 2010 | PUGACE, a cellular Evolutionary Algorithm framework on GPUsabstractMetaheuristics are used for solving optimization problems since they are able to compute near optimal solutions in reasonable times. However, solving large instances it may pose a challenge even for these techniques. For this reason, metaheuristics parallelization is an interesting alternative in order to decrease the execution time and to provide a different search pattern. In the last years, GPUs have evolved at a breathtaking pace. Originally, they were specific-purpose devices, but in a few years they became general-purpose shared memory multiprocessors. Nowadays, these devices are a powerful low cost platform for implementing parallel algorithms. In this paper, we present a preliminary version of PUGACE, a cellular Evolutionary Algorithm framework implemented on GPU. PUGACE was designed with the goal of providing a tool for easily developing this kind of algorithms. The experimental results when solving the Quadratic Assignment Problem are presented to show the potential of the proposed framework. Nicolas Soca, Jose Luis Blengio, Martín Pedemonte, Pablo Ezzatti |
IEEE Congress on Evolutionary Computation | 4 |