EDBT 2026 Demo / reviewers in the wild / expert
Ernesto Dufrechu
dblp:132/7493 · also Ernesto Dufrechou
· DBLP profile ↗
31ranked-venue papers
8as first author
10since 2021 · last 2026
0000-0003-4971-340XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 6 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optimizing SpMV Kernel for Banded Matrices on FPGAs Using High-Level Synthesis
Federico Favaro, Ernesto Dufrechu, Juan P. Oliver, Pablo Ezzatti |
Euro-Par (1) | 2 |
| 2025 | Load-balanced SpMV kernel for the bmSparse matrix formatabstractThe design of sparse matrix storage formats is essential to achieve high-performance sparse kernels in modern parallel architectures. The bitmap-based bmSparse format was designed with the SPGEMM operation as its main focus, but it shows potential for other operations as well when the sparse matrix has a convenient structure. In this paper, we propose a new SPMV kernel that greatly improves the load balance of previous open-source implementations and utilizes the GPU resources more efficiently. The results show speedups of up to $100 \times$ regarding other SPMV kernels for bmSparse. Finally, we leverage a hybrid implementation between the new and existing approach, selecting the kernel which will probably result in a better performance depending on each matrix characteristics. Gonzalo Berger, Ernesto Dufrechu, Pablo Ezzatti |
PDP | 2 |
| 2025 | A synchronization-free incomplete LU factorization for GPUs with level-set analysisabstractIncomplete factorization methods are powerful algebraic preconditioners widely used to accelerate the convergence of linear solvers. The parallelization of ILU methods has been extensively studied, particularly for GPUs, which are ubiquitous parallel computing devices. In recent years, synchronization-free methods have become the mainstream approach for solving sparse triangular linear systems.Although the sparse triangular solver and ILU factorization are closely related, the application of synchronization-free strategies to ILU factorization has not been explored in the literature to the same extent as the triangular solver. In this work, we present synchronization-free implementations of the ILU-0 preconditioner on GPUs. Specifically, we propose three implementations that vary in how row updates are handled after each coefficient elimination, as well as an additional approach that leverages a prior level-set analysis to optimize the execution schedule. Manuel Freire 0002, Ernesto Dufrechu, Pablo Ezzatti |
PDP | 2 |
| 2024 | A new level-set analysis and sparse storage format for the SPTRSV in GPUsabstractDue to its relevant role in many numerical methods, the solution of sparse triangular linear systems (SpTRSV) in parallel platforms is continuously studied to extract as much performance as possible from the latest hardware architectures. In the case of GPUs, the latest solvers use the synchronization-free paradigm. When the problem involves several system solutions for the same matrix, they often pre-process it through a levelset analysis to improve the equation solution scheduling in the solution phase. In addition, other optimizations address the load balancing issues and irregular memory access of the SpTRSV. In this work, we modify the classical approach to compute the level sets used in the parallel SpTRSV computation, and we show that the new strategy generally reduces the computation time of the solver. Furthermore, we design an internal matrix representation that can significantly accelerate the solution stage at the cost of increasing the memory storage requirements of the algorithm. The experimental evaluation shows that the proposed modifications can improve the performance of a recent levelset and synchronization-free solver by up to 70%, significantly outperforming other state-of-the-art solvers, especially when several linear systems must be solved for each analysis phase. Manuel Freire 0002, Ernesto Dufrechu, Pablo Ezzatti |
SBAC-PAD | 2 |
| 2023 | Evaluation of architecture-aware optimization techniques for Convolutional Neural NetworksabstractThe growing need to perform Neural network inference with low latency is giving place to a broad spectrum of heterogeneous devices with deep learning capabilities. Therefore, obtaining the best performance from each device and choosing the most suitable platform for a given problem has become challenging. This paper evaluates multiple inference platforms using architecture-aware optimizations for convolutional neural networks. Specifically, we use TensorRT and OpenVINO frameworks for hardware optimizations on top of the platform-aware NetAdapt algorithm. The experimental evaluation shows that on MobileNet and AlexNet, using NetAdapt with TensorRT or Open-VINO can improve latency up to 10 x and 5.3 x, respectively. Moreover, a throughput test using different batch sizes showed variable performance improvement on the different devices. Discussing the experimental results can guide the selection of devices and optimizations for different AI solutions. Raúl Marichal, Guillermo Toyos, Ernesto Dufrechu, Pablo Ezzatti |
PDP | 3 |
| 2023 | Advancing on an efficient sparse matrix multiplication kernel for modern GPUsabstractSummary The sparse matrix multiplication (SpGeMM) increased its importance in the last years due to its data science and machine learning applications. Consequently, considerable research has focused on accelerating this kernel in GPUs. Designing massively‐parallel algorithms for the SpGeMM is a challenging task since the computation pattern is highly irregular, and the required memory and operations depend on the interaction between the nonzero layout of the inputs. One strategy to attack this kernel consists of proposing new sparse matrix storage formats that contribute to mitigating this irregularity. In previous work, we commenced a study of the recently proposed bmSparse matrix format, suggesting several modifications to the SpGeMM algorithm. This work integrates the previous extensions and proposes new improvements to unleash bmSparse's full potential before comparing it to more consolidated options. In particular, we enhance one of the most computationally demanding stages with an adaptive technique, apply optimizations to achieve more efficient data accesses, and analyze the effect of using Tensor Cores to accelerate the multiplication stage of the algorithm. The experimental results on a set of real‐world sparse matrices show that the optimized implementation largely outperforms vendor implementations such as NVIDIA cuSparse Intel MKL‐CSR variant, while being competitive with MKL's‐BSR. Gonzalo Berger, Manuel Freire 0002, Renzo Marini, Ernesto Dufrechu, Pablo Ezzatti |
Concurr. Comput. Pract. Exp. | 4 |
| 2023 | A GPU method for the analysis stage of the SPTRSV kernel
Manuel Freire 0002, Juan Ferrand, Franco Seveso, Ernesto Dufrechu, Pablo Ezzatti |
J. Supercomput. | 4 |
| 2021 | Assessing the solution of one sparse triangular linear system on multi-many core platformsabstractThe solution of sparse triangular linear systems is an important building block for a large number of numerical methods used in science and engineering. It is then crucial to count with implementations of this operation that can execute efficiently in the most recent hardware platforms. In the case of GPUs, several methods have been proposed in the last years. These methods belong to two main categories. On the one hand, there are the methods that rely on a previous analysis of the sparse matrix to determine a better execution schedule and, on the other hand, there are methods that decide this scheduling dynamically. The experimental results in the literature are not conclussive in favour of any of these strategies. However, the experimental evaluations usually focus on the use case where many systems have to be solved with the same sparse matrix, where the analysis phase needs to be performed only once and its cost is not important in relation with the total runtime. In this work we are interested in determining which is the best strategy, according to the degree of parallelism of the problem, when only one sytem is to be solved. The experimental evaluation performed on NVIDIA P100 accelerators shows that the self-scheduled routines present important advantages when the degree of parallelism of the problem allows it. Raúl Marichal, Ernesto Dufrechu, Pablo Ezzatti |
CLEI | 2 |
| 2021 | Machine learning for optimal selection of sparse triangular system solvers on GPUs
Ernesto Dufrechu, Pablo Ezzatti, Manuel Freire 0002, Enrique S. Quintana-Ortí |
J. Parallel Distributed Comput. | 1 |
| 2021 | Factorized solution of generalized stable Sylvester equations using many-core GPU accelerators
Peter Benner, Ernesto Dufrechu, Pablo Ezzatti, Rodrigo Gallardo, Enrique S. Quintana-Ortí |
J. Supercomput. | 2 |
| 2020 | Using analysis information in the synchronization-free GPU solution of sparse triangular systemsabstractSummary The solution of sparse triangular linear systems is one of the most important building blocks for a large number of science and engineering problems. For these reasons, it has been studied steadily for several decades, principally in order to take advantage of emerging parallel platforms. In the context of massively parallel platforms such as GPUs, the standard strategy of parallel solution is based on performing a level‐set analysis of the sparse matrix, and the kernel included in the nVidia cuSparse library is the most prominent example of this approach. However, a weak spot of this implementation is the costly analysis phase and the constant synchronizations with the CPU during the solution stage. In previous work, we presented a self‐scheduled and synchronization‐free GPU algorithm that avoided the analysis phase and the synchronizations of the standard approach. Here, we extend this proposal and show how the level‐set information can be leveraged to improve its performance. In particular, we present new GPU solution routines that attack some of the weak spots of the self‐scheduled solver, such as the under‐utilization of the GPU resources in the case of highly sparse matrices. The experimental evaluation reveals a sensible runtime reduction over cuSparse and the state‐of‐the‐art synchronization‐free method. Ernesto Dufrechu, Pablo Ezzatti |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | Automatic Selection of Sparse Triangular Linear System Solvers on GPUs through Machine Learning TechniquesabstractThe solution of sparse triangular linear systems is often the most time-consuming stage of preconditioned iterative methods to solve general sparse linear systems, where it has to be applied several times for the same sparse matrix. For this reason, its computational performance has a strong impact on a wide range of scientific and engineering applications, which has motivated the study of its efficient execution on massively parallel platforms. In this sense, several methods have been proposed to tackle this operation on graphics processing units (GPUs), which can be classified under either the level-set or the self-scheduling paradigms. The results obtained from the experimental evaluation of the different methods suggest that both paradigms perform well for certain problems but poorly for others. Additionally, the relation between the properties of the linear systems and the performance of the different solvers is not evident a-priori. In this context, techniques that allow to predict inexpensively which is be the best solver for a particular linear system can lead to important runtime reductions. Our approach leverages machine learning techniques to select the best sparse triangular solver for a given linear system, with focus on the case where a small number of triangular systems has to be solved for the same matrix. We study the performance of several methods using different features derived from the sparse matrices, obtaining models with more than 80% of accuracy and acceptable prediction speed. These results are an important advance towards the automatic selection of the best GPU solver for a given sparse triangular linear system, and the characterization of the performance of these kernels. Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí |
SBAC-PAD | 1 |
| 2019 | Avoiding Synchronization to Accelerate a CFD Solver in GPUabstractThe caffa3d.MBRi is an open source, GPU-aware, general purpose incompressible flow solver, aimed at providing a useful tool for numerical simulation of real world fluid flow problems that require both geometrical flexibility and parallel computation capabilities to afford tens and hundreds million cells simulations. At the core of this tool there are a number of linear solvers that can be selected according to the characteristics of the problem to solve. For band matrices, the most efficient linear solver included in caffa3d.MBRi is the Strongly Implicit Procedure (SIP) solver. The parallelization of this solver follows the hyper-planes strategy, where the computations in one hyper-plane bare no dependencies and can be executed in parallel, while the hyper-planes have to be processed sequentially. In this work, we analyze this strategy to reach an efficient GPU implementation of the SIP solver for the caffa3d.MBRi. In particular, we design and implement a self-scheduling procedure to avoid the overhead of CPU-GPU synchronization implied by the hyper-planes strategy, outperforming the standard GPU implementation of the SIP by approximately 2×. Ernesto Dufrechu, Pablo Ezzatti, Gabriel Usera |
SBAC-PAD | 1 |
| 2019 | A GPU-aware mixed-precision solver for low-rank algebraic Riccati equationsabstractSummary We investigate different alternatives for the solution of algebraic Riccati equations on hybrid hardware platforms (ie, CPUs+GPUs). We evaluate a mixed‐precision approach that uses single‐precision arithmetic to obtain an approximation to the solution and later improve it to the desired precision, applying some steps of an economic iterative refinement. This method exploits the higher performance of the hardware to accelerate the solver when single‐precision arithmetic is employed and simultaneously obtains a high‐accuracy solution with the iterative refinement. We extend this approach to exploit the low‐rank property of the equation, when possible, to further improve its efficiency. The experimental evaluation shows that the mixed‐precision approach reports time and energy savings and provides similar or even more accurate solutions than well‐known methods like the sign function iteration or the structure‐preserving doubling algorithm. Peter Benner, Ernesto Dufrechu, Pablo Ezzatti, Alfredo Remón, Jens Saak |
Concurr. Comput. Pract. Exp. | 2 |
| 2019 | Accelerating the task/data-parallel version of ILUPACK's BiCG in multi-CPU/GPU configurations
José Ignacio Aliaga, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí |
Parallel Comput. | 2 |
| 2019 | An efficient GPU version of the preconditioned GMRES method
José Ignacio Aliaga, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí |
J. Supercomput. | 2 |
| 2018 | Extending ILUPACK with a GPU Version of the BiCGStab MethodabstractThe solution of sparse linear systems of large dimension is a important stage in problems that span a diverse kind of applications. For this reason, a number of iterative solvers have been developed, among which ILUPACK integrates an inverse-based multilevel ILU preconditioner with appealing numerical properties. In this work we extend the iterative methods available in ILUPACK. Concretely, we develop a data-parallel implementation of the BiCGStab method for GPUs hardware platforms that completes the functionality of ILUPACK-preconditioned solvers for general linear systems. The experimental evaluation carried out in a hybrid hardware platform, including a multicore CPU and a Nvidia GPU, shows that our novel proposal reaches speedups values between 5 and 10× when is compared with the CPU counterpart and values of up to 8.2× runtime reduction over other GPU solvers. José Ignacio Aliaga, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí |
CLEI | 2 |
| 2018 | A New GPU Algorithm to Compute a Level Set-Based Analysis for the Parallel Solution of Sparse Triangular SystemsabstractA myriad of problems in science and engineering, involve the solution of sparse triangular linear systems. They arise frequently as part of direct and iterative solvers for linear systems and eigenvalue problems, and hence can be considered as a key building block of sparse numerical linear algebra. This is why, since the early days, their parallel solution has been exhaustively studied, and efficient implementations of this kernel can be found for almost every hardware platform. In the GPU context, the most widespread implementation of this kernel is the one distributed in NVIDIA CUSPARSE library, which relies on a preprocessing stage to aggregate the unknowns of the triangular system into level sets. This determines an execution schedule for the solution of the system, where the level sets have to be processed sequentially while the unknowns that belong to one level set can be solved in parallel. One of the disadvantages of the CUSPARSE implementation is that this preprocessing stage is often extremely slow in comparison to the runtime of the solving phase. In this work, we present a parallel GPU algorithm that is able to compute the same level sets as CU S PARSE but takes significantly less runtime. Our experiments on a set of matrices from the SuiteSparse collection show acceleration factors of up to 44×. Additionally, we provide a routine capable of solving a triangular linear system on the same pass used to calculate the level sets, yielding important performance benefits. Ernesto Dufrechu, Pablo Ezzatti |
IPDPS | 1 |
| 2018 | Solving Sparse Triangular Linear Systems in Modern GPUs: A Synchronization-Free AlgorithmabstractSparse triangular linear systems are ubiquitous in a wide range of science and engineering fields, and represent one of the most important building blocks of Sparse Numerical Lineal Algebra methods. For this reason, their parallel solution has been subject of exhaustive study, and efficient implementations of this kernel can be found for almost every hardware platform. However, the strong data dependencies that serialize a great deal of the execution and the load imbalance inherent to the triangular structure poses serious difficulties for its parallel performance, specially in the context of massively- parallel processors such as GPUs. To this day, the most widespread GPU implementation of this kernel is the one distributed in NVIDIA CUSPARSE library, which relies on a preprocessing stage to determine the parallel execution schedule. Although the solution phase is highly efficient, this strategy pays the cost of constant synchronizations with the CPU. In this work, we present a synchronization-free GPU al- gorithm to solve sparse triangular linear systems for the CSR format. The experimental evaluation shows performance improvements over CUSPARSE and a recently proposed synchronization-free method for the CSC matrix format. Ernesto Dufrechu, Pablo Ezzatti |
PDP | 1 |
| 2017 | Solving Sparse Differential Riccati Equations on Hybrid CPU-GPU Platforms
Peter Benner, Ernesto Dufrechu, Pablo Ezzatti, Hermann Mena, Enrique S. Quintana-Ortí, Alfredo Remón |
ICCSA (1) | 2 |
| 2017 | Overcoming Memory-Capacity Constraints in the Use of ILUPACK on Graphics ProcessorsabstractAn important number of scientific and engineering problems currently require the solution of large and sparse linear systems of equations. In previous work, we applied a GPU accelerator to the solution of sparse linear systems of moderate dimension via ILUPACK, showing important reductions in the execution time while maintaining the quality of the solution. Unfortunately, the use of GPUs attached to only one compute node strongly limits the memory available to solve the systems, and thus the size of the problems that can be tackled with this approach. In this work we introduce a distributed-parallel version of ILUPACK that overcomes these limitations. The results of the evaluation show that the inclusion of multiple GPUs, located on distinct nodes of a cluster, yields relevant reductions in the execution time for large problems and, more importantly, allows to increase the dimension of the problems, showing interesting scaling properties. José Ignacio Aliaga, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí |
SBAC-PAD | 2 |
| 2016 | Taking advantage of HPC techniques in the operational forecast of the Río de la PlataabstractIn this paper we address the use of high performance computing techniques with the aim of accelerating the runtime of a numerical model for the South Atlantic Ocean. This numerical model is a component of a larger system, that includes higher precision models to compute the hidrodynamics of the Río de la Plata river. Our work includes, in a first stage, a thorough study of the use of traditional HPC techniques on the CPU, and in a second stage we perform a preliminar study about the use of GPUs to accelerate the optimized version of the model. Specifically, for the first stage we made a profiling of the original model to identify the routines with larger computational cost. After that, we design and implement some variants for the three most expensive ones. Finally, we validate the proposals with an experimental evaluation. The results obtained show important acceleration values for the routines (up to 11 x), and these accelerations impact in the whole model with a runtime reduction of more than three times. Finally, we migrated one of the most costly routines of the optimized version to the GPU as a proof of concept. Rodrigo Baya, Ernesto Dufrechu, Pablo Ezzatti, Michelle Jackson, Mónica Fossati |
CLEI | 2 |
| 2016 | Exploiting task and data parallelism in ILUPACK's preconditioned CG solver on NUMA architectures and many-core accelerators
José Ignacio Aliaga, Rosa M. Badia, Maria Barreda, Matthias Bollhöfer, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí |
Parallel Comput. | 5 |
| 2015 | Solving dense linear systems with hybrid ARM+GPU platformsabstractThe necessity of reducing the energy consumption while improving the computational performance has encouraged the development of new hardware platforms. In this line, hybrid architectures that integrate ARM processors with graphics accelerators offer a positive balance between computing capabilities and energy requirements. However, in order to make an efficient use of this hardware, it is necessary to develop new methods and computational kernels, as well as to adapt existing ones. The solution of linear systems of equations is a basic operation in the solution of different problems. Its relevance and computational cost has motivated an important amount of work, and in consequence, it is possible to find high performance solvers for most hardware platforms. In this work we study the solution of dense linear systems of equations in an NVIDIA Jetson TK1 device via the Gauss-Huard method. The experimental evaluation shows that the new solvers outperform the ones available in the MAGMA library for systems of dimesion n ≤ 6,000. Juan Pablo Silva, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí, Peter Benner, Alfredo Remón |
CLEI | 2 |
| 2015 | Extending lyapack for the solution of band Lyapunov equations on hybrid CPU-GPU platforms
Peter Benner, Alfredo Remón, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí |
J. Supercomput. | 3 |
| 2014 | Accelerating the general band matrix multiplication using graphics processorsabstractIn this paper, we leverage the intrinsic data-parallelism of the band matrix-matrix product to accelerate this operation on Graphics Processing Units (GPUs). In particular, we propose a Level-3 BLAS style algorithm to tackle the band matrix-matrix product and implement two GPU-based versions that off-load the most expensive computations - i.e., general dense matrix-matrix multiplication, triangular matrixmatrix multiplication and matrix addition - to the hardware accelerator. Results collected using GPUs for the two most recent generations of NVIDIA (“Fermi” and “Kepler”) and a complete set of benchmark cases (which differ in the matrix dimensions and bandwidth) show that the GPU-enabled implementations deliver a notable reduction of the execution time. Peter Benner, Alfredo Remón, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí |
CLEI | 3 |
| 2014 | Accelerating Band Linear Algebra Operations on GPUs with Application in Model Reduction
Peter Benner, Ernesto Dufrechu, Pablo Ezzatti, Pablo Igounet, Enrique S. Quintana-Ortí, Alfredo Remón |
ICCSA (6) | 2 |
| 2014 | Leveraging Data-Parallelism in ILUPACK using Graphics ProcessorsabstractIn this paper, we address the exploitation of data parallelism for the solution of sparse symmetric positive definite linear systems via iterative methods on Graphics Processing Units (GPUs). In particular, we accelerate the preconditioned CG-based iterative solver underlying the incomplete LU decomposition package (ILUPACK) by off-loading the most expensive computations i.e., The solution of sparse triangular systems and sparse matrix-vector products-to the hardware accelerator. The results collected using GPUs from the two most recent generations from NVIDIA ("Fermi" and "Kepler") and a benchmark test bed of sparse linear systems show that the GPU-enabled implementations deliver a notable reduction of the execution time, while maintaining the convergence rate and numerical properties of the original ILUPACK solver. José Ignacio Aliaga, Matthias Bollhöfer, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí |
ISPDC | 3 |
| 2014 | Another step to the full GPU implementation of the weather research and forecasting model
Juan Pablo Silva, José Hagopian, Marcel Burdiat, Ernesto Dufrechu, Martín Pedemonte, Alejandro Gutiérrez Arce, Gabriel Cazes, Pablo Ezzatti |
J. Supercomput. | 4 |
| 2013 | Accelerating the Lyapack library using GPUs
Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí, Alfredo Remón |
J. Supercomput. | 1 |
| 2012 | Accelerating radiative heat transfer calculations on modern hardwareabstractWidely used for the resolution of radiative heat transfer calculations, the radiosity method entails the use of view factors which may require significant calculation efforts in complex geometries. In this work, we study the heat transfer of the filament of an incandescent light bulb using the radiosity method. Due to the high computational cost of the Monte Carlo method used for computing the view factors, two high-performance computing (HPC) techniques, namely, a parallel multi-core approach based on OpenMP and a massively parallel implementation over two different graphics processing units (GPUs) in CUDA were study. The use of such techniques enabled to reduce the calculation time up to 10× in two Quad Core INTEL Xeon processors, 112× in an NVIDIA Tesla C1060 GPU and 199× in an NVIDIA Tesla C2070. Ernesto Dufrechu, Federico Favre, Martín Pedemonte, Pedro Curto, Pablo Ezzatti |
CLEI | 1 |