José Ignacio Aliaga

dblp:29/5218 · DBLP profile ↗
← Back
25ranked-venue papers
20as first author
6since 2021 · last 2026
0000-0001-8469-764XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 15 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Resource optimization with MPI process malleability for dynamic workloads in HPC clusters
abstract
Dynamic resource management is essential for optimizing computational efficiency in modern high-performance computing (HPC) environments, particularly as systems scale. While research has demonstrated the benefits of malleability in resource management systems (RMS), the adoption of such techniques in production environments remains limited due to challenges in standardization, interoperability, and usability. Addressing these gaps, this paper extends our prior work on the Dynamic Management of Resources (DMR) framework, which provides a modular and user-friendly approach to dynamic resource allocation. Building upon the original DMRlib reconfiguration runtime, this work integrates new methodology from the Malleability Module (MaM) of the Proteo framework, further enhancing reconfiguration capabilities with new spawning strategies and data redistribution methods. In this paper, we explore new malleability strategies in HPC dynamic workloads, such as merging MPI communicators and asynchronous reconfigurations, which offer new opportunities for dramatically reducing memory overhead. The proposed enhancements are rigorously evaluated on a world-class supercomputer, demonstrating improved resource utilization and workload efficiency. Results show that dynamic resource management can reduce the workload completion time by 40% and increase the resource utilization by over 20%, compared to static resource allocation.
Sergio Iserte, Iker Martín-Álvarez, Krzysztof Rojek, José Ignacio Aliaga, María Isabel Castillo, Weronika Folwarska, Antonio J. Peña
Future Gener. Comput. Syst.4
2024 Proteo: a framework for the generation and evaluation of malleable MPI applications
abstract
Abstract Applying malleability to HPC systems can increase their productivity without degrading or even improving the performance of running applications. This paper presents Proteo, a configurable framework that allows to design benchmarks to study the effect of malleability on a system, and also incorporates malleability into a real application. Proteo consists of two modules: SAM allows to emulate the computational behavior of iterative scientific MPI applications, and MaM is able to reconfigure a job during execution, adjusting the number of processes, redistributing data, and resuming execution. An in-depth study of all the possibilities shows that Proteo is able to behave like a real malleable or non-malleable application in the range [0.85, 1.15]. Furthermore, the different methods defined in MaM for process management and data redistribution are analyzed, concluding that asynchronous malleability, where reconfiguration and application execution overlap, results in a 1.15 $$\times$$ × speedup.
Iker Martín-Álvarez, José Ignacio Aliaga, María Isabel Castillo, Sergio Iserte
J. Supercomput.2
2023 Configurable synthetic application for studying malleability in HPC
abstract
Nowadays, the throughput improvement in large clusters of computers recommends the development of malleable applications. Thus, during the execution of these applications in a job, the resource management system (RMS) can modify its resource allocation, in order to increase the global throughput. There are different alternatives to complete the different steps in which the reallocation of resources is decomposed. To find the best alternatives, this paper introduces a configurable synthetic iterative MPI malleable application capable of modifying, in execution time, the number of MPI processes according to several parameters. The application includes a performance module to measure stages time within steps, from processes management to data redistribution. In this way, the analysis of different scenarios will allow to conclude how the reconfiguration of application has to be made in different circumstances. At the same time, this tool can be used to create workloads that will allow to analyse the impact of malleability on a system and the work in progress.
Iker Martín-Álvarez, José Ignacio Aliaga, María Isabel Castillo, Sergio Iserte
PDP2
2023 Sparse matrix-vector and matrix-multivector products for the truncated SVD on graphics processors
abstract
Summary Many practical algorithms for numerical rank computations implement an iterative procedure that involves repeated multiplications of a vector, or a collection of vectors, with both a sparse matrix and its transpose. Unfortunately, the realization of these sparse products on current high performance libraries often deliver much lower arithmetic throughput when the matrix involved in the product is transposed. In this work, we propose a hybrid sparse matrix layout, named CSRC, that combines the flexibility of some well‐known sparse formats to offer a number of appealing properties: (1) CSRC can be obtained at low cost from the popular CSR (compressed sparse row) format; (2) CSRC has similar storage requirements as CSR; and especially, (3) the implementation of the sparse product kernels delivers high performance for both the direct product and its transposed variant on modern graphics accelerators thanks to a significant reduction of atomic operations compared to a conventional implementation based on CSR. This solution thus renders considerably higher performance when integrated into an iterative algorithm for the truncated singular value decomposition (SVD), such as the randomized SVD or, as demonstrated in the experimental results, the block Golub–Kahan–Lanczos algorithm.
José Ignacio Aliaga, Hartwig Anzt, Enrique S. Quintana-Ortí, Andrés Tomás
Concurr. Comput. Pract. Exp.1
2022 Compression and load balancing for efficient sparse matrix-vector product on multicore processors and graphics processing units
abstract
Summary We contribute to the optimization of the sparse matrix‐vector product by introducing a variant of the coordinate sparse matrix format that balances the workload distribution and compresses both the indexing arrays and the numerical information. Our approach is multi‐platform, in the sense that the realizations for (general‐purpose) multicore processors as well as graphics accelerators (GPUs) are built upon common principles, but differ in the implementation details, which are adapted to avoid thread divergence in the GPU case or maximize compression element‐wise (i.e., for each matrix entry) for multicore architectures. Our evaluation on the two last generations of NVIDIA GPUs as well as Intel and AMD processors demonstrate the benefits of the new kernels when compared with the optimized implementations of the sparse matrix‐vector product in NVIDIA's cuSPARSE and Intel's MKL, respectively.
José Ignacio Aliaga, Hartwig Anzt, Thomas Grützmacher, Enrique S. Quintana-Ortí, Andrés Tomás
Concurr. Comput. Pract. Exp.1
2021 Malleability Implementation in a MPI Iterative Method
abstract
In this poster is evaluated the data redistribution stage for two malleable versions of the Conjugate Gradient. One version is based on synchronous communications, while the other one uses asynchronous communications to overlap computation and data redistribution. Both improve execution time when adding more processes, but there is not a noticeable difference between them, because the asynchronous method lowers the performance of the iterations due to the method’s own communications. When both versions are compared, the synchronous version is preferred when resizing to more processes, while the asynchronous one achieves better times when resizing to fewer processes.
Iker Martín-Álvarez, José Ignacio Aliaga, María Isabel Castillo, Rafael Mayo 0002, Sergio Iserte
CLUSTER2
2020 Iteration-fusing conjugate gradient for sparse linear systems with MPI + OmpSs
Maria Barreda, José Ignacio Aliaga, Vicenç Beltran 0001, Marc Casas
J. Supercomput.2
2019 Energy-aware strategies for task-parallel sparse linear system solvers
abstract
Summary We present several energy‐aware strategies to improve the energy efficiency of a task‐parallel preconditioned Conjugate Gradient (PCG) iterative solver on a Haswell‐EP Intel Xeon. These techniques leverage the power‐saving states of the processor, promoting the hardware into a more energy‐efficient C‐state and modifying the CPU frequency (P‐states of the processors) of some operations of the PCG. We demonstrate that the application of these strategies during the main operations of the iterative solver can reduce its energy consumption considerably, especially for memory‐bound computations.
José Ignacio Aliaga, Maria Barreda, M. Asunción Castaño
Concurr. Comput. Pract. Exp.1
2019 Accelerating the task/data-parallel version of ILUPACK's BiCG in multi-CPU/GPU configurations
José Ignacio Aliaga, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí
Parallel Comput.1
2019 An efficient GPU version of the preconditioned GMRES method
José Ignacio Aliaga, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí
J. Supercomput.1
2018 Extending ILUPACK with a GPU Version of the BiCGStab Method
abstract
The solution of sparse linear systems of large dimension is a important stage in problems that span a diverse kind of applications. For this reason, a number of iterative solvers have been developed, among which ILUPACK integrates an inverse-based multilevel ILU preconditioner with appealing numerical properties. In this work we extend the iterative methods available in ILUPACK. Concretely, we develop a data-parallel implementation of the BiCGStab method for GPUs hardware platforms that completes the functionality of ILUPACK-preconditioned solvers for general linear systems. The experimental evaluation carried out in a hybrid hardware platform, including a multicore CPU and a Nvidia GPU, shows that our novel proposal reaches speedups values between 5 and 10× when is compared with the CPU counterpart and values of up to 8.2× runtime reduction over other GPU solvers.
José Ignacio Aliaga, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí
CLEI1
2017 Overcoming Memory-Capacity Constraints in the Use of ILUPACK on Graphics Processors
abstract
An important number of scientific and engineering problems currently require the solution of large and sparse linear systems of equations. In previous work, we applied a GPU accelerator to the solution of sparse linear systems of moderate dimension via ILUPACK, showing important reductions in the execution time while maintaining the quality of the solution. Unfortunately, the use of GPUs attached to only one compute node strongly limits the memory available to solve the systems, and thus the size of the problems that can be tackled with this approach. In this work we introduce a distributed-parallel version of ILUPACK that overcomes these limitations. The results of the evaluation show that the inclusion of multiple GPUs, located on distinct nodes of a cluster, yields relevant reductions in the execution time for large problems and, more importantly, allows to increase the dimension of the problems, showing interesting scaling properties.
José Ignacio Aliaga, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí
SBAC-PAD1
2017 Communication in task-parallel ILU-preconditioned CG solvers using MPI + OmpSs
abstract
Summary We target the parallel solution of sparse linear systems via iterative Krylov subspace–based methods enhanced with incomplete LU (ILU)‐type preconditioners on clusters of multicore processors. In order to tackle large‐scale problems, we develop task‐parallel implementations of the classical iteration for the CG method, accelerated via ILUPACK and ILU(0) preconditioners, using MPI + OmpSs. In addition, we integrate several communication‐avoiding strategies into the codes, including the butterfly communication scheme and Eijkhout's formulation of the CG method. For all these implementations, we analyze the communication patterns and perform a comparative analysis of their performance and scalability on a cluster consisting of 16 nodes, with 16 cores each.
José Ignacio Aliaga, Maria Barreda, Goran Flegar, Matthias Bollhöfer, Enrique S. Quintana-Ortí
Concurr. Comput. Pract. Exp.1
2017 Adapting concurrency throttling and voltage-frequency scaling for dense eigensolvers
José Ignacio Aliaga, Maria Barreda, M. Asunción Castaño, Manuel F. Dolz, Enrique S. Quintana-Ortí
J. Supercomput.1
2016 Exploiting Task-Parallelism in Message-Passing Sparse Linear System Solvers Using OmpSs
José Ignacio Aliaga, Maria Barreda, Matthias Bollhöfer, Enrique S. Quintana-Ortí
Euro-Par1
2016 Exploiting task and data parallelism in ILUPACK's preconditioned CG solver on NUMA architectures and many-core accelerators
José Ignacio Aliaga, Rosa M. Badia, Maria Barreda, Matthias Bollhöfer, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí
Parallel Comput.1
2015 Systematic Fusion of CUDA Kernels for Iterative Sparse Linear System Solvers
José Ignacio Aliaga, Enrique S. Quintana-Ortí
Euro-Par1
2015 Unveiling the performance-energy trade-off in iterative linear system solvers for multithreaded processors
abstract
Summary In this paper, we analyze the interactions occurring in the triangle performance‐power‐energy for the execution of a pivotal numerical algorithm, the iterative conjugate gradient (CG) method, on a diverse collection of parallel multithreaded architectures. This analysis is especially timely in a decade where the power wall has arisen as a major obstacle to build faster processors. Moreover, the CG method has recently been proposed as a complement to the LINPACK benchmark, as this iterative method is argued to be more archetypical of the performance of today's scientific and engineering applications. To gain insights about the benefits of hands‐on optimizations we include runtime and energy efficiency results for both out‐of‐the‐box usage relying exclusively on compiler optimizations, and implementations manually optimized for target architectures, that range from general‐purpose and digital signal multicore processors to manycore graphics processing units, all representative of current multithreaded systems. Copyright © 2014 John Wiley & Sons, Ltd.
José Ignacio Aliaga, Hartwig Anzt, María Isabel Castillo, Juan Carlos Fernández 0002, German Leon, Enrique S. Quintana-Ortí
Concurr. Comput. Pract. Exp.1
2015 Out-of-core macromolecular simulations on multithreaded architectures
abstract
Summary We address the solution of large‐scale eigenvalue problems that appear in the motion simulation of complex macromolecules on multithreaded platforms, consisting of multicore processors and possibly a graphics processor (graphics processing unit). In particular, we compare specialized implementations of several high‐performance eigensolvers that, by relying on disk storage and out‐of‐core techniques, can in principle tackle the large memory requirements of these biological problems, which in general do not fit into the main memory of current desktop machines. All these out‐of‐core eigensolvers, except for one, are composed of compute‐bound (i.e., arithmetically intensive) operations, which we accelerate by exploiting the performance of current multicore processors and, in some cases, by additionally off‐loading certain parts of the computation to a graphics processing unit accelerator. One of the eigensolvers is a memory‐bound algorithm, which strongly constrains its performance when the data is on disk. However, this method exhibits a much lower arithmetic cost compared with its compute‐bound alternatives for this particular application. Experimental results on a desktop platform, representative of current server technology, illustrate the potential of these methods to address the simulation of biological activity. Copyright © 2014 John Wiley & Sons, Ltd.
José Ignacio Aliaga, José M. Badía, María Isabel Castillo, Davor Davidovic, Rafael Mayo 0002, Enrique S. Quintana-Ortí
Concurr. Comput. Pract. Exp.1
2014 Leveraging Data-Parallelism in ILUPACK using Graphics Processors
abstract
In this paper, we address the exploitation of data parallelism for the solution of sparse symmetric positive definite linear systems via iterative methods on Graphics Processing Units (GPUs). In particular, we accelerate the preconditioned CG-based iterative solver underlying the incomplete LU decomposition package (ILUPACK) by off-loading the most expensive computations i.e., The solution of sparse triangular systems and sparse matrix-vector products-to the hardware accelerator. The results collected using GPUs from the two most recent generations from NVIDIA ("Fermi" and "Kepler") and a benchmark test bed of sparse linear systems show that the GPU-enabled implementations deliver a notable reduction of the execution time, while maintaining the convergence rate and numerical properties of the original ILUPACK solver.
José Ignacio Aliaga, Matthias Bollhöfer, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí
ISPDC1
2014 Leveraging Task-Parallelism with OmpSs in ILUPACK's Preconditioned CG Method
abstract
In this paper we describe how to efficiently exploit task parallelism for the solution of sparse linear systems on multithreaded processors via ILUPACK's multi-level preconditioned CG method. Using a pair of data structures, we capture the task dependencies that appear in the two most challenging operations in the method (calculation of the preconditioned and its application), passing this information to the OmpSs runtime which can then implement a correct and efficient schedule of the entire solver. Our results with high-end multicore platforms equipped with Intel and AMD processors report significant performance gains, demonstrating that OmpSs provides an efficient and close-to seamless means to leverage the concurrency in a complex scientific code like ILUPACK.
José Ignacio Aliaga, Rosa M. Badia, Maria Barreda, Matthias Bollhöfer, Enrique S. Quintana-Ortí
SBAC-PAD1
2013 Reformulated Conjugate Gradient for the Energy-Aware Solution of Linear Systems on GPUs
abstract
In this paper we introduce a redesign of the conjugate gradient method for the iterative solution of sparse linear systems on heterogeneous systems accelerated by graphics processing units (GPUs). Reshaping the GPU kernels induced by the classical formulation of the CG method into algorithm-specific routines results in a slight increase of performance and, more importantly, enables the efficient exploitation of power-saving techniques implicit in the hardware, like the processor C-states, that produce remarkable energy savings. Numerical experiments using data matrices from a popular sparse matrix collection show that the time overhead naturally associated with the application of these energy-aware techniques is no longer crucial to the overall runtime performance.
José Ignacio Aliaga, Enrique S. Quintana-Ortí, Hartwig Anzt
ICPP1
2011 Exploiting thread-level parallelism in the iterative solution of sparse linear systems
José Ignacio Aliaga, Matthias Bollhöfer, Alberto F. Martín, Enrique S. Quintana-Ortí
Parallel Comput.1
2009 Toward the parallelization of GSL
José Ignacio Aliaga, Francisco Almeida, José M. Badía, Sergio Barrachina 0001, Vicente Blanco 0001, María Isabel Castillo, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí, Alfredo Remón, Casiano Rodríguez, Francisco de Sande, Adrián Santos
J. Supercomput.1
2006 Parallelization of GSL: The Web Service Interface
abstract
We present our joint effort to develop a Web based interface for the GNU Scientific library and its parallelization. The interface has been developed using standard Web services technology to enable the use of non local resources to execute parallel programs. The final result is a computing service where sequential and parallel routines demanding high performance computing are supplied. The design allows to incorporate new servers and platforms with a small number of software requirements.
José Ignacio Aliaga, José M. Badía, Sergio Barrachina 0001, María Isabel Castillo, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí, Francisco Almeida, Vicente Blanco 0001, Casiano Rodríguez, Francisco de Sande, Adrián Santos
PDP1