Jesús Cámara

dblp:119/4430 · DBLP profile ↗
← Back
6ranked-venue papers
6as first author
3since 2021 · last 2026
0000-0003-2176-8273ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 6 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Towards a hierarchical approach for autotuning task-based libraries
abstract
Abstract This work proposes a hierarchical approach to reduce the training time of task-based routines by reusing previously obtained autotuning information. This approach has been integrated into a working prototype of Chameleon, a dense linear algebra software whose tile-based routines are executed on the available computational resources by means of a runtime system. The results show that this approach provides a high degree of scalability to the entire self-optimization process, achieving a reduction in training time of up to 80% and an appropriate selection of values for the adjustable parameters.
Jesús Cámara, Javier Cuenca 0001, Murilo Boratto
J. Supercomput.1
2025 An autotuning approach to select the inter-GPU communication library on heterogeneous systems
abstract
Abstract In this work, an automatic optimisation approach for parallel routines on multi-GPU systems is presented. Several inter-GPU communication libraries (such as CUDA-Aware MPI or NCCL) are used with a set of routines to perform the numerical operations among the GPUs located on the compute nodes. The main objective is the selection of the most appropriate communication library, the number of GPUs to be used and the workload to be distributed among them in order to reduce the cost of data movements, which represent a large percentage of the total execution time. To this end, a hierarchical modelling of the execution time of each routine to be optimised is proposed, combining experimental and theoretical approaches. The results show that near-optimal decisions are taken in all the scenarios analysed.
Jesús Cámara, Javier Cuenca 0001, Victor Galindo, Arturo Vicente, Murilo Boratto
J. Supercomput.1
2022 PARCSIM: a parallel computing simulator for scalable software optimization
abstract
Abstract PARCSIM is a parallel software simulator that allows a user to capture, through a graphical interface, matrix algorithm schemes that solve scientific problems. With this tool, the user can analyse the execution times that would be obtained by using different spatio-temporal mapping of computational tasks on available computational units, parallelism parameters and computational libraries. Furthermore, for complex problem models, the self-optimization engine incorporated in this tool analyses the huge tree of possible calculations grouping and mapping strategies in search of the choice that makes the best use of the available hardware resources. This tool also offers polyalgorithmic resolution by making automatically the best decision between different software approaches to solve a given problem on the hardware system available. This work shows the usefulness of this simulator to efficiently solve hierarchical problems constructed from previously modelled subproblems. This task is performed by reusing, in a scalable way, the optimization information of these subproblems to establish the best execution configuration for the composite problem.
Jesús Cámara, José-Carlos Cano, Javier Cuenca 0001, Mariano Saura-Sánchez
J. Supercomput.1
2020 Integrating software and hardware hierarchies in an autotuning method for parallel routines in heterogeneous clusters
Jesús Cámara, Javier Cuenca 0001, Domingo Giménez
J. Supercomput.1
2014 Auto-tuned nested parallelism: A way to reduce the execution time of scientific software in NUMA systems
Jesús Cámara, Javier Cuenca 0001, Luis-Pedro García, Domingo Giménez
Parallel Comput.1
2012 Empirical Autotuning of Two-level Parallel Linear Algebra Routines on Large cc-NUMA Systems
abstract
In large cc-NUMA systems the efficient use of the different levels of the memory hierarchy is not an easy task, and the performance of multithreading implementations of the libraries decreases when the number of cores used increases, so producing an important lost of efficiency. To alleviate this problem, routines with multilevel parallelism can be developed by combining OpenMP and BLAS parallelism. In that way, higher performance can be achieved, but it is necessary to develop some autotuning technique for the appropriate selection of the number of threads to use at each level. The selection can be made through theoretical models of the execution time or some installation methodology. This work analyses some installation techniques for a two-level matrix multiplication routine, with the aim of developing a valid methodology for other linear algebra routines in large cc-NUMA systems. The basic ideas of the two-level parallelisation and the installation methodology are discussed and some experimental results are commented on.
Jesús Cámara, Javier Cuenca 0001, Domingo Giménez, Antonio M. Vidal
ISPA1