EDBT 2026 Demo / reviewers in the wild / expert
Javier Cuenca 0001
dblp:14/4433-1
· DBLP profile ↗
18ranked-venue papers
7as first author
3since 2021 · last 2026
0000-0002-8763-756XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 5 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards a hierarchical approach for autotuning task-based librariesabstractAbstract This work proposes a hierarchical approach to reduce the training time of task-based routines by reusing previously obtained autotuning information. This approach has been integrated into a working prototype of Chameleon, a dense linear algebra software whose tile-based routines are executed on the available computational resources by means of a runtime system. The results show that this approach provides a high degree of scalability to the entire self-optimization process, achieving a reduction in training time of up to 80% and an appropriate selection of values for the adjustable parameters. Jesús Cámara, Javier Cuenca 0001, Murilo Boratto |
J. Supercomput. | 2 |
| 2025 | An autotuning approach to select the inter-GPU communication library on heterogeneous systemsabstractAbstract In this work, an automatic optimisation approach for parallel routines on multi-GPU systems is presented. Several inter-GPU communication libraries (such as CUDA-Aware MPI or NCCL) are used with a set of routines to perform the numerical operations among the GPUs located on the compute nodes. The main objective is the selection of the most appropriate communication library, the number of GPUs to be used and the workload to be distributed among them in order to reduce the cost of data movements, which represent a large percentage of the total execution time. To this end, a hierarchical modelling of the execution time of each routine to be optimised is proposed, combining experimental and theoretical approaches. The results show that near-optimal decisions are taken in all the scenarios analysed. Jesús Cámara, Javier Cuenca 0001, Victor Galindo, Arturo Vicente, Murilo Boratto |
J. Supercomput. | 2 |
| 2022 | PARCSIM: a parallel computing simulator for scalable software optimizationabstractAbstract PARCSIM is a parallel software simulator that allows a user to capture, through a graphical interface, matrix algorithm schemes that solve scientific problems. With this tool, the user can analyse the execution times that would be obtained by using different spatio-temporal mapping of computational tasks on available computational units, parallelism parameters and computational libraries. Furthermore, for complex problem models, the self-optimization engine incorporated in this tool analyses the huge tree of possible calculations grouping and mapping strategies in search of the choice that makes the best use of the available hardware resources. This tool also offers polyalgorithmic resolution by making automatically the best decision between different software approaches to solve a given problem on the hardware system available. This work shows the usefulness of this simulator to efficiently solve hierarchical problems constructed from previously modelled subproblems. This task is performed by reusing, in a scalable way, the optimization information of these subproblems to establish the best execution configuration for the composite problem. Jesús Cámara, José-Carlos Cano, Javier Cuenca 0001, Mariano Saura-Sánchez |
J. Supercomput. | 3 |
| 2020 | Integrating software and hardware hierarchies in an autotuning method for parallel routines in heterogeneous clusters
Jesús Cámara, Javier Cuenca 0001, Domingo Giménez |
J. Supercomput. | 2 |
| 2019 | A self-optimized software tool for quantifying the degree of left ventricle hyper-trabeculation
Gregorio Bernabé, José D. Casanova, Javier Cuenca 0001, Josefa González-Carrillo |
J. Supercomput. | 3 |
| 2019 | A parallel simulator for multibody systems based on group equations
José-Carlos Cano, Javier Cuenca 0001, Domingo Giménez, Mariano Saura-Sánchez, Pablo Segado-Cabezos |
J. Supercomput. | 2 |
| 2017 | Guided installation of basic linear algebra routines in a cluster with manycore componentsabstractSummary Computational systems are nowadays composed of basic computational components that share multiprocessors and coprocessors of different types, typically several graphics processing units (GPUs) or many integrated cores (MICs), and those computational components are combined in heterogeneous clusters of nodes with different characteristics, including coprocessors of different types, with varying numbers of nodes at different speeds. The software previously developed and optimized for simpler system needs to be redesigned and reoptimized for these new, more complex systems. The adaptation to hybrid multicore + multiGPU and multicore + multiMIC of autotuning techniques for basic linear algebra routines is analyzed. The matrix‐matrix multiplication kernel, which is optimized for different computational system components through guided experimentation, is studied. The routine is installed for each node in the cluster, and the information generated from individual installations may be used for a hierarchical installation in a cluster. The basic matrix‐matrix multiplication may, in turn, be used inside higher level routines, which delegate their efficient execution to the optimization of the lower level routine. Experimental results are satisfactory in different multicore + multiGPU and multicore + multiMIC systems. So the guided search of execution configurations for satisfactory execution times proves to be a useful tool for heterogeneous systems, where the complexity of the system means a correct use of highly efficient routines and libraries is difficult. Javier Cuenca 0001, Luis-Pedro García, Domingo Giménez, Francisco-José Herrera |
Concurr. Comput. Pract. Exp. | 1 |
| 2014 | Auto-tuned nested parallelism: A way to reduce the execution time of scientific software in NUMA systems
Jesús Cámara, Javier Cuenca 0001, Luis-Pedro García, Domingo Giménez |
Parallel Comput. | 2 |
| 2014 | Improving an autotuning engine for 3D Fast Wavelet Transform on manycore systems
Gregorio Bernabé, Javier Cuenca 0001, Luis-Pedro García, Domingo Giménez |
J. Supercomput. | 2 |
| 2013 | Optimizing a 3D-FWT Code in a Heterogeneous Cluster of Multicore CPUs and Manycore GPUsabstractClusters of nodes composed of many core GPUs and multicore CPUs are used to solve scientific problems with high computational requirements. The development and optimization of parallel-heterogeneous codes for these systems is a complex task which requires a deep knowledge of the different components of the hybrid, heterogeneous and hierarchical computational system, and also of the scientific problem to be solved and the different programing paradigms to be used for its efficient solution. Techniques for efficient development and optimization of scientific codes for these systems are needed. This paper presents an analysis of the development and optimization of the 3D-Fast Wavelet Transform (3D-FWT) for a heterogeneous cluster of multicores+GPUs. Different parallel programming paradigms (message passing, shared memory and SIMD GPU) are combined to fully exploit the computing capacity of the different computational elements of the cluster, so resulting in an efficient combination of basic codes developed previously for individual components (individual nodes, multicore or GPU) and an important reduction of the compression time of long video sequences. Gregorio Bernabé, Javier Cuenca 0001, Domingo Giménez |
SBAC-PAD | 2 |
| 2012 | Empirical Autotuning of Two-level Parallel Linear Algebra Routines on Large cc-NUMA SystemsabstractIn large cc-NUMA systems the efficient use of the different levels of the memory hierarchy is not an easy task, and the performance of multithreading implementations of the libraries decreases when the number of cores used increases, so producing an important lost of efficiency. To alleviate this problem, routines with multilevel parallelism can be developed by combining OpenMP and BLAS parallelism. In that way, higher performance can be achieved, but it is necessary to develop some autotuning technique for the appropriate selection of the number of threads to use at each level. The selection can be made through theoretical models of the execution time or some installation methodology. This work analyses some installation techniques for a two-level matrix multiplication routine, with the aim of developing a valid methodology for other linear algebra routines in large cc-NUMA systems. The basic ideas of the two-level parallelisation and the installation methodology are discussed and some experimental results are commented on. Jesús Cámara, Javier Cuenca 0001, Domingo Giménez, Antonio M. Vidal |
ISPA | 2 |
| 2012 | Improving Linear Algebra Computation on NUMA Platforms through Auto-tuned Nested ParallelismabstractThe most computationally demanding scientific and engineering problems are solved with large parallel systems. In some cases those systems are Non-Uniform Memory Access multiprocessors made up of a large number of cores which share a hierarchically organized memory. Basic linear algebra routines of the type of BLAS typically constitute the kernel of the computation for those problems, and the efficient use of these routines in those systems would contribute to a faster solution of a large range of scientific problems. Normally some multithreaded BLAS library optimized for the system is used, but when the number of cores increases the degradation in the performance is significant, and this can produce a misuse of the large, expensive systems. This paper empirically analyses the behaviour in large NUMA systems of the matrix multiplication of the BLAS library, and its combination with OpenMP to obtain nested parallelism. With the auto-tuning method proposed in this work, a reduction in the execution time is achieved with respect to the matrix multiplication of the library. Javier Cuenca 0001, Luis-Pedro García, Domingo Giménez |
PDP | 1 |
| 2012 | A framework for the application of metaheuristics to tasks-to-processors assignation problems
Francisco Almeida, Javier Cuenca 0001, Domingo Giménez, Antonio Llanes, Juan-Pedro Martínez-Gallar |
J. Supercomput. | 2 |
| 2010 | Analysis of the Influence of the Compiler on Multicore PerformanceabstractThe possibility of connecting several nodes in a network of processors has popularized parallel programming in the scientific community, but its use has been limited by the difficulty of message-passing programming. With the arrival of multicore processors, parallel programming has regained popularity. The use of an OpenMP compiler optimized for the multicore system in question is a good option, but it is possible to have access in a system to more than one compiler and different compilers can appropriately optimize different parts of the code. In this paper we study theoretically and experimentally the influence of the compiler on performance of routines. We conclude that a poly-compiling approach that decides the best compiler for each situation is necessary. Javier Cuenca 0001, Luis-Pedro García, Domingo Giménez, Manuel Quesada-Martínez |
PDP | 1 |
| 2007 | A proposal of metaheuristics to schedule independent tasks in heterogeneous memory-constrained systemsabstractThis paper proposes some metaheuristics for the solution of a tasks scheduling problem. Independent tasks with different computational costs and memory requirements are scheduled in a heterogeneous system with computational and communication heterogeneity and memory constraints. Versions of the scheduling problem with and without communications and with constant and variable computation and communication costs are considered. Javier Cuenca 0001, Domingo Giménez, Jose-Juan López-Espín, Juan-Pedro Martínez-Gallar |
CLUSTER | 1 |
| 2005 | Processes Distribution of Homogeneous Parallel Linear Algebra Routines on Heterogeneous ClustersabstractThis paper presents a self-optimization methodology for parallel linear algebra routines on heterogeneous systems. For each routine, a series of decisions is taken automatically in order to obtain an execution time close to the optimum (without rewriting the routine's code). Some of these decisions are: the number of processes to generate, the heterogeneous distribution of these processes over the network of processors, the logical topology of the generated processes, ... To reduce the search space of such decisions, different heuristics have been used. The experiments have been performed with a parallel LU factorization routine similar to the ScaLAPACK one, and good results have been obtained on different heterogeneous platforms. Javier Cuenca 0001, Luis-Pedro García, Domingo Giménez, Jack J. Dongarra |
CLUSTER | 1 |
| 2005 | Heuristics for work distribution of a homogeneous parallel dynamic programming scheme on heterogeneous systems
Javier Cuenca 0001, Domingo Giménez, Juan-Pedro Martínez-Gallar |
Parallel Comput. | 1 |
| 2004 | Architecture of an automatically tuned linear algebra library
Javier Cuenca 0001, Domingo Giménez, José González 0002 |
Parallel Comput. | 1 |