EDBT 2026 Demo / reviewers in the wild / expert
Pedro Alonso 0002
dblp:89/8296 · also Pedro Alonso-Jordá
· DBLP profile ↗
53ranked-venue papers
18as first author
18since 2021 · last 2026
0000-0002-6882-6592ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 43 · 16 first-author · 15 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accelerated border tracking in binary images with GPUsabstractAbstract This work presents an optimized algorithm for contour detection and extraction (i.e., border tracking) in binary images, aiming to improve performance in computer vision scenarios that require real-time processing. The approach divides the image into rectangular blocks, processing each block in parallel to extract “triads” (structures representing three interconnected and ordered points). Subsequently, the triads are connected both within each block and between adjacent blocks to form complete, closed contours. The algorithm is composed of three steps, each implemented as CUDA kernels. The main objective of the proposed algorithm is to avoid costly data transfers between the CPU and GPU, while maintaining performance at a level similar to that of the CPU. This objective is, particularly, beneficial when the algorithm is part of industrial workflows with high efficiency requirements. Pedro Alonso 0002, Roberto Díaz-Cano Lozano, Enrique S. Quintana-Ortí, Francesc Folch |
J. Supercomput. | 1 |
| 2026 | Enhancing transformer performance and portability through auto-tuning frameworksabstractAbstract Transformer-based models such as BERT and GPT2 have become the foundation of many modern applications, yet their execution requires substantial computational and memory resources. To address these challenges, recent advances in compiler technology and hardware accelerators have introduced new opportunities for performance portability. In this work, we evaluate JAX and TVM as high-level frameworks that combine a NumPy-like programming model with Just-In-Time (JIT) or Ahead-of-Time (AOT) code optimization and compilation, enabling efficient execution across CPUs or GPUs, and in the case of JAX, on TPUs as well. We present systematic implementations of the core Transformer encoder and decoder blocks in JAX and TVM and compare their automatically optimized code against NumPy and CuPy baselines. Our experimental study covers heterogeneous hardware platforms (AMD CPU, NVIDIA GPUs, and Google TPUs) and multiple arithmetic precisions (FP32, BF16, INT8, and INT32). Results show that JAX and TVM deliver significant performance improvements over standard libraries, while reducing the programming effort required to adapt to different hardware. These findings demonstrate the potential of JIT- and AOT-oriented frameworks to serve as a portable and efficient solution for deploying Transformer workloads in diverse computing environments. Patricia Siwinska, Jie Lei 0007, Adrián Castelló 0001, Pedro Alonso 0002, Enrique S. Quintana-Ortí |
J. Supercomput. | 4 |
| 2025 | Acceleration of the MVS workflow using graphics processors
Roberto Díaz-Cano Lozano, Francesc Folch, Enrique S. Quintana-Ortí, Pedro Alonso 0002 |
J. Supercomput. | 4 |
| 2025 | The evolution of high-performance computing: how AI and quantum computing are reshaping supercomputing
Sandra Ranilla-Cortina, Pedro Alonso 0002, Jesús Vigo-Aguiar, José Ranilla |
J. Supercomput. | 2 |
| 2024 | Automatic generation of ARM NEON micro-kernels for matrix multiplicationabstractAbstract General matrix multiplication ( gemm ) is a fundamental kernel in scientific computing and current frameworks for deep learning. Modern realisations of gemm are mostly written in C, on top of a small, highly tuned micro-kernel that is usually encoded in assembly. The high performance realisation of gemm in linear algebra libraries in general include a single micro-kernel per architecture, usually implemented by an expert. In this paper, we explore a couple of paths to automatically generate gemm micro-kernels, either using C++ templates with vector intrinsics or high-level Python scripts that directly produce assembly code. Both solutions can integrate high performance software techniques, such as loop unrolling and software pipelining, accommodate any data type, and easily generate micro-kernels of any requested dimension. The performance of this solution is tested on three ARM-based cores and compared with state-of-the-art libraries for these processors: BLIS, OpenBLAS and ArmPL. The experimental results show that the auto-generation approach is highly competitive, mainly due to the possibility of adapting the micro-kernel to the problem dimensions. Guillermo Alaejos, Héctor Martínez 0002, Adrián Castelló 0001, Manuel F. Dolz, Francisco D. Igual, Pedro Alonso 0002, Enrique S. Quintana-Ortí |
J. Supercomput. | 6 |
| 2024 | Algorithm 1039: Automatic Generators for a Family of Matrix Multiplication Routines with Apache TVMabstractWe explore the utilization of the Apache TVM open source framework to automatically generate a family of algorithms that follow the approach taken by popular linear algebra libraries, such as GotoBLAS2, BLIS, and OpenBLAS, to obtain high-performance blocked formulations of the general matrix multiplication ( gemm ). In addition, we fully automatize the generation process by also leveraging the Apache TVM framework to derive a complete variety of the processor-specific micro-kernels for gemm . This is in contrast with the convention in high-performance libraries, which hand-encode a single micro-kernel per architecture using Assembly code. In global, the combination of our TVM-generated blocked algorithms and micro-kernels for gemm (1) improves portability, maintainability, and, globally, streamlines the software life cycle; (2) provides high flexibility to easily tailor and optimize the solution to different data types, processor architectures, and matrix operand shapes, yielding performance on a par (or even superior for specific matrix shapes) with that of hand-tuned libraries; and (3) features a small memory footprint. Guillermo Alaejos, Adrián Castelló 0001, Pedro Alonso 0002, Francisco D. Igual, Héctor Martínez 0002, Enrique S. Quintana-Ortí |
ACM Trans. Math. Softw. | 3 |
| 2023 | Leveraging State-of-the-Art Engines for Large-Scale Data Analysis in High Energy PhysicsabstractAbstract The Large Hadron Collider (LHC) at CERN has generated a vast amount of information from physics events, reaching peaks of TB of data per day which are then sent to large storage facilities. Traditionally, data processing workflows in the High Energy Physics (HEP) field have leveraged grid computing resources. In this context, users have been responsible for manually parallelising the analysis, sending tasks to computing nodes and aggregating the partial results. Analysis environments in this field have had a common building block in the ROOT software framework. This is the de facto standard tool for storing, processing and visualising HEP data. ROOT offers a modern analysis tool called RDataFrame, which can parallelise computations from a single machine to a distributed cluster while hiding most of the scheduling and result aggregation complexity from users. This is currently done by leveraging Apache Spark as the distributed execution engine, but other alternatives are being explored by HEP research groups. Notably, Dask has rapidly gained popularity thanks to its ability to interface with batch queuing systems, widespread in HEP grid computing facilities. Furthermore, future upgrades of the LHC are expected to bring a dramatic increase in data volumes. This paper presents a novel implementation of the Dask backend for the distributed RDataFrame tool in order to address the aforementioned future trends. The scalability of the tool with both the new backend and the already available Spark backend is demonstrated for the first time on more than two thousand cores, testing a real HEP analysis. Vincenzo Eduardo Padulano, Ivan Donchev Kabadzhov, Enric Tejedor, Enrico Guiraud, Pedro Alonso 0002 |
J. Grid Comput. | 5 |
| 2023 | Micro-kernels for portable and efficient matrix multiplication in deep learningabstractAbstract We provide a practical demonstration that it is possible to systematically generate a variety of high-performance micro-kernels for the general matrix multiplication (gemm) via generic templates which can be easily customized to different processor architectures and micro-kernel dimensions. These generic templates employ vector intrinsics to exploit the SIMD (single instruction, multiple data) units in current general-purpose processors and, for the particular type of gemm problems encountered in deep learning, deliver a floating-point throughput rate on par with or even higher than that obtained with conventional, carefully tuned implementations of gemm in current linear algebra libraries (e.g., BLIS, AMD AOCL, ARMPL). Our work exposes the structure of the template-based micro-kernels for ARM Neon (128-bit SIMD), ARM SVE (variable-length SIMD) and Intel AVX512 (512-bit SIMD), showing considerable performance for an NVIDIA Carmel processor (ARM Neon), a Fujitsu A64FX processor (ARM SVE) and on an AMD EPYC 7282 processor (256-bit SIMD). Guillermo Alaejos, Adrián Castelló 0001, Héctor Martínez 0002, Pedro Alonso 0002, Francisco D. Igual, Enrique S. Quintana-Ortí |
J. Supercomput. | 4 |
| 2023 | Efficient and portable Winograd convolutions for multi-core processorsabstractAbstract We take a step forward towards developing high-performance codes for the convolution operator, based on the Winograd algorithm, that are easy to customise for general-purpose processor architectures. In our approach, augmenting the portability of the solution is achieved via the introduction of vector instructions from Intel SSE/AVX2/AVX512 and ARM NEON/SVE to exploit the single-instruction multiple-data capabilities of current processors as well as OpenMP pragmas to exploit multi-threaded parallelism. While this comes at the cost of sacrificing a fraction of the computational performance, our experimental results on three distinct processors, with Intel Xeon Skylake, ARM Cortex A57 and Fujitsu A64FX processors, show that the impact is affordable and still renders a Winograd-based solution that is competitive when compared with the lowering gemm-based convolution. Manuel F. Dolz, Héctor Martínez 0002, Adrián Castelló 0001, Pedro Alonso 0002, Enrique S. Quintana-Ortí |
J. Supercomput. | 4 |
| 2023 | Parallel border tracking in binary images for multicore computersabstractAbstract Border tracking in binary images is an important operation in many computer vision applications. The problem consists in finding borders in a 2D binary image (where all of the pixels are either 0 or 1). There are several algorithms available for this problem, but most of them are sequential. In a former paper, a parallel border tracking algorithm was proposed. This algorithm was designed to run in Graphics Processing units, and it was based on the sequential algorithm known as the Suzuki algorithm. In this paper, we adapt the previously proposed GPU algorithm so that it can be executed in multicore computers. The resulting algorithm is evaluated against its GPU counterpart. The results show that the performance of the GPU algorithm worsens (or even fails) for very large images or images with many borders. On the other hand, the proposed multicore algorithm can efficiently cope with large images. Víctor M. García 0001, Pedro Alonso 0002 |
J. Supercomput. | 2 |
| 2023 | Leveraging an open source serverless framework for high energy physics computingabstractAbstract CERN (Centre Europeen pour la Recherce Nucleaire) is the largest research centre for high energy physics (HEP). It offers unique computational challenges as a result of the large amount of data generated by the large hadron collider. CERN has developed and supports a software called ROOT , which is the de facto standard for HEP data analysis. This framework offers a high-level and easy-to-use interface called RDataFrame , which allows managing and processing large data sets. In recent years, its functionality has been extended to take advantage of distributed computing capabilities. Thanks to its declarative programming model, the user-facing API can be decoupled from the actual execution backend . This decoupling allows physical analysis to scale automatically to thousands of computational cores over various types of distributed resources. In fact, the distributed RDataFrame module already supports the use of established general industry engines such as Apache Spark or Dask. Notwithstanding the foregoing, these current solutions will not be sufficient to meet future requirements in terms of the amount of data that the new projected accelerators will generate. It is of interest, for this reason, to investigate a different approach, the one offered by serverless computing. Based on a first prototype using AWS Lambda , this work presents the creation of a new backend for RDataFrame distributed over the OSCAR tool, an open source framework that supports serverless computing. The implementation introduces new ways, relative to the AWS Lambda -based prototype, to synchronize the work of functions. Vincenzo Eduardo Padulano, Pablo Oliver Cortés, Pedro Alonso 0002, Enric Tejedor, Sebastián Risco, Germán Moltó |
J. Supercomput. | 3 |
| 2023 | Efficient GPU implementation of a Boltzmann-Schrödinger-Poisson solver for the simulation of nanoscale DG MOSFETsabstractAbstract A previous study by Mantas and Vecil (Int J High Perform Comput Appl 34(1): 81–102, 2019) describes an efficient and accurate solver for nanoscale DG MOSFETs through a deterministic Boltzmann-Schrödinger-Poisson model with seven electron–phonon scattering mechanisms on a hybrid parallel CPU/GPU platform. The transport computational phase, i.e. the time integration of the Boltzmann equations, was ported to the GPU using CUDA extensions, but the computation of the system’s eigenstates, i.e. the solution of the Schrödinger-Poisson block, was parallelized only using OpenMP due to its complexity. This work fills the gap by describing a port to GPU for the solver of the Schrödinger-Poisson block. This new proposal implements on GPU a Scheduled Relaxation Jacobi method to solve the sparse linear systems which arise in the 2D Poisson equation. The 1D Schrödinger equation is solved on GPU by adapting a multi-section iteration and the Newton-Raphson algorithm to approximate the energy levels, and the Inverse Power Iterative Method is used to approximate the wave vectors. We want to stress that this solver for the Schrödinger-Poisson block can be thought as a module independent of the transport phase (Boltzmann) and can be used for solvers using different levels of description for the electrons; therefore, it is of particular interest because it can be adapted to other macroscopic, hence faster, solvers for confined devices exploited at industrial level. Francesco Vecil, José Miguel Mantas, Pedro Alonso 0002 |
J. Supercomput. | 3 |
| 2022 | A Serverless Engine for High Energy Physics Distributed AnalysisabstractThe Large Hadron Collider (LHC) at CERN has generated in the last decade an unprecedented volume of data for the High-Energy Physics (HEP) field. Scientific collaborations interested in analysing such data very often require computing power beyond a single machine. This issue has been tackled traditionally by running analyses in distributed environments using stateful, managed batch computing systems. While this approach has been effective so far, current estimates for future computing needs of the field present large scaling challenges. Such a managed approach may not be the only viable way to tackle them and an interesting alternative could be provided by serverless architectures, to enable an even larger scaling potential. This work describes a novel approach to running real HEP scientific applications through a distributed serverless computing engine. The engine is built upon ROOT, a well-established HEP data analysis software, and distributes its computations to a large pool of concurrent executions on Amazon Web Services Lambda Serverless Platform. Thanks to the developed tool, physicists are able to access datasets stored at CERN (also those that are under restricted access policies) and process it on remote infrastructures outside of their typical environment. The analysis of the serverless functions is monitored at runtime to gather performance metrics, both for data- and computation-intensive workloads. Jacek Kusnierz, Vincenzo Eduardo Padulano, Maciej Malawski, Kamil Burkiewicz, Enric Tejedor, Pedro Alonso 0002, Michael Pitt, Valentina Avati |
CCGRID | 6 |
| 2022 | Convolution Operators for Deep Learning Inference on the Fujitsu A64FX ProcessorabstractThe convolution operator is a crucial kernel for many computer vision and signal processing applications that rely on deep learning (DL) technologies. As such, the efficient implementation of this operator has received considerable attention in the past few years for a fair range of processor architectures. In this paper, we follow the technology trend toward integrating long SIMD (single instruction, multiple data) arithmetic units into high performance multicore processors to analyse the benefits of this type of hardware acceleration for latency-constrained DL workloads. For this purpose, we implement and optimise for the Fujitsu processor A64FX, three distinct methods for the calculation of the convolution, namely, the lowering approach, a blocked variant of the direct convolution algorithm, and the Winograd minimal filtering algorithm. Our experimental results include an extensive evaluation of the parallel scalability of these three methods and a comparison of their global performance using three popular DL models and a representative dataset. Manuel F. Dolz, Héctor Martínez 0002, Pedro Alonso 0002, Enrique S. Quintana-Ortí |
SBAC-PAD | 3 |
| 2022 | Parallel border tracking in binary images using GPUsabstractAbstract Border tracking in binary images is an important kernel for many applications. There are very efficient sequential algorithms, most notably, the algorithm proposed by Suzuki et al., which has been implemented for CPUs in well-known libraries. However, under some circumstances, it would be advantageous to perform the border tracking in GPUs as efficiently as possible. In this paper, we propose a parallel version of the Suzuki algorithm that is designed to be executed in GPUs and implemented in CUDA. The proposed algorithm is based on splitting the image into small rectangles. Then, a thread is launched for each rectangle, which tracks the borders in its associated rectangle. The final step is to perform the connection of the borders belonging to several rectangles. The parallel algorithm has been compared with a state-of-the-art sequential CPU version, using two different CPUs and two different GPUs for the evaluation. The computing times obtained show that in these experiments with the GPUs and CPUs that we had available, the proposed parallel algorithm running in the fastest GPU is more than 10 times faster than the sequential CPU routine running in the fastest CPU. Víctor M. García 0001, Pedro Alonso 0002, Ricardo García-Laguía |
J. Supercomput. | 2 |
| 2022 | Parallel signal detection for generalized spatial modulation MIMO systemsabstractAbstract Generalized Spatial Modulation is a recently developed technique that is designed to enhance the efficiency of transmissions in MIMO Systems. However, the procedure for correctly retrieving the sent signal at the receiving end is quite demanding. Specifically, the computation of the maximum likelihood solution is computationally very expensive. In this paper, we propose a parallel method for the computation of the maximum likelihood solution using the parallel computing library OpenMP. The proposed parallel algorithm computes the maximum likelihood solution faster than the sequential version, and substantially reduces the worst-case computing times. Víctor M. García 0001, M. Ángeles Simarro, Francisco-Jose Martínez-Zaldívar, Murilo Boratto, Pedro Alonso 0002, Alberto González 0001 |
J. Supercomput. | 5 |
| 2021 | High Performance and Energy Efficient Integer Matrix Multiplication for Deep LearningabstractWe present a multi-threaded implementation of the matrix multiplication for deep learning on ARM multicore processors. Following standard practice for inference with convolutional neural networks, our GEMM kernel operates with 16-bit integer arithmetic, yielding significant performance acceleration and cutting the memory requirements with respect to IEEE (floating point) single precision by half, allowing the deployment of larger neural network models on low power devices with limited storage capacity. Pau San Juan, Pedro Alonso 0002, Enrique S. Quintana-Ortí |
PDP | 2 |
| 2021 | Low precision matrix multiplication for efficient deep learning in NVIDIA Carmel processors
Pablo San Juan, Rafael Rodríguez-Sánchez 0001, Francisco D. Igual, Pedro Alonso 0002, Enrique S. Quintana-Ortí |
J. Supercomput. | 4 |
| 2020 | High Performance and Portable Convolution Operators for Multicore ProcessorsabstractThe considerable impact of Convolutional Neural Networks on many Artificial Intelligence tasks has led to the development of various high performance algorithms for the convolution operator present in this type of networks. One of these approaches leverages the IM2COL transform followed by a general matrix multiplication (GEMM) in order to take advantage of the highly optimized realizations of the GEMM kernel in many linear algebra libraries. The main problems of this approach are 1) the large memory workspace required to host the intermediate matrices generated by the IM2COL transform; and 2) the time to perform the IM2COL transform, which is not negligible for complex neural networks. This paper presents a portable high performance convolution algorithm based on the BLIS realization of the GEMM kernel that avoids the use of the intermediate memory by taking advantage of the BLIS structure. In addition, the proposed algorithm eliminates the cost of the explicit IM2COL transform, while maintaining the portability and performance of the underlying realization of GEMM in BLIS. Pablo San Juan, Adrián Castelló 0001, Manuel F. Dolz, Pedro Alonso 0002, Enrique S. Quintana-Ortí |
SBAC-PAD | 4 |
| 2020 | Performance modeling of the sparse matrix-vector product via convolutional neural networks
Maria Barreda, Manuel F. Dolz, M. Asunción Castaño, Pedro Alonso 0002, Enrique S. Quintana-Ortí |
J. Supercomput. | 4 |
| 2019 | Computing matrix trigonometric functions with GPUs through Matlab
Pedro Alonso 0002, Jesús Peinado-Pinilla, Jacinto Javier Ibáñez, Jorge Sastre, Emilio Defez |
J. Supercomput. | 1 |
| 2019 | Fast block QR update in digital signal processing
Fran J. Alventosa, Pedro Alonso 0002, Antonio M. Vidal, Gema Piñero, Enrique S. Quintana-Ortí |
J. Supercomput. | 2 |
| 2019 | Exploring hybrid parallel systems for probabilistic record linkage
Murilo Boratto, Pedro Alonso 0002, Clícia Pinto, Pedro Melo, Marcos E. Barreto, Spiros C. Denaxas |
J. Supercomput. | 2 |
| 2019 | HReMAS: hybrid real-time musical alignment system
Pablo Cabañas Molero, Raquel Cortina, Elías F. Combarro, Pedro Alonso 0002, F. J. Bris-Peñalver |
J. Supercomput. | 4 |
| 2019 | A pipeline structure for the block QR update in digital signal processing
Manuel F. Dolz, Fran J. Alventosa, Pedro Alonso 0002, Antonio M. Vidal |
J. Supercomput. | 3 |
| 2019 | Real-time Soundprism
Antonio Jesús Muñoz-Montoro, José Ranilla, Pedro Vera-Candeas, Elías F. Combarro, Pedro Alonso 0002 |
J. Supercomput. | 5 |
| 2018 | Two-sided orthogonal reductions to condensed forms on asymmetric multicore processors
Pedro Alonso 0002, Sandra Catalán, José R. Herrero 0001, Enrique S. Quintana-Ortí, Rafael Rodríguez-Sánchez 0001 |
Parallel Comput. | 1 |
| 2017 | Parallel online time warping for real-time audio-to-score alignment in multi-core systems
Pedro Alonso 0002, Raquel Cortina, Francisco J. Rodríguez-Serrano, Pedro Vera-Candeas, M. Alonso-González, José Ranilla |
J. Supercomput. | 1 |
| 2017 | High-performance computing: the essential tool and the essential challenge
Pedro Alonso 0002, José Ranilla, Jesús Vigo-Aguiar |
J. Supercomput. | 1 |
| 2017 | An efficient musical accompaniment parallel system for mobile devices
Pedro Alonso 0002, Pedro Vera-Candeas, Raquel Cortina, José Ranilla |
J. Supercomput. | 1 |
| 2017 | Accelerating multi-channel filtering of audio signal on ARM processors
Jose A. Belloch, Fran J. Alventosa, Pedro Alonso 0002, Enrique S. Quintana-Ortí, Antonio M. Vidal |
J. Supercomput. | 3 |
| 2017 | Automatic tuning to performance modelling of matrix polynomials on multicore and multi-GPU systems
Murilo Boratto, Pedro Alonso 0002, Domingo Giménez, Alexey L. Lastovetsky |
J. Supercomput. | 2 |
| 2014 | Enhancing performance and energy consumption of runtime schedulers for dense linear algebraabstractSUMMARY The road towards Exascale Computing requires a holistic effort to address three different challenges simultaneously: high performance, energy efficiency, and programmability. The use of runtime task schedulers to orchestrate parallel executions with minimal developer intervention has been introduced in recent years to tackle the programmability issue while maintaining, or even improving, performance. In this paper, we enhance the SuperMatrix runtime task scheduler integrated in the libflame library in two different directions that address high performance and energy efficiency. First, we extend the runtime by accommodating hybrid parallel executions and managing task priorities for dense linear algebra operations, with remarkable performance improvements. Second, we introduce techniques to reduce energy consumption during idle times inherent to parallel executions, attaining important energy savings. In addition, we propose a power consumption model that can be leveraged by runtime task schedulers to make decisions based not only on performance but also on energy considerations. Copyright © 2014 John Wiley & Sons, Ltd. Pedro Alonso 0002, Manuel F. Dolz, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
Concurr. Comput. Pract. Exp. | 1 |
| 2014 | Modeling power and energy consumption of dense matrix factorizations on multicore processorsabstractSUMMARY In this paper, we propose a model for the energy consumption of the concurrent execution of three key dense matrix factorizations, with task parallelism leveraged via the Symmetric Multi‐Processing Superscalar (SMPSs) runtime, on a multicore processor. Our model decomposes the power dissipation into the system, static and dynamic components, with the former two being estimated from basic, off‐line experiments. The dynamic power, on the other hand, requires significantly more care, and we introduce a contention‐aware model that accommodates for the variability of power consumption due to memory contention. Experimental results on an Intel Xeon E5504 processor with four cores, using an internal powermeter that samples the power drawn by the mainboard with a frequency of 1 KHz, show the reliability of the energy model for the Cholesky, LU, and QR factorizations on this platform. Copyright © 2013 John Wiley & Sons, Ltd. Pedro Alonso 0002, Manuel F. Dolz, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
Concurr. Comput. Pract. Exp. | 1 |
| 2014 | Block pivoting implementation of a symmetric Toeplitz solver
Pedro Alonso 0002, Manuel F. Dolz, Antonio M. Vidal |
J. Parallel Distributed Comput. | 1 |
| 2014 | Automatic routine tuning to represent landform attributes on multicore and multi-GPU systems
Murilo Boratto, Pedro Alonso 0002, Domingo Giménez, Marcos E. Barreto |
J. Supercomput. | 2 |
| 2014 | Solving time-invariant differential matrix Riccati equations using GPGPU computing
Jesús Peinado-Pinilla, Pedro Alonso 0002, Jacinto Javier Ibáñez, Vicente Hernández, Murilo Boratto |
J. Supercomput. | 2 |
| 2013 | A multicore solution to Block-Toeplitz linear systems of equations
Pedro Alonso 0002, Daniel Argüelles, José Ranilla, Antonio M. Vidal |
J. Supercomput. | 1 |
| 2012 | Parallel Algorithm for Landform Attributes Representation on Multicore and Multi-GPU Systems
Murilo Boratto, Pedro Alonso 0002, Carla Ramiro, Marcos E. Barreto, Leandro dos Santos Coelho |
ICCSA (1) | 2 |
| 2012 | Tools for Power-Energy Modelling and Analysis of Parallel Scientific ApplicationsabstractUnderstanding power usage in parallel workloads is crucial to develop the energy-aware software that will run in future Exascale systems. In this paper, we contribute towards this goal by introducing an integrated framework to profile, monitor, model and analyze power dissipation in parallel MPI and multi-threaded scientific applications. The framework includes an own-designed device to measure internal DC power consumption and a package offering a simple interface to interact with this design as well as commercial power meters. Combined with the instrumentation package Extrae and the graphical analysis tool Paraver, the result is a useful environment to identify sources of power inefficiency directly in the source application code. For task-parallel codes, we also offer a statistical software module that inspects the execution trace of the application to calculate the parameters of an accurate model for the global energy consumption, which can be then decomposed into the average power usage per task or the nodal power dissipated per core. Pedro Alonso 0002, Rosa M. Badia, Jesús Labarta, Maria Barreda, Manuel F. Dolz, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Ruymán Reyes |
ICPP | 1 |
| 2012 | Reducing Energy Consumption of Dense Linear Algebra Operations on Hybrid CPU-GPU PlatformsabstractWe investigate the balance between the time-to-solution and the energy consumption of a task-parallel execution of the Cholesky and LU factorizations on a hybrid platform, equipped with a multi-core processor and several GPUs. To improve energy efficiency, we incorporate two energy-saving techniques in the runtime in charge of scheduling the computations, to block idle threads and enable the transition to a more energy-friendly state of the general-purpose cores. Experiments on an Intel Xeon-based platform connected to an NVIDIA Tesla server report an average reduction of the energy consumption close to 9% (38% when only the consumption associated with the application is considered), for a minor increase in the execution time of the algorithm. Pedro Alonso 0002, Manuel F. Dolz, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
ISPA | 1 |
| 2012 | Saving Energy in the LU Factorization with Partial Pivoting on Multi-core ProcessorsabstractIn this paper we analyze the trade-off between energy and performance for a data-parallel execution of the LU factorization with partial pivoting on a multi-core processor. To improve energy efficiency, we adapt the runtime in charge of controlling the concurrent execution of the algorithm to leverage DVFS and block idle threads. For a CPU-bounded operation like the LU factorization, experiments on an AMD 8-core processor report a reduction around 5% in energy consumption for the largest problem sizes in exchange for a minor increase in the execution time. Pedro Alonso 0002, Manuel F. Dolz, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
PDP | 1 |
| 2011 | Implementation and tuning of a parallel symmetric Toeplitz eigensolver
Pedro Alonso 0002, Miguel O. Bernabeu, Víctor M. García 0001, Antonio M. Vidal |
J. Parallel Distributed Comput. | 1 |
| 2010 | Experimental Study of Six Different Implementations of Parallel Matrix Multiplication on Heterogeneous Computational Clusters of Multicore ProcessorsabstractTwo strategies of distribution of computations can be used to implement parallel solvers for dense linear algebra problems for Heterogeneous Computational Clusters of Multicore Processors (HCoMs). These strategies are called Heterogeneous Process Distribution Strategy (HPS) and Heterogeneous Data Distribution Strategy (HDS). They are not novel and have been researched thoroughly. However, the advent of multicores necessitates enhancements to them. In this paper, we present these enhancements. Our study is based on experiments using six applications to perform Parallel Matrix-matrix Multiplication (PMM) on an HCoM employing the two distribution strategies. Pedro Alonso 0002, Ravi Reddy, Alexey L. Lastovetsky |
PDP | 1 |
| 2009 | Parallel solvers for dense linear systems for heterogeneous computational clustersabstractThis paper describes the design and the implementation of parallel routines in the heterogeneous ScaLAPACK library that solve a dense system of linear equations. This library is written on top of HeteroMPI and ScaLAPACK whose building blocks, the de facto standard kernels for matrix and vector operations (BLAS and its parallel counterpart PBLAS) and message passing communication (BLACS), are optimized for heterogeneous computational clusters. We show that the efficiency of these parallel routines is due to the most important feature of the library, which is the automation of the difficult optimization tasks of parallel programming on heterogeneous computing clusters. They are the determination of the accurate values of the platform parameters such as the speeds of the processors and the latencies and bandwidths of the communication links connecting different pairs of processors, the optimal values of the algorithmic parameters such as the total number of processes, the 2D process grid arrangement and the efficient mapping of the processes executing the parallel algorithm to the executing nodes of the heterogeneous computing cluster. We describe this process of automation followed by presentation of experimental results on a local heterogeneous computing cluster demonstrating the efficiency of these solvers. Ravi Reddy, Alexey L. Lastovetsky, Pedro Alonso 0002 |
IPDPS | 3 |
| 2008 | Scalable Dense Factorizations for Heterogeneous Computational ClustersabstractThis paper discusses the design and the implementation of the LU factorization routines included in the Heterogeneous ScaLAPACK library, which is built on top of ScaLAPACK. These routines are used in the factorization and solution of a dense system of linear equations. They are implemented using optimized PBLAS, BLACS and BLAS libraries for heterogeneous computational clusters. We present the details of the implementation as well asperformance results on a heterogeneous computingcluster. Ravi Reddy, Alexey L. Lastovetsky, Pedro Alonso 0002 |
ISPDC | 3 |
| 2008 | Heterogeneous PBLAS: Optimization of PBLAS for Heterogeneous Computational ClustersabstractThis paper presents a package, called Heterogeneous PBLAS (HeteroPBLAS), which is built on top of PBLAS and provides optimized parallel basic linear algebra subprograms for heterogeneous computational clusters. We present the user interface and the software hierarchy of the first research implementation of HeteroPBLAS. This is the first step towards the development of a parallel linear algebra package for heterogeneous computational clusters. We demonstrate the efficiency of the HeteroPBLAS programs on a homogeneous computing cluster and a heterogeneous computing cluster. Ravi Reddy, Alexey L. Lastovetsky, Pedro Alonso 0002 |
ISPDC | 3 |
| 2008 | A Threaded Divide and Conquer Symmetric Tridiagonal Eigensolver on Multicore SystemsabstractThe increasing power of computation of modern processors rely on the increasing number of cores per chip. The challenge of software developers is to keep this power with the legacy code. Although commercial and non commercial libraries are improving their codes step by step, there exits probably insurmountable scalability issues for standard programming models due to the fact that using locks to implement synchronisation is inherently a bottleneck. We propose an implementation of the divide and conquer algorithm to compute the eigenpairs of symmetric tridiagonal matrices on multicore systems. We take advantage of the natural parallelism of the method by using pthreads. We avoided as much as possible the negative impact of synchronisation in the performance by overlapping operations of different classes. Furthermore, the unevenly workload distribution of the computational cost of the elemental tasks yields in a speedup even larger than expected. Antonio M. Vidal, Murilo Boratto, Pedro Alonso 0002 |
ISPDC | 3 |
| 2008 | Parallel computation of the eigenvalues of symmetric Toeplitz matrices through iterative methods
Antonio M. Vidal, Víctor M. García 0001, Pedro Alonso 0002, Miguel O. Bernabeu |
J. Parallel Distributed Comput. | 3 |
| 2008 | A multilevel parallel algorithm to solve symmetric Toeplitz linear systems
Miguel O. Bernabeu, Pedro Alonso 0002, Antonio M. Vidal |
J. Supercomput. | 2 |
| 2006 | A Parallel Algorithm for the Solution of the Deconvolution Problem on Heterogeneous NetworksabstractIn this work we present a parallel algorithm for the solution of a least squares problem with structured matrices. This problem arises in many applications mainly related to digital signal processing. The parallel algorithm is designed to speed up the sequential one on heterogeneous networks of computers. The parallel algorithm follows the HeHo strategy (Heterogeneous distribution of processes over processors with homogeneous distribution of computations over the processes) and is implemented using HeteroMPI, a recently developed extension of MPI for programming high performance computations on heterogeneous networks of computers. The obtained results validate HeteroMPI as a very useful tool for portable implementation of parallel algorithms for heterogeneous environments Pedro Alonso 0002, Antonio M. Vidal, Alexey L. Lastovetsky |
CLUSTER | 1 |
| 2005 | Solving the block-Toeplitz least-squares problem in parallelabstractIn this paper we present two versions of a parallel algorithm to solve the block–Toeplitz least-squares problem on distributed-memory architectures. We derive a parallel algorithm based on the seminormal equations arising from the triangular decomposition of the product T TT . Our parallel algorithm exploits the displacement structure of the Toeplitz-likematrices using theGeneralized SchurAlgorithm to obtain the solution in O(mn) flops instead of O(mn2) flops of the algorithms for non-structured matrices. The strong regularity of the previous product of matrices and an appropriate computation of the hyperbolic rotations improve the stability of the algorithms. We have reduced the communication cost of previous versions, and have also reduced the memory access cost by appropriately arranging the elements of the matrices. Furthermore, the second version of the algorithm has a very low spatial cost, because it does not store the triangular factor of the decomposition. The experimental results show a good scalability of the parallel algorithm on two different clusters of personal computers. Copyright c © 2005 John Wiley & Sons, Ltd. Pedro Alonso 0002, José M. Badía, Antonio M. Vidal |
Concurr. Pract. Exp. | 1 |
| 2005 | An Efficient Parallel Algorithm to Solve Block-Toeplitz Systems
Pedro Alonso 0002, José M. Badía, Antonio M. Vidal |
J. Supercomput. | 1 |