Pedro Alonso 0002

dblp:89/8296 · also Pedro Alonso-Jordá · DBLP profile ↗
← Back
53ranked-venue papers
18as first author
18since 2021 · last 2026
0000-0002-6882-6592ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 43 · 16 first-author · 15 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Accelerated border tracking in binary images with GPUs
abstract
Abstract This work presents an optimized algorithm for contour detection and extraction (i.e., border tracking) in binary images, aiming to improve performance in computer vision scenarios that require real-time processing. The approach divides the image into rectangular blocks, processing each block in parallel to extract “triads” (structures representing three interconnected and ordered points). Subsequently, the triads are connected both within each block and between adjacent blocks to form complete, closed contours. The algorithm is composed of three steps, each implemented as CUDA kernels. The main objective of the proposed algorithm is to avoid costly data transfers between the CPU and GPU, while maintaining performance at a level similar to that of the CPU. This objective is, particularly, beneficial when the algorithm is part of industrial workflows with high efficiency requirements.
Pedro Alonso 0002, Roberto Díaz-Cano Lozano, Enrique S. Quintana-Ortí, Francesc Folch
J. Supercomput.1
2026 Enhancing transformer performance and portability through auto-tuning frameworks
abstract
Abstract Transformer-based models such as BERT and GPT2 have become the foundation of many modern applications, yet their execution requires substantial computational and memory resources. To address these challenges, recent advances in compiler technology and hardware accelerators have introduced new opportunities for performance portability. In this work, we evaluate JAX and TVM as high-level frameworks that combine a NumPy-like programming model with Just-In-Time (JIT) or Ahead-of-Time (AOT) code optimization and compilation, enabling efficient execution across CPUs or GPUs, and in the case of JAX, on TPUs as well. We present systematic implementations of the core Transformer encoder and decoder blocks in JAX and TVM and compare their automatically optimized code against NumPy and CuPy baselines. Our experimental study covers heterogeneous hardware platforms (AMD CPU, NVIDIA GPUs, and Google TPUs) and multiple arithmetic precisions (FP32, BF16, INT8, and INT32). Results show that JAX and TVM deliver significant performance improvements over standard libraries, while reducing the programming effort required to adapt to different hardware. These findings demonstrate the potential of JIT- and AOT-oriented frameworks to serve as a portable and efficient solution for deploying Transformer workloads in diverse computing environments.
Patricia Siwinska, Jie Lei 0007, Adrián Castelló 0001, Pedro Alonso 0002, Enrique S. Quintana-Ortí
J. Supercomput.4
2025 Acceleration of the MVS workflow using graphics processors
Roberto Díaz-Cano Lozano, Francesc Folch, Enrique S. Quintana-Ortí, Pedro Alonso 0002
J. Supercomput.4
2025 The evolution of high-performance computing: how AI and quantum computing are reshaping supercomputing
Sandra Ranilla-Cortina, Pedro Alonso 0002, Jesús Vigo-Aguiar, José Ranilla
J. Supercomput.2
2024 Automatic generation of ARM NEON micro-kernels for matrix multiplication
abstract
Abstract General matrix multiplication ( gemm ) is a fundamental kernel in scientific computing and current frameworks for deep learning. Modern realisations of gemm are mostly written in C, on top of a small, highly tuned micro-kernel that is usually encoded in assembly. The high performance realisation of gemm in linear algebra libraries in general include a single micro-kernel per architecture, usually implemented by an expert. In this paper, we explore a couple of paths to automatically generate gemm micro-kernels, either using C++ templates with vector intrinsics or high-level Python scripts that directly produce assembly code. Both solutions can integrate high performance software techniques, such as loop unrolling and software pipelining, accommodate any data type, and easily generate micro-kernels of any requested dimension. The performance of this solution is tested on three ARM-based cores and compared with state-of-the-art libraries for these processors: BLIS, OpenBLAS and ArmPL. The experimental results show that the auto-generation approach is highly competitive, mainly due to the possibility of adapting the micro-kernel to the problem dimensions.
Guillermo Alaejos, Héctor Martínez 0002, Adrián Castelló 0001, Manuel F. Dolz, Francisco D. Igual, Pedro Alonso 0002, Enrique S. Quintana-Ortí
J. Supercomput.6
2024 Algorithm 1039: Automatic Generators for a Family of Matrix Multiplication Routines with Apache TVM
abstract
We explore the utilization of the Apache TVM open source framework to automatically generate a family of algorithms that follow the approach taken by popular linear algebra libraries, such as GotoBLAS2, BLIS, and OpenBLAS, to obtain high-performance blocked formulations of the general matrix multiplication ( gemm ). In addition, we fully automatize the generation process by also leveraging the Apache TVM framework to derive a complete variety of the processor-specific micro-kernels for gemm . This is in contrast with the convention in high-performance libraries, which hand-encode a single micro-kernel per architecture using Assembly code. In global, the combination of our TVM-generated blocked algorithms and micro-kernels for gemm (1) improves portability, maintainability, and, globally, streamlines the software life cycle; (2) provides high flexibility to easily tailor and optimize the solution to different data types, processor architectures, and matrix operand shapes, yielding performance on a par (or even superior for specific matrix shapes) with that of hand-tuned libraries; and (3) features a small memory footprint.
Guillermo Alaejos, Adrián Castelló 0001, Pedro Alonso 0002, Francisco D. Igual, Héctor Martínez 0002, Enrique S. Quintana-Ortí
ACM Trans. Math. Softw.3
2023 Leveraging State-of-the-Art Engines for Large-Scale Data Analysis in High Energy Physics
abstract
Abstract The Large Hadron Collider (LHC) at CERN has generated a vast amount of information from physics events, reaching peaks of TB of data per day which are then sent to large storage facilities. Traditionally, data processing workflows in the High Energy Physics (HEP) field have leveraged grid computing resources. In this context, users have been responsible for manually parallelising the analysis, sending tasks to computing nodes and aggregating the partial results. Analysis environments in this field have had a common building block in the ROOT software framework. This is the de facto standard tool for storing, processing and visualising HEP data. ROOT offers a modern analysis tool called RDataFrame, which can parallelise computations from a single machine to a distributed cluster while hiding most of the scheduling and result aggregation complexity from users. This is currently done by leveraging Apache Spark as the distributed execution engine, but other alternatives are being explored by HEP research groups. Notably, Dask has rapidly gained popularity thanks to its ability to interface with batch queuing systems, widespread in HEP grid computing facilities. Furthermore, future upgrades of the LHC are expected to bring a dramatic increase in data volumes. This paper presents a novel implementation of the Dask backend for the distributed RDataFrame tool in order to address the aforementioned future trends. The scalability of the tool with both the new backend and the already available Spark backend is demonstrated for the first time on more than two thousand cores, testing a real HEP analysis.
Vincenzo Eduardo Padulano, Ivan Donchev Kabadzhov, Enric Tejedor, Enrico Guiraud, Pedro Alonso 0002
J. Grid Comput.5
2023 Micro-kernels for portable and efficient matrix multiplication in deep learning
abstract
Abstract We provide a practical demonstration that it is possible to systematically generate a variety of high-performance micro-kernels for the general matrix multiplication (gemm) via generic templates which can be easily customized to different processor architectures and micro-kernel dimensions. These generic templates employ vector intrinsics to exploit the SIMD (single instruction, multiple data) units in current general-purpose processors and, for the particular type of gemm problems encountered in deep learning, deliver a floating-point throughput rate on par with or even higher than that obtained with conventional, carefully tuned implementations of gemm in current linear algebra libraries (e.g., BLIS, AMD AOCL, ARMPL). Our work exposes the structure of the template-based micro-kernels for ARM Neon (128-bit SIMD), ARM SVE (variable-length SIMD) and Intel AVX512 (512-bit SIMD), showing considerable performance for an NVIDIA Carmel processor (ARM Neon), a Fujitsu A64FX processor (ARM SVE) and on an AMD EPYC 7282 processor (256-bit SIMD).
Guillermo Alaejos, Adrián Castelló 0001, Héctor Martínez 0002, Pedro Alonso 0002, Francisco D. Igual, Enrique S. Quintana-Ortí
J. Supercomput.4
2023 Efficient and portable Winograd convolutions for multi-core processors
abstract
Abstract We take a step forward towards developing high-performance codes for the convolution operator, based on the Winograd algorithm, that are easy to customise for general-purpose processor architectures. In our approach, augmenting the portability of the solution is achieved via the introduction of vector instructions from Intel SSE/AVX2/AVX512 and ARM NEON/SVE to exploit the single-instruction multiple-data capabilities of current processors as well as OpenMP pragmas to exploit multi-threaded parallelism. While this comes at the cost of sacrificing a fraction of the computational performance, our experimental results on three distinct processors, with Intel Xeon Skylake, ARM Cortex A57 and Fujitsu A64FX processors, show that the impact is affordable and still renders a Winograd-based solution that is competitive when compared with the lowering gemm-based convolution.
Manuel F. Dolz, Héctor Martínez 0002, Adrián Castelló 0001, Pedro Alonso 0002, Enrique S. Quintana-Ortí
J. Supercomput.4
2023 Parallel border tracking in binary images for multicore computers
abstract
Abstract Border tracking in binary images is an important operation in many computer vision applications. The problem consists in finding borders in a 2D binary image (where all of the pixels are either 0 or 1). There are several algorithms available for this problem, but most of them are sequential. In a former paper, a parallel border tracking algorithm was proposed. This algorithm was designed to run in Graphics Processing units, and it was based on the sequential algorithm known as the Suzuki algorithm. In this paper, we adapt the previously proposed GPU algorithm so that it can be executed in multicore computers. The resulting algorithm is evaluated against its GPU counterpart. The results show that the performance of the GPU algorithm worsens (or even fails) for very large images or images with many borders. On the other hand, the proposed multicore algorithm can efficiently cope with large images.
Víctor M. García 0001, Pedro Alonso 0002
J. Supercomput.2
2023 Leveraging an open source serverless framework for high energy physics computing
abstract
Abstract CERN (Centre Europeen pour la Recherce Nucleaire) is the largest research centre for high energy physics (HEP). It offers unique computational challenges as a result of the large amount of data generated by the large hadron collider. CERN has developed and supports a software called ROOT , which is the de facto standard for HEP data analysis. This framework offers a high-level and easy-to-use interface called RDataFrame , which allows managing and processing large data sets. In recent years, its functionality has been extended to take advantage of distributed computing capabilities. Thanks to its declarative programming model, the user-facing API can be decoupled from the actual execution backend . This decoupling allows physical analysis to scale automatically to thousands of computational cores over various types of distributed resources. In fact, the distributed RDataFrame module already supports the use of established general industry engines such as Apache Spark or Dask. Notwithstanding the foregoing, these current solutions will not be sufficient to meet future requirements in terms of the amount of data that the new projected accelerators will generate. It is of interest, for this reason, to investigate a different approach, the one offered by serverless computing. Based on a first prototype using AWS Lambda , this work presents the creation of a new backend for RDataFrame distributed over the OSCAR tool, an open source framework that supports serverless computing. The implementation introduces new ways, relative to the AWS Lambda -based prototype, to synchronize the work of functions.
Vincenzo Eduardo Padulano, Pablo Oliver Cortés, Pedro Alonso 0002, Enric Tejedor, Sebastián Risco, Germán Moltó
J. Supercomput.3
2023 Efficient GPU implementation of a Boltzmann-Schrödinger-Poisson solver for the simulation of nanoscale DG MOSFETs
abstract
Abstract A previous study by Mantas and Vecil (Int J High Perform Comput Appl 34(1): 81–102, 2019) describes an efficient and accurate solver for nanoscale DG MOSFETs through a deterministic Boltzmann-Schrödinger-Poisson model with seven electron–phonon scattering mechanisms on a hybrid parallel CPU/GPU platform. The transport computational phase, i.e. the time integration of the Boltzmann equations, was ported to the GPU using CUDA extensions, but the computation of the system’s eigenstates, i.e. the solution of the Schrödinger-Poisson block, was parallelized only using OpenMP due to its complexity. This work fills the gap by describing a port to GPU for the solver of the Schrödinger-Poisson block. This new proposal implements on GPU a Scheduled Relaxation Jacobi method to solve the sparse linear systems which arise in the 2D Poisson equation. The 1D Schrödinger equation is solved on GPU by adapting a multi-section iteration and the Newton-Raphson algorithm to approximate the energy levels, and the Inverse Power Iterative Method is used to approximate the wave vectors. We want to stress that this solver for the Schrödinger-Poisson block can be thought as a module independent of the transport phase (Boltzmann) and can be used for solvers using different levels of description for the electrons; therefore, it is of particular interest because it can be adapted to other macroscopic, hence faster, solvers for confined devices exploited at industrial level.
Francesco Vecil, José Miguel Mantas, Pedro Alonso 0002
J. Supercomput.3
2022 A Serverless Engine for High Energy Physics Distributed Analysis
abstract
The Large Hadron Collider (LHC) at CERN has generated in the last decade an unprecedented volume of data for the High-Energy Physics (HEP) field. Scientific collaborations interested in analysing such data very often require computing power beyond a single machine. This issue has been tackled traditionally by running analyses in distributed environments using stateful, managed batch computing systems. While this approach has been effective so far, current estimates for future computing needs of the field present large scaling challenges. Such a managed approach may not be the only viable way to tackle them and an interesting alternative could be provided by serverless architectures, to enable an even larger scaling potential. This work describes a novel approach to running real HEP scientific applications through a distributed serverless computing engine. The engine is built upon ROOT, a well-established HEP data analysis software, and distributes its computations to a large pool of concurrent executions on Amazon Web Services Lambda Serverless Platform. Thanks to the developed tool, physicists are able to access datasets stored at CERN (also those that are under restricted access policies) and process it on remote infrastructures outside of their typical environment. The analysis of the serverless functions is monitored at runtime to gather performance metrics, both for data- and computation-intensive workloads.
Jacek Kusnierz, Vincenzo Eduardo Padulano, Maciej Malawski, Kamil Burkiewicz, Enric Tejedor, Pedro Alonso 0002, Michael Pitt, Valentina Avati
CCGRID6
2022 Convolution Operators for Deep Learning Inference on the Fujitsu A64FX Processor
abstract
The convolution operator is a crucial kernel for many computer vision and signal processing applications that rely on deep learning (DL) technologies. As such, the efficient implementation of this operator has received considerable attention in the past few years for a fair range of processor architectures. In this paper, we follow the technology trend toward integrating long SIMD (single instruction, multiple data) arithmetic units into high performance multicore processors to analyse the benefits of this type of hardware acceleration for latency-constrained DL workloads. For this purpose, we implement and optimise for the Fujitsu processor A64FX, three distinct methods for the calculation of the convolution, namely, the lowering approach, a blocked variant of the direct convolution algorithm, and the Winograd minimal filtering algorithm. Our experimental results include an extensive evaluation of the parallel scalability of these three methods and a comparison of their global performance using three popular DL models and a representative dataset.
Manuel F. Dolz, Héctor Martínez 0002, Pedro Alonso 0002, Enrique S. Quintana-Ortí
SBAC-PAD3
2022 Parallel border tracking in binary images using GPUs
abstract
Abstract Border tracking in binary images is an important kernel for many applications. There are very efficient sequential algorithms, most notably, the algorithm proposed by Suzuki et al., which has been implemented for CPUs in well-known libraries. However, under some circumstances, it would be advantageous to perform the border tracking in GPUs as efficiently as possible. In this paper, we propose a parallel version of the Suzuki algorithm that is designed to be executed in GPUs and implemented in CUDA. The proposed algorithm is based on splitting the image into small rectangles. Then, a thread is launched for each rectangle, which tracks the borders in its associated rectangle. The final step is to perform the connection of the borders belonging to several rectangles. The parallel algorithm has been compared with a state-of-the-art sequential CPU version, using two different CPUs and two different GPUs for the evaluation. The computing times obtained show that in these experiments with the GPUs and CPUs that we had available, the proposed parallel algorithm running in the fastest GPU is more than 10 times faster than the sequential CPU routine running in the fastest CPU.
Víctor M. García 0001, Pedro Alonso 0002, Ricardo García-Laguía
J. Supercomput.2
2022 Parallel signal detection for generalized spatial modulation MIMO systems
abstract
Abstract Generalized Spatial Modulation is a recently developed technique that is designed to enhance the efficiency of transmissions in MIMO Systems. However, the procedure for correctly retrieving the sent signal at the receiving end is quite demanding. Specifically, the computation of the maximum likelihood solution is computationally very expensive. In this paper, we propose a parallel method for the computation of the maximum likelihood solution using the parallel computing library OpenMP. The proposed parallel algorithm computes the maximum likelihood solution faster than the sequential version, and substantially reduces the worst-case computing times.
Víctor M. García 0001, M. Ángeles Simarro, Francisco-Jose Martínez-Zaldívar, Murilo Boratto, Pedro Alonso 0002, Alberto González 0001
J. Supercomput.5
2021 High Performance and Energy Efficient Integer Matrix Multiplication for Deep Learning
abstract
We present a multi-threaded implementation of the matrix multiplication for deep learning on ARM multicore processors. Following standard practice for inference with convolutional neural networks, our GEMM kernel operates with 16-bit integer arithmetic, yielding significant performance acceleration and cutting the memory requirements with respect to IEEE (floating point) single precision by half, allowing the deployment of larger neural network models on low power devices with limited storage capacity.
Pau San Juan, Pedro Alonso 0002, Enrique S. Quintana-Ortí
PDP2
2021 Low precision matrix multiplication for efficient deep learning in NVIDIA Carmel processors
Pablo San Juan, Rafael Rodríguez-Sánchez 0001, Francisco D. Igual, Pedro Alonso 0002, Enrique S. Quintana-Ortí
J. Supercomput.4
2020 High Performance and Portable Convolution Operators for Multicore Processors
abstract
The considerable impact of Convolutional Neural Networks on many Artificial Intelligence tasks has led to the development of various high performance algorithms for the convolution operator present in this type of networks. One of these approaches leverages the IM2COL transform followed by a general matrix multiplication (GEMM) in order to take advantage of the highly optimized realizations of the GEMM kernel in many linear algebra libraries. The main problems of this approach are 1) the large memory workspace required to host the intermediate matrices generated by the IM2COL transform; and 2) the time to perform the IM2COL transform, which is not negligible for complex neural networks. This paper presents a portable high performance convolution algorithm based on the BLIS realization of the GEMM kernel that avoids the use of the intermediate memory by taking advantage of the BLIS structure. In addition, the proposed algorithm eliminates the cost of the explicit IM2COL transform, while maintaining the portability and performance of the underlying realization of GEMM in BLIS.
Pablo San Juan, Adrián Castelló 0001, Manuel F. Dolz, Pedro Alonso 0002, Enrique S. Quintana-Ortí
SBAC-PAD4
2020 Performance modeling of the sparse matrix-vector product via convolutional neural networks
Maria Barreda, Manuel F. Dolz, M. Asunción Castaño, Pedro Alonso 0002, Enrique S. Quintana-Ortí
J. Supercomput.4
2019 Computing matrix trigonometric functions with GPUs through Matlab
Pedro Alonso 0002, Jesús Peinado-Pinilla, Jacinto Javier Ibáñez, Jorge Sastre, Emilio Defez
J. Supercomput.1
2019 Fast block QR update in digital signal processing
Fran J. Alventosa, Pedro Alonso 0002, Antonio M. Vidal, Gema Piñero, Enrique S. Quintana-Ortí
J. Supercomput.2
2019 Exploring hybrid parallel systems for probabilistic record linkage
Murilo Boratto, Pedro Alonso 0002, Clícia Pinto, Pedro Melo, Marcos E. Barreto, Spiros C. Denaxas
J. Supercomput.2
2019 HReMAS: hybrid real-time musical alignment system
Pablo Cabañas Molero, Raquel Cortina, Elías F. Combarro, Pedro Alonso 0002, F. J. Bris-Peñalver
J. Supercomput.4
2019 A pipeline structure for the block QR update in digital signal processing
Manuel F. Dolz, Fran J. Alventosa, Pedro Alonso 0002, Antonio M. Vidal
J. Supercomput.3
2019 Real-time Soundprism
Antonio Jesús Muñoz-Montoro, José Ranilla, Pedro Vera-Candeas, Elías F. Combarro, Pedro Alonso 0002
J. Supercomput.5
2018 Two-sided orthogonal reductions to condensed forms on asymmetric multicore processors
Pedro Alonso 0002, Sandra Catalán, José R. Herrero 0001, Enrique S. Quintana-Ortí, Rafael Rodríguez-Sánchez 0001
Parallel Comput.1
2017 Parallel online time warping for real-time audio-to-score alignment in multi-core systems
Pedro Alonso 0002, Raquel Cortina, Francisco J. Rodríguez-Serrano, Pedro Vera-Candeas, M. Alonso-González, José Ranilla
J. Supercomput.1
2017 High-performance computing: the essential tool and the essential challenge
Pedro Alonso 0002, José Ranilla, Jesús Vigo-Aguiar
J. Supercomput.1
2017 An efficient musical accompaniment parallel system for mobile devices
Pedro Alonso 0002, Pedro Vera-Candeas, Raquel Cortina, José Ranilla
J. Supercomput.1
2017 Accelerating multi-channel filtering of audio signal on ARM processors
Jose A. Belloch, Fran J. Alventosa, Pedro Alonso 0002, Enrique S. Quintana-Ortí, Antonio M. Vidal
J. Supercomput.3
2017 Automatic tuning to performance modelling of matrix polynomials on multicore and multi-GPU systems
Murilo Boratto, Pedro Alonso 0002, Domingo Giménez, Alexey L. Lastovetsky
J. Supercomput.2
2014 Enhancing performance and energy consumption of runtime schedulers for dense linear algebra
abstract
SUMMARY The road towards Exascale Computing requires a holistic effort to address three different challenges simultaneously: high performance, energy efficiency, and programmability. The use of runtime task schedulers to orchestrate parallel executions with minimal developer intervention has been introduced in recent years to tackle the programmability issue while maintaining, or even improving, performance. In this paper, we enhance the SuperMatrix runtime task scheduler integrated in the libflame library in two different directions that address high performance and energy efficiency. First, we extend the runtime by accommodating hybrid parallel executions and managing task priorities for dense linear algebra operations, with remarkable performance improvements. Second, we introduce techniques to reduce energy consumption during idle times inherent to parallel executions, attaining important energy savings. In addition, we propose a power consumption model that can be leveraged by runtime task schedulers to make decisions based not only on performance but also on energy considerations. Copyright © 2014 John Wiley & Sons, Ltd.
Pedro Alonso 0002, Manuel F. Dolz, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí
Concurr. Comput. Pract. Exp.1
2014 Modeling power and energy consumption of dense matrix factorizations on multicore processors
abstract
SUMMARY In this paper, we propose a model for the energy consumption of the concurrent execution of three key dense matrix factorizations, with task parallelism leveraged via the Symmetric Multi‐Processing Superscalar (SMPSs) runtime, on a multicore processor. Our model decomposes the power dissipation into the system, static and dynamic components, with the former two being estimated from basic, off‐line experiments. The dynamic power, on the other hand, requires significantly more care, and we introduce a contention‐aware model that accommodates for the variability of power consumption due to memory contention. Experimental results on an Intel Xeon E5504 processor with four cores, using an internal powermeter that samples the power drawn by the mainboard with a frequency of 1 KHz, show the reliability of the energy model for the Cholesky, LU, and QR factorizations on this platform. Copyright © 2013 John Wiley & Sons, Ltd.
Pedro Alonso 0002, Manuel F. Dolz, Rafael Mayo 0002, Enrique S. Quintana-Ortí
Concurr. Comput. Pract. Exp.1
2014 Block pivoting implementation of a symmetric Toeplitz solver
Pedro Alonso 0002, Manuel F. Dolz, Antonio M. Vidal
J. Parallel Distributed Comput.1
2014 Automatic routine tuning to represent landform attributes on multicore and multi-GPU systems
Murilo Boratto, Pedro Alonso 0002, Domingo Giménez, Marcos E. Barreto
J. Supercomput.2
2014 Solving time-invariant differential matrix Riccati equations using GPGPU computing
Jesús Peinado-Pinilla, Pedro Alonso 0002, Jacinto Javier Ibáñez, Vicente Hernández, Murilo Boratto
J. Supercomput.2
2013 A multicore solution to Block-Toeplitz linear systems of equations
Pedro Alonso 0002, Daniel Argüelles, José Ranilla, Antonio M. Vidal
J. Supercomput.1
2012 Parallel Algorithm for Landform Attributes Representation on Multicore and Multi-GPU Systems
Murilo Boratto, Pedro Alonso 0002, Carla Ramiro, Marcos E. Barreto, Leandro dos Santos Coelho
ICCSA (1)2
2012 Tools for Power-Energy Modelling and Analysis of Parallel Scientific Applications
abstract
Understanding power usage in parallel workloads is crucial to develop the energy-aware software that will run in future Exascale systems. In this paper, we contribute towards this goal by introducing an integrated framework to profile, monitor, model and analyze power dissipation in parallel MPI and multi-threaded scientific applications. The framework includes an own-designed device to measure internal DC power consumption and a package offering a simple interface to interact with this design as well as commercial power meters. Combined with the instrumentation package Extrae and the graphical analysis tool Paraver, the result is a useful environment to identify sources of power inefficiency directly in the source application code. For task-parallel codes, we also offer a statistical software module that inspects the execution trace of the application to calculate the parameters of an accurate model for the global energy consumption, which can be then decomposed into the average power usage per task or the nodal power dissipated per core.
Pedro Alonso 0002, Rosa M. Badia, Jesús Labarta, Maria Barreda, Manuel F. Dolz, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Ruymán Reyes
ICPP1
2012 Reducing Energy Consumption of Dense Linear Algebra Operations on Hybrid CPU-GPU Platforms
abstract
We investigate the balance between the time-to-solution and the energy consumption of a task-parallel execution of the Cholesky and LU factorizations on a hybrid platform, equipped with a multi-core processor and several GPUs. To improve energy efficiency, we incorporate two energy-saving techniques in the runtime in charge of scheduling the computations, to block idle threads and enable the transition to a more energy-friendly state of the general-purpose cores. Experiments on an Intel Xeon-based platform connected to an NVIDIA Tesla server report an average reduction of the energy consumption close to 9% (38% when only the consumption associated with the application is considered), for a minor increase in the execution time of the algorithm.
Pedro Alonso 0002, Manuel F. Dolz, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí
ISPA1
2012 Saving Energy in the LU Factorization with Partial Pivoting on Multi-core Processors
abstract
In this paper we analyze the trade-off between energy and performance for a data-parallel execution of the LU factorization with partial pivoting on a multi-core processor. To improve energy efficiency, we adapt the runtime in charge of controlling the concurrent execution of the algorithm to leverage DVFS and block idle threads. For a CPU-bounded operation like the LU factorization, experiments on an AMD 8-core processor report a reduction around 5% in energy consumption for the largest problem sizes in exchange for a minor increase in the execution time.
Pedro Alonso 0002, Manuel F. Dolz, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí
PDP1
2011 Implementation and tuning of a parallel symmetric Toeplitz eigensolver
Pedro Alonso 0002, Miguel O. Bernabeu, Víctor M. García 0001, Antonio M. Vidal
J. Parallel Distributed Comput.1
2010 Experimental Study of Six Different Implementations of Parallel Matrix Multiplication on Heterogeneous Computational Clusters of Multicore Processors
abstract
Two strategies of distribution of computations can be used to implement parallel solvers for dense linear algebra problems for Heterogeneous Computational Clusters of Multicore Processors (HCoMs). These strategies are called Heterogeneous Process Distribution Strategy (HPS) and Heterogeneous Data Distribution Strategy (HDS). They are not novel and have been researched thoroughly. However, the advent of multicores necessitates enhancements to them. In this paper, we present these enhancements. Our study is based on experiments using six applications to perform Parallel Matrix-matrix Multiplication (PMM) on an HCoM employing the two distribution strategies.
Pedro Alonso 0002, Ravi Reddy, Alexey L. Lastovetsky
PDP1
2009 Parallel solvers for dense linear systems for heterogeneous computational clusters
abstract
This paper describes the design and the implementation of parallel routines in the heterogeneous ScaLAPACK library that solve a dense system of linear equations. This library is written on top of HeteroMPI and ScaLAPACK whose building blocks, the de facto standard kernels for matrix and vector operations (BLAS and its parallel counterpart PBLAS) and message passing communication (BLACS), are optimized for heterogeneous computational clusters. We show that the efficiency of these parallel routines is due to the most important feature of the library, which is the automation of the difficult optimization tasks of parallel programming on heterogeneous computing clusters. They are the determination of the accurate values of the platform parameters such as the speeds of the processors and the latencies and bandwidths of the communication links connecting different pairs of processors, the optimal values of the algorithmic parameters such as the total number of processes, the 2D process grid arrangement and the efficient mapping of the processes executing the parallel algorithm to the executing nodes of the heterogeneous computing cluster. We describe this process of automation followed by presentation of experimental results on a local heterogeneous computing cluster demonstrating the efficiency of these solvers.
Ravi Reddy, Alexey L. Lastovetsky, Pedro Alonso 0002
IPDPS3
2008 Scalable Dense Factorizations for Heterogeneous Computational Clusters
abstract
This paper discusses the design and the implementation of the LU factorization routines included in the Heterogeneous ScaLAPACK library, which is built on top of ScaLAPACK. These routines are used in the factorization and solution of a dense system of linear equations. They are implemented using optimized PBLAS, BLACS and BLAS libraries for heterogeneous computational clusters. We present the details of the implementation as well asperformance results on a heterogeneous computingcluster.
Ravi Reddy, Alexey L. Lastovetsky, Pedro Alonso 0002
ISPDC3
2008 Heterogeneous PBLAS: Optimization of PBLAS for Heterogeneous Computational Clusters
abstract
This paper presents a package, called Heterogeneous PBLAS (HeteroPBLAS), which is built on top of PBLAS and provides optimized parallel basic linear algebra subprograms for heterogeneous computational clusters. We present the user interface and the software hierarchy of the first research implementation of HeteroPBLAS. This is the first step towards the development of a parallel linear algebra package for heterogeneous computational clusters. We demonstrate the efficiency of the HeteroPBLAS programs on a homogeneous computing cluster and a heterogeneous computing cluster.
Ravi Reddy, Alexey L. Lastovetsky, Pedro Alonso 0002
ISPDC3
2008 A Threaded Divide and Conquer Symmetric Tridiagonal Eigensolver on Multicore Systems
abstract
The increasing power of computation of modern processors rely on the increasing number of cores per chip. The challenge of software developers is to keep this power with the legacy code. Although commercial and non commercial libraries are improving their codes step by step, there exits probably insurmountable scalability issues for standard programming models due to the fact that using locks to implement synchronisation is inherently a bottleneck. We propose an implementation of the divide and conquer algorithm to compute the eigenpairs of symmetric tridiagonal matrices on multicore systems. We take advantage of the natural parallelism of the method by using pthreads. We avoided as much as possible the negative impact of synchronisation in the performance by overlapping operations of different classes. Furthermore, the unevenly workload distribution of the computational cost of the elemental tasks yields in a speedup even larger than expected.
Antonio M. Vidal, Murilo Boratto, Pedro Alonso 0002
ISPDC3
2008 Parallel computation of the eigenvalues of symmetric Toeplitz matrices through iterative methods
Antonio M. Vidal, Víctor M. García 0001, Pedro Alonso 0002, Miguel O. Bernabeu
J. Parallel Distributed Comput.3
2008 A multilevel parallel algorithm to solve symmetric Toeplitz linear systems
Miguel O. Bernabeu, Pedro Alonso 0002, Antonio M. Vidal
J. Supercomput.2
2006 A Parallel Algorithm for the Solution of the Deconvolution Problem on Heterogeneous Networks
abstract
In this work we present a parallel algorithm for the solution of a least squares problem with structured matrices. This problem arises in many applications mainly related to digital signal processing. The parallel algorithm is designed to speed up the sequential one on heterogeneous networks of computers. The parallel algorithm follows the HeHo strategy (Heterogeneous distribution of processes over processors with homogeneous distribution of computations over the processes) and is implemented using HeteroMPI, a recently developed extension of MPI for programming high performance computations on heterogeneous networks of computers. The obtained results validate HeteroMPI as a very useful tool for portable implementation of parallel algorithms for heterogeneous environments
Pedro Alonso 0002, Antonio M. Vidal, Alexey L. Lastovetsky
CLUSTER1
2005 Solving the block-Toeplitz least-squares problem in parallel
abstract
In this paper we present two versions of a parallel algorithm to solve the block–Toeplitz least-squares problem on distributed-memory architectures. We derive a parallel algorithm based on the seminormal equations arising from the triangular decomposition of the product T TT . Our parallel algorithm exploits the displacement structure of the Toeplitz-likematrices using theGeneralized SchurAlgorithm to obtain the solution in O(mn) flops instead of O(mn2) flops of the algorithms for non-structured matrices. The strong regularity of the previous product of matrices and an appropriate computation of the hyperbolic rotations improve the stability of the algorithms. We have reduced the communication cost of previous versions, and have also reduced the memory access cost by appropriately arranging the elements of the matrices. Furthermore, the second version of the algorithm has a very low spatial cost, because it does not store the triangular factor of the decomposition. The experimental results show a good scalability of the parallel algorithm on two different clusters of personal computers. Copyright c © 2005 John Wiley & Sons, Ltd.
Pedro Alonso 0002, José M. Badía, Antonio M. Vidal
Concurr. Pract. Exp.1
2005 An Efficient Parallel Algorithm to Solve Block-Toeplitz Systems
Pedro Alonso 0002, José M. Badía, Antonio M. Vidal
J. Supercomput.1