Francisco D. Igual

dblp:88/501 · also Francisco Daniel Igual Peña · DBLP profile ↗
← Back
54ranked-venue papers
3as first author
20since 2021 · last 2026
0000-0003-4480-9517ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 36 · 3 first-author · 12 since 2021Theory of computation · 7 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 The cambrian explosion of mixed-precision matrix multiplication for quantized deep learning inference
abstract
Recent advances in deep learning (DL) have promoted to a shift from traditional 64-bit floating point (FP64) arithmetic for scientific computing toward reduced-precision formats–such as FP16, BF16, or even 8-bit integers–combined with mixed-precision arithmetic. This transition enhances computational throughput, reduces memory and bandwidth usage, and improves energy efficiency, offering significant advantages for resource-constrained edge devices. To support this shift, hardware architectures have evolved accordingly, now including adapted ISAs (Instruction Set Architectures) that expose mixed-precision vector units and matrix engines tailored for DL workloads. At the heart of many DL and scientific computing tasks is the general matrix-matrix multiplication ( GEMM ), a fundamental kernel historically optimized using fused multiply-add (FMA) vector instructions on SIMD (single instruction, multiple data) units. However, as hardware moves toward mixed-precision dot (or inner)-product-centric operations optimized for quantized inference, these legacy approaches are being phased out. In response to this, our paper revisits the conventional, high-performance implementation of GEMM and describes strategies for adapting it to mixed integer precision (MIP) arithmetic across modern ISAs, including x86_64, Arm, and RISC-V. Concretely, we illustrate novel micro-kernel designs and data layouts that better exploit today’s specialized hardware and demonstrate significant performance gains from MIP arithmetic over floating-point implementations across three representative CPUs. These contributions highlight a new era of GEMM optimization-driven by the demands of DL inference on heterogeneous architectures, marking what we term as the “Cambrian period” for matrix multiplication.
Héctor Martínez 0002, Adrián Castelló 0001, Francisco D. Igual, Enrique S. Quintana-Ortí
Future Gener. Comput. Syst.3
2026 Solving the task scheduling and GPU reconfiguration problem on MIG devices via deep reinforcement learning
abstract
• Multi-Instance GPU (MIG) technology enables adaptive co-execution of tasks, greatly improving the efficiency of computational resources in a flexible manner. • Prior methods simplify the MIG scheduling and dynamic reconfiguration challenge by reducing problem complexity, but can yield markedly suboptimal solutions in certain scenarios. • Modeling the problem with Reinforcement Learning (RL), and training with Deep Learning techniques, allows to approach it successfully without great simplifications, despite its high dimensionality. The design of the RL agent also needs to be carefully refined, with analysis such as that detailed in the manuscript, which can serve as a guide for similar resource management work. • Our refined RL agent reduces makespan by 2–7% over state-of-the-art on a wide set of benchmarks and synthetic workloads, with improvements of up to 30% in specific cases. Additional benefits include enhanced flexibility and adaptability of the scheduling framework. Recent advances in dynamic GPU partitioning, such as NVIDIA’s Multi-Instance GPU (MIG) technology, have enhanced resource utilization by enabling task co-execution without contention. However, existing MIG schedulers remain limited to static or task-agnostic methods that sacrifice optimality for tractability. This paper presents a Deep Reinforcement Learning framework that seeks to minimize the completion time of a task queue by holistically addressing the dimensions of the problem: task molding, GPU reconfiguration and execution order. To manage the vast solution space, we apply optimizations such as discrete and canonical representation of states, unification of equivalent configurations, action masking, or promoting the exploration of reconfigurations; this offers insights for similar resource management scenarios. The proposed models are extensively evaluated with widely used benchmarks of the Rodinia and Altis suites, and synthetic workloads generated to emulate a wide range of plausible real situations. The final model improves to the state-of-the-art, especially in workloads that clearly contradict the assumptions of previous proposals, achieving a difference of less than 20% to the optimum. Additionally, two different approaches to the problem are faced (offline vs. online), discussing their theoretical advantages and disadvantages, and evaluating them experimentally for the final model.
Jorge Villarrubia, Luis Costero, Francisco D. Igual, Katzalin Olcoz
Future Gener. Comput. Syst.3
2026 A comprehensive evaluation of spatial co-execution on GPUs using MPS and MIG technologies
Jorge Villarrubia, Luis Costero, Francisco D. Igual, Katzalin Olcoz
J. Supercomput.3
2025 QoS-aware workload scheduling on heterogeneous and dynamic edge-to-cloud deployments
abstract
With the advent of the edge-to-cloud continuum paradigm, the heterogeneity of compute entities imposes new challenges to achieve proper levels of Quality of Service (QoS) in the mapping of workloads and requests to nodes, mainly in terms of response times, but also in terms of application-specific metrics. In addition, for deployments in which computing elements are dynamic by nature (in terms of availability, but also compute capabilities or latency, to name only a few parameters), placing workloads in the most suitable compute element becomes a major challenge. Current generic orchestrators, such as Kubernetes, have demonstrated to be a valid option in homogeneous and static deployments, in which QoS-oblivious scheduling policies typically pursue exclusively load balancing, and ignore other parameters such as latency reduction or limitation of application-level metrics. In this paper, we demonstrate that QoS-oblivious policies in a generic orchestrator such as Kubernetes are not enough for heterogeneous and dynamic edge-to-cloud deployments. We propose a new framework that integrates seamlessly into Kubernetes by means of a service. This framework is equipped with a set of QoS-aware scheduling policies to tackle the heterogeneity and dynamic character of many edge-to-cloud deployments. Our experimental results for a specific use case (a deployment of inference servers across heterogeneous nodes) reveal significant gains in terms of response times under different dynamic scenarios that include computing devices with different capabilities (multi-core CPUs and different types of GPUs).
Julián Cámara-Miró, Luis Costero, Francisco D. Igual
PDP3
2025 Portable, High Performance Matrix Multiplication Micro-Kernels for RISC-V with ExO
abstract
The proliferation of RISC-V platforms and their use in a wide variety of scientific applications, including deep learning scenarios, has dramatically increased the interest to generate optimized code for them. In the field of HPC (High Performance Computing), the RISCV ISA (Instruction Set Architecture) has been adopted by a wide variety of designs with different micro-architecture; as a result, performance portability of existing codes is a major endeavor. Code generators and compilers such as Apache TVM, MLIR, or EXO provide a hardware abstraction for implementing optimized hardware-aware codes, thus reducing development time and potential errors. These generators can handle the full software stack, from basic micro-kernels to complex operations. In this work, we focus on the optimization of GEMM (general matrix-matrix multiplication), a key operation on top of which dense linear algebra libraries and deep learning frameworks are built. Specifically, we present an EXO-based GEMM microkernel generator for the RISC-V ISA with RVV vector extensions that addresses the lack of high-performance and portable GEMM micro-kernels. Our results demonstrate that, by generating a wide range of micro-kernels, one can obtain GEMM realizations that outperform those in the state-of-the-art high performance libraries.
Adrián Castelló 0001, Héctor Martínez 0002, Sandra Catalán, Jie Lei 0007, Yuka Ikarashi, Grace Dinh, Francisco D. Igual, Enrique S. Quintana-Ortí
PDP7
2025 Leveraging Multi-Instance GPUs through moldable task scheduling
abstract
NVIDIA MIG (Multi-Instance GPU) allows partitioning a physical GPU into multiple logical instances with fully-isolated resources, which can be dynamically reconfigured. This work highlights the untapped potential of MIG through moldable task scheduling with dynamic reconfigurations . Specifically, we propose a makespan minimization problem for multi-task execution under MIG constraints. Our profiling shows that assuming monotonicity in task work with respect to resources is not viable, as is usual in multicore scheduling. Relying on a state-of-the-art proposal that does not require such an assumption, we present FAR , a 3-phase algorithm to solve the problem. Phase 1 of FAR builds on a classical task moldability method, phase 2 combines Longest Processing Time First and List Scheduling with a novel repartitioning tree heuristic tailored to MIG constraints, and phase 3 employs local search via task moves and swaps. FAR schedules tasks in batches offline, concatenating their schedules on the fly in an improved way that favors resource reuse. Excluding reconfiguration costs, the List Scheduling proof shows an approximation factor of 7/4 on the NVIDIA A30 model. We adapt the technique to the particular constraints of an NVIDIA A100/H100 to obtain an approximation factor of 2. Including the reconfiguration cost, our real-world experiments reveal a makespan with respect to the optimum no worse than 1.22× for a well-known suite of benchmarks, and 1.10× for synthetic inputs inspired by real kernels. We obtain good experimental results for each batch of tasks, but also in the concatenation of batches, with large improvements over the state-of-the-art and proposals without GPU reconfiguration. Moreover, we show that the proposed heuristics allow a correct adaptation to tasks of very different characteristics. Beyond the specific algorithm, the paper demonstrates the research potential of the MIG technology and suggests useful metrics, workload characterizations and evaluation techniques for future work in this field.
Jorge Villarrubia, Luis Costero, Francisco D. Igual, Katzalin Olcoz
J. Parallel Distributed Comput.3
2025 Experience-guided, mixed-precision matrix multiplication with apache TVM for ARM processors
abstract
Abstract Deep learning (DL) generates new computational tasks that are different from those encountered in classical scientific applications. In particular, DL training and inference require general matrix multiplications (gemm) with matrix operands that are far from large and square as in other scientific fields. In addition, DL models gain arithmetic/storage complexity, and as a result, reduced precision via quantization is now mainstream for inferring DL models in edge devices. Automatic code generation addresses these new types of gemm by (1) improving portability between different hardware with only one base code; (2) supporting mixed and reduced precision; and (3) enabling auto-tuning methods that, given a base operation, perform a (costly) optimization search for the best schedule. In this paper, we rely on Apache TVM to generate an experience-guided gemm that provides performance competitive with the TVM auto-scheduler, while reducing tuning time by a factor of 48×.
Adrián Castelló 0001, Héctor Martínez 0002, Sandra Catalán, Francisco D. Igual, Enrique S. Quintana-Ortí
J. Supercomput.4
2025 Balanced segmentation of CNNs for multi-TPU inference
John S. Villarrubia, Luis Costero, Francisco D. Igual, Katzalin Olcoz
J. Supercomput.3
2024 Inference with Transformer Encoders on ARM and RISC-V Multicore Processors
abstract
Abstract We delve into the performance of transformer encoder inference on low-power multi-core processors from two perspectives: First, we conduct a detailed profile of the inference process for two members of the BERT family on a modern multi-core processor, identifying the main bottlenecks and opportunities for improvement. Second, we propose a number of accumulative optimisations for their primary building blocks. For that, we elaborate our own implementation of the general matrix multiplication (), which dynamically tunes several key parameters yielding relevant performance gains for transformer encoders. Additionally, we introduce a number of strategies to also improve the parallel execution of the transformer block. Our implementations for ARMv8a and RISC-V multi-core processors with SIMD units, taking as a reference state-of-the-art implementations (BLIS for ARM and OpenBLAS for RISC-V) reveal accelerations of up to $$2.5\times $$ 2.5 × for natural language processing tasks.
Héctor Martínez 0002, Francisco D. Igual, Rafael Rodríguez-Sánchez 0001, Sandra Catalán, Adrián Castelló 0001, Enrique S. Quintana-Ortí
Euro-Par (2)2
2024 Automatic generation of ARM NEON micro-kernels for matrix multiplication
abstract
Abstract General matrix multiplication ( gemm ) is a fundamental kernel in scientific computing and current frameworks for deep learning. Modern realisations of gemm are mostly written in C, on top of a small, highly tuned micro-kernel that is usually encoded in assembly. The high performance realisation of gemm in linear algebra libraries in general include a single micro-kernel per architecture, usually implemented by an expert. In this paper, we explore a couple of paths to automatically generate gemm micro-kernels, either using C++ templates with vector intrinsics or high-level Python scripts that directly produce assembly code. Both solutions can integrate high performance software techniques, such as loop unrolling and software pipelining, accommodate any data type, and easily generate micro-kernels of any requested dimension. The performance of this solution is tested on three ARM-based cores and compared with state-of-the-art libraries for these processors: BLIS, OpenBLAS and ArmPL. The experimental results show that the auto-generation approach is highly competitive, mainly due to the possibility of adapting the micro-kernel to the problem dimensions.
Guillermo Alaejos, Héctor Martínez 0002, Adrián Castelló 0001, Manuel F. Dolz, Francisco D. Igual, Pedro Alonso 0002, Enrique S. Quintana-Ortí
J. Supercomput.5
2024 Algorithm 1039: Automatic Generators for a Family of Matrix Multiplication Routines with Apache TVM
abstract
We explore the utilization of the Apache TVM open source framework to automatically generate a family of algorithms that follow the approach taken by popular linear algebra libraries, such as GotoBLAS2, BLIS, and OpenBLAS, to obtain high-performance blocked formulations of the general matrix multiplication ( gemm ). In addition, we fully automatize the generation process by also leveraging the Apache TVM framework to derive a complete variety of the processor-specific micro-kernels for gemm . This is in contrast with the convention in high-performance libraries, which hand-encode a single micro-kernel per architecture using Assembly code. In global, the combination of our TVM-generated blocked algorithms and micro-kernels for gemm (1) improves portability, maintainability, and, globally, streamlines the software life cycle; (2) provides high flexibility to easily tailor and optimize the solution to different data types, processor architectures, and matrix operand shapes, yielding performance on a par (or even superior for specific matrix shapes) with that of hand-tuned libraries; and (3) features a small memory footprint.
Guillermo Alaejos, Adrián Castelló 0001, Pedro Alonso 0002, Francisco D. Igual, Héctor Martínez 0002, Enrique S. Quintana-Ortí
ACM Trans. Math. Softw.4
2023 Improving inference time in multi-TPU systems with profiled model segmentation
abstract
In this paper, we systematically evaluate the inference performance of the Edge TPU by Google for neural networks with different characteristics. Specifically, we determine that, given the limited amount of on-chip memory on the Edge TPU, accesses to external (host) memory rapidly become an important performance bottleneck. We demonstrate how multiple devices can be jointly used to alleviate the bottleneck introduced by accessing the host memory. We propose a solution combining model segmentation and pipelining on up to four TPUs, with remarkable performance improvements that range from 6x for neural networks with convolutional layers to 46x for fully connected layers, compared with single-TPU setups.
Jorge Villarrubia, Luis Costero, Francisco D. Igual, Katzalin Olcoz
PDP3
2023 Fine-grain task-parallel algorithms for matrix factorizations and inversion on many-threaded CPUs
abstract
Abstract We extend a two‐level task partitioning previously applied to the inversion of dense matrices via Gauss–Jordan elimination to the more challenging QR factorization as well as the initial orthogonal reduction to band form found in the singular value decomposition. Our new task‐parallel algorithms leverage the tasking mechanism currently available in OpenMP to exploit “nested” task parallelism, with a first outer level that operates on matrix panels and a second inner level that processes the matrix either by ‐panels or by tiles, in order to expose a large number of independent tasks. We present a detailed performance analysis, including execution traces, which shows that the two‐level refinement into fine grain tasks allows for an improved load balancing and delivers high performance on current general‐purpose many‐core processors (CPUs) from Intel and AMD.
Sandra Catalán, José R. Herrero 0001, Francisco D. Igual, Enrique S. Quintana-Ortí, Rafael Rodríguez-Sánchez 0001
Concurr. Comput. Pract. Exp.3
2023 Programming parallel dense matrix factorizations and inversion for new-generation NUMA architectures
abstract
We propose a methodology to address the programmability issues derived from the emergence of new-generation shared-memory NUMA architectures. For this purpose, we employ dense matrix factorizations and matrix inversion (DMFI) as a use case, and we target two modern architectures (AMD Rome and Huawei Kunpeng 920) that exhibit configurable NUMA topologies. Our methodology pursues performance portability across different NUMA configurations by proposing multi-domain implementations for DMFI plus a hybrid task- and loop-level parallelization that configures multi-threaded executions to fix core-to-data binding, exploiting locality at the expense of minor code modifications. In addition, we introduce a generalization of the multi-domain implementations for DMFI that offers support for virtually any NUMA topology in present and future architectures. Our experimentation on the two target architectures for three representative dense linear algebra operations validates the proposal, reveals insights on the necessity of adapting both the codes and their execution to improve data access locality, and reports performance across architectures and inter- and intra-socket NUMA configurations competitive with state-of-the-art message-passing implementations, maintaining the ease of development usually associated with shared-memory programming.
Sandra Catalán, Francisco D. Igual, José R. Herrero 0001, Rafael Rodríguez-Sánchez 0001, Enrique S. Quintana-Ortí
J. Parallel Distributed Comput.2
2023 Micro-kernels for portable and efficient matrix multiplication in deep learning
abstract
Abstract We provide a practical demonstration that it is possible to systematically generate a variety of high-performance micro-kernels for the general matrix multiplication (gemm) via generic templates which can be easily customized to different processor architectures and micro-kernel dimensions. These generic templates employ vector intrinsics to exploit the SIMD (single instruction, multiple data) units in current general-purpose processors and, for the particular type of gemm problems encountered in deep learning, deliver a floating-point throughput rate on par with or even higher than that obtained with conventional, carefully tuned implementations of gemm in current linear algebra libraries (e.g., BLIS, AMD AOCL, ARMPL). Our work exposes the structure of the template-based micro-kernels for ARM Neon (128-bit SIMD), ARM SVE (variable-length SIMD) and Intel AVX512 (512-bit SIMD), showing considerable performance for an NVIDIA Carmel processor (ARM Neon), a Fujitsu A64FX processor (ARM SVE) and on an AMD EPYC 7282 processor (256-bit SIMD).
Guillermo Alaejos, Adrián Castelló 0001, Héctor Martínez 0002, Pedro Alonso 0002, Francisco D. Igual, Enrique S. Quintana-Ortí
J. Supercomput.5
2023 Algorithm 1033: Parallel Implementations for Computing the Minimum Distance of a Random Linear Code on Distributed-memory Architectures
abstract
The minimum distance of a linear code is a key concept in information theory. Therefore, the time required by its computation is very important to many problems in this area. In this article, we introduce a family of implementations of the Brouwer–Zimmermann algorithm for distributed-memory architectures for computing the minimum distance of a random linear code over 𝔽 2 . Both current commercial and public-domain software only work on either unicore architectures or shared-memory architectures, which are limited in the number of cores/processors employed in the computation. Our implementations focus on distributed-memory architectures, thus being able to employ hundreds or even thousands of cores in the computation of the minimum distance. Our experimental results show that our implementations are much faster, even up to several orders of magnitude, than current implementations widely used nowadays.
Gregorio Quintana-Ortí, Fernando Hernando, Francisco D. Igual
ACM Trans. Math. Softw.3
2022 Anatomy of the BLIS Family of Algorithms for Matrix Multiplication
abstract
The efforts of the scientific community and hardware vendors to develop and optimize linear algebra codes have historically led to highly-tuned libraries, carefully adapted to the underlying processor architecture, with excellent (near-peak) performance. These optimization efforts, however, are commonly focused on obtaining the best performance possible when the involved operands are large and “squarish” matrices. New computationally-intensive applications (e.g., in deep learning) are increasingly demanding high-performance BLAS (Basic Linear Algebra Subprograms) also for small operands in any of their dimensions. In this paper, we tackle this problem by refactoring the general matrix-matrix multiplication (GEMM) algorithm within a specific high-performance implementation of BLAS, named BLIS, proposing a complete family of algorithmic variants to implement GEMM with different strategies to exploit the target cache hierarchy, together with the changes to be applied to architecture-specific codes to instantiate a complete GEMM implementation. Experimental results on an ARM processor (NVIDIA Carmel) reveal significant performance differences between the members of the GEMM family, depending on the shape and dimension of the matrix operands.
Adrián Castelló 0001, Enrique S. Quintana-Ortí, Francisco D. Igual
PDP3
2022 NUMA-Aware Dense Matrix Factorizations and Inversion with Look-Ahead on Multicore Processors
abstract
We address the efficient design and implementation of dense matrix factorizations and inversion (DMFI) on modern multicore processors with several NUMA (non-uniform memory access) nodes. Our approach enhances the DMFI routines with a look-ahead strategy, in order to overcome the “panel factorization bottleneck”. In addition, it exploits both hybrid task- and loop-level parallelizations while taking into account the NUMA organization of the memory hierarchy. The experiments on a Huawei Kunpeng-based server, with two sockets and 48 cores per socket, for three representative dense linear algebra operations, expose the necessity of adapting both the codes and their execution environment parameters to improve data access locality. The results of these changes deliver performance across inter- and intra-socket NUMA configurations superior to that of reference implementations from state-of-the-art libraries for this platform.
Sandra Catalán, Francisco D. Igual, Rafael Rodríguez-Sánchez 0001, José R. Herrero 0001, Enrique S. Quintana-Ortí
SBAC-PAD2
2022 Algorithm 1022: Efficient Algorithms for Computing a Rank-Revealing UTV Factorization on Parallel Computing Architectures
abstract
Randomized singular value decomposition (RSVD) is by now a well-established technique for efficiently computing an approximate singular value decomposition of a matrix. Building on the ideas that underpin RSVD, the recently proposed algorithm “randUTV” computes a full factorization of a given matrix that provides low-rank approximations with near-optimal error. Because the bulk of randUTV is cast in terms of communication-efficient operations such as matrix-matrix multiplication and unpivoted QR factorizations, it is faster than competing rank-revealing factorization methods such as column-pivoted QR in most high-performance computational settings. In this article, optimized randUTV implementations are presented for both shared-memory and distributed-memory computing environments. For shared memory, randUTV is redesigned in terms of an algorithm-by-blocks that, together with a runtime task scheduler, eliminates bottlenecks from data synchronization points to achieve acceleration over the standard blocked algorithm based on a purely fork-join approach. The distributed-memory implementation is based on the ScaLAPACK library. The performance of our new codes compares favorably with competing factorizations available on both shared-memory and distributed-memory architectures.
Nathan Heavner, Francisco D. Igual, Gregorio Quintana-Ortí, Per-Gunnar Martinsson
ACM Trans. Math. Softw.2
2021 Low precision matrix multiplication for efficient deep learning in NVIDIA Carmel processors
Pablo San Juan, Rafael Rodríguez-Sánchez 0001, Francisco D. Igual, Pedro Alonso 0002, Enrique S. Quintana-Ortí
J. Supercomput.3
2020 Leveraging knowledge-as-a-service (KaaS) for QoS-aware resource management in multi-user video transcoding
Luis Costero, Francisco D. Igual, Katzalin Olcoz, Francisco Tirado
J. Supercomput.2
2020 STEEL-RT: combining single task-single executor model and expanded scheduling to ease heterogeneity exploitation
Antón Rey, Francisco D. Igual, Manuel Prieto 0001
J. Supercomput.2
2020 Integration and exploitation of intra-routine malleability in BLIS
Rafael Rodríguez-Sánchez 0001, Francisco D. Igual, Enrique S. Quintana-Ortí
J. Supercomput.2
2020 Resource Management for Power-Constrained HEVC Transcoding Using Reinforcement Learning
abstract
The advent of online video streaming applications and services along with the users' demand for high-quality contents require High Efficiency Video Coding (HEVC), which provides higher video quality and more compression at the cost of increased complexity. On one hand, HEVC exposes a set of dynamically tunable parameters to provide trade-offs among Quality-of-Service (QoS), performance, and power consumption of multi-core servers on the video providers' data center. On the other hand, resource management of modern multi-core servers is in charge of adapting system-level parameters, such as operating frequency and multithreading, to deal with concurrent applications and their requirements. Therefore, efficient multi-user HEVC streaming necessitates joint adaptation of application-and system-level parameters. Nonetheless, dealing with such a large and dynamic design space is challenging and difficult to address through conventional resource management strategies. Thus, in this work, we develop a multi-agent Reinforcement Learning framework to jointly adjust application-and system-level parameters at runtime to satisfy the QoS of multi-user HEVC streaming in power-constrained servers. In particular, the design space, composed of all design parameters, is split into smaller independent sub-spaces. Each design sub-space is assigned to a particular agent so that it can explore it faster, yet accurately. The benefits of our approach are revealed in terms of adaptability and quality (with up to to 4× improvements in terms of QoS when compared to a static resource management scheme), and learning time (6× fasterthan an equivalent mono-agent implementation). Finally, we show that the power-capping techniques formulated outperform the hardware-based power capping with respect to quality.
Luis Costero, Arman Iranfar, Marina Zapater, Francisco D. Igual, Katzalin Olcoz, David Atienza 0001
IEEE Trans. Parallel Distributed Syst.4
2019 MAMUT: Multi-Agent Reinforcement Learning for Efficient Real-Time Multi-User Video Transcoding
abstract
Real-time video transcoding has recently raised as a valid alternative to address the ever-increasing demand for video contents in servers' infrastructures in current multi-user environments. High Efficiency Video Coding (HEVC) makes efficient online transcoding feasible as it enhances user experience by providing the adequate video configuration, reduces pressure on the network, and minimizes inefficient and costly video storage. However, the computational complexity of HEVC, together with its myriad of configuration parameters, raises challenges for power management, throughput control, and Quality of Service (QoS) satisfaction. This is particularly challenging in multi-user environments where multiple users with different resolution demands and bandwidth constraints need to be served simultaneously. In this work, we present MAMUT, a multi-agent machine learning approach to tackle these challenges. Our proposal breaks the design space composed of run-time adaptation of the transcoder and system parameters into smaller sub-spaces that can be explored in a reasonable time by individual agents. While working cooperatively, each agent is in charge of learning and applying the optimal values for internal HEVC and system-wide parameters. In particular, MAMUT dynamically tunes Quantization Parameter, selects number of threads per video, and sets the operating frequency with throughput and video quality objectives under compression and power consumption constraints. We implement MAMUT on an enterprise multicore server and compare equivalent scenarios to state-of-the-art alternative approaches. The obtained results reveal that MAMUT consistently attains up to 8× improvement in terms of FPS violations (and thus Quality of Service), 24% power reduction, as well as faster and more accurate adaptation both to the video contents and available resources.
Luis Costero, Arman Iranfar, Marina Zapater, Francisco D. Igual, Katzalin Olcoz, David Atienza 0001
DATE4
2019 Practical Considerations for Acoustic Source Localization in the IoT Era: Platforms, Energy Efficiency, and Performance
abstract
The rapid development of the Internet of Things (IoT) has posed important changes in the way emerging acoustic signal processing applications are conceived. While traditional acoustic processing applications have been developed taking into account high-throughput computing platforms equipped with expensive multichannel audio interfaces, the IoT paradigm is demanding the use of more flexible and energy-efficient systems. In this context, algorithms for source localization and ranging in wireless acoustic sensor networks can be considered an enabling technology for many IoT-based environments, including security, industrial, and health-care applications. This paper is aimed at evaluating important aspects dealing with the practical deployment of IoT systems for acoustic source localization. Recent systems-on-chip composed of low-power multicore processors, combined with a small graphics accelerator (or GPU), yield a notable increment of the computational capacity needed in intensive signal processing algorithms while partially retaining the appealing low power consumption of embedded systems. Different algorithms and implementations over several state-of-the-art platforms are discussed, analyzing important aspects, such as the tradeoffs between performance, energy efficiency, and exploitation of parallelism by taking into account real-time constraints.
Jose A. Belloch, José M. Badía, Francisco D. Igual, Maximo Cobos
IEEE Internet Things J.3
2019 Portability Study of an OpenCL Algorithm for Automatic Target Detection in Hyperspectral Images
abstract
In the last decades, the problem of target detection has received considerable attention in remote sensing applications. When this problem is tackled using hyperspectral images with hundreds of bands, the use of high-performance computing (HPC) is essential. One of the most popular algorithms in the hyperspectral image analysis community for this purpose is the automatic target detection and classification algorithm (ATDCA). Previous research has already investigated the mapping of ATDCA on HPC platforms such as multicore processors, graphics processing units (GPUs), and field-programmable gate arrays (FPGAs), showing impressive speedup factors (after careful fine-tuning) that allow for its exploitation in time-critical scenarios. However, the lack of standardization resulted in most implementations being too specific to a given architecture, eliminating (or at least making extremely difficult) code reusability across different platforms. In order to address this issue, we present a portability study of an implementation of ATDCA developed using the open computing language (OpenCL). We focus on cross-platform parameters such as performance, energy consumption, and code design complexity, as compared to previously developed (hand-tuned) implementations. Our portability study analyzes different strategies to expose data parallelism as well as enable the efficient exploitation of complex memory hierarchies in heterogeneous devices. We also conduct an assessment of energy consumption and discuss metrics to analyze the quality of our code. The conducted experiments-using synthetic and real hyperspectral data sets collected by the Hyperspectral Digital Imagery Collection Experiment (HYDICE) and NASA's Airborne Visible Infra-Red Imaging Spectrometer (AVIRIS)-demonstrate, for the first time in the literature, that portability across different HPC platforms can be achieved for real-time target detection in hyperspectral missions.
Sergio Bernabé, Carlos García 0001, Francisco D. Igual, Guillermo Botella Juan, Manuel Prieto 0001, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.3
2019 Accelerating the SRP-PHAT algorithm on multi- and many-core platforms using OpenCL
José M. Badía, Jose A. Belloch, Maximo Cobos, Francisco D. Igual, Enrique S. Quintana-Ortí
J. Supercomput.4
2019 Variable intra-task threading for power-constrained performance and energy optimization in DAG scheduling
Antón Rey, Francisco D. Igual, Manuel Prieto 0001
J. Supercomput.2
2019 Algorithm 994: Fast Implementations of the Brouwer-Zimmermann Algorithm for the Computation of the Minimum Distance of a Random Linear Code
abstract
The minimum distance of an error-correcting code is an important concept in information theory. Hence, computing the minimum distance of a code with a minimum computational cost is crucial to many problems in this area. In this article, we present and assess a family of implementations of both the brute-force algorithm and the Brouwer-Zimmermann algorithm for computing the minimum distance of a random linear code over F 2 that are faster than current implementations, both in the commercial and public domain. In addition to the basic sequential implementations, we present parallel and vectorized implementations that produce high performances on modern architectures. The attained performance results show the benefits of the developed optimized algorithms, which obtain remarkable improvements compared with state-of-the-art implementations widely used nowadays.
Fernando Hernando, Francisco D. Igual, Gregorio Quintana-Ortí
ACM Trans. Math. Softw.2
2017 Performance-Power Evaluation of an OpenCL Implementation of the Simplex Growing Algorithm for Hyperspectral Unmixing
abstract
Over the last few years, several new strategies for spectral unmixing of remotely sensed hyperspectral data have been proposed. Many of them have been developed to solve the most time-consuming and relevant step: endmember extraction. However, unmixing algorithms can be computationally very expensive in terms of processing time and energy consumption, a fact that compromises their use in applications under real-time and energy/power constraints. In this letter, we present a new parallel simplex growing algorithm (SGA) for hyperspectral data which exploits the memory hierarchy with operations in single-precision floating point. Those optimizations accelerate the most time-consuming parts of this method using the open computing language (OpenCL) standard. We have evaluated the performance versus energy consumption using the same open standard for parallel programming over a diverse set of heterogeneous platforms. Experiments have been conducted using real hyperspectral images collected by NASA's Airborne Visible Infrared Imaging Spectrometer and a collection of 24 synthetic hyperspectral images simulated with different sizes and number of endmembers (10-30). Considering the power consumption and OpenCL across all the proposed devices, the analysis presented indicates that the SGA can now be executed in computationally efficient fashion, which was not possible before introducing the parallel implementation described in this letter.
Sergio Bernabé, Guillermo Botella Juan, Jose M. R. Navarro, Carlos Orueta, Francisco D. Igual, Manuel Prieto 0001, Antonio Plaza
IEEE Geosci. Remote. Sens. Lett.5
2017 Revisiting conventional task schedulers to exploit asymmetry in multi-core architectures for dense linear algebra operations
Luis Costero, Francisco D. Igual, Katzalin Olcoz, Sandra Catalán, Rafael Rodríguez-Sánchez 0001, Enrique S. Quintana-Ortí
Parallel Comput.2
2017 Solving Weighted Least Squares (WLS) problems on ARM-based architectures
Jose A. Belloch, Balázs Bank, Francisco D. Igual, Enrique S. Quintana-Ortí, Antonio M. Vidal
J. Supercomput.3
2017 Time and energy modeling of a high-performance multi-threaded Cholesky factorization
Sandra Catalán, Francisco D. Igual, Rafael Mayo 0002, Rafael Rodríguez-Sánchez 0001, Enrique S. Quintana-Ortí
J. Supercomput.2
2016 HeSP: A Simulation Framework for Solving the Task Scheduling-Partitioning Problem on Heterogeneous Architectures
Antón Rey, Francisco D. Igual, Manuel Prieto 0001
Euro-Par2
2016 Analytical Modeling Is Enough for High-Performance BLIS
abstract
We show how the BLAS-like Library Instantiation Software (BLIS) framework, which provides a more detailed layering of the GotoBLAS (now maintained as OpenBLAS) implementation, allows one to analytically determine tuning parameters for high-end instantiations of the matrix-matrix multiplication. This is of both practical and scientific importance, as it greatly reduces the development effort required for the implementation of the level-3 BLAS while also advancing our understanding of how hierarchically layered memories interact with high-performance software. This allows the community to move on from valuable engineering solutions (empirically autotuning) to scientific understanding (analytical insight).
Tze Meng Low, Francisco D. Igual, Tyler M. Smith, Enrique S. Quintana-Ortí
ACM Trans. Math. Softw.2
2016 The BLIS Framework: Experiments in Portability
abstract
BLIS is a new software framework for instantiating high-performance BLAS-like dense linear algebra libraries. We demonstrate how BLIS acts as a productivity multiplier by using it to implement the level-3 BLAS on a variety of current architectures. The systems for which we demonstrate the framework include state-of-the-art general-purpose, low-power, and many-core architectures. We show, with very little effort, how the BLIS framework yields sequential and parallel implementations that are competitive with the performance of ATLAS, OpenBLAS (an effort to maintain and extend the GotoBLAS), and commercial vendor implementations such as AMD’s ACML, IBM’s ESSL, and Intel’s MKL libraries. Although most of this article focuses on single-core implementation, we also provide compelling results that suggest the framework’s leverage extends to the multithreaded domain.
Field G. Van Zee, Tyler M. Smith, Bryan Marker, Tze Meng Low, Robert A. van de Geijn, Francisco D. Igual, Mikhail Smelyanskiy, Xianyi Zhang, Michael Kistler, Vernon Austel, John A. Gunnels, Lee Killough
ACM Trans. Math. Softw.6
2014 Enhancing performance and energy consumption of runtime schedulers for dense linear algebra
abstract
SUMMARY The road towards Exascale Computing requires a holistic effort to address three different challenges simultaneously: high performance, energy efficiency, and programmability. The use of runtime task schedulers to orchestrate parallel executions with minimal developer intervention has been introduced in recent years to tackle the programmability issue while maintaining, or even improving, performance. In this paper, we enhance the SuperMatrix runtime task scheduler integrated in the libflame library in two different directions that address high performance and energy efficiency. First, we extend the runtime by accommodating hybrid parallel executions and managing task priorities for dense linear algebra operations, with remarkable performance improvements. Second, we introduce techniques to reduce energy consumption during idle times inherent to parallel executions, attaining important energy savings. In addition, we propose a power consumption model that can be leveraged by runtime task schedulers to make decisions based not only on performance but also on energy considerations. Copyright © 2014 John Wiley & Sons, Ltd.
Pedro Alonso 0002, Manuel F. Dolz, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí
Concurr. Comput. Pract. Exp.3
2013 Non-negative matrix factorization on low-power architectures: a comparative study
abstract
Power consumption is emerging as one of the main concerns in the High Performance Computing (HPC) field. Many bioinformatics applications require HPC techniques and parallel architectures to meet performance requirements, but at the same time they can be severely limited by energy consumption restrictions. In this paper, we perform an empirical study of an optimized implementation of the Nonnegative Matrix Factorization (NMF), that is widely used in many fields of bioinformatics. We target different types of architectures, including general-purpose, low-power embedded processors and specific-purpose architectures like graphics processors and digital signal processors. From our study, we gain insights in both performance and energy consumption for each one of them under given experimental conditions, and conclude that the most appropriate architecture is usually a trade-off between performance and power consumption for a given experiment and dataset.
Carlos García 0001, Francisco D. Igual, Guillermo Botella Juan, Manuel Prieto 0001, Francisco Tirado
EuroMPI2
2013 Scheduling algorithms-by-blocks on small clusters
abstract
SUMMARY The arrival of multicore architectures has generated an interest in reformulating dense matrix computations as algorithms‐by‐blocks, where submatrices are units of data and computations with those blocks are units of computation. Rather than directly executing such an algorithm, a directed acyclic graph is generated at runtime that is then scheduled by a runtime system such as SuperMatrix. The benefit is a clear separation of concerns between the library and the heuristics for scheduling. In this paper, we show that this approach can be taken one step further using the same methodology and an ad hoc runtime to map algorithms‐by‐blocks to small clusters. With no change to the library code, and the application that uses it, the computational power of such small clusters can be utilized. An impressive performance on a number of small clusters is reported. As a proof of the flexibility of the solution, we report performance results on accelerated clusters based on graphics processors. We believe this to be a possible step towards programming many‐core architectures, as demonstrated by a port of the solution to Intel's Single‐chip Cloud Computer (Intel, Santa Clara, CA, USA). Copyright © 2012 John Wiley & Sons, Ltd.
Francisco D. Igual, Gregorio Quintana-Ortí, Robert A. van de Geijn
Concurr. Comput. Pract. Exp.1
2012 Reducing Energy Consumption of Dense Linear Algebra Operations on Hybrid CPU-GPU Platforms
abstract
We investigate the balance between the time-to-solution and the energy consumption of a task-parallel execution of the Cholesky and LU factorizations on a hybrid platform, equipped with a multi-core processor and several GPUs. To improve energy efficiency, we incorporate two energy-saving techniques in the runtime in charge of scheduling the computations, to block idle threads and enable the transition to a more energy-friendly state of the general-purpose cores. Experiments on an Intel Xeon-based platform connected to an NVIDIA Tesla server report an average reduction of the energy consumption close to 9% (38% when only the consumption associated with the application is considered), for a minor increase in the execution time of the algorithm.
Pedro Alonso 0002, Manuel F. Dolz, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí
ISPA3
2012 Saving Energy in the LU Factorization with Partial Pivoting on Multi-core Processors
abstract
In this paper we analyze the trade-off between energy and performance for a data-parallel execution of the LU factorization with partial pivoting on a multi-core processor. To improve energy efficiency, we adapt the runtime in charge of controlling the concurrent execution of the algorithm to leverage DVFS and block idle threads. For a CPU-bounded operation like the LU factorization, experiments on an AMD 8-core processor report a reduction around 5% in energy consumption for the largest problem sizes in exchange for a minor increase in the execution time.
Pedro Alonso 0002, Manuel F. Dolz, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí
PDP3
2012 Level-3 BLAS on the TI C6678 Multi-core DSP
abstract
Digital Signal Processors (DSP) are commonly employed in embedded systems. The increase of processing needs in cellular base-stations, radio controllers and industrial/medical imaging systems, has led to the development of multi-core DSPs as well as inclusion of floating point operations while maintaining low power dissipation. The eight-core DSP from Texas Instruments, codenamed TMS320C6678, provides a peak performance of 128 GFLOPS (single precision) and an effective 32 GFLOPS(double precision) for only 10 watts. In this paper, we present the first complete implementation and report performance of the Level-3 Basic Linear Algebra Subprograms(BLAS) routines for this DSP. These routines are first optimized for single core and then parallelized over the different cores using OpenMP constructs. The results show that we can achieve about 8 single precision GFLOPS/watt and 2.2double precision GFLOPS/watt for General Matrix-Matrix multiplication (GEMM). The performance of the rest of theLevel-3 BLAS routines is within 90% of the corresponding GEMM routines.
Eric Stotzer, Francisco D. Igual, Robert A. van de Geijn
SBAC-PAD3
2012 Unleashing the high-performance and low-power of multi-core DSPs for general-purpose HPC
abstract
Take a multicore Digital Signal Processor (DSP) chip designed for cellular base stations and radio network controllers, add floating-point capabilities to support 4G networks, and out of thin air a HPC engine is born. The potential for HPC is clear: It promises 128 GFLOPS (single precision) for 10 Watts; It is used in millions of network related devices and hence benefits from economies of scale; It should be simpler to program than a GPU. Simply put, it is fast, green, and cheap. But is it easy to use? In this paper, we show how this potential can be applied to general-purpose high performance computing, more specifically to dense matrix computations, without major changes in existing codes and methodologies, and with excellent performance and power consumption numbers.
Francisco D. Igual, Arnon Friedmann, Eric Stotzer, Timothy Wentz, Robert A. van de Geijn
SC1
2012 The FLAME approach: From dense linear algebra algorithms to high-performance multi-accelerator implementations
Francisco D. Igual, Ernie Chan, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí, Robert A. van de Geijn, Field G. Van Zee
J. Parallel Distributed Comput.1
2012 A Runtime System for Programming Out-of-Core Matrix Algorithms-by-Tiles on Multithreaded Architectures
abstract
Out-of-core implementations of algorithms for dense matrix computations have traditionally focused on optimal use of memory so as to minimize I/O, often trading programmability for performance. In this article we show how the current state of hardware and software allows the programmability problem to be addressed without sacrificing performance. This comes from the realizations that memory is cheap and large, making it less necessary to optimally orchestrate I/O, and that new algorithms view matrices as collections of submatrices and computation as operations with those submatrices. This enables libraries to be coded at a high level of abstraction, leaving the tasks of scheduling the computations and data movement in the hands of a runtime system. This is in sharp contrast to more traditional approaches that leverage optimal use of in-core memory and, at the expense of introducing considerable programming complexity, explicit overlap of I/O with computation. Performance is demonstrated for this approach on multicore architectures as well as platforms equipped with hardware accelerators.
Gregorio Quintana-Ortí, Francisco D. Igual, Mercedes Marqués 0001, Enrique S. Quintana-Ortí, Robert A. van de Geijn
ACM Trans. Math. Softw.2
2011 Condensed forms for the symmetric eigenvalue problem on multi-threaded architectures
abstract
Abstract We investigate the performance of the routines in LAPACK and the Successive Band Reduction (SBR) toolbox for the reduction of a dense matrix to tridiagonal form, a crucial preprocessing stage in the solution of the symmetric eigenvalue problem, on general‐purpose multi‐core processors. In response to the advances of hardware accelerators, we also modify the code in the SBR toolbox to accelerate the computation by off‐loading a significant part of the operations to a graphics processor (GPU). The performance results illustrate the parallelism and scalability of these algorithms on current high‐performance multi‐core and many‐core architectures. Copyright © 2010 John Wiley & Sons, Ltd.
Paolo Bientinesi, Francisco D. Igual, Daniel Kressner, Matthias Petschow, Enrique S. Quintana-Ortí
Concurr. Comput. Pract. Exp.2
2009 An Extension of the StarSs Programming Model for Platforms with Multiple GPUs
Eduard Ayguadé, Rosa M. Badia, Francisco D. Igual, Jesús Labarta, Rafael Mayo 0002, Enrique S. Quintana-Ortí
Euro-Par3
2009 Fast development of dense linear algebra codes on graphics processors
abstract
We present an application programming interface (API) for the C programming language that facilitates the development of dense linear algebra algorithms on graphics processors applying the FLAME methodology. The interface, built on top of the NVIDIA CUBLAS library, implements all the computational functionality of the FLAME/C interface. In addition, the API includes data transference routines to explicitly handle communication between the CPU and GPU memory spaces. The flexibility and simplicity-of-use of this tool are illustrated using a complex operation of dense linear algebra: the Cholesky factorization. For this operation, we implement and evaluate all existing variants on an NVIDIA G80 processor, attaining speedups 7times compared with the CPU implementations.
M. Jesús Zafont, Alberto F. Martín, Francisco D. Igual, Enrique S. Quintana-Ortí
IPDPS3
2009 Solving dense linear systems on platforms with multiple hardware accelerators
abstract
In a previous PPoPP paper we showed how the FLAME methodology, combined with the SuperMatrix runtime system, yields a simple yet powerful solution for programming dense linear algebra operations on multicore platforms. In this paper we provide further evidence that this approach solves the programmability problem for this domain by targeting a more complex architecture, composed of a multicore processor and multiple hardware accelerators (GPUs, Cell B.E., etc.), each with its own local memory, resulting in a platform more reminiscent of a heterogeneous distributed-memory system. In particular, we show that the FLAME programming model accommodates this new situation effortlessly so that no significant change needs to be made to the codebase. All complexity is hidden inside the SuperMatrix runtime scheduling mechanism, which incorporates software implementations of standard cache/memory coherence techniques in computer architecture to improve the performance. Our experimental evaluation on a Intel Xeon 8-core host linked to an NVIDIA Tesla S870 platform with four GPUs delivers peak performances around 550 and 450 (single-precision) GFLOPS for the matrix-matrix product and the Cholesky factorization, respectively, which we believe to be the best performance numbers posted on this new architecture for such operations.
Gregorio Quintana-Ortí, Francisco D. Igual, Enrique S. Quintana-Ortí, Robert A. van de Geijn
PPoPP2
2009 Exploiting the capabilities of modern GPUs for dense matrix computations
abstract
Abstract We present several algorithms to compute the solution of a linear system of equations on a graphics processor (GPU), as well as general techniques to improve their performance, such as padding and hybrid GPU‐CPU computation. We compare single and double precision performance of a modern GPU with unified architecture, and show how iterative refinement with mixed precision can be used to regain full accuracy in the solution of linear systems, exploiting the potential of the processor for single precision arithmetic. Experimental results on a GTX280 using CUBLAS 2.0, the implementation of BLAS for NVIDIA® GPUs with unified architecture, illustrate the performance of the different algorithms and techniques proposed. Copyright © 2009 John Wiley & Sons, Ltd.
Sergio Barrachina 0001, María Isabel Castillo, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí
Concurr. Comput. Pract. Exp.3
2008 Solving Dense Linear Systems on Graphics Processors
Sergio Barrachina 0001, María Isabel Castillo, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí
Euro-Par3
2008 Biomedical image analysis on a cooperative cluster of GPUs and multicores
abstract
We are currently witnessing the emergence of two paradigms in parallel computing: streaming processing and multi-core CPUs. Represented by solid commercial products widely available in commodity PCs, GPUs and multi-core CPUs bring together an unprecedented combination of high performance at low cost. The scientific computing community needs to keep pace with application models and middleware which scale efficiently to hundreds of internal processing units. The purpose of the work we present here is twofold: first, a cooperative environment is designed so that both parallel models can coexist and complement one another. Second, beyond the parallelism of multiple internal cores, further parallelism is introduced when multiple CPU sockets, multiple GPUs, and multiple nodes are combined within a unique multi-processor platform which exceeds 10 TFLOPS when using 16 nodes. We illustrate our cooperative parallelization approach by implementing a large-scale, biomedical image analysis application which contains a number of assorted kernels including typical streaming operators, co-occurrence matrices, convolutions, and histograms. Experimental results are compared among different implementation strategies and almost linear speed-up is achieved when all coexisting methods in CPUs and GPUs are combined.
Timothy D. R. Hartley, Ümit V. Çatalyürek, Antonio Ruiz 0001, Francisco D. Igual, Rafael Mayo 0002, Manuel Ujaldon
ICS4
2008 Evaluation and tuning of the Level 3 CUBLAS for graphics processors
abstract
The increase in performance of the last generations of graphics processors (GPUs) has made this class of platform a coprocessing tool with remarkable success in certain types of operations. In this paper we evaluate the performance of the Level 3 operations in CUBLAS, the implementation of BIAS for NVIDIAreg GPUs with unified architecture. From this study, we gain insights on the quality of the kernels in the library and we propose several alternative implementations that are competitive with those in CUBLAS. Experimental results on a GeForce 8800 Ultra compare the performance of CUBLAS and the new variants.
Sergio Barrachina 0001, María Isabel Castillo, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí
IPDPS3