Rafael Rodríguez-Sánchez 0001

dblp:51/8228 · DBLP profile ↗
← Back
24ranked-venue papers
9as first author
6since 2021 · last 2026
0000-0001-8789-3953ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author
YearPublicationVenuePosition
2026 A comparative performance and efficiency analysis of Apple's M architectures: A GEMM case study
abstract
This paper evaluates the performance and energy efficiency of Apple processors across multiple ARM-based M-series generations and models (standard and Pro). The study is motivated by the increasing heterogeneity of Apple’s SoC architectures, which integrate multiple computing engines raising the scientific question of which hardware components are best suited for executing general-purpose and domain-specific computations such as the GEneral Matrix Multiply ( GEMM ). The analysis focuses on four key components: the Central Processing Unit (CPU), the Graphics Processing Unit (GPU), the matrix calculation accelerator (AMX), and the Apple Neural Engine (ANE). The assessments use the GEMM as benchmark to characterize the performance of the CPU and GPU, alongside tests on AMX, which is specialized in handling large-scale mathematical operations, and tests on the ANE, which is specifically designed for Deep Learning purposes. Additionally, energy consumption data has been collected to analyze the energy efficiency of the aforementioned resources. Results highlight notable improvements in computational capacity and energy efficiency over successive generations. On one hand, the AMX stands out as the most efficient component for FP32 and FP64 workloads, significantly boosting overall system performance. In the M4 Pro, which integrates two matrix accelerators, it achieves up to 68% of the GPU’s FP32 performance while consuming only 42% of its power. On the other hand, the ANE, although limited to FP16 precision, excels in energy efficiency for low-precision tasks, surpassing other accelerators with over 700 GFLOPs/Watt under batched workloads. This analysis offers a clear understanding of how Apple’s custom ARM designs optimize both performance and energy use, particularly in the context of multi-core processing and specialized acceleration units. In addition, a significant contribution of this study is the comprehensive comparative analysis of Apple’s accelerators, which have previously been poorly documented and scarcely studied. The analysis spans different generations and compares the accelerators against both CPU and GPU performance.
Sandra Catalán, Rafael Rodríguez-Sánchez 0001, Carlos García 0001, Luis Piñuel Moreno
Future Gener. Comput. Syst.2
2024 Inference with Transformer Encoders on ARM and RISC-V Multicore Processors
abstract
Abstract We delve into the performance of transformer encoder inference on low-power multi-core processors from two perspectives: First, we conduct a detailed profile of the inference process for two members of the BERT family on a modern multi-core processor, identifying the main bottlenecks and opportunities for improvement. Second, we propose a number of accumulative optimisations for their primary building blocks. For that, we elaborate our own implementation of the general matrix multiplication (), which dynamically tunes several key parameters yielding relevant performance gains for transformer encoders. Additionally, we introduce a number of strategies to also improve the parallel execution of the transformer block. Our implementations for ARMv8a and RISC-V multi-core processors with SIMD units, taking as a reference state-of-the-art implementations (BLIS for ARM and OpenBLAS for RISC-V) reveal accelerations of up to $$2.5\times $$ 2.5 × for natural language processing tasks.
Héctor Martínez 0002, Francisco D. Igual, Rafael Rodríguez-Sánchez 0001, Sandra Catalán, Adrián Castelló 0001, Enrique S. Quintana-Ortí
Euro-Par (2)3
2023 Fine-grain task-parallel algorithms for matrix factorizations and inversion on many-threaded CPUs
abstract
Abstract We extend a two‐level task partitioning previously applied to the inversion of dense matrices via Gauss–Jordan elimination to the more challenging QR factorization as well as the initial orthogonal reduction to band form found in the singular value decomposition. Our new task‐parallel algorithms leverage the tasking mechanism currently available in OpenMP to exploit “nested” task parallelism, with a first outer level that operates on matrix panels and a second inner level that processes the matrix either by ‐panels or by tiles, in order to expose a large number of independent tasks. We present a detailed performance analysis, including execution traces, which shows that the two‐level refinement into fine grain tasks allows for an improved load balancing and delivers high performance on current general‐purpose many‐core processors (CPUs) from Intel and AMD.
Sandra Catalán, José R. Herrero 0001, Francisco D. Igual, Enrique S. Quintana-Ortí, Rafael Rodríguez-Sánchez 0001
Concurr. Comput. Pract. Exp.5
2023 Programming parallel dense matrix factorizations and inversion for new-generation NUMA architectures
abstract
We propose a methodology to address the programmability issues derived from the emergence of new-generation shared-memory NUMA architectures. For this purpose, we employ dense matrix factorizations and matrix inversion (DMFI) as a use case, and we target two modern architectures (AMD Rome and Huawei Kunpeng 920) that exhibit configurable NUMA topologies. Our methodology pursues performance portability across different NUMA configurations by proposing multi-domain implementations for DMFI plus a hybrid task- and loop-level parallelization that configures multi-threaded executions to fix core-to-data binding, exploiting locality at the expense of minor code modifications. In addition, we introduce a generalization of the multi-domain implementations for DMFI that offers support for virtually any NUMA topology in present and future architectures. Our experimentation on the two target architectures for three representative dense linear algebra operations validates the proposal, reveals insights on the necessity of adapting both the codes and their execution to improve data access locality, and reports performance across architectures and inter- and intra-socket NUMA configurations competitive with state-of-the-art message-passing implementations, maintaining the ease of development usually associated with shared-memory programming.
Sandra Catalán, Francisco D. Igual, José R. Herrero 0001, Rafael Rodríguez-Sánchez 0001, Enrique S. Quintana-Ortí
J. Parallel Distributed Comput.4
2022 NUMA-Aware Dense Matrix Factorizations and Inversion with Look-Ahead on Multicore Processors
abstract
We address the efficient design and implementation of dense matrix factorizations and inversion (DMFI) on modern multicore processors with several NUMA (non-uniform memory access) nodes. Our approach enhances the DMFI routines with a look-ahead strategy, in order to overcome the “panel factorization bottleneck”. In addition, it exploits both hybrid task- and loop-level parallelizations while taking into account the NUMA organization of the memory hierarchy. The experiments on a Huawei Kunpeng-based server, with two sockets and 48 cores per socket, for three representative dense linear algebra operations, expose the necessity of adapting both the codes and their execution environment parameters to improve data access locality. The results of these changes deliver performance across inter- and intra-socket NUMA configurations superior to that of reference implementations from state-of-the-art libraries for this platform.
Sandra Catalán, Francisco D. Igual, Rafael Rodríguez-Sánchez 0001, José R. Herrero 0001, Enrique S. Quintana-Ortí
SBAC-PAD3
2021 Low precision matrix multiplication for efficient deep learning in NVIDIA Carmel processors
Pablo San Juan, Rafael Rodríguez-Sánchez 0001, Francisco D. Igual, Pedro Alonso 0002, Enrique S. Quintana-Ortí
J. Supercomput.2
2020 Integration and exploitation of intra-routine malleability in BLIS
Rafael Rodríguez-Sánchez 0001, Francisco D. Igual, Enrique S. Quintana-Ortí
J. Supercomput.1
2019 Dynamic look-ahead in the reduction to band form for the singular value decomposition
Andrés Tomás, Rafael Rodríguez-Sánchez 0001, Sandra Catalán, Rocío Carratalá-Sáez, Enrique S. Quintana-Ortí
Parallel Comput.2
2018 Two-sided orthogonal reductions to condensed forms on asymmetric multicore processors
Pedro Alonso 0002, Sandra Catalán, José R. Herrero 0001, Enrique S. Quintana-Ortí, Rafael Rodríguez-Sánchez 0001
Parallel Comput.5
2018 Energy balance between voltage-frequency scaling and resilience for linear algebra routines on low-power multicore architectures
Sandra Catalán, José R. Herrero 0001, Enrique S. Quintana-Ortí, Rafael Rodríguez-Sánchez 0001
Parallel Comput.4
2018 Static scheduling of the LU factorization with look-ahead on asymmetric multicore processors
Sandra Catalán, José R. Herrero 0001, Enrique S. Quintana-Ortí, Rafael Rodríguez-Sánchez 0001
Parallel Comput.4
2017 Revisiting conventional task schedulers to exploit asymmetry in multi-core architectures for dense linear algebra operations
Luis Costero, Francisco D. Igual, Katzalin Olcoz, Sandra Catalán, Rafael Rodríguez-Sánchez 0001, Enrique S. Quintana-Ortí
Parallel Comput.5
2017 Time and energy modeling of a high-performance multi-threaded Cholesky factorization
Sandra Catalán, Francisco D. Igual, Rafael Mayo 0002, Rafael Rodríguez-Sánchez 0001, Enrique S. Quintana-Ortí
J. Supercomput.4
2015 Time and energy modeling of an INTRA-ONLY HEVC encoder
abstract
In this paper, we present precise time and energy models for an intra-only HEVC video encoder. These models are a step forward to understand and estimate the computational complexity and energy demands of an HEVC encoder, which in turn opens the path to finely tuning the computational resources that are dedicated to this purpose. Our models estimate the complexity and energy consumed by the HEVC encoder, in a frame by frame basis, considering two factors: the quantification parameter used to encode each frame and the spatial information of that frame. Our experimental validation demonstrates the accuracy of these models, which report errors that are, on average, below 10% for full HD videos, and 5% for 832 × 480 videos.
Rafael Rodríguez-Sánchez 0001, Maria Teresa Alonso, José Luis Martínez 0001, Rafael Mayo 0002, Enrique S. Quintana-Ortí
VCIP1
2015 HEVC to VP9 transcoder
abstract
HEVC and VP9 are the current state-of-the-art in video compression, thus, it is expected that in the near future these new codecs will replace their predecessors. However, the process of converting video contents compressed with one standard to those using another standard is highly computationally expensive, since a priori the video contents must be decompressed and compressed with the target video encoder. Nevertheless, it is known that some information can be extracted from the decoding process in order to accelerate the encoding process. In this paper, a technique for transcoding from HEVC to VP9 is presented. By using some information extracted from the HEVC decoding process some coding computations will be discarded from being checked in the VP9 encoder. By applying the proposed approaches, a reduction of about 38% of the coding complexity is achieved with acceptable Rate Distortion penalties.
Enrique de la Torre, Rafael Rodríguez-Sánchez 0001, José Luis Martínez 0001
VCIP2
2014 An smpUMHexagonS-based motion estimation algorithm for heterogeneous architectures
abstract
Simplified Unsymmetrical Multi-Hexagon Search is one of the faster motion estimation algorithms for video coding available in the literature. It achieves a very good tradeoff between computational complexity and coding efficiency. In fact, the H.264/AVC reference software includes it, and it could easily be extended to the new video coding standard: High Efficiency Video Coding. This paper proposes a parallel algorithm which implements the Simplified Unsymmetrical Multi-Hexagon motion search in a graphic processing unit which serves as co-processor for the central processing unit. The results show a negligible rate distortion drop with a speed-up of up to 8× for the motion estimation module compared to the sequential version. Furthermore, the proposed algorithm is evaluated against some related proposals available in the literature, outperforming all of them in terms of time reduction as well as coding efficiency.
Rafael Rodríguez-Sánchez 0001, José Luis Martínez 0001
ICIP2
2013 Low delay H.264/AVC bidirectional inter prediction on a GPU
abstract
The H.264/AVC video coding standard introduces some improved tools in order to increase compression efficiency. One of these new features is the variable block-size motion estimation. Moreover, H.264/AVC defines different prediction structures which include bi-directional predictions and, more recently, the hierarchical one. These structures also includes new Group Of Picture patterns which outperform the traditional ones. In the literature, several techniques have been proposed over the last few years which are aimed at accelerating the traditional inter prediction process, but there are no many works focusing on bidirectional and hierarchical predictions. In this paper, with the emergence of manycore processors or accelerators, a step forward is taken towards an implementation of an H.264/AVC inter prediction algorithm on a Graphics Processing Unit. The results show a negligible rate distortion drop with a time reduction of up to 92% for the complete H.264/AVC encoder.
Rafael Rodríguez-Sánchez 0001, José Luis Martínez 0001, Jan De Cock, José L. Sánchez 0002, José M. Claver, Rik Van de Walle
ICIP1
2013 Fast transrating for high efficiency video coding based on machine learning
abstract
To incorporate the newly developed High Efficiency Video Coding (HEVC) standard in real-life network applications, efficient transrating algorithms are required. We propose a fast transrating scheme, based on the early prediction of the partition split-flags in P pictures. Using machine learning techniques, the correlation between co-located partitions at different quantizations is investigated. This results in a model which predicts the split-flag and gives the associated prediction accuracy so that the splitting process in the transcoder is optimized. At each partition depth, the model indicates whether the full rate-distortion cost evaluations should be performed at the current depth, or if the partition can be split immediately. Experimental results show that the proposed transcoder reduces the complexity of the transrating process by 76.04%, while maintaining the coding efficiency of a cascaded decoder-encoder.
Luong Pham Van, Jan De Cock, Glenn Van Wallendael, Sebastiaan Van Leuven, Rafael Rodríguez-Sánchez 0001, José Luis Martínez 0001, Peter Lambert, Rik Van de Walle
ICIP5
2013 H.264/AVC inter prediction on accelerator-based multi-core systems
Rafael Rodríguez-Sánchez 0001, José Luis Martínez 0001, Gerardo Fernández-Escribano, José L. Sánchez 0002, José M. Claver
Multim. Tools Appl.1
2013 H.264/AVC inter prediction for heterogeneous computing systems
Rafael Rodríguez-Sánchez 0001, José Luis Martínez 0001, Gerardo Fernández-Escribano, José M. Claver, José L. Sánchez 0002
J. Supercomput.1
2012 A Fast GPU-Based Motion Estimation Algorithm for HD 3D Video Coding
abstract
H.264/AVC is the commercial standard currently most in use for video and is based on single view (mono view). Recently, the video community has also standardized an H.264/AVC extension for supporting 3D video sensation which is referred to as Multiview Video Coding (MVC). Like H.264/AVC, MVC includes temporal and spatial prediction but also includes inter-view prediction as well as disparity estimation. Until now, in H.264/AVC the inter prediction techniques have been the most time-consuming tasks and, thus, in MVC with its new interview predictions the encoding time is even higher. This paper proposes a GPU-based algorithm for both temporal and interview prediction. The algorithm proposed is able to perform this complex prediction task by means of an efficient distribution of all the computations over the GPU and also tries to mitigate the sequential dependencies. The approach can achieve a remarkable time reduction of up to 98% with only a negligible loss in coding efficiency. Moreover, this paper shows that the proposed GPU-based algorithm is more energy efficient and thus, requires less energy than the sequential reference.
Rafael Rodríguez-Sánchez 0001, José Luis Martínez 0001, Gerardo Fernández-Escribano, José L. Sánchez 0002, José M. Claver
ISPA1
2012 A Fast GPU-Based Motion Estimation Algorithm for H.264/AVC
Rafael Rodríguez-Sánchez 0001, José Luis Martínez 0001, Gerardo Fernández-Escribano, José L. Sánchez 0002, José M. Claver
MMM1
2012 Optimizing H.264/AVC interprediction on a GPU-based framework
abstract
SUMMARY H.264/MPEG‐4 part 10 is the latest standard for video compression and promises a significant advance in terms of quality and distortion compared with the commercial standards currently most in use such as MPEG‐2 or MPEG‐4. To achieve this better performance, H.264 adopts a large number of new/improved compression techniques compared with previous standards, albeit at the expense of higher computational complexity. In addition, in recent years new hardware accelerators have emerged, such as graphics processing units (GPUs), which provide a new opportunity to reduce complexity for a large variety of algorithms. However, current GPUs suffer from higher power consumption requirements because of its design. Up to now, GPU‐based software developers have not taken this into account. In this paper, we present a detailed procedure to implement the H.264 motion estimation for a GPU, with the aim of reducing time and, as a consequence, the energy consumption. The results show a negligible drop in rate distortion with a time reduction of over 91.5% on average and it reduces the energy consumption by a factor of 11.78 compared with the reference implementation. Copyright © 2011 John Wiley & Sons, Ltd.
Rafael Rodríguez-Sánchez 0001, José Luis Martínez 0001, Gerardo Fernández-Escribano, José L. Sánchez 0002, José M. Claver, Pedro Diaz
Concurr. Comput. Pract. Exp.1
2011 Reducing complexity in H.264/AVC motion estimation by using a GPU
abstract
H.264/AVC applies a complex mode decision technique that has high computational complexity in order to reduce the temporal redundancies of video sequences. Several algorithms have been proposed in the literature in recent years with the aim of accelerating this part of the encoding process. Recently, with the emergence of many-core processors or accelerators, a new approach can be adopted for reducing the complexity of the H.264/AVC encoding algorithm. This paper focuses on reducing the inter prediction complexity adopted in H.264/AVC and proposes a GPU-based implementation using CUDA. Experimental results show that the proposed approach reduces the complexity by as much as 99% (100x of speedup) while maintaining the coding efficiency.
Rafael Rodríguez-Sánchez 0001, José Luis Martínez 0001, Gerardo Fernández-Escribano, José M. Claver, José L. Sánchez 0002
MMSP1