Héctor Martínez 0002

dblp:40/9854-2 · also Héctor Martínez Pérez 0002 · DBLP profile ↗
← Back
24ranked-venue papers
9as first author
20since 2021 · last 2026
0000-0001-5891-4479ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 5 first-author · 12 since 2021Computer networks · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 The cambrian explosion of mixed-precision matrix multiplication for quantized deep learning inference
abstract
Recent advances in deep learning (DL) have promoted to a shift from traditional 64-bit floating point (FP64) arithmetic for scientific computing toward reduced-precision formats–such as FP16, BF16, or even 8-bit integers–combined with mixed-precision arithmetic. This transition enhances computational throughput, reduces memory and bandwidth usage, and improves energy efficiency, offering significant advantages for resource-constrained edge devices. To support this shift, hardware architectures have evolved accordingly, now including adapted ISAs (Instruction Set Architectures) that expose mixed-precision vector units and matrix engines tailored for DL workloads. At the heart of many DL and scientific computing tasks is the general matrix-matrix multiplication ( GEMM ), a fundamental kernel historically optimized using fused multiply-add (FMA) vector instructions on SIMD (single instruction, multiple data) units. However, as hardware moves toward mixed-precision dot (or inner)-product-centric operations optimized for quantized inference, these legacy approaches are being phased out. In response to this, our paper revisits the conventional, high-performance implementation of GEMM and describes strategies for adapting it to mixed integer precision (MIP) arithmetic across modern ISAs, including x86_64, Arm, and RISC-V. Concretely, we illustrate novel micro-kernel designs and data layouts that better exploit today’s specialized hardware and demonstrate significant performance gains from MIP arithmetic over floating-point implementations across three representative CPUs. These contributions highlight a new era of GEMM optimization-driven by the demands of DL inference on heterogeneous architectures, marking what we term as the “Cambrian period” for matrix multiplication.
Héctor Martínez 0002, Adrián Castelló 0001, Francisco D. Igual, Enrique S. Quintana-Ortí
Future Gener. Comput. Syst.1
2026 In-memory operators for medical image processing
abstract
Medical-image processing (MIP) frequently faces challenges related to computational efficiency and memory bandwidth, primarily due to the intensive data movement between processing units and memory. This work explores the emerging paradigm of Processing-in-Memory (PIM) to alleviate these data movement bottlenecks in MIP. It presents the first PIM implementation of five fundamental algorithms widely used in MIP: voxel-counting, thresholding, histogram computation, convolution, and interpolation, outlining specific PIM patterns. The algorithms, implemented using the UPMEM PIM architecture, were evaluated in real non-commercial PIM hardware (20 DDR4-2400 PIM modules providing 160 GB PIM memory), using both synthetic and real data sets of varying image sizes and underlying datatypes (INT8, INT32, FP32), thus covering a wide range of applications. The evaluation results indicate that, for data-intensive tasks, the PIM prototype can improve significantly the computational efficiency over traditional commercial CPU 20 × , and GPU 3 × . This research highlights the potential of PIM for revolutionizing MIP applications by enabling faster and more energy-efficient processing of medical images, thereby addressing critical needs in clinical and research applications.
Héctor Martínez 0002, Juan Gómez-Luna, Rafael Palomar, Joaquín Olivares 0001
Future Gener. Comput. Syst.1
2025 Distributed Fog Computing for Real-Time Surveillance
Joaquín Olivares 0001, Héctor Martínez 0002, Fernando León-García, José M. Palomares
Networking2
2025 Portable, High Performance Matrix Multiplication Micro-Kernels for RISC-V with ExO
abstract
The proliferation of RISC-V platforms and their use in a wide variety of scientific applications, including deep learning scenarios, has dramatically increased the interest to generate optimized code for them. In the field of HPC (High Performance Computing), the RISCV ISA (Instruction Set Architecture) has been adopted by a wide variety of designs with different micro-architecture; as a result, performance portability of existing codes is a major endeavor. Code generators and compilers such as Apache TVM, MLIR, or EXO provide a hardware abstraction for implementing optimized hardware-aware codes, thus reducing development time and potential errors. These generators can handle the full software stack, from basic micro-kernels to complex operations. In this work, we focus on the optimization of GEMM (general matrix-matrix multiplication), a key operation on top of which dense linear algebra libraries and deep learning frameworks are built. Specifically, we present an EXO-based GEMM microkernel generator for the RISC-V ISA with RVV vector extensions that addresses the lack of high-performance and portable GEMM micro-kernels. Our results demonstrate that, by generating a wide range of micro-kernels, one can obtain GEMM realizations that outperform those in the state-of-the-art high performance libraries.
Adrián Castelló 0001, Héctor Martínez 0002, Sandra Catalán, Jie Lei 0007, Yuka Ikarashi, Grace Dinh, Francisco D. Igual, Enrique S. Quintana-Ortí
PDP2
2025 Latency-Critical Quantized Inference With Transformer Decoders on ARM and RISC-V CPUs
abstract
Large language models are transforming industries but face challenges due to their high computational and energy demands. Model compression via quantization mitigates these barriers by reducing the bit precision of parameters and arithmetic operations, enabling deployment on resource-constrained devices like smartphones and edge platforms. This paper focuses on quantization applied to transformer decoders, which are critical for tasks such as text generation and conversational artificial intelligence. Unlike encoders, decoders are constrained by memory due to their sequential processing nature and low arithmetic intensity. We propose optimizations targeting inference on low-power CPUs, emphasizing efficient linear layers with quantized data/arithmetic and cache optimization. Using two representative ARM and RISC-V platforms, we present optimized mixed-precision implementations of the matrix multiplication that outperform the instance of that computational kernel in popular libraries such as BLIS, XNNPACK and ARMCL. This work thus advances the understanding of the impact of quantization on transformer decoder efficiency, energy consumption and precision in edge environments.
Héctor Martínez 0002, Sandra Catalán, Adrián Castelló 0001, José I. Mestre, Enrique S. Quintana-Ortí
IEEE Internet Things J.1
2025 Experience-guided, mixed-precision matrix multiplication with apache TVM for ARM processors
abstract
Abstract Deep learning (DL) generates new computational tasks that are different from those encountered in classical scientific applications. In particular, DL training and inference require general matrix multiplications (gemm) with matrix operands that are far from large and square as in other scientific fields. In addition, DL models gain arithmetic/storage complexity, and as a result, reduced precision via quantization is now mainstream for inferring DL models in edge devices. Automatic code generation addresses these new types of gemm by (1) improving portability between different hardware with only one base code; (2) supporting mixed and reduced precision; and (3) enabling auto-tuning methods that, given a base operation, perform a (costly) optimization search for the best schedule. In this paper, we rely on Apache TVM to generate an experience-guided gemm that provides performance competitive with the TVM auto-scheduler, while reducing tuning time by a factor of 48×.
Adrián Castelló 0001, Héctor Martínez 0002, Sandra Catalán, Francisco D. Igual, Enrique S. Quintana-Ortí
J. Supercomput.2
2024 Tackling the Matrix Multiplication Micro-Kernel Generation with Exo
abstract
The optimization of the matrix multiplication (or GEMM) has been a need during the last decades. This operation is considered the flagship of current linear algebra libraries such as BLIS, OpenBLAS, or Intel OneAPI because of its widespread use in a large variety of scientific applications. The GEMM is usually implemented following the GotoBLAS philosophy, which tiles the GEMM operands and uses a series of nested loops for performance improvement. These approaches extract the maximum computational power of the architectures through small pieces of hardware-oriented, high-performance code called micro-kernel. However, this approach forces developers to generate, with a nonnegligible effort, a dedicated micro-kernel for each new hardware. In this work, we present a step-by-step procedure for generating micro-kernels with the Exo compiler that perform close to (or even better than) manually developed microkernels written with intrinsic functions or assembly language. Our solution also improves the portability of the generated code, since a hardware target is fully specified by a concise library-based description of its instructions.
Adrián Castelló 0001, Julian Bellavita, Grace Dinh, Yuka Ikarashi, Héctor Martínez 0002
CGO5
2024 Inference with Transformer Encoders on ARM and RISC-V Multicore Processors
abstract
Abstract We delve into the performance of transformer encoder inference on low-power multi-core processors from two perspectives: First, we conduct a detailed profile of the inference process for two members of the BERT family on a modern multi-core processor, identifying the main bottlenecks and opportunities for improvement. Second, we propose a number of accumulative optimisations for their primary building blocks. For that, we elaborate our own implementation of the general matrix multiplication (), which dynamically tunes several key parameters yielding relevant performance gains for transformer encoders. Additionally, we introduce a number of strategies to also improve the parallel execution of the transformer block. Our implementations for ARMv8a and RISC-V multi-core processors with SIMD units, taking as a reference state-of-the-art implementations (BLIS for ARM and OpenBLAS for RISC-V) reveal accelerations of up to $$2.5\times $$ 2.5 × for natural language processing tasks.
Héctor Martínez 0002, Francisco D. Igual, Rafael Rodríguez-Sánchez 0001, Sandra Catalán, Adrián Castelló 0001, Enrique S. Quintana-Ortí
Euro-Par (2)1
2024 Communication-Avoiding Fusion of GEMM-Based Convolutions for Deep Learning in the RISC-V GAP8 MCU
abstract
Incorporating deep learning (DL) technologies to the edge is crucial for improving the security, privacy, and energy efficiency of the Internet of Things (IoT). In this scenario, the limitations of edge devices in terms of power dissipation, memory capacity, and processing power require a careful selection and optimization of algorithms for IoT DL applications. In this line, our work focuses on the convolution operator, a key component in deep neural networks for signal processing and computer vision. Specifically, the work aims at the efficient implementation of the lowering-based implementation of this operator, on the GAP8 parallel ultra-low power platform (PULP), with the goal of mitigating the data transfer costs across the memory hierarchy. Our contributions include 1) an analytical model for estimating the parallel execution time, 2) the exploration of different configuration options, and 3) four variants of the algorithm that fuse several components to address memory bottlenecks in the method. Overall, our best-fused variant provides a speedup of up to 1.25× over the baseline algorithm when applied to infer MobileNet-v1+ImageNet and VGG9+CIFAR10 using 8 threads, and up to 1.34× for ResNet18+ImageNet.
Cristián Ramírez, Adrián Castelló 0001, Héctor Martínez 0002, Enrique S. Quintana-Ortí
IEEE Internet Things J.3
2024 Distributed Fog computing system for weapon detection and face recognition
Héctor Martínez 0002, Francisco J. Rodríguez-Lozano, Fernando León-García, José M. Palomares, Joaquín Olivares 0001
J. Netw. Comput. Appl.1
2024 Parallel GEMM-based convolutions for deep learning on multicore ARM and RISC-V architectures
Héctor Martínez 0002, Sandra Catalán, Adrián Castelló 0001, Enrique S. Quintana-Ortí
J. Syst. Archit.1
2024 Automatic generation of ARM NEON micro-kernels for matrix multiplication
abstract
Abstract General matrix multiplication ( gemm ) is a fundamental kernel in scientific computing and current frameworks for deep learning. Modern realisations of gemm are mostly written in C, on top of a small, highly tuned micro-kernel that is usually encoded in assembly. The high performance realisation of gemm in linear algebra libraries in general include a single micro-kernel per architecture, usually implemented by an expert. In this paper, we explore a couple of paths to automatically generate gemm micro-kernels, either using C++ templates with vector intrinsics or high-level Python scripts that directly produce assembly code. Both solutions can integrate high performance software techniques, such as loop unrolling and software pipelining, accommodate any data type, and easily generate micro-kernels of any requested dimension. The performance of this solution is tested on three ARM-based cores and compared with state-of-the-art libraries for these processors: BLIS, OpenBLAS and ArmPL. The experimental results show that the auto-generation approach is highly competitive, mainly due to the possibility of adapting the micro-kernel to the problem dimensions.
Guillermo Alaejos, Héctor Martínez 0002, Adrián Castelló 0001, Manuel F. Dolz, Francisco D. Igual, Pedro Alonso 0002, Enrique S. Quintana-Ortí
J. Supercomput.2
2024 Parallel GEMM-based convolution for deep learning on multicore RISC-V processors
abstract
Abstract We address the efficient implementation of the convolution operator on the GAP8 parallel ultra-low power platform (PULP), a heterogeneous multi-core processor equipped with a fabric controller (FC); a cluster of eight compute cores; and a four-level memory hierarchy with scratchpads instead of conventional, hardware-assisted cache memories. Our solution for this platform transforms the convolution into a general matrix–matrix multiplication ( gemm ) via the lowering approach, demonstrating that it is possible to attain reasonable performance on the GAP8 by carefully adapting techniques such as tiling and loop parallelism, which are mainstream in the multi-threaded, cache-aware realization of gemm .
Cristián Ramírez, Adrián Castelló 0001, Héctor Martínez 0002, Enrique S. Quintana-Ortí
J. Supercomput.3
2024 Algorithm 1039: Automatic Generators for a Family of Matrix Multiplication Routines with Apache TVM
abstract
We explore the utilization of the Apache TVM open source framework to automatically generate a family of algorithms that follow the approach taken by popular linear algebra libraries, such as GotoBLAS2, BLIS, and OpenBLAS, to obtain high-performance blocked formulations of the general matrix multiplication ( gemm ). In addition, we fully automatize the generation process by also leveraging the Apache TVM framework to derive a complete variety of the processor-specific micro-kernels for gemm . This is in contrast with the convention in high-performance libraries, which hand-encode a single micro-kernel per architecture using Assembly code. In global, the combination of our TVM-generated blocked algorithms and micro-kernels for gemm (1) improves portability, maintainability, and, globally, streamlines the software life cycle; (2) provides high flexibility to easily tailor and optimize the solution to different data types, processor architectures, and matrix operand shapes, yielding performance on a par (or even superior for specific matrix shapes) with that of hand-tuned libraries; and (3) features a small memory footprint.
Guillermo Alaejos, Adrián Castelló 0001, Pedro Alonso 0002, Francisco D. Igual, Héctor Martínez 0002, Enrique S. Quintana-Ortí
ACM Trans. Math. Softw.5
2023 3D reconstruction system and multiobject local tracking algorithm designed for billiards
abstract
Abstract The use of virtual reality or augmented reality systems in billiards sports are useful tools for pure entertainment or improving the player’s skills. Depending on the purpose of these systems, tracking algorithms based on computer vision must be used. These algorithms are especially useful in systems aiming to reconstruct the trajectories followed by the balls after a strike. However, depending on the billiard modality, the problem of tracking multiple small identical objects, such as balls, is a complex task. In addition, when an amateur or nontop professional player uses low-frame-rate and low-resolution devices, problems such as blurred balls, blurred contours, or fuzzy edges, among others, arise. These effects have a negative impact on ball-tracking accuracy and reconstruction quality. Thus, this work proposes two contributions. The first contribution is a new tracking algorithm called“multiobject local tracking (MOLT)”. This algorithm can track balls with high precision and accuracy even with motion blur caused by low-resolution and low-frame-rate devices. Moreover, the proposed MOLT algorithm is compared with nine tracking methods and four different metrics, outperforming the rest of the methods in the majority of the cases and providing a robust solution. The second contribution is a whole system to track (using the MOLT algorithm) and reconstruct the movements of the balls on a billiard table in a 3D virtual world using computer vision. The proposed system covers all steps from image capture to 3D reconstruction. The 3D reconstruction results have been qualitatively evaluated by different users through a series of questionnaires, obtaining an overall score of 7.6 (out of 10), which indicates that the system is a promising and useful tool for training. Finally, both the MOLT algorithm and the reconstruction system are tested in three billiard modalities: blackball, carom billiards, and snooker.
Francisco J. Rodríguez-Lozano, Juan Carlos Gámez, Héctor Martínez 0002, José M. Palomares, Joaquín Olivares 0001
Appl. Intell.3
2023 Reformulating the direct convolution for high-performance deep learning inference on ARM processors
abstract
We present two high-performance implementations of the convolution operator via the direct algorithm that outperform the so-called lowering approach based on the im2col transform plus the gemm kernel on an ARMv8-based processor. One of our methods presents the additional advantage of zero-memory overhead while the other employs an additional yet rather moderate workspace, substantially smaller than that required by the im2col+gemm solution. In contrast with a previous implementation of a similar zero-memory overhead direct convolution, this work exhibits the key advantage of preserving the conventional NHWC data layout for the input/output activations of the convolution layers.
Sergio Barrachina 0001, Adrián Castelló 0001, Manuel F. Dolz, Tze Meng Low, Héctor Martínez 0002, Enrique S. Quintana-Ortí, Upasana Sridhar, Andrés Tomás
J. Syst. Archit.5
2023 Micro-kernels for portable and efficient matrix multiplication in deep learning
abstract
Abstract We provide a practical demonstration that it is possible to systematically generate a variety of high-performance micro-kernels for the general matrix multiplication (gemm) via generic templates which can be easily customized to different processor architectures and micro-kernel dimensions. These generic templates employ vector intrinsics to exploit the SIMD (single instruction, multiple data) units in current general-purpose processors and, for the particular type of gemm problems encountered in deep learning, deliver a floating-point throughput rate on par with or even higher than that obtained with conventional, carefully tuned implementations of gemm in current linear algebra libraries (e.g., BLIS, AMD AOCL, ARMPL). Our work exposes the structure of the template-based micro-kernels for ARM Neon (128-bit SIMD), ARM SVE (variable-length SIMD) and Intel AVX512 (512-bit SIMD), showing considerable performance for an NVIDIA Carmel processor (ARM Neon), a Fujitsu A64FX processor (ARM SVE) and on an AMD EPYC 7282 processor (256-bit SIMD).
Guillermo Alaejos, Adrián Castelló 0001, Héctor Martínez 0002, Pedro Alonso 0002, Francisco D. Igual, Enrique S. Quintana-Ortí
J. Supercomput.3
2023 Performance-energy trade-offs of deep learning convolution algorithms on ARM processors
abstract
Abstract In this work, we assess the performance and energy efficiency of high-performance codes for the convolution operator, based on the direct, explicit/implicit lowering and Winograd algorithms used for deep learning (DL) inference on a series of ARM-based processor architectures. Specifically, we evaluate the NVIDIA Denver2 and Carmel processors, as well as the ARM Cortex-A57 and Cortex-A78AE CPUs as part of a recent set of NVIDIA Jetson platforms. The performance–energy evaluation is carried out using the ResNet-50 v1.5 convolutional neural network (CNN) on varying configurations of convolution algorithms, number of threads/cores, and operating frequencies on the tested processor cores. The results demonstrate that the best throughput is obtained on all platforms with the Winograd convolution operator running on all the cores at their highest frequency. However, if the goal is to reduce the energy footprint, there is no rule of thumb for the optimal configuration.
Manuel F. Dolz, Sergio Barrachina 0001, Héctor Martínez 0002, Adrián Castelló 0001, Antonio M. Vidal, Germán Fabregat, Andrés Tomás
J. Supercomput.3
2023 Efficient and portable Winograd convolutions for multi-core processors
abstract
Abstract We take a step forward towards developing high-performance codes for the convolution operator, based on the Winograd algorithm, that are easy to customise for general-purpose processor architectures. In our approach, augmenting the portability of the solution is achieved via the introduction of vector instructions from Intel SSE/AVX2/AVX512 and ARM NEON/SVE to exploit the single-instruction multiple-data capabilities of current processors as well as OpenMP pragmas to exploit multi-threaded parallelism. While this comes at the cost of sacrificing a fraction of the computational performance, our experimental results on three distinct processors, with Intel Xeon Skylake, ARM Cortex A57 and Fujitsu A64FX processors, show that the impact is affordable and still renders a Winograd-based solution that is competitive when compared with the lowering gemm-based convolution.
Manuel F. Dolz, Héctor Martínez 0002, Adrián Castelló 0001, Pedro Alonso 0002, Enrique S. Quintana-Ortí
J. Supercomput.2
2022 Convolution Operators for Deep Learning Inference on the Fujitsu A64FX Processor
abstract
The convolution operator is a crucial kernel for many computer vision and signal processing applications that rely on deep learning (DL) technologies. As such, the efficient implementation of this operator has received considerable attention in the past few years for a fair range of processor architectures. In this paper, we follow the technology trend toward integrating long SIMD (single instruction, multiple data) arithmetic units into high performance multicore processors to analyse the benefits of this type of hardware acceleration for latency-constrained DL workloads. For this purpose, we implement and optimise for the Fujitsu processor A64FX, three distinct methods for the calculation of the convolution, namely, the lowering approach, a blocked variant of the direct convolution algorithm, and the Winograd minimal filtering algorithm. Our experimental results include an extensive evaluation of the parallel scalability of these three methods and a comparison of their global performance using three popular DL models and a representative dataset.
Manuel F. Dolz, Héctor Martínez 0002, Pedro Alonso 0002, Enrique S. Quintana-Ortí
SBAC-PAD2
2017 Accelerating FaST-LMM for Epistasis Tests
Héctor Martínez 0002, Sergio Barrachina 0001, María Isabel Castillo, Enrique S. Quintana-Ortí, Jordi Rambla De Argila, Xavier Farré, Arcadi Navarro
ICA3PP1
2015 Concurrent and Accurate Short Read Mapping on Multicore Processors
abstract
We introduce a parallel aligner with a work-flow organization for fast and accurate mapping of RNA sequences on servers equipped with multicore processors. Our software, HPG Aligner SA (HPG Aligner SA is an open-source application. The software is available at http://www.opencb.org, exploits a suffix array to rapidly map a large fraction of the RNA fragments (reads), as well as leverages the accuracy of the Smith-Waterman algorithm to deal with conflictive reads. The aligner is enhanced with a careful strategy to detect splice junctions based on an adaptive division of RNA reads into small segments (or seeds), which are then mapped onto a number of candidate alignment locations, providing crucial information for the successful alignment of the complete reads. The experimental results on a platform with Intel multicore technology report the parallel performance of HPG Aligner SA, on RNA reads of 100-400 nucleotides, which excels in execution time/sensitivity to state-of-the-art aligners such as TopHat 2+Bowtie 2, MapSplice, and STAR.
Héctor Martínez 0002, Joaquín Tárraga, Ignacio Medina, Sergio Barrachina 0001, María Isabel Castillo, Joaquín Dopazo, Enrique S. Quintana-Ortí
IEEE ACM Trans. Comput. Biol. Bioinform.1
2014 Acceleration of short and long DNA read mapping without loss of accuracy using suffix array
abstract
UNLABELLED: HPG Aligner applies suffix arrays for DNA read mapping. This implementation produces a highly sensitive and extremely fast mapping of DNA reads that scales up almost linearly with read length. The approach presented here is faster (over 20× for long reads) and more sensitive (over 98% in a wide range of read lengths) than the current state-of-the-art mappers. HPG Aligner is not only an optimal alternative for current sequencers but also the only solution available to cope with longer reads and growing throughputs produced by forthcoming sequencing technologies. AVAILABILITY AND IMPLEMENTATION: https://github.com/opencb/hpg-aligner.
Joaquín Tárraga, Vicente Arnau, Héctor Martínez 0002, Raul Moreno, Diego Cazorla, José Salavert Torres, Ignacio Blanquer, Joaquín Dopazo, Ignacio Medina
Bioinform.3
2013 A dynamic pipeline for RNA sequencing on multicore processors
abstract
We present a concurrent algorithm for mapping short and long RNA sequences on multicore processors. Our solution processes the data, initially stored on disk, in batches of reads which are passed between the consecutive stages of a pipeline. A major operational reorganization of the original static pipeline, combined with a complete reimplementation based on POSIX threads, renders a dissociated execution between threads and stages/task types, so that threads can compute any type of pending task resulting in a dynamic pipeline. The experiments on a multicore platform reveal that this reorganization yields significantly higher performance, specially for architectures equipped with a small to moderate number of cores.
Héctor Martínez 0002, Joaquín Tárraga, Ignacio Medina, Sergio Barrachina 0001, María Isabel Castillo, Joaquín Dopazo, Enrique S. Quintana-Ortí
EuroMPI1