VLDB 2026 Research / reviewers in the wild / expert
Sandra Catalán
dblp:135/0181
· DBLP profile ↗
26ranked-venue papers
12as first author
13since 2021 · last 2026
0000-0002-9321-2728ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 10 first-author · 10 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A comparative performance and efficiency analysis of Apple's M architectures: A GEMM case studyabstractThis paper evaluates the performance and energy efficiency of Apple processors across multiple ARM-based M-series generations and models (standard and Pro). The study is motivated by the increasing heterogeneity of Apple’s SoC architectures, which integrate multiple computing engines raising the scientific question of which hardware components are best suited for executing general-purpose and domain-specific computations such as the GEneral Matrix Multiply ( GEMM ). The analysis focuses on four key components: the Central Processing Unit (CPU), the Graphics Processing Unit (GPU), the matrix calculation accelerator (AMX), and the Apple Neural Engine (ANE). The assessments use the GEMM as benchmark to characterize the performance of the CPU and GPU, alongside tests on AMX, which is specialized in handling large-scale mathematical operations, and tests on the ANE, which is specifically designed for Deep Learning purposes. Additionally, energy consumption data has been collected to analyze the energy efficiency of the aforementioned resources. Results highlight notable improvements in computational capacity and energy efficiency over successive generations. On one hand, the AMX stands out as the most efficient component for FP32 and FP64 workloads, significantly boosting overall system performance. In the M4 Pro, which integrates two matrix accelerators, it achieves up to 68% of the GPU’s FP32 performance while consuming only 42% of its power. On the other hand, the ANE, although limited to FP16 precision, excels in energy efficiency for low-precision tasks, surpassing other accelerators with over 700 GFLOPs/Watt under batched workloads. This analysis offers a clear understanding of how Apple’s custom ARM designs optimize both performance and energy use, particularly in the context of multi-core processing and specialized acceleration units. In addition, a significant contribution of this study is the comprehensive comparative analysis of Apple’s accelerators, which have previously been poorly documented and scarcely studied. The analysis spans different generations and compares the accelerators against both CPU and GPU performance. Sandra Catalán, Rafael Rodríguez-Sánchez 0001, Carlos García 0001, Luis Piñuel Moreno |
Future Gener. Comput. Syst. | 1 |
| 2026 | Cross-platform characterisation and performance analysis of homomorphic matrix multiplicationabstractAbstract Fully Homomorphic Encryption (FHE) enables computation over encrypted data while preserving strong security and privacy guarantees. However, its high computational cost remains a major challenge. This study therefore evaluates the performance of homomorphic matrix multiplication using three Homomorphic Encryption (HE) libraries, Microsoft SEAL, HElib and OpenFHE, across two platforms with AMD EPYC and Intel Xeon CPUs, with a particular focus on the impact of different compiler flags. The results indicate that compiler configurations and hardware selection significantly influence runtime in libraries such as Microsoft SEAL, whereas HElib and OpenFHE show negligible variation under different compilation settings. Furthermore, the analysis reveals that SEAL makes more efficient use of memory bandwidth while OpenFHE achieves the highest overall performance, the lowest mean absolute error and the shortest execution time. Franklin Espinoza, Justo Molina, Darwin Quezada-Gaibor, Sandra Catalán, Manuel F. Dolz |
J. Supercomput. | 4 |
| 2026 | Exploiting mixed-precision redundancy for soft-error detection in LU decomposition
Nima Sahraneshinsamani, Sandra Catalán, José R. Herrero 0001 |
J. Supercomput. | 2 |
| 2025 | Portable, High Performance Matrix Multiplication Micro-Kernels for RISC-V with ExOabstractThe proliferation of RISC-V platforms and their use in a wide variety of scientific applications, including deep learning scenarios, has dramatically increased the interest to generate optimized code for them. In the field of HPC (High Performance Computing), the RISCV ISA (Instruction Set Architecture) has been adopted by a wide variety of designs with different micro-architecture; as a result, performance portability of existing codes is a major endeavor. Code generators and compilers such as Apache TVM, MLIR, or EXO provide a hardware abstraction for implementing optimized hardware-aware codes, thus reducing development time and potential errors. These generators can handle the full software stack, from basic micro-kernels to complex operations. In this work, we focus on the optimization of GEMM (general matrix-matrix multiplication), a key operation on top of which dense linear algebra libraries and deep learning frameworks are built. Specifically, we present an EXO-based GEMM microkernel generator for the RISC-V ISA with RVV vector extensions that addresses the lack of high-performance and portable GEMM micro-kernels. Our results demonstrate that, by generating a wide range of micro-kernels, one can obtain GEMM realizations that outperform those in the state-of-the-art high performance libraries. Adrián Castelló 0001, Héctor Martínez 0002, Sandra Catalán, Jie Lei 0007, Yuka Ikarashi, Grace Dinh, Francisco D. Igual, Enrique S. Quintana-Ortí |
PDP | 3 |
| 2025 | Latency-Critical Quantized Inference With Transformer Decoders on ARM and RISC-V CPUsabstractLarge language models are transforming industries but face challenges due to their high computational and energy demands. Model compression via quantization mitigates these barriers by reducing the bit precision of parameters and arithmetic operations, enabling deployment on resource-constrained devices like smartphones and edge platforms. This paper focuses on quantization applied to transformer decoders, which are critical for tasks such as text generation and conversational artificial intelligence. Unlike encoders, decoders are constrained by memory due to their sequential processing nature and low arithmetic intensity. We propose optimizations targeting inference on low-power CPUs, emphasizing efficient linear layers with quantized data/arithmetic and cache optimization. Using two representative ARM and RISC-V platforms, we present optimized mixed-precision implementations of the matrix multiplication that outperform the instance of that computational kernel in popular libraries such as BLIS, XNNPACK and ARMCL. This work thus advances the understanding of the impact of quantization on transformer decoder efficiency, energy consumption and precision in edge environments. Héctor Martínez 0002, Sandra Catalán, Adrián Castelló 0001, José I. Mestre, Enrique S. Quintana-Ortí |
IEEE Internet Things J. | 2 |
| 2025 | Experience-guided, mixed-precision matrix multiplication with apache TVM for ARM processorsabstractAbstract Deep learning (DL) generates new computational tasks that are different from those encountered in classical scientific applications. In particular, DL training and inference require general matrix multiplications (gemm) with matrix operands that are far from large and square as in other scientific fields. In addition, DL models gain arithmetic/storage complexity, and as a result, reduced precision via quantization is now mainstream for inferring DL models in edge devices. Automatic code generation addresses these new types of gemm by (1) improving portability between different hardware with only one base code; (2) supporting mixed and reduced precision; and (3) enabling auto-tuning methods that, given a base operation, perform a (costly) optimization search for the best schedule. In this paper, we rely on Apache TVM to generate an experience-guided gemm that provides performance competitive with the TVM auto-scheduler, while reducing tuning time by a factor of 48×. Adrián Castelló 0001, Héctor Martínez 0002, Sandra Catalán, Francisco D. Igual, Enrique S. Quintana-Ortí |
J. Supercomput. | 3 |
| 2025 | Mixed-precision pre-pivoting strategy for the LU factorizationabstractAbstract This paper investigates the efficient application of half-precision floating-point (FP16) arithmetic on GPUs for boosting LU decompositions in double (FP64) precision. Addressing the motivation to enhance computational efficiency, we introduce two novel algorithms: Pre-Pivoted LU (PRP) and Mixed-precision Panel Factorization (MPF). Deployed in both hybrid CPU-GPU setups and native GPU-only configurations, PRP identifies pivot lists through LU decomposition computed in reduced precision and subsequently reorders matrix rows in FP64 precision before executing LU decomposition without pivoting. Two variants of PRP, namely hPRP and xPRP, are introduced, differing in their computation of pivot lists in full half-precision or mixed half-single precision. The MPF algorithm generates FP64 LU factorization while internally utilizing hPRP for panel factorization, showcasing accuracy on par with standard DGETRF but with superior speed. The study further explores auxiliary functions required for the native mode implementation of PRP variants and MPF. Nima Sahraneshinsamani, Sandra Catalán, José R. Herrero 0001 |
J. Supercomput. | 2 |
| 2024 | Inference with Transformer Encoders on ARM and RISC-V Multicore ProcessorsabstractAbstract We delve into the performance of transformer encoder inference on low-power multi-core processors from two perspectives: First, we conduct a detailed profile of the inference process for two members of the BERT family on a modern multi-core processor, identifying the main bottlenecks and opportunities for improvement. Second, we propose a number of accumulative optimisations for their primary building blocks. For that, we elaborate our own implementation of the general matrix multiplication (), which dynamically tunes several key parameters yielding relevant performance gains for transformer encoders. Additionally, we introduce a number of strategies to also improve the parallel execution of the transformer block. Our implementations for ARMv8a and RISC-V multi-core processors with SIMD units, taking as a reference state-of-the-art implementations (BLIS for ARM and OpenBLAS for RISC-V) reveal accelerations of up to $$2.5\times $$ 2.5 × for natural language processing tasks. Héctor Martínez 0002, Francisco D. Igual, Rafael Rodríguez-Sánchez 0001, Sandra Catalán, Adrián Castelló 0001, Enrique S. Quintana-Ortí |
Euro-Par (2) | 4 |
| 2024 | Parallel GEMM-based convolutions for deep learning on multicore ARM and RISC-V architectures
Héctor Martínez 0002, Sandra Catalán, Adrián Castelló 0001, Enrique S. Quintana-Ortí |
J. Syst. Archit. | 2 |
| 2023 | Fine-grain task-parallel algorithms for matrix factorizations and inversion on many-threaded CPUsabstractAbstract We extend a two‐level task partitioning previously applied to the inversion of dense matrices via Gauss–Jordan elimination to the more challenging QR factorization as well as the initial orthogonal reduction to band form found in the singular value decomposition. Our new task‐parallel algorithms leverage the tasking mechanism currently available in OpenMP to exploit “nested” task parallelism, with a first outer level that operates on matrix panels and a second inner level that processes the matrix either by ‐panels or by tiles, in order to expose a large number of independent tasks. We present a detailed performance analysis, including execution traces, which shows that the two‐level refinement into fine grain tasks allows for an improved load balancing and delivers high performance on current general‐purpose many‐core processors (CPUs) from Intel and AMD. Sandra Catalán, José R. Herrero 0001, Francisco D. Igual, Enrique S. Quintana-Ortí, Rafael Rodríguez-Sánchez 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2023 | Programming parallel dense matrix factorizations and inversion for new-generation NUMA architecturesabstractWe propose a methodology to address the programmability issues derived from the emergence of new-generation shared-memory NUMA architectures. For this purpose, we employ dense matrix factorizations and matrix inversion (DMFI) as a use case, and we target two modern architectures (AMD Rome and Huawei Kunpeng 920) that exhibit configurable NUMA topologies. Our methodology pursues performance portability across different NUMA configurations by proposing multi-domain implementations for DMFI plus a hybrid task- and loop-level parallelization that configures multi-threaded executions to fix core-to-data binding, exploiting locality at the expense of minor code modifications. In addition, we introduce a generalization of the multi-domain implementations for DMFI that offers support for virtually any NUMA topology in present and future architectures. Our experimentation on the two target architectures for three representative dense linear algebra operations validates the proposal, reveals insights on the necessity of adapting both the codes and their execution to improve data access locality, and reports performance across architectures and inter- and intra-socket NUMA configurations competitive with state-of-the-art message-passing implementations, maintaining the ease of development usually associated with shared-memory programming. Sandra Catalán, Francisco D. Igual, José R. Herrero 0001, Rafael Rodríguez-Sánchez 0001, Enrique S. Quintana-Ortí |
J. Parallel Distributed Comput. | 1 |
| 2022 | NUMA-Aware Dense Matrix Factorizations and Inversion with Look-Ahead on Multicore ProcessorsabstractWe address the efficient design and implementation of dense matrix factorizations and inversion (DMFI) on modern multicore processors with several NUMA (non-uniform memory access) nodes. Our approach enhances the DMFI routines with a look-ahead strategy, in order to overcome the “panel factorization bottleneck”. In addition, it exploits both hybrid task- and loop-level parallelizations while taking into account the NUMA organization of the memory hierarchy. The experiments on a Huawei Kunpeng-based server, with two sockets and 48 cores per socket, for three representative dense linear algebra operations, expose the necessity of adapting both the codes and their execution environment parameters to improve data access locality. The results of these changes deliver performance across inter- and intra-socket NUMA configurations superior to that of reference implementations from state-of-the-art libraries for this platform. Sandra Catalán, Francisco D. Igual, Rafael Rodríguez-Sánchez 0001, José R. Herrero 0001, Enrique S. Quintana-Ortí |
SBAC-PAD | 1 |
| 2021 | Leveraging teaching on demand: Approaching HPC to undergrads
Sandra Catalán, Rocío Carratalá-Sáez, Sergio Iserte |
J. Parallel Distributed Comput. | 1 |
| 2020 | sLASs: A fully automatic auto-tuned linear algebra library based on OpenMP extensions implemented in OmpSs (LASs Library)
Pedro Valero-Lara, Sandra Catalán, Xavier Martorell, Tetsuzo Usui, Jesús Labarta |
J. Parallel Distributed Comput. | 2 |
| 2019 | Accelerating Conjugate Gradient using OmpSsabstractIn this paper, we present the benefits of using the clause concurrent of OmpSs when performing reductions, more specifically, when applied to the dot product (DOT) operations. We analyze its benefits through the implementation of different versions of the Conjugate Gradient (CG) method. We start from a parallel version of the code based on tasks and dependencies; later, we introduce the use of the concurrent clause, which allows to overlap the execution of tasks that have data dependencies among them. In this way, we want to show the benefits of the concurrent clause, which might be included in OpenMP standard as previously done with other OmpSs features. Our tests, performed on a single node of the (Intel-based) Marenostrum 4 Supercomputer and a single socket of the (ARM-based) Dibona cluster, show that the use of the concurrent clause may improve performance with respect to the version where only tasks and dependencies are used around 37% and 23% respectively. Sandra Catalán, Xavier Martorell, Jesús Labarta, Tetsuzo Usui, Leonel Toledo, Pedro Valero-Lara |
PDCAT | 1 |
| 2019 | Tasking in Accelerators: Performance EvaluationabstractIn this work, we analyze the implications and results of implementing dynamic parallelism, concurrent kernels and CUDA Graphs to solve task-oriented problems. As a benchmark we propose three different methods for solving DGEMM operation on tiled-matrices; which might be the most popular benchmark for performance analysis. For the algorithms that we study, we present significant differences in terms of data dependencies, synchronization and granularity. The main contribution of this work is determining which of the previous approaches work better for having multiple task running concurrently in a single GPU, as well as stating the main limitations and benefits of every technique. Using dynamic parallelism and CUDA Streams we were able to achieve up to 30% speedups and for CUDA Graph API up to 25x acceleration outperforming state of the art results. Leonel Toledo, Antonio J. Peña, Sandra Catalán, Pedro Valero-Lara |
PDCAT | 3 |
| 2019 | BLAS-3 Optimized by OmpSs Regions (LASs Library)abstractIn this paper we propose a set of optimizations for the BLAS-3 routines of LASs library (Linear Algebra routines on OmpSs) and perform a detailed analysis of the impact of the proposed changes in terms of performance and execution time. OmpSs allows to use regions in the dependences of the tasks. This helps not only in the programming of the algorithmic optimizations, but also in the reduction of the execution time achieved by such optimizations. Different strategies are implemented in order to reduce the amount of tasks created (when there is enough parallelism) during the execution of BLAS-3 operations in the original LASs. Also a better IPC is obtained thanks to a better memory hierarchy exploitation. More specifically, we increase the performance, in particular on big matrices, about 12% for TRSM, and 17% for GEMM with respect to the original version of LASs, even using less cores in the case of GEMM/SYMM. Moreover, when LASs is compared to the OpenMP reference dense linear algebra library PLASMA, performance is increased up to 12.5% for GEMM/SYMM, while for TRSM/TRMM this value raises to 15%. Pedro Valero-Lara, Sandra Catalán, Xavier Martorell, Jesús Labarta |
PDP | 2 |
| 2019 | Dynamic look-ahead in the reduction to band form for the singular value decomposition
Andrés Tomás, Rafael Rodríguez-Sánchez 0001, Sandra Catalán, Rocío Carratalá-Sáez, Enrique S. Quintana-Ortí |
Parallel Comput. | 3 |
| 2018 | Two-sided orthogonal reductions to condensed forms on asymmetric multicore processors
Pedro Alonso 0002, Sandra Catalán, José R. Herrero 0001, Enrique S. Quintana-Ortí, Rafael Rodríguez-Sánchez 0001 |
Parallel Comput. | 2 |
| 2018 | Energy balance between voltage-frequency scaling and resilience for linear algebra routines on low-power multicore architectures
Sandra Catalán, José R. Herrero 0001, Enrique S. Quintana-Ortí, Rafael Rodríguez-Sánchez 0001 |
Parallel Comput. | 1 |
| 2018 | Static scheduling of the LU factorization with look-ahead on asymmetric multicore processors
Sandra Catalán, José R. Herrero 0001, Enrique S. Quintana-Ortí, Rafael Rodríguez-Sánchez 0001 |
Parallel Comput. | 1 |
| 2017 | Revisiting conventional task schedulers to exploit asymmetry in multi-core architectures for dense linear algebra operations
Luis Costero, Francisco D. Igual, Katzalin Olcoz, Sandra Catalán, Rafael Rodríguez-Sánchez 0001, Enrique S. Quintana-Ortí |
Parallel Comput. | 4 |
| 2017 | Time and energy modeling of a high-performance multi-threaded Cholesky factorization
Sandra Catalán, Francisco D. Igual, Rafael Mayo 0002, Rafael Rodríguez-Sánchez 0001, Enrique S. Quintana-Ortí |
J. Supercomput. | 1 |
| 2016 | The Impact of Voltage-Frequency Scaling for the Matrix-Vector Product on the IBM POWER8
Sandra Catalán, Cristiano Malossi, Costas Bekas, Enrique S. Quintana-Ortí |
Euro-Par | 1 |
| 2016 | The Impact of Panel Factorization on the Gauss-Huard Algorithm for the Solution of Linear Systems on Modern Architectures
Sandra Catalán, Pablo Ezzatti, Enrique S. Quintana-Ortí, Alfredo Remón |
ICA3PP | 1 |
| 2014 | Analyzing the Energy Efficiency of the Memory Subsystem in Multicore ProcessorsabstractIn this paper we analyze the energy overhead incurred when operating with data stored in different levels of the memory subsystem (cache levels and DDR chips) of current multicore architectures. Our approach builds upon servet, a portable framework for the memory characterization of multicore processors, extending this suite with a power-related test that, when applied to a platform equipped with a power measurement mechanism, provides information on the efficiency of memory energy usage. As additional contributions, i) we provide a complete experimental study of the impact that the CPU performance states (also known as P-states) exert on the memory energy efficiency of a collection of recent server-oriented and low-power cores, and ii) we show how this framework carries over to cover also the scalability analysis of the memory energy performance on multicore processors. Sandra Catalán, Jorge González-Domínguez, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
ISPA | 1 |