VLDB 2026 Research / reviewers in the wild / expert
Adrián Castelló 0001
dblp:154/7966 · also Adrián Castelló Gimeno
· DBLP profile ↗
41ranked-venue papers
16as first author
26since 2021 · last 2026
0000-0002-8576-8451ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 11 first-author · 18 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The cambrian explosion of mixed-precision matrix multiplication for quantized deep learning inferenceabstractRecent advances in deep learning (DL) have promoted to a shift from traditional 64-bit floating point (FP64) arithmetic for scientific computing toward reduced-precision formats–such as FP16, BF16, or even 8-bit integers–combined with mixed-precision arithmetic. This transition enhances computational throughput, reduces memory and bandwidth usage, and improves energy efficiency, offering significant advantages for resource-constrained edge devices. To support this shift, hardware architectures have evolved accordingly, now including adapted ISAs (Instruction Set Architectures) that expose mixed-precision vector units and matrix engines tailored for DL workloads. At the heart of many DL and scientific computing tasks is the general matrix-matrix multiplication ( GEMM ), a fundamental kernel historically optimized using fused multiply-add (FMA) vector instructions on SIMD (single instruction, multiple data) units. However, as hardware moves toward mixed-precision dot (or inner)-product-centric operations optimized for quantized inference, these legacy approaches are being phased out. In response to this, our paper revisits the conventional, high-performance implementation of GEMM and describes strategies for adapting it to mixed integer precision (MIP) arithmetic across modern ISAs, including x86_64, Arm, and RISC-V. Concretely, we illustrate novel micro-kernel designs and data layouts that better exploit today’s specialized hardware and demonstrate significant performance gains from MIP arithmetic over floating-point implementations across three representative CPUs. These contributions highlight a new era of GEMM optimization-driven by the demands of DL inference on heterogeneous architectures, marking what we term as the “Cambrian period” for matrix multiplication. Héctor Martínez 0002, Adrián Castelló 0001, Francisco D. Igual, Enrique S. Quintana-Ortí |
Future Gener. Comput. Syst. | 2 |
| 2026 | Enhancing transformer performance and portability through auto-tuning frameworksabstractAbstract Transformer-based models such as BERT and GPT2 have become the foundation of many modern applications, yet their execution requires substantial computational and memory resources. To address these challenges, recent advances in compiler technology and hardware accelerators have introduced new opportunities for performance portability. In this work, we evaluate JAX and TVM as high-level frameworks that combine a NumPy-like programming model with Just-In-Time (JIT) or Ahead-of-Time (AOT) code optimization and compilation, enabling efficient execution across CPUs or GPUs, and in the case of JAX, on TPUs as well. We present systematic implementations of the core Transformer encoder and decoder blocks in JAX and TVM and compare their automatically optimized code against NumPy and CuPy baselines. Our experimental study covers heterogeneous hardware platforms (AMD CPU, NVIDIA GPUs, and Google TPUs) and multiple arithmetic precisions (FP32, BF16, INT8, and INT32). Results show that JAX and TVM deliver significant performance improvements over standard libraries, while reducing the programming effort required to adapt to different hardware. These findings demonstrate the potential of JIT- and AOT-oriented frameworks to serve as a portable and efficient solution for deploying Transformer workloads in diverse computing environments. Patricia Siwinska, Jie Lei 0007, Adrián Castelló 0001, Pedro Alonso 0002, Enrique S. Quintana-Ortí |
J. Supercomput. | 3 |
| 2025 | Portable, High Performance Matrix Multiplication Micro-Kernels for RISC-V with ExOabstractThe proliferation of RISC-V platforms and their use in a wide variety of scientific applications, including deep learning scenarios, has dramatically increased the interest to generate optimized code for them. In the field of HPC (High Performance Computing), the RISCV ISA (Instruction Set Architecture) has been adopted by a wide variety of designs with different micro-architecture; as a result, performance portability of existing codes is a major endeavor. Code generators and compilers such as Apache TVM, MLIR, or EXO provide a hardware abstraction for implementing optimized hardware-aware codes, thus reducing development time and potential errors. These generators can handle the full software stack, from basic micro-kernels to complex operations. In this work, we focus on the optimization of GEMM (general matrix-matrix multiplication), a key operation on top of which dense linear algebra libraries and deep learning frameworks are built. Specifically, we present an EXO-based GEMM microkernel generator for the RISC-V ISA with RVV vector extensions that addresses the lack of high-performance and portable GEMM micro-kernels. Our results demonstrate that, by generating a wide range of micro-kernels, one can obtain GEMM realizations that outperform those in the state-of-the-art high performance libraries. Adrián Castelló 0001, Héctor Martínez 0002, Sandra Catalán, Jie Lei 0007, Yuka Ikarashi, Grace Dinh, Francisco D. Igual, Enrique S. Quintana-Ortí |
PDP | 1 |
| 2025 | Latency-Critical Quantized Inference With Transformer Decoders on ARM and RISC-V CPUsabstractLarge language models are transforming industries but face challenges due to their high computational and energy demands. Model compression via quantization mitigates these barriers by reducing the bit precision of parameters and arithmetic operations, enabling deployment on resource-constrained devices like smartphones and edge platforms. This paper focuses on quantization applied to transformer decoders, which are critical for tasks such as text generation and conversational artificial intelligence. Unlike encoders, decoders are constrained by memory due to their sequential processing nature and low arithmetic intensity. We propose optimizations targeting inference on low-power CPUs, emphasizing efficient linear layers with quantized data/arithmetic and cache optimization. Using two representative ARM and RISC-V platforms, we present optimized mixed-precision implementations of the matrix multiplication that outperform the instance of that computational kernel in popular libraries such as BLIS, XNNPACK and ARMCL. This work thus advances the understanding of the impact of quantization on transformer decoder efficiency, energy consumption and precision in edge environments. Héctor Martínez 0002, Sandra Catalán, Adrián Castelló 0001, José I. Mestre, Enrique S. Quintana-Ortí |
IEEE Internet Things J. | 3 |
| 2025 | Experience-guided, mixed-precision matrix multiplication with apache TVM for ARM processorsabstractAbstract Deep learning (DL) generates new computational tasks that are different from those encountered in classical scientific applications. In particular, DL training and inference require general matrix multiplications (gemm) with matrix operands that are far from large and square as in other scientific fields. In addition, DL models gain arithmetic/storage complexity, and as a result, reduced precision via quantization is now mainstream for inferring DL models in edge devices. Automatic code generation addresses these new types of gemm by (1) improving portability between different hardware with only one base code; (2) supporting mixed and reduced precision; and (3) enabling auto-tuning methods that, given a base operation, perform a (costly) optimization search for the best schedule. In this paper, we rely on Apache TVM to generate an experience-guided gemm that provides performance competitive with the TVM auto-scheduler, while reducing tuning time by a factor of 48×. Adrián Castelló 0001, Héctor Martínez 0002, Sandra Catalán, Francisco D. Igual, Enrique S. Quintana-Ortí |
J. Supercomput. | 1 |
| 2024 | Tackling the Matrix Multiplication Micro-Kernel Generation with ExoabstractThe optimization of the matrix multiplication (or GEMM) has been a need during the last decades. This operation is considered the flagship of current linear algebra libraries such as BLIS, OpenBLAS, or Intel OneAPI because of its widespread use in a large variety of scientific applications. The GEMM is usually implemented following the GotoBLAS philosophy, which tiles the GEMM operands and uses a series of nested loops for performance improvement. These approaches extract the maximum computational power of the architectures through small pieces of hardware-oriented, high-performance code called micro-kernel. However, this approach forces developers to generate, with a nonnegligible effort, a dedicated micro-kernel for each new hardware. In this work, we present a step-by-step procedure for generating micro-kernels with the Exo compiler that perform close to (or even better than) manually developed microkernels written with intrinsic functions or assembly language. Our solution also improves the portability of the generated code, since a hardware target is fully specified by a concise library-based description of its instructions. Adrián Castelló 0001, Julian Bellavita, Grace Dinh, Yuka Ikarashi, Héctor Martínez 0002 |
CGO | 1 |
| 2024 | Inference with Transformer Encoders on ARM and RISC-V Multicore ProcessorsabstractAbstract We delve into the performance of transformer encoder inference on low-power multi-core processors from two perspectives: First, we conduct a detailed profile of the inference process for two members of the BERT family on a modern multi-core processor, identifying the main bottlenecks and opportunities for improvement. Second, we propose a number of accumulative optimisations for their primary building blocks. For that, we elaborate our own implementation of the general matrix multiplication (), which dynamically tunes several key parameters yielding relevant performance gains for transformer encoders. Additionally, we introduce a number of strategies to also improve the parallel execution of the transformer block. Our implementations for ARMv8a and RISC-V multi-core processors with SIMD units, taking as a reference state-of-the-art implementations (BLIS for ARM and OpenBLAS for RISC-V) reveal accelerations of up to $$2.5\times $$ 2.5 × for natural language processing tasks. Héctor Martínez 0002, Francisco D. Igual, Rafael Rodríguez-Sánchez 0001, Sandra Catalán, Adrián Castelló 0001, Enrique S. Quintana-Ortí |
Euro-Par (2) | 5 |
| 2024 | Optimization of One-to-Many Communication Primitives for Dragonfly TopologiesabstractCollective communication primitives (CCPs), such as multicast and broadcast, are essential for many parallel and distributed applications. In response, this study compares a topology-oblivious algorithm underlying the implementation of CCPs in standard instances of MPI with two topology-aware implementations, based on the LLF and GLF algorithms, and an ideal hardware-assisted approach. By using real scientific applications, instead of synthetic traffic, our study reveals workload-dependent performance variations among CCP implementations; and highlights the importance of CCP algorithm selection in optimizing application performance in supercomputing environments. Jose Duro, Adrián Castelló 0001, María Engracia Gómez, Julio Sahuquillo, Enrique S. Quintana-Ortí |
ICPADS | 2 |
| 2024 | Communication-Avoiding Fusion of GEMM-Based Convolutions for Deep Learning in the RISC-V GAP8 MCUabstractIncorporating deep learning (DL) technologies to the edge is crucial for improving the security, privacy, and energy efficiency of the Internet of Things (IoT). In this scenario, the limitations of edge devices in terms of power dissipation, memory capacity, and processing power require a careful selection and optimization of algorithms for IoT DL applications. In this line, our work focuses on the convolution operator, a key component in deep neural networks for signal processing and computer vision. Specifically, the work aims at the efficient implementation of the lowering-based implementation of this operator, on the GAP8 parallel ultra-low power platform (PULP), with the goal of mitigating the data transfer costs across the memory hierarchy. Our contributions include 1) an analytical model for estimating the parallel execution time, 2) the exploration of different configuration options, and 3) four variants of the algorithm that fuse several components to address memory bottlenecks in the method. Overall, our best-fused variant provides a speedup of up to 1.25× over the baseline algorithm when applied to infer MobileNet-v1+ImageNet and VGG9+CIFAR10 using 8 threads, and up to 1.34× for ResNet18+ImageNet. Cristián Ramírez, Adrián Castelló 0001, Héctor Martínez 0002, Enrique S. Quintana-Ortí |
IEEE Internet Things J. | 2 |
| 2024 | Parallel GEMM-based convolutions for deep learning on multicore ARM and RISC-V architectures
Héctor Martínez 0002, Sandra Catalán, Adrián Castelló 0001, Enrique S. Quintana-Ortí |
J. Syst. Archit. | 3 |
| 2024 | Automatic generation of ARM NEON micro-kernels for matrix multiplicationabstractAbstract General matrix multiplication ( gemm ) is a fundamental kernel in scientific computing and current frameworks for deep learning. Modern realisations of gemm are mostly written in C, on top of a small, highly tuned micro-kernel that is usually encoded in assembly. The high performance realisation of gemm in linear algebra libraries in general include a single micro-kernel per architecture, usually implemented by an expert. In this paper, we explore a couple of paths to automatically generate gemm micro-kernels, either using C++ templates with vector intrinsics or high-level Python scripts that directly produce assembly code. Both solutions can integrate high performance software techniques, such as loop unrolling and software pipelining, accommodate any data type, and easily generate micro-kernels of any requested dimension. The performance of this solution is tested on three ARM-based cores and compared with state-of-the-art libraries for these processors: BLIS, OpenBLAS and ArmPL. The experimental results show that the auto-generation approach is highly competitive, mainly due to the possibility of adapting the micro-kernel to the problem dimensions. Guillermo Alaejos, Héctor Martínez 0002, Adrián Castelló 0001, Manuel F. Dolz, Francisco D. Igual, Pedro Alonso 0002, Enrique S. Quintana-Ortí |
J. Supercomput. | 3 |
| 2024 | Parallel GEMM-based convolution for deep learning on multicore RISC-V processorsabstractAbstract We address the efficient implementation of the convolution operator on the GAP8 parallel ultra-low power platform (PULP), a heterogeneous multi-core processor equipped with a fabric controller (FC); a cluster of eight compute cores; and a four-level memory hierarchy with scratchpads instead of conventional, hardware-assisted cache memories. Our solution for this platform transforms the convolution into a general matrix–matrix multiplication ( gemm ) via the lowering approach, demonstrating that it is possible to attain reasonable performance on the GAP8 by carefully adapting techniques such as tiling and loop parallelism, which are mainstream in the multi-threaded, cache-aware realization of gemm . Cristián Ramírez, Adrián Castelló 0001, Héctor Martínez 0002, Enrique S. Quintana-Ortí |
J. Supercomput. | 2 |
| 2024 | Algorithm 1039: Automatic Generators for a Family of Matrix Multiplication Routines with Apache TVMabstractWe explore the utilization of the Apache TVM open source framework to automatically generate a family of algorithms that follow the approach taken by popular linear algebra libraries, such as GotoBLAS2, BLIS, and OpenBLAS, to obtain high-performance blocked formulations of the general matrix multiplication ( gemm ). In addition, we fully automatize the generation process by also leveraging the Apache TVM framework to derive a complete variety of the processor-specific micro-kernels for gemm . This is in contrast with the convention in high-performance libraries, which hand-encode a single micro-kernel per architecture using Assembly code. In global, the combination of our TVM-generated blocked algorithms and micro-kernels for gemm (1) improves portability, maintainability, and, globally, streamlines the software life cycle; (2) provides high flexibility to easily tailor and optimize the solution to different data types, processor architectures, and matrix operand shapes, yielding performance on a par (or even superior for specific matrix shapes) with that of hand-tuned libraries; and (3) features a small memory footprint. Guillermo Alaejos, Adrián Castelló 0001, Pedro Alonso 0002, Francisco D. Igual, Héctor Martínez 0002, Enrique S. Quintana-Ortí |
ACM Trans. Math. Softw. | 2 |
| 2023 | Reformulating the direct convolution for high-performance deep learning inference on ARM processorsabstractWe present two high-performance implementations of the convolution operator via the direct algorithm that outperform the so-called lowering approach based on the im2col transform plus the gemm kernel on an ARMv8-based processor. One of our methods presents the additional advantage of zero-memory overhead while the other employs an additional yet rather moderate workspace, substantially smaller than that required by the im2col+gemm solution. In contrast with a previous implementation of a similar zero-memory overhead direct convolution, this work exhibits the key advantage of preserving the conventional NHWC data layout for the input/output activations of the convolution layers. Sergio Barrachina 0001, Adrián Castelló 0001, Manuel F. Dolz, Tze Meng Low, Héctor Martínez 0002, Enrique S. Quintana-Ortí, Upasana Sridhar, Andrés Tomás |
J. Syst. Archit. | 2 |
| 2023 | Micro-kernels for portable and efficient matrix multiplication in deep learningabstractAbstract We provide a practical demonstration that it is possible to systematically generate a variety of high-performance micro-kernels for the general matrix multiplication (gemm) via generic templates which can be easily customized to different processor architectures and micro-kernel dimensions. These generic templates employ vector intrinsics to exploit the SIMD (single instruction, multiple data) units in current general-purpose processors and, for the particular type of gemm problems encountered in deep learning, deliver a floating-point throughput rate on par with or even higher than that obtained with conventional, carefully tuned implementations of gemm in current linear algebra libraries (e.g., BLIS, AMD AOCL, ARMPL). Our work exposes the structure of the template-based micro-kernels for ARM Neon (128-bit SIMD), ARM SVE (variable-length SIMD) and Intel AVX512 (512-bit SIMD), showing considerable performance for an NVIDIA Carmel processor (ARM Neon), a Fujitsu A64FX processor (ARM SVE) and on an AMD EPYC 7282 processor (256-bit SIMD). Guillermo Alaejos, Adrián Castelló 0001, Héctor Martínez 0002, Pedro Alonso 0002, Francisco D. Igual, Enrique S. Quintana-Ortí |
J. Supercomput. | 2 |
| 2023 | Performance-energy trade-offs of deep learning convolution algorithms on ARM processorsabstractAbstract In this work, we assess the performance and energy efficiency of high-performance codes for the convolution operator, based on the direct, explicit/implicit lowering and Winograd algorithms used for deep learning (DL) inference on a series of ARM-based processor architectures. Specifically, we evaluate the NVIDIA Denver2 and Carmel processors, as well as the ARM Cortex-A57 and Cortex-A78AE CPUs as part of a recent set of NVIDIA Jetson platforms. The performance–energy evaluation is carried out using the ResNet-50 v1.5 convolutional neural network (CNN) on varying configurations of convolution algorithms, number of threads/cores, and operating frequencies on the tested processor cores. The results demonstrate that the best throughput is obtained on all platforms with the Winograd convolution operator running on all the cores at their highest frequency. However, if the goal is to reduce the energy footprint, there is no rule of thumb for the optimal configuration. Manuel F. Dolz, Sergio Barrachina 0001, Héctor Martínez 0002, Adrián Castelló 0001, Antonio M. Vidal, Germán Fabregat, Andrés Tomás |
J. Supercomput. | 4 |
| 2023 | Efficient and portable Winograd convolutions for multi-core processorsabstractAbstract We take a step forward towards developing high-performance codes for the convolution operator, based on the Winograd algorithm, that are easy to customise for general-purpose processor architectures. In our approach, augmenting the portability of the solution is achieved via the introduction of vector instructions from Intel SSE/AVX2/AVX512 and ARM NEON/SVE to exploit the single-instruction multiple-data capabilities of current processors as well as OpenMP pragmas to exploit multi-threaded parallelism. While this comes at the cost of sacrificing a fraction of the computational performance, our experimental results on three distinct processors, with Intel Xeon Skylake, ARM Cortex A57 and Fujitsu A64FX processors, show that the impact is affordable and still renders a Winograd-based solution that is competitive when compared with the lowering gemm-based convolution. Manuel F. Dolz, Héctor Martínez 0002, Adrián Castelló 0001, Pedro Alonso 0002, Enrique S. Quintana-Ortí |
J. Supercomput. | 3 |
| 2022 | RED-SEA: Network Solution for Exascale ArchitecturesabstractIn order to enable Exascale computing, next generation interconnection networks must scale to hundreds of thousands of nodes, and must provide features to also allow the HPC, HPDA, and AI applications to reach Exascale, while benefiting from new hardware and software trends. RED-SEA will pave the way to the next generation of European Exascale interconnects, including the next generation of BXI, as follows: (i) specify the new architecture using hardware-software co-design and a set of applications representative of the new terrain of converging HPC, HPDA, and AI; (ii) test, evaluate, and/or implement the new architectural features at multiple levels, according to the nature of each of them, ranging from mathematical analysis and modeling, to simulation, or to emulation or implementation on FPGA testbeds; (iii) enable seamless communication within and between resource clusters, and therefore development of a high-performance low latency gateway, bridging seamlessly with Ethernet; (iv) add efficient network resource management, thus improving congestion resiliency, virtualization, adaptive routing, collective operations; (v) open the interconnect to new kinds of applications and hardware, with enhancements for end-to-end network services - from programming models to reliability, security, low- latency, and new processors; (vi) leverage open standards and compatible APIs to develop innovative reusable libraries and Fabrics management solutions. Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Pier Stanislao Paolucci, Elena Pastorelli, Francesco Simula, Matteo Turisini, Piero Vicini, Roberto Ammendola, Pascale Bernier-Bruna, Said Derradji, Stéphane Guez, Pierre-Axel Lagadec, Gregoire Pichon, Etienne Walter, Gaetan De Gassowski, Matthieu Hautreaux, Stephane Mathieu, Gilles Moreau, Marc Pérache, Hugo Taboada, Torsten Hoefler, Timo Schneider, Matteo Barnaba, Giuseppe Piero Brandino, Francesco De Giorgi, Matteo Poggi, Iakovos Mavroidis, Ioannis Papaefstathiou, Nikolaos Tampouratzis, Benjamin Kalisch, Ulrich Krackhardt, Mondrian Nüssle, Pantelis Xirouchakis, Vangelis Mageiropoulos, Michalis Gianioudis, Harisis Loukas, Aggelos Ioannou, Nikolaos D. Kallimanis, Nikolaos Chrysos, Manolis Katevenis, Wolfgang Frings, Dominik Gottwald, Felime Guimaraes, Max Holicki, Volker Marx, Yannik Müller, Carsten Clauss, Hugo Falter, Xu Huang 0010, Jennifer Lopez Barillao, Thomas Moschny, Simon Pickartz, Francisco J. Alfaro, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José L. Sánchez 0002, Adrián Castelló 0001, Jose Duro, María Engracia Gómez, Enrique S. Quintana-Ortí, Julio Sahuquillo, Eugenio Stabile |
DSD | 63 |
| 2022 | Anatomy of the BLIS Family of Algorithms for Matrix MultiplicationabstractThe efforts of the scientific community and hardware vendors to develop and optimize linear algebra codes have historically led to highly-tuned libraries, carefully adapted to the underlying processor architecture, with excellent (near-peak) performance. These optimization efforts, however, are commonly focused on obtaining the best performance possible when the involved operands are large and “squarish” matrices. New computationally-intensive applications (e.g., in deep learning) are increasingly demanding high-performance BLAS (Basic Linear Algebra Subprograms) also for small operands in any of their dimensions. In this paper, we tackle this problem by refactoring the general matrix-matrix multiplication (GEMM) algorithm within a specific high-performance implementation of BLAS, named BLIS, proposing a complete family of algorithmic variants to implement GEMM with different strategies to exploit the target cache hierarchy, together with the changes to be applied to architecture-specific codes to instantiate a complete GEMM implementation. Experimental results on an ARM processor (NVIDIA Carmel) reveal significant performance differences between the members of the GEMM family, depending on the shape and dimension of the matrix operands. Adrián Castelló 0001, Enrique S. Quintana-Ortí, Francisco D. Igual |
PDP | 1 |
| 2022 | Towards Portable Realizations of Winograd-based Convolution with Vector Intrinsics and OpenMPabstractWe take a step forward in the direction of developing high performance codes for the convolution, based on the Winograd transformation, that are easy to customize for different processor architectures. In our approach, augmenting the portability of the solution is achieved via the introduction of vector intrinsics to exploit the SIMD (single-instruction multiple-data) capabilities of current processors as well as OpenMP pragmas to exploit multi-thread parallelism. While this comes at the cost of sacrificing a fraction of the computational performance, our experimental results on two distinct processors, with Intel Xeon Skylake and ARM Cortex A57 architectures, show that the impact is affordable, and still renders a Winograd-based solution that is competitive with the general method for the convolution based on the so-called im2col transform followed by a matrix-matrix multiplication. Manuel F. Dolz, Adrián Castelló 0001, Enrique S. Quintana-Ortí |
PDP | 2 |
| 2022 | High performance and energy efficient inference for deep learning on multicore ARM processors using general optimization techniques and BLISabstractWe evolve PyDTNN, a framework for distributed parallel training of Deep Neural Networks (DNNs), into an efficient inference tool for convolutional neural networks. Our optimization process on multicore ARM processors involves several high-level transformations of the original framework, such as the development and integration of Cython routines to exploit thread-level parallelism; the design and development of micro-kernels for the matrix multiplication, vectorized with ARM’s NEON intrinsics, that can accommodate layer fusion; and the appropriate selection of several cache configuration parameters tailored to the memory hierarchy of the target ARM processors. Our experiments evaluate both inference throughput (measured in processed images/s) and inference latency (i.e., time-to-response) as well as energy consumption per image when varying the level of thread parallelism and the processor power modes. The experiments with the new inference engine are reported for the ResNet50 v1.5 model on the ImageNet dataset from the MLPerf suite using the ARM v8.2 cores in the NVIDIA Jetson AGX Xavier board. These results show superior performance compared with the well-spread TFLite from Google and slightly inferior results when compared with ArmNN, the native library from ARM for DNN inference. Adrián Castelló 0001, Sergio Barrachina 0001, Manuel F. Dolz, Enrique S. Quintana-Ortí, Pau San Juan, Andrés Tomás |
J. Syst. Archit. | 1 |
| 2022 | BestOf: an online implementation selector for the training and inference of deep neural networksabstractAbstract Tuning and optimising the operations executed in deep learning frameworks is a fundamental task in accelerating the processing of deep neural networks (DNNs). However, this optimisation usually requires extensive manual efforts in order to obtain the best performance for each combination of tensor input size, layer type, and hardware platform. In this work, we present , a novel online auto-tuner that optimises the training and inference phases of DNNs. automatically selects at run time, and among the provided alternatives, the best performing implementation in each layer according to gathered profiling data. The evaluation of is performed on multi-core architectures for different DNNs using , a lightweight library for distributed training and inference. The experimental results reveal that the auto-tuner delivers the same or higher performance than that achieved using a static selection approach. Sergio Barrachina 0001, Adrián Castelló 0001, Manuel F. Dolz, Andrés Tomás |
J. Supercomput. | 2 |
| 2022 | A BLIS-like matrix multiplication for machine learning in the RISC-V ISA-based GAP8 processorabstractAbstract We address the efficient realization of matrix multiplication (gemm), with application in the convolution operator for machine learning, for the RISC-V core present in the GreenWaves GAP8 processor. Our approach leverages BLIS (Basic Linear Algebra Instantiation Software) to develop an implementation that (1) re-organizes the gemm algorithm adapting its micro-kernel to exploit the hardware-supported dot product kernel in the GAP8; (2) explicitly orchestrates the data transfers across the hierarchy of scratchpad memories via DMA (direct memory access); and (3) operates with integer arithmetic. Cristián Ramírez, Adrián Castelló 0001, Enrique S. Quintana-Ortí |
J. Supercomput. | 2 |
| 2021 | Performance Modeling for Distributed Training of Convolutional Neural NetworksabstractWe perform a theoretical analysis comparing the scalability of data versus model parallelism, applied to the distributed training of deep convolutional neural networks (CNNs), along five axes: batch size, node (floating-point) arithmetic performance, node memory bandwidth, network link bandwidth, and cluster dimension. Our study relies on analytical performance models that can be configured to reproduce the components and organization of the CNN model as well as the hardware configuration of the target distributed platform. In addition, we provide evidence of the accuracy of the analytical models by performing a validation against a Python library for distributed deep learning training. Adrián Castelló 0001, Mar Catalán, Manuel F. Dolz, José I. Mestre, Enrique S. Quintana-Ortí, José Duato |
PDP | 1 |
| 2021 | Evaluation of MPI Allreduce for Distributed Training of Convolutional Neural NetworksabstractTraining deep neural networks is a costly procedure, often performed via sophisticated deep learning frameworks on clusters of computers. As faster processor technologies are integrated into these cluster facilities (e.g., NVIDIA's graphics accelerators or Google's tensor processing units), the communication component of the training process rapidly becomes a performance bottleneck. In this paper, we offer a complete analysis of the key collective communication primitive for the distributed data-parallel training of convolutional network networks (CNNs) focused on three relevant instances of the Message Passing Interface (MPI): MPICH, OpenMPI, and IntelMPI. In addition, our experimental evaluation is extended to expose the practical impact of this collective primitive when the training is performed using TensorFlow+ Horovod on a 16-node cluster. Finally, the theoretical analysis is further refined to a number of accelerated cluster configurations that are emulated by adjusting the communication-arithmetic ratio of the training process. Adrián Castelló 0001, Mar Catalán, Manuel F. Dolz, José I. Mestre, Enrique S. Quintana-Ortí, José Duato |
PDP | 1 |
| 2021 | PyDTNN: A user-friendly and extensible framework for distributed deep learning
Sergio Barrachina 0001, Adrián Castelló 0001, Mar Catalán, Manuel F. Dolz, José I. Mestre |
J. Supercomput. | 2 |
| 2020 | High Performance and Portable Convolution Operators for Multicore ProcessorsabstractThe considerable impact of Convolutional Neural Networks on many Artificial Intelligence tasks has led to the development of various high performance algorithms for the convolution operator present in this type of networks. One of these approaches leverages the IM2COL transform followed by a general matrix multiplication (GEMM) in order to take advantage of the highly optimized realizations of the GEMM kernel in many linear algebra libraries. The main problems of this approach are 1) the large memory workspace required to host the intermediate matrices generated by the IM2COL transform; and 2) the time to perform the IM2COL transform, which is not negligible for complex neural networks. This paper presents a portable high performance convolution algorithm based on the BLIS realization of the GEMM kernel that avoids the use of the intermediate memory by taking advantage of the BLIS structure. In addition, the proposed algorithm eliminates the cost of the explicit IM2COL transform, while maintaining the portability and performance of the underlying realization of GEMM in BLIS. Pablo San Juan, Adrián Castelló 0001, Manuel F. Dolz, Pedro Alonso 0002, Enrique S. Quintana-Ortí |
SBAC-PAD | 2 |
| 2020 | Analysis of Threading Libraries for High Performance ComputingabstractWith the appearance of multi-/many core machines, applications and runtime systems have evolved in order to exploit the new on-node concurrency brought by new software paradigms. POSIX threads (Pthreads) was widely-adopted for that purpose and it remains as the most used threading solution in current hardware. Lightweight thread (LWT) libraries emerged as an alternative offering lighter mechanisms to tackle the massive concurrency of current hardware. In this article, we analyze in detail the most representative threading libraries including Pthread- and LWT-based solutions. In addition, to examine the suitability of LWTs for different use cases, we develop a set of microbenchmarks consisting of OpenMP patterns commonly found in current parallel codes, and we compare the results using threading libraries and OpenMP implementations. Moreover, we study the semantics offered by threading libraries in order to expose the similarities among different LWT application programming interfaces and their advantages over Pthreads. This article exposes that LWT libraries outperform solutions based on operating system threads when tasks and nested parallelism are required. Adrián Castelló 0001, Rafael Mayo 0002, Pavan Balaji, Enrique S. Quintana-Ortí, Antonio J. Peña |
IEEE Trans. Computers | 1 |
| 2019 | Theoretical Scalability Analysis of Distributed Deep Convolutional Neural NetworksabstractWe analyze the asymptotic performance of the training process of deep neural networks (NN) on clusters in order to determine the scalability. For this purpose, i) we assume a data parallel implementation of the training algorithm, which distributes the batches among the cluster nodes and replicates the model; ii) we leverage the roofline model to inspect the performance at the node level, taking into account the floating-point unit throughput and memory bandwidth; and iii) we consider distinct collective communication schemes that are optimal depending on the message size and underlying network interconnection topology. We then apply the resulting performance model to analyze the scalability of several well-known deep convolutional neural networks as a function of the batch size, node floating-point throughput, node memory bandwidth, cluster dimension, and link bandwidth. Adrián Castelló 0001, Manuel F. Dolz, Enrique S. Quintana-Ortí, José Duato |
CCGRID | 1 |
| 2019 | Analysis of model parallelism for distributed neural networksabstractWe analyze the performance of model parallelism applied to the training of deep neural networks on clusters. For this study, we elaborate a parameterized analytical performance model that captures the main computational and communication stages in distributed model parallel training. This model is then leveraged to assess the impact on the performance of four representative convolutional neural networks (CNNs) when varying the node throughput in terms of operations per second and memory bandwidth, the number of nodes of the cluster, the bandwidth of the network links, and algorithmic parameters such as the dimension of the batch. Adrián Castelló 0001, Manuel F. Dolz, Enrique S. Quintana-Ortí, José Duato |
EuroMPI | 1 |
| 2018 | On the adequacy of lightweight thread approaches for high-level parallel programming models
Adrián Castelló 0001, Rafael Mayo 0002, Kevin Sala, Vicenç Beltran 0001, Pavan Balaji, Antonio J. Peña |
Future Gener. Comput. Syst. | 1 |
| 2018 | Exploring the interoperability of remote GPGPU virtualization using rCUDA and directive-based programming models
Adrián Castelló 0001, Antonio J. Peña, Rafael Mayo 0002, Judit Planas, Enrique S. Quintana-Ortí, Pavan Balaji |
J. Supercomput. | 1 |
| 2018 | Argobots: A Lightweight Low-Level Threading and Tasking FrameworkabstractIn the past few decades, a number of user-level threading and tasking models have been proposed in the literature to address the shortcomings of OS-level threads, primarily with respect to cost and flexibility. Current state-of-the-art user-level threading and tasking models, however, either are too specific to applications or architectures or are not as powerful or flexible. In this paper, we present Argobots, a lightweight, low-level threading and tasking framework that is designed as a portable and performant substrate for high-level programming models or runtime systems. Argobots offers a carefully designed execution model that balances generality of functionality with providing a rich set of controls to allow specialization by end users or high-level programming models. We describe the design, implementation, and performance characterization of Argobots and present integrations with three high-level models: OpenMP, MPI, and colocated I/O services. Evaluations show that (1) Argobots, while providing richer capabilities, is competitive with existing simpler generic threading runtimes; (2) our OpenMP runtime offers more efficient interoperability capabilities than production OpenMP runtimes do; (3) when MPI interoperates with Argobots instead of Pthreads, it enjoys reduced synchronization costs and better latency-hiding capabilities; and (4) I/O services with Argobots reduce interference with colocated applications while achieving performance competitive with that of a Pthreads approach. Abdelhalim Amer, Pavan Balaji, Cyril Bordage, George Bosilca, Alex Brooks, Philip H. Carns, Adrián Castelló 0001, Damien Genet, Thomas Hérault, Shintaro Iwasaki, Prateek Jindal, Laxmikant V. Kalé, Sriram Krishnamoorthy, Jonathan Lifflander, Huiwei Lu, Esteban Meneses, Marc Snir, Yanhua Sun, Kenjiro Taura, Pete Beckman |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2017 | GLT: A Unified API for Lightweight Thread Libraries
Adrián Castelló 0001, Rafael Mayo 0002, Pavan Balaji, Enrique S. Quintana-Ortí, Antonio J. Peña |
Euro-Par | 1 |
| 2017 | GLTO: On the Adequacy of Lightweight Thread Approaches for OpenMP ImplementationsabstractOpenMP is the de facto standard application programming interface (API) for on-node parallelism. The most popular OpenMP runtimes rely on POSIX threads (pthreads) implementations that offer an excellent performance for coarse-grained parallelism and match perfectly with the current hardware. However, a recent trend in runtimes/applications points in the direction of leveraging massive on-node parallelism in conjunction with fine-grained and dynamic scheduling paradigms. It has been demonstrated that lightweight thread (LWT) solutions are more appropriate for these new parallel paradigms. We have developed GLTO, an OpenMP implementation over the recently-emerged Generic Lightweight Threads (GLT) API. GLT exports a common API for LWT libraries that offers the possibility of running the same application over different native LWT solutions. In this paper we use GLTO to analyze different scenarios where OpenMP implementations may benefit from the use of either LWT or pthreads. Our study reveals that none of the threading approaches obtains the best performance in all the scenarios, but that there are important gaps among them. Adrián Castelló 0001, Rafael Mayo 0002, Pavan Balaji, Enrique S. Quintana-Ortí, Antonio J. Peña |
ICPP | 1 |
| 2016 | Enabling GPU Virtualization in Cloud EnvironmentsabstractThe use of accelerators, such as graphics processing units (GPUs), to reduce the execution time of compute-intensive applications has become popular during the past few years. These devices increment the computational power of a node thanks to their parallel architecture. This trend has led cloud service providers as Amazon or middlewares such as OpenStack to add virtual machines (VMs) including GPUs to their facilities instances. To fulfill these needs, the guest hosts must be equipped with GPUs which, unfortunately, will be barely utilized if a non GPU-enabled VM is running in the host. The solution presented in this work is based on GPU virtualization and shareability in order to reach an equilibrium between service supply and the applicationsâ?? demand of accelerators. Concretely, we propose to decouple real GPUs from the nodes by using the virtualization technology rCUDA. With this software configuration, GPUs can be accessed from any VM avoiding the need of placing a physical GPUs in each guest host. Moreover, we study the viability of this approach using a public cloud service configuration, and we develop a module for OpenStack in order to add support for the virtualized devices and the logic to manage them. The results demonstrate this is a viable configuration which adds flexibility to current and well-known cloud solutions. Sergio Iserte, Francisco J. Clemente-Castelló, Adrián Castelló 0001, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
CLOSER (2) | 3 |
| 2016 | A Review of Lightweight Thread Approaches for High Performance ComputingabstractHigh-level, directive-based solutions are becoming the programming models (PMs) of the multi/many-core architectures. Several solutions relying on operating system (OS) threads perfectly work with a moderate number of cores. However, exascale systems will spawn hundreds of thousands of threads in order to exploit their massive parallel architectures and thus conventional OS threads are too heavy for that purpose. Several lightweight thread (LWT) libraries have recently appeared offering lighter mechanisms to tackle massive concurrency. In order to examine the suitability of LWTs in high-level runtimes, we develop a set of microbenchmarks consisting of commonly-found patterns in current parallel codes. Moreover, we study the semantics offered by some LWT libraries in order to expose the similarities between different LWT application programming interfaces. This study reveals that a reduced set of LWT functions can be sufficient to cover the common parallel code patterns andthat those LWT libraries perform better than OS threads-based solutions in cases where task and nested parallelism are becoming more popular with new architectures. Adrián Castelló 0001, Antonio J. Peña, Rafael Mayo 0002, Pavan Balaji, Enrique S. Quintana-Ortí |
CLUSTER | 1 |
| 2015 | Exploring the Suitability of Remote GPGPU Virtualization for the OpenACC Programming Model Using rCUDAabstractOpenACC is an application programming interface (API) that aims to unleash the power of heterogeneous systems composed of CPUs and accelerators such as graphic processing units (GPUs) or Intel Xeon Phi coprocessors. This directive-based programming model is intended to enable developers to accelerate their application's execution with much less effort. Coprocessors offer significant computing power but in many cases these devices remain largely underused because not all parts of applications match the accelerator architecture. Remote accelerator virtualization frameworks introduce a means to address this problem. In particular, the remote CUDA virtualization middleware rCUDA provides transparent remote access to any GPU installed in a cluster. Combining these two technologies, OpenACC and rCUDA, in a single scenario is naturally appealing. In this work we explore how the different OpenACC directives behave on top of a remote GPGPU virtualization technology in two different hardware configurations. Our experimental evaluation reveals favorable performance results when the two technologies are combined, showing low overhead and similar scaling factors when executing OpenACC-enabled directives. Adrián Castelló 0001, Antonio J. Peña, Rafael Mayo 0002, Pavan Balaji, Enrique S. Quintana-Ortí |
CLUSTER | 1 |
| 2015 | Improving the user experience of the rCUDA remote GPU virtualization frameworkabstractSummary Graphics processing units (GPUs) are being increasingly embraced by the high‐performance computing community as an effective way to reduce execution time by accelerating parts of their applications. remote CUDA (rCUDA) was recently introduced as a software solution to address the high acquisition costs and energy consumption of GPUs that constrain further adoption of this technology. Specifically, rCUDA is a middleware that allows a reduced number of GPUs to be transparently shared among the nodes in a cluster. Although the initial prototype versions of rCUDA demonstrated its functionality, they also revealed concerns with respect to usability, performance, and support for new CUDA features. In response, in this paper, we present a new rCUDA version that (1) improves usability by including a new component that allows an automatic transformation of any CUDA source code so that it conforms to the needs of the rCUDA framework, (2) consistently features low overhead when using remote GPUs thanks to an improved new communication architecture, and (3) supports multithreaded applications and CUDA libraries. As a result, for any CUDA‐compatible program, rCUDA now allows the use of remote GPUs within a cluster with low overhead, so that a single application running in one node can use all GPUs available across the cluster, thereby extending the single‐node capability of CUDA. Copyright © 2014 John Wiley & Sons, Ltd. Carlos Reaño, Federico Silla, Adrián Castelló 0001, Antonio J. Peña, Rafael Mayo 0002, Enrique S. Quintana-Ortí, José Duato |
Concurr. Comput. Pract. Exp. | 3 |
| 2014 | Boosting the performance of remote GPU virtualization using InfiniBand connect-IB and PCIe 3.0abstractA clear trend has emerged involving the acceleration of scientific applications by using GPUs. However, the capabilities of these devices are still generally underutilized. Remote GPU virtualization techniques can help increase GPU utilization rates, while reducing acquisition and maintenance costs. The overhead of using a remote GPU instead of a local one is introduced mainly by the difference in performance between the internode network and the intranode PCIe link. In this paper we show how using the new InfiniBand Connect-IB network adapters (attaining similar throughput to that of the most recently emerged GPUs) boosts the performance of remote GPU virtualization, reducing the overhead to a mere 0.19% in the application tested. Carlos Reaño, Federico Silla, Antonio J. Peña, Gilad Shainer, Scot Schultz, Adrián Castelló 0001, Enrique S. Quintana-Ortí, José Duato |
CLUSTER | 6 |
| 2014 | SLURM Support for Remote GPU Virtualization: Implementation and Performance StudyabstractSLURM is a resource manager that can be leveraged to share a collection of heterogeneous resources among the jobs in execution in a cluster. However, SLURM is not designed to handle resources such as graphics processing units (GPUs). Concretely, although SLURM can use a generic resource plugin (GRes) to manage GPUs, with this solution the hardware accelerators can only be accessed by the job that is in execution on the node to which the GPU is attached. This is a serious constraint for remote GPU virtualization technologies, which aim at providing a user-transparent access to all GPUs in cluster, independently of the specific location of the node where the application is running with respect to the GPU node. In this work we introduce a new type of device in SLURM, "rgpu", in order to gain access from any application node to any GPU node in the cluster using rCUDA as the remote GPU virtualization solution. With this new scheduling mechanism, a user can access any number of GPUs, as SLURM schedules the tasks taking into account all the graphics accelerators available in the complete cluster. We present experimental results that show the benefits of this new approach in terms of increased flexibility for the job scheduler. Sergio Iserte, Adrián Castelló 0001, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Federico Silla, José Duato, Carlos Reaño, Javier Prades |
SBAC-PAD | 2 |