EDBT 2026 Demo / reviewers in the wild / expert
Andrés Tomás
dblp:118/5331 · also Andrés E. Tomás, Andrés E. Tomás Dominguez, Andrés Tomás Dominguez
· DBLP profile ↗
19ranked-venue papers
5as first author
6since 2021 · last 2023
0000-0003-3969-2174ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 4 first-author · 6 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Sparse matrix-vector and matrix-multivector products for the truncated SVD on graphics processorsabstractSummary Many practical algorithms for numerical rank computations implement an iterative procedure that involves repeated multiplications of a vector, or a collection of vectors, with both a sparse matrix and its transpose. Unfortunately, the realization of these sparse products on current high performance libraries often deliver much lower arithmetic throughput when the matrix involved in the product is transposed. In this work, we propose a hybrid sparse matrix layout, named CSRC, that combines the flexibility of some well‐known sparse formats to offer a number of appealing properties: (1) CSRC can be obtained at low cost from the popular CSR (compressed sparse row) format; (2) CSRC has similar storage requirements as CSR; and especially, (3) the implementation of the sparse product kernels delivers high performance for both the direct product and its transposed variant on modern graphics accelerators thanks to a significant reduction of atomic operations compared to a conventional implementation based on CSR. This solution thus renders considerably higher performance when integrated into an iterative algorithm for the truncated singular value decomposition (SVD), such as the randomized SVD or, as demonstrated in the experimental results, the block Golub–Kahan–Lanczos algorithm. José Ignacio Aliaga, Hartwig Anzt, Enrique S. Quintana-Ortí, Andrés Tomás |
Concurr. Comput. Pract. Exp. | 4 |
| 2023 | Reformulating the direct convolution for high-performance deep learning inference on ARM processorsabstractWe present two high-performance implementations of the convolution operator via the direct algorithm that outperform the so-called lowering approach based on the im2col transform plus the gemm kernel on an ARMv8-based processor. One of our methods presents the additional advantage of zero-memory overhead while the other employs an additional yet rather moderate workspace, substantially smaller than that required by the im2col+gemm solution. In contrast with a previous implementation of a similar zero-memory overhead direct convolution, this work exhibits the key advantage of preserving the conventional NHWC data layout for the input/output activations of the convolution layers. Sergio Barrachina 0001, Adrián Castelló 0001, Manuel F. Dolz, Tze Meng Low, Héctor Martínez 0002, Enrique S. Quintana-Ortí, Upasana Sridhar, Andrés Tomás |
J. Syst. Archit. | 8 |
| 2023 | Performance-energy trade-offs of deep learning convolution algorithms on ARM processorsabstractAbstract In this work, we assess the performance and energy efficiency of high-performance codes for the convolution operator, based on the direct, explicit/implicit lowering and Winograd algorithms used for deep learning (DL) inference on a series of ARM-based processor architectures. Specifically, we evaluate the NVIDIA Denver2 and Carmel processors, as well as the ARM Cortex-A57 and Cortex-A78AE CPUs as part of a recent set of NVIDIA Jetson platforms. The performance–energy evaluation is carried out using the ResNet-50 v1.5 convolutional neural network (CNN) on varying configurations of convolution algorithms, number of threads/cores, and operating frequencies on the tested processor cores. The results demonstrate that the best throughput is obtained on all platforms with the Winograd convolution operator running on all the cores at their highest frequency. However, if the goal is to reduce the energy footprint, there is no rule of thumb for the optimal configuration. Manuel F. Dolz, Sergio Barrachina 0001, Héctor Martínez 0002, Adrián Castelló 0001, Antonio M. Vidal, Germán Fabregat, Andrés Tomás |
J. Supercomput. | 7 |
| 2022 | Compression and load balancing for efficient sparse matrix-vector product on multicore processors and graphics processing unitsabstractSummary We contribute to the optimization of the sparse matrix‐vector product by introducing a variant of the coordinate sparse matrix format that balances the workload distribution and compresses both the indexing arrays and the numerical information. Our approach is multi‐platform, in the sense that the realizations for (general‐purpose) multicore processors as well as graphics accelerators (GPUs) are built upon common principles, but differ in the implementation details, which are adapted to avoid thread divergence in the GPU case or maximize compression element‐wise (i.e., for each matrix entry) for multicore architectures. Our evaluation on the two last generations of NVIDIA GPUs as well as Intel and AMD processors demonstrate the benefits of the new kernels when compared with the optimized implementations of the sparse matrix‐vector product in NVIDIA's cuSPARSE and Intel's MKL, respectively. José Ignacio Aliaga, Hartwig Anzt, Thomas Grützmacher, Enrique S. Quintana-Ortí, Andrés Tomás |
Concurr. Comput. Pract. Exp. | 5 |
| 2022 | High performance and energy efficient inference for deep learning on multicore ARM processors using general optimization techniques and BLISabstractWe evolve PyDTNN, a framework for distributed parallel training of Deep Neural Networks (DNNs), into an efficient inference tool for convolutional neural networks. Our optimization process on multicore ARM processors involves several high-level transformations of the original framework, such as the development and integration of Cython routines to exploit thread-level parallelism; the design and development of micro-kernels for the matrix multiplication, vectorized with ARM’s NEON intrinsics, that can accommodate layer fusion; and the appropriate selection of several cache configuration parameters tailored to the memory hierarchy of the target ARM processors. Our experiments evaluate both inference throughput (measured in processed images/s) and inference latency (i.e., time-to-response) as well as energy consumption per image when varying the level of thread parallelism and the processor power modes. The experiments with the new inference engine are reported for the ResNet50 v1.5 model on the ImageNet dataset from the MLPerf suite using the ARM v8.2 cores in the NVIDIA Jetson AGX Xavier board. These results show superior performance compared with the well-spread TFLite from Google and slightly inferior results when compared with ArmNN, the native library from ARM for DNN inference. Adrián Castelló 0001, Sergio Barrachina 0001, Manuel F. Dolz, Enrique S. Quintana-Ortí, Pau San Juan, Andrés Tomás |
J. Syst. Archit. | 6 |
| 2022 | BestOf: an online implementation selector for the training and inference of deep neural networksabstractAbstract Tuning and optimising the operations executed in deep learning frameworks is a fundamental task in accelerating the processing of deep neural networks (DNNs). However, this optimisation usually requires extensive manual efforts in order to obtain the best performance for each combination of tensor input size, layer type, and hardware platform. In this work, we present , a novel online auto-tuner that optimises the training and inference phases of DNNs. automatically selects at run time, and among the provided alternatives, the best performing implementation in each layer according to gathered profiling data. The evaluation of is performed on multi-core architectures for different DNNs using , a lightweight library for distributed training and inference. The experimental results reveal that the auto-tuner delivers the same or higher performance than that achieved using a static selection approach. Sergio Barrachina 0001, Adrián Castelló 0001, Manuel F. Dolz, Andrés Tomás |
J. Supercomput. | 4 |
| 2020 | Tall-and-skinny QR factorization with approximate Householder reflectors on graphics processors
Andrés Tomás, Enrique S. Quintana-Ortí |
J. Supercomput. | 1 |
| 2019 | Cholesky and Gram-Schmidt Orthogonalization for Tall-and-Skinny QR Factorizations on Graphics Processors
Andrés Tomás, Enrique S. Quintana-Ortí |
Euro-Par | 1 |
| 2019 | Dynamic look-ahead in the reduction to band form for the singular value decomposition
Andrés Tomás, Rafael Rodríguez-Sánchez 0001, Sandra Catalán, Rocío Carratalá-Sáez, Enrique S. Quintana-Ortí |
Parallel Comput. | 1 |
| 2019 | FloatX: A C++ Library for Customized Floating-Point ArithmeticabstractWe present FloatX (Float eXtended), a C ++ framework to investigate the effect of leveraging customized floating-point formats in numerical applications. FloatX formats are based on binary IEEE 754 with smaller significand and exponent bit counts specified by the user. Among other properties, FloatX facilitates an incremental transformation of the code, relies on hardware-supported floating-point types as back-end to preserve efficiency, and incurs no storage overhead. The article discusses in detail the design principles, programming interface, and datatype casting rules behind FloatX. Furthermore, it demonstrates FloatX’s usage and benefits via several case studies from well-known numerical dense linear algebra libraries, such as BLAS and LAPACK; the Ginkgo library for sparse linear systems; and two neural network applications related with image processing and text recognition. Goran Flegar, Florian Scheidegger, Vedran Novakovic, Giovanni Mariani, Andrés Tomás, Cristiano Malossi, Enrique S. Quintana-Ortí |
ACM Trans. Math. Softw. | 5 |
| 2018 | The transprecision computing paradigm: Concept, design, and applicationsabstractGuaranteed numerical precision of each elementary step in a complex computation has been the mainstay of traditional computing systems for many years. This era, fueled by Moore's law and the constant exponential improvement in computing efficiency, is at its twilight: from tiny nodes of the Internet-of-Things, to large HPC computing centers, sub-picoJoule/operation energy efficiency is essential for practical realizations. To overcome the power wall, a shift from traditional computing paradigms is now mandatory. In this paper we present the driving motivations, roadmap, and expected impact of the European project OPRECOMP. OPRECOMP aims to (i) develop the first complete transprecision computing framework, (ii) apply it to a wide range of hardware platforms, from the sub-milliWatt up to the MegaWatt range, and (iii) demonstrate impact in a wide range of computational domains, spanning IoT, Big Data Analytics, Deep Learning, and HPC simulations. By combining together into a seamless design transprecision advances in devices, circuits, software tools, and algorithms, we expect to achieve major energy efficiency improvements, even when there is no freedom to relax end-to-end application quality of results. Indeed, OPRECOMP aims at demolishing the ultra-conservative “precise” computing abstraction, replacing it with a more flexible and efficient one, namely transprecision computing. Cristiano Malossi, Michael Schaffner, Anca Mariana Molnos, Luca Gammaitoni, Giuseppe Tagliavini, Andrew P. J. Emerson, Andrés Tomás, Dimitrios S. Nikolopoulos, Eric Flamand, Norbert Wehn |
DATE | 7 |
| 2018 | Fast Blocking of Householder Reflectors on Graphics ProcessorsabstractWe revisit an alternative representation to the compact WY transform for the accumulation (blocking) of Householder reflectors that exhibits the same numerical stability and is composed of efficient computational kernels from Level-3 Basic Linear Algebra Subprograms (BLAS) in contrast with the Level-2 BLAS that are utilized for the construction of the conventional compact WY representation. For the orthogonal reduction to condensed forms on multicore platforms equipped with a fast graphics processing unit (GPU), (or when there is a notable gap in performance between the multicore processors and the graphics accelerator,) our approach removes the assembly of the accumulation from the critical path of the algorithm. This comes as a consequence of accelerating this operation via the use of Level-3 BLAS, moving this computation to the GPU, and allowing the use of larger algorithmic block sizes. Our experiments with the alternative orthogonal representation show considerable speed-ups, which can be in the range 20-40% on recent GPUs when compared with the codes in MAGMA. Andrés Tomás, Enrique S. Quintana-Ortí |
PDP | 1 |
| 2017 | Selecting the optimal buffer management for opportunistic networks both in pedestrian and vehicular contextsabstractOpportunistic networks are a form of mobile ad-hoc networks that exploit users mobility to provide data sharing. Individual nodes store, carry and forward messages using direct communication between devices; no fixed infrastructure of communication is required. The final goal is to reach the totality of the participating users. The efficiency of message diffusion depends on the users' mobility, the transfer time, and the device buffer management. User's mobility determines the number of contacts and their duration while the transfer time depends on the size of the interchanged messages and the channel throughput. The device buffer management is a critical factor since, due to its limited size, the chosen policy for forwarding messages when a connection is available and for message dropping when the buffer is full can impact greatly on the diffusion of the messages. In this paper we focus on this issue, evaluating the impact of different buffer management approaches in two different scenarios: pedestrian and vehicular. The results show that the best buffer management is the combination of smallest message forwarding and largest message dropping. We also show that a TTL (Time to Live) of 24 hours allows the best diffusion because it corresponds to the daily movement pattern of the nodes. Jorge Herrera-Tapia, Enrique Hernández-Orallo, Andrés Tomás, Pietro Manzoni, Carlos T. Calafate, Juan-Carlos Cano |
CCNC | 3 |
| 2017 | On the impact of urban intersection characteristics in vehicular to vehicular (V2V) communicationsabstractIntersection management is one of the challenges of vehicular communications in urban environments. Maximizing message delivery at intersections requires gaining awareness about when the vehicle is at the center of intersection, as buildings tend to obstruct wireless signals especially in the 5 GHz band. This paper addresses the issue of communications on different types of intersections by characterizing each intersection type through both real experiments and analytic modeling. The study has found that antenna location, distance to the intersection and line of sight conditions are all important factors for vehicular communications at intersections. Furthermore, this communication success also depends on the intersections characteristics, as different degrees of obstruction cause different delivery probabilities of sending messages. Seilendria A. Hadiwardoyo, Andrés Tomás, Enrique Hernández-Orallo, Carlos T. Calafate, Juan-Carlos Cano, Pietro Manzoni |
IWCMC | 2 |
| 2016 | MuffinEc: Error correction for de Novo assembly via greedy partitioning and sequence alignment
Andy S. Alic, Andrés Tomás, Ignacio Medina, Ignacio Blanquer |
Inf. Sci. | 2 |
| 2014 | A Fast Sparse Block Circulant Matrix Vector Product
Eloy Romero, Andrés Tomás, Antonio Soriano-Asensi, Ignacio Blanquer |
Euro-Par | 2 |
| 2012 | Advancing Large Scale Many-Body QMC Simulations on GPU Accelerated Multicore SystemsabstractThe Determinant Quantum Monte Carlo (DQMC) method is one of the most powerful approaches for understanding properties of an important class of materials with strongly interacting electrons, including magnets and superconductors. It treats these interactions exactly, but the solution of a system of N electrons must be extrapolated to bulk values. Currently N 500 is state-of-the-art. Increasing N is required before DQMC can be used to model newly synthesized materials like functional multilayers. DQMC requires millions of linear algebra computations of order N matrices and scales as N3. DQMC cannot exploit parallel distributed memory computers efficiently due to limited scalability with the small matrix sizes and stringent procedures for numerical stability. Today, the combination of multisocket multicore processors and GPUs provides widely available platforms with new opportunities for DQMC parallelization. The kernel of DQMC, the calculation of the Green's function, involves long products of matrices. For numerical stability, these products must be computed using graded decompositions generated by the QR decomposition with column pivoting. The high communication overhead of pivoting limits parallel efficiency. In this paper, we propose a novel approach that exploits the progressive graded structure to reduce the communication costs of pivoting. We show that this method preserves the same numerical stability and achieves 70% performance of highly optimized DGEMM on a two-socket six-core Intel processor. We have integrated this new method and other parallelization techniques into QUEST, a modern DQMC simulation package. Using 36 hours on this Intel processor, we are able to compute accurately the magnetic properties and Fermi surface of a system of N = 1024 electrons. This simulation is almost an order of magnitude more difficult than N 500, owing to the N3 scaling. This increase in system size will allow, for the first time, the computation of the magnetic and transport properties of layered materials with DQMC. In addition, we show preliminary results which further accelerate DQMC simulations by using GPU processors. Andrés Tomás, Chia-Chen Chang, Richard Scalettar, Zhaojun Bai |
IPDPS | 1 |
| 2012 | Using GPUs for the Exact Alignment of Short-Read Genetic Sequences by Means of the Burrows-Wheeler TransformabstractGeneral Purpose Graphic Processing Units (GPGPUs) constitute an inexpensive resource for computing-intensive applications that could exploit an intrinsic fine-grain parallelism. This paper presents the design and implementation in GPGPUs of an exact alignment tool for nucleotide sequences based on the Burrows-Wheeler Transform. We compare this algorithm with state-of-the-art implementations of the same algorithm over standard CPUs, and considering the same conditions in terms of I/O. Excluding disk transfers, the implementation of the algorithm in GPUs shows a speedup larger than 12, when compared to CPU execution. This implementation exploits the parallelism by concurrently searching different sequences on the same reference search tree, maximizing memory locality and ensuring a symmetric access to the data. The paper describes the behavior of the algorithm in GPU, showing a good scalability in the performance, only limited by the size of the GPU inner memory. José Salavert Torres, Ignacio Blanquer, Andrés Tomás, Vicente Hernández, Ignacio Medina, Joaquín Tárraga, Joaquín Dopazo |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2007 | Parallel Arnoldi eigensolvers with enhanced scalability via global communications rearrangement
Vicente Hernández, José E. Román, Andrés Tomás |
Parallel Comput. | 3 |