José R. Herrero 0001

dblp:60/2853 · also Jose Ramón Herrero Zaragoza, José Ramon Herrero 0001 · DBLP profile ↗
← Back
23ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0002-4060-367XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
YearPublicationVenuePosition
2026 Exploiting mixed-precision redundancy for soft-error detection in LU decomposition
Nima Sahraneshinsamani, Sandra Catalán, José R. Herrero 0001
J. Supercomput.3
2025 Mixed-precision pre-pivoting strategy for the LU factorization
abstract
Abstract This paper investigates the efficient application of half-precision floating-point (FP16) arithmetic on GPUs for boosting LU decompositions in double (FP64) precision. Addressing the motivation to enhance computational efficiency, we introduce two novel algorithms: Pre-Pivoted LU (PRP) and Mixed-precision Panel Factorization (MPF). Deployed in both hybrid CPU-GPU setups and native GPU-only configurations, PRP identifies pivot lists through LU decomposition computed in reduced precision and subsequently reorders matrix rows in FP64 precision before executing LU decomposition without pivoting. Two variants of PRP, namely hPRP and xPRP, are introduced, differing in their computation of pivot lists in full half-precision or mixed half-single precision. The MPF algorithm generates FP64 LU factorization while internally utilizing hPRP for panel factorization, showcasing accuracy on par with standard DGETRF but with superior speed. The study further explores auxiliary functions required for the native mode implementation of PRP variants and MPF.
Nima Sahraneshinsamani, Sandra Catalán, José R. Herrero 0001
J. Supercomput.3
2023 Fine-grain task-parallel algorithms for matrix factorizations and inversion on many-threaded CPUs
abstract
Abstract We extend a two‐level task partitioning previously applied to the inversion of dense matrices via Gauss–Jordan elimination to the more challenging QR factorization as well as the initial orthogonal reduction to band form found in the singular value decomposition. Our new task‐parallel algorithms leverage the tasking mechanism currently available in OpenMP to exploit “nested” task parallelism, with a first outer level that operates on matrix panels and a second inner level that processes the matrix either by ‐panels or by tiles, in order to expose a large number of independent tasks. We present a detailed performance analysis, including execution traces, which shows that the two‐level refinement into fine grain tasks allows for an improved load balancing and delivers high performance on current general‐purpose many‐core processors (CPUs) from Intel and AMD.
Sandra Catalán, José R. Herrero 0001, Francisco D. Igual, Enrique S. Quintana-Ortí, Rafael Rodríguez-Sánchez 0001
Concurr. Comput. Pract. Exp.2
2023 Programming parallel dense matrix factorizations and inversion for new-generation NUMA architectures
abstract
We propose a methodology to address the programmability issues derived from the emergence of new-generation shared-memory NUMA architectures. For this purpose, we employ dense matrix factorizations and matrix inversion (DMFI) as a use case, and we target two modern architectures (AMD Rome and Huawei Kunpeng 920) that exhibit configurable NUMA topologies. Our methodology pursues performance portability across different NUMA configurations by proposing multi-domain implementations for DMFI plus a hybrid task- and loop-level parallelization that configures multi-threaded executions to fix core-to-data binding, exploiting locality at the expense of minor code modifications. In addition, we introduce a generalization of the multi-domain implementations for DMFI that offers support for virtually any NUMA topology in present and future architectures. Our experimentation on the two target architectures for three representative dense linear algebra operations validates the proposal, reveals insights on the necessity of adapting both the codes and their execution to improve data access locality, and reports performance across architectures and inter- and intra-socket NUMA configurations competitive with state-of-the-art message-passing implementations, maintaining the ease of development usually associated with shared-memory programming.
Sandra Catalán, Francisco D. Igual, José R. Herrero 0001, Rafael Rodríguez-Sánchez 0001, Enrique S. Quintana-Ortí
J. Parallel Distributed Comput.3
2022 NUMA-Aware Dense Matrix Factorizations and Inversion with Look-Ahead on Multicore Processors
abstract
We address the efficient design and implementation of dense matrix factorizations and inversion (DMFI) on modern multicore processors with several NUMA (non-uniform memory access) nodes. Our approach enhances the DMFI routines with a look-ahead strategy, in order to overcome the “panel factorization bottleneck”. In addition, it exploits both hybrid task- and loop-level parallelizations while taking into account the NUMA organization of the memory hierarchy. The experiments on a Huawei Kunpeng-based server, with two sockets and 48 cores per socket, for three representative dense linear algebra operations, expose the necessity of adapting both the codes and their execution environment parameters to improve data access locality. The results of these changes deliver performance across inter- and intra-socket NUMA configurations superior to that of reference implementations from state-of-the-art libraries for this platform.
Sandra Catalán, Francisco D. Igual, Rafael Rodríguez-Sánchez 0001, José R. Herrero 0001, Enrique S. Quintana-Ortí
SBAC-PAD4
2022 A distributed Monte Carlo based linear algebra solver applied to the analysis of large complex networks
Filipe Magalhães, José Monteiro 0001, Juan A. Acebrón, José R. Herrero 0001
Future Gener. Comput. Syst.4
2018 Two-sided orthogonal reductions to condensed forms on asymmetric multicore processors
Pedro Alonso 0002, Sandra Catalán, José R. Herrero 0001, Enrique S. Quintana-Ortí, Rafael Rodríguez-Sánchez 0001
Parallel Comput.3
2018 Energy balance between voltage-frequency scaling and resilience for linear algebra routines on low-power multicore architectures
Sandra Catalán, José R. Herrero 0001, Enrique S. Quintana-Ortí, Rafael Rodríguez-Sánchez 0001
Parallel Comput.2
2018 Static scheduling of the LU factorization with look-ahead on asymmetric multicore processors
Sandra Catalán, José R. Herrero 0001, Enrique S. Quintana-Ortí, Rafael Rodríguez-Sánchez 0001
Parallel Comput.2
2017 Low-latency multi-threaded ensemble learning for dynamic big data streams
abstract
Real-time mining of evolving data streams involves new challenges when targeting today's application domains such as the Internet of the Things: increasing volume, velocity and volatility requires data to be processed on-the-fly with fast reaction and adaptation to changes. This paper presents a high performance scalable design for decision trees and ensemble combinations that makes use of the vector SIMD and multicore capabilities available in modern processors to provide the required throughput and accuracy. The proposed design offers very low latency and good scalability with the number of cores on commodity hardware when compared to other state-of-the art implementations. On an Intel i7-based system, processing a single decision tree is 6× faster than MOA (Java), and 7× faster than StreamDM (C++), two well-known reference implementations. On the same system, the use of the 6 cores (and 12 hardware threads) available allow to process an ensemble of 100 learners 85× faster that MOA while providing the same accuracy. Furthermore, our solution is highly scalable: on an Intel Xeon socket with large core counts, the proposed ensemble design achieves up to 16× speedup when employing 24 cores with respect to a single threaded execution.
Diego Marron, Eduard Ayguadé, José R. Herrero 0001, Jesse Read, Albert Bifet
IEEE BigData3
2016 Echo State Hoeffding Tree Learning
abstract
Nowadays, real-time classification of Big Data streams is becoming essential in a variety of application domains. While decision trees are powerful and easy-to-deploy approaches for accurate and fast learning from data streams, they are unable to capture the strong temporal dependences typically present in the input data. Recurrent Neural Networks are an alternative solution that include an internal memory to capture these temporal dependences; however their training is computationally very expensive and with slow convergence, requiring a large number of hyper-parameters to tune. Reservoir Computing was proposed to reduce the computation requirements of the training phase but still include a feed-forward layer which requires a large number of parameters to tune. In this work we propose a novel architecture for real-time classification based on the combination of a Reservoir and a decision tree. This combination reduces the number of hyper-parameters while still maintaining the good temporal properties of recurrent neural networks. The capabilities of the proposed architecture to learn some typical string-based functions with strong temporal dependences are evaluated in the paper. We show how the new architecture is able to incrementally learn these functions in real-time with fast adaptation to unknown sequences. And we study the influence of the reduced number of hyper-parameters in the behaviour of the proposed solution.
Diego Marron, Jesse Read, Albert Bifet, Talel Abdessalem, Eduard Ayguadé, José R. Herrero 0001
ACML6
2015 Parallel computing on graphics processing units and heterogeneous platforms
abstract
This special issue contributes to the field of parallel computing on graphics processing units and heterogeneous platforms with extended versions of selected papers from two workshops, namely the 3rd Minisymposium on GPU Computing—held as part of the 10th International Conference on Parallel Processing and Applied Mathematics (PPAM 2013) in Warsaw, Poland—and the 11th International Workshop on Algorithms, Models and Tools for Parallel Computing on Heterogeneous Platforms (HeteroPar'2013)—held in conjunction with the Euro-Par 2013 conference in Aachen, Germany. During the past decade, high-performance computing evolved toward multi-core and many-core architectures. General-purpose processors feature now dozens of coarse-grain (complex) cores each with four to eight SIMD lanes for parallel computation and multi-channel memory buses for high bandwidth. Hardware accelerators such as graphics processing units (GPUs) have also a two-stage design with multiple coarse units that contain an even higher number of SIMD lanes and wider memory buses for higher bandwidth. Therefore, the adoption of hardware accelerators is rapidly advancing in performance sensitive areas. They are particularly relevant in high-throughput disciplines such as high-quality 3D computer graphics and vision, real-time data stream processing, and high-performance scientific computing. The main reason behind this trend is that these accelerators can potentially yield speedups and energy savings orders of magnitude higher than those obtained with optimized implementations for general-purpose CPU cores. A clear indicator of this trend is the prevalence of these accelerators in the supercomputing systems in the top positions of both the TOP500 and Green500 lists. As a result, during the past few years, these architectures have become powerful, capable, and inexpensive mainstream coprocessors, useful for a wide variety of applications. Furthermore, they are nowadays present in a large variety of machines, ranging from low-end single user-platforms to supercomputers. However, the benefits of heterogeneous systems do not come ‘for free’: scientists using these platforms have to deal not only with multiple parallelism levels, but also with the programmability differences of available accelerators. To address these challenges, we observe the development of a very rich environment for their programming, particularly in comparison with the restricted landscape of only a few years ago. A key criterion to characterize the new high-level programming tools and libraries for these devices is their positioning within the triangle of performance, coding comfort and specialization. The spectrum ranges from high-performance building blocks for common numeric or discrete transformations, to domain-specific libraries that facilitate the solution of a certain class of problems, and to general high-level abstractions targeted toward increasing programmers' productivity. In summary, the advances both in the hardware and in the programmability of accelerators, coupled with their potentially appealing performance/power ratio for a wide range of applications, have pushed organizations to invest in heterogeneous systems that include accelerators and have motivated researchers to port their algorithms to such systems and develop novel tools to facilitate their usage. This special issue contributes to this important field with extended and carefully reviewed versions of selected papers from two workshops, namely the 3rd Minisymposium on GPU Computing, which was held as part of the 10th International Conference on Parallel Processing and Applied Mathematics (PPAM 2013) in Warsaw, and the 11th International Workshop on Algorithms, Models and Tools for Parallel Computing on Heterogeneous Platforms (HeteroPar'2013), which was held in conjunction with the Euro-Par 2013 conference in Aachen. Nineteen papers were published in the Conference Proceedings of these two events, after one or two review rounds. Extended versions of selected papers went through two new review rounds, resulting in the acceptance of the nine papers contained in this special issue. The topics offer a good cross section of current challenges on heterogeneous computing: further abstractions in the programming model, advances in the scheduling of tasks or their communications, improvements of basic parallel algorithms in discrete mathematics and linear algebra, and the utilization of the parallel processing power of GPUs for real-world applications. In 1, the authors propose the design of a directory, along with a reduced runtime application binary interface, to handle data management between a host and accelerators in the OpenMP 4.0 and OpenACC standards. Some extensions were added to the directory to allow more flexibility when handling subarrays in the data clauses, including support for unstructured data lifetime. With these modifications, one can use multiple parts of the same array in a nested data environment, keeping the coherence in the accelerator memory between all the subparts. In 2, the paper addresses the solution of large-scale eigenvalue problems that appear in the motion simulation of complex macromolecules on multi-threaded platforms. They compare implementations of three high-performance eigensolvers using out-of-core techniques, enhancing their performance by leveraging hybrid CPU-GPU routines. They show that a Krylov subspace-based eigensolver presents a much lower theoretical cost and outperforms the GPU alternatives for macromolecular simulations. In 3, the authors describe how to conduct high-performance tracking of 3D human motion in real-time using multi-view images and particle swarm optimization. The tracking involves configuring the 3D human model, in the pose described by each particle, and then rasterizing it in each particle's 2D plane. Image acquisition and image processing are multi-threaded and run on CPU in parallel with particle swarm optimization-based searching which is GPU-accelerated, obtaining more precise tracking. In 4, an automated approach to estimate the memory footprint of non-linear data objects is presented. This is a novel method to build a graph-based static data type descriptions that allow to create code for injectable functions that automatically determine the memory footprint of data objects at run-time. This is useful in the context of current programming models for heterogeneous devices with disjoint physical memory spaces which require explicit allocation of device memory and explicit data transfers. This task becomes difficult for non-linear objects, for example, linked lists or multiple inherited classes, due to memory requirements known only at run-time and the composition of complex data structures from basic types. In 5, the authors derive an asymmetric network property on TCP layer for concurrent bidirectional communications on Ethernet clusters and develop a communication model to characterize the communication times accordingly. They show that if the asymmetric network property is excluded from the model, the communication time predictions will be significantly less accurate than those made by using the asymmetric network property. In 6, the authors examine the possibilities of using a GPU for complex 3D finite difference computation. Parallel simulation algorithms using shared and surface memory for relativistic hydrodynamics problems are implemented. Their main objective is to design an efficient algorithm that can benefit from the properties of surface memory optimized for 2D spatial locality and compare it to the best known approach working on shared memory. Their results expose that surface memory is a promising approach for complex 3D finite difference methods. In 7, we are concerned with the challenges underlying the translation of advanced magnetic resonance imaging protocols into a clinical environment. Specifically, rapid online reconstructions require significant computational power. The authors address this problem by developing an external, online, heterogeneous image reconstruction system for magnetic resonance data. The system integrates an external computer equipped with a GPU card into the magnetic resonance scanners image reconstruction pipeline. The system promotes fast online reconstruction for computationally intensive algorithms turning them feasible in a busy clinical service. In 8, the authors compare the performance of various algorithms for the reduction of collective operations in a non-clairvoyant setting, that is, when the algorithms are oblivious to the communication and computation costs. Communication times can rarely be predicted with high accuracy, and may vary significantly over time. The paper assesses how classical static algorithms, where the tree is built before the actual reduction, perform in such settings and quantifies the potential advantage of dynamic algorithms, where the tree is built at run-time and depends on the actual duration of the operations. The study includes both commutative and non-commutative reductions. In 9, a new method for scheduling efficiently parallel applications on hybrid architectures (multi-core machine with GPUs) with m CPUs and k GPUs is presented. Thereby each task of the application can be processed either on a core (CPU) or on a GPU. The objective is to minimize the maximum completion time (makespan). The corresponding scheduling problem is NP-hard, and the authors propose an efficient approximation of a generic methodology. The main idea of the approach is to determine an adequate partition of the set of tasks on the CPUs and the GPUs using a dual approximation scheme. We would like to thank the authors for their excellent contributions to this special issue. The anonymous reviewers who helped to greatly improve the quality of the papers deserve also special acknowledgement; without their selfless effort, this special issue would not have been possible. We hope that the here assembled body of work inspires future research in the area of parallel computing on GPUs and heterogeneous platforms.
Paolo Bientinesi, José R. Herrero 0001, Enrique S. Quintana-Ortí, Robert Strzodka
Concurr. Comput. Pract. Exp.2
2014 Evaluation and assessment of professional skills in the Final Year Project
abstract
In this paper, we present a methodology for Final Year Project (FYP) monitoring and assessment that considers the inclusion of the professional skills required in the particular engineering degree. This proper monitoring and clear evaluation framework provides the student with valuable support for the project implementation as well as for improving the quality of the projects, thereby reducing the academic drop-out rate. The proposed methodology has been implemented at the Barcelona School of Informatics at the Universität Politècnica de Catalunya - BarcelonaTech. The FYP is structured around three milestones: project definition, project monitoring and project completion. Skills are assigned to each milestone according to the tasks required in that phase, and a list of indicators is defined for each phase. The evaluation criteria for each indicator at each phase are specified in a rubric, and are made public both to students and teachers. Thus, the FYP includes an exhaustive evaluation method distributed throughout the whole project implementation, thereby facilitating project organization for the student as well as providing a clear and homogeneous assessment framework. The methodology for the FYP organization, assessment and evaluation was launched and piloted over two semesters. We believe the experience to be general in the sense that it has been conducted as part of an ICT engineering degree, but may easily be extended to any other engineering degree.
Fermín Sánchez, Joan Climent, Julita Corbalán, Pau Fonseca i Casas, Jordi Garcia 0001, José R. Herrero 0001, Xavier Llinas, Horacio Rodríguez, Maria-Ribera Sancho, Marc Alier Forment, Jose Cabré, David López 0001
FIE6
2014 Tuning and hybrid parallelization of a genetic-based multi-point statistics simulation code
Oscar Peredo, Julián M. Ortiz, José R. Herrero 0001, Cristóbal Samaniego
Parallel Comput.3
2013 Graphics processing unit computing and exploitation of hardware accelerators
abstract
SUMMARY This special issue contributes to this promising field with extended and carefully reviewed versions of selected papers from two workshops, namely the 2nd Minisymposium on GPU Computing, which was held as part of the 9th International Conference on Parallel Processing and Applied Mathematics (PPAM 2011) in Torun (Poland); and the Workshop on Exploitation of Hardware Accelerators (WEHA 2011), which was held in conjunction with The 2011 International Conference on High Performance Computing & Simulation in Istanbul (Turkey). Copyright © 2012 John Wiley & Sons, Ltd.
Margarita Amor, Ramón Doallo, Basilio B. Fraguela, José R. Herrero 0001, Enrique S. Quintana-Ortí, Robert Strzodka
Concurr. Comput. Pract. Exp.4
2013 Level-3 Cholesky Factorization Routines Improve Performance of Many Cholesky Algorithms
abstract
Four routines called DPOTF3i, i = a,b,c,d, are presented. DPOTF3i are a novel type of level-3 BLAS for use by BPF ( B locked P acked F ormat) Cholesky factorization and LAPACK routine DPOTRF. Performance of routines DPOTF3i are still increasing when the performance of Level-2 routine DPOTF2 of LAPACK starts decreasing. This is our main result and it implies, due to the use of larger block size nb , that DGEMM, DSYRK, and DTRSM performance also increases! The four DPOTF3i routines use simple register blocking. Different platforms have different numbers of registers. Thus, our four routines have different register blocking sizes. BPF is introduced. LAPACK routines for POTRF and PPTRF using BPF instead of full and packed format are shown to be trivial modifications of LAPACK POTRF source codes. We call these codes BPTRF. There are two variants of BPF: lower and upper. Upper BPF is “identical” to Square Block Packed Format (SBPF). “LAPACK” implementations on multicore processors use SBPF. Lower BPF is less efficient than upper BPF. Vector inplace transposition converts lower BPF to upper BPF very efficiently. Corroborating performance results for DPOTF3i versus DPOTF2 on a variety of common platforms are given for n ≈ nb as well as results for large n comparing DBPTRF versus DPOTRF.
Fred G. Gustavson, Jerzy Wasniewski, Jack J. Dongarra, José R. Herrero 0001, Julien Langou
ACM Trans. Math. Softw.4
2011 Special Issue: GPU computing
abstract
The combined hurdles of power consumption, limited instruction-level parallelism, and memory latency have led hardware manufacturers to design power aware multi-core processors and specialized many-core hardware accelerators in order to further exploit the increasing number of transistors dictated by Moore's Law. The race is open: As of today, Intel's top-of-the-line designs include 8 cores in its Xeon Nehalem architecture, AMD raises this number to 12 cores in the Opteron Magny-Cours processor, and the road-maps of the two companies indicate that these quantities will increase to 10–12 (Intel) and 16 (AMD) cores in 2011. At the same time, specialized hardware architectures, such as graphics processing units (GPUs), with tens of cores are already widely deployed. Core numbers are only one level of parallelism that is on the rise. Each core in current processors contains multiple processing elements enabling parallel processing within it. Power efficiency requires that the parallel processing elements are assembled in SIMD (Single Instruction Multiple Data) units. In the current CPU cores SSE instructions are supported enabling up to 4 parallel multiply-add operations on single precision floating point numbers. Soon this figure will rise to 8 with the AVX instruction set and the Intel Larrabee design featured already 16-wide SIMD units in each core. The number of instructions that can be executed in parallel on each GPU core varies strongly between 16 and 80 because of the different designs of GPU cores by different manufacturers; e.g. NVIDIA Fermi architecture has up to 16 cores times 32 processing elements, whereas AMD Cypress architecture contains up to 20 cores times 80 processing elements, in both cases 2 × increase over the previous generation. However, counting the processing elements on a GPU allows only a comparison within the same GPU family. Across families at least the differing factors of shader clock and arrangement of the processing elements must be taken into account. Although these new multi-core and many-core architectures can potentially deliver a revolutionary boost in raw performance, the efficient utilization of the growing SIMD and many-core parallelism is the key that will determine their success or failure. In this line, the recent advances in the hardware, functionality, and programmability of graphics processors (GPUs) have greatly increased their appeal as add-on co-processors for general-purpose computing. With the involvement of the largest processor manufacturers, NVIDIA, AMD, and Intel, and the strong interest from researchers of various disciplines, this approach has moved from a research niche to a forward-looking technique for heterogeneous parallel computing. Scientific and industry researchers are constantly finding new applications for GPUs in a wide variety of areas, including image and video processing, molecular dynamics, seismic simulation, computational biology and chemistry, fluid dynamics, weather forecast, computational finance, quantum physics, and many others. GPU hardware has evolved over many years from graphics pipelines with many heterogeneous fixed-function components over partially programmable architectures toward a more homogeneous general-purpose design (though some fixed-function hardware has remained because of its efficiency). The general-purpose computing on GPU (GPGPU) revolution started with programmable shaders. NVIDIA Compute Unified Device Architecture (CUDA) and, to a smaller extent, AMD CAL/Brook+ have brought GPUs into the mainstream of computing, developing what has been recently coinedGPU computing* . The great advantage of CUDA is that it defines an abstraction that presents the underlying hardware architecture as a sea of hundreds of fine-grained computational units with synchronization primitives on multiple levels. With OpenCL, there is now also a vendor-independent high-level parallel programming language and an application programming interface that offers the same type of hardware abstraction. GPUs are very versatile accelerators because besides the high hardware parallelism they also feature a high bandwidth connection to dedicated device memory. The latency problem of DRAM is tackled via a sophisticated thread scheduling and switching mechanism on-chip that continues the processing of the next thread as soon as the previous stalls on a data read. These characteristics make GPUs suitable for both compute- and data-intensive parallel processing. All together, the advances in the GPU hardware, the improvements in their programmability, as well as a potentially appealing performance/power ratio, have pushed organizations to invest in heterogeneous systems that include GPUs, and have motivated researchers to port their algorithms to such systems. This special issue contains extended versions of selected papers from the Minisymposium on GPU Computing, which was held as part of the Eight International Conference on Parallel Processing and Applied Mathematics—PPAM 2009 in Wroclaw (Poland). Ten papers were published in the Conference Proceedings, after two review rounds. Extended versions of some of these papers went through a new review process, resulting in the selection of papers contained in this special issue. The topics offer a good cross-section of the current GPU challenges: further abstraction of the hardware and the programming model, improvements of basic parallel algorithms in discrete mathematics and linear algebra, and the utilization of the parallel processing power of GPUs for real-world applications. Michael Repplinger and Philipp Slusallek (‘Stream processing on GPUs using distributed multimedia middleware’, Concurrency and Computation: Practice and Experience [this issue]) introduce an open distributed middleware for the development of applications in multi-GPU systems. In particular, the solution contributed by the authors can seamlessly integrate processing components, hide architecture-specific issues, combine GPUs and CPUs in a heterogeneous computational system, and use local and remote GPUs for distributed processing. Hagens Peters et al. (‘Fast in-place, comparison-based sorting with CUDA: a study with bitonic sort’, Concurrency and Computation: Practice and Experience [this issue]) present their work on a comparison-based in-place implementation of bitonic sort on CUDA-enabled GPUs. They identify and minimize the access to global memory as the main bottleneck and obtain remarkable sorting rates for a large number of sorting elements. Paolo Bientinesi et al. (‘Condensed forms for the symmetric eigenvalue problems on multithreaded architectures’, Concurrency and Computation: Practice and Experience [this issue]) analyze an alternative blocked algorithm for the solution of symmetric eigenvalue problems that can be efficiently cast in terms of efficient matrix–matrix products that attain high performance on a graphics processors. The experimental study of the authors using an accelerated version of this algorithm on NVIDIA GT200 generation of graphics processors demonstrates its superior performance compared with the traditional Level-2 BLAS-based approach on Intel Xeon E5520 (Nehalem) and E7640 (Dunnington) processors. Bernardo Rocha et al. (‘Accelerating cardiac excitation spread simulations using GPUs’, Concurrency and Computation: Practice and Experience [this issue]) employ a graphics processor to significantly accelerate the simulation of electrical activity in the heart. The authors' experiments with the solution of the ordinary differential equations modeling 2D cardiac tissues on a NVIDIA GeForce GT200 show a performance acceleration of 20–180 times with respect to a Quad-core processor. We thank the authors for their excellent contributions to this special issue as well as the anonymous reviewers who helped the authors and the editors of this issue to greatly improve the quality of the papers. We hope that it inspires future research in the area of GPU computing.
José R. Herrero 0001, Enrique S. Quintana-Ortí, Robert Strzodka
Concurr. Comput. Pract. Exp.1
2009 Parallelizing dense and banded linear algebra libraries using SMPSs
abstract
Abstract The promise of future many‐core processors, with hundreds of threads running concurrently, has led the developers of linear algebra libraries to rethink their design in order to extract more parallelism, further exploit data locality, attain better load balance, and pay careful attention to the critical path of computation. In this paper we describe how existing serial libraries such as (C)LAPACK and FLAME can be easily parallelized using the SMPSs tools, consisting of a few OpenMP‐like pragmas and a run‐time system. In the LAPACK case, this usually requires the development of blocked algorithms for simple BLAS‐level operations, which expose concurrency at a finer grain. For better performance, our experimental results indicate that column‐major order, as employed by this library, needs to be abandoned in benefit of a block data layout. This will require a deeper rewrite of LAPACK or, alternatively, a dynamic conversion of the storage pattern at run‐time. The parallelization of FLAME routines using SMPSs is simpler as this library includes blocked algorithms (or algorithms‐by‐blocks in the FLAME argot) for most operations and storage‐by‐blocks (or block data layout) is already in place. Copyright © 2009 John Wiley & Sons, Ltd.
Rosa M. Badia, José R. Herrero 0001, Jesús Labarta, Josep M. Pérez, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí
Concurr. Comput. Pract. Exp.2
2008 Hypermatrix oriented supernode amalgamation
José R. Herrero 0001, Juan J. Navarro
J. Supercomput.1
2007 Exploiting computer resources for fast nearest neighbor classification
José R. Herrero 0001, Juan J. Navarro
Pattern Anal. Appl.1
2006 Compiler-Optimized Kernels: An Efficient Alternative to Hand-Coded Inner Kernels
José R. Herrero 0001, Juan J. Navarro
ICCSA (5)1
2003 Improving Performance of Hypermatrix Cholesky Factorization
José R. Herrero 0001, Juan J. Navarro
Euro-Par1
1996 Data Prefetching and Multilevel Blocking for Linear Algebra Operations
abstract
Much effort has been directed towards obtaining near peak performance for linear algebra operations on current high performance workstations. The large amounts of data accesses however, make performance highly dependent on the behavior of the memory hierarchy. Techniques such as Multilevel Blocking (Tiling), Data Precopying, Software Pipelining and Software Prefetching have been applied in order to improve performance. Nevertheless, to our knowledge, no other work has been done considering the relation between these techniques when applied together. In this paper we analyze the behavior of matrix multiplication algorithms for large matrices on a superscalar and superpipelined processor with a multilevel memory hierarchy when these techniques are applied together. We study and model the performance and limitations of different codes. We also compare two different approaches to data prefetching, binding versus non-binding, and find the latter remarkably more effective than the former due...
Juan J. Navarro, Elena García-Diego, José R. Herrero 0001
International Conference on Supercomputing3