VLDB 2026 Research / reviewers in the wild / expert
Pedro Valero-Lara
dblp:07/9764
· DBLP profile ↗
31ranked-venue papers
20as first author
10since 2021 · last 2026
0000-0002-1479-4310ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 12 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accl++ : A high-productivity programming language for performance and code portability on heterogeneous systemsabstractThis work describes the Accl++ programming language for heterogeneous computing. Accl++ is embedded in the C++ language and implemented as a C++ library. The language allows for describing device code and execution libraries and includes primitives for runtime compilation (RTC), thereby enabling code portability across disparate devices. Here, we demonstrate Accl++’s capability by coding different benchmarks from different domains. This work also describes an analysis of the overheads introduced by the Accl++ RTC support as well as the performance of the generated code for the Accl++ kernels. Accl++ improves heterogeneous code portability without incurring high levels of overhead. In two different heterogeneous systems—one composed of two 32-core AMD EPYC 7513 CPUs and two NVIDIA A100 GPUs and the other composed of two 12-core AMD EPYC 7272 CPUs and two AMD MI100 GPUs—Accl++ enables the execution of the same binary application with observed RTC overheads in the range of 3%–10% of the total kernel execution time, resulting in performance levels similar to those of native CUDA/HIP/OpenCL code. Marc González 0001, Pedro Valero-Lara, Mohammad Alaul Haque Monil, Seyong Lee, Beau Johnston, Aaron R. Young, Narasinga Rao Miniskar, Keita Teranishi, Jeffrey S. Vetter |
Future Gener. Comput. Syst. | 2 |
| 2026 | High-performance computing heterogeneous systems and subsystems
Sergio Iserte, Pedro Valero-Lara, Kevin A. Brown |
Future Gener. Comput. Syst. | 2 |
| 2025 | Enabling Scientific Applications with Performance-Portability and High-Productivity for Multi-GPU Programming with JACC.MultiabstractThis work bridges the gap between multi-GPU computing and high-productivity, performance-portable programming solutions. Our goal is to enhance scientific applications with a productive and portable solution—program once, deploy everywhere—for multi-GPU programming with no cost to programmability. To accomplish this, we implemented JACC.Multi, which is part of the Julia for ACCelerators (JACC) performance-portable framework. JACC. Multi is the only high-level, portable metaprogramming solution that targets multi-GPU environments and is integrated in a readily accessible programming language (e.g., Julia language). With transparent GPU-to-GPU communication, JACC. Multi is optimized for scientific application workloads and is portable for NVIDIA and AMD accelerators. For the evaluation, we use two modern multi-GPU systems: Hudson, which features two NVIDIA H100 Hopper GPUs per node, and Frontier, which features four AMD MI250X GPUs per node, each with two Graphics Compute Dies (GCDs) for a total of eight GCDs per node. Additionally, as part of the evaluation, we use JACC (one GPU), MPI+JACC, and JACC. Multi codes that implement well-known and widely used scientific algorithms/kernels such as the conjugate gradient algorithm and an explicit forward Euler solver that requires GPU-to-GPU communication. Overall, JACC. Multi codes achieve better performance than MPI+JACC codes and significant speedups over JACC (one GPU), with up to 1.9× on Hudson and 6× on Frontier. Pedro Valero-Lara, William F. Godoy, Philip W. Fackler, Keita Teranishi, Jeffrey S. Vetter |
eScience | 1 |
| 2025 | LLM-Driven Fortran-to-C/C++ Portability for Parallel Scientific CodesabstractWe define the fundamental practices and criteria for evaluating and using the Meta Llama 3 and OpenAI ChatGPT 3.5 and 4o large language models (LLMs) to translate parallel scientific Fortran + OpenMP and Fortran + OpenACC codes to C/C++ codes that can leverage vendor-specific libraries (CUDA, HIP) for GPU acceleration in addition to other performance-portable programming models (e.g., Kokkos, OpenMP, OpenACC). In this study, LLMs are used to translate 11 different parallel Fortran codes with some of the most popular and widely used kernels/proxies in high-performance computing (HPC): AXPY, GEMV, GEMM, Jacobi, SpMV, and the >200-line Hartree-Fock application proxy, which implements a solver for quantum many-body systems. In all, we analyze the correctness and reproducibility of more than 1,650 AI-generated parallel C/C++ codes. Additionally, we evaluate the performance of Fortran codes and AI-generated C/C++ codes on two modern HPC architectures—one AMD EPYC Rome CPU with 64 cores and one NVIDIA Ampere A100 GPU. We use multi-modal prompting and fine-tuning techniques for LLMs to produce parallel scientific C/C++ codes with high levels of correctness (more than 95% of the codes are well ported) and speedups of up to an order of magnitude versus Fortran + OpenMP and Fortran + OpenACC codes on the same system. Pedro Valero-Lara, William F. Godoy, Jose Gonzalez, Alexis Huante, Hallyma Gauthier-Chaparro, Jhonny Gonzalez, Yuguo Kelly Tang, Keita Teranishi, Jeffrey S. Vetter |
eScience | 1 |
| 2025 | ChatHPC: Building the Foundations for a Productive and Trustworthy AI-Assisted HPC EcosystemabstractChatHPC democratizes large language models for the high-performance computing (HPC) community by providing the infrastructure, ecosystem, and knowledge needed to apply modern generative AI technologies to rapidly create specific capabilities for critical HPC components while using relatively modest computational resources. Our divide-and-conquer approach focuses on creating a collection of reliable, highly specialized, and optimized AI assistants for HPC based on the cost-effective and fast Code Llama fine-tuning processes and expert supervision. We target major components of the HPC software stack, including programming models, runtimes, I/O, tooling, and math libraries. Thanks to AI, ChatHPC provides a more productive HPC ecosystem by boosting important tasks related to portability, parallelization, optimization, scalability, and instrumentation, among others. With relatively small datasets (on the order of KB), the AI assistants, which are created in a few minutes by using one node with two NVIDIA H100 GPUs and the ChatHPC library, can create new capabilities with Meta’s 7-billion parameter Code Llama base model to produce high-quality software with a level of trustworthiness of up to 90% higher than the 1.8-trillion parameter OpenAI ChatGPT-4o model for critical programming tasks in the HPC software stack. Pedro Valero-Lara, Aaron R. Young, Jeffrey S. Vetter, Zheming Jin, Swaroop Pophale, Mohammad Alaul Haque Monil, Keita Teranishi, William F. Godoy |
SC | 1 |
| 2024 | sKokkos: Enabling Kokkos with Transparent Device Selection on Heterogeneous Systems using OpenACCabstractThis paper presents a new feature to enable Kokkos with transparent device selection. For application developers, it is not easy to identify which device is the most appropriate to use in a heterogeneous system, since this depends on the characteristics of both the application and the hardware. In Kokkos, a backend is associated with one specific programming model/hardware. Programmers decide which backend to use at compilation time. This new feature implemented on the OpenACC backend eliminates the burden of deciding which device to use, providing a highly productive programming solution for Kokkos applications. This work includes implementation details and a performance study conducted with a set of mini-benchmarks (i.e., AXPY and dot product), kernels (Lattice-Bolzmann method), and two mini-apps (LULESH and miniFE) on two heterogeneous systems with different hardware capabilities. This new Kokkos feature provides high accelerations of up to 35 × thanks to automatic and transparent device selection. Pedro Valero-Lara, Seyong Lee, Joel E. Denny, Keita Teranishi, Jeffrey S. Vetter, Marc González 0001 |
HPC Asia | 1 |
| 2024 | Large language model evaluation for high-performance computing software developmentabstractAbstract We apply AI‐assisted large language model (LLM) capabilities of GPT‐3 targeting high‐performance computing (HPC) kernels for (i) code generation, and (ii) auto‐parallelization of serial code in C ++ , Fortran, Python and Julia. Our scope includes the following fundamental numerical kernels: AXPY, GEMV, GEMM, SpMV, Jacobi Stencil, and CG, and language/programming models: (1) C ++ (e.g., OpenMP [including offload], OpenACC, Kokkos, SyCL, CUDA, and HIP), (2) Fortran (e.g., OpenMP [including offload] and OpenACC), (3) Python (e.g., numpy, Numba, cuPy, and pyCUDA), and (4) Julia (e.g., Threads, CUDA.jl, AMDGPU.jl, and KernelAbstractions.jl). Kernel implementations are generated using GitHub Copilot capabilities powered by the GPT‐based OpenAI Codex available in Visual Studio Code given simple + + prompt variants. To quantify and compare the generated results, we propose a proficiency metric around the initial 10 suggestions given for each prompt. For auto‐parallelization, we use ChatGPT interactively giving simple prompts as in a dialogue with another human including simple “prompt engineering” follow ups. Results suggest that correct outputs for C ++ correlate with the adoption and maturity of programming models. For example, OpenMP and CUDA score really high, whereas HIP is still lacking. We found that prompts from either a targeted language such as Fortran or the more general‐purpose Python can benefit from adding language keywords, while Julia prompts perform acceptably well for its Threads and CUDA.jl programming models. We expect to provide an initial quantifiable point of reference for code generation in each programming model using a state‐of‐the‐art LLM. Overall, understanding the convergence of LLMs, AI, and HPC is crucial due to its rapidly evolving nature and how it is redefining human‐computer interactions. William F. Godoy, Pedro Valero-Lara, Keita Teranishi, Prasanna Balaprakash, Jeffrey S. Vetter |
Concurr. Comput. Pract. Exp. | 2 |
| 2022 | IRIS-BLAS: Towards a Performance Portable and Heterogeneous BLAS LibraryabstractThis paper presents IRIS-BLAS, a novel heterogeneous and performance portable BLAS library. IRIS-BLAS is built on top of the IRIS runtime and multiple vendor and open-source BLAS libraries. It can transparently use all the architectures/devices available in a heterogeneous system, using the appropriate BLAS library based on the task mapping at run time. Thus, IRIS-BLAS is portable across a broad spectrum of architectures and BLAS libraries, alleviating the worry of application developers about modifying the application source code. Even though the emphasis is on portability, IRIS-BLAS provides competitive or even better performance than other state-of-the-art references. Moreover, IRIS-BLAS offers new features such as efficiently using extremely heterogeneous systems composed of multiple GPUs from different hardware vendors. Narasinga Rao Miniskar, Mohammad Alaul Haque Monil, Pedro Valero-Lara, Frank Liu 0001, Jeffrey S. Vetter |
HIPC | 3 |
| 2022 | Propagation Pattern for Moment Representation of the Lattice Boltzmann MethodabstractA propagation pattern for the moment representation of the regularized lattice Boltzmann method (LBM) in three dimensions is presented. Using effectively lossless compression, the simulation state is stored as a set of moments of the lattice Boltzmann distribution function, instead of the distribution function itself. An efficient cache-aware propagation pattern for this moment representation has the effect of substantially reducing both the storage and memory bandwidth required for LBM simulations. This paper extends recent work with the moment representation by expanding the performance analysis on central processing unit (CPU) architectures, considering how boundary conditions are implemented, and demonstrating the effectiveness of the moment representation on a graphics processing unit (GPU) architecture. John Gounley, Madhurima Vardhan, Erik W. Draeger, Pedro Valero-Lara, Shirley V. Moore, Amanda Randles |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2021 | Static Graphs for Coding Productivity in OpenACCabstractThe main contribution of this work is to increase the coding productivity for GPU programming by using the concept of Static Graphs. To do so, we have combined the new CUDA Graph API with the OpenACC programming model. We use as test cases a well-known and widely used problems in HPC and AI: the Particle Swarm Optimization. We complement the OpenACC functionality with the use of CUDA Graph, achieving accelerations of more than one order of magnitude, and a performance very close to a reference and optimized CUDA code. Finally, we propose a new specification to incorporate the concept of Static Graphs into the OpenACC specification. Leonel Toledo, Pedro Valero-Lara, Jeffrey S. Vetter, Antonio J. Peña |
HiPC | 2 |
| 2020 | sLASs: A fully automatic auto-tuned linear algebra library based on OpenMP extensions implemented in OmpSs (LASs Library)
Pedro Valero-Lara, Sandra Catalán, Xavier Martorell, Tetsuzo Usui, Jesús Labarta |
J. Parallel Distributed Comput. | 1 |
| 2019 | Accelerating Conjugate Gradient using OmpSsabstractIn this paper, we present the benefits of using the clause concurrent of OmpSs when performing reductions, more specifically, when applied to the dot product (DOT) operations. We analyze its benefits through the implementation of different versions of the Conjugate Gradient (CG) method. We start from a parallel version of the code based on tasks and dependencies; later, we introduce the use of the concurrent clause, which allows to overlap the execution of tasks that have data dependencies among them. In this way, we want to show the benefits of the concurrent clause, which might be included in OpenMP standard as previously done with other OmpSs features. Our tests, performed on a single node of the (Intel-based) Marenostrum 4 Supercomputer and a single socket of the (ARM-based) Dibona cluster, show that the use of the concurrent clause may improve performance with respect to the version where only tasks and dependencies are used around 37% and 23% respectively. Sandra Catalán, Xavier Martorell, Jesús Labarta, Tetsuzo Usui, Leonel Toledo, Pedro Valero-Lara |
PDCAT | 6 |
| 2019 | Tasking in Accelerators: Performance EvaluationabstractIn this work, we analyze the implications and results of implementing dynamic parallelism, concurrent kernels and CUDA Graphs to solve task-oriented problems. As a benchmark we propose three different methods for solving DGEMM operation on tiled-matrices; which might be the most popular benchmark for performance analysis. For the algorithms that we study, we present significant differences in terms of data dependencies, synchronization and granularity. The main contribution of this work is determining which of the previous approaches work better for having multiple task running concurrently in a single GPU, as well as stating the main limitations and benefits of every technique. Using dynamic parallelism and CUDA Streams we were able to achieve up to 30% speedups and for CUDA Graph API up to 25x acceleration outperforming state of the art results. Leonel Toledo, Antonio J. Peña, Sandra Catalán, Pedro Valero-Lara |
PDCAT | 4 |
| 2019 | BLAS-3 Optimized by OmpSs Regions (LASs Library)abstractIn this paper we propose a set of optimizations for the BLAS-3 routines of LASs library (Linear Algebra routines on OmpSs) and perform a detailed analysis of the impact of the proposed changes in terms of performance and execution time. OmpSs allows to use regions in the dependences of the tasks. This helps not only in the programming of the algorithmic optimizations, but also in the reduction of the execution time achieved by such optimizations. Different strategies are implemented in order to reduce the amount of tasks created (when there is enough parallelism) during the execution of BLAS-3 operations in the original LASs. Also a better IPC is obtained thanks to a better memory hierarchy exploitation. More specifically, we increase the performance, in particular on big matrices, about 12% for TRSM, and 17% for GEMM with respect to the original version of LASs, even using less cores in the case of GEMM/SYMM. Moreover, when LASs is compared to the OpenMP reference dense linear algebra library PLASMA, performance is increased up to 12.5% for GEMM/SYMM, while for TRSM/TRMM this value raises to 15%. Pedro Valero-Lara, Sandra Catalán, Xavier Martorell, Jesús Labarta |
PDP | 1 |
| 2019 | MPI+OpenMP tasking scalability for multi-morphology simulations of the human brain
Pedro Valero-Lara, Raül Sirvent, Antonio J. Peña, Jesús Labarta |
Parallel Comput. | 1 |
| 2018 | Variable Batched DGEMMabstractMany scientific applications are in need to solve a high number of small-size independent problems. These individual problems do not provide enough parallelism and then, these must be computed as a batch. Today, vendors such as Intel and NVIDIA are developing their own suite of batch routines. Although most of the works focus on computing batches of fixed size, in real applications we can not assume a uniform size for all set of problems. We explore and analyze different strategies based on parallel for, task and taskloop OpenMP pragmas. Although these strategies are straightforward from a programmer's point of view, they have a different impact on performance. We also analyze a new prototype provided by Intel (MKL), which deals with batch operations (cblas_dgemm_batch). We propose a new approach called grouping. It basically groups a set of problems until filling a limit in terms of memory occupancy or number of operations. In this way, groups composed by different number of problems are distributed on cores, achieving a more balanced distribution in terms of computational cost. This strategy is able to be up to 6× faster than the Intel (MKL) batch routine. Pedro Valero-Lara, Ivan Martínez-Pérez, Sergi Mateo, Raül Sirvent, Vicenç Beltran 0001, Xavier Martorell, Jesús Labarta |
PDP | 1 |
| 2018 | MPI+OpenMP Tasking Scalability for the Simulation of the Human Brain: Human Brain ProjectabstractThe simulation of the behavior of the Human Brain is one of the most ambitious challenges today with a non-end of important applications. We can find many different initiatives in the USA, Europe and Japan which attempt to achieve such a challenging target. In this work we focus on the most important European initiative (Human Brain Project) and on one of the tools (Arbor). This tool simulates the spikes triggered in a neuronal network by computing the voltage capacitance on the neurons' morphology, being one of the most precise simulators today. In the present work, we have evaluated the use of MPI+OpenMP tasking on top of the Arbor simulator. In this paper, we present the main characteristics of the Arbor tool and how these can be efficiently managed by using MPI+OpenMP tasking. We prove that this approach is able to achieve a good scaling even when computing a relatively low workload (number of neurons) per node using up to 32 nodes. Our target consists of achieving not only a highly scalable implementation based on MPI, but also to develop a tool with a high degree of abstraction without losing control and performance by using MPI+OpenMP tasking. Pedro Valero-Lara, Raül Sirvent, Antonio J. Peña, Xavier Martorell, Jesús Labarta |
EuroMPI | 1 |
| 2018 | cuThomasBatch and cuThomasVBatch, CUDA Routines to compute batch of tridiagonal systems on NVIDIA GPUsabstractSummary The solving of tridiagonal systems is one of the most computationally expensive parts in many applications, so that multiple studies have explored the use of NVIDIA GPUs to accelerate such computation. However, these studies have mainly focused on using parallel algorithms to compute such systems, which can efficiently exploit the shared memory and are able to saturate the GPUs capacity with a low number of systems, presenting a poor scalability when dealing with a relatively high number of systems. The gtsvStridedBatch routine in the cuSPARSE NVIDIA package is one of these examples, which is used as reference in this article. We propose a new implementation (cuThomasBatch) based on the Thomas algorithm. Unlike other algorithms, the Thomas algorithm is sequential, and so a coarse‐grained approach is implemented where one CUDA thread solves a complete tridiagonal system instead of one CUDA block as in gtsvStridedBatch. To achieve a good scalability using this approach, it is necessary to carry out a transformation in the way that the inputs are stored in memory to exploit coalescence (contiguous threads access to contiguous memory locations). Different variants regarding the transformation of the data are explored in detail. We also explore some variants for the case of variable batch, when the size of the systems of the batch has different size (cuThomasVBatch). The results given in this study prove that the implementations carried out in this work are able to beat the reference code, being up to 5× (in double precision) and 6× (in single precision) faster using the latest NVIDIA GPU architecture, the Pascal P100. Pedro Valero-Lara, Ivan Martínez-Pérez, Raül Sirvent, Xavier Martorell, Antonio J. Peña |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | Reducing memory requirements for large size LBM simulations on GPUsabstractSummary The scientific community in its never‐ending road of larger and more efficient computational resources is in need of more efficient implementations that can adapt efficiently on the current parallel platforms. Graphics processing units are an appropriate platform that cover some of these demands. This architecture presents a high performance with a reduced cost and an efficient power consumption. However, the memory capacity in these devices is reduced and so expensive memory transfers are necessary to deal with big problems. Today, the lattice‐Boltzmann method (LBM) has positioned as an efficient approach for Computational Fluid Dynamics simulations. Despite this method is particularly amenable to be efficiently parallelized, it is in need of a considerable memory capacity, which is the consequence of a dramatic fall in performance when dealing with large simulations. In this work, we propose some initiatives to minimize such demand of memory, which allows us to execute bigger simulations on the same platform without additional memory transfers, keeping a high performance. In particular, we present 2 new implementations, LBM‐Ghost and LBM‐Swap, which are deeply analyzed, presenting the pros and cons of each of them. Pedro Valero-Lara |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | Heterogeneous CPU+GPU approaches for mesh refinement over Lattice-Boltzmann simulationsabstractSummary The use of mesh refinement in CFD is an efficient and widely used methodology to minimize the computational cost by solving those regions of high geometrical complexity with a finer grid. In this work, the author focuses on studying two methods, one based on Multi‐Domain and one based on Irregular meshing, to deal with mesh refinement over LBM simulations. The numerical formulation is presented in detail. It is proposed two approaches, homogeneous GPU and heterogeneous CPU+GPU, on each of the refinement methods. Obviously, the use of the two architectures, CPU and GPU, to compute the same problem involves more important challenges with respect to the homogeneous counterpart. These challenges and the strategies to deal with them are described in detail into the present work. We pay a particular attention to the differences among both methodologies/implementations in terms of programmability, memory management, and performance. The size of the refined sub‐domain has important consequences over both methodologies; however, the influence on Multi‐Domain approach is much higher. For instance, when dealing with a big refined sub‐domain, the Multi‐Domain approach achieves an important fall in performance with respect to other cases, where the size of the refined sub‐domain is smaller. Otherwise, using the Irregular approach, there is no such a dramatic fall in performance when increasing the size of the refined sub‐domain. Copyright © 2016 John Wiley & Sons, Ltd. Pedro Valero-Lara, Johan Jansson |
Concurr. Comput. Pract. Exp. | 1 |
| 2016 | Leveraging the Performance of LBM-HPC for Large Sizes on GPUs Using Ghost Cells
Pedro Valero-Lara |
ICA3PP | 1 |
| 2015 | LBM-HPC - An Open-Source Tool for Fluid Simulations. Case Study: Unified Parallel C (UPC-PGAS)abstractThe main motivation of this work is the evaluation of the Unified Parallel C (UPC) model, for Boltzmann-fluid simulations. UPC is one of the current models in the so-called Partitioned Global Address Space paradigm. This paradigm attempts to increase the simplicity of codes and achieve a better efficiency and scalability. Two different UPC-based implementations, explicit and implicit, are presented and evaluated. We compare the fundamental features of our UPC implementations with other parallel programming model, MPI-OpenMP. In particular each of the major steps of any LBM code, i.e., Boundary Conditions, Communication, and LBM solver, are analyzed. Pedro Valero-Lara, Johan Jansson |
CLUSTER | 1 |
| 2014 | Multi-GPU acceleration of DARTEL (early detection of Alzheimer)abstractMedical image processing is becoming a significant discipline within the bioinformatic community. In particular, deformable registration methods are one of the most sophisticate and important lines of research within biomedical image processing, due to the valuable information provided. However, these methods consume considerable processing time, power consumption and require high amounts of memory. Current Graphics Processing Units (GPU) have a high number of cores and high memory bandwidth, providing an excellent platform for reducing the cost of these methods in terms of processing time and power consumption. This work proposes several Graphics Processing Units GPU-based implementations of one of the most sophisticated deformable registration algorithms, DARTEL. The main contribution consists of a new GPU approach, which considerably reduces the overhead caused by memory transfers and the computational cost required by the parallelization of DARTEL. Furthermore, the use of multiple (2 and 4) GPUs is studied, achieving favorable results. This new approach provides a high speedup with respect to the sequential counterpart. Finally, the experimental results show a processing time reduction of more than 3 hours in typical cases of study. Additionally, this new approach significantly reduces power consumption. Pedro Valero-Lara |
CLUSTER | 1 |
| 2014 | hLCS. A Hybrid GPGPU Approach for Solving Multiple Short and Unbalanced LCS Problems
Pedro Valero-Lara |
ICCSA (6) | 1 |
| 2014 | Accelerating solid-fluid interaction based on the immersed boundary method on multicore and GPU architectures
Pedro Valero-Lara |
J. Supercomput. | 1 |
| 2013 | GPU Powered ROSA AnalyzerabstractIn this work we present the first version of ROSAA, Rosa Analyzer, using a GPU architecture. ROSA is a Markovian Process Algebra able to capture pure non-determinism, probabilities and timed actions, Over it, a tool has been developed for getting closer to a fully automatic process of analyzing the behaviour of a system specified as a process of ROSA, so that, ROSAA is able to automatically generate the part of the Labeled Transition System (occasionally the whole one), LTS in the sequel, in which we could be interested, but, since this is a very computationally expensive task, a GPU powered version of ROSAA which includes parallel processing capabilities, has been created to better deal with such generating process. As the conventional GPU processing loads are mainly focused on data parallelization over quite similar types of data, this work means a quite novel use of these kind of architectures, moreover the authors do not know any other formal model tool running over GPUs. ROSAA running starts with the Syntactic analysis so generating a layered structure suitable to, afterwards, apply the Operational Semantics transition rules in the easiest way. Since from each specification/state more than one rule could be applied, this is the key point at which GPU should provide its benefits, i.e., allowing to generate all the new states reachable in a single-semantics-step from a given one, at the same time through a simultaneous launching of a set of threads over the GPU platform. Although this establishes a step forward to the practical usefulness of such type of tools, the state-explosion problem arises indeed, so we are aware that reducing the size of the LTS will be sooner or later required, in this line the authors are working on an heuristics to properly prune an enough number of branches of the LTS, so making the task of generating it, more tractable. Raúl Pardo, Fernando López Pelayo, Pedro Valero-Lara |
ICPP | 3 |
| 2013 | A GPU approach for accelerating 3D deformable registration (DARTEL) on brain biomedical imagesabstractMedical image processing is becoming a significant discipline in bioinformatic. Particularly, deformable registration methods are one of the field most important in the biomedical image processing, due to the valuable information provided. However, these methods consume a considerable processing time and memory requirements. Current GPUs have a high number of cores and high memory bandwidth providing an excellent platform for reducing the cost of these methods in terms of processing time. In this work, it is proposed a Graphics Processing Units (GPU)-based implementation of one of the most sophisticated deformable registration algorithms, DARTEL. The experimental results show a processing time reduction higher than 2 hours in typical cases of study. Moreover, the power consumption is also reduced in a significant amount. Pedro Valero-Lara |
EuroMPI | 1 |
| 2012 | Improving the Performance for the Range Search on Metric Spaces Using a Multi-GPU Platform
Roberto Uribe, Enrique Arias-Antúnez, José L. Sánchez 0002, Diego Cazorla, Pedro Valero-Lara |
DEXA (2) | 5 |
| 2012 | Block Tridiagonal Solvers on Heterogeneous ArchitecturesabstractModern multi-core and many-core systems offer a very impressive cost/performance ratio. In this paper a set of new parallel implementations for the solution of linear systems with block-tridiagonal coefficient matrix on current parallel architectures is proposed and evaluated: one of them on multi-core, others on many-core and finally, a new heterogeneous implementation on both architectures. The results show a speedup higher than 6 on certain parts of the problem, being the heterogeneous implementation the fastest. Pedro Valero-Lara, Alfredo Pinelli, Julien Favier, Manuel Prieto 0001 |
ISPA | 1 |
| 2011 | A GPU-Based Implementation for Range Queries on Spaghettis Data Structure
Roberto Uribe, Pedro Valero-Lara, Enrique Arias-Antúnez, José L. Sánchez 0002, Diego Cazorla |
ICCSA (1) | 2 |
| 2011 | A GPU-based implementation of the MRF algorithm in ITK package
Pedro Valero-Lara, José L. Sánchez 0002, Diego Cazorla, Enrique Arias-Antúnez |
J. Supercomput. | 1 |