EDBT 2026 Demo / reviewers in the wild / expert
Yuri Torres
dblp:117/6986 · also Yuri Torres De La Sierra
· DBLP profile ↗
14ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-3037-3567ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 4 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On the development of high-performance, multi-GPU applications on heterogeneous systems leveraging SYCLabstractComputational platforms for high-performance scientific applications are increasingly heterogeneous, incorporating multiple GPU accelerators. However, differences in GPU vendors, architectures, and programming models challenge performance portability and ease of development. SYCL provides a unified programming approach, enabling applications to target NVIDIA and AMD GPUs simultaneously while offering higher-level abstractions for data and task management. This paper evaluates SYCL’s performance and development effort using the Finite Time Lyapunov Exponent (FTLE) calculation as a case study. We compare SYCL’s AdaptiveCpp (Ahead-Of-Time and Just-In-Time) and Intel oneAPI compilers, along with different data management strategies (Unified Shared Memory and buffers), against equivalent CUDA and HIP implementations. Our analysis considers single and multi-GPU execution, including heterogeneous setups with GPUs from different vendors. Results show that, while SYCL introduces additional development effort compared to native CUDA and HIP implementations, it enables multi-vendor portability with minimal performance overhead when using specific design options. Based on our findings, we provide development guidelines to help programmers decide when to use SYCL versus vendor-specific alternatives. Francisco J. Andujar, Rocío Carratalá-Sáez, Yuri Torres, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Parallel Distributed Comput. | 3 |
| 2026 | Pipelined FPGA Implementation of a Differential Evolution Engine for Optimization of Scientific ModelsabstractCustom computing machines implemented on FPGAs have emerged as a powerful solution for tackling computationally intensive tasks, leveraging their capacity for deep pipelining and parallel memory access. Differential Evolution (DE), a robust optimization algorithm, combined with adaptive numerical integration methods, is widely used to optimize parameter values in diverse scientific models. These tasks involve extensive floating-point computations, making FPGAs an ideal platform for efficiently accelerating their execution. In this work, we present a flexible and scalable FPGA architecture optimized for DE. This architecture is tailored to solve complex, resource-intensive optimization problems and is easily customizable for various models and integration methods. To demonstrate its efficacy, we evaluate two case studies: The Hodgkin–Huxley model for neuron action potentials and the Circadian clock model of Arabidopsis thaliana . Our architecture integrates adaptive numerical methods with DE and achieves significant performance and energy efficiency gains over CPU and GPU implementations while maintaining versatility across applications. Our architecture’s modular design enables seamless adaptation across different scientific contexts, enabling further optimization of resource utilization and expansion of application domains. The results underline the potential of FPGAs as a superior platform for large-scale scientific computation, offering unmatched energy efficiency and computational throughput for highly demanding tasks. The code developed to carry out this work is publicly available at https://github.com/mdccUVa/de-fpga . Manuel de Castro, Roberto R. Osorio, Yuri Torres, Diego R. Llanos Ferraris |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2025 | Accelerating Scientific Model Optimization with a Pipelined FPGA-Based Differential Evolution EngineabstractCustom computing machines on FPGAs excel in solving computationally intensive tasks by leveraging deep pipelines and parallel memory access. This paper introduces a flexible FPGA-based architecture for Differential Evolution (DE), optimized for a variety of scientific models and numerical integration methods. The architecture's modular design allows seamless customization for diverse applications. Two case studies demonstrate the architecture's capabilities: the Hodgkin-Huxley model for neuron action potentials and the Circadian model of Arabidopsis thaliana. These implementations employ double-precision floating-point arithmetic and adaptive numerical integration techniques, addressing the challenges of complex, stiff differential equations. The proposed design outperforms CPUs and GPUs in computational speed and energy efficiency, achieving up to 3.8x faster processing and significant reductions in energy consumption. This work highlights the potential of FPGA platforms for accelerating complex scientific computations while providing insights for future optimizations in resource utilization and broader applicability. Manuel de Castro, Roberto R. Osorio, Yuri Torres, Diego R. Llanos Ferraris |
FCCM | 3 |
| 2024 | Performance improvement of the triangular matrix product in commodity clustersabstractAbstract There are many works devoted to improving the matrix product computation, as it is used in a wide variety of scientific applications arising from many different fields. In this work, we propose alternative data distribution policies and communication patterns to reduce the elapsed time when computing triangular matrix products in distributed memory environments. In particular, we focus on commodity clusters, where the number of nodes is limited, proposing alternatives to traditional approaches in order to improve this operation’s performance. Our proposal overcomes the performance results associated with the state-of-the-art libraries, such as ScaLAPACK and SLATE, offering execution times that are up to 30% faster. Inmaculada Santamaria-Valenzuela, Rocío Carratalá-Sáez, Yuri Torres, Diego R. Llanos Ferraris, Arturo González-Escribano |
J. Supercomput. | 3 |
| 2023 | Task-based preemptive scheduling on FPGAs leveraging partial reconfigurationabstractSummary Field‐programmable gate arrays (FPGAs) are an attractive type of accelerator for all‐purpose high performance computing computing systems due to the possibility of deploying tailored hardware on demand. However, the common tools for programming and operating FPGAs are still complex to use, especially in scenarios where diverse types of tasks should be dynamically executed. In this work, we present a programming abstraction with a simple interface that internally leverages high‐level synthesis, dynamic partial reconfiguration and synchronization mechanisms to use an FPGA as a multi‐tasking server with preemptive scheduling and priority queues. This leads to an improved use of the FPGA resources, allowing the execution of several different kernels concurrently and deploying the most urgent ones as fast as possible. The results of our experimental study show that our approach incurs only a 10 5% overhead in the worst case when using two reconfigurable regions, whilst providing a significant performance improvement of at least 24 21% over the traditional full reconfiguration approach. Gabriel Rodriguez-Canal, Nick Brown 0002, Yuri Torres, Arturo González-Escribano |
Concurr. Comput. Pract. Exp. | 3 |
| 2023 | Supporting efficient overlapping of host-device operations for heterogeneous programming with CtrlEventsabstractHeterogeneous systems with several kinds of devices, such as multi-core CPUs, GPUs, FPGAs, among others, are now commonplace. Exploiting all these devices with device-oriented programming models, such as CUDA or OpenCL, requires expertise and knowledge about the underlying hardware to tailor the application to each specific device, thus degrading performance portability. Higher-level proposals simplify the programming of these devices, but their current implementations do not have an efficient support to solve problems that include frequent bursts of computation and communication, or input/output operations. In this work we present CtrlEvents, a new heterogeneous runtime solution which automatically overlaps computation and communication whenever possible, simplifying and improving the efficiency of data-dependency analysis and the coordination of both device computations and host tasks that include generic I/O operations. Our solution outperforms other state-of-the-art implementations for most situations, presenting a good balance between portability, programmability and efficiency. Yuri Torres, Francisco J. Andujar, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Parallel Distributed Comput. | 1 |
| 2023 | UVaFTLE: Lagrangian finite time Lyapunov exponent extraction for fluid dynamic applicationsabstractAbstract The determination of Lagrangian Coherent Structures (LCS) is becoming very important in several disciplines, including cardiovascular engineering, aerodynamics, and geophysical fluid dynamics. From the computational point of view, the extraction of LCS consists of two main steps: The flowmap computation and the resolution of Finite Time Lyapunov Exponents (FTLE). In this work, we focus on the design, implementation, and parallelization of the FTLE resolution. We offer an in-depth analysis of this procedure, as well as an open source C implementation (UVaFTLE) parallelized using OpenMP directives to attain a fair parallel efficiency in shared-memory environments. We have also implemented CUDA kernels that allow UVaFTLE to leverage as many NVIDIA GPU devices as desired in order to reach the best parallel efficiency. For the sake of reproducibility and in order to contribute to open science, our code is publicly available through GitHub. Moreover, we also provide Docker containers to ease its usage. Rocío Carratalá-Sáez, Yuri Torres, José Sierra-Pallares, Sergio López-Huguet, Diego R. Llanos Ferraris |
J. Supercomput. | 2 |
| 2023 | EPSILOD: efficient parallel skeleton for generic iterative stencil computations in distributed GPUsabstractAbstract Iterative stencil computations are widely used in numerical simulations. They present a high degree of parallelism, high locality and mostly-coalesced memory access patterns. Therefore, GPUs are good candidates to speed up their computation. However, the development of stencil programs that can work with huge grids in distributed systems with multiple GPUs is not straightforward, since it requires solving problems related to the partition of the grid across nodes and devices, and the synchronization and data movement across remote GPUs. In this work, we present EPSILOD, a high-productivity parallel programming skeleton for iterative stencil computations on distributed multi-GPUs, of the same or different vendors that supports any type of n-dimensional geometric stencils of any order. It uses an abstract specification of the stencil pattern (neighbors and weights) to internally derive the data partition, synchronizations and communications. Computation is split to better overlap with communications. This paper describes the underlying architecture of EPSILOD, its main components, and presents an experimental evaluation to show the benefits of our approach, including a comparison with another state-of-the-art solution. The experimental results show that EPSILOD is faster and shows good strong and weak scalability for platforms with both homogeneous and heterogeneous types of GPU. Manuel de Castro, Inmaculada Santamaria-Valenzuela, Yuri Torres, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Supercomput. | 3 |
| 2021 | Efficient heterogeneous programming with FPGAs using the Controller model
Gabriel Rodriguez-Canal, Yuri Torres, Francisco J. Andujar, Arturo González-Escribano |
J. Supercomput. | 2 |
| 2014 | Optimizing an APSP implementation for NVIDIA GPUs using kernel characterization criteria
Hector Ortega-Arranz, Yuri Torres, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Supercomput. | 2 |
| 2014 | An Extensible System for Multilevel Automatic Data Partition and MappingabstractAutomatic data distribution is a key feature to obtain efficient implementations from abstract and portable parallel codes. We present a highly efficient and extensible runtime library that integrates techniques for automatic data partition and mapping. It uses a novel approach to define an abstract interface and a plug-in system to encapsulate different types of regular and irregular techniques, helping to generate codes which are independent of the exact mapping functions selected. Currently, it supports hierarchical tiling of arrays with dense and stride domains, that allows the implementation of both data and task parallelism using a SPMD model. It automatically computes appropriate domain partitions for a selected virtual topology, mapping them to available processors with static or dynamic load-balancing techniques. Our library also allows the construction of reusable communication patterns that efficiently exploit MPI communication capabilities. The use of our library greatly reduces the complexity of data distribution and communication, hiding the details of the underlying architecture. The library can be used as an abstract layer for building generic tiling operations as well. Our experimental results show that the use of this library allows to achieve similar performance as carefully-implemented manual versions for several, well-known parallel kernels and benchmarks in distributed and multicore systems, and substantially reduces programming effort. Arturo González-Escribano, Yuri Torres, Javier Fresno, Diego R. Llanos Ferraris |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2013 | uBench: exposing the impact of CUDA block geometry in terms of performance
Yuri Torres, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Supercomput. | 1 |
| 2012 | Encapsulated Synchronization and Load-Balance in Heterogeneous Programming
Yuri Torres, Arturo González-Escribano, Diego R. Llanos Ferraris |
Euro-Par | 1 |
| 2012 | Using Fermi Architecture Knowledge to Speed up CUDA and OpenCL ProgramsabstractThe NVIDIA graphics processing units (GPUs) are playing an important role as general purpose programming devices. The implementation of parallel codes to exploit the GPU hardware architecture is a task for experienced programmers. The threadblock size and shape choice is one of the most important user decisions when a parallel problem is coded. The threadblock configuration has a significant impact on the global performance of the program. While in CUDA parallel programming model it is always necessary to specify the threadblock size and shape, the OpenCL standard also offers an automatic mechanism to take this delicate decision. In this paper we present a study of these criteria for Fermi architecture, introducing a general approach for threadblock choice, and showing that there is considerable room for improvement in OpenCL automatic strategy. Yuri Torres, Arturo González-Escribano, Diego R. Llanos Ferraris |
ISPA | 1 |