EDBT 2026 Demo / reviewers in the wild / expert
Diego R. Llanos Ferraris
dblp:l/DiegoRLlanosFerraris · also Diego R. Llanos
· DBLP profile ↗
42ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0001-6240-9109ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 37 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On the development of high-performance, multi-GPU applications on heterogeneous systems leveraging SYCLabstractComputational platforms for high-performance scientific applications are increasingly heterogeneous, incorporating multiple GPU accelerators. However, differences in GPU vendors, architectures, and programming models challenge performance portability and ease of development. SYCL provides a unified programming approach, enabling applications to target NVIDIA and AMD GPUs simultaneously while offering higher-level abstractions for data and task management. This paper evaluates SYCL’s performance and development effort using the Finite Time Lyapunov Exponent (FTLE) calculation as a case study. We compare SYCL’s AdaptiveCpp (Ahead-Of-Time and Just-In-Time) and Intel oneAPI compilers, along with different data management strategies (Unified Shared Memory and buffers), against equivalent CUDA and HIP implementations. Our analysis considers single and multi-GPU execution, including heterogeneous setups with GPUs from different vendors. Results show that, while SYCL introduces additional development effort compared to native CUDA and HIP implementations, it enables multi-vendor portability with minimal performance overhead when using specific design options. Based on our findings, we provide development guidelines to help programmers decide when to use SYCL versus vendor-specific alternatives. Francisco J. Andujar, Rocío Carratalá-Sáez, Yuri Torres, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Parallel Distributed Comput. | 5 |
| 2026 | Pipelined FPGA Implementation of a Differential Evolution Engine for Optimization of Scientific ModelsabstractCustom computing machines implemented on FPGAs have emerged as a powerful solution for tackling computationally intensive tasks, leveraging their capacity for deep pipelining and parallel memory access. Differential Evolution (DE), a robust optimization algorithm, combined with adaptive numerical integration methods, is widely used to optimize parameter values in diverse scientific models. These tasks involve extensive floating-point computations, making FPGAs an ideal platform for efficiently accelerating their execution. In this work, we present a flexible and scalable FPGA architecture optimized for DE. This architecture is tailored to solve complex, resource-intensive optimization problems and is easily customizable for various models and integration methods. To demonstrate its efficacy, we evaluate two case studies: The Hodgkin–Huxley model for neuron action potentials and the Circadian clock model of Arabidopsis thaliana . Our architecture integrates adaptive numerical methods with DE and achieves significant performance and energy efficiency gains over CPU and GPU implementations while maintaining versatility across applications. Our architecture’s modular design enables seamless adaptation across different scientific contexts, enabling further optimization of resource utilization and expansion of application domains. The results underline the potential of FPGAs as a superior platform for large-scale scientific computation, offering unmatched energy efficiency and computational throughput for highly demanding tasks. The code developed to carry out this work is publicly available at https://github.com/mdccUVa/de-fpga . Manuel de Castro, Roberto R. Osorio, Yuri Torres, Diego R. Llanos Ferraris |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2025 | Accelerating Scientific Model Optimization with a Pipelined FPGA-Based Differential Evolution EngineabstractCustom computing machines on FPGAs excel in solving computationally intensive tasks by leveraging deep pipelines and parallel memory access. This paper introduces a flexible FPGA-based architecture for Differential Evolution (DE), optimized for a variety of scientific models and numerical integration methods. The architecture's modular design allows seamless customization for diverse applications. Two case studies demonstrate the architecture's capabilities: the Hodgkin-Huxley model for neuron action potentials and the Circadian model of Arabidopsis thaliana. These implementations employ double-precision floating-point arithmetic and adaptive numerical integration techniques, addressing the challenges of complex, stiff differential equations. The proposed design outperforms CPUs and GPUs in computational speed and energy efficiency, achieving up to 3.8x faster processing and significant reductions in energy consumption. This work highlights the potential of FPGA platforms for accelerating complex scientific computations while providing insights for future optimizations in resource utilization and broader applicability. Manuel de Castro, Roberto R. Osorio, Yuri Torres, Diego R. Llanos Ferraris |
FCCM | 4 |
| 2024 | Performance improvement of the triangular matrix product in commodity clustersabstractAbstract There are many works devoted to improving the matrix product computation, as it is used in a wide variety of scientific applications arising from many different fields. In this work, we propose alternative data distribution policies and communication patterns to reduce the elapsed time when computing triangular matrix products in distributed memory environments. In particular, we focus on commodity clusters, where the number of nodes is limited, proposing alternatives to traditional approaches in order to improve this operation’s performance. Our proposal overcomes the performance results associated with the state-of-the-art libraries, such as ScaLAPACK and SLATE, offering execution times that are up to 30% faster. Inmaculada Santamaria-Valenzuela, Rocío Carratalá-Sáez, Yuri Torres, Diego R. Llanos Ferraris, Arturo González-Escribano |
J. Supercomput. | 4 |
| 2023 | Supporting efficient overlapping of host-device operations for heterogeneous programming with CtrlEventsabstractHeterogeneous systems with several kinds of devices, such as multi-core CPUs, GPUs, FPGAs, among others, are now commonplace. Exploiting all these devices with device-oriented programming models, such as CUDA or OpenCL, requires expertise and knowledge about the underlying hardware to tailor the application to each specific device, thus degrading performance portability. Higher-level proposals simplify the programming of these devices, but their current implementations do not have an efficient support to solve problems that include frequent bursts of computation and communication, or input/output operations. In this work we present CtrlEvents, a new heterogeneous runtime solution which automatically overlaps computation and communication whenever possible, simplifying and improving the efficiency of data-dependency analysis and the coordination of both device computations and host tasks that include generic I/O operations. Our solution outperforms other state-of-the-art implementations for most situations, presenting a good balance between portability, programmability and efficiency. Yuri Torres, Francisco J. Andujar, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Parallel Distributed Comput. | 4 |
| 2023 | UVaFTLE: Lagrangian finite time Lyapunov exponent extraction for fluid dynamic applicationsabstractAbstract The determination of Lagrangian Coherent Structures (LCS) is becoming very important in several disciplines, including cardiovascular engineering, aerodynamics, and geophysical fluid dynamics. From the computational point of view, the extraction of LCS consists of two main steps: The flowmap computation and the resolution of Finite Time Lyapunov Exponents (FTLE). In this work, we focus on the design, implementation, and parallelization of the FTLE resolution. We offer an in-depth analysis of this procedure, as well as an open source C implementation (UVaFTLE) parallelized using OpenMP directives to attain a fair parallel efficiency in shared-memory environments. We have also implemented CUDA kernels that allow UVaFTLE to leverage as many NVIDIA GPU devices as desired in order to reach the best parallel efficiency. For the sake of reproducibility and in order to contribute to open science, our code is publicly available through GitHub. Moreover, we also provide Docker containers to ease its usage. Rocío Carratalá-Sáez, Yuri Torres, José Sierra-Pallares, Sergio López-Huguet, Diego R. Llanos Ferraris |
J. Supercomput. | 5 |
| 2023 | Implementation of a motion estimation algorithm for Intel FPGAs using OpenCLabstractMotion Estimation is one of the main tasks behind any video encoder. It is a computationally costly task; therefore, it is usually delegated to specific or reconfigurable hardware, such as FPGAs. Over the years, multiple FPGA implementations have been developed, mainly using hardware description languages such as Verilog or VHDL. Since programming using hardware description languages is a complex task, it is desirable to use higher-level languages to develop FPGA applications.The aim of this work is to evaluate OpenCL, in terms of expressiveness, as a tool for developing this kind of FPGA applications. To do so, we present and evaluate a parallel implementation of the Block Matching Motion Estimation process using OpenCL for Intel FPGAs, usable and tested on an Intel Stratix 10 FPGA. The implementation efficiently processes Full HD frames completely inside the FPGA. In this work, we show the resource utilization when synthesizing the code on an Intel Stratix 10 FPGA, as well as a performance comparison with multiple CPU implementations with varying levels of optimization and vectorization capabilities. We also compare the proposed OpenCL implementation, in terms of resource utilization and performance, with estimations obtained from an equivalent VHDL implementation. Manuel de Castro, Roberto R. Osorio, David López Vilariño, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Supercomput. | 5 |
| 2023 | EPSILOD: efficient parallel skeleton for generic iterative stencil computations in distributed GPUsabstractAbstract Iterative stencil computations are widely used in numerical simulations. They present a high degree of parallelism, high locality and mostly-coalesced memory access patterns. Therefore, GPUs are good candidates to speed up their computation. However, the development of stencil programs that can work with huge grids in distributed systems with multiple GPUs is not straightforward, since it requires solving problems related to the partition of the grid across nodes and devices, and the synchronization and data movement across remote GPUs. In this work, we present EPSILOD, a high-productivity parallel programming skeleton for iterative stencil computations on distributed multi-GPUs, of the same or different vendors that supports any type of n-dimensional geometric stencils of any order. It uses an abstract specification of the stencil pattern (neighbors and weights) to internally derive the data partition, synchronizations and communications. Computation is split to better overlap with communications. This paper describes the underlying architecture of EPSILOD, its main components, and presents an experimental evaluation to show the benefits of our approach, including a comparison with another state-of-the-art solution. The experimental results show that EPSILOD is faster and shows good strong and weak scalability for platforms with both homogeneous and heterogeneous types of GPU. Manuel de Castro, Inmaculada Santamaria-Valenzuela, Yuri Torres, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Supercomput. | 5 |
| 2021 | Distributed programming of a hyperspectral image registration algorithm for heterogeneous GPU clusters
Jorge Fernández-Fabeiro, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Parallel Distributed Comput. | 3 |
| 2019 | High-level parallel programming in a heterogeneous worldabstractDuring the last decade, parallel programming has evolved in an unprecedent way. Fifteen years ago, the future of parallel computing seemed to consist on the advent of multicore processors composed by an ever‐increasing number in the core count per CPU, and their interconnection to form larger clusters. Programming models, such as OpenMP1 that allows to transform sequential C and Fortran codes into parallel versions with low effort, without requiring to explicitly handle threads or to share memory among them, seemed to be the winner choice in a world where computers would include more and more cores inside a single chip. Message‐passing paradigms, such as MPI,2 provided a solution to interconnect these computers in larger facilities. José Daniel García, Diego R. Llanos Ferraris |
Concurr. Comput. Pract. Exp. | 2 |
| 2019 | Computational and mathematical models meet heterogeneous computing
Diego R. Llanos Ferraris, Jesús Vigo-Aguiar |
J. Supercomput. | 1 |
| 2019 | Toward a BLAS library truly portable across different accelerator types
Eduardo Rodriguez-Gutiez, Ana Moreton-Fernandez, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Supercomput. | 4 |
| 2017 | Supporting the Xeon Phi Coprocessor in a Heterogeneous Programming Model
Ana Moreton-Fernandez, Eduardo Rodriguez-Gutiez, Arturo González-Escribano, Diego R. Llanos Ferraris |
Euro-Par | 4 |
| 2017 | TORMENT OpenACC2016: A Benchmarking Tool for OpenACC CompilersabstractOpenACC is a parallel programming model for hardware accelerators, such as GPUs or Xeon Phi, which has been in development for several years by now. During this time, different compilers have appeared, both commercial and open source, which are still on development stage. Due to the fact that both the OpenACC standard and its implementations are relatively recent, we propose a benchmark suite specifically designed to check the performance of the OpenACC features in the code generated by different compilers on different architectures. Our benchmark suite is named TORMENT OpenACC2016. Along with this tool we have developed an adequate metric for the comparison of performance among different machine-compiler pairs which we have named TORMENT ACC2016 Score. The version 1 of TORMENT OpenACC2016 presented in this paper, contains six benchmarks, and is available online. Daniel Barba, Arturo González-Escribano, Diego R. Llanos Ferraris |
PDP | 3 |
| 2017 | A technique to automatically determine Ad-hoc communication patterns at runtime
Ana Moreton-Fernandez, Arturo González-Escribano, Diego R. Llanos Ferraris |
Parallel Comput. | 3 |
| 2017 | BFCA+: automatic synthesis of parallel code with TLS capabilities
Sergio Aldea, Diego R. Llanos Ferraris, Arturo González-Escribano |
J. Supercomput. | 2 |
| 2016 | An OpenMP Extension that Supports Thread-Level SpeculationabstractOpenMP directives are the de-facto standard for shared-memory parallel programming. However, OpenMP does not guarantee the correctness of the parallel execution of a given loop if runtime data dependences arise. Consequently, many highly-parallel regions cannot be safely parallelized with OpenMP due to the possibility of a dependence violation. In this paper, we propose to augment OpenMP capabilities, by adding thread-level speculation (TLS) support. Our contribution is threefold. First, we have defined a new speculative clause for variables inside parallel loops. This clause ensures that all accesses to these variables will be carried out according to sequential semantics. Second, we have created a new, software-based TLS runtime library to ensure correctness in the parallel execution of OpenMP loops that include speculative variables. Third, we have developed a new GCC plugin, which seamlessly translates our OpenMP speculative clause into calls to our TLS runtime engine. The result is the ATLaS C Compiler framework, which takes advantage of TLS techniques to expand OpenMP functionalities, and guarantees the sequential semantics of any parallelized loop. Sergio Aldea, Alvaro Estebanez, Diego R. Llanos Ferraris, Arturo González-Escribano |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2015 | Moody Scheduling for Speculative Parallelization
Alvaro Estebanez, Diego R. Llanos Ferraris, David Orden, Belén Palop |
Euro-Par | 2 |
| 2014 | A New GCC Plugin-Based Compiler Pass to Add Support for Thread-Level Speculation into OpenMP
Sergio Aldea, Alvaro Estebanez, Diego R. Llanos Ferraris, Arturo González-Escribano |
Euro-Par | 3 |
| 2014 | Squashing Alternatives for Software-Based Speculative ParallelizationabstractSpeculative parallelization is a runtime technique that optimistically executes sequential code in parallel, checking that no dependence violations arise. In the case of a dependence violation, all mechanisms proposed so far either switch to sequential execution, or conservatively stop and restart the offending thread and all its successors, potentially discarding work that does not depend on this particular violation. In this work we systematically explore the design space of solutions for this problem, proposing a new mechanism that reduces the number of threads that should be restarted when a data dependence violation is found. Our new solution, called exclusive squashing, keeps track of inter-thread dependencies at runtime, selectively stopping and restarting offending threads, together with all threads that have consumed data from them. We have compared this new approach with existent solutions on a real system, executing different applications with loops that are not analyzable at compile time and present as much as 10% of inter-thread dependence violations at runtime. Our experimental results show a relative performance improvement of up to 14%, together with a reduction of one-third of the numbers of squashed threads. The speculative parallelization scheme and benchmarks described in this paper are available under request. Álvaro García-Yágüez, Diego R. Llanos Ferraris, Arturo González-Escribano |
IEEE Trans. Computers | 2 |
| 2014 | The BonaFide C Analyzer: automatic loop-level characterization and coverage measurement
Sergio Aldea, Diego R. Llanos Ferraris, Arturo González-Escribano |
J. Supercomput. | 2 |
| 2014 | Optimizing an APSP implementation for NVIDIA GPUs using kernel characterization criteria
Hector Ortega-Arranz, Yuri Torres, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Supercomput. | 4 |
| 2014 | Blending Extensibility and Performance in Dense and Sparse Parallel Data ManagementabstractDealing with both dense and sparse data in parallel environments usually leads to two different approaches: To rely on a monolithic, hard-to-modify parallel library, or to code all data management details by hand. In this paper we propose a third approach, that delivers good performance while the underlying library structure remains modular and extensible. Our solution integrates dense and sparse data management using a common interface, that also decouples data representation, partitioning, and layout from the algorithmic and parallel strategy decisions of the programmer. Our experimental results in different parallel environments show that this new approach combines the flexibility obtained when the programmer handles all the details with a performance comparable to the use of a state-of-the-art, sparse matrix parallel library. Javier Fresno, Arturo González-Escribano, Diego R. Llanos Ferraris |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2014 | An Extensible System for Multilevel Automatic Data Partition and MappingabstractAutomatic data distribution is a key feature to obtain efficient implementations from abstract and portable parallel codes. We present a highly efficient and extensible runtime library that integrates techniques for automatic data partition and mapping. It uses a novel approach to define an abstract interface and a plug-in system to encapsulate different types of regular and irregular techniques, helping to generate codes which are independent of the exact mapping functions selected. Currently, it supports hierarchical tiling of arrays with dense and stride domains, that allows the implementation of both data and task parallelism using a SPMD model. It automatically computes appropriate domain partitions for a selected virtual topology, mapping them to available processors with static or dynamic load-balancing techniques. Our library also allows the construction of reusable communication patterns that efficiently exploit MPI communication capabilities. The use of our library greatly reduces the complexity of data distribution and communication, hiding the details of the underlying architecture. The library can be used as an abstract layer for building generic tiling operations as well. Our experimental results show that the use of this library allows to achieve similar performance as carefully-implemented manual versions for several, well-known parallel kernels and benchmarks in distributed and multicore systems, and substantially reduces programming effort. Arturo González-Escribano, Yuri Torres, Javier Fresno, Diego R. Llanos Ferraris |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2013 | Extending a hierarchical tiling arrays library to support sparse data partitioning
Javier Fresno, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Supercomput. | 3 |
| 2013 | uBench: exposing the impact of CUDA block geometry in terms of performance
Yuri Torres, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Supercomput. | 3 |
| 2012 | Encapsulated Synchronization and Load-Balance in Heterogeneous Programming
Yuri Torres, Arturo González-Escribano, Diego R. Llanos Ferraris |
Euro-Par | 3 |
| 2012 | Using Fermi Architecture Knowledge to Speed up CUDA and OpenCL ProgramsabstractThe NVIDIA graphics processing units (GPUs) are playing an important role as general purpose programming devices. The implementation of parallel codes to exploit the GPU hardware architecture is a task for experienced programmers. The threadblock size and shape choice is one of the most important user decisions when a parallel problem is coded. The threadblock configuration has a significant impact on the global performance of the program. While in CUDA parallel programming model it is always necessary to specify the threadblock size and shape, the OpenCL standard also offers an automatic mechanism to take this delicate decision. In this paper we present a study of these criteria for Fermi architecture, introducing a general approach for threadblock choice, and showing that there is considerable room for improvement in OpenCL automatic strategy. Yuri Torres, Arturo González-Escribano, Diego R. Llanos Ferraris |
ISPA | 3 |
| 2012 | Using SPEC CPU2006 to evaluate the sequential and parallel code generated by commercial and open-source compilers
Sergio Aldea, Diego R. Llanos Ferraris, Arturo González-Escribano |
J. Supercomput. | 2 |
| 2011 | Robust thread-level speculationabstractRobustness is a key issue on any runtime system that aims to speed up the execution of a program. However, robustness considerations are commonly overlooked when new software-based, thread-level speculation (STLS) systems are proposed. This paper highlights the relevance of the problem, showing different situations when the use of incorrect data can irreversibly alter the speculative execution of an algorithm, despite the efforts of a given STLS system to maintain sequential consistency. We show that the management of speculative exceptions is a common factor to these problems. Based on this fact, we propose a novel solution to handle speculative exceptions. Our solution eagerly tries to solve the issue before the non-speculative thread arrives to the instruction that rose the exception. We compare our solution to a more conservative approach found in the bibliography. The comparison is done both qualitatively, through a detailed analysis of the tradeoffs involved, and quantitatively, evaluating the effects of both solutions in the execution of three different benchmarks on a real system. Both studies conclude that our solution handles the occurrence of speculative exceptions more efficiently. Under heavy loads intended to push to its limits a STLS system, our solution leads to execution times reduced by up to 52.02% with respect to earlier proposals. Our solution does not affect the performance when speculative exceptions do not appear. We believe that our proposal makes STLS systems robust enough to be used in production environments. Álvaro García-Yágüez, Diego R. Llanos Ferraris, Arturo González-Escribano |
HiPC | 2 |
| 2011 | Exclusive squashing for thread-level speculationabstractSpeculative parallelization is a runtime technique that optimistically executes sequential code in parallel, checking that no dependence violations appear. In this paper, we address the problem of minimizing the number of threads that should be restarted when a data dependence violation is found. We present a new mechanism that keeps track of inter-thread dependencies in order to selectively stop and restart offending threads, and all threads that have consumed data from them. Results show a reduction of 38.5% to 81.8% in the number of restarted threads for real application loops and up to a 10% speedup, depending on the amount of local computation. Álvaro García-Yágüez, Diego R. Llanos Ferraris, Arturo González-Escribano |
HPDC | 2 |
| 2011 | Towards a Compiler Framework for Thread-Level SpeculationabstractSpeculative parallelization techniques allow to extract parallelism of fragments of code that can not be analyzed at compile time. However, research on software-based, thread-level speculation will greatly benefit from an appropriate compiler framework for easy prototyping and further development of new techniques. This paper presents an experimental XML-based compilation framework to handle speculative parallelization of C code. The framework extends Cetus, a source-to-source C compiler, to build an XML tree based on the Cetus Internal Representation of the source code. Other modules of our framework rely on XPath and XSLT capabilities to process the XML tree generated, to perform analysis on the use of variables and to augment the original code for software-based, speculative parallel execution. The use of the current version of our framework allows a fast prototyping of new analysis and transformation solutions, with a reduction of around 83% on the number of code lines needed with respect to the direct use of Cetus for the same purpose. To show the possibilities of this framework, we present an automatically-generated classification of loops for several SPEC CPU2006 C benchmarks. This classification is useful to better understand the potential benefits derived from the use of speculative parallelization techniques. The development framework presented here is freely available under request. Sergio Aldea, Diego R. Llanos Ferraris, Arturo González-Escribano |
PDP | 2 |
| 2011 | Automatic Data Partitioning Applied to Multigrid PDE SolversabstractThis paper studies the impact of using automatic data-layout techniques on the process of coding the well-known multigrid MG NAS parallel benchmark. We describe the sequential problem in detail, and discuss the parallel version and its optimizations. Then, we implement the parallel algorithm using Hit map, a highly-efficient modular library for hierarchical tiling and mapping of arrays. We describe how to use the library plug-in system to add a new data-layout module that encapsulates a generalization of the data-alignment policy of the MG benchmark. The module system applies this policy to automatically adapt the data distribution and communication code to any grain level. The impact of using these techniques is qualitatively and quantitatively described in terms of development effort and performance. Our results show that it is possible to introduce flexible automatic data-layout techniques in current parallel compiler technology, without sacrificing performance. Javier Fresno, Arturo González-Escribano, Diego R. Llanos Ferraris |
PDP | 3 |
| 2011 | Trasgo: a nested-parallel programming system
Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Supercomput. | 2 |
| 2010 | Effortless and Efficient Distributed Data-Partitioning in Linear AlgebraabstractThis paper introduces a new technique to exploit compositions of different data-layout techniques with Hit map, a library for hierarchical-tiling and automatic mapping of arrays. We show how Hit map is used to implement block-cyclic layouts for a parallel LU decomposition algorithm. The paper compares the well-known ScaLAPACK implementation of LU, as well as other carefully optimized MPI versions, with a Hit map implementation. The comparison is made in terms of both performance and code length. Our results show that the Hit map version outperforms the ScaLAPACK implementation and is almost as efficient as our best manual MPI implementation. The insertion of this composition technique in the automatic data-layouts of Hit map allows the programmer to develop parallel programs with both a significant reduction of the development effort and a negligible loss of efficiency. Carlos de Blas Carton, Arturo González-Escribano, Diego R. Llanos Ferraris |
HPCC | 3 |
| 2008 | Just-In-Time Scheduling for Loop-based Speculative ParallelizationabstractScheduling for speculative parallelization is a problem that remained unsolved despite its importance. Simple methods such as Fixed-Size Chunking (FSC) need several 'dry-runs' before an acceptable chunk size is found. Other traditional scheduling methods were originally designed for loops with no dependences, so they are primarily focused in the problem of load balancing. In general, all these methods perform poorly when used for speculative parallelization, where loops may present unexpected dependences that adversely affect performance. In this work we address the problem of scheduling loops with and without dependences for speculative execution. We have found that a trade-off between minimizing the number of re-executions and reducing overheads can be found if the size of the scheduled block of iterations is calculated at runtime. We introduce here a scheduling method called Just-In- Time (JIT) scheduling that uses the information available during the execution of the loop in order to dynamically compute the size of the next block to be scheduled. The results show a 10% to 26% speedup improvement in real applications with dependences with respect to a carefully- tuned FSC strategy, and a 9% to 39% speedup improvement in real applications without dependences. With our proposal, the number of dependence violations that lead to squashes can be reduced by up to 62%. Moreover, in applications where the cost of dependence violations is too high to obtain speedups with FSC, our runtime scheduling mechanism avoids performance degradation. Diego R. Llanos Ferraris, David Orden, Belén Palop |
PDP | 1 |
| 2007 | New Scheduling Strategies for Randomized Incremental Algorithms in the Context of Speculative ParallelizationabstractIn this work, we address the problem of scheduling loops with dependences in the context of speculative parallelization. We show that the scheduling alternatives are highly influenced by the dependence violation pattern the code presents. We center our analysis in those algorithms where dependences are less likely to appear as the execution proceeds. Particularly, we focus on randomized incremental algorithms, widely used as a much more efficient solution to many problems than their deterministic counterparts. These important algorithms are, in general, hard to parallelize by hand and represent a challenge for any automatic parallelization scheme. Our analysis led us to the development of MESETA, a new scheduling strategy that takes into account the probability of a dependence violation to determine the number of iterations being scheduled. MESETA is compared with existing techniques, including fixed-size chunking (FSC), the only scheduling alternative used so far in the context of speculative parallelization. Our experimental results show a 5.5 percent to 36.25 percent speedup improvement over FSC, leading to a better extraction of the parallelism inherent to randomized incremental algorithms. Moreover, when the cost of dependence violations is too high to obtain speedups, MESETA curves the performance degradation Diego R. Llanos Ferraris, David Orden, Belén Palop |
IEEE Trans. Computers | 1 |
| 2006 | TPCC-UVa: an open-source TPC-C implementation for parallel and distributed systemsabstractThis paper presents TPCC-UVa, an open-source implementation of the TPC-C benchmark intended to be used in parallel and distributed systems. TPCC-UVa is written entirely in C language and it uses the Post-greSQL database engine. This implementation includes all the functionalities described by the TPC-C standard specification for the measurement of both uni- and multiprocessor systems performance. The major characteristics of the TPC-C specification are discussed, together with a description of the TPCC-UVa implementation and architecture and real examples of performance measurements Diego R. Llanos Ferraris, Belén Palop |
IPDPS | 1 |
| 2005 | Design Space Exploration of a Software Speculative Parallelization SchemeabstractWith speculative parallelization, code sections that cannot be fully analyzed by the compiler are optimistically executed in parallel. Hardware schemes are fast but expensive and require modifications to the processors and/or memory system. Software schemes require no changes to the hardware of existing shared-memory systems, but can suffer from significant overheads involved with the speculative execution. In fact, the performance of software schemes is highly dependent on application characteristics, the design and implementation of the scheme, and the system configuration and size. This paper explores the design space of a recently proposed software speculative parallelization scheme. In the process, we gain insight into the most beneficial features of software schemes for speculative parallelization, as well as the most influential application characteristics. For instance, experimental results show that, contrary to intuition, checking for data dependence violations on every speculative store, as opposed to at commit time, leads to little performance degradation in the worst case and to significantly better performance with large configurations. Also, scheduling policies based on windows can perform very close to fully dynamic policies with a fraction of the memory overhead. Finally, experimental results show consistent speedups in the execution of loops that cannot be parallelized at compile time, both with and without RAW data dependences, for 4 to 32 processors. Marcelo H. Cintra, Diego R. Llanos Ferraris |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2004 | Speculative Parallelization of a Randomized Incremental Convex Hull Algorithm
Marcelo H. Cintra, Diego R. Llanos Ferraris, Belén Palop |
ICCSA (3) | 2 |
| 2003 | Toward efficient and robust software speculative parallelization on multiprocessorsabstractWith speculative parallelization, code sections that cannot be fully analyzed by the compiler are aggressively executed in parallel. Hardware schemes are fast but expensive and require modifications to the processors and memory system. Software schemes require no extra hardware but can be inefficient.This paper proposes a new software-only speculative parallelization scheme. The scheme is developed after a systematic evaluation of the design options available and is shown to be efficient and robust and to outperform previously proposed schemes. The novelty and performance advantage of the scheme stem from the use of carefully tuned data structures, synchronization policies, and scheduling mechanisms. Experimental results show that our scheme has small overheads and, for applications with few or no data dependence violations, realizes on average 71% of the speedup of a manually parallelized version of the code, outperforming two recently proposed software-only speculative parallelization schemes. For applications with many data dependence violations, our performance monitors and switches can effectively curb the performance degradation. Marcelo H. Cintra, Diego R. Llanos Ferraris |
PPoPP | 2 |
| 2000 | Reducing the Replacement Overhead on COMA Protocols for Workstation-Based Architectures
Diego R. Llanos Ferraris, Benjamín Sahelices Fernández, Agustín De Dios Hernández |
Euro-Par | 1 |