VLDB 2026 Research / reviewers in the wild / expert
Arturo González-Escribano
dblp:35/7039
· DBLP profile ↗
37ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0003-1309-9321ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 4 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On the development of high-performance, multi-GPU applications on heterogeneous systems leveraging SYCLabstractComputational platforms for high-performance scientific applications are increasingly heterogeneous, incorporating multiple GPU accelerators. However, differences in GPU vendors, architectures, and programming models challenge performance portability and ease of development. SYCL provides a unified programming approach, enabling applications to target NVIDIA and AMD GPUs simultaneously while offering higher-level abstractions for data and task management. This paper evaluates SYCL’s performance and development effort using the Finite Time Lyapunov Exponent (FTLE) calculation as a case study. We compare SYCL’s AdaptiveCpp (Ahead-Of-Time and Just-In-Time) and Intel oneAPI compilers, along with different data management strategies (Unified Shared Memory and buffers), against equivalent CUDA and HIP implementations. Our analysis considers single and multi-GPU execution, including heterogeneous setups with GPUs from different vendors. Results show that, while SYCL introduces additional development effort compared to native CUDA and HIP implementations, it enables multi-vendor portability with minimal performance overhead when using specific design options. Based on our findings, we provide development guidelines to help programmers decide when to use SYCL versus vendor-specific alternatives. Francisco J. Andujar, Rocío Carratalá-Sáez, Yuri Torres, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Parallel Distributed Comput. | 4 |
| 2025 | Programming and Mapping for Mixed Heterogeneous Devices: The Case of Optical FlowabstractCurrent parallel systems are increasingly heterogeneous, mixing devices of different types and computing capabilities. Exploiting multiple different devices for the same application continues to be a challenge that ranges from technical problems related to synchronizing and communicating diverse devices to problems of load distribution and flexibility to adjust the computation to the platform resources. In this work, we study the problem of using and extending a heterogeneous portability layer to program and adapt HSOpticalFlow to heterogeneous platforms. HSOpticalFlow is a streaming application to estimate the apparent movement of objects in a sequence of images. It is a simple but characteristic example of the structure of applications based on multilevel ILS (Iterative Loop Stencil), also known as multi-grid methods, applied to a sequence of inputs. Starting from the original CUDA reference code, we present a methodology and programming techniques based on the Controller programming model to implement it as a pipeline among multiple devices. We discuss a technique to determine a proper work partition and mapping for a set of devices. This allows for building very efficient parallel solutions, using similar devices or taking advantage of devices with lower computing power, to reduce the load and increase the productivity of more powerful ones. We present the results of an experimental study using several GPUs of different vendors, architectures, and generations, showing that this solution allows combinations of devices to be efficiently exploited to improve performance. Specifically, the results include speedups of 1.91x using two NVIDIA A100 GPUs and 1.21x using one NVIDIA V100 GPU and one AMD WX9100 GPU, which is about $3 x$ slower than the NVIDIA GPU for this application. Sergio Alonso Pascual, Arturo González-Escribano |
PDP | 2 |
| 2024 | Performance improvement of the triangular matrix product in commodity clustersabstractAbstract There are many works devoted to improving the matrix product computation, as it is used in a wide variety of scientific applications arising from many different fields. In this work, we propose alternative data distribution policies and communication patterns to reduce the elapsed time when computing triangular matrix products in distributed memory environments. In particular, we focus on commodity clusters, where the number of nodes is limited, proposing alternatives to traditional approaches in order to improve this operation’s performance. Our proposal overcomes the performance results associated with the state-of-the-art libraries, such as ScaLAPACK and SLATE, offering execution times that are up to 30% faster. Inmaculada Santamaria-Valenzuela, Rocío Carratalá-Sáez, Yuri Torres, Diego R. Llanos Ferraris, Arturo González-Escribano |
J. Supercomput. | 5 |
| 2023 | Task-based preemptive scheduling on FPGAs leveraging partial reconfigurationabstractSummary Field‐programmable gate arrays (FPGAs) are an attractive type of accelerator for all‐purpose high performance computing computing systems due to the possibility of deploying tailored hardware on demand. However, the common tools for programming and operating FPGAs are still complex to use, especially in scenarios where diverse types of tasks should be dynamically executed. In this work, we present a programming abstraction with a simple interface that internally leverages high‐level synthesis, dynamic partial reconfiguration and synchronization mechanisms to use an FPGA as a multi‐tasking server with preemptive scheduling and priority queues. This leads to an improved use of the FPGA resources, allowing the execution of several different kernels concurrently and deploying the most urgent ones as fast as possible. The results of our experimental study show that our approach incurs only a 10 5% overhead in the worst case when using two reconfigurable regions, whilst providing a significant performance improvement of at least 24 21% over the traditional full reconfiguration approach. Gabriel Rodriguez-Canal, Nick Brown 0002, Yuri Torres, Arturo González-Escribano |
Concurr. Comput. Pract. Exp. | 4 |
| 2023 | Supporting efficient overlapping of host-device operations for heterogeneous programming with CtrlEventsabstractHeterogeneous systems with several kinds of devices, such as multi-core CPUs, GPUs, FPGAs, among others, are now commonplace. Exploiting all these devices with device-oriented programming models, such as CUDA or OpenCL, requires expertise and knowledge about the underlying hardware to tailor the application to each specific device, thus degrading performance portability. Higher-level proposals simplify the programming of these devices, but their current implementations do not have an efficient support to solve problems that include frequent bursts of computation and communication, or input/output operations. In this work we present CtrlEvents, a new heterogeneous runtime solution which automatically overlaps computation and communication whenever possible, simplifying and improving the efficiency of data-dependency analysis and the coordination of both device computations and host tasks that include generic I/O operations. Our solution outperforms other state-of-the-art implementations for most situations, presenting a good balance between portability, programmability and efficiency. Yuri Torres, Francisco J. Andujar, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Parallel Distributed Comput. | 3 |
| 2023 | Implementation of a motion estimation algorithm for Intel FPGAs using OpenCLabstractMotion Estimation is one of the main tasks behind any video encoder. It is a computationally costly task; therefore, it is usually delegated to specific or reconfigurable hardware, such as FPGAs. Over the years, multiple FPGA implementations have been developed, mainly using hardware description languages such as Verilog or VHDL. Since programming using hardware description languages is a complex task, it is desirable to use higher-level languages to develop FPGA applications.The aim of this work is to evaluate OpenCL, in terms of expressiveness, as a tool for developing this kind of FPGA applications. To do so, we present and evaluate a parallel implementation of the Block Matching Motion Estimation process using OpenCL for Intel FPGAs, usable and tested on an Intel Stratix 10 FPGA. The implementation efficiently processes Full HD frames completely inside the FPGA. In this work, we show the resource utilization when synthesizing the code on an Intel Stratix 10 FPGA, as well as a performance comparison with multiple CPU implementations with varying levels of optimization and vectorization capabilities. We also compare the proposed OpenCL implementation, in terms of resource utilization and performance, with estimations obtained from an equivalent VHDL implementation. Manuel de Castro, Roberto R. Osorio, David López Vilariño, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Supercomput. | 4 |
| 2023 | EPSILOD: efficient parallel skeleton for generic iterative stencil computations in distributed GPUsabstractAbstract Iterative stencil computations are widely used in numerical simulations. They present a high degree of parallelism, high locality and mostly-coalesced memory access patterns. Therefore, GPUs are good candidates to speed up their computation. However, the development of stencil programs that can work with huge grids in distributed systems with multiple GPUs is not straightforward, since it requires solving problems related to the partition of the grid across nodes and devices, and the synchronization and data movement across remote GPUs. In this work, we present EPSILOD, a high-productivity parallel programming skeleton for iterative stencil computations on distributed multi-GPUs, of the same or different vendors that supports any type of n-dimensional geometric stencils of any order. It uses an abstract specification of the stencil pattern (neighbors and weights) to internally derive the data partition, synchronizations and communications. Computation is split to better overlap with communications. This paper describes the underlying architecture of EPSILOD, its main components, and presents an experimental evaluation to show the benefits of our approach, including a comparison with another state-of-the-art solution. The experimental results show that EPSILOD is faster and shows good strong and weak scalability for platforms with both homogeneous and heterogeneous types of GPU. Manuel de Castro, Inmaculada Santamaria-Valenzuela, Yuri Torres, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Supercomput. | 4 |
| 2021 | Carrot and Stick approaches revisited when managing Technical Debt in an educational contextabstractTechnical Debt management is an important aspect in the training of Software Engineering students. In this paper we study the effect of two assessment strategies in an educational context: One based on penalisation, the other based on rewards. Both are applied to assignments where the students develop a project focusing on keeping a low technical debt level, and obtaining a high quality code. We describe the design, tools and context of the strategies applied. SonarQube, a tool commonly used in production environments, is used for measuring the metrics. The penalisation strategy is based on a SonarQube quality gate. The reward strategy is based on a contest, where an automatic judge tool is devised to provide an online leaderboard with a classification based on the SonarQube metrics. An empirical study is conducted to determine which of the strategies works better to help the students/trainees keep the Technical Debt low. Statistically significant results are obtained in 5 of the 8 analysed metrics, showing that the reward strategy works much better. The effect size of the executed statistical tests is analysed, resulting in medium and large effect size in the majority of the analysed metrics. Yania Crespo, Arturo González-Escribano, Mario Piattini |
TechDebt@ICSE | 2 |
| 2021 | Distributed programming of a hyperspectral image registration algorithm for heterogeneous GPU clusters
Jorge Fernández-Fabeiro, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Parallel Distributed Comput. | 2 |
| 2021 | Efficient heterogeneous programming with FPGAs using the Controller model
Gabriel Rodriguez-Canal, Yuri Torres, Francisco J. Andujar, Arturo González-Escribano |
J. Supercomput. | 4 |
| 2019 | Automatic runtime calculation of communications for data-parallel expressions with periodic conditionsabstractSummary Many real‐world applications feature data accesses on periodic domains. Manually implementing the synchronizations and communications associated to the data dependences on each case is cumbersome and error‐prone. It is increasingly interesting to support these applications in high‐level parallel programming languages or parallelizing compilers. In this paper, we present a technique that, for distributed‐memory systems, calculates the specific communications derived from data‐parallel codes with or without periodic boundary conditions on affine access expressions. It makes transparent to the programmer the management of aggregated communications for the chosen data partition. Our technique moves to runtime part of the compile‐time analysis typically used to generate the communication code for affine expressions, introducing a complete new technique that also supports the periodic boundary conditions. We present an experimental study to evaluate our proposal using several study cases. Our experimental results show that our approach can automatically obtain communication codes as efficient as those found in MPI reference codes, reducing the development effort. Ana Moreton-Fernandez, Arturo González-Escribano |
Concurr. Comput. Pract. Exp. | 2 |
| 2019 | A multi-device version of the HYFMGPU algorithm for hyperspectral scenes registration
Jorge Fernández-Fabeiro, Álvaro Ordóñez, Arturo González-Escribano, Dora Blanco Heras |
J. Supercomput. | 3 |
| 2019 | Toward a BLAS library truly portable across different accelerator types
Eduardo Rodriguez-Gutiez, Ana Moreton-Fernandez, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Supercomput. | 3 |
| 2017 | Supporting the Xeon Phi Coprocessor in a Heterogeneous Programming Model
Ana Moreton-Fernandez, Eduardo Rodriguez-Gutiez, Arturo González-Escribano, Diego R. Llanos Ferraris |
Euro-Par | 3 |
| 2017 | TORMENT OpenACC2016: A Benchmarking Tool for OpenACC CompilersabstractOpenACC is a parallel programming model for hardware accelerators, such as GPUs or Xeon Phi, which has been in development for several years by now. During this time, different compilers have appeared, both commercial and open source, which are still on development stage. Due to the fact that both the OpenACC standard and its implementations are relatively recent, we propose a benchmark suite specifically designed to check the performance of the OpenACC features in the code generated by different compilers on different architectures. Our benchmark suite is named TORMENT OpenACC2016. Along with this tool we have developed an adequate metric for the comparison of performance among different machine-compiler pairs which we have named TORMENT ACC2016 Score. The version 1 of TORMENT OpenACC2016 presented in this paper, contains six benchmarks, and is available online. Daniel Barba, Arturo González-Escribano, Diego R. Llanos Ferraris |
PDP | 2 |
| 2017 | A technique to automatically determine Ad-hoc communication patterns at runtime
Ana Moreton-Fernandez, Arturo González-Escribano, Diego R. Llanos Ferraris |
Parallel Comput. | 2 |
| 2017 | BFCA+: automatic synthesis of parallel code with TLS capabilities
Sergio Aldea, Diego R. Llanos Ferraris, Arturo González-Escribano |
J. Supercomput. | 3 |
| 2016 | An OpenMP Extension that Supports Thread-Level SpeculationabstractOpenMP directives are the de-facto standard for shared-memory parallel programming. However, OpenMP does not guarantee the correctness of the parallel execution of a given loop if runtime data dependences arise. Consequently, many highly-parallel regions cannot be safely parallelized with OpenMP due to the possibility of a dependence violation. In this paper, we propose to augment OpenMP capabilities, by adding thread-level speculation (TLS) support. Our contribution is threefold. First, we have defined a new speculative clause for variables inside parallel loops. This clause ensures that all accesses to these variables will be carried out according to sequential semantics. Second, we have created a new, software-based TLS runtime library to ensure correctness in the parallel execution of OpenMP loops that include speculative variables. Third, we have developed a new GCC plugin, which seamlessly translates our OpenMP speculative clause into calls to our TLS runtime engine. The result is the ATLaS C Compiler framework, which takes advantage of TLS techniques to expand OpenMP functionalities, and guarantees the sequential semantics of any parallelized loop. Sergio Aldea, Alvaro Estebanez, Diego R. Llanos Ferraris, Arturo González-Escribano |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2014 | A New GCC Plugin-Based Compiler Pass to Add Support for Thread-Level Speculation into OpenMP
Sergio Aldea, Alvaro Estebanez, Diego R. Llanos Ferraris, Arturo González-Escribano |
Euro-Par | 4 |
| 2014 | Squashing Alternatives for Software-Based Speculative ParallelizationabstractSpeculative parallelization is a runtime technique that optimistically executes sequential code in parallel, checking that no dependence violations arise. In the case of a dependence violation, all mechanisms proposed so far either switch to sequential execution, or conservatively stop and restart the offending thread and all its successors, potentially discarding work that does not depend on this particular violation. In this work we systematically explore the design space of solutions for this problem, proposing a new mechanism that reduces the number of threads that should be restarted when a data dependence violation is found. Our new solution, called exclusive squashing, keeps track of inter-thread dependencies at runtime, selectively stopping and restarting offending threads, together with all threads that have consumed data from them. We have compared this new approach with existent solutions on a real system, executing different applications with loops that are not analyzable at compile time and present as much as 10% of inter-thread dependence violations at runtime. Our experimental results show a relative performance improvement of up to 14%, together with a reduction of one-third of the numbers of squashed threads. The speculative parallelization scheme and benchmarks described in this paper are available under request. Álvaro García-Yágüez, Diego R. Llanos Ferraris, Arturo González-Escribano |
IEEE Trans. Computers | 3 |
| 2014 | The BonaFide C Analyzer: automatic loop-level characterization and coverage measurement
Sergio Aldea, Diego R. Llanos Ferraris, Arturo González-Escribano |
J. Supercomput. | 3 |
| 2014 | Optimizing an APSP implementation for NVIDIA GPUs using kernel characterization criteria
Hector Ortega-Arranz, Yuri Torres, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Supercomput. | 3 |
| 2014 | Blending Extensibility and Performance in Dense and Sparse Parallel Data ManagementabstractDealing with both dense and sparse data in parallel environments usually leads to two different approaches: To rely on a monolithic, hard-to-modify parallel library, or to code all data management details by hand. In this paper we propose a third approach, that delivers good performance while the underlying library structure remains modular and extensible. Our solution integrates dense and sparse data management using a common interface, that also decouples data representation, partitioning, and layout from the algorithmic and parallel strategy decisions of the programmer. Our experimental results in different parallel environments show that this new approach combines the flexibility obtained when the programmer handles all the details with a performance comparable to the use of a state-of-the-art, sparse matrix parallel library. Javier Fresno, Arturo González-Escribano, Diego R. Llanos Ferraris |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2014 | An Extensible System for Multilevel Automatic Data Partition and MappingabstractAutomatic data distribution is a key feature to obtain efficient implementations from abstract and portable parallel codes. We present a highly efficient and extensible runtime library that integrates techniques for automatic data partition and mapping. It uses a novel approach to define an abstract interface and a plug-in system to encapsulate different types of regular and irregular techniques, helping to generate codes which are independent of the exact mapping functions selected. Currently, it supports hierarchical tiling of arrays with dense and stride domains, that allows the implementation of both data and task parallelism using a SPMD model. It automatically computes appropriate domain partitions for a selected virtual topology, mapping them to available processors with static or dynamic load-balancing techniques. Our library also allows the construction of reusable communication patterns that efficiently exploit MPI communication capabilities. The use of our library greatly reduces the complexity of data distribution and communication, hiding the details of the underlying architecture. The library can be used as an abstract layer for building generic tiling operations as well. Our experimental results show that the use of this library allows to achieve similar performance as carefully-implemented manual versions for several, well-known parallel kernels and benchmarks in distributed and multicore systems, and substantially reduces programming effort. Arturo González-Escribano, Yuri Torres, Javier Fresno, Diego R. Llanos Ferraris |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2013 | Extending a hierarchical tiling arrays library to support sparse data partitioning
Javier Fresno, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Supercomput. | 2 |
| 2013 | uBench: exposing the impact of CUDA block geometry in terms of performance
Yuri Torres, Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Supercomput. | 2 |
| 2012 | Encapsulated Synchronization and Load-Balance in Heterogeneous Programming
Yuri Torres, Arturo González-Escribano, Diego R. Llanos Ferraris |
Euro-Par | 2 |
| 2012 | Using Fermi Architecture Knowledge to Speed up CUDA and OpenCL ProgramsabstractThe NVIDIA graphics processing units (GPUs) are playing an important role as general purpose programming devices. The implementation of parallel codes to exploit the GPU hardware architecture is a task for experienced programmers. The threadblock size and shape choice is one of the most important user decisions when a parallel problem is coded. The threadblock configuration has a significant impact on the global performance of the program. While in CUDA parallel programming model it is always necessary to specify the threadblock size and shape, the OpenCL standard also offers an automatic mechanism to take this delicate decision. In this paper we present a study of these criteria for Fermi architecture, introducing a general approach for threadblock choice, and showing that there is considerable room for improvement in OpenCL automatic strategy. Yuri Torres, Arturo González-Escribano, Diego R. Llanos Ferraris |
ISPA | 2 |
| 2012 | Using SPEC CPU2006 to evaluate the sequential and parallel code generated by commercial and open-source compilers
Sergio Aldea, Diego R. Llanos Ferraris, Arturo González-Escribano |
J. Supercomput. | 3 |
| 2011 | Robust thread-level speculationabstractRobustness is a key issue on any runtime system that aims to speed up the execution of a program. However, robustness considerations are commonly overlooked when new software-based, thread-level speculation (STLS) systems are proposed. This paper highlights the relevance of the problem, showing different situations when the use of incorrect data can irreversibly alter the speculative execution of an algorithm, despite the efforts of a given STLS system to maintain sequential consistency. We show that the management of speculative exceptions is a common factor to these problems. Based on this fact, we propose a novel solution to handle speculative exceptions. Our solution eagerly tries to solve the issue before the non-speculative thread arrives to the instruction that rose the exception. We compare our solution to a more conservative approach found in the bibliography. The comparison is done both qualitatively, through a detailed analysis of the tradeoffs involved, and quantitatively, evaluating the effects of both solutions in the execution of three different benchmarks on a real system. Both studies conclude that our solution handles the occurrence of speculative exceptions more efficiently. Under heavy loads intended to push to its limits a STLS system, our solution leads to execution times reduced by up to 52.02% with respect to earlier proposals. Our solution does not affect the performance when speculative exceptions do not appear. We believe that our proposal makes STLS systems robust enough to be used in production environments. Álvaro García-Yágüez, Diego R. Llanos Ferraris, Arturo González-Escribano |
HiPC | 3 |
| 2011 | Exclusive squashing for thread-level speculationabstractSpeculative parallelization is a runtime technique that optimistically executes sequential code in parallel, checking that no dependence violations appear. In this paper, we address the problem of minimizing the number of threads that should be restarted when a data dependence violation is found. We present a new mechanism that keeps track of inter-thread dependencies in order to selectively stop and restart offending threads, and all threads that have consumed data from them. Results show a reduction of 38.5% to 81.8% in the number of restarted threads for real application loops and up to a 10% speedup, depending on the amount of local computation. Álvaro García-Yágüez, Diego R. Llanos Ferraris, Arturo González-Escribano |
HPDC | 3 |
| 2011 | Towards a Compiler Framework for Thread-Level SpeculationabstractSpeculative parallelization techniques allow to extract parallelism of fragments of code that can not be analyzed at compile time. However, research on software-based, thread-level speculation will greatly benefit from an appropriate compiler framework for easy prototyping and further development of new techniques. This paper presents an experimental XML-based compilation framework to handle speculative parallelization of C code. The framework extends Cetus, a source-to-source C compiler, to build an XML tree based on the Cetus Internal Representation of the source code. Other modules of our framework rely on XPath and XSLT capabilities to process the XML tree generated, to perform analysis on the use of variables and to augment the original code for software-based, speculative parallel execution. The use of the current version of our framework allows a fast prototyping of new analysis and transformation solutions, with a reduction of around 83% on the number of code lines needed with respect to the direct use of Cetus for the same purpose. To show the possibilities of this framework, we present an automatically-generated classification of loops for several SPEC CPU2006 C benchmarks. This classification is useful to better understand the potential benefits derived from the use of speculative parallelization techniques. The development framework presented here is freely available under request. Sergio Aldea, Diego R. Llanos Ferraris, Arturo González-Escribano |
PDP | 3 |
| 2011 | Automatic Data Partitioning Applied to Multigrid PDE SolversabstractThis paper studies the impact of using automatic data-layout techniques on the process of coding the well-known multigrid MG NAS parallel benchmark. We describe the sequential problem in detail, and discuss the parallel version and its optimizations. Then, we implement the parallel algorithm using Hit map, a highly-efficient modular library for hierarchical tiling and mapping of arrays. We describe how to use the library plug-in system to add a new data-layout module that encapsulates a generalization of the data-alignment policy of the MG benchmark. The module system applies this policy to automatically adapt the data distribution and communication code to any grain level. The impact of using these techniques is qualitatively and quantitatively described in terms of development effort and performance. Our results show that it is possible to introduce flexible automatic data-layout techniques in current parallel compiler technology, without sacrificing performance. Javier Fresno, Arturo González-Escribano, Diego R. Llanos Ferraris |
PDP | 2 |
| 2011 | Trasgo: a nested-parallel programming system
Arturo González-Escribano, Diego R. Llanos Ferraris |
J. Supercomput. | 1 |
| 2010 | Effortless and Efficient Distributed Data-Partitioning in Linear AlgebraabstractThis paper introduces a new technique to exploit compositions of different data-layout techniques with Hit map, a library for hierarchical-tiling and automatic mapping of arrays. We show how Hit map is used to implement block-cyclic layouts for a parallel LU decomposition algorithm. The paper compares the well-known ScaLAPACK implementation of LU, as well as other carefully optimized MPI versions, with a Hit map implementation. The comparison is made in terms of both performance and code length. Our results show that the Hit map version outperforms the ScaLAPACK implementation and is almost as efficient as our best manual MPI implementation. The insertion of this composition technique in the automatic data-layouts of Hit map allows the programmer to develop parallel programs with both a significant reduction of the development effort and a negligible loss of efficiency. Carlos de Blas Carton, Arturo González-Escribano, Diego R. Llanos Ferraris |
HPCC | 2 |
| 2009 | Performance implications of synchronization structure in parallel programming
Arturo González-Escribano, Arjan J. C. van Gemund, Valentín Cardeñoso-Payo |
Parallel Comput. | 1 |
| 2005 | SPC-XML: A Structured Representation for Nested-Parallel Programming Languages
Arturo González-Escribano, Arjan J. C. van Gemund, Valentín Cardeñoso-Payo |
Euro-Par | 1 |