EDBT 2026 Demo / reviewers in the wild / expert
Rafael Mayo 0002
dblp:272/5869 · also Rafael Mayo Gual
· DBLP profile ↗
49ranked-venue papers
1as first author
2since 2021 · last 2021
0000-0003-1552-3069ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 41 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Cloud and datacenter computing · 61% Parallel and multicore computing · 37% Performance modeling and evaluation · 3% | |
| Software engineering, system software, and programming languages
1 paper |
Operating systems · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing › resource allocation
elastic resource allocation |
0.5 | 1 | 2021 | DMRlib: Easy-Coding and Efficient Resource Management for Job Malleability · IEEE Trans. Computers 2021 |
Cloud and datacenter computing
job scheduling |
0.5 | 1 | 2021 | DMRlib: Easy-Coding and Efficient Resource Management for Job Malleability · IEEE Trans. Computers 2021 |
Cloud and datacenter computing
resource management |
0.5 | 1 | 2021 | DMRlib: Easy-Coding and Efficient Resource Management for Job Malleability · IEEE Trans. Computers 2021 |
Parallel and multicore computing › parallel computing › parallel optimization
parallel code optimization |
0.4 | 1 | 2020 | Analysis of Threading Libraries for High Performance Computing · IEEE Trans. Computers 2020 |
Parallel and multicore computing
parallel programming models |
0.4 | 1 | 2020 | Analysis of Threading Libraries for High Performance Computing · IEEE Trans. Computers 2020 |
Parallel and multicore computing
parallel programming runtimes |
0.4 | 1 | 2020 | Analysis of Threading Libraries for High Performance Computing · IEEE Trans. Computers 2020 |
Cloud and datacenter computing
cloud bursting |
0.3 | 1 | 2018 | Performance Model of MapReduce Iterative Applications for Hybrid Cloud Bursting · IEEE Trans. Parallel Distributed Syst. 2018 |
Cloud and datacenter computing › cloud deployment
hybrid cloud |
0.3 | 1 | 2018 | Performance Model of MapReduce Iterative Applications for Hybrid Cloud Bursting · IEEE Trans. Parallel Distributed Syst. 2018 |
Performance modeling and evaluation › performance prediction
execution time prediction |
0.1 | 1 | 2018 | Performance Model of MapReduce Iterative Applications for Hybrid Cloud Bursting · IEEE Trans. Parallel Distributed Syst. 2018 |
Methods — techniques the papers use, named apart from their topics
microbenchmarking · 0.9MPI · 0.5performance modeling · 0.3mapreduce · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Malleability Implementation in a MPI Iterative MethodabstractIn this poster is evaluated the data redistribution stage for two malleable versions of the Conjugate Gradient. One version is based on synchronous communications, while the other one uses asynchronous communications to overlap computation and data redistribution. Both improve execution time when adding more processes, but there is not a noticeable difference between them, because the asynchronous method lowers the performance of the iterations due to the method’s own communications. When both versions are compared, the synchronous version is preferred when resizing to more processes, while the asynchronous one achieves better times when resizing to fewer processes. Iker Martín-Álvarez, José Ignacio Aliaga, María Isabel Castillo, Rafael Mayo 0002, Sergio Iserte |
CLUSTER | 4 |
| 2021 | DMRlib: Easy-Coding and Efficient Resource Management for Job MalleabilityabstractProcess malleability has proved to have a highly positive impact on the resource utilization and global productivity in data centers compared with the conventional static resource allocation policy. However, the non-negligible additional development effort this solution imposes has constrained its adoption by the scientific programming community. In this work, we present DMRlib, a library designed to offer the global advantages of process malleability while providing a minimalist MPI-like syntax. The library includes a series of predefined communication patterns that greatly ease the development of malleable applications. In addition, we deploy several scenarios to demonstrate the positive impact of process malleability featuring different scalability patterns. Concretely, we study two job submission modes (rigid and moldable) in order to identify the best-case scenarios for malleability using metrics such as resource allocation rate, completed jobs per second, and energy consumption. The experiments prove that our elastic approach may improve global throughput by a factor higher than 3x compared to the traditional workloads of non-malleable jobs. Sergio Iserte, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Antonio J. Peña |
IEEE Trans. Computers | 2 |
| 2020 | Analysis of Threading Libraries for High Performance ComputingabstractWith the appearance of multi-/many core machines, applications and runtime systems have evolved in order to exploit the new on-node concurrency brought by new software paradigms. POSIX threads (Pthreads) was widely-adopted for that purpose and it remains as the most used threading solution in current hardware. Lightweight thread (LWT) libraries emerged as an alternative offering lighter mechanisms to tackle the massive concurrency of current hardware. In this article, we analyze in detail the most representative threading libraries including Pthread- and LWT-based solutions. In addition, to examine the suitability of LWTs for different use cases, we develop a set of microbenchmarks consisting of OpenMP patterns commonly found in current parallel codes, and we compare the results using threading libraries and OpenMP implementations. Moreover, we study the semantics offered by threading libraries in order to expose the similarities among different LWT application programming interfaces and their advantages over Pthreads. This article exposes that LWT libraries outperform solutions based on operating system threads when tasks and nested parallelism are required. Adrián Castelló 0001, Rafael Mayo 0002, Pavan Balaji, Enrique S. Quintana-Ortí, Antonio J. Peña |
IEEE Trans. Computers | 2 |
| 2019 | Noise estimation for hyperspectral subspace identification on FPGAs
German Leon, Carlos González 0002, Rafael Mayo 0002, Daniel Mozos, Enrique S. Quintana-Ortí |
J. Supercomput. | 3 |
| 2018 | On the adequacy of lightweight thread approaches for high-level parallel programming models
Adrián Castelló 0001, Rafael Mayo 0002, Kevin Sala, Vicenç Beltran 0001, Pavan Balaji, Antonio J. Peña |
Future Gener. Comput. Syst. | 2 |
| 2018 | DMR API: Improving cluster productivity by turning applications into malleable
Sergio Iserte, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Vicenç Beltran 0001, Antonio J. Peña |
Parallel Comput. | 2 |
| 2018 | Exploring the interoperability of remote GPGPU virtualization using rCUDA and directive-based programming models
Adrián Castelló 0001, Antonio J. Peña, Rafael Mayo 0002, Judit Planas, Enrique S. Quintana-Ortí, Pavan Balaji |
J. Supercomput. | 3 |
| 2018 | Performance Model of MapReduce Iterative Applications for Hybrid Cloud BurstingabstractHybrid cloud bursting (i.e., leasing temporary off-premise cloud resources to boost the overall capacity during peak utilization) can be a cost-effective way to deal with the increasing complexity of big data analytics, especially for iterative applications. However, the low throughput, high latency network link between the on-premise and off-premise resources (“weak link”) makes maintaining scalability difficult. While several data locality techniques have been designed for big data bursting on hybrid clouds, their effectiveness is difficult to estimate in advance. Yet such estimations are critical, because they help users decide whether the extra pay-as-you-go cost incurred by using the off-premise resources justifies the runtime speed-up. To this end, the current paper presents a performance model and methodology to estimate the runtime of iterative MapReduce applications in a hybrid cloud-bursting scenario. The paper focuses on the overhead incurred by the weak link at fine granularity, for both the map and the reduce phases. This approach enables high estimation accuracy, as demonstrated by extensive experiments at scale using a mix of real-world iterative MapReduce applications from standard big data benchmarking suites that cover a broad spectrum of data patterns. Not only are the produced estimations accurate in absolute terms compared with experimental results, but they are also up to an order of magnitude more accurate than applying state-of-art estimation approaches originally designed for single-site MapReduce deployments. Francisco J. Clemente-Castelló, Bogdan Nicolae, Rafael Mayo 0002, Juan Carlos Fernández 0002 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | Evaluation of Data Locality Strategies for Hybrid Cloud Bursting of Iterative MapReduceabstractHybrid cloud bursting (i.e., leasing temporary off-premise cloud resources to boost the overall capacity during peak utilization) is a popular and cost-effective way to deal with the increasing complexity of big data analytics. It is particularly promising for iterative MapReduce applications that reuse massive amounts of input data at each iteration, which compensates for the high overhead and cost of concurrent data transfers from the on-premise to the off-premise VMs over a weak inter-site link that is of limited capacity. In this paper we study how to combine various MapReduce data locality techniques designed for hybrid cloud bursting in order to achieve scalability for iterative MapReduce applications in a cost-effective fashion. This is a non trivial problem due to the complex interaction between the data movements over the weak link and the scheduling of computational tasks that have to adapt to the shifting data distribution. We show that using the right combination of techniques, iterative MapReduce applications can scale well in a hybrid cloud bursting scenario and come even close to the scalability observed in single sites. Francisco J. Clemente-Castelló, Bogdan Nicolae, M. Mustafa Rafique, Rafael Mayo 0002, Juan Carlos Fernández 0002 |
CCGrid | 4 |
| 2017 | Cost Model and Analysis of Iterative MapReduce Applications for Hybrid Cloud BurstingabstractA popular and cost-effective way to deal with the increasing complexity of big data analytics is hybrid cloud bursting that leases temporary off-premise cloud resources to boost the overall capacity during peak utilization. The main challenge of hybrid cloud bursting is that the network link between the on-premise and the off-premise computational resources often exhibit high latency and low throughput ("weak link") compared to the links within the same data-center. This paper introduces a cost model that is specifically designed for iterative MapReduce applications running in a hybrid cloud bursting scenario, which are a popular class of large-scale data-intensive applications that provides near real-time responsiveness. Using this cost model, users can discover trends that can be leveraged to reason about how to balance performance, accuracy and cost such that it op-timizes their requirements. We illustrated this approach through a cost analysis that focuses on two real-life iterative MapReduce applications using extensive horizontal scalability experiments that involve multiple hybrid cloud bursting strategies. Francisco J. Clemente-Castelló, Rafael Mayo 0002, Juan Carlos Fernández 0002 |
CCGrid | 2 |
| 2017 | GLT: A Unified API for Lightweight Thread Libraries
Adrián Castelló 0001, Rafael Mayo 0002, Pavan Balaji, Enrique S. Quintana-Ortí, Antonio J. Peña |
Euro-Par | 3 |
| 2017 | GLTO: On the Adequacy of Lightweight Thread Approaches for OpenMP ImplementationsabstractOpenMP is the de facto standard application programming interface (API) for on-node parallelism. The most popular OpenMP runtimes rely on POSIX threads (pthreads) implementations that offer an excellent performance for coarse-grained parallelism and match perfectly with the current hardware. However, a recent trend in runtimes/applications points in the direction of leveraging massive on-node parallelism in conjunction with fine-grained and dynamic scheduling paradigms. It has been demonstrated that lightweight thread (LWT) solutions are more appropriate for these new parallel paradigms. We have developed GLTO, an OpenMP implementation over the recently-emerged Generic Lightweight Threads (GLT) API. GLT exports a common API for LWT libraries that offers the possibility of running the same application over different native LWT solutions. In this paper we use GLTO to analyze different scenarios where OpenMP implementations may benefit from the use of either LWT or pthreads. Our study reveals that none of the threading approaches obtains the best performance in all the scenarios, but that there are important gaps among them. Adrián Castelló 0001, Rafael Mayo 0002, Pavan Balaji, Enrique S. Quintana-Ortí, Antonio J. Peña |
ICPP | 3 |
| 2017 | Time and energy modeling of a high-performance multi-threaded Cholesky factorization
Sandra Catalán, Francisco D. Igual, Rafael Mayo 0002, Rafael Rodríguez-Sánchez 0001, Enrique S. Quintana-Ortí |
J. Supercomput. | 3 |
| 2016 | On exploiting data locality for iterative mapreduce applications in hybrid cloudsabstractHybrid cloud bursting (i.e., leasing temporary off-premise cloud resources to boost the capacity during peak utilization), has made significant impact especially for big data analytics, where the explosion of data sizes and increasingly complex computations frequently leads to insufficient local data center capacity. Cloud bursting however introduces a major challenge to runtime systems due to the limited throughput and high latency of data transfers between on-premise and off-premise resources (weak link). This issue and how to address it is not well understood. We contribute with a comprehensive study on what challenges arise in this context, what potential strategies can be applied to address them and what best practices can be leveraged in real-life. Specifically, we focus our study on iterative MapReduce applications, which are a class of large-scale data intensive applications particularly popular on hybrid clouds. In this context, we study how data locality can be leveraged over the weak link both from the storage layer perspective (when and how to move it off-premise) and from the scheduling perspective (when to compute off-premise). We conclude with a brief discussion on how to set up an experimental framework suitable to study the effectiveness of our proposal in future work. Francisco J. Clemente-Castelló, Bogdan Nicolae, Rafael Mayo 0002, Juan Carlos Fernández 0002, M. Mustafa Rafique |
BDCAT | 3 |
| 2016 | Enabling GPU Virtualization in Cloud EnvironmentsabstractThe use of accelerators, such as graphics processing units (GPUs), to reduce the execution time of compute-intensive applications has become popular during the past few years. These devices increment the computational power of a node thanks to their parallel architecture. This trend has led cloud service providers as Amazon or middlewares such as OpenStack to add virtual machines (VMs) including GPUs to their facilities instances. To fulfill these needs, the guest hosts must be equipped with GPUs which, unfortunately, will be barely utilized if a non GPU-enabled VM is running in the host. The solution presented in this work is based on GPU virtualization and shareability in order to reach an equilibrium between service supply and the applicationsâ?? demand of accelerators. Concretely, we propose to decouple real GPUs from the nodes by using the virtualization technology rCUDA. With this software configuration, GPUs can be accessed from any VM avoiding the need of placing a physical GPUs in each guest host. Moreover, we study the viability of this approach using a public cloud service configuration, and we develop a module for OpenStack in order to add support for the virtualized devices and the logic to manage them. The results demonstrate this is a viable configuration which adds flexibility to current and well-known cloud solutions. Sergio Iserte, Francisco J. Clemente-Castelló, Adrián Castelló 0001, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
CLOSER (2) | 4 |
| 2016 | A Review of Lightweight Thread Approaches for High Performance ComputingabstractHigh-level, directive-based solutions are becoming the programming models (PMs) of the multi/many-core architectures. Several solutions relying on operating system (OS) threads perfectly work with a moderate number of cores. However, exascale systems will spawn hundreds of thousands of threads in order to exploit their massive parallel architectures and thus conventional OS threads are too heavy for that purpose. Several lightweight thread (LWT) libraries have recently appeared offering lighter mechanisms to tackle massive concurrency. In order to examine the suitability of LWTs in high-level runtimes, we develop a set of microbenchmarks consisting of commonly-found patterns in current parallel codes. Moreover, we study the semantics offered by some LWT libraries in order to expose the similarities between different LWT application programming interfaces. This study reveals that a reduced set of LWT functions can be sufficient to cover the common parallel code patterns andthat those LWT libraries perform better than OS threads-based solutions in cases where task and nested parallelism are becoming more popular with new architectures. Adrián Castelló 0001, Antonio J. Peña, Rafael Mayo 0002, Pavan Balaji, Enrique S. Quintana-Ortí |
CLUSTER | 4 |
| 2015 | Exploring the Suitability of Remote GPGPU Virtualization for the OpenACC Programming Model Using rCUDAabstractOpenACC is an application programming interface (API) that aims to unleash the power of heterogeneous systems composed of CPUs and accelerators such as graphic processing units (GPUs) or Intel Xeon Phi coprocessors. This directive-based programming model is intended to enable developers to accelerate their application's execution with much less effort. Coprocessors offer significant computing power but in many cases these devices remain largely underused because not all parts of applications match the accelerator architecture. Remote accelerator virtualization frameworks introduce a means to address this problem. In particular, the remote CUDA virtualization middleware rCUDA provides transparent remote access to any GPU installed in a cluster. Combining these two technologies, OpenACC and rCUDA, in a single scenario is naturally appealing. In this work we explore how the different OpenACC directives behave on top of a remote GPGPU virtualization technology in two different hardware configurations. Our experimental evaluation reveals favorable performance results when the two technologies are combined, showing low overhead and similar scaling factors when executing OpenACC-enabled directives. Adrián Castelló 0001, Antonio J. Peña, Rafael Mayo 0002, Pavan Balaji, Enrique S. Quintana-Ortí |
CLUSTER | 3 |
| 2015 | Time and energy modeling of an INTRA-ONLY HEVC encoderabstractIn this paper, we present precise time and energy models for an intra-only HEVC video encoder. These models are a step forward to understand and estimate the computational complexity and energy demands of an HEVC encoder, which in turn opens the path to finely tuning the computational resources that are dedicated to this purpose. Our models estimate the complexity and energy consumed by the HEVC encoder, in a frame by frame basis, considering two factors: the quantification parameter used to encode each frame and the spatial information of that frame. Our experimental validation demonstrates the accuracy of these models, which report errors that are, on average, below 10% for full HD videos, and 5% for 832 × 480 videos. Rafael Rodríguez-Sánchez 0001, Maria Teresa Alonso, José Luis Martínez 0001, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
VCIP | 4 |
| 2015 | Out-of-core macromolecular simulations on multithreaded architecturesabstractSummary We address the solution of large‐scale eigenvalue problems that appear in the motion simulation of complex macromolecules on multithreaded platforms, consisting of multicore processors and possibly a graphics processor (graphics processing unit). In particular, we compare specialized implementations of several high‐performance eigensolvers that, by relying on disk storage and out‐of‐core techniques, can in principle tackle the large memory requirements of these biological problems, which in general do not fit into the main memory of current desktop machines. All these out‐of‐core eigensolvers, except for one, are composed of compute‐bound (i.e., arithmetically intensive) operations, which we accelerate by exploiting the performance of current multicore processors and, in some cases, by additionally off‐loading certain parts of the computation to a graphics processing unit accelerator. One of the eigensolvers is a memory‐bound algorithm, which strongly constrains its performance when the data is on disk. However, this method exhibits a much lower arithmetic cost compared with its compute‐bound alternatives for this particular application. Experimental results on a desktop platform, representative of current server technology, illustrate the potential of these methods to address the simulation of biological activity. Copyright © 2014 John Wiley & Sons, Ltd. José Ignacio Aliaga, José M. Badía, María Isabel Castillo, Davor Davidovic, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
Concurr. Comput. Pract. Exp. | 5 |
| 2015 | Improving the user experience of the rCUDA remote GPU virtualization frameworkabstractSummary Graphics processing units (GPUs) are being increasingly embraced by the high‐performance computing community as an effective way to reduce execution time by accelerating parts of their applications. remote CUDA (rCUDA) was recently introduced as a software solution to address the high acquisition costs and energy consumption of GPUs that constrain further adoption of this technology. Specifically, rCUDA is a middleware that allows a reduced number of GPUs to be transparently shared among the nodes in a cluster. Although the initial prototype versions of rCUDA demonstrated its functionality, they also revealed concerns with respect to usability, performance, and support for new CUDA features. In response, in this paper, we present a new rCUDA version that (1) improves usability by including a new component that allows an automatic transformation of any CUDA source code so that it conforms to the needs of the rCUDA framework, (2) consistently features low overhead when using remote GPUs thanks to an improved new communication architecture, and (3) supports multithreaded applications and CUDA libraries. As a result, for any CUDA‐compatible program, rCUDA now allows the use of remote GPUs within a cluster with low overhead, so that a single application running in one node can use all GPUs available across the cluster, thereby extending the single‐node capability of CUDA. Copyright © 2014 John Wiley & Sons, Ltd. Carlos Reaño, Federico Silla, Adrián Castelló 0001, Antonio J. Peña, Rafael Mayo 0002, Enrique S. Quintana-Ortí, José Duato |
Concurr. Comput. Pract. Exp. | 5 |
| 2014 | Evaluating the Impact of Virtualization on Performance and Power DissipationabstractIn this paper we assess the impact of virtualization in both performance-oriented environments, like high
performance computing facilities, and throughput-oriented systems, like data processing centers, e.g., for
web search and data serving. In particular, our work-in-progress analyzes the power consumption required
to dynamically migrate virtual machines at runtime, a technique that is crucial to consolidate underutilized
servers, reducing energy costs while maintaining service level agreements. Preliminary experimental results
are reported for two different applications, using the KVM virtualization solution for Linux, on an Intel Xeon-
based cluster. Francisco J. Clemente-Castelló, Sonia Cervera, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
CLOSER | 3 |
| 2014 | Analyzing the Energy Efficiency of the Memory Subsystem in Multicore ProcessorsabstractIn this paper we analyze the energy overhead incurred when operating with data stored in different levels of the memory subsystem (cache levels and DDR chips) of current multicore architectures. Our approach builds upon servet, a portable framework for the memory characterization of multicore processors, extending this suite with a power-related test that, when applied to a platform equipped with a power measurement mechanism, provides information on the efficiency of memory energy usage. As additional contributions, i) we provide a complete experimental study of the impact that the CPU performance states (also known as P-states) exert on the memory energy efficiency of a collection of recent server-oriented and low-power cores, and ii) we show how this framework carries over to cover also the scalability analysis of the memory energy performance on multicore processors. Sandra Catalán, Jorge González-Domínguez, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
ISPA | 3 |
| 2014 | SLURM Support for Remote GPU Virtualization: Implementation and Performance StudyabstractSLURM is a resource manager that can be leveraged to share a collection of heterogeneous resources among the jobs in execution in a cluster. However, SLURM is not designed to handle resources such as graphics processing units (GPUs). Concretely, although SLURM can use a generic resource plugin (GRes) to manage GPUs, with this solution the hardware accelerators can only be accessed by the job that is in execution on the node to which the GPU is attached. This is a serious constraint for remote GPU virtualization technologies, which aim at providing a user-transparent access to all GPUs in cluster, independently of the specific location of the node where the application is running with respect to the GPU node. In this work we introduce a new type of device in SLURM, "rgpu", in order to gain access from any application node to any GPU node in the cluster using rCUDA as the remote GPU virtualization solution. With this new scheduling mechanism, a user can access any number of GPUs, as SLURM schedules the tasks taking into account all the graphics accelerators available in the complete cluster. We present experimental results that show the benefits of this new approach in terms of increased flexibility for the job scheduler. Sergio Iserte, Adrián Castelló 0001, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Federico Silla, José Duato, Carlos Reaño, Javier Prades |
SBAC-PAD | 3 |
| 2014 | Enhancing performance and energy consumption of runtime schedulers for dense linear algebraabstractSUMMARY The road towards Exascale Computing requires a holistic effort to address three different challenges simultaneously: high performance, energy efficiency, and programmability. The use of runtime task schedulers to orchestrate parallel executions with minimal developer intervention has been introduced in recent years to tackle the programmability issue while maintaining, or even improving, performance. In this paper, we enhance the SuperMatrix runtime task scheduler integrated in the libflame library in two different directions that address high performance and energy efficiency. First, we extend the runtime by accommodating hybrid parallel executions and managing task priorities for dense linear algebra operations, with remarkable performance improvements. Second, we introduce techniques to reduce energy consumption during idle times inherent to parallel executions, attaining important energy savings. In addition, we propose a power consumption model that can be leveraged by runtime task schedulers to make decisions based not only on performance but also on energy considerations. Copyright © 2014 John Wiley & Sons, Ltd. Pedro Alonso 0002, Manuel F. Dolz, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
Concurr. Comput. Pract. Exp. | 4 |
| 2014 | Modeling power and energy consumption of dense matrix factorizations on multicore processorsabstractSUMMARY In this paper, we propose a model for the energy consumption of the concurrent execution of three key dense matrix factorizations, with task parallelism leveraged via the Symmetric Multi‐Processing Superscalar (SMPSs) runtime, on a multicore processor. Our model decomposes the power dissipation into the system, static and dynamic components, with the former two being estimated from basic, off‐line experiments. The dynamic power, on the other hand, requires significantly more care, and we introduce a contention‐aware model that accommodates for the variability of power consumption due to memory contention. Experimental results on an Intel Xeon E5504 processor with four cores, using an internal powermeter that samples the power drawn by the mainboard with a frequency of 1 KHz, show the reliability of the energy model for the Cholesky, LU, and QR factorizations on this platform. Copyright © 2013 John Wiley & Sons, Ltd. Pedro Alonso 0002, Manuel F. Dolz, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
Concurr. Comput. Pract. Exp. | 3 |
| 2014 | A complete and efficient CUDA-sharing solution for HPC clusters
Antonio J. Peña, Carlos Reaño, Federico Silla, Rafael Mayo 0002, Enrique S. Quintana-Ortí, José Duato |
Parallel Comput. | 4 |
| 2013 | Influence of InfiniBand FDR on the performance of remote GPU virtualizationabstractThe use of GPUs to accelerate general-purpose scientific and engineering applications is mainstream today, but their adoption in current high-performance computing clusters is impaired primarily by acquisition costs and power consumption. Therefore, the benefits of sharing a reduced number of GPUs among all the nodes of a cluster can be remarkable for many applications. This approach, usually referred to as remote GPU virtualization, aims at reducing the number of GPUs present in a cluster, while increasing their utilization rate. The performance of the interconnection network is key to achieving reasonable performance results by means of remote GPU virtualization. To this end, several networking technologies with throughput comparable to that of PCI Express have appeared recently. In this paper we analyze the influence of InfiniBand FDR on the performance of remote GPU virtualization, comparing its impact on a variety of GPU-accelerated applications with other networking technologies, such as Infini-Band QDR and Gigabit Ethernet. Given the severe limitations of freely available remote GPU virtualization solutions, the rCUDA framework is used as the case study for this analysis. Results show that the new FDR interconnect, featuring higher bandwidth than its predecessors, allows the reduction of the overhead of using GPUs remotely, thus making this approach even more appealing. Carlos Reaño, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Federico Silla, José Duato, Antonio J. Peña |
CLUSTER | 2 |
| 2012 | CU2rCU: Towards the complete rCUDA remote GPU virtualization and sharing solutionabstractGPUs are being increasingly embraced by the high performance computing and computational communities as an effective way of considerably reducing execution time by accelerating significant parts of their application codes. However, despite their extraordinary computing capabilities, the adoption of GPUs in current HPC clusters may present certain negative side-effects. In particular, to ease job scheduling in these platforms, a GPU is usually attached to every node of the cluster. In addition to increasing acquisition costs this favors that GPUs may frequently remain idle, as applications usually do not fully utilize them. On the other hand, idle GPUs consume non-negligible amounts of energy, which translates into very poor energy efficiency during idle cycles. rCUDA was recently developed as a software solution to address these concerns. Specifically, it is a middleware that allows transparently sharing a reduced number of GPUs among the nodes in a cluster. rCUDA thus increases the GPU-utilization rate, taking care of job scheduling. While the initial prototype versions of rCUDA demonstrated its functionality, they also revealed several concerns related with usability and performance. With respect to usability, in this paper we present a new component of the rCUDA suite that allows an automatic transformation of any CUDA source code, so that it can be effectively accommodated within this technology. In response to performance, we briefly show some interesting results, which will be deeply analyzed in future publications. The net outcome is a new version of rCUDA that allows, for any CUDA-compatible program, to use remote GPUs in a cluster with minimum overhead. Carlos Reaño, Antonio J. Peña, Federico Silla, José Duato, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
HiPC | 5 |
| 2012 | Tools for Power-Energy Modelling and Analysis of Parallel Scientific ApplicationsabstractUnderstanding power usage in parallel workloads is crucial to develop the energy-aware software that will run in future Exascale systems. In this paper, we contribute towards this goal by introducing an integrated framework to profile, monitor, model and analyze power dissipation in parallel MPI and multi-threaded scientific applications. The framework includes an own-designed device to measure internal DC power consumption and a package offering a simple interface to interact with this design as well as commercial power meters. Combined with the instrumentation package Extrae and the graphical analysis tool Paraver, the result is a useful environment to identify sources of power inefficiency directly in the source application code. For task-parallel codes, we also offer a statistical software module that inspects the execution trace of the application to calculate the parameters of an accurate model for the global energy consumption, which can be then decomposed into the average power usage per task or the nodal power dissipated per core. Pedro Alonso 0002, Rosa M. Badia, Jesús Labarta, Maria Barreda, Manuel F. Dolz, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Ruymán Reyes |
ICPP | 6 |
| 2012 | Reducing Energy Consumption of Dense Linear Algebra Operations on Hybrid CPU-GPU PlatformsabstractWe investigate the balance between the time-to-solution and the energy consumption of a task-parallel execution of the Cholesky and LU factorizations on a hybrid platform, equipped with a multi-core processor and several GPUs. To improve energy efficiency, we incorporate two energy-saving techniques in the runtime in charge of scheduling the computations, to block idle threads and enable the transition to a more energy-friendly state of the general-purpose cores. Experiments on an Intel Xeon-based platform connected to an NVIDIA Tesla server report an average reduction of the energy consumption close to 9% (38% when only the consumption associated with the application is considered), for a minor increase in the execution time of the algorithm. Pedro Alonso 0002, Manuel F. Dolz, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
ISPA | 4 |
| 2012 | Binding Performance and Power of Dense Linear Algebra OperationsabstractIn this paper we combine a powerful tracing framework with a power measurement setup to perform a visual analysis of the computational performance and the power consumption of tuned implementations for three key dense linear algebra operations: the LU factorization, the Cholesky factorization, and the reduction to tridiagonal form. Our results using 6 and 12 cores of an AMD Opteron-based platform reveal the serial/concurrent phases of the algorithms, and their connection to periods of low/high power consumption, as well as the linear dependency between execution time and energy for this class of operations. Maria Barreda, Manuel F. Dolz, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Ruymán Reyes |
ISPA | 3 |
| 2012 | Saving Energy in the LU Factorization with Partial Pivoting on Multi-core ProcessorsabstractIn this paper we analyze the trade-off between energy and performance for a data-parallel execution of the LU factorization with partial pivoting on a multi-core processor. To improve energy efficiency, we adapt the runtime in charge of controlling the concurrent execution of the algorithm to leverage DVFS and block idle threads. For a CPU-bounded operation like the LU factorization, experiments on an AMD 8-core processor report a reduction around 5% in energy consumption for the largest problem sizes in exchange for a minor increase in the execution time. Pedro Alonso 0002, Manuel F. Dolz, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
PDP | 4 |
| 2012 | Analysis of Strategies to Save Energy for Message-Passing Dense Linear Algebra KernelsabstractIn this paper we analyze the impact that energy-saving strategies, like the application of DVFS via Linux governors and the MPI communication mode, have on the performance and energy consumption of message-passing dense linear algebra operations. In the study, we employ codes from ScaLAPACK for three matrix kernels, the matrix-matrix and matrix-vector products and the Cholesky factorization, which exhibit different levels of concurrency and CPU/memory activity. Following a recent trend, we also include an accelerated version of the matrix-matrix product that off-loads all computation to a graphics processor and study the energy gains of this hybrid solver when the general-purpose cores of the system are promoted to a low consuming mode. Experimental results on a cluster equipped with state-of-the-art computation and communication hardware illustrate the results of this study. María Isabel Castillo, Juan Carlos Fernández 0002, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Vicente Roca |
PDP | 3 |
| 2011 | Enabling CUDA acceleration within virtual machines using rCUDAabstractThe hardware and software advances of Graphics Processing Units (GPUs) have favored the development of GPGPU (General-Purpose Computation on GPUs) and its adoption in many scientific, engineering, and industrial areas. Thus, GPUs are increasingly being introduced in high-performance computing systems as well as in datacenters. On the other hand, virtualization technologies are also receiving rising interest in these domains, because of their many benefits on acquisition and maintenance savings. There are currently several works on GPU virtualization. However, there is no standard solution allowing access to GPGPU capabilities from virtual machine environments like, e.g., VMware, Xen, VirtualBox, or KVM. Such lack of a standard solution is delaying the integration of GPGPU into these domains. In this paper, we propose a first step towards a general and open source approach for using GPGPU features within VMs. In particular, we describe the use of rCUDA, a GPGPU (General-Purpose Computation on GPUs) virtualization framework, to permit the execution of GPU-accelerated applications within virtual machines (VMs), thus enabling GPGPU capabilities on any virtualized environment. Our experiments with rCUDA in the context of KVM and VirtualBox on a system equipped with two NVIDIA GeForce 9800 GX2 cards illustrate the overhead introduced by the rCUDA middleware and prove the feasibility and scalability of this general virtualizing solution. Experimental results show that the overhead is proportional to the dataset size, while the scalability is similar to that of the native environment. José Duato, Antonio J. Peña, Federico Silla, Juan Carlos Fernández 0002, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
HiPC | 5 |
| 2011 | Performance of CUDA Virtualized Remote GPUs in High Performance ClustersabstractIn a previous work we presented the architecture of rCUDA, a middleware that enables CUDA remoting over a commodity network. That is, the middleware allows an application to use a CUDA-compatible Graphics Processor (GPU) installed in a remote computer as if it were installed in the computer where the application is being executed. This approach is based on the observation that GPUs in a cluster are not usually fully utilized, and it is intended to reduce the number of GPUs in the cluster, thus lowering the costs related with acquisition and maintenance while keeping performance close to that of the fully-equipped configuration. In this paper we model rCUDA over a series of high throughput networks in order to assess the influence of the performance of the underlying network on the performance of our virtualization technique. For this purpose, we analyze the traces of two different case studies over two different networks. Using this data, we calculate the expected performance for these same case studies over a series of high throughput networks, in order to characterize the expected behavior of our solution in high performance clusters. The estimations are validated using real 1 Gbps Ethernet and 40 Gbps InfiniBand networks, showing an error rate in the order of 1% for executions involving data transfers above 40 MB. In summary, although our virtualization technique noticeably increases execution time when using a 1 Gbps Ethernet network, it performs almost as efficiently as a local GPU when higher performance interconnects are used. Therefore, the small overhead incurred by our proposal because of the remote use of GPUs is worth the savings that a cluster configuration with less GPUs than nodes reports. José Duato, Antonio J. Peña, Federico Silla, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
ICPP | 4 |
| 2009 | An Extension of the StarSs Programming Model for Platforms with Multiple GPUs
Eduard Ayguadé, Rosa M. Badia, Francisco D. Igual, Jesús Labarta, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
Euro-Par | 5 |
| 2009 | Exploiting the capabilities of modern GPUs for dense matrix computationsabstractAbstract We present several algorithms to compute the solution of a linear system of equations on a graphics processor (GPU), as well as general techniques to improve their performance, such as padding and hybrid GPU‐CPU computation. We compare single and double precision performance of a modern GPU with unified architecture, and show how iterative refinement with mixed precision can be used to regain full accuracy in the solution of linear systems, exploiting the potential of the processor for single precision arithmetic. Experimental results on a GTX280 using CUBLAS 2.0, the implementation of BLAS for NVIDIA® GPUs with unified architecture, illustrate the performance of the different algorithms and techniques proposed. Copyright © 2009 John Wiley & Sons, Ltd. Sergio Barrachina 0001, María Isabel Castillo, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí |
Concurr. Comput. Pract. Exp. | 4 |
| 2009 | Toward the parallelization of GSL
José Ignacio Aliaga, Francisco Almeida, José M. Badía, Sergio Barrachina 0001, Vicente Blanco 0001, María Isabel Castillo, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí, Alfredo Remón, Casiano Rodríguez, Francisco de Sande, Adrián Santos |
J. Supercomput. | 7 |
| 2008 | Solving Dense Linear Systems on Graphics Processors
Sergio Barrachina 0001, María Isabel Castillo, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
Euro-Par | 4 |
| 2008 | Biomedical image analysis on a cooperative cluster of GPUs and multicoresabstractWe are currently witnessing the emergence of two paradigms in parallel computing: streaming processing and multi-core CPUs. Represented by solid commercial products widely available in commodity PCs, GPUs and multi-core CPUs bring together an unprecedented combination of high performance at low cost. The scientific computing community needs to keep pace with application models and middleware which scale efficiently to hundreds of internal processing units. The purpose of the work we present here is twofold: first, a cooperative environment is designed so that both parallel models can coexist and complement one another. Second, beyond the parallelism of multiple internal cores, further parallelism is introduced when multiple CPU sockets, multiple GPUs, and multiple nodes are combined within a unique multi-processor platform which exceeds 10 TFLOPS when using 16 nodes. We illustrate our cooperative parallelization approach by implementing a large-scale, biomedical image analysis application which contains a number of assorted kernels including typical streaming operators, co-occurrence matrices, convolutions, and histograms. Experimental results are compared among different implementation strategies and almost linear speed-up is achieved when all coexisting methods in CPUs and GPUs are combined. Timothy D. R. Hartley, Ümit V. Çatalyürek, Antonio Ruiz 0001, Francisco D. Igual, Rafael Mayo 0002, Manuel Ujaldon |
ICS | 5 |
| 2008 | Evaluation and tuning of the Level 3 CUBLAS for graphics processorsabstractThe increase in performance of the last generations of graphics processors (GPUs) has made this class of platform a coprocessing tool with remarkable success in certain types of operations. In this paper we evaluate the performance of the Level 3 operations in CUBLAS, the implementation of BIAS for NVIDIAreg GPUs with unified architecture. From this study, we gain insights on the quality of the kernels in the library and we propose several alternative implementations that are competitive with those in CUBLAS. Experimental results on a GeForce 8800 Ultra compare the performance of CUBLAS and the new variants. Sergio Barrachina 0001, María Isabel Castillo, Francisco D. Igual, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
IPDPS | 4 |
| 2007 | Stabilizing large-scale generalized systems on parallel computers using multithreading and message-passingabstractAbstract We discuss the parallelization of an efficient algorithm for the partial stabilization of large‐scale linear control systems in generalized state‐space form. The algorithm is composed of highly parallel iterative schemes that appear in the computation of certain matrix functions. Here we evaluate different approaches to exploit parallelism at two levels, based on threads and processes. Our experimental results on a cluster of symmetric multiprocessors and a CC‐NUMA platform show that the efficiency of the matrix operations underlying the iterative schemes carry over to the parallel implementation of the stabilization algorithm. Copyright © 2006 John Wiley & Sons, Ltd. Peter Benner, María Isabel Castillo, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí |
Concurr. Comput. Pract. Exp. | 3 |
| 2006 | Parallel Solution of Large-Scale and Sparse Generalized Algebraic Riccati Equations
José M. Badía, Peter Benner, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
Euro-Par | 3 |
| 2006 | Parallelization of GSL: The Web Service InterfaceabstractWe present our joint effort to develop a Web based interface for the GNU Scientific library and its parallelization. The interface has been developed using standard Web services technology to enable the use of non local resources to execute parallel programs. The final result is a computing service where sequential and parallel routines demanding high performance computing are supplied. The design allows to incorporate new servers and platforms with a small number of software requirements. José Ignacio Aliaga, José M. Badía, Sergio Barrachina 0001, María Isabel Castillo, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí, Francisco Almeida, Vicente Blanco 0001, Casiano Rodríguez, Francisco de Sande, Adrián Santos |
PDP | 5 |
| 2005 | Parallel Order Reduction via Balanced Truncation for Optimal Cooling of Steel Profiles
José M. Badía, Peter Benner, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí, Jens Saak |
Euro-Par | 3 |
| 2002 | Solving Large Sparse Lyapunov Equations on Parallel Computers (Research Note)
José M. Badía, Peter Benner, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
Euro-Par | 3 |
| 2002 | Parallel Algorithms for LQ Optimal Control of Discrete-Time Periodic Linear Systems
Peter Benner, Ralph Byers, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Vicente Hernández |
J. Parallel Distributed Comput. | 3 |
| 2001 | Parallel solvers for discrete-time algebric Riccati equationsabstractAbstract We investigate the numerical solution of discrete‐time algebraic Riccati equations on a parallel distributed architecture. Our solvers obtain an initial solution of the Riccati equation via the disc function method, and then refine this solution using Newton's method. The Smith iteration is employed to solve the Stein equation that arises at each step of Newton's method. The numerical experiments on an Intel Pentium‐II cluster, connected via a Myrinet switch, report the performance and scalability of the new algorithms. Copyright © 2001 John Wiley & Sons, Ltd. Rafael Mayo 0002, Enrique S. Quintana-Ortí, Gregorio Quintana-Ortí, Vicente Hernández |
Concurr. Comput. Pract. Exp. | 1 |
| 2000 | Solving Discrete-Time Periodic Riccati Equations on a Cluster (Research Note)
Peter Benner, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Vicente Hernández |
Euro-Par | 2 |