EDBT 2026 Demo / reviewers in the wild / expert
Carlos Reaño
dblp:129/5476
· DBLP profile ↗
36ranked-venue papers
17as first author
14since 2021 · last 2026
0000-0001-7871-9152ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 28 · 16 first-author · 10 since 2021Computer networks · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing the performance of GPU acceleration in virtual environments: Thoroughly benchmarking the rigidity of mediated device passthroughabstractVirtualization has been the key element for the growth of cloud computing. Historically, GPUs have been complex devices to virtualize efficiently. The mediated passthrough mechanism is usually leveraged. However, it implies a rigid association between the virtual domain and the virtual GPU, which impairs overall system GPU performance. In this paper we propose the use of Network GPGPU system (NGS) to improve the performance of GPUs in virtual domains. Our proposal is compared to NVIDIA vGPU, the most widely used mechanism for virtualizing CUDA-enabled GPUs today. Results show throughput benefits of approximately 20%, speedup of up to 28% in the execution time of the applications, improved overall GPU utilization (over 85%), and lower energy consumption per job (up to 15.34%). Javier Prades, Carlos Reaño, Federico Silla |
Future Gener. Comput. Syst. | 2 |
| 2025 | SaaS-Enabled RGB to Hyperspectral Imaging: A Novel Paradigm in Image Processing Technology
Robert Williamson, Carlos Reaño, Jesús Martínez del Rincón, Anastasios Koidis |
AINA (4) | 2 |
| 2025 | Efficient Integer-Only-Inference of Gradient Boosting Decision Trees on Low-Power DevicesabstractThere is increasingly interest in developing embedded machine learning hardware as it can offer better performance in terms of privacy, bandwidth efficiency, and scalability. Gradient-boosted decision trees (GBDT) represent a strong candidate as they employ less complex logic, but their efficient implementation in field programmable gate array (FPGA) needs to be explored in detail. In this paper, we propose sophisticated quantisation approaches to balance the dual goals of efficiency and performance. In particular, we introduce quantisation-aware training of GBDT for integer-only and binary arithmetic. Results are presented for implementations on a Zynq UltraScale+ MPSoC FPGA with the best design using only 170 Look-up Tables and 233 flip-flops at a clock speed of 724 MHz. Implementations focused on network intrusion detection and jet substructure classification for large-scale physics experiments are explored. An order of magnitude less FPGA resources are used whilst offering extremely high throughput rate and maintaining accuracy. Code is available athttps://github.com/malsharari/QATGBDT. Majed Alsharari, Son T. Mai, Roger F. Woods, Carlos Reaño |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | Accelerator virtualizationabstractWelcome to this special issue on accelerator virtualization in Concurrency and Computation: Practice and Experience. Virtualization is a key technique developed for sharing the underlying physical resources of a computer, such as the processor or memory and the network. This enables effective use of resources by improving their utilization and reduces costs. A well-known example of virtualization are virtual machines that have become prominent with the advent of cloud technologies. Virtual machines are an abstraction of the underlying physical computer that can be made available to different users. The underpinning technology ensures data security by isolating the environment in which each user works. Creating and executing virtual machines requires both software and hardware support. It is thought that one of the earliest forms of virtualization was time-multiplexing a single processor for different applications. Although this significantly varies from the current notion of virtualization, processor multiplexing laid the groundwork for modern operating systems to facilitate the concurrent sharing of an expensive hardware resource among several applications as if they each used the resource exclusively. A network file system is another example of virtualization in which the file system is physically available to several nodes of a computer cluster while it is simultaneously accessed by different client nodes. In this case, storage is the common resource shared via multiplexing mechanisms. Recently, virtualization has been adopted for hardware accelerators, specialized hardware such as graphics processing units (GPUs), field programmable gate arrays (FPGAs), and tensor processing units (TPUs). Accelerators reduce the execution time of certain applications by allowing programmers to offload compute intensive components of their applications to accelerators. They are also known to improve energy efficiency. We hope that the contents of this special issue are useful to you and that you enjoy them as much as we did. Carlos Reaño, Federico Silla, Blesson Varghese |
Concurr. Comput. Pract. Exp. | 1 |
| 2024 | NGS: A network GPGPU system for orchestrating remote and virtual acceleratorsabstractIn General-Purpose computing on Graphics Processing Unit (GPGPU), the use of CPUs is combined with that of GPUs. CPUs are used for sequential code, while GPUs are used for parallel code. GPGPU has been enabled by two key factors: (i) the massively parallel architecture of GPUs, which allows thousands of single cores to run parallel code; and (ii) the development of platforms, such as CUDA, that simplify implementing code for GPUs. GPGPU has established itself as the standard computing system in most computing fields due to the great improvements it brings. However, its use is not without problems, such as GPU underutilization, high cost, power consumption, etc. In this paper we present NGS (Network GPGPU System) to address the underutilization of GPUs in computing centers. NGS orchestrates the concurrent access to GPGPU resources from different nodes of the cluster by leveraging the remote GPU virtualization mechanism and the NVML library by NVIDIA. In this way, NGS enables different nodes of the cluster to access remote GPUs as if they were local at the same time that this access is guaranteed to be carried out without collisions. The main novelty is that NGS offers a global and standard solution independent of the computing environment used. Experimental results show up to 4x improvements compared to popular approaches. Javier Prades, Carlos Reaño, Federico Silla |
J. Syst. Archit. | 2 |
| 2024 | Accelerating the detection of DNA differentially methylated regions using multiple GPUsabstractAbstract DNA methylation analysis has become an important topic in the study of human health. In previous work, we developed a suite of tools to perform this analysis. It includes HPG-Dhunter, a web-based tool for automatic detection of differentially methylated regions (DMRs) between different samples. The back-end of that tool receives an undefined number of simultaneous requests to detect DMRs on different datasets. Currently, simultaneous requests are queued and processed one at a time. This paper proposes a parallel architecture where multiple daemons serve requests simultaneously. Daemons can also share the same physical GPUs. A scheduler manages requests and forwards them to daemons. The number of daemons per GPU is configurable, thus adapting the architecture to the available hardware. Results show that the proposed parallel architecture hugely reduces the execution time. Furthermore, the speedup increases proportionally to the number of available GPUs (up to 7.47x in our experimental setup). Carlos Reaño, Ricardo Olanda, Elvira Baydal, Mariano Pérez, Juan M. Orduña |
J. Supercomput. | 1 |
| 2023 | RGB-2-Hyper-Spectral Image Reconstruction for Food Science Using Encoder/Decoder Neural ArchitecturesabstractHyper-spectral imaging captures spatial and spectral information of a subject. This is used for the identification of substances within a scene, and food analysis. Presented is an investigation into the capabilities of encoder/decoder deep learning architectures for hyper-spectral image reconstruction from RGB images. For this analysis state-of-the-art (SOTA) techniques for hyper-spectral image reconstruction and other architectures from different fields have been used. Our approach examines a food science case study, using a CPU-based server and different accelerators. An in-house multi-sensor setup was used to capture the dataset which contains hyper-spectral images of twenty slices of different Spanish Ham in the range of 400-100∼nm and their analogous RGB images. The results show no degradation in the output when moving outside of the visual range. This study shows that the SOTA methods for reconstructing from RGB do not produce the most accurate reconstruction of the spectral domain within the range of 400-1000∼nm. Robert Williamson, Jesús Martínez del Rincón, Anastasios Koidis, Carlos Reaño |
ISCC | 4 |
| 2023 | Exploring the use of data compression for accelerating machine learning in the edge with remote virtual graphics processing unitsabstractSummary Internet of Things (IoT) devices are usually low performance nodes connected by low bandwidth networks. To improve performance in such scenarios, some computations could be done at the edge of the network. However, edge devices may not have enough computing power to accelerate applications such as the popular machine learning ones. Using remote virtual graphics processing units (GPUs) can address this concern by accelerating applications leveraging a GPU installed in a remote device. However, this requires exchanging data with the remote GPU across the slow network. To address the problem with the slow network, the data to be exchanged with the remote GPU could be compressed. In this article, we explore the suitability of using data compression in the context of remote GPU virtualization frameworks in edge scenarios executing machine learning applications. We use popular machine learning applications to carry out such exploration. After characterizing the GPU data transfers of these applications, we analyze the usage of existing compression libraries for compressing those data transfers to/from the remote GPU. Our exploration shows that transferring compressed data becomes more beneficial as networks get slower, reducing transfer time by up to 10 times. Our analysis also reveals that efficient integration of compression into remote GPU virtualization frameworks is strongly required. Cristian Peñaranda, Carlos Reaño, Federico Silla |
Concurr. Comput. Pract. Exp. | 2 |
| 2023 | Multi-Tier GPU Virtualization for Deep Learning in Cloud-Edge SystemsabstractAccelerator virtualization offers several advantages in the context of cloud-edge computing. Relatively weak user devices can enhance performance when running workloads by accessing virtualized accelerators available on other resources in the cloud-edge continuum. However, cloud-edge systems are heterogeneous, often leading to compatibility issues arising from various hardware and software stacks present in the system. One mechanism to alleviate this issue is using containers for deploying workloads. Containers isolate applications and their dependencies and store them as images that can run on any device. In addition, user devices may move during the course of application execution, and thus mechanisms such as container migration are required to move running workloads from one resource to another in the network. Furthermore, an optimal destination will need to be determined when migrating between virtual accelerators. Scheduling and placement strategies are incorporated to choose the best possible location depending on the workload requirements. This paper presentsAVEC, a framework for accelerator virtualization in cloud-edge computing. The AVEC framework enables the offloading of deep learning workloads for inference from weak user devices to computationally more powerful devices in a cloud-edge network. AVEC incorporates a mechanism that efficiently manages and schedules the virtualization of accelerators. It also supports migration between accelerators to enable stateless container migration. The experimental analysis highlights that AVEC can achieve up to 7x speedup by offloading applications to remote resources. Furthermore, AVEC features a low migration downtime that is less than 5 seconds. Jason Kennedy, Vishal Sharma 0001, Blesson Varghese, Carlos Reaño |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2021 | AVEC: Accelerator Virtualization in Cloud-Edge Computing for Deep Learning LibrariesabstractEdge computing offers the distinct advantage of harnessing compute capabilities on resources located at the edge of the network to run workloads of relatively weak user devices. This is achieved by offloading computationally intensive workloads, such as deep learning from user devices to the edge. Using the edge reduces the overall communication latency of applications as workloads can be processed closer to where data is generated on user devices rather than sending them to geographically distant clouds. Specialised hardware accelerators, such as Graphics Processing Units (GPUs) available in the cloud-edge network can enhance the performance of computationally intensive workloads that are offloaded from devices on to the edge. The underlying approach required to facilitate this is virtualization of GPUs. This paper therefore sets out to investigate the potential of GPU accelerator virtualization to improve the performance of deep learning workloads in a cloud-edge environment. The AVEC accelerator virtualization framework is proposed that incurs minimum overheads and requires no source-code modification of the workload. AVEC intercepts local calls to a GPU on a device and forwards them to an edge resource seamlessly. The feasibility of AVEC is demonstrated on a real-world application, namely OpenPose using the Caffe deep learning library. It is observed that on a lab-based experimental test-bed AVEC delivers up to 7.48x speedup despite communication overheads incurred due to data transfers. Jason Kennedy, Blesson Varghese, Carlos Reaño |
ICFEC | 3 |
| 2021 | GPU-Accelerated Discrete Event Simulations: Towards Industry 4.0 ManufacturingabstractDiscrete Event Simulations (DES) are the most commonplace tools for modelling today's manufacturing factories and their processes. DES are becoming steadfastly integrated into their corresponding physical counterparts to administer greater avenues for their analysis, control, forecasts and optimisations in real-time. However, this growth does not materialise without a penalty in the form of computational burden. The demand for flexible and alternate approaches to accelerate DES is made necessary. Hence, the utilisation of GPUs to comply with such acceleration presents a research topic of growing interest. This work investigates the use of the Machine Learning platform TensorFlow with GPUs to accelerate a variety of manufacturing-domain DES using the SimPy simulation framework. A range of results were gathered, of speed-ups spanning between x1.4 and x3.21, paving the way for further enhancements towards the vision of real-time communication between simulation and physical system in the form of a complete Digital Twin. Moustafa Faheem, Adrian Murphy, Carlos Reaño |
ISCC | 3 |
| 2021 | Improving the management efficiency of GPU workloads in data centers through GPU virtualizationabstractSummary Graphics processing units (GPUs) are currently used in data centers to reduce the execution time of compute‐intensive applications. However, the use of GPUs presents several side effects, such as increased acquisition costs and larger space requirements. Furthermore, GPUs require a nonnegligible amount of energy even while idle. Additionally, GPU utilization is usually low for most applications. In a similar way to the use of virtual machines, using virtual GPUs may address the concerns associated with the use of these devices. In this regard, the remote GPU virtualization mechanism could be leveraged to share the GPUs present in the computing facility among the nodes of the cluster. This would increase overall GPU utilization, thus reducing the negative impact of the increased costs mentioned before. Reducing the amount of GPUs installed in the cluster could also be possible. However, in the same way as job schedulers map GPU resources to applications, virtual GPUs should also be scheduled before job execution. Nevertheless, current job schedulers are not able to deal with virtual GPUs. In this paper, we analyze the performance attained by a cluster using the remote Compute Unified Device Architecture middleware and a modified version of the Slurm scheduler, which is now able to assign remote GPUs to jobs. Results show that cluster throughput, measured as jobs completed per time unit, is doubled at the same time that the total energy consumption is reduced up to 40%. GPU utilization is also increased. Sergio Iserte, Javier Prades, Carlos Reaño, Federico Silla |
Concurr. Comput. Pract. Exp. | 3 |
| 2021 | Redesigning the rCUDA communication layer for a better adaptation to the underlying hardwareabstractSummary The use of Graphics Processing Units (GPUs) has become a very popular way to accelerate the execution of many applications. However, GPUs are not exempt from side effects. For instance, GPUs are expensive devices which additionally consume a non‐negligible amount of energy even when they are not performing any computation. Furthermore, most applications present low GPU utilization. To address these concerns, the use of GPU virtualization has been proposed. In particular, remote GPU virtualization is a promising technology that allows applications to transparently leverage GPUs installed in any node of the cluster. In this paper, the remote GPU virtualization mechanism is comparatively analyzed across three different generations of GPUs. The first contribution of this study is an analysis about how the performance of the remote GPU virtualization technique is impacted by the underlying hardware. To that end, the Tesla K20, Tesla K40, and Tesla P100 GPUs along with FDR and EDR InfiniBand fabrics are used in the study. The analysis is performed in the context of the rCUDA middleware. It is clearly shown that the GPU virtualization middleware requires a comprehensive design of its communication layer, which should be perfectly adapted to every hardware generation in order to avoid a reduction in performance. This is precisely the second contribution of this work, ie, redesigning the rCUDA communication layer in order to improve the management of the underlying hardware. Results show that it is possible to improve bandwidth up to 29.43%, which translates into up to 4.81% average less execution time in the performance of the analyzed applications. Carlos Reaño, Federico Silla |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | Editorial
Carlos Reaño, Federico Silla, Blesson Varghese |
J. Parallel Distributed Comput. | 1 |
| 2020 | Improving the performance of physics applications in atom-based clusters with rCUDA
Federico Silla, Javier Prades, Elvira Baydal, Carlos Reaño |
J. Parallel Distributed Comput. | 4 |
| 2019 | Analyzing the performance/power tradeoff of the rCUDA middleware for future exascale systems
Carlos Reaño, Javier Prades, Federico Silla |
J. Parallel Distributed Comput. | 1 |
| 2019 | On the support of inter-node P2P GPU memory copies in rCUDA
Carlos Reaño, Federico Silla |
J. Parallel Distributed Comput. | 1 |
| 2018 | Intra-Node Memory Safe GPU Co-SchedulingabstractGPUs in High-Performance Computing systems remain under-utilised due to the unavailability of schedulers that can safely schedule multiple applications to share the same GPU. The research reported in this paper is motivated to improve the utilisation of GPUs by proposing a framework, we refer to as schedGPU, to facilitate intra-node GPU co-scheduling such that a GPU can be safely shared among multiple applications by taking memory constraints into account. Two approaches, namely a client-server and a shared memory approach are explored. However, the shared memory approach is more suitable due to lower overheads when compared to the former approach. Four policies are proposed in schedGPU to handle applications that are waiting to access the GPU, two of which account for priorities. The feasibility of schedGPU is validated on three real-world applications. The key observation is that a performance gain is achieved. For single applications, a gain of over 10 times, as measured by GPU utilisation and GPU memory utilisation, is obtained. For workloads comprising multiple applications, a speed-up of up to 5x in the total execution time is noted. Moreover, the average GPU utilisation and average GPU memory utilisation is increased by 5 and 12 times, respectively. Carlos Reaño, Federico Silla, Dimitrios S. Nikolopoulos, Blesson Varghese |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | Enhancing the rCUDA Remote GPU Virtualization Framework: from a Prototype to a Production SolutionabstractThe use of hardware accelerators to increase the performance of parallel applications is very common nowadays. For a number of reasons, however, the access to local accelerators is not always feasible (e.g., lack of space or cost). It would also be the case that some applications benefit from having access to more accelerators than the physically possible. To address all these concerns, middleware offering access not only to local but also to remote accelerators appeared. This paper presents a high-level summary of a dissertation focused on enhancing one of these middleware, called rCUDA. Carlos Reaño, Federico Silla, José Duato |
CCGrid | 1 |
| 2017 | On the benefits of the remote GPU virtualization mechanism: The rCUDA caseabstractSummary Graphics processing units (GPUs) are being adopted in many computing facilities given their extraordinary computing power, which makes it possible to accelerate many general purpose applications from different domains. However, GPUs also present several side effects, such as increased acquisition costs as well as larger space requirements. They also require more powerful energy supplies. Furthermore, GPUs still consume some amount of energy while idle, and their utilization is usually low for most workloads. In a similar way to virtual machines, the use of virtual GPUs may address the aforementioned concerns. In this regard, the remote GPU virtualization mechanism allows an application being executed in a node of the cluster to transparently use the GPUs installed at other nodes. Moreover, this technique allows to share the GPUs present in the computing facility among the applications being executed in the cluster. In this way, several applications being executed in different (or the same) cluster nodes can share 1 or more GPUs located in other nodes of the cluster. Sharing GPUs should increase overall GPU utilization, thus reducing the negative impact of the side effects mentioned before. Reducing the total amount of GPUs installed in the cluster may also be possible. In this paper, we explore some of the benefits that remote GPU virtualization brings to clusters. For instance, this mechanism allows an application to use all the GPUs present in the computing facility. Another benefit of this technique is that cluster throughput, measured as jobs completed per time unit, is noticeably increased when this technique is used. In this regard, cluster throughput can be doubled for some workloads. Furthermore, in addition to increase overall GPU utilization, total energy consumption can be reduced up to 40%. This may be key in the context of exascale computing facilities, which present an important energy constraint. Other benefits are related to the cloud computing domain, where a GPU can be easily shared among several virtual machines. Finally, GPU migration (and therefore server consolidation) is one more benefit of this novel technique. Federico Silla, Sergio Iserte, Carlos Reaño, Javier Prades |
Concurr. Comput. Pract. Exp. | 3 |
| 2017 | Multi-tenant virtual GPUs for optimising performance of a financial risk application
Javier Prades, Blesson Varghese, Carlos Reaño, Federico Silla |
J. Parallel Distributed Comput. | 3 |
| 2016 | Increasing the Performance of Data Centers by Combining Remote GPU Virtualization with SlurmabstractThe use of Graphics Processing Units (GPUs) presents several side effects, such as increased acquisition costs as well as larger space requirements. Furthermore, GPUs require a non-negligible amount of energy even while idle. Additionally, GPU utilization is usually low for most applications. Using the virtual GPUs provided by the remote GPU virtualization mechanism may address the concerns associated with the use of these devices. However, in the same way as workload managers map GPU resources to applications, virtual GPUs should also be scheduled before job execution. Nevertheless, current workload managers are not able to deal with virtual GPUs. In this paper we analyze the performance attained by a cluster using the rCUDA remote GPU virtualization middleware and a modified version of the Slurm workload manager, which is now able to map remote virtual GPUs to jobs. Results show that cluster throughput is doubled at the same time that total energy consumption is reduced up to 40%. GPU utilization is also increased. Sergio Iserte, Javier Prades, Carlos Reaño, Federico Silla |
CCGrid | 3 |
| 2016 | Providing CUDA Acceleration to KVM Virtual Machines in InfiniBand Clusters with rCUDAabstractThere is a trend towards using graphics processing units (GPUs) not only for graphics visualization, but also for accelerating scientific applications. But their use for this purpose is not without disadvantages: GPUs increase costs and energy consumption. Furthermore, GPUs are generally underutilized. Using virtual machines could be a possible solution to address these problems, however, current solutions for providing GPU acceleration to virtual machines environments, such as KVM or Xen, present some issues. In this paper we propose the use of remote GPUs to accelerate scientific applications running inside KVM virtual machines. Our analysis shows that this approach could be a possible solution, with low overhead when used over InfiniBand networks. Ferran Perez, Carlos Reaño, Federico Silla |
DAIS | 2 |
| 2016 | Reducing the performance gap of remote GPU virtualization with InfiniBand Connect-IBabstractGPU accelerators may provide great performance improvements in the context of parallel applications. However, their use in HPC clusters may also present some disadvantages such as their high cost and high power consumption. In addition, this kind of accelerators are generally underutilized. Remote GPU virtualization could be a solution to overcome these drawbacks, but its performance is usually impaired because the network bandwidth is lower than the PCIe one. In this paper we analyze how the InfiniBand Connect-IB network adapters (with performance similar to that of PCIe 3.0) reduce the overhead of remote GPU virtualization. We show that this overhead is decreased to 1.5% in terms of bandwidth, and to 0.51% in the tested application. Carlos Reaño, Federico Silla |
ISCC | 1 |
| 2016 | CUDA acceleration for Xen virtual machines in infiniband clusters with rCUDAabstractMany data centers currently use virtual machines (VMs) to achieve a more efficient usage of hardware resources. However, current virtualization solutions, such as Xen, do not easily provide graphics processing unit (GPU) accelerators to applications running in the virtualized domain with the flexibility usually required in data centers (i.e., managing virtual GPU instances and concurrently sharing them among several VMs). Remote GPU virtualization frameworks such as the rCUDA solution may address this problem. Javier Prades, Carlos Reaño, Federico Silla |
PPoPP | 2 |
| 2016 | Tuning remote GPU virtualization for InfiniBand networks
Carlos Reaño, Federico Silla |
J. Supercomput. | 1 |
| 2015 | On the Design of a Demo for Exhibiting rCUDAabstractCUDA is a technology developed by NVIDIA which provides a parallel computing platform and programming model for NVIDIA GPUs and compatible ones. It takes benefit from the enormous parallel processing power of GPUs in order to accelerate a wide range of applications, thus reducing their execution time. rCUDA (remote CUDA) is a middleware which grants applications concurrent access to CUDA-compatible devices installed in other nodes of the cluster in a transparent way so that applications are not aware of being accessing a remote device. In this paper we present a demo which shows, in real time, the overhead introduced by rCUDA in comparison to CUDA when running image filtering applications. The approach followed in this work is to develop a graphical demo which contains both an appealing design and technical contents. Carlos Reaño, Ferran Perez, Federico Silla |
CCGRID | 1 |
| 2015 | A Performance Comparison of CUDA Remote GPU Virtualization FrameworksabstractUsing GPUs reduces execution time of many applications but increases acquisition cost and power consumption. Furthermore, GPUs usually attain a relatively low utilization. In this context, remote GPU virtualization solutions were recently created to overcome the drawbacks of using GPUs. Currently, many different remote GPU virtualization frameworks exist, all of them presenting very different characteristics. These differences among them may lead to differences in performance. In this work we present a performance comparison among the only three CUDA remote GPU virtualization frameworks publicly available at no cost. Results show that performance greatly depends on the exact framework used, being the rCUDA virtualization solution the one that stands out among them. Furthermore, rCUDA doubles performance over CUDA for pageable memory copies. Carlos Reaño, Federico Silla |
CLUSTER | 1 |
| 2015 | InfiniBand Verbs Optimizations for Remote GPU VirtualizationabstractThe use of InfiniBand networks to interconnect high performance computing clusters has considerably increased during the last years. So much so that the majority of the supercomputers included in the TOP500 list either use Ethernet or InfiniBand interconnects. Regarding the latter, due to the complexity of the InfiniBand programming API (i.e., InfiniBand Verbs) and the lack of documentation, there are not enough recent available studies explaining how to optimize applications to get the maximum performance from this fabric. In this paper we expose two different optimizations to be used when developing applications using InfiniBand Verbs, each providing an average bandwidth improvement of 3.68% and 217.14%, respectively. In addition, we show that when combining both optimizations, the average bandwidth gain is 43.29%. This bandwidth increment is key for remote GPU virtualization frameworks. Actually, this noticeable gain translates into a reduction of up to 35% in execution time of applications using remote GPU virtualization frameworks. Carlos Reaño, Federico Silla |
CLUSTER | 1 |
| 2015 | Acceleration-as-a-Service: Exploiting Virtualised GPUs for a Financial ApplicationabstractHow can GPU acceleration be obtained as a service in a cluster? This question has become increasingly significant due to the inefficiency of installing GPUs on all nodes of a cluster. The research reported in this paper is motivated to address the above question by employing rCUDA (remote CUDA), a framework that facilitates Acceleration-as-a-Service (AaaS), such that the nodes of a cluster can request the acceleration of a set of remote GPUs on demand. The rCUDA framework exploits virtualisation and ensures that multiple nodes can share the same GPU. In this paper we test the feasibility of the rCUDA framework on a real-world application employed in the financial risk industry that can benefit from AaaS in the production setting. The results confirm the feasibility of rCUDA and highlight that rCUDA achieves similar performance compared to CUDA, provides consistent results, and more importantly, allows for a single application to benefit from all the GPUs available in the cluster without loosing efficiency. Blesson Varghese, Javier Prades, Carlos Reaño, Federico Silla |
e-Science | 3 |
| 2015 | Improving the user experience of the rCUDA remote GPU virtualization frameworkabstractSummary Graphics processing units (GPUs) are being increasingly embraced by the high‐performance computing community as an effective way to reduce execution time by accelerating parts of their applications. remote CUDA (rCUDA) was recently introduced as a software solution to address the high acquisition costs and energy consumption of GPUs that constrain further adoption of this technology. Specifically, rCUDA is a middleware that allows a reduced number of GPUs to be transparently shared among the nodes in a cluster. Although the initial prototype versions of rCUDA demonstrated its functionality, they also revealed concerns with respect to usability, performance, and support for new CUDA features. In response, in this paper, we present a new rCUDA version that (1) improves usability by including a new component that allows an automatic transformation of any CUDA source code so that it conforms to the needs of the rCUDA framework, (2) consistently features low overhead when using remote GPUs thanks to an improved new communication architecture, and (3) supports multithreaded applications and CUDA libraries. As a result, for any CUDA‐compatible program, rCUDA now allows the use of remote GPUs within a cluster with low overhead, so that a single application running in one node can use all GPUs available across the cluster, thereby extending the single‐node capability of CUDA. Copyright © 2014 John Wiley & Sons, Ltd. Carlos Reaño, Federico Silla, Adrián Castelló 0001, Antonio J. Peña, Rafael Mayo 0002, Enrique S. Quintana-Ortí, José Duato |
Concurr. Comput. Pract. Exp. | 1 |
| 2014 | Boosting the performance of remote GPU virtualization using InfiniBand connect-IB and PCIe 3.0abstractA clear trend has emerged involving the acceleration of scientific applications by using GPUs. However, the capabilities of these devices are still generally underutilized. Remote GPU virtualization techniques can help increase GPU utilization rates, while reducing acquisition and maintenance costs. The overhead of using a remote GPU instead of a local one is introduced mainly by the difference in performance between the internode network and the intranode PCIe link. In this paper we show how using the new InfiniBand Connect-IB network adapters (attaining similar throughput to that of the most recently emerged GPUs) boosts the performance of remote GPU virtualization, reducing the overhead to a mere 0.19% in the application tested. Carlos Reaño, Federico Silla, Antonio J. Peña, Gilad Shainer, Scot Schultz, Adrián Castelló 0001, Enrique S. Quintana-Ortí, José Duato |
CLUSTER | 1 |
| 2014 | SLURM Support for Remote GPU Virtualization: Implementation and Performance StudyabstractSLURM is a resource manager that can be leveraged to share a collection of heterogeneous resources among the jobs in execution in a cluster. However, SLURM is not designed to handle resources such as graphics processing units (GPUs). Concretely, although SLURM can use a generic resource plugin (GRes) to manage GPUs, with this solution the hardware accelerators can only be accessed by the job that is in execution on the node to which the GPU is attached. This is a serious constraint for remote GPU virtualization technologies, which aim at providing a user-transparent access to all GPUs in cluster, independently of the specific location of the node where the application is running with respect to the GPU node. In this work we introduce a new type of device in SLURM, "rgpu", in order to gain access from any application node to any GPU node in the cluster using rCUDA as the remote GPU virtualization solution. With this new scheduling mechanism, a user can access any number of GPUs, as SLURM schedules the tasks taking into account all the graphics accelerators available in the complete cluster. We present experimental results that show the benefits of this new approach in terms of increased flexibility for the job scheduler. Sergio Iserte, Adrián Castelló 0001, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Federico Silla, José Duato, Carlos Reaño, Javier Prades |
SBAC-PAD | 7 |
| 2014 | A complete and efficient CUDA-sharing solution for HPC clusters
Antonio J. Peña, Carlos Reaño, Federico Silla, Rafael Mayo 0002, Enrique S. Quintana-Ortí, José Duato |
Parallel Comput. | 2 |
| 2013 | Influence of InfiniBand FDR on the performance of remote GPU virtualizationabstractThe use of GPUs to accelerate general-purpose scientific and engineering applications is mainstream today, but their adoption in current high-performance computing clusters is impaired primarily by acquisition costs and power consumption. Therefore, the benefits of sharing a reduced number of GPUs among all the nodes of a cluster can be remarkable for many applications. This approach, usually referred to as remote GPU virtualization, aims at reducing the number of GPUs present in a cluster, while increasing their utilization rate. The performance of the interconnection network is key to achieving reasonable performance results by means of remote GPU virtualization. To this end, several networking technologies with throughput comparable to that of PCI Express have appeared recently. In this paper we analyze the influence of InfiniBand FDR on the performance of remote GPU virtualization, comparing its impact on a variety of GPU-accelerated applications with other networking technologies, such as Infini-Band QDR and Gigabit Ethernet. Given the severe limitations of freely available remote GPU virtualization solutions, the rCUDA framework is used as the case study for this analysis. Results show that the new FDR interconnect, featuring higher bandwidth than its predecessors, allows the reduction of the overhead of using GPUs remotely, thus making this approach even more appealing. Carlos Reaño, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Federico Silla, José Duato, Antonio J. Peña |
CLUSTER | 1 |
| 2012 | CU2rCU: Towards the complete rCUDA remote GPU virtualization and sharing solutionabstractGPUs are being increasingly embraced by the high performance computing and computational communities as an effective way of considerably reducing execution time by accelerating significant parts of their application codes. However, despite their extraordinary computing capabilities, the adoption of GPUs in current HPC clusters may present certain negative side-effects. In particular, to ease job scheduling in these platforms, a GPU is usually attached to every node of the cluster. In addition to increasing acquisition costs this favors that GPUs may frequently remain idle, as applications usually do not fully utilize them. On the other hand, idle GPUs consume non-negligible amounts of energy, which translates into very poor energy efficiency during idle cycles. rCUDA was recently developed as a software solution to address these concerns. Specifically, it is a middleware that allows transparently sharing a reduced number of GPUs among the nodes in a cluster. rCUDA thus increases the GPU-utilization rate, taking care of job scheduling. While the initial prototype versions of rCUDA demonstrated its functionality, they also revealed several concerns related with usability and performance. With respect to usability, in this paper we present a new component of the rCUDA suite that allows an automatic transformation of any CUDA source code, so that it can be effectively accommodated within this technology. In response to performance, we briefly show some interesting results, which will be deeply analyzed in future publications. The net outcome is a new version of rCUDA that allows, for any CUDA-compatible program, to use remote GPUs in a cluster with minimum overhead. Carlos Reaño, Antonio J. Peña, Federico Silla, José Duato, Rafael Mayo 0002, Enrique S. Quintana-Ortí |
HiPC | 1 |