Federico Silla

dblp:75/3178 · DBLP profile ↗
← Back
76ranked-venue papers
9as first author
9since 2021 · last 2026
0000-0002-6435-1200ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 68 · 9 first-author · 8 since 2021Software engineering, systems software and programming languages · 3Computer networks · 2 · 1 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Enhancing the performance of GPU acceleration in virtual environments: Thoroughly benchmarking the rigidity of mediated device passthrough
abstract
Virtualization has been the key element for the growth of cloud computing. Historically, GPUs have been complex devices to virtualize efficiently. The mediated passthrough mechanism is usually leveraged. However, it implies a rigid association between the virtual domain and the virtual GPU, which impairs overall system GPU performance. In this paper we propose the use of Network GPGPU system (NGS) to improve the performance of GPUs in virtual domains. Our proposal is compared to NVIDIA vGPU, the most widely used mechanism for virtualizing CUDA-enabled GPUs today. Results show throughput benefits of approximately 20%, speedup of up to 28% in the execution time of the applications, improved overall GPU utilization (over 85%), and lower energy consumption per job (up to 15.34%).
Javier Prades, Carlos Reaño, Federico Silla
Future Gener. Comput. Syst.3
2024 Accelerator virtualization
abstract
Welcome to this special issue on accelerator virtualization in Concurrency and Computation: Practice and Experience. Virtualization is a key technique developed for sharing the underlying physical resources of a computer, such as the processor or memory and the network. This enables effective use of resources by improving their utilization and reduces costs. A well-known example of virtualization are virtual machines that have become prominent with the advent of cloud technologies. Virtual machines are an abstraction of the underlying physical computer that can be made available to different users. The underpinning technology ensures data security by isolating the environment in which each user works. Creating and executing virtual machines requires both software and hardware support. It is thought that one of the earliest forms of virtualization was time-multiplexing a single processor for different applications. Although this significantly varies from the current notion of virtualization, processor multiplexing laid the groundwork for modern operating systems to facilitate the concurrent sharing of an expensive hardware resource among several applications as if they each used the resource exclusively. A network file system is another example of virtualization in which the file system is physically available to several nodes of a computer cluster while it is simultaneously accessed by different client nodes. In this case, storage is the common resource shared via multiplexing mechanisms. Recently, virtualization has been adopted for hardware accelerators, specialized hardware such as graphics processing units (GPUs), field programmable gate arrays (FPGAs), and tensor processing units (TPUs). Accelerators reduce the execution time of certain applications by allowing programmers to offload compute intensive components of their applications to accelerators. They are also known to improve energy efficiency. We hope that the contents of this special issue are useful to you and that you enjoy them as much as we did.
Carlos Reaño, Federico Silla, Blesson Varghese
Concurr. Comput. Pract. Exp.2
2024 NGS: A network GPGPU system for orchestrating remote and virtual accelerators
abstract
In General-Purpose computing on Graphics Processing Unit (GPGPU), the use of CPUs is combined with that of GPUs. CPUs are used for sequential code, while GPUs are used for parallel code. GPGPU has been enabled by two key factors: (i) the massively parallel architecture of GPUs, which allows thousands of single cores to run parallel code; and (ii) the development of platforms, such as CUDA, that simplify implementing code for GPUs. GPGPU has established itself as the standard computing system in most computing fields due to the great improvements it brings. However, its use is not without problems, such as GPU underutilization, high cost, power consumption, etc. In this paper we present NGS (Network GPGPU System) to address the underutilization of GPUs in computing centers. NGS orchestrates the concurrent access to GPGPU resources from different nodes of the cluster by leveraging the remote GPU virtualization mechanism and the NVML library by NVIDIA. In this way, NGS enables different nodes of the cluster to access remote GPUs as if they were local at the same time that this access is guaranteed to be carried out without collisions. The main novelty is that NGS offers a global and standard solution independent of the computing environment used. Experimental results show up to 4x improvements compared to popular approaches.
Javier Prades, Carlos Reaño, Federico Silla
J. Syst. Archit.3
2023 Exploring the use of data compression for accelerating machine learning in the edge with remote virtual graphics processing units
abstract
Summary Internet of Things (IoT) devices are usually low performance nodes connected by low bandwidth networks. To improve performance in such scenarios, some computations could be done at the edge of the network. However, edge devices may not have enough computing power to accelerate applications such as the popular machine learning ones. Using remote virtual graphics processing units (GPUs) can address this concern by accelerating applications leveraging a GPU installed in a remote device. However, this requires exchanging data with the remote GPU across the slow network. To address the problem with the slow network, the data to be exchanged with the remote GPU could be compressed. In this article, we explore the suitability of using data compression in the context of remote GPU virtualization frameworks in edge scenarios executing machine learning applications. We use popular machine learning applications to carry out such exploration. After characterizing the GPU data transfers of these applications, we analyze the usage of existing compression libraries for compressing those data transfers to/from the remote GPU. Our exploration shows that transferring compressed data becomes more beneficial as networks get slower, reducing transfer time by up to 10 times. Our analysis also reveals that efficient integration of compression into remote GPU virtualization frameworks is strongly required.
Cristian Peñaranda, Carlos Reaño, Federico Silla
Concurr. Comput. Pract. Exp.3
2023 Using remote GPU virtualization techniques to enhance edge computing devices
José M. Cecilia, Juan Morales-García, Baldomero Imbernon, Javier Prades, Juan-Carlos Cano, Federico Silla
Future Gener. Comput. Syst.6
2022 AI-Enabled Autonomous Drones for Fast Climate Change Crisis Assessment
abstract
Climate change is one of the greatest challenges for modern societies. Its consequences, often associated with extreme events, have dramatic results worldwide. New synergies between different disciplines, including artificial intelligence (AI), Internet of Things (IoT), and edge computing can lead to radically new approaches for the real-time tracking of natural disasters that are also designed to reduce the environmental footprint. In this article, we propose an AI-based pipeline for processing natural disaster images taken from drones. The purpose of this pipeline is to reduce the number of images to be processed by the first responders of the natural disaster. It consists of three main stages: 1) a lightweight autoencoder based on deep learning; 2) a dimensionality reduction using the$t$-distributed stochastic neighbor embedding algorithm; and 3) a fuzzy clustering procedure. This pipeline is evaluated on several edge computing platforms with low-power accelerators to assess the design of intelligent autonomous drones to provide this service in real time. Our experimental evaluation focuses on flooding, showing that the amount of information to be processed is substantially reduced, whereas edge computing platforms with low-power graphics accelerators are placed as a compelling alternative for processing these heavy computational workloads, obtaining a performance loss of only$2.3\times $compared to its cloud counterpart version, running both the training and inference steps.
Daniel Hernández 0009, Juan-Carlos Cano, Federico Silla, Carlos T. Calafate, José M. Cecilia
IEEE Internet Things J.3
2021 Improving the management efficiency of GPU workloads in data centers through GPU virtualization
abstract
Summary Graphics processing units (GPUs) are currently used in data centers to reduce the execution time of compute‐intensive applications. However, the use of GPUs presents several side effects, such as increased acquisition costs and larger space requirements. Furthermore, GPUs require a nonnegligible amount of energy even while idle. Additionally, GPU utilization is usually low for most applications. In a similar way to the use of virtual machines, using virtual GPUs may address the concerns associated with the use of these devices. In this regard, the remote GPU virtualization mechanism could be leveraged to share the GPUs present in the computing facility among the nodes of the cluster. This would increase overall GPU utilization, thus reducing the negative impact of the increased costs mentioned before. Reducing the amount of GPUs installed in the cluster could also be possible. However, in the same way as job schedulers map GPU resources to applications, virtual GPUs should also be scheduled before job execution. Nevertheless, current job schedulers are not able to deal with virtual GPUs. In this paper, we analyze the performance attained by a cluster using the remote Compute Unified Device Architecture middleware and a modified version of the Slurm scheduler, which is now able to assign remote GPUs to jobs. Results show that cluster throughput, measured as jobs completed per time unit, is doubled at the same time that the total energy consumption is reduced up to 40%. GPU utilization is also increased.
Sergio Iserte, Javier Prades, Carlos Reaño, Federico Silla
Concurr. Comput. Pract. Exp.4
2021 Redesigning the rCUDA communication layer for a better adaptation to the underlying hardware
abstract
Summary The use of Graphics Processing Units (GPUs) has become a very popular way to accelerate the execution of many applications. However, GPUs are not exempt from side effects. For instance, GPUs are expensive devices which additionally consume a non‐negligible amount of energy even when they are not performing any computation. Furthermore, most applications present low GPU utilization. To address these concerns, the use of GPU virtualization has been proposed. In particular, remote GPU virtualization is a promising technology that allows applications to transparently leverage GPUs installed in any node of the cluster. In this paper, the remote GPU virtualization mechanism is comparatively analyzed across three different generations of GPUs. The first contribution of this study is an analysis about how the performance of the remote GPU virtualization technique is impacted by the underlying hardware. To that end, the Tesla K20, Tesla K40, and Tesla P100 GPUs along with FDR and EDR InfiniBand fabrics are used in the study. The analysis is performed in the context of the rCUDA middleware. It is clearly shown that the GPU virtualization middleware requires a comprehensive design of its communication layer, which should be perfectly adapted to every hardware generation in order to avoid a reduction in performance. This is precisely the second contribution of this work, ie, redesigning the rCUDA communication layer in order to improve the management of the underlying hardware. Results show that it is possible to improve bandwidth up to 29.43%, which translates into up to 4.81% average less execution time in the performance of the analyzed applications.
Carlos Reaño, Federico Silla
Concurr. Comput. Pract. Exp.2
2021 Editorial
Carlos Reaño, Federico Silla, Blesson Varghese
J. Parallel Distributed Comput.2
2020 Improving the performance of physics applications in atom-based clusters with rCUDA
Federico Silla, Javier Prades, Elvira Baydal, Carlos Reaño
J. Parallel Distributed Comput.1
2019 Analyzing the performance/power tradeoff of the rCUDA middleware for future exascale systems
Carlos Reaño, Javier Prades, Federico Silla
J. Parallel Distributed Comput.3
2019 On the support of inter-node P2P GPU memory copies in rCUDA
Carlos Reaño, Federico Silla
J. Parallel Distributed Comput.2
2019 GPU-Job Migration: The rCUDA Case
abstract
Virtualization techniques have been shown to report benefits to data centers and other computing facilities. In this regard, not only virtual machines allow to reduce the size of the computing infrastructure while increasing overall resource utilization, but also virtualizing individual components of computers may provide significant benefits. This is the case, for instance, for the remote GPU virtualization technique, implemented in several frameworks during the recent years. The large degree of flexibility provided by the remote GPU virtualization technique can be further increased by applying the migration mechanism to it, so that the GPU part of applications can be live-migrated to another GPU elsewhere in the cluster during execution time in a transparent way. In this paper we present the implementation of the migration mechanism within the rCUDA remote GPU virtualization middleware. Furthermore, we present a thorough performance analysis of the implementation of the migration mechanism within rCUDA. To that end, we leverage both synthetic and real production applications as well as three different generations of NVIDIA GPUs. Additionally, two different versions of the InfiniBand interconnect are used in this study. Several use cases are provided in order to show the extraordinary benefits that the GPU-job migration mechanism can report to data centers.
Javier Prades, Federico Silla
IEEE Trans. Parallel Distributed Syst.2
2018 Heterogeneous and unconventional cluster architectures and applications
abstract
Recent trends in cluster computing and related topics, including processor and memory design, demonstrate a continuous growing need for more processing and memory, both in terms of capacity and performance. Cluster computing has been traditionally at the forefront of such computing systems and, usually, is one of the earliest adopters of future and emerging technologies. With this special issue, we gear to gather recent related works, hoping that the reader finds these contributions helpful for a better understanding of future directions in cluster computing. Before that, we will briefly present a short rationale on our view of cluster computing. This rationale is based on two trends that are most important for such cluster architectures: first, the end of Dennard scaling has led to an era in which the growing amount of transistors, as described by Moore's law, cannot be simultaneously active because of an increasing power density. Second, applications continue to demand more processing and memory capacity. However, economy and technology laws imply that horizontal scaling is usually more cost effective than vertical scaling. Furthermore, he also described power scaling rules for CMOS silicon dies, defined by the observation that a transition to a new processing technology will decrease the feature size (ie, the size of a transistor gate) by a factor α. Based on this factor, characteristics including voltage, current, and capacity scale inversely. Formula 1 then shows that, for a given power budget, one can implement α2 more components ”a” and even increase operating frequency f by a factor of α. Thus, Dennard scaling actually gave Moore's law its teeth by enabling constant power budgets and frequency scaling. Unfortunately, since early 2015, this law is no longer applicable, mainly because voltage scaling is no longer possible because of saturated threshold voltages and because leakage power became a major contributor to the overall power consumption. As a result, a still growing amount of transistors, as described by Moore's law, now results in a growing power budget, which results intype a hard technical constraint. The usual escape path for post-Dennard performance scaling is two-fold: first, one can observe that frequency usually behaves linearly with regard to voltage, effectively making power consumption proportional to frequency cubed. For instance, reducing frequency by half would result in a relative power consumption of 1/8th, allowing to replicate one computational core 8 times while maintaining the power budget. Second, it is a common first-order approximation to assume that performance behaves proportional to frequency. To continue the previous example, such an 8-core design at half the initial frequency would now result in a relative performance improvement of a factor of 4. However, this is obviously only feasible if the application workloads exhibit enough parallelism. Given extreme examples of many-core processors with 1000s of vector units, heterogeneity is usually the chosen solution to also support the sequential parts of the workloads. As a result, we are seeing a huge interest in many-core processors, which, however, only excel in performance for massively parallel workloads. Given that not all workloads fulfill this requirement, heterogeneous architectures are being deployed. Applications continue to demand for an increasing amount of processing power and memory capacity, in particular pushed by Big Data and Machine Learning. For instance, training a recurrent deep neural network requires about 20 ExaFLOPs and still is not being trained with all data available. Similarly, in particular, deep learning required a plethora of data, leading to huge data collections for various tasks. However, the costs of resources like processors or memory do not scale linearly with capability. On the other hand, while the manufacturers are not very candid about the reasons behind, it is well known that the yield of a silicon die production is a function of the die area. While small dies are less likely to contain a manufacturing error, this probability will increase with die size. Process variation can have similar effects on operating frequency, making high-speed designs more sensitive to such variations. With the series on Heterogeneous and Unconventional Cluster Architectures and Applications, we gear to gather most recent insights and ideas from the wide area of cluster computing. This special issue of the "International Journal of Concurrency and Computation: Practice and Experience" resembles our most recent selection. In particular, this selection includes four interesting works.1-4 Two of them were contributions from the last two workshop editions (HUCAA 2015 and 2016, both collocated with the International Conference on Parallel Processing - ICPP'15 and ICPP'16). Furthermore, the two other articles are contributed by the authors based on an open call for contributions of this special issue. All contributions were peer reviewed and received in between three and five reviews. GPU accelerators have been established in the state-of-the-art clusters by offering high performance and energy efficiency. In GPU, such efficient communication among processes with their data residing in the GPU memory is of paramount importance to the application performance. This paper investigates various algorithms in conjunction with the latest GPU features to improve GPU collective operations. For clusters with multi-GPU nodes, the authors of this paper propose a hierarchical framework that allows different algorithms at each hierarchy level. By studying various combination of algorithms, the authors of this article highlight the importance of choosing the right algorithm within each level. They evaluate their framework on MPI Allreduce and show promising performance results, specifically for large message sizes which are highly in-use in deep learning and big data applications. They also show the benefit of using the Hyper-Q feature and the MPS service in jointly using different copy types to perform multiple inter-process communications. However, the authors of this paper show that efficient designs are required to further harness this potential. Accordingly, they propose Hyper-Q aware algorithms for GPU collectives. They evaluate our algorithms on MPI Allgather and MPI Allreduce operations. While their experimental results show the benefit of their algorithms, their profiling results indicate that this benefit is mainly rooted in overlapping different copy types. Virtual Screening methods (VS) simulate molecular interactions in silico to look for the best chemical compound that interacts with a given molecular target. VS are becoming increasingly popular to accelerate the drug discovery process and constitute hard optimization problems with a huge computational cost. To deal with these two challenges, the authors of this paper have created METADOCK, an application that (1) enables a wide range of metaheuristics through a parametrized schema, and (2) promotes the use of a multi-GPU environment within a heterogeneous cluster. Metaheuristics provide approximate solutions in a reasonable time frame, but given the stochastic nature of real-life procedures, the energy budget goes hand in hand with acceleration to validate the proposed solution. This paper evaluates energy trade-offs and correlations with performance for a set of metaheuristics derived from METADOCK. The authors of this paper establish a solid inference from minimal power to maximal performance in GPUs and from there to optimal energy consumption. This way, ideal heuristics can be chosen according not only to best accuracy and performance but also to energy requirements. Their study starts with a preselection of parameterized metaheuristic functions, building blocks where we will find optimal patterns from power criteria while preserving parallelism through a GPU execution. They then establish a methodology to figure out the best instances of the parameterized kernels based on energy patterns obtained, which are analyzed from different viewpoints: performance, average power, and total energy consumed. The authors of this paper also compare the best workload distributions for optimal performance and power efficiency among Pascal and Maxwell GPUs on popular Titan models. The experimental results in this paper demonstrate that the most power efficient GPU can be overloaded in order to reduce the total amount of energy required by as much as 20%, finding unique scenarios where Maxwell does it better in execution time, but with Pascal always ahead in performance per watt, reaching peaks of up to 40%. Faster, lower power, and/or less expensive computation will be a software problem forever. Hardware can only make the challenge simpler or harder, and heterogeneous approaches exacerbate it. For emerging alternative computational technologies like quantum, optical, resistive (and other forms of analog computation), and/or biological computing (among others), to be successful, they must be integrated into the existing computational infrastructure (both hardware and software) if they are to realize their full potential. The increasingly main-stream options that reconfigurable logic represents (both fine and coarse grained) will also be most useful within an infrastructure that is sympathetic to legacy memory and storage models. SAHARA is a reduction of computation into data wavefronts that, independent of the underlying technology, employs memory as the fundamental unit of computation within a simple data-flow model, essentially turning processing into a side-effect of the relevant data being made available to the logic that manipulates that data. No single aspect of SAHARA is “new”. Its foundations are more than 50 years old and started with Minsky's 1961 paper on Turing equivalence. SAHARA is an eminently useful abstraction of computation that has the potential of seamlessly integrating many disparate forms of computation behind a simple, common, architectural interface. The rise of heterogeneous systems has given place to great challenges for users, as they involve new concepts, restrictions and frameworks. Their exploitation is further complicated in the context of distributed memory systems, which require the usage of additional different programming paradigms and tools. In this paper, the authors propose a novel approach to program heterogeneous clusters that is based on high-level abstractions such as tiles and hierarchical decomposition combined with the powerful APIs that data types and embedded languages can provide in languages such as C++. Rather than building their proposal from scratch, they have implemented it as a natural integration of the existing Hierarchically Tiled Arrays (HTA) and Heterogeneous Programming Library (HPL) projects, the first one being focused on distributed computing and the second one on heterogeneous processing. The result, called Heterogeneous Hierarchically Tiled Arrays (H2TA), is very intuitive and easy to use thanks to the global view of the data and the single-threaded view of the execution that it provides at cluster level together with the transparency it provides with respect to the management of the heterogeneous devices. An evaluation comparing the proposal in this paper with previous MPI-based implementations shows its large programmability advantages and the reasonable overhead incurred. Cluster computing is currently facing a pivotal point in time as we are hitting hard constraints about the future of CMOS processors. While, currently, most energy is still spent on computations, first research results show that an increasing fraction of overall energy is spent for data movements. Given the hard constraints on power consumption, one can imagine how influential this fundamental transition will be. Still, CMOS replacements like quantum computing, neuromorphic computing and many other candidates are either still nascent or will only be helpful for certain workloads. While it seems safe to assume that this will further increase heterogeneity in the future, we will have to find out if the currently narrow workload spectrum for these architectures can be extended or if generic computing in the future will have so solely rely on CMOS processors. In this context, we hope the readers of this special issue will find the contributions interesting and inspiring for future research. We in particular acknowledge the thorough work of our review board, which did an excellent and timely work on proving opinions and helpful feedback to the submitted articles. Furthermore, we are also thankful to the authors, which submitted their research contributions to our special edition. Last, but not least, we are also indebted to the continuous support by the editor-in-chief Geoffrey Fox, who is always available for assistance and recommendations regarding the publication procedure.
Holger Fröning, Federico Silla
Concurr. Comput. Pract. Exp.2
2018 Enhancing large-scale docking simulation on heterogeneous systems: An MPI vs rCUDA study
Baldomero Imbernon, Javier Prades, Domingo Giménez, José M. Cecilia, Federico Silla
Future Gener. Comput. Syst.5
2018 Intra-Node Memory Safe GPU Co-Scheduling
abstract
GPUs in High-Performance Computing systems remain under-utilised due to the unavailability of schedulers that can safely schedule multiple applications to share the same GPU. The research reported in this paper is motivated to improve the utilisation of GPUs by proposing a framework, we refer to as schedGPU, to facilitate intra-node GPU co-scheduling such that a GPU can be safely shared among multiple applications by taking memory constraints into account. Two approaches, namely a client-server and a shared memory approach are explored. However, the shared memory approach is more suitable due to lower overheads when compared to the former approach. Four policies are proposed in schedGPU to handle applications that are waiting to access the GPU, two of which account for priorities. The feasibility of schedGPU is validated on three real-world applications. The key observation is that a performance gain is achieved. For single applications, a gain of over 10 times, as measured by GPU utilisation and GPU memory utilisation, is obtained. For workloads comprising multiple applications, a speed-up of up to 5x in the total execution time is noted. Moreover, the average GPU utilisation and average GPU memory utilisation is increased by 5 and 12 times, respectively.
Carlos Reaño, Federico Silla, Dimitrios S. Nikolopoulos, Blesson Varghese
IEEE Trans. Parallel Distributed Syst.2
2017 A Live Demo for Showing the Benefits of Applying the Remote GPU Virtualization Technique to Cloud Computing
abstract
Cloud computing has become pervasive nowadays. Additionally, cloud computing customers increasingly demand the use of accelerators such as CUDA GPUs. This has motivated that Amazon, for example, provides virtual machine instances comprising up to 16 NVIDIA GPUs. However, the use of GPUs in cloud computing deployments is not exempt from important concerns. In order to overcome many of these concerns, the remote GPU virtualization technique can be used. In this paper we present the design of a live demo to be used in exhibitions in order to show the benefits of using such a virtualization technique in the context of cloud computing systems. The demo was designed to be technically sound at the same time that it draws the attention of attendees. The demo was successfully used in the recent SuperComputing'16 exhibition, attracting more than 100 people to the booth.
Javier Prades, Federico Silla
CCGrid2
2017 Enhancing the rCUDA Remote GPU Virtualization Framework: from a Prototype to a Production Solution
abstract
The use of hardware accelerators to increase the performance of parallel applications is very common nowadays. For a number of reasons, however, the access to local accelerators is not always feasible (e.g., lack of space or cost). It would also be the case that some applications benefit from having access to more accelerators than the physically possible. To address all these concerns, middleware offering access not only to local but also to remote accelerators appeared. This paper presents a high-level summary of a dissertation focused on enhancing one of these middleware, called rCUDA.
Carlos Reaño, Federico Silla, José Duato
CCGrid2
2017 On the benefits of the remote GPU virtualization mechanism: The rCUDA case
abstract
Summary Graphics processing units (GPUs) are being adopted in many computing facilities given their extraordinary computing power, which makes it possible to accelerate many general purpose applications from different domains. However, GPUs also present several side effects, such as increased acquisition costs as well as larger space requirements. They also require more powerful energy supplies. Furthermore, GPUs still consume some amount of energy while idle, and their utilization is usually low for most workloads. In a similar way to virtual machines, the use of virtual GPUs may address the aforementioned concerns. In this regard, the remote GPU virtualization mechanism allows an application being executed in a node of the cluster to transparently use the GPUs installed at other nodes. Moreover, this technique allows to share the GPUs present in the computing facility among the applications being executed in the cluster. In this way, several applications being executed in different (or the same) cluster nodes can share 1 or more GPUs located in other nodes of the cluster. Sharing GPUs should increase overall GPU utilization, thus reducing the negative impact of the side effects mentioned before. Reducing the total amount of GPUs installed in the cluster may also be possible. In this paper, we explore some of the benefits that remote GPU virtualization brings to clusters. For instance, this mechanism allows an application to use all the GPUs present in the computing facility. Another benefit of this technique is that cluster throughput, measured as jobs completed per time unit, is noticeably increased when this technique is used. In this regard, cluster throughput can be doubled for some workloads. Furthermore, in addition to increase overall GPU utilization, total energy consumption can be reduced up to 40%. This may be key in the context of exascale computing facilities, which present an important energy constraint. Other benefits are related to the cloud computing domain, where a GPU can be easily shared among several virtual machines. Finally, GPU migration (and therefore server consolidation) is one more benefit of this novel technique.
Federico Silla, Sergio Iserte, Carlos Reaño, Javier Prades
Concurr. Comput. Pract. Exp.1
2017 Multi-tenant virtual GPUs for optimising performance of a financial risk application
Javier Prades, Blesson Varghese, Carlos Reaño, Federico Silla
J. Parallel Distributed Comput.4
2016 Increasing the Performance of Data Centers by Combining Remote GPU Virtualization with Slurm
abstract
The use of Graphics Processing Units (GPUs) presents several side effects, such as increased acquisition costs as well as larger space requirements. Furthermore, GPUs require a non-negligible amount of energy even while idle. Additionally, GPU utilization is usually low for most applications. Using the virtual GPUs provided by the remote GPU virtualization mechanism may address the concerns associated with the use of these devices. However, in the same way as workload managers map GPU resources to applications, virtual GPUs should also be scheduled before job execution. Nevertheless, current workload managers are not able to deal with virtual GPUs. In this paper we analyze the performance attained by a cluster using the rCUDA remote GPU virtualization middleware and a modified version of the Slurm workload manager, which is now able to map remote virtual GPUs to jobs. Results show that cluster throughput is doubled at the same time that total energy consumption is reduced up to 40%. GPU utilization is also increased.
Sergio Iserte, Javier Prades, Carlos Reaño, Federico Silla
CCGrid4
2016 Providing CUDA Acceleration to KVM Virtual Machines in InfiniBand Clusters with rCUDA
abstract
There is a trend towards using graphics processing units (GPUs) not only for graphics visualization, but also for accelerating scientific applications. But their use for this purpose is not without disadvantages: GPUs increase costs and energy consumption. Furthermore, GPUs are generally underutilized. Using virtual machines could be a possible solution to address these problems, however, current solutions for providing GPU acceleration to virtual machines environments, such as KVM or Xen, present some issues. In this paper we propose the use of remote GPUs to accelerate scientific applications running inside KVM virtual machines. Our analysis shows that this approach could be a possible solution, with low overhead when used over InfiniBand networks.
Ferran Perez, Carlos Reaño, Federico Silla
DAIS3
2016 Reducing the performance gap of remote GPU virtualization with InfiniBand Connect-IB
abstract
GPU accelerators may provide great performance improvements in the context of parallel applications. However, their use in HPC clusters may also present some disadvantages such as their high cost and high power consumption. In addition, this kind of accelerators are generally underutilized. Remote GPU virtualization could be a solution to overcome these drawbacks, but its performance is usually impaired because the network bandwidth is lower than the PCIe one. In this paper we analyze how the InfiniBand Connect-IB network adapters (with performance similar to that of PCIe 3.0) reduce the overhead of remote GPU virtualization. We show that this overhead is decreased to 1.5% in terms of bandwidth, and to 0.51% in the tested application.
Carlos Reaño, Federico Silla
ISCC2
2016 CUDA acceleration for Xen virtual machines in infiniband clusters with rCUDA
abstract
Many data centers currently use virtual machines (VMs) to achieve a more efficient usage of hardware resources. However, current virtualization solutions, such as Xen, do not easily provide graphics processing unit (GPU) accelerators to applications running in the virtualized domain with the flexibility usually required in data centers (i.e., managing virtual GPU instances and concurrently sharing them among several VMs). Remote GPU virtualization frameworks such as the rCUDA solution may address this problem.
Javier Prades, Carlos Reaño, Federico Silla
PPoPP3
2016 Heterogeneous cluster architectures and applications
abstract
Welcome to the Special Issue on Heterogeneous and Unconventional Cluster Architectures and Applications! Why such a topic for a special issue? Which is the reason for fostering recent work on cluster architectures that do not fit into the widely established categories? Are unconventional applications especially interesting so that they are worth a special issue? Do these disruptive pieces of work deserve such a focus in a journal? From the point of view of the guest editors of this special issue, the progress of technology can be seen as a two-step process. One of the steps is performing research in order to achieve a better version of the designs we are currently enjoying. In this regard, for instance, when an interconnection network is enhanced so that its bandwidth is doubled, datacenters can deliver a new level of performance to their customers. In a similar way, new generations of processors help to significantly reduce energy consumption of the servers used in such datacenter. Also, installing new versions of accelerators in such servers allows applications to experience important reductions in their execution time at the time that the overall power consumption of the facility is not impacted. All these new developments require an enormous amount of research. For example, faster interconnects require to glue together research on microelectronics, physics, communication protocols, encoding schemes, and thermal issues, to name only a few. The creation of new generations of processors and accelerators also involves large amounts of research from a variety of areas, thus making that a big team of researchers is required in order to bring all those ideas to market. Nevertheless, if those developments are observed from a very broad and long-term perspective, they could be thought to be incremental, despite of their clear significance and despite that they enable new levels of performance. Hence, which should be the non-incremental developments? Our understanding of this fundamental problem is that the non-incremental developments would be those based on ideas that could cause surprise to those researchers initially listening to them. For instance, the very first time that several computers were interconnected in order to create a cluster of computers that cooperatively work together in order to solve a problem, at that point in time an unconventional idea appeared. From a long-term perspective, all the later developments around clusters are incremental; despite those developments are the result of many hours of work and huge amounts of effort from very smart researchers that probably devoted their lives to create those improved versions of the technology behind that initial idea of putting together a few computers into a cluster so that they collaborate in order to solve a problem. One more example of these non-incremental developments could be the use of Graphics Processing Unit (GPU) or any other accelerator in general, in order to reduce the execution time of applications. In this regard, new versions of these accelerators are required so that technology makes progress. Furthermore, these new versions of these accelerators are the result of big investments in time as well as funding resources, which involve hundreds of trained researchers. However, the qualitative difference was performed when someone had the initial idea of attaching an accelerator to the system. That was the point in time that really impacted technology and made it evolve because that allowed a performance increase of several orders of magnitude for the datacenters. Hence, all these non-incremental developments are the other step of the two-step process described previously regarding the progress of technology. Actually, in this two-step process, these non-incremental ideas would be the first step, and the second step would be composed of all those developments that enhance and enrich the initial idea. In fact, the first step cannot live without the second, because both of them are equally important and both of them are required. Given that the two steps are required, and given that most of current conferences and journals include a large percentage of work devoted to the second step, in the opinion of the guest editors of this special issue, it is worth to make an effort to gather innovative work on unconventional cluster architectures and applications, which might have a big impact on future cluster architectures. This includes more or less any cluster architecture that is not based on the usual commodity components and therefore makes use of some special hardware or software components, or that is used for very special and unconventional applications. Within this fostering effort, it is particularly encouraged to work on disruptive approaches, which may show inferior performance today but can already point out their performance potential. Obviously, innovative and disruptive ideas like putting together a bunch of computers into a cluster so that they collaborate to solve a problem only happen from time to time. Actually, they happen seldom. Most innovative ideas are not so disruptive. However, this obvious fact will not discourage the editors of this special issue on their hunt. This special issue contains five pieces of work selected from eleven submissions. Four independent reviews were produced for each of the submissions. From these lines, the guest editors profoundly thank the review board members for their excellent work. We know that the review task is performed on a volunteer basis and many times taking time from other areas such as personal leisure or time with family and friends. This is why the support of our reviewers is so appreciated by us. Actually, without their time and effort, this special issue could not be completed. We are happy about receiving papers from a broad range of topics, allowing this special issue to cover aspects from hardware architecture design, from low-level software and run-times up to new system architecture concepts. The first paper shows how accelerators as traffic sources and sinks impact collective operations 1. The authors point out a set of optimizations and analyze their effectiveness. They also show that a variety of algorithms is required for best performance and that accelerators as sources and sinks significantly increase the overhead for collective operations when compared with general-purpose processors like CPUs. In summary, these improvements lead to a higher utilization of clusters composed of heterogeneous components. In the next paper, the authors propose several optimizations for schedulers of heterogeneous systems 2. First, they propose the possibility of resource range specifications to improve resource specification within batch jobs. Second, they introduce a new integer programming formulation that allows reducing the number of variables, thereby enabling faster solutions. In the third paper, the authors introduce Gaspar, an aspect-oriented framework that facilitates the development of Java applications for heterogeneous systems and clusters 3. They propose to exploit parallelism by aspect modules and a data layout that is dependent on the target platform. Their proposed aspect modules for Java have many similarities with OpenMP formulations for C/C++. The impact of different data structures and layouts for heterogeneous resources is well known, and they address this fact including support in their framework for different variants. Their evaluation shows how such an approach allows for a high-performance portability and general competitive performance. In the fourth paper, the authors show how they use a custom cluster based on reconfigurable logic to design a highly optimized solution for certain tasks, including communication-intensive molecular dynamics 4. Reconfigurable logic opens up a new degree of freedom when designing computing systems; however, this huge amount of freedom comes with significant complexity increase. The authors show how they approached the problem with a combination of modeling and abstractions, demonstrating how in the future reconfigurable logic could change the computing landscape. Last, the fifth paper describes the approach of the European Union-funded DEEP project to design a new cluster architecture for Exascale computing 5. Opposed to other systems, their concept decouples the ratio of general-purpose CPUs and specialized accelerators, giving additional freedom when upgrading such systems. Besides this, they also nicely point out the importance of accelerators directly sourcing and sinking traffic. The guest editors of this special issue hope that you enjoy the read as much as we did enjoy composing it. We believe this collection to be representative of some important problems related to heterogeneous systems and clusters today and that such a reading is highly inspiring for our own research and general interest.
Federico Silla, Holger Fröning
Concurr. Comput. Pract. Exp.1
2016 Tuning remote GPU virtualization for InfiniBand networks
Carlos Reaño, Federico Silla
J. Supercomput.2
2015 On the Design of a Demo for Exhibiting rCUDA
abstract
CUDA is a technology developed by NVIDIA which provides a parallel computing platform and programming model for NVIDIA GPUs and compatible ones. It takes benefit from the enormous parallel processing power of GPUs in order to accelerate a wide range of applications, thus reducing their execution time. rCUDA (remote CUDA) is a middleware which grants applications concurrent access to CUDA-compatible devices installed in other nodes of the cluster in a transparent way so that applications are not aware of being accessing a remote device. In this paper we present a demo which shows, in real time, the overhead introduced by rCUDA in comparison to CUDA when running image filtering applications. The approach followed in this work is to develop a graphical demo which contains both an appealing design and technical contents.
Carlos Reaño, Ferran Perez, Federico Silla
CCGRID3
2015 A Performance Comparison of CUDA Remote GPU Virtualization Frameworks
abstract
Using GPUs reduces execution time of many applications but increases acquisition cost and power consumption. Furthermore, GPUs usually attain a relatively low utilization. In this context, remote GPU virtualization solutions were recently created to overcome the drawbacks of using GPUs. Currently, many different remote GPU virtualization frameworks exist, all of them presenting very different characteristics. These differences among them may lead to differences in performance. In this work we present a performance comparison among the only three CUDA remote GPU virtualization frameworks publicly available at no cost. Results show that performance greatly depends on the exact framework used, being the rCUDA virtualization solution the one that stands out among them. Furthermore, rCUDA doubles performance over CUDA for pageable memory copies.
Carlos Reaño, Federico Silla
CLUSTER2
2015 InfiniBand Verbs Optimizations for Remote GPU Virtualization
abstract
The use of InfiniBand networks to interconnect high performance computing clusters has considerably increased during the last years. So much so that the majority of the supercomputers included in the TOP500 list either use Ethernet or InfiniBand interconnects. Regarding the latter, due to the complexity of the InfiniBand programming API (i.e., InfiniBand Verbs) and the lack of documentation, there are not enough recent available studies explaining how to optimize applications to get the maximum performance from this fabric. In this paper we expose two different optimizations to be used when developing applications using InfiniBand Verbs, each providing an average bandwidth improvement of 3.68% and 217.14%, respectively. In addition, we show that when combining both optimizations, the average bandwidth gain is 43.29%. This bandwidth increment is key for remote GPU virtualization frameworks. Actually, this noticeable gain translates into a reduction of up to 35% in execution time of applications using remote GPU virtualization frameworks.
Carlos Reaño, Federico Silla
CLUSTER2
2015 On the Execution of Computationally Intensive CPU-Based Libraries on Remote Accelerators for Increasing Performance: Early Experience with the OpenBLAS and FFTW Libraries
abstract
Virtualization techniques have shown to report benefits to data centers and other computing facilities. In this regard, virtual machines not only allow reducing the size of the computing infrastructure while increasing overall resource utilization but virtualizing individual components of computers may also provide significant benefits. This is the case, for example, for the remote GPU virtualization technique, implemented in several frameworks during the last years. In this paper we present an initial implementation of a new middleware for the remote virtualization of another component of computers: the CPU itself. Our proposal uses remote accelerators to perform computations that were initially intended to be carried out in the local CPUs, doing so transparently to the application and without having to modify its source code. By making use of the OpenBLAS and FFTW libraries as case studies to show the performance gains of our proposal, we carry out a performance evaluation targeting several system configurations comprising Xeon processors as well as Ethernet and InfiniBand QDR, FDR, and EDR network adapters in addition to NVIDIA Tesla K40 GPUs. Results not only demonstrate that the new middleware is feasible, but they also show that mathematical libraries may experience a significant speed up, despite of having to move data forth and back to/from remote servers.
Santiago Mislata Valero, Federico Silla
CLUSTER2
2015 Acceleration-as-a-Service: Exploiting Virtualised GPUs for a Financial Application
abstract
How can GPU acceleration be obtained as a service in a cluster? This question has become increasingly significant due to the inefficiency of installing GPUs on all nodes of a cluster. The research reported in this paper is motivated to address the above question by employing rCUDA (remote CUDA), a framework that facilitates Acceleration-as-a-Service (AaaS), such that the nodes of a cluster can request the acceleration of a set of remote GPUs on demand. The rCUDA framework exploits virtualisation and ensures that multiple nodes can share the same GPU. In this paper we test the feasibility of the rCUDA framework on a real-world application employed in the financial risk industry that can benefit from AaaS in the production setting. The results confirm the feasibility of rCUDA and highlight that rCUDA achieves similar performance compared to CUDA, provides consistent results, and more importantly, allows for a single application to benefit from all the GPUs available in the cluster without loosing efficiency.
Blesson Varghese, Javier Prades, Carlos Reaño, Federico Silla
e-Science4
2015 Improving the user experience of the rCUDA remote GPU virtualization framework
abstract
Summary Graphics processing units (GPUs) are being increasingly embraced by the high‐performance computing community as an effective way to reduce execution time by accelerating parts of their applications. remote CUDA (rCUDA) was recently introduced as a software solution to address the high acquisition costs and energy consumption of GPUs that constrain further adoption of this technology. Specifically, rCUDA is a middleware that allows a reduced number of GPUs to be transparently shared among the nodes in a cluster. Although the initial prototype versions of rCUDA demonstrated its functionality, they also revealed concerns with respect to usability, performance, and support for new CUDA features. In response, in this paper, we present a new rCUDA version that (1) improves usability by including a new component that allows an automatic transformation of any CUDA source code so that it conforms to the needs of the rCUDA framework, (2) consistently features low overhead when using remote GPUs thanks to an improved new communication architecture, and (3) supports multithreaded applications and CUDA libraries. As a result, for any CUDA‐compatible program, rCUDA now allows the use of remote GPUs within a cluster with low overhead, so that a single application running in one node can use all GPUs available across the cluster, thereby extending the single‐node capability of CUDA. Copyright © 2014 John Wiley & Sons, Ltd.
Carlos Reaño, Federico Silla, Adrián Castelló 0001, Antonio J. Peña, Rafael Mayo 0002, Enrique S. Quintana-Ortí, José Duato
Concurr. Comput. Pract. Exp.2
2015 On the design of a new dynamic credit-based end-to-end flow control mechanism for HPC clusters
Javier Prades, Federico Silla, Holger Fröning, Mondrian Nüssle, José Duato
Parallel Comput.2
2014 Boosting the performance of remote GPU virtualization using InfiniBand connect-IB and PCIe 3.0
abstract
A clear trend has emerged involving the acceleration of scientific applications by using GPUs. However, the capabilities of these devices are still generally underutilized. Remote GPU virtualization techniques can help increase GPU utilization rates, while reducing acquisition and maintenance costs. The overhead of using a remote GPU instead of a local one is introduced mainly by the difference in performance between the internode network and the intranode PCIe link. In this paper we show how using the new InfiniBand Connect-IB network adapters (attaining similar throughput to that of the most recently emerged GPUs) boosts the performance of remote GPU virtualization, reducing the overhead to a mere 0.19% in the application tested.
Carlos Reaño, Federico Silla, Antonio J. Peña, Gilad Shainer, Scot Schultz, Adrián Castelló 0001, Enrique S. Quintana-Ortí, José Duato
CLUSTER2
2014 SLURM Support for Remote GPU Virtualization: Implementation and Performance Study
abstract
SLURM is a resource manager that can be leveraged to share a collection of heterogeneous resources among the jobs in execution in a cluster. However, SLURM is not designed to handle resources such as graphics processing units (GPUs). Concretely, although SLURM can use a generic resource plugin (GRes) to manage GPUs, with this solution the hardware accelerators can only be accessed by the job that is in execution on the node to which the GPU is attached. This is a serious constraint for remote GPU virtualization technologies, which aim at providing a user-transparent access to all GPUs in cluster, independently of the specific location of the node where the application is running with respect to the GPU node. In this work we introduce a new type of device in SLURM, "rgpu", in order to gain access from any application node to any GPU node in the cluster using rCUDA as the remote GPU virtualization solution. With this new scheduling mechanism, a user can access any number of GPUs, as SLURM schedules the tasks taking into account all the graphics accelerators available in the complete cluster. We present experimental results that show the benefits of this new approach in terms of increased flexibility for the job scheduler.
Sergio Iserte, Adrián Castelló 0001, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Federico Silla, José Duato, Carlos Reaño, Javier Prades
SBAC-PAD5
2014 A complete and efficient CUDA-sharing solution for HPC clusters
Antonio J. Peña, Carlos Reaño, Federico Silla, Rafael Mayo 0002, Enrique S. Quintana-Ortí, José Duato
Parallel Comput.3
2013 Influence of InfiniBand FDR on the performance of remote GPU virtualization
abstract
The use of GPUs to accelerate general-purpose scientific and engineering applications is mainstream today, but their adoption in current high-performance computing clusters is impaired primarily by acquisition costs and power consumption. Therefore, the benefits of sharing a reduced number of GPUs among all the nodes of a cluster can be remarkable for many applications. This approach, usually referred to as remote GPU virtualization, aims at reducing the number of GPUs present in a cluster, while increasing their utilization rate. The performance of the interconnection network is key to achieving reasonable performance results by means of remote GPU virtualization. To this end, several networking technologies with throughput comparable to that of PCI Express have appeared recently. In this paper we analyze the influence of InfiniBand FDR on the performance of remote GPU virtualization, comparing its impact on a variety of GPU-accelerated applications with other networking technologies, such as Infini-Band QDR and Gigabit Ethernet. Given the severe limitations of freely available remote GPU virtualization solutions, the rCUDA framework is used as the case study for this analysis. Results show that the new FDR interconnect, featuring higher bandwidth than its predecessors, allows the reduction of the overhead of using GPUs remotely, thus making this approach even more appealing.
Carlos Reaño, Rafael Mayo 0002, Enrique S. Quintana-Ortí, Federico Silla, José Duato, Antonio J. Peña
CLUSTER4
2013 Silicon-aware distributed switch architecture for on-chip networks
Antoni Roca 0001, Carles Hernández 0001, José Flich, Federico Silla, José Duato
J. Syst. Archit.4
2012 A New End-to-End Flow-Control Mechanism for High Performance Computing Clusters
abstract
High Performance Computing usually leverages messaging libraries such as MPI or GASNet in order to exchange data among processes in large-scale clusters. Furthermore, these libraries make use of specialized low-level networking layers in order to retrieve as much performance as possible from hardware interconnects such as Infini Band or Myrinet, for example. EXTOLL is another emerging technology targeted for high performance clusters. These specialized low-level networking layers require some kind of flow control in order to prevent buffer overflows at the received side. In this paper we present a new flow control mechanism that is able to adapt the buffering resources used by a process according to the parallel application communication pattern and the varying activity among communicating peers. The tests carried out in a 64-node 1024-core EXTOLL cluster show that our new dynamic flow-control mechanism provides extraordinarily high buffer efficiency along with very low overhead, which is reduced between 8 and 10 times.
Javier Prades, Federico Silla, José Duato, Holger Fröning, Mondrian Nüssle
CLUSTER2
2012 CU2rCU: Towards the complete rCUDA remote GPU virtualization and sharing solution
abstract
GPUs are being increasingly embraced by the high performance computing and computational communities as an effective way of considerably reducing execution time by accelerating significant parts of their application codes. However, despite their extraordinary computing capabilities, the adoption of GPUs in current HPC clusters may present certain negative side-effects. In particular, to ease job scheduling in these platforms, a GPU is usually attached to every node of the cluster. In addition to increasing acquisition costs this favors that GPUs may frequently remain idle, as applications usually do not fully utilize them. On the other hand, idle GPUs consume non-negligible amounts of energy, which translates into very poor energy efficiency during idle cycles. rCUDA was recently developed as a software solution to address these concerns. Specifically, it is a middleware that allows transparently sharing a reduced number of GPUs among the nodes in a cluster. rCUDA thus increases the GPU-utilization rate, taking care of job scheduling. While the initial prototype versions of rCUDA demonstrated its functionality, they also revealed several concerns related with usability and performance. With respect to usability, in this paper we present a new component of the rCUDA suite that allows an automatic transformation of any CUDA source code, so that it can be effectively accommodated within this technology. In response to performance, we briefly show some interesting results, which will be deeply analyzed in future publications. The net outcome is a new version of rCUDA that allows, for any CUDA-compatible program, to use remote GPUs in a cluster with minimum overhead.
Carlos Reaño, Antonio J. Peña, Federico Silla, José Duato, Rafael Mayo 0002, Enrique S. Quintana-Ortí
HiPC3
2012 Enabling High-Performance Crossbars through a Floorplan-Aware Design
abstract
Networks-on-Chip (NoC) with low-radix switches forming a simple and planar topology is typically accepted as the right interconnection infrastructure for current Chip Multi Processor and high-end Multi Processor System-on-Chip. This is mainly due to its simplicity in the physical mapping on the chip. However, as the network diameter increases, latency and power consumption are increased due to the rapidly growing queuing delay in each switch do not scale with system size. In this context, topologies with high-radix switches have been recently proposed in the NoC scenario to keep message latency low when interconnecting a large number of devices. However, the use of high-radix switches present several well-known drawbacks, being the most important the scalability in area and frequency when implemented. In addition, average and maximum wire length is increased. In this paper we present a distributed crossbar NoC architecture that reduces network latency, increases network throughput significantly. For the distributed crossbar implementation trees of 2-to-1 multiplexers with arbitration and buffer capabilities are spread over the chip, avoiding the negative impact of a high radix switch degree on NoC operating frequency, and minimizing the impact of long wires. Results show that in a 64-node NoC our most aggressive distributed crossbar configuration reduces flit latency by 42% and increases throughput by 544% with respect to the low latency flattened butterfly architecture, meanwhile area is increased a 110%. A more conservative distributed crossbar configuration obtains an increment in throughput of 276.1%, latency is decreased a 29.7%, but area is also decreased by 7%.
Antoni Roca 0001, Carles Hernández 0001, José Flich, Federico Silla, José Duato
ICPP4
2012 On the Impact of Within-Die Process Variation in GALS-Based NoC Performance
abstract
Current integration scales allow designing chip multiprocessors (CMP), where cores are interconnected by means of a network-on-chip (NoC). Unfortunately, the small feature size of current integration scales causes some unpredictability in manufactured devices because of process variation. In NoCs, variability may affect links and routers causing them not to match the parameters established at design time. In this paper, we first analyze the way that manufacturing deviations affect the components of a NoC by applying a new comprehensive and detailed within-die variability model to 200 instances of an 8×8 mesh NoC synthesized using 45 nm technology. Later, we show that GALS-based NoCs present communication bottlenecks under process variation which cannot be avoided by using just device-level solutions but higher level architectural approaches are required. Therefore, to overcome this performance reduction, we draft a novel architectural approach, called performance domains, intended to reduce the negative impact of variability on application execution time. This mechanism is suitable when several applications are simultaneously running in the CMP chip.
Carles Hernández 0001, Antoni Roca 0001, Federico Silla, José Flich, José Duato
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2011 MEMSCALE: in-cluster-memory databases
abstract
We have developed a new memory architecture for clusters that allows automatic access from any processor to any memory module in the cluster completely by hardware. Thus, with a single assembly instruction a processor can retrieve (or update) a memory location in a remote node. The efficiency of this new paradigm makes it possible to speed-up the execution of shared-memory applications with very large memory footprints by running them across the entire cluster, thus providing them a true shared-memory environment (contrary to the emulation typically carried out by software-based distributed shared memory).
Héctor Montaner, Federico Silla, Holger Fröning, José Duato
CIKM2
2011 Enabling CUDA acceleration within virtual machines using rCUDA
abstract
The hardware and software advances of Graphics Processing Units (GPUs) have favored the development of GPGPU (General-Purpose Computation on GPUs) and its adoption in many scientific, engineering, and industrial areas. Thus, GPUs are increasingly being introduced in high-performance computing systems as well as in datacenters. On the other hand, virtualization technologies are also receiving rising interest in these domains, because of their many benefits on acquisition and maintenance savings. There are currently several works on GPU virtualization. However, there is no standard solution allowing access to GPGPU capabilities from virtual machine environments like, e.g., VMware, Xen, VirtualBox, or KVM. Such lack of a standard solution is delaying the integration of GPGPU into these domains. In this paper, we propose a first step towards a general and open source approach for using GPGPU features within VMs. In particular, we describe the use of rCUDA, a GPGPU (General-Purpose Computation on GPUs) virtualization framework, to permit the execution of GPU-accelerated applications within virtual machines (VMs), thus enabling GPGPU capabilities on any virtualized environment. Our experiments with rCUDA in the context of KVM and VirtualBox on a system equipped with two NVIDIA GeForce 9800 GX2 cards illustrate the overhead introduced by the rCUDA middleware and prove the feasibility and scalability of this general virtualizing solution. Experimental results show that the overhead is proportional to the dataset size, while the scalability is similar to that of the native environment.
José Duato, Antonio J. Peña, Federico Silla, Juan Carlos Fernández 0002, Rafael Mayo 0002, Enrique S. Quintana-Ortí
HiPC3
2011 Highly scalable barriers for future high-performance computing clusters
abstract
Although large scale high performance computing today typically relies on message passing, shared memory can offer significant advantages, as the overhead associated with MPI is completely avoided. In this way, we have developed an FPGA-based Shared Memory Engine that allows to forward memory transactions, like loads and stores, to remote memory locations in large clusters, thus providing a single memory address space. As coherency protocols do not scale with system size we completely avoid a global coherency across the cluster. However, we maintain local coherency domains, thus keeping the cores within one node coherent. In this paper, we show the suitability of our approach by analyzing the performance of barriers, a very common synchronization primitive in parallel programs. Experiments in a real cluster prototype show that our approach allows synchronization among 1024 cores spread over 64 nodes in less than 15us, several times faster than other highly optimized barriers. We show the feasibility of this approach by executing a shared-memory implementation of FFT. Finally, note that this barrier can also be leveraged by MPI applications running on our shared memory architecture for clusters. This ensures the usefulness of this work for applications already written.
Holger Fröning, Alexander Giese, Héctor Montaner, Federico Silla, José Duato
HiPC4
2011 Unleash Your Memory-Constrained Applications: A 32-Node Non-coherent Distributed-Memory Prototype Cluster
abstract
Improvements in hardware for parallel shared-memory computing usually involve increments in the number of computing cores and in the amount of memory available for a given application. However, many shared-memory applications do not require more computing cores than available in current motherboards because their scalability is bounded to a few tens of parallel threads. Nevertheless, they may still benefit from having more memory resources. Additionally, the performance of extended systems involving more cores is typically constrained by the glueing coherency protocol, whose overhead lowers the performance of the final system. In this paper we present a 32-node prototype of a new non-coherent distributed-memory architecture for clusters, aimed to provide applications additional memory borrowed from other nodes without providing them more cores, thus avoiding the penalty of maintaining coherency among nodes of the cluster. Results from the execution of real applications in this prototype demonstrate that our proposal truly works, as well as its performance is assessed.
Héctor Montaner, Federico Silla, Holger Fröning, José Duato
HPCC2
2011 MEMSCALETM: A Scalable Environment for Databases
abstract
In this paper we propose a new memory architecture for clusters referred to as MEMSCALE. This architecture provides a distributed non-coherent shared-memory view of the memory resources present in the cluster. With this aggregation technique, a given processor can directly access any memory address located at other nodes in the cluster and, therefore, the whole memory present in the cluster can be granted to a single application. In this study we focus on in-memory databases as a memory-hungry application in order to show the possibilities of our new architecture. To prove the feasibility of our idea, a 16-node prototype cluster serves as a demonstrator. Part of the memory in each node is used to create a global memory pool of 128GB which hosts an entire database. First we show that providing more memory than usually available in a typical commodity node for a database server makes the execution of queries more than one order of magnitude faster than using regular SSD drives. After that, we go one step further and show that simultaneously accessing the database from all the nodes in the cluster converts our prototype into a powerful database server capable of beating current commercial solutions in terms of latency and throughput.
Héctor Montaner, Federico Silla, Holger Fröning, José Duato
HPCC2
2011 Performance of CUDA Virtualized Remote GPUs in High Performance Clusters
abstract
In a previous work we presented the architecture of rCUDA, a middleware that enables CUDA remoting over a commodity network. That is, the middleware allows an application to use a CUDA-compatible Graphics Processor (GPU) installed in a remote computer as if it were installed in the computer where the application is being executed. This approach is based on the observation that GPUs in a cluster are not usually fully utilized, and it is intended to reduce the number of GPUs in the cluster, thus lowering the costs related with acquisition and maintenance while keeping performance close to that of the fully-equipped configuration. In this paper we model rCUDA over a series of high throughput networks in order to assess the influence of the performance of the underlying network on the performance of our virtualization technique. For this purpose, we analyze the traces of two different case studies over two different networks. Using this data, we calculate the expected performance for these same case studies over a series of high throughput networks, in order to characterize the expected behavior of our solution in high performance clusters. The estimations are validated using real 1 Gbps Ethernet and 40 Gbps InfiniBand networks, showing an error rate in the order of 1% for executions involving data transfers above 40 MB. In summary, although our virtualization technique noticeably increases execution time when using a 1 Gbps Ethernet network, it performs almost as efficiently as a local GPU when higher performance interconnects are used. Therefore, the small overhead incurred by our proposal because of the remote use of GPUs is worth the savings that a cluster configuration with less GPUs than nodes reports.
José Duato, Antonio J. Peña, Federico Silla, Rafael Mayo 0002, Enrique S. Quintana-Ortí
ICPP3
2011 Energy and Performance Efficient Thread Mapping in NoC-Based CMPs under Process Variations
abstract
Within-die process variation causes cores, memories, and network resources in NoC-based CMPs to present different speeds and leakage power. In this context, thread mapping strategies that consider the effects of process variability on chip resources arise as a suitable choice to maximize performance while energy consumption constraints are satisfied. However, other factors, as the location of memory controllers and the concurrent execution of several applications in the chip, can bound the possible benefits of such mapping strategies. In this paper we propose a mapping strategy, named as uniform regions, that takes variability effects into account when assigning application threads to cores in the chip. More specifically, uniform regions, in terms of operating frequency, that additionally present the highest available frequency, are selected so that the benefits of such a variation-aware mapping strategy in a NoC-based CMP are maximized. We additionally present two different ways of configuring the frequency and voltage of the cores in the selected region. The first one is intended to provide the maximum performance while keeping energy as low as possible, while the second one is much more for energy-aware. The first one reduces the execution time up to a 23% while reducing the energy up to 24% whereas the second one provides smaller speed ups while reduces energy up to 33%.
Carles Hernández 0001, Federico Silla, José Duato
ICPP2
2011 A Distributed Switch Architecture for On-Chip Networks
abstract
It is well-known that current Chip Multiprocessor (CMP) and high-end MultiProcessor System-on-Chip (MPSoC) designs are growing in their number of components. Networks-on-Chip (NoC) provide the required connectivity for such CMP and MPSoC designs at reasonable costs. However, as technology advances, links become the critical component in the NoC. First, because the power consumption of the link is extremely high with respect the power consumption of the rest of components (mainly switches), becoming unacceptable for long global interconnects. Second, the delay of a link does not scale with technology, thus, degrading the performance of the network. To solve both problems, several solutions have been previously proposed. In this paper, we present a new switch architecture that reduces the negative impact of links on the NoC. We call our proposal distributed switch. The distributed switch moves the circuitry of a standard switch onto the links. Then, packets are buffered, routed, and forwarded at the same time they are crossing the link. Distributing a standard switch onto the link improves the trade off between the power consumption and the operating frequency of the entire network. In contrast, area requirements are increased. The distributed switch reduces up to 14.8% the peak power consumption while increases its area up to 22%. Furthermore, the distributed switch is able to increase the maximum achievable frequency with respect to the standard switch. In particular, the maximum operating frequency of the distributed switch can be increased up to 14.3%.
Antoni Roca 0001, Carles Hernández 0001, José Flich, Federico Silla, José Duato
ICPP4
2011 Characterizing the impact of process variation on 45 nm NoC-based CMPs
Carles Hernández 0001, Antoni Roca 0001, José Flich, Federico Silla, José Duato
J. Parallel Distributed Comput.4
2011 Cost-Efficient On-Chip Routing Implementations for CMP and MPSoC Systems
abstract
The high-performance computing domain is enriching with the inclusion of networks-on-chip (NoCs) as a key component of many-core (CMPs or MPSoCs) architectures. NoCs face the communication scalability challenge while meeting tight power, area, and latency constraints. Designers must address new challenges that were not present before. Defective components, the enhancement of application-level parallelism, or power-aware techniques may break topology regularity, thus, efficient routing becomes a challenge. This paper presents universal logic-based distributed routing (uLBDR), an efficient logic-based mechanism that adapts to any irregular topology derived from 2-D meshes, instead of using routing tables. uLBDR requires a small set of configuration bits, thus being more practical than large routing tables implemented in memories. Several implementations of uLBDR are presented highlighting the tradeoff between routing cost and coverage. The alternatives span from the previously proposed LBDR approach (with 30% of coverage) to the uLBDR mechanism achieving full coverage. This comes with a small performance cost, thus exhibiting the tradeoff between fault tolerance and performance. Power consumption, area, and delay estimates are also provided highlighting the efficiency of the mechanism. To do this, different router models (one for CMPs and one for MPSoCs) have been designed as a proof concept.
Samuel Rodrigo, José Flich, Antoni Roca 0001, Simone Medardoni, Davide Bertozzi, Jesús Camacho Villanueva, Federico Silla, José Duato
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2010 Getting Rid of Coherency Overhead for Memory-Hungry Applications
abstract
Current commercial solutions intended to provide additional resources to an application being executed in a cluster usually aggregate processors and memory from different nodes. In this paper we present a 16-node prototype for a shared-memory cluster architecture that follows a different approach by decoupling the amount of memory available to an application from the processing resources assigned to it. In this way, we provide a new degree of freedom so that the memory granted to a process can be expanded with the memory from other nodes in the cluster without increasing the number of processors used by the program. This feature is especially suitable for memory-hungry applications that demand large amounts of memory but present a parallelization level that prevents them from using more cores than available in a single node. The main advantage of this approach is that an application can use more memory from other nodes without involving the processors, and caches, from those nodes. As a result, using more memory no longer implies increasing the coherence protocol overhead because the number of caches involved in the coherent domain has become independent from the amount of available memory. The prototype we present in this paper leverages this idea by sharing 128GB of memory among the cluster. Real executions show the feasibility of our prototype and its scalability.
Héctor Montaner, Federico Silla, Holger Fröning, José Duato
CLUSTER2
2010 A methodology for the characterization of process variation in NoC links
abstract
Associated with the ever growing integration scales is the increase in process variability. In the context of network-on-chip, this variability affects the maximum frequency that could be sustained by each link that interconnects two cores in a chip multiprocessor. In this paper we present a methodology to model delay variations in NoC links. We also show its application to several technologies, namely 45nm, 32nm, 22nm, and 16nm. Simulation results show that conclusions about variability greatly depend on the implementation context.
Carles Hernández 0001, Federico Silla, José Duato
DATE2
2010 A Latency-Efficient Router Architecture for CMP Systems
abstract
As technology advances, the number of cores in Chip Multi Processor systems (CMPs) and Multi Processor Systems-on-Chips (MPSoCs) keeps increasing. Current test chips and products reach tens of cores, and it is expected to reach hundreds of cores in the near future. Such complexity demands for an efficient network-on-chip (NoC). The common choice to build such networks is the 2D mesh topology (as it matches the regular tile-based design) and the Dimension-Order Routing (DOR) algorithm (because its simplicity). The network in such systems must provide sustained throughput and ultra low latencies. One of the key components in the network is the router, and thus, it plays a major role when designing for such performance levels. In this paper we propose a new pipelined router design focused in reducing the router latency. As a first step we identify the router components that take most of the critical path, and thus limit the router frequency. In particular, the arbiter is the one limiting the performance of the router. Based on this fact, we simplify the arbiter logic by using multiple smaller arbiters. The initial set of requests in the initial arbiter is then distributed over the smaller arbiters that operate in parallel. With this design procedure, and with a proper internal router organization, different router architectures are evolved. All of them enable the use of smaller arbiters in parallel by replicating ports and assuming the use of the DOR algorithm. The net result of such changes is a faster router. Preliminary results demonstrate a router latency reduction ranging from 10% to 21% with an increase of the router area. Network latency is reduced in a range from 11% to 15%.
Antoni Roca 0001, José Flich, Federico Silla, José Duato
DSD3
2010 VCTlite: Towards an efficient implementation of virtual cut-through switching in on-chip networks
abstract
On-chip networks have rapidly emerged as the best interconnection choice for high-core count chip multiprocessors (CMPs) because of the good scalability properties they present. Their fast evolution has been accelerated by the large inheritance from the offchip network domain. Many of the mechanisms and techniques previously developed in that area have been directly applied to the on-chip domain due to the perfect match between the features provided by those techniques and the requirements of on-chip networks. Other mechanisms have been adapted in order to fit the new environment needs. In this paper we present a new example of such an adaptation. Although wormhole switching was initially chosen as the switching mechanism that best fits the on-chip domain characteristics because of its well-known low input buffer requirements, in this paper we show that an efficient implementation of virtual cut-through switching, specially adapted to the particular characteristics of the CMP domain, is feasible as well. Our implementation of virtual cut-through switching, carried out in a 45nm technology, demonstrates to be faster than a wormhole one, at the same time that does not require more area and reduces power consumption.
Antoni Roca 0001, José Flich, Federico Silla, José Duato
HiPC3
2010 A practical way to extend shared memory support beyond a motherboard at low cost
abstract
Improvements in parallel computing hardware usually involve increments in the number of available resources for a given application such as the number of computing cores and the amount of memory. In the case of shared-memory computers, the increase in computing resources and available memory is usually constrained by the coherency protocol, whose overhead rises with system size, limiting the scalability of the final system. In this paper we propose an efficient and cost-effective way to increase the memory available for a given application by leveraging free memory in other computers in the cluster.
Héctor Montaner, Federico Silla, José Duato
HPDC2
2010 Improving the Performance of GALS-Based NoCs in the Presence of Process Variation
abstract
Current integration scales allow designing chip multiprocessors (CMP) where cores are interconnected by means of a network-on-chip (NoC). Unfortunately, the small feature size of current integration scales cause some unpredictability in manufactured devices because of process variation. In NoCs,variability may affect links and routers causing that they do not match the parameters established at design time. In this paper we first analyze the way that manufacturing deviations affect the components of a NoC by applying a comprehensive and detailed variability model to 200 instances of an 8x8 mesh NoC synthesized using 45nm technology. A second contribution of this paper is showing that GALS-based NoCs present communication bottlenecks under process variation. To overcome this performance reduction we draft a novel approach, called performance domains, intended to reduce the negative impact of variability on application execution time. This mechanism is suitable when several applications are simultaneously running in the CMP chip.
Carles Hernández 0001, Antoni Roca 0001, Federico Silla, José Flich, José Duato
NOCS3
2010 Addressing Manufacturing Challenges with Cost-Efficient Fault Tolerant Routing
abstract
The high-performance computing domain is enriching with the inclusion of Networks-on-chip (NoCs) as a key component of many-core (CMPs or MPSoCs) architectures. NoCs face the communication scalability challenge while meeting tight power, area and latency constraints. Designers must address new challenges that were not present before. Defective components, the enhancement of application-level parallelism or power-aware techniques may break topology regularity, thus, efficient routing becomes a challenge.In this paper, uLBDR (Universal Logic-Based Distributed Routing) is proposed as an efficient logic-based mechanism that adapts to any irregular topology derived from 2D meshes, being an alternative to the use of routing tables (either at routers or at end-nodes). uLBDR requires a small set of configuration bits, thus being more practical than large routing tables implemented in memories. Several implementations of uLBDR are presented highlighting the trade-off between routing cost and coverage. The alternatives span from the previously proposed LBDR approach (with 30\% of coverage) to the uLBDR mechanism achieving full coverage. This comes with a small performance cost, thus exhibiting the trade-off between fault tolerance and performance.
Samuel Rodrigo, José Flich, Antoni Roca 0001, Simone Medardoni, Davide Bertozzi, Jesús Camacho Villanueva, Federico Silla, José Duato
NOCS7
2009 A new mechanism to deal with process variability in NoC links
abstract
Associated with the ever growing integration scale of VLSI technologies is the increase in process variability, which makes silicon devices to become less predictable. In the context of network-on-chip (NoC), this variability affects the maximum frequency that could be sustained by each wire of the link that interconnects two cores in a CMP system. Reducing the clock frequency so that all wires can properly work is a trivial solution but, as variability increases, this approach causes an unacceptable performance penalty. In this paper, we propose a new technique to deal with the effects of variability on the links of the NoC that interconnects cores in a CMP system. This technique, called Phit Reduction (PR), retrieves most of the bandwidth still available in links containing wires that are not able to operate at the designed operating frequency. More precisely, our mechanism discards these slow wires and uses all the wires that can work at the design frequency. Two implementations are presented: Local Phit Reduction (LPR), oriented to fabrication processes with very high variability, which requires more hardware but provides higher performance; and Global Phit Reduction (GPR), that requires less additional hardware but is not able to extract all the available bandwidth. The performance evaluation presented in the paper confirms that LPR obtains good results both for low and high variability scenarios. Moreover, in most of our experiments LPR practically achieves the same performance than the ideal network. On the other hand, GPR is appropriate for systems where whithin-die variations are expected to be low.
Carles Hernández 0001, Federico Silla, Vicente Santonja, José Duato
IPDPS2
2008 Network Reconfiguration Suitability for Scientific Applications
abstract
This paper analyzes the communication pattern of several scientific applications and how they can make profit of network reconfiguration in order to adapt network topology to the communication needs so that total execution time is reduced. By using an analysis methodology based on real application executions, we study the variation of the required communication bandwidth with time and also the global interprocedural communication patterns. Results show that required bandwidth between each pair of processes does not significantly fluctuates, leading to a constant use of the links and therefore discouraging dynamic reconfigurations of the network during execution time. Nevertheless, the group of busy links changes with each application showing a different communication graph for each of them. Thus, execution time may be accelerated by using an ad-hoc topology, that is, reconfiguring the network before the execution of the application in order to adapt it to the application needs.
Héctor Montaner, Federico Silla, Vicente Santonja, José Duato
ICPP2
2004 On the development of a communication-aware task mapping technique
Juan M. Orduña, Federico Silla, José Duato
J. Syst. Archit.2
2003 LSOM: A Link State Protocol Over Mac Addresses for Metropolitan Backbones Using Optical Ethernet Switches
abstract
This paper presents a new protocol named "Link State Over MAC" (LSOM) for Optical Ethernet switches to allow the use of active loop topologies, like meshes, in Metropolitan Area Networks (MAN) or even Wide Area Networks (WAN) backbone. In this respect, LSOM is an alternative to a ring topology as proposed in draft IEEE 802.17 Resilient Packet Ring (RPR) or a tree topology using IEEE802. 1D Rapid Spanning Tree Protocol (RSTP). LSOM provides higher scalability and is able to achieve better bandwidth utilization and lower latency than RSTP and RPR. Simulation results for 4-node and 9-node topologies show that LSOM can improve throughput over RPR by a factor of up to 1.7. Furthermore, full freedom to choose any MAN active topology allows an effective use of the available dark fiber resources.
Román García, José Duato, Federico Silla
NCA3
2002 A comparative study of arbitration algorithms for the Alpha 21364 pipelined router
abstract
Interconnection networks usually consist of a fabric of interconnected routers, which receive packets arriving at their input ports and forward them to appropriate output ports. Unfortunately, network packets moving through these routers are often delayed due to conflicting demand for resources, such as output ports or buffer space. Hence, routers typically employ arbiters that resolve conflicting resource demands to maximize the number of matches between packets waiting at input ports and free output ports. Efficient design and implementation of the algorithm running on these arbiters is critical to maximize network performance.This paper proposes a new arbitration algorithm called SPAA (Simple Pipelined Arbitration Algorithm), which is implemented in the Alpha 21364 processor's on-chip router pipeline. Simulation results show that SPAA significantly outperforms two earlier well-known arbitration algorithms: PIM (Parallel Iterative Matching) and WFA (Wave-Front Arbiter) implemented in the SGI Spider switch. SPAA outperforms PIM and WFA because SPAA exhibits matching capabilities similar to PIM and WFA under realistic conditions when many output ports are busy, incurs fewer clock cycles to perform the arbitration, and can be pipelined effectively. Additionally, we propose a new prioritization policy called the Rotary Rule, which prevents the network's adverse performance degradation from saturation at high network loads by prioritizing packets already in the network over new packets generated by caches or memory.
Shubhendu S. Mukherjee, Federico Silla, Peter J. Bannon, Joel S. Emer, Steven Lang, David Webb
ASPLOS2
2001 Improving Network Performance by Efficiently Dealing with Short Control Messages in Fibre Channel SANs
Xavier Molero, Federico Silla, Vicente Santonja, José Duato
Euro-Par2
2001 On the Switch Architecture for Fibre Channel Storage Area Networks
abstract
The fast growth of data intensive applications has caused a change in the traditional storage model. The server-to-disk approach is being replaced by storage area networks (SANs), which enable storage to be externalized from servers, thus allowing storage devices to be shared among multiple servers. Nowadays, the majority of SANs use fibre channel. The standard for fibre channel defines several issues related to the switch interface, but does not make any suggestion about the internal switch architecture to be implemented by manufacturers. We analyze the key architectural switch characteristics for building fibre channel storage area networks. To do so, our starting point is the performance analysis of two different switch architectures, identifying their strongest and weakest points, and thus taking advantage of the best features from both of them. After this first analysis, we introduce several other features in the switch, concluding with a proposed architecture that doubles network throughput while reducing response delay.
Xavier Molero, Federico Silla, Vicente Santonja, José Duato
ICPADS2
2001 On the Interconnection Topology for Storage Area Networks
abstract
Department d’Informatica de Sistemes i Computadors Clusters of workstations are becoming an interesting al-ternative to parallel computers for those applications with high needs of resources such as memory, processing powe< and input/output storage capacity. Also, the fast growth of data intensive applications has caused a change in the tra-ditional storage model. The server-to-disk approach is be-ing replaced by storage area networks (SANS). SANS are a separate network for storage, isolated from the messaging network and optimized for the movement of data between servers and storage devices. Depending on the required network size and the environment targeted for the SAN, different interconnection topologies may be advis-able, affecting both performance and cost. Moreove ~ for a given topology, the routing algorithm used by messages also influences network performance. In this paper we analyze the impact of network topol-ogy on both performance and cost of storage area net-works. This analysi,s is pe~ormedfor up */down * and mini-mal adaptive routin,g in the context of two different environ-ments: buildings and departments. We show that depending on the network topology and the routing scheme, differences in the pe~ormance/cost ratio may increase by a factor of up to 6. Moreove ~ we demonstrate that slightly modifiing the network topology, such as adding a few new links, we can noticeably improve the overall performance without signif-icantly affecting the total cost. 1.
Xavier Molero, Federico Silla, Vicente Santonja, José Duato
IPDPS2
2001 On the Scalability of Topologies for Storage Area Networks in Building Environments
abstract
Nowadays, the fast growth of data intensive applications is changing the way storage is devised. The traditional server-to-disk approach is being replaced by storage area networks (SANs), which are a separate network for storage, isolated from the messaging network and optimized for the movement of data between servers and storage devices (usually disks). We analyze the performance and cost scalability of a family of network topologies devised to be used in building environments. Performance simulation results combined with cost estimations have revealed that slight modifications in network topology can affect the overall scalability. In particular wraparound links connecting the lowest and highest floors in the building significantly affect the scalability of the network. Anyway, the use of this kind of links by itself does not provide the best solution. It is also necessary to have a good interconnection pattern in the backbone.
Xavier Molero, Federico Silla, Vicente Santonja, José Duato
NCA2
2001 A Comparison of Router Architectures for Virtual Cut-Through and Wormhole Switching in a NOW Environment
José Duato, Antonio Robles, Federico Silla, Ramón Beivide
J. Parallel Distributed Comput.3
2000 Modeling and Simulation of Storage Area Networks
abstract
Storage area networks (SANs) are an emerging data communications platform which interconnects servers and storage devices (such as disks, disk arrays, and tape drives) to create a pool of storage that users can access directly. This networking approach reports benefits such as computer clustering, topological flexibility, fault tolerance, high availability, and remote management. In order to evaluate the performance of these systems it is necessary to have the adequate tools. Usually, performance evaluation may be based on analytical modeling or simulation. Each of them differs in their scope and applicability. However the simulation modeling technique offers more freedom, flexibility, and accuracy than the analytical methods. Thus, when evaluating the performance of SANs, simulation modeling should be used. In this paper the issues involved in the modeling and design of a very flexible and easy to use SAN simulator are presented. This tool is able to consider among others, both real-world I/O traces and synthetic I/O traffic, message packetization, faults in links and switches, virtual channels, different routing algorithms, etc. We describe its main internal organization, the basic modeling mechanisms the simulator is based on, the main input parameters and output performance variables. Also, the analysis of preliminary results using I/O traces is presented, showing that the storage network increases self-similarity of the traffic received by servers, latency variations are more important for control messages than for data messages, and links have a low utilization.
Xavier Molero, Federico Silla, Vicente Santonja, José Duato
MASCOTS2
2000 High-Performance Routing in Networks of Workstations with Irregular Topology
abstract
Networks of workstations are rapidly emerging as a cost-effective alternative to parallel computers. Switch-based interconnects with irregular topology allow the wiring flexibility, scalability, and incremental expansion capability required in this environment. However, the irregularity also makes routing and deadlock avoidance on such systems quite complicated. In current proposals, many messages are routed following nonminimal paths, increasing latency and wasting resources. In this paper, we propose two general methodologies for the design of adaptive routing algorithms for networks with irregular topology. Routing algorithms designed according to these methodologies allow messages to follow minimal paths in most cases, reducing message latency and increasing network throughput. As an example of application, we propose two adaptive routing algorithms for ANI (previously known as Autonet). They can be implemented either by duplicating physical channels or by splitting each physical channel into two virtual channels. In the former case, the implementation does not require a new switch design. It only requires changing the routing tables and adding links in parallel with existing ones, taking advantage of spare switch ports. In the latter case, a new switch design is required, but the network topology is not changed. Evaluation results for several different tapologies and message distributions show that the new routing algorithms are able to increase throughput for random traffic by a factor of up to 4 with respect to the original up*/down* algorithm, also reducing latency significantly. For other message distributions, throughput is increased more than seven times. We also show that most of the improvement comes from the use of minimal routing.
Federico Silla, José Duato
IEEE Trans. Parallel Distributed Syst.1
2000 On the Use of Virtual Channels in Networks of Workstations with Irregular Topology
abstract
Networks of workstations are becoming increasingly popular as a cost-effective alternative to parallel computers. Typically, these networks connect workstations using irregular topologies, providing the wiring flexibility, scalability, and incremental expansion capability required in this environment. Recently, we proposed two methodologies for the design of adaptive routing algorithms for networks with irregular topology, as well as fully adaptive routing algorithms for these networks. These algorithms increase throughput considerably with respect to previously existing ones, but require the use of at least two virtual channels. In this paper, we propose a very efficient flow control protocol to support virtual channels when link wires are very long and/or have different lengths. This flow control protocol relies on the use of channel pipelining and control flits. Control traffic is minimized by assigning physical bandwidth to virtual channels until the corresponding message blocks or it is completely transmitted. Simulation results show that this flow control protocol performs as efficiently as an ideal network with short wires and flit-by-flit multiplexing. The effect of additional virtual channels per physical channel has also been studied, revealing that the optimal number of virtual channels varies with network size. The use of virtual channel priorities is also analyzed. The proposed flow control protocol may increase short message latency, due to long messages monopolizing channels and hindering the progress of short messages. Therefore, we have analyzed the impact of limiting the number of flits (block size) that a virtual channel may forward once it gets the link. Simulation results show that limiting the maximum block size causes the overall network performance to decrease.
Federico Silla, José Duato
IEEE Trans. Parallel Distributed Syst.1
1998 Virtual channel multiplexing in networks of workstations with irregular topology
abstract
Networks of workstations are becoming a cost-effective alternative for small-scale parallel computing. Although they may not provide the closely coupled environment of multicomputers and multiprocessors, they meet the needs of a great variety of parallel computing problems at a lower cost. However in order to achieve a high efficiency, the interconnects used to build the network of workstations must provide a very high bandwidth and low latencies, making their design a critical issue. Recently, a very efficient flow control protocol for networks of workstations has been proposed by the authors. This protocol multiplexes physical channels between several virtual channels and minimizes the use of control flits by transmitting several data flits each time a virtual channel gets the link. In this protocol, a virtual channel sends data flits until the message blocks or is completely transmitted. However it can reduce network throughput, by increasing short message latency, due to long messages monopolizing channels and hindering the progress of short messages. In this paper, we analyze the impact of limiting the number of flits (block size) that a virtual channel can send once it gets the link. We propose a new version of the previous flow control protocol that is easily, implementable on hardware. Simulation results show that limiting the maximum block size is not a good design decision, because the overall network performance decreases. Only when short message latency is crucial is it is acceptable to limit the block size.
Federico Silla, José Duato, Anand Sivasubramaniam, Chita R. Das
HiPC1
1998 Impact of Adaptivity on the Behaviour of Networks of Workstations under Bursty Traffic
abstract
Networks of workstations (NOWs) are becoming increasingly popular as an alternative to parallel computers. Typically, these networks present irregular topologies, providing the wiring flexibility, scalability, and incremental expansion capability required in this environment. Similar to the evolution of parallel computers, NOWs are also evolving from distributed memory to shared memory. However distances between processors are longer in NOWs, leading to higher message latency and lower network bandwidth. Therefore, one can expect the network to be a bottleneck when executing some parallel applications on a NOW supporting a shared-memory programming paradigm. The authors analyze whether the interconnection network in a NOW is able to efficiently handle the traffic generated in a DSM with the same number of processors. They evaluate the behavior of a NOW using application traces captured during the execution of several SPLASH2 applications on a DSM simulator. They show through simulation that the adaptive routing algorithm previously proposed by them almost eliminates network saturation due to its ability to support a higher sustained throughput. Therefore, adaptive routing becomes a key design issue to achieve similar performance in NOWs and tightly-coupled DSMs.
Federico Silla, Manuel P. Malumbres, José Duato, Donglai Dai, Dhabaleswar K. Panda 0001
ICPP1
1998 Improving Performance of Networks of Workstations by using Disha Concurrent
abstract
Networks of workstations are currently emerging as a cost-effective alternative to parallel computers. Recently, deadlock recovery techniques have been shown to be an alternative to deadlock avoidance. Disha Concurrent is a progressive deadlock recovery scheme able to simultaneously redirect several deadlocked messages through a deadlock-free lane. Unlike deadlock avoidance techniques, Disha provides true fully adaptive routing without using virtual channels to guarantee deadlock freedom. In this paper, we analyze the application of Disha to networks of workstations. We propose an implementation of Disha on irregular networks that allows concurrent deadlock recovery proving that this implementation is always able to recover from deadlock. A new switch organization and a new flow control protocol are proposed to support Disha. Performance evaluation results show that applying Disha to irregular networks increases network throughput by a factor of up to 3.5, and also reduces latency with regard to other routing algorithms based on deadlock avoidance techniques.
Federico Silla, Antonio Robles, José Duato
ICPP1
1997 Improving the efficiency of adaptive routing in networks with irregular topology
abstract
Networks of workstations are emerging as a cost-effective alternative to parallel computers. The interconnection between workstations usually relies on switch-based networks with irregular topologies. This irregularity makes routing and deadlock avoidance quite complicated. Current proposals avoid deadlock by removing cyclic dependencies between channels and therefore, many messages are routed along non-minimal paths, increasing latency and wasting resources. We propose a general methodology for the design of adaptive routing algorithms for networks with irregular topology that improves a previously proposed one by reducing the probability of routing over non-minimal paths. The resulting routing algorithms allow messages to follow minimal paths in most cases, reducing message latency and increasing network throughput. As an example of application, we propose an improved adaptive routing algorithm for Autonet.
Federico Silla, José Duato
HiPC1