VLDB 2026 Research / reviewers in the wild / expert
Francisco J. Alfaro
dblp:115/4710 · also Francisco J. Alfaro-Cortes, Francisco J. Alfaro-Cortés, Francisco José Alfaro
· DBLP profile ↗
65ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0002-4430-4482ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 56 · 5 first-author · 10 since 2021Computer networks · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On the power saving in high-speed Ethernet-based networks for supercomputers and data centersabstractThe increase in computation and storage has led to a significant growth in the scale of systems powering applications and services, raising concerns about sustainability and operational costs. In this paper, we explore power-saving techniques in high-performance computing (HPC) and data center networks, and their relation with performance degradation. From this premise, we propose leveraging the Energy Efficient Ethernet (EEE) protocol, with the flexibility to extend to conventional Ethernet or upcoming Ethernet-derived interconnect versions of BXI and Omnipath. We analyze the PerfBound power-saving mechanism, identifying possible improvements and modeling it into a simulation framework. Through different experiments, we examine its impact on performance and determine the most appropriate interconnect. We also study traffic patterns generated by selected HPC and machine learning applications to evaluate the behavior of power-saving techniques. From these experiments, we provide an analysis of how applications affect system and network energy consumption. Based on this, we disclose the weakness of dynamic power-down mechanisms and propose an approach that improves energy reduction with minimal or no performance penalty. This work presents a thorough analysis of PerfBound and an enhancement to the technique, while also targeting emerging post-exascale networks. Miguel Sánchez de la Rosa, Francisco J. Andujar, Jesús Escudero-Sahuquillo, José L. Sánchez 0002, Francisco J. Alfaro |
J. Syst. Archit. | 5 |
| 2025 | Quality-of-service provision for BXIv3-based interconnection networksabstractAbstract Supercomputers (SCs) enable advanced research for a variety of scientific fields, and data centers (DCs) power our day-to-day services. These two massive systems work at scales, in terms of storage and computing power, which are not comparable to our everyday devices. As such, they require state-of-the-art technology to constantly evolve and meet our increasing demand. The interconnection network is the backbone of these systems, since it must provide efficient communication among the nodes that compose the whole system, otherwise becoming the entire system bottleneck. As multiple applications and services may use subsets of the system at the same time, interconnection networks must prevent excessive degradation for latency-sensitive applications. To this end, differentiated services are used to provide fair network access that considers bandwidth and latency requirements for each application. In this paper, we extend the switch architecture of next-generation BXI networks (hereafter called BXIv3) to incorporate arbitration tables so these networks can provide quality of service (QoS) to applications and services running on both SCs and DCs. Our proposal has been implemented in a network simulator, which models the behavior of a BXIv3 network. We have used several traffic patterns and arbitration table configurations to conduct a set of simulation experiments for the evaluation of our solution. The obtained results show that our proposal achieves accurate bandwidth allocation with differentiated latencies. Moreover, a study of memory requirements shows that our solution is quite feasible for hardware implementation. Miguel Sánchez de la Rosa, Gabriel Gomez-Lopez, Francisco J. Andujar, Jesús Escudero-Sahuquillo, José L. Sánchez 0002, Francisco J. Alfaro, Pierre-Axel Lagadec |
J. Supercomput. | 6 |
| 2024 | Quality-of-Service Provision for BXI3-Based Interconnection NetworksabstractThe ever-increasing demand for computational power and storage capacity has led to massive Supercomputers and Data Centers running highly parallel applications and services commonly utilized in fields such as Physics, Biology, Robotics, Medicine, or generative AI. The interconnection network is the backbone of these systems since it allows processing and storage nodes to communicate with high bandwidth and low latency, otherwise becoming the entire system bottleneck. These systems commonly run several applications simultaneously, which may have specific network requirements due to technical or contractual reasons. Indeed, the traffic flows from different applications are expected to need different bandwidth and latency levels predetermined before execution. Therefore, Quality of Service (QoS) has been a recurrent design aspect for high-performance interconnection networks, as it happens for different technologies, such as Slingshot (Cray) or InfiniBand (NVIDIA). In this regard, as far as we know, no proposals have been made yet to provide applications with QoS for the upcoming generation of BXI (BXI3). We propose using arbitration tables to assign different priorities to packets when they are injected by NICs or forwarded by switches. We have conducted simulation experiments to evaluate our proposal, comparing several QoS configurations using synthetic workloads. The obtained results show that the proposed QoS approach is efficient and feasible so that it can be applied to the upcoming BXI3. Miguel Sánchez de la Rosa, Gabriel Gomez-Lopez, Francisco J. Andujar, Jesús Escudero-Sahuquillo, José L. Sánchez 0002, Francisco J. Alfaro, Pierre-Axel Lagadec |
HOTI | 6 |
| 2023 | Energy efficient HPC network topologies with on/off linksabstractEnergy efficiency is a must in today HPC systems. To achieve this goal, a holistic design based on the use of power-aware components should be performed. One of the key components of an HPC system is the high-speed interconnect. In this paper, we compare and evaluate several design options for the interconnection network of an HPC system, including torus, fat-trees and dragonflies. State of the art low power modes are also used in the interconnection networks. The paper does not only consider energy efficiency at the interconnection network level but also at the system as a whole. The analysis is performed by using a simple yet realistic power model of the system. The model has been adjusted using actual power consumption values measured on a real system. Using this model, realistic multi-job trace-based workloads have been used, obtaining the execution time and energy consumed. The results are presented to ease choosing a system, depending on which parameter, performance or energy consumption, receives the most importance. Francisco J. Andujar, Salvador Coll, Marina Alonso, Juan-Miguel Martinez-Rubio, Pedro López 0001, José L. Sánchez 0002, Francisco J. Alfaro |
Future Gener. Comput. Syst. | 7 |
| 2022 | RED-SEA: Network Solution for Exascale ArchitecturesabstractIn order to enable Exascale computing, next generation interconnection networks must scale to hundreds of thousands of nodes, and must provide features to also allow the HPC, HPDA, and AI applications to reach Exascale, while benefiting from new hardware and software trends. RED-SEA will pave the way to the next generation of European Exascale interconnects, including the next generation of BXI, as follows: (i) specify the new architecture using hardware-software co-design and a set of applications representative of the new terrain of converging HPC, HPDA, and AI; (ii) test, evaluate, and/or implement the new architectural features at multiple levels, according to the nature of each of them, ranging from mathematical analysis and modeling, to simulation, or to emulation or implementation on FPGA testbeds; (iii) enable seamless communication within and between resource clusters, and therefore development of a high-performance low latency gateway, bridging seamlessly with Ethernet; (iv) add efficient network resource management, thus improving congestion resiliency, virtualization, adaptive routing, collective operations; (v) open the interconnect to new kinds of applications and hardware, with enhancements for end-to-end network services - from programming models to reliability, security, low- latency, and new processors; (vi) leverage open standards and compatible APIs to develop innovative reusable libraries and Fabrics management solutions. Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Pier Stanislao Paolucci, Elena Pastorelli, Francesco Simula, Matteo Turisini, Piero Vicini, Roberto Ammendola, Pascale Bernier-Bruna, Said Derradji, Stéphane Guez, Pierre-Axel Lagadec, Gregoire Pichon, Etienne Walter, Gaetan De Gassowski, Matthieu Hautreaux, Stephane Mathieu, Gilles Moreau, Marc Pérache, Hugo Taboada, Torsten Hoefler, Timo Schneider, Matteo Barnaba, Giuseppe Piero Brandino, Francesco De Giorgi, Matteo Poggi, Iakovos Mavroidis, Ioannis Papaefstathiou, Nikolaos Tampouratzis, Benjamin Kalisch, Ulrich Krackhardt, Mondrian Nüssle, Pantelis Xirouchakis, Vangelis Mageiropoulos, Michalis Gianioudis, Harisis Loukas, Aggelos Ioannou, Nikolaos D. Kallimanis, Nikolaos Chrysos, Manolis Katevenis, Wolfgang Frings, Dominik Gottwald, Felime Guimaraes, Max Holicki, Volker Marx, Yannik Müller, Carsten Clauss, Hugo Falter, Xu Huang 0010, Jennifer Lopez Barillao, Thomas Moschny, Simon Pickartz, Francisco J. Alfaro, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José L. Sánchez 0002, Adrián Castelló 0001, Jose Duro, María Engracia Gómez, Enrique S. Quintana-Ortí, Julio Sahuquillo, Eugenio Stabile |
DSD | 58 |
| 2022 | Improving Congestion Control through Fine-Grain Monitoring of InfiniBand NetworksabstractCongestion situations are a serious threat to the performance of the interconnection networks of High-Performance Computing and Data-Center systems. Hence, the specifications of the main interconnect technologies, such as InfiniBand, define some mechanisms to deal with congestion and its effects. However, these standard mechanisms may not be suitable to detect or track accurately the actual status of network congestion, as congestion dynamics indeed can be very complex and varied. Moreover, achieving an optimal configuration of the parameters that drive the different functionalities of congestion-control mechanisms is often a difficult task, as some configurations may be suitable for some traffic scenarios, but not for others. In this paper, we propose combining an existing light-weight platform monitoring tool (LIMITLESS) with the InfiniBand control software (OpenSM), such that the metrics about communication volumes in the network provided by the former allow the latter having a more precise image of congestion status, then being able to react more efficiently in these situations. The main contributions of this paper are the methodology to link the monitor and OpenSM, as well as modifications in the InfiniBand standard congestion-control mechanism so that its reaction is modulated based on the enhanced knowledge about congestion provided by the monitor. These improvements are ready to be integrated into any InfiniBand-based system. According to the results from our experiments (performed in a real InfiniBand-based cluster where we run a widely used benchmark), the proposed approach reduces significantly the number of wrong detections of congestion, and so the number of times that the congestion-control mechanisms react unnecessarily, hence improving system performance up to 74%. The overhead of this monitoring tool is 0.1% in our experiments, collecting data each 200ms. Alberto Cascajo, Gabriel Gomez-Lopez, Jesús Escudero-Sahuquillo, Pedro Javier García, David E. Singh, Francisco J. Alfaro, Francisco J. Quiles 0001, Jesús Carretero 0001 |
HOTI | 6 |
| 2022 | Providing quality of service in omni-path networksabstractAbstract New hierarchical crossbar switch architectures, such as Omni-Path (OPA) and Cray X2, have appeared to improve packet latency, reduce overall cost and increase fault tolerance of the high-performance interconnection networks in supercomputing and data center systems. These and other interconnect technologies (Infiniband or 40/100 Gigabit Ethernet) include support to provide quality of service (QoS) to the applications. In this paper, we show how this QoS support can be enabled to achieve bandwidth and/or latency differentiation in Omni-Path interconnection networks, as a representative case of hierarchical switches. To do that, three different table-based schedulers are used. We include the description of these schedulers and a comparative study by using the results obtained when we evaluate them with Hiperion, a simulation tool that implements an OPA model. Javier Cano-Cano, Francisco J. Andujar, Francisco J. Alfaro, José L. Sánchez 0002, Gaspar Mora |
J. Supercomput. | 3 |
| 2021 | QoS provision in hierarchical and non-hierarchical switch architectures
Javier Cano-Cano, Francisco J. Andujar, Francisco J. Alfaro, José L. Sánchez 0002 |
J. Parallel Distributed Comput. | 3 |
| 2021 | A methodology to enable QoS provision on InfiniBand hardware
Javier Cano-Cano, Francisco J. Andujar, Jesús Escudero-Sahuquillo, Francisco J. Alfaro, José L. Sánchez 0002 |
J. Supercomput. | 4 |
| 2021 | UPR: deadlock-free dynamic network reconfiguration by exploiting channel dependency graph compatibility
Juan-José Crespo, José L. Sánchez 0002, Francisco J. Alfaro, José Flich, José Duato |
J. Supercomput. | 3 |
| 2019 | Constructing virtual 5-dimensional tori out of lower-dimensional network cardsabstractSummary In the Top500 and Graph500 lists of the last years, some of the most powerful systems implement a torus topology to interconnect the millions of computing nodes they include. Some of these torus networks are of five or six dimensions, which implies an additional difficulty as the node degree increases. In previous works, we proposed and evaluated the nD Twin (nDT) torus topology to virtually increase the dimensions a torus is able to implement. We showed that this new topology reduces the distances between nodes, increasing, therefore, global network performance. In this work, we present how to build a 5DT torus network using a specific commercial 6‐port network card (EXTOLL card) to interconnect those nodes. We show, using the same number of cards, that the performance of the 5DT torus network we are able to implement using our proposal is higher than the performance of the 3D torus network for the same number of compute nodes. Francisco J. Andujar, Juan A. Villar, José L. Sánchez 0002, Francisco J. Alfaro, José Duato, Holger Fröning |
Concurr. Comput. Pract. Exp. | 4 |
| 2019 | Silicon photonic networks: Signal loss and power challengesabstractSummary Exascale systems are in need for alternative interconnection technologies. Electrical interconnects are not likely to scale well to a large number of computing nodes in terms of energy efficiency and latency. Silicon photonic networks stand as the main alternative to solve this problem, but are we there yet? In this paper, we exhibit some challenges to be solved for this technology to become a viable solution. Signal loss sources play a critical role in photonic network designs as they restrict the ability to perform data transmission in an effective way. Juan-José Crespo, José L. Sánchez 0002, Francisco J. Alfaro |
Concurr. Comput. Pract. Exp. | 3 |
| 2019 | Energy efficient torus networks with on/off links
Francisco J. Andujar, Salvador Coll, Marina Alonso, Juan-Miguel Martinez-Rubio, Pedro López 0001, José L. Sánchez 0002, Francisco J. Alfaro, Raúl Martínez |
J. Parallel Distributed Comput. | 7 |
| 2019 | Speeding up exascale interconnection network simulations with the VEF3 trace framework
Javier Cano-Cano, Francisco J. Andujar, Francisco J. Alfaro, José L. Sánchez 0002 |
J. Parallel Distributed Comput. | 3 |
| 2019 | Combining Source-adaptive and Oblivious Routing with Congestion Control in High-performance Interconnects using Hybrid and Direct TopologiesabstractHybrid and direct topologies are cost-efficient and scalable options to interconnect thousands of end nodes in high-performance computing (HPC) systems. They offer a rich path diversity, high bisection bandwidth, and a reduced diameter guaranteeing low latency. In these topologies, efficient deterministic routing algorithms can be used to balance smartly the traffic flows among the available routes. Unfortunately, congestion leads these networks to saturation, where the HoL blocking effect degrades their performance dramatically. Among the proposed solutions to deal with HoL blocking, the routing algorithms selecting alternative routes, such as adaptive and oblivious, can mitigate the congestion effects. Other techniques use queues to separate congested flows from non-congested ones, thus reducing the HoL blocking. In this article, we propose a new approach that reduces HoL blocking in hybrid and direct topologies using source-adaptive and oblivious routing. This approach also guarantees deadlock-freedom as it uses virtual networks to break potential cycles generated by the routing policy in the topology. Specifically, we propose two techniques, called Source-Adaptive Solution for Head-of-Line Blocking Avoidance (SASHA) and Oblivious Solution for Head-of-Line Blocking Avoidance (OSHA). Experiment results, carried out through simulations under different traffic scenarios, show that SASHA and OSHA can significantly reduce the HoL blocking. Pedro Yébenes, Jose Rocher-Gonzalez, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Alfaro, Francisco J. Quiles 0001, Crispín Gómez Requena, José Duato |
ACM Trans. Archit. Code Optim. | 5 |
| 2017 | Applying search algorithms to obtain the optimal configuration of nDT torus nodesabstractSummary An nDT torus is a topology where each node comprises 2 identical (n+1)‐port communication cards interconnected by 1 port. By using the current switches or communication cards, this node architecture allows to build torus networks having a greater number of dimensions than networks including only 1 card per node. There are multiple ways to use the ports of the 2 cards to connect a node to other nodes on the nDT torus, and therefore, checking all the configurations is only an affordable problem for small values of n. In this paper, we use artificial intelligence and data mining techniques to obtain the optimal port configuration of all the nodes in the network. We include a performance evaluation that shows nDT torus effectively increases the performance compared with the equivalent torus in resources, with synthetic and application trace–based workloads. We also apply these techniques to 3DT and 5DT tori to confirm the increase in the number of dimensions that does not affect to performance of the nDT torus. Francisco J. Andujar, Juan A. Villar, José L. Sánchez 0002, Francisco J. Alfaro |
Concurr. Comput. Pract. Exp. | 4 |
| 2017 | Providing differentiated services, congestion management, and deadlock freedom in dragonfly networks with adaptive routingabstractSummary The number of endnodes in high‐performance computing systems has grown significantly in the last years. Hence, the interconnection network has become an essential issue as it may end up being the system bottleneck if it is not properly designed. In that sense, the Dragonfly topology has become very popular for interconnecting high‐performance computing systems in the last years because it offers high performance at an affordable cost. However, when using deterministic minimal‐path routing, this topology is not able to offer a high performance under certain traffic conditions. This problem can be solved by using oblivious or adaptive routing. However, there are no congestion management techniques specially tailored to Dragonfly topologies using oblivious or adaptive routing. Note that in congestion situations, the Dragonfly performance may drop because of the head‐of‐line blocking effect. This effect could be even more dangerous in systems where several applications with different priorities coexist. In this work we propose several techniques especially designed for providing differentiated services and congestion management in Dragonfly networks using oblivious or adaptive routing. First, we propose thehierarchical 3‐level queuingqueuing scheme, which configures several virtual channels distributed into 3 virtual networks to reduce the head‐of‐line blocking while deadlocks derived from the routing algorithm are prevented. Second, we extendhierarchical 3‐level queuingto provide differentiated services through 2 different solutions. Finally, some experiments are performed to show the benefits obtained by using the proposed techniques. Pedro Yébenes, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Alfaro, Francisco J. Quiles 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2017 | Providing differentiated services, congestion management, and deadlock freedom in dragonfly networks with adaptive routingabstractIn this article,1 an error in one of the author names has been found subsequent to the publication. “Jesus Escudero-Sahuquilllo” should be “Jesus Escudero-Sahuquillo.” The correct name is now presented above. The author's name has also been corrected in the original published article. Pedro Yébenes, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Alfaro, Francisco J. Quiles 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2016 | Adaptive Routing for N-Dimensional Twin TorusabstractTorus topology is one of the most common topologies used in the current largest supercomputers due to its properties related to cost, implementation or scalability. N-dimensional twin torus (nDT) topology has been proposed to increase the number of dimensions of the torus networks when port-limited low cost expansion cards are available. These topologies have been characterized and evaluated considering only deterministic routing. Adaptive routing algorithms improve communication performance exploiting the path diversity of the torus networks. Due to the particular properties of the nDT torus, designing an adaptive routing algorithm presents a challenge. The peculiarities of the internal link, which interconnects the two communication cards of an nDT torus node, complicate the design of the adaptive routing. In this paper, we study these peculiarities and propose an adaptive routing for nDT tori. Moreover, we show that, by using cards with the same number of ports, we can improve the network performance by building an adaptive nDT torus instead of an adaptive nD torus. Francisco J. Andujar, Juan A. Villar, José L. Sánchez 0002, Francisco J. Alfaro, José Duato |
IEEE Trans. Computers | 4 |
| 2016 | An open-source family of tools to reproduce MPI-based workloads in interconnection network simulators
Francisco J. Andujar, Juan A. Villar, Francisco J. Alfaro, José L. Sánchez 0002, Jesús Escudero-Sahuquillo |
J. Supercomput. | 3 |
| 2015 | VEF Traces: A Framework for Modelling MPI Traffic in Interconnection Network SimulatorsabstractSimulation is often used to evaluate the behaviour and measure the performance of computing systems. Specifically, in high-performance interconnection networks, the simulation has been extensively considered to verify the behaviour of the network itself and to evaluate its performance. In this context, network simulation must be fed with network traffic, also referred to as network workload, whose nature has been traditionally synthetic. These workloads can be used for the purpose of driving studies on network performance, but often such workloads are not accurate enough if a realistic evaluation is pursued. For this reason, other non-synthetic workloads have gained popularity over last decades since they are best to capture the realistic behaviour of existing applications. In this paper, we present the VEF traces framework, a self-related trace model, and all their associated tools. The main novelty of this framework is that, unlike existing ones, it does not provide a network simulation framework, but only offers an MPI task simulation framework, which allows one to use the MPI-based network traffic by any third-party network simulator, since this framework does not depend on any specific simulation platform. Francisco J. Andujar, Juan A. Villar, José L. Sánchez 0002, Francisco J. Alfaro, Jesús Escudero-Sahuquillo |
CLUSTER | 4 |
| 2015 | N-Dimensional Twin Torus TopologyabstractTorus topology is one of the preferred topologies for the interconnection network in high-performance clusters and supercomputers. Cost and scalability are some of the properties that make torus suitable for systems with a large number of nodes. The 3D torus is the version more extended due to its excellent nearest neighbor. However, some of the last supercomputers have been built using a torus network with five or six dimensions. To obtain an nD torus, 2n ports per node are needed, which can be offered by a single or several cards per node. In the second case, there are multiple ways of assigning the dimension and direction of the card ports. In previous work we defined and characterized the 3D Twin (3DT) torus which uses two four-port cards per node. In this paper we extend that previous work to define the n-dimensional Twin (nDT) torus topology. In this case, we formally obtain the optimal port configuration when (n + 1)-port cards are used instead of 2n-port cards. Moreover, we explain how deadlock problem can appear and propose a simple solution. Finally, we include evaluation results which show performance increases when an nDT torus is used instead of an nD torus with fewer dimensions and with the same computational resources. Francisco J. Andujar, Juan A. Villar, José L. Sánchez 0002, Francisco J. Alfaro, José Duato |
IEEE Trans. Computers | 4 |
| 2015 | Optimizing the configuration of combined high-radix switches
Juan A. Villar, Francisco J. Andujar, Francisco J. Alfaro, José L. Sánchez 0002, José Duato |
J. Supercomput. | 3 |
| 2014 | Combining HoL-blocking avoidance and differentiated services in high-speed interconnectsabstractCurrent high-performance platforms such as Datacenters or High-Performance Computing systems rely on highspeed interconnection networks able to cope with the ever-increasing communication requirements of modern applications. In particular, in high-performance systems that must offer differentiated services to applications which involve traffic prioritization, it is almost mandatory that the interconnection network provides some type of Quality-of-Service (QoS) and Congestion-Management mechanism in order to achieve the required network performance. Most current QoS and Congestion-Management mechanisms for high-speed interconnects are based on using the same kind of resources, but with different criteria, resulting in disjoint types of mechanisms. By contrast, we propose in this paper a novel, straightforward solution that leverages the resources already available in InfiniBand components (basically Service Levels and Virtual Lanes) to provide both QoS and Congestion Management at the same time. This proposal is called CHADS (Combined HoL-blocking Avoidance and Differentiated Services), and it could be applied to any network topology. From the results shown in this paper for networks configured with the novel, cost-efficient KNS hybrid topology, we can conclude that CHADS is more efficient than other schemes in reducing the interferences among packet flows that have the same or different priorities. Pedro Yébenes, Jesús Escudero-Sahuquillo, Crispín Gómez Requena, Pedro Javier García, Francisco J. Alfaro, Francisco J. Quiles 0001, José Duato |
HiPC | 5 |
| 2014 | Optimal Configuration for N-Dimensional Twin Torus NetworksabstractTorus topology is one of the most common topologies used in the current largest supercomputers. Although 3D torus is widely used, recently some supercomputers in the Top500 list have been built using networks with topologies of five or six dimensions. To obtain an nD torus, 2n ports per node are needed. These ports can be offered by a single or several cards per node. In the second case, there are multiple ways of assigning the dimension and direction of the card ports. In a previous work we proposed the 3D Twin (3DT) torus which uses two 4-port cards per node, and obtained the optimal port configuration. This paper extends and generalizes that work in order to obtain the optimal port configuration when n dimensions are considered. Thus, the nDT torus topology is presented and defined, and a detailed formal analysis leads to the optimal port configuration. Finally, performance results are included. Francisco J. Andujar, Juan A. Villar, José L. Sánchez 0002, Francisco J. Alfaro, José Duato |
NCA | 4 |
| 2014 | Building 3D Torus Using Low-Profile Expansion CardsabstractTorus is a subclass of direct topologies that was defined in theory to support${\mbi {n}}$dimensions. Although recently some supercomputers have been built on a network with five and six dimensions, the most common case is when only three dimensions are implemented. In the market, there are low-profile communication expansion cards that have a reduced number of ports which is not enough to build tori of a certain number of dimensions. In this paper, we will deal with four-port expansion cards. By means of one of these cards per node, a 2-D torus topology could be built, but not a 3-D torus topology. However, two of these cards could be used to build each node of a 3-D torus topology. In this case, two ports are used to interconnect both cards each other, and the other six ports to connect to six neighbor nodes in the 3-D torus. Theoretically, there are several ways of assigning the dimension and direction of the ports. This paper presents a detailed study of the possible port configurations, and under specific network conditions, the best of them is obtained. Francisco J. Andujar, Juan A. Villar, José L. Sánchez 0002, Francisco J. Alfaro, José Duato |
IEEE Trans. Computers | 4 |
| 2014 | Formalization and configuration methodology for high-radix combined switches
Juan A. Villar, Francisco J. Andujar, Francisco J. Alfaro, José L. Sánchez 0002, José Duato |
J. Supercomput. | 3 |
| 2013 | Obtaining the optimal configuration of high-radix Combined switches
Juan A. Villar, Francisco J. Andujar, José L. Sánchez 0002, Francisco J. Alfaro, José A. Gámez 0001, José Duato |
J. Parallel Distributed Comput. | 4 |
| 2013 | A complete self-testing and self-configuring NoC infrastructure for cost-effective MPSoCsabstractNetworks-on-chip need to survive to manufacturing faults in order to sustain yield. An effective testing and configuration strategy however implies two opposite requirements. One one hand, a fast and scalable built-in self-testing and self-diagnosis procedure has to be carried out concurrently at NoC switches. On the other hand, programming the NoC routing mechanism to go around faulty links and switches can be optimally performed by a centralized controller with global network visibility. To the best of our knowledge, this article proposes for the first time a global network testing and configuration strategy that meets the opposite requirements by means of a fault-tolerant dual network architecture and a fast configuration algorithm for the most common failure patterns. Experimental results report an area overhead as low as 12.5% with respect to the baseline switch architecture while achieving a high degree of fault tolerance. In fact, even when multiple stuck-at faults are considered, the capability of fault masking by the dual network is always over 80%, and the support for multiple link failures is more than 90% in presence of two unusable links in the main network with minimum set-up times. Alberto Ghiribaldi, Daniele Ludovici, Francisco Triviño, Alessandro Strano, José Flich, José L. Sánchez 0002, Francisco J. Alfaro, Michele Favalli, Davide Bertozzi |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2013 | An integrated solution for QoS provision and congestion management in high-performance interconnection networks using deterministic source-based routing
Juan A. Villar, Pedro Javier García, Francisco J. Alfaro, José L. Sánchez 0002, Francisco J. Quiles 0001 |
J. Supercomput. | 3 |
| 2012 | Exploring NoC Virtualization Alternatives in CMPsabstractChip Multiprocessor systems (CMPs) contain more and more cores in every new generation. However, applications for these systems do not scale at the same pace. Thus, in order to obtain a good utilization several applications will need to coexist in the system and in those cases virtualization of the CMP system will become mandatory. In this paper we analyze two virtualization strategies at NoC-level aiming to isolate the traffic generated by each application to reduce or even eliminate interferences among messages belonging to different applications. The first model handles most interferences among messages with a virtual-channels (VCs) implementation minimizing both execution time and network latency. However, using VCs results in area and power overhead due to the cost of control and buffer implementation. In contrast, the second model is based on the resource partitioning which results in a space partitioning of the CMP chip in several regions. The paper shows a comparison of both models and identifies their main advantages and disadvantages. Francisco Triviño, José L. Sánchez 0002, Francisco J. Alfaro, José Flich |
PDP | 3 |
| 2012 | Optimal Configuration of High-Radix Combined SwitchesabstractHigh-radix switches are an attractive option to improve network performance and to reduce network cost, especially in large switch-based interconnection networks. However, there are some problems related to the integration scale to design such single-chip switches. In this paper we describe an interesting alternative for building high-radix switches which basically consists in combining several current smaller single-chip switches to obtain switches having greater number of ports. This approach is independent of the evolution of single-chip switches and will remain valid as integration scale keeps evolving. We discuss about key design issues of this kind of switches and focus on their internal structure. In order to show the relevance of this issue, we obtain the optimal internal configuration of switches for several networks and evaluate the network performance considering different conditions. Simulation results show that with a correct internal switch design, a network based on these high-radix switches achieves similar performance to a network based on single-chip switches, which have the same number of ports as high-radix switches, and which would be unfeasible with the current integration scale. Juan A. Villar, Francisco J. Andujar, José L. Sánchez 0002, Francisco J. Alfaro, José Duato |
PDP | 4 |
| 2012 | Hardware implementation study of several new egress link scheduling algorithms
Raúl Martínez, José M. Claver, Francisco J. Alfaro, José L. Sánchez 0002 |
J. Parallel Distributed Comput. | 3 |
| 2012 | Network-on-Chip virtualization in Chip-Multiprocessor Systems
Francisco Triviño, José L. Sánchez 0002, Francisco J. Alfaro, José Flich |
J. Syst. Archit. | 3 |
| 2011 | A fast centralized computation routing algorithm for self-configuring NoC systemsabstractAs technology evolves, networks-on-chip will need to survive to manufacturing faults in order to sustain yield. An effective configuration strategy implies the design of an efficient routing infrastructure, that enables a fast and efficient configuration of the NoC system to go around faulty links and switches. The strategy must minimize the overhead in resources and guarantee the entire system to be deadlock free. A centralized approach, through a monitoring controller is appealing as will get global network visibility. This paper proposes a centralized routing configuration strategy that meets the requirements by means of a fast configuration algorithm for the most common failure patterns. The strategy is designed towards the goals of reduced configuration time and high coverage support (maximum number of supported failure patterns). No extra resources (virtual channels) are needed for the effective final configuration of the system. Results show the effectiveness of the proposed configuration algorithm. Francisco Triviño, Francisco J. Alfaro, José L. Sánchez 0002, José Flich |
HiPC | 2 |
| 2011 | C-Switches: Increasing Switch Radix with Current Integration ScaleabstractIn large switch-based interconnection networks, increasing the switch radix results in a decrease in the total number of network components, and consequently the overall cost of the network can be significantly reduced. Moreover, high-radix switches are an attractive option to improve the network performance in terms of latency, since hop count is also reduced. However, there are some problems related to the integration scale to design such single-chip switches. In this paper we discuss key issues and evaluate an interesting alternative for building high-radix switches going beyond the integration scale bounds. The idea basically consists in combining several current smaller single-chip switches to obtain switches having greater number of ports. This approach is independent of the evolution of single-chip switches and remains valid as integration scale keeps evolving. Simulation results show that with a correct internal switch design, this alternative achieves almost the same performance as single-chip switches with the same number of ports, which would be unfeasible with the current integration scale. Juan A. Villar, Francisco J. Andujar, José L. Sánchez 0002, Francisco J. Alfaro, José Duato |
HPCC | 4 |
| 2011 | NoC Reconfiguration for CMP VirtualizationabstractAt NoC level, the traffic interferences can be drastically reduced by using virtualization mechanisms. An effective strategy to virtualize a NoC consists in dividing the network in different partitions, each one serving different applications and traffic flows. In this paper, we propose a NoC reconfiguration mechanism to support NoC virtualization under real scenarios. Dynamic reassignment of network resources to different partitions is allowed in order to NoC dynamically adapts to application needs. Evaluation results show a good behavior of CMP virtualization. Francisco Triviño, Francisco J. Alfaro, José L. Sánchez 0002, José Flich |
NCA | 2 |
| 2011 | Evaluation of an Alternative for Increasing Switch RadixabstractIn large switch-based interconnection networks, increasing the switch radix results in a decrease in the total number of network components. In this paper we evaluate an interesting strategy for building high-radix switches going beyond the integration scale bounds. This approach is independent of the evolution of single-chip switches and will remain valid as integration scale keeps evolving. Simulation results show that with a correct internal switch design, this kind of switches achieves almost the same performance as single-chip switches with the same radix, which would be unfeasible with current integration scale. Juan A. Villar, Francisco J. Andujar, José L. Sánchez 0002, Francisco J. Alfaro, José Duato |
NCA | 4 |
| 2010 | Providing QoS with the Deficit Table SchedulerabstractA key component for networks with Quality of Service (QoS) support is the egress link scheduling algorithm. An ideal scheduling algorithm implemented in a high-performance network with QoS support should satisfy two main properties: good end-to-end delay and implementation simplicity. Table-based schedulers try to offer a simple implementation and good latency bounds. Some of the latest proposals of network technologies, like Advanced Switching and InfiniBand, include in their specifications one of these schedulers. However, these table-based schedulers do not work properly with variable packet sizes, as is usually the case in current network technologies. We have proposed a new table-based scheduler, which we have called Deficit Table (DTable) scheduler, that works properly with variable packet sizes. Moreover, we have proposed a methodology to configure this table-based scheduler in such a way that it permits us to decouple the bounding between the bandwidth and latency assignments. In this paper, we thoroughly review the provision of QoS with the DTable scheduler and our configuration methodology, and evaluate the performance of our proposals in a multimedia scenario. Simulation results show that our proposals are able to provide a similar latency performance than more complex scheduling algorithms. Moreover, we show the advantages of our decoupling configuration methodology over the usual ways of configuring this kind of table-based schedulers. Raul Martinez-Morais, Francisco J. Alfaro, José L. Sánchez 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2009 | Hardware Implementation Study of the SCFQ-CA and DRR-CA Scheduling Algorithms
Raúl Martínez, Francisco J. Alfaro, José L. Sánchez 0002, José M. Claver |
Euro-Par | 2 |
| 2009 | Hardware Implementation Study of the Deficit Table Egress Link Scheduling AlgorithmabstractThe provision of quality of service (QoS) in computing and communication environments has increasingly focused the attention from academia and industry during the last decades. Some of the current interconnection technologies include hardware support that, adequately used, allows to offer QoS guarantees to the applications. The egress link scheduling algorithm is a key part of that support. Apart from providing a good performance in terms of, for example, good end-to-end delay (also called latency) and fair bandwidth allocation, an ideal scheduling algorithm implemented in a high-performance network with QoS support should satisfy other important property which is to have a low computational and implementation complexity. In this paper, we propose a specific implementation of the DTable scheduling algorithm and show estimates about its complexity in terms of silicon area and computation delay. In order to obtain these estimates, we have performed our own hardware implementation using the Handel-C language and employed the DK design suite tool from Celoxica. Raúl Martínez, José M. Claver, Francisco J. Alfaro, José L. Sánchez 0002 |
ICPP | 3 |
| 2009 | A new strategy to manage the InfiniBand arbitration tables
Francisco J. Alfaro, José L. Sánchez 0002, José Duato |
J. Parallel Distributed Comput. | 1 |
| 2009 | A Switch Architecture Guaranteeing QoS Provision and HOL Blocking EliminationabstractBoth QoS support and congestion management techniques become essential to achieve good network performance in current high-speed interconnection networks. The most effective techniques traditionally considered for both issues, however, require too many resources for being implemented. In this paper we propose a new cost-effective switch architecture able to face the challenges of congestion management and, at the same time, to provide QoS. The efficiency of our proposal is based on using the resources (queues) used by RECN (an efficient Head-Of-Line blocking elimination technique) also for QoS support, without increasing queue requirements. Provided results show that the new switch architecture is able to guarantee QoS levels without any degradation due to congestion situations. Alejandro Martínez, Pedro Javier García, Francisco J. Alfaro, José L. Sánchez 0002, José Flich, Francisco J. Quiles 0001, José Duato |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2008 | Efficient Deadline-Based QoS Algorithms for High-Performance NetworksabstractQuality of service (QoS) is becoming an attractive feature for high-performance networks and parallel machines because, in those environments, there are different traffic types, each one having its own requirements. In that sense, deadline-based algorithms can provide powerful QoS provision. However, the cost associated with keeping ordered lists of packets makes these algorithms impractical for high-performance networks. In this paper, we explore how to efficiently adapt the Earliest Deadline First family of algorithms to high-speed network environments. The results show excellent performance using just two virtual channels, FIFO queues, and a cost feasible with today's technology. Alejandro Martínez-Vicente, George Apostolopoulos, Francisco J. Alfaro, José L. Sánchez 0002, José Duato |
IEEE Trans. Computers | 3 |
| 2008 | A Framework to Provide Quality of Service over Advanced SwitchingabstractAdvanced Switching (AS) is a network technology that expands the capabilities of PCI-Express adding new features like peer-to-peer communication. Together, PCI Express and AS have the potential for building the next generation interconnects. Furthermore, the provision of Quality of Service (QoS) in computing and communication environments is currently the focus of much discussion and research in industry and academia. In this paper we propose a framework to provide QoS based on bandwidth, latency, and jitter over AS employing the mechanisms provided by AS. We also present several implementations for the output scheduling mechanism. Finally, we evaluate our proposals by simulation, comparing the performance of the schedulers that we propose and their implementation complexity. Raúl Martínez, Francisco J. Alfaro, José L. Sánchez 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2007 | Integrated QoS Provision and Congestion Management for Interconnection Networks
Alejandro Martínez-Vicente, Pedro Javier García, Francisco J. Alfaro, José L. Sánchez 0002, José Flich, Francisco J. Quiles 0001, José Duato |
Euro-Par | 3 |
| 2007 | Comparing the latency performance of the DTable and DRR schedulersabstractA key component for networks with quality of service (QoS) support is the egress link scheduling algorithm. An ideal scheduling algorithm implemented in a high performance network with QoS support should satisfy two main properties: good end-to-end delay and implementation simplicity. The deficit round robin (DRR) algorithm is known to have a very little implementation complexity. However, depending on the situation, its latency performance can be very bad. On the other hand, table-based schedulers try to offer a simple implementation and good latency bounds. Some of the latest proposals of network technologies, like advanced switching and infiniband, include in their specifications one of these schedulers. However, these table-based schedulers do not work properly with variable packet sizes and face the problem of bounding the bandwidth and latency assignments. We have proposed a new table-based scheduler, which we have called deficit table (DTable) scheduler, that works properly with variable packet sizes. Moreover, we have proposed a methodology to configure this table-based scheduler to decouple the bounding of bandwidth and latency assignments. In this paper, we review these proposals and present simulation results that show that the DTable scheduler is able to provide a better latency performance than the DRR scheduler, with only a slightly higher implementation and computational complexity. Raúl Martínez, Francisco J. Alfaro, José L. Sánchez 0002 |
IPDPS | 2 |
| 2007 | Deadline-based QoS Algorithms for High-performance NetworksabstractQuality of service (QoS) is becoming an attractive feature for high-performance networks and parallel machines because it could allow a more efficient use of resources. Deadline-based algorithms can provide powerful QoS provision. However, the cost associated with keeping ordered lists of packets makes them impractical for high-performance networks. In this paper, we explore how to adapt efficiently the earliest deadline first family of algorithms to the high-speed networks environments. The results show excellent performance using just two virtual channels, FIFO queues, and a cost feasible with today's technology. Alejandro Martínez, Francisco J. Alfaro, José L. Sánchez 0002, José Duato |
IPDPS | 2 |
| 2007 | Efficient Switches with QoS Support for ClustersabstractCurrent interconnect standards providing hardware support for quality of service (QoS) consider up to 16 virtual channels (VCs) for this purpose. However, most implementations do not offer so many VCs because they increase the complexity of the switch and the scheduling delays. We have shown that this number of VCs can be significantly reduced, because it is enough to use two VCs for QoS purposes at each switch port. In this paper, we cover the weaknesses of that proposal and, not only we reduce VCs, but we also improve performance due to the flexibility assigning buffer memory. Alejandro Martínez, Francisco J. Alfaro, José L. Sánchez 0002, José Duato |
IPDPS | 2 |
| 2007 | A low-cost strategy to provide full QoS support in Advanced Switching networks
Alejandro Martínez, Raúl Martínez, Francisco J. Alfaro, José L. Sánchez 0002 |
J. Syst. Archit. | 3 |
| 2007 | A Formal Model to Manage the InfiniBand Arbitration Tables Providing QoSabstractThe InfiniBand architecture (IBA) is an industry-standard architecture for server I/O and interprocessor communication. IBA enables quality-of-service (QoS) support with certain mechanisms. These mechanisms are basically the service levels, the virtual lanes, and the table-based arbitration of those virtual lanes. In previous papers, we have examined these mechanisms and described how we can apply them to the requirements requested by the applications. We have also tested our proposals, showing that the applications achieve the level of QoS requested. In this paper, we present a formal model for the techniques previously proposed. According to this model, each application needs a sequence of entries in the IBA arbitration tables based on its requirements. These requirements are related to the mean bandwidth needed and the maximum latency tolerated by the application. Specifically, each request requires a number of entries with a maximum separation between any consecutive pair. In order to manage the requests, we propose certain algorithms and we prove some propositions and theorems, showing that our method achieves good behavior. Francisco J. Alfaro, José L. Sánchez 0002, M. Menduiña, José Duato |
IEEE Trans. Computers | 1 |
| 2007 | A New Cost-Effective Technique for QoS Support in ClustersabstractVirtual channels (VCs) are a popular solution for the provision of quality of service (QoS). Current interconnect standards propose 16 or even more VCs for this purpose. However, most implementations do not offer so many VCs because it is too expensive in terms of silicon area. Therefore, a reduction of the number of VCs necessary to support QoS can be very helpful in the switch design and implementation. In this paper, we show that this number of VCs can be reduced if the system is considered as a whole rather than each element being taken separately. The scheduling decisions made at network interfaces can be easily reused at switches without significantly altering the global behavior. In this way, we obtain a noticeable reduction of silicon area, component count and, thus, power consumption, and we can provide similar performance to a more complex architecture. We also show that this is a scalable technique, suitable for the foreseen demands of traffic. Alejandro Martínez, Francisco J. Alfaro, José L. Sánchez 0002, Francisco J. Quiles 0001, José Duato |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2006 | Towards a Cost-Effective Interconnection Network Architecture with QoS and Congestion Management Support
Alejandro Martínez, Pedro Javier García, Francisco J. Alfaro, José L. Sánchez 0002, José Flich, Francisco J. Quiles 0001, José Duato |
Euro-Par | 3 |
| 2006 | Improving the Flexibility of the Deficit Table Scheduler
Raúl Martínez, Francisco J. Alfaro, José L. Sánchez 0002 |
HiPC | 2 |
| 2006 | QoS Support for Video Transmission in High-Speed Interconnects
Alejandro Martínez, George Apostolopoulos, Francisco J. Alfaro, José L. Sánchez 0002, José Duato |
HPCC | 3 |
| 2006 | Evaluating several implementations for the AS Minimum Bandwidth Egress Link SchedulerabstractThe relevance of the provision of QoS is taken into account in the definition of the new network technologies like for example Advanced Switching (AS). AS is a new fabric-interconnect technology that further enhances the capabilities of PCI Express, which is the next PCI generation. In this paper we discuss the aspects that must be considered for implementing a specific mechanism for the AS minimum bandwidth egress link scheduler, or just MinBW scheduler. We also propose several implementations for this scheduler, analyze their computational complexity, and compare their performance by simulation. The main differentiating aspect from other interconnection technologies that must be taken into account when implementing the AS MinBW scheduler is that both the link-level flow control and the scheduling are made at a Virtual Channel (VC) level. This means that the scheduler must have the ability to enable or disable the selection of a given VC based on the flow control information. Raul Martinez-Morais, Francisco J. Alfaro, José L. Sánchez 0002 |
ICCCN | 2 |
| 2006 | Decoupling the Bandwidth and Latency Bounding for Table-based SchedulersabstractThe provision of quality of service (QoS) in computing and communication environments is currently the focus of much discussion and research in industry and academia. A key component for networks with QoS support is the output scheduling algorithm. Some of the latest network technology proposals define scheduling algorithms that use an arbitration table to select the next packet to be transmitted. These table-based schedulers are simple to implement and can offer good latency performance. However, the versions proposed until now do not work properly with variable packet sizes. Moreover, they face the problem of bounding the bandwidth and latency assignments. In this paper, we propose a new table-based scheduler, which we call deficit table (DTable), that works properly with variable packet sizes. We also propose a methodology to decouple the bandwidth and latency assignments Raúl Martínez, Francisco J. Alfaro, José L. Sánchez 0002 |
ICPP | 2 |
| 2006 | Studying Several Proposals for the Adaptation of the DTable Scheduler to Advanced Switching
Raúl Martínez, Francisco J. Alfaro, José L. Sánchez 0002 |
ISPA | 2 |
| 2006 | Implementing the Advanced Switching Minimum Bandwidth Egress Link SchedulerabstractAdvanced switching (AS) is a new fabric-interconnect technology that further enhances the capabilities of PCI Express, which is the next PCI generation. On the other hand, the provision of quality of service (QoS) in computing and communication environments is currently the focus of much discussion and research in industry and academia. One of the mechanisms that AS provides to support QoS is the minimum bandwidth egress link scheduler, or just MinBW scheduler. In this paper, we propose several implementations of the MinBW scheduler and compare their performance by simulation. These implementations fulfill all the properties that an AS MinBW scheduler must have, including the interaction with the AS link layer flow control Raúl Martínez, Francisco J. Alfaro, José L. Sánchez 0002 |
NCA | 2 |
| 2006 | Full QoS Support with 2 VCs for Single-chip SwitchesabstractCurrent interconnection standards providing hardware support for quality of service (QoS) consider up to 16 virtual channels (VCs) for this purpose. However, most implementations do not offer so many because VCs increase the complexity of the switch and the scheduling delays. We have shown that this number of VCs can be significantly reduced, because it is enough to use two VCs for QoS purposes at each switch port. In this paper, we explore two alternative switch designs that take advantage of this reduction Alejandro Martínez, Francisco J. Alfaro, José L. Sánchez 0002, José Duato |
NCA | 2 |
| 2005 | Providing Full QoS Support in Clusters Using Only Two VCs at the Switches
Alejandro Martínez, Francisco J. Alfaro, José L. Sánchez 0002, José Duato |
HiPC | 2 |
| 2005 | Studying the Influence of the InfiniBand Packet Size to Guarantee QoSabstractInfiniBand (IBA) has been proposed as an industry-standard architecture both for I/O server and interprocessor communication. IBA employs a switched point-to-point network, instead of using a shared bus. IBA is being developed by the InfiniBand/sub SM/ Trade Association to provide present and future server systems with the required levels of reliability, availability, performance, scalability, and quality of service (QoS). In previous papers we have proposed an effective strategy for configuring the IBA networks to provide users with the required levels of QoS. This strategy is based on the proper configuration of the mechanisms IBA carries to support QoS. Specifically, our methodology configures the InfiniBand arbitration tables and uses the different service levels and virtual lanes that are available, in order to segregate the different traffic flows. Thus, each flow receives the treatment it has previously requested. Moreover, by using our methodology, applications can be assured that their requirements will be satisfied. In this paper, we review the basis of our methodology and we study the influence of the packet size on the QoS guaranteed to the applications. Francisco J. Alfaro, José L. Sánchez 0002, José Duato |
ISCC | 1 |
| 2004 | Tuning Buffer Size in InfiniBand to Guarantee QoS
Francisco J. Alfaro, José L. Sánchez 0002 |
Euro-Par | 1 |
| 2004 | QoS in InfiniBand SubnetworksabstractThe InfiniBand architecture (IBA) has been proposed as an industry standard both for communication between processing nodes and I/O devices and for interprocessor communication. It replaces the traditional bus-based interconnect with a switch-based network for connecting processing nodes and I/O devices. It is being developed by the InfiniBand/sup SM/ Trade Association (IBTA) in the aim to provide the levels of reliability, availability, performance, scalability, and quality of service (QoS) required by present and future server systems. For this purpose, IBA provides a series of mechanisms that are able to guarantee QoS to the applications. In previous papers, we have proposed a strategy to compute the InfiniBand arbitration tables. In one of these, we presented and evaluated our proposal to treat traffic with bandwidth requirements. In another, we evaluated our strategy to compute the InfiniBand arbitration tables for traffic with delay requirements, which is a more complex task. In this paper, we evaluate both these proposals together. Furthermore, we also adapt these proposals in order to treat VBR traffic without QoS guarantees, but achieving very good results. Performance results show that, with a correct treatment of each traffic class in the arbitration of the output port, all traffic classes reach their QoS requirements. Francisco J. Alfaro, José L. Sánchez 0002, José Duato |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2003 | A New Proposal to Fill in the InfiniBand Arbitration TablesabstractThe InfiniBand architecture (IBA) is a new industry-standard architecture for server I/O and interprocessor communication. InfiniBand is very likely to become the de facto standard in a few years. It is being developed by the InfiniBandSMTrade Association (IBTA) to provide the levels of reliability, availability, performance, scalability, and quality of service (QoS) necessary for present and future server systems. We propose a simple and effective strategy for configuring the IBA networks to provide the required levels of QoS. This is a global frame that allows one to do a different treatment to each kind of traffic based on its QoS requirements. It is based on the correct configuration of the mechanisms IBA provides to support QoS. We also propose a simple algorithm to maximize the number of requests to be allocated in the arbitration table that the output ports have. This proposal is evaluated and the results show that every traffic class meets its QoS requirements Francisco J. Alfaro, José L. Sánchez 0002, José Duato |
ICPP | 1 |