EDBT 2026 Demo / reviewers in the wild / expert
Francisco J. Quiles 0001
dblp:30/2177 · also Francisco José Quiles Flor
· DBLP profile ↗
80ranked-venue papers
0as first author
12since 2021 · last 2026
0000-0002-8966-6225ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 58 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13Computer networks · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On the impact of intra- and inter-node communication in the performance of interconnection networks in HPC and AI systemsabstractAbstract In the last decade, specialized computing and storage devices, such as GPUs, TPUs, and high-speed storage, have been incorporated into server nodes of HPC and AI systems. The development of high-bandwidth memory (HBM) enabled a much more compact form factor for these devices, allowing the interconnection of several devices within a server node, typically using an intra-node interconnection network (e.g., PCIe, NVLink, or infinity fabric). Intra-node networks must allow efficient scale-up of the number of these devices within (and even beyond) a single node. Similarly, inter-node networks must enable efficient communication among hundreds of thousands of devices across thousands of server nodes when scale-up domains cannot be larger. Unfortunately, the intra- and inter-node networks may become the system’s bottleneck as communication demand among accelerators increases, driven by emerging applications such as generative AI. Although current intra-node network designs alleviate this bottleneck by increasing intra-node network bandwidth, intra-node communication may hinder inter-node communication performance when traffic from outside the server node arrives at the intra-node network. To evaluate the impact of this interference, we have analyzed intra- and inter-node communication operations generated under realistic traffic scenarios. We have developed a generic intra- and inter-node simulation model in OMNeT++ and have modeled the representative communication operations. We have also performed extensive simulation experiments confirming that increasing the intra-node network bandwidth and the number of computing devices per node (i.e., accelerators) may be counterproductive to the inter-node communication performance. Joaquín Tárraga, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001 |
J. Supercomput. | 4 |
| 2024 | A Hybrid Solution to Provide End-to-End Flow Control and Congestion Management in High-Performance Interconnection NetworksabstractCongestion seriously threatens high-performance interconnection networks in supercomputers and data centers, where thousands of server nodes generate massive communication operations when running highly parallel and distributed applications and services. In recent years, numerous solutions have been proposed to address congestion and its effects, including flow control (e.g., priority flow control, PFC) to prevent packet dropping at congested buffers, injection throttling to detect congested points and notify source server nodes to reduce the injection rate of congesting flows, and congestion isolation (as defined in the IEEE 802.1Qcz standard) that stores the congesting flows in separate queues or virtual channels (VCs) at switch buffers. Unfortunately, these solutions have exhibited important drawbacks, such as the prohibitive latency generated by flow control during congestion situations, the slow and ineffective response of injection throttling, or the excessive resources required to identify and isolate congesting flows. In this paper, we propose a hybrid congestion management solution, called 3SC (from three strategies combined), which combines end-to-end flow control, injection throttling, and congesting-flow isolation. 3SC significantly reduces Head-of-Line (HoL) blocking by isolating congesting flows in special queues at some switches, swiftly throttles the injection of these congesting flows, and significantly decreases the number of flow control messages compared to PFC. To evaluate our proposal, we have conducted a large set of simulation experiments for different network configurations and realistic traffic patterns. The results demonstrate that 3SC is efficient and feasible, making it a promising solution for future interconnection network designs. Alberto Merino, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, Yunping Lyu, José Duato |
CCGrid | 4 |
| 2024 | Hybrid Congestion Control for BXI-Based Interconnection Networks
Gabriel Gomez-Lopez, Miguel Sánchez de la Rosa, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, Pierre-Axel Lagadec |
Euro-Par (2) | 5 |
| 2024 | A New Mechanism to Identify Congesting Packets in High-Performance Interconnection NetworksabstractInterconnection networks are key components in Data Centers and Supercomputers, as they must guarantee high communication bandwidth and low latency under very demanding communication patterns generated by computing-and data-hungry applications and services. These traffic patterns may generate congestion, clogging different parts of the intercon-nection network, and impact the overall system performance if no countermeasures are taken. Unfortunately, congestion detection mechanisms used by congestion control techniques in current interconnection networks, such as DCQCN, do not precisely identify which packets contribute to generating congestion, so false-positive congestion detection events are possible. To overcome these problems, in this paper, we propose a new mechanism, called Enhanced Congestion Point (ECP), which accurately identifies packets that truly contribute to congestion. Specifically, ECP monitors packets at the head of the switch ingress queues and identifies them as congesting when a queue occupancy is over a given threshold and a crossing request for that packet within the switch is rejected. In addition, to solve the false-positive congestion-detection events, ECP defines are-evaluation mechanism that cancels the identification of congesting packets, if they no longer contribute to congestion after congestion areas have been sidestepped. We have evaluated ECP through different experiments using a network simulator that models different interconnection network configurations and realistic traffic patterns. This simulator also provides specific metrics for measuring the quality of the congestion detection mechanism. The obtained results show that ECP precisely identifies contesting packets with a low error margin, improving the DCQCN performance under congestion scenarios. Cristina Olmedilla, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, Yunping Lyu, José Duato |
HOTI | 4 |
| 2024 | A smart and novel approach for managing incast and in-network congestion through adaptive routingabstractHigh-Performance Computing and Datacenter systems, with numerous endnodes, demand an efficient interconnection network to prevent performance bottlenecks. Fat-Tree topologies are preferred for their high bisection bandwidth and multiple shortest-path routes. While existing adaptive routing excels in light or in-network congestion, it struggles with incast congestion. This paper proposes a new technique, called Congestion-Aware Adaptive Routing (SCAR), which addresses both in-network and incast congestion. SCAR limits adaptivity for incast congestion, using deterministic routing, while employing adaptive routing for non-congesting flows. It also resolves in-network congestion by routing traffic flows through alternative routes. Simulation experiments on large Fat-Trees using synthetic and trace-based traffic patterns modeling realistic applications demonstrate SCAR’s immediate reaction on mitigating in-network congestion, and a reasonable delay during incast situations, while other state-of-the-art solutions are not able to cope with incast and in-network situations at the same time. Jose Rocher-Gonzalez, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José Duato |
Future Gener. Comput. Syst. | 4 |
| 2024 | Implementation and testing of a KNS topology in an InfiniBand clusterabstractAbstract The InfiniBand (IB) interconnection technology is widely used in the networks of modern supercomputers and data centers. Among other advantages, the IB-based network devices allow for building multiple network topologies, and the IB control software (subnet manager) supports several routing engines suitable for the most common topologies. However, the implementation of some novel topologies in IB-based networks may be difficult if suitable routing algorithms are not supported, or if the IB switch or NIC architectures are not directly applicable for that topology. This work describes the implementation of the network topology known as KNS in a real HPC cluster using an IB network. As far as we know, this is the first implementation of this topology in an IB-based system. In more detail, we have implemented the KNS routing algorithm in the OpenSM software distribution of the subnet manager, and we have adapted the available IB-based switches to the particular structure of this topology. We have evaluated the correctness of our implementation through experiments in the real cluster, using well-known benchmarks. The obtained results, which match the expected performance for the KNS topology, show that this topology can be implemented in IB-based clusters as an alternative to other interconnection patterns. Gabriel Gomez-Lopez, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001 |
J. Supercomput. | 4 |
| 2023 | Congestion management in high-performance interconnection networks using adaptive routing notifications
Jose Rocher-Gonzalez, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001 |
J. Supercomput. | 4 |
| 2022 | Adaptive Routing in InfiniBand HardwareabstractInterconnection networks are the communication backbone of modern high-performance computing systems and an optimised interconnection network is crucial for the performance and utilisation of the system as a whole. One element of the interconnection network is the routing algorithm, which directly influences how we are able to utilise the physical network topology. InfiniBand is one of the most common network architectures used in high-performance computing and traditionally it only supported static routing. For multi-path networks such as Fat-trees, static routing is inefficient because it cannot balance traffic in real-time nor utilise multiple paths efficiently under adversarial traffic. This again potentially leads to unnecessary contention and an underutilised network, which has led to numerous proposals on how to avoid this by using adaptive routing. Adaptive routing has recently been introduced in InfiniBand and in this paper we evaluate to what extent the expected benefits of adaptive routing is true for InfiniBand. Through a set of experiments on HDR InfiniBand equipment we describe the basic behaviour of adaptive routing in InfiniBand, its benefits in Fat tree topologies and the unfortunate side effects related to unfairness that adaptive routing in general might introduce, including such phenomena as the reverse parking lot problem and congestion spreading. Jose Rocher-Gonzalez, Ernst Gunnar Gran, Sven-Arne Reinemo, Tor Skeie, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001 |
CCGRID | 7 |
| 2022 | RED-SEA: Network Solution for Exascale ArchitecturesabstractIn order to enable Exascale computing, next generation interconnection networks must scale to hundreds of thousands of nodes, and must provide features to also allow the HPC, HPDA, and AI applications to reach Exascale, while benefiting from new hardware and software trends. RED-SEA will pave the way to the next generation of European Exascale interconnects, including the next generation of BXI, as follows: (i) specify the new architecture using hardware-software co-design and a set of applications representative of the new terrain of converging HPC, HPDA, and AI; (ii) test, evaluate, and/or implement the new architectural features at multiple levels, according to the nature of each of them, ranging from mathematical analysis and modeling, to simulation, or to emulation or implementation on FPGA testbeds; (iii) enable seamless communication within and between resource clusters, and therefore development of a high-performance low latency gateway, bridging seamlessly with Ethernet; (iv) add efficient network resource management, thus improving congestion resiliency, virtualization, adaptive routing, collective operations; (v) open the interconnect to new kinds of applications and hardware, with enhancements for end-to-end network services - from programming models to reliability, security, low- latency, and new processors; (vi) leverage open standards and compatible APIs to develop innovative reusable libraries and Fabrics management solutions. Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Pier Stanislao Paolucci, Elena Pastorelli, Francesco Simula, Matteo Turisini, Piero Vicini, Roberto Ammendola, Pascale Bernier-Bruna, Said Derradji, Stéphane Guez, Pierre-Axel Lagadec, Gregoire Pichon, Etienne Walter, Gaetan De Gassowski, Matthieu Hautreaux, Stephane Mathieu, Gilles Moreau, Marc Pérache, Hugo Taboada, Torsten Hoefler, Timo Schneider, Matteo Barnaba, Giuseppe Piero Brandino, Francesco De Giorgi, Matteo Poggi, Iakovos Mavroidis, Ioannis Papaefstathiou, Nikolaos Tampouratzis, Benjamin Kalisch, Ulrich Krackhardt, Mondrian Nüssle, Pantelis Xirouchakis, Vangelis Mageiropoulos, Michalis Gianioudis, Harisis Loukas, Aggelos Ioannou, Nikolaos D. Kallimanis, Nikolaos Chrysos, Manolis Katevenis, Wolfgang Frings, Dominik Gottwald, Felime Guimaraes, Max Holicki, Volker Marx, Yannik Müller, Carsten Clauss, Hugo Falter, Xu Huang 0010, Jennifer Lopez Barillao, Thomas Moschny, Simon Pickartz, Francisco J. Alfaro, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José L. Sánchez 0002, Adrián Castelló 0001, Jose Duro, María Engracia Gómez, Enrique S. Quintana-Ortí, Julio Sahuquillo, Eugenio Stabile |
DSD | 61 |
| 2022 | Improving Congestion Control through Fine-Grain Monitoring of InfiniBand NetworksabstractCongestion situations are a serious threat to the performance of the interconnection networks of High-Performance Computing and Data-Center systems. Hence, the specifications of the main interconnect technologies, such as InfiniBand, define some mechanisms to deal with congestion and its effects. However, these standard mechanisms may not be suitable to detect or track accurately the actual status of network congestion, as congestion dynamics indeed can be very complex and varied. Moreover, achieving an optimal configuration of the parameters that drive the different functionalities of congestion-control mechanisms is often a difficult task, as some configurations may be suitable for some traffic scenarios, but not for others. In this paper, we propose combining an existing light-weight platform monitoring tool (LIMITLESS) with the InfiniBand control software (OpenSM), such that the metrics about communication volumes in the network provided by the former allow the latter having a more precise image of congestion status, then being able to react more efficiently in these situations. The main contributions of this paper are the methodology to link the monitor and OpenSM, as well as modifications in the InfiniBand standard congestion-control mechanism so that its reaction is modulated based on the enhanced knowledge about congestion provided by the monitor. These improvements are ready to be integrated into any InfiniBand-based system. According to the results from our experiments (performed in a real InfiniBand-based cluster where we run a widely used benchmark), the proposed approach reduces significantly the number of wrong detections of congestion, and so the number of times that the congestion-control mechanisms react unnecessarily, hence improving system performance up to 74%. The overhead of this monitoring tool is 0.1% in our experiments, collecting data each 200ms. Alberto Cascajo, Gabriel Gomez-Lopez, Jesús Escudero-Sahuquillo, Pedro Javier García, David E. Singh, Francisco J. Alfaro, Francisco J. Quiles 0001, Jesús Carretero 0001 |
HOTI | 7 |
| 2021 | Leveraging InfiniBand controller to configure deadlock-free routing engines for Dragonflies
German Maglione Mathey, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, Eitan Zahavi |
J. Parallel Distributed Comput. | 4 |
| 2021 | Towards an efficient combination of adaptive routing and queuing schemes in Fat-Tree topologies
Jose Rocher-Gonzalez, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, Gaspar Mora |
J. Parallel Distributed Comput. | 4 |
| 2019 | Efficient Congestion Management for High-Speed Interconnects using Adaptive RoutingabstractThe interconnection network is the central element in high-performance computing (HPC) clusters and Datacenters, where thousands of end nodes must communicate in a fast and reliable manner. The network performance depends on several design choices, such as the topology, the routing algorithm, the switch architecture, etc. Highly efficient routing algorithms, either deterministic or adaptive, have been proposed to smartly balance traffic flows in cost-effective network topologies, but their performance is reduced in scenarios where congestion and their negative effects (e.g. the HoL blocking) appear. In particular, in scenarios where congestion is intense and persistent, the HoL blocking may degrade dramatically the performance of adaptive routing algorithms, since they may spread congested traffic flows through all the available routes. In addition, as we have shown in previous studies, this spreading of congested flows may spoil the performance of the static queuing schemes that are used to reduce HoL blocking by separating flows into different queues at switch buffers. Indeed, as these schemes are based on a static criterion defined prior to the traffic injection in the network, they are unable to avoid that congested and non-congested flows share queues when paired with adaptive routing. In this paper, we propose to use some existing static queuing schemes and dynamic allocation of virtual channels (VCs) to isolate into a single VC the flows whose routes have been adaptively routed, in order to prevent the impact of the congestion spreading through several routes. Basically, adapted flows are moved to a special adapted-flow channel (AFC), so that they do not interact with flows mapped to other VCs by the static queuing scheme. In this way, the HoL blocking that adaptively routed flows could cause to non-adaptive flows is prevented, even if congested flows have been spread through several routes. On the other hand, the static queuing scheme will reduce without any interference the HoL blocking that may appear among non-adaptive flows. To evaluate our proposal we have conducted extensive simulation experiments modeling large interconnection networks based on the fat-tree topology. From the obtained results, we can conclude that our approach efficiently and significantly reduces the HoL blocking impact in interconnection networks using adaptive routing and queuing schemes when congestion appears. Jose Rocher-Gonzalez, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, Gaspar Mora |
CCGRID | 4 |
| 2019 | Head-of-line blocking avoidance in Slim Fly networks using deadlock-free non-minimal and adaptive routingabstractSummary Interconnection network performance is a key issue in HPC systems and datacenters, especially as their number of end nodes grows, to cope with application needs. The network topology and the routing algorithm are important factors for performance and cost. Topologies such as fat‐tree or Dragonfly were proposed to maximize network performance while reducing network resources. One of the most promising topologies is Slim Fly, which offers high network bandwidth assuring low network diameter. However, adversarial traffic and/or congestion situations may degrade Slim Fly's performance dramatically. Non‐minimal routings, such as Valiant or UGAL, can mitigate the former problem while queuing schemes can handle the latter one. In this paper, we proposed a combined mechanism to provide Slim Fly network with both non‐minimal routing and queuing schemes by using several virtual networks to guarantee deadlock freedom. Each virtual network consists of a set of virtual channels to store packets separately according to a mapping policy. This diminishes the interaction among traffic flows, thus reducing head‐of‐line blocking. The results obtained from a simulation‐based evaluation show that our proposal enhances the performance in all the traffic cases, in contrast to other mechanisms whose performance drops in certain scenarios. Pedro Yébenes, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, Torsten Hoefler |
Concurr. Comput. Pract. Exp. | 4 |
| 2019 | Combining Source-adaptive and Oblivious Routing with Congestion Control in High-performance Interconnects using Hybrid and Direct TopologiesabstractHybrid and direct topologies are cost-efficient and scalable options to interconnect thousands of end nodes in high-performance computing (HPC) systems. They offer a rich path diversity, high bisection bandwidth, and a reduced diameter guaranteeing low latency. In these topologies, efficient deterministic routing algorithms can be used to balance smartly the traffic flows among the available routes. Unfortunately, congestion leads these networks to saturation, where the HoL blocking effect degrades their performance dramatically. Among the proposed solutions to deal with HoL blocking, the routing algorithms selecting alternative routes, such as adaptive and oblivious, can mitigate the congestion effects. Other techniques use queues to separate congested flows from non-congested ones, thus reducing the HoL blocking. In this article, we propose a new approach that reduces HoL blocking in hybrid and direct topologies using source-adaptive and oblivious routing. This approach also guarantees deadlock-freedom as it uses virtual networks to break potential cycles generated by the routing policy in the topology. Specifically, we propose two techniques, called Source-Adaptive Solution for Head-of-Line Blocking Avoidance (SASHA) and Oblivious Solution for Head-of-Line Blocking Avoidance (OSHA). Experiment results, carried out through simulations under different traffic scenarios, show that SASHA and OSHA can significantly reduce the HoL blocking. Pedro Yébenes, Jose Rocher-Gonzalez, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Alfaro, Francisco J. Quiles 0001, Crispín Gómez Requena, José Duato |
ACM Trans. Archit. Code Optim. | 6 |
| 2018 | Feasible enhancements to congestion control in InfiniBand-based networks
Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, German Maglione Mathey, José Duato |
J. Parallel Distributed Comput. | 3 |
| 2018 | Scalable Deadlock-Free Deterministic Minimal-Path Routing Engine for InfiniBand-Based Dragonfly NetworksabstractDragonfly topologies are gathering great interest nowadays as one of the most promising interconnect options for High-Performance Computing (HPC) systems. However, Dragonflies contain physical cycles that may lead to traffic deadlocks unless the routing algorithm prevents them properly. In general, existing deadlock-free routing algorithms, either deterministic or adaptive, proposed for Dragonflies, use Virtual Channels (VCs) to prevent cyclic dependencies. However, these topology-aware algorithms are difficult to implement, or even unfeasible, in systems based on the InfiniBand (IB) architecture, which is nowadays the most widely used network technology in HPC systems. This is due to some limitations in the IB specification, specifically regarding the way Virtual Lanes (VLs), which are considered as similar to VCs, can be assigned to traffic flows. Indeed, none of the routing engines currently available in the official releases of the IB control software has been specifically proposed for Dragonflies. In this paper, we present a new deterministic, minimal-path routing for Dragonfly that prevents deadlocks using VLs according to the IB specification, so that it can be straightforwardly implemented in IB-based networks. We have called this proposal D3R (Deterministic Deadlock-free Dragonfly Routing). Specifically, D3R maps each route to a single, specific VL depending on the destination group, and according to a specific order, so that cyclic dependencies (so deadlocks) are prevented. D3R is scalable as it requires only 2 VLs to prevent deadlocks regardless of network size, i.e., fewer VLs than the required by the deadlock-free routing engines available in IB that are suitable for Dragonflies. Alternatively, D3R achieves higher throughput if an additional VL is used to reduce internal contention in the Dragonfly groups. We have implemented D3R as a new routing engine in OpenSM, the control software including the subnet manager in IB. We have evaluated D3R by means of simulation and by experiments performed in a real IB-based cluster, the results showing that, in general, D3R outperforms other routing engines. German Maglione Mathey, Pedro Yébenes, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, Eitan Zahavi |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2017 | Providing differentiated services, congestion management, and deadlock freedom in dragonfly networks with adaptive routingabstractSummary The number of endnodes in high‐performance computing systems has grown significantly in the last years. Hence, the interconnection network has become an essential issue as it may end up being the system bottleneck if it is not properly designed. In that sense, the Dragonfly topology has become very popular for interconnecting high‐performance computing systems in the last years because it offers high performance at an affordable cost. However, when using deterministic minimal‐path routing, this topology is not able to offer a high performance under certain traffic conditions. This problem can be solved by using oblivious or adaptive routing. However, there are no congestion management techniques specially tailored to Dragonfly topologies using oblivious or adaptive routing. Note that in congestion situations, the Dragonfly performance may drop because of the head‐of‐line blocking effect. This effect could be even more dangerous in systems where several applications with different priorities coexist. In this work we propose several techniques especially designed for providing differentiated services and congestion management in Dragonfly networks using oblivious or adaptive routing. First, we propose thehierarchical 3‐level queuingqueuing scheme, which configures several virtual channels distributed into 3 virtual networks to reduce the head‐of‐line blocking while deadlocks derived from the routing algorithm are prevented. Second, we extendhierarchical 3‐level queuingto provide differentiated services through 2 different solutions. Finally, some experiments are performed to show the benefits obtained by using the proposed techniques. Pedro Yébenes, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Alfaro, Francisco J. Quiles 0001 |
Concurr. Comput. Pract. Exp. | 5 |
| 2017 | Providing differentiated services, congestion management, and deadlock freedom in dragonfly networks with adaptive routingabstractIn this article,1 an error in one of the author names has been found subsequent to the publication. “Jesus Escudero-Sahuquilllo” should be “Jesus Escudero-Sahuquillo.” The correct name is now presented above. The author's name has also been corrected in the original published article. Pedro Yébenes, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Alfaro, Francisco J. Quiles 0001 |
Concurr. Comput. Pract. Exp. | 5 |
| 2016 | Straightforward solutions to reduce HoL blocking in different Dragonfly fully-connected interconnection patterns
Pedro Yébenes, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001 |
J. Supercomput. | 4 |
| 2015 | Efficient Queuing Schemes for HoL-Blocking Reduction in Dragonfly Topologies with Minimal-Path RoutingabstractHPC systems are growing in number of connected endnodes, making the network a main issue in their design. In order to interconnect large systems, dragonfly topologies have become very popular in the latest years as they achieve high scalability by exploiting high-radix switches. However, dragonfly high performance may drop severely due to the Head-of-Line (HoL) blocking effect derived from congestion situations. Many techniques have been proposed for dealing with this harmful effect, the most effective ones being those especially designed for a specific topology and a specific routing algorithm. In this paper we present a queuing scheme called Hierarchical Two-Levels Queuing, designed specially to reduce HoL blocking in fully-connected dragonfly networks that use minimal-path routing. This proposal boosts network performance compared with other techniques and requires fewer network resources than the others. Besides, an upgrade for existing queuing schemes for improving their performance is explained. Pedro Yébenes, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001 |
CLUSTER | 4 |
| 2015 | Efficient and Cost-Effective Hybrid Congestion Control for HPC Interconnection NetworksabstractInterconnection networks are key components in high-performance computing (HPC) systems, their performance having a strong influence on the overall system one. However, at high load, congestion and its negative effects (e.g., Head-of-line blocking) threaten the performance of the network, and so the one of the entire system. Congestion control (CC) is crucial to ensure an efficient utilization of the interconnection network during congestion situations. As one major trend is to reduce the effective wiring in interconnection networks to reduce cost and power consumption, the network will operate very close to its capacity. Thus, congestion control becomes essential. Existing CC techniques can be divided into two general approaches. One is to throttle traffic injection at the sources that contribute to congestion, and the other is to isolate the congested traffic in specially designated resources. However, both approaches have different, but non-overlapping weaknesses: injection throttling techniques have a slow reaction against congestion, while isolating traffic in special resources may lead the system to run out of those resources. In this paper we propose EcoCC, a new Efficient and Cost-Effective CC technique, that combines injection throttling and congested-flow isolation to minimize their respective drawbacks and maximize overall system performance. This new strategy is suitable for current commercial switch architectures, where it could be implemented without requiring significant complexity. Experimental results, using simulations under synthetic and real trace-based traffic patterns, show that this technique improves by up to 55 percent over some of the most successful congestion control techniques. Jesús Escudero-Sahuquillo, Ernst Gunnar Gran, Pedro Javier García, José Flich, Tor Skeie, Olav Lysne, Francisco J. Quiles 0001, José Duato |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2014 | Combining HoL-blocking avoidance and differentiated services in high-speed interconnectsabstractCurrent high-performance platforms such as Datacenters or High-Performance Computing systems rely on highspeed interconnection networks able to cope with the ever-increasing communication requirements of modern applications. In particular, in high-performance systems that must offer differentiated services to applications which involve traffic prioritization, it is almost mandatory that the interconnection network provides some type of Quality-of-Service (QoS) and Congestion-Management mechanism in order to achieve the required network performance. Most current QoS and Congestion-Management mechanisms for high-speed interconnects are based on using the same kind of resources, but with different criteria, resulting in disjoint types of mechanisms. By contrast, we propose in this paper a novel, straightforward solution that leverages the resources already available in InfiniBand components (basically Service Levels and Virtual Lanes) to provide both QoS and Congestion Management at the same time. This proposal is called CHADS (Combined HoL-blocking Avoidance and Differentiated Services), and it could be applied to any network topology. From the results shown in this paper for networks configured with the novel, cost-efficient KNS hybrid topology, we can conclude that CHADS is more efficient than other schemes in reducing the interferences among packet flows that have the same or different priorities. Pedro Yébenes, Jesús Escudero-Sahuquillo, Crispín Gómez Requena, Pedro Javier García, Francisco J. Alfaro, Francisco J. Quiles 0001, José Duato |
HiPC | 6 |
| 2014 | A new proposal to deal with congestion in InfiniBand-based fat-trees
Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, Sven-Arne Reinemo, Tor Skeie, Olav Lysne, José Duato |
J. Parallel Distributed Comput. | 3 |
| 2014 | Toward fast Wyner-Ziv video decoding on multicore processors
Alberto Corrales-García, José Luis Martínez 0001, Gerardo Fernández-Escribano, Francisco J. Quiles 0001 |
Multim. Tools Appl. | 4 |
| 2013 | BBQ: A Straightforward Queuing Scheme to Reduce HoL-Blocking in High-Performance Hybrid Networks
Pedro Yébenes, Jesús Escudero-Sahuquillo, Crispín Gómez Requena, Pedro Javier García, Francisco J. Quiles 0001, José Duato |
Euro-Par | 5 |
| 2013 | Towards Modeling Interconnection Networks of Exascale Systems with OMNet++abstractOne of the objectives of the decade for High-Performance Computing systems is to reach the exascale level of computing power before 2018, hence this will require strong efforts in their design. In that sense, High-speed low-latency interconnection networks are essential elements for exascale HPC systems. Indeed, the performance of the whole system depends on that of the interconnection network. In order to develop and test new techniques, suited to exascale HPC systems, software-based networks simulators are commonly used. As developing a network simulator from scratch is a difficult task, several platforms help the developers, OMNeT++ being one of the most popular. In this paper, we propose a new generic network simulator, exploiting the features of the OMNeT++ framework. The proposed tool is the first step to model HPC high-performance interconnection networks of exascale HPC systems: the message switching layer, routing and arbitration algorithms and buffer organizations have been modeled according to the current and expected characteristics of these systems. In addition, the tool has been designed so that it is possible to simulate networks of large size. Simulation results, validated against real systems, show the accuracy of the model. Pedro Yébenes, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001 |
PDP | 4 |
| 2013 | An integrated solution for QoS provision and congestion management in high-performance interconnection networks using deterministic source-based routing
Juan A. Villar, Pedro Javier García, Francisco J. Alfaro, José L. Sánchez 0002, Francisco J. Quiles 0001 |
J. Supercomput. | 5 |
| 2013 | An Effective and Feasible Congestion Management Technique for High-Performance MINs with Tag-Based Distributed RoutingabstractAs parallel computing systems increase in size, the interconnection network is becoming a critical subsystem. The current trend in network design is to use as few components as possible to interconnect the end nodes, thereby reducing cost and power consumption. However, this increases the probability of congestion appearing in the network. As congestion may severely degrade network performance, the use of a congestion management mechanism is becoming mandatory in modern interconnects. One of the most cost-effective proposals to deal with the problems derived from congestion situations is the Regional Explicit Congestion Notification (RECN) strategy, based on using special queues to totally isolate the packet flows which contribute to congestion, thereby preventing the Head-of-Line (HoL) blocking effect that these flows may cause to others. Unfortunately, RECN requires the use of source-based routing, thus not being suitable for interconnects with distributed routing, like InfiniBand. Although some RECN-like mechanisms have been proposed for distributed-routing networks, they are not scalable due to the huge amount of control memory that they require in medium-size or large networks. In this paper, we propose Distributed-Routing-Based Congestion Management (DRBCM), a new scalable technique which, following the RECN principles, totally prevents congestion from producing HoL-blocking in multistage interconnection networks (MINs) using tag-based distributed routing. Simulation results indicate that, regardless of network size, DRBCM presents small resource requirements to keep network performance at maximum level even in scenarios of heavy congestion, where it utterly outperforms (with a gain up to 70 percent) current solutions for distributed-routing networks, like the InfiniBand congestion-control mechanism based on injection throttling. Thus, DRBCM is an efficient, cost-effective, and scalable solution for congestion management. Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José Flich, José Duato |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2012 | Multi-Core Parallel Algorithm for Wyner-Ziv Video DecodingabstractWyner-Ziv video coding paradigm provides a framework where most of complexity is moved from the encoder to the decoder. In this way, Wyner-Ziv coding support efficiently multimedia services for low complexity devices which have to capture, encode and send video. Aplications such as mobile phones, sensor, Personal Digital Assistant (PDA) could benefice to this paradigm. On contraty to tradicional video paradigms, the complexity of the decoder is quite high and it should be reduced. This work presents a parallel Wyner-Ziv decoding algorithm in aims to reduce its high complexity. The present approach efficiently distribute the burden of the complexity over the number of cores which are available in the architecture. By using this parallel approach, the decoding time is reduced around 73%. The proposed methods are scalable for any multicore architecture and adaptable for different Wyner-Ziv decoding schemes. Alberto Corrales-García, José Luis Martínez 0001, Gerardo Fernández-Escribano, Francisco J. Quiles 0001 |
ISPA | 4 |
| 2012 | Scalable Mobile-to-Mobile Video Communications Based on an Improved WZ-to-SVC Transcoder
Alberto Corrales-García, José Luis Martínez 0001, Gerardo Fernández-Escribano, Francisco J. Quiles 0001 |
MMM | 4 |
| 2012 | Forward Wyner-Ziv Fast Video Decoding Using Multicore Processors
Alberto Corrales-García, José Luis Martínez 0001, Gerardo Fernández-Escribano, Francisco J. Quiles 0001 |
MMM | 4 |
| 2011 | Combining Congested-Flow Isolation and Injection Throttling in HPC Interconnection NetworksabstractExisting congestion control mechanisms in interconnects can be divided into two general approaches. One is to throttle traffic injection at the sources that contribute to congestion, and the other is to isolate the congested traffic in specially designated resources. These two approaches have different, but non-overlapping weaknesses. In this paper we present in detail a method that combines injection throttling and congested-flow isolation. Through simulation studies we first demonstrate the respective flaws of the injection throttling and of flow isolation. Thereafter we show that our combined method extracts the best of both approaches in the sense that it gives fast reaction to congestion, it is scalable and it has good fairness properties with respect to the congested flows. Jesús Escudero-Sahuquillo, Ernst Gunnar Gran, Pedro Javier García, José Flich, Tor Skeie, Olav Lysne, Francisco J. Quiles 0001, José Duato |
ICPP | 7 |
| 2011 | Wyner-Ziv frame parallel decoding based on multicore processorsabstractWyner-Ziv video coding presents a new paradigm which offers low-complexity video encoding. However, the Wyner-Ziv paradigm accumulates high complexity at the decoder side and this could involve difficulties for applications which have delay requisites. On the other hand, technological advances provide us with new hardware which supports parallel data processing. In this paper, a faster Wyner-Ziv video decoding scheme based on multicore processors is proposed. In this way, each frame is decoded by means of the collaboration between several processing units, achieving a time reduction up to 71% without significant rate-distortion drop penalty. Alberto Corrales-García, José Luis Martínez 0001, Gerardo Fernández-Escribano, Francisco J. Quiles 0001, Warnakulasuriya Anil Chandana Fernando |
MMSP | 4 |
| 2011 | Cost-effective queue schemes for reducing head-of-line blocking in fat-treesabstractSUMMARY The fat‐tree is one of the most common topologies among the interconnection networks of the systems currently used for high‐performance parallel computing. Among other advantages, fat‐trees allow the use of simple but very efficient routing schemes. One of them is a deterministic routing algorithm that has been recently proposed, offering a similar (or better) performance than adaptive routing while reducing complexity and guaranteeing in‐order packet delivery. However, as other deterministic routing proposals, this deterministic routing algorithm cannot react when high traffic loads or hot‐spot traffic scenarios produce severe contention for the use of network resources, leading to the appearance of Head‐of‐Line (HoL) blocking, which spoils the network performance. In that sense, we describe in this paper two simple, cost‐effective strategies for dealing with the HoL‐blocking problem that may appear in fat‐trees with the aforementioned deterministic routing algorithm. From the results presented in the paper, we conclude that, in the mentioned environment, these proposals considerably reduce HoL‐blocking without significantly increasing switch complexity and the required silicon area. Copyright © 2011 John Wiley & Sons, Ltd. Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José Flich, José Duato |
Concurr. Comput. Pract. Exp. | 3 |
| 2011 | OBQA: Smart and cost-efficient queue scheme for Head-of-Line blocking elimination in fat-trees
Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José Flich, José Duato |
J. Parallel Distributed Comput. | 3 |
| 2011 | Variable and constant bitrate in a DVC to H.264/AVC transcoder
Alberto Corrales-García, José Luis Martínez 0001, Gerardo Fernández-Escribano, Francisco J. Quiles 0001 |
Signal Process. Image Commun. | 4 |
| 2010 | Mapping GOPS in an Improved DVC to H.264 Video Transcoder
Alberto Corrales-García, Gerardo Fernández-Escribano, Francisco J. Quiles 0001 |
ACIVS (2) | 3 |
| 2010 | An Efficient Strategy for Reducing Head-of-Line Blocking in Fat-Trees
Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José Duato |
Euro-Par (2) | 3 |
| 2010 | Cost-Effective Congestion Management for Interconnection Networks Using Distributed Deterministic RoutingabstractThe Interconnection networks are essential elements in current computing systems. For this reason, achieving the best network performance, even in congestion situations, has been a primary goal in recent years. In that sense, there exist several techniques focused on eliminating the main negative effect of congestion: the Head of Line (HOL) blocking. One of the most successful HOL blocking elimination techniques is RECN, which can be applied in source routing networks. FBICM follows the same approach as RECN, but it has been developed for distributed deterministic routing networks. Although FBICM effectively eliminates HOL blocking, it requires too much resources to be implemented. In this paper we present a new FBICM version, based on a new organization of switch memory resources, that significantly reduces the required silicon area, complexity and cost. Moreover, we present new results about FBICM, in network topologies not yet analyzed. From the experiment results we can conclude that a far less complex and feasible FBICM implementation can be achieved by using the proposed improvements, while not losing efficiency. Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José Flich, José Duato |
ICPADS | 3 |
| 2010 | Flexible GOP transcoding between DVC and H.264abstractMobile to mobile video telephony is being one of the most attractive services that the newest generations of mobile communications system (such as 4G) are offering at the present time. This kind of service needs special requirements in terms of low complexity in both sides of the communication. By using traditional video encoders, such as H.264, those requirements are not satisfied due to the complexity of the encoder. Distributed Video Coding (DVC) deals with the problem of higher complexity constraints encoding algorithms at the expense of increasing the decoder complexity. In order to efficiently support video communications, this paper proposes an improved DVC to H.264 transcoder that maps different kind of GOPs, as well as different GOP lengths, between both paradigms. Moreover, the H.264 motion estimation is adjusted by using information gathered in the first step of the transcoders, offering a considerable reduction of the transcoding total time with a negligible rate-distortion penalty. Alberto Corrales-García, José Luis Martínez 0001, Francisco J. Quiles 0001 |
MoMM | 3 |
| 2009 | Effiecient WZ-to-H264 transcoding using motion vector information sharingabstractIn mobile-to-mobile video communications, both the sender and the receiver devices should not have higher complexity requirements to perform complex video compression tasks. The traditional video coding solutions are not suitable to support this communications due to its extremely complex encoding algorithm. On the other hand, the new Wyner-Ziv video coding paradigm reduces the complexity of the encoder at the expenses of a more complex decoder. In this paper, we propose an improved WZ/H.264 video transcoder to support this mobile-to-mobile communications, using the low complexity Wyner-Ziv encoding and the traditional H.264 decoding to be implemented in the end-user devices. The improved transcoder converts the video from the Wyner-Ziv to H.264 and reuses the motion vectors generated in the Wyner-Ziv decoding, in order to reduce the computational complexity of the motion estimation process in the H.264 encoding. Simulations results show a complexity reduction up to 55% with negligible rate-distortion drop. José Luis Martínez 0001, Hari Kalva, Warnakulasuriya Anil Chandana Fernando, Pedro Cuenca 0001, Francisco J. Quiles 0001 |
ICME | 5 |
| 2009 | A Switch Architecture Guaranteeing QoS Provision and HOL Blocking EliminationabstractBoth QoS support and congestion management techniques become essential to achieve good network performance in current high-speed interconnection networks. The most effective techniques traditionally considered for both issues, however, require too many resources for being implemented. In this paper we propose a new cost-effective switch architecture able to face the challenges of congestion management and, at the same time, to provide QoS. The efficiency of our proposal is based on using the resources (queues) used by RECN (an efficient Head-Of-Line blocking elimination technique) also for QoS support, without increasing queue requirements. Provided results show that the new switch architecture is able to guarantee QoS levels without any degradation due to congestion situations. Alejandro Martínez, Pedro Javier García, Francisco J. Alfaro, José L. Sánchez 0002, José Flich, Francisco J. Quiles 0001, José Duato |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2008 | FBICM: Efficient Congestion Management for High-Performance Networks Using Distributed Deterministic Routing
Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José Flich, José Duato |
HiPC | 3 |
| 2008 | DVC using a half-feedback based approachabstractDistributed video coding has become increasingly popular in recent years among the researchers in video coding due to its attractive and promising features. DVC proposed a dramatic structural change to video coding by shifting the majority of complexity conventionally residing in the encoder towards the decoder. Nevertheless, these kinds of architectures have some serious limitations that hinder its practical application. The uses of a feedback channel between the encoder and the decoder requires an interactive decoding procedure which is a limitation for certain applications such as offline processing. On the other hand, the decoder needs an efficient way to estimate the probability of error without assuming the availability of the original video at the decoder. In this paper we investigate a first approximation to solve both problems based on the use of machine learning to extract the knowledge that exits between the residual frame and the number of requests over this feedback channel. Exploiting this correlation gives us a more practical architecture without higher complexity encoders. We apply these concepts to pixel-domain Wyner-Ziv coding and the results show a loss of 0.21 dB in the rate-distortion performance. José Luis Martínez 0001, Christopher Holder, Gerardo Fernández-Escribano, Hari Kalva, Francisco J. Quiles 0001 |
ICME | 5 |
| 2008 | Transform Domain Wyner-Ziv Codec Based on Turbo Trellis Codes Modulation
José Luis Martínez 0001, W. A. Rajitha Jayaruwan Weerakkody, Pedro Cuenca 0001, Francisco J. Quiles 0001, Warnakulasuriya Anil Chandana Fernando |
MMM | 4 |
| 2008 | Evaluation of a Fabric Management Mechanism for Advanced Switching in Presence of TrafficabstractRecent years, computer performance has been significantly increased. As a consequence, data I/O systems have become bottlenecks within systems. In order to alleviate this problem, Advanced Switching Interconnect was proposed as a new standard for future high-performance interconnects. The Advanced Switching specification establishes a fabric management infrastructure, which is in charge of maintaining the connectivity between fabric endpoints after fabric initialization and each time a topological change takes place. Once the change has been detected, a fabric manager must discover the new topology, obtain a new set of valid fabric paths, and distribute them to the fabric endpoints. This paper evaluates a first implementation for this management mechanism, by analyzing the influence of application traffic on its performance, and the way in which network service can be affected by the change assimilation process. Antonio Robles-Gómez, Aurelio Bermúdez, Rafael Casado, Francisco J. Quiles 0001 |
PDP | 4 |
| 2008 | A proposal for managing ASI fabrics
Antonio Robles-Gómez, Aurelio Bermúdez, Rafael Casado, Francisco J. Quiles 0001, Tor Skeie, José Duato |
J. Syst. Archit. | 4 |
| 2007 | Integrated QoS Provision and Congestion Management for Interconnection Networks
Alejandro Martínez-Vicente, Pedro Javier García, Francisco J. Alfaro, José L. Sánchez 0002, José Flich, Francisco J. Quiles 0001, José Duato |
Euro-Par | 6 |
| 2007 | An Iterative Refinement Technique for Side Information Generation in DVCabstractDistributed video coding (DVC) is an increasingly popular approach among the researchers in video coding during past few years due to its attractive and promising features. In DVC, the majority of the computational complexity has been shifted from encoder to the decoder in comparison to its conventional counterparts, including MPEG and H.26 x enabling a dramatically low cost encoder implementation. Side information generation, carried out at the decoder, is a major function in the DVC coding algorithm and plays a key-role in determining the performance of the codec. In this paper, a novel iterative refinement technique is proposed for the side information generation process. Simulation results of the proposed technique depict a consistent improvement in performance in comparison to the state-of-the-art in pixel domain DVC. W. A. Rajitha Jayaruwan Weerakkody, Warnakulasuriya Anil Chandana Fernando, José Luis Martínez 0001, Pedro Cuenca 0001, Francisco J. Quiles 0001 |
ICME | 5 |
| 2007 | Implementing the Advanced Switching Fabric Discovery ProcessabstractAdvanced switching is a new high-speed industrial standard serial interconnect. It is defined as a switching fabric architecture based on the PCI express technology. The advanced switching specification establishes a management infrastructure which maintains the fabric operation. The topology discovery process is triggered after fabric initialization and every time a topological change is detected. The information gathered by this process is used to build a set of paths between fabric endpoints. This work analyzes the performance of several possible implementations for this management task. Antonio Robles-Gómez, Aurelio Bermúdez, Rafael Casado, Francisco J. Quiles 0001 |
IPDPS | 4 |
| 2007 | A Complete Topology Management Mechanism for the Advanced Switching Interconnect TechnologyabstractThe advanced switching technology is a new high-performance standard serial inter-connect. Its specification establishes a management infrastructure in charge of maintaining the fabric operation after the occurrence of a topological change. When the change is detected, the management mechanism discovers the new topology, obtains a set of fabric paths, and finally distributes them to the endpoints. The main contribution of this paper is a completely functional mechanism that implements all these management tasks involved in handling topological changes in source routing switched networks, as advanced switching. We also analyze the behavior of this first proposal, in order to identify its bottlenecks and define future improvements. Antonio Robles-Gómez, Aurelio Bermúdez, Rafael Casado, Francisco J. Quiles 0001 |
ISCC | 4 |
| 2007 | Low-Complexity TTCM Based Distributed Video Coding Architecture
José Luis Martínez 0001, Warnakulasuriya Anil Chandana Fernando, W. A. Rajitha Jayaruwan Weerakkody, José Oliver 0001, Otoniel López, Miguel Martínez-Rach, M. Pérez, Pedro Cuenca 0001, Francisco J. Quiles 0001 |
PSIVT | 9 |
| 2007 | Handling Topology Changes in InfiniBandabstractInfiniBand is a high-performance switched network. Its topology may change due to devices being turned on/off, hot expansion, link remapping, and component failures. The InfiniBand specification defines a management infrastructure which is responsible for detecting and assimilating any change in the network. When a change occurs, management entities must update switch forwarding tables, in order to maintain the connectivity among end nodes. This implies the acquisition of the current topology and the computation of a new set of routes accordingly. It is desirable that the execution of this process does not affect the performance of the upper-level applications that are using the network. In previous works, we have proposed enhanced implementations for the main tasks involved in the assimilation of a change. Now, we present a detailed performance evaluation of a management mechanism which incorporates all our proposals Aurelio Bermúdez, Rafael Casado, Francisco J. Quiles 0001, José Duato |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2007 | A New Cost-Effective Technique for QoS Support in ClustersabstractVirtual channels (VCs) are a popular solution for the provision of quality of service (QoS). Current interconnect standards propose 16 or even more VCs for this purpose. However, most implementations do not offer so many VCs because it is too expensive in terms of silicon area. Therefore, a reduction of the number of VCs necessary to support QoS can be very helpful in the switch design and implementation. In this paper, we show that this number of VCs can be reduced if the system is considered as a whole rather than each element being taken separately. The scheduling decisions made at network interfaces can be easily reused at switches without significantly altering the global behavior. In this way, we obtain a noticeable reduction of silicon area, component count and, thus, power consumption, and we can provide similar performance to a more complex architecture. We also show that this is a scalable technique, suitable for the foreseen demands of traffic. Alejandro Martínez, Francisco J. Alfaro, José L. Sánchez 0002, Francisco J. Quiles 0001, José Duato |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2006 | Towards a Cost-Effective Interconnection Network Architecture with QoS and Congestion Management Support
Alejandro Martínez, Pedro Javier García, Francisco J. Alfaro, José L. Sánchez 0002, José Flich, Francisco J. Quiles 0001, José Duato |
Euro-Par | 6 |
| 2006 | A Model for the Development of AS Fabric Management Protocols
Antonio Robles-Gómez, Eva M. García, Aurelio Bermúdez, Rafael Casado, Francisco J. Quiles 0001 |
Euro-Par | 5 |
| 2006 | RECN-DD: A Memory-Efficient Congestion Management Technique for Advanced SwitchingabstractAs VLSI technology advances, the interconnection network represents a larger percentage of the total system cost and power consumption. In fact, a current trend in network design is to reduce the number of components. However, this leads to systems working closer to saturation point, and therefore an efficient congestion management technique is required. In that sense, RECN has been recently proposed for advanced switching (AS). RECN detects the formation of congestion trees and dynamically allocates queues for storing congested packets, thus, eliminating the HOL blocking introduced by congestion trees. These queues are deallocated when congestion vanishes. We have identified two shortcomings that may affect RECN scalability and implementation. Firstly, although RECN allocates queues in an efficient way, resource deallocation is performed in-order, thus losing efficiency and wasting resources. This leads to an excessive requirement of memory at switch ports. Secondly, both allocation and deallocation mechanisms involve the use of specific control packets not supported by the AS standard, thus preventing RECN implementation. In this sense we provide a detailed description of the current RECN deallocation mechanism. In this paper we present an enhanced RECN version (RECN-DD) where these problems have been eliminated. Specifically, we propose a new distributed queue deallocation mechanism that reduces the number of required resources and does not require the use of control packets. Moreover, we propose a new congestion notification mechanism that does not require non-standard AS packets. Instead, flow control packets are used to notify congestion, thus simplifying the implementation of RECN-DD in AS Pedro Javier García, Francisco J. Quiles 0001, José Flich, José Duato, Ian Johnson, Finbar Naven |
ICPP | 2 |
| 2006 | MMR: A MultiMedia Router architecture to support hybrid workloads
María Blanca Caminero, Carmen Carrión 0001, Francisco J. Quiles 0001, José Duato, Sudhakar Yalamanchili |
J. Parallel Distributed Comput. | 3 |
| 2005 | A New Hardware Efficient Link Scheduling Algorithm to Guarantee QoS on Clusters
José M. Claver, Carmen Carrión 0001, Manel Canseco, María Blanca Caminero, Francisco J. Quiles 0001 |
Euro-Par | 5 |
| 2005 | On the Correct Sizing on Meshes Through an Effective Congestion Management Strategy
Pedro Javier García, José Flich, José Duato, Francisco J. Quiles 0001, Ian Johnson, Finbar Naven |
Euro-Par | 4 |
| 2005 | Dynamic Evolution of Congestion Trees: Analysis and Impact on Switch Architecture
Pedro Javier García, José Flich, José Duato, Ian Johnson, Francisco J. Quiles 0001, Finbar Naven |
HiPEAC | 5 |
| 2005 | Traffic Scheduling Solutions with QoS Support for an Input-Buffered MultiMedia RouterabstractQuality of service (QoS) support in local and cluster area environments has become an issue of great interest in recent years. Most current high-performance interconnection solutions for these environments have been designed to enhance conventional best-effort traffic performance, but are not well-suited to the special requirements of the new multimedia applications. The multimedia router (MMR) aims at offering hardware-based QoS support within a compact interconnection component. One of the key elements in the MMR architecture is the algorithms used in traffic scheduling. These algorithms are responsible for the order in which information is forwarded through the internal switch. Thus, they are closely related to the QoS-provisioning mechanisms. In this paper, several traffic scheduling algorithms developed for the MMR architecture are described. Their general organization is motivated by chances for parallelization and pipelining, while providing the necessary support both to multimedia flows and to best-effort traffic. Performance evaluation results show that the QoS requirements of different connections are met, in spite of the presence of best-effort traffic, while achieving high link utilizations. María Blanca Caminero, Carmen Carrión 0001, Francisco J. Quiles 0001, José Duato, Sudhakar Yalamanchili |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2004 | Distributing InfiniBand Forwarding Tables
Aurelio Bermúdez, Rafael Casado, Francisco J. Quiles 0001 |
Euro-Par | 3 |
| 2004 | Use of Provisional Routes to Speed-up Change Assimilation in InfiniBand NetworksabstractSummary form only given. The InfiniBand architecture has been proposed as a technology both for communication between processing nodes and I/O devices, and for interprocessor communication. The InfiniBand specification defines a basic management infrastructure that is responsible for subnet configuration, activation, and fault tolerance. Each time a topology change is detected, management entities collect the current subnet topology. After that, new forwarding tables have to be computed and uploaded to routing devices. The time required to compute these tables is a critical issue, due to application traffic being negatively affected by the temporary lack of connectivity. We present a way to compute a valid set of subnet routes in a short period of time. These provisional routes can be immediately distributed to routing devices. After that, final routes can be later uploaded without affecting user traffic. Aurelio Bermúdez, Rafael Casado, Francisco J. Quiles 0001, José Duato |
IPDPS | 3 |
| 2003 | On the InfiniBand Subnet Discovery ProcessabstractInfiniBand is becoming an industry standard both for communication between processing nodes and I/O devices, and for interprocessor communication. Instead of using a shared bus, InfiniBand employs an arbitrary (possibly irregular) switched point-to-point network. InfiniBand specification defines a basic management infrastructure that is responsible for subnet configuration, activation, and fault tolerance. After the detection of a topology change, management entities collect the current subnet topology. The topology discovery algorithm is one of the management issues that are outside the scope of the current specification. Preliminary implementations obtain the entire topological information each time a change is detected. In this work, we present and analyze an optimized implementation, based on exploring only the region that has been affected by the change. Aurelio Bermúdez, Rafael Casado, Francisco J. Quiles 0001, Timothy M. Pinkston, José Duato |
CLUSTER | 3 |
| 2003 | Evaluation of a Subnet Management Mechanism for InfiniBand NetworksabstractThe InfiniBand architecture is a high-performance network technology for the interconnection of processor nodes and I/O devices using a point-to-point switch-based fabric. The InfiniBand specification defines a basic management infrastructure that is responsible for subnet configuration, activation, and fault tolerance. Subnet management entities and functions are described, but the specifications do not impose any particular implementation. We present and analyze a complete subnet management mechanism for this architecture. We allow to anticipate future directions to obtain efficient management protocols Aurelio Bermúdez, Rafael Casado, Francisco J. Quiles 0001, Timothy M. Pinkston, José Duato |
ICPP | 3 |
| 2002 | A multimedia router architecture to provide high performance and QoS guarantees to mixed trafficabstractThe explosive growth in using scalable and cost-effective clusters and local area environments involve the design of high performance networks aimed at providing QoS to multimedia flows. Thus, the main goal pursued by the Multi-Media (MMR) project is to design a single-chip router able to efficiently handle multimedia flows and best-effort traffic. In this paper we focus on the performance evaluation of the MMR architecture using a mix of CBR, VBR and best effort workload. Preliminary simulation results show that, by using simple link and switch scheduling algorithms, the router is able to achieve a link bandwidth utilization of 80%, while still providing QoS guarantees to both CBR and VBR traffic in the presence of best-effort traffic. María Blanca Caminero, Carmen Carrión 0001, Francisco J. Quiles 0001, José Duato, Sudhakar Yalamanchili |
ICME (1) | 3 |
| 2002 | Improving the robustness of MPEG-4 video communications over wireless/3G mobile networksabstractTwo major issues in providing true end-to-end wireless/mobile video capabilities are: interoperability among network platforms and robustness of video compression algorithms in error-prone environments. In this paper, we mainly focus on the second issue and show how error resilience techniques can be used to improve the video quality. We argue that the error resilient tools provided within the MPEG-4 standard are not sufficient to provide acceptable quality in wireless/mobile networks, but that this quality can be significantly improved by the inclusion of hierarchical MPEG-4 video coding techniques. We present a novel hierarchical MPEG-4 video scheme particularly designed for video communications over QoS-capable wireless/mobile network. Francisco M. Delicado Martínez, Antonio Jose Garrido del Solo, Pedro Cuenca 0001, Luis Orozco-Barbosa, Francisco J. Quiles 0001 |
PIMRC | 5 |
| 2001 | Tuning Buffer Size in the Multimedia Router (MMR)abstractThe primary objective of the Multimedia Router (MMR) project is the design and implementation of a compact router optimized for multimedia applications. The router is targeted for use in cluster and LAN interconnection networks, which offer different constraints and therefore differing router solutions than WANs. One of the key design parameters is the amount of buffer space, which is closely related to the silicon area required to implement the router. In this paper, the MMR performance obtained when varying the size of the input buffers is explored. Preliminary results show that buffers as small as one flit large suffice to guarantee QoS to both CBR and VBR traffic, thanks to the use of flow control and short links. María Blanca Caminero, Carmen Carrión 0001, Francisco J. Quiles 0001, José Duato, Sudhakar Yalamanchili |
IPDPS | 3 |
| 2001 | Influence of Network Size and Load on the Performance of Reconfiguration ProtocolsabstractSwitched point-to-point interconnection networks provide the high bandwidth and low latency required by current distributed applications. When the topology changes, a reconfiguration of the routing tables is performed to maintain network connectivity. In order to prevent deadlock, traditional reconfiguration schemes discard application traffic during the reconfiguration process. The consequence is that the network cannot provide the bandwidth demanded by user applications. In order to solve this problem, we proposed two deadlock-free schemes that allow traffic through the network while the reconfiguration is being performed By using these schemes, the network is able to fulfill the applications requirements. In this paper, we evaluate these traditional and novel reconfiguration schemes. In particular, we analyze the impact of network size and load on their behavior. Application traffic has been modeled by means of a self-similar pattern. Simulation results clearly show the large performance degradation associated with the traditional approach and the significant benefits that can be obtained by using dynamic reconfiguration techniques. Rafael Casado, Aurelio Bermúdez, Francisco J. Quiles 0001, José Duato |
NCA | 3 |
| 2001 | A Protocol for Deadlock-Free Dynamic Reconfiguration in High-Speed Local Area NetworksabstractHigh-speed local area networks (LANs) consist of a set of switches interconnected by point-to-point links, and hosts linked to those switches through a network interface card. High-speed LANs may change their topology due to switches being turned on/off, hot expansion, link remapping, and component failures. In these cases, a distributed reconfiguration protocol analyzes the topology, computes the new routing tables, and downloads them to the corresponding switches. Unfortunately, in most cases, user traffic is stopped during the reconfiguration process to avoid deadlock. These strategies are called static reconfiguration techniques. Although network reconfigurations are not frequent, static reconfiguration such as this may take hundreds of milliseconds to execute, thus degrading system availability significantly. Several distributed real-time applications have strict communication requirements; Distributed multimedia applications have similar, although less strict, quality of service (QoS) requirements. Both stopping packet transmission and discarding packets due to the reconfiguration process prevent the system from satisfying the above requirements. Therefore, in order to support hard real-time and distributed multimedia applications over a high-speed LAN, we need to avoid stopping user traffic and discarding packets when the topology changes. In this paper, we propose a new deadlock-free distributed reconfiguration protocol that is able to asynchronously update routing tables without stopping user traffic. This protocol is valid for any topology, including regular as well as irregular topologies. It is also valid for packet switching as well as for cut-through switching techniques and does not rely on the existence of virtual channels to work. Simulation results show that the behavior of our protocol is significantly better than for other protocols based on stopping user traffic. Rafael Casado, Aurelio Bermúdez, José Duato, Francisco J. Quiles 0001, José L. Sánchez 0002 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2000 | Performance Evaluation of Dynamic Reconfiguration in High-Speed Local Area NetworksabstractHigh-speed local area networks (LANs) consist of a set of switches connected by point-to-point links, and hosts linked to switches through a network interface card. High-speed LANs may change their topology due to switches and hosts being turned on/off, link remapping, and component failures. In these cases, a distributed reconfiguration algorithm analyzes the topology, computes the new routing tables, and downloads them to the corresponding switches. Unfortunately, in most cases, user traffic is stopped during the reconfiguration process to avoid deadlock. Although network reconfigurations are not frequent, static reconfiguration such as this may take hundreds of milliseconds to execute, thus degrading system availability significantly. In this paper, we propose a new deadlock-free distributed reconfiguration algorithm that is able to asynchronously update routing tables without stopping user traffic. This algorithm is valid for any topology, including regular as well as irregular topologies. Simulation results show that the behavior of our algorithm is significantly better than for other algorithms based on a spanning-tree formation. Rafael Casado, Aurelio Bermúdez, Francisco J. Quiles 0001, José L. Sánchez 0002, José Duato |
HPCA | 3 |
| 2000 | Switch Scheduling in the Multimedia Router (MMR)abstractThe primary goal of the Multimedia Router (MMR) project is the design and implementation of a router optimized for multimedia applications. The router is targeted for use in cluster and LAN interconnection networks which offer different constraints and therefore differing router solutions than WANs. This paper describes and evaluates a switch scheduling algorithm based on a priority biasing scheme for dynamically updating the priorities of the connections established through the router. Unlike existing schemes that simply use the age of a flit as its priority, the novel feature of the proposed approach is that the priority is biased using the measured quality of service (QoS) values for the connection. Furthermore, the structure of the switch scheduling algorithm is motivated by opportunities for pipelined and concurrent operation so that scheduling decisions could be made at switching speeds. The performance of two of the many possible biasing functions is evaluated. Damon S. Love, Sudhakar Yalamanchili, José Duato, María Blanca Caminero, Francisco J. Quiles 0001 |
IPDPS | 5 |
| 2000 | Loss-resilient ATM protocol architecture for MPEG-2 video communicationsabstractMPEG-2 video communications over ATM networks is one of the most active research areas in the field of computer communications. In the transmission process of a variable bit rate video signal over an ATM network, cells are inevitably exposed to delays, errors, and losses due to the statistical multiplexing used in these networks. These phenomena affect the quality of the video signal, and without adequate measures to control the propagation of the impairments, the quality of the service may fall below acceptable levels. In this paper, we study the impact of cell losses on the quality of an MPEG-2 video sequence encoded in a variable bit rate mode. We introduce a set of control mechanisms at different levels of the protocol architecture to be used in MPEG-2-based video communications systems using ATM networks as their underlying transmission mechanism. Our results (using different video sequences) show the effectiveness to improve the video quality by using a structured set of control mechanisms to overcome for the loss of cells carrying VBR MPEG-2 video streams. We argue that in order to be able to create video systems able to cope with cell losses encountered in computer communications systems, a structured set of error-resilient protocol mechanisms is needed. Pedro Cuenca 0001, Luis Orozco-Barbosa, Francisco J. Quiles 0001, Antonio Jose Garrido del Solo |
IEEE J. Sel. Areas Commun. | 3 |
| 1999 | MMR: A High-Performance Multimedia Router - Architecture and Design Trade-OffsabstractThis paper presents the architecture of a router designed to efficiently support traffic generated by multimedia applications. The router is targeted for use in clusters and LANs rather than in WANs, the latter being served by communication substrates such as ATM. The distinguishing features of the proposed router architecture are the use of small fixed-size buffers, a large number of virtual channels, link-level virtual channel flow control, support for dynamic modification of connection bandwidth and priorities, and coordinated scheduling of connections across all output channels. The paper begins with a discussion of the design choices and architectural trade-offs made in the current MultiMedia Router (MMR) project. The performance evaluation section presents some preliminary results of the coordinated scheduling of constant bit rate (CBR) traffic streams. José Duato, Sudhakar Yalamanchili, María Blanca Caminero, Damon S. Love, Francisco J. Quiles 0001 |
HPCA | 5 |
| 1998 | Techniques to increase MPEG-2 error resilience in the VBR video transmission over ATM networksabstractWe introduce a set of control mechanisms at different levels of the protocol architecture to be used in MPEG-2-based video communications systems using ATM networks as their underlying transmission mechanism. Our results show the effectiveness of using a structured set of control mechanisms to overcome for the loss of cells carrying VBR video streams. Pedro Cuenca 0001, Teresa Olivares, Antonio Jose Garrido del Solo, Francisco J. Quiles 0001, Luis Orozco-Barbosa |
ICC | 4 |
| 1998 | Some Proposals to Improve Error Resilience in the MPEG-2 Video Transmission over ATM NetworksabstractMPEG-2 video communications over ATM networks is one of the most active research areas in the field of computer communications. In this paper, we introduce a set of control mechanisms at different levels of the protocol architecture to be used in MPEG-2-based video communications systems using ATM networks as their underlying transmission mechanism. We argue that in order to be able to create video systems able to cope with some of the errors encountered in computer communications systems, a structured set of error-resilient protocol mechanisms is needed. Our results show the effectiveness of using a structured set of control mechanisms to overcome for the loss of cells carrying VBR MPEG-2 video streams. Pedro Cuenca 0001, Antonio Jose Garrido del Solo, Francisco J. Quiles 0001, Luis Orozco-Barbosa |
INFOCOM | 3 |
| 1997 | Interconnection network behavior on a multicomputer in the parallelization of the MPEG coding algorithm. Worm-hole vs. packet-switching routingabstractWe propose the implementation of a MPEG encoder developed by the University of California at Berkeley on a multicomputer system. Since this application is in real time, we present a mapping of the video sequence between the EPs of the architecture, where the communication between EPs is minimized. We also propose the necessary load/store process with a simple mechanism input/output, where the global distribution process latency is compensated. Idonety of the topology of the system is analyzed, together with the most adequate commutation technique for the interconnection network. Finally the incidence of the frame format on the system communication performance is analyzed. Teresa Olivares, Pedro Cuenca 0001, Francisco J. Quiles 0001, Antonio Jose Garrido del Solo, José L. Sánchez 0002, José Duato |
HiPC | 3 |
| 1995 | A simulation tool of parallel architectures for digital image processing applications based on DLX processorsabstractWe present a simulation tool of parallel architectures for image digital treatment applications, which is characterized by the simulation of different architectures based on RISC DLX processors and an interconnection network based on wormhole routing. Each node is provided with a computational DLX processor and another of similar features, devoted to control the communication network. Valentín Valero Ruiz, Fernando Cuartero, Antonio Jose Garrido del Solo, Francisco J. Quiles 0001 |
ICIP (3) | 4 |