EDBT 2026 Demo / reviewers in the wild / expert
Jesús Escudero-Sahuquillo
dblp:72/1341
· DBLP profile ↗
44ranked-venue papers
13as first author
18since 2021 · last 2026
0000-0003-0835-8624ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 43 · 13 first-author · 18 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On the power saving in high-speed Ethernet-based networks for supercomputers and data centersabstractThe increase in computation and storage has led to a significant growth in the scale of systems powering applications and services, raising concerns about sustainability and operational costs. In this paper, we explore power-saving techniques in high-performance computing (HPC) and data center networks, and their relation with performance degradation. From this premise, we propose leveraging the Energy Efficient Ethernet (EEE) protocol, with the flexibility to extend to conventional Ethernet or upcoming Ethernet-derived interconnect versions of BXI and Omnipath. We analyze the PerfBound power-saving mechanism, identifying possible improvements and modeling it into a simulation framework. Through different experiments, we examine its impact on performance and determine the most appropriate interconnect. We also study traffic patterns generated by selected HPC and machine learning applications to evaluate the behavior of power-saving techniques. From these experiments, we provide an analysis of how applications affect system and network energy consumption. Based on this, we disclose the weakness of dynamic power-down mechanisms and propose an approach that improves energy reduction with minimal or no performance penalty. This work presents a thorough analysis of PerfBound and an enhancement to the technique, while also targeting emerging post-exascale networks. Miguel Sánchez de la Rosa, Francisco J. Andujar, Jesús Escudero-Sahuquillo, José L. Sánchez 0002, Francisco J. Alfaro |
J. Syst. Archit. | 3 |
| 2026 | On the impact of intra- and inter-node communication in the performance of interconnection networks in HPC and AI systemsabstractAbstract In the last decade, specialized computing and storage devices, such as GPUs, TPUs, and high-speed storage, have been incorporated into server nodes of HPC and AI systems. The development of high-bandwidth memory (HBM) enabled a much more compact form factor for these devices, allowing the interconnection of several devices within a server node, typically using an intra-node interconnection network (e.g., PCIe, NVLink, or infinity fabric). Intra-node networks must allow efficient scale-up of the number of these devices within (and even beyond) a single node. Similarly, inter-node networks must enable efficient communication among hundreds of thousands of devices across thousands of server nodes when scale-up domains cannot be larger. Unfortunately, the intra- and inter-node networks may become the system’s bottleneck as communication demand among accelerators increases, driven by emerging applications such as generative AI. Although current intra-node network designs alleviate this bottleneck by increasing intra-node network bandwidth, intra-node communication may hinder inter-node communication performance when traffic from outside the server node arrives at the intra-node network. To evaluate the impact of this interference, we have analyzed intra- and inter-node communication operations generated under realistic traffic scenarios. We have developed a generic intra- and inter-node simulation model in OMNeT++ and have modeled the representative communication operations. We have also performed extensive simulation experiments confirming that increasing the intra-node network bandwidth and the number of computing devices per node (i.e., accelerators) may be counterproductive to the inter-node communication performance. Joaquín Tárraga, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001 |
J. Supercomput. | 2 |
| 2025 | Quality-of-service provision for BXIv3-based interconnection networksabstractAbstract Supercomputers (SCs) enable advanced research for a variety of scientific fields, and data centers (DCs) power our day-to-day services. These two massive systems work at scales, in terms of storage and computing power, which are not comparable to our everyday devices. As such, they require state-of-the-art technology to constantly evolve and meet our increasing demand. The interconnection network is the backbone of these systems, since it must provide efficient communication among the nodes that compose the whole system, otherwise becoming the entire system bottleneck. As multiple applications and services may use subsets of the system at the same time, interconnection networks must prevent excessive degradation for latency-sensitive applications. To this end, differentiated services are used to provide fair network access that considers bandwidth and latency requirements for each application. In this paper, we extend the switch architecture of next-generation BXI networks (hereafter called BXIv3) to incorporate arbitration tables so these networks can provide quality of service (QoS) to applications and services running on both SCs and DCs. Our proposal has been implemented in a network simulator, which models the behavior of a BXIv3 network. We have used several traffic patterns and arbitration table configurations to conduct a set of simulation experiments for the evaluation of our solution. The obtained results show that our proposal achieves accurate bandwidth allocation with differentiated latencies. Moreover, a study of memory requirements shows that our solution is quite feasible for hardware implementation. Miguel Sánchez de la Rosa, Gabriel Gomez-Lopez, Francisco J. Andujar, Jesús Escudero-Sahuquillo, José L. Sánchez 0002, Francisco J. Alfaro, Pierre-Axel Lagadec |
J. Supercomput. | 4 |
| 2025 | Distributed fast and accurate simulation platform for advanced ARM- and RISC-V-based HPC systems
Nikolaos Tampouratzis, Ioannis Papaefstathiou, Gabriel Gomez-Lopez, Miguel Sánchez de la Rosa, Jesús Escudero-Sahuquillo, Pedro Javier García |
J. Supercomput. | 5 |
| 2024 | A Hybrid Solution to Provide End-to-End Flow Control and Congestion Management in High-Performance Interconnection NetworksabstractCongestion seriously threatens high-performance interconnection networks in supercomputers and data centers, where thousands of server nodes generate massive communication operations when running highly parallel and distributed applications and services. In recent years, numerous solutions have been proposed to address congestion and its effects, including flow control (e.g., priority flow control, PFC) to prevent packet dropping at congested buffers, injection throttling to detect congested points and notify source server nodes to reduce the injection rate of congesting flows, and congestion isolation (as defined in the IEEE 802.1Qcz standard) that stores the congesting flows in separate queues or virtual channels (VCs) at switch buffers. Unfortunately, these solutions have exhibited important drawbacks, such as the prohibitive latency generated by flow control during congestion situations, the slow and ineffective response of injection throttling, or the excessive resources required to identify and isolate congesting flows. In this paper, we propose a hybrid congestion management solution, called 3SC (from three strategies combined), which combines end-to-end flow control, injection throttling, and congesting-flow isolation. 3SC significantly reduces Head-of-Line (HoL) blocking by isolating congesting flows in special queues at some switches, swiftly throttles the injection of these congesting flows, and significantly decreases the number of flow control messages compared to PFC. To evaluate our proposal, we have conducted a large set of simulation experiments for different network configurations and realistic traffic patterns. The results demonstrate that 3SC is efficient and feasible, making it a promising solution for future interconnection network designs. Alberto Merino, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, Yunping Lyu, José Duato |
CCGrid | 2 |
| 2024 | Hybrid Congestion Control for BXI-Based Interconnection Networks
Gabriel Gomez-Lopez, Miguel Sánchez de la Rosa, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, Pierre-Axel Lagadec |
Euro-Par (2) | 3 |
| 2024 | A New Mechanism to Identify Congesting Packets in High-Performance Interconnection NetworksabstractInterconnection networks are key components in Data Centers and Supercomputers, as they must guarantee high communication bandwidth and low latency under very demanding communication patterns generated by computing-and data-hungry applications and services. These traffic patterns may generate congestion, clogging different parts of the intercon-nection network, and impact the overall system performance if no countermeasures are taken. Unfortunately, congestion detection mechanisms used by congestion control techniques in current interconnection networks, such as DCQCN, do not precisely identify which packets contribute to generating congestion, so false-positive congestion detection events are possible. To overcome these problems, in this paper, we propose a new mechanism, called Enhanced Congestion Point (ECP), which accurately identifies packets that truly contribute to congestion. Specifically, ECP monitors packets at the head of the switch ingress queues and identifies them as congesting when a queue occupancy is over a given threshold and a crossing request for that packet within the switch is rejected. In addition, to solve the false-positive congestion-detection events, ECP defines are-evaluation mechanism that cancels the identification of congesting packets, if they no longer contribute to congestion after congestion areas have been sidestepped. We have evaluated ECP through different experiments using a network simulator that models different interconnection network configurations and realistic traffic patterns. This simulator also provides specific metrics for measuring the quality of the congestion detection mechanism. The obtained results show that ECP precisely identifies contesting packets with a low error margin, improving the DCQCN performance under congestion scenarios. Cristina Olmedilla, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, Yunping Lyu, José Duato |
HOTI | 2 |
| 2024 | Quality-of-Service Provision for BXI3-Based Interconnection NetworksabstractThe ever-increasing demand for computational power and storage capacity has led to massive Supercomputers and Data Centers running highly parallel applications and services commonly utilized in fields such as Physics, Biology, Robotics, Medicine, or generative AI. The interconnection network is the backbone of these systems since it allows processing and storage nodes to communicate with high bandwidth and low latency, otherwise becoming the entire system bottleneck. These systems commonly run several applications simultaneously, which may have specific network requirements due to technical or contractual reasons. Indeed, the traffic flows from different applications are expected to need different bandwidth and latency levels predetermined before execution. Therefore, Quality of Service (QoS) has been a recurrent design aspect for high-performance interconnection networks, as it happens for different technologies, such as Slingshot (Cray) or InfiniBand (NVIDIA). In this regard, as far as we know, no proposals have been made yet to provide applications with QoS for the upcoming generation of BXI (BXI3). We propose using arbitration tables to assign different priorities to packets when they are injected by NICs or forwarded by switches. We have conducted simulation experiments to evaluate our proposal, comparing several QoS configurations using synthetic workloads. The obtained results show that the proposed QoS approach is efficient and feasible so that it can be applied to the upcoming BXI3. Miguel Sánchez de la Rosa, Gabriel Gomez-Lopez, Francisco J. Andujar, Jesús Escudero-Sahuquillo, José L. Sánchez 0002, Francisco J. Alfaro, Pierre-Axel Lagadec |
HOTI | 4 |
| 2024 | A smart and novel approach for managing incast and in-network congestion through adaptive routingabstractHigh-Performance Computing and Datacenter systems, with numerous endnodes, demand an efficient interconnection network to prevent performance bottlenecks. Fat-Tree topologies are preferred for their high bisection bandwidth and multiple shortest-path routes. While existing adaptive routing excels in light or in-network congestion, it struggles with incast congestion. This paper proposes a new technique, called Congestion-Aware Adaptive Routing (SCAR), which addresses both in-network and incast congestion. SCAR limits adaptivity for incast congestion, using deterministic routing, while employing adaptive routing for non-congesting flows. It also resolves in-network congestion by routing traffic flows through alternative routes. Simulation experiments on large Fat-Trees using synthetic and trace-based traffic patterns modeling realistic applications demonstrate SCAR’s immediate reaction on mitigating in-network congestion, and a reasonable delay during incast situations, while other state-of-the-art solutions are not able to cope with incast and in-network situations at the same time. Jose Rocher-Gonzalez, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José Duato |
Future Gener. Comput. Syst. | 2 |
| 2024 | Implementation and testing of a KNS topology in an InfiniBand clusterabstractAbstract The InfiniBand (IB) interconnection technology is widely used in the networks of modern supercomputers and data centers. Among other advantages, the IB-based network devices allow for building multiple network topologies, and the IB control software (subnet manager) supports several routing engines suitable for the most common topologies. However, the implementation of some novel topologies in IB-based networks may be difficult if suitable routing algorithms are not supported, or if the IB switch or NIC architectures are not directly applicable for that topology. This work describes the implementation of the network topology known as KNS in a real HPC cluster using an IB network. As far as we know, this is the first implementation of this topology in an IB-based system. In more detail, we have implemented the KNS routing algorithm in the OpenSM software distribution of the subnet manager, and we have adapted the available IB-based switches to the particular structure of this topology. We have evaluated the correctness of our implementation through experiments in the real cluster, using well-known benchmarks. The obtained results, which match the expected performance for the KNS topology, show that this topology can be implemented in IB-based clusters as an alternative to other interconnection patterns. Gabriel Gomez-Lopez, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001 |
J. Supercomput. | 2 |
| 2023 | Extending the VEF traces framework to model data center network workloadsabstractAbstract Data centers are a fundamental infrastructure in the Big-Data era, where applications and services demand a high amount of data and minimum response times. The interconnection network is an essential subsystem in the data center, as it must guarantee high communication bandwidth and low latency to the communication operations of applications, otherwise becoming the system bottleneck. Simulation is widely used to model the network functionality and to evaluate its performance under specific workloads. Apart from the network modeling, it is essential to characterize the end-nodes communication pattern, which will help identify bottlenecks and flaws in the network architecture. In previous works, we proposed the VEF traces framework: a set of tools to capture communication traffic of MPI-based applications and generate traffic traces used to feed network simulator tools. In this paper, we extend the VEF traces framework with new communication workloads such as deep-learning training applications and online data-intensive workloads. Francisco J. Andujar, Miguel Sánchez de la Rosa, Jesús Escudero-Sahuquillo, José L. Sánchez 0002 |
J. Supercomput. | 3 |
| 2023 | Congestion management in high-performance interconnection networks using adaptive routing notifications
Jose Rocher-Gonzalez, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001 |
J. Supercomput. | 2 |
| 2022 | Adaptive Routing in InfiniBand HardwareabstractInterconnection networks are the communication backbone of modern high-performance computing systems and an optimised interconnection network is crucial for the performance and utilisation of the system as a whole. One element of the interconnection network is the routing algorithm, which directly influences how we are able to utilise the physical network topology. InfiniBand is one of the most common network architectures used in high-performance computing and traditionally it only supported static routing. For multi-path networks such as Fat-trees, static routing is inefficient because it cannot balance traffic in real-time nor utilise multiple paths efficiently under adversarial traffic. This again potentially leads to unnecessary contention and an underutilised network, which has led to numerous proposals on how to avoid this by using adaptive routing. Adaptive routing has recently been introduced in InfiniBand and in this paper we evaluate to what extent the expected benefits of adaptive routing is true for InfiniBand. Through a set of experiments on HDR InfiniBand equipment we describe the basic behaviour of adaptive routing in InfiniBand, its benefits in Fat tree topologies and the unfortunate side effects related to unfairness that adaptive routing in general might introduce, including such phenomena as the reverse parking lot problem and congestion spreading. Jose Rocher-Gonzalez, Ernst Gunnar Gran, Sven-Arne Reinemo, Tor Skeie, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001 |
CCGRID | 5 |
| 2022 | RED-SEA: Network Solution for Exascale ArchitecturesabstractIn order to enable Exascale computing, next generation interconnection networks must scale to hundreds of thousands of nodes, and must provide features to also allow the HPC, HPDA, and AI applications to reach Exascale, while benefiting from new hardware and software trends. RED-SEA will pave the way to the next generation of European Exascale interconnects, including the next generation of BXI, as follows: (i) specify the new architecture using hardware-software co-design and a set of applications representative of the new terrain of converging HPC, HPDA, and AI; (ii) test, evaluate, and/or implement the new architectural features at multiple levels, according to the nature of each of them, ranging from mathematical analysis and modeling, to simulation, or to emulation or implementation on FPGA testbeds; (iii) enable seamless communication within and between resource clusters, and therefore development of a high-performance low latency gateway, bridging seamlessly with Ethernet; (iv) add efficient network resource management, thus improving congestion resiliency, virtualization, adaptive routing, collective operations; (v) open the interconnect to new kinds of applications and hardware, with enhancements for end-to-end network services - from programming models to reliability, security, low- latency, and new processors; (vi) leverage open standards and compatible APIs to develop innovative reusable libraries and Fabrics management solutions. Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Pier Stanislao Paolucci, Elena Pastorelli, Francesco Simula, Matteo Turisini, Piero Vicini, Roberto Ammendola, Pascale Bernier-Bruna, Said Derradji, Stéphane Guez, Pierre-Axel Lagadec, Gregoire Pichon, Etienne Walter, Gaetan De Gassowski, Matthieu Hautreaux, Stephane Mathieu, Gilles Moreau, Marc Pérache, Hugo Taboada, Torsten Hoefler, Timo Schneider, Matteo Barnaba, Giuseppe Piero Brandino, Francesco De Giorgi, Matteo Poggi, Iakovos Mavroidis, Ioannis Papaefstathiou, Nikolaos Tampouratzis, Benjamin Kalisch, Ulrich Krackhardt, Mondrian Nüssle, Pantelis Xirouchakis, Vangelis Mageiropoulos, Michalis Gianioudis, Harisis Loukas, Aggelos Ioannou, Nikolaos D. Kallimanis, Nikolaos Chrysos, Manolis Katevenis, Wolfgang Frings, Dominik Gottwald, Felime Guimaraes, Max Holicki, Volker Marx, Yannik Müller, Carsten Clauss, Hugo Falter, Xu Huang 0010, Jennifer Lopez Barillao, Thomas Moschny, Simon Pickartz, Francisco J. Alfaro, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José L. Sánchez 0002, Adrián Castelló 0001, Jose Duro, María Engracia Gómez, Enrique S. Quintana-Ortí, Julio Sahuquillo, Eugenio Stabile |
DSD | 59 |
| 2022 | Improving Congestion Control through Fine-Grain Monitoring of InfiniBand NetworksabstractCongestion situations are a serious threat to the performance of the interconnection networks of High-Performance Computing and Data-Center systems. Hence, the specifications of the main interconnect technologies, such as InfiniBand, define some mechanisms to deal with congestion and its effects. However, these standard mechanisms may not be suitable to detect or track accurately the actual status of network congestion, as congestion dynamics indeed can be very complex and varied. Moreover, achieving an optimal configuration of the parameters that drive the different functionalities of congestion-control mechanisms is often a difficult task, as some configurations may be suitable for some traffic scenarios, but not for others. In this paper, we propose combining an existing light-weight platform monitoring tool (LIMITLESS) with the InfiniBand control software (OpenSM), such that the metrics about communication volumes in the network provided by the former allow the latter having a more precise image of congestion status, then being able to react more efficiently in these situations. The main contributions of this paper are the methodology to link the monitor and OpenSM, as well as modifications in the InfiniBand standard congestion-control mechanism so that its reaction is modulated based on the enhanced knowledge about congestion provided by the monitor. These improvements are ready to be integrated into any InfiniBand-based system. According to the results from our experiments (performed in a real InfiniBand-based cluster where we run a widely used benchmark), the proposed approach reduces significantly the number of wrong detections of congestion, and so the number of times that the congestion-control mechanisms react unnecessarily, hence improving system performance up to 74%. The overhead of this monitoring tool is 0.1% in our experiments, collecting data each 200ms. Alberto Cascajo, Gabriel Gomez-Lopez, Jesús Escudero-Sahuquillo, Pedro Javier García, David E. Singh, Francisco J. Alfaro, Francisco J. Quiles 0001, Jesús Carretero 0001 |
HOTI | 3 |
| 2021 | Leveraging InfiniBand controller to configure deadlock-free routing engines for Dragonflies
German Maglione Mathey, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, Eitan Zahavi |
J. Parallel Distributed Comput. | 2 |
| 2021 | Towards an efficient combination of adaptive routing and queuing schemes in Fat-Tree topologies
Jose Rocher-Gonzalez, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, Gaspar Mora |
J. Parallel Distributed Comput. | 2 |
| 2021 | A methodology to enable QoS provision on InfiniBand hardware
Javier Cano-Cano, Francisco J. Andujar, Jesús Escudero-Sahuquillo, Francisco J. Alfaro, José L. Sánchez 0002 |
J. Supercomput. | 3 |
| 2019 | Efficient Congestion Management for High-Speed Interconnects using Adaptive RoutingabstractThe interconnection network is the central element in high-performance computing (HPC) clusters and Datacenters, where thousands of end nodes must communicate in a fast and reliable manner. The network performance depends on several design choices, such as the topology, the routing algorithm, the switch architecture, etc. Highly efficient routing algorithms, either deterministic or adaptive, have been proposed to smartly balance traffic flows in cost-effective network topologies, but their performance is reduced in scenarios where congestion and their negative effects (e.g. the HoL blocking) appear. In particular, in scenarios where congestion is intense and persistent, the HoL blocking may degrade dramatically the performance of adaptive routing algorithms, since they may spread congested traffic flows through all the available routes. In addition, as we have shown in previous studies, this spreading of congested flows may spoil the performance of the static queuing schemes that are used to reduce HoL blocking by separating flows into different queues at switch buffers. Indeed, as these schemes are based on a static criterion defined prior to the traffic injection in the network, they are unable to avoid that congested and non-congested flows share queues when paired with adaptive routing. In this paper, we propose to use some existing static queuing schemes and dynamic allocation of virtual channels (VCs) to isolate into a single VC the flows whose routes have been adaptively routed, in order to prevent the impact of the congestion spreading through several routes. Basically, adapted flows are moved to a special adapted-flow channel (AFC), so that they do not interact with flows mapped to other VCs by the static queuing scheme. In this way, the HoL blocking that adaptively routed flows could cause to non-adaptive flows is prevented, even if congested flows have been spread through several routes. On the other hand, the static queuing scheme will reduce without any interference the HoL blocking that may appear among non-adaptive flows. To evaluate our proposal we have conducted extensive simulation experiments modeling large interconnection networks based on the fat-tree topology. From the obtained results, we can conclude that our approach efficiently and significantly reduces the HoL blocking impact in interconnection networks using adaptive routing and queuing schemes when congestion appears. Jose Rocher-Gonzalez, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, Gaspar Mora |
CCGRID | 2 |
| 2019 | Trends in High-Performance Interconnection Networks in the Exascale and Big-Data Era (HiPINEB 2017)abstractThe interconnection network plays an important role in the technological revolution brought by Exascale and Big-Data challenges to high-performance computing (HPC) systems and Datacenters. In these systems, the number of processing and storage nodes is growing significantly in order to meet the increasingly higher computing-power and storage demands. On the other hand, the overall power consumption of these systems must remain approximately as it is nowadays; otherwise, the cost of these systems in terms of acquisition and power consumption would be excessive. Therefore, the interconnection network must provide low latency and high-communication bandwidth, while requiring low power, otherwise becoming the bottleneck of the entire HPC system. In order to face these challenges and improve the network performance, several design aspects should be considered: network topology and connectivity, routing algorithm, reliability and fault tolerance, virtualization, congestion control, network configuration and control, development of the software stack, etc. This Special Issue of the Journal on Concurrency and Computation: Practice and Experience gathers several research works related to the latest advances and efforts in the design of high-performance interconnection networks, like those intended to be part of the near-future Exascale and Big-Data systems. In response to the Call-For-Papers of this Special Issue, we received very interesting submissions. Specifically, seven papers were submitted, which have been assessed by 21 reviewers, so that each paper has received three reviews. As can be seen in the list included in this editorial paper, all the reviewers are experts of the highest level, from both industry and academia, and their collaboration has been essential for the success of this Special Issue. We would like to thank the reviewers, which have provided not only detailed evaluations of the submissions but also valuable suggestions to enhance them. Based on the review and assessment process performed by the reviewers, we have been able to select five papers, which reflect prominent efforts and advances in the design and development of scalable high-performance interconnection networks for HPC systems and Datacenters. In that sense, Zahn et al1 analyze different aspects of energy proportionality in interconnection networks for systems designed within current technical constraints, but also for future systems that might be designed with different parameters. Benito et al2 explore a range of routing solutions based on exploiting explicit congestion notification messages (in particular, 802.1Qau) to adapt the number of packets using non-minimal paths. Andujar et al3 present how to build a 5DT torus network using a specific commercial 6-port network card (EXTOLL) to interconnect the system nodes. Tasoulas et al4 suggest solutions to rectify the space-domain scalability issues that are present in vSwitch-enabled subnets as a result of the VMs using dedicated layer-two addresses, also discussing routing strategies for virtualized environments using vSwitches and presenting a routing algorithm for Fat-Trees. Finally, Yébenes et al5 propose a combined mechanism to provide Slim Fly network with both non-minimal routing and queuing schemes by using several virtual networks to guarantee deadlock freedom. In our opinion, both the scope and the high technical quality of these papers make them very relevant for anyone involved, or just interested, in the Exascale and Big-data challenges, especially from the point-of-view of high-performance interconnection networks. List of reviewers We would like to thank the following reviewers for their support assessing the submitted papers to this Special Issue. Their reviews and discussions provided invaluable help for the guest-editors final decision. Dhabaleswar K. Panda, The Ohio State University, USA Enrique Vallejo, University of Cantabria, Spain Ernst Gunnar Gran, NTNU, Norway Evangelos Tasoulas, Schibsted Media Group, Norway Francisco J. Alfaro, University of Castilla-La Mancha, Spain Francisco J. Andújar, Technical University of Valencia, Spain Gaspar Mora, Intel Corporation, USA John Kim, KAIST, South Korea Jose Cano-Reyes, University of Glasgow, United Kingdom Jose Miguel Montañana, University of York, United Kingdom Julio Ortega, University of Granada, Spain Maria Engracia Gomez, Technical University of Valencia, Spain Matthieu Perotin, ATOS BULL, France Michihiro Koibuchi, National Institute of Informatics, Japan Nikos Chrysos, FORTH, Greece Pedro Lopez, Technical University of Valencia, Spain Peng Zhang, Stony Brook University, USA Ryan E. Grant, Sandia National Laboratories, USA Samuel Rodrigo, Graphcore, Norway Sebastien Rumley, Columbia University, USA Yuho Jin, New Mexico State University, USA Jesús Escudero-Sahuquillo, Pedro Javier García |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | Head-of-line blocking avoidance in Slim Fly networks using deadlock-free non-minimal and adaptive routingabstractSummary Interconnection network performance is a key issue in HPC systems and datacenters, especially as their number of end nodes grows, to cope with application needs. The network topology and the routing algorithm are important factors for performance and cost. Topologies such as fat‐tree or Dragonfly were proposed to maximize network performance while reducing network resources. One of the most promising topologies is Slim Fly, which offers high network bandwidth assuring low network diameter. However, adversarial traffic and/or congestion situations may degrade Slim Fly's performance dramatically. Non‐minimal routings, such as Valiant or UGAL, can mitigate the former problem while queuing schemes can handle the latter one. In this paper, we proposed a combined mechanism to provide Slim Fly network with both non‐minimal routing and queuing schemes by using several virtual networks to guarantee deadlock freedom. Each virtual network consists of a set of virtual channels to store packets separately according to a mapping policy. This diminishes the interaction among traffic flows, thus reducing head‐of‐line blocking. The results obtained from a simulation‐based evaluation show that our proposal enhances the performance in all the traffic cases, in contrast to other mechanisms whose performance drops in certain scenarios. Pedro Yébenes, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, Torsten Hoefler |
Concurr. Comput. Pract. Exp. | 2 |
| 2019 | Combining Source-adaptive and Oblivious Routing with Congestion Control in High-performance Interconnects using Hybrid and Direct TopologiesabstractHybrid and direct topologies are cost-efficient and scalable options to interconnect thousands of end nodes in high-performance computing (HPC) systems. They offer a rich path diversity, high bisection bandwidth, and a reduced diameter guaranteeing low latency. In these topologies, efficient deterministic routing algorithms can be used to balance smartly the traffic flows among the available routes. Unfortunately, congestion leads these networks to saturation, where the HoL blocking effect degrades their performance dramatically. Among the proposed solutions to deal with HoL blocking, the routing algorithms selecting alternative routes, such as adaptive and oblivious, can mitigate the congestion effects. Other techniques use queues to separate congested flows from non-congested ones, thus reducing the HoL blocking. In this article, we propose a new approach that reduces HoL blocking in hybrid and direct topologies using source-adaptive and oblivious routing. This approach also guarantees deadlock-freedom as it uses virtual networks to break potential cycles generated by the routing policy in the topology. Specifically, we propose two techniques, called Source-Adaptive Solution for Head-of-Line Blocking Avoidance (SASHA) and Oblivious Solution for Head-of-Line Blocking Avoidance (OSHA). Experiment results, carried out through simulations under different traffic scenarios, show that SASHA and OSHA can significantly reduce the HoL blocking. Pedro Yébenes, Jose Rocher-Gonzalez, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Alfaro, Francisco J. Quiles 0001, Crispín Gómez Requena, José Duato |
ACM Trans. Archit. Code Optim. | 3 |
| 2018 | Feasible enhancements to congestion control in InfiniBand-based networks
Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, German Maglione Mathey, José Duato |
J. Parallel Distributed Comput. | 1 |
| 2018 | Scalable Deadlock-Free Deterministic Minimal-Path Routing Engine for InfiniBand-Based Dragonfly NetworksabstractDragonfly topologies are gathering great interest nowadays as one of the most promising interconnect options for High-Performance Computing (HPC) systems. However, Dragonflies contain physical cycles that may lead to traffic deadlocks unless the routing algorithm prevents them properly. In general, existing deadlock-free routing algorithms, either deterministic or adaptive, proposed for Dragonflies, use Virtual Channels (VCs) to prevent cyclic dependencies. However, these topology-aware algorithms are difficult to implement, or even unfeasible, in systems based on the InfiniBand (IB) architecture, which is nowadays the most widely used network technology in HPC systems. This is due to some limitations in the IB specification, specifically regarding the way Virtual Lanes (VLs), which are considered as similar to VCs, can be assigned to traffic flows. Indeed, none of the routing engines currently available in the official releases of the IB control software has been specifically proposed for Dragonflies. In this paper, we present a new deterministic, minimal-path routing for Dragonfly that prevents deadlocks using VLs according to the IB specification, so that it can be straightforwardly implemented in IB-based networks. We have called this proposal D3R (Deterministic Deadlock-free Dragonfly Routing). Specifically, D3R maps each route to a single, specific VL depending on the destination group, and according to a specific order, so that cyclic dependencies (so deadlocks) are prevented. D3R is scalable as it requires only 2 VLs to prevent deadlocks regardless of network size, i.e., fewer VLs than the required by the deadlock-free routing engines available in IB that are suitable for Dragonflies. Alternatively, D3R achieves higher throughput if an additional VL is used to reduce internal contention in the Dragonfly groups. We have implemented D3R as a new routing engine in OpenSM, the control software including the subnet manager in IB. We have evaluated D3R by means of simulation and by experiments performed in a real IB-based cluster, the results showing that, in general, D3R outperforms other routing engines. German Maglione Mathey, Pedro Yébenes, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, Eitan Zahavi |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | Special issue on trends in high-performance interconnection networks in the exascale and big-data eraabstractSpecial issue on trends in high-performance interconnection networks in the exascale and big-data era Guest EditorsThe interconnection network plays an important role in the technological revolution brought by exascale and big-data challenges to high-performance computing (HPC) systems and datacenters.In these systems, the number of processing and storage nodes is growing significantly to meet the higher-computing power and storage demands.Therefore, the interconnection network of these systems must provide low latency and high-communication bandwidth; otherwise, it would become the bottleneck of the entire system.In addition, the power consumption of these systems should not increase significantly with respect to current values; otherwise, the cost of these systems in terms of power consumption would be excessive.Therefore, the interconnection network of these systems must provide the required high performance while reducing as much as possible its power consumption.To face these challenges, there are several design activities that should be considered: network topology and connectivity, routing algorithm, reliability and fault tolerance, virtualization, congestion control, network configuration and control, development of the software stack, etc. Jesús Escudero-Sahuquillo, Pedro Javier García |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | Providing differentiated services, congestion management, and deadlock freedom in dragonfly networks with adaptive routingabstractSummary The number of endnodes in high‐performance computing systems has grown significantly in the last years. Hence, the interconnection network has become an essential issue as it may end up being the system bottleneck if it is not properly designed. In that sense, the Dragonfly topology has become very popular for interconnecting high‐performance computing systems in the last years because it offers high performance at an affordable cost. However, when using deterministic minimal‐path routing, this topology is not able to offer a high performance under certain traffic conditions. This problem can be solved by using oblivious or adaptive routing. However, there are no congestion management techniques specially tailored to Dragonfly topologies using oblivious or adaptive routing. Note that in congestion situations, the Dragonfly performance may drop because of the head‐of‐line blocking effect. This effect could be even more dangerous in systems where several applications with different priorities coexist. In this work we propose several techniques especially designed for providing differentiated services and congestion management in Dragonfly networks using oblivious or adaptive routing. First, we propose thehierarchical 3‐level queuingqueuing scheme, which configures several virtual channels distributed into 3 virtual networks to reduce the head‐of‐line blocking while deadlocks derived from the routing algorithm are prevented. Second, we extendhierarchical 3‐level queuingto provide differentiated services through 2 different solutions. Finally, some experiments are performed to show the benefits obtained by using the proposed techniques. Pedro Yébenes, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Alfaro, Francisco J. Quiles 0001 |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | Providing differentiated services, congestion management, and deadlock freedom in dragonfly networks with adaptive routingabstractIn this article,1 an error in one of the author names has been found subsequent to the publication. “Jesus Escudero-Sahuquilllo” should be “Jesus Escudero-Sahuquillo.” The correct name is now presented above. The author's name has also been corrected in the original published article. Pedro Yébenes, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Alfaro, Francisco J. Quiles 0001 |
Concurr. Comput. Pract. Exp. | 2 |
| 2016 | An open-source family of tools to reproduce MPI-based workloads in interconnection network simulators
Francisco J. Andujar, Juan A. Villar, Francisco J. Alfaro, José L. Sánchez 0002, Jesús Escudero-Sahuquillo |
J. Supercomput. | 5 |
| 2016 | High-performance interconnection networks in the Exascale and Big-Data Era
Jesús Escudero-Sahuquillo, Pedro Javier García |
J. Supercomput. | 1 |
| 2016 | Straightforward solutions to reduce HoL blocking in different Dragonfly fully-connected interconnection patterns
Pedro Yébenes, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001 |
J. Supercomput. | 2 |
| 2015 | VEF Traces: A Framework for Modelling MPI Traffic in Interconnection Network SimulatorsabstractSimulation is often used to evaluate the behaviour and measure the performance of computing systems. Specifically, in high-performance interconnection networks, the simulation has been extensively considered to verify the behaviour of the network itself and to evaluate its performance. In this context, network simulation must be fed with network traffic, also referred to as network workload, whose nature has been traditionally synthetic. These workloads can be used for the purpose of driving studies on network performance, but often such workloads are not accurate enough if a realistic evaluation is pursued. For this reason, other non-synthetic workloads have gained popularity over last decades since they are best to capture the realistic behaviour of existing applications. In this paper, we present the VEF traces framework, a self-related trace model, and all their associated tools. The main novelty of this framework is that, unlike existing ones, it does not provide a network simulation framework, but only offers an MPI task simulation framework, which allows one to use the MPI-based network traffic by any third-party network simulator, since this framework does not depend on any specific simulation platform. Francisco J. Andujar, Juan A. Villar, José L. Sánchez 0002, Francisco J. Alfaro, Jesús Escudero-Sahuquillo |
CLUSTER | 5 |
| 2015 | Efficient Queuing Schemes for HoL-Blocking Reduction in Dragonfly Topologies with Minimal-Path RoutingabstractHPC systems are growing in number of connected endnodes, making the network a main issue in their design. In order to interconnect large systems, dragonfly topologies have become very popular in the latest years as they achieve high scalability by exploiting high-radix switches. However, dragonfly high performance may drop severely due to the Head-of-Line (HoL) blocking effect derived from congestion situations. Many techniques have been proposed for dealing with this harmful effect, the most effective ones being those especially designed for a specific topology and a specific routing algorithm. In this paper we present a queuing scheme called Hierarchical Two-Levels Queuing, designed specially to reduce HoL blocking in fully-connected dragonfly networks that use minimal-path routing. This proposal boosts network performance compared with other techniques and requires fewer network resources than the others. Besides, an upgrade for existing queuing schemes for improving their performance is explained. Pedro Yébenes, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001 |
CLUSTER | 2 |
| 2015 | Efficient and Cost-Effective Hybrid Congestion Control for HPC Interconnection NetworksabstractInterconnection networks are key components in high-performance computing (HPC) systems, their performance having a strong influence on the overall system one. However, at high load, congestion and its negative effects (e.g., Head-of-line blocking) threaten the performance of the network, and so the one of the entire system. Congestion control (CC) is crucial to ensure an efficient utilization of the interconnection network during congestion situations. As one major trend is to reduce the effective wiring in interconnection networks to reduce cost and power consumption, the network will operate very close to its capacity. Thus, congestion control becomes essential. Existing CC techniques can be divided into two general approaches. One is to throttle traffic injection at the sources that contribute to congestion, and the other is to isolate the congested traffic in specially designated resources. However, both approaches have different, but non-overlapping weaknesses: injection throttling techniques have a slow reaction against congestion, while isolating traffic in special resources may lead the system to run out of those resources. In this paper we propose EcoCC, a new Efficient and Cost-Effective CC technique, that combines injection throttling and congested-flow isolation to minimize their respective drawbacks and maximize overall system performance. This new strategy is suitable for current commercial switch architectures, where it could be implemented without requiring significant complexity. Experimental results, using simulations under synthetic and real trace-based traffic patterns, show that this technique improves by up to 55 percent over some of the most successful congestion control techniques. Jesús Escudero-Sahuquillo, Ernst Gunnar Gran, Pedro Javier García, José Flich, Tor Skeie, Olav Lysne, Francisco J. Quiles 0001, José Duato |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2014 | Combining HoL-blocking avoidance and differentiated services in high-speed interconnectsabstractCurrent high-performance platforms such as Datacenters or High-Performance Computing systems rely on highspeed interconnection networks able to cope with the ever-increasing communication requirements of modern applications. In particular, in high-performance systems that must offer differentiated services to applications which involve traffic prioritization, it is almost mandatory that the interconnection network provides some type of Quality-of-Service (QoS) and Congestion-Management mechanism in order to achieve the required network performance. Most current QoS and Congestion-Management mechanisms for high-speed interconnects are based on using the same kind of resources, but with different criteria, resulting in disjoint types of mechanisms. By contrast, we propose in this paper a novel, straightforward solution that leverages the resources already available in InfiniBand components (basically Service Levels and Virtual Lanes) to provide both QoS and Congestion Management at the same time. This proposal is called CHADS (Combined HoL-blocking Avoidance and Differentiated Services), and it could be applied to any network topology. From the results shown in this paper for networks configured with the novel, cost-efficient KNS hybrid topology, we can conclude that CHADS is more efficient than other schemes in reducing the interferences among packet flows that have the same or different priorities. Pedro Yébenes, Jesús Escudero-Sahuquillo, Crispín Gómez Requena, Pedro Javier García, Francisco J. Alfaro, Francisco J. Quiles 0001, José Duato |
HiPC | 2 |
| 2014 | A new proposal to deal with congestion in InfiniBand-based fat-trees
Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, Sven-Arne Reinemo, Tor Skeie, Olav Lysne, José Duato |
J. Parallel Distributed Comput. | 1 |
| 2013 | BBQ: A Straightforward Queuing Scheme to Reduce HoL-Blocking in High-Performance Hybrid Networks
Pedro Yébenes, Jesús Escudero-Sahuquillo, Crispín Gómez Requena, Pedro Javier García, Francisco J. Quiles 0001, José Duato |
Euro-Par | 2 |
| 2013 | Towards Modeling Interconnection Networks of Exascale Systems with OMNet++abstractOne of the objectives of the decade for High-Performance Computing systems is to reach the exascale level of computing power before 2018, hence this will require strong efforts in their design. In that sense, High-speed low-latency interconnection networks are essential elements for exascale HPC systems. Indeed, the performance of the whole system depends on that of the interconnection network. In order to develop and test new techniques, suited to exascale HPC systems, software-based networks simulators are commonly used. As developing a network simulator from scratch is a difficult task, several platforms help the developers, OMNeT++ being one of the most popular. In this paper, we propose a new generic network simulator, exploiting the features of the OMNeT++ framework. The proposed tool is the first step to model HPC high-performance interconnection networks of exascale HPC systems: the message switching layer, routing and arbitration algorithms and buffer organizations have been modeled according to the current and expected characteristics of these systems. In addition, the tool has been designed so that it is possible to simulate networks of large size. Simulation results, validated against real systems, show the accuracy of the model. Pedro Yébenes, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001 |
PDP | 2 |
| 2013 | An Effective and Feasible Congestion Management Technique for High-Performance MINs with Tag-Based Distributed RoutingabstractAs parallel computing systems increase in size, the interconnection network is becoming a critical subsystem. The current trend in network design is to use as few components as possible to interconnect the end nodes, thereby reducing cost and power consumption. However, this increases the probability of congestion appearing in the network. As congestion may severely degrade network performance, the use of a congestion management mechanism is becoming mandatory in modern interconnects. One of the most cost-effective proposals to deal with the problems derived from congestion situations is the Regional Explicit Congestion Notification (RECN) strategy, based on using special queues to totally isolate the packet flows which contribute to congestion, thereby preventing the Head-of-Line (HoL) blocking effect that these flows may cause to others. Unfortunately, RECN requires the use of source-based routing, thus not being suitable for interconnects with distributed routing, like InfiniBand. Although some RECN-like mechanisms have been proposed for distributed-routing networks, they are not scalable due to the huge amount of control memory that they require in medium-size or large networks. In this paper, we propose Distributed-Routing-Based Congestion Management (DRBCM), a new scalable technique which, following the RECN principles, totally prevents congestion from producing HoL-blocking in multistage interconnection networks (MINs) using tag-based distributed routing. Simulation results indicate that, regardless of network size, DRBCM presents small resource requirements to keep network performance at maximum level even in scenarios of heavy congestion, where it utterly outperforms (with a gain up to 70 percent) current solutions for distributed-routing networks, like the InfiniBand congestion-control mechanism based on injection throttling. Thus, DRBCM is an efficient, cost-effective, and scalable solution for congestion management. Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José Flich, José Duato |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2011 | Combining Congested-Flow Isolation and Injection Throttling in HPC Interconnection NetworksabstractExisting congestion control mechanisms in interconnects can be divided into two general approaches. One is to throttle traffic injection at the sources that contribute to congestion, and the other is to isolate the congested traffic in specially designated resources. These two approaches have different, but non-overlapping weaknesses. In this paper we present in detail a method that combines injection throttling and congested-flow isolation. Through simulation studies we first demonstrate the respective flaws of the injection throttling and of flow isolation. Thereafter we show that our combined method extracts the best of both approaches in the sense that it gives fast reaction to congestion, it is scalable and it has good fairness properties with respect to the congested flows. Jesús Escudero-Sahuquillo, Ernst Gunnar Gran, Pedro Javier García, José Flich, Tor Skeie, Olav Lysne, Francisco J. Quiles 0001, José Duato |
ICPP | 1 |
| 2011 | Cost-effective queue schemes for reducing head-of-line blocking in fat-treesabstractSUMMARY The fat‐tree is one of the most common topologies among the interconnection networks of the systems currently used for high‐performance parallel computing. Among other advantages, fat‐trees allow the use of simple but very efficient routing schemes. One of them is a deterministic routing algorithm that has been recently proposed, offering a similar (or better) performance than adaptive routing while reducing complexity and guaranteeing in‐order packet delivery. However, as other deterministic routing proposals, this deterministic routing algorithm cannot react when high traffic loads or hot‐spot traffic scenarios produce severe contention for the use of network resources, leading to the appearance of Head‐of‐Line (HoL) blocking, which spoils the network performance. In that sense, we describe in this paper two simple, cost‐effective strategies for dealing with the HoL‐blocking problem that may appear in fat‐trees with the aforementioned deterministic routing algorithm. From the results presented in the paper, we conclude that, in the mentioned environment, these proposals considerably reduce HoL‐blocking without significantly increasing switch complexity and the required silicon area. Copyright © 2011 John Wiley & Sons, Ltd. Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José Flich, José Duato |
Concurr. Comput. Pract. Exp. | 1 |
| 2011 | OBQA: Smart and cost-efficient queue scheme for Head-of-Line blocking elimination in fat-trees
Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José Flich, José Duato |
J. Parallel Distributed Comput. | 1 |
| 2010 | An Efficient Strategy for Reducing Head-of-Line Blocking in Fat-Trees
Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José Duato |
Euro-Par (2) | 1 |
| 2010 | Cost-Effective Congestion Management for Interconnection Networks Using Distributed Deterministic RoutingabstractThe Interconnection networks are essential elements in current computing systems. For this reason, achieving the best network performance, even in congestion situations, has been a primary goal in recent years. In that sense, there exist several techniques focused on eliminating the main negative effect of congestion: the Head of Line (HOL) blocking. One of the most successful HOL blocking elimination techniques is RECN, which can be applied in source routing networks. FBICM follows the same approach as RECN, but it has been developed for distributed deterministic routing networks. Although FBICM effectively eliminates HOL blocking, it requires too much resources to be implemented. In this paper we present a new FBICM version, based on a new organization of switch memory resources, that significantly reduces the required silicon area, complexity and cost. Moreover, we present new results about FBICM, in network topologies not yet analyzed. From the experiment results we can conclude that a far less complex and feasible FBICM implementation can be achieved by using the proposed improvements, while not losing efficiency. Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José Flich, José Duato |
ICPADS | 1 |
| 2008 | FBICM: Efficient Congestion Management for High-Performance Networks Using Distributed Deterministic Routing
Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José Flich, José Duato |
HiPC | 1 |