Jose Rocher-Gonzalez

dblp:207/3465 · also José Rocher-González · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
4since 2021 · last 2024
0000-0002-3179-5905ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 5 first-author · 4 since 2021
YearPublicationVenuePosition
2024 A smart and novel approach for managing incast and in-network congestion through adaptive routing
abstract
High-Performance Computing and Datacenter systems, with numerous endnodes, demand an efficient interconnection network to prevent performance bottlenecks. Fat-Tree topologies are preferred for their high bisection bandwidth and multiple shortest-path routes. While existing adaptive routing excels in light or in-network congestion, it struggles with incast congestion. This paper proposes a new technique, called Congestion-Aware Adaptive Routing (SCAR), which addresses both in-network and incast congestion. SCAR limits adaptivity for incast congestion, using deterministic routing, while employing adaptive routing for non-congesting flows. It also resolves in-network congestion by routing traffic flows through alternative routes. Simulation experiments on large Fat-Trees using synthetic and trace-based traffic patterns modeling realistic applications demonstrate SCAR’s immediate reaction on mitigating in-network congestion, and a reasonable delay during incast situations, while other state-of-the-art solutions are not able to cope with incast and in-network situations at the same time.
Jose Rocher-Gonzalez, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José Duato
Future Gener. Comput. Syst.1
2023 Congestion management in high-performance interconnection networks using adaptive routing notifications
Jose Rocher-Gonzalez, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001
J. Supercomput.1
2022 Adaptive Routing in InfiniBand Hardware
abstract
Interconnection networks are the communication backbone of modern high-performance computing systems and an optimised interconnection network is crucial for the performance and utilisation of the system as a whole. One element of the interconnection network is the routing algorithm, which directly influences how we are able to utilise the physical network topology. InfiniBand is one of the most common network architectures used in high-performance computing and traditionally it only supported static routing. For multi-path networks such as Fat-trees, static routing is inefficient because it cannot balance traffic in real-time nor utilise multiple paths efficiently under adversarial traffic. This again potentially leads to unnecessary contention and an underutilised network, which has led to numerous proposals on how to avoid this by using adaptive routing. Adaptive routing has recently been introduced in InfiniBand and in this paper we evaluate to what extent the expected benefits of adaptive routing is true for InfiniBand. Through a set of experiments on HDR InfiniBand equipment we describe the basic behaviour of adaptive routing in InfiniBand, its benefits in Fat tree topologies and the unfortunate side effects related to unfairness that adaptive routing in general might introduce, including such phenomena as the reverse parking lot problem and congestion spreading.
Jose Rocher-Gonzalez, Ernst Gunnar Gran, Sven-Arne Reinemo, Tor Skeie, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001
CCGRID1
2021 Towards an efficient combination of adaptive routing and queuing schemes in Fat-Tree topologies
Jose Rocher-Gonzalez, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, Gaspar Mora
J. Parallel Distributed Comput.1
2019 Efficient Congestion Management for High-Speed Interconnects using Adaptive Routing
abstract
The interconnection network is the central element in high-performance computing (HPC) clusters and Datacenters, where thousands of end nodes must communicate in a fast and reliable manner. The network performance depends on several design choices, such as the topology, the routing algorithm, the switch architecture, etc. Highly efficient routing algorithms, either deterministic or adaptive, have been proposed to smartly balance traffic flows in cost-effective network topologies, but their performance is reduced in scenarios where congestion and their negative effects (e.g. the HoL blocking) appear. In particular, in scenarios where congestion is intense and persistent, the HoL blocking may degrade dramatically the performance of adaptive routing algorithms, since they may spread congested traffic flows through all the available routes. In addition, as we have shown in previous studies, this spreading of congested flows may spoil the performance of the static queuing schemes that are used to reduce HoL blocking by separating flows into different queues at switch buffers. Indeed, as these schemes are based on a static criterion defined prior to the traffic injection in the network, they are unable to avoid that congested and non-congested flows share queues when paired with adaptive routing. In this paper, we propose to use some existing static queuing schemes and dynamic allocation of virtual channels (VCs) to isolate into a single VC the flows whose routes have been adaptively routed, in order to prevent the impact of the congestion spreading through several routes. Basically, adapted flows are moved to a special adapted-flow channel (AFC), so that they do not interact with flows mapped to other VCs by the static queuing scheme. In this way, the HoL blocking that adaptively routed flows could cause to non-adaptive flows is prevented, even if congested flows have been spread through several routes. On the other hand, the static queuing scheme will reduce without any interference the HoL blocking that may appear among non-adaptive flows. To evaluate our proposal we have conducted extensive simulation experiments modeling large interconnection networks based on the fat-tree topology. From the obtained results, we can conclude that our approach efficiently and significantly reduces the HoL blocking impact in interconnection networks using adaptive routing and queuing schemes when congestion appears.
Jose Rocher-Gonzalez, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, Gaspar Mora
CCGRID1
2019 Combining Source-adaptive and Oblivious Routing with Congestion Control in High-performance Interconnects using Hybrid and Direct Topologies
abstract
Hybrid and direct topologies are cost-efficient and scalable options to interconnect thousands of end nodes in high-performance computing (HPC) systems. They offer a rich path diversity, high bisection bandwidth, and a reduced diameter guaranteeing low latency. In these topologies, efficient deterministic routing algorithms can be used to balance smartly the traffic flows among the available routes. Unfortunately, congestion leads these networks to saturation, where the HoL blocking effect degrades their performance dramatically. Among the proposed solutions to deal with HoL blocking, the routing algorithms selecting alternative routes, such as adaptive and oblivious, can mitigate the congestion effects. Other techniques use queues to separate congested flows from non-congested ones, thus reducing the HoL blocking. In this article, we propose a new approach that reduces HoL blocking in hybrid and direct topologies using source-adaptive and oblivious routing. This approach also guarantees deadlock-freedom as it uses virtual networks to break potential cycles generated by the routing policy in the topology. Specifically, we propose two techniques, called Source-Adaptive Solution for Head-of-Line Blocking Avoidance (SASHA) and Oblivious Solution for Head-of-Line Blocking Avoidance (OSHA). Experiment results, carried out through simulations under different traffic scenarios, show that SASHA and OSHA can significantly reduce the HoL blocking.
Pedro Yébenes, Jose Rocher-Gonzalez, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Alfaro, Francisco J. Quiles 0001, Crispín Gómez Requena, José Duato
ACM Trans. Archit. Code Optim.2