EDBT 2026 Demo / reviewers in the wild / expert
Pedro López 0001
dblp:15/6063 · also Pedro Juan López Rodríguez
· DBLP profile ↗
106ranked-venue papers
7as first author
5since 2021 · last 2026
0000-0003-4544-955XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 95 · 7 first-author · 5 since 2021Software engineering, systems software and programming languages · 3Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TinyATD: Compressing Auxiliary Tag Directories using tag hashing and set-samplingabstractAuxiliary Tag Directories (ATD) are hardware structures widely analyzed in academia to estimate the interference that a task experiences when sharing a cache in a multithreaded system. ATDs are used for execution time inflation calculations, reducing power consumption in multiprocessor systems, improving cache replacement, and improving cache prefetching accuracy. However, the commercial adoption of ATD-based approaches seems to be limited due to the area overheads that they impose. We propose TinyATD, a mechanism that aims to improve state-of-the-art ATD area reduction by combining ATD set sampling and tag hashing area reduction techniques to build smaller and more precise ATDs. TinyATD achieves a 41% area reduction for the same error or a 23% error reduction for the same area over the baseline set-sampling when sampling 32 out of 2048 sets. Furthermore, TinyATD offers the stated reduction over a wide array of commonly used LLC replacement policies. Pablo Andreu, Pedro López 0001, Carles Hernández 0001 |
J. Syst. Archit. | 2 |
| 2026 | HashTAG With CALM: Low-Overhead Hardware Support for Inter-Task Eviction MonitoringabstractMulticore processors have emerged as the preferred architecture for safetycritical systems due to their significant performance advantages. However, concurrent access by multiple cores to a shared cache induces intercore evictions that generate nondeterministic interference and compromise timing predictability. Static partitioning of the cache among cores is a wellestablished countermeasure that effectively eliminates such evictions but reduces flexibility and system throughput. To accurately estimate inter-core cache contention, Auxiliary Tag Directories (ATDs) are widely adopted. However, ATDs incur substantial hardware area costs, which often motivates the use of heuristic-based reductions. These reduced ATD designs, while more compact, compromise accuracy and therefore are not suitable for safety-critical domains. This paper extends the proposal of HashTAG, a novel approach to accurately upper-bound inter-core eviction interference. HashTAG introduces a safe and lightweight Auxiliary Tag Directory mechanism that tracks which cores are responsible for evicting cache lines used by others, thus measuring contention. We further refine the proposed HashTAG approach by creating CALM, a custom-made memory allocator that significantly improves HashTAG performance in multicore systems. Our results show that no inter-task interference underprediction is possible with HashTAG, making it suitable for the safety domain. HashTAG provides a 47% reduction in the Auxiliary Tag Directory area, presenting perfect measurements on 80% of cases and only a 1% error on maximum inter-core eviction measurements for a HashTAG tag size of ten bits. Pablo Andreu, Pedro López 0001, Carles Hernández 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2025 | Expanding SafeSU capabilities by leveraging security frameworks for contention monitoring in complex SoCsabstractThe increased performance requirements of applications running on safety-critical systems have led to the use of complex platforms with several CPUs, GPUs, and AI accelerators. However, higher platform and system complexity challenge performance verification and validation since timing interference across tasks occurs in unobvious ways, hence defeating attempts to optimize application consolidation informedly during design phases and validating that mutual interference across tasks is within bounds during test phases. In that respect, the SafeSU has been proposed to extend inter-task interference monitoring capabilities in simple systems. However, modern mixed-criticality systems are complex, with multilayered interconnects, shared caches, and hardware accelerators. To that end, this paper proposes a non-intrusive add-on approach for monitoring interference across tasks in multilayer heterogeneous systems implemented by leveraging existing security frameworks and the SafeSU infrastructure. The feasibility of the proposed approach has been validated in an RTL RISC-V-based multicore SoC with support for AI hardware acceleration. Our results show that our approach can safely track contention and properly break down contention cycles across the different sources of interference, hence guiding optimization and validation processes. • Enables inter-core interference monitoring in complex SoCs. • Tool to check inter-task timing independence at all NoC levels. • Proposes unified initiator naming to solve initiator dilution across SoC layers. Pablo Andreu, Sergi Alcaide, Pedro López 0001, Jaume Abella 0001, Carles Hernández 0001 |
Future Gener. Comput. Syst. | 3 |
| 2023 | Energy efficient HPC network topologies with on/off linksabstractEnergy efficiency is a must in today HPC systems. To achieve this goal, a holistic design based on the use of power-aware components should be performed. One of the key components of an HPC system is the high-speed interconnect. In this paper, we compare and evaluate several design options for the interconnection network of an HPC system, including torus, fat-trees and dragonflies. State of the art low power modes are also used in the interconnection networks. The paper does not only consider energy efficiency at the interconnection network level but also at the system as a whole. The analysis is performed by using a simple yet realistic power model of the system. The model has been adjusted using actual power consumption values measured on a real system. Using this model, realistic multi-job trace-based workloads have been used, obtaining the execution time and energy consumed. The results are presented to ease choosing a system, depending on which parameter, performance or energy consumption, receives the most importance. Francisco J. Andujar, Salvador Coll, Marina Alonso, Juan-Miguel Martinez-Rubio, Pedro López 0001, José L. Sánchez 0002, Francisco J. Alfaro |
Future Gener. Comput. Syst. | 5 |
| 2021 | Improving the Robustness of Redundant Execution with Register File RandomizationabstractStaggered Redundant execution (SRE) is a fault-tolerance mechanism that has been widely deployed in the context of safety-critical applications. SRE not only protects the system in the presence of faults but also helps relaxing safety requirements of individual elements. However, in this paper, we show that SRE does not effectively protect the system against a wide range of faults and thus, new mechanisms to increase the diversity of homogeneous cores are needed. In this paper, we propose Register File Randomization (RFR), a low-cost diversity mechanism that significantly increases the robustness of homogeneous multicores in front of common-cause faults (CCFs) and register file wearout. Our results show that RFR completely removes the failure rate for register file CCFs for certain workloads and reduces by a factor of 5X the impact of stress related register file aging for the workloads analysed. Our implementation requires less than 50 RTL lines of code and the area (FPGA logic) overhead of RFR is less than 0.2% of a 64-bit RISC-V core FPGA implementation. Ilya Tuzov, Pablo Andreu, Laura Medina, Tomás Picornell, Antonio Robles, Pedro López 0001, José Flich, Carles Hernández 0001 |
ICCAD | 6 |
| 2019 | Energy efficient torus networks with on/off links
Francisco J. Andujar, Salvador Coll, Marina Alonso, Juan-Miguel Martinez-Rubio, Pedro López 0001, José L. Sánchez 0002, Francisco J. Alfaro, Raúl Martínez |
J. Parallel Distributed Comput. | 5 |
| 2019 | POWAR: Power-Aware Routing in HPC Networks with On/Off LinksabstractIn order to save energy in HPC interconnection networks, one usual proposal is to switch idle links into a low-power mode after a certain time without any transmission, as IEEE Energy Efficient Ethernet standard proposes. Extending the low-power mode mechanism, we propose POW er- A ware R outing ( POWAR ), a simple power-aware routing and selection function for fat-tree and torus networks. POWAR adapts the amount of network links that can be used, taking into account the network load, and obtaining great energy savings in the network (55%--65%) and the entire system (9%--10%) with negligible performance overhead. Francisco J. Andujar, Salvador Coll, Marina Alonso, Pedro López 0001, Juan-Miguel Martinez-Rubio |
ACM Trans. Archit. Code Optim. | 4 |
| 2018 | Teaching high-performance service in a cluster computing course
Pedro López 0001, Elvira Baydal |
J. Parallel Distributed Comput. | 1 |
| 2017 | A fault-tolerant routing strategy for k-ary n-direct s-indirect topologies based on intermediate nodesabstractSummary Exascale computing systems are being built with thousands of nodes. The high number of components of these systems significantly increases the probability of failure. A key component for them is the interconnection network. If failures occur in the interconnection network, they may isolate a large fraction of the machine. For this reason, an efficient fault‐tolerant mechanism is needed to keep the system interconnected, even in the presence of faults. A recently proposed topology for these large systems is the hybridk‐aryn‐directs‐indirect family that provides optimal performance and connectivity at a reduced hardware cost. This paper presents a fault‐tolerant routing methodology for thek‐aryn‐directs‐indirect topology that degrades performance gracefully in presence of faults and tolerates a large number of faults without disabling any healthy computing node. In order to tolerate network failures, the methodology uses a simple mechanism. For any source‐destination pair, if necessary, packets are forwarded to the destination node through a set of intermediate nodes (without being ejected from the network) with the aim of circumventing faults. The evaluation results shows that the proposed methodology tolerates a large number of faults. For instance, it is able to tolerate more than 99.5% of fault combinations when there are 10 faults in a 3‐D network with 1000 nodes using only 1 intermediate node and more than 99.98% if 2 intermediate nodes are used. Furthermore, the methodology offers a gracious performance degradation. As an example, performance degrades only by 1% for a 2‐D network with 1024 nodes and 1% faulty links. Roberto Peñaranda, María Engracia Gómez, Pedro López 0001, Ernst Gunnar Gran, Tor Skeie |
Concurr. Comput. Pract. Exp. | 3 |
| 2017 | On a course on computer cluster configuration and administration
Pedro López 0001, Elvira Baydal |
J. Parallel Distributed Comput. | 1 |
| 2017 | XOR-based HoL-blocking reduction routing mechanisms for direct networks
Roberto Peñaranda, Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001 |
Parallel Comput. | 4 |
| 2016 | Embedded GPU and multicore processors for emotional-based mobile robotic agents
Francisco Almenar, Carlos Domínguez, Houcine Hassan, Juan-Miguel Martinez-Rubio, Pedro López 0001 |
Future Gener. Comput. Syst. | 5 |
| 2016 | The k-ary n-direct s-indirect family of topologies for large-scale interconnection networks
Roberto Peñaranda, Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, José Duato |
J. Supercomput. | 4 |
| 2016 | A Family of Fault-Tolerant Efficient Indirect TopologiesabstractOn the one hand, performance and fault-tolerance of interconnection networks are key design issues for high performance computing (HPC) systems. On the other hand, cost should be also considered. Indirect topologies are often chosen in the design of HPC systems. Among them, the most commonly used topology is the fat-tree. In this work, we focus on getting the maximum benefits from the network resources by designing a simple indirect topology with very good performance and fault-tolerance properties, while keeping the hardware cost as low as possible. To do that, we propose some extensions to the fat-tree topology to take full advantage of the hardware resources consumed by the topology. In particular, we propose three new topologies with different properties in terms of cost, performance and fault-tolerance. All of them are able to achieve a similar or better performance results than the fat-tree, providing also a good level of fault-tolerance and, contrary to most of the available topologies, these proposals are able to tolerate also faults in the links that connect to end nodes. Diego F. Bermúdez Garzón, Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, José Duato |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2015 | XORAdap: A HoL-Blocking Aware Adaptive Routing AlgorithmabstractRouting is a key parameter in the design of the interconnection network of large parallel computers. Depending on the number of routing options available for each packet, routing algorithms can be deterministic (one available path) or adaptive (several ones). Adaptive routing usually outperforms deterministic routing but it also may increase the Head-of-Line blocking effect. Usually, adaptive routing uses virtual channels to provide routing flexibility and to guarantee deadlock freedom. On the other hand, deterministic routing is simpler and therefore it has lower routing delay. In this paper, we take the challenge of developing new routing algorithms for direct topologies that exploit virtual channels in an efficient way combining the good properties of both routing algorithms types: flexibility and reduced HoL blocking. To do that, this paper proposes several hybrid (combination of adaptivity and determinism) simple mechanisms to perform an efficient distribution of packets among virtual channels based on their destination. The resulting routing mechanisms are able to adapt to the different traffic patterns to obtain the best performance while keeping the simplicity of routing. Roberto Peñaranda, Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001 |
PDP | 4 |
| 2015 | Power consumption management in fat-tree interconnection networks
Marina Alonso, Salvador Coll, Juan-Miguel Martinez-Rubio, Vicente Santonja, Pedro López 0001 |
Parallel Comput. | 5 |
| 2015 | Design of Hybrid Second-Level CachesabstractIn recent years, embedded dynamic random-access memory (eDRAM) technology has been implemented in last-level caches due to its low leakage energy consumption and high density. However, the fact that eDRAM presents slower access time than static RAM (SRAM) technology has prevented its inclusion in higher levels of the cache hierarchy. This paper proposes to mingle SRAM and eDRAM banks within the data array of second-level (L2) caches. The main goal is to achieve the best trade-off among performance, energy, and area. To this end, two main directions have been followed. First, this paper explores the optimal percentage of banks for each technology. Second, the cache controller is redesigned to deal with performance and energy. Performance is addressed by keeping the most likely accessed blocks in fast SRAM banks. In addition, energy savings are further enhanced by avoiding unnecessary destructive reads of eDRAM blocks. Experimental results show that, compared to a conventional SRAM L2 cache, a hybrid approach requiring similar or even lower area speedups the performance on average by 5.9 percent, while the total energy savings are by 32 percent. For a 45 nm technology node, the energy-delay-area product confirms that a hybrid cache is a better design than the conventional SRAM cache regardless of the number of eDRAM banks, and also better than a conventional eDRAM cache when the number of SRAM banks is an eighth of the total number of cache banks. Alejandro Valero, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
IEEE Trans. Computers | 4 |
| 2015 | A HoL-blocking aware mechanism for selecting the upward path in fat-tree topologies
Crispín Gómez Requena, Francisco Gilabert Villamón, María Engracia Gómez, Pedro López 0001, José Duato |
J. Supercomput. | 4 |
| 2014 | FT-RUFT: A Performance and Fault-Tolerant Efficient Indirect TopologyabstractAlthough performance is a key design issue of interconnection networks, fault-tolerance is becoming more important due to the large amount of components of large machines. In this paper, we focus on designing a simple indirect topology with both good performance and fault-tolerance properties. The idea is to take full advantage of the network resources consumed by the topology. To do that, starting from the RUFT topology, which is a simple UMIN topology that does not tolerate any link fault, we first duplicate injection and ejection links connecting these extra links in a particular way. The resulting topology tolerates 3 network link faults and also slightly increases performance with marginal increase in the network hardware cost. Most important, contrary to most of the available topologies, the topology is able to tolerate also faults in the links that connect to end-nodes. We also propose another topology that also duplicates network links, achieving 2x performance improvements and tolerating up to 7 network link faults. These results are better than the ones obtained by a BMIN with a similar amount of resources. Diego F. Bermúdez Garzón, Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, José Duato |
PDP | 4 |
| 2014 | Efficient Register Renaming and Recovery for High-Performance ProcessorsabstractModern superscalar processors implement register renaming using either random access memory (RAM) or content-addressable memories (CAM) tables. The design of these structures should address both access time and misprediction recovery penalty. Although direct-mapped RAMs provide faster access times, CAMs are more appropriate to avoid recovery penalties. The presence of associative ports in CAMs, however, prevents them from scaling with the number of physical registers and pipeline width, negatively impacting performance, area, and energy consumption at the rename stage. In this paper, we present a new hybrid RAM-CAM register renaming scheme, which combines the best of both approaches. In a steady state, a RAM provides fast and energy-efficient access to register mappings. On misspeculation, a low-complexity CAM enables immediate recovery. Experimental results show that in a four-way state-of-the-art superscalar processor, the new approach provides almost the same performance as an ideal CAM-based renaming scheme, while dissipating only between 17% and 26% of the original energy and, in some cases, consuming less energy than purely RAM-based renaming schemes. Overall, the silicon area required to implement the hybrid RAM-CAM scheme does not exceed the area required by conventional renaming mechanisms. Salvador Petit, Rafael Ubal, Julio Sahuquillo, Pedro López 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2013 | Combining RAM technologies for hard-error recovery in L1 data caches working at very-low power modesabstractLow-power modes in modern microprocessors rely on low frequencies and low voltages to reduce the energy budget. Nevertheless, manufacturing induced parameter variations can make SRAM cells unreliable producing hard errors at supply voltages below Vccmin. Vicente Lorente, Alejandro Valero, Julio Sahuquillo, Salvador Petit, Ramon Canal, Pedro López 0001, José Duato |
DATE | 6 |
| 2013 | Topic 13: High-Performance Networks and Communication - (Introduction)
Olav Lysne, Torsten Hoefler, Pedro López 0001, Davide Bertozzi |
Euro-Par | 3 |
| 2013 | Hardware-Based Generation of Independent Subtraces of Instructions in Clustered ProcessorsabstractMulticore chips are currently dominating the microprocessor market as designs that improve performance and sustain power consumption. However, complex core features must be still considered to provide good performance for existing sequential applications. An effective approach to reduce core complexity without dramatically sacrificing performance is to distribute critical processor structures by using clustered microarchitectures. In these designs, communication latency among clusters is a critical performance bottleneck, and a good steering algorithm is required to reduce intercluster communication. In this paper, we propose a new energy-efficient microarchitectural approach that reduces intercluster communication by detecting and generating independent chains of instructions, referred to as subtraces, from the execution of sequential programs. The devised mechanism has been modeled on an x86-based trace-cache processor, where subtraces are built in the fill unit, stored in a trace cache, and individually steered to different clusters. Experimental results show that the proposal reaches performance speedups around 7 and 15 percent for point-to-point and bus-based interconnects, respectively, while achieving energy savings of up to 12 percent. Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
IEEE Trans. Computers | 4 |
| 2012 | Towards an Efficient Fat-Tree like Topology
Diego F. Bermúdez Garzón, Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, José Duato |
Euro-Par | 4 |
| 2012 | Analyzing the optimal ratio of SRAM banks in hybrid cachesabstractCache memories have been typically implemented with Static Random Access Memory (SRAM) technology. This technology presents a fast access time but high energy consumption and low density. As opposite, the recently appeared embedded Dynamic RAM (eDRAM) technology allows caches to be built with lower energy and area, although with a slower access time. The eDRAM technology provides important leakage and area savings, especially in huge Last-Level Caches (LLCs), which occupy almost half the silicon area in some recent microprocessors. This paper proposes a novel hybrid LLC, which combines SRAM and eDRAM banks to address the trade-off among performance, energy, and area. To this end, we explore the optimal percentage of SRAM and eDRAM banks that achieves the best target trade-off. Architectural mechanisms have been devised to keep the most likely accessed blocks in fast SRAM banks as well as to avoid unnecessary destructive reads. Experimental results show that, compared to a conventional SRAM LLC with the same storage capacity, performance degradation does not surpass, on average, 2.9% (even with 12.5% of banks built with SRAM technology), whereas area savings can be as high as 46% for a 1MB-16way LLC. For a 45nm technology node, the energy-delay squared product confirms that a hybrid cache is a better design than the conventional SRAM cache regardless the number of eDRAM banks, and also better than a conventional eDRAM cache when the number of SRAM banks is a quarter or an eighth of the cache banks. Alejandro Valero, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
ICCD | 4 |
| 2012 | IODET: A HoL-blocking-aware Deterministic Routing Algorithm for Direct TopologiesabstractIn large parallel computers routing is a key design point to obtain the maximum possible performance out of the interconnection network. Routing can be classified into two categories depending on the number of routing options that a packet can use to go from its source to its destination. If the packet can only use a single predetermined path then the routing is deterministic, whereas if several paths are possible it is adaptive. It is a well-known fact that adaptive routing usually outperforms deterministic routing; but in this paper we take the challenge of developing a HOL-blocking-aware deterministic routing algorithm that can obtain a similar or even better performance than adaptive routing, while decreasing its implementation complexity and providing some inherent advantages to deterministic routing such as in-order delivery of packets. In this large computers regular direct topologies are widely-used, so in this paper we focus on meshes and tori. Roberto Peñaranda, Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, José Duato |
ICPADS | 4 |
| 2012 | A New Family of Hybrid Topologies for Large-Scale Interconnection NetworksabstractIn large supercomputers the topology of the interconnection network is a key design issue that impacts the performance and cost of the whole system. Direct topologies provide a reduced hardware cost, but as the number of dimensions is conditioned by 3D wiring restrictions, a high number of nodes per dimension is used, which increases communication latency and reduces network throughput. On the other hand, indirect topologies can provide better performance for large network sizes, but at the cost of a high amount of switches and links. In this paper we propose a new family of topologies that combines the best features of both direct and indirect topologies to efficiently connect an extremely high number of nodes. In particular, we propose an n-dimensional topology where the nodes of each dimension are connected through a small indirect topology. This combination results in a family of topologies that provides high performance, with latency and throughput figures of merit close to indirect topologies, but with a lower hardware cost. In particular, it is able to double the throughput obtained per switching element of indirect topologies. Moreover, the layout of the topology is much simpler than in indirect topologies. Indeed, its fault-tolerance degree is equal or higher than the one for direct and indirect topologies. Roberto Peñaranda, Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, José Duato |
NCA | 4 |
| 2012 | Combining recency of information with selective random and a victim cache in last-level cachesabstractMemory latency has become an important performance bottleneck in current microprocessors. This problem aggravates as the number of cores sharing the same memory controller increases. To palliate this problem, a common solution is to implement cache hierarchies with large or huge Last-Level Cache (LLC) organizations. LLC memories are implemented with a high number of ways (e.g., 16) to reduce conflict misses. Typically, caches have implemented the LRU algorithm to exploit temporal locality, but its performance goes away from the optimal as the number of ways increases. In addition, the implementation of a strict LRU algorithm is costly in terms of area and power. This article focuses on a family of low-cost replacement strategies, whose implementation scales with the number of ways while maintaining the performance. The proposed strategies track the accessing order for just a few blocks, which cannot be replaced. The victim is randomly selected among those blocks exhibiting poor locality. Although, in general, the random policy helps improving the performance, in some applications the scheme fails with respect to the LRU policy leading to performance degradation. This drawback can be overcome by the addition of a small victim cache of the large LLC. Experimental results show that, using the best version of the family without victim cache, MPKI reduction falls in between 10% and 11% compared to a set of the most representative state-of-the-art algorithms, whereas the reduction grows up to 22% with respect to LRU. The proposal with victim cache achieves speedup improvements, on average, by 4% compared to LRU. In addition, it reduces dynamic energy, on average, up to 8%. Finally, compared to the studied algorithms, hardware complexity is largely reduced by the baseline algorithm of the family. Alejandro Valero, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
ACM Trans. Archit. Code Optim. | 4 |
| 2012 | Progressive Congestion Management Based on Packet Marking and Validation TechniquesabstractCongestion management in multistage interconnection networks is a serious problem, which is not solved completely. In order to avoid the degradation of network performance when congestion appears, several congestion management mechanisms have been proposed. Most of these mechanisms are based on explicit congestion notification. For this purpose, switches detect congestion and depending on the applied strategy, packets are marked to warn the source hosts. In response, source hosts apply some corrective actions to adjust their packet injection rate. Although these proposals seem quite effective, they either exhibit some drawbacks or are partial solutions. Some of them introduce some penalties over the flows not responsible for congestion, whereas others can cope only with congestion situations that last for a short time. In this paper, we present an overview of the different strategies to detect and correct congestion in multistage interconnection networks, and propose a new mechanism referred to as Marking and Validation Congestion Management (MVCM), targeted to this kind of lossless networks, and based on a more refined packet marking strategy combined with a fair set of corrective actions, that makes the mechanism able to effectively manage congestion regardless of the congestion degree. Evaluation results show the effectiveness and robustness of the proposed mechanism. Joan-Lluís Ferrer, Elvira Baydal, Antonio Robles, Pedro López 0001, José Duato |
IEEE Trans. Computers | 4 |
| 2012 | Design, Performance, and Energy Consumption of eDRAM/SRAM Macrocells for L1 Data CachesabstractSRAM and DRAM have been the predominant technologies used to implement memory cells in computer systems, each one having its advantages and shortcomings. SRAM cells are faster and require no refresh since reads are not destructive. In contrast, DRAM cells provide higher density and minimal leakage energy since there are no paths within the cell from Vdd to ground. Recently, DRAM cells have been embedded in logic-based technology (eDRAM), thus overcoming the speed limit of typical DRAM cells. In this paper, we propose a hybrid n-bit macrocell that implements one SRAM cell and n-1 eDRAM cells. This cell is aimed at being used in an n-way set-associative first-level data cache. Architectural mechanisms (e.g., special writeback policies) have been devised to completely avoid refresh logic. Performance, energy, and area have been analyzed in detail. Experimental results show that using typical eDRAM capacitors, and compared to a conventional cache, a 4-way set-associative hybrid cache reduces both energy consumption and area up to 54 and 29 percent, respectively, while having negligible impact on performance (less than 2 percent). Alejandro Valero, Salvador Petit, Julio Sahuquillo, Pedro López 0001, José Duato |
IEEE Trans. Computers | 4 |
| 2012 | A Survey and Evaluation of Topology-Agnostic Deterministic Routing AlgorithmsabstractMost standard cluster interconnect technologies are flexible with respect to network topology. This has spawned a substantial amount of research on topology-agnostic routing algorithms, which make no assumption about the network structure, thus providing the flexibility needed to route on irregular networks. Actually, such an irregularity should be often interpreted as minor modifications of some regular interconnection pattern, such as those induced by faults. In fact, topology-agnostic routing algorithms are also becoming increasingly useful for networks on chip (NoCs), where faults may make the preferred 2D mesh topology irregular. Existing topology-agnostic routing algorithms were developed for varying purposes, giving them different and not always comparable properties. Details are scattered among many papers, each with distinct conditions, making comparison difficult. This paper presents a comprehensive overview of the known topology-agnostic routing algorithms. We classify these algorithms by their most important properties, and evaluate them consistently. This provides significant insight into the algorithms and their appropriateness for different on- and off-chip environments. José Flich, Tor Skeie, Andres Mejia, Olav Lysne, Pedro López 0001, Antonio Robles, José Duato, Michihiro Koibuchi, Tomas Rokicki, José Carlos Sancho |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2012 | A Sequentially Consistent Multiprocessor Architecture for Out-of-Order Retirement of InstructionsabstractOut-of-order retirement of instructions has been shown to be an effective technique to increase the number of in-flight instructions. This form of runtime scheduling can reduce pipeline stalls caused by head-of-line blocking effects in the reorder buffer (ROB). Expanding the width of the instruction window can be highly beneficial to multiprocessors that implement a strict memory model, especially when both loads and stores encounter long latencies due to cache misses, and whose stalls must be overlapped with instruction execution to overcome the memory latencies. Based on the Validation Buffer (VB) architecture (a previously proposed out-of-order retirement, checkpoint-free architecture for single processors), this paper proposes a cost-effective, scalable, out-of-order retirement multiprocessor, capable of enforcing sequential consistency without impacting the design of the memory hierarchy or interconnect. Our simulation results indicate that utilizing a VB can speed up both relaxed and sequentially consistent in-order retirement in future multiprocessor systems by between 3 and 20 percent, depending on the ROB size. Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001, David R. Kaeli |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2012 | Impact on Performance and Energy of the Retention Time and Processor Frequency in L1 Macrocell-Based Data CachesabstractCache memories dissipate an important amount of the energy budget in current microprocessors. This is mainly due to cache cells are typically implemented with six transistors. To tackle this design concern, recent research has focused on the proposal of new cache cells. Ann-bit cache cell, namely macrocell, has been proposed in a previous work. This cell combines SRAM and eDRAM technologies with the aim of reducing energy consumption while maintaining the performance. The capacitance of eDRAM cells impacts on energy consumption and performance since these cells lose their state once the retention time expires. On such a case, data must be fetched from a lower level of the memory hierarchy, so negatively impacting on performance and energy consumption. As opposite, if the capacitance is too high, energy would be wasted without bringing performance benefits. This paper identifies the optimal capacitance for a given processor frequency. To this end, the tradeoff between performance and energy consumption of a macrocell-based cache has been evaluated varying the capacitance and frequency. Experimental results show that, compared to a conventional cache, performance losses are lower than 2% and energy savings are up to 55% for a cache with 10 fF capacitors and frequencies higher than 1 GHz. In addition, using trench capacitors, a 4-bit macrocell reduces by 29% the area of four conventional SRAM cells. Alejandro Valero, Julio Sahuquillo, Vicente Lorente, Salvador Petit, Pedro López 0001, José Duato |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2011 | Improving Last-Level Cache Performance by Exploiting the Concept of MRU-TourabstractLast-Level Caches (LLCs) implement the LRU algorithm to exploit temporal locality, but its performance is quite far of Belady's optimal algorithm as the number of ways increases. One of the main reasons because of LRU does not reach good performance in LLCs is that this policy forces a block to descend until the bottom of the stack before eviction. Nevertheless, most of the blocks that leave the MRU position are not referenced again before eviction. This work pursues to select candidate blocks to be victimized before reaching the bottom of the stack. To this end, this work defines the number of MRU-Tours (MRUTs) of a block as the number of times that a block enters in the MRU position during its live time. Based on the fact that most of the blocks exhibit a single MRUT, this work presents the family of MRUT-based algorithms aimed at exploiting this block behavior to improve performance. Alejandro Valero, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
PACT | 4 |
| 2011 | MRU-Tour-based Replacement Algorithms for Last-Level CachesabstractMemory hierarchy design is a major concern in current microprocessors. Many research work focuses on the Last-Level Cache (LLC), which is designed to hide the long miss penalty of accessing to main memory. To reduce both capacity and conflict misses, LLCs are implemented as large memory structures with high associativities. To exploit temporal locality, LRU is the replacement algorithm usually implemented in caches. However, for a high-associative cache, its implementation is costly in terms of area and power consumption. Indeed, LRU is not well suited for the LLC, because as this cache level does not see all memory accesses, it cannot cope with temporal locality. In addition, blocks must descend down to the LRU position of the stack before eviction, even when they are not longer useful. In this paper, we show that most of the blocks are not referenced again once they leave the MRU position. Moreover, the probability of being referenced again does not depend on the location on the LRU stack. Based on these observations, we define the number of MRU-Tours (MRUTs) of a block as the number of times that a block occupies the MRU position while it is stored in the cache, and propose the MRUT replacement algorithm, which selects the block to be replaced among the blocks that show only one MRUT. Variations of this algorithm have been also proposed to exploit both MRUT behavior and recency of information. Experimental results show that, compared to LRU, the proposal reduces the MPKI up to 22%, while IPC is improved by 48%. Alejandro Valero, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
SBAC-PAD | 4 |
| 2011 | How to reduce packet dropping in a bufferless NoCabstractAbstract Networks on‐chip (NoCs) interconnect the components located inside a chip. In multicore chips, NoCs have a strong impact on the overall system performance. NoC bandwidth is limited by the critical path delay. Recent works show that the critical path delay is heavily affected by switch port buffer size. Therefore, by removing buffers, switch clock frequency can be increased. Recently, a new switching technique for NoCs called Blind Packet Switching (BPS) has been proposed, which is based on removing the switch port buffers. Since buffers consume a high percentage of switch power and area, BPS not only improves performance but also reduces power and area. In BPS, as there are no buffers at the switch ports, packets cannot be stopped and stored on them. If contention arises packets are dropped and later reinjected, negatively affecting performance. In order to prevent packet dropping, some techniques based on resource replication have been proposed. In this paper, we propose some alternative and complementary techniques that do not rely on resource replication. By using them, packet dropping is highly reduced. In particular, packet dropping is completely removed for a very wide network traffic range. Moreover, network throughput is increased and packet latency is reduced. Copyright © 2010 John Wiley & Sons, Ltd. Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, José Duato |
Concurr. Comput. Pract. Exp. | 3 |
| 2010 | Exploiting subtrace-level parallelism in clustered processorsabstractNo abstract available. Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
PACT | 4 |
| 2010 | Out-of-order retirement of instructions in sequentially consistent multiprocessorsabstractOut-of-order retirement of instructions has been shown to be an effective technique to increase the number of in-flight instructions. This form of runtime scheduling can reduce pipeline stalls caused by head-of-line blocking effects in the reorder buffer (ROB). Wide instruction windows are very beneficial to multiprocessors that implement a strict memory model, especially when both loads and stores encounter long latencies due to cache misses, and whose stalls must be overlapped with instruction execution to overcome the memory gap. In this paper, the Validation Buffer (VB) multiprocessor architecture is proposed as a cost-effective, checkpoint-free, scalable approach to retire instructions out of program order, while still enforcing sequential consistency, and without impacting the memory hierarchy or interconnect. Experimental results show that utilizing the Validation Buffer can speed up both release and sequentially consistent in-order retirement in future multiprocessor systems by between 3% and 20%, depending on the ROB size. Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001, David R. Kaeli |
ICCD | 4 |
| 2010 | A Scalable and Early Congestion Management Mechanism for MINsabstractSeveral packet marking-based mechanisms have been proposed to manage congestion in multistage interconnection networks. One of them, the MVCM mechanism obtains very good results for different network configurations and traffic loads. However, as MVCM applies full virtual output queuing at origin, its memory requirements may jeopardize its scalability. Additionally, the applied packet marking technique introduces certain delay to detect congestion. In this paper, we propose and evaluate the Scalable Early Congestion Management mechanism which eliminates the drawbacks exhibited by MVCM. The new mechanism replaces the full virtual output queuing at origin by either a partial virtual output queuing or a shared buffer, in order to reduce its memory requirements, thus making the mechanism scalable. Also, it applies an improved packet marking technique based on marking packets at output buffers regardless of their marking at input buffers, which simplifies the marking technique, allowing also a sooner detection of the root of a congestion tree. Joan-Lluís Ferrer, Elvira Baydal, Antonio Robles, Pedro López 0001, José Duato |
PDP | 4 |
| 2010 | Power saving in regular interconnection networks
Marina Alonso, Salvador Coll, Juan-Miguel Martinez-Rubio, Vicente Santonja, Pedro López 0001, José Duato |
Parallel Comput. | 5 |
| 2009 | Anaphase: A Fine-Grain Thread Decomposition Scheme for Speculative MultithreadingabstractIndustry is moving towards multi-core designs as we have hit the memory and power walls. Multi-core designs are very effective to exploit thread-level parallelism (TLP) but do not provide benefits when executing serial code (applications with low TLP, serial parts of a parallel application and legacy code). In this paper we propose Anaphase, a novel approach for speculative multithreading to improve single-thread performance in a multi-core design. The proposed technique is based on a graph partitioning technique which performs a decomposition of applications into speculative threads at instruction granularity. Moreover, the proposed technique leverages communications and pre-computation slices to deal with inter-thread dependences. Results presented in this paper show that this approach improves single-thread performance by 32% on average and up to 2.15x for some selected applications of the Spec2006 suite. In addition, the proposed technique outperforms by 21% on average schemes in which thread decomposition is performed at a coarser granularity. Carlos Madriles, Pedro López 0001, Josep M. Codina, Enric Gibert, Fernando Latorre, Alejandro Martínez, Raúl Martínez, Antonio González 0001 |
PACT | 2 |
| 2009 | Assessing fat-tree topologies for regular network-on-chip design under nanoscale technology constraintsabstractMost of past evaluations of fat-trees for on-chip interconnection networks rely on oversimplifying or even irrealistic architecture and traffic pattern assumptions, and very few layout analyses are available to relieve practical feasibility concerns in nanoscale technologies. This work aims at providing an in-depth assessment of physical synthesis efficiency of fat-trees and at extrapolating silicon-aware performance figures to back-annotate in the system-level performance analysis. A 2D mesh is used as a reference architecture for comparison, and a 65 nm technology is targeted by our study. Finally, in an attempt to mitigate the implementation cost of k-ary n-tree topologies, we also review an alternative unidirectional multi-stage interconnection network which is able to simplify the fat-tree architecture and to minimally impact performance. Daniele Ludovici, Francisco Gilabert Villamón, Simone Medardoni, Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, Georgi Gaydadjiev, Davide Bertozzi |
DATE | 6 |
| 2009 | An Efficient Low-Complexity Alternative to the ROB for Out-of-Order Retirement of InstructionsabstractCurrent superscalar processors use a reorder buffer (ROB) to support speculation, precise exceptions, and register reclamation. Instructions are retired from this structure in program order, which may lead to significant performance degradation if a long latency operation blocks the ROB head. In this paper, a checkpoint-free out-of-order commit architecture is proposed, which replaces the ROB with a small structure called validation buffer (VB) from which instructions are retired as soon as their speculative state is resolved. An aggressive register reclamation mechanism targeted to this microarchitecture is also devised. Experimental results show that the VB microarchitecture is much more efficient than a ROB-based microprocessor. For example, a 32-entry VB provides similar performance to a 256-entry ROB, while reducing the utilization of other major processor structures. Salvador Petit, Rafael Ubal, Julio Sahuquillo, Pedro López 0001, José Duato |
DSD | 4 |
| 2009 | Paired ROBs: A Cost-Effective Reorder Buffer Sharing Strategy for SMT Processors
Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001 |
Euro-Par | 4 |
| 2009 | A power-aware hybrid RAM-CAM renaming mechanism for fast recoveryabstractModern superscalar processors implement register renaming by using either RAM or CAM tables. The design of these structures should address their access time and misprediction recovery penalty. While direct-mapped RAMs provide faster access times, CAMs are more appropriate to avoid recovery penalties. Although they are more complex and slower, CAMs usually match the processor cycle in current designs. However, they do not scale with the number of physical registers and the pipeline width. In this paper we present a new hybrid RAM-CAM register renaming scheme, which combines the best of both approaches. In a steady state, a RAM provides the current mappings quickly; on mispeculation, a low-complexity CAM enables immediate recovery and further register renaming. Compared to an ideal CAM in a 4-way state-of-the-art superscalar microprocessor, and for almost the same performance (1% slowdown) and area (95% of the ideal CAM size), the proposed scheme consumes about 90% less dynamic energy. Salvador Petit, Rafael Ubal, Julio Sahuquillo, Pedro López 0001 |
ICCD | 4 |
| 2009 | Boosting single-thread performance in multi-core systems through fine-grain multi-threadingabstractIndustry has shifted towards multi-core designs as we have hit the memory and power walls. However, single thread performance remains of paramount importance since some applications have limited thread-level parallelism (TLP), and even a small part with limited TLP impose important constraints to the global performance, as explained by Amdahl's law. Carlos Madriles, Pedro López 0001, Josep M. Codina, Enric Gibert, Fernando Latorre, Alejandro Martínez, Raúl Martínez, Antonio González 0001 |
ISCA | 2 |
| 2009 | An hybrid eDRAM/SRAM macrocell to implement first-level data cachesabstractSRAM and DRAM cells have been the predominant technologies used to implement memory cells in computer systems, each one having its advantages and shortcomings. SRAM cells are faster and require no refresh since reads are not destructive. In contrast, DRAM cells provide higher density and minimal leakage energy since there are no paths within the cell from Vdd to ground. Recently, DRAM cells have been embedded in logic-based technology, thus overcoming the speed limit of typical DRAM cells. Alejandro Valero, Julio Sahuquillo, Salvador Petit, Vicente Lorente, Ramon Canal, Pedro López 0001, José Duato |
MICRO | 6 |
| 2009 | A Complexity-Effective Out-of-Order Retirement MicroarchitectureabstractCurrent superscalar processors commit instructions in program order by using a reorder buffer (ROB). The ROB provides support for speculation, precise exceptions, and register reclamation. However, committing instructions in program order may lead to significant performance degradation if a long latency operation blocks the ROB head. Several proposals have been published to deal with this problem. Most of them retire instructions speculatively. However, as speculation may fail, checkpoints are required in order to rollback the processor to a precise state, which requires both extra hardware to manage checkpoints and the enlargement of other major processor structures, which, in turn, might impact the processor cycle. This paper focuses on out-of-order commit in a nonspeculative way, thus, avoiding checkpointing. To this end, we replace the ROB with a validation buffer (VB) structure. This structure keeps dispatched instructions until they are nonspeculative or mispeculated, which allows an early retirement. By doing so, the performance bottleneck is largely alleviated. An aggressive register reclamation mechanism targeted to this microarchitecture is also devised. As experimental results show, the VB structure is much more efficient than a typical ROB since, with only 32 entries, it achieves a performance close to an in-order commit microprocessor using a 256-entry ROB. Salvador Petit, Julio Sahuquillo, Pedro López 0001, Rafael Ubal, José Duato |
IEEE Trans. Computers | 3 |
| 2009 | FT2EI: A Dynamic Fault-Tolerant Routing Methodology for Fat Trees with Exclusion IntervalsabstractFault tolerance in the interconnection network of large clusters of PCs is an issue of growing importance, since their increasing size also increases the failure probability. The fat-tree topology is usually used in these machines since it has become very popular among high-speed interconnect manufacturers. This paper proposes a new distributed fault-tolerant routing methodology for fat trees. Unlike other previous proposals, it does not require additional network hardware, and its memory requirements, switch hardware, and routing delay scales up with the network size. Indeed, it nullifies only the strictly necessary paths, allowing adaptive routing through the healthy paths. The methodology is based on enhancing the interval routing scheme with exclusion intervals. Exclusion intervals are associated to each switch output port and represent the nodes that are unreachable from this port after a fault. We propose a methodology to identify the links where the exclusion intervals must be updated after a fault, the values to write on them, and a very efficient mechanism to distribute the required information through the network without stopping the system activity. Our methodology can tolerate a high number of network failures with a low degradation in performance. Moreover, it can achieve zero packet losing during the updating period. Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, J. F. D. Marin |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2009 | Region-Based Routing: A Mechanism to Support Efficient Routing Algorithms in NoCsabstractAn efficient routing algorithm is important for large on-chip networks [network-on-chip (NoC)] to provide the required communication performance to applications. Implementing NoC using table-based switches provide many advantages, including possibility of changing routing algorithms and fault tolerance, due to the option of table reconfigurations. However, table-based switches have been considered unsuitable for NoCs due to their perceived high area and power consumption. In this paper, we describe the region-based routing (RBR) mechanism which groups destinations into network regions allowing an efficient implementation with logic blocks. RBR can also be viewed as a mechanism to reduce the number of entries in routing tables. RBR is general and can be used in conjunction with any adaptive routing algorithm. In particular, we have evaluated the proposed scheme in conjunction with a general routing algorithm, namely segment-based routing (SR) and an application specific routing algorithm (APSRA) using regular and irregular mesh topologies. Our study shows that the number of entries in the table is significantly reduced, especially for large networks. Evaluation results show that RBR requires only four regions to support several routing algorithms in a 2-D mesh with no performance degradation. Considering link failures, our results indicate that RBR combined with SR is able to tolerate up to 7 link failures in an 8times8 mesh. RBR also reduces area and power dissipation of an equivalent table-based implementation by factors of 8 and 10, respectively. Moreover, the degradation in performance of the network is insignificant when using APSRA combined with RBR. Andres Mejia, Maurizio Palesi, José Flich, Shashi Kumar, Pedro López 0001, Rickard Holsmark, José Duato |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2008 | On the Influence of the Packet Marking and Injection Control Schemes in Congestion Management for MINs
Joan-Lluís Ferrer, Elvira Baydal, Antonio Robles, Pedro López 0001, José Duato |
Euro-Par | 4 |
| 2008 | Reducing Packet Dropping in a Bufferless NoC
Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, José Duato |
Euro-Par | 3 |
| 2008 | Reducing the Number of Bits in the BTB to Attack the Branch Predictor Hot-Spot
Noel Tomás, Julio Sahuquillo, Salvador Petit, Pedro López 0001 |
Euro-Par | 4 |
| 2008 | An Efficient Switching Technique for NoCs with Reduced Buffer RequirementsabstractNetworks on chip (NoCs) communicate the components located inside a chip. Overall system performance depends on NoC performance, that is affected by several factors. One of them is the network clock frequency, imposed by the critical path delay. Recent works show that switch critical path includes buffer control logic. Consequently, by removing switch buffers, switch frequency can be doubled. In this paper, we exploit this idea, proposing a new switching technique for NoCs which requires a reduced amount of storage at the switches. It is based on replacing switch port buffers by single latches. By doing so, network cycle can be reduced, which reduces packet latency. On the other hand, power and area consumption requirements can be reduced. However, since there are no buffers at the switch ports, packets can not be stopped. Stopped packets due to contention are dropped and reinjected from their senders via negative acknowledgments. Packet dropping is strongly reduced by exploiting NoCs wiring capability. Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, José Duato |
ICPADS | 3 |
| 2008 | RUFT: Simplifying the Fat-Tree TopologyabstractThe fat-tree is one of the most widely-used topologies by interconnection network manufacturers. Recently, a deterministic routing algorithm that optimally balances the network traffic in fat--trees was proposed. It can not only achieve almost the same performance than adaptive routing, but also outperforms it for some traffic patterns. Nevertheless, fat--trees require a high number of switches with a non-negligible wiring complexity. In this paper, we propose replacing the fat--tree by an unidirectional multistage interconnection network referred to as Reduced Unidirectional Fat--tree (RUFT) that uses a a simplified version of the aforementioned deterministic routing algorithm. As a consequence, switch hardware is almost reduced to the half, decreasing, in this way, power consumption, arbitration complexity, switch size, and network cost. Evaluation results show that RUFT obtains lower latency than fat--tree for low and medium traffic loads. Furthermore, in large networks, it obtains almost the same throughput than the classical fat-tree. Crispín Gómez Requena, Francisco Gilabert Villamón, María Engracia Gómez, Pedro López 0001, José Duato |
ICPADS | 4 |
| 2008 | The impact of out-of-order commit in coarse-grain, fine-grain and simultaneous multithreaded architecturesabstractMultithreaded processors in their different organizations (simultaneous, coarse grain and fine grain) have been shown as effective architectures to reduce the issue waste. On the other hand, retiring instructions from the pipeline in an out-of-order fashion helps to unclog the ROB when a long latency instruction reaches its head. This further contributes to maintain a higher utilization of the available issue bandwidth. In this paper, we evaluate the impact of retiring instructions out of order on different multithreaded architectures and different instruction fetch policies, using the recently proposed Validation Buffer microarchitecture as baseline out-of-order commit technique. Experimental results show that, for the same performance, out-of-order commit permits to reduce multithread hardware complexity (e.g., fine grain multithreading with a lower number of supported threads). Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
IPDPS | 4 |
| 2008 | Exploring High-Dimensional Topologies for NoC Design Through an Integrated Analysis and Synthesis Framework
Francisco Gilabert Villamón, Simone Medardoni, Davide Bertozzi, Luca Benini, María Engracia Gómez, Pedro López 0001, José Duato |
NOCS | 6 |
| 2008 | Exploiting Wiring Resources on Interconnection Network: Increasing Path DiversityabstractOn-chip networks are the answer to the growing demands for high communication performance of chip multiprocessors. These networks have a number of characteristics that make their design quite different to off-chip networks. In particular, wires are an abundant available resource inside the chip. In this paper, we explore how to organize the huge wiring capabilities available in on-chip networks. In particular, we analyze the option of distributing the wires among several parallel links connecting the same two switches. This technique is known as Space Division Multiplexing (SDM). The number of parallel sub-links and their width are two key parameters that are studied together with the relationship with the mean packet size. The paper shows that SDM is a technique to take into account in on-chip networks since it allows to highly increase the network accepted traffic at the expense of a small latency increase or even no increase. Moreover, in some networks, it allows to reduce the network hardware, providing simiar performance results, which results in a reduction in the consumption of area and power. Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, José Duato |
PDP | 3 |
| 2007 | VB-MT: Design Issues and Performance of the Validation Buffer Microarchitecture for Multithreaded Processors
Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001, José Duato |
PACT | 4 |
| 2007 | Power-Aware Fat-Tree Networks Using On/Off Links
Marina Alonso, Salvador Coll, Vicente Santonja, Juan-Miguel Martinez-Rubio, Pedro López 0001, José Duato |
HPCC | 5 |
| 2007 | Deterministic versus Adaptive Routing in Fat-TreesabstractClusters of PCs have become very popular to build high performance computers. These machines use commodity PCs linked by a high speed interconnect. Routing is one of the most important design issues of interconnection networks. Adaptive routing usually better balances network traffic, thus allowing the network to obtain a higher throughput. However, adaptive routing introduces out-of-order packet delivery, which is unacceptable for some applications. Concerning topology, most of the commercially available interconnects are based on fat-tree. Fat-trees offer a rich connectivity among nodes, making possible to obtain paths between all source-destination pairs that do not share any link. We exploit this idea to propose a deterministic routing algorithm for fat-trees, comparing it with adaptive routing in several workloads. The results show that deterministic routing can achieve a similar, and in some scenarios higher, level of performance than adaptive routing, while providing in-order packet delivery. Crispín Gómez Requena, Francisco Gilabert Villamón, María Engracia Gómez, Pedro López 0001, José Duato |
IPDPS | 4 |
| 2007 | An Efficient Fault-Tolerant Routing Methodology for Fat-Tree Interconnection Networks
Crispín Gómez Requena, María Engracia Gómez, Pedro López 0001, José Duato |
ISPA | 3 |
| 2007 | Region-Based Routing: An Efficient Routing Mechanism to Tackle Unreliable Hardware in Network on ChipsabstractThe design of scalable and reliable interconnection networks for system on chips (SoCs) introduce new design constraints not present in current multicomputer systems. Although regular topologies are preferred for building NoCs, heterogeneous blocks, fabrication faults and reliability issues derived from the high integration scale may lead to irregular topologies. In this situation, efficient routing becomes a challenge. Although table-based routing allows the use of most routing algorithms on any topology, it does not scale in terms of latency and area. In this paper we propose the region-based routing mechanism that avoids the scalability problems of table-based solutions. From an initial topology and routing algorithm, the mechanism groups, at every switch, destinations into different regions based on the output ports. By doing this, redundant routing information typically found in routing tables is eliminated. Evaluation results show that the mechanism requires only four regions to support several routing algorithms in a 2D mesh with no performance degradation. Moreover, when dealing with link failures, our results indicate that the mechanism combined with the segment-based routing algorithm is able to pack all the routing information into eight regions providing high throughput. The paper provides also a simple and efficient hardware implementation of the mechanism requiring only 240 logic gates per switch to support eight regions in a 2D mesh topology José Flich, Andres Mejia, Pedro López 0001, José Duato |
NOCS | 3 |
| 2007 | Congestion Management in MINs through Marked and Validated PacketsabstractCongestion management is a very critical problem tackled in interconnection networks for years but not solved yet. Although several mechanisms have been recently proposed for lossless multistage interconnection networks (MINs), they either have drawbacks or are partial solutions. Some of them introduce penalty over packets not really addressed to the hot-spots, whereas others can cope only with congestion situations that last a short time. In this paper, we propose an effective and efficient congestion management mechanism for lossless interconnection networks based on explicit congestion notification. The mechanism uses two different flags in ACK packets, a Marking Bit (MB) and a Validation Bit (VB), to detect congestion and warn the origin hosts. In this way, packets belonging to "coldflows" but stopped because of head-of-line (HOL) blocking can be distinguished from "hotflow" packets which are really causing congestion. In response, origin hosts can apply corrective actions only to the "hotflows", minimizing the negative impact on "coldflows"performance. Evaluation results show that the proposed congestion management strategy is able to avoid the degradation of network performance, regardless of traffic load and the location of the congestion in the network. Joan-Lluís Ferrer, Elvira Baydal, Antonio Robles, Pedro López 0001, José Duato |
PDP | 4 |
| 2007 | Multi2Sim: A Simulation Framework to Evaluate Multicore-Multithreaded ProcessorsabstractCurrent microprocessors are based in complex designs, integrating different components on a single chip, such as hardware threads, processor cores, memory hierarchy or interconnection networks. The permanent need of evaluating new designs on each of these components motivates the development of tools which simulate the system working as a whole. In this paper, we present the Multi2Sim simulation framework, which models the major components of incoming systems, and is intended to cover the limitations of existing simulators. A set of simulation examples is also included for illustrative purposes. Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001 |
SBAC-PAD | 4 |
| 2006 | Towards an efficient switch architecture for high-radix switchesabstractThe interconnection network plays a key role in the overall performance achieved by high performance computing systems, also contributing an increasing fraction of its cost and power consumption. Current trends in interconnection network technology suggest that high-radix switches will be preferred as networks will become smaller (in terms of switch count) with the associated savings in packet latency, cost, and power consumption. Unfortunately, current switch architectures have scalability problems that prevent them from being effective when implemented with a high number of ports. In this paper, an efficient and cost-effective architecture for high-radix switches is proposed. The architecture, referred to as Partitioned Crossbar Input Queued (PCIQ), relies on three key components: a partitioned crossbar organization that allows the use of simple arbiters and crossbars, a packet-based arbiter, and a mechanism to eliminate the switch-level HOL blocking. Under uniform traffic, maximum switch efficiency is achieved. Furthermore, switch-level HOL blocking is completely eliminated under hot-spot traffic, again delivering maximum throughput. Additionally, PCIQ inherently implements an efficient congestion management technique that eliminates all the network-wide HOL blocking. On the contrary, the previously proposed architectures either show poor performance or they require significantly higher costs than PCIQ (in both components and complexity). Gaspar Mora, José Flich, José Duato, Pedro López 0001, Elvira Baydal, Olav Lysne |
ANCS | 4 |
| 2006 | On the Influence of the Selection Function on the Performance of Fat-Trees
Francisco Gilabert Villamón, María Engracia Gómez, Pedro López 0001, José Duato |
Euro-Par | 3 |
| 2006 | Dynamic power saving in fat-tree interconnection networks using on/off linksabstractCurrent trends in high-performance parallel computers show that fat-tree interconnection networks are one of the most popular topologies. The particular characteristics of this topology, that provide multiple alternative paths for each source/destination pair, make it an excellent candidate for applying power consumption reduction techniques. Such techniques are being increasingly applied in computer systems and the interconnection network is not an exception, since its contribution to the system power budget is not negligible. In this paper, we present a mechanism that dynamically switches on and off network links as a function of traffic. The mechanism is designed to guarantee network connectivity, according to the underlying routing algorithm. In this way, the default routing algorithm can be used regardless of the power saving actions taken, thus simplifying router design. Our simulation results show that significant network power consumption reductions can be obtained at no cost. Latency remains the same although the number of operating network links is dynamically adjusted. Marina Alonso, Salvador Coll, Juan-Miguel Martinez-Rubio, Vicente Santonja, Pedro López 0001, José Duato |
IPDPS | 5 |
| 2006 | Applying the zeros switch-off technique to reduce static energy in data cachesabstractZeros switch-off is a leakage energy reduction technique applicable to cache memories. It works at the cache word level by removing the power supply of all or part of its most significant bytes when they store a zero, taking advantage of the high percentage of zero data bits in common programs. Experimental results, obtained by using the SPEC2000 benchmarks suite, show that the average leakage energy savings reach 60.3% with no IPC loss indeed. The proposed technique can be combined with other existing energy reduction techniques, reaching, on average, 67.3% savings with 0.5% IPC losses Rafael Ubal, Julio Sahuquillo, Salvador Petit, Pedro López 0001 |
SBAC-PAD | 4 |
| 2006 | FIR: An efficient routing strategy for tori and meshes
María Engracia Gómez, Pedro López 0001, José Duato |
J. Parallel Distributed Comput. | 2 |
| 2006 | A Routing Methodology for Achieving Fault Tolerance in Direct NetworksabstractMassively parallel computing systems are being built with thousands of nodes. The interconnection network plays a key role for the performance of such systems. However, the high number of components significantly increases the probability of failure. Additionally, failures in the interconnection network may isolate a large fraction of the machine. It is therefore critical to provide an efficient fault-tolerant mechanism to keep the system running, even in the presence of faults. This paper presents a new fault-tolerant routing methodology that does not degrade performance in the absence of faults and tolerates a reasonably large number of faults without disabling any healthy node. In order to avoid faults, for some source-destination pairs, packets are first sent to an intermediate node and then from this node to the destination node. Fully adaptive routing is used along both subpaths. The methodology assumes a static fault model and the use of a checkpoint/restart mechanism. However, there are scenarios where the faults cannot be avoided solely by using an intermediate node. Thus, we also provide some extensions to the methodology. Specifically, we propose disabling adaptive routing and/or using misrouting on a per-packet basis. We also propose the use of more than one intermediate node for some paths. The proposed fault-tolerant routing methodology is extensively evaluated in terms of fault tolerance, complexity, and performance. María Engracia Gómez, Nils Agne Nordbotten, José Flich, Pedro López 0001, Antonio Robles, José Duato, Tor Skeie, Olav Lysne |
IEEE Trans. Computers | 4 |
| 2005 | A Memory-Effective Fault-Tolerant Routing Strategy for Direct Interconnection NetworksabstractHigh-performance interconnection networks are crucial in massively parallel computers. Routing is one of the most important design issues of interconnection networks. Moreover, the huge amount of hardware of these machines makes fault-tolerance another important design issue. In this paper, we propose a mechanism that combines scalable routing and fault-tolerance for commercial switches to build direct regular topologies, which are the topologies used in large machines. The hardware required is not complex. Furthermore, it allows a high degree of fault-tolerance inflicting a minimal decrease of performance María Engracia Gómez, Pedro López 0001, José Duato |
ISPDC | 2 |
| 2005 | Enforcing in-order packet delivery in system area networks with adaptive routing
Michihiro Koibuchi, José Flich, Antonio Robles, Pedro López 0001, José Duato |
J. Parallel Distributed Comput. | 5 |
| 2005 | A Family of Mechanisms for Congestion Control in Wormhole NetworksabstractMultiprocessor interconnection networks may reach congestion with high traffic loads, which prevents reaching the wished performance. Unfortunately, many of the mechanisms proposed in the literature for congestion control either suffer from a lack of robustness, being unable to work properly with different traffic patterns or message lengths, or detect congestion relying on global information that wastes some network bandwidth. This paper presents a family of mechanisms to avoid network congestion in wormhole networks. All of them need only local information, applying message throttling when it is required. The proposed mechanisms use different strategies to detect network congestion and also apply different corrective actions. The mechanisms are evaluated and compared for several network loads and topologies, noticeably improving network performance with high loads but without penalizing network behavior for low and medium traffic rates, where no congestion control is required. Elvira Baydal, Pedro López 0001, José Duato |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2004 | Reducing Power Consumption in Interconnection Networks by Dynamically Adjusting Link Width
Marina Alonso, Juan-Miguel Martinez-Rubio, Vicente Santonja, Pedro López 0001 |
Euro-Par | 4 |
| 2004 | A New Adaptive Fault-Tolerant Routing Methodology for Direct Networks
María Engracia Gómez, José Duato, José Flich, Pedro López 0001, Antonio Robles, Nils Agne Nordbotten, Tor Skeie, Olav Lysne |
HiPC | 4 |
| 2004 | LASH-TOR: A Generic Transition-Oriented Routing Algorithm
Tor Skeie, Olav Lysne, José Flich, Pedro López 0001, Antonio Robles, José Duato |
ICPADS | 4 |
| 2004 | An Effective Fault-Tolerant Routing Methodology for Direct NetworksabstractCurrent massively parallel computing systems are being built with thousands of nodes, which significantly affect the probability of failure. M. E. Gomex proposed a methodology to design fault-tolerant routing algorithms for direct interconnection networks. The methodology uses a simple mechanism: for some source-destination pairs, packets are first forwarded to an intermediate node, and later, from this node to the destination node. Minimal adaptive routing is used along both subpaths. For those cases where the methodology cannot find a suitable intermediate node, it combines the use of intermediate nodes with two additional mechanisms: disabling adaptive routing and using misrouting on a per-packet basis. While the combination of these three mechanisms tolerates a large number of faults, each one requires adding some hardware support in the network and also introduces some overhead. In this paper, we perform an in-depth detailed analysis of the impact of these mechanisms on network behaviour. We analyze the impact of the three mechanisms separately and combined. The ultimate goal of this paper is to obtain a suitable combination of mechanisms that is able to meet the trade-off between fault-tolerance degree, routing complexity, and performance. María Engracia Gómez, José Flich, Pedro López 0001, Antonio Robles, José Duato, Nils Agne Nordbotten, Olav Lysne, Tor Skeie |
ICPP | 3 |
| 2004 | A Transition-Based Fault-Tolerant Routing Methodology for InfiniBand NetworksabstractSummary form only given. Currently, clusters of PCs are considered a cost-effective alternative to large parallel computers. As the number of elements increases in these systems, the probability of faults increases dramatically. Therefore, it is critical to keep the system running even in the presence of faults. The interconnection network plays a key role in its performance. InfiniBand (IBA) is a new standard interconnect suitable for clusters. Most of the fault-tolerant routing strategies proposed for massively parallel computers cannot be applied to IBA because routing and virtual channel transitions are deterministic, which prevents packets from avoiding the faults. A possible approach to provide fault-tolerance in IBA consists of using several disjoint paths between every source-destination pair of nodes and selecting the appropriate path at the source host. However, to this end, a routing algorithm able to provide enough disjoint paths, while still guaranteeing deadlock freedom, is required. We propose a simple and effective fault-tolerant methodology for IBA networks that can be applied to any network topology and meets the trade-off between fault-tolerance degree and the number of network resources devoted to it. Preliminary results show that the proposed methodology scales well and supports up to three faults in 2D and five in 3D tori using only two virtual channels. José Miguel Montañana, José Flich, Antonio Robles, Pedro López 0001, José Duato |
IPDPS | 4 |
| 2004 | A Fully Adaptive Fault-Tolerant Routing Methodology Based on Intermediate Nodes
Nils Agne Nordbotten, María Engracia Gómez, José Flich, Pedro López 0001, Antonio Robles, Tor Skeie, Olav Lysne, José Duato |
NPC | 4 |
| 2003 | A Robust Mecahnism for Congestion Control: INC
Elvira Baydal, Pedro López 0001 |
Euro-Par | 2 |
| 2003 | Low-Fragmentation Mapping Strategies for Linear Forwarding Tables in InfiniBandTM
Pedro López 0001, José Flich, Antonio Robles |
Euro-Par | 1 |
| 2003 | Routing in InfiniBandTM Torus Network TopologieabstractInfiniBand is an interconnect standard for communication between processing nodes and I/O devices as well as for interprocessor communication (NOWs). The InfiniBand architecture (IBA) defines a switch-based network with point-to-point links whose topology can be established by the customer. When the performance is the primary concern regular topologies are preferred. Low-dimensional tori (2D and 3D) are some of the regular topologies most widely used in commercial parallel computers. Routing in torus requires the use of virtual channels. Although InfiniBand provides support for deterministic routing and virtual channels, they are selected at each switch by service level (SL) identifiers associated to packets and do not depend on packet destination. This makes routing algorithm implementation more complex. In particular, a large number of SLs may be required, which is a scarce resource. We analyze the way several routing strategies can be applied in tori InfiniBand networks, also evaluating their resource requirements. In particular, we analyze and compare the well-known e-cube and up*/down* routing algorithms and the flexible routing algorithm recently proposed José Carlos Sancho, Antonio Robles, Pedro López 0001, José Flich, José Duato |
ICPP | 3 |
| 2003 | Supporting adaptive routing in IBA switches
José Flich, Antonio Robles, Pedro López 0001, José Duato |
J. Syst. Archit. | 4 |
| 2003 | Applying In-Transit Buffers to Boost the Performance of Networks with Source RoutingabstractIn this paper, we analyze in depth the effect of using ITB in the network, showing that they not only serve for guaranteeing minimal routing, but also that they are a powerful mechanism able to balance network traffic and reduce network contention. To demonstrate these capabilities, we apply the ITB mechanism to improved routing schemes, such as DFS and smart-routing. These routing algorithms (without ITB) are able to improve the performance of up*/down* by 30 percent and 90 percent, respectively, for a 32-switch network. The evaluation results show that, when ITB are used together with these improved routing algorithms, network throughput achieved by DFS and smart-routing can still be improved by 56 percent and 23 percent, respectively. However, smart-routing requires a time to compute the routing tables that rapidly grows with network size, it being impossible in practice to build networks with more than 32 switches. This high computational cost is mainly motivated by the need of obtaining deadlock-free routing tables. However, when ITB are used, one can decouple the stages of computing routing tables and breaking cycles. Moreover, as stated above, ITB can be used to reduce network contention. In this way, in this paper, we also propose a completely new routing algorithm that tries to balance network traffic by using a simple and low time consuming strategy. The proposed algorithm guarantees deadlock freedom and reduces network contention with the use of ITB. The evaluation results show that our algorithm obtains unprecedented throughputs in 32-switch networks, tripling the original up*/down* and almost doubling smart-routing. José Flich, Pedro López 0001, Manuel P. Malumbres, José Duato, Tomas Rokicki |
IEEE Trans. Computers | 2 |
| 2003 | FC3D: Flow Control-Based Distributed Deadlock Detection Mechanism for True Fully Adaptive Routing in Wormhole NetworksabstractTwo general approaches have been proposed for deadlock handling in wormhole networks. Traditionally, deadlock-avoidance strategies have been used. In this case, either routing is restricted so that there are no cyclic dependencies between channels or cyclic dependencies between channels are allowed provided that there are some escape paths to avoid deadlock. More recently, deadlock recovery strategies have begun to gain acceptance. These strategies allow the use of unrestricted fully adaptive routing, usually outperforming deadlock avoidance techniques. However, they require a deadlock detection mechanism and a deadlock recovery mechanism that is able to recover from deadlocks faster than they occur. In particular, progressive deadlock recovery techniques are very attractive because they allocate a few dedicated resources to quickly deliver deadlocked messages, instead of killing them. Unfortunately, distributed deadlock detection is usually based on crude time-outs, which detect many false deadlocks. As a consequence, messages detected as deadlocked may saturate the bandwidth offered by recovery resources, thus degrading performance. Additionally, the threshold required by the detection mechanism (the time-out) strongly depends on network load, which is not known in advance at the design stage. This limits the applicability of deadlock recovery on actual networks. We propose a novel distributed deadlock detection mechanism that uses only local information, detects all the deadlocks, considerably reduces the probability of false deadlock detection over previously proposed techniques, and is not significantly affected by variations in message length and/or message destination distribution. Juan-Miguel Martinez-Rubio, Pedro López 0001, José Duato |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2002 | Congestion Control Based on Transmission Times
Elvira Baydal, Pedro López 0001, José Duato |
Euro-Par | 2 |
| 2002 | Evaluation of Routing Algorithms for InfiniBand Networks (Research Note)
María Engracia Gómez, José Flich, Antonio Robles, Pedro López 0001, José Duato |
Euro-Par | 4 |
| 2002 | Effective Methodology for Deadlock-Free Minimal Routing in InfiniBand NetworksabstractThe InfiniBand Architecture (IBA) defines a switch-based network with point-to-point links whose topology is arbitrarily established by the customer. We propose a simple and effective methodology for designing deadlock-free routing strategies that are able to route packets through minimal paths in InfiniBand networks. This methodology can meet the trade-off between network performance and the number of resources dedicated to deadlock avoidance. Evaluation results show that the resulting routing strategies significantly outperform up*/down* routing. In particular, throughput improvement ranges, on average, from 1.33 for small networks to 4.05 for large networks. Also, it is shown that just two virtual lanes and three service levels are enough to achieve more than 80% of the throughput improvement achieved by the best proposed routing strategy (the one that always provides minimal paths without limiting the number of resources). José Carlos Sancho, Antonio Robles, José Flich, Pedro López 0001, José Duato |
ICPP | 4 |
| 2002 | Boosting the Performance of Myrinet NetworksabstractNetworks of workstations (NOWs) are becoming increasingly popular as a cost-effective alternative to parallel computers. These networks allow the customer to connect processors using irregular topologies, providing the wiring flexibility, scalability and incremental expansion capability required in this environment. Some of these networks use source routing and wormhole switching. In particular, we are interested in Myrinet networks because they are a well-known commercial product and their behavior can be controlled by the software running on the network interfaces (the Myrinet Control Program, MCP). Usually, the Myrinet network uses up*/down* routing for computing the paths for every source-destination pair. In this paper, we propose an in-transit buffer (ITB) mechanism to improve the network performance. We apply the ITB mechanism to NOWs with up*/down* source routing, like the Myrinet, analyzing its behavior on networks with both regular and irregular topologies. The proposed scheme can be implemented on Myrinet networks by simply modifying the MCP, without changing the network hardware. We evaluate by simulation several networks with different traffic patterns using timing parameters taken from the Myrinet network. The results show that the current routing schemes used in Myrinet networks can be strongly improved by applying the ITB mechanism. In general, our proposed scheme is able to double the network throughput on medium and large NOWs. Finally, we present a first implementation of the ITB mechanism on a Myrinet network. José Flich, Pedro López 0001, Manuel P. Malumbres, José Duato |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2002 | Boosting the Performance of Myrinet NetworksabstractNetworks of workstations (NOWs) are becoming increasingly popular as a cost-effective alternative to parallel computers. These networks allow the customer to connect processors using irregular topologies, providing the wiring flexibility, scalability, and incremental expansion capability required in this environment. Some of these networks use source routing and wormhole switching. In particular, we are interested in Myrinet networks because it is a well-known commercial product and its behavior can be controlled by the software running in network interfaces (Myrinet Control Program, MCP). Usually, the Myrinet network uses up*/down* routing for computing the paths for every source-destination pair. We propose the In-Transit Buffer (ITB) mechanism to improve network performance. We apply the ITB mechanism to NOWs with up*/down* source routing, like Myrinet, analyzing its behavior on both networks with regular and irregular topologies. The proposed scheme can be implemented on Myrinet networks by only modifying the MCP, without changing the network hardware. We evaluate by simulation several networks with different traffic patterns using timing parameters taken from the Myrinet network. Results show that the current routing schemes used in Myrinet networks can be strongly improved by applying the ITB mechanism. In general, our proposed scheme is able to double the network throughput on medium and large NOWs. Finally, we present a first implementation of the ITB mechanism on a Myrinet network. José Flich, Pedro López 0001, Manuel P. Malumbres, José Duato |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2001 | Deadlock-Free Routing in InfiniBand through Destination RenamingabstractThe InfiniBand Architecture (IBA) defines a switch-based network with point-to-point links that supports any topology defined by the user including irregular ones, in order to provide flexibility and incremental expansion capability. Routing in IBA is distributed, based on forwarding tables, and only considers the packet destination ID for routing within subnets in order to drastically reduce forwarding table size. Unfortunately, the forwarding tables for most of the previously proposed routing algorithms for irregular topologies consider both the destination ID and the input channel. Therefore, these popular routing algorithms for irregular topologies may not be usable in InfiniBand networks because they do nor conform to the IBA specifications. In this paper we propose an easy-to-implement strategy to adapt the forwarding tables already computed following any routing algorithm that considers the destination ID and the input channel into the required IBA forwarding table format. The resulting routing algorithm is deadlock-free on IBA. Indeed, the originally computed paths are not modified at all. Hence, the proposed strategy does not degrade performance with respect to the original routing scheme. Pedro López 0001, José Flich, José Duato |
ICPP | 1 |
| 2001 | A First Implementation of In-Transit Buffers on Myrinet GM SoftwareabstractClusters of workstations (COWs) are becoming increasingly popular as a cost-effective alternative to parallel computers. In these systems, the interconnection network connects hosts using irregular topologies, providing the wiring flexibility, scalability, and incremental expansion capability required in this environment. Myrinet is the most popular network used to build COWs. It uses source routing with the up*/down * routing algorithm. In previous papers we proposed the In-Transit Buffer (ITB) mechanism that improves network performance by allowing minimal routing, balancing network traffic, and reducing network contention. The mechanism is based on ejecting packets at some intermediate hosts and later re-injecting them into the network. Moreover, the ITB mechanism does not require additional hardware as it can be implemented on the software running at Myrinet network adapters. In this paper, we present a first implementation of the ITB mechanism on Myrinet GM software. We show the changes required in packet format and the modifications performed in the Myrinet Control Program (MCP). In addition, both the overhead introduced by the new code and the cost of extracting and re-injecting packets are measured. Results show that, even for this simple implementation, code overhead is only about 125 ns per packet and the message latency increase for messages that use the ITB mechanism is around 1.3 s per ITB. This is the first attempt to implement this mechanism, showing that a real implementation of ITBs is feasible on Myrinet COWs, and the associated overhead does not restrict the potential benefits of this mechanism. 1. Salvador Coll, José Flich, Manuel P. Malumbres, Pedro López 0001, José Duato, Francisco J. Mora |
IPDPS | 4 |
| 2001 | Improving Network Performance by Reducing Network Contention in Source-Based COWs with a Low Path-Computation OverheadabstractIn previous papers, we have proposed the in-transit buffer mechanism (ITB) to improve network performance in COWs with irregular topology and source routing. This mechanism allows the use of minimal paths among all hosts, breaking cyclic dependences between channels by storing and later re-injecting packets at some intermediate hosts. However it also has two additional features that can improve even more network performance. First, the ITB mechanism reduces network contention because some messages are ejected from the network freeing network links. Second the ITB mechanism allows the use of any path between each source-destination pair improving traffic balance. In this paper we present a new routing algorithm that takes advantage of ITB by exploiting both issues: traffic balance and network contention reduction. The evaluation results show that network throughput can be considerably improved. On average, network throughput increases with respect to up*/down* by factors of 2.51 and 3.77 in 32 and 64-switch networks, respectively. José Flich, Pedro López 0001, Manuel P. Malumbres, José Duato, Tomas Rokicki |
IPDPS | 2 |
| 2001 | A Cost-Effective Approach to Deadlock Handling in Wormhole NetworksabstractWormhole networks have traditionally used deadlock avoidance strategies. More recently, deadlock recovery strategies have begun to gain acceptance. In particular, progressive deadlock recovery techniques allocate a few dedicated resources to quickly deliver deadlocked packets. Deadlock recovery is based on the assumption that deadlocks are rare; otherwise, recovery techniques are not efficient. Measurements of deadlock occurrence frequency show that deadlocks are highly unlikely when enough routing freedom is provided. However, networks are more prone to deadlocks when the network is close to or beyond saturation, causing some network performance degradation. Similar performance degradation behavior at saturation was also observed in networks using deadlock avoidance strategies. In this paper, we take a different approach to handling deadlocks and performance degradation. We propose the use of an injection limitation mechanism that prevents performance degradation near the saturation point and, at the same time, reduces the probability of deadlock to negligible values. We also propose an improved deadlock detection mechanism that uses only local information, detects all deadlocks, and considerably reduces the probability of false deadlock detection over previous proposals. In the rare case when impending deadlock is detected, our proposal consists of using a simple recovery technique that absorbs the deadlocked message at the current node and later reinjects it for continued routing toward its destination. Performance evaluation results show that our new approach to handling deadlock is more efficient than previously proposed techniques. Juan-Miguel Martinez-Rubio, Pedro López 0001, José Duato |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2000 | Improving the Performance of Regular Networks with Source RoutingabstractNetworks of workstations (NOWs) are becoming increasingly popular as a cost-effective alternative to parallel computers. In these machines, the network connects processors using irregular topologies, providing the wiring flexibility, scalability, and incremental expansion capability required in this environment. Also, when performance is the primary concern, these network products are being used to build large commodity clusters with regular topologies. In previous papers, we have proposed the in-transit buffer mechanism to improve network performance, applying it to NOWs with irregular topology and source routing. This mechanism allows the use of minimal paths among all hosts, breaking cyclic dependencies between channels by storing and later re-injecting packers at some intermediate hosts. In this paper we apply the in-transit buffer mechanism to regular networks with source routing in order to improve their performance. Also, two path selection policies are evaluated. The first one will always choose the same minimal path from source to destination, whereas the second one will choose from different alternative minimal paths in a round-robin fashion. The evaluation results show that the overall network throughput can be doubled for large networks. José Flich, Pedro López 0001, Manuel P. Malumbres, José Duato |
ICPP | 2 |
| 2000 | Performance evaluation of a new routing strategy for irregular networks with source routingabstractNetworks of workstations (NOWs) are becoming increasingly popular as a cost-effective alternative to parallel computers. Typically, these networks connect processors using irregular topologies, providing the wiring flexibility, scalability, and incremental expansion capability required in this environment. In some of these networks, messages are delivered using the up*/down* routing algorithm [9]. However, the up*/down* routing scheme is often non-minimal. Also, some of these networks use source routing [1]. With this technique, the entire path to destination is generated at the source host before the message is sent. José Flich, Manuel P. Malumbres, Pedro López 0001, José Duato |
ICS | 3 |
| 2000 | A Simple and Efficient Mechanism to Prevent Saturation in Wormhole NetworksabstractBoth deadlock avoidance and recovery techniques suffer from severe performance degradation when the network is close to or beyond saturation. This performance degradation appears because messages block in the network faster than they are drained by the escape paths in the deadlock avoidance strategies or the deadlock recovery mechanism. Many parallel applications produce bursty traffic that may saturate the network during some intervals, significantly increasing execution time. Therefore, the use of techniques that prevent network saturation are of crucial importance. Although several mechanisms have been proposed in the literature to reach this goal, some of them introduce some penalty when the network is not fully saturated, require complex hardware to be implemented or do not behave well under all network load conditions. In this paper we propose a new mechanism to avoid network saturation that overcomes these drawbacks. Elvira Baydal, Pedro López 0001, José Duato |
IPDPS | 2 |
| 2000 | Improving Routing Performance in Myrinet NetworksabstractNetworks of workstations (NOWs) are becoming increasingly popular as a cost-effective alternative to parallel computers. Typically, these networks connect processors using irregular topologies, providing the wiring flexibility, scalability, and incremental expansion capability required in this environment. In some of these networks, packets are delivered using source routing. Due to the irregular topology, the routing scheme is often non-minimal. In this paper we analyze the routing scheme used in Myrinet networks in order to improve its performance. We propose new routing algorithms that balance the utilization of the available routes and always use minimal paths. We show through simulation that the current routing schemes used in Myrinet networks can be improved by modifying only the routing software without increasing the software overhead significantly. The overall throughput can be doubled without modifying the network hardware. José Flich, Manuel P. Malumbres, Pedro López 0001, José Duato |
IPDPS | 3 |
| 1999 | Impact of Buffer Size on the Efficiency of Deadlock DetectionabstractDeadlock detection is one of the most important design issues in recovery strategies for routing in interconnection networks. In a previous paper, we presented an efficient deadlock detection mechanism. This mechanism requires that when a message header blocks it must be quickly notified to all the channels reserved by that message. To achieve this goal, the detection mechanism uses the information provided by flow control. Some recent commercial multiprocessors use deep buffers, since they may increase network throughput and efficiently allow transmission over long wires. However, deep buffers may increase the elapsed time between header blocking at a router and the propagation of flow control signals, thus negatively affecting the behavior of our deadlock detection mechanism. On the other hand, deeper buffers reduce deadlock frequency. As a consequence, buffer size has opposing effects on deadlock detection. In this paper, we analyze by simulation the influence of these effects on the efficiency of our deadlock detection mechanism, showing that overall performance improves with buffer size. Juan-Miguel Martinez-Rubio, Pedro López 0001, José Duato |
HPCA | 2 |
| 1999 | Performance Evaluation of Networks of Workstations with Hardware Shared Memory Model Using Execution-Driven SimulationabstractNetworks of workstations (NOWs) are becoming increasingly popular as a cost-effective alternative to parallel computers. Typically, these networks connect processors using irregular topologies, providing the wiring flexibility, scalability, and incremental expansion capability required in this environment. Similar to the evolution of parallel computers, NOWs are also evolving from distributed memory to shared memory programming model. However, physical distances between processors are longer in NOWs than in tightly-coupled distributed shared-memory multiprocessors (DSMs), leading to higher message latency and lower network bandwidth. Therefore, the network may be a bottleneck when executing some parallel applications in a NOW supporting a shared-memory programming paradigm. In this paper we analyze whether the interconnection network is able to efficiently handle the traffic generated in a NOW with the shared memory model. In particular, we are interested in analyzing the influence of the routing mechanism in the performance of the system. We evaluate the behavior of a NOW with irregular topology by means of an execution-driven simulator using SPLASH-2 applications as the input load. The results show that the routing algorithm can considerably reduce the total execution time of applications. In particular routing adaptivity can reduce the total execution time by 58% in some applications. These results confirm the behavior observed in previous works using synthetic traffic loads. José Flich, Manuel P. Malumbres, Pedro López 0001, José Duato |
ICPP | 3 |
| 1998 | A Very Efficient Distributed Deadlock Detection Mechanism for Wormhole NetworksabstractNetworks using wormhole switching have traditionally relied upon deadlock avoidance strategies for the design of routing algorithms. More recently, deadlock recovery strategies have begun to gain acceptance. Progressive deadlock recovery techniques are very attractive because they allocate a few dedicated resources to quickly deliver deadlocked messages, instead of killing them. However, the distributed deadlock detection techniques proposed up to now detect many false deadlocks, especially when the network is heavily loaded and messages have different lengths. As a consequence, messages detected as deadlocked may saturate the bandwidth offered by recovery resources, thus degrading performance considerably. In this paper we propose an improved distributed deadlock detection mechanism that uses only local information, detects all the deadlocks, considerably reduces the probability of false deadlock detection and is not strongly affected by variations in message length and message destination distribution. Pedro López 0001, Juan-Miguel Martinez-Rubio, José Duato |
HPCA | 1 |
| 1998 | DRIL: Dynamically Reduced Message Injection Limitation Mechanism for Wormhole NetworksabstractDeadlock avoidance and recovery techniques are alternatives to deal with the interconnection network deadlock problem. Both techniques allow fully adaptive routing on some set of resources while providing dedicated resources to escape from deadlock. They mainly differ in the way they supply escape paths and when those paths are used. As the escape paths only provide limited bandwidth to escape from deadlocks, both techniques suffer from severe performance degradation when the network is close to saturation. On the other hand, deadlock recovery is based on the assumption that deadlocks are rare. Several studies show that deadlock are more prone when the network is close to or beyond saturation. In this paper we propose a new mechanism that prevents network saturation by dynamically adjusting message injection limitation into the network. As a consequence, this mechanism will avoid the performance degradation problem that typically occurs in both deadlock avoidance and recovery techniques, making fully adaptive feasible. Also, it will guarantee that the frequency of deadlock is really negligible, allowing the use of simple low-cost recovery strategies. Pedro López 0001, Juan-Miguel Martinez-Rubio, José Duato |
ICPP | 1 |
| 1998 | A cost-effective methodology for the evaluation of interconnection networks
Pedro López 0001, Rosa Alcover, José Duato, Luisa Zúnica |
J. Syst. Archit. | 1 |
| 1997 | LIFE: a limited injection, fully adaptive, recovery-based routing algorithmabstractNetworks using wormhole switching have traditionally relied upon deadlock avoidance strategies for the design of deadlock-free algorithms. The past few years have seen a rise in popularity of deadlock recovery strategies, that are based on the property that deadlocks are quite rare in practice and happen only at or beyond the network saturation point. In fact, recovery-based routing algorithms have a higher potential performance over the deadlock avoidance-based ones which allow less routing freedom. We present a recovery-based fully adaptive routing algorithm, LIFE, which is based on an innovative injection policy that reduces the probability of deadlocks to negligible values, both with uniform and non-uniform traffic patterns. The experimental results, conducted on an 8-ary 3-cube with 512 nodes, show that it is possible to implement true fully adaptive routing using only two virtual channels. Also, LIFE outperforms state-of-the-art avoidance- and recovery-based algorithms of the same cost both in terms of throughput and message latency under uniform traffic and provides stable throughput under non-uniform traffic patterns. Fabrizio Petrini, José Duato, Pedro López 0001, Juan-Miguel Martinez-Rubio |
HiPC | 3 |
| 1997 | Software-Based Deadlock Recovery Technique for True Fully Adaptive Routing in Wormhole NetworksabstractIn this paper, we take a different approach to handle deadlocks and performance degradation. We propose the use of an injection limitation mechanism that prevents performance degradation near the saturation point and reduces the probability of deadlock to negligible values even when fully adaptive routing is used. We also propose an improved deadlock detection mechanism that only uses local information, detects all the deadlocks, and considerably reduces the probability of false deadlock detection over previous proposals. In the rare case when impending deadlock is detected, our proposed recovery technique absorbs the deadlocked message at the current node and later re-injects it for continued routing towards its destination. Performance evaluation results show that our new approach to deadlock handling is more efficient than previously proposed techniques. Juan-Miguel Martinez-Rubio, Pedro López 0001, José Duato, Timothy M. Pinkston |
ICPP | 2 |