EDBT 2026 Demo / reviewers in the wild / expert
Javier Navaridas
dblp:66/5660
· DBLP profile ↗
49ranked-venue papers
13as first author
10since 2021 · last 2026
0000-0001-7272-6597ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 10 first-author · 4 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 3Software engineering, systems software and programming languages · 2 · 2 since 2021Security and privacy · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GLow - A Novel, Flower-Based Simulated Gossip Learning Strategyabstract• Creation of a Decentralized Federated Learning (Gossip Learning) strategy to simulate fully distributed agent configurations. • Deploy and evaluate custom network scenarios and assess how interconnection among agents affect distributed systems convergence. • Experimentation with MNIST and CIFAR10 datasets and 8, 16 network agents to second the viability of the designed Gossip Learning system. • Real-world application in the cybersecurity domain - Network Intrusion Detection Systems. Fully decentralized learning algorithms are still in an early stage of development. Creating modular Decentralized Federated Learning strategies, as Gossip Learning, is not trivial due to convergence challenges and Byzantine faults intrinsic in systems of decentralized nature. Our contribution provides a novel means to simulate custom Gossip Learning systems by leveraging the state-of-the-art Flower Framework. Specifically, we introduce GLow, allows researchers to train and assess scalability and convergence of devices, across custom network topologies, before making a physical deployment. The Flower Framework is selected for being a simulation featured library with a very active community on Federated Learning research. However, Flower exclusively includes vanilla Federated Learning strategies and, thus, is not originally designed to perform simulations without a centralized authority. GLow is presented to fill this gap and make simulation of Gossip Learning systems possible. The results achieved by GLow on the MNIST and CIFAR10 datasets show accuracies above 0.98 and 0.75, respectively, using double ring or denser topologies. More importantly, GLow performs similarly in terms of accuracy and convergence to its analogous Centralized and Federated approaches. Additional evidence is provided including irregular topologies as well as a cybersecurity use case, where a potential application of Decentralized Federated Learning is deployed using the TON_IOT dataset. Aitor Belenguer, Jose Antonio Pascual, Javier Navaridas |
J. Parallel Distributed Comput. | 3 |
| 2025 | A Review of Federated Learning Applications in Intrusion Detection SystemsabstractIntrusion detection systems are evolving into sophisticated systems that perform data analysis while searching for anomalies in their environment. The development of deep learning technologies paved the way to build more complex and effective threat detection models. However, training those models may be computationally infeasible in most Internet of Things devices. Current approaches rely on powerful centralized servers that receive data from all their parties — substantially affecting response times and operational costs due to the huge communication overheads and violating basic privacy constraints. To mitigate these issues, Federated Learning emerged as a promising approach, where different agents collaboratively train a shared model, without exposing training data to others or requiring a compute-intensive centralized infrastructure. This paper focuses on the application of Federated Learning approaches in the field of Intrusion Detection. Both technologies are described in detail and current scientific progress is reviewed and taxonomized. Finally, the paper highlights the limitations present in recent works and proposes some future directions for this technology. Aitor Belenguer, Jose Antonio Pascual, Javier Navaridas |
Comput. Networks | 3 |
| 2025 | Improving the performance of Dragonfly networks through restrictive Proxy routing strategiesabstractDragonfly has become the network of choice for large-scale high-performance computing systems and, indeed, it dominates the top positions of supercomputer rankings. The reason for this is that it offers a sweet spot in terms of cost, simplicity, performance, fault-tolerance and power consumption. In this work, we propose a collection of routing strategies which restrict proxies to be adjacent to either the local or the remote router. This way, it features shorter paths than the standard Valiant routing. We carry out an extensive simulation-based evaluation to assess their performance. Our experiments found latency reductions of up to 76% and throughput improvements of up to 26% when compared with standard Valiant routing when using synthetic traffic from independent sources at different scales. Furthermore, when using realistic application-inspired workloads, we found the strategies required between 5% and 20% less time to perform communications. In general, we observe that selecting proxies that are adjacent to the sender is more beneficial than those adjacent to the destination because the latter tends to generate backpressure in the last level of the interconnect. Interestingly, we found that the most restrictive proxy routing strategies obtain the best results in all scenarios and show that counterintuitively, the lower the path diversity, the more balanced the use of network resources. Our study includes investigating the interplay between routing and Dragonfly parameters and provide optimal parameters for proxy-based routing algorithms. Finally, we discuss some practical considerations related to the deployment of our strategies. Javier Navaridas, Jose Antonio Pascual |
Comput. Networks | 1 |
| 2024 | On the parallelization of multipacting simulation codes for the design of particle accelerator componentsabstractAbstract Particle trajectory and collision simulation is a critical step of the design and construction of novel particle accelerator components. However it requires a huge computational effort which can slow down the design process. We started from a sequential simulation program which is used to study an event called Multipacting. Our work explains the physical problem that is simulated and the implications it can have on the behavior of the components. Then we analyze the original program’s operation to find the best options for parallelization. We first developed a parallel version of the Multipacting simulation and were able to accelerate the execution up to $$\sim 35\times $$ ∼ 35 × with 48 or 56 cores. In the best cases, parallelization efficiency was maintained up to 16 cores ( $$\sim 95$$ ∼ 95 %) and the speed-up plateaus at around 40–48 cores. When this first parallelization effort was tried for multi-power simulations, we found that parallelism was severely limited with a maximum of $$20\times $$ 20 × speed-up. For this reason, we introduced a new method to improve the parallelization efficiency for this second use case. This method uses a shared processor pool for all simulations of electrons (OnePool). OnePool improved scalability by pushing the speed-up to over $$32\times $$ 32 × . Javier Navaridas, Jose Antonio Pascual, Julen Galarza, Txomin Romero, Juan L. Muñoz, Ibon Bustinduy |
J. Supercomput. | 1 |
| 2024 | Understanding the Impact of Arbitration in MZI-Based Beneš Switching FabricsabstractTop-of-rack switches based on photonic switching fabrics (PSF) could provide higher bandwidth and energy efficiency for datacenters (DC) and high-performance computers (HPC) than these with traditional electronic crossbars. However, because of their bufferless nature, PFS are affected by contention much more drastically than traditional packet-switched electronic networks where traffic can advance towards its destination, getting buffered upon encountering contention and resuming transmission once resources are freed. In contrast, PSFs stop the injection of all traffic that generate contention. Consequently, it is important to understand how the order in which flows are serviced affects performance metrics. Our contribution is to quantify this impact through a comprehensive simulation-based evaluation focusing on a recently fabricated PSF prototype. Our experiments include configurations with three routing algorithms, two switching methods, three ToR switch sizes and 9 representative workloads from the DC and HPC domains. We found that the effect of arbitration on raw throughput is negligible but, when considering more realistic loads, selecting an appropriate arbitration policy can improve communication time and energy efficiency. Indeed, the communication time can be reduced by between 10% and 30% by employing appropriate arbitration. Switching energy efficiency can also be improved between 4% and 13%. Finally, insertion loss is barely affected, with differences below 2%. LFU and ARR were found to obtain the best results. LFU is very good with regular workloads but one of the worse with irregular workloads. ARR obtains good results regardless of the type of workload. Javier Navaridas, Markos Kynigos, Jose Antonio Pascual, Mikel Luján, José Miguel-Alonso, John Goodacre |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2023 | SiliconBurmuin: A Horizon Europe propelled Neurocomputing Initiative in the Basque CountryabstractSiliconBurmuin is aimed at creating a multi-disciplinary neurocomputing community in the Basque Country, bringing together technology and scientific research centres and industry companies. This community will: (1) identify key biological structures and mechanisms that play a major role in vision across species, and (2) transform this knowledge into novel mathematical formalisms, neuromorphic designs and algorithms to solve industry challenges and enable new experiments of interest in neuroscience and clinical research. To achieve the latter objective in a time-effective manner, SiliconBurmuin will draw strong connections with the ongoing Horizon Europe NimbleAI project, with which it shares coordination. This is expected to allow reinforcement of ideas, knowledge and technology via a common prototyping platform where to implement IP from both projects. In addition to describing the research objectives and direction of SiliconBurmuin, this paper posits that co-coordination and co-funding of aligned projects at EU and regional levels might well be a catalyst for raising regional self-awareness of own potential and develop it to help fulfill global challenges, such as semiconductor sovereignty. Xabier Iturbe, Xabier Alberdi, Ander Aramburu, Armando Astarloa, Iñigo Barandiaran, Koldo Basterretxea, Angélica Dávila, Asier Erramuzpe, Iñigo Gabilondo, Garikoitz Lerma-Usabiaga, Lisandro Gabriel Monsalve, Libe Mori, Javier Navaridas, Jose Antonio Pascual, Joaquin Piriz, Serafim Rodrigues, Oscar Seijo, Ander Soraluze, Edgar Soria, Ignacio Torres, Nerea Uriarte, Juan Luis Valerdi |
SEAA | 13 |
| 2023 | A Novel Simulation Methodology for Silicon Photonic Switching FabricsabstractOptical communication based on silicon photonics is a promising candidate for future networks. However, a key component that still presents challenges is a practical, silicon photonics-based, high performance switch with a high port count. The impracticality of buffering traffic in the optical domain mandates the use of circuit switching at the transmission level. This renders the photonic power penalty dependent on many factors, including architectural aspects and, most importantly, the switch load. Since the latter changes dynamically with network traffic we argue that simulating silicon photonics-based switches requires considering the photonic power penalty under dynamic workloads, which is not supported by state-of-the-art techniques. In this paper, we show how to simultaneously simulate both the overall switch as well as the photonic power penalty, by proposing a novel combination of the bufferless nature of photonic fabrics, flow-level simulation and optical beam propagation modelling. This approach enables a simulator to consider different kinds of switching fabrics and photonic components. We focus on how to model Beneš photonic switching fabrics formed with Mach-Zehnder Interferometers and consider their deployment as switching cores for top-of-rack switches. We compare our simulation with the published data from two fabricated chips and found accuracy is within 0. 5dB with respect to insertion loss, and within 3dB with respect to crosstalk. As a use-case, we evaluate the impact of routing algorithms on the photonic power penalty and found this can reduce the worst-case photonic power penalty by up to 4dB. Markos Kynigos, Javier Navaridas, Jose Antonio Pascual, Mikel Luján |
ISPASS | 2 |
| 2023 | Parallelizing Multipacting Simulation for the Design of Particle Accelerator ComponentsabstractParticle trajectory and collision simulation is a critical step of the design and construction of novel particle accelerator components. However it requires a huge computational effort which can slow down the design process. We started from a sequential simulation program which is used to study an event called “Multipacting”. Our work explains the physical problem that is simulated and the implications it can have on the behavior of the components. Then we analyze the original program's operation to find the best options for parallelization. We first developed a parallel version of the Multipacting simulation and were able to accelerate the execution up to ~ 35× with 48 or 56 cores. In the best cases, parallelization efficiency was maintained up to 16 cores (~ 95%) and the speed-up plateaus at around 40 to 48 cores. When this first parallelization effort was tried for multi-power simulations, we found that parallelism was severely limited with a maximum of 20× speed-up. For this reason, we introduced a new method to improve the parallelization efficiency for this second use case. This method uses a shared processor pool for all simulations of electrons (OnePool). OnePool improved scalability by pushing the speed-up to over 32×. Julen Galarza, Javier Navaridas, Jose Antonio Pascual, Txomin Romero, Juan L. Muñoz, Ibon Bustinduy |
PDP | 2 |
| 2023 | GöwFed: A novel federated network intrusion detection systemabstractNetwork intrusion detection systems are evolving into intelligent systems that perform data analysis while searching for anomalies in their environment. Indeed, the development of deep learning techniques paved the way to build more complex and effective threat detection models. However, training those models may be computationally infeasible in most Edge or IoT devices. Current approaches rely on powerful centralized servers that receive data from all their parties — violating basic privacy constraints and substantially affecting response times and operational costs due to the huge communication overheads. To mitigate these issues, Federated Learning emerged as a promising approach, where different agents collaboratively train a shared model, without exposing training data to others or requiring a compute-intensive centralized infrastructure. This work presents GöwFed, a novel network threat detection system that combines the usage of Gower Dissimilarity matrices and Federated averaging. Different approaches of GöwFed have been developed based on state-of the-art knowledge: (1) a vanilla version — achieving a median point of [0.888, 0.960] in the PR space and a median accuracy of 0.930; and (2) a version instrumented with an attention mechanism — achieving comparable results when 0.8 of the best performing nodes contribute to the model. Furthermore, each variant has been tested using simulation oriented tools provided by TensorFlow Federated framework. In the same way, a centralized analogous development of the Federated systems is carried out to explore their differences in terms of scalability and performance — the median point of the experiments is [0.987, 0.987]) and the median accuracy is 0.989. Overall, GöwFed intends to be the first stepping stone towards the combined usage of Federated Learning and Gower Dissimilarity matrices to detect network threats in industrial-level networks. Aitor Belenguer, Jose Antonio Pascual, Javier Navaridas |
J. Netw. Comput. Appl. | 3 |
| 2021 | Power and energy efficient routing for Mach-Zehnder interferometer based photonic switchesabstractSilicon Photonic top-of-rack (ToR) switches are highly desirable for the datacenter (DC) and high-performance computing (HPC) domains for their potential high-bandwidth and energy efficiency. Recently, photonic Beneš switching fabrics based on Mach-Zehnder Interferometers (MZIs) have been proposed as a promising candidate for the internals of high-performance switches. However, state-of-the-art routing algorithms that control these switching fabrics are either computationally complex or unable to provide non-blocking, energy efficient routing permutations.To address this, we propose for the first time a combination of energy efficient routing algorithms and time-division multiplexing (TDM). We evaluate this approach by conducting a simulation-based performance evaluation of a 16x16 Beneš fabric, deployed as a ToR switch, when handling a set of 8 representative workloads from the DC and HPC domains. Our results show that state-of-the-art approaches (circuit switched energy efficient routing algorithms) introduce up to 23% contention in the switching fabric for some workloads, thereby increasing communication time. We show that augmenting the algorithms with TDM can ameliorate switch fabric contention by segmenting communication data and gracefully interleaving the segments, thus reducing communication time by up to 20% in the best case. We also discuss the impact of the TDM segment size, finding that although a 10KB segment size is the most beneficial in reducing communication time, a 100KB segment size offers similar performance while requiring a less stringent path-computation time window. Finally, we assess the impact of TDM on path-dependent insertion loss and switching energy consumption, finding it to be minimal in all cases. Markos Kynigos, Jose Antonio Pascual, Javier Navaridas, John Goodacre, Mikel Luján |
ICS | 3 |
| 2020 | Message from the Conference Chairs - ASAP 2020abstractPresents the introductory welcome message from the conference proceedings. May include the conference officers' congratulations to all involved with the conference event and publication of the proceedings record. Dirk Koch, Frank Hannig, Javier Navaridas |
ASAP | 3 |
| 2020 | Relating the bisection width of dual-port, server-centric datacenter networks and the solution of edge isoperimetric problems in graphsabstractStellar datacenter networks are a recent generic construction designed to transform a base-graph into a dual-port, server-centric datacenter network. We prove that the S-bisection width of any stellar datacenter network can be obtained from the solution of isoperimetric problems on the base-graph, provided that the base-graph is regular. We extend previous research on the stellar datacenter networks GQ⁎, instantiated with generalized hypercubes, and show that with respect to S-bisection width, GQ⁎ performs well in comparison with the dual-port datacenter network FiConn. Our work develops a strong combinatorial link between graph bisection width and throughput metrics for stellar datacenter networks. Alejandro Erickson, Javier Navaridas, Iain A. Stewart |
J. Comput. Syst. Sci. | 2 |
| 2019 | Design Exploration of Multi-tier Interconnection Networks for Exascale SystemsabstractInterconnection networks are one of the main limiting factors when it comes to scale out computing systems. In this paper, we explore what role the hybridization of topologies has on the design of an state-of-the-art exascale-capable computing system. More precisely we compare several hybrid topologies and compare with common single-topology ones when dealing with large-scale applicationlike traffic. In addition we explore how different aspects of the hybrid topology can affect the overall performance of the system. In particular, we found that hybrid topologies can outperform state-of-the-art torus and fattree networks as long as the density of connections is high enough--one connection every two or four nodes seems to be the sweet spot--and the size of the subtori is limited to a few nodes per dimension. Moreover, we explored two different alternatives to use in the upper tiers of the interconnect, a fattree and a generalised hypercube, and found little difference between the topologies, mostly depending on the workload to be executed. Javier Navaridas, Joshua Lant, Jose Antonio Pascual, Mikel Luján, John Goodacre |
ICPP | 1 |
| 2019 | Robust Covert Channels Based on DRAM Power Consumption
Thales Bandiera Paiva, Javier Navaridas, Routo Terada |
ISC | 2 |
| 2019 | Enabling shared memory communication in networks of MPSoCsabstractSummary Ongoing transistor scaling and the growing complexity of embedded system designs has led to the rise of MPSoCs (Multi‐Processor System‐on‐Chip), combining multiple hard‐core CPUs and accelerators (FPGA, GPU) on the same physical die. These devices are of great interest to the supercomputing community, who are increasingly reliant on heterogeneity to achieve power and performance goals in these closing stages of the race to exascale. In this paper, we present a network interface architecture and networking infrastructure, designed to sit inside the FPGA fabric of a cutting‐edge MPSoC device, enabling networks of these devices to communicate within both a distributed and shared memory context, with reduced need for costly software networking system calls. We will present our implementation and prototype system and discuss the main design decisions relevant to the use of the Xilinx Zynq Ultrascale+, a state‐of‐the‐art MPSoC, and the challenges to be overcome given the device's limitations and constraints. We demonstrate the working prototype system connecting two MPSoCs, with communication between processor and remote memory region and accelerator. We then discuss the limitations of the current implementation and highlight areas of improvement to make this solution production‐ready. Joshua Lant, Caroline Concatto, Andrew Attwood, Jose Antonio Pascual, Mike Ashworth, Javier Navaridas, Mikel Luján, John Goodacre |
Concurr. Comput. Pract. Exp. | 6 |
| 2019 | On the effects of allocation strategies for exascale computing systems with distributed storage and unified interconnectsabstractSummary The convergence between computing‐ and data‐centric workloads and platforms is imposing new challenges on how to best use the resources of modern computing systems. In this paper, we investigate alternatives for the storage subsystem of a novel exascale‐capable system with special emphasis on how allocation strategies would affect the overall performance. We consider several aspects of data‐aware allocation such as the effect of spatial and temporal locality, the affinity of data to storage sources, and the network‐level traffic prioritization for different types of flows. In our experimental set‐up, temporal locality can have a substantial effect on application runtime (up to a 10% reduction), whereas spatial locality can be even more significant (up to one order of magnitude faster with perfect locality). The use of structured access patterns to the data and the allocation of bandwidth at the network level can also have a significant impact (up to 20% and 17% reduction of runtime, respectively). These results suggest that scheduling policies exposing data‐locality information can be essential for the appropriate utilization of future large‐scale systems. Finally, we found that the distributed storage system we are implementing can outperform traditional SAN architectures, even with a much smaller (in terms of I/O servers) back‐end. Jose Antonio Pascual, Joshua Lant, Caroline Concatto, Andrew Attwood, Javier Navaridas, Mikel Luján, John Goodacre |
Concurr. Comput. Pract. Exp. | 5 |
| 2019 | INRFlow: An interconnection networks research flow-level simulation frameworkabstractThis paper presents INRFlow, a mature, frugal, flow-level simulation framework for modelling large-scale networks and computing systems. INRFlow is designed to carry out performance-related studies of interconnection networks for both high performance computing systems and datacentres. It features a completely modular design in which adding new topologies, routings or traffic models requires minimum effort. Moreover, INRFlow includes two different simulation engines: a static engine that is able to scale to tens of millions of nodes and a dynamic one that captures temporal and causal relationships to provide more realistic simulations. We will describe the main aspects of the simulator, including system models, traffic models and the large variety of topologies and routings implemented so far. We conclude the paper with a case study that analyses the scalability of several typical topologies. INRFlow has been used to conduct a variety of studies including evaluation of novel topologies and routings (both in the context of graph theory and optimization), analysis of storage and bandwidth allocation strategies and understanding of interferences between application and storage traffic. • We present our flow-level simulation framework INRFlow. • It is a mature, flexible and efficient tool for simulating large scale systems. • It models network, storage, scheduler and applications. • It has been used extensively for our research in the past. • INRFlow is open source and programmed in C. Javier Navaridas, Jose Antonio Pascual, Alejandro Erickson, Iain A. Stewart, Mikel Luján |
J. Parallel Distributed Comput. | 1 |
| 2018 | Network-on-chip evaluation for a novel neural architectureabstractThis paper provides a performance evaluation and trade-off analysis of a novel chip architecture for neuromorphic computing, especially focused on the memory subsystems and the Network-On-Chip (NoC). More precisely, we study the performance-related effect of the number of memory modules, as well as that of allowing direct core-to-core communication. Our simulation-based experimental work throws many interesting results on the above aspects and allows to ensure that congestion at the NoC-level is unlikely to degrade performance. Markos Kynigos, Javier Navaridas, Luis A. Plana, Steve Furber |
CF | 2 |
| 2018 | High-Performance, Low-Complexity Deadlock Avoidance for Arbitrary Topologies/RoutingsabstractRecently, the use of graph-based network topologies has been proposed as an alternative to traditional networks such as tori or fat-trees due to their very good topological characteristics. However they pose practical implementation challenges such as the lack of deadlock avoidance strategies. Previous proposals either lack flexibility, underutilise network resources or are exceedingly complex. We propose--and prove formally--three generic, low-complexity deadlock avoidance mechanisms that only require local information. Our methods are topology- and routing-independent and their virtual channel count is bounded by the length of the longest path. We evaluate our algorithms through an extensive simulation study to measure the impact on the performance using both synthetic and realistic traffic. First we compare against a well-known HPC mechanism for dragonfly and achieve similar performance level. Then we moved to Graph-based networks and show that our mechanisms can greatly outperform traditional, spanning-tree based mechanisms, even if these use a much larger number of virtual channels. Overall, our proposal provides a simple, flexible and high performance deadlock-avoidance solution. Jose Antonio Pascual, Javier Navaridas |
ICS | 2 |
| 2017 | Designing an exascale interconnect using multi-objective optimizationabstractExascale performance will be delivered by systems composed of millions of interconnected computing cores. The way these computing elements are connected with each other (network topology) has a strong impact on many performance characteristics. In this work we propose a multi-objective optimization-based framework to explore possible network topologies to be implemented in the EU-funded ExaNeSt project. The modular design of this system's interconnect provides great flexibility to design topologies optimized for specific performance targets such as communications locality, fault tolerance or energy-consumption. The generation procedure of the topologies is formulated as a three-objective optimization problem (minimizing some topological characteristics) where solutions are searched using evolutionary techniques. The analysis of the results, carried out using simulation, shows that the topologies meet the required performance objectives. In addition, a comparison with a well-known topology reveals that the generated solutions can provide better topological characteristics and also higher performance for parallel applications. Jose Antonio Pascual, Joshua Lant, Andrew Attwood, Caroline Concatto, Javier Navaridas, Mikel Luján, John Goodacre |
CEC | 5 |
| 2017 | The Next Generation of Exascale-Class Systems: The ExaNeSt ProjectabstractThe ExaNeSt project started on December 2015 and is funded by EU H2020 research framework (call H2020-FETHPC-2014, n. 671553) to study the adoption of low-cost, Linux-based power-efficient 64-bit ARM processors clusters for Exascale-class systems. The ExaNeSt consortium pools partners with industrial and academic research expertise in storage, interconnects and applications that share a vision of an Euro-pean Exascale-class supercomputer. Their goal is designing and implementing a physical rack prototype together with its cooling system, the storage non-volatile memory (NVM) architecture and a low-latency interconnect able to test different options for interconnection and storage. Furthermore, the consortium is to provide real HPC applications to validate the system. Herein we provide a status report of the project initial developments. Roberto Ammendola, Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Pier Stanislao Paolucci, Elena Pastorelli, Francesco Simula, Piero Vicini, Giuliano Taffoni, Jose Antonio Pascual, Javier Navaridas, Mikel Luján, John Goodacre, Nikolaos Chrysos, Manolis Katevenis |
DSD | 14 |
| 2017 | Designing Low-Power, Low-Latency Networks-on-Chip by Optimally Combining Electrical and Optical LinksabstractOptical on-chip communication is considered a promising candidate to overcome latency and energy bottlenecks of electrical interconnects. Although recently proposed hybrid Networks-on-chip (NoCs), which implement both electrical and optical links, improve power efficiency, they often fail to combine these two interconnect technologies efficiently and suffer from considerable laser power overheads caused by high-bandwidth optical links. We argue that these overheads can be avoided by inserting a higher quantity of low-bandwidth optical links in a topology, as this yields lower optical loss and in turn laser power. Moreover, when optimally combined with electrical links for short distances, this can be done without trading off latency. We present the effectiveness of this concept with Lego, our hybrid, mesh-based NoC that provides high power efficiency by utilizing electrical links for local traffic, and low-bandwidth optical links for long distances. Electrical links are placed systematically to outweigh the serialization delay introduced by the optical links, simplify router microarchitecture, and allow to save optical resources. Our routing algorithm always chooses the link that offers the lowest latency and energy. Compared to state-of-the-art proposals, Lego increases throughput-per-watt by at least 40%, and lowers latency by 35% on average for synthetic traffic. On SPLASH-2/PARSEC workloads, Lego improves power efficiency by at least 37% (up to 3.5×). Sebastian Werner 0002, Javier Navaridas, Mikel Luján |
HPCA | 2 |
| 2017 | The stellar transformation: From interconnection networks to datacenter networksabstractThe first dual-port server-centric datacenter network, FiConn, was introduced in 2009 and there are several others now in existence; however, the pool of topologies to choose from remains small. We propose a new generic construction, the stellar transformation, that dramatically increases the size of this pool by facilitating the transformation of well-studied topologies from interconnection networks, along with their networking properties and routing algorithms, into viable dual-port server-centric datacenter network topologies. We demonstrate that under our transformation, numerous interconnection networks yield datacenter network topologies with potentially good, and easily computable, baseline properties. We instantiate our construction so as to apply it to generalized hypercubes and obtain the datacenter networks GQ⋆. Our construction automatically yields routing algorithms for GQ⋆ and we empirically compare GQ⋆ (and its routing algorithms) with the established datacenter networks FiConn and DPillar (and their routing algorithms); this comparison is with respect to network throughput, latency, load balancing, fault-tolerance, and cost to build, and is with regard to all-to-all, many all-to-all, butterfly, random, hot-region, and hot-spot traffic patterns. We find that GQ⋆ outperforms both FiConn and DPillar (sometimes significantly so) and that there is substantial scope for our stellar transformation to yield new dual-port server-centric datacenter networks that are a considerable improvement on existing ones. Alejandro Erickson, Iain A. Stewart, Javier Navaridas, Abbas Eslami Kiasari |
Comput. Networks | 3 |
| 2017 | Improved routing algorithms in the dual-port datacenter networks HCN and BCNabstractWe present significantly improved one-to-one routing algorithms in the datacenter networks HCN and BCN in that our routing algorithms result in much shorter paths when compared with existing routing algorithms. We also present a much tighter analysis of HCN and BCN by observing that there is a very close relationship between the datacenter networks HCN and the interconnection networks known as WK-recursive networks. We use existing results concerning WK-recursive networks to prove the optimality of our new routing algorithm for HCN and also to significantly aid the implementation of our routing algorithms in both HCN and BCN. Furthermore, we empirically evaluate our new routing algorithms for BCN, against existing ones, across a range of metrics relating to path-length, throughput, and latency for the traffic patterns all-to-one, bisection, butterfly, hot-region, many-all-to-all, and uniform-random, and we also study the completion times of workloads relating to MapReduce, stencil and sweep, and unstructured applications. Not only do our results significantly improve routing in our datacenter networks for all of the different scenarios considered but they also emphasize that existing theoretical research can impact upon modern computational platforms. Alejandro Erickson, Iain A. Stewart, Jose Antonio Pascual, Javier Navaridas |
Future Gener. Comput. Syst. | 4 |
| 2017 | An Optimal Single-Path Routing Algorithm in the Datacenter Network DPillarabstractDPillar has recently been proposed as a server-centric datacenter network and is combinatorially related to (but distinct from) the well-known wrapped butterfly network. We explain the relationship between DPillar and the wrapped butterfly network before proving that the underlying graph of DPillar is a Cayley graph; hence, the datacenter network DPillar is node-symmetric. We use this symmetry property to establish a single-path routing algorithm for DPillar that computes a shortest path and has time complexity O(k), where k parameterizes the dimension of DPillar (we refer to the number of ports in its switches as n). Our analysis also enables us to calculate the diameter of DPillar exactly. Moreover, our algorithm is trivial to implement, being essentially a conditional clause of numeric tests, and improves significantly upon a routing algorithm earlier employed for DPillar. Furthermore, we provide empirical data in order to demonstrate this improvement. In particular, we empirically show that our routing algorithm improves the average length of paths found, the aggregate bottleneck throughput, and the communication latency. A secondary, yet important, effect of our work is that it emphasises that datacenter networks are amenable to a closer combinatorial scrutiny that can significantly improve their computational efficiency and performance. Alejandro Erickson, Abbas Eslami Kiasari, Javier Navaridas, Iain A. Stewart |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | Handling Physical-Layer Deadlock Caused by Permanent Faults in Quasi-Delay-Insensitive Networks-on-ChipabstractNetworks-on-Chip (NoCs) are promising fabrics to provide scalable and efficient on-chip communication for large-scale many-core systems. In place of the well-studied synchronous NoCs, the event-driven asynchronous ones have emerged as promising replacement thanks to their strong timing robustness especially when implemented in quasi-delay-insensitive (QDI) circuits. However, their fault tolerance has rarely been studied. The QDI NoCs show complicated failure scenarios and behave differently from synchronous ones. As the scaling semiconductor technology is expected with the accelerated aging process, permanent faults become more likely to happen at runtime. These faults can break the handshake, leading to physical-layer deadlocks which can spread and paralyze the whole QDI NoC. This physical-layer deadlock cannot be resolved using conventional fault-tolerant or deadlock management techniques. This paper systematically studies the impact of permanent faults on QDI NoCs, and presents novel deadlock detection and recovery techniques to handle the fault-caused physical-layer deadlock. The proposed detection technique has been implemented to protect the NoC data paths that occupy~90% of the logic. Employing the detection and recovery techniques to protect interrouter links (~60% of the logic), a permanently faulty link is precisely located and the network function can be recovered with graceful performance degradation. Guangda Zhang, Wei Song 0002, Jim D. Garside, Javier Navaridas, Zhiying Wang 0003 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | The ExaNeSt Project: Interconnects, Storage, and Packaging for Exascale SystemsabstractExaNest is one of three European projects that support a ground-breaking computing architecture for exascale-class systems built upon power-efficient 64-bit ARM processors. This group of projects share an "everything-close" and "share-anything" paradigm, which trims down the power consumption -- by shortening the distance of signals for most data transfers -- as well as the cost and footprint area of the installation -- by reducing the number of devices needed to meet performance targets. In ExaNeSt, we will design and implement: (i) a physical rack prototype and its liquid-cooling subsystem providing ultra-dense compute packaging, (ii) a storage architecture with distributed (in-node) non-volatile memory (NVM) devices, (iii) a unified, low-latency interconnect, designed to efficiently uphold desired Quality-of-Service guarantees for a mix of storage with inter-processor flows, and (iv) efficient rack-level memory sharing, where each page is cacheable at only a single node. Our target is to test alternative storage and interconnect options on actual hardware, using real-world HPC applications. The ExaNeSt consortium brings together technology, skills, and knowledge across the entire value chain, from computing IP, packaging, and system deployment, all the way up to operating systems, storage, HPC, big data frameworks, and cutting-edge applications. Manolis Katevenis, Nikolaos Chrysos, Manolis Marazakis, Iakovos Mavroidis, Fabien Chaix, Nikolaos D. Kallimanis, Javier Navaridas, John Goodacre, Piero Vicini, Andrea Biagioni, Pier Stanislao Paolucci, Alessandro Lonardo, Elena Pastorelli, Francesca Lo Cicero, Roberto Ammendola, P. Hopton, P. Coates, Giuliano Taffoni, Stefano Cozzini, Martin L. Kersten, Julio Sahuquillo, Sergio Lechago, C. Pinto, Bernd Lietzow, D. Everett, Gino Perna |
DSD | 7 |
| 2015 | An Efficient Shortest-Path Routing Algorithm in the Data Centre Network DPillar
Alejandro Erickson, Abbas Eslami Kiasari, Javier Navaridas, Iain A. Stewart |
COCOA | 3 |
| 2015 | Accelerating Interconnect Analysis Using High-Level HDLs and FPGA, SpiNNaker as a Case StudyabstractArchitectural simulation is a fundamental tool for modern computing system design. Computer architects can choose from a large set of software simulators that provide a robust and efficient platform for design exploration although, if the new architecture incorporates unconventional or novel features not supported by the existing simulators, the designers must develop their own. As reconfigurable hardware platforms grow more computationally capable we observe a move from software simulators towards these hardware platforms. The introduction of high-level HDLs, such as Blue spec System Verilog (BSV), offers improved productivity while still providing a tool flow capable of exploiting reconfigurable platforms. This paper focuses on understanding how to accelerate the simulation of the interconnection network of SpiNNaker [1], a massively-parallel computer for neural simulation. We analysed the modelling choices and trade-offs made during the implementation of the software (SW) model as well as those made when developing a new hardware model (HW) built on a Xilinx FPGA. Mohsen Ghasempour, Jonathan Heathcote, Javier Navaridas, Luis A. Plana, Jim D. Garside, Mikel Luján |
FCCM | 3 |
| 2015 | On Routing Algorithms for the DPillar Data Centre Networks
Abbas Eslami Kiasari, Javier Navaridas, Iain A. Stewart |
ICA3PP (4) | 2 |
| 2015 | SpiNNaker: Enhanced multicast routingabstractThe human brain is a complex biological neural network characterised by high degrees of connectivity among neurons. Any system designed to simulate large-scale spiking neuronal networks needs to support such connectivity and the associated communication traffic in the form of spike events. This paper investigates how best to generate multicast routes for SpiNNaker, a purpose-built, low-power, massively-parallel architecture. The discussed algorithms are an essential ingredient for the efficient operation of SpiNNaker since generating multicast routes is known to be an NP-complete problem. In fact, multicast communications have been extensively studied in the literature, but we found no existing algorithm adaptable to SpiNNaker. The proposed algorithms exploit the regularity of the two-dimensional triangular torus topology and the availability of selective multicast at hardware level. A comprehensive study of the parameters of the algorithms and their effectiveness is carried out in this paper considering different destination distributions ranging from worst-case to a real neural application. The results show that two novel proposed algorithms can reduce significantly the pressure exerted onto the interconnection infrastructure while remaining effective to be used in a production environment. Javier Navaridas, Mikel Luján, Luis A. Plana, Steve Temple, Steve Furber |
Parallel Comput. | 1 |
| 2014 | An empirical evaluation of High-Level Synthesis languages and tools for database accelerationabstractHigh Level Synthesis (HLS) languages and tools are emerging as the most promising technique to make FPGAs more accessible to software developers. Nevertheless, picking the most suitable HLS for a certain class of algorithms depends on requirements such as area and throughput, as well as on programmer experience. In this paper, we explore the different trade-offs present when using a representative set of HLS tools in the context of Database Management Systems (DBMS) acceleration. More specifically, we conduct an empirical analysis of four representative frameworks (Bluespec SystemVerilog, Altera OpenCL, LegUp and Chisel) that we utilize to accelerate commonly-used database algorithms such as sorting, the median operator, and hash joins. Through our implementation experience and empirical results for database acceleration, we conclude that the selection of the most suitable HLS depends on a set of orthogonal characteristics, which we highlight for each HLS framework. Oriol Arcas-Abella, Geoffrey Ndu, Nehir Sönmez, Mohsen Ghasempour, Adrià Armejach, Javier Navaridas, Wei Song 0002, John Mawer, Adrián Cristal, Mikel Luján |
FPL | 6 |
| 2013 | Transient Fault Tolerant QDI Interconnects Using Redundant Check CodeabstractAsynchronous logic is a promising technology for building the chip-level interconnect of multi-core systems. However, asynchronous circuits are vulnerable to faults. This paper presents a novel scheme to improve the robustness of asynchronous systems. Our first contribution is a fault tolerant delay-insensitive redundant check coding scheme named DIRC. Using DIRC in 4-phase 1-of-n quasi-delay-insensitive (QDI) interconnects, all 1-bit and some multi-bit transient faults can be tolerated. The DIRC and the basic 4-phase 1-of-n pipeline stages are mutually exchangeable so that arbitrary basic stages can be replaced by DIRC stages to strengthen the fault-tolerance of long wires. Our second contribution, RPA, is a redundant technique to protect the acknowledge wires from transient faults - an issue that has long been disregarded by the community. The DIRC pipelines (using DIRC plus RPA) were simulated using the UMC 0.13μm standard cell library and compared with the basic pipelines. Detailed experimental results show that the 128-bit DIRC 1-of-4 pipeline is only 13% slower than the basic one but increases fault-tolerance hundred-folds when multi-bit transient faults are considered. Guangda Zhang, Wei Song 0002, Jim D. Garside, Javier Navaridas, Zhiying Wang 0003 |
DSD | 4 |
| 2013 | SpiNNaker: Fault tolerance in a power- and area- constrained large-scale neuromimetic architectureabstractSpiNNaker is a biologically-inspired massively-parallel computer designed to model up to a billion spiking neurons in real-time. A full-fledged implementation of a SpiNNaker system will comprise more than 105 integrated circuits (half of which are SDRAMs and half multi-core systems-on-chip). Given this scale, it is unavoidable that some components fail and, in consequence, fault-tolerance is a foundation of the system design. Although the target application can tolerate a certain, low level of failures, important efforts have been devoted to incorporate different techniques for fault tolerance. This paper is devoted to discussing how hardware and software mechanisms collaborate to make SpiNNaker operate properly even in the very likely scenario of component failures and how it can tolerate system-degradation levels well above those expected. Javier Navaridas, Steve Furber, Jim D. Garside, Xin Jin 0003, Mukaram M. Khan, David R. Lester, Mikel Luján, José Miguel-Alonso, Eustace Painkras, Cameron Patterson, Luis A. Plana, Alex Rast, Dominic Richards, Yebin Shi, Steve Temple, Shufan Yang |
Parallel Comput. | 1 |
| 2012 | Population-based routing in the SpiNNaker neuromorphic architectureabstractSpiNNaker is a hardware-based massively-parallel real-time universal neural network simulator designed to simulate large-scale spiking neural networks. Spikes are distributed across the system using a multicast packet router. Each packet represents an event (spike) generated by a neuron. On the basis of the source of the spike (chip, core and neuron), the routers distribute the network packet across the system towards the destination neuron(s). This paper describes a novel approach to the projection routing problem that shows advantages in both the size of the routing tables generated and the computational complexity for the generation of routing tables. To achieve this, spikes are routed on the basis of the source population, leaving to the destination core the duty to propagate the received spike to the appropriate neuron(s). Sergio Davies, Javier Navaridas, Francesco Galluppi, Steve Furber |
IJCNN | 2 |
| 2012 | Reservation-based Network-on-Chip Timing Models for Large-scale Architectural SimulationabstractArchitectural simulation is an essential tool when it comes to evaluating the design of future many-core chips. However, reproducing all the components of such complex systems precisely would require unreasonable amounts of computing power. Hence, a trade off between accuracy and compute time is needed. For this reason most state-of-the-art tools do not have accurate models for the networks-on-chip, and rely on timing models that permit fast-simulation. Generally, these models are very simplistic and disregard contention for the use of network resources. As the number of nodes in the network-on-chip grows, fluctuations with contention and other parameters can considerably affect the accuracy of such models. In this paper we present and evaluate a collection of timing models based on a reservation scheme which consider the contention for the use of network resources. These models provide results quickly while being more accurate than simple no-contention approaches. Javier Navaridas, Behram Khan, Salman Khan 0002, Paolo Faraboschi, Mikel Luján |
NOCS | 1 |
| 2012 | Scalable communications for a million-core neural processing architecture
Cameron Patterson, Jim D. Garside, Eustace Painkras, Steve Temple, Luis A. Plana, Javier Navaridas, Thomas Sharp, Steve Furber |
J. Parallel Distributed Comput. | 6 |
| 2011 | Event-driven configuration of a neural network CMP system over an homogeneous interconnect fabric
Mukaram M. Khan, Alex Rast, Javier Navaridas, X. Jin, Luis A. Plana, Mikel Luján, Steve Temple, Cameron Patterson, Dominic Richards, John V. Woods, José Miguel-Alonso, Steve Furber |
Parallel Comput. | 3 |
| 2010 | Reducing complexity in tree-like computer interconnection networks
Javier Navaridas, José Miguel-Alonso, Francisco Javier Ridruejo, Wolfgang E. Denzel |
Parallel Comput. | 1 |
| 2010 | Twisted Torus Topologies for Enhanced Interconnection NetworksabstractMany current parallel computers are built around a torus interconnection network. Machines from Cray, HP, and IBM, among others, make use of this topology. In terms of topological advantages, square (2D) or cubic (3D) tori would be the topologies of choice. However, for different practical reasons, 2D and 3D tori with different number of nodes per dimension have been used. These mixed-radix topologies are not edge symmetric, which translates into poor performance due to an unbalanced use of network resources. In this work, we analyze twisted 2D and 3D mixed-radix tori that remove the network bottlenecks present in nontwisted ones. Such topologies recover edge symmetry, and consequently, balance the utilization of their links. The distance-related properties of twisted tori together with a full characterization of their bisection bandwidth are described in this paper. A simulation-based performance evaluation has been carried out to assess the network performance under synthetic and trace-driven workloads. The obtained results show noticeable and consistent performance gains (up to an increase of 74 percent in accepted load). In addition, we propose scalable and practicable packet routing mechanisms and wiring layouts for these interconnection systems. The complexity of the architectural proposals is similar to the one exhibited by routing and folding mechanisms in standard tori. José M. Cámara, Miquel Moretó, Enrique Vallejo 0001, Ramón Beivide, José Miguel-Alonso, Carmen Martínez 0001, Javier Navaridas |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2009 | Understanding the interconnection network of SpiNNakerabstractSpiNNaker is a massively parallel architecture designed to model large-scale spiking neural networks in (biological) real-time. Its design is based around ad-hoc multi-core System-on-Chips which are interconnected using a two-dimensional toroidal triangular mesh. Neurons are modeled in software and their spikes generate packets that propagate through the on- and inter-chip communication fabric relying on custom-made on-chip multicast routers. This paper models and evaluates large-scale instances of its novel interconnect (more than 65 thousand nodes, or over one million computing cores), focusing on real-time features and fault-tolerance. The key contribution can be summarized as understanding the properties of the feasible topologies and establishing the stable operation of the SpiNNaker under different levels of degradation. First we derive analytically the topological characteristics of the network, which are later confirmed by experimental work. With the computational model developed, we investigate the topology of SpiNNaker, and compare it with a standard 3-dimensional torus. The novel emergency routing mechanism, implemented within the routers, allows the topology of SpiNNaker to be more robust than the 3-dimensional torus, regardless of the latter having better topological characteristics. Furthermore, we obtain optimal values of two router parameters related with livelock and deadlock avoidance mechanisms. Javier Navaridas, Mikel Luján, José Miguel-Alonso, Luis A. Plana, Steve Furber |
ICS | 1 |
| 2009 | Event-Driven Configuration of a Neural Network CMP System over a Homogeneous Interconnect FabricabstractConfiguring a million-core parallel system at boot time is a difficult process when the system has neither specialised hardware support for the configuration process nor a preconfigured default state that puts it in operating condition. SpiNNaker is a parallel chip multiprocessor (CMP) system for neural network (NN) simulation. Where most large CMP systems feature a sideband network to complete the boot process, SpiNNaker has a single homogeneous network interconnect for both application inter-processor communications and system control functions such as boot load and run-time user-system interaction. This network improves fault tolerance and makes it easier to support dynamic run-time reconfiguration, however, it requires a boot process that is transaction-level compatible with the application's communications model. Since SpiNNaker uses event-driven asynchronous communications throughout, the loader operates with purely local control: there is no global synchronisation, state information, or transition sequence. A novel two-stage ldquounfoldingrdquo boot-up process efficiently configures the SpiNNaker hardware and loads the application using a high-speed flood-fill technique with support for run-time re-configuration. SystemC simulation of a multi-CMP SpiNNaker system indicates an error-free CMP configuration time of 1.3 ms, while a high-level simulation of a full-scale system (64 K CMPs) indicates a mean application-loading time of ~20 ms (for a 100 KB application), which is virtually independent of the size of the system. We verified the CMP configuration process with hardware-level Verilog simulation. Muhammad Mukaram Khan, Javier Navaridas, Alex Rast, Xin Jin 0003, Luis A. Plana, Mikel Luján, John V. Woods, José Miguel-Alonso, Steve Furber |
ISPDC | 2 |
| 2009 | Realistic Evaluation of Interconnection Networks Using Synthetic TrafficabstractEvaluation of high performance parallel systems is a delicate issue, due to the difficulty of generating workloads that represent, those that will run on actual systems. We overview the most usual workloads for performance evaluation purposes, in the scope of interconnection networks simulation. Aiming to fill the gap between purely synthetic and application-driven workloads, we present a set of synthetic communication micro-kernels that enhance regular synthetic traffic by adding point-to-point causality. They are conceived to stress the interconnection architecture. As an example of the proposed methodology, we use these micro-kernels to evaluate a topological improvement of k-ary n-cubes. Javier Navaridas, José Miguel-Alonso |
ISPDC | 1 |
| 2009 | Effects of Topology-Aware Allocation Policies on Scheduling Performance
Jose Antonio Pascual, Javier Navaridas, José Miguel-Alonso |
JSSPP | 2 |
| 2009 | Effects of Job and Task Placement on Parallel Scientific Applications PerformanceabstractThis paper studies the influence that task placement may have on the performance of applications, mainly due to the relationship between communication locality and overhead. This impact is studied for torus and fat-tree topologies. A simulation-based performance study is carried out, using traces of applications and application kernels, to measure the time taken to complete one or several concurrent instances of a given workload. As the purpose of the paper is not to offer a miraculous task placement strategy, but to measure the impact that placement have on performance, we selected simple strategies, including random placement. The quantitative results of these experiments show that different workloads present different degrees of responsiveness to placement. Furthermore, both the number of concurrent parallel jobs sharing a machine and the size of its network has a clear impact on the time to complete a given workload. We conclude that the efficient exploitation of a parallel computer requires the utilization of scheduling policies aware of application behavior and network topology. Javier Navaridas, Jose Antonio Pascual, José Miguel-Alonso |
PDP | 1 |
| 2008 | On synthesizing workloads emulating MPI applicationsabstractEvaluation of high performance parallel systems is a delicate issue, due to the difficulty of generating workloads that represent, with fidelity, those that will run on actual systems. In this paper we make an overview of the most usual methodologies used to generate workloads for performance evaluation purposes, focusing on the network: random traffic, patterns based on permutations, traces, execution-driven, etc. In order to fill the gap between purely synthetic and application- driven workloads, we present a set of pseudo-synthetic workloads that mimic applications behavior, emulating some widely-used implementations of MPI collectives and some communication patterns commonly used in scientific applications. This mimicry is done in terms of spatial distribution of messages as well as in the causal relationship among them. As an example of the proposed methodology, we use a subset of these workloads to confront tori and fat-trees. Javier Navaridas, José Miguel-Alonso, Francisco Javier Ridruejo |
IPDPS | 1 |
| 2007 | Concepts and components of full-system simulation of distributed memory parallel computersabstractIn this work we discuss a range of approaches to full-system simulation of distributed memory parallel computers, with emphasis on the interconnection network. We present our environment, based on Simics®, and discuss how unforeseen interactions and fine tuning of components can affect results. Francisco Javier Ridruejo, José Miguel-Alonso, Javier Navaridas |
HPDC | 3 |
| 2007 | Mixed-radix Twisted Torus Interconnection NetworksabstractMany parallel computers use Tori interconnection networks. Machines from Cray, HP and IBM, among others, exploit these topologies. In order to maintain full network symmetry, 2D and 3D Tori must have the same number of nodes (k) per dimension resulting in square or cubic topologies. Nevertheless, for practical reasons, computer engineers have designed and built 2D and 3D Tori having a different number of nodes per dimension. These mixed-radix topologies are not edge-symmetric which translates into poor performance provoked by an unbalanced use of the network links. In this paper, we propose and analyze twisted 2D and 3D Tori which remove the network bottlenecks present in mixed-radix standard Tori. These new topologies recover edge-symmetry and, consequently, balance the utilization of their links. We describe the distance-related parameters of these twisted networks and use simulation to asses their performance under synthetic loads. The obtained results show noticeable and consistent performance gains. In addition, we propose scalable and practicable packet routing and folding techniques for these interconnection subsystems. The complexity of the resulting architectural solutions is similar to the one exhibited by traditional routing and folding mechanisms employed in standard Tori. This fact together with the performance improvements obtained could justify the use of these twisted topologies in the future. José M. Cámara, Miquel Moretó, Enrique Vallejo 0001, Ramón Beivide, José Miguel-Alonso, Carmen Martínez 0001, Javier Navaridas |
IPDPS | 7 |
| 2007 | Realistic Evaluation of Interconnection Network Performance at High LoadsabstractAny simulation-based evaluation of an interconnection network proposal requires a good characterization of the workload. Synthetic traffic patterns based on independent traffic sources are commonly used to measure performance in terms of average latency and peak throughput. As they do not capture the level of self-throttling that occurs in most parallel applications, they can produce inaccurate throughput estimates at high loads. Thus, workloads that resemble the varying levels of synchronization of actual applications are needed to study the performance of interconnection networks. One approach is to use simple, burst-synchronized synthetic workloads that emulate the self-throttling of many parallel applications. To validate this approach, we compare the gains achieved by a restrictive injection mechanism under this workload with those obtained using traces from the NAS Parallel Benchmarks. This study confirms that the burst-synchronized traffic model provides reasonable performance estimates, which could be improved by taking into account dependency chains between messages. Francisco Javier Ridruejo, Javier Navaridas, José Miguel-Alonso, Cruz Izu |
PDCAT | 2 |