EDBT 2026 Demo / reviewers in the wild / expert
Nikolaos Chrysos
dblp:52/6055 · also Nikos Chrysos
· DBLP profile ↗
31ranked-venue papers
12as first author
6since 2021 · last 2025
0009-0002-8497-6985ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 1 first-author · 5 since 2021Computer networks · 13 · 9 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NET4EXA: Pioneering the Future of Interconnects for Supercomputing and AIabstractNET4EXA aims to develop a next-generation high-performance interconnect for HPC and AI systems, addressing the increasing demands of large-scale infrastructures, such as those required for training Large Language Models. Building upon the proven BXI (Bull eXascale Interconnect) European technology used in TOP15 supercomputers, NET4EXA will deliver the new BXI release, BXIv3, a complete hardware and software interconnect solution, including switch and network interface components. The project will integrate a fully functional pilot system at TRL 8, ready for deployment into upcoming exascale and post-exascale systems from 2025 onward. Leveraging prior research from European initiatives like RED-SEA, the previous achievements of consortium partners and over 20 years of expertise from BULL, NET4EXA also lays the groundwork for the future generation of BXI, BXIv4, providing analysis and preliminary design. The project will use a hybrid development and co-design approach, combining commercial switch technology with custom IP and FPGA-based NICs. Performances of NET4EXA BXIv3 interconnect will be evaluated using a broad portfolio of benchmarks, scientific scalable applications, and AI workloads. Michele Martinelli, Roberto Ammendola, Andrea Biagioni, Carlotta Chiarini, Ottorino Frezza, Francesca Lo Cicero, Alessandro Lonardo, Pier Stanislao Paolucci, Elena Pastorelli, Pierpaolo Perticaroli, Luca Pontisso, Cristian Rossi, Francesco Simula, Piero Vicini, David Colin, Gregoire Pichon, Alexandre Louvet, John Gliksberg, Matteo Turisini, Andrea Monterubbiano, Jean-Philippe Nomine, Denis Dutoit, Hugo Taboada, Lilia Zaourar, Mohamed Benazouz, Angelos Bilas, Fabien Chaix, Manolis Katevenis, Nikolaos Chrysos, Evangelos Mageiropoulos, Christos Kozanitis, Thomas Moen, Steffen Persvold, Einar Rustad, Sandro Fiore, Fabrizio Granelli, Simone Pezzuto, Raffaello Potestio, Luca Tubiana, Philippe Velha, Flavio Vella, Daniele De Sensi, Salvatore Pontarelli |
DSD | 30 |
| 2025 | The ExaNeSt Prototype: Evaluation of Efficient HPC Communication Hardware in an ARM-based Multi-FPGA RackabstractWe present and evaluate the ExaNeSt prototype, which compactly packages 128 Xilinx ZU9EG MPSoCs, two TBytes of DRAM, and eight TBytes of SSD into a liquid-cooled rack, using a custom interconnection hardware based on 10 GB/s links. We developed this testbed in 2016–2019 in order to leverage the flexibility of FPGAs for experimenting with efficient hardware support for HPC communication among tens of thousands of processors and accelerators in the quest toward Exascale systems and beyond. In the years since then, we carefully studied this system, and we present our key design choices and insights resulting from our measurement and analysis. We developed this testbed, from architecture to the PCBs and the run-time software, within the ExaNeSt project. It is fully operational in configurations with up to 8 × 4 × 4 MPSoC nodes. It achieves high density through tight board design, while also leveraging state-of-the-art liquid cooling technology. In this article, we present a thorough architectural analysis, along with important aspects of our infrastructure development. Our custom interconnect includes a low-cost low-latency network interface, offering user-level, zero-copy RDMA, which we coupled with the ARMv8 processors in the MPSoCs. We further developed the corresponding runtimes that allow us to test real MPI applications on the large-scale testbed. We evaluated our platform through MPI microbenchmarks, mini application, and full MPI applications. Single-hop, one-way latency is 1.3 μs; approximately 0.47 μs out of these are attributed to network interface and the user-space library that exposes its functionality to the runtime. Latency over longer paths increases as expected, reaching 2.55 μs for a five-hop path. Bandwidth tests show that, for single-hop, link utilization reaches \(82\%\) of the theoretical capacity. Microbenchmarks based on MPI collectives reveal that broadcast latency scales as expected when the number of participating ranks increases. We also implemented a custom MPI_Allreduce accelerator in the network interface, which reduces the latency of such collectives by up to \(88\%\) . We assess performance scaling through weak and strong scaling tests for HPCG, LAMMPS, and the miniFE mini application; for all these tests, parallelization efficiency is at least \(69\%\) , or better. Manolis Ploumidis, Fabien Chaix, Nikolaos Chrysos, Marios Assiminakis, Nikolaos D. Kallimanis, Nikolaos Kossifidis, Michael Nikoloudakis, Nikolaos Dimou, Michalis Gianioudis, Giorgos Ieronymakis, Aggelos Ioannou, George Kalokerinos, Pantelis Xirouchakis, Astrinos Damianakis, Michael Ligerakis, Theocharis Vavouris, Manolis Katevenis, Vassilis Papaefstathiou, Manolis Marazakis, Iakovos Mavroidis |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2024 | Low-latency Communication in RISC-V ClustersabstractLow-latency inter-node communication is important in HPC clusters. In this work, we design and integrate a low-cost interconnect, capable for low-latency user-level communication with open-source RISC-V processors, obviating the need for bulky and expensive network interface cards connected over the PCI. Our lean network interface is connected next to the Load/Store (LD/ST) stage of the RISC-V processor, which we modify to achieve back-to-back stores for the address range dedicated to the the NI. The primitives that we examine are suitable for many-to-one communication and optimized for small messages, while offering reliable delivery and hardware-level protection using protection domains. We also describe our runtime system that presents the NI to user processes with minimal overheads. Our design achieves sub-microsecond (720 ns) user-level latency for small packet generation and transmission between adjacent FPGA nodes containing Ariane RISC-V soft-cores running at 100 MHz. We also present an analytical latency breakdown including key hardware and software components. Michalis Gianioudis, Pantelis Xirouchakis, Charisios Loukas, Evangelos Mageiropoulos, Orestis Mousouros, Sokratis Mpartzis, Aggelos Ioannou, Vassilis Papaefstathiou, Manolis Katevenis, Nikolaos Chrysos |
HPC Asia | 10 |
| 2022 | RED-SEA: Network Solution for Exascale ArchitecturesabstractIn order to enable Exascale computing, next generation interconnection networks must scale to hundreds of thousands of nodes, and must provide features to also allow the HPC, HPDA, and AI applications to reach Exascale, while benefiting from new hardware and software trends. RED-SEA will pave the way to the next generation of European Exascale interconnects, including the next generation of BXI, as follows: (i) specify the new architecture using hardware-software co-design and a set of applications representative of the new terrain of converging HPC, HPDA, and AI; (ii) test, evaluate, and/or implement the new architectural features at multiple levels, according to the nature of each of them, ranging from mathematical analysis and modeling, to simulation, or to emulation or implementation on FPGA testbeds; (iii) enable seamless communication within and between resource clusters, and therefore development of a high-performance low latency gateway, bridging seamlessly with Ethernet; (iv) add efficient network resource management, thus improving congestion resiliency, virtualization, adaptive routing, collective operations; (v) open the interconnect to new kinds of applications and hardware, with enhancements for end-to-end network services - from programming models to reliability, security, low- latency, and new processors; (vi) leverage open standards and compatible APIs to develop innovative reusable libraries and Fabrics management solutions. Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Pier Stanislao Paolucci, Elena Pastorelli, Francesco Simula, Matteo Turisini, Piero Vicini, Roberto Ammendola, Pascale Bernier-Bruna, Said Derradji, Stéphane Guez, Pierre-Axel Lagadec, Gregoire Pichon, Etienne Walter, Gaetan De Gassowski, Matthieu Hautreaux, Stephane Mathieu, Gilles Moreau, Marc Pérache, Hugo Taboada, Torsten Hoefler, Timo Schneider, Matteo Barnaba, Giuseppe Piero Brandino, Francesco De Giorgi, Matteo Poggi, Iakovos Mavroidis, Ioannis Papaefstathiou, Nikolaos Tampouratzis, Benjamin Kalisch, Ulrich Krackhardt, Mondrian Nüssle, Pantelis Xirouchakis, Vangelis Mageiropoulos, Michalis Gianioudis, Harisis Loukas, Aggelos Ioannou, Nikolaos D. Kallimanis, Nikolaos Chrysos, Manolis Katevenis, Wolfgang Frings, Dominik Gottwald, Felime Guimaraes, Max Holicki, Volker Marx, Yannik Müller, Carsten Clauss, Hugo Falter, Xu Huang 0010, Jennifer Lopez Barillao, Thomas Moschny, Simon Pickartz, Francisco J. Alfaro, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José L. Sánchez 0002, Adrián Castelló 0001, Jose Duro, María Engracia Gómez, Enrique S. Quintana-Ortí, Julio Sahuquillo, Eugenio Stabile |
DSD | 44 |
| 2022 | Optimized Page Fault Handling During RDMAabstractRemote Direct Memory Access (RDMA) is widely used in High-Performance Computing (HPC) while making inroads in datacenters and accelerators. State-of-the-art RDMA engines typically do not endure page faults, therefore users are forced to pin their buffers, which complicates the programming model, limits the memory utilization, and moves the pressure to the Network Interface Cards (NICs). In this article we introduce a mechanism for handling dynamic page faults during RDMA, named PART, suitable for emerging processors that also integrate the Network Interface. PART leverages the IOMMU already present in modern processors for translations. PART avoids the pinning overheads, allows any buffer to be used for communication, and enables overlapping page fault handling with serving subsequent RDMA transfers. We implement and optimize PART for a cluster of ARMv8 cores with tightly-coupled network interfaces. Handling a minor page-fault of a small transfer at the destination takes approximately 38 $\mu$ secs, while there is no performance degradation when running three full MPI applications in 16 nodes and 64 cores. Detailed breakdown uncovers the hardware and system software components of this overhead and was used to further optimize the system. A 4MB RDMA transfer performs 1.46x better over pinning. Antonis Psistakis, Nikolaos Chrysos, Fabien Chaix, Marios Asiminakis, Michalis Gianioudis, Pantelis Xirouchakis, Vassilis Papaefstathiou, Manolis Katevenis |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2021 | Using hls4ml to Map Convolutional Neural Networks on Interconnected FPGA DevicesabstractIn this paper, we demonstrate this architecture by partitioning the SqueezeNet CNN in six (6) Ultrascale+ FPGAs and use existing tools in order to simplify the path from network definition to HLS. In particular, we use Keras in order to define arbitrary convolutional neural networks and hIs4m1 to generate HLS kernels that can be used in a Vivado HLS project. An unexpected finding is that original HLS code generated by hls4m1 is suboptimal for our purpose. Therefore, in order to achieve satisfactory results, we optimize each of its kernels appropriately by applying optimization directives available by Vivado HLS and changing the original kernel code when needed. Evangelos Mageiropoulos, Nikolaos Chrysos, Nikolaos Dimou, Manolis Katevenis |
FCCM | 2 |
| 2020 | PART: Pinning Avoidance in RDMA TechnologiesabstractState-of-the-art Remote Direct Memory Access (RDMA) engines pin communication buffers, complicating the programming model, limiting the memory utilization, and mandating a separate memory translation subsystem spanning the network interface card and the OS. In this paper, we introduce PART, a page fault handling mechanism suitable for emerging nodes that integrate the NI with the main processor. PART does not need to pin pages, thus any process buffer can be used for communication, and resolves occasional page-faults dynamically, when the network accesses the memory, by reusing the RDMA transport. Additionally, PART leverages the I/O Memory Management Unit (IOMMU) which is next to the processor in order to translate virtual to physical addresses, thus reducing cost and complexity. We implement and evaluate PART in a cluster of 16 nodes and 64 ARM cores. We evaluate the performance of transfers for varying page fault frequency, and examine optimizations that proactively page-in all pages upon the first page fault or ahead of the transfer, providing useful insights that can be used to optimize runtimes. Our results show that PART completes one-page transfers with a minor page-fault at the destination in approximately 38 μsecs, while the slowdown on 1MB transfers that experience faults in all pages is as little as 2.6x compared to the no-page-fault case. Page faults are expected to be rare in HPC setups: the performance of LAMMPS in our cluster is virtually unaffected when pages are handled dynamically using PART. Antonis Psistakis, Nikolaos Chrysos, Fabien Chaix, Marios Asiminakis, Michalis Giannioudis, Pantelis Xirouchakis, Vassilis Papaefstathiou, Manolis Katevenis |
NOCS | 2 |
| 2018 | GPU Provisioning: The 80 - 20 80 - 20 Rule
Eleni Kanellou, Nikolaos Chrysos, Stelios Mavridis, Yannis Sfakianakis, Angelos Bilas |
Euro-Par | 2 |
| 2018 | Accurate Congestion Control for RDMA TransfersabstractHigh-performance interconnects need congestion control to deal with traffic bursts. In this paper, we propose ACCurate, a congestion control protocol that assigns exact max-min fair rates to flows, without relying on costly per-flow state inside the network. ACCurate keeps the backlogs outside of the network, protects innocent flows, and promptly recovers the flows' rates after congestive episodes. Comparisons with TCP and PAUSE-only RDMA under datacenter-resembling workloads further show that ACCurate provides up to 10× faster flow completion times. ACCurate relies on simple hardware that can be readily implemented inside switches. In our implementation, the additional circuitry needed in a 16×16 switch occupies less than 2% of FPGA resources. Dimitris Giannopoulos, Nikolaos Chrysos, Evangelos Mageiropoulos, Ioannis Vardas, Leandros Tzanakis, Manolis Katevenis |
NOCS | 2 |
| 2017 | The Next Generation of Exascale-Class Systems: The ExaNeSt ProjectabstractThe ExaNeSt project started on December 2015 and is funded by EU H2020 research framework (call H2020-FETHPC-2014, n. 671553) to study the adoption of low-cost, Linux-based power-efficient 64-bit ARM processors clusters for Exascale-class systems. The ExaNeSt consortium pools partners with industrial and academic research expertise in storage, interconnects and applications that share a vision of an Euro-pean Exascale-class supercomputer. Their goal is designing and implementing a physical rack prototype together with its cooling system, the storage non-volatile memory (NVM) architecture and a low-latency interconnect able to test different options for interconnection and storage. Furthermore, the consortium is to provide real HPC applications to validate the system. Herein we provide a status report of the project initial developments. Roberto Ammendola, Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Pier Stanislao Paolucci, Elena Pastorelli, Francesco Simula, Piero Vicini, Giuliano Taffoni, Jose Antonio Pascual, Javier Navaridas, Mikel Luján, John Goodacre, Nikolaos Chrysos, Manolis Katevenis |
DSD | 17 |
| 2017 | VineTalk: Simplifying software access and sharing of FPGAs in datacentersabstractFPGA-based accelerators are becoming first class citizens in data centers. Adding FPGAs in data centers can lead to higher compute densities with improved energy efficiency for latency critical workloads, such as financial applications. However FPGA deployment in datacenters brings difficulties both to application developers, and cloud providers. Application writers need to deal with the interfacing of FPGAs on top of application logic/algorithms. On the other hand, cloud providers are reluctant face the risk that their hardware remains underutilized, due to the lack of a sharing mechanism for FPGAs. In this paper, we introduce VineTalk, a framework that reduces the programming effort associated with FPGA-based accelerators and FPGA virtualization. We integrate VineTalk with the Xilinx SDAccel development framework and we map it to the Kintex UltraScale FPGA. Our preliminary evaluation with a use-case of financial applications shows that VineTalk can offer effective FPGA sharing introducing less than 4% overhead to application execution time. Stelios Mavridis, Emmanouil Pavlidakis, Ioannis Stamoulias, Christos Kozanitis, Nikolaos Chrysos, Christoforos Kachris, Dimitrios Soudris, Angelos Bilas |
FPL | 5 |
| 2016 | The ExaNeSt Project: Interconnects, Storage, and Packaging for Exascale SystemsabstractExaNest is one of three European projects that support a ground-breaking computing architecture for exascale-class systems built upon power-efficient 64-bit ARM processors. This group of projects share an "everything-close" and "share-anything" paradigm, which trims down the power consumption -- by shortening the distance of signals for most data transfers -- as well as the cost and footprint area of the installation -- by reducing the number of devices needed to meet performance targets. In ExaNeSt, we will design and implement: (i) a physical rack prototype and its liquid-cooling subsystem providing ultra-dense compute packaging, (ii) a storage architecture with distributed (in-node) non-volatile memory (NVM) devices, (iii) a unified, low-latency interconnect, designed to efficiently uphold desired Quality-of-Service guarantees for a mix of storage with inter-processor flows, and (iv) efficient rack-level memory sharing, where each page is cacheable at only a single node. Our target is to test alternative storage and interconnect options on actual hardware, using real-world HPC applications. The ExaNeSt consortium brings together technology, skills, and knowledge across the entire value chain, from computing IP, packaging, and system deployment, all the way up to operating systems, storage, HPC, big data frameworks, and cutting-edge applications. Manolis Katevenis, Nikolaos Chrysos, Manolis Marazakis, Iakovos Mavroidis, Fabien Chaix, Nikolaos D. Kallimanis, Javier Navaridas, John Goodacre, Piero Vicini, Andrea Biagioni, Pier Stanislao Paolucci, Alessandro Lonardo, Elena Pastorelli, Francesca Lo Cicero, Roberto Ammendola, P. Hopton, P. Coates, Giuliano Taffoni, Stefano Cozzini, Martin L. Kersten, Julio Sahuquillo, Sergio Lechago, C. Pinto, Bernd Lietzow, D. Everett, Gino Perna |
DSD | 2 |
| 2016 | Discharging the Network From Its Flow Control Headaches: Packet Drops and HOL BlockingabstractCongestion control becomes indispensable in highly utilized consolidated networks running demanding applications. In this paper, proactive congestion management schemes for Clos networks are described and evaluated. The key idea is to move the congestion avoidance burden from the data fabric to a scheduling network, which isolates flows using per-flow request counters. The scheduling network comprises per-output arbiters that grant data packets after reserving space for them in the buffer memories in front of fabric outputs. Computer simulations show that this strategy eliminates head-of-line (HOL) blocking and its adversarial effects throughout the fabric, without having to drop packets. In particular, a simplified model describes this result as a synergy between proactive buffer reservations and fine-grained multipath routing. Two alternative designs are presented. The first one places all arbiters in a central control unit, is simpler, and has superior performance. The second is more scalable by distributing the arbiters over the switching elements of the Clos network and by routing the control messages to and from endpoint adapters via multiple paths. Computer simulations of the complete system demonstrate high throughput and low latency under any number of congested outputs. Weighted max-min fair allocation of fabric-output link bandwidth is also demonstrated. Furthermore, delay breakdowns show that the time that packets wait in fabric and resequencing buffers is minimized as a result of the reduced (and equalized across all fabric paths) in-fabric contention. Finally, the high throughput capability of the system is corroborated by a Markov chain analysis of output buffer credits. Nikolaos Chrysos, Lydia Y. Chen, Christoforos Kachris, Manolis Katevenis |
IEEE/ACM Trans. Netw. | 1 |
| 2015 | A Systematic Evaluation of Emerging Mesh-like CMP NoCsabstractThis paper studies alternative Network-on-Chip architectures for emerging many-core chip multiprocessors, by exploring the following design options on mesh-based networks: Multiple physical networks (P), cores concentration (C), express channels (X), it widths (W), and virtual channels (V). We exhaustively evaluate all combinations of the afore-mentioned parameters (P, C, X, W, V), using the energy-throughput ratio (ETR) as a metric to classify network congurations. Our experimental results show that, on one hand, with an appropriate selection of parameters (V,W), an optimized baseline 2D mesh offers the best possible ETR for NoCs with up to a few tens of cores (64-core NoC). More complicated networks, using concentration and express channels, can reduce the zero-load latency, but do not necessarily help to improve ETR. On the other hand, for larger CMPs, a 2D mesh with multiple physical networks is a better option: once optimized, this architectural choice can reduce the ETR by up to 46% for 256 cores. Antonis Psathakis, Vassilis Papaefstathiou, Nikolaos Chrysos, Fabien Chaix, Evangelos Vasilakis, Dionisios N. Pnevmatikatos, Manolis Katevenis |
ANCS | 3 |
| 2015 | SCOC: High-radix switches made of bufferless clos networksabstractIn today's datacenters handling big data and for exascale computers of tomorrow, there is a pressing need for high-radix switches to economically and efficiently unify the computing and storage resources that are dispersed across multiple racks. In this paper, we present SCOC, a switch architecture suitable for economical IC implementation that can efficiently replace crossbars for high-radix switch nodes. SCOC is a multi-stage bufferless network with O(N2/m) cost, where m is a design parameter, practically ranging between 4-16. We identify and resolve more than five fairness violations that are pertinent to hierarchical scheduling. Effectively, from a performance perspective, SCOC is indistinguishable from efficient flat crossbars. Computer simulations show that it competes well or even outperforms flat crossbars and hierarchical switches. We report data from our ASIC implementation at 32 nm of a SCOC 136×136 switch, with shallow buffers, connecting 25 Gb/s links. In this first incarnation, SCOC is used at the spines of a server-rack, fat-tree network. Internally, it runs at 9.9 Tb/s, thus offering a speedup of 1.45 ×, and provides a fall-through latency of just 61 ns. Nikolaos Chrysos, Cyriel Minkenberg, Mark Rudquist, Claude Basso, Brian Vanderpool |
HPCA | 1 |
| 2015 | Large switches or blocking multi-stage networks? An evaluation of routing strategies for datacenter fabrics
Nikolaos Chrysos, Fredy D. Neeser, Mitchell Gusat, Cyriel Minkenberg, Wolfgang E. Denzel, Claude Basso, Mark Rudquist, Kenneth M. Valk, Brian Vanderpool |
Comput. Networks | 1 |
| 2014 | Integration and QoS of multicast traffic in a server-rack fabric with 640 100g portsabstractFlexible datacenters rely on high-bandwidth server-rack fabrics to allocate their distributed computing and storage resources anywhere,anyhow, and anytime demanded. We describe the multicast architecture of a distributed server-rack fabric, which is arranged around a spine-leaf topology and connects 640 Ethernet ports running at 100G. To cope with the immense fabric speed, we resort to hierarchical, tree-based replication, facilitated by specially commissioned fabric-end ports. At each (port-to-port) leg of the tree, a frame copy is forwarded after a request-grant admission phase and is ACKed by the receiver. To save on bandwidth, we use a packet cache in our input-queued switching-nodes, which replicates asynchronously forwarded frames thus tolerating the variable-delay in the admission phase. Because the cache has limited size, we loosely synchronize the multicast subflows to protect the cache from thrashing. We describe our policies for lossy classes, which segregate and provide fair treatment to multicast subflows. Finally, we show that industry-standard Level2 congestion control does not adapt well to one-to-many flows, and demonstrate that the methods that we implement achieve the best performance. Nikolaos Chrysos, Fredy D. Neeser, Brian Vanderpool, Mark Rudquist, Kenneth M. Valk, Todd Greenfield, Claude Basso |
ANCS | 1 |
| 2014 | zFabric: How to virtualize lossless ethernet?abstractConverged Enhanced Ethernet (CEE) is a crucial step in embracing storage, cluster, and high-performance computing fabrics under a common network. However, the adoption of lossless CEE in virtualized clusters is hindered by the lack of network hypervisor software that addresses the major issues of losslessness, i.e., head-of-line blocking and saturation trees. Our objective is to design a hypervisor that prevents miscon-figured or malicious virtual machines from filling the lossless network with stalled packets, thus compromising tenant isolation. Furthermore, we observe that current hypervisors perform compulsory isolation, management, and mobility functions, but introduce new bottlenecks on the data-path. By taking advantage of the lossless fabric, we deconstruct the existing virtualized networking stack into its core functions and consolidate them into zFabric, an efficient hypervisor that meets our aforementioned goals. To demonstrate zFabric's benefits, we evaluate a prototype implementation on a datacenter testbed. Besides resolving head-of-line blocking, zFabric improves throughputs for long flows by up to 56%, lowers CPU utilization by up to 63%, and shortens completion times by up to 7x for partition-aggregate queries when compared with current virtualized TCP stacks. Daniel Crisan, Robert Birke, Nikolaos Chrysos, Cyriel Minkenberg, Mitchell Gusat |
CLUSTER | 3 |
| 2014 | High performance multipath routing for datacentersabstractPerformance-optimized datacenter networks aim to handle more efficiently the growing East-West intra-cluster traffic of BigData applications. The demanding latency constraints and traffic patterns of these applications expose the inherent bottlenecks of the often oversubscribed datacenter network topologies, favoring in stead the full-bisectional bandwidth fat-trees. And yet their topological benefits may remain unrealized in practical deployments, if such fabrics use single path or flow-level (ECMP hashing) multipath routing. Here we model in detail on Layer 2 the routing performance of modern fat-tree networks using stochastic permutations of bursty traffic. We first analytically simplify and then validate by accurate simulation models that the throughputs for ‘static’ d-mod-k and for ECMP-like multipath routing are 63% and 47%, respectively. We also find that ECMP routing results in a wide spread of link loads under random permutation traffic, which manifests as a 3x throughput reduction for 30% of the flows. Furthermore, ECMP can lead to collisions of mouse and elephant flows, often increasing the flow completion time (FCT) of delay-sensitive flows by a factor of 10. In contrast, packet-based multipath outperforms all the others in this study. Nikolaos Chrysos, Mitchell Gusat, Fredy D. Neeser, Cyriel Minkenberg, Wolfgang E. Denzel, Claude Basso |
HPSR | 1 |
| 2011 | Distributed WFQ scheduling converging to weighted max-min fairness
Nikolaos Chrysos, Manolis Katevenis |
Comput. Networks | 1 |
| 2010 | Throughput of random arbitration for approximate matchingsabstractModern switches and switching fabrics typically employ virtual output queues (VOQs) at network adapters, in order to mitigate head-of-line blocking. The core of the switch, often a crossbar, can either be bufferless or buffered. Previous research on random arbitration for bufferless crossbars found that the throughput for large switch sizes is bounded at 63% [1]. Although crossbars containing crosspoint buffers are shown to reach 100% throughput under uniform traffic, their scalability is restricted by the quadratic growth of the number of crosspoint buffers, and the size of each, which depends on the intra-switch (VOQ-crossbar) round trip time (RTT). In this paper, we study an alternative architecture [2], which can reduce buffer requirements, by (a) having a linear growth of the number of buffers, and (b) making the size of each independent of the RTT. As the scheduler in such an architecture may match multiple inputs to the same output at the same time, we refer to such matchings as approximate, as opposed to the exact matchings required in bufferless crossbars. Lydia Y. Chen, Nikolaos Chrysos |
ANCS | 2 |
| 2010 | End-to-end congestion management for non-blocking multi-stage switching fabricsabstractPacket-switched networks are encountered at the heart of scalable network routers and high-performance computer (or data center) interconnects. As these networks scale to larger port counts, and their utilization increases, congestion management becomes indispensable. At the same time, technology constraints rule out monolithic bufferless switches with centralized schedulers, and impose buffered multi-stage switching fabrics with distributed control. These trends have for some time now called forth research and products [1, 2, 3, 4, 5], which applied the request-grant philosophy of bufferless crossbars to make buffered multi-stage switching fabrics practical and efficient. In such proactive schemes, saturation-tree congestion is avoided by having inputs inform outputs of their demand, and inject data only after receiving output permission (grants). The per output admission arbiters can be located in a central scheduling unit, as in [3] [2] [4], can be distributed across the edge switches of the fabric [4], or can be placed in the respective output adapters [1] [5]. Such schemes have been quite successful, especially because they can be combined conveniently with reorder/reassembly buffer management, as well as with end-to-end reliable-delivery schemes. Nikolaos Chrysos, Lydia Y. Chen, Cyriel Minkenberg, Christoforos Kachris, Manolis Katevenis |
ANCS | 1 |
| 2010 | Towards low-cost high-performance all-optical interconnection networksabstractWe consider low-complexity, all-optical multi-stage networks with distributed arbitration achieved through minimal per-node buffering. To enhance the saturation throughput of such networks while maintaining low latency at low loads, we examine a novel combination of deterministic (prescheduled) and speculative (eager) packet injections. Prescheduled injections are performed in a time-division-multiplexing (TDM) manner, whereas eager injections follow a packet multiplexing paradigm. Prescheduled injections aim at reducing contention in the fabric, and sustain network throughput when the load is high. Eager injections, on the other hand, ignore the TDM schedule, thus allowing low latency communication when contention is low. In our set of rules that govern the interaction between prescheduled and eager packets, eager packets can be intentionally dropped when they block the progress of prescheduled ones. At the same time a highly efficient end-to-end reliable delivery scheme, implemented at the host adapters, deals with random packet losses in the optical domain, and also recovers the dropped eager packets. Computer simulations demonstrate that our approach can render low-complexity networks, which are amenable to an all-optical implementation, attractive for use in computer interconnects. Nikolaos Chrysos, Cyriel Minkenberg, Jens Hofrichter, Folkert Horst, Bert J. Offrein |
HPSR | 1 |
| 2008 | Fast arbiters for on-chip network switchesabstractThe need for efficient implementation of simple crossbar schedulers has increased in the recent years due to the advent of on-chip interconnection networks that require low latency message delivery. The core function of any crossbar scheduler is arbitration that resolves conflicting requests for the same output. Since, the delay of the arbiters directly determine the operation speed of the scheduler, the design of faster arbiters is of paramount importance. In this paper, we present a new bit-level algorithm and new circuit techniques for the design of programmable priority arbiters that offer significantly more efficient implementations compared to already-known solutions. From the experimental results it is derived that the proposed circuits are more than 15% faster than the most efficient previous implementations, which under equal delay comparisons, translates to 40% less energy. Giorgos Dimitrakopoulos, Nikolaos Chrysos, Costas Galanopoulos |
ICCD | 2 |
| 2007 | Congestion management for non-blocking clos networksabstractWe propose a distributed congestion management scheme for non-blocking, 3-stage Clos networks, comprising plain buffered crossbar switches. VOQ requests are routed using multipath routing to the switching elements of the 3rd-stage, and grants travel back to the linecards the other way around. The fabric elements contain independent single-resource schedulers, that serve requests and grants in a pipeline. As any other network with limited capacity, this scheduling network may suffer from oversubscribed links, hotspot contention, etc., which we identify and tackle. We also reduce the cost of internal buffers, by reducing the data RTT, and by allowing sub-RTT crosspoint buffers. Performance simulations demonstrate that, with almost all outputs congested, packets destined to non-congested outputs experience very low delays ( ow isolation). For applications requiring very low communication delays, we propose a second, parallel operation mode, wherein linecards can forward a few packets eagerly, each, bypassing the request-grant latency overhead. Nikolaos Chrysos |
ANCS | 1 |
| 2007 | A buffered crossbar-based chip interconnection framework supporting quality of serviceabstractAs Systems-on-a-Chip (SoCs) become larger, the problem of interconnecting the various subsystems becomes more complicated. In this framework, certain alternatives to the standard buses, based on Network Technologies, have emerged as innovative approaches for SoC's interconnect. One of the main advantages of such an alternative, is that it can offer certain Quality of Service (QoS) over the internal cross-connects while at the same time it supports higher transfer rates than the existing on-chip buses. This paper presents the first chip interconnection architecture, which is based on a buffered crossbar switch. The main advantage of the proposed system is that it efficiently supports different priority levels; it also provides several Gigabits per Second of aggregate bandwidth, while it introduces very low latency. Moreover, the hardware complexity of this highly scalable scheme is minimal. All those facts make this framework ideal for SoCs that contain IP cores with diverse speed/throughput requirements. Ioannis Papaefstathiou, Nikolaos Chrysos |
ACM Great Lakes Symposium on VLSI | 3 |
| 2007 | Performance evaluation of the Data Vortex photonic switchabstractThe data vortex photonic packet-switching architecture features an all-optical transparent data path, highly distributed control, low latency, and a high degree of scalability. These characteristics make it attractive as a routing fabric in future photonic packet switches. We analyze the performance of the data vortex architecture as a function of its height and angle dimensions, H and A. The investigation is based on two performance measures: the average delay and the maximum throughput of the switch. We present an analytical model assuming uniform traffic and derive closed-form expressions for these measures. Our results obtained demonstrate that as H increases, the saturation throughput decreases and approaches2/9= 0.22 when A is small and H is large. Furthermore, for fixed switch size, the saturation throughput is maximized when A is minimal. We also present simulation results for the maximum throughput under uniform and nonuniform traffic, as well as for the mean number of hops and the mean input-queue packet delay as a function of input load, and address the issue of resequencing delay. The results obtained advocate that to support more ports, it is preferable to increase the height dimension and to keep the angle dimension as small as possible. Ilias Iliadis, Nikolaos Chrysos, Cyriel Minkenberg |
IEEE J. Sel. Areas Commun. | 2 |
| 2006 | Scheduling in Non-Blocking Buffered Three-Stage Switching FabricsabstractAbstract — Three-stage non-blocking switching fabrics are the next step in scaling current crossbar switches to many hundreds or few thousands of ports. Congestion management, however, is the central open problem; without it, performance suffers heavily under real-world traffic patterns. Schedulers for bufferless crossbars perform congestion management but are not scalable to high valencies and to multi-stage fabrics. Distributed scheduling, as used in buffered crossbars, is scalable but has never been scaled beyond crossbar valencies. We combine ideas from central and distributed schedulers, from request-grant protocols and from credit-based flow control, to propose a novel, practical architecture for scheduling in non-blocking buffered switching fabrics. The new architecture relies on multiple, independent, single-resource schedulers, operating in a pipeline. It: (i) isolates well-behaved against congested flows; (ii) provides throughput in excess of 95 % under unbalanced traffic, and delays that successfully compete again output queueing; (iii) provides weighted max-min fairness; (iv) directly operates on variable-size packets or multi-packet segments; (v) resequences cells or segments using very small buffers; and (vi) can be realistically implemented for a 1024×1024 reference fabric made out of 32×32 buffered crossbar switch elements. This paper carefully studies the many intricacies of the problem and the solution, discusses implementation, and provides performance simulation results. 1 Nikolaos Chrysos, Manolis Katevenis |
INFOCOM | 1 |
| 2005 | Scheduling in switches with small internal buffersabstractUnbuffered crossbars or switching fabrics contain no internal buffers, and function using only input (VOQ) and possibly output queues. Schedulers for such switches are complex, and introduce increased delay at medium loads, because they have to admit at most one cell per input and per output, during each time slot. Buffered crossbars, on the other hand, contain Q sufficient internal buffering (N2buffers) to allow independent schedulers to concurrently forward packets to the same output from any number of inputs. These architectures represent the two extremes in a range of solutions, which we examine here; although intermediate points in this range are of reduced practical interest for crossbars, they are nevertheless quite interesting for switching fabrics, and they may be of interest for optical switches. We find that tolerating two cells per-output per time-slot, using small buffers inside the switch or fabric, suffices for independent and efficient scheduling. First, we introduce a novel "request-grant" credit protocol, enabling N inputs to share a small switch buffer. Then, we apply this protocol to a switch with N such buffers, one per output, and we consider the resulting scheduling problem. Interestingly, this looks like unbuffered crossbar schedulers, but it is much simpler because it comprises independent schedulers that can be pipelined. We show that individual buffer sizes do not need to grow, neither with switch size nor with propagation delay. Through simulations, we study performance as a function of the number of cells allowed per-output per-time-slot. For one cell, the switch performs very close to the iSLIP unbuffered crossbar with one iteration. For more cells, performance improves quickly; for 12 cells, packet delay under (smooth) uniform load is practically as low as ideal output queueing. Under unbalanced load, throughput is superior to buffered crossbars, due to better buffer sharing Nikolaos Chrysos, Manolis Katevenis |
GLOBECOM | 1 |
| 2004 | Multiple priorities in a two-lane buffered crossbarabstractA significant advantage of buffered crossbar switches is that they can directly operate on variable-size packets. However, in order to support multiple priority levels, separate queues per priority are needed at each crosspoint, in order to prevent HOL blocking and buffer hogging; these queues are expensive because they each need a size of at least one maximum-size packet. In this paper, we propose a scheme that uses only two queues per crosspoint to effectively support multiple priorities. We adaptively adjust the priority levels of the two queues so that most traffic goes through the "lower" queue, while the "upper" queue remains usually available for higher priority packets to overtake the former. Through simulation, and assuming 8 priority levels, we compare our scheme to an ideal system that uses 8 queues per crosspoint. For realistic traffic, the two systems perform almost identically, although ours uses 4 times less memory in the crossbar. Even under a highly irregular traffic pattern, Bursts60, our system does not increase the average delay of any priority level by more than 75% compared to the ideal system. Nikolaos Chrysos, Manolis Katevenis |
GLOBECOM | 1 |
| 2004 | Variable packet size buffered crossbar (CICQ) switchesabstractOne of the most widely used architectures for packet switches is the crossbar. A special version of it is the buffered crossbar, where small buffers are associated with the crosspoints; this simplifies scheduling and improves its efficiency and QoS capabilities to the point where the switch needs no internal speedup. Furthermore, by supporting variable length packets throughout a buffered crossbar: (a) there is no need for segmentation and reassembly (SAR) circuits; (b) no speedup is necessary to support SAR; and (c) synchronization between the input and output clock domains is simplified. In turn, the lack of SAR and speedup mean that no output queues are needed, either. In this paper we present an architecture, a chip layout and cost analysis, and a performance evaluation of such a 300 Gbps buffered crossbar operating on variable-size packets. The proposed organization is simple yet powerful, it can be implemented using modern technology, and, as the performance results demonstrate, it clearly outperforms unbuffered crossbars. Manolis Katevenis, Giorgos Passas, Dimitrios Simos, Ioannis Papaefstathiou, Nikolaos Chrysos |
ICC | 5 |