EDBT 2026 Demo / reviewers in the wild / expert
Manolis Katevenis
dblp:k/ManolisKatevenis · also Manolis G. H. Katevenis
· DBLP profile ↗
49ranked-venue papers
10as first author
6since 2021 · last 2025
0009-0008-5437-4709ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 28 · 4 first-author · 5 since 2021Computer networks · 18 · 6 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NET4EXA: Pioneering the Future of Interconnects for Supercomputing and AIabstractNET4EXA aims to develop a next-generation high-performance interconnect for HPC and AI systems, addressing the increasing demands of large-scale infrastructures, such as those required for training Large Language Models. Building upon the proven BXI (Bull eXascale Interconnect) European technology used in TOP15 supercomputers, NET4EXA will deliver the new BXI release, BXIv3, a complete hardware and software interconnect solution, including switch and network interface components. The project will integrate a fully functional pilot system at TRL 8, ready for deployment into upcoming exascale and post-exascale systems from 2025 onward. Leveraging prior research from European initiatives like RED-SEA, the previous achievements of consortium partners and over 20 years of expertise from BULL, NET4EXA also lays the groundwork for the future generation of BXI, BXIv4, providing analysis and preliminary design. The project will use a hybrid development and co-design approach, combining commercial switch technology with custom IP and FPGA-based NICs. Performances of NET4EXA BXIv3 interconnect will be evaluated using a broad portfolio of benchmarks, scientific scalable applications, and AI workloads. Michele Martinelli, Roberto Ammendola, Andrea Biagioni, Carlotta Chiarini, Ottorino Frezza, Francesca Lo Cicero, Alessandro Lonardo, Pier Stanislao Paolucci, Elena Pastorelli, Pierpaolo Perticaroli, Luca Pontisso, Cristian Rossi, Francesco Simula, Piero Vicini, David Colin, Gregoire Pichon, Alexandre Louvet, John Gliksberg, Matteo Turisini, Andrea Monterubbiano, Jean-Philippe Nomine, Denis Dutoit, Hugo Taboada, Lilia Zaourar, Mohamed Benazouz, Angelos Bilas, Fabien Chaix, Manolis Katevenis, Nikolaos Chrysos, Evangelos Mageiropoulos, Christos Kozanitis, Thomas Moen, Steffen Persvold, Einar Rustad, Sandro Fiore, Fabrizio Granelli, Simone Pezzuto, Raffaello Potestio, Luca Tubiana, Philippe Velha, Flavio Vella, Daniele De Sensi, Salvatore Pontarelli |
DSD | 29 |
| 2025 | The ExaNeSt Prototype: Evaluation of Efficient HPC Communication Hardware in an ARM-based Multi-FPGA RackabstractWe present and evaluate the ExaNeSt prototype, which compactly packages 128 Xilinx ZU9EG MPSoCs, two TBytes of DRAM, and eight TBytes of SSD into a liquid-cooled rack, using a custom interconnection hardware based on 10 GB/s links. We developed this testbed in 2016–2019 in order to leverage the flexibility of FPGAs for experimenting with efficient hardware support for HPC communication among tens of thousands of processors and accelerators in the quest toward Exascale systems and beyond. In the years since then, we carefully studied this system, and we present our key design choices and insights resulting from our measurement and analysis. We developed this testbed, from architecture to the PCBs and the run-time software, within the ExaNeSt project. It is fully operational in configurations with up to 8 × 4 × 4 MPSoC nodes. It achieves high density through tight board design, while also leveraging state-of-the-art liquid cooling technology. In this article, we present a thorough architectural analysis, along with important aspects of our infrastructure development. Our custom interconnect includes a low-cost low-latency network interface, offering user-level, zero-copy RDMA, which we coupled with the ARMv8 processors in the MPSoCs. We further developed the corresponding runtimes that allow us to test real MPI applications on the large-scale testbed. We evaluated our platform through MPI microbenchmarks, mini application, and full MPI applications. Single-hop, one-way latency is 1.3 μs; approximately 0.47 μs out of these are attributed to network interface and the user-space library that exposes its functionality to the runtime. Latency over longer paths increases as expected, reaching 2.55 μs for a five-hop path. Bandwidth tests show that, for single-hop, link utilization reaches \(82\%\) of the theoretical capacity. Microbenchmarks based on MPI collectives reveal that broadcast latency scales as expected when the number of participating ranks increases. We also implemented a custom MPI_Allreduce accelerator in the network interface, which reduces the latency of such collectives by up to \(88\%\) . We assess performance scaling through weak and strong scaling tests for HPCG, LAMMPS, and the miniFE mini application; for all these tests, parallelization efficiency is at least \(69\%\) , or better. Manolis Ploumidis, Fabien Chaix, Nikolaos Chrysos, Marios Assiminakis, Nikolaos D. Kallimanis, Nikolaos Kossifidis, Michael Nikoloudakis, Nikolaos Dimou, Michalis Gianioudis, Giorgos Ieronymakis, Aggelos Ioannou, George Kalokerinos, Pantelis Xirouchakis, Astrinos Damianakis, Michael Ligerakis, Theocharis Vavouris, Manolis Katevenis, Vassilis Papaefstathiou, Manolis Marazakis, Iakovos Mavroidis |
ACM Trans. Reconfigurable Technol. Syst. | 17 |
| 2024 | Low-latency Communication in RISC-V ClustersabstractLow-latency inter-node communication is important in HPC clusters. In this work, we design and integrate a low-cost interconnect, capable for low-latency user-level communication with open-source RISC-V processors, obviating the need for bulky and expensive network interface cards connected over the PCI. Our lean network interface is connected next to the Load/Store (LD/ST) stage of the RISC-V processor, which we modify to achieve back-to-back stores for the address range dedicated to the the NI. The primitives that we examine are suitable for many-to-one communication and optimized for small messages, while offering reliable delivery and hardware-level protection using protection domains. We also describe our runtime system that presents the NI to user processes with minimal overheads. Our design achieves sub-microsecond (720 ns) user-level latency for small packet generation and transmission between adjacent FPGA nodes containing Ariane RISC-V soft-cores running at 100 MHz. We also present an analytical latency breakdown including key hardware and software components. Michalis Gianioudis, Pantelis Xirouchakis, Charisios Loukas, Evangelos Mageiropoulos, Orestis Mousouros, Sokratis Mpartzis, Aggelos Ioannou, Vassilis Papaefstathiou, Manolis Katevenis, Nikolaos Chrysos |
HPC Asia | 9 |
| 2022 | RED-SEA: Network Solution for Exascale ArchitecturesabstractIn order to enable Exascale computing, next generation interconnection networks must scale to hundreds of thousands of nodes, and must provide features to also allow the HPC, HPDA, and AI applications to reach Exascale, while benefiting from new hardware and software trends. RED-SEA will pave the way to the next generation of European Exascale interconnects, including the next generation of BXI, as follows: (i) specify the new architecture using hardware-software co-design and a set of applications representative of the new terrain of converging HPC, HPDA, and AI; (ii) test, evaluate, and/or implement the new architectural features at multiple levels, according to the nature of each of them, ranging from mathematical analysis and modeling, to simulation, or to emulation or implementation on FPGA testbeds; (iii) enable seamless communication within and between resource clusters, and therefore development of a high-performance low latency gateway, bridging seamlessly with Ethernet; (iv) add efficient network resource management, thus improving congestion resiliency, virtualization, adaptive routing, collective operations; (v) open the interconnect to new kinds of applications and hardware, with enhancements for end-to-end network services - from programming models to reliability, security, low- latency, and new processors; (vi) leverage open standards and compatible APIs to develop innovative reusable libraries and Fabrics management solutions. Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Pier Stanislao Paolucci, Elena Pastorelli, Francesco Simula, Matteo Turisini, Piero Vicini, Roberto Ammendola, Pascale Bernier-Bruna, Said Derradji, Stéphane Guez, Pierre-Axel Lagadec, Gregoire Pichon, Etienne Walter, Gaetan De Gassowski, Matthieu Hautreaux, Stephane Mathieu, Gilles Moreau, Marc Pérache, Hugo Taboada, Torsten Hoefler, Timo Schneider, Matteo Barnaba, Giuseppe Piero Brandino, Francesco De Giorgi, Matteo Poggi, Iakovos Mavroidis, Ioannis Papaefstathiou, Nikolaos Tampouratzis, Benjamin Kalisch, Ulrich Krackhardt, Mondrian Nüssle, Pantelis Xirouchakis, Vangelis Mageiropoulos, Michalis Gianioudis, Harisis Loukas, Aggelos Ioannou, Nikolaos D. Kallimanis, Nikolaos Chrysos, Manolis Katevenis, Wolfgang Frings, Dominik Gottwald, Felime Guimaraes, Max Holicki, Volker Marx, Yannik Müller, Carsten Clauss, Hugo Falter, Xu Huang 0010, Jennifer Lopez Barillao, Thomas Moschny, Simon Pickartz, Francisco J. Alfaro, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José L. Sánchez 0002, Adrián Castelló 0001, Jose Duro, María Engracia Gómez, Enrique S. Quintana-Ortí, Julio Sahuquillo, Eugenio Stabile |
DSD | 45 |
| 2022 | Optimized Page Fault Handling During RDMAabstractRemote Direct Memory Access (RDMA) is widely used in High-Performance Computing (HPC) while making inroads in datacenters and accelerators. State-of-the-art RDMA engines typically do not endure page faults, therefore users are forced to pin their buffers, which complicates the programming model, limits the memory utilization, and moves the pressure to the Network Interface Cards (NICs). In this article we introduce a mechanism for handling dynamic page faults during RDMA, named PART, suitable for emerging processors that also integrate the Network Interface. PART leverages the IOMMU already present in modern processors for translations. PART avoids the pinning overheads, allows any buffer to be used for communication, and enables overlapping page fault handling with serving subsequent RDMA transfers. We implement and optimize PART for a cluster of ARMv8 cores with tightly-coupled network interfaces. Handling a minor page-fault of a small transfer at the destination takes approximately 38 $\mu$ secs, while there is no performance degradation when running three full MPI applications in 16 nodes and 64 cores. Detailed breakdown uncovers the hardware and system software components of this overhead and was used to further optimize the system. A 4MB RDMA transfer performs 1.46x better over pinning. Antonis Psistakis, Nikolaos Chrysos, Fabien Chaix, Marios Asiminakis, Michalis Gianioudis, Pantelis Xirouchakis, Vassilis Papaefstathiou, Manolis Katevenis |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2021 | Using hls4ml to Map Convolutional Neural Networks on Interconnected FPGA DevicesabstractIn this paper, we demonstrate this architecture by partitioning the SqueezeNet CNN in six (6) Ultrascale+ FPGAs and use existing tools in order to simplify the path from network definition to HLS. In particular, we use Keras in order to define arbitrary convolutional neural networks and hIs4m1 to generate HLS kernels that can be used in a Vivado HLS project. An unexpected finding is that original HLS code generated by hls4m1 is suboptimal for our purpose. Therefore, in order to achieve satisfactory results, we optimize each of its kernels appropriately by applying optimization directives available by Vivado HLS and changing the original kernel code when needed. Evangelos Mageiropoulos, Nikolaos Chrysos, Nikolaos Dimou, Manolis Katevenis |
FCCM | 4 |
| 2020 | PART: Pinning Avoidance in RDMA TechnologiesabstractState-of-the-art Remote Direct Memory Access (RDMA) engines pin communication buffers, complicating the programming model, limiting the memory utilization, and mandating a separate memory translation subsystem spanning the network interface card and the OS. In this paper, we introduce PART, a page fault handling mechanism suitable for emerging nodes that integrate the NI with the main processor. PART does not need to pin pages, thus any process buffer can be used for communication, and resolves occasional page-faults dynamically, when the network accesses the memory, by reusing the RDMA transport. Additionally, PART leverages the I/O Memory Management Unit (IOMMU) which is next to the processor in order to translate virtual to physical addresses, thus reducing cost and complexity. We implement and evaluate PART in a cluster of 16 nodes and 64 ARM cores. We evaluate the performance of transfers for varying page fault frequency, and examine optimizations that proactively page-in all pages upon the first page fault or ahead of the transfer, providing useful insights that can be used to optimize runtimes. Our results show that PART completes one-page transfers with a minor page-fault at the destination in approximately 38 μsecs, while the slowdown on 1MB transfers that experience faults in all pages is as little as 2.6x compared to the no-page-fault case. Page faults are expected to be rare in HPC setups: the performance of LAMMPS in our cluster is virtually unaffected when pages are handled dynamically using PART. Antonis Psistakis, Nikolaos Chrysos, Fabien Chaix, Marios Asiminakis, Michalis Giannioudis, Pantelis Xirouchakis, Vassilis Papaefstathiou, Manolis Katevenis |
NOCS | 8 |
| 2019 | Towards Exascale: Measuring the Energy Footprint of Astrophysics HPC SimulationsabstractThe increasing amount of data produced in Astronomy by observational studies and the size of theoretical problems to be tackled in the next future pushes the need of HPC (High Performance Computing) resources towards the "Exascale". The HPC sector is undergoing a profound phase of transition, in which one of the toughest challenges to cope with is the energy efficiency that is one of the main blocking factors to the achievement of "Exascale". Since ideal peak-performance is unlikely to be achieved in realistic scenarios, the aim of this work is to give some insights about the energy consumption of contemporary architectures with real scientific applications in a HPC context. We use two state-of-the-art applications from the astrophysical domain, that we optimized in order to fully exploit the underlying hardware: a direct N-body code and a semi-analytical code for Cosmic Structure formation simulations. For these two applications, we quantitatively evaluate the impact of computation on the energy consumption when running on three different systems: one that represents the present of current HPC systems (an Intel-based cluster), one that (possibly) represents the future of HPC systems (a prototype of an Exascale supercomputer) and a micro-cluster based on Arm MPSoC. We provide a comparison of the time-to-solution, energy-to-solution and energy delay product (EDP) metrics, for different software configurations. ARM-based HPC systems have lower energy consumption albeit running ≈10 times slower. Giuliano Taffoni, Manolis Katevenis, Renato Panchieri, Gino Perna, Luca Tornatore, David Goz, Antonio Ragagnin, Sara Bertocco, Igor Coretti, Manolis Marazakis, Fabien Chaix, Manolis Ploumidis |
eScience | 2 |
| 2019 | Efficient Convolutional Neural Network Weight Compression for Space Data Classification on Multi-fpga PlatformsabstractConvolutional Neural Networks (CNNs) represent the cutting edge in signal analysis tasks like classification and regression. Realization of such architectures in hardware capable of performing high throughput computations, with minimal energy consumption, is a key enabling factor towards the proliferation of analysis immediately after acquisition. Our driving problem is a satellite-based remote sensing platform in which onboard signal processing and classification tasks must take place, given strict bandwidth and energy limitations. In this work, we exploit the implementation of a CNN on Field Programmable Gate Array (FPGA) platforms and explore different ways to minimize the impact of different hardware restrictions to performance. We compare our results against competing technologies such as Graphics Processing Units (GPU) in terms of throughput, latency and energy consumption. In actual experimental runs we demonstrate competitive latency and throughput of the FPGA platform vs. GPU technology at an order-of-magnitude energy savings, which is especially important for space-borne computing. George Pitsis, Grigorios Tsagkatakis, Christos Kozanitis, Ioannis Kalomoiris, Aggelos Ioannou, Apostolos Dollas, Manolis Katevenis, Panagiotis Tsakalides |
ICASSP | 7 |
| 2018 | Accurate Congestion Control for RDMA TransfersabstractHigh-performance interconnects need congestion control to deal with traffic bursts. In this paper, we propose ACCurate, a congestion control protocol that assigns exact max-min fair rates to flows, without relying on costly per-flow state inside the network. ACCurate keeps the backlogs outside of the network, protects innocent flows, and promptly recovers the flows' rates after congestive episodes. Comparisons with TCP and PAUSE-only RDMA under datacenter-resembling workloads further show that ACCurate provides up to 10× faster flow completion times. ACCurate relies on simple hardware that can be readily implemented inside switches. In our implementation, the additional circuitry needed in a 16×16 switch occupies less than 2% of FPGA resources. Dimitris Giannopoulos, Nikolaos Chrysos, Evangelos Mageiropoulos, Ioannis Vardas, Leandros Tzanakis, Manolis Katevenis |
NOCS | 6 |
| 2017 | The Next Generation of Exascale-Class Systems: The ExaNeSt ProjectabstractThe ExaNeSt project started on December 2015 and is funded by EU H2020 research framework (call H2020-FETHPC-2014, n. 671553) to study the adoption of low-cost, Linux-based power-efficient 64-bit ARM processors clusters for Exascale-class systems. The ExaNeSt consortium pools partners with industrial and academic research expertise in storage, interconnects and applications that share a vision of an Euro-pean Exascale-class supercomputer. Their goal is designing and implementing a physical rack prototype together with its cooling system, the storage non-volatile memory (NVM) architecture and a low-latency interconnect able to test different options for interconnection and storage. Furthermore, the consortium is to provide real HPC applications to validate the system. Herein we provide a status report of the project initial developments. Roberto Ammendola, Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Pier Stanislao Paolucci, Elena Pastorelli, Francesco Simula, Piero Vicini, Giuliano Taffoni, Jose Antonio Pascual, Javier Navaridas, Mikel Luján, John Goodacre, Nikolaos Chrysos, Manolis Katevenis |
DSD | 18 |
| 2016 | The ExaNeSt Project: Interconnects, Storage, and Packaging for Exascale SystemsabstractExaNest is one of three European projects that support a ground-breaking computing architecture for exascale-class systems built upon power-efficient 64-bit ARM processors. This group of projects share an "everything-close" and "share-anything" paradigm, which trims down the power consumption -- by shortening the distance of signals for most data transfers -- as well as the cost and footprint area of the installation -- by reducing the number of devices needed to meet performance targets. In ExaNeSt, we will design and implement: (i) a physical rack prototype and its liquid-cooling subsystem providing ultra-dense compute packaging, (ii) a storage architecture with distributed (in-node) non-volatile memory (NVM) devices, (iii) a unified, low-latency interconnect, designed to efficiently uphold desired Quality-of-Service guarantees for a mix of storage with inter-processor flows, and (iv) efficient rack-level memory sharing, where each page is cacheable at only a single node. Our target is to test alternative storage and interconnect options on actual hardware, using real-world HPC applications. The ExaNeSt consortium brings together technology, skills, and knowledge across the entire value chain, from computing IP, packaging, and system deployment, all the way up to operating systems, storage, HPC, big data frameworks, and cutting-edge applications. Manolis Katevenis, Nikolaos Chrysos, Manolis Marazakis, Iakovos Mavroidis, Fabien Chaix, Nikolaos D. Kallimanis, Javier Navaridas, John Goodacre, Piero Vicini, Andrea Biagioni, Pier Stanislao Paolucci, Alessandro Lonardo, Elena Pastorelli, Francesca Lo Cicero, Roberto Ammendola, P. Hopton, P. Coates, Giuliano Taffoni, Stefano Cozzini, Martin L. Kersten, Julio Sahuquillo, Sergio Lechago, C. Pinto, Bernd Lietzow, D. Everett, Gino Perna |
DSD | 1 |
| 2016 | Discharging the Network From Its Flow Control Headaches: Packet Drops and HOL BlockingabstractCongestion control becomes indispensable in highly utilized consolidated networks running demanding applications. In this paper, proactive congestion management schemes for Clos networks are described and evaluated. The key idea is to move the congestion avoidance burden from the data fabric to a scheduling network, which isolates flows using per-flow request counters. The scheduling network comprises per-output arbiters that grant data packets after reserving space for them in the buffer memories in front of fabric outputs. Computer simulations show that this strategy eliminates head-of-line (HOL) blocking and its adversarial effects throughout the fabric, without having to drop packets. In particular, a simplified model describes this result as a synergy between proactive buffer reservations and fine-grained multipath routing. Two alternative designs are presented. The first one places all arbiters in a central control unit, is simpler, and has superior performance. The second is more scalable by distributing the arbiters over the switching elements of the Clos network and by routing the control messages to and from endpoint adapters via multiple paths. Computer simulations of the complete system demonstrate high throughput and low latency under any number of congested outputs. Weighted max-min fair allocation of fabric-output link bandwidth is also demonstrated. Furthermore, delay breakdowns show that the time that packets wait in fabric and resequencing buffers is minimized as a result of the reduced (and equalized across all fabric paths) in-fabric contention. Finally, the high throughput capability of the system is corroborated by a Markov chain analysis of output buffer credits. Nikolaos Chrysos, Lydia Y. Chen, Christoforos Kachris, Manolis Katevenis |
IEEE/ACM Trans. Netw. | 4 |
| 2015 | A Systematic Evaluation of Emerging Mesh-like CMP NoCsabstractThis paper studies alternative Network-on-Chip architectures for emerging many-core chip multiprocessors, by exploring the following design options on mesh-based networks: Multiple physical networks (P), cores concentration (C), express channels (X), it widths (W), and virtual channels (V). We exhaustively evaluate all combinations of the afore-mentioned parameters (P, C, X, W, V), using the energy-throughput ratio (ETR) as a metric to classify network congurations. Our experimental results show that, on one hand, with an appropriate selection of parameters (V,W), an optimized baseline 2D mesh offers the best possible ETR for NoCs with up to a few tens of cores (64-core NoC). More complicated networks, using concentration and express channels, can reduce the zero-load latency, but do not necessarily help to improve ETR. On the other hand, for larger CMPs, a 2D mesh with multiple physical networks is a better option: once optimized, this architectural choice can reduce the ETR by up to 46% for 256 cores. Antonis Psathakis, Vassilis Papaefstathiou, Nikolaos Chrysos, Fabien Chaix, Evangelos Vasilakis, Dionisios N. Pnevmatikatos, Manolis Katevenis |
ANCS | 7 |
| 2014 | EUROSERVER: Energy Efficient Node for European Micro-ServersabstractEUROSERVER is a collaborative project that aims to dramatically improve data centre energy-efficiency, cost, and software efficiency. It is addressing these important challenges through the coordinated application of several key recent innovations: 64-bit ARM cores, 3D heterogeneous silicon-on-silicon integration, and fully-depleted silicon-on-insulator (FD SOI) process technology, together with new software techniques for efficient resource management, including resource sharing and workload isolation. We are pioneering a system architecture approach that allows specialized silicon devices to be built even for low-volume markets where NRE costs are currently prohibitive. The EUROSERVER device will embed multiple silicon "chiplets" on an active silicon interposer. Its system architecture is being driven by requirements from three use cases: data centres and cloud computing, telecom infrastructures, and high-end embedded systems. We will build two fully integrated full-system prototypes, based on a common micro-server board, and targeting embedded servers and enterprise servers. Yves Durand, Paul M. Carpenter, Stefano Adami, Angelos Bilas, Denis Dutoit, Alexis Farcy, Georgi Gaydadjiev, John Goodacre, Manolis Katevenis, Manolis Marazakis, Emil Matús, Iakovos Mavroidis, John Thomson |
DSD | 9 |
| 2014 | Design trade-offs in energy efficient NoC architecturesabstractThis paper studies design trade-offs in energy efficient Networks-on-Chip by evaluating every network architecture that derives when we apply all possible variations of design-configuration parameters on a baseline 2D mesh. Network separation (P), concentration (C), express channels (X), flit widths (W), and virtual channels (V). Our comperative analysis selects the network architecture configuration that gives the best energy delay product (EDP) while allowing a maximum area margin of 15% over the most energy efficient configuration of the baseline. Antonis Psathakis, Vassilis Papaefstathiou, Manolis Katevenis, Dionisios N. Pnevmatikatos |
NOCS | 3 |
| 2014 | FPGA prototyping of emerging manycore architectures for parallel programming research using Formic boards
Spyros Lyberis, George Kalokerinos, Michalis Lygerakis, Vassilis Papaefstathiou, Iakovos Mavroidis, Manolis Katevenis, Dionisios N. Pnevmatikatos, Dimitrios S. Nikolopoulos |
J. Syst. Archit. | 6 |
| 2013 | Prefetching and cache management using task lifetimesabstractTask-based dataflow programming models and runtimes emerge as promising candidates for programming multicore and manycore architectures. These programming models analyze dynamically task dependencies at runtime and schedule independent tasks concurrently to the processing elements. In such models, cache locality, which is critical for performance, becomes more challenging in the presence of fine-grain tasks, and in architectures with many simple cores. Vassilis Papaefstathiou, Manolis Katevenis, Dimitrios S. Nikolopoulos, Dionisios N. Pnevmatikatos |
ICS | 2 |
| 2013 | NP-SARC: Scalable network processing in the SARC multi-core FPGA platform
Christoforos Kachris, George Nikiforos, Vassilis Papaefstathiou, Stamatis G. Kavadias, Manolis Katevenis |
J. Syst. Archit. | 5 |
| 2012 | Formic: Cost-efficient and Scalable Prototyping of Manycore ArchitecturesabstractModeling emerging multicore architectures is challenging and imposes a tradeoff between simulation speed and accuracy. An effective practice that balances both targets well is to map the target architecture on FPGA platforms. We find that accurate prototyping of hundreds of cores on existing FPGA boards faces at least one of the following problems: (i) limited fast memory resources (SRAM) to model caches, (ii) insufficient inter-board connectivity for scaling the design or (iii) the board is too expensive. We address these shortcomings by designing a new FPGA board for multicore architecture prototyping, which explicitly targets scalability and cost-efficiency. Formic has a 35% bigger FPGA, three times more SRAM, four times more links and costs at most half as much when compared to the popular Xilinx XUPV5 prototyping platform. We build and test a 64-board system by developing a 512-core, Micro Blaze-based, non-coherent hardware prototype with DMA capabilities, with full network on-chip in a 3D-mesh topology. We believe that Formic offers significant advantages over existing academic and commercial platforms that can facilitate hardware prototyping for future many core architectures. Spyros Lyberis, George Kalokerinos, Michalis Lygerakis, Vassilis Papaefstathiou, Dimitrios Tsaliagkos, Manolis Katevenis, Dionisios N. Pnevmatikatos, Dimitrios S. Nikolopoulos |
FCCM | 6 |
| 2012 | Crossbar NoCs Are Scalable Beyond 100 NodesabstractWe describe the design and layout of a radix-128 crossbar in 90 nm CMOS. The data path is 32 bits wide and runs at 750 MHz using a three-stage pipeline, while fitting in a silicon area as small as 6.6 mm2by filling it at the 90% level. The control path occupies 7 mm2next to the data path by filling it at 35% level, and reconfigures the data path once every three clock cycles. Next, we arrange 128 1 mm2“user tiles” around the crossbar, forming a 150 mm2die, and we connect all tiles to the crossbar via global links running on top of the tiles. Including the overhead of repeaters and flip flops on global links, the area cost of the crossbar is 11% of the die. Thus, we prove that crossbar networks-on-chips (NoCs) are small enough for radices exceeding by far the few tens of ports, that were believed to be the practical limit up to now, and reaching above 100 ports. We also attempt a first-order comparison between our crossbar and a model of a popular mesh NoC, and we find that our crossbar NoC increases performance when traffic is global and stressed, at the cost of worse performance when traffic is local and benign. Finally, we present an experimental cost analysis showing that crossbar area practically grows asO(N2W), as all wiring of the crossbar fits over its standard cells, while crossbar delay grows as O(N√W) , as wire length increases with the perimeter of the crossbar. Giorgos Passas, Manolis Katevenis, Dionisios N. Pnevmatikatos |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | LP-NUCA: Networks-in-Cache for High-Performance Low-Power Embedded ProcessorsabstractHigh-end embedded processors demand complex on-chip cache hierarchies satisfying several contradicting design requirements such as high-performance operation and low energy consumption. This paper introduces light-power (LP) nonuniform cache architecture (NUCA), a tiled-cache addressing both goals. LP-NUCA places a group of small and low-latency tiles between the L1 and the last level cache (LLC) that adapt better to the application working sets and keep most recently evicted blocks close to L1. LP-NUCA is built around three specialized “networks-in-cache,” each aimed at a separate cache operation. To prove the design feasibility, we have fully implemented LP-NUCA in a 90-nm technology. From the VLSI implementation, we observe that the proposed networks-in-cache incur minimal area, latency, and power overhead. To further reduce the energy consumption, LP-NUCA employs two network-wide techniques (miss wave stopping and sectoring) that together reduce the dynamic cache energy by 35% without degrading performance. Our evaluations also show that LP-NUCA improves performance with respect to cache hierarchies similar to those found in high-end embedded processors. Similar results have been obtained after scaling to a 32-nm technology. Darío Suárez Gracia, Giorgos Dimitrakopoulos, Teresa Monreal Arnal, Manolis Katevenis, Víctor Viñals |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2011 | Fine-grain OpenMP runtime support with explicit communication hardware primitivesabstractWe present a runtime system that uses the explicit on-chip communication mechanisms of the SARC multi-core architecture, to implement efficiently the OpenMP programming model and enable the exploitation of fine-grain parallelism in OpenMP programs. We explore the design space of implementation of OpenMP directives and runtime intrinsics, using a family of hardware primitives; remote stores, remote DMAs, hardware counters and hardware event queues with automatic responses, to support static and dynamic scheduling and data transfers in local memories. Using an FPGA prototype with four cores, we achieve OpenMP task creation latencies of 30-35 processor clock cycles, initiation of parallel contexts in 50 cycles and synchronization primitives in 65-210 cycles. Pranav Tendulkar, Vassilis Papaefstathiou, George Nikiforos, Stamatis G. Kavadias, Dimitrios S. Nikolopoulos, Manolis Katevenis |
DATE | 6 |
| 2011 | An efficient sequential iterative matching algorithm for CIOQ switchesabstractThis paper presents a sequential mode that can be used to improve the efficiency of iterative matching algorithm for CIOQ crossbar switches. The proposed matching algorithm is collectively called SIM-δ, and four different configuration are presented: restricted iteration, pipelined iteration, exhaustive transmission, and GSI (general state information)-based matching arbitration. The implementation-based simulation results show that switches with SIM-δ outperform switches with other matching algorithms, such as iSLIP, under uniform Bernoulli and bursty, non-uniform local hotspot and diagonal traffic models. We also explore the impact of speedup to throughput and delay. For uniform Bernoulli traffic, without or with small speedup, SIM-δ can achieve higher aggregate throughput and lower average delay than iSLIP. Yanping Gao, Christoforos Kachris, Manolis Katevenis |
ISCC | 3 |
| 2011 | VLSI micro-architectures for high-radix crossbar schedulersabstractWe study the scaling of parallel-matching crossbar schedulers to radices above 100. First, we examine a traditional microarchitecture that implements the matching decision of each input and each output of the crossbar in a separate arbiter block and communicates the matching decisions between the input and the output arbiters through global point-to-point links. Using simple models and experimentation with 90nm CMOS layouts, we show that this architecture is expensive because the global point-to-point links take up O(N4) area, where N the radix of the crossbar. Next, by observing that the wiring of an arbiter fits in a minimal O(NlogN) area, we propose a novel microarchitecture that inverts the locality of wires by orthogonally interleaving the input with the output arbiters, thus lowering the wiring area of the scheduler down to O(N2log2N). Using this architecture, the scheduler for a radix-128 FIFO, VOQ, or 2-VC crossbar becomes gate limited, fitting in 3.6, 7.2, and 7.2mm2 respectively, which is a 40, 50, and 70% improvement compared to the traditional. Moreover, the proposed schedulers find a new match in less than 10ns, thus allowing a minimum packet below 30Bytes at 24Gb/s line rate. Based on these findings, we conclude that crossbar schedulers are feasible even for radices above 100. Giorgos Passas, Manolis Katevenis, Dionisios N. Pnevmatikatos |
NOCS | 2 |
| 2011 | Distributed WFQ scheduling converging to weighted max-min fairness
Nikolaos Chrysos, Manolis Katevenis |
Comput. Networks | 2 |
| 2010 | End-to-end congestion management for non-blocking multi-stage switching fabricsabstractPacket-switched networks are encountered at the heart of scalable network routers and high-performance computer (or data center) interconnects. As these networks scale to larger port counts, and their utilization increases, congestion management becomes indispensable. At the same time, technology constraints rule out monolithic bufferless switches with centralized schedulers, and impose buffered multi-stage switching fabrics with distributed control. These trends have for some time now called forth research and products [1, 2, 3, 4, 5], which applied the request-grant philosophy of bufferless crossbars to make buffered multi-stage switching fabrics practical and efficient. In such proactive schemes, saturation-tree congestion is avoided by having inputs inform outputs of their demand, and inject data only after receiving output permission (grants). The per output admission arbiters can be located in a central scheduling unit, as in [3] [2] [4], can be distributed across the edge switches of the fabric [4], or can be placed in the respective output adapters [1] [5]. Such schemes have been quite successful, especially because they can be combined conveniently with reorder/reassembly buffer management, as well as with end-to-end reliable-delivery schemes. Nikolaos Chrysos, Lydia Y. Chen, Cyriel Minkenberg, Christoforos Kachris, Manolis Katevenis |
ANCS | 5 |
| 2010 | Efficient implementation of CIOQ switches with sequential iterative matching algorithmsabstractThis paper presents the implementation of a crossbar-based CIOQ switch with a novel matching algorithm that can be used for efficient interprocessor communication. The proposed matching algorithm, called SIM, uses a sequential mode to replace the traditional parallel mode to improve the efficiency of iterative matching operation. Moreover, the paper presents a FIFO-based request scheduler for the implementation of SIM. The proposed algorithm and its implementation architecture are evaluated in an FPGA. The implementation-based simulation results show that for switches without speedup, SIM can achieve higher aggregate throughput and lower average delay than iSLIP. Moreover, SIM achieves a fair trade-off among performance and logic area. Christoforos Kachris, Manolis Katevenis |
FPT | 3 |
| 2010 | A 128 x 128 x 24Gb/s Crossbar Interconnecting 128 Tiles in a Single Hop and Occupying 6% of Their AreaabstractWe describe the implementation of a 128×128 crossbar switch in 90 nm CMOS standard-cell ASIC technology. The crossbar operates at 750 MHz and is 32-bits for a port capacity above 20Gb/s, while fitting in a silicon area as small as 6.6 mm2by filling it at the 90% level (control not included). Next, we arrange 128 1 mm2"user tiles" around the crossbar, forming a 150 mm2die, and we connect all tiles to the crossbar via global links that run on top of SRAM blocks that we assume to occupy three fourths of each user tile. Including the overhead of repeaters and pipeline registers on the global links, the area cost of the crossbar is 6% of the total tile area. Thus, we prove that crossbars are dense enough and can be connected "for free" for valencies exceeding by far the few tens of ports, that were believed to be the practical limit up to now, and reaching above one hundred ports. Applications include Combined Input-Qutput Queued switch chips for Internet routers and data-center interconnects and the replacement of mesh-type NoC for many-core chips. Giorgos Passas, Manolis Katevenis, Dionisios N. Pnevmatikatos |
NOCS | 2 |
| 2007 | Approaching Ideal NoC Latency with Pre-Configured RoutesabstractIn multi-core ASICs, processors and other compute engines need to communicate with memory blocks and other cores with latency as close as possible to the ideal of a direct buffered wire. However, current state of the art networks-on-chip (NoCs) suffer, at best, latency of one clock cycle per hop. We investigate the design of a NoC that offers close to the ideal latency in some preferred, run-time configurable paths. Processors and other compute engines may perform network reconfiguration to guarantee low latency over different sets of paths as needed. Flits in non-preferred paths are given lower priority than flits in preferred ones, and suffer a delay of one clock cycle per hop when there is no contention. To achieve our goal, we use the "mad-postman" technique: every incoming flit is eagerly (i.e. speculatively) forwarded to the input's preferred output, if any. This is accomplished with the mere delay of a single pre-enabled tri-state driver. We later check if that decision was correct, and if not, we forward the flit to the proper output. Incorrectly forwarded flits are classified as dead and eliminated in later hops. We use a 2D mesh topology tailored for processor-memory communication, and a modified version of XY routing that remains deadlock-free. Performance gains are significant and can be proven greatly useful in other application domains as well George Michelogiannakis, Dionisios N. Pnevmatikatos, Manolis Katevenis |
NOCS | 3 |
| 2007 | Pipelined heap (priority queue) management for advanced scheduling in high-speed networks
Aggelos Ioannou, Manolis Katevenis |
IEEE/ACM Trans. Netw. | 2 |
| 2006 | Scheduling in Non-Blocking Buffered Three-Stage Switching FabricsabstractAbstract — Three-stage non-blocking switching fabrics are the next step in scaling current crossbar switches to many hundreds or few thousands of ports. Congestion management, however, is the central open problem; without it, performance suffers heavily under real-world traffic patterns. Schedulers for bufferless crossbars perform congestion management but are not scalable to high valencies and to multi-stage fabrics. Distributed scheduling, as used in buffered crossbars, is scalable but has never been scaled beyond crossbar valencies. We combine ideas from central and distributed schedulers, from request-grant protocols and from credit-based flow control, to propose a novel, practical architecture for scheduling in non-blocking buffered switching fabrics. The new architecture relies on multiple, independent, single-resource schedulers, operating in a pipeline. It: (i) isolates well-behaved against congested flows; (ii) provides throughput in excess of 95 % under unbalanced traffic, and delays that successfully compete again output queueing; (iii) provides weighted max-min fairness; (iv) directly operates on variable-size packets or multi-packet segments; (v) resequences cells or segments using very small buffers; and (vi) can be realistically implemented for a 1024×1024 reference fabric made out of 32×32 buffered crossbar switch elements. This paper carefully studies the many intricacies of the problem and the solution, discusses implementation, and provides performance simulation results. 1 Nikolaos Chrysos, Manolis Katevenis |
INFOCOM | 2 |
| 2005 | Scheduling in switches with small internal buffersabstractUnbuffered crossbars or switching fabrics contain no internal buffers, and function using only input (VOQ) and possibly output queues. Schedulers for such switches are complex, and introduce increased delay at medium loads, because they have to admit at most one cell per input and per output, during each time slot. Buffered crossbars, on the other hand, contain Q sufficient internal buffering (N2buffers) to allow independent schedulers to concurrently forward packets to the same output from any number of inputs. These architectures represent the two extremes in a range of solutions, which we examine here; although intermediate points in this range are of reduced practical interest for crossbars, they are nevertheless quite interesting for switching fabrics, and they may be of interest for optical switches. We find that tolerating two cells per-output per time-slot, using small buffers inside the switch or fabric, suffices for independent and efficient scheduling. First, we introduce a novel "request-grant" credit protocol, enabling N inputs to share a small switch buffer. Then, we apply this protocol to a switch with N such buffers, one per output, and we consider the resulting scheduling problem. Interestingly, this looks like unbuffered crossbar schedulers, but it is much simpler because it comprises independent schedulers that can be pipelined. We show that individual buffer sizes do not need to grow, neither with switch size nor with propagation delay. Through simulations, we study performance as a function of the number of cells allowed per-output per-time-slot. For one cell, the switch performs very close to the iSLIP unbuffered crossbar with one iteration. For more cells, performance improves quickly; for 12 cells, packet delay under (smooth) uniform load is practically as low as ideal output queueing. Under unbalanced load, throughput is superior to buffered crossbars, due to better buffer sharing Nikolaos Chrysos, Manolis Katevenis |
GLOBECOM | 2 |
| 2005 | Variable-size multipacket segments in buffered crossbar (CICQ) architecturesabstractBuffered crossbars can directly switch variable size packets, but require large cross point buffers to do so, especially when jumbo frames are to he supported. When this is not feasible, segmentation and reassembly (SAR) must be used. We propose a novel SAR scheme for buffered crossbars that uses variable-size segments while merging multiple packets (or fragments thereof) into each segment. This scheme eliminates padding overhead, reduces header overhead, reduces crosspoint buffer size and is suitable for use with external, modern DRAM buffer memory in the ingress line cards. We evaluate the new scheme using simulation, and show that it outperforms existing segmentation schemes in buffered as well as unbuffered crossbars. We also study how the size of the maximum segment affects system performance. Manolis Katevenis, Giorgos Passas |
ICC | 1 |
| 2004 | Multiple priorities in a two-lane buffered crossbarabstractA significant advantage of buffered crossbar switches is that they can directly operate on variable-size packets. However, in order to support multiple priority levels, separate queues per priority are needed at each crosspoint, in order to prevent HOL blocking and buffer hogging; these queues are expensive because they each need a size of at least one maximum-size packet. In this paper, we propose a scheme that uses only two queues per crosspoint to effectively support multiple priorities. We adaptively adjust the priority levels of the two queues so that most traffic goes through the "lower" queue, while the "upper" queue remains usually available for higher priority packets to overtake the former. Through simulation, and assuming 8 priority levels, we compare our scheme to an ideal system that uses 8 queues per crosspoint. For realistic traffic, the two systems perform almost identically, although ours uses 4 times less memory in the crossbar. Even under a highly irregular traffic pattern, Bursts60, our system does not increase the average delay of any priority level by more than 75% compared to the ideal system. Nikolaos Chrysos, Manolis Katevenis |
GLOBECOM | 2 |
| 2004 | Variable packet size buffered crossbar (CICQ) switchesabstractOne of the most widely used architectures for packet switches is the crossbar. A special version of it is the buffered crossbar, where small buffers are associated with the crosspoints; this simplifies scheduling and improves its efficiency and QoS capabilities to the point where the switch needs no internal speedup. Furthermore, by supporting variable length packets throughout a buffered crossbar: (a) there is no need for segmentation and reassembly (SAR) circuits; (b) no speedup is necessary to support SAR; and (c) synchronization between the input and output clock domains is simplified. In turn, the lack of SAR and speedup mean that no output queues are needed, either. In this paper we present an architecture, a chip layout and cost analysis, and a performance evaluation of such a 300 Gbps buffered crossbar operating on variable-size packets. The proposed organization is simple yet powerful, it can be implemented using modern technology, and, as the performance results demonstrate, it clearly outperforms unbuffered crossbars. Manolis Katevenis, Giorgos Passas, Dimitrios Simos, Ioannis Papaefstathiou, Nikolaos Chrysos |
ICC | 1 |
| 2002 | Web-conscious storage management for web proxiesabstractMany proxy servers are limited by their file I/O needs. Even when a proxy is configured with sufficient I/O hardware, the file system software often fails to provide the available bandwidth to the proxy processes. Although specialized file systems may offer a significant improvement and overcome these limitations, we believe that user-level disk management on top of industry-standard file systems can offer similar performance advantages. We study the overheads associated with file I/O in Web proxies, we investigate their underlying causes, and we propose Web-conscious storage management, a set of techniques that exploit the unique reference characteristics of Web-page accesses in order to allow Web proxies to overcome file I/O limitations. Using realistic trace-driven simulations, we show that these techniques can improve the proxy's secondary storage I/O throughput by a factor of 15 over traditional open-source proxies, enabling a single disk to serve over 400 (URL-get) operations per second. We implement Foxy, a Web proxy which incorporates our techniques. Experimental evaluation suggests that Foxy outperforms traditional proxies, such as SQUID, by more than a factor of four in throughput, without sacrificing response latency. Evangelos P. Markatos, Dionisios N. Pnevmatikatos, Michail Flouris, Manolis Katevenis |
IEEE/ACM Trans. Netw. | 4 |
| 2001 | Pipelined heap (priority queue) management for advanced scheduling in high-speed networksabstractQuality-of-service (QoS) guarantees in networks are increasingly based on per-flow queueing and sophisticated scheduling. Most advanced scheduling algorithms rely on a common computational primitive: priority queues. Large priority queues are built using calendar queue or heap data structures. To support advanced scheduling at OC-192 (10 Gbps) rates and above, pipelined management of the priority queue is needed. We present a pipelined heap manager that we have designed as a core integratable into ASICs, in synthesizable Verilog form. We discuss how to use it in switches and routers, its advantages over calendar queues, and we present cost-performance tradeoffs. Our design can be configured to any heap size. We have verified and synthesized our design and present cost and performance analysis information. Aggelos Ioannou, Manolis Katevenis |
ICC | 2 |
| 2001 | Efficient per-flow queueing in DRAM at OC-192 line rate using out-of-order execution techniquesabstractModern switches and routers often use dynamic RAM (DRAM) in order to provide large buffer space. For advanced quality of service (QoS), per-flow queueing is desirable. We study the architecture of a queue manager for many thousands of queues at OC-192 (10 Gbps) line rate. It forms the core of the "datapath" chip in an efficient chip partitioning for the line cards of switches and routers that we propose. To effectively deal with bank conflicts in the DRAM buffer, we use pipelining and out-of-order execution techniques, like the ones originating in the supercomputers of the 60s. To avoid off-chip SRAM, we maintain the pointers in the DRAM, using free buffer preallocation and free list bypassing. We have described our architecture using behavioral Verilog (a hardware description language), at the clock-cycle accuracy level, assuming Rambus DRAM. We estimate the complexity of the queue manager at roughly 60 thousand gates, 80 thousand flip-flops, and 4180 Kbits of on-chip SRAM, for 64 K flows. Aggelos Nikologiannis, Manolis Katevenis |
ICC | 2 |
| 2001 | Wormhole IP over (connectionless) ATMabstractHigh-speed switches and routers internally operate using fixed-size cells or segments; variable-size packets are segmented and later reassembled. Connectionless ATM was proposed to quickly carry IP packets segmented into cells (AAL5) using a number of hardware-managed ATM VCs. We show that this is analogous to wormhole routing. We modify this architecture to make it applicable to existing ATM equipment: we propose a low-cost, single-input, single-output wormhole IP router that functions as a VP/VC translation filter between ATM subnetworks. When compared to IP routers, the proposed architecture features simpler hardware and lower latency. When compared to software-based IP-over-ATM techniques, the new architecture avoids the overheads of a large number of labels, or the delays of establishing new flows in software after the first few packets have suffered considerable latencies. We simulated a wormhole IP routing filter, showing that a few tens of hardware-managed VCs per outgoing VP usually suffice. We built and successfully tested a prototype, operating at 2/spl times/155 Mb/s, using one field programmable gate array (FPGA) and DRAM. Simple analysis shows that operation at 10 Gb/s and beyond is feasible today. Manolis Katevenis, Iakovos Mavroidis, Georgios Sapountzis, Evangelia Kalyvianaki, Ioannis Mavroidis, Georgios Glykopoulos |
IEEE/ACM Trans. Netw. | 1 |
| 1998 | Credit-Flow-Controlled ATM for MP Interconnection: The ATLAS I Single-Chip ATM SwitchabstractMultiprocessing (MP) on networks of workstations (NOW) is a high-performance computing architecture of growing importance. In traditional MP's, wormhole routing interconnection networks use fixed-size flits and backpressure. In NOW's, ATM-one of the major contending interconnection technologies-uses fixed-size cells, while backpressure can be added to it. We argue that ATM with backpressure has interesting similarities with wormhole routing. We are implementing ATLAS I, a single-chip gigabit ATM switch, which includes credit flow control (backpressure), according to a protocol resembling Quantum Flow Control (QFC). We show by simulation that this protocol performs better than the traditional multi-lane wormhole protocol: high throughput and low latency are provided with less buffer space. Also, ATLAS I demonstrates little sensitivity to bursty traffic, and, unlike wormhole, it is fair in terms of latency in hot-spot configurations. We use detailed switch models, operating at clock-cycle granularity. Manolis Katevenis, Dimitrios Serpanos, Emmanuel Spyridakis |
HPCA | 1 |
| 1997 | User-Level DMA without Operating System Kernel ModificationabstractDirect Memory Access (DMA) is frequently used to transfer data between the main memory of a host computer and the interconnection network, in order to free the host processor from the burden of the transfer. DMA operations are traditionally initiated by the operating system kernel, mainly to prevent one application from tampering with another applications' data. Recent architecture trends suggest that interconnection networks get faster, while operating systems get slower (compared to processor speeds). These trends imply that the initiation of a DMA operation becomes slower (due to operating system involvement), while the DMA data transfer itself becomes faster with time. Soon, the operating system overhead associated with starting a DMA will be larger than the data transfer itself, esp. for small data transfers. This paper proposes several algorithms that allow user-level applications to start DMA operating without the involvement of the operating system. Our algorithms allow user applications to have direct (but controlled) access to the DMA engine registers. Overhead user-level DMA is achieved without compromising protection, and without requiring changes to the underlying operating system kernel. Using our proposed algorithms, a DMA operation can be initiated in 2 to 5 assembly instructions. By comparison, operating system-based initiation of DMA requires thousands of assembly instructions. Evangelos P. Markatos, Manolis Katevenis |
HPCA | 2 |
| 1997 | Telegraphos: A Substrate for High-Performance Computing on Workstation Clusters
Manolis Katevenis, Evangelos P. Markatos, George Kalokerinos, Apostolos Dollas |
J. Parallel Distributed Comput. | 1 |
| 1996 | Telegraphos: High-Performance Networking for Parallel Processing on Workstation ClustersabstractNetworks of workstations and high-performance microcomputers have been rarely used for running high-performance applications like multimedia, simulations, scientific and engineering applications, because, although they have significant aggregate computing power, they lack the support for efficient message-passing and shared-memory communication. In this paper we present Telegraphos, a distributed system that provides efficient shared-memory support on top of a workstation cluster. We focus on the network interface of Telegraphos that provides a variety of shared-memory operations like remote reads, remote writes, remote atomic operations, all launched from user level without any intervention of the operating system. Telegraphos I, the first Telegraphos prototype has been implemented. Emphasis was put on rapid prototyping, so the technology used was conservative: FPGA's, SRAM's, and TTL buffers. Telegraphos II, is the single-chip version of the Telegraphos architecture; its switch was implemented and its network interface is being debugged. Evangelos P. Markatos, Manolis Katevenis |
HPCA | 2 |
| 1995 | Pipelined Memory Shared Buffer for VLSI SwitchesabstractSwitch chips are building blocks for computer and communication systems. Switches need internal buffering, because of output contention; shared buffering is known to perform better than multiple input queues or buffers, and the VLSI implementation of the former is not more expensive than the latter. We present a new organization for a shared buffer with its associated switching and cut-through functions. It is simpler and smaller than wide or interleaved organizations, and it is particularly suitable for VLSI technologies. It is based on multiple memory banks, addressed in a pipelined fashion. The first word of a packet is transferred to/from the first bank, followed by a "wave" of similar operations for the remaining words in the remaining banks. An FPGA-based prototype is operational, while standard-cell and full-custom chips are being submitted for fabrication. Simulation of the full-custom version indicates that, even in a conservative 1-micron CMOS technology, a 64 Kbit central buffer for an 8×8 switch operates at 1 Gbps/link (worst case) and fits in 45 mm2 including crossbar and cut-through. Manolis Katevenis, Panagiota Vatsolaki, Aristides Efthymiou |
SIGCOMM | 1 |
| 1991 | Reducing the Branch Penalty by Rearranging Instructions in Double-Width MemoryabstractArticle Free Access Share on Reducing the branch penalty by rearranging instructions in a double-width memory Authors: Manolis Katevenis View Profile , Nestoras Tzartzanis View Profile Authors Info & Claims ASPLOS IV: Proceedings of the fourth international conference on Architectural support for programming languages and operating systemsApril 1991 Pages 15–27https://doi.org/10.1145/106972.106977Published:01 April 1991Publication History 5citation533DownloadsMetricsTotal Citations5Total Downloads533Last 12 Months125Last 6 weeks20 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Manolis Katevenis, Nestoras Tzartzanis |
ASPLOS | 1 |
| 1991 | Weighted Round-Robin Cell Multiplexing in a General-Purpose ATM Switch ChipabstractThe authors present the architecture of a general-purpose broadband-ISDN (B-ISDN) switch chip and, in particular, its novel feature: the weighted round-robin cell (packet) multiplexing algorithm and its implementation in hardware. The flow control and buffer management strategies that allow the chip to operate at top performance under congestion are given, and the reason why this multiplexing scheme should be used under those circumstances is explained. The chip architecture and how the key choices were made are discussed. The statistical performance of the switch is analyzed. The critical parts of the chip have been laid out and simulated, thus proving the feasibility of the architecture. Chip sizes of four to ten links with link throughput of 0.5 to 1 Gb/s and with about 1000 virtual circuits per switch have been realized. The results of simulations of the chip are presented.> Manolis Katevenis, Stefanos Sidiropoulos, Costas Courcoubetis |
IEEE J. Sel. Areas Commun. | 1 |
| 1987 | A Vector Hardware Accelerator with Circuit Simulation EmphasisabstractA floating-point vector accelerator has been built which runs circuit simulation efficiently. The design considerations of the accelerator are based on the time-consuming parts of SPICE2, available off-the-shelf parts, advanced software tools experience and cost/performance. The three board accelerator can run the entire application program complied from a high-level language. A personal workstation, such as the PC-AT, is used for the general I/O tasks such as file handling and network support. The processor has a Single-Instruction Multiple-Data 64-bit floating-point pipelined architecture. It can achieve a maximum speed of 8 Mips and 8 MFlops. A floating-point processor based on two functional units, a multiplier and an ALU, and an integer processor work in parallel to achieve the high performance. The accelerator attached to a PC-AT runs SPICE2 60 times faster than the personal workstation alone and achieves double the performance of a VAX 8650. Andrei Vladimirescu, Manolis Katevenis, Zvika Bronstein, Alon Kifir, Karja Danuwidjaja, K. C. Ng, Niraj Jain, Steven Lass |
DAC | 3 |
| 1987 | Fast switching and fair control of congested flow in broadband networksabstractA new switching architecture is proposed based on the tradeoffs of modern VLSI technology-inexpensive memory and 2-dimensional layout structures. Today, it is economically feasible to preallocate buffer space individually to each virtual circuit in every node, so that "congestion" ceases to have negative effects. On the contrary, when some low-priority circuits offer more traffic than the network can carry, full utilization of the link bandwidth is achieved. In this context, the allocation of bandwidth can be done automatically and in a "fair" way, if packets are multiplexed by circularly scanning all virtual circuits and transmitting one packet from each "ready" circuit. This multiplexing algorithm equally distributes all the available link BW to all the VC's that can use it (other than equal distribution is also possible), while it also guarantees an upper bound for the total packet delay through noncongested VC's (VC's that use less than their share of BW). We present methods for hardware implementation of such fast circular scans, and propose a structure for the switching nodes of such networks, consisting of a cross-bar arrangement like a systolic array that performs merge sorting. It is ideally suited for physical layout on printed-circuit boards or with wafer-scale integration. Manolis Katevenis |
IEEE J. Sel. Areas Commun. | 1 |