EDBT 2026 Demo / reviewers in the wild / expert
Manolis Ploumidis
dblp:04/4735
· DBLP profile ↗
8ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0003-2173-062XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The ExaNeSt Prototype: Evaluation of Efficient HPC Communication Hardware in an ARM-based Multi-FPGA RackabstractWe present and evaluate the ExaNeSt prototype, which compactly packages 128 Xilinx ZU9EG MPSoCs, two TBytes of DRAM, and eight TBytes of SSD into a liquid-cooled rack, using a custom interconnection hardware based on 10 GB/s links. We developed this testbed in 2016–2019 in order to leverage the flexibility of FPGAs for experimenting with efficient hardware support for HPC communication among tens of thousands of processors and accelerators in the quest toward Exascale systems and beyond. In the years since then, we carefully studied this system, and we present our key design choices and insights resulting from our measurement and analysis. We developed this testbed, from architecture to the PCBs and the run-time software, within the ExaNeSt project. It is fully operational in configurations with up to 8 × 4 × 4 MPSoC nodes. It achieves high density through tight board design, while also leveraging state-of-the-art liquid cooling technology. In this article, we present a thorough architectural analysis, along with important aspects of our infrastructure development. Our custom interconnect includes a low-cost low-latency network interface, offering user-level, zero-copy RDMA, which we coupled with the ARMv8 processors in the MPSoCs. We further developed the corresponding runtimes that allow us to test real MPI applications on the large-scale testbed. We evaluated our platform through MPI microbenchmarks, mini application, and full MPI applications. Single-hop, one-way latency is 1.3 μs; approximately 0.47 μs out of these are attributed to network interface and the user-space library that exposes its functionality to the runtime. Latency over longer paths increases as expected, reaching 2.55 μs for a five-hop path. Bandwidth tests show that, for single-hop, link utilization reaches \(82\%\) of the theoretical capacity. Microbenchmarks based on MPI collectives reveal that broadcast latency scales as expected when the number of participating ranks increases. We also implemented a custom MPI_Allreduce accelerator in the network interface, which reduces the latency of such collectives by up to \(88\%\) . We assess performance scaling through weak and strong scaling tests for HPCG, LAMMPS, and the miniFE mini application; for all these tests, parallelization efficiency is at least \(69\%\) , or better. Manolis Ploumidis, Fabien Chaix, Nikolaos Chrysos, Marios Assiminakis, Nikolaos D. Kallimanis, Nikolaos Kossifidis, Michael Nikoloudakis, Nikolaos Dimou, Michalis Gianioudis, Giorgos Ieronymakis, Aggelos Ioannou, George Kalokerinos, Pantelis Xirouchakis, Astrinos Damianakis, Michael Ligerakis, Theocharis Vavouris, Manolis Katevenis, Vassilis Papaefstathiou, Manolis Marazakis, Iakovos Mavroidis |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2023 | Impact of Cache Coherence on the Performance of Shared-Memory based MPI Primitives: A Case Study for Broadcast on Intel Xeon Scalable ProcessorsabstractRecent processor advances have made feasible HPC nodes with high core counts, capable of hosting tens or even, hundreds of processes. Therefore, designing MPI collective operations at the intra-node level has received significant attention over the past years. Deriving efficient algorithms for modern HPC nodes, with complex internal topologies and memory hierarchies, is challenging. Moreover, the cache coherency protocol, and its impact on performance, further complicate algorithm design for MPI collectives. This latter concern is often only partially addressed. George Katevenis, Manolis Ploumidis, Manolis Marazakis |
ICPP | 2 |
| 2022 | A framework for hierarchical single-copy MPI collectives on multicore nodesabstractCollective operations are widely used by MPI applications to realize their communication patterns. Their efficiency is crucial for both performance and scalability of parallel applications. For deriving efficient MPI implementations, significant effort is put to keep pace with advances and capabilities of the underlying hardware and interconnect. Recent processor advances have led to nodes with higher core counts and complex internal structures and memory hierarchies. Such nodes are able to host tens to hundreds of processes and thus, performance of MPI collectives at the intra-node level becomes critical. In this work, we propose a framework for collective operations at the intra-node level, that aims to lower latency and increase bandwidth. Our approach utilizes knowledge of internal node structure to construct hierarchical algorithms, and XPMEM to achieve single-copy transfers. Pipelining is used to overlap communication at different levels of the hierarchy. We evaluate the proposed approach through several microbenchmarks and real-world MPI applications. For evaluation purposes, we compare the proposed approach with implementations of similar schemes from two recent studies. Our evaluation with microbenchmarks for Broadcast and Allreduce shows speedup up to 2.$5x$and$3x_{2}$respectively, over UCC and OpenMPI's default collectives implementation. Compared to recent research studies, we improve Broadcast by up to$5x_{2}$and Allreduce by up to$7x$. We reduce the time of three applications PiSvM, miniAMR and CNTK, by up to 12%, 52% and 12%, respectively, over the next best-performing alternative. George Katevenis, Manolis Ploumidis, Manolis Marazakis |
CLUSTER | 2 |
| 2020 | Towards Communication Profile, Topology and Node Failure Aware Process PlacementabstractHPC systems need to keep growing in size to meet the ever-increasing demand for high levels of capability and capacity, often in tight time windows for urgent computation. However, increasing the size, complexity and heterogeneity of HPC systems also increases the risk and impact of system failures, that result in resource waste and aborted jobs. A major contributor to job completion time is the cost of interprocess communication. To address performance and energy efficiency, several prior studies have targeted improvements of communication locality. To meet this goal, they derive a mapping of MPI processes to system nodes in a way that reduces communication cost. However, such approaches disregard the effect of system failures. In this work, we propose a resource allocation approach for MPI jobs, considering both high performance and error resilience. Our approach, named Communication Profile, Topology and node Failure (CPTF), takes into account the application's communication profile, system topology and node failure probability for assigning job processes to nodes. We evaluate variants of CPTF through simulations of two MPI applications, one with a regular communication pattern (LAMMPS) and one with an irregular one (NPB-DT). In both cases, the variant of CPTF that strives to avoid failure-prone nodes and communication paths achieves lower time to complete job batches when compared to the default resource allocation policy of Slurm. It also exhibits the lowest ratio of aborted jobs. The average improvement in batch completion time is 67% for NPB-DT and 34% for LAMMPS. Ioannis Vardas, Manolis Ploumidis, Manolis Marazakis |
SBAC-PAD | 2 |
| 2019 | Towards Exascale: Measuring the Energy Footprint of Astrophysics HPC SimulationsabstractThe increasing amount of data produced in Astronomy by observational studies and the size of theoretical problems to be tackled in the next future pushes the need of HPC (High Performance Computing) resources towards the "Exascale". The HPC sector is undergoing a profound phase of transition, in which one of the toughest challenges to cope with is the energy efficiency that is one of the main blocking factors to the achievement of "Exascale". Since ideal peak-performance is unlikely to be achieved in realistic scenarios, the aim of this work is to give some insights about the energy consumption of contemporary architectures with real scientific applications in a HPC context. We use two state-of-the-art applications from the astrophysical domain, that we optimized in order to fully exploit the underlying hardware: a direct N-body code and a semi-analytical code for Cosmic Structure formation simulations. For these two applications, we quantitatively evaluate the impact of computation on the energy consumption when running on three different systems: one that represents the present of current HPC systems (an Intel-based cluster), one that (possibly) represents the future of HPC systems (a prototype of an Exascale supercomputer) and a micro-cluster based on Arm MPSoC. We provide a comparison of the time-to-solution, energy-to-solution and energy delay product (EDP) metrics, for different software configurations. ARM-based HPC systems have lower energy consumption albeit running ≈10 times slower. Giuliano Taffoni, Manolis Katevenis, Renato Panchieri, Gino Perna, Luca Tornatore, David Goz, Antonio Ragagnin, Sara Bertocco, Igor Coretti, Manolis Marazakis, Fabien Chaix, Manolis Ploumidis |
eScience | 12 |
| 2015 | On the performance of network coding and forwarding schemes with different degrees of redundancy for wireless mesh networks
Manolis Ploumidis, Nikolaos Pappas 0001, Vasilios A. Siris, Apostolos Traganitis |
Comput. Commun. | 1 |
| 2007 | Multi-level application-based traffic characterization in a large-scale wireless networkabstractWith the increasing deployment of wireless networks, network management and configuration of wireless Access Points (APs) has become one of the main concerns of network operators. While statistics and measurements regarding the overall usage of individual APs are readily available, the limited knowledge of the wireless traffic demand in terms of the type of application hinders efficient network provisioning. This paper provides an extensive application-based characterization of a large-scale wireless network, going beyond the port-number limitation, across three levels, namely, network, clients, and APs. We found that the most popular application types, in terms of the number of flows, bytes, and clients, are web and peer-to-peer; and while the majority of APs is dominated by them, APs of the same building type have large differences in their traffic mix. File transfer flows, such as FTP and P2P, are heavier in wired than in wireless networks. Finally, an interesting dichotomy among APs, in terms of their dominant application type and downloading and uploading behavior was observed. Manolis Ploumidis, Maria Papadopouli, Thomas Karagiannis |
WOWMOM | 1 |
| 2005 | Short-Term Traffic Forecasting in a Campus-Wide Wireless NetworkabstractOur goal is to characterize the traffic load in an IEEE802.11 infrastructure. This can be beneficial in many domains, including coverage planning, resource reservation, network monitoring for anomaly detection, and producing more accurate simulation models. The key issue that drives this study is traffic forecasting at each wireless access point (AP) in an hourly timescale. We conducted an extensive measurement study of wireless users on a major university campus using the IEEE802.11 wireless infrastructure. We propose several traffic models that take into account the periodicity and recent traffic history for each AP and present a time-series forecasting methodology. Finally, we build and evaluate these forecasting algorithms and discuss our findings. Maria Papadopouli, Haipeng Shen, Elias Raftopoulos, Manolis Ploumidis, Félix Hernández-Campos |
PIMRC | 4 |