EDBT 2026 Demo / reviewers in the wild / expert
Nikolaos D. Kallimanis
dblp:57/2392 · also Nikos Kallimanis
· DBLP profile ↗
23ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0002-0331-1475ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 3 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Deep Reinforcement Learning Based Uplink Configuration in Multi-RIS Assisted MEC NetworkabstractMulti-access edge computing (MEC) is an important enabler of computationally demanding and delay sensitive intelligent applications, placing significant strain on traditional cloud computing systems. As massive multiple access and high offloading data rates are envisioned for the 6th generation (6G) networks, new paradigms that can support next generation multiple access systems have appeared, e.g., the reconfigurable intelligent surfaces (RIS). Particularly, a time-division multiple access (TDMA) MEC system can leverage the integration of multiple RISs. By regulating the interplay between multiple access control and RIS phase configuration, the inclusion of multiple RISs can improve the signal quality providing better channel conditions for uplink (UL) offloading. To that end, aiming to address the offloading sum-rate maximization problem in a multi-RIS assisted TDMA MEC network, we employ deep reinforcement learning (DRL) and propose a deep deterministic policy gradient (DDPG) based algorithm, i.e., UL-RIS-DDPG. DDPG is a DRL variant designed for high-dimensional continuous action and state spaces induced by the considered problem. The UL-RIS-DDPG algorithm jointly optimizes key parameters, i.e., the users' transmission powers, the RISs phase shift design, the user-RIS association, and the UL transmission scheduling and improves significantly the network's offloading sum-rate. The simulation results demonstrate that the algorithm adapts to the network environment offering flexible RIS selection and efficient time slot allocation, outperforming the benchmarks. Eftychia G. Datsika, Nikolaos D. Kallimanis |
WCNC | 2 |
| 2025 | The ExaNeSt Prototype: Evaluation of Efficient HPC Communication Hardware in an ARM-based Multi-FPGA RackabstractWe present and evaluate the ExaNeSt prototype, which compactly packages 128 Xilinx ZU9EG MPSoCs, two TBytes of DRAM, and eight TBytes of SSD into a liquid-cooled rack, using a custom interconnection hardware based on 10 GB/s links. We developed this testbed in 2016–2019 in order to leverage the flexibility of FPGAs for experimenting with efficient hardware support for HPC communication among tens of thousands of processors and accelerators in the quest toward Exascale systems and beyond. In the years since then, we carefully studied this system, and we present our key design choices and insights resulting from our measurement and analysis. We developed this testbed, from architecture to the PCBs and the run-time software, within the ExaNeSt project. It is fully operational in configurations with up to 8 × 4 × 4 MPSoC nodes. It achieves high density through tight board design, while also leveraging state-of-the-art liquid cooling technology. In this article, we present a thorough architectural analysis, along with important aspects of our infrastructure development. Our custom interconnect includes a low-cost low-latency network interface, offering user-level, zero-copy RDMA, which we coupled with the ARMv8 processors in the MPSoCs. We further developed the corresponding runtimes that allow us to test real MPI applications on the large-scale testbed. We evaluated our platform through MPI microbenchmarks, mini application, and full MPI applications. Single-hop, one-way latency is 1.3 μs; approximately 0.47 μs out of these are attributed to network interface and the user-space library that exposes its functionality to the runtime. Latency over longer paths increases as expected, reaching 2.55 μs for a five-hop path. Bandwidth tests show that, for single-hop, link utilization reaches \(82\%\) of the theoretical capacity. Microbenchmarks based on MPI collectives reveal that broadcast latency scales as expected when the number of participating ranks increases. We also implemented a custom MPI_Allreduce accelerator in the network interface, which reduces the latency of such collectives by up to \(88\%\) . We assess performance scaling through weak and strong scaling tests for HPCG, LAMMPS, and the miniFE mini application; for all these tests, parallelization efficiency is at least \(69\%\) , or better. Manolis Ploumidis, Fabien Chaix, Nikolaos Chrysos, Marios Assiminakis, Nikolaos D. Kallimanis, Nikolaos Kossifidis, Michael Nikoloudakis, Nikolaos Dimou, Michalis Gianioudis, Giorgos Ieronymakis, Aggelos Ioannou, George Kalokerinos, Pantelis Xirouchakis, Astrinos Damianakis, Michael Ligerakis, Theocharis Vavouris, Manolis Katevenis, Vassilis Papaefstathiou, Manolis Marazakis, Iakovos Mavroidis |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2023 | Exploring Trade-Offs in Partial Snapshot Implementations
Nikolaos D. Kallimanis, Eleni Kanellou, Charidimos Kiosterakis, Vasiliki Liagkou |
SSS | 1 |
| 2022 | RED-SEA: Network Solution for Exascale ArchitecturesabstractIn order to enable Exascale computing, next generation interconnection networks must scale to hundreds of thousands of nodes, and must provide features to also allow the HPC, HPDA, and AI applications to reach Exascale, while benefiting from new hardware and software trends. RED-SEA will pave the way to the next generation of European Exascale interconnects, including the next generation of BXI, as follows: (i) specify the new architecture using hardware-software co-design and a set of applications representative of the new terrain of converging HPC, HPDA, and AI; (ii) test, evaluate, and/or implement the new architectural features at multiple levels, according to the nature of each of them, ranging from mathematical analysis and modeling, to simulation, or to emulation or implementation on FPGA testbeds; (iii) enable seamless communication within and between resource clusters, and therefore development of a high-performance low latency gateway, bridging seamlessly with Ethernet; (iv) add efficient network resource management, thus improving congestion resiliency, virtualization, adaptive routing, collective operations; (v) open the interconnect to new kinds of applications and hardware, with enhancements for end-to-end network services - from programming models to reliability, security, low- latency, and new processors; (vi) leverage open standards and compatible APIs to develop innovative reusable libraries and Fabrics management solutions. Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Pier Stanislao Paolucci, Elena Pastorelli, Francesco Simula, Matteo Turisini, Piero Vicini, Roberto Ammendola, Pascale Bernier-Bruna, Said Derradji, Stéphane Guez, Pierre-Axel Lagadec, Gregoire Pichon, Etienne Walter, Gaetan De Gassowski, Matthieu Hautreaux, Stephane Mathieu, Gilles Moreau, Marc Pérache, Hugo Taboada, Torsten Hoefler, Timo Schneider, Matteo Barnaba, Giuseppe Piero Brandino, Francesco De Giorgi, Matteo Poggi, Iakovos Mavroidis, Ioannis Papaefstathiou, Nikolaos Tampouratzis, Benjamin Kalisch, Ulrich Krackhardt, Mondrian Nüssle, Pantelis Xirouchakis, Vangelis Mageiropoulos, Michalis Gianioudis, Harisis Loukas, Aggelos Ioannou, Nikolaos D. Kallimanis, Nikolaos Chrysos, Manolis Katevenis, Wolfgang Frings, Dominik Gottwald, Felime Guimaraes, Max Holicki, Volker Marx, Yannik Müller, Carsten Clauss, Hugo Falter, Xu Huang 0010, Jennifer Lopez Barillao, Thomas Moschny, Simon Pickartz, Francisco J. Alfaro, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José L. Sánchez 0002, Adrián Castelló 0001, Jose Duro, María Engracia Gómez, Enrique S. Quintana-Ortí, Julio Sahuquillo, Eugenio Stabile |
DSD | 43 |
| 2022 | The performance power of software combining in persistenceabstractThe availability of Non-Volatile Main Memory (known as NVMM) enables the design of recoverable concurrent algorithms. We study the power of software combining in achieving recoverable synchronization and designing persistent data structures. Software combining is a general synchronization approach, which attempts to simulate the ideal world when executing synchronization requests (i.e., requests that must be executed in mutual exclusion). A single thread, called the combiner, executes all active requests, while the rest of the threads are waiting for the combiner to notify them that their requests have been applied. Software combining significantly decreases the synchronization cost and outperforms many other synchronization techniques in various cases. Panagiota Fatourou, Nikolaos D. Kallimanis, Eleftherios Kosmas |
PPoPP | 2 |
| 2021 | Brief Announcement: Persistent Software CombiningabstractWe study the performance power of software combining in designing recoverable algorithms and data structures. We present two recoverable synchronization protocols, one blocking and another wait-free, which illustrate how to use software combining to achieve both low persistence and synchronization cost. Our experiments show that these protocols outperform by far state-of-the-art recoverable universal constructions and transactional memory systems. We built recoverable queues and stacks, based on these protocols, that exhibit much better performance than previous such implementations. Panagiota Fatourou, Nikolaos D. Kallimanis, Eleftherios Kosmas |
DISC | 2 |
| 2020 | The RedBlue family of universal constructions
Panagiota Fatourou, Nikolaos D. Kallimanis |
Distributed Comput. | 2 |
| 2019 | An Efficient Universal Construction for Large ObjectsabstractConcurrency has been a subject of study for more than 50 years. Still, many developers struggle to adapt their sequential code to be accessed concurrently. This need has pushed for generic solutions and specific concurrent data structures. Wait-free universal constructs are attractive as they can turn a sequential implementation of any object into an equivalent, yet concurrent and wait-free, implementation. While highly relevant from a research perspective, these techniques are of limited practical use when the underlying object or data structure is sizable. The copy operation can consume much of the CPU's resources and significantly degrade performance. To overcome this limitation, we have designed CX, a multi-instance-based wait-free universal construct that substantially reduces the amount of copy operations. The construct maintains a bounded number of instances of the object that can potentially be brought up to date. We applied CX to several sequential implementations of data structures, including STL implementations, and compared them with existing wait-free constructs. Our evaluation shows that CX performs significantly better in most experiments, and can even rival with hand-written lock-free and wait-free data structures, simultaneously providing wait-free progress, safe memory reclamation and high reader scalability. Panagiota Fatourou, Nikolaos D. Kallimanis, Eleni Kanellou |
OPODIS | 2 |
| 2018 | Mainstream vs. Emerging HPC: Metrics, Trade-Offs and Lessons LearnedabstractVarious servers with different characteristics and architectures are hitting the market, and their evaluation and comparison in terms of HPC features is complex and multidimensional. In this paper, we share our experience of evaluating a diverse set of HPC systems, consisting of three mainstream and five emerging architectures. We evaluate the performance and power efficiency using prominent HPC benchmarks, High-Performance Linpack (HPL) and High Performance Conjugate Gradients (HPCG), and expand our analysis using publicly available specialized kernel benchmarks, targeting specific system components. In addition to a large body of quantitative results, we emphasize six usually overlooked aspects of the HPC platforms evaluation, and share our conclusions and lessons learned. Overall, we believe that this paper will improve the evaluation and comparison of HPC platforms, making a first step towards a more reliable and uniform methodology. Milan Radulovic, Kazi Asifuzzaman, Darko Zivanovic, Nikola Rajovic, Guillaume Colin de Verdière, Dirk Pleiter, Manolis Marazakis, Nikolaos D. Kallimanis, Paul M. Carpenter, Petar Radojkovic, Eduard Ayguadé |
SBAC-PAD | 8 |
| 2018 | An Efficient Wait-free Resizable Hash TableabstractThis paper presents an efficient wait-free resizable hash table. To achieve high throughput at large core counts, our algorithm is specifically designed to retain the natural parallelism of concurrent hashing, while providing wait-free resizing. An extensive evaluation of our hash table shows that in the common case where resizing actions are rare, our implementation outperforms all existing lock-free hash table implementations while providing a stronger progress guarantee. Panagiota Fatourou, Nikolaos D. Kallimanis, Thomas Ropars |
SPAA | 2 |
| 2017 | Lock Oscillation: Boosting the Performance of Concurrent Data StructuresabstractIn combining-based synchronization, two main parameters that affect performance are the com- bining degree of the synchronization algorithm, i.e. the average number of requests that each com- biner serves, and the number of expensive synchronization primitives (like CAS, Swap, etc.) that it performs. The value of the first parameter must be high, whereas the second must be kept low. In this paper, we present Osci, a new combining technique that shows remarkable perform- ance when paired with cheap context switching. We experimentally show that Osci significantly outperforms all previous combining algorithms. Specifically, the throughput of Osci is higher than that of previously presented combining techniques by more than an order of magnitude. Notably, Osci’s throughput is much closer to the ideal than all previous algorithms, while keep- ing the average latency in serving each request low. We evaluated the performance of Osci in two different multiprocessor architectures, namely AMD and Intel. Based on Osci, we implement and experimentally evaluate implementations of concurrent queues and stacks. These implementations outperform by far all current state-of-the-art concur- rent queue and stack implementations. Although the current version of Osci has been evaluated in an environment supporting user-level threads, it would run correctly on any threading library, preemptive or not (including kernel threads). Panagiota Fatourou, Nikolaos D. Kallimanis |
OPODIS | 2 |
| 2017 | Lower and upper bounds for single-scanner snapshot implementations
Panagiota Fatourou, Nikolaos D. Kallimanis |
Distributed Comput. | 2 |
| 2016 | The ExaNeSt Project: Interconnects, Storage, and Packaging for Exascale SystemsabstractExaNest is one of three European projects that support a ground-breaking computing architecture for exascale-class systems built upon power-efficient 64-bit ARM processors. This group of projects share an "everything-close" and "share-anything" paradigm, which trims down the power consumption -- by shortening the distance of signals for most data transfers -- as well as the cost and footprint area of the installation -- by reducing the number of devices needed to meet performance targets. In ExaNeSt, we will design and implement: (i) a physical rack prototype and its liquid-cooling subsystem providing ultra-dense compute packaging, (ii) a storage architecture with distributed (in-node) non-volatile memory (NVM) devices, (iii) a unified, low-latency interconnect, designed to efficiently uphold desired Quality-of-Service guarantees for a mix of storage with inter-processor flows, and (iv) efficient rack-level memory sharing, where each page is cacheable at only a single node. Our target is to test alternative storage and interconnect options on actual hardware, using real-world HPC applications. The ExaNeSt consortium brings together technology, skills, and knowledge across the entire value chain, from computing IP, packaging, and system deployment, all the way up to operating systems, storage, HPC, big data frameworks, and cutting-edge applications. Manolis Katevenis, Nikolaos Chrysos, Manolis Marazakis, Iakovos Mavroidis, Fabien Chaix, Nikolaos D. Kallimanis, Javier Navaridas, John Goodacre, Piero Vicini, Andrea Biagioni, Pier Stanislao Paolucci, Alessandro Lonardo, Elena Pastorelli, Francesca Lo Cicero, Roberto Ammendola, P. Hopton, P. Coates, Giuliano Taffoni, Stefano Cozzini, Martin L. Kersten, Julio Sahuquillo, Sergio Lechago, C. Pinto, Bernd Lietzow, D. Everett, Gino Perna |
DSD | 6 |
| 2016 | Efficient Distributed Data Structures for Future Many-Core ArchitecturesabstractWe study general techniques for implementing distributed data structures on top of future many-core architectures with non cache-coherent or partially cache-coherent memory. With the goal of contributing towards what might become, in the future, the concurrency utilities package in Java collections for such architectures, we end up with a comprehensive collection of data structures by considering different variants of these techniques. To achieve scalability, we study a generic scheme which makes all our implementations hierarchical. We also describe a collection of techniques for further improving scalability in most implementations. We have performed experiments which illustrate nice scalability characteristics for some of the proposed techniques and reveal the performance and scalability power of the hierarchical approach. We distill the experimental observations into a metric that expresses the scalability potential of such implementations. We finally present experiments to study energy consumption aspects of the proposed techniques by using an energy model recently proposed for such architectures. Panagiota Fatourou, Nikolaos D. Kallimanis, Eleni Kanellou, Odysseas Makridakis, Christi Symeonidou |
ICPADS | 2 |
| 2015 | Wait-Free Concurrent Graph Objects with Dynamic TraversalsabstractGraphs are versatile data structures that allow the implementation of a variety of applications, such as computer-aided design and manufacturing, video gaming, or scientific simulations. However, although data structures such as queues, stacks, and trees have been widely studied and implemented in the concurrent context, multi-process applications that rely on graphs still largely use a sequential implementation where accesses are synchronized through the use of global locks or partitioning, thus imposing serious performance bottlenecks. In this paper we introduce an innovative concurrent graph model that provides addition and removal of any edge of the graph, as well as atomic traversals of a part (or the entirety) of the graph. We further present Dense, a concurrent graph implementation that aims at mitigating the two aforementioned implementation drawbacks. Dense achieves wait-freedom by relying on helping and provides the inbuilt capability of performing a partial snapshot on a dynamically determined subset of the graph. Nikolaos D. Kallimanis, Eleni Kanellou |
OPODIS | 1 |
| 2014 | The Power of Scheduling-Aware Synchronization
Panagiota Fatourou, Nikolaos D. Kallimanis |
DISC | 2 |
| 2014 | Highly-Efficient Wait-Free Synchronization
Panagiota Fatourou, Nikolaos D. Kallimanis |
Theory Comput. Syst. | 2 |
| 2012 | Speeding Up OpenMP Tasking
Spiros N. Agathos, Nikolaos D. Kallimanis, Vassilios V. Dimakopoulos |
Euro-Par | 2 |
| 2012 | Revisiting the combining synchronization techniqueabstractFine-grain thread synchronization has been proved, in several cases, to be outperformed by efficient implementations of the combining technique where a single thread, called the combiner, holding a coarse-grain lock, serves, in addition to its own synchronization request, active requests announced by other threads while they are waiting by performing some form of spinning. Efficient implementations of this technique significantly reduce the cost of synchronization, so in many cases they exhibit much better performance than the most efficient finely synchronized algorithms. Panagiota Fatourou, Nikolaos D. Kallimanis |
PPoPP | 2 |
| 2011 | A highly-efficient wait-free universal constructionabstractWe present a new simple wait-free universal construction, called Sim, that uses just a Fetch&Add and an LL/SC object and performs a constant number of shared memory accesses. We have implemented SIM in a real shared-memory machine. In theory terms, our practical version of SIM, called P-SIM, has worse complexity than its theoretical analog; in practice though, we experimentally show that P-SIM outperforms several state-of-the-art lock-based and lock-free techniques, and this given that it is wait-free, i.e., that it satisfies a stronger progress condition than all the algorithms it outperforms. Panagiota Fatourou, Nikolaos D. Kallimanis |
SPAA | 2 |
| 2009 | The RedBlue Adaptive Universal Constructions
Panagiota Fatourou, Nikolaos D. Kallimanis |
DISC | 2 |
| 2007 | Time-optimal, space-efficient single-scanner snapshots & multi-scanner snapshots using CASabstractSnapshots are fundamental shared objects which provide consistent views of blocks of shared memory. A snapshot object consists of an array of m memory cells and allows processes to execute UPDATES to write new values in any of the snapshot cells, and SCANS to return consistent views of all m cells. An interesting (weaker) form of snapshot with several applications is a single-scanner snapshot which allows to only one process, called scanner, to execute SCANS (UPDATES can still be executed concurrently). Panagiota Fatourou, Nikolaos D. Kallimanis |
PODC | 2 |
| 2006 | Single-scanner multi-writer snapshot implementations are fast!abstractSnapshot objects allow processes to obtain consistent global views of shared memory. A snapshot object consists of m components each capable of storing a value. The processes execute UPDATE operations to write new values in any of the components and SCANS to obtain consistent views of the snapshot contents. A single-scanner snapshot object supports only one active SCAN at any point in time (although UPDATES can still be executed concurrently).This paper studies wait-free, linearizable, single-scanner, multi-writer snapshot implementations from registers in an asynchronous shared-memory system, and presents a collection of upper and lower bounds on their complexity. We provide the first such implementations with time complexities that are (linear or quadratic) functions only of the number of snapshot components and not of the number of processes. Moreover, we argue that single-scanner implementations require at least m registers and we prove that, for such implementations which are space-optimal, SCANS execute Ω(m2) steps.Our results reveal that a lower bound derived for the multi-scanner case can be beaten when we restrict the number of concurrently active SCANS. When m is constant, our algorithms exhibit constant time complexity (while one of them requires a constant number of registers as well). For the design of our algorithms we employ new ideas, while the proof of their correctness is a complicated task. Panagiota Fatourou, Nikolaos D. Kallimanis |
PODC | 2 |