EDBT 2026 Demo / reviewers in the wild / expert
Ioannis Vardas
dblp:277/3290 · also Giannis Vardas
· DBLP profile ↗
6ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0001-5461-556XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mpisee: Communicator-Centric Profiling of MPI ApplicationsabstractABSTRACT mpisee is a lightweight profiling tool designed to track MPI communication operations per communicator, providing fine‐grained insights into MPI applications that use communicators to partition MPI communication. While existing profiling tools offer valuable information, they may limit detailed analysis and optimization for such MPI applications, as they do not associate MPI communication with their communicator. Additionally, mpisee categorizes MPI communication operations based on message size, offering more granular information. It uses an SQLite database to efficiently store the profiling data, enabling users to analyze the application's profile from various perspectives, focusing on specific MPI ranks, operations, and more. Our analysis shows that mpisee incurs less than 5% overhead, performing on par with other state‐of‐the‐art profilers. We demonstrate mpisee 's effectiveness by profiling and analyzing an FFT application, revealing potential performance bottlenecks related to the MPI_Alltoallv collective operation on small communicators and insights not available by other profilers. Leveraging this detailed information, we improved the application's overall performance by selecting different algorithms for MPI_Alltoallv and measuring their performance on different communicators with mpisee . This study illustrates mpisee 's utility and highlights the significant advantages of a communicator‐centric approach in MPI profiling. Ioannis Vardas, Jesper Larsson Träff, Ruben Laso, Sascha Hunold |
Concurr. Comput. Pract. Exp. | 1 |
| 2024 | Improved Parallel Application Performance and Makespan by Colocation and Topology-aware Process MappingabstractIn modern, deeply hierarchical HPC systems shared resource congestion can hinder the efficient use of many cores by parallel applications and degrade performance. Such congestion is often caused when parallel processes within an application that execute similar operations share the same resources. Previous research suggests using fewer cores with better process-to-core mapping can improve applications’ performance but leaves many cores unused. To utilize these cores, we colocate additional applications and map them using a topology-aware process-to-core, application-agnostic mapping algorithm. We show that these mappings significantly impact memory bandwidth and communication latency. We evaluate our approach using eight parallel applications on an HPC system with 128-core nodes, demonstrating the performance effects of mappings combined with colocation. Our goal is to determine whether colocation with topology-aware mapping is a viable alternative to typical exclusive node allocation. Our results show makespan improvements of 2.4x over exclusive allocation in an HPC system, demonstrating the potential benefits of colocation with optimized mappings. Ioannis Vardas, Sascha Hunold, Philippe Swartvagher, Jesper Larsson Träff |
CCGrid | 1 |
| 2023 | Uniform Algorithms for Reduce-scatter and (most) other Collectives for MPIabstractWe explore the use of a regular, circulant graph communication pattern for the implementation of the reduction-to-all (MPI_Allreduce), by specialization the reduction-to-root (MPI_Reduce), the reduce-scatter (MPI_Reduce_scatter_block), the all-to-all-broadcast (MPI_Allgather) and the rooted gather and scatter (MPI_Gather and MPI_Scatter) collective operations, all as found in MPI (the Message-Passing Interface), for commutative operators and for any number of processes. The reduction-to-all algorithm reconstructs the little known algorithm by Bar-Noy, Kipnis and Schieber (1993), which the paper considerably extends.We experiment with extensions and combinations of the algorithms for these operations, and examine their performance from the perspective of performance guidelines, and in direct comparison to the implementations in common MPI libraries. On a small cluster with 36 × 32 cores and two larger HPC production systems, we show that we can especially for MPI_Reduce_scatter_block achieve considerably better performance than standard MPI library implementations. Our algorithms can perform consistently, which the implementations in standard MPI libraries sometimes do not.In a homogeneous, one-ported communication system with linear transmission costs, reduction-to-all, reduce-scatter and all-to-all-broadcast can all be implemented in O(log p + m) time steps for problems of size m with small constants which we analyze and discuss. Jesper Larsson Träff, Sascha Hunold, Ioannis Vardas, Nikolaus Manes Funk |
CLUSTER | 3 |
| 2023 | Library Development with MPI: Attributes, Request Objects, Group Communicator Creation, Local Reductions, and DatatypesabstractA major design objective of MPI is to enable support for the construction of safe parallel libraries that can be used and mixed freely in complex applications. In this respect, MPI has been extremely successful; but may nevertheless lack elementary supporting functionality for some situations, and may have made design choices that are difficult to accommodate in certain libraries. We discuss several cases of library construction requiring different kinds of supporting MPI functionality, and propose concrete improvements for library implementations and future MPI versions to alleviate the problems that were encountered. Specifically, we pinpoint (performance) issues with MPI object attributes, caching and lookup, request objects, partly collective and non-blocking communicator creation, process local reductions, type correct process local copying, and user-defined datatypes. Jesper Larsson Träff, Ioannis Vardas |
EuroMPI | 2 |
| 2020 | Towards Communication Profile, Topology and Node Failure Aware Process PlacementabstractHPC systems need to keep growing in size to meet the ever-increasing demand for high levels of capability and capacity, often in tight time windows for urgent computation. However, increasing the size, complexity and heterogeneity of HPC systems also increases the risk and impact of system failures, that result in resource waste and aborted jobs. A major contributor to job completion time is the cost of interprocess communication. To address performance and energy efficiency, several prior studies have targeted improvements of communication locality. To meet this goal, they derive a mapping of MPI processes to system nodes in a way that reduces communication cost. However, such approaches disregard the effect of system failures. In this work, we propose a resource allocation approach for MPI jobs, considering both high performance and error resilience. Our approach, named Communication Profile, Topology and node Failure (CPTF), takes into account the application's communication profile, system topology and node failure probability for assigning job processes to nodes. We evaluate variants of CPTF through simulations of two MPI applications, one with a regular communication pattern (LAMMPS) and one with an irregular one (NPB-DT). In both cases, the variant of CPTF that strives to avoid failure-prone nodes and communication paths achieves lower time to complete job batches when compared to the default resource allocation policy of Slurm. It also exhibits the lowest ratio of aborted jobs. The average improvement in batch completion time is 67% for NPB-DT and 34% for LAMMPS. Ioannis Vardas, Manolis Ploumidis, Manolis Marazakis |
SBAC-PAD | 1 |
| 2018 | Accurate Congestion Control for RDMA TransfersabstractHigh-performance interconnects need congestion control to deal with traffic bursts. In this paper, we propose ACCurate, a congestion control protocol that assigns exact max-min fair rates to flows, without relying on costly per-flow state inside the network. ACCurate keeps the backlogs outside of the network, protects innocent flows, and promptly recovers the flows' rates after congestive episodes. Comparisons with TCP and PAUSE-only RDMA under datacenter-resembling workloads further show that ACCurate provides up to 10× faster flow completion times. ACCurate relies on simple hardware that can be readily implemented inside switches. In our implementation, the additional circuitry needed in a 16×16 switch occupies less than 2% of FPGA resources. Dimitris Giannopoulos, Nikolaos Chrysos, Evangelos Mageiropoulos, Ioannis Vardas, Leandros Tzanakis, Manolis Katevenis |
NOCS | 4 |