EDBT 2026 Demo / reviewers in the wild / expert
Petros Anastasiadis
dblp:291/5114
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0001-7821-3610ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PaHiS: A Hierarchical Synchronous Parallel Model for Irregular WorkloadsabstractLarge-scale solvers are critical in scientific and engineering domains, where achieving portable and near-optimal performance remains a challenge. Existing high-level parallel performance models provide coarse-grain communication cost estimates but fail to capture the hierarchical memory behavior and irregular access patterns prevalent in modern sparse solver workloads. Conversely, sparse modeling techniques focus on low-level kernel details and lack support for solver-level reasoning. To bridge this gap, we introduce PaHiS, a hierarchical cost model that represents hardware and algorithms through a multi-level abstraction of their communication, computation, and synchronization characteristics. We integrate this model into a GraphBLAS library and couple it with an automated microbenchmarking pipeline for hardware characterization, enabling solver cost estimation and thread autotuning for algorithms expressed in GraphBLAS. We evaluate the model on three CPU architectures for two solvers, demonstrating i) strong correlation between predicted and measured performance, ii) high-performance thread-level autotuning in ALP GraphBLAS, and iii) reliable guidance for high-level HW-SW co-design decisions such as hardware selection and algorithm comparison. Petros Anastasiadis, Denis Jelovina, Albert-Jan Nicholas Yzelman |
SPAA | 1 |
| 2025 | DIV: An Index & Value compression method for SpMV on large matricesabstractSpMV on large matrices is a heavily memory-bound kernel, a characteristic attributed to its extremely low computational intensity.To address this, research has mainly focused on compressing the matrix indices.Nevertheless, the values of a matrix usually occupy up to two thirds of the total size.Research on value compression, on the other hand, has been limited to specific matrix types.In this paper, we propose DIV, a combined index and value lossless compression scheme, based on variations of delta and run-length encoding, that achieves substantially improved SpMV performance for large matrices, i.e., those that exceed the CPU cache.We evaluate its performance against other state-of-the-art matrix formats, on an Intel Xeon and an AMD EPYC platform.Our format achieves 77% and 115% geometric mean speedup respectively versus the Intel MKL library.We finally demonstrate the applicability of DIV on a Biconjugate Gradient Stabilized solver, where we also achieve significant speedups. Dimitrios Galanopoulos, Panagiotis Mpakos, Petros Anastasiadis, Nectarios Koziris, Georgios I. Goumas |
ICS | 3 |
| 2024 | Uncut-GEMMs: Communication-Aware Matrix Multiplication on Multi-GPU NodesabstractGeneral Matrix Multiplication (GEMM) is one of the most common kernels in high-performance computing (HPC) and machine-learning (ML) applications, frequently dominating their execution time, rendering its performance vital. As multi-GPU nodes have become common in modern HPC systems, GEMM is usually offloaded on GPUs as its compute-intensive nature is a good match for their architecture. On the other hand, despite the GEMM kernel itself being usually compute-bound, execution on multi-GPU systems also requires fine-grained communication and task scheduling to achieve optimal performance. While numerous multi-GPU level-3 BLAS libraries have faced these issues in the past, they are bound by older design concepts that are not necessarily applicable to modern multi-GPU clusters, resulting in considerable deviation from peak performance. In this work, we thoroughly analyze the current challenges regarding data movement, caching, and overlap of multi-GPU GEMM, and the shortcomings of previous solutions, and provide a fresh approach to multi-GPU GEMM optimization. We devise a static scheduler for GEMM, enabling a variety of algorithmic, communication, and auto-tuning optimizations, and integrate those in an end-to-end open-source multi-GPU GEMM library. Our library is evaluated on a multi-GPU NVIDIA HGX system with 8 NVIDIA A100 GPUs, achieving on average a 1.37x and 1.29x performance improvement over the state-of-the-art multi-GPU GEMM libraries, for double and single precision, respectively. Petros Anastasiadis, Nikela Papadopoulou, Nectarios Koziris, Georgios I. Goumas |
CLUSTER | 1 |
| 2024 | Large-Scale Parallelization of Human Migration SimulationabstractForced displacement of people worldwide, for example, due to violent conflicts, is common in the modern world, and today more than 82 million people are forcibly displaced. This puts the problem of migration at the forefront of the most important problems of humanity. The Flee simulation code is an agent-based modeling tool that can forecast population displacements in civil war settings, but performing accurate simulations requires nonnegligible computational capacity. In this article, we present our approach to Flee parallelization for fast execution on multicore platforms, as well as discuss the computational complexity of the algorithm and its implementation. We benchmark parallelized code using supercomputers equipped with AMD EPYC Rome 7742 and Intel Xeon Platinum 8268 processors and investigate its performance across a range of alternative rule sets, different refinements in the spatial representation, and various numbers of agents representing displaced persons. We find that Flee scales excellently to up to 8192 cores for large cases, although very detailed location graphs can impose a large initialization time overhead. Derek Groen, Nikela Papadopoulou, Petros Anastasiadis, Marcin Lawenda, Lukasz Szustak, Sergiy Gogolenko, Hamid Arabnejad, Alireza Jahani |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2023 | Feature-based SpMV Performance Analysis on Contemporary DevicesabstractThe SpMV kernel is characterized by high performance variation per input matrix and computing platform. While GPUs were considered State-of-the-Art for SpMV, with the emergence of advanced multicore CPUs and low-power FPGA accelerators, we need to revisit its performance and energy efficiency. This paper provides a high-level SpMV performance analysis based on structural features of matrices related to common bottlenecks of memory-bandwidth intensity, low ILP, load imbalance and memory latency overheads. Towards this, we create a wide artificial matrix dataset that spans these features and study the performance of different storage formats in nine modern HPC platforms; five CPUs, three GPUs and an FPGA. After validating our proposed methodology using real-world matrices, we analyze our extensive experimental results and draw key insights on the competitiveness of different target architectures for SpMV and the impact of each feature/bottleneck on its performance. Panagiotis Mpakos, Dimitrios Galanopoulos, Petros Anastasiadis, Nikela Papadopoulou, Nectarios Koziris, Georgios I. Goumas |
IPDPS | 3 |
| 2023 | PARALiA: A Performance Aware Runtime for Auto-tuning Linear Algebra on Heterogeneous SystemsabstractDense linear algebra operations appear very frequently in high-performance computing (HPC) applications, rendering their performance crucial to achieve optimal scalability. As many modern HPC clusters contain multi-GPU nodes, BLAS operations are frequently offloaded on GPUs, necessitating the use of optimized libraries to ensure good performance. Unfortunately, multi-GPU systems are accompanied by two significant optimization challenges: data transfer bottlenecks as well as problem splitting and scheduling in multiple workers (GPUs) with distinct memories. We demonstrate that the current multi-GPU BLAS methods for tackling these challenges target very specific problem and data characteristics, resulting in serious performance degradation for any slightly deviating workload. Additionally, an even more critical decision is omitted because it cannot be addressed using current scheduler-based approaches: the determination of which devices should be used for a certain routine invocation. To address these issues we propose a model-based approach: using performance estimation to provide problem-specific autotuning during runtime. We integrate this autotuning into an end-to-end BLAS framework named PARALiA. This framework couples autotuning with an optimized task scheduler, leading to near-optimal data distribution and performance-aware resource utilization. We evaluate PARALiA in an HPC testbed with 8 NVIDIA-V100 GPUs, improving the average performance of GEMM by 1.7× and energy efficiency by 2.5× over the state-of-the-art in a large and diverse dataset and demonstrating the adaptability of our performance-aware approach to future heterogeneous systems. Petros Anastasiadis, Nikela Papadopoulou, Georgios I. Goumas, Nectarios Koziris, Dennis Hoppe, Li Zhong 0008 |
ACM Trans. Archit. Code Optim. | 1 |
| 2021 | CoCoPeLia: Communication-Computation Overlap Prediction for Efficient Linear Algebra on GPUsabstractGraphics Processing Units (GPUs) are well established in HPC systems and frequently used to accelerate linear algebra routines. Since data transfers pose a severe bottleneck for GPU offloading, modern GPUs provide the ability to overlap communication with computation by splitting the problem to fine-grained sub-kernels that are executed in a pipelined manner. This optimization is currently underutilized by GPU BLAS libraries, since it requires an approach to select an efficient tiling size, which in turn leads to a challenging problem that needs to consider routine, system, data, and problem-specific characteristics. In this work, we introduce an elaborate 3-way concurrency model for GPU BLAS offload time that considers previously neglected features regarding data access and machine behavior. We then incorporate our model in an automated, end-to-end framework (called CoCoPeLia) that supports overlap prediction, tile selection and effective tile scheduling. We validate our model's efficacy for dgemm, sgemm, and daxpy on two testbeds, with our experimental results showing that it achieves significantly lower prediction error than previous models and provides near-optimal tiling sizes for all problems. We also demonstrate that CoCoPeLia leads to considerable performance improvements compared to the state of the art BLAS routine implementations for GPUs. Petros Anastasiadis, Nikela Papadopoulou, Georgios I. Goumas, Nectarios Koziris |
ISPASS | 1 |