Konstantinos Iliakis

dblp:228/3827 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0002-1403-6851ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Multi-Partner Project: Multi-GPU Performance Portability Analysis for CFD Simulations at Scale
abstract
As heterogeneous supercomputing architectures leveraging GPUs become increasingly central to high-performance computing (HPC), it is crucial for computational fluid dynamics (CFD) simulations, a de-facto HPC workload, to efficiently utilize such hardware. One of the key challenges of HPC codes is performance portability, i.e. the ability to maintain near-optimal performance across different accelerators. In the context of the REFMAP project, which targets scalable, GPU-enabled multi-fidelity CFD for urban airflow prediction, this paper analyzes the performance portability of SOD2D, a state-of-the-art Spectral Elements simulation framework across AMD and NVIDIA GPU architectures. We first discuss the physical and numerical models underlying SOD2D, highlighting its computational hotspots. Then, we examine its performance and scalability in a multi-level manner, i.e. defining and characterizing an extensive full-stack design space spanning across application, software and hardware infrastructure related parameters. Single-GPU performance characterization across server-grade NVIDIA and AMD GPU architectures and vendor-specific compiler stacks, show the potential as well as the diverse effect of memory access optimizations, i.e. 0.69× - 3.91× deviations in acceleration speedup. Performance variability of SOD2D at scale is further examined on the LUMI multi-GPU cluster, where profiling reveals similar throughput variations, highlighting the limits of performance projections and the need for multi-level, informed tuning.
Panagiotis-Eleftherios Eleftherakis, George Anagnostopoulos, Anastassis Kapetanakis, Mohammad Umair, Jean-Yves Vet, Konstantinos Iliakis, Jonathan Vincent, Akshay Patil, Clara García-Sánchez, Gerardo Zampino, Ricardo Vinuesa, Sotirios Xydis
DATE6
2026 sCROOGe: Circuit-level Design and Optimization Framework for RISC-V Out-of-Order GPUs
Maria Zerva, Panagiotis-Eleftherios Eleftherakis, Alexis Maras, Konstantinos Iliakis, Alexandros Moiras, Sotirios Xydis
ISCA4
2026 AccelHSA: Modeling Single-ISA Heterogeneous GPU Architectures
abstract
Initially branded as dedicated graphics processing accelerators, GPUs now find applications in an ever-growing range of domains, including artificial intelligence, high-performance computing, self-driving vehicles, and bioinformatics. However, this diversity comes at the cost of reduced resource efficiency and micro-architectural affinity. Evidently, the homogeneity of the GPU hardware struggles to cope with the vast heterogeneity of GPU applications. Motivated by the aforementioned observations, this article introduces the concept of Single-ISA Heterogeneous GPU architectures. In order to explore the efficiency of the new GPU architectural paradigm, we extend Accel-Sim, the state-of-the-art, cycle-accurate GPU simulator to support single-ISA heterogeneous cores within the GPU chip. The proposed implementation, called AccelHSA, supports independently tuning the micro-architectural characteristics of the cores, unlocking a wide design space. The CUDA API is extended to allow control of the kernel-to-core-type mapping along with a newly developed kernel launching model that supports concurrent execution, aimed at, albeit not limited to, the context of the simulator. We showcase the impact of single-ISA heterogeneous GPU architectures via a case study targeting the collocation of resource sensitive and insensitive HPC kernels. Finally, the heterogeneous GPU architecture is evaluated against homogeneous GPU baselines, demonstrating a 27.07% average speedup with a marginal 0.47% area overhead.
Alexandros Moiras, Konstantinos Iliakis, Dimitrios Soudris, Sotirios Xydis
ACM Trans. Archit. Code Optim.2
2025 POSTER: Performance Portability in GPU-Accelerated Spectral Finite Element Fluid Simulations: A Cross-layer Exploration Approach
abstract
As heterogeneous supercomputing architectures leveraging GPUs become increasingly central to high-performance computing (HPC), it is crucial for computational fluid dynamics (CFD) simulations to maintain performance portability.In this paper, we examine the performance and scalability of CFD framework SOD2D in a crosslayer manner, i.e. across application, software and hardware infrastructure related parameters.Single-GPU performance characterization across server-grade NVIDIA and AMD GPU architectures and vendor-specific compiler stacks, show the potential as well as the diverse effect of memory access optimizations, i.e. 0.69× -3.96× deviations in acceleration speedup.Performance variability of SOD2D at scale is then further examined on the LUMI multi-GPU cluster, showcasing analogous diverse effects on throughput, demonstrating the ineffectiveness of adopting performance projections, thus underscoring the importance and necessity of cross-layer informed performance analysis and tuning for multi-GPU configurations.
Panagiotis-Eleftherios Eleftherakis, George Anagnostopoulos, Anastassis Kapetanakis, Mohammad Umair, Jean-Yves Vet, Konstantinos Iliakis, Jonathan Vincent, Ricardo Vinuesa, Sotirios Xydis
CF6
2024 GhOST: a GPU Out-of-Order Scheduling Technique for Stall Reduction
abstract
Graphics Processing Units (GPUs) use massive multi-threading coupled with static scheduling to hide instruction latencies. Despite this, memory instructions pose a challenge as their latencies vary throughout the application’s execution, leading to stalls. Out-of-order (OoO) execution has been shown to effectively mitigate these types of stalls. However, prior OoO proposals involve costly techniques such as reordering loads and stores, register renaming, or two-phase execution, amplifying implementation overhead and consequently creating a substantial barrier to adoption in GPUs. This paper introduces GhOST, a minimal yet effective OoO technique for GPUs. Without expensive components, GhOST can manifest a substantial portion of the instruction reorderings found in an idealized OoO GPU. GhOST leverages the decode stage’s existing pool of decoded instructions and the existing issue stage’s information about instructions in the pipeline to select instructions for OoO execution with little additional hardware. A comprehensive evaluation of GhOST and the prior state-of-the-art OoO technique across a range of diverse GPU benchmarks yields two surprising insights: (1) Prior works utilized Nvidia’s intermediate representation PTX for evaluation; however, the optimized static instruction scheduling of the final binary form negates many purported improvements from OoO execution; and (2) The prior state-of-the-art OoO technique results in an average slowdown across this set of benchmarks. In contrast, GhOST achieves a $\mathbf{3 6 \%}$ maximum and $6.9 \%$ geometric mean speedup on GPU binaries with only a $0.007 \%$ area increase, surpassing previous techniques without slowing down any of the measured benchmarks.
Ishita Chaturvedi, Bhargav Reddy Godala, Yucan Wu, Konstantinos Iliakis, Panagiotis-Eleftherios Eleftherakis, Sotirios Xydis, Dimitrios Soudris, Tyler Sorensen 0001, Simone Campanoni, Tor M. Aamodt, David I. August
ISCA5
2022 Enabling Large Scale Simulations for Particle Accelerators
abstract
International high-energy particle physics research centers, like CERN and Fermilab, require excessive studies and simulations to plan for the upcoming upgrades of the world's largest particle accelerators, and the design of future machines given the technological challenges and tight budgetary constraints. The Beam Longitudinal Dynamics (BLonD) simulator suite incorporates the most detailed and complex physics phenomena in the field of longitudinal beam dynamics, required for providing extremely accurate predictions. Modern challenges in beam dynamics dictate for longer, larger and numerous simulation studies to draw meaningful conclusions that will drive the baseline choices for the daily operation of current machines and the design choices of future projects. These studies are extremely time consuming, and would be impractical to perform without a High-Performance Computing oriented simulator framework. In this article, at first, we design and evaluate a highly-optimized distributed version of BLonD. We combine approximate computing techniques, and leverage a dynamic load-balancing scheme to relax synchronization and improve scalability. In addition, we employ GPUs to accelerate the distributed implementation. We evaluate the highly optimized distributed beam longitudinal dynamics simulator in a supercomputing system and demonstrate speedups of more than two orders of magnitude when run on 32 GPU platforms, w.r.t. the previous state-of-art. By driving a wide range of new studies, the proposed high performance beam longitudinal dynamics simulator forms an invaluable tool for accelerator physicists.
Konstantinos Iliakis, Helga Timko, Sotirios Xydis, Panagiotis Tsapatsaris, Dimitrios Soudris
IEEE Trans. Parallel Distributed Syst.1
2022 Repurposing GPU Microarchitectures with Light-Weight Out-Of-Order Execution
abstract
GPU is the dominant platform for accelerating general-purpose workloads due to its computing capacity and cost-efficiency. GPU applications cover an ever-growing range of domains. To achieve high throughput, GPUs rely on massive multi-threading and fast context switching to overlap computations with memory operations. We observe that among the diverse GPU workloads, there exists a significant class of kernels that fail to maintain a sufficient number of active warps to hide the latency of memory operations, and thus suffer from frequent stalling. We argue that the dominant Thread-Level Parallelism model is not enough to efficiently accommodate the variability of modern GPU applications. To address this inherent inefficiency, we propose a novel micro-architecture with lightweight Out-Of-Order execution capability enabling Instruction-Level Parallelism to complement the conventional Thread-Level Parallelism model. To minimize the hardware overhead, we carefully design our extension to highly re-use the existing micro-architectural structures and study various design trade-offs to contain the overall area and power overhead, while providing improved performance. We show that the proposed architecture outperforms traditional platforms by 23 percent on average for low-occupancy kernels, with an area and power overhead of 1.29 and 10.05 percent, respectively. Finally, we establish the potential of our proposal as a micro-architecture alternative by providing 16 percent speedup over a wide collection of 60 general-purpose kernels.
Konstantinos Iliakis, Sotirios Xydis, Dimitrios Soudris
IEEE Trans. Parallel Distributed Syst.1
2020 Scale-out beam longitudinal dynamics simulations
abstract
Excessive studies and simulations are required to plan for the upcoming upgrades of the world's largest particle accelerators, and the design of future machines, given the technological challenges and tight budgetary constraints. The Beam Longitudinal Dynamics (BLonD) simulator suite incorporates the most detailed and complex physics phenomena in the field of longitudinal beam dynamics, required for providing extremely accurate predictions. These predictions are invaluable to the operation of existing accelerators, upcoming upgrades, and future studies. To undertake this agenda, and enable for the first time scale-out beam longitudinal dynamics simulations, we implement Hybrid-BLond, a distributed version of BLonD, that efficiently combines horizontal and vertical scaling. We propose a series of techniques that minimize the inter-node communication overhead and improve scalability. Firstly, we exploit mixed data and task parallelism opportunities. Secondly, we discuss two traffic optimisation techniques motivated by the properties of the simulated physics phenomena. Finally, we build a dynamic load-balancing scheme that coordinates effectively all the above features. We evaluate experimentally Hybrid-BLonD in an HPC cluster built with cutting-edge Intel servers and Infiniband interconnection network. Our fully-optimised implementation demonstrates an average 25.7X speedup over the previous state-of-the-art simulator when run on 32 computing nodes, across three real-world testcases.
Konstantinos Iliakis, Helga Timko, Sotirios Xydis, Dimitrios Soudris
CF1
2020 Resource-Aware MapReduce Runtime for Multi/Many-core Architectures
abstract
Modern multi/many-core processors exhibit high integration densities, e.g. up to several dozens or hundreds of cores. To ease the application development burden for such systems, various programming frameworks have emerged. The MapReduce programming model, after having demonstrated its usability in the area of distributed systems, has been adapted to the needs of shared-memory many-core and multi-processor systems, showing promising results in comparison with conventional multi-threaded libraries, e.g. pthreads. In this paper, we propose a novel resource-aware MapReduce architecture. The proposed runtime decouples map and combine phases in order to enhance the parallelism degree, while it effectively overlaps the memory-intensive combine with the compute-intensive map operation resulting in superior resource utilization and performance improvements. A detailed sensitivity analysis to the framework's tuning knobs is provided. The decoupled MapReduce architecture is evaluated against the state-of-art library into two diverse systems, i.e. a Haswell server and a Xeon Phi co-processor, demonstrating speedups on average up-to 2.2x and 2.9x respectively.
Konstantinos Iliakis, Sotirios Xydis, Dimitrios Soudris
DATE1