EDBT 2026 Demo / reviewers in the wild / expert
Iakovos Stamoulis
dblp:42/1530
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2025
0009-0004-4624-6089ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A CNN Compression Methodology for Layer-Wise Rank Selection Considering Inter-Layer InteractionsabstractConvolutional Neural Networks (CNNs) achieve state-of-the-art performance across various application domains but are often resource-intensive, limiting their use on resource-constrained devices. Low-rank factorization (LRF) has emerged as a promising technique to reduce the computational complexity and memory footprint of CNNs, enabling efficient deployment without significant performance loss. However, challenges still remain in optimizing the rank selection problem, balancing memory reduction and accuracy, and integrating LRF into the training process of CNNs. In this paper, a novel and generic methodology for layer-wise rank selection is presented, considering inter-layer interactions. Our approach is compatible with any decomposition method and does not require additional retraining. The proposed methodology is evaluated in thirteen widely-used, CNN models, significantly reducing model parameters and Floating-Point Operations (FLOPs). In particular, our approach achieves up to a 94.6% parameter reduction (82.3% on average) and up to 90.7% FLOPs reduction (59.6% on average), with less than a 1.5% drop in validation accuracy, demonstrating superior performance and scalability compared to existing techniques. Milad Kokhazadeh, Georgios Keramidas, Vasilios I. Kelefouras, Iakovos Stamoulis |
DATE | 4 |
| 2025 | Register Blocking: A Source-to-Source Analytical Modelling Approach for Affine Loop KernelsabstractRegister Blocking (RB), also known as ‘Register-level Tiling’ or ‘unroll-and-jam,’ is a key compiler optimization for developing efficient micro-kernels. However, applying RB effectively is a complex task due to several challenges. First, the exploration space of possible RB configurations is vast. Second, RB and loop permutation are interdependent; therefore, addressing both optimizations simultaneously further inflates the exploration space. Third, the effectiveness of RB is highly dependent on the target hardware platform and the specific loop kernel being optimized. As a result, an extensive and time-consuming fine-tuning process is necessary for achieving an efficient implementation. To address these challenges, a source-to-source analytical modelling approach is proposed. The RB factors, the loops to apply RB, the number of allocated variables/registers per array reference, and the loops’ ordering are generated by an analytical model, leveraging the target hardware architecture details and loop kernel characteristics. The proposed methodology has been evaluated on both embedded and general-purpose CPUs, using seven well-known loop kernels and three machine learning applications. The results show significant speedups over the GCC compiler, the Pluto tool, and related work. Theologos Anthimopoulos, Georgios Keramidas, Vasilios I. Kelefouras, Iakovos Stamoulis |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2024 | Register Blocking: An Analytical Modelling Approach for Affine Loop KernelsabstractFor the past several decades, optimizing compilers have been a primary area of focus in both industry and academia. This continued research interest is a testament to the complexity of this task, primarily stemming from the vast number of parameters that must be explored to attain near-optimal results. One of the key compiler optimizations is "Register Blocking (RB)" also known as "Register-level Tiling" or "unroll-and-jam". RB can strongly reduce the number of executed Load/Store (L/S) instructions, and as a consequence the number of data accesses in memory hierarchy, but due to its inherent complexities, fine-tuning is essential for its effective implementation. To address this problem, in this work a new methodology is proposed for RB. The RB factors, the loops to apply RB, the number of allocated variables/registers per array reference, and the loops' ordering are generated by an analytical model, leveraging the target hardware (HW) architecture details and loop kernel characteristics. The proposed methodology has been evaluated on both embedded and general-purpose CPUs across seven well-known loop kernels, achieving high speedups and L/S instruction gains over GCC compiler, handwritten optimized codes, and the popular Pluto tool. Theologos Anthimopoulos, Georgios Keramidas, Vasilios I. Kelefouras, Iakovos Stamoulis |
CF | 4 |
| 2024 | Denseflex: A Low Rank Factorization Methodology for Adaptable Dense Layers in DNNsabstractLow-Rank Factorization (LRF) is a popular compression technique used in Deep Neural Networks (DNNs). LRF can reduce both the memory size and the arithmetic operations in a DNN layer by approximating a weight tensor/matrix by two or more smaller tensors/matrices. Employing LRF to DNN is a challenging task for several reasons. First, the exploration space is massive and different solutions provide different trade-offs among memory, FLOPs, inference time, and validation accuracy; second, multiple DNN layers and multiple LRF algorithms must be considered; third, every extracted solution must undergo through a calibration phase and this makes the LRF process time-consuming. In this paper, a methodology, called Denseflex, is presented that formulates the LRF problem as an inference time vs. FLOPs vs. memory vs. validation accuracy Design Space Exploration (DSE) problem. Moreover, to the best of our knowledge, this is the first work that proposes a methodology to efficiently combine two different LRF methods (Singular Value Decomposition -SVD- and Tensor Train Decomposition -TTD-) in the same framework. Denseflex is formulated as a design tool in which the user can provide specific memory, FLOPs, and/or execution time constraints and the tool will output a set of solutions that meet the given constraints avoiding the time-consuming re-training phases. Our results indicate that our approach is able to prune the design space by 62% (on average) over related works for nine DNN models (up to 88% in AlexNet), while the extracted LRF solutions exhibit both lower memory footprints and lower execution times compared to the initial model. Milad Kokhazadeh, Georgios Keramidas, Vasilios I. Kelefouras, Iakovos Stamoulis |
CF | 4 |
| 2021 | Architectures for SLAM and Augmented Reality ComputingabstractIn the next few years, new demanding applications will be supported on mobile platforms by reconciling two conflicting requirements: high performance (often with real-time limitations) and low power consumption. The objective of the vipGPU project is to develop hardware and software technology to provide efficient support for two such application scenarios, namely (a) simultaneous localization and mapping (SLAM) in mobile robotics systems, and (b) virtual reality (VR) in portable devices to simulate serious games with emphasis on simulating surgical interventions and medical training in general. In this project, we aim at developing a new heterogeneous platform consisting of hardware accelerators for low power embedded systems optimized (at the hardware and software level) for the implementation of the two applications mentioned above. Nikolaos Bellas, Christos D. Antonopoulos, Spyros Lalis, Maria Rafaela Gkeka, Alexandros Patras, Georgios Keramidas, Iakovos Stamoulis, Nikolaos Tavoularis, Stylianos Piperakis, Emmanouil Hourdakis, Panos E. Trahanias, Paul Zikas, George Papagiannakis, Ioanna Kartsonaki |
FPL | 7 |