EDBT 2026 Demo / reviewers in the wild / expert
Georgios Keramidas
dblp:65/1103 · also Giorgos Keramidas
· DBLP profile ↗
38ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0003-0460-6061ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 33 · 6 first-author · 18 since 2021Software engineering, systems software and programming languages · 10 · 2 first-author · 5 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sensor Placement and Transformer-Based Thermal Map Generation for Reusable Interposersabstract2.5D integration has been a promising packaging approach intrinsically underpinning heterogeneous integration. The physical proximity of diverse components (e.g., chiplets) on interposers entails multi-physics, including thermal coupling, which affects the performance and reliability of the entire system. Consequently, interposer-level thermal monitoring is required to avoid overheating during run-time. Furthermore, reusable interposers have also recently been proposed in the literature, implying that a specific interposer is used for multiple systems. Therefore, conventional thermal sensor placement methods, developed for a specific system, are incompatible with this emerging design concept. A new flow focusing on thermal sensor allocation and thermal map reconstruction for reusable interposers is proposed. The flow utilizes a transformer neural network to reconstruct the thermal map of the interposer and hyperparameter tuning to select the appropriate thermal sensor locations that minimize the reconstruction error across the entire set of available floorplans for a specific transformer architecture. The benchmarks used to train the transformer are produced through gem5, McPat, HotSpot and TAP-2.5D for ten different floorplans, showcasing the effectiveness and generality of the approach compared with prior art and achieving an average maximum error of less than 1K. Aristotelis Tsekouras, Theodoros Papavasileiou, Panagiotis Petrantonakis, Georgios Keramidas, Vasilis F. Pavlidis |
DATE | 4 |
| 2026 | Project Highlights - Reliability Evaluation for ARCHYTAS AI hardware accelerators
Angeliki Kritikakou, Fernando Santos 0001, Marcello Traiola, Rafael Billig Tonetto, Olivier Sentieys, Paolo Rech, Haralampos-G. D. Stratigopoulos, Georgios Keramidas |
IOLTS | 8 |
| 2025 | A CNN Compression Methodology for Layer-Wise Rank Selection Considering Inter-Layer InteractionsabstractConvolutional Neural Networks (CNNs) achieve state-of-the-art performance across various application domains but are often resource-intensive, limiting their use on resource-constrained devices. Low-rank factorization (LRF) has emerged as a promising technique to reduce the computational complexity and memory footprint of CNNs, enabling efficient deployment without significant performance loss. However, challenges still remain in optimizing the rank selection problem, balancing memory reduction and accuracy, and integrating LRF into the training process of CNNs. In this paper, a novel and generic methodology for layer-wise rank selection is presented, considering inter-layer interactions. Our approach is compatible with any decomposition method and does not require additional retraining. The proposed methodology is evaluated in thirteen widely-used, CNN models, significantly reducing model parameters and Floating-Point Operations (FLOPs). In particular, our approach achieves up to a 94.6% parameter reduction (82.3% on average) and up to 90.7% FLOPs reduction (59.6% on average), with less than a 1.5% drop in validation accuracy, demonstrating superior performance and scalability compared to existing techniques. Milad Kokhazadeh, Georgios Keramidas, Vasilios I. Kelefouras, Iakovos Stamoulis |
DATE | 2 |
| 2025 | Optimizing Tensor Train Decomposition in DNNs for RISC-V Architectures Using Design Space Exploration and Compiler OptimizationsabstractDeep neural networks (DNNs) have become indispensable in many real-life applications like natural language processing, and autonomous systems. However, deploying DNNs on resource-constrained devices, e.g., in RISC-V platforms, remains challenging due to the high computational and memory demands of fully connected (FC) layers, which dominate resource consumption. Low-rank factorization (LRF) offers an effective approach to compressing FC layers, but the vast design space of LRF solutions involves complex tradeoffs among FLOPs, memory size, inference time, and accuracy, making the LRF process complex and time-consuming. This article introduces an end-to-end LRF design space exploration methodology and a specialized design tool for optimizing FC layers on RISC-V processors. Using Tensor Train Decomposition (TTD) offered by TensorFlow T3F library, the proposed work prunes the LRF design space by excluding first, inefficient decomposition shapes and second, solutions with poor inference performance on RISC-V architectures. Compiler optimizations are then applied to enhance custom T3F layer performance, minimizing inference time and boosting computational efficiency. On average, our TT-decomposed layers run 3× faster than IREE and 8× faster than Pluto on the same compressed model. This work provides an efficient solution for deploying DNNs on edge and embedded devices powered by RISC-V architectures. Theologos Anthimopoulos, Milad Kokhazadeh, Vasilios I. Kelefouras, Benjamin Himpel, Georgios Keramidas |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2025 | Register Blocking: A Source-to-Source Analytical Modelling Approach for Affine Loop KernelsabstractRegister Blocking (RB), also known as ‘Register-level Tiling’ or ‘unroll-and-jam,’ is a key compiler optimization for developing efficient micro-kernels. However, applying RB effectively is a complex task due to several challenges. First, the exploration space of possible RB configurations is vast. Second, RB and loop permutation are interdependent; therefore, addressing both optimizations simultaneously further inflates the exploration space. Third, the effectiveness of RB is highly dependent on the target hardware platform and the specific loop kernel being optimized. As a result, an extensive and time-consuming fine-tuning process is necessary for achieving an efficient implementation. To address these challenges, a source-to-source analytical modelling approach is proposed. The RB factors, the loops to apply RB, the number of allocated variables/registers per array reference, and the loops’ ordering are generated by an analytical model, leveraging the target hardware architecture details and loop kernel characteristics. The proposed methodology has been evaluated on both embedded and general-purpose CPUs, using seven well-known loop kernels and three machine learning applications. The results show significant speedups over the GCC compiler, the Pluto tool, and related work. Theologos Anthimopoulos, Georgios Keramidas, Vasilios I. Kelefouras, Iakovos Stamoulis |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2024 | Register Blocking: An Analytical Modelling Approach for Affine Loop KernelsabstractFor the past several decades, optimizing compilers have been a primary area of focus in both industry and academia. This continued research interest is a testament to the complexity of this task, primarily stemming from the vast number of parameters that must be explored to attain near-optimal results. One of the key compiler optimizations is "Register Blocking (RB)" also known as "Register-level Tiling" or "unroll-and-jam". RB can strongly reduce the number of executed Load/Store (L/S) instructions, and as a consequence the number of data accesses in memory hierarchy, but due to its inherent complexities, fine-tuning is essential for its effective implementation. To address this problem, in this work a new methodology is proposed for RB. The RB factors, the loops to apply RB, the number of allocated variables/registers per array reference, and the loops' ordering are generated by an analytical model, leveraging the target hardware (HW) architecture details and loop kernel characteristics. The proposed methodology has been evaluated on both embedded and general-purpose CPUs across seven well-known loop kernels, achieving high speedups and L/S instruction gains over GCC compiler, handwritten optimized codes, and the popular Pluto tool. Theologos Anthimopoulos, Georgios Keramidas, Vasilios I. Kelefouras, Iakovos Stamoulis |
CF | 2 |
| 2024 | Denseflex: A Low Rank Factorization Methodology for Adaptable Dense Layers in DNNsabstractLow-Rank Factorization (LRF) is a popular compression technique used in Deep Neural Networks (DNNs). LRF can reduce both the memory size and the arithmetic operations in a DNN layer by approximating a weight tensor/matrix by two or more smaller tensors/matrices. Employing LRF to DNN is a challenging task for several reasons. First, the exploration space is massive and different solutions provide different trade-offs among memory, FLOPs, inference time, and validation accuracy; second, multiple DNN layers and multiple LRF algorithms must be considered; third, every extracted solution must undergo through a calibration phase and this makes the LRF process time-consuming. In this paper, a methodology, called Denseflex, is presented that formulates the LRF problem as an inference time vs. FLOPs vs. memory vs. validation accuracy Design Space Exploration (DSE) problem. Moreover, to the best of our knowledge, this is the first work that proposes a methodology to efficiently combine two different LRF methods (Singular Value Decomposition -SVD- and Tensor Train Decomposition -TTD-) in the same framework. Denseflex is formulated as a design tool in which the user can provide specific memory, FLOPs, and/or execution time constraints and the tool will output a set of solutions that meet the given constraints avoiding the time-consuming re-training phases. Our results indicate that our approach is able to prune the design space by 62% (on average) over related works for nine DNN models (up to 88% in AlexNet), while the extracted LRF solutions exhibit both lower memory footprints and lower execution times compared to the initial model. Milad Kokhazadeh, Georgios Keramidas, Vasilios I. Kelefouras, Iakovos Stamoulis |
CF | 2 |
| 2024 | XANDAR: An X-by-Construction Framework for Safety, Security, and Real-Time Behavior of Embedded Software SystemsabstractThe safe and secure implementation of increasingly complex features is a major challenge in the development of autonomous and distributed embedded systems. Automated design-time procedures that guarantee the fulfillment of critical system properties are a promising approach to tackle this challenge. In the European project XANDAR, which took place from 2021 to 2023, eight partners developed an X-by-Construction (XbC) design framework to support developers in the creation of embedded software systems with certain safety, security, and real-time properties. The design framework combines a model-based toolchain with a hypervisor-based runtime architecture. It targets modern high-performance hardware, facilitates the integration of machine learning applications, and employs a library of trusted safety and security patterns to reduce the implementation and verification effort. This paper describes the concepts developed during the project, the prototypical implementation of the design framework, and its application in both an automotive and an avionics use case. Tobias Dörr, Florian Schade, Jürgen Becker 0001, Georgios Keramidas, Nikos Petrellis, Vasilios I. Kelefouras, Michail Mavropoulos, Konstantinos Antonopoulos, Christos P. Antonopoulos, Nikos S. Voros, Alexander Ahlbrecht, Wanja Zaeske, Vincent Janson, Phillip Nöldeke, Umut Durak, Christos Panagiotou, Dimitris Karadimas, Nico Adler, Clemens Reichmann, Andreas Sailer, Raphael Weber, Thomas Wilhelm 0005, Wolfgang Gabler, Katrin Weiden, Xavier Anzuela Recasens, Sakir Sezer, Fahad Siddiqui 0001, Rafiullah Khan, Kieran McLaughlin, Sena Yengec Tasdemir, Balmukund Sonigara, Henry Hui, Esther Soriano Viguer, Aridane Álvarez Suárez, Vicente Nicolau Gallego, Manuel Muñoz Alcobendas, Miguel Masmano Tello |
DATE | 4 |
| 2024 | Introduction to the FPL 2021 Special SectionabstractThe International Conference on Field-Programmable Logic and Applications (FPL) was the first and remains the largest conference covering the rapidly growing area of field-programmable logic and reconfigurable computing.During the past 30 years, many of the advances in reconfigurable system architectures, applications, embedded processors, and design automation methods and tools were first published in the proceedings of the FPL conference series.The conference objective is to bring together researchers and practitioners from both academia and industry and from around the world.The 31st edition of the FPL (2021) took place from August 30 till September 3, 2021.It is the second FPL conference that had to be organized as a virtual event due to the COVID-19 pandemic.The purpose of this Special Section is to provide an insight into current research and development in aspects related to Field-Programmable Gate Array (FPGA) applications, FPGA technology, and FPGA programming models and tools.This Special Section includes three papers that were presented in the 2021 edition of the conference.The three articles were appropriately selected (based on their quality) to cover various topics of the conference. Diana Göhringer, Georgios Keramidas, Akash Kumar 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2023 | Enhanced Sound Recognition and Classification Through Spectrogram Analysis, MEMS Sensors, and PyTorch: A Comprehensive Approach
Alexandros Spournias, Nikolaos Nanos, Evanthia Faliagka, Christos P. Antonopoulos, Nikos S. Voros, Georgios Keramidas |
CollaborateCom (1) | 6 |
| 2023 | Design and Implementation of Deep Learning 2D Convolutions on Modern CPUsabstractIn this article, a new method is provided for accelerating the execution of convolution layers in Deep Neural Networks. This research work provides the theoretical background to efficiently design and implement the convolution layers on x86/x64 CPUs, based on the target layer parameters, quantization level and hardware architecture. The proposed work is general and can be applied to other processor families too, e.g., Arm. The proposed work achieves high speedup values over the state of the art, which is Intel oneDNN library, by applying compiler optimizations, such as vectorization, register blocking and loop tiling, in a more efficient way. This is achieved by developing an analytical modelling approach for finding the optimization parameters. A thorough experimental evaluation has been applied on two Intel CPU platforms, for DenseNet-121, ResNet-50 and SqueezeNet (including 112 different convolution layers), and for both FP32 and int8 input/output tensors (quantization). The experimental results show that the convolution layers of the aforementioned models are executed from$x1.1$up to$x7.2$times faster. Vasilios I. Kelefouras, Georgios Keramidas |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | XANDAR: Exploiting the X-by-Construction Paradigm in Model-based Development of Safety-critical SystemsabstractRealizing desired properties “by construction” is a highly appealing goal in the design of safety-critical embedded systems. As verification and validation tasks in this domain are often both challenging and time-consuming, the by-construction paradigm is a promising solution to increase design productivity and reduce design errors. In the XANDAR project, partners from industry and academia develop a toolchain that will advance current development processes by employing a modelbased X-by-Construction (XbC) approach. XANDAR defines a development process, metamodel extensions, a library of safety and security patterns, and investigates many further techniques for design automation, verification, and validation. The developed toolchain will use a hypervisor-based platform, targeting future centralized, AI-capable high-performance embedded processing systems. It is co-developed and validated in both an avionics use case for situation perception and pilot assistance as well as an automotive use case for autonomous driving. Leonard Masing, Tobias Dörr, Florian Schade, Jürgen Becker 0001, Georgios Keramidas, Christos P. Antonopoulos, Michail Mavropoulos, Efstratios Tiganourias, Vasilios I. Kelefouras, Konstantinos Antonopoulos, Nikos S. Voros, Umut Durak, Alexander Ahlbrecht, Wanja Zaeske, Christos Panagiotou, Dimitris Karadimas, Nico Adler, Andreas Sailer, Raphael Weber, Thomas Wilhelm 0005, Géza Németh, Fahad Siddiqui 0001, Rafiullah Khan, Vahid Garousi, Sakir Sezer, Victor Morales |
DATE | 5 |
| 2022 | XANDAR: A holistic Cybersecurity Engineering Process for Safety-critical and Cyber-physical SystemsabstractThe integration of connected and autonomous technologies in safety-critical and cyber-physical systems offers great potential in the vital application domains of transportation, manufacturing and aerospace. These technological advancements are necessary to meet the increasing demand for intelligent services, as they open doors to new business models by analysing and sharing the generated data. However, where this sharing of mix-critical data and broader connectivity brings opportunities, it simultaneously presents serious cybersecurity and safety risks due to the cyber-physical nature of these systems. Hence, delivering these intelligent services securely, safely, and reliably to its consumers is a complex engineering and design problem. One of the ways to approach this engineering problem is to consider both system functional and non-functional properties (safety, security, reliability) and systematically integrate them across system design and operational life cycle. The XANDAR project investigates this approach and aims to develop holistic software design methods and architectures for safety-critical and cyber-physical systems that guarantee functional and non-functional properties “byconstruction”. This paper focuses on the non-functional aspects of the project and discusses the preliminary work. by presenting the core cybersecurity principles and uses them as a baseline to propose a holistic cybersecurity engineering process. The tasks of the proposed cybersecurity engineering process are also map onto relevant clauses of ISO 21434. In future, proposed work will be integrated into the XANDAR software toolchain and validated for an avionics situation perception pilot assistance and automotive autonomous driving use cases. Fahad Siddiqui 0001, Rafiullah Khan, Sakir Sezer, Kieran McLaughlin, Leonard Masing, Tobias Dörr, Florian Schade, Jürgen Becker 0001, Alexander Ahlbrecht, Wanja Zaeske, Umut Durak, Nico Adler, Andreas Sailer, Raphael Weber, Thomas Wilhelm 0005, Géza Németh, Victor Morales, Paco Gomez, Georgios Keramidas, Christos P. Antonopoulos, Michail Mavropoulos, Vasilios I. Kelefouras, Konstantinos Antonopoulos, Nikos S. Voros, Christos Panagiotou, Dimitris Karadimas |
VTC Spring | 19 |
| 2022 | Design and Implementation of 2D Convolution on x86/x64 ProcessorsabstractIn this paper, a new method for accelerating the 2D direct Convolution operation on x86/x64 processors is presented. It includes efficient vectorization by using SIMD intrinsics, bit-twiddling optimizations, the optimization of the division operation, multi-threading using OpenMP, register blocking and the shortest possible bit-width value of the intermediate results. The proposed method, which is provided as open-source, is general and can be applied to other processor families too, e.g., Arm. The proposed method has been evaluated on two different multi-core Intel CPUs, by using twenty different image sizes, 8-bit integer computations and the most commonly used kernel sizes (3x3, 5x5, 7x7, 9x9). It achieves from$2.8\times$to$40\times$speedup over the Intel IPP library (OpenCV GaussianBlur and Filter2D routines), from$105 \times$to$400 \times$speedup over the gemm-based convolution method (by using Intel MKL int8 matrix multiplication routine), and from$8.5\times$to$618\times$speedup over the vslsConvExec Intel MKL direct convolution routine. The proposed method is superior as it achieves far fewer arithmetical and load/store instructions. Vasilios I. Kelefouras, Georgios Keramidas |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2021 | High Speed Implementation of the Deformable Shape Tracking Face Alignment AlgorithmabstractThe 2D facial landmark alignment method, implemented in C++ in the open source libraries DLIB and Deformable Shape Tracking (DEST), is used in several applications such as driver drowsiness detection. The most challenging of these applications require fast video frame processing. Therefore, the alignment of the facial landmarks in a single video frame has to be performed with the minimum possible latency without precision loss. In this paper, the DEST implementation of the face alignment method that is based on regression trees is heavily restructured to reduce latency. The resulting face alignment predictor is implemented in C. The elimination of multiple nested routine calls, excessive argument copying, type conversions and integrity checks lead to a software implementation that is 240 times faster than the one provided in the DEST library. Moreover, the structure of the new face alignment predictor is appropriate for hardware implementation on a Field Programmable Gate Array (FPGA) for further acceleration1. Nikos Petrellis, Stavros Zogas, Panagiotis Christakos, Georgios Keramidas, Panagiotis Mousouliotis, Nikos S. Voros, Christos P. Antonopoulos |
DSD | 4 |
| 2021 | Run Time Management of Faulty Data CachesabstractAs the technology continuous to shrink, power consumption appears to be the main design parameter. Operation on low voltage negatively affects mainly the operation of on-chip memories, resulting in multiple malfunctioning memory cells. As a reaction many cache fault tolerance (CFT) mechanisms have been proposed targeting the mitigation of performance degradation. The challenge is to devise mechanisms that are tailored to the memory access patterns of the executing applications. In this work we initially investigate the impact of the granularity of cache line disabling scheme in the first level data caches. Based on our analysis, we propose a run time adaptive mechanism that is able to opt the cache (sub-)block taking into account the diverse memory characteristics of the application. The proposed mechanism is based on the widely used block (sub-block) disabling scheme, and dynamically selects the appropriate sub-block granularity during the execution of the applications. Our evaluation results reveal that the proposed dynamic approach is able to offer significant benefits over a faulty cache design with a monolithic (sub-)block granularity. Michail Mavropoulos, Georgios Keramidas, Dimitris Nikolos |
ETS | 2 |
| 2021 | XANDAR: X-by-Construction Design framework for Engineering Autonomous & Distributed Real-time Embedded Software SystemsabstractThe next generation of networked embedded systems (ES) necessitates rapid prototyping and high performance while maintaining key qualities like trustworthiness and safety. However, development of safety-critical ES suffers from complex software (SW) toolchains and engineering processes. Moreover, the current trend in autonomous systems, which relies on Machine Learning (ML) and AI applications when combined with fail-operational requirements renders the Verification and Validation (V&V) of these new systems a challenging endeavor. Prime examples are Advanced Driver-Assistance Systems (ADAS) that are prone to various safety/security vulnerabilities. The XANDAR project aims at developing a mature SW toolchain (from requirements analysis to the actual code integration on target including V&V) fulfilling the needs of industry for rapid prototyping of interoperable and autonomous ES. Starting from a model-based system architecture, XANDAR will leverage automatic model synthesis and software parallelization techniques to achieve specific non-functional requirements setting the foundation for a novel (real-time, safety-, and security)-by-Construction paradigm. Jürgen Becker 0001, Leonard Masing, Tobias Dörr, Florian Schade, Georgios Keramidas, Christos P. Antonopoulos, Michail Mavropoulos, Efstratios Tiganourias, Vasilios I. Kelefouras, Konstantinos Antonopoulos, Nikos S. Voros, Umut Durak, Alexander Ahlbrecht, Wanja Zaeske, Christos Panagiotou, Dimitris Karadimas, Nico Adler, Andreas Sailer, Raphael Weber, Thomas Wilhelm 0005, Florian Oszwald, Dominik Reinhardt, Mohamad Chamas, Adnan Bekan, Graham Smethurst, Fahad Siddiqui 0001, Rafiullah Khan, Vahid Garousi, Sakir Sezer, Victor Morales |
FPL | 5 |
| 2021 | Architectures for SLAM and Augmented Reality ComputingabstractIn the next few years, new demanding applications will be supported on mobile platforms by reconciling two conflicting requirements: high performance (often with real-time limitations) and low power consumption. The objective of the vipGPU project is to develop hardware and software technology to provide efficient support for two such application scenarios, namely (a) simultaneous localization and mapping (SLAM) in mobile robotics systems, and (b) virtual reality (VR) in portable devices to simulate serious games with emphasis on simulating surgical interventions and medical training in general. In this project, we aim at developing a new heterogeneous platform consisting of hardware accelerators for low power embedded systems optimized (at the hardware and software level) for the implementation of the two applications mentioned above. Nikolaos Bellas, Christos D. Antonopoulos, Spyros Lalis, Maria Rafaela Gkeka, Alexandros Patras, Georgios Keramidas, Iakovos Stamoulis, Nikolaos Tavoularis, Stylianos Piperakis, Emmanouil Hourdakis, Panos E. Trahanias, Paul Zikas, George Papagiannakis, Ioanna Kartsonaki |
FPL | 6 |
| 2021 | The SMART4ALL High Performance Computing Infrastructure: Sharing high-end hardware resources via cloud-based microservicesabstractThe goal of the SMART4ALL EU funded project is to build capacity amongst European stakeholders via the development of self-sustained, cross-border experiments that transfer knowledge and technology between academia and industry. It targets Customized Low Energy Computing (CLEC) in the Cyber-Physical (CPS) and the Internet of Things (IoT) domains. The vision of the project will be mainly realized through funded (via open calls) Pathfinder Application Experiments (PAEs) that will enable the transformation of academic knowledge into products.It is important to mention that SMART4ALL sets forward the concept of marketplace which is offered as a service (Marketplace-as-a-Servtce or MaaS) that acts as one-stopsmart-stop by offering tools, HPC services, and platforms to the PAE grantees. The purpose of this paper is to present the SMART4ALL High-Performance Computing (HPC) infrastructure and services realized according to the Hardware-as-a-Service paradigm. The SMART4A11 HPC infrastructure includes CPU, Memory, GPU, storage as well as FPGA resources connected with high-speed links. Angelos S. Voros, Christos Panagiotou, Stavros Zogas, Georgios Keramidas, Christos P. Antonopoulos, Michael Hübner 0001, Nikos S. Voros |
FPL | 4 |
| 2021 | Challenges Towards Hardware Acceleration of the Deformable Shape Tracking ApplicationabstractIn the context of this paper, a shape tracking application based on landmark alignment is transformed to support implementation in Field Programmable Gate Arrays (FPGAs). Towards this direction, several challenges are posed since a) computational intensive operations have to be replaced by faster ones, b) specific loops have to be modified (e.g., unrolled) to support the implementation of operations in parallel with different hardware resources, c) multiple pretrained models have to be compared in terms of speed and accuracy, d) partial loading of the pre-trained models has to be examined in order to fit their parameters in the Block Random Access Memories (BRAMs) of the FPGA for faster access, and e) alternative arithmetic representations have to be evaluated for higher speed and reduced resources.The C++ Deformable Shape Tracking (DEST) implementation of face alignment that is based on an Ensemble of Regression Trees is employed in our approach. The DEST application uses Eigen library routines to implement algebraic operations which are proved to be quite slow. The achievements of this paper, concern the replacement of appropriate Eigen calls in time critical paths with fast C code that can be directly used to synthesize reconfigurable hardware implementations. The elimination of the computational intensive Eigen calls has already improved the speed of the face alignment application by more than 240 times. In this paper we examine how the modified source code structure of the DEST application can be used to address the challenges described above. Nikos Petrellis, Panagiotis Christakos, Stavros Zogas, Panagiotis Mousouliotis, Georgios Keramidas, Nikos S. Voros, Christos P. Antonopoulos |
VLSI-SoC | 5 |
| 2018 | A novel fault tolerant cache architecture based on orthogonal latin squares theoryabstractAggressive dynamic voltage and frequency scaling is widely used to reduce the power consumption of microprocessors. Unfortunately, voltage scaling increases the impact of process variations on memory cells resulting in an exponential increase in the number of malfunctioning memory cells. As a result, various cache fault-tolerant (CFT) techniques have been proposed. In this work, we propose a new CFT technique which applies a systematic redistribution (permutation) of the cache blocks (assuming various block granularity levels) within the cache structure using the orthogonal Latin Square concept and taking as input the location of the malfunctioning cells in the cache array. The aim of the redistribution is twofold. First, to uniformly distribute the faulty blocks to sets and second, to gather the faulty subblocks to a minimum number of blocks, so as the fault free blocks are maximized. Our evaluation results using the benchmarks of SPEC2006 suite, 100 memory fault maps, and four percentages of malfunctioning cells show that our proposal exhibits strong capability to reduce cache performance degradation especially in situations with high percentages of faulty cells and compares favorably to already known techniques. Filippos Filippou, Georgios Keramidas, Michail Mavropoulos, Dimitris Nikolos |
DATE | 2 |
| 2018 | Combining Software Cache Partitioning and Loop Tiling for Effective Shared Cache ManagementabstractOne of the biggest challenges in multicore platforms is shared cache management, especially for data-dominant applications. Two commonly used approaches for increasing shared cache utilization are cache partitioning and loop tiling. However, state-of-the-art compilers lack efficient cache partitioning and loop tiling methods for two reasons. First, cache partitioning and loop tiling are strongly coupled together, and thus addressing them separately is simply not effective. Second, cache partitioning and loop tiling must be tailored to the target shared cache architecture details and the memory characteristics of the corunning workloads. To the best of our knowledge, this is the first time that a methodology provides (1) a theoretical foundation in the above-mentioned cache management mechanisms and (2) a unified framework to orchestrate these two mechanisms in tandem (not separately). Our approach manages to lower the number of main memory accesses by an order of magnitude keeping at the same time the number of arithmetic/addressing instructions to a minimal level. We motivate this work by showcasing that cache partitioning, loop tiling, data array layouts, shared cache architecture details (i.e., cache size and associativity), and the memory reuse patterns of the executing tasks must be addressed together as one problem, when a (near)-optimal solution is requested. To this end, we present a search space exploration analysis where our proposal is able to offer a vast deduction in the required search space. Vasilios I. Kelefouras, Georgios Keramidas, Nikos S. Voros |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2017 | The LPGPU2 Project: Low-Power Parallel Computing on GPUs: Extended AbstractabstractThe LPGPU2 project is a 30-month-project (Innovation Action) funded by the European Union. Its overall goal is to develop an analysis and visualization framework that enables GPU application developers to improve the performance and power consumption of their applications. To achieve this overall goal, several key objectives need to be achieved. First, several applications (use cases) need to be developed for or ported to low-power GPUs. Thereafter, these applications need to be optimized using the tooling framework. In addition, power measurement devices and power models need to be developed that are 10x more accurate than the state of the art. The project consortium actively promotes open vendor-neutral standards via the Khronos group. This paper briefly reports on the achievements made in the first half of the project, and focuses on the progress made in applications; in power measurement, estimation, and modelling; and in the analysis and visualization tool suite. Ben H. H. Juurlink, Jan Lucas, Nadjib Mammeri, Martyn Bliss, Georgios Keramidas, Chrysa Kokkala, Andrew Richards |
SCOPES | 5 |
| 2016 | Computation and communication challenges to deploy robots in assisted living environments
Georgios Keramidas, Christos P. Antonopoulos, Nikos S. Voros, Fynn Schwiegelshohn, Philipp Wehner, Jens Rettkowski, Diana Göhringer, Michael Hübner 0001, Stasinos Konstantopoulos, Theodoros Giannakopoulos, Vangelis Karkaletsis, Evaggelinos P. Mariatos |
DATE | 1 |
| 2016 | Recovery of performance degradation in defective branch target buffersabstractDynamic voltage and frequency scaling (DVFS) is a commonly-used power-management technique. Unfortunately, voltage scaling increases the impact of process variations on memory cells reliability resulting in an exponential increase in the number of malfunctioning memory cells. In this work, we systematically investigate the behavior of branch target buffers (BTB) with faulty memory cells. Although being an intrinsically fault-tolerant unit (i.e., it does not affect correctness of the system), as we show in this work for several fault probabilities and core configurations, disabling the faulty parts of BTBs can damage the performance of the executing applications. To remedy the negative impact of malfunctioning BTB memory cells in contemporary BTB organizations, we present an ultra lightweight performance recovery mechanism. The proposed mechanism introduces minimal hardware overheads and practically-zero delays. Using cycle-accurate simulations, the benchmarks of SPEC2006 suite, a plethora of memory fault maps, and two fault probabilities corresponding to low supply voltages, we show the effectiveness of the proposed recovery mechanism. Filippos Filippou, Georgios Keramidas, Michail Mavropoulos, Dimitris Nikolos |
IOLTS | 2 |
| 2015 | A defect-aware reconfigurable cache architecture for low-vccmin DVFS-enabled systems
Michail Mavropoulos, Georgios Keramidas, Dimitris Nikolos |
DATE | 2 |
| 2015 | A Holistic Approach for Advancing Robots in Ambient Assisted Living EnvironmentsabstractDue to the demographic change in western society, new challenges regarding healthcare of the elderly population are at the verge of surfacing. Since young people are not capable of sustaining an adequate healthcare for elderly people, new healthcare fields have to be devised. Recent advances in information and communication technology enable the support of elderly people in their domestic environment. The EU project RADIO will design of an old age compliant smart home environment which specializes in fulfilling the needs of elderly people. This is partially achieved through a mobile robot platform which serves as an assistant to the respective elderly person. Apart from this, the robot also functions as a mobile sensor platform. Under this context, unobtrusiveness is of paramount importance since the robot should be a natural participant of patients' daily life. This paper discusses such a healthcare facility, analyses its requirements and poses the challenges towards this direction. Fynn Schwiegelshohn, Philipp Wehner, Jens Rettkowski, Diana Göhringer, Michael Hübner 0001, Georgios Keramidas, Christos P. Antonopoulos, Nikos S. Voros |
EUC | 6 |
| 2015 | Reconfigurable: Self Adaptive Fault Tolerant Cache Memory for DVS enabled SystemsabstractProcessor caches play a critical role in the performance of today"s computer systems. As technology scales, due to manufacturing defects and process variations a large number of cells in a cache is expected to be faulty. The number of faulty cells varies from die to die and in the field of the application depends on the operating conditions (e.g., supply voltage, frequency). Several techniques have been proposed to tolerate faults in caches. A drawback of the redundancy based techniques is that the amount of redundancy is decided at the design time targeting a maximum number of faults, so in cases of a small number of faults (e.g., in the nominal supply voltage in a system with DVS) only a part of the redundant resources is used. In this paper we propose a new reconfigurable-self adaptive fault tolerant cache scheme. The unique characteristic of our scheme is that it uses its resources for both the reduction of the misses caused by the faulty blocks as well as for the reduction of conflict misses, depending on the number of faults, their distribution in the cache, and the running application. Our experimental results for a wide range of scientific applications and a plethora of fault maps with different SRAM failure probabilities reveal that our proposal can achieve significant benefits. Michail Mavropoulos, Georgios Keramidas, Grigorios Adamopoulos, Dimitris Nikolos |
ACM Great Lakes Symposium on VLSI | 2 |
| 2014 | Spatial pattern prediction based management of faulty data cachesabstractTechnology scaling leads to significant faulty bit rates in on-chip caches. In this work, we propose a methodology to mitigate the impact of defective bits (due to permanent faults) in first-level set-associative data caches. Our technique assumes that faulty caches are enhanced with the ability of disabling their defective parts at cache subblock granularity. Our experimental findings reveal that while the occurrence of hard-errors in faulty caches may have a significant impact in performance, a lot of room for improvement exists, if someone is able to take into account the spatial reuse patterns of the to-be-referenced blocks (not all the data fetched into the cache is accessed). To this end, we propose frugal PC-indexed spatial predictors (with very small storage requirements) to orchestrate the (re)placement decisions among the fully and partially unusable faulty blocks. Using cycle-accurate simulations, a wide range of scientific applications, and a plethora of cache fault maps, we showcase that our approach is able to offer significant benefits in cache performance. Georgios Keramidas, Michail Mavropoulos, Anna Karvouniari, Dimitris Nikolos |
DATE | 1 |
| 2012 | Embedded reconfigurable architecturesabstractIn current-day embedded systems design, one is faced with cut-throat competition to deliver new functionalities in increasingly shorter time frames. This is now achieved by incorporating processor cores into embedded systems through (re-)programmability. However, this is not always beneficial for the performance or energy consumption. Therefore, adaptable embedded systems have been proposed to deal with these negative effects by reconfiguring the critical sections of an embedded system. In these proposals, we are clearly witnessing a trend that is moving from static configurations to dynamic (re)configurations. Stephan Wong, Luigi Carro, Stamatios Kavvadias, Georgios Keramidas, Francesco Papariello, Claudio Scordino, Roberto Giorgi, Stefanos Kaxiras |
CASES | 4 |
| 2011 | Multicore Cache Simulations Using Heterogeneous Computing on General Purpose and Graphics ProcessorsabstractTraditional trace-driven memory system simulation is a very time consuming process while the advent of multicores simply exacerbates the problem. We propose a framework for accelerating trace-driven multicore cache simulations by utilizing the capabilities of the modern many core GPUs. A straightforward way towards this direction is to rely on the inherent parallelism in cache simulations: communicating cache sets can be simulated independently and concurrently to other sets. Based on this, we map collections of communicating cache sets (each belonging to a different target cache) on the same GPU block so that the simulated coherence traffic is local traffic in the GPU. However, this is not enough due to the great imbalance in the activity in the different cache sets: some sets receive a flurry of activity while others do not. Our solution is to load balance the simulated sets (based on activity) on the computing element (host-CPU or GPU) that can manage them in the most efficient way. We propose a heterogeneous computing approach in which the host-CPU simulates the few but most active sets, while the GPU is responsible for the many more but less active sets. Our experimental findings using the SPLASH-2 suite demonstrate that our cache simulator based on the CPU-GPU cooperation achieves on average 5.88x speedup over alternative implementations running on CPU, speedups which scale well with the size of the simulated system. Georgios Keramidas, Nikolaos Strikos, Stefanos Kaxiras |
DSD | 1 |
| 2011 | Poster: DVFS management in real-processorsabstractWe describe a framework for run-time adaptive dynamic voltage-frequency scaling in Linux systems. Our underlying methodology is based on a simple first-order processor performance model in which frequency scaling is expressed as a change (in cycles) of the main memory latency. Utilizing available performance monitoring hardware, we show that our model is powerful enough to i) predict with reasonable accuracy the effect of frequency scaling, and ii) predict the energy consumed by the core under different V/f combinations. To validate our approach we perform highly accurate, fine grained power measurements directly on the processor off-chip voltage regulator. Vasileios Spiliopoulos 0001, Georgios Keramidas, Stefanos Kaxiras, Konstantinos Efstathiou 0002 |
ICS | 2 |
| 2007 | Using value locality to reduce memory encryption overhead in embedded processorsabstractMemory encryption has gained much attention lately as a way to offer a secure environment to fight against software and hardware attacks. Many researchers provided memory encryption schemes whereby one or more levels of the memory hierarchy were encrypted using a cryptographic algorithm such as AES. Counter mode (CM) encryption, also called one-time-pad (OTP) encryption, is proven to be quite effective for main memory encryption. However, CM encryption requires an extra sequence number (counter) to be associated with every memory location (L2 block cacheline granularity is used). The per-block counters must be updated every time a block is written back to memory otherwise known-plaintext attacks may occur. Thus, the size of those counters is a critical parameter in the system design. In this work, we propose the use of silent stores as a method of providing the CM encryption with less overhead. Silent stores, i.e. stores, to memory that write the same value as already stored in that memory location, have been observed to occur frequently. These stores create redundant memory write-backs (and counter updates), so eliminating them will lower performance overheads introduced by the encyption/decryption process. Our initial results show significant benefits across the board indicating the promising nature of the proposed idea. Georgios Keramidas, Pavlos Petoumenos, Alexandros Antonopoulos, Stefanos Kaxiras, Dimitrios Serpanos |
ETFA | 1 |
| 2007 | Applying Decay to Reduce Dynamic Power in Set-Associative Caches
Georgios Keramidas, Polychronis Xekalakis, Stefanos Kaxiras |
HiPEAC | 1 |
| 2007 | Cache replacement based on reuse-distance predictionabstractSeveral cache management techniques have been proposed that indirectly try to base their decisions on cacheline reuse-distance, like Cache Decay which is a postdiction of reuse-distances: if a cacheline has not been accessed for some ldquodecay intervalrdquo we know that its reuse-distance is at least as large as this decay interval. In this work, we propose to directly predict reuse-distances via instruction-based (PC) prediction and use this information for cache level optimizations. In this paper, we choose as our target for optimization the replacement policy of the L2 cache, because the gap between the LRU and the theoretical optimal replacement algorithm is comparatively large for L2 caches. This indicates that, in many situations, there is ample room for improvement. We evaluate our reusedistance based replacement policy using a subset of the most memory intensive SPEC2000 and our results show significant benefits across the board. Georgios Keramidas, Pavlos Petoumenos, Stefanos Kaxiras |
ICCD | 1 |
| 2005 | IPStash: a set-associative memory approach for efficient IP-lookupabstractIP-lookup is a challenging problem because of the increasing routing table sizes, increased traffic and higher speed links. These characteristics lead to the prevalence of hardware solutions such as TCAMs (ternary content addressable memories), despite their high power consumption, low update rate and increased board area requirements. We propose a memory architecture called IPStash to act as a TCAM replacement, offering at the same time, high update rate, higher performance and significant power savings. The premise of our work is that full associativity is not necessary for IP-lookup. Rather, we show that the required associativity is simply a function of the routing table size, Thus, we propose a memory architecture similar to set-associative caches but enhanced with mechanisms to facilitate IP-lookup and in particular longest prefix match (LPM). To reach a minimum level of required associativity we introduce an iterative method to perform LPM in a small number of iterations. This allows us to insert route prefixes of different lengths in IPStash very efficiently, selecting the most appropriate index in each case. Orthogonal to this, we use skewed associativity to increase the effective capacity of our devices. We thoroughly examine different choices in partitioning routing tables for the iterative LPM and the design space for the IPStash devices. The proposed architecture is also easily expandable. Using the Cacti 3.2 access time and power consumption simulation tool we explore the design space for IPStash devices and we compare them with the best blocked commercial TCAMs. Stefanos Kaxiras, Georgios Keramidas |
INFOCOM | 2 |
| 2005 | A simple mechanism to adapt leakage-control policies to temperatureabstractLeakage power reduction in cache memories continues to be a critical area of research because of the promise of a significant pay-off. Various techniques have been developed so far that can be broadly categorized into state-preserving (e.g., Drowsy Caches) and non-state preserving (e.g., Cache Decay). Decay saves more leakage but also incurs dynamic power overhead in the form of induced misses. Previous work has shown that depending on the leakage vs. dynamic power trade-off, one or the other technique can be better. Several factors such as cache architecture, technology parameters and temperature, affect this trade-off. Our work proposes the first mechanism ---to the best of our knowledg--- that takes into account temperature in adjusting the leakage control policy at run time. At very low temperatures, leakage is relatively weak so the need to tightly control it is not as important as the need to minimize extra dynamic power (e.g., decay-induced misses) or performance loss. We use a hybrid decay+drowsy policy where the main benefit comes from decaying cache lines while the drowsy mode is used to save leakage in long decay intervals. To adapt the decay mode to temperature, we propose a simple triggering mechanism that is based on the principles of decaying 4T thermal sensors and, as such, tied to temperature. The hotter the cache is, the faster cache lines are decayed since it is beneficial to do so with very high leakage currents.Conversely, when the cache temperature is low, our mechanism defers putting cache lines in decay mode to avoid dynamic power overhead but still saves a significant amount of leakage using the drowsy mode. Our study shows that across a wide range of temperatures, the simple adaptability of our proposal yields consistently better results than either the decay mode, or drowsy mode alone, improving over the best by as much as 33% Stefanos Kaxiras, Polychronis Xekalakis, Georgios Keramidas |
ISLPED | 3 |
| 2003 | IPStash: a Power-Efficient Memory Architecture for IP-lookupabstractHigh-speed routers often use commodity, fully-associative, TCAMs (ternary content addressable memories) to perform packet classification and routing (IP-lookup). We propose a memory architecture called IPStash to act as a TCAM replacement, offering at the same time, better functionality, higher performance, and significant power savings. The premise of our work is that full associativity is not necessary for IP-lookup. Rather, we show that the required associativity is simply a function of the routing table size. We propose a memory architecture similar to set-associative caches but enhanced with mechanisms to facilitate IP-lookup and in particular longest prefix match. To perform longest prefix match efficiently in a set-associative array, we restrict routing table prefixes to a small number of lengths using a controlled prefix expansion technique. Since this inflates the routing tables, we use skewed associativity to increase the effective capacity of our devices. Compared to previous proposals, IPStash does not require any complicated routing table transformations but more importantly, it makes incremental updates to the routing tables effortless. The proposed architecture is also easily expandable. Our simulations show that IPStash is both fast and power efficient compared to TCAMs. Specifically, IPStash devices - built in the same technology as TCAMS - can run at speeds in excess of 600 MHz, offer more than twice the search throughput (>200Msps), and consume up to 35% less power (for the same throughput) than the best commercially available TCAMs when testes with real routing tables and IP traffic. Stefanos Kaxiras, Georgios Keramidas |
MICRO | 2 |