EDBT 2026 Demo / reviewers in the wild / expert
Vasilios I. Kelefouras
dblp:52/10471
· DBLP profile ↗
29ranked-venue papers
12as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 26 · 11 first-author · 12 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Energy efficiency support for software defined networks: a serverless computing approachabstractAutomatic network management strategies have become paramount for meeting the needs of innovative real-time and data-intensive applications, such as those in the Internet of Things. However, the ever-growing and fluctuating demands for data and services in such applications require more than ever an efficient, scalable, and energy-aware network resource management. To address these challenges, this paper introduces a novel approach that leverages a modular architecture based on serverless functions within an energy-aware environment. By deploying SDN services as Functions as a Service (FaaS), the proposed approach enables dynamic, on-demand network function deployment, achieving significant cost and energy savings through fine-grained resource provisioning. Unlike previous monolithic SDN approaches, this work disaggregates SDN control plane into modular, serverless components, transforming tightly integrated functionalities into independent, on-demand services while ensuring performance, scalability, and energy efficiency. An analytical model is presented to approximate the service delivery time and power consumption, as well as an open source prototype implementation supported by an extensive experimental evaluation. Experimental results demonstrate significant improvement in energy efficiency compared to traditional approaches, highlighting the potential of this approach for sustainable network environments. Fatemeh Banaie, Karim Djemame, Abdulaziz Alhindi, Vasilios I. Kelefouras |
Future Gener. Comput. Syst. | 4 |
| 2025 | A CNN Compression Methodology for Layer-Wise Rank Selection Considering Inter-Layer InteractionsabstractConvolutional Neural Networks (CNNs) achieve state-of-the-art performance across various application domains but are often resource-intensive, limiting their use on resource-constrained devices. Low-rank factorization (LRF) has emerged as a promising technique to reduce the computational complexity and memory footprint of CNNs, enabling efficient deployment without significant performance loss. However, challenges still remain in optimizing the rank selection problem, balancing memory reduction and accuracy, and integrating LRF into the training process of CNNs. In this paper, a novel and generic methodology for layer-wise rank selection is presented, considering inter-layer interactions. Our approach is compatible with any decomposition method and does not require additional retraining. The proposed methodology is evaluated in thirteen widely-used, CNN models, significantly reducing model parameters and Floating-Point Operations (FLOPs). In particular, our approach achieves up to a 94.6% parameter reduction (82.3% on average) and up to 90.7% FLOPs reduction (59.6% on average), with less than a 1.5% drop in validation accuracy, demonstrating superior performance and scalability compared to existing techniques. Milad Kokhazadeh, Georgios Keramidas, Vasilios I. Kelefouras, Iakovos Stamoulis |
DATE | 3 |
| 2025 | Optimizing Tensor Train Decomposition in DNNs for RISC-V Architectures Using Design Space Exploration and Compiler OptimizationsabstractDeep neural networks (DNNs) have become indispensable in many real-life applications like natural language processing, and autonomous systems. However, deploying DNNs on resource-constrained devices, e.g., in RISC-V platforms, remains challenging due to the high computational and memory demands of fully connected (FC) layers, which dominate resource consumption. Low-rank factorization (LRF) offers an effective approach to compressing FC layers, but the vast design space of LRF solutions involves complex tradeoffs among FLOPs, memory size, inference time, and accuracy, making the LRF process complex and time-consuming. This article introduces an end-to-end LRF design space exploration methodology and a specialized design tool for optimizing FC layers on RISC-V processors. Using Tensor Train Decomposition (TTD) offered by TensorFlow T3F library, the proposed work prunes the LRF design space by excluding first, inefficient decomposition shapes and second, solutions with poor inference performance on RISC-V architectures. Compiler optimizations are then applied to enhance custom T3F layer performance, minimizing inference time and boosting computational efficiency. On average, our TT-decomposed layers run 3× faster than IREE and 8× faster than Pluto on the same compressed model. This work provides an efficient solution for deploying DNNs on edge and embedded devices powered by RISC-V architectures. Theologos Anthimopoulos, Milad Kokhazadeh, Vasilios I. Kelefouras, Benjamin Himpel, Georgios Keramidas |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2025 | Register Blocking: A Source-to-Source Analytical Modelling Approach for Affine Loop KernelsabstractRegister Blocking (RB), also known as ‘Register-level Tiling’ or ‘unroll-and-jam,’ is a key compiler optimization for developing efficient micro-kernels. However, applying RB effectively is a complex task due to several challenges. First, the exploration space of possible RB configurations is vast. Second, RB and loop permutation are interdependent; therefore, addressing both optimizations simultaneously further inflates the exploration space. Third, the effectiveness of RB is highly dependent on the target hardware platform and the specific loop kernel being optimized. As a result, an extensive and time-consuming fine-tuning process is necessary for achieving an efficient implementation. To address these challenges, a source-to-source analytical modelling approach is proposed. The RB factors, the loops to apply RB, the number of allocated variables/registers per array reference, and the loops’ ordering are generated by an analytical model, leveraging the target hardware architecture details and loop kernel characteristics. The proposed methodology has been evaluated on both embedded and general-purpose CPUs, using seven well-known loop kernels and three machine learning applications. The results show significant speedups over the GCC compiler, the Pluto tool, and related work. Theologos Anthimopoulos, Georgios Keramidas, Vasilios I. Kelefouras, Iakovos Stamoulis |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2024 | Register Blocking: An Analytical Modelling Approach for Affine Loop KernelsabstractFor the past several decades, optimizing compilers have been a primary area of focus in both industry and academia. This continued research interest is a testament to the complexity of this task, primarily stemming from the vast number of parameters that must be explored to attain near-optimal results. One of the key compiler optimizations is "Register Blocking (RB)" also known as "Register-level Tiling" or "unroll-and-jam". RB can strongly reduce the number of executed Load/Store (L/S) instructions, and as a consequence the number of data accesses in memory hierarchy, but due to its inherent complexities, fine-tuning is essential for its effective implementation. To address this problem, in this work a new methodology is proposed for RB. The RB factors, the loops to apply RB, the number of allocated variables/registers per array reference, and the loops' ordering are generated by an analytical model, leveraging the target hardware (HW) architecture details and loop kernel characteristics. The proposed methodology has been evaluated on both embedded and general-purpose CPUs across seven well-known loop kernels, achieving high speedups and L/S instruction gains over GCC compiler, handwritten optimized codes, and the popular Pluto tool. Theologos Anthimopoulos, Georgios Keramidas, Vasilios I. Kelefouras, Iakovos Stamoulis |
CF | 3 |
| 2024 | Denseflex: A Low Rank Factorization Methodology for Adaptable Dense Layers in DNNsabstractLow-Rank Factorization (LRF) is a popular compression technique used in Deep Neural Networks (DNNs). LRF can reduce both the memory size and the arithmetic operations in a DNN layer by approximating a weight tensor/matrix by two or more smaller tensors/matrices. Employing LRF to DNN is a challenging task for several reasons. First, the exploration space is massive and different solutions provide different trade-offs among memory, FLOPs, inference time, and validation accuracy; second, multiple DNN layers and multiple LRF algorithms must be considered; third, every extracted solution must undergo through a calibration phase and this makes the LRF process time-consuming. In this paper, a methodology, called Denseflex, is presented that formulates the LRF problem as an inference time vs. FLOPs vs. memory vs. validation accuracy Design Space Exploration (DSE) problem. Moreover, to the best of our knowledge, this is the first work that proposes a methodology to efficiently combine two different LRF methods (Singular Value Decomposition -SVD- and Tensor Train Decomposition -TTD-) in the same framework. Denseflex is formulated as a design tool in which the user can provide specific memory, FLOPs, and/or execution time constraints and the tool will output a set of solutions that meet the given constraints avoiding the time-consuming re-training phases. Our results indicate that our approach is able to prune the design space by 62% (on average) over related works for nine DNN models (up to 88% in AlexNet), while the extracted LRF solutions exhibit both lower memory footprints and lower execution times compared to the initial model. Milad Kokhazadeh, Georgios Keramidas, Vasilios I. Kelefouras, Iakovos Stamoulis |
CF | 3 |
| 2024 | XANDAR: An X-by-Construction Framework for Safety, Security, and Real-Time Behavior of Embedded Software SystemsabstractThe safe and secure implementation of increasingly complex features is a major challenge in the development of autonomous and distributed embedded systems. Automated design-time procedures that guarantee the fulfillment of critical system properties are a promising approach to tackle this challenge. In the European project XANDAR, which took place from 2021 to 2023, eight partners developed an X-by-Construction (XbC) design framework to support developers in the creation of embedded software systems with certain safety, security, and real-time properties. The design framework combines a model-based toolchain with a hypervisor-based runtime architecture. It targets modern high-performance hardware, facilitates the integration of machine learning applications, and employs a library of trusted safety and security patterns to reduce the implementation and verification effort. This paper describes the concepts developed during the project, the prototypical implementation of the design framework, and its application in both an automotive and an avionics use case. Tobias Dörr, Florian Schade, Jürgen Becker 0001, Georgios Keramidas, Nikos Petrellis, Vasilios I. Kelefouras, Michail Mavropoulos, Konstantinos Antonopoulos, Christos P. Antonopoulos, Nikos S. Voros, Alexander Ahlbrecht, Wanja Zaeske, Vincent Janson, Phillip Nöldeke, Umut Durak, Christos Panagiotou, Dimitris Karadimas, Nico Adler, Clemens Reichmann, Andreas Sailer, Raphael Weber, Thomas Wilhelm 0005, Wolfgang Gabler, Katrin Weiden, Xavier Anzuela Recasens, Sakir Sezer, Fahad Siddiqui 0001, Rafiullah Khan, Kieran McLaughlin, Sena Yengec Tasdemir, Balmukund Sonigara, Henry Hui, Esther Soriano Viguer, Aridane Álvarez Suárez, Vicente Nicolau Gallego, Manuel Muñoz Alcobendas, Miguel Masmano Tello |
DATE | 6 |
| 2023 | Design and Implementation of Deep Learning 2D Convolutions on Modern CPUsabstractIn this article, a new method is provided for accelerating the execution of convolution layers in Deep Neural Networks. This research work provides the theoretical background to efficiently design and implement the convolution layers on x86/x64 CPUs, based on the target layer parameters, quantization level and hardware architecture. The proposed work is general and can be applied to other processor families too, e.g., Arm. The proposed work achieves high speedup values over the state of the art, which is Intel oneDNN library, by applying compiler optimizations, such as vectorization, register blocking and loop tiling, in a more efficient way. This is achieved by developing an analytical modelling approach for finding the optimization parameters. A thorough experimental evaluation has been applied on two Intel CPU platforms, for DenseNet-121, ResNet-50 and SqueezeNet (including 112 different convolution layers), and for both FP32 and int8 input/output tensors (quantization). The experimental results show that the convolution layers of the aforementioned models are executed from$x1.1$up to$x7.2$times faster. Vasilios I. Kelefouras, Georgios Keramidas |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2022 | Evaluation of Language Runtimes in Open-source Serverless PlatformsabstractServerless computing is revolutionising cloud application development as it offers the ability to create modular,\nhighly-scalable, fault-tolerant applications, with minimal operational management. In order to contribute\nto its widespread adoption of serverless platforms, the design and performance of language runtimes that\nare available in Function-as-a-Service (FaaS) serverless platforms is key. This paper aims to investigate the\nperformance impact of language runtimes in open-source serverless platforms, deployable on local clusters.\nA suite of experiments is developed and deployed on two selected platforms: OpenWhisk and Fission. The\nresults show a clear distinction between compiled and dynamic languages in cold starts but a pretty close\noverall performance in warm starts. Comparisons with similar evaluations for commercial platforms reveal\nthat warm start performance is competitive for certain languages, while cold starts are lagging behind by a wide\nmargin. Overall, the evaluation yielded usable results in regards to preferable choice of language runtime for\neach platform Karim Djemame, Daniel Datsev, Vasilios I. Kelefouras |
CLOSER | 3 |
| 2022 | XANDAR: Exploiting the X-by-Construction Paradigm in Model-based Development of Safety-critical SystemsabstractRealizing desired properties “by construction” is a highly appealing goal in the design of safety-critical embedded systems. As verification and validation tasks in this domain are often both challenging and time-consuming, the by-construction paradigm is a promising solution to increase design productivity and reduce design errors. In the XANDAR project, partners from industry and academia develop a toolchain that will advance current development processes by employing a modelbased X-by-Construction (XbC) approach. XANDAR defines a development process, metamodel extensions, a library of safety and security patterns, and investigates many further techniques for design automation, verification, and validation. The developed toolchain will use a hypervisor-based platform, targeting future centralized, AI-capable high-performance embedded processing systems. It is co-developed and validated in both an avionics use case for situation perception and pilot assistance as well as an automotive use case for autonomous driving. Leonard Masing, Tobias Dörr, Florian Schade, Jürgen Becker 0001, Georgios Keramidas, Christos P. Antonopoulos, Michail Mavropoulos, Efstratios Tiganourias, Vasilios I. Kelefouras, Konstantinos Antonopoulos, Nikos S. Voros, Umut Durak, Alexander Ahlbrecht, Wanja Zaeske, Christos Panagiotou, Dimitris Karadimas, Nico Adler, Andreas Sailer, Raphael Weber, Thomas Wilhelm 0005, Géza Németh, Fahad Siddiqui 0001, Rafiullah Khan, Vahid Garousi, Sakir Sezer, Victor Morales |
DATE | 9 |
| 2022 | XANDAR: A holistic Cybersecurity Engineering Process for Safety-critical and Cyber-physical SystemsabstractThe integration of connected and autonomous technologies in safety-critical and cyber-physical systems offers great potential in the vital application domains of transportation, manufacturing and aerospace. These technological advancements are necessary to meet the increasing demand for intelligent services, as they open doors to new business models by analysing and sharing the generated data. However, where this sharing of mix-critical data and broader connectivity brings opportunities, it simultaneously presents serious cybersecurity and safety risks due to the cyber-physical nature of these systems. Hence, delivering these intelligent services securely, safely, and reliably to its consumers is a complex engineering and design problem. One of the ways to approach this engineering problem is to consider both system functional and non-functional properties (safety, security, reliability) and systematically integrate them across system design and operational life cycle. The XANDAR project investigates this approach and aims to develop holistic software design methods and architectures for safety-critical and cyber-physical systems that guarantee functional and non-functional properties “byconstruction”. This paper focuses on the non-functional aspects of the project and discusses the preliminary work. by presenting the core cybersecurity principles and uses them as a baseline to propose a holistic cybersecurity engineering process. The tasks of the proposed cybersecurity engineering process are also map onto relevant clauses of ISO 21434. In future, proposed work will be integrated into the XANDAR software toolchain and validated for an avionics situation perception pilot assistance and automotive autonomous driving use cases. Fahad Siddiqui 0001, Rafiullah Khan, Sakir Sezer, Kieran McLaughlin, Leonard Masing, Tobias Dörr, Florian Schade, Jürgen Becker 0001, Alexander Ahlbrecht, Wanja Zaeske, Umut Durak, Nico Adler, Andreas Sailer, Raphael Weber, Thomas Wilhelm 0005, Géza Németh, Victor Morales, Paco Gomez, Georgios Keramidas, Christos P. Antonopoulos, Michail Mavropoulos, Vasilios I. Kelefouras, Konstantinos Antonopoulos, Nikos S. Voros, Christos Panagiotou, Dimitris Karadimas |
VTC Spring | 22 |
| 2022 | Workflow simulation and multi-threading aware task scheduling for heterogeneous computingabstractEfficient application scheduling is critical for achieving high performance in heterogeneous computing systems . This problem has proved to be NP-complete even for the homogeneous case , heading research efforts in obtaining low complexity heuristics that produce good quality schedules. Such an example is HEFT, one of the most efficient list scheduling heuristics in terms of makespan and robustness. In this paper, we propose two task scheduling methods for heterogeneous computing systems that can be integrated to several task scheduling algorithms . First, a method that improves the scheduling time (the time for obtaining the output schedule) of a family of task scheduling algorithms is delivered without sacrificing the schedule length, when the computation costs of the application tasks are unknown. Second, a method that improves the scheduling length (makespan) of several task scheduling algorithms is proposed, by identifying which tasks are going to be executed as single-threaded and which as multi-threaded implementations, as well as the number of the threads used. We showcase both methods by using HEFT popular algorithm, but they can be integrated to other algorithms too, such as HCPT, HPS, PETS and CPOP. The experimental results, which consider 14580 random synthetic graphs and five real world applications , show that by enhancing HEFT algorithm with the two proposed methods, significant makespan gains and high scheduling time gains, are achieved. Vasilios I. Kelefouras, Karim Djemame |
J. Parallel Distributed Comput. | 1 |
| 2022 | Design and Implementation of 2D Convolution on x86/x64 ProcessorsabstractIn this paper, a new method for accelerating the 2D direct Convolution operation on x86/x64 processors is presented. It includes efficient vectorization by using SIMD intrinsics, bit-twiddling optimizations, the optimization of the division operation, multi-threading using OpenMP, register blocking and the shortest possible bit-width value of the intermediate results. The proposed method, which is provided as open-source, is general and can be applied to other processor families too, e.g., Arm. The proposed method has been evaluated on two different multi-core Intel CPUs, by using twenty different image sizes, 8-bit integer computations and the most commonly used kernel sizes (3x3, 5x5, 7x7, 9x9). It achieves from$2.8\times$to$40\times$speedup over the Intel IPP library (OpenCV GaussianBlur and Filter2D routines), from$105 \times$to$400 \times$speedup over the gemm-based convolution method (by using Intel MKL int8 matrix multiplication routine), and from$8.5\times$to$618\times$speedup over the vslsConvExec Intel MKL direct convolution routine. The proposed method is superior as it achieves far fewer arithmetical and load/store instructions. Vasilios I. Kelefouras, Georgios Keramidas |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | XANDAR: X-by-Construction Design framework for Engineering Autonomous & Distributed Real-time Embedded Software SystemsabstractThe next generation of networked embedded systems (ES) necessitates rapid prototyping and high performance while maintaining key qualities like trustworthiness and safety. However, development of safety-critical ES suffers from complex software (SW) toolchains and engineering processes. Moreover, the current trend in autonomous systems, which relies on Machine Learning (ML) and AI applications when combined with fail-operational requirements renders the Verification and Validation (V&V) of these new systems a challenging endeavor. Prime examples are Advanced Driver-Assistance Systems (ADAS) that are prone to various safety/security vulnerabilities. The XANDAR project aims at developing a mature SW toolchain (from requirements analysis to the actual code integration on target including V&V) fulfilling the needs of industry for rapid prototyping of interoperable and autonomous ES. Starting from a model-based system architecture, XANDAR will leverage automatic model synthesis and software parallelization techniques to achieve specific non-functional requirements setting the foundation for a novel (real-time, safety-, and security)-by-Construction paradigm. Jürgen Becker 0001, Leonard Masing, Tobias Dörr, Florian Schade, Georgios Keramidas, Christos P. Antonopoulos, Michail Mavropoulos, Efstratios Tiganourias, Vasilios I. Kelefouras, Konstantinos Antonopoulos, Nikos S. Voros, Umut Durak, Alexander Ahlbrecht, Wanja Zaeske, Christos Panagiotou, Dimitris Karadimas, Nico Adler, Andreas Sailer, Raphael Weber, Thomas Wilhelm 0005, Florian Oszwald, Dominik Reinhardt, Mohamad Chamas, Adnan Bekan, Graham Smethurst, Fahad Siddiqui 0001, Rafiullah Khan, Vahid Garousi, Sakir Sezer, Victor Morales |
FPL | 9 |
| 2019 | A methodology correlating code optimizations with data memory accesses, execution time and energy consumption
Vasilios I. Kelefouras, Karim Djemame |
J. Supercomput. | 1 |
| 2018 | A methodology for efficient code optimizations and memory managementabstractThe key to optimizing software is the correct choice, order as well parameters of optimizations-transformations, which has remained an open problem in compilation research for decades for various reasons. First, most of the compilation subproblems-transformations are interdependent and thus addressing them separately is not effective. Second, it is very hard to couple the transformation parameters to the processor architecture (e.g., cache size and associativity) and algorithm characteristics (e.g. data reuse); therefore compiler designers and researchers either do not take them into account at all or do it partly. Third, the search space (all different transformation parameters) is very large and thus searching is impractical. Vasilios I. Kelefouras, Karim Djemame |
CF | 1 |
| 2018 | Workflow Simulation Aware and Multi-threading Effective Task Scheduling for Heterogeneous ComputingabstractEfficient application scheduling is critical for achieving high performance in heterogeneous computing systems. This problem has proved to be NP-complete, heading research efforts in obtaining low complexity heuristics that produce good quality schedules. Although this problem has been extensively studied in the past, all the related works assume the computation costs of application tasks on processors are available a priori, ignoring the fact that the time needed to run/simulate all these tasks is orders of magnitude higher than finding a good quality schedule, especially in heterogeneous systems. In this paper, we propose two new methods applicable to several task scheduling algorithms for heterogeneous computing systems. We showcase both methods by using HEFT well known and popular algorithm, but they are applicable to other algorithms too, such as HCPT, HPS, PETS and CPOP. First, we propose a methodology to reduce the scheduling time of HEFT when the computation costs are unknown, without sacrificing the length of the output schedule (monotonic computation costs); this is achieved by reducing the number of computation costs required by HEFT and as a consequence the number of simulations applied. Second, we give heuristics to find which tasks are going to be executed as Single-Thread and which as Multi-Thread CPU implementations, as well as the number of the threads used. The experimental results considering both random graphs and real world applications show that extending HEFT with the two proposed methods achieves better schedule lengths, while at the same time requires from 4.5 up to 24 less simulations. Vasilios I. Kelefouras, Karim Djemame |
HiPC | 1 |
| 2018 | Combining Software Cache Partitioning and Loop Tiling for Effective Shared Cache ManagementabstractOne of the biggest challenges in multicore platforms is shared cache management, especially for data-dominant applications. Two commonly used approaches for increasing shared cache utilization are cache partitioning and loop tiling. However, state-of-the-art compilers lack efficient cache partitioning and loop tiling methods for two reasons. First, cache partitioning and loop tiling are strongly coupled together, and thus addressing them separately is simply not effective. Second, cache partitioning and loop tiling must be tailored to the target shared cache architecture details and the memory characteristics of the corunning workloads. To the best of our knowledge, this is the first time that a methodology provides (1) a theoretical foundation in the above-mentioned cache management mechanisms and (2) a unified framework to orchestrate these two mechanisms in tandem (not separately). Our approach manages to lower the number of main memory accesses by an order of magnitude keeping at the same time the number of arithmetic/addressing instructions to a minimal level. We motivate this work by showcasing that cache partitioning, loop tiling, data array layouts, shared cache architecture details (i.e., cache size and associativity), and the memory reuse patterns of the executing tasks must be addressed together as one problem, when a (near)-optimal solution is requested. To this end, we present a search space exploration analysis where our proposal is able to offer a vast deduction in the required search space. Vasilios I. Kelefouras, Georgios Keramidas, Nikos S. Voros |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2016 | A high-performance matrix-matrix multiplication methodology for CPU and GPU architectures
Vasilios I. Kelefouras, Angeliki Kritikakou, Iosif Mporas, Vasileios Kolonias |
J. Supercomput. | 1 |
| 2016 | Array Size Computation under Uniform Overlapping and Irregular AccessesabstractThe size required to store an array is crucial for an embedded system, as it affects the memory size, the energy per memory access, and the overall system cost. Existing techniques for finding the minimum number of resources required to store an array are less efficient for codes with large loops and not regularly occurring memory accesses. They have to approximate the accessed parts of the array leading to overestimation of the required resources. Otherwise, their exploration time is increased with an increase over the number of the different accessed parts of the array. We propose a methodology to compute the minimum resources required for storing an array which keeps the exploration time low and provides a near-optimal result for regularly and non-regularly occurring memory accesses and overlapping writes and reads. Angeliki Kritikakou, Francky Catthoor, Vasilios I. Kelefouras, Constantinos E. Goutis |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2015 | A methodology for speeding up loop kernels by exploiting the software information and the memory architecture
Vasilios I. Kelefouras, Angeliki Kritikakou, Constantinos E. Goutis |
Comput. Lang. Syst. Struct. | 1 |
| 2015 | A methodology for speeding up matrix vector multiplication for single/multi-core architectures
Vasilios I. Kelefouras, Angeliki Kritikakou, Elissavet Papadima, Constantinos E. Goutis |
J. Supercomput. | 1 |
| 2014 | A scalable and near-optimal representation of access schemes for memory managementabstractMemory management searches for the resources required to store the concurrently alive elements. The solution quality is affected by the representation of the element accesses: a sub-optimal representation leads to overestimation and a non-scalable representation increases the exploration time. We propose a methodology to near-optimal and scalable represent regular and irregular accesses. The representation consists of a set of pattern entries to compactly describe the behavior of the memory accesses and of pattern operations to consistently combine the pattern entries. The result is a final sequence of pattern entries which represents the global access scheme without unnecessary overestimation. Angeliki Kritikakou, Francky Catthoor, Vasilios I. Kelefouras, Constantinos E. Goutis |
ACM Trans. Archit. Code Optim. | 3 |
| 2014 | A methodology for speeding up edge and line detection algorithms focusing on memory architecture utilization
Vasilios I. Kelefouras, Angeliki Kritikakou, Constantinos E. Goutis |
J. Supercomput. | 1 |
| 2014 | A Matrix-Matrix Multiplication methodology for single/multi-core architectures using SIMD
Vasilios I. Kelefouras, Angeliki Kritikakou, Constantinos E. Goutis |
J. Supercomput. | 1 |
| 2013 | Near-Optimal Microprocessor and Accelerators Codesign with Latency and Throughput ConstraintsabstractA systematic methodology for near-optimal software/hardware codesign mapping onto an FPGA platform with microprocessor and HW accelerators is proposed. The mapping steps deal with the inter-organization, the foreground memory management, and the datapath mapping. A step is described by parameters and equations combined in a scalable template. Mapping decisions are propagated as design constraints to prune suboptimal options in next steps. Several performance-area Pareto points are produced by instantiating the parameters. To evaluate our methodology we map a real-time bio-imaging application and loop-dominated benchmarks. Angeliki Kritikakou, Francky Catthoor, George Athanasiou, Vasilios I. Kelefouras, Constantinos E. Goutis |
ACM Trans. Archit. Code Optim. | 4 |
| 2013 | Near-optimal and scalable intrasignal in-place optimization for non-overlapping and irregular access schemesabstractStorage-size management techniques aim to reduce the resources required to store elements and to concurrently provide efficient addressing during element accessing. Existing techniques are less appropriate for large iteration spaces with increased numbers of irregularly spread holes. They either have to approximate the accessed regions, leading to overestimation of the final resources, or they require prohibited exploration time to find the storage size. In this work, we present a near-optimal and scalable methodology for storage-size, intrasignal, in-place optimization, that is, to compute the minimum amount of resources required to store the elements of a group (array), for irregular complex access schemes in the target domain of non-overlapping store and load accesses. Angeliki Kritikakou, Francky Catthoor, Vasilios I. Kelefouras, Constantinos E. Goutis |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2012 | A data locality methodology for matrix-matrix multiplication algorithm
Nikolaos Alachiotis 0002, Vasilios I. Kelefouras, George Athanasiou, Harris E. Michail, Angeliki Kritikakou, Constantinos E. Goutis |
J. Supercomput. | 2 |
| 2012 | On the exploitation of a high-throughput SHA-256 FPGA design for HMACabstractHigh-throughput and area-efficient designs of hash functions and corresponding mechanisms for Message Authentication Codes (MACs) are in high demand due to new security protocols that have arisen and call for security services in every transmitted data packet. For instance, IPv6 incorporates the IPSec protocol for secure data transmission. However, the IPSec's performance bottleneck is the HMAC mechanism which is responsible for authenticating the transmitted data. HMAC's performance bottleneck in its turn is the underlying hash function. In this article a high-throughput and small-size SHA-256 hash function FPGA design and the corresponding HMAC FPGA design is presented. Advanced optimization techniques have been deployed leading to a SHA-256 hashing core which performs more than 30% better, compared to the next better design. This improvement is achieved both in terms of throughput as well as in terms of throughput/area cost factor. It is the first reported SHA-256 hashing core that exceeds 11Gbps (after place and route in Xilinx Virtex 6 board). Harris E. Michail, George Athanasiou, Vasilios I. Kelefouras, George Theodoridis, Constantinos E. Goutis |
ACM Trans. Reconfigurable Technol. Syst. | 3 |