EDBT 2026 Demo / reviewers in the wild / expert
George Theodoridis
dblp:78/3760
· DBLP profile ↗
25ranked-venue papers
0as first author
4since 2021 · last 2027
0000-0002-2015-108XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 4 since 2021Security and privacy · 2Software engineering, systems software and programming languages · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | CECAIServe: Facilitating ML inference serving across the cloud-edge-continuumabstractMachine learning (ML) inference-serving has become a core operational component of MLOps, particularly as ML services are increasingly deployed across the cloud–edge continuum. However, existing inference-serving tools typically provide limited support for heterogeneous hardware, weak energy observability, and cumbersome integration of device-specific accelerated inference frameworks and model-specific pre-/post-processing. In this work, we present CECAIServe, an open-source, unified, and vendor-neutral framework that automatically generates deployment-ready Accelerated Inference Serving Containers (AISCs) from high-level TensorFlow and PyTorch models. Through its modular design, CECAIServe abstracts device- and framework-specific complexity while supporting CPUs, GPUs, edge accelerators, and FPGA-based systems. The framework further incorporates MLOps-oriented functionality, including fine-grained latency instrumentation, integrated power monitoring, and an interface for pre-/post-processing. Our comprehensive evaluation on 12 models demonstrates that CECAIServe can automatically generate AISCs across 7 diverse devices in less than 10 min. Additionally, we demonstrate that CECAIServe facilitates effective benchmarking and design space exploration on HW-accelerated inference-serving on devices across the cloud–edge-continuum. Finally, we validate the extensibility of the framework by adapting it to support Large Language Model inference-serving. Aimilios Leftheriotis, Achilleas Tzenetopoulos, George Lentaris, Dimitrios Soudris, George Theodoridis |
Future Gener. Comput. Syst. | 5 |
| 2025 | Multi-Partner Project: Secure Hardware Accelerated Data Analytics for 6G Networks: The PRIVATEER ApproachabstractNext generation 6G networks are designed to meet the requirements of modern applications, including the need for higher bandwidth and ultra-low latency services. While these networks show significant potential to fulfill these evolving connectivity needs, they also bring new challenges, particularly in the area of security. Meanwhile, ensuring the privacy is paramount in 6G network development, demanding robust solutions following “privacy-by-design” principles. To address these challenges, PRIVATEER project strengthens existing security mechanisms, introducing privacy-centric enablers tailored for 6G networks. This work, evaluates key enablers within PRIVATEER, focusing on the development and acceleration of AI -driven anomaly detection models, as well as attestation mechanisms for both hardware accelerators and containerized applications. Ilias Papalamprou, Aimilios Leftheriotis, Apostolis Garos, Georgios Gardikis, Maria Christopoulou, Georgios Xilouris, Lampros Argyriou, Antonia Karamatskou, Emmanouil Kalotychos, Nikolaos Chatzivasileiadis, Dimosthenis Masouros, George Theodoridis, Dimitrios Soudris |
DATE | 13 |
| 2022 | A Multi-stage Hybrid Approach for Mapping Applications on Heterogeneous Multi-core PlatformsabstractDue to the incorporation of heterogeneous cores in modern multi-core systems, the exploitation of their full potential strongly depends on the proper mapping of an application to the platform. This work presents an approach to map static applications on heterogeneous platforms minimizing their makespan based on the Benders decomposition principle combined with an Integer Linear Programming (ILP) model. The proposed approach adopts a three-stage decomposition scheme, finding permutations of infeasible solutions to generate multiple cuts in every iteration. The first stage deals with the assignment of the tasks to the cores and the last one with their scheduling, whereas the second stage propagates new bounds based on the current assignment and provides an explanation of the infeasibility in the form of subsets of assignment variables. Based on that, other infeasible combinations are computed by checking their permutations and more Benders cuts are produced. The proposed method is compared with a two-stage decomposition approach and an ILP model and exhibits better performance in terms of solution time and number solved instances to optimality. Andreas Emeretlis, George Theodoridis, Panayiotis Alefragis, Nikos S. Voros |
VLSI-SoC | 2 |
| 2022 | An FPGA implementation of the VESA Display Stream Compression decoderabstractA HW architecture for the implementation of the DSC decoder is proposed. It demands 1 cycle/pixel (cpp) for 4:4:4 chroma sub-sampling format and 0.5 cpp for 4:2:2 or 4:2:0 formats. Also, it supports 3:1 lossless compression ratio and operates with sub-line latency. To achieve the above, specific optimizations have been applied at the algorithmic and design levels to efficiently distribute the operations on the pipeline stages and improve frequency. Its implementation on a Kintex Ultrascale+ device achieves at 262 MHz for 10-bit input sample, and it can process 32 High Dynamic Range (HDR) frames per second (Fps) for 4:4:4 chroma format and 60 HDR Fps for 4:2:2 or 4:2:0 formats. Nikolaos Kefalas, George Theodoridis |
VLSI-SoC | 2 |
| 2020 | Exploring the FPGA Implementations of the LBlock, Piccolo, Twine, and Klein CiphersabstractIn this work, the implementations of the LBlock, Piccolo, Twine, and Klein lightweight ciphers in FPGA technology are studied in terms of area, frequency, throughput, and throughput/area. To accomplish this, loop unrolling and pipelining were employed in two phases. In the first phase, different loop unrolling factors were used to implement the round function of each cipher, while in the second phase, 2-stage pipelining with loop unrolling per stage was applied. The produced designs were implemented in Xilinx (Kintex-7) FPGA technology. Based on the implementation results, a detailed study on the above-mentioned design metrics was performed and important outcomes were derived. S. Moraitis, D. Seitanidis, George Theodoridis, Odysseas G. Koufopavlou |
VLSI-SOC | 3 |
| 2019 | Low-memory and high-performance architectures for the CCSDS 122.0-B-1 compression standard
Nikolaos Kefalas, George Theodoridis |
Integr. | 2 |
| 2018 | Static Mapping of Applications on Heterogeneous Multi-Core Platforms Combining Logic-Based Benders Decomposition with Integer Linear ProgrammingabstractThe proper mapping of an application on a multi-core platform and the scheduling of its tasks are key elements to achieve the maximum performance. In this article, a novel hybrid approach based on integrating the Logic-Based Benders Decomposition (LBBD) principle with a pure Integer Linear Programming (ILP) model is introduced for mapping applications described by Directed Acyclic Graphs (DAGs) on platforms consisting of heterogeneous cores. The LBBD approach combines two optimization techniques with complementary strengths, namely ILP and Constraint Programming (CP), and is employed as a cut generation scheme. The generated constraints are utilized by the ILP model to cut possible assignment combinations aiming at improving the solution or proving the optimality of the best-found one. The introduced approach was applied both on synthetic DAGs and on DAGs derived from real applications. Through the proposed approach, many problems were optimally solved that could not be solved by any of the above methods (ILP, LBBD) alone within a time limit of 2 hours, while the overall solution time was also significantly decreased. Specifically, the hybrid method exhibited speedups equal to 4.2× for the synthetic instances and 10× for the real-application DAGs over the LBBD approach and two orders of magnitude over the ILP model. Andreas Emeretlis, George Theodoridis, Panayiotis Alefragis, Nikos S. Voros |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2016 | A Logic-Based Benders Decomposition Approach for Mapping Applications on Heterogeneous Multicore PlatformsabstractThe development of efficient methods for mapping applications on heterogeneous multicore platforms is a key issue in the field of embedded systems. In this article, a novel approach based on the Logic-Based Benders decomposition principle is introduced for mapping complex applications on these platforms, aiming at optimizing their execution time. To provide optimal solutions for this problem in a short time, a new hybrid model that combines Integer Linear Programming (ILP) and Constraint Programming (CP) models is introduced. Also, to reduce the complexity of the model and its solution time, a set of novel techniques for generating additional constraints called Benders cuts is proposed. An extensive set of experiments has been performed in which synthetic applications described by Directed Acyclic Graphs (DAGs) were mapped to a number of heterogeneous multicore platforms. Moreover, experiments with DAGs that correspond to two real-life applications have also been performed. Based on the experimental results, it is proven that the proposed approach outperforms the pure ILP model in terms of the solution time and quality of the solution. Specifically, the proposed approach is able to find an optimal solution within a time limit of 2 hours in the vast majority of performed experiments, while the pure ILP model fails. Also, for the cases where both methods fail to find an optimal solution within the time limit, the solution of the proposed approach is systematically better than the solution of the ILP model. Andreas Emeretlis, George Theodoridis, Panayiotis Alefragis, Nikos S. Voros |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2015 | A high performance 5 stage pipeline architecture for the H.264/AVC deblocking filter
Nikolaos Kefalas, George Theodoridis |
Integr. | 2 |
| 2014 | On the development of high-throughput and area-efficient multi-mode cryptographic hash designs in FPGAs
Harris E. Michail, George Athanasiou, George Theodoridis, Constantinos E. Goutis |
Integr. | 3 |
| 2012 | High-throughput Hardware Architectures of the JH Round-three SHA-3 Candidate - An FPGA Design and Implementation Approach
George Athanasiou, Chara I. Chalkou, D. Bardis, Harris E. Michail, George Theodoridis, Constantinos E. Goutis |
SECRYPT | 5 |
| 2012 | On the Development of Totally Self-checking Hardware Design for the SHA-1 Hash Function
Harris E. Michail, George Athanasiou, Andreas Gregoriades, George Theodoridis, Constantinos E. Goutis |
SECRYPT | 4 |
| 2012 | On the exploitation of a high-throughput SHA-256 FPGA design for HMACabstractHigh-throughput and area-efficient designs of hash functions and corresponding mechanisms for Message Authentication Codes (MACs) are in high demand due to new security protocols that have arisen and call for security services in every transmitted data packet. For instance, IPv6 incorporates the IPSec protocol for secure data transmission. However, the IPSec's performance bottleneck is the HMAC mechanism which is responsible for authenticating the transmitted data. HMAC's performance bottleneck in its turn is the underlying hash function. In this article a high-throughput and small-size SHA-256 hash function FPGA design and the corresponding HMAC FPGA design is presented. Advanced optimization techniques have been deployed leading to a SHA-256 hashing core which performs more than 30% better, compared to the next better design. This improvement is achieved both in terms of throughput as well as in terms of throughput/area cost factor. It is the first reported SHA-256 hashing core that exceeds 11Gbps (after place and route in Xilinx Virtex 6 board). Harris E. Michail, George Athanasiou, Vasilios I. Kelefouras, George Theodoridis, Constantinos E. Goutis |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2009 | An Application Development Framework for ARISE Reconfigurable ProcessorsabstractCoupling reconfigurable hardware accelerators with processors is an effective way to meet the performance and flexibility required to cope with modern embedded applications. The ARISE framework provides a systematic approach to extend a processor once. It will thereafter support the coupling of arbitrary hardware accelerators. The accelerators can be coupled as coprocessors or functional units of the processor’s datapath, and therefore exploited as a hybrid, which includes both loose and tight computational models. This article presents a complete framework for developing applications on such hybrid reconfigurable ARISE machines. The framework integrates the automatic identification of custom instructions and the semiautomatic/profiling-driven identification of coprocessors supporting the hybrid computational model. Moreover, it supports a modular design approach where the software and the hardware modules are developed independently and later ported into any ARISE machine with reconfigurable technology. To evaluate efficiency, a set of benchmarks is implemented on an ARISE evaluation machine utilizing the proposed framework. In addition, the ARISE machine is compared against a well-established processor paradigm that utilizes reconfigurable accelerators following only the typical coprocessor approach. Experimental results prove that the framework can be used to exploit the hybrid computational model and achieve significant performance improvements over the typical coprocessor acceleration approach. Moreover, results demonstrate how the framework can be used to trade off performance, silicon area, and application development time. Nikolaos Vassiliadis, George Theodoridis, Spiridon Nikolaidis 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2009 | The ARISE Approach for Extending Embedded Processors With Arbitrary Hardware AcceleratorsabstractARISE introduces a systematic approach for extending once an embedded processor to support thereafter the coupling of an arbitrary number of custom computing units (CCUs). A CCU can be a hardwired or a reconfigurable unit, which can be utilized following a tight and/or loose model of computation. By selecting the appropriate model of computation for each part of the application, the complete application space is considered for acceleration, resulting in significant performance improvements. Also, ARISE offers modularity and scalability and is not restricted by the opcode space and operands limitation problems that exist in such type of machines. To support these features we introduce a machine organization that allows the cooperation of a processor and a set of CCUs. To control the CCUs we extend once the instruction set of the processor with eight instructions. To efficiently incorporate these features to an embedded processor, we propose a micro-architecture implementation that minimizes the control and communication overhead between the processor and the CCUs. To evaluate our proposal, we extended a MIPS processor with the ARISE infrastructure and implemented it on a Xilinx field-programmable gate array (FPGA). Implementation results, demonstrate that the timing model of the processor is not affected. Also, we implemented a set of benchmarks on the ARISE evaluation machine. Performance results prove significant improvements and reduced communication overhead compared to a typical coprocessor approach. Nikolaos Vassiliadis, George Theodoridis, Spiridon Nikolaidis 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2006 | An automated development framework for a RISC processor with reconfigurable instruction set extensionsabstractBy coupling a reconfigurable hardware to a standard processor, high levels of flexibility and adaptability are achieved. However, this approach requires modifications to the compiler of the processor to take into account reconfigurable aspects. In this paper, a development framework for a RISC processor with reconfigurable instruction set extensions is presented. The framework is fully automated, hiding all reconfigurable related issues from the user and can be used for both program and fine-tune the architecture at design time. We demonstrate the above issues using a set of benchmarks. Experimental results show an x2.9 average speedup in addition to potential energy reduction Nikolaos Vassiliadis, George Theodoridis, Spiridon Nikolaidis 0001 |
IPDPS | 2 |
| 2006 | A high-performance data path for synthesizing DSP kernelsabstractA high-performance data path to implement digital signal processing (DSP) kernels is introduced in this paper. The data path is realized by a flexible computational component (FCC), which is a pure combinational circuit and it can implement any 2 times 2 template (cluster) of primitive resources. Thus, the data path's performance benefits from the intracomponent chaining of operations. Due to the flexible structure of the FCC, the data path is implemented by a small number of such components. This allows for direct connections among FCCs and for exploiting intercomponent chaining, which further improves performance. Due to the universality and flexibility of the FCC, simple and efficient algorithms perform scheduling and binding of the data flow graph (DFG). DSP benchmarks synthesized with the FCC data path method show significant performance improvements when compared with template-based data path designs. Detailed results on execution time, FCC utilization, and area are presented Michalis D. Galanis, George Theodoridis, Spyros Tragoudas, Constantinos E. Goutis |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2006 | High-Speed FPGA Implementation of Secure Hash Algorithm for IPSec and VPN Applications
Athanasios Kakarountas, Harris E. Michail, Athanasios Milidonis, Constantinos E. Goutis, George Theodoridis |
J. Supercomput. | 5 |
| 2004 | A Partitioning Methodology for Accelerating Applications in Hybrid Reconfigurable PlatformsabstractIn this paper, we propose a methodology for partitioning and mapping computational intensive applications in reconfigurable hardware blocks of different granularity. A generic hybrid reconfigurable architecture is considered so as the methodology can be applicable to a large number of heterogeneous reconfigurable platforms. The methodology mainly consists of two stages, the analysis and the mapping of the application onto fine and coarse-grain hardware resources. A prototype framework consisting of analysis, partitioning and mapping tools has been also developed. For the coarse-grain reconfigurable hardware, we use our previously developed high-performance coarse-grain datapath. In this work, the methodology is validated using two real-world applications, an OFDM transmitter and a JPEG encoder. In the case of the OFDM transmitter, a maximum clock cycle decrease of 82 % relative to the ones in an all fine-grain mapping solution is achieved. The corresponding performance improvement for the JPEG is 43 %. Michalis D. Galanis, Athanasios Milidonis, George Theodoridis, Dimitrios Soudris, Constantinos E. Goutis |
DATE | 3 |
| 2004 | Accelerating DSP Applications on a Mixed Granularity Platform with a New Reconfigurable Coarse-Grain Data-PathabstractIn this paper, a high performance reconfigurable coarse-grain data-path, part of a mixed-granularity reconfigurable platform, is presented. The computational resources are coarse grain components of the same type. An automated methodology for mapping DSP applications on the data-path is also presented, and it is based on unsophisticated, yet efficient, algorithms. Results on DSP benchmarks show the performance improvements over previously published high-performance data-paths. Michalis D. Galanis, George Theodoridis, Spyros Tragoudas, Dimitrios Soudris, Constantinos E. Goutis |
FCCM | 2 |
| 2004 | A novel coarse-grain reconfigurable data-path for accelerating DSP kernelsabstractIn this paper, an efficient implementation of a high performance coarse-grain reconfigurable data-path on a mixed-granularity reconfigurable platform is presented. It consists of several coarse grain components of the same type, a reconfigurable inter-component network, and a centralized register bank. The universal type of coarse grain component is shown to increase the system's performance due to significant reductions in the latency. A flexible interconnection network facilitates the data transfers between the coarse grain components and also from or to the register bank. An automated methodology for mapping DSP and multimedia kernels on the data-path is also presented. Chaining of operations is optimally exploited, and the architecture allows for simple and efficient algorithms for scheduling, live signal reduction, and component binding. Experimental results verify the impact of our architectural decisions and design automation methods. Michalis D. Galanis, George Theodoridis, Spyros Tragoudas, Dimitrios Soudris, Constantinos E. Goutis |
FPGA | 2 |
| 2004 | Mapping DSP Applications to a High-Performance Reconfigurable Coarse-Grain Data-Path
Michalis D. Galanis, George Theodoridis, Spyros Tragoudas, Dimitrios Soudris, Constantinos E. Goutis |
FPL | 2 |
| 2004 | An Automated C++ Code and Data Partitioning Framework for Data Management of Data-Intensive Applications
Athanasios Milidonis, Grigoris Dimitroulakos, Michalis D. Galanis, George Theodoridis, Constantinos E. Goutis, Francky Catthoor |
SCOPES | 4 |
| 2002 | A fast and accurate delay dependent method for switching estimation of large combinational circuits
Spyros Theoharis, George Theodoridis, Dimitrios Soudris, Constantinos E. Goutis, Adonios Thanailakis |
J. Syst. Archit. | 2 |
| 2000 | Low power design of a multi-mode transceiverabstractRecent advances in electronic technology integration coupled with increasing needs for more services in portable communications favors the development of high performance dual-mode terminals. We present the complete architecture implementation of the GMSK/GFSK modulator/demodulator including the FIR filters design. The main features of the modulator/demodulator and the architectural implementation of FIR filters are described. The interface with ASPIS processor and A/D & D/A converters is also described in detail manner. The whole architecture of the modulator/demodulator was described by VHDL hardware language, synthesised and implemented in Xilinx environment. Dimitrios Soudris, Minas Perakis, Haris Mizas, Vasilios A. Mardiris, Kosfas Katis, Chrissavgi Dre, A. E. Tzimas, E. G. Metaxakis, Grigorios Kalivas, Nikolaos D. Zervas, Spyros Theoharis, George Theodoridis, Adonios Thanailakis, Constantinos E. Goutis |
ISCAS | 12 |