VLDB 2026 Research / reviewers in the wild / expert
Ioannis Papaefstathiou
dblp:19/4489 · also Yannis Papaefstathiou
· DBLP profile ↗
101ranked-venue papers
8as first author
16since 2021 · last 2026
0000-0001-6386-5616ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 64 · 3 first-author · 11 since 2021Computer networks · 20 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 13 · 2 first-author · 1 since 2021Security and privacy · 7 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Open-Source Distributed Simulation Framework for RISC-V Systems incorporating Vector and Cryptographic Extensions
Nikolaos Tampouratzis, Ioannis Papaefstathiou |
CF | 2 |
| 2026 | An efficient open-source design and implementation framework for non-quantized CNNs on FPGAsabstractThe growing demand for real-time processing in artificial intelligence applications, particularly those involving Convolutional Neural Networks (CNNs), has highlighted the need for efficient computational solutions. Conventional processors and graphical processing units (GPUs), very often, fall short in balancing performance, power consumption, and latency, especially in embedded systems and edge computing platforms. Field-Programmable Gate Arrays (FPGAs) offer a promising alternative, combining high performance with energy efficiency and reconfigurability. This paper presents a design and implementation framework for implementing CNNs seamlessly on FPGAs that maintains full precision in all neural network parameters thus addressing a niche, that of non-quantized NNs. The presented framework extends Darknet, which is very widely used for the design of CNNs, and allows the designer, by effectively using a Darknet NN description, to efficiently implement CNNs in a heterogeneous system comprising of CPUs and FPGAs. Our framework is evaluated on the implementation of a number of different CNNs and as part of a real world application utilizing UAVs; in all cases it outperforms the CPU and GPU systems in terms of performance and/or power consumption. When compared with the FPGA frameworks that support quantization, our solution offers similar performance and/or energy efficiency without any degradation on the NN accuracy. • An open-source library to accelerate CNN algorithms. • High design productivity, flexibility and adaptability of DarkNet design framework. • Achieve high performance by exploitation of the full parallelism of any FPGA. • Power efficiency of CNN inference on FPGAs for power-sensitive applications. Angelos Athanasiadis, Nikolaos Tampouratzis, Ioannis Papaefstathiou |
Integr. | 3 |
| 2025 | CyberNEMO: Enhancing End-to-End Cybersecurity and Privacy in the IoT-Edge-Cloud ContinuumabstractAccording to the EU State of Cybersecurity report by the European Union Agency for Cybersecurity (ENISA), the number of cybersecurity-related incidents will increase by 24 percent by 2025, with ransomware and DoS/DDoS attacks being the most common. The emergence of new threats [1] and the consolidation of existing ones require doubling of efforts in proactive prevention and a decisive increase in research dedicated to cybersecurity. CyberNEMO (End-to-end Cybersecurity to NEMO meta-OS) project emerges as an evolution of the NEMO (Next Generation Meta Operating System) platform, designed to provide a secure, trustworthy, and robust execution environment across the IoT-Edge-Cloud computing continuum. Leveraging NEMO modular meta-operating system (mOS) framework, CyberNEMO introduces advanced cybersecurity and privacy-preserving mechanisms, emphasizing Zero Trust principles. This paper presents the CyberNEMO architecture, details its core innovative technologies, and describes its validation strategy through diverse living labs-including Smart Energy, Smart Water, Smart Manufacturing, Healthcare, Multimedia Distribution, and Smart Farming scenarios-demonstrating end-to-end cybersecurity and real-time threat mitigation capabilities, aligned with Europe's strategic cybersecurity goals. Theodore B. Zahariadis, Artemis C. Voulkidis, Ilias Nektarios Seitanidis, Andreas E. Papadakis, Alberto del Río, Javier Serrano 0003, David Jiménez, Antonio Pastor 0001, Diego R. López, Alejandro Muñiz, Mattin Antartiko Elorza Forcada, Ana Méndez, Wafa Ben Jaballah, Rosaria Rossini, Maria Belesioti, Ioannis P. Chochliouros, Marco Angelini, Vasileios Megalooikonomou, Carmela Occhipinti, Luigi Briguglio, Alexandru Plesa, Vladut Dinu, Mohammad Ghoreishi, Mostafa Jabari, Dimitrios Skias, Konstantinos Sakatis, Ioannis Papaefstathiou |
SRDS | 27 |
| 2025 | Operationalizing cybersecurity knowledge: Design, implementation & evaluation of a knowledge management system for CACAO playbooks
Orestis Tsirakis, Konstantinos Fysarakis, Vasileios Mavroeidis, Ioannis Papaefstathiou |
Comput. Secur. | 4 |
| 2025 | Distributed fast and accurate simulation platform for advanced ARM- and RISC-V-based HPC systems
Nikolaos Tampouratzis, Ioannis Papaefstathiou, Gabriel Gomez-Lopez, Miguel Sánchez de la Rosa, Jesús Escudero-Sahuquillo, Pedro Javier García |
J. Supercomput. | 2 |
| 2024 | REBECCA: Reconfigurable Heterogeneous Highly Parallel Processing Platform for Safe and Secure AI
Andreas Brokalakis, Iakovos Mavroidis, Konstantinos Georgopoulos, Pavlos Malakonakis, Konstantinos Harteros, Dimitris Andronikou, Yannis Galanomatis, Charalampos Savvakos, Grigorios Chrysos 0001, Sotiris Ioannidis, Ioannis Papaefstathiou |
DSD | 11 |
| 2024 | Fast, Accurate and Distributed Simulation of novel HPC systems incorporating ARM and RISC-V CPUsabstractThe growing developments of HPC systems used in a plethora of domains (healthcare, financial services, government and defense, energy) triggers an urgent demand for simulation frameworks that can simulate, in an integrated manner, both processing and network components of an HPC system-under-design (SuD). The main problem, however, is that, currently, there is a shortage of simulation frameworks that can handle the simulation of actual HPC systems, including the hardware, complete software stack and network dynamics in an integrated manner. In this work we start from the first known, open-source, fully-distributed Cloud simulation framework, COSSIM, and, as part of the RED-SEA1 and Vitamin-V2 European projects, we extend it so as to be able to accurately simulate HPC systems. The extended simulator has been evaluated when executing the very-widely used HPCG & LAMMPS benchmarks on both ARM & RISC-V architectures; the results demonstrate that the presented approach has up to 95% accuracy in the reported SuD aspects. Nikolaos Tampouratzis, Ioannis Papaefstathiou |
HPDC | 2 |
| 2023 | eProcessor: European, Extendable, Energy-Efficient, Extreme-Scale, Extensible, Processor EcosystemabstractThe eProcessor project aims at creating a RISC-V full stack ecosystem. The eProcessor architecture combines a high-performance out-of-order core with energy-efficient accelerators for vector processing and artificial intelligence with reduced-precision functional units. The design of this architecture follows a hardware/software co-design approach with relevant application use cases from the high-performance computing, bioinformatics and artificial intelligence domains. Two eProcessor prototypes will be developed based on two fabricated eProcessor ASICs integrated into a computer-on-module. Lluc Alvarez, Abraham Ruiz, Arnau Bigas-Soldevilla, Pavel Kuroedov, Alberto González 0004, Hamsika Mahale, Noe Bustamante, Albert Aguilera, Francesco Minervini, Javier Salamero, Oscar Palomar, Vassilis Papaefstathiou, Antonis Psathakis, Nikolaos Dimou, Michalis Giaourtas, Iasonas Mastorakis, Giorgos Ieronymakis, Georgios-Michail Matzouranis, Vassilis Flouris, Nikolaos Kossifidis, Manolis Marazakis, Bhavishya Goel, Madhavan Manivannan, Ahsen Ejaz, Panagiotis Strikos, Mateo Vázquez, Ioannis Sourdis, Pedro Trancoso, Per Stenström, Jens Hagemeyer, Lennart Tigges, Nils Kucza, Jean-Marc Philippe, Ioannis Papaefstathiou |
CF | 34 |
| 2023 | Early Results of Mapping Industrial Applications on Heterogeneous HPC Systems: The OPTIMA ProjectabstractThe OPTIMA project aims to port and optimize industrial applications and a set of open-source libraries into two novel FPGA-populated HPC systems. Target applications are from the domains of robotics simulation, underground analysis and computational fluid dynamics (CFD), where data processing is based on differential equations, matrix-matrix and matrix-vector operations. Moreover, the OPTIMA OPen Source (OOPS) library will support basic linear algebraic operations, sparse matrix-vector arithmetic, as well as computer-aided engineering (CAE) solvers. The OPTIMA target platforms are JUMAX, an HPC system that couples an AMD Epyc Server with Maxeler FPGA-based Dataflow Engines (DFEs), and server class machines with Alveo FPGA cards installed. Experimental results show that performance on robotic simulation can be enhanced up to 1.2x, and CFD calculations up to 4.7x. Finally, BLAS L1 routines are improved up to 7x, with a performance-per-Watt ratio boost of more than 40x compared to multi-threaded software routines from the Intel Math Kernel Library (MKL) suite when executed on an Intel Xeon server-class machine. Dimitris Theodoropoulos 0001, Giorgos Pekridis, Panagiotis Miliadis, Chloe Alverti, Panagiotis Mpakos, Dionisios N. Pnevmatikatos, Pavlos Malakonakis, Konstantinos Georgopoulos, Iakovos Mavroidis, Gino Perna, Marisa Zanotti, Giovanni Isotton, Max Engelen, Aggelos Ioannou, Ioannis Papaefstathiou, Albert Kahira, Andreas Herten |
CF | 15 |
| 2023 | Optimizing Industrial Applications for Heterogeneous HPC Systems: The OPTIMA Project Intermediate stageabstractOPTIMA is an SME-driven project (intermediate stage) that aims to port and optimize industrial applications and a set of open-source libraries into two novel FPGA-populated HPC systems. Target applications are from the domain of robotics simulation, underground analysis and computational fluid dy-namics (CFD), where data processing is based on differential equations, matrix-matrix and matrix-vector operations. Moreover, the OPTIMA OPen Source (OOPS) library will support basic linear algebraic operations, sparse matrix-vector arithmetic, as well as computer-aided engineering (CAE) solvers. The OPTIMA target platforms are JUMAX, an HPC system that couples an AMD Epyc Server with Maxeler FPGA-based Dataflow Engines (DFEs), and server-class machines with Alveo FPGA cards in-stalled. Experimental results on applications up to now, show that performance on robotic simulation can be enhanced up to 1.2x, CFD calculations up to 4.7x, and BLAS routines up to 7x compared to optimized software implementations from OpenBLAS. Dimitris Theodoropoulos 0001, Pavlos Malakonakis, Konstantinos Georgopoulos, Giovanni Isotton, Dionisios N. Pnevmatikatos, Ioannis Papaefstathiou, Gino Perna, Marisa Zanotti, Panagiotis Miliadis, Panagiotis Mpakos, Chloe Alverti, Aggelos Ioannou, Max Engelen, Albert Kahira, Iakovos Mavroidis |
DATE | 7 |
| 2023 | VITAMIN-V: Virtual Environment and Tool-Boxing for Trustworthy Development of RISC-V Based Cloud ServicesabstractVITAMIN-V is a 2023–2025 Horizon Europe project that aims to develop a complete RISC-V open-source software stack for cloud services with comparable performance to the cloud-dominant x86 counterpart and a powerful virtual execution environment for software development, validation, verification, and testing that considers the relevant RISC-VISA extensions for cloud deployment. VITAMIN-V will specifically support the RISC-V extensions for virtualization, cryptography, and vec-torization in three virtual environments: QEMU, gem5, and cloud FPGA prototype platforms. The project will focus on European Processor Initiative (EPI) based RISC-V designs and accelerators. VITAMIN-V will also support the ISA extensions by adding the compiler and toolchain support. Furthermore, it will develop novel software validation, verification, and testing approaches to ensure software trustworthiness. To enable the execution of complete cloud stacks, VITAMIN-V will port all necessary machine-dependent modules in relevant open-source cloud software distributions, focusing on three cloud setups. Finally, VITAMIN-V will demonstrate and benchmark these three cloud setups using relevant AI, big-data, and serverless applications. VITAMIN-V aims to match the software performance of its x86 equivalent while contributing to RISC-V open-source virtual environments, software validation, and cloud software suites. Ramon Canal, Cristiano Pegoraro Chenet, Aggelos Arelakis, José-María Arnau, Josep Lluís Berral, Aaron Call, Stefano Di Carlo, Juan José Costa, Dimitris Gizopoulos, Vasileios Karakostas, Francesco Lubrano, Konstantinos Nikas, Yiannis Nikolakopoulos, Beatriz Otero, George Papadimitriou 0001, Ioannis Papaefstathiou, Dionisios N. Pnevmatikatos, Daniel Raho, Alvise Rigo, Eva Rodríguez, Alessandro Savino 0001, Alberto Scionti, Nikolaos Tampouratzis, Alex Torregrosa |
DSD | 16 |
| 2023 | A Novel Integrated Simulation Framework for Cyber-Physical Systems ModellingabstractThe growing use of Cyber-Physical Systems (CPS) in a plethora of domains (e.g. healthcare, industry, smart homes, transportation, etc.) triggers an urgent demand for simulation frameworks that can simulate in an integrated manner all the components (i.e. CPUs, Memories, Networks, Physical Environment) of a system-under-design(SuD). By utilizing such a simulator, software design can proceed in parallel with physical development which results in the reduction of the so important time-to-market. The main problem, however, is that currently there is a shortage of such simulation frameworks; most simulators used for modelling the digital aspects of CPS applications (i.e. full-system CPU/Mem/Peripheral simulators) lack any support of the CPS physical aspects and vice versa. The presented fully-distributed simulation framework (APOLLON) is the first known open-source, high-performance simulator that can handle holistically complex CPSs including processors, peripherals, networks and physical aspects of them. APOLLON is an extension of the COSSIM simulation framework and it integrates, in a novel and efficient way, a combined processing and network simulator with the widely-used Ptolemy physical simulator, in a transparent way. Our highly integrated approach is further augmented with Machine Learning capabilities by implementing Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) recurrent neural networks in both the Cyber and Physical domains, enabling users to develop their complex recurrent neural networks significantly fast and accurately. APOLLON has been evaluated when executing a number of benchmarks and real-world use cases; the end results demonstrate that the presented approach has up to 99% accuracy in the reported SuD aspects. Nikolaos Tampouratzis, Panagiotis Mousouliotis, Ioannis Papaefstathiou |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | RED-SEA: Network Solution for Exascale ArchitecturesabstractIn order to enable Exascale computing, next generation interconnection networks must scale to hundreds of thousands of nodes, and must provide features to also allow the HPC, HPDA, and AI applications to reach Exascale, while benefiting from new hardware and software trends. RED-SEA will pave the way to the next generation of European Exascale interconnects, including the next generation of BXI, as follows: (i) specify the new architecture using hardware-software co-design and a set of applications representative of the new terrain of converging HPC, HPDA, and AI; (ii) test, evaluate, and/or implement the new architectural features at multiple levels, according to the nature of each of them, ranging from mathematical analysis and modeling, to simulation, or to emulation or implementation on FPGA testbeds; (iii) enable seamless communication within and between resource clusters, and therefore development of a high-performance low latency gateway, bridging seamlessly with Ethernet; (iv) add efficient network resource management, thus improving congestion resiliency, virtualization, adaptive routing, collective operations; (v) open the interconnect to new kinds of applications and hardware, with enhancements for end-to-end network services - from programming models to reliability, security, low- latency, and new processors; (vi) leverage open standards and compatible APIs to develop innovative reusable libraries and Fabrics management solutions. Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Pier Stanislao Paolucci, Elena Pastorelli, Francesco Simula, Matteo Turisini, Piero Vicini, Roberto Ammendola, Pascale Bernier-Bruna, Said Derradji, Stéphane Guez, Pierre-Axel Lagadec, Gregoire Pichon, Etienne Walter, Gaetan De Gassowski, Matthieu Hautreaux, Stephane Mathieu, Gilles Moreau, Marc Pérache, Hugo Taboada, Torsten Hoefler, Timo Schneider, Matteo Barnaba, Giuseppe Piero Brandino, Francesco De Giorgi, Matteo Poggi, Iakovos Mavroidis, Ioannis Papaefstathiou, Nikolaos Tampouratzis, Benjamin Kalisch, Ulrich Krackhardt, Mondrian Nüssle, Pantelis Xirouchakis, Vangelis Mageiropoulos, Michalis Gianioudis, Harisis Loukas, Aggelos Ioannou, Nikolaos D. Kallimanis, Nikolaos Chrysos, Manolis Katevenis, Wolfgang Frings, Dominik Gottwald, Felime Guimaraes, Max Holicki, Volker Marx, Yannik Müller, Carsten Clauss, Hugo Falter, Xu Huang 0010, Jennifer Lopez Barillao, Thomas Moschny, Simon Pickartz, Francisco J. Alfaro, Jesús Escudero-Sahuquillo, Pedro Javier García, Francisco J. Quiles 0001, José L. Sánchez 0002, Adrián Castelló 0001, Jose Duro, María Engracia Gómez, Enrique S. Quintana-Ortí, Julio Sahuquillo, Eugenio Stabile |
DSD | 33 |
| 2022 | Making Citizens' systems more Secure: Practical Encryption Bypassing and CountermeasuresabstractCryptography is used to protect the confidentiality, integrity, and authenticity of information by preventing unauthorized users from accessing or modifying them. Encryption techniques are used to protect personal or company data. This work demonstrates practical scenarios where, under certain conditions, encryption may be bypassed. Bypassing encryption, either by recovering the encryption key, a password used to generate the encryption key, or a plaintext copy of the encrypted data, allows for accessing data which appear to be inaccessible in the first place. There are six categories for bypassing encryption: find the key, guess the key, compel the key, exploit a flaw in the encryption scheme, access unencrypted message when the device is in use and locate an unencrypted copy of the message. In this study we utilize publicly available software to demonstrate real-world scenarios that fall into most of the aforementioned categories and show how, in those specific cases, encryption may be successfully bypassed. Moreover, we underline that bypassing encryption is possible only when certain conditions are met (e.g., software misconfiguration, physical access to the target device, etc.) and we highlight each one of them so as to effectively suggest countermeasures to the demonstrated techniques for encryption bypassing. The main aim of this paper is to highlight how encryption can be bypassed and thus make citizens set up their system in such a way that it would be more difficult to be hacked. This is especially important for citizens that may have limited knowledge/exposure to technology as they can be, for example. people from certain diversity groups such as elderly and/or people of very low income. Marios Adam Sirgiannis, Charalampos Manifavas, Ioannis Papaefstathiou |
ISCC | 3 |
| 2021 | An Open-source Implementation of LSTM and GRU in the Ptolemy Simulation FrameworkabstractPtolemy II [1] is an open-source software framework for modelling, simulation and design of concurrent, heterogeneous, real-time systems, including distributed/parallel systems [2]. These systems can be described using the combination of different formal as well as computation models. Ptolemy also includes machine learning libraries providing support for particle filtering, model-predictive control, hidden Markov models, and various statistical analysis tools. However, one of the main problems Ptolemy users face is the lack of fundamental recurrent neural network structures. In this paper, an LSTM and a GRU recurrent neural network are implemented in Ptolemy II framework, in order to extend its machine learning library, enabling users to develop their complex recurrent neural networks in significantly less time. The presented work has been verified through a real-world weather forecasting use case; the results demonstrate that our approach has identical accuracy with one of the most widely used machine learning library (i.e. Keras) in all cases. To further increase the impact of our approach, the complete source code is freely distributed to the community. Vasilis Daoulas, Nikolaos Tampouratzis, Panagiotis Mousouliotis, Ioannis Papaefstathiou |
DS-RT | 4 |
| 2021 | Novel Reconfigurable Hardware Systems for Tumor Growth PredictionabstractAn emerging trend in biomedical systems research is the development of models that take full advantage of the increasing available computational power to manage and analyze new biological data as well as to model complex biological processes. Such biomedical models require significant computational resources, since they process and analyze large amounts of data, such as medical image sequences. We present a family of advanced computational models for the prediction of the spatio-temporal evolution of glioma and their novel implementation in state-of-the-art FPGA devices. Glioma is a rapidly evolving type of brain cancer, well known for its aggressive and diffusive behavior. The developed system simulates the glioma tumor growth in the brain tissue, which consists of different anatomic structures, by utilizing MRI slices. The presented models have been proved highly accurate in predicting the growth of the tumor, whereas the developed innovative hardware system, when implemented on a low-end, low-cost FPGA, is up to 85% faster than a high-end server consisting of 20 physical cores (and 40 virtual ones) and more than 28× more energy-efficient than it; the energy efficiency grows up to 50× and the speedup up to 14× if the presented designs are implemented in a high-end FPGA. Moreover, the proposed reconfigurable system, when implemented in a large FPGA, is significantly faster than a high-end GPU (i.e., from 80% and up to 250% faster), for the majority of the models, while it is also significantly better (i.e., from 80% to over 1,600%) in terms of power efficiency, for all the implemented models. Konstantinos Malavazos, Maria Papadogiorgaki, Pavlos Malakonakis, Ioannis Papaefstathiou |
ACM Trans. Comput. Heal. | 4 |
| 2020 | A novel FPGA-based system for Tumor Growth PredictionabstractAn emerging trend in the biomedical community is to create models that take advantage of the increasing available computational power, in order to manage and analyze new biological data as well as to model complex biological processes. Such biomedical software applications require significant computational resources since they process and analyze large amounts of data, such as medical image sequences. This paper presents a novel FPGA-based system that implements a novel model for the prediction of the spatio-temporal evolution of glioma. Glioma is a rapidly evolving type of brain cancer, well known for its aggressive and diffusive behavior. The developed system simulates the glioma tumor growth in the brain tissue, which consists of different anatomic structures, by utilizing individual MRI slices. The presented innovative hardware system is more than 60% faster than a high-end server consisting of 20 physical cores (and 40 virtual ones) and more than 28x more energy efficient. Konstantinos Malavazos, Maria Papadogiorgaki, Pavlos Malakonakis, Ioannis Papaefstathiou |
DATE | 4 |
| 2020 | SqueezeJet-3: An Accelerator Utilizing FPGA MPSoCs for Edge CNN ApplicationsabstractMost FPGA-based Convolutional Neural Network (CNN) hardware accelerators target the datacenter rather than edge processing units. To further fill this gap, this work presents SqueezeJet-3 - a novel FPGA-based embedded system, consisting of software and hardware, for accelerating edge CNN inference. Even though SqueezeJet-3 is optimized for accelerating small ImageNet class CNNs, such as SqueezeNet v1.1, on low-end lowcost FPGA SoC devices, it can also be used for the acceleration of larger CNNs, such as the VGG16. Evaluation of our accelerator reveals better or comparable performance with that triggered by the current state-of-the-art similar tools and systems. Panagiotis Mousouliotis, Ioannis Papaefstathiou, Loukas Petrou |
FCCM | 2 |
| 2020 | A Novel, Highly Integrated Simulator for Parallel and Distributed SystemsabstractIn an era of complex networked parallel heterogeneous systems, simulating independently only parts, components, or attributes of a system-under-design is a cumbersome, inaccurate, and inefficient approach. Moreover, by considering each part of a system in an isolated manner, and due to the numerous and highly complicated interactions between the different components, the system optimization capabilities are severely limited. The presented fully-distributed simulation framework (called as COSSIM) is the first known open-source, high-performance simulator that can handle holistically system-of-systems including processors, peripherals and networks; such an approach is very appealing to both Cyber Physical Systems (CPS) and Highly Parallel Heterogeneous Systems designers and application developers. Our highly integrated approach is further augmented with accurate power estimation and security sub-tools that can tap on all system components and perform security and robustness analysis of the overall system under design—something that was unfeasible up to now. Additionally, a sophisticated Eclipse-based Graphical User Interface (GUI) has been developed to provide easy simulation setup, execution, and visualization of results. COSSIM has been evaluated when executing the widely used Netperf benchmark suite as well as a number of real-world applications. Final results demonstrate that the presented approach has up to 99% accuracy (when compared with the performance of the real system), while the overall simulation time can be accelerated almost linearly with the number of CPUs utilized by the simulator. Nikolaos Tampouratzis, Ioannis Papaefstathiou, Antonis Nikitakis, Andreas Brokalakis, Stamatis Andrianakis, Apostolos Dollas, Marco Marcon, Emanuele Plebani |
ACM Trans. Archit. Code Optim. | 2 |
| 2020 | UNILOGIC: A Novel Architecture for Highly Parallel Reconfigurable SystemsabstractOne of the main characteristics of High-performance Computing (HPC) applications is that they become increasingly performance and power demanding, pushing HPC systems to their limits. Existing HPC systems have not yet reached exascale performance mainly due to power limitations. Extrapolating from today’s top HPC systems, about 100–200 MWatts would be required to sustain an exaflop-level of performance. A promising solution for tackling power limitations is the deployment of energy-efficient reconfigurable resources (in the form of Field-programmable Gate Arrays (FPGAs)) tightly integrated with conventional CPUs. However, current FPGA tools and programming environments are optimized for accelerating a single application or even task on a single FPGA device. In this work, we present UNILOGIC (Unified Logic), a novel HPC-tailored parallel architecture that efficiently incorporates FPGAs. UNILOGIC adopts the Partitioned Global Address Space (PGAS) model and extends it to include hardware accelerators, i.e., tasks implemented on the reconfigurable resources. The main advantages of UNILOGIC are that (i) the hardware accelerators can be accessed directly by any processor in the system, and (ii) the hardware accelerators can access any memory location in the system. In this way, the proposed architecture offers a unified environment where all the reconfigurable resources can be seamlessly used by any processor/operating system. The UNILOGIC architecture also provides hardware virtualization of the reconfigurable logic so that the hardware accelerators can be shared among multiple applications or tasks. The FPGA layer of the architecture is implemented by splitting its reconfigurable resources into (i) a static partition, which provides the PGAS-related communication infrastructure, and (ii) fixed-size and dynamically reconfigurable slots that can be programmed and accessed independently or combined together to support both fine and coarse grain reconfiguration. 1 Finally, the UNILOGIC architecture has been evaluated on a custom prototype that consists of two 1U chassis, each of which includes eight interconnected daughter boards, called Quad-FPGA Daughter Boards (QFDBs); each QFDB supports four tightly coupled Xilinx Zynq Ultrascale+ MPSoCs as well as 64 Gigabytes of DDR4 memory, and thus, the prototype features a total of 64 Zynq MPSoCs and 1 Terabyte of memory. We tuned and evaluated the UNILOGIC prototype using both low-level (baremetal) performance tests, as well as two popular real-world HPC applications, one compute-intensive and one data-intensive. Our evaluation shows that UNILOGIC offers impressive performance that ranges from being 2.5 to 400 times faster and 46 to 300 times more energy efficient compared to conventional parallel systems utilizing only high-end CPUs, while it also outperforms GPUs by a factor ranging from 3 to 6 times in terms of time to solution, and from 10 to 20 times in terms of energy to solution. Aggelos Ioannou, Konstantinos Georgopoulos, Pavlos Malakonakis, Dionisios N. Pnevmatikatos, Vassilis Papaefstathiou, Ioannis Papaefstathiou, Iakovos Mavroidis |
ACM Trans. Reconfigurable Technol. Syst. | 6 |
| 2019 | An Open-Source High-Throughput, Reduced Memory Footprint, Face Detection, Pose Estimation and Landmark Localization SystemabstractFace Detection, Pose Estimation and Landmark Localization are all considered important vision processes and are very widely utilized in several applications ranging from security/safety to automotive and assisted living. In this paper we present an open-source optimized implementation of a system addressing all those processes. In particular, we optimized both the memory requirements and the performance of the very widely utilized Tree Structure Model (TSM), which is the main core in all those tasks. Several optimizations have been proposed so as to increase the performance in both uni-and multi-processor systems, while also reducing the memory footprint so as to allow for the implementation of those schemes, for the first time, in embedded systems. The proposed system is at least 100% faster (and up to 300%) and requires more than 10 time less memory than the existing implementations. Since the algorithm implemented is one of the most widely used in the area of face detection and we distribute the optimized code in an open-source manner, we believe that it can act as an important reference implementation for any similar system proposed. Panos Kalodimas, Antonis Nikitakis, Ioannis Papaefstathiou |
DSD | 3 |
| 2019 | Accelerating Physics Engine Components with Embedded FPGAsabstractIn recent years there has been a steady increase in the use of physics engines, deployed in applications such as video games, scientific simulations, computer graphics and film productions. Their main purpose is to simulate the motions of objects based on real-world physics rules. As the complexity of the simulated scenes increases with the use of multiple objects and desirable effects, the computational cost of the physics-related calculations explodes. Typically, physics engines make use of the general-purpose computational capabilities of modern GPUs in order to take advantage of their massively parallel resources. In this paper, we consider the use of FPGAs to accelerate certain demanding components of the physics simulation pipeline aiming to provide better performing solutions at significantly lower energy cost. The results of our work demonstrate that by employing Zynq UltraScale+ devices featuring embedded ARM cores and FPGA fabric, we can accelerate physics computations of the popular Bullet library on highly demanding scenes up to 2.2x compared to high-end GPUs at a fraction of the energy required (up to 44x better energy efficiency). Petros Toupas, Andreas Brokalakis, Ioannis Papaefstathiou |
FPL | 3 |
| 2018 | COSSIM: An Open-Source Integrated Solution to Address the Simulator Gap for Systems of SystemsabstractIn an era of complex networked heterogeneous systems, simulating independently only parts, components or attributes of a system under design is not a viable, accurate or efficient option. The interactions are too many and too complicated to produce meaningful results and the optimization opportunities are severely limited when considering each part of a system in an isolated manner. The presented COSSIM simulation framework is the first known open-source, high-performance simulator that can handle holistically system-of-systems including processors, peripherals and networks; such an approach is very appealing to both CPS/IoT and Highly Parallel Heterogeneous Systems designers and application developers. Our highly integrated approach is further augmented with accurate power estimation and security sub-tools that can tap on all system components and perform security and robustness analysis of the overall networked system. Additionally, a GUI has been developed to provide easy simulation set-up, execution and visualization of results. COSSIM has been evaluated using real-world applications representing cloud (mobile visual search) and CPS systems (building management) demonstrating high accuracy and performance that scales almost linearly with the number of CPUs dedicated to the simulator. Andreas Brokalakis, Nikolaos Tampouratzis, Antonis Nikitakis, Ioannis Papaefstathiou, Stamatis Andrianakis, Danilo Pau, Emanuele Plebani, Marco Paracchini, Marco Marcon, Ioannis Sourdis, Prajith Ramakrishnan Geethakumari, Maria Carmen Palacios, Miguel Ángel Antón, Attila Szasz |
DSD | 4 |
| 2018 | SeMIBIoT: Secure Multi-Protocol Integration Bridge for the IoTabstractThe Internet of Things (IoT) is gradually becoming a reality, supported by an assortment of heterogeneous devices, varying from resource- starved wireless sensors to embedded devices and resource-rich backend systems, which are supplemented by a range of networking technologies and protocols. Nevertheless, this diverse ecosystem of platforms and protocols, along with the inherent limitations in processing power, energy, memory and communications bandwidth for some of the involved devices, render secure and interoperable interactions a primary concern, and an important obstacle to the introduction of novel applications, and the adoption of IoT in general. Motivated by the above, this work presents SeMIBIoT, a Secure Multi-protocol Integration Bridge for the IoT. Acting as a gateway, SeMIBIoT is able to provide hop-by-hop or end-to-end secure communications between an array of heterogeneous nodes and standardized IoT protocols, guaranteeing seamless interactions and thus alleviating security and interoperability concerns. Using a realistic testbed featuring a number of diverse platforms, the bridge is evaluated in a variety of scenarios, validating the feasibility of the proposed approach. Emmanouil Palavras, Konstantinos Fysarakis, Ioannis Papaefstathiou, Ioannis G. Askoxylakis |
ICC | 3 |
| 2018 | AmbISPDM - Managing embedded systems in ambient environments and disaster mitigation planning
George Hatzivasilis, Ioannis Papaefstathiou, Dimitris Plexousakis, Charalampos Manifavas, Nikos Papadakis |
Appl. Intell. | 2 |
| 2018 | The Industrial Internet of Things as an enabler for a Circular Economy Hy-LP: A novel IIoT protocol, evaluated on a wind park's SDN/NFV-enabled 5G industrial network
George Hatzivasilis, Konstantinos Fysarakis, Othonas Sultatos, Ioannis G. Askoxylakis, Ioannis Papaefstathiou, Giorgos Demetriou |
Comput. Commun. | 5 |
| 2018 | XSACd - Cross-domain resource sharing & access control for smart environments
Konstantinos Fysarakis, Othonas Sultatos, Charalampos Manifavas, Ioannis Papaefstathiou, Ioannis G. Askoxylakis |
Future Gener. Comput. Syst. | 4 |
| 2018 | Data-Driven Background Subtraction Algorithm for In-Camera Acceleration in Thermal ImageryabstractDetection of moving objects in videos is a crucial step toward successful surveillance and monitoring applications. A key component for such tasks is called background subtraction and tries to extract regions of interest from the image background for further processing or action. For this reason, its accuracy and real-time performance are of great significance. Although effective background subtraction methods have been proposed, only a few of them take into consideration the special characteristics of thermal imagery. In this paper, we propose a background subtraction scheme, which models the thermal responses of each pixel as a mixture of Gaussians with unknown number of components. Following a Bayesian approach, our method automatically estimates the mixture structure, while simultaneously it avoids over-/underfitting. The pixel density estimate is followed by an efficient and highly accurate updating mechanism, which permits our system to be automatically adapted to dynamically changing operation conditions. We propose a reference implementation of our method in reconfigurable hardware achieving both adequate performance and low-power consumption. Adopting a high-level synthesis design and demanding floating point arithmetic operations are mapped in reconfigurable hardware, demonstrating fast prototyping and on-field customization at the same time. Konstantinos Makantasis, Antonis Nikitakis, Anastasios Doulamis, Nikolaos D. Doulamis, Ioannis Papaefstathiou |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2017 | A novel way to efficiently simulate complex full systems incorporating hardware acceleratorsabstractThe breakdown of Dennard scaling coupled with the persistently growing transistor counts severally increased the importance of application-specific hardware acceleration; such an approach offers significant performance and energy benefits compared to general-purpose solutions. In order to thoroughly evaluate such architectures, the designer should perform a quite extensive design space exploration so as to evaluate the trade-offs across the entire system. The design, until recently, has been predominantly done using Register Transfer Level (RTL) languages such as Verilog and VHDL, which, however, lead to a prohibitively long and costly design effort. In order to reduce the design time a wide range of both commercial and academic High-Level Synthesis (HLS) tools have emerged; most of those tools, handle hardware accelerators that are described in synthesisable SystemC. The problem today, however, is that most simulators used for evaluating the complete user applications (i.e. full-system CPU/Mem/Peripheral simulators) lack any type of SystemC accelerator support. Within this context this paper presents a novel simulation environment comprised of a generic SystemC accelerator and probably the most widely known fullsystem simulator (i.e. GEM5). The proposed system is the only solution supporting the very important feature of global synchronization across the integrated simulation; furthermore it has been evaluated based on two different computationally-intensive use cases and the final results demonstrate that the presented approach is orders of magnitude faster than the existing ones. Nikolaos Tampouratzis, Konstantinos Georgopoulos, Ioannis Papaefstathiou |
DATE | 3 |
| 2017 | An Open-Source Extendable, Highly-Accurate and Security Aware CPS SimulatorabstractIn this paper, we present an open-source Cyber Physical Systems (CPS) simulation framework that aims to address the limitations of currently available tools. Our solution models the computing devices of the processing nodes and the network that comprise the CPS system and thus provides cycle accurate results, realistic communications and power/energy consumption estimates based on the actual dynamic usage scenarios. The simulator provides the necessary hooks to security testing software and can be extended through an IEEE standardized interface to include additional tools, such as simulators of physical models. Andreas Brokalakis, Nikolaos Tampouratzis, Antonis Nikitakis, Stamatis Andrianakis, Ioannis Papaefstathiou, Apostolos Dollas |
DCOSS | 5 |
| 2017 | An Architecture for the Acceleration of a Hybrid Leaky Integrate and Fire SNN on the Convey HC-2ex FPGA-Based ProcessorabstractNeuromorphic computing is expanding by leaps and bounds through custom integrated circuits (digital and analog), and large scale platforms developed by industry or government-funded projects (e.g. TrueNorth and BrainScaleS, respectively). Whereas the trend is for massive parallelism and neuromorphic computation in order to solve problems, such as those that may appear in machine learning and deep learning algorithms, there is substantial work on brain-like highly accurate neuromorphic computing in order to model the human brain. In such a form of computing, spiking neural networks (SNN) such as the Hodgkin and Huxley model are mapped to various technologies, including FPGAs. In this work, we present a highly efficient FPGA-based architecture for the detailed hybrid Leaky Integrate and Fire SNN that can simulate generic characteristics of neurons of the cerebral cortex. This architecture supports arbitrary, sparse O(n2) interconnection of neurons without need to re-compile the design, and plasticity rules, yielding on a four-FPGA Convey 2ex hybrid computer a speedup of 923x for a non-trivial data set on 240 neurons vs. the same model in the software simulator BRAIN on a Intel(R) Xeon(R) CPU E5-2620 v2 @ 2.10GHz, i.e. the reference state-of-the-art software. Although the reference, official software is single core, the speedup demonstrates that the application scales well among multiple FPGAs, whereas this would not be the case in general-purpose computers due to the arbitrary interconnect requirements. The FPGA-based approach leads to highly detailed models of parts of the human brain up to a few hundred neurons vs. a dozen or fewer neurons on the reference system. Emmanouil Kousanakis, Apostolos Dollas, Euripides Sotiriades, Ioannis Papaefstathiou, Dionisios N. Pnevmatikatos, Athanasia Papoutsi, Panagiotis Petrantonakis, Panayiota Poirazi, Spyridon Chavlis, George Kastellakis |
FCCM | 4 |
| 2017 | SecRoute: End-to-end secure communications for wireless ad-hoc networksabstractRailways constitute a main means of mass transportation, used by public, private, and military entities to traverse long distances every day. Railway control software must collect spatial information and effectively manage these systems. Wireless sensor networks (WSNs) are an attractive solution to cover the area along-side the railway routes. In-carriage WSNs are also studied in cases of dangerous cargo transportation. The secure communication of all these devices becomes important as successful attacks can harm the railway's business operation or cause serious injuries and deaths. This paper presents SecRoute - an end-to-end secure communications scheme for wireless ad hoc networks. The scheme implements mechanisms for cryptographic communication, trusted-based routing, and policy-based access control. SecRoute and alternative schemes are modelled on the NS-2 network simulator and a comparative analysis is conducted, indicating that the proposed scheme provides enhanced protection. A proof of concept of SecRoute is deployed on real embedded platforms and exhibits good overall performance, demonstrating that attacks on the route and carriage WSNs are effectively countered. George Hatzivasilis, Ioannis Papaefstathiou, Konstantinos Fysarakis, Ioannis G. Askoxylakis |
ISCC | 2 |
| 2017 | Lightweight & secure industrial IoT communications via the MQ telemetry transport protocolabstractMassive advancements in computing and communication technologies have enabled the ubiquitous presence of interconnected computing devices in all aspects of modern life, forming what is typically referred to as the “Internet of Things”. These major changes could not leave the industrial environment unaffected, with “smart” industrial deployments gradually becoming a reality; a trend that is often referred to as the 4th industrial revolution or Industry 4.0. Nevertheless, the direct interaction of the smart devices with the physical world and their resource constraints, along with the strict performance, security, and reliability requirements of industrial infrastructures, necessitate the adoption of lightweight as well as secure communication mechanisms. Motivated by the above, this paper highlights the Message Queue Telemetry Transport (MQTT) as a lightweight protocol suitable for the industrial domain, presenting a comprehensive evaluation of different security mechanisms that could be used to protect the MQTT-enabled interactions on a real testbed of wireless sensor motes. Moreover, the applicability of the proposed solutions is assessed in the context of a real industrial application, analyzing the network characteristics and requirements of an actual, operating wind park, as a representative use case of industrial networks. Sotirios Katsikeas, Konstantinos Fysarakis, Andreas I. Miaoudakis, Amaury Van Bemten, Ioannis G. Askoxylakis, Ioannis Papaefstathiou, Anargyros Plemenos |
ISCC | 6 |
| 2017 | SCOTRES: Secure Routing for IoT and CPSabstractWireless ad-hoc networks are becoming popular due to the emergence of the Internet of Things and cyber-physical systems (CPSs). Due to the open wireless medium, secure routing functionality becomes important. However, the current solutions focus on a constrain set of network vulnerabilities and do not provide protection against newer attacks. In this paper, we propose SCOTRES-a trust-based system for secure routing in ad-hoc networks which advances the intelligence of network entities by applying five novel metrics. The energy metric considers the resource consumption of each node, imposing similar amount of collaboration, and increasing the lifetime of the network. The topology metric is aware of the nodes' positions and enhances load-balancing. The channel-health metric provides tolerance in periodic malfunctioning due to bad channel conditions and protects the network against jamming attacks. The reputation metric evaluates the cooperation of each participant for a specific network operation, detecting specialized attacks, while the trust metric estimates the overall compliance, safeguarding against combinatorial attacks. Theoretic analysis validates the security properties of the system. Performance and effectiveness are evaluated in the network simulator 2, integrating SCOTRES with the DSR routing protocol. Similar schemes are implemented using the same platform in order to provide a fair comparison. Moreover, SCOTRES is deployed on two typical embedded system platforms and applied on real CPSs for monitoring environmental parameters of a rural application on olive groves. As is evident from the above evaluations, the system provides the highest level of protection while retaining efficiency for real application deployments. George Hatzivasilis, Ioannis Papaefstathiou, Charalampos Manifavas |
IEEE Internet Things J. | 2 |
| 2016 | ECOSCALE: Reconfigurable computing and runtime system for future exascale systems
Iakovos Mavroidis, Ioannis Papaefstathiou, Luciano Lavagno, Dimitrios S. Nikolopoulos, Dirk Koch, John Goodacre, Ioannis Sourdis, Vassilis Papaefstathiou, Marcello Coppola, Manuel Palomino |
DATE | 2 |
| 2016 | Highly efficient reconfigurable parallel graph cuts for embedded vision
Antonis Nikitakis, Ioannis Papaefstathiou |
DATE | 2 |
| 2016 | A novel background subtraction scheme for in-camera acceleration in thermal imagery
Antonis Nikitakis, Ioannis Papaefstathiou, Konstantinos Makantasis, Anastasios Doulamis |
DATE | 2 |
| 2016 | IoT design course using open-source toolsabstractOne of the most promising areas in Computer Engineering is certainly that of the Internet of Things (IoT) Systems. Moreover, such systems can be efficiently implemented in current Field Programmable Gate Arrays (FPGAs) comprising of either soft-core or hard-core CPUs. The design of an IoT system comprises of developments in both hardware and software whereas a productive IoT Systems Designer should have numerous competencies in various independent computer engineering fields together with a number of non-technical competencies such as interpersonal competencies, transferable skills, etc. This paper describes a course which is based on open-source tools and methodologies that has been utilized efficiently in a number of post-graduate programs in two European countries (Greece, and Spain). The course lasts only one week (5 hours a day) and it allows computer engineering graduates with different specialties to implement an FPGA-based complete real-world IoT system as a course project. Ioannis Papaefstathiou |
EDUCON | 1 |
| 2016 | Which IoT Protocol? Comparing Standardized Approaches over a Common M2M ApplicationabstractComputing devices already permeate working and living environments, while researchers and engineers aim to exploit the potential of pervasive systems in order to introduce new types of services and address inveterate and emerging problems. This process will lead us eventually to the era of urban computing and the Internet of Things (IoT). However, the long-promised improvements require overcoming some significant obstacles introduced by these technological advancements. One such obstacle is the lack of interoperable solutions, to facilitate the use, monitoring and management of the plethora of devices and their services. While seamless machine-to-machine (M2M) and human-to-machine (H2M) interactions are a necessity for secure and truly ubiquitous computing, the current status quo is that of a segregated and incompatible assortment of devices. The resource-constraints of the platforms integrated into smart environments, and their heterogeneity in hardware, network and overlaying technologies, only exacerbate these interoperability issues. Motivated by the above, this paper identifies three promising, standardized protocols, each following a different approach in addressing the above concerns. We evaluate the selected protocols in the context of designing and implementing an application requiring various M2M interactions, namely a policy-based access control framework for IoT devices. Thus, three variants of the application are developed, considering each protocol's intrinsic characteristics and features. Finally, the developed applications are evaluated on a common testbed of embedded devices, allowing us to extract useful conclusions concerning the protocols' performance, their intricacies and their applicability in similar applications. Konstantinos Fysarakis, Ioannis G. Askoxylakis, Othonas Sultatos, Ioannis Papaefstathiou, Charalampos Manifavas, Vasilios Katos |
GLOBECOM | 4 |
| 2016 | A survey of lightweight stream ciphers for embedded systemsabstractPervasive computing constitutes a growing trend, aiming to embed smart devices into everyday objects. The limited resources of these devices and the ever-present need for lower production costs, lead to the research and development of lightweight cryptographic mechanisms. Block ciphers, the main symmetric key cryptosystems, perform well in this field. Nevertheless, stream ciphers are also relevant in ubiquitous computing applications, as they can be used to secure the communication in applications where the plaintext length is either unknown or continuous, like network streams. This paper provides the latest survey of stream ciphers for embedded systems. Lightweight implementations of stream ciphers in embedded hardware and software are examined as well as relevant authenticated encryption schemes. Their speed and simplicity enable compact and low-power implementations, allow them to excel in applications pertaining to resource-constrained devices. The outcomes of the International Organization for Standardization/International Electrotechnical Commission 29192-3 standard and the cryptographic competitions eSTREAM and Competition for Authenticated Encryption: Security, Applicability, and Robustness are summarized along with the latest results in the field. However, cryptanalysis has proven many of these schemes are actually insecure. From the 31 designs that are examined, only six of them have been found to be secure by independent cryptanalysis. A constrained benchmark analysis is performed on low-cost embedded hardware and software platforms. The most appropriate and secure solutions are then mapped in different types of applications. Copyright © 2015 John Wiley & Sons, Ltd. Charalampos Manifavas, George Hatzivasilis, Konstantinos Fysarakis, Ioannis Papaefstathiou |
Secur. Commun. Networks | 4 |
| 2016 | Accelerating Intercommunication in Highly Parallel SystemsabstractEvery HPC system consists of numerous processing nodes interconnect using a number of different inter-process communication protocols such as Messaging Passing Interface (MPI) and Global Arrays (GA). Traditionally, research has focused on optimizing these protocols and identifying the most suitable ones for each system and/or application. Recently, there has been a proposal to unify the primitive operations of the different inter-processor communication protocols through the Portals library. Portals offer a set of low-level communication routines which can be composed in order to implement the functionality of different intercommunication protocols. However, Portals modularity comes at a performance cost, since it adds one more layer in the actual protocol implementation. This work aims at closing the performance gap between a generic and reusable intercommunication layer, such as Portals, and the several monolithic and highly optimized intercommunication protocols. This is achieved through the development of a novel hardware offload engine efficiently implementing the basic Portals’ modules. Our innovative system is up to two2 orders of magnitude faster than the conventional software implementation of Portals’ while the speedup achieved over the conventional monolithic software implementations of MPI and GAs is more than an order of magnitude. The power consumption of our hardware system is less than 1/100th of what a low-power CPU consumes when executing the Portal's software while its silicon cost is less than 1/10th of that of a very simple RISC CPU. Moreover, our design process is also innovative since we have first modeled the hardware within an untimed virtual prototype which allowed for rapid design space exploration; then we applied a novel methodology to transform the untimed description into an efficient timed hardware description, which was then transformed into a hardware netlist through a High-Level Synthesis (HLS) tool. Nikolaos Tampouratzis, Pavlos M. Mattheakis, Ioannis Papaefstathiou |
ACM Trans. Archit. Code Optim. | 3 |
| 2015 | WSACd - A Usable Access Control Framework for Smart Home Devices
Konstantinos Fysarakis, Charalampos Konstantourakis, Konstantinos Rantos, Charalampos Manifavas, Ioannis Papaefstathiou |
WISTP | 5 |
| 2015 | Lightweight Password Hashing Scheme for Embedded Systems
George Hatzivasilis, Ioannis Papaefstathiou, Charalampos Manifavas, Ioannis G. Askoxylakis |
WISTP | 2 |
| 2015 | Embedded Systems Security: A Survey of EU Research EffortsabstractAbstract Embedded systems security is a recurring theme in current research efforts, brought in the limelight by the wide adoption of ubiquitous devices. Significant funding has been allocated to various European projects on this subject area, in order to investigate and overcome the various security challenges. This paper provides an overview of recent EU research efforts pertaining to embedded systems security, where several prominent security issues and the respective proposed approaches are presented. Surveying relatively recent EU research projects, the authors identify 20 such projects that focus on embedded systems security aspects. The investigated technologies are categorised using a layered approach, to facilitate the presentation of the results; the categories comprise the node, network, and middleware and overlay layers, as well as architectures, frameworks and formal validation of the security of embedded systems. From this survey, certain patterns emerge regarding the issues investigated and the technologies researchers focus on, in order to address the said issues. Finally, the existing open issues are summarised, and directions for future research are given. Copyright © 2014 John Wiley & Sons, Ltd. Charalampos Manifavas, Konstantinos Fysarakis, Alexandros Papanikolaou, Ioannis Papaefstathiou |
Secur. Commun. Networks | 4 |
| 2014 | ModConTR: A modular and configurable trust and reputation-based system for secure routing in ad-hoc networksabstractDistributed wireless networks have become popular due to the evolution of the Internet-of-Things. These networks utilize ad-hoc routing protocols to interconnecting all nodes. Each peer forwards data for other nodes on the basis of network connectivity and a set of conventions that is determined by the routing protocol. Still, these protocols fail to protect legitimate nodes against several types of selfish and malicious activity. Thus, trust and reputation schemes are integrated with pure routing protocols to provide secure routing functionality. In this paper we propose ModConTR - a modular and adaptable trust and reputation-based system for secure routing. The system is composed of 11 different components which can be configured at runtime to adjust to each application's security and performance requirements. Presented work includes three possible configurations of ModConTR, considering ultra-lightweight, low-cost and lightweight implementations. Moreover, predefined configurations permit the implementation of the reasoning process for well-known secure routing protocols. Thus, we present a security and performance analysis for each of the components, including a comparative analysis of 10 complete trust and reputation schemes under identical attack scenarios. ModConTR is implemented using the NS2 simulator and is integrated with the DSR routing protocol. George Hatzivasilis, Ioannis Papaefstathiou, Charalampos Manifavas |
AICCSA | 2 |
| 2014 | A novel embedded system for vision trackingabstractOne of the most important challenges in the field of Computer Vision is the implementation of low-power embedded systems that will execute very accurate, yet real-time, algorithms. In the visual tracking sector one of the most promising approaches is the recently introduced OpenTLD algorithm which uses a random forest classification method. While it is very robust, it cannot be efficiently parallelized in its native form as its memory access pattern has certain characteristics that make it hard to take advantage of the conventional memory hierarchies. In this paper, we present a novel embedded system implementing this algorithm. We accelerate the bottleneck of the algorithm by designing and implementing a high bandwidth distributed memory sub-system which is independent of the various software parameters. We demonstrate the applicability and efficiency of this novel approach by implementing our scheme in a modern FPGA. Antonis Nikitakis, Theofilos Paganos, Ioannis Papaefstathiou |
DATE | 3 |
| 2014 | Policy-based access control for DPWS-enabled ubiquitous devicesabstractAs computing becomes ubiquitous, researchers and engineers aim to exploit the potential of the pervasive systems in order to introduce new types of services and address inveterate and emerging problems. This process will, eventually, lead us to the era of urban computing and the Internet of Things; the ultimate goal being to improve our quality of life. But these concepts typically require direct and constant interaction of computing systems with the physical world in order to be realized, which inevitably leads to the introduction of a range of safety and privacy issues that must be addressed. One such important aspect is the fine-grained control of access to the resources of these pervasive embedded systems, in a secure and scalable manner. This paper presents an implementation of such a secure policy-based access control scheme, focusing on the use of well-established, standardized technologies and considering the potential resource-constraints of the target heterogeneous embedded devices. The proposed framework adopts a DPWS-compliant approach for smart devices and introduces XACML-based access control mechanisms. The proof-of-concept implementation is presented in detail, along with a performance evaluation on typical embedded platforms. Konstantinos Fysarakis, Ioannis Papaefstathiou, Charalampos Manifavas, Konstantinos Rantos, Othonas Sultatos |
ETFA | 2 |
| 2014 | HPC-gSpan: An FPGA-based parallel system for frequent subgraph miningabstractGraph mining is an important research area within the domain of data mining. One of the most challenging tasks of graph mining is frequent subgraph mining. This work presents the first FPGA-based implementation, to the best of our knowledge, of the most efficient and well-known algorithm for the Frequent Subgraph Mining (FSM) problem, i.e. gSpan. The proposed system, named High Performance Computing-gSpan (HPC-gSpan), achieves manyfold speedup vs. the official software solution of the gboost library when executed on a high-end CPU for various real-world datasets. Athanasios Stratikopoulos, Grigorios Chrysos 0001, Ioannis Papaefstathiou, Apostolos Dollas |
FPL | 3 |
| 2014 | Evaluation of RPL with a transmission count-efficient and trust-aware routing metricabstractWireless Sensor Networks (WSNs) often need to operate under strict requirements on energy consumption and be capable of self-adapting to the presence of non-trusted nodes which do not fully cooperate in the packet forwarding operation. In such an environment, the mechanism employed for the calculation of routing paths of minimum cost in terms of the number of transmissions executed for the reliable communication between the data source and the data sink is essential to prevent unnecessary nodes' energy depletion and help prolong the network's lifetime. In this paper, we study a routing metric, TXPFI, that captures the expected number of frame transmissions - including retransmissions - needed for the successful delivery of data from the source to the destination in the presence of malicious nodes and lossy links, and validate its applicability to the IETF RPL routing protocol. Through extensive simulations we evaluate the capability of TXPFI to compute routing paths of minimum transmission count and compare it against a number of metrics suitable for transmission count-efficient and trust-aware routing in WSNs. The results show that the use of TXPFI in RPL can lead to significant transmission count savings. Panagiotis Karkazis, Ioannis Papaefstathiou, Lambros Sarakis, Theodore B. Zahariadis, Terpsichori Helen Velivassaki, Dimitrios Bargiotas |
ICC | 2 |
| 2014 | Policy-Based Access Control for Body Sensor Networks
Charalampos Manifavas, Konstantinos Fysarakis, Konstantinos Rantos, Konstantinos Kagiambakis, Ioannis Papaefstathiou |
WISTP | 5 |
| 2013 | Parallelizing bioinformatics and security applications on a low-cost multi-core systemabstractUtilizing multi-cores is now the norm in order to increase performance while also saving energy. The need to break the physical limits of uniprocessing (by branch prediction or RAW dependencies etc.) while being cost and power effective at the same time were the motivation for the scientific and industrial communities to focus on multi-processor architectures. However, the parallelization of existing applications has very frequently proved to be a cumbersome task and in many cases the parallel application is slower than the original serial one. This work demonstrates the parallelization of one high-end bioinformatics application (multiple sequence alignment for amino acids or nucleotide sequences “MAFFT”) as well as a novel security application (fingerprinting recognition “NBIS”) on a highly parallel, yet a very low cost, system. We initially demonstrate the method for parallelizing the applications and then we focus on the end performance. One application is significantly accelerated when the 7 cores of the system are utilized whereas the other cannot get any gain when being ported to more than one cores; we also demonstrate certain optimization techniques. We believe that this paper can act as a guideline for programmers that need to port their serial code to a parallel machine. Teodor Tzanoudakis, Ioannis Papaefstathiou, Charalampos Manifavas |
AICCSA | 2 |
| 2013 | Fast, FPGA-based Rainbow Table creation for attacking encrypted mobile communicationsabstractEncryption algorithms utilized in mobile communication systems have been under attack since their introduction, and many of these attacks have been successful in practical settings. One such example, A5/1 used in GSM, was attacked using “Rainbow Tables”, i.e. pre-computed tables that trade long offline computation and large storage for runtime efficiency when cracking the code. Traditionally, Rainbow Tables were used to reverse password hashes. Their application against A5/1 opened up a new domain of exploitation. In this paper, we present an FPGA-based architecture for the efficient creation of Rainbow Tables for the A5/3 block cipher that is used in 2ndand 3rdgeneration mobile communication systems. The overall goal is to extract the encryption key, provided we have a ciphertext block under a known plaintext attack. The presented architecture exploits the parallelism in the Rainbow Table creation process, and using a Virtext5 LX330T achieves speedups around 9x and 550x for one and 64 compute engines respectively. We show that due to the limited available memory in our experimental setup, our approach achieves high success rates for a key space reduced to 242. We then demonstrate how we can seamlessly extend the proposed architecture to efficiently create much larger Rainbow Tables for the full key-space. Panagiotis Papantonakis, Dionisios N. Pnevmatikatos, Ioannis Papaefstathiou, Charalampos Manifavas |
FPL | 3 |
| 2013 | HC-CART: A parallel system implementation of data mining classification and regression tree (CART) algorithm on a multi-FPGA systemabstractData mining is a new field of computer science with a wide range of applications. Its goal is to extract knowledge from massive datasets in a human-understandable structure, for example, the decision trees. In this article we present an innovative, high-performance, system-level architecture for the Classification And Regression Tree (CART) algorithm, one of the most important and widely used algorithms in the data mining area. Our proposed architecture exploits parallelism at the decision variable level, and was fully implemented and evaluated on a modern high-performance reconfigurable platform, the Convey HC-1 server, that features four FPGAs and a multicore processor. Our FPGA-based implementation was integrated with the widely used “rpart” software library of the R project in order to provide the first fully functional reconfigurable system that can handle real-world large databases. The proposed system, named HC-CART system, achieves a performance speedup of up to two orders of magnitude compared to well-known single-threaded data mining software platforms, such as WEKA and the R platform. It also outperforms similar hardware systems which implement parts of the complete application by an order of magnitude. Finally, we show that the HC-CART system offers higher performance speedup than some other proposed parallel software implementations of decision tree construction algorithms. Grigorios Chrysos 0001, Panagiotis Dagritzikos, Ioannis Papaefstathiou, Apostolos Dollas |
ACM Trans. Archit. Code Optim. | 3 |
| 2013 | Significantly reducing MPI intercommunication latency and power overhead in both embedded and HPC systemsabstractHighly parallel systems are becoming mainstream in a wide range of sectors ranging from their traditional stronghold high-performance computing, to data centers and even embedded systems. However, despite the quantum leaps of improvements in cost and performance of individual components over the last decade (e.g., processor speeds, memory/interconnection bandwidth, etc.), system manufacturers are still struggling to deliver low-latency, highly scalable solutions. One of the main reasons is that the intercommunication latency grows significantly with the number of processor nodes. This article presents a novel way to reduce this intercommunication delay by implementing, in custom hardware, certain communication tasks. In particular, the proposed novel device implements the two most widely used procedures of the most popular communication protocol in parallel systems the Message Passing Interface (MPI). Our novel approach has initially been simulated within a pioneering parallel systems simulation framework and then synthesized directly from a high-level description language (i.e., SystemC) using a state-of-the-art synthesis tool. To the best of our knowledge, this is the first article presenting the complete hardware implementation of such a system. The proposed novel approach triggers a speedup from one to four orders of magnitude when compared with conventional software-based solutions and from one to three orders of magnitude when compared with a sophisticated software-based approach. Moreover, the performance of our system is from one to two orders of magnitude higher than the simulated performance of a similar but, relatively simpler hardware architecture; at the same time the power consumption of our device is about two orders of magnitude lower than that of a low-power CPU when executing the exact same intercommunication tasks. Pavlos M. Mattheakis, Ioannis Papaefstathiou |
ACM Trans. Archit. Code Optim. | 2 |
| 2013 | A novel low-power embedded object recognition system working at multi-frames per secondabstractOne very important challenge in the field of multimedia is the implementation of fast and detailed Object Detection and Recognition systems. In particular, in the current state-of-the-art mobile multimedia systems, it is highly desirable to detect and locate certain objects within a video frame in real time. Although a significant number of Object Detection and Recognition schemes have been developed and implemented, triggering very accurate results, the vast majority of them cannot be applied in state-of-the-art mobile multimedia devices; this is mainly due to the fact that they are highly complex schemes that require a significant amount of processing power, while they are also time consuming and very power hungry. In this article, we present a novel FPGA-based embedded implementation of a very efficient object recognition algorithm called Receptive Field Cooccurrence Histograms Algorithm (RFCH). Our main focus was to increase its performance so as to be able to handle the object recognition task of today's highly sophisticated embedded multimedia systems while keeping its energy consumption at very low levels. Our low-power embedded reconfigurable system is at least 15 times faster than the software implementation on a low-voltage high-end CPU, while consuming at least 60 times less energy. Our novel system is also 88 times more energy efficient than the recently introduced low-power multi-core Intel devices which are optimized for embedded systems. This is, to the best of our knowledge, the first system presented that can execute the complete complex object recognition task at a multi frame per second rate while consuming minimal amounts of energy, making it an ideal candidate for future embedded multimedia systems. Antonis Nikitakis, Savvas Papaioannou, Ioannis Papaefstathiou |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2013 | Evaluating routing metric composition approaches for QoS differentiation in low power and lossy networks
Panagiotis Karkazis, Panagiotis Trakadas, Helen-Catherine Leligou, Lambros Sarakis, Ioannis Papaefstathiou, Theodore B. Zahariadis |
Wirel. Networks | 5 |
| 2012 | An FPGA-based parallel processor for Black-Scholes option pricing using finite differences schemesabstractFinancial engineering is a very active research field as a result of the growth of the derivative markets and the complexity of the mathematical models utilized in pricing the numerous financial products. In this paper, we present an FPGA-based parallel processor optimized for solving the Black-Scholes partial derivative equation utilized in option pricing which employs the two most widely used finite difference schemes: Crank-Nicholson and explicit differences. As our measurements demonstrate, the presented architecture is expandable and the speedup triggered is increased almost linearly with the available silicon resources. Although the processor is optimized for this specific application, it is highly programmable and thus it can significantly accelerate all applications that use finite differences computations. Performance measurements show that our FPGA prototype triggers a 5× speedup when compared with a 2GHz dual-core Intel CPU (Core2Duo). Moreover, for the explicit scheme, our FPGA processor provides an 8× speedup over the same Intel processor. Georgios Chatziparaskevas, Andreas Brokalakis, Ioannis Papaefstathiou |
DATE | 3 |
| 2012 | HEAP: A Highly Efficient Adaptive Multi-processor FrameworkabstractWriting parallel code is difficult, especially when starting from a sequential reference implementation. Our research efforts, as demonstrated in this paper, face this challenge directly by providing an innovative toolset that helps software developers profile and parallelize an existing sequential implementation, by exploiting top-level pipeline-style parallelism. The innovation of our approach is based on the facts that a) we use both automatic and profiling-driven estimates of the available parallelism, b) we refine those estimates using metric-driven verification techniques, and c) we support dynamic recovery of excessively optimistic parallelization. The proposed toolset has been utilized to find an efficient parallel code organization for a number of real-world representative applications, and a version of the toolset is provided in an open-source manner. Luciano Lavagno, Mihai T. Lazarescu, Ioannis Papaefstathiou, Andreas Brokalakis, Johan Walters, Bart Kienhuis, Florian Schäfer 0001 |
DSD | 3 |
| 2012 | FASTCUDA: Open Source FPGA Accelerator & Hardware-Software Codesign Toolset for CUDA KernelsabstractUsing FPGAs as hardware accelerators that communicate with a central CPU is becoming a common practice in the embedded design world but there is no standard methodology and toolset to facilitate this path yet. On the other hand, languages such as CUDA and OpenCL provide standard development environments for Graphical Processing Unit (GPU) programming. FASTCUDA is a platform that provides the necessary software toolset, hardware architecture, and design methodology to efficiently adapt the CUDA approach into a new FPGA design flow. With FASTCUDA, the CUDA kernels of a CUDA-based application are partitioned into two groups with minimal user intervention: those that are compiled and executed in parallel software, and those that are synthesized and implemented in hardware. A modern low power FPGA can provide the processing power (via numerous embedded micro-CPUs) and the logic capacity for both the software and hardware implementations of the CUDA kernels. This paper describes the system requirements and the architectural decisions behind the FASTCUDA approach. Iakovos Mavroidis, Ioannis Mavroidis, Ioannis Papaefstathiou, Luciano Lavagno, Mihai T. Lazarescu, Eduardo de la Torre, Florian Schäfer 0001 |
DSD | 3 |
| 2012 | FASTER: Facilitating Analysis and Synthesis Technologies for Effective ReconfigurationabstractThe FASTER project aims to ease the definition, implementation and use of dynamically changing hardware systems. Our motivation stems from the promise reconfigurable systems hold for achieving better performance and extending product functionality and lifetime via the addition of new features that work at hardware speed. This is a clear advantage over the more straightforward software component adaptivity. However, designing a changing hardware system is both challenging and time consuming. The FASTER project will facilitate the use of reconfigurable technology by providing a complete methodology that enables designers to easily specify, analyse, implement and verify applications on platforms with general-purpose processors and acceleration modules implemented in the latest reconfigurable technology. To better adapt to different application requirements, the tool-chain will support both region-based and micro-reconfiguration and provide a flexible run-time system that will efficiently manage the reconfigurable resources. We will use applications from the embedded, high performance computing, and desktop domains to demonstrate the potential benefits of the FASTER tools on metrics such as performance, power consumption and total ownership cost. Dionisios N. Pnevmatikatos, Tobias Becker, Andreas Brokalakis, Karel Bruneel, Georgi Gaydadjiev, Wayne Luk, Kyprianos Papademetriou, Ioannis Papaefstathiou, Oliver Pell, Christian Pilato, M. Robart, Marco D. Santambrogio, Donatella Sciuto, Dirk Stroobandt, Tim Todman |
DSD | 8 |
| 2012 | Breaking the GSM A5/1 cryptography algorithm with rainbow tables and high-end FPGASabstractA5 is the basic cryptographic algorithm used in GSM cell-phones to ensure that the user communication is protected against illicit acts. The A5/1 version was developed in 1987 and has since been under attack. The most recent attack on A5/1 is the “A51 security project”, led by Karsten Nohl that consists of the creation of rainbow tables that map the internal state of the algorithm with the keystream. Rainbow tables are efficient structures that allow the tradeoff between run-time (computations performed to crack a conversation) and space (memory to hold pre-computed information). In this paper we describe a very effective parallel architecture for the creation of the A5/1 rainbow tables in reconfigurable hardware. Rainbow table creation is the most expensive portion of cracking a particular encrypted information exchange. Our approach achieves almost 3000× speedup over a single processor, and 2.5× speedup compared to GPUs. This performance is achieved with less than 5 Watt power consumption, achieving an energy efficiency in the order of 150x better that the GPU approach. Maria Kalenderi, Dionisios N. Pnevmatikatos, Ioannis Papaefstathiou, Charalampos Manifavas |
FPL | 3 |
| 2012 | Using hardware-based forward error correction to reduce the overall energy consumption of WSNsabstractProviding an energy-efficient communication scheme is highly desirable in Wireless Sensor Networks (WSNs), however this is often constrained by the processing and energy limitations of the wireless nodes. In this paper, we propose the use of a Turbo Code scheme to increase the robustness and energy efficiency of the communication between end nodes and base stations in single-hop topologies. Using several real-world energy consumption measurements from a widely used WSN platform, we demonstrate the operational environment in which the end-user can take full advantage of the proposed scheme. We propose, for the first time, the use of a reconfigurable hardware device that executes the encoding scheme; this approach can reduce the overall energy consumption of a node by more than 40%, when compared with a Turbo Code scheme implemented in software, as well as by more than 70% when compared with the traditional WSN transmission schemes that do not support any kind of Forward Error Correction. Andreas Brokalakis, Ioannis Papaefstathiou |
WCNC | 2 |
| 2012 | Fast and power-efficient hardware implementation of a routing scheme for WSNsabstractOne of the most rapidly expanding areas, nowadays, in networking systems is the Wireless Sensor Network (WSN). Typically WSNs rely on multi-hop routing protocols which must be able to establish communication among nodes and guarantee packet deliveries. In this paper, we present a novel approach for the implementation of the WSN routing protocols, which takes advantage of modern FPGAs in order to provide faster routing decisions while consuming significantly less energy than existing systems. Despite our focus on a particular routing protocol (GPSR), the platform developed has the additional advantage that due to the reconfigurability feature of the FPGA it can efficiently execute different routing protocols based on the requirements of the different WSN applications. As our real world experiments demonstrate, we accelerated the execution of the most widely used WSN routing protocol (GPSR) by at least 31 times when compared to the speed achieved when the exact same protocol is executed on a low power Intel Atom processor. More importantly by utilizing a high-end FPGA the overall energy consumption was reduced by more than 90%. Georgios-Grigorios Mplemenos, Ioannis Papaefstathiou |
WCNC | 2 |
| 2011 | Parallel accelerators for GlimmerHMM bioinformatics algorithmabstractIn the last decades there is an exponential growth in the amount of genomic data that need to be analyzed. A very important problem in biology is the extraction of the biologically functional genomic DNA from the actual genome of the organisms. There have been proposed many computational biology algorithms that solve the gene finding problem which utilize various approaches; GlimmerHMM is considered one of the most efficient such algorithms. This paper presents two different accelerators for the GlimmerHMM algorithm. One of them is implemented on a modern FPGA platform exploiting the parallelism that reconfigurable logic offers and the other one utilizes a GPU (Graphic Processing Unit) taking advantage of a highly multithreaded operational environment. The performance of the implemented systems is compared against the one achieved when the official distribution of the algorithm is executed on a high-end multi-core server; the speedup initiated, for the most compute intensive part, is up to 200× for the FPGA-based system and up to 34× for the GPU-based system. Nafsika Chrysanthou, Grigorios Chrysos 0001, Euripides Sotiriades, Ioannis Papaefstathiou |
DATE | 4 |
| 2011 | Architecture, Design, and Experimental Evaluation of a Lightfield Descriptor Depth Buffer Algorithm on Reconfigurable Logic and on a GPUabstractThe Lightfield descriptor method for 3D computer graphics offers the highest quality object retrieval from a database at the expense of higher storage and computational cost vs. other methods. This paper presents two special purpose architectures, based on FPGAs and GPUs, for the depth buffer extraction algorithm which is used by the Light field Descriptor method. The two architectures were fully designed and implemented in hardware on a Virtex 5 FPGA Device and on a GeForce GPU. The FPGA-based design offers a measured average speedup of 50x vs. software. The corresponding GPU results were by comparison less promising, but still better than software solutions. Results reported in this paper are from actual runs on hardware. Matina Lakka, Grigorios Chrysos 0001, Ioannis Papaefstathiou, Apostolos Dollas |
FCCM | 3 |
| 2011 | Novel and Highly Efficient Reconfigurable Implementation of Data Mining Classification TreeabstractThe available e-data throughout the Web are growing at such a high rate that data mining on the web is considered the biggest challenge of information technology. As a result it is crucial to find new and innovative ways for classifying and mining those huge amounts of data. In this paper we present an implementation of a state-of-the-art data mining algorithm on a modern FPGA. This is one of the first approaches utilizing the resources of an FPGA to accelerate certain very CPU intensive data-mining/data classification schemes and our real-world results from actual runs on hardware demonstrate that it is a highly promising one. In particular, our FPGA-based system achieves, depending on the data classified, a speedup from 4x and up to 50x (on average 25x) when compared with a state-of-the art multi-core CPU, including I/O overhead. Grigorios Chrysos 0001, Panagiotis Dagritzikos, Ioannis Papaefstathiou, Apostolos Dollas |
FPL | 3 |
| 2011 | FPGA power consumption measurements and estimations under different implementation parametersabstractThis paper investigates the effects of different design tool (Xilinx ISE) optimisation schemes on FPGA power consumption. Specifically, on-the-bench measurements are presented for eight highly popular security algorithms, which have been tested under a number of different synthesis and implementation optimisation scenarios. The algorithms under investigation are the BasicRSA, BasicDES, Camellia (with two distinct variations), TripleDES, AES, DES, and MD5. Finally, the efficiency of the design tool in generating accurate predictions on the power consumption of a specific design is also addressed. Results show that power consumption figures may vary from a 306mW reduction (compared to nominal design effort) to a 33mW increase and average improvement on the power consumption measured values ranges between 9.21% and -0.94% when different optimisation schemes are utilised. The Xilinx XPower Analyzer is also scrutinised; It provides estimates that are well above what is actually measured and the estimation error ranges between 17.5% to more than 200%. In the worst case, the XPower estimate is 307mW greater than the least power consumption value measured on-the-bench whereas it remains 258mW higher than the average value measured for the same algorithm and under all different optimisation scenarios. The least deviation in results is measured between 23mW and 31mW, however, the XPower estimate remains greater than that measured on-the-bench. Dimitrios Meidanis, Konstantinos Georgopoulos, Ioannis Papaefstathiou |
FPT | 3 |
| 2011 | RESENSE: An Innovative, Reconfigurable, Powerful and Energy Efficient WSN NodeabstractWireless Sensor Networks (WSNs) have recently enjoyed a tremendous rise in popularity. The current WSN node offerings, however, need both increased processing power and lower energy consumption in order to enable the full potential of such networks. To address these requirements, we explore the benefits of an innovative platform which combines a standard wireless node with very low cost reconfigurable hardware. In order to evaluate the efficiency of this pioneering approach three different networking and security protocols have been implemented on the present system: a) Turbo coding, b) Blowfish encryption and c) XMesh routing. Our real-world experiments demonstrate that our prototype system provides comparable performance to the existing microcontroller-based schemes (while in its productized version it could potentially be much faster) whereas, and more importantly, its overall energy consumption is from 70% to 93% lower than that triggered when a very widely used commercial WSN node is executing the exact same processing tasks. Andreas Brokalakis, Georgios-Grigorios Mplemenos, Konstantinos Papadopoulos 0003, Ioannis Papaefstathiou |
ICC | 4 |
| 2010 | Fast and Efficient FPGA-Based Feature Detection Employing the SURF AlgorithmabstractFeature detectors are schemes that locate and describe points or regions of `interest' in an image. Today there are numerous machine vision applications needing efficient feature detectors that can work on Real-time; moreover, since this detection is one of the most time consuming tasks in several vision devices, the speed of the feature detection schemes severally affects the effectiveness of the complete systems. As a result, feature detectors are increasingly being implemented in state-of-the-art FPGAs. This paper describes an FPGA-based implementation of the SURF (Speeded-Up Robust Features) detector introduced by Bay, Ess, Tuytelaars and Van Gool; this algorithm is considered to be the most efficient feature detector algorithm available. Moreover, this is, to the best of our knowledge, the first implementation of this scheme in an FPGA. Our innovative system can support processing of standard video (640 x 480 pixels) at up to 56 frames per second while it outperforms a state-of-the-art dual-core Intel CPU by at least 8 times. Moreover, the proposed system, which is clocked at 200 MHz and consumes less than 20W, supports constantly a frame rate only 20% lower than the peak rate of a high-end GPU executing the same basic algorithm; the specified GPU consists of 128 floating point CPUs, clocked at 1.35 GHz and consumes more than 200W. Dimitris Bouris, Antonis Nikitakis, Ioannis Papaefstathiou |
FCCM | 3 |
| 2010 | Implementing Rainbow Tables in High-End FPGAs for Super-Fast Password CrackingabstractOne of the most efficient methods for cracking passwords, which are hashed based on different cryptographic algorithms, is the one based on “Rainbow Tables”. Those lookup tables offer an almost optimal time-memory tradeoff in the process of recovering the plaintext password from a password hash, generated by a cryptographic hash function. In this paper, the first known such generic system is demonstrated. It is implemented in a state-of-the-art reconfigurable device that cracks passwords, which are encrypted with a number of different cryptographic algorithms. The proposed FPGA-based system is up to 1000 times faster than the corresponding software approach. This is achieved by using a highly parallel architecture employing a fine-grained pipeline. Kostas Theocharoulis, Ioannis Papaefstathiou, Charalampos Manifavas |
FPL | 2 |
| 2010 | Using Reconfigurable Hardware Devices in WSNs for Reducing the Energy Consumption of Routing and Security TasksabstractAs Wireless Sensor Networks (WSNs) expand their reach, their applications require lower energy consumption than those provided by current offerings. In this paper, a novel platform is introduced, which employs a reconfigurable device (CPLD) in order to enhance the processing power of typical sensor nodes and, more importantly, reduce the overall energy consumption in common tasks such as routing and security. Our real-world measurements, demonstrate that the proposed system can reduce the energy consumption of the Cost Estimation algorithm of the widely used XMesh routing protocol by 71.5%. Higher energy conservations (more than 90%) can be achieved at Blowfish Encryption, when this is implemented on our new platform. Georgios-Grigorios Mplemenos, Konstantinos Papadopoulos 0003, Ioannis Papaefstathiou |
GLOBECOM | 3 |
| 2010 | Titan-R: A Multigigabit Reconfigurable Combined Compression/Decompression UnitabstractData compression techniques can alleviate bandwidth problems in even multigigabit networks and are especially useful when combined with encryption. This article demonstrates a reconfigurable hardware compressor/decompressor core, the Titan-R, which can compress/decompress data streams at 8.5 Gb/sec, making it the fastest reconfigurable such device ever proposed; the presented full-duplex implementation allows for fully symmetric compression and decompression rates at 8.5 Gbps each. Its compression algorithm is a variation of the most widely used and efficient such scheme, the Lempel-Ziv (LZ) algorithm that uses part of the previous input stream as the dictionary. In order to support this high network throughput, the Titan-R utilizes a very fine-grained pipeline and takes advantage of the high bandwidth provided by the distributed on-chip RAMs of state-of-the-art FPGAs. Konstantinos Papadopoulos 0003, Ioannis Papaefstathiou |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2009 | On the Power Consumption of Security Algorithms Employed in Wireless NetworksabstractSupporting high levels of security in wireless networks is a challenging issue because of the specific problems this environment poses; the provided security by small mobile systems, such as PDAs and mobile phones, is often restricted by their limited battery power and their limited processing power. Driven by these restrictions, the designer will have to decide whether to implement the wireless network security schemes in software or to add special purpose hardware units to the system, executing those CPU intensive tasks. This paper demonstrates and compares the Hardware and Software implementations of a number of widely used security applications employed in wireless networks. We measured the total energy consumption for each security algorithm when implemented in reconfigurable hardware devices and we compared it with the total energy consumption of the equivalent software applications. We demonstrate, that the hardware implementations on a state-of-the-art FPGA are significantly faster while they consume three orders of magnitude less power when compared with the software implementations executed on a state-of-the-art hard-core CPU which is embedded in the same FPGA device. Dimitrios Meintanis, Ioannis Papaefstathiou |
CCNC | 2 |
| 2009 | Design and implementation of a database filter for BLAST accelerationabstractBLAST is a very popular computational biology algorithm. Since it is computationally expensive it is a natural target for acceleration research, and many reconfigurable architectures have been proposed offering significant improvements. In this paper we approach the same problem with a different approach: we propose a BLAST algorithm preprocessor that efficiently identifies the portions of the database that must be processed by the full algorithm in order to find the complete set of desired results. We show that this preprocessing is feasible and quick, and requires minimal FPGA resources, while achieving a significant reduction in the size of the database that needs to be processed by BLAST. We also determine the parameters under which prefiltering is guaranteed to identify the same set of solutions as the original NCBI software. We model our preprocessor in VHDL and implement it in reconfigurable architecture. To evaluate the performance, we use a large set of datasets and compare against the original (NCBI) software. Prefiltering is able to determine that between 80 and 99.9% of the database will not produce matches and can be safely ignored. Processing only the remaining portions using software such as NCBI-BLAST improves the system performance (reduces execution time) by 3 to 15 times. Since our prefiltering technique is generic, it can be combined with any other software or reconfigurable acceleration technique. Panagiotis Afratis, Constantinos Galanakis, Euripides Sotiriades, Georgios-Grigorios Mplemenos, Grigorios Chrysos 0001, Ioannis Papaefstathiou, Dionisios N. Pnevmatikatos |
DATE | 6 |
| 2009 | High-End Reconfigurable Systems for Fast Windows' Password CrackingabstractOne of the most efficient methods for cracking passwords is the one based on ldquorainbow tablesrdquo; those lookup tables are offering an almost optimal time-memory tradeoff in the process of recovering the plaintext password from a password hash generated by a cryptographic hash function. In this paper, we demonstrate the first known system, implemented in a state-of-the-art reconfigurable device that cracks passwords up to 1000 times faster than the software approach. This is achieved by using a highly parallel architecture employing a fine-grained pipeline. Kostas Theocharoulis, Charalampos Manifavas, Ioannis Papaefstathiou |
FCCM | 3 |
| 2009 | Implementation of a genetic algorithm on a virtex-ii pro FPGAabstractThis paper presents the implementation of a Genetic Algorithm on a XUPV2P platform with a Virtex-II Pro FPGA. A Genetic Algorithm (GA) is a search technique finding exact or approximate solutions to optimization and search problems. It is a computer simulation approach in which a population of abstract representations of candidate solutions to an optimization problem evolves toward better solutions. The aim is the optimization of a given function, called fitness function, which is evaluated upon the initial population as well as upon the solutions after successive generations. The motivation for implementing GAs in hardware stems from the fact that they are very CPU intensive while they are also intrinsically parallel algorithms and the basic operations of a GA can execute in a pipelining fashion. Our architecture incorporates a Power PC, and built-in hardcore resources like multiplier blocks and BRAMs in order to create an efficient hardware-based genetic algorithm. We have fine tuned the architecture so as to be more parallel, and added complex fitness functions. The design executes on 100 MHz and although complex in logic, it has low silicon requirements as it utilizes 16% slices, 7% BRAMs and 11% multiplier blocks, plus the Power PC of the XC2VP30 FPGA. The result is a functional prototype on which experiments for a range of different genetic parameters can be conducted. We explore the GA's behavior with real-world experiments for different fitness functions and different number of generations. Our design is the first single-chip fully embedded approach that optimizes six fitness functions, which is more than any other proposed solution, while its current implementation supports populations of up to 32 members. The experiments show that our system outperforms in terms of execution time the existing and proposed hardware systems from 23% up to 5895%. Michalis Vavouras, Kyprianos Papademetriou, Ioannis Papaefstathiou |
FPGA | 3 |
| 2009 | A FPGA based coprocessor for gene finding using Interpolated Markov Model (IMM)abstractAn important biology problem is the decoding of the DNA and the extraction of useful genetic information. There are many bioinformatics algorithms that try to solve the gene finding problem and one of the most efficient is the Glimmer algorithm. In this paper, we present a hardware architecture that implements the Glimmer algorithm. The architecture was developed specifically for the capabilities of present-day FPGAs. In addition, this paper presents an efficient hardware method to construct a huge but very sparse lookup table by taking advantage the tree-like structure of memories. Grigorios Chrysos 0001, Euripides Sotiriades, Ioannis Papaefstathiou, Apostolos Dollas |
FPL | 3 |
| 2009 | A self-reconfiguring architecture supporting multiple objective functions in genetic algorithmsabstractGenetic algorithms (GA) are search algorithms based on the mechanism of natural selection and genetics. FPGAs have been widely used to implement hardware-based genetic algorithms (HGA) and have provided speedups of up to three orders of magnitude as compared to their software counterparts. In this paper, we propose a parameterized partially reconfigurable HGA architecture (PPR-HGA). The novelty of this architecture is that it allows for the objective function to be updated through partial reconfiguration, and supports various genetic parameters. Charalampos Effraimidis, Kyprianos Papademetriou, Apostolos Dollas, Ioannis Papaefstathiou |
FPL | 4 |
| 2009 | Design space exploration of reconfigurable systems for calculating flying object's optimal noise reduction pathsabstractDespite improved aerodynamic designs that decrease sound emission, the noise produced by flying objects is a problem because it propagates in large distances in the atmosphere. However, it is possible to describe sound propagation effectively with a parabolic differential equation and determine a path that minimizes noise emissions taking into consideration atmospheric and geographic data. This approach calculates noise propagation progressively in the propagation direction and gives accurate results even for large distances. This paper presents a reconfigurable system that solves the tridiagonal problem that results from the Crank-Nicolson function of the 2ndorder parabolic equation. Generally tridiagonal algorithms do not allow parallelism in every level, and complicate parallel and/or reconfigurable hardware implementations. We show that reconfigurable hardware technology allows the fast and accurate implementation of such systems. We explore and present several architecture alternatives that pose different tradeoffs. We consider and evaluate different implementations in order to achieve the best possible parallelism and speed with reasonable cost. We find that a large Virtex-5 device is between 9 and 76 times faster than a current desktop PC, depending on the architecture used. Dimitrios Kontos, Ioannis Papaefstathiou, Dionisios N. Pnevmatikatos |
FPL | 2 |
| 2009 | A fast parallel matrix multiplication reconfigurable unit utilized in face recognitions systemsabstractIn this paper we present a reconfigurable device which significantly improves the execution time of the most computational intensive functions of three of the most widely used face recognition algorithms; those tasks multiply very large dense matrices. The presented architecture utilizes numerous digital signal processing units (DSPs) organized in a parallel manner within a state-of-the-art FPGA device. In order to accelerate those functions we have implemented a ldquoblockedrdquo matrix multiplication algorithm which multiplies certain sub-matrices of fixed-point 32-bit numbers; the size of the sub-matrices has been selected so as to fully exploit the resources of the underlying reconfigurable device. Our system is up to 550 times faster than a conventional general purpose processor when implementing the most CPU intensive parts of a number of very widely used face identification schemes, whereas it is more than 40 times faster than the similar schemes implemented in reconfigurable devices. Moreover, our system is general enough so as to be efficiently utilized in any application incorporating fixed-point matrix multiplications. Joannis Sotiropoulos, Ioannis Papaefstathiou |
FPL | 2 |
| 2008 | MPLEM: An 80-processor FPGA Based Multiprocessor SystemabstractMultiprocessor embedded systems (MESes) are a very promising approach for high performance yet relatively low-cost computing. At the same time modern FPGAs provide the silicon capacity to build multiprocessor systems containing 10-100 processors, complex memory systems, heterogeneous interconnection schemes and custom engines executing the performance-critical operations. In this work we present a MES implemented in a state-of-the-art FPGA consisting of up to eighty 32-bit processors. The efficiency of our approach is demonstrated by the fact that our system can execute the BLAST CPU-intensive application, which is the prevalent tool used by molecular biologists for DNA sequence matching and database search, many times faster than a simple PC. Georgios-Grigorios Mplemenos, Ioannis Papaefstathiou |
FCCM | 2 |
| 2008 | A Memory-Efficient FPGA-based Classification EngineabstractPacket classification is one of the most important enabling technologies for next generation network services. Even though many multi-dimensional classification algorithms have been proposed, most of them are precluded from commercial equipments due to their high memory requirements. In this paper, we present an efficient packet classification scheme, implemented in reconfigurable hardware, called Dual Stage Bloom Filter Classification Engine (2sBFCE). 2sBFC comprises of an innovative 5-field search scheme that decomposes multi-field classification rules into internal single-field rules which are combined using multi-level Bloom filters. The design of 2sBFCE is optimized for the common case based on analysis of real world classification databases. The FPGA implementation of the proposed scheme handles 4K rules, with very small memory requirements, while supporting network streams at a rate of 2Gbps in the worst case, and more than 6Gbps in the average case. Antonis Nikitakis, Ioannis Papaefstathiou |
FCCM | 2 |
| 2008 | Titan-R: A Reconfigurable Hardware Implementation of a High-Speed CompressorabstractData compression techniques can alleviate low bandwidth problems in multigigabit networks and are especially useful when combined with encryption. This paper presents a reconfigurable hardware compressor core, the Titan-R, which can compress data streams at 8.5 Gb/sec making it the fastest reconfigurable such device ever proposed. Its compression algorithm is a variation of the most widely used and efficient such scheme the Lempel-Ziv (LZ) algorithm that uses part of the previous input stream as the dictionary. In order to support this high network throughput the Titan-Rutilizes a very fine-grained pipeline and takes advantages of the high-bandwidth provided by the distributed on-chip RAMs of the state-of-the-art FPGAs. Konstantinos Papadopoulos 0003, Ioannis Papaefstathiou |
FCCM | 2 |
| 2008 | Accelerating hardware simulation: Testbench code emulationabstractTodaypsilas verification challenges require high-performance simulation solutions, such as hardware simulation accelerators and emulators, that have been in use in hardware and electronic system design centers for approximately the last decade. In particular, in order to accelerate functional simulation, hardware emulation is used so as to offload calculation-intensive tasks from the software simulator. However, the communication overhead between the software simulator and the hardware emulator is becoming a new critical bottleneck. In our work we introduce a novel way of repartitioning the simulation between software and hardware in order to minimize this communication bottleneck. Using the techniques described in this paper we are able to offload a big part of the work that is traditionally done by the software simulator, onto the hardware emulator. Our experiments, using real-world designs, demonstrate that the proposed method reduces significantly the communication overhead and outperforms the conventional hardware emulation systems by a factor of more than 7. Finally, we provide a way of observing and modifying the internal state of the hardware emulator while the test is running. Iakovos Mavroidis, Ioannis Papaefstathiou |
FPT | 2 |
| 2008 | A Multi Gigabit FPGA-Based 5-tuple Classification SystemabstractPacket classification is one of the most important enabling technologies for next generation network services. Even though many multi-dimensional classification algorithms have been proposed, most of them are precluded from commercial equipments due to their high memory requirements. In this paper, we present an efficient packet classification scheme, called dual stage bloom filter classification engine (2sBFCE). 2sBFC comprises of an innovative 5- field search scheme that decomposes multi-field classification rules into internal single-field rules which are combined using multi-level Bloom filters. The design of 2sBFCE is optimized for the common case based on analysis of real world classification databases. The hardware implementation of this scheme handles 4 K rules while supporting network streams at a rate of 2 Gbps even in the worst case, and more than 6 Gbps in the average case when implemented in an off-the-shelf FPGA. Antonis Nikitakis, Ioannis Papaefstathiou |
ICC | 2 |
| 2007 | Efficient testbench code synthesis for a hardware emulator systemabstractThe rising complexity of modern embedded systems is causing a significant increase in the verification effort required by hardware designers and software developers, leading to the "design verification crisis ", as it is known among engineers. Today's verification challenges require powerful testbenches and high-performance simulation solutions such as Hardware Simulation Accelerators and Hardware Emulators that have been in use in hardware and electronic system design centers for approximately the last decade. In particular, in order to accelerate functional simulation, hardware emulation is used so as to offload calculation-intensive tasks from the software simulator. However, the communication overhead between the software simulator and hardware emulator is becoming a new critical bottleneck. We tackle this problem by partitioning the code running on the software simulator into two sections: the testbench HDL (hardware description language) code that communicates directly with the design under test (DUT) and the rest C-like testbench code. The former section is transformed into synthesizable code while the latter runs in a general purpose CPU. Our experiments demonstrate that the proposed method reduces the communication overhead by a factor of about 5 compared to a conventional hardware emulated simulation Ioannis Mavroidis, Ioannis Papaefstathiou |
DATE | 2 |
| 2007 | A Fast FPGA-Based 2-Opt Solver for Small-Scale Euclidean Traveling Salesman ProblemabstractIn this paper we discuss and analyze the FPGA-based implementation of an algorithm for the traveling salesman problem (TSP), and in particular of 2-Opt, one of the most famous local optimization algorithms, for Euclidean TSP instances up to a few hundred cities. We introduce the notion of "symmetrical 2-Opt moves" which allows us to uncover fine-grain parallelism when executing the specified algorithm. We propose a novel architecture that exploits this parallelism, and demonstrate its implementation in reconfigurable hardware. We evaluate our proposed architecture and its implementation on a state-of-the-art FPGA using a subset of the TSPLIB benchmark, and find that our approach exhibits better quality of final results and an average speedup of 600% when compared with the state-of-the-art software implementation. Our approach produces, to the best of our knowledge, the fastest to date TSP 2-Opt solver for small-scale Euclidean TSP instances. Ioannis Mavroidis, Ioannis Papaefstathiou, Dionisios N. Pnevmatikatos |
FCCM | 2 |
| 2007 | An efficient FPGA-based implementation of Pollard's (ρ - 1) factorization algorithmabstractDue to the widespread use of public key cryptosystems whose security depends on the presumed difficulty of the factorization problem, the algorithms for finding the prime factors of large composite numbers are becoming extremely important. In recent years the limits of the best integer factorization algorithms have been extended greatly, due in part to Moore's law and in part to algorithmic improvements. Furthermore, new silicon devices, such as FPGAs, give us the advantage of custom hardware architectures for minimizing execution time for such difficult computations. This paper demonstrates a very efficient FPGA-based design executing Pollard's (ρ - 1) factorization algorithm. The proposed device offers a speedup from 20 to 231 when compared to the software implementation of the same algorithm in a state-of-the-art CPU. Dimitrios Meintanis, Ioannis Papaefstathiou |
FPT | 2 |
| 2007 | A buffered crossbar-based chip interconnection framework supporting quality of serviceabstractAs Systems-on-a-Chip (SoCs) become larger, the problem of interconnecting the various subsystems becomes more complicated. In this framework, certain alternatives to the standard buses, based on Network Technologies, have emerged as innovative approaches for SoC's interconnect. One of the main advantages of such an alternative, is that it can offer certain Quality of Service (QoS) over the internal cross-connects while at the same time it supports higher transfer rates than the existing on-chip buses. This paper presents the first chip interconnection architecture, which is based on a buffered crossbar switch. The main advantage of the proposed system is that it efficiently supports different priority levels; it also provides several Gigabits per Second of aggregate bandwidth, while it introduces very low latency. Moreover, the hardware complexity of this highly scalable scheme is minimal. All those facts make this framework ideal for SoCs that contain IP cores with diverse speed/throughput requirements. Ioannis Papaefstathiou, Nikolaos Chrysos |
ACM Great Lakes Symposium on VLSI | 1 |
| 2007 | Memory-Efficient 5D Packet Classification At 40 GbpsabstractPacket classification is one of the most important enabling technologies for next generation network services. Even though many multi-dimensional classification algorithms have been proposed, most of them are precluded from commercial equipments due to their high memory requirements. In this paper, we present an efficient packet classification scheme, called Bloom Based Packet Classification (B2PC). B2PC comprises of an innovative 5-field search algorithm that decomposes multifield classification rules into internal single field rules which are combined using multi-level Bloom filters. The design of B2PC is optimized for the common case based on analysis of real world classification databases. The hardware implementation of this scheme handles 4K rules by involving only 530KB of memory for its data structures, while it supports network streams at a rate of 15Gbps even in the worst case, and more than 40Gbps in the average case. This system covers 1.3 mm in a 0.18mum CMOS technology. We show that given a certain memory budget and silicon cost, the B2PC is the most efficient hardware-based approach to the classification problem. Ioannis Papaefstathiou, Vassilis Papaefstathiou |
INFOCOM | 1 |
| 2007 | An Embedded Networking SoC for purely Ethernet MANs/WANsabstractEthernet technology is lo longer used only in Local Area Networks (LANs); it is continuously gaining momentum in the Metropolitan Area Networks (MAN s) and Wide Area Networks (WANs). This paper presents a multi-service access concentrator core that has been designed specifically for multi-service, purely Ethernet, access nodes. In particular the presented system is optimised for Ethernet traffic aggregation over MPLS-based optical backbone networks. Moreover, we also demonstrate the bottlenecks that have been identified when such networking applications are executed in a general-purpose network processing device, and the techniques we used in order to bypass them. Our experiments results, executed in a state-of-the-art FPGA-based platform, strongly support that the combination of general purpose processing units with powerful specialized hardware modules, is the most cost-effective approach for designing systems that (i) can support today's and future network speeds and applications, in purely Ethernet networks, (ii) provide the end-user with the required programmability. Theofanis Orphanoudakis, Ioannis Mavroidis, Aristides Nikologiannis, Ioannis Papaefstathiou |
ISCC | 5 |
| 2006 | An innovative low-cost Classification Scheme for combined multi-Gigabit IP and Ethernet NetworksabstractIP is certainly the most popular wide area network protocol while Ethernet is the most common Layer-2 network protocol, and it is currently being deployed beyond the tight borders of LANs. In order to accommodate the needs of MANs and WANs, several QoS mechanisms employed either at the IP layer or the MAC sublayer have been proposed. These QoS mechanisms require identification of network flows and the classification of network packets according to certain packet header fields. In this paper, we propose a classification engine employed either at the MAC sublayer or the IP layer, which is the successor of a scheme already successfuly implemented which is only employed at the MAC sublayer. This new scheme uses an innovative hashing scheme combined with an efficient trie-based structure. By using such techniques, the extremely high speed decisions - at a rate of more than 100Gb/sec- are supported, while the memory needs of the proposed engine are significantly lower compared to those of the similar schemes currently used. This engine has been implemented in hardware utilizing less than 0.2mm2 in a state of the art CMOS technology. As a result the proposed scheme is a very promising candidate for both the next-generation IP classification engines(probably incorporated within the high-end network processors) as well as for the Ethernet equipments that need to support classification at multi-Gigabit per second network speeds, while also employing the minimum amount of memory. Ioannis Papaefstathiou, Vassilis Papaefstathiou |
ICC | 1 |
| 2005 | A Low-Power Processor Architecture Optimized forWireless DevicesabstractThe advantages of power-aware processors are well known. This paper presents an innovative processor architecture optimized for wireless environments. The presented architecture incorporates a certain power-aware microarchitectural technique, called pipeline depth adaptation and it is tailored to self-timed processors. With this technique, a processor is able to alter its pipeline depth, while in operation, trading speed and energy use. The pipeline depth is changed by making selected pipeline registers transparent. A shallow pipeline has lower energy consumption for two reasons: the capacitance driven by the load signal of the 'collapsed' pipeline registers is not switched and the reduction in branch latency and data-dependent stalls reduce the cycles per instruction (CPI) of the processor. An analysis of the advantages of using pipeline depth adaptation in an asynchronous processor is given, supported by simulation results based on a real asynchronous processor and on applications that are frequently executed on a wireless environment. Finally, a method of dynamically adapting the pipeline depth is described and evaluated which only reduces the pipeline depth when a branch instruction is expected. The presented architecture has a relatively lower power consumption than a conventional similar architecture, therefore it can be useful in wireless environments. Aristides Efthymiou, Jim D. Garside, Ioannis Papaefstathiou |
ASAP | 3 |
| 2005 | Queue Management in Network ProcessorsabstractOne of the main bottlenecks when designing a network processing system is very often its memory subsystem. This is mainly due to the state-of-the-art network links operating at very high speeds and to the fact that in order to support advanced quality of service (QoS), a large number of independent queues is desirable. In this paper we analyze the performance bottlenecks of various data memory managers integrated in typical network processing units (NPU). We expose the performance limitations of software implementations utilizing the RISC processing cores typically found in most NPU architectures and we identify the requirements for hardware assisted memory management in order to achieve wire-speed operation at gigabit per second rates. Furthermore, we describe the architecture and performance of a hardware memory manager that fulfills those requirements. This memory manager, although it is implemented in a reconfigurable technology, can provide up to 6.2 Gbit/s of aggregate throughput, while handling 32 K independent queues. Ioannis Papaefstathiou, Theofanis Orphanoudakis, Christoforos Kachris, Ioannis Mavroidis, Aristides Nikologiannis |
DATE | 1 |
| 2005 | A Memory Efficient, 100 Gb/sec MAC Classification EngineabstractIn this paper, we propose a classification engine employed at Ethernet's MAC Layer which uses an innovative hashing scheme and internal replacement of MAC vendor IDs; the hash based classification engine (HBCE) compacts the MAC address tables and supports extremely high speed decisions. Its very low memory requirements make it a very promising candidate for Ethernet equipments that would need to support classification at data link layer and at multi-Gigabit per second network speeds. Moreover HBCE can also so be used very efficiently in lower-bandwidth wireless environments Vassilis Papaefstathiou, Ioannis Papaefstathiou |
LCN | 2 |
| 2004 | Software Processing Performance in Network ProcessorsabstractTo meet the demand for higher performance, flexibility, and economy in today's state-of-the-art networks, an alternative to the ASICs that traditionally were used to implement packet-processing functions in hardware, called network processors (NPs), has emerged. In this paper, we briefly outline the architecture of such an innovative network processor aiming at the acceleration of protocol processing in high-speed network interfaces, and we use this architecture as a case study for our measurements. We focus on the performance of the general purpose processors used for executing high level protocol processing, since this part proves to be the bottleneck of the design. The performance is analyzed by executing a set of widely used, real applications and by applying network traffic according to certain stochastic criteria. The performance of the RISC used is compared with that of other well-known CPU architectures so as to verify that our results are applicable to the general network processors era. As our results demonstrate, the bottleneck of the majority of the network processors is the general-purpose processing units used, since today's network protocols need a great amount of high-level processing. On the other hand the specific purpose processors or co-processors, optimized for certain part of the network packet processing, involved in such systems, can provide the power needed, even at today's ultra high network speeds. Ioannis Papaefstathiou, Nikolaos A. Zervos |
DATE | 1 |
| 2004 | Variable packet size buffered crossbar (CICQ) switchesabstractOne of the most widely used architectures for packet switches is the crossbar. A special version of it is the buffered crossbar, where small buffers are associated with the crosspoints; this simplifies scheduling and improves its efficiency and QoS capabilities to the point where the switch needs no internal speedup. Furthermore, by supporting variable length packets throughout a buffered crossbar: (a) there is no need for segmentation and reassembly (SAR) circuits; (b) no speedup is necessary to support SAR; and (c) synchronization between the input and output clock domains is simplified. In turn, the lack of SAR and speedup mean that no output queues are needed, either. In this paper we present an architecture, a chip layout and cost analysis, and a performance evaluation of such a 300 Gbps buffered crossbar operating on variable-size packets. The proposed organization is simple yet powerful, it can be implemented using modern technology, and, as the performance results demonstrate, it clearly outperforms unbuffered crossbars. Manolis Katevenis, Giorgos Passas, Dimitrios Simos, Ioannis Papaefstathiou, Nikolaos Chrysos |
ICC | 4 |
| 2003 | GFS: An Efficient Implementation of Fair Scheduling for Mult-Gigabit Packet NetworksabstractIn order to address the challenge of providing quality of service guarantees in today's network processing systems, the use of efficient scheduling algorithms is required. The efficiency of a scheduler is determined by several factors including its fairness, capability to operate at high speeds, as well as the resources required for its implementation. We present an architecture to support fair scheduling (gigabit FS) considering variable length packets (i. e. for packet forwarding/switching networks) over gigabit links. This high speed scheduler is designed to manage 32 K flows based on an algorithm that yields efficient implementation in hardware by avoiding the complexity of computing the system virtual time function that many packet fair queueing (PFQ) algorithms have proposed. Further, we demonstrate the critical factors in designing an effective scheduling engine at gigabit rates and we present several enhancements together with their associated cost. Theofanis Orphanoudakis, Ioannis Papaefstathiou |
ASAP | 3 |
| 2003 | A fully-programmable memory management system optimizing queue handling at multi-gigabit ratesabstractTwo of the main bottlenecks when designing a network embedded system are very often the memory bandwidth and its capacity. This is mainly due to the extremely high speed of the state-of-the-art network links and to the fact that in order to support advanced quality of service (QoS), per-flow queueing is desirable. In this paper we describe the architecture of a memory manager that can provide up to 10Gbs of aggregate throughput while handling 512K queues. The presented system supports a complete instruction set and thus we believe it can be used as a hardware component in any suitable embedded system, particularly network SoCs that implement per flow queuing. When designing this scheme several optimisation techniques have been evaluated and the most cost and performance effective ones used. These techniques minimize both the memory bandwidth and the memory capacity needed, which is considered a main advantage of the proposed scheme. The proposed architecture uses a simple DRAM for data storage and a typical SRAM for keeping data structures-pointers, therefore minimising the system's cost. The device has been fabricated within a novel programmable network processor designed for efficient protocol processing in high speed networking applications. It consists of 155K Gates and occupies 5.23 mm in UMC 0.18υ CMOS. Ioannis Papaefstathiou, Aristides Nikologiannis, Nikolaos A. Zervos |
DAC | 2 |
| 1999 | Compressing ATM Streams On-LineabstractSummary form only given. Asynchronous transfer mode (ATM) is one of the state-of-the art network protocols nowadays. A main characteristic of an ATM network is that its switches should be fairly simple and inexpensive. As a result, a significant part of the network cost is in the cost of the links. So by increasing the network traffic we can send over a given link, we will certainly increase the effectiveness of the whole network. A way of increasing this traffic is to send compressed ATM cells. This idea, although it seems very simple, is a new one and we have proved that it is very important. Our compression scheme is particularly useful in LAN since it will increase the bandwidth of a typical ATM LAN by at least 250%. This increase comes at a minimum additional cost (the cost of a pair of our inexpensive hardware chips per link) and with no bad side effects. This scheme is also useful in WAN. The effective bandwidth of such a network can be increased by about 50% at a minimum additional cost and without any side effects, again. This increase is particularly important in the case of WAN since the cost of the link bandwidth in such a network dominates by far the cost of the whole network. But the kind of networks that would be perfect for applying this compression scheme to, will be the LAN interconnection networks over WAN. In these networks the increase in the effective bandwidth of the expensive WAN links will be equal to that of a typical LAN (e.g., at least 250%), since the traffic sent over these WAN links will be mainly LAN traffic. Ioannis Papaefstathiou |
Data Compression Conference | 1 |
| 1999 | Accelerating ATM: on-line compression of ATM streamsabstractSince ATM switches intended to be simple and inexpensive, a significant part of the network cost is in the cost of the links. A way of increasing the traffic we can send over these expensive links is to transmit compressed ATM cells. This idea, although it seems very simple, is a new one for ATM and in this paper we show that it can significantly increase the bandwidth of a typical ATM network. We also show that the buffer size needed for accommodating the peaks introduced by the compression is not prohibitively large, the delay our compression scheme introduces is very low, and that this scheme can be very effective when used with IP traffic. Ioannis Papaefstathiou |
IPCCC | 1 |