EDBT 2026 Demo / reviewers in the wild / expert
Apostolos P. Fournaris
dblp:83/4622
· DBLP profile ↗
32ranked-venue papers
16as first author
11since 2021 · last 2026
0000-0002-4758-2349ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 12 first-author · 7 since 2021Security and privacy · 5 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MEDIATE: Multi-Faceted Implementation of a Mixed Software/Hardware-Based Zero Trust Framework for the Computing Continuum
Apostolos P. Fournaris, Evangelos Haleplidis, Shahin Abdoul-Soukour, Chih-Kai Huang 0001, Niemat Khoder, Georgios Bouloukakis, Andreas Brokalakis, Konstantinos Georgopoulos, Sotiris Ioannidis |
MDM | 1 |
| 2024 | CCSW 2024 - Cloud Computing Security WorkshopabstractThe CCSW workshop aims to bring together researchers and practitioners to explore all aspects of security in cloud-centric and outsourced computing. In CCSW 2024, based on the papers that have been accepted, the workshop is focused on applied cryptographic schemes and protocols for the cloud, cloud-based leakage related attacks and their countermeasures, trusted computing technology in clouds, binary analysis of software for cloud protection, network security mechanisms using AI anomaly detection as well as security for emerging cloud programming models (like extended Berkeley Packet Filter, eBPF, programs). Throughout the years, the workshop particularly encourages novel paradigms and controversial ideas not covered by traditional cloud security research, serving as a fertile ground for creative debate and interaction in security-sensitive areas of computing impacted by cloud technologies. The workshop received 20 submissions, 15 of which passed through a thorough review process of at least 3 reviewers, and 7 of them were accepted for publication and presentation. Apostolos P. Fournaris, Paolo Palmieri 0001 |
CCS | 1 |
| 2024 | SECURED for Health: Scaling Up Privacy to Enable the Integration of the European Health Data SpaceabstractIn this paper, we present the SECURED project11Funded in part by the European Union (EU), Grant Agreement no. 10109571. Views and opinions expressed are those of the authors and do not necessarily reflect those of the EU or the Health and Digital Executive Agency. Neither the EU nor the granting authority are responsible for them., aimed at improving privacy-preserving processing of data in the health domain. The technologies developed in the project will be demonstrated in four health-related use cases and with the involvement of SME's selected through an open funding call. Francesco Regazzoni 0001, Gergely Ács, Albert Zoltan Aszalos, Christos Avgerinos, Nikolaos Bakalos, Josep Lluís Berral, Joppe W. Bos, Marco Brohet, Andrés G. Castillo, Gareth T. Davies, Stefanos Florescu, Pierre-Elisée Flory, Alberto Gutierrez-Torre, Evangelos Haleplidis, Alice Héliou, Sotiris Ioannidis, Alexander El-Kady, Katarzyna Kapusta, Konstantina Karagianni, Pieter Kruizinga, Kyrian Maat, Zoltán Ádám Mann, Kalliopi Mastoraki, SeoJeong Moon, Maja Nisevic, Balazs Pejo, Kostas Papagiannopoulos, Vassilis Paliouras, Paolo Palmieri 0001, Francesca Palumbo, Juan Carlos Pérez Baun, Péter Pollner, Eduard Porta-Pardo, Luca Pulina, Muhammad Ali Siddiqi, Daniela Spajic, Christos Strydis, George Tasopoulos, Vincent Thouvenot, Christos Tselios, Apostolos P. Fournaris |
DATE | 41 |
| 2023 | CCSW '23: Cloud Computing Security WorkshopabstractClouds and massive-scale computing infrastructures are starting to dominate computing and will likely continue to do so for the foreseeable future. Major cloud operators are now comprising millions of cores hosting substantial fractions of corporate and government IT infrastructure. CCSW is the world's premier forum bringing together researchers and practitioners in all security aspects of cloud-centric and outsourced computing, including: Francesco Regazzoni 0001, Apostolos P. Fournaris |
CCS | 2 |
| 2023 | Energy Consumption Evaluation of Post-Quantum TLS 1.3 for Resource-Constrained Embedded DevicesabstractPost-Quantum cryptography (PQC), in the past few years, constitutes the main driving force of the quantum resistance transition for security primitives, protocols and tools. TLS is one of the widely used security protocols that needs to be made quantum safe. However, PQC algorithms integration into TLS introduce various implementation overheads compared to traditional TLS that in battery powered embedded devices with constrained resources, cannot be overlooked. While there exist several works, evaluating the PQ TLS execution time overhead in embedded systems there are only a few that explore the PQ TLS energy consumption cost. In this paper, a thorough power/energy consumption evaluation and analysis of PQ TLS 1.3 on embedded systems has been made. A WolfSSL PQ TLS 1.3 custom implementation is used that integrates all the NIST PQC algorithms selected for standardisation as well as 2 out of 3 of those evaluated in NIST Round 4. Also 1 out of 2 of the BSI recommendations have been included. The PQ TLS 1.3 with the various PQC algorithms is deployed in a STM Nucleo evaluation board under a mutual and a unilateral client-server authentication scenario. The power and energy consumption collected results are analyzed in detail. The performed comparisons and overall analysis provide very interesting results indicating that the choice of the PQC algorithms in TLS 1.3 to be deployed on an embedded system may be very different depending on the device use as an authenticated or not authenticated, client or server. Also, the results indicate that in some cases, PQ TLS 1.3 implementations can be equally or more energy consumption efficient compared to traditional TLS 1.3. George Tasopoulos, Charis Dimopoulos, Apostolos P. Fournaris, Raymond K. Zhao, Amin Sakzad, Ron Steinfeld |
CF | 3 |
| 2023 | Invited Paper: Dilithium Hardware-Accelerated Application Using OpenCL-Based High-Level SynthesisabstractPost-quantum cryptography (PQC) has been gaining attention in the last few years due to the evolution of quantum computers and the need to replace traditional, quantum-attack-insecure cryptography schemes with quantum-attack-resistant schemes. Lattice-based cryptography (LBC) constitutes a highly promising post-quantum solution (Quantum Resistant), but implementations in software or hardware are challenging due to the use of operations based on large-size polynomials. LBC selected schemes for standardization by the National Institute of Standards and Technology (NIST), rely on matrix-to-matrix multiplications of high-order polynomials, having performance bottlenecks that are solved using the Number-Theoretic Transform (NTT). One of the NIST-selected schemes for digital signature (DS) is the CRYSTALS-Dilithium scheme, which uses$N=256$degee polynomials. In this paper a Hardware/Software (HW/SW) co-design solution is proposed for all security levels of Dilithium, utilizing the OpenCL framework and Vitis High-Level Synthesis (HLS) tool. In our work, the HW/SW co-design OpenCL mechanisms are analyzed extensively and communication overheads between the hardware kernel and an ARM processor are identified, while appropriate techniques are proposed in order to bypass the I/O time-bottleneck on a real-world application. The proposed implementation runs on the ARM Processing System (PS) of a Xilinx Multi-Processor System on Chip (MPSoC) system, utilizing the MPSoC FPGA Programmable Logic (PL) in order to accelerate the calculations relative to the heavy matrix-multiplication operation. Finally, the proposed HW/SW codesigned solution is realized as a real-world Linux-based Dilithium DS executable and manages to achieve realistic performance gain, in terms of time execution, versus a CPU-only execution ranging from 2-23% (depending on the utilized CPU Clock Frequency). Alexander El-Kady, Apostolos P. Fournaris, Vassilis Paliouras |
ICCAD | 2 |
| 2023 | A Design Approach and Prototype Implementation for Factory Monitoring Based on Virtual and Augmented Reality at the Edge of Industry 4.0abstractVirtual and augmented reality are currently enjoying a great deal of attention from the research community and the industry towards their adoption within industrial spaces and processes. However, the current design and implementation landscape is still very fluid, while the community as a whole has not yet consolidated into concrete design directions, other than basic patterns. Other open issues include the choice over a cloud or edge-based architecture when designing such systems. Within this work, we present our approach for a monitoring intervention inside a factory space utilizing both Virtual Reality (VR) and Augmented Reality (AR), based primarily on edge computing, while also utilizing the cloud. We discuss its main design directions, as well as a basic ontology to aid in simple description of factory assets. In order to highlight the design aspects of our approach, we present a prototype implementation, based on a use case scenario in a factory site, within the context of the EnerMan H2020 project. Georgios Mylonas, Apostolos P. Fournaris, Christos Koulamas |
INDIN | 3 |
| 2022 | Performance Evaluation of Post-Quantum TLS 1.3 on Resource-Constrained Embedded Systems
George Tasopoulos, Apostolos P. Fournaris, Raymond K. Zhao, Amin Sakzad, Ron Steinfeld |
ISPEC | 3 |
| 2022 | High-Level Synthesis design approach for Number-Theoretic MultiplierabstractLattice-based cryptography (LBC) performs polynomial multiplication using the Number Theoretic Transform (NTT), in order to reduce the polynomial multiplication complexity from O(n2) to O(n log n). Although NTT-based multipliers offer the fastest way to compute a polynomial multiplication product for high-degree polynomials (with non-trivial bit-length coefficients), they constitute a significant part of the overall LBC scheme delay thus becoming the main LBC efficiency bottleneck. Therefore, the need to optimize the NTT-based multiplication in an easy, automatic yet efficient manner is significant. High-Level synthesis (HLS) tools offer such a capability since they can hide the Register Transfer Level (RTL)-based design complexity (typically realized by hardware description languages) using high level descriptions in C, C++ or openCL. However, this design approach requires careful modifications for high-level description code like loop reordering, loop flattening, removing dependencies, loop pipelining and loop unrolling in order to produce through an HLS tool a design with performance comparable to RTL hand-crafted designs. In this paper, extending the work in [1] we propose a complete NTT-based polynomial multiplier that combines an HLS optimized Cooley-Tukey (CT) NTT design with a proposed, HLS optimized, Gentleman-Sande (GS) Inverse-NTT design to create a highly efficient multiplier design that can benefit from the HLS flexibility yet still achieve significant high speed. More specifically, in the paper, the read and write access of the NTT processing elements (PE) to the memory is significantly increased though appropriate code redesign and the use of the dependence HLS pragma is proposed in order to reduce the dependencies between PEs. The proposed work has been evaluated by introducing the proposed NTT multiplier in the LBC Dilithium digital-signature scheme (polynomial degree n = 256, coefficient modulus Q = 8380417) and managed to achieve significantly higher speed compared to other similar works. Alexander El-Kady, Apostolos P. Fournaris, Evangelos Haleplidis, Vassilis Paliouras |
VLSI-SoC | 2 |
| 2021 | Studying OpenCL-based Number Theoretic Transform for heterogeneous platformsabstractLattice based cryptography can be considered a candidate alternative for post-quantum cryptosystems offering key exchange, digital signature and encryption functionality. Number Theoretic Transform (NTT) can be utilized to achieve better performance for these functionalities, where polynomials are needed to be multiplied. NTT simplifies the multiplication overhead allowing point-wise multiplication by transforming the polynomials into the spectral domain and then inversing the result to the original domain. It is important to optimize this technique that is used in a wide range of computing systems. In this paper we study the feasibility of using OpenCL, a portable framework, to implement a parallelized version of NTT which allows deployment on heterogeneous platforms, such as Graphic Processing Units (GPUs) and Field Programmable Gate Arrays (FPGAs). We measure the performance of our implementation on a GPU and evaluate when and where such a deployment is beneficial. Our results showed that the proposed parallel implementation is a viable acceleration approach for these algorithms for lattice-based cryptography solutions. Evangelos Haleplidis, Thanasis Tsakoulis, Alexander El-Kady, Charis Dimopoulos, Odysseas G. Koufopavlou, Apostolos P. Fournaris |
DSD | 6 |
| 2021 | High-Level Synthesis design approach for Number-Theoretic Transform ImplementationsabstractLattice-based cryptography performs polynomial multiplication using the Number Theoretic Transform (NTT), in order to reduce the polynomial multiplication complexity from $O\left(n^{2}\right)$ to $O(n \log n)$. NTT has been in the center of investigation in cryptography space, as it is applied in many cryptography schemes such as hash functions, homomorphic encryption, key-encapsulation mechanisms, and digital signatures. A common approach for rapid production of hardware designs commences from semi-automatic software production, as supported by the Xilinx High-Level Synthesis (HLS) toolchain or similar tools. Most of the times this approach requires careful modifications (e.g. code modification, loop reordering, loop flattening, removing dependencies, loop pipelining, loop unrolling) in order to achieve a design with performance comparable to a Register-Transfer Level (RTL) hand-crafted design. In this paper a design solution is proposed that solves the data and loop-carry dependencies of the Cooley-Tukey NTT algorithm, by assisting the HLS synthesizer to produce efficient designs, in terms of latency and resources. The proposed work has been evaluated using the Dilithium digital-signature scheme NTT version ($n=256, Q$ of 23 bits), and is shown to achieve a 20-50 % improvement in terms of latency (without really affecting the resources) compared to other existing HLS-based NTT solutions in the literature. Alexander El-Kady, Apostolos P. Fournaris, Thanasis Tsakoulis, Evangelos Haleplidis, Vassilis Paliouras |
VLSI-SoC | 2 |
| 2020 | Privacy Preservation in Industrial IoT via Fast Adaptive Correlation Matrix CompletionabstractThe Industrial Internet of Things (IIoT) is a key element of industry 4.0, bringing together modern sensor technology, fog and cloud computing platforms, and artificial intelligence to create smart, self-optimizing industrial equipment and facilities. Though, the scale and sensitivity degree of information continuously increases, giving rise to serious privacy concerns. The scope of this article is to provide efficient privacy preservation techniques, by tracking the correlation of multivariate streams recorded in a network of IIoT devices. The time-varying data covariance matrix is used to add noise that cannot be easily removed by filtering, generating obfuscated measurements and, thus, preventing unauthorized access to the original data. To improve communication efficiency between connected IoT devices, we exploit inherent properties of the correlation matrices, and track the essential correlations from a small subset of correlation values. Extensive simulation studies using constrained IIoT devices validate the robustness, efficiency, and effectiveness of our approach. Aris S. Lalos, Evangelos Vlachos, Kostas Berberidis, Apostolos P. Fournaris, Christos Koulamas |
IEEE Trans. Ind. Informatics | 4 |
| 2019 | Robust and Efficient Privacy Preservation in Industrial IoT via correlation completion and trackingabstractThe Industrial IoT (IIoT) is a key element of Industry 4.0, bringing together modern sensor technology, fog - cloud computing platforms, and artificial intelligence (AI) to create smart, self-optimizing industrial equipment and facilities. Though, the scale and sensitivity degree of information continuously increases, giving rise to serious privacy concerns. In this work we address the problem of efficiently and effectively tracking the structure of multivariate streams recorded in a network of IIoT devices. The time varying correlation data values are used to add noise which maximally preserves privacy, in the sense that it is very hard to be removed. To improve communication efficiency between connected IoT devices, we exploit low rank properties of the correlation matrices, and track the essential correlations from a small subset of correlation values estimated by a subset of network nodes. Extensive simulation studies, validate the correctness, efficiency, and effectiveness of our approach in terms of computational complexity, transmission energy efficiency and privacy preservation. Aris S. Lalos, Evangelos Vlachos, Kostas Berberidis, Apostolos P. Fournaris, Christos Koulamas |
INDIN | 4 |
| 2018 | An FPGA Hardware Trojan Detection Approach Based on Multiple Parameter AnalysisabstractIn this paper an FPGA Hardware Trojan (HT) detection approach based on multiple parameter analysis is proposed. In this direction, we apply a logic testing method, a run-time method and a side-channel analysis method. Logic testing and side-channel analysis methods are non-invasive while the run-time method is invasive in the sense that on-chip digital sensors are used to detect unexpected differentiations in the layout of the IC. The introduced methods do not rely on the presence of a "Golden chip" for detecting the HT. The proposed approach is implemented and evaluated on an actual FPGA board, providing practical results that validate our assumptions. To the best of our knowledge this is the first attempt for combining three different parameter analysis for HT detection in an of-the-shelf FPGA board. Apostolos P. Fournaris, Lampros Pyrgas, Paris Kitsos |
DSD | 1 |
| 2018 | IoT Integration for Adaptive ManufacturingabstractThe Industrial Internet of Things (IIoT) has emerged as a concept for the integration of the Internet of Things (IoT) in the industrial domain, posing a variety of challenges and emphasizing among others in interoperability and integration aspects with Shop Floor and Plant layer systems. Present work builds on top of previous work related to the integration of the different systems in the manufacturing environment automation pyramid. This classical hierarchy involves three layers: the Enterprise Resource Planning (ERP), the Manufacturing Execution System (MES) and the Shop Floor. A further issue is relevant to the semantic integration of IoT sensors with manufacturing systems. The present paper presents a complex system architecture which allows both the aggregation of data generated by IIoT devices, and the automated reaction of the manufacturing environment to either detected and diagnosed, or predicted events based on this data. Christos Alexakos, Apostolos P. Fournaris, Christos Koulamas, Athanasios P. Kalogeras |
ISORC | 3 |
| 2017 | A Design Strategy for Digit Serial Multiplier Based Binary Edwards Curve Scalar Multiplier ArchitecturesabstractBinary Edwards Curves (BEC) constitute an alternative to the standardized Weierstrass elliptic curve (EC) equations since the latter have intrinsic side channel attack vulnerabilities due to their lack of point operation uniformity. Thus, BECs have gained popularity over the past few years due to their uniformity, operation regularity, completeness and implementation attractiveness. However, BEC Scalar multiplication hardware implementations are still lacking in performance when compared to their Weierstrass equivalent. In this paper, a design strategy/methodology is proposed in order to realizeBEC Scalar multipliers with a good trade-off between computation speed and utilized hardware resources. The strategy is based on a GF(2k) operations parallelism mechanism that aims at the minimization of idle states in the utilized processing elements as well as the minimization of employed storage andcontrol elements. A BEC SM architecture is proposed in order to describe the realization of the proposed strategy and its benefits are analyzed. In order to evaluate the efficiency of the proposed architecture an implementation was made in FPGA technology for GF(2^233) fields with BECs using d1 = d2. Whencompared to other works, the implementation expressed very balanced results. Apostolos P. Fournaris, Charalambos Dimopoulos, Odysseas G. Koufopavlou |
DSD | 1 |
| 2017 | Production process adaptation to IoT triggered manufacturing resource failure eventsabstractUsage of raw data as an asset, from which value can be created to support business and manufacturing decision making, motivated a lot of scientists to explore the challenges on how to exploit this value. Such efforts are concentrated on the framework of “data value chain”, where approaches of architectures and applications aim at equipping enterprises with tools that gather, process and extract knowledge from raw data generated by their internal or external processes. In this context, the present paper deals with the way the production processes in a factory can be adapted to changes that are detected by the processing of raw data which are aggregated by IoT devices installed in the manufacturing environment. The proposed approach introduces a complex system that combines a network of IoT sensors and a high-level multi-agent system that contributes to the vertical integration of all the systems residing in the Enterprise/Factory. Christos Alexakos, Apostolos P. Fournaris, Athanasios P. Kalogeras, Christos Koulamas |
ETFA | 3 |
| 2015 | Affine Coordinate Binary Edwards Curve Scalar Multiplier with Side Channel Attack ResistanceabstractTaking into account the high regularity and completeness of Binary Edwards Curves (BEC), BEC point operation efficient implementation in hardware becomes a need especially since such curves tend to be more resistant against side channel attacks than the classical Weierstrass Elliptic Curves. However, BECs require more GF(2k) operations for a single scalar multiplication. This constitutes a deterring factor for their wide adoption and standardization. In this paper, a design methodology, hardware architecture and implementation is proposed on the efficient implementation of BEC scalar multiplication accelerators. To achieve that, a parallelism approach is introduced on affine coordinate representation BECs supporting fast GF(2k) inversion through a GF(2k) inversion algorithm capable of realizing also GF(2k) multiplication. The resulting architecture using 4 parallel operating GF(2k) arithmetic units when implemented in FPGA technology provide better results than similar Weierstrass Curves following parallelism techniques, indicated that BECs support parallelism better than their Weierstrass equivalent. Apostolos P. Fournaris, Odysseas G. Koufopavlou |
DSD | 1 |
| 2015 | Designing efficient elliptic Curve Diffie-Hellman accelerators for embedded systemsabstractIn this paper, a methodology towards a hardware/software implementation of an Elliptic Curve Diffie Hellman (ECDH) scheme is proposed in an effort to overcome the design problems of Elliptic Curve Cryptography (ECC) systems stemming from the highly constrained embedded system hardware and software environment (restricted RAM, storage and processing power). To achieve that, instead of the excessively slow software ECDH implementations or monolithic, not flexible hardware implementations, we propose the use of a flexible, scalar multiplication (SM) accelerator connected to the main embedded system processor in order to speed up ECDH functionality without downgrading the overall main processor performance. The proposed solution can be used for a wide variety of GF(2k) based Elliptic Curves (EC) and is capable of shifting from one EC to another EC at runtime (flexibility). The proposed architecture was implemented and tested in Xilinx Virtex 5 technology by realizing the proposed SM accelerator unit interconnected with a Xilinx microblaze softcore processor. Apostolos P. Fournaris, Ioannis Zafeirakis, Christos Koulamas, Nicolas Sklavos 0001, Odysseas G. Koufopavlou |
ISCAS | 1 |
| 2015 | Challenges in designing trustworthy cryptographic co-processorsabstractSecurity is becoming ubiquitous in our society. However, the vulnerability of electronic devices that implement the needed cryptographic primitives has become a major issue. This paper starts by presenting a comprehensive overview of the existing attacks to cryptography implementations. Thereafter, the state-of-the-art on some of the most critical aspects of designing cryptographic co-processors are presented. This analysis starts by considering the design of asymmetrical and symmetrical cryptographic primitives, followed by the discussion on the design and online testing of True Random Number Generation. To conclude, techniques for the detection of Hardware Trojans are also discussed. Ricardo Chaves, Giorgio Di Natale, Lejla Batina, Shivam Bhasin, Baris Ege, Apostolos P. Fournaris, Nele Mentens, Stjepan Picek, Francesco Regazzoni 0001, Vladimir Rozic, Nicolas Sklavos 0001, Bohan Yang 0001 |
ISCAS | 6 |
| 2014 | Designing and Evaluating High Speed Elliptic Curve Point MultipliersabstractPoint Multiplication (PM) is considered the most computationally complex and resource hungry Elliptic Curve Cryptography (ECC) related mathematic operation. The design of PM hardware accelerators follows approaches that have a trade off between utilized hardware resources and computation speed. In this paper, the above trade-off and its relation with the operations of the GF(2k) defining the Elliptic Curve (EC) is highlighted and investigated. Following this direction, a point operation design methodology based on the parallelization and scheduling of GF(2k) operations is proposed. This design approach is adapted to the PM employed GF(2k) multiplication algorithm and associated implementation in an effort to increase PM accelerator speed with an acceptable cost on chip covered area (hardware resources). Using the proposed methodology, two PM accelerator hardware architectures were proposed based on bit serial and bit parallel GF(2k) multipliers that, when implemented in FPGA technology, proved to be very fast in comparison to other similar works. Apostolos P. Fournaris, John Zafeirakis, Odysseas G. Koufopavlou |
DSD | 1 |
| 2012 | CRT RSA Hardware Architecture with Fault and Simple Power Attack CountermeasuresabstractRSA cryptographic algorithm has long achieved cryptographic and market maturity. However, RSA implementations, after the discovery of Side Channel Attacks (SCA), are susceptible to a variety of different attacks that target the hardware structure rather than the algorithm itself. There are a wide range of countermeasures that can be applied on the RSA structure in order to protect the algorithm from SCAs, however few of them are efficient in hardware since they add extensive performance cost to an SCA resistant RSA implementation. In this paper, a hardware architecture is proposed based on a Fault attack (FA) and Simple Power attack (SPA) resistant algorithm for Chinese Remainder Theorem (CRT) RSA that through the principles of parallelism and component reusability can guarantee hardware efficiency. We describe an implementation approach based on Montgomery modular multiplication and also propose a testing hardware architecture to simulate the security chip environment that our FA-SPA resistant CRT RSA can be integrated in. The designed architecture is implemented in FPGA technology and results on its time and space complexity are extracted and evaluated. Apostolos P. Fournaris, Odysseas G. Koufopavlou |
DSD | 1 |
| 2012 | Distributed Threshold Certificate based Encryption Scheme with No Trusted Dealer
Apostolos P. Fournaris |
SECRYPT | 1 |
| 2011 | Efficient CRT RSA with SCA CountermeasuresabstractRSA cryptographic algorithm, working as a security tool for many years, has long achieved cryptographic and market maturity. However, as all crypto algorithms, RSA implementations, after the discovery and wide spread of Side Channel Attacks (SCA), are susceptible to a wide variety of different attacks that target the hardware structure rather than the algorithm itself. While there are a wide range of countermeasures that can be applied on the RSA structure in order to protect the algorithm from SCAs, combining several such measures in order to guarantee an SCA resistant RSA design is not an easy job. There are many incompatibility issues among SCA protection methods as well as an extensive performance cost added to an SCA secure RSA implementation. In this paper, we address some very popular and potent SCAs against RSA like Fault attacks (FA), Simple Power attacks (SPA), Doubling attacks (DA) and Differential Power attacks (DPA), and propose an algorithmic modification of RSA based on Chinese Remainder Theorem (CRT) that can thwart those attacks. We describe an implementation approach based on Montgomery modular multiplication and propose a hardware architecture for a SCA resistant CRT RSA that is structured on our proposed algorithm. The designed architecture is implemented in FPGA technology and results on its time and space complexity are extracted and evaluated. Apostolos P. Fournaris, Odysseas G. Koufopavlou |
DSD | 1 |
| 2011 | Distributed Threshold Cryptography Certification with No Trusted Dealer
Apostolos P. Fournaris |
SECRYPT | 1 |
| 2010 | Fault and simple power attack resistant RSA using Montgomery modular multiplicationabstractSide channel attacks and more specifically fault, simple power attacks, constitute a pragmatic, potent mean of braking a cryptographic algorithm like RSA. For this reason, many researchers have proposed modifications on the arithmetic operation functions required for RSA in order to thwart those attacks. However, these modifications are applied on theoretic - algorithmic level and do not necessary result in high performance RSA designs. This paper constitute the first complete attempt for an efficient design approach on a fault and simple power attack resistant RSA based on the well known, for its high performance, Montgomery multiplication algorithm. To achieve this, a fault and simple power attack resistant modular exponentiation algorithm is proposed that is based on the Montgomery modular multiplication. In order to optimize this algorithm's performance we also propose a modified version of Montgomery modular multiplication algorithm that employs value precomputation and carry save logic in all input, output and intermediate values. We introduce a hardware architecture based on the proposed Montgomery modular multiplication algorithm and use it as a building block for the design of a fault and simple power attack resistant modular exponentiation unit. This unit is optimized by taking advantage of the inherit parallelism in the proposed fault and simple power attack resistant modular exponentiation algorithm. Realizing the proposed unit in FPGA technology very advantageous results are found when compared against other well known designs even though our design bears an extra computation cost due to its fault and simple power attack resistance characteristic. Apostolos P. Fournaris |
ISCAS | 1 |
| 2009 | One Dimensional Systolic Inversion Architecture Based on Modified GF(2^k) Extended Euclidean AlgorithmabstractThe need for small chip covered area in most handheld devices with out sacrifices in computational power introduces an interesting problem concerning expensive, computational intensive operations, like GF(2k) inversion which is widely used in cryptography. This paper addresses this problem by proposing a systolic inversion architecture for GF(2k) fields. This architecture is based on an extended analysis on an optimized version of modified extended Euclidean algorithm (OMEEA) that is using signal reusability and simplification of the control signals with regard to hardware design and manages to make the inversion process less complex. The proposed one dimensional systolic inversion architecture based on OMEEA was measured in terms of hardware components number, latency and critical path delay with very interesting results when compared to other well known designs thus proving the efficiency of the analysis on OMEEA algorithm. Apostolos P. Fournaris, Odysseas G. Koufopavlou |
DSD | 1 |
| 2009 | Low Area Elliptic Curve Arithmetic UnitabstractIn this paper, a generic elliptic curve (EC) arithmetic unit with high flexibility and small chip covered area is proposed. This EC arithmetic unit is based on the one dimensional systolic architectural realization of a proposed modified multiplication - inversion algorithm that through appropriate initialization uses the algorithmic structure of inversion to also perform multiplication. The proposed architecture is realized on FPGA for GF(2163) and is compared with similar up-to-date designs. The proposed EC arithmetic unit has very small chip covered area without any serious penalty in calculation delay and since is designed on the affine coordinate plane, it offers a good side channel attack resistance base for further optimizations on this field. Apostolos P. Fournaris, Odysseas G. Koufopavlou |
ISCAS | 1 |
| 2009 | Improved throughput bit-serial multiplier for GF(2m) fields
Georgios N. Selimis, Apostolos P. Fournaris, Harris E. Michail, Odysseas G. Koufopavlou |
Integr. | 2 |
| 2008 | Creating an Elliptic Curve arithmetic unit for use in elliptic curve cryptographyabstractElliptic curve cryptography (ECC) is a very promising cryptographic method, offering the same security level as traditional public key cryptosystems (RSA, El Gamal) but with considerably smaller key lengths. To increase the performance of an EC Cryptosystem, dedicated hardware is employed for all EC point operations. However, the computational complexity and hardware resources of an Elliptic Curve processing unit are very high and depend on the efficient design of the Elliptic Curvepsilas underlined GF(2k) Field. In this paper, we propose an EC arithmetic unit that is structured over a high peformance, low gate number GF(2k) arithmetic unit. This proposed GF(2k) arithmetic unit is based on one dimensional systolic architecture that can perform GF(2k) multiplication and inversion with only the performance cost of inversion. This is achieved by utilizing a multiplication/inversion algorithm based on the modified extended Euclidean algorithm. Apostolos P. Fournaris, Odysseas G. Koufopavlou |
ETFA | 1 |
| 2008 | Versatile multiplier architectures in GF(2k) fields using the Montgomery multiplication algorithm
Apostolos P. Fournaris, Odysseas G. Koufopavlou |
Integr. | 1 |
| 2006 | An RNS architecture of an Fp elliptic curve point multiplierabstractAn elliptic curve point multiplier (ECPM) is the main part of all elliptic curve cryptography (ECC) systems and its performance is decisive for the performance of the overall cryptosystem. A VLSI residue number system (RNS) architecture of an ECPM is presented in this paper. In the proposed approach, the necessary mathematical conditions that need to be satisfied, in order to replace typical finite field circuits with RNS ones, are investigated. It is shown that such an application is feasible and that it leads to a significant improvement in the execution time of a scalar point multiplication Dimitrios M. Schinianakis, Apostolos P. Fournaris, Athanasios Kakarountas, Thanos Stouraitis |
ISCAS | 2 |