Alexander El-Kady

dblp:303/8092 · also Alexander Islam El-Kady · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2024 SECURED for Health: Scaling Up Privacy to Enable the Integration of the European Health Data Space
abstract
In this paper, we present the SECURED project11Funded in part by the European Union (EU), Grant Agreement no. 10109571. Views and opinions expressed are those of the authors and do not necessarily reflect those of the EU or the Health and Digital Executive Agency. Neither the EU nor the granting authority are responsible for them., aimed at improving privacy-preserving processing of data in the health domain. The technologies developed in the project will be demonstrated in four health-related use cases and with the involvement of SME's selected through an open funding call.
Francesco Regazzoni 0001, Gergely Ács, Albert Zoltan Aszalos, Christos Avgerinos, Nikolaos Bakalos, Josep Lluís Berral, Joppe W. Bos, Marco Brohet, Andrés G. Castillo, Gareth T. Davies, Stefanos Florescu, Pierre-Elisée Flory, Alberto Gutierrez-Torre, Evangelos Haleplidis, Alice Héliou, Sotiris Ioannidis, Alexander El-Kady, Katarzyna Kapusta, Konstantina Karagianni, Pieter Kruizinga, Kyrian Maat, Zoltán Ádám Mann, Kalliopi Mastoraki, SeoJeong Moon, Maja Nisevic, Balazs Pejo, Kostas Papagiannopoulos, Vassilis Paliouras, Paolo Palmieri 0001, Francesca Palumbo, Juan Carlos Pérez Baun, Péter Pollner, Eduard Porta-Pardo, Luca Pulina, Muhammad Ali Siddiqi, Daniela Spajic, Christos Strydis, George Tasopoulos, Vincent Thouvenot, Christos Tselios, Apostolos P. Fournaris
DATE17
2023 Invited Paper: Dilithium Hardware-Accelerated Application Using OpenCL-Based High-Level Synthesis
abstract
Post-quantum cryptography (PQC) has been gaining attention in the last few years due to the evolution of quantum computers and the need to replace traditional, quantum-attack-insecure cryptography schemes with quantum-attack-resistant schemes. Lattice-based cryptography (LBC) constitutes a highly promising post-quantum solution (Quantum Resistant), but implementations in software or hardware are challenging due to the use of operations based on large-size polynomials. LBC selected schemes for standardization by the National Institute of Standards and Technology (NIST), rely on matrix-to-matrix multiplications of high-order polynomials, having performance bottlenecks that are solved using the Number-Theoretic Transform (NTT). One of the NIST-selected schemes for digital signature (DS) is the CRYSTALS-Dilithium scheme, which uses$N=256$degee polynomials. In this paper a Hardware/Software (HW/SW) co-design solution is proposed for all security levels of Dilithium, utilizing the OpenCL framework and Vitis High-Level Synthesis (HLS) tool. In our work, the HW/SW co-design OpenCL mechanisms are analyzed extensively and communication overheads between the hardware kernel and an ARM processor are identified, while appropriate techniques are proposed in order to bypass the I/O time-bottleneck on a real-world application. The proposed implementation runs on the ARM Processing System (PS) of a Xilinx Multi-Processor System on Chip (MPSoC) system, utilizing the MPSoC FPGA Programmable Logic (PL) in order to accelerate the calculations relative to the heavy matrix-multiplication operation. Finally, the proposed HW/SW codesigned solution is realized as a real-world Linux-based Dilithium DS executable and manages to achieve realistic performance gain, in terms of time execution, versus a CPU-only execution ranging from 2-23% (depending on the utilized CPU Clock Frequency).
Alexander El-Kady, Apostolos P. Fournaris, Vassilis Paliouras
ICCAD1
2022 High-Level Synthesis design approach for Number-Theoretic Multiplier
abstract
Lattice-based cryptography (LBC) performs polynomial multiplication using the Number Theoretic Transform (NTT), in order to reduce the polynomial multiplication complexity from O(n2) to O(n log n). Although NTT-based multipliers offer the fastest way to compute a polynomial multiplication product for high-degree polynomials (with non-trivial bit-length coefficients), they constitute a significant part of the overall LBC scheme delay thus becoming the main LBC efficiency bottleneck. Therefore, the need to optimize the NTT-based multiplication in an easy, automatic yet efficient manner is significant. High-Level synthesis (HLS) tools offer such a capability since they can hide the Register Transfer Level (RTL)-based design complexity (typically realized by hardware description languages) using high level descriptions in C, C++ or openCL. However, this design approach requires careful modifications for high-level description code like loop reordering, loop flattening, removing dependencies, loop pipelining and loop unrolling in order to produce through an HLS tool a design with performance comparable to RTL hand-crafted designs. In this paper, extending the work in [1] we propose a complete NTT-based polynomial multiplier that combines an HLS optimized Cooley-Tukey (CT) NTT design with a proposed, HLS optimized, Gentleman-Sande (GS) Inverse-NTT design to create a highly efficient multiplier design that can benefit from the HLS flexibility yet still achieve significant high speed. More specifically, in the paper, the read and write access of the NTT processing elements (PE) to the memory is significantly increased though appropriate code redesign and the use of the dependence HLS pragma is proposed in order to reduce the dependencies between PEs. The proposed work has been evaluated by introducing the proposed NTT multiplier in the LBC Dilithium digital-signature scheme (polynomial degree n = 256, coefficient modulus Q = 8380417) and managed to achieve significantly higher speed compared to other similar works.
Alexander El-Kady, Apostolos P. Fournaris, Evangelos Haleplidis, Vassilis Paliouras
VLSI-SoC1
2021 Studying OpenCL-based Number Theoretic Transform for heterogeneous platforms
abstract
Lattice based cryptography can be considered a candidate alternative for post-quantum cryptosystems offering key exchange, digital signature and encryption functionality. Number Theoretic Transform (NTT) can be utilized to achieve better performance for these functionalities, where polynomials are needed to be multiplied. NTT simplifies the multiplication overhead allowing point-wise multiplication by transforming the polynomials into the spectral domain and then inversing the result to the original domain. It is important to optimize this technique that is used in a wide range of computing systems. In this paper we study the feasibility of using OpenCL, a portable framework, to implement a parallelized version of NTT which allows deployment on heterogeneous platforms, such as Graphic Processing Units (GPUs) and Field Programmable Gate Arrays (FPGAs). We measure the performance of our implementation on a GPU and evaluate when and where such a deployment is beneficial. Our results showed that the proposed parallel implementation is a viable acceleration approach for these algorithms for lattice-based cryptography solutions.
Evangelos Haleplidis, Thanasis Tsakoulis, Alexander El-Kady, Charis Dimopoulos, Odysseas G. Koufopavlou, Apostolos P. Fournaris
DSD3
2021 High-Level Synthesis design approach for Number-Theoretic Transform Implementations
abstract
Lattice-based cryptography performs polynomial multiplication using the Number Theoretic Transform (NTT), in order to reduce the polynomial multiplication complexity from $O\left(n^{2}\right)$ to $O(n \log n)$. NTT has been in the center of investigation in cryptography space, as it is applied in many cryptography schemes such as hash functions, homomorphic encryption, key-encapsulation mechanisms, and digital signatures. A common approach for rapid production of hardware designs commences from semi-automatic software production, as supported by the Xilinx High-Level Synthesis (HLS) toolchain or similar tools. Most of the times this approach requires careful modifications (e.g. code modification, loop reordering, loop flattening, removing dependencies, loop pipelining, loop unrolling) in order to achieve a design with performance comparable to a Register-Transfer Level (RTL) hand-crafted design. In this paper a design solution is proposed that solves the data and loop-carry dependencies of the Cooley-Tukey NTT algorithm, by assisting the HLS synthesizer to produce efficient designs, in terms of latency and resources. The proposed work has been evaluated using the Dilithium digital-signature scheme NTT version ($n=256, Q$ of 23 bits), and is shown to achieve a 20-50 % improvement in terms of latency (without really affecting the resources) compared to other existing HLS-based NTT solutions in the literature.
Alexander El-Kady, Apostolos P. Fournaris, Thanasis Tsakoulis, Evangelos Haleplidis, Vassilis Paliouras
VLSI-SoC1