Eslam Yahya Tawfik

dblp:295/8274 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 8 since 2021
YearPublicationVenuePosition
2025 An Efficient Number Theoretic Transform Implementation for FIPS-203 on FPGA
abstract
The Number Theoretic Transform (NTT) is a key component in modern Post-Quantum Cryptography (PQC) systems, known for its efficient polynomial multiplication capabilities. This paper presents an innovative conflict-free NTT memory architecture that features a mathematically optimized Barrett-based modular reduction with bit correction, tailored for FIPS-203. The primary goal is to significantly reduce hardware resource usage while maintaining high performance and superior efficiency (frequency of operation/area) compared to prior work. Our NTT design, implemented on a Xilinx FPGA Virtex-7, achieves a clock speed of 300 MHz, to the best of our knowledge, the highest reported in the field, outperforming the best-known implementation by 15%. Additionally, it utilizes a single RAMB, one DSP block, 374 LUTs, and 270 registers, making it the most resource-efficient design in the literature, with less than half the hardware usage of prior work. Compared to the leading designs, the proposed architecture reduces LUT utilization by 38%, register usage by 58%, DSP utilization by 50%, and RAMB usage by 75%. As a result, the overall efficiency measured by the area-time product (ATP) surpasses the best design in the literature by a factor of 1.38. These results establish our approach as an optimal solution for applications that require minimal resources and low power consumption without compromising performance.
Sherif Maher Elewa, Eslam Yahya Tawfik
ISCAS2
2025 Optimized and Reconfigurable Hardware Design for ASCON AEAD and Hash Standards
abstract
NIST has finalized LWC standardization process by selecting ASCON as the new standard. ASCON is a versatile algorithm supporting two primary functions: AEAD and HASH. The standard includes three AEAD variants "ASCON-128, ASCON-128a, ASCON-80pq" and two HASH variants "Hash, Hasha". In this work an efficient programable hardware solution that supports all the three AEAD and two HASH variants is designed targeting GF12nm ASIC technology and Spartan-7 FPGA. The programmable design is compared to individual hardware implementations of each variant, demonstrating significant enhancements. By utilizing a single programmable design instead of five separate ones, 61% reduction in total silicon area and 75% energy savings are achieved. In addition, a comprehensive comparison of hardware implementations for each ASCON variant is provided. Performance metrics such as area, throughput, and energy consumption are evaluated, revealing that ASCON-128a offers the lowest energy consumption (0.45 pJ/bit) with the highest throughput (12.15 Gbps), while ASCON-128 achieves the smallest silicon footprint (1242 μm2on GF12nm).
Islam Elsadek, Eslam Yahya Tawfik
ISCAS2
2025 State-of-the-Art ASCON ASIC Achieving 4.3 [email protected] and 3.5Tb/[email protected] Outperforming AES by 25 Times
abstract
NIST selected ASCON as the standard Lightweight Cryptography (LWC) algorithm in 2023. ASCON’s implementations promise bringing lightweight Authenticated Encryption (AEAD) to resource-constrained devices surpassing Advanced Encryption Standard (AES) implementations. In this work, a standard compliant ASCON Application Specific Integrated Circuit (ASIC) hardware (HW) is designed and fabricated using CMOS GF22FDx technology. This study provides a quantitative assessment of the HW with two other standard compliant implementations of ASCON. One is software implementation (SW), and the other is a hardware accelerated (HW/SW co-design) implementation. The assessment shows that HW outperforms the SW implementation by up to three orders of magnitude in energy efficiency and throughput, whereas HW/SW co-design throughput and energy efficiency falls in the middle between HW and software. ASCON ASIC HW is also compared with a standard HW implementation of the AES fabricated over the same chip. ASCON uses only 39% of AES’s area and boosts energy efficiency by up to 25 times. To the best of our knowledge, this work is the first work providing a silicon-based analysis for ASCON ASIC implementation reaching a throughput of 4.3 Gbps @ 0.8V and 2 Gbps @ 0.6V, and energy efficiency of 1.9 Tb/J @ 0.8V and 3.5 Tb/J @ 0.6V in$2505~\mu $m2 on GF22FDx at 620MHz @ 0.8V. Furthermore, the comparative assessments between different implementations of ASCON guides the implementation choice for specific deployments to meet the demands of secure processing in dust-size sensors, edge and IoT.
Islam Elsadek, Elsayed Elgendy, Sherif Abouzeid, Ahmed Zaky Ghonem, John Ross Wallrabenstein, Erik MacLean, Douglas Gardner, Sohrab Aftabjahani, Rosario Cammarota, Eslam Yahya Tawfik
IEEE Trans. Circuits Syst. I Regul. Pap.10
2024 ASIC Implementation of Efficient 512-Neuron 256K-Synapses Digital Neuromorphic Processor with On-Chip Encoding in 22nmFDX
abstract
There is a rising demand for AI workloads running on the edge that are increasingly complex, calling for more efficient computing platforms. Neuromorphic hardware is a class of electronic circuits and devices that aim to mimic a biological brain’s functionality and computational efficiency. In this work, we propose a 512 neuron 256K synapses LIF-based digital neuromorphic processor. Each neurosynaptic core can be configured to act as either an encoding neuron or as a processing neuron, chosen to be Leaky-Integrate-and-Fire (LIF). We have also implemented a 1-1 software code for the processor, in which Pytorch and SNNtorch are used to build the spiking neural network, and surrogate gradients are used for training. An image classification task is used to evaluate the performance of the neuromorphic processor using the MNIST dataset, achieving an accuracy of 96.12% for 4-bit quantized synaptic weights. Implemented using GF22nmFDX technology, the processor is 0.72mm2in area, of which synaptic memories occupy 68% of the area, and achieves an average power efficiency of 14.8 nJ/SOP. This large area is justified by the configuration flexibility of the processor, where each neuron can be connected to any other. This allows the mapping of different and more sophisticated tasks.
Ahmed Zaky Ghonem, Eslam Yahya Tawfik
ISCAS2
2022 Hardware and Energy Efficiency Evaluation of NIST Lightweight Cryptography Standardization Finalists
abstract
Current cryptographic algorithms are designed for server environments prioritizing security with no limitations on hardware resources. They may not be suitable for emerging resource-constrained devices used in areas such as Edge computing, UAV, and IoT. For such constrained devices, many LWC algorithms have been proposed, however, there is no FIPS standard yet. So, NIST initiated a standardization process for a LWC FIPS standard. Finalists are announced with 10 algorithms after two rounds of evaluation. The aim of this work is to design and evaluate the hardware of these candidates using ASIC synthesis over GF 22nm CMOS technology. The evaluation focuses on energy efficiency using bit/joule as the main metric. Other metrics such as throughput and area are evaluated as well. Results show a great discrepancy in the energy efficiency between the finalists. For example, TinyJambu, Xoodyak and ASCON achieved 10-25 times better energy efficiency compared to ISAP, Elephant, and Grain-128AEAD while processing the same number of bits.
Islam Elsadek, Sohrab Aftabjahani, Doug Gardner, Erik MacLean, John Ross Wallrabenstein, Eslam Yahya Tawfik
ISCAS6
2022 Energy Efficiency Enhancement Of Parallelized Implementation of NIST Lightweight Cryptography Standardization Finalists
abstract
Parallelism and pipelining are widely used to improve the performance and throughput of systems. However, its effect on energy consumption needs to be studied. In this paper the alteration in energy consumption that results from using parallel architecture is studied over LWC algorithms from NIST standardization process. Ten algorithms are currently in the final round of the standardization process. Two algorithms out of the ten final round candidates can be parallelized which are Elephant and ISAP algorithms. For these algorithms, both iterative looping and parallel architectures are designed and synthesized over ASIC GF22nm technology. Then both architectures are compared in terms of area, throughput and energy. Results showed an enhancement in energy efficiency up to 49% and 28% and throughput improvement reaches up to 96% and 45% in Elephant and ISAP, respectively.
Islam Elsadek, Sohrab Aftabjahani, Doug Gardner, Erik MacLean, John Ross Wallrabenstein, Eslam Yahya Tawfik
ISCAS6
2022 Low-Complexity AES Architectures Resilient to Power Analysis Attacks
abstract
The advanced encryption standard (AES) is the current standard for symmetric-key cipher. To protect AES implementations from correlation power analysis (CPA) side-channel attacks (SCAs), many countermeasures have been proposed. However, existing approaches have large area overheads. This paper proposes two low-complexity techniques for AES to resist CPA attacks. By exploiting generalized dual ciphers, alternative conversions are utilized to substantially simplify all the involved constant matrix multiplications. Additionally, a multiplicative masking scheme utilizing simple constant multipliers is developed for AES designs used in resource-constraint applications. FPGA implementation results on a Xilinx XC7a200t device show that the proposed design with four Sboxes achieves 11.1% area reduction and 89.2% clock frequency increase compared to the best prior architecture without sacrificing CPA-attack resistance. In addition, the proposed multiplicative masking scheme not only keeps the resistance to CPA attacks, but also reduces the area by 52.6% and increases the clock frequency by 29.7% in a single-Sbox design compared to the previous scheme.
Elsayed Elgendy, Eslam Yahya Tawfik, Xinmiao Zhang 0001
ISCAS3
2021 Impact of Physical Design on PUF Behavior: A Statistical Study
abstract
FPGA-based designs dominate many applications for their faster implementation, configurability, and low design cost. These applications need a root of trust to be secured against malicious activities. Physical unclonable functions (PUF) are promising security primitive that can be used to identify silicon dies. To effectively distinguish between different dies, PUF should satisfy a set of quality metrics. FPGA-based PUFs are susceptible to parameters such as systematic variation and placement and routing. In this work, we conduct a statistical analysis to quantify the effect of physical layout on the randomness of multiple copies of the same PUF structure that deployed relatively close on the same FPGA die. As a case study, we have adopted an FPGA-based ring PUF structure known as bistable ring PUF. The results show that only 2 out of the 64 PUFs structure can show good randomness behavior. Even by considering 40%-60% randomness as an acceptable range, only 10 out of 64 PUFs can pass this criterion. Moreover, most of the rest PUFs are extremely biased toward a single state. As a result, 84.4% of PUFs under test end up non-functional PUFs due to the physical design of the FPGA.
Sayed Elgendy, Eslam Yahya Tawfik
ISCAS2
2019 Hardware Obfuscation of AES through Finite Field Construction Variation
abstract
To protect intellectual property, hardware obfuscation is necessary to conceal the implemented function. Besides logic-level approaches, hardware obfuscation can be done through algorithmic modifications. Prior algorithmic obfuscations address signal processing systems and those with variable data flow. This paper focuses on the obfuscation of systems based on finite field arithmetic, which are broadly adopted in digital communications. Netlists of hardware units with different field constructions are first analyzed to evaluate possible attacks. Taking into account the specifics of the computations in the Advanced Encryption Standard (AES) algorithm, optimized schemes are proposed to efficiently introduce obfuscation keys utilizing the variation of finite field construction. For an example pipelined fully-unrolled AES encryptor, the proposed scheme leads to 480 bits of obfuscation key with 3% area overhead without sacrificing the throughput. The proposed obfuscation method can be also extended to other algorithms involving finite field arithmetic.
Xinmiao Zhang 0001, Phillip Shvartsman, Eslam Yahya Tawfik
ISCAS4