Doaa Ashmawy

dblp:66/7548 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
4since 2021 · last 2026
0000-0003-1506-7534ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 2 since 2021Theory of computation · 3 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 From Compact to Fast: Exploring 32 and 64-Bit AES Datapaths for IoT Systems
Doaa Ashmawy, Arash Reyhani-Masoleh
WAIFI1
2026 Efficient Hardware Architectures for AES-128, AES-192, and AES-256 Encryption
abstract
This paper presents efficient hardware architectures for the Advanced Encryption Standard (AES) cipher supporting AES-128, AES-192, and AES-256 with On-The-Fly (OTF) key expansion units. The new designs feature low-area, low-latency OTF key expansion units that generate round keys concurrently with the data path, eliminating the need for key storage and reducing hardware overhead. By sharing logic within the key expansion circuitry, the proposed design achieves a 46% reduction in XOR gate count compared to previous works.To optimize the architecture, we analyzed eight state-of-the-art AES S-box designs and evaluated their ASIC implementation results. Four S-boxes were selected for integration, including the most compact and the fastest designs reported in the literature. ASIC synthesis results of the AES-128 architecture shows superior performance, with the fastest design achieving 29%-79% lower Area-Delay-Power Product (ADPP) than prior designs when synthesized using the same standard cell library to ensure a fair comparison.To support AES-192 and AES-256, we scaled the architecture by designing corresponding OTF key expansion units, while keeping the proposed AES-128 data path unchanged. The number of rounds was increased to 12 and 14, respectively. ASIC synthesis shows a proportional increase in ADPP for larger key sizes. Despite the higher hardware cost and reduced throughput for larger key sizes, the Fast architecture consistently delivers the highest performance and lowest ADPP, while the compact architecture minimizes area and power consumption. These results demonstrate flexible design trade-offs suitable for both high-performance and resource-constrained cryptographic applications.
Doaa Ashmawy, Arash Reyhani-Masoleh
IEEE Trans. Computers1
2025 A New 16-Bit IoT ASIC Design for the AES Encryption Algorithm
abstract
Previous works to secure IoT devices have mainly focused on 8-bit hardware architectures for AES encryption. In this paper, we present a new 16-bit ASIC design for AES encryption optimized for IoT systems. Our design includes a new 16-bit key derivation circuit that generates keys dynamically in parallel with the datapath, enhancing security by avoiding key exposure and protecting against existing attacks. Our design employs column-wise byte ordering for both the datapath and key derivation, eliminating the need for external reordering and reducing hardware resource usage. Additionally, we design a lightweight 16-bit serial MixColumns circuit that supports higher data rates compared to existing designs. ASIC implementation results using a 65nm CMOS technology library demonstrate a 50% increase in throughput with a 21% increase in area over previous 8-bit based designs. Our lightweight and fast AES ASIC design offers a tailored solution for securing IoT systems.
Doaa Ashmawy, Arash Reyhani-Masoleh
ISCAS1
2021 A Faster Hardware Implementation of the AES S-box
abstract
In this paper, we propose a very fast, yet compact, AES S-box, by applying two techniques to a composite field$GF((2^{4})^{2})$fast AES S-box. The composite field fast S-box has three main components, namely the input transformation matrix, the inversion circuit, and the output transformation matrix. The core inversion circuit computes the multiplicative inverse over the composite field$GF((2^{4})^{2})$and consists of three arithmetic blocks over subfield$GF(2^{{4}})$, namely exponentiation, subfield inverter, and output multipliers. For the first technique, we consider multiplication of the input of the composite field fast S-box by 255 nonzero 8-bit binary field elements. The multiplication constant increases the variety of the input and output transformation matrices of the S-box by a factor of 255, hence increasing the search space of the logic minimization algorithm correspondingly. For the second technique, we reduce the delay of the composite field fast S-box, by combining the output multipliers and the output transformation matrix. Moreover, we modify the architecture of the input transformation matrix and re-design the exponentiation block and the subfield inverter for lower delay and area. We find that 8 unique binary transformation matrices could be used to change from the binary field$GF(2^{8})$to the composite field$GF(({2}^{{4}})^{2})$at the input of the composite field S-box. We use Matla$\mathbf{b}$® to derive all$(255\times 8=2040)$new input transformation matrices. We search the matrices for the fastest and lowest complexity implementation and the minimal one is selected for the proposed fast S-box. The proposed fast S-box is 24% faster (with 5% increase in area) than the composite field fast design and 10% faster (with about 1% increase in area) than the fastest S-box available in the literature, to the best of our knowledge.
Doaa Ashmawy, Arash Reyhani-Masoleh
ARITH1
2020 New Low-Area Designs for the AES Forward, Inverse and Combined S-Boxes
abstract
The implementation of AES S-boxes is one of the most extensively studied areas of cryptography. In this paper, we propose three new hardware designs for the AES S-box that can serve in the forward, inverse and combined data paths. Each of these designs represents the smallest AES S-box ever proposed in its respective category. We achieve this goal by using new tower field representation over normal bases and optimizing each and every block inside the three proposed architectures. Our complexity analysis and ASIC synthesis results in the CMOS STM 65 nm, as well as the NanGate 15 nm technologies, show that our designs outperform their counterparts in terms of area and power.
Arash Reyhani-Masoleh, Mostafa M. I. Taha, Doaa Ashmawy
IEEE Trans. Computers3
2018 New Area Record for the AES Combined S-Box/Inverse S-Box
abstract
The AES combined S-box/inverse S-box is a single construction that is shared between the encryption and decryption data paths of the AES. The currently most compact implementation of the AES combined S-box/inverse S-box is Canright's design, introduced back in 2005. Since then, the research community has introduced several optimizations over the S-box only, however the combined S-boxlinverse S-box received little attention. In this paper, we propose a new AES combined S-boxlinverse S-box design that is both smaller and faster than Canright's design. We achieve this goal by proposing to use new tower field and optimizing each and every block inside the combined architecture for this field. Our complexity analysis and ASIC implementation results in the CMOS STM 65nm and NanGate 15nm technologies show that our design outperforms the counterparts in terms of area and speed.
Arash Reyhani-Masoleh, Mostafa M. I. Taha, Doaa Ashmawy
ARITH3