EDBT 2026 Demo / reviewers in the wild / expert
Hayssam El-Razouk
dblp:68/2680
· DBLP profile ↗
10ranked-venue papers
6as first author
4since 2021 · last 2025
0000-0002-3452-5281ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 5 first-author · 4 since 2021Theory of computation · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Low-Clock Latency SDRR-Based AESabstractRecent hardware designs of the Advanced encryption standard (AES) employ randomized Double Rate Registers (DRRs) to counter Power Analysis Attacks (PAAs). This paper introduces a low output latency hardware architecture of AES based on DRRs. The proposed design does not duplicate the data-path and reduces the clock latency by utilizing the rising and falling edges of the clock. The proposed architecture reduces encryption clock latency by 25% and achieves the best hardware efficiency of 0.87 Mbit/s × LUT compared to the existing DRR based AES counterparts when realized on Xilinx xc7a100tftg256-2 FPGA device. In addition to its superior hardware and time complexities, the proposed AES hardware architecture demonstrates similar protection to the existing DRR based AES designs when tested against Correlation power analysis attacks (CPAs) using ChipWhisperer. Mehmet Ali Cetin, Hayssam El-Razouk |
ISCAS | 2 |
| 2023 | Squeezing Area of the Versatile GF (2m) GNB Arithmetic OperatorsabstractCryptography primitives have a prominent role in securing applications that may require low-area realizations, for example portable devices and other resource constrained devices. A given system may require support for different cryptography based protocols/ primitives. Many standardized and/or published primitives rely on arithmetic operations over$GF\left ({2^{m}}\right)$that occupy major area footprint. Therefore, versatile operators have been of interest to reduce the area penalty, in particular bit-serial multipliers. This paper introduces a novel scheme for versatile multiplication by the normal element in the Gaussian Normal Basis (GNB) leading to new low-area versatile GNB multiplier and inverter architectures that are presented for the first time, as far as we know. Specifically, the proposed inverters are the first versatile GNB inversion in open literature, to the best of our knowledge. Field Programmable Gate Arrays (FPGA) implementation results demonstrate that the proposed versatile multiplication and inversion techniques save almost 30% and 46%, and for Application Specific Integrated Circuits (ASIC) implementation the savings are up-to 29% and 35% respectively, in terms of area when compared to other counterparts. Mahidhar Puligunta, Hayssam El-Razouk |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | Input-Latency Free Versatile Bit-Serial GF(2m) Polynomial Basis MultiplicationabstractCryptography and error correction codes are widely used for information security and data integrity services in modern digital computing and communications systems. A number of standardized and published cryptography, and error correction algorithms utilize arithmetic operations over$\text {GF}(2^{m})$. The polynomial basis (PB) representation of$\text {GF}(2^{m})$elements is suitable for the support of multiple$\text {GF}(2^{m})$fields (that is, versatility). This article considers the problem of reduced hardware efficiency of versatile$\text {GF}(2^{m})$PB multipliers as a result of increased loading latency of the inputs and irreducible polynomial. In this context, to the best of our knowledge, this article proposes the first input-latency free versatile bit-serial$\text {GF}(2^{m})$multiplier using PB. The superiority of the proposed versatile PB multiplier is demonstrated for the case where the multiplication’s inputs arrive serially in terms of higher output throughput and hardware efficiency compared to existing versatile PB and Gaussian normal basis (GNB) counterparts based on the application-specific integrated circuit (ASIC) and field-programmable gate array (FPGA) realizations, respectively. The proposed versatile PB multiplier improves hardware efficiency by factors of 1.8 and 2.5, respectively, compared to the versatile PB and versatile GNB counterparts. Hayssam El-Razouk |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2021 | Novel $GF\left(2^{m}\right)$GF2m Digit-Serial PISO Multipliers for the Self-Dual Gaussian Normal BasesabstractSecurity protocols such as Transport Layer Security implement the Elliptic curve digital signature algorithm (ECDSA) over different binary extension fields defined by the National Institute of Standards and Technology (NIST). Specifically, such multiple cipher-suite support is a security recommendation. Binary extension field arithmetic processors are expensive, especially if more than one field is supported. In this context, this article introduces a novel lightweight digit-serial parallel-in-serial-out (DS-PISO) design for versatile multiplication (DS-VPISO) targeting the NIST-like fields in resource constrained embedded systems where the crypto module allocation is limited. The proposed DS-VPISO multiplier offers competitive area (up to 40 percent area savings) compared to existing multiple field multiplier schemes based on results conducted on bit-serial implementations using Intel's Field programmable gate arrays. The article first presents a DS-PISO self-dual Gaussian normal basis multiplication architecture based on the trace mapping. After this, the article extends the new trace based DS-PISO multiplier to construct an architecture for the first versatile DS-PISO multiplication (DS-VPISO) targeting NIST's binary fields. The latter extension to versatile multiplication is based on novel architectures for versatile cyclic shifts and versatile multiplication by normal elements. Hayssam El-Razouk, Kirthi Kotha, Mahidhar Puligunta |
IEEE Trans. Computers | 1 |
| 2019 | New Multiplicative Inverse Architectures Using Gaussian Normal BasisabstractThe multiplicative inverse over binary fields is one of main arithmetic operations used in cryptography. This paper presents two new inversion architectures. First, an improved architecture for classic inversion scheme using single multiplier is presented. The new improved inverter achieves lower latency through loading input registers during last multiplication cycle, at the expense of higher propagation delay. After this, a novel inversion architecture which uses half the latency to process the classic-based addition chains (or improved ones) is presented. The latter architecture, named Classical-Interleaved, is constructed based on a novel fully-serial-in square-multiply processor (FSISM). The FSISM, squares one operand, and multiply it to the second one, concurrently while the two inputs are absorbed serially digit-by-digit. The new classical inverter and the new classical-interleaved inverter reduce the latency compared to other schemes. In addition, the proposed classical and interleaved inverters outperform the original Itoh-Tsujii algorithm (ITA) and Ternary Itoh-Tsujii / optimal 3-chain algorithms in terms of its higher throughput and improved hardware efficiency for a number of digit sizes. The efficiency of the proposed field inverters are demonstrated by comparisons based on application specific integrated circuits (ASIC) implementations results using the standard 65 nm CMOS technology libraries. Arash Reyhani-Masoleh, Hayssam El-Razouk, Amin Monfared |
IEEE Trans. Computers | 2 |
| 2017 | A New Multiplicative Inverse Architecture in Normal Basis Using Novel Concurrent Serial Squaring and MultiplicationabstractItoh and Tsujii proposed a fast algorithm for computing multiplicative inverses (inversions) over GF(2m) using normal bases by iterating single multiplications and cyclic shifts. Recently, the Itoh-Tsujii algorithm (ITA) has been modified to use two digit-level single multiplications. The improvements of the modified Itoh-Tsujii and its variant algorithms are based on reducing the computational latency at the expense of more area requirements. In this paper, we propose a new inversion architecture based on the classical IT algorithm (or improved one) utilizing a novel interleaved computations of two single multiplications and squarings at the digit-level. The new inverter outperforms previous modified Itoh-Tsujii algorithms (such as the Ternary Itoh-Tsujii and optimal 3-chain algorithms) in terms of its lower latency, higher throughput, and improved hardware efficiency. The efficiency of the proposed field inverter is demonstrated by comparisons based on application specific integrated circuits (ASIC) implementations results using the standard 65nm CMOS technology libraries. Amin Monfared, Hayssam El-Razouk, Arash Reyhani-Masoleh |
ARITH | 2 |
| 2016 | New Architectures for Digit-Level Single, Hybrid-Double, Hybrid-Triple Field Multiplications and Exponentiation Using Gaussian Normal BasesabstractGaussian normal bases (GNBs) are special set of normal bases (NBs) which yield low complexity$GF\left(2^{m}\right)$arithmetic operations. In this paper, we present new architectures for the digit-level single, hybrid-double, and hybrid-triple multiplication of$GF\left(2^{m}\right)$elements based on the GNB representation for odd values of$m > 1$. The proposed fully-serial-in single multipliers perform multiplication of two field elements and offer high throughput when the data-path capacity for entering inputs is limited. The proposed hybrid-double and hybrid-triple digit-level GNB multipliers perform, respectively, two and three field multiplications using the same latency required for a single digit-level multiplier, at the expense of increased area. In addition, we present a new eight-ary field exponentiation architecture which does not require precomputed or stored intermediate values. Hayssam El-Razouk, Arash Reyhani-Masoleh |
IEEE Trans. Computers | 1 |
| 2015 | New Bit-Level Serial GF (2m) Multiplication Using Polynomial BasisabstractThe Polynomial basis (PB) representation offers efficient hardware realizations of GF(2m) multipliers. Bit-level serial multiplication over GF(2m) trades-off the computational latency for lower silicon area, and hence, is favored in resource constrained applications. In such area critical applications, extra clock cycles might take place to read the inputs of the multiplication if the data-path has limited capacity. In this paper, we present a new bit-level serial PB multiplication scheme which generates its output bits in parallel after m clock cycles without requiring any preloading of the inputs, for the first time in the open literature. The proposed architecture, referred to as fully-serial-in-parallel-out (FSIPO), is useful for achieving higher throughput in resource constrained environments if the data-path for entering inputs has limited capacity, especially, for large dimensions of the field GF (2m). Hayssam El-Razouk, Arash Reyhani-Masoleh |
ARITH | 1 |
| 2015 | New Hardware Implementations of WG(29, 11) and WG-16 Stream Ciphers Using Polynomial BasisabstractThe WG stream ciphers are based on the WG (Welch-Gong) transformation and possess proved randomness properties. In this paper we propose nine new hardware designs for the two classes of WG(29,11) and WG-16. For each class, we design and implement three versions of standard, pipelined and serial. For the first time, we use the polynomial basis (PB) representation to design and implement the WG(29,11) and WG-16. We consider traditional PB multiplier for the WG(29,11), and, the traditional and Karatsuba multipliers for the WG-16. For efficient field operations, we propose an irreducible trinomial for the WG(29,11). For the WG-16, a new formulation of its permutation which requires only 8 multipliers is introduced. In these designs, the multipliers in the transforms are further reduced by utilizing a novel computation for the trace of the multiplication of two field elements. We have implemented the proposed designs in ASIC using CMOS 65 nm technology. The results show that the proposed standard WG(29,11) consumes less area and slightly enhances the normalized throughput, compared to the existing counterparts. For the WG-16, throughput of the proposed pipelined instance outperforms the previous designs. Moreover, the speed of the proposed WG-16 designs meet the peak bit rates for the 4 G specifications. Hayssam El-Razouk, Arash Reyhani-Masoleh, Guang Gong |
IEEE Trans. Computers | 1 |
| 2014 | New Implementations of the WG Stream CipherabstractThis paper presents two new hardware designs of the Welch-Gong (WG)-128 cipher, one for the multiple output WG (MOWG) version, and the other for the single output version WG based on type-II optimal normal basis representation. The proposed MOWG design uses signal reuse techniques to reduce hardware cost in the MOWG transformation, whereas it increases the speed by eliminating the inverters from the critical path. This is accomplished through reconstructing the key and initial vector loading algorithm and the feedback polynomial of the linear feedback shift register. The proposed WG design uses properties of the trace function to optimize the hardware cost in the WG transformation. The application-specific integrated circuit and field-programmable gate array implementations of the proposed designs show that their areas and power consumptions outperform the existing implementations of the WG cipher. Hayssam El-Razouk, Arash Reyhani-Masoleh, Guang Gong |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |