Reza Azarderakhsh

dblp:14/7448 · DBLP profile ↗
← Back
86ranked-venue papers
14as first author
36since 2021 · last 2026
0000-0002-6921-6868ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 53 · 8 first-author · 26 since 2021Security and privacy · 22 · 4 first-author · 5 since 2021Theory of computation · 6 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Partial Recomputation Fault Detection Architecture for Multiple-Precision Montgomery Modular Multiplication
abstract
Detection of soft errors and faults are one of the most critical factors in ensuring the reliability of algorithm implementations. Multiplication, as a fundamental and computationally intensive operation, is particularly vulnerable to such errors. Given its widespread use in cryptography and coding applications, detecting these errors is crucial. For example, in hash functions, even a single-bit change in the input can completely alter the output (ideally, each bit of the output changes with a probability of 12). Montgomery multiplication as an efficient multiplication method is an integral part of numerous cryptographic applications expanding both classical and post quantum cryptography. For that reason, this paper introduces a fault detection method for the multiple-precision Montgomery modular multiplication algorithm based on partial recomputation. Through extensive simulations and implementations, we demonstrate that our approach efficiently detects both permanent and transient errors with a high success rate, while imposing modest area and time overhead on the system.
Saeed Aghapour, Kasra Ahmadi, Mehran Mozaffari Kermani, Reza Azarderakhsh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 ParallelNTT: Maximizing Performance of Forward and Inverse NTT on FPGA for ML-DSA and ML-KEM
Bardia Taghavi, Reza Azarderakhsh, Mehran Mozaffari Kermani
ACM Great Lakes Symposium on VLSI2
2025 Guest Editorial Special Issue on Emerging Hardware Security and Trust Technologies - AsianHOST 2023
abstract
If no abstract provided do not include one in the JATS XML
Xinmiao Zhang 0001, Chongyan Gu, Mengmei Ye, Reza Azarderakhsh, Weiqiang Liu 0001
IEEE Trans. Circuits Syst. I Regul. Pap.5
2025 PUF-Dilithium: Design of a PUF-Based Dilithium Architecture Benchmarked on ARM Processors
abstract
Addressing the looming threat posed by quantum computers capable of breaching current public key cryptography schemes has become imperative. To this end, the National Institute of Standards and Technology (NIST) initiated a competition in Post-Quantum Cryptography, resulting in the selection of four schemes as the new standardized replacements, while a fourth round and an additional signature round is still ongoing. Notably, CRYSTALS-Dilithium, a lattice-based signature scheme, has exhibited promising resilience due to its efficiency and simplicity. Despite the finalization of standardization for these new four schemes, transitioning from classical cryptography to these alternatives necessitates further investigation and analysis. Comprehensive scrutiny of these newly standardized schemes is imperative, including considerations of implementation efficiency across various platforms and side-channel vulnerability analysis. This article introduces a novel design leveraging physical unclonable functions to bolster the physical security of CRYSTALS-Dilithium. Physical security is paramount in scenarios where network nodes are exposed to public scrutiny, potentially making them targets for adversaries. After discussing the advantages of our design compared to the original design, we implemented it on two different architectures, ARMv7 and ARMv8. Our results indicate substantial improvements in both security and performance compared to existing references. Moreover, noting the new competition initiated by the NIST in 2023 for new signatures (first round finalized in October 2024), potentially the proposed schemes can be adopted to the new standards set to be finalized in the coming years. These make our scheme not solely confined to the current standards and would be an important merit of the presented approaches.
Saeed Aghapour, Kasra Ahmadi, Mila Anastasova, Reza Azarderakhsh, Mehran Mozaffari Kermani
ACM Trans. Embed. Comput. Syst.4
2025 Efficient Algorithm-Level Error Detection for Number-Theoretic Transform Used for Kyber Assessed on FPGAs and ARM
abstract
Polynomial multiplication stands out as a highly demanding arithmetic process in the development of post-quantum cryptosystems. The importance of the number-theoretic transform (NTT) extends beyond post-quantum cryptosystems, proving valuable in enhancing existing security protocols such as digital signature schemes and hash functions. CRYSTALS-KYBER stands out as the sole public key encryption (PKE) algorithm chosen by the National Institute of Standards and Technology (NIST) in its third round selection, making it highly regarded as a leading post-quantum cryptography (PQC) solution. Faults have the potential to disrupt cryptographic systems, compromise data integrity, and enable side-channel attacks, making the incorporation of robust error detection mechanisms essential. This article introduces algorithm-level fault detection schemes in the NTT multiplication using Negative Wrapped Convolution ( NWC ) and the NTT tailored for Kyber Round 3, representing a significant enhancement compared with previous research. We evaluate this through the simulation of a fault model, ensuring that the conducted assessments accurately mirror the obtained results. Our fault detection scheme is designed to address both malicious fault injection attacks on Kyber and naturally occurring faults. Furthermore, we assessed the effectiveness of the proposed error detection scheme for the NTT implemented in both NWC and Kyber , using AMD/Xilinx Artix-7 FPGA, HLS and processor-based approaches. In our FPGA implementation of NWC , the integration of our error detection approach achieves near-100% fault coverage with minimal area overhead and results in only a 12% increase in latency compared with the original hardware design. Finally, we attained an error detection ratio of nearly 100% for the NTT operation in Kyber , with a clock cycle overhead of 16% on the Cortex-A72 processor.
Kasra Ahmadi, Saeed Aghapour, Mehran Mozaffari Kermani, Reza Azarderakhsh
ACM Trans. Embed. Comput. Syst.4
2025 Efficient Partial Recomputation-Based Fault Detection Approaches for Z-transform
abstract
The Z-transform is a fundamental and strong tool being widely utilized in signal processing and various other applications such as communications and networking. By analyzing the Z-transform of a signal, one can extract critical information about its stability, causality, frequency response, energy and power, and overall behavior of the signal. However, errors caused either by environmental changes or malicious injections in large-scale integration (VLSI) implementations can critically compromise the integrity and reliability of its output. Failure to detect such faults may result in unpredictable, erroneous, and misleading function analyses. Therefore, the ability to detect soft errors and faults before accepting the results is of paramount importance. In this article, we propose an efficient fault detection method that combines algorithmic-level checks with partial recomputation to identify both transient and permanent faults with a high error coverage rate across various injection scenarios. The AMD/Xilinx field-programmable gate array (FPGA) implementation of our design demonstrated only a modest increase in time and area overhead. To the best of our knowledge, fault detection for the Z-transform function has not been previously studied.
Saeed Aghapour, Kasra Ahmadi, Mehran Mozaffari Kermani, Reza Azarderakhsh
IEEE Trans. Very Large Scale Integr. Syst.4
2024 Integrating Post-Quantum TLS into the Control Plane of 5G Networks
abstract
Significant performance improvements in bandwidth and latency make 5G a suitable candidate for a wide range of applications, particularly those requiring real-time communication, such as Industrial Control Systems (ICS) and autonomous vehicles. However, today’s security, including modern cryptographic systems, is prone to different attacks caused by the high computational power of quantum computing, highlighting the need for integrating quantum-resistant security measures. To accommodate attacks targeted at 5G networks, there are efforts to move towards TLS-based security, which is the widely accepted standard across networks. However, integrating post-quantum algorithms must also be considered in such a transition. Thus, this paper is the first to perform the integration of Post-quantum TLS (PQ-TLS) protocols into 5G networks and offer a realistic performance evaluation. Our approach focuses on integrating PQ-TLS into the 5G control plane (CP) without requiring a major overhaul, thus ensuring communications’ interoperability even with legacy components of 5G, which may not support TLS. Specifically, we have updated the registration and authentication protocols for both core network functions and user equipment (UE) by implementing a TLS tunneling approach through virtualization. We then evaluate the performance and feasibility of PQ-TLS in enhancing the security of 5G communications on an actual testbed. Our results demonstrate that while PQ algorithms introduce some overhead, they remain viable for 5G applications, particularly for protocols that can run on the core network.
Yacoub Hanna, Diana Pineda, Maryna Veksler, Manish Paudel, Kemal Akkaya, Mila Anastasova, Reza Azarderakhsh
IPCCC7
2024 PUF-Kyber: Design of a PUF-Based Kyber Architecture Benchmarked on Diverse ARM Processors
abstract
It is well-studied that quantum computing breaks the security of the current worldwide implemented public key cryptosystems. This forces us toward post quantum cryptography (PQC) whose security remains solid even against adversaries having access to quantum computers. For this matter, National Institute of Standards and Technology (NIST) announced four winners in 2022. Among them, CRYSTALS-Kyber which is the only KEM/PKE algorithm, is the aim of this paper. In this paper, through using physical unclonable functions (PUF) and true random number generators (TRNG), we improve the overall security of Kyber and provide physical security to it. Our implementation results on ARMv7 and ARMv8 architectures, indicate significant speedup, compared to the reference work. For example, for the CCA.KEM-KeyGen() algorithm, we achieved roughly 26%, 13%, and 10% speedup at security levels of 512, 768, and 1024 on ARMv7 implementation, and 25%, 12%, and 10% for ARMv8 implementation. Comparing the implementation results of our design with the reference work indicates that both the security and the system performance are improved.
Saeed Aghapour, Kasra Ahmadi, Mila Anastasova, Mehran Mozaffari Kermani, Reza Azarderakhsh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2024 Hardware Constructions for Error Detection in WG-29 Stream Cipher Benchmarked on FPGA
abstract
WG-29 is a Welch-Gong (WG) stream cipher, implemented in$GF (2^{29})$and an 11-stage LFSR, whose polynomial-basis (PB)-based architecture is utilized in diverse applications. This work, for the first time, presents low-cost normal signature, interleaved signature, and Hamming code-based error detection mechanisms for the hardware implementations of PB-based WG-29 stream cipher. The presented schemes are benchmarked on field-programmable gate array (FPGA) hardware platform using Kintex-7 and Spartan-7 FPGA families for area$(< 40\%)$, power$(< 12\%)$, and delay$(< 10\%)$overheads. Using a faulty module to inject stuck-at single bit and multiple bit upsets, the error coverage for these presented schemes is evaluated via simulations performed in Xilinx Vivado for 80 000 faults and shown to be over 99.99%. The overhead and error simulation results for the presented schemes show that they provide high-error coverage with acceptable overheads to make hardware constructions of WG-29 more reliable. Other WG ciphers that have similar underlying primitives can also benefit from the presented work, with slight modifications, for secure hardware implementations.
Jasmin Kaur, Alvaro Cintas Canto, Mehran Mozaffari Kermani, Reza Azarderakhsh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 Efficient and Side-Channel Resistant Ed25519 on ARM Cortex-M4
abstract
An estimated 14.7 billion Internet of Things (IoT) devices will be connected to the Internet by 2023. The ubiquity of these devices creates exciting new opportunities, while at the same time introducing new concerns about privacy and security. To address these concerns, efficient cryptographic algorithms are needed to secure communication between IoT devices. In this work, we present an optimized implementation of one such algorithm, the Edwards Curve Digital Signature Algorithm (EdDSA) with operations Keygen, Sign, and Verify using the Ed25519 parameter on the ARM Cortex-M4 implemented in assembly code. The ARM Cortex-M4 is used in millions of devices world-wide, and is a popular choice for a wide range of IoT applications. We discuss the optimization of field and group arithmetic on this platform to produce high-throughput cryptographic primitives. Then, we present the first SCA-resistant implementation of the Signed Comb method, and Test Vector Leakage Assessment (TVLA) measurements. Our fastest implementation performs Ed25519 Keygen in 200,000 cycles, Sign in 240,000 cycles, and Verify in 720,000 cycles on the ARM Cortex-M4.
Daniel Owens, Rabih El Khatib, Mojtaba Bisheh-Niasar, Reza Azarderakhsh, Mehran Mozaffari Kermani
IEEE Trans. Circuits Syst. I Regul. Pap.4
2024 Cryptographic Engineering a Fast and Efficient SIKE in FPGA
abstract
Recent attacks have shown that SIKE is not secure and should not be used in its current state. However, this work was completed before these attacks were discovered and might be beneficial to other cryptosystems such as SQISign. The primary downside of SIKE is its performance. However, this work achieves new SIKE speed records even using less resources than the state-of-the-art. Our approach entails designing and optimizing a new field multiplier, SIKE-optimized Keccak unit, and high-level controller. On a Xilinx Virtex-7 FPGA, this architecture performs the NIST Level 1 SIKE scheme key encapsulation and key decapsulation functions in 2.23 and 2.39 ms, respectively. The combined key encapsulation and decapsulation time is 4.62 ms, which outperforms the next best Virtex-7 implementation by nearly 2 ms. Our implementation achieves speed records for the NIST Level 1, 2, and 3 parameter sets. Only our NIST Level 5 parameter set was beat by an all-out performance implementation. Our implementations also efficiently utilize the FPGA resources, achieving new records in area-time product metrics for all parameter sets. Overall, this work continues to push the bar for accelerating SIKE computations to make a stronger case for SIKE standardization.
Rami El Khatib, Brian Koziel, Reza Azarderakhsh, Mehran Mozaffari Kermani
ACM Trans. Embed. Comput. Syst.3
2024 Efficient Error Detection Schemes for ECSM Window Method Benchmarked on FPGAs
abstract
Elliptic curve scalar multiplication (ECSM) stands as a crucial subblock in elliptic curve cryptography (ECC), which represents the most widely used prequantum public key cryptography. Hardware constructions of cryptographic systems utilizing ECSM have been subject to permanent or transient errors. In cryptographic systems, it is important to validate the correctness of the underlying computation performed on hardware or software to identify such errors. In this article, we present new fault detection schemes in window method scalar multiplication, which, to the best of our knowledge, has not been previously investigated. Our approach involves introducing refined algorithms and implementations that can effectively counter both permanent and transient errors. We assess this by simulating a fault model, ensuring that the evaluations conducted reflect the obtained results. As a result, we achieve a significantly extensive coverage of errors. Finally, we benchmark our proposed error detection scheme on ARMv8 and field-programmable gate array (FPGA) to demonstrate the implementation and resource overhead. On Cortex-A72 processors, we maintain a clock cycle overhead of under 3%. In addition, when implementing our error detection method on different FPGAs, including Zynq Ultrascale+, Artix-7, and Kintex Ultrascale+, we achieve comparable throughput while introducing a mere 2% increase in area compared with the original hardware implementations.
Kasra Ahmadi, Saeed Aghapour, Mehran Mozaffari Kermani, Reza Azarderakhsh
IEEE Trans. Very Large Scale Integr. Syst.4
2024 Efficient Error Detection Cryptographic Architectures Benchmarked on FPGAs for Montgomery Ladder
abstract
Elliptic curve scalar multiplication (ECSM) is a fundamental element of public key cryptography. The ECSM implementations on deeply embedded architectures and Internet-of-nano-Things have been vulnerable to both permanent and transient errors, as well as fault attacks. Consequently, error detection is crucial. In this work, we present a novel algorithm-level error detection scheme on Montgomery Ladder often used for a number of elliptic curves featuring highly efficient point arithmetic, known as Montgomery curves. Our error detection simulations achieve high error coverage on loop abort and scalar bit flipping fault model using binary tree data structure. Assuming n is the size of the private key, the overhead of our error detection scheme is$O(n)$. Finally, we conduct a benchmark of our proposed error detection scheme on both ARMv8 and field-programmable gate array (FPGA) platforms to illustrate the implementation and resource utilization. Deployed on Cortex-A72 processors, our proposed error detection scheme maintains a clock cycle overhead of less than 5.2%. In addition, integrating our error detection approach into FPGAs, including AMD/Xilinx Zynq Ultrascale+ and Artix Ultrascale+, results in a comparable throughput and less than 2% increase in area compared with the original hardware implementation. We note that we envision using adoptions of the proposed architectures in the postquantum cryptography (PQC) based on elliptic curves.
Kasra Ahmadi, Saeed Aghapour, Mehran Mozaffari Kermani, Reza Azarderakhsh
IEEE Trans. Very Large Scale Integr. Syst.4
2023 A Flexible Shared Hardware Accelerator for NIST-Recommended Algorithms CRYSTALS-Kyber and CRYSTALS-Dilithium with SCA Protection
Luke Beckwith, Abubakr Abdulgadir, Reza Azarderakhsh
CT-RSA3
2023 Reliable Constructions for the Key Generator of Code-based Post-quantum Cryptosystems on FPGA
abstract
Advances in quantum computing have urged the need for cryptographic algorithms that are low-power, low-energy, and secure against attacks that can be potentially enabled. For this post-quantum age, different solutions have been studied. Code-based cryptography is one feasible solution whose hardware architectures have become the focus of research in the NIST standardization process and has been advanced to the final round (to be concluded by 2022–2024). Nevertheless, although these constructions, e.g., McEliece and Niederreiter public key cryptography, have strong error correction properties, previous studies have proved the vulnerability of their hardware implementations against faults product of the environment and intentional faults, i.e., differential fault analysis. It is previously shown that depending on the codes used, i.e., classical or reduced (using either quasi-dyadic Goppa codes or quasi-cyclic alternant codes), flaws in error detection could be observed. In this work, efficient fault detection constructions are proposed for the first time to account for such shortcomings. Such schemes are based on regular parity, interleaved parity, and two different cyclic redundancy checks (CRC), i.e., CRC-2 and CRC-8. Without losing the generality, we experiment on the McEliece variant, noting that the presented schemes can be used for other code-based cryptosystems. We perform error detection capability assessments and implementations on field-programmable gate array Kintex-7 device xc7k70tfbv676-1 to verify the practicality of the presented approaches. To demonstrate the appropriateness for constrained embedded systems, the performance degradation and overheads of the presented schemes are assessed.
Alvaro Cintas Canto, Mehran Mozaffari Kermani, Reza Azarderakhsh
ACM J. Emerg. Technol. Comput. Syst.3
2023 Error Detection Architectures for Hardware/Software Co-Design Approaches of Number-Theoretic Transform
abstract
Number-theoretic transform (NTT) is an efficient polynomial multiplication technique of lattice-based post-quantum cryptography including Kyber which is standardized as the NIST key encapsulation mechanism (KEM) in 2022. Prominent NTT architectures have recently been implemented on hardware/software coprocessors. In this article, we introduce new error detection schemes embedded efficiently in the NTT accelerator architecture, detecting both transient and permanent faults. By encoding the operands with two approaches, i.e., negating and swapping, we detect the faults in such constructions after recomputing and decoding. Through simulation, our schemes show high error coverage for the stuck-at fault model. Moreover, we implement the schemes on field-programmable gate array (FPGA) and assure that acceptable overhead is achieved for performance and implementation metrics. The low overhead and high efficiency of our schemes make them suitable for various constrained usage models. Additionally, our schemes are also applicable to similar classical and post-quantum sub-blocks to obtain more reliable respective hardware constructions.
Ausmita Sarker, Alvaro Cintas Canto, Mehran Mozaffari Kermani, Reza Azarderakhsh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2023 Error Detection Constructions for ITA Finite Field Inversions Over $\text{GF}(2^{m})$ on FPGA Using CRC and Hamming Codes
abstract
Finite field arithmetic operations over$\text{GF}(2^{m})$are widely used in critical applications, such as cryptography, coding theory, error-correcting codes, and digital signal processing. Finite field inversions are the most time-consuming operations among other widely-used ones and require a large footprint as well as power/energy to be performed. To reduce such complexity, the Itoh–Tsujii algorithm (ITA) has received prominent attention in the literature; however, implementations using ITA are still vulnerable to natural very-large-scale integration defects. To overcome the challenge of detecting naturally-induced faults in such constructions, for the first time, we propose error detection schemes based on Hamming codes for architectures performing finite field inversions using the ITA algorithm over$\text{GF}(2^{m})$with polynomial basis. Additionally, CRC-oriented error detection schemes for inversions in$\text{GF}(2^{m})$with normal basis are also studied and new approaches to protect them are presented. In this article, general formulations are provided along with different case studies to show the feasibility of our schemes with any finite field size. Moreover, field-programmable gate array (FPGA) implementations are performed on two Xilinx FPGA families, i.e., Kintex UltraScale+ and Xilinx Virtex-7 UltraScale+, to verify that the overheads added by the error detection architectures to provide reliability are suitable for deeply-constrained embedded systems.
Alvaro Cintas Canto, Mehran Mozaffari Kermani, Reza Azarderakhsh
IEEE Trans. Reliab.3
2023 Reliable Architectures for Finite Field Multipliers Using Cyclic Codes on FPGA Utilized in Classic and Post-Quantum Cryptography
abstract
Fault detection is becoming greatly important in protecting cryptographic designs that can suffer from both natural or malicious faults. Finite fields over$\text {GF}(2^{m})$are widely used in such designs, since their data are coded in binary form for practical reasons. Among the different finite field arithmetic, multiplication is the bottleneck operation for many cryptosystems due to its complexity. Therefore, in this work, fault detection schemes based on cyclic codes for finite field multipliers using different fields found in traditional and post-quantum cryptography are derived. Moreover, we implement such schemes by embedding them into the original architectures to perform an exhaustive study, benchmark the different overheads obtained, and prove their suitability for deeply constrained embedded systems. These implementations are performed on advanced micro devices (AMD)/Xilinx field-programmable gate array (FPGA) and provide a very high error coverage with acceptable overhead.
Alvaro Cintas Canto, Mehran Mozaffari Kermani, Reza Azarderakhsh
IEEE Trans. Very Large Scale Integr. Syst.3
2022 Faster Isogenies for Post-quantum Cryptography: SIKE
Rami El Khatib, Brian Koziel, Reza Azarderakhsh
CT-RSA3
2022 High-Performance FPGA Accelerator for SIKE
abstract
New primes were proposed for Supersingular Isogeny Key Encapsulation (SIKE) in NIST standardization process of Round 2 after further cryptanalysis research showed that the security levels of the initial primes chosen were over-estimated [#AdjCCMR18, #SIKESubmission2]. In this paper, we develop a highly optimized \mathbb{F}_{p} Montgomery multiplication algorithm and architecture that further utilizes the special form of SIKE prime compared to previous implementations available in the literature. We furthermore improve the scheduling for our program Rom. As SIKE Stays as an alternate candidate in Round 3 of standardization it has high potential to be standardized in Round 4. Therefore, we implement SIKE for all Round 2 NIST security levels (SIKEp434 for NIST security level 1, SIKEp503 for NIST security level 2, SIKEp610 for NIST security level 3, and SIKEp751 for NIST security level 5) on Xilinx Artix 7 and Xilinx Virtex 7 using the proposed multiplier. Our best implementation (NIST security level 1) runs 38% faster and occupies 30% less hardware resources in comparison to the leading counterpart available in the literature [#brian-2019] and implementations for other security levels achieved similar improvement.
Rami El Khatib, Reza Azarderakhsh, Mehran Mozaffari Kermani
IEEE Trans. Computers2
2022 Accelerated RISC-V for Post-Quantum SIKE
abstract
In this work, we present a fast and area-efficient software-hardware implementation of the supersingular isogeny key encapsulation (SIKE) mechanism. Our software-hardware design achieves both the flexibility of software as well as the efficient performance of intense computations of hardware. In particular, our implementation takes advantage of new and highly optimized hardware modules for addition, multiplication, and hardware-software control, targeted at Xilinx FPGAs. In conjunction with a small RISC-V processor, we can support all four SIKE parameter sets. On a Virtex-7 FPGA, this implementation occupies 3,492 slices, 78 DSPs, and 29 BRAMs, to perform encapsulation and decapsulation over SIKEp434, SIKEp503, SIKEp610, and SIKEp751 in 14.5, 19.2, 29.8, and 42.7 ms, respectively. Despite supporting all four parameter sets, this design has the best area-time product of all isogeny accelerators in the literature.
Rami El Khatib, Brian Koziel, Reza Azarderakhsh, Mehran Mozaffari Kermani
IEEE Trans. Circuits Syst. I Regul. Pap.3
2022 Efficient Error Detection Architectures for Postquantum Signature Falcon's Sampler and KEM SABER
abstract
Among the National Institute for Standards and Technology (NIST) postquantum cryptography (PQC) standardization Round 3 finalists (announced in 2020 and anticipated to conclude in 2022–2024), SABER and Falcon are efficient key encapsulation mechanism (KEM) and compact signature scheme, respectively. SABER is a simple and flexible cryptographic scheme, highly suitable for thwarting potential attacks in the postquantum era. Implementing SABER can be performed solely in hardware (HW) or on HW/software coprocessors. On the other hand, the compact key size, efficient design, and strong reliability proof in the quantum random oracle model (QROM) make Falcon a highly suitable signature algorithm for PQC. Although Falcon is crucial as a PQC signature scheme, the utilization of the Gaussian sampler makes it vulnerable to malicious attacks, e.g., fault attacks. This is the first work to present error detection schemes embedded efficiently in SABER as well as Falcon’s sampler architectures, which can detect both transient and permanent faults. Moreover, we implement HW design for the ModFalcon signature algorithm as well as the Gaussian sampler. These schemes are implemented on a formerly Xilinx field-programmable gate array (FPGA) family, for both SABER and Falcon variants, where we assess the error coverage and the performance. The proposed schemes incur low overhead (the area, delay, and power overheads being 22.59%, 19.77%, and 10.67%, respectively, in the worst case) while providing a high fault detection rate (99.9975% in the worst case scenario), making them suitable for high efficiency and compact HW implementations of constrained applications.
Ausmita Sarker, Mehran Mozaffari Kermani, Reza Azarderakhsh
IEEE Trans. Very Large Scale Integr. Syst.3
2021 Accelerated RISC-V for SIKE
abstract
Software implementations of cryptographic algorithms are slow but highly flexible and relatively easy to implement. On the other hand, hardware implementations are usually faster but provide little flexibility and require a lot of time to implement efficiently. In this paper, we develop a hybrid software-hardware implementation of the third round of Supersingular Isogeny Key Encapsulation (SIKE), a post-quantum cryptography algorithm candidate for NIST. We implement an isogeny field accelerator for the hardware and integrate it with a RISC-V processor which also acts as the main control unit for the field accelerator. The main advantage of this design is the high performance gain from the hardware implementation and the flexibility and fast development the software implementation provides. This is the first hybrid RISC-V and accelerator of SIKE. Furthermore, we provide one implementation for all NIST security levels of SIKE. Our design has the best area-time at NIST security levels 3 and 5 out of all hardware and hybrid designs provided in the literature.
Rami El Khatib, Reza Azarderakhsh, Mehran Mozaffari Kermani
ARITH2
2021 High-Speed NTT-based Polynomial Multiplication Accelerator for Post-Quantum Cryptography
abstract
This paper demonstrates an architecture for accelerating the polynomial multiplication using number theoretic transform (NTT). Kyber is one of the finalists in the third round of the NIST post-quantum cryptography standardization process. Simultaneously, the performance of NTT execution is its main challenge, requiring large memory and complex memory access pattern. In this paper, an efficient NTT architecture is presented to improve the respective computation time. We propose several optimization strategies for efficiency improvement targeting different performance requirements for various applications. Our NTT architecture, including four butterfly cores, occupies only 798 LUTs and 715 FFs on a small Artix-7 FPGA, showing more than 44% improvement compared to the best previous work. We also implement a coprocessor architecture for Kyber KEM benefiting from our high-speed NTT core to accomplish three phases of the key exchange in 9, 12, and$\mathbf{19}\mu \mathbf{s}$, respectively, operating at 200 MHz.
Mojtaba Bisheh-Niasar, Reza Azarderakhsh, Mehran Mozaffari Kermani
ARITH2
2021 Compressed SIKE Round 3 on ARM Cortex-M4
Mila Anastasova, Mojtaba Bisheh-Niasar, Reza Azarderakhsh, Mehran Mozaffari Kermani
SecureComm (2)3
2021 Hardware Deployment of Hybrid PQC: SIKE+ECDH
Reza Azarderakhsh, Rami El Khatib, Brian Koziel, Brandon Langenberg
SecureComm (2)1
2021 Kyber on ARM64: Compact Implementations of Kyber on 64-Bit ARM Cortex-A Processors
Pakize Sanal, Emrah Karagoz, Hwajeong Seo, Reza Azarderakhsh, Mehran Mozaffari Kermani
SecureComm (2)4
2021 Supersingular Isogeny Key Encapsulation (SIKE) Round 2 on ARM Cortex-M4
abstract
We present the first practical software implementation of Supersingular Isogeny Key Encapsulation (SIKE) round 2, targeting NIST's 1, 2, 3, and 5 security levels on 32-bit ARM Cortex-M4 microcontrollers. The proposed library introduces a new speed record of all SIKE Round 2 protocols with reasonable memory consumption on the low-end target platform. We achieved this record by adopting several state-of-the-art engineering techniques as well as highly-optimized hand-crafted assembly implementation of finite field arithmetic. In particular, we carefully redesign the previous optimized implementations of finite field arithmetic on the 32-bit ARM Cortex-M4 platform and propose a set of novel techniques which are explicitly suitable for SIKE primes. The benchmark result on STM32F4 Discovery board equipped with 32-bit ARM Cortex-M4 microcontrollers shows that entire key encapsulation and decapsultation over SIKEp434 take about 184 million clock cycles (i.e., 1.09 seconds @168 MHz). In contrast to the previous optimized implementation of the isogeny-based key exchange on low-end 32-bit ARM Cortex-M4, our performance evaluation shows feasibility of using SIKE mechanism on the low-end platform. In comparison to the most of the post-quantum candidates, SIKE requires an excessive number of arithmetic operations, resulting in significantly slower timings. However, its small key size makes this scheme as a promising candidate on low-end microcontrollers in the quantum era by ensuring the lower energy consumption for key transmission than other schemes.
Hwajeong Seo, Mila Anastasova, Amir Jalali, Reza Azarderakhsh
IEEE Trans. Computers4
2021 Reliable Architectures for Composite-Field-Oriented Constructions of McEliece Post-Quantum Cryptography on FPGA
abstract
Code-based cryptography based on binary Goppa codes is a promising solution for thwarting attacks based on quantum computers. The McEliece cryptosystem is a code-based public-key cryptosystem which is believed to be resistant against quantum attacks. In fact, it is successfully advanced to the second round of the post-quantum cryptography standardization competition early 2019. Due to its very large key size, different variants of binary Goppa codes have been proposed. Nevertheless, research has shown that such codes can be thwarted through the injection of faults, causing erroneous outputs. In this work, we present countermeasures for the implementation of different composite field arithmetic units used in the McEliece cryptosystem. The proposed architectures use overhead-aware and tailored signatures. We apply these error detection signatures to the McEliece cryptosystem and perform field-programmable gate array (FPGA) implementations to show the feasibility of adopting the proposed schemes. We benchmark the overhead and performance degradation of the proposed approaches and show their suitability for constrained embedded systems.
Alvaro Cintas Canto, Mehran Mozaffari Kermani, Reza Azarderakhsh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2021 Fast Strategies for the Implementation of SIKE Round 3 on ARM Cortex-M4
abstract
The Supersingular Isogeny Key Encapsulation mechanism (SIKE) is the only post-quantum key encapsulation protocol based on elliptic curves and isogeny maps between them. Despite the quantum security of the protocol, SIKE requires a greater number of clock cycles and hence does not provide competitive timing and energy consumption results. However, it is more attractive offering the smallest public key as well as ciphertext sizes, which considering the impact of the communication costs and storage of the keys could become a good fit for resource-constrained devices. In this work, we present the fastest practical implementation of SIKE, targeting the platform Cortex-M4 based on the ARMv7-M architecture. We performed our measurements on the STM32F407VG microcontroller for benchmarking the clock cycles and on Nucleo-F411RE attached to X-NUCLEO-LPM01A (Power Shield) for measuring the energy consumption of the protocol. The low-level finite field arithmetic operations play main role in determining the efficiency of SIKE. Therefore, we mainly focus on their optimization and apply them to all NIST-required security levels. Our SIKEp434 implementation for NIST security level 1 is about 22.97% faster than the counterparts appeared in Seoet al.(2020), where for the SIKEp503, SIKEp610 and SIKEp751 the speedup reaches 21.10%, 19.21% and 19.08%. Finally, we benchmark energy consumption and report optimization of up to 11.9% depending on the NIST security level implementation.
Mila Anastasova, Reza Azarderakhsh, Mehran Mozaffari Kermani
IEEE Trans. Circuits Syst. I Regul. Pap.2
2021 Instruction-Set Accelerated Implementation of CRYSTALS-Kyber
abstract
Large scale quantum computers will break classical public-key cryptography protocols by quantum algorithms such as Shor’s algorithm. Hence, designing quantum-safe cryptosystems to replace current classical algorithms is crucial. Luckily there are some post-quantum candidates that are assumed to be resistant against future attacks from quantum computers, and NIST is considering standardizing them. Among these candidates, lattice-based cryptography sounds more interesting than others due to the performance results as well as confidence in the security. There are few works in the literature evaluating the performance of lattice-based cryptography in hardware. In this paper, we focus on Cryptographic Suite for Algebraic Lattices (CRYSTALS) key exchange mechanisms known as Kyber and provide an instruction-set hardware architecture and implement on Xilinx Artix-7 FPGA for performance evaluation and testing. Our proposed architecture provides an efficient and high-performance set of components to perform polynomial sampling, number-theoretic transform (NTT), and point-wise multiplication to speed up lattice-based post-quantum cryptography (PQC). This architecture implemented on ASIC outperforms state-of-the-art implementations.
Mojtaba Bisheh-Niasar, Reza Azarderakhsh, Mehran Mozaffari Kermani
IEEE Trans. Circuits Syst. I Regul. Pap.2
2021 SIKE in 32-bit ARM Processors Based on Redundant Number System for NIST Level-II
abstract
We present an optimized implementation of the post-quantum Supersingular Isogeny Key Encapsulation (SIKE) for 32-bit ARMv7-A processors supporting NEON engine (i.e., SIMD instruction). Unlike previous SIKE implementations, finite field arithmetic is efficiently implemented in a redundant representation, which avoids carry propagation and pipeline stall. Furthermore, we adopted several state-of-the-art engineering techniques as well as hand-crafted assembly implementation for high performance. Optimized implementations are ported to Microsoft SIKE library written in “a non-redundant representation” and evaluated in high-end 32-bit ARMv7-A processors, such as ARM Cortex-A5, A7, and A15. A full key-exchange execution of SIKEp503 is performed in about 109 million cycles on ARM Cortex-A15 processors (i.e., 54.5 ms @2.0 GHz), which is about 1.58× faster than previous state-of-the-art work presented in CHES’18.
Hwajeong Seo, Pakize Sanal, Reza Azarderakhsh
ACM Trans. Embed. Comput. Syst.3
2021 Error Detection Architectures for Ring Polynomial Multiplication and Modular Reduction of Ring-LWE in $\boldsymbol{\frac{\mathbb{Z}/p\mathbb{Z}[x]}{x^{n}+1}}$ Benchmarked on ASIC
abstract
Ring learning with error (ring-LWE) within lattice-based cryptography is a promising cryptographic scheme for the post-quantum era. In this article, we explore efficient error detection approaches for implementing ring-LWE encryption. For achieving accurate operation of the ring-LWE problem and thwarting active side-channel attacks, error detection schemes need to be devised so that the induced overhead is not a burden to deeply embedded and constrained applications. This article, for the first time, investigates error detection schemes for both stages of the ring-LWE encryption operation, i.e., ring polynomial multiplication and modular reduction. Our schemes exploit recomputing with encoded operands, which successfully counter both natural faults (for the stuck-at model). We implement our schemes on an application-specific integrated circuit. As performance metrics show hardware overhead, our schemes prove to be low complexity with high error coverage. The proposed efficient architectures can be tailored and utilized for post-quantum cryptographic schemes in different usage models with diverse constraints.
Ausmita Sarker, Mehran Mozaffari Kermani, Reza Azarderakhsh
IEEE Trans. Reliab.3
2021 Reliable CRC-Based Error Detection Constructions for Finite Field Multipliers With Applications in Cryptography
abstract
Finite-field multiplication has received prominent attention in the literature with applications in cryptography and error-detecting codes. For many cryptographic algorithms, this arithmetic operation is a complex, costly, and time-consuming task that may require millions of gates. In this work, we propose efficient hardware architectures based on cyclic redundancy check (CRC) as error-detection schemes for postquantum cryptography (PQC) with case studies for the Luov cryptographic algorithm. Luov was submitted for the National Institute of Standards and Technology (NIST) PQC standardization competition and was advanced to the second round. The CRC polynomials selected are in-line with the required error-detection capabilities and with the field sizes as well. We have developed verification codes through which software implementations of the proposed schemes are performed to verify the derivations of the formulations. Additionally, hardware implementations of the original multipliers with the proposed error-detection schemes are performed over a Xilinx field-programmable gate array (FPGA), verifying that the proposed schemes achieve high error coverage with acceptable overhead.
Alvaro Cintas Canto, Mehran Mozaffari Kermani, Reza Azarderakhsh
IEEE Trans. Very Large Scale Integr. Syst.3
2021 CRC-Based Error Detection Constructions for FLT and ITA Finite Field Inversions Over GF(2m)
abstract
Binary extension finite fields GF(2m) have received prominent attention in the literature due to their application in many modern public-key cryptosystems and error-correcting codes. In particular, the inversion over GF(2m) is crucial for current and postquantum cryptographic applications. Schemes such as Fermat's little theorem (FLT) and the Itoh-Tsujii algorithm (ITA) have been studied to achieve better performance; however, this arithmetic operation is a complex, expensive, and time-consuming task that may require thousands of gates, increasing its vulnerability chance to natural defects. In this work, we propose efficient hardware architectures based on cyclic redundancy check (CRC) as error detection schemes for state-of-the-art finite field inversion over GF(2m) for a polynomial basis. To verify the derivations of the formulations, software implementations are performed. Likewise, hardware implementations of the original finite field inversions with the proposed error detection schemes are performed over Xilinx field-programmable gate array (FPGA) verifying that the proposed schemes achieve high error coverage with acceptable overhead.
Alvaro Cintas Canto, Mehran Mozaffari Kermani, Reza Azarderakhsh
IEEE Trans. Very Large Scale Integr. Syst.3
2021 Cryptographic Accelerators for Digital Signature Based on Ed25519
abstract
This article presents highly optimized implementations of the Ed25519 digital signature algorithm [Edwards curve digital signature algorithm (EdDSA)]. This algorithm significantly improves the execution time without sacrificing security, compared to exiting digital signature algorithms. Although EdDSA is employed in many widely used protocols, such as TLS and SSH, there appear to be extremely few hardware implementations that focus only on EdDSA. Hence, we propose two different field-programmable gate array (FPGA)-based EdDSA implementations, i.e., efficient and high-performance Ed25519 architectures applicable for a security level comparable to AES-128. Our proposed efficient Ed25519 scheme achieves an improvement of more than 84% compared to the best previous work by reducing the required area. It also incorporates more than 8× speedup. Furthermore, our proposed high-performance architecture shows a 21× speedup with more than 6200 digital signature algorithms per second, showing a significant improvement in terms of utilized area × time on a Xilinx Zynq-7020 FPGA. Finally, the effective side-channel countermeasures are embedded in our proposed designs, which also outperform the previous works.
Mojtaba Bisheh-Niasar, Reza Azarderakhsh, Mehran Mozaffari Kermani
IEEE Trans. Very Large Scale Integr. Syst.2
2020 How Not to Create an Isogeny-Based PAKE
Reza Azarderakhsh, David Jao, Brian Koziel, Jason T. LeGrow, Vladimir Soukharev, Oleg Taraskin
ACNS (1)1
2020 Further Optimizations of CSIDH: A Systematic Approach to Efficient Strategies, Permutations, and Bound Vectors
Aaron Hutchinson, Jason T. LeGrow, Brian Koziel, Reza Azarderakhsh
ACNS (1)4
2020 Highly Optimized Montgomery Multiplier for SIKE Primes on FPGA
abstract
New primes were proposed for Supersingular Isogeny Key Encapsulation (SIKE) in NIST standardization process of Round 2 after further cryptanalysis research showed that the security levels of the initial primes chosen were overestimated [1], [2]. In this paper, we develop a highly optimized EpMontgomery multiplication algorithm and architecture that further utilizes the special form of SIKE prime compared to previous implementations available in the literature. We then implement SIKE for all Round 2 NIST security levels (SIKEp434 for NIST security level 1, SIKEp503 for NIST security level 2, SIKEp610 for NIST security level 3, and SIKEp751 for NIST security level 5) on Xilinx Virtex 7 using the proposed multiplier. Our best implementation (NIST security level 1) runs 29% faster and occupies 30% less hardware resources in comparison to the leading counterpart available in the literature [3] and implementations for other security levels achieved similar improvement.
Rami El Khatib, Reza Azarderakhsh, Mehran Mozaffari Kermani
ARITH2
2020 Fast, Small, and Area-Time Efficient Architectures for Key-Exchange on Curve25519
abstract
This paper demonstrates fast and compact implementations of Elliptic Curve Cryptography (ECC) for efficient key agreement over Curve25519. Curve25519 has been recently adopted as a key exchange method for several applications and included in the National Institute of Standards and Technology (NIST) recommendations for public key cryptography. This paper presents three different performance level designs including lightweight, area-time efficient, and high-performance architectures. Lightweight hardware implementations are used for several Internet of Things (IoT) applications due to their resources being at premium. Our lightweight architecture utilizes 90% less resources compared to the best previous work while it is still more optimized in term of A middot; T (area×time). For efficient implementation from either time or utilized resources, our area-time efficient architecture can establish almost 7,000 key sessions per second which is 64% faster than the previous works. The area-time efficient architecture uses well scheduled interleaved multiplication combined with a reduction algorithm. Additionally, we offer a fast architecture for high performance applications based on the 4-level Karatsuba method and Carry-Compact Addition (CCA). Our high-performance architecture also outperforms previous work in terms of A middot; T. The results show 9% and 29% improvement in A middot; T and Admiddot;T (DSP_count×time), respectively. All architectures are variable-base-point implemented on the Xilinx Zynq-7020 FPGA family where performance and implementation metrics are reported and compared. Finally, various side-channel attack countermeasures are embedded in the proposed architectures.
Mojtaba Bisheh-Niasar, Rami El Khatib, Reza Azarderakhsh, Mehran Mozaffari Kermani
ARITH3
2020 Efficient Software Implementation of Ring-LWE Encryption on IoT Processors
abstract
Embedded processors have been widely used for building up Internet of Things (IoT) platforms, in which the security issue is becoming critical. This paper studies efficient techniques of lattice-based cryptography on these processors and presents the first implementation of ring-LWE encryption on ARM NEON and MSP430 architectures. For ARM NEON architecture, we propose a vectorized version of Iterative Number Theoretic Transform (NTT) for high-speed computation of polynomial multiplication on ARM NEON platforms and a 32-bit variant of SAMS2 technique for fast reduction. For MSP430 architecture, we propose an optimized SWAMS2 reduction technique, which consists of five different basic operations, including Shifting, Swapping, Addition, and two Multiplication-Subtractions. Regarding of the sampling from the discrete Gaussian distribution, we adopt Knuth-Yao sampler, accompanied with optimized methods such as Look-Up Table (LUT) and byte-scanning. Subsequently, a full-fledged implementation of Ring-LWE is presented by both taking advantage of our proposed method and previous optimization techniques re-designed for desired platforms. Our ring-LWE implementation of encryption/decryption at a classical security level of 128 bits requires only 149:4k=32:8k clock cycles on ARM NEON, and 2126:3k=244:5k clock cycles on MSP430. These results are roughly 7 times faster than the fastest ECC implementation on desired platforms with same security level.
Zhe Liu 0001, Reza Azarderakhsh, Howon Kim 0001, Hwajeong Seo
IEEE Trans. Computers2
2019 Optimized Algorithms and Architectures for Montgomery Multiplication for Post-quantum Cryptography
Rami El Khatib, Reza Azarderakhsh, Mehran Mozaffari Kermani
CANS2
2019 SIKE Round 2 Speed Record on ARM Cortex-M4
Hwajeong Seo, Amir Jalali, Reza Azarderakhsh
CANS3
2019 Improved Digital Signatures Based on Elliptic Curve Endomorphism Rings
Xiu Xu, Christopher Leonardi, Anzo Teh, David Jao, Kunpeng Wang 0001, Wei Yu 0008, Reza Azarderakhsh
ISPEC7
2019 Supersingular Isogeny Diffie-Hellman Key Exchange on 64-Bit ARM
abstract
We present an efficient implementation of the supersingular isogeny Diffie-Hellman (SIDH) key exchange protocol on 64-bit ARMv8 processors for 125and 160-bit post-quantum security levels. We analyze the use of both affine and projective SIDH formulas and provide a comprehensive analysis of both approaches based on the inversion-to-multiplication ratio. Implementation results show that regardless of security concerns, affine SIDH is competitive with the projective coordinates implementation, and even outperforms projective implementation in the final round of SIDH; however, projective SIDH shows better overall performance for the whole key exchange protocol. Notably, over larger finite fields, using optimized field multiplication leads to the much better performance of projective compared to affine formulas. We integrate our optimized software into the open quantum-safe OpenSSL library and compare our software with other available post-quantum primitives. The benchmark results on ARMv8 demonstrate speedup of up to 5X over the generic version of SIDH implementation which is available inside the OQS library for the same quantum security level. We observe that our highly-optimized implementation still suffers from a large number of operations for computing isogenies of elliptic curves. However, in terms of communication overhead, supersingular isogeny-based cryptosystem provides significantly smaller key size compared to its counterparts.
Amir Jalali, Reza Azarderakhsh, Mehran Mozaffari Kermani, David Jao
IEEE Trans. Dependable Secur. Comput.2
2019 Reliable Architecture-Oblivious Error Detection Schemes for Secure Cryptographic GCM Structures
abstract
To augment the confidentiality property provided by block ciphers with authentication, the Galois Counter Mode (GCM) has been standardized by the National Institute of Standards and Technology. The GCM is used as an add-on to 128-bit block ciphers, such as the Advanced Encryption Standard (AES), SMS4, or Camellia, to verify the integrity of data. Prior works on the error detection of the GCM either use linear codes to protect the GCM architectures or are based on AES-GCM architectures, confining the mechanisms to the AES block cipher. Although such structures are efficient, they are not only confined to specific architectures of the GCM but might also not fully take advantage of the parallel architectures of the GCM. Moreover, linear codes have been shown to be potentially ineffective with respect to biased faults. In this paper, we propose algorithm-oblivious constructions through recomputing with swapped ciphertext and additional authenticated blocks, which can be applied to the GCM architectures using different finite field multipliers in GF (2128). Such obliviousness for the proposed constructions used in the GCM gives freedom to the designers. We present the results of error simulations and application-specific integrated circuit implementations to demonstrate the utility of the presented schemes. Based on the overhead/degradation tolerance for implementation/performance metrics, one can fine-tune the proposed method to achieve more reliable architectures for the GCM.
Mehran Mozaffari Kermani, Reza Azarderakhsh
IEEE Trans. Reliab.2
2019 Hardware Constructions for Error Detection of Number-Theoretic Transform Utilized in Secure Cryptographic Architectures
abstract
Polynomial multiplication is one of the most rigorous arithmetic construction of postquantum cryptosystems. Utilizing number-theoretic transformations, the product of such multiplication can be efficiently computed in quasi-linear time O(n.lgn). Error detection schemes of number-theoretic transform (NTT) architectures are essential to ensure correct mathematical operations, improved security, and thwart active side-channel attacks mounted through faults. NTT is not only significant to post-quantum cryptosystems, but the structure is also valuable to the already existing security protocols, e.g., signature schemes, hash functions, and the like. This paper, for the first time, introduces new error detection schemes of NTT architectures, successfully detecting both permanent and transient faults. Our schemes are based on recomputing with negated, scaled, and swapped operands. We have implemented the proposed schemes on the application-specific integrated circuit (ASIC). Performance and implementation metrics on this hardware platform show acceptable hardware overhead. As our schemes provide acceptable complexity and high efficiency, they can be utilized in compact hardware implementations of constrained applications, e.g., deeply embedded architectures.
Ausmita Sarker, Mehran Mozaffari Kermani, Reza Azarderakhsh
IEEE Trans. Very Large Scale Integr. Syst.3
2018 Comparative realization of error detection schemes for implementations of mixcolumns in lightweight cryptography
abstract
In this paper, through considering lightweight cryptography, we present a comparative realization of MDS matrices used in the VLSI implementations of lightweight cryptography. We verify the MixColumn/MixNibble transformation using MDS matrices and propose reliability approaches for thwarting natural and malicious faults. We note that one other contribution of this work is to consider not only linear error detecting codes but also recomputation mechanisms as well as fault space transformation (FST) adoption for lightweight cryptographic algorithms. Our intention in this paper is to propose reliability and error detection mechanisms (through linear codes, recomputations, and FST adopted for lightweight cryptography) to consider the error detection schemes in designing beforehand taking into account such algorithmic security. We also posit that the MDS matrices applied in the MixColumn (or MixNibble) transformation of ciphers to protect ciphers against linear and differential attacks should be incorporated in the cipher design in order to reduce the overhead of the applied error detection schemes. Finally, we present a comparative implementation framework on ASIC to benchmark the VLSI hardware implementation presented in this paper.
Anita Aghaie, Mehran Mozaffari Kermani, Reza Azarderakhsh
CF3
2018 Reliable hardware architectures for efficient secure hash functions ECHO and fugue
abstract
In cryptographic engineering, extensive attention has been devoted to ameliorating the performance and security of the algorithms within. Nonetheless, in the state-of-the-art, the approaches for increasing the reliability of the efficient hash functions ECHO and Fugue have not been presented to date. We propose efficient fault detection schemes by presenting closed formulations for the predicted signatures of different transformations in these algorithms. These signatures are derived to achieve low overhead for the specific transformations and can be tailored to include byte/word-wide predicted signatures. Through simulations, we show that the proposed fault detection schemes are highly-capable of detecting natural hardware failures and are capable of deteriorating the effectiveness of malicious fault attacks. The proposed reliable hardware architectures are implemented on the application-specific integrated circuit (ASIC) platform using a 65-nm standard technology to benchmark their hardware and timing characteristics. The results of our simulations and implementations show very high error coverage with acceptable overhead for the proposed schemes.
Mehran Mozaffari Kermani, Reza Azarderakhsh, Siavash Bayat Sarmadi
CF2
2018 An Exposure Model for Supersingular Isogeny Diffie-Hellman Key Exchange
Brian Koziel, Reza Azarderakhsh, David Jao
CT-RSA2
2018 A High-Performance and Scalable Hardware Architecture for Isogeny-Based Cryptography
abstract
In this work, we present a high-performance and scalable architecture for isogeny-based cryptosystems. In particular, we use the architecture in a fast, constant-time FPGA implementation of the quantum-resistant supersingular isogeny Diffie-Hellman (SIDH) key exchange protocol. On a Virtex-7 FPGA, we show that our architecture is scalable by implementing at 83, 124, 168, and 252-bit quantum security levels. This is the first SIDH implementation at close to the 256-bit quantum security level to appear in literature. Further, our implementation completes the SIDH protocol 2 times faster than performance-optimized software implementations and 1.34 times faster than the previous best FPGA implementation, both running a similar set of formulas. Our implementation employs inversion-free projective isogeny formulas. By replicating multipliers and utilizing an efficient scheduling methodology, we can heavily parallelize quadratic extension field arithmetic and the isogeny evaluation stage of the large-degree isogeny computation. For a constant-time implementation of 124-bit quantum security SIDH on a Virtex-7 FPGA, we generate ephemeral public keys in 8.0 and 8.6 ms and generate the shared secret key in 7.1 and 7.9 ms for Alice and Bob, respectively. Finally, we show that this architecture could also be used to efficiently generate undeniable and digital signatures based on supersingular isogenies.
Brian Koziel, Reza Azarderakhsh, Mehran Mozaffari Kermani
IEEE Trans. Computers2
2018 Reliable and Fault Diagnosis Architectures for Hardware and Software-Efficient Block Cipher KLEIN Benchmarked on FPGA
abstract
Security-sensitive usage models, such as implantable and wearable medical devices are prone to algorithmic cryptographic attacks as well as malicious implementation attacks. The cryptographic algorithms in these cryptosystems, such as lightweight block ciphers utilized in constrained applications, face significant tradeoff between the high level of security and the efficiency of their implementation metrics. For thwarting fault analysis attacks, among effective variants of active implementation attacks, and also to detect natural faults, efficient fault diagnosis schemes for these lightweight block ciphers are essential. In this paper, for the first time, we propose error detection schemes for lightweight block cipher, KLEIN, to ameliorate its error resiliency with low hardware complexity. The proposed fault diagnosis architectures are for linear and nonlinear components of this cipher. We also consider the notion of fault space transformation for lightweight cryptography and present its potential complications. The implementation of the proposed schemes through variants of Xilinx field-programmable gate arrays and the error coverage assessed with fault injection simulations show the effectiveness of the proposed schemes with acceptable footprints.
Anita Aghaie, Mehran Mozaffari Kermani, Reza Azarderakhsh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2018 Reliable Inversion in GF(28) With Redundant Arithmetic for Secure Error Detection of Cryptographic Architectures
abstract
In secure cryptographic primitives, such as block ciphers, the reliability of hardware implementations needs to be closely considered because faults in the hardware implementations can potentially reduce or impact on the underlying security. In this paper, we present approaches to detect errors in hardware implementations of the inversion in GF(28). The proposed approaches are based on both nonredundant and redundant arithmetic, utilizing normal basis (nonredundant) and two redundant Galois field representations, i.e., polynomial ring representation and redundantly represented basis through tower fields. To the best of our knowledge, this is the first work focusing on the error detection architectures for redundant arithmetic-based inversion in GF(28). The presented signature-based schemes in this paper are general and can be applied to block ciphers with 8-bit S-boxes, such as Camellia, SMS4, the advanced encryption standard, and CLEFIA. We present the results of error simulations and application-specific integrated circuit implementations to demonstrate the utility of the presented schemes. Based on the specific implementation's security/reliability objectives and the overhead/degradation tolerance for implementation/performance metrics, one can fine-tune and tailor the proposed work to achieve more reliable inversions in GF(28).
Mehran Mozaffari Kermani, Amir Jalali, Reza Azarderakhsh, Jiafeng Xie, Kim-Kwang Raymond Choo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2018 Efficient and Reliable Error Detection Architectures of Hash-Counter-Hash Tweakable Enciphering Schemes
abstract
Through pseudorandom permutation, tweakable enciphering schemes (TES) constitute block cipher modes of operation which perform length-preserving computations. The state-of-the-art research has focused on different aspects of TES, including implementations on hardware [field-programmable gate array (FPGA)/ application-specific integrated circuit (ASIC)] and software (hard/soft-core microcontrollers) platforms, algorithmic security, and applicability to sensitive, security-constrained usage models. In this article, we propose efficient approaches for protecting such schemes against natural and malicious faults. Specifically, noting that intelligent attackers do not merely get confined to injecting multiple faults, one major benchmark for the proposed schemes is evaluation toward biased and burst fault models. We evaluate a variant of TES, i.e., the Hash-Counter-Hash scheme, which involves polynomial hashing as other variants are either similar or do not constitute finite field multiplication which, by far, is the most involved operation in TES. In addition, we benchmark the overhead and performance degradation on the ASIC platform. The results of our error injection simulations and ASIC implementations show the suitability of the proposed approaches for a wide range of applications including deeply embedded systems.
Mehran Mozaffari Kermani, Reza Azarderakhsh, Ausmita Sarker, Amir Jalali
ACM Trans. Embed. Comput. Syst.2
2017 Evidence-Based Trust Mechanism Using Clustering Algorithms for Distributed Storage Systems (Short Paper)
abstract
In distributed storage systems, documents are shared among multiple Cloud providers and stored within their respective storage servers. In social secret sharing-based distributed storage systems, shares of the documents are allocated according to the trustworthiness of the storage servers. This paper proposes a trust mechanism using machine learning techniques to compute evidence-based trust values. Our mechanism mitigates the effect of colluding storage servers. More precisely, it becomes possible to detect unreliable evidence and establish countermeasures in order to discourage the collusion of storage servers. Furthermore, this trust mechanism is applied to the social secret sharing protocol AS^3, showing that this new evidence-based trust mechanism enhances the protection of the stored documents.
Giulia Traverso, Carlos Garcia Cordero, Mehrdad Nojoumian, Reza Azarderakhsh, Denise Demirel, Sheikh Mahbub Habib, Johannes Buchmann 0001
PST4
2017 Post-Quantum Static-Static Key Agreement Using Multiple Protocol Instances
Reza Azarderakhsh, David Jao, Christopher Leonardi
SAC1
2017 Efficient Post-Quantum Undeniable Signature on 64-Bit ARM
Amir Jalali, Reza Azarderakhsh, Mehran Mozaffari Kermani
SAC2
2017 Side-Channel Attacks on Quantum-Resistant Supersingular Isogeny Diffie-Hellman
Brian Koziel, Reza Azarderakhsh, David Jao
SAC2
2017 Reliable Hardware Architectures for Cryptographic Block Ciphers LED and HIGHT
abstract
Cryptographic architectures provide different security properties to sensitive usage models. However, unless reliability of architectures is guaranteed, such security properties can be undermined through natural or malicious faults. In this paper, two underlying block ciphers which can be used in authenticated encryption algorithms are considered, i.e., light encryption device and high security and lightweight block ciphers. The former is of the Advanced Encryption Standard type and has been considered area-efficient, while the latter constitutes a Feistel network structure and is suitable for low-complexity and low-power embedded security applications. In this paper, we propose efficient error detection architectures including variants of recomputing with encoded operands and signature-based schemes to detect both transient and permanent faults. Authenticated encryption is applied in cryptography to provide confidentiality, integrity, and authenticity simultaneously to the message sent in a communication channel. In this paper, we show that the proposed schemes are applicable to the case study of simple lightweight CFB for providing authenticated encryption with associated data. The error simulations are performed using Xilinx Integrated Synthesis Environment tool and the results are benchmarked for the Xilinx FPGA family Virtex-7 to assess the reliability capability and efficiency of the proposed architectures.
Srivatsan Subramanian, Mehran Mozaffari Kermani, Reza Azarderakhsh, Mehrdad Nojoumian
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2017 Fast Software Implementations of Bilinear Pairings
abstract
Advancement in pairing-based protocols has had a major impact on the applicability of cryptography to the solution of more complex real-world problems. However, the computation of pairings in software still needs to be optimized for different platforms including emerging embedded systems and high-performance PCs. Few works in the literature have considered implementations of pairings on the former applications despite their growing importance in a post-PC world. In this paper, we investigate the efficient computation of the Optimal-Ate pairing over special class of pairing friendly Barreto-Naehrig curves in software at different security levels. We target both applications and perform our implementations on ARM-powered processors (with and without NEON instructions) and PC processors. We exploit state-of-the-art techniques and propose new optimizations to speed up the computation in the different levels including tower field and curve arithmetic. In particular, we extend the concept of lazy reduction to inversion in extension fields, analyze an efficient alternative for the sparse multiplication used inside the Miller’s algorithm and reduce further the cost of point/line evaluation formulas in affine and projective homogeneous coordinates. In addition, we study the efficiency of using M-type and D-type sextic twists in the pairing computation and carry out a detailed comparison between affine, Jacobian, and homogeneous coordinate systems. Our implementations on various mass-market emerging embedded devices significantly improve the state-of-the-art of pairing computation on ARM-powered devices and x86-64 PC platforms. For ARM implementations we achieved considerably faster computations in comparison to the counterparts.
Reza Azarderakhsh, Dieter Fishbein, Gurleen Grewal, Shi Hu, David Jao, Patrick Longa, Rajeev Verma
IEEE Trans. Dependable Secur. Comput.1
2017 Emerging Embedded and Cyber Physical System Security Challenges and Innovations
abstract
The papers in this special issue focus on deeply embedded systems and the challenges of cyper-physical security systems. Deeply-embedded systems (deployed in human body, with computer programs sending and receiving sensitive data and performing data mining for the decisions) are increasingly popular, but the security and privacy issues are not fully understood and studied. For example, issues relating to the confidentiality/integrity/availability/privacy of implantable and wearable medical devices, secure and private big data analytics, acquisition, and storage, privacy-preserving data mining, secure machine learning, cyber physical systems security, and security of hardware and software systems used for databases (with diverse societal contexts) are critical, and can be challenging to address due to their unique constraints and usage model. Existing systems for such computations would need to be transparently integrated into sensitive environments-the consequent size and energy constraints imposed on any security solutions are demanding. Thus, unique challenges arise due to the sensitivity of computation processing, need for security in implementations, and assurance “gaps.” This special issue is dedicated to the identification of techniques designed for embedded systems and cyber-physical systems, such as emerging cryptographic solutions applicable to extremely-constrained, sensitive infrastructures.
Kim-Kwang Raymond Choo, Mehran Mozaffari Kermani, Reza Azarderakhsh, G. Manimaran
IEEE Trans. Dependable Secur. Comput.3
2017 Lightweight Architectures for Reliable and Fault Detection Simon and Speck Cryptographic Algorithms on FPGA
abstract
The widespread use of sensitive and constrained applications necessitates lightweight (low-power and low-area) algorithms developed for constrained nano-devices. However, nearly all of such algorithms are optimized for platform-based performance and may not be useful for diverse and flexible applications. The National Security Agency (NSA) has proposed two relatively recent families of lightweight ciphers, that is, Simon and Speck, designed as efficient ciphers on both hardware and software platforms. This article proposes concurrent error detection schemes to provide reliable architectures for these two families of lightweight block ciphers. The research work on analyzing the reliability of these algorithms and providing fault diagnosis approaches has not been undertaken to date to the best of our knowledge. The main aim of the proposed reliable architectures is to provide high error coverage while maintaining acceptable area and power consumption overheads. To achieve this, we propose a variant of recomputing with encoded operands. These low-complexity schemes are suited for low-resource applications such as sensitive, constrained implantable and wearable medical devices. We perform fault simulations for the proposed architectures by developing a fault model framework. The architectures are simulated and analyzed on recent field-programmable grate array (FPGA) platforms, and it is shown that the proposed schemes provide high error coverage. The proposed low-complexity concurrent error detection schemes are a step forward toward more reliable architectures for Simon and Speck algorithms in lightweight, secure applications.
Prashant Ahir, Mehran Mozaffari Kermani, Reza Azarderakhsh
ACM Trans. Embed. Comput. Syst.3
2017 Fault Detection Architectures for Post-Quantum Cryptographic Stateless Hash-Based Secure Signatures Benchmarked on ASIC
abstract
Symmetric-key cryptography can resist the potential post-quantum attacks expected with the not-so-faraway advent of quantum computing power. Hash-based, code-based, lattice-based, and multivariate-quadratic equations are all other potential candidates, the merit of which is that they are believed to resist both classical and quantum computers, and applying “Shor’s algorithm”—the quantum-computer discrete-logarithm algorithm that breaks classical schemes—to them is infeasible. In this article, we propose, assess, and benchmark reliable constructions for stateless hash-based signatures. Such architectures are believed to be one of the prominent post-quantum schemes, offering security proofs relative to plausible properties of the hash function; however, it is well known that their confidentiality does not guarantee reliable architectures in the presence natural and malicious faults. We propose and benchmark fault diagnosis methods for this post-quantum cryptography variant through case studies for hash functions and present the simulations and implementations results (through application-specific integrated circuit evaluations) to show the applicability of the presented schemes. The proposed approaches make such hash-based constructions more reliable against natural faults and help protecting them against malicious faults and can be tailored based on the resources available and for different reliability objectives.
Mehran Mozaffari Kermani, Reza Azarderakhsh, Anita Aghaie
ACM Trans. Embed. Comput. Syst.2
2017 Fault Diagnosis Schemes for Low-Energy Block Cipher Midori Benchmarked on FPGA
abstract
Achieving secure high-performance implementations for constrained applications such as implantable and wearable medical devices are a priority in efficient block ciphers. However, security of these algorithms is not guaranteed in the presence of malicious and natural faults. Recently, a new lightweight block cipher, Midori, has been proposed that optimizes the energy consumption besides having low latency and hardware complexity. In this paper, fault diagnosis schemes for variants of Midori are proposed. To the best of the authors' knowledge, there has been no fault diagnosis scheme presented in the literature for Midori to date. The fault diagnosis schemes are provided for the nonlinear S-box layer and for the round structures with both 64-bit and 128-bit Midori symmetric key ciphers. The proposed schemes are benchmarked on a field-programmable gate array and their error coverage is assessed with fault-injection simulations. These proposed error detection architectures make the implementations of this new low-energy lightweight block cipher more reliable.
Anita Aghaie, Mehran Mozaffari Kermani, Reza Azarderakhsh
IEEE Trans. Very Large Scale Integr. Syst.3
2017 FPGA Realization of Low Register Systolic All-One-Polynomial Multipliers Over $GF(2^{m})$ and Their Applications in Trinomial Multipliers
abstract
Systolic all-one-polynomial (AOP) multipliers usually suffer from the problem of high register complexity, especially in field-programmable gate array (FPGA) platforms where the register resources are not that abundant. In this paper, we have shown that the AOP-based systolic multipliers can easily achieve low register-complexity implementations and the proposed architectures can be employed as computation cores to derive efficient implementations of systolic Montgomery multipliers based on trinomials. First, we propose a novel data broadcasting scheme in which the register complexity involved within existing AOP-based systolic multipliers is significantly reduced. We have found out that the modified AOP-based structure can be packed as a standard computation core. Next, we propose a novel Montgomery multiplication algorithm that can fully employ the proposed AOP-based computation core. The proposed Montgomery algorithm employs a novel precomputed-modular operation, and the systolic structures based on this algorithm fully inherit the advantages brought from the AOP-based core (low register complexity, low critical-path delay, and low latency) except some marginal hardware overhead brought by a precomputation unit. The proposed architectures are then implemented by Xilinx ISE 14.1 and it is shown that compared with the existing designs, the proposed designs achieve at least 61.8% and 47.6% less area-delay product and power-delay product than the best of competing designs, respectively.
Pingxiuqi Chen, Shaik Nazeem Basha, Mehran Mozaffari Kermani, Reza Azarderakhsh, Jiafeng Xie
IEEE Trans. Very Large Scale Integr. Syst.4
2016 NEON-SIDH: Efficient Implementation of Supersingular Isogeny Diffie-Hellman Key Exchange Protocol on ARM
Brian Koziel, Amir Jalali, Reza Azarderakhsh, David Jao, Mehran Mozaffari Kermani
CANS3
2016 Four ℚ on FPGA: New Hardware Speed Records for Elliptic Curve Cryptography over Large Prime Characteristic Fields
Kimmo Järvinen 0001, Andrea Miele, Reza Azarderakhsh, Patrick Longa
CHES3
2016 On Fast Calculation of Addition Chains for Isogeny-Based Cryptography
Brian Koziel, Reza Azarderakhsh, David Jao, Mehran Mozaffari Kermani
Inscrypt2
2016 Efficient error detection architectures for CORDIC through recomputing with encoded operands
abstract
Various optimized coordinate rotation digital computer (CORDIC) designs have been proposed to date. Nonetheless, in the presence of natural faults, such architectures could lead to erroneous outputs. In this paper, we propose error detection schemes for CORDIC architectures used vastly in applications such as complex number multiplication, and singular value decomposition for signal and image processing. To the best of our knowledge, this work is the first in providing reliable architectures for these variants of CORDIC. We present three variants of recomputing with encoded operands to detect both transient and permanent faults. The overheads and effectiveness of the proposed designs are benchmarked through Xilinx FPGA implementations and error simulations. The proposed approaches can be tailored based on overhead tolerance and the reliability constraints to achieve.
Mehran Mozaffari Kermani, Rajkumar Ramadoss, Reza Azarderakhsh
ISCAS3
2016 Guest Editorial: Introduction to the Special Section on Emerging Security Trends for Biomedical Computations, Devices, and Infrastructures
abstract
The four papers in this special section address new and emerging security trends for biomedial computations, devices, and infrastructures.
Mehran Mozaffari Kermani, Reza Azarderakhsh, Kui Ren 0001, Jean-Luc Beuchat
IEEE ACM Trans. Comput. Biol. Bioinform.2
2015 A Generalization of Addition Chains and Fast Inversions in Binary Fields
abstract
In this paper, we study a generalization of addition chains where$k$previous values are summed together on each step instead of only two values as in traditional addition chains. Such chains are called$k$-chains and we show that they have applications in finding efficient parallelizations in problems that are known to be difficult to parallelize. In particular, 3-chains improve computations of inversions in finite fields using hybrid-double multipliers. Recently, it was shown that this operation can be efficiently computed using a ternary algorithm but we show that 3-chains provide a significantly more efficient solution.
Kimmo Järvinen 0001, Vassil S. Dimitrov, Reza Azarderakhsh
IEEE Trans. Computers3
2015 High-Performance Two-Dimensional Finite Field Multiplication and Exponentiation for Cryptographic Applications
abstract
Finite field arithmetic operations have been traditionally used in different applications ranging from error control coding to cryptographic computations. Among these computations are normal basis multiplications and exponentiations which are utilized in efficient applications due to their advantageous characteristics and the fact that squaring (and subsequent powering by two) of elements can be obtained with no hardware complexity. In this paper, we present 2-D decomposition systolic-oriented algorithms to develop systolic structures for digit-level Gaussian normal basis multiplication and exponentiation over GF(2m). The proposed high-performance architectures are suitable for a number of applications, e.g., architectures for elliptic curve Diffie-Hellman key agreement scheme in cryptography. The results of the benchmark of efficiency, performance, and implementation metrics of such architectures through a 65-nm application-specific integrated circuit platform confirm high-performance structures for the multiplication and exponentiation architectures presented in this paper are suitable for high-speed architectures, including cryptographic applications.
Reza Azarderakhsh, Mehran Mozaffari Kermani
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2015 Secure and Efficient Architectures for Single Exponentiations in Finite Fields Suitable for High-Performance Cryptographic Applications
abstract
High performance implementation of single exponentiation in finite field is crucial for cryptographic applications such as those used in embedded systems and industrial networks. In this paper, we propose a new architecture for performing single exponentiations in binary finite fields. For the first time, we employ a digit-level hybrid-double multiplier proposed by Azarderakhsh and Reyhani-Masoleh for computing exponentiations based on square-and-multiply scheme. In our structure, the computations for squaring and multiplication are uniform and independent of the Hamming weight of the exponent; considered to have built-in resistance against simple power analysis attacks. The presented structure reduces the latency of exponentiation in binary finite field considerably and thus can be utilized in applications exhibiting high-performance computations including sensitive and constrained ones in embedded systems used in industrial setups and networks.
Reza Azarderakhsh, Mehran Mozaffari Kermani, Kimmo Järvinen 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2015 Reliable Radix-4 Complex Division for Fault-Sensitive Applications
abstract
Complex division is commonly used in various applications in signal processing and control theory including astronomy and nonlinear RF measurements. Nevertheless, unless reliability and assurance are embedded into the architectures of such structures, the sub-optimal (and thus erroneous) results could undermine the objectives of such applications. As such, in this paper, we present schemes to provide complex number division architectures based on Sweeney, Robertson, and Tocher-division with error detection mechanisms. Different error detection architectures are proposed in this paper which can be tailored based on the eventual objectives of the designs in terms of area and time requirements, among which we pinpoint carefully the schemes based on recomputing with shifted operands to be able to detect faults based on recomputations for different operands in addition to the unified parity (simplified detecting code) and hardware redundancy approach. The design also implements a minimized look up table approach which favors in error detection based designs and provides high fault coverage with relatively-low overhead. Additionally, to benchmark the effectiveness of the proposed schemes, extensive error detection assessments are performed for the proposed designs through fault simulations and field-programmable gate array (FPGA) implementations; the design is implemented on Xilinx Spartan-6 and Xilinx Virtex-6 FPGA families.
Mehran Mozaffari Kermani, Niranjan Manoharan, Reza Azarderakhsh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2015 Common Subexpression Algorithms for Space-Complexity Reduction of Gaussian Normal Basis Multiplication
abstract
The use of normal bases for representing elements in a binary field is attractive in some applications because it is easy to perform squaring operations in hardware. In such cases, the costs of implementing the multiplication operation become a primary concern. We present new algorithms for reducing the space complexity of Gaussian normal basis multipliers over binary fields GF(2m), where m is odd. Compared with previous results, our approach incurs no additional costs in time complexity, and achieves improvements in space complexity over a wide range of finite fields and digit sizes. For the binary fields specified in the NIST FIPS 186-3 elliptic curve digital signature algorithm standards document, our algorithms reduce by 16% (respectively, 27%) the number of XOR gates needed for the implementation of a digit-level parallel-input parallel-output multiplier over a 163-bit (respectively, 409 bit) binary field.
Reza Azarderakhsh, David Jao, Hao Lee
IEEE Trans. Inf. Theory1
2015 Parallel and High-Speed Computations of Elliptic Curve Cryptography Using Hybrid-Double Multipliers
abstract
High-performance and fast implementation of point multiplication is crucial for elliptic curve cryptographic systems. Recently, considerable research has investigated the implementation of point multiplication on different curves over binary extension fields. In this paper, we propose efficient and high speed architectures to implement point multiplication on binary Edwards and generalized Hessian curves. We perform a data-flow analysis and investigate maximum number of parallel multipliers to be employed to reduce the latency of point multiplication on these curves. Then, we modify the addition and doubling formulations and employ a newly proposed digit-level hybrid-double Gaussian normal basis multiplier to remove the data dependencies and hence reduce the latency of point multiplication. To the best of our knowledge, this is the first time that one employs hybrid-double multiplication technique to reduce the computation time of point multiplication. Moreover, we have implemented our proposed architectures for point multiplication on FPGA and obtained the results of timing and area. Our results indicate that the proposed scheme is one step forward to improve the performance of point multiplication on binary Edward and generalized Hessian curves.
Reza Azarderakhsh, Arash Reyhani-Masoleh
IEEE Trans. Parallel Distributed Syst.1
2015 Systolic Gaussian Normal Basis Multiplier Architectures Suitable for High-Performance Applications
abstract
Normal basis multiplication in finite fields is vastly utilized in different applications, including error control coding and the like due to its advantageous characteristics and the fact that squaring of elements can be obtained without hardware complexity. In this brief, we present decomposition algorithms to develop novel systolic structures for digit-level Gaussian normal basis multiplication over GF(2m). The proposed architectures are suitable for high-performance applications, which require fast computations in finite fields with high throughputs. We also present the results of our application-specific integrated circuit synthesis using a 65-nm standard-cell library to benchmark the effectiveness of the proposed systolic architectures. The presented architectures for multiplication can result in more efficient and high-performance VLSI systems.
Reza Azarderakhsh, Mehran Mozaffari Kermani, Siavash Bayat Sarmadi, Chiou-Yng Lee
IEEE Trans. Very Large Scale Integr. Syst.1
2015 Reliable and Error Detection Architectures of Pomaranch for False-Alarm-Sensitive Cryptographic Applications
abstract
Efficient cryptographic architectures are used extensively in sensitive smart infrastructures. Among these architectures are those based on stream ciphers for protection against eavesdropping, especially when these smart and sensitive applications provide life-saving or vital mechanisms. Nevertheless, natural defects call for protection through design for fault detection and reliability. In this paper, we present implications of fault detection cryptographic architectures (Pomaranch in the hardware profile of European Network of Excellence for Cryptology) for smart infrastructures. In addition, we present low-power architectures for its nine-to-seven uneven substitution box [tower field architectures in GF(33)]. Through error simulations, we assess resiliency against false-alarms which might not be tolerated in sensitive intelligent infrastructures as one of our contributions. We further benchmark the feasibility of the proposed approaches through application-specific integrated circuit realizations. Based on the reliability objectives, the proposed architectures are a step-forward toward reaching the desired objective metrics suitable for intelligent, emerging, and sensitive applications.
Mehran Mozaffari Kermani, Reza Azarderakhsh, Anita Aghaie
IEEE Trans. Very Large Scale Integr. Syst.2
2014 Fast Inversion in ${\schmi{GF(2^m)}}$ with Normal Basis Using Hybrid-Double Multipliers
abstract
Fast inversion in finite fields is crucial for high-performance cryptography and codes. We present techniques to exploit the recently proposed hybrid-double multipliers for fast inversions in binary fields GF(2m) with normal bases. A hybrid-double multiplier computes a double multiplication, the product of three elements in GF(2m), with a latency comparable to the latency of single multiplication of two elements. Traditional approaches, such as Itoh-Tsujii, cannot utilize hybrid-double multipliers. We devise a new inversion algorithm based on ternary representations that exploits their potential. The algorithm reduces the latency of inversion significantly for the fields recommended by NIST if hybrid-double multipliers are employed. For example, the algorithm computes an inversion in GF(2163) with only five double multiplications whereas the Itoh-Tsujii algorithm requires nine single or double multiplications. We propose a new inverter architecture using this new algorithm and a hybrid-double multiplier. We show that it is faster than the existing techniques by providing ASIC synthesis results using 65-nm CMOS technology. For example, our inverter for GF(2163) achieves about 34 percent shorter computation time than an inverter using the Itoh-Tsujii algorithm and a single multiplier.
Reza Azarderakhsh, Kimmo Järvinen 0001, Vassil S. Dimitrov
IEEE Trans. Computers1
2014 A New Double Point Multiplication Algorithm and Its Application to Binary Elliptic Curves with Endomorphisms
abstract
We present a new double point multiplication algorithm based on differential addition chains. Our proposed scheme has a uniform structure and has some degree of built-in resistance against side channel analysis attacks. We discuss deploying our scheme in a hardware implementation of single point multiplication on binary elliptic curves with efficiently computable endomorphisms. Based on operation counts, we expect to gain accelerations of 30% and 18% for computing single point multiplication with and without availability of parallel multipliers, respectively, and these results are verified in our implementations.
Reza Azarderakhsh, Koray Karabina
IEEE Trans. Computers1
2014 Low-Latency Digit-Serial Systolic Double Basis Multiplier over $\mbi GF{(2^m})$ Using Subquadratic Toeplitz Matrix-Vector Product Approach
abstract
Recently in cryptography and security, the multipliers with subquadratic space complexity for trinomials and some specific pentanomials have been proposed. For such kind of multipliers, alternatively, we use double basis multiplication which combines the polynomial basis and the modified polynomial basis to develop a new efficient digit-serial systolic multiplier. The proposed multiplier depends on trinomials and almost equally space pentanomials (AESPs), and utilizes the subquadratic Toeplitz matrix-vector product scheme to derive a low-latency digit-serial systolic architecture. If the selected digit-size is$d$bits, the proposed digit-serial multiplier for both polynomials, i.e., trinomials and AESPs, requires the latency of$2\left\lceil \!{\sqrt {{m \over d}}} \right\rceil\! $, while traditional ones take at least$O\left({\left\lceil {{m \over d}} \right\rceil} \right)$clock cycles. Analytical and application-specific integrated circuit (ASIC) synthesis results indicate that both the area and the time$ \times $area complexities of our proposed architecture are significantly lower than the existing digit-serial systolic multipliers.
Jeng-Shyang Pan 0001, Reza Azarderakhsh, Mehran Mozaffari Kermani, Chiou-Yng Lee, Wen-Yo Lee, Che Wun Chiou, Jim-Min Lin
IEEE Trans. Computers2
2014 Reliable Concurrent Error Detection Architectures for Extended Euclidean-Based Division Over GF(2m)
abstract
The extended Euclidean algorithm (EEA) is an important scheme for performing the division operation in finite fields. Many sensitive and security-constrained applications such as those using the elliptic curve cryptography for establishing key agreement schemes, augmented encryption approaches, and digital signature algorithms utilize this operation in their structures. Although much study is performed to realize the EEA in hardware efficiently, research on its reliable implementations needs to be done to achieve fault-immune reliable structures. In this regard, this paper presents a new concurrent error detection (CED) scheme to provide reliability for the aforementioned sensitive and constrained applications. Our proposed CED architecture is a step forward toward more reliable architectures for the EEA algorithm architectures. Through simulations and based on the number of parity bits used, the error detection capability of our CED architecture is derived to be 100% for single-bit errors and close to 99% for the experimented multiple-bit errors. In addition, we present the performance degradations of the proposed approach, leading to low-cost and reliable EEA architectures. The proposed reliable architectures are also suitable for constrained and fault-sensitive embedded applications utilizing the EEA hardware implementations.
Mehran Mozaffari Kermani, Reza Azarderakhsh, Chiou-Yng Lee, Siavash Bayat Sarmadi
IEEE Trans. Very Large Scale Integr. Syst.2
2013 Low-Complexity Multiplier Architectures for Single and Hybrid-Double Multiplications in Gaussian Normal Bases
abstract
The extensive rise in the number of resource constrained wireless devices and the needs for secure communications with the servers imply fast and efficient cryptographic computations for both parties. Efficient hardware implementation of arithmetic operations over finite field using Gaussian normal basis is attractive for public key cryptography as it provides free squarings. In this paper, we first present two low-complexity digit-level multiplier architectures. It is shown that the proposed multipliers outperform the existing Gaussian normal basis (GNB) multiplier structures available in the literature. Then, for the first time, using these two architectures, we propose a new digit-level hybrid multiplier which performs two successive multiplications with the same latency as the one for one multiplication. We have studied the efficiency of the proposed hybrid architecture in terms of area and time delay for different digit sizes. The main advantage of this new hybrid architecture is to speed up exponentiation and point multiplication whenever double-multiplication is required and the traditional schemes fail due to the data dependencies. We have investigated the applicability of the proposed hybrid structure to reduce the latency of exponentiation-based cryptosystems. Our analysis and timing results show that the expected acceleration in double-exponentiation is considerable. Prototypes of the presented low-complexity multiplier architectures and the proposed hybrid architecture are implemented and experimental results are presented.
Reza Azarderakhsh, Arash Reyhani-Masoleh
IEEE Trans. Computers1
2012 Efficient Implementation of Bilinear Pairings on ARM Processors
Gurleen Grewal, Reza Azarderakhsh, Patrick Longa, Shi Hu, David Jao
Selected Areas in Cryptography2
2012 Efficient FPGA Implementations of Point Multiplication on Binary Edwards and Generalized Hessian Curves Using Gaussian Normal Basis
abstract
Efficient implementation of point multiplication is crucial for elliptic curve cryptographic systems. This paper presents the implementation results of an elliptic curve crypto-processor over binary fields GF(2m) on binary Edwards and generalized Hessian curves using Gaussian normal basis (GNB). We demonstrate how parallelization in higher levels can be performed by full resource utilization of computing point addition and point-doubling formulas for both binary Edwards and generalized Hessian curves. Then, we employ the ω-coordinate differential formulations for computing point multiplication. Using a lookup-table (LUT)-based pipelined and efficient digit-level GNB multiplier, we evaluate the LUT complexity and time-area tradeoffs of the proposed crypto-processor on an FPGA. We also compare the implementation results of point multiplication on these curves with the ones on the traditional binary generic curve. To the best of the authors' knowledge, this is the first FPGA implementation of point multiplication on binary Edwards and generalized Hessian curves represented by ω-coordinates.
Reza Azarderakhsh, Arash Reyhani-Masoleh
IEEE Trans. Very Large Scale Integr. Syst.1
2010 A Modified Low Complexity Digit-Level Gaussian Normal Basis Multiplier
Reza Azarderakhsh, Arash Reyhani-Masoleh
WAIFI1