Dur-e-Shahwar Kundi

dblp:37/8031 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
9since 2021 · last 2023
0000-0001-5120-0887ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 first-author · 6 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2023 HPKA: A High-Performance CRYSTALS-Kyber Accelerator Exploring Efficient Pipelining
abstract
CRYSTALS-Kyber (Kyber) was recently chosen as the first quantum resistant Key Encapsulation Mechanism (KEM) scheme for standardisation, after three rounds of the National Institute of Standards and Technology (NIST) initiated PQC competition which begin in 2016 and search of the best quantum resistant KEMs and digital signatures. Kyber is based on the Module-Learning with Errors (M-LWE) class of Lattice-based Cryptography, that is known to manifest efficiently on FPGAs. This work explores several architectural optimizations and proposes a high-performance and area-time (AT) product efficient hardware accelerator for Kyber. The proposed architectural optimizations include inter-module and intra-module pipelining, that are designed and balanced via FIFO based buffering to ensure maximum parallelisation. The implementation results show that compared to state-of-the-art designs, the proposed architecture delivers 25–51% speedups for Kyber's three different security levels on Artix-7 and Zynq UltraScale+ devices, and a 50–75% reduction in DSPs at comparable security level. Consequently, the proposed design achieve higher AT product efficiencies of 19–33%.
Ziying Ni, Ayesha Khalid, Dur-e-Shahwar Kundi, Máire O'Neill, Weiqiang Liu 0001
IEEE Trans. Computers3
2022 Horizontal Correlation Analysis without Precise Location on Schoolbook Polynomial Multiplication of Lattice-based Cryptosystem
abstract
Most cryptographic systems are secure in theory; however, the implementation of cryptographic system on embedded devices can be attacked by analyzing the power consumption of specific operation to reveal the key. The classic vertical correlation power analysis (CPA) attack requires a large number of power traces for analysis. Using transient secret-key scheme significantly weakens such an attack as insufficient data could be obtained. On the other hand, the horizontal CPA requires at least a single power trace and can make full use of multiple intermediate values to analyze the correlation of power consumption. In this work, we devised a horizontal CPA attack on schoolbook polynomial multiplication of hardware-implemented lattice-based cryptosystem without precise location. The accuracy of correctly recovering any one sub secret-key using only a single trace is 99.90%, and the accuracy of correctly recovering the secret-key is 76.41%. The powerful attack capability of horizontal CPA exposes the vulnerability of unprotected schoolbook polynomial multiplication against the attack of side-channel analysis (SCA).
Chuanchao Lu, Yijun Cui, Dur-e-Shahwar Kundi, Chenghua Wang, Weiqiang Liu 0001
ISCAS4
2022 A Lightweight and Efficient Schoolbook Polynomial Multiplier for Saber
abstract
Saber is a lattice-based post-quantum cryptography (PQC) algorithm, which is still a candidate in the 3rdRound of National Institute of Standards and Technology (NIST) PQC standardization process. Saber provides a great advantage of being lightest among all the candidates, so a suitable choice for resource-constraint platforms. Polynomial multiplication occupies most of the resources in hardware implementation of Saber, which needs to be optimized for the efficient hardware implementation. In this work, a lightweight and efficient schoolbook polynomial multiplier is proposed. The architecture includes an efficient multiplication strategy that compute four coefficient-wise multiplication per cycle along with the multiplication operand loading technique being designed for the compact multiplier. The proposed multiplier on Artix-7 FPGA, achieves a frequency of 130 MHz and fits into 201 slices. Compared with the state-of-the-art lightweight schoolbook implementations for Saber, our design has a 30% improved frequency and saves 15.8% of the clock counts at the cost of only 3.7% more LUTs.
Yuantuo Zhang, Yijun Cui, Ziying Ni, Dur-e-Shahwar Kundi, Weiqiang Liu 0001
ISCAS4
2022 Stacked Ensemble Model for Enhancing the DL based SCA
abstract
Deep learning (DL) has proven to be very effective for image recognition tasks, with a large body of research on various models for object classification. The application of DL to side-channel analysis (SCA) has already shown promising results, with experimentation on open-source variable key datasets showing that secret keys for block ciphers like Advanced Encryption Standard (AES)-128 can be revealed with 40 traces even in the presence of countermeasures. This paper aims to further improve the application of DL in SCA, by enhancing the power of DL when targeting the secret key of cryptographic algorithms when protected with SCA countermeasures. We propose a stacked ensemble model, which trains the output probabilities and Maximum likelihood score of multiple traces and/or sub-models to improve the performance of Convolutional Neural Network (CNN)-based models. Our model generates state-of-the art results when attacking the ASCAD variable-key database, which has a restricted number of training traces per key, recovering the key within 20 attack traces in comparison to 40 traces as required by the state-of-the-art CNN-based model with Plaintext feature extension (CNNP)-based model. During the profiling stage an attacker needs no additional knowledge of the implementation, such as the masking scheme or random mask values, only the ability to record the power consumption or electromagnetic field traces, plaintext/ciphertext and the key is needed. However, a two step training procedure is required. Additionally, no heuristic pre-processing is required in order to break the multiple masking countermeasures of the target implementation.
Anh-Tuan Hoang, Neil Hanley, Ayesha Khalid, Dur-e-Shahwar Kundi, Máire O'Neill
SECRYPT4
2022 AxRLWE: A Multilevel Approximate Ring-LWE Co-Processor for Lightweight IoT Applications
abstract
This work presents a multilevel approximation exploration undertaken on the Ring-Learning-with-Errors (R-LWE)-based public-key cryptographic (PKC) schemes that belong to quantum-resilient cryptography algorithms. Among the various quantum-resilient cryptography schemes proposed in the currently running NIST’s post-quantum cryptography (PQC) standardization plan, the lattice-based learning-with-error (LWE) schemes have emerged as the most viable and preferred class for the Internet of Things (IoT) applications due to their compact area and memory footprint compared to other alternatives. However, compared to the classical schemes used today, R-LWE is much harder a challenge to fit on embedded IoT (end-node) devices, due to their stricter resource constraints (lower area, memory, and energy budgets) as well as their limited computational capabilities. To the best of our knowledge, this is the first endeavor exploring the inherent approximate nature of the LWE problem to undertake a multilevel approximate R-LWE (AxRLWE) architecture with respective security estimates opt for lightweight IoT devices. Undertaking AxRLWE on field-programmable gate arrays (FPGAs), we benchmarked a 64% area reduction cost compared to earlier accurate R-LWE designs at the cost of reduced quantum security. For the application-specific integrated circuits (ASICs) with 45-nm CMOS technology, AxRLWE was benchmarked to fit well within the same area budget of a lightweight ECC processor and consume a third of energy compared to special class of R-Binary LWE (R-BLWE) designs being proposed for an IoT, with a better security level.
Dur-e-Shahwar Kundi, Ayesha Khalid, Song Bian 0001, Chenghua Wang, Máire O'Neill, Weiqiang Liu 0001
IEEE Internet Things J.1
2022 A High-Performance SIKE Hardware Accelerator
abstract
Supersingular isogeny key encapsulation (SIKE) is a promising candidate in the NIST postquantum cryptography (PQC) standardization process, which has the smallest key lengths. It is the only isogeny-based cryptographic scheme in the NIST list that leverages the traditional elliptic curve cryptography (ECC) arithmetic; however, the high computational complexity is one of its limiting factors. In this work, we proposed a high-performance hardware architecture for the SIKE protocol. The architecture includes an improved multiplier based on the high-performance finite field multiplication (HFFM) algorithm which is 15%–20.7% faster than the previous multiplier based on the HFFM algorithm and a unified adder/subtractor with radix$3^{b}$. In addition, it also comprises an efficient scheduler strategy that decomposes all the functions of SIKE into finite field$F_{p}$and then effectively schedules through optimized multiplication chains for maximal performance. The proposed architecture is synthesized and implemented on Xilinx Virtex-7 FPGA for all the four variants of SIKE having security levels from 1 to 5 and achieved 2.6%–7.8% faster speeds as well as consumed less equivalent number of slices (ENS) than the state-of-the-art designs. In the comparison of area and time (AT), the proposed architecture is 14.2%–34.5% lower than the previous architecture.
Ziying Ni, Dur-e-Shahwar Kundi, Máire O'Neill, Weiqiang Liu 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2021 High-Performance Systolic Array Montgomery Multiplier for SIKE
abstract
In theory, the speed of quantum computers is much faster than classical computers, which poses a threat to the Public Key Cryptography (PKC) that are currently in use. Post Quantum Cryptography (PQC) is a class of cryptography based on complex mathematical problems that are difficult to be attacked by quantum computers. The Supersingular Isogeny Key Encapsulation (SIKE) protocol is one of candidate algorithms for the US National Institute of Standards and Technology (NIST) PQC standardization process and survived to the Round 3. In this paper, we reconstruct the systolic array based Montgomery multiplier architecture for SIKE, using a three-stage pipeline that results in frequency improvement of 21.4%. The proposed multiplier consumed fewer DSP resources than the state-of-the-art SIKE designs and has a speed increase up to 12.7%.
Ziying Ni, Dur-e-Shahwar Kundi, Máire O'Neill, Weiqiang Liu 0001
ISCAS2
2021 Towards CRYSTALS-Kyber: A M-LWE Cryptoprocessor with Area-Time Trade-Off
abstract
CRYSTALS-Kyber is a quantum-resistant and promising lattice-based cryptography (LBC) in the finalists of the third round post-quantum cryptography (PQC) standardization, which is based on the hardness of Module-Learning with Errors (M-LWE). The variadic parameters make M-LWE obtain a more flexible security-performance trade-off than Ring-LWE. In this paper, we propose a M-LWE cryptoprocessor targeting CRYSTALS-Kyber with area-time trade-off for the first time. This balanced design includes a fast and low-cost Binomial Sampler and vector-polynomials multiplication structure based on pipelined decimation-in-frequency (DIF) based Number Theoretic Transform (NTT) technique. The M-LWE cryptoprocessor achieve 27,708 encryption operations per second using only 690 slices and 106,716 decryption operations per second using only 571 slices. Our proposed design achieved the lowest area-time product (ATP) with at least 2 χ performance improvement than the state-of-the-art LBC designs with a similar security level and complexity of polynomials.
Kan Yao, Dur-e-Shahwar Kundi, Chenghua Wang, Máire O'Neill, Weiqiang Liu 0001
ISCAS2
2021 APAS: Application-Specific Accelerators for RLWE-Based Homomorphic Linear Transformations
abstract
Recently, the application of multi-party secure computing schemes based on homomorphic encryption in the field of machine learning attracts attentions across the research fields. Previous studies have demonstrated that secure protocols adopting packed additive homomorphic encryption (PAHE) schemes based on the ring learning with errors (RLWE) problem exhibit significant practical merits, and are particularly promising in enabling efficient secure inference in machine-learning-as-a-service applications. In this work, we introduce a new technique for performing homomorphic linear transformation (HLT) over PAHE ciphertexts. Using the proposed HLT technique, homomorphic convolutions and inner products can be executed without the use of number theoretic transform and the rotate-and-add algorithms that were proposed in existing works. To maximize the efficiency of the HLT technique, we propose APAS, a hardware-software co-design framework consisting of approximate arithmetic units for the hardware acceleration of HLT. In the experiments, we use actual neural network architectures as benchmarks to show that APAS can improve the computational and communicational efficiency of homomorphic convolution by 8× and 3×, respectively, with an energy reduction of up to 26× as compared to the ASIC implementations of existing methods.
Song Bian 0001, Dur-e-Shahwar Kundi, Kazuma Hirozawa, Weiqiang Liu 0001, Takashi Sato 0001
IEEE Trans. Inf. Forensics Secur.2
2020 AxMM: Area and Power Efficient Approximate Modular Multiplier for R-LWE Cryptosystem
abstract
Amongst various Post-Quantum Cryptographic (PQC) schemes, Lattice-Based Cryptography (LBC) stands out as the most viable substitute to the classical cryptographic schemes due to its efficiency, versatility and solid foundations on hard mathematical problems. Ring Learning With Errors (R-LWE) is a Public Key Encryption (PKE) scheme of LBC, in which the modular polynomial multiplication in a ring is the main bottleneck in the realization of a practical resource-constraint design for the embedded IoT devices. This work explores novel Approximate Computing (AC) technique for the design of area/power efficient modular multiplier (so called AxMM) for R-LWE, exploiting the inherent approximate structure of the scheme. The proposed AxMM on 45nm ASIC library achieved an area and power reduction of 36% and 23%, respectively, along with a speed increase of 1.34× as compared to state-of-art smallest exact R-LWE modular multiplier.
Dur-e-Shahwar Kundi, Song Bian 0001, Ayesha Khalid, Chenghua Wang, Máire O'Neill, Weiqiang Liu 0001
ISCAS1
2015 An efficient single unit T-box/T-1-box implementation for 128-bit AES on FPGA
abstract
Abstract In this paper, we present an area efficient Block RAM (BRAM)‐based single unit design of T‐box/T−1‐box on a Field Programmable Gate Array (FPGA) for combined Advanced Encryption Standard (AES) encryption and decryption. Conventional FPGA designs for T‐box module not only utilize several BRAMs for 16 bytes parallel look‐up operations but also because of asymmetric nature of AES, use separate hardware for T−1‐box unit in decryption process. Alternatively in iterative architecture, BRAM in dual‐port mode configuration takes eight clock cycles to access 16 look‐up operations from a T‐box because of its synchronous nature. Thus, resulting in an unoptimized solution not only in term of FPGA resources but also results in high latency for iterative architecture. Our proposed design uses single symmetric T‐box/T−1‐box table with same set of single resource‐shared hardware for both the encryptor and decryptor and at the same time performs eight look‐up operations from single BRAM in one clock cycle using efficient BRAM switching technique instead of using multirated clocking. Our complete 128‐bit symmetric T‐box/T−1‐box design fits into just 2 BRAMs and 136 Slices. It occupies lowest area reported to date with 50% power saving and highest Throughput Per Slice (TPS) of 10.77. Copyright © 2014 John Wiley & Sons, Ltd.
Dur-e-Shahwar Kundi, Arshad Aziz, Majida Kazmi
Secur. Commun. Networks1
2010 Resource efficient implementation of T-Boxes in AES on Virtex-5 FPGA
Dur-e-Shahwar Kundi, Arshad Aziz, Nassar Ikram
Inf. Process. Lett.1