Khalid Javeed

dblp:152/9692 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
4since 2021 · last 2024
0000-0003-4645-4043ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Efficient Soft Core Multiplier for Post Quantum Digital Signatures
abstract
Multiplication is a core operation in various applications such as cryptography and machine learning. Dedicated DSP blocks are provided by FPGA vendors for multiplication. However, these DSP blocks are limited in number and their location on FPGA is fixed, resulting in routing delays that affects the performance for small size multipliers. In this paper, a high performance and resource efficient 5 × 5 multiplier is presented that utilizes lookup tables (LUTs) and fast carry chain of the FPGA. The proposed multiplier offers 30% reduction in LUTs compared to Vivado DSP-less inferred multiplier at the cost of a slight increase in critical path delay (CPD). The proposed multiplier requires lesser power consumption and has better area- delay product (ADP) and power-delay product (PDP) metrics. Based on the proposed multiplier, a finite field multiplier is developed for post quantum digital signatures such as QR-UOV, MAYO and MQOM. The matrix-vector architecture is the core operation in multivariate digital signatures and integration of our finite field multiplier in a matrix-vector architecture shows that area is almost halved compared to state-of-the-art.
Yasir Ali Shah, Ciara Rafferty, Ayesha Khalid, Safiullah Khan, Khalid Javeed, Máire O'Neill
ISCAS5
2024 GMC-crypto: Low latency implementation of ECC point multiplication for generic Montgomery curves over GF(p)
Khalid Javeed, Yasir Ali Shah, David Gregg
J. Parallel Distributed Comput.1
2024 Privacy-preserving collaborative AI for distributed deep learning with cross-sectional data
Saeed Iqbal, Adnan N. Qureshi, Musaed Alhussein, Khursheed Aurangzeb, Khalid Javeed, Rizwan Ali Naqvi
Multim. Tools Appl.5
2024 E2CSM: efficient FPGA implementation of elliptic curve scalar multiplication over generic prime field GF(p)
Khalid Javeed, Ali El-Moursy, David Gregg
J. Supercomput.1
2015 Serial and parallel interleaved modular multipliers on FPGA platform
abstract
Modular multiplication is a core operation in all public key based cryptosystems. The performance of these cryptosystems can be enhanced substantially by incorporating an optimized modular multiplier. This paper presents serial and parallel radix-4 modular multipliers based on interleaved multiplication algorithm and Montgomery power laddering technique. A serial radix-4 interleaved modular multiplier provides 50% reduction in the required clock cycles. In addition to the reduction in clock cycles, a parallel modular multiplier maintains a critical path delay comparable to the bit serial interleaved multipliers. The proposed designs are implemented in Verilog HDL and synthesized targeting virtex-6 FPGA platform using Xilinx ISE 14.2 Design suite. The serial radix-4 multiplier computes a 256-bit modular multiplication in 1.3μs, occupies 3.9K LUTs, and runs at 96 MHz. The parallel radix-4 multiplier takes 0.77μs, occupies 5.3K LUTs, and runs at 166 MHz. The results show that the parallel radix-4 modular multiplier provides 62% and 49% speed-up over the corresponding bit serial and bit parallel versions, respectively. Thus, these designs are suitable to accelerate modular multiplication in many cryptographic processors.
Khalid Javeed, Xiaojun Wang 0001, Mike Scott
FPL1
2014 Radix-4 and radix-8 booth encoded interleaved modular multipliers over general Fp
abstract
This paper presents radix-4 and radix-8 Booth encoded modular multipliers over general Fpbased on inter-leaved multiplication algorithm. An existing bit serial interleaved multiplication algorithm is modified using radix-4, radix-8 and Booth recoding techniques. The modified radix-4 and radix-8 versions of interleaved multiplication result in 50% and 75% reduction in required number of clock cycles for one modular multiplication over the corresponding bit serial interleaved multipliers, while maintaining a competitive critical path delay. The proposed architectures are implemented in Verilog HDL and synthesized by targeting virtex-6 FPGA platform. Due to an efficient utilization of optimized addition chains available in FPGAs and exploiting the parallelism among operations, the proposed radix-4 and radix-8 multipliers compute one 256 × 256 bit modular multiplication in 1.49μs and 0.93μs respectively, which are 35% and 94% improvement over the corresponding bit serial version. Further, this work also presents a thorough comparison on basis of area, throughput, and area × time per bit value. Which shows that these designs are efficiently optimized for area × time per bit value with a high throughput rate. Thus, these designs are suitable to construct most of the elliptic curve and pairing based cryptographic processors.
Khalid Javeed, Xiaojun Wang 0001
FPL1