EDBT 2026 Demo / reviewers in the wild / expert
Khalid Javeed
dblp:152/9692
· DBLP profile ↗
6ranked-venue papers
4as first author
4since 2021 · last 2024
0000-0003-4645-4043ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Efficient Soft Core Multiplier for Post Quantum Digital SignaturesabstractMultiplication is a core operation in various applications such as cryptography and machine learning. Dedicated DSP blocks are provided by FPGA vendors for multiplication. However, these DSP blocks are limited in number and their location on FPGA is fixed, resulting in routing delays that affects the performance for small size multipliers. In this paper, a high performance and resource efficient 5 × 5 multiplier is presented that utilizes lookup tables (LUTs) and fast carry chain of the FPGA. The proposed multiplier offers 30% reduction in LUTs compared to Vivado DSP-less inferred multiplier at the cost of a slight increase in critical path delay (CPD). The proposed multiplier requires lesser power consumption and has better area- delay product (ADP) and power-delay product (PDP) metrics. Based on the proposed multiplier, a finite field multiplier is developed for post quantum digital signatures such as QR-UOV, MAYO and MQOM. The matrix-vector architecture is the core operation in multivariate digital signatures and integration of our finite field multiplier in a matrix-vector architecture shows that area is almost halved compared to state-of-the-art. Yasir Ali Shah, Ciara Rafferty, Ayesha Khalid, Safiullah Khan, Khalid Javeed, Máire O'Neill |
ISCAS | 5 |
| 2024 | GMC-crypto: Low latency implementation of ECC point multiplication for generic Montgomery curves over GF(p)
Khalid Javeed, Yasir Ali Shah, David Gregg |
J. Parallel Distributed Comput. | 1 |
| 2024 | Privacy-preserving collaborative AI for distributed deep learning with cross-sectional data
Saeed Iqbal, Adnan N. Qureshi, Musaed Alhussein, Khursheed Aurangzeb, Khalid Javeed, Rizwan Ali Naqvi |
Multim. Tools Appl. | 5 |
| 2024 | E2CSM: efficient FPGA implementation of elliptic curve scalar multiplication over generic prime field GF(p)
Khalid Javeed, Ali El-Moursy, David Gregg |
J. Supercomput. | 1 |
| 2015 | Serial and parallel interleaved modular multipliers on FPGA platformabstractModular multiplication is a core operation in all public key based cryptosystems. The performance of these cryptosystems can be enhanced substantially by incorporating an optimized modular multiplier. This paper presents serial and parallel radix-4 modular multipliers based on interleaved multiplication algorithm and Montgomery power laddering technique. A serial radix-4 interleaved modular multiplier provides 50% reduction in the required clock cycles. In addition to the reduction in clock cycles, a parallel modular multiplier maintains a critical path delay comparable to the bit serial interleaved multipliers. The proposed designs are implemented in Verilog HDL and synthesized targeting virtex-6 FPGA platform using Xilinx ISE 14.2 Design suite. The serial radix-4 multiplier computes a 256-bit modular multiplication in 1.3μs, occupies 3.9K LUTs, and runs at 96 MHz. The parallel radix-4 multiplier takes 0.77μs, occupies 5.3K LUTs, and runs at 166 MHz. The results show that the parallel radix-4 modular multiplier provides 62% and 49% speed-up over the corresponding bit serial and bit parallel versions, respectively. Thus, these designs are suitable to accelerate modular multiplication in many cryptographic processors. Khalid Javeed, Xiaojun Wang 0001, Mike Scott |
FPL | 1 |
| 2014 | Radix-4 and radix-8 booth encoded interleaved modular multipliers over general FpabstractThis paper presents radix-4 and radix-8 Booth encoded modular multipliers over general Fpbased on inter-leaved multiplication algorithm. An existing bit serial interleaved multiplication algorithm is modified using radix-4, radix-8 and Booth recoding techniques. The modified radix-4 and radix-8 versions of interleaved multiplication result in 50% and 75% reduction in required number of clock cycles for one modular multiplication over the corresponding bit serial interleaved multipliers, while maintaining a competitive critical path delay. The proposed architectures are implemented in Verilog HDL and synthesized by targeting virtex-6 FPGA platform. Due to an efficient utilization of optimized addition chains available in FPGAs and exploiting the parallelism among operations, the proposed radix-4 and radix-8 multipliers compute one 256 × 256 bit modular multiplication in 1.49μs and 0.93μs respectively, which are 35% and 94% improvement over the corresponding bit serial version. Further, this work also presents a thorough comparison on basis of area, throughput, and area × time per bit value. Which shows that these designs are efficiently optimized for area × time per bit value with a high throughput rate. Thus, these designs are suitable to construct most of the elliptic curve and pairing based cryptographic processors. Khalid Javeed, Xiaojun Wang 0001 |
FPL | 1 |