Mohammed A. S. Khalid

dblp:34/2645 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
4since 2021 · last 2024
0000-0003-3903-8789ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 2 first-author · 4 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2024 A High Speed and Area Efficient Processor for Elliptic Curve Scalar Point Multiplication for GF(2m)
abstract
Binary polynomial multipliers impact the overall performance and cost of elliptic curve cryptography (ECC) systems. Multiplication algorithms with subquadratic computational complexity are widely used to reduce area requirements and improve the delay of ECC cryptographic hardware. This work presents an elliptic curve scalar point multiplication (SPM) processor implementation using a novel classification of improved overlap-free multipliers targeting applications in the Internet of Things (IoT) devices. The proposed multipliers combine the advantages of fewer partial products and the overlap-free reconstructions which results in better recurrence and improved performance. The proposed multipliers and point multiplication hardware were designed, implemented, and tested on FPGA. The implemented processor presents a reasonable trade-off between speed and area consumption, and the design compares favorably with the previous designs in terms of area-delay product.
Madhan Thirumoorthi, Alexander J. Leigh, Moslem Heidarpur, Mitra Mirhassani, Mohammed A. S. Khalid
IEEE Trans. Very Large Scale Integr. Syst.5
2023 Novel Formulations of M-Term Overlap-Free Karatsuba Binary Polynomial Multipliers and Their Hardware Implementations
abstract
Novel binary polynomial multipliers have been designed using M-term overlap-free Karatsuba multiplication (OFKM), where$M$is 5–8. The proposed designs were realized in digital hardware and implemented on field-programmable gate array (FPGA) and the best value of$M$was selected and presented for common National Institute of Standards and Technology (NIST) operand sizes from 64 to 571 bits. The implemented hardware designs use a hybrid approach that combines a given M-term overlap-free Karatsuba multipliers with two-term splitting to reduce the need for zero-padding in the final recurrent stages. Compared to the traditional M-term Karatsuba multipliers, the proposed overlap-free implementations offer reductions in delay and area-delay product (ADP). The proposed designs also compare favorably to previous implementations of binary polynomial multipliers. Their favorable characteristics make the proposed overlap-free Karatsuba polynomial multipliers viable options for use in cryptographic systems where speed is a significant consideration and hardware resource consumption must be limited.
Madhan Thirumoorthi, Alexander J. Leigh, Moslem Heidarpur, Mohammed A. S. Khalid, Mitra Mirhassani
IEEE Trans. Very Large Scale Integr. Syst.4
2022 An Optimized M-Term Karatsuba-Like Binary Polynomial Multiplier for Finite Field Arithmetic
abstract
Finite field multiplication is a fundamental and frequently used operation in various cryptographic circuits and systems. Because of its high complexity, this operation generally determines the overall complexity and cost of these systems. Therefore, finite field multipliers and their hardware implementation have received considerable attention from researchers. This article proposes a methodology to design an efficient Galois field multiplier. First, space and time complexities for theoretical and field-programmable gate array (FPGA) implementations of M-term Karatsuba-like finite field multipliers were obtained. In addition, an algorithm was developed to obtain an efficient design based on a composite M-term Karatsuba-like multiplier. Furthermore, the proposed multipliers were verified and implemented on various FPGA devices, and implementation results were presented. Reported device utilization and latency indicated that the proposed multiplier is roughly 26% faster and 15% more efficient in the area–delay product compared to the standard Karatsuba multiplier. Moreover, comparison with state of the art also indicated that the proposed design is leading in terms of effectiveness and speed.
Madhan Thirumoorthi, Moslem Heidarpur, Mitra Mirhassani, Mohammed A. S. Khalid
IEEE Trans. Very Large Scale Integr. Syst.4
2021 Design and Evaluation of a Hybrid Chaotic-Bistable Ring PUF
abstract
A physical unclonable function (PUF) is a promising lightweight circuit that provides security and authentication capability for electronic devices with low computational resources. Among various PUFs, the bistable ring PUF (BR-PUF) is considered one of the robust configurations. However, it has been shown that the challenge-response pairs (CRPs) from BR-PUF are vulnerable to statistical machine learning (ML) attacks, such as k-junta learning, support vector machine (SVM), and logistic regression (LR). In this article, we first show that the k-junta attack can break CRPs from the BR-PUF. Then, we present a hybrid chaotic-BR-PUF structure that obfuscates the BR-PUF response with the nonlinearized chaotic response. The proposed PUF structure has been implemented and experimentally evaluated on Xilinx Artix-7 FPGA, and the PUF measurements were captured. The proposed PUF was tested with a powerful statistical method developed using k-junta-based learning to confirm its strength against such attacks and evaluated using CRPs collected. The proposed PUF provides better resistance against ML attacks and reduces the learning accuracy to 50%–60% compared with previously proposed PUFs.
Madhan Thirumoorthi, Marko Jovanovic, Mitra Mirhassani, Mohammed A. S. Khalid
IEEE Trans. Very Large Scale Integr. Syst.4
2016 An FPGA-Based Controller for a 77 GHz MEMS Tri-Mode Automotive Radar (Abstract Only)
abstract
No abstract available.
Sabrina Zereen, Sundeep Lal, Mohammed A. S. Khalid, Sazzadur Chowdhury
FPGA3
2016 Acceleration of k-Means Algorithm Using Altera SDK for OpenCL
abstract
A K-means clustering algorithm involves partitioning of data iteratively into k clusters. It is one of the most popular data-mining algorithms [Wu et al. 2007], and is widely used in other applications, such as image processing and machine learning. However, k-means is highly time-consuming when data or cluster size is large. Traditionally, FPGAs have shown great promise for accelerating computationally intensive algorithms, but they are harder to use for acceleration if we rely on traditional HD-based design methods. The recent introduction of Altera SDK for the OpenCL high-level synthesis tool allows developers to utilize FPGA's potential without long development periods and extensive hardware knowledge. This article presents an optimized implementation of a k-means clustering algorithm on an FPGA using Altera SDK for OpenCL. Performance and power consumption is measured with various data, cluster, and dimension sizes. When compared to state-of-the-art solutions, this implementation supports larger cluster sizes, offers up to 21x speed over a CPU and is more power efficient than a GPU. Unlike previous implementations, it can deliver consistently high throughput across large or small feature dimensions given reasonable cluster sizes and large enough data size.
Qing Y. Tang, Mohammed A. S. Khalid
ACM Trans. Reconfigurable Technol. Syst.2
2010 Design and evaluation of a parameterizable NoC router for FPGAs (abstract only)
abstract
The Network-on-Chip (NoC) approach for designing (System-on-Chip) SoCs is currently emerging as an advanced concept for overcoming the scalability and efficiency problems of traditional on-chip interconnection schemes, such as shared buses and point-to-point links. NoC design draws on concepts from computer networks to interconnect Intellectual Property (IP) cores in a structured and scalable way, promoting design re-use. We present the design and evaluation of a parameterizable NoC router for FPGAs. The importance of low area overhead for NoC components is crucial in FPGAs, which have fixed logic and routing resources. We achieve a low area router design through optimizations in switching fabric and dual purpose buffer/connection signals. We use a store and forward flow control with input and output buffering. We propose a component library to increase re-use and allow tailoring of parameters for application specific NoCs of various sizes. Our router supports the mesh architecture which is well known for its scalability and simple XY routing algorithm. We introduce IP-core-to-router mapping strategies for multi-local port routers that provide ample opportunity to optimize the NoC for application specific data traffic. A set of experiments were conducted to explore the design space of the proposed NoC router using different values of key router parameters: channel width (flit size), arbitration scheme and IP-core-to-router mapping strategy. Area and latency results from the experiments are presented and analyzed. These results will be useful to designers who want to implement NoC on FPGAs.
Mike Brugge, Mohammed A. S. Khalid
FPGA2
2006 A novel radius-adjusted approach for blind adaptive equalization
abstract
A new radius-adjusted approach for blind adaptive equalization for quadrature amplitude modulation (QAM) signals is introduced. Static circular contours are defined around an estimated symbol point in a QAM signal constellation, which creates regions that can be mapped to adaptation phases. The equalizer tap update consists of a linearly weighted sum of adaptation criteria that is scaled by a variable step size. Each region corresponds to a fixed step size and weighting factor, which creates a time-varying tap update based on the equalizer output radius. Two new algorithms are proposed based on this new approach and the multimodulus algorithm (MMA). The first algorithm trades off MMA and constellation-matched errors to reduce the time-to-convergence and mean-squared error (MSE), while the second trades off MMA and decision-directed errors to achieve reliable transfer between error modes and to obtain low MSE. A method to tune the proposed algorithms is developed based on statistics of the radius. The proposed algorithms are compared with related blind algorithms, and simulation results confirm that the proposed algorithms lead to enhanced performance.
Kevin Banovic, Esam Abdel-Raheem, Mohammed A. S. Khalid
IEEE Signal Process. Lett.3
2005 QPF: Efficient Quadratic Placement for FPGAs
abstract
In this paper we present QPF, a quadratic placement tool for FPGAs. Quadratic placement algorithms try to minimize total squared wire length by solving linear equations. The resulting placement tends to locate all cells near the center of the chip with a large amount of overlap. Also, since squared wire length is only an indirect measure of linear wire length, the resulting total wire length may not be minimized. We propose methods to alleviate the above two problems that give high quality results while minimizing the total run time. We incorporate multiple iterations of equation solving process together with a technique for pulling nodes out of the dense area while minimizing linear wire length. Experimental results using twenty MCNC benchmark circuits show that, on average, QPF is 5.8 times faster compared to a well known FPGA placement tool VPR, while providing almost comparable estimated total wire length.
Yonghong Xu, Mohammed A. S. Khalid
FPL2
2005 Hybrid Methods for Blind Adaptive Equalization: New Results and Comparisons
abstract
This paper proposes two new hybrid blind algorithms based on a new radius-adjusted approach for QAM signal constellations and presents a comprehensive survey of hybrid methods for blind adaptive equalization. The proposed hybrid blind algorithms define static circular regions around symbol points that correspond to a specific weighting factor and stepsize, which optimize the equalizer tap update based on the adaptation phase. Hybrid methods are discussed for the constant modulus algorithm (CMA), improved transfer to the decision-directed (DD) algorithm, and dual-mode hybrid algorithms. Comparisons are made between the proposed algorithms and related hybrid methods, and it is shown that the new algorithms lead to enhanced performance with minimal added complexity.
Kevin Banovic, Esam Abdel-Raheem, Mohammed A. S. Khalid
ISCC3
2000 A novel and efficient routing architecture for multi-FPGA systems
abstract
Multi-FPGA systems (MFSs) are used as custom computing machines, logic emulators and rapid prototyping vehicles. A key aspect of these systems is their programmable routing architecture which is the manner in which wires, FPGAs and field-programmable interconnect devices (FPIDs) are connected. Several routing architectures for MFSs have been proposed, and previous research has shown that the partial crossbar is one of the best existing architectures. In this paper, we propose a new routing architecture, called the hybrid complete-graph and partial-crossbar (HCGP) which has superior speed and cost compared to a partial crossbar. The new architecture uses both hard-wired and programmable connections between the FPGAs. We compare the performance and cost of the HCGP and partial crossbar architectures experimentally, by mapping a set of 15 large benchmark circuits into each architecture. A customized set of partitioning and interchip routing tools were developed, with particular attention paid to architecture-appropriate interchip routing algorithms. We show that the cost of the partial crossbar (as measured by the number of pins on all FPGAs and FPIDs required to fit a design), is on average 20% more than the new HCGP architecture and as much as 25% more. Furthermore, the critical path delay for designs implemented on the partial crossbar were on average 20% more than the HCGP architecture and up to 43% more. Using our experimental approach, we also explore a key architecture parameter associated with the HCGP architecture-the proportion of hard-wired connections versus programmable connections-to determine its best value.
Mohammed A. S. Khalid, Jonathan Rose
IEEE Trans. Very Large Scale Integr. Syst.1
1998 A Hybrid Complete-Graph Partial-Crossbar Routing Architecture for Multi-FPGA Systems
abstract
Multi-FPGA systems (MFSs) are used as custom computing machines, logic emulators and rapid prototyping vehicles. A key aspect of these systems is their programmable routing architecture; the manner in which wires, FPGAs and Field-Programmable Interconnect Devices (FPIDs) are connected. Several routing architectures for MFSs have been proposed [Arno92] [Butt92] [Hauc94] [Apti96] [Vuil96] and previous research has shown that the partial crossbar is one of the best existing architectures [Kim96] [Khal97]. In this paper we propose a new routing architecture, called the Hybrid Complete-Graph and Partial-Crossbar (HCGP) which has superior speed and cost compared to a partial crossbar. The new architecture uses both hard-wired and programmable connections between the FPGAs.We compare the performance and cost of the HCGP and partial crossbar architectures experimentally, by mapping a set of 15 large benchmark circuits into each architecture. A customized set of partitioning and inter-chip routing tools were developed, with particular attention paid to architecture-appropriate inter-chip routing algorithms. We show that the cost of the partial crossbar (as measured by the number of pins on all FPGAs and FPIDs required to fit a design), is on average 20% more than the new HCGP architecture and as much as 35% more. Furthermore, the critical path delay for designs implemented on the partial crossbar increased, and were on average 9% more than the HCGP architecture and up to 26% more.Using our experimental approach, we also explore a key architecture parameter associated with the HCGP architecture: the proportion of hard-wired connections versus programmable connections, to determine its best value.
Mohammed A. S. Khalid, Jonathan Rose
FPGA1