Esam El-Araby

dblp:98/82 · DBLP profile ↗
← Back
23ranked-venue papers
12as first author
3since 2021 · last 2023
0000-0002-4575-1049ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 11 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorArtificial intelligence and machine learning · 1
YearPublicationVenuePosition
2023 Towards Complete and Scalable Emulation of Quantum Algorithms on High-Performance Reconfigurable Computers
abstract
Contemporary quantum computers face many critical challenges that limit their usefulness for practical applications. A primary limiting factor is classical-to-quantum (C2Q) data encoding, which requires specific circuits for quantum state initialization. The required state initialization circuits are often complex and violate decoherence constraints, particularly for I/O intensive applications. Existing Noisy Intermediate-Scale Quantum (NISQ) devices are noise-sensitive and have low quantum bit (qubit) counts, thus limiting the applicability of C2Q circuits for encoding large and realistic datasets. This has made the study of complete and realistic circuits that include data encoding challenging and has also led to a heavy dependency on costly and resource-intensive simulations on classical platforms. In this work, we propose a cost-effective, classical-hardware-accelerated framework for realistic and complete emulation of quantum algorithms. The emulation framework incorporates components for the critical C2Q data encoding process, as well as architectures for quantum algorithms such as the quantum Haar transform (QHT). The framework is used to investigate optimizations for C2Q and QHT algorithms, and the corresponding optimized quantum circuits are presented. The framework is implemented on a High-Performance Reconfigurable Computer (HPRC) which emulates the proposed QHT circuits combined with proposed C2Q data encoding methods. For performance benchmarks, CPU-based emulations and simulations on a state-of-the-art quantum computing simulator are also carried out. Results show that the proposed hardware-accelerated emulation framework is more efficient in terms of speed and scalability compared to CPU-based emulation and simulation.
Esam El-Araby, Naveed Mahmud, Mingyoung Jessica Jeng, Andrew MacGillivray, Manu Chaudhary, Md. Alvir Islam Nobel, S. M. Ishraq Ul Islam, Dylan Kneidel, Madeline R. Watson, Jack G. Bauer, Andrew E. Riachi
IEEE Trans. Computers1
2023 Improving quantum-to-classical data decoding using optimized quantum wavelet transform
Mingyoung Jessica Jeng, S. M. Ishraq Ul Islam, Andrew E. Riachi, Manu Chaudhary, Md. Alvir Islam Nobel, Dylan Kneidel, Vinayak Jha, Jack G. Bauer, Anshul Maurya, Naveed Mahmud, Esam El-Araby
J. Supercomput.12
2022 Quantum Dimension Reduction for Pattern Recognition in High-Resolution Spatio-Spectral Data
abstract
The promises of advanced quantum computing technology have driven research in the simulation of quantum computers on classical hardware, where the feasibility of quantum algorithms for real-world problems can be investigated. In domains such as High Energy Physics (HEP) and Remote Sensing Hyperspectral Imagery, classical computing systems are held back by enormous readouts of high-resolution data. Due to the multi-dimensionality of the readout data, processing and performing pattern recognition operations for this enormous data are both computationally intensive and time-consuming. In this article, we propose a methodology that utilizes Quantum Haar Transform (QHT) and a modified Grover's search algorithm for time-efficient dimension reduction and dynamic pattern recognition in data sets that are characterized by high spatial resolution and high dimensionality. QHT is performed on the data to reduce its dimensionality at preserved spatial locality, while the modified Grover's search algorithm is used to search for dynamically changing multiple patterns in the reduced data set. By performing search operations on the reduced data set, processing overheads are minimized. Moreover, quantum techniques produce results in less time than classical dimension reduction and search methods. The feasibility of the proposed methodology is verified by emulating the quantum algorithms on classical hardware based on field programmable gate arrays (FPGAs). We present designs of the quantum circuits for multi-dimensional QHT and multi-pattern Grover's search. We also present two emulation techniques and the corresponding hardware architectures for this methodology. A high performance reconfigurable computer (HPRC) was used for the experimental evaluation, and high-resolution images were used as the input data set. Analysis of the methods and implications of the experimental results are discussed.
Naveed Mahmud, Bennett Haase-Divine, Andrew MacGillivray, Esam El-Araby
IEEE Trans. Computers4
2019 Improving Emulation of Quantum Algorithms using Space-Efficient Hardware Architectures
abstract
With rapid advancement in quantum computing technology, continuous efforts are being directed to simulation and emulation of quantum algorithms on classical platforms. A well-known limitation to classical emulation of quantum circuits is scalability. Existing hardware emulators implement gate-based circuit models of quantum circuits that result in heavy resource utilization and degrade the scalability of the system. Also, current quantum emulation hardware use fixedpoint arithmetic, which has an adverse effect on accuracy when the system is scaled up. In this work, we employ a complexmultiply-and-accumulate (CMAC) and lookup-based emulation approach that greatly reduces resource utilization and improves system scalability in terms of number of emulated qubits. We demonstrate emulation of up to 16 fully-entangled qubits which is highest among existing work. We design fully-pipelined, highthroughput hardware architectures that use floating-point precision for higher accuracy. Experimental evaluation and analysis of the architectures in terms of speed and area is also provided. The emulator is prototyped on a high-performance reconfigurable computing (HPRC) system and our results demonstrate quantitative improvement over existing Field-Programmable-Gate-Array (FPGA)-based hardware emulators.
Naveed Mahmud, Esam El-Araby
ASAP2
2016 Chaotic architectures for secure free-space optical communication
abstract
Free-Space Optical (FSO) communication provides very large bandwidth, relatively low cost, low power, low mass of implementation, and improved security when compared to conventional Free-Space Radio-Frequency systems. Secure FSO communications has usually been proposed through the use of laser N-slit-interferometers. This technique, however, works over relatively short propagation distances, particularly for deep-space communication. In this paper, we propose the use of chaotic systems combined with FSO in order to achieve robust longer-range communication while maintaining the inherent security in chaotic systems targeting both space and terrestrial applications. Chaotic systems show particular characteristics such as broadband noise-like signals with multi-path fading resistance, unpredictability, and sensitivity to initial conditions which make it difficult for unintentional receivers to synchronize to the chaotic signal. We also propose the use of FPGAs as the implementation technology that could meet the demanding real-time requirements of FSO communications. The experimental results show favorable characteristics of our approach.
Esam El-Araby, Nader M. Namazi
FPL1
2014 GPU acceleration of nonlinear diffusion tensor estimation using CUDA and MPI
Lin-Ching Chang, Esam El-Araby, Vinh Q. Dang, Lam H. Dao
Neurocomputing2
2013 Accelerating nonlinear diffusion tensor estimation for medical image processing using high performance GPU clusters
abstract
Diffusion Tensor Imaging (DTI) is a non-invasive magnetic resonance technique that produces in vivo images of biological tissues with local microstructural characteristics such as water diffusion. It can be used, for example, to localize white matter lesions, or in neuro-navigation surgery of brain tumors. Diffusion tensor maps are usually computed on a voxel-by-voxel basis by fitting the signal intensities of diffusion weighted images as a function of their corresponding data acquisition parameters. This processing is highly computation-intensive and can be time-consuming which constraints the clinical use of DTI. This study presents the application of using high performance GPU clusters in diffusion tensor estimation by accelerating the multivariate non-linear regression. The results are tested in simulated DTI brain datasets and show significant performance gain in tensor fitting in addition to favorable scalability characteristics. The proposed GPU implementation framework can further promote the clinical use of DTI, and can be used to accelerate statistical analysis of DTI where Monte Carlo simulations are employed, or readily applied to quantitative assessment of DTI using bootstrap analysis.
Vinh Q. Dang, Esam El-Araby, Lam H. Dao, Lin-Ching Chang
ASAP2
2011 Modelling the performance of an SSD-Aware storage system using least squares regression
abstract
Flash memory has lately been used as a cache located between the system main memory and the magnetic hard drives in order to create robust and cost effective hybrid storage systems. The reason comes from the growing density of Solid State Devices (SSDs) at lower prices with main advantage of high random read efficiency compared to magnetic hard drives. When predicting the performance of such hybrid storage systems, it is inevitable to study the trade-off in selecting the different storage elements such as the main memory and the SSD cache relative to the capacity of the magnetic hard drive. The parameters of such prediction model are determined based on the application that the storage system would serve. In this paper, a prediction model that uses experimental evaluation of a hybrid storage system and real applications/benchmarks is used. The model utilizes parameters of both the storage system and applications in order to predict system performance based on metrics that are commonly used in storage system evaluation. The model allows the designer to select the best hybrid storage system parameters that satisfy certain application performance requirements. The model is highly accurate with a minimal error and a high prediction confidence level (95%) in reference to the experimental data collected from real applications using the proposed SSD-aware hybrid storage system.
Abdullah Aldahlawi, Esam El-Araby, Suboh A. Suboh, Tarek A. El-Ghazawi
AICCSA2
2011 GPU Resource Sharing and Virtualization on High Performance Computing Systems
abstract
Modern Graphic Processing Units (GPUs) are widely used as application accelerators in the High Performance Computing (HPC) field due to their massive floating-point computational capabilities and highly data-parallel computing architecture. Contemporary high performance computers equipped with co-processors such as GPUs primarily execute parallel applications using the Single Program Multiple Data (SPMD) model, which requires balanced computing resources of both microprocessor and co-processors to ensure full system utilization. While the inclusion of GPUs in HPC systems provides more computing resources and significant performance improvements, the asymmetrical distribution of the number of GPUs relative to the microprocessors can result in an underutilization of overall system computing resources. In this paper, we propose a GPU resource virtualization approach to allow underutilized microprocessors to share the GPUs. We analyze factors affecting the parallel execution performance on GPUs and conduct a theoretical performance estimation based on the most recent GPU architectures as well as the SPMD model. Then we present the implementation details of the virtualization infrastructure, followed by an experimental verification of the proposed concepts using an NVIDIA Fermi GPU computing node. The results demonstrate a considerable performance gain over the traditional SPMD execution without virtualization. Furthermore, the proposed solution enables full utilization of the asymmetrical system resources, through the sharing of the GPUs among microprocessors, while incurring low overheads due to the virtualization layer.
Teng Li 0009, Vikram K. Narayana, Esam El-Araby, Tarek A. El-Ghazawi
ICPP3
2011 A Framework for Evaluating High-Level Design Methodologies for High-Performance Reconfigurable Computers
abstract
High-performance reconfigurable computers have potential to provide substantial performance improvements over traditional supercomputers. Their acceptance, however, has been hindered by productivity challenges arising from increased design complexity, a wide array of custom design languages and tools, and often overblown sales literature. This paper presents a review and taxonomy of High-Level Languages (HLLs) and a framework for the comparative analysis of their features. It also introduces new metrics and a model based on computational effort. The proposed concepts are inspired by Netwon's equations of motion and the notion of work and power in an abstract multidimensional space of design specifications. The metrics are devised to highlight two aspects of the design process: the total time-to-solution and the efficient utilization of user and computing resources at discrete time steps along the development path. The study involves analytical and experimental evaluations demonstrating the applicability of the proposed model.
Esam El-Araby, Saumil G. Merchant, Tarek A. El-Ghazawi
IEEE Trans. Parallel Distributed Syst.1
2009 Exploiting Partial Runtime Reconfiguration for High-Performance Reconfigurable Computing
abstract
Runtime Reconfiguration (RTR) has been traditionally utilized as a means for exploiting the flexibility of High-Performance Reconfigurable Computers (HPRCs). However, the RTR feature comes with the cost of high configuration overhead which might negatively impact the overall performance. Currently, modern FPGAs have more advanced mechanisms for reducing the configuration overheads, particularly Partial Runtime Reconfiguration (PRTR). It has been perceived that PRTR on HPRC systems can be the trend for improving the performance. In this work, we will investigate the potential of PRTR on HPRC by formally analyzing the execution model and experimentally verifying our analytical findings by enabling PRTR for the first time, to the best of our knowledge, on one of the current HPRC systems, Cray XD1. Our approach is general and can be applied to any of the available HPRC systems. The paper will conclude with recommendations and conditions, based on our conceptual and experimental work, for the optimal utilization of PRTR as well as possible future usage in HPRC.
Esam El-Araby, Iván González 0004, Tarek A. El-Ghazawi
ACM Trans. Reconfigurable Technol. Syst.1
2008 Portable library development for reconfigurable computing systems: A case study
Proshanta Saha, Esam El-Araby, Miaoqing Huang, Mohamed Taher, Sergio López-Buedo, Tarek A. El-Ghazawi, Chang Shu 0003, Kris Gaj, Alan Michalski, Duncan A. Buell
Parallel Comput.2
2007 Bringing High-Performance Reconfigurable Computing to Exact Computations
abstract
Numerical non-robustness is a recurring phenomenon in scientific computing. It is primarily caused by numerical errors arising because of fixed-precision arithmetic in integer and/or floating-point computations. Exact computation, based on arbitrary-precision arithmetic, has been developed over the last decade as an emerging numerical computation paradigm in response to this problem of numerical non-robustness. Exact arithmetic, specifically arbitrary-precision arithmetic, has been traditionally implemented using efficient software libraries such as GNU Multi-Precision (GMP). However, this results in a slower arithmetic performance when compared to fixed-precision arithmetic. In this paper we present a first effort, to the best of our knowledge, of reconfigurable hardware support for arbitrary-precision arithmetic. The proposed hardware architectures are based on virtual convolution sche1duling which is derived from a formal representation of the problem. Targeting high performance and efficiency, dynamic (non-linear) pipelines techniques were exploited to eliminate the effects of deeply-pipelined operators. Referenced to GMP, our experiments showed promising results.
Esam El-Araby, Iván González 0004, Tarek A. El-Ghazawi
FPL1
2007 Productivity of High-Level Languages on Reconfigurable Computers: An HPC Perspective
abstract
Productivity on high-performance reconfigurable computers (HPRCs) is becoming a concern given the complexity of today's applications and development flows. Furthermore, the plethora of options from which application developers need to select their development environments has recently become another productivity obstacle. High-level languages (HLLs) for developing reconfigurable computing applications trade performance with ease-of-use. However, it is hard to know in a general sense how much performance one is giving up and how much ease-of-use he/she is gaining. More importantly, given the lack of standards and the uncertainty generated by sales literature, it is very hard to know the real differences that exist among different high-level programming paradigms. In order to do so, one needs a classification of HLLs programming models from a general high-performance computing (HPC) perspective. In this work, we consider a number of representative high-level tools that were selected to represent imperative programming, functional programming and graphical programming, and thereby demonstrate the applicability of our methodology. It will be shown that in spite of the disparity in concepts behind those tools, our methodology will be able to uncover the basic differences among them and assess their comparative productivity in terms of performance, and ease-of-use.
Esam El-Araby, Preetham Nosum, Tarek A. El-Ghazawi
FPT1
2005 Reconfigurable computers: an empirical analysis (abstract only)
abstract
Reconfigurable Computers are parallel systems that are designed around multiple general-purpose processors and multiple field programmable gate array (FPGA) chips. These systems can leverage the synergism between conventional processors and FPGAs to provide low-level hardware functionality at the same level of programmability as general-purpose computers. In this work we conduct an experimental study using one of the state-of-the-art reconfigurable computers and a representative set of applications to assess the field, uncover the challenges, propose solutions, and conceive a realistic evolution path. We consider issues of concern including performance/cost. We also consider productivity in the sense of development, compiling, running, and system reliability. It will be shown that for some applications, the performance/cost can be orders of magnitude better than conventional computers. It will be also shown that programming such machines may still require some hardware knowledge, similar to hardware knowledge computer programmers must acquire to write scalable programs.
Tarek A. El-Ghazawi, Kris Gaj, Nikitas A. Alexandridis, Allen Michalski, Osman Devrim Fidanci, Mohamed Taher, Esam El-Araby, Esmail Chitalwala, Proshanta Saha
FPGA7
2005 Image processing library for reconfigurable computers (abstract only)
abstract
Reconfigurable Computers (RCs) are parallel systems that are designed around multiple general-purpose processors and multiple field programmable gate array (FPGA) chips. These systems can leverage the synergism between conventional processors and FPGAs to provide low-level hardware functionality at the same level of programmability as general-purpose computers. RCs have proposed very high processing capabilities for computationally intensive applications such as Image Processing. This is due to the inherently parallel operation paradigm of the FPGA hardware.In this paper we present the design and implementation of image processing kernels for RCs. This library of kernels have been tested and verified for performance on one of the state-of-the-art reconfigurable computers, SRC-6E. This paper shows that RCs are between 8 to 400 times faster than comparable Pentiums for image based tasks.
Mohamed Taher, Esam El-Araby, Tarek A. El-Ghazawi, Kris Gaj
FPGA2
2005 A System-Level Design Methodology for Reconfigurable Computing Applications
Esam El-Araby, Tarek A. El-Ghazawi, Kris Gaj
FPT1
2005 Prototyping Automatic Cloud Cover Assessment (ACCA) Algorithm for Remote Sensing On-Board Processing on a Reconfigurable Computer
Esam El-Araby, Mohamed Taher, Tarek A. El-Ghazawi, Jacqueline LeMoigne-Stewart
FPT1
2005 Performance of Sorting Algorithms on the SRC 6 Reconfigurable Computer
John Harkins, Tarek A. El-Ghazawi, Esam El-Araby, Miaoqing Huang
FPT3
2005 An efficient implementation of on-board cloud detection on a reconfigurable computer
abstract
The presence of cloud contamination can hinder the use of satellite data, and this requires a cloud detection process to mask out cloudy pixels from further processing. The trend for remote sensing satellite missions has always been towards smaller size, lower cost, more flexibility, and higher computational power. Reconfigurable Computers (RCs) combine the flexibility of traditional microprocessors with the power of Field Programmable Gate Arrays (FPGAs). Therefore, RCs are a promising candidate for on-board preprocessing. This paper presents the design and implementation of an RC-based real-time cloud detection system. We investigate the potential of using RCs for on-board preprocessing by prototyping the Landsat 7 ETM+ ACCA algorithm on one of the state-of-the art reconfigurable platforms, SRC-6E. Although a reasonable amount of investigations of the ACCA cloud detection algorithm using FPGAs has been reported in the literature, very few details/results were provided and/or limited contributions were accomplished. Our work has been proven to provide higher performance and higher detection accuracy.
Esam El-Araby, Mohamed Taher, Tarek A. El-Ghazawi, Jacqueline LeMoigne-Stewart
IGARSS1
2004 Wavelet spectral dimension reduction of hyperspectral imagery on a reconfigurable computer
abstract
Hyperspectral imagery, by definition, provides valuable remote sensing observations at hundreds of frequency bands. Conventional image classification (interpretation) methods may not be used without dimension reduction preprocessing. Automatic wavelet reduction has been proven to yield better or comparable classification accuracy, while achieving substantial computational savings. However, the large hyperspectral data volumes remain to present a challenge for traditional processing techniques. Reconfigurable computers (RCs) can leverage the synergism between conventional processors and FPGAs to provide low-level hardware functionality at the same level of programmability as general-purpose computers. We investigate the potential of using RCs for on-board, i.e. aboard airborne/spaceborne carriers, preprocessing of hyperspectral imagery by prototyping for the first time the automatic wavelet dimension reduction algorithm. Our investigation exploits the fine and coarse grain parallelism provided by the RCs and has been experimentally verified on one of the state-of the art reconfigurable platforms, SRC-6E. An order of magnitude speedup over traditional processing techniques has been reported.
Esam El-Araby, Tarek A. El-Ghazawi, Jacqueline LeMoigne-Stewart, Kris Gaj
FPT1
2004 System-Level Parallelism and Throughput Optimization in Designing Reconfigurable Computing Applications
abstract
Summary form only given. Reconfigurable computers (RCs) can leverage the synergism between conventional processors and FPGAs to provide low-level hardware functionality at the same level of programmability as general-purpose computers. In a large class of applications, the total I/O time is comparable or even greater than the computations time. As a result, the rate of the DMA transfer between the microprocessor memory and the on-board memory of the FPGA-based processor becomes the performance bottleneck. We perform a theoretical and experimental study of this specific performance limitation. The mathematical formulation of the problem has been experimentally verified on the state-of-the art reconfigurable platform, SRC-6E. We demonstrate and quantify the possible solution to this problem that exploits the system-level parallelism within reconfigurable machines.
Esam El-Araby, Mohamed Taher, Kris Gaj, Tarek A. El-Ghazawi, David Caliga, Nikitas A. Alexandridis
IPDPS1
2003 Exploiting system-level parallelism in the application development on a reconfigurable computer
abstract
Reconfigurable Computers (RCs) can leverage the synergism between conventional processors and FPGAs to provide low-level hardware functionality at the same level of programmability as general-purpose computers. In a large class of applications, the total I/O time is comparable or even greater than the computations time. As a result, the rate of the DMA transfer between the microprocessor memory and the on-board memory becomes the performance bottleneck even on RCs. In this paper, we perform a theoretical and experimental study of this specific performance limitation for the state-of-the art reconfigurable platform, SRC-6E. We demonstrate and quantify the possible solution to this problem that exploits the system-level parallelism within the reconfigurable machine.
Esam El-Araby, Mohamed Taher, Kris Gaj, Tarek A. El-Ghazawi, David Caliga, Nikitas A. Alexandridis
FPT1