Mohamed El-Hadedy 0001

dblp:03/7947 · also Mohamed Ezzat El-Hadedy Aly · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
2since 2021 · last 2023
0000-0002-3823-0712ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1Security and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Reconfigurable computing and FPGAs · 74% Hardware accelerators and domain-specific architectures · 26%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs › FPGA accelerator
short read mapping
0.722019
ASAP: Accelerated Short-Read Alignment on Programmable Hardware · IEEE Trans. Computers 2019
ASAP: Accelerated Short Read Alignment on Programmable Hardware (Abstract Only) · FPGA 2017
Hardware accelerators and domain-specific architectures
bioinformatics accelerator
0.412019
ASAP: Accelerated Short-Read Alignment on Programmable Hardware · IEEE Trans. Computers 2019
Reconfigurable computing and FPGAs
FPGA accelerator
0.312017
ASAP: Accelerated Short Read Alignment on Programmable Hardware (Abstract Only) · FPGA 2017
Bioinformatics and computational biology
sequence alignment
0.222019
ASAP: Accelerated Short-Read Alignment on Programmable Hardware · IEEE Trans. Computers 2019
ASAP: Accelerated Short Read Alignment on Programmable Hardware (Abstract Only) · FPGA 2017
Bioinformatics and computational biology › sequence alignment
edit distance
0.112019
ASAP: Accelerated Short-Read Alignment on Programmable Hardware · IEEE Trans. Computers 2019
Reconfigurable computing and FPGAs
FPGA implementation
0.112019
ASAP: Accelerated Short-Read Alignment on Programmable Hardware · IEEE Trans. Computers 2019
Bioinformatics and computational biology › sequence analysis
read mapping
0.112017
ASAP: Accelerated Short Read Alignment on Programmable Hardware (Abstract Only) · FPGA 2017

Methods — techniques the papers use, named apart from their topics

programmable hardware acceleration · 0.6
YearPublicationVenuePosition
2023 RECO-ASCON: Reconfigurable ASCON hash functions for IoT applications
Mohamed El-Hadedy 0001, Xinfei Guo, Kazutomo Yoshii, Yichen Cai 0004, Robert Herndon, Bryan Banta, Wen-Mei W. Hwu
Integr.1
2022 Agile-AES: Implementation of configurable AES primitive with agile design approach
Xinfei Guo, Mohamed El-Hadedy 0001, Sergiu Mosanu, Xiangdong Wei, Kevin Skadron, Mircea R. Stan
Integr.2
2020 Ensemble Hyperspectral Band Selection for Detecting Nitrogen Status in Grape Leaves
abstract
The large data size and dimensionality of hyperspectral data demands complex processing and data analysis. Multispectral data do not suffer the same limitations, but are normally restricted to blue, green, red, red edge, and near infrared bands. This study aimed to identify the optimal set of spectral bands for nitrogen detection in grape leaves using ensemble feature selection on hyperspectral data from over 3,000 leaves from 150 ‘Flame Seedless’ table grapevines. Six machine learning base rankers were included in the ensemble: random forest, LASSO, SelectKBest, ReliefF, SVM-RFE, and chaotic crow search algorithm (CCSA). The pipeline identified less than 0.45% of the bands as most informative about grape nitrogen status. The selected violet, yellow-orange, and shortwave infrared bands lie outside of the typical blue, green, red, red edge, and near infrared bands of commercial multispectral cameras, so the potential improvement in remote sensing of nitrogen in grapevines brought forth by a customized multispectral sensor centered at the selected bands is promising and worth further investigation. The proposed pipeline may also be used for application-specific multispectral sensor design in domains other than agriculture.
Ryan Omidi, Ali Moghimi, Alireza Pourreza, Mohamed El-Hadedy 0001, Anas Salah Eddin
ICMLA4
2019 Flexi-AES: A Highly-Parameterizable Cipher for a Wide Range of Design Constraints
abstract
Interconnected devices communicate efficiently and securely over untrusted networks via security protocols that employ various encryption algorithms, often as hardware modules. State-of-the-art hardware implementations typically focus on optimizing a single metric and are tedious to adapt to a wider set of design constraints. In this work, we develop an open-source, flexible and parameterizable hardware implementation of the Advanced Encryption Standard (AES). We present a feature-rich implementation in Chisel that is simple to employ to any architectures and to fine-tune to specific design requirements. Despite the larger design space, we use 50% fewer lines of code than existing Verilog versions, thus enabling a higher level of development productivity.
Sergiu Mosanu, Xinfei Guo, Mohamed El-Hadedy 0001, Lorena Anghel, Mircea R. Stan
FCCM3
2019 Analysis and Optimization of I/O Cache Coherency Strategies for SoC-FPGA Device
abstract
Unlike traditional PCIe-based FPGA accelerators, heterogeneous SoC-FPGA devices provide tighter integrations between software running on CPUs and hardware accelerators. Modern heterogeneous SoC-FPGA platforms support multiple I/O cache coherence options between CPUs and FPGAs, but these options can have inadvertent effects on the achieved bandwidths depending on applications and data access patterns. To provide the most efficient communications between CPUs and accelerators, understanding the data transaction behaviors and selecting the right I/O cache coherence method is essential. In this paper, we use Xilinx Zynq UltraScale+ as the SoC platform to show how certain I/O cache coherence method can perform better or worse in different situations, ultimately affecting the overall accelerator performances as well. Based on our analysis, we further explore possible software and hardware modifications to improve the I/O performances with different I/O cache coherence options. With our proposed modifications, the overall performance of SoC design can be averagely improved by 20%.
Seungwon Min, Sitao Huang, Mohamed El-Hadedy 0001, Jinjun Xiong, Deming Chen, Wen-Mei W. Hwu
FPL3
2019 Analysis and Modeling of Collaborative Execution Strategies for Heterogeneous CPU-FPGA Architectures
abstract
Heterogeneous CPU-FPGA systems are evolving towards tighter integration between CPUs and FPGAs for improved performance and energy efficiency. At the same time, programmability is also improving with High Level Synthesis tools (e.g., OpenCL Software Development Kits), which allow programmers to express their designs with high-level programming languages, and avoid time-consuming and error-prone register-transfer level (RTL) programming. In the traditional loosely-coupled accelerator mode, FPGAs work as offload accelerators, where an entire kernel runs on the FPGA while the CPU thread waits for the result. However, tighter integration of the CPUs and the FPGAs enables the possibility of fine-grained collaborative execution, i.e., having both devices working concurrently on the same workload. Such collaborative execution makes better use of the overall system resources by employing both CPU threads and FPGA concurrency, thereby achieving higher performance. In this paper, we explore the potential of collaborative execution between CPUs and FPGAs using OpenCL High Level Synthesis. First, we compare various collaborative techniques (namely, data partitioning and task partitioning), and evaluate the tradeoffs between them. We observe that choosing the most suitable partitioning strategy can improve performance by up to 2x. Second, we study the impact of a common optimization technique, kernel duplication, in a collaborative CPU-FPGA context. We show that the general trend is that kernel duplication improves performance until the memory bandwidth saturates. Third, we provide new insights that application developers can use when designing CPU-FPGA collaborative applications to choose between different partitioning strategies. We find that different partitioning strategies pose different tradeoffs (e.g., task partitioning enables more kernel duplication, while data partitioning has lower communication overhead and better load balance), but they generally outperform execution on conventional CPU-FPGA systems where no collaborative execution strategies are used. Therefore, we advocate even more integration in future heterogeneous CPU-FPGA systems (e.g., OpenCL 2.0 features, such as fine-grained shared virtual memory).
Sitao Huang, Li-Wen Chang, Izzat El Hajj, Simon Garcia de Gonzalo, Juan Gómez-Luna, Sai Rahul Chalamalasetti, Mohamed El-Hadedy 0001, Dejan S. Milojicic, Onur Mutlu, Deming Chen, Wen-Mei W. Hwu
ICPE7
2019 Reco-Pi: A reconfigurable Cryptoprocessor for π-Cipher
Mohamed El-Hadedy 0001, Amit Kulkarni 0002, Dirk Stroobandt, Kevin Skadron
J. Parallel Distributed Comput.1
2019 ASAP: Accelerated Short-Read Alignment on Programmable Hardware
abstract
The proliferation of high-throughput sequencing machines ensures rapid generation of up to billions of short nucleotide fragments in a short period of time. This massive amount of sequence data can quickly overwhelm today's storage and compute infrastructure. This paper explores the use of hardware acceleration to significantly improve the runtime of short-read alignment, a crucial step in preprocessing sequenced genomes. We focus on the Levenshtein distance (edit-distance) computation kernel and propose the ASAP accelerator, which utilizes the intrinsic delay of circuits for edit-distance computation elements as a proxy for computation. Our design is implemented on an Xilinx Virtex 7 FPGA in an IBM POWER8 system that uses the CAPI interface for cache coherence across the CPU and FPGA. Our design is$200\times$faster than an equivalent Smith-Waterman-C implementation of the kernel running on the host processor,$40-60\times$faster than an equivalent Landau-Vishkin-C++ implementation of the kernel running on the IBM Power8 host processor, and$2\times$faster for an end-to-end alignment tool for 120–150 base-pair short-read sequences. Further the design represents a$3760\times$improvement over the CPU in performance/Watt terms.
Subho S. Banerjee, Mohamed El-Hadedy 0001, Jong Bin Lim, Zbigniew T. Kalbarczyk, Deming Chen, Steven S. Lumetta, Ravishankar K. Iyer
IEEE Trans. Computers2
2017 ASAP: Accelerated Short Read Alignment on Programmable Hardware (Abstract Only)
Subho S. Banerjee, Mohamed El-Hadedy 0001, Jong Bin Lim, Daniel Chen 0001, Zbigniew T. Kalbarczyk, Deming Chen, Ravishankar K. Iyer
FPGA2
2017 On accelerating pair-HMM computations in programmable hardware
abstract
This paper explores hardware acceleration to significantly improve the runtime of computing the forward algorithm on Pair-HMM models, a crucial step in analyzing mutations in sequenced genomes. We describe 1) the design and evaluation of a novel accelerator architecture that can efficiently process real sequence data without performing wasteful work; and 2) aggressive memoization techniques that can significantly reduce the number of invocations of, and the amount of data transferred to the accelerator. We describe our demonstration of the design on a Xilinx Virtex 7 FPGA in an IBM Power8 system. Our design achieves a 14.85× higher throughput than an 8-core CPU baseline (that uses SIMD and multi-threading) and a 147.49 × improvement in throughput per unit of energy expended on the NA12878 sample.
Subho S. Banerjee, Mohamed El-Hadedy 0001, Ching Y. Tan, Zbigniew T. Kalbarczyk, Steven S. Lumetta, Ravishankar K. Iyer
FPL2
2016 Generating efficient and high-quality pseudo-random behavior on Automata Processors
abstract
Micron's Automata Processor (AP) efficiently emulates non-deterministic finite automata and has been shown to provide large speedups over traditional von Neumann execution for massively parallel, rule-based, data-mining and pattern matching applications. We demonstrate the AP's ability to generate high-quality and energy efficient pseudo-random behavior for use in pseudo-random number generation or in chip simulation. By recognizing that transition rules become probabilistic when input characters are randomized, the AP is also capable of simulating Markov chains. Combining hundreds of parallel Markov chains creates high-quality, high-throughput pseudo-random number sequences with greater power efficiency than state-of-the-art CPU and GPU algorithms. This indicates that the AP could potentially accelerate other Markov Chain-based applications such as agent-based simulation. We explore how to achieve throughputs upwards of 40GB/s per AP chip, with power efficiency 6.8x greater than state-of-the-art pseudo-random number generation on GPUs.
Jack Wadden, Nathan Brunelle, Ke Wang 0011, Mohamed El-Hadedy 0001, Gabriel Robins, Mircea R. Stan, Kevin Skadron
ICCD4
2011 An Efficient Authorship Protection Scheme for Shared Multimedia Content
abstract
Many electronic content providers today like Flickr and Google, offer space to users to publish their electronic media(e.g. photos and videos) in their cloud infrastructures so that they can be publicly accessed. Features like including other information, such as keywords or owner information into the digital material is already offered by existing providers. Despite the useful features made available to users by such infrastructures, the authorship of the published content is not protected against various attacks such as compression. In this paper we propose a robust scheme that uses digital invisible watermarking and hashing to protect the authorship of the digital content and provide resistance against malicious manipulation of multimedia content. The scheme is enhanced by an algorithm called MMBEC, that is an extension of an established scheme MBEC towards higher resistance.
Mohamed El-Hadedy 0001, Georgios Pitsilis, Svein J. Knapskog
ICIG1
2010 Resource-efficient implementation of Blue Midnight Wish-256 hash function on Xilinx FPGA platform
abstract
This paper presents the design and analysis of an area efficient Blue Midnight Wish compression function with digest size of 256 bits (BMW-256) on FPGA platforms. The proposed architecture achieves significant improvements in system throughput with reduced area. We demonstrate the performance of the proposed BMW hash function core using VIRTEX 5 FPGA implementation. The new BMW hash function design allows for 16X speed up in performance while consuming significantly lower area than previously reported (i.e. just 445 slices).
Mohamed El-Hadedy 0001, Martin Margala, Danilo Gligoroski, Svein J. Knapskog
IAS1