VLDB 2026 Research / reviewers in the wild / expert
Mohamed Abdelghany
dblp:227/8627 · also Mohamed A. Abd El Ghany, Mohamed A. Abd El-Ghany, Mohamed AbdelGhany
· DBLP profile ↗
26ranked-venue papers
8as first author
12since 2021 · last 2026
0000-0002-6282-7738ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 4 first-author · 8 since 2021Software engineering, systems software and programming languages · 4 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Theory of computation · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A First-Settle, First-Convert Scheduling Scheme for Low-Latency AI Inference in Crossbar ArraysabstractResistive random access memory (RRAM) crossbar arrays enable analog in-memory computing, promising high parallelism and reduced data transfer for deep neural networks (DNNs) inference. However, when many columns share a single time-multiplexed analog-to-digital converter (ADC), readout latency, dominated by serialized conversion, can become a performance bottleneck. This article presents the First-Settle, First-Convert (FSFC) readout architecture, a readiness-driven ADC scheduling scheme for RRAM-based fully connected neural networks (FCNNs). Conventional crossbar readout relies on global settling, where all bitlines must wait for the slowest column to stabilize before sequential conversion begins. FSFC replaces this schedule with per-column readiness detection and a priority-queue encoder that continuously routes the earliest-settled column to the shared ADC, overlapping bitline settling and conversion without adding extra ADCs or changing the crossbar. Software experiments and SPICE-level circuit simulations in SkyWater 130 nm show that FSFC reduces inference latency by 35%–67% across 1- to 8-bit quantized MNIST FCNNs, with no loss of classification accuracy. A transistor-level implementation of the differentiator, hysteresis comparator, and digital priority-queue encoder incurs about 80 pJ of control energy per matrix-vector multiplication (MVM) for a$30\times 10$array, scaling to approximately 1 nJ per MVM and about 5%–8% tile area overhead for a 128-column tile. Because FSFC shortens ADC active time while introducing only modest peripheral circuitry, it provides a low-cost, drop-in latency improvement for RRAM compute-in-memory (CIM) tiles. The technique is orthogonal to existing ADC pooling, precision-reduction, and analog-cascading schemes, and can be integrated as a readout front-end in future crossbar-based accelerators. Maryam M. Atia, Mohamed Abdelghany, Mohammed Ismail 0001, Mohammad Alhawari |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2025 | $\mathbb {FETMA}$: A Tool for Functional Block Diagram and Event Tree Based Safety Analysis
Mohamed Abdelghany, Adnan Rashid, Sofiène Tahar |
VECoS | 1 |
| 2025 | Domain-Specific Hyperdimensional RISC-V Processor for Edge-AI TrainingabstractEdge AI has become the cornerstone of many applications. Yet, progress is limited by the large complexity of training a deep neural network (a DNN). hyperdimensional computing (HDC) is positioned as an alternative approach for Edge AI that is compact enough to enable training. The main challenge for an HDC model is to maintain its key features while balancing high inference accuracy with efficiency. A simple binary HDC model lacks accuracy, while the computational complexity of a floating-point model is too high. This work presents FixedHD, a novel 16-bit fixed-point HDC model enabling training at the Edge. FixedHD achieves an accuracy similar to floating-point model while lowering computational complexity. The model is supported by a customized RISC-V processor tailored to speedup both training and inference. The processor is extended with advanced HDC-specific instructions, a vector unit to utilize HDC’s parallel nature, and, for the first time, approximate computing to exploit its robustness. Further, memory requirements are reduced by quantizing mathematical functions and reducing the large HDC encoding matrix by up to 390 x. Compared to the baseline processor, inference and training are accelerated on average by 6.9 x and 3 x, respectively. The energy consumption is reduced by 4.6 x and 1.9 x at the cost of an increase in area by 45 %. The inference accuracy remains at the high level of floating-point models despite the heavy quantization and approximation. Sandy A. Wasif, Miran Wael, Paul R. Genssler, Eman Azab, Maggie Mashaly, Mohamed Abdelghany, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2025 | Timekeepers: ML-Driven SDF Analysis for Power-Wasters Detection in FPGAsabstractAs the integration of FPGAs into cloud computing platforms accelerates, the risk of fault injection attacks - especially through power-wasting designs - becomes increasingly critical. Malicious tenants can upload FPGA designs that, under specific input stimuli, generate excessive power consumption, jeopardizing the integrity of the shared power delivery network (PDN) and enabling denial-of-service or side-channel attacks. Traditional detection techniques relying on netlist and bitstream analysis struggle with generalization and can be evaded through circuit obfuscation and seemingly benign designs. In contrast to these netlist-based approaches, we introduce Timekeepers, a novel detection method that utilizes Standard Delay Format (SDF) timing data combined with machine learning to detect anomalous power behavior in synthesized FPGA designs. Our method trains a decision tree classifier on SDF files generated from both benign and malicious designs, focusing on timing characteristics such as propagation delays and setup/hold violations to identify power wasters at the primitive level. By abstracting away from circuit connectivity and emphasizing timing patterns, our framework is both scalable and robust across different FPGA architectures. The classifier independently evaluates each FPGA component and aggregates the results using a threshold-based voting system to improve detection granularity and reduce false positives. Timekeepers achieves 99.6% accuracy and demonstrates superior performance compared to state-of-the-art solutions. Furthermore, our approach is platform-agnostic and does not require access to netlists or bitstreams, preserving intellectual property confidentiality while enhancing pre-deployment security checks. Mohamed Fathy, Hassan Nassar, Mohamed Abdelghany, Jörg Henkel |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2024 | A Framework for Formal Probabilistic Risk Assessment Using HOL Theorem Proving
Mohamed Abdelghany, Adnan Rashid, Sofiène Tahar |
CICM | 1 |
| 2023 | Auto-DOK: Compiler-Assisted Automatic Detection of Offload Kernels for FPGA-HBM ArchitecturesabstractThe bandwidth improvement provided by high-bandwidth memory (HBM), and the capability of FPGAs to customize the processing and memory hierarchy, results in a considerable performance increase for memory-intensive work-loads such as graph processing, sorting, machine learning, and database analytics. Modern systems integrating 3D-stacked DRAM memory can be leveraged to realize the Near-Memory Computing (NMC) paradigm by offloading some computations to accelerators placed near the HBM. Although numerous studies have investigated efficient accelerators for FPGA-HBM platforms, researchers have not proposed a systematic way for identifying which application kernels are suitable for execution near the HBM. In this article, we propose compiler support for recognizing offloading candidates without any burden on programmers. Auto-DOK analyzes an application code based on criteria derived from the hardware design goals of FPGA-HBM platforms, and automatically identifies kernels suitable for offloading. We evaluate Auto-DOK on benchmarks ranging from microbenchmarks to real-world kernels. Our results show that Auto-DOK can correctly identify kernels and input sizes suitable for execution near the HBM, and prevents slowdown caused by incorrect offloading decisions for other workloads. Moreover, Auto-DOK operates at compile time with negligible overhead and without the need for expensive profiling. Veronia Iskandar, Mohamed Abdelghany, Diana Göhringer |
DSD | 2 |
| 2023 | Compiler-Assisted Kernel Selection for FPGA-based Near-Memory Computing PlatformsabstractThe speed of modern computing systems has improved significantly, thanks to advances in CMOS technology. However, the memory bandwidth of DRAM has not kept pace with these improvements in terms of latency and energy consumption, which is known as the memory wall [1]. FPGAs with high-bandwidth memory (HBM) provide significantly improved performance on memory-intensive tasks, such as graph processing and machine learning. By leveraging 3D-stacked DRAM memory on FPGAs, it is possible to realize the Near-Memory Computing (NMC) paradigm, which involves offloading some kernels to be processed close to the memory. While there have been many studies on NMC accelerators, there is no established method for determining which application kernels are suitable for execution near the HBM. To fully realize the potential of FPGA-HBM architectures, it is important to identify offloading candidates without relying on programmers' knowledge. However, this is a non-trivial task due to the complexity of modern applications. To address this issue, we propose a compiler-assisted tool-flow for the automatic selection of kernels to be offloaded. Veronia Iskandar, Mohamed Abdelghany, Diana Göhringer |
FCCM | 2 |
| 2023 | Performance Estimation and Prototyping of Reconfigurable Near-Memory Computing SystemsabstractThe concept of near-memory computing (NMC) has emerged as a promising solution to address the memory wall challenges faced by future computing architectures. By utilizing modern systems that integrate 3D-stacked DRAM memory, the NMC paradigm minimizes unnecessary data movement between the memory subsystem and the CPU. FPGA vendors have incorporated 3D-stacked memories into their products to meet the increasing bandwidth requirements of memory-intensive applications, enabling FPGAs to compete with GPU solutions in terms of speed and energy efficiency. Recent NMC proposals focus on different data processing workloads, including graph processing and machine learning. This work addresses the research questions of how to leverage the full bandwidth of 3D-stacked high-bandwidth memory and how to facilitate the adoption of the near-memory computing paradigm. Veronia Iskandar, Mohamed Abdelghany, Diana Göhringer |
FPL | 2 |
| 2023 | Near-memory Computing on FPGAs with 3D-stacked Memories: Applications, Architectures, and OptimizationsabstractThe near-memory computing (NMC) paradigm has transpired as a promising method for overcoming the memory wall challenges of future computing architectures. Modern systems integrating 3D-stacked DRAM memory can be leveraged to prevent unnecessary data movement between the main memory and the CPU. FPGA vendors have started introducing 3D memories to their products in an effort to remain competitive on bandwidth requirements of modern memory-intensive applications. Recent NMC proposals target various types of data processing workloads such as graph processing, MapReduce, sorting, machine learning, and database analytics. In this article, we conduct a literature survey on previous proposals of NMC systems on FPGAs integrated with 3D memories. By leveraging the high bandwidth offered from such memories together with specifically designed hardware, FPGA architectures have become a competitor to GPU solutions in terms of speed and energy efficiency. Various FPGA-based NMC designs have been proposed with software and hardware optimization methods to achieve high performance and energy efficiency. Our review investigates various aspects of NMC designs such as platforms, architectures, workloads, and tools. We identify the key challenges and open issues with future research directions. Veronia Iskandar, Mohamed Abdelghany, Diana Göhringer |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2021 | Near-Data-Processing Architectures Performance Estimation and Ranking using Machine Learning PredictorsabstractThe near-data processing (NDP) paradigm has emerged as a promising solution for the memory wall challenges of future computing architectures. Modern 3D-stacked DRAM systems can be exploited to prevent unnecessary data movement between the main memory and the CPU. To date, no standardized simulation frameworks or benchmarks are available for the systematic evaluation of NDP systems. Identifying which type of high-performance 3D memory is suitable to use in an NDP system remains a challenge. This is mainly due to the fact that understanding the interactions between modern workloads and the memory subsystem is not a trivial task. Each memory type has its advantages and drawbacks. Additionally, memory access patterns vary greatly across applications. As a result, the performance of a given application on a given memory type is difficult to intuitively predict. There is no specific memory type that can effectively provide high performance for all applications.In this work, we propose a machine learning framework that can efficiently decide which NDP system is suitable for an application. The framework relies on performance prediction based on an input set of application characteristics. For each NDP system we are examining, we build a machine learning model that can accurately predict performance of previously unseen applications on this system. Our models are on average 200x faster than architectural simulation. They can accurately predict performance with coefficients of determination ranging between 0.88 and 0.92, and root mean square errors ranging between 0.08 and 0.19. Veronia Iskandar, Mohamed Abdelghany, Diana Göhringer |
DSD | 2 |
| 2021 | Formalization of RBD-Based Cause Consequence Analysis in HOL
Mohamed Abdelghany, Sofiène Tahar |
CICM | 1 |
| 2021 | A Review on Charging Systems for Electric Vehicles in Smart Cities
Mohamed Abdelghany |
VEHITS | 1 |
| 2020 | Deep Learning Utilization in Beamforming Enhancement for Medical UltrasoundabstractUltrasound imaging offers a low cost, noninvasive and portable system, which allowed it to be an invaluable tool for medical imaging. However, the quality of the reconstructed images depends significantly on the beamforming technique utilized. Although advanced data-adaptive methods of reconstruction such as Minimum Variance (MV) beamforming can recover image quality much higher than conventional techniques, their implementation also entails a heavy computational burden. This dichotomy hinders the ultrasound imaging use as a standalone device in some applications such as early breast cancer detection. Deep neural networks (DNNs) have shown a huge potential when applied to many Artificial intelligence (AI) research fields. In this work, the use of Deep learning in improving the quality of the beamforming technique Delay and Sum (DAS) normally used for ultrasound (US) images reconstruction is explored. Three different architectures are implemented: Convolutional AutoEncoder (CAE), Fully Connected network (FC) and U-Net-like architecture. They were trained on datasets simulated using field II. The dataset consists of input-output pairs where the input is Noisy DAS beamformed scan lines and the output is MV beamformed non-noisy scan lines. The networks show a great ability in predicting the beamformed signals along with significantly reducing noise in the reconstructed images. Additionally, the proposed networks improve other image characteristics such as scatterer size and position along with reducing tail characteristic normally found in DAS beamformed ultrasound images. US images constructed by the networks achieved better quality metrics that surpass conventional DAS beamformed images. The CAE, U-net-like architecture, and FC enhanced the signal to noise ratio (SNR) compared to DAS by 218%, 165% and 136% respectively. Additionally, the networks showed higher Contrast to Noise Ratio (CNR) and Contrast Ratio (CR) metrics than DAS beamformed signals. Finally, the proposed approach achieves a 60% enhancement in time consumption of image reconstruction compared to MV technique, which allows higher possible frame rate with a comparable outcome. Mariam M. Fouad, Yousef Metwally, Georg Schmitz, Michael Hübner 0001, Mohamed Abdelghany |
COMPSAC | 5 |
| 2020 | HPPT-NoC: A Dark-Silicon Inspired Hierarchical TDM NoC with Efficient Power-Performance TradingabstractNetworks-on-chip (NoCs) acquired substantial advancements as the typical solution for a modular, flexible and high performance communication infrastructure coping with the scalable Multi-/Manycores technology. However, the increasing chip complexity heading towards thousand cores, together with the approaching dark-silicon era, puts energy efficiency as an integral design key for future NoC-based multicores, where NoCs are significantly contributing to the total chip power. In this paper, we propose HPPT-NoC, a dark-silicon inspired energy-efficient hierarchical TDM NoC with online distributed setup-scheme. The proposed network makes use of the dim silicon parts of the chip to hierarchically connect quad-routers units. Normal routers operate at full-chip-frequency at high supply level, and hierarchical routers operate at half-chip-frequency and lower supply voltage with adequate synchronization. Routers follow a proposed TDM architecture that separates the datapath from the control-setup planes. This allows separate clocking and operating supplies between data and control and to keep the control-setup as a single-slot-cycle design independent of the datapath slot size. The proposed NoC architecture is evaluated versus a base NoC from the state-of-the-art in terms of performance and hardware results using Synopsys VCS and Synopsys Design Compiler for SAED90nm and SAED32nm technologies. The obtained results highlight the power-frequency-trading feature supported by the proposed hierarchical NoC through the configurable data-control clock relation and maintained over the different technology nodes. With the same power budget of the base NoC, the proposed architecture provides up to 74% setup latency enhancement, 32% increased NoC saturation load, and 21% higher success rates, offering up to 78% improved power delay product. On the other hand, with 38% power savings, the proposed NoC provides up to 37% enhanced latency and 15% higher success rates, with 72% enhanced power delay product. The proposed design consumes almost double the area of the base NoC, however with an average of 56% under-clocked (dim) silicon area operating at half to quarter the maximum chip frequency. This results in reduced power density as a main concern in the dark-silicon era down to 24% of the base NoC. Salma Hesham, Diana Göhringer, Mohamed Abdelghany |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | FPGA implementation of dynamically reconfigurable IoT security module using algorithm hopping
Shady Mohamed Soliman, Mohammed A. Jaela, Abdelrhman Mohamed Abotaleb, Youssef Hassan, Mohamed Abdelghany, Amr Talaat Abdel-Hamid, Khaled N. Salama, Hassan Mostafa |
Integr. | 5 |
| 2018 | Integrated Sensors for Early Breast Cancer DiagnosticsabstractA wearable, low cost and power efficient early breast cancer detection device is proposed. Bioimpedance spectroscopy (BIS) and near infrared spectroscopy (NIRS) are used. The NIRS and BIS sensors differentiate between normal and cancerous breasts according to their optical and electrical properties respectively. The bioimpedance spectroscopy sensor measurements are carried out by a multi-step frequency sweep. The near infrared spectroscopy sensor uses multi-wavelengths LEDs with optical filters. The results obtained by NIRS and BIS sensors are combined together in the control unit. A custom designed mobile application is connected with the device through Bluetooth. The proposed device was tested in vitro using tissue like mimicking phantoms stimulating the electrical and the optical properties of the normal and cancerous breast tissues. The proposed system shows accuracy of 99.3% and low power consumption of 80mW. Omar Farag, Mariam Mohamed, Mohamed Abdelghany, Klaus Hofmann |
DDECS | 3 |
| 2017 | Survey on Real-Time Networks-on-ChipabstractMulti-Processor Systems-on-Chip (MPSoCs) have emerged as an evolution trend to meet the growing complexity of embedded applications with increasing computation parallelism. Particularly, real-time applications make out a significant portion of the embedded field. Networks-on-Chip (NoCs) are the backbone of communications in an MPSoC platform. However, the use of NoCs in real-time systems imposes complex constraints on the overall design. This paper discusses the challenges faced, when designing NoCs for real-time applications. Contributions in this area are surveyed on the level of guaranteed Quality-of-Service (QoS) support, adaptivity, and energy efficient techniques. Furthermore, the evaluation methodologies and experimental performance measurements of real-time NoCs are examined. This survey provides a comprehensive overview of existing endeavors in real-time NoCs and gives an insight towards future promising research points in this field. Salma Hesham, Jens Rettkowski, Diana Göhringer, Mohamed Abdelghany |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | Real-time sleep detection and warning system to ensure driver's safety based on EEGabstractA Real-Time Sleep Detection and Warning System for Driver's Safety Based on EEG is proposed and implemented to ensure the safety for the drivers and pilots. This system is implemented to estimate and measure the driver attention, the percentage of oxygen in the blood of the driver and to check if driver is failing a sleep. The design and implementation of oxygen saturation sensor is also provided. In addition this system contains a Real time vital signs monitoring system to measure the vital signs values. The proposed system achieved an Accuracy of 96.3%, 100% of sensitivity, 92.4% of Predictability and 93% Specificity. The accuracy, predictability and specificity of the vital signs monitoring system is increased by 2%, 3% and 1%, respectively. Michael S. Saleab, Mohamed Abdelghany, Ramez M. Toma, Klaus Hofmann |
DDECS | 2 |
| 2014 | High throughput architecture for the Advanced Encryption Standard AlgorithmabstractA high throughput architecture is proposed for an efficient implementation of the Advanced Encryption Standard (AES) Algorithm. The presented architecture is adapted for AES encryptor-only as well as integrated AES encryptor/decryptor designs. The SubBytes/InvSubBytes operations are implemented using composite field arithmetic in order to exploit the sub-pipelining advantage within the loop-unrolling methodology. The proposed architecture minimizes the critical path delay through the modification of the SubBytes/InvSubBytes as well as the KeyExpansion modules. Compared to previously reported AES encryptors and integrated AES encryptors/decryptors designs, the proposed architecture provides an efficiency improvement of 61% and 29% respectively. Salma Hesham, Mohamed Abdelghany, Klaus Hofmann |
DDECS | 2 |
| 2013 | Hybrid Mesh-Ring wireless NoC for multi-core systemabstractHybrid network on chip architecture is proposed for high system performance, so with the increase in the number of IP blocks it improves the three main parameters which are throughput, latency and power dissipation better than traditional wired network on chip. Hybrid Mesh-Ring architecture has shown the advantages of using both wired and wireless links in the same network which shows a great performance in Network on Chip (NoC). Two models of Mesh-Ring Architecture are proposed one is based on wired and wireless links which is hybrid model and the other model is based on wired links only which is wired model. The hybrid model has improved the performance in Latency by 20 % as compared to wired model. Throughput has increased by 31 % compared to throughput of wired model and Power-dissipation has decreased by 11 % compared to wired model. Mohamed A. Wanas, Mohamed Abdelghany, Klaus Hofmann |
DDECS | 2 |
| 2012 | CDMA technique for Network-on-ChipabstractA Code-Division Multiple Access (CDMA) based on-chip communication network is proposed in this paper. The proposed design features a novel encoding and decoding scheme for CDMA transmission which improves area, latency and power dissipation of the network on Chip (NoC). The orthogonal and balance properties of Walsh codes are used for the routing of data between the resources on the network. The proposed CDMA encoding and decoding schemes are compared with the conventional schemes. The overall area required to implement the proposed CDMA NoC design is reduced by 54%. The design decreases the latency of the network by 48.2%. The total power consumption required to achieve the proposed design is decreased by 54.8%. Ahmed A. El Badry, Mohamed Abdelghany |
DDECS | 2 |
| 2012 | A simulation framework for 3-dimension Networks-on-chip with different vertical channel density configurationsabstract3D ICs are emerging as a promising solution for scalability, power and performance demands of next generation Systems-on-Chip (SoCs). Along with the advantages, it also imposes a number of challenges with respect to cost, technological reliability, thermal budget and so forth. Networks-on-chip (NoCs), which is thoroughly investigated in 2D SoCs design as scalable interconnects, is also well relevant to 3D IC Design. The cost of moving from 2D to 3D should be justified with improvements in performance, power or latency. To solve this problem, this paper presents a new simulation framework for 3D NoCs. We established a new Generic Scalable Pseudo Application (GSPA), where user can generate their own scalable pseudo applications. We have also integrated the state-of-the-art benchmarks to evaluate the 3D NoC system. In the framework, the 3D NoC with different vertical channel densities (VD) (i.e. number of Through-Silicon-Vias (TSVs)) can be generated according to the preference of users. After the simulation, the power consumption and system performance are evaluated. We have compared 2D NoC architecture with 3D NoC architecture with different VDs. The experimental results show that 3D architectures have significant advantage (Avg. 51%, 44%, 35% for 100%, 50%, 25% VD, respectively) in the aspect of interconnect power delay product in comparison to 2D mesh architecture. The 25% VD architecture is the best choice with 17% advantage over full connection (100% VD) 3D NoC architecture in the aspect of Figure of Merit which takes area and TSV connection yield into account among all the experiments for the given constrains. Haoyuan Ying, Ashok Jaiswal, Mohamed Abdelghany, Thomas Hollstein, Klaus Hofmann |
DDECS | 3 |
| 2010 | Asynchronous BFT for low power networks on chipabstractAsynchronous Butterfly Fat Tree (BFT) architecture is proposed to achieve low power Network on Chip (NoC). Asynchronous design could reduce the power dissipation of the network if the activity factor of the data transfer between two switches (αdata) satisfies a certain condition. The area of Asynchronous BFT switch is increased by 25% as compared to Synchronous switch. However, the power dissipation of the Asynchronous architecture could be decreased by up to 33% as compared to the power dissipation of the conventional Synchronous architecture when the αdataequals 0.2 and the activity factor of the control signals is equal to 1/64 of the αdata. The total metal resources required to implement Asynchronous design is decreased by 12%. Mohamed Abdelghany, Magdy A. El-Moursy, Darek Korzec, Mohammed Ismail 0001 |
ISCAS | 1 |
| 2010 | Power characteristics of Networks on ChipabstractPower characteristics of different Network on Chip (NoC) topologies are developed. Among different NoC topologies, the Butterfly Fat Tree (BFT) dissipates the minimum power. With the advance in technology, the relative power consumption of the interconnects and the associate repeaters of the BFT decreases as compared to the power consumption of the network switches. The power dissipation of interswitch links and repeaters for BFT represents only 1% of the total power dissipation of the network. In addition of providing high throughput, the BFT is a power efficient topology for NoCs. Mohamed Abdelghany, Magdy A. El-Moursy, Darek Korzec, Mohammed Ismail 0001 |
ISCAS | 1 |
| 2009 | High Throughput Architecture for High Performance NoCabstractHigh Throughput Butterfly Fat Tree (HTBFT) architecture to achieve high performance Networks on Chip (NoC) is proposed. The architecture increases the throughput of the network by 38% while preserving the average latency. The area of HTBFT switch is decreased by 18% as compared to Butterfly Fat Tree switch. The total metal resources required to implement HTBFT design is increased by 5% as compared to the total metal resources required to implement BFT design. The extra power consumption required to achieve the proposed architecture is 3% of the total power consumption of the BFT architecture. Mohamed Abdelghany, Magdy A. El-Moursy, Mohammed Ismail 0001 |
ISCAS | 1 |
| 2007 | Design and Implementation of FPGA-based Systolic Array for LZ Data CompressionabstractHardware Implementation of Data compression algorithms is receiving increasing attention due to exponentially expanding network traffic and digital data storage usage. Among lossless data compression algorithms for hardware implementation, Lempel -Ziv algorithm is one of the most widely used. The main objective of this paper is to enhance the efficiency of systolic- array approach for implementation of Lempel- Ziv algorithm. The proposed implementation is area and speed efficient. The compression rate is increased by more than 40% and the design area is decreased by more than 30%. The effect of the selected buffer's size on the compression ratio is analyzed. An FPGA implementation of the proposed design is carried out. It verifies that data can be compressed and decompressed on-the-fly. Mohamed Abdelghany, Aly E. Salama, Ahmed Hussein Khalil |
ISCAS | 1 |