Ali El-Moursy

dblp:99/6710 · also Ali A. El-Moursy · DBLP profile ↗
← Back
20ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0002-3660-6544ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 6 first-author · 4 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 VFCkM: a federated clustering framework based on k-means algorithm for vertically partitioned data with shared attributes
Oruba Alfawaz, Ali El-Moursy, Mohamed Saad 0001, Ahmed Khedr 0001
J. Supercomput.2
2025 MOTO-MASSA: multi-objective task offloading based on modified sparrow search algorithm for fog-assisted IoT applications
Ahmed Khedr 0001, Oruba Alfawaz, Marya Alseid, Ali El-Moursy
Wirel. Networks4
2024 E2CSM: efficient FPGA implementation of elliptic curve scalar multiplication over generic prime field GF(p)
Khalid Javeed, Ali El-Moursy, David Gregg
J. Supercomput.2
2023 Fog-ROCL: A Fog based RSU Optimum Configuration and Localization in VANETs
Rehab Shahin, Sherif M. Saif, Ali El-Moursy, Hazem M. Abbas, Salwa M. Nassar
Pervasive Mob. Comput.3
2023 MSSAMTO-IoV: modified sparrow search algorithm for multi-hop task offloading for IoV
Marya Alseid, Ali El-Moursy, Oruba Alfawaz, Ahmed Khedr 0001
J. Supercomput.2
2022 Wireless link scheduling via parallel genetic algorithm
abstract
Abstract With the advent of fifth generation (5G) systems and the Internet‐of‐Things (IoT), the number of interconnected wireless devices is increasing significantly. Protocols that allow these deceives to interconnect peer‐to‐peer through wireless links are becoming of interest. The major challenge is the inevitable interference among the simultaneously activated wireless links. Given a set of wireless links, this article addresses the non‐deterministic polynomial‐time (NP) hard problem of selecting the maximum subset of links that can be simultaneously activated at their respective signal‐to‐interference‐plus‐noise‐ratio (SINR) targets. The contribution of this article is two‐fold. First, we introduce a new genetic algorithm (GA) constraint‐handling mechanism, and prove analytically that finding optimal link schedules is guaranteed. Second, we develop a novel parallelized GA to solve the problem. Through serial algorithm analysis, we utilize data decomposition as well as exploratory decomposition in order to achieve significant running time speedup, which scales well with problem size. Our numerical results for openMP parallelization illustrate 6.5 and 5.4 reduction in computation time as compared to the serial versions of the GA and hybrid genetic algorithm (HGA), respectively. Moreover, the parallelization of the GA and HGA result in a speedup of 10.4 and 5.4 , respectively, using master‐slave multithreading.
Mohamed Saad 0001, Ali El-Moursy, Oruba Alfawaz, Khawla Alnajjar, Saeed Abdallah
Concurr. Comput. Pract. Exp.2
2022 Hardware Acceleration of the STRIKE String Kernel Algorithm for Estimating Protein to Protein Interactions
abstract
Protein-protein interaction (PPI) is an important field in bioinformatics which helps in understanding diseases and devising therapy. PPI aims at estimating the similarity of protein sequences and their common regions. STRIKE was introduced as a PPI algorithm which was able to achieve reasonable improvement over existing PPI prediction methods. Although it consumes a lower execution time than most of other state-of the-art PPI prediction methods, its compute-intensive nature and the large volume of protein sequences in protein databases necessitate further time acceleration. In this paper, we develop hardware accelerator designs for the STRIKE algorithm. Results indicate that the weighted STRIKE accelerator execution times are about 10x longer than the unweighted STRIKE accelerator execution times. To further accelerate the performance of the weighted STRIKE, a parallel module accelerator organization duplicating the weighted STRIKE modules is introduced, achieving near linear speedups for long sequences of 100 or more characters. As demonstrated by Verilog simulations and FPGA runs, the weighted STRIKE module accelerator exhibits three orders of magnitude speed improvement over multi-core and cluster computers. Much higher speedups are possible with the parallel module accelerator.
Fadi N. Sibai, Ali El-Moursy, Abu Asaduzzaman, Sohaib Majzoub
IEEE ACM Trans. Comput. Biol. Bioinform.2
2021 Joint timing-offset and channel estimation for physical layer network coding in frequency selective environments
abstract
Abstract This paper considers the problem of joint timing‐offset and channel estimation for physical‐layer network coding systems operating in frequency‐selective environments. Three different algorithms are investigated for the joint estimation of the channel coefficients and the fractional timing offset. The first algorithm is based on the maximum‐likelihood (ML) criterion assuming baud‐rate (BR) sampling. The second algorithm also assumes BR sampling and is based on the special properties of Zadoff‐Chu training sequences. In the third algorithm, oversampling at double the baud‐rate (DBR) is used and the least‐squares (LS) estimation criterion applied. While the above algorithms assume that the integer timing offset is known, three generalized‐likelihood‐ratio tests (GLRTs) are also considered for integer offset error correction that integrate very well with the proposed estimation algorithms. Our simulation studies show that the DBR‐LS estimator provides the highest estimation accuracy, significantly outperforming both BR estimators and performing very close to the corresponding Cramer–Rao bound. A gain of 4 dB is observed in symbol‐error‐rate performance using the DBR‐LS algorithm. The DBR‐GLRT also provides substantially higher probability of error correction.
Saeed Abdallah, Mohamed Saad 0001, Khawla Alnajjar, Ali El-Moursy
IET Commun.4
2020 A Dynamic Clustering Approach for Increasing User Throughput in 5G Wireless Networks
abstract
In cellular networks, user density and activity across network area vary from cell to another, and even at the same area over the day hours. This results in an irregular utilization of the network resources (i.e., resource blocks "RBs") across the network cells. One of the solutions to improve network utilization is the coordinated multi-point (CoMP) approach which coordinates signal transmission and reception for a user to and from multiple base stations (BSs). However, providing CoMP to all network users could waste RBs. To solve this problem, a novel clustering technique is proposed in this paper to guarantee a wise CoMP service to users based on their area density state (i.e., overloaded or under-loaded) and across different day periods. Through simulations, the proposed dynamic clustering algorithm outperforms other clustering algorithms in improving overall user throughput while maintaining the total consumed energy to acceptable levels.
Wael S. Afifi, Ali El-Moursy, Mohamed Saad 0001, Salwa M. Nassar, Hadia M. El-Hennawy
NOMS2
2020 PMSMC: Priority-based Multi-requestor Scheduler for Embedded System Memory Controller
Ali El-Moursy, Fadi N. Sibai, Magdy A. El-Moursy, Ahmed S. S. Mohamed
J. Parallel Distributed Comput.1
2019 Adaptive TB-LMI: An efficient memory controller and scheduler design
abstract
Summary In the modern multi‐core systems, concurrently executing applications share common resource such as main memory. Memory scheduling algorithms are developed to resolve memory contention among competing applications so that throughput is high and fairness of the overall multi‐core system is guaranteed. In this paper, we present Adaptive Time‐based Least Memory Intensive (Adaptive TB‐LMI) scheduling, a new memory scheduling algorithm that addresses both fairness and system performance. Adaptive TB‐LMI is based on TB‐LMI which prioritizes applications according to their memory contention every pre‐defined CPU cycle. Adaptive TB‐LMI dynamically prioritizes applications according to their memory contention. Considering the previous algorithms with the best performance, for 16‐core system, TB‐LMI improves system throughput on average by 2.25X and 36% comparing to FCFS and TCM respectively. Adaptive TB‐LMI is 6% better than the TB‐LMI with static threshold. In terms of fairness and slowdown metrics, TB‐LMI show improvements of 30% and 3X, respectively, compared to FCFS and, 18% and 8%, respectively, compared to TCM. Adaptive TB‐LMI and TB‐LMI are 15% more efficient in energy‐delay product, although they are within 5% in terms of area overhead. Moreover, TCM has an area overhead of about 45% more than Adaptive TB‐LMI. In terms of Energy‐Delay Product, Adaptive TB‐LMI is 10% and 24% better than TB‐LMI and TCM, respectively. This is due to the dynamic capabilities of the adaptive algorithm to change the rate of the SQ, hence reducing the energy consumption.
Ali El-Moursy, Amr Saleh Elhelw
Concurr. Comput. Pract. Exp.1
2019 CFPA: Congestion aware, fault tolerant and process variation aware adaptive routing algorithm for asynchronous Networks-on-Chip
Sayed Taha Muhammad, Mohamed Saad 0001, Ali El-Moursy, Magdy A. El-Moursy, Hesham F. A. Hamed
J. Parallel Distributed Comput.3
2017 Architecture level analysis for process variation in synchronous and asynchronous Networks-on-Chip
Sayed Taha Muhammad, Magdy A. El-Moursy, Ali El-Moursy, Hesham F. A. Hamed
J. Parallel Distributed Comput.3
2015 Traffic-Based Virtual Channel Activation for Low-Power NoC
abstract
A large amount of leakage power could be saved by increasing the number of idle virtual channels (VCs) in a network-on-chip (NoC). Low-leakage power switch is proposed to allow saving in power dissipation of the NoC. The proposed NoC switch employs power supply gating to reduce the power dissipation. Two power reduction techniques are exploited to design the proposed switch. Adaptive virtual channel technique is proposed as an efficient technique to reduce the active area using hierarchical multiplexing tree. Moreover, power gating (PG) reduces the average leakage power consumption of proposed switch. The proposed techniques save up to 97% of the switch leakage power. In addition, the dynamic power is reduced by 40%. The traffic-based virtual channel activation (TVA) algorithm is used to determine the traffic status and send adaptation signals to PG units to activate/deactivate the VCs. The TVA algorithm optimally utilizes VCs by deactivating idle VCs to guarantee high-leakage power saving with high throughput. TVA is an efficient and flexible algorithm that defines a set of parameters to be used to achieve minimum degradation in NoC throughput with maximum reduction in leakage power. The whole network average leakage power has been reduced by up to 80% for 2-D-mesh NoC with throughput degradation within only 1%. For 2-D-torus NoC, a saving in power of up to 84% is achieved with <;2% degradation in throughput. The implementation overhead of TVA is negligible.
Sayed Taha Muhammad, Rabab Ezz-Eldin, Magdy A. El-Moursy, Ali El-Moursy, Amr M. Refaat
IEEE Trans. Very Large Scale Integr. Syst.4
2014 High-accuracy hierarchical parallel technique for hidden Markov model-based 3D magnetic resonance image brain segmentation
abstract
ABSTRACT Scientific applications represent a dominant sector of compute‐intensive applications. Using massively parallel processing systems increases the feasibility to automate such applications because of the cooperation among multiple processors to perform the designated task. This paper proposes a parallel hidden Markov model (HMM) algorithm for 3D magnetic resonance image brain segmentation using two approaches. In the first approach, a hierarchical/multilevel parallel technique is used to achieve high performance for the running algorithm. This approach can speed up the computation process up to 7.8× compared with a serial run. The second approach is orthogonal to the first and tries to help in obtaining a minimum error for 3D magnetic resonance image brain segmentation using multiple processes with different randomization paths for cooperative fast minimum error convergence. This approach achieves minimum error level for HMM training not achievable by the serial HMM training on a single node. Then both approaches are combined to achieve both high accuracy and high performance simultaneously. For 768 processing nodes of a Blue Gene system, the combined approach, which uses both methods cooperatively, can achieve high‐accuracy HMM parameters with 98% of the error level and 2.6× speedup compared with the pure accuracy‐oriented approach alone. Copyright © 2012 John Wiley & Sons, Ltd.
Ali El-Moursy, Hanan H. Elazhary, Akmal Younis
Concurr. Comput. Pract. Exp.1
2011 Image processing applications performance study on Cell BE and Blue Gene/L
abstract
Abstract Two image processing applications, edge detection and image resizing, are studied in this paper on two HPC platforms namely the Cell BE and the Blue Gene/L machines. In this paper we focus on the performance scalability of the studied applications. Our results show that the scale of the problem to be solved highly affects the fitness of the platform. If the data set size is to fit into the Cell core, the fast on‐chip inter‐core communication of a multi‐core system pays back for its high technology design. On the other hand, the overhead of the distant communication in the massively parallel Blue Gene/L machine will only show its benefits for huge data set size that otherwise mandates multiple round‐trip data communications between the local memory of a core and main memory. Copyright © 2010 John Wiley & Sons, Ltd.
Ali El-Moursy, Fadi N. Sibai
Concurr. Comput. Pract. Exp.1
2006 Compatible phase co-scheduling on a CMP of multi-threaded processors
abstract
The industry is rapidly moving towards the adoption of chip multi-processors (CMPs) of simultaneous multi-threaded (SMT) cores for general purpose systems. The most prominent use of such processors, at least in the near term, is as job servers running multiple independent threads on the different contexts of the various SMT cores. In such an environment, the co-scheduling of phases from different threads plays a significant role in the overall throughput. Less throughput is achieved when phases from different threads that conflict for particular hardware resources are scheduled together, compared with the situation where compatible phases are co-scheduled on the same SMT core. Achieving the latter requires precise per-phase hardware statistics that the scheduler can use to rapidly identify possible incompatibilities among phases of different threads, thereby avoiding the potentially high performance cost of inter-thread contention. In this paper, we devise phase co-scheduling policies for a dual-core CMP of dual-threaded SMT processors. We explore a number of approaches and find that the use of ready and in-flight instruction metrics permits effective co-scheduling of compatible phases among the four contexts. This approach significantly outperforms the worst static grouping of threads, and very closely matches the best static grouping, even outperforming it by as much as 7%.
Ali El-Moursy, Rajeev Garg, David H. Albonesi, Sandhya Dwarkadas
IPDPS1
2005 Partitioning Multi-Threaded Processors with a Large Number of Threads
abstract
Today's general-purpose processors are increasingly using multithreading in order to better leverage the additional on-chip real estate available with each technology generation. Simultaneous multi-threading (SMT) was originally proposed as a large dynamic superscalar processor with monolithic hardware structures shared among all threads. Inters hyper-threaded Pentium 4 processor partitions the queue structures among two threads, demonstrating more balanced performance by reducing the hoarding of structures by a single thread. IBM's Power5 processor is a 2-way chip multiprocessor (CMP) of SMT processors, each supporting 2 threads, which significantly reduces design complexity and can improve power efficiency. This paper examines processor partitioning options for larger numbers of threads on a chip. While growing transistor budgets permit four and eight-thread processors to be designed, design complexity, power dissipation, and wire scaling limitations create significant barriers to their actual realization. We explore the design choices of sharing, or of partitioning and distributing, the front end (instruction cache, instruction fetch, and dispatch), the execution units and associated state, as well as the L1 Dcache banks, in a clustered multi-threaded (CMT) processor. We show that the best performance is obtained by restricting the sharing of the L1 Dcache banks and the execution engines among threads. On the other hand, significant sharing of the front-end resources is the best approach. When compared against large monolithic SMT processors, a CMT processor provides very competitive IPC performance on average, 90-96% of that of partitioned SMT while being more scalable and much more power efficient. In a CMP organization, the gap between SMT and CMT processors shrinks further, making a CMP of CMT processors a highly viable alternative for the future
Ali El-Moursy, Rajeev Garg, David H. Albonesi, Sandhya Dwarkadas
ISPASS1
2003 Front-End Policies for Improved Issue Efficiency in SMT Processors
abstract
The performance and power optimization of dynamic superscalar microprocessors requires striking a careful balance between exploiting parallelism and hardware simplification. Hardware structures which are needlessly complex may exacerbate critical timing paths and dissipate extra power. One such structure requiring careful design is the issue queue. In a simultaneous multi-threading (SMT) processor it is particularly challenging to achieve issue queue simplification due to the increased utilization of the queue afforded by multi-threading. In this paper we propose new front-end policies that reduce the required integer and floating point issue queue sizes in SMT processors. We explore both general policies as well as those directed towards alleviating a particular cause of issue queue inefficiency. For the same level of performance, the most effective policies reduce the issue queue occupancy by 33% for an SMT processor with appropriately sized issue queue resources.
Ali El-Moursy, David H. Albonesi
HPCA1
2003 1-V ADPCM Processor for Low-Power Wireless Applications
Martin Margala, Magdy A. El-Moursy, Ali El-Moursy, Junmou Zhang, Wendi B. Heinzelman
VLSI-SOC3