Ahmad Salah

dblp:128/1447 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
8since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Systems, architecture and hardware · 5 · 5 first-authorDatabases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2023 Real-Time and Automatic System for Performance Evaluation of Karate Skills Using Motion Capture Sensors and Continuous Wavelet Transform
abstract
In sports science, the automation of performance analysis and assessment is urgently required to increase the evaluation accuracy and decrease the performance analysis time of a subject. Existing methods of performance analysis and assessment are either performed manually based on human experts’ opinions or using motion analysis software, i.e., biomechanical analysis software, to assess only one side of a subject. Therefore, we propose an automated system for performance analysis and assessment that can be used for any human movement. The performance of any skill can be described by a curve depicting the joint angle over the time required to perform a skill. In this study, we focus on only 14 body joints, and each joint comprises three angles. The proposed system comprises three main stages. In the first stage, data are obtained using motion capture inertial measurement unit sensors from top professional fighters/players while they are performing a certain skill. In the second stage, the collected sensor data obtained are input to the biomechanical software to extract the player’s joint angle curve. Finally, each joint angle curve is processed using a continuous wavelet transform to extract the main curve points (i.e., peaks and valleys). Finally, after extracting the joint curves from several top players, we summarize the players’ curves based on five statistical indicators, i.e., the minimum, maximum, mean, and mean ± standard deviation. These five summarized curves are regarded as standard performance curves for the joint angle. When a player’s joint curve is surrounded by the five summarized curves, the performance is considered acceptable. Otherwise, the performance is considered unsatisfactory. The proposed system is evaluated based on four different karate skills. The results of the proposed system are identical to the decisions of the expert panels and are thus suitable for real‐time decisions.
Ahmed Fathalla, Ahmad Salah, Mahmoud Bekhit, Esraa Eldesouky, Ahmed Talha, Abdalla Zenhom
Int. J. Intell. Syst.2
2023 Virtual Machine Replica Placement Using a Multiobjective Genetic Algorithm
abstract
Virtual machine (VM) replication is a critical task in any cloud computing platform to ensure the availability of the cloud service for the end user. In this task, one primary VM resides on a physical machine (PM) and one or more replicas reside on separate PMs. In cloud computing, VM placement (VMP) is a well‐studied problem in terms of different goals, such as power consumption reduction. The VMP problem can be solved by using heuristics, namely, first‐fit and meta‐heuristics such as the genetic algorithm. Despite extensive research into the VMP problem, there are few works that consider VM replication when choosing a VMP. In this context, we proposed studying the problem of optimal VMP considering VM replication requirements. The proposed work frames the problem at hand as a multiobjective problem and adapts a nondominated sorting genetic algorithm (NSGA‐III) to address the problem. VM replicas’ placement should consider several dimensions such as the geographical distance between the PM hosting the primary VM and the other PMs hosting the replicas. In addition, to this end, the proposed model aims to minimize (1) power consumption, (2) performance degradation, and (3) the distance between the PMs hosting the primary VM and its replica(s). The proposed method is thoroughly tested on a variety of computing environments with various heterogeneous VMs and PMs, including compute‐intensive and memory‐intensive environments. The obtained results illustrate the performance disparity between the adapted NSGA‐III and MOEA/D methods and other methods of comparison, including heuristic and meta‐heuristic approaches, with NSGA‐III outperforming other comparison methods. For instance, in memory‐intensive and in heterogeneous environments, the NSGA‐III method’s performance was superior to the first‐fit, next‐fit, best‐fit, PSO, and MOEA/D methods by 58%, 62%, 64%, 55%, and 31%, respectively.
Marwa F. Mohamed, Mai Dahshan, Kenli Li 0001, Ahmad Salah
Int. J. Intell. Syst.4
2023 Forecasting the Friction Coefficient of Rubbing Zirconia Ceramics by Titanium Alloy
abstract
The thermal issues generated from friction are the key obstacle in the high‐performance machining of titanium alloys. The friction between the workpiece being cut and the cutting tool is the dominant parameter that affects the heat generation during the machining processes, i.e., the temperature inside the cutting zone and the consumed cutting energy. Besides, the complexity is associated with the nature of the friction phenomenon. However, there are limited efforts to forecast the friction coefficient during the machining operations. In this work, the friction coefficients between the titanium alloy against zirconia ceramics lubricated by minimum quantity lubrication were recorded and measured using a universal mechanical tester pin‐on‐disc tribometer. Then, we proposed two models for forecasting the friction coefficient which are trained and tested on the recorded data. The two predictive models are based on autoregressive integrated moving average and gated recurrent unit deep neural network methods. The proposed models are evaluated through a set of exhaustive experiments. These experiments demonstrated that the proposed models can efficiently be used to reduce power consumption dedicated to monitoring the friction coefficients. Besides, they can reduce or avoid surface thermal damage by predicting the high level of friction coefficients in advance, which can be used as an alert to enable or readjust the lubrication parameters (fluid pressure, fluid flow rate, etc.) to maintain lower ranges of friction coefficients and power consumption.
Ahmad Salah, Ahmed Fathalla, Esraa Eldesouky, Wei Li 0333, Ahmed Mohamed Mahmoud Ibrahim
Int. J. Intell. Syst.1
2023 Virtual machine placement in service-oriented computing environments
Asma Alkalbani, Khalil Al Ruqeishi, Ahmad Salah, Marwa F. Mohamed
Serv. Oriented Comput. Appl.3
2022 An LSTM-based distributed scheme for data transmission reduction of IoT systems
Ahmed Fathalla, Kenli Li 0001, Ahmad Salah, Marwa F. Mohamed
Neurocomputing3
2022 Robust color image watermarking using multi-core Raspberry pi cluster
abstract
Abstract Image authentication approaches have gotten a lot of interest recently as a way to safeguard transmitted images. Watermarking is one of the many ways used to protect transmitted images. Watermarking systems are pc-based that have limited portability that is difficult to use in harsh environments as military use. We employ embedded devices like Raspberry Pi to get around the PC’s mobility limitations. Digital image watermarking technology is used to secure and ensure digital images’ copyright by embedding hidden information that proves its copyright. In this article, the color images Parallel Robust watermarking algorithm using Quaternion Legendre-Fourier Moment (QLFM) in polar coordinates is implemented on Raspberry Pi (RPi) platform with parallel computing and C++ programming language. In the host image, a binary Arnold scrambled image is embedded. Watermarking algorithm is implemented and tested on Raspberry Pi model 4B. We can combine many Raspberry Pi’s into a ‘cluster’ (many computers working together as one) for high-performance computation. Message Passing Interface (MPI) and OpenMP for parallel programming to accelerate the execution time for the color image watermarking algorithm implemented on the Raspberry Pi cluster.
Khalid M. Hosny, Amal Magdi, Nabil A. Lashin, Osama El-Komy, Ahmad Salah
Multim. Tools Appl.5
2021 A whale optimization system for energy-efficient container placement in data centers
Almoalmi Ammar, Juan Luo, Ahmad Salah, Kenli Li 0001, Luxiu Yin
Expert Syst. Appl.3
2021 Efficient index-independent approaches for the collective spatial keyword queries
Zhibang Yang, Yifu Zeng, Jiayi Du, Fangmin Li, Ahmad Salah
Neurocomputing5
2020 Accelerated CPU-GPUs implementations for quaternion polar harmonic transform of color images
Ahmad Salah, Kenli Li 0001, Khalid M. Hosny, Mohamed M. Darwish, Qi Tian 0001
Future Gener. Comput. Syst.1
2020 Efficient scientific workflow scheduling for deadline-constrained parallel tasks in cloud computing environments
Longxin Zhang, Liqian Zhou, Ahmad Salah
Inf. Sci.3
2020 Deep end-to-end learning for price prediction of second-hand items
Ahmed Fathalla, Ahmad Salah, Kenli Li 0001, Keqin Li 0001, Francesco Piccialli
Knowl. Inf. Syst.2
2020 Editorial introduction: special issue on advances in parallel and distributed computing for neural computing
Ahmad Salah
Neural Comput. Appl.2
2016 Impedance matrix analysis technique in wound rotor induction machines including general rotor asymmetry
abstract
A steady state analysis is developed for a wound rotor induction machine. In this paper the authors derive simple expressions for the mutual and coupling impedance in order to build impedance matrix. The analysis can be used in two ways when dealing with a symmetrical and asymmetrical rotor. The analysis is verified using a wound-rotor motor; the analysis can then be used to investigate asymmetry faults in a wound rotor machine. Experimental results (torque and current characteristic) are compared with computer predictions for the machine with both open-circuit and short-circuit faults.
Ahmad Salah, Youguang Guo, David G. Dorrell
IECON1
2016 Lazy-Merge: A Novel Implementation for Indexed Parallel K-Way In-Place Merging
abstract
Merging sorted segments is a core topic of fundamental computer science that has many different applications, such as n-body simulation. In this research, we propose Lazy-Merge, a novel implementation of sequential in-place k-way merging algorithms, that can be utilized in their parallel counterparts. The implementation divides the k-way merging problem into t ordered and independent smaller k-way merging tasks (partitions), but each merging task includes a set of scattered ranges to be merged by an existing merging algorithm. The final merged list includes ranges with ordered elements, but the ranges themselves are not ordered. Lazy-Merge utilizes a novel usage of indexes to access the entire set of merged elements in order. Its merging time complexity is O(k log (n/k) + merge(n/p)), where k, n, and pare the number of segments, the list size and the number of processors (partitions), respectively. Here, merge(n/p) represents the time needed to merge n/p elements by the used in-place merging algorithm. The time complexity of accessing an element in the merged list is O(log k), that time can be constant if k processors are used. The results of the proposed work are compared with those of bitonic merge and the best time-space optimal algorithms on number of moves and execution time. In comparison with the existing algorithms, significant speedup and reasonable reduction factor for number of moves have been achieved.
Ahmad Salah, Kenli Li 0001, Keqin Li 0001
IEEE Trans. Parallel Distributed Syst.1
2015 A Framework to Accelerate Protein Structure Comparison Tools
abstract
At the center of computational structural biology, protein structure comparison is a key problem. The steady increase in the number of protein structures encourages the development of massively parallel tools. While the focus of research is to propose data-analytical methods to tackle this problem, there are limited research proposing generic tools to run these methods in parallel environments. Herein, we propose a scalable framework to handle this steady increase. The proposed framework runs the sequential tools on parallel environments. It is a GUI-based and requiring no scripting or installation procedures. The framework includes optimally distributing protein structure database over the existing computing resources, tracking the remote processes course of execution, and merging the results to form the final output. The first stage realizes the biological database distribution as an optimization problem in order to maximize the cluster resources utilization and minimize the execution time. The experimental results show linear and nearly optimal speedups with no loss in accuracy. The framework is available at http://biocloud.hnu.edu.cn/ppsc/.
Ahmad Salah, Kenli Li 0001, Tarek F. Gharib
CCGRID1
2015 Parallel Implementation of MAFFT on CUDA-Enabled Graphics Hardware
abstract
Multiple sequence alignment (MSA) constitutes an extremely powerful tool for many biological applications including phylogenetic tree estimation, secondary structure prediction, and critical residue identification. However, aligning large biological sequences with popular tools such as MAFFT requires long runtimes on sequential architectures. Due to the ever increasing sizes of sequence databases, there is increasing demand to accelerate this task. In this paper, we demonstrate how graphic processing units (GPUs), powered by the compute unified device architecture (CUDA), can be used as an efficient computational platform to accelerate the MAFFT algorithm. To fully exploit the GPU's capabilities for accelerating MAFFT, we have optimized the sequence data organization to eliminate the bandwidth bottleneck of memory access, designed a memory allocation and reuse strategy to make full use of limited memory of GPUs, proposed a new modified-run-length encoding (MRLE) scheme to reduce memory consumption, and used high-performance shared memory to speed up I/O operations. Our implementation tested in three NVIDIA GPUs achieves speedup up to 11.28 on a Tesla K20m GPU compared to the sequential MAFFT 7.015.
Xiangyuan Zhu, Kenli Li 0001, Ahmad Salah, Keqin Li 0001
IEEE ACM Trans. Comput. Biol. Bioinform.3
2014 PAR-3D-BLAST: A parallel tool for searching and aligning protein structures
abstract
SUMMARY Protein structure comparison is a vital process in several tasks like the prediction of protein structures and functions and detecting the proteins evolutionary relationships. The expansion of both the parallel computational hardware and the discovered protein structures stimulates the growth of the parallel computational tools to handle this massive data of proteome. Here, we present a parallel tool, parallel 3D‐BLAST (PAR‐3D‐BLAST), which lists the similar structures to the query protein. Each protein in the result list has a structural similarity score and an alignment to the query structure. The presented tool is implemented to fit both the standalone multi‐core computers and clusters of multi‐core nodes. The achieved speedup is linear and scalable. The experimental results outline that the speedup increases as the size of the database increases. Using a cluster of 35 computing cores, the tool constructs the database of the entire structural classification of proteins dataset, 108,116 protein entries, in less than 6min and with average query time of 1.45s. The obtained speed up is 20 times for database construction and 17 times for searching the query. The tool is an open source and free to use, distribute, and share; it is available at http://aca.hnu.cn/par3dblast . Copyright © 2013 John Wiley & Sons, Ltd.
Ahmad Salah, Kenli Li 0001
Concurr. Comput. Pract. Exp.1
2013 A data decomposition middleware tool with a generic built-in work-flow
abstract
The steady increase of the biological data encourages computer scientists to develop data-analytical methods in order to study the biological systems. Most of these methods are firstly presented in the sequential version and then it is converted to a parallel version due to the constant expanding of data size. In this paper, we present a tool to eliminate the time gap between the two versions, sequential and parallel. The proposed tool acts as a middleware with a generic built in work-flow. The middleware manages a set of instances of the same sequential method; it manages to assign the data to each instance, to merge the results of all instances, and to form the final results. The proposed tool is applicable for problems with data decomposition nature (e.g., protein structure comparison, sequence comparison and CpG Islands search). It targets two levels of parallelism, which are the homogenous multi-core processor architectures and networks of heterogeneous computational nodes. The experimental results, of different computational biology problems, show a speedup close to the optimal with identical accuracy.
Ahmad Salah, Kenli Li 0001
EuroMPI1