Walid Gomaa 0001

dblp:37/7589 · also Walid E. Gomaa, Walid El-Sayed Ali Gomaa · DBLP profile ↗
← Back
66ranked-venue papers
7as first author
24since 2021 · last 2025
0000-0002-8518-8908ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 50 · 6 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorTheory of computation · 4Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Media Bias in the Guardian: Sentiment and Semantics in Israeli-Palestinian Coverage
Ahmed Gamea, Ahmed Hasan, Abdelaziz Serour, Moustafa Rezk, Walid Gomaa 0001
CHIRA (2)5
2025 Video-Based Vibration Analysis for Predictive Maintenance: A Motion Magnification and Random Forest Approach
Walid Gomaa 0001, Abdelrahman Wael Ammar, Ismael Abbo, Mohamed Galal Nassef, Tetsuji Ogawa, Mohab Hossam
ICINCO (1)1
2025 ALEX-GYM-1: A Novel Dataset and Hybrid 3D Pose Vision Model for Automated Exercise Evaluation
Abdelaziz Serour, Ahmed Gamea, Walid Gomaa 0001
ICINCO (2)4
2025 Context-Based Meta Reinforcement Learning for Robust and Adaptable Peg-in-Hole Assembly Tasks
abstract
Autonomous assembly is an essential capability for industrial and service robots, with Peg-in-Hole (PiH) insertion being one of the core tasks. However, PiH assembly in unknown environments is still challenging due to uncertainty in task parameters, such as the hole position and orientation, resulting from sensor noise. Although context-based meta reinforcement learning (RL) methods have been previously presented to adapt to unknown task parameters in PiH assembly tasks, the performance depends on a sample-inefficient procedure or human demonstrations. Thus, to enhance the applicability of meta RL in real-world PiH assembly tasks, we propose to train the agent to use information from the robot’s forward kinematics and an uncalibrated camera. Furthermore, we improve the applicability by efficiently adapting the meta-trained agent to use data from force/torque sensor. Finally, we propose an adaptation procedure for out-of-distribution tasks whose parameters are different from the training tasks. Experiments on simulated and real robots prove that our modifications enhance the sample efficiency during meta training, real-world adaptation performance, and generalization of the context-based meta RL agent in PiH assembly tasks compared to previous approaches.
Ahmed Shokry, Walid Gomaa 0001, Tobias Zaenker, Murad Dawood, Rohit U. Menon, Shady A. Maged, Mohammed I. Awad, Maren Bennewitz
IROS2
2024 Blockchain-Based Framework for Stock Market Operations: IPOs, Trading, and Dividend Distribution
abstract
This study examines blockchain technology's potential to enhance the stock market by offering a decentralized alternative to the traditional system. The study explores the challenges of the current traditional stock market, such as long settlement times, weak transparency, dependency on centralized clearing houses, and heightened fees. The proposed novel system leverages blockchain technology to create a fully decentralized stock market encompassing all fundamental operations from Initial Public Offerings (IPOs) to stock trading and extending to distributing profits or dividends. It connects all market participants through an Ethereum consortium blockchain fuelled by smart contracts. This approach replaces traditional intermediaries like brokers with automated regulations, ensuring more rapid transaction settlements, reduced operational errors and fees, and strengthened security measures. The proposed blockchain-based system, consisting of four smart contracts for IPO, stock exchange, dividend distribution, and participant management, demonstrates remarkable performance, showing promise in efficiently handling the current load of the largest stock markets like the London Stock Exchange (LSE). Through theoretical and practical testing, the system achieved a remarkable throughput of 120 transactions per second and latency or finality of just 4 seconds, underscoring its capability to significantly enhance the efficiency of stock market operations.11The implementation code is available via GitHub upon request.
Muhammad Sharshar, Kareem Elzeky, Menna Othman, SeifALdin Khaled, Hosam Moubarak, Walid Gomaa 0001
CloudCom6
2024 Automated Detection of Defects on Metal Surfaces Using Vision Transformers
Toqa Alaa, Mostafa Kotb, Arwa Zakaria, Mariam Diab, Walid Gomaa 0001
ICINCO (2)5
2024 Video Summarization Techniques: A Comprehensive Review
Toqa Alaa, Ahmad Mongy, Assem Bakr, Mariam Diab, Walid Gomaa 0001
ICINCO (2)5
2024 Precision Aquaculture: An Integrated Computer Vision and IoT Approach for Optimized Tilapia Feeding
Rania Hossam, Ahmed Heakl, Walid Gomaa 0001
ICINCO (1)3
2024 Amina: an Arabic multi-purpose integral news articles dataset
abstract
Abstract Electronic newspapers are one of the most common sources of Modern Standard Arabic. Existing datasets of Arabic news articles typically provide a title, body, and single label. Ignoring important features, like the article author, image, tags, and publication date, can degrade the efficacy of classification models. In this paper, we propose the Arabic multi-purpose integral news articles (AMINA) dataset. AMINA is a large-scale Arabic news corpus with over 1,850,000 articles collected from 9 Arabic newspapers from different countries. It includes all the article features: title, tags, publication date and time, location, author, article image and its caption, and the number of visits. To test the efficacy of the proposed dataset, three tasks were developed and validated: article textual content (classification and generation) and article image classification. For content classification, we experimented the performance of several state-of-the-art Arabic NLP models including AraBERT and CAMeL-BERT, etc. For content generation, the reformer architecture is adopted as a character text generation model. For image classification applied on Al-Sharq and Youm7 news portals, we have compared the performance of 10 pre-trained models including ConvNeXt, MaxViT, ResNet18, etc. The overall study verifies the significance and contribution of our newly introduced Arabic articles dataset. The AMINA dataset has been released at https://huggingface.co/datasets/MohamedZayton/AMINA .
Mohamed Zaytoon, Muhannad Bashar, Mohamed Abdelaziz Khamis, Walid Gomaa 0001
Neural Comput. Appl.4
2023 Non-Parallel Training Approach for Emotional Voice Conversion Using CycleGAN
Mohamed El-Sayed 0011, Sama Hadhoud, Alaa Elsetohy, Menna Osman, Walid Gomaa 0001
ICINCO (2)5
2023 Masry: A Text-to-Speech System for the Egyptian Arabic
Ahmed Hammad Azab, Ahmed Bayoumy Zaki, Tetsuji Ogawa, Walid Gomaa 0001
ICINCO (2)4
2023 Network Analysis of the Egyptian Reddit Community
Samy Shaawat, Adham Hammad, Karim Farhat, Mina Thabet, Walid Gomaa 0001
ICINCO (2)5
2023 Camera Coach: Activity Recognition and Assessment Using Thermal and RGB Videos
abstract
Human Activity Recognition (HAR) has sparked a lot of attention because of its wide range of applications in sports, animation, simulation, and entertainment. There are many applications based on HAR, like healthcare monitoring, entertainment applications, and athletes' performance analysis. During the Corona pandemic, people took remote training as a safe and easy solution to do sports activities during the closure period, so there was a need to use smart systems to help people to train alone. In this paper, we expand our work on our dataset Multi-Modal Dataset of Sports (MMDOS); this dataset contains a variety of data forms from different sensors, including RGB videos, inertial motion data, depth, and thermal data, all synchronized with regard to the activities/exercises performed. The dataset consists of four workout exercises: free squats, shoulder press, push-ups, and lunges. The exercises are performed by 50 participants indoors at a fitness center. Professional trainers label the data with the type of activity and the participant's mistakes. Data are collected from different sensors, including Red-Green-Blue (RGB) videos, depth videos and skeleton data, inertial motion units (IMU) data, and thermal data; all are synchronized. From these modalities, we specifically used RGB, thermal videos, and 3-D skeleton of depth videos to recognize and asses the four activities. We built a smart Gym coaching model using classical machine learning (ML) and deep learning (DL) methods. The proposed system achieved an outstanding performance of 99% for activity recognition and an error of 0.75 out of 10 for activity assessment.
Ahmed Sharshar, Ahmed Abo Eitta, Ahmed Fayez, Mohamed Abdelaziz Khamis, Ahmed Bayoumy Zaki, Walid Gomaa 0001
IJCNN6
2023 Thermal Gait Dataset for Deep Learning-Oriented Gait Recognition
abstract
This study attempted to construct a thermal dataset of human gait in diverse environments suitable for building and evaluating sophisticated deep learning models (e.g., vision transformers) for gait recognition. Gait is a behavioral biometric to identify a person and requires no cooperation from the person, making it suitable for security and surveillance applications. For security purposes, it is desirable to be able to recognize a person in darkness or other inadequate lighting conditions, in which thermal imagery is advantageous over visible light imagery. Despite the importance of such nighttime person identification, available thermal gait datasets captured in the dark are scarce. This study, therefore, collected a relatively large set of thermal gait data in both indoor and outdoor environments with several walking styles, e.g., walking normally, walking while carrying a bag, and walking fast. This dataset was utilized in multiple gait recognition tasks, such as gender classification and person verification, using legacy convolutional neural networks (CNNs) and modern vision transformers (ViTs). Experiments using this dataset revealed the effective training method for person ver-ification, the effectiveness of ViT on gait recognition, and the robustness of the models against the difference in walking styles; it suggests that the developed dataset enables various studies on gait recognition using state-of-the-art deep learning models.
Fatma Youssef, Ahmed El-Mahdy 0002, Tetsuji Ogawa, Walid Gomaa 0001
IJCNN4
2023 A perspective on human activity recognition from inertial motion data
abstract
Abstract Human activity recognition (HAR) using inertial motion data has gained a lot of momentum in recent years both in research and industrial applications. From the abstract perspective, this has been driven by the rapid dynamics for building intelligent, smart environments, and ubiquitous systems that cover all aspects of human life including healthcare, sports, manufacturing, commerce, etc., which necessitate and subsume activity recognition aiming at recognizing the actions, characteristics, and goals of one or more agent(s) from a temporal series of observations streamed from one or more sensors. From a more concrete and seemingly orthogonal perspective, such momentum has been driven by the ubiquity of inertial motion sensors on-board mobile and wearable devices including smartphones, smartwatches, etc. In this paper we give an introductory and a comprehensive survey to the subject from a given perspective. We focus on a subset of topics, that we think are major, that will have significant and influential impacts on the future research and industrial-scale deployment of HAR systems. These include: (1) a comprehensive and detailed description of the inertial motion benchmark datasets that are publicly available and/or accessible, (2) feature selection and extraction techniques and the corresponding learning methods used to build workable HAR systems; we survey classical handcrafted datasets as well as data-oriented automatic representation learning approach to the subject, (3) transfer learning as a way to overcome many hurdles in actual deployments of HAR systems on a large scale, (4) embedded implementations of HAR systems on mobile and/or wearable devices, and finally (5) we touch on adversarial attacks, a topic that is essentially related to the security and privacy of HAR systems. As the field is very huge and diverse, this article is by no means comprehensive; it is though meant to provide a logically and conceptually rather complete picture to advanced practitioners, as well as to present a readable guided introduction to newcomers. Our logical and conceptual perspectives mimic the typical data science pipeline for state-of-the-art AI-based systems.
Walid Gomaa 0001, Mohamed Abdelaziz Khamis
Neural Comput. Appl.1
2023 Lyrics Analysis of the Arab Singer Abdel ElHalim Hafez
abstract
In this work we analyze the lyrics of one of the most famous and influential Arab artists in the twentieth century, namely, (Abdel ElHalim Hafez). Lyrics analysis provides a deep insight into the artist’s career evolution and his interactions with the surrounding environment including the social, political, and economic conditions. In order to perform such analysis we had to collect and compile the lyrics of Abdel ElHalim accompanied with the necessary metadata into an organized and structured form. The data are preprocessed by removing stop words and doing some normalization operations over the songs’ prose. We did not perform any lemmatization or stemming as the original form of the tokens convey much more information than the source words. We performed a lexical analysis in order to study both the lexical density and diversity over the course of Abdel ElHalim’s career life. We have as well studied the most significant words, idioms, and terms played in the songs using tools such as word clouds and more quantitative measures such as term frequency-inverse document frequency. We have divided the career life of Abdel ElHalim into sub-decades of length 5 years and all analyses are done both in a yearly fashion and more coarsely over such sub-decades. We have found a strong correlation between our statistical analysis and the socio-political status in Egypt and the Arab world during that time. This is especially relevant knowing that Abdel ElHalim is very much truly considered the son of the generation of the 1952 revolution in Egypt. The significance of Abdel ElHalim and his lyrics stem essentially from being contemporaneous to radical changes in Egypt across all sectors including political (support of liberation movements across the world, and the conflict with Israel), and socio-economic (especially changing the social class structure in Egypt). We also investigated the potential effectiveness of PoS (Part of Speech) tagging in genre analysis and classification.
Walid Gomaa 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2022 Analysis of the Squat Exercise from Visual Data
Fatma Youssef, Ahmed Bayoumy Zaki, Walid Gomaa 0001
ICINCO3
2022 Violence Detection and Recognition from Diverse Video Sources
abstract
The world is witnessing a vast increase in the crime rate, accompanied by severe losses of lives and property. Surveillance cameras are present everywhere. However, human supervision is becoming inefficient and incapable of early detection of violent acts, due to the diversity and the unexpected scenarios of violence. Accordingly, contributing to automatic violence prevention is becoming exceedingly important. Fulfilling this urgent demand, we turn to Machine Learning to help detect and classify violent events from video streams. In this paper, we propose a framework for detecting violence in video captures and streams, followed by categorizing the violent act in case the video clip is classified as violent. Supervised Learning is applied for both, the binary classification problem and the multi-class violence classification problem. The detection model relies on the usage of 3D Convolutional Neural Networks. The classification model utilizes the pre-trained Inception-v3 model for feature extraction, followed by Gated Recurrent Units (GRUs) for temporal processing. The models are trained on multiple datasets, whose frame-level annotations are available. We meant to use videos from various sources such as surveillance cameras, human recordings, movies, and public websites such as YouTube, to demonstrate the effectiveness of our models on different data sources. Furthermore, Transfer Learning from pre-trained models is applied, where each model trained on one of the datasets, is lightly re-trained on a different one, mostly demonstrating better performance than the original model, in terms of computational resources demands and accuracy.
Maria Gadelkarim, Mazen Khodier, Walid Gomaa 0001
IJCNN3
2022 MM-DOS: A Novel Dataset Of Workout Activities
abstract
Human Activity Recognition (HAR) is now one of the most widespread fields in research and application, especially within the area of artificial intelligence. A wide variety of applications are based on HAR, including healthcare, entertainment, and sports performance analysis. Such systems can be used as a reliable source of performance feedback for sports. Hence, coaching and the athletes' performance enhancement became the main application for HAR. Considering a data-oriented approach, the need for quality datasets increased to cover these springing needs. In this paper, we present a new multi-modal dataset MMDOS; the dataset contains a variety of data forms from different sensors, including Red-Green-Blue (RGB) videos, inertial motion data, depth, and thermal data, all synchronized with regard to the activities/exercises performed. The dataset consists of four workout exercises: free squats, shoulder presses, push-ups, and lunges. The exercises are performed by 50 participants indoors at a fitness center. Professional trainers label the data with the type of activity and the mistakes that the participant made. In this paper, we offer further details about the nature of the data and how the dataset was collected and post-processed. Finally, the paper shows some analysis performed to evaluate the dataset quality to recognize the different exercises. We reached accuracy up to 96% for RGB videos, 90% for depth videos, and 91% for IMU data.
Ahmed Sharshar, Ahmed Fayez, Ahmed Abo Eitta, Walid Gomaa 0001
IJCNN4
2022 Investigating strategies towards adversarially robust time series classification
abstract
Deep neural networks have been shown to be vulnerable against specifically-crafted perturbations designed to affect their predictive performance. Such perturbations, formally termed ‘adversarial attacks’ have been designed for various domains in the literature, most prominently in computer vision and more recently, in time series classification. Therefore there is a need to derive robust strategies to defend deep networks from such attacks. In this work we propose to establish axioms of robustness against adversarial attacks in time series classification. We subsequently design a suitable experimental methodology and empirically validate the hypotheses put forth. Results obtained from our investigations confirm the proposed hypotheses, and provide a strong empirical baseline with a view to mitigating the effects of adversarial attacks in deep time series classification.
Mubarak G. Abdu-Aguye, Walid Gomaa 0001, Yasushi Makihara, Yasushi Yagi
Pattern Recognit. Lett.2
2021 Location Determination of On-body Inertial Sensors
Hisham Madcor, Osama Adel, Walid Gomaa 0001
ICINCO3
2021 BioDeep: A Deep Learning System for IMU-based Human Biometrics Recognition
Abeer Mostafa, Samir A. Elsagheer Mohamed, Walid Gomaa 0001
ICINCO3
2021 Inertial Gait-based Person Authentication Using Siamese Networks
abstract
Gait-based person authentication using smartphones and wearable devices has attracted a lot of attention in recent years due to its unobtrusive nature and the widespread inertial sensors embedded in such devices. While traditional methods of authentication relied on hand-crafted feature extraction tailored for specific conditions, recent deep learning algorithms have been used to automatically discover hidden features and patterns in the raw data. However, current deep learning methods of gait authentication are not portable. They train models to classify a specific user as legitimate and the rest as imposters without being able to change that user to a different one without retraining the model from scratch. Furthermore, they are impractical, since they require a large volume of data for each class to train properly. In this work, we propose using a Siamese Network-based model which (1) Can be trained using only a few samples per subject. (2) Automatically extract features from raw inertial data. (3) Once trained, can authenticate any new user given only one gait instance of that new user. We apply our model to three publicly available datasets and show that it can achieve very promising results. For example, our model achieves 3.42% EER on the world's largest inertial gait dataset (OU-ISIR) after training on raw data of 592 subjects and testing on the raw data of completely different 78 subjects, where each subject provided an average of 5.9 seconds of gait data only. We show some performance analysis on different architectural variations. and finally, we show the ability of our model to generalize to new unseen datasets using transfer learning.
Osama Adel, Mostafa Soliman, Walid Gomaa 0001
IJCNN3
2021 Automatic Detection of Violence in Video Scenes
abstract
Violence detection from video stream using machine learning is a very rising field of inspection due to its great contribution to achieving peace and security and saving people's lives by automatically early-detecting violent acts and alarming those responsible to interfere. Knowing its significance, A Convolutional Neural Network (CNN) has been designed and implemented to tackle this problem, along with two classical classifiers, namely, support vector machine and random forest, in order to detect violence in video streams. The CNN is used both as a classifier and as a feature extractor whose outputs are fed into the other two classifiers. A data structure called ‘packets' was developed to refer to the input of our CNN network, packets are used to train the model on short clips where a packet consists of 15 sampled frames that make up one second of a video. So the input to the CNN is a subsampling of 3D volume of video frames. The problem is posed as a binary classification. Using four different datasets one of which is a combination of YouTube videos collected and annotated by the authors mixed with a dataset found on kaggle but also filtered by the authors. In addition, three other benchmark datasets is used. Our model was trained using supervised learning on a set of normal and violent videos, and tested using three separate classifiers whose results are compared together and then compared to other state-of-the-art approaches. Finally, transfer learning was implemented via cross validating models that where trained on a certain dataset and tested on another.
Ahmed Abo Eitta, Toka Barabash, Yousef Nafea, Walid Gomaa 0001
IJCNN4
2020 Detecting Adversarial Attacks In Time-Series Data
abstract
In recent times, deep neural networks have seen increased adoption in highly critical tasks. They are also susceptible to adversarial attacks, which are specifically crafted changes made to input samples which lead to erroneous output from such models. Such attacks have been shown to affect different types of data such as images and more recently, time-series data. Such susceptibility could have catastrophic consequences, depending on the domain.We propose a method for detecting Fast Gradient Sign Method (FGSM) and Basic Iterative Method (BIM) adversarial attacks as adapted for time-series data. We frame the problem as an instance of outlier detection and construct a normalcy model based on information and chaos-theoretic measures, which can then be used to determine whether unseen samples are normal or adversarial. Our approach shows promising performance on several datasets from the 2015 UCR Time Series Archive, reaching up to 97% detection accuracy in the best case.
Mubarak G. Abdu-Aguye, Walid Gomaa 0001, Yasushi Makihara, Yasushi Yagi
ICASSP2
2020 Gait-based Person Identification using Multiple Inertial Sensors
Osama Adel, Yousef Nafea, Ahmed Hesham, Walid Gomaa 0001
ICINCO4
2020 Improved IMU-based Human Activity Recognition using Hierarchical HMM Dissimilarity
Sara Ashry Mohammed, Walid Gomaa 0001, Mubarak G. Abdu-Aguye, Nahla El-borae
ICINCO2
2020 Multi-sensor Gait Analysis for Gender Recognition
Abeer Mostafa, Toka Ossama Barghash, Asmaa Al-Sayed Assaf, Walid Gomaa 0001
ICINCO4
2020 Adaptive Pooling Is All You Need: An Empirical Study on Hyperparameter-insensitive Human Action Recognition Using Wearable Sensors
abstract
A plethora of techniques have been proposed in human action recognition fields, and particularly deep learning-based methods such as convolutional neural networks (CNNs) have achieved impressive results. Usually, there is need to tune hyper-parameters in the deep neural network (e.g., filter size, stride) to achieve reasonable results. Such hyper-parameter tuning is, however, extremely time and resource-intensive even for small models. In this paper, we posit that the inclusion of an adaptive pooling in CNNs used for human action recognition largely eliminates the need for hyper-parameter tuning. Specifically, we demonstrated our idea for human action recognition using inertial sensor data (i.e., a temporal sequence) with a one-dimensional adaptive pooling. We compared the adaptive pooling to conventional CNNs with randomly chosen hyper-parameters using a publicly available data set for human action recognition. Experimental results showed that the adaptive pooling achieved better accuracy than the conventional CNNs.
Mubarak G. Abdu-Aguye, Walid Gomaa 0001, Yasushi Makihara, Yasushi Yagi
IJCNN2
2020 Identifying motion pathways in highly crowded scenes: A non-parametric tracklet clustering approach
Allam S. Hassanein, Mohamed E. Hussein 0001, Walid Gomaa 0001, Yasushi Makihara, Yasushi Yagi
Comput. Vis. Image Underst.3
2020 Gate and common pathway detection in crowd scenes and anomaly detection using motion units and LSTM predictive models
Abdullah N. Moustafa, Walid Gomaa 0001
Multim. Tools Appl.2
2020 Efficient and Robust Skeleton-Based Quality Assessment and Abnormality Detection in Human Action Performance
abstract
Elderly people can be provided with safer and more independent living by the early detection of abnormalities in their performing actions and the frequent assessment of the quality of their motion. Low-cost depth sensing is one of the emerging technologies that can be used for unobtrusive and inexpensive motion abnormality detection and quality assessment. In this study, we develop and evaluate vision-based methods to detect and assess neuromusculoskeletal disorders manifested in common daily activities using three-dimensional skeletal data provided by the SDK of a depth camera (e.g., MS Kinect and Asus Xtion PRO). The proposed methods are based on extracting medically -justified features to compose a simple descriptor. Thereafter, a probabilistic normalcy model is trained on normal motion patterns. For abnormality detection, a test sequence is classified as either normal or abnormal based on its likelihood, which is calculated from the trained normalcy model. For motion quality assessment, a linear regression model is built using the proposed descriptor in order to quantitatively assess the motion quality. The proposed methods were evaluated on four common daily actions-sit to stand, stand to sit, flat walk, and gait on stairs-from two datasets, a publicly released dataset and our dataset that was collected in a clinic from 32 patients suffering from different neuromusculoskeletal disorders and 11 healthy individuals. Experimental results demonstrate promising results, which is a step toward having convenient in-home automatic health care services.
Amr Elkholy, Mohamed E. Hussein 0001, Walid Gomaa 0001, Dima Damen, Emmanuel Saba
IEEE J. Biomed. Health Informatics3
2019 VersaTL: Versatile Transfer Learning for IMU-based Activity Recognition using Convolutional Neural Networks
abstract
The advent of Deep Learning has, together with massive gains in predictive accuracy, made it possible to reuse knowledge learnt from solving one problem in solving related problems. This is described as Transfer Learning, and has seen wide adoption especially in computer vision problems, where Convolutional Neural Networks have shown great flexibility and performance. On the other hand, transfer learning for sequences or timeseries data is typically made possible through the use of recurrent neural networks, which are difficult to train and prone to overfitting. In this work we present VersaTL, a novel approach to transfer learning for fixed and variable-length activity recognition timeseries data. We train a Convolutional Neural Network and use its convolutional filters as a feature extractor, then subsequently train a feedforward neural network as a classifier over the extracted features for other datasets. Our experiments on five different activity recognition datasets show the promise of this method, yielding results typically within 5% of trained-from-scratch networks while obtaining between a 24-52x reduction in the training time.
Mubarak G. Abdu-Aguye, Walid Gomaa 0001
ICINCO (1)2
2019 Robust Human Activity Recognition based on Deep Metric Learning
abstract
In the domain of Activity Recognition, the proliferation of low-cost and sensor-enabled personal devices has led to significant heterogeneity in the data generated by users. Traditional approaches to this problem have previously relied on handcrafted features and template-matching methods, which have limited flexibility and performance with high variability. In this work we investigate the use of Deep Metric Learning in the domain of activity recognition. We use a deep Triplet Network to generate fixed-length descriptors from activity samples for purposes of classification. We carry out evaluation of our proposed method on five datasets from different sources with differing activities. We obtain classification accuracies of up to 96% in self-testing scenarios and up to 91% accuracy in cross-dataset testing without retraining. We also show that our method performs similarly to traditional Convolutional Neural Networks. The obtained results indicate the promise of this approach.
Mubarak G. Abdu-Aguye, Walid Gomaa 0001
ICINCO (1)2
2019 On the Feasibility of On-body Roaming Models in Human Activity Recognition
Mubarak G. Abdu-Aguye, Walid Gomaa 0001, Yasushi Makihara, Yasushi Yagi
ICINCO (1)2
2019 Estimating Vehicle Speed on Highway Roads from Smartphone Sensors Using Deep Learning Models
abstract
Speed estimation is an open research area because of its importance in many applications and the necessity of replacing GPS due to smartphones battery drainage. Relying on integrating accelerometer values is challenging and generally requires continual utilization of external references to correct speed, because of error accumulation. Therefore, highway speed estimation, in particular, is a difficult problem due to the maintenance of high speeds by vehicles for a long time and the scarcity of reference points like uneven road surfaces, turns, and stops. In this paper, we investigate exploiting micro road surface unevenness that results in vibrations on the accelerometer readings without integrating the acceleration signal. In particular, we employ deep 1D convolutional neural networks to learn and extract robust features that learn the relation between such complex vibrations and speed. Also, the use of bidirectional LSTMs is investigated to benefit from both forward and backward dependencies in the sensed data, and allow a form of integration. Specifically, two highway speed estimation models are proposed. The first uses a deep convolutional neural network with 5 layers and the second uses a deep bidirectional LSTM neural network. The inputs to both networks are the readings from the accelerometer and gyroscope sensors of a smartphone. The methods achieved mean absolute error results of 5.53 km/hr and 3.71 km/hr, respectively; whereas a related LSTM based method, resulted in a high error rate of 68.05 km/hr. Finally, an implementation of the proposed CNN model on an android and iOS smartphones is described and analyzed.
Norhan Elsayed Amer Abdelgawad, Ahmed El-Mahdy 0002, Walid Gomaa 0001, Amin A. Shoukry
ICTAI3
2019 Competitive Feature Extraction for Activity Recognition based on Wavelet Transforms and Adaptive Pooling
abstract
Any application of machine learning requires feature extraction, whereupon the source data is processed to yield representations that are germane to obtaining the desired output. Traditional approaches to feature extraction include the estimation of statistical and structural properties of the data, the choice of which is mainly influenced by domain knowledge. However, Deep Learning has made it possible to learn the best features directly from the data itself, when sufficient data is available. For domains such as activity recognition where such plentiful data may be lacking, the application of deep learning may be limited. Therefore, methods which can yield deep learning-like performance without the need for massive amounts of data remain of interest.In this work we present a novel approach to feature extraction for sensor-generated activity recognition data. We first process the data using wavelet transforms, and subsequently use an adaptive pooling operator on the generated decomposition to obtain a compact, fixed-length representation of the data. Our experiments on seven different activity recognition datasets yield results comparable to those obtained from a deep neural network for all the considered datasets without the need for large amounts of data or any training overhead.
Mubarak G. Abdu-Aguye, Walid Gomaa 0001
IJCNN2
2019 Trans-Sense: Real Time Transportation Schedule Estimation Using Smart Phones
abstract
Developing countries suffer from traffic congestion, poorly planned road/rail networks, and lack of access to public transportation facilities. This context results in an increase in fuel consumption, pollution level, monetary losses, massive delays, and less productivity. On the other hand, it has a negative impact on the commuters feelings and moods. Availability of real-time transit information - by providing public transportation vehicles locations using GPS devices - helps in estimating a passenger's waiting time and addressing the above issues. However, such solution is expensive for developing countries. This paper aims at designing and implementing a crowd-sourced mobile phones-based solution to estimate the expected waiting time of a passenger in public transit systems, the prediction of the remaining time to get on/off a vehicle, and to construct a real time public transit schedule.Trans-Sense has been evaluated using real data collected for over 800 hours, on a daily basis, by different Android phones, and using different light rail transit lines at different time spans. The results show that Trans-Sense can achieve an average recall and precision of 95.35% and 90.1%, respectively, in discriminating lightrail stations. Moreover, the empirical distributions governing the different time delays affecting a passenger's total trip time enable predicting the right time of arrival of a passenger to her destination with an accuracy of 91.81%. In addition, the system estimates the stations dimensions with an accuracy of 95.71%.
Ali Mohamed AbdelAziz, Amin A. Shoukry, Walid Gomaa 0001, Moustafa Youssef 0001
SECON3
2018 Zero-Calibration Device-Free Localization for the IoT Based on Participatory Sensing
abstract
Device-free localization (DFL) is an emerging technology for estimating the position of a human or object that is not equipped with any electronic tag, nor participate actively in the localization process. Similar to device-based localization, the initial phase in DFL is to build the fingerprint database which is usually done manually using site surveying. This process is tedious, time-consuming, and vulnerable to environmental dynamics. Motivated by the recent advances in the Internet of Things (IoT), this paper introduces RadioGrapher; a system that automates the process of device-free fingerprint calibration in IoT environments. RadioGrapher leverages the device-based locations of entities in the area of interest in a crowd-sensing manner, aided with Fresnel zones of the wirelessly connected IoT devices to automatically construct a device-free fingerprint. Experimental evaluation of RadioGrapher in an IoT testbed using multiple entities shows that it can construct DFL fingerprints with high accuracy. Moreover, its median localization accuracy is comparable to that of manual fingerprinting. This comes with no calibration overhead, highlighting the promise of RadioGrapher as a crowdsourcing device-free fingerprint constructor in IoT environments.
Osama T. Ibrahim, Walid Gomaa 0001, Moustafa Youssef 0001
GLOBECOM2
2018 Deep 3D Pose Dictionary: 3D Human Pose Estimation from Single RGB Image Using Deep Convolutional Neural Network
Reda Elbasiony, Walid Gomaa 0001, Tetsuya Ogata
ICANN (3)2
2018 An LSTM-based Descriptor for Human Activities Recognition using IMU Sensors
Sara Ashry Mohammed, Reda Elbasiony, Walid Gomaa 0001
ICINCO (1)3
2018 Novel Approaches to Activity Recognition Based on Vector Autoregression and Wavelet Transforms
abstract
The recognition of daily activities has been a long-running research domain, which has received increasing attention over the past few years. This is due to the proliferation of personal devices which are capable of reporting the physical signals generated during these activities. Being a classification problem, the primary focus is on suitable modalities for feature extraction and proper choice of classifiers. In this work we investigate the performance of two novel approaches to feature extraction based on Vector Autoregression and Wavelet Transforms together with four different classifiers. The results indicate that the two proposed feature extraction methods are suitable for this domain. In addition, the Canonical Correlation Forests classifier has been found to be a promising candidate for inference in the domain of Activity Recognition.
Mubarak G. Abdu-Aguye, Walid Gomaa 0001
ICMLA2
2018 On the complexity of bounded time and precision reachability for piecewise affine systems
Hugo Bazille, Olivier Bournez, Walid Gomaa 0001, Amaury Pouly
Theor. Comput. Sci.3
2017 Exploration of Unknown Map for Safety Purposes using Wheeled Mobile Robots
Sara Ashry Mohammed, Walid Gomaa 0001
ICINCO (2)2
2017 ADL Classification Based on Autocorrelation Function of Inertial Signals
abstract
Recognition of human activities is one of the most promising research areas in artificial intelligence. This has come along with the technological advancement in sensing technologies as well as the high demand for applications that are mobile, context-aware, and real-time. In this paper, we use a smart watch to collect sensory data for 14 ADL activities. We collect three types of sensory signals: acceleration, angular velocity, and rotation displacement; each is a tri-axial signal. From each given signal we compute the autocorrelation function up to certain lag and take these computed values as representative features of the given signal. We then feed these features to a random-forest-based classifier for training and prediction. We experiment with different combinations of sensory data. The joint use of acceleration with angular velocity has achieved the best performance in prediction accuracy which reaches about 80% for the whole set of 14 activities.
Walid Gomaa 0001, Reda Elbasiony, Sara Ashry Mohammed
ICMLA1
2017 On the extension of computable real functions
abstract
We investigate interrelationships among different notions from mathematical analysis, effective topology, and classical computability theory. Our main object of study is the class of computable functions defined over an interval with the boundary being a left-c.e. real number. We investigate necessary and sufficient conditions under which such functions can be computably extended. It turns out that this depends on the behavior of the function near the boundary as well as on the class of left-c.e. real numbers to which the boundary belongs, that is, how it can be constructed. Of particular interest a class of functions is investigated: sawtooth functions constructed from computable enumerations of c.e. sets.
Mathieu Hoyrup, Walid Gomaa 0001
LICS2
2016 Exploration of Unknown Map for Gas Searching and Picking Up Objects using Khepera Mobile Robots
abstract
In this paper, the integrated approach for searching, detecting the source of gas leakage and capturing object using Khepera III mobile robot is introduced. It was tested in natural grid environments as shown in experiments but it is usable in vast areas; e.g. houses and labs to pick up gas victims. Experiments are performed using a master Khepera robot equipped with a gas sensor and two slave robots equipped with grippers to take the best path to pick up the object. Moreover, this paper proposes an improvement of Albers exploration algorithm to reduce the time required to explore an unknown map with different polygon obstacles. The proposed approach aims at minimizing the overall exploration time, making it possible to localize gas source in an effi- cient way, as demonstrated in V-REP simulation as well as real world experiments. A comparison among both algorithms has shown the effectiveness of the proposed one, where the percentage of performance speedup is about 30% to 57% depending on the size of the map and number of obstacles.
Sara Ashry Mohammed, Walid Gomaa 0001
ICINCO (2)2
2016 Semantic Analysis for Crowded Scenes Based on Non-Parametric Tracklet Clustering
Allam S. Hassanein, Mohamed E. Hussein 0001, Walid Gomaa 0001
IJCAI3
2015 Car Following Markov Regime Classification and Calibration
abstract
The car following behavior has recently gained much attention due to its wide variety of applications. This includes accident analysis, driver assessment, support systems, and road design. In this paper, we present a model that leverages Markov regime switching models to classify various car following regimes. The detected car following regimes are then mined to calibrate the parameters of drivers to be dependent on the driver's current driving regime. A two stage Markov regime switching model is utilized to detect different car following regimes. The first stage discriminates normal car following regimes from abnormal ones, while the second stage classifies normal car following regimes to their fine-grained regimes like braking, accelerating, standing, free-flowing, and normal following. A genetic algorithm is then employed to the observed driver data in each car following regime to optimize car following model parameter values of the driver in each regime. Experimental evaluation of the proposed model using a real dataset shows that it can detect up-normal (rare and short time) events. In addition, it can infer the switching process dynamics such as the expected duration, the probability of moving from one regime to another and the switching parameters of each regime. Finally, the model is able to accurately calibrate the parameters of drivers according to their driving regimes, so we can achieve a better understanding of drivers behavior and better simulation of driving situation.
Ahmed Bayoumy Zaki, Walid Gomaa 0001, Mohamed Abdelaziz Khamis
ICMLA2
2015 Machine learning in computational docking
Mohamed Abdelaziz Khamis, Walid Gomaa 0001, Walaa Fathy Ahmed
Artif. Intell. Medicine2
2015 Comparative assessment of machine-learning scoring functions on PDBbind 2013
abstract
Computational docking is the core process of computer-aided drug design (CADD); it aims at predicting the best orientation and conformation of a small molecule (drug ligand) when bound to a target large receptor molecule (protein) in order to form a stable complex molecule. The docking quality is typically measured by a scoring function: a mathematical predictive model that produces a score representing the binding free energy and hence the stability of the resulting complex molecule. An effective scoring function should produce promising drug candidates which can then be synthesized and physically screened using high throughput screening (HTS) process. Therefore, the key to CADD is the design of an efficient highly accurate scoring function. Many traditional techniques have been proposed, however, the performance was generally poor. Only in the last few years the application of the machine learning (ML) technology has been applied in the design of scoring functions; and the results have been very promising. In this paper, we propose 12 scoring functions based on a wide range of ML techniques. We analyze the performance of each on the scoring power (binding affinity prediction), ranking power (relative ranking prediction), docking power (identifying the native binding poses among computer-generated decoys), and screening power (classifying true binders versus negative binders) using the PDBbind 2013 database. We compare our results with the recently published comparative assessment of scoring functions (CASF-2013) of 20 classical scoring functions most of which are implemented in main-stream commercial software. For the scoring and ranking powers, the proposed ML scoring functions depend on a wide range of features (energy terms, pharmacophore, geometrical) that entirely characterize the protein–ligand complexes (about 108 features); these features are extracted from several docking software available in the literature; a feature-space reduction technique, namely, principal component analysis is then applied and the performance is studied accordingly. For the docking and screening powers, the proposed ML scoring functions depend on the geometrical features of the RF-Score (36 features) to utilize a larger number of training complexes (relative to the large number of decoys in the testing set). For the scoring power, the best ML scoring function (RF) achieves a Pearson correlation coefficient between the predicted and experimentally determined binding affinities of 0.704 versus 0.614 achieved by the best classical scoring function (X-ScoreHM). For the ranking power, the best ML scoring function (RF) achieves a Spearman correlation coefficient between the ranks based on the predicted and experimentally determined binding affinities of 0.697 versus 0.626 achieved by the best classical scoring function (X-ScoreHM). For the docking power, the best ML scoring function (BRT) has a success rate in identifying the top best-scored ligand binding pose within 2 Å root-mean-square deviation from the native pose of 13.8% versus 81.0% achieved by the best classical scoring function (ChemPLP@GOLD). For the screening power, the best ML scoring function (SVM) has an average enrichment factor and success rate at the top 1% level of 3.76 and 6.45% versus 19.54 and 60% respectively achieved by the best classical scoring function (GlideScore-SP).
Mohamed Abdelaziz Khamis, Walid Gomaa 0001
Eng. Appl. Artif. Intell.2
2014 Ideal Quantification of Chaos at Finite Resolution
Islam A. Elshaarawy, Walid Gomaa 0001
ICCSA (1)2
2014 Complex Motion Planning for NAO Humanoid Robot
abstract
In this paper, we introduce an integrated approach that enables a humanoid robot to plan and robustly execute whole body motions including stepping over, climbing up or down obstacles as well as climbing up straight staircase using only onboard sensing. Reliable and accurate sequence of motions for humanoid robots operating in complex indoor environments is a prerequisite for robots to fulfill high level tasks. The design of complex dynamic motions is achievable only through the use of robot kinematics. Based on the recognized object from the robot database, using the robot camera, a sequence of actions for avoiding that object is executed. As demonstrated in simulation as well as real world experiments with NAO humanoid, NAO can reliably execute robustly whole body movements in cluttered, multi-level environments containing objects of various shapes and sizes.
Walaa Gouda, Walid Gomaa 0001
ICINCO (2)2
2014 WiFi Localization for Mobile Robots Based on Random Forests and GPLVM
abstract
The proliferation of WiFi networks has attracted many research communities to employ WiFi signals in estimating the location of mobile devices in indoor environments. In this paper, we propose a localization framework that is capable of determining the location of mobile robots in indoor limited areas. The proposed framework exploits the random forests algorithm in both classification and regression techniques, which are used to build cooperative supervised localization models. The localization models are trained offline based on training data that contains measurements of WiFi signal strengths and the location of these measurements. We also propose an extension of our framework using the gaussian process latent variable model (GPLVM), which gives our framework the ability to build subjective localization models which do not require any prior knowledge about ground truth on the localization place. Our experimental evaluation of the proposed framework using the Khepera III mobile robot in one test bed show that it gives high accuracy, where the calculated mean localization error is ±36 cm.
Reda Elbasiony, Walid Gomaa 0001
ICMLA2
2014 Multi-Agent Reinforcement Learning Control for Ramp Metering
Ahmed Fares, Walid Gomaa 0001
ICSEng2
2014 Aggregate Reinforcement Learning for multi-agent territory division: The Hide-and-Seek game
Mohamed K. Gunady, Walid Gomaa 0001, Ikuo Takeuchi
Eng. Appl. Artif. Intell.2
2014 Adaptive multi-objective reinforcement learning with hybrid exploration for traffic signal control based on cooperative multi-agent framework
Mohamed Abdelaziz Khamis, Walid Gomaa 0001
Eng. Appl. Artif. Intell.2
2014 Analytical properties of resource-bounded real functionals
Hugo Férée, Walid Gomaa 0001, Mathieu Hoyrup
J. Complex.2
2013 On the Query Complexity of Real Functionals
abstract
Recently Kawamura and Cook developed a framework to define the computational complexity of operators arising in analysis. Our goal is to understand the effects of complexity restrictions on the analytical properties of the operator. We focus on the case of norms over C[0,1] and introduce the notion of dependence of a norm on a point and relate it to the query complexity of the norm. We show that the dependence of almost every point is of the order of the query complexity of the norm. A norm with small complexity depends on a few points but, as compensation, highly depends on them. We characterize the functionals that are computable using one oracle call only and discuss the uniformity of that characterization.
Hugo Férée, Mathieu Hoyrup, Walid Gomaa 0001
LICS3
2012 Enhanced Multiagent Multi-objective Reinforcement Learning for Urban Traffic Light Control
abstract
Traffic light control is one of the major problems in urban areas. This is due to the increasing number of vehicles and the high dynamics of the traffic network. Ordinary methods for traffic light control cause high rate of accidents, waste in time, and affect the environment negatively due to the high rates of fuel consumption. In this paper, we develop an enhanced version of our multiagent multi-objective traffic light control system that is based on a Reinforcement Learning (RL) approach. As a testbed framework for our traffic light controller, we use the open source Green Light District (GLD) vehicle traffic simulator. We analyze and fix some implementation problems in GLD that emerged when applying a more realistic continuous time acceleration model. We propose a new cooperation method between the neighboring traffic light agent controllers using specific learning and exploration rates. Our enhanced traffic light controller minimizes the trip time in major arteries and increases safety in residential areas. In addition, our traffic light controller satisfies green waves for platoons traveling in major arteries and considers as well the traffic environmental impact by keeping the vehicles speeds within the desirable thresholds for lowest fuel consumption. In order to evaluate the enhancements and new methods proposed in this paper, we have added new performance indices to GLD.
Mohamed Abdelaziz Khamis, Walid Gomaa 0001
ICMLA (1)2
2012 A survey of recursive analysis and Moore's notion of real computation
Walid Gomaa 0001
Nat. Comput.1
2011 Approaches to prediction of protein structure
abstract
Protein structure prediction (PSP) is one of the most important and challenging problems in bioinformatics today. This is due to the fact that the biological function of the protein is determined by its structure. While there is a gap between the number of known protein sequences and the number of known structures, protein structure prediction aims at reducing this sequence-structure gap. Protein structure can be experimentally determined using either X-ray crystallography or Nuclear Magnetic Resonance (NMR). However, these empirical techniques are very time consuming, so computational models and the associated algorithms have been developed for protein structure prediction. Such computational models fall into three categories: the homology approach, threading, and the ab initio approach. In this paper we give a general introductory background to the area and a literature survey about the ab initio approach in particular. This approach depends on the chemical and physical properties of the constituent amino acids. Not all ab initio algorithms have the same performance, so we represent the general success keys for any such algorithm.
Walaa Fathy Ahmed, Walid Gomaa 0001
AICCSA2
2011 Modeling gene regulatory networks: A survey
abstract
One of the major missions of the health care enterprise in Egypt is the proper control of local endemic diseases, such as liver and breast cancers, through the development and maintenance of integrated quality programs in patient care, research, education, and prevention. This was particularly manifested by the establishment of the National Cancer Institute NCI in the 1950's [1]. NCI is the largest comprehensive cancer center in Egypt, and is committed to the multidisciplinary approach to cancer management in which different treatment modalities cooperate rather than compete. Given a particular kind of disease one of the major problems to its management and treatment can be summarized by the following: (1) to isolate the genes that are involved in that cancer, (2) to model the interactions among these genes to derive what is called Gene Regulatory Network GRN, and (3) to detect the topological changes among the GRN of the normal state, the GRN of the cancer state, and the GRN that results as a pharmacologic response to a therapeutic intervention; these changes are called biomarkers. Modeling of GRNs is a multidisciplinary task that poses many scientific challenges. The subject has gained much interest in the past few years. In this article we give a quick survey of the current modeling techniques of gene regulatory networks.
Walid Gomaa 0001
AICCSA1
2011 Efficient Parallel Implementations of Controlled Optimization of Traffic Phases
Sameh Samra, Ahmed El-Mahdy 0002, Walid Gomaa 0001, Yasutaka Wada, Amin A. Shoukry
ICA3PP (1)3
2008 Active logic semantics for a single agent in a static world
Michael L. Anderson, Walid Gomaa 0001, John Grant, Donald Perlis
Artif. Intell.2
2003 Learning Joint Coordinated Plans in Multi-agent Systems
Walid Gomaa 0001, Amani A. Saad, Mohamed A. Ismail
IEA/AIE1