Mohamed Medhat Gaber

dblp:75/5251 · DBLP profile ↗
← Back
71ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0003-0339-4474ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 51 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 22 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5Human-computer interaction and ubiquitous computing · 3Theory of computation · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2025 CURVETE: Curriculum Learning and Progressive Self-supervised Training for Medical Image Classification
Asmaa Abbas, Mohamed Medhat Gaber, Mohammed M. Abdelsamea
ICONIP (3)2
2025 HITSQL: Human-In-The-Loop Techniques for Enhancing Text-to-SQL Query Generation with Large Language Models
abstract
This paper presents a novel model we called HIT-SQL to enhance text-to-SQL query generation using large language models (LLMs) with progressive active learning. HITSQL is specifically tailored for SAP Business One and has been validated on SAP HANA databases. Using an iterative refinement process, HITSQL combines expert feedback, systematic evaluation, and adaptive fine-tuning to achieve significant performance improvements. Starting with an initial dataset of 20 training samples, the model iteratively incorporated 180 evaluation samples over four iterations, expanding the training set to 200 samples while reducing the low-confidence pool to zero. The evaluation framework, which employs execution-based evaluation, component matching, and semantic evaluation, ensured robust assessment and classification of SQL queries. External testing demonstrated high generalisability, with an average accuracy of 97% on unseen queries. The results highlight the efficacy of the framework in generating reliable and accurate SQL queries for enterprise databases. Future work will focus on expanding datasets, automating feedback mechanisms, and extending applicability to cross-domain environments. This study establishes a strong foundation for LLM-based natural language interfaces for enterprise resource planning systems, with HITSQL contributing to improved data accessibility and decision-making for SMEs.
Dhoyazan Al-Turki, Shadi Basurra, Mohamed Medhat Gaber, Bushra Attasi, Mohammed M. Abdelsamea
IJCNN3
2023 AdaDeepStream: streaming adaptation to concept evolution in deep neural networks
abstract
Abstract Typically, Deep Neural Networks (DNNs) are not responsive to changing data. Novel classes will be incorrectly labelled as a class on which the network was previously trained to recognise. Ideally, a DNN would be able to detect changing data and adapt rapidly with minimal true-labelled samples and without catastrophically forgetting previous classes. In the Online Class Incremental (OCI) field, research focuses on remembering all previously known classes. However, real-world systems are dynamic, and it is not essential to recall all classes forever. The Concept Evolution field studies the emergence of novel classes within a data stream. This paper aims to bring together these fields by analysing OCI Convolutional Neural Network (CNN) adaptation systems in a concept evolution setting by applying novel classes in patterns. Our system, termed AdaDeepStream, offers a dynamic concept evolution detection and CNN adaptation system using minimal true-labelled samples. We apply activations from within the CNN to fast streaming machine learning techniques. We compare two activation reduction techniques. We conduct a comprehensive experimental study and compare our novel adaptation method with four other state-of-the-art CNN adaptation methods. Our entire system is also compared to two other novel class detection and CNN adaptation methods. The results of the experiments are analysed based on accuracy, speed of inference and speed of adaptation. On accuracy, AdaDeepStream outperforms the next best adaptation method by 27% and the next best combined novel class detection/CNN adaptation method by 24%. On speed, AdaDeepStream is among the fastest to process instances and adapt.
Lorraine Chambers, Mohamed Medhat Gaber, Hossein Ghomeshi
Appl. Intell.2
2023 SwinCup: Cascaded swin transformer for histopathological structures segmentation in colorectal cancer
abstract
Transformer models have recently become the dominant architecture in many computer vision tasks, including image classification, object detection, and image segmentation. The main reason behind their success is the ability to incorporate global context information into the learning process. By utilising self-attention, recent advancements in the Transformer architecture design enable models to consider long-range dependencies. In this paper, we propose a novel transformer, named Swin Transformer with Cascaded UPsampling (SwinCup) model for the segmentation of histopathology images. We use a hierarchical Swin Transformer with shifted windows as an encoder to extract global context features. The multi-scale feature extraction in a Swin transformer enables the model to attend to different areas in the image at different scales. A cascaded up-sampling decoder is used with an encoder to improve its feature aggregation. Experiments on GLAS and CRAG histopathology colorectal cancer datasets were used to validate the model, achieving an average 0.90 (F1 score) and surpassing the state-of-the-art by (23%).
Usama Zidan, Mohamed Medhat Gaber, Mohammed M. Abdelsamea
Expert Syst. Appl.2
2023 WhatsUp: An event resolution approach for co-occurring events in social media
abstract
The rapid growth of social media networks has resulted in the generation of a vast data amount, making it impractical to conduct manual analyses to extract newsworthy events. Thus, automated event detection mechanisms are invaluable to the community. However, a clear majority of the available approaches rely only on data statistics without considering linguistics. A few approaches involved linguistics, only to extract textual event details without the corresponding temporal details. Since linguistics define words’ structure and meaning, a severe information loss can happen without considering them. Targeting this limitation, we propose a novel method named WhatsUp to detect temporal and fine-grained textual event details, using linguistics captured by self-learned word embeddings and their hierarchical relationships and statistics captured by frequency-based measures. We evaluate our approach on recent social media data from two diverse domains and compare the performance with several state-of-the-art methods. Evaluations cover temporal and textual event aspects, and results show that WhatsUp notably outperforms state-of-the-art methods. We also analyse the efficiency, revealing that WhatsUp is sufficiently fast for (near) real-time detection. Further, the usage of unsupervised learning techniques, including self-learned embedding, makes our approach expandable to any language, platform and domain and provides capabilities to understand data-specific linguistics.
Hansi Hettiarachchi, Mariam Adedoyin-Olowe, Jagdev Bhogal, Mohamed Medhat Gaber
Inf. Sci.4
2022 Embed2Detect: temporally clustered embedded words for event detection in social media
abstract
Abstract Social media is becoming a primary medium to discuss what is happening around the world. Therefore, the data generated by social media platforms contain rich information which describes the ongoing events. Further, the timeliness associated with these data is capable of facilitating immediate insights. However, considering the dynamic nature and high volume of data production in social media data streams, it is impractical to filter the events manually and therefore, automated event detection mechanisms are invaluable to the community. Apart from a few notable exceptions, most previous research on automated event detection have focused only on statistical and syntactical features in data and lacked the involvement of underlying semantics which are important for effective information retrieval from text since they represent the connections between words and their meanings. In this paper, we propose a novel method termedEmbed2Detectfor event detection in social media by combining the characteristics in word embeddings and hierarchical agglomerative clustering. The adoption of word embeddings givesEmbed2Detectthe capability to incorporate powerful semantical features into event detection and overcome a major limitation inherent in previous approaches. We experimented our method on two recent real social media data sets which represent the sports and political domain and also compared the results to several state-of-the-art methods. The obtained results show thatEmbed2Detectis capable of effective and efficient event detection and it outperforms the recent event detection methods. For the sports data set, Embed2Detect achieved 27% higher F-measure than the best-performed baseline and for the political data set, it was an increase of 29%.
Hansi Hettiarachchi, Mariam Adedoyin-Olowe, Jagdev Bhogal, Mohamed Medhat Gaber
Mach. Learn.4
2022 DeepStreamOS: Fast open-Set classification for convolutional neural networks
Lorraine Chambers, Mohamed Medhat Gaber
Pattern Recognit. Lett.2
2021 Embed2Detect: Temporally Clustered Embedded Words for Event Detection in Social Media: Extended Abstract
abstract
This paper is an extended abstract for work [1]. We propose a novel method termed Embed2Detect for event detection in social media by mainly combining the characteristics in word embeddings and dendrograms. The adoption of word embeddings incorporates powerful semantical features into event detection to overcome a major limitation inherent in previous approaches.
Hansi Hettiarachchi, Mariam Adedoyin-Olowe, Jagdev Bhogal, Mohamed Medhat Gaber
DSAA4
2021 Towards Intrusion Detection Of Previously Unknown Network Attacks
abstract
Advances in telecommunication network technologies have led to an ever more interconnected world. Accordingly, the types of threats and attacks to intrude or disable such networks or portions of it are continuing to develop likewise. Thus, there is a need to detect previously unknown attack types. Supervised techniques are not suitable to detect previously not encountered attack types. This paper presents a new ensemble-based Unknown Network Attack Detector (UNAD) system. UNAD proposes a training workflow composed of heterogeneous and unsupervised anomaly detection techniques, trains on attack-free data and can distinguish normal network flow from (previously unknown) attacks. This scenario is more realistic for detecting previously unknown attacks than supervised approaches and is evaluated on telecommunication network data with known ground truth. Empirical results reveal that UNAD can detect attacks on which the workflows have not been trained on with a precision of 75% and a recall of 80%. The benefit of UNAD with existing network attack detectors is, that it can detect completely new attack types that have never been encountered before.
Saif Alzubi, Frederic T. Stahl, Mohamed Medhat Gaber
ECMS3
2021 On The Effect Of Decomposition Granularity On DeTraC For COVID-19 Detection Using Chest X-Ray Images
abstract
Covid-19 is a growing issue in society and there is a need for resources to manage the disease. This paper looks at studying the effect of class decomposition in our previously proposed deep Convolutional Neural Network, called DeTraC (Decompose, Transfer and Compose). DeTraC has the ability to robustly detect and predict Covid-19 from chest X-ray images. The experimental results showed that changing the number of clusters affected the performance of DeTraC and influenced the accuracy of the model. As the number of clusters increased, the accuracy decreased for the shallow tuning mode but increased for the deep tuning mode. This shows the importance of using suitable hyperparameter settings in order to get the best results from a deep learning model. The highest accuracy obtained, in this study, was 98.33% from the deep tuning model.
Nicole P. Mugova, Mohammed M. Abdelsamea, Mohamed Medhat Gaber
ECMS3
2021 Classification of COVID-19 in chest X-ray images using DeTraC deep convolutional neural network
abstract
Abstract Chest X-ray is the first imaging technique that plays an important role in the diagnosis of COVID-19 disease. Due to the high availability of large-scale annotated image datasets, great success has been achieved using convolutional neural networks ( CNN s) for image recognition and classification. However, due to the limited availability of annotated medical images, the classification of medical images remains the biggest challenge in medical diagnosis. Thanks to transfer learning, an effective mechanism that can provide a promising solution by transferring knowledge from generic object recognition tasks to domain-specific tasks. In this paper, we validate and a deep CNN , called Decompose, Transfer, and Compose ( DeTraC ), for the classification of COVID-19 chest X-ray images. DeTraC can deal with any irregularities in the image dataset by investigating its class boundaries using a class decomposition mechanism. The experimental results showed the capability of DeTraC in the detection of COVID-19 cases from a comprehensive image dataset collected from several hospitals around the world. High accuracy of 93.1% (with a sensitivity of 100%) was achieved by DeTraC in the detection of COVID-19 X-ray images from normal, and severe acute respiratory syndrome cases.
Asmaa Abbas, Mohammed M. Abdelsamea, Mohamed Medhat Gaber
Appl. Intell.3
2021 4S-DT: Self-Supervised Super Sample Decomposition for Transfer Learning With Application to COVID-19 Detection
abstract
Due to the high availability of large-scale annotated image datasets, knowledge transfer from pretrained models showed outstanding performance in medical image classification. However, building a robust image classification model for datasets with data irregularity or imbalanced classes can be a very challenging task, especially in the medical imaging domain. In this article, we propose a novel deep convolutional neural network, which we called self-supervised super sample decomposition for transfer learning (4S-DT) model. The 4S-DT encourages a coarse-to-fine transfer learning from large-scale image recognition tasks to a specific chest X-ray image classification task using a generic self-supervised sample decomposition approach. Our main contribution is a novel self-supervised learning mechanism guided by a super sample decomposition of unlabeled chest X-ray images. 4S-DT helps in improving the robustness of knowledge transformation via a downstream learning strategy with a class-decomposition (CD) layer to simplify the local structure of the data. The 4S-DT can deal with any irregularities in the image dataset by investigating its class boundaries using a downstream CD mechanism. We used 50000 unlabeled chest X-ray images to achieve our coarse-to-fine transfer learning with an application to COVID-19 detection, as an exemplar. The 4S-DT has achieved a high accuracy of 99.8% on the larger of the two datasets used in the experimental study and an accuracy of 97.54% on the smaller dataset, which was enriched by augmented images, out of which all real COVID-19 cases were detected.
Asmaa Abbas, Mohammed M. Abdelsamea, Mohamed Medhat Gaber
IEEE Trans. Neural Networks Learn. Syst.3
2020 A non-canonical hybrid metaheuristic approach to adaptive data stream classification
Hossein Ghomeshi, Mohamed Medhat Gaber, Yevgeniya Kovalchuk
Future Gener. Comput. Syst.2
2020 Co-eye: a multi-resolution ensemble classifier for symbolically approximated time series
abstract
Abstract Time series classification (TSC) is a challenging task that attracted many researchers in the last few years. One main challenge in TSC is the diversity of domains where time series data come from. Thus, there is no “one model that fits all” in TSC. Some algorithms are very accurate in classifying a specific type of time series when the whole series is considered, while some only target the existence/non-existence of specific patterns/shapelets. Yet other techniques focus on the frequency of occurrences of discriminating patterns/features. This paper presents a new classification technique that addresses the inherent diversity problem in TSC using a nature-inspired method. The technique is stimulated by how flies look at the world through “compound eyes” that are made up of thousands of lenses, called ommatidia. Each ommatidium is an eye with its own lens, and thousands of them together create a broad field of vision. The developed technique similarly uses different lenses and representations to look at the time series, and then combines them for broader visibility. These lenses have been created through hyper-parameterisation of symbolic representations (Piecewise Aggregate and Fourier approximations). The algorithm builds a random forest for each lens, then performs soft dynamic voting for classifying new instances using the most confident eyes, i.e., forests. We evaluate the new technique, coined Co-eye, using the recently released extended version of UCR archive, containing more than 100 datasets across a wide range of domains. The results show the benefits of bringing together different perspectives reflecting on the accuracy and robustness of Co-eye in comparison to other state-of-the-art techniques.
Zahraa Said Abdallah, Mohamed Medhat Gaber
Mach. Learn.2
2019 DeepHist: Towards a Deep Learning-based Computational History of Trends in the NIPS
abstract
Research in analysis of big scholarly data has increased in the recent past and it aims to understand research dynamics and forecast research trends. The ultimate objective in this research is to design and implement novel and scalable methods for extracting knowledge and computational history. While citations are highly used to identify emerging/rising research topics, they can take months or even years to stabilise enough to reveal research trends. Consequently, it is necessary to develop faster yet accurate methods for trend analysis and computational history that dig into content and semantics of an article. Therefore, this paper aims to conduct a fine-grained content analysis of scientific corpora from the domain of Machine Learning. This analysis uses DeepHist, a deep learning-based computational history approach; the approach relies on a dynamic word embedding that aims to represent words with low-dimensional vectors computed by deep neural networks. The scientific corpora come from 5991 publications from Neural Information Processing Systems (NIPS) conference between 1987 and 2015 which are divided into six 5-year timespans. The analysis of these corpora generates visualisations produced by applying t-distributed stochastic neighbor embedding (t-SNE) for dimensionality reduction. The qualitative and quantitative study reported here reveals the evolution of the prominent Machine Learning keywords; this evolution supports the popularity of current research topics in the field. This support is evident given how well the popularity of the detected keywords correlates with the citation counts received by their corresponding papers: Spearman's positive correlation is 100%. With such a strong result, this work evidences the utility of deep learning techniques for determining the computational history of science.
Amna Dridi, Mohamed Medhat Gaber, R. Muhammad Atif Azad, Jagdev Bhogal
IJCNN2
2019 EnSyth: A Pruning Approach to Synthesis of Deep Learning Ensembles
abstract
Deep neural networks have achieved state-of-art performance in many domains including computer vision, natural language processing and self-driving cars. However, they are very computationally expensive and memory intensive which raises significant challenges when it comes to deploy or train them on strict latency applications or resource-limited environments. As a result, many attempts have been introduced to accelerate and compress deep learning models, however the majority were not able to maintain the same accuracy of the baseline models. In this paper, we describe EnSyth, a deep learning ensemble approach to enhance the predictability of compact neural network's models. First, we generate a set of diverse compressed deep learning models using different hyperparameters for a pruning method, after that we utilise ensemble learning to synthesise the outputs of the compressed models to compose a new pool of classifiers. Finally, we apply backward elimination on the generated pool to explore the best performing combinations of models. On CIFAR-10, CIFAR-5 data-sets with LeNet-5, EnSyth outperforms the predictability of the baseline model.
Besher Alhalabi, Mohamed Medhat Gaber, Shadi Basurra
SMC2
2019 EACD: evolutionary adaptation to concept drifts in data streams
abstract
This paper presents a novel ensemble learning method based on evolutionary algorithms to cope with different types of concept drifts in non-stationary data stream classification tasks. In ensemble learning, multiple learners forming an ensemble are trained to obtain a better predictive performance compared to that of a single learner, especially in non-stationary environments, where data evolve over time. The evolution of data streams can be viewed as a problem of changing environment, and evolutionary algorithms offer a natural solution to this problem. The method proposed in this paper uses random subspaces of features from a pool of features to create different classification types in the ensemble. Each such type consists of a limited number of classifiers (decision trees) that have been built at different times over the data stream. An evolutionary algorithm (replicator dynamics) is used to adapt to different concept drifts; it allows the types with a higher performance to increase and those with a lower performance to decrease in size. Genetic algorithm is then applied to build a two-layer architecture based on the proposed technique to dynamically optimise the combination of features in each type to achieve a better adaptation to new concepts. The proposed method, called EACD, offers both implicit and explicit mechanisms to deal with concept drifts. A set of experiments employing four artificial and five real-world data streams is conducted to compare its performance with that of the state-of-the-art algorithms using the immediate and delayed prequential evaluation methods. The results demonstrate favourable performance of the proposed EACD method in different environments.
Hossein Ghomeshi, Mohamed Medhat Gaber, Yevgeniya Kovalchuk
Data Min. Knowl. Discov.2
2018 k-NN Embedding Stability for word2vec Hyper-Parametrisation in Scientific Text
Amna Dridi, Mohamed Medhat Gaber, R. Muhammad Atif Azad, Jagdev Bhogal
DS2
2018 An Agent-Based Collective Model to Simulate Peer Pressure Effect on Energy Consumption
Fatima Abdallah, Shadi Basurra, Mohamed Medhat Gaber
ICCCI (1)3
2018 Adaptive One-Class Ensemble-based Anomaly Detection: An Application to Insider Threats
abstract
The malicious insider threat is getting increased concern by organisations, due to the continuously growing number of insider incidents. The absence of previously logged insider threats shapes the insider threat detection mechanism into a one-class anomaly detection approach. A common shortcoming in the existing data mining approaches to detect insider threats is the high number of False Positives (FP) (i.e. normal behaviour predicted as anomalous). To address this shortcoming, in this paper, we propose an anomaly detection framework with two components: one-class modelling component, and progressive update component. To allow the detection of anomalous instances that have a high resemblance with normal instances, the one-class modelling component applies class decomposition on normal class data to create k clusters, then trains an ensemble of k base anomaly detection algorithms (One-class Support Vector Machine or Isolation Forest), having the data in each cluster used to construct one of the k base models. The progressive update component updates each of the k models with sequentially acquired FP chunks; segments of a predetermined capacity of FPs. It includes an oversampling method to generate artificial samples for FPs per chunk, then retrains each model and adapts the decision boundary, with the aim to reduce the number of future FPs. A variety of experiments is carried out, on synthetic data sets generated at Carnegie Mellon University, to test the effectiveness of the proposed framework and its components. The results show that the proposed framework reports the highest F1 measure and less number of FPs compared to the base algorithms, as well as it attains to detect all the insider threats in the data sets.
Diana Haidar, Mohamed Medhat Gaber
IJCNN2
2018 Deep imitation learning for 3D navigation tasks
abstract
Deep learning techniques have shown success in learning from raw high-dimensional data in various applications. While deep reinforcement learning is recently gaining popularity as a method to train intelligent agents, utilizing deep learning in imitation learning has been scarcely explored. Imitation learning can be an efficient method to teach intelligent agents by providing a set of demonstrations to learn from. However, generalizing to situations that are not represented in the demonstrations can be challenging, especially in 3D environments. In this paper, we propose a deep imitation learning method to learn navigation tasks from demonstrations in a 3D environment. The supervised policy is refined using active learning in order to generalize to unseen situations. This approach is compared to two popular deep reinforcement learning techniques: deep-Q-networks and Asynchronous actor-critic (A3C). The proposed method as well as the reinforcement learning methods employ deep convolutional neural networks and learn directly from raw visual input. Methods for combining learning from demonstrations and experience are also investigated. This combination aims to join the generalization ability of learning by experience with the efficiency of learning by imitation. The proposed methods are evaluated on 4 navigation tasks in a 3D simulated environment. Navigation tasks are a typical problem that is relevant to many real applications. They pose the challenge of requiring demonstrations of long trajectories to reach the target and only providing delayed rewards (usually terminal) to the agent. The experiments show that the proposed method can successfully learn navigation tasks from raw visual input while learning from experience methods fail to learn an effective policy. Moreover, it is shown that active learning can significantly improve the performance of the initially learned policy using a small number of active samples.
Ahmed Hussein 0001, Eyad Elyan, Mohamed Medhat Gaber, Chrisina Jayne
Neural Comput. Appl.3
2017 A Hybrid Agent-Based and Probabilistic Model for Fine-Grained Behavioural Energy Waste Simulation
abstract
Several agent-based and probabilistic models were proposed to simulate human behaviour, which is an important cause of high energy consumption in buildings. However, some of these models ignore behavioural energy waste at occupant level, and when they model it, they are based on small case studies and produce high level energy consumption data. This paper proposes a hybrid approach that integrates agent-based and probabilistic models to simulate behavioural energy waste at occupant level. The combination of the two approaches helps produce fine-grained data, and is based on large real data samples. The developed model was validated against realistic data. The results show that employment type have an effect on the energy consumption of households, which needs further investigation to quantify the effect and test other social parameters.
Fatima Abdallah, Shadi Basurra, Mohamed Medhat Gaber
ICTAI3
2017 Deep reward shaping from demonstrations
abstract
Deep reinforcement learning is rapidly gaining attention due to recent successes in a variety of problems. The combination of deep learning and reinforcement learning allows for a generic learning process that does not consider specific knowledge of the task. However, learning from scratch becomes more difficult when tasks involve long trajectories with delayed rewards. The chances of finding the rewards using trial and error become much smaller compared to tasks where the agent continuously interacts with the environment. This is the case in many real life applications which poses a limitation to current methods. In this paper we propose a novel method for combining learning from demonstrations and experience to expedite and improve deep reinforcement learning. Demonstrations from a teacher are used to shape a potential reward function by training a deep supervised convolutional neural network. The shaped function is added to the reward function used in deep-Q-learning (DQN) to perform off-policy training through trial and error. The proposed method is demonstrated on navigation tasks that are learned from raw pixels without utilizing any knowledge of the problem. Navigation tasks represent a typical AI problem that is relevant to many real applications and where only delayed rewards (usually terminal) are available to the agent. The results show that using the proposed shaped rewards significantly improves the performance of the agent over standard DQN. This improvement is more pronounced the sparser the rewards are.
Ahmed Hussein 0001, Eyad Elyan, Mohamed Medhat Gaber, Chrisina Jayne
IJCNN3
2017 On expressiveness and uncertainty awareness in rule-based classification for data streams
abstract
Mining data streams is a core element of Big Data Analytics. It represents the velocity of large datasets, which is one of the four aspects of Big Data, the other three being volume, variety and veracity. As data streams in, models are constructed using data mining techniques tailored towards continuous and fast model update. The Hoeffding Inequality has been among the most successful approaches in learning theory for data streams. In this context, it is typically used to provide a statistical bound for the number of examples needed in each step of an incremental learning process. It has been applied to both classification and clustering problems. Despite the success of the Hoeffding Tree classifier and other data stream mining methods, such models fall short of explaining how their results (i.e., classifications) are reached (black boxing). The expressiveness of decision models in data streams is an area of research that has attracted less attention, despite its paramount of practical importance. In this paper, we address this issue, adopting Hoeffding Inequality as an upper bound to build decision rules which can help decision makers with informed predictions (white boxing). We termed our novel method Hoeffding Rules with respect to the use of the Hoeffding Inequality in the method, for estimating whether an induced rule from a smaller sample would be of the same quality as a rule induced from a larger sample. The new method brings in a number of novel contributions including handling uncertainty through abstaining, dealing with continuous data through Gaussian statistical modelling, and an experimentally proven fast algorithm. We conducted a thorough experimental study using benchmark datasets, showing the efficiency and expressiveness of the proposed technique when compared with the state-of-the-art.
Thien Le, Frederic T. Stahl, Mohamed Medhat Gaber, João Bártolo Gomes, Giuseppe Di Fatta
Neurocomputing3
2017 A genetic algorithm approach to optimising random forests applied to class engineered data
Eyad Elyan, Mohamed Medhat Gaber
Inf. Sci.2
2017 A SOM-based Chan-Vese model for unsupervised image segmentation
Mohammed M. Abdelsamea, Giorgio Gnecco, Mohamed Medhat Gaber
Soft Comput.3
2016 An Outlier Ranking Tree Selection Approach to Extreme Pruning of Random Forests
Khaled Fawagreh, Mohamed Medhat Gaber, Eyad Elyan
EANN2
2016 Deep Active Learning for Autonomous Navigation
Ahmed Hussein 0001, Mohamed Medhat Gaber, Eyad Elyan
EANN2
2016 A Statistical Learning Method to Fast Generalised Rule Induction Directly from Raw Measurements
abstract
Induction of descriptive models is one of the most important technologies in data mining. The expressiveness of descriptive models are of paramount importance in applications that examine the causality of relationships between variables. Most of the work on descriptive models has concentrated on less expressive approaches such as clustering algorithms or rule-based approaches that are limited to a particular type of data, such as association rule mining for binary data. However, in many applications its important to understand the structure of the produced model for further human evaluation. In this research we present a novel generalised rule induction method that allows the induction of descriptive and expressive rules directly from both categorical and numerical features.
Thien Le, Frederic T. Stahl, Chris Wrench, Mohamed Medhat Gaber
ICMLA4
2016 A rule dynamics approach to event detection in Twitter with its application to sports and politics
Mariam Adedoyin-Olowe, Mohamed Medhat Gaber, Carlos J. Martín-Dancausa, Frederic T. Stahl, João Bártolo Gomes
Expert Syst. Appl.2
2016 A fine-grained Random Forests using class decomposition: an application to medical diagnosis
Eyad Elyan, Mohamed Medhat Gaber
Neural Comput. Appl.2
2015 Spatio-temporal analysis of Greenhouse Gas data via clustering techniques
abstract
Data mining allows for hidden patterns to be brought to light in large data sets. This paper aims to apply data mining on real-life data showing Greenhouse Gas Emissions for countries within the European Union. Greenhouse Gasses are gasses which are released into the atmosphere, trapping the infrared radiation from the sun and causing an effect called the Greenhouse Effect. This effect contributes to global Climate Change and is a topical issue. Using the K-means clustering algorithm, a model is produced in order to provide a deeper insight into the emissions of the industrial sectors of the UK, France and Italy. The model is intended to be of use to those in governmental authority when decisions on emissions within individual industries are to be made.
Alfredo Cuzzocrea, Mohamed Medhat Gaber, Staci Lattimer
CSCWD2
2015 Distributed Classification of Data Streams: An Adaptive Technique
Alfredo Cuzzocrea, Mohamed Medhat Gaber, Ary Mazharuddin Shiddiqi
DaWaK2
2015 Adaptive mobile activity recognition system with evolving data streams
Zahraa Said Abdallah, Mohamed Medhat Gaber, Bala Srinivasan 0002, Shonali Krishnaswamy
Neurocomputing2
2015 An efficient Self-Organizing Active Contour model for image segmentation
Mohammed M. Abdelsamea, Giorgio Gnecco, Mohamed Medhat Gaber
Neurocomputing3
2014 Extraction of Unexpected Rules from Twitter Hashtags and its Application to Sport Events
abstract
Twitter has become a dependable microblogging tool for real time information dissemination and newsworthy events broadcast. Its users sometimes break news on the network faster than traditional newsagents due to their presence at ongoing real life events at most times. Different topic detection methods are currently used to match Twitter posts to real life news of mainstream media. In this paper, we analyse tweets relating to the English FA Cup finals 2012 by applying our novel method named TRCM to extract association rules present in hash tag keywords of tweets in different time-slots. Our system identify evolving hash tag keywords with strong association rules in each time-slot. We then map the identified hash tag keywords to event highlights of the game as reported in the ground truth of the main stream media. The performance effectiveness measure of our experiments show that our method perform well as a Topic Detection and Tracking approach.
Mariam Adedoyin-Olowe, Mohamed Medhat Gaber, Carlos J. Martín-Dancausa, Frederic T. Stahl
ICMLA2
2014 Diversified Random Forests Using Random Subspaces
Khaled Fawagreh, Mohamed Medhat Gaber, Eyad Elyan
IDEAL2
2014 Automatic Content Related Feedback for MOOCs Based on Course Domain Ontology
Safwan Shatnawi, Mohamed Medhat Gaber, Ella Haig
IDEAL2
2014 Adaptive data stream mining for wireless sensor networks
abstract
Data stream mining in wireless sensor networks has many important applications. Realizing these applications is faced by resource constraints of the sensor nodes that form the network. Adaptation to availability of resources is crucial to the success of these applications. In this paper, we propose a distributed data stream classification technique that has been tested on a real sensor network platform, namely, Sun SPOT. Experimental results evidenced the applicability of our technique to operate in such an environment of scarce resources.
Alfredo Cuzzocrea, Mohamed Medhat Gaber, Ary Mazharuddin Shiddiqi
IDEAS2
2014 Mining Recurring Concepts in a Dynamic Feature Space
abstract
Most data stream classification techniques assume that the underlying feature space is static. However, in real-world applications the set of features and their relevance to the target concept may change over time. In addition, when the underlying concepts reappear, reusing previously learnt models can enhance the learning process in terms of accuracy and processing time at the expense of manageable memory consumption. In this paper, we propose mining recurring concepts in a dynamic feature space (MReC-DFS), a data stream classification system to address the challenges of learning recurring concepts in a dynamic feature space while simultaneously reducing the memory cost associated with storing past models. MReC-DFS is able to detect and adapt to concept changes using the performance of the learning process and contextual information. To handle recurring concepts, stored models are combined in a dynamically weighted ensemble. Incremental feature selection is performed to reduce the combined feature space. This contribution allows MReC-DFS to store only the features most relevant to the learnt concepts, which in turn increases the memory efficiency of the technique. In addition, an incremental feature selection method is proposed that dynamically determines the threshold between relevant and irrelevant features. Experimental results demonstrating the high accuracy of MReC-DFS compared with state-of-the-art techniques on a variety of real datasets are presented. The results also show the superior memory efficiency of MReC-DFS.
João Bártolo Gomes, Mohamed Medhat Gaber, Pedro A. C. Sousa, Ernestina Menasalvas Ruiz
IEEE Trans. Neural Networks Learn. Syst.2
2013 An Information-Theoretic Approach for Setting the Optimal Number of Decision Trees in Random Forests
abstract
Data Classification is a process within the Data Mining and Machine Learning field which aims at annotating all instances of a dataset by so-called class labels. This involves in creating a model from a training set of data instances which are already labeled, possibly being this model also used to define the class of data instances which are not classified already. A successful way of performing the classification process is provided by the algorithm Random Forests (RF), which is itself a type of Ensemble-based Classifier. An ensemble-based classifier increases the accuracy of the class label assigned to a data instance by using a set of classifiers that are modeled on different, but possibly overlapping, instance sets, and then combining the so-obtained intermediate classification results. To this end, RF particularly makes use of a number of decision trees to classify an instance, then taking the majority of votes from these trees as the final classifier. The latter one is a critical task of algorithm RF, which heavily impacts on the accuracy of the final classifier. In this paper, we propose a variation of algorithm RF, namely adjusting one of the two parameters that RF takes, the number of decision trees, dependant on a meaningful relation between the dataset predictive power rating and the number of trees itself, with the goal of improving accuracy and performance of the algorithm. This is finally demonstrated by our comprehensive experimental evaluation on several clean datasets.
Alfredo Cuzzocrea, Shane Leo Francis, Mohamed Medhat Gaber
SMC3
2013 Interactive self-adaptive clutter-aware visualisation for mobile data mining
Mohamed Medhat Gaber, Shonali Krishnaswamy, Brett Gillick, Hasnain AlTaiar, Nicholas Nicoloudis, Jonathan Liono, Arkady B. Zaslavsky
J. Comput. Syst. Sci.1
2012 Mobile Activity Recognition Using Ubiquitous Data Stream Mining
João Bártolo Gomes, Shonali Krishnaswamy, Mohamed Medhat Gaber, Pedro A. C. Sousa, Ernestina Menasalvas Ruiz
DaWaK3
2012 GARF: Towards Self-optimised Random Forests
Mohamed Bahy Bader-El-Den, Mohamed Medhat Gaber
ICONIP (2)2
2012 StreamAR: Incremental and Active Learning with Evolving Sensory Data for Activity Recognition
abstract
Activity recognition focuses on inferring current user activities by leveraging sensory data available on todayÕs sensor rich environment. Supervised learning has been applied pervasively for activity recognition. Typical activity recognition techniques process sensory data based on point-by-point approaches. In this paper, we propose a novel cluster-based classification for activity recognition Systems, termed StreamAR. The system incorporates incremental and active learning for mining user activities in data streams. The novel approach processes activities as clusters to build a robust classification framework. StreamAR integrates supervised, unsupervised and active learning and applies hybrid similarity measures technique for recognising activities. Extensive experimental results using real activity recognition datasets have evidenced that our new approach shows improved performance over other existing state-of-the-art learning methods.
Zahraa Said Abdallah, Mohamed Medhat Gaber, Bala Srinivasan 0002, Shonali Krishnaswamy
ICTAI2
2012 Deploying Mobile Software Agents for Distributed Data Mining on Wireless Sensor Networks: A Comparative Analysis
abstract
In this paper, we investigate the deployment of mobile software agents in the field of distributed data mining in wireless sensor networks. The paper provides a survey of various current mobile software agents based Distributed Data Mining systems like BODHI, PADMA, Papyrus, JAM, InfoSleuth and DKN and discusses their salient features. Flexibility and autonomy of mobile software agents make them a suitable tool for software deployment in wireless sensor networks. We then discuss the deployment of mobile software agents on two popular sensor architectures Crossbow Mica2 motes and Sun SPOT. The limitation of the Sun SPOT environment in supporting mobile software agents due to the absence of serializable interface on Sun SPOTs is discussed. The common pitfalls while deploying mobile software agents using Agilla agent framework on Crossbow motes are also discussed.
Ranjani Nagarajan, Alfredo Cuzzocrea, Mohamed Medhat Gaber
ICTAI3
2012 Mobile Sentiment Analysis
abstract
Mobile devices play a significant part in a user’s communication methods and much data that they read and write is received and sent via mobile phones, for instance SMS messages, e-mails, Twitter tweets and social media networking feeds. One of the main goals is to make people aware of how much negative and positive content they read and write via their mobile phones. Existing sentiment analysis applications perform sentiment analysis on downloaded data from mobile phones or use an application installed on another computer to perform the analysis. The sentiment analysis described in this paper is to be performed locally on the mobile phone enabling immediate and private analysis of personal messages and social media contents, allowing the users to be able to reason about their mood and stress level that may be affected by what they had been receiving. Experimental results showed the effectiveness of the proposed system on Android smartphones with varying computational capabilities.
Lorraine Chambers, Erik Tromp, Mykola Pechenizkiy, Mohamed Medhat Gaber
KES4
2012 Optimisation of Ensemble Classifiers using Genetic Algorithm
Mohamed Medhat Gaber, Mohamed Bahy Bader-El-Den
KES1
2012 Incorporating Farthest Neighbours in Instance Space Classification
abstract
The nearest neighbour (NN) classifier is often known as a ‘lazy’ approach but it is still widely used particularly in the systems that require pattern matching. Many algorithms have been developed based on NN in an attempt to improve classification accuracy and to reduce the time taken, especially in large data sets. This paper proposes a new classification technique based on k-Nearest Neighbour (k-NN), called k-Nearest & Farthest Neighbours (k-NFN). Farthest neighbours are used to identify classes that an unseen record may not belong to and are considered with the nearest neighbours in the classification decision. Two neighbour voting systems are also proposed to further improve k-NN and k-NFN accuracy. The first uses a ranking system and the second uses a spectrum to consider how near or far a neighbour actually is. The accuracy of our three proposed k-NFN techniques and k-NN are compared using the standard ten cross fold validation experiments on a number of real data sets, evidencing the superiority of our proposed techniques in terms of accuracy.
Daniel Vaccaro-Senna, Mohamed Medhat Gaber
KES2
2012 MARS: A Personalised Mobile Activity Recognition System
abstract
Mobile activity recognition focuses on inferring the current activities of a mobile user by leveraging the sensory data that is available on today's smart phones. The state of the art in mobile activity recognition uses traditional classification techniques. Thus, the learning process typically involves: i) collection of labelled sensory data that is transferred and collated in a centralised repository, ii) model building where the classification model is trained and tested using the collected data, iii) a model deployment stage where the learnt model is deployed on-board a mobile device for identifying activities based on new sensory data. In this paper, we demonstrate the Mobile Activity Recognition System (MARS) where for the first time the model is built and continuously updated on-board the mobile device itself using data stream mining. The advantages of the on-board approach are that it allows model personalisation and increased privacy as the data is not sent to any external site. Furthermore, when the user or its activity profile changes MARS enables quick model adaptation. One of the stand out features of MARS is that training/updating the model takes less than 30 seconds per activity. MARS has been implemented on the Android platform to demonstrate that it can achieve accurate mobile activity recognition. Moreover, we can show in practice that MARS quickly adapts to user profile changes while at the same time being scalable and efficient in terms of consumption of the device resources.
João Bártolo Gomes, Shonali Krishnaswamy, Mohamed Medhat Gaber, Pedro A. C. Sousa, Ernestina Menasalvas Ruiz
MDM3
2012 Mobile Data Stream Mining: From Algorithms to Applications
abstract
This paper presents an overview of the current state-of-the-art in mobile data stream mining. This area of mobile data stream mining is significant for a number of new application domains such as mobile crowd sensing and mobile activity recognition. The paper presents the strategies and techniques for adaptation that are essential in order to perform real-time, continuous data mining on mobile devices. We present an overview of the algorithms research in this area. Finally, we discuss the key toolkits, systems and applications of mobile data stream mining.
Shonali Krishnaswamy, João Gama 0001, Mohamed Medhat Gaber
MDM3
2011 KB-CB-N classification: Towards unsupervised approach for supervised learning
abstract
Data classification has attracted considerable research attention in the field of computational statistics and data mining due to its wide range of applications. K Best Cluster Based Neighbour (KB-CB-N) is our novel classification technique based on the integration of three different similarity measures for cluster based classification. The basic principle is to apply unsupervised learning on the instances of each class in the dataset and then use the output as an input for the classification algorithm to find the K best neighbours of clusters from the density, gravity and distance perspectives. Clustering is applied as an initial step within each class to find the inherent in-class grouping in the dataset. Different data clustering techniques use different similarity measures. Each measure has its own strength and weakness. Thus, combining the three measures can benefit from the strength of each one and eliminate encountered problems of using an individual measure. Extensive experimental results using eight real datasets have evidenced that our new technique typically shows improved or equivalent performance over other existing state-of-the-art classification methods.
Zahraa Said Abdallah, Mohamed Medhat Gaber
CIDM2
2011 Advances in data stream mining for mobile and ubiquitous environments
abstract
The tutorial presents the state-of-the-art in mobile and ubiquitous data stream mining and discusses open research problems, issues, and challenges in this area.
Shonali Krishnaswamy, João Gama 0001, Mohamed Medhat Gaber
CIKM3
2011 Context-Aware Collaborative Data Stream Mining in Ubiquitous Devices
João Bártolo Gomes, Mohamed Medhat Gaber, Pedro A. C. Sousa, Ernestina Menasalvas Ruiz
IDA2
2011 Distributed Classification for Pocket Data Mining
Frederic T. Stahl, Mohamed Medhat Gaber, Han Liu 0002, Max Bramer, Philip S. Yu
ISMIS2
2011 RA-SAX: Resource-Aware Symbolic Aggregate Approximation for Mobile ECG Analysis
abstract
There is a growing focus on 24/7 cardiac monitoring that leverages state of the art mobile phones and commercial-off-the-shelf (COTS) wearable bio-sensors. While many signal processing techniques for mobile ECG analysis have been developed, these techniques tend to be computationally intensive. In this paper, we propose, develop and evaluate a resource-aware and energy-efficient time series analysis technique for real-time ECG analysis on mobile devices based on the well-known SAX (Symbolic Aggregate Approximation) representation for time series termed RA-SAX.
Hossein Tayebi, Shonali Krishnaswamy, Agustinus Borgy Waluyo, Abhijat Sinha, Mohamed Medhat Gaber
Mobile Data Management (1)5
2011 Energy conservation in wireless sensor networks: a rule-based approach
Suan Khai Chong, Mohamed Medhat Gaber, Shonali Krishnaswamy, Seng W. Loke
Knowl. Inf. Syst.2
2010 Corona: Energy-Efficient Multi-query Processing in Wireless Sensor Networks
Raymes Khoury, Tim Dawborn, Bulat Gafurov, Glen Pink, Edmund Tse, Quincy Tse, Khaled Almiani, Mohamed Medhat Gaber, Uwe Röhm, Bernhard Scholz
DASFAA (2)8
2010 Adaptive Clutter-Aware Visualization for Mobile Data Stream Mining
abstract
There is an emerging focus on real-time data stream analysis on mobile devices. A wide range of data stream processing applications are targeted to run on mobile handheld devices with limited computational capabilities such as patient monitoring, driver monitoring, providing real-time analysis and visualization for emergency and disaster management, real-time optimization for courier pick-up and delivery etc. There are many challenges in visualization of the analysis/data stream mining results on a mobile device. These include coping with the small screen real-estate and effective presentation of highly dynamic and real-time analysis. This paper proposes a generic theory for visualization on small screens that we term Adaptive Clutter Reduction ACR. Based on ACR, we have developed and experimentally validated a novel data stream clustering result visualization technique that we term Clutter-Aware Clustering Visualizer (CACV). Experimental results on both synthetic and real datasets using the Google Andriod platform are presented proving the effectiveness of the proposed techniques.
Mohamed Medhat Gaber, Shonali Krishnaswamy, Brett Gillick, Nicholas Nicoloudis, Jonathan Liono, Hasnain AlTaiar, Arkady B. Zaslavsky
ICTAI (2)1
2010 Pocket Data Mining: Towards Collaborative Data Mining in Mobile Computing Environments
abstract
Pocket Data Mining (PDM) is our new term describing collaborative mining of streaming data in mobile and distributed computing environments. With sheer amounts of data streams are now available for subscription on our smart mobile phones, the potential of using this data for decision making using data stream mining techniques has now been achievable owing to the increasing power of these handheld devices. Wireless communication among these devices using Bluetooth and WiFi technologies has opened the door wide for collaborative mining among the mobile devices within the same range that are running data mining techniques targeting the same application. This paper proposes a new architecture that we have prototyped for realizing the significant applications in this area. We have proposed using mobile software agents in this application for several reasons. Most importantly the autonomic intelligent behaviour of the agent technology has been the driving force for using it in this application. Other efficiency reasons are discussed in details in this paper. Experimental results showing the feasibility of the proposed architecture are presented and discussed.
Frederic T. Stahl, Mohamed Medhat Gaber, Max Bramer, Philip S. Yu
ICTAI (2)2
2009 A Weighted Approach to Partial Matching for Mobile Reasoning
Luke Steller, Shonali Krishnaswamy, Mohamed Medhat Gaber
ISWC3
2009 Knowledge discovery from data streams
João Gama 0001, Auroop R. Ganguly, Olufemi A. Omitaomu, Ranga Raju Vatsavai, Mohamed Medhat Gaber
Intell. Data Anal.5
2009 Context-aware adaptive data stream mining
abstract
In resource-constrained devices, adaptation of data stream processing to variations of data rates and availability of resources is crucial for consistency and continuity of running applications. However, to enhance and maximize the benefits of adapta
Pari Delir Haghighi, Arkady B. Zaslavsky, Shonali Krishnaswamy, Mohamed Medhat Gaber, Seng W. Loke
Intell. Data Anal.4
2009 Enabling Scalable Semantic Reasoning for Mobile Services
abstract
With the emergence of high-end smart phones/PDAs there is a growing opportunity to enrich mobile/pervasive services with semantic reasoning. This article presents novel strategies for optimising semantic reasoning for realizing semantic applications and services on mobile devices. We have developed the mTableaux algorithm which optimizes the reasoning process to facilitate service selection. We present comparative experimental results which show that mTableaux improves the performance and scalability of semantic reasoning for mobile devices.
Luke Steller, Shonali Krishnaswamy, Mohamed Medhat Gaber
Int. J. Semantic Web Inf. Syst.3
2009 An analytical study of central and in-network data processing for wireless sensor networks
Mohamed Medhat Gaber, Uwe Röhm, Karel Herink
Inf. Process. Lett.1
2008 Clustering Distributed Time Series in Sensor Networks
abstract
Event detection is a critical task in sensor networks, especially for environmental monitoring applications. Traditional solutions to event detection are based on analyzing one-shot data points, which might incur a high false alarm rate because sensor data is inherently unreliable and noisy. To address this issue, we propose a novel Distributed Single-pass Incremental Clustering (DSIC) technique to cluster the time series obtained at sensor nodes based on their underlying trends. In order to achieve scalability and energy-efficiency, our DSIC technique uses a hierarchical structure of sensor networks as the underlying infrastructure. The algorithm first compresses the time series produced at individual sensor nodes into a compact representation using Haar wavelet transform, and then, based on dynamic time warping distances, hierarchically groups the approximate time series into a global clustering model in an incremental manner. Experimental results on both real data and synthetic data demonstrate that our DSIC algorithm is accurate, energy-efficient and robust with respect to network topology changes.
Jie Yin 0001, Mohamed Medhat Gaber
ICDM2
2007 Resource-aware Online Data Mining in Wireless Sensor Networks
abstract
Data processing in wireless sensor networks often relies on high-speed data stream input, but at the same time is inherently constrained by limited resource availability. Thus, energy efficiency and good resource management are vital for in-network processing techniques. We propose enabling resource-awareness for in-network processing algorithms by means of a resource monitoring component and designed a corresponding framework. As proof of concept, we implement an online clustering algorithm, which uses the resource monitor to adapt to resource availability, on the Sun SPOT sensor nodes from Sun Microsystem. We refer to this adaptive clustering algorithm as extended resource-aware cluster (ERA-cluster). Finally, we report on the outcome of several experiments to evaluate the validity of our approach in terms of resource adaptiveness and accuracy of the ERA-cluster. Results show that ERA-cluster can effectively adapt to resource availability while maintaining acceptable level of accuracy.
Nhan Duc Phung, Mohamed Medhat Gaber, Uwe Röhm
CIDM2
2007 On the Integration of Data Stream Clustering into a Query Processor for Wireless Sensor Networks
abstract
We discuss the integration of on-line data stream clustering into a distributed query processor for wireless sensor networks (WSNs). Our approach is to combine an adaptive clustering algorithm with in-network data processing by introducing specialised query operators that implement stateful stream processing. We have implemented a testbed for an on-line clustering algorithm as part of a query processing system for the Sun SPOT sensor network platform. The paper discusses the design alternatives for continuous stream clustering in WSNs and for the integration of the resource-awareness to be able to trade result accuracy for resource consumption.
Uwe Röhm, Bernhard Scholz, Mohamed Medhat Gaber
MDM3
2007 A fuzzy approach for interpretation of ubiquitous data stream clustering and its application in road safety
Osnat Horovitz, Shonali Krishnaswamy, Mohamed Medhat Gaber
Intell. Data Anal.3
2005 Making Sense of Ubiquitous Data Streams - A Fuzzy Logic Approach
Osnat Horovitz, Mohamed Medhat Gaber, Shonali Krishnaswamy
KES (2)2
2004 Towards an Adaptive Approach for Mining Data Streams in Resource Constrained Environments
Mohamed Medhat Gaber, Arkady B. Zaslavsky, Shonali Krishnaswamy
DaWaK1