VLDB 2026 Research / reviewers in the wild / expert
Ali Bou Nassif
dblp:07/8331 · also Ali J. Bou Nassif
· DBLP profile ↗
72ranked-venue papers
16as first author
39since 2021 · last 2025
0000-0003-1570-0897ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 13 first-author · 20 since 2021Software engineering, systems software and programming languages · 12 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 10 since 2021Computer networks · 3 · 1 since 2021Security and privacy · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Deep Learning Framework for Melanoma Subtype Classification from mRNA Expression ProfilesabstractMelanoma is the most aggressive form of skin cancer, and its early and accurate diagnosis is crucial for patient treatment. Gene expression profiling has become a powerful tool to capture the molecular portraits of tumors, but it suffers from the difficulties of small sample size and high dimensionality. In this study, we propose a lightweight one-dimensional convolutional neural network (1D-CNN) for melanoma sample classification from mRNA expression data. The proposed framework involves mutual information for feature selection to reduce dimensionality, supported by preprocessing steps such as log transformation, low-variance gene removal, and filtering of weakly expressed genes to improve data quality before modeling. The proposed CNN is benchmarked against classical machine learning (ML) methods- K-nearest neighbors (KNN), Random Forest (RF), and AdaBoost-across multiple gene subsets (150, 300, and 450). Experimental results indicate that when 300 genes are selected, the proposed CNN demonstrates a consistent improvement with an F 1 -score of 97% and a peak accuracy of 96%, while providing robustness across different feature sizes. Compared with classical models, the proposed CNN delivers a more stable trade-off between precision and recall. The results show that streamlined deep learning (DL) models enable accurate and efficient melanoma subtype classification, which highlights potential for clinical application. Abdulaziz Abdullatif, Ali Bou Nassif, Bassel Al Homssi |
DeSE | 2 |
| 2025 | A Feedback-Enhanced Effort-Aware Just in Time Defect Prediction PipelineabstractSoftware defect prediction aims to identify faults in code before they cause failures. It helps developers locate problematic modules early in the development cycle. Early detection reduces the cost of fixing defects and prevents them from affecting users. Just-in-Time defect prediction is the concept of detecting buggy commits early on in the development stage rather than in the release. While recent research focuses on improving the accuracy of pre-diction faulty commits, fewer studies explore how to do that task with minimal effort. In this study, we propose an effort-aware JIT defect prediction pipeline that utilizes code pre-trained models finetuned with Low Rank Adaptation (LoRA) and a feedback loop mechanism. The aim is to improve fault detection while reducing effort. Our experiments evaluate four different transformer-based models: CodeBERT, JavaBERT, UniXcoder, and RoBERTa on ApacheJIT. Our results illustrate CodeBERT and UniXcoder are benefiting the most from the feedback loop with the highest gain in F1 score, where CodeBERT highest gain is 0.14 and UniXcoder is 0.02. RoBERTa performed the best across experiments with an F1 score of 0.73. However, Popt metric remains static with limited gain that reached maximum 0.02, suggesting that further experiments are needed to improve performance. Huda AlGhussein, Abdullah Mustafa, Manar Abu Talib, Mohammed Azzeh, Ali Bou Nassif |
DeSE | 5 |
| 2025 | Real-Time Low-Cost Automatic Collision Detection with Owner Notification for Parked VehiclesabstractThis study introduces a fully automated collision detection and notification system specifically engineered to safeguard parked vehicles against accidental impacts in densely populated areas such as commercial parking lots and urban streets. The system employs an integrated network of front and rear cameras, proximity sensors, and vibration sensors to provide continuous environmental monitoring around a stationary vehicle. When a foreign object or vehicle encroaches within a predefined proximity, the system initiates real-time surveillance by activating on-board cameras. Simultaneously, visual alert mechanisms, such as high-intensity flashing lights, are triggered to attract the attention of nearby drivers and prevent potential collisions. In the event of physical contact, the system immediately begins continuous video recording, capturing high-resolution footage of the incident. This evidence is securely transmitted to the vehicle owner's mobile device via a dedicated application, delivering instant notification and remote access to the recorded material. The design emphasizes affordability and accessibility, ensuring that advanced vehicle protection is available to a broad user base. By combining proactive collision deterrence with post-incident documentation and real-time communication, the proposed system offers a comprehensive and practical solution to mitigate the risk and consequences of parked vehicle collisions. Experimental validation confirms the system's reliability, responsiveness, and effectiveness in real-world parking scenarios, demonstrating its value as a robust enhancement to vehicular safety infrastructure. Antanios Kaissar, Sam Ansari, Soliman Mahmoud, Khawla Alnajjar, Eqab R. F. Almajali, Anwar Jarndal, Ali Bou Nassif, Youssef Mansour, Abir Jaafar Hussain |
DeSE | 7 |
| 2025 | Semi-Supervised Learning-Based Genetic Biomarkers Dataset for Multiple-Stage Hepatocellular Carcinoma PredictionabstractLiver cancer is a complex disease responsible for a high number of deaths across the globe each year, making automated solutions for liver cancer classification urgent. The most common form of liver cancer is hepatocellular carcinoma (HCC), accounting for over 90 % of liver cancer cases. There is a distinct lack of publicly available HCC datasets utilizing genomic data, which is necessary for training artificial intelligence (AI) models for automated HCC classification. This study proposes constructing a multi-stage HCC dataset using XGBoost and Semi-Supervised learning on three separate datasets of genomic biomarkers, utilizing their existing labels in the Semi-Supervised learning process to label the proposed dataset. The proposed dataset consists of 770 patient samples in total, categorized into five classes that represent normal tissue alongside different stages of HCC. Each sample in the dataset consists of$\mathbf{1 1, 1 5 0}$different gene expression levels. The XGBoost model demonstrated a final classification accuracy of 96.5 % during the Semi-Supervised learning process. Ahmed Ammar Kubba, Manar Abu Talib, Jibran Sualeh Muhammad, Ali Bou Nassif, Abdalla Sayed Mohamed, Darko Castven, Jens U. Marquardt |
DeSE | 4 |
| 2025 | Enhancing RLHF in LLMs: Comparing BERT, XGBoost, and Deep Models for Bias DetectionabstractThis study explores bias detection in large language models (LLMs) to support Reinforcement Learning from Human Feedback (RLHF) pipelines aimed at reducing social bias. We compare three classifiers-a frozen BERT-based model with a classification head, a custom deep learning model, and an XGBoost model-on the StereoSet dataset. This dataset includes five bias categories: gender, profession, race, religion, and no bias. To address class imbalance, we generated synthetic samples using Google's gemma-2-9b-it model. Results show that the BERT-based model achieved the highest overall accuracy 0.92 and outperformed others across all categories, particularly in detecting subtle biases. The deep learning model was ranked second with 0.85 accuracy, while XGBoost lagged at 0.75, struggling with no-bias detection. These findings confirm that architecture choice and training strategy significantly impact bias classification and reinforce the importance of integrating reliable detection models into RLHF workflows. Walid Torfa, Omar Elgendy, Ali Bou Nassif |
DeSE | 3 |
| 2025 | Unlocking language boundaries: AraCLIP - transforming Arabic language and image understanding through cross-lingual models
Muhammad Al-Barham, Imad Afyouni, Khalid Almubarak, Ayad Mashaan Turky, Ibrahim Abaker Targio Hashem, Ali Bou Nassif, Ismail Shahin, Ashraf Elnagar |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | Chrominance and Luminance: a study to detect deepfakesabstractA deepfake is a fast attack technique that has evolved rapidly over the past several years. It only takes one person to synthesize thousands of photorealistic images in a few hours or to manipulate a large number of videos. The creation of fake faces through image tampering has been identified as abusive media and could result in major ethical, legal, or political consequences. Existing known methods, such as Generative Adversarial Network (GAN), have simplified the synthetization of such images and can be used for detection purposes. However, current solutions and existing methods remain vulnerable and cannot stop deepfakes from spreading. In this paper, two novel solutions are proposed, named ColorDense and LightDense. They both use the DenseNet network backbone to determine fake face images from real face images. The two proposed models incorporate luminance and color signals from multi-color spaces to spot manipulated images obtained from the most recent deepfake datasets. Chrominance components of face images are considered to give a deep representation, while luminance depends on measuring the perceived brightness levels of the image. The system extracts the important features using the DenseNet network. Afterwards, it then passes the feature vector to a classification system based on neuronal techniques such as Multilayer Perceptron (MLP), Random Forest (RF), Support Vector Machine (SVM), and Radial Basis Function (RBF) as deepfake image detectors. Several experiments were conducted in this paper and the proposed system achieved over 98% accuracy in classifying deepfake images. Manar Abu Talib, Qassim Nasir, Ali Bou Nassif, Norah Ba Fadhl, Omar Mohamed Gouda |
Multim. Tools Appl. | 3 |
| 2025 | Correction: Software effort estimation using convolutional neural network and fuzzy clustering
Mohammad Azzeh, Abedalrhman Alkhateeb, Ali Bou Nassif |
Neural Comput. Appl. | 3 |
| 2025 | Axolotl inspired metaheuristic for software implementation effort prediction
Cuauhtémoc López Martín, Yenny Villuendas-Rey, Ali Bou Nassif, Noé Oswaldo Rodríguez-Rodríguez |
J. Supercomput. | 3 |
| 2024 | Enhanced Intrusion Detection in Cloud Security by Optimizing Classification AlgorithmsabstractThe primary threat to network security is detecting intrusive activities over the network. Because network models and bandwidths are rapidly evolving, the effectiveness of Intrusion Detection Systems (IDSs) requires intuitive Artificial Intelligence (AI). Furthermore, studies in recent years have shown that parsing network traffic data into an IDS is a productive method for detecting security breaches. Detecting malicious assaults on any system is the first step in securing it. To detect whether a network traffic is intrusive or normal, classification algorithms support the IDSs in classifying network traffic types. In this paper, we display the comparative results obtained from conducting the CSE-CICIDS2018 dataset base, which represents in information security research the most reliable and up-to-date dataset. Applying the new AI trends to evaluate different feature selections and machine learning techniques. The classification algorithms such as Multilayer Perceptron (MLP), Naïve Bayes, Radial Basis Neural Network (RBF), Support Vector Machine (SVM), k-Nearest Neighbours (KNN), and Decision Tree (J48) algorithms come up to show the unwanted features to be removed from nonbehavioral intrusions and identify the most beneficial combination of qualities for a specific statistical report depending on the accuracy, precision, recall, and ROC curve, leading to the best model used with clarification. Results figured that the best classification algorithm is KNN with an accuracy of 100%. Mohammad Ziad Mizher, Ali Bou Nassif |
BDCAT | 2 |
| 2024 | Deep Learning Algorithms in Aircraft Detection and Classification: An Analytical SurveyabstractAircraft detection has become a very crucial aspect in the aviation industry due to its vital role in identifying unauthorized aircraft trespassing into restricted areas. Due to their significant role in decision-making, these detection systems must have a high level of accuracy. In this paper, we address deep learning feature extraction and aircraft detection methods. This study aims to explore deep learning-based feature extraction techniques and airplane identification techniques. The applications examined stop unauthorized aircraft from entering restricted airspace by utilizing a range of deep learning models. The study results and basic information on feature types, image types, and detection techniques are presented in this paper. Noora Alkharji, Hessa Almazrouei, Shaikha Alzaabi, Ali Bou Nassif, Mariame Elsalhy, Manar Abu Talib |
DeSE | 4 |
| 2024 | Impact of Outliers on Regression and Classification Models: An Empirical AnalysisabstractIn recent years, the proliferation of data and sensor measurements in various scientific fields, particularly within the realm of the Internet of Things, has opened new avenues for knowledge extraction through advanced data analysis techniques. However, the presence of outliers and anomalies poses significant challenges, leading to inaccuracies that can compromise analytical outcomes. Outliers are defined as data points that deviate markedly from other observations, often resulting from measurement errors or inconsistencies within the dataset. Their detection and removal during the data cleaning process are crucial for enhancing data quality and ensuring robust analysis. This study systematically investigates the impact of outliers and their detection on the accuracy and performance of various machine learning algorithms and statistical models in regression and classification tasks. A series of MATLAB simulations is conducted on standard datasets to evaluate the effects of outliers and validate the performance of different methodologies. The findings highlight the critical importance of effective outlier detection, demonstrating a marked improvement in the accuracy and reliability of analytical results. Sam Ansari, Ali Bou Nassif, Soliman A. Mahmoud, Sohaib Majzoub, Eqab R. F. Almajali, Anwar Jarndal, Talal Bonny, Khawla Alnajjar, Abir Jaafar Hussain |
DeSE | 2 |
| 2024 | Aircraft-Type Classification Using Deep Learning AlgorithmsabstractIn the field of artificial intelligence (AI), deep learning (DL) has become an essential methodology with many uses, particularly in the aviation sector. The efficiency and effectiveness of tasks connected to aircraft recognition and classification are considerably improved by DL. The goal of this study is to examine and deal with problems related to efficiency and accuracy in DL models for identifying and classifying aircrafts. In this article, we suggest a solution that utilizes a customized convolutional neural network (CNN) model that our team has created and assess its performance in comparison to other deep learning (DL) models like VGG16 and ResNet50. Accuracy, recall, precision, and f1-score are among the metrics used for evaluation. The outcomes of this study demonstrate significant gains and offer insightful information for future endeavors in the aviation field. Noora Alkharji, Hessa Almazrouei, Shaikha Alzaabi, Ali Bou Nassif, Mariame Elsalhy, Manar Abu Talib |
IS | 4 |
| 2024 | Effect of feature optimization on performance of machine learning models for predicting traffic incident duration
Lubna Obaid, Khaled Hamad, Mohamad Ali Khalil, Ali Bou Nassif |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Software effort estimation using convolutional neural network and fuzzy clustering
Mohammad Azzeh, Abedalrhman Alkhateeb, Ali Bou Nassif |
Neural Comput. Appl. | 3 |
| 2024 | Comparative study of ML models for IIoT intrusion detection: impact of data preprocessing and balancing
Abdulrahman Mahmoud Eid, Bassel Soudan, Ali Bou Nassif, MohammadNoor Injadat |
Neural Comput. Appl. | 3 |
| 2024 | Correction: Comparative study of ML models for IIoT intrusion detection: impact of data preprocessing and balancing
Abdulrahman Mahmoud Eid, Bassel Soudan, Ali Bou Nassif, MohammadNoor Injadat |
Neural Comput. Appl. | 3 |
| 2024 | Enhancing intrusion detection in IIoT: optimized CNN model with multi-class SMOTE balancing
Abdulrahman Mahmoud Eid, Bassel Soudan, Ali Bou Nassif, MohammadNoor Injadat |
Neural Comput. Appl. | 3 |
| 2024 | ViT-LSTM synergy: a multi-feature approach for speaker identification and mask detection
Ali Bou Nassif, Ismail Shahin, Mohamed Bader, Abdelfatah Hassan Ahmed, Naoufel Werghi |
Neural Comput. Appl. | 1 |
| 2024 | Emotional speaker identification using PCAFCM-deepforest with fuzzy logic
Ali Bou Nassif, Ismail Shahin, Nawel Nemmour |
Neural Comput. Appl. | 1 |
| 2024 | Parameter-efficient fine-tuning of pre-trained code models for just-in-time defect prediction
Manar Abu Talib, Ali Bou Nassif, Mohammad Azzeh, Yaser Alesh, Yaman Afadar |
Neural Comput. Appl. | 2 |
| 2023 | Speaker identification from emotional and noisy speech using learned voice segregation and speech VGG
Shibani Hamsa, Ismail Shahin, Youssef Iraqi, Ernesto Damiani, Ali Bou Nassif, Naoufel Werghi |
Expert Syst. Appl. | 5 |
| 2023 | Examining the performance of kernel methods for software defect prediction based on support vector machine
Mohammad Azzeh, Yousef Elsheikh, Ali Bou Nassif, Lefteris Angelis |
Sci. Comput. Program. | 3 |
| 2022 | On Selection of Optimal Kernel Function for Software Defect PredictionabstractSoftware defect prediction is used to identify the most likely defect-prone modules. This activity can help the quality assurance team to distribute limited resources efficiently as well as reduce the time and cost of the testing phase. Support Vector Machine (SVM) has been widely used to build software defect prediction models. Nevertheless, the accuracy of such method depends on multiple factors such as the type of SVM method, choice of kernel functions and dimensionality of the dataset. This paper continues along that research line and examines the impact and stability of four kernel functions and feature dimensionality reduction on the performance of SVM for software defect pre- diction. Comprehensive experiments have been conducted using four kernel functions, ten feature selection thresholds based on the Information gain algorithm, and seven public datasets. This has resulted in 280 experiments. The findings demonstrate that there is no kernel function that can show stable performance across different experimental settings. However, both Radial basis and Sigmoid kernel functions showed more stability than Polynomial and Linear kernel functions. In addition, changing the feature set from high to low dimensionality for each kernel function did not always improve the accuracy of SVM, which contradicts the findings that say using a small set of features can be as effective as using a large set of features. Mohammad Azzeh, Ali Bou Nassif, Shadi Banitaan |
ICMLA | 2 |
| 2022 | Breast cancer detection using artificial intelligence techniques: A systematic literature review
Ali Bou Nassif, Manar Abu Talib, Qassim Nasir, Yaman Afadar, Omar Elgendy |
Artif. Intell. Medicine | 1 |
| 2022 | APT beaconing detection: A systematic review
Manar Abu Talib, Qassim Nasir, Ali Bou Nassif, Takua Mokhamed, Nafisa Ahmed, Bayan Mahfood |
Comput. Secur. | 3 |
| 2022 | Emotional speaker identification using a novel capsule nets model
Ali Bou Nassif, Ismail Shahin, Ashraf Elnagar, Divya Velayudhan, Adi Alhudhaif, Kemal Polat |
Expert Syst. Appl. | 1 |
| 2022 | Novel dual-channel long short-term memory compressed capsule networks for emotion recognition
Ismail Shahin, Noor Ahmad Al Hindawi, Ali Bou Nassif, Adi Alhudhaif, Kemal Polat |
Expert Syst. Appl. | 3 |
| 2022 | 64-bit quantization: taking payload capacity of speech steganography to the limits
Mohammed Baziyad, Ismail Shahin, Tamer Rabie, Ali Bou Nassif |
Multim. Tools Appl. | 4 |
| 2022 | Arabic fake news detection based on deep contextualized embedding models
Ali Bou Nassif, Ashraf Elnagar, Omar Elgendy, Yaman Afadar |
Neural Comput. Appl. | 1 |
| 2022 | On the value of project productivity for early effort estimation
Mohammad Azzeh, Ali Bou Nassif, Yousef Elsheikh, Lefteris Angelis |
Sci. Comput. Program. | 2 |
| 2022 | Empirical Evaluation of Shallow and Deep Learning Classifiers for Arabic Sentiment AnalysisabstractThis work presents a detailed comparison of the performance of deep learning models such as convolutional neural networks, long short-term memory, gated recurrent units, their hybrids, and a selection of shallow learning classifiers for sentiment analysis of Arabic reviews. Additionally, the comparison includes state-of-the-art models such as the transformer architecture and the araBERT pre-trained model. The datasets used in this study are multi-dialect Arabic hotel and book review datasets, which are some of the largest publicly available datasets for Arabic reviews. Results showed deep learning outperforming shallow learning for binary and multi-label classification, in contrast with the results of similar work reported in the literature. This discrepancy in outcome was caused by dataset size as we found it to be proportional to the performance of deep learning models. The performance of deep and shallow learning techniques was analyzed in terms of accuracy and F1 score. The best performing shallow learning technique was Random Forest followed by Decision Tree, and AdaBoost. The deep learning models performed similarly using a default embedding layer, while the transformer model performed best when augmented with araBERT. Ali Bou Nassif, Abdollah Masoud Darya, Ashraf Elnagar |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2021 | The exploitation of Multiple Feature Extraction Techniques for Speaker Identification in Emotional States under Disguised VoicesabstractDue to improvements in artificial intelligence, speaker identification (SI) technologies have brought a great direction and are now widely used in a variety of sectors. One of the most important components of SI is feature extraction, which has a substantial impact on the SI process and performance. As a result, numerous feature extraction strategies are thoroughly investigated, contrasted, and analyzed. This article exploits five distinct feature extraction methods for speaker identification in disguised voices under emotional environments. To evaluate this work significantly, three effects are used: high-pitched, low-pitched, and Electronic Voice Conversion (EVC). Experimental results reported that the concatenated Mel-Frequency Cepstral Coefficients (MFCCs), MFCCs-delta, and MFCCs-delta-delta is the best feature extraction method. Noor Ahmad Al Hindawi, Ismail Shahin, Ali Bou Nassif |
DeSE | 3 |
| 2021 | COVID-19 Electrocardiograms Classification using CNN ModelsabstractWith the periodic rise and fall of COVID-19 and numerous countries being affected by its ramifications, there has been a tremendous amount of work that has been done by scientists, researchers, and doctors all over the world. Prompt intervention is keenly needed to tackle the unconscionable dissemination of the disease. The implementation of Artificial Intelligence (AI) has made a significant contribution to the digital health district by applying the fundamentals of deep learning algorithms. In this study, a novel approach is proposed to automatically diagnose the COVID-19 by the utilization of Electrocardiogram (ECG) data with the integration of deep learning algorithms, specifically the Convolutional Neural Network (CNN) models. Several CNN models have been utilized in this proposed framework, including VGG16, VGG19, InceptionResnetv2, Inception V3, Resnet50, and Densenet201. The VGG16 model has outperformed the rest of the models, with an accuracy of 85.92%. Our results show a relatively low accuracy in the rest of the models compared to the VGG16 model, which is due to the small size of the utilized dataset, in addition to the exclusive utilization of the Grid search hyperparameters optimization approach for the VGG16 model only. Moreover, our results are preparatory, and there is a possibility to enhance the accuracy of all models by further expanding the dataset and adapting a suitable hyperparameters optimization technique. Ismail Shahin, Ali Bou Nassif, Mohamed Bader Alsabek |
DeSE | 2 |
| 2021 | Maximizing embedding capacity for speech steganography: a segment-growing approach
Mohammed Baziyad, Ismail Shahin, Tamer Rabie, Ali Bou Nassif |
Multim. Tools Appl. | 4 |
| 2021 | Novel hybrid DNN approaches for speaker verification in emotional and stressful talking environments
Ismail Shahin, Ali Bou Nassif, Nawel Nemmour, Ashraf Elnagar, Adi Alhudhaif, Kemal Polat |
Neural Comput. Appl. | 2 |
| 2021 | Predicting software effort from use case points: A systematic review
Mohammad Azzeh, Ali Bou Nassif, Imtinan Basem Attili |
Sci. Comput. Program. | 2 |
| 2021 | Empirical analysis on productivity prediction and locality for use case points method
Mohammad Azzeh, Ali Bou Nassif, Cuauhtémoc López Martín |
Softw. Qual. J. | 2 |
| 2021 | Multi-Stage Optimized Machine Learning Framework for Network Intrusion DetectionabstractCyber-security garnered significant attention due to the increased dependency of individuals and organizations on the Internet and their concern about the security and privacy of their online activities. Several previous machine learning (ML)-based network intrusion detection systems (NIDSs) have been developed to protect against malicious online behavior. This paper proposes a novel multi-stage optimized ML-based NIDS framework that reduces computational complexity while maintaining its detection performance. This work studies the impact of oversampling techniques on the models’ training sample size and determines the minimal suitable training sample size. Furthermore, it compares between two feature selection techniques, information gain and correlation-based, and explores their effect on detection performance and time complexity. Moreover, different ML hyper-parameter optimization techniques are investigated to enhance the NIDS’s performance. The performance of the proposed framework is evaluated using two recent intrusion detection datasets, the CICIDS 2017 and the UNSW-NB 2015 datasets. Experimental results show that the proposed model significantly reduces the required training sample size (up to 74%) and feature set size (up to 50%). Moreover, the model performance is enhanced with hyper-parameter optimization with detection accuracies over 99% for both datasets, outperforming recent literature works by 1-2% higher accuracy and 1-2% lower false alarm rate. MohammadNoor Injadat, Abdallah Moubayed, Ali Bou Nassif, Abdallah Shami |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2020 | Multi-split optimized bagging ensemble model selection for multi-class educational data mining
MohammadNoor Injadat, Abdallah Moubayed, Ali Bou Nassif, Abdallah Shami |
Appl. Intell. | 3 |
| 2020 | Arabic audio clips: Identification and discrimination of authentic Cantillations from imitations
Mohammed Lataifeh, Ashraf Elnagar, Ismail Shahin, Ali Bou Nassif |
Neurocomputing | 4 |
| 2020 | Transformed k-nearest neighborhood output distance minimization for predicting the defect density of software projects
Cuauhtémoc López Martín, Yenny Villuendas-Rey, Mohammad Azzeh, Ali Bou Nassif, Shadi Banitaan |
J. Syst. Softw. | 4 |
| 2020 | Systematic ensemble model selection approach for educational data mining
MohammadNoor Injadat, Abdallah Moubayed, Ali Bou Nassif, Abdallah Shami |
Knowl. Based Syst. | 3 |
| 2020 | Novel cascaded Gaussian mixture model-deep neural network classifier for speaker identification in emotional talking environments
Ismail Shahin, Ali Bou Nassif, Shibani Hamsa |
Neural Comput. Appl. | 2 |
| 2019 | Dimensionality reduction with IG-PCA and ensemble classifier for network intrusion detection
Fadi Salo, Ali Bou Nassif, Aleksander Essex |
Comput. Networks | 2 |
| 2018 | Bayesian Optimization with Machine Learning Algorithms Towards Anomaly DetectionabstractNetwork attacks have been very prevalent as their rate is growing tremendously. Both organization and individuals are now concerned about their confidentiality, integrity and availability of their critical information which are often impacted by network attacks. To that end, several previous machine learning-based intrusion detection methods have been developed to secure network infrastructure from such attacks. In this paper, an effective anomaly detection framework is proposed utilizing Bayesian Optimization technique to tune the parameters of Support Vector Machine with Gaussian Kernel (SVM-RBF), Random Forest (RF), and k-Nearest Neighbor (k-NN) algorithms. The performance of the considered algorithms is evaluated using the ISCX 2012 dataset. Experimental results show the effectiveness of the proposed framework in term of accuracy rate, precision, low-false alarm rate, and recall. MohammadNoor Injadat, Fadi Salo, Ali Bou Nassif, Aleksander Essex, Abdallah Shami |
GLOBECOM | 3 |
| 2018 | Ensemble of Learning Project Productivity in Software Effort Based on Use Case PointsabstractIt is well recognized that the project productivity is a key driver in estimating software project effort from Use Case Point size metric at early software development stages. Although, there are few proposed models for predicting productivity, there is no consistent conclusion regarding which model is the superior. Therefore, instead of building a new productivity prediction model, this paper presents a new ensemble construction mechanism applied for software project productivity prediction. Ensemble is an effective technique when performance of base models is poor. We proposed a weighted mean method to aggregate predicted productivities based on average of errors produced by training model. The obtained results show that the using ensemble is a good alternative approach when accuracies of base models are not consistently accurate over different datasets, and when models behave diversely. Mohammad Azzeh, Ali Bou Nassif, Shadi Banitaan, Cuauhtémoc López Martín |
ICMLA | 2 |
| 2018 | Upsilon-SVR Polynomial Kernel for Predicting the Defect Density in New Software ProjectsabstractAn important product measure to determine the effectiveness of software processes is the defect density (DD). In this study, we propose the application of support vector regression (SVR) to predict the DD of new software projects obtained from the International Software Benchmarking Standards Group (ISBSG) Release 2018 data set. Two types of SVR (i.e., ε-SVR and υ-SVR) were applied to train and test these projects. Each SVR used four types of kernels. The prediction accuracy of each SVR was compared to that of a statistical regression (i.e., a simple linear regression, SLR). Statistical significance test showed that υ-SVR with polynomial kernel was better than that of SLR when new software projects were developed on mainframes and coded in programming languages of third generation Cuauhtémoc López Martín, Mohammad Azzeh, Ali Bou Nassif, Shadi Banitaan |
ICMLA | 3 |
| 2018 | Performance Analysis of Hyperledger Fabric PlatformsabstractBlockchain is a key technology that has the potential to decentralize the way we store, share, and manage information and data. One of the more recent blockchain platforms that has emerged is Hyperledger Fabric, an open source, permissioned blockchain that was introduced by IBM, first as Hyperledger Fabric v0.6, and then more recently, in 2017, IBM released Hyperledger Fabric v1.0. Although there are many blockchain platforms, there is no clear methodology for evaluating and assessing the different blockchain platforms in terms of their various aspects, such as performance, security, and scalability. In addition, the new version of Hyperledger Fabric was never evaluated against any other blockchain platform. In this paper, we will first conduct a performance analysis of the two versions of Hyperledger Fabric, v0.6 and v1.0. The performance evaluation of the two platforms will be assessed in terms of execution time, latency, and throughput, by varying the workload in each platform up to 10,000 transactions. Second, we will analyze the scalability of the two platforms by varying the number of nodes up to 20 nodes in each platform. Overall, the performance analysis results across all evaluation metrics, scalability, throughput, execution time, and latency, demonstrate that Hyperledger Fabric v1.0 consistently outperforms Hyperledger Fabric v0.6. However, Hyperledger Fabric v1.0 platform performance did not reach the performance level in current traditional database systems under high workload scenarios. Qassim Nasir, Ilham A. Qasse, Manar Abu Talib, Ali Bou Nassif |
Secur. Commun. Networks | 4 |
| 2018 | Project productivity evaluation in early software effort estimationabstractAbstract The productivity factor has long been a key driver to estimate effort from Use Case Points (UCP) size measure, especially when historical dataset is absent. But, no one questions: Does productivity still matter when historical data are also available? To facilitate answering this question, the present paper studies the role of productivity from 2 perspectives. First, does learning productivity from historical data lead to better accuracy than using fixed productivity ratios? Second, what is the impact of ignoring productivity when estimating the effort from UCP? Five different models that use productivity factor have been used under different experimental settings and compared with some regression models that use only UCP size metrics. We found that dynamically learning and adjusting productivity from historical data are more efficient than using fixed productivity values. Moreover, using UCP size variables to estimate effort tends to be more accurate than using productivity and UCP variables. We also did not find any significant improvement when using UCP adjustment factors for measuring productivity. Finally, we conclude that the productivity factor is a good driver to generate effort estimate from UCP in the presence and absence of historical datasets. But using UCP size variables alone for predicting effort is more accurate than using productivity. Mohammad Azzeh, Ali Bou Nassif |
J. Softw. Evol. Process. | 2 |
| 2018 | On the value of parameter tuning in heterogeneous ensembles effort estimation
Mohamed Hosni, Ali Idri, Alain Abran, Ali Bou Nassif |
Soft Comput. | 4 |
| 2017 | A training process for improving the quality of software projects developed by a practitioner
Cuauhtémoc López Martín, Ali Bou Nassif, Alain Abran |
J. Syst. Softw. | 2 |
| 2017 | Analyzing the relationship between project productivity and environment factors in the use case points methodabstractProject productivity is a key factor for producing effort estimates from use case points (UCP), especially when the historical dataset is absent. The first versions of UCP effort estimation models used a fixed number or very limited numbers of productivity ratios for all new projects. These approaches have not been well examined over a large number of projects, so the validity of these studies was a matter for criticism. The newly available large software datasets allow us to perform further research on the usefulness of productivity for effort estimation of software development. Specifically, we studied the relationship between project productivity and UCP environmental factors, as they have a significant impact on the amount of productivity needed for a software project. Therefore, we designed 4 studies, using various classification and regression methods, to examine the usefulness of that relationship and its impact on UCP effort estimation. The results we obtained are encouraging and show potential improvement in effort estimation. Furthermore, the efficiency of that relationship is better over a dataset that comes from industry because of the quality of data collection. Our comment on the findings is that it is better to exclude environmental factors from calculating UCP and make them available only for computing productivity. The study also encourages project managers to understand how to better assess the environmental factors, as they do have a significant impact on productivity. Mohammad Azzeh, Ali Bou Nassif |
J. Softw. Evol. Process. | 2 |
| 2016 | User Movement Prediction: The Contribution of Machine Learning TechniquesabstractAmbient Assisted Living (AAL) aims to increase the time older people or disabled people can live in their home environment by assisting them in performing activities of daily living by the use of intelligent products. Localization and tracking of users in indoor environment are the main components of AAL. Wireless sensor networks is an effective technology to accomplish these services by using Received Signal Strength (RSS) information. This work seeks to investigate the effect of machine learning techniques on the accuracy of user movement prediction. Five base classifiers and two ensemble learning approaches are employed and the results are evaluated in terms of precision recall, and F-measure. A real-life benchmark dataset in the area of AAL is used for evaluation. The results show that J48 is the best performing model compared to the other base-level classifiers. It also shows that Bagged J48 achieves the best performance. Shadi Banitaan, Mohammad Azzeh, Ali Bou Nassif |
ICMLA | 3 |
| 2016 | Data mining techniques in social media: A survey
MohammadNoor Injadat, Fadi Salo, Ali Bou Nassif |
Neurocomputing | 3 |
| 2016 | Pareto efficient multi-objective optimization for local tuning of analogy-based estimation
Mohammad Azzeh, Ali Bou Nassif, Shadi Banitaan, Fadi Almasalha |
Neural Comput. Appl. | 2 |
| 2016 | Guest editorial: special issue on predictive analytics using machine learning
Ali Bou Nassif, Mohammad Azzeh, Shadi Banitaan, Daniel Neagu |
Neural Comput. Appl. | 1 |
| 2016 | Neural network models for software development effort estimation: a comparative study
Ali Bou Nassif, Mohammad Azzeh, Luiz Fernando Capretz, Danny Ho |
Neural Comput. Appl. | 1 |
| 2015 | An Application of Classification and Class Decomposition to Use Case Point Estimation MethodabstractUse Case Points (UCP) estimation method describes the process of computing the software project size and productivity from use case diagram elements. These metrics are then used to predict the project effort at early stage of software development. The main challenges with previous models are that they were constructed based on a very limited number of observations, and using limited productivity ratios. This paper presents a new approach to predict productivity from UCP environmental factors by applying classification with decomposition technique. A class decomposition provides a number of advantages to supervised learning algorithms through segmenting classes into more homogenous classes, and therefore, increase their diversity. The proposed model is constructed and validated over two datasets that have relatively sufficient number of observations. The accuracy results are promising and have potential to increase accuracy of early effort estimation. Mohammad Azzeh, Ali Bou Nassif, Shadi Banitaan |
ICMLA | 2 |
| 2015 | Class Decomposition Using K-Means and Hierarchical ClusteringabstractThis paper presents a clustering-based class decomposition approach to improve the performance of classifiers. Class decomposition works by dividing each class into clusters, and by relabeling instances contained by each cluster with a new class. Several case studies used class decomposition combined with linear classifiers. While there is an essential improvement in classification accuracy because of class decomposition, the most effective clustering algorithm is not obvious. The aim of this work is to investigate the effect of two clustering algorithms, K-means and hierarchical, on class decomposition. In this work, we study class decomposition when combined with the Naive Bayes classifier using four real-world datasets. Experimental results show an improvement in classification accuracy for most of the datasets when class decomposition using both K-means and hierarchical clustering is performed. The results also show that class decomposition is not suitable for all datasets. Shadi Banitaan, Ali Bou Nassif, Mohammad Azzeh |
ICMLA | 2 |
| 2015 | An empirical evaluation of ensemble adjustment methods for analogy-based effort estimation
Mohammad Azzeh, Ali Bou Nassif, Leandro L. Minku |
J. Syst. Softw. | 2 |
| 2014 | A Better Case Adaptation Method for Case-Based Effort Estimation Using Multi-objective OptimizationabstractCase-Based Reasoning (CBR) is considered as one of the efficient methods in the area of software effort estimation because of its outstanding performance and capability of handling noisy datasets. This study examines the performance of multi-objective Particle Swarm Optimization algorithm to find the best configuration parameters for the adaptation process. Particularly, we propose a new adaptation method for which its parameters can be optimized by making trade off between multiple accuracy measures. The proposed adaptation is fully automated and able to dynamically adapt each case in the dataset individually. Based on empirical validation over 8 datasets, the performance figures have seen good improvements against conventional CBR and some adapted versions of CBR. Ali Bou Nassif, Shadi Banitaan |
ICMLA | 1 |
| 2013 | Fuzzy Model Tree for Early Effort EstimationabstractUse Case Points (UCP) is a well-known method to estimate the project size, based on Use Case diagram, at early phases of software development. Although the Use Case diagram is widely accepted as a de-facto model for analyzing object oriented software requirements over the world, UCP method did not take sufficient amount of attention because, as yet, there is no consensus on how to produce software effort from UCP. This paper aims to study the potential of using Fuzzy Model Tree to derive effort estimates based on UCP size measure using a dataset collected for that purpose. The proposed approach has been validated against Tree boost model, Multiple Linear Regression and classical effort estimation based on the UCP model. The obtained results are promising and show better performance than those obtained by classical UCP, Multiple Linear Regression and slightly better than those obtained by Tree boost model. Mohammad Azzeh, Ali Bou Nassif |
ICMLA (2) | 2 |
| 2013 | Reliability Prediction of Smartphone Applications through Failure Data AnalysisabstractSmart phones have become the most used electronic devices. They carry out most of the functionalities of desktops, offering various useful applications that suit the user's needs. Therefore, instead of being an operator, the user has become the main controller of the device and its applications. Therefore, its reliability becomes an emergent requirement. In this paper, we shortly highlight the shortcomings of Software Reliability Growth Models (SRGMs) when applied to Smartphone applications and suggest future directions for reliability prediction in the mobile area. Sonia Meskini, Ali Bou Nassif, Luiz Fernando Capretz |
PRDC | 2 |
| 2013 | Towards an early software estimation using log-linear regression and a multilayer perceptron model
Ali Bou Nassif, Danny Ho, Luiz Fernando Capretz |
J. Syst. Softw. | 1 |
| 2012 | Fuzzy-ExCOM Software Project Risk AssessmentabstractA software development project is considered to be risky due to the uncertainty of the information (customer requirements), the complexity of the process, and the intangible nature of the product. Under these conditions, risk management in software development projects is mandatory, but often it is difficult and expensive to implement. Expert COCOMO is an efficient approach to software project risk management, which leverages existing knowledge and expertise from previous effort estimation activities to assess the risks in new software projects. However, the original method has limitation because it cannot effectively deal with imprecise and uncertain inputs in the form of linguistic terms such as: Very Low (VL), Low (L), Nominal (N), High (H), Very High (VH) and Extra High (XH). This paper introduces the fuzzy-ExCOM methodology that combines the advantages of a fuzzy technique with Expert COCOMO methodology for risk assessment in software projects. The validation of this approach with industrial data shows that fuzzy-ExCOM provides better risk assessment results with a higher level of sensitivity with respect to risk identification compared to the original Expert COCOMO methodology. Ekananta Manalif, Luiz Fernando Capretz, Ali Bou Nassif, Danny Ho |
ICMLA (2) | 3 |
| 2012 | Estimating Software Effort Using an ANN Model Based on Use Case PointsabstractIn this paper, we propose a novel Artificial Neural Network (ANN) to predict software effort from use case diagrams based on the Use Case Point (UCP) model. The inputs of this model are software size, productivity and complexity, while the output is the predicted software effort. A multiple linear regression model with three independent variables (same inputs of the ANN) and one dependent variable (effort) is also introduced. Our data repository contains 240 data points in which, 214 are industrial and 26 are educational projects. Both the regression and ANN models were trained using 168 data points and tested using 72 data points. The ANN model was evaluated using the MMER and PRED criteria against the regression model, as well as the UCP model that estimates effort from use cases. Results show that the ANN model is a competitive model with respect to other regression models and can be used as an alternative to predict software effort based on the UCP method. Ali Bou Nassif, Luiz Fernando Capretz, Danny Ho |
ICMLA (2) | 1 |
| 2012 | A Treeboost Model for Software Effort Estimation Based on Use Case PointsabstractSoftware effort prediction is an important task in the software development life cycle. Many models including regression models, machine learning models, algorithmic models, expert judgment and estimation by analogy have been widely used to estimate software effort and cost. In this work, a Tree boost (Stochastic Gradient Boosting) model is put forward to predict software effort based on the Use Case Point method. The inputs of the model include software size in use case points, productivity and complexity. A multiple linear regression model was created and the Tree boost model was evaluated against the multiple linear regression model, as well as the use case point model by using four performance criteria: MMRE, PRED, MdMRE and MSE. Experiments show that the Tree boost model can be used with promising results to estimate software effort. Ali Bou Nassif, Luiz Fernando Capretz, Danny Ho, Mohammad Azzeh |
ICMLA (2) | 1 |
| 2012 | Software Effort Estimation in the Early Stages of the Software Life Cycle Using a Cascade Correlation Neural Network ModelabstractSoftware cost estimation is a crucial element in project management. Failing to use a proper cost estimation method might lead to project failures. According to the Standish Chaos Report, 65% of software projects are delivered over budget or after the delivery deadline. Conducting software cost estimation in the early stages of the software life cycle is important and this would be helpful to project managers to bid on projects. In this paper, we propose a novel model to predict software effort from use case diagrams using a cascade correlation neural network approach. The proposed model was evaluated based on the MMER and PRED criteria using 214 industrial and 26 educational projects against a multiple linear regression model and the Use Case Point model. The results show that the proposed cascade correlation neural network can be used with promising results as an alternative approach to predict software effort. Ali Bou Nassif, Luiz Fernando Capretz, Danny Ho |
SNPD | 1 |
| 2011 | Measuring the Usage of SaaS Applications based on Utilized Features
Ali Bou Nassif, Hanan Lutfiyya |
CLOSER | 1 |
| 2011 | Estimating Software Effort Based on Use Case Point Model Using Sugeno Fuzzy Inference SystemabstractSoftware effort estimation is one of the most important tasks in software engineering. Software developers conduct software estimation in the early stages of the software life cycle to derive the required cost and schedule for a project. In the requirements stage, where most software estimation is conducted, the available information is usually imprecise or incomplete. In this paper, a new regression model is created for software effort estimation based on use case point model. Furthermore, a Sugeno Fuzzy Inference System (FIS) approach is applied on this model to improve the estimation. Results show that an improvement of 11% can be achieved in MMRE after applying the Sugeno fuzzy logic approach. Ali Bou Nassif, Luiz Fernando Capretz, Danny Ho |
ICTAI | 1 |
| 2010 | Moving from SaaS Applications towards SOA ServicesabstractThis paper presents a brief introduction of Software as a Service (SaaS) and Service Oriented Architecture (SOA). Specifically, the paper introduces a five-step model to show how SaaS can be offered as SOA services. Furthermore, a real-life scenario is provided to demonstrate the benefits of using the proposed model. Ali Bou Nassif, Miriam A. M. Capretz |
SERVICES | 1 |