Mohammad Azzeh

dblp:61/376 · DBLP profile ↗
← Back
35ranked-venue papers
22as first author
17since 2021 · last 2026
0000-0002-0323-6452ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 10 first-author · 8 since 2021Software engineering, systems software and programming languages · 16 · 12 first-author · 9 since 2021
YearPublicationVenuePosition
2026 Cross-project software defects prediction using fuzzy embedding and deep learning
abstract
Context: Cross-project defect prediction (CPDP) aims to predict software defects in a target project using data from related source projects, especially when defect data for the target project is limited or unavailable. A key challenge in CPDP is data heterogeneity and distributional differences across projects, which often lead to poor performance and unreliable predictions. Objectives: This study proposes a new CPDP model that improves prediction accuracy and robustness by introducing fuzzy embedding and deep learning to better capture similarities and differences across projects. The method is designed to mitigate mismatches in data distribution that hinder existing transfer learning and transformation-based approaches. Methods: The fuzzy embedding technique is built on fuzzy clustering and fuzzy set theory, which map each data point into a two-dimensional space of membership degrees across clusters. This representation models complex relationships with partial memberships and preserves contextual information that is often lost in conventional transformations. A deep learning model based on convolutional neural networks (CNN) processes the embedding matrices to learn discriminative defect patterns. The framework is evaluated against state-of-the-art CPDP models across multiple datasets, and sensitivity analysis is conducted on the number of clusters used in fuzzy embeddings. Results: Empirical evaluation shows that the proposed model consistently outperforms advanced CPDP approaches that rely on standard transformation or weighting methods. Improvements are observed in AUC and other key metrics across diverse datasets, demonstrating that fuzzy embeddings enhance the ability of deep learning to generalize knowledge across projects with varying characteristics. Conclusion: This work contributes a practical and effective solution for addressing heterogeneity in CPDP. By combining fuzzy embeddings with deep learning, the model not only achieves higher predictive accuracy but also improves reliability in real-world scenarios where project data distributions differ significantly. These findings highlight the potential of fuzzy embedding to support more resilient software quality assurance practices and provide actionable insights for practitioners dealing with limited or imbalanced defect data.
Mohammad Azzeh, Mohammad Abdel-Rahman
Inf. Softw. Technol.1
2026 A Systematic Review on Code Smell Detection Approaches in Open Source Projects
abstract
ABSTRACT Background Code smells are an indicator that something is misplaced in software systems, which can reduce maintainability and quality of software applications. In the literature, there is much research about code smells where numerous strategies for automating code smell detection have been developed to improve software quality. Objective The purpose of this work is to provide a review of search‐based, heuristic‐based, machine learning‐based, deep learning‐based, and hybrid‐based code smell detection algorithms. Method Concerning the main goals of this research, we have found 38 primary studies. We gathered relevant studies published on this topic between 2017 and 2024. These articles' data were extracted according to some criteria, including code smells, machine learning methods, programming languages, dataset size, evaluation strategy, and statistical tests. Results Machine learning‐based code smell detection methods have been suggested in most empirical investigations. We found that machine learning and deep learning are the most popular approaches for predicting code smells. Researchers typically employ support vector machine and decision tree algorithms. The Arcelli Fontana and Zanoni benchmark dataset was the most frequently investigated dataset. Most of the research community's attention has been focused on a small number of smells, including blob, feature envy, long method, and data class. Researchers also pay more attention to code smells like Long Method and Feature Envy. Deep learning techniques are increasingly used, and most scientists utilize source code metrics as indicators. The standard performance measurements mostly used are F‐measure, recall, and precision. Conclusions The study provides an overview of existing approaches and highlights current research trends in code smell detection, particularly the increasing use of machine learning and deep learning techniques.
Sawsan Alodibat, Mohammad Azzeh
Softw. Pract. Exp.2
2025 Correction: Software effort estimation using convolutional neural network and fuzzy clustering
Mohammad Azzeh, Abedalrhman Alkhateeb, Ali Bou Nassif
Neural Comput. Appl.1
2025 Integrative multi-omics approach for identifying and classifying high-risk patients in prostate cancer management
Mohammad Azzeh, Noor Afeshat, Hazem Qattous, Abedrahman Alkhateeb
Neural Comput. Appl.1
2024 Software effort estimation using convolutional neural network and fuzzy clustering
Mohammad Azzeh, Abedalrhman Alkhateeb, Ali Bou Nassif
Neural Comput. Appl.1
2024 Parameter-efficient fine-tuning of pre-trained code models for just-in-time defect prediction
Manar Abu Talib, Ali Bou Nassif, Mohammad Azzeh, Yaser Alesh, Yaman Afadar
Neural Comput. Appl.3
2024 A soft computing approach for software defect density prediction
abstract
Abstract Defect density is an essential software testing and maintenance aspect that determines the quality of software products. It is used as a management factor to distribute limited human resources successfully. The availability of public defect datasets facilitates building defect density prediction models using established static code metrics. Since the data gathered for software modules are often subject to uncertainty, it becomes difficult to deliver accurate and reliable predictions. To alleviate this issue, we propose a new prediction model that integrates gray system theory and fuzzy logic to handle the imprecision in software measurement. We propose a new similarity measure that combines the benefits of fuzzy logic and gray relational analysis. The proposed model was validated against defect density prediction models using public defect datasets. The defect density variable is frequently sparse because of the vast number of none‐defected modules in the datasets. Therefore, we also check our proposed model's performance against the sparsity level. The findings reveal that the developed model surpasses other defect density prediction models over the datasets with high and very high sparsity ratios. The ensemble learning techniques are competitive choices to the proposed model when the sparsity ratio is relatively small. On the other hand, the statistical regression models were the most inadequate methods for such problems and datasets. Finally, the proposed model was evaluated against different degrees of uncertainty using a sensitivity analysis procedure. The results showed that our model behaves stably under different degrees of uncertainty.
Mohammad Azzeh, Yousef Alqasrawi, Yousef Elsheikh
J. Softw. Evol. Process.1
2024 Software defect density prediction using grey system theory and fuzzy logic
Mohammad Azzeh, Yousef Elsheikh, Yousef Alqasrawi
Soft Comput.1
2023 Truth Seeker of the Largest Social Media Content using Machine Learning Algorithms
abstract
Social media is the source of news, information, and opinions for many people. So, the pursuit of truth in the realm of social media is a noble and essential endeavor because it underpins informed decision-making, counters misinformation, nurtures critical thinking and builds trust. We used machine learning(ML) algorithms and Natural Language Processing(NLP) to understand and process human language used in social media posts of the Largest Social Media Ground-Truth Dataset for Real/Fake Content to create a model that could be used as a reference to detect fake news from Twitter posts, which is a binary classification problem. We created two different datasets from the original large dataset using different NLP vectorization techniques (word2vec and TF -IDF), generating different features per dataset. Then we applied seven ML algorithms on each dataset and compared the results by extracting the confusion matrix and ROC per model. Random forest and AdBoost algorithms achieved the best F1 scores.
Maysa Khalil, Mohammad Azzeh
ICMLA2
2023 An optimized case-based software project effort estimation using genetic algorithm
Shaima Hameed, Yousef Elsheikh, Mohammad Azzeh
Inf. Softw. Technol.3
2023 Examining the performance of kernel methods for software defect prediction based on support vector machine
Mohammad Azzeh, Yousef Elsheikh, Ali Bou Nassif, Lefteris Angelis
Sci. Comput. Program.1
2022 Multi-omics Data Integration Model based on Isomap and Convolutional Neural Network
abstract
The recent advances in genetics technologies have led to high-throughput characterized different biological molecules’ functionalities. The availability of heterogeneous omics sparked the challenge of integrating them for further analysis. This work incorporates the Isomap technique to embed multi-omic data into a convolutional neural network (CNN). The deep learning model fuses three omics data, which are gene expression, copy number alteration (CNA), and DNA methylation data, for breast cancer stage prediction. Isomap is utilized to convert the high-dimensional data into 2-dimensional maps. The gene similarity network (GSN) map is created based on gene expression data to preserve gene relationships. The values from three omics for each sample are used to color the GSN map based on the RGB system. The created GSN maps for all samples are fed into the CNN for classification.The model was applied to TCGA breast invasive carcinoma data set to predict the stage of breast cancer. It outperformed the state-of-art iSOM-GSN model in performance metrics, including accuracy, precision, recall, f1-measure, and area under the curve (AUC). The results indicate that a combination of Isomap embedding technique and CNN can successfully integrate a multi-omics data set for cancer outcome prediction, including the diagnosis and prognosis of the complex disease.
Abedalrhman Alkhateeb, Bashier ElKarami, Hazem Qattous, Abdullah Al-Refai, Noor AlAfeshat, Behnam Shahrrava, Mohammad Azzeh
ICMLA7
2022 On Selection of Optimal Kernel Function for Software Defect Prediction
abstract
Software defect prediction is used to identify the most likely defect-prone modules. This activity can help the quality assurance team to distribute limited resources efficiently as well as reduce the time and cost of the testing phase. Support Vector Machine (SVM) has been widely used to build software defect prediction models. Nevertheless, the accuracy of such method depends on multiple factors such as the type of SVM method, choice of kernel functions and dimensionality of the dataset. This paper continues along that research line and examines the impact and stability of four kernel functions and feature dimensionality reduction on the performance of SVM for software defect pre- diction. Comprehensive experiments have been conducted using four kernel functions, ten feature selection thresholds based on the Information gain algorithm, and seven public datasets. This has resulted in 280 experiments. The findings demonstrate that there is no kernel function that can show stable performance across different experimental settings. However, both Radial basis and Sigmoid kernel functions showed more stability than Polynomial and Linear kernel functions. In addition, changing the feature set from high to low dimensionality for each kernel function did not always improve the accuracy of SVM, which contradicts the findings that say using a small set of features can be as effective as using a large set of features.
Mohammad Azzeh, Ali Bou Nassif, Shadi Banitaan
ICMLA1
2022 Locally weighted regression with different kernel smoothers for software effort estimation
Yousef Alqasrawi, Mohammad Azzeh, Yousef Elsheikh
Sci. Comput. Program.2
2022 On the value of project productivity for early effort estimation
Mohammad Azzeh, Ali Bou Nassif, Yousef Elsheikh, Lefteris Angelis
Sci. Comput. Program.1
2021 Predicting software effort from use case points: A systematic review
Mohammad Azzeh, Ali Bou Nassif, Imtinan Basem Attili
Sci. Comput. Program.1
2021 Empirical analysis on productivity prediction and locality for use case points method
Mohammad Azzeh, Ali Bou Nassif, Cuauhtémoc López Martín
Softw. Qual. J.1
2020 Transformed k-nearest neighborhood output distance minimization for predicting the defect density of software projects
Cuauhtémoc López Martín, Yenny Villuendas-Rey, Mohammad Azzeh, Ali Bou Nassif, Shadi Banitaan
J. Syst. Softw.3
2018 Ensemble of Learning Project Productivity in Software Effort Based on Use Case Points
abstract
It is well recognized that the project productivity is a key driver in estimating software project effort from Use Case Point size metric at early software development stages. Although, there are few proposed models for predicting productivity, there is no consistent conclusion regarding which model is the superior. Therefore, instead of building a new productivity prediction model, this paper presents a new ensemble construction mechanism applied for software project productivity prediction. Ensemble is an effective technique when performance of base models is poor. We proposed a weighted mean method to aggregate predicted productivities based on average of errors produced by training model. The obtained results show that the using ensemble is a good alternative approach when accuracies of base models are not consistently accurate over different datasets, and when models behave diversely.
Mohammad Azzeh, Ali Bou Nassif, Shadi Banitaan, Cuauhtémoc López Martín
ICMLA1
2018 Upsilon-SVR Polynomial Kernel for Predicting the Defect Density in New Software Projects
abstract
An important product measure to determine the effectiveness of software processes is the defect density (DD). In this study, we propose the application of support vector regression (SVR) to predict the DD of new software projects obtained from the International Software Benchmarking Standards Group (ISBSG) Release 2018 data set. Two types of SVR (i.e., ε-SVR and υ-SVR) were applied to train and test these projects. Each SVR used four types of kernels. The prediction accuracy of each SVR was compared to that of a statistical regression (i.e., a simple linear regression, SLR). Statistical significance test showed that υ-SVR with polynomial kernel was better than that of SLR when new software projects were developed on mainframes and coded in programming languages of third generation
Cuauhtémoc López Martín, Mohammad Azzeh, Ali Bou Nassif, Shadi Banitaan
ICMLA2
2018 Project productivity evaluation in early software effort estimation
abstract
Abstract The productivity factor has long been a key driver to estimate effort from Use Case Points (UCP) size measure, especially when historical dataset is absent. But, no one questions: Does productivity still matter when historical data are also available? To facilitate answering this question, the present paper studies the role of productivity from 2 perspectives. First, does learning productivity from historical data lead to better accuracy than using fixed productivity ratios? Second, what is the impact of ignoring productivity when estimating the effort from UCP? Five different models that use productivity factor have been used under different experimental settings and compared with some regression models that use only UCP size metrics. We found that dynamically learning and adjusting productivity from historical data are more efficient than using fixed productivity values. Moreover, using UCP size variables to estimate effort tends to be more accurate than using productivity and UCP variables. We also did not find any significant improvement when using UCP adjustment factors for measuring productivity. Finally, we conclude that the productivity factor is a good driver to generate effort estimate from UCP in the presence and absence of historical datasets. But using UCP size variables alone for predicting effort is more accurate than using productivity.
Mohammad Azzeh, Ali Bou Nassif
J. Softw. Evol. Process.1
2017 Analyzing the relationship between project productivity and environment factors in the use case points method
abstract
Project productivity is a key factor for producing effort estimates from use case points (UCP), especially when the historical dataset is absent. The first versions of UCP effort estimation models used a fixed number or very limited numbers of productivity ratios for all new projects. These approaches have not been well examined over a large number of projects, so the validity of these studies was a matter for criticism. The newly available large software datasets allow us to perform further research on the usefulness of productivity for effort estimation of software development. Specifically, we studied the relationship between project productivity and UCP environmental factors, as they have a significant impact on the amount of productivity needed for a software project. Therefore, we designed 4 studies, using various classification and regression methods, to examine the usefulness of that relationship and its impact on UCP effort estimation. The results we obtained are encouraging and show potential improvement in effort estimation. Furthermore, the efficiency of that relationship is better over a dataset that comes from industry because of the quality of data collection. Our comment on the findings is that it is better to exclude environmental factors from calculating UCP and make them available only for computing productivity. The study also encourages project managers to understand how to better assess the environmental factors, as they do have a significant impact on productivity.
Mohammad Azzeh, Ali Bou Nassif
J. Softw. Evol. Process.1
2016 User Movement Prediction: The Contribution of Machine Learning Techniques
abstract
Ambient Assisted Living (AAL) aims to increase the time older people or disabled people can live in their home environment by assisting them in performing activities of daily living by the use of intelligent products. Localization and tracking of users in indoor environment are the main components of AAL. Wireless sensor networks is an effective technology to accomplish these services by using Received Signal Strength (RSS) information. This work seeks to investigate the effect of machine learning techniques on the accuracy of user movement prediction. Five base classifiers and two ensemble learning approaches are employed and the results are evaluated in terms of precision recall, and F-measure. A real-life benchmark dataset in the area of AAL is used for evaluation. The results show that J48 is the best performing model compared to the other base-level classifiers. It also shows that Bagged J48 achieves the best performance.
Shadi Banitaan, Mohammad Azzeh, Ali Bou Nassif
ICMLA2
2016 Pareto efficient multi-objective optimization for local tuning of analogy-based estimation
Mohammad Azzeh, Ali Bou Nassif, Shadi Banitaan, Fadi Almasalha
Neural Comput. Appl.1
2016 Guest editorial: special issue on predictive analytics using machine learning
Ali Bou Nassif, Mohammad Azzeh, Shadi Banitaan, Daniel Neagu
Neural Comput. Appl.2
2016 Neural network models for software development effort estimation: a comparative study
Ali Bou Nassif, Mohammad Azzeh, Luiz Fernando Capretz, Danny Ho
Neural Comput. Appl.2
2015 An Application of Classification and Class Decomposition to Use Case Point Estimation Method
abstract
Use Case Points (UCP) estimation method describes the process of computing the software project size and productivity from use case diagram elements. These metrics are then used to predict the project effort at early stage of software development. The main challenges with previous models are that they were constructed based on a very limited number of observations, and using limited productivity ratios. This paper presents a new approach to predict productivity from UCP environmental factors by applying classification with decomposition technique. A class decomposition provides a number of advantages to supervised learning algorithms through segmenting classes into more homogenous classes, and therefore, increase their diversity. The proposed model is constructed and validated over two datasets that have relatively sufficient number of observations. The accuracy results are promising and have potential to increase accuracy of early effort estimation.
Mohammad Azzeh, Ali Bou Nassif, Shadi Banitaan
ICMLA1
2015 Class Decomposition Using K-Means and Hierarchical Clustering
abstract
This paper presents a clustering-based class decomposition approach to improve the performance of classifiers. Class decomposition works by dividing each class into clusters, and by relabeling instances contained by each cluster with a new class. Several case studies used class decomposition combined with linear classifiers. While there is an essential improvement in classification accuracy because of class decomposition, the most effective clustering algorithm is not obvious. The aim of this work is to investigate the effect of two clustering algorithms, K-means and hierarchical, on class decomposition. In this work, we study class decomposition when combined with the Naive Bayes classifier using four real-world datasets. Experimental results show an improvement in classification accuracy for most of the datasets when class decomposition using both K-means and hierarchical clustering is performed. The results also show that class decomposition is not suitable for all datasets.
Shadi Banitaan, Ali Bou Nassif, Mohammad Azzeh
ICMLA3
2015 An empirical evaluation of ensemble adjustment methods for analogy-based effort estimation
Mohammad Azzeh, Ali Bou Nassif, Leandro L. Minku
J. Syst. Softw.1
2013 Fuzzy Model Tree for Early Effort Estimation
abstract
Use Case Points (UCP) is a well-known method to estimate the project size, based on Use Case diagram, at early phases of software development. Although the Use Case diagram is widely accepted as a de-facto model for analyzing object oriented software requirements over the world, UCP method did not take sufficient amount of attention because, as yet, there is no consensus on how to produce software effort from UCP. This paper aims to study the potential of using Fuzzy Model Tree to derive effort estimates based on UCP size measure using a dataset collected for that purpose. The proposed approach has been validated against Tree boost model, Multiple Linear Regression and classical effort estimation based on the UCP model. The obtained results are promising and show better performance than those obtained by classical UCP, Multiple Linear Regression and slightly better than those obtained by Tree boost model.
Mohammad Azzeh, Ali Bou Nassif
ICMLA (2)1
2012 A Treeboost Model for Software Effort Estimation Based on Use Case Points
abstract
Software effort prediction is an important task in the software development life cycle. Many models including regression models, machine learning models, algorithmic models, expert judgment and estimation by analogy have been widely used to estimate software effort and cost. In this work, a Tree boost (Stochastic Gradient Boosting) model is put forward to predict software effort based on the Use Case Point method. The inputs of the model include software size in use case points, productivity and complexity. A multiple linear regression model was created and the Tree boost model was evaluated against the multiple linear regression model, as well as the use case point model by using four performance criteria: MMRE, PRED, MdMRE and MSE. Experiments show that the Tree boost model can be used with promising results to estimate software effort.
Ali Bou Nassif, Luiz Fernando Capretz, Danny Ho, Mohammad Azzeh
ICMLA (2)4
2012 A replicated assessment and comparison of adaptation techniques for analogy-based effort estimation
Mohammad Azzeh
Empir. Softw. Eng.1
2011 Adjusted Case-Based Software Effort Estimation Using Bees Optimization Algorithm
Mohammad Azzeh
KES (2)1
2011 Analogy-based software effort estimation using Fuzzy numbers
Mohammad Azzeh, Daniel Neagu, Peter I. Cowling
J. Syst. Softw.1
2010 Fuzzy grey relational analysis for software effort estimation
Mohammad Azzeh, Daniel Neagu, Peter I. Cowling
Empir. Softw. Eng.1